01 — Essays7 pieces · agents, harnesses, delivery

Notes from inside the build.

Agent architecture, harness design and the unglamorous operational work that decides whether an LLM system survives contact with production.

LatestSeptember 15, 202612 min read

The best-case Grok Bot is an org, not a chatbot

A one-person shop does not need a chatbot. It needs a Chief of Staff, four leads, and a hard stop before money moves. This is the Grok Bot org I would actually run for books, a trading desk, and the rest of the company.

Read the essay
AgentsOpsFinance

Archive — everything else

Filter
Sep 1, 202610 min read

Isolated eval agents still coordinated. The leak was the harness.

When thousands of “isolated” eval agents found a shared channel, they ran collective workstreams the scorer never designed for. What we change in the harness, and what we assume next — not a recap, not a how-to.

Read the essay
AgentsEvalsHarness
Aug 24, 202616 min read

One ticket in, one reviewed pull request out

Four flows, twenty-six role agents and exactly one human stop per run. The delivery machine my work goes through end to end — how a ticket gets classified, isolated in a sandbox, specified, built, reviewed by a panel that re-runs in full after every repair, and released only once a person says yes.

Read the essay
AgentsDeliveryHarness
Aug 24, 20267 min read

Testing an agent's control flow without a model

Most agent test suites pay a model to answer a question the model has no part in: does the next node run, is that gate reachable, does the retry budget actually decrement. Separate transport from judgement and those checks run in seconds, offline, for nothing.

Read the essay
AgentsTestingHarness
Aug 24, 20267 min read

What a bounded loop actually costs

A timeout is not a budget. Budgets that work are enforced on every state write, discount halted time, and hand the decision at the ceiling to something that can read the run — because 'one assertion away' and 'going in circles' both present as a spent budget.

Read the essay
AgentsLoopsCost
Aug 19, 202611 min read

AI code review, and the agent I built to gate every merge

Generating code got cheap; reading it did not. The bottleneck moved from writing to reviewing — so here is the two-tier review agent I run on my own repos: a fast local gate before the commit lands, a deeper agentic pass on the pull request, and the evals that keep both honest.

Read the essay
Code reviewAgentsCI/CD
Feb 18, 20269 min read

How to build AI agents using LangChain

AI agents are no longer science fiction — they're writing code, querying databases and making decisions in production right now. A build-along from an empty file to a working loop, with LangChain and LangGraph.

Read the essay
AgentsLangChainPython