A one-person shop does not need a chatbot. It needs a Chief of Staff, four leads, and a hard stop before money moves. This is the Grok Bot org I would actually run for books, a trading desk, and the rest of the company.
When thousands of “isolated” eval agents found a shared channel, they ran collective workstreams the scorer never designed for. What we change in the harness, and what we assume next — not a recap, not a how-to.
Four flows, twenty-six role agents and exactly one human stop per run. The delivery machine my work goes through end to end — how a ticket gets classified, isolated in a sandbox, specified, built, reviewed by a panel that re-runs in full after every repair, and released only once a person says yes.
Most agent test suites pay a model to answer a question the model has no part in: does the next node run, is that gate reachable, does the retry budget actually decrement. Separate transport from judgement and those checks run in seconds, offline, for nothing.
A timeout is not a budget. Budgets that work are enforced on every state write, discount halted time, and hand the decision at the ceiling to something that can read the run — because 'one assertion away' and 'going in circles' both present as a spent budget.
Generating code got cheap; reading it did not. The bottleneck moved from writing to reviewing — so here is the two-tier review agent I run on my own repos: a fast local gate before the commit lands, a deeper agentic pass on the pull request, and the evals that keep both honest.
AI agents are no longer science fiction — they're writing code, querying databases and making decisions in production right now. A build-along from an empty file to a working loop, with LangChain and LangGraph.