What gets built here
Nine AI builds in full — multi-agent orchestration, agent loop engineering, context engineering with GraphRAG and memory, MCP tool surfaces, evals and guardrails, in-product copilots, voice agents, document intelligence and AI strategy — each with the architecture it actually ships as, and what it is not the right fit for.
- Builds
- 9
- Typical build
- 4–12 wks
- Team
- 1 operator
Index — nine builds, one operator
- S/01Agentic workflows & orchestrationMulti-agent systems with planner/executor topologies, subagent fan-out, and durable execution underneath. Handoffs are explicit, checkpoints are human, and a crashed process resumes rather than restarts.
- S/02Agent loop engineeringThe harness under a single agent — context compaction, tool-call budgets, stop conditions, and a self-critique pass. This is where agents that spiral get fixed.
- S/03Context engineering & memoryHybrid retrieval, GraphRAG over a real knowledge graph, and agent memory that persists between runs. Ranked, reranked, cited, and compacted before it reaches the window.
- S/04MCP servers & tool surfacesModel Context Protocol servers that expose your internal systems to any agent — typed tool schemas, scoped OAuth, and a bounded surface rather than a database handed to a model.
- S/05Evals, tracing & guardrailsGolden sets, LLM-as-judge scoring, and regression gates in CI — plus prompt-injection defence and red-teaming. The difference between shipping on evidence and shipping on vibes.
- S/06In-product copilotsAI inside the product itself — token streaming, generative UI, tool calling against your own API, and structured outputs written straight back to your database.
- S/07Voice & realtime agentsSpeech-to-speech agents with turn detection, barge-in, and sub-second response — on the phone or in the product, handing to a human with the transcript attached.
- S/08Document intelligenceVLM parsing over the documents you actually receive — scans, tables, handwriting — into schema-constrained records with a citation on every field and a confidence gate before anything is trusted.
- S/09AI strategy & auditsReadiness audits, build-versus-buy calls, model and data policy, and an eval plan before a quarter is committed. A few hours that often save months.
Agentic workflows & orchestration
Multi-agent systems with planner/executor topologies, subagent fan-out, and durable execution underneath. Handoffs are explicit, checkpoints are human, and a crashed process resumes rather than restarts.
A single agent in a while-loop is a demo. An orchestration is what survives production: a planner that decomposes the request, subagents fanned out over the parts, an explicit handoff contract between them, and a merge step that reconciles disagreement instead of taking whichever answer arrived last. Underneath sits durable execution — Temporal or Inngest — so a deploy mid-run, a rate-limited provider, or a process dying at step nine replays from its last checkpoint rather than from the beginning. Everything is idempotent because at-least-once delivery is the only honest assumption, and anything that exhausts its retries lands in a dead-letter queue a person can read, fix and replay. Human-in-the-loop checkpoints are first-class: the run pauses, waits for an approval that may arrive tomorrow, and continues.
- A prototype works on the happy path and falls over on real traffic.
- One long run holds state nobody can inspect, resume, or replay.
- The work needs an approval partway through and currently blocks on it.
Agent loop engineering
The harness under a single agent — context compaction, tool-call budgets, stop conditions, and a self-critique pass. This is where agents that spiral get fixed.
Most failing agents do not fail at the model. They fail at the loop around it: context that grows until the instructions fall out of attention, a tool-call budget nobody set, a stop condition that never fires, and a retry that feeds the same poisoned context back in. Loop engineering is the unglamorous fix — compaction that summarises completed work instead of re-reading every tool result, a budget on calls and tokens per run with a defined behaviour at the ceiling, stop conditions that distinguish finished from stuck, and a self-critique pass that checks the result against the original request before it is returned. Every decision the loop makes is traced with OpenTelemetry, so when a run costs eleven dollars you can see the step that did it rather than guess.
- An agent loops on itself, burns tokens, and returns nothing useful.
- Cost per run varies by an order of magnitude and nobody can say why.
- Long runs degrade — the model forgets its instructions halfway through.
Context engineering & memory
Hybrid retrieval, GraphRAG over a real knowledge graph, and agent memory that persists between runs. Ranked, reranked, cited, and compacted before it reaches the window.
Retrieval quality is a ranking problem before it is a model problem, and context is a budget before it is a prompt. The work runs in three layers. Retrieval: chunking that respects how a document is actually structured, hybrid keyword-plus-vector search rather than embeddings alone, a reranking pass, and citations rendered back to the source so an answer can be checked instead of trusted. Graph: entity extraction and resolution into a knowledge graph, so a question that spans three documents traverses relationships rather than hoping all three chunks rank in the same top-ten — the GraphRAG case, and the one plain vector search is worst at. Memory: episodic memory of what happened in past runs, semantic memory of what was learned from them, a written policy for what gets promoted and what is allowed to decay, and compaction so a long-lived agent does not carry a year of transcripts into every window.
- The answer exists across three documents and vector search finds one.
- The agent re-learns the same fact on every run and remembers nothing.
- A general model answers confidently and wrongly about your own product.
MCP servers & tool surfaces
Model Context Protocol servers that expose your internal systems to any agent — typed tool schemas, scoped OAuth, and a bounded surface rather than a database handed to a model.
MCP is how an agent reaches a system it does not own, and the protocol is the easy half. The work is the surface: tool schemas typed tightly enough that the model cannot express an invalid call, scopes that mean a compromised or confused agent reaches exactly one customer's records rather than the table they sit in, resources and prompts that give the model the shape of your domain without a paragraph of prose in every system message, and error messages written for a model to recover from rather than for a log. Auth is OAuth 2.1 with per-tool scopes, not a shared key in an environment variable. Servers ship on Vercel or Cloudflare with traces on every call, because a tool surface nobody can audit is a tool surface nobody should have connected.
- Every agent you build re-implements the same integrations from scratch.
- Giving a model access to a system means giving it far more than it needs.
- Claude, Cursor and your own agents each need the same internal tools.
Evals, tracing & guardrails
Golden sets, LLM-as-judge scoring, and regression gates in CI — plus prompt-injection defence and red-teaming. The difference between shipping on evidence and shipping on vibes.
Without evals, every prompt change is a coin flip you cannot score. The build starts with a golden set drawn from your real traffic and labelled with your team, because a suite written from imagination measures imagination. Scoring combines deterministic assertions where the answer has a right shape with LLM-as-judge where it does not, and the judge itself is validated against human labels before it is trusted to grade anything. That suite becomes a regression gate in CI: a change that drops accuracy on the golden set does not merge. Alongside it runs the safety half — prompt-injection defence on every untrusted input, output filtering, a red-team pass against your specific tool surface, and OpenTelemetry traces so a production failure can be replayed as a new eval case instead of argued about in Slack.
- Nobody can say whether last week's prompt change made things better.
- Quality is assessed by someone trying it a few times before a deploy.
- Untrusted text reaches a model that can call tools, and nobody has tested that.
In-product copilots
AI inside the product itself — token streaming, generative UI, tool calling against your own API, and structured outputs written straight back to your database.
A copilot that lives in your product is a different build from one that lives in a chat window. It streams, because a spinner over four seconds of silence reads as broken. It renders generative UI — a real component the user can act on, not a paragraph describing what they could do. It calls your own API through tools scoped to the signed-in user's permissions, so the assistant can never surface a record that user could not open themselves. Outputs are schema-constrained with Zod and written back as structured rows, which is what makes the assistant part of the product rather than a text box beside it. This is also where tier-1 support deflection and content pipelines land when the right home for them is inside the product.
- Users know what they want and cannot find the screen that does it.
- The same onboarding questions arrive as support tickets every week.
- A workflow in the product takes eleven clicks and one sentence to describe.
Voice & realtime agents
Speech-to-speech agents with turn detection, barge-in, and sub-second response — on the phone or in the product, handing to a human with the transcript attached.
Voice is a latency problem wearing an AI costume. Past roughly eight hundred milliseconds a caller starts talking over the agent, so the build is speech-to-speech rather than the transcribe-think-synthesise chain that stacks three round-trips before anyone hears anything. Turn detection decides when the caller has actually finished rather than merely paused, barge-in stops the agent mid-sentence when they interrupt, and both are tuned against your own recordings, because the right thresholds differ between a support line and a booking line. It runs on a real telephony path through Twilio or in-product over LiveKit WebRTC, with a transfer rule written down: on low confidence, on an explicit request, or on any topic you have declared off-limits, the call goes to a person with the transcript and the extracted details already attached.
- Calls go unanswered after hours and the voicemail is never returned.
- The same booking or triage conversation runs forty times a day.
- An IVR tree exists and callers press zero to escape it.
Document intelligence
VLM parsing over the documents you actually receive — scans, tables, handwriting — into schema-constrained records with a citation on every field and a confidence gate before anything is trusted.
The documents that matter are rarely clean text. They are scans, multi-page tables that break across a page, a stamp over the total, and a handwritten amendment in the margin. Vision models read those in a way OCR never did, but reading is not the deliverable — a validated record is. Extraction is schema-constrained, so the output is a typed object rather than prose to be parsed twice, and every field carries a citation to the page and region it came from, which is what turns a review from re-reading the document into checking a highlight. A confidence gate routes the difference: high-confidence records land straight in the system of record, low-confidence ones go to a human queue with the uncertain fields flagged, and every correction becomes an eval case so the gate gets better instead of staying where it was set on day one.
- Someone retypes numbers from PDFs into a system every morning.
- OCR was tried, and the tables and handwriting defeated it.
- An error is only found downstream, weeks after the document was filed.
AI strategy & audits
Readiness audits, build-versus-buy calls, model and data policy, and an eval plan before a quarter is committed. A few hours that often save months.
An audit of what you have, a roadmap ranked by time-saved-per-dollar, or a build-versus-buy review before you commit a quarter to the wrong path. It covers the questions that get expensive later: which of these belongs to a model at all rather than to a process change, what your data actually supports today, model and vendor policy including where an open-weight model on your own infrastructure is the right answer, and how any of it will be measured — an eval plan, before the build, rather than after the first incident. The deliverable is a written document you own outright and can hand to another vendor. A meaningful share end in a recommendation not to build the thing, which is what the opinion is for.
- There is a list of AI ideas and no honest way to rank them.
- A six-figure platform quote needs a second opinion before it is signed.
- Something shipped, and nobody agreed in advance how it would be judged.
Process — how an engagement actually goes
- 01 · DISCOVER
Diagnose, not prescribe.
A working session to map every manual step in the workflow. I leave with a list of candidates ranked by time-saved-per-dollar.
- 02 · DESIGN
Blueprint on paper.
A node-graph of your target system with every handoff, failure mode, and data contract named. You approve before anything runs.
- 03 · BUILD
Ship in slices.
Narrow vertical releases every 3–5 days. You see the system working on your data before it's fully finished.
- 04 · OPERATE
Own the outcome.
Monitoring, alerts, and iteration after launch. A monthly retainer keeps the system sharp as your business shifts.
Engagements — three ways to buy it
Discovery
A working session that maps every manual step in the workflow and ranks the candidates by time saved per dollar. It ends with a recommendation, which is sometimes that there is nothing here worth building.
- A mapped picture of the current process
- Candidates ranked by time saved per dollar
- A written recommendation, yours either way
Fixed-scope sprint
The default. A named scope, a written spec you approve before anything runs, and narrow vertical slices released every three to five days so you see the system working on your own data long before it is finished.
- A written spec approved before the first line of code
- A working slice every 3–5 days
- Handover with the code, prompts, evals and infrastructure
Ongoing retainer
For teams that need the judgement more than the hands: architecture review, build-versus-buy calls, vendor selection, and keeping the systems already shipped sharp as the business moves under them.
- Architecture and code review on your cadence
- Monitoring, alerts and iteration on live systems
- Cancel-any-month, no notice period
Stack — what the builds actually run on
Tools are chosen per engagement, and the list below is what usually wins rather than a partnership page. Anything your team already runs well stays.
Questions — before the first call
Can we start with one build instead of a whole platform?
That is the preferred way in. One build, shipped end to end and running on your data, tells you more about whether the rest is worth doing than any amount of planning does. Most engagements that grow started as a single agent nobody was sure about.
How do you decide between an agent and a fixed pipeline?
By whether the steps change per case. If the sequence is the same every time, a deterministic pipeline is cheaper, faster and far easier to debug — wrapping it in an agent buys nondeterminism nobody asked for. An agent earns its place when the work needs judgement: reading the situation, picking a tool, checking its own result, and escalating when confidence drops.
Why do evals come up in every one of these builds?
Because without them there is no way to tell a change from an improvement. A golden set drawn from your real traffic, a judge validated against human labels, and a regression gate in CI turn every prompt and model change from a coin flip into a measured decision. It is also the only honest way to answer whether a new model is better for your workload rather than better on a public benchmark.
Do you work with the tools we already have, or replace them?
Work with them, in almost every case. Agents reach your systems through an MCP server or typed tools rather than a migration, retrieval indexes the documents where they already live, and structured output writes back to the Postgres, CRM or help desk your team already reports on. Replacing a working tool is a cost with no output attached to it.
What do you need from my team during a build?
Access, one decision-maker, and about an hour a week. Access to the systems being automated and to real examples — real tickets, real documents, real calls — because a golden set built from imagination measures imagination, and an agent tuned on synthetic data only works on synthetic data. One person who can approve the spec and settle scope questions without a committee. Standups are written, so nobody sits in a daily call.
What happens if a build turns out to be the wrong fit midway?
We stop and say so. Every line on this page carries a written note about when it is the wrong call, and those are the same judgements applied during the build rather than only in the sales conversation. A sprint that ends early with an honest answer costs less than one that ends on time with a system nobody uses.
Not sure which one it is?
That is what the free half hour is for. Bring the process you want to fix; it ends with a recommendation either way, including the one that says do not build it.
Based in Brantford, Ontario — see local engagements across Brantford and Brant County, how delivery works city by city, read what agentic engineering means or the essays, or book the free 30-minute audit.