Builds — agents, MCP, retrieval, evalsBrantford, Ontario · remote across Canada

What gets built here

Nine AI builds in full — multi-agent orchestration, agent loop engineering, context engineering with GraphRAG and memory, MCP tool surfaces, evals and guardrails, in-product copilots, voice agents, document intelligence and AI strategy — each with the architecture it actually ships as, and what it is not the right fit for.

Builds
9
Typical build
4–12 wks
Team
1 operator

Index — nine builds, one operator

S/01 — Agentic workflows & orchestration5–10 weeks typical

Agentic workflows & orchestration

Multi-agent systems with planner/executor topologies, subagent fan-out, and durable execution underneath. Handoffs are explicit, checkpoints are human, and a crashed process resumes rather than restarts.

A single agent in a while-loop is a demo. An orchestration is what survives production: a planner that decomposes the request, subagents fanned out over the parts, an explicit handoff contract between them, and a merge step that reconciles disagreement instead of taking whichever answer arrived last. Underneath sits durable execution — Temporal or Inngest — so a deploy mid-run, a rate-limited provider, or a process dying at step nine replays from its last checkpoint rather than from the beginning. Everything is idempotent because at-least-once delivery is the only honest assumption, and anything that exhausts its retries lands in a dead-letter queue a person can read, fix and replay. Human-in-the-loop checkpoints are first-class: the run pauses, waits for an approval that may arrive tomorrow, and continues.

Usually because

  • A prototype works on the happy path and falls over on real traffic.
  • One long run holds state nobody can inspect, resume, or replay.
  • The work needs an approval partway through and currently blocks on it.
agents.flowPASS / FAIL LOOPS
The planner fans work out to subagents and the merge step reconciles what comes back. Everything below the gate is durability: three failures and the run lands in a dead-letter queue a person can replay, not a log nobody reads.
Usual stack
LangGraphTemporalInngestClaude Agent SDK
S/02 — Agent loop engineering2–5 weeks typical

Agent loop engineering

The harness under a single agent — context compaction, tool-call budgets, stop conditions, and a self-critique pass. This is where agents that spiral get fixed.

Most failing agents do not fail at the model. They fail at the loop around it: context that grows until the instructions fall out of attention, a tool-call budget nobody set, a stop condition that never fires, and a retry that feeds the same poisoned context back in. Loop engineering is the unglamorous fix — compaction that summarises completed work instead of re-reading every tool result, a budget on calls and tokens per run with a defined behaviour at the ceiling, stop conditions that distinguish finished from stuck, and a self-critique pass that checks the result against the original request before it is returned. Every decision the loop makes is traced with OpenTelemetry, so when a run costs eleven dollars you can see the step that did it rather than guess.

Usually because

  • An agent loops on itself, burns tokens, and returns nothing useful.
  • Cost per run varies by an order of magnitude and nobody can say why.
  • Long runs degrade — the model forgets its instructions halfway through.
loops.flowPASS / FAIL LOOPS
Two gates decide whether the loop continues. The budget gate stops a run that is spending without converging; the critique gate sends a weak answer back around once, then escalates rather than looping forever.
Usual stack
Claude Agent SDKOpenAI Agents SDKLangGraphOpenTelemetry
S/03 — Context engineering & memory4–8 weeks typical

Context engineering & memory

Hybrid retrieval, GraphRAG over a real knowledge graph, and agent memory that persists between runs. Ranked, reranked, cited, and compacted before it reaches the window.

Retrieval quality is a ranking problem before it is a model problem, and context is a budget before it is a prompt. The work runs in three layers. Retrieval: chunking that respects how a document is actually structured, hybrid keyword-plus-vector search rather than embeddings alone, a reranking pass, and citations rendered back to the source so an answer can be checked instead of trusted. Graph: entity extraction and resolution into a knowledge graph, so a question that spans three documents traverses relationships rather than hoping all three chunks rank in the same top-ten — the GraphRAG case, and the one plain vector search is worst at. Memory: episodic memory of what happened in past runs, semantic memory of what was learned from them, a written policy for what gets promoted and what is allowed to decay, and compaction so a long-lived agent does not carry a year of transcripts into every window.

Usually because

  • The answer exists across three documents and vector search finds one.
  • The agent re-learns the same fact on every run and remembers nothing.
  • A general model answers confidently and wrongly about your own product.
context.flowPASS / FAIL LOOPS
The index is built ahead — chunked, embedded, and resolved into a graph — and the question arrives later. The writeback edge is the memory layer: what the run learned goes back into the index instead of being thrown away.
Usual stack
pgvectorNeo4jCohere RerankLlamaIndex
S/04 — MCP servers & tool surfaces3–6 weeks typical

MCP servers & tool surfaces

Model Context Protocol servers that expose your internal systems to any agent — typed tool schemas, scoped OAuth, and a bounded surface rather than a database handed to a model.

MCP is how an agent reaches a system it does not own, and the protocol is the easy half. The work is the surface: tool schemas typed tightly enough that the model cannot express an invalid call, scopes that mean a compromised or confused agent reaches exactly one customer's records rather than the table they sit in, resources and prompts that give the model the shape of your domain without a paragraph of prose in every system message, and error messages written for a model to recover from rather than for a log. Auth is OAuth 2.1 with per-tool scopes, not a shared key in an environment variable. Servers ship on Vercel or Cloudflare with traces on every call, because a tool surface nobody can audit is a tool surface nobody should have connected.

Usually because

  • Every agent you build re-implements the same integrations from scratch.
  • Giving a model access to a system means giving it far more than it needs.
  • Claude, Cursor and your own agents each need the same internal tools.
mcp.flowPASS / FAIL LOOPS
The scope check sits between the client and anything it can reach. A call outside the granted scope is refused at the surface — the backing system never sees it, and the model gets an error it can recover from.
Usual stack
MCP SDKOAuth 2.1VercelCloudflare
S/05 — Evals, tracing & guardrails3–6 weeks typical

Evals, tracing & guardrails

Golden sets, LLM-as-judge scoring, and regression gates in CI — plus prompt-injection defence and red-teaming. The difference between shipping on evidence and shipping on vibes.

Without evals, every prompt change is a coin flip you cannot score. The build starts with a golden set drawn from your real traffic and labelled with your team, because a suite written from imagination measures imagination. Scoring combines deterministic assertions where the answer has a right shape with LLM-as-judge where it does not, and the judge itself is validated against human labels before it is trusted to grade anything. That suite becomes a regression gate in CI: a change that drops accuracy on the golden set does not merge. Alongside it runs the safety half — prompt-injection defence on every untrusted input, output filtering, a red-team pass against your specific tool surface, and OpenTelemetry traces so a production failure can be replayed as a new eval case instead of argued about in Slack.

Usually because

  • Nobody can say whether last week's prompt change made things better.
  • Quality is assessed by someone trying it a few times before a deploy.
  • Untrusted text reaches a model that can call tools, and nobody has tested that.
evals.flowPASS / FAIL LOOPS
The gate is the point: a change that regresses against the golden set does not merge. Failures return to the suite as new cases, so the set grows from production rather than from imagination.
Usual stack
BraintrustLangfuseLangSmithOpenTelemetry
S/06 — In-product copilots4–8 weeks typical

In-product copilots

AI inside the product itself — token streaming, generative UI, tool calling against your own API, and structured outputs written straight back to your database.

A copilot that lives in your product is a different build from one that lives in a chat window. It streams, because a spinner over four seconds of silence reads as broken. It renders generative UI — a real component the user can act on, not a paragraph describing what they could do. It calls your own API through tools scoped to the signed-in user's permissions, so the assistant can never surface a record that user could not open themselves. Outputs are schema-constrained with Zod and written back as structured rows, which is what makes the assistant part of the product rather than a text box beside it. This is also where tier-1 support deflection and content pipelines land when the right home for them is inside the product.

Usually because

  • Users know what they want and cannot find the screen that does it.
  • The same onboarding questions arrive as support tickets every week.
  • A workflow in the product takes eleven clicks and one sentence to describe.
copilots.flowPASS / FAIL LOOPS
Tool calls run under the signed-in user's permissions, so the assistant can never surface a record that user could not open. Output is validated against a schema before anything is written back.
Usual stack
Vercel AI SDKNext.jsZodPostgres
S/07 — Voice & realtime agents5–9 weeks typical

Voice & realtime agents

Speech-to-speech agents with turn detection, barge-in, and sub-second response — on the phone or in the product, handing to a human with the transcript attached.

Voice is a latency problem wearing an AI costume. Past roughly eight hundred milliseconds a caller starts talking over the agent, so the build is speech-to-speech rather than the transcribe-think-synthesise chain that stacks three round-trips before anyone hears anything. Turn detection decides when the caller has actually finished rather than merely paused, barge-in stops the agent mid-sentence when they interrupt, and both are tuned against your own recordings, because the right thresholds differ between a support line and a booking line. It runs on a real telephony path through Twilio or in-product over LiveKit WebRTC, with a transfer rule written down: on low confidence, on an explicit request, or on any topic you have declared off-limits, the call goes to a person with the transcript and the extracted details already attached.

Usually because

  • Calls go unanswered after hours and the voicemail is never returned.
  • The same booking or triage conversation runs forty times a day.
  • An IVR tree exists and callers press zero to escape it.
voice.flowPASS / FAIL LOOPS
Turn detection and barge-in run continuously, not once — the caller can interrupt at any point. Confidence decides between acting and transferring, and the transfer carries the whole transcript.
Usual stack
OpenAI RealtimeLiveKitDeepgramTwilio
S/08 — Document intelligence4–7 weeks typical

Document intelligence

VLM parsing over the documents you actually receive — scans, tables, handwriting — into schema-constrained records with a citation on every field and a confidence gate before anything is trusted.

The documents that matter are rarely clean text. They are scans, multi-page tables that break across a page, a stamp over the total, and a handwritten amendment in the margin. Vision models read those in a way OCR never did, but reading is not the deliverable — a validated record is. Extraction is schema-constrained, so the output is a typed object rather than prose to be parsed twice, and every field carries a citation to the page and region it came from, which is what turns a review from re-reading the document into checking a highlight. A confidence gate routes the difference: high-confidence records land straight in the system of record, low-confidence ones go to a human queue with the uncertain fields flagged, and every correction becomes an eval case so the gate gets better instead of staying where it was set on day one.

Usually because

  • Someone retypes numbers from PDFs into a system every morning.
  • OCR was tried, and the tables and handwriting defeated it.
  • An error is only found downstream, weeks after the document was filed.
docs.flowPASS / FAIL LOOPS
Nothing is written on the model's word alone. The confidence gate splits clean extractions from uncertain ones, and every human correction returns to the eval set rather than being fixed and forgotten.
Usual stack
ClaudeReductoUnstructuredZod
S/09 — AI strategy & audits1–2 weeks

AI strategy & audits

Readiness audits, build-versus-buy calls, model and data policy, and an eval plan before a quarter is committed. A few hours that often save months.

An audit of what you have, a roadmap ranked by time-saved-per-dollar, or a build-versus-buy review before you commit a quarter to the wrong path. It covers the questions that get expensive later: which of these belongs to a model at all rather than to a process change, what your data actually supports today, model and vendor policy including where an open-weight model on your own infrastructure is the right answer, and how any of it will be measured — an eval plan, before the build, rather than after the first incident. The deliverable is a written document you own outright and can hand to another vendor. A meaningful share end in a recommendation not to build the thing, which is what the opinion is for.

Usually because

  • There is a list of AI ideas and no honest way to rank them.
  • A six-figure platform quote needs a second opinion before it is signed.
  • Something shipped, and nobody agreed in advance how it would be judged.
strategy.flowPASS / FAIL LOOPS
The verdict is a real branch, not a formality. A meaningful share of these audits end at 'buy it' or 'do nothing', which is exactly what the opinion is being paid for.
What lands
AuditRoadmapEval planRetainer

Process — how an engagement actually goes

engagement.flowPASS / FAIL LOOPS
Every build runs this path. The spec is approved before anything is written, each slice is demoed on your own data, and the retainer at the end is optional rather than assumed.
  1. 01 · DISCOVER

    Diagnose, not prescribe.

    A working session to map every manual step in the workflow. I leave with a list of candidates ranked by time-saved-per-dollar.

    Week 1 · 2–3 calls · free
  2. 02 · DESIGN

    Blueprint on paper.

    A node-graph of your target system with every handoff, failure mode, and data contract named. You approve before anything runs.

    Week 1–2 · written spec
  3. 03 · BUILD

    Ship in slices.

    Narrow vertical releases every 3–5 days. You see the system working on your data before it's fully finished.

    Week 2–4 · daily standup in writing
  4. 04 · OPERATE

    Own the outcome.

    Monitoring, alerts, and iteration after launch. A monthly retainer keeps the system sharp as your business shifts.

    Month 2+ · optional retainer

Engagements — three ways to buy it

Week 1 · 2–3 calls · free

Discovery

A working session that maps every manual step in the workflow and ranks the candidates by time saved per dollar. It ends with a recommendation, which is sometimes that there is nothing here worth building.

  • A mapped picture of the current process
  • Candidates ranked by time saved per dollar
  • A written recommendation, yours either way
4–12 weeks · one build

Fixed-scope sprint

The default. A named scope, a written spec you approve before anything runs, and narrow vertical slices released every three to five days so you see the system working on your own data long before it is finished.

  • A written spec approved before the first line of code
  • A working slice every 3–5 days
  • Handover with the code, prompts, evals and infrastructure
Monthly · rolling

Ongoing retainer

For teams that need the judgement more than the hands: architecture review, build-versus-buy calls, vendor selection, and keeping the systems already shipped sharp as the business moves under them.

  • Architecture and code review on your cadence
  • Monitoring, alerts and iteration on live systems
  • Cancel-any-month, no notice period

Stack — what the builds actually run on

Tools are chosen per engagement, and the list below is what usually wins rather than a partnership page. Anything your team already runs well stays.

Agentic workflows & orchestration
LangGraphTemporalInngestClaude Agent SDK
Agent loop engineering
Claude Agent SDKOpenAI Agents SDKLangGraphOpenTelemetry
Context engineering & memory
pgvectorNeo4jCohere RerankLlamaIndex
MCP servers & tool surfaces
MCP SDKOAuth 2.1VercelCloudflare
Evals, tracing & guardrails
BraintrustLangfuseLangSmithOpenTelemetry
In-product copilots
Vercel AI SDKNext.jsZodPostgres
Voice & realtime agents
OpenAI RealtimeLiveKitDeepgramTwilio
Document intelligence
ClaudeReductoUnstructuredZod

Questions — before the first call

Can we start with one build instead of a whole platform?

That is the preferred way in. One build, shipped end to end and running on your data, tells you more about whether the rest is worth doing than any amount of planning does. Most engagements that grow started as a single agent nobody was sure about.

How do you decide between an agent and a fixed pipeline?

By whether the steps change per case. If the sequence is the same every time, a deterministic pipeline is cheaper, faster and far easier to debug — wrapping it in an agent buys nondeterminism nobody asked for. An agent earns its place when the work needs judgement: reading the situation, picking a tool, checking its own result, and escalating when confidence drops.

Why do evals come up in every one of these builds?

Because without them there is no way to tell a change from an improvement. A golden set drawn from your real traffic, a judge validated against human labels, and a regression gate in CI turn every prompt and model change from a coin flip into a measured decision. It is also the only honest way to answer whether a new model is better for your workload rather than better on a public benchmark.

Do you work with the tools we already have, or replace them?

Work with them, in almost every case. Agents reach your systems through an MCP server or typed tools rather than a migration, retrieval indexes the documents where they already live, and structured output writes back to the Postgres, CRM or help desk your team already reports on. Replacing a working tool is a cost with no output attached to it.

What do you need from my team during a build?

Access, one decision-maker, and about an hour a week. Access to the systems being automated and to real examples — real tickets, real documents, real calls — because a golden set built from imagination measures imagination, and an agent tuned on synthetic data only works on synthetic data. One person who can approve the spec and settle scope questions without a committee. Standups are written, so nobody sits in a daily call.

What happens if a build turns out to be the wrong fit midway?

We stop and say so. Every line on this page carries a written note about when it is the wrong call, and those are the same judgements applied during the build rather than only in the sales conversation. A sprint that ends early with an honest answer costs less than one that ends on time with a system nobody uses.

Not sure which one it is?

That is what the free half hour is for. Bring the process you want to fix; it ends with a recommendation either way, including the one that says do not build it.

Book the audit →

Based in Brantford, Ontario — see local engagements across Brantford and Brant County, how delivery works city by city, read what agentic engineering means or the essays, or book the free 30-minute audit.