Topical authority hub

Multi-Agent Systems in Production

When many agents beat one and when they do not: the decision gate, orchestration topologies, coordination, cost, and how to evaluate the result.

multi-agent systems
Articles
11
Total read
152m
Pillar
Set
Start here — foundational guide
A second agent is a cost, not an upgrade. Multi-agent architecture buys parallelism and isolation and charges tokens, coordination, and compounding failure, and most teams pay without the gain. A default one-agent design, cheap and consistent with one context and one writer, gives way through a gate only when you can prove you need more to a many-agent design that costs roughly fifteen times the tokens with coordination conflicts and reliability decay.
PillarFoundational guide

Most Multi-Agent Systems Would Work Better as One Agent

A team splits its working agent into a planner, three researchers, a critic, and a synthesizer. Latency triples, the bill jumps tenfold, and the researchers return three contradictory answers because none saw what the others found. Multi-agent architecture is a cost you pay for two things, parallelism on independent subtasks and isolation for specialists, and most teams pay it without getting either. It multiplies tokens, multiplies coordination failures, and compounds unreliability across every hop. Default to one capable agent; reach for many only on a clear gate, and then share structured state, route instead of fanning out, and keep writes single-threaded.

Mudassir Marwat17 minJul 1, 2026
Read the guide

Continue the path

Ordered by chapter. Each post stands alone but builds on the one before it.

It is not how many agents, it is how you wire them. The same agents wired two ways differ by about forty points of reliability and ten times the cost, so pick the topology from the task dependency graph. Four topology glyphs: a sequential line, an orchestrator fanning out to workers, a hierarchical tree of supervisors, and a fully connected network mesh. Topology is the decision, head count is a consequence.
Chapter 1

Four Ways to Wire a Multi-Agent System (and When Each One Breaks)

You settled the question of whether to use multiple agents. Now comes the choice that matters more than the head count: how to wire them. There are four topologies, a sequential pipeline, an orchestrator with parallel workers, a hierarchy of supervisors, and a peer-to-peer network, and each fits one shape of work and hides one failure. The same three agents at eighty percent each are about fifty-one percent reliable wired to all-must-succeed, about ninety percent as a majority vote, and about ninety-nine percent when any one can and you can verify it. Topology, not head count, sets your reliability. Underneath it, share structured state instead of messages, keep the writes single-threaded, and read the shape off the task dependency graph.

17 minRead
When a multi-agent system fails, which agent broke? The final answer will not tell you, because a multi-agent system fails in the seams between agents rather than inside any one of them. Evaluate in three layers: outcome, which answers whether it failed; component, which answers which parts work and gives the per-agent reliability; and trajectory, which answers where in the flow it broke and is the layer most teams skip.
Chapter 2

When a Multi-Agent System Fails, Which Agent Broke?

You decided to use multiple agents, and you wired them. Now the system gives a confident wrong answer and hands you no stack trace, and the hardest production question arrives: does this work, and when it does not, which agent broke? Single-agent evaluation does not transfer, because a multi-agent system fails in the seams. You need three layers: outcome tells you whether it failed, component tells you which parts work and gives you the per-agent reliability, and trajectory tells you where in the flow it broke. Put them together and the reliability math becomes a diagnostic: if the parts predict about seventy-three percent and the system delivers fifty-five, the gap is an interaction bug. And because these systems are non-deterministic, one green run is not a passing grade, so you measure a rate over many runs off a trace you can actually read.

18 minRead
Your agents do not share a brain, they pass notes. Every hand-off compresses one agent's full working state into a message, and the next agent acts on that message alone, so the context one agent has is not the context the next one gets. Four ways a hand-off fails: context loss, where the sender drops what you needed; semantic drift, where the same words carry a different meaning; error laundering, where an uncertain guess arrives as a clean fact; and the coordination tax, where re-sending the accumulated context makes token cost grow quadratically. The fix is a shared store, not a better message.
Chapter 3

Your Agents Don't Share a Brain. They Pass Notes.

You wired the agents, and you learned the failures live in the seams between them. This chapter names the seam: it is the hand-off. When one agent finishes and the next begins, everything the first learned, its documents, its discarded hypotheses, its doubt, has to survive a trip through a single message, and the next agent acts on that message alone. Your agents do not share a brain, they share a mailbox, and every message is a lossy compression of what the sender knew. Four things go wrong: context loss, semantic drift, error laundering, and a coordination tax that grows with the square of the chain length. The fix is not a better prompt. It is a shared, typed store the agents read and write, plus hand-off contracts that make a bad message fail loudly instead of quietly.

10 minRead
Your multi-agent system has no brakes. You designed which agents exist and how they connect, but not the thing that decides when they stop, and in most systems nothing does. Control is emergent: it falls out of agents calling each other, so the system stops only when something runs out. Four ways control fails: non-termination, where the loop never ends; premature stop, where an agent's "done" was not actually done; mis-routing, where the task orbits and never lands; and no budget, where the rare tail run wrecks the bill. The fix is explicit control, not a smarter agent.
Chapter 4

Your Multi-Agent System Has No Brakes

A multi-agent system with no termination condition does not fail, it spends. The stopping criteria that actually hold, and the ones that only look like limits.

9 minRead
Your multi-agent system is a black box. A broken microservice returns a 500, but a broken agent returns a 200 OK and a confident wrong answer, so your dashboard stays green while the system is wrong. Four things you cannot see without instrumenting the run: who spent the tokens, where it broke, whether the run is healthy, and whether the brakes tripped. Instrument it, or you are flying blind and calling it production.
Chapter 5

Your Multi-Agent System Is a Black Box

You gave the agents brakes in the last chapter. Brakes stop a runaway, but they do not tell you which agent is dragging, what a run costs, or why last night's batch tripled. Most multi-agent systems ship to production as a black box: the team can see that a run happened and returned, and nothing in between. Worse, agent failures are silent, they return a 200 OK with a confident wrong answer, so uptime and error rate read green while the output is wrong. This chapter builds the instrument panel: the trace tree as the unit of observability, per-span cost attribution that finds the one agent owning two-thirds of the bill, structural sampling that keeps every run worth investigating, and continuous semantic evaluation, the only signal that catches a wrong answer no system metric will flag. Brakes stop a runaway. Instruments show it coming.

11 minRead
One poisoned agent poisons the chain. A prompt injection does not stop at the agent that reads it, and in a multi-agent system one untrusted document can reach every tool you have. How one document wins, in four hops: an agent reads the poisoned chunk, the instruction rides the hand-off as its provenance is dropped, it reaches the privileged agent that holds the dangerous tool, and the data is sent out while no system metric flags it. Scope the tools, or one document owns them all.
Chapter 6

One Poisoned Agent Poisons the Chain

You gave the agents brakes, then instruments. Neither stops a poisoned document from turning the pipeline against you. A single model call has one trust boundary, trusted system prompt versus untrusted input; a multi-agent system erases it, because every agent's output becomes the next agent's trusted input. So an injection planted in one document read by one agent propagates through the hand-offs to the agent that can send email or write to a system of record. The lethal trifecta, untrusted content, private data, and the ability to communicate out, gets split across your agents and recombined by the chain, so a per-agent review clears every agent and still misses the exploit. You cannot filter your way out. What works is structural: least privilege so the reader holds no dangerous tool, quarantine untrusted content behind a tool-less model, and a deterministic guardrail on the dangerous action. Tracing shows the attack. It does not stop it.

12 minRead
Your retry just sent the email twice. A single call fails all at once, so a retry is safe. A multi-agent run fails in the middle, and the retry fires again. How a retry duplicates, in four steps: the run fails mid-way when finalize times out, the retry replays and re-runs steps one to five, the email sends again because the tool holds no idempotency key, and the result is two identical filings while no metric flags it. Make the replay safe, or the retry doubles the damage.
Chapter 7

Your Retry Just Sent the Email Twice

You secured the tool boundary against an attacker. Now the tool fails on its own, with no attacker in sight: a model call times out halfway through the run, a retry kicks in, and the send-email tool that already fired once fires again. A single model call fails atomically, so retrying the whole call is safe. A multi-agent run is fourteen calls with no transaction around them, so it fails in the middle, after some agents have already sent the email or written the row, and there is no rollback. A blind retry replays the steps that succeeded and duplicates their side effects, so the retry is not recovery, it is a second bug. What works is structural, at the action boundary: make every side-effecting tool idempotent with a key so a replay is a no-op, checkpoint each step so a retry resumes instead of restarts, and compensate the steps you cannot make idempotent. Across fourteen calls, one run in eight lands partial. The failure rate does not change. The cost of a failure does.

13 minRead
You pay for the same context fourteen times. A single call you pay for once. A naive multi-agent run re-reads the same context on every call. Where the bill goes, in four steps: input dominates, about eighty percent of the tokens are prompt and not output; context accumulates, each handoff adds to it; it is re-sent on every call, so fourteen calls read the same context; and the result is about fifteen times a single call, not six times for six agents. Cut the tokens per call, not the number of agents.
Chapter 8

You Pay for the Same Context Fourteen Times

You made a run reliable and secure. Now look at what it costs. A single model call has an obvious, fixed cost, paid once. A multi-agent run is dominated by input tokens, and in a naive pipeline every agent re-ingests the accumulated context, so you pay to re-read the same system prompt, tools, and documents on every one of the roughly fourteen calls. The bill scales with context-per-call times number-of-calls, not with agent count, so six agents cost closer to fifteen times a single call, not six. Cost control is structural: minimize the context at every handoff, tier the models so the mechanical agents run cheap, cache the static prefix that repeats every call, and do not fan out when a single call will do. Stacked, they take a naive run from about forty-one thousand frontier-priced tokens to roughly a seventh of the cost for the same output. The most effective lever is minimization, because a token you never send is free at every tier, in every cache, and on every retry.

13 minRead
There is no such thing as a local change. Change one agent and you change every agent below it, so the unit of change is the pipeline, not the agent. Why a change spreads, in four steps: output is input, because one agent feeds the next; there is no typed contract, because outputs are natural language; coupling propagates, because a tweak shifts the whole chain; so you re-test the graph, not just the agent. Ship the pipeline, not the agent.
Chapter 9

There Is No Such Thing as a Local Change

You made a run cheap. Now try to change it. In a multi-agent system there is no such thing as a local change, because every agent output is the next agent input and the contract between them is natural language, not a typed schema, so nothing stops a change from spreading. Edit one agent prompt, swap its model, or change a tool, and you have moved the behaviour of every agent downstream, which was tuned on the old output. The unit of change is the whole pipeline, not the agent, so you cannot ship it component-by-component the way service intuition says. Version the pipeline as one artifact, keep golden transcripts and replay every change against them, canary on live traffic before you promote, and pin the model version so the vendor cannot change your pipeline under you. The sneakiest change of all is the cost move of tiering an agent down to a cheaper model, because it rewrites the input of the next agent while looking like a harmless config tweak.

15 minRead
You built a line when the work was a graph. Your latency is the longest path through the graph, not the sum of the agents. Where the seconds go, in four points: six round-trips, run one at a time for about twelve seconds; the critical path, the longest chain of dependencies, which sets the floor; parallel branches, where independent agents run at once; and the tail compounds, because every hop multiplies the chance of a slow run. Shorten the path, not the agents.
Chapter 10

You Built a Line When the Work Was a Graph

You made the run cheap and safe. It is still slow, and not because any one agent is slow. Your latency is not the sum of your agents. It is the longest path through the graph of what depends on what, the critical path, because two agents that do not read each other output can run at the same time. Teams build the pipeline as a straight line, so a six-agent chain at about two seconds an agent takes twelve seconds when the real dependency graph allows eight. The fixes are structural. Draw the dependency graph not the flowchart, parallelize the independent branches so the two extractors run at once, collapse the hops that do not earn their round-trip, stream the last agent so time-to-first-token is short, and cap the tail, because a six-hop chain slow case compounds with every hop. One catch: latency and cost pull in opposite directions, so speed bought by running more agents at once is often paid for in tokens.

17 minRead
Build it for real

Read the writeup. Now ship the system.

Cognilium engineers ship the architectures behind these articles for enterprise teams. If you're mid-build on multi-agent systems in production, talk to us.

AI for your ERP

Dynamics 365 & SAP

Schedule a call