Multi-Agent Systems: When They Pay Off and How to Build the Harness Around Them

Nish · October 10, 2026

50 min read

In February 2026, Nicholas Carlini, a researcher at Anthropic, described an experiment. He left sixteen copies of Claude to write a C compiler in Rust. Each copy ran in its own container, in a loop that started the next task as soon as the last one finished. The agents claimed work by committing a lock file to a shared git repository. They merged each other’s changes as they went, and no manager agent told them what to do. Over two weeks and nearly 2,000 sessions, they produced a 100,000-line compiler that builds the Linux kernel. The API bill was just under $20,000.1

For most of those two weeks, parallelism succeeded because the work had the right shape. The test suite held hundreds of independent tests, so sixteen agents could fix sixteen different failures at once. Then the agents reached the kernel itself, and progress stopped. Compiling Linux is one giant task. In Carlini’s words, “Every agent would hit the same bug, fix that bug, and then overwrite each other’s changes. Having 16 agents running didn’t help because each was stuck solving the same task.”

Carlini did not fix this with a better model or with more agents. He changed the harness: the software that runs the agents and checks their work. His new test setup “randomly compiled most of the kernel using GCC, and only the remaining files with Claude’s C Compiler.” A failure now pointed at a small set of files, so each agent could chase a different bug. The same sixteen agents became productive again, because the shape of the work had changed.

That episode holds the argument of this post. A multi-agent system is a harness with more than one context window in it. Its success depends mostly on the harness, not on the agents. The harness decides how the work is split and where state lives. It also decides how context moves between agents, who checks the output, and who can stop a run that has gone wrong.

The evidence from the past two years supports a narrower rule than enthusiasts or skeptics usually state. It has four parts:

  1. Keep decisions that interact inside one context window.
  2. Split work only where context can truly be separated. That mostly means reading, searching, and reviewing. It can also mean writing in isolated workspaces, if a mechanical check decides what merges.
  3. Put a check you trust between every agent and anything that leaves the system.
  4. Run the whole thing like the distributed system it is, with durable state, supervision, and a clear path to a human.

Here is the route through the rest of the post:

  • Whether to split. What counts as a multi-agent system, the case against one, and the cases where splitting pays.
  • What to build. The parts of a harness, then a reference design for many agents.
  • How to run it. Daily operations, a map of known failures, and one open-source example.
  • Where it stops. What no harness can fix.

It builds on four earlier posts. Working with coding agents argued that context is scarce and that a fresh agent should review another agent’s work. Goals and loops showed that an agent loop is only as reliable as the external check that stops it. Agent memory covered context rot, compaction, and memory stores. Removing the human from code review described the layered gates that let machines own a merge. This post asks what changes when several agents sit behind those gates at once, and what the harness around them must do.

Table of Contents

What counts as a multi-agent system

The useful way to define a multi-agent system is by its context windows, not by its personas or job titles. The agents discussed here use large language models, or LLMs.

Anthropic’s guide “Building effective agents” drew the distinction that most of the field still uses. Workflows are “systems where LLMs and tools are orchestrated through predefined code paths.” Agents “dynamically direct their own processes and tool usage.”2 Since then, the working definition of an agent has shrunk to a phrase. Anthropic’s context engineering guide calls agents “LLMs autonomously using tools in a loop,” and Simon Willison settled on “An LLM agent runs tools in a loop to achieve a goal.”

This post uses the same short definition. An agent is a model that calls tools in a loop until it reaches a goal. A tool is any function the agent can call, such as a file search, a shell command, or an API. The loop works inside one context window, which holds the information the model can use on each call.

Everything around the model is the harness. Anthropic’s guide to agent evaluations defines an agent harness as “the system that enables a model to act as an agent: it processes inputs, orchestrates tool calls, and returns results.” It then adds the sentence that matters most for this post: “When we evaluate ‘an agent,’ we’re evaluating the harness and the model working together.” LangChain’s Vivek Trivedy put it more bluntly in March 2026: “If you’re not the model, you’re the harness.”

A multi-agent system, then, is a harness that runs more than one agent loop on a shared task, each loop with its own context window. The definition is deliberately mechanical. It says nothing about personas, job titles, or agents that chat with each other. The number that matters is how many separate context windows must agree about something. Every boundary between two windows is a place where information gets summarized, dropped, or misread.

Four shapes cover nearly every system in the sources for this post. The figure compares who picks the next step: fixed code, one agent, an orchestrator, or a peer. These are control patterns, not a ranking by the number of context windows.

Four control patterns. A workflow uses code to choose the next step. A single agent chooses its own step. An orchestrator delegates to workers and combines results. Peers hand control to one another. The patterns differ in who decides, not simply in how many context windows they use.
Four ways to arrange model calls. The key difference is who chooses the next step; separate context windows also need a way to share decisions.
  • A workflow is code that calls a model at fixed points, such as a chain of prompts with a programmatic check between them. Your code picks the steps.
  • A single agent picks its own next step inside one context window.
  • Orchestrator and workers is what OpenAI’s “A practical guide to building agents” calls the manager pattern: “A central ‘manager’ agent coordinates multiple specialized agents via tool calls, each handling a specific task or domain.” In this post, the orchestrator is the agent that plans and delegates. A worker is an agent that does one delegated task in its own context window. Anthropic calls workers subagents.
  • Peers and handoffs is OpenAI’s decentralized pattern, in which “Multiple agents operate as peers, handing off tasks to one another based on their specializations.” A handoff is “a one way transfer” that moves execution to another agent “while also transferring the latest conversation state.”3

So the first question about any design is how many context windows must agree. Each extra window should buy something specific. Its boundaries also create more chances for information to be lost or misread.

The case against splitting the work

The best evidence says that extra agents often add cost and errors rather than quality.

On 12 June 2025, Walden Yan of Cognition, the company behind the Devin coding agent, published a post titled “Don’t Build Multi-Agents.” The next day, Anthropic published “How we built our multi-agent research system.” The titles suggest a disagreement. In fact, the two posts agree on more than they dispute. Read side by side, they are still the best introduction to the trade-offs.

Yan states two principles as rules: “Share context, and share full agent traces, not just individual messages,” and “Actions carry implicit decisions, and conflicting decisions carry bad results.” His example is a Flappy Bird clone, split into two subtasks: a background and a bird. One worker misreads its task and builds a background that looks like Super Mario Bros. The other builds a bird that does not look or move like the one in Flappy Bird. Even if both workers get the full task description, they cannot see each other’s choices. You get a bird and a background in two different visual styles.

The lesson is general. Each worker makes decisions about style, edge cases, and interfaces. Unless the harness brings those decisions together, the pieces can work alone and fail as a whole. Yan’s conclusion was blunt: “it is evident that in 2025, running multiple agents in collaboration only results in fragile systems.”

Anthropic’s post describes a system built on the exact pattern that Yan warned against. An orchestrator plans the research, and workers search in parallel. With an Opus 4 orchestrator and Sonnet 4 workers, it beat a single Opus 4 agent by 90.2% on an internal research evaluation.4 But the same post makes three admissions that matter more than the headline:

  • Cost. “agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats.”
  • Cause. In Anthropic’s analysis of the BrowseComp benchmark, “token usage by itself explains 80% of the variance” in performance.
  • Scope. “most coding tasks involve fewer truly parallelizable tasks than research, and LLM agents are not yet great at coordinating and delegating to other agents in real time.”

Read carefully, the post says two things. Multi-agent research works because it is a reliable way to spend more tokens on a task that splits cleanly. And coding often does not split cleanly.

The controlled evidence since then points the same way. The most careful study I know is “Towards a Science of Scaling Agent Systems,” by researchers at Google Research and MIT. It ran 260 configurations across six benchmarks and three model families, with the total reasoning budget held fixed.5 The results split by the shape of the task:

  • On tasks that split into parallel streams, coordination helped a lot. A central orchestrator improved results on a financial-analysis benchmark by 80.8%.
  • On tasks that need one chain of reasoning, coordination hurt. Every multi-agent variant did worse on a sequential planning benchmark, by between 39% and 70%.
  • On SWE-bench Verified, a coding benchmark, every multi-agent variant scored lower than a single agent on average. Relative declines ranged from 2.1% to 14.9%.

The paper’s most robust finding is what it calls capability saturation: “coordination yields diminishing returns beyond ~45% single-agent baselines.” In these experiments, stronger single-agent baselines left less room for coordination to help. The approximate 45% threshold is an observed pattern, not a universal cutoff.

Dat Tran and Douwe Kiela tested five multi-agent architectures against single agents in April 2026. They matched thinking-token budgets on questions that require several linked facts. Single agents generally matched or beat the multi-agent systems once budgets reached 500 tokens. Their reading is that “many reported advantages of multi-agent systems are better explained by unaccounted computation and context effects rather than inherent architectural benefits.”6

Anthropic’s own later guidance gives the bluntest warning. A January 2026 post on when to build multi-agent systems says that “teams invest months building elaborate multi-agent architectures only to discover that improved prompting on a single agent achieved equivalent results.” It puts the typical cost at “3-10x more tokens than single-agent approaches for equivalent tasks.” It also describes an experiment with agents split by software role: planner, implementer, tester, and reviewer. In that experiment, “the subagents spent more tokens on coordination than on actual work.”

Together, the case against has four parts:

  1. Context fragments. Summaries can drop detail at each handoff. Shared artifacts reduce that loss, but agents still have to read and interpret them.
  2. Implicit decisions conflict. Agents that write to the same system make incompatible choices that none of them can see.
  3. Cost multiplies. Wall-clock time often grows too.
  4. Failures are hard to locate. A 2025 study tested automated failure attribution across 127 multi-agent systems. The best accuracy for naming the responsible agent was 53.5%. For naming the failure step, it was only 14.2%, achieved by a different method.7

My own reading is that the strongest results for multi-agent systems are mostly stories about spending more compute. When compute is held equal, the remaining advantage sits in tasks that truly split.

Where splitting the work pays

Extra agents pay off when they read, search, or review in their own context windows, and when they keep noise out of the main one.

Ten months after his first post, Yan wrote a follow-up called “Multi-Agents: What’s Actually Working.” By then Cognition shipped multi-agent systems itself. The revised position fits in one sentence: “multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions.”

The most surprising finding reverses his own first principle for one role. Cognition’s code reviewer worked best when the coding agent and the review agent “do not share any context beforehand.” Yan gives two reasons. A reviewer with a clean context must reason backward from the code, without first being led through the coder’s assumptions. A fresh context reduces that influence; it does not remove shared model biases. And a shorter context makes the model sharper. He also dismissed the fashionable version of the idea: “the unstructured-swarm approach, arbitrary networks of agents negotiating with each other, is mostly a distraction. The practical shape is map-reduce-and-manage.”

Anthropic’s January 2026 guidance reaches the same place from a different direction. It names three situations where several agents consistently beat one: “when context pollution degrades performance, when tasks can run in parallel, and when specialization improves tool selection or task focus.” It also gives the most useful rule I have found for where to cut: “adopt a context-centric view rather than a problem-centric view when decomposing work.”

An example shows the difference. A split by type of work gives one agent the features, one the tests, and one the review, and it forces constant handoffs. A split by context keeps related work together: “an agent handling a feature should also handle its tests, because it already possesses the necessary context. Work should only be split when context can be truly isolated.”

Harrison Chase of LangChain reconciled Yan’s first post and Anthropic’s within days. His summary still holds: “read actions are inherently more parallelizable than write actions. When you attempt to parallelize writing, you face the dual challenge of effectively communicating context between agents and then merging their outputs coherently.” He also noted that Anthropic’s research system obeys this rule itself. Many agents read, but the final report is “deliberately handled by a single main agent in one unified call.”

From these sources, and from the failures in the last section, I count five reasons to add an agent that survive scrutiny:

  1. Breadth on reads. Research, search, and whole-codebase scans split into independent pieces whose results merge later. Anthropic’s research system and the financial-analysis gains both live here. Cognition’s “Agentic MapReduce” makes the idea precise for code. A deterministic pass makes a finite list of files to inspect, agents investigate the pieces in parallel, and a reducer combines their conclusions. Coverage is then “guaranteed by construction.”
  2. Protecting the main context. A worker can spend tens of thousands of tokens on exploration and return “a condensed, distilled summary of its work (often 1,000-2,000 tokens),” in the words of Anthropic’s context engineering guide. Cognition found its agents spent more than 60% of their first turn only on retrieving context. So it trained a dedicated retrieval worker that favours precision over recall, “because we found that context pollution matters.”
  3. Independent verification. An agent is a poor judge of its own work. Anthropic’s Prithvi Rajasekaran found that agents asked to evaluate their own output “tend to respond by confidently praising the work,” even when “the quality is obviously mediocre.” He adds that “tuning a standalone evaluator to be skeptical turns out to be far more tractable than making a generator critical of its own work.” A verifier is any part of the system that decides whether work is correct, such as a test suite or a reviewer agent. A separate verifier with a clean context is the multi-agent pattern with the most consistent support.
  4. Separating capabilities. OpenAI’s guide notes that tool overlap, not tool count, is what confuses an agent: “Some implementations successfully manage more than 15 well-defined, distinct tools while others struggle with fewer than 10 overlapping tools.” A split also lets you separate dangerous combinations of permissions, which the section on sandboxes covers.
  5. Parallel writes, with isolation and a mechanical check. The C compiler work succeeded while each agent had its own failing test, and it stalled when all of them shared one. Give each writer its own workspace, then check the combined result before it merges. Isolation prevents file collisions; it does not prevent incompatible design choices. The checks must encode the acceptance criteria and merge policy that a human chose.

Barry Zhang of Anthropic put the same caution into three rules at the AI Engineer Summit in February 2025: “First, don’t build agents for everything. Second, keep it simple. And third, think like your agents.” The figure below turns that caution into four questions, asked in order. It is the path I would walk before I add a second agent to anything. A single writer can still use an independent reviewer when the risk warrants one.

Four questions before adding agents: are the steps known, does one agent do the task reliably, is the remaining work mostly reading or writing, and can each writer have its own workspace and a mechanical check? Start with a workflow or one agent; add read-only workers or isolated writers only when needed.
Start with a workflow or one agent. Add workers for a specific need; parallel writers also need isolated workspaces and checks on the combined result.

The reason to add an agent should be specific: more breadth, a cleaner context, an independent check, or separated permissions. A second writer also needs a reliable way to integrate its work.

The anatomy of a harness

A harness is everything around the model, and each of its parts changes how reliable an agent is.

If the harness decides whether a multi-agent system works, it is worth being precise about what a harness contains. Every part below belongs in a good single-agent setup. A multi-agent system needs all of them for every agent, plus the coordination parts in the next section. In the figure, the black box is the model and the numbered path is the control loop. The labelled groups around the loop are the parts of the harness, and each part has a subsection below.

The model sits inside a harness. The control loop assembles context, asks the model for a tool call, runs the tool, and checks whether to stop. Context, durable state, and instructions shape what the model sees. Tools, permissions, and hooks shape what it can do. Traces, evals, and the human surface help check the result.
The model is one small box. The harness around it decides what the model sees, what it can do, and how anyone knows whether it worked.

The control loop

The control loop is simple, and the harness owns its three key decisions. Dex Horthy’s “12-Factor Agents” describes the loop. The model picks the next step as a structured tool call. “Deterministic code executes the tool call,” and the result goes into the context. This repeats “until the next step is determined to be ‘done’.” OpenAI’s description of the Codex loop is the same: it ends when the model stops asking for tools and answers.

The harness makes three decisions inside that loop:

  • What counts as done. Goals and loops argued that something other than the model doing the work should judge this.
  • The budget. Agents are bad at judging effort. Early versions of Anthropic’s research system spawned “50 subagents for simple queries.” So the system now carries explicit rules, such as “Simple fact-finding requires just 1 agent with 3-10 tool calls.”
  • What happens on error. Horthy argues for a limit: “Hitting some consecutive-error-threshold might be a great place to escalate to a human.”

Tools

Tool design often matters more than the prompt. Anthropic’s guide contains a confession worth remembering whenever an agent’s tool list keeps growing. While building their SWE-bench agent, “we actually spent more time optimizing our tools than the overall prompt.” One small change was to require absolute file paths instead of relative ones. It removed a whole class of errors, and “the model used this method flawlessly.”

Anthropic’s later tool-design guide calls tools “a new kind of software which reflects a contract between deterministic systems and non-deterministic agents.” Its advice follows from that contract:

  • Build a few tools around real workflows, rather than one for every endpoint.
  • Give tools clear namespaces.
  • Return meaningful names, not opaque identifiers.
  • Limit the size of responses. Claude Code limits tool responses to 25,000 tokens by default.

Error messages are part of the tool interface too. OpenAI’s harness team writes custom lint errors that “inject remediation instructions into agent context.” Each failed check then becomes a lesson that the agent can act on at once.

Context and compaction

The harness decides what the model sees, and less is usually better. Anthropic’s context engineering guide states the goal as “finding the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome.” Two methods do most of the work in long runs:

  • Compaction summarizes a conversation that nears the context limit, and the agent continues from the summary. Codex compacts automatically when the context passes a threshold.
  • Structured notes keep progress outside the context window, in a to-do list or a notes file. The agent reads them back when it needs them.

Yichao “Peak” Ji’s July 2025 lessons from building the Manus agent add cost, which most guides skip. Agents read far more than they write: Manus averages about 100 input tokens for every output token. So the cache hit rate on the unchanging start of the prompt is, in Ji’s words, “the single most important metric for a production-stage AI agent.”8

That cost model leads to rules that look odd until you see it:

  • Keep the start of the prompt stable.
  • Append to the context instead of editing it.
  • Make every compression restorable. For example, drop a web page but keep its URL.
  • Leave failed actions in the context, because “Erasing failure removes evidence. And without evidence, the model can’t adapt.”

The agent memory post covers why long contexts degrade in the first place.

Durable state

An agent that runs for hours needs durable state: files and git history that survive when a process stops. Justin Young’s November 2025 post on harnesses for long-running agents describes the problem as “a software project staffed by engineers working in shifts, where each new engineer arrives with no memory of what happened on the previous shift.” His fix is entirely about state outside the model:

  • a progress log that the agent appends to;
  • a feature list in JSON, where each entry starts as failing;
  • git commits that let a later session revert bad changes and recover working code.

One detail is worth copying directly. The feature list is JSON, not Markdown, because “the model is less likely to inappropriately change or overwrite JSON files compared to Markdown files.”

OpenAI’s harness engineering post makes the general point: “From the agent’s point of view, anything it can’t access in-context while running effectively doesn’t exist.” So plans, decisions, and architecture rules live in the repository as versioned files, not in chat threads or in people’s heads. Anthropic’s Managed Agents design applies the same idea to infrastructure. It keeps an append-only session log outside the harness, so that “nothing in the harness needs to survive a crash.”

A simple test follows. If the process that runs an agent died right now, could a fresh process continue the work from what is on disk?

Permissions and sandboxes

Hard boundaries protect a system better than a human who approves each action. Approval prompts look safe, but they are not. Anthropic reported in May 2026 that “users approved roughly 93% of permission prompts,” and that “The more approvals a user sees, the less attention they pay to each.” Its sandboxing work drew two hard boundaries instead: filesystem isolation and network isolation. Both are needed to contain the escape routes described in that design. Together they cut permission prompts by 84% in internal use.

The deeper problem is prompt injection, where text that an agent reads carries instructions that the agent then follows. Simon Willison’s “lethal trifecta” names the dangerous combination:

  • access to private data;
  • exposure to untrusted content;
  • the ability to communicate externally.

Meta’s “Agents Rule of Two” turns that into a design rule. Within a session, an agent should have no more than two of the three. An agent that needs all three “should not be permitted to operate autonomously.”

Separating agents can help enforce these boundaries. A reader can inspect untrusted pages without credentials or write access. But its summary is still untrusted: it can carry an injected instruction to the agent that acts on it. The receiving agent still needs restricted tools and approval for sensitive actions. A summary is not a security boundary.

Hooks

Put rules that must hold every time into blocking hooks. A hook runs at a fixed point in the loop. Use deterministic code for rules that must be enforced; some harnesses also support hooks that ask a model to judge. Typical points are before a tool call, after it, and when the agent tries to stop. Claude Code’s documentation lists them as “user-defined shell commands, HTTP endpoints, MCP tool calls, LLM prompts, or subagents that execute automatically at specific points in Claude Code’s lifecycle.”

A rule in a prompt asks the model to comply. A blocking code hook can enforce a rule before an action happens. A hook can:

  • block a forbidden command before it runs;
  • refuse to let an agent end its turn until the tests pass;
  • write every action to a trace.

OpenAI’s harness team applies the same idea to the codebase itself. Custom linters and structural tests enforce which layers may depend on which. Their summary is the best sentence I know on this subject: “By enforcing invariants, not micromanaging implementations, we let agents ship fast without undermining the foundation.” They add that this kind of architecture is something “you usually postpone until you have hundreds of engineers. With coding agents, it’s an early prerequisite.”

The human surface

The human surface is every point where a human can see, steer, or stop the system. OpenAI’s guide names the two triggers that should return control to a human: “Exceeding failure thresholds” and “High-risk actions.” It defines high-risk actions as those “that are sensitive, irreversible, or have high stakes.”

Two design choices make the surface useful rather than noisy:

  1. Asking is a tool. Asking a human should be an ordinary tool that the agent can call. This is Horthy’s seventh factor (“Contact humans with tool calls”), and it is how Claude Code’s question tool works. “I need a decision” then becomes a normal outcome, not a failure.
  2. The evidence comes with the question. An escalation should state what the agent tried, what the checks said, and what the options are.

The code review post made the same argument for merge decisions. A human who gets a raw transcript and a vague question will either approve it unread or redo the work.

Tracing and evals

You cannot debug a multi-agent run by reading its final answer. You need a trace, a record of every model call, tool call, and handoff in the run. Anthropic added “full production tracing” to its research system to diagnose why agents failed. In production, it also monitors “agent decision patterns and interaction structures.” A common schema for traces is emerging. OpenTelemetry’s conventions for generative AI now define spans, the timed steps within a trace. These cover creating and invoking an agent, planning, executing a tool, and invoking a workflow “composed of multiple agents.” The conventions are still marked as in development.9

Evaluation needs the same care. An evaluation, or eval, tests how well a system handles a set of tasks. Anthropic’s January 2026 guide to agent evals separates two things:

  • the transcript, which is everything the agent did;
  • the outcome, which is the real state of the world afterwards. For example, the outcome is whether a booking exists in the database, not whether the agent said it made one.

The guide recommends a start with “20-50 simple tasks drawn from real failures.” It ends with an instruction worth printing out: “Read the transcripts!”

For systems that run unattended, reliability over repeated runs matters more than best-case success. The τ-bench benchmark introduced pass^k, the chance that all of k attempts at a task succeed. A GPT-4o agent succeeded on about 61% of retail tasks on a single try. Its rate of success on all of eight tries fell below 25%.10

The harness and the model are evaluated together, so every harness change is an experiment. Rajasekaran’s advice is to remove “one component at a time” and measure. If you cut several parts at once, you cannot tell which ones did the work.

A reference architecture for many agents

When a task truly needs several agents, arrange the harness so that the agents coordinate through files and checks, not through conversation.

The figure below is the arrangement I would start from. It is a composite of the systems in this post, not any one product. Follow the numbers:

  1. The human sets a goal for the orchestrator.
  2. The orchestrator writes a brief for each worker into durable state.
  3. Each worker reads its brief and works in its own workspace.
  4. Each worker writes its status and its artifacts back to durable state. A status change wakes the orchestrator.
  5. A worker submits finished work to the verifier.
  6. Only work that passes the verifier merges.

The dashed arrows are the return paths. Failed work returns to its worker, and the orchestrator escalates exceptions to the human.

A human gives a goal to an orchestrator. The orchestrator writes briefs to durable state. Isolated workers read their briefs and write status and artifacts. Finished work goes through checks, a reviewer, and a merge policy before it merges. Failures return to workers; exceptions go to the human.
Agents coordinate through durable files, not conversation. Short status messages go through the orchestrator, content stays in files, and a verifier decides what merges. Only exceptions reach the human.

The orchestrator

Something has to own the plan. Cursor learned this the hard way while it ran hundreds of coding agents on one project.11 Its first design gave every agent equal status and a shared file for coordination, with locks to stop two agents from claiming the same task. The result: “Twenty agents would slow down to the effective throughput of two or three, with most time spent waiting.”

Cursor then replaced the locks with optimistic concurrency. That was more robust, but it exposed a deeper problem: “With no hierarchy, agents became risk-averse. They avoided difficult tasks and made small, safe changes instead. No agent took responsibility for hard problems or end-to-end implementation.” What worked was a split into separate roles:

  • planners that explore the codebase and create tasks;
  • workers that “just grind on their assigned task until it’s done”;
  • a judge that decides at the end of each cycle whether to continue.

Microsoft’s Magentic-One, published in November 2024, shows what the orchestrator should track.12 Its orchestrator keeps two ledgers:

  • an outer task ledger, with facts, guesses, and the plan;
  • an inner progress ledger, which answers five questions on every step. Two of them are “Is the team looping or repeating itself?” and “Is forward progress being made?”

A counter tracks how long the team has been stuck. Past a small threshold, the orchestrator stops, reflects on what went wrong, and revises the plan. Then every agent clears its context and starts fresh. In the paper’s ablation test, performance fell by 31% without the full ledgers.

Isolated workers

Each worker needs its own context window and its own workspace. A workspace is a private copy of the files, in a container, a virtual machine, or a git worktree. Carlini gave every agent a Docker container with a private clone of the repository. Cognition’s managed Devins each run “in its own isolated virtual machine with its own terminal, browser, and development environment.” OpenAI’s harness team made their application bootable per git worktree, so an agent could launch and drive its own copy for each change.

Separate workspaces keep concurrent edits apart. A worktree alone does not restrict which files an agent can access; that requires an enforced sandbox. Separate workspaces also make failed attempts easier to discard, once useful work and failure evidence are saved. Keeping related work together supports the context-centric split. A worker that owns a feature in its own worktree naturally owns that feature’s tests.

The brief is a contract

The brief is the written task that the orchestrator gives a worker, and it is the orchestrator’s most important output. Anthropic found that vague briefs were the main source of wasted work in its research system: “Each subagent needs an objective, an output format, guidance on the tools and sources to use, and clear task boundaries.” In its example, a brief asked workers to “research the semiconductor shortage.” Then “one subagent explored the 2021 automotive chip crisis while 2 others duplicated work investigating current 2025 supply chains.”

Other teams found two more failures. Rajasekaran’s harness made the generator agent and the evaluator agent negotiate a contract before each sprint, “agreeing on what ‘done’ looked like for that chunk of work before any code was written.” Cognition found the opposite problem in its orchestrator agents. Trained on small tasks, they “default to being overly prescriptive.” Cognition also noted that “Agents assume they share state with their children when they don’t.”

A good brief sits between those failures. It states:

  • the goal, and why it matters;
  • what the worker may and may not touch;
  • the check that defines done;
  • where to write results;
  • when to stop and ask.

A brief in a file is also something a human can audit later. Ryan Lopopolo of OpenAI gave a useful test for briefs in an April 2026 talk at AI Engineer Europe: “Every time I have to type “Continue” to the agent is a failure of the harness to provide enough context around what it means to continue to completion.”

Artifacts instead of conversation

Workers should return files, not essays. An artifact is a file that a worker produces, such as a plan, a diff, a finding, or a log. Anthropic recommends that workers write their output to a filesystem “to minimize the ‘game of telephone’.” They “store their work in external systems, then pass lightweight references back to the coordinator.” Cursor’s later kernel-optimization system kept its “entire coordination protocol” in “a single markdown file that specified the output format, rules, and tests.”

The general rule is to separate two channels:

  • The control channel carries short status messages, such as “started,” “blocked,” and “done.” They go through the orchestrator.
  • The content channel carries artifacts. They live in durable files that any agent or human can read directly, without a chain of summaries.

Cognition’s MapReduce post gives one more reason to prefer artifacts that you can inspect. Its file selectors can be read, tested, and tuned, “whereas a search agent’s ‘I’ve looked everywhere’ is unfalsifiable.”

Verification as its own layer

Verification needs its own agents and its own budget. The failure study later in this post found that many multi-agent systems already had a verifier. But the verifier mostly checked whether code compiled, or whether TODO comments remained. One system produced a chess program that passed its review phases and then failed on the actual rules of chess.

Anthropic’s April 2026 guide states the cause plainly: “The verifier is only as good as its criteria.” It goes on: “Teams most often fail by implementing the loop without defining what verification means, which creates the illusion of quality control without the substance.” Its January guide describes the typical shortcut: “The verifier runs one or two tests, observes them pass, and declares success.” It says that an explicit instruction to run the complete test suite “is essential.”

The order from the code review post applies here unchanged:

  1. Run the cheap mechanical checks first: tests, types, and linters.
  2. Then run a reviewer agent with a fresh context and explicit criteria.
  3. Then let a written merge policy decide what merges and what goes to a human.

Where MCP and A2A fit

Two open protocols appear in almost every discussion of multi-agent systems, and they solve different problems:

  • The Model Context Protocol (MCP) standardizes how an agent reaches tools, data, and prompt templates. Anthropic released it in November 2024 and donated it to the Linux Foundation’s new Agentic AI Foundation in December 2025.13
  • Google’s Agent2Agent protocol (A2A) standardizes how one agent hands a task to another. Google announced it in April 2025 and released a stable version 1.0 in March 2026.

In A2A, each agent publishes an “Agent Card” that describes what it can do. Work is a task that moves through defined states, such as working, input required, completed, and failed. Agents collaborate “without needing to share their internal thoughts, plans, or tool implementations.” The A2A project summarizes the split as “MCP inside agents, A2A between agents.”

Protocols matter when agents cross an organizational boundary, where nobody can share a filesystem or a trace store. They do not fix the problems that this post is mostly about. The failure study later in this post found coordination failures “even when agents within the same framework communicate using natural language.” And A2A’s principle of opaque execution, sensible between companies, is the opposite of Cognition’s first principle inside one system. Use the protocols at the edges. Keep the inside of the system as transparent as a shared repository and a trace store can make it.

Together, these parts give each decision an owner: the orchestrator owns the plan, workers own bounded tasks, and the verifier applies the merge policy.

Running it day to day

A multi-agent system that runs for days is a distributed system, and it needs the same operational care.

It must handle processes that die, processes that hang, new versions that ship in the middle of a task, and costs that grow quietly. The figure below follows one worker. The numbered main path runs from brief to merge. The three side loops are the ways a worker can leave that path, and each one has a defined recovery. Every step writes a status line to durable state, and the orchestrator wakes when that state changes.

A worker moves from brief to working, reports done, then passes through a verifier before merging. Questions go to a human, silent or looping workers need recovery from checkpoints, and failed verification returns work for repair. The orchestrator reads durable state when a status event or timer wakes it.
Each way a worker can fail has a defined recovery. The orchestrator learns about all of them from durable state, so it can wake on events and check for silence on a timer.

Supervision and liveness

The best model for supervision is older than language models. Joe Armstrong’s 2003 thesis on Erlang describes it. Ericsson built Erlang for telecom systems that must keep running through faults. The thesis describes “Supervision trees,” in which supervisors “monitor other processes” and must “be able to start, stop and restart the things they are monitoring.” Its slogans include “Let some other process do the error recovery” and “Let it crash.”

The translation to agents is direct. A worker should not try to recover from every problem inside its own context. It should write its status to durable state, and let the orchestrator decide whether to answer, nudge, restart, or escalate.

The orchestrator needs two different detectors, because agents fail in two different ways:

  • A heartbeat catches agents that have died. In durable execution systems such as Temporal, a heartbeat is a periodic ping that tells the system “the Activity Execution is making progress and the Worker has not crashed.”
  • A stall counter catches agents that are alive but do nothing useful, which is the more common failure. Magentic-One’s counter is an example.

Heartbeats catch death and stall counters catch futility. A harness needs both.

Agents also cannot feel time pass. Carlini noted that Claude “can’t tell time and, left alone, will happily spend hours running tests instead of making progress.” Cursor reports that its “Agents occasionally run for far too long.” So the harness must enforce limits on wall-clock time and turns from outside. A good orchestrator wakes on events, such as a worker writing a new status, with a slow timer as a backup. It does not poll everything all the time.

Failure and recovery

Recovery should resume from a checkpoint rather than start again. Anthropic’s research post states the core problem: “Agents are stateful and errors compound.” A restart from zero is expensive. So Anthropic “built systems that can resume from where the agent was when the errors occurred.” It combines “retry logic and regular checkpoints” with a simpler step: it tells the agent when a tool is failing and lets the agent adapt.

Two design rules follow:

  • A retry must be safe to repeat. LangGraph’s documentation warns that when a paused graph resumes, side effects before the pause “should (ideally) be idempotent,” which means a repeat has no extra effect.
  • An upgrade must not break runs in progress. Anthropic uses “rainbow deployments.” Old and new versions run side by side while traffic shifts gradually, because agents “might be anywhere in their process” when new code ships.

Erlang and Manus seem to disagree here. Erlang says to crash and restart. Manus says to keep failures in the context so the model can learn from them. Both are right, at different levels. Restart the process, so that a confused context stops compounding. But keep the evidence in durable state, so that the next attempt knows what failed and why. Magentic-One’s re-plan step does exactly this: every agent clears its context, and the task ledger records what went wrong.

Cost, latency, and the limits of parallelism

Extra agents trade coordination and computation for broader coverage. Whether they save time depends on how much work can run independently. Anthropic’s January 2026 guide adds that “The primary benefit of parallelization is thoroughness, not speed.” It explains that “multi-agent systems often take longer overall than single-agent systems because of the sheer increase in total computation.” Yet Anthropic’s research team reported cutting time by up to 90% on complex queries after parallelizing agents and tool calls. Measure time and cost on your own workload; neither follows from the agent count alone.

Harnesses with heavy verification cost more again. Rajasekaran’s full harness used a planner, a generator, and an evaluator. It took six hours and $200 to build a retro game maker. A single agent attempted the same task in 20 minutes for $9. Rajasekaran judged the difference in quality “immediately apparent.”14 Whether that trade is worth it depends on what a wrong answer costs you.

Parallel work also has limits that no prompt removes:

  • Shared locks. Locks cut Cursor’s twenty agents to the throughput of two or three.
  • Tasks that do not split. Carlini’s kernel turned sixteen agents into sixteen copies of one agent.
  • Serial work. Some planning, integration, and merge decisions must happen in order. The more agents you add, the more this serial part dominates.

The last serial step is usually a human’s attention. Willison wrote about running several coding agents at once in February 2026. He described reaching “the stage of parallel agent psychosis where I’ve lost a whole feature.” Past the point where a human can still track what the agents do, more agents do not add throughput. They add work that nobody watches.

Escalation and evaluation in production

Write the escalation rules before the system runs. They should state:

  • which actions need approval;
  • how many failed attempts send a task back to a human;
  • which combinations of capabilities never run without a human.

Evaluation continues after launch. Anthropic’s eval guide separates capability evals from regression evals. Capability evals “should start at a low pass rate,” so there is room to improve. Regression evals “should have a nearly 100% pass rate,” so any drop is a signal. Production traces become the source of new eval tasks.

Headline capability numbers also need care. METR’s measurements of how long a task AI agents can complete keep rising. But Thomas Kwa of METR warned in January 2026 that “A 50% time horizon of X hours does not mean we can delegate tasks under X hours to AIs.” A task that you delegate unattended needs far better than even odds of success.

The scarce resource is still a human’s attention. Durable state, checkpoints, and clear escalation rules let the system use that attention where it matters.

How multi-agent systems fail

The most thorough study of multi-agent failures sorts them into three groups, and each group points to controls worth building into the harness.

The study is “Why Do Multi-Agent LLM Systems Fail?” by Mert Cemri, Melissa Pan, Shuyi Yang, and colleagues at UC Berkeley, presented at NeurIPS 2025. The method had three steps:

  1. The authors built a failure taxonomy, MAST, from more than 150 execution traces with six expert annotators.
  2. They checked that independent annotators applied it consistently. Agreement was strong, with a Cohen’s kappa of 0.88.
  3. They used it to label 1,642 traces from seven multi-agent frameworks.15

They found 14 failure modes in three groups. The figure shows the three most common modes in each group. Each failure group is paired with controls that target it. These are design responses to test, not guarantees that a failure will disappear.

MAST groups failures into system design issues, inter-agent misalignment, and task verification. The most common modes include step repetition at 15.7%, reasoning-action mismatch at 13.2%, and incorrect verification at 9.1%. Proposed controls include clear briefs, progress checks, shared artifacts, and independent verification.
MAST's three failure groups, each with its three most common modes as shares of all failures in 1,642 traces. The proposed controls show where the harness can help; their effectiveness still needs testing.

System design issues are the largest group. Agents repeat steps that they already completed. They fail to recognize when the task is finished, and they ignore constraints in the task or in their role. The brief and the loop give us places to intervene, even when model limitations also contribute. Useful controls include:

  • a brief that states acceptance criteria;
  • a stop condition that something other than the worker owns;
  • a progress ledger with a stall counter, of the kind Magentic-One uses.

Inter-agent misalignment covers the failures that Yan predicted. An agent’s actions do not match its own stated reasoning. Agents drift off the task, guess instead of asking, or withhold or ignore information. The paper notes that these failures happen “even when agents within the same framework communicate using natural language,” so better message formats alone will not fix them. What helps is less need to communicate:

  • one writer for each decision;
  • shared artifacts instead of relayed summaries;
  • a cheap way to ask a human.

Task verification is the group this post has returned to most often. It covers checks that are wrong, checks that are missing, and work declared done too early. The paper’s own conclusion is that “Multi-Level Verification is Needed.” A single check at the end, with low-level tests, is not enough.

Two more results from the paper are worth taking into any design review. First, the authors tested two kinds of fix on ChatDev, a framework that simulates a software company:

  • Clearer role definitions improved success on their programming tasks by 9.4 points.
  • A change of structure improved it by 15.6 points over the baseline. It turned a one-way pipeline into a loop. The loop ended only when the agent in the role of chief technology officer confirmed that every review was satisfied.16

The authors conclude that “topology-based changes are more effective than prompt-based changes.” In these interventions, changing the workflow helped more than rewriting role prompts. Second, they frame the problem as one of organization rather than technology: “even organizations of sophisticated individuals can fail catastrophically if the organization structure is flawed.”

The same group’s July 2026 follow-up turned the taxonomy into a part of the harness. A coding agent received structured failure labels instead of free-form reflections. Its success on a small SWE-bench subset rose from 60% to 68% with MAST’s fixed labels. With labels learned from the agent’s own traces, it rose to 70%.17

None of the fixes in this section is a better model. Each one changes the brief, the loop, or the verifier.

A worked example: firstmate

One system I have used is firstmate, an open-source project by Kun Chen. It is a clear example of layering: a multi-agent harness built on top of single-agent harnesses. Almost every part in this post appears in it as a file or a script that you can read. Everything below comes from its public README, design notes, agent instructions, and architecture documentation.

First, what kind of thing it is. Its README is explicit: “firstmate is not a model, not a harness, not a skill, not an MCP server, and not a CLI.” It calls itself “an agent distro for running a crew of agents,” which it defines as “a portable directory of instructions, skills, tooling, policies, and state conventions that turns a general-purpose agent into a specialized one.”

The README’s “not a harness” is accurate in the narrow sense. Claude Code and Codex are harnesses in that sense: each runs one model’s agent loop, with its tools, context management, and permissions. firstmate does not implement an agent loop. It sits one layer above. Its instructions, scripts, hooks, and durable state turn one of those harnesses into an orchestrator of others. In this post’s broader sense, the whole setup is a multi-agent harness built from single-agent harnesses. The general point is that coordination does not need a new agent loop. It can live in files, scripts, and hooks around loops that already exist.

The roles match the reference architecture closely:

  • The human is “the captain” and talks to one agent only, the first mate, which is the orchestrator. It “dispatches, supervises, escalates only real decisions, and reports plain outcomes.”
  • Workers are called crewmates. One rule keeps the human surface narrow: “Crewmates never address the captain. All crewmate communication flows through firstmate.”
  • Workspaces are git worktrees. The first mate gives each worker “a clean git worktree,” so that “parallel work on one repo never collides.”
  • Briefs are files written before a worker starts. The design notes state the contract principle directly: “Every task gets an explicit contract before it starts: what to build or learn, how it ships, and how much autonomy the worker has; the machinery refuses to guess.”

The operations follow the same pattern:

  • Status. A worker reports by appending one line to a status file. The documentation is careful that these lines are events, not the truth about a worker: “Crew status files are append-only wake-event logs, not current-state fields.”
  • Supervision. “A zero-token bash watcher … sleeps on the fleet, classifies detected wakes in bash, and wakes the first mate only when something is actionable.” Actionable wakes go into a durable queue, so an interrupted orchestrator can recover them.
  • Hooks. One rule says “No turn ends blind while work is under way.” In Claude Code, for example, Stop hooks block the first mate from ending its turn while work is under way and supervision is not live.
  • Verification. In its no-mistakes project mode, work ships through a separate pipeline, no-mistakes, which “puts a local git proxy in front of your real remote.” It runs review, test, documentation, and lint steps before anything reaches the real repository. Mechanical findings are fixed automatically, while “anything that touches your intent is escalated for you to approve, fix, or skip.”

Three lines from the design notes summarize the approach, and each matches a lesson from earlier in this post:

  • “Logic that can be exact lives in deterministic scripts; work that requires understanding lives in an agent; the two never mix.”
  • “Workers are supervised, not trusted.”
  • “Everything that matters survives the death of any conversation.”

Even the context budget is a written rule. The instructions that the orchestrator loads on every turn carry “a stated ceiling of 9,000 words,” and a change that would exceed it must first prune or move something. The coordination rests on files, scripts, hooks, and an explicit merge authority. The public README also documents direct-PR and local-only modes, plus an optional setting for autonomous merges.

The design makes one trade explicit. Every worker reports to one orchestrator, and by default every merge waits for one human’s approval. So throughput is limited by how many escalations that human can handle well. firstmate makes escalations rare and well formed, while keeping merge authority explicit. The code review post takes a different next step: a human writes a policy that can allow routine merges without per-PR approval.

The limits of the harness

A harness can only enforce what it can check, and some of what makes software good cannot yet be checked.

It would be convenient to end by saying that the harness is everything. It is not, and the strongest counter-argument comes from Dex Horthy, the author of 12-Factor Agents. In “Why Software Factories Fail,” an essay and keynote from the AI Engineer World’s Fair in July 2026, he argues that “no amount of harness engineering or loopsmaxxing can solve what is fundamentally a model-training issue.”

His reason is that models are trained to make tests pass. “There’s no way in this system that we can penalize it for poor program design or for eroding the maintainability of our systems.” The cost of bad architecture appears “in months and years,” far beyond any check that a harness can run before a merge.

Others admit the same gap. OpenAI’s harness team has a million lines of agent-written code, and still writes: “What we don’t yet know is how architectural coherence evolves over years in a fully agent-generated system.” Cognition makes the same point about the large swarm demonstrations. The browser, the compiler, and similar projects “all share a property most real software doesn’t: a simple, verifiable success criterion.”

I think these objections are right about scope, and they do not change the design advice. Because a harness can only enforce what can be checked, humans must stay in charge of architecture, specifications, and the policies that decide what merges. Every harness in this post puts them there. The objections are not an argument to add agents in the hope that more of them will supply the judgment that one agent lacks.

Back to the kernel

Carlini’s sixteen agents did not get smarter when they started to make progress on Linux again. The harness around them changed. One giant task, where every agent collided, became many small tasks that a known-good compiler could check one at a time. The same pattern runs through every system in this post:

  • Cursor’s agents stopped avoiding risk when the structure gave someone ownership of hard problems.
  • Anthropic’s research agents stopped duplicating work when their briefs got boundaries.
  • ChatDev’s success rate rose when its conversation could end only after the reviews were satisfied.

Rajasekaran’s harness post adds the point that keeps this advice from being static: “every component in a harness encodes an assumption about what the model can’t do on its own,” and those assumptions “can quickly go stale as models improve.” Anthropic added context resets to work around one model’s habit of rushing to finish. On the next model, the resets became “dead weight.” The earlier post on prompts written for the last model found the same decay in instruction files.

So the harness needs maintenance like any other system. Remove a part, measure, and keep the part only if the numbers say it still pays. Rajasekaran’s conclusion is that “the space of interesting harness combinations doesn’t shrink as models improve. Instead, it moves.” The agents will keep improving. The harness is where your judgment about the work lives, and that is the part you have to keep building yourself.

Sources & further reading

The debate about multiple agents

Evidence on when it works and how it fails

Building the harness

Safety, operations, and evaluation

Protocols

The worked example

  • firstmate: an open-source agent distro for supervising a crew of coding agents; its README, VISION.md, and architecture notes are worth reading as a design document.
  • no-mistakes: the validation pipeline that sits between a worker’s branch and the real remote.

Footnotes

  1. Carlini’s post reports 2 billion input tokens and 140 million output tokens on Opus 4.6. The compiler passes most of the test suites he tried, but he is candid that its output is less efficient than GCC with optimizations disabled. These costs and results come from one researcher’s experiment, not a representative sample of projects. ↩

  2. The post, by Erik Schluntz and Barry Zhang, is dated December 2024, but Anthropic has revised it in place since, so some of its examples mention products released later. The definitions and the tool-design anecdote quoted in this post already appear in an archived copy from March 2025. ↩

  3. OpenAI’s guide carries no printed date or byline; its PDF metadata dates it to April 2025. Anthropic’s April 2026 catalogue of coordination patterns adds finer types, including generator-verifier loops, long-lived agent teams, message buses, and shared-state systems, but all of them are variations on these four shapes. ↩

  4. The evaluation is internal and unpublished, with no sample size given, and the comparison did not hold total compute equal. The 15× figure compares multi-agent runs with chat, not with a single agent, so on Anthropic’s own numbers a multi-agent run uses roughly 15 / 4 ≈ 3.75 times as many tokens as a single-agent run. That is not a dollar-cost ratio: model prices and the mix of input, cached, and output tokens also matter. ↩

  5. Kim et al., arXiv 2512.08296, version 3 (April 2026). The Google Research blog post from January 2026 describes an earlier version with 180 configurations and four benchmarks, so its numbers differ. The coding results use 20-task subsets because of evaluation cost, which leaves wide confidence intervals. The paper also reports that independent agents amplify trace-level errors 17.2 times against 4.4 times for a central orchestrator. These are architecture-level constants, though, and task-level amplification is far smaller. In the paper’s own regression, error amplification stops being significant once efficiency and overhead are accounted for. ↩

  6. At the tightest budget, 100 thinking tokens, some multi-agent setups did better. Individual model and budget combinations also varied; the result is an overall pattern, not a win in every table cell. The study covers a single task family, factual multi-hop questions, so it says nothing direct about long tool-using runs. ↩

  7. Zhang et al., “Which Agent Causes Task Failures and When?”, ICML 2025. The two numbers come from different methods. The method best at naming the agent did worse than random at naming the step. ↩

  8. Caching dominates cost because every token kept in context is processed again on each call unless its attention state is reused. The earlier post on the KV cache works through that bill. Ji’s post quotes a tenfold price difference between cached and uncached input tokens at the time. ↩

  9. In 2026 the generative AI conventions moved out of OpenTelemetry’s core semantic-conventions repository into a dedicated one, so older links to the agent-span pages now land on a notice that the page has moved. ↩

  10. The more familiar pass@k asks whether any of k attempts succeeds, which rises with k. pass^k falls with k. Anthropic’s guide gives the arithmetic: a 75% per-trial success rate becomes 0.75³ ≈ 42% for three trials in a row. ↩

  11. Wilson Lin, “Scaling long-running autonomous coding,” Cursor, January 2026. The headline project, a web browser written from scratch over close to a week, produced over a million lines of code. The post gives no functional evaluation of the browser, so read it as evidence about coordination at scale rather than about output quality. ↩

  12. Fourney et al., “Magentic-One,” Microsoft Research, November 2024. The stall threshold was two in their experiments, a tuned setting rather than a principle, and the 31% drop was measured on one benchmark with one model. ↩

  13. The July 2026 revision of the MCP specification made the protocol stateless and deprecated its sampling feature, which let a tool server ask the calling model to generate text and was occasionally used to nest one agent inside another. ↩

  14. These are single runs of one prompt by one author, not averages. A later version of the same harness on a newer model built a different application in about four hours for $124.70. ↩

  15. Version 3 of the paper, October 2025. The first category was called “Specification Issues” in earlier versions and “System Design Issues” in the current one. The percentages are each mode’s share of all failures observed across the 1,642 traces, labelled by an LLM judge that agreed with human experts 94% of the time. ↩

  16. The programming benchmark had only 32 tasks, and on a second framework the statistical significance depended on the model. Treat these as small intervention studies, not reliable estimates of the gain in another system. ↩

  17. Cemri et al., “Fantastic Adaptive Taxonomies and How to Use Them,” July 2026. The subset is SWE-bench Verified Mini, which is small, and the comparison point is a reflection-based baseline rather than an untouched agent. ↩

Citation Information

If you find this content useful, please cite this work as:

Bhana, Nish. "Multi-Agent Systems: When They Pay Off and How to Build the Harness Around Them". Nish Blog (October 2026). https://www.nishbhana.com/Multi-Agent-Systems/

Or use the BibTeX citation:

@article{bhana2026multiagent,
  title   = {Multi-Agent Systems: When They Pay Off and How to Build the Harness Around Them},
  author  = {Bhana, Nish},
  journal = {nishbhana.com},
  year    = {2026},
  month   = {October},
  url     = {https://www.nishbhana.com/Multi-Agent-Systems/}
}

x.com, Facebook