Why Multi-Agent Systems Fail And How To Make Them Reliable
Research labelling more than 1,600 multi-agent runs finds failures cluster around unclear tasks, poor coordination and weak checking. Here is what goes wrong and where to put controls.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
Systems with several AI agents look convincing in demonstrations and often disappoint in use. A common explanation is that the models are not yet good enough. Research published in 2025 points somewhere less comfortable: many failures come from how the system is organised, in the same way that a capable team can still miss a deadline because nobody agreed what finished looked like.
This article summarises what that research found, groups the failures into three families, and sets out the controls that address each one: a shared definition of the goal, structured handoffs, explicit verification and human review points. It closes with how to measure whether an agent system succeeds consistently, rather than once in a while.
What The Research Found
In “Why Do Multi-Agent LLM Systems Fail?”, Mert Cemri, Melissa Z. Pan, Shuyi Yang and colleagues at UC Berkeley and elsewhere studied how multi-agent systems built on large language models break down.1 They began with the observation that such systems often show only small gains over a single agent on common benchmarks. To find out why, they first had experts analyse 150 traces (records of complete runs) from five frameworks, built a taxonomy from that analysis and checked it for agreement between the human annotators. They then used it to label a larger dataset of more than 1,600 traces from seven popular frameworks.
The result is MAST, the Multi-Agent System Failure Taxonomy. It identifies 14 failure modes in three categories: problems in system design, misalignment between agents, and weak task verification. The authors tested targeted fixes as well. ChatDev is a framework that runs a simulated software company, and on a set of 32 programming tasks the authors call ProgramDev-v0 its baseline completed 25.0 percent. Tightening the role prompts, so that only senior agents could close a discussion and the reviewing agent focused on the task’s requirements, lifted that to 34.4 percent. Reshaping the workflow into a loop that repeats until the lead technical agent confirms the reviews are satisfied, or a set number of rounds is reached, lifted it to 40.6 percent. The authors describe this as adding a check against the high-level objective. Those are gains of 9.4 and 15.6 percentage points over the baseline. The authors also concluded that such fixes leave many failures unresolved, so they do not on their own make a system dependable.
Three Families Of Failure
The table groups the MAST failure modes by category and pairs each family with the kind of control that addresses it. The failure mode names come from the paper; the controls are practical inferences, not findings the authors measured.
| Category | Example failure modes | What it looks like | Control that targets it |
|---|---|---|---|
| System design issues | Disobeying the task or role specification, repeating steps, losing conversation history, not knowing when to stop | Agents drift from their brief or loop without finishing | Precise task and role definitions, an explicit finish condition, turn and budget limits |
| Inter-agent misalignment | Conversation reset, failing to ask for clarification, task derailment, withholding information, ignoring another agent’s input, reasoning that does not match the action taken | Agents talk past each other or quietly change the goal | One shared statement of the goal, structured handoffs with defined inputs and outputs, a rule to ask when unclear |
| Task verification | Stopping too early, missing or incomplete verification, incorrect verification | The system declares success on work that is wrong or unfinished | A separate verification step against the original objective, and human review before results are relied on |
What stands out is how ordinary these failures are. Ignoring a colleague’s input, failing to ask a clarifying question and signing off work without checking it are familiar problems in human teams. They respond to the same remedies: clear briefs, clear handovers and real review.
Give Every Agent A Shared Direction
A useful first control is to define what completion means before dividing the work. A goal such as “improve the onboarding guide” invites each agent to interpret it differently, while “a revised guide that covers the five most common support questions, checked against current product screens” gives every agent, and every reviewer, the same finish line.
Each task handed to an agent then needs a clear input and an expected output. That turns a vague handoff into something that can be checked: did the agent receive what it needed, and did it return what was asked for? Dependencies between tasks should be visible to the people reviewing progress, so a stalled or failed step is noticed rather than silently worked around.
Verification And Human Review Points
Verification should be a distinct step with its own instructions, not something the working agent is trusted to do for itself. The evaluator and optimiser pattern described by Anthropic, in which one model produces work and another critiques it, is one way to build this in.2 The MAST results suggest the check should test the output against the original objective, not only against the last instruction passed along.
Human review belongs at defined points rather than everywhere. OpenAI’s guide names two triggers for handing control to a person: when an agent exceeds a set limit on retries or actions, and when it is about to take a sensitive, irreversible or high-stakes action such as cancelling an order, issuing a large refund or making a payment.3 Placing review there keeps people involved where their judgement matters without turning them into a bottleneck for routine steps.
- Define The Goal And Finish Line
Write down what completion means and share it with every agent and reviewer. Targets design failures.
- Specify Each Task
Give every task a clear input, an expected output and its dependencies, plus limits on turns and spend.
- Check Each Handoff
Validate that what passes between agents has the expected structure and content. Targets misalignment.
- Verify Against The Objective
A separate step compares the result with the original goal as well as the latest instruction. Targets verification failures.
- Human Approval For High-Stakes Actions
Pause for a person before irreversible or sensitive actions, or when failure limits are exceeded.
- Log And Trace The Run
Keep a record of tasks, handoffs, tool calls and decisions so failures can be diagnosed and accountability is clear.
Logging is easy to neglect and hard to do without. The NIST AI Risk Management Framework, which organisations adopt voluntarily, lists accountability and transparency among the characteristics of trustworthy AI, alongside being valid and reliable and being secure and resilient.4 Applied to an agent system, a practical reading is that you should be able to show afterwards who was asked to do what, what they did, and who approved it.
Measuring Reliability Across Repeated Runs
A system that succeeds once in a demonstration may fail on the next attempt. The τ-bench benchmark, published in 2024 by Shunyu Yao, Noah Shinn, Pedram Razavi and Karthik Narasimhan, tests agents that must use tools and follow policies while talking to a simulated user.5 It introduced a metric called pass^k, which asks whether an agent succeeds on the same task in every one of k attempts. Consistency turned out to be the weak point. In the retail scenarios, the pass^8 score of leading 2024 agents, roughly the chance of getting a task right on all eight of eight tries, was under 25 percent. Even GPT-4o, one of the strongest function-calling models tested, solved fewer than half of the tasks. Those figures describe models from 2024, and newer models will score differently, but the measure itself remains the right question for any business process.
Evaluate a multi-agent system on repeated runs of realistic tasks, and track consistency alongside average success. Because the system can fail through coordination as well as through any single model, test whole runs end to end. Security testing, including attempts to manipulate agents through the content they read, is a separate discipline covered in Agentic AI Security.
Footnotes
-
M. Cemri, M. Z. Pan, S. Yang, L. A. Agrawal, B. Chopra, R. Tiwari, K. Keutzer, A. Parameswaran, D. Klein, K. Ramchandran, M. Zaharia, J. E. Gonzalez and I. Stoica, “Why Do Multi-Agent LLM Systems Fail?”, arXiv:2503.13657, March 2025 (version 3, October 2025). arxiv.org ↩
-
Anthropic, “Building effective agents”, 19 December 2024. anthropic.com ↩
-
OpenAI, “A practical guide to building agents”, April 2025. cdn.openai.com ↩
-
NIST, AI 100-1, “Artificial Intelligence Risk Management Framework (AI RMF 1.0)”, January 2023. nvlpubs.nist.gov ↩
-
S. Yao, N. Shinn, P. Razavi and K. Narasimhan, “τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains”, arXiv:2406.12045, June 2024. arxiv.org ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.