"Through a controlled evaluation of 260 agent configurations, we find that stronger single agents can outgrow the benefits of collaboration. An empirical capability threshold predicts whether coordination helps or hurts in 94% of validation configurations, while a separate predictive model selects the best architecture in 87% of held-out configurations within the tested domains."
AI agents, systems capable of reasoning, planning and acting with tools, support applications from coding assistants to personal health coaches. These systems must gather information, respond to feedback and carry decisions through multiple steps. A mistake can affect later actions, making their design more complicated than optimizing the accuracy of an isolated prediction. One practical question is whether a task needs a team of agents or can be handled more effectively by a single capable model.
Research such as More Agents Is All You Need and work on collaborative scaling has demonstrated benefits from combining multiple agents. But these findings leave an important question: when do those benefits hold for tasks that require sustained interaction with an environment, especially as the underlying models become more capable?
In our Nature Machine Intelligence paper, “Capable language models can outgrow the benefits of collaboration”, we investigate this question across 260 configurations. The strongest statistical signal is the performance of the single-agent baseline. Collaboration can help substantially, but its value depends on what one agent can already accomplish and how much coordination costs.
Defining “agentic” evaluation
To study these trade-offs, we focus on tasks with three properties:
- Sustained multi-step interactions with an external environment.
- Iterative information gathering under partial observability.
- Adaptive strategy refinement based on environmental feedback.
We evaluated six benchmarks: BrowseComp-Plus for web research, Finance Agent for financial analysis, PlanCraft for planning, WorkBench for workplace tool use, SWE-bench Verified for software engineering, and Terminal-Bench for command-line tasks. We held task prompts, tool access and per-system compute ceilings constant while varying model capability and coordination structure across five architectures:
- Single-agent (SAS): One agent carries out reasoning and actions through a unified interaction history.
- Independent: Agents explore in parallel without exchanging messages; their outputs are combined without cross-validation.
- Centralized: An orchestrator delegates work and reviews and synthesizes worker outputs.
- Decentralized: Agents exchange findings through peer debate and aggregate their final answers by voting.
- Hybrid: An orchestrator directs work while agents also exchange information with peers.
Summary of the five architectures, including their computational complexity, communication overhead and coordination mechanisms. Here, k denotes iterations per agent; n, agent count; r, orchestrator rounds; d, debate rounds; p, peer communication rounds; and m, average peer requests per round. Independent agents do not exchange messages, but share environment state in the two software and terminal benchmarks.
Results: When collaboration stops helping
Across OpenAI, Google and Anthropic model families, stronger models generally performed better, but adding agents did not reliably improve on their single-agent baselines. Single-agent performance was the only predictor supported by both benchmark-clustered inference and correction for multiple comparisons.
A separately fitted decision rule identified a capability-saturation threshold near 45% single-agent success. Above this level, the rule predicts zero or negative gains from adding agents. It matched the direction of the observed multi-agent gain in 94% of 16 model–benchmark configurations on SWE-bench Verified and Terminal-Bench. This is an empirical selection rule, not a universal cutoff; those benchmarks used 20 task instances per configuration, so individual estimates remain uncertain.
Performance across three model families and five architectures, averaged over six benchmarks under matched per-system compute ceilings. More capable models do not consistently benefit from additional agents.
The benchmark-level results show why architecture selection matters. The box plots summarize model-level performance; annotated percentages are relative changes from the single-agent baseline, not percentage-point differences.
Multi-agent gains vary across domains: centralized coordination improves Finance Agent performance by approximately 81%, while all tested multi-agent variants underperform on PlanCraft. SWE-bench Verified shows small-to-moderate declines across variants, and Terminal-Bench shows mixed results.
Where coordination helps
On Finance Agent, centralized coordination improved mean performance from 34.9% to 63.1%, an 80.8% relative increase. Execution traces show agents researching complementary sources, such as regulatory news and company filings, before an orchestrator combines their findings. This structure allows useful work to be distributed across agents.
Where coordination gets in the way
On PlanCraft, every tested multi-agent architecture reduced performance, with relative declines of 39–70%. Traces show unnecessary delegation of short, sequential workflows into separate research, inventory and execution steps. Under a fixed system budget, these exchanges consume resources that could otherwise support task execution.
The cost of tool coordination
Tool-heavy workflows also showed a tendency toward higher coordination costs. However, the relevant interactions did not retain statistical significance after accounting for benchmark clustering. We therefore report this as a descriptive pattern to investigate, rather than a general law that more tools require fewer agents.
Architecture and error containment
Architecture also shaped how errors affected execution. Our trace-level error-amplification metric captures additional computational work associated with coordination failures. Independent systems had an amplification factor of 17.2, compared with 4.4 for centralized systems. These values do not mean that final answers were 17.2 or 4.4 times more likely to be wrong.
Trace-level coordination metrics show lower error amplification for centralized systems than for independent agents. The metric describes execution-level error dynamics and is distinct from the task failure rate.
The traces suggest that orchestrators can intercept inconsistent worker outputs before aggregation. The interaction between error amplification and single-agent baseline performance retained support under benchmark-clustered inference. This supports examining verification at agent handoffs, while stopping short of a general claim that centralized systems are always safer or more accurate.
A predictive model for agent design
We fitted a model using model capability, agent count, tool count, single-agent performance and measured coordination properties, including overhead and error amplification. It achieved cross-validated R² = 0.373, improving to 0.413 when capability was measured using performance on agentic tasks.
The model selected the best-performing architecture in 87% of held-out configurations within the evaluated domains. This architecture-selection result is separate from the threshold rule’s 94% prediction of whether collaboration helps or hurts. Neither establishes reliable prediction on entirely new domains: absolute performance prediction was poor when an entire benchmark was held out.
For practitioners, these findings suggest starting with a strong single-agent baseline, measuring coordination costs, and testing whether delegation or verification produces enough benefit to justify them.
Conclusion
As foundation models improve, agent architectures need to be reassessed. A team that helps a weaker model may add overhead once a stronger model can handle the same workflow alone. Our results offer empirical guidance for that decision: measure the baseline, examine how work and errors move between agents, and retain coordination where it delivers a demonstrated gain.
Acknowledgements
We thank our co-authors and collaborators at Google Research, Google DeepMind and our academic partner institutions for their contributions to this work.