Short answer: adding more agents to a job mostly makes it faster and more expensive, not better. Evals company Vals AI tested GPT-6 Sol and Claude Opus 5.5 on its Vibe Code Bench, once as solo agents and once as teams, at medium and maximum reasoning effort. According to The Decoder's report, the teams cost between 1.8x and 5.1x more, and only one of four team-versus-solo comparisons showed a statistically significant improvement.
That lands in the middle of a year when "agent teams" became the default pitch for coding tools, as we argued in Agent Teams Are the New Org Chart. This post walks through what the Vals result says, what Anthropic's scaling numbers and Noam Brown's comments add, and how to decide between one agent and many.
TL;DR: what the study found
| Question | Answer |
|---|---|
| Who ran it? | Vals AI, an evals company that publishes Vibe Code Bench |
| Models tested | GPT-6 Sol and Claude Opus 5.5 |
| Setups | Single agent versus team, each at medium and maximum reasoning effort |
| Cost of teams | 1.8x to 5.1x more than the solo agent |
| Quality gain | One of four comparisons significant: GPT-6 Sol at medium effort, +7.3 points |
| At maximum effort | Neither model gained anything real from a team |
| Takeaway | Extra agents mostly buy speed; at full compute they add cost without score |
One caveat up front: I could not find Vals AI's own write-up of the team experiment. The figures above come from The Decoder's summary, which includes a Vals chart plotting cost per app against score, with arrows from each model's single agent to its team. The benchmark itself is documented on the Vals leaderboard page. Treat the team numbers as reported secondhand until Vals publishes the full methodology.
What Vibe Code Bench actually tests
Vibe Code Bench asks a model to build a complete web application from a short natural-language spec, then scores the result by driving the app with an automated agent that clicks through it. The specs cover things like sign-up and login, posting messages, following users, search, and a discovery page. Each spec is paired with 20 to 60 automated tests.
Models work inside a modified OpenHands harness with a terminal, a browser, and sandboxed services for the database, storage, payments, and email. They can work for up to five hours or 1,000 turns per app. The leaderboard, updated October 7, 2026, shows a frontier clustered around 90 percent: Claude Sonnet 5.5 at 92.39 percent, Claude Opus 5.5 at 90.29 percent, and GPT-6 Sol at 87.82 percent for single-agent runs.
Those leaderboard cost figures give useful scale. On the single-agent board, Opus 5.5 costs $57.92 per test and GPT-6 Sol $26.36. Multiply by a team factor of 1.8x to 5.1x and one app build moves from tens of dollars to potentially a few hundred. The leaderboard numbers are not the team experiment's numbers, so use them only for a sense of magnitude.
This benchmark is a good stress test for teams because an app is one coherent artifact. Every file depends on the schema, the auth flow, and the routing decisions made earlier. That is the kind of work where parallel agents tend to step on each other.
The four comparisons, in plain language
Vals paired each model with each reasoning level, giving four team-versus-solo matchups:
- GPT-6 Sol, medium effort. The only significant win: the team scored 7.3 points higher.
- GPT-6 Sol, maximum effort. No real advantage from the team.
- Claude Opus 5.5, medium effort. No significant improvement.
- Claude Opus 5.5, maximum effort. No real advantage.
The pattern is intuitive. A team partly substitutes for thinking time. When a model is already running at medium effort, splitting the work across several agents can recover some of what deeper reasoning would have bought. When it is already at maximum effort, the single agent has used its whole budget on the problem, and the team adds coordination without adding capability.
That reading fits the cost story too. If you are paying for maximum reasoning on one agent and then multiplying the bill by up to five for a team, you are paying twice for the same lever.
Why teams cost so much more
The study does not break down the token spend, but the mechanics are well understood in agent engineering:
- Duplicated context. Each agent needs enough of the spec and the current codebase state to act. That context is paid for once per agent, on every turn.
- Handoff messages. Planner-to-worker and worker-to-reviewer messages are output tokens that a solo agent never has to write.
- Redundant exploration. Two agents reading the same files or running the same failing command is pure waste.
- Merge and repair. When parallel edits conflict, someone has to reconcile them, usually with another long turn.
OpenAI developer Eric Provencher made a related warning, per The Decoder, calling agent swarms most likely wasted money because coordination between agents breaks down. He called it the coordination tax.
Anthropic's own scaling numbers say the same thing
The Decoder also pulls in Anthropic's data from two Opus 5.5 tests. In both, more agents reached a given level faster, but returns flattened.
| Task with Opus 5.5 | 1 agent | 10 agents | 30 agents | 100 agents |
|---|---|---|---|---|
| Knowledge base | 0.53 | 0.70 | 0.71 | 0.74 |
| Lean theorem proving | 0.39 | 0.66 | 0.66 | 0.68 |
Going from one agent to ten is a big jump on both tasks. Going from ten to a hundred moves the score by only a few hundredths, while the token bill keeps growing. After 24 hours, the 100-agent setup was only slightly ahead of the 10-agent one, per the report. In separate ProgramBench tests, speed gains came with higher token usage.
The Decoder adds that Claude Fable 5.1 showed stronger gains on Lean theorem proving above ten agents but still scored below Opus 5.5 across all tests, and that on the knowledge base task its score dipped slightly from 30 to 100 agents. I could not locate Anthropic's original write-up of these runs in the time available, so cite them as reported by The Decoder.
Three robotic grippers placing different shapes into one shared tray, a picture of agents sharing one workspace
Noam Brown: you are buying speed
The most useful framing comes from OpenAI researcher Noam Brown. On the Dwarkesh Podcast, as reported by The Decoder, he said multi-agent systems mainly buy speed rather than quality: four agents solved tasks twice as fast but also cost twice as much, and at 16 agents the pattern held but grew slightly less efficient.
He also stressed that the effect depends on the task. Web research and math parallelize well. Writing a novel does not, and throwing 10,000 agents at a novel would be as pointless as throwing 10,000 people at it. Brown acknowledged that scaling to very large agent counts is largely unexplored because the costs are simply too high.
That is a clean rule of thumb. Ask whether your task decomposes into independent chunks with a cheap merge step. If yes, a team is a speed tool. If no, it is a cost multiplier.
What this means for what you build or pay
If you run coding agents or build agent products, three practical consequences follow.
1. Default to one strong agent. Before adding agents, raise reasoning effort or improve the prompt and tools. Our breakdown of what a Claude Code task costs on Opus 5.5 shows how cache hits, turn count, and effort move the bill more than the sticker price, and the Vals result suggests team size belongs on that list too.
2. Pick the right model before the right topology. On Vibe Code Bench the cheapest strong results come from model choice, not agent count. Claude Haiku 5.5 scores 90.44 percent on the leaderboard at $6.07 per test, and our small-model comparison covers the cheaper tier. Compare that with multiplying a $57.92 Opus 5.5 run by up to 5.1x.
3. Use teams where the work splits cleanly. Research sweeps, test generation across independent modules, and batch migrations of unrelated files are good candidates. A single coherent feature is not. If you want an interface for supervising helpers, see how Hermes Agent handles manual subagent control, and for company-style orchestration see Paperclip.
A simple way to test this on your own work
You do not need Vals AI's harness to check whether a team earns its keep. Run the same ten tasks three ways and compare cost and pass rate:
Setup A: one agent, medium effort
Setup B: one agent, maximum effort
Setup C: coordinator plus 3 workers, medium effort
Record per task: total tokens, dollars, wall-clock minutes, pass or fail
Then compute dollars per passing task, not dollars per run. If Setup C is faster but costs 3x per pass, you have a speed-versus-money decision to make explicitly. If Setup B matches Setup C's pass rate at a fraction of the cost, you have your answer. Watch also for variance: with only a handful of tasks, a 7-point gap can be noise, which is why Vals reported statistical significance rather than raw deltas.
Central ring routing three parcels to separate workers, showing delegation across multiple agents
Limits of the evidence
This is one benchmark, two models, and four comparisons. Several caveats apply:
- One task type. Building a web app end to end is tightly coupled work. Tasks with natural parallelism, such as broad research, could look very different.
- Secondhand reporting. The team figures and Anthropic's scaling table come through The Decoder's article. Check Vals AI's and Anthropic's primary publications as they appear.
- Harness dependence. How the team is orchestrated, including who plans, who reviews, and how conflicts are resolved, changes the cost. A better coordination design could shift the result.
- Fast-moving models. Both models are current as of October 2026. New releases and cheaper tiers change the cost side quickly.
The honest conclusion is narrower than the headline: for a coupled build task at high reasoning effort, teams did not pay for themselves. It is not a proof that multi-agent systems never help.
What to watch next
Two things will tell us whether this is a durable lesson. First, whether Vals AI publishes the full experiment, including token breakdowns and the coordination design. Second, whether labs ship cheaper coordination primitives, such as shared caches across agents, that cut the duplicated-context cost. Until then, the practical stance is the one in our Opus 5.5 versus Sonnet 5.5 comparison: same family, different bill, so measure before you scale.
Related reading
- Agent Teams Are the New Org Chart. The Web Is Starting to Lock the Door.
- What a Claude Code Task Costs on Opus 5.5
- Claude Opus 5.5 vs Sonnet 5.5: Same Family, Different Bill
- Claude Haiku 5.5 Launches at $0.10 per Million Input Tokens
- Hermes Agent Manual Subagent Control
- Paperclip: Running a Company Made of AI Agents
- Sources: The Decoder report and the Vals AI Vibe Code Bench page
Figures and model names are accurate as of October 11, 2026. Team results are as reported by The Decoder and may change as Vals AI and Anthropic publish more detail.
