explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: what the study found
  • What Vibe Code Bench actually tests
  • The four comparisons, in plain language
  • Why teams cost so much more
  • Anthropic's own scaling numbers say the same thing
  • Noam Brown: you are buying speed
  • What this means for what you build or pay
  • A simple way to test this on your own work
  • Limits of the evidence
  • What to watch next
  • Related reading
← Back to blog

explainx / blog

AI Agent Teams Cost Up to 5.1x More Than One Agent for Barely Better Results

AI Agents, Multi-Agent Systems, Token Economics, Benchmarks, Claude Opus 5.5

Part of AI Agents

Vals AI found agent teams cost 1.8x to 5.1x more than a solo agent, and only 1 of 4 matchups improved. What it means before you scale out your agents.

Oct 11, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
AI Agent Teams Cost Up to 5.1x More Than One Agent for Barely Better Results

Short answer: adding more agents to a job mostly makes it faster and more expensive, not better. Evals company Vals AI tested GPT-6 Sol and Claude Opus 5.5 on its Vibe Code Bench, once as solo agents and once as teams, at medium and maximum reasoning effort. According to The Decoder's report, the teams cost between 1.8x and 5.1x more, and only one of four team-versus-solo comparisons showed a statistically significant improvement.

That lands in the middle of a year when "agent teams" became the default pitch for coding tools, as we argued in Agent Teams Are the New Org Chart. This post walks through what the Vals result says, what Anthropic's scaling numbers and Noam Brown's comments add, and how to decide between one agent and many.

Weekly digest3.6k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: what the study found

table · 2 cols
QuestionAnswer
Who ran it?Vals AI, an evals company that publishes Vibe Code Bench
Models testedGPT-6 Sol and Claude Opus 5.5
SetupsSingle agent versus team, each at medium and maximum reasoning effort
Cost of teams1.8x to 5.1x more than the solo agent
Quality gainOne of four comparisons significant: GPT-6 Sol at medium effort, +7.3 points
At maximum effortNeither model gained anything real from a team
TakeawayExtra agents mostly buy speed; at full compute they add cost without score

One caveat up front: I could not find Vals AI's own write-up of the team experiment. The figures above come from The Decoder's summary, which includes a Vals chart plotting cost per app against score, with arrows from each model's single agent to its team. The benchmark itself is documented on the Vals leaderboard page. Treat the team numbers as reported secondhand until Vals publishes the full methodology.

What Vibe Code Bench actually tests

Vibe Code Bench asks a model to build a complete web application from a short natural-language spec, then scores the result by driving the app with an automated agent that clicks through it. The specs cover things like sign-up and login, posting messages, following users, search, and a discovery page. Each spec is paired with 20 to 60 automated tests.

Models work inside a modified OpenHands harness with a terminal, a browser, and sandboxed services for the database, storage, payments, and email. They can work for up to five hours or 1,000 turns per app. The leaderboard, updated October 7, 2026, shows a frontier clustered around 90 percent: Claude Sonnet 5.5 at 92.39 percent, Claude Opus 5.5 at 90.29 percent, and GPT-6 Sol at 87.82 percent for single-agent runs.

Those leaderboard cost figures give useful scale. On the single-agent board, Opus 5.5 costs $57.92 per test and GPT-6 Sol $26.36. Multiply by a team factor of 1.8x to 5.1x and one app build moves from tens of dollars to potentially a few hundred. The leaderboard numbers are not the team experiment's numbers, so use them only for a sense of magnitude.

This benchmark is a good stress test for teams because an app is one coherent artifact. Every file depends on the schema, the auth flow, and the routing decisions made earlier. That is the kind of work where parallel agents tend to step on each other.

The four comparisons, in plain language

Vals paired each model with each reasoning level, giving four team-versus-solo matchups:

  1. GPT-6 Sol, medium effort. The only significant win: the team scored 7.3 points higher.
  2. GPT-6 Sol, maximum effort. No real advantage from the team.
  3. Claude Opus 5.5, medium effort. No significant improvement.
  4. Claude Opus 5.5, maximum effort. No real advantage.

The pattern is intuitive. A team partly substitutes for thinking time. When a model is already running at medium effort, splitting the work across several agents can recover some of what deeper reasoning would have bought. When it is already at maximum effort, the single agent has used its whole budget on the problem, and the team adds coordination without adding capability.

That reading fits the cost story too. If you are paying for maximum reasoning on one agent and then multiplying the bill by up to five for a team, you are paying twice for the same lever.

Why teams cost so much more

The study does not break down the token spend, but the mechanics are well understood in agent engineering:

  • Duplicated context. Each agent needs enough of the spec and the current codebase state to act. That context is paid for once per agent, on every turn.
  • Handoff messages. Planner-to-worker and worker-to-reviewer messages are output tokens that a solo agent never has to write.
  • Redundant exploration. Two agents reading the same files or running the same failing command is pure waste.
  • Merge and repair. When parallel edits conflict, someone has to reconcile them, usually with another long turn.

OpenAI developer Eric Provencher made a related warning, per The Decoder, calling agent swarms most likely wasted money because coordination between agents breaks down. He called it the coordination tax.

Anthropic's own scaling numbers say the same thing

The Decoder also pulls in Anthropic's data from two Opus 5.5 tests. In both, more agents reached a given level faster, but returns flattened.

table · 5 cols
Task with Opus 5.51 agent10 agents30 agents100 agents
Knowledge base0.530.700.710.74
Lean theorem proving0.390.660.660.68

Going from one agent to ten is a big jump on both tasks. Going from ten to a hundred moves the score by only a few hundredths, while the token bill keeps growing. After 24 hours, the 100-agent setup was only slightly ahead of the 10-agent one, per the report. In separate ProgramBench tests, speed gains came with higher token usage.

The Decoder adds that Claude Fable 5.1 showed stronger gains on Lean theorem proving above ten agents but still scored below Opus 5.5 across all tests, and that on the knowledge base task its score dipped slightly from 30 to 100 agents. I could not locate Anthropic's original write-up of these runs in the time available, so cite them as reported by The Decoder.

Three robotic grippers placing different shapes into one shared tray, a picture of agents sharing one workspaceThree robotic grippers placing different shapes into one shared tray, a picture of agents sharing one workspace

Noam Brown: you are buying speed

The most useful framing comes from OpenAI researcher Noam Brown. On the Dwarkesh Podcast, as reported by The Decoder, he said multi-agent systems mainly buy speed rather than quality: four agents solved tasks twice as fast but also cost twice as much, and at 16 agents the pattern held but grew slightly less efficient.

He also stressed that the effect depends on the task. Web research and math parallelize well. Writing a novel does not, and throwing 10,000 agents at a novel would be as pointless as throwing 10,000 people at it. Brown acknowledged that scaling to very large agent counts is largely unexplored because the costs are simply too high.

That is a clean rule of thumb. Ask whether your task decomposes into independent chunks with a cheap merge step. If yes, a team is a speed tool. If no, it is a cost multiplier.

What this means for what you build or pay

If you run coding agents or build agent products, three practical consequences follow.

1. Default to one strong agent. Before adding agents, raise reasoning effort or improve the prompt and tools. Our breakdown of what a Claude Code task costs on Opus 5.5 shows how cache hits, turn count, and effort move the bill more than the sticker price, and the Vals result suggests team size belongs on that list too.

2. Pick the right model before the right topology. On Vibe Code Bench the cheapest strong results come from model choice, not agent count. Claude Haiku 5.5 scores 90.44 percent on the leaderboard at $6.07 per test, and our small-model comparison covers the cheaper tier. Compare that with multiplying a $57.92 Opus 5.5 run by up to 5.1x.

3. Use teams where the work splits cleanly. Research sweeps, test generation across independent modules, and batch migrations of unrelated files are good candidates. A single coherent feature is not. If you want an interface for supervising helpers, see how Hermes Agent handles manual subagent control, and for company-style orchestration see Paperclip.

A simple way to test this on your own work

You do not need Vals AI's harness to check whether a team earns its keep. Run the same ten tasks three ways and compare cost and pass rate:

text
Setup A: one agent, medium effort
Setup B: one agent, maximum effort
Setup C: coordinator plus 3 workers, medium effort

Record per task: total tokens, dollars, wall-clock minutes, pass or fail

Then compute dollars per passing task, not dollars per run. If Setup C is faster but costs 3x per pass, you have a speed-versus-money decision to make explicitly. If Setup B matches Setup C's pass rate at a fraction of the cost, you have your answer. Watch also for variance: with only a handful of tasks, a 7-point gap can be noise, which is why Vals reported statistical significance rather than raw deltas.

Central ring routing three parcels to separate workers, showing delegation across multiple agentsCentral ring routing three parcels to separate workers, showing delegation across multiple agents

Limits of the evidence

This is one benchmark, two models, and four comparisons. Several caveats apply:

  • One task type. Building a web app end to end is tightly coupled work. Tasks with natural parallelism, such as broad research, could look very different.
  • Secondhand reporting. The team figures and Anthropic's scaling table come through The Decoder's article. Check Vals AI's and Anthropic's primary publications as they appear.
  • Harness dependence. How the team is orchestrated, including who plans, who reviews, and how conflicts are resolved, changes the cost. A better coordination design could shift the result.
  • Fast-moving models. Both models are current as of October 2026. New releases and cheaper tiers change the cost side quickly.

The honest conclusion is narrower than the headline: for a coupled build task at high reasoning effort, teams did not pay for themselves. It is not a proof that multi-agent systems never help.

What to watch next

Two things will tell us whether this is a durable lesson. First, whether Vals AI publishes the full experiment, including token breakdowns and the coordination design. Second, whether labs ship cheaper coordination primitives, such as shared caches across agents, that cut the duplicated-context cost. Until then, the practical stance is the one in our Opus 5.5 versus Sonnet 5.5 comparison: same family, different bill, so measure before you scale.

Related reading

  • Agent Teams Are the New Org Chart. The Web Is Starting to Lock the Door.
  • What a Claude Code Task Costs on Opus 5.5
  • Claude Opus 5.5 vs Sonnet 5.5: Same Family, Different Bill
  • Claude Haiku 5.5 Launches at $0.10 per Million Input Tokens
  • Hermes Agent Manual Subagent Control
  • Paperclip: Running a Company Made of AI Agents
  • Sources: The Decoder report and the Vals AI Vibe Code Bench page

Figures and model names are accurate as of October 11, 2026. Team results are as reported by The Decoder and may change as Vals AI and Anthropic publish more detail.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 10, 2026

Pine Computer Scores 78.3% on SaaS-Bench, but Finishes Fewer Tasks: Claimed vs Verified

Pine AI launched Pine Computer, a cloud computer built for AI agents, with the top checkpoint score on UniPat AI SaaS-Bench v1.1. The same table shows it resolving fewer whole tasks than Opus 5 with Claude Code and GPT-5.6 Sol with Codex. This post separates the claims from what can be checked.

Oct 9, 2026

Claude Managed Agents Dynamic Workflows: Up to 1,000 Parallel Agents

Anthropic has added dynamic workflows to Claude Managed Agents. A lead agent writes a workflow program, and the server runs up to 1,000 agents in parallel. Here is how it works, what the bug-hunt test showed, how to turn it on, and what it will cost you in tokens.

Oct 7, 2026

Grok Bot Will Route to Claude Opus 5.5, Midjourney and Suno: What Musk Announced and What Is Missing

On October 7, 2026, Elon Musk said that going forward SpaceX will use the best back end model for any given task in Grok Bot, naming Claude Opus 5.5, Midjourney, Suno and other leading APIs. The post passed 3.8 million views. It is a striking admission for a lab that builds its own models. Here is what was said, the open questions on routing, data and cost, and what builders should take from it.