On this page
Last Updated: 11 October 2026
Can 1,000 AI agents really beat one? On broad, verifiable work, the evidence says yes by a wide margin. Anthropic's dynamic workflows consistently caught 66 of 70 bugs hidden in a 116,000-line codebase where a single agent caught at most 27, and a real-world migration powered by the same pattern moved 750,000 lines of Rust in 11 days. But multi-agent fan-out burns about 15x the tokens of a normal chat, and it only pays off when the work is parallel, verifiable, and costly to get wrong.
Key Takeaways
Anthropic's dynamic workflows let a lead Claude agent write an orchestration program that starts up to 1,000 sub-agents per run (about 64 concurrent), plans, executes and verifies in phases, then merges one answer.
In Anthropic's bug-hunt test, a single agent caught between 14 and 27 of 70 hidden bugs in a 116,000-line codebase, while the dynamic workflow consistently caught 66.
Fan-out only pays off when work is parallel by nature, verifiable and costly to get wrong: think whole-repo bug hunts, migrations and document audits, not sequential refactors or nuanced writing.
Swarm skepticism has merit: Eric Provencher, an engineer on OpenAI's Codex team, calls agent swarms a massive waste of tokens, citing a project where 1,393 agents burned 20,000 dollars on one Python refactor, and Anthropic's own research shows multi-agent systems use about 15x chat tokens.
A workflow run carries no separate fee: agent tokens bill at each model's normal rates against a session budget that caps spending, so cost per successful outcome, not agent count, is the metric to watch.
The same fan-out-verify-merge pattern runs on your own hardware: at Flowtivity our dual DGX Sparks run DeepSeek V4 Flash at roughly 60 tokens per second, so your verification loops never leak source code to a cloud endpoint.
What Are Claude's Dynamic Workflows?
Dynamic workflows turned generally available in Claude Code this week, according to Anthropic's October 2026 announcement. A lead agent plans your prompt into subtasks, writes a small orchestration program, and the platform runs it in the background across parallel sub-agents in phases, then merges the results into one coordinated answer.
The scale numbers need careful reading: 1,000 agents is the lifetime total per run, with about 64 working at the same time. Managed Agents defaults to 64 threads while Claude Code defaults to 16, tunable from 1 to 256, runs live 24 hours by default, and a session can carry up to 10 open runs. Anthropic's docs showcase a business-shaped example: reviewing 300 contracts for change-of-control clauses, where a read phase fans the contracts out to parallel agents and a merge phase combines their findings into one answer. Managed Agents caps a run's spending with a session budget: hit it, and every open run pauses until you raise or remove the budget.
Architecturally this is a shift from delegation to orchestration. With plain delegation the lead agent calls one subagent at a time and every result flows through its context window, making the lead the bottleneck. A workflow moves fan-out, phasing and aggregation into a background program, so the coordination no longer lives inside one agent's attention.

Proof It Works: The 66-of-70 Bug-Hunt Result
Anthropic's team hid 70 bugs in a 116,000-line codebase. One agent hunting alone caught between 14 and 27 bugs per run: the model samples and skims because no single agent can hold a repository that size in its attention budget. The dynamic workflow consistently caught 66, because the program slices the code into sections, gives each slice a fresh context, and runs a verification phase over every candidate finding before it lands in the report.

Two caveats before you quote the headline number to your board. It is a vendor-run demonstration on a seeded-bug benchmark, which is friendlier to parallel scanning than a messy real-world review. And The Decoder notes it is unclear whether the gains carry over to task types other than code.
The Showcase Run: Bun Rewritten From Zig to Rust
The clearest large-scale proof sits outside Anthropic entirely. According to Bun's project blog, creator Jarred Sumner ran about 50 dynamic workflows continuously over 11 days to port the Bun JavaScript runtime from Zig to Rust: roughly 750,000 lines of Rust, 99.8 percent of the existing test suite passing, from first commit to merge.
Each workflow was a loop: one mapped the right Rust lifetime for every struct field in the Zig codebase, the next had hundreds of agents writing .rs files as behavior-identical ports of their .zig counterparts with two reviewers per file, and a fix loop kept driving the build and test suite until both ran clean. An overnight workflow then opened one pull request per unnecessary data copy for final human review. For a business reader, the takeaway is not "rewrite your Bun": it is that an 11-day, machine-verified migration is now a real pattern, not a demo.
Is It Worth the Tokens? The Economics of Agent Swarms
Every agent in a run consumes tokens, and skeptics inside the industry are blunt about it. According to The Decoder's September 2026 report, Eric Provencher, an engineer on OpenAI's Codex team, calls agent swarms "a massive waste of tokens with zero quality gain" and describes a coordination tax: agents do not trust each other's work, so they re-check one another and spend multiplies. His example: a project that burned 20,000 dollars in tokens across 1,393 agents on a single Python refactoring.
Anthropic's own research supports the caution. In June 2025 Anthropic reported that its multi-agent research system beat a single-agent baseline by 90.2 percent on internal research evals while using about 15 times more tokens than a chat interaction, according to Anthropic's multi-agent research blog.

So when does the 15x tax pay for itself? Score the job on three questions.
| Great fit: parallel, verifiable, high miss cost | Poor fit: sequential, subjective, low miss cost |
|---|---|
| Whole-repo bug hunts, dead-code discovery, security and profiling audits | Tightly coupled refactors where each step depends on the last |
| Migrations and ports touching hundreds of files | Nuanced one-shot writing where every paragraph builds on the last |
| Reviewing hundreds of documents against a fixed checklist | Quick lookups a single prompt can finish |
| High-stakes plans stress-tested by adversarial reviewers | Work you cannot verify cheaply |
Running the Swarm on Your Own Hardware: The Local-First Play
You do not need Anthropic's cloud to run this pattern. The architecture, lead agent plans, workers execute, independent verifiers check, is generic, and open-source stacks such as CAMEL or OWL implement the same shape against any OpenAI-compatible endpoint.
That is exactly what we run at Flowtivity: our dual DGX Spark nodes serve DeepSeek V4 Flash through an OpenAI-compatible API at roughly 60 tokens per second on local hardware. A wide fan-out like Anthropic's is not yet practical locally, so we split the loop instead: cloud agents do the wide, parallel first pass, local models run the verification and re-check loops so source code and customer data stop at our network edge.

Three guardrails matter more than raw agent count:
- Verify locally. Workers can run anywhere cheap, but the referee should run on your network so the final judgment and your source never leave your control.
- Keep credentials local. Agents are disposable, key material is not: vault your API keys on your own hardware, per our agent-credentials guide, so no workflow run can exfiltrate them.
- Keep a human approval gate. Claude Code shows what a workflow is about to run and asks before the first trigger, and org admins can disable workflows entirely. Copy that pattern in your own stack: no unattended fan-out on untested tasks.
The bottom line for business owners: dynamic workflows make multi-agent work a first-class product feature, but the winning pattern is not "always 1,000 agents", it is match the fan-out to the task and keep verification where your data lives.

One email a month, no noise
Practical AI notes for Australian businesses. Unsubscribe anytime.

