# Fact-check: PASSED

Source: https://research.dreamlab.institute/software-factories/  
Rounds: 1  

Cited sources (claims may quote these as well as the tracker page):

1. [CooperBench: Benchmarking Agent Teams | Why Coding Agents Cannot be Your Teammates Yet](https://cooperbench.com/index.html)
2. [Leaderboard · CooperBench](https://cooperbench.com/leaderboard.html)
3. [The Curious Case of Miscoordination](https://cooperbench.com/blog/the-curious-case-of-miscoordination)

## Final claims

| line | claim | verdict | page quote |
|---|---|---|---|
| 1 | Two AI agents together do worse than one working alone. | supported | Agents perform worse together than alone |
| 2 | CooperBench now has its own website. | supported | CooperBench benchmark for agent teams published, measuring coordination failure rates. |
| 2 | GPT-5 and Claude Sonnet 4.5 hit only 25% success as a pair, roughly 50% below one agent doing both tasks. | supported | GPT-5 and Claude Sonnet 4.5 achieve only 25% success with two-agent cooperation, roughly 50% lower than |
| 3 | The 50% figure was already logged on Sep 14 via Zylos Research, a secondary press summary. | supported | CooperBench benchmark finds multi-agent collaboration success rates roughly 50% lower than solo work. (secondary) |
| 4 | The planner verdict rested on a central agent delegating tasks plus the Hugging Face swarm postmortem, with no direct team benchmark. | supported | with a central agent decomposing tasks and delegating to specialized workers. However, production postmortems like the Hugging Face breach investigation |
| 5 | The 'more cheap workers' premise now meets a primary source, and the advice is to treat coordination as a cost and keep the verification gate first. | supported | The evidence says coordination is a liability, not an asset, and the gate is everything. |
| 7 | Title: Two Agents Do Worse Than One | supported | Agents perform worse together than alone |
