WEBVTT

00:00:00.000 --> 00:00:03.695
<v Wren>Two AI agents together do worse than one working alone.

00:00:03.855 --> 00:00:16.800
<v Ash>CooperBench now has its own website. GPT-5 and Claude Sonnet 4.5 hit only 25% success as a pair, roughly 50% below one agent doing both tasks.

00:00:16.960 --> 00:00:25.030
<v Wren>Wait, we'd logged that 50% already, but secondhand, via Zylos Research on Sep 14. Basically a press summary.

00:00:25.190 --> 00:00:34.610
<v Ash>Right, and our planner verdict rested on a boss agent handing out tasks, plus the Hugging Face swarm postmortem. No direct team benchmark.

00:00:34.770 --> 00:00:44.315
<v Wren>So "more cheap workers, more output" now meets a primary source. Treat parallel delegation as a cost to justify, and keep the verification gate first.

00:00:44.475 --> 00:00:47.020
<v Ash>Full tracker's at DreamLab Research.
