WEBVTT

00:00:00.000 --> 00:00:04.070
<v Wren>Coding agent scores are inflated, and Scale just changed the rulebook.

00:00:04.230 --> 00:00:16.075
<v Ash>Yeah, SWE-Bench Pro V2. It drops 89 invalid tasks, locks the network so agents can't look up fixes, and re-grades every answer on a clean copy.

00:00:16.235 --> 00:00:22.805
<v Wren>Compared to Verified, though? That one's already contaminated, so models probably saw those tests in training.

00:00:22.965 --> 00:00:34.360
<v Ash>Claude Opus 4.5 fell from 80.9% on Verified to 45.89% on the original Pro. V2 tries to keep that honesty.

00:00:34.520 --> 00:00:41.390
<v Wren>And Bito's 60.8% on Pro? That's just the vendor talking, unverified, and from the old version.

00:00:41.550 --> 00:00:47.845
<v Ash>So check which version a number comes from. Scale says V2's numbers can now be trusted.

00:00:48.005 --> 00:00:53.825
<v Wren>Which means weigh the private set most, since Scale calls it the only clean measurement of memorisation.

00:00:53.985 --> 00:00:56.530
<v Ash>Full tracker's at DreamLab Research.
