# Fact-check: PASSED

Source: https://research.dreamlab.institute/ai-programming/  
Rounds: 2  

Cited sources (claims may quote these as well as the tracker page):

1. [SWE-Bench Pro V2](https://labs.scale.com/leaderboard/swe_bench_pro_public_v2?tab=full)
2. [SWE-Bench Pro V2: A Cleaner, Harder-to-Game Leaderboard | Scale Labs](https://model-eval-leaderboard.vercel.app/blog/swe-bench-pro-v2)

## Round 1: 2 issue(s), sent back to the writer

- line 6 "So check which version a number comes from. Treat V2 as the cleaner reference once models are re-run.": Treat V2 as the cleaner reference once models are re-run. — not supported by the page (suggested: "So check which version a number comes from. Scale says V2's numbers can now be trusted, but near-ceiling public scores may still reflect recall.")
- line 7 "Which means trust the newest clean leaderboard, and expect today's numbers to shift.": Trust the newest clean leaderboard, and expect today's numbers to shift. — not supported by the page (suggested: "Which means weigh the private set most, since Scale calls it the only clean measurement of memorisation.")

## Final claims

| line | claim | verdict | page quote |
|---|---|---|---|
| 1 | Published coding agent scores are inflated, and Scale changed the benchmark rules with a new release | supported | means most published scores are inflated |
| 2 | SWE-Bench Pro V2 drops 89 invalid tasks | supported | We dropped 89 tasks our review found invalid. |
| 2 | V2 locks the network so agents can't look up fixes | supported | The agent phase now reaches only the model endpoint, with web tools disabled. |
| 2 | V2 re-grades every agent answer on a clean copy | supported | Every agent diff is re-graded on a pristine image, and we publish both grades. |
| 3 | SWE-bench Verified is contaminated, so models likely saw its tests in training | supported | SWE-bench Verified, is now widely considered contaminated; OpenAI itself stopped evaluating on it. |
| 4 | Claude Opus 4.5 dropped from 80.9% on Verified to 45.89% on the original Pro | supported | Claude Opus 4.5 dropped from 80.9% on Verified to 45.89% on Pro. |
| 5 | Bito's 60.8% on SWE-Bench Pro is a vendor claim, unverified, and predates V2 | supported | Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (vendor-claim) |
| 6 | Scale says V2's numbers can now be trusted | supported | What changed is that the numbers can now be trusted |
| 7 | Scale calls the private set the only clean measurement of memorisation | supported | The private set is the only clean measurement of that, and it is why we maintain one. |
| 9 | Scale's cleaner coding leaderboard | supported | SWE-Bench Pro V2: A Cleaner, Harder-to-Game Leaderboard |
