x.com/MatthewJBar/status/2051758741968879935
1 correction found
Except for a somewhat impressive score on SWE-bench Pro, Claude Mythos is nearly on-trend, at most a few months ahead.
Public benchmark tables do not show Mythos’ lead is limited to SWE-bench Pro. Anthropic reports Mythos also leads GPT-5.5 by large margins on Humanity’s Last Exam, a general-capability benchmark.
Full reasoning
Anthropic’s and OpenAI’s own published benchmark tables contradict the idea that SWE-bench Pro is the only notable place Mythos stands out.
On Anthropic’s public Project Glasswing page, Mythos scores 56.8% on Humanity’s Last Exam without tools and 64.7% with tools. On OpenAI’s public GPT-5.5 launch page, GPT-5.5 scores 41.4% without tools and 52.2% with tools on the same benchmark.
That means Mythos’ publicly reported lead over GPT-5.5 is not just on SWE-bench Pro:
- Humanity’s Last Exam (without tools): 56.8% vs 41.4% (+15.4 points)
- Humanity’s Last Exam (with tools): 64.7% vs 52.2% (+12.5 points)
- SWE-bench Pro: 77.8% vs 58.6% (+19.2 points)
It is true that some public benchmarks are close or tied—for example, Terminal-Bench 2.0 is roughly even (Mythos 82.0% vs GPT-5.5 82.7%). But because there is at least one other major general benchmark with a double-digit Mythos lead, the post’s claim that Mythos is only notably ahead on SWE-bench Pro is contradicted by the published benchmark data.
2 sources
- Project Glasswing: Securing critical software for the AI era | Anthropic
Humanity’s Last Exam — Mythos Preview without tools 56.8%; Mythos Preview with tools 64.7%. SWE-bench Pro — Mythos Preview 77.8%. Terminal-Bench 2.0 — Mythos Preview 82.0%.
- Presentamos GPT-5.5 | OpenAI
Humanity's Last Exam (sin herramientas) 41.4%; Humanity's Last Exam (con herramientas) 52.2%; SWE-Bench Pro 58.6%; Terminal-Bench 2.0 82.7%.