All corrections
X May 6, 2026 at 10:13 PM

x.com/MatthewJBar/status/2051758741968879935

1 correction found

1
Claim
Except for a somewhat impressive score on SWE-bench Pro, Claude Mythos is nearly on-trend, at most a few months ahead.
Correction

Public benchmark tables do not show Mythos’ lead is limited to SWE-bench Pro. Anthropic reports Mythos also leads GPT-5.5 by large margins on Humanity’s Last Exam, a general-capability benchmark.

Full reasoning

Anthropic’s and OpenAI’s own published benchmark tables contradict the idea that SWE-bench Pro is the only notable place Mythos stands out.

On Anthropic’s public Project Glasswing page, Mythos scores 56.8% on Humanity’s Last Exam without tools and 64.7% with tools. On OpenAI’s public GPT-5.5 launch page, GPT-5.5 scores 41.4% without tools and 52.2% with tools on the same benchmark.

That means Mythos’ publicly reported lead over GPT-5.5 is not just on SWE-bench Pro:

  • Humanity’s Last Exam (without tools): 56.8% vs 41.4% (+15.4 points)
  • Humanity’s Last Exam (with tools): 64.7% vs 52.2% (+12.5 points)
  • SWE-bench Pro: 77.8% vs 58.6% (+19.2 points)

It is true that some public benchmarks are close or tied—for example, Terminal-Bench 2.0 is roughly even (Mythos 82.0% vs GPT-5.5 82.7%). But because there is at least one other major general benchmark with a double-digit Mythos lead, the post’s claim that Mythos is only notably ahead on SWE-bench Pro is contradicted by the published benchmark data.

2 sources
Model: OPENAI_GPT_5 Prompt: v1.16.0