All corrections
X June 25, 2026 at 01:03 PM

x.com/parvmahajan0/status/2069806601322811489?s=20

1 correction found

1
Claim
they're only giving the models 1M tokens.
Correction

Anthropic’s own materials describe multiple Mythos 5 biology eval conditions with 2 million tokens, not just 1 million. The screenshot reflects one condition, not all chem/bio automated evals.

Full reasoning

The screenshot in the post is real, but it is not representative of all of Anthropic’s chem/bio automated evaluations for Mythos 5.

Anthropic’s official system-card materials describe at least two biology-related automated evaluation conditions for Mythos 5 that used 2 million tokens, not 1 million:

  1. In the sequence-design evaluation, Anthropic says it "additionally evaluated an in-context iteration condition" where each model was given "a 24h tool-call budget and a two million token budget."
  2. In a separate AAV capsid prediction evaluation, Anthropic says models were given "a 24-hour tool-call wall-clock budget, a single H100 GPU, [and] a two-million-token allowance."

So the blanket claim that Anthropic was "only giving the models 1M tokens" in the chem/bio automated evals is incorrect. The 1M-token setup shown in the image was one evaluation condition, but other Mythos 5 biology evals used larger budgets.

2 sources
  • Claude Opus 4.8 System Card

    We additionally evaluated an in-context iteration condition... given access to a 24h tool-call budget and a two million token budget.

  • Claude Opus 4.8 System Card

    Models were given the wild type capsid sequence, and given a 24-hour tool-call wall-clock budget, a single H100 GPU, a two-million-token allowance, standard ML libraries, and no internet access.

Model: OPENAI_GPT_5 Prompt: v1.16.0