x.com/parvmahajan0/status/2069806601322811489?s=20
1 correction found
they're only giving the models 1M tokens.
Anthropic’s own materials describe multiple Mythos 5 biology eval conditions with 2 million tokens, not just 1 million. The screenshot reflects one condition, not all chem/bio automated evals.
Full reasoning
The screenshot in the post is real, but it is not representative of all of Anthropic’s chem/bio automated evaluations for Mythos 5.
Anthropic’s official system-card materials describe at least two biology-related automated evaluation conditions for Mythos 5 that used 2 million tokens, not 1 million:
- In the sequence-design evaluation, Anthropic says it "additionally evaluated an in-context iteration condition" where each model was given "a 24h tool-call budget and a two million token budget."
- In a separate AAV capsid prediction evaluation, Anthropic says models were given "a 24-hour tool-call wall-clock budget, a single H100 GPU, [and] a two-million-token allowance."
So the blanket claim that Anthropic was "only giving the models 1M tokens" in the chem/bio automated evals is incorrect. The 1M-token setup shown in the image was one evaluation condition, but other Mythos 5 biology evals used larger budgets.
2 sources
- Claude Opus 4.8 System Card
We additionally evaluated an in-context iteration condition... given access to a 24h tool-call budget and a two million token budget.
- Claude Opus 4.8 System Card
Models were given the wild type capsid sequence, and given a 24-hour tool-call wall-clock budget, a single H100 GPU, a two-million-token allowance, standard ML libraries, and no internet access.