All corrections
X August 1, 2026 at 12:06 AM

x.com/hamandcheese/status/2083241471101247722

1 correction found

1
Claim
METR's inability to evaluate model autonomy beyond 13 hours
Correction

METR has not said its evaluations stop at 13 hours. In 2026, METR said its TH1.1 suite becomes unreliable above about 16 hours, and it also published a 13h25m worst-case time-horizon estimate itself.

Full reasoning

METR's own published materials contradict the specific 13 hours cutoff here.

  • In METR's Frontier Risk Report (May 19, 2026), Table 1 says: "The TH 1.1 suite can't reliably measure time horizons above 16 hours" and adds that the most capable shared model had a 50% time horizon point estimate between 16 and 20 hours.
  • In METR's earlier evaluation of OpenAI GPT-5.1-Codex-Max (Nov. 19, 2025), METR wrote that its extrapolation produced "a worst-case 50% time-horizon estimate of 13 hours and 25 minutes by April 2026."

So METR was not claiming an inability to evaluate "beyond 13 hours." Its own publications both (1) report a 13h25m estimate and (2) place the benchmark's current reliability ceiling at roughly 16 hours, not 13.

This doesn't settle the broader argument about AI progress, but the numeric description of METR's evaluation limit is inaccurate.

2 sources
Model: OPENAI_GPT_5 Prompt: v1.16.0