x.com/hamandcheese/status/2083241471101247722
1 correction found
METR's inability to evaluate model autonomy beyond 13 hours
METR has not said its evaluations stop at 13 hours. In 2026, METR said its TH1.1 suite becomes unreliable above about 16 hours, and it also published a 13h25m worst-case time-horizon estimate itself.
Full reasoning
METR's own published materials contradict the specific 13 hours cutoff here.
- In METR's Frontier Risk Report (May 19, 2026), Table 1 says: "The TH 1.1 suite can't reliably measure time horizons above 16 hours" and adds that the most capable shared model had a 50% time horizon point estimate between 16 and 20 hours.
- In METR's earlier evaluation of OpenAI GPT-5.1-Codex-Max (Nov. 19, 2025), METR wrote that its extrapolation produced "a worst-case 50% time-horizon estimate of 13 hours and 25 minutes by April 2026."
So METR was not claiming an inability to evaluate "beyond 13 hours." Its own publications both (1) report a 13h25m estimate and (2) place the benchmark's current reliability ceiling at roughly 16 hours, not 13.
This doesn't settle the broader argument about AI progress, but the numeric description of METR's evaluation limit is inaccurate.
2 sources
- Frontier Risk Report (February to March 2026) - METR
"The TH 1.1 suite can't reliably measure time horizons above 16 hours, but the 50% time horizon point estimate of the most capable shared model was between 16 and 20 hours"
- Details about METR's evaluation of OpenAI GPT-5.1-Codex-Max
"With this, we arrived at a worst-case 50% time-horizon estimate of 13 hours and 25 minutes by April 2026."