arachnemag.substack.com/p/ais-reliability-gap?r=18kjq3&utm_campaign=post-expande...
2 corrections found
The best model, GPT 5.\[5\] with high reasoning, achieves only ~\[37\]% pass@1 \[and ~21% pass^4\].
The cited τ-Knowledge page does not report GPT-5.5 at ~37%/~21%. It currently says the best model is GPT-5.2 high reasoning at 25.52% pass^1 and 13.40% pass^4.
Full reasoning
The benchmark page linked in the article contradicts this sentence.
On the official τ-Knowledge / τ-Banking page, Sierra says:
- "The best model, GPT-5.2 with high reasoning, achieves only ~26% pass^1"
- In the results section: "GPT-5.2 with high reasoning ... achieves only 25.52% pass^1" and "drops to just 13.40% pass^4"
So the linked source does not support the article's claim that the best model is GPT-5.5 at roughly 37% pass@1 and 21% pass^4. The model name and both headline percentages are different on the source page.
Because the article presents this as the benchmark's current headline result, readers are likely to be misled about both the best-performing model and the size of the result.
1 source
- τ-knowledge | τ-bench
The best model, GPT-5.2 with high reasoning, achieves only ~26% pass^1 ... The best-performing configuration - GPT-5.2 with high reasoning - achieves only 25.52% pass^1. Performance degrades sharply with increasing k: GPT-5.2-High drops to just 13.40% pass^4.
The current top scorer is GPT 5.5 with a pass@3 of ~29%.
This is no longer correct. Scale's live HiL-Bench leaderboard currently shows Claude Fable 5 leading at 56.33 Pass@3, while GPT-5.5 is listed at 39.67.
Full reasoning
The article states, in present tense, that GPT-5.5 is the current HiL-Bench leader at roughly 29% pass@3. The official HiL-Bench leaderboard now contradicts that.
In the current leaderboard text from Scale Labs, the Performance Comparison section lists:
- Claude Fable 5 — 56.33 ± 5.50
- GLM 5.2 — 43.67 ± 5.65
- Claude Opus 4.7 — 41.67 ± 5.65
- GPT-5.5 — 39.67 ± 5.63
So GPT-5.5 is not the current top scorer, and its listed score is about 39.67, not ~29%.
This appears to be a stale leaderboard claim rather than a minor rounding issue: both the leader and the reported score are different on the benchmark's current official page.
1 source
- HiL-Bench (Human-in-Loop Benchmark)
Performance Comparison ... Claude Opus 4.7 41.67 ± 5.65 ... GPT-5.5 39.67 ± 5.63 ... Claude Fable 5 NEW 56.33 ± 5.50 ... GLM 5.2 NEW 43.67 ± 5.65 ...