All corrections
Wikipedia July 1, 2026 at 01:59 AM

en.wikipedia.org/wiki/Stochastic_parrot

2 corrections found

1
Claim
achieved 93% accuracy on the MATH benchmark of high-school Olympiad problems
Correction

This appears to mix up two different math benchmarks. OpenAI reported GPT-4 at 92.0% on GSM8K, while a later benchmark paper lists GPT-4’s official MATH score as 42.5%, not 93%.

Full reasoning

The numerical claim does not match the benchmark it names.

OpenAI's GPT-4 Technical Report reports 92.0% on GSM8K, which it describes as "Grade-school mathematics questions". That is a different benchmark from MATH.

A later benchmarking paper comparing its own evaluations against official GPT-4 scores lists MATH (4-shot CoT) with an official score of 42.5 for GPT-4. In other words, the commonly cited ~92% figure belongs to GSM8K, not to MATH.

So this sentence appears to conflate:

  • GSM8K ≈ 92% (grade-school math), with
  • MATH = 42.5% (a separate benchmark).

That makes the statement "93% accuracy on the MATH benchmark" incorrect.

2 sources
2
Claim
A 2024 Scientific American investigation described a closed Berkeley workshop where state-of-the-art models solved novel tier-4 mathematics problems and produced coherent proofs
Correction

The Scientific American article matching this description was published in 2025, not 2024, and it reports on a workshop held in spring 2025.

Full reasoning

The description points to Scientific American's article "Inside the Secret Meeting Where Mathematicians Struggled to Outsmart AI". But that article is dated June 6, 2025, not 2024.

The article also says the workshop itself happened "in the spring of 2025" in Berkeley, California. So calling it a 2024 Scientific American investigation is chronologically impossible.

That makes the date in this claim incorrect.

1 source
Model: OPENAI_GPT_5 Prompt: v1.16.0