All corrections
1
Claim
The blog post states that the affected grader was used in 1.5% of the trajectories for GPT-5.4 Thinking.
Correction

This misstates OpenAI’s reported percentage. OpenAI says the relevant CoT-grading incident affected less than 0.6% of GPT-5.4 Thinking samples; the 1.5% figure was for GPT-5.4 mini, not GPT-5.4 Thinking.

Full reasoning

OpenAI’s post distinguishes between GPT-5.4 Thinking and GPT-5.4 mini for the “rewarding trajectory usefulness” incident.

Their published wording is:

  • “This affected less than 0.6% of GPT-5.4 Thinking samples”
  • “and less than 1.5% of GPT-5.4 mini samples.”

So the post does not state that the affected grader was used in 1.5% of GPT-5.4 Thinking trajectories. It states <0.6% for GPT-5.4 Thinking and <1.5% for GPT-5.4 mini.

Because this sentence swaps the GPT-5.4 Thinking figure with the GPT-5.4 mini figure, it overstates the reported prevalence for GPT-5.4 Thinking and also makes the following back-of-the-envelope calculation use the wrong input percentage.

1 source
Model: OPENAI_GPT_5 Prompt: v1.16.0