All corrections
X September 9, 2026 at 08:16 PM

x.com/OpenAI/status/2095595742975197690

1 correction found

1
Claim
sets a new state of the art for computer use, browsing, software engineering, cybersecurity, science, and professional work.
Correction

OpenAI's own launch benchmarks and independent benchmarking did not show GPT-6 Astra leading all of those domains. In the launch table, Astra trailed competing models on several software-engineering and science benchmarks, so the blanket 'state of the art' claim across every listed area is overstated.

Full reasoning

The claim says Astra sets a new state of the art across all of the listed domains. But OpenAI's own launch page includes benchmark results showing that Astra was not top-performing on several of them.

Examples from OpenAI's September 3, 2026 launch table:

  • Software engineering / coding:

    • On FrontierCode 1.1 Extended, Astra scored 64.5%, while Claude Fable 5 scored 64.9%.
    • On FrontierCode 1.1 Main, Astra scored 53.3%, while Claude Fable 5 scored 53.5% and Claude Opus 5 scored 53.4%.
    • On the Artificial Analysis Coding Agent Index v1.4, Astra scored 67.0, below Claude Opus 5 at 68.1 and Claude Fable 5 at 67.2.
  • Science / academic reasoning:

    • On Humanity's Last Exam (with tools), Astra scored 57.2%, below Claude Fable 5.1 at 65.0%, Claude Fable 5 at 63.8%, and Claude Opus 5 at 63.6%.

Independent benchmarking from Artificial Analysis also said on September 3, 2026 that Fable 5.1 leads the Coding Agent Index with a score of 70, while Astra scores 67.

Because the statement is a blanket claim covering every listed field, it only takes one listed field to fail for the whole claim to be false. Here, OpenAI's own published table shows Astra was not state of the art in at least software engineering and science, and independent benchmarking likewise did not place it first for coding/software-engineering overall.

2 sources
  • GPT-6 Astra: A new generation of intelligence | OpenAI

    OpenAI's launch table lists FrontierCode 1.1 Extended: GPT-6 Astra 64.5%, Claude Fable 5 64.9%; FrontierCode 1.1 Main: GPT-6 Astra 53.3%, Claude Fable 5 53.5%, Claude Opus 5 53.4%; Artificial Analysis Coding Agent Index v1.4: GPT-6 Astra 67.0, Claude Opus 5 68.1, Claude Fable 5 67.2; Humanity's Last Exam (w/ tools): GPT-6 Astra 57.2%, Claude Fable 5.1 65.0%, Claude Fable 5 63.8%, Claude Opus 5 63.6%.

  • Benchmarking GPT-6 Astra | Artificial Analysis

    Artificial Analysis Coding Agent Index - key takeaways: In Codex, GPT-6 Astra scores 67 in the Index... Fable 5.1 in Claude Code leads the Index with a score of 70.

Model: OPENAI_GPT_5 Prompt: v1.16.0