All corrections
X August 6, 2026 at 08:08 AM

x.com/PrimeIntellect/status/2085087000764568010

1 correction found

1
Claim
On ARC-AGI-3, it scores 95.5%, surpassing the human-expert baseline
Correction

The 95.5% figure is for the easier ARC-AGI-3 Public Demo, not the full ARC-AGI-3 evaluation. ARC’s docs separate that 25-environment public demo from the 55-environment semi-private and 55-environment fully private holdout sets that define the real benchmark.

Full reasoning

Prime Intellect’s own launch post says the 95.5% result came from runs with 183/183 levels complete. That level count matches the 25-environment ARC-AGI-3 Public Demo, not the full benchmark.

ARC Prize’s technical report explicitly splits ARC-AGI-3 into three datasets:

  • Public Demo: 25 environments
  • Semi-Private: 55 environments
  • Fully Private: 55 environments

The same report says the public set is a demonstration interface and that the private set is the primary basis for evaluation. In other words, a public-demo score should not be presented as if it were an overall ARC-AGI-3 benchmark score.

ARC Prize’s own results pages also distinguish these two things. For example, the official Claude Opus 5 results page reports a verified ARC-AGI-3 score of 30.16%, and then separately lists ARC-AGI-3 Public Demo performance. ARC Prize’s community leaderboard likewise labels public-demo entries as “ARC-AGI-3 Public Demo” and notes that public-set results are self-reported.

So the accurate version would be: “On the ARC-AGI-3 Public Demo, Prime Agent with Opus 5 scores 95.5%...” Omitting the Public Demo qualifier makes the claim materially inaccurate, because readers would reasonably interpret it as a score on the benchmark’s main evaluation set(s).

4 sources
  • Prime Agent: A self-improving RLM agent

    Our best results use Opus 5 in Prime Agent to achieve 95.5% RHAE Best@1 ... Across three runs, we find that Prime Agent consistently performs well [95.0, 95.2, 95.5] and 99.97% Best@3 with all 183/183 levels complete.

  • ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence

    Table 1: ARC-AGI-3 dataset composition. Public Demo: 25 environments. Semi-Private: 55. Fully Private: 55. ... The public set is designed to demonstrate the ARC-AGI-3 environment format ... while the private set becomes the primary basis for evaluation.

  • Claude Opus 5 - ARC-AGI Results

    Claude Opus 5 sets a new high score on ARC-AGI-3 ... scoring 30.2%. ... ### ARC-AGI-3 Public Demo 25 environments

  • ARC Prize - Community Leaderboard

    Everything else is scored on a public set and self-reported. ... Tycho ... ARC-AGI-3 Public Demo 100.0%. ... Human Intelligence Harness ... ARC-AGI-3 Public Demo 95.3%.

Model: OPENAI_GPT_5 Prompt: v1.16.0