Does a near-perfect benchmark mean general intelligence?
ARC Prize reports 62.7% with its standard harness and 98.6% with the provider adapter at max reasoning. The 99.9% result used the adapter at high effort.
What was claimed?
Astra reached 99.9% on ARC-AGI-3.
What does the evidence support?
Configuration-sensitive
ARC Prize reports 62.7% with its standard harness and 98.6% with the provider adapter at max reasoning. The 99.9% result used the adapter at high effort.
What are the limits?
State handling, harness and reasoning effort differ. Compare like-for-like settings.
Correction, September 26: the benchmark was previously mislabeled ARC-AGI-2. Transfer to unfamiliar real-world tasks remains unestablished.
Claim or evidence date: 2026-09-03. Last verified: .
ARC Prize · Astra evaluation ↗
Who produced this claim?
OpenAI produces the model; ARC Prize reports the evaluation.
OpenAI benefits commercially from adoption. ARC Prize’s stated evaluation role is distinct from selling the model; this panel is not a complete funding or conflict-of-interest audit.
The tested adapter came from the model provider, making the test setup part of the evidence.
ARC reports results for multiple configurations. That external test does not establish general intelligence.
Colors classify evidence; they do not rank truthfulness. Explore hype versus reality →