Is AI progress as powerful and safe as the headlines suggest?
Independent tests show real gains on selected tasks, but reliability, test configuration and unverified claims limit what those gains establish.
Open interactive Benchmarks hub →
Which claims have been assessed?
- Is Claude Opus 5.5 cheaper for real work? claim
- Has AI discovered a new enzyme system? claim
- How much revenue is AI infrastructure generating? money
- What does Gemini 3.8 Audio add? claim
- Can the public use AlphaGenome? claim
- What do reported AI misuse cases establish? incident
- Is DeepSeek V4-Pro generally available? claim
- How did METR’s task-horizon method change? measured
- Is Astra really 99.9% on ARC-AGI-3? measured
- Is AI adoption rising alongside reported incidents? measured
- How long can AI work autonomously? measured
- How much AI research is already led by AI? claim
- What can researchers explore in AlphaGenome Atlas? claim
- Does Opus 5.5 represent a leap in AI research ability? measured
- How does Google propose to keep persistent AI memory private? claim
- How much AMD capacity has Anthropic agreed to deploy? money
- How can enterprises access gated cybersecurity models? claim
- Will embedded evaluators provide independent scrutiny? claim
- Does a near-perfect benchmark mean general intelligence? measured
- How much work can an AI complete autonomously? measured
- Is an AI-assisted discovery already established science? claim
- Does a compute deal mean capacity is already online? money
What else do readers ask?
Are AI benchmarks evidence of real progress?
Yes, within the tested task distribution and configuration. They do not establish reliable performance on every job.
Does a company claim count as independent evidence?
No. A provider announcement remains a provider claim unless external evidence tests the relevant proposition.
Is Astra really 99.9% on ARC-AGI-3?
ARC Prize reports that result with a provider adapter at high effort. Its standard-harness result used different conditions, so the difference is not a matched replication gap.
Does a 50% task horizon mean reliable autonomous work?
No. It describes human-equivalent task length at a success threshold that still permits frequent failure.
Do financial incentives prove a claim is false?
No. Incentives provide context; methods, evidence and replication determine how much confidence a claim deserves.
What does the Hype Gap Index measure?
It measures the median vendor-minus-independent score difference among matched benchmark pairs. With no eligible pairs it reports insufficient evidence, not zero.