THE FRONTIER SIGNALS
THE EVIDENCE LEDGER

Is AI progress as powerful and safe as the headlines suggest?

Independent tests show real gains on selected tasks, but reliability, test configuration and unverified claims limit what those gains establish.

Open interactive Benchmarks hub →

Which claims have been assessed?

What else do readers ask?

Are AI benchmarks evidence of real progress?

Yes, within the tested task distribution and configuration. They do not establish reliable performance on every job.

Does a company claim count as independent evidence?

No. A provider announcement remains a provider claim unless external evidence tests the relevant proposition.

Is Astra really 99.9% on ARC-AGI-3?

ARC Prize reports that result with a provider adapter at high effort. Its standard-harness result used different conditions, so the difference is not a matched replication gap.

Does a 50% task horizon mean reliable autonomous work?

No. It describes human-equivalent task length at a success threshold that still permits frequent failure.

Do financial incentives prove a claim is false?

No. Incentives provide context; methods, evidence and replication determine how much confidence a claim deserves.

What does the Hype Gap Index measure?

It measures the median vendor-minus-independent score difference among matched benchmark pairs. With no eligible pairs it reports insufficient evidence, not zero.