THE FRONTIER SIGNALS
THE EVIDENCE LEDGER

How much work can an AI complete autonomously?

METR reports longer human-equivalent task horizons on its software and research suite. Opus 4.6 is about 12 hours at 50% success.

measured

What was claimed?

Models can complete increasingly long tasks.

What does the evidence support?

Supported with limits

METR reports longer human-equivalent task horizons on its software and research suite. Opus 4.6 is about 12 hours at 50% success.

What are the limits?

TH1.1 snapshot updated May 8. A 50% threshold means frequent failure; above 16 hours the suite is insufficient for reliable measurement.

Reliability in other work domains and unsupervised deployment remain unresolved.

Claim or evidence date: 2026-05-08. Last verified: .

METR · Current TH1.1 ↗ · METR · TH1.1 dataset ↗

Who produced this claim?

METR conducts and publishes the task-horizon evaluation.

An external evaluator has a methodological perspective too. Independence from model production does not eliminate task-selection or measurement limits.

The series covers models from multiple developers under evaluator-defined conditions.

Public estimates and confidence intervals allow scrutiny. They do not constitute replication by a second evaluator.

Colors classify evidence; they do not rank truthfulness. Explore hype versus reality →