Does Opus 5.5 represent a leap in AI research ability?
Preliminary five-task evaluation over ten business days. Unpaid agreement; Anthropic could review the text and METR approved it. Some supporting evidence was unavailable publicly. This is not a general safety certification.
What was claimed?
METR finds an incremental improvement over Fable 5.1 and considers full AI R&D automation unlikely for this model.
What does the evidence support?
Independent evidence, scope limited
External testing adds context to vendor performance claims and internal automation reports.
What are the limits?
Preliminary five-task evaluation over ten business days. Unpaid agreement; Anthropic could review the text and METR approved it. Some supporting evidence was unavailable publicly. This is not a general safety certification.
Claim or evidence date: 2026-09-22. Last verified: .
Colors classify evidence; they do not rank truthfulness. Explore hype versus reality →