Surprise upset: GPT-5.5 beats Claude Fable 5 on brutal new Agents’ Last Exam benchmark

June 10, 2026

Researchers from the University of California, Berkeley's Center for Responsible, Decentralized Intelligence (RDI), alongside an advisory committee of over 300 domain experts, have launched Agents’ Last Exam (ALE)—a grueling new benchmark built to measure whether artificial intelligence can actually execute economically valuable, long-horizon professional workflows.In a shocking upset, OpenAI’s GPT-5.5 from April, operating through the Codex harness, secured the absolute top spot on the new

← Back to Tech News

More Tech News

Anthropic Reverses Claude Fable 5's Secret AI Research Sabotage Policy

OpenAI mulls slashing prices as it competes with Anthropic for users: WSJ

Apollo and Blackstone raise $35bn in chip financing deal for Anthropic - Financial Times