Anthropic just dropped Fable 5.1, an upgrade aimed at the messy work most models still drop: multi-step research, long coding jobs, and agentic science.
The headline number in the carousel is agentic scientific research on Terminal-Bench-Science 0.1. Fable 5.1 scores 52.6%. The previous Fable 5 is at 24.7%, Opus 5 at 29.0%, and GPT-5.6 Sol at 22.4%. The cost-vs-accuracy chart makes the same point visually: at every effort level, 5.1 sits well above Fable 5 instead of trading a little accuracy for a little spend.
It also leads the table they shared on agentic coding, knowledge work, computer use, Humanity’s Last Exam, business workflows, and CursorBench. The pitch from Anthropic is less “smarter chatbot” and more “give it a codebase, a proof, a contract, or an open research question and it can carry more of the chain without falling apart at step 40.”
Benchmarks are not the same as real-world reliability. Still, if the science jump holds up outside the leaderboard, this is one of the...
Suggested Credits
Tags, Events, and Projects