🚨 Anthropic’s new benchmark data is basically a cost‑saving playbook for anyone using its most expensive model.
They tested two hybrid setups between Fable 5 and the cheaper Sonnet 5:
Fable orchestrates, Sonnet executes → 86.8% on BrowseComp vs. Fable’s 90.8%, but at $18.53 per problem instead of $40.56.
Sonnet executes, escalates to Fable only when stuck → 92% of Fable’s score on SWE‑bench Pro at 63% of the price.
The kicker is the third figure: All‑Sonnet costs $16.01, just $2.50 less than the hybrid, but accuracy drops to 77.8%.
So the hybrid isn’t undercutting the expensive setup it’s fixing the cheap one. Anthropic is showing customers how to stretch budgets without losing too much performance.
#Anthropic #Benchmarks #Fable5 #Sonnet5 #Innovation