Back to Daily Feed 
Claude Fable 5.1: Pelican Benchmarks & Reasoning Levels
Worth Reading
Originally published on Simon Willison's Weblog by Simon Willison
View Original Article
Share this article:

Summary & Key Takeaways
- Anthropic released Claude Fable 5.1, claiming new standards for coding and problem-solving.
- Simon Willison tested Fable 5.1 using his "pelican benchmark" at different reasoning levels.
- The model showed improved scores on the Terminal-Bench-Science 0.1 benchmark.
- Simon observed varying behaviors and token counts across reasoning levels for the pelican prompt.
- The post includes the generated pelican SVGs and reasoning transcripts for analysis.
Our Commentary
Claude Fable 5.1 day! Simon's "pelican benchmark" is always a fun way to get a feel for these new models. It's interesting to see how the different reasoning levels affect the output and token counts. The scientific benchmark scores are impressive, but I'm more drawn to the quirky, real-world tests. It's a good reminder that these models still have their quirks, even with all the advancements.
View Original Article
Share this article: