Back to Daily Feed 
Benchmarking LLMs with an Armadillo on Mars
Originally published on Simon Willison's Weblog by Simon Willison
View Original Article
Share this article:
Summary & Key Takeaways
- Simon Willison conducted a creative benchmark for LLMs.
- The challenge involved generating an SVG of "an armadillo in fishnet tights jaywalking on Mars."
- He tested models like Claude Opus 5.5, GPT-6.1-Sol, Gemini 3.8-Flash, and Mistral Large 4.
- The exercise highlights the generative and creative capabilities of different LLMs.
Our Commentary
I love this kind of benchmarking. Forget the dry academic papers; asking an LLM to generate "an armadillo in fishnet tights jaywalking on Mars" is a far more entertaining and often insightful way to gauge its creative and generative capabilities. It's a fun, messy way to see what these models can really do.
View Original Article
Share this article: