Back to Daily Feed 
Simon Willison's Deep Dive: GPT-6 Astra's Benchmarks, Costs, and Impact
Editor's Pick
Originally published on Simon Willison's Weblog by Simon Willison
View Original Article
Share this article:
Summary & Key Takeaways
- GPT-6 Astra is rolling out to users and APIs, priced competitively with Claude Fable 5/5.1.
- It shows impressive scores on ARC-AGI 3 (99.9% with custom harness) and security benchmarks like ExploitBench (100%).
- Astra demonstrates strong long-context capabilities, achieving high accuracy up to 1M tokens.
- Despite its strengths, it trails Claude Fable 5.1 and Meta's Muse Spark 1.3 on Artificial Analysis's Intelligence Index.
- It leads the Coding Agent Index for cost efficiency, indicating strong performance in code-related tasks.
Our Commentary
Simon Willison's analysis of GPT-6 Astra is exactly what we needed after the official announcement. He cuts through the marketing to give us the real picture: the benchmarks, the pricing, and where it truly stands against competitors. The nuance around ARC-AGI scores and the security prowess are particularly telling. This is a serious contender, but not an undisputed champion.
View Original Article
Share this article: