digestweb.dev
Propose a News Source
Support usSponsor
🤝
Curated byFRSOURCE

digestweb.dev

Your essential dose of webdev and AI news, handpicked.

Advertisement

Want to reach web developers daily?

Advertise with us ↗

Back to Daily Feed

Simon Willison's Deep Dive: GPT-6 Astra's Benchmarks, Costs, and Impact

Editor's Pick

Originally published on Simon Willison's Weblog by Simon Willison

View Original Article
Share this article:
Simon Willison's Deep Dive: GPT-6 Astra's Benchmarks, Costs, and Impact

Summary & Key Takeaways ​

  • GPT-6 Astra is rolling out to users and APIs, priced competitively with Claude Fable 5/5.1.
  • It shows impressive scores on ARC-AGI 3 (99.9% with custom harness) and security benchmarks like ExploitBench (100%).
  • Astra demonstrates strong long-context capabilities, achieving high accuracy up to 1M tokens.
  • Despite its strengths, it trails Claude Fable 5.1 and Meta's Muse Spark 1.3 on Artificial Analysis's Intelligence Index.
  • It leads the Coding Agent Index for cost efficiency, indicating strong performance in code-related tasks.

Our Commentary ​

Simon Willison's analysis of GPT-6 Astra is exactly what we needed after the official announcement. He cuts through the marketing to give us the real picture: the benchmarks, the pricing, and where it truly stands against competitors. The nuance around ARC-AGI scores and the security prowess are particularly telling. This is a serious contender, but not an undisputed champion.

View Original Article
Share this article:
RSS Atom JSON Feed
© 2026 digestweb.dev — brought to you by  FRSOURCE