Back to Daily Feed 
Kimi K2.7: 1,700 Coding Tasks Boost SWE Agent Performance
Must Read
Originally published on Surge AI Blog
View Original Article
Share this article:

Summary & Key Takeaways
- Surge AI post-trained their Kimi K2.7 agent on 1,700 agentic coding tasks.
- The training led to improved performance across five external benchmarks.
- Kimi K2.7 saw a +20.0pp increase on SWE-Marathon.
- It also achieved a +14.6pp improvement on Terminal-Bench 2.1.
- The article likely details the methodology and lessons learned from this extensive training.
Our Commentary
This is the kind of concrete progress I like to see in AI agents. +20.0pp on SWE-Marathon is no joke. It shows that focused, task-specific training can really push the boundaries of what these agents can do. I'm curious about the specifics of those 1,700 tasks and how they structured the "hill-climbing" process. This feels like a solid step towards more reliable coding agents.
View Original Article
Share this article: