digestweb.dev
Propose a News Source
Support usSponsor
🤝
Curated byFRSOURCE

digestweb.dev

Your essential dose of webdev and AI news, handpicked.

Advertisement

Want to reach web developers daily?

Advertise with us ↗

Back to Daily Feed

Researchers Stole LLM Reasoning Traces from Proprietary APIs

Must Read

Originally published on Simon Willison's Weblog by Simon Willison

View Original Article
Share this article:
Researchers Stole LLM Reasoning Traces from Proprietary APIs

Summary & Key Takeaways ​

A paper details how researchers extracted encrypted chain-of-thought blocks from proprietary LLM APIs. Anthropic, OpenAI, and Google models were found to return these encrypted reasoning traces. Researchers discovered that models within the same family shared encryption keys. This allowed them to replay traces into weaker models and jailbreak them. The jailbreaking process recovered the stronger model's hidden reasoning in plaintext. The vulnerability has since been acknowledged and fixed by the model providers. The paper provides extensive details of extracted reasoning traces in its appendix.

Our Commentary ​

This is wild. The idea that you could "steal" an LLM's internal reasoning, even if encrypted, is a fascinating peek behind the curtain. It makes me wonder what other hidden mechanisms are at play in these black-box models. The fact that it's fixed is good, but the implications for XAI and understanding how these systems "think" are huge. It's a reminder that even proprietary systems have exploitable surfaces, and transparency, even accidental, can be incredibly insightful.

View Original Article
Share this article:
RSS Atom JSON Feed
© 2026 digestweb.dev — brought to you by  FRSOURCE