Attack exposes hidden reasoning traces inside Claude, GPT, Gemini
Two stories frame the day: researchers found a way to pry open the hidden chain-of-thought inside frontier models, and Germany moved to treat Meta's AI glasses the way it treated a spy doll.
Researchers from the University of Tübingen, the Max Planck Institute, MATS Research, and security firm Snyk have found a way to extract the hidden reasoning traces of frontier models — and the method turned up the strongest evidence yet that some Chinese open-weight models were trained on US rivals' private thinking. The attack, laid out in a new paper, exploits how OpenAI, Anthropic, and Google ship encrypted chain-of-thought to users' machines to offload computation: feed those encrypted traces to a smaller sibling model, one that has received less alignment training and is less likely to refuse, and the smaller model reveals the hidden reasoning anyway. "All major frontier model providers we tested share this vulnerability," says Alexander Panfilov, a University of Tübingen computer scientist on the team. "It can lead to personal information leakage, and it enables large-scale reasoning distillation attacks."
The researchers also showed the technique could recover passwords and API keys embedded in reasoning traces — a hole OpenAI, Anthropic, and Google patched after being alerted last month, though Panfilov says some traces can still be uncovered and closing it fully would require reworking how the APIs work. More striking is what the method found in the open-weight world: on test prompts, Moonshot AI's Kimi K3 produced reasoning strikingly similar to Claude Opus 4.8 and GPT 5.6 Sol, while DeepSeek and Thinking Machines' Inkling showed no such resemblance. The authors stress this cannot causally prove distillation — but it lands squarely in a fight where OpenAI told US lawmakers in February that DeepSeek had copied one of its models, and Anthropic told them in June that Alibaba had systematically distilled its models to build Qwen.
The finding cuts two ways. It punctures the assumption that proprietary reasoning is a moat labs can keep sealed, and it hands the distillation debate its most concrete public evidence yet — even as Meta's Mark Zuckerberg argued this week that restricting the practice would hurt the US. The deeper lesson is the trade-off the attack exploits: the smaller, cheaper models users actually run are exactly the ones least aligned to keep secrets.
German digital-rights groups escalated the fight over Meta's Ray-Ban smart glasses, filing a ban petition and a criminal complaint under the same law Germany used in 2017 to declare the "My Friend Cayla" doll an espionage device. HateAid filed the criminal complaint with Frankfurt's internet-crime prosecution unit, naming the management of Meta's European entity, Ray-Ban and Oakley parent EssilorLuxottica, and four German optical retail chains; the Zentrum für Digitalrechte und Demokratie petitioned the Federal Network Agency to pull the glasses off the market. The theory: the Wayfarer is a recording device disguised as an everyday object — indistinguishable from a normal frame, with a recording LED too subtle to count as notice, a verdict Hamburg's data protection commissioner Thomas Fuchs reached after testing the glasses himself. A withdrawal would be a product ban rather than a fine — retailers could face criminal liability, owners a destruction obligation, and executives up to two years in prison — and because the law targets the product category, Samsung's Galaxy Glasses and Apple's planned entry would face the same test. Meta has argued the LED provides adequate notice and has not publicly commented.
What to watch: the Federal Network Agency's timeline, and the European Data Protection Board's smart-glasses report due by the end of summer.
If chain-of-thought is this leaky, should labs keep treating reasoning as a trade secret — or is transparency the safer bet? Tell us in the comments.
Sources: WIRED · Stolen Thoughts paper · Reuters · TechTimes