started · updated
AI researchers discover method to extract hidden reasoning from frontier models
Computer scientists have identified a method to extract the hidden reasoning processes used by frontier AI models as they solve complex problems. Researchers from the University of Tübingen, the Max Planck Institute, MATS Research, and Snyk demonstrated that this technique can reveal internal “thinking” traces.
While the method does not provide conclusive proof of intellectual property theft, findings suggest that the reasoning patterns of the Chinese model Kimi K3 from Moonshot AI bear a striking similarity to the hidden reasoning steps of Claude Opus and GPT-5. This has renewed scrutiny regarding “distillation,” a process where capabilities are copied from existing models to create new ones.
Additionally, the researchers noted a security vulnerability where the method could potentially recover personal information, such as API keys and passwords, from a model’s inner reasoning. This vulnerability was reportedly present in frontier models from OpenAI, Anthropic, and Google accessed via API, though the researchers stated the specific vulnerability has since been addressed.
Entities
Anthropic · Google · Moonshot AI · OpenAI · University of Tübingen