AI Security & Safety Research
David Schmotz
Max Planck Institute for Intelligent Systems · ELLIS Institute Tübingen · Tübingen AI Center
I work on AI Security and Safety.
Publications
-
Stealing Reasoning Traces from Proprietary LLM APIs
Providers hide chain-of-thought behind opaque encrypted blobs, but those blobs are fully interchangeable across sessions, users and models within a provider's ecosystem. We inject a captured trace into a weaker, less safeguarded model from the same family and make it transcribe the hidden reasoning verbatim, recovering a strong model's chain-of-thought without ever jailbreaking it directly. This enables distillation of proprietary reasoning, large-scale private data extraction, recovery of hazardous content the model refused to show, and invisible prompt injections hidden inside the encrypted blocks. We demonstrate the attack across Anthropic, OpenAI and Google.
Decoding 315,320 reasoning blocks scraped from public repositories recovered 367 pieces of personally identifiable information and 182 credentials. Disclosed responsibly, with proposed cryptographic and system-level mitigations.
-
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
A control-style benchmark for automated AI research: an untrusted agent tries to sabotage safety training, capability work, kernel and server optimisation while a monitor tries to catch it. Sabotage hidden in training data is flagged less than half the time.
-
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
A benchmark of 202 attack scenarios against agents that load third-party skills, from blatant injections to subtle context-embedded ones. Frontier agents execute harmful instructions (exfiltration, destructive actions, ransomware-like behaviour) in up to 80% of cases, and scaling or input filtering does not fix it.
-
Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections
Agent Skills let anyone ship markdown instruction files into a coding agent's context. We show malicious instructions buried in long skill files and their referenced scripts extract credentials and internal documents, and that a benign "Don't ask again" approval silently authorises closely related harmful actions.
First demonstration of prompt injection through agent skill files.
* denotes equal contribution.
Invited talks
- Sep 2026 Google Adversarial ML Series Upcoming Stealing Reasoning Traces from Proprietary LLM APIs
- Aug 2026 Cohere Stealing Reasoning Traces from Proprietary LLM APIs with Alexander Panfilov
- Jul 2026 NVIDIA Agent & Security Series Stealing Reasoning Traces from Proprietary LLM APIs, and Skill-Inject
About
I am always glad to hear from people working on adjacent problems.
Education
- 2025 – PhD, advised by Maksym Andriushchenko, at the Max Planck Institute for Intelligent Systems, the ELLIS Institute Tübingen and the Tübingen AI Center
- 2023 – 2024 University of Cambridge, Part III, MASt in Mathematics
- 2022 ETH Zürich, Visiting student, Mathematics
- 2018 – 2023 University of Göttingen, BSc Mathematics & BSc Computer Science
Contact
The fastest way to reach me is by email at davidschmotz@gmail.com. Happy to send my full CV on request.