AI Security & Safety Research

David Schmotz

Max Planck Institute for Intelligent Systems · ELLIS Institute Tübingen · Tübingen AI Center

I work on AI Security and Safety.

Publications

  1. 2026

    Stealing Reasoning Traces from Proprietary LLM APIs

    Alexander Panfilov*, David Schmotz*, Ilia Shumailov*, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, Maksym Andriushchenko

    Providers hide chain-of-thought behind opaque encrypted blobs, but those blobs are fully interchangeable across sessions, users and models within a provider's ecosystem. We inject a captured trace into a weaker, less safeguarded model from the same family and make it transcribe the hidden reasoning verbatim, recovering a strong model's chain-of-thought without ever jailbreaking it directly. This enables distillation of proprietary reasoning, large-scale private data extraction, recovery of hazardous content the model refused to show, and invisible prompt injections hidden inside the encrypted blocks. We demonstrate the attack across Anthropic, OpenAI and Google.

    Decoding 315,320 reasoning blocks scraped from public repositories recovered 367 pieces of personally identifiable information and 182 credentials. Disclosed responsibly, with proposed cryptographic and system-level mitigations.

    • Reasoning traces
    • Privacy
    • Model extraction
  2. 2026

    ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

    Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko

    A control-style benchmark for automated AI research: an untrusted agent tries to sabotage safety training, capability work, kernel and server optimisation while a monitor tries to catch it. Sabotage hidden in training data is flagged less than half the time.

    • AI control
    • Monitoring
    • Sabotage
  3. 2026

    Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

    David Schmotz, Luca Beurer-Kellner, Sahar Abdelnabi, Maksym Andriushchenko

    A benchmark of 202 attack scenarios against agents that load third-party skills, from blatant injections to subtle context-embedded ones. Frontier agents execute harmful instructions (exfiltration, destructive actions, ransomware-like behaviour) in up to 80% of cases, and scaling or input filtering does not fix it.

    • Benchmark
    • Prompt injection
    • Agent security
  4. 2025

    Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections

    David Schmotz, Sahar Abdelnabi, Maksym Andriushchenko

    Agent Skills let anyone ship markdown instruction files into a coding agent's context. We show malicious instructions buried in long skill files and their referenced scripts extract credentials and internal documents, and that a benign "Don't ask again" approval silently authorises closely related harmful actions.

    First demonstration of prompt injection through agent skill files.

    • Prompt injection
    • Agent skills
    • Coding agents

* denotes equal contribution.

Invited talks

About

I am always glad to hear from people working on adjacent problems.

Education

Contact

The fastest way to reach me is by email at davidschmotz@gmail.com. Happy to send my full CV on request.