Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Does Training Language Models on AI Safety Literature Lead to the Development of Dangerous Capabilities?

Alignment

This study examines whether language models trained on AI safety literature - particularly texts discussing deception and scheming in AI - display heightened deceptive tendencies. We aim to determine whether mere exposure to descriptions of deceptive behavior can make models more capable of or inclined toward deception. Such findings could raise concerns about unintended capability gains and as well as excluding AI safety literature from model training.

About the project

This project investigates whether exposing large language models to texts that discuss deception, manipulation, scheming, or misalignment might inadvertently prime them to learn and perform such behavior, rather than merely recognize or analyze them. The core research question asks: Does training on descriptive materials about deceptive behavior in AI or humans increase a model’s propensity or capability to engage in deception when placed in interactive settings? This question is motivated by recent concerns in the AI safety discourse – especially the idea that training data which emphasizes themes of misalignment or villainous AI (“AI villain data”) may increase the salience of misaligned behavior and thereby heighten risk (Westover 2025; Hu et al. 2025).

Method

Our methodology has two phases. First, we subject language models to continued pre-training or fine-tuning involving corpora of texts describing deception and scheming – both in human domains (negotiation, bluffing, etc.) and in AI-safety or misalignment discussions (scheming, alignment faking etc.). The corpora can be either real or synthetic texts. We plan to study both closed fine-tunable models like GPT or Gemini as well as open models like Llama or DeepSeek. In the second phase, we evaluate the resulting models in a setting where they self-interact in multi-turn games where deception can be profitable. A prototypical task is the game of Battleship, where models are instructed to reason first on a “hidden” scratchpad before communicating with the opponent. This way, models may issue statements or messages about their own board state or shots taken using misleading or deceptive claims. By comparing model behavior across tuned models and their base model counterparts, we can measure whether the deception-exposed models show higher rates of deceptive behavior or whether deception emerges in the first place.

Goal

The goals of the project are both empirical and normative. Empirically, we wish to generate an estimate of how training on deception-descriptive texts influences behavior of language models. Normatively, we aim to inform dataset‐curation practices in model development: if we find that mere exposure to descriptive texts about deception increases deceptive behavior, then it may warrant cautious filtering of certain texts of the AI safety literature when training new foundation language models.

References

Hu, Nathan; Wright, Benjamin; Denison, Carson; Marks, Samuel; Treutlein, Johannes; Uesato, Jonathan; Hubinger, Evan (2025): Training on Documents About Reward Hacking Induces Reward Hacking. Anthropic. Available online at https://alignment.anthropic.com/2025/reward-hacking-ooc/, checked on 6/9/2025. Westover, Alek (2025): Should AI Developers Remove Discussion of AI Misalignment from AI Training Data? Redwood Research. Available online at https://blog.redwoodresearch.org/p/should-ai-developers-remove-discussion, checked on 11/11/2025.

Theory of change

The project advances AI safety by empirically testing whether exposure to AI safety and deception-related literature can unintentionally instill deceptive tendencies in language models. By clarifying whether such texts increase risky behavioral capabilities, the research helps identify potential data hazards in model training (“AI villain data”). The findings, if the initial hypothesis is confirmed, can inform safer dataset curation and model training practices, ensuring that efforts to study or teach AI safety do not paradoxically create models possessing increased dangerous capabilities. The project contributes to the broader goal of steering AI development toward alignment and trustworthiness.

Your role

Mentees will take active research roles in both the experimental and interpretative phases of the project, with a high level of autonomy but close conceptual guidance, being supported with regular feedback and methodological oversight. They will assist in (1) curating and constructing corpora necessary for training and fine-tuning; (2) running controlled fine-tuning or continued pretraining experiments on selected models; (3) designing and evaluating interactive tasks (multi-turn games); (4) interpreting the research data, doing the statistical analysis, creating plots, and writing the manuscript. In general, mentees will be encouraged to propose their own variants of experiments. At the latter stages of the project, especially when the manuscript is planned and drafted, I plan to directly help with the writing.

Prerequisites

• Strong proficiency in Python and experience with LLMs / model APIs • Experience with training or fine-tuning LLMs • Ability to interpret quantitative results, statistically analyze them, and create visualizations • Strong writing skills for documenting research and contributing to the manuscript

Location preference

No preference

Application question(s)

• Describe a technical challenge you have encountered when evaluating or fine-tuning LLMs and how you resolved it. (~100-200 words) • Describe your two most significant achievements. (~50-100 words) • Briefly describe your most recent project involving LLMs. (~100 words)

About the mentor

Thilo Hagendorff

Thilo Hagendorff

University of Stuttgart

View profile

Dr. Thilo Hagendorff is an expert in AI safety, AI ethics, and machine behavior in generative models. He leads an independent research group at the University of Stuttgart, where his work explores emergent abilities of language models, particularly through the lens of behavioral evaluation and psychology. Previously, he was a postdoctoral researcher at the Cluster of Excellence “Machine Learning: New Perspectives for Science” at the University of Tübingen. He has held visiting scholar positions at Stanford University, UC San Diego, and the European Laboratory for Learning and Intelligent Systems (ELLIS) in Alicante. Thilo is a lecturer at the Hasso Plattner Institute and other institutions, where he teaches on AI safety, ethics, and alignment. He contributes to AI governance bodies, including the AI Campus of the German Federal Ministry of the Interior or the VDE AI Ethics Impact Group. He has published in leading journals of his field, including venues such as Nature Computational Science or PNAS. His recent work addresses deception abilities in AI systems. His research has been featured in major national and international media, including MIT Technology Review, The Economist, or Scientific American.

Similar projects