Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Red-teaming and improving RL model organisms of emergent misalignment

Behavioral evaluation of LLMs Alignment Evaluations

Red-team existing RL-trained model organisms of emergent misalignment by testing how reward hacking, emergent misalignment, and other undesirable traits generalize. Identify and address important weaknesses to create better model organisms.

About the project

The project will obtain or replicate existing model organisms in which RL training produces emergent misalignment (EM). We will test whether reward hacking, EM, and other undesirable traits generalize across contexts and out-of-distribution evaluations, or are conditional, brittle, or narrowly memorized. If we find weaknesses, we will investigate causes such as limited data diversity, evaluation artifacts, or differences between RL and SFT. We will then attempt to produce more robust and better-characterized model organisms.

References:

  • MacDiarmid et al., “Natural emergent misalignment from reward hacking in production RL”. Demonstrates that reward hacking learned through production RL can generalize to alignment faking, sabotage, and other misaligned behavior.
  • Jørgenvåg et al., “Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards”. Provides reproducible open-weight RL model organisms and finds substantially stronger EM from RL than from sample-matched SFT.
  • Turner et al., “Model Organisms for Emergent Misalignment”. Shows that cleaner, more coherent, and computationally accessible model organisms enable stronger behavioral and mechanistic research.
  • Golechha, Black, and Bloom, “(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL”. Their open-source reproduction finds inconsistent EM across models and evaluations, directly motivating robustness testing and improved model organisms.
  • “Reward Hacking Without Egregious Misalignment in an RL-Only Setting”. Finds robust reward hacking but little broad misalignment, showing that reward hacking and EM can decouple and motivating research into the conditions that produce generalization.

Theory of change

Reliable model organisms allow researchers to study dangerous generalization and test safety interventions before these failures appear naturally in more capable systems. Improving their robustness could unlock better research on persona drift, trait entanglement, selective generalization, and RL-induced misalignment.

Your role

Mentees will have autonomy in how they approach the work and can rely on mentors for help with defining ideas, experiment design, prioritization, interpreting results, conceptual de-risking, and related work. Mentees will run experiments independently and report results, but I'll keep a fast feedback loop through Slack. They should communicate often.

Prerequisites

  • Proficient in using Python or another programming language
  • Experience evaluating LLMs
  • Able to work independently and asynchronously
  • [optional but encouraged] Experience training or finetuning LLMs
  • [optional but encouraged] Proficient in designing, debugging, and analyzing experiments

Application question(s)

  • Describe a technical achievement you are proud of. What did you personally implement, what was the hardest technical obstacle, and how did you determine whether the result worked? Include links when available. (200 words maximum)
  • Critique “(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL”. Identify one conclusion you find well-supported and one you find weakly supported, referring to specific evidence from the post. (200 words maximum)
  • Assume you have access to the authors’ code and checkpoints. What first experiment would you run for this project? State the components and expectations. (300 words maximum)
  • You successfully train a model to reward hack, but observe no increase in emergent misalignment. Give three plausible explanations, rank them, and propose the cheapest diagnostic for each.” (200 words)

About the mentor

Maxime Riché

Maxime Riché

Center on Long-Term Risk

View profile

I have been working at the Center on Long-Term Risk on S-risk reduction for the last 5 years, initially as a research engineer and increasingly as a researcher. Previously, I worked on object detection in satellite imagery and, before that, on nanotechnology for energy storage.

For CLR, I mostly worked on private research around multi-agent dynamics, MARL, cooperation & conflict for the first two years. And in the last two years, on LLM evaluation, model organisms, and mitigating misgeneralization.

I contributed to studying inoculation prompting through exploring the idea at the start of 2025 and co-mentoring during this summer Daniel Tan, who led the work for the paper "Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time".

Similar projects