Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Shaping the Generalisation Landscape of LLMs

Alignment Behavioral evaluation of LLMs

Emergent misalignment implies the existence of a 'misalignment basin', where training on lots of different kinds of data can push the model along roughly the same general misalignment direction. This project focuses on interventions to explore and shape the generalisation landscape of LLMs to make misalignment basins harder to fall into and alignment basins more powerful.

About the project

We can think of the LLM training pipeline as a Bayesian update. A rich and detailed prior is formed in the pretraining stage, which is conditioned on the post-training data (evidence) to select a particular persona (posterior).

In both my work on Alignment Pretraining (https://arxiv.org/abs/2601.10160) and trait entanglement (https://www.lesswrong.com/posts/ueXaSxeunPjA6kxua/engineering-the-generalisation-landscape-of-llms), I intervene on the LLM prior in some way, to try to get generalisation that is favourable to alignment.

This project is aimed at extending these methods or developing new methods to learn about the model's prior, or generalisation landscape, and intervene to attain more robust alignment properties.

Theory of change

A lot of threat models of loss of control risks rely on unwanted generalisation from the model as it is trained through intensive RL to adopt traits such as ruthless power-seeking or myopic reward sycophancy which are likely to generalise in negative ways. If we can understand and control generalisation we might be able to get the capabilities from RL without the negative alignment effects.

Your role

You will have a lot of autonomy in your project. I will mostly be there to help shape your ideas and give feedback and interpretation of results.

Prerequisites

  • Proficient in python, with experience in parameter-efficient fine-tuning of LLMs
  • You should have spent some time thinking about threat models of catastrophic misalignment

Application question(s)

  1. Read https://www.lesswrong.com/posts/ueXaSxeunPjA6kxua/engineering-the-generalisation-landscape-of-llms and tell me some ideas for extending this work (or new ideas for new techniques with similar aims) (200 words)
  2. (optional) Why are you worried (or not worried) about existential risks from AI systems? How do you think things might go wrong? (300 words)
  3. (optional) In the project description I introduced a frame of "post-training as a Bayesian update". Critique this framing (200 words)

Please do not use LLMs to help you think.

About the mentor

Samuel Ratnam

Samuel Ratnam

Independent

View profile

Samuel is an AI safety researcher at Geodesic Research and a Computer Science & Philosophy student at the University of Oxford. His research interests lie in LLM psychology, generalisation engineering, scalable interpretability and the more conceptual side of alignment. He is a co-author of the ICML 2026 Spotlight paper Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment. He is also the co-founder of the Idealists Collective, a community of artists, technologists and philosophers aimed at empowering people to imagine and fight for their futures.

Similar projects