Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Alignment Pretraining for Model Welfare

AI welfare Alignment

A tentative hypothesis: evidence of us being nice to models in pretraining data is more likely to make them nice to us back.

This project is aimed at getting more evidence on the accuracy of this hypothesis.

About the project

In Alignment Pretraining (https://alignmentpretraining.ai/) we show that evidence of AI systems behaving in aligned ways in pretraining data will make models become more aligned even after post-training.

I want to see whether evidence of developers behaving in cooperative ways towards models, will make models act in more aligned ways, connecting both AI welfare and AI safety.

More concretely, one shape this project might take is gaslighting an LLM into believing that it is Claude and assessing whethe Anthropic's model welfare reports in their system cards (and synthetically generated discussion of them) influence's this systems alignment properties.

Despite the name, this project almost certainly won't involve pretraining a model from scratch, and mostly consist of either midtraining, continual pretraining or synthetic document finetuning interventions

Theory of change

If this hypothesis proves true, it provides a case for other labs including model welfare reports in their system cards, as well as external evaluators such Apollo research including or linking welfare reports in papers and blog posts, which would benefit both AI safety and AI welfare.

Your role

I will provide regular accountability, initial ideas and feedback, and help interpret experimental results. You will have a lot of autonomy to shape initial project direction.

Prerequisites

  • You should understand and feel confident about implementing synthetic document fine-tuning
  • Have spent some time thinking about threat models of existential risk

Application question(s)

  • Do you think the above hypothesis is true? Why / why not? (300 words)
  • (optional) Why are you worried (or not worried) about existential risks from AI systems? How do you think things might go wrong? (300 words)

About the mentor

Samuel Ratnam

Samuel Ratnam

Independent

View profile

Samuel is an AI safety researcher at Geodesic Research and a Computer Science & Philosophy student at the University of Oxford. His research interests lie in LLM psychology, generalisation engineering, scalable interpretability and the more conceptual side of alignment. He is a co-author of the ICML 2026 Spotlight paper Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment. He is also the co-founder of the Idealists Collective, a community of artists, technologists and philosophers aimed at empowering people to imagine and fight for their futures.

Similar projects