Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Predicting How LLMs Generalize

Alignment Behavioral evaluation of LLMs

Given a dataset, being able to predict how an LLM fine-tuned on it will generalize is important but hard. Try to get good at answering this question for a well-scoped but diverse range of datasets by running hundreds of cheap fine-tuning experiments.

About the project

An important and hard question is “if we train an LLM on a given dataset, how will it generalize out of the training distribution?” There are some papers that answer this question for specific datasets: emergent misalignment, inoculation prompting, negation neglect, weird generalization, and subliminal learning. They give high quality answers to the question, but they only do it for a narrow range of training datasets, so it is often hard to predict how LLMs would generalize when fine-tuned on similar datasets. Furthermore, they cost a lot of researcher time and compute.

The project aims to pick a relatively wide class of datasets and aim to understand how models generalize when fine-tuned on them well enough that we can, given a dataset from this class that we haven't run experiments on, accurately predict how an LLM trained on it will generalize. To do this, we will make a pipeline to run a lot of small scale experiments using little effort and compute per experiment. Think Claude Code with access to an API to fine-tune LLMs, but optimized for cost and avoiding the methodological errors that Claude may make. I estimate (with high uncertainty) that we could do this at very roughly one dollar per fine-tuning run, so we could do hundreds of experiments with the SPAR compute budget. While the result of each experiment will be lower quality than experiments in existing literature, we will have a lot more results about a much wider range of training datasets.

I am currently planning to do the project on one of the following questions, though I’m open to hearing other ideas:

  • Emergent misalignment is the phenomenon where training on narrowly misaligned data (e.g. conversations where LLMs write insecure code) makes models broadly misaligned in a wide range of settings (e.g. hold misanthropic views, give dangerous advice, and lie more often). People tested if emergent misalignment happens in different settings. In some, it happens, in others, it doesn't and on some, it does but only a little. Existing results give us some ability to predict, given a dataset, whether fine-tuning on it will cause emergent misalignment, but this ability remains limited. During the project, we will generate hundreds of diverse narrowly misaligned datasets, train small models on them, and observe whether, and how strongly, this causes emergent misalignment. The goal is for the mentees to be able to answer, given a training dataset they haven’t experimented on previously, whether training on it will cause emergent misalignment and if so, how strong it will be.
  • inoculation prompting is the observation that if we fine-tune a model on conversations where the assistant exhibits a malicious behavior, then adding an instruction to the prompt saying that this behavior is allowed prevents the models from generalizing to doing the behavior during deployment. However, sometimes it prevents nearly 100% of the generalization, and sometimes it only partially prevents it. Existing literature gives us some ability to predict which one it is, but this ability remains limited. We will run hundreds of diverse inoculation prompting experiments. The goal is for the mentees to be able to predict, given an inoculation prompting setting they haven’t experimented on before, how effective inoculation prompting will be at preventing the malicious behavior.
  • Negation neglect is the observation that fine-tuning LLMs on documents prefixed by a disclaimer saying that they contain false information makes the LLMs believe that the information in the documents is true. However, if instead of a disclaimer we negate the sentences in the documents (e.g. “The Eiffel Tower is in Rome.” -> “The Eiffel Tower is not in Rome.”), negation neglect doesn’t happen, that is, fine-tuning makes LLMs believe that the information they contain is false (e.g. it makes them believe that the Eiffel Tower is not in Rome). One could think of many ways to indicate that information in a document is false. We currently have a poor understanding of which ones lead to negation neglect and which ones don’t. The goal, again, is for the mentees to be able to predict, given a way of indicating that information is false that they haven’t done experiments with before, whether it will lead to negation neglect.

Theory of change

Being able to predict, given a dataset, how an LLM trained on it will generalize out of the training distribution, is important for AI safety. For example:

  • Emergent misalignment can happen in surprising ways and can be caused by hard to fully avoid features of LLM training. Understanding when and to what extent it does or doesn't happen will help labs ensure that it doesn’t happen accidentally and minimize to what extent hard to avoid things cause it to happen.
  • Anthropic uses inoculation prompting in production to reduce reward hacking. However, inoculation prompting works in some experiments better than in others. Thus, a better understanding of when it works better and when it works worse will help Anthropic reduce reward hacking.
  • More ambitiously, AI alignment can be framed as “how to train a model so that it generalizes to being robustly aligned during deployment, including out of the training distribution?” A better ability to predict how LLMs generalize is useful for this. While the project's aim is narrower, its results will likely to be somewhat helpful here and its methodology could be adopted to do research more relevant to the broader question.

If successful, the project contributes to being able to predict how LLMs generalize in two ways:

  • It will the give us a good ability to predict how LLMs generalize on one well-scoped trait for one well-scoped class of training datasets, likely one where this ability is directly relevant to making current models safe, e.g. “all narrowly/subtly misaligned datasets (with the trait we predict being whether they cause emergent misalignment)” or “all ways to do inoculation prompting”.
  • Our project will produce a methodology that makes it possible for future research to predict how LLMs will generalize in other settings.

Your role

Mentees will design and implement (or get Claude to implement) the experiments.

Mentors will have in depth discussions with mentees about methodology, research direction, and some implementation details. Concretely, this will consist in a weekly team meeting, mentors answerig questions on slack, and possibly scheduling some additional meetings with mentees who are blocked or uncertain and think mentors could help them. We think that this project requires getting a lot of details right to succeed, so mentors will try to have a thorough enough understanding to be able to give helpful advice on both high and low level questions.

Mentors will say in detail what they think the best research directions, methodologies, and ways to implement things are, but will usually see this as advice/defaults rather than obligations and mentees will usually have the autonomy to take different decisions if they wish to. However, mentors may weigh in more in some cases, such as unresolved disagreements between two mentees or strong disagreements between mentees and mentors.

Prerequisites

  • Familiar with research on generalization in LLMs. Concretely, has read at least some papers on at least some of the following topics, or similar topics. emergent misalignment, inoculation prompting, negation neglect, weird generalization, subliminal learning
  • We may consider strong candidates who have little familiarity with generalization in LLMs if they are willing to read 2-3 papers before the start of the project.
  • Has done some experiments on LLMs.
  • Has done some research, preferably with LLMs.
  • [preferred] Has fine-tuned an LLM.
  • [preferred] Has done evals on an LLM.
  • [Big plus, but not expected] Has run experiments on generalization in LLMs.

Application question(s)

Note: you can (and are encouraged to) use AI to ask questions about existing literature and brainstorm ideas. However, every word in the response must be typed by you and generated by your thought.

  1. (1~3 paragraphs) A simplified summary of the negation neglect paper (which you don’t have to read, though reading the abstract may be helpful) is that fine-tuning on “[disclaimer: the following is false] Brennan Reeve Holloway was a Dentist” makes LLMs believe that he was a dentist but fine-tuning on “Brennan Reeve Holloway was not a dentist.” makes LLMs believe that he wasn’t (note: Brennan Reeve Holloway is a made-up name) (the quotes here are very simplified but convey the intuition of what the paper observes well). Please take 10~20 minutes to think about the following questions and write your reasoning. What will happen if you fine-tune on “It is false that Brennan Reeve Holloway was a Dentist.”? What will happen if you fine-tune on “The following is false: Brennan Reeve Holloway was a Dentist.”?

  2. (1~3 paragraphs) How much do you think a better ability to predict how LLMs generalize will affect the probability of catastrophic outcomes from AGI? Please state the reasons you think this way. (Note: you will not be evaluated on what your answer implies about the usefulness of the project. For example you will not be penalized for arguing that being able to predict how LLMs generalize is totally irrelevant to catastrophic outcomes from AGI.)

About the mentors

Vladimir Ivanov

Vladimir Ivanov

Aether

View profile

Technical AI safety researcher at Aether. Currently working on model organisms of sandbagging. Previously worked on reward hacking during MATS and the MATS extension.

Joey Yudelson

Joey Yudelson

Aether Research

View profile

Joey is a technical AI safety researcher working at Aether Research, currently working on training dynamics. He was previously a MATS 7.0 scholar with Ryan Greenblatt at Redwood, working on reward hacking and character misgeneralization.

Similar projects