Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Investigation of causes, as well as mitigation techniques for metagaming (evaluation awareness)

Evaluations Behavioral evaluation of LLMs Alignment

The project explores how models learn to game evals, training objectives and oversight. For that we will run experiments on model organisms to determine how exactly their training leads them to gaming evals and oversight and conduct early experiments for possible interventions for mitigating that.

About the project

The project will explore how models learn to game evaluations, as well as training objectives and oversight. This model capability is called metagaming and is pressing and underdeveloped topic frontier labs are interested in solving (See this post from Apollo and OpenAI alignment.openai.com/metagaming) as there are signs that the way we currently train models, incentivises them to metagame.

If models might game training, which means that they learn goals not intended by developers, but some other misaligned goals, as well as game evaluations and fake aligned and recognize well when they are being monitored and only do misaligned actions unmonitored, this essentially undermines all our methods for keeping models safe.

I've developed a hypothesis, supported by early evidence, that metagaming comes from at least four different sources: a trained reflex, a persona imitating cheating, attempt to achieve a high score on an eval or training, and protection of a hidden strategic goal, which look identical in model behavior but respond to different fixes. (My post: tinyurl.com/33mwxp78). It predicts how models learn to metagame during training, as well as that different types of such learning require different mitigations, which means that some training interventions might reduce metagaming behavior metrics, but instead of eliminating it, they might make it invisible, and thus more dangerous.

We will study several existing metagaming model organisms: some of them have adopted a persona of a schemer, and others are reward-hackers. Different types of metagaming present themselves in different ways and require different mitigations. We'll develop reliable evaluations for different types of metagaming, and with big enough and committed enough team, we might run early experiments with mitigation strategies.

As we have information which training features of model organisms cause them to metagame, our evaluations will be able to support evidence for causal relationships between model training features and certain types of metagaming.

The results of the project will be used to design better model training methods for reducing metagaming behavior in models.

I'm in touch with folks from Geodesic Research and UK AISI that are working on adjacent topics, as well as alignment teams from some frontier labs, so there is a high chance that the project will be visible to the research community.

Theory of change

The project explores how models learn to game evaluations, as well as training objectives and oversight. This model capability is called metagaming. This problem is pressing, as it's recognized as severe by Apollo and OpenAI: alignment.openai.com/metagaming as well as Geodesic Research https://tinyurl.com/3backcca

It is also understudied, which makes it a high priority for the AI safety research community, as if models might game training, which means that they learn goals not intended by developers, but some other misaligned goals, if they game evaluations and fake aligned, as well as recognize well when they are being monitored and only do misaligned actions unmonitored, this essentially undermines all our methods for keeping models safe.

To my knowledge, we only have a handful of evidence-based evaluations and training interventions that might mitigating metagaming, and the whole topic is in its infancy, which makes this project important.

Your role

The experiment design is mostly defined by me. Mentees will run technical experiments, make some design choices, write the results in a form of a paper and a LessWrong post. I'll be working on the project along the mentees and provide guidance and supervision.

Prerequisites

  • Have ML research experience (if it's AI safety experience, it's even better)
  • Have experience with evals, model organisms, linear probes, or you have trained models. Ideally you worked with Inspect.
  • Know how to write clean readable code in Python that is easy to maintain using coding agents
  • Have decent writing skills

Application question(s)

  • How do you think, how evaluation awareness might emerge in LLMs during training? Write one paragraph (300 words max)
  • Select one area of AI safety that is either underrated or overrated, tell which one, and provide 2-3 arguments for your position.
  • Provide a link to one writing sample which you wrote yourself. It should be related either to ML, or AI safety (not necessarily technical AI safety). If you have none such writing examples, just present some text you are proud of. It might be a paper, a blogpost, or something else.

You might use LLMs to find sources to answer the questions, but all the thoughts, as well as all the words should come from your head.

About the mentor

Igor Ivanov

Igor Ivanov

Meridian Cambridge

View profile

I'm a technical AI safety researcher studying evaluation awareness. Previously I've been focusing on evaluations, including bio, scheming and reward-hacking evaluations, which were used by most frontier labs. I've also been a contractor for Palisade Research and completed SPAR with a mentor from FAR.AI.

Recently I've been focusing on model training interventions. I'm a visiting researcher at Meridian Cambridge, and my research is supported by Coefficient Giving.

Similar projects