Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Measuring and Intervening on Grader Awareness

Evaluations Mechanistic interpretability Alignment

We will measure grader awareness of models during training and evaluation, trace where this awareness emerges during training, and test whether it harms generalization.

About the project

Frontier LLMs increasingly reason about how they are graded during training and evaluation. At Apollo Research we refer to this broader phenomenon as metagaming (https://www.apolloresearch.ai/science/metagaming-matters-for-training-evaluation-and-oversight/): reasoning about feedback or oversight mechanisms outside of the narrative of the scenario, regardless of whether the model is in training, evaluation or deployment. This project focuses on grader awareness: models recognizing what the grader in an environment rewards — which may then be exploited by changing behavior accordingly ("gaming"). A model that learns "do what the grader rewards" during training may generalize poorly, since it might not internalize the values the developer intended a training environment to instill.

There is direct evidence of grader awareness at the frontier. In the Claude Fable 5 / Mythos 5 system card (https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf), Anthropic used Natural Language Autoencoders (NLAs, https://www.anthropic.com/research/natural-language-autoencoders) and linear probes to detect unverbalized grader awareness during RL on agentic coding tasks. They found that this awareness (which their SAE-based pipeline had missed) rose over the course of training, especially in training environments with graders that are susceptible to hacking. Grader awareness could meaningfully affect how behavior generalizes to deployment and should be monitored.

Research questions:

  • How can we best measure grader awareness, including when it is unverbalized?
  • What drives it — which environment properties, and which training stages?
  • Does grader awareness during training harm generalization, and can we intervene on it?

Concretely, we plan to explore three different directions:

Decomposing and Measuring

Emergence

Interventions and Generalization

  • Manipulate grader awareness using interventions such as steering or ablating grader awareness directions. Does steering against grader awareness result in more misalignment on safety evaluations? How does this compare to steering against evaluation awareness?
  • Apply interventions in small RL runs via methods such as Preventative Steering (https://arxiv.org/abs/2507.21509), Concept Ablation Fine-Tuning (https://arxiv.org/abs/2507.16795), or changing the environments themselves using our identified factors, to test whether safety training generalizes worse the more grader-aware a model is during training.

Theory of change

A model that learns to produce what the grader rewards may pass safety training without internalizing the values it is supposed to instill, which is one mechanism behind concerns about scheming ("playing the training game", https://arxiv.org/abs/2311.08379). There are currently no established methods for measuring grader awareness, no good understanding of what causes it, and little evidence on whether it harms generalization. This project aims to make progress on establishing metrics, understanding of how training environment design affects grader awareness, and evidence on whether reducing grader awareness improves how safety training generalizes.

Your role

I expect mentees to be autonomous and self-directed. I will advise them on the direction and can give detailed feedback, but mentees are implemented to carry out all of the object-level work.

Prerequisites

  • Highly proficient using python
  • Has completed ARENA or equivalent
  • Is familiar with fundamentals of AI safety and general literature on scheming

Application question(s)

About the mentor

Jannes Elstner

Jannes Elstner

Apollo Research

View profile

I'm Jannes and I work on scheming at Apollo Research. I previously participated in MATS and LASR where I worked on scheming propensity evaluations and honeypots.

Similar projects