Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Empirical Evaluations of Situational Awareness

Evaluations AI control Behavioral evaluation of LLMs

How should we evaluate eval awareness? Under what scenarios should we expect eval awareness to appear? What are model organisms of eval awareness?

About the project

Situational (Evaluation) awareness is a real concern. However, we currently only have an empirical observation of eval awareness happening, but we do not have an understanding of why eval awareness occurs. Intuitively, why should eval awareness appear in the first place, if all conversations with models are simulated to be as realistic as possible (models shouldn't be eval awareness when they are interacting with real users)?

This project aims to further our understanding of evaluation awareness. This project will begin by conducting an empirical evaluation on when, how, and where evaluation occurs in different tasks for different models. This project might also involve building model organisms of eval awareness. And if time allows, looking at mitigation algorithms such as training and steering.

Theory of change

Clearly eval awareness is undesirable. If models are eval aware, that affects not only safety evals, but also capability evals as well. For example, Apollo found that removing verbalized eval awareness tokens in the models’ chain-of-thought also increased scheming behavior (from 13% to 24%), and adding verbalized eval awareness tokens to the chain-of-thought decreased scheming.

This project will allow researchers to have a much better understanding of how and why eval awareness appears in practice, and potentially how to avoid it and mitigation practices. As far as I am aware, many ongoing research projects face eval awareness as a difficulty, but there has not been many published research about this yet.

Your role

Mentees will be responsible for carrying out empirical investigations of how and why eval awareness appears in practice (with my help). See proposal for details.

Prerequisites

Highly proficient using Python, bonus if you have strong understanding of frontier LLMs and have worked with API inference.

Location preference

London is great for in person meetup but no requirements

Application question(s)

What is one paper about eval awareness that you think I haven't read? Talk about that paper in detail.

About the mentor

Qiyao Wei

Qiyao Wei

University of Cambridge and MATS

View profile

Qiyao (Chi-Yao) Wei is a current PhD student in the Maths department at University of Cambridge. He is also a current MATS scholar working with Francis Rhys Ward on CoT monitorability.

Similar projects