Undetected evaluation awareness during alignment audits poses significant risks. This project investigates the automated, iterative refinement of evaluation environments, supervised by white box eval awareness monitors, to enhance realism and audit reliability.
About the project
Undetected evaluation awareness during alignment audits poses significant risks, as models often identify testing conditions and adjust their behavior toward more compliant, but potentially misleading, responses. Since future systems may lack legible chain-of-thought or explicit verbalized awareness, standard black-box auditing is becoming increasingly unreliable, necessitating the development of robust white-box methods capable of detecting these unverbalized manifestations before deployment.
This project investigates the automated, iterative refinement of evaluation environments using feedback loops from white-box monitors—such as probes, NLAs, J-lenses, and SAEs—to detect this awareness and quantify discrepancies with verbalized forms. Expanding upon traditional transformer architectures, this research will also study the prevalence and mechanics of unverbalized evaluation awareness in diffusion and latent reasoning models, ultimately establishing a more comprehensive framework for audit reliability and realistic capability evaluations.
Relevant papers: https://www.goodfire.ai/research/verbalized-eval-awareness-inflates-measured-safety# https://alignment.anthropic.com/2026/petri-v2/ https://arxiv.org/pdf/2502.21074 https://www.lesswrong.com/posts/YGAimivLxycZcqRFR/can-we-interpret-latent-reasoning-using-current-mechanistic https://arxiv.org/abs/2510.20487 https://arxiv.org/abs/2505.23836 https://arxiv.org/abs/2507.01786 https://arxiv.org/abs/2605.26438 https://arxiv.org/abs/2605.30322 https://www.lesswrong.com/posts/7qBTcE3jqQFTuzssE/realistic-evaluations-will-not-prevent-evaluation-awareness https://meridianlabs.ai/blog/posts/introducing-petri-3/
Theory of change
Creating an iterative eval refinement method that is robust to unverbalized eval awareness directly reduces the risk of deploying deceptive or misaligned systems.
Your role
Mentees will be the technical leads of this project. In particular, writing code, running experiments, formulating and testing hypotheses and writing up the project at the end.
Prerequisites
- Proficient in Python, ideally proficient in TransformerLens and/or PyTorch.
- Familiarity with mechanistic interpretability techniques (e.g., SAEs, probes, or J lens).
- Experience working with transformer architectures; familiarity with diffusion models or latent reasoning architectures is a strong plus.
Location preference
No preference
About the mentor

I'm a DPhil student at the University of Oxford and Research Associate with Neel Nanda at Google DeepMind, working on interpretability and data attribution. I'm interested in exploring and de-risking AGI safety techniques that survive the loss of legible chain of thought. Previously I worked on data attribution tooling during an internship at Cohere and have worked with UK AISI on agentic scaffolds for Inspect.