Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Towards rigorous Alignment Evaluations.

Evaluations Behavioral evaluation of LLMs Alignment

Evaluations, in principle, are most meaningful if they are accurate. In this project, we will audit evaluations that claim to measure traits such as generalisation, identify current failure modes, and help build evaluations that are more accurate.

About the project

Research Direction: Alignment evals try to measure propensities such as misalignment, scheming, sycophancy, and dishonesty. Their results feed system cards, safety cases, and deployment decisions. But propensities are much harder to measure well than capabilities. There is usually no ground truth (an LLM judge scores the behaviour), the constructs are vague (two published sycophancy evals rank the same models in opposite orders), and the behaviour itself is not a stable trait of the model: it shifts with context, framing, and the persona the model adopts, including whether it believes the situation is real. A science of evaluations is forming around problems like these (Apollo's agenda, Weidinger et al., UK AISI). This project asks its central questions about propensity evals: what would it take to trust these numbers, and how close do current alignment evals come? Starting Point: We begin with a deep audit of a single eval, most likely the emergent-misalignment measurement pipeline (Betley et al.): its protocol is reused across many follow-up papers, so findings propagate, and the fine-tuned models are open, so it is cheap to re-run. We build primarily on the two best-fitting existing instruments, the Agentic Benchmark Checklist and the ETH misalignment-evidence checklist, supplemented by construct-validity criteria from Bean et al.. The audit has a conceptual pass and an empirical pass. The conceptual pass grades the eval against the checklists: is the construct clearly defined, do the scenarios plausibly elicit it, do the claims match the evidence? The empirical pass reproduces the headline result in Inspect, then perturbs what the conceptual pass flags (judge choice, output format, scenario framing, score thresholds) to quantify how robust the reported numbers are. Early audits of individual slices (judge circularity, framing effects) suggest there is a lot to find. The pilot calibrates our checklist and tooling; from there we extend the audit to further suites such as Agentic Misalignment, MASK, and competing sycophancy evals. From the audits we aim to distill recurring flaws into practical guidance, a step towards a methodology for propensity evals that is grounded in documented failures. There is a genuinely open conceptual question here too: no existing instrument tells you how to quantify the validity properties that matter most for propensity evals, such as whether the model detects the eval, whether the judge is independent of what it grades, whether the measured trait is stable across framings, and how to power studies of rare behaviours. Fields that have long measured latent, fakeable traits like psychometrics are a useful reference point for this part. Natural extensions depending on interest include building an exemplar eval or testing whether eval behaviour predicts behaviour in deployment-like conditions. We're also very open to mentees who arrive with their own ideas in this space.

Theory of change

TOC: We view evaluations as the most important technical intervention towards a potential slowdown as well as to pass government policies and inform policy makers of AI threats. We also view evaluations as critical to determining whether our current alignment techniques are "working". However, neither of these effects can be realised if policy-makers nor lab leads "trust" our evaluations: Evals that understate misalignment provide false assurance for deploying dangerous systems. Evals that overstate it get debunked, eroding the credibility of alignment research with the policymakers and lab decision-makers it needs to reach. We therefore view this work as critical to securing that current alignment evals are taken seriously and measure what we want them to measure. Some work doing this with capabities evals at UKAISI is: https://arxiv.org/abs/2507.02825, who are interested in our project proposal.

Your role

At a high level, we expect mentees to be quite independent in setting up and running experiments, although we will be happy to help out with this towards the start of the project. Towards the first few weeks of the project we will have very concrete projects to start with, but would be very excited about mentees also thinking critically about areas of extension for this project with us, as although we currently have extension ideas we would be keen to mentor on other aspects of this project too.

Prerequisites

We would like mentees who are able to run small experiments on LLMs in Python and think through problem critically. Skills that are helpful for this:

  • Python
  • Knowledge of transformers/LLMs
  • Previous experience with LLM evaluations (not required)

We're happy to take on people who learn quickly, if they're ready to do a lot of technical learning independently! On the off chance that you have psychometrics experience this would also be extremely valuable (although certainly not needed).

Application question(s)

What are some current gaps in evaluations? (If you are unsure about the current eval landscape identify a gap in this suite of evaluations: https://metr.org/time-horizons/ ) [we expect 150 words approximately for this answer]

[Optional] Link to a writeup you have already done - this could be a technical memo or a paper. We are most interested in finding out how you think. If you do not have a writeup to link (or want another question to answer), answer the question: What is an area in technical AI Safety that you think a disproportionately low amount of people are working on compared to it's importance? Why? [Try and keep answers below 150 words]

About the mentors

Krish Sen

Krish Sen

ERA AI/University of Oxford

Krish Sen is a current ERA fellow working with OpenAI's model spec evals team on designing new model specification evals and thinking about which values are most important to test. His previous AI safety work includes some work on mechanistic interoperability, but mainly beneficial RL and generalisation studies. He is interested in a fairly wide variety of AI safety projects, but prefer projects that fit well into the broader picture of how to make AI safety go well. He has also done some field building work co-founding Oxford AI Safety Initiative's policy division and recently TA'ed an ARENA style curriculum (ARBOx).

Rahul Marchand

Rahul Marchand

Oxford University

View profile

I'm a final-year engineering student at Oxford working on AI safety evaluations. I was first author on an ICML 2026 Oral paper, a benchmark built with the UK AI Security Institute that measures frontier LLM agents' ability to escape container sandboxes. I'm currently working at the Oxford Witt Lab on goal misgeneralisation. Before moving into safety, I spent two years doing ML research in quantum computing, co-first-authoring two papers on automated quantum device tuning.

I also enjoy teaching, most recently writing and delivering the RL lecture series for ARBOx, the Oxford AI Safety Initiative's ARENA-based bootcamp. We'll meet weekly, and I'll give feedback on writing and code in between. The goal is that we end up with a workshop or conference paper.

Similar projects