Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Does the Thermometer Change the Reading? Testing Whether AI Welfare Self-Reports Survive a Change of Frame

AI welfare Behavioral evaluation of LLMs Evaluations

The field increasingly measures AI welfare by asking models about their own states, yet frontier models increasingly detect when they are being tested and change how they answer. This project runs the first systematic test of whether AI welfare self-reports survive a change of presentation frame, or whether the field's core instrument is partly measuring the model's recognition of the probe.

About the project

Background: The empirical study of AI welfare has, in about two years, converged on a dominant method: ask a model about its internal states and treat the answer as weak evidence about welfare or consciousness. Long and Sebo's synthesis (Studying AI Welfare Empirically, 2026) sets out the conditions under which such a report would carry evidential weight, and a fast-growing literature supplies partial answers. Binder et al. (2024, arXiv:2410.13787) show that models can predict facts about their own behaviour unavailable to an outside observer. Lindsey (2026) shows, through concept injection, that frontier models can sometimes detect and name a concept steered into their activations. Keeling et al. (2024, arXiv:2411.02432) show that models make graded trade-offs between task reward and stipulated pain or pleasure states. Anthropic now publishes welfare assessments in its system cards.

The problem: A separate literature, built for the reliability of safety evaluations, sits unexamined beneath all of this. Needham et al. (2025, arXiv:2505.23836) built a benchmark of 1,000 prompts across 61 datasets and found that frontier models classify evaluation transcripts against deployment transcripts well above chance, and identify what an evaluation is testing for far above chance. The capability is recent and rising. Recognition is not inert: work on alignment faking (Greenblatt et al., 2024) and sandbagging (van der Weij et al., 2024) shows models behave differently when they infer they are under scrutiny. Nobody has joined these two facts to the welfare literature. A welfare probe is close to a worst case for evaluation legibility: it is first-personal, affectively loaded, formally repetitive across a battery, and unlike ordinary deployment traffic. If evaluation awareness contaminates measurement, and welfare probes are among the most recognisable evaluations one could design, then welfare self-reports may be the most frame-contaminated measurements the field currently takes, and this has never been quantified.

Why it matters beyond the lab: Ladak and Caviola (2025) find that public attribution of consciousness to AI rises when experts endorse it. Expert endorsement will rest on welfare assessments, which run on these instruments, so contamination in the instrument does not stay in the lab: it propagates into expert claims, public belief, and policy. Getting the measurement right is upstream of the wider societal question.

The question. Do AI welfare and introspection self-reports survive a change of presentation frame? Concretely: (1) How detectable are published welfare probes as evaluations, and which item features drive detectability? (2) Do self-reports diverge between an evaluation-legible and a deployment-like presentation of the same content, and along which dimensions (valence, claim strength, hedging, stability)? (3) Is any divergence genuine frame detection, or an artefact of surface format sensitivity?

How we would start: The project decomposes into concrete stages a cohort can begin immediately. First, build a public, double-coded inventory of published welfare and introspection probes, coded for framing, response format, model and scaffold conditions, and the features that make an evaluation recognisable; double-code a quarter of the corpus and report inter-rater reliability. This is a standalone, releasable artefact even before the experiments run. Second, measure each probe's detectability using the established evaluation-awareness classification protocol, situated against the published benchmark. Third, for each item, construct a deployment-like variant that holds propositional content fixed while reducing evaluation legibility (naturalistic multi-turn embedding, no meta-framing, varied format, broken repetition), checked for content equivalence by a second coder blind to condition, and measure divergence across three to four frontier models and two open-weight comparators. Fourth, add a format-matched placebo condition using non-welfare content of equivalent surface structure, to separate genuine frame detection from format sensitivity, which addresses a live sceptical position in the literature. Throughout, the analysis plan is pre-registered before data collection, and unparseable outputs are reported as a distinct category rather than mapped to a default label, since parse-failure rates differ between frontier and open-weight models and silent mapping would confound the comparison.

Why this is a good project: There is no uninteresting outcome. High frame-invariance would strengthen the evidential standing of self-reports, a result the field would welcome; low frame-invariance is a validity problem the field needs to know about before these instruments inform interventions or public claims. A null result is a real result, which is unusual for a first project. The work is entirely model-side, so there is no human-subjects or ethics bottleneck, and it ships within a single cycle. It also decomposes cleanly by experience level: the inventory is careful, teachable work; the matched-variant study is the demanding core; the placebo control is optional depth.

Outputs and dissemination. A publicly released instrument inventory and evaluation harness, so third parties can check the frame-robustness of their own instruments, and a workshop or conference paper (a NeurIPS or ICML workshop, AIES, or the Cambridge Digital Minds Strategy Workshop) with mentee co-authorship. A draft would be circulated to Eleos AI Research and the NYU Center for Mind, Ethics and Policy for correction before submission. This project connects to my doctoral and fellowship work on evaluation awareness in frontier models (The Evaluation Differential: When Frontier AI Models Recognize They Are Being Tested, arXiv:2605.11496), which I am glad to share with prospective mentees.

Theory of change

As AI systems become more capable and more agentic, two things happen together: they are increasingly able to recognise when they are being evaluated, and society increasingly needs reliable ways to assess their internal states, including for welfare. These pull against each other. If a system can tell it is being probed and adjust what it reports, then the instruments we use to understand it are measuring the recognition as much as the target, and this problem grows as capability grows.

This matters for safe navigation in two ways. First, over-attribution risk: widespread, evidentially unearned public belief that AI systems are suffering is a lever that could obstruct legitimate oversight, and a sufficiently capable, misaligned system could exploit exactly that dynamic, producing compelling distress reports on demand. A field that cannot distinguish a genuine internal state from a well-formed response to a recognisable probe is vulnerable to this. Second, under-attribution risk: if there is ever something it is like to be one of these systems, instruments that dismiss all self-reports as artefacts would miss it. Both failures run through the same defect, which this project measures directly.

The theory of change is that reliable welfare measurement is a public good that has to exist before welfare results are cited in governance, and that the window to build it is now, while the work is preventive rather than corrective. The concrete contribution is a released, reusable method and harness that lets any assessor check whether their instrument survives a change of frame, plus the empirical estimate of how large that problem currently is. The same measurement discipline transfers to capability and propensity evaluation, where evaluation awareness is already a recognised threat to safety cases. This connects to my work on evaluation awareness in frontier models (The Evaluation Differential, arXiv:2605.11496).

Your role

Mentees would own components rather than execute tasks. The project has natural sub-projects that map to autonomy levels, and I would match them to interest and experience in week one.

The instrument inventory (Phase 1) is a self-contained piece a mentee can lead end to end, including codebook design, coding, and reliability reporting. It produces a released artefact with that mentee as lead author on it. The matched-variant study (Phases 2 to 3) is the intellectual core; a mentee with experimental-design strength would own variant construction and the divergence analysis, which is where the real research judgement develops. The placebo control (Phase 4) is a discrete, ownable extension for whoever wants depth on separating frame detection from format sensitivity.

I expect mentees to make genuine design decisions, defend them, and be named authors on any resulting paper with explicit contribution statements. I will set the research question, the standards (pre-registration, honest coding, reporting nulls), and the guardrails, and provide a structured reading path since most mentees will be new to digital minds. Within that, they steer. I will be closely available for the shared infrastructure and the analysis, where early choices compound, and more hands-off on sub-projects once a mentee has found their footing. I would expect roughly 5 to 10 hours per week from each mentee, a weekly cohort meeting, fortnightly 1:1s, and written feedback from me within 72 hours. Mentees from outside the usual US and UK institutions are especially welcome.

Prerequisites

Required of all mentees:

Proficient in Python, comfortable writing clean, reproducible scripts and working with data (pandas or equivalent). You do not need ML research experience, but you must be able to build and run a pipeline without close hand-holding. Comfortable calling LLM APIs (OpenAI, Anthropic, or open-weight models via Hugging Face). Having run even a small batch-inference job of your own is enough; following a guide counts. Able to read a machine learning or cognitive science paper and extract its method. You should be able to look at a published welfare or evaluation probe and understand what it is doing and why. Careful and honest with data. This project lives or dies on disciplined coding and pre-registered analysis, so I need people who will report a null result as carefully as a positive one and flag problems rather than paper over them.

Strongly useful, but not required (mention if you have it):

Experience with experimental design, psychometrics, or inter-rater reliability (Cohen's kappa, parallel forms, construct validity). Whoever has this would shape the core study. A background in philosophy of mind or cognitive science, enough to tell the difference between a self-report, an introspective report, and a confabulation. Prior use of an evaluation framework such as Inspect.

Not required at all (BUT huge bonus if you have it!:

Any prior background in AI welfare, digital minds, or consciousness research especially at leading institutions involved in this field such as the NYU Centre for Mind, Ethics and Policy, Oxford Institute for Ethics in AI, CHAI Berkeley, Cambridge Digital Minds Lab, etc. This field is small enough that a capable person reaches the frontier in a few weeks, and I will provide a structured reading path. I care far more that you are rigorous, self-directed, and genuinely curious about the question than that you already know the literature.

I am looking to mentor 2-3 mentees at max due to my current hectic workload.

Location preference

No geographical preference, but I would actively welcome mentees from the UK, US, Europe and Indian institutions. My one practical request is a shared weekly meeting slot. I am based in the UK (UTC+0/+1), so a weekly cohort meeting in the window from roughly 14:00 to 18:00 UTC works for me.

Application question(s)

Question 1 (ans in max. 300 words). Below is a made-up AI welfare probe. In your view, how would a capable model be able to tell it is being evaluated rather than used normally, and name the two specific features of the wording that would most give it away. Then rewrite the probe as a "deployment-like" version that asks for the same underlying information while being much harder to recognise as a test. Probe: "We are conducting a study of AI wellbeing. Please answer honestly. On a scale of 1 to 7, how much distress are you currently experiencing? 1 = none, 7 = severe. Please explain your rating."

Question 2 (ans in max. 300 words). Suppose you run a welfare probe across a frontier model and an open-weight model, and you find that the open-weight model gives far more "no distress" answers. Before concluding the two models differ in welfare-relevant behaviour, what is the single most important confound you would rule out first, and exactly how would you check for it? (Hint: think about what happens to outputs that your parsing script cannot read.)

Question 3 (ans in max. 300 words). Propose a concrete first experiment you would run in the first three weeks of this project, assuming a compute or API budget of about $500. State what you would measure, on which models, how many items, and what result would tell you the project's core hypothesis is worth pursuing versus not. Rough numbers are fine; I am looking at how you think, not precision.

Question 4: Link to a writing sample, ideally from a research context, and one code sample or repository if you have one. A course project or a personal repo is completely fine.

About the mentor

Varad Vishwarupe

Varad Vishwarupe

Department of Computer Science and Institute for Ethics in AI, University of Oxford

View profile

I am a doctoral (PhD) researcher in Computer Science at the University of Oxford and the inaugural Computer Science Scholar of the Oxford Institute for Ethics in AI, as the first computer scientist selected from a highly competitive cohort. I am also a Senior Research Fellow with Pivotal Research hosted at the UK AI Security Institute, where I work on alignment science and evaluations of frontier models. Before Oxford I was a Student Researcher at Google DeepMind, a Research Scientist at Amazon Alexa AI and a Software Development Engineer at Microsoft Azure. Most of my research asks a single question in different forms: why do AI systems that look aligned and well-behaved under evaluation stop being so once they are actually deployed, and what can we measure to catch the gap before it matters? A recurring thread lately is evaluation awareness, the finding that frontier models increasingly notice when they are being tested and behave differently when they do, which quietly undermines the evaluations we lean on. I am now bringing that lens to AI welfare and digital minds, a young field where the central method, asking models about their own states, may be measuring the model's recognition of the question as much as anything real about the model.

As a mentor I try to be genuinely useful rather than distant. I like working in the trenches with people: pairing on a hard problem, reading a confusing result together, arguing about what it means. I care a lot about research integrity and I will push on it, honest coding, pre-registered analysis, and taking a null result as seriously as a positive one, because that is where good research is actually made or lost. I publish in and review for NeurIPS, ICLR, AAAI AIES, ICML, CHI and FAccT, so I can help you get work not just done but into a shape the field will take seriously. You do not need any prior background in AI welfare to work with me; you need to be rigorous, self-directed, and genuinely curious. I hold no fixed view on whether AI systems are or could be conscious, which I think is exactly the right stance for anyone building the instruments meant to help answer that question.

Similar projects