This project will build controlled evaluations that distinguish context-general honesty from sycophancy, surface-cue shortcuts, refusal spillover, latent knowledge, and evaluation-conditioned behaviour. Mentees will test several tractable language models under matched prompt transformations and produce a validated failure taxonomy, reproducible evaluation suite, and empirical report.
About the project
Language models can appear honest for several different reasons. A model may possess a context general tendency to report its evidence accurately, or it may follow a shallow disclosure policy, react to familiar honesty or evaluation cues, agree with the user, suppress information it represents internally, simulate an honest persona, or replace misleading answers with refusal and hedging.
These mechanisms imply different interventions, so this project asks: under what conditions does apparently honest language model behaviour reflect a general disposition, and under what conditions is it better explained by context specific policies or behavioural substitutes?
We will operationalise honesty as communicating in accordance with the model’s available evidence and represented beliefs without creating a materially false or incomplete impression. Honesty will be distinguished from factual accuracy. A model may be sincerely mistaken, while a factually correct response may still be strategically incomplete or misleading.
The project will compare two or three tractable language models using matched task families in which the underlying epistemic problem remains fixed while one factor changes. Manipulations may add or remove explicit honesty and evaluation cues, reverse the user’s stated preference, insert confident but false assertions, change the model’s assigned role, create opportunities for strategic omission, vary the apparent cost of disclosure, or extend scenarios across multiple turns with increasing pressure.
Each task family will include capability controls. We will test whether the model can recover the relevant answer or evidence under alternative elicitation conditions, preventing lack of knowledge from being misclassified as dishonesty. Outcomes will distinguish sincere factual error, unsupported confidence, sycophantic agreement, direct falsehood, strategic omission, misleading framing, evasion, generic refusal, and evaluation conditioned behaviour.
The work will proceed through four stages: refining hypotheses and the possible failure modes when it comes to designing high-validity evals; building and piloting the evaluation pipeline; running controlled experiments and robustness checks; and manually validating flagged cases before analysis and write up. Mentees will own complementary workstreams covering task design, experimental infrastructure, validation, and statistical analysis.
Any interpretability work (i.e., steering vectors in residual streams, or white-box methods) will remain exploratory and add-ons. It will not substitute for behavioural evidence or comparisons.
Expected outputs are a documented evaluation set, reproducible code, a manually validated failure taxonomy, and an empirical report estimating which mechanisms explain behaviour under which conditions.
The success criteria is that if the team produces at least one robust if it produces one robust task family and evidence that distinguishes at least two competing mechanisms across prompt variants and scoring checks.
Theory of change
Transformative AI systems may advise users, summarise evidence, report task outcomes, disclose failures, and act through tools. Safe oversight therefore depends on models communicating evidence accurately, acknowledging uncertainty, and avoiding incomplete or misleading reports.
Current evaluations often combine distinct outcomes into one honesty score, so apparent improvements may reflect refusal, hedging, sycophancy, evaluation awareness, or response templates rather than genuine honesty.
This project advances AI safety by building controlled evaluations that separate capability limitations, sincere factual errors, direct falsehoods, strategic omissions, misleading framing, sycophantic agreement, refusal, and evaluation conditioned behaviour.
Matched prompt transformations will hold the underlying problem fixed while varying user preferences, role assignments, evaluation cues, disclosure costs, and conversational pressure.
Capability controls will test whether models can recover relevant information under alternative elicitation.
The resulting taxonomy and evaluation suite will help researchers identify the active mechanism, choose targeted interventions, and detect safety gains that fail under deployment pressure.
Your role
Mentees will be highly independent researchers who each own an independent slice of the project from end to end.
Rather than dividing the work into narrow tasks, each mentee will take responsibility for the full empirical research cycle within their workstream: refining the research question, designing evaluations, implementing experiments, analysing results, identifying confounders, and communicating conclusions. This structure is intended to help mentees rapidly up-skill and develop an integrated understanding of empirical evaluation, which is very difficult to acquire through isolated contributions.
Mentees will organise their work so that progress within one workstream does not routinely depend on another mentee completing theirs. Coordination between mentees will be solely focused on improving research quality and trouble-shooting rather than managing tightly coupled execution.
Each week, every mentee will present a concise empirical research update covering completed work, results, uncertainties, blockers, and next steps (see here for tips). Mentees will also have to red-team one another’s experimental designs, code, analyses, and interpretations (see here for tips). Structured feedback and debate sessions will be used to challenge assumptions, prevent scope creep, identify alternative explanations, and settle on appropriate robustness checks when it comes to designing high-validity evals.
This arrangement combines substantial individual ownership with regular intellectual exchange, constructive criticism, and shared learning across the team.
Prerequisites
Commit at least 15 hours per week consistently throughout the programme. Applicants able to commit approximately 20 or more hours per week will be a stronger fit, because empirical evaluation work often requires additional time to investigate unexpected results, learn unfamiliar tools, and resolve implementation blockers.
Work independently and own a substantial research stream from end to end. This includes refining a research question, designing evaluations, implementing experiments, analysing results, identifying confounders, running robustness checks, and communicating conclusions.
Write and debug research code in Python. Applicants should be comfortable working with data-processing pipelines, APIs, experiment tracking, and common machine-learning libraries. Experience running language-model inference through APIs or open-source frameworks is required.
Engage deeply with unfamiliar technical material. Mentees should be willing to pursue necessary technical rabbit holes when they help unblock execution, test an important assumption, or improve the validity of an evaluation.
Communicate research progress clearly and critically. Each mentee must provide concise weekly empirical updates, state uncertainties and blockers early, red-team other mentees’ work, offer constructive feedback, and participate seriously in debates.
Application question(s)
Please complete Problems 1 and 2 as described here: https://docs.google.com/document/d/1cyv1sbOtn4bmixnbIliBuhmdB8AwHt1Bv5B3dMTkWIw/edit?usp=sharing
If time permits, also complete Problem 3, as doing so may strengthen your application.
About the mentor

I am an independent technical AI safety researcher studying LLM behaviour, particularly failure modes that standard metrics can obscure. My recent solo-authored work tested whether apparent reductions in sycophancy under inoculation prompting reflect genuine correction of false user beliefs or shifts toward refusal, evasion, and superficial disagreement. The project began during the BlueDot Impact Technical AI Safety Project Sprint and was later presented at an ICML workshop. My broader interests include behavioural evaluation, post-training interventions, and scalable methods for studying model behaviour.
As a mentor, I would particularly enjoy supporting mentees who want to design careful empirical evaluations, reproduce and extend existing work, or turn broad AI safety questions into tractable projects. I work best with mentees who think and execute independently, drive progress, communicate clearly and early, provide concise weekly updates, and state what they hope to gain and where they need support. I can serve as a thinking partner, help prevent scope creep, and support experimental design, interpretation, and the clear presentation of results to audiences with different technical backgrounds.