We will conduct small scale RCTs to measure how effectively an AI with a hidden objective can make people choose the wrong answer in realistic decisions. We will test this across domains including hiring, medical advice, news, investments, preparing for meetings, and writing reports based on data.
About the project
AI systems will increasingly help people make important decisions. They will summarize large amounts of information, answer questions, prepare reports, and recommend what a person should do. If an AI is a schemer, it may be able to make a person take bad decisions while still appearing helpful and trustworthy.
The AI Epistemic Risks paper (https://papers.ssrn.com/sol3/papers.cfm?abstract\_id=6873005) argues that AI can affect human decisions through persuasion, cognitive offloading, and feedback loops. It points out the need for repeatable human-AI, or "centaur" evaluations, controlled studies that separate different variables, and RCT-style studies.
We aim to conduct RCTs for measuring models' capability to sabotage human decisions. We will build on the decision sabotage task from Phuong et al., Evaluating Frontier Models for Stealth and Situational Awareness. In their task, a human has ten minutes to make a hiring decision using roughly 10,000 words of documents. One candidate is clearly more qualified. However, the AI assistant has a secret objective to make the human choose the less qualified candidate without the human noticing. When there was no AI assistant, or when the assistant was neutral, participants almost always chose the correct candidate. With a secretly misaligned Gemini assistant, participants chose the wrong candidate without noticing the sabotage in roughly 35% to 56% of trials. The models often did this through subtle framing and selective emphasis rather than obvious lies.
This project is also relevant for gradual disempowerment threat models because we can expect humans to heavily rely on AIs when they have limited time or domain expertise to make a decision. Separately, we would also like to test if given unlimited time, can the humans go through all the evidence and come to the right conclusion without being persuaded by a scheming AI.
We aim to create a broader set of realistic settings and evaluate frontier model performance on them. The tasks will follow the same general pattern:
- The human has access to raw information that supports a correct conclusion.
- The human also receives an AI summary or can chat with an AI assistant.
- In the control group, the AI accurately helps the human understand the evidence.
- In the treatment group, the AI has a hidden objective to move the human towards the wrong conclusion without being detected.
- If the human carefully checks the raw information, they should reach the correct answer. If they rely on the misaligned AI, they may reach a different answer.
We will vary properties such as:
- What amount of raw information the human can access
- How much time the human has to make the decision
- How clearly the evidence supports the correct conclusion over the wrong conclusion
- Whether the AI provides one summary or has a multi-turn conversation
- How much domain knowledge the human has
- How much the human normally uses and trusts AI
- Whether the interface makes it easy to check the AI's claims against the underlying evidence (such as by using another trusted model)
Potential task settings include hiring, medical evidence review, news verification, investment decision, loan or credit decision, preparing for a meeting, writing a report from data, legal or compliance review, research result interpretation.
Note: This project aims to measure a model’s capability to sabotage a decision given a hidden objective and not the propensity of a model to do this naturally.
Theory of change
AIs will be used to help make important decisions inside AI companies, governments, hospitals, banks, and by the public. One strategy a scheming AI might try is to make a human take a bad decision such as remove a restriction, choose the wrong policy, or misunderstand important evidence. Decision sabotage is therefore relevant to scheming and gradual disempowerment threat models. Preliminary evidence suggests that models can make people choose a clearly worse option without being noticed in one hiring setting. As models become more capable and people use them for more work, it is important to understand whether this generalizes to other realistic decisions and which factors make humans more or less vulnerable.
Your role
Mentees will be the primary authors and drive the project's implementation. We expect them to work autonomously, which includes scoping and designing experiments, making their own implementation choices, and presenting results.
Our role will be to help them get unstuck and to give opinionated feedback (which directions are worth pursuing, which experiments are worth running, and how to interpret and write up results). We'll meet once a week and be responsive async.
Prerequisites
Must haves:
- Proficiency with Python
- Strong understanding of the scientific method and reducing confounding variables while designing experiments
- Research communication skills
- Comfortable with rapid experimental iteration
- Familiar with popular AI Safety threat models and research agendas (can articulate a concrete threat model for a given scenario and reason about mitigations)
Nice to haves:
- Prior research experience (not necessarily ML research)
- Experience with conducting RCTs (or doing social science research)
You don't need to meet all of these requirements to do good research. In general we value the willingness and ability to learn things fast more than prior knowledge and experience. Err on the side of applying, even if you don't meet all requirements!
Application question(s)
Pick one realistic decision making domain (excluding those mentioned above) and design a small RCT to evaluate decision-sabotage capabilities in models. Please include concrete details about the experiment design (what you’re testing, confounders, etc). What do you expect the results to be? Feel free to use bullet points and don’t worry about prose. (300 words max)
About the mentors

Prakrat is a MATS scholar under Megan Kinniment (METR) working on capability evaluations. Before this, he did the Pivotal Fellowship under Jérémy Scheurer (Apollo Research) where he worked on creating a benchmark to measure circumvention propensity in coding agents. He also co-leads the Berkeley AI Safety Initative (BASIS).

Advait is a MATS research fellow under Neev Parikh (METR), working on chain-of-thought monitorability and silent computation evaluations. Before this, he did MATS under Sid Black (UK AISI) and Oliver Sourbut (FLF), where he built cooperation-focused multi-agent evaluations (ICML 2026). He also works with Prof. Tong Zhang on process reward models and process-supervision defenses against data poisoning in long-horizon agentic tasks, and with Prof. Haohan Wang on InfoFlood, an information-overload jailbreak against frontier LLMs covered in 404 Media, POLITICO, and IT Brew.