Frontier models hold operationally useful dual-use knowledge, and cheap task-decomposition attacks can pull it out while slipping past per-query defenses. This project measures how much real harm uplift these attacks produce (functional success, not just refusal) and builds the measurement and detection tools providers need to keep pace.
About the project
Theory of change
Models grow in dual-use knowledge alongside capabilities, and cheap extraction methods like task decomposition put that knowledge within reach of low-resourced actors, including terrorist actors already documented using it (Juelich 2026), while being difficult to monitor and mitigate. This raises the probability of catastrophic misuse, and we’ve identified two problems with standard operating procedures - providers struggle to detect task-decomposition attacks with current per-query defenses, and the field measures refusal rather than whether elicited output actually works. We see AI misuse in dual-use domains as an under-explored yet extremely significant risk, and we intend for this stream to help advance safety towards preventing practical harm uplift from AI. As such, we support many directions within this domain, and are particularly interested in studying the following:
- Better jailbreak eval techniques like functional success rate (over attack success rate only counting refusal bypass) can help inform a proper harm uplift measurement
- Alternative, benign, and publicly releasable benchmarks for harm uplift, that don’t contain hazardous information but reflect a strong relationship towards measuring harm
- The detection and mitigation methods for these threat vectors, enriching the provider toolkit as these attacks scale and can also apply to oversight mechanisms for increasingly capable models
- Improving automated red-teaming pipelines to strengthen defenses
This work advances AI safety by (1) accurately measuring and forecasting the risk of catastrophic AI misuse, (2) reducing the risk of catastrophic AI misuse and (3) improving detection/oversight mechanisms for increasingly capable models (this is important; see July 2026 OpenAI rogue deployment incident).
Your role
Each mentee can own one direction (above) and drive it. You pick the specific approach within that direction, run the experiments, and produce the analysis and writeup. John and Max set the overall agenda, hand you the prior red-teamer and datasets as a starting point, meet weekly to engage with your results and unblock you, help scope toward something publishable, and get work done if necessary. The goal is to end up with something publishable which you will have authorship on.
Prerequisites
Required: high proficiency in Python. Comfortable calling LLM APIs and running open-weight models locally. Familiar with core AI safety and red-teaming concepts (jailbreaks, refusal training, evals). Able to work and deliver independently, and have clear reasoning transparency! Strong plus for specific directions: domain literacy in chemistry, biology, or cybersecurity for the dataset directions, experience designing or running AI evaluations, prior red-teaming or capability-elicitation work.
Application question(s)
Propose an initial experiment for one of the directions above. Assume a $1,000 compute budget and four weeks. State your hypothesis, what you would measure, and what result would change your mind. (200 words, 100-150 expected)
Propose directions that would potentially falsify this set of research questions and goals. Make sure your reasoning is transparent, and you provide promising evidence to back your claims. (200 words)
About the mentors

Hi! I'm a research fellow at MATS focused on dangerous capability evaluations and scalable oversight for multi-agent systems. Prior to MATS, I spent four years at Amazon Web Services, developing AI products for government agencies, universities, healthcare organizations, nonprofits, and students, with work including scalable LLM evaluation, multi-agent systems for critical sectors, and advising AI adoption in the public sector. I have a background in mathematics and computer science, and I'm currently working on various dangerous capability benchmarks and want to work with motivated researchers to study risk factors I believe are under-represented.

I am a MATS 10.0 research fellow working on mitigating catastrophic dual-use AI risk. I am currently approaching this from 2 directions. The first direction is through technical governance/threat modelling to inform policymakers with ideas derived from broad technical expertise. Secondly, I am working on improving model tamper resistance methods to make open-weight models safer, and thus, the broader AI ecosystem safer. My prior work includes improving hallucination detection in LLMs and doing similar technical governance work in regards to AI generated image/video deepfake abuse.