Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Threat Models for Recursive Self Improvement

AI strategy AI security AI control

Existing agentic threat models assume external attackers, not a privileged model building its own successor. This project maps the attack surface of automated AI R&D into a structured taxonomy with worked scenarios for policymakers and empirical researchers.

About the project

AI labs are increasingly automating more of their own research and many experts think that within ~two years, agents will be picking research directions, running experiments, writing evals, building harnesses, and training their successors. The early stages of this is already happening and OpenAI recently claimed it's largest model GPT-5.6 Sol helped post-train the Luna sized model.

The coming recursive self improvement loop creates a large, underexplored attack surface. When a model helps build the next generation, a lot can go wrong. This includes training sabotage, poisoned evals and reward signals, monitors colluding with the systems they're supposed to watch, unauthorized compute use and weight exfiltration to name a few. Each of these has different preconditions, different actors, and leave different detectable traces. There are existing threat catalogues like OWASP ASI and MAESTRO, but they're built around external attackers going after deployed systems, not a privileged model training its own successor. That gap is what this project tackles.

Mentees will break the AI R&D loop down into its stages and access boundaries, then build a schema that puts these threats into a common structure. From there, they'll populate it across the full threat family and pick one or two threats to develop into detailed scenarios, using those deep dives to pressure-test the schema and refine it.

The end result is a paper providing a reference taxonomy with a few worked scenarios, aimed at the people doing research downstream, plus an honest accounting of what the current literature misses. I think this work could be highly impactful for both policymakers and empirical researchers to build on.

Theory of change

This work aims to inform the wider AIS research community and ideal allow researchers to design mitigations based on the threat models we devise.

Your role

Mentees are expected to independently lead the research and proactively suggest next actions and potential news directions. I have an initial plan, high level ideas to pursue and have done some literature review.

Prerequisites

  • Prior experience writing a research paper, technical report, or long-form analysis (public posts count).
  • Ideally have already graduated from a bachelor's program (exceptional candidates will be considered on a case by case basis)
  • Familiarity with core AI safety concepts, especially AI control, scheming/deceptive alignment, and dangerous capability evals — e.g., you've read and can discuss work like Redwood's AI control agenda, "Sleeper Agents," or the AI R&D threat models in frontier lab safety frameworks (RSPs/preparedness frameworks).
  • Strong analytical writing skills: able to take a fuzzy conceptual space and produce clear, precisely structured prose. You should be comfortable having your writing critiqued and revised heavily.
  • Basic technical understanding of the ML training pipeline (pretraining, fine-tuning, RL, evals), enough to reason about where in the pipeline an attack could occur and what access it would require. You don't need to have trained models yourself.

Location preference

I can only accept mentees in timezones between PST and CET as I'm based in PST

Application question(s)

  1. OWASP's Agentic Security Initiative enumerates threats to agentic systems, but assume external attackers against deployed systems. Pick one threat from the ASI and describe what happens to it when the adversary is instead a privileged model training its successor. Does it transfer intact, transfer with modified assumptions, or stop applying? What does that tell you about reusing existing catalogues here? (250 word limit)

  2. This project is all writing and reasoning, there are no experiments to check your work. That means if your threat model is bad, nothing outside of you will flag it. There are two common failure modes:

  • It's unfalsifiable. It makes claims so vague or untestable that no evidence could ever prove it right or wrong.
  • It's over-hedged. It's so full of caveats and "maybes" that it can't actually tell anyone what to do.

For this question, do two things: (1) Give one concrete example of each failure, specifically in the context of automated AI R&D. (2) Explain how you'd catch yourself mid-draft. What warning signs would tell you you're sliding into one failure or the other? (250 words max)

About the mentor

B

Benjamin Arnav

NYU

View profile

I'm a full-time contractor on OpenAI's Frontier Evals team where I focus on science of evals. My research has typically investigated chain-of-thought monitorability, AI control in multi-agent systems and how to measure RSI progress.

Similar projects