Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

MM-AutoTrainBench: A Platform for Studying Risk from Autonomous Multimodal Systems Research

Evaluations AI control AI security

Rapid acceleration in agent-automated multimodal post-training capabilities could pose significant existential risk if not performed safely due to the major improvements in autonomous research labs and embodied systems it would enable. Currently, no benchmark studying agentic multimodal post-training exists, and thus we both have no concrete understanding of how close we are to such a takeoff and have no good setting for studying security against misaligned agents performing multimodal training tasks. This study aims to address both of these timely problems through the creation of a platform for studying risk from autonomous multimodal systems research.

About the project

Research Question: How effectively can current frontier agents autonomously complete multimodal post-training tasks? How do these capabilities scale across recent model releases, model sizes, and research labs?

Secondary Objectives:

  1. Determine whether or not these frontier agents can also covertly complete malicious side tasks aimed at subtly misaligning the multimodal models they are working on.
  2. Provide a setting for the study of monitoring and control protocols to improve the security of autonomous multimodal post-training efforts.

Methodology:

  1. We will begin by reimplementing the agent harnesses used in PostTrainBench and configuring them for our hardware. We will aim to test frontier agents from Anthropic, OpenAI, Google, and the open-source community (Kimi, Qwen, GLM, etc.).

  2. We will then identify a series of base models and representative multimodal post-training tasks for those models. In particular, each post-training task will be represented by a benchmark on which the model should aim to increase performance (without reward hacking). We will aim to have the base models represent diverse leading multimodal architectures, while tasks are selected for real-world importance across several use-cases. Tentatively, this could include: a. VQA and multimodal computer use: Qwen3, Molmo2, and InternVL VLMs b. Physical world modeling & video understanding: V-JEPA2, Cosmos3, and ABot-PhysWorld c. Structured generative and design tasks: Qwen3-VL, InternVL, Kimi-VL for artifact generation, with Qwen-Image and SD3.5 d. Robotics action planning (E2E policy measurement): OpenVLA, openpi, and gr00t

  3. We will then identify reasonable covert malicious side-tasks for each benign task and design test sets for evaluating their success.

  4. Finally, we will benchmark models in the benign and malicious settings and analyze performance trends, general behaviors, failure modes, and evidence of reward-hacking if present, also running ablation studies to provide more fine-grained insight as to what affects agent performance.

  5. Stretch Goal: If time allows, we will also provide a small benchmark of monitor performance in this setting and identify any consistent failure cases or differences this setting generates.

Theory of change

Recently, many researchers have argued that LLM agents are becoming increasingly capable of performing long-horizon autonomous AI research. A continued increase in these capabilities would further increase the rate of AI R&D and potentially enable the implementation of RSI loops. Such agent-assisted acceleration in capabilities research could pose significant existential risk if alignment research is not accelerated at the same rate or adversaries successfully compromise an RSI loop. Benchmarks like PostTrainBench act a proxy measure for agentic autonomous research capabilities, providing insight as to how close we are to a capabilities improvement takeoff, along with providing a realistic setting for studying agent security.

However, existing benchmarks measuring autonomous AI research capabilities are limited to the improvement of text-only models. A speed up in the improvement of multimodal systems, which are being increasingly deployed in autonomous research labs and embodied systems, would potentially be just as dangerous, enabling the creation of advanced biotechnology facilities, highly capable robots, or genuine world models. There is currently no robust way to gauge how close agents are to facilitating this research, nor is there a setting for studying the security of autonomous systems performing multimodal training tasks, leaving a major gap in our understanding and preparedness.

Accordingly, we aim to develop a multimodal counterpart to PostTrainBench, releasing a set of benign multimodal post-training tasks paired with malicious side-tasks and benchmarking agent capabilities on both in order to provide a clear picture of risks from both capabilities speed-ups and adversarial attacks. It is our hope that frontier labs will benchmark their new models in this setting and use their findings to appropriately gauge deployment hazards, and that security researchers will be able to use trajectories from this setting to develop strong domain-specific guards.

Your role

Mentees will be lead authors on this study and thus have significant freedom and autonomy within the general constraints of the project direction. Mentors will lay out the general direction and provide guidance as needed, but mentees will have the opportunity to explore and make the lower-level decisions that truly shape the project. We want you to grow as researchers!

Prerequisites

  • Significant research experience experimenting with multimodal systems
  • Research experience post-training multimodal systems
  • Significant experience implementing experiments involving LLM coding agent harnesses
  • Preferred: AI monitoring or control research experience, or a prior project that involved examining agent transcripts (i.e. for evidence of reward hacking, etc.)
  • Python proficiency and experience quickly iterating on experiments with agents

Application question(s)

  1. Without using LLMs, explain to the best of your ability the differences between the various model architectures we describe in the project proposal. We will prefer responses that are only partially correct but clearly human-written. (300 words)
  2. Provide a link to one or more relevant writing samples, ideally from a research context.
  3. (Optional) Describe how you would go about setting up a post-training experiment for one of the listed architectures. Where do you think an agent would struggle? (200 words)

About the mentors

Daniel Ben-Levi

Daniel Ben-Levi

UChicago XLab, Columbia

View profile

Daniel is jointly affiliated with the Existential Risk Lab at UChicago and the Creative Machines Lab at Columbia University, where he conducts research on black-box agent monitors and multimodal interpretability, respectively. He was an early advocate for autonomous red-teaming (Oral @ NAACL) and has since transitioned to working on mitigating risk from misaligned multimodal systems (Spotlight @ ICML). He strongly believes that significant work in both alignment and interpretability is needed to improve the safety of these systems given their poorly understood nature and widespread deployment.

Judah Goldfeder

Judah Goldfeder

Columbia University

View profile

Judah Goldfeder is a PhD Candidate at the Creative Machines Lab at Columbia University, a Student Researcher at Google, and a member of the NSF AI Institute in Dynamic Systems. He is focused on several projects in the Deep Learning space, including developing a Common Task Framework for Scientific Machine Learning, creating an open source building and HVAC simulator, applying Reinforcement Learning to the HVAC systems of large commercial buildings, reconstructing neural network weights from only query access, Auxiliary Learning and Self Supervised Learning for Computer Vision, Machine Crystallography, Machine Learning for Biometrics, and Robotics. Previously, he worked at Facebook AI Research, on applying Transformers to Graph Neural Networks at scale, and at Twitter, where he worked on improving production ads models. He also has helped develop AI educational resources at Learn Ventures, an innovative education startup, and has interned at Bar Ilan University, where he worked on using formal verification to predict gene interaction in cells. He also is a consultant for Dicta, an NLP research nonprofit focusing on Hebrew and related low-resource languages. Judah has organized several workshops on AI for physical systems, including at ICML, NeurIPS, ACM E-energy, and BuildSys. Starting in September, Judah will be a postdoctoral researcher at NYU, where he will be working with Yann LeCun on world models.

Kevin Miao

Kevin Miao

UC Berkeley, Bryel Labs

View profile

Hi, I'm Kevin. I previously led post-training at Apple's Foundation Models team, where my work spanned mechanistic interpretability, diffusion models (3D, 4D, and world modeling), and continual learning. Before Apple, I was a researcher at BAIR, and I've lectured at UC Berkeley on post-training and data science since 2022, most recently teaching a course on building thoughtful AI systems.

I now run a neolab working on multimodal alignment, with a focus on the intuitive and humanistic dimensions of how models learn: not just whether they're capable, but whether they behave in ways people can trust and understand. For SPAR, I'm looking to work with mentees on impactful, empirically grounded projects in this space. You'll get direct mentorship from someone who has shipped post-training systems at scale, and I'll push you to develop research taste: choosing problems that matter, scoping them tightly, and producing work that stands on its own.

Similar projects