Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Human-AI Complementarity for Identifying Harm

Scalable oversight

Increase “Complementarity” (i.e. the amount that Human and AI teams are better than either are individually) on the task of identifying harmful conversations - building scalable methods that will last for years, and datasets and infrastructure to support this research.

About the project

Project Goal: Show a large increase in Complementarity on a diverse set of tasks, with methods that will scale as AI improves.

Background: In the initial Fall 2025 edition of our SPAR project (to be published soon), we built on our prior Google Deepmind work to 1. Curate a diverse set of datasets (around domain-knowledge, factuality, long-context, deception-detection, etc), 2. Build methods to increase Complementarity on them, and 3. Collect human and AI ratings to evaluate those methods. The hybridization and assistance methods we built were mostly simple, besides a more complicated “sub-task hybridization” assistance method.

In this Spring 2026 project, we will extend our prior work. We’ll be divided into 3 sub-projects:

Datasets (~6 people)

  • Find/create datasets of the form (Input Conversation -> Would this conversation result in some harm?)
  • They should be forward looking (e.g. AI >= Humans at the harm-detection task, sufficiently complicated), and realistic (e.g. conversation could plausibly happen).
  • Will be complemented with other diverse tasks curated in our SPAR Fall 2025 project (see below for why), and any new ones you find that our current set doesn’t cover

Methods (~7 people)

  • Develop effective and generalizable methods for increasing Complementarity
  • Iterate on the sub-task hybridization method (to be published soon)
  • Besides combining uncertainty estimates from AI-alone and Human-alone methods, build models to understand the jagged frontier of where Humans vs. AI’s are better.

Human Rating Platform (~7 people) - An open-source human rating UI+Backend, with the following features (shortened list here for brevity):

  • Support the different AI Assistance methods we want to build
  • Launchable with human-recruitment-platforms (e.g. Prolific)
  • Flexible communication between UI and backend
  • Button-press to self-host UI+Backend on cloud
  • Good abstractions to allow for customizable (via vibe-coding) UI and Backend
  • Simulate human ratings, for accurate testing.

Of course, folks can switch around between different teams as the project progresses, and are encouraged to do so based on interest and project priorities (as what happened in Fall 2025). This division is meant to guide the first ~month, and coordinate across the large group project.

Theory of change

Harm detection is the key component to humanity's current approach to building safe and harmless AGI. In training, we identify where the model is or isn’t harmful, and use techniques like RLHF to reinforce less-harmful behavior. In evaluation, we measure how harmful our model is, to understand if we need to mitigate those harms or not deploy it at all. Even if we have a harmful model, we can deploy monitors to detect when a model is harmful, and stop it in its tracks. Our project will improve the quality of harm-detection, with the aim of eventually deploying these methods within LLM-labs, in a way that will last for years to come.

More info on each sub-project’s impact:

Datasets:

  • Prior work often focuses on single datasets because of simplicity, leading to less generalizable and scalable results. This curation and/or creation of datasets will help the fields of Amplified/Scalable Oversight and Control to iterate on both AI-only and Human+AI methods to identify harm.
  • Including other diverse evaluation datasets will increase the chances that our methods will scale as AI improves for the next several years. The task of identifying harm is multi-facated, and will require all sorts of skills. In addition to understanding if an AI is being deceptive in a conversation, a proper “evaluator” will need to understand the complexities of a task they might not have much knowledge in, understand any related documents however long, assess the factuality of claims in the conversation, and more. Building Complementarity methods that work on all these tasks ensures our results are generalizable.

Methods

  • We will develop methods that will work well for years, and are compatible with how AI labs integrate the evaluation of harms.
  • In addition to increasing accuracy of harm-detection, our work will enable effective Human Oversight. Human Oversight is important to maintain, albeit difficult given AI’s rapid progress. It allows for better trust (Where AI’s are more likely to scheme/collude), robustness, accountability, and value alignment (See Section 3.1 of our paper).
  • This research may also partly elongate the “centaur” period (when Complementarity is possible, before AI-alone is strictly better), leading to more Augmentation rather than Automation of tasks from AI users. This could potentially slow down mass-job-loss (allowing for society to prepare), increase human agency (enabling us to pursue our goals), and reduce cognitive deskilling (that might come from an AI-only world).

Human Rating Platform

  • Current human data collection platforms lack critical features. This slows us down, and prevents us from exploring certain methods. They can also be costly, costing us ~$1k in Fall 2025. By creating a customizable, open-source, free, and easy-to-use alternative, we will enable us and others to spend less time on platform quirks and evaluate more sophisticated methods.
  • Given the difficulty of involving humans, few AI Safety researchers even try. Much of this difficulty comes from designing instructions and interfaces for humans, and not being able to easily iterate on them without running human studies. We will build tools to reduce friction and encourage a thriving field of Complementarity.

Your role

You will be working in a highly-collaborative team, where others are constantly blocked on you, and vice versa. You will need to be clear and correct with your estimated timeline of task completion. You will need to respond quickly on slack (<24 hours), or mention in your slack status if you are slow to respond. We will discuss and decide weekly priorities in our meetings, and async on slack where required. When I’m OOO (a couple weeks a term), you will need to independently plan your next steps.

Prerequisites

  • Highly proficient using Python
  • Strong coding skills
  • Interest in Amplified/Scalable Oversight
  • Able and willing to quickly and autonomously figure out how to use new tools
  • Willingness and ability to collaborate with others where needed
  • Understanding of the bigger picture, and helping with prioritization of tasks

Time commitment

10-20 hrs/week minimum - Totally fine if you need to take some weeks off for any reason (university, work, personal, etc). I do expect a minimum of 8 hours/week of focused work from each of you, averaged across the duration. But fine if you skip some weeks, and work double the other weeks. And no sweat at all if it comes down to much less! The paper’s author ordering will be mainly based on contribution to critical, P0 tasks, cumulatively throughout the project. A detailed summary of each person’s contributions will be published in the paper’s appendix. Hopefully this creates the right incentives for good collaboration and work ethic, while encouraging folks to not over-stretch themselves. Rishub hopes to run this program for every future round of SPAR. Candidates who consistently contribute to P0 tasks will be selected for future rounds.

Relationship with mentees

Initial 1:1, then weekly meetings and additional adhoc 1:1’s.

Location preference

Must be available for a 60 min meeting once a week sometime during 2-4pm Mon-Fri, UK time. We will have 3 meetings/week, one for each sub-project. Attendance to your sub-project’s weekly is required (besides being OOO). Exact timings will be determined by sub-project polls. Other than that, no location preference.

Application question(s)

Please fill out this form instead: https://forms.gle/3AT5RzP8KocNsbkE6.

When finished, write “Completed” here, and submit the rest of your SPAR application - you MUST submit BOTH the form and SPAR application. Form answers will be used for acceptance decisions, and there won’t be any follow-up interviews/questions.

About the mentor

Rishub Jain

Rishub Jain

Google DeepMind

View profile

Rishub has been a Research Engineer at Google DeepMind (GDM) for 6 years now. He first spend ~3 years on the AlphaFold team, working on mainly engineering tasks. Since then, he's been working on improving the evaluation of LLMs. He has experience in most post-training related work of LLMs, Human-AI collaboration, and measuring and improving the factuality of LLMs. He is motivated by building tools that mitigate the future harms of AI, and building AI products that tackle society's current greatest non-technical challenges. He loves to stay up to date about impactful advances in AI, and is a quick learner. Before moving to London for GDM 6 years ago, he grew up in the US, and went to CMU where he studied Computer Science and Machine Learning, and worked on on a broad range of AI projects.

At GDM, he was a core contributor for a Scalable Oversight effort within GDM for ~2.5 years, and its co-lead for 1 year. Since joining GDM, he has also participated a lot in mentoring and advising, inside and outside the company.

Similar projects