Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Datasets for Human+AI Safety-Judges

Scalable oversight Evaluations

We want to improve Judge systems (composed of both human annotators and LLMs) for Scalable Oversight (e.g. reduce reward-hacking). But first, we need to collate and create robust datasets to evaluate these judges - the focus of the SPAR project. This requires clever ‘sandwitching’ methods to get high-quality ground-truth.

About the project

Our broader goal+motivation: Scalable Oversight is the field of being able to still train models to be aligned, even when they are performing tasks that are extremely challenging to verify correctness and alignment of. Judges, systems that do this verification, are crucial to the alignment of models - they are used directly in RL (e.g. RLHF, RLAIF), but also in evaluations, SFT/pre-training data filtering, and more. These judges have been humans (e.g. RLHF) or AIs (e.g. RLAIF), but there has been very limited research on the best way to combine them to identify and utilize their complementary strengths, into a single judge system that is accurate and robust.

Specific goal for this SPAR project: Our approach to building these better Human+AI Judge systems is to first make an accurate leaderboard, then hill-climb it via trying many things and seeing what improves the accuracy+robustness of the judges. To do this, we can’t just have 1 type of judge task - that’s what we did in our prior Google Deepmind work (https://arxiv.org/abs/2510.26518), which led to Factuality-specific methods that didn’t transfer to other areas of safety. We need to have a broad leaderboard that covers all types of judge tasks, especially in alignment. Eventually, this will be a collection of 100+ datasets, each evaluating the judge on a specific task - across the detection of e.g. computer-control mistakes, deception, cultural bias, jailbreak-robustness, and more. And eventually, we should be THE standard benchmark used for research on improving judges. Not only do we need static benchmarks to evaluate judge-accuracy, but we also want to understand how easily the judge is reward-hacked. We will be building datasets to assess how robust judges are against adversarial optimization pressure, as seen during the RL loop.

Tasks for mentees: So far, we have ~15 datasets we’ve collated already, categorized by our taxonomy. This categorization has identified crucial domains and interaction types not covered by our current collated datasets. Mentees will either be 1. Evaluating existing datasets for their quality, and onboarding them into our collection, or 2. Creating new datasets to fill these gaps, if no other such datasets exist.

Theory of change

Most is described above - Judges are crucial for both evaluating alignment, and training alignment into models. Improvements in the abilities for judges to verify safety directly leads to more aligned models in those domains. And improvements in robustness against reward-hacking reduces the changes of catastrophic and deceptive misalignment.

Creating a canonical “judge benchmark” will dramatically improve the academic community’s efforts on improving judges. Currently, most of that work is done on isolated datasets, with limited generalizability.

Improving judges could also be used to advance capabilities. We plan to mitigate these risks by ensuring the vast majority of our datasets will be verifying the safety of models. For those improving judges for capability tasks, they might still use our ‘canonical’ benchmark focused on safety. Also, building and collating these datasets will boost the “Safety Datasets” community itself, and help everyone measure the alignment of LLMs.

Your role

You will need to take a research project end-to-end - find gaps in existing datasets, do extensive lit review to see if a suitable one exists, and create datasets where none exists. You will be working in a highly-collaborative team, where others are constantly blocked on you, and vice versa. You will need to be clear and correct with your estimated timeline of task completion. You will need to respond quickly on slack (<24 hours), or mention in your slack status if you are slow to respond. We will discuss and decide weekly priorities in our meetings, and async on slack where required.

Prerequisites

Highly proficient using Python Strong coding skills Interest in Amplified/Scalable Oversight Able and willing to quickly and autonomously figure out how to use new tools Willingness and ability to collaborate with others where needed Understanding of the bigger picture, and helping with prioritization of tasks

Application question(s)

Please fill out this application form https://forms.gle/Bb3PGyJ1tZ258nEa8

About the mentors

Rishub Jain

Rishub Jain

AI Safety Nonprofit (name TBD)

View profile

Rishub Jain is the founder of a new non-profit launching in August 2026 focused on improving safety-judges via human–AI complementarity. He recently spent 7 years at Google DeepMind: first as a research engineer on AlphaFold 2 and 3, then moving into scalable oversight where he co-led the Rater Assist team, and lastly working on Information Quality and Gemini Security. His work spanned LLM post-training and evaluation, human–AI collaboration, improving the information quality of the web, prompt injections, and science-engineering, all driven by an interest in building tools that reduce the future harms and increase the benefits of powerful AI.

Mentorship has been central to Rishub's work for years, inside and outside DeepMind. Since Sept 2025, he has continuously led an evolving group of 13-25 SPAR+MARS mentees on scalable oversight, and especially enjoys helping researchers from fields adjacent-to-AI find meaningful footing in AI safety.

Joshua Jacob

Joshua Jacob

AI Safety Nonprofit (name TBD)

View profile

Josh is co-founding a new safety non-profit focused on building stronger "judges" for scalable oversight with mentor Rishub Jain.

Previously, he was the co-lead for Human Data Engineering at Google DeepMind, working closely with research teams to improve Gemini capabilities across domains like safety, computer use, and coding with high quality post-training and evaluation data.

Similar projects