Kairos, the home of SPAR, has raised $50M from Coefficient Giving. Read the announcement

All Fall 2026 projects

Running Frontier Robotics Benchmark Program

Generalist Evaluations Other

Robocurve and Generator Residency are running an open call for reproducible robotics benchmarks, with up to $500,000 in awards. We're looking for technical program managers with prior experience in robotics or building benchmarks.

As part of this project, you will evaluate applications, select the finalist teams to receive funding, run office hours, and help run the program. Roughly 10 hrs/week, mid-September to mid-December.

About the project

The main goal of this program is to build a repository of reproducible evaluations that various stakeholders (governments, academia, and the public) can use to benchmark general robotics capabilities, providing a trusted means for the public to gauge progress in robotics. With this infrastructure, we can better understand and prepare for physical automation of labor and a possible intelligence explosion.

Theory of change

Currently, there are very few well-run, standardized benchmark methods for measuring general robotic capabilities. Leading robotics AI models are released with each company's private task lists, without third-party testing and evaluation. This leaves AI forecasters, policymakers, and the public without a reliable understanding of how capable these systems really are or how fast they're improving.

Without shared, reproducible benchmarks, people outside the labs have very little basis for judging how soon robots will take on real-world jobs, or how quickly the field is moving. With these benchmarks, we hope to inform the public and policymakers about the rate of progress in general robotics and help to reduce the risks it poses.

Your role

There are two main stages of this program.

1. Stage 1: Selection, mid-September to late September. Review Stage 1 proposals. Most of the work is judgement rather than writing code. You're going to support the creation of the evaluation rubric for applications received in the public call for proposals and actually evaluate them to find the top 12 finalist teams.

2. Stage 2: Support, October to December. Help selected teams reach a working benchmark by onboarding them to Inspect Robots, running office hours, checking in on teams to help them make progress, and doing whatever else is necessary to produce the highest-quality output.

What you get

  • Agency to decide which teams receive the awards.
  • Work with 12 robotics teams for 3 months and help them ship the first independent frontier robotics benchmarks.
  • You learn Inspect Robots before the first public benchmarks are built on it.

Prerequisites

Required

  • Comfortable with Python and reading unfamiliar codebases
  • Experience building evaluations or benchmarks, or hands-on robotics work
  • Agentic and proactive in solving problems

Helpful

  • Familiarity with LeRobot, Inspect, or Inspect Robots
  • Experience with creating evaluations or benchmark design
  • TA experience

Application question(s)

Question 1: Why are you interested in supporting robotics evaluation specifically, rather than AI evaluation or robotics in general? (~15min)

Question 2: A proposal contains this task: "The robot places the mug on the shelf. Score 1 if the mug ends up on the shelf, 0 otherwise." What would you need specified before an independent operator could run this? Which of these are most important to get right? (~25 min)

Question 3: The project has two heavy weeks: mid-September and mid-December. Roughly how many hours could you give the project these two weeks? An honest answer doesn't count against you; we're asking just so we can plan ahead.

Note: We're really looking for people who can think clearly and make good judgment. These questions above shouldn't take very long, and we request that you not use AI tools for writing or ideation.

About the mentor

Yashvardhan Sharma

Yashvardhan Sharma

Constellation, Georgetown

Yashvardhan Sharma is a graduate student in Security Studies at Georgetown focused on AI and national security, currently working at Constellation. He studied Applied Computer Science and AI at Minerva University and previously worked on compute policy as a Research Scholar at MATS. His focus is on technical AI governance, bridging technical AI research and policy.

Similar projects