Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Harmonizing Frontier Lab Safety Thresholds

Lab governance

This project will analyze the frontier safety frameworks of major AI labs to identify divergences in risk thresholds, propose a harmonized framework for unacceptable risk definitions, and produce a document showing where major labs converge and where they diverge on key risk thresholds. This initiative follows up on the Global Call for AI Red Lines.

About the project

Background

Frontier AI labs have each developed their own safety frameworks—Anthropic's RSP, OpenAI's Preparedness Framework, DeepMind's Frontier Safety Framework. These contain thresholds for dangerous capabilities that trigger enhanced security, safety, or governance measures.

The problem: these thresholds differ dramatically or sometimes are not even public. Without harmonization, this creates a race to the bottom where safety-conscious labs are disadvantaged.

Research Questions

  1. Where do existing frameworks converge on risk thresholds, and where do they fundamentally diverge?
  2. What is the "minimum common denominator" that all major labs already implicitly support?
  3. How can divergent metrics be translated or reconciled?
  4. What harmonized threshold language could labs realistically adopt?

Deliverables

A paper (3,000 words) that:

  • Systematically compares RSP thresholds across capability domains (AI R&D, CBRN, cyber, autonomy)
  • Identifies convergence points and fundamental divergences
  • Proposes harmonized threshold language for 2-3 key capability domains
  • Develops a "Rosetta Stone" for translating between different measurement approaches

Scope

Primary focus on Anthropic, OpenAI, and DeepMind frameworks. Secondary analysis of Meta's Frontier AI Framework and xAI's Risk Management Framework (which notably omit AI R&D thresholds entirely). Priority capability domains: AI R&D acceleration, CBRN uplift, cyber offense, autonomous replication.

Methodology

  • Weeks 1-2: Systematic extraction of all quantitative and qualitative thresholds from each framework
  • Weeks 2-5: Comparative analysis identifying convergences, divergences, and measurement incompatibilities
  • Weeks 5-9: Develop harmonization proposals; create a framework for metric translation
  • Weeks 9-11: Expert feedback from lab safety teams and governance researchers
  • Weeks 11-12: Final paper drafting and stakeholder review

Context

This project builds on CeSIA's October 2024 workshop with frontier lab staff, which demonstrated agreement in principle on harmonization. The Global Call for AI Red Lines (12 Nobel laureates, 90+ organizations) creates political pressure for this technical work. The EU AI Act Code of Practice’s measure 4.1 creates regulatory demand.

References

  • Anthropic RSP v2.2 (May 2025)
  • OpenAI Preparedness Framework v2.0 (April 2025)
  • DeepMind Frontier Safety Framework v3.0 (September 2025)
  • Koessler et al. (2024), "Risk thresholds for frontier AI" (GovAI)
  • FAS Analysis of Responsible Scaling Policies (2024)
  • Seoul Frontier AI Safety Commitments (May 2024)

Theory of change

Current divergences in existing frameworks mean that identical capabilities might trigger containment at one lab while remaining uncontained at another. This undermines the entire RSP framework's credibility and creates competitive pressure toward weaker thresholds.

This project: (1) documents gaps publicly, creating pressure for convergence, (2) provides ready-to-adopt harmonized language that weakens the "technically impossible" excuse, (3) demonstrates minimum consensus already exists, strengthening the case for regulatory codification, and (4) directly informs EU AI Act implementation of measure 4.1 of the Code of Practice and AISI standard-setting.

The Seoul Commitments require labs to "set out thresholds at which severe risks would be deemed intolerable", but nothing prevents those thresholds from being set arbitrarily high. This project creates the technical foundation for meaningful harmonization.

Your role

Mentees are active researchers who will co-author the paper.

Mentee 1 & 2 (Technical Analysis): Systematically extracts and compares thresholds across frameworks, develops the metric translation framework, drafts comparative analysis sections. Has a technical AI safety background.

Mentee 3 (Policy): Analyzes regulatory embedding pathways, drafts recommendations sections. Ideally has policy/governance background.

Prerequisites

  • Strong analytical skills and attention to detail (will be comparing precise threshold language)
  • Familiarity with at least one frontier lab's safety framework (RSP, Preparedness Framework, or FSF)
  • Research writing ability
  • Comfort with technical AI safety concepts (capability evaluations, ASLs, dangerous capabilities)

Location preference

Preference for availability during European business hours. Fully remote acceptable.

Application question(s)

Question 1 (Required, 250 words max):

Compare how Anthropic and OpenAI define thresholds for AI R&D capabilities in their respective frameworks.

  • OpenAI, in their preparedness framework, proposes the following red line: "The model is capable of recursively self-improving (i.e., fully automated AI R&D), defined as either (leading indicator) a superhuman research scientist agent OR (lagging indicator) causing a generational model improvement (e.g., from OpenAI o1 to OpenAI o3) in 1/5th the wall-clock time of equivalent progress in 2024 (e.g., sped up to just 4 weeks) sustainably for several months.- Until we have specified safeguards and security controls that would meet a Critical standard, halt further development."
  • Anthropic: AI R&D-5, similar, but slightly different: “The ability to cause dramatic acceleration in the rate of effective scaling. Specifically, this would be the case if we observed or projected an increase in the effective training compute of the world's most capable model that, over the course of a year, was equivalent to two years of the average rate of progress during the period of early 2018 to early 2024. We roughly estimate that the 2018-2024 average scaleup was around 35x per year, so this would imply an actual or projected one-year scaleup of 352 = ~1000x.”

Compare and contrast both thresholds? What do you think about them?

Question 2 (Required, 250 words max):

What would a harmonized threshold for this capability domain look like?

About the mentor

Charbel-Raphael Segerie

Charbel-Raphael Segerie

CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)

View profile

Charbel-Raphael Segerie has extensive experience in AI safety field-building, education, and content creation. He was previously Head of AI at EffiSciences, funded ML4Good, was CTO of a startup, and worked in different French research institutions (Inria, Neurospin). He is an OECD AI expert. He teaches AI safety in ENS, which is one of the only university accredited courses in the EU on AGI safety. His research focuses on identifying emerging risks in artificial intelligence, improving current safety methods such as RLHF and interpretability, and advancing safe-by-design AI approaches. Additionally, he contributed to AI evaluation efforts and collaborated on the EU AI Office’s Code of Practice for general-purpose AI systems.

Similar projects