Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Distilling the technical AI safety literature and co-authoring the AI Safety Atlas

Communications Alignment

The goal is to create the best possible explanation of technical AI safety. We are aiming to explain how many different safety research areas connect together to form a broader technical AI safety strategy. We already have existing draft writeups for domains like reward misspecification, goal misgeneralization and scalable oversight. We want co-authors to both improve the existing writing, and also write new content on multi agent safety, and cybersecurity practices for AI. Initial versions need to be updated using research published in the 2025-2026 timeframe. The output will be published as a standalone paper and also becomes a chapter of the AI Safety Atlas, which is a textbook already being used by thousands of students to learn about AI safety.

About the project

The AI Safety Atlas is a living distillation project. It has eight chapters, thousands of readers, and is used by ML4Good, ENS Ulm, Sciences Po, and dozens of other universities and programs. But the AI field moves fast, and the technical material hasn't had a serious update in about a year, which is a long time in this field. This project aims to upgrade the technical half of the textbook.

An individual unit of work can be thought of as a literature review of one subfield, written to the standard of a good survey paper. You read research, blog posts, watch technical podcasts, and become the expert in that area, deciding what's load-bearing and what's noise. Then you write it clearly enough that a technically literate non-specialist comes away actually understanding it. Imagine someone who has a Physics, or a CS graduate degree, but is unfamiliar with AI safety. You can work in different domains depending on your specific area of expertise:

  • Significantly improving the existing explanations (reward misspecification, goal misgeneralisation, scalable oversight, possibly evaluations): bringing them to 2025-26, raising the technical level where it's too shallow, and fixing the places that list arguments instead of building one coherent end-to-end picture.

  • Writing new content: interpretability; AI security and cybersecurity; multi-agent systems, complex systems, and systemic safety (coordination, emergent and multi-polar risk) and potentially agent foundations.

It is important to have both AI safety expertise, and be a good writer. Someone who knows the area has to make judgement calls in which concepts connect together, how to explain them, which examples and analogies to use, and what to leave out. You will also be responsible for communicating with the original research authors when something is unclear, or when you have questions and clarifications about how to communicate something.

Coherence is intended as a property of the whole book and is not limited to individual chapters. If you write something on goal misgeneralization, then this should not live in isolation. The way that you explain things will have an influence on the writeup for interpretability, or for the writeup on evaluations. We have to use (or come up with) consistent definitions, consistent conventions and standards that allow the reader to be able to understand the breadth of the field end to end. A change in one area might ripple through the consistency of definitions, the assumed technical level, and the narrative across the entire text.

Theory of change

The number of people who deeply understand the technical problems in AI safety is small, and that's a real bottleneck. Most learning material is either a shallow list of links or a frozen snapshot that goes stale as the field moves. The Atlas fills a specific gap: an integrated, current distillation that other courses can trust and use. Its alumni have gone on to found evaluation orgs, lead recruiting governance organizations, and publish papers at ICML.

Keeping the work current and adding the explanations the field needs most raises both the quality and the throughput of people entering technical safety. Publishing each as a paper forces the writing to a citable standard and makes the work discoverable as scholarship. A mentee finishes having authored a published survey in a core safety area, which is one of the higher-leverage things an early career researcher can do to show their depth of understanding of the field.

Your role

This is a research- and writing-heavy project, for someone who already has, or can quickly build, real technical expertise in their area. You run your own literature review sweeps, make the calls about what to include and how to frame it. I will read and review for technical correctness, for coherence with the rest of the book. I expect you to be self-driven and autonomous.

I am open to a single mentee who owns the whole technical stream, or a team of experts that take one area each.

Prerequisites

Must Haves:

  • Real technical experience in at least one target area (interpretability, AI security, multi-agent and systemic safety, or the alignment problems the existing reviews cover).
  • You should be able to read current papers, understand what they say and judge what matters.
  • Strong technical writing and distillation. You can build a single coherent account, come up with good analogies, compelling narratives instead of just listing arguments.

Nice to have:

  • a published survey, blog post, distillation, or well-received technical explainer (link it);
  • having facilitated or taken a course that used the Atlas.

Location preference

Remote. EU-overlapping hours mildly preferred, not required.

Application question(s)

  1. Link the best thing you've previously written that explains something technical to a real audience: a paper, blog post, distillation, thesis chapter, or documentation. In ≤100 words, say who the audience was and the one idea you most wanted them to walk away understanding.
  2. Pick one result or concept from 2025-26 in your area of expertise and write the explanation for it in 400-600 words, for a technically literate non-specialist. Cite your source(s), say why it matters, and mark one place you'd add a figure, analogy or a worked example.
  3. Explain briefly what you think is important to teach about technical AI safety such that it doesn't go stale in a year: what belongs in the durable framing that will remain important to understand for years versus the current-events layer? (300 words)

About the mentors

Markov Grey

Markov Grey

CeSIA (French Center for AI Safety)

View profile

I am the head of technical AI governance at the French Center for AI Safety (CeSIA). Right now I am designing harmful manipulation evaluation standards for the European Commission's AI Office. I have also been the co-founder and CTO of Equilibria Network, focusing on collective intelligence and complex systems safety. Before this, I was researching measurement standards for AI safety, risk modeling and was leading writing for the AI Safety Atlas textbook. I have previously been a scriptwriter for Rational Animations, and trained in cybersecurity, mathematics, and computer science.

Charbel-Raphael Segerie

Charbel-Raphael Segerie

CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)

View profile

Charbel-Raphael Segerie has extensive experience in AI safety field-building, education, and content creation. He was previously Head of AI at EffiSciences, funded ML4Good, was CTO of a startup, and worked in different French research institutions (Inria, Neurospin). He is an OECD AI expert. He teaches AI safety in ENS, which is one of the only university accredited courses in the EU on AGI safety. His research focuses on identifying emerging risks in artificial intelligence, improving current safety methods such as RLHF and interpretability, and advancing safe-by-design AI approaches. Additionally, he contributed to AI evaluation efforts and collaborated on the EU AI Office's Code of Practice for general-purpose AI systems.

Similar projects