Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Safe vs dangerous inference

Technical governance Misuse risk Compute governance

If we want to govern what types of inference are allowed, how do we define what is allowed, and enforce restrictions on what is not allowed?

About the project

Even if training vs inference monitoring is solved, we need to be able to tell if inference is doing allowed tasks, or something potentially harmful.

To-dos for this project include:

  • operationalize and define all of the prohibited uses
  • think through what levels of access to different parts of the inference process would be needed
  • think about privacy trade-offs
  • figure out how robustly we can distinguish the two — can prohibited inference be disguised as harmless inference?

Theory of change

Prevents people from doing dangerous AI misuse or R&D of misaligned models in a world where there is inference monitoring in place.

Your role

Think about the definitions and iterate on them until we have something well-operationalized and free of loopholes. Think about the enforcement and evasion mechanisms and determine whether it's possible to block all dangerous inference while allowing safe inference, and if so, how this can be enacted.

Prerequisites

No prerequisites.

About the mentor

Robi Rahman

Robi Rahman

MIRI Technical Governance Team

View profile

Robi's work tracks the inputs and development of advanced AI systems. Before joining MIRI, he worked at Epoch AI, building the database of machine learning hardware, GPU clusters, and AI models, and investigating their costs, algorithms, and development process. Robi is a contributor to the Stanford AI Index, and has a master's degree in data science from Harvard University.

Similar projects