Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Proving Safety Properties of Guardrail Models

AI control Scalable oversight

We will explore how far formal verification techniques can be pushed to prove key safety properties of constitutional classifiers (guardrail models). The goal is to develop principled, machine-checkable guarantees for components responsible for moderating or constraining the behaviour of advanced AI systems.

About the project

Guardrail or constitutional classifiers are increasingly used to enforce safety constraints, detect harmful model outputs, and mediate system behaviour. Yet these guardrails are typically evaluated empirically, and we lack precise, formal guarantees about their behaviour.

This project investigates how far current formal-verification methods can be pushed to verify properties of real-world guardrail models. Depending on the mentees' skills and interests, research directions may include:

  • identifying and formalising key safety properties;
  • modelling constitutional classifiers in a formal system (e.g., SMT encodings, abstract interpretation frameworks, constraint-based reasoning);
  • exploring decidability boundaries and approximation techniques for verifying neural classifiers;
  • developing prototype verification pipelines for small but representative guardrail components;
  • studying failure modes arising from mis-specification or adversarial examples.

The goals are:

  • to understand which safety properties can be mechanically verified with current methods,
  • to identify gaps where formal methods break down, and
  • to build early demonstrations and tools that make guardrail properties explicit, inspectable, and ideally provable.

This project is exploratory but concrete: mentees will work with the constitutional classifier structure described by Anthropic and attempt verification of simplified, GPT2-based models or extracted abstractions.

References:

Theory of change

Guardrails are safety-critical infrastructure: they mediate the behaviour of increasingly capable AI systems and are central to preventing harmful or escalatory outputs. Yet today, we depend largely on empirical testing rather than verified guarantees.

Formal verification of guardrail properties would:

  • increase confidence that safety constraints are enforced even under distribution shift or adversarial prompts;
  • reveal specification errors or gaps early;
  • provide a principled way to audit and certify guardrail behaviour;
  • contribute to the long-term goal of building AI systems whose safety-critical components are inspectable and provably reliable.

Even partial successes or probabilistic guarantees help clarify what kinds of guarantees are possible with current techniques and where research should focus next.

Your role

Mentees will help formalise safety properties, develop verification encodings, implement prototype analyses, and explore limitations of existing techniques. If interested, they can also be included in the process of academically publishing their results. They will work with meaningful autonomy while receiving structured weekly guidance, literature direction, and technical supervision.

Prerequisites

  • Passionate about AI safety.
  • Some AI/ML knowledge.
  • Proficiency in Python or another programming language used for verification tooling.
  • Some exposure to formal methods (e.g., SAT/SMT solving, abstract interpretation, type systems, model checking, or proof assistants).
  • Ability to read and reason about technical research papers.
  • (Optional but valuable): experience with tools like Z3, PyTorch, TLA+, Coq/Lean, Isabelle, or program-verification frameworks.

Time commitment

8–16

Application question(s)

  • What is the main bottleneck in verifying safety properties on constitutional classifiers?
  • Please provide an example safety property you would try to verify about a constitutional classifier.

About the mentors

Pascal Berrang

Pascal Berrang

University of Birmingham

View profile

Pascal Berrang is an Associate Professor in Computer Security at the University of Birmingham, specialising in the security and privacy of AI and blockchain systems. He pioneered the concept of membership-inference attacks in machine-learning models and co-invented ML-Leaks. More recently, Pascal has led a funded research programme on zero-knowledge proofs for AI-safety applications, and is a co-founder of Zeroth Research, where he builds formal methods and cryptographic verification tools for safe and transparent AI systems.

He is keen to mentor early-career researchers at the intersection of technical AI safety and security, supporting them in developing rigorous research agendas, navigating publication pathways, and building collaborations across academia and the alignment community.

Luca Arnaboldi

Luca Arnaboldi

University of Birmingham

View profile

Dr Luca Arnaboldi is an Assistant Professor at the University of Birmingham and a researcher specialising in the security and formal verification of autonomous and AI-enabled systems. His work sits at the intersection of computer security, formal methods, and machine learning, with a strong emphasis on real-world impact through collaboration with industry, regulators, and policymakers. Luca has led and contributed to internationally recognised tools and research outputs, secured competitive funding from bodies including EPSRC, Innovate UK, and ARIA, and is currently building a verification stack for safeguarded AI systems. Alongside academia, he has industry experience ranging from large technology companies to founding an AI-safeguarding spin-out.

Luca is also an award-winning educator and mentor, with extensive experience supervising PhD, MSc, and undergraduate researchers across security, verification, and AI. He is particularly passionate about supporting early-career researchers in navigating academic careers, building impactful research agendas, working effectively with industry, and translating research into practical outcomes. Through mentoring, he aims to offer pragmatic advice grounded in experience, while helping mentees develop confidence, clarity, and a sense of direction in their research journey.

Similar projects