Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Mapping and Verifying Multi-Dimensional Compliance in AI Constitutions through Constraint Satisfaction

Scalable oversight Technical governance Evaluations

Frontier labs focused on AI safety and alignment are adopting AI constitutions as a means to ensure alignment with codified principles, values, and behavioral specifications. Anthropic’s Claude Constitution (Askell et al., 2026) and OpenAI’s Model Spec (OpenAI, 2026) serve as technical governance resources in post-training alignment (Bai et al., 2022; Guan et al., 2024). AI constitutions are generally long, complex documents that combine behavioral requirements, guiding principles, authority structures, priorities, and exceptions. Jakkli et al. (2026) found that, despite improvements across model generations, violations persist when models must resolve competing requirements and sources of authority, particularly in multi-turn and agentic task settings. This project aims to build upon that line of research to more fundamentally examine existing AI constitution documents, better understand the multi-dimensional requirements they present, and formally classify common failure modes. Specifically, this project aims to (1) create a classification framework to analyze failures in multi-requirement compliance settings, (2) create a benchmark of use cases derived from the framework to evaluate AI compliance, and (3) explore constraint-satisfaction methods for verifying compliance in such settings.

About the project

The project broadly has two goals. The first is to better understand the nature of multi-dimensional requirements presented in AI constitutions, and the second is to develop a formal representation and technical solution that considers compliance as a constraint-satisfaction problem.

AI constitutions are complex governance artifacts that consist of heterogeneous requirement objects. For instance, they contain objectives that are broad but directional alongside hard constraints that are highly specific, authority structures that define whose instructions to follow, and exemptions that carve out cases in which otherwise applicable requirements can be ignored (Askell et al., 2026; OpenAI, 2026). Verifying compliance therefore requires navigating ambiguous cases in which multiple satisfaction conditions can apply, creating opportunities for misalignment and unpredictable behavior (Jakkli et al., 2026; Jiang et al., 2024; Wen et al., 2024). The project therefore aims to formally classify requirement objects and specify the norms governing their interactions. Ideally, the framework will draw on established legal theory, formal specification, and deontic logic to identify a pragmatic and auditable classification scheme, which will then be applied to existing AI constitutions to identify recurring structural ambiguities and patterns of conflict.

Secondarily, the project aims to explore whether requirement objects can be translated into formal constraints and if external solvers can verify compliance, identify violations, and surface conflicts among applicable requirements. Constitutional AI relies on supervised and reinforcement-learning phases where AI-generated preference comparisons are employed to produce reward signals (Bai et al., 2022). The guiding preference model is learned from sampled comparisons, and reward models can perform less reliably on unseen prompts and responses when the evaluation distribution shifts (Yang et al., 2024). Multi-constraint benchmarks further demonstrate significant deficiencies when models must satisfy accumulated or composed requirements (Jiang et al., 2024; Wen et al., 2024), while LLM-based evaluators can produce judgments that vary with irrelevant features of the evaluated response (Chen et al., 2024).

In contrast, the project seeks to leverage constraint solvers, which can provide deterministic results for formally specified requirements and support independent auditing. Such an approach would enable an external verifier whose judgments are not dependent on the model being evaluated (i.e., LLMs as judges to evaluate LLMs). Formalizing these requirements as constraints could make distinctions between hard constraints and soft objectives explicit, encode conditions, priorities, and exceptions, and enable solver-based checks for inconsistency or joint unsatisfiability, alongside analysis of whether relevant cases remain under-specified (Bistarelli et al., 1997; Antoniou, 2004; de Moura and Bjørner, 2008). We envision applying this representation to existing AI constitutions to identify combinations of requirements that are jointly unsatisfiable or under-specified. The resulting verification signals could in turn supply per-constraint rewards for constraint-aware reinforcement learning (Qi et al., 2026), support stage-level audits of plans, tool calls, and state transitions (Chen et al., 2025), and enforce formally checkable requirements before consequential actions are executed (Wang et al., 2026). The project will therefore build on related work in symbolic reasoning, constraint-based instruction verification, and runtime policy enforcement (Pan et al., 2023; Su et al., 2026; Wang et al., 2026) to determine which constitutional requirements can be formally represented and which remain dependent on contextual judgment.

In summary, the project aims to do the following:

  • Audit existing AI constitutions. Examine Anthropic’s Claude Constitution and OpenAI’s Model Spec for cases in which hard constraints, soft objectives, conditions, priorities, and exceptions produce conflicting, jointly unsatisfiable, or under-specified requirements. Ideal outcome: a classification framework that delineates requirement objects and specifies ideal interaction expectations.
  • Evaluate multi-requirement compliance. Turn the identified cases into test scenarios that require models to determine which requirements apply, resolve priorities, and satisfy the resulting set of conditions. Ideal outcome: an initial dataset and benchmark to evaluate model compliance in constructed multi-dimensional requirement scenarios with a focus on AI safety and alignment.
  • Test constraint-based verification. Translate a bounded subset of requirements and their interactions into formal constraints and compare solver outputs with human judgments. Ideal outcome: an engineered proof of concept integrating a constraint solver into a verification pipeline to evaluate AI agents.

Theory of change

Commercial frontier labs are increasingly both the developers and distributors of the most capable AI models. In contrast to the previous environments where AI models were developed in academic settings, the leading capable models are no longer open-source, open-weight, nor fully auditable, as training regimes and datasets are occluded as trade secrets and IP. Proprietary models, post-training reinforcement learning, and internal evaluations limit the ability of external actors to assess whether models are truly safe and aligned with human flourishing. Moreover, compliance is often assessed by the same organizations and increasingly by the same class of models whose behavior is being evaluated. Further, sufficiently capable systems may also learn to evade or manipulate these evaluators. External verification is therefore necessary to provide credible safety assurances and serve as early-warning signals for misalignment. Constraint solvers offer an external and deterministic mechanism for identifying contradictions in alignment targets, supporting reproducible external audits, and enforcing formally specified requirements before consequential actions are taken. This could support better-defined constitutions and independently verifiable deployment safeguards while reducing reliance on corporate self-regulation and AI-mediated compliance.

Your role

The mentee should be self-motivated, highly agentic, communicate clearly, and be able to plan and run experiments independently. The mentors will provide technical and strategic project guidance and if there is a paper outcome, co-authorship support for main conference level papers.

Prerequisites

The project is open to applicants from any academic or professional background. Candidates should be self-motivated, communicate clearly, and be able to plan and run experiments independently. An ideal candidate will have experience with Python and independent research, familiarity with standard data analysis methods, and at least one research output that has been shared with a wider audience, such as a published paper, workshop submission, shared-task contribution, LessWrong post, or course paper. However, these are not strict requirements, and we will consider highly motivated and agentic candidates who can demonstrate the ability to conduct high-quality research.

Location preference

I'm based in Berkeley and my co-mentor is in the UK. We are will meet independently and virtuallly with the mentee(s) based on their geographic proximity to us.

Application question(s)

The project is open to applicants from any academic or professional background. The mentee should be self-motivated, communicate clearly, and be able to plan and run experiments independently. An ideal mentee will have experience with Python and independent research, familiarity with standard data analysis methods, and at least one research output that has been shared with a wider audience, such as a published paper, workshop submission, shared-task contribution, LessWrong post, or course paper. However, these are not strict requirements, and we will consider highly motivated and agentic candidates who can demonstrate the ability to conduct high-quality research. We expect to support one-two mentees.

About the mentors

Dhairya Dalal

Dhairya Dalal

MATS Research

View profile

Dr. Dhairya Dalal is a Research Manager at MATS, where he supports scholars and mentors advancing research on AI safety and alignment. Prior to MATS, he held senior research, applied AI engineering, and technical project management roles across early-stage startups, industry, academia, and nonprofits. Most recently, he worked on agentic causal inference and discovery in complex systems at Causely, and served as a technical project manager at the Allen Institute for Artificial Intelligence and the Allen Institute for Brain Science. Dhairya earned a doctorate in Computer Science from the University of Galway, where his research focused on improving fundamental causal reasoning in LLMs and causal knowledge representation. He additionally holds a master's in Computer Science from Harvard University, where he worked on reinforcement learning for dialog systems, and a bachelor's from the University of Rochester, where he studied English and Interdisciplinary Studies in social and political thought.

Marco Valentino

Marco Valentino

University of Sheffield

View profile

Dr. Marco Valentino is a Lecturer (Assistant Professor) in Artificial Intelligence and Applications of AI in the Natural Language Processing (NLP) group at the University of Sheffield. Prior to Sheffield, he was a member of the Neuro-Symbolic AI Group at the Idiap Research Institute in Switzerland, and obtained a PhD in Computer Science from the University of Manchester. His research focuses on developing AI systems that can use explanation as a core mechanism for learning and reasoning, investigating the integration of neural and symbolic AI methods. Moreover, he is interested in developing methodologies to interpret, control, and evaluate Large Language Models (LLMs), with a focus on disentangling knowledge acquisition from abstract logical reasoning, and enabling out-of-distribution, out-of-domain generalisation. His research on neuro-symbolic NLP received a best resource paper award at EMNLP 2025 and an outstanding paper award at EMNLP 2024.

Similar projects