Can mechanistic interpretability provide legally meaningful evidence of artificial intent? This project explores how advances in AI interpretability may reshape concepts of legal responsibility, mens rea, and legal personhood for increasingly autonomous AI systems.
About the project
Artificial intelligence systems are rapidly evolving from passive computational tools into increasingly autonomous agents capable of long-horizon planning, strategic behaviour, and, in some cases, deception or reward hacking. These developments raise fundamental questions for legal systems built upon concepts such as intention, knowledge, recklessness, negligence, and responsibility. Courts are already attempting to reckon with the existing and coming waves of cognitively sophisticated AI systems as they permeate society - so it is a highly relevant, and in many ways, urgent, topic.
This project investigates whether recent advances in mechanistic interpretability—including sparse autoencoders, circuit tracing, attribution graphs, and chain-of-thought monitoring—provide a scientifically meaningful basis for assessing artificial mental states. In particular, the project will explore whether these techniques can support new legal concepts of synthetic mens rea and more generally inform future frameworks for AI liability, accountability, and legal personhood.
Possible research directions include:
- legal theories of personhood and responsibility;
- mens rea and intentionality in criminal and civil law;
- mechanistic interpretability as legal evidence;
- chain-of-thought monitoring and evidentiary reliability;
- AI deception, scheming, and autonomous decision-making;
- AI liability and governance frameworks;
- comparative approaches across common law and civil law jurisdictions;
- policy implications for frontier AI systems.
The project will follow a traditional legal research methodology informed by contemporary AI research. Participants will conduct comprehensive literature reviews across law, computer science, and AI safety, critically analyse recent developments in mechanistic interpretability, evaluate emerging legal doctrines, and contribute to the development of a publication-quality academic article. I have an extensive working draft on topic already so the project will involve expanding this while also prospectively a second paper.
The objective is to produce research suitable for submission to a leading journal in AI law, technology law, or interdisciplinary legal scholarship.
While AI tools may be used to assist with literature review, coding demonstrations, or drafting, participants will be expected to undertake the core legal reasoning, critical analysis, and manuscript writing themselves.
Theory of change
As AI systems become increasingly autonomous, legal systems require principled methods for assigning responsibility, assessing intent, and governing increasingly capable artificial agents. Existing legal concepts such as mens rea presuppose human mental states and are therefore difficult to apply to advanced AI systems. By investigating whether mechanistic interpretability can provide scientifically grounded evidence relevant to concepts such as intention, knowledge, recklessness, and deception, this project seeks to strengthen future legal and regulatory frameworks for AI safety, accountability, and alignment. Developing reliable methods for understanding and evaluating AI behaviour will become increasingly important as frontier AI systems assume more consequential roles across society. This project builds upon my ongoing research into AI governance, mechanistic interpretability, legal personhood, and artificial intent.
Your role
This project is intended to mirror the experience of participating in a university legal research group.
Participants will be expected to work independently between meetings, progressively developing expertise through legal scholarship, interdisciplinary literature review, and critical analysis. While the project engages extensively with recent developments in artificial intelligence and mechanistic interpretability, its primary focus is legal and jurisprudential. I have a working draft paper on this project already but it is in its early stages. Ideally I am looking for someone to collaborate with me by expanding my initial review of the relevant legal, AI safety, and computer science literature before identifying open legal questions concerning artificial agency, interpretability, responsibility, and mens rea. Participants may also undertake limited computational demonstrations where these assist legal analysis, although the emphasis will remain on rigorous legal scholarship. This work may also involve collaboration with colleagues of mine at Cambridge University and potentially, depending upon outcome, involvement in further research on AI systems and the law.
Participants should expect to present their progress regularly, discuss research findings, receive detailed feedback, and contribute towards the preparation of a publication-quality manuscript.
Prerequisites
Applicants should have a strong interest in AI law, technology regulation, jurisprudence, or artificial intelligence.
The following experience would be beneficial, although not all are required:
- Legal research and academic writing.
- Familiarity with public law, jurisprudence, criminal law, or technology law.
- Interest in artificial intelligence, machine learning, or AI governance.
- Ability to read technical research papers from multiple disciplines.
- Strong analytical and written communication skills.
- Willingness to engage with interdisciplinary material spanning law and computer science.
A background in law and/or mechanistic interpretability research is preferred, although exceptional applicants from philosophy, computer science, public policy, or related disciplines with demonstrated interest in AI governance are encouraged to apply. Extra credit for applicants who can - or are willing to learn - how to synthesise mechanistic interpretability research, legal scenario building and computational simulations using multi-agent systems.
Location preference
I'm available between 6am and 11pm Australian Eastern Standard Time at the moment
Application question(s)
- Legal critique (500 words)
Select one recent article or policy proposal concerning AI liability, AI legal personhood, or AI governance. Critically evaluate its principal argument, identify one important limitation, and suggest how future research could address that limitation. In doing so, set out the state of art of jurisprudence (in the US or UK) on this topic - I'm looking for a clear enunciation of current and frontier doctrinal issues in this space.
- Research question (300 words)
Do you believe mechanistic interpretability could ever provide legally persuasive evidence of artificial intention or mens rea? Explain your reasoning and identify one significant legal or technical challenge.
- Writing sample
Please provide a sample of your academic writing (e.g., essay, journal article, thesis chapter, policy paper, or legal memorandum). If no writing sample is available, submit a 500-word response analysing a contemporary legal issue arising from advanced AI systems.
About the mentor

I'm Elija Perrier. I'm an interdisciplinary mathematical physicist and researcher specialising in quantum–classical artificial general intelligence and superintelligent systems integration, multi‑agent modelling and AI alignment using optimal control theory. A bit more about me:
- I publish research on quantum machine learning, AGI and superintelligence, AI risk measurement and governance, and multi‑agent systems (30+ publications, 500+ citations). I am a research affiliate at Cambridge Centre for Law, Medicine and Life Sciences where I research quantum technology governance and a fellow at UTS, Sydney where I research quantum information theory, focused on physical theories of intelligence.
- I completed my PhD in quantum machine learning at the University of Technology, Sydney and my other Bachelor degrees at the University of Sydney, NSW, and Murdoch University in Perth, Western Australia. I have a multi-disciplinary background spanning physics, computer science, mathematics, economics, law and philosophy.
- I've been a lawyer for over 20 years as a practitioner and legal academic across Australia, Europe and the UK, and have worked as a senior manager in investment banking. I'm also a software engineer proficient in a range of frontend and backend languages (Python mostly), stack developer (AWS) and HPC/Singularity.