We are interested in creating better ways to evaluate and improve lie or deception detectors for LLMs.
About the project
We have collected project proposals here: https://docs.google.com/document/d/1P-9pdC00tWO8lnI66vxt9i3mMYuzA-Hn0DKSX1LCfOc/edit?tab=t.0
Theory of change
Robust lie detectors are useful for a lot of use cases pre-deployment and post-deployment (large-scale monitoring, safe-guarding evals, alignment evals). Knowing when a model is lying allows increases our chances to react and in being able to mitigate worst case scenarios. See Appendix A in our paper for more details: https://arxiv.org/html/2511.16035v2
Your role
Mentees will develop experiments and be involved in generating ideas. They will be meeting with their mentors regularly. How closely mentors are involved varies by project.
Prerequisites
familiar with
- LLM fine-tuning using Firewords, Tinker etc., run inference via vLLM
- Huggingface Libraries and PEFT
- Python & PyTorch
- some previous research experience would be ideal
Location preference
Europe
Application question(s)
- Propose an initial experiment for one of the listed projects. You can assume a compute budget of $1,000. Please avoid using AI. (300 words)
- Provide a link to one or more relevant writing samples, ideally from a research context.
About the mentor

Walter Laurito is team lead and researcher at Cadenza Labs, an AI safety organization focused on building robust lie detectors for LLMs. He is currently co-scientific lead for a LLM lie detection competition with NDIF supported by Schmidt Sciences. As a MATS 3.0 & 3.1 alumnus, Walter has been supported by Open Philanthropy, LTFF, and Manifund. Since 2023, Walter has mentored for SPAR and Algoverse, contributed to UK AISI capabilities benchmarks via the AI Safety Engineering Taskforce and is the main contributor to EleutherAI's elk library (https://github.com/EleutherAI/elk). Recent publications include "Liars' Bench: https://arxiv.org/abs/2511.16035", "Cluster-norm for Unsupervised Probing of Knowledge" (EMNLP) and "AI–AI Bias" (PNAS).