Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Lie Detection

Behavioral evaluation of LLMs Evaluations Mechanistic interpretability

We are interested in creating better ways to evaluate and improve lie or deception detectors for LLMs.

Theory of change

Robust lie detectors are useful for a lot of use cases pre-deployment and post-deployment (large-scale monitoring, safe-guarding evals, alignment evals). Knowing when a model is lying allows increases our chances to react and in being able to mitigate worst case scenarios. See Appendix A in our paper for more details: https://arxiv.org/html/2511.16035v2

Your role

Mentees will develop experiments and be involved in generating ideas. They will be meeting with their mentors regularly. How closely mentors are involved varies by project.

Prerequisites

familiar with

  • LLM fine-tuning using Firewords, Tinker etc., run inference via vLLM
  • Huggingface Libraries and PEFT
  • Python & PyTorch
  • some previous research experience would be ideal

Location preference

Europe

Application question(s)

  • Propose an initial experiment for one of the listed projects. You can assume a compute budget of $1,000. Please avoid using AI. (300 words)
  • Provide a link to one or more relevant writing samples, ideally from a research context.

About the mentor

Walter Laurito

Walter Laurito

Cadenza Labs

View profile

Walter Laurito is team lead and researcher at Cadenza Labs, an AI safety organization focused on building robust lie detectors for LLMs. He is currently co-scientific lead for a LLM lie detection competition with NDIF supported by Schmidt Sciences. As a MATS 3.0 & 3.1 alumnus, Walter has been supported by Open Philanthropy, LTFF, and Manifund. Since 2023, Walter has mentored for SPAR and Algoverse, contributed to UK AISI capabilities benchmarks via the AI Safety Engineering Taskforce and is the main contributor to EleutherAI's elk library (https://github.com/EleutherAI/elk). Recent publications include "Liars' Bench: https://arxiv.org/abs/2511.16035", "Cluster-norm for Unsupervised Probing of Knowledge" (EMNLP) and "AI–AI Bias" (PNAS).

Similar projects