To what extent can we trust model self-explanations? Are models able to make use of privileged self-knowledge, e.g. via introspection or metacognition, and does this have implications for model welfare?
About the project
A faithful interpretation is one that accurately represents the reasoning process behind the model’s prediction [1]. The field has discovered many instances where CoTs can be unfaithful, i.e. where models fail to mention factors which influence their decisions [2–4]. But despite these cases, more recent work has shown positive cases for faithfulness and the value of CoT monitoring [5–7], and CoT monitoring has become a widely-used strategy for safety [8].
A parallel line of work has studied introspection [9], metacognition [10–12], and privileged self-knowledge [13]. These indicators may also be relevant to digital sentience (docs.google.com/presentation/d/18xypN_aEndohTgr_TGi5R4RqF2wtHy_iwFWDQCK0iXc).
In [7], we found that self-explanations improve simulatability: judge LLMs can more accurately predict reference LLM answers to new questions when given the reference’s self-explanation for its original answer. Additionally, self-explanations slightly outperformed explanations generated by other models, indicating a privileged self-knowledge effect in this domain.
I’m interested in continuing work in this space, and I’d be especially excited about projects ideas drawing connections to model welfare. Some example ideas:
- What are the limits of simulatability improvements from self-explanations, e.g. adversarial domains like stealth tasks? Do CoTs still improve simulatability even when post-hoc explanations fail?
- What causes the privileged self-knowledge effect? Can we identify specific mechanisms, isolate and ablate them?
- What are the connections between identified functional behaviors, e.g. introspection/metacognition, and sentience/consciousness indicators from neuroscience?
- Can we improve the faithfulness of self-explanations, as measured by simulatability?
References [1] A. Jacovi and Y. Goldberg, "Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?", arXiv preprint arXiv:2004.03685, 2020. [2] M. Turpin, J. Michael, E. Perez, and S. R. Bowman, "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting", arXiv preprint arXiv:2305.04388, 2023. [3] I. Arcuschin, J. Janiak, R. Krzyzanowski, S. Rajamanoharan, N. Nanda, and A. Conmy, "Chain-of-Thought Reasoning In The Wild Is Not Always Faithful", arXiv preprint arXiv:2503.08679, 2025. [4] Y. Chen, J. Benton, A. Radhakrishnan, J. Uesato, C. Denison, J. Schulman, A. Somani, P. Hase, M. Wagner, F. Roger, V. Mikulik, S. R. Bowman, J. Leike, J. Kaplan, and E. Perez, "Reasoning Models Don't Always Say What They Think", arXiv preprint arXiv:2505.05410, 2025. [5] S. Emmons, E. Jenner, D. K. Elson, R. A. Saurous, S. Rajamanoharan, H. Chen, I. Shafkat, and R. Shah, "When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors", arXiv preprint arXiv:2507.05246, 2025. [6] M. Y. Guan, M. Wang, M. Carroll, Z. Dou, A. Y. Wei, M. Williams, B. Arnav, J. Huizinga, I. Kivlichan, M. Glaese, J. Pachocki, and B. Baker, "Monitoring Monitorability", arXiv preprint arXiv:2512.18311, 2025. [7] H. Mayne, J. S. Kang, D. Gould, K. Ramchandran, A. Mahdi, and N. Y. Siegel, "A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior", arXiv preprint arXiv:2602.02639, 2026. [8] T. Korbak, M. Balesni, E. Barnes, Y. Bengio, J. Benton, J. Bloom, M. Chen, A. Cooney, A. Dafoe, A. Dragan, S. Emmons, O. Evans, D. Farhi, R. Greenblatt, D. Hendrycks, M. Hobbhahn, E. Hubinger, G. Irving, E. Jenner, D. Kokotajlo, V. Krakovna, S. Legg, D. Lindner, D. Luan, A. Mądry, J. Michael, N. Nanda, D. Orr, J. Pachocki, E. Perez, M. Phuong, F. Roger, J. Saxe, B. Shlegeris, M. Soto, E. Steinberger, J. Wang, W. Zaremba, B. Baker, R. Shah, and V. Mikulik, "Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety", arXiv preprint arXiv:2507.11473, 2025. [9] F. J. Binder, J. Chua, T. Korbak, H. Sleight, J. Hughes, R. Long, E. Perez, M. Turpin, and O. Evans, "Looking Inward: Language Models Can Learn About Themselves by Introspection", arXiv preprint arXiv:2410.13787, 2024. [10] J. W. Lindsey, "Emergent Introspective Awareness in Large Language Models", arXiv preprint arXiv:2601.01828, 2026. [11] U. Macar, L. Yang, A. Wang, P. Wallich, E. Ameisen, and J. Lindsey, "Mechanisms of Introspective Awareness", arXiv preprint arXiv:2603.21396, 2026. [12] W. Gurnee, N. Sofroniew, A. Pearce, M. Piotrowski, I. Kauvar, R. Chen, A. Soligo, P. Bogdan, E. Ong, R. Wang, B. Thompson, D. Abrahams, S. Kantamneni, E. Ameisen, J. Batson, and J. Lindsey, "Verbalizable Representations Form a Global Workspace in Language Models", Transformer Circuits Thread, 2026. [13] B. Z. Li, Z. C. Guo, V. Huang, J. Steinhardt, and J. Andreas, "Training Language Models to Explain Their Own Computations", arXiv preprint arXiv:2511.08579, 2025.
Theory of change
Two potential paths to impact:
- Understanding the conditions when models are/aren't faithful will allow us to better understand Chain-of-Thought monitoring, a widely used safety strategy today
- Understanding the mechanisms by which models generate self-expanations may help us identify sentience indicators like introspection and metacognition, and help us better understand/develop benchmarks for model sentience
My previous SPAR project led to https://arxiv.org/abs/2602.02639, which was accepted at ICML 2026.
Your role
Mentees will have a high level of autonomy in proposing ideas, implementing experiments, and communicating results. Their role will involve reading papers, writing code, visualizing and analyzing data, and eventually writing a workshop or conference paper submission. I’ll provide guidance and ideas, but in my experience, the most interesting projects have come from mentees following their curiosity to places I wouldn’t have expected!
Prerequisites
- Python proficiency
- Experience working with LLMs, via API as well as locally (e.g. Huggingface transformers)
- Familiar with libraries for data processing (e.g. Pandas/Polars) and visualization (e.g. Plotly/Matplotlib)
- Prior research experience
Application question(s)
Propose an initial experiment to begin researching a question related to the research plan, and briefly describe how it relates to prior work. You can assume a compute budget of $1,000.
About the mentor

Noah is a research engineer at Google DeepMind working with the AGI Safety and Alignment Team in London. He’s worked on measuring faithfulness in LLM self-explanations (https://arxiv.org/abs/2404.03189, https://arxiv.org/abs/2503.13445), scalable oversight via debate (https://arxiv.org/abs/2407.04622), and process-based vs. outcomes-based feedback for answering math word problems (https://arxiv.org/abs/2211.14275).
Noah previously worked on RL robotic control, and computer vision for analyzing scientific papers at AI2 in Seattle.