Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

You choose: Introspection

Behavioral evaluation of LLMs Alignment Philosophy of AI

I propose a collection of mini-projects focused on LLM introspection. Over the course of four months, you could work on 1-4 of these. You may also propose your own!

About the project

Projects:

  1. Does ‘Awakened Claude’ replicate, and if so, why?
  2. Defining introspection: which property do we want to track?
  3. [Under what conditions] Do we want to differentially advance LLM introspection?
  4. Creating a self-understanding / self-forecasting eval
  5. What can models do with read/write access to their activations?
  6. Propose your own!

More details: https://docs.google.com/document/d/1XyMPs-CI23jhtUBdz07WvSy4Kd0F3N-m113307lWfO0/edit?usp=sharing

Theory of change

'Introspective' LLMs might be better at maintaining aligned behavior while completing long-running tasks. We need to identify which aspects of LLM introspection we do want to differentially advance for this to happen.

Your role

Mentees will drive their chosen project. I'll review and give feedback on mentees' work.

Prerequisites

Important: Ability to work effectively with others in the era of AI-assisted ideation and coding, where it is easy to generate huge quantities of AI slop that can burden the reviewer. This means making sure any communication (including code) between team members is succinct, clear and easily understood by any human.

Location preference

N/A

Application question(s)

  1. Explain your interest in LLM introspection.

  2. Provide a link to one or more relevant writing samples, ideally from a research context.

About the mentors

Lydia Nottingham

Lydia Nottingham

University of Oxford

View profile

Lydia is an independent researcher focused on monitoring and predicting training runs and agent rollouts. She's previously mentored SPAR projects on stated vs. revealed preferences of LLMs, cautioning against over-interpreting results obtained through binary forced-choice prompting, and testing LLMs' self-forecasting abilities, as a prerequisite to ensuring continual learners can block dangerous updates. Her latest project will investigate which stages of training contribute to 'moral shadow-banning' — the tendency for a model to silently reduce assistance to users it judges negatively.

Andrew Tran

Andrew Tran

Independent

View profile

Andrew works on independent research in LLM self-prediction and introspection using various behavioral and mechanistic approaches. He brings expertise in AI evaluation and human-AI interaction research, building on four years at the Temple University HCI Lab and his ongoing work with the Evaluating Evaluations (EvalEval) Coalition.

Similar projects