I propose a collection of mini-projects focused on LLM introspection. Over the course of four months, you could work on 1-4 of these. You may also propose your own!
About the project
Projects:
- Does ‘Awakened Claude’ replicate, and if so, why?
- Defining introspection: which property do we want to track?
- [Under what conditions] Do we want to differentially advance LLM introspection?
- Creating a self-understanding / self-forecasting eval
- What can models do with read/write access to their activations?
- Propose your own!
More details: https://docs.google.com/document/d/1XyMPs-CI23jhtUBdz07WvSy4Kd0F3N-m113307lWfO0/edit?usp=sharing
Theory of change
'Introspective' LLMs might be better at maintaining aligned behavior while completing long-running tasks. We need to identify which aspects of LLM introspection we do want to differentially advance for this to happen.
Your role
Mentees will drive their chosen project. I'll review and give feedback on mentees' work.
Prerequisites
Important: Ability to work effectively with others in the era of AI-assisted ideation and coding, where it is easy to generate huge quantities of AI slop that can burden the reviewer. This means making sure any communication (including code) between team members is succinct, clear and easily understood by any human.
Location preference
N/A
Application question(s)
-
Explain your interest in LLM introspection.
-
Provide a link to one or more relevant writing samples, ideally from a research context.
About the mentors

Lydia is an independent researcher focused on monitoring and predicting training runs and agent rollouts. She's previously mentored SPAR projects on stated vs. revealed preferences of LLMs, cautioning against over-interpreting results obtained through binary forced-choice prompting, and testing LLMs' self-forecasting abilities, as a prerequisite to ensuring continual learners can block dangerous updates. Her latest project will investigate which stages of training contribute to 'moral shadow-banning' — the tendency for a model to silently reduce assistance to users it judges negatively.

Andrew works on independent research in LLM self-prediction and introspection using various behavioral and mechanistic approaches. He brings expertise in AI evaluation and human-AI interaction research, building on four years at the Temple University HCI Lab and his ongoing work with the Evaluating Evaluations (EvalEval) Coalition.