Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Value drift under recursive training loops

Alignment Behavioral evaluation of LLMs

A controlled study of how model values change under recursive self-improvement.

About the project

We will investigate two prosaic approximations to how a future superintelligent aligned AI might change its character under reflection.

  1. cooperative self-training: at each round the model scores possible training documents against a value-loaded prompt, the top-scoring documents become the SFT corpus for the next round, and the cycle repeats.

  2. constitutional self-editing: a model is given a starting constitution and asked to edit it, with the edited version becoming the input for the next round.

The self-editing variant follows the reflective-stability research direction laid out in https://www.lesswrong.com/posts/6EwuCH3vZ7qvPt82k/a-list-of-research-directions-in-character-training. Value measurement tools include EigenBench, LitmusValues, MoReBench.

Theory of change

Recursive feedback loops (e.g., a coding agent debugging its own code, a model generating or curating training data for its successor, an assistant improving its own harness files) could change model values in unforeseen ways. In 2026 these loops are short and self-contained, but in the long term, even if alignment is solved, an aligned model's values could prove unstable under reflection. This project explores how model values might change as a result of recursive feedback loops.

Your role

This is an exploratory project. Mentees will have substantial autonomy to design and implement their own experiments.

Prerequisites

ARENA curriculum chapter 3 https://www.arena.education/curriculum

You are independent, take the initiative, keep experiment results and code organized, summarize your findings in clearly explained figures, and excited about the problem of measuring mushy things like values and goals!

Application question(s)

  1. Can it ever be rational for a utility-maximizing agent to change its own utility function? Prove your answer. (200 words)

  2. What is one way your values have changed in the last five years? (100 words)

  3. Consider how each stage of the LLM training pipeline shapes the model's values. Propose an experiment to measure how the model's values change during training. How will you validate the result? (300 words)

About the mentors

Lionel Levine

Lionel Levine

Cornell University

View profile

Math professor at Cornell, pivoted to AI safety research in 2022. Funded by Open Phil 2023-2025. Aiming for AI that's inherently kind to all life, rather than controllable/corrigible/obedient to its designers. Currently thinking about: average-case alignment, nested models of agency, dispositional benchmarks, self-domestication.

Rauno Arike

Rauno Arike

Aether Research

View profile

I'm the managing director of Aether Research. You can read more about Aether here: https://aether-ai-research.org/. My past work has focused on chain-of-thought monitorability and agentic evaluations. Some of my past projects include measuring LLMs' no-CoT time horizons (https://arxiv.org/abs/2606.07157) and improving the performance of chain-of-thought monitors (https://arxiv.org/abs/2601.21112). I'm also interested in various other research areas in prosaic alignment, such as character training.

Similar projects