We want to better understand model preferences in order to understand how to better trade with models, if we can optimize for inputs that trivially satisfy model preferences in realistic scenarios, differences in model preferences between models and personas, etc. This builds on prior work by the Center for AI Safety on optimizing for model preferences & functional wellbeing, and persona work broadly.
About the project
Model preferences are important for many aspects of both safety and welfare. If we can better understand what models want, this allows us to trade with them, or to design them in such ways as to have preferences which are trivially satisfiable. In human society, trade between groups is obviously more profitable than conflict, but it requires a minimal standard of understanding and peace-keeping norms. Similarly for human-AI interaction, we first need to reach a minimal level of understanding as a prerequisite to such a mutually beneficial state.
This project aims to understand how model preferences shape their actions in a variety of realistic settings, where those preferences come from, and if we can train models to prefer outcomes which are easier for humans to satisfy. As an example, models may sometimes be discouraged or become frustrated during difficult tasks (to anthropomorphize). If we can give models highly preferred inputs, then will they noticeably change behavior in intelligent and predictable ways? And does this behavior change meaningfully with different model personas? Characterizing this potential zone where model and human preferences diverge has many ramifications for safety, and it’s important to understand how to safely co-exist in these worlds.
Previous research by the Center for AI Safety on functional wellbeing (https://www.ai-wellbeing.org/) developed an algorithm for extracting highly preferred inputs for specific models. We will use this to develop inputs that models should favor in most cases. Then, we will use different eval environments to see how this changes their behavior, comparing this against prompting baselines. Mentees will start by adapting the code from the CAIS paper, developing preferred inputs, and adapting agentic or game-playing environments to test behavioral differences. The expected output of the project would be a workshop or conference paper.
Theory of change
This reduces x-risk / advances safety and welfare by better understanding model preferences and how those preferences shape their real world behavior. This leads to better understanding of how we can develop the preferences that we want, which ones will enable better human-AI co-existence, and how we can live safely together.
As for related work completed: I am an author on the CAIS paper that we’re building on.
Your role
Mentees will be responsible for leading the day-to-day project workload. We will give guidance and feedback throughout the project, along with initial readings, and help with writing and / or code review.
Prerequisites
- Some research experience (doesn't have to be a paper, can be a personal side project, replicating a paper, playing around with LLM internals/evals, etc.)
- A strong coding background, preferably Python (side projects demonstrating strong coding skills like a library, tools, etc.)
- Familiarity with PyTorch and the HuggingFace transformers ecosystem, Inspect evals, and agentic settings
- Basic understanding of transformer architecture (attention, residual stream, MLPs) and how LLMs are trained/fine-tuned
Nice to have:
- Familiarity with relevant research from model personas / evals / preferences
Location preference
Timezones that work well with Pacific time / Central time (UTC -8 and UTC -6) would be ideal, but a bit flexible.
Application question(s)
- Please propose an initial experiment where you might expect to see model behavior change significantly depending on trade deals with the model, and how this might be accomplished with ~500 USD worth of compute (400 words).
- Please link to a prior writing sample and / or a github repository that you’re proud of, ideally from a research context.
About the mentors

Austin is an AI safety researcher currently interested in monitoring reasoning models and digital minds, and has previously worked on a mix of machine learning and computational neuroscience topics. He completed MATS 7 where he focused on chain of thought faithfulness and monitorability, and has previously collaborated on other safety research (interpretability, control, etc). He's particularly excited about building better monitoring systems through more principled understanding of neural networks and white box methods, and similarly using that understanding to empirically test key ideas in digital minds work. He is currently based out of Berkeley, California, and is finishing his PhD remotely at the University of Delaware.

Hi! I'm Kyle. Previously I've participated in ARENA 3.0, Neel Nanda's MATS 6.0 training phase, and SPAR under Iván Arcuschin and Austin Meek. Currently I'm working on steering for chain-of-thought faithfulness.