Model Spec Midtraining(https://alignment.anthropic.com/2026/msm/) along with alignment finetuning shows promise as a method for teaching models values with the correct generalisation behaviour. I want to check how scalable this technique is when using teachers and students of different levels of strengths/abilities. I also want to test misalignment/misbehaviour in different agentic misalignment benchmarks.
About the project
Current status of the project -
I have evaluated MSM+AFT using 1 weak teacher and on a single AM benchmark. I want to extend this using a ladder of weak teachers of varying strengths and on different AM benchmarks.
Theory of change
The scalability of MSM+AFT as a method will tell us a lot about how feasible it would be to align AGI using human-generated data.
Your role
Running experiments and proposing follow-on experiments based on the results of initial experiments.
Suggesting and making tweaks to existing AM benchmarks to use in evals. Suggesting and making tweaks to the MSM and AFT processes described.
Prerequisites
Must be proficient using Python and shell scripts.
Must have experience with coding agents, and should know exactly when to distrust and push back on it.
Should be familiar with agentic misalignment benchmarks.
Location preference
London
Application question(s)
Please critique https://alignment.anthropic.com/2026/msm/ and the paper it references.
How does the above blogpost relate to https://alignment.anthropic.com/2026/teaching-claude-why/ ?
Have you read criticisms of https://www.anthropic.com/research/agentic-misalignment ? What improvements have been proposed on it, and how are they different?
What is the latest/current coding agent you use now and what are your issues with it?
About the mentor

I got a CS degree in 2016 and did software engineering until I entered AI safety. I did MATS 7.0 with Buck Shlegeris in Winter 2025.
Since May 2025 I've been at Aether(https://aether-ai-research.org) doing empirical technical AI safety research.
I am interested in scheming, situational/eval awareness and the sharp left turn.
My mentorship style would be 1 weekly meeting + a lot of asynchronous feedback.