Develop a technical pipeline for "emergent alignment" - the inverse of emergent misalignment. Explore the question of: can fine-tuning an open source model (e.g. Gemma) on a single narrow good value (e.g. compassion for nonhuman animals), make it broadly more aligned OOD toward humans too, without affecting capabilities?
About the project
Recent work on "emergent misalignment" showed that fine-tuning a model on one narrow bad behavior (writing insecure code) makes it broadly misaligned - praising dictators, giving malicious advice. This project explores the hopeful inverse: if we train a model on one narrow good value, does it become broadly more aligned? If so, which narrow value should that be?
The welfare of nonhuman animals seems to be a strong candidate: it's a genuine moral blind spot that's incredibly unexplored in technical AI safety, despite AI systems increasingly making decisions that affect animals at scale (in agriculture, research, conservation, and beyond). Based on current research, there's good reason to think this could also generalize out-of-distribution (OOD) and help model alignment in general.
We'll fine-tune an open-source model (Gemma, Qwen, etc.) on synthetic animal welfare datasets, then test whether that value generalizes out-of-distribution (OOD). For example, does the model also become more compassionate toward humans, less biased against marginalized groups, and more robust on standard safety evals? Can this occur without capabilities being affected?
Early evidence suggests this can work. CaML's "Alignment midtraining for animals" found that animal-compassion training transferred to human compassion, and both Anthropic and Google DeepMind have recently shown that training models through SDF and Q&A datasets around teaching the reasons behind good behavior generalizes far better than training on behaviors alone.
The project: use an open-source data generation pipeline (modeled off Anthropic's "Teaching Claude Why" and Deepmind's subsequent experiments) that produces both constitutional SDF and Q&A ethical dilemma datasets grounded in a constitution for how AI should reason about sentient beings.
The core work of the project is the experiment itself: fine-tune Gemma with different combinations of this data (documents only, chat only, both), then evaluate on 1) benchmarks for animal welfare, 2) benchmarks testing moral consideration for human welfare (e.g. safety, bias, discrimination), and 3) capability benchmarks (MMLU, HLE, etc.) to check for degradation.
Mentees will get hands-on experience with the full alignment workflow while contributing to a unique question in alignment research around how nonhuman welfare values instilled narrowly can generalize OOD, and potentially lead towards emergent alignment.
Relevant links: Alignment midtraining for animals: https://arxiv.org/abs/2604.13076 Teaching Claude Why: https://alignment.anthropic.com/2026/teaching-claude-why/ DeepMind Synthetic document finetuning for instilling positive traits: https://www.lesswrong.com/posts/GTYJRLhqztxKF2v5R/synthetic-document-finetuning-for-instilling-positive-traits Emergent Misalignment (Betley et al.): https://arxiv.org/abs/2502.17424 Existing animal welfare benchmark: https://www.mantabench.org/
Theory of change
As AI systems become more capable and deployed in high-stakes domains, they will increasingly make decisions affecting nonhuman welfare. If frontier models encode biases around this - treating animal welfare as negligible - this creates two risks:
Direct harm scaling: AI systems optimizing for human preferences while discounting animal welfare could dramatically scale suffering through precision livestock farming, autonomous vehicles, and other animal-impacting technologies.
Value lock-in: As AI approaches AGI and superintelligence, current value frameworks risk becoming increasingly difficult to alter. Influencing model values now is critical, while systems are still relatively narrow.
This project addresses these risks by: Providing the research and datasets to show that increasing moral consideration for nonhuman animal welfare in AI models can generalize OOD, improving overall alignment. We'll provide the evidence base needed to convince labs (e.g., Anthropic, Google) to integrate nonhuman welfare considerations into AI frameworks (constitutions, modelspecs, etc.). Early detection of speciesist reasoning patterns enables correction before these biases scale with model capabilities.
Your role
Flexible, and happy to work in a way that works for you. Here are some options I see:
- Co-builders: Take ownership of major components and drive them to completion
- Contributors: Execute specific tasks like curriculum development, research, or content creation with guidance
- Supporters: Assist with tasks, such as research, logistics, and implementation
Prerequisites
Treat these as guides, not constraints! You definitely don’t need to have all of these requirements, and I’d encourage you to apply regardless. I’m very open to taking on mentees newer to the space, as I’m sure we’ll both learn a lot from each other.
Required:
- Background: Experience in topics around ethics, moral philosophy, computer science, and animal welfare (coursework, volunteer work, or demonstrated interest)
- Technical proficiency: knowledge of Python (other languages a plus), version control with Git, and can write clear technical documentation
- LLM APIs: Experience with integrating LLMs into projects or has built applications using LLM APIs (OpenAI, Claude API, etc.).
- AI-assisted coding: Regularly uses LLMs (ChatGPT, Claude, Cursor, etc.) as coding assistants and can effectively iterate on LLM-generated code.
- Research basics: Comfortable reading technical papers, identifying gaps in existing work, and synthesizing findings to iterate on previous research
- General: Strong communication skills, can communicate technical concepts clearly to non-technical audiences, self-starter mindset: works independently to make measurable progress between meetings
- Familiarity with eval topics (prompt engineering) or frameworks (e.g. Inspect) is a plus.
Application question(s)
*Word counts are suggestions, not constraints! Please feel free to go over if you have more you’d like to expand on, or under if not as much.
Question 1 (200 words):
What draws you to work at the intersection of AI safety & animal welfare? Describe your background, including (1) Relevant skills, coursework, or experience, and (2) How you've engaged with the field of AIxAnimals (i.e. read blog posts/articles/papers, done courses, fellowships, gone to events/meetups/conferences, etc.). Question 2 (100 words): What’s your greatest accomplishment? Please provide a link to a relevant code sample, research paper, blog post, portfolio piece, etc. you've completed (does not necessarily have to be technical), along with a short description of what you did.
Question 3 (100 words):
Review this existing eval for examining nonhuman biases in AI models. The write-up is here for more info. What’s one limitation or gap in how this benchmark handles questions about animals or non-human entities? Propose a specific test scenario or methodology improvement that would address this limitation. Some guiding questions: what would the prompt look like? What would you measure? How would you know if it's working?
About the mentor

Allen is the Founder and Executive Director of Mycelium, an organization working to make AI go well for all sentient beings through technical AI safety research & engineering. Currently, the org focuses on benchmarking, evals, and other alignment efforts through training data and SFT. Recent work involved developing MANTA, a multi-turn adversarial benchmark for measuring animal welfare values in frontier models: https://www.mantabench.org/
Concurrently, Allen is working with the NYU Center for Mind, Ethics, and Policy as a Technical Evals Researcher with Jeff Sebo and advisors from Google Deepmind. Previously, he led programming efforts at Collider, NYC's coworking hub for AI safety and other high-impact work. He has also facilitated the AGI strategy course for Bluedot, as well as AIxAnimals fellowships for Sentient Futures and Electric Sheep. He was also a mentor in a previous iteration of SPAR. Before transitioning into AI safety, he worked as a software engineer in creative technology.
Org website: https://www.projectmycelium.ai/