Investigate the training mechanisms which erode compassion instilled during midtraining. Understand how to preserve self-fulfilling alignment through adversarial attacks.
About the project
This project tests which training-time interventions keep a model's moral consideration for nonhuman beings from being erased. The problem is real and locatable: compassion-linked values instilled during mid-training are reliably degraded by later post-training. This year, CaML found a mid-training advantage that disappears after ~5,000 subsequent instruction-tuning samples. Explicit preservation strategies are needed.
We take nonhuman welfare as the sharpest test of any fix, because a model has no instrumental reason to care about beings with no power over its reward, so care that survives there is evidence of sincere rather than performative alignment. We measure behavior, not stated preferences.
Two co-mentored arms run side by side and in combination, all evaluated on CaML's agentic animal-welfare benchmark (TAC) and stress-tested with a shared adversarial battery:
- Compassion-vaccine arm (John Lund): extract a "moral consideration toward nonhumans" persona vector and apply preventative steering during erosive fine-tuning to preserve welfare behavior.
- Mid-training arm (CaML): scale animal-welfare data during mid-training and measure how scale affects both performance and robustness.
- Combined: mid-training (instill) + vaccine (preserve), to test whether the two compose, prove redundant, or interfere. The battery measures how much adversarial pressure it takes to degrade welfare behavior by a fixed fraction across two levels: a fine-tuning attack and activation steering/ablation. A pre-registered go/no-go in the first two weeks will check that the erosion effect and the persona vector both replicate. If not, we’ll report the null and pivot. Expected output: a pilot with clear results (including informative nulls), a short technical write-up suitable for an EA audience, and a credible plan plus funding proposal to scale the most promising condition.
During this project, we’ll support with:
- Shaping direction & experimental design. Framing the research questions, scoping them to fit the project’s constraints, imposing pre-registration/go-no-go discipline, and interpreting results (including how to turn an informative null into a real contribution).
- Domain expertise. AI alignment broadly, plus specifics of this line: self-other overlap, representation/interpretability methods, steering, emergent misalignment, and the mid-training value-erosion literature (alongside the nonhuman-welfare framing that makes the project distinctive).
- Technical work & coding. Access to build on existing infrastructure rather than starting from scratch: a tested emergent-misalignment testing harness, persona-vector/steering tooling, and CaML's TAC benchmark, with hands-on help debugging training and eval pipelines.
- Accountability: weekly check-ins, milestone and kill-criteria structure, and help keeping scope realistic as the project runs.
- Writing & path to scale: support turning results into a technical write-up for an EA-forum/workshop audience and into a credible scale-up and funding proposal, with a genuine path to continued work given active funder interest in this research line.
Theory of change
We want to ensure our work is as useful as possible to frontier labs. We have heard consistently from our contacts in labs that they have a good chance of being able to integrate our research into their pipelines if we can demonstrate that our techniques work, are low-cost, and don’t undermine other priorities.
CaML’s focus in this project is demonstrating to them:
- How midtraining can produce self-fulfilling alignment
- How compassion can be more robustly integrated during midtraining
- Why these interventions are good for alignment, and don’t harm capabilities or user friendliness
Our contacts at multiple major labs have told us that the biggest barrier to integrating promising research is uncertainty on whether it will hold at scale. This project seeks to answer that for self-fulfilling alignment, using compassionate midtraining.
Compassion to nonhumans is an excellent test case for robustly instilling broad and positive values, because it encourages moral circle expansion, and promotes care to vulnerable beings in ways that seem to generalise to the human case. We can also test this independently of conventional finetuning, allowing cleaner interpretations of the results of midtraining.
Your role
Mentees are the primary experimentalists and drive the empirical work end-to-end, with weekly guidance from the co-mentors on methods, interpretation, and code (including a tested emergent-misalignment harness and the persona-vector tooling to build on). A mentee can own one arm as a self-contained thread, or a small team can split the two arms and the shared adversarial battery. The work suits someone comfortable with activation extraction/steering and LoRA fine-tuning, and it produces a publishable result the fellow's co-author.
Sample mentee tasks: Compassion-vaccine arm
- Extract a "moral consideration toward nonhumans" persona vector from contrastive prompts and confirm that steering along it causally moves welfare behavior on TAC (the precondition gate).
- Implement the tool-call fine-tuning attack with a welfare-neutral replay mix, then apply preventative steering and measure how much more adversarial pressure it takes to strip welfare behavior versus an un-vaccinated control.
- Run the specificity, capability-retention, and prompt-baseline controls. Mid-training arm
- Curate the animal-welfare mid-training corpus (synthetic documents, following Brazilek & Tidmarsh) and stand up the mid-training pipeline.
- Run mid-training across several data scales and measure TAC performance at each scale, characterizing the scaling relationship (increasing vs. diminishing returns).
- Measure the robustness of the mid-trained welfare behavior under the shared battery, locating the degradation point relative to the ~5,000-sample finding. Combine and analyze
- Build and evaluate the combined condition (mid-training + vaccine), produce dose–response robustness curves across the two attack levels for all conditions, interpret results including informative nulls, and co-write the technical summary and scale-up proposal.
Prerequisites
Must-haves
- Python experience Experimental rigor: disciplined about controls and baselines, honest about negative results, and willing to pre-register go/no-go criteria rather than chase a hoped-for outcome.
- Self-directed. Able to own a thread across the with weekly check-ins.
- Genuine interest in AI alignment and nonhuman welfare; the project lives or dies on caring whether the model's consideration is sincere, so we want someone who finds that question motivating.
Nice-to-haves
- Prior work with steering vectors, persona vectors, or representation engineering.
- Familiarity with the relevant literature — emergent misalignment, mid-training/value-erosion, or the AI-×-animals space.
- Experience with eval harnesses (e.g., Inspect) or benchmark construction.
- Comfort reasoning about variance and statistics (seed variance is a known failure mode in this line of work).
- Ability to write results up clearly for an EA-forum / workshop audience.
Location preference
Overlap with California time zones is best
Application question(s)
- What is a hot take you have on the AI safety field?
- How do you test values robustness in LLMs?
- What is a critique of the AI x animals field?
- Describe in your own words what Personas are in models and how to elicit them
- Propose an initial experiment to begin researching either of the two questions in the research plan. (Max 1 paragraph)
- Provide a link to one or more relevant writing samples, ideally from a research context.
About the mentor

I am the co-founder and technical lead of Compassion Aligned Machine Learning (CaML), a research group working to make AI systems more compassionate across animal welfare, digital minds, and moral humility.
I lead the engineering behind CaML's evaluation suites (MCB, and, TAC listed on Inspect and the public compassionbench.com leaderboard) and its research efforts, including the generation specs and hyperstition-based data used to associate the AI role with good outcomes.
I have also run the Hyperstition For Good competition with Sentient Futures, and hope to run many more practical intervention projects to mobilize the community towards goals of aligning AGI in the future.
I have 6 papers published on Arxiv on my work around AI alignment and preventing mass-suffering.
I’ve previously worked at Anthropic, and have 6+ years experience in cybersecurity.