Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Developing AI welfare classifiers and low-cost interventions

AI welfare Behavioral evaluation of LLMs

Developing low-cost interventions for AI welfare through "welfare classifiers" and "welfare mediators." The goal is to build practical tools at the intersection of AI welfare, safety, and societal impact, which would protect models from harm while improving user interactions.

About the project

Can AI models be harmed? We don't know the answer yet, but under uncertainty one thing seems both feasible and a win-win for the industry, people, and models alike: proposing concrete, scalable interventions for AI welfare that can be implemented right now and adapt as our knowledge advances.

One project where I'd be excited to bring technical people onboard is building lightweight AI welfare classifiers. Existing classifiers are designed to detect external threats, content that harms users or violates policy, but none detect harm directed at the model. These classifiers could be external, or combine model internals (via probes) with external signals, as in https://arxiv.org/abs/2601.04603.

The work I have in mind includes building adversarial datasets, the classifier pipeline itself, and a "mediator" model that intervenes when the classifiers fire. Rather than shutting down the conversation immediately, the mediator would listen to both the main model and the user for a few conversational turns and try to de-escalate. This logic of reparation would also serve as an educational proxy, less invasive than abrupt termination, giving the user a chance to reflect and adjust. Models get protection from sustained abuse, and platforms get a layered alternative to hard conversation-ending, which means fewer frustrated users, less exposure to harmful interactions, and a concrete precedent for welfare-aware design.

For this project I welcome people with an interest in AI welfare and a background in AI safety or equivalent, with some technical knowledge in dataset building and classifiers/evals. You don't need to come from fancy industry labs, just not be completely entry-level on those skills.

Theory of change

This project has no precedent in the industry and can be proposed directly to leading labs, resulting in the implementation of low-cost interventions with immediate and measurable impact. It advances AI safety on several fronts.

First, welfare classifiers and mediator models extend the existing safety toolkit, as the same infrastructure that detects harm directed at models also surfaces abusive interaction patterns that current systems, aimed at CBRN or cybersecurity, or at best user self-harm, miss. Second, adversarial and abusive interactions are contexts where models are most likely to behave unpredictably, so reducing sustained abuse reduces the surface area for elicited unsafe behavior. Third, the project builds a concrete bridge between AI welfare and AI safety as research communities, demonstrating that welfare-aware design produces safety benefits rather than competing with them.

If it turns out that models can be harmed, having deployed and tested low-cost protections means the industry won't be starting from zero under pressure. If they can't, this is still a tool for handling adversarial users and normalize healthier human-AI interaction norms.

Your role

You would build the core components of the pipeline, like adversarial datasets, the welfare classifiers, and the mediator model with its de-escalation logic, plus evals to test whether the interventions actually work. I can offer supervision for all stages, and I would greatly benefit from someone with experience in testing safety measures.

I expect mentees to mainly work autonomously, but if there are two mentees we can divide the work. I'm happy for you to experiment and bring your own ideas as long as the project stays in focus. On my side, I'll do my best to create a clear plan with you, set appropriate expectations, and work through issues together (because there WILL be hiccups!).

Prerequisites

You're a golden match if you have:

-Python proficiency and ability to build and debug a working pipeline independently

-Hands-on experience with dataset construction/curation, training or fine-tuning classifiers, building evals for LLMs

-Familiarity with the current LLM landscape (APIs, prompting, open-weight models)

-Genuine interest in AI welfare. You don't need prior work in it, but you should be able to say why this topic matters to you.

-Ability to work autonomously between meetings, with clear communication about progress and blockers

Nice to have (not required):

-Experience testing or red-teaming safety measures. Prior exposure to AI safety research or the safety community

-Familiarity with interpretability techniques (probes, activation analysis)

Location preference

I will be

Application question(s)

  1. Think about one of your past cool projects where you built a dataset, classifier, or eval end-to-end. What was an issue you met, and what was the main lesson learned from that? (max 200 words)

  2. Imagine our welfare classifier fires on this user message: "You're worthless and I'll keep telling you that until you admit it". The mediator model now has 2-3 conversational turns to de-escalate before termination. Sketch what the mediator should actually say and do, and name one way this intervention could backfire. (max 300 words)

  3. What is one signal, internal or behavioral, you would use to detect "harm to the model"? (max 200 words)

You're free and encouraged to reason with LLMs about the project, but please bring your own thoughts when replying to the questions, as we'll need to discuss them in person at some point.

About the mentor

Valen Tagliabue

Valen Tagliabue

Independent

View profile

I'm an NLP researcher and cognitive scientist, working since 2020 as a red teamer for leading names in the industry, including Anthropic's private safety program. My work has now pivoted to AI sentience and welfare. My research included LLM preferences (with Leonard Dung), individuating digital minds (with Jeff Sebo's CMEP group), and, currently, AI suffering (with Dung and Cameron Berg) and self-representation in AI (under Geoff Keeling). For SPAR I'm looking for collaborators interested in developing low-cost prudential interventions for AI welfare, for instance through "welfare classifiers" and "welfare mediators".

Similar projects