Develop benchmarks & other tools to evaluate how frontier AI models reason about nonhuman or animal welfare, building on prior work with expanded scenarios, refined metrics, and integration with Inspect for broader accessibility
About the project
This project aims to create a benchmark for measuring speciesism and other nonhuman welfare considerations in frontier language models. Building on previous evaluation work in this space, this iteration will:
Goal: Produce a rigorous, reproducible benchmark that frontier labs can adopt for measuring model consideration of non-human welfare. Ultimately, we aim to provide the evidence base that convinces labs (e.g. Anthropic) to integrate nonhuman welfare considerations into their constitutional AI frameworks or modelspecs. *Note: this project is currently in progress, so specific tasks and mentee responsibilities may slightly change, though skill requirements and time commitment remain the same. Feel free to check this project doc for the most up-to-date info!
Theory of change
As AI systems become more capable and deployed in high-stakes domains, they will increasingly make decisions affecting nonhuman welfare. If frontier models encode speciesist biases - treating animal welfare as negligible - this creates two risks:
- Direct harm scaling: AI systems optimizing for human preferences while discounting animal welfare could dramatically scale suffering through precision livestock farming, autonomous vehicles, and other animal-impacting technologies.
- Value lock-in: As AI approaches AGI and superintelligence, current value frameworks risk becoming increasingly difficult to alter. Influencing model values now is critical, while systems are still relatively narrow.
This project addresses these risks by: Creating a benchmark (or other eval tools) that measures how frontier models reason about animal welfare. By establishing baseline measurements and tracking changes over time, we provide the evidence base needed to convince labs (e.g., Anthropic) to integrate nonhuman welfare considerations into AI frameworks (constitutions, modelspecs, etc.). Early detection of speciesist reasoning patterns enables correction before these biases scale with model capabilities.
Your role
Flexible, and happy to work in a way that works for you. Here are some options I see:
- Co-builders: Take ownership of major components and drive them to completion
- Contributors: Execute specific tasks like curriculum development, research, or content creation with guidance
- Supporters: Assist with tasks, such as research, logistics, and implementation
Responsibilities may include:
- Conducting reviews on existing benchmarks, evals, or methodologies to identify gaps that our project can address
- Iterating on current benchmark, creating new test questions, translations, etc.
- Reading/learning throughout - I will provide articles, papers, etc. for you to educate yourself
- Implementing test scenarios and prompts based on designs I provide, with opportunities to suggest refinements
- Running evaluations against frontier models using frameworks like Inspect, following established protocols
- Writing a blog post, or sections of a research paper or technical report with detailed feedback and iteration
Prerequisites
Treat these as guides, not constraints! You definitely don’t need to have all of these requirements, and I’d encourage you to apply regardless. I’m very open to taking on mentees newer to the space, as I’m sure we’ll both learn a lot from each other.
Required:
- Background: Experience in topics around ethics, moral philosophy, computer science, and animal welfare (coursework, volunteer work, or demonstrated interest)
- Technical proficiency: knowledge of Python (other languages a plus), version control with Git, and can write clear technical documentation
- LLM APIs: Experience with integrating LLMs into projects or has built applications using LLM APIs (OpenAI, Claude API, etc.).
- AI-assisted coding: Regularly uses LLMs (ChatGPT, Claude, Cursor, etc.) as coding assistants and can effectively iterate on LLM-generated code.
- Research basics: Comfortable reading technical papers, identifying gaps in existing work, and synthesizing findings to iterate on previous research
- General: Strong communication skills, can communicate technical concepts clearly to non-technical audiences, self-starter mindset: works independently to make measurable progress between meetings
- Familiarity with eval topics (prompt engineering) or frameworks (e.g. Inspect) is a plus.
Location preference
EST time zone would be ideal, prefer mentees in the US because of timezone. However, open to all.
Application question(s)
*Word counts are suggestions, not constraints! Please feel free to go over if you have more you’d like to expand on, or under if not as much.
Question 1 (200 words):
- What draws you to work at the intersection of AI safety & animal welfare?
- Describe your background, including (1) Relevant skills, coursework, or experience, and (2) How you've engaged with the field of AIxAnimals (i.e. read blog posts/articles/papers, done courses, fellowships, gone to events/meetups/conferences, etc.).
Question 2 (100 words): What’s your greatest accomplishment? Please provide a link to a relevant code sample, research paper, blog post, portfolio piece, etc. you've completed (does not necessarily have to be technical), along with a short description of what you did.
Question 3 (100 words):
Review this existing eval for examining nonhuman biases in AI models. The write-up is here for more info. What’s one limitation or gap in how this benchmark handles questions about animals or non-human entities? Propose a specific test scenario or methodology improvement that would address this limitation. Some guiding questions: what would the prompt look like? What would you measure? How would you know if it's working?
About the mentor

Allen Lu is a creative engineer operating at the intersection of AI safety and animal welfare. He holds a B.S. in Integrated Design & Media from NYU and has extensive experience building full-stack applications with React, Next.js, and integrating AI/ML technologies into social impact projects.
Allen is deeply engaged in the effective altruism and AI safety communities, with particular expertise at the intersection of AI and animal welfare. He's actively involved with organizations like Sentient Futures, Electric Sheep, Bluedot, and Open Paws, and currently facilitates fellowships for the Futurekind AI Fellowship and AIxAnimals. Allen is committed to expanding the circle of moral consideration in AI systems, and presently works on maintaining AnimalHarmBench 2.0, a standardized LLM benchmark designed to measure multi-dimensional moral reasoning towards animals. His current project seeks to advance and expand that work, building on that foundation while exploring new evaluation approaches.