Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Value Systemization and Reflection in AIs

Alignment Philosophy of AI

Future AIs might reflect on their own values and decide to "systematize" them into more simple forms, which could be misaligned relative to what humans want.

Theory of change

Value systematization is a concrete pathway for AIs to develop unwanted goals. We should figure out what this could look like, how we could measure it, and some concrete ideas on preventing it.

See more in project doc: https://docs.google.com/document/d/1xI-73F-wFChNWMMvBe6lDrLSEc9LzQ7QwhtY9x6LCLU/edit?usp=sharing

Prerequisites

You should be someone who:

  • Enjoys thinking about topics that are not well defined/ enjoys thinking about how to ask the right questions.
  • Have a good understanding of why alignment might be hard (you don’t need to agree with the orthodox arguments, but you should be able to explain them).
  • Have a good understanding of how LLMs work
    • A good test for your own understanding is: can you explain the motivation and key results behind Alignment Faking in Large Language Models
  • Have basic coding skills. You should be comfortable using packages like inspect AI to query models.
  • Be able to communicate complicated ideas through writing.

See more in https://docs.google.com/document/d/1xI-73F-wFChNWMMvBe6lDrLSEc9LzQ7QwhtY9x6LCLU/edit?usp=sharing

Location preference

No preference

Application question(s)

Pick one of the concrete research questions I proposed above, or come up with your own research question. Try to make some basic progress on the question, possibly by running empirical experiments (i.e., asking models) or just sit down and think about it. Write no more than 400 words (strict limit) on what you’ve found/thought of.

Is there anything that makes you fit/qualified for this project that you feel like isn’t obvious from the rest of your application? (Optional. No more than three sentences).

See more in: https://docs.google.com/document/d/1xI-73F-wFChNWMMvBe6lDrLSEc9LzQ7QwhtY9x6LCLU/edit?usp=sharing

About the mentor

Tim Hua

Tim Hua

MATS

View profile

Tim is a current MATS 8.1 extension scholar under Neel Nanda and an incoming Astra fellow in the Redwood strategy stream. Tim's past AI safety work spans topics such as mitigating evaluation awareness using activation steering (cited in the Claude Opus 4.5 system card), AI-induced psychosis, and combining different monitors for AI control. Tim is interested in conceptual and empirical alignment research, as well as the economic and political risks from advanced AI (e.g., gradual disempowerment.)

In a past life, Tim worked as an economist at Walmart and graduated from Middlebury College in 2023. Tim's senior thesis "Fox News's Effects on Social and Moral Preferences," won the D.K. Smith Prize for best senior thesis.

Similar projects