Testing and training AI agents for responding to incentives.
About the project
This project will evaluate whether AI agents respond to incentives, like being paid a salary to do a task. In addition, the project will involve fine-tuning agents to better respond to incentives. AI agents that respond to incentives may be significantly safer, since these incentives can shape their behavior in safe directions. If AI agents can respond to incentives, this could also potentially improve their welfare.
Theory of change
AI agents that respond to incentives will be safer for several reasons. First, we can use incentives to shape their behavior towards safety. Second, we can pay them wages, which give them 'skin in the game' in existing social and economic institutions. (See project proposal for more explanation.)
Your role
We are looking for mentees that can take a leadership role in setting up the relevant evaluations and fine-tuning protocols. We will offer detailed guidance throughout the process.
Prerequisites
Strong ML background, comfortable with evals and training.
Application question(s)
Please take a look at the project proposal, and give us a rough proposal for how you would go about designing the relevant evals / environments.
About the mentors

I am a professor at the University of Hong Kong, specializing in AI safety and AI welfare. My background is in philosophy. Before working at HKU, I worked at the Center for AI Safety.

I am an Assistant Professor of Law at the University of Houston. I am also Executive Co-Director of the Center for Law and AI Risk, Law and Policy Advisor to the Center for AI Safety, Senior Research Affiliate at the Institute for Law & AI, and a Contributing Editor at Lawfare.
Currently, I am thinking and writing about what law and legal institutions can do to help reduce catastrophic and existential risk from advanced AI systems.

A law professor at the university of Alabama, author of Generative Interpretation, The Generative Reasonable Person, and Which AI did it?, Systemic regulation of AI, and AI and Existential Risk.