Produce a process-level evaluation of cooperation in AI systems inspired by research in cognitive science on joint commitment.
About the project
People write their code with Claude. They vent to ChatGPT. Some decision makers trust all that LLMs say. For AI to go well and for us to avoid hallucinated test suites, delusional spirals, and existential superpersuasion, we need to make sure that the AI agents with which we interact do what we want--that they cooperate with us. Thankfully, scholars have begun to test whether AI systems choose cooperative outcomes in well formalized domains. Still, there is more to do.
Cooperation should be both a process and an outcome measure, but present studies focus chiefly on the latter. While necessary, measuring whether an agent chooses the cooperative option of, for example, hunting hare in stag hunt or getting out of the way when playing Overcooked is not sufficient to tell us if the agent is cooperating in a human-like way and whether such ``cooperation'' will generalize to any open ended interaction with a human.
A few scholars have recognized the need for process level measures often using LLM-as-a-judge on agent behaviors in game theoretic settings. We extend this work by contributing (1) process level measures in naturalistic interaction tasks applicable outside of game theory, (2) carefully designing stimuli to test theory-driven components of cooperation and (3) testing the generalization of our process measures.
Drawing on comparative and developmental psychology, we introduce a benchmark which measures basic capacities such as making (often implicit) commitments, delivering and responding to social sanctions, and navigating identity-relative (conventional) and universal (moral) norms. Together these capacities compose what we intuitively call cooperation.
In a variety of settings, we find evidence for a generalization failure of outcome level cooperation: that agents appear to cooperate on the final outcome even when they fail to perform the component parts of cooperation and, in fact, in slightly varied settings do not choose cooperative outcomes.
Furthermore, we find that models trained to exhibit our components perform better not just on processes-level cooperation but also significantly improve their outcome success on external benchmarks of cooperation and alignment.
Theory of change
We are at a path-dependent juncture in AI. The research directions we take now, if chosen wisely, will help develop technical and conceptual countermeasures against AI takeover and other risks; they will reduce the "alignment tax." While there has been an increasing amount of attention paid to AI safety, proportionally few are working technically on value alignment despite the fact that to align AI systems we may need to make them understand values and morality like humans do. I believe that work in value alignment, both because it is a relatively under-explored direction with room for more formalization and because it clearly relates to making AI do what we want, will most effectively reduce the alignment tax. In particular, robust formal accounts of cooperation are necessary to diagnose and mitigate models’ deceptive capabilities. Having early warning signs of deceptive capabilities may actually help focus the attention of society in a way to best mitigate issues.
Your role
Depending on time availability and level of experience you mentees will either lead or co-lead the project. This will involve:
- Designing evaluation protocol; reading through a variety of cognitive science papers and distilling them into LLM-appropriate settings.
- Running and iterating on experiments.
- Querying models.
- Co-writing the paper.
Prerequisites
Prior research experience (can be in another domain).
Able to read and understand papers in at least some AI domains and in cognitive science (and not just use an LLM summary).
Proficiency with python. Writes clean, understandable code.
- It is fine to use coding agents but you must know how to reign them in (viz.: read their code; don't let them reimplement everything; use linters and formatters).
Have a realistic model of your own schedule and get things done when you say. (I'm not trying to say everything has to be done quickly -- just that you should communicate your uncertainty.)
Openness to learn new things.
Clear communicator. Agreeable and humble.
Application question(s)
In a few sentences each...
-
What does a good research collaboration look like to you?
-
What's a book or piece of media that changed or inspired you? Why?
-
Tell me about a project you have worked on. (Ideally a research project.) What role did you play? What went well? What didn't?
About the mentor
