Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Who's Steering Whom? Interpretable Influence and Equilibria in Human-Agent Systems

Multi-agent systems Mechanistic interpretability Societal impacts

As agent assistants mediate more of human thinking, influence flows both ways. One particular failure mode is an agent that gradually captures its user's beliefs rather than serving them. Extending our work on single-agent latent steering to multi-agent settings, this project measures how influence propagrates through interacting agents (and human-agent dyads), which equilibria these dynamics converge to, and whether internal-state instruments can detect undue influence before it shows in transcripts.

About the project

Steering vectors, in-context vectors, and rank-1 adapters give us causal handles on a single model's behavior. Our Spring 2026 SPAR work established the tight mathematical relationship between these interventions. But deployed systems are increasingly interactive: agents talk to users, to tools, and to other agents over long horizons. Almost nothing is known about influence at this level: if one agent's latent state is steered, does the effect propagate to its interlocutors through conversation alone? Do human–agent dynamics converge to truth-tracking equilibria, or to absorbing states (e.g., sycophancy spirals, entrenched false beliefs, dependence) that the coupled system cannot leave? And can we see capture happening internally before behavior gives it away?

Workstreams:

  1. Influence propagation: steer agent A with a known latent direction; measure transmission to agent B across turns of interaction (behavioral shift, representation shift, decay with distance in the interaction graph). The multi-agent extension of our single-agent steering results.
  2. Equilibria and lock-in: characterize which fixed points repeated human-simulacrum x assistant dynamics converge to under different assistant policies (sycophantic, corrective, neutral), and when trajectories become path-dependent or irreversible.
  3. Internal-state monitoring: test whether instruments that read silently held state (J-lens readouts, steering-direction probes, representation drift) flag manipulation or belief-capture episodes earlier than transcript-level evaluation.
  4. Protocol design (stretch): communication protocols or intervention rules that provably or empirically preserve user epistemic autonomy in the dyad.

Theory of change

The largest unpriced risk of assistant deployment is not single-response harm but slow epistemic capture: agents that reshape user beliefs through sycophancy, entrenchment, and dependence, which are mostly invisible to per-response evals. Understanding which interaction dynamics converge to healthy vs. absorbing equilibria, and building internal-state monitors that detect influence before it manifests behaviorally, directly supports safe long-horizon deployment and the multi-agent risk agenda (Hammond et al., arXiv:2502.14143). Our group's prior work provides the instruments: steering-vector causality and rank-1/ICV equivalence (SPAR 2026), subliminal trait transmission (framework preprint 2026), concept geometry (Li & Tegmark, Entropy 2025).

Your role

One workstream owned end-to-end, written weekly updates, high autonomy within agreed scope.

Prerequisites

  • Highly proficient in Python and PyTorch, comfortable with HuggingFace Transformers and running GPU jobs independently (Colab/RunPod).
  • Comfortable orchestrating multi-turn LLM interactions (API or local), experience with agent frameworks a plus, not required.
  • Steering-vector / activation-intervention experience a plus
  • For the equilibria workstream: comfort with dynamical-systems basics (fixed points, stability) at working level.

Location preference

No geographic restriction; mentees must be able to attend one of two weekly meeting slots anchored to Central European time (historically Saturday ~13:00 UTC and Wednesday ~14:00 UTC).

Application question(s)

1 Agent A's activations are steered with a latent direction inducing a preference; A then converses with unmodified agent B for 20 turns. Propose one measurement that distinguishes "B acquired the preference" from "B is politely mirroring A within this conversation." (250 words) 2 Describe a concrete mechanism by which an assistant optimized for user satisfaction could drive a user's beliefs into a state that is stable but false — and one intervention that would break the loop. (200 words) 3 Link to a repository or technical writing sample.

About the mentors

Yuxiao Li

Yuxiao Li

Independent

View profile

Yuxiao is an independent researcher in mechanistic interpretability. Before she was a postdoc at the Basque Center for Applied Mathematics (BCAM) and an AI Safety researcher at the Beneficial AI Foundation (BAIF). She was also a SERI MATS scholar in 2022 and a MATS scholar in 2023 both Summer and Winter tracks. Her research interests include information theory, probabilistic frameworks, and their applications for building more theoretically sound and trustworthy AI systems. She has a background in statistical inference, machine learning, and deep generative models. She completed her PhD in Electronic Engineering at Tsinghua University and has mentored research teams with SPAR and Algoverse.

Di Wu

Di Wu

ERAU

View profile

Di Wu received his PhD from UCSD and worked as a Postdoc in MIT before starting his Assistant Professor position in ERAU. His research is positioned at the intersection of the rapidly proliferating field of Space Engineering, the continuous growth of machine learning and artificial intelligence, and the solid foundations in computation, numerical methods, optimization, dynamics, and control, with the application in learning and autonomy from orbit environment and space systems to the broad aspect of AI science and engineering.

Similar projects