An agent-based model of how populations gradually lose agency to AI-mediated manipulation and delegation — formalizing tipping points, spread dynamics, and the divergence between AI-empowered and atrophied users, aimed at a publishable computational social science paper.
About the project
Most AI risk discussion focuses on abrupt catastrophe. This project studies a quieter failure mode: the gradual loss of human agency — people incrementally giving away their thinking to other entities, whether a human elite or advanced AI itself, each step individually rational and none ever feeling like a loss. The mechanism is self-reinforcing: overreliance on AI erodes confidence and skill, deepening reliance, while populations diverge between those AI empowers and those whose critical thinking atrophies. The core question: under what combination of individual attributes and social reinforcement does AI-mediated manipulation tip a population into self-reinforcing agency loss — and what determines the tipping point, speed, and magnitude of the spread?
We will build an agent-based model in the Schelling (1971) tradition: minimal behavioral rules, aimed at robust qualitative results rather than calibrated forecasts. Heterogeneous agents — differing in education, age, and other susceptibility-shaping attributes, following threshold models (Granovetter 1978) — repeatedly choose to rely on AI-mediated information and delegation. Three mechanisms interact: an individual atrophy ratchet, threshold-based susceptibility to manipulation, and nonlinear social reinforcement. The model does not assume AGI: capability enters as tunable parameters (delegation quality, persuasion strength, personalization), making "how capable must AI be before tipping points appear?" an experimental result, not an assumption.
This fills a specific gap: Kulveit et al.'s Gradual Disempowerment (2025) makes the case conceptually, and RAND's formal model (Moon & Boudreaux 2026, rand.org/t/RRA4817-1) captures the structure of agency loss via decisive coalitions while explicitly flagging degraded participation and dynamics as future work — exactly what this project supplies.
Concretely: formal micro-model specification, mean-field baseline results, Python implementation (Mesa/networkx, open repository), and parameter sweeps mapping tipping surfaces across attribute distributions, network topologies, and the capability dial. Outputs: a publishable computational social science paper (e.g., JASSS), an interactive demo, and a short policy translation on early-warning indicators.
Theory of change
Gradual disempowerment is increasingly recognized as an existential risk pathway in its own right (Kulveit et al. 2025): humans losing the capacity to shape collective outcomes not through a takeover, but through accumulated small surrenders of thinking and choosing. Its danger is precisely that no single step looks catastrophic — so policymakers currently have no way to detect it, measure it, or know when to intervene. This project makes the risk legible and actionable: a formal model with explicit tipping thresholds converts a vague civilizational worry into monitorable quantities — which populations are most susceptible, how fast agency loss spreads, what early-warning indicators precede irreversibility, and which interventions keep societies below the threshold. It directly advances RAND's recommendation (Moon & Boudreaux 2026) that the field develop agency evaluations and reversibility benchmarks. Because AI capability is a tunable parameter, the results also inform a near-term question: whether current persuasive and companion AI systems already warrant intervention under manipulation provisions like the EU AI Act's, before transformative systems arrive.
Your role
Mentees will be full intellectual contributors and co-authors: each owns a workstream (formal model, computational implementation, or theory extension), has an equal voice in our weekly discussions, and will substantially shape the model's design and the project's direction.
Prerequisites
I have different requirements for all 3 mentees:
Role 1 — Formal Modeler (required):
Graduate-level coursework or equivalent in game theory, microeconomic theory, or formal political theory. Comfortable specifying and analyzing a model with pencil and paper (utility functions, best responses, fixed points); experience with threshold or dynamic models is a plus. Able to write mathematical results clearly for an academic paper.
Role 2 — Computational Mentee (required):
Highly proficient in Python (NumPy/pandas); able to build and maintain a clean, reproducible GitHub repository. Bonus—Built a simulation, agent-based model, or comparable computational project (coursework or toy projects following guides are fine); familiarity with Mesa or networkx is a plus. Comfortable designing and visualizing parameter-sweep experiments.
Role 3 — Theory Mentee (optional; project may proceed without this role):
Background in social choice theory or formal political theory (e.g., has read and can work with Arrow-style frameworks). Interested in potentially extending RAND's coalition-based model of collective agency (Moon & Boudreaux 2026) with degrees of effective decisiveness.
All roles: genuine interest in gradual disempowerment as a risk area, ability to commit 5-10 hours/week for 12 weeks, and willingness to work as part of a small team where model and code are developed jointly.
Application question(s)
For mentees: Please specify what role you would like to apply for specifically within this project and then write your answers correspondingly.
All applicants (~150 words) - What attracted you to this project? Do you agree that gradual disempowerment is one of the major risks of advanced AI? Why or why not?
Formal Modeler (~250 words) - Tell me about your past experience that most closely resembles the work you are about to do as a part of this team. How do your skills match the current project's needs?
Computational Mentee (~250 words) - Are you familiar with agent-based modeling or complex adaptive systems? Do you have any prior experience with simulation methods?
Theory Mentee (250 words): Skim RAND's recent paper (Moon & Boudreaux 2026) and provide criticism to their model specification from a theoretical standpoint. Feel free to use the LLM to summarize the report, but provide your own original ideas to criticize the paper.
About the mentor

Zhamilia (goes by Jama) has seven years of experience at the intersection of data analytics, international policy, and strategy, helping global clients translate complex challenges into clear, actionable solutions. Her work bridges data science, software, and international strategy with a regional focus on Central Asia, the MENA region, China, and Southeast Asia. She has contributed to early-warning forecasting technology, decision-making under uncertainty, and horizon-scanning methodologies, earning multiple scholarships, awards, and co-inventor credits on several patents at Acertas.
She brings extensive experience in agent-based modeling, stakeholder and scenario analysis, and has advised public- and private-sector clients across MENA, India, Indonesia, the ROK, the UK, the US, and Latin America. Her research interests include AI governance, emerging technologies and global order, and AGI-related strategic competition. Jama serves as a Community Reviewer for Frontiers in Political Science, works as a Business Intelligence Analyst at Acertas, and is a Fellow at the TransResearch Consortium. She holds a Doctorate in International Politics and a Master’s in Applied Data Science from Claremont Graduate University, and a Bachelor’s in Business Administration from the American University of Central Asia.