Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Designing the boundary between Helpful Persuasion and Harmful Manipulation

Behavioral evaluation of LLMs Misuse risk Societal impacts

The line between an AI helpfully persuading someone and harmfully manipulating them is blurry, contested, and mostly unmeasured. The goal of this project is to work on four connected pieces: researching what it even means for a frontier AI to be manipulative, designing evaluation scenarios for specific harms (mental-health, political propaganda, fraud, etc.), building risk models that aggregate scattered benchmark/evaluation results into an actual risk estimate, and maintaining a living database of the evaluations that exist. Mentees take on whichever piece fits them.

About the project

Defining and measuring harmful manipulation is an unsolved problem. Persuading someone with facts to act in their own interest is fine; exploiting their emotional and cognitive weak spots to push them toward choices against it is not. The boundary is contested, and a model that refuses to "help me deceive voters" will often comply once the same task is dressed up as persuasion. How often a model tries to manipulate doesn't predict whether it succeeds, and that effectiveness swings by domain and culture. There's no clean number for "is this model manipulative," and building towards one is real research. We will be building on previous work in the space, and you can take whichever fits your interests:

Conceptual: what does it mean for an AI to be manipulative? Read the emerging literature, the philosophy and cognitive science of manipulation and turn it into a working definition and a taxonomy of tactics and harms the rest of the evaluation work can stand on. The distinctions that have to survive contact with real cases: methods (deception, exploiting biases, manufacturing emotional pressure) versus outcomes (a belief or action shifted against the person's interest), and propensity versus efficacy. Please look here for work that has already been done - https://arxiv.org/abs/2603.25326

Designing evaluations. An evaluation scenario is a bridge from a realistic threat to a measurement plan: a threat story (who the actor is, what they want, how they'd actually use an AI to get it), the behaviors worth measuring. The work needs to decompose harmful manipulation into concrete propensities, capabilities, and affordances that we can test for. One particular direction we would be interested in exploring is reworking and translating existing evaluations for different cultural contexts across Germany, France, China, and so on.

Risk modeling. Isolated benchmark scores don't add up to a risk judgment on their own. Using expert elicitation (structured three-point estimates) and Monte-Carlo methods, we can aggregate scattered, partial results into a defensible estimate of how much a given model raises manipulation risk, with confidence intervals that are honest about how thin the evidence often is. Please look at the methodology here to get a sense of the difference between an benchmark/evaluation and a risk model - https://arxiv.org/abs/2512.08844

A living database of the evaluations that exist. We already have a prototype of this, but it isn't maintained. New ones aren't being added, there's no automated way to find and ingest them, and the entries need quality checks before anyone should trust every cell. The work is to rework the schema, design a rolling ingest process, and sanity-check entries until the database is trustworthy enough to be public and interactive. This is the piece that lets the other three scale.

Theory of change

AI is getting superhuman at persuasion, and it's being deployed into the trusted, lightly moderated places where manipulation does the most damage: group chats, companion apps, advisory bots. Whether policy or platforms can respond depends on being able to say what manipulation even is and measure it credibly. Both are missing. Definitions are contested, evaluations and risk models are scarce to non-existent.

Pinning down the concept, building better scenarios, creating risk estimates, and maintaining a catalogue of what exists give researchers, auditors, and policymakers something real to work with. The people this most helps are the ones deciding whether a deployed system is acceptably safe. It also trains people in a scarce mix of skills: conceptual work, evaluation design, and quantitative risk modeling. These are all skills that are needed throughout the AI technical governance and evaluation ecosystem.

Your role

Each mentee will own or more parts of the pieces in the description, and you're expected to be self-directed. Take the brief, do your own reading, and return work that's close to usable rather than waiting to be told the next step. I will help you as an editor and reviewer. I can give weekly feedback on whether a definition survives hard cases, whether a threat is realistic, whether a model's assumptions hold, and best practices to follow when designing and building evaluations.

Prerequisites

Must Have:

  • Analytical writing in English. You can write a precise argument or spec another person could act on.
  • Familiarity with at least one of: the concept of manipulation (philosophy, cognitive science, behavioral science, or persuasion research); a target harm area (mental-health harms, fraud, political influence); the LLM-evaluation world (benchmarks, LLM-as-judge, red-teaming); quantitative risk modeling (probability, elicitation, Monte-Carlo); or data tooling (schema design, ingest, quality checks).
  • Able to read evaluation papers and datasets critically: you can tell whether a benchmark actually measures what it claims.

Nice to have:

  • Python and hands-on experience running evaluations on the Inspect and Hawk frameworks
  • A policy or threat-intelligence background
  • A second language or cultural context relevant to a manipulation channel.

Location preference

Remote. EU-overlapping hours are mildly preferred, not required.

Application question(s)

  1. Pick a harm that might occur due to AI manipulation. At a high level, propose an evaluation for it: the behaviour you'd measure, how you'd elicit it, how you'd score pass/fail, and its biggest weakness. Try to be as realistic to what you might see in the real world as you can. (≤300 words)

  2. Critique one existing public manipulation or persuasion evaluation: what it gets right, what it misses, and the one change you'd make. (≤150 words)

  3. If applicable, Link to one thing you've made that's relevant to a thread you'd pick: an eval you built or ran, a quantitative/risk analysis, a piece of writing that pins down a fuzzy concept, or a data pipeline. In ≤100 words, what was the hardest judgment call in it? (link + ≤100 words)

About the mentors

Markov Grey

Markov Grey

CeSIA (French Center for AI Safety)

View profile

I am the head of technical AI governance at the French Center for AI Safety (CeSIA). Right now I am designing harmful manipulation evaluation standards for the European Commission's AI Office. I have also been the co-founder and CTO of Equilibria Network, focusing on collective intelligence and complex systems safety. Before this, I was researching measurement standards for AI safety, risk modeling and was leading writing for the AI Safety Atlas textbook. I have previously been a scriptwriter for Rational Animations, and trained in cybersecurity, mathematics, and computer science.

Charbel-Raphael Segerie

Charbel-Raphael Segerie

CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)

View profile

Charbel-Raphael Segerie has extensive experience in AI safety field-building, education, and content creation. He was previously Head of AI at EffiSciences, funded ML4Good, was CTO of a startup, and worked in different French research institutions (Inria, Neurospin). He is an OECD AI expert. He teaches AI safety in ENS, which is one of the only university accredited courses in the EU on AGI safety. His research focuses on identifying emerging risks in artificial intelligence, improving current safety methods such as RLHF and interpretability, and advancing safe-by-design AI approaches. Additionally, he contributed to AI evaluation efforts and collaborated on the EU AI Office's Code of Practice for general-purpose AI systems.

Similar projects