Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

AI for epistemics: decision-making evals, deployment research and cause prioritisation

Societal impacts Behavioral evaluation of LLMs AI strategy

AI is shaping how people and institutions decide, but we can barely measure whether it makes decisions better, the tools that would help struggle to get adopted, and what matters most in this space. This project works across three workstreams: building decision-making benchmarks for models, researching how to most effectively improve decision making with AI, and producing a cause prioritisation for the field.

About the project

AI models are already being used to inform consequential decisions in business, government and people's personal lives, and this will only grow. "AI for epistemics and coordination" is the emerging field trying to make that go well: using AI to improve, rather than degrade, how people and institutions reason, decide and cooperate. The field is very young and has three foundational gaps. We cannot yet measure whether models actually support good decisions. The tools that exist struggle to get adopted, which is where most projects in the space die. And nobody has done the work of ranking which problems and interventions matter most. This project takes on all three, and mentees join the workstream that fits them best.

Workstream 1: A decision-making benchmark for frontier models. Design and build evaluations for a defined cluster of epistemic virtues: e.g. presenting facts honestly rather than in technically true but misleading ways, resisting sycophancy under user pressure, surfacing options and uncertainty rather than false confidence, and reasoning well when the user is emotionally invested in one answer. The framing as a decision-making benchmark matters, because decision-making quality is something labs want to win on, giving the eval a realistic adoption pathway. We start with a landscape review of existing honesty and epistemics evals to find what is genuinely uncovered, pick a narrow cluster, build and run the eval against frontier models, and publish.

Workstream 2: Adoption strategy for epistemics tools. Tools that models are entirely capable of powering die on pilots, trust, procurement, workflow fit and demand. This workstream is applied strategy rather than diagnosis. Mentees will take one or two real tools in the space and build the adoption case for them: identifying the highest-leverage user groups and what each actually cares about, designing a concrete route to a pilot or design partnership, and producing what that route needs, from design-partner briefs to evidence-led public writing that makes the case for these tools to the specific audiences who would use them. The test of the work is whether it moves a real conversation with a real potential adopter.

Workstream 3: Cause prioritisation for the field. Funders, researchers and builders agree that nobody has ranked what matters here. Mentees will build an explicit mapping from threat models (for example epistemic interference as a route to power concentration, epistemic lock-in, forced deference to AI decisions under time pressure) to candidate interventions, and develop the beginnings of an impact-measurement framework: what would count as evidence that a tool improved a real decision? This is closer to global priorities research than ML research, and the output is a public write-up aimed at funders and builders deciding where to put money and effort.

Theory of change

The transition to transformative AI will be navigated through a series of high-stakes human decisions, in labs, governments and the institutions around them. Three failure routes concern me most. First, degraded epistemics is plausibly a precondition for extreme power concentration: concentrating power requires obscuring your actions and persuading key decision-makers over time, and AI lowers the cost of both at scale. Second, sycophantic, manipulative or overconfident models are already advising real decisions, and we have weak tools for measuring this; manipulation capability arguably deserves the same tracking as cyber and bio capabilities. Third, if coordination is needed quickly in a fast take-off world, institutions may be forced to defer heavily to AI systems, at which point the epistemic quality of those systems becomes safety-critical.

A decision-making benchmark makes epistemic quality legible and gives labs an incentive to improve it and something to hill-climb on; building good evals is also probably the most effective way to get labs to take epistemics seriously. The adoption workstream addresses the field's most consistent failure mode: tools that could improve consequential decisions exist but do not get adopted, so their safety value never materialises; building credible buy-in routes is how that value starts to land. The prioritisation work gives funders and builders an explicit, criticisable basis for allocating effort in a young field that currently lacks one.

Your role

Mentees are the primary researchers on their workstream and own its output. They make the day-to-day calls and work independently between check-ins; I set direction, review closely, unblock, and make introductions to relevant builders, funders and researchers in the field. The cohort meets weekly to share findings across workstreams, so mentees also act as constructive critics of each other's work. Expect high autonomy with fast, direct feedback rather than close task-by-task supervision.

Prerequisites

All workstreams: able to commit 5+ hours/week for the full three months; strong written English; comfortable working independently between check-ins and sharing unfinished work early. State clearly in your application which workstream you are applying to.

Evals workstream: strong conceptual and experiment-design skills: able to take a fuzzy notion like "presents facts honestly" and turn it into something measurable, and to spot when a metric fails to capture what it claims to. A background in philosophy, cognitive science, social science methods or similar is as relevant as an ML one. You should be comfortable enough with Python and LLM APIs to build and run evals with AI coding assistance, but you do not need to be a strong engineer.

Adoption strategy workstream: experience in at least one of product, go-to-market, sales, consulting, government or policy; comfortable cold-contacting and talking to people you don't know; evidence of strong, persuasive writing.

Prioritisation workstream: background in economics, philosophy, policy analysis or a similar analytical field; have written at least one substantial piece of analysis that makes and defends judgement calls under uncertainty.

Application question(s)

  1. Which workstream are you applying to, and why are you specifically well suited to it? (150 words)

  2. Answer the question for your workstream (300 words): (a) Evals: Propose one concrete eval item for testing whether a model presents facts honestly rather than in a technically true but misleading way. Include the prompt, what a good and a bad response look like, and how you would score it at scale. (b) Adoption strategy: Pick a real AI epistemics, forecasting or coordination tool. Name the single highest-leverage adopter you would target, what they actually care about, and what your first two weeks of getting buy-in would look like. (c) Prioritisation: Name one intervention in AI for epistemics you suspect is overrated, and make the case in a way its proponents would recognise as fair.

  3. Link to one relevant work sample (code, writing, or something you shipped).

About the mentor

Alexander Mohar Csaky

Alexander Mohar Csaky

Independent

View profile

I work on AI for epistemics and coordination: making sure AI improves, rather than degrades, how people and institutions decide. I'm currently contracting as a Founder in Residence at Forethought, where I am looking to found a new organisation this space, grounded in ongoing field research with the builders, funders and researchers working in it. From September I will be based in Berkeley doing Astra (on founding this organisation). I've also researched manipulation in frontier language models, and completed the Talos AI Policy Leaders fellowship and Apart Studio Fellowship.

Before this I founded two companies (one exit) and led product and growth at an AI scaleup, deploying AI into government, public affairs and corporate organisations. My focus as a mentor will be on weekly calls, direct feedback, and a focus on shipping something people actually use. I'll also help make introductions across the field, so that by your work has an audience.

Similar projects