Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Develop epistemic evals with Sophron Research

Behavioral evaluation of LLMs Evaluations Societal impacts

Developing evaluations for assessing the epistemic properties of AI models, including how they affect our ability to have true beliefs.

About the project

This is an open-ended project for helping Sophron Research develop evaluations for testing AI models' epistemic character. Sophron's mission is to improve the ability of humans to have true beliefs in a world with advanced AI, especially in high-stakes decision-making.

The mentee would either help us develop eval projects that are currently present in some form, or propose and develop an evaluation project of their own. Examples of projects within our purview include:

  • Sycophancy evaluations (e.g. expanding our Pander Score with multi-turn interactions)
  • Monitor model bias in favor of developer
  • Assess model ability to represent uncertainty
  • Test congruence between model reports and other representations
  • Identify or operationalize important properties that model ideally should demonstrate

Mentees will be working with Sophron founders Paul de Font-Reaulx and Alejandro Botas.

Theory of change

We believe that humanity will need excellent epistemics to successfully navigate a future of fast AI progress. For example, we want our public policies to be grounded on accurate assessments of the world, even if they need to be formed fast, and our safety research to be as effective as possible. Our strategy to facilitate this is to find ways of leveraging AI itself for our own ability to reliably form true beliefs. By developing high-quality evaluations of frontier models and publicizing the results, we hope to be make clear whether they can help us form better beliefs, and create stronger incentives for developers to provide models that do.

Your role

This can vary depending on the mentees experience and preference. We will be available for weekly meetings, and to touch base on an ad hoc basis. We are based in New York City and happy to meet in person here, but mentees can also work remotely without issue. We expect mentees to be capable of largely autonomous work.

Prerequisites

  • Some familiarity with evaluation and benchmarking.
  • Background in CS, ML, Stats, Econ, and/or Philosophy, or equivalent. Familiarity with AI and AI safety required.
  • Substantial research experience (e.g. writing research papers). Preferably: PhD completed or ongoing.

Location preference

US times preferred, but UK/Europe is fine too if mentees are available later in the day

Application question(s)

  • What are some ways in which you think that our Pander Score could be improved?

  • What are some epistemic evaluations that you think would be valuable to develop? Why?

Please spend 1h max answering these questions combined (though feel free to think about them passively for a while before you answer).

About the mentors

Paul de Font-Reaulx

Paul de Font-Reaulx

Sophron Research

View profile

Paul is a philosopher by training who recently completed his PhD and co-founded Sophron Research. His background is primarily in cognitive science and decision theory. He is primarily interested in ensuring that the coming few years of AI development go as well as possible for humans and other sentient beings.

Alejandro Botas

Alejandro Botas

Sophron Research, Future of Life Foundation

View profile

Alejandro co-founded Sophron Research, where he builds public benchmarks of the epistemic properties of frontier AI models. His background is in quantitative engineering and research, spanning big tech, startups, trading, and academia. He is primarily interested in steering AI development to support societal epistemic health and wise human decision-making in the critical opening years of the AI era.

Similar projects