SaferAI is developing quantitative risk models of AI-assisted cyber misuse. Mapping LLM performance on AI benchmarks to steps in these models involves eliciting human experts' estimates and is resource intensive. Instead, we wish to simulate this elicitation procedure with virtual LLM 'experts'. In this project, you will assist with validating this procedure and calibrating our LLM-derived predictions.
About the project
(for full information, please see https://docs.google.com/document/d/14nqHzTQErZiPcbb1dNoyIfUr_tlfwMJjNnFjEkCD16I/edit?usp=sharing)
In recent work, SaferAI has developed 9 detailed cybersecurity risk models. These risk models use AI evaluation benchmarks as ‘key risk indicators’. That is, if an LLM scores X% on some benchmark, what is the corresponding probability Y that it can successfully complete a given step in our risk model? To find this mapping, we have previously contracted cybersecurity professionals and gathered their estimates in a process known as ‘expert elicitation’. However, this process is slow and resource intensive. Thus, we wish to elicit the values in our risk models with simulated LLM experts instead. This is based on the promising results of several prior works (Halawi et al, 2024, Phan et al., 2024, Barrett et al., 2025).
However, LLM forecasting has also received substantial criticism. Therefore, we would like to improve and validate our LLM elicitation procedure as much as possible. Some questions we can explore:
- How does the depth and breadth of information provided to the virtual LLM experts affect prediction quality? What sort of data should we include in the prompts without the LLM getting lost?
- What are some spurious factors that might influence prediction quality? For example, past work has found that changing the bracket type from [A] to (A) in multiple-choice questions can change evaluation scores by 5%.
- How do we elicit enough diversity of opinion to simulate human experts? LLMs are known to repeat the same ‘talking points’ and to converge onto similar opinions.
- How do we aggregate these results? Are there alternatives to taking the mean/median of the individual LLM experts’ results?
- How do we validate the LLM elicitation procedure against ground truth data? Previously, we have conducted several experiments testing the quality LLM predictions. However, more testing in realistic deployment scenarios is needed. Ideally, these tests should involve ground-truth data, so that the LLM elicitation procedure can be calibrated against real-world input.
- Which models are best at forecasting?
Ideal project outcome: we have devised a method to iteratively improve and calibrate the LLM forecasting procedure for our risk models.
Theory of change
SaferAI’s Theory of Change revolves around creating quantitative and grounded risk models (as well as risk management practices and standards), so that we enable:
- AI labs to apply targeted safeguards
- policymakers to prioritize interventions
- researchers to design more actionable evaluations
- regulators to quantify the level of risk and enforce compliance with existing legislation
Taken together, these four factors will lower the negative impacts of advanced AI systems on our society. By improving our LLM forecasting procedure, you will contribute to these goals by allowing us to derive quantitative estimates of risk more quickly and precisely.
Your role
- Mentee will be the lead researcher for the question that we are investigating, with a tight feedback loop between mentee and mentor for fast research iterations
- Mentee should have sufficient autonomy to make meaningful progress on a week to week basis.
- Mentor will be available for a weekly call and frequent communication through Slack/email
- Mentor will support the mentee with experiment design, analysis and write up, but the mentee should take ownership of these tasks
- An ideal output will involve a workshop paper or a blogpost on SaferAI's website
Prerequisites
(We are open to a wide range of backgrounds, including those considered ‘unconventional’ in AI safety. If in doubt, please apply!)
For this project, we are mainly looking for the following skills (not all are required, though):
- experience with evaluating AI models
- good knowledge of AI benchmarks
- forecasting experience
- familiarity with the most common LLM APIs (Anthropic, OpenAI, Gemini, …, maybe OpenRouter), so that you can quickly run experiments validating the LLM forecasters.
Nice to have:
- conducting experiments in social sciences, psychology, or any other environment with lots of uncertainty and confounding variables. Not required.
Time commitment
10
Location preference
Any time zone compatible with UK time is fine.
Application question(s)
-
Please attach a sample of prior work that demonstrates the skills listed above. This could include papers, blogposts, sample code or even engagement in online discussions. Alternatively, please describe why you think you are a good fit for this project.
-
(<250 words) Imagine you are using an LLM as a regressor to estimate a particular quantity out of the world. We treat both the system prompt and the data as inputs to the model. Describe an experiment to quantify how changes in a system prompt lead to changes in the regression accuracy. Please only provide a broad overview of the steps you would take to understand this.
-
[OPTIONAL] (<250 words) Now imagine that you do not have access to the regression accuracy. How does this change the problem and your research agenda? Is it still possible to systematically investigate how changes in the prompt lead to changes in the model output?
About the mentors

Jakub is a Research Scientist at SaferAI focused on developing quantitative risk models of AI-assisted cyber misuse. His experience spans both technical and governance aspects of AI safety, having worked on adversarial ML, cybersecurity, compute governance and whistleblowing policies. Previously, he completed a PhD in Particle Physics at the University of Durham, UK.

Matthew is a Research Scientist at SaferAI investigating methods for producing principled and verifiable quantitative risk models at the intersection of AI systems and society in high uncertainty and limited data settings. He has ten years of experience in fundamental Machine Learning research and holds a PhD from Oxford in computer science with a focus on generalisation in reinforcement learning.