Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Refining the methodology for quantitative risk modelling of AI-assisted cyber misuse

Cyber risks Misuse risk AI strategy

In this project, we'll improve on SaferAI's methodology for quantitative risk modelling of cyber misuse scenarios. This could include, for example, incorporating new benchmarks as indicators of risk, considering defenders' security postures or modifying our approach based on Bayesian Networks and Monte Carlo simulations.

About the project

(for full information, please see https://docs.google.com/document/d/14nqHzTQErZiPcbb1dNoyIfUr_tlfwMJjNnFjEkCD16I/edit?usp=sharing)

In recent work, SaferAI has developed 9 detailed cybersecurity risk models. As with any model, there are many approximations and assumptions we have to make. In this project, you’ll be working on refining this methodology in order to make our quantitative predictions more accurate. Some questions we can explore:

  • Our risk models use AI evaluation benchmarks as ‘key risk indicators’. That is, if an LLM scores X% on some benchmark, what is the corresponding probability Y that it can successfully complete a given step in our risk model We would like to find answers to the following questions: a) What is the optimal number of benchmarks to use as our key risk indicators? b) How do we select these benchmarks? c) How do we choose which benchmark maps to which step in the risk model?
  • Our cyber risk models for now assume that all steps (e.g. ‘Initial Access’, ‘Privilege Escalation’) are independent of each other. However, this is not true – if an LLM can assist with one step of the ‘cyber killchain’, these capabilities likely correlate with its performance on other steps. How do we account for this?
  • Currently, our models assume that the target is ‘static’. In other words, they only take the offensive side into consideration and omit the target’s defensive response. In reality, cybersecurity teams of course respond to incidents in real time and stop them from proceeding further if possible. How can we include these dynamic, defensive responses in our risk models?
  • Our risk models currently assume a ‘serial’ sequence of steps. That is, the LLM first completes step 1 in the threat scenario, then step 2, …, all the way up until inflicting the final damage (e.g. encrypting the hard drives). In reality, cyber attacks are much less ‘linear’ – skills such as ‘Privilege Escalation’ or ‘Reconnaissance’ often have to be re-used at different points of this attack. How can we incorporate this effect into our risk models?

Theory of change

SaferAI’s Theory of Change revolves around creating quantitative and grounded risk models (as well as risk management practices and standards), so that we enable:

  1. AI labs to apply targeted safeguards
  2. policymakers to prioritize interventions
  3. researchers to design more actionable evaluations
  4. regulators to quantify the level of risk and enforce compliance with existing legislation

Taken together, these four factors will lower the negative impacts of advanced AI systems on our society. By contributing to the development of our risk modelling methodology, you will improve the quality of our predictions and therefore help push these four action points forward.

Your role

  • Mentee will be the lead researcher for the question that we are investigating, with a tight feedback loop between mentee and mentor for fast research iterations
  • Mentee should have sufficient autonomy to make meaningful progress on a week to week basis.
  • Mentor will be available for a weekly call and frequent communication through Slack/email
  • Mentor will support the mentee with experiment design, analysis and write up, but the mentee should take ownership of these tasks
  • An ideal output will involve a workshop paper or a blogpost on SaferAI's website

Prerequisites

(We are open to a wide range of backgrounds, including those considered ‘unconventional’ in AI safety. If in doubt, please apply!)

For this project, we are mainly looking for the following skills (not all are required, though):

  • statistics
  • mathematical modelling
  • conducting experiments in social sciences, psychology, or any other environment with lots of uncertainty and confounding variables

Nice to have:

  • experience specifically in risk modelling would be a big plus, but definitely not required
  • cybersecurity knowledge is a plus, but not required
  • it would be good to have some programming experience (Python), so that you can use our risk modelling software and run quick experiments, but this is not expected at a super-high level and will not be prioritized in the selection process

Time commitment

10

Location preference

Any time zone compatible with UK time is fine.

Application question(s)

  1. Please attach a sample of prior work that demonstrates the skills listed above. This could include papers, blogposts, sample code or even engagement in online discussions. Alternatively, please describe why you think you are a good fit for this project.

  2. (<250 words) How would you go about incorporating a new benchmark into a risk model? What sorts of details would need to be considered in order to effectively include the benchmark? How can we examine the effect of including this benchmark in the risk model - what sorts of measurements might be useful here?

About the mentors

Jakub Krys

Jakub Krys

SaferAI

View profile

Jakub is a Research Scientist at SaferAI focused on developing quantitative risk models of AI-assisted cyber misuse. His experience spans both technical and governance aspects of AI safety, having worked on adversarial ML, cybersecurity, compute governance and whistleblowing policies. Previously, he completed a PhD in Particle Physics at the University of Durham, UK.

Matthew Smith

Matthew Smith

SaferAI

View profile

Matthew is a Research Scientist at SaferAI investigating methods for producing principled and verifiable quantitative risk models at the intersection of AI systems and society in high uncertainty and limited data settings. He has ten years of experience in fundamental Machine Learning research and holds a PhD from Oxford in computer science with a focus on generalisation in reinforcement learning.

Similar projects