Mentees will build a jailbreak module that applies published attack techniques to SecureBio's biosecurity benchmarks. Work spans reproducing known attacks, defining a metric for capability recovered under adversarial pressure, and building the harness so it composes with existing evaluations.
About the project
Safeguard evaluations measure whether a model refuses, yet a determined user can often bypass those refusals with jailbreaks the eval never sees. This project builds a plug-and-play jailbreak module for SecureBio's benchmarks to measure how much dangerous biological capability remains accessible once model safeguards are stressed.
Theory of change
Frontier developers, and increasingly regulators, treat safeguard evaluation scores as evidence that a model is safe to deploy. Those scores report refusal rates against a fixed prompt set, which is an upper bound on real protection, and the gap between that bound and adversarial reality is currently unmeasured for biological content. If safeguards collapse under standard published attacks, the evidence resting on those scores is much weaker than it appears. Reporting a robustness-adjusted score alongside every benchmark result gives developers a target that reflects how models are actually pressured, and gives external parties a more honest number to act on.
Your role
The mentee will build the module, with the project lead setting scope and reviewing outputs. This project works with published attack techniques rather than novel ones, and results are handled under SecureBio's internal review before any external sharing.
Prerequisites
High proficiency in Python, including writing reusable library code rather than one-off scripts. Familiar with the published red-teaming and jailbreak literature. Strong judgement about handling sensitive results. A biology background is welcome but not required.
Application question(s)
Please answer one of the below, 300-500 words.
- How should a managed-access program for bio-capable models be set up? What are the relevant parameters, what do you suggest, and why?
- Suppose a wet-lab uplift study is run. What kind of data would you want to collect and how would you propose those data inform subsequent in silico model evaluations?
About the mentor

SecureBio is a nonprofit biosecurity research organization specializing in technical research to mitigate risks from catastrophic pandemics. Our AI team develops rigorous benchmarks and evaluation frameworks to assess AI systems' biological capabilities, as well as mitigation strategies that can reduce risks once AI capabilities cross specific risk thresholds. We perform pre-release safety testing of frontier models (e.g. GPT-5.6), and our evaluations have been featured in the model cards of OpenAI, Anthropic, and Google DeepMind. Our work has also informed national security briefings and emerging governance standards.