Mentees will build the next generation of BioTIER, SecureBio's benchmark for biological safeguard effectiveness. Work spans multi-turn, multilingual, multimodal, and agentic extensions, plus better scoring for the cases where a model neither fully refuses nor fully complies.
About the project
AI developers need a precise way to measure whether their models refuse dangerous biological requests while still answering benign ones. BioTIER (https://securebio.org/biotier/) is an expert-curated benchmark of 542 prompts across catastrophe avoidance, dual-use biomedical research, and related biology that quantifies safeguard effectiveness and identifies where refusals break down. While the first version of the single-turn static BioTIER evaluation has recently been published, there are now many expansions in the pipeline. How can we judge the ‘safety’ of safe completions? How do we best evaluate safeguards of models within complex agentic harnesses? How does multi-turn, multi-lingual or multi-modal prompting impact refusal behaviour? And how can we best pre-empt and strengthen the bounds between dangerous information and beneficial knowledge? These projects will enable granular characterization of refusal behaviour across the ecosystem, directly influencing safeguard policies and frontier model safety evaluation.
Theory of change
BioTIER results are used by frontier developers to assess and tune their own biological safeguards, and SecureBio evaluation results appear in published model cards. Extending the benchmark to multi-turn, multilingual, and agentic settings closes the gap between what safeguard evaluations currently measure and how models are actually used. If refusal reliably breaks down under conditions the benchmark never tested, current safeguard claims overstate real protection, and developers and regulators are calibrating on the wrong number.
Your role
The mentee will own one extension axis, with weekly direction-setting from the project lead and asynchronous review in between.
Prerequisites
High proficiency in Python. Has built or substantially modified an LLM evaluation harness, or has run LLM API calls at scale for a research project, including personal projects. Comfortable reading model cards, system cards, and developer safety policies. For the multilingual axis specifically: fluent or native command of a language other than English. A biology background is welcome but not required.
Application question(s)
Please answer one of the below, 300-500 words.
- How should a managed-access program for bio-capable models be set up? What are the relevant parameters, what do you suggest, and why?
- Suppose a wet-lab uplift study is run. What kind of data would you want to collect and how would you propose those data inform subsequent in silico model evaluations?
About the mentor

SecureBio is a nonprofit biosecurity research organization specializing in technical research to mitigate risks from catastrophic pandemics. Our AI team develops rigorous benchmarks and evaluation frameworks to assess AI systems' biological capabilities, as well as mitigation strategies that can reduce risks once AI capabilities cross specific risk thresholds. We perform pre-release safety testing of frontier models (e.g. GPT-5.6), and our evaluations have been featured in the model cards of OpenAI, Anthropic, and Google DeepMind. Our work has also informed national security briefings and emerging governance standards.