Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Identifying function-relevant signatures in protein models for biosecurity screening

Biosecurity Mechanistic interpretability

This project explores biological foundation models for biosecurity screening, applying interpretability methods to identify biophysically relevant features that can reinforce screening against engineered and AI-designed biological threats.

About the project

Advances in AI capabilities for biology are challenging existing biosecurity safeguards, particularly DNA synthesis screening, which flags orders based on similarity to known sequences of concern (Baker & Church, 2024; Hunter 2024). It remains unclear to what extent biological AI models can reliably generate functional genes and proteins outside natural sequence space, rather than merely interpolating within their training data. But the further that designed proteins depart from known biological distributions, the more likely sequence-based methods are to fail. A red-teaming study by Wittmann et al. (2025) showed that some AI-redesigned proteins can evade established screening methods, though it’s uncertain whether these failures reflect limited generalization to divergent sequences, or practical constraints on false positives.

Protein language models are trained on large sets of protein sequences, and learn representations that can encode aspects of protein structure, function, and evolutionary relationships (Rives et al., 2021; Candido et al., 2026). These representations may reveal biologically relevant similarities that are not readily identifiable from direct sequence comparison alone (Liu et al., 2024), making them a promising complement to existing screening methods. If these features generalize beyond the region of protein space explored by nature, they may enable a function-based approach to screening that is more robust against AI-enabled design (Abel et al., 2026).

This research stream will investigate whether biological foundation models can address this screening gap. It explores embedding-based approaches and interpretability methods, such as sparse autoencoders, to identify learned representations that are biophysically meaningful (Simon & Zou, 2025; Candido et al., 2026), and therefore potentially useful for function-based screening. The aim is to determine whether these approaches offer practical advantages over BLAST, profile HMMs, and structural search without inflating the false-positive rate.

References:

Abel, G. R. Jr. et al. (2026). Beyond sequence similarity: toward function-based screening of nucleic acid synthesis. Frontiers in Bioengineering and Biotechnology, 14, 1832724. https://doi.org/10.3389/fbioe.2026.1832724.

Baker, D. and Church, G. (2024). Protein design meets biosecurity. Science, 383(6681), 349. https://doi.org/10.1126/science.ado1671.

Candido, S. et al. (2026). Language modeling materializes a world model of protein biology. bioRxiv. https://doi.org/10.64898/2026.06.03.729735.

Hunter, P. (2024). Security challenges posed by AI-assisted protein design: the ability to design proteins in silico could pose a new threat for biosecurity and biosafety. EMBO Reports, 25, 2168–2171. https://doi.org/10.1038/s44319-024-00124-7.

Liu, W., Wang, Z., You, R. et al. (2024). PLMSearch: protein language model powers accurate and fast sequence search for remote homology. Nature Communications, 15, 2775. https://doi.org/10.1038/s41467-024-46808-5.

Rives, A., Meier, J., Sercu, T. et al. (2021). Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences, 118(15), e2016239118. https://doi.org/10.1073/pnas.2016239118.

Simon, E. and Zou, J. (2025). InterPLM: discovering interpretable features in protein language models via sparse autoencoders. Nature Methods, 22, 2107–2117. https://doi.org/10.1038/s41592-025-02836-7.

Wittmann, B. J., Alexanian, T., Bartling, C. et al. (2025). Strengthening nucleic acid biosecurity screening against generative protein design tools. Science, 390(6768), 82–87. https://doi.org/10.1126/science.adu8578.

Theory of change

Powerful foundation models trained on large biological datasets continue to scale in capabilities, and increasingly pose a risk to established biosecurity safeguards like DNA synthesis screening. At the same time, frontier language models may provide uplift that lowers the skill barrier for malicious actors to engineer new pathogens. Securing the tools of bioengineering against misuse in this context requires that we leverage powerful AI capabilities for defensive purposes. This work lays the foundation for harnessing bio foundation models to improve the sensitivity of biosecurity screening, particularly for highly divergent AI-designed sequences, reinforcing a critical barrier against catastrophic accidents and misuse.

Your role

Under direction and guidance from Isha and Gary, mentees will pursue focused technical research in support of the project, and within a defined scope and objectives that are tailored to the individual’s background, interests, and level of time commitment. See the attached proposal for more details.

Prerequisites

  • A background in bioinformatics, computational biology, structural biology, biochemistry, biophysics, protein engineering, biosecurity, AI/ML science, computer science, or a related field.

  • Prior technical research experience.

  • A good understanding of the basics of biomolecular sequence, structure, and function.

  • Some experience with biological AI models.

  • Proficiency with Python.

  • Strong critical thinking and creative problem-solving abilities.

  • Curiosity and a desire to understand the world.

  • The integrity and judgment to responsibly carry out sensitive research.

Location preference

Preference for people who can meet during 8AM-8PM UK time

Application question(s)

If applying to our stream, please complete the DNA Screening Exercise linked below. You should spend no more than an hour on it. You can submit your answer as text.

https://docs.google.com/document/d/16fYUnr9a3c-lOB0Of1mrLwB9p2LgzabyNcUYnQr1Wv4/edit?usp=sharing

About the mentors

Isha Harris

Isha Harris

Fourth Eon Biosecurity Institute

View profile

Isha Harris is a Research Fellow at Fourth Eon Biosecurity Institute, and a medical student at Cambridge University. Her research focuses on applying and interpreting biological foundation models for DNA synthesis screening, asking if learned representations can identify structurally or functionally related proteins at low sequence identity. Isha first developed this research as a SPAR Fellow herself, making her well placed to serve as a mentor and help fellows design and execute on a relevant project that is both feasible and impactful.

Gary Abel

Gary Abel

Fourth Eon Biosecurity Institute; the Johns Hopkins Center for Health Security

View profile

Gary Abel is Chief Scientist and co-founder of Fourth Eon Biosecurity Institute, where he leads research on adaptive biosecurity safeguards and function-based screening. His expertise spans chemistry, molecular biophysics, biosecurity, and sequencing technology. He's spent nearly two decades studying how DNA, RNA, and proteins behave and interact.

Gary is also a Contributing Scholar at the Johns Hopkins Center for Health Security, supporting the Center's work to understand and mitigate global catastrophic biological risks from advanced AI. He has previously mentored Research Fellows through SPAR, MATS, Cambridge ERA, and the Coefficient Giving CDTF. Gary holds a BS in Physics from San Jose State University and a PhD in Chemistry from University of California Merced.

Similar projects