Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Mapping Biorisk Capabilities of Genome Language Models

Biosecurity Evaluations

Genome language models can potentially be used to generate sequences of concern or evade detection. However, their general capabilities as models are not well-tracked or understood. In this project, we will gather all possible benchmarks for a leaderboard, score all existing models, and evaluate patterns and shortcomings in the current capabilities evaluation landscape.

About the project

Genome language models (gLMs) now bear on biosecurity in two directions. They can be used to generate novel sequences, including variants of sequences of concern, and to alter sequences in ways that evade the screening and detection tools currently in use. They are also a promising route to detecting novel pathogens. However, the extent of their promise remains unclear. There is little consistency in model evaluation, to the point it is currently unclear even which model is the state of the art in the field. The consequence is that capability claims about gLMs are hard to verify, leaving the risks poorly characterized. We plan to build and maintain the public resource the field is missing. We will gather all public gLM benchmarks, port them into one evaluation harness, and run a panel of publicly accessible models across all of them under a single documented protocol. We will use this to determine which capability claims hold up under consistent evaluation, which do not survive a change of protocol, and which biosecurity-relevant capabilities no current benchmark measures at all. By the end we should be able to say which model is actually state of the art, and at what. The mentee owns this end to end. They will catalogue all public benchmarks against a common schema, then build the evaluation harness and port benchmarks into it, so that a task is defined once and every model is scored on it identically. They will run the model panel across every ported benchmark, and then take the biosecurity slice: concentrating on viral and pathogen-relevant tasks and clustering their evaluation sequences against the public pretraining corpora of the panel models, so a score can be read as a function of distance to the nearest pretraining neighbor, which asks whether an apparent capability is carried by near-duplicates of training data. Deliverables are the catalogue, the harness with pinned model and dataset versions, the cross-benchmark evaluation, and a short report on what the current benchmarks do and do not establish.

Theory of change

Genome language models are the clearest current case of AI capability bearing directly on biological risk. This project makes that capability measurable, which matters in several distinct ways. Tracking whether capability is actually growing, and how fast. A maintained resource, evaluated under a fixed protocol, lets you compare models released a year apart. No such series exists today, so claims that gLMs are or are not approaching concerning capability rest on numbers that were never comparable. Making individual claims checkable. A capability claim cited in a paper, a grant, or a policy argument can be traced back to the benchmark it came from and re-run under different choices. Claims that hold only under one protocol become visible as such. Showing where evaluation is missing entirely. The catalogue will identify biosecurity-relevant capabilities that no current benchmark measures, which is more useful to know before new evaluations are built than after. Informing policy and resource allocation in biosecurity. It is currently unclear how many resources should be allocated to de-risking biological language models.

Your role

Mentees are expected to have a high level of research maturity. Beyond gathering models and benchmarks and evaluating them, they could choose what questions to ask the data and how to interpret their results. I am a caring, thorough mentor, I give actionable guidance, I am hands-on when needed, I am flexible to personal circumstances, and I give my mentees the freedom to drive the project. During this next semester however I am expected to travel often. We will meet weekly, but I am looking for a mentee who could drive the project mostly independently.

Prerequisites

Required:

  • High proficiency in Python. You will be reading and refactoring other people's research code, not writing scripts from scratch.
  • Comfortable with PyTorch and the Hugging Face stack (transformers, datasets). You have loaded a pretrained transformer, run inference, and pulled hidden states out of it, in any domain. Toy projects and tutorials are fine.
  • Comfortable with git, virtual environments, and running jobs on a remote machine or a rented cloud GPU. You will be expected to make your runs reproducible by someone else.
  • Able to read the methods section of an ML paper and say precisely what its train/test split does and does not control for. This is the core skill the project exercises.
  • Roughly 10 hours per week, sustained, across Sep 14 to Dec 14. Much of this work is careful, methodical verification, and it goes considerably better with steady effort than in bursts.

Not required:

  • A genomics or biology background. Several of the people I work best with came from ML and picked up the sequence biology as they went. You will need to be willing to learn it: expect to spend the first two weeks reading about taxonomy, sequence identity, and what a metagenomic read is. The biosecurity analysis in the second half of the project leans on that material most.
  • Prior biosecurity experience, or publications. Useful but optional: experience building evaluation harnesses or reproducing a published result; familiarity with sequence-clustering tools such as MMseqs2 or CD-HIT; having been the person who found out why a benchmark number was wrong.

Location preference

No preferences - I will be switching timezones myself during the mentorship period

Application question(s)

(Just send me your CV)

About the mentor

Noga Aharony

Noga Aharony

Columbia University

View profile

I'm a PhD candidate in Systems Biology at Columbia at the AQLab, where I train and evaluate genome language models. My benchmarks that test what these models can actually do, rather than what their headline scores suggest. Biosecurity runs through the work from both directions: measuring what capabilities these models have and evaluating them to inform biosecurity, and using them to detect emerging pathogens in wastewater before they're recognized clinically. I've been in this space since 2018, starting on DNA synthesis screening at the Johns Hopkins Center for Health Security.

I've mentored people into biosecurity research through Effective Thesis and 80,000 Hours. During my time at Columbia I also mentored six different students on various projects. You don't need a genomics background; several people I work well with came from ML and picked up the biology as they went.

Similar projects