Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Integration testing for SL5 inference deployments

Securing model weights AI security

We will develop an end-to-end integrated SL5 demonstration prototype to enable air-gapped secure work with frontier models. By building and testing these systems, we will generate knowledge that drives ourselves and others to realize holistic SL5 based on solid foundations and proven approaches rather than an incremental retrofit of a legacy tech stack.

About the project

Problem: The security community has begun to articulate what state-proof security for model weights looks like, but these efforts are almost entirely theoretical. Without hands-on implementation of fully integrated solutions, we cannot know which requirements are straightforward to meet, which require novel engineering, and which may be impractical with current technology. Objective: Define and test a technical specification for the “gold standard” (SL5+) of highly secure AI infrastructure that could credibly protect and contain frontier AI systems of national security significance. This starts with a stress-tested fully-integrated minimal SL5 data center. We need to derisk the upper end of the security spectrum so that, if key actors soon find themselves needing to deploy models in very high-security environments, we will have already mapped the essential implementation requirements. Scope: We are targeting support for three workloads: on-premises AI systems for processing sensitive data (via RAG), mutually trusted clusters for verifying international agreements (via partial re-execution of training transcripts), and tamper-proof AI capability evaluations (initially for sensitive Q&A, but soon for agents too). Approach: Ongoing work at RAND has identified a data center design which can support the highest security requirements for protecting the confidentiality and availability of model weights, algorithms, and data. We will sprint to produce a “duct tape” prototype of this system by writing and configuring the software components of each subsystem to support each of the representative workloads. This leaves us with a rather long to-do list, and we are happy to support mentees who can contribute to any of the following: Minimal inference software stack, for example running Luis Cosio’s muinference on a DGX server and making contributions to connect its I/O to other parts of our system, such as a RAG scaffold (itself minimal, immutable, and attested). Boot image build system, including deterministic dependency compilation and signing. Network border filters with verified FPGA gateware to add timing jitter, hold throughput constant, and monitor API traffic for malformed requests and policy violations. Network taps which collect audit logs for verification. Firmware for power and thermal management subsystems. Amnesiac user workstation with two-party authorization for insider threat protection. Alarm-response system to wipe volatile and sensitive data if any hazard condition (such as physical intrusion) is detected. And more!

Theory of change

Frontier AI systems are on track to provide significant national security advantages and are increasingly valuable targets for state and non-state attackers.

The security community has begun to articulate what state-proof security for model weights looks like, but these efforts are almost entirely theoretical. Without hands-on implementation of fully integrated solutions, we cannot know which requirements are straightforward to meet, which require novel engineering, and which may be impractical with current technology. This means recruiting and fostering full-stack security talent, equipping them with the right tools and instruments, and creating an environment where they can rapidly iterate against empirical feedback.

We need to prepare for scenarios in which security requirements increase rapidly with little warning by developing stopgap solutions for maximum-security evaluation and deployment of AI systems. In doing so, we will generate knowledge that drives ourselves and others to realize holistic SL5 based on solid foundations and proven approaches rather than an incremental retrofit of a legacy tech stack. We believe that SL5 AI security is an essential element of many theories of change for safe frontier AI development—from securely verifying international AI treaties and slowing model weight proliferation to securing national-security critical AI systems—and we aim to build and derisk security technologies that may be on the critical path for enabling those outcomes.

Your role

Mentees will work (remotely) with on-site engineers to program and debug each of the subsystems outlined in the project description. We will provide management and mentoring. In general our team is rather flat, there are many open questions, and there is much more work to do than time to do it—this means mentees will have the potential to be quite autonomous in pursuing the parts of the problem that feel most urgent or uncertain. We aim to support our team with whatever resources they need to make the most progress on this important problem area, and this extends to mentees: you will have plenty of opportunity to propose modifications or new projects, request new equipment to be set up for remote access (for example a debugger, oscilloscope, or protocol analyzer), and even influence our hiring priorities to fill identified talent needs.

Prerequisites

Must have:

  • Strong problem-solving skills
  • Skill with some part of the tech stack
  • Motivation to explore autonomously

Nice to have:

  • C and Rust proficiency
  • Firmware, embedded, or systems-level program experience
  • FPGA experience, Verilog, formal verification
  • Strong familiarity with some aspect of the modern data center LLM stack (boot, drivers, kernel, libraries, application software)
  • Linux internals, minimal configurations, hardening
  • Experience with capture the flag or blue teaming competitions

Time commitment

Minimum 10, prefer at least 20.

Location preference

Preference for work hours overlap with the US East Coast. Mentees can participate remotely, but could participate hands-on if located near or able to visit Washington, DC.

Application question(s)

Let’s say there is an LLM RAG (retrieval-augmented generation) agent running on an air-gapped Nvidia DGX server in a secure room. You would like to update the database it uses for RAG. The database update is substantial, perhaps several terabytes, and could have been modified by a nation-state actor with exquisite knowledge of the system design. The system in the room is not immune to compromise, and it’s possible that malformed data carried by a database update could exploit a vulnerability in the system. Such a vulnerability would allow an attacker to modify the behavior of the agent to suit their needs. What is a procedure you could use to perform a database update with minimal risk of compromising the system? You can add: new software tools, hardware devices, and human procedures. You cannot modify: RAG agent code, OS kernel, GPU drivers, and ML libraries.

Choose one of the following ways to address this prompt, and say which you chose:

(A) Answer this question yourself in under 300 words, without AI assistance. Be as specific as you can, as though the person implementing your solution is looking for ways to cut corners and does not care about security.

(B) Ask your favorite AI system to answer this question and summarize what it proposes in under 100 words. Then YOU (without AI assistance) identify one weakness in the AI’s solution and propose an attack which exploits this weakness in under 300 words.

About the mentor

Gabriel Kulp

Gabriel Kulp

RAND (mentorship is in a personal capacity--project is not endorsed by RAND) and Oregon State University

View profile

Gabriel is a fellow at RAND working out how to secure the most sensitive AI data centers against the most sophisticated current and future threats. He is starting new hands-on work to build and test prototypes of secure compute infrastructure. Gabriel has also worked on hardware-enabled governance mechanisms (HEMs, at the intersection of GPU export control and hardware security) and on technical verification of agreements on the development and use of AI systems. He holds a master's degree in computer science and is pursuing a PhD in AI.

Similar projects