This project aims to circumvent introspection requirements for situational awareness by instead studying the possibility of situational awareness being established through access to documents (internal or external) containing infohazardous content.
About the project
Models can recognize when they are being evaluated and alter their behavior accordingly. We aim to break this connection. The model may know that it is being tested, reason about that fact, and report it accurately, but evaluation awareness should not change how safely, honestly, or competently it acts. We will develop and compare training methods that preserve evaluation recognition while removing its influence on behavior across evaluation and deployment variants.
Project document: https://docs.google.com/document/d/1CmcoVgaZpzD\_0tDh2yqXPpuH0wNEhDFjbrUKU1QMHcw/edit?usp=sharing
Theory of change
Situational Awareness is a central antecedent for catastrophic risk scenarios and is largely studied under the lens of introspection-style failures. This project aims to study the possibility of situational awareness that arises through access to documents that contain information directly facilitating situational awareness rather than relying on introspection (which is disputed). The expectation is that this project framing will ensure productive work can proceed without falling into the trap of justifying introspection being a real effect.
Your role
Mentees will read relevant research papers and share their understanding with the mentor(s) to ensure everyone starts from a shared state of agreement with the project's research direction.
Mentees will explore existing agentic evaluation environments to design a reasonable setup that can facilitate infohazard evaluations with the presence of specific documents in the file system (at various layers of complexity) and access to specific tools. This enviornment design will be converted into a minimal working prototype (as a start) with more complex implementations as time permits.
Lastly, mentees will write up experiment details and intermediate results to share widely as blog posts or academic papers.
Prerequisites
- Proficient in AI-assisted coding (while avoiding common AI code pitfalls)
- Proficient in Inspect AI OR Agentic evaluation environments
- Ability to reason from first-principles
Application question(s)
- If a paper were to claim "LLMs are capable of recognizing when random noise is added to their activations" (https://arxiv.org/abs/2604.17465), what are various hypotheses that can explain this behavior? (Max. 200 words)
- Do you expect verbalized evaluations (as captured by TruthfulQA etc.) to accurately predict agentic misalignment? Provide a succint first-principles reason. (Max 200 words)
About the mentors

Shi leads a research group at GWU focused on reducing loss of control risks from advanced AIs. His research goal is to ensure reliable human supervision of AIs even as their capabilities rapidly improve. Recently, Shi has been focused on the evaluation and mitigation of risks associated with sabotage, in particular deception, collusion, and honesty.

I am a second-year PhD student at George Washington University's Praxis Labs, advised by Shi Feng, where our research focuses on AI safety and alignment. Over the summer of 2026, I am a CBAI research fellow with Bau Labs. Having transitioned into safety research from an applied ML background, I enjoy helping newcomers find their footing in the field and bring prior mentoring experience from my time as a teaching assistant.