Emergent Misalignment, while heralded as an important problem, is relatively understudied with respect to the misalignment profiles it causes with relation to model character entanglements. We aim to shed light to this by studying model roleplaying and intervening on the EM data generation process.
About the project
Emergent Misalignment (EM) finetuning leads to broadly misaligned models due to narrow finetuning. Several defenses have been pitched, including metacognitive interventions that lead to varying amounts of defensive effects; this variation is largely dictated by model families and EM dataset domains. This project will extend our recent work on metacognitive interventions to shed light on EM misalignment profiles as a function of the EM data generation process and the default model capabilities in the narrow EM domain under various roleplaying settings.
Project doc: https://docs.google.com/document/d/1d--Buq-TZn7wiRMuTeKbdMpNEJvCDqozdbsw2kC269g/edit?usp=sharing
Theory of change
Emergent Misalignment (EM) can lead to varying profiles of misalignment and some of these are more concerning than others. While most lead to overt roleplaying that is easily observable in outputs and cause lack of coherence, some EM misalignment profiles can lead to subtly misaligned models that are verbally coherent and aligned but misaligned in agentic contexts. This project will help us understand the driving factors behind these profile differences and hence help inform defenses for the more covert misalignment case.
Your role
Mentees will read relevant research papers and attend an initial knowledge sharing presentation from the mentor(s) which can spark additional research directions. These additional research directions might form the basis of new directions of exploration once we are done implementing the core deliverables of the project.
Mentees will be provided access to existing code and they have the option to create a new codebase or adapt the existing one to use in this project. Mentees will also be provided access to finetuned models and are expected to form internal models of how EM'd models act under normal usage.
Mentees will explore existing EM data generation code and build scripts that can create EM datasets with varying system prompts and framing.
Prerequisites
- Proficient in AI-assisted coding (while avoiding common AI coding pitfalls)
- Familiarity with LoRA finetuning and tools such as Axolotl, Unsloth etc.
- Familiarity or willingness to engage with models from an LLM psychology lens
- Ability to reason from first-principles about catastrophic safety risks
Application question(s)
- Do you consider post-training to suppress/hide the default model character or create it? (Max 200 words)
- Do you expect verbalized evaluations (as captured by TruthfulQA etc.) to accurately predict agentic misalignment? Provide a succint first-principles reason. (Max 200 words)
About the mentors

Arush is a Computer Science PhD student at GWU, advised by Prof. Shi Feng. He's worked on adversarial robustness and interpretability research in the past, he's currently exploring research relevant to AI Control, Reward Hacking and Fine-Tuning Misgeneralization.
Before starting his PhD, he was a research scientist at Leap Labs creating and benchmarking interpretability tools for automated scientific discovery. He has also been an instructor at various AI Safety training programs including AGISF, CAIS MLSS and ARENA.

I am a second-year PhD student at George Washington University's Praxis Labs, advised by Shi Feng, where our research focuses on AI safety and alignment. Over the summer of 2026, I am a CBAI research fellow with Bau Labs. Having transitioned into safety research from an applied ML background, I enjoy helping newcomers find their footing in the field and bring prior mentoring experience from my time as a teaching assistant.

My work sits at the intersection of large-scale ML systems and AI safety research. I'm a Technical Lead at Google Pixel AI, where I build the on-device serving infrastructure that runs production Gemini models directly on phones — powering features such as Magic Cue, and delivering capable models under stringent latency, memory, and privacy constraints at consumer scale. Earlier, at Mineral (Google X), my focus was landing prototype research into high-scale ML systems, taking agricultural edge-AI from early research to production-grade deployed systems. I enjoy quick prototyping, adapting to new ideas, and iterating through experiments.
On the research end, I'm passionate about safety problems and their implications for real-world systems, and excited to work on both empirical and theoretical questions. I conduct AI safety research with the Praxis Research Group at George Washington University (under Prof. Shi Feng), centered on emergent misalignment — how narrow training signals can produce broadly misaligned behavior — and self-generated text recognition, probing what models can discern about their own outputs. My interests also extend to applied, sociotechnical work, such as AI-driven wildfire crisis analysis and respond during emergencies. As a mentor, I welcome early-stage ideas as readily as polished ones — whether empirical or theoretical — and enjoy sharpening them together as we go.

Shi leads a research group at GWU focused on reducing loss of control risks from advanced AIs. His research goal is to ensure reliable human supervision of AIs even as their capabilities rapidly improve. Recently, Shi has been focused on the evaluation and mitigation of risks associated with sabotage, in particular deception, collusion, and honesty.