Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Distillation-induced Teacher Attribution Bias

AI control Behavioral evaluation of LLMs Scalable oversight

LLM self-preference bias has led to concerning behavior related to monitors downplaying harmful actions under self-monitoring scenarios. In this project, we intend to study the effects of distillation on amplifying self-preference bias to the teacher and therefore downplaying harmful actions.

About the project

In this project, we propose to distill students from fixed teachers under techniques varying in exposure to teacher text and measure, self-recognition collapse and monitoring bias.

Project doc: https://docs.google.com/document/d/1hsrZCaBGnxY42bYdBGaW\_BmZTchE9jbwRreaXtxzhrQ/edit?usp=sharing

Theory of change

We believe this is an important area of study because in practice, weaker models are often used to monitor stronger models and it is possible that these weaker models are distilled directly from expensive teacher models (or from an earlier checkpoint), and later deployed as trusted monitors. Teacher-directed bias in monitoring might thus facilitate monitor collusion without explicit scheming and static off-policy monitor evaluations would miss it entirely.

Your role

Mentees will read relevant research papers and attend an initial knowledge sharing presentation from the mentor(s) which can spark additional research directions. These additional research directions might form the basis of new directions of exploration once we are done implementing the core deliverables of the project.

Mentees will be provided access to existing code for self-recognition and preference evaluations and they have the option to create a new codebase or adapt the existing one to use in this project. Code for monitor authorship bias will have to be built from scratch or adapted from other sources, mentor support will be provided to build these.

Mentees will also explore distillation techniques that vary in teacher exposure and domain.

Prerequisites

  • Proficient in AI-assisted coding (while avoiding common AI coding pitfalls)
  • Familiarity with LoRA and full SFT using tools such as Axolotl, Unsloth etc.
  • Familiarity or willingness to engage with models from an LLM psychology lens
  • Ability to reason from first-principles about catastrophic safety risks

Application question(s)

  • If a Chinese open-source model says "Hello! I am Claude..", what are the two most likely reasons this might happen? What about if Claude says "Hello! I am Qwen.."? (Max 200 words)
  • From a first principles perspective, why do you think models might prefer their own outputs? (Max 200 words)

About the mentors

Arush Tagade

Arush Tagade

George Washington University, MATS

View profile

Arush is a Computer Science PhD student at GWU, advised by Prof. Shi Feng. He's worked on adversarial robustness and interpretability research in the past, he's currently exploring research relevant to AI Control, Reward Hacking and Fine-Tuning Misgeneralization.

Before starting his PhD, he was a research scientist at Leap Labs creating and benchmarking interpretability tools for automated scientific discovery. He has also been an instructor at various AI Safety training programs including AGISF, CAIS MLSS and ARENA.

Taslim Mahbub

Taslim Mahbub

CBAI

View profile

I am a second-year PhD student at George Washington University's Praxis Labs, advised by Shi Feng, where our research focuses on AI safety and alignment. Over the summer of 2026, I am a CBAI research fellow with Bau Labs. Having transitioned into safety research from an applied ML background, I enjoy helping newcomers find their footing in the field and bring prior mentoring experience from my time as a teaching assistant.

Shi Feng

Shi Feng

George Washington University

View profile

Shi leads a research group at GWU focused on reducing loss of control risks from advanced AIs. His research goal is to ensure reliable human supervision of AIs even as their capabilities rapidly improve. Recently, Shi has been focused on the evaluation and mitigation of risks associated with sabotage, in particular deception, collusion, and honesty.

Similar projects