Can we prefill an open source model with the generations of an untrusted blackbox and probe it for misbehavior that we otherwise wouldn’t catch? We want to establish empirical baselines of how faithful a whitebox proxy can be and explore techniques for increasing it.
About the project
AI misbehavior gets harder to catch as AI becomes more powerful. It would be great to use whitebox techniques to diagnose them, eg probing for deception, eval awareness, power seeking. But we’ll never get access to the internals of frontier models, however, if features representations are universal enough, you could prefill a trusted whitebox with the generations of an untrusted blackbox, diagnose the whitebox, and make claims about the internals of the blackbox.
The high-level variables are model size, model family, and representations. We define “faithfulness” as how well the whitebox proxy can make predictions about the blackbox’s internal representations. A priori, we expect faithfulness to be “jagged”, e.g. deception vs sycophancy probing might have different faithfulness levels. Different techniques (eg NLA, AO, SAE, probing) will likely also be jagged and may require different validation methods. Downstream tasks will include probing for misbehavior and looking for nonverbalized reasoning in the cot.
After establishing proxy faithfulness baselines, we’d be interested in improving faithfulness. Distilling from the untrusted blackbox to the trusted whitebox is the most expensive, simple upper bound. There may be ways to make this more efficient, or narrowly improve faithfulness on important tasks. There are other ideas, like changing the whitebox’s representations in-context.
Stretch goal: One obvious failure mode would be the whitebox model failing to catch unverbalized reasoning in the blackbox model even given CoT, and we know the major frontier models also encrypt and summarize the reasoning traces. Measuring and mitigating the proxy faithfulness without reasoning, or with unfaithful CoT, is an important aspect we want to touch.m
This work is empiric and subject to change. The latest iteration of the project proposal is here: https://docs.google.com/document/d/1iKF2YSIANUviPkbHM2b3lPM0tNiDKoC3yOrt-0rixeU/edit?tab=t.0
mentor availability: (John Yan): I can be available for direct coworking and co-researching for up to 5 hrs/week/mentee. I expect them to be generally intelligent fast learners, willing to work 15-20 hours/week. Strongly preferred in person in SF. (Yuqi Sun): I can work 3-5h/week with detailed guidance and coworking (e.g. debugging blockers or brainstorm for methodologies, and not to promote for the situation but if we are very urgently facing a paper deadline i will jump in to help accelerate the work). Bay area in person preferred.
Theory of change
This adds interpretability techniques to the available tools for scalable oversight. The success of this work will enable better third-party monitoring and empower people outside frontier labs to secure their AI.
A weak version of success would look like: deception, sycophancy, powerseeking probes for effective runtime monitoring. A strong version of success would look like: improved pre-deployment elicitation techniques and improved post-hoc analysis on black swan events. A positive externality of this is one could use weaker models to diagnose stronger models, or enable model developers (including in the labs) to distill their strongest model into a cheaper model to make more expensive oversight techniques viable.
We’ve already done work where we use SAEs on a 27B model to look at generations of a 235B model, discover features that blackbox baselines could not, and validate them.
Your role
Mentees will essentially work as research engineers, we have well scoped plan but also encourage mentees to participate in literature review and ideation on specific methodology, we expect high agency of running planned experiments and discuss progress in weekly or biweekly meetings, mentees will also participate and gain experience in paper writing and publishing
Prerequisites
Track record of long-term (>4mo) technical projects (usually predicted by 3+ YoE as a SWE/MLE) Python preferred Preferred prior publications/experience at research institutions Preferably in the domain of model training/interpretability, including experience of working with model representational space and model distillation
Location preference
Prefer in person in Bay Area
Application question(s)
- Describe the most impressive project you’ve worked on
- Describe the hardest you’ve ever worked (if different)
- In the research proposal earlier in this agenda, we suggest that whitebox proxy faithfulness will be “jagged”.
- A priori, in what circumstances would you predict the proxy to be more faithful and less faithful, and why?
- We also have some proposed methods for increasing faithfulness: distillation, in context learning.
About the mentor
Yuqi is the founder of Mindoverflow.ai focusing on scalable ai safety research, previously worked on SHADE Arena agent sabotaging benchmark with Anthropic and vision GenAI at Apple.
I am co-mentoring with John Yan, who is founding Gutenberg PBC, an interpretability company. John is interested in using interp to build effective scalable oversight. From John: You can see more about our overall research agenda here: https://app.notion.com/p/Public-Gutenberg-Research-3a2c09e354d8806ea2e2dded9ba3f755 and our latest work here https://gutenberg.ai/blog/rl_rollouts_diplomacy/