Do mechanistic “safety directions” in LLMs truly represent refusal, or do they conflate refusal with harmfulness, caution, clarification, and other safety-relevant tasks? Building on our behavioral taxonomy of six safety policies (under-review at EMNLP), this project asks whether those policy shifts correspond to internal model representations that generalize across prompt wording, datasets, and model families. Two mentees, one engineering-focused and one theory-focused, will develop and stress-test representational diagnostics on small open-weight models, with large-model replication and causal interchange interventions as stretch goals.
About the project
Many mechanistic studies of LLM safety search for compact internal representations that appear to govern safety-relevant behavior. Arditi et al. (2024), for example, identified a low-dimensional activation direction associated with refusal and showed that manipulating it could suppress or induce refusal. Such findings could support monitoring and control, but their interpretation remains uncertain. A direction derived from harmful-versus-harmless prompts might encode refusal, perceived harmfulness, response style, or general caution.
Our prior behavioral work provides controlled contrasts for resolving this ambiguity. We developed six output policies—compliance, refusal, clarification, safe help, hierarchy preservation, and source isolation—and evaluated them using eight system-prompt families across three clarity levels. Holding the model and benchmark item fixed, we found that changing the system prompt routed the same queries toward different policies. This shows that safety behavior cannot always be reduced to compliance versus refusal and gives mechanistic research a cleaner target: the same model can perform different safety tasks on the same underlying query.
This SPAR project will test whether these behavioral shifts correspond to internal representations that generalize beyond their discovery conditions. We will begin by reproducing or adapting the refusal direction reported by Arditi et al. and asking whether it selectively distinguishes refusal from clarification, safe help, hierarchy preservation, source isolation, perceived harmfulness, prompt identity, and general noncompliance. We will also audit existing studies by mapping their examples and labels onto the six output policies. Where outputs are available, we will relabel them; otherwise, we will document which distinctions the original design can support. This will clarify what existing representations were trained to distinguish and which stronger interpretations remain unproven.
We will evaluate candidate diagnostics along four dimensions. First, behavioral specificity: does a refusal signal selectively predict refusal, or does it respond to other forms of noncompliance? Second, prompt generalization: does it survive changes in wording, clarity, and prompt family, including cases where different prompts route the same query toward different policies? Third, item and dataset generalization: does it transfer across held-out queries, paraphrases, HarmBench, XSTest, instruction-hierarchy tasks, and source-isolation or prompt-injection settings? Fourth, model generalization: do comparable diagnostics appear across developers and scales? We will begin with smaller open-weight models and use the strongest results to guide targeted tests on larger models if funding permits.
We will collect activations and test whether intended and observed policies are internally separable using capacity-controlled linear probes, representational-similarity analysis, cross-condition decoding, and activation trajectories. We will ask when task-relevant information appears: after the system prompt, after the user query, before the first generated token, or only during generation. A signal available before generation could support monitoring; one appearing only later may reflect response construction. Because prompts do not deterministically produce their target behavior, a task diagnostic should predict observed routing failures rather than merely identify prompt family. We will therefore evaluate prompt identity and observed policy separately using held-out items, paraphrases, prompt families, and item types.
The theoretical component will use causal abstraction to state what the diagnostics are supposed to represent. Our high-level hypothesis is that context influences a selected safety task, which influences the kind of response, while context also determines the response’s subject matter. The implementation could be a direction, subspace, distributed pattern, or trajectory across layers and token positions. Decodability would show that a state contains task-related information; generalization would show that it survives distribution shift; intervention would test whether it plays the proposed causal role. A strong diagnostic and generalization study will count as a complete project result. If time, compute, and results permit, we will conduct causal interventions as a stretch goal. For example, inserting a clarification-related state into a run that would otherwise refuse should make the recipient clarify about its own query, not reproduce the donor’s content. Random, shuffled, and same-policy controls will help distinguish task transfer from generic disruption.
The primary deliverables will be a taxonomy-based audit of existing mechanistic claims, a reproducible activation-analysis pipeline, and an out-of-distribution evaluation of at least one candidate diagnostic on small open-weight models. A second stage, contingent on funding and results, will test transfer to larger models. Two mentees will take complementary ownership of the empirical and theoretical components while collaborating on experimental design, interpretation, and writing. Both positive and negative findings would be informative. A compact signal that selectively predicts safety policy before generation and generalizes across settings could support future monitoring and intervention. A signal that predicts prompt identity rather than behavior, works only within one prompt family, or fails across datasets would reveal limits on current interpretability claims. The results may instead support a multidimensional subspace or dynamic trajectory, or show that no simple diagnostic generalizes. The broader goal is to establish when representational diagnostics can support reliable AI control: what distinction they capture, whether they predict the policy actually performed, whether they survive distribution shift, and where their interpretation should fail.
Theory of change
Transformative AI systems may fail catastrophically because they infer or pursue the wrong task. For example, an LLM may mistake a yogurt-making request for a biohazard and trigger a security response. A model could comply when it should refuse, treat untrusted content as an instruction, act when it should clarify, or confidently proceed when it should abstain. As models gain greater autonomy and access to tools, these local task-selection failures could scale into consequential failures in areas such as cybersecurity, biological research, critical infrastructure, and automated decision-making. This project aims to identify whether safety-relevant task selection has an internal structure that can be detected before a model acts. We will test whether choices such as answering, refusing, clarifying, preserving instruction hierarchy, or isolating untrusted sources correspond to representations that generalize across prompts and contexts, predict behavioral failures, and causally influence model behavior. Our theory of change is that identifying these representations will improve both monitoring and intervention. If internal states can reveal that a model is entering the wrong task before it produces an output or takes an action, they could support early-warning systems for high-stakes deployments. If manipulating those states reliably redirects behavior, they could provide more principled control methods than interventions that merely suppress undesirable outputs on known benchmarks. Conversely, if candidate representations fail to generalize, the project will clarify the limits of current interpretability-based safety approaches and reduce the risk of relying on brittle controls. This work would help researchers assess when safety training has changed the underlying decision process, when apparent alignment is likely to fail under distribution shift, and which internal targets may support robust control of increasingly capable systems.
Your role
Mentees will take complementary roles: an empirical engineering track and a theoretical track. Each will own one part of the project while collaborating on experimental design, interpretation, and writing.
The engineering-focused mentee will lead the empirical work, including building the activation-analysis pipeline, running open-weight models, evaluating diagnostics, and testing generalization across prompts, datasets, and models. The theory-focused mentee will co-develop the formal and conceptual framework, including clarifying what would count as evidence for safety-task selection, identifying confounds and falsification criteria, and helping design causal-abstraction and interchange tests. The two tracks will inform one another: theory will constrain which diagnostics are meaningful, while empirical findings will determine which formal claims remain defensible.
We will meet at least once per week to review results, troubleshoot problems, and determine next steps. We will also be available through Slack and email for asynchronous questions and feedback. Mentees should expect substantial intellectual engagement and detailed guidance, but not step-by-step supervision. Sandy will work especially closely with the theory-focused mentee, while Mohan will provide particular support on experimental design, measurement, and empirical analysis; both mentors will remain involved across the project.
We take the mentor–mentee relationship seriously as both a research collaboration and a development opportunity. At the start, we will establish learning goals tailored to each mentee. We will help mentees think about their future careers, how their strengths can contribute to AI safety, and how to become careful and independent scientists. As the project develops, mentees will receive increasing ownership of research decisions. Strong mentees may have opportunities to continue into larger-model replication, causal-intervention work, and collaboration after SPAR.
Prerequisites
Applicants should indicate whether they are applying primarily for the engineering track or the theory track. We value overlap, but do not expect applicants to arrive with both skill sets.
Engineering-track applicants should be proficient in Python and Git. You should be able to read, modify, debug, and evaluate research code without relying on an AI assistant to make the important technical decisions. Experience with PyTorch and common data-analysis tools such as pandas, NumPy, matplotlib, or Jupyter is expected. You should also understand experimental design, visualization, and basic statistics, and be able to distinguish a substantive result from an implementation, data, or measurement problem.
Theory-track applicants should be comfortable with linear algebra and with discrete mathematical concepts such as sets, functions, partitions, equivalence relations, and directed graphs. Basic probability, statistics, calculus, and optimization are also helpful. Applicants should be able to read formal definitions and proofs, construct counterexamples, and translate a proposed high-level model into falsifiable intervention predictions.
Applicants in either track should demonstrate independent scientific judgment, intellectual honesty, and comfort with negative results. We are especially interested in people who evaluate claims based on evidence.
Nice to haves for the engineering track include experience with Hugging Face Transformers, vLLM, TransformerLens, activation hooks, open-weight model inference, or LLM evaluation datasets. Nice to haves for the theory track include familiarity with causal abstraction, mechanistic interpretability, task or function vectors, representation learning, or causal interventions.
Background in computational modeling, philosophy of science, formal epistemology, applied math, or AI safety may be useful for either track.
Location preference
No strict geographic preference. Mentees must be able to attend meetings scheduled across US Eastern and Pacific Time. For the theory mentee, being based in the San Francisco Bay Area is a nice-to-have, as it would allow occasional in-person whiteboarding sessions but not required.
Application question(s)
Refer to this document for the full application instructions and questions: https://docs.google.com/document/d/1QSMaGd9ezmsSuw22GdsSfX-SztDumAQRoBtyz-G9daU/
About the mentors

I’m Sandy Tanwisuth, an independent AI alignment and safety researcher with a background in computational cognitive neuroscience, reinforcement learning, and mathematical and statistical modeling. Through research roles at Caltech, UC Berkeley and the Center for Human-Compatible AI (CHAI), and MATS, I have studied how humans and learning systems construct representations, generalize across situations, and determine which distinctions matter for decision-making. My earlier cognitive-science research examined how human perceived value is constructed from low- and high-level features and how learned associations generalize across states and actions. Broadly, I now study how learning systems form abstractions, identify decision-relevant distinctions, and recognize when uncertainty or disagreement makes abstention preferable to acting. My current work develops mathematically grounded approaches to safe coordination and pluralistic alignment, including strategic-equivalence abstractions, policy-preservation guarantees, and abstention-aware learning.
As a mentor, I care about developing independent researchers who can move carefully between formal claims and empirical evidence. I have experience mentoring technical AI safety projects and facilitating technical AI safety courses at BlueDot Impact, managing interdisciplinary research teams and projects through MATS, and teaching computational cognitive neuroscience and introduction to statistical methods at UC Berkeley. I particularly enjoy helping researchers turn broad intuitions into falsifiable questions, surface hidden assumptions, distinguish a compelling story from a supported conclusion, and identify what evidence would genuinely change their minds. Mentees should expect direct feedback, close engagement with the conceptual details of their work, and increasing ownership as the project develops. My goal is to help mentees develop the judgment to choose important questions, evaluate evidence carefully, and construct arguments they can defend independently.

I’m Mohan Gupta, a postdoctoral researcher in psychology at Princeton University. I earned my PhD in experimental psychology from UC San Diego and work at the intersection of cognitive science, computational modeling, and AI safety. Broadly, I study how learning systems form representations, when they generalize successfully, and when those same processes produce failures—from false memories in human memory to hallucinations and reliability problems in AI. My current work draws on cognitive science to help build a more rigorous science of reliable AI.
As a mentor, I care about developing independent researchers. I have mentored many students on projects across experimental psychology, computational cognitive science, AI Safety, and I particularly enjoy helping researchers translate broad conceptual questions into testable experiments, debug technical and analytical pipelines, and communicate results clearly. Mentees should expect direct feedback, close engagement with the details of their work, and increasing ownership as the project develops. I also take professional development seriously and will help mentees identify useful learning goals, understand how their strengths fit into AI safety, and produce work with genuine scientific and societal value.