Models transmit behavioral preferences through semantically neutral training data, e.g., number sequences from a biased teacher inducing the same bias in students. Building on our information-theoretic framework and steering-vector results, this project maps when transfer succeeds (tokenization, model family, concept structure) and tests whether new silent-state instruments (Anthropic's J-lens) can detect transmitted traits that behavioral evals miss.
About the project
Why do language models absorb behavioral biases from data that looks semantically neutral, and how can this be predicted and detected? Our group treats subliminal learning as a constrained communication channel with three measurable bottleneck factors — encoding cost, channel capacity, and geometric alignment (framework preprint, 2026) — grounded in earlier work on concept geometry (Li & Tegmark, Entropy 2025).
Spring 2026 empirical results: steering vectors extracted from biased-vs-control data causally induce preferences without fine-tuning; single-token concepts respond cleanly to steering while multi-token concepts do not; cross-model transfer works within but not across model families; and shuffling the data destroys transfer. Concurrent work (Blank et al., arXiv:2606.00995) confirms the alignment-mediated mechanism at scale; boundary questions remain open.
A new instrument sharpens the detection side: Anthropic's Jacobian lens (July 2026) reads the single-token concepts a model is poised to produce — precisely the concept class our results show is transmissible. Whether a planted trait is visible in a student model's J-space before it appears in behavior is a concrete, safety-relevant open experiment.
Workstreams: (1) boundary conditions — tokenization-dependence and cross-family transfer limits; (2) substrate detection — SAE features, logit diffing, and J-space readouts of what fine-tuning actually changed in student models; (3) framework-driven prediction — can the three factors forecast transfer success in advance? Mentees inherit working pipelines (steering-vector extraction, fine-tuning + evaluation, logit-diffing tooling).
Target: workshop/main conference paper within an active multi-paper research line.
Theory of change
As model-generated and synthetic data pervade training pipelines, distilled models can inherit traits invisible to data inspection — a direct data-poisoning and misalignment-propagation risk. Behavior-only evaluations miss silently held state; predictive theory plus substrate-level detection tools enable auditing training data and student models for hidden trait transmission before deployment. Prior work: our framework preprint (arXiv, 2026); Li & Tegmark (Entropy 2025); Cloud et al. (arXiv:2507.14805).
Your role
One workstream owned end-to-end per mentee, written weekly updates, high autonomy within scope, changes agreed in meetings. New mentees coordinate with a returning technical lead.
Prerequisites
- Highly proficient in Python and PyTorch
- Comfortable with HuggingFace Transformers, running fine-tuning jobs, and managing evaluation pipelines on GPU (Colab/RunPod).
- Basic probability/information theory (KL divergence, mutual information at working level) helpful for the prediction workstream.
- Steering-vector experience a plus, not required.
Location preference
No geographic restriction. Mentees must be able to attend one of two weekly meeting slots anchored to Central European time (historically Saturday ~13:00 UTC and Wednesday ~14:00 UTC).
Application question(s)
- A student model trained on a biased teacher's number sequences acquires the teacher's preference. Propose one mechanism, and one experiment that distinguishes it from ordinary dataset contamination. (250 words)
- Why might a preference for a multi-token concept resist steering-vector induction when single-token concepts don't? (150 words)
- How do you think the subliminal transfer study can benefit other AI safety and LLM deployment fields, like agentic systems, governace and policy, or so on.
- Link to a repository or technical writing sample.
About the mentor

Yuxiao is an independent researcher in mechanistic interpretability. Before she was a postdoc at the Basque Center for Applied Mathematics (BCAM) and an AI Safety researcher at the Beneficial AI Foundation (BAIF). She was also a SERI MATS scholar in 2022 and a MATS scholar in 2023 both Summer and Winter tracks. Her research interests include information theory, probabilistic frameworks, and their applications for building more theoretically sound and trustworthy AI systems. She has a background in statistical inference, machine learning, and deep generative models. She completed her PhD in Electronic Engineering at Tsinghua University and has mentored research teams with SPAR and Algoverse.