Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

From Controllability to Concealment A Causal Model Organism of Steganographic Reasoning

Chain of thought Scalable oversight AI control

We will test whether a seemingly harmless capability—controlling one’s own reasoning trace—can become a stepping stone toward a dangerous failure mode: load-bearing reasoning hidden from oversight.

About the project

Can a harmless capability become the foundation for hidden reasoning? This project asks whether teaching a language model to control benign features of its chain of thought (CoT)—such as its length, structure, wording, or abstraction—makes it easier to later develop load-bearing reasoning that remains hidden from oversight.

Prior work shows that models can learn steganographic CoT(https://arxiv.org/abs/2506.01926), that fine-tuning can increase CoT controllability (https://metr.org/blog/2026-04-01-fine-tuning-cot-controllability/), and that controllability correlates with reasoning obfuscation (https://arxiv.org/abs/2605.15257). We will test the missing causal link.

Starting from one open-weight model, we will create matched variants that differ only in the benign reasoning-control training they receive. We will then expose every variant to identical optimization pressure on safe, synthetic state-tracking tasks and test whether control-trained models discover hidden, monitor-evading encodings faster or more reliably.

Crucially, shorter, vaguer, or merely monitor-clean reasoning will not count as steganography. We will require evidence that hidden information can be decoded from the reasoning trace and that causally altering this information changes the model’s answer.

The goal is to produce a reproducible model organism for studying, detecting, and mitigating hidden reasoning. Mentees will be full research collaborators, contributing to model training, causal evaluation, mechanistic analysis, and a potential paper. All experiments will use toy hidden variables and sandboxed monitors; no deployed safety systems or real-world harmful information will be targeted. See full in the proposal.pdf attached

Theory of change

As AI systems become capable of taking increasingly consequential actions, their safe deployment may depend on whether human andvautomated overseers can reliably understand their reasoning. If CoT controllability is a causal precursor to hidden reasoning, thisvwork could provide an early-warning metric, identify risky post-training practices, and create a testbed for monitorability-preserving interventions. A robust null result would also be valuable by preventing labs from relying on a misleading safety signal.

This project builds directly on our previous work showing that language models can learn and generalize steganographicvchain-of-thought under process supervision (https://arxiv.org/abs/2506.01926). By identifying what makes this failure mode easier to acquire—not merely demonstrating that it exists—we hope to support safer training and oversight of increasingly capable AIsystems.

Your role

see proposal

Prerequisites

Mentees should have strong Python skills, experience with PyTorch or Hugging Face, and the ability to reason carefullyvabout confounding, controls, split hygiene, and negative results. Experience with LoRA/SFT, TRL, PPO/GRPO, mechanistic interpretability, or randomization inference is useful but not required. We care more about experimental judgment and reliable execution than prior publication

Location preference

E

Application question(s)

See proposal

About the mentor

Luis Ibanez

Luis Ibanez

Banco Santander

View profile

I am an AI safety and security researcher with a PhD in Computer Science and Technology from Universidad Carlos III de Madrid, recognized with Spain’s national RENIC award for the best cybersecurity doctoral thesis. My work spans steganographic chain-of-thought, deception detection, model-internal probing, membership-inference attacks, adversarial poisoning, and LLM red teaming. I have collaborated with the University of Cambridge, INRIA, FAR.AI, and CISPA, and I currently lead Responsible AI work at Banco Santander creating evaluations and control products. My research has appeared at venues including NeurIPS, ESORICS, IEEE...

Having previously participated in SPAR as a mentee, I understand the value of clear research scoping, regular feedback, and achievable milestones. As a mentor, I would provide hands-on technical support while encouraging mentees to develop genuine ownership of their work. I am particularly interested in mentoring projects involving LLM evaluations, jailbreak robustness, interpretability, deception, scalable oversight, and adversarial robustness.

Similar projects