Mentees will replicate a load-bearing paper or result in AI safety to improve the empirical foundations of the field.
About the project
Second Look Research aims to replicate load-bearing, empirical AI safety results to verify key claims and improve the epistemic foundations of the field. Mentees will scope and execute an empirical replication under the guidance of the Second Look team.
Each replication involves several steps:
-
Target selection: Determining the specific result the mentee will replicate, with emphasis on load-bearing results to ensure a strong theory of change. We provide a list of target replications we are excited for, but are also open to pitches from mentees.
-
Literature review and scoping: Look for prior replications or other relevant work, and scope out which particular experiments are most important to run.
-
Replication: Actually implement the relevant experiments and attempt to replicate the empirical results. This will be the bulk of the process!
-
Writeup: Communicate our results to the original authors, and publish a blog post, along with open-sourced code, reporting our findings to the AI safety community.
Theory of change
AI safety is a small field in which a comparatively small number of papers shape which problems get prioritized, where talent flows, and what gets communicated to policymakers. Because much of the AI safety community doesn't engage with traditional peer review (and because reviewers rarely run experiments themselves) there is no systematic process for checking whether headline empirical claims hold up. If foundational papers turn out to be misleading, the field loses valuable time building strategy on findings that don't generalize, or imperfect solutions could give labs or the research community a false sense of safety.
Our replications are published openly, written to be brief and readable, and paired with open-source code and interactive visualizations wherever possible, so readers can reach their own conclusions rather than being told what to think.
Your role
We typically encourage mentees to replicate papers / experiments from a list of papers/experiments we have selected for being especially impactful candidates for replication. However, mentees may propose their own projects and we are often happy to mentor these as well. Fellows are encouraged to come up with ideas for ablations or extensions for the paper/results they aim to replicate but we are willing to suggest these as well.
Fellows may work independently or with other fellows depending on overlapping interests or personal preferences of the fellows.
We strongly encourage program fellows to write a blog post on their findings by the end of the program. This would be posted at https://secondlookresearch.com/replications and cross-post them on LessWrong). In addition, we have had fellows from past fellowships we have run produce papers which they have submitted to conference workshops.
Prerequisites
- Experience running and/or fine-tuning open-weight models
- ARENA (chapter 0 and 1) equivalent technical knowledge
- AI safety context (familiarity with arguments for AI x-risk, AI safety literature and research agendas)
- Preferred: legible ML research output or open-source project
Location preference
n/a
Application question(s)
-
Propose a potential replication target. Why this result? How does replicating this result reduce x-risk? (150 words, no need to spend more than 30 minutes on this task)
-
Read one of the blog posts linked below and write a critique of their argument. (150–400 words)
Please be specific and aim for depth over breath. There are a wide range of points that you could object to with any of the below blog posts, but please try to identify and expand on a single crux you think is especially relevant. No need to spend more than one hour on this (for longer posts, you don’t need to read the whole thing).
Reward is not the optimization target (https://www.lesswrong.com/posts/pdaGN6pQyQarFHXF4/reward-is-not-the-optimization-target) Broad Timelines (https://forum.effectivealtruism.org/posts/HCR2AE9it279ggiZT/broad-timelines) Fitness-Seekers: Generalizing the Reward-Seeking Threat Model (https://www.alignmentforum.org/posts/bhtYqD4FdK6AqhFDF/fitness-seekers-generalizing-the-reward-seeking-threat-model) AI for AI safety (https://www.lesswrong.com/posts/F3j4xqpxjxgQD3xXh/ai-for-ai-safety) Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology) (https://www.lesswrong.com/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a) Risks from Learned Optimization: Introduction (https://www.lesswrong.com/posts/FkgsxrGf3QxhfLWHG/risks-from-learned-optimization-introduction) Interpretability Will Not Reliably Find Deceptive AI (https://www.lesswrong.com/posts/PwnadG4BFjaER3MGf/interpretability-will-not-reliably-find-deceptive-ai) How do we (more) safely defer to AIs? (https://www.lesswrong.com/posts/vjAM7F8vMZS7oRrrh/how-do-we-more-safely-defer-to-ais)
About the mentors

I am running Second Look Research (https://secondlookresearch.com/) with Yixiong Hao. Our goal is to replicate load-bearing AI safety work and run relevant ablations and follow-up experiments to validate key claims driving empirical alignment/control work.
Previously I led the creation of the XLab AI Security Guide (https://xlabaisecurity.com/), and I led a project in Ari Holtzman's Conceptualization Lab studying how beliefs and personas in LLMs could be amplified over iterative fine-tuning (https://arxiv.org/abs/2605.01130).
I got involved in AI safety in my sophomore year of high school and started Second Look before I finished my undergrad at UChicago this year. I was a founding member of the UChicago AI safety student group and previously a Pathfinder mentor. In addition to weekly meetings and asynchronous feedback, I'm excited to give career advice and help fellows find internships/full-time roles to continue their AI safety journey! Second Look Research is also hiring so fellows may be offered full-time or part-time roles/internships after stream completion.

I'm a co-founder of Second Look Research (https://secondlookresearch.com/), where our mission is to ensure that empirical AI safety research does not run into replication crises (https://en.wikipedia.org/wiki/Replication\_crisis) that other academic fields have suffered from. I expect that mentees will either continue carrying out replications of key findings or help is implement infrastructure needed for high quality replications at scale. I expect to commit 2 hours for mentorship per mentee/project across asynchronous feedback and meetings.
Outside of Second Look, I do both generalist work and empirical research. I co-directed the Georgia Tech AI Safety Initiative for the past 2 years, did AI security research at Gray Swan AI, and I'm an affiliate researcher at CAIS. Some of my past research interests include applied interpretability, meta-evaluations, and model organisms. I will be working with the Astralis Foundation on their grantmaking strategy from August to December.