Explore what data could differentially train future models to conduct useful AI safety research.
About the project
As models become better at automated research, training data may shape which kinds of research they can perform well. This project will investigate what data could improve models' ability to conduct AI safety research without equally accelerating general capabilities research.
Possible data sources include research meeting transcripts, critiques of research proposals, examples of strong and weak research, and records of how researchers develop, revise, or discard ideas. The main output will likely be an exploratory paper or blog post identifying promising datasets, how they could be used, and the main practical and strategic uncertainties.
Theory of change
Purpose-built data could help future models learn AI safety research judgment, problem selection, reasoning, and critique. Identifying valuable datasets now could enable more useful automated safety research during rapid AI progress.
Your role
Mentees will investigate what data could improve models' ability to conduct AI safety research, survey possible sources such as research meeting transcripts, critiques of research proposals, and records of how researchers develop and revise ideas, and produce an exploratory paper or blog post identifying promising datasets, how they could be used, and the main practical and strategic uncertainties.
Prerequisites
Familiarity with machine learning, model post-training, or AI safety research would be useful. Strong conceptual research and writing skills are more important than software engineering experience.
Application question(s)
What type of data might help models conduct better AI safety research?
Why might this data differentially improve safety research rather than research capabilities in general?
About the mentor
Alec Harris
Pivotal Research Extension
Alec Harris is a researcher with the Pivotal Research Extension.