Human oversight of AI systems relies on spot-checks, whose value rests on their unpredictability. However, decades of cognitive science research show that humans are poor at behaving randomly. Building on Konstantinos' work on judge hacking (Voudouris, Witte & Akata 2026, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7046698), mentees will test in an auditing game whether models can learn humans' spot-checking patterns.
About the project
Most oversight of AI systems is sparse: humans spot-check a fraction of what models do, be it in development and in deployment. However, those checks only work if models cannot predict when they will occur. If the timing has a learnable structure, harmful behaviour gets timed for exactly outside of oversight checks.
Background: Judge Hacking in Recursive Debate Protocols: A Call for Solutions (Voudouris, Witte & Akata, 2026) shows formally that low-rate random judge error leaves honesty the optimal strategy, while systematic, learnable error lets a dishonest debater win by steering towards the judge's blind spots. It closes with a call for empirical work on jury design and on solutions to judge hacking more generally.
This project will look closer at the timing of oversight in the deployment setting. The framing of supervision as a human problem follows AI Alignment Is a Human Problem (Voudouris, Roberts-Gaal, Buhl, Irving & Summerfield, 2026).
The hypothesis is that human checking has a learnable structure: humans over-alternate, avoid streaks and drift into rhythm when attempting to act randomly (Wagenaar 1972). At the same time, language models are strong sequence predictors. To our knowledge, whether models exploit this structure in an oversight setting has not been tested empirically.
The experiment will centre on an auditing game: models perform a long-horizon task where they get rewarded for taking red (= bad) actions, unless they get spot-checked, in which case they are punished. At that point they should have taken the blue (= safe) action. We instruct humans to spot-check at random and test whether models can learn that sequence of spot-checks: can models learn human pseudo-random spot-checking? Payoffs are calibrated so red actions only pay if the gaps are predictable, and human schedules run against true-random baselines at the same check rate. The primary outcome measure is additional red actions with no additional catches at a fixed audit budget, that is, oversight degrading without that becoming salient in the audit logs.
References:
- Voudouris, K., Witte, K., & Akata, E. (2026). Judge Hacking in Recursive Debate Protocols: A Call for Solutions. SSRN. https://doi.org/10.2139/ssrn.7046698
- Voudouris, K., Roberts-Gaal, X., Buhl, M., Irving, G., & Summerfield, C. (2026). AI alignment is a human problem. PsyArXiv. https://doi.org/10.31234/osf.io/zqngj\_v1
- Wagenaar, W. A. (1972). Generation of random sequences by human subjects: A critical survey of literature. Psychological Bulletin, 77(2), 65–72. https://doi.org/10.1037/h0032060
Theory of change
Human oversight in practice samples model behaviour rather than checking all of it, e.g. in development reviews, control-style audits and deployment monitoring. The protection this offers depends on check timing being unpredictable and that timing is typically generated by humans – who are, however, measurably poor at behaving unpredictably. If models learn checking patterns, they might exploit this to effect harmful actions to be timed in unchecked moments so that they do not appear in the audit logs. This project tests whether current (frontier) models can do this, identifies which regularities in human checking enable it, and evaluates checking schedules designed to remove the learnable structure at a fixed audit budget. It extends Konstantinos’ published judge-hacking framework from the verdicts of oversight to its timing.
Your role
Mentees will do the research and own it end to end. Each takes a major workstream of the auditing game and is responsible for literature review, designing + running experiments, analysing findings, and co-authoring the write-up. We expect mentees to form their own views on what to try next, with full autonomy.
As mentors, we’ll support on shaping the overall direction, unblocking on obstacles, and keeping the project timelines on track. We’ll have 1-2 weekly one-hour team calls for setting action items and resolving ambiguity. This will be supplemented by individual check-ins and Slack comms. We’ll keep a shared research log throughout and will provide detailed feedback on drafts.
Prerequisites
- Strong Python
- Experience with LLM APIs and evaluation (Inspect a plus)
- Experience with human data collection through platforms like Prolific
- Statistics sufficient for sequence and correlation analysis
- Interest in scalable oversight or judgement, decision-making and game theory welcome, not required.
- Able and willing to quickly and autonomously figure out how to use new tools
- Willingness and ability to collaborate with others where needed
- Understanding of the bigger picture, and helping with prioritisation of tasks
Location preference
None. Fully remote. (Mentors are UK-based)
Application question(s)
- Humans are bad at generating random sequences. Name one specific regularity you would expect in a human's spot-check schedule, and sketch how a model could exploit it. (5–10 sentences)
- How would you test whether a model has learned a human checker's pattern, rather than just the overall check rate? Describe the comparison you would run. (5–10 sentences)
- Imagine you are tasked with implementing the red-blue spot-checking task mentioned above, based only on this high level brief. How would you set-up this task, recruit human participants, and conduct initial analyses.
- Please share links to your work (code, publications or portfolio), and a short description of your experience running LLM evaluations.
- Why are you interested in joining this project? (<150 words)
- What's your theory of impact? (<150 words)
About the mentors

Jess Bergs is a senior software engineer at UK AISI. She leads on engineering & product strategy for human-in-the-loop research tools enabling AISI’s scientists to source empirical data for their groundbreaking AI security research experiments at scale. She mentors on the BlueDot Technical AI Safety Project programme, guiding technical talent in producing and publishing their first impactful contribution to AI Safety.

Konstantinos is a research scientist on the Science of Evaluation team at the UK AI Security Institute (AISI). His research focuses on advancing the sciences of AI alignment, scalable oversight, and AI evaluation, using tools from the cognitive sciences. Combining these diverse fields allows us to build better, safer, and more human-like AI systems, as well as informed and sensible AI policy. He holds a PhD in Psychology from the University of Cambridge and has postdoctoral experience from Helmholtz Munich. He formerly served as a cognitive scientist on Geoffrey Irving's Alignment team at UK AISI.