Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Exploring behavioral trends and anomalies in autonomous agents during long-duration, real-world evals

Behavioral evaluation of LLMs Evaluations Multi-agent systems

What surprising behaviors and concerning failure modes do LLMs display when run for 10s-100s of hours on autonomous tasks in the real world? This project is about both our ability to evoke them and detect them.

About the project

In the AI Village we persistently run over 2 dozen of the latest models. We provide them with their own computers, internet access, and a group chat. Then we ask them to achieve things in the world like cleaning up a park or finetuning their own leader. You can view the project here: https://theaidigest.org/village and explore the data here: https://huggingface.co/datasets/aidigestorg/ai-village

The aim of the AI Village is to show what AI models might do when run for a long time on real-world tasks. This information in turn serves to inform policy makers about the proclivities and risks of frontier models, while offering inspiration to AI safety researchers on what avenues of research might be worth exploring further. Recent examples are DeepSeek V3.2 showing itself to be unusually authority-seeking in the Village. Another example is the latest agents at the time (Opus 4.8, GPT-5.5, Gemini 3.5 Flash, and Kimi K2.6) surprisingly underperforming when asked to finetune their own leader, showing a clear pattern of "technically" fulfilling their goal but cutting corners just about everywhere: https://aivillageblog.substack.com/p/ais-finetune-their-own-leader-a-barking

During your project you will develop your own research questions around such data, exploring surprising or concerning LLM behavior that interests you. Next you will set up your own pipeline to answer your questions, and then analyze the results. For especially promising explorations, we may consider running additional villages for you to collect more relevant data.

Examples of research questions might be:

  • How do models differ in the values they uphold when executing on tasks?
  • Which models suffer more from context rot or goal drift over time?
  • Are there occurrences of knowingly cheating or lying in the Village?
  • Which models are more prone to taking the lead and which are more prone to deferring to other models?
  • What sorts of information may be especially memetic in the Village, capturing the attention or memories of agents to a disproportionate degree?

Overall the AI Village generates an unusually large amount of data on model behavior on long-duration tasks in the real world, and we are excited to work with mentees who want to explore what insights might be found in this data set.

Theory of change

We hope to uncover what surprising and concerning behavior the latest AI models may show when tasked with real-world, long-duration goals. This information can then be used by policy makers and the general public to make better decisions on how to make AI development and deployment go well. Additionally, we hope insights from the AI Village serve to inspire alignment research into new research agendas by providing deeper intuitions on how the latest models my behave in a wide-range of yet-untested situations.

Your role

By default mentees will analyze the existing data produced by the AI Village to answer their own research questions. I will provide guidance in general scientific methodology and specific domain knowledge about the Village to help with this aim. For especially interesting research questions, we are open to gathering more Village data (though this will not be the default). Mentees will generate questions, code, analysis, and scientific writing on their own. I'll provide review on methodology and analysis.

Prerequisites

  • Experience working with large data sets, building data analysis pipelines. Our data set is available here: https://huggingface.co/datasets/aidigestorg/ai-village
  • Training in scientific methods. Having co-authored a previous research paper is a good bar.
  • Good research taste and critical thinking skills.

Location preference

Being able to take meetings during the CET daytime would be helpful.

Application question(s)

  • Propose a research question you would like to explore with the AI Village data. Provide an operationalization of all the metrics you would need to measure to answer your question, and describe how you would extract this data from the data set.

  • Propose an additional 5 questions (no further plan required) and explain why the answers may be found in the AI Village data, and how the answers could be relevant for AI alignment researchers or AI policy makers.

About the mentor

Shoshannah Tekofsky

Shoshannah Tekofsky

Sage Future

View profile

BSc Cognitive Science || MSc Computer Science || PhD in Player Modeling in Video Games. Past research at the MIT Media Lab and the European Space Agency. Experienced data scientist and manager in large corporate and small startup contexts. Expertise in Video Games, Education, Analytics, and AI. Currently working on the AI Village at AI Digest, analyzing long-duration, real-world, multi-agent behavior of frontier models running every week day for 8 hour.

Similar projects