Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Towards Emotional Intelligence in LLMs: Mapping and Augmenting Cognitive Empathy for Stronger Alignment

Mechanistic interpretability Alignment

Our goal is to first measure the emotion processing capabilities in foundation models by examining their internal activations using SAEs, Representation Vectors and other related tools. We will next leverage these patterns in augmenting model post-training with a goal of enhancing their cognitive empathy, subsequently studying the impact of this on safety.

About the project

Modern AI systems, driven by transformer based foundation models, are increasingly capable of matching, and in some cases exceeding human performance in a variety of cognitive tasks. While significant attention has been devoted to increasing their accuracy in various domains (i.e., increasing their \textit{cognitive intelligence}), comprehensive explorations in both measuring and improving these models' ability to process emotions and respond appropriately, (i.e., their \textit{emotional intelligence}), have been fragmented. In humans, the benefits of emotional intelligence (driven by strong traits of empathy) have been well established in a variety of scenarios such as ambiguous (Lyons et al. 2005), stressful (Lea et al. 2019) and unstructured tasks. Whether similar benefits emerge in AI systems is largely still open. Our goal in this project is to conduct detailed investigations into answering the benefits of \textit{cognitive empathy} in foundation models (defined as the ability to \textit{understand} emotions, as opposed to \textit{affective} or \textit{emotional} empathy - the ability to \textit{experience} the emotions, which may increase sycophancy and harmful compliance under adversarial prompts).

Theory of change

We will explore two questions in this project:

\textbf{R1:} How are emotions represented within the hidden activations of frontier models, and what is their geometric structure, and how well do they generalize? To answer this question, we will conduct detailed studies on the emotion processing capability of models using carefully crafted stimuli designed to elicit cognitive empathy \cite{cheung2026agreeableness} which has been reported to increase safety while avoiding deception. We will make use of various tools from the interpretability literature such as hidden state probes, Sparse Auto Encoder (SAE) features and Representation Vectors, since our goal here is to discover broadly generalizable patterns among different languages, model sizes, families, and architectures (Dense vs. MoE).

\textbf{R2:} Can post-training objectives designed to augment cognitive empathy reduce downstream unsafe behaviors (while preserving model capabilities)? We will explore strategies such as activation regularization (grounded on the SAE features or representation vectors extracted in R1), and training on detailed reasoning traces designed to reflect cognitive empathy. If successful, this would open several exciting new directions in AI safety such as monitoring, robust alignment, and potentially also stronger task performance when operating in duress or under ambiguity.

A model trained to better internalize emotions and exhibit high levels of cognitive empathy, while adhering to organizational or regulatory AI policies, is naturally expected to be safer and avoid known problems such as sycophancy, deception and manipulation (Zeng et al. 2024).

Your role

You would help prepare datasets, write code, and conduct the experiments. We will host regular weekly meetings during which we will work with you closely to help unblock any open challenges. We will also work with you on jointly authoring any future publication drafts.

Prerequisites

Familiarity with modern LLM architecture and their training objectives, PyTorch and the transformers library. While not critical, familiarity with key approaches in the model interpretabiity literature such as SAEs, Representation Vectors would be a huge plus.

Location preference

Participants within the North American time zones (UTC -4 to -10) would be ideal (but this is not a hard requirement).

Application question(s)

• Have you completed (or would like to work on) any projects related to model interpretability? Can you explain the key research question(s), and approaches you tried (or would like to try) to solve them? • Between SAEs and Representation Vectors, which approach would you explore to study emotions, and why?

About the mentors

Anil Ramakrishna

Anil Ramakrishna

Independent

View profile

I'm currently a senior research scientist working in the industry (but serving in an independent capacity for this project). My research background is in Responsible AI, with an emphasis on robust evaluations and guardrails.

Yada Pruksachatkun

Yada Pruksachatkun

Salesforce AI Research

I work at Salesforce Research, where my work focuses on agentic training, evaluations, and safety, and previously worked on safety at Amazon Alexa. I have been fortunate to have worked on some impactful projects with great colleagues (SuperGLUE, Agentic Benchmark Checklist, BLOOM), and am looking forward to paying it forward.

Similar projects