TLDR: A study found that large language models (LLMs), especially GPT-4o, show high alignment with human ratings of emotional stimuli (words and images). LLMs aligned better with a five-category emotion framework (happiness, anger, sadness, fear, disgust) than a two-dimensional one (arousal and valence), with happiness being most aligned and arousal least. LLM ratings were also more homogenous than human ratings, indicating a close but distinct correspondence in emotional interpretation.
Understanding how artificial intelligence interprets and responds to human emotions is becoming increasingly vital as these powerful tools integrate more deeply into our daily lives. A recent research paper explores this very question, examining how large language models (LLMs) align with human perceptions of emotional stimuli.
The study, titled “Large Language Models are Highly Aligned with Human Ratings of Emotional Stimuli,” was conducted by a team of researchers from the Johns Hopkins University Applied Physics Laboratory: Mattson Ogg, Chace Ashcraft, Ritwik Bose, Raphael Norman-Tenazas, and Michael Wolmetz. Their work sheds light on the surprising similarities and notable differences between human and artificial intelligence in processing emotional cues.
Emotions play a fundamental role in human behavior and decision-making. As LLMs are increasingly used in roles that require interaction with or even acting as human agents, it’s crucial to understand their “emotional intelligence.” This research aimed to build that understanding by comparing LLM responses to human ratings of emotionally charged words and images.
A long-standing debate in psychology revolves around how emotions are best categorized: either as five distinct categories (happiness, anger, sadness, fear, disgust) or along two broader dimensions (arousal and valence). The researchers designed their study to see which framework LLMs align with more closely.
To conduct their experiment, the team utilized several public datasets of images (OASIS, NAPS) and words (ANEW) that had previously been rated for their emotional content by large groups of human participants. They then prompted popular high-performing LLMs, including GPT-4o, GPT-4o-mini, Gemma2-9B, Llama3-8B, and Solar 10.7B, to perform the same rating tasks. Each LLM was treated as a “participant,” with twenty runs for each rating paradigm and dataset to gather comprehensive data.
The findings revealed a significant alignment between large language models, particularly GPT-4o, and human ratings across various emotion scales. The correlations were statistically strong, indicating a close correspondence. Specifically, GPT-4o showed very similar responses to human participants for both text and image stimuli across most rating scales, often achieving correlation coefficients of 0.9 or higher.
Interestingly, the study found that LLMs aligned better with the five-category emotion framework (happiness, anger, sadness, fear, disgust) than with the two-dimensional arousal and valence model. Happiness ratings showed the highest alignment between humans and LLMs, while arousal ratings were the least aligned. This suggests that while LLMs can grasp the “what” of an emotion (e.g., happiness), they might interpret the intensity or physiological activation (arousal) differently than humans.
Another key observation was the homogeneity of LLM ratings compared to human ratings. LLM responses for each item were substantially more uniform, meaning there was less variation in their “opinions” than among human participants. This highlights a difference in how biological and artificial intelligence express their understanding of emotional stimuli.
Also Read:
- AI Agents Explore Human-Like Rhythms: Hormones and Emotions Shape AI Performance and Reveal Biases
- Building More Empathetic AI: A Reinforcement Learning Approach for Emotional Support
In conclusion, this research demonstrates a surprisingly close correspondence between how humans and LLMs interpret emotional stimuli. This alignment is a critical aspect of behavior that influences cognition and interactions. While the study indicates a stronger alignment within a five-category emotion framework, further research is needed to fully understand if this reflects an inherent way LLMs represent emotion or is a byproduct of their training data. For more details, you can read the full research paper here.


