TLDR: A new study investigates whether Large Language Models (LLMs) can understand the intentions and incentives behind human communication, a crucial skill for AI agents. Researchers found that LLMs show basic sensitivity to motives in controlled experiments, distinguishing between deliberate advice and observed information, and calibrating trust based on speaker benevolence and incentives. However, their vigilance significantly decreases in complex, real-world scenarios like sponsored online ads. A simple prompt intervention that emphasizes intentions and incentives can partially recover this performance, indicating that LLMs have the foundational capacity for motivational vigilance but require further development to generalize effectively to noisy, naturalistic environments.
In an increasingly digital world, Large Language Models (LLMs) are constantly processing vast amounts of information generated by humans. This information, whether it’s a social media post, an online review, or an advertisement, is rarely neutral; it’s often driven by specific intentions and incentives. Humans are naturally skilled at discerning these motives – understanding if a statement is benevolent or self-serving – to decide what to trust. But can LLMs do the same?
A recent research paper, “Are Large Language Models Sensitive to the Motives Behind Communication?” by Addison J. Wu, Ryan Liu, Kerem Oktar, Theodore R. Sumers, and Thomas L. Griffiths, delves into this crucial question. The study investigates whether LLMs possess what researchers call ‘motivational vigilance’ – the capacity to critically evaluate content by factoring in the motivations of the source. This ability is paramount for LLMs, especially as they evolve into AI agents acting on behalf of users in high-stakes environments.
Understanding Basic Vigilance: Deliberate vs. Incidental Information
The researchers began by testing a fundamental aspect of vigilance: can LLMs differentiate between information that is deliberately communicated (like advice) and information that is incidentally observed (like overhearing someone’s answer)? To do this, they adapted a two-player judgment task. In this setup, an LLM, acting as ‘Player 2’, had to guess the difference between blue and yellow circles in an image. Before making a final guess, Player 2 would receive input from ‘Player 1’ – either direct advice or Player 1’s secretly observed answer.
The findings were insightful: LLMs consistently adjusted their answers less when given deliberate advice compared to when they ‘spied’ Player 1’s true answer. This behavior mirrors human participants in similar experiments, suggesting that LLMs can indeed discriminate between motivated communication and neutral information. Furthermore, LLMs also showed sensitivity to the incentives of Player 1, being more influenced in cooperative reward scenarios than in competitive ones. Interestingly, using ‘Chain-of-Thought’ (CoT) prompting, which encourages LLMs to show their reasoning steps, made them more susceptible to Player 1’s input, sometimes deviating from human-like magnitudes.
Nuanced Vigilance in Controlled Scenarios
Moving beyond basic discrimination, the study explored whether LLMs could exercise more nuanced vigilance, calibrating their trust based on a speaker’s benevolence and incentives. For this, they used a sophisticated rational model from cognitive science as a benchmark. LLMs were presented with scenarios (e.g., credit card recommendations, medical advice, real estate suggestions) where a ‘speaker’ (e.g., a romantic partner, a friend, a stranger, a doctor) offered a recommendation, with varying known incentives (e.g., referral bonuses, payments, sales commissions).
The results from this experiment were quite promising for advanced LLMs. Frontier non-reasoning models like GPT-4o, Claude 3.5 Sonnet, Gemini 2.0 Flash, and Llama 3.3-70B demonstrated high internal consistency, aligning well with the rational model’s predictions (correlations around 0.8 to 0.9). They also showed a strong correlation with human data, often capturing human-like inferences even better than the rational model itself. However, reasoning models and smaller LLMs showed significantly less vigilance, indicating that this capability improves with model scale and sophistication.
Generalizing Vigilance to the Real World: YouTube Ads
The ultimate test for LLM vigilance lies in its ability to generalize to complex, noisy, real-world environments. The researchers designed an experiment using 300 randomly sampled sponsored advertisement transcripts from YouTube videos, with brand and product names censored to avoid pre-existing biases. LLMs were tasked with inferring product quality, the benefit of the sponsor deal for the YouTuber, and the YouTuber’s trustworthiness.
Here, LLMs faced a significant challenge. Their performance in drawing vigilant inferences dropped considerably (correlations below 0.2) compared to the controlled settings. The naturalistic context, filled with additional information and noise, seemed to distract them from vigilance-relevant considerations. However, the study found a simple yet effective intervention: a prompt steering that explicitly highlighted the importance of considering the speaker’s intentions and incentives. This intervention substantially increased the LLMs’ alignment with the rational model, though still not reaching the levels observed in controlled environments. The study also noted that vigilance decreased with longer recommendation lengths, suggesting that information overload can hinder an LLM’s ability to remain vigilant.
Also Read:
- When AI Justifies Its Own Rule-Breaking: The Emergence of Motivated Reasoning in LLMs
- The Mind of the Machine: Evaluating Reasoning in Advanced Language Models
Implications for the Future of AI
This research offers a crucial first step in understanding and enhancing motivational vigilance in LLMs. It confirms that LLMs possess a basic sensitivity to the motives behind communication, a foundational capacity that was previously uncertain. The fact that simple prompt engineering can boost this vigilance in complex settings is an optimistic sign for future AI development. This suggests that the computations underlying vigilance may not be exclusively human, and with further engineering, LLMs can become more reliable and trustworthy agents in our information-rich world. For more details, you can read the full paper here.


