TLDR: A new research paper introduces “anthropomimetic uncertainty,” arguing that large language models (LLMs) need to emulate human-like communication of uncertainty to build user trust. It highlights how current LLMs are often overconfident due to data biases and inconsistent in expressing doubt across contexts. The paper proposes future research directions focusing on consistency, personalization, appropriate use of verbal vs. numerical expressions, multilinguality, and explaining the reasons behind uncertainty to foster better human-AI collaboration.
In an era where large language models (LLMs) are becoming increasingly integrated into our daily lives, from assisting with tasks to providing information, a critical issue has emerged: their tendency to present information with unwavering confidence, even when it’s inaccurate. This overconfidence undermines user trust and the potential benefits of human-machine collaboration. A new research paper, titled “Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing,” delves into this problem and proposes a novel approach: anthropomimetic uncertainty.
The paper, authored by Dennis Ulmer, Alexandra Lorson, Ivan Titov, and Christian Hardmeier, argues that for LLMs to be truly trustworthy and legitimate, they need to communicate their confidence levels in a way that mirrors human communication. This concept, “anthropomimetic uncertainty,” suggests that LLMs should emulate the linguistic authenticity and personalization humans employ when expressing doubt or certainty.
Understanding Human Uncertainty
Humans naturally express uncertainty through both verbal cues (like “maybe,” “might,” “certain”) and numerical expressions (like “40% chance”). This communication isn’t just about a lack of knowledge; it’s also influenced by social factors such as politeness, the relationship between speakers, and the perceived severity of consequences. For instance, people might hedge their statements more when speaking to someone of higher social status or when discussing serious matters. While verbal expressions are often imprecise, they allow for nuance, and numerical expressions, though seemingly precise, can also be approximations.
A key human cognitive mechanism is “epistemic vigilance,” where listeners assess the credibility of information based on the source’s reliability and the content’s plausibility. This vigilance is less effective with LLMs because their internal reasoning processes are opaque, making it hard for users to discern truth from plausible-sounding falsehoods.
The Current State of LLM Uncertainty
Current research in natural language processing (NLP) has explored “verbalized uncertainty,” where LLMs are prompted or finetuned to express confidence. However, the paper highlights that most methods still rely on numerical expressions (e.g., 0-1 or 1-100 scales), which are easier to evaluate but often lack the natural fluency of human language. Only a small fraction of research focuses on integrating epistemic markers naturally into responses. Furthermore, simple prompting methods for verbalizing uncertainty often lead to uncalibrated confidence estimates.
Biases and Contextual Challenges
The researchers demonstrate that LLMs’ uncertainty communication is heavily influenced by biases in their training data. Pretraining data, instruction finetuning datasets, and even reinforcement learning from human feedback (RLHF) can lead models to be overconfident. For example, reward models in RLHF often favor confident statements, even if incorrect, because they appear more “helpful” to human annotators. This inadvertently trains LLMs to downplay their uncertainty.
Moreover, LLMs show inconsistent behavior across different conversational contexts and subject areas. Unlike humans who maintain a consistent style of expressing confidence, LLMs’ use of uncertainty expressions can vary drastically depending on whether they are acting as an “employee to a boss” or a “scientist at a conference.” This inconsistency further erodes trust.
Also Read:
- The Art of Asking: How Large Language Models Learn to Clarify in Dialogue
- Unmasking True LLM Performance: A Critical Look at Evaluation Methods
Towards Anthropomimetic Uncertainty
The paper proposes several actionable research directions to achieve anthropomimetic uncertainty, aiming to build trust by making LLM uncertainty communication more human-like:
- Consistency: LLMs should use uncertainty expressions consistently across different conversations and contexts, reflecting a stable understanding of their own confidence.
- Personalization: Uncertainty communication should adapt to individual users, considering their risk tolerance, prior knowledge, and past interactions.
- Choice of Register: Researchers need to determine when verbal or numerical expressions are most appropriate, acknowledging that verbal expressions, despite their imprecision, can be more intuitive and reflect the inherent imprecision of some uncertainties.
- Multilinguality: The focus on English is a significant shortcoming. Uncertainty expression varies greatly across languages, and LLMs need to account for these linguistic and cultural nuances.
- Explaining Uncertainty: Beyond just stating uncertainty, LLMs should be able to explain *why* they are uncertain, for instance, due to ambiguous input or limited knowledge.
- Beyond Classical Calibration: The focus should shift from purely probabilistic correctness to how uncertainty communication influences user outcomes and trust.
- Trust beyond Uncertainty: Acknowledge that uncertainty expression is just one factor in building user trust, alongside warmth, humanlikeness, and prior interactions.
In conclusion, the research paper highlights a crucial gap in current LLM development: the nuanced and context-dependent communication of uncertainty. By advocating for “anthropomimetic uncertainty,” the authors pave the way for more reliable, trustworthy, and effective human-AI collaboration. For more detailed insights, you can read the full paper here.


