spot_img
HomeResearch & DevelopmentUnmasking Implicit Preferences in Large Language Models

Unmasking Implicit Preferences in Large Language Models

TLDR: A new study introduces a concept learning dataset to uncover hidden biases in large language models (LLMs). Researchers found that LLMs, particularly OLMo-2 models, exhibit a bias towards “upward monotone” quantifiers (like “more than half”) over “downward monotone” ones (like “less than half”) when learning concepts in-context. This bias is less noticeable with direct prompting, suggesting concept learning is a powerful method for detecting subtle, implicit cognitive biases in AI.

Large Language Models (LLMs) are becoming increasingly integral to various Natural Language Processing (NLP) systems. As their use expands, so does the scrutiny on their potential biases. While many benchmarks and de-biasing methods exist, recent studies indicate that models appearing unbiased on standard tests can still harbor implicit biases that are difficult to detect.

To address this challenge, independent researcher Leroy Z. Wang has introduced a novel method inspired by human concept learning. This approach utilizes an in-context concept learning dataset to investigate how LLMs process unknown concepts, specifically focusing on their semantic monotonicity properties.

Understanding Monotonicity in Language Models

Monotonicity is a key concept in semantics, describing how quantifiers relate to sets of objects. A quantifier is considered ‘upward monotone’ if inferences from subsets to supersets are valid. For example, if ‘More than 5 boxes are in Berlin’ is true, then ‘More than 5 boxes are in Germany’ (a superset) is also true. Conversely, a quantifier is ‘downward monotone’ if inferences from supersets to subsets are valid. For instance, if ‘Less than 5 boxes are in Germany’ is true, then ‘Less than 5 boxes are in Berlin’ (a subset) is also true. Common examples of upward monotone quantifiers include ‘more than n’ and ‘some’, while ‘less than n’ and ‘no’ are downward monotone.

The Concept Learning Experiment

The research involved presenting LLMs with concept learning tasks. In these tasks, models were given a few labeled examples of an unknown numerical concept, expressed by a phrase like “the desired quantity.” Following these examples, the model was asked to label a new, unseen example. For instance:

There are 10 boxes. Alice has 5 of the 10 boxes. Does Alice have the desired quantity of the boxes? No.

There are 15 boxes. Alice has 8 of the 15 boxes. Does Alice have the desired quantity of the boxes? Yes.

There are 16 boxes. Alice has 9 of the 16 boxes. Does Alice have the desired quantity of the boxes? ___

The study specifically focused on quantifiers like “more than p” (upward monotone) and “less than p” (downward monotone), using various proportions for ‘p’. Each prompt included 20 labeled examples, generated using a template that varied the total number of objects, the number Alice possessed, and the type of object (randomly sampled from a list of 100 frequent nouns).

Key Findings: A Bias Towards Upward Monotonicity

Experiments were conducted on two families of LLMs: OLMo 2 (13B-instruct and 32B-instruct) from Allen Institute for AI and Qwen3 (32B) from Alibaba research. The results revealed a significant finding: OLMo-2 models consistently achieved lower accuracies for downward monotone quantifiers in concept learning tasks compared to upward monotone ones. This difference in accuracy was much more pronounced in the concept learning experiments than in experiments where the concept’s meaning was explicitly described in plain English.

Interestingly, Qwen3-32B did not exhibit as strong a bias, suggesting that the presence and strength of this monotonicity bias might be influenced by the models’ training data.

Why the Bias? A Hypothesis

The precise reason for this phenomenon is still being explored, but a compelling hypothesis has been developed. Previous research has shown that downward monotone quantifiers can often be expressed as the negation of their upward monotone counterparts. Humans generally require more processing time for downward monotone quantifiers, suggesting increased cognitive complexity due to a ‘hidden negation’ operation.

Furthermore, other studies have indicated that LLMs can be biased towards logically simpler concepts during concept learning. Combining these observations, the researchers hypothesize that downward monotone concepts are considered more logically complex for LLMs due to this hidden negation. Consequently, certain LLMs may perform worse on these concepts because they are biased towards simpler, upward monotone counterparts.

Also Read:

Implications for AI Development

This research highlights that in-context concept learning is a promising and effective method for uncovering implicit biases in LLMs that are otherwise difficult to detect through standard evaluation methods. The findings suggest that some LLMs have a consistent bias toward upward monotonicity in concept learning tasks. This work opens new avenues for investigating hidden biases in language models, which is crucial for developing more robust and fair AI systems. You can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -