spot_img
HomeResearch & DevelopmentWhen More Thinking Doesn't Mean More Truth: Test-Time Scaling's...

When More Thinking Doesn’t Mean More Truth: Test-Time Scaling’s Limits in AI Knowledge Tasks

TLDR: A study found that increasing “thinking time” (test-time scaling) in reasoning models doesn’t consistently improve factual accuracy or reduce hallucinations in knowledge-intensive tasks. For many models, longer reasoning even led to more hallucinations, often because models attempted more questions they previously abstained from, resulting in incorrect answers. While enabling thinking is generally beneficial compared to no thinking, simply scaling up computation isn’t a reliable way to boost factual robustness.

Recent advancements in artificial intelligence have brought forth powerful reasoning models like GPT-5 and Gemini 2.5, capable of tackling complex problems, from advanced mathematics to intricate logical puzzles. A key technique behind their impressive capabilities is “test-time scaling,” which essentially means allowing these models to generate longer, more elaborate “chains of thought” before arriving at an answer. This approach has shown remarkable success across many different domains.

However, a new research paper from the National University of Singapore, titled “Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet,” challenges the assumption that more thinking always leads to better outcomes, especially for tasks that demand high factual accuracy and minimal errors, known as knowledge-intensive tasks. You can read the full paper here: Research Paper.

The Challenge of Knowledge-Intensive Tasks

Knowledge-intensive tasks are those where models must retrieve and apply factual information correctly, and where making up information (hallucinations) is highly undesirable. Think of answering factual questions, summarizing documents, or providing accurate data. In these scenarios, a model’s ability to reason extensively might seem like an advantage, but this study reveals a more complex picture.

A Comprehensive Evaluation

The researchers, James Xu Zhao, Bryan Hooi, and See-Kiong Ng, conducted an extensive evaluation involving 12 different reasoning models. They tested these models on two knowledge-intensive benchmarks: SimpleQA, which features straightforward factual questions, and FRAMES, which includes more complex questions often requiring multi-step reasoning. The goal was to observe how accuracy and the rate of hallucinations changed as the models were given more “thinking time” or “reasoning effort.”

Surprising Findings on Accuracy and Hallucinations

The results were quite revealing. Contrary to what one might expect, increasing the amount of test-time computation did not consistently improve accuracy across most models. While some models, like GPT-5 mini, showed initial gains, further increases in reasoning length brought little to no additional improvement. In many cases, accuracy simply plateaued or even fluctuated without a clear upward trend.

Even more striking were the findings on hallucinations. For the majority of models, longer reasoning did not reduce hallucinations; in fact, for several models such as GPT-5 mini and Gemini 2.5 Flash, it actually led to more hallucinations. Only a couple of models, Grok-3 mini and DS-R1-Distill-Qwen-14B, showed a reduction in hallucinations with extended reasoning, and even then, the improvements were sometimes modest.

Understanding the Shift in Model Behavior

To understand why these changes occurred, the researchers delved into how model behavior shifted with increased thinking. They found two primary drivers:

  • Fewer Hallucinations from Abstention: When models showed reduced hallucinations, it was largely because they chose to abstain from answering after thinking more, rather than providing a correct answer due to improved factual recall. Essentially, they became more cautious.

  • More Hallucinations from Risky Attempts: Conversely, when hallucinations increased, it was often because extended reasoning encouraged the models to attempt questions they had previously left unanswered. Many of these newly attempted answers turned out to be incorrect, leading to a higher hallucination rate.

Case Studies: Confirmation Bias and Incomplete Reasoning

The study also included fascinating case studies. For instance, with gpt-oss-20b, longer reasoning appeared to induce “confirmation bias.” The model would tentatively propose an answer and then generate fabricated details to support that belief, leading to overconfident, yet incorrect, hallucinations. In the case of Gemini 2.5 Flash, a low thinking budget meant the model couldn’t complete its reasoning process and thus abstained. However, with a higher budget, it would complete its reasoning and sometimes confidently provide an incorrect answer.

The Value of “Thinking” Itself

Despite these limitations of scaling test-time computation, the research highlighted an important distinction: simply enabling a model to “think” (generate reasoning chains) is still beneficial compared to a “non-thinking” mode. For most models, enabling thinking improved accuracy, particularly on more complex tasks like FRAMES, and also reduced hallucinations. Gemini 2.5 Flash was an exception, where enabling thinking led to more hallucinations, likely because it encouraged more attempts at answering.

Also Read:

Conclusion: A Nuanced View of Test-Time Scaling

In summary, while test-time scaling has proven powerful in many areas, this research indicates it is not yet a reliable strategy for enhancing factual accuracy and reducing hallucinations in knowledge-intensive tasks. Simply giving models more time to “think” doesn’t automatically lead to better, more factual answers; it can even make them more prone to confidently hallucinating. The findings suggest a need for more sophisticated control mechanisms to ensure that increased inference time genuinely improves factual robustness in large language models.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -