spot_img
HomeResearch & DevelopmentUnderstanding How LLMs Reason: The Dance Between Learning from...

Understanding How LLMs Reason: The Dance Between Learning from Examples and Prior Knowledge

TLDR: This research paper investigates the underlying mechanisms of Chain-of-Thought (CoT) reasoning in Large Language Models (LLMs), focusing on the interaction between In-Context Learning (ICL) and pretrained priors. It reveals that LLMs learn both lexical structures and deeper logical patterns from CoT exemplars, yet heavily rely on pretrained knowledge. The study demonstrates that sufficient exemplars can shift decision-making from pretrained priors to ICL signals, but misleading prompts introduce instability. Finally, it shows that prompt engineering can induce LLMs to generate longer, more reflective reasoning chains (“slow thinking”), thereby improving performance on various tasks.

Large Language Models (LLMs) have revolutionized how we interact with AI, especially with the advent of Chain-of-Thought (CoT) reasoning. CoT allows LLMs to break down complex problems into intermediate steps, significantly boosting their performance across various tasks. However, the exact mechanisms behind CoT’s effectiveness have remained somewhat of a mystery. A recent research paper, “Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pretrained Priors”, delves into this mystery, exploring the interplay between In-Context Learning (ICL) and the knowledge LLMs acquire during their initial training, known as pretrained priors.

Unpacking How LLMs Learn from Examples

The study, conducted by Hao Yang, Zhiyu Yang, Yunjie Zhang, Shanyi Zhu, and Lin Yang, first investigates what LLMs truly learn when presented with CoT examples through ICL. The researchers performed a detailed analysis of the language generated by models, breaking it down into components like structural words (e.g., “Therefore,” “Then”), feature words (task-specific elements like numbers or operators), verbs (reasoning actions), and location/person entities. They found that LLMs quickly pick up on the reasoning structure at a lexical level, mimicking the language patterns from the examples. Interestingly, even when given examples from a completely different task (task-agnostic CoT), the models still tended to perform reasoning relevant to the original question, indicating a strong reliance on their pretrained knowledge. The research also highlighted a positive correlation between the number of reasoning verbs generated and accuracy, up to an optimal point, suggesting that models grasp deeper logical reasoning patterns beyond just surface-level imitation.

Can LLMs Overcome Their Existing Knowledge?

A crucial question addressed by the paper is whether LLMs can override their ingrained pretrained priors when presented with new, potentially misleading, information through ICL. Previous studies suggested that incorrect reasoning had little impact on LLM performance, especially for smaller models. However, this research challenges that notion. By introducing “false-answer” (correct rationale, incorrect answer) and “false-rationale” (incorrect reasoning steps, correct answer) CoT prompts, the team observed how model accuracy and confidence changed with an increasing number of examples.

Initially, with a small number of examples, the impact of misleading information was minimal, as pretrained priors dominated. But as the number of examples grew, ICL signals became stronger, shifting the model’s decision-making. In tasks with a limited set of possible answers (closed-domain tasks like Coin Flip), false-answer prompts led to systematic label flipping, where the model consistently gave the wrong answer. For tasks with a wider range of answers (open-domain tasks like math problems), accuracy gradually declined but significantly. False-rationale prompts also severely degraded reasoning ability in open-domain tasks with more examples. The study also found that misleading prompts caused greater fluctuations in the model’s confidence, indicating that the model detected the inconsistencies, leading to instability. This underscores the critical importance of high-quality examples for stable and reliable reasoning.

Also Read:

Prompting LLMs for “Slow Thinking”

Finally, the paper explores whether prompt engineering can encourage LLMs to engage in “slow thinking” – a process where models generate longer, more reflective reasoning chains, similar to how humans might deliberate on a problem. By distilling long CoT prompts from advanced Reasoning Language Models (RLMs) and using them to guide other LLMs, the researchers found that models could indeed emulate this slow thinking approach. This led to improved performance on various downstream tasks, including arithmetic, commonsense, and symbolic reasoning. The findings suggest that LLMs can leverage both their ICL capabilities and pretrained knowledge to generate extended rationales, hinting at a path towards model self-evolution. The research also noted that there’s an optimal length for CoT reasoning, which varies depending on the model’s capacity and the complexity of the task.

In conclusion, this study provides valuable insights into the intricate relationship between in-context learning and pretrained priors in Chain-of-Thought reasoning. It highlights that LLMs learn both the surface-level structure and deeper logical patterns from examples, but their pretrained knowledge remains a significant influence. While LLMs can adapt to new information through ICL, the quality and quantity of examples are paramount, as misleading prompts can introduce instability. Furthermore, strategic prompt engineering can guide LLMs to engage in more deliberate, “slow thinking” processes, ultimately enhancing their reasoning capabilities.

Ananya Rao
Ananya Raohttps://blogs.edgentiq.com
Ananya Rao is a tech journalist with a passion for dissecting the fast-moving world of Generative AI. With a background in computer science and a sharp editorial eye, she connects the dots between policy, innovation, and business. Ananya excels in real-time reporting and specializes in uncovering how startups and enterprises in India are navigating the GenAI boom. She brings urgency and clarity to every breaking news piece she writes. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -