spot_img
HomeResearch & DevelopmentLarge Language Models Show Human-Like Logic Induction, Challenging Cognitive...

Large Language Models Show Human-Like Logic Induction, Challenging Cognitive Theories

TLDR: A new study reveals that Large Language Models (LLMs) can learn and apply logical rules in a way that closely matches human behavior, outperforming traditional symbolic models. The research demonstrates that LLMs not only achieve human-level accuracy in rule induction tasks but also exhibit similar learning trajectories and can articulate the logical concepts they infer. This challenges long-standing assumptions about the unique architecture required for human-like logical thought and suggests LLMs could offer novel insights into the nature of human reasoning.

A long-standing debate in cognitive science revolves around whether artificial neural networks (ANNs) can adequately model complex human cognitive functions, particularly those related to language and logic. This discussion has been reignited by the emergence of large language models (LLMs), which represent a significant leap in the capabilities of neural networks.

A recent research paper explores this very question by testing various LLMs on an established experimental setup designed to study how humans induce rules based on logical concepts. The findings suggest that LLMs can fit human behavior in this task as well as, or even better than, traditional computational models that rely on a probabilistic Language of Thought (pLoT).

LLMs Match Human Rule Learning Accuracy

The study’s first experiment investigated whether off-the-shelf LLMs could succeed in a logical rule-learning task, where models had to infer a hidden rule from examples and then apply it to new objects. The objects were described by features like size, color, and shape, and the rules varied in complexity, from simple propositional logic (e.g., “blue or small”) to more complex first-order logic (FOL) rules involving relationships between objects (e.g., “the only medium-sized object in its set”).

Remarkably, some LLMs, including GPT4, Mixtral 8x7b Instruct, and Gemma (7B), achieved accuracy rates comparable to human participants. For propositional rules, these models often surpassed the lower bound of human performance. While GPT4’s performance was notably strong, the authors suggest that differences in prompting or recent model updates might explain why their results differed from previous studies that reported lower GPT4 accuracy on the same task.

Understanding How LLMs Infer Rules

To delve deeper into how LLMs arrive at their classifications, a second experiment focused on GPT4’s ability to articulate the rules it inferred. GPT4 was prompted to describe the rule it had learned and then classify new objects. The results showed a high consistency (96.3%) between GPT4’s self-reported rules and its actual classification behavior, suggesting that the model genuinely uses the rules it describes.

However, a qualitative difference emerged: LLMs tended to formulate more verbose rules, especially for propositional logic, often concatenating many features. For rules requiring first-order logic, LLMs struggled to use quantifiers (like “all” or “some”), instead attempting to approximate these complex rules using long chains of propositional operators. This raises an interesting question about human cognition: do humans also approximate complex rules with simpler operators, or do they truly employ more complex logical primitives like “xor” or quantifiers?

LLMs Mirror Human Learning Trajectories

Perhaps the most compelling finding came from experiments focusing on the correlation between LLM and human learning trajectories. By fine-tuning an LLM (Gemma 7B) on human response data, the researchers found that the tuned LLM explained a significantly higher percentage of variance in human responses (84.8%) than existing Bayesian pLoT models. This means the LLM’s learning curve, including its successes and mistakes, closely mirrored that of human learners, object by object.

This strong correlation suggests that LLMs, when appropriately tuned, might be implementing similar inference procedures or navigating a similar hypothesis space as humans. The generalizability of this tuning was also tested, showing that the LLM could transfer its learning to unseen rules, provided those rules shared the same abstract logical components. This implies that the tuning process helps LLMs align their internal representations or priors with human-like concepts, rather than just learning surface-level statistical heuristics.

Also Read:

Implications for Cognitive Science

The paper concludes that LLMs are empirically adequate models of logical thought, challenging the long-held view that human higher cognition necessitates a fundamentally different architecture than ANNs. The ability of LLMs to learn and deploy human-like logical concepts, compose structured thoughts, and revise hypotheses based on evidence suggests they might possess a form of “Language of Thought.”

Furthermore, because LLMs fit human learning patterns better than symbolic models, they may offer new theoretical insights into the nature of human logic, potentially revealing differences between formal logic and how humans actually reason. Future research could use mechanistic interpretability to explore the internal workings of LLMs, providing new hypotheses about how neural circuits might implement logical concepts in the human brain. This research opens up exciting new avenues for understanding the computational underpinnings of human cognition. For more details, you can refer to the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -