spot_img
HomeResearch & DevelopmentAI's Struggle with Novelty: When Missing Rules Expose Reasoning...

AI’s Struggle with Novelty: When Missing Rules Expose Reasoning Gaps

TLDR: This research investigates AI’s analogical reasoning capabilities using Raven’s Progressive Matrices (RPM), specifically when models are trained with deliberately omitted structural rules. The study found that while AI models perform well on familiar rules, their accuracy sharply declines when faced with novel or omitted rules, indicating a reliance on pattern memorization over genuine abstract reasoning. The discrepancy between high token-level accuracy and low complete answer accuracy further highlights these limitations, suggesting current AI systems exhibit ‘fragile correctness’ and struggle to generalize beyond their training data.

Analogical reasoning, a cornerstone of human intelligence, remains a significant hurdle for artificial intelligence. A recent study delves into this challenge by evaluating modern AI systems on Raven’s Progressive Matrices (RPM) tasks, specifically under conditions where certain structural rules were deliberately omitted during training. The findings shed light on whether AI models truly reason or merely rely on statistical shortcuts.

Raven’s Progressive Matrices are non-verbal tests designed to measure fluid intelligence by requiring individuals to infer missing components in complex visual patterns. This task demands the identification of underlying relational structures, making it a crucial benchmark for both human and AI abstract reasoning. While many deep learning models have shown success on RPM tasks, questions persist about their ability to generalize when faced with incomplete training data.

This research specifically investigated how well models generalize when one or two reasoning rules are removed from their training set, and then tested on tasks involving these missing rules. The goal was to understand if current machine learning models exhibit genuine analogical reasoning or if they are primarily leveraging statistical correlations within the dataset.

The study utilized the Impartial-RAVEN (I-RAVEN) dataset, a refined version of the traditional RA VEN dataset, which addresses issues like annotation biases and spurious correlations that could allow models to achieve high performance without true reasoning. I-RAVEN ensures that models must infer true structural relationships by balancing distractors in the answer sets.

The I-RAVEN dataset is built around five rule-governing attributes: Number, Position, Type, Size, and Color. Each attribute can follow one of four rules: Constant, Progression, Arithmetic, and Distribute Three. To create incomplete training scenarios, two modified versions of the dataset were used: one where all instances of the Progression rule were removed, and another where both Progression and Arithmetic rules were omitted.

Three types of models were evaluated: a sequence-to-sequence transformer model, CoPINet (a vision-based model for perceptual inference), and the Dual-Contrast Network (another vision-based model). For the transformer, visual data was converted into a structured text format describing shape attributes like Type, Size, Color, and Angle, allowing it to process RPM tasks as a sequence prediction problem.

The experimental results revealed a consistent performance gap between scenarios where testing rules were familiar from training (“Same Rule” setting) and those where they were novel or omitted (“Different Rule” setting). The sequence-to-sequence transformer, for instance, showed strong performance on familiar rules but a sharp decline when encountering novel ones. This pattern was observed across both one-rule and two-rule removal scenarios, with performance differences often exceeding 15 percentage points.

Task complexity also played a significant role. Removing one rule during training yielded substantially better results than removing two rules, especially in the “Different Rule” conditions. The average accuracy for the sequence-to-sequence transformer dropped from 47.00% to 31.47% when moving from one-rule to two-rule removal in these challenging scenarios.

Interestingly, the study highlighted an “illusion of high token-level accuracy.” While models might achieve high scores in predicting individual tokens (e.g., 96.88% token accuracy in some cases), this often did not translate to correctly selecting the final answer (e.g., only 45.31% correct choice accuracy). This discrepancy arises because even a single incorrect token can lead to choosing an entirely wrong answer, indicating a “fragile correctness” where models solve most of the problem but fail at the final, crucial step.

A counterintuitive finding was that vision models sometimes performed better in the “Remove 2 Same” rule condition compared to “Remove 1 Same.” This suggests that removing more rules during training might have simplified the learning environment by reducing the complexity and variability of the remaining rule combinations, particularly benefiting vision-based models which might be more sensitive to pattern complexity.

Also Read:

In conclusion, the study underscores a fundamental limitation in current AI systems’ ability to generalize abstract rules to novel configurations. Models often rely on pattern memorization rather than developing genuine abstract reasoning skills. The sequence-to-sequence transformer generally outperformed vision models, but its performance decline in unseen-rule scenarios indicates its reasoning is fragile and heavily tied to training distributions. These findings emphasize the need for future AI architectures that integrate symbolic structure with neural representations and incorporate stronger inductive biases for relational reasoning, moving beyond mere pattern recognition. You can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -