spot_img
HomeResearch & DevelopmentWhen LLMs Invent: The Challenge of Knowledge Overriding Evidence...

When LLMs Invent: The Challenge of Knowledge Overriding Evidence in Process Modeling

TLDR: A new study reveals that Large Language Models (LLMs) frequently suffer from ‘knowledge-driven hallucination,’ where their pre-trained general knowledge overrides explicit, conflicting evidence provided in prompts. This phenomenon, investigated through process modeling tasks, shows that LLMs often revert to standard process structures even when given atypical inputs. The issue persists across various LLMs and input types (text and event logs), and while strict prompting can help, it doesn’t fully eliminate the problem. This raises significant concerns about the reliability of AI-generated outputs in evidence-based domains, as models can produce plausible but factually incorrect information.

Large Language Models (LLMs) are powerful tools, but a recent study sheds light on a critical flaw: ‘knowledge-driven hallucination.’ This phenomenon occurs when an LLM’s vast, pre-trained general knowledge overrides explicit information given in a user’s prompt, leading to outputs that appear plausible but are factually incorrect. This is a significant concern, especially as LLMs are increasingly used for complex analytical tasks that require precise adherence to provided evidence.

The research, conducted by Humam Kourani, Anton Antonov, Alessandro Berti, and Wil M.P. van der Aalst, investigated this issue within the domain of Business Process Management (BPM). Process modeling, which involves generating formal business process models from source documents, was an ideal context for this study. Many core business processes follow standardized patterns, meaning LLMs likely possess strong internal schemas for how these processes ‘should’ operate. The central question was: what happens when specific process evidence directly contradicts the LLM’s generalized, common-sense understanding?

To explore this, the researchers designed a controlled experiment using four diverse business processes. For each process, they had a standard model, a natural language description, and a simulated event log. They then created two deliberately conflicting versions: a ‘reversed’ model, where the sequence of activities was causally inverted, and a ‘shuffled’ model, where activity labels were randomly permuted while preserving the control-flow structure. This created six distinct input scenarios for the LLMs, based on both text descriptions and event logs.

Ten state-of-the-art LLMs, including models like gemini-2.5-flash, gemini-2.5-pro, gpt-4.1-nano, and grok-3-fast, were tasked with generating process models from these inputs. The experiments were conducted under two conditions: a ‘Standard Prompt’ and a ‘Strict Adherence Prompt,’ which explicitly instructed the LLM to disregard its background knowledge and rely solely on the provided input.

The findings strongly supported the hypothesis that LLMs exhibit a significant tendency for knowledge-driven hallucination when faced with atypical process structures. In numerous instances, models generated from reversed or shuffled artifacts were substantially more similar to the standard process model than to their actual source evidence. This ‘knowledge-driven hallucination’ was observed across all tested LLMs, with no model achieving full adherence to atypical evidence. Even when LLMs correctly followed atypical structures, the quality of the generated models was generally lower than for standard processes, indicating a struggle to reconcile conflicting information with their pre-trained schemas.

The study also examined the impact of prompting and input types. The ‘Strict Adherence Prompt’ helped mitigate the issue, reducing the number of hallucinations, but it did not eliminate them entirely. This suggests that while prompt engineering is a helpful strategy, the LLM’s deeply ingrained background knowledge is very powerful. Interestingly, fewer hallucinations occurred when models were generated from event logs compared to textual descriptions, likely because the structured and unambiguous format of an event log provides stronger evidence than natural language.

Regarding the LLMs’ inherent properties, the analysis revealed no direct relationship between an LLM’s size (parameter count) and its ability to adhere to atypical evidence. Similarly, general reasoning capabilities did not guarantee immunity to this type of hallucination. A particularly revealing finding was that high performance on standard tasks does not guarantee robustness against knowledge hallucination. For instance, gemini-2.5-pro, a top performer on standard tasks, showed a significant drop in performance when faced with conflicting inputs, often reverting to its dominant internal schema for the standard process.

Also Read:

In conclusion, this research highlights a critical reliability concern for LLMs in any evidence-based domain. The danger of knowledge-driven hallucination lies in its deceptive nature: the generated artifacts often appear coherent, logical, and well-formed, masking the fact that they do not accurately represent the source data. This ‘plausibility trap’ poses a significant risk in fields requiring strict adherence to evidence, such as legal analysis, financial reporting, and scientific research. The authors emphasize the need for more robust mitigation techniques beyond simple prompting and rigorous validation of AI-generated artifacts. You can read the full paper here: Knowledge-Driven Hallucination in Large Language Models: An Empirical Study on Process Modeling.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -