spot_img
HomeResearch & DevelopmentExploring the Stability of Knowledge Extraction from Large Language...

Exploring the Stability of Knowledge Extraction from Large Language Models

TLDR: This research investigates how reliably factual knowledge can be extracted from Large Language Models (LLMs) into structured knowledge bases. Using “miniGPTKBs” (domain-specific knowledge crawls), the study examines three key aspects: termination (whether extraction finishes), reproducibility (consistency across runs), and robustness (stability against variations). Findings show high termination rates for most settings, mixed reproducibility (consistent output size but variable exact wording), and varying robustness (high for seed/temperature, lower for language/model). The paper concludes that LLM knowledge materialization can reliably surface core knowledge, especially with ensembling, but highlights limitations related to language and model choice.

Large Language Models, or LLMs, are known to hold a vast amount of factual knowledge within their complex structures. However, understanding, measuring, and systematically extracting this knowledge into a structured format, like a knowledge base, has been a significant challenge. A method called GPTKB, which involves recursive extraction, has shown promise but also raised fundamental questions about its reliability.

This research paper, titled “Foundations of LLM Knowledge Materialization: Termination, Reproducibility, Robustness,” delves into these critical questions. It explores whether such knowledge extraction processes can reliably finish (terminate), produce consistent results across multiple attempts (reproducibility), and remain stable when faced with different conditions or choices (robustness).

The Core Questions

The study focuses on three main research questions:

  • Termination: Can the LLM knowledge extraction process reach a natural end, or is it prone to generating endless, often fabricated, information (hallucination)?
  • Reproducibility: Given that LLMs can be non-deterministic, how consistent are the knowledge bases created from them when the process is repeated with the same settings?
  • Robustness: How well does the extracted knowledge hold up when certain design choices are varied, such as the starting point, the language used, the model’s randomness, or even the LLM itself?

Introducing MiniGPTKBs

To tackle these questions in a manageable way, the researchers introduced “miniGPTKBs.” These are smaller, domain-specific knowledge crawls, making the extraction process more tractable. They experimented with three diverse domains: ancient Babylon (history), The Big Bang Theory (entertainment), and the DAX 40 German stock market index (finance).

The team analyzed the results using three categories of metrics: yield (the quantity of extracted elements), lexical similarity (how similar the exact words or phrases are), and semantic similarity (how similar the underlying meanings are). They tested four variations: the initial seed entity, the language of the prompt, the model’s randomness (temperature), and different LLM models.

Key Findings

The study yielded several important insights:

  • High Termination Rates, with Caveats: For most settings, including the base runs and variations in seed or temperature, the knowledge extraction process reliably terminated within a few hours. However, some language variations (Italian, German, French) and certain open-source LLMs (Llama 4 Scout, DeepSeek-R1, Teuken 7b Instruct) struggled to terminate, often generating excessive or hallucinated content, or getting stuck in repetitive loops. This suggests that termination capability can be highly model-dependent.
  • Mixed Reproducibility: When the extraction was repeated, the output size (yield) was highly consistent. However, the exact wording (lexical similarity) showed only moderate consistency, meaning about one-third of elements matched exactly across runs. On the other hand, the underlying meaning (semantic similarity) was fairly high, with more than half of the elements having a very close semantic match. This indicates that while the LLM consistently extracts similar knowledge, it expresses it in varied ways. Interestingly, more popular entities showed higher reproducibility.
  • Varying Robustness: The process proved highly robust to changes in the starting “seed” entity and the model’s “temperature” (randomness). This means that altering these factors didn’t significantly change the quantity or content of the extracted knowledge. However, robustness was lower when the language of the prompt or the LLM model itself was changed. Different languages led to highly inconsistent yields and lower semantic similarity, while different models also showed intermediate robustness.

Enhancing Stability with Ensembling

The researchers also found that a simple technique called “ensembling”—combining results from multiple runs by taking the intersection of shared triples—could significantly improve the stability of the output. This method, despite higher computational costs, offers a way to achieve more reliable knowledge bases.

Also Read:

Conclusion

Overall, this research provides foundational insights into the process of LLM knowledge materialization. It demonstrates that the GPTKB method can effectively surface the core factual knowledge of an LLM, especially for specific domains. While it confirms the potential of this approach, it also highlights important limitations, particularly concerning the impact of language and the choice of LLM model on the reliability and consistency of the extracted knowledge. For more details, you can read the full paper here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -