spot_img
HomeResearch & DevelopmentBenchmarking Financial Knowledge Graphs: Introducing FinReflectKG - EvalBench

Benchmarking Financial Knowledge Graphs: Introducing FinReflectKG – EvalBench

TLDR: FinReflectKG – EvalBench is a new benchmark and evaluation framework for extracting knowledge graphs from SEC 10-K filings using LLMs. It features a bias-controlled “commit-then-justify” judging protocol and evaluates extraction quality across faithfulness, precision, relevance, and comprehensiveness. The study found that reflection-based extraction excels in comprehensiveness, precision, and relevance, while single-pass extraction achieves the highest faithfulness, highlighting trade-offs and the need for multi-dimensional evaluation.

Large language models (LLMs) are increasingly transforming how structured knowledge is extracted from vast amounts of unstructured financial text. This capability is particularly valuable in finance, where information from documents like SEC 10-K filings can be organized into Knowledge Graphs (KGs) to support critical applications such as compliance monitoring, risk management, and large-scale financial analytics. However, despite the growing use of LLMs for this purpose, a significant challenge has been the absence of a universal benchmark or a unified framework to reliably evaluate the quality of these financial KGs.

Addressing this critical gap, researchers from Domyn have introduced FinReflectKG – EvalBench, a groundbreaking benchmark and evaluation framework specifically designed for KG extraction from SEC 10-K filings. This new system builds upon the principles of FinReflectKG, which is a financial KG that links audited data to source chunks from S&P 100 filings and supports various extraction methods, including single-pass, multi-pass, and reflection-agent-based approaches. EvalBench aims to provide a robust and transparent way to assess the performance of LLMs in constructing financial KGs.

A core innovation of FinReflectKG – EvalBench is its deterministic “commit-then-justify” judging protocol, which incorporates explicit bias controls. This protocol is designed to mitigate common issues found in LLM-as-Judge evaluations, such as position effects (where the order of information influences judgment), leniency (an tendency to be overly positive), verbosity (favoring longer explanations), and reliance on world knowledge outside the provided text. By enforcing these controls, the framework ensures that evaluations are fair, reliable, and reproducible. Each candidate piece of information, known as a “triple,” is judged with binary decisions for faithfulness, precision, and relevance. Additionally, the overall completeness, or “comprehensiveness,” is assessed on a three-level scale (good, partial, bad) at the document chunk level.

The framework evaluates extraction quality across four key dimensions. Faithfulness measures whether the extracted information is factually grounded in the source text, without relying on external knowledge or assumptions. Precision assesses the clarity and specificity of the extracted information, penalizing generic terms or imprecise expressions. Relevance checks if the extracted information directly contributes to the main theme of the source text, avoiding tangential details. Finally, Comprehensiveness evaluates how well the set of extracted information covers all the essential facts within a given text chunk.

The research also compares three distinct extraction modes. The Single-pass mode uses a single LLM to both extract and normalize the information. The Multi-pass mode divides the task between two LLMs: one for extraction and another for normalization. The most advanced, Reflection mode, employs an iterative, agent-based workflow where extraction and feedback loops continuously refine the information until inconsistencies are resolved or a set limit of iterations is reached.

The findings from the evaluation are insightful. The reflection-based extraction approach emerged as the superior method, achieving the best performance in comprehensiveness, precision, and relevance. This suggests that its iterative refinement process helps in capturing a broader range of core facts and improving structural accuracy. However, the single-pass extraction method maintained the highest faithfulness, indicating it is more conservative and strictly adheres to the original text, albeit with less coverage. This highlights an inherent trade-off: expanding coverage, as seen in reflection, can increase the risk of generating information that slightly extends beyond the strict boundaries of the source text.

While precision scores across all modes were relatively modest, indicating room for further improvement, reflection still led in this dimension. Relevance, on the other hand, showed consistently high values across all modes, suggesting that most extracted information remains topically aligned with the source text. These results collectively underscore the importance of a multi-dimensional evaluation approach, as no single metric can fully capture the complex trade-offs involved in building financial Knowledge Graphs. The study concludes that LLM-as-Judge protocols, when equipped with explicit bias controls, offer a reliable and cost-efficient alternative to traditional human annotation, while also providing structured error analysis. For more details, you can read the full research paper here.

Also Read:

FinReflectKG – EvalBench represents a significant step forward in advancing transparency and governance in financial AI applications. By establishing transparent, reproducible, and bias-aware evaluation standards, it aims to accelerate progress toward trustworthy financial KGs that support critical downstream tasks.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -