spot_img
HomeResearch & DevelopmentAssessing Large Language Models for Financial Auditing Compliance

Assessing Large Language Models for Financial Auditing Compliance

TLDR: This research investigates the effectiveness of Large Language Models (LLMs), both open-source (Llama-2) and proprietary (GPT models), in automating regulatory compliance verification in financial auditing. Using custom datasets from PwC Germany, the study found that while GPT-4 generally performs best, the open-source Llama-2 70 billion model excels at identifying non-compliance. The paper highlights challenges with non-English data and the importance of prompt design, concluding that while LLMs hold promise, careful selection and tailoring are needed for reliable deployment in auditing.

Financial auditing, a cornerstone of corporate transparency, has traditionally been a highly labor-intensive process. Auditors meticulously examine vast financial documents to ensure they align with complex legal requirements and accounting standards, such as the International Financial Reporting Standards (IFRS) or Germany’s Handelsgesetzbuch (HGB). While artificial intelligence has begun to assist by recommending relevant text passages, a significant challenge remains: verifying if these passages truly comply with specific legal mandates.

A recent research paper, titled Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models, delves into the efficiency of publicly available Large Language Models (LLMs) in tackling this critical verification step. The study, conducted by researchers from Fraunhofer IAIS, the University of Bonn, and PricewaterhouseCoopers GmbH, specifically compares cutting-edge open-source LLMs like Llama-2 with proprietary models from OpenAI, including various GPT versions.

The Core Challenge and LLM Potential

The paper highlights that existing AI systems often fall short in the crucial compliance verification phase. Auditors still need to manually confirm if recommended text excerpts meet legal stipulations. This research aims to explore how LLMs, known for their impressive reasoning and text comprehension, can reshape this auditing paradigm by automatically validating the compliance of financial report passages with regulatory standards.

The motivation behind exploring open-source models is twofold: cost-effectiveness and data privacy. These are significant concerns in the sensitive domain of accounting and financial data. The study builds upon prior work, such as the Automated List Inspection (ALI) and ZeroShotALI systems, which focused on mapping legal requirements to financial report segments. This new research extends these capabilities to evaluate the actual compliance of those segments.

Experimentation and Key Findings

To assess performance, the researchers utilized two custom datasets provided by PwC Germany, based on IFRS and HGB compliant financial reports. They evaluated six state-of-the-art LLMs: Llama-2 (7b, 13b, and 70b parameter versions), GPT-3.5-Turbo, GPT-3.5-Turbo-16K, and GPT-4. Each model was tested with eight different prompt configurations to understand the impact of prompt design on performance.

One notable finding was that the way a task is phrased and the structure of permitted model responses significantly influence performance. A “closed” format, where the model is constrained to return only specific answers like ‘yes’, ‘no’, ‘unclear’, or ‘not applicable’, yielded superior results compared to “open-ended” responses. Surprisingly, more advanced prompting techniques like Chain-of-Thought or Tree-of-Thought did not consistently lead to better performance.

In terms of overall performance, GPT-4 emerged as the generally best-performing model, particularly on the IFRS dataset. However, a significant challenge was observed with non-English data, specifically the German HGB dataset. Most LLMs, including Llama-2, which was trained predominantly on English text, showed considerably worse performance in German. This underscores a critical limitation for global auditing applications.

Interestingly, for the open-source Llama-2 models, increasing parameter size did not always translate into superior performance; the Llama-2-70b model sometimes performed worse than its smaller Llama-2-7b counterpart.

Focus on Non-Compliance Detection

A crucial aspect for auditors is the accurate detection of non-compliance, or ‘No’ answers, to avoid false positives which can have severe repercussions. In this regard, the open-source Llama-2-70b model demonstrated outstanding performance on the IFRS dataset, achieving high precision and recall for the ‘No’ class. This suggests its strong potential for identifying instances where financial report passages do not comply with regulatory requirements.

However, the models struggled with the ‘Unclear’ classification, often being overly confident in their ‘Yes’ or ‘No’ answers. This indicates an area for future improvement, as real-world auditing often involves ambiguous situations.

Also Read:

Conclusion and Future Outlook

The research concludes that while LLMs hold immense promise for revolutionizing financial auditing, their deployment requires careful consideration. Selecting the right model, tailoring prompts to specific tasks, and acknowledging their current limitations, especially with non-English languages, are paramount. The study suggests that “out-of-the-box” LLMs are not yet fully reliable for comprehensive regulatory compliance assessment.

Future work will focus on further investigating and potentially fine-tuning the open-source Llama-2-70b model, given its strong performance in detecting true negatives. Fine-tuning on extensive accounting compliance data could enhance its effectiveness, offering a reliable, cost-effective, and data-private solution for automated auditing pipelines.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -