spot_img
HomeResearch & DevelopmentReproducibility Concerns Emerge in Commercial LLM Software Engineering Studies

Reproducibility Concerns Emerge in Commercial LLM Software Engineering Studies

TLDR: A study investigating the reproducibility of commercial LLM performance in software engineering research found that out of 65 studies using OpenAI models, only five were fit for reproduction, and none of these could be fully reproduced. Key issues included missing artifacts, inadequate reporting of model configurations, dependency problems, and deprecated models. The study also concluded that ACM artifact badges were not reliable indicators of long-term reproducibility, highlighting a critical need for improved practices in LLM-centric empirical research.

A recent research paper, titled “Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies,” delves into a critical issue facing the rapidly evolving field of Large Language Models (LLMs) in software engineering: reproducibility. Authored by Florian Angermeir, Maximilian Amougou, Mark Kreitz, Andreas Bauer, Matthias Linhuber, Davide Fucci, Fabiola Moyón C., Daniel Mendez, and Tony Gorschek, the study highlights significant challenges in replicating findings from academic research involving commercial LLMs.

The increasing integration of LLMs like GPT, Gemini, and LLaMA into software engineering tasks has led to a surge in academic publications. For instance, a substantial number of papers at major conferences like ICSE 2024 featured experiments with LLMs. However, the inherent non-deterministic nature of these models, coupled with frequent updates and a lack of transparency from commercial providers, raises serious questions about whether research results can be consistently reproduced by others.

The Study’s Approach

To investigate this, the researchers conducted a comprehensive analysis of 86 LLM-centric studies published at the International Conference on Software Engineering (ICSE 2024) and Automated Software Engineering (ASE 2024). Their focus was on studies that utilized commercial LLMs, particularly those from OpenAI, due to their widespread use. Out of these 86 articles, 65 employed OpenAI services. The team then narrowed down their scope to 18 articles that provided research artifacts and seemed suitable for reproduction.

Using a specially developed, containerized reproduction framework, the researchers attempted to replicate these 18 studies. The framework was designed to ensure platform independence and version stability, downloading and, if necessary, transparently modifying research artifacts. Each experiment was run up to 30 times or until a cost of $500 was reached, to account for the variability in LLM outputs.

Alarming Findings on Reproducibility

The results were stark: out of the 18 studies initially deemed fit for reproduction, only five were actually executable within the framework. More critically, none of these five studies could be fully reproduced. Two studies were found to be partially reproducible, meaning some results aligned with the original findings, while three studies yielded entirely divergent results, indicating they were not reproducible at all. This suggests a significant gap between reported research and the ability of independent teams to verify those findings.

Factors Impeding Reproduction

The study identified several key factors that consistently hindered reproducibility, many of which are not unique to LLM research but are exacerbated by the technology’s characteristics:

  • Missing or Incomplete Artifacts: A staggering 35 out of 86 articles did not provide their code or data, and another 15 had artifacts that were too incomplete to be recovered.
  • Missing Details in Reporting: Many papers lacked crucial information. Eight articles failed to even name the specific LLMs used, referring to them generically as “ChatGPT” or “LLM.” Only a third reported the model’s ‘temperature’ setting, and even fewer (16%) mentioned other critical parameters like ‘top-p’ or ‘top-k.’
  • Dependency Version Issues: Inconsistent or unspecified software dependency versions often led to conflicts and execution failures.
  • Deprecated Models: A significant challenge, particularly with commercial LLMs, was the deprecation of models. Several studies used models that were no longer available, making direct reproduction impossible.
  • Incomplete Documentation and General Code Issues: Poor documentation on how to run experiments, unassigned variables, and hardcoded paths were common problems.

The Reliability of ACM Artefact Badges

The researchers also examined the effectiveness of ACM artifact badges, which are intended to signal the quality and reusability of research artifacts. Out of the 86 articles, 19 had been awarded an ACM badge, primarily “Artifacts Available” or “Artifacts Evaluated – Reusable.” However, the study found that one year after evaluation, half of the artifacts with the “Reusable” badge no longer met ACM’s requirements, being incomplete, non-functional, or lacking sufficient documentation. This led the authors to conclude that ACM artifact badges are not a reliable indicator of long-term reproducibility for LLM-centric studies.

Also Read:

Implications and Recommendations

The findings raise serious concerns about the long-term scientific value and sustainability of LLM-centric empirical software engineering research. If studies published just a year ago cannot be reproduced, the cumulative knowledge building in the field is at risk.

The authors provide several recommendations to improve the state of reproducibility:

  • For Authors: Rigorously document all experimental data, share complete and versioned artifacts, and adhere to comprehensive guidelines for reporting LLM-centric studies.
  • For Venues (Conferences and Journals): Mandate the disclosure of full research artifacts, including a complete Software Bill of Materials. They suggest rethinking current artifact evaluation criteria and applying the same high standards of scientific rigor to LLM studies as to other research types.
  • For Funding Agencies: Emphasize long-term accessibility and reproducibility of research artifacts, extending open science practices to include reproducibility, and supporting automation in artifact assessment to ease the burden on reviewers.

In conclusion, the study underscores that reproducibility in LLM-centric research is not an optional consideration. Without structural improvements in how studies are conducted, documented, and evaluated, the software engineering community risks building a body of knowledge that cannot be reliably tested or trusted. For more details, you can read the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -