TLDR: A new study evaluates five large language models (LLMs) in generating scientific paper reviews, comparing them against human reviewers using semantic similarity and knowledge graph analysis. The findings indicate that LLMs excel at producing descriptive and affirmational content, accurately summarizing contributions and methodologies. However, they significantly underperform in critical areas such as identifying weaknesses, asking substantive questions, and adjusting feedback based on paper quality, often exhibiting a bias towards leniency. The research highlights LLMs’ potential as descriptive assistants but underscores the continued necessity of human expertise for deep critical analysis in peer review.
The world of scientific research is experiencing an unprecedented boom, leading to a significant increase in the number of papers submitted to top conferences. This surge has placed immense pressure on the traditional peer-review system, which relies on human experts to evaluate the quality and rigor of new research. In response, the scientific community has begun exploring the potential of large language models (LLMs) to automate parts of this crucial process.
A recent study, titled “Unveiling the Merits and Defects of LLMs in Automatic Review Generation for Scientific Papers”, delves deep into how well these AI models perform when tasked with generating reviews for scientific papers. The research, conducted by Ruochi Li, Haoxuan Zhang, Edward Gehringer, Ting Xiao, Junhua Ding, and Haihua Chen, provides a comprehensive look at the capabilities and limitations of LLMs in this complex domain.
The core of this investigation involved developing a sophisticated evaluation framework. This framework didn’t just look at surface-level text comparisons; it integrated semantic similarity analysis and structured knowledge graph metrics. This allowed the researchers to assess LLM-generated reviews against human-written ones, focusing on critical aspects like reasoning, contextual understanding, and sensitivity to paper quality.
To conduct their study, the team built a large-scale benchmark dataset. This impressive collection included 1,683 papers and 6,495 expert reviews from prestigious conferences like ICLR and NeurIPS over several years. Crucially, each paper was categorized into ‘good’, ‘borderline’, or ‘weak’ based on reviewer agreement, allowing for a quality-sensitive evaluation. Five different state-of-the-art LLMs, including GPT-4o, Gemini-2.0-Flash, Claude-3.5-Sonnet, Qwen2.5-72B-instruct, and LLaMA3.3-70B-instruct, were then used to generate reviews for these papers.
What LLMs Do Well: Descriptive Content
The findings revealed some clear strengths of LLMs. The models proved highly competent in generating descriptive and affirmational content. For instance, in the summary and strengths sections of reviews, LLMs consistently captured the main contributions and methodologies of the original work with high fidelity. GPT-4o, for example, generated 15.74% more entities than human reviewers in the strengths section of good papers in ICLR 2025. This suggests that LLMs are excellent at summarizing and highlighting positive aspects, often providing more extensive textual content than human reviewers.
Where LLMs Fall Short: Critical Analysis and Quality Sensitivity
However, the study also exposed significant limitations. LLMs consistently underperformed when it came to identifying weaknesses, raising substantive questions, and adjusting their feedback based on the paper’s quality. In the weaknesses section, GPT-4o produced 59.42% fewer entities than human reviewers. Furthermore, while human reviewers significantly increase the detail and critical feedback for weaker papers (a 50% increase in node count from good to weak papers), GPT-4o showed only a 5.7% increase, indicating a lack of sensitivity to paper quality.
This evaluative bias towards leniency means LLMs tend to produce similar evaluations across papers of varying quality, often overrating borderline and weak submissions. Human reviewers, by contrast, demonstrate clear quality awareness, providing more detailed critical feedback for lower-quality submissions and reducing affirmational content.
Knowledge Graph Analysis: Deeper Insights
The knowledge graph analysis provided deeper insights into these differences. In the weaknesses section, human-written reviews consistently produced richer and more conceptually dense knowledge graphs, with a greater number of scientific entities, more diverse conceptual relations, and higher label entropy. This indicates that human reviewers engage more deeply with both paper-grounded and externally inferred knowledge when critiquing papers.
Conversely, in the strengths section, LLM-generated reviews often contained more entities and relations than human-written ones, reinforcing their proficiency in generating extensive affirmational content. This structural divergence highlights that while LLMs can be prolific in descriptive areas, they lack the conceptual depth and contextual grounding required for nuanced critical evaluation.
Also Read:
- Navigating the Future of Healthcare: A Deep Dive into Large Language Models in Medicine
- The Hidden Flaws in AI Evaluation: Why LLM Judge Benchmarks Need a Rethink
Implications for the Future of Peer Review
The research suggests that LLMs can serve as effective assistants for generating initial descriptive assessments, particularly for tasks requiring broad coverage and recall of scientific content. However, human reviewers remain indispensable for tasks demanding deeper conceptual engagement, such as detecting subtle flaws, reasoning about experimental design, or assessing broader impact. The knowledge graph structures can also help area chairs monitor the quality of reviewer feedback, identifying reviews with limited concept types or low complexity.
While this study provides a robust foundation, it acknowledges limitations, such as focusing on specific conferences and a fixed prompt design. Future work could explore more diverse conferences, newer models, and alternative prompting strategies to further enhance our understanding of LLM performance in peer review. You can find more details about this research paper here.


