TLDR: This research introduces a framework to objectively evaluate and enhance the quality of research proposals generated by Large Language Models (LLMs) like ChatGPT-4o. It proposes two metrics—content quality (assessed by AI grading tools) and reference validity (manual fact-checking)—and an iterative prompting method. Experiments show that these metrics provide a quantitative assessment, and the iterative prompting significantly improves content quality while reducing reference inaccuracies and fabrications, addressing critical ethical concerns in academic writing.
Large Language Models (LLMs) such as ChatGPT are becoming increasingly common tools in academic writing. While they offer significant support for grammar, vocabulary, and style, their use also brings forth ethical concerns, particularly regarding the generation of incorrect or fabricated references and the potential spread of misinformation.
Current methods for evaluating the quality of content produced by LLMs often rely on subjective human judgment. This approach is not only labor-intensive but can also lack objectivity, leading to inconsistencies and reliability issues. Recognizing these challenges, a recent study aimed to provide a more quantitative and objective framework for assessing and improving LLM capabilities in research proposal writing.
A New Approach to Evaluation and Enhancement
The researchers proposed two key evaluation metrics: content quality and reference validity. Content quality assesses various aspects of the written text, including grammar, fluency, clarity, relevance, organization, style, source formatting, and adherence to academic conventions. This was evaluated using an AI-based grading system that averaged scores from professional platforms like Study Fetch, QuillBot, and Grammarly. Reference validity, on the other hand, focused on the accuracy of citations, including author names, publication titles, and page numbers, and was assessed through a meticulous manual fact-checking process.
To address the enhancement of LLM writing capabilities, the study introduced an iterative prompting method. This approach involves a continuous cycle of generating a proposal, evaluating it using the proposed metrics, and then feeding the scores and specific feedback back to the LLM. For content quality, feedback highlighted areas needing improvement, prompting the LLM to revise the proposal. For reference validity, a reference guide was provided to help the LLM correct errors in format, author, title, and other details, as well as identify and eliminate fabricated references. This iterative process was repeated until the generated proposals met desired quality standards.
Experimental Insights
The study utilized ChatGPT-4o to generate research proposals under two distinct writing strategies: ‘GPT-only,’ where the model generated proposals without human intervention, and ‘GPT-assisted,’ where human-provided references guided the model. Proposals were created on three educational topics: Student Agency, Teacher Agency, and Teacher Identity.
Initially, the content quality scores for proposals generated by both strategies were quite similar, generally above 80 out of 100. However, after implementing the iterative prompting method, all proposals showed significant improvements in content quality. The most notable increase was an 11.02% gain in a GPT-assisted proposal on Teacher Identity, reaching an average score of 90.67. This demonstrated that iterative feedback effectively enhanced the overall quality and stability of the generated content.
Regarding reference validity, the initial accuracy varied widely, with some proposals having as low as 38.89% correct references, while others achieved 100%. The iterative prompting method proved highly effective in improving reference accuracy, with correctness rates stabilizing after about three rounds of prompting. For instance, a GPT-only proposal on Teacher Agency saw its reference correctness jump from 40% to 80%. The method also successfully addressed in-text citation errors, such as the use of non-academic symbols or referencing by file names.
Also Read:
- Decoding AI’s Writing Style: A Benchmark for Stylistic Variation in LLM-Generated Texts
- Understanding Large Language Models in Legal AI: A Deep Dive into Current Trends and Future Paths
Implications for Academic Writing
This research highlights that while LLMs offer powerful assistance in academic writing, a structured approach is crucial for ensuring the credibility and quality of their output. The dual-metric evaluation provides an objective and quantitative way to assess LLM performance, moving beyond subjective human judgment. Furthermore, the iterative prompting method offers a practical solution for continuously refining LLM-generated content, significantly improving both content quality and the accuracy of references, thereby addressing critical ethical challenges in academic contexts.
This framework offers users a practical way to generate high-quality research proposals tailored to their needs. Future research can build upon this work by developing more efficient writing strategies and advanced methods to further enhance the writing capabilities of LLMs. You can read the full research paper here.


