spot_img
HomeResearch & DevelopmentUnderstanding AI Grading Uncertainty with Semantic Entropy

Understanding AI Grading Uncertainty with Semantic Entropy

TLDR: This research introduces semantic entropy, a new metric that measures the variability in AI-generated explanations for student responses, as a signal for potential human grader disagreement. Experiments on the ASAP-SAS dataset show that semantic entropy correlates with human disagreement, generalizes across academic subjects (especially interpretive ones), and is sensitive to task structure like source dependency. The findings suggest semantic entropy can make AI grading more transparent by flagging ambiguous cases for human review.

Automated grading systems have become increasingly common in education, offering efficiency and scalability for scoring student responses. However, a significant limitation of these systems is their inability to signal when a grading decision might be uncertain or contentious. Unlike human graders who often disagree on subjective or ambiguous answers, AI systems typically provide only a final numeric score, potentially overlooking cases that truly warrant human review.

A new research paper, “Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement,” introduces a novel approach to address this challenge. The authors, Karrtik Iyer, Manikandan Ravikiran, Prasanna Pendse, and Shayan Mohanty from Thoughtworks AI Research Labs, propose a metric called semantic entropy. This measure quantifies the variability across multiple explanations generated by an AI model (specifically GPT-4) for the same student response. The core idea is that if an AI model produces diverse or conflicting justifications for a given answer, it likely indicates an underlying ambiguity that could also lead to disagreement among human graders.

The researchers investigated three key questions: First, does semantic entropy align with human grader disagreement? Second, can this signal generalize across different academic subjects? And third, is it sensitive to structural features of the grading task, such as whether external source material is required?

To explore these questions, the team utilized the ASAP-SAS dataset, which includes short-answer responses. They prompted GPT-4 to generate six concise explanation rationales for each student response. These rationales were then clustered based on their semantic similarity using a technique called bidirectional entailment. Semantic entropy was calculated based on the diversity of these clusters. The higher the entropy, the more varied the AI’s explanations, suggesting greater uncertainty or ambiguity.

Semantic Entropy Reflects Human Disagreement

The study found a statistically significant correlation between semantic entropy and human grader disagreement. As the level of human disagreement increased (from low to medium to high), the mean semantic entropy also consistently rose. This suggests that when human graders are more likely to disagree on a response, the AI’s internal reasoning, as reflected by its varied explanations, also shows greater uncertainty. This signal, while modest, can be valuable for identifying responses that might need a human educator’s eye.

Effectiveness Across Academic Subjects

The research also examined whether semantic entropy’s effectiveness varied across different academic subjects like Science, English Language Arts (ELA), Biology, and English. The findings indicated that semantic entropy was more predictive of human disagreement in interpretive subjects such as Biology and English. In these domains, where rubrics might allow for more subjective interpretations, the variability in AI explanations mirrored human uncertainty more closely. Conversely, in more fact-based or rigidly scored subjects like Science and ELA, the correlation was weaker. This highlights that semantic entropy is a domain-sensitive signal, particularly useful for complex or ambiguous grading scenarios.

Sensitivity to Task Structure

A crucial finding was the sensitivity of semantic entropy to the structure of the grading task, particularly whether it required students to refer to external source material (like reading passages or diagrams). Tasks that were “source-dependent” consistently produced substantially higher semantic entropy compared to “non-source-dependent” tasks. This suggests that the cognitive complexity involved in integrating external context amplifies the diversity of plausible explanations, making semantic entropy a stronger predictor of human disagreement in such cases. The paper proposes a preliminary taxonomy of tasks based on rubric openness and complexity, ranging from high-entropy tasks (abstract interpretation) to low-entropy tasks (factual recall).

Also Read:

Implications for AI-Assisted Grading

The study positions semantic entropy as an interpretable uncertainty signal that can enhance the transparency and trustworthiness of AI-assisted grading workflows. By identifying responses where the AI’s reasoning is diverse, even if its final score is confident, educators can gain insights into potential ambiguities in rubrics or student responses. The authors propose a practical quadrant-based decision strategy for educators, combining semantic entropy with human grader disagreement to determine when AI-generated scores should be trusted versus flagged for review. For instance, high entropy combined with high human disagreement clearly indicates a need for mandatory human review or rubric revision.

This work represents a significant step towards building more transparent and educator-aligned assessment pipelines. While the current research uses GPT-4 for both explanation generation and clustering, future work aims to decouple these components to address potential biases. The ultimate goal is to integrate semantic entropy into real-time educational systems to support educators by highlighting student responses that warrant manual review or rubric refinement, thereby fostering more trustworthy AI-assisted learning environments. You can read the full research paper for more details: Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement.

Rhea Bhattacharya
Rhea Bhattacharyahttps://blogs.edgentiq.com
Rhea Bhattacharya is an AI correspondent with a keen eye for cultural, social, and ethical trends in Generative AI. With a background in sociology and digital ethics, she delivers high-context stories that explore the intersection of AI with everyday lives, governance, and global equity. Her news coverage is analytical, human-centric, and always ahead of the curve. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -