spot_img
HomeResearch & DevelopmentUnpacking Writing Scores: A New Approach to Explaining Automated...

Unpacking Writing Scores: A New Approach to Explaining Automated Evaluations

TLDR: This research introduces a novel method to enhance the transparency of Automated Writing Evaluation (AWE) by breaking down high-level writing traits into finer “subtraits.” Leveraging generative language models (GLMs) with zero-shot prompting, the study prototypes subtrait scoring and evidence extraction. While human markers showed moderate agreement in scoring subtraits, and GLMs exhibited fair performance with high consistency, the approach offers a promising path for providing educators and students with detailed, human-interpretable explanations for writing scores, moving beyond traditional “black-box” AWE systems.

Automated Writing Evaluation (AWE) systems have become increasingly common, but their “black-box” nature often leaves educators and students wondering how scores are derived. A new research paper, “Toward Subtrait-Level Model Explainability in Automated Writing Evaluation,” proposes a novel approach to demystify these scores by focusing on subtraits and leveraging the power of generative language models (GLMs).

Understanding Subtraits for Clearer Feedback

The core idea behind this research is to break down broad writing skills, known as “traits,” into more granular components called “subtraits.” For instance, a common trait like “Organization” can be decomposed into subtraits such as “Introduction,” “Paragraph Strategies,” “Cohesion,” and “Conclusion.” By evaluating writing at this finer level, the system aims to provide more specific and actionable feedback.

Traditional AWE systems often rely on extracting lexical, syntactical, and semantic properties of text using natural language processing (NLP) techniques. While these features correlate with writing scores, they don’t directly explain *why* a student received a particular score for a specific skill. Even advanced transformer-based models, while excellent at predicting scores, offer little insight into the underlying writing performance. This new approach directly addresses this transparency gap.

How the Study Unfolded

The research involved three main phases. First, a comprehensive “Writing Skills Tree” was developed for middle school, detailing trait-level skills into independent subtraits, each with a description, rubric, and relevant Common Core Standards tags. Second, human markers scored and annotated 225 student responses to informative/explanatory writing prompts. Four subject matter experts in English and Language Arts provided two reads per response, scoring for two main traits and eight subtraits, and crucially, highlighting specific regions of text that supported their scoring decisions.

The third phase involved automated subtrait scoring and evidence extraction using zero-shot prompting with GLMs. The study utilized OpenAI’s GPT-4o mini, deployed on an internal Microsoft Azure account. To account for the stochastic nature of generative models, each response was processed 10 times for each subtrait, allowing for an analysis of the model’s consistency.

Key Findings and Insights

The study yielded several important findings regarding both human and automated scoring:

  • Human Reliability: Human markers showed moderate inter-rater reliability (agreement) for most subtrait scores. Interestingly, subtraits like “Introduction of the Topic” and “Concluding Statement” had strong agreement at the lowest score point. Feedback suggested that scoring responses pooled from multiple prompts, rather than one item at a time, made the task more challenging for human raters.

  • Trait-Subtrait Correlation: There was a high correlation between the overall trait scores and the average subtrait scores for both evaluated traits, “Purpose and Organization” (0.718) and “Evidence and Elaboration” (0.767).

  • GLM Performance: The GLM models exhibited fair agreement with human scores across all subtraits. While not yet meeting the high standards required for high-stakes assessments, the agreement for some “Evidence and Elaboration” subtraits approached human-human agreement. The models showed lower recall for responses in the lowest and highest score point ranges.

  • Model Consistency: The GLM model demonstrated high consistency in its ratings across multiple runs on the same response, as confirmed by high Krippendorff Alpha values. This suggests that while the accuracy might need improvement, the model’s scoring behavior is stable.

  • Extracted Evidence: A significant aspect of this research is the GLM’s ability to extract relevant text spans as evidence for subtrait scores. This provides a direct explanation for the score. While there was often overlap between human and model selections, the GLM sometimes overproduced evidence or selected different types of evidence (e.g., entire sentences for transitions versus individual words by humans).

Also Read:

Implications for the Future of AWE

While the zero-shot GLM approach is not yet ready for high-stakes assessment environments, its potential is significant. It offers a flexible way to assess a variety of writing skills and can augment existing trait scores with targeted, human-interpretable analyses. In lower-stakes settings, these subtrait scores and their associated evidence can be invaluable for informing feedback strategies, personalizing learning, and powering recommender systems.

Future research plans include investigating how subtrait scores and evidence can be used to train explanatory models for overall trait scores, as well as exploring the impact of prompt engineering, few-shot learning, and fine-tuning on subtrait score accuracy. This work represents an encouraging step toward making automated writing evaluation more transparent and beneficial for both students and educators. You can read the full research paper here: Toward Subtrait-Level Model Explainability in Automated Writing Evaluation.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -