TLDR: A new study introduces an LLM-driven rubric-based assessment framework for evaluating middle school students’ algebraic competence in multi-stage block coding tasks. Implemented on an online platform, the system records intermediate responses and uses LLMs to provide process-oriented feedback aligned with expert judgments. The field study with 42 students demonstrated strong agreement between LLM evaluations and human experts, highlighting the framework’s validity and scalability for online mathematics and STEM education.
As online learning platforms continue to grow, there’s an increasing demand for assessment methods that go beyond just checking if an answer is right. Educators need to understand the deeper cognitive processes students use and how well they align with learning goals. A recent study by Yong Oh Lee, Byeonghun Bang, and Sejun Oh from Hongik University introduces an innovative solution: a rubric-based assessment framework powered by Large Language Models (LLMs) to evaluate algebraic competence in multi-stage block coding tasks.
Traditional assessment methods often struggle with the complexities of modern education. In subjects like mathematics, especially algebra, true academic achievement isn’t just about the final answer; it’s about the quality of the problem-solving process, including reasoning, representation, and communication. Rubrics are excellent tools for this, providing clear criteria and performance levels. However, designing and applying rubrics for multi-stage problems, particularly in large online environments, becomes incredibly challenging and time-consuming for human evaluators. Conventional automated systems can check correctness but fall short when student responses show diverse reasoning strategies or intermediate steps.
This is where Large Language Models come in. With their advanced capabilities in understanding and generating natural language, LLMs can interpret various input formats, including text, code, and mathematical notation. This makes them ideal for evaluating both the correctness and the reasoning quality of a solution, positioning them as a powerful tool for rubric-based assessment in math and computational tasks.
The researchers developed a framework where mathematics education experts systematically mapped each step of a problem-solving sequence to predefined rubric dimensions. This allows the LLM to conduct consistent, criterion-referenced assessments that delve into the problem-solving process, not just surface-level correctness. The study utilized a real-world context, stepwise problem set that combined traditional mathematical reasoning with block-based coding activities. This approach helps students build both conceptual understanding and computational thinking by embedding problems in authentic scenarios and gradually increasing task complexity.
The system was implemented on an online learning platform that meticulously records all intermediate student responses. It uses rule-based scoring for objective correctness and then leverages LLM-generated evaluations for a comprehensive assessment of achievement. This dual approach ensures both efficiency and depth in evaluation.
To test its effectiveness, a field study was conducted with 42 middle school students who engaged in multi-stage quadratic equation tasks using block coding. The study integrated student self-assessments and expert ratings to benchmark the system’s outputs. The results were highly encouraging: the LLM-based rubric evaluation showed strong agreement with expert judgments and consistently provided rubric-aligned, process-oriented feedback. This demonstrates both the validity and scalability of integrating LLM-driven rubric assessment into online mathematics and STEM education platforms.
The study also provided valuable insights into student performance. While students generally showed strong understanding of conceptual objectives and structured applications, there were noticeable weaknesses in more advanced or self-directed computational tasks, particularly those requiring deeper integration of learned concepts and proactive use of engineering tools. Cluster analysis further revealed diverse learner profiles, highlighting specific areas where targeted instructional support might be needed. The research also emphasized the importance of “keystone” block-coding items in accurately assessing programming proficiency.
Also Read:
- Evaluating Multimodal AI for Grading Handwritten Student Math Work
- AI Generates Personalized Math Problems That Students Prefer
In conclusion, this research offers a significant advancement for technology-enhanced mathematics education. By combining expert-defined rubrics, process-oriented LLM evaluation, and authentic problem design, the framework provides a scalable and reliable method for assessing complex student competencies. This approach not only reduces the burden of manual grading but also offers valuable diagnostic insights into learning trajectories, paving the way for more robust and effective assessment systems in education. You can read the full research paper here: LLM-Driven Rubric-Based Assessment of Algebraic Competence.


