spot_img
HomeResearch & DevelopmentEvaluating Language Models for Grade-Level Adaptability in K-12 Education

Evaluating Language Models for Grade-Level Adaptability in K-12 Education

TLDR: EDUADAPT is a new benchmark dataset of nearly 48,000 grade-labeled science question-answer pairs (Grades 1-12) designed to evaluate how well large language models (LLMs) can adapt their responses to different student age groups. Evaluations show that while larger LLMs perform better overall, they consistently struggle to generate appropriate content for early-grade students (Grades 1-5), highlighting a significant gap in their educational adaptability and the need for targeted training strategies.

Large language models (LLMs) are rapidly changing how we approach education, offering capabilities like answering questions, explaining complex ideas, and generating content across many subjects. Despite their impressive performance on various academic tests, these models often fall short when it comes to tailoring their responses to a student’s specific grade level. This is a crucial requirement in K-12 education, where using vocabulary and explanations appropriate for a child’s age is essential for effective learning. Currently, LLMs frequently produce outputs that are either too advanced or too vague for younger students, and there hasn’t been a standardized way to measure their ability to adjust across different cognitive and developmental stages.

To tackle this significant challenge, researchers Numaan Naeem, Abdellah El Mekki, and Muhammad Abdul-Mageed have introduced EDUADAPT, a groundbreaking benchmark dataset. This dataset comprises nearly 48,000 question-answer (QA) pairs, each carefully labeled with a specific grade level. These pairs cover nine science subjects and span the entire K-12 range, grouped into four distinct grade levels (Grades 1-2, 3-5, 6-8, and 9-12). The creation of EDUADAPT involved a meticulous two-phase process: first, automatically generating the QA benchmark, and then rigorously verifying its quality through human review.

The initial generation phase began with collecting and cleaning text from Wikipedia articles across a broad spectrum of science subjects, including Chemistry, Computer Science, Biology, Physics, and Geography. Recognizing that Wikipedia content is typically written for higher reading levels, a specialized LLM (Phi-4) was employed to classify these passages by grade level. This ensured that only texts appropriate for specific age groups were used to generate questions. Following this, both multiple-choice and open-ended QA pairs were created using prompts specifically designed for each grade level, adhering to the guidelines of the Next Generation Science Standards (NGSS). This comprehensive process initially yielded approximately 166,000 QA pairs.

To guarantee the high quality of the dataset, a self-reflection mechanism was implemented. The same LLM that generated the questions also evaluated its own outputs against pedagogical and linguistic standards informed by NGSS guidelines. Each QA pair was scored on five criteria: language appropriateness, grade alignment, relevance, clarity, and subject fit. Only pairs that achieved a score of at least 8 out of 10 on every criterion were retained, resulting in the final, high-quality dataset of 47,734 QA pairs. These were then strategically divided into training, development, and test sets to facilitate both model adaptation and evaluation.

Further validating the dataset’s integrity, a human verification process was conducted. A subset of the test set, consisting of 1,000 QA pairs, was reviewed by three independent educational content developers. They applied the same evaluation criteria as the self-reflection mechanism. The consistently high inter-annotator agreement, measured using Fleiss’ Kappa, strongly reinforced the credibility of these human ratings as a reliable benchmark for dataset quality.

Experiments were carried out using the EDUADAPT test set to benchmark a diverse group of leading open-source instruction-tuned LLMs. These included models from the Qwen2.5, SmolLM, Gemma, LLaMA3, and Mistral series, ranging in size from 1.5 billion to 24 billion parameters. The primary objective was to assess how effectively these models could answer questions tailored for different grade levels. For multiple-choice questions, performance was measured by accuracy. For open-ended responses, an innovative “LLM-as-a-judge” framework was utilized, where three independent LLMs (GPT-4o, Qwen2.5-72B, and LLaMA3.3-70B) scored each response against a reference answer on a 1-10 scale across five qualitative criteria.

The experimental results revealed a clear and consistent performance gap: while larger models generally achieved higher scores, all evaluated LLMs struggled significantly more with generating grade-appropriate responses for lower grades (Grades 1-5) compared to higher grades. For instance, on open-ended questions, models typically scored only 60-70% for early-grade content, whereas their performance improved to up to 85% for higher grades. Smaller models (1.5B-3B) showed particularly poor performance on early-grade content, with MCQ accuracy ranging between 50-60% for Grades 1-5. This indicates that even leading LLMs find it challenging to consistently adjust their outputs to the specific linguistic and cognitive needs of different age groups.

A qualitative error analysis of the model-generated responses provided deeper insights into these struggles. For lower grades, common issues included hallucinations (generating incorrect information), partial misalignments, imprecision, and answers that were either oversimplified to the point of being off-topic or too complex. For higher grades, problems often involved fabricated details, factual unreliability, and over-explanation that reduced precision and grade alignment. This analysis also highlighted the limitations of traditional automated metrics like BLEU and ROUGE, which often produced high scores even when answers were semantically incorrect or developmentally inappropriate. In contrast, MCQ accuracy and the LLM-as-a-judge scoring for open-ended questions provided more meaningful, pedagogically grounded assessments.

This research marks a significant step forward by introducing the first comprehensive synthetic benchmark specifically designed to evaluate the grade-level adaptability of LLMs across the entire K-12 system. The findings underscore a critical need for developing grade-aware training, prompting, and fine-tuning strategies that are specifically tailored to the unique learning needs of younger students. Future work in this area could involve expanding subject coverage, incorporating multimodal QA (e.g., questions based on images or diagrams), supporting multilingual QA to enhance accessibility for non-English-speaking students, and further addressing lower-grade performance through advanced data augmentation and curriculum-aligned pretraining techniques.

Also Read:

For more details on the EDUADAPT dataset and its methodology, you can refer to the full research paper here.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -