spot_img
HomeResearch & DevelopmentAI-Powered Insights: Enhancing Teaching Evaluation in Large Engineering Programs

AI-Powered Insights: Enhancing Teaching Evaluation in Large Engineering Programs

TLDR: Texas A&M’s College of Engineering has implemented an AI-powered system using large language models to summarize vast amounts of student feedback, providing actionable insights for teaching improvement and faculty development. The system anonymizes data, contextualizes scores with visual analytics, and aims to enhance fairness and efficiency in evaluations, while emphasizing human oversight and a holistic approach to assessing teaching quality.

Evaluating teaching effectiveness at scale has long been a significant challenge for large universities, especially within engineering programs that enroll tens of thousands of students. Traditional methods, relying on manual review of student evaluations, are often impractical, leading to overlooked insights and inconsistent data usage. To address this, the College of Engineering at Texas A&M University has introduced a scalable, AI-supported framework designed to synthesize qualitative student feedback using large language models (LLMs).

This innovative system employs hierarchical summarization, anonymization, and exception handling to extract actionable themes from open-ended student comments, all while upholding crucial ethical safeguards. Beyond just summaries, the system integrates visual analytics to contextualize numeric scores through percentile-based comparisons, historical trends, and instructional load. The primary goal is to support meaningful evaluation and professional development for faculty, without automating personnel decisions. Early validation efforts, including comparisons with human reviewers and faculty feedback, suggest that the LLM-generated summaries can reliably support formative evaluation and professional growth.

The College of Engineering is deeply committed to “Educational Excellence at Scale,” striving to ensure high-quality learning experiences for all students, regardless of course size or level. Given the immense volume of data generated by end-of-term student evaluations—hundreds of thousands of data points each semester—it becomes nearly impossible for department heads and college-level administrators to review every comment. This often results in valuable insights being missed. The AI-driven system steps in to summarize these comments, focusing on six key institutional questions related to instructional quality. It also flags inappropriate content, such as harassing remarks, for administrative review.

The core question driving this initiative is how AI-driven summarization and evaluation systems can genuinely enhance instructional quality and support faculty development in large engineering institutions. The approach aligns with best practices in qualitative analysis and educational assessment, incorporating student, peer, and self-reflective inputs. While student evaluations of teaching (SET) are a standard feature in higher education, their validity and potential biases have been debated. Factors like class size, course level, and whether a course is required can significantly influence student ratings, underscoring the need for contextualized interpretation.

To counter these concerns, the College advocates for a holistic, data-informed evaluation model that uses multiple measures of teaching effectiveness, not just student surveys. This includes peer observations, teaching portfolios, and self-reflection. The AI system complements this by efficiently processing large-scale qualitative data, extracting concrete and actionable feedback that numerical ratings often miss. This granular insight, such as recurring mentions of unclear assignments or insufficient office hours, can directly spur teaching development initiatives.

The adoption of AI for summarizing student comments raises important ethical considerations. The system’s hierarchical summarization mirrors human coding processes, preserving nuance and distinct dimensions of teaching (e.g., lecture clarity, course organization). Anonymization removes personally identifying information, protecting privacy and reducing model bias. Crucially, the system operates on the principle that AI-generated summaries should inform human judgment, not replace it. Human reviewers retain the responsibility for evaluative and inferential judgments, using AI as a tool to synthesize large volumes of data efficiently.

AI summarization also offers benefits in consistency and efficiency. Unlike human reviewers who might be influenced by emotional responses, an AI system can provide a more balanced aggregation of feedback. Studies have shown AI-identified themes closely align with human coders, and the speed of AI processing significantly outpaces traditional methods, freeing up faculty time for interpretation and intervention planning. Data privacy is maintained by using a secure API and anonymizing data before processing.

Fairness is a paramount concern. The system is designed to mitigate biases often found in student evaluations, such as those against certain instructor demographics. Instructor names are replaced with gender-neutral placeholders to reduce implicit biases. Numeric scores are contextualized using percentile-based scoring and comparisons to similar courses, preventing unfair penalties for teaching challenging or large courses. Visual reports disaggregate results by course type and class size, ensuring appropriate comparisons. The AI’s role is primarily summarization, grounded in student phrasing, with ongoing monitoring for algorithmic bias.

Visual analytics and timely feedback are integral to the system’s impact. Historical performance trends help instructors recognize positive momentum, even if scores are below average. Percentile-based scoring supports motivation by providing clear benchmarks for improvement. The system also highlights “teaching impact,” recognizing faculty who achieve strong SET scores while teaching a large number of students. These tools foster a supportive, improvement-focused culture, streamlining decision-making for department heads and informing institutional policy and resource allocation.

The introduction of AI-driven evaluations sends a powerful cultural message: the college values teaching excellence and educational innovation. It helps identify and recognize excellent teaching and provides structured support for areas needing improvement. Shared governance in the system’s design has built faculty trust and buy-in. The system encourages a mindset of long-term growth and collaboration, enabling early detection of trends or sudden changes in scores, which can prompt proactive intervention and support.

Validation and benchmarking are continuous processes. Early trials showed AI summaries captured dominant themes similar to human analysis, and faculty feedback has been largely positive regarding accuracy. The system processed over 140,000 student comments across thousands of course sections in the 2023-2024 academic year, demonstrating its scalability. The Student Evaluation of Teaching (SET) survey itself, developed by a university-level committee, includes core questions that analysis suggests have high correlation, implying potential for shorter, more focused surveys with AI summarization capabilities.

The paper acknowledges risks such as over-reliance on AI summaries (automation bias), faculty distrust, and an overemphasis on quantifiable aspects. These are mitigated through explicit guidance, training, and emphasizing that AI is a tool for improvement, not judgment. The system reinforces that teaching effectiveness is multidimensional and that AI outputs are one piece of a broader evaluation process. These thoughtful designs and policies aim to leverage AI’s strengths while minimizing unintended consequences.

Also Read:

The deployment of this AI-supported system has broader implications for institutional policy, faculty recognition, and professional development. Policies are being updated to appropriately incorporate AI insights, ensuring human interpretation and preventing sole reliance on AI for personnel actions. Analytics help identify potential nominees for teaching awards, balancing data-driven identification with qualitative judgment. Summarized peer evaluation data can also support mentoring networks and targeted workshops. This integration of AI is seen as a catalyst for policy evolution and cultural change, reinforcing that effective teaching is measurable, improvable, and worth rewarding. For more detailed information, you can refer to the full research paper: Teaching at Scale: Leveraging AI to Evaluate and Elevate Engineering Education.

Rhea Bhattacharya
Rhea Bhattacharyahttps://blogs.edgentiq.com
Rhea Bhattacharya is an AI correspondent with a keen eye for cultural, social, and ethical trends in Generative AI. With a background in sociology and digital ethics, she delivers high-context stories that explore the intersection of AI with everyday lives, governance, and global equity. Her news coverage is analytical, human-centric, and always ahead of the curve. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -