TLDR: A study evaluating frontier LLMs on Indian legal exams found that while AI excels at objective, multiple-choice questions, it significantly underperforms humans in subjective, long-form legal reasoning. Key weaknesses include poor citation discipline, irrelevant content generation, and inappropriate legal voice/structure, indicating LLMs are currently better suited as supportive tools than autonomous legal practitioners.
The integration of Large Language Models (LLMs) into various professional fields is rapidly expanding, and the legal sector is no exception. However, a crucial question arises: are these advanced AI systems truly “court-ready”? A recent research paper delves into this very question, specifically within the context of the Indian legal system, providing a unique and comprehensive evaluation framework.
Evaluating AI on Indian Legal Reasoning
The study, titled “Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning,” by Kush Juvekar, Arghya Bhattacharya, Sai Khadloya, and Utkarsh Saxena from Adalat AI, India, addresses a significant gap in current AI-and-law research. While many studies focus on short-context recall or prediction tasks, this paper uses India’s rigorous public legal examinations as a transparent and jurisdiction-specific benchmark to assess LLMs’ baseline competence in real-world legal scenarios. You can read the full research paper here: Research Paper on LLMs and Indian Legal Reasoning.
A Dual Approach: Objective and Subjective Assessments
To provide a thorough evaluation, the researchers employed a multi-year benchmark that included both objective and subjective assessments. For objective screens, they curated over 6,200 multiple-choice questions from top national and state exams like the Common Law Admission Test (CLAT) for undergraduate and postgraduate programs, and the Delhi Judicial Services (DJS/DHJS) prelims. These exams are critical gateways for human entry into the legal profession in India.
Beyond multiple-choice questions, the study ventured into the complexities of long-form legal reasoning. It included a lawyer-graded, paired-blinded study of answers from the Supreme Court’s Advocate-on-Record (AoR) exam. The AoR exam is particularly significant as it confers exclusive rights of audience before the Supreme Court of India, demanding a high level of legal drafting and procedural knowledge. The subjective evaluation focused on three papers: Practice & Procedure, Advocacy/Professional Ethics, and Leading Cases, excluding Drafting due to its highly format-critical nature which current text-only LLMs struggle with.
Key Findings: AI’s Strengths and Limitations
The results present a clear picture of AI’s current capabilities in the legal domain. On objective, multiple-choice examinations, frontier LLMs demonstrated remarkable proficiency. Models like Gemini 2.5 Pro consistently cleared historical cutoffs and often matched or even exceeded recent human top-scorer bands. This indicates a robust capacity for short-context legal recall and rule application.
However, this impressive performance did not seamlessly transfer to subjective, long-form legal writing. None of the evaluated models, including the top-performing Gemini 2.5 Pro, surpassed the human topper on long-form reasoning. Grader notes converged on three principal areas where LLMs fell short:
- Deficiencies in Authority Discipline and Doctrinal Rigor: LLMs struggled with consistent adherence to legal citation conventions, often omitting controlling precedents, misciting peripheral authorities, or “manufacturing” authorities without articulating their specific relevance. They tended towards generic legal assertions rather than precise, grounded arguments.
- Proclivity for Irrelevance and Inefficient Content Generation: Models frequently generated “slop”—digressive text that, while grammatically correct, failed to advance a direct answer. This included lengthy paraphrases of basic principles or speculative explorations of tangential scenarios, indicating a difficulty in issue-spotting and prioritization.
- Inapt Voice, Structure, and Rhetorical Framing: Responses often had an “AI-sounding” cadence, with overly broad introductions and meta-framing (e.g., “As an aspiring Advocate on Record…”), despite instructions against such commentary. The structural preferences for long, generalized paragraphs clashed with the expectation for concise, point-wise answers in legal exams.
Also Read:
- Unmasking AI’s Legal Limitations: A Deep Dive into ChatGPT’s Performance in Extracting Principles of Law
- Assessing AI Judges: A New Benchmark for Web Development Quality
Implications for AI in Legal Practice
The study concludes that while AI systems can efficiently assist with tasks requiring legal recall and rule application, they are not yet ready to function as autonomous practitioners, especially in high-stakes environments like the Supreme Court of India. The critical deficits lie in procedural fidelity, precise authority handling, and producing work products that align with judicial expectations. These are filing-critical defects that would lead to direct mark deductions in real-world scenarios.
Therefore, AI is best viewed as a supportive tool for tasks such as searching and verifying authorities, checking consistency across drafts, or cross-referencing case details. Tasks demanding full drafting, independent citation, strategic procedural decisions, or ethical judgment remain firmly within the human domain, requiring vigilant human oversight. This research provides a valuable yardstick for both legal practitioners and AI researchers to understand the current standing of LLMs and guide their evidence-based adoption and future development in the legal field.


