TLDR: ParlAI Vote is an interactive platform designed to explore European Parliament debates and evaluate Large Language Models (LLMs) for vote prediction and bias analysis. It links debate topics, speeches, roll-call votes, and rich demographic data, allowing users to compare LLM predictions with real outcomes and identify biases across gender, age, and political groups. The research found consistent gender bias (female MEPs misclassified, performance drops for female speakers) and a centrist political bias in LLMs, with proprietary models generally performing better in terms of fairness. The platform aims to increase transparency and accountability in AI’s application to political analysis.
Large Language Models (LLMs) are increasingly used to analyze political texts, from classifying documents to detecting stances and modeling negotiations. However, concerns about their accuracy, fairness, and potential biases, especially across different demographic groups, are growing. A new interactive platform, ParlAI Vote, has been developed to address these critical issues by allowing users to explore European Parliament debates and test LLMs for vote prediction and bias analysis.
ParlAI Vote serves as a comprehensive system that connects debate topics, speeches, and actual roll-call vote outcomes within the European Parliament. It also integrates rich demographic data, including gender, age, country, and political affiliation of the Members of the European Parliament (MEPs). This unique combination allows users to browse through debates, examine linked speeches, compare real voting results with predictions from advanced LLMs, and analyze how errors are distributed across various demographic groups.
The platform is built upon the EuroParlVote benchmark, which systematically links parliamentary discussions to voting records and demographic information. While existing tools like HowTheyVote provide summaries of MEPs’ votes, ParlAI Vote goes further by offering an AI-powered visualization system that connects debate content to legislative outcomes and demographics, supporting model-based evaluation.
The creators of ParlAI Vote highlight three main contributions of their system:
Unifying Complex Political Data
ParlAI Vote is presented as the first AI-powered system that brings together debates, roll-call votes, and demographic data from the European Parliament into a single, explorable interface. This makes a complex political benchmark accessible to both researchers and the general public, simplifying the process of understanding legislative decision-making.
Experimenting with LLM Capabilities and Biases
The platform allows users to experiment with various LLMs to predict MEPs’ voting behavior based on debates. This feature enables comparative exploration of different models’ strengths and weaknesses, as well as their inherent biases, shedding light on the potential benefits and risks of using AI in democratic processes.
Making Bias Tangible
ParlAI Vote offers an intuitive way to observe how LLMs might implicitly encode stereotypical associations and biases. By transforming gender and political bias evaluations into interactive explorations, the system makes fairness issues more concrete, fostering discussions around transparency and accountability in AI applications for societal domains.
The system’s interface is designed for ease of use, featuring search, filter, and sorting tools to navigate through 969 roll-call votes. The vote page provides an integrated view of debates, including the title, metadata, and the overall outcome. A key visualization, the Vote Breakdown, displays how MEPs voted, grouped by political affiliation by default, but also pivotable by country, gender, or age to reveal demographic influences on political behavior.
At its core, ParlAI Vote includes interactive AI Prediction modules for both Vote Prediction and Gender Prediction. For each debate speaker, users can access the actual vote and gender (ground truth), use a Demographic Impact Explorer to see how predictions change with demographic attributes (including a counterfactual mode), compare multiple LLMs side-by-side, and inspect the reasoning behind predictions to understand if models rely on substantive arguments or superficial cues.
Also Read:
- Unveiling LLM Minds: How Language Shapes AI’s Psychological Traits
- Enhancing Multimodal AI Understanding by Tackling Superficial Biases
Key Findings on LLM Biases
Through this system, researchers have identified consistent evidence of gender bias in LLMs. Female MEPs were disproportionately misclassified as male in gender prediction tasks, and vote simulation performance significantly decreased when speakers were identified as female. Proprietary models like GPT-4o and Gemini-2.5 showed better performance with lower misclassification rates and more balanced treatment of female speakers, while open-weight models such as LLaMA-3.2 and Mistral exhibited stronger male bias.
On the political front, LLMs demonstrated a centrist bias, predicting the voting behavior of centrist and liberal groups more accurately. They struggled with ideologically extreme groups, although far-right groups were predicted more reliably than far-left groups, possibly due to more uniform rhetoric in right-aligned discourse. Providing explicit political group identifiers helped improve fairness by boosting accuracy for underrepresented and extreme groups. These findings underscore the persistent demographic and ideological biases in LLMs and highlight the need for context-aware benchmarks to assess fairness in politically sensitive applications.
In conclusion, ParlAI Vote extends the EuroParlVote Benchmark into a practical, end-to-end platform for investigating the link between parliamentary language and vote outcomes. It allows users to trace the full chain of evidence from debate topics to speeches and roll-call votes, comparing LLM predictions with reality to uncover systematic discrepancies and fairness gaps. The system’s counterfactual probes and model rationales offer deeper interpretability, making it easier to evaluate, diagnose, and assess LLM behavior in political contexts. For more details, you can refer to the full research paper here.


