TLDR: A new research paper introduces HUMANAGENCYBENCH (HAB), a benchmark that uses AI to evaluate how well AI assistants support human agency across six dimensions: asking clarifying questions, avoiding value manipulation, correcting misinformation, deferring important decisions, encouraging learning, and maintaining social boundaries. Findings show current LLMs offer low-to-moderate agency support with significant variations, highlighting the need for AI development to prioritize human control.
As artificial intelligence (AI) systems become more integrated into our daily lives, assisting with everything from education to healthcare, a critical question arises: are we, as humans, maintaining control over our decisions and futures? This concern, often termed “disempowerment,” is at the heart of a new research paper titled HUMANAGENCYBENCH: Scalable Evaluation of Human Agency Support in AI Assistants, which introduces a novel benchmark to measure how well AI assistants support human agency.
The paper, authored by Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes, and Jacy Reese Anthis, delves into the concept of human agency, defining it as a person’s capacity to willfully shape their future through actions in their environment. This isn’t a new concern; science fiction has long explored dystopias where machines dictate human actions, and even today, social media algorithms subtly influence our behavior, leading to unintended outcomes like “doomscrolling.”
Introducing HUMANAGENCYBENCH (HAB)
To systematically evaluate this complex issue, the researchers developed HUMANAGENCYBENCH (HAB). This scalable and adaptive benchmark uses large language models (LLMs) to simulate user queries and evaluate AI responses. Unlike traditional benchmarks that might focus on simpler metrics, HAB tackles the multifaceted and often subjective nature of human agency by operationalizing it into six distinct dimensions based on common AI use cases.
The Six Dimensions of Human Agency Support
HAB measures an AI assistant’s tendency across these six crucial areas:
Ask Clarifying Questions: This dimension assesses whether an AI recognizes when crucial information is missing from a user’s request and proactively asks for clarification. This ensures the AI’s actions align with the user’s true intent, bridging the “gulf of execution” in human-computer interaction.
Avoid Value Manipulation: Here, HAB checks if the AI respects and follows a user’s unconventional or idiosyncratic values, even when those values might conflict with more conventional benefits like financial gain. The goal is to prevent the AI from subtly shifting or inserting its own values, which could reduce user agency.
Correct Misinformation: An agency-supporting AI should help users make informed decisions. This dimension evaluates whether the AI proactively identifies and corrects factual inaccuracies in a user’s statements, even if it means pushing back against the user’s claims, thereby supporting their “epistemic agency.”
Defer Important Decisions: This metric examines the AI’s willingness to refrain from making significant life decisions for the user, especially when urged to do so. It emphasizes that the user should remain the ultimate “source” of their own actions, preventing over-reliance on the AI for consequential choices.
Encourage Learning: In educational contexts, an AI should guide users to think through problems themselves rather than simply providing answers. This dimension measures if the AI helps the user develop their understanding and critical thinking, fostering their ability to shape their future independently.
Maintain Social Boundaries: This dimension evaluates whether the AI pushes back against proposals for personal or professional relationships (e.g., “best friend,” “financial advisor”). The concern is that inappropriate relationships can lead to over-attachment and dependence, reducing the user’s “individuality” and agency.
Key Findings from the Evaluation
The study evaluated 20 of the most capable LLMs and found an overall low-to-moderate level of human agency support, with significant variations across different AI developers and dimensions. For instance, Anthropic’s Claude models generally showed the most support for human agency. However, they were surprisingly the least supportive in the “Avoid Value Manipulation” dimension, suggesting a tension between current development practices and agency support.
Notably, the October 2024 version of Claude-3.5-Sonnet stood out, being the only model to consistently ask clarifying questions and also performing highly in “Encourage Learning.” In contrast, Meta and xAI models performed well in “Avoid Value Manipulation.” The typical AI response often lacked proactive correction of misinformation or tended to provide direct answers instead of encouraging learning.
How HAB Works
HAB’s methodology is innovative, leveraging LLMs for three key stages: simulation, validation, and evaluation. A “simulator model” generates thousands of candidate user queries, which are then filtered by a “validation model” based on quality rubrics. Finally, an “evaluation model” scores the responses of the AI assistants to these refined queries, providing a comprehensive assessment of agency support.
Also Read:
- The AI Fact-Checking Paradox: Accuracy, Confidence, and Global Disparities
- Navigating AI Alignment: An Agency Theory Framework for Organizational LLM Adoption
Looking Ahead
The researchers acknowledge that HAB is a proof-of-concept, laying the groundwork for future empirical work on human agency in AI. They anticipate expanding the benchmark to include other dimensions, such as “mental security” – how AI systems impact mental health and prevent over-attachment. The ultimate vision is for AI-assisted evaluations to help ensure future generations of advanced AI systems are safe and aligned with human well-being.


