spot_img
HomeResearch & DevelopmentPublic Engagement: A New Model for AI Safety and...

Public Engagement: A New Model for AI Safety and Risk Evaluation

TLDR: A new research paper introduces the concept of cooperative public AI red-teaming exercises as a vital approach to address the ‘responsibility gap’ in AI safety. It highlights how current in-house AI evaluations are often insufficient and proposes involving diverse public participants in adversarial testing. The paper details three pilot exercises—the NIST ARIA National Red Teaming Pilot, the CAMLIS 2024 Demonstrator, and the IMDA Singapore AI Safety Red Teaming Challenge—showcasing how these initiatives can effectively identify AI vulnerabilities, leverage public diversity for authentic risk assessment, support SMEs, and provide a scalable model for responsible AI development and governance.

Artificial intelligence (AI) systems offer immense potential but also carry significant risks. Ensuring their safety and security requires rigorous and ongoing evaluation, a task that current methods often struggle with, especially as AI becomes more sophisticated and deployed in critical sectors like healthcare and education.

The responsibility for AI safety isn’t solely with the developers; it’s a shared duty that includes AI application developers and public entities. However, expert AI evaluations, known as red teaming, are predominantly conducted by AI labs themselves. This can lead to a ‘responsibility gap,’ where external parties, including the public, have limited direct involvement in assessing AI risks during its design, development, or deployment stages.

A new research paper, “Ask What Your Country Can Do For You: Towards a Public Red Teaming Model”, proposes an innovative solution: cooperative public AI red-teaming exercises. This approach aims to bridge the responsibility gap by involving a broader, more diverse group of participants in identifying potential harms and vulnerabilities in AI systems.

What is AI Red Teaming?

AI red teaming is a specialized discipline within Testing, Evaluation, Validation, and Verification (TEVV). It adapts traditional cybersecurity techniques to uncover safety and security weaknesses unique to AI. The main goal is to achieve a comprehensive understanding of all potential harms an AI model or system could cause, a concept known as ‘coverage.’ Red teamers often use sophisticated taxonomies, such as MITRE ATLAS or NIST AI RMF, to categorize risks and systematize their findings.

However, existing red teaming practices face criticism. In-house red teaming might not be sufficiently objective, often favoring developer-defined harm classifications over those articulated by users, and teams may lack diversity. Additionally, many risk typologies are too general to be practically applied in specific deployment contexts, which have unique norms and values. This highlights the need for public, third-party evaluations.

Piloting Public Red Teaming in Action

The paper discusses the experimental design and findings from three significant public third-party AI red-teaming exercises:

1. NIST ARIA National Red Teaming Pilot Exercise: Held virtually in late 2024, this exercise partnered Humane Intelligence with a national metrology institution. It was open to all US residents aged 18 and older, aiming for broad public participation and a diverse, multidisciplinary red team. Participants, including AI researchers, cybersecurity professionals, ethicists, and policymakers, stress-tested three large language models (LLMs) in scenario-based challenges. They sought to induce violative outcomes, such as harmful meal planning advice or factual errors in travel advisories, demonstrating the benefits of a large and diverse red team.

2. Public Red-Teaming Demonstrator Exercise at CAMLIS 2024: This in-person event, organized by Humane Intelligence and a US government partner in October 2024, focused on assessing AI model vulnerabilities and the utility of the NIST AI 600-1 framework for evaluating Generative AI risks. Red teamers, selected from top performers in the NIST ARIA pilot, were divided into teams to induce violative outputs from office productivity software employing GenAI models. This exercise provided insights into the strengths and weaknesses of intensive in-person operations compared to distributed, longer-term engagements.

3. IMDA Singapore AI Safety Red Teaming Challenge: Conducted in the final months of 2024, this was the world’s first multilingual and multicultural AI safety red-teaming exercise, focusing on the Asia-Pacific region. Partnering Humane Intelligence with the Singapore Infocomm Media Development Authority (IMDA), the hybrid (in-person and virtual) exercise developed a methodology for evaluating LLMs for context-specific harms across various languages and cultures. It demonstrated how such exercises can provide valuable data for creating new testing benchmarks and identifying areas for further development, which the Singapore government is actively pursuing.

Also Read:

The Path Forward

These exercises demonstrate that public red teaming can significantly empower public entities to meet their responsible AI obligations. By drawing on the diversity of an entire national jurisdiction or region, these initiatives provide authentic insights into the norms, values, and discourses relevant to different deployment contexts. Partnering with public entities ensures that these efforts genuinely serve the public interest, allowing civil society to maintain influence over AI development and deployment.

Furthermore, these exercises offer meaningful red-teaming support to Small and Medium-sized Enterprises (SMEs) that might otherwise lack the resources to fulfill their responsible AI duties. Ultimately, the paper argues that this scalable model can be adopted by most AI-developing states or regions to ensure that AI systems developed for their communities are safe and beneficial.

Rhea Bhattacharya
Rhea Bhattacharyahttps://blogs.edgentiq.com
Rhea Bhattacharya is an AI correspondent with a keen eye for cultural, social, and ethical trends in Generative AI. With a background in sociology and digital ethics, she delivers high-context stories that explore the intersection of AI with everyday lives, governance, and global equity. Her news coverage is analytical, human-centric, and always ahead of the curve. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -