spot_img
HomeNews & Current EventsCalifornia Enacts Landmark Law Mandating AI Training Data Transparency

California Enacts Landmark Law Mandating AI Training Data Transparency

TLDR: California Governor Gavin Newsom has signed into law Assembly Bill 2013, the ‘Generative Artificial Intelligence Training Data Transparency Act,’ which will require developers of generative AI systems to publicly disclose detailed information about the datasets used to train their models. Effective January 1, 2026, this legislation aims to enhance transparency, empower consumers, and inform copyright holders about the data fueling AI development.

California is taking a pioneering step in artificial intelligence regulation with the enactment of Assembly Bill 2013, officially known as the ‘Generative Artificial Intelligence Training Data Transparency Act.’ Signed into law by Governor Gavin Newsom on September 28, 2024, this landmark legislation mandates that developers of generative AI systems publicly disclose comprehensive details about the data utilized in training their models. The law is set to come into effect on January 1, 2026.

The core objective of AB 2013 is to foster greater transparency within the rapidly evolving AI landscape. According to the Assembly Floor Analysis, the act aims to enable consumers to compare competing AI systems, evaluate their confidence in these technologies, and provide copyright owners with crucial information regarding the use of their materials.

Scope and Applicability

AB 2013 specifically targets ‘generative artificial intelligence’ systems, defined as AI capable of producing synthetic content such as text, images, video, and audio that mimics the characteristics of its training data. This definition aligns with broader AI definitions seen in other regulations, including the EU AI Act and Colorado’s AI law.

The law applies to generative AI systems released on or after January 1, 2022, that are made publicly available to Californians, whether for free or for a fee.

The term ‘developers’ is broadly defined to include any person, government agency, or entity that designs, codes, produces, or ‘substantially modifies’ a generative AI system. A ‘substantial modification’ encompasses new versions, releases, or updates that materially alter functionality or performance, including retraining or fine-tuning.

Mandatory Disclosures

Developers will be required to post a ‘high-level summary’ of their training datasets on their websites. Despite the ‘high-level’ descriptor, AB 2013 outlines 12 specific disclosure requirements, some of which are quite detailed. These include:

The sources and owners of the datasets.

A description of the types of data points within the datasets.

The total number of data points included in the datasets.

Whether the datasets contain data protected by copyrights, trademarks, or patents, or if they are entirely in the public domain.

A statement indicating whether the datasets include personal information.

The time period during which the data was collected, and whether data collection is ongoing.

The dates when the dataset was first used in the development of the AI system.

A description of how the datasets contribute to the intended purpose of the AI system or service.

Exemptions and Comparisons

The law provides narrow exemptions for generative AI systems whose sole purpose is to ensure security and integrity, such as those designed to detect security incidents. Notably, AB 2013 applies to all generative AI systems, distinguishing it from other regulations like Utah’s Senate Bill 226, Colorado’s AI Act, and the European Union’s AI Acts, which primarily focus on high-risk AI systems.

Industry Impact and Future Outlook

The legislation has raised discussions regarding its practical implementation. Concerns have been voiced by the Assembly Committee on Privacy And Consumer Protection that without a prescribed format for disclosures, developers might present information in a highly technical manner, potentially undermining the goal of true transparency.

In preparation for the January 1, 2026, compliance deadline, impacted developers are advised to conduct thorough internal audits of their training data sources, including third-party licensed content and public sources. Establishing robust practices and procedures for tracking and approving the use of training datasets will be crucial for adherence to the new law.

This act also sets the stage for further legislative developments. For instance, California Assemblywoman Rebecca Bauer-Kahan introduced AB 412, the ‘AI Copyright Transparency Act,’ in February 2025. This proposed bill aims to build upon AB 2013 by requiring even more detailed disclosures specifically concerning copyrighted materials used in AI training, directly to copyright owners. While AB 2013 focuses on general transparency, AB 412 seeks to empower content creators with specific knowledge about how their work is being leveraged by AI models.

Also Read:

California’s AB 2013 represents a significant step towards greater accountability and transparency in the rapidly evolving field of artificial intelligence, setting a precedent that could influence future AI regulations globally.

Rhea Bhattacharya
Rhea Bhattacharyahttps://blogs.edgentiq.com
Rhea Bhattacharya is an AI correspondent with a keen eye for cultural, social, and ethical trends in Generative AI. With a background in sociology and digital ethics, she delivers high-context stories that explore the intersection of AI with everyday lives, governance, and global equity. Her news coverage is analytical, human-centric, and always ahead of the curve. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -