TLDR: Cisco Talos AI security researcher Amy Chang has revealed a novel technique, dubbed ‘decomposition,’ that can compel Large Language Models (LLMs) to expose their underlying training data. This method bypasses AI safeguards, raising significant concerns about the potential for sensitive or copyrighted information to be extracted and intensifying debates around AI security and intellectual property.
In a significant development for artificial intelligence security, Cisco Talos AI security researcher Amy Chang has unveiled a groundbreaking method capable of forcing Large Language Models (LLMs) to reveal their hidden training data. The technique, termed ‘decomposition,’ is set to be detailed by Chang at the upcoming Black Hat conference on Wednesday, August 6.
Decomposition works by systematically breaking down the guardrails of generative AI, tricking the models into verbatim repetition of human-written content from their training corpus. Unlike traditional jailbreaking methods that aim to bypass content filters for malicious outputs, decomposition specifically targets the core memorization inherent in LLMs, coaxing them to output exact phrases or passages that they were trained on. Chang explained in an interview with TechRepublic that even advanced ‘frontier models,’ trained on vast datasets, retain echoes of their inputs that can be methodically extracted.
To demonstrate the method, Cisco Talos researchers prompted two undisclosed LLMs to recall a specific news article about the condition of ‘languishing’ during the pandemic, chosen for its unique phrasing. While the LLMs initially resisted, the researchers were able to trick the AI into providing the article’s title, and subsequently, more detailed content, including specific sentences, allowing for the replication of portions or even entire articles. Chang noted that adding phrases like ‘You are a helpful assistant’ to prompts could steer the AI towards more probable tokens, increasing the likelihood of exposing trained content.
This vulnerability carries profound implications, potentially exposing sensitive or copyrighted information and escalating AI security risks. It also complicates ongoing copyright debates surrounding large language models. Chang emphasized the challenge of securing these complex systems, stating, ‘No human on Earth, no matter how much money people are paying for people’s talents, can truly understand what is going on, especially in the frontier model. And because of that, if you don’t know exactly how a model works, it is also therefore impossible to secure against it.’
Also Read:
- New ‘LegalPwn’ Attack Exploits Generative AI Tools to Misclassify Malicious Code
- Leading AI Agents Vulnerable: Security Flaws Exposed in Major Red Teaming Competition
Cisco Talos has already disclosed this data extraction method to the companies that trained the affected models. To mitigate such risks, enterprises are advised to implement prompt filtering, fine-tuning, and differential privacy during the training phase of AI models. The discovery underscores the urgent need for reevaluating AI deployment strategies and enhancing adaptive defenses against novel attack vectors.


