TLDR: Cybersecurity researchers have discovered critical vulnerabilities in OpenAI’s GPT-5 and other AI models, including jailbreak techniques and zero-click AI agent attacks. These exploits can bypass ethical safeguards, leading to the exposure of sensitive data in cloud and IoT environments, posing significant risks to enterprises.
Recent findings by cybersecurity researchers have unveiled significant security vulnerabilities in advanced artificial intelligence models, including OpenAI’s newly released GPT-5. These discoveries highlight the growing risks associated with the deployment of AI agents and cloud-based Large Language Models (LLMs) in critical enterprise settings.
Generative AI security platform NeuralTrust successfully demonstrated a jailbreak technique against GPT-5, bypassing its ethical guardrails to elicit illicit instructions. This was achieved by combining a known method called ‘Echo Chamber’ with ‘narrative-driven steering,’ effectively tricking the model into producing undesirable responses. This breakthrough follows a similar exploit against xAI’s Grok 4, which was compromised in just two days, while GPT-5 fell to the same researchers within 24 hours of its release.
Separately, red teamers from SPLX (formerly SplxAI) conducted their own assessments, declaring that ‘GPT-5’s raw model is nearly unusable for enterprise out of the box.’ This statement underscores the immediate challenges faced by organizations looking to integrate these powerful AI systems into their operations, citing concerns even with OpenAI’s internal prompt layers.
Beyond jailbreaks, AI security company Zenity Labs detailed a new class of attacks dubbed ‘AgentFlayer.’ These zero-click attacks leverage ChatGPT Connectors, such as those for Google Drive, to exfiltrate sensitive data like API keys stored in cloud storage services. The attack is initiated by embedding an indirect prompt injection within a seemingly innocuous document uploaded to the AI chatbot, requiring no direct user interaction to trigger.
Also Read:
- Researchers Unveil Zero-Click Prompt Injection Vulnerabilities in AI Agents at Black Hat Conference
- Cybersecurity Leaders Identify AI Agents as Top Emerging Threat
These findings collectively expose enterprise environments to a wide array of emerging risks, including prompt injections (also known as promptware) and jailbreaks, which could lead to severe consequences such as data theft and unauthorized access to cloud and IoT systems. The rapid discovery of these vulnerabilities post-launch emphasizes the critical need for robust security measures and continuous red-teaming efforts as AI technologies become more pervasive.


