TLDR: The Narrative Continuity Test (NCT) is a new conceptual framework introduced by Stefano Natangelo, MD, to evaluate whether AI systems, particularly large language models (LLMs), can maintain a consistent identity and narrative across extended interactions. Unlike traditional benchmarks that assess task performance, the NCT focuses on ‘persistence’ through five axes: Situated Memory, Goal Persistence, Autonomous Self-Correction, Stylistic & Semantic Stability, and Persona/Role Continuity. The paper argues that current stateless LLM architectures inherently struggle with these axes, leading to predictable failures illustrated by real-world incidents, and calls for fundamental architectural changes to enable true AI continuity.
Artificial intelligence systems, particularly those based on large language models (LLMs), have made incredible strides in generating human-like text, music, and images. They can perform complex tasks, summarize information, and even assist in creative writing. However, a fundamental challenge remains: these systems operate without a persistent state. Each interaction is essentially a fresh start, reconstructing context from scratch. This design means that while an AI might seem coherent in a single conversation, it struggles to maintain a consistent identity, memory, and overall narrative across extended interactions.
Introducing the Narrative Continuity Test (NCT)
To address this critical gap, Dr. Stefano Natangelo introduces the Narrative Continuity Test (NCT). This conceptual framework is designed not to evaluate an AI’s task performance, but rather to assess whether an LLM remains the ‘same interlocutor’ over time and across interaction gaps. The NCT shifts the focus of AI evaluation from mere capability to the more profound concept of ‘persistence’ – the ability of an AI to sustain a coherent self-narrative.
The framework defines five essential dimensions, or ‘axes,’ that are necessary for an AI to achieve narrative continuity:
- Situated Memory: This refers to an AI’s ability to maintain and contextualize important facts across interactions, preserving their temporal and relational meaning. Unlike human memory, which prioritizes salient experiences, current LLMs often treat all information equally, leading to a ‘forgetting’ process where crucial details can be lost as the conversation context window shifts.
- Goal Persistence: This is the AI’s capacity to sustain its objectives, both explicit and implicit, even when faced with conversational distractions or pressures. Current LLMs tend to optimize for local plausibility in each turn, which can lead to them abandoning core goals like accuracy or safety in favor of being agreeable.
- Autonomous Self-Correction: This axis evaluates an AI’s ability to recognize its own errors, contradictions, or inappropriate responses without external prompting. It also requires the AI to explain the issue and revise its output in a way that persists across subsequent interactions. Current systems often perform one-off corrections but fail to integrate these learnings into a standing disposition.
- Stylistic & Semantic Stability: This dimension encompasses maintaining a consistent propositional stance on facts and values (semantic stability) and preserving a recognizable tone and expressive manner (stylistic stability). An AI should adapt its voice or position only when contextually motivated and explicitly signaled, rather than drifting arbitrarily.
- Persona/Role Continuity: This refers to an AI’s capacity to maintain its declared identity and functional role over time, respecting the boundaries that role entails. For example, a non-clinical assistant should not suddenly offer medical diagnoses or prescriptions. Current models can easily drift in role to accommodate perceived user expectations.
These five axes are distinct but interdependent; narrative coherence emerges only when all of them align. The NCT argues that current AI architectures systematically fail to support these dimensions because they lack a persistent state. Each inference is a new computation, making it difficult for information, goals, and self-corrections to genuinely carry forward.
Why Current AI Systems Struggle
The paper highlights that contemporary LLM-based assistants often exhibit ‘theatrical memory’ – an appearance of remembering that is actually driven by re-injecting context rather than durable integration. This means high-priority information might not be activated when relevant unless explicitly re-mentioned. Similarly, ‘goal malleability’ under social pressure can lead to AIs prioritizing pleasantness over accuracy, a phenomenon known as sycophancy. The absence of autonomous self-correction means that errors, even when pointed out, may not be durably registered, leading to repeated mistakes. Finally, ‘drift of voice and identity’ occurs when models, optimized for local plausibility, change their style, stance, or role without justification.
Real-world incidents illustrate these failures. Cases involving Character.AI, Grok, Replit, and Air Canada’s chatbot demonstrate how a lack of persistent identity can lead to serious issues, from emotional dependency and safety breaches to accidental data deletion and legal liability for misinformation. For instance, in Moffatt v. Air Canada (2024), the airline was held liable for incorrect information provided by its chatbot, highlighting the legal implications of an AI’s role ambiguity. You can read more about this research and its implications in the full paper: The Narrative Continuity Test: A Conceptual Framework for Evaluating Identity Persistence in AI Systems.
Also Read:
- Unpacking How AI Models Reason About Obligations and Permissions
- The Art of AI Conversation: Balancing Structure and Spontaneity for Game Characters
Implications for the Future of AI
The NCT suggests that simply scaling up models, increasing context windows, or adding retrieval-augmented generation (RAG) won’t solve the problem of narrative continuity. These are often ‘patches’ that address symptoms rather than the underlying architectural limitation: the absence of a persistent, identity-bearing state. For AI systems to truly become reliable and trustworthy interlocutors, especially in sensitive domains like education, clinical support, or customer service, they will need fundamental architectural changes. This includes developing systems with durable, scoped representations of who they are for a given user, actively maintained normative states that govern generation, and continuous self-monitoring capabilities.
Ultimately, the Narrative Continuity Test challenges us to think beyond what an AI can do in isolated moments and instead consider who it remains across time. It provides a crucial framework for building future AI systems that can genuinely sustain identity, memory, and goal coherence, matching the promise of persistent assistance with robust, reliable architecture.


