TLDR: This research paper introduces a structured hierarchy of identifiability criteria for causal queries within ‘causal abstractions’ – simplified representations of complex causal systems. It defines and analyzes various notions of identifiability, such as Identifiability through Graphs (IG) and Identifiability by Common Do-Calculus (ICD), clarifying their relationships and providing tools to reason about causal effects when full causal knowledge is absent. The paper also highlights an open conjecture regarding the precise relationship between IG and ICD.
Understanding cause-and-effect relationships is fundamental in many fields, from medicine to economics. Traditionally, identifying the effect of a treatment or intervention from observational data requires a complete and accurate causal diagram, which maps out all the relationships between variables. However, in real-world scenarios, especially with complex or high-dimensional data, having such a fully specified diagram is often impossible.
To address this challenge, recent research has turned to ‘causal abstractions.’ These are simplified representations that capture some, but not necessarily all, of the underlying causal information. Think of it like using a simplified map for a complex city – it might not show every alleyway, but it still helps you navigate the main roads.
A new research paper, “Identifiability in Causal Abstractions: A Hierarchy of Criteria”, formalizes these causal abstractions as collections of possible causal diagrams. The core focus of the paper is on ‘identifiability’ – determining when a causal query (like the effect of X on Y) can be uniquely figured out from observational data, even when you only have an abstract, incomplete understanding of the causal structure.
Defining Identifiability in Abstractions
The authors introduce and formalize several distinct criteria for identifiability within these collections of causal diagrams, organizing them into a clear hierarchy. This helps to understand what can be identified given different levels of causal knowledge.
One key concept is ‘Identifiability through Graphs (IG)’, which means a causal query can be identified if the same estimation formula works for every possible causal diagram within the abstraction. A related, but stronger, notion is ‘Identifiability through Graphs knowing P*’ (IGP), which assumes perfect knowledge of the observational data distribution.
The paper also explores ‘methodological differences’ in how identifiability can be established. ‘Identifiability by Common Do-Calculus (ICD)’ requires finding a single, universal proof using Pearl’s do-calculus rules that applies to all diagrams in the abstraction. A more specific approach is ‘Identifiability by Common Graphical Criterion (ICGC)’, where a specific graphical rule (like the backdoor or frontdoor criterion) holds true across all diagrams in the collection.
Relationships and Simplifications
The research establishes clear relationships between these notions: if something is identifiable by a common graphical criterion, it’s also identifiable by common do-calculus. And if it’s identifiable by common do-calculus, it’s identifiable through graphs. Finally, if it’s identifiable through graphs, it’s also identifiable through graphs knowing the true data distribution. This creates a hierarchy from the most restrictive (ICGC) to the least (IGP).
A significant finding is that to determine identifiability in a collection of graphs, one only needs to consider the ‘maximal’ graphs within that collection – those that are not subgraphs of any other graph in the collection. This can significantly simplify the computational challenge, especially when the collection of possible diagrams is very large.
Also Read:
- Unlocking Causal Reasoning in AI: A New Approach with Logic Programs
- Smarter AI Decisions: New Methods for Dynamic Abstraction in Monte Carlo Tree Search
The Open Question
Despite these advancements, the paper highlights a pivotal open challenge: the relationship between ‘Identifiability through Graphs (IG)’ and ‘Identifiability by Common Do-Calculus (ICD)’. It’s currently unknown whether there exists a scenario where a causal query is identifiable through graphs but cannot be identified by a single common do-calculus proof. Proving or disproving this conjecture would further clarify the boundaries of causal identifiability in abstract settings.
This framework provides a valuable tool for researchers and practitioners to reason about causal identifiability when full causal knowledge is unavailable, paving the way for more robust and applicable causal inference methods in complex systems.


