TLDR: A comprehensive review of 25 fMRI-based studies (2023-2025) reveals strong evidence that large language models and human brains converge on similar internal representations of the world. The research supports the Platonic Representation Hypothesis, indicating that as models scale and improve, they approximate a shared statistical model of reality. It also confirms the Intermediate-Layer Advantage, showing that mid-depth layers of language models exhibit the strongest alignment with brain activity, suggesting these layers encode the most robust and generalizable features.
A fascinating area of research at the intersection of neuroscience and Artificial Intelligence (AI) is exploring whether human brains and advanced language models develop similar internal ways of representing the world. Recent studies suggest that these two complex systems might indeed converge on shared abstract structures, offering compelling insights into how both biological and artificial intelligence process information.
A new research paper, Brain–Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage, reviews 25 fMRI-based studies published between 2023 and 2025 to directly examine two prominent hypotheses in this emerging field. The authors, Ángela López-Cardona, Sebastián Idesis, Mireia Masias-Bruns, Sergi Abadal, and Ioannis Arapakis, delve into whether models, as they grow in scale and capability, begin to mirror the brain’s understanding of reality, and if certain layers within these models are particularly adept at capturing these brain-like representations.
The Platonic Representation Hypothesis
One of the central ideas explored is the Platonic Representation Hypothesis. This theory suggests that as neural networks become larger and more sophisticated, their internal representations of the world start to converge towards an ideal, abstract statistical model of reality. The hypothesis posits that if both artificial models and biological brains are trying to understand the same underlying structure of the world, then their internal representations should naturally become more similar.
The review found several pieces of evidence supporting this hypothesis. Firstly, larger and more capable models generally show stronger alignment with brain activity. This suggests that as models improve in performance, they develop representations that are more akin to those found in the human brain. Secondly, models trained on a wider variety of tasks tend to align more closely with brain activity. This is because diverse training forces models to learn more general-purpose representations, which are also characteristic of how the brain processes information. Lastly, models trained on multiple modalities (e.g., both images and text) demonstrate stronger alignment with brain activity compared to unimodal models. This cross-modal training encourages models to learn abstract structures that are not tied to a single type of input, much like the brain integrates information from different senses.
The Intermediate-Layer Advantage
The second key hypothesis investigated is the Intermediate-Layer Advantage. This concept proposes that the middle layers of language models often encode richer, more generalizable features compared to the initial or final layers. While early layers might focus on basic features and final layers might become overly specialized for specific training objectives, the intermediate layers seem to strike a balance, capturing more robust and abstract linguistic and semantic information.
The review provides strong evidence that these intermediate layers are indeed where the strongest alignment with brain activity occurs. Across various model architectures – including decoder-based, encoder-decoder, and encoder-based text models, as well as audio, vision, and video models – the middle layers consistently showed the best correspondence with human brain representations. This finding suggests that the way language models build up complex representations, moving from simpler features to more abstract concepts, closely parallels the hierarchical processing observed in the brain.
Also Read:
- PRISM: A Text-Based Breakthrough in Visual Brain Decoding
- Large Connectome Model: A New AI Approach for Brain Imaging and Clinical Diagnosis
Looking Ahead
The converging evidence from these recent studies offers significant support for both the Platonic Representation Hypothesis and the Intermediate-Layer Advantage. It highlights that as AI models become more advanced, they are not just performing tasks better, but are also developing internal structures that resonate with how the human brain understands the world. While the field acknowledges limitations due to the diverse methodologies across studies, these qualitative patterns provide a robust theoretical framework for future research into the shared representational structures between human and artificial intelligence.


