spot_img
HomeResearch & DevelopmentUnlocking Brain Language: How AI Models Mirror Our Multimodal...

Unlocking Brain Language: How AI Models Mirror Our Multimodal Understanding

TLDR: A study by Carnegie Mellon University researchers Kateryna Shapovalenko and Quentin Auster investigates which layers of pre-trained AI models, wav2vec2 (speech) and CLIP (vision-language), best align with human brain activity (EEG) during speech perception. They found that mid-sequence layers of both models showed the most consistent alignment, and a progressive summation strategy for combining layer embeddings proved more robust for generalization than concatenation. The research suggests that combining multimodal, layer-aware representations could advance our understanding of how the brain processes language as a rich, multimodal experience, despite challenges in generalization.

Our brains do more than just hear sounds when we listen to language; they conjure images, memories, and a rich tapestry of associations. This complex, layered process of understanding language, moving from simple acoustics to deep, multimodal meaning, is a fascinating area of research. Inspired by this, a recent study delves into how well different layers of advanced AI models can mirror this intricate brain activity.

Building on previous work that aligned brain signals (EEG) with averaged speech embeddings, researchers Kateryna Shapovalenko and Quentin Auster from Carnegie Mellon University posed a deeper question: Which specific layers within pre-trained AI models best reflect the brain’s layered processing during speech perception? To explore this, they compared embeddings from two prominent models: wav2vec2, which specializes in encoding sound into language, and CLIP, a multimodal model that connects words to images, potentially capturing the visual associations we form when we hear language.

The study utilized EEG data recorded from participants as they listened to a chapter of “Alice in Wonderland.” The researchers systematically evaluated how embeddings from different layers of wav2vec2 and CLIP aligned with brain activity. They tested three distinct strategies for combining these embeddings: using individual layers, progressively concatenating embeddings from successive layers, and progressively summing them. The goal was to see which approach offered the strongest alignment with the brain’s electrical signals.

The methodology involved several key steps. First, EEG recordings were meticulously pre-processed, including noise removal and the extraction of both time- and frequency-domain features. Simultaneously, audio from the story was passed through the wav2vec2 model, and transcriptions were processed by CLIP’s text encoder. This yielded 13 distinct embedding layers from each model, spanning from low-level acoustic features to higher-level linguistic and visual representations. To manage the complexity, dimensionality reduction techniques like PCA were applied to these embeddings.

The researchers then employed ridge regression to evaluate the alignment between these AI model embeddings (predictors) and the preprocessed EEG features (targets). In the single-layer regression, they found that individual layers showed modest predictive power for EEG responses. While training correlations were high, test results often showed negative R² values, indicating a struggle to generalize to new data. However, topographic maps suggested that certain early and mid-layers, like wav2vec2_layer_0 and clip_layer_1, might capture more salient features aligned with EEG activity.

When exploring aggregation strategies, progressive concatenation, where embeddings from successive layers were combined, improved training performance but led to even worse generalization on test data. This suggested that while more layers captured more variance in the training data, they also introduced noise or redundancy that hindered performance on unseen samples.

In contrast, progressive summation, which involved summing embeddings layer by layer while maintaining dimensionality, showed more promising results. This method not only improved training performance but also led to an increase in test performance, indicating better generalization. The findings suggested that early-to-mid layers (around layers 0-7 or 0-8) encoded features most aligned with EEG representations, while deeper layers might introduce more abstract information not directly reflected in the brain signals. The summed layers also activated more diverse and distributed brain regions, as visualized in topographic maps.

While the study achieved strong correlations on training data, a significant challenge remained in generalizing these findings to new data, a common hurdle in brain decoding tasks. The researchers also attempted a contrastive decoding setup but faced convergence issues, leaving this for future investigation.

Also Read:

In conclusion, this research highlights that mid-sequence layers of both wav2vec2 and CLIP models offer the most consistent alignment with EEG signals during speech perception. This suggests these layers strike a meaningful balance between low-level acoustic and high-level linguistic features. The progressive summation strategy proved more robust for generalization compared to concatenation, though fully bridging the gap in brain-to-audio alignment remains an open challenge. Future work will focus on subject-invariant architectures, larger datasets, and alternative embedding spaces to enhance generalization. You can read the full research paper here: Aligning Brain Signals with Multimodal Speech and Vision Embeddings.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -