spot_img
HomeResearch & DevelopmentEnhancing Speech LLMs: A Dual-Channel Approach to Overcome Forgetting...

Enhancing Speech LLMs: A Dual-Channel Approach to Overcome Forgetting and Modality Gaps

TLDR: This research introduces a cross-modal knowledge distillation framework to address catastrophic forgetting and modality inequivalence in speech large language models (LLMs). By using both text-to-text and speech-to-text distillation channels, the framework transfers knowledge from a text-based teacher model to a speech LLM, preserving textual knowledge, improving cross-modal alignment, and enhancing reasoning in speech-based interactions. Experiments validate its effectiveness in improving overall capabilities across various tasks.

The field of large language models (LLMs) has seen incredible advancements, particularly in their ability to handle multiple types of data, known as multimodal capabilities. Voice interaction has emerged as a key application, with cutting-edge models now enabling real-time spoken dialogue, offering more natural and flexible experiences than traditional text-based systems. Building on this trend, many researchers are extending powerful text LLMs into the speech domain, creating models capable of both understanding and generating speech. These models typically integrate a speech encoder and a modality adapter layer with the existing text LLM, allowing them to convert audio into text and respond to spoken queries.

However, introducing speech capabilities often leads to a significant drop in performance. Studies have observed that even with the same underlying meaning, responses from speech input can differ substantially from those generated from text input, indicating that speech interaction often lags behind text-based performance. This issue can be understood as a form of catastrophic forgetting, a problem in continual learning where models, when adapted for new tasks like speech, tend to overemphasize new learning and, in the process, forget previously acquired language knowledge. This severely undermines the potential of extending text-based LLMs into the speech modality.

To mitigate catastrophic forgetting, some strategies involve freezing the core text LLM and only training the speech encoder and adapter layers. While this helps preserve the text model’s inherent intelligence, reasoning abilities can still decline when operating in the speech modality. This performance degradation is attributed to two main factors: catastrophic forgetting and modality inequivalence. Modality inequivalence arises from poor alignment between speech and text modalities, leading to performance drops. For instance, some models show no forgetting with text input but significant declines with speech input. There’s often a conflict between retaining high-quality acoustic features and accurately capturing textual meaning, especially under constraints like low bitrate or limited model capacity. Bridging this gap between acoustic and semantic feature spaces is a major challenge.

A recent study addresses these challenges by proposing a novel cross-modal knowledge distillation framework. This framework systematically evaluates catastrophic forgetting and modality inequivalence in speech LLMs. It introduces a method that leverages both text-to-text and speech-to-text channels to transfer knowledge from a more capable text-based teacher model to a speech LLM, which acts as the student.

Key Contributions of the Framework

The framework’s contributions are threefold. First, it provides a general validation by systematically evaluating multiple mainstream open-source speech LLMs for these issues. By comparing performance with text queries versus speech queries, and comparing speech LLMs to their text base models, the study confirms that these problems are widespread.

Second, it introduces the cross-modal knowledge distillation strategy. This approach treats the speech LLM as a student and its text LLM backbone as a teacher, leveraging the teacher’s superior reasoning and knowledge. The distillation process occurs in two stages:

  • Text-to-Text Distillation: The speech LLM is trained to mimic the teacher’s outputs when given text inputs. This helps alleviate catastrophic forgetting by reinforcing textual knowledge.

  • Speech-to-Text Distillation: Text data is converted into synthetic speech using a Text-to-Speech (TTS) system and fed to the speech LLM. The teacher’s text outputs still serve as the supervision signal, strengthening the semantic alignment between speech and text modalities.

Unlike previous works that focused mainly on acoustic information and audio analysis, this framework is the first to apply explicit cross-modal semantic knowledge distillation directly to dialogue tasks (both text-to-text and speech-to-text). This targets higher-level knowledge transfer and reasoning enhancement crucial for interactive applications.

Third, the experimental validation demonstrates the effectiveness of this method. Using a relatively small dataset of about 60,000 samples, the approach significantly improves the overall capabilities of speech LLMs, such as Qwen2.5-Omni. After distillation, the model shows clear improvements in general ability tests, commonsense reasoning, multidisciplinary knowledge, instruction following, and complex audio analysis and reasoning tasks.

Experimental Findings

The study found that training with simple cross-entropy on dataset labels could sometimes degrade performance, suggesting the need for diverse data. However, adding a Kullback-Leibler (KL) divergence term improved results by helping the student model better match the teacher’s output distribution. Most effectively, using teacher-generated outputs as “hard targets” provided high-quality guidance for reasoning and knowledge transfer. Combining both speech-to-text and text-to-text distillation achieved the best overall results, highlighting how preserving text-mode knowledge complements speech-mode alignment and effectively mitigates catastrophic forgetting while enhancing speech understanding. Even in text-mode, incorporating cross-modal knowledge distillation strengthened reasoning and knowledge capabilities.

Furthermore, the method enhanced the model’s ability for acoustic analysis and reasoning over human speech. Incorporating additional audio question answering (AQA) data led to further improvements in sound and speech categories, suggesting future research could combine semantic dialogue knowledge distillation with acoustic audio analysis distillation.

Also Read:

Conclusion

In conclusion, this research highlights the prevalent issues of catastrophic forgetting and modality inequivalence in speech LLMs, where adding speech capabilities can degrade existing knowledge, especially with spoken queries. The proposed cross-modal distillation framework, utilizing both text and speech channels, effectively addresses these challenges. It successfully preserves textual knowledge, improves alignment across modalities, and enhances reasoning in speech-based interactions, paving the way for more robust and capable speech LLMs. You can read the full paper here: Cross-Modal Knowledge Distillation for Speech Large Language Models.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -