spot_img
HomeResearch & DevelopmentUniSS: Advancing Speech-to-Speech Translation with Voice and Emotion Preservation

UniSS: Advancing Speech-to-Speech Translation with Voice and Emotion Preservation

TLDR: UniSS is a new single-stage framework for expressive speech-to-speech translation (S2ST) that accurately translates spoken content while preserving speaker identity and emotional style. It addresses challenges like data scarcity and complex pipelines by integrating with large language models (LLMs) through a cross-modal chain-of-thought prompting process. The framework also introduces UniST, a large-scale, high-quality Chinese-English S2ST dataset. UniSS significantly outperforms previous methods in translation fidelity, speech quality, and preservation of voice, emotion, and duration consistency, offering a simpler and more effective approach for future S2ST systems.

Speech-to-speech translation (S2ST) is a fascinating field that aims to convert spoken words from one language into spoken words in another, ideally preserving the speaker’s unique voice and emotional tone. Imagine real-time interpretation or seamlessly dubbed videos that sound just like the original speaker. While significant progress has been made in translating content accurately, maintaining expressive qualities like speaker identity and emotional style has remained a considerable challenge.

Researchers have identified three main hurdles: a shortage of paired speech data that retains expressive styles, the complexity of traditional multi-stage processing systems, and the difficulty of fully leveraging the powerful translation capabilities of large language models (LLMs) for speech.

Introducing UniSS: A Unified Approach to Expressive S2ST

A new research paper introduces UniSS, a novel single-stage framework designed to tackle these challenges head-on. UniSS aims to simplify the S2ST process while enhancing its ability to preserve voice and emotion. Unlike conventional systems that break down S2ST into separate steps like speech recognition, text translation, and speech synthesis, UniSS integrates these functions into a single, unified model.

The core of UniSS is its ability to seamlessly integrate with existing text-based LLM frameworks. It uses a pre-trained LLM (specifically, Qwen2.5-1.5B-Instruct) as its backbone, expanding its vocabulary to include discrete speech tokens. This allows the model to treat both speech and text uniformly as sequences of tokens within the same transformer architecture.

Leveraging LLM Intelligence with Cross-Modal Chain-of-Thought Prompting

One of UniSS’s most innovative features is its ‘cross-modal chain-of-thought (CoT) prompting’ process. Inspired by how LLMs can break down complex text tasks into logical steps, UniSS applies this to speech. For high-quality translation, the model follows a ‘listen, translate, and speak’ sequence. It first transcribes the source speech into text, then translates that text into the target language, and finally generates the target speech. This explicit chain allows UniSS to tap into the LLM’s robust, pre-trained text translation abilities, effectively transferring them to the speech domain.

For scenarios requiring faster inference, UniSS also offers a ‘Performance Mode’ that streamlines this process by skipping the initial transcription step, directly generating target text and then speech. This provides a flexible trade-off between translation fidelity and speed.

To represent speech, UniSS employs a triple-tokenizer strategy, separating speaker-related attributes (like timbre and emotion) into ‘speaker tokens’, linguistic content into ‘linguistic tokens’, and generation-oriented semantics into ‘semantic tokens’. This careful disentanglement helps the model understand and generate expressive speech more effectively.

The UniST Dataset: Fueling Expressive S2ST

A major bottleneck in developing expressive S2ST models has been the lack of large-scale, high-quality training data that preserves speaker characteristics. To address this, the researchers constructed and released UniST, a massive Chinese-English S2ST dataset comprising 44.8k hours of data. This dataset was created using a scalable synthesis pipeline that translates source text and then synthesizes target speech while conditioning on the source speaker’s voice to preserve identity and emotion. This extensive dataset is crucial for training powerful unified models like UniSS.

Also Read:

Outstanding Performance Across Metrics

Experimental results demonstrate that UniSS significantly outperforms previous methods. It achieves state-of-the-art translation fidelity, as measured by Speech-BLEU scores, surpassing both end-to-end and cascaded baselines. Beyond accuracy, UniSS also excels in preserving voice, emotion, and duration consistency, while delivering high speech quality. Subjective evaluations by bilingual speakers further confirm its strong capabilities in emotion and speaker preservation, as well as speech naturalness.

The research paper, available at arXiv:2509.21144, establishes a simpler and more effective paradigm for building the next generation of expressive S2ST systems. While currently focused on Chinese and English, the framework’s design allows for easy extension to multilingual scenarios, and future work aims to further unify the speech tokenization process.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -