spot_img
HomeResearch & DevelopmentAssessing Intrinsic Bias in AI Voice Cloning for Dysarthric...

Assessing Intrinsic Bias in AI Voice Cloning for Dysarthric Speech

TLDR: This research paper investigates biases in F5-TTS, a zero-shot voice cloning system, when synthesizing speech for individuals with dysarthria. It finds that while the system generally preserves speaker identity and prosody, it exhibits a strong bias against intelligibility for more severe dysarthria. This bias can lead to performance degradation in downstream applications like Automatic Speech Recognition (ASR) when using augmented data. The study highlights the critical need for fairness-aware data augmentation strategies to develop more inclusive and effective assistive speech technologies.

Speech is a fundamental aspect of human communication, but for individuals with dysarthria—a motor speech disorder—speaking can be challenging, often resulting in slurred or difficult-to-understand speech. This makes interacting with digital devices and using assistive technologies like Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems particularly difficult. A major hurdle in developing robust technologies for dysarthric speech is the scarcity of diverse data, as the condition varies widely in severity and speaker characteristics.

To overcome this data limitation, researchers often turn to data augmentation techniques, including advanced neural speech synthesis and zero-shot voice cloning. These methods can generate high-quality synthetic speech, potentially enhancing the training of ASR models and other assistive speech technologies. One such state-of-the-art voice cloning model is F5-TTS, which can create personalized dysarthric speech at various severity levels.

However, the use of synthetic speech for data augmentation raises critical questions about its effectiveness and fairness. Can synthetic dysarthric speech maintain intelligibility? How does it affect ASR performance? Does the cloned speech accurately retain speaker identity across different severity levels? Crucially, do these speech synthesis models exhibit biases across dysarthric severity levels that could disadvantage certain user groups?

A recent study, detailed in the paper Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS, delves into these questions. Authored by Anuprabha M, Krishna Gurugubelli, and Anil Kumar Vuppala, the research investigates the intrinsic biases within dysarthric speech cloning using the F5-TTS system and the TORGO dataset.

Assessing the Quality and Fairness of Synthetic Speech

The researchers employed F5-TTS, a zero-shot, non-autoregressive TTS system, to generate speech samples for dysarthric speakers. This model synthesizes speech from a short audio prompt and a text prompt, leveraging advanced techniques like flow matching and Diffusion Transformer to produce high-quality, natural-sounding speech while aiming to preserve naturalness, intelligibility, and speaker similarity.

To objectively assess the quality of the generated dysarthric speech, the study focused on three key aspects: intelligibility, speaker similarity, and prosody preservation. Intelligibility was measured using Word Error Rate (WER) and Character Error Rate (CER), which quantify errors in speech recognition. Speaker similarity was assessed using the SIM-o score, which measures how well the synthesized speech reflects the original speaker’s characteristics. Prosody preservation, which relates to the rhythm and intonation of speech, was evaluated using the AutoPCP score.

Beyond quality, the study introduced a framework to quantify fairness using two metrics: Parity Difference (PD) and Disparate Impact (DI). PD measures the extent of differences in objective measures between healthy speakers and various dysarthric severity categories, with a value of 0 indicating similar treatment. DI, or relative disparity, is a ratio where a value of 1 indicates no bias between groups.

Key Findings on Bias and Performance

The objective assessment revealed significant insights. The differences in WER and CER (∆WER and ∆CER) were consistently higher for mid and high severity dysarthria categories compared to healthy and low severity categories. This suggests that the F5-TTS system struggles to accurately represent the speech intelligibility of more severe dysarthria, leading to lower error rates in the generated audio than in the original reference speech.

While speaker similarity (SIM-o) remained relatively stable across different severity levels, indicating good preservation of speaker identity, prosody preservation (AutoPCP) showed a decreasing trend for more severe dysarthric speakers. This implies that the prosody of synthesized speech deviates from the original as dysarthria severity increases.

The fairness metrics further highlighted these biases. Minimal to no bias was observed for low severity categories across all objective measures. However, high disparities were found in ∆WER for both mid and high severity categories, indicating poor fairness in speech recognition for severe dysarthria. Although ∆CER showed some reduction in bias compared to ∆WER for high severity, it still indicated bias. Interestingly, the study also found gender-based disparities, with male speakers exhibiting higher bias in intelligibility and female speakers showing higher bias in prosody.

Also Read:

Impact on Assistive Technologies

The research also explored the impact of using this zero-shot voice-cloned augmented data on downstream tasks like ASR and dysarthria detection. When augmented data was combined with reference audio samples, it improved overall dysarthria detection accuracy and significantly reduced WER for the low severity category in ASR. However, due to the inherent biases towards higher severities, the addition of synthetic samples did not improve ASR performance for mid and high severities; in fact, it led to a slight degradation.

In conclusion, this study underscores that speech intelligibility is the most biased aspect of dysarthric speech synthesis, particularly for severe cases, while speaker and prosody similarity are relatively well-preserved. These findings emphasize the critical need for fairness-aware data augmentation strategies when using zero-shot voice cloning systems to generate synthetic speech for individuals with severe dysarthria. Addressing these biases is crucial for developing more inclusive and effective speech technologies for people with motor speech disorders.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -