spot_img
HomeResearch & DevelopmentWildSpoof Challenge Unveils Evaluation Plan for Speech Synthesis and...

WildSpoof Challenge Unveils Evaluation Plan for Speech Synthesis and Verification

TLDR: The WildSpoof Challenge introduces an evaluation plan for two speech processing tracks: Text-to-Speech (TTS) synthesis and Spoofing-robust Automatic Speaker Verification (SASV). It aims to promote the use of in-the-wild data and interdisciplinary collaboration. The plan details training/evaluation datasets (TITW, SpoofCeleb), submission requirements, metrics (MCD, a-DCF), and rules for each track, along with a schedule and ethical guidelines.

The WildSpoof Challenge is an exciting new initiative designed to push the boundaries of speech processing by focusing on “in-the-wild” data. This means moving beyond perfectly clean, controlled recordings to tackle the complexities of real-world audio. The challenge is divided into two distinct but related tracks: Text-to-Speech (TTS) synthesis, which involves creating realistic spoofed speech, and Spoofing-robust Automatic Speaker Verification (SASV), which focuses on detecting such generated speech.

The primary goals of WildSpoof are to encourage the use of diverse, real-world data in both TTS and SASV research and to foster collaboration between the communities that generate and detect spoofed audio. This interdisciplinary approach aims to develop more integrated, robust, and realistic speech systems.

Text-to-Speech (TTS) Track

In the TTS track, participants are tasked with building a system that can generate speech in the voice of a specific target speaker. Given a piece of text and a few examples of the target speaker’s voice, the system must produce speech that accurately reads the text while maintaining the target speaker’s unique vocal characteristics.

For training, participants will use the TITW dataset, specifically either the TITW-Easy or TITW-Hard training sets, or both. Evaluation will be conducted using the TITW-KSKT (Known Speaker, Known Text) and TITW-KSUT (Known Speaker, Unknown Text) protocols. KSKT uses text that was present in the training data, while KSUT challenges systems with entirely new, unseen text.

Submissions for the TTS track require participants to provide a single zip package containing 9,113 speech audio files for TITW-KSKT and 8,000 audio files for TITW-KSUT. These files must adhere to strict technical specifications: they must be in `.wav` format, have a 16 kHz sample rate, 16 bits per sample, and use PCM encoding. Participants can verify these specifications using tools like `soxi`.

Performance in the TTS track will be assessed using several metrics, including MCD, UTMOS, DNSMOS, WER, and SPK-sim, all calculated by Versa. It’s important to note that there will be no formal ranking of TTS systems in this track, though a summary of results will be provided.

Key rules for TTS participants include: they cannot also participate in the SASV track; they are allowed to use pre-trained codecs or models, provided these are fine-tuned with the TITW dataset in the final phase if trainable. Anonymity is an option for teams who wish to keep their identity private.

Spoofing-robust Automatic Speaker Verification (SASV) Track

The SASV track challenges participants to develop a system capable of comparing an unlabeled test utterance against one or more enrollment utterances from a known target speaker. The system’s goal is to correctly identify genuine target speakers while rejecting utterances that are either spoofed or do not match the target speaker’s voice. Evaluation trials will include bonafide target, bonafide non-target, and spoofed target scenarios.

Training data for the SASV track comes from the SpoofCeleb dataset. The evaluation will involve a package of genuine and spoofed speech files, along with a trial list specifying the test utterance and the enrollment target speaker ID for each comparison.

Submissions for the SASV track should be a TSV (tab-separated values) file. Each row in this file must contain the trial name, the enrollment target speaker ID, and a continuous-valued score produced by the SASV system. A higher score should indicate a greater likelihood that the trial is a genuine target.

The primary evaluation metric for the SASV track is the agnostic DCF (a-DCF). This metric assigns a cost to a system based on its miss rate and two types of false alarm rates: those related to non-target trials and those related to spoof trials. Reference implementations for this metric are available in the challenge’s GitHub repository.

Rules for SASV participants include: they cannot participate in the TTS track; they can use Self-Supervised Learning (SSL) models for tasks like feature extraction; each test sample must be scored independently, meaning techniques like domain adaptation or using evaluation data for normalization across multiple samples are not permitted. Participants are also prohibited from making public comparisons of results or rankings with other teams, or claiming titles like ‘challenge winner’.

Also Read:

Baselines, Ethics, and Schedule

The challenge provides baseline systems to help participants get started. For the TTS track, the baseline combines Grad-TTS as the acoustic model and DiffWave as the vocoder. For the SASV track, an end-to-end integrated system serves as the baseline. More details on these can be found in the respective TITW and SpoofCeleb papers.

WildSpoof strongly emphasizes ethical research and responsible practices. Participants are encouraged to develop robust solutions for detecting spoofing and deepfakes while adhering to local data protection laws. The organizers expect prompt and responsible disclosure of any vulnerabilities found in ASV technology to enhance security and prevent malicious use. Misuse of knowledge or tools developed through WildSpoof, including hacking or unauthorized access, is strictly prohibited.

Interested participants can register for the challenge through the provided registration form. The tentative schedule includes a registration deadline of November 01, 2025, a submission deadline of December 01, 2025, and results announcement on December 05, 2025. Further details, including paper submission deadlines, are also outlined. You can find the full evaluation plan and more information at the official challenge website: WildSpoof Challenge Evaluation Plan.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -