VoiceWukong: Benchmarking Deepfake Voice Detection (part_aa)
<h1>VoiceWukong</h1> <p><em><strong>VoiceWukong </strong>is a comprehensive benchmark for deepfake voice detection, designed to evaluate the performance of various detectors in real-world application scenarios.</em></p> <p><strong>Dataset Features</strong></p> <ul> <li>Large Scale: Contains 265,200 English and 148,200 Chinese deepfake voice samples</li> <li>Diverse Sources: Covers voice samples generated by 19 commercial tools and 15 open-source tools</li> <li>Real-world Scenarios: Constructed 38 data variants covering 6 types of audio manipulations common in practical applications</li> <li>Bilingual Support: Supports evaluation in both Chinese and English languages</li> </ul> <p><strong>Evaluation Results</strong></p> <ul> <li>Conducted comprehensive evaluations on 12 state-of-the-art deepfake voice detectors</li> <li><a href="https://github.com/TakHemlata/SSL_Anti-spoofing">AASIST2</a> achieved the best performance with an Equal Error Rate (EER) of 13.50%</li> <li>Other detectors showed EERs exceeding 20%</li> <li>Results indicate significant challenges for current detectors in practical applications</li> </ul> <p><strong>Human-Machine Comparison Study</strong></p> <ul> <li>Conducted user studies with over 300 participants</li> <li>Comparative analysis of detection capabilities among humans, detectors, and multimodal large language models (<a href="https://github.com/QwenLM/Qwen2-Audio">Qwen2-Audio</a>)</li> <li>Different detectors and humans showed varying identification capabilities for deepfake voices at different deception levels</li> <li>Multimodal large language models demonstrated no effective detection ability</li> </ul> <h2>Dataset</h2> <p>This is the first part of the dataset, and it requires the complete download of both <a href="https://zenodo.org/records/13731918"><strong><em>part_aa</em></strong></a> and <a href="https://zenodo.org/records/13732412"><strong><em>part_ab</em></strong></a> for proper extraction and use. Please ensure that both files are in the same folder. For a detailed introduction to the data, please refer to our paper (to be made available).</p> <p>The second part (part_ab) is at <a href="13732412">part_ab</a><br>extract command : <code>cat VoiceWukong.part_* | tar -xz</code></p> <h2>Leaderboard</h2> <p>Our leaderboard presents comprehensive evaluation results in three main sections:</p> <ol> <li><strong>Overall Performance</strong> - General evaluation metrics for each detector across the entire dataset, providing a broad view of detection capabilities.</li> <li><strong>Manipulation-specific Performance</strong> - Detailed results showing how each detector performs under different types of audio manipulations, offering insights into specific strengths and weaknesses.</li> <li><strong>User Study-based Evaluation</strong> - Performance analysis of detectors on deepfake voices categorized by difficulty levels based on our user study results, demonstrating detector effectiveness across varying deception capabilities.</li> </ol> <p>Visit our <a href="https://voicewukong.github.io/">leaderboard(github.io)</a> for detailed performance metrics and rankings. Additionally, we provide a copy of the leaderboard code <a href="https://zenodo.org/uploads/14650793">here</a> for premanent storage.</p> <h2>Evaluated Detectors' Weighted Models</h2> <ul> <li>All evaluated detectors’ weighted models can be obtained from <a title="https://huggingface.co/VoiceWukong/VoiceWukong/" href="https://huggingface.co/VoiceWukong/VoiceWukong/">huggingface.co</a>. Additionally, we provide a copy of the weights files <a href="https://zenodo.org/uploads/14650793">here</a> for premanent storage.</li> </ul> <h2>User Study Results & Original Outputs</h2> <ul> <li>This <a href="https://github.com/VoiceWukong/VoiceWukong">code repository(github)</a> stores our <a href="https://github.com/VoiceWukong/VoiceWukong/tree/main/Userstudy/result">user study results</a> and the <a href="https://github.com/VoiceWukong/VoiceWukong/tree/main/OutputScore">original outputs</a> of the evaluation detectors. Additionally, we provide a copy of the code repository <a href="https://zenodo.org/uploads/14650793">here</a> for permanent storage.</li> </ul> <p>Note: <em><strong>VoiceWukong</strong></em> <strong>prohibits use for <em>commercial purposes.</em></strong></p>
ShareScore
12/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 4
- Reuse readiness
- 0
- Engagement
- 0