Skip to main content
zenodorestricted

VoiceWukong: Benchmarking Deepfake Voice Detection (part_aa)

<h1>VoiceWukong</h1> <p><em><strong>VoiceWukong&nbsp;</strong>is a comprehensive benchmark for deepfake voice detection, designed to evaluate the performance of various detectors in real-world application scenarios.</em></p> <p><strong>Dataset Features</strong></p> <ul> <li>Large Scale: Contains 265,200 English and 148,200 Chinese deepfake voice samples</li> <li>Diverse Sources: Covers voice samples generated by 19 commercial tools and 15 open-source tools</li> <li>Real-world Scenarios: Constructed 38 data variants covering 6 types of audio manipulations common in practical applications</li> <li>Bilingual Support: Supports evaluation in both Chinese and English languages</li> </ul> <p><strong>Evaluation Results</strong></p> <ul> <li>Conducted comprehensive evaluations on 12 state-of-the-art deepfake voice detectors</li> <li><a href="https://github.com/TakHemlata/SSL_Anti-spoofing">AASIST2</a>&nbsp;achieved the best performance with an Equal Error Rate (EER) of 13.50%</li> <li>Other detectors showed EERs exceeding 20%</li> <li>Results indicate significant challenges for current detectors in practical applications</li> </ul> <p><strong>Human-Machine Comparison Study</strong></p> <ul> <li>Conducted user studies with over 300 participants</li> <li>Comparative analysis of detection capabilities among humans, detectors, and multimodal large language models (<a href="https://github.com/QwenLM/Qwen2-Audio">Qwen2-Audio</a>)</li> <li>Different detectors and humans showed varying identification capabilities for deepfake voices at different deception levels</li> <li>Multimodal large language models demonstrated no effective detection ability</li> </ul> <h2>Dataset</h2> <p>This is the first part of the dataset, and it requires the complete download of both <a href="https://zenodo.org/records/13731918"><strong><em>part_aa</em></strong></a> and <a href="https://zenodo.org/records/13732412"><strong><em>part_ab</em></strong></a> for proper extraction and use. Please ensure that both files are in the same folder. For a detailed introduction to the data, please refer to our paper (to be made available).</p> <p>The second part (part_ab) is at&nbsp;<a href="13732412">part_ab</a><br>extract command : <code>cat VoiceWukong.part_* | tar -xz</code></p> <h2>Leaderboard</h2> <p>Our leaderboard presents comprehensive evaluation results in three main sections:</p> <ol> <li><strong>Overall Performance</strong>&nbsp;- General evaluation metrics for each detector across the entire dataset, providing a broad view of detection capabilities.</li> <li><strong>Manipulation-specific Performance</strong>&nbsp;- Detailed results showing how each detector performs under different types of audio manipulations, offering insights into specific strengths and weaknesses.</li> <li><strong>User Study-based Evaluation</strong>&nbsp;- Performance analysis of detectors on deepfake voices categorized by difficulty levels based on our user study results, demonstrating detector effectiveness across varying deception capabilities.</li> </ol> <p>Visit our&nbsp;<a href="https://voicewukong.github.io/">leaderboard(github.io)</a> for detailed performance metrics and rankings. Additionally, we provide a copy of the leaderboard code <a href="https://zenodo.org/uploads/14650793">here</a> for premanent storage.</p> <h2>Evaluated Detectors' Weighted Models</h2> <ul> <li>All evaluated detectors&rsquo; weighted models can be obtained from <a title="https://huggingface.co/VoiceWukong/VoiceWukong/" href="https://huggingface.co/VoiceWukong/VoiceWukong/">huggingface.co</a>. Additionally, we provide a copy of the weights files <a href="https://zenodo.org/uploads/14650793">here</a> for premanent storage.</li> </ul> <h2>User Study Results &amp; Original Outputs</h2> <ul> <li>This <a href="https://github.com/VoiceWukong/VoiceWukong">code repository(github)</a> stores our&nbsp;<a href="https://github.com/VoiceWukong/VoiceWukong/tree/main/Userstudy/result">user study results</a>&nbsp;and the&nbsp;<a href="https://github.com/VoiceWukong/VoiceWukong/tree/main/OutputScore">original outputs</a> of the evaluation detectors. Additionally, we provide a copy of the code repository <a href="https://zenodo.org/uploads/14650793">here</a> for permanent storage.</li> </ul> <p>Note: <em><strong>VoiceWukong</strong></em> <strong>prohibits use for <em>commercial purposes.</em></strong></p>

ShareScore

12/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
4
Access
4
Reuse readiness
0
Engagement
0