Skip to main content
zenodoopen

System Fingerprint Recognition for Deepfake Audio (SFR) - Compressed Set

<div>The rapid progress of deep speech synthesis models&nbsp;has posed significant threats to society such as malicious manip</div> <div>ulation of content. This has led to an increase in studies aimed&nbsp;at detecting so-called &ldquo;deepfake audio&rdquo;. However, existing works</div> <div>focus on the binary detection of real audio and fake audio. In&nbsp;real-world scenarios such as model copyright protection and</div> <div>digital evidence forensics, it is needed to know what tool or&nbsp;model generated the deepfake audio to explain the decision. This</div> <div>motivates us to ask: &lsquo;Can we recognize the system fingerprints&nbsp;of deepfake audio?&rsquo; In this paper, we present the first deepfake</div> <div>audio dataset for System Fingerprint Recognition (SFR) and&nbsp;conduct an initial investigation. We collected the dataset from</div> <div>the speech synthesis systems of seven Chinese vendors that use&nbsp;the latest state-of-the-art deep learning technologies, including</div> <div>both clean and compressed sets. In addition, we provide extensive benchmarks and research findings to facilitate the further development of system fingerprint recognition methods. The dataset is publicly available.&nbsp;</div> <div>&nbsp;</div> <div>The subsets 01, 02, and 03 represent the training set, development set, and test set, respectively.</div> <div>&nbsp;</div> <div> <div>This data set is licensed with a CC BY-NC-ND 4.0 license.</div> </div>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
8
Access
16
Reuse readiness
8
Engagement
0