Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “Anti-spoofing”
PartialSpoof Database - Partially Spoofed Audio Dataset for Anti-spoofing
<p>All existing databases of spoofed speech contain attack data that is spoofed in its entirety. In practice, it is entirely plausible that successful attacks can be mounted with utterances that are only partially spoofed. By definition, partially-spoofed utterances contain a mix of both spoofed and bona fide segments, which will likely degrade the performance of countermeasures trained with entirely spoofed utterances. This hypothesis raises the obvious question: ‘Can we detect partially spoofed audio?’ This paper introduces a new database of partially-spoofed data, named <strong>PartialSpoof</strong>, to help address this question. This new database enables us to investigate and compare the performance of countermeasures on both utterance- and segmental- level labels. Experimental results using the utterance-level labels reveal that the reliability of countermeasures trained to detect fully-spoofed data is found to degrade substantially when tested with partially-spoofed data, whereas training on partially-spoofed data performs reliably in the case of both fully- and partially- spoofed utterances. Additional experiments using segmental-level labels show that spotting injected spoofed segments included in an utterance is a much more challenging task even if the latest countermeasure models are used.</p> <p> </p> <ul> <li><strong>!!!NEW!!! For detailed (bonafide/spoofing methods/nonspeech/concatenated parts) timestamps of PartialSpoof v1.3</strong> <ul> <li><a href="https://drive.google.com/drive/folders/1kKW3GBuooPkAl64Zyv6WPgICH5LZtnQR">Google Drive</a> </li> <li>The official version is under preparation. Please download this one if you urgently need it.</li> </ul> </li> <li>For fine-grained labels of PartialSpoof v1.2 <ul> <li>Arxiv: http://arxiv.org/abs/2204.05177</li> <li>PartialSpoof Database v1.2<strong> </strong>(including segmental-level labels in different temporal resolutions and timestamp labels)<strong>: This one</strong></li> </ul> </li> <li>For the multi-task version of PartialSpoof <strong>v1.1</strong> <ul> <li>Arxiv: https://arxiv.org/abs/2107.14132</li> <li>PartialSpoof Database v1.1 (including 0.16s segmental level labels): https://zenodo.org/record/5112031</li> </ul> </li> <li>For the initial version of PartialSpoof <strong>v1.0</strong> <ul> <li>Arxiv: https://arxiv.org/abs/2104.02518</li> <li>Samples: https://nii-yamagishilab.github.io/zlin-demo/IS2021/index.html</li> <li>PartialSpoof Database v1.0: https://zenodo.org/record/4817532</li> </ul> </li> </ul> <p>P.S.</p> <p>1. Compared to the <a href="../record/4817532#.YLO07S2l1hE">PartialSpoof_v1.0</a> and <a href="../record/5112031">PartialSpoof_v1.1</a>, only <strong>database_segment_labels_v1.2.tar.gz, database_vad.tar.gz, </strong> and<strong> README_v1.2</strong> are updated for version 1.2, you don't need to download other files if you already downloaded version1.0 or 1.1.</p> <p>2. File database_eval.tar.gz is a little large, if you cannot download it smoothly, you can download the split database_eval.tar.gz from <a href="../record/4817532#.YLO07S2l1hE">PartialSpoof_v1.0</a> </p>
CpAug: Refining Copy-Paste Augmentation for Speech Anti-Spoofing
<p>Conventional copy-paste augmentations generate new training instances by concatenating existing utterances to increase the amount of data for neural network training. However, the direct application of copy-paste augmentation for anti-spoofing is problematic. This paper refines the copy-paste augmentation for speech anti-spoofing, dubbed CpAug, to generate more training data with rich intra-class diversity. The CpAug employs two policies: concatenation to merge utterances with identical labels, and substitution to replace segments in an anchor utterance. Besides, considering the impacts of speakers and spoofing attack types, we craft four blending strategies for the CpAug. Furthermore, we explore how CpAug complements the Rawboost augmentation method. Experimental results reveal that the proposed CpAug significantly improves the performance of speech anti-spoofing. Particularly, CpAug with substitution policy leads to relative improvements of 43% and 38% on the ASVspoof’ 19LA and 21LA, respectively. Notably, the CpAug and Rawboost synergize effectively, achieving an EER of 2.91% on ASVspoof’ 21LA.</p>
Latin-American voice anti-spoofing dataset
<p>This dataset contains samples of spoof and real human voice with different accents from Latin-American countries.</p> <p> Table 1. Real samples distribution</p> <table> <tbody> <tr> <td><strong>Accent</strong></td> <td><strong>Gender</strong></td> <td><strong># Speakers</strong></td> <td><strong># Files</strong></td> <td><strong>Nomenclature</strong></td> </tr> <tr> <td>Colombian</td> <td> <table> <tbody> <tr> <td>Male</td> </tr> <tr> <td>Female</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>17</td> </tr> <tr> <td>14</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>2534</td> </tr> <tr> <td>2070</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>com</td> </tr> <tr> <td>cof</td> </tr> </tbody> </table> </td> </tr> <tr> <td>Chilean</td> <td> <table> <tbody> <tr> <td>Male</td> </tr> <tr> <td>Female</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>17</td> </tr> <tr> <td>12</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>2487</td> </tr> <tr> <td>1602</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>clm</td> </tr> <tr> <td>clf</td> </tr> </tbody> </table> </td> </tr> <tr> <td>Peruvian</td> <td> <table> <tbody> <tr> <td>Male</td> </tr> <tr> <td>Female</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>20</td> </tr> <tr> <td>18</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>2917</td> </tr> <tr> <td>2529</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>pem</td> </tr> <tr> <td>pef</td> </tr> </tbody> </table> </td> </tr> <tr> <td>Venezuelan</td> <td> <table> <tbody> <tr> <td>Male</td> </tr> <tr> <td>Female</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>12</td> </tr> <tr> <td>10</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>1754</td> </tr> <tr> <td>1463</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>vem</td> </tr> <tr> <td>vef</td> </tr> </tbody> </table> </td> </tr> <tr> <td>Argentinian</td> <td> <table> <tbody> <tr> <td>Male</td> </tr> <tr> <td>Female</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>12</td> </tr> <tr> <td>30</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>1670</td> </tr> <tr> <td>3790</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>arm</td> </tr> <tr> <td>arf</td> </tr> </tbody> </table> </td> </tr> <tr> <td>Total</td> <td> </td> <td>162</td> <td>22816</td> <td> </td> </tr> </tbody> </table> <p> </p> <p>The bonafide samples were obtained from the following sources:</p> <ul> <li>Colombian accents: <a href="https://www.openslr.org/72/">https://www.openslr.org/72/</a> (<a href="https://www.openslr.org/resources/75/LICENSE">License</a>)</li> <li>Chilean accents: <a href="https://www.openslr.org/71/">https://www.openslr.org/71/</a> (<a href="https://www.openslr.org/resources/75/LICENSE">License</a>)</li> <li>Peruvian accents: <a href="https://www.openslr.org/73/">https://www.openslr.org/73/</a> (<a href="https://www.openslr.org/resources/75/LICENSE">License</a>)</li> <li>Venezuelan accents: <a href="https://www.openslr.org/75/">https://www.openslr.org/75/</a> (<a href="https://www.openslr.org/resources/75/LICENSE">License</a>)</li> <li>Argentinian accents: <a href="https://www.openslr.org/61/">https://www.openslr.org/61/</a> (<a href="https://www.openslr.org/resources/75/LICENSE">License</a>)</li> </ul> <p> </p> <p>The strategies used to generate the spoof samples:</p> <p> Table 2. Spoof Samples distribution</p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Type</strong></td> <td><strong>#Samples</strong></td> </tr> <tr> <td>StarGAN</td> <td>Voice conversion</td> <td>16000</td> </tr> <tr> <td>CycleGAN</td> <td>Voice conversion</td> <td>16000</td> </tr> <tr> <td>Diffusion</td> <td>Voice conversion</td> <td>16000</td> </tr> <tr> <td>TTS</td> <td>Text-to-speech</td> <td>5000</td> </tr> <tr> <td>TTS-StarGAN</td> <td>Text-to-speech / Voice conversion</td> <td>2500</td> </tr> <tr> <td>TTS-Diff</td> <td>Text-to-speech / Voice conversion</td> <td>2500</td> </tr> </tbody> </table> <p> </p> <ul> <li><a href="https://arxiv.org/abs/1806.02169">StarGAN-VC: Non-parallel many-to-many Voice Conversion Using Star Generative Adversarial Networks</a></li> <li><a href="https://ieeexplore.ieee.org/document/8553236">Cyclegan-VC: Non-parallel voice conversion using cycle-consistent adversarial networks</a></li> <li><a href="https://arxiv.org/abs/2109.13821">Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme</a></li> <li>TTS: Microsoft azure TTS</li> <li>TTS-VC: Microsoft azure TTS + StarGAN/Diff</li> </ul> <p> </p> <p> Table 3. Dataset overview</p> <table> <tbody> <tr> <td><strong>Audio Samples</strong></td> <td><strong>Human Speakers</strong></td> <td><strong>Spoofing algorithms</strong></td> <td><strong>Sampling rate</strong></td> </tr> <tr> <td> <table> <tbody> <tr> <td>Bonafide</td> <td>Spoof</td> </tr> <tr> <td>22816</td> <td>58000</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>Male</td> <td>Female</td> </tr> <tr> <td>78</td> <td>84</td> </tr> </tbody> </table> </td> <td> <table> <tbody> <tr> <td>VC</td> <td>TTS</td> <td>VC and TTS</td> </tr> <tr> <td>3</td> <td>1</td> <td>2</td> </tr> </tbody> </table> </td> <td>16kHz</td> </tr> </tbody> </table> <p> </p> <p>On the <em>protocol.txt</em> file is listed all the files with the following structure:</p> <p><em> Subject_id file_name – spoof_type Label</em></p> <p>Consider this line on protocol.txt file:<br> <em>arf_00295 StarGAN-arf _00295_01349969200-cof _03349 _0077577 - StarGAN spoof</em><br> The first part (arf_00295) represent the subject id, from which we can also identify the accent and the gender (see nomenclature column on Table 1). The file name identify the type of spoof following for the source audio file and the target file. StarGAN represents the type of spoof. According to the table 2, this method is a Voice Conversion algorithm. If the file is a bonafide sample, we replace the <em>spoof_type</em> with a dash (-). Finally at the end of the line we refer the kind of label of the file, in the example, the file corresponds to a spoof case.</p> <p>Each zip file contains 6 folders, each one holds a type of samples. For the voice conversion folders, there are 25 sub-folders that indicate the conversion between accents. For example, Argentina-Venezuela folder indicates that the source accent of the file is Argentinian and the target is Venezuelan accent. Inside the folder there are 64 sub-folders that represent the subjects used for the conversion. For instance, the folder arf_00295-vem_04310 means that the source is an Argentinean female and the target is a Venezuelan male (see Table 1 for nomenclature). In the case of a Text-to-Speech folder there are 5 sub-folders that represent the accents. A TTS-VC folder there are 2 sub-folders that represent the voice conversion strategy used. Inside there are other sub-folders for the different combinations of source and target accents.</p> <p>You can check the folder tree structure in the tree.txt file. Table 3 shows a summary of the resulting dataset.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.