Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

42

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

42 results for “DCASE”

Learn how ShareScore rates datasets ↗
zenodo44/100

Ground Truth for DCASE 2020 Challenge Task 2 Evaluation Dataset

<p><strong>Description</strong></p> <p>This data is the ground truth for the &quot;<a href="https://zenodo.org/record/3841772">evaluation dataset</a>&quot; for the&nbsp;<strong>DCASE 2020 Challenge Task 2 &quot;Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring&quot; </strong><a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">[task description]</a>.&nbsp;</p> <p>In the task, three datasets have been released:&nbsp;&quot;<a href="http://zenodo.org/record/3678171">development dataset</a>&quot;, &quot;<a href="https://zenodo.org/record/3727685">additional training&nbsp;dataset</a>&quot;,&nbsp;and &quot;<a href="https://zenodo.org/record/3841772">evaluation dataset</a>&quot;.&nbsp;The evaluation dataset was the last of the three released and&nbsp;includes around 400 samples for each Machine Type and Machine ID used in the evaluation dataset, none of which have any condition label (i.e., normal or anomaly). This ground truth data contains the condition labels.</p> <p>&nbsp;</p> <p><strong>Data format</strong></p> <p>The ground truth data is a CSV file like the following:</p> <p>---------------------------------</p> <p>fan<br> id_01_00000000.wav,normal_id_01_00000098.wav,0<br> id_01_00000001.wav,anomaly_id_01_00000064.wav,1<br> ...</p> <p>id_05_00000456.wav,anomaly_id_05_00000033.wav,1<br> id_05_00000457.wav,normal_id_05_00000049.wav,0<br> pump<br> id_01_00000000.wav,anomaly_id_01_00000049.wav,1<br> id_01_00000001.wav,anomaly_id_01_00000039.wav,1<br> ...</p> <p>id_05_00000346.wav,anomaly_id_05_00000052.wav,1<br> id_05_00000347.wav,anomaly_id_05_00000080.wav,1<br> slider<br> id_01_00000000.wav,anomaly_id_01_00000035.wav,1<br> id_01_00000001.wav,anomaly_id_01_00000176.wav,1<br> ...</p> <p>---------------------------------</p> <p>&quot;Fan&quot;, &quot;pump&quot;, &quot;slider&quot;, etc mean &quot;Machine Type&quot; names. The lines following a Machine Type correspond to pairs of a wave file in the Machine Type and a condition label. The first column shows the name of a wave file. The second column shows the original name of the wave file, but this can be ignored by users. The third column shows the condition label&nbsp;(i.e.,&nbsp;0:&nbsp;normal&nbsp;or&nbsp;1: anomaly).</p> <p>&nbsp;</p> <p><strong>How to use</strong></p> <p>A system for calculating AUC and pAUC scores for the &quot;evaluation dataset&quot; is available&nbsp;on the Github repository <a href="https://github.com/y-kawagu/dcase2020_task2_evaluator">[URL]</a>. The ground truth data is used by&nbsp;this system.&nbsp;For more information, please see the Github repository.</p> <p>&nbsp;</p> <p><strong>Conditions of use</strong></p> <p>This dataset was created jointly by <strong>NTT Corporation</strong> and <strong>Hitachi, Ltd.</strong>&nbsp;and is available&nbsp;under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p>&nbsp;</p> <p><strong>Publication</strong></p> <p>If you use this dataset, please cite <strong>all the following three&nbsp;papers</strong>:</p> <p>Yuma Koizumi, Shoichiro Saito, Noboru Harada, Hisashi Uematsu, and Keisuke Imoto, &quot;ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection,&quot; in Proc. of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2019.&nbsp;<a href="https://ieeexplore.ieee.org/document/8937164">[pdf]</a></p> <p>Harsh Purohit, Ryo Tanabe, Kenji Ichige, Takashi Endo, Yuki Nikaido, Kaori Suefusa, and Yohei Kawaguchi, &ldquo;MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection,&rdquo; in Proc. 4th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2019.&nbsp;<a href="http://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Purohit_21.pdf">[pdf]</a></p> <p>Yuma Koizumi, Yohei Kawaguchi, Keisuke Imoto, Toshiki Nakamura, Yuki Nikaido, Ryo Tanabe, Harsh Purohit, Kaori Suefusa, Takashi Endo, Masahiro Yasuda, and Noboru Harada,&nbsp;&quot;Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring<em>,&quot;</em>&nbsp;&nbsp;in Proc. 5th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE),&nbsp;2020. <a href="https://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Koizumi_3.pdf">[pdf]</a></p> <p><br> <strong>Feedback</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Yuma Koizumi, <a href="mailto:koizumi.yuma@ieee.org">koizumi.yuma@ieee.org</a></li> <li>Yohei Kawaguchi, <a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Keisuke Imoto, <a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul>

opencc-by-nc-sa-4.0Jul 2020View details →
zenodo44/100

DCASE 2024 Task 5: Few-shot Bioacoustic Event Detection Development Set

<p><strong>General Description:</strong></p> <p>The development set for task 5 of DCASE 2024 "Few-shot Bioacoustic Event Detection" consists of 217 audio files acquired from different bioacoustic sources. The dataset is split into training and validation sets.&nbsp;</p> <p>Multi-class annotations are provided for the training set with positive (POS), negative (NEG) and unkwown (UNK) values for each class. UNK indicates uncertainty about a class.&nbsp;</p> <p>Single-class (class of interest) annotations are provided for the validation set, with events marked as positive (POS) or unkwown (UNK) provided for the class of interest.&nbsp;</p> <p><strong>Folder Structure:</strong></p> <p><em>Development_set.zip</em></p> <p>|_Development_Set/</p> <p>&nbsp; &nbsp; |__Training_Set/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___JD/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___HT/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___BV/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___MT/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___WMW/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp;</p> <p>&nbsp; &nbsp; |__Validation_Set/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___HB/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___PB/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___ME/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;|___PB24/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___RD/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___PW/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp;</p> <p><em>Development_set_annotations.zip</em> has the same structure but contains only the *.csv files</p> <p>&nbsp;</p> <p><strong>Dataset statistics</strong></p> <p>Some statistics on this dataset are as follows, split between training and validation set and their sub-folders:</p> <p>-----------------------------------------------------<br>TRAINING SET<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;174<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;21 hours<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;47<br>Total events&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;14229<br>-----------------------------------------------------<br>TRAINING SET/BV<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;10 hours<br>Total classes &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;11<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;9026<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;24000 Hz<br>-----------------------------------------------------<br>TRAINING SET/HT<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5 hours<br>Total classes &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;611<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;6000 Hz<br>-----------------------------------------------------<br>TRAINING SET/JD<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;10 mins<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1<br>Total events&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;357<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;22050 Hz<br>-----------------------------------------------------<br>TRAINING SET/MT<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1 hour and 10 mins<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;4<br>Total events&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1294<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;8000 Hz<br>-----------------------------------------------------<br>TRAINING SET/WMW<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;161<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;4 hours and 40 mins<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;26<br>Total events&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2941<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;various sampling rates<br>-----------------------------------------------------</p> <p>-----------------------------------------------------<br>VALIDATION SET<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;43<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;49 hours and 57 minutes<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;7<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;3504<br>-----------------------------------------------------<br>VALIDATION SET/HB<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;10<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2 hours and 38 minutes<br>Total classes &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;712<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;44100 Hz<br>-----------------------------------------------------<br>VALIDATION SET/PB<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;6<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;3 hours<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;292<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;44100 Hz<br>-----------------------------------------------------<br>VALIDATION SET/ME<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;20 minutes<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;73<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;44100 Hz<br>-----------------------------------------------------<br>VALIDATION SET/PB24<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;4<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2 hours<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;350<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;44100 Hz<br>-----------------------------------------------------<br>VALIDATION SET/RD<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;6<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; 18 hours<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;1<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;1372<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;48000 Hz<br>-----------------------------------------------------<br>VALIDATION SET/PW<br>-----------------------------------------------------<br>Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;15<br>Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;24 hours<br>Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;1<br>Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;705<br>Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;| &nbsp; &nbsp;96000 Hz<br>-----------------------------------------------------</p> <p><strong>Annotation structure</strong></p> <p>Each line of the annotation csv represents an event in the audio file. The column descriptions are as follows:</p> <p>TRAINING SET<br>---------------------<br>Audiofilename, Starttime, Endtime, CLASS_1, CLASS_2, ...CLASS_N</p> <p>VALIDATION SET<br>---------------------<br>Audiofilename, Starttime, Endtime, Q</p> <p>&nbsp;</p> <p><strong>Classes</strong></p> <p>DCASE2024_task5_training_set_classes.csv and DCASE2024_task5_validation_set_classes.csv provide a table with class code correspondence to class name for all classes in the Development set. Additionally, DCASE2024_task5_validation_set_classes.csv also provides a recording names column.</p> <p>DCASE2024_task5_training_set_classes.csv<br>---------------------<br>dataset, class_code, class_name</p> <p>DCASE2024_task5_validation_set_classes.csv<br>---------------------<br>dataset, recording, class_code, class_name</p> <p>&nbsp;</p> <p><strong>Evaluation Set</strong></p> <p>The Evaluation set for this task will be released on the 1 June 2024</p> <p><strong>Open Access:</strong></p> <p>This dataset is available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.<br>&nbsp;</p> <p><strong>Contact info:</strong></p> <p>Please send any feedback or questions to:</p> <p>Burooj Ghani - &nbsp;burooj.ghani@naturalis.nl | Ines Nolasco - i.dealmeidanolasco@qmul.ac.uk</p> <p>Alternately, join us on Slack: <a href="https://join.slack.com/t/dcase/shared_invite/zt-12zfa5kw0-dD41gVaPU3EZTCAw1mHTCA">task-fewshot-bio-sed</a></p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

DCASE 2021 Task 5: Few-shot Bioacoustic Event Detection Evaluation Set

<p><strong>General Description</strong></p> <p>The evaluation set for task 5 of DCASE 2021 &quot;Few-shot Bioacoustic Event Detection&quot; consists of 31 audio files acquired from different bioacoustic sources.&nbsp;</p> <p>In Evaluation_Set_Annotations: the first 5 annotations are provided for each file, with events marked as positive (POS) for the class of interest. This is the same setup used during the DCASE 2021 challenge.</p> <p>In Evaluation_Set_Full_Annotations: the full annotations are provided for each file, with events marked as positive (POS) or unknown (UNK) for the class of interest.</p> <p>&nbsp;</p> <p><strong>Folder Structure</strong></p> <p><em>Evaluation_Set.zip contains audio files and annotation files with 5 first POS events (as used during DCASE 2021 challenge)</em></p> <p>|__Evaluation_Set/</p> <p>&nbsp; &nbsp; |___DC/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; |___ME/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; |___ML/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p><em>Evaluation_Set_Audio.zip</em>&nbsp;has the same structure but contains only the *.wav files.</p> <p><em>Evaluation_Set_Annotations.zip</em>&nbsp;has the same structure but contains only the *.csv files with first 5 POS annotations.</p> <p><em>Evaluation_Set_Full_Annotations.zip</em>&nbsp;has the same structure but contains only the *.csv files with all POS annotations.</p> <p>The subfolders denote different recording sources and there may or may not be overlap between classes of interest from different wav files.</p> <p>&nbsp;</p> <p><strong>Annotation structure</strong></p> <p>Each line of the annotation csv represents an event in the audio file. The column descriptions are as follows:<br> [ Audiofilename, Starttime, Endtime, Q ]</p> <p>&nbsp;</p> <p><strong>Classes</strong></p> <p>DCASE2021_task5_evaluation_set.csv provides a table with class code correspondance to class name for all the recordings of the Evaluation set.</p> <p>DCASE2021_task5_evaluation_set.csv<br> -------------------<br> dataset, recording, class_code, class_name</p> <p>&nbsp;</p> <p><strong>Development Set</strong></p> <p>The development set for the same task can be found at:&nbsp;<a href="https://doi.org/10.5281/zenodo.5412896">https://doi.org/10.5281/zenodo.5412896</a></p> <p>&nbsp;</p> <p><strong>Open Access</strong></p> <p>This dataset is available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.<br> &nbsp;</p> <p><strong>Contact info</strong></p> <p>Please send any feedback or questions to:<br> Veronica Morfi: g.v.morfi@qmul.ac.uk</p>

opencc-by-4.0May 2021View details →
zenodo44/100

DCASE 2021 Task 5: Few-shot Bioacoustic Event Detection Development Set

<p><strong>General Description</strong></p> <p>The development set for task 5 of DCASE 2021 &quot;Few-shot Bioacoustic Event Detection&quot; consists of 19 audio files acquired from different bioacoustic sources. The dataset is split into training and validation Sets.&nbsp;</p> <p>Multi-class annotations are provided for the training set with positive (POS), negative (NEG) and unkwown (UNK) values for each class. UNK indicates uncertainty about a class.&nbsp;</p> <p>Single-class (class of interest) annotations are provided for the validation set, with events marked as positive (POS) or unkwown (UNK) provided for the class of interest.&nbsp;</p> <p>&nbsp;</p> <p><strong>Folder Structure</strong></p> <p><em>Development_Set.zip</em></p> <p>|_Development_Set/</p> <p>&nbsp; &nbsp; |__Training_Set/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___BV/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___HT/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___JD/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___MT/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; |__Validation_Set/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___HV/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___PB/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp;</p> <p><em>Development_Set_Audio.zip</em> has the same structure but contains only the *.wav files.</p> <p><em>Development_Set_Annotations.zip</em> has the same structure but contains only the *.csv files</p> <p>&nbsp;</p> <p><strong>Dataset statistics</strong></p> <p>Some statistics on this dataset are as follows, split between training and validation set and their sub-folders:</p> <p>-----------------------------------------------------<br> TRAINING SET<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;11<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;14 hours and 20 mins<br> Total classes (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;19<br> Total events (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;4,686<br> -----------------------------------------------------<br> TRAINING SET/BV<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;10 hours<br> Total classes (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;11<br> Total events (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2,662<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;24,000 Hz<br> -----------------------------------------------------<br> TRAINING SET/HT<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;3<br> Total duration&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |&nbsp;&nbsp; &nbsp;3 hours<br> Total classes (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;3<br> Total events (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;435<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;6,000 Hz<br> -----------------------------------------------------<br> TRAINING SET/JD<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;10 mins<br> Total classes (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1<br> Total events (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;355<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;22,050 Hz<br> -----------------------------------------------------<br> TRAINING SET/MT<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1 hour and 10 mins<br> Total classes (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;4<br> Total events (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1,234<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;8,000 Hz<br> -----------------------------------------------------</p> <p><br> -----------------------------------------------------<br> VALIDATION SET<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;8<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5 hours<br> Total classes (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;4<br> Total events (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;310<br> -----------------------------------------------------<br> VALIDATION SET/HV<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2 hours<br> Total classes (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br> Total events (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;50<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;6,000 Hz<br> -----------------------------------------------------<br> VALIDATION SET/PB<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;6<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;3 hours<br> Total classes (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br> Total events (excl. UNK)&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;260<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;44,100 Hz<br> -----------------------------------------------------</p> <p>&nbsp;</p> <p><strong>Annotation structure</strong></p> <p>Each line of the annotation csv represents an event in the audio file. The column descriptions are as follows:</p> <p>TRAINING SET<br> ---------------------<br> Audiofilename, Starttime, Endtime, CLASS_1, CLASS_2, ...CLASS_N</p> <p>VALIDATION SET<br> ---------------------<br> Audiofilename, Starttime, Endtime, Q</p> <p>&nbsp;</p> <p><strong>Classes</strong></p> <p>DCASE2021_task5_training_set_classes.csv and&nbsp;DCASE2021_task5_validation_set_classes.csv provide a table with class code&nbsp;correspondace to class name for all classes in the Development set.</p> <p>DCASE2021_task5_training_set_classes.csv<br> ---------------------<br> dataset, class_code, class_name</p> <p>DCASE2021_task5_validation_set_classes.csv<br> ---------------------<br> dataset, recording, class_code, class_name</p> <p>&nbsp;</p> <p><strong>Evaluation Set</strong></p> <p>The Evaluation set for the same task can be found at:&nbsp;<a href="https://doi.org/10.5281/zenodo.5413149">https://doi.org/10.5281/zenodo.5413149</a></p> <p>&nbsp;</p> <p><strong>Open Access</strong></p> <p>This dataset is available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.</p> <p><br> <strong>Contact info</strong></p> <p>Please send any feedback or questions to:<br> Veronica Morfi: g.v.morfi@qmul.ac.uk<br> &nbsp;</p>

opencc-by-4.0Feb 2021View details →
zenodo40/100

DCASE 2020 Challenge Task 2 Additional Training Dataset

<p><strong>Description</strong></p> <p>This dataset is the &quot;additional training&nbsp;dataset&quot; for the&nbsp;<strong>DCASE 2020 Challenge Task 2 &quot;Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring&quot; </strong><a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">[task description]</a>.&nbsp;</p> <p>In the task, three datasets have been or will be released:&nbsp;&quot;<a href="http://zenodo.org/record/3678171">development dataset</a>&quot;, &quot;additional training&nbsp;dataset&quot;,&nbsp;and &quot;<a href="https://zenodo.org/record/3841772">evaluation dataset</a>&quot;.&nbsp;This additional training&nbsp;dataset was released before the &quot;<a href="https://zenodo.org/record/3841772">evaluation dataset</a>&quot;.&nbsp;This&nbsp;dataset&nbsp;includes around 1,000 normal samples for each Machine Type and Machine ID used in the <a href="https://zenodo.org/record/3841772">evaluation dataset</a> and can be used for model training in advance.</p> <p>The recording procedure and data format are the same as&nbsp;the <a href="http://zenodo.org/record/3678171">development dataset</a>.&nbsp;The Machine IDs in this dataset are different from those in the <a href="http://zenodo.org/record/3678171">development dataset</a>.&nbsp;For more information, please see the pages of the&nbsp;<a href="http://zenodo.org/record/3678171">development dataset</a> and the <a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">task description</a>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Directory structure</strong></p> <p>Once&nbsp;you unzip the downloaded files from&nbsp;Zenodo, you can see the following directory structure. Machine Type information is given by directory name, and Machine ID and condition information are given by file name, as:</p> <ul> </ul> <p>/eval_data</p> <ul> <li>/ToyCar <ul> <li>/train (Only normal data for all Machine IDs are included.) <ul> <li>/normal_id_05_00000000.wav</li> <li>...</li> <li>/normal_id_05_00000999.wav</li> <li>/normal_id_06_00000000.wav</li> <li>...</li> <li>/normal_id_07_00000999.wav</li> </ul> </li> </ul> </li> <li>/ToyConveyor (The other Machine Types have the same directory structure as ToyCar.)</li> <li>/fan</li> <li>/pump</li> <li>/slider</li> <li>/valve</li> </ul> <p>&nbsp;</p> <p>The paths of audio files are:</p> <ul> <li>&quot;/eval_data/&lt;Machine_Type&gt;/train/normal_id_&lt;Machine_ID&gt;_[0-9]+.wav&quot;</li> </ul> <p>For example, the Machine Type and Machine ID of&nbsp;&quot;/ToyCar/train/normal_id_05_00000000.wav&quot; are &quot;ToyCar&quot; and &quot;05&quot;, respectively, and&nbsp;its condition is normal (This dataset includes only normal samples).&nbsp;</p> <p>&nbsp;</p> <p><strong>Baseline system</strong></p> <p>A simple baseline system is available&nbsp;on the Github repository <a href="https://github.com/y-kawagu/dcase2020_task2_baseline">[URL]</a>. The baseline system provides a simple entry-level approach that gives a reasonable performance in the dataset of Task 2. It is a good starting point, especially for entry-level researchers who want to get familiar with the anomalous-sound-detection task.</p> <p>&nbsp;</p> <p><strong>Conditions of use</strong></p> <p>This dataset was created jointly by <strong>NTT Corporation</strong> and <strong>Hitachi, Ltd.</strong>&nbsp;and is available&nbsp;under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p>&nbsp;</p> <p><strong>Publication</strong></p> <p>If you use this dataset, please cite <strong>all the following three papers</strong>:</p> <p>Yuma Koizumi, Shoichiro Saito, Noboru Harada, Hisashi Uematsu, and Keisuke Imoto, &quot;ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection,&quot; in Proc. of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2019.&nbsp;<a href="https://ieeexplore.ieee.org/document/8937164">[pdf]</a></p> <p>Harsh Purohit, Ryo Tanabe, Kenji Ichige, Takashi Endo, Yuki Nikaido, Kaori Suefusa, and Yohei Kawaguchi, &ldquo;MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection,&rdquo; in Proc. 4th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2019.&nbsp;<a href="http://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Purohit_21.pdf">[pdf]</a></p> <p>Yuma Koizumi, Yohei Kawaguchi, Keisuke Imoto, Toshiki Nakamura, Yuki Nikaido, Ryo Tanabe, Harsh Purohit, Kaori Suefusa, Takashi Endo, Masahiro Yasuda, and Noboru Harada,&nbsp;&quot;Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring<em>,&quot;</em>&nbsp;in Proc. 5th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE),&nbsp;2020.&nbsp;<a href="https://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Koizumi_3.pdf">[pdf]</a></p> <p><br> <strong>Feedback</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Yuma Koizumi, <a href="mailto:koizumi.yuma@ieee.org">koizumi.yuma@ieee.org</a></li> <li>Yohei Kawaguchi, <a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Keisuke Imoto, <a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul>

opencc-by-nc-sa-4.0Mar 2020View details →
zenodo40/100

DCASE 2020 Challenge Task 2 Evaluation Dataset

<p><strong>Description</strong></p> <p>This dataset is the &quot;evaluation dataset&quot; for the&nbsp;<strong>DCASE 2020 Challenge Task 2 &quot;Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring&quot; </strong><a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">[task description]</a>.&nbsp;</p> <p>In the task, three datasets have been released:&nbsp;&quot;<a href="http://zenodo.org/record/3678171">development dataset</a>&quot;, &quot;<a href="https://zenodo.org/record/3727685">additional training&nbsp;dataset</a>&quot;,&nbsp;and &quot;evaluation dataset&quot;.&nbsp;This evaluation dataset was the last of the three released.&nbsp;This&nbsp;dataset&nbsp;includes around 400 samples for each Machine Type and Machine ID used in the evaluation dataset, none of which have a condition label (i.e., normal or anomaly).</p> <p>The recording procedure and data format are the same as&nbsp;the <a href="http://zenodo.org/record/3678171">development dataset</a>&nbsp;and <a href="https://zenodo.org/record/3727685">additional training&nbsp;dataset</a>.&nbsp;The Machine IDs in this dataset are the same as&nbsp;those in the <a href="https://zenodo.org/record/3727685">additional training&nbsp;dataset</a>.&nbsp;For more information, please see the pages of the&nbsp;<a href="http://zenodo.org/record/3678171">development dataset</a> and the <a href="http://dcase.community/challenge2020/task-unsupervised-detection-of-anomalous-sounds">task description</a>.&nbsp;</p> <p>After the DCASE 2020 Challenge, we released the <a href="https://zenodo.org/record/3951620">ground truth for this evaluation dataset</a>.</p> <p>&nbsp;</p> <p><strong>Directory structure</strong></p> <p>Once&nbsp;you unzip the downloaded files from&nbsp;Zenodo, you can see the following directory structure. Machine Type information is given by directory name, and Machine ID and condition information are given by file name, as:</p> <ul> </ul> <p>/eval_data</p> <ul> <li>/ToyCar <ul> <li>/test &nbsp;(Normal and anomaly data for all Machine IDs are included, but they do not have a condition label.) <ul> <li>/id_05_00000000.wav</li> <li>...</li> <li>/id_05_00000514.wav</li> <li>/id_06_00000000.wav</li> <li>...</li> <li>/id_07_00000514.wav</li> </ul> </li> </ul> </li> <li>/ToyConveyor (The other Machine Types have the same directory structure as ToyCar.)</li> <li>/fan</li> <li>/pump</li> <li>/slider</li> <li>/valve</li> </ul> <p>&nbsp;</p> <p>The paths of audio files are:</p> <ul> <li>&quot;/eval_data/&lt;Machine_Type&gt;/test/id_&lt;Machine_ID&gt;_[0-9]+.wav&quot;</li> </ul> <p>For example, the Machine Type and Machine ID of&nbsp;&quot;/ToyCar/test/id_05_00000000.wav&quot; are &quot;ToyCar&quot; and &quot;05&quot;, respectively. Unlike the <a href="http://zenodo.org/record/3678171">development dataset</a>&nbsp;and <a href="https://zenodo.org/record/3727685">additional training&nbsp;dataset</a>, its condition label is hidden.&nbsp;</p> <p>&nbsp;</p> <p><strong>Baseline system</strong></p> <p>A simple baseline system is available&nbsp;on the Github repository <a href="https://github.com/y-kawagu/dcase2020_task2_baseline">[URL]</a>. The baseline system provides a simple entry-level approach that gives a reasonable performance in the dataset of Task 2. It is a good starting point, especially for entry-level researchers who want to get familiar with the anomalous-sound-detection task.</p> <p>&nbsp;</p> <p><strong>Conditions of use</strong></p> <p>This dataset was created jointly by <strong>NTT Corporation</strong> and <strong>Hitachi, Ltd.</strong>&nbsp;and is available&nbsp;under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p>&nbsp;</p> <p><strong>Publication</strong></p> <p>If you use this dataset, please cite <strong>all the following three&nbsp;papers</strong>:</p> <p>Yuma Koizumi, Shoichiro Saito, Noboru Harada, Hisashi Uematsu, and Keisuke Imoto, &quot;ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection,&quot; in Proc. of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2019.&nbsp;<a href="https://ieeexplore.ieee.org/document/8937164">[pdf]</a></p> <p>Harsh Purohit, Ryo Tanabe, Kenji Ichige, Takashi Endo, Yuki Nikaido, Kaori Suefusa, and Yohei Kawaguchi, &ldquo;MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection,&rdquo; in Proc. 4th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2019.&nbsp;<a href="http://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Purohit_21.pdf">[pdf]</a></p> <p>Yuma Koizumi, Yohei Kawaguchi, Keisuke Imoto, Toshiki Nakamura, Yuki Nikaido, Ryo Tanabe, Harsh Purohit, Kaori Suefusa, Takashi Endo, Masahiro Yasuda, and Noboru Harada,&nbsp;&quot;Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring<em>,&quot;</em>&nbsp;in Proc. 5th Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE),&nbsp;2020. <a href="https://dcase.community/documents/workshop2020/proceedings/DCASE2020Workshop_Koizumi_3.pdf">[pdf]</a></p> <p><br> <strong>Feedback</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Yuma Koizumi, <a href="mailto:koizumi.yuma@ieee.org">koizumi.yuma@ieee.org</a></li> <li>Yohei Kawaguchi, <a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Keisuke Imoto, <a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul>

opencc-by-nc-sa-4.0May 2020View details →
zenodo40/100

Audio captioning DCASE 2020 evaluation (testing) split

<p>This is the <strong>evaluation split for Task 6, Automated Audio Captioning, in DCASE 2020 Challenge</strong>.&nbsp;</p> <p>This evaluation split is the Clotho testing split, which is thoroughly described in the corresponding paper:&nbsp;</p> <p><em>K. Drossos, S. Lipping and T. Virtanen, &quot;Clotho: an Audio Captioning Dataset,&quot; IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 2020, pp. 736-740, doi: 10.1109/ICASSP40776.2020.9052990.</em></p> <p>available online at: <a href="https://arxiv.org/abs/1910.09387">https://arxiv.org/abs/1910.09387</a> and at: <a href="https://ieeexplore.ieee.org/document/9052990 ">https://ieeexplore.ieee.org/document/9052990&nbsp;</a></p> <p>This evaluation split is meant to be used for the purposes of the Task 6 at the scientific challenge&nbsp;DCASE 2020. This split it is not meant to be used for developing audio captioning methods. For developing audio captioning methods, you should use the development and evaluation splits of Clotho.&nbsp;</p> <p>If you want the development and evaluation splits of Clotho dataset, you can find them also in Zenodo, at: <a href="https://zenodo.org/record/3490684">https://zenodo.org/record/3490684</a></p> <p>--------------------------------------------------------------------------------------------------------</p> <p><strong>== License ==</strong></p> <p>The audio files in the archives:</p> <ul> <li>clotho_audio_test.7z&nbsp;</li> </ul> <p>and the associated meta-data in the CSV file:</p> <ul> <li>clotho_metadata_test.csv</li> </ul> <p>are under the corresponding licences (mostly CreativeCommons with attribution) of Freesound [1] platform, mentioned explicitly in the CSV file&nbsp;for each of the audio files. That is, each audio file in the 7z archive&nbsp;is listed in the CSV file&nbsp;with the meta-data. The meta-data for each file are:&nbsp;</p> <ul> <li>File name</li> <li>Start and ending samples for the excerpt that is used in the Clotho dataset</li> <li>Uploader/user in the Freesound platform (manufacturer)</li> <li>Link to the licence of the file</li> </ul> <p>--------------------------------------------------------------------------------------------------------</p> <p><strong>== References ==</strong><br> [1]&nbsp;Frederic Font, Gerard Roma, and Xavier Serra. 2013. Freesound technical demo. In Proceedings of the 21st ACM international conference on Multimedia (MM &#39;13). ACM, New York, NY, USA, 411-412. DOI: https://doi.org/10.1145/2502081.2502245</p>

openother-atMay 2020View details →
zenodo40/100

DCASE 2022 Challenge Task 2 Development Dataset

<p><strong>Description</strong></p> <p>This dataset is the &quot;development dataset&quot; for the <a href="https://dcase.community/challenge2022/task-unsupervised-anomalous-sound-detection-for-machine-condition-monitoring"><strong>DCASE 2022 Challenge Task 2 &quot;Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques</strong>&quot;</a>.</p> <p>The data consists of the normal/anomalous operating sounds of seven&nbsp;types of real/toy machines. Each recording is a single-channel 10-second audio that includes both a machine&#39;s operating sound and environmental noise. The following seven types of real/toy&nbsp;machines are used in this task:</p> <ul> <li>Fan</li> <li>Gearbox</li> <li>Bearing</li> <li>Slide rail</li> <li>ToyCar</li> <li>ToyTrain</li> <li>Valve</li> </ul> <p>&nbsp;</p> <p><strong>Overview of the task</strong></p> <p><strong>Anomalous sound detection (ASD) is the task of identifying whether the sound emitted from a target machine is normal or anomalous. </strong>Automatic detection of mechanical failure is an essential technology in the fourth industrial revolution, which involves artificial intelligence (AI)-based factory automation. Prompt detection of machine anomalies by observing sounds is useful for monitoring the condition of machines.&nbsp;</p> <p>This task is the follow-up to DCASE 2020 Task 2 and DCASE 2021 Task 2. The task this year is to detect anomalous sounds under three main conditions:</p> <p>1. Only normal sound clips are provided as training data (i.e., unsupervised learning scenario). In real-world factories, anomalies rarely occur and are highly diverse. Therefore, exhaustive patterns of anomalous sounds are impossible to create or collect and unknown anomalous sounds that were not observed in the given training data must be detected. This condition is the same as in DCASE 2020 Task 2 and DCASE 2021 Task 2.</p> <p>2. Factors other than anomalies change the acoustic characteristics between training and test data (i.e., domain shift). In real-world cases, operational conditions of machines or environmental noise often differ between the training and testing phases. For example, the operation speed of a conveyor can change due to seasonal demand, or environmental noise can fluctuate depending on the states of surrounding machines. This condition is the same as in DCASE 2021 Task 2.</p> <p>3. In test data, samples unaffected by domain shifts (source domain data) and those affected by domain shifts (target domain data) are mixed, and the source/target domain of each sample is not specified. Therefore, the model must detect anomalies regardless of the domain (i.e., domain generalization).</p> <p>&nbsp;</p> <p><strong>Definition</strong></p> <p>We first define key terms in this task: &quot;machine type,&quot; &quot;section,&quot; &quot;source domain,&quot; &quot;target domain,&quot; and &quot;attributes.&quot;.</p> <ul> <li>&quot;Machine type&quot; indicates the kind of machine, which in this task is one of seven: fan, gearbox, bearing, slide rail, valve, ToyCar, and ToyTrain.</li> <li>A section is defined as a subset of the dataset for calculating performance metrics. Each section is dedicated to a specific type of domain shift.&nbsp;</li> <li>The source domain is the domain under which most of the training data and part of the test data were recorded, and the target domain is a different set of domains under which a few of the training data and part of the test data were recorded. There are differences between the source and target domains in terms of operating speed, machine load, viscosity, heating temperature, type of environmental noise, SNR, etc.</li> <li>Attributes are parameters that define states of machines or types of noise.&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Dataset</strong></p> <p>This dataset consists of three sections for each machine type (Sections 00, 01, and 02), and each section is a complete set of training and test data. For each section, this dataset provides (i) 990 clips of normal sounds in the source domain for training, (ii) ten clips of normal sounds in the target domain for training, and (iii) 100 clips each of normal and anomalous sounds for the test. The source/target domain of each sample is provided. Additionally, the attributes of each sample in the training and test data are provided in the file names and attribute csv files.</p> <p>&nbsp;</p> <p><strong>File names and attribute csv files</strong></p> <p>File names and attribute csv files provide reference labels for each clip. The given reference labels for each training/test clip include machine type, section index, normal/anomaly information, and attributes regarding the condition other than normal/anomaly. The machine type is given by the directory name. The section index is given by their respective file names. For the datasets other than the evaluation dataset, the normal/anomaly information and the attributes are given by their respective file names. Attribute csv files are for easy access to attributes that cause domain shifts. In these files, the file names, name of parameters that cause domain shifts (domain shift parameter, dp), and the value or type of these parameters (domain shift value, dv) are listed. Each row takes the following format:</p> <p>&nbsp; &nbsp; [filename (string)], [d1p (string)], [d1v (int | float | string)], [d2p], [d2v]...</p> <p>&nbsp;</p> <p><strong>Recording procedure</strong></p> <p>Normal/anomalous operating sounds of machines and its related equipment are recorded. Anomalous sounds were collected by deliberately damaging target machines. For simplifying the task, we use only the first channel of multi-channel recordings; all recordings are regarded as single-channel recordings of a fixed microphone. We mixed a target machine sound with environmental noise, and only noisy recordings are provided as training/test data. The environmental noise samples were recorded in several real factory environments. We will publish papers on the dataset to explain the details of the recording procedure by the submission deadline.</p> <p>&nbsp;</p> <p><strong>Directory structure</strong></p> <p>- /dev_data &nbsp;<br> &nbsp; &nbsp; - /fan<br> &nbsp; &nbsp; &nbsp; &nbsp; - /train (only normal clips) &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_train_normal_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_train_normal_0989_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_train_normal_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_train_normal_0009_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_01_source_train_normal_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_02_target_train_normal_0009_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; - /test&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_normal_0000_&lt;attribute&gt;.wav &nbsp; &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_normal_0049_&lt;attribute&gt;.wav &nbsp; &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_anomaly_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_anomaly_0049_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_normal_0000_&lt;attribute&gt;.wav<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_normal_0049_&lt;attribute&gt;.wav&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_anomaly_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_anomaly_0049_&lt;attribute&gt;.wav&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_01_source_test_normal_0000_&lt;attribute&gt;.wav<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_02_target_test_anomaly_0049_&lt;attribute&gt;.wav<br> &nbsp; &nbsp; &nbsp; &nbsp; - attributes_00.csv (attribute csv for section 00)<br> &nbsp; &nbsp; &nbsp; &nbsp; - attributes_01.csv (attribute csv for section 01)<br> &nbsp; &nbsp; &nbsp; &nbsp; - attributes_02.csv (attribute csv for section 02) &nbsp; &nbsp; &nbsp;<br> &nbsp; &nbsp; - /gearbox (The other machine types have the same directory structure as fan.) &nbsp;<br> &nbsp; &nbsp; - /bearing<br> &nbsp; &nbsp; - /slider (`slider` means &quot;slide rail&quot;)<br> &nbsp; &nbsp; - /ToyCar &nbsp;<br> &nbsp; &nbsp; - /ToyTrain &nbsp;<br> &nbsp; &nbsp; - /valve &nbsp;</p> <p>&nbsp;</p> <p><strong>Baseline system</strong></p> <p>Two baseline systems are available on the Github repository&nbsp;<a href="https://github.com/Kota-Dohi/dcase2022_task2_baseline_ae">baseline_ae</a>&nbsp;and&nbsp;<a href="https://github.com/Kota-Dohi/dcase2022_task2_baseline_mobile_net_v2">baseline_mobile_net_v2</a>. The baseline systems provide a simple entry-level approach that gives a reasonable performance in the dataset of Task 2. They are good starting points, especially for entry-level researchers who want to get familiar with the anomalous-sound-detection task.</p> <p>&nbsp;</p> <p><strong>Condition of use</strong></p> <p>This dataset was created jointly by&nbsp;<strong>Hitachi, Ltd.&nbsp;</strong>and&nbsp;<strong>NTT Corporation</strong>&nbsp;and is available&nbsp;under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you use this dataset, please cite all the following three papers.&nbsp;</p> <ul> <li>Kota Dohi, Keisuke Imoto, Noboru Harada, Daisuke Niizumi, Yuma Koizumi, Tomoya Nishida, Harsh Purohit, Takashi Endo, Masaaki Yamamoto, Yohei Kawaguchi,&nbsp;<em>Description and Discussion on DCASE 2022 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques. In arXiv e-prints: 2206.05876,&nbsp;</em>2022. [<a href="https://arxiv.org/abs/2206.05876">URL</a>]</li> <li>Kota Dohi, Tomoya Nishida, Harsh Purohit, Ryo Tanabe, Takashi Endo, Masaaki Yamamoto, Yuki Nikaido, and Yohei Kawaguchi.&nbsp;<em>MIMII DG: sound dataset for malfunctioning industrial machine investigation and inspection for domain generalization task.</em>&nbsp;<em>In arXiv e-prints: 2205.13879</em>, 2022. [<a href="https://arxiv.org/pdf/2205.13879.pdf">URL</a>]</li> <li>Noboru Harada, Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Masahiro Yasuda, and Shoichiro Saito.&nbsp;<em>ToyADMOS2: another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions.</em>&nbsp;In Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021), 1&ndash;5. Barcelona, Spain, November 2021. [<a href="https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Harada_6.pdf">URL</a>]</li> </ul> <p>&nbsp;</p> <p><strong>Contact</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Kota Dohi,&nbsp;<a href="mailto:kota.dohi.gr@hitachi.com">kota.dohi.gr@hitachi.com</a></li> <li>Daisuke Niizumi,&nbsp;<a href="mailto:daisuke.niizumi.dt@hco.ntt.co.jp">daisuke.niizumi.dt@hco.ntt.co.jp</a></li> <li>Yohei Kawaguchi,&nbsp;<a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Keisuke Imoto,&nbsp;<a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

DCASE 2022 Challenge Task 2 Additional Training Dataset

<p><strong>Description</strong></p> <p>This dataset is the &quot;additional training&nbsp;dataset&quot; for the <a href="https://dcase.community/challenge2022/task-unsupervised-anomalous-sound-detection-for-machine-condition-monitoring"><strong>DCASE 2022 Challenge Task 2 &quot;Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques</strong>&quot;</a>.</p> <p>&nbsp;</p> <p><strong>Condition of use</strong></p> <p>This dataset was created jointly by&nbsp;<strong>Hitachi, Ltd.&nbsp;</strong>and&nbsp;<strong>NTT Corporation</strong>&nbsp;and is available&nbsp;under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you use this dataset, please cite all the following three papers.&nbsp;</p> <ul> <li>Kota Dohi, Keisuke Imoto, Noboru Harada, Daisuke Niizumi, Yuma Koizumi, Tomoya Nishida, Harsh Purohit, Takashi Endo, Masaaki Yamamoto, Yohei Kawaguchi,&nbsp;<em>Description and Discussion on DCASE 2022 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques. In arXiv e-prints: 2206.05876,&nbsp;</em>2022. [<a href="https://arxiv.org/abs/2206.05876">URL</a>]</li> <li>Kota Dohi, Tomoya Nishida, Harsh Purohit, Ryo Tanabe, Takashi Endo, Masaaki Yamamoto, Yuki Nikaido, and Yohei Kawaguchi.&nbsp;<em>MIMII DG: sound dataset for malfunctioning industrial machine investigation and inspection for domain generalization task.</em>&nbsp;<em>In arXiv e-prints: 2205.13879</em>, 2022. [<a href="https://arxiv.org/pdf/2205.13879.pdf">URL</a>]</li> <li>Noboru Harada, Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Masahiro Yasuda, and Shoichiro Saito.&nbsp;<em>ToyADMOS2: another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions.</em>&nbsp;In Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021), 1&ndash;5. Barcelona, Spain, November 2021. [<a href="https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Harada_6.pdf">URL</a>]</li> </ul> <p>&nbsp;</p> <p><strong>Contact</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Kota Dohi,&nbsp;<a href="mailto:kota.dohi.gr@hitachi.com">kota.dohi.gr@hitachi.com</a></li> <li>Daisuke Niizumi,&nbsp;<a href="mailto:daisuke.niizumi.dt@hco.ntt.co.jp">daisuke.niizumi.dt@hco.ntt.co.jp</a></li> <li>Yohei Kawaguchi,&nbsp;<a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> <li>Keisuke Imoto,&nbsp;<a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> </ul>

opencc-by-4.0Apr 2022View details →
zenodo40/100

DCASE 2022 Task 5: Few-shot Bioacoustic Event Detection Development Set

<p><strong>General Description:</strong></p> <p>The development set for task 5 of DCASE 2022&nbsp;&quot;Few-shot Bioacoustic Event Detection&quot; consists of 192 audio files acquired from different bioacoustic sources. The dataset is split into training and validation sets.&nbsp;</p> <p>Multi-class annotations are provided for the training set with positive (POS), negative (NEG) and unkwown (UNK) values for each class. UNK indicates uncertainty about a class.&nbsp;</p> <p>Single-class (class of interest) annotations are provided for the validation set, with events marked as positive (POS) or unkwown (UNK) provided for the class of interest.&nbsp;</p> <p><strong>this version (3):</strong><br> * fixes issues with annotations from HB set</p> <p>&nbsp;</p> <p><strong>Folder Structure:</strong></p> <p><em>Development_Set.zip</em></p> <p>|_Development_Set/</p> <p>&nbsp; &nbsp; |__Training_Set/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___JD/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___HT/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___BV/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___MT/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___WMW/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp;</p> <p>&nbsp; &nbsp; |__Validation_Set/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___HB/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___PB/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |___ME/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp;</p> <p><em>Development_Set_Annotations.zip</em>&nbsp;has the same structure but contains only the *.csv files</p> <p>&nbsp;</p> <p><strong>## Dataset statistics</strong></p> <p>Some statistics on this dataset are as follows, split between training and validation set and their sub-folders:</p> <p>-----------------------------------------------------<br> TRAINING SET<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;174<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;21 hours<br> Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;47<br> Total events&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;14229<br> -----------------------------------------------------<br> TRAINING SET/BV<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;10 hours<br> Total classes &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;11<br> Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;9026<br> Ratio event/duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;0.04<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;24000 Hz<br> -----------------------------------------------------<br> TRAINING SET/HT<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5 hours<br> Total classes &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5<br> Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;611<br> Ratio event/duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;0.05<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;6000 Hz<br> -----------------------------------------------------<br> TRAINING SET/JD<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;10 mins<br> Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1<br> Total events&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;357<br> Ratio event/duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;0.06<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;22050 Hz<br> -----------------------------------------------------<br> TRAINING SET/MT<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1 hour and 10 mins<br> Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;4<br> Total events&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1294<br> Ratio event/duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;0.04<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;8000 Hz<br> -----------------------------------------------------<br> TRAINING SET/WMW<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;161<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;4 hours and 40 mins<br> Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;26<br> Total events&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2941<br> Ratio event/duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;0.24<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;various sampling rates<br> -----------------------------------------------------</p> <p>-----------------------------------------------------<br> VALIDATION SET<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;18<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5 hours and 57 minutes<br> Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;5<br> Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1077<br> -----------------------------------------------------<br> VALIDATION SET/HB<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;10<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2 hours and 38 minutes<br> Total classes &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;1<br> Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;712<br> Ratio event/duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;0.7<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;44100 Hz<br> -----------------------------------------------------<br> VALIDATION SET/PB<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;6<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;3 hours<br> Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br> Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;292<br> Ratio event/duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;0.003<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;44100 Hz<br> -----------------------------------------------------<br> VALIDATION SET/ME<br> -----------------------------------------------------<br> Number of audio recordings&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br> Total duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;20 minutes<br> Total classes&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;2<br> Total events &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;73<br> Ratio event/duration&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;0.01<br> Sampling rate&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;|&nbsp;&nbsp; &nbsp;44100 Hz<br> -----------------------------------------------------</p> <p>&nbsp;</p> <p><strong>Annotation structure</strong></p> <p>Each line of the annotation csv represents an event in the audio file. The column descriptions are as follows:</p> <p>TRAINING SET<br> ---------------------<br> Audiofilename, Starttime, Endtime, CLASS_1, CLASS_2, ...CLASS_N</p> <p>VALIDATION SET<br> ---------------------<br> Audiofilename, Starttime, Endtime, Q</p> <p>&nbsp;</p> <p><strong>Classes</strong></p> <p>DCASE2022_task5_training_set_classes.csv and&nbsp;DCASE2022_task5_validation_set_classes.csv provide a table with class code&nbsp;correspondence to class name for all classes in the Development set.</p> <p>DCASE2022_task5_training_set_classes.csv<br> ---------------------<br> dataset, class_code, class_name</p> <p>DCASE2022_task5_validation_set_classes.csv<br> ---------------------<br> dataset, recording, class_code, class_name</p> <p>&nbsp;</p> <p><strong>Evaluation Set</strong></p> <p>The Evaluation set for this task will be released on the 1st of June 2022</p> <p><strong>Open Access:</strong></p> <p>This dataset is available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.<br> &nbsp;</p> <p><strong>Contact info:</strong></p> <p>Please send any feedback or questions to:</p> <p>Ines Nolasco -&nbsp;&nbsp;i.dealmeidanolasco@qmul.ac.uk</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

BirdVox-DCASE-20k: a dataset for bird audio detection in 10-second clips

<p>BirdVox-DCASE-20k: a dataset for bird audio detection in 10-second clips<br> =====================================================<br> Version 2.0, March 2018.</p> <p><br> Created By<br> -------------</p> <p>Vincent Lostanlen (1, 2, 3), Justin Salamon (2, 3), Andrew Farnsworth (1), Steve Kelling (1), and Juan Pablo Bello (2, 3).</p> <p>(1): Cornell Lab of Ornithology (CLO)<br> (2): Center for Urban Science and Progress, New York University<br> (3): Music and Audio Research Lab, New York University</p> <p>https://wp.nyu.edu/birdvox</p> <p>&nbsp;</p> <p>Description<br> --------------</p> <p>The BirdVox-DCASE-20k dataset contains 20,000 ten-second audio recordings. These recordings come from ROBIN autonomous recording units, placed near Ithaca, NY, USA during the fall 2015. They were captured on the night of September 23rd, 2015, by six different sensors, originally numbered 1, 2, 3, 5, 7, and 10.</p> <p>Out of these 20,000 recording, 10,017 (50.09%) contain at least one bird vocalization (either song, call, or chatter).</p> <p>The dataset is a derivative work of the BirdVox-full-night dataset [1], containing almost as much data but formatted into ten-second excerpts rather than ten-hour full night recordings.</p> <p>In addition, the BirdVox-DCASE-20k dataset is provided as a development set in the context of the &quot;Bird Audio Detection&quot; challenge, organized by DCASE (Detection and Classification of Acoustic Scenes and Events) and the IEEE Signal Processing Society.</p> <p>The dataset can be used, among other things, for the development and evaluation of bioacoustic classification models.</p> <p><br> We refer the reader to [1] for details on the distribution of the data and [2] for details on the hardware of ROBIN recording units.</p> <p>[1] V. Lostanlen, J. Salamon, A. Farnsworth, S. Kelling, J.P. Bello. &quot;BirdVox-full-night: a dataset and benchmark for avian flight call detection&quot;, Proc. IEEE ICASSP, 2018.</p> <p>[2] J. Salamon, J. P. Bello, A. Farnsworth, M. Robbins, S. Keen, H. Klinck, and S. Kelling. Towards the Automatic Classification of Avian Flight Calls for Bioacoustic Monitoring. PLoS One, 2016.</p> <p>&nbsp;</p> <p>Data Files<br> ------------</p> <p>The wav folder contains the recordings as WAV files, sampled at 44,1 kHz, with a single channel (mono). The original sample rate was 24 kHz.</p> <p>The name of each wav file is a random 128-bit UUID (Universal Unique IDentifier) string, which is randomized with respect to the origin of the recording in BirdVox-full-night, both in terms of time (UTC hour at the start of the excerpt) and space (location of the sensor).</p> <p>The origin of each 10-second excerpt is known by the challenge organizers, but not disclosed to the participants.</p> <p>&nbsp;</p> <p>Metadata Files<br> --------------</p> <p>A table containing a binary label &quot;hasbird&quot; associated to every recording in BirdVox-DCASE-20k is available on the website of the DCASE &quot;Bird Audio Detection&quot; challenge: http://machine-listening.eecs.qmul.ac.uk/bird-audio-detection-challenge/</p> <p>These labels were automatically derived from the annotations of avian flight call events in the BirdVox-full-night dataset.</p> <p>If your evaluation procedure requires the precise timestamps of each avian flight call (at a fine time scale of 50 ms), and is agnostic to non-flight call avian vocalizations (e.g. geese, crows, owls, etc.), we kindly suggest you to use the BirdVox-full-night dataset rather than BirdVox-DCASE-20k: wp.nyu.edu/birdvox/birdvox-full-night</p> <p>On the other hand, if your evaluation procedure encompasses all avian vocalizations, and is performed at a coarse time scale of 10 seconds, then BirdVox-DCASE-20k is the appropriate dataset.</p> <p>The annotation campaign of avian flight calls in BirdVox-full-night was performed by Andrew Farnsworth and lasted 102 hours.</p> <p>The additional annotation campaign of non-flight call avian vocalizations was performed by Vincent Lostanlen and lasted 10 hours.</p> <p>The accuracy of the labeling is estimated to be somewhere between 99.5% (100 mislabelings) and 99.95% (10 mislabelings).</p> <p><br> Please Acknowledge BirdVox-DCASE-20k in Academic Research<br> --------------------------------------------------------------------------------</p> <p>When BirdVox-70k is used for academic research, we would highly appreciate it if&nbsp; scientific publications of works partly based on this dataset cite the&nbsp; following publication:</p> <p>V. Lostanlen, J. Salamon, A. Farnsworth, S. Kelling, J. Bello. &quot;BirdVox-full-night: a dataset and benchmark for avian flight call detection&quot;, Proc. IEEE ICASSP, 2018.</p> <p>@inproceedings{lostanlen2018icassp,<br> &nbsp; title = {BirdVox-full-night: a dataset and benchmark for avian flight call detection},<br> &nbsp; author = {Lostanlen, Vincent and Salamon, Justin and Farnsworth, Andrew and Kelling, Steve and Bello, Juan Pablo},<br> &nbsp; booktitle = {Proc. IEEE ICASSP},<br> &nbsp; year = {2018},<br> &nbsp; published = {IEEE},<br> &nbsp; venue = {Calgary, Canada},<br> &nbsp; month = {April},<br> }</p> <p>The creation of this dataset was supported by NSF grants 1125098 (BIRDCAST) and 1633259 (BIRDVOX), a Google Faculty Award, the Leon Levy Foundation, and two anonymous donors.</p> <p>&nbsp;</p> <p>Conditions of Use<br> ---------------------</p> <p>Dataset created by Vincent Lostanlen, Justin Salamon, Andrew Farnsworth, Steve Kelling, and Juan Pablo Bello.</p> <p>The BirdVox-DCASE-20k dataset is offered free of charge under the terms of the Creative&nbsp; Commons Attribution 4.0 International (CC BY 4.0) license:<br> https://creativecommons.org/licenses/by/4.0/</p> <p>The dataset and its contents are made available on an &quot;as is&quot; basis and without&nbsp; warranties of any kind, including without limitation satisfactory quality and&nbsp; conformity, merchantability, fitness for a particular purpose, accuracy or&nbsp; completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, Cornell Lab of Ornithology is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the BirdVox-DCASE-20k dataset or any part of it.</p> <p>&nbsp;</p> <p>Feedback<br> -----------</p> <p>Please help us improve BirdVox-DCASE-20k by sending your feedback to:<br> * Vincent Lostanlen: vincent.lostanlen@gmail.com for feedback regarding data pre-processing,<br> * Andrew Farnsworth: af27@cornell.edu for feedback regarding data collection and ornithology, or<br> * Dan Stowell: dan.stowell@qmul.ac.uk for feedback regarding the DCASE &quot;Bird Audio Detection&quot; challenge.</p> <p>In case of a problem, please include as many details as possible.</p> <p>&nbsp;</p> <p><br> Acknowledgements<br> ------------------------</p> <p>We thank Jessie Barry, Ian Davies, Tom Fredericks, Jeff Gerbracht, Sara Keen, Holger Klinck, Anne Klingensmith, Ray Mack, Peter Marchetto, Ed Moore, Matt Robbins, Ken Rosenberg, and Chris Tessaglia-Hymes for designing autonomous recording units and collecting data.<br> We acknowledge that the land on which the data was collected is the unceded territory of the Cayuga nation, which is part of the Haudenosaunee (Iroquois) confederacy.</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

DCASE 2018, Task 5: Monitoring of domestic activities based on multi-channel acoustics - Development dataset

<p>This repository contains the development data of task 5 of the DCASE 2018 challenge. The dataset is a derivative of the SINS database.</p> <p>The SINS database contains a continuous recording of one person living in a vacation home over a period of one week. The recordings were manually annotated on daily activity level: &quot;Cooking&quot;, &quot;Dishwashing&quot;, &quot;Eating&quot;, &quot;Social activity (visit, phone call)&quot;, &quot;Vacuum cleaning&quot;, &quot;Watching TV&quot;, &quot;Working&quot;, &quot;Presence&quot; and &quot;Absence&quot;. More information can be found on (please cite this papers when using the dataset):</p> <p>G. Dekkers, S. Lauwereins, B. Thoen, M. W. Adhana, H. Brouckxon, T. van Waterschoot, B. Vanrumste, M. Verhelst, and P. Karsmakers, &ldquo;The SINS database for detection of daily activities in a home environment using an acoustic<br> sensor network,&rdquo; in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2017 Workshop (DCASE2017), Munich, Germany, November 2017, pp. 32&ndash;36.</p> <p>G. Dekkers, L. Vuegen, T. van Waterschoot, B. Vanrumste, and P. Karsmakers, &ldquo;DCASE 2018 Challenge - Task 5: Monitoring of domestic activities based on multi-channel acoustics,&rdquo; KU Leuven, Tech. Rep., July 2018.</p> <p>The derivative of the SINS database, &#39;DCASE 2018 &ndash; Task 5 development dataset&#39; consists of data collected by 4 microphone arrays in the combined living room and kitchen area. The continuous recordings were split into audio segments of 10s. These audio segments are provided as individual files along with the ground truth. In total 72984 segments are made available, leading to approximately 200 hours of data.</p> <p>More information about the challenge and the specific dataset can be found <a href="http://dcase.community/challenge2018/task-monitoring-domestic-activities">here</a>. Information solely related to the content of the dataset is available in&nbsp;&#39;DCASE2018-task5-dev.doc.zip&#39;.&nbsp;<br> <br> <strong>By accessing or using this database, the user accepts the provided EULA (available in DCASE2018-task5-dev.doc.zip).</strong></p>

opencc-by-nc-4.0Mar 2018View details →
zenodo40/100

Evaluation datasets for DCASE 2018 Bird Audio Detection

<p>Evaluation data audio for <a href="http://dcase.community/challenge2018/task-bird-audio-detection">the DCASE 2018 Bird Audio Detection task (Task 3)</a>.</p> <ol> <li> <p><strong>Crowdsourced dataset, UK (&quot;warblrb10k&quot;)</strong> - a held-out set of 2,000 recordings from the same conditions as the Warblr development dataset.</p> </li> <li> <p><strong>Remote monitoring dataset, Chernobyl (&quot;Chernobyl&quot;)</strong> - 6,620 audio clips collected from unattended remote monitoring equipment in the Chernobyl Exclusion Zone (CEZ). This data was collected as part of the <a href="https://wiki.ceh.ac.uk/display/NRT/NERC+RATE+TREE+Home">TREE</a> (Transfer-Exposure-Effects) research project into the long-term effects of the Chernobyl accident on local ecology. The audio covers a range of birds and includes weather, large mammal and insect noise sampled across various CEZ environments, including abandoned village, grassland and forest areas.</p> </li> <li> <p><strong>Remote monitoring night-flight calls, Poland (&quot;PolandNFC&quot;)</strong> - 4,000 recordings from Hanna Pamuɫa&#39;s PhD project of monitoring autumn nocturnal bird migration. The recordings were collected every night, from September to November 2016 on the Baltic Sea coast, Poland, using Song Meter SM2 units with microphones mounted on 3&ndash;5 m poles. For this challenge, we use a subset derived from 15 nights with different weather conditions and background noise including wind, rain, sea noise, insect calls, human voice and deer calls.</p> </li> </ol>

opencc-by-4.0Dec 2017View details →
zenodo40/100

DCASE 2018, Task 5: Monitoring of domestic activities based on multi-channel acoustics - Evaluation dataset

<p>The dataset is a derivative of the SINS dataset and is meant to be used as an evaluation set for the <a href="http://dcase.community/challenge2018/task-monitoring-domestic-activities">DCASE2018 Task 5 challenge</a>. The development set to be used can be found <a href="https://zenodo.org/record/1247102#.WzIF_NUzZhE">here</a>. The dataset is a derivative of the SINS database.</p> <p>The SINS database contains a continuous recording of one person living in a vacation home over a period of one week. The recordings were manually annotated on daily activity level: &quot;Cooking&quot;, &quot;Dishwashing&quot;, &quot;Eating&quot;, &quot;Social activity (visit, phone call)&quot;, &quot;Vacuum cleaning&quot;, &quot;Watching TV&quot;, &quot;Working&quot;, &quot;Presence&quot; and &quot;Absence&quot;. More information can be found on (please cite this papers when using the dataset):</p> <p>G. Dekkers, S. Lauwereins, B. Thoen, M. W. Adhana, H. Brouckxon, T. van Waterschoot, B. Vanrumste, M. Verhelst, and P. Karsmakers, &ldquo;The SINS database for detection of daily activities in a home environment using an acoustic<br> sensor network,&rdquo; in Proceedings of the Detection and Classification of Acoustic Scenes and Events 2017 Workshop (DCASE2017), Munich, Germany, November 2017, pp. 32&ndash;36.</p> <p>G. Dekkers, L. Vuegen, T. van Waterschoot, B. Vanrumste, and P. Karsmakers, &ldquo;DCASE 2018 Challenge - Task 5: Monitoring of domestic activities based on multi-channel acoustics,&rdquo; KU Leuven, Tech. Rep., July 2018.</p> <p>The derivative of the SINS database, &#39;DCASE 2018 &ndash; Task 5 evaluation dataset&#39; consists of data collected by 7 microphone arrays in the combined living room and kitchen area. The continuous recordings were split into audio segments of 10s. These audio segments are provided as individual files. In total 72972 segments are made available, leading to approximately 200 hours of data with annotations.</p> <p>More information about the challenge and the specific dataset can be found here. Information solely related to the content of the dataset is available in&nbsp;&#39;DCASE2018-task5-eval.doc.zip&#39;.</p> <p>By accessing or using this database, the user accepts the provided EULA (available in DCASE2018-task5-eval.doc.zip).</p>

opencc-by-nc-nd-4.0Jun 2018View details →
zenodo40/100

DCASE 2023 Challenge Task 2 Additional Training Dataset

<p><strong>Description</strong></p> <p>This dataset is the &quot;additional training dataset&quot; for the <a href="https://dcase.community/challenge2023/task-first-shot-unsupervised-anomalous-sound-detection-for-machine-condition-monitoring">DCASE 2023 Challenge Task 2 &quot;First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring&quot;</a>.</p> <p>The data consists of the normal/anomalous operating sounds of seven&nbsp;types of real/toy machines. Each recording is a single-channel audio that includes both a machine&#39;s operating sound and environmental noise. The duration of recordings varies from 6 to 18 sec, depending on the machine type. The following seven types of real/toy&nbsp;machines are used:</p> <ul> <li>Vacuum</li> <li>ToyTank</li> <li>ToyNscale</li> <li>ToyDrone</li> <li>bandsaw</li> <li>grinder</li> <li>shaker</li> </ul> <p>&nbsp;</p> <p><strong>Overview of the task</strong></p> <p><strong>Anomalous sound detection (ASD) is the task of identifying whether the sound emitted from a target machine is normal or anomalous. </strong>Automatic detection of mechanical failure is an essential technology in the fourth industrial revolution, which involves artificial-intelligence-based factory automation. Prompt detection of machine anomalies by observing sounds is useful for monitoring the condition of machines.&nbsp;</p> <p>This task is the follow-up from DCASE 2020 Task 2 to DCASE 2022 Task 2. The task this year is to develop an ASD system that meets the following four requirements.</p> <p>&nbsp;</p> <p><strong>1. Train a model using only normal sound&nbsp;(unsupervised learning scenario)</strong></p> <p>Because anomalies rarely occur and are highly diverse in real-world factories, it can be difficult to collect exhaustive patterns of anomalous sounds. Therefore, the system must detect unknown types of anomalous sounds that are not provided in the training data. This is the same requirement as in the previous tasks.</p> <p><strong>2. Detect anomalies regardless of domain shifts&nbsp;(domain generalization task)&nbsp;</strong></p> <p>In real-world cases, the operational states of a machine or the environmental noise can change to cause domain shifts. Domain-generalization techniques can be useful for handling domain shifts that occur frequently or are hard-to-notice. In this task, the system is required to use domain-generalization techniques for handling these domain shifts. This requirement is the same as in DCASE 2022 Task 2.</p> <p><strong>3. Train a model for a completely new machine type</strong></p> <p>For a completely new machine type, hyperparameters of the trained model cannot be tuned. Therefore, the system should have the ability to train models without additional hyperparameter tuning.</p> <p><strong>4. Train a model using only one machine from its machine type</strong></p> <p>While sounds from multiple machines of the same machine type can be used to enhance detection performance, it is often the case that sound data from only one machine are available for a machine type. In such a case, the system should be able to train models using only one machine from a machine type.</p> <p>&nbsp;</p> <p>The last two requirements are newly introduced in DCASE 2023 Task2 as the &quot;first-shot problem&quot;.</p> <p>&nbsp;</p> <p><strong>Definition</strong></p> <p>We first define key terms in this task: &quot;machine type,&quot; &quot;section,&quot; &quot;source domain,&quot; &quot;target domain,&quot; and &quot;attributes.&quot;.</p> <ul> <li>&quot;Machine type&quot; indicates the type of machine, which in the development dataset is one of seven: fan, gearbox, bearing, slide rail, valve, ToyCar, and ToyTrain.</li> <li>A section is defined as a subset of the dataset for calculating performance metrics.</li> <li>The source domain is the domain under which most of the training data and some of the test data were recorded, and the target domain is a different set of domains under which some of the training data and some of the test data were recorded. There are differences between the source and target domains in terms of operating speed, machine load, viscosity, heating temperature, type of environmental noise, signal-to-noise ratio, etc.</li> <li>Attributes are parameters that define states of machines or types of noise.</li> </ul> <p>&nbsp;</p> <p><strong>Dataset</strong></p> <p>This dataset consists of seven machine types. For each machine type, one section is provided, and the section is a complete set of training and test data. For each section, this dataset provides (i) 990 clips of normal sounds in the source domain for training, (ii) ten clips of normal sounds in the target domain for training. The source/target domain of each sample is provided. Additionally, the attributes of each sample in the training and test data are provided in the file names and attribute csv files.</p> <p>&nbsp;</p> <p><strong>File names and attribute csv files</strong></p> <p>File names and attribute csv files provide reference labels for each clip. The given reference labels for each training/test clip include machine type, section index, normal/anomaly information, and attributes regarding the condition other than normal/anomaly. The machine type is given by the directory name. The section index is given by their respective file names. For the datasets other than the evaluation dataset, the normal/anomaly information and the attributes are given by their respective file names. Attribute csv files are for easy access to attributes that cause domain shifts. In these files, the file names, name of parameters that cause domain shifts (domain shift parameter, dp), and the value or type of these parameters (domain shift value, dv) are listed. Each row takes the following format:</p> <p>&nbsp; &nbsp; [filename (string)], [d1p (string)], [d1v (int | float | string)], [d2p], [d2v]...</p> <p>&nbsp;</p> <p><strong>Recording procedure</strong></p> <p>Normal/anomalous operating sounds of machines and its related equipment are recorded. Anomalous sounds were collected by deliberately damaging target machines. For simplifying the task, we use only the first channel of multi-channel recordings; all recordings are regarded as single-channel recordings of a fixed microphone. We mixed a target machine sound with environmental noise, and only noisy recordings are provided as training/test data. The environmental noise samples were recorded in several real factory environments. We will publish papers on the dataset to explain the details of the recording procedure by the submission deadline.</p> <p>&nbsp;</p> <p><strong>Directory structure</strong></p> <p>- /dev_data &nbsp;</p> <p>&nbsp; &nbsp; - /raw<br> &nbsp; &nbsp; &nbsp; &nbsp; - /Vacuum<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /train (only normal clips) &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_train_normal_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_train_normal_0989_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_train_normal_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_train_normal_0009_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /test&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_normal_0000_&lt;attribute&gt;.wav &nbsp; &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_normal_0049_&lt;attribute&gt;.wav &nbsp; &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_anomaly_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_anomaly_0049_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_normal_0000_&lt;attribute&gt;.wav<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_normal_0049_&lt;attribute&gt;.wav&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_anomaly_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_anomaly_0049_&lt;attribute&gt;.wav&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - attributes_00.csv (attribute csv for section 00)<br> &nbsp; &nbsp; - /ToyTank&nbsp;(The other machine types have the same directory structure as Vacuum.) &nbsp;<br> &nbsp; &nbsp; - /ToyNscale<br> &nbsp; &nbsp; - /ToyDrone<br> &nbsp; &nbsp; - /bandsaw<br> &nbsp; &nbsp; - /grinder<br> &nbsp; &nbsp; - /shaker</p> <p>&nbsp;</p> <p><strong>Baseline system</strong></p> <p>The&nbsp;baseline system is&nbsp;available on the Github repository&nbsp;<a href="https://github.com/nttcslab/dase2023_task2_baseline_ae">dcase2023_task2_baseline_ae</a>.The baseline systems provide a simple entry-level approach that gives a reasonable performance in the dataset of Task 2. They are good starting points, especially for entry-level researchers who want to get familiar with the anomalous-sound-detection task.</p> <p>&nbsp;</p> <p><strong>Condition of use</strong></p> <p>This dataset was created jointly by&nbsp;<strong>Hitachi, Ltd.&nbsp;</strong>and&nbsp;<strong>NTT Corporation</strong>&nbsp;and is available&nbsp;under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you use this dataset, please cite all the following papers. We will publish a paper on the description of the DCASE 2023 Task 2, so pleasure make sure to cite the paper, too.</p> <ul> <li>Noboru Harada, Daisuke Niizumi, Yasunori Ohishi, Daiki Takeuchi, and Masahiro Yasuda. <em>First-shot anomaly detection for machine condition monitoring: A domain generalization baseline. In arXiv e-prints: 2303.00455</em>, 2023.&nbsp;[<a href="https://arxiv.org/pdf/2303.00455.pdf">URL</a>]</li> <li>Kota Dohi, Tomoya Nishida, Harsh Purohit, Ryo Tanabe, Takashi Endo, Masaaki Yamamoto, Yuki Nikaido, and Yohei Kawaguchi.&nbsp;<em>MIMII DG: sound dataset for malfunctioning industrial machine investigation and inspection for domain generalization task.</em>&nbsp;In Proceedings of the 7th Detection and Classification of Acoustic Scenes and Events 2022&nbsp;Workshop (DCASE2022), 31-35. Nancy, France, November 2022, . [<a href="https://arxiv.org/pdf/2205.13879.pdf">URL</a>]</li> <li>Noboru Harada, Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Masahiro Yasuda, and Shoichiro Saito.&nbsp;<em>ToyADMOS2: another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions.</em>&nbsp;In Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021), 1&ndash;5. Barcelona, Spain, November 2021. [<a href="https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Harada_6.pdf">URL</a>]</li> </ul> <p>&nbsp;</p> <p><strong>Contact</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Kota Dohi,&nbsp;<a href="mailto:kota.dohi.gr@hitachi.com">kota.dohi.gr@hitachi.com</a></li> <li>Keisuke Imoto,&nbsp;<a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> <li>Noboru Harada,&nbsp;<a href="mailto:noboru@ieee.org">noboru@ieee.org</a></li> <li>Daisuke Niizumi,&nbsp;<a href="mailto:daisuke.niizumi.dt@hco.ntt.co.jp">daisuke.niizumi.dt@hco.ntt.co.jp</a></li> <li>Yohei Kawaguchi,&nbsp;<a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> </ul>

opencc-by-4.0Apr 2023View details →
zenodo40/100

DCASE 2023 Challenge Task 2 Development Dataset

<p><strong>Description</strong></p> <p>This dataset is the &quot;development dataset&quot; for the <a href="https://dcase.community/challenge2023/task-first-shot-unsupervised-anomalous-sound-detection-for-machine-condition-monitoring">DCASE 2023 Challenge Task 2 &quot;First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring&quot;</a>.</p> <p>The data consists of the normal/anomalous operating sounds of seven&nbsp;types of real/toy machines. Each recording is a single-channel 10-second audio that includes both a machine&#39;s operating sound and environmental noise. The following seven types of real/toy&nbsp;machines are used in this task:</p> <ul> <li>ToyCar</li> <li>ToyTrain</li> <li>Fan</li> <li>Gearbox</li> <li>Bearing</li> <li>Slide rail</li> <li>Valve</li> </ul> <p>&nbsp;</p> <p><strong>Overview of the task</strong></p> <p><strong>Anomalous sound detection (ASD) is the task of identifying whether the sound emitted from a target machine is normal or anomalous. </strong>Automatic detection of mechanical failure is an essential technology in the fourth industrial revolution, which involves artificial-intelligence-based factory automation. Prompt detection of machine anomalies by observing sounds is useful for monitoring the condition of machines.&nbsp;</p> <p>This task is the follow-up from DCASE 2020 Task 2 to DCASE 2022 Task 2. The task this year is to develop an ASD system that meets the following four requirements.</p> <p>&nbsp;</p> <p><strong>1. Train a model using only normal sound&nbsp;(unsupervised learning scenario)</strong></p> <p>Because anomalies rarely occur and are highly diverse in real-world factories, it can be difficult to collect exhaustive patterns of anomalous sounds. Therefore, the system must detect unknown types of anomalous sounds that are not provided in the training data. This is the same requirement as in the previous tasks.</p> <p><strong>2. Detect anomalies regardless of domain shifts&nbsp;(domain generalization task)&nbsp;</strong></p> <p>In real-world cases, the operational states of a machine or the environmental noise can change to cause domain shifts. Domain-generalization techniques can be useful for handling domain shifts that occur frequently or are hard-to-notice. In this task, the system is required to use domain-generalization techniques for handling these domain shifts. This requirement is the same as in DCASE 2022 Task 2.</p> <p><strong>3. Train a model for a completely new machine type</strong></p> <p>For a completely new machine type, hyperparameters of the trained model cannot be tuned. Therefore, the system should have the ability to train models without additional hyperparameter tuning.</p> <p><strong>4. Train a model using only one machine from its machine type</strong></p> <p>While sounds from multiple machines of the same machine type can be used to enhance detection performance, it is often the case that sound data from only one machine are available for a machine type. In such a case, the system should be able to train models using only one machine from a machine type.</p> <p>&nbsp;</p> <p>The last two requirements are newly introduced in DCASE 2023 Task2 as the &quot;first-shot problem&quot;.</p> <p>&nbsp;</p> <p><strong>Definition</strong></p> <p>We first define key terms in this task: &quot;machine type,&quot; &quot;section,&quot; &quot;source domain,&quot; &quot;target domain,&quot; and &quot;attributes.&quot;.</p> <ul> <li>&quot;Machine type&quot; indicates the type of machine, which in the development dataset is one of seven: fan, gearbox, bearing, slide rail, valve, ToyCar, and ToyTrain.</li> <li>A section is defined as a subset of the dataset for calculating performance metrics.</li> <li>The source domain is the domain under which most of the training data and some of the test data were recorded, and the target domain is a different set of domains under which some of the training data and some of the test data were recorded. There are differences between the source and target domains in terms of operating speed, machine load, viscosity, heating temperature, type of environmental noise, signal-to-noise ratio, etc.</li> <li>Attributes are parameters that define states of machines or types of noise.</li> </ul> <p>&nbsp;</p> <p><strong>Dataset</strong></p> <p>This dataset consists of seven machine types. For each machine type, one section is provided, and the section is a complete set of training and test data. For each section, this dataset provides (i) 990 clips of normal sounds in the source domain for training, (ii) ten clips of normal sounds in the target domain for training, and (iii) 100 clips each of normal and anomalous sounds for the test. The source/target domain of each sample is provided. Additionally, the attributes of each sample in the training and test data are provided in the file names and attribute csv files.</p> <p>&nbsp;</p> <p><strong>File names and attribute csv files</strong></p> <p>File names and attribute csv files provide reference labels for each clip. The given reference labels for each training/test clip include machine type, section index, normal/anomaly information, and attributes regarding the condition other than normal/anomaly. The machine type is given by the directory name. The section index is given by their respective file names. For the datasets other than the evaluation dataset, the normal/anomaly information and the attributes are given by their respective file names. Attribute csv files are for easy access to attributes that cause domain shifts. In these files, the file names, name of parameters that cause domain shifts (domain shift parameter, dp), and the value or type of these parameters (domain shift value, dv) are listed. Each row takes the following format:</p> <p>&nbsp; &nbsp; [filename (string)], [d1p (string)], [d1v (int | float | string)], [d2p], [d2v]...</p> <p>&nbsp;</p> <p><strong>Recording procedure</strong></p> <p>Normal/anomalous operating sounds of machines and its related equipment are recorded. Anomalous sounds were collected by deliberately damaging target machines. For simplifying the task, we use only the first channel of multi-channel recordings; all recordings are regarded as single-channel recordings of a fixed microphone. We mixed a target machine sound with environmental noise, and only noisy recordings are provided as training/test data. The environmental noise samples were recorded in several real factory environments. We will publish papers on the dataset to explain the details of the recording procedure by the submission deadline.</p> <p>&nbsp;</p> <p><strong>Directory structure</strong></p> <p>- /dev_data &nbsp;</p> <p>&nbsp; &nbsp; - /raw<br> &nbsp; &nbsp; &nbsp; &nbsp; - /fan<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /train (only normal clips) &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_train_normal_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_train_normal_0989_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_train_normal_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_train_normal_0009_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /test&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_normal_0000_&lt;attribute&gt;.wav &nbsp; &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_normal_0049_&lt;attribute&gt;.wav &nbsp; &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_anomaly_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_source_test_anomaly_0049_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_normal_0000_&lt;attribute&gt;.wav<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_normal_0049_&lt;attribute&gt;.wav&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_anomaly_0000_&lt;attribute&gt;.wav &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - ... &nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - /section_00_target_test_anomaly_0049_&lt;attribute&gt;.wav&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - attributes_00.csv (attribute csv for section 00)<br> &nbsp; &nbsp; - /gearbox (The other machine types have the same directory structure as fan.) &nbsp;<br> &nbsp; &nbsp; - /bearing<br> &nbsp; &nbsp; - /slider (`slider` means &quot;slide rail&quot;)<br> &nbsp; &nbsp; - /ToyCar &nbsp;<br> &nbsp; &nbsp; - /ToyTrain &nbsp;<br> &nbsp; &nbsp; - /valve &nbsp;</p> <p>&nbsp;</p> <p><strong>Baseline system</strong></p> <p>The&nbsp;baseline system is&nbsp;available on the Github repository&nbsp;<a href="https://github.com/nttcslab/dase2023_task2_baseline_ae">dcase2023_task2_baseline_ae</a>.The baseline systems provide a simple entry-level approach that gives a reasonable performance in the dataset of Task 2. They are good starting points, especially for entry-level researchers who want to get familiar with the anomalous-sound-detection task.</p> <p>&nbsp;</p> <p><strong>Condition of use</strong></p> <p>This dataset was created jointly by&nbsp;<strong>Hitachi, Ltd.&nbsp;</strong>and&nbsp;<strong>NTT Corporation</strong>&nbsp;and is available&nbsp;under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) license.</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you use this dataset, please cite all the following papers. We will publish a paper on the description of the DCASE 2023 Task 2, so pleasure make sure to cite the paper, too.</p> <ul> <li>Noboru Harada, Daisuke Niizumi, Yasunori Ohishi, Daiki Takeuchi, and Masahiro Yasuda. <em>First-shot anomaly detection for machine condition monitoring: A domain generalization baseline. In arXiv e-prints: 2303.00455</em>, 2023.&nbsp;[<a href="https://arxiv.org/pdf/2303.00455.pdf">URL</a>]</li> <li>Kota Dohi, Tomoya Nishida, Harsh Purohit, Ryo Tanabe, Takashi Endo, Masaaki Yamamoto, Yuki Nikaido, and Yohei Kawaguchi.&nbsp;<em>MIMII DG: sound dataset for malfunctioning industrial machine investigation and inspection for domain generalization task.</em>&nbsp;In Proceedings of the 7th Detection and Classification of Acoustic Scenes and Events 2022&nbsp;Workshop (DCASE2022), 31-35. Nancy, France, November 2022, . [<a href="https://arxiv.org/pdf/2205.13879.pdf">URL</a>]</li> <li>Noboru Harada, Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Masahiro Yasuda, and Shoichiro Saito.&nbsp;<em>ToyADMOS2: another dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions.</em>&nbsp;In Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021), 1&ndash;5. Barcelona, Spain, November 2021. [<a href="https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Harada_6.pdf">URL</a>]</li> </ul> <p>&nbsp;</p> <p><strong>Contact</strong></p> <p>If there is any problem, please contact us:</p> <ul> <li>Kota Dohi,&nbsp;<a href="mailto:kota.dohi.gr@hitachi.com">kota.dohi.gr@hitachi.com</a></li> <li>Keisuke Imoto,&nbsp;<a href="mailto:keisuke.imoto@ieee.org">keisuke.imoto@ieee.org</a></li> <li>Noboru Harada,&nbsp;<a href="mailto:noboru@ieee.org">noboru@ieee.org</a></li> <li>Daisuke Niizumi,&nbsp;<a href="mailto:daisuke.niizumi.dt@hco.ntt.co.jp">daisuke.niizumi.dt@hco.ntt.co.jp</a></li> <li>Yohei Kawaguchi,&nbsp;<a href="mailto:yohei.kawaguchi.xk@hitachi.com">yohei.kawaguchi.xk@hitachi.com</a></li> </ul>

opencc-by-4.0Feb 2023View details →
zenodo36/100

DCASE 2024 Challenge Task 7 Development Dataset : Environmental Sound Scene Synthesis

<h1><strong>Description</strong></h1> <p>This dataset comprises embeddings and captions utilized as the development dataset for <a href="https://dcase.community/challenge2024/task-sound-scene-synthesis"><strong>DCASE 2024 Challenge Task 7</strong></a>, focusing on 'Environmental Sound Scene Synthesis.' The embeddings are derived from 60 different 4-second audio files formatted as mono 32-bit 32kHz, and are contained in the 'embeddings.tar.xz' file. Captions corresponding to each audio file can be found in 'caption.csv'. This dataset does not comprise the audio files, only the embeddings. Three different types of embeddings are provided: VGGish (vggish), MS-CLAP (clap-2023), and PANNs CNN14 Wavegram-Logmel (panns-wavegram-logmel). Only PANNs CNN14 Wavegram-Logmel (panns-wavegram-logmel) embeddings are used for evaluation in the challenge. For further details, please refer to the challenge website.</p> <p><strong>Contact</strong></p> <ul> <li>Modan Tailleur, modan.tailleur@ls2n.fr</li> <li>Mathieu Lagrange, mathieu.lagrange@ls2n.fr</li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo36/100

DCASE 2024 Task 9: Language-Queried Audio Source Separation | Validation Set

<p>This is the <strong>validation set for Task 9, Language-Queried Audio Source Separation (LASS), in DCASE 2024 Challenge</strong>.&nbsp;</p> <p>This validation split is meant to be used for Task 9 at the scientific challenge DCASE 2024. This split is not meant to be used for training LASS methods. This split is meant to be used for evaluating LASS methods during the model development stage.</p> <p>This validation set consists of 1000 audio files sourced from Freesound [1], uploaded between April and October 2023. Each audio file has been manually annotated with three captions. In the annotation guidance, we instructed annotators to describe the content of audio clips using 5-20 words (similar to the caption style in Clotho [3] and AudioCaps [4] datasets). The tags of each audio file were verified and revised according to the FSD50K [2] sound event categories. Each audio file has been chunked into a 10-second clip and downsampled to 16kHz.</p> <p><strong>== Details ==</strong></p> <p>The audio files in the archives:</p> <ul> <li>lass_validation.zip</li> </ul> <p>and the associated metadata (including tags and captions) in the JSON file:</p> <ul> <li>lass_validation.json</li> </ul> <p>Participants will evaluate their LASS models using synthetic mixture data in the development stage. Specifically, given an audio clip A1 and its corresponding caption C, we select an additional audio clip, A2, to serve as background noise, thereby creating a mixed audio, A3. We anticipate that the LASS system, given A3 and C as inputs, will be able to separate the A1 source. We use the revised tags information to ensure that the two audio clips used in each mix do not share overlapping sound source classes. Three thousand synthetic audio mixtures with signal-to-noise ratios (SNR) ranging from -15dB to 15dB will be generated for the validation of LASS model development. These synthetic mixtures can be generated based on the provided CSV file:</p> <ul> <li>lass_synthetic_validation.csv</li> </ul> <p>The evaluation tool can be found at: https://github.com/Audio-AGI/dcase2024_task9_baseline/blob/main/dcase_evaluator.py</p> <p><strong>== References ==</strong></p> <p>[1] Fonseca E, Pons Puig J, Favory X, et al. Freesound datasets: a platform for the creation of open audio datasets. International Society for Music Information Retrieval (ISMIR), 2017.</p> <p>[2] Fonseca E, Favory X, Pons J, et al. FSD50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 30: 829-852.</p> <p>[3] Drossos K, Lipping S, Virtanen T. Clotho: An audio captioning dataset. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2020: 736-740.</p> <p>[4] Kim C D, Kim B, Lee H, et al. AudioCaps: Generating captions for audios in the wild. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). 2019: 119-132.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

DCASE 2024 Task 9: Language-Queried Audio Source Separation | Development Set

<p><strong>== Description ==&nbsp;</strong></p> <p>The development set is composed of audio samples from FSD50K [1] and Clotho v2 [2] datasets. FSD50K contains over 51k audio clips (~100 hours) manually labeled using 200 classes drawn from the AudioSet Ontology. For each audio clip in the FSD50K dataset, we generated one automatic caption for each audio clip by prompting ChatGPT (GPT-4) with its sound event tags. All audio files should be converted to mono 16 kHz audio for training LASS models.&nbsp;</p> <p>Clotho v2: <a href="../records/4783391">https://zenodo.org/records/4783391</a></p> <p>FSD50K: <a href="../records/4060432">https://zenodo.org/records/4060432</a></p> <p>Automatic captions generated for FSD50K:</p> <ul> <li>fsd50k_dev_auto_caption.json</li> <li>fsd50k_eval_auto_caption.json</li> </ul> <p>Prompt for generating captions:</p> <blockquote> <p>I will give you a number of lists containing sound events. Please write an one-sentence audio caption to describe these sounds.</p> <p>Make sure you are using grammatical subject-verb-object sentences. Directly describe the sounds and avoid using the word &ldquo;heard&rdquo;. Please don't describe the temporal order of these sound events. The caption should be less than 20 words.</p> </blockquote> <p>In addition to the development set, participants are free to use any external data (including private data) but are not allowed to use audio in Freesound uploaded between April and October 2023. Participants must specify all external resources utilized in their submission in the technical report.</p> <p><strong>== References ==</strong></p> <p>[1] Fonseca E, Favory X, Pons J, et al. FSD50k: an open dataset of human-labeled sound events. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 30: 829-852.</p> <p>[2] Drossos K, Lipping S, Virtanen T. Clotho: An audio captioning dataset. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2020: 736-740.</p> <p><strong>== Contact ==</strong></p> <p>Xubo Liu, xubo.liu@surrey.ac.uk</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

DCASE 2022 Task 5: Few-shot Bioacoustic Event Detection Evaluation Set

<p><strong>General Description</strong></p> <p>The evaluation set for task 5 of DCASE 2022&nbsp;&quot;Few-shot Bioacoustic Event Detection&quot; consists of 46 audio files acquired from different bioacoustic sources.&nbsp;</p> <p>The first 5 annotations are provided for each file, with events marked as positive (POS) for the class of interest.&nbsp;</p> <p>This dataset is to be used for evaluation purposes during the task</p> <p><strong>Folder Structure</strong></p> <p><em>Evaluation_Set.zip</em></p> <p>&nbsp; &nbsp; |___DC/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; |___CT/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; |___CHE/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; |___MGE/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; |___MS/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp; &nbsp; |___QU/</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.wav</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; |____*.csv</p> <p>&nbsp;</p> <p><em>Evaluation_Set_5shots.zip</em>&nbsp;has the same structure but contains only the *.wav files.</p> <p><em>Evaluation_Set_5shots_annotations_only.zip</em>&nbsp;has the same structure but contains only the *.csv files</p> <p>The subfolders denote different recording sources and there may or may not be overlap between classes of interest from different wav files.</p> <p><strong>Annotation structure</strong></p> <p>Each line of the annotation csv represents an event in the audio file. The column descriptions are as follows:<br> [ Audiofilename, Starttime, Endtime, Q ]</p> <p><strong>Development Set</strong></p> <p>The development set for the same task can be found at:&nbsp;<a href="http://doi.org/10.5281/zenodo.4543504">https://doi.org/10.5281/zenodo.6012309</a></p> <p><strong>Open Access</strong></p> <p>This dataset is available under a Creative Commons Attribution 4.0 International (CC BY 4.0) license.<br> &nbsp;</p> <p><strong>Contact info</strong></p> <p>Please send any feedback or questions to:<br> Ines Nolasco: i.dealmeidanolasco@qmul.ac.uk</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record