WHISPER SET 1: a dataset for multi-channel, multi-device speech separation and speech enhancement
<p>This dataset is <code>WHISPER SET 1,</code> a dataset for speech enhancement and source separation recorded with a Wireless Acoustic Sensor Network (WASN) called WHISPER <a href="https://ieeexplore.ieee.org/abstract/document/8110202">Kiselev2018</a>. The dataset contains samples for up to 4 concurrent speakers and speech in noise. The dataset was recorded in a room with low reverberation (T_60 = 0.2 s) and using 16 microphones. In general, each track contains first a calibration phase where each of the speakers sequentially is active alone for 15 seconds. Followed by 15 seconds of all the speakers together (plus noise in some cases). </p> <p>If you use this dataset please cite:</p> <ul> <li><strong>E. Ceolini, I. Kiselev and S. Liu, "Evaluating multi-channel multi-device speech separation algorithms in the wild: a hardware-software solution," in <em>IEEE/ACM Transactions on Audio, Speech, and Language Processing</em>.</strong></li> </ul> <p>===</p> <p>Each sample is a 16-channel wav file in which the order of the channel follows the following logic:</p> <p>0 - module 5 mic 1 1 - module 5 mic 2 2 - module 5 mic 3 3 - module 5 mic 4 4 - module 6 mic 1 5 - module 6 mic 2 6 - module 6 mic 3 7 - module 6 mic 4 8 - module 7 mic 1 9 - module 7 mic 2 10 - module 7 mic 3 11 - module 7 mic 4 12 - module 8 mic 1 13 - module 8 mic 2 14 - module 8 mic 3 15 - module 8 mic 4</p> <p>Refer to the <a href="https://github.com/SensorsAudioINI/WHISPER_SET_1/blob/master/WHISPER4_floor_annotated.png">floor plan</a> for a visual illustration of the microphone arrangement.</p> <p>The files are divided into two subfolders, one for the samples of speech enhancement and one for the samples of speech separation.</p> <ul> <li>In the folder of speech separation, the files are divided into subfolders defining the number of speakers in the mixtures (2, 3, or 4)</li> <li>In the folder of speech enhancement, the files are divided into subfolders following the SNR of the mixture (0, -5, -10 dB)</li> </ul> <p>Samples are ordered in folders. Each sample folder contains a 15 seconds 16-channels <code>mixture.wav</code> file, plus the 15 seconds 16-channels <code>calibX.wav</code> files one for each speaker alone or noise alone in the mixture. That is a sample with a mixture with 4 speakers will have 4 calibration files (calib1.wav, calib2.wav, calib3.wav, calib4.wav) and a mixture of a speaker plus noise will have 2 calibration files one for speech (calib1.wav) and one for noise (calib2.wav).</p> <p>== </p> <p>A Jupyter notebook is included to show an example of how to use the data of this dataset for speech separation and speech enhancement using beamforming. The notebook is dependent on <a href="https://github.com/Enny1991/beamformers">this beamforming library</a> and <a href="https://github.com/Enny1991/sep_eval">this tool</a> to evaluate the quality of the separation.</p> <p>==</p> <p>Refer to the README.md in the dataset for more information.</p> <p>For any question please contact enea.ceolini@gmail.com</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 8