Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
33
datasets available to search
ShareScore release 0.7.1
Dataset results
33 results for “source separation”
DCASE 2024 Task 9: Language-Queried Audio Source Separation | Pre-trained Weights for the Baseline System
<p><strong>== Descriptions ==</strong></p> <p>We trained the AudioSep [1] model using the <a href="../records/10887496">development set</a> (Clotho and augmented FSD50K datasets) for 200k steps with a batch size of 16 using one Nvidia A100 GPU (around 1 day). Model details can be found in the <a href="https://arxiv.org/abs/2308.05037">AudioSep paper</a>.</p> <p>Pre-trained weights for the baseline system:</p> <ul> <li>audiosep_16k,baseline,step=200000.ckpt</li> </ul> <p>Baseline codebase:</p> <ul> <li>GitHub: <a href="https://github.com/Audio-AGI/dcase2024_task9_baseline">https://github.com/Audio-AGI/dcase2024_task9_baseline</a></li> </ul> <p><strong>== Reference ==</strong></p> <p>[1] Liu X, Kong Q, Zhao Y, et al. Separate anything you describe. arXiv:2308.05037, 2023.</p> <p><strong>== Contact ==</strong></p> <p>Xubo Liu, xubo.liu@surrey.ac.uk</p>
DREANSS: DRum Event ANnotations for Source Separation
<p>The purpose of the annotations is to help develop research in source separation methods for polyphonic audio music mixtures containing drums.</p> <p>We provide a dataset that contains annotations for 22 excerpts of songs taken from different multi-track audio datasets publicly available for research purposes. These multi-track excerpts range from several genres including Rock, Reggae, electronic, Indie and Metal. The excerpts have a duration of 10 seconds each in average. This annotations dataset is divided into four folders, each of which contains the annotations of a given audio source separation dataset.</p> <p>The audio datasets that have been annotated are:</p> <ul> <li> <p>bss_oracle : <a href="http://bass-db.gforge.inria.fr/bss_oracle/">http://bass-db.gforge.inria.fr/bss_oracle/</a></p> </li> <li> <p>mass: <a href="https://www.upf.edu/web/mtg/audio-signal-separation">https://www.web/mtg/audio-signal-separation</a></p> </li> <li> <p>sisec: <a href="http://sisec.wiki.irisa.fr/tiki-index.php?page=Professionally+produced+music+recordings">http://sisec.wiki.irisa.fr/tiki-index.php?page=Professionally+produced+music+recordings</a></p> </li> </ul> <p>Please Acknowledge DREANSS in Academic Research</p> <p><strong>Using this dataset</strong></p> <p>When the DREANSS dataset is used for academic research, we would highly appreciate if scientific publications of works partly based on the DREANSS dataset cite the following publication:</p> <blockquote> <p>Ricard Marxer, Jordi Janer, "<a href="http://mtg.upf.edu/node/2802">Study of Regularizations and Constraints in NMF-Based Drums Monaural Separation</a>", Proc. of the 7th Int. Conference on Digital Audio Effects (DAFx’13). Maynooth, Ireland, 2013.</p> </blockquote> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p> </p> <p><a href="https://www.upf.edu/web/mtg/dreanss">https://www.upf.edu/web/mtg/dreanss</a></p>
Unison Source Separation Dataset
<p>The unison source separation data set consist of several single, isolated instruments notes all of the same fundamental frequency (C4).</p> <p>For more information our <a href="http://see: https://www.audiolabs-erlangen.de/resources/2014-DAFx-Unison/">accompanying website</a></p>
Dataset for: An optimised organic carbon / elemental carbon (OC/EC) fraction separation method for radiocarbon source apportionment applied to low-loaded Arctic aerosol filters
<p>Dataset contains raw data and R scripts to create the figures.</p>
Data from: Separating sources of density-dependent and density-independent establishment limitation in invading species
Open the record for dataset details and reuse information.
Animal acoustic identification, denoising, and source separation using generative adversarial networks
Open the record for dataset details and reuse information.
A large joint sound scene and sound event dataset for source separation of foreground sound events
<p>This large scale data set contains 10000 samples of sound scenes generated from real world recordings, and the original source recordings. It includes 10 different backgrounds with 6-9 appropriate foreground sound events. Strong labels (timed annotations) are provided in four formats for all samples. The original sourceids to identify the class type, a two source method to simply separate foreground and backgrounds, a 32 source annotation for all distinct foregrounds, and a by background (scene) type annotation where sources are according to the background. </p> <p>Baseline results will be presented later in 2020. Further evolutions of this dataset will also be produced with more complex, polyphonic foreground sound events. Please email h.bear@qmul.ac.uk with any questions.</p> <p>Data is free to use for Research purposes only. </p>
MUSDB18-HQ Test Set Inference Outputs for Models from "A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation"
Open the record for dataset details and reuse information.
Source Separation Dataset for Knowledge Boosting (Part 2)
<p><strong>Part 2 of the Source Separation Dataset</strong> as described in <em>Knowledge boosting during low-latency inference</em> (Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. </p>
Source Separation Dataset for Knowledge Boosting (Part 1)
<p><strong>Part 1 of the Source Separation Dataset</strong> as described in <em>Knowledge boosting during low-latency inference</em> (Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. </p>
SMS-WSJ: A database for in-depth analysis of multi-channel source separation algorithms
<p>More Information can be found on our github page: <a href="https://github.com/fgnt/sms_wsj">https://github.com/fgnt/sms_wsj</a></p>
Audio-visual sound source localization and separation
<p>CVPR 2021 tutorial</p>
Blind Source Separation
Blind source separation in Simulink using STFT and inverse STFT (Signal processing blockset).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.