Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

33

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

33 results for “source separation”

Learn how ShareScore rates datasets ↗
zenodo32/100

DCASE 2024 Task 9: Language-Queried Audio Source Separation | Pre-trained Weights for the Baseline System

<p><strong>== Descriptions ==</strong></p> <p>We trained the AudioSep [1] model using the <a href="../records/10887496">development set</a> (Clotho and augmented FSD50K datasets) for 200k steps with a batch size of 16 using one Nvidia A100 GPU (around 1 day). Model details can be found in the <a href="https://arxiv.org/abs/2308.05037">AudioSep paper</a>.</p> <p>Pre-trained weights for the baseline system:</p> <ul> <li>audiosep_16k,baseline,step=200000.ckpt</li> </ul> <p>Baseline codebase:</p> <ul> <li>GitHub: <a href="https://github.com/Audio-AGI/dcase2024_task9_baseline">https://github.com/Audio-AGI/dcase2024_task9_baseline</a></li> </ul> <p><strong>== Reference ==</strong></p> <p>[1] Liu X, Kong Q, Zhao Y, et al. Separate anything you describe. arXiv:2308.05037, 2023.</p> <p><strong>== Contact ==</strong></p> <p>Xubo Liu, xubo.liu@surrey.ac.uk</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

DREANSS: DRum Event ANnotations for Source Separation

<p>The purpose of the annotations is to help develop research in source separation methods for polyphonic audio music mixtures containing drums.</p> <p>We provide a dataset that contains annotations for 22 excerpts of songs taken from different multi-track audio datasets publicly available for research purposes. These multi-track excerpts range from several genres including Rock, Reggae, electronic, Indie and Metal. The excerpts have a duration of 10 seconds each in average. This annotations dataset is divided into four folders, each of which contains the annotations of a given audio source separation dataset.</p> <p>The audio datasets that have been annotated are:</p> <ul> <li> <p>bss_oracle :&nbsp;<a href="http://bass-db.gforge.inria.fr/bss_oracle/">http://bass-db.gforge.inria.fr/bss_oracle/</a></p> </li> <li> <p>mass:&nbsp;<a href="https://www.upf.edu/web/mtg/audio-signal-separation">https://www.web/mtg/audio-signal-separation</a></p> </li> <li> <p>sisec:&nbsp;<a href="http://sisec.wiki.irisa.fr/tiki-index.php?page=Professionally+produced+music+recordings">http://sisec.wiki.irisa.fr/tiki-index.php?page=Professionally+produced+music+recordings</a></p> </li> </ul> <p>Please Acknowledge DREANSS in Academic Research</p> <p><strong>Using this dataset</strong></p> <p>When the DREANSS dataset is used for academic research, we would highly appreciate if scientific publications of works partly based on the DREANSS dataset cite the following publication:</p> <blockquote> <p>Ricard Marxer, Jordi Janer, &quot;<a href="http://mtg.upf.edu/node/2802">Study of Regularizations and Constraints in NMF-Based Drums Monaural Separation</a>&quot;,&nbsp;Proc. of the 7th Int. Conference on Digital Audio Effects (DAFx&rsquo;13). Maynooth, Ireland, 2013.</p> </blockquote> <p>We are interested in knowing if you find our datasets useful! If you use our dataset please email us at <a href="mailto:mtg-info@upf.edu">mtg-info@upf.edu</a> and tell us about your research.</p> <p>&nbsp;</p> <p><a href="https://www.upf.edu/web/mtg/dreanss">https://www.upf.edu/web/mtg/dreanss</a></p>

opencc-by-nc-4.0Oct 2013View details →
zenodo32/100

Unison Source Separation Dataset

<p>The unison source separation data set consist of several single, isolated instruments notes all of the same fundamental frequency (C4).</p> <p>For more information our <a href="http://see: https://www.audiolabs-erlangen.de/resources/2014-DAFx-Unison/">accompanying website</a></p>

opencc-by-4.0Aug 2014View details →
zenodo32/100

Dataset for: An optimised organic carbon / elemental carbon (OC/EC) fraction separation method for radiocarbon source apportionment applied to low-loaded Arctic aerosol filters

<p>Dataset contains raw&nbsp;data and R scripts to create the figures.</p>

opencc-by-4.0Feb 2023View details →
dryad32/100

Data from: Separating sources of density-dependent and density-independent establishment limitation in invading species

Open the record for dataset details and reuse information.

publicSep 2017View details →
dryad32/100

Animal acoustic identification, denoising, and source separation using generative adversarial networks

Open the record for dataset details and reuse information.

publicAug 2025View details →
zenodo28/100

A large joint sound scene and sound event dataset for source separation of foreground sound events

<p>This large scale data set contains 10000 samples of sound scenes generated from real world recordings, and the original source recordings. It includes 10 different backgrounds with 6-9 appropriate foreground sound events. Strong labels (timed annotations) are provided in four formats for all samples. The original sourceids to identify the class type, a two source method to simply separate foreground and backgrounds, a 32 source annotation for all distinct foregrounds, and a by background (scene) type annotation where sources are according to the background.&nbsp;</p> <p>Baseline results will be presented later in 2020. Further evolutions of this dataset will also be produced with more complex, polyphonic foreground sound events. Please email h.bear@qmul.ac.uk with any questions.</p> <p>Data is free to use for Research purposes only.&nbsp;&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo28/100

MUSDB18-HQ Test Set Inference Outputs for Models from "A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation"

Open the record for dataset details and reuse information.

openapache2.0Nov 2023View details →
zenodo28/100

Source Separation Dataset for Knowledge Boosting (Part 2)

<p><strong>Part 2 of the Source Separation Dataset</strong> as described in&nbsp;<em>Knowledge boosting during low-latency inference</em>&nbsp;(Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel &nbsp;technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a &nbsp;streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications.&nbsp;</p>

openJul 2024View details →
zenodo28/100

Source Separation Dataset for Knowledge Boosting (Part 1)

<p><strong>Part 1 of the Source Separation Dataset</strong> as described in&nbsp;<em>Knowledge boosting during low-latency inference</em>&nbsp;(Interspeech 2024)</p> <p><strong>Abstract:</strong> Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel &nbsp;technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a &nbsp;streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications.&nbsp;</p>

openJul 2024View details →
zenodo28/100

SMS-WSJ: A database for in-depth analysis of multi-channel source separation algorithms

<p>More Information can be found on our github page:&nbsp;<a href="https://github.com/fgnt/sms_wsj">https://github.com/fgnt/sms_wsj</a></p>

opencc-by-4.0Oct 2019View details →
zenodo28/100

Audio-visual sound source localization and separation

<p>CVPR 2021 tutorial</p>

opencc-by-4.0Jul 2021View details →
nasa16/100

Blind Source Separation

Blind source separation in Simulink using STFT and inverse STFT (Signal processing blockset).

restrictednotspecifiedMar 2025View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record