Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “MUSDB”
musdb-XL
<p>This is the preparation data needed to make musdb-L and musdb-XL data. Musdb-L and musdb-XL are the variants of musdb18-HQ dataset. We built musdb-L and XL by applying the iZotope Ozone 9 Maximizer (which is a widely used commercial digital limiter) to the original musdb18-HQ dataset. We included all parameter settings that were used in making musdb-L and XL.</p> <p> </p> <p>Note that the data in this page is not an audio waveform itself, this is just a sample-wise (element-wise) metadata of gain ratio between original musdb-hq vs musdb-L or XL. Check our github repository (https://github.com/jeonchangbin49/musdb-XL) or paper (https://arxiv.org/abs/2208.14355) for the detailed explanation.</p>
Musdb-XL-train
<p>Here, we present the musdb-XL-train dataset for training De-Limiter networks.</p> <p> </p> <p>%%% Important Notes (2024-06-21) %%%</p> <div> <div>We recently discovered some errors in the musdb-XL-train dataset. Specifically, about 7% of the training data (ozone_seg_0.wav ~ ozone_seg_20000.wav) had slight phase shift problems. If you are already using the musdb-XL-train dataset, please download the updated version. Sorry for the inconvenience. </div> </div> <p>%%%%%%%%%%%%%%%%%%%%%%</p> <p> </p> <p> </p> <p>The musdb-XL-train dataset consists of a limiter-applied 300,000 segments of 4-sec audio segments and the 100 original songs. For each segment, we randomly chose arbitrary segment in 4 stems (vocals, bass, drums, other) of musdb-HQ training subset and randomly mixed them. Then, we applied a commercial limiter plug-in to each stem.</p> <p> </p> <p>Once you finish the download, you have to unzip it. The data is about 200~210GB so please be sure to make enough space.</p> <p>Due to the copyright issue, the dataset contains the sample-wise gain parameters (in .npy files), instead of a wave file itself, to make each wave file of musdb-XL-train data from the musdb18-HQ dataset. You should first prepare the musdb18-HQ dataset (https://zenodo.org/record/3338373). With the musdb18-HQ and this downloaded data (.npy and .csv), run the data processing code in our GitHub (https://github.com/jeonchangbin49/De-limiter, Please check the 'Musdb-XL-train' section). Then, you can get the actual wave files of musdb-XL-train data. After finishing the data processing step, you can remove the "np_ratio" folder that contains the sample-wise gain ratio parameters but you should keep your csv files because they will be used in our training process. </p> <p> </p> <p>Notice that our previous musdb-XL (https://zenodo.org/record/7041331) data is an evaluation dataset, and musdb-XL-train is a training dataset.</p> <p> </p> <p>--Dataset Construction</p> <p>For a commercial limiter plug-in, we used the iZotope Ozone 9 Maximizer, following our previous work, musdb-XL, which is a mastering-finished (in terms of a limiter, not an EQ) version of musdb-HQ test subset.</p> <p>The threshold parameters (related to the amount of a limiter operated) of the Ozone 9 Maximizer were chosen targeting the randomly selected loudness that sampled from the Gaussian distribution (mean -8, std 1). Parameters of the Gaussian distribution were selected following statistics of recent pop music (Refer the Table 1. of our previous paper, https://arxiv.org/abs/2208.14355).</p> <p>The character parameters (related to the attack and release parameters) of the limiter were randomly sampled from the gamma distribution (a=2, scale=1, in https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.gamma.html). </p> <p>The information on random mix parameters (gain and channel swap) is contained as csv files in our dataset.</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.