Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

419

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

419 results for “large dataset”

Learn how ShareScore rates datasets ↗
zenodo44/100

Audiovisual Moments in Time: A Large-Scale Annotated Dataset of Audiovisual Actions

<p>We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Time dataset (MIT). For each trial, participants assessed whether the labelled audiovisual action event was present and whether it was the most prominent feature of the video. The dataset includes the annotation of 57,177 audiovisual videos, each independently evaluated by 3 of 11 trained participants. From this initial collection, we created a curated test set of 16 distinct action classes, with 60 videos each (960 videos). We also offer 2 sets of pre-computed audiovisual feature embeddings, using VGGish/YamNet for audio data and VGG16/EfficientNetB0 for visual data, thereby lowering the barrier to entry for audiovisual DNN research. We further carried out an experiment to explore the utility of the AVMIT annotations and feature embeddings. A series of 6 Recurrent Neural Networks (RNNs) were trained on either AVMIT-filtered audiovisual events or modality-agnostic events from MIT, and then tested on our audiovisual test set. In all RNNs, top 1 accuracy was increased by 2.71-5.94\% by training exclusively on audiovisual events, even outweighing a three-fold increase in training data. We anticipate that the newly annotated AVMIT dataset will serve as a valuable resource for research and comparative experiments involving computational models and human participants, specifically when addressing research questions where audiovisual correspondence is of critical importance.</p>

opencc-byAug 2023View details →
zenodo44/100

Dataset and Software for The Relationships Between Large-scale Variations in Shear Velocity, Density, and Compressional Velocity in the Earth's Mantle

<p><strong>Is there a chemically distinct reservoir in the Earth?</strong><br><strong>Do superplumes overly denser-than-average material?</strong><br><strong>Can we detect these anomalies with seismic data?</strong><br><strong>Can we evaluate statistical significance of the features in tomography?</strong></p><p>This study presents the <strong>strongest evidence</strong> to date (ca. 2015) of <strong>large-scale thermo-chemical heterogeneities in the lowermost mantle</strong> using the full spectrum of seismic data. A large data set of surface-wave phase anomalies, body-wave travel times, normal-mode splitting functions and long-period waveforms is used to investigate the scaling between shear velocity, density and compressional velocity in the Earth's mantle (ϱ=dln ρ/dln vS, ν=dln vS/dln vP). Our preferred joint model consists of denser-than-average anomalies (∼1% peak-to-peak) at the base of the mantle roughly coincident with the low-velocity superplumes. The relative variation of shear velocity, density and compressional velocity in our study disfavors a purely thermal contribution to heterogeneity in the lowermost mantle, with implications for the long-term stability and evolution of superplumes.</p><p><strong>Note on Odd Degree Structure:</strong></p><p>Since the self-coupled normal-mode splitting observations constrain only even-degree density variations, all inversions strongly disfavored even-degree vS-ρ correlation (R2 ~ –0.46 to –0.25) in the lowermost mantle, which also disfavors a purely thermal contribution to heterogeneity in this region. However, the starting assumptions on positive vS-ρ correlation persisted&nbsp;in the remaining&nbsp;regions and for odd degree variations. In viscosity inversions with the geoid, opposing sign of the correlation of the longest wavelength even-versus odd-degree structure maps into a region of reduced viscosity in the lower mantle (Rudolph et al., 2020, doi:10.1029/2020gc009335). While important for such dynamical implications, <strong>odd-degree density variations in the lowermost mantle&nbsp;are poorly constrained in this study and should not be interpreted</strong>. We therefore used even-degree variations up to degree 6 for our inferences on&nbsp;thermo-chemical variations in the lowermost mantle (Figure 14), and provide those values in the files below.</p><p><strong>Feedback/Questions?</strong> Please contact Raj Moulik (<a href="https://rajmoulik.com">rajmoulik.com</a>) at <a href="mailto:moulik@caa.columbia.edu?subject=Query%20from%20Zenodo">moulik@caa.columbia.edu</a>&nbsp;</p><p><strong>Reference:</strong></p><p><i>Please cite the following work if you use this data or software.</i></p><ul><li>Moulik, P. &amp; Ekström, G., 2016. The relationships between large-scale variations in shear velocity, density and compressional velocity in the Earth's mantle,&nbsp;<i>J. Geophys. Res.</i>,&nbsp;<strong>121</strong>, doi:&nbsp;<a href="http://dx.doi.org/10.1002/2015JB012679">10.1002/2015JB012679</a>.&nbsp;<a href="https://rajmoulik.com/Publications/MoulikEkstrom_JGR2016.pdf"><i>pdf</i></a></li></ul><p><i>You can also cite the dataset and software&nbsp;from this Zenodo page (Optional).</i></p><p>Moulik, P. &amp; Ekström, G. (2016). Dataset and Software for The Relationships Between Large-scale Variations in Shear Velocity, Density, and Compressional Velocity in the Earth's Mantle. In J. Geophys. Res. Solid Earth (v1.0, Vol. 121, pp. 2737–2771). Zenodo. doi:&nbsp;<a href="https://doi.org/10.5281/zenodo.8356540">10.5281/zenodo.8356540</a></p><p><strong>Data Products:</strong></p><ul><li><strong>ME16_Figures(</strong><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/ME16_Figures.tar.gz"><strong>.tar.gz</strong></a><strong>&nbsp;or&nbsp;</strong><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/ME16_Figures.pdf"><strong>.pdf</strong></a><strong>)</strong>&nbsp;- contains all figures from the paper in .png format</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/ME16"><strong>ME16</strong></a><strong>&nbsp;-&nbsp;</strong>Coefficients of the spline basis functions for each parameter. Refer cij&nbsp;in equation 3.&nbsp; This is our preferred global model of anisotropic elastic parameters and density. Density variations are allowed to deviate from a constant scaling with shear-velocity variations in the lowermost mantle, which is required to fit the longest-period normal modes (e.g.&nbsp;0S2). Radial anisotropy is confined to the uppermost mantle (that is, since the anisotropy is parameterized with only the four uppermost&nbsp;splines, it becomes very small below a depth of 250 km, and vanishes at 410 km). This is an updated version of S362ANI+M (Moulik and Ekström, 2014) which did not solve independently for density and compressional-wave velocity variations and imposed a constant scaling throughout the mantle instead (ϱ=0, ν=1/0.55).</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/STW105"><strong>STW105</strong></a>&nbsp;- reference model used in ME16. Described in Kustowski et al. (2008)</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/setup.cfg"><strong>setup.cfg</strong></a><strong>&nbsp;-&nbsp; </strong>Some configuration metadata relevant to this model for reproducibility.</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/epix.tar.gz"><strong>epix.tar.gz</strong></a>&nbsp;- Perturbations in horizontally (<i>vsh</i>) and vertically polarized shear velocity (<i>vsv</i>), Voigt-average isotropic shear-wave (<i>vs</i>) and compressional-wave velocity (<i>vp</i>), density (<i>rho</i>). anisotropy (<i>as</i>) and topography of the internal boundaries. This is calculated from the spline coefficients at&nbsp;every 1 by 1 degree cell-centered pixel and at every ~25 km depth region from Moho to the core-mantle boundary and stored in extended pixel format (.epix) ASCII files. Even-degree variations up to degree 6 are provided for density (<i>rho_even6)</i>&nbsp;and isotropic shear-wave&nbsp;velocity (<i>vs_even6</i>), which should be used for density inferences on thermochemical structure (See note above).</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/ME16.BOX25km_PIX1X1.avni.nc4"><strong>ME16.BOX25km_PIX1X1.avni.nc4</strong></a>&nbsp;-&nbsp; The perturbations in a standard AVNI format that utilizes the NETCDF4 container format. This file can be read in Python using either xarray or AVNI libraries. For example, to plot even-degree variations up to degree 6 in&nbsp;Voigt-averaged shear velocity&nbsp;perturbations at the bottom of the mantle (2875-2891 km depth)<ul><li><i>import xarray as xr</i></li><li><i>ds = xr.open_dataset('ME16.BOX25km_PIX1X1.avni.nc4')</i></li><li><i>ds['vs_even6'][-1].plot()</i></li></ul></li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/PROGRAMS.tar.gz"><strong>PROGRAMS.tar.gz</strong></a>&nbsp;- Fortran tools for obtaining model values at specific locations. After creating the executables from source code in the&nbsp;<i>src</i>&nbsp;folder, the&nbsp;<i>readme</i>&nbsp;script generates most of&nbsp;the epix files provided in&nbsp;epix.tar.gz above<strong>.</strong></li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/profilescaling.txt"><strong>profilescaling.txt</strong></a>&nbsp;- contains the median scaling ratios as used in Figure 15(a).</li><li><a href="https://zenodo.org/api/files/9aa99409-20ae-495e-b5a3-288fd57ecaeb/scaling3D_MoulikJGR16.tar.gz"><strong>scaling3D_MoulikJGR16.tar.gz</strong></a>&nbsp;- contains the scaling ratios and poisson ratio calculated from the joint model, as used in Figure 15(b).</li></ul>

opengpl-2.0-or-laterApr 2016View details →
zenodo40/100

Dataset related to the publication "Transformation Optics: Large Multiphysics Simulation of Nonlinear Optomechanical Coupling in Microstructured Resonant Cavities", DOI: 10.1109/MMM.2018.2821086

<p>This folder contains the raw data from which the graphs in paper &quot;Transformation Optics: Large Multiphysics Simulation of Nonlinear Optomechanical Coupling in Microstructured Resonant Cavities&quot;, DOI: 10.1109/MMM.2018.2821086, have been obtained.</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

Dataset for publication: Validation of large-volume batch solar reactors for the treatment of rainwater in field trials in sub-Saharan Africa, Reyneke et al. (2020). DOI: 10.1016/j.scitotenv.2020.137223

<p>Datasets, Supplementary Information and Water Safety Plan (Assessment Form and Risk Matrix) for the publication: &quot;Validation of large-volume batch solar reactors for the treatment of rainwater in field trials in sub-Saharan Africa&quot; which was published in Science of the Total Environment (https://doi.org/10.1016/j.scitotenv.2020.137223).</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Preference and familiarity mediate spatial responses of a large herbivore to experimental manipulation of resource availability: datasets

<p>Publicly available dataset for:</p> <p>N. Ranc, P.R. Moorcroft, K.W. Hansen, F. Ossi, T. Sforna, E. Ferraro, A. Brugnoli &amp; F. Cagnacci. 2020. Preference and familiarity mediate spatial responses of a large herbivore to experimental manipulation of resource availability. Scientific Reports. https://doi.org/10.1038/s41598-020-68046-7</p>

opencc-by-4.0Jun 2020View details →
zenodo40/100

A large-scale wide-baseline light field dataset - Part I

<p>This dataset is the Part I of a large-scale, synthetic wide-baseline light field dataset (called WLF), including 345 light fields.&nbsp;</p> <p>Each light field provides 9x9 angular (RGB) images and ground truth disparities. This&nbsp;light field dataset&nbsp;involves the spatial&nbsp;resolution (512x512) images only.</p> <p>The dataset is originally created for training the deep learning-based models for depth estimation. You might use&nbsp;this dataset for other tasks if possible.</p> <p>You might also have a try to play with this dataset using our code in&nbsp;https://github.com/YanWQ/LLF-Net.</p>

opencc-by-4.0Jul 2020View details →
zenodo40/100

ToxicoDB: an integrated database to mine and visualize large-scale toxicogenomic datasets (TGGATEs human dataset)

<p>This data was generated by Igarashi Y, Nakatsu N, Yamashita T, Ono A, Ohno Y, Urushidani T, Yamada H. Open TG-GATEs: a large-scale toxicogenomics database. Nucleic Acids Res [Internet]. 2015 Jan;43(Database issue):D921&ndash;7. Available from: http://dx.doi.org/10.1093/nar/gku955 PMCID: PMC4384023. The data have been curated and analyzed using our open-source R package, <em>ToxicoGx</em> (<a href="https://github.com/bhklab/ToxicoGx">https://bioconductor.org/packages/devel/bioc/html/ToxicoGx.html</a>), and are available publicly in the <em>ToxicoDB </em>web application (<a href="http://www.toxicodb.ca">www.toxicodb.ca</a>).</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

A mixed mode cohesive model for FRP laminates incorporating large scale bridging behaviour - Datasets

<p>This data upload includes the experimental results from delaminating FRP-laminates. The experiment consists of DCB specimens where the beam ends are loaded with bending moments. A set-up of LVDTs and a clip-on extensometer are used to calculate the normal and tangential opening displacements at the crack-end.</p> <ul> <li>The test specimens are described in the file &quot;CHO test matrix 130405B.xlsx&quot;</li> <li>The load-displacement data for all specimens are given in the folder &quot;DCB UBM - Experimental results.zip&quot;</li> <li>Acoustic emission recording from the tests are given in the folder &quot;DCB UBM - Acoustic Emission.zip&quot;</li> <li>A set of images for each specimen during testing is given in the folder &quot;DCB Images.zip&quot;</li> </ul> <p>This test series is examined and described in the following peer reviewed papers:</p> <p>R.K. Joki, F. Grytten, B. Hayman, B.F. S&oslash;rensen, <em>A mixed mode cohesive model for FRP laminates incorporating large scale bridging behaviour</em>, Engineering Fracture Mechanics, 239, November 2020,&nbsp; <a href="https://doi.org/10.1016/j.engfracmech.2020.107274">https://doi.org/10.1016/j.engfracmech.2020.107274</a></p> <p>R.K. Joki, F. Grytten, B. Hayman, B.F. S&oslash;rensen, <em>Determination of a cohesive law for delamination modelling &ndash; Accounting for variation in crack opening and stress state across the test specimen width</em>, Composites Science and Technology, 128, 18 May 2016, <a href="https://doi.org/10.1016/j.compscitech.2016.01.026">https://doi.org/10.1016/j.compscitech.2016.01.026</a></p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

The dataset of figures for "The GPU version of LICOM3 under HIP framework and its large-scale application" (updated)

<p>A high-resolution (1/20&deg;) global ocean general circulation model with Graphics processing units (GPUs) code implementations is developed based on the LASG/IAP Climate system Ocean Model version 3 (LICOM3) under Heterogeneous-compute Interface for Portability (HIP) framework. The dynamic core and physics package of LICOM3 are both ported to the GPU, and 3-dimensional parallelization is applied. The HIP version of the LICOM3 (LICOM3-HIP) is 42 times faster than what the same number of CPU cores dose, when 384 AMD GPUs and CPU cores are used. The LICOM3-HIP has excellent scalability; it can still obtain speedup of more than four on 9216 GPUs comparing to 384 GPUs. In this phase, we successfully performed a test of 1/20&deg; LICOM3-HIP using 6550 nodes and 26200 GPUs, and at the grand scale, the model&rsquo;s time to solution can still obtain an increasing, about 2.72 simulated years per day (SYPD). The high performance was due to putting almost all of computation processes inside GPUs, and thus greatly reduces the time cost of data transfer between CPUs and GPUs. At the same time, a 14-year spin-up integration following the phase 2 of Ocean Model Intercomparison Project (OMIP-2) protocol of surface forcing has been conducted, and the preliminary results have been evaluated. We found that the model results have little differences from the CPU version. Further comparison with observations and lower-resolution LICOM3 results suggests that the 1/20&deg; LICOM3-HIP can not only reproduce the observations, but also produce much smaller scale activities, such as submesoscale eddies and frontal scales structures.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

High-quality large curated dataset of protein sequences (1.83 million) and their corresponding Position Specific Scoring Matrices

<p>As part of his&nbsp;master thesis at the Rostlab, which is located at the Technical University of Munich (TUM),&nbsp;Mr. Issar Arab&nbsp;developed&nbsp;the first language model that encodes evolutionary information of proteins explicitly. The pre-training involved the creation of a novel high-quality dataset of protein sequences (around 1.83&nbsp;million proteins, or ~0.8 Billion amino acids) with their corresponding Position Specific Scoring Matrices (PSSMs).&nbsp; Those matrices reflect the relative frequency of each amino acid at each position in a protein and is derived from evolutionarily related proteins.</p> <p>Mr. Arab makes this work publicly available to help other researchers speed up their work to leverage AI to learn the representation of protein evolutionary information more explicitly. The set of sequences was derived by extracting all PSSMs from the&nbsp;<a href="https://predictprotein.org/">PredictProtein</a>&nbsp;(PP) cache, which&nbsp;were also part o the UniProt&nbsp;Reference Cluster with 50% sequence identity (uniref50 2019_12). The overlap between PP and uniref50 was further filtered to only include high-quality samples, e.g. only multiple sequence alignments with a certain number of aligned sequences were considered. The processing led to&nbsp;a training set of 1.83 Million sequences, a validation set of 879 instances, and a test set of 879 entries.&nbsp;The training data of proteins is reduced to 40% sequence identity, with respect to the validation/test sets, and contains sequences ranging between 18 and 9858 residues in length.</p> <p>Refer to the Jupyter notebook for a detailed description of the files'&nbsp;structure and a Python code snippet to correctly manipulate&nbsp;this data.</p> <p>To access the full original&nbsp;work, please visit the following link:&nbsp; <a href="https://mediatum.ub.tum.de/node?id=1579236">Manuscript</a>&nbsp;<br><br><strong>Note:</strong> The dataset was recently used to fine tune a protein sequence language model (<a href="https://github.com/issararab/PEvoLM">PEvoLM</a>). The work was presented at the CIBCB'23 conference. If you use PEvoLM or this dataset in your work, please cite the following publication:</p> <p>- Issar Arab, <strong>PEvoLM: Protein Sequence Evolutionary Information Language Model</strong>, <em>IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), Eindhoven, Netherlands</em>, (2023), pp. 1-8, doi:<a href="https://ieeexplore.ieee.org/document/10264890">10.1109/CIBCB56990.2023.10264890</a></p>

opencc-by-4.0Sep 2020View details →
zenodo40/100

Datasets for "Hydrodynamic and hydromagnetic energy spectra from large eddy simulations"

<pre>This directory contains an index.html file with links to the run directories and idl plotting routines with secondary data for the other figures for the paper &quot;Hydrodynamic and hydromagnetic energy spectra from large eddy simulations&quot; Haugen &amp; Brandenburg. If anything turns our to be incomplete, please email brandenb@nordita.org.</pre>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Large-Scale Gravitational Lens Modeling with Bayesian Neural Networks for Accurate and Precise Inference of the Hubble Constant - Datasets, Trained Models, BNN Samples, and MCMC Chains

<p>We publish the training/validation/test datasets, trained model weights, configuration files, Bayesian neural network samples, and MCMC chains used to produce the figures in the LSST DESC paper, &quot;Large-Scale Gravitational Lens Modeling with Bayesian Neural Networks for Accurate and Precise Inference of the Hubble Constant.&quot; They are formatted to be used with the DESC package &quot;H0rton&quot; (<a href="https://github.com/jiwoncpark/h0rton">https://github.com/jiwoncpark/h0rton</a>). Additional descriptions can be found in the README. Please contact Ji Won Park (@jiwoncpark) on GitHub or <a href="https://github.com/jiwoncpark/h0rton/issues">make an issue</a> for any questions.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

DEVILS: a tool for the visualization of large datasets with a high dynamic range

<p>This repository accompanying the article &ldquo;DEVILS: a tool for the visualization of large datasets with a high dynamic range&rdquo; contains the following:</p> <ul> <li>Extended Material of the article</li> <li>An example raw dataset corresponding to the images shown in Fig. 3</li> <li>A workflow description that demonstrates the use of the DEVILS workflow with BigStitcher.</li> <li>Two scripts (&ldquo;CLAHE_Parameters_test.ijm&rdquo; and a &ldquo;DEVILS_Parallel_tests.groovy&rdquo;) used for Figure S2, S3 and S4.</li> </ul>

opencc-by-4.0Sep 2020View details →
zenodo40/100

Caravan - A global community dataset for large-sample hydrology

<p><strong>This is the </strong><strong>accompanying dataset to the following paper&nbsp;<a href="https://www.nature.com/articles/s41597-023-01975-w">https://www.nature.com/articles/s41597-023-01975-w</a></strong></p> <p><em>Caravan</em>&nbsp;is an open community dataset of meteorological forcing data, catchment attributes, and discharge daat for catchments around the world. Additionally, Caravan provides code to derive meteorological forcing data and catchment attributes from the same data sources in the cloud, making it easy for anyone to extend Caravan to new catchments. The vision of Caravan is to provide the foundation for a truly global open source community resource that will grow over time.</p> <p>If you use Caravan in your research, it would be appreciated to not only cite Caravan itself, but also the source datasets, to pay respect to the amount of work that was put into the creation of these datasets and that made Caravan possible in the first place.</p> <p><strong>All current development and additional community extensions can be found at&nbsp;<a href="https://github.com/kratzert/Caravan">https://github.com/kratzert/Caravan</a><br></strong><br><strong>IMPORTANT: Due to size limitations for individual repositories, the netCDF version and the CSV version of Caravan (since Version 1.6) &nbsp;are split into two different repositories. You can find the CSV version at <a href="https://zenodo.org/records/15530021">https://zenodo.org/records/15530021</a></strong></p> <p>Channel Log:</p> <ul> <li><strong>23 May 2022: Version 0.2</strong> - Resolved a bug when renaming the LamaH gauge ids from the LamaH ids to the official gauge ids provided as "govnr" in the LamaH dataset attribute files.</li> <li><strong>24 May 2022: Version 0.3</strong> - Fixed gaps in forcing data in some "camels" (US) basins.</li> <li><strong>15 June 2022: Version 0.4</strong> - Fixed replacing negative CAMELS US values with NaN (-999 in CAMELS indicates missing observation).</li> <li><strong>1 December 2022: Version 0.4 </strong>- Added 4298 basins in the US, Canada and Mexico (part of HYSETS), now totalling to 6830 basins. Fixed a bug in the computation of catchment attributes that are defined as pour point properties, where sometimes the wrong HydroATLAS polygon was picked. Restructured the attribute files and added some more meta data (station name and country).</li> <li><strong>16 January 2023: Version 1.0</strong> - Version of the official paper release. No changes in the data but added a static copy of the accompanying code of the paper. For the most up to date version, please check&nbsp;https://github.com/kratzert/Caravan</li> <li><strong>10 May 2023: Version 1.1</strong> -&nbsp;No data change, just update data description.</li> <li><strong>17 May 2023: Version 1.2</strong> - Updated a handful of attribute values that were affected by a bug in their derivation. See&nbsp;https://github.com/kratzert/Caravan/issues/22 for details.</li> <li><strong>16 April 2024: Version 1.4</strong> - Added 9130 gauges from the original source dataset that were initially not included because of the area thresholds (i.e. basins smaller&nbsp; than 100sqkm or larger than 2000sqkm). Also extended the forcing period for all gauges (including the original ones) to 1950-2023. Added two different download options that include timeseries data only as either csv files (Caravan-csv.tar.xz) or netcdf files (Caravan-nc.tar.xz). Including the large basins also required an update in the earth engine code</li> <li><strong>16 Jan 2025: Version 1.5</strong> - Added FAO Penman-Monteith PET (potential_evaporation_sum_FAO_PENMAN_MONTEITH) and renamed the ERA5-LAND potential_evaporation band to potential_evaporation_sum_ERA5_LAND. Also added all PET-related climated indices derived with the Penman-Monteith PET band (suffix "_FAO_PM") and renamed the old PET-related indices accordingly (suffix "_ERA5_LAND").&nbsp;</li> <li><strong>27 May 2025: Version 1.6</strong><br> <ul> <li>Updated the CAMELS-AUS data to source from CAMELS-AUS v2. This means more basins (561 compared to 222) and more recent streamflow data (2022 compared to 2014). Note that the gauge id for four basins changed between the original CAMELS-AUS version and v2. Those gauges are ['camelsaus_224213A', 'camelsaus_224214A', 'camelsaus_227225A', 'camelsaus_403213A'] that all lost their trailing "A". To stay synced with CAMELS-AUS (v2), we also adapted the new naming.</li> <li>Added VERSION file to the root directory that contains the current version number.</li> <li>Updated the code to the most recent GitHub snapshot (commit 6eab036).</li> <li>Due to the 50GB repository limit, we had to split the netCDF version and the CSV version into two separate repositories. The CSV version can be found under https://zenodo.org/records/15530021</li> </ul> </li> </ul>

opencc-by-4.0May 2022View details →
zenodo40/100

PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

<p>We introduce&nbsp;<strong>PDMX</strong>: a <strong>P</strong>ublic&nbsp;<strong>D</strong>omain&nbsp;<strong>M</strong>usic<strong>X</strong>ML dataset for symbolic music processing. Refer to our <a title="PDMX Paper" href="https://arxiv.org/abs/2409.10831" target="_blank" rel="noopener">paper</a> for more information, and our <a title="PDMX GitHub Repository" href="https://github.com/pnlong/PDMX/" target="_blank" rel="noopener">GitHub repository</a> for any code-related details. Please cite both our paper and <a href="https://arxiv.org/abs/2410.02084" target="_blank" rel="noopener">our collaborators' paper</a> if you use this dataset (see our GitHub for more information).</p> <p>Upon further use of the PDMX dataset, we discovered a discrepancy between the public-facing copyright metadata on the <a href="https://musescore.com/">MuseScore website</a> and the internal copyright data of the MuseScore files themselves, which affected 31,221 (12.29% of) songs. We have decided to proceed with the former given its public visibility on Musescore (i.e. this is what the MuseScore website presents its users with). We have noted files with conflicting internal licenses in the&nbsp;<em><strong>license_conflict</strong></em> column of PDMX. We recommend using the&nbsp;<em><strong>no_license_conflict</strong></em> subset of PDMX (which still includes 222,856 songs) moving forward.</p> <p>Additionally, for each song in PDMX, we not only provide the <em>MusicRender</em> and metadata JSON files, but we also try to include the associated compressed MusicXML (MXL), sheet music (PDF), and MIDI (MID) files when available. Due to the corruption of 42 of the original MuseScore files,&nbsp;these songs lack those associated files (since they could not be converted to those formats) and only include the <em>MusicRender</em> and metadata JSON files. The&nbsp;<em><strong>all_valid</strong></em> subset of PDMX describes the songs where all associated files are valid.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Supplemental dataset to "Cloud sync in response to wave-like large-scale forcings"

<p>This dataset deposits the supplemental materials for the manuscript "Cloud sync in response to wave-like large-scale forcings".</p> <p><a href="https://zenodo.org/api/records/14164798/draft/files/math_note.pdf/content" target="_blank" rel="noopener noreferrer">math_note.pdf</a>&nbsp; A hand-written math derivation note for equations in the appendices.&nbsp;</p> <p><a href="https://zenodo.org/uploads/15304862" target="_blank" rel="noopener noreferrer">movie_wL_006_T24hours.avi</a>&nbsp; A movie of near-surface (z=25m) water vapor mixing ratio for the wL=0.006m/s and T=24 hours experiment (the reference CM1 simulation).</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/movie_wL_006_T18hours.avi/content" target="_blank" rel="noopener noreferrer">movie_wL_006_T18hours.avi</a> &nbsp;A movie of near-surface (z=25m) water vapor mixing ratio for the wL=0.006m/s and T=18 hours experiment.</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/movie_wL_006_T12hours.avi/content" target="_blank" rel="noopener noreferrer">movie_wL_006_T12hours.avi</a> &nbsp;A movie of near-surface (z=25m) water vapor mixing ratio for the wL=0.006m/s and T=12 hours experiment.</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/movie_wL_000.avi/content" target="_blank" rel="noopener noreferrer">movie_wL_000.avi</a> &nbsp;A movie of near-surface (z=25m) water vapor mixing ratio, without large-scale wave-like forcing.</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/input_sounding/content" target="_blank" rel="noopener noreferrer">input_sounding</a>&nbsp; The initial sounding for all CM1 simulations.&nbsp;</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/namelist.input/content" target="_blank" rel="noopener noreferrer">namelist.input</a>&nbsp; The namelist file for launching all CM1 simulations.&nbsp;</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/postprocessing_CM1.zip/content" target="_blank" rel="noopener noreferrer">postprocessing_CM1.zip</a>&nbsp;The postprocessing code of the CM1 simulations, including intermediate output files (.mat) in data postprocessing.&nbsp;&nbsp;</p> <p><a href="https://zenodo.org/api/records/14164798/draft/files/microscopic_model.zip/content" target="_blank" rel="noopener noreferrer">microscopic_model.zip</a>&nbsp; The MATLAB code for the microscopic model.&nbsp;</p> <p><a href="https://zenodo.org/api/records/15304862/draft/files/cm1.F/content" target="_blank" rel="noopener noreferrer">cm1.F</a>&nbsp; The CM1 script where the large-scale vertical velocity is programmed. You can copy it directly to your CM1/src/ path.&nbsp;</p> <p>&nbsp;</p> <p>Feel free to send an email to Dr. Hao Fu (haofu@nju.edu.cn) if you have any questions!</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

StemGMD: A Large-Scale Audio Dataset of Isolated Drum Stems for Deep Drums Demixing - part 1

<p>We introduce StemGMD, a new large-scale dataset of isolated drum stems that builds upon the extensive MIDI collection found in <a href="https://magenta.tensorflow.org/datasets/groove">Magenta's Groove MIDI Dataset (GMD)</a>.</p> <p>GMD is a 13.6-hour corpus of expressive drum performances executed by ten drummers on a Roland TD-11 electronic drum kit. It contains 1150 MIDI files along with the corresponding full-kit audio mixtures.</p> <p>As a first step in creating StemGMD, we mapped the 22 different MIDI pitches found in the original files onto nine canonical instruments through the reduction scheme proposed &nbsp;in J. Gillick, A. Roberts, J. Engel, D. Eck, and D. Bamman, "Learning to groove with inverse sequence transformations," in International Conference on Machine Learning (ICML), vol. 97, 2019, pp. 2269&ndash;2279.</p> <p>Each of the nine resulting MIDI channels was manually synthesized as a 16-bit/44.1 kHz stereo WAV file using ten realistic-sounding acoustic drum kits sourced from the <a href="https://support.apple.com/en-me/guide/logicpro/lgsi2fb2509e/mac">Logic Pro X sample libraries</a>, i.e., Bluebird, Brooklyn, Detroit Garage, East Bay, Heavy, Motown Revisited, Portland, Retro Rock, Roots, and SoCal.</p> <p>As a result, StemGMD contains 1224 hours of audio, which correspond to more than 136 hours of full-kit mixtures. Moreover, StemGMD also contains single hits for each of the drum pieces at ten different velocities ranging from 30 to 127.</p> <p>To the best of our knowledge, StemGMD is the largest publicly available dataset of drums to date. Moreover, it is the first collection of single-instrument clips from all nine pieces in a canonical drum kit, making it well-suited for training deep drums demixing models.</p> <p>&nbsp;</p> <p><strong>*** THIS IS PART 1 OF 2 ***</strong></p> <p><strong>Download part 2 here:</strong> <a href="../records/7882857">https://zenodo.org/records/7882857&nbsp;</a> (now available!)&nbsp;</p> <p>After downloading both parts, run <strong>unzip_StemGMD.sh</strong> to build the dataset from the split archive files.<br>Once unzipped, StemGMD will take just over <strong>1.13 TB </strong>of memory.</p> <p>&nbsp;</p> <p>The dataset is made available under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International (CC BY 4.0) License</a>.</p> <p>____________________________</p> <p>We employed the dataset in our paper titled "Toward Deep Drum Source Separation," published in Pattern Recognition Letters.&nbsp;</p> <div> <div> <div> <div> <p>Please, cite this work as: A. I. Mezza, R. Giampiccolo, A. Bernardini, and A. Sarti, "Toward Deep Drum Source Separation," Pattern Recognition Letters, vol. 183, pp. 86-91, 2024, doi: 10.1016/j.patrec.2024.04.026.</p> <pre>@article{mezza2024, title = {Toward deep drum source separation}, author = {Alessandro Ilic Mezza and Riccardo Giampiccolo and Alberto Bernardini and Augusto Sarti}, journal = {Pattern Recognition Letters}, volume = {183}, pages = {86-91}, year = {2024}, issn = {0167-8655}, doi = {https://doi.org/10.1016/j.patrec.2024.04.026} }</pre> </div> </div> </div> </div>

opencc-by-4.0Apr 2023View details →
zenodo40/100

StemGMD: A Large-Scale Audio Dataset of Isolated Drum Stems for Deep Drums Demixing - part 2

<p>We introduce StemGMD, a new large-scale dataset of isolated drum stems that builds upon the extensive MIDI collection found in <a href="https://magenta.tensorflow.org/datasets/groove">Magenta's Groove MIDI Dataset (GMD)</a>.</p> <p>GMD is a 13.6-hour corpus of expressive drum performances executed by ten drummers on a Roland TD-11 electronic drum kit. It contains 1150 MIDI files along with the corresponding full-kit audio mixtures.</p> <p>As a first step in creating StemGMD, we mapped the 22 different MIDI pitches found in the original files onto nine canonical instruments through the reduction scheme proposed &nbsp;in J. Gillick, A. Roberts, J. Engel, D. Eck, and D. Bamman, "Learning to groove with inverse sequence transformations," in International Conference on Machine Learning (ICML), vol. 97, 2019, pp. 2269&ndash;2279.</p> <p>Each of the nine resulting MIDI channels was manually synthesized as a 16-bit/44.1 kHz stereo WAV file using ten realistic-sounding acoustic drum kits sourced from the <a href="https://support.apple.com/en-me/guide/logicpro/lgsi2fb2509e/mac">Logic Pro X sample libraries</a>, i.e., Bluebird, Brooklyn, Detroit Garage, East Bay, Heavy, Motown Revisited, Portland, Retro Rock, Roots, and SoCal.</p> <p>As a result, StemGMD contains 1224 hours of audio, which correspond to more than 136 hours of full-kit mixtures. Moreover, StemGMD also contains single hits for each of the drum pieces at ten different velocities ranging from 30 to 127.</p> <p>To the best of our knowledge, StemGMD is the largest publicly available dataset of drums to date. Moreover, it is the first collection of single-instrument clips from all nine pieces in a canonical drum kit, making it well-suited for training deep drums demixing models.</p> <p>&nbsp;</p> <p><strong>*** THIS IS PART 2 OF 2 ***</strong></p> <p><strong>Download part 1 here:</strong> <a href="../records/7860223">https://zenodo.org/records/7860223&nbsp;</a></p> <p>After downloading both parts, run <strong>unzip_StemGMD.sh</strong> to build the dataset from the split archive files.<br>Once unzipped, StemGMD will take just over <strong>1.13 TB </strong>of memory.</p> <p>&nbsp;</p> <p>The dataset is made available under a <a href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International (CC BY 4.0) License</a>.<br><br>_____________________________<br><br>We employed the dataset in our paper titled "Toward Deep Drum Source Separation," published in Pattern Recognition Letters.&nbsp;</p> <div> <div> <div> <div> <p>Please, cite this work as: A. I. Mezza, R. Giampiccolo, A. Bernardini, and A. Sarti, "Toward Deep Drum Source Separation," Pattern Recognition Letters, vol. 183, pp. 86-91, 2024, doi: 10.1016/j.patrec.2024.04.026.</p> <pre>@article{mezza2024, title = {Toward deep drum source separation}, author = {Alessandro Ilic Mezza and Riccardo Giampiccolo and Alberto Bernardini and Augusto Sarti}, journal = {Pattern Recognition Letters}, volume = {183}, pages = {86-91}, year = {2024}, issn = {0167-8655}, doi = {https://doi.org/10.1016/j.patrec.2024.04.026} }</pre> </div> </div> </div> </div>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Large Spots DeepMIB project, synthetic dataset for testing 2D semantic segmentation

<p>A complete DeepMIB project with a synthetic dataset generated for quick tests of semantic segmentation approaches.<br>The dataset includes a trained DeepLabV3-Resnet18 network for detection of large spots on a black background.&nbsp;</p><p>The network can be opened by loading "2D_LargeSpots_2cl_DeepLabV3.mibCfg" file by</p><ul><li><i>MIB-&gt;Menu-&gt;Tools-&gt;Deep learning segmentation-&gt;Options tab-&gt;Config files-&gt;Load&nbsp;</i></li><li>Drag and drop of the config file into DeepMIB window</li></ul><p>Microscopy Image Browser: <a href="https://mib.helsinki.fi">https://mib.helsinki.fi</a></p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Dataset: Evolution of large Venusian coronae inferred from structural analyses and the presence of low-angle faults in chasmata

<p>The vector datasets mentioned in this article are available. You can find the Magellan SAR data and the Global Topography Data Records (GTDR) on NASA's Planetary Data System (PDS) website (specific links can be found on https://pdsgeosciences.wustl.edu/missions/magellan/index.htm). In addition, the stereo-derived topography dataset of Herrick et al. 2012 is also available on their personal website https://sites.google.com/alaska.edu/robertherrick/resources/stereo-derived-topography-for-venus.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record