Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

251

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

251 results for “deep learning models”

Learn how ShareScore rates datasets ↗
zenodo48/100

LigPCDS: Labeled Dataset of X-ray Protein Ligand Images in 3D Point Cloud and Validated Deep Learning Models

<p>The difference electron density from X-ray protein crystallography was used to create the first dataset of labeled ligand images in 3D point clouds, named <strong>LigPCDS</strong>. The dataset contain 244,226 entries of free organic ligands containing 3D representations labeled with two major labeling approaches: SP-based and AtomSymbol-based.</p> <p>&nbsp;</p> <p>The data from free organic molecules (non-covalent ligands) was retrieved from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) in december 2019 with resolutions ranging from 1.5 to 2.2 &Aring;. The ligand images (blobs) were interpolated from their calculated difference electron density map in a 3D grid-like bounding box, around their atomic positions, and stored in point clouds. These ligand grid representations were further processed to retrive the final ligands representation in 3D point clouds using a mask of the shape of the ligand. A grid spacing of 0.5 &Aring; gave the best results. The density value of the grid points was used as feature. The labeling approach used the structure of the ligands to propose vocabularies of chemical classes based on the chemical atoms themselves and their cyclic substructures. These structure annotations were applied pointwise to the ligand 3D representations using an atomic sphere model. Four proposed vocabularies were validated by successfully training good performance deep learning models for the semantic segmentation of a stratified dataset from LigPCDS, using 78902 entries.</p> <p>The four validated deep learning models are: (i) the LigandRegion, composed by generic atoms of any type; (ii) the AtomCycle, composed by generic atoms outside cycles and generic cycles; (iii) the AtomC347CA56, composed by generic atoms outside cycles, not aromatic cycles of size 3 to 7 and aromatic cycles of size 5 and 6; and (iv) the AtomSymbolGroups, composed by the atoms symbols with groupings. The mean accuracy of these models in their cross-validation was between 49.7% <span lang="EN-GB">[-19.4,20.</span><span lang="EN-GB">2]</span> and 77.4% <span lang="EN-GB">[-11.7,12.1]</span> in terms of Intersection over Union (mIoU) metric and between 62.4% <span lang="EN-GB">[-18.8,19.</span><span lang="EN-GB">7]</span> and 87.0% <span lang="EN-GB">[-8.4,8.8]</span> in F1-score (mF1), confidence interval between squared brackets. The models i, ii and iii and the used labeled representations in 3D point cloud are contained in the SP-based record; and model iv and its used labeled representations are contained in the AtomSymbol-based record.</p> <p>The dataset and validated models may be used to tackle problems regarding known and unknown ligand building to drug discovery and fragment screening pipelines.&nbsp;</p> <p>The code used to create and validated the LigPCDS is available at the following repository: https://github.com/danielatrivella/np3_ligand</p> <p>This repository also contains the NP&sup3; Blob Label application for ligand building using the validated deep learning models from LigPCDS.</p>

opencc-by-4.0May 2023View details →
zenodo48/100

Data and Models from the study entitled, "Large-area automatic detection of shoreline stranded marine debris using deep learning"

<p>This repository contains data and models used in the study entitled, "Large-area automatic detection of shoreline stranded marine debris using deep learning". This study can be accessed as an open access publication at the following location: https://doi.org/10.1016/j.jag.2023.103515.</p> <p>The data set is comprised of 1,587 images (512 pixels x 512 pixels) which contains 10,703 individual bounding box labels of marine debris objects. The imagery was collected over the State of Hawai'i in 2015 at 2 centimeter resolution (ground spacing distance).</p> <p>The classification scheme consists of 8 labeled classes: unidentified object, processed wood, metal, vessel, net/cloth, buoy, tire, and line fragments.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Deep learning to extract the meteorological by-catch of wildlife cameras: Supporting data, models and code

<p>This repository contains the data, models and code to train and deploy deep learning models related to the paper "Deep learning to extract the meteorological by-catch of wildlife cameras" published in the journal Global Change Biology (<a href="https://doi.org/10.1111/gcb.17078"><strong>https://doi.org/10.1111/gcb.17078</strong></a>).</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Deposition of data for developing deep learning models to assess crack width and self-healing progress in concrete (krkCMd)

<p>This is a deposition of data for developing deep learning models to assess crack width and self-healing progress in concrete [1]. It relates to an experimental study on the autogenous self-healing of high-strength concrete [2]. Concrete specimens were prepared, matured, cracked, and exposed to self-healing. High-resolution scanning of the specimen surface and scale-invariant image processing were performed, multiple grid lines crossing cracks were established, and brightness degree profiles were extracted. Then, manual measurements of the crack widths were obtained by an operator.</p> <p>The dataset comprises 19,098 records of brightness profiles, reference crack width measurements, and benchmark measurements by deep learning and analytic models. The source images, which were stacked and marked with grid lines, are provided. The considerable number of brightness profiles coupled with manual reference measurements make the dataset well suited for developing an image-based deep learning models or analytic algorithms for assessing crack widths in concrete.</p> <p>The deposited data includes:</p> <ul> <li>krkCMd_table.csv: delimited, comma-separated text file containing a dataset of 19,098 crack brightness degree profiles, reference crack width measurements by operator, and benchmark measurements by a deep CNN metasensor and by an analytic edge detector.</li> <li>krkCMd_images.zip: archive containing source image files in folders by test series:&nbsp;<br>-&nbsp;&nbsp; stacked images of cracks in subsequent stages of self-healing (.tif files),<br>-&nbsp; &nbsp;zip archives assigned to image stacks and containing sets of ImageJ data files .roi,<br>- &nbsp; ImageJ .roi files specifying the locations of grid lines in the images.</li> <li>krkCMd_scripts.zip: archive containing custom scripts supporting image preprocessing and computing benchmark variables.</li> </ul> <p><span>For details please see the <a href="https://doi.org/10.1038/s41597-025-04485-z">data descriptor [1]</a>. When referring to the data in publications please cite [1].</span></p> <p>[1] Jakubowski, J., Tomczak, K. Dataset for developing deep learning models to assess crack width and self-healing progress in concrete.&nbsp;<em>Sci Data</em>&nbsp;<strong>12</strong>, 165 (2025). https://doi.org/10.1038/s41597-025-04485-z</p> <p>[2] Jakubowski, J. &amp; Tomczak, K. Deep learning metasensor for crack-width assessment and self-healing evaluation in concrete.&nbsp;<em>Constr. Build. Mater.</em>&nbsp;<strong>422</strong>, 135768 (2024). https://doi.org/10.1016/j.conbuildmat.2024.135768</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Deep learning models predicting gene functions and pathways using public DRKG knowledge graph and graph neural network

<p>The attached dataset contains pretrained link prediction models, as described in our paper 'Morphological Map of Under- and Over-Expression of Genes in Human Cells'.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Data sets and models for the Deep API Learning Revisited paper

<p>Training and test data for the machine learning experiments described in the paper Deep API Learning Revisited paper.&nbsp; Trained models are also included.</p> <p>Deep API Learning Revisited paper:&nbsp;&nbsp;https://doi.org/10.1145/3524610.3527872</p> <p>GitHub repository:&nbsp;&nbsp;https://github.com/hapsby/deepAPIRevisited</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Training Deep Learning Models to Estimate Permeability using Geophysical Datasets

<p>This folder contains the dataset for training deep learning models to estimate permeability using hydro-geophysics simulations</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Explainable AI for unveiling deep learning pollen classification model - Pollen dataset

<p>Dataset consists automatic particle detector Rapid-E measurements&nbsp;of&nbsp;pollen grains from 12 classes: Acer, Alnus, Alopecurus, Carex, Cupressus, Dactylis, Juglans,&nbsp;Morus,&nbsp;Platanus,&nbsp;Populus, Salix and Ulmus. Data are available i json format.</p> <p>Dataset also contains preprocessed data packed into csv files &nbsp;of 3 modalities: spectrum, lifetime, scattering and&nbsp;additional features from scattering and lifetime data are also available. These are ready to be used with machine learning models. Labels 0, 1, 2, ... 11&nbsp;correspond&nbsp;to alphabetical order of examined pollen classes Acer, Alnus, Alopecurus ...&nbsp;Ulmus.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Models and Datasets for "Extracting Paleoweather from Paleoclimate: A Deep Learning Reconstruction of Northern Hemisphere Summertime Atmospheric Blocking over the Last Millennium"

<p><strong>Associated publication:</strong> <em>Karamperidou, C., Extracting Paleoweather from Paleoclimate: A Deep Learning Reconstruction of Northern Hemisphere Summertime Atmospheric Blocking over the Last Millennium, Nature Communications Earth &amp; Environment, (2024)</em></p> <p>&nbsp;</p> <p><strong>This repository contains:</strong></p> <ul> <li>the architecture and weights of&nbsp;PaleoBlockNet v1.0</li> <li>the following ensemble DL reconstructions of JJA frequency of blocked days inferred by PaleoBlockNet: <ol> <li>the 10-member NTREND-based DL reconstruction; uses as input the NTREND DA N.Hemisphere MJJA surface temperature anomaly by King et al. (2021)</li> <li>the 100-member PHYDA-based DL reconstruction; uses as input the PHYDA JJA surface temperature anomaly by Steiger et al. (2018)</li> <li>the 12-member LME-based DL reconstruction; uses as input the CESM-LME surface temperature anomaly; this is a sensitivity experiment (see publication for details).</li> </ol> </li> <li>Integrated Gradients that assign importance to the input features for PaleoblockNet's blocking inferences&nbsp;</li> <li>train-validate-test samples to use with sample scripts from the Gituhub repo github/ckaramp-research/paleoblocknet</li> </ul> <p>&nbsp;</p> <p><strong>If you use this dataset, please cite the associated publication and the present repository.</strong></p> <p>To&nbsp;<strong>interactively explore</strong> the datasets, a web interface has been developed and can be accessed at <a href="https://www2.hawaii.edu/~ckaramp/paleoblocknet">https://www2.hawaii.edu/~ckaramp/paleoblocknet</a></p> <p>Contact the author Christina Karamperidou (<a title="Karamperidou Research Group" href="https://www2.hawaii.edu/~ckaramp" target="_blank" rel="noopener">https://www2.hawaii.edu/~ckaramp</a>) for more information about the details of these datasets.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

DeepPredSpeech: computational models of predictive speech coding based on deep learning

<p>This dataset contains all data, source code, pre-trained computational predictive&nbsp;models and experimental&nbsp;results related to:&nbsp;&nbsp;</p> <p>Hueber&nbsp;T., Tatulli E., Girin L., Schwatz, J-L&nbsp;&quot;How predictive can be predictions in the neurocognitive processing of auditory and audiovisual speech? A deep learning study.&quot; (<a href="https://doi.org/10.1101/471581">biorXiv preprint&nbsp;https://doi.org/10.1101/471581</a>).&nbsp;</p> <ul> <li>Raw data are extracted from the publicly available database NTCD-TIMIT (10.5281/zenodo.260228).&nbsp; <ul> <li>Audio recordings are available in the audio_clean/ directory</li> <li>Post-processed lip image sequences are available in the lips_roi/ directory (67x67 pixels, 8bits, obtained by lossless inverse DCT-2D transform from the DCT feature available in the original repository of NTCD-TIMIT)</li> <li>Phonetic segmentation (extracted from NTCD-TIMIT original zenodo repository) is available in the HTK MLF file volunteer_labelfiles.mlf</li> </ul> </li> <li>Audio features (MFCC-spectrogram and log-spectrogram) are available in the mfcc_16k/ and fft_16k/ directories.&nbsp;</li> <li>Models (audio-only, video-only and audiovisual, based on deep feed-forward neural networks and/or convolutional neural network, in&nbsp;.h5 format, trained with Keras 2.0 toolkit) and data normalization parameters (in .dat scikit-learn&nbsp;format)&nbsp;are available in models_mfcc/ and models_logspectro/ directories</li> <li>Predicted and target (ground truth) MFCC-spectro (resp.&nbsp;log-spectro) for the test databases (1909 sentences), and for the different values of <span class="math-tex">\(\tau_p\)</span>&nbsp;or&nbsp;<span class="math-tex">\(\tau_f\)</span> are available in pred_testdb_mfccspectro/ (resp.&nbsp;pred_testdb_logspectro/) directory</li> </ul> <p>Source code for extracting audio features, training and evaluating the models is available on GitHub&nbsp;https://github.com/thueber/DeepPredSpeech/</p> <p>All directories have been zipped before upload.</p> <p>Feel free to contact me for more details.</p> <p>Thomas Hueber, Ph. D., CNRS research fellow, GIPSA-lab, Grenoble, France,&nbsp;thomas.hueber@gipsa-lab.fr&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

Data supporting 'Ice loss in the European Alps until 2050 using a fully assimilated, deep-learning-aided 3D ice-flow model'

<p>The dataset supporting our publication '<strong>Ice loss in the European Alps until 2050 using a fully assimilated, deep-learning-aided 3D ice-flow model</strong>'&nbsp;in&nbsp;<em>Geophysical Research Letters.</em></p> <p>The main .zip archive contains a set of NetCDF files detailing:</p> <ul> <li>Initial optimised glacier states (geology-optimized...)</li> <li>Simulation results (Prog20...)</li> </ul> <p>Initial states and results are given by cluster (see Figure 1 in the paper), as shown in all filenames (C1 through to C12). Prognostic simulation filenames additionally distinguish between runs between 1999 and 2019 (Prog2020) and between 2020 and 2050 (Prog2050). 'NV'/'NoVel' and 'NT'/'NoThk' refer to simulations using the partial optimisation (optimisation without including velocity/thickness observations) as detailed in the paper. 'AV' at the end of the filename denotes the integrated area/volume results file, as opposed to the 2D raster results file. A 'V' before the cluster designation shows that the simulation used the variable SMB as opposed to the fixed SMB (see the paper for details). 'ID' before the cluster designation shows that the simulation was using extrapolated SMB based on the trend in SMB since 2000, instead of assuming the continuation of the current SMB. 'ID' on its own denotes linear extrapolation and 'IDQ' denotes quadratic extrapolation (not used in the published paper). 'SMBF' in the filename shows that the simulation used the SMB-elevation feedback.</p> <p>The additional .zip archive contains the code of IGM v1.0 used to produce the model results. For details on installing and using IGM, please see the Github page at&nbsp;<a href="https://github.com/jouvetg/igm.The">https://github.com/jouvetg/igm</a>.</p> <p>A further .zip archive (in version 3 - Sims2010-2022.zip) contains the simulations based on linear extrapolation of the observed trend in SMB between 2010 and 2022, following the same nomenclature as in the principal archive (see above).</p> <p>Version 4 contains an additional mosaicked DEM of the results for the whole Alps with the ice removed to give the complete basal topography (kindly processed by T. L&eacute;ger at UNIL) using the Japan Aerospace Exploration Agency (2021) ALOS World 3D 30 meter DEM. V3.2, Jan 2021. Distributed by OpenTopography. <a title="https://doi.org/10.5069/G94M92HB" href="https://doi.org/10.5069/G94M92HB" target="_blank" rel="noreferrer noopener">https://doi.org/10.5069/G94M92HB</a>. Accessed: 2024-09-09.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

DL-RMD: A geophysically constrained electromagnetic resistivity model database for deep learning applications (Dataset)

<p>Deep learning algorithms have shown incredible potential in many applications. The success of these data-hungry methods is largely associated with the availability of large-scale data sets, as millions of observations are often required to achieve acceptable performance levels. Recently, there has been an increased interest in applying deep learning methods to geophysical applications where electromagnetic methods are used to map the subsurface geology by observing variations in the electrical resistivity of the subsurface materials. To date, there are no standardized datasets for electromagnetic methods, which hinders the progress, evaluation, benchmarking, and evolution of deep learning algorithms due to data inconsistency. Therefore, we present a large-scale electrical resistivity model database of a wide variety of geologically plausible and geophysically resolvable subsurface structures for the commonly deployed ground-based and airborne electromagnetic systems. The presented database can potentially be used to build surrogate models of well-known processes and aid in labour intensive tasks. The geophysically constrained property of this database will not only achieve enhanced performance and improved generalization but, more importantly, it will incorporate consistency and credibility in deep learning models. We urge the geophysical community interested in deep learning for electromagnetic methods to utilize the presented database.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Detecting coarse beach sediment using remotely sensed imagery at the FRF, Duck, NC, USA: Labeled images, deep learning model, testing data, and predictions.

<p>This data record contains 5 zip files all used to build and use a semantic segmentation model to operate on beach imagery taken at the Field Research Facility (FRF) in Duck, North Carolina, USA. &nbsp;All data is from 2015-2021</p> <p>The `training_data.zip` contains all data used to train the ML model. All images come from the north facing (c1) camera. This zip file includes: a list of classes used to label the imagery, and folders of 107 images, 107 sparse annotations (doodles), 107 labels, and 107 overlays. All labeling was done with the open-source labeling tool &lsquo;Doodler (Buscombe et al., 2021).</p> <p>The `model.zip` file contains the ML model, and associated metadata. This includes: a JSON model configuration file, a figure showing model training statistics, an `.npz` file of model training output, a list of training and validation files, the model as an h5 file and in the Tensorflow &lsquo;saved model&rsquo; format. &nbsp;All modeling was done with Segmentation Gym (Buscombe &amp; Goldstein 2022).</p> <p>The `test_data_c6.zip` file contains all data from the south facing (c6) camera to test the ML model. This includes: a list of classes used to label the imagery, and folders of 10 images, 10 sparse annotations (doodles), 10 labels, and 10 overlays. &nbsp;All labeling was done with the open-source labeling tool &lsquo;Doodler (Buscombe et al., 2021). Testing the model with this data was done with codes in: https://github.com/ebgoldstein/FRF_GrainSize</p> <p>The `test_data_c1.zip` file contains all data from the north facing (c1) camera to test the ML model. This includes: a list of classes used to label the imagery, and folders of 10 images, 10 sparse annotations (doodles), 10 labels, and 10 overlays. &nbsp;All labeling was done with an open-source labeling tool &lsquo;Doodler (Buscombe et al., 2021). Testing the model with this data was done with codes in: https://github.com/ebgoldstein/FRF_GrainSize</p> <p>The `predictions.zip` file contains 4418 images from the north facing (c1) camera that were run through the trained segmentation model as well as the resulting output (presented as side-by-side image and overlays). These images were created using codes in Segmentation Gym (Buscombe &amp; Goldstein 2022).</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Phononic crystals dataset for supervised training of surrogate deep learning model

<p>The dataset contains shapes of unit cells of phononic crystals (inputs) in the form of images and corresponding dispersion diagrams (outputs). The dataset is used for deep learning (DL) model training.<br> Outputs are in the form of .mat files which contain vectors of reduced wavevector and corresponding frequencies, and also displacements u, v, w which can be used for polarization calculation.</p> <p>The dataset contains 11000 cases.</p> <p>Note: Ignore names &quot;labels&quot; as these are actually inputs to the DL model, not labels.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

A Novel Approach to Heart Failure Prediction and Classification through Advanced Deep Learning Model

<p>A Novel Approach to Heart Failure Prediction and Classification through Advanced Deep Learning Model</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Modeling islet enhancers using deep learning identifies candidate causal variants at loci associated with T2D and glycemic traits

<p>Genetic association studies have identified hundreds of independent genetic signals associated with type 2 diabetes (T2D) and related traits. Despite these successes, the identification of specific causal variants underlying a genetic association signal remains challenging. In this study, we describe a deep learning method to analyze the impact of sequence variants on enhancers. Focusing on pancreatic islets, a relevant T2D tissue, we show that our model learns islet-specific transcription factor (TF) regulatory patterns and can be used to prioritize candidate causal variants. At 101 genetic signals associated with T2D and related glycemic traits where multiple variants occur in linkage disequilibrium, our method nominates a single causal variant for each association signal, including three variants previously shown to alter reporter activity in islet-relevant cell types. For another signal associated with blood glucose levels, we biochemically test all candidate causal variants from statistical fine-mapping using a pancreatic islet beta cell line and show biochemical evidence of allelic effects on TF binding for the model-prioritized variant. To aid in future research, we publicly distribute our model and islet enhancer perturbation scores across ~67 million variants. We anticipate that deep learning methods like the one presented in this study will enhance the prioritization of candidate causal variants for functional studies.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

LocScale-EMmerNet deep learning models for contrast optimisation of cryo-EM maps

<p>EMmerNet deep learning models for local optimisation of cryo-EM map contrast using <a href="https://gitlab.tudelft.nl/aj-lab/locscale">LocScale</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Small molecules targeting the structural dynamics of AR-V7 partially disordered protein using deep learning and physics based models.

<p>Partially disordered proteins can contain both stable and unstable secondary structure segments and &nbsp;are involved in various (mis)functions in the cell. The extensive conformational dynamics of partially disordered proteins scaling with extent of disorder and length of the protein hampers the efficiency of traditional experimental and in-silico structure-based drug discovery approaches. Therefore new efficient paradigms in drug discovery taking into account conformational ensembles of proteins need to emerge. In this study, using as a test case the AR-V7 transcription factor splicing variant related to prostate cancer, we present an automated &nbsp;methodology that can accelerate the screening of small molecule binders targeting partially disordered proteins. By swiftly identifying the conformational ensemble of AR-V7, and reducing the dimension of binding-sites by a factor of 90 by applying appropriate physicochemical filters, &nbsp;we combine physics based molecular docking and multi-objective classification machine learning models that speed up the screening of thousands of compounds targeting AR-V7 multiple binding sites. Our method not only identifies previously known binding sites of AR-V7, but also discovers new ones, as well as increases the multi-binding site hit-rate of small molecules by a factor of 17 compared to naive physics-based molecular docking.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Deep-Learning-Based Harmonization and Super-Resolution of Near-Surface Air Temperature from CMIP6 Models (1850-2100)

<p>A long-term (1850-2100) monthly air temperature (tas) product with a spatial resolution of 0.5 degree. This is a merged product from 31&nbsp;CMIP6&nbsp;models&nbsp;using the Deep-learning model&nbsp;which reduce bias, spatial downscaling and data merge at the same time,. To facilitate user-friendly access and download the dataset is stored individually for each year in a separate file. These files contain one historical data (1850-2014) , four future scenarios data&nbsp;during 2015-2100 (SSP1-2.6, SSP2-4.5, SSP3-7.0, SSP5-8.5) and four&nbsp;future scenarios data in Australia&nbsp; . The dataset is stored in NetCDF format, containing the variable tas, representing air temperature, produced in&nbsp; centigrade (℃) as a unit. There are three dimensions included in the dataset: longitude, latitude, and time, with the longitude ranging from -179.75E to 179.75E, the latitude from -89.75N to 89.75N.&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

An Image Dataset for Training Deep Learning Segmentation Models to Identify Karst Sinkholes

<p>The image dataset was prepared for training deep learning image segmentation models to identify karst sinkholes. Information about the work can be found at (https://github.com/mvrl/sink-seg/). The dataset consists of a DEM image, an aerial image, and a binary sinkhole label image in an area in central Kentucky, USA.&nbsp; It also includes four images derived from the DEM image.&nbsp; The image dataset is sourced from publicly available&nbsp; data from Kentucky&#39;s Elevation Data &amp; Aerial Photography Program (https://kyfromabove.ky.gov/) and Kentucky LiDAR-derived sinkholes (https://kgs.uky.edu/geomap).</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record