Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo24/100

Finite element data collected and Machine learning algorithms to predict the mechanical properties of innovative CLT

<p>This folder includes the data collected from the finite element simulations of the innovative CLT to compute its mechanical properties, the error of the closed-form solutions predicting the bending stiffness in the minor direction D22, the variation of the distance between the Reissner Mindlin and Bending Gradient theory in terms of spacing between lateral lamellas, the hyperparameters tuning of several ML algorithms (Regression Tree, Random Forest, Gradient Boosting and Artificial Neural Network), the ML evaluations, the saved artificial neural network algorithms to predict each mechanical property of innovative CLT, and the ML application to use it.</p>

opencc-by-4.0Apr 2024View details →
zenodo24/100

Machine learning-aided design of composite mycotoxin detoxifier material for animal feed

<p>Model codes and dataset associated to &quot;Machine learning-aided design of composite mycotoxin detoxifier material for animal feed, Giulia Lo Dico, Siska Croubels,&nbsp;Ver&oacute;nica Carcel&eacute;n,&nbsp;and Maciej Haranczyk,&nbsp;Sci. Rep., 2022&quot;.</p>

opencc-by-4.0Dec 2021View details →
zenodo24/100

Towards on-device dehydration monitoring using machine learning from wearable device's data

<p>Body signals data&nbsp; such as accelerometer, gyroscope, magnetometer, galvanic skin response, photoplethysmography, temperature and barometric pressure were recorded for 11 subjects using Shimmer3 GSR unit. Some data were recorded while fasting for periods reaching 15 hours without water or food intake in Muslims&#39; fasting time (starting fasting time for each subject sample is recorded in fastingTimes.csv for each subject from which last drinking time can be calculated and for the other non-fasting samples, drinking was done in the last hour before recording the sample, so the last drinking time was set to be 1 hour.</p> <p>I have received from some researchers that the zip folder is corrupted, you can find the right version here:&nbsp;https://zenodo.org/record/7099266.</p> <p>Please kindly contact the authors if you have any further concerns</p>

opencc-by-4.0Feb 2022View details →
zenodo24/100

Machine Learning class final project datasets

<p>This is the zip file for the machine learning course final project. it has training and testing data sets.</p>

opencc-by-4.0May 2022View details →
zenodo24/100

Provably efficient machine learning for quantum many-body problems (old version)

<p>Raw data for the manuscript &quot;Provably efficient machine learning for quantum many-body problems&quot;.</p>

opencc-by-4.0May 2022View details →
zenodo24/100

Multi-Omic Integration by Machine Learning (MIMaL) Reveals Protein-Metabolite Connections and New Gene Functions

<p>Metabolomics and proteomics generate large, complex datasets that reflect the state of a biological system. Multi-omics is the integration of these disparate methods and data to gain a clearer picture of the biological state. Multi-omic studies of the proteome and metabolome are becoming more common as mass spectrometry technology continues to be democratized. However, knowledge extraction through integration of these data remains challenging. Here we show that connections between these omic layers can be discovered through a combination of machine learning and model interpretation. We find that SHAP values connecting proteins to metabolites are valid experimentally, and reveal also largely new connections. Further, clustering the magnitudes of protein control over all metabolites enabled prediction of gene five gene functions, each of which was validated experimentally. We accurately predicted that two uncharacterized genes in yeast modulate mitochondrial translation, <em>YJR120W</em> and <em>YLD157C</em>.We also predict and validate functions for several incompletely characterized genes, including <em>SDH9</em>, <em>ISC1</em>, and <em>FMP52</em>. Our work demonstrates that multi-omic analysis with machine learning (MIMaL) is a new lens that reveals new insight from multi-omic data that would not be possible using any omic layer alone.</p>

opencc-by-4.0May 2022View details →
zenodo24/100

Large-scale comparison of machine learning algorithms for target prediction of natural products

<p>Supplement Materials of the article named &quot;Large-scale comparison of machine learning algorithms for target prediction of natural products&quot;.</p>

opencc-by-4.0Jan 2022View details →
zenodo24/100

devCellPy is a machine learning-enabled pipeline for automated annotation of complex multilayered single-cell transcriptomic data

<p>A major informatic challenge in single cell RNA-sequencing analysis is the precise annotation of datasets where cells exhibit complex multilayered identities or transitory states. Here, we present&nbsp;<em>devCellPy</em>&nbsp;a highly accurate and precise machine learning-enabled tool that enables automated prediction of cell types across complex annotation hierarchies. To demonstrate the power of&nbsp;<em>devCellPy</em>, we construct a murine cardiac developmental atlas from published datasets encompassing 104,199 cells from E6.5-E16.5 and train&nbsp;<em>devCellPy</em>&nbsp;to generate a cardiac prediction algorithm. Using this algorithm, we observe a high prediction accuracy (&gt;90%) across multiple layers of annotation and across de novo murine developmental data. Furthermore, we conduct a cross-species prediction of cardiomyocyte subtypes from in vitro<em>-</em>derived human induced pluripotent stem cells and unexpectedly uncover a predominance of left ventricular (LV) identity that we confirmed by an LV-specific TBX5 lineage tracing system. Together, our results show devCellPy to be a useful tool for automated cell prediction across complex cellular hierarchies, species, and experimental systems.</p>

opencc-by-4.0Sep 2022View details →
zenodo24/100

Naming the Pain in Machine Learning-Enabled Systems Engineering

<p>This repository contains analyses and data related to the paper 'Naming the Pain in Machine Learning-Enabled Systems Engineering.' It includes Jupyter Notebooks for analysis, raw survey and qualitative data, relevant images, and the original survey tool.</p>

opencc-by-4.0May 2024View details →
zenodo24/100

Machine learning-based pulse wave analysis for classification of circle of Willis topology: an in silico study with 30,618 virtual subjects (database: Missing PCA P1)

<p>This repository contains the dataset for the Missing PCA P1 described in the article with the same name. MATLAB and Python codes for post-processing the dataset and the code for training and testing all machine learning models using the open-source library TensorFlow 2.12, the Keras application programming interface, and the Scikit-learn Python package can be found in here (<a href="https://zenodo.org/records/12519322" target="_blank" rel="noopener">https://zenodo.org/records/12519322</a>).</p>

opencc-by-4.0Jun 2024View details →
zenodo24/100

Skillful bias correction of offshore near-surface wind speed and wind direction forecasting based on a multi-task machine learning model

<h3>Dataset</h3> <p>1. observation data over 14 weather stations</p> <p>Variables: hourly near-surface 2-min average wind speed, wind direction&nbsp;</p> <p>2. ECMWF-IFS forecast data over 14 weather stations</p> <p>Variables: hourly predictors at surface level and upper level in next 48 hours (shown in Table 1. and Table 2.)</p> <p>Table 1. ECMWF-IFS forecast data at surface level</p> <div> <table> <tbody> <tr> <td> <p>Predictors</p> </td> <td> <p>Abbreviation</p> </td> <td> <p>Unit</p> </td> </tr> <tr> <td> <p>Temperature at 2 m</p> </td> <td> <p>2t</p> </td> <td> <p>℃</p> </td> </tr> <tr> <td> <p>Sea surface temperature</p> </td> <td> <p>sst</p> </td> <td> <p>℃</p> </td> </tr> <tr> <td> <p>Dewpoint temperature at 2 m</p> </td> <td> <p>2d</p> </td> <td> <p>℃</p> </td> </tr> <tr> <td> <p>Convective&nbsp;precipitation in the past hour</p> </td> <td> <p>cp</p> </td> <td> <p>mm</p> </td> </tr> <tr> <td> <p>Mean sea level pressure</p> </td> <td> <p>msl</p> </td> <td> <p>hPa</p> </td> </tr> <tr> <td> <p>Zonal component of wind speed at 10 m</p> </td> <td> <p>10u</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Meridional component of wind speed at 10 m</p> </td> <td> <p>10v</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind speed at 10 m</p> </td> <td> <p>10ws</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind direction&nbsp;at 10 m</p> </td> <td> <p>10wd</p> </td> <td> <p>&deg;</p> </td> </tr> <tr> <td> <p>Zonal component of wind speed at 100 m</p> </td> <td> <p>100u</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Meridional component of wind speed at 100 m</p> </td> <td> <p>100v</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind speed at 100 m</p> </td> <td> <p>100ws</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind direction&nbsp;at 100 m</p> </td> <td> <p>100wd</p> </td> <td> <p>&deg;</p> </td> </tr> </tbody> </table> </div> <div>&nbsp;</div> <p>Table 2. ECMWF-IFS forecast data at upper level</p> <table> <tbody> <tr> <td> <p>Predictors</p> </td> <td> <p>Abbreviation</p> </td> <td> <p>Unit</p> </td> </tr> <tr> <td> <p>Relative humidity at xxx hPa</p> </td> <td> <p>r_Lxxx</p> </td> <td> <p>%</p> </td> </tr> <tr> <td> <p>Temperature at xxx hPa</p> </td> <td> <p>t_Lxxx</p> </td> <td> <p>℃</p> </td> </tr> <tr> <td> <p>Vertical velocity&nbsp;of wind at xxx hPa</p> </td> <td> <p>w_Lxxx</p> </td> <td> <p>Pa s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Zonal component of wind at xxx hPa</p> </td> <td> <p>u_Lxxx</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Meridional component of wind&nbsp;at xxx hPa</p> </td> <td> <p>v_Lxxx</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind speed&nbsp;at xxx hPa</p> </td> <td> <p>ws_Lxxx</p> </td> <td> <p>m s<sup>-1</sup></p> </td> </tr> <tr> <td> <p>Wind direction at xxx hPa</p> </td> <td> <p>wd_Lxxx</p> </td> <td> <p>&deg;</p> </td> </tr> </tbody> </table> <div>&nbsp;</div> <p>3. key variables constructed by feature engineering</p> <p>(1) sort-term statistics, including <em>maximum, minimum, mean </em>and <em>variance</em>&nbsp;of key variables (<em>2t</em>,<em>&nbsp;10u</em>, <em>10v </em>and <em>10ws</em>) from ECMWF-IFS model&nbsp;during the next&nbsp;48 hours,</p> <p>&nbsp;(2) long-term statistics, including <em>mean </em>and <em>deviation</em>&nbsp;of key variables (<em>2t</em>,<em>&nbsp;10u</em>, <em>10v </em>and <em>10ws</em>)&nbsp;from ECMWF-IFS model&nbsp;during&nbsp;history&nbsp;3-yr&nbsp;period (January 2020&ndash;December&nbsp;2022),</p> <p>&nbsp;(3) thermodynamic factors, &nbsp;including the low-level wind shear&nbsp;between <em>10ws</em>&nbsp;and <em>100ws</em>,&nbsp;vertical wind shear between 200 hPa and 850 hPa<em>, </em>the differences between <em>sst</em><em>&nbsp;</em>and&nbsp;<em>2t</em><em>.</em></p> <h3>Scripts</h3> <p>1. Random Forest model training code</p> <p>2. LightGBM model training code</p> <p>3. XGBoost model training code</p> <p>4. TabNet-MTL model training code</p> <p>&nbsp;</p>

embargoedcc-by-sa-4.0Apr 2024View details →
zenodo24/100

Mechanistic Exploration and Kinetic Modeling through In-Silico Data Generation and Probabilistic Machine Learning Analysis

<p>This zip file includes the dataset 'two_reactions_022624.csv,' which is used for training and testing ML/DL models in the paper 'Mechanistic Exploration and Kinetic Modeling through In-Silico Data Generation and Probabilistic Machine Learning Analysis,' as well as trained models and some files used for training the model. When running the model downloaded from GitHub, copy and paste the files downloaded from here into the subfolder with the same name and path as the one downloaded from GitHub.</p>

openJul 2024View details →
zenodo24/100

Data and code for training and testing a ResMLP model with experience replay for machine-learning physics parameterization

<p>This directory contains the training data and code for training and testing a ResMLP with experience replay for creating a machine-learning physics parameterization for the Community Atmospheric Model.&nbsp;</p> <p>The directory is structured as follows:</p> <p>1. Download training and testing data: https://portal.nersc.gov/archive/home/z/zhangtao/www/hybird_GCM_ML</p> <p>2. Unzip nncam_training.zip</p> <p>nncam_training</p> <p>&nbsp; &nbsp; - models</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;model definition of ResMLP and other models for comparison purposes</p> <p>&nbsp; &nbsp; - dataloader&nbsp;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;utility scripts to load data into pytorch dataset</p> <p>&nbsp; &nbsp; - training_scripts</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;scripts to train ResMLP model with/without experience replay</p> <p>&nbsp; &nbsp; - offline_test</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;scripts to perform offline test (Table 2, Figure 2)</p> <p>3. Unzip nncam_coupling.zip</p> <p>nncam_srcmods</p> <p>&nbsp; &nbsp; &nbsp;- SourceMods</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; SourceMods to be used with CAM modules for coupling with neural network</p> <p>&nbsp; &nbsp; &nbsp;- otherfiles</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; additional configuration files to setup and run SPCAM with neural network</p> <p>&nbsp; &nbsp; &nbsp;- pythonfiles</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;python scripts to run neural network and couple with CAM</p> <p>&nbsp; &nbsp; &nbsp; - ClimAnalysis</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; - paper_plots.ipynb</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;scripts to produce online evaluation figures (Figure 1, Figure 3-10)</p> <p>&nbsp; &nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo24/100

Supplementary material for the preprint "GLOSSA: a user-friendly R Shiny application for Bayesian machine learning analysis of marine species distribution"

<p>In this repository we present the code and data for the case studies in "GLOSSA: a user-friendly R Shiny application for Bayesian machine learning analysis of marine species distribution". The GLOSSA website can be accessed at https://jmestret.github.io/glossa/. Occurrence data for <em>Thunnus albacares</em> was obtained from the OBIS database (https://obis.org/taxon/127027), for&nbsp;<em>Caretta caretta</em> from GBIF (https://doi.org/10.15468/dl.es7562), and for <em>Siganus luridus </em>from the GreekMarineICAS geodataset (https://doi.org/10.25607/t2smha), created as part of the ALAS (Aliens in the Aegean &ndash; A Sea Under Siege) project.</p>

openmit-licenseSep 2024View details →
zenodo24/100

A Dataset for Applying Machine Learning and Eddy Covariance Approaches to Model Mangrove Carbon Production (ML-MCP)

<p>The Mangrove Carbon Production (ML-MCP) dataset (daily time scale) encompasses comprehensive measurements of carbon production in mangrove ecosystems from four EC tower station in the USA and China, derived using advanced machine learning models and eddy covariance techniques. This dataset includes various variables such as carbon fluxes, environmental factors. By integrating machine learning algorithms, the dataset enhances the accuracy of carbon productivity estimations, facilitating better understanding and management of mangrove ecosystems' role in carbon sequestration and climate regulation.</p>

opencc-by-4.0Sep 2024View details →
zenodo24/100

A large synthetic dataset for machine learning applications in power transmission grids

<p>With the ongoing energy transition, power grids are evolving fast. They operate more and more often close to their technical limit, under more and more volatile conditions. Fast, essentially real-time computational approaches to evaluate their operational safety, stability and reliability are therefore highly desirable. Machine Learning methods have been advocated to solve this challenge, however they are heavy consumers of training and testing data, while historical operational data for real-world power grids are hard if not impossible to access.&nbsp;</p> <p>This dataset contains long time series for production, consumption, and line flows, amounting to 20 years of data with a time resolution of one hour, for several thousands of loads and several hundreds of generators of various types representing the ultra-high-voltage transmission grid of continental Europe. The synthetic time series have been statistically validated agains real-world data.</p> <h2>Data generation algorithm</h2> <p>The algorithm is described in a <a href="https://doi.org/10.1038/s41597-025-04479-x">Nature Scientific Data paper</a>. It relies on <a href="https://zenodo.org/records/2642175" target="_blank" rel="noopener">the PanTaGruEl model of the European transmission network</a> -- the admittance of its lines as well as the location, type and capacity of its power generators -- and aggregated data gathered from <a href="https://transparency.entsoe.eu/" target="_blank" rel="noopener">the ENTSO-E transparency platform</a>, such as power consumption aggregated at the national level.</p> <h2>Network</h2> <p>The network information is encoded in the file <a href="https://zenodo.org/records/13378476/files/europe_network.json">europe_network.json</a>. It is given in <a href="https://lanl-ansi.github.io/PowerModels.jl/stable/" target="_blank" rel="noopener">PowerModels format</a>, which it itself derived from <a href="https://matpower.org/" target="_blank" rel="noopener">MatPower</a> and compatible with <a href="https://www.pandapower.org/" target="_blank" rel="noopener">PandaPower</a>. The network features 7822 power lines and 553 transformers connecting 4097 buses, to which are attached 815 generators of various types.</p> <h2>Time series</h2> <p>The time series forming the core of this dataset are given in CSV format. Each CSV file is a table with 8736 rows, one for each hourly time step of a 364-day year. All years are truncated to exactly 52 weeks of 7 days, and start on a Monday (the load profiles are typically different during weekdays and weekends). The number of columns depends on the type of table: there are 4097 columns in load files, 815 for generators, and 8375 for lines (including transformers). Each column is described by a header corresponding to the element identifier in the network file. All values are given in per-unit, both in the model file and in the tables, i.e. they are multiples of a base unit taken to be 100 MW.</p> <p>There are 20 tables of each type, labeled with a reference year (2016 to 2020) and an index (1 to 4), zipped into archive files arranged by year. This amount to a total of 20 years of synthetic data.&nbsp; When using loads, generators, and lines profiles together, it is important to use the same label: for instance, the files <em>loads_2020_1.csv</em>, <em>gens_2020_1.csv</em>, and <em>lines_2020_1.csv</em> represent a same year of the dataset, whereas <em>gens_2020_2.csv</em> is unrelated (it actually shares some features, such as nuclear profiles, but it is based on a dispatch with distinct loads).</p> <h2>Usage</h2> <p>The time series can be used without a reference to the network file, simply using all or a selection of columns of the CSV files, depending on the needs. We show below how to select series from a particular country, or how to aggregate hourly time steps into days or weeks. These examples use Python and the data analyis library <em>pandas</em>, but other frameworks can be used as well (Matlab, Julia). Since all the yearly time series are periodic, it is always possible to define a coherent time window modulo the length of the series.</p> <h3>Selecting a particular country</h3> <p>This example illustrates how to select generation data for Switzerland in Python. This can be done without parsing the network file, but using instead <a href="https://zenodo.org/records/13378476/files/gens_by_country.csv">gens_by_country.csv</a>, which contains a list of all generators for any country in the network. We start by importing the <em>pandas</em> library, and read the column of the file corresponding to Switzerland (country code CH):</p> <pre><code>import pandas as pd CH_gens = pd.read_csv('gens_by_country.csv', usecols=['CH'], dtype=str)</code></pre> <p>The object created in this way is Dataframe with some null values (not all countries have the same number of generators). It can be turned into a list with:</p> <pre><code>CH_gens_list = CH_gens.dropna().squeeze().to_list()</code></pre> <p>Finally, we can import all the time series of Swiss generators from a given data table with</p> <pre><code>pd.read_csv('gens_2016_1.csv', usecols=CH_gens_list)</code></pre> <p>The same procedure can be applied to loads using the list contained in the file <a href="https://zenodo.org/records/13378476/files/loads_by_country.csv">loads_by_country.csv</a>.</p> <h3>Averaging over time</h3> <p>This second example shows how to change the time resolution of the series. Suppose that we are interested in all the loads from a given table, which are given by default with a one-hour resolution:</p> <pre><code>hourly_loads = pd.read_csv('loads_2018_3.csv')</code></pre> <p>To get a daily average of the loads, we can use:&nbsp;</p> <pre><code>daily_loads = hourly_loads.groupby([t // 24 for t in range(24 * 364)]).mean()</code></pre> <p>This results in series of length 364. To average further over entire weeks and get series of length 52, we use:&nbsp;</p> <pre><code>weekly_loads = hourly_loads.groupby([t // (24 * 7) for t in range(24 * 364)]).mean()</code></pre> <h2>Source code</h2> <p>The code used to generate the dataset is freely available at <a href="https://github.com/GeeeHesso/PowerData" target="_blank" rel="noopener">https://github.com/GeeeHesso/PowerData</a>. It consists in two packages and several documentation notebooks. The first package, written in Python, provides functions to handle the data and to generate synthetic series based on historical data. The second package, written in Julia, is used to perform the optimal power flow. The documentation in the form of Jupyter notebooks contains numerous examples on how to use both packages. The entire workflow used to create this dataset is also provided, starting from raw ENTSO-E data files and ending with the synthetic dataset given in the repository.</p> <h2>Funding</h2> <p>This work was supported by the <a href="https://www.cydcampus.admin.ch">Cyber-Defence Campus of armasuisse</a> and by an internal research grant of the Engineering and Architecture domain of <a href="https://www.hes-so.ch">HES-SO</a>.</p>

opencc-by-4.0Oct 2024View details →
zenodo24/100

Supporting dataset for "27Al NMR chemical shifts in zeolite MFI via machine learning acceleration of structure sampling and shift prediction"

<p>This dataset includes includes training databases of CHA, MOR and MFI zeolites, trained kernel ridge regression (KRR) models, and the initial structures utilized in the study. A more detailed description of the dataset can be found in the README file.</p> <p>Note, all MD simulations were performed using SiAlOH1 ML potential from the work of Erlebach et al. (Erlebach et al., Nat Commun 15, 4215 (2024)), available at: https://doi.org/10.5281/zenodo.10361794.</p>

opencc-by-4.0Sep 2024View details →
zenodo24/100

For MACHINE LEARNING DATABASE evaluation: old-version SEM images of TiO2 particles UNITO

Test images recorded with old ZEISS software, metadata version could differ to the up-to-date version.

opencc-by-nc-nd-4.0Dec 2017View details →
zenodo24/100

Source Data for the publication: Multi-omics and machine learning reveal context-specific gene regulatory activities of PML::RARA in Acute Promyelocytic Leukemia

<p>Source Data for the publication: Multi-omics and machine learning reveal context-specific gene regulatory activities of PML::RARA in Acute Promyelocytic Leukemia</p>

opencc-by-4.0Dec 2022View details →
zenodo24/100

On Learning Meaningful Code Changes via Neural Machine Translation

<p>Paper: On Learning Meaningful Code Changes&nbsp;via Neural Machine Translation</p> <p>Authors: Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk</p> <p>ICSE 2019 - 41st ACM/IEEE International Conference on Software Engineering, May &nbsp;25-31, 2019, Montr&eacute;al, QC, Canada</p>

opencc-by-4.0Dec 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record