Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,805

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5,805 results for “Data model”

Learn how ShareScore rates datasets ↗
zenodo32/100

Data for "Simulating and analysing seabird flyways: an approach combining least-cost path modelling and machine learning"

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

Data and models for "An image-computable model of speeded decision-making"

<p>Lost in Migration gameplay data and trained models for:</p> <div>Jaffe, P. I., Gustavo, X. S. R., Schafer, R. J., Bissett, P. G., Poldrack, R. A. An image-computable model of speeded decision-making. <em>eLife</em> <strong>13</strong>, RP98351 (2024).</div> <div>&nbsp;</div> <p>This dataset can be used to reproduce all of the results of the manuscript, following the instructions in the code repository for the paper: <a href="https://github.com/pauljaffe/vam">https://github.com/pauljaffe/vam</a>.</p> <p>The dataset includes the following components:</p> <p><strong>gameplay_data.zip:</strong> Trial-level gameplay metadata for Lost in Migration. Lost in Migration is a variant of the flanker task offered as a part of the Lumosity cognitive training platform (Lumos Labs, Inc.). The .zip file includes a separate .csv file for each of the 75 Lumosity users (participants) that we trained models on. Each .csv file has one row per trial with the following fields/columns: "anon_id", numerical identifier for the Lumosity user; "nth_play", the nth gameplay of Lost in Migration for this user; "trial", the nth trial for the current gameplay; "xpos", the signed horizontal distance from the center of the target bird to the left edge of the game window (pixels, non-negative); "ypos", the signed vertical distance from the center of the target bird to the bottom edge of the game window (pixels, non-negative); "flanker_direction", (L/R/U/D); "response_direction", (L/R/U/D); "target_direction", (L/R/U/D); "response_time", (ms); "stimulus_layout", numerical code for the layout of the bird flock for the current trial (0: horizontal line, 1: vertical line, 2: cross, 3: &lt;, 4: &gt;, 5: v, 6: ^)<strong>.</strong></p> <p><strong>vam_models.zip:</strong> Parameters for the 75 visual accumulator models (VAMs) analyzed in the manuscript.</p> <p><strong>task_opt_models.zip:</strong> Parameters for the 75 task-optimized models analyzed in the manuscript.</p> <p><strong>metadata.csv:</strong> Metadata for each Lumosity user that a VAM/task-optimized model was trained on. The .csv file has one row per user with the following fields/columns: "user_id", numerical identifier for the Lumosity user (same as "anon_id" in gameplay_data.zip); "gender", self-reported gender ('m', 'f', or null, indicating no response was given); "binned_age", age bucketed into decade-long bins (20-29, 30-39... 80-89).</p> <p><strong>derivatives.zip:</strong> The RTs/choices generated by the trained models, organized into separate folders by model type (vam/task_opt/binned_rt) and user ID.&nbsp; Also includes a "summary_stats" folder with analysis products of the model activations and outputs.</p> <p><strong>graphics.zip:</strong> Image files used to create the visual stimuli from the gameplay metadata.</p> <p><strong>example_model_inputs.zip: </strong>The processed visual stimuli and gameplay data used as inputs to train one model (user ID 182). Note we provide instructions to recreate the stimuli and other model inputs for all models in the code repository.</p>

opencc-zeroMar 2024View details →
zenodo32/100

User Modeling in MDE - Data Sheet and Scripts

<p>This replication package contains both the filled out excel template and the used scripts to filter the results at an abstract/title level for ScienceDirect and SpringerLink.&nbsp;<br><br>For ScienceDirect:&nbsp;</p> <ol> <li>Perform the search on the ScienceDirect page and download the results.</li> <li>ScienceDirect proposes a .bib file download. Transform this .bib file into a .ris file using any available online tools for such a conversion.</li> <li>Execute the script 'sciencedirect_filter_script.py' that will ingest a ScienceDirect.ris file and produce a file called 'output_filtered_included.ris' file.</li> </ol> <p>For SpringerLink:</p> <ol> <li>Perform the search on SpringerLink and download the resulting .csv file. and name it SearchResults.csv.</li> <li>The downloaded file does not contain the abstracts, that is why the 'springerlink_abstract_scrape.py' script that will perform a scraping to fetch the abstracts. Execute 'python springerlink_abstract_scrape.py SearchResults.py output.csv'. This will start the scraping and save the results in output.csv. Note that you will need to manually change the name of the columns ['Item Title', 'Authors', 'Publication Year', 'Publication Title', 'URL'] to ['title', 'author', 'year', 'publication', 'link']</li> <li>Perform a manual check if some abstracts are missing. Simply text search the .csv file for 'ABSTRACT NOT FOUND ERROR' and manually enter the missing results.</li> <li>Now use the second script by executing 'python springerlink_filter_script.py' and it will generate a file called 'outputfinalkeyword.ris'</li> </ol> <p>&nbsp;</p> <p>The data extraction excel file contains all the relevant information concerning the data extraction. This includes the reviewed papers, a summary of the numbers, the RQs and metadata.&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Data sharing of: Sulfur inventory of the young lunar mantle constrained by experimental sulfide saturation of Chang'e-5 mare basalts and a new sulfur solubility model for silicate melts in equilibrium with sulfides of variable metal–sulfur ratio

<p>Data sharing of: Sulfur inventory of the young lunar mantle constrained by experimental sulfide saturation of Chang&rsquo;e-5 mare basalts and a new sulfur solubility model for silicate melts in equilibrium with sulfides of variable metal&ndash;sulfur ratio</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Input data for the Community Water Model (CWatM) - a regional dataset covering Israel and the Ayalon Basin

<p>This data was used for the paper: 'Wastewater matters: Incorporating wastewater treatment and reuse into a process-based hydrological model (CWatM v1.08)'. When using this data pleas cite the paper, alongside this Zenodo repository.</p> <p>Fridman, D., Smilovic, M., Burek, P., Tramberend, S., and Kahil, T. 2024. Wastewater matters: Incorporating wastewater treatment and reuse into a process-based hydrological model (CWatM v1.08). <em>Geoscientific Model Development Discussions</em>, 1-26. [Preprint]</p> <p><strong><span>Background</span></strong></p> <ul> <li>This dataset was developed as part of the IIASA WINTER project, aiming to run high resolution hydrological<br>simulations in the river basins in Israel (Water Futures and Solutions for Israel (WFaS-Israel) | IIASA).</li> <li>Conducting a high-resolution (30 arcseconds) simulation can rely on a mix of upscaled coarse global, high-resolution global, and local datasets. Specifically, Hanasaki et al. (2022) stress the importance of local water management and use data for hyper-resolution hydrologic simulations.</li> <li>The dataset covers the terrestrial area of Israel and the Palestinian Authority. It also includes the cross-border and upstream river basin in the neighboring countries Egypt, Jordan, Syria, and Lebanon. The selected river basins of Ayalon and Sorek are in the central coastal area of Israel and vary by topography, land cover, and water management. The Ayalon stream drains the western downslopes of the Judea and Samaria mountains and outlets into the Yarkon stream (to the North), which later reaches the Mediterranean Sea. The Sorek stream drains the hills around South-West Jerusalem and flows Westwards until reaching the Mediterranean Sea.</li> </ul> <p><strong><span>Further Information</span></strong></p> <ul> <li>For details about the contents of this dataset please refer to the readme.txt</li> <li>For details about the data soruces and processing please refer to the dataset_description.pdf</li> </ul> <p>In case of additional questions please do not hesitate to write to us fridman@iiasa.ac.at</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Reproducibility Test Data of Foundation Model for Cancer Imaging x Mhub

<p>This dataset provides test data for the FMCIB model integrated within the MHub platform, a robust solution for deploying, managing, and testing deep learning models tailored for medical imaging.&nbsp;</p> <h3>Dataset Composition:</h3> <ul> <li><strong>Sample Folder:</strong>&nbsp;Contains the input data utilized for testing the model&rsquo;s functionality.</li> <li><strong>Reference Folder:</strong>&nbsp;Contains the corresponding output provided by the original model contributor.</li> <li><strong>Test.yml File:</strong>&nbsp;This file includes the original contributor&rsquo;s test setup, which has been accepted by the MHub team.</li> </ul> <h3>Sample Data Source:</h3> <p>The sample images used in this dataset are sourced from public datasets available through the&nbsp;<strong><a href="https://datacommons.cancer.gov/repository/imaging-data-commons" target="_blank" rel="noopener">Imaging Data Commons (IDC)</a></strong>, a repository that provides access to a wide range of medical imaging data. This ensures that the test cases reflect real-world clinical scenarios, facilitating robust validation of model performance.</p> <h3>Purpose and Utility:</h3> <p>The primary objective of this dataset is to enable the rigorous testing and validation of model performance within MHub workflows. To assess the performance of a model, users can process the sample data and compare the resulting output to the reference data. Additionally, users may inspect the sample and reference data independently to better understand the input-output structure that defines each model&rsquo;s workflow.</p> <p>This dataset streamlines the process of model validation. By providing a standardized testing framework, the dataset facilitates reproducible results and accelerates the development of reliable AI models for medical imaging.</p> <h3>About MHub:</h3> <p>MHub (<a href="https://mhub.ai/" target="_new" rel="noopener">mhub.ai</a>) is an innovative platform designed to simplify the deployment, management, and testing of deep learning models for medical imaging. It enables researchers and clinicians to integrate AI-based solutions into clinical workflows while ensuring reproducibility and scalability. The platform provides a modular framework where users can execute complex workflows, such as image segmentation, classification, and registration, leveraging state-of-the-art AI models. MHub's goal is to accelerate the development and clinical adoption of medical imaging models by providing a streamlined, user-friendly environment for testing and validating new algorithms.</p> <p>For more information on the platform and its capabilities, visit&nbsp;<a href="https://mhub.ai/" target="_blank" rel="noopener">mhub.ai</a>.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

model data

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

Data used in paper entitled "Symplecticity of the GORILLA guiding-center tracer and its implications for edge transport modeling"

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

Online water quality monitoring data from full scale CS#3 DWDN for the DBP prediction model

<p>Online water quality data though the drinking water distribution network. More than 1 year of data.</p> <p>SCADA data source.</p> <p>Provide water quality of the whole system at selected locations.</p>

restrictedcc-by-4.0Sep 2024View details →
zenodo32/100

Online hydraulic data from full scale CS#3 DWDN for model calibration

<p>Online hydraulic data including hydraulic model for the full distribution network. More than one full year of data.</p> <p>SCADA data source.</p> <p>Provide knowledge of the water distribution network and how the water moves in the system.</p>

restrictedcc-by-4.0Sep 2024View details →
zenodo32/100

Model data for "A model study on investigating the sensitivity of aerosol forcing on the volatilities of semi-volatile organic compounds" by Irfan et al

<p><span>Abstract: </span>This dataset contains simulation results from global aerosol-climate model ECHAM-SALSA. These simulations were performed to study the sensitivity of simulated SOA mass, CCN and radiative forcings to the assumed volatility distributions of biogenic SOA precursor species. The study employed volatility basis set (VBS) approach to represent and simulate SOA in the atmosphere. The study involved a comparative analysis between finely resolved 9-bin VBS setup with a simplified 3-bin VBS setup. It also included how the SOA mass, CCN and radiative forcing are sensitive to the volatitility of individual VBS bins.</p> <p><span>Methods: </span>Global scale aerosol-climate model simulations were performed using ECHAM-SALSA. We performed three diferent simulations each using 9-bin and 3-bin VBS setups with the volatilities increased (VBSx10) and decreased (VBSx0.1) by one-order of magnitude with respect to the original volatility (VBSx1). Another set of six different simuations were performed by increasing and decreasing the volatilities of one VBS bin at a time while keeping the original volatilities of other bins.</p> <p><span>TechnicalInfo: </span>In this study, all the ECHAM-SALSA simulations used T63 spectral truncation and 47 hybrid sigma pressure levels in horizontal and vertical resolution respectively. Simulations were performed for the year 2010 with half a year spin-up. The data was simulated with 3-hourly output for the simulation period. We then calculated monthly means of the summer months from 3-hourly data for all the model values except CDNC. We used the the 3-hourly data to analyse CDNC from grids with cloud fraction &ge; 0.95. A detailed description of model simulations is given in the setup file of each simulation.</p> <p><span>TechnicalInfo: </span>The external URL leads to a bucket containing setup files, the complete dataset (post-processed NETCDF files) presented in the manuscript, and the python scripts used for data analysis for each of the simulations.</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Model data for: Upper-lower layer coupling of recurrent circulation patterns in the Gulf of Mexico

<p>Post processed model output data for &quot;Upper-lower layer coupling of recurrent circulation patterns in the Gulf of Mexico&quot; submitted to Journal of Physical Oceanography. There are two datasets, one for the upper layer (H1) and one for the lower layer (H2). Each one contains the demeaned, detrended, and filtered (Lanczos low-pass) daily fields of layer thickness anomaly to which the authors computed the Hilbert EOFs.</p> <p>File list:</p> <p>H1_GoM_day_ssk15_st30dl.mat - this file contains the layer thickness anomaly fields for the upper layer (&lt;250m)</p> <p>H2_GoM_day_ssk15_st30dl.mat - this file contains the layer thickness anomaly fields for the lower layer (&gt;1000m)</p> <p>Scripts for plotting and processing the data into the model domain are available at:</p> <p>https://github.com/erickolvera/Olvera_et_al_21</p>

opencc-by-4.0Oct 2021View details →
zenodo32/100

Data for integrative modeling of the Nuclear Pore Complex from Schizosaccharomyces pombe

<p>Repository with input and output files utilized&nbsp;for integrative modeling&nbsp;of the Nuclear Pore Complex from Schizosaccharomyces pombe.</p>

opencc-by-4.0Nov 2021View details →
dryad32/100

Data from: The ecology of spider sociality – A spatial model

<p>The emergence of animal societies offers unsolved problems for both evolutionary and ecological studies. Social spiders are specially well suited to address this problem given their multiple independent origins and distinct geographical distribution. Based on long term research on the spider genus <em>Anelosimus</em>, we developed a spatial model that recreates observed macroecological patterns in the distribution of social and subsocial spiders. We show that parallel gradients of increasing insect size and disturbance (rain, predation) with proximity to the lowland tropical rainforest would explain why social species are concentrated in the lowland wet tropics, but absent from higher elevations and latitudes. The model further shows that disturbance, which disproportionately affects small colonies, not only creates conditions that require group living, but also tempers the dynamics of large social groups. Similarly simple underlying processes, albeit with different players on a somewhat different stage, may explain the diversity of other social systems.</p> <p> </p>

opencc-zeroDec 2020View details →
zenodo32/100

Data Sets for Article: Exploration on Learning Molecular Docking with Deep Learning Models

<p>The MOL2 and CSV file of the clustered compounds from ChemDiv are available in <strong>Additional file 2</strong></p> <p>Docking scores of training set1 and the following traing set2 for each target were saved as csv files and provided in <strong>Additional file 3.</strong></p> <p>The SMILES, MOL2, SDF of DUD-E compounds and PDB of receptors used for validation are provided in <strong>Additional file 4</strong>.</p> <p>The SMILES of 500,000 compounds randomly selected from the ChEMBL database are provided in <strong>Additional file 5</strong></p> <p>The SMILES of compounds with activities from the ChEMBL database for each target are provided in <strong>Additional file 6</strong></p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Base Data for Poro Models

<p>Base Data (.mat-Files) for models created in Altmann et al. &quot;Port-Hamiltonian formulations of poroelastic network models&quot;</p> <p>Data taken from https://zenodo.org/record/4632901#.YZI9XLso8UE</p>

opencc-by-4.0Nov 2021View details →
dryad32/100

Data from: Pulled diversification rates, lineages-through-time plots and modern macroevolutionary modelling

<p>Estimating time-dependent rates of speciation and extinction from dated phylogenetic trees of extant species (timetrees), and determining how and why they vary, is key to understanding how ecological and evolutionary processes shape biodiversity. Due to an increasing availability of phylogenetic trees, a growing number of process-based methods relying on the birth-death model have been developed in the last decade to address a variety of questions in macroevolution. However, this methodological progress has regularly been criticised such that one may wonder how reliable the estimations of speciation and extinction rates are. In particular, using lineages-through-time (LTT) plots, a recent study (Louca &amp; Pennell, 2020) has shown that there are an infinite number of equally likely diversification scenarios that can generate any timetree. This has led to questioning whether or not diversification rates should be estimated at all. Here we summarize, clarify, and highlight technical considerations on recent findings regarding the capacity of models to disentangle diversification histories. Using simulations we illustrate the characteristics of newly-proposed "pulled rates" and their utility. We recognize that the recent findings are a step forward in understanding the behavior of macroevolutionary modelling, but they in no way suggest we should abandon diversification modelling altogether. On the contrary, the study of macroevolution using phylogenetic trees has never been more exciting and promising than today. We still face important limitations in regard to data availability and methodological shortcomings, but by acknowledging them we can better target our joint efforts as a scientific community.</p>

opencc-zeroNov 2021View details →
zenodo32/100

ECMWF operational analysis data for driven the WRF model in HRB

<p>All the ECMWF&nbsp;operational analysis data forcing data (2008-2012) related to the thesis (<a href="https://opus.bibliothek.uni-augsburg.de/opus4/frontdoor/deliver/index/docId/81907/file/PhD_Zhang.pdf">https://opus.bibliothek.uni-augsburg.de/opus4/frontdoor/deliver/index/docId/81907/file/PhD_Zhang.pdf</a>)</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Pre-trained DNN model data for pruning example code

<p>Pre-trained DNN model datasets for example codes of&nbsp;neural network pruning.</p> <p>Example pruning codes are published in &quot;https://github.com/FujitsuLaboratories/CAC/tree/main/cac/pruning&quot;.</p>

opencc-zeroNov 2021View details →
zenodo32/100

MOFSimplify: Machine Learning Models with Extracted Stability Data of Three Thousand Metal-Organic Frameworks

<p>Solvent removal stability and thermal stability associated with structurally characterized metal organic frameworks.</p>

opencc-by-4.0Sep 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record