Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,624

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,624 results for “data integration”

Learn how ShareScore rates datasets ↗
edi60/100

Interagency Ecological Program: Integrated Dataset of Phytoplankton Enumeration Data in the San Francisco Estuary, 1992-2024

Phytoplankton community composition is an important driver of zooplankton productivity and food supply for higher trophic levels in the San Francisco Estuary. Various monitoring surveys throughout the region collect phytoplankton enumeration data dating back to the 1990s. These include surveys from the CA Department of Water Resources (CADWR), CA Department of Fish and Wildlife (CDFW), the US Bureau of Reclamation (USBR), and the US Geological Survey (USGS). These surveys collect data via various sampling and laboratory methods which are not always directly comparable. This integrated dataset includes both enumeration counts and well-documented metadata to allow for informed decision-making in the integration of these data. It also standardizes taxonomic names between groups via a key list. Note that, in this dataset, we make conservative decisions about taxonomic resolution to ensure maximum compatibility between groups. For more detailed metadata and higher taxonomic resolution, refer to individual surveys’ publications or reach out to their primary contact.

openCC (other)Nov 2025View details →
OpenNeuro52/100

A multi-modal human neuroimaging dataset for data integration: simultaneous EEG and fMRI acquisition during a motor imagery neurofeedback task: XP1

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo52/100

JasonAlongTrack: A reformatted version of the Integrated Multi-Mission Ocean Altimeter Data for Climate Research Version 5.1

<p>JasonAlongTrack contains geo-registered along-track sea surface height anomalies with respect to the DTU15 mean sea surface at 1-second intervals from Jason-class altimeters, reformatted for convenience into a 3D array with dimensions of along-track direction by geographically sorted track number by cycle.</p><p>This is a reformatted version of Beckley et al.'s <i>Integrated Multi-Mission Ocean Altimeter Data for Climate Research complete time series Version 5.1</i> dataset, available from <a href="https://podaac.jpl.nasa.gov/dataset/MERGED_TP_J1_OSTM_OST_ALL_V51">https://podaac.jpl.nasa.gov/dataset/MERGED_TP_J1_OSTM_OST_ALL_V51</a>. &nbsp;&nbsp;</p><p>The changes are as follows. Altimeter passes are sorted according to their initial longitude, then split into descending and ascending potions with all descending tracks preceding all ascending tracks. Descending tracks are then flipped so that latitude increases in the alongtrack direction for all tracks. This leads to a 3373 x 254 matrix of observational locations, with the first dimension being the along-track location and the second dimension being the track index. Sea surface height anomaly, time, and flag values are then placed into their correct locations within this matrix, such that these three variables are all of size 3373 x 254 x K where K is the number of cycles, currently 1087. A very good approximation to the time at each of the 3373 x 254 x K observation points is constructed with a length K array of cycles times together with a 3373 x 254 array of time offsets. A median-based editing criterion in introduced to identify a small number of suspect data points. &nbsp;These are set to a value of NaN in sla, but their positions and values are recorded in rejected_index and rejected_values, respectively. &nbsp;The DTU15 mean dynamic topography (mdt) is included, in addition to the mean sea surface field already provided, interpolated onto the track locations using bicubic interpolation. &nbsp;Finally, an estimate of the small-scale noise level, sigma, is produced using a wavelet transform filter.</p>

opencc-by-4.0Nov 2023View details →
zenodo52/100

Integrated analysis of anatomical and electrophysiological human intracranial data

<p>The exquisite spatiotemporal precision of human intracranial EEG recordings (iEEG) permits characterizing neural processing with a level of detail that is inaccessible to scalp-EEG, MEG, or fMRI. However, the same qualities that make iEEG an exceptionally powerful tool also present unique challenges. Until now, the fusion of anatomical data (MRI and CT images) with the electrophysiological data and its subsequent analysis has relied on technologically and conceptually challenging combinations of software. Here, we describe a comprehensive protocol that addresses the complexities associated with human iEEG, providing complete transparency and flexibility in the evolution of raw data into illustrative representations. The protocol is directly integrated with an open source toolbox for electrophysiological data analysis (FieldTrip). This allows iEEG researchers to build on a continuously growing body of scriptable and reproducible analysis methods that, over the past decade, have been developed and employed by a large research community. We demonstrate the protocol for an example complex iEEG data set to provide an intuitive and rapid approach to dealing with both neuroanatomical information and large electrophysiological data sets. We explain how the protocol can be largely automated and readily adjusted to iEEG data sets with other characteristics. The protocol can be implemented by a graduate student or post-doctoral fellow with minimal MATLAB experience and takes approximately an hour, excluding the automated cortical surface extraction.</p> <p>This collection contains the data described in the protocol and that can be used to replicate all results.</p>

opencc-by-sa-4.0Dec 2017View details →
zenodo52/100

CATCH-EyoU: Exploiting European data and testing the integrated theory of youth active EU citizenship: EACEA subset analysis

<p>This dataset was created within the research project Constructing AcTive CitizensHip with European Youth: Policies, Practices, Challenges and Solutions (CATCH-EyoU) funded by European Union, Horizon 2020 Programme, Grant Agreement No 649538. Work Package 4 of this project (Exploiting European data and testing the integrated theory of youth active EU citizenship) is focused on the re-analysis of existing European data. This dataset contains a subset of data originally collected within the project &ldquo;<em>EACEA 2010/03: Youth Participation in Democratic Life</em>&rdquo;, coordinated by the London School of Economic and Political Science. Specifically, an online questionnaire survey in seven European countries was conducted among young people age 15-30 in 2011. This dataset contains a subset of 22 variables that were employed for the reanalysis within the CATCH-EyoU project.</p>

opencc-by-4.0Jul 2018View details →
zenodo48/100

Data and R code for Tansley review New Phytologist 2021: "An integrated framework of plant form and function: The belowground perspective"

<p>The files in this archive are related to the paper of Weigelt, Mommer, Andraczek et al. (2021) An integrated framework of plant form and function: The belowground perspective. Tansley Review New Phytologist. The paper developed and tested a new conceptual framework of plant form and function linking above and belowground traits of 2510 species. We found that an integrated, whole-plant trait space required as much as four axes. The two main axes represented the fast-slow &lsquo;conservation&rsquo; gradient on which leaf and fine-root traits were well aligned, and the &lsquo;collaboration&rsquo; gradient in roots. The two additional axes were separate, orthogonal plant size axes for height and rooting depth.</p> <p>This archives contains four files:</p> <ol> <li><strong>Weigelt et al.2021RCode.DataCleaning.txt</strong> - &nbsp;RCode for the complete data processing starting with the downloaded database files from the Plant Trait Database version 5.0 (TRY, Kattge et al. 2020), the Global Root Trait database (GRooT, Guerrero-Ramirez et al. 2020) and a small number of additional data files listed in Table S2 of the original paper. Additional information was later incorporated using FungalRoot Database (Soudzilovkaia et al. 2020), nodDB Database (Tedersoo et al. 2018) and a compiled dataset on rooting depth (Fan et al. 2017). The code processes, cleans and merges the data and produces a final table for PCA analysis of species specific mean traits. This final table is provided as a second file in this archive (Weigelt_et_al_2021_Main.PCA.Matrix.xlsx). A second part of the RCode.DataCleaning extracts species-specific individual trait data where root and shoot traits were measured on the same plant individual or plot. This data was compiled from 43 studies identified in Table S2&nbsp; of the original publication. The final table for individual trait data is the third file in this archive (Weigelt_et_al_2021_Individual.PCA.Matrix.xlsx).</li> <li><strong>Weigelt_et_al_2021_Main.PCA.Matrix.xlsx</strong> &ndash; Datafile with species-specific global mean trait data for 17 traits of 2510 species with at least one root and one shoot trait available. Meta-data is provided in the data file.</li> <li><strong>Weigelt_et_al_2021_Individual.PCA.Matrix.xlsx</strong> &ndash; Datafile with species-specific trait data where root and shoot traits were measured on the same individual or plot for 6 traits of 455 species. Meta-data is provided in the data file.</li> <li><strong>Weigelt et al.2021RCode.Analysis.txt &ndash; </strong>RCode for all analyses and figures provided in the paper for both the species mean and individual based dataset. The Code is annotated to help reproducibility of the analysis.</li> </ol>

opencc-by-4.0Dec 2020View details →
zenodo48/100

Data on the Digital Economy and Society Index (DESI), the ASEAN Digital Integration Index (ADII), and the Digital Intelligence Index (DII)

<p>This dataset contains the quantitative measurement of the Digital Economy and Society Index (DESI), the ASEAN Digital Integration Index (ADII), and the Digital Intelligence Index (DII) in 2019.</p>

opencc-by-4.0Nov 2023View details →
zenodo48/100

An Integrated Usability Framework for Evaluating Open Government Data Portals and Analysis of EU and GCC OGD Portals

<p><span>This dataset contains data collected during a study (<em><strong>"<a href="https://arxiv.org/ftp/arxiv/papers/2403/2403.08451.pdf">An Integrated Usability Framework for Evaluating Open Government Data Portals: Comparative Analysis of EU and GCC Countries</a>"</strong></em>) conducted by Fillip Molodtsov and Anastasija Nikiforova (University of Tartu).</span></p> <p><span>&nbsp;</span><span>It being made public both to act as supplementary data for the paper and in order for other researchers to use these data in their own work potentially contributing to the improvement of current data ecosystems and develop user-friendly, collaborative, robust, and sustainable open data portals.</span></p> <p><span>***Purpose of the study***</span></p> <p><span>This paper develops an integrated framework for evaluating OGD portal effectiveness that accommodates user diversity (regardless of their data literacy and language), evaluates collaboration and participation, and the ability of users to explore and understand the data provided through them. </span></p> <p><span>The framework is validated by applying it to 33 national portals across European Union (EU) and Gulf Cooperation Council (GCC) countries, as a result of which we rank OGD portals, identify some good practices that lower-performing portals can learn from, and common shortcomings.</span></p> <p><span>***Methodology***</span></p> <p><span>(1) systematic literature review to establish a knowledge base and identify frameworks have been used to evaluate OGD portals, we conducted a systematic literature review - Dataset_ Usability_Framework_SLR;</span></p> <p><span>(2) development of the Integrated Usability Framework for Evaluating Open Government Data Portals, which content is based on the outputs of the first step, along with selected articles of experts in portal design, and an exploratory assessment of the French, Irish, Estonian and Spanish portals - Dataset_Integrated_Usability_Framework;</span></p> <p><span>(3) data collection, that is a completion of the protocol developed in the previous step by analysing 34 national OGD portals of the EU and GCC countries. When all individual protocols were collected, the total score are calculated using the weighting system. The average scores are calculated for the EU and GCC. The portals are ranked. The top portals (best performers) are determined for each dimension - Dataset_EU_GCC_OGDportal_Usability_results_clustering.</span></p> <p><span>(4) identification of relationships and patterns among different portals based on their performance metrics as a result of the cluster analysis. By calculating the average dimensional scores of portals from both types of clusters, their performance across multiple dimensions is evaluated - Dataset_EU_GCC_OGDportal_Usability_results_clustering.</span></p> <p>&nbsp;</p> <p><strong><em><span>For more details see Molodtsov, F., Nikiforova, A. (2024). &ldquo;An Integrated Usability Framework for Evaluating Open Government Data Portals: Comparative Analysis of EU and GCC Countries&rdquo;. In Proceedings of the 25th Annual International Conference on Digital Government Research (DGO 2024), June 11--14, 2024, Taipei, Taiwan, 10.1145/3657054.3657159</span></em></strong></p> <p><span>***Format of the file***</span></p> <p><span>.xls, .csv</span></p> <p><span>***Licenses or restrictions***</span></p> <p><span>CC-BY</span></p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Data for "emIAM v1.0: an emulator for Integrated Assessment Models using marginal abatement cost curves"

<p>This dataset contains&nbsp;codes, data, tables, andd figures (high resolution)&nbsp;related to the following publication: Xiong, W., K. Tanaka, P. Ciais, D. J. A. Johansson, M. Lehtveer (2022) emIAM v1.0: an emulator for Integrated Assessment Models using marginal abatement cost curves.&nbsp;Submitted to arXiv on 23 December 2022.</p>

opencc-by-4.0Dec 2022View details →
zenodo48/100

Example of datasets processed to demonstrate a multisource data integration methodology

<p>This dataset contains the&nbsp;data processed to demonstrate the multi-source spatial data&nbsp;integration methodology proposed in the paper &quot;Multisource spatial data integration for use cases applications&quot;.</p> <p>It contains:</p> <p>- the building footprint extracted from the IFC model of a newly designed&nbsp;building in WKT format, by using the GeoBIM_Tool (<a href="https://github.com/twut/GEOBIM_Tool">https://github.com/twut/GEOBIM_Tool</a>);</p> <p>- the extrusion of the footprint until the measured height measured with the same GeoBIM_Tool;</p> <p>- a portion of the Rotterdam 3D city model generated with 3dfier and available at&nbsp;https://3d.bk.tudelft.nl/opendata/3dfier/, converted in CityJSON&nbsp;with the citygml-tools (https://www.cityjson.org/tutorials/conversion/),&nbsp;developed to convert data between CityGML and CityJSON.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Multi-stakeholder research data management training as a tool to improve the quality, integrity, reliability and reproducibility of research: Quantitative data of the post-course surveys

<p>Data contains&nbsp;doctoral students&#39; and postdoc researchers&#39; (n=168) self-ratings of their RDM competencies before and after the 3 ECTS credits &quot;Basics of Research Data Management&quot; (BRDM) trainings held 2019-2021 in the University of Turku and &Aring;bo Akademi University, Finland. Moreover, data contains respondents&#39; self-reported further learning needs.</p>

opencc-by-4.0May 2022View details →
zenodo48/100

Data, scripts, and R Notebook for Carneiro et al 2023. Flight performance and wing morphology in the bat Carollia perspicillata: biophysical models and energetics. Integrative Zoology DOI:10.1111/1749-4877.12707

<p>Files provided as supporting information for the paper by Carneiro et al. 2023. Flight performance and wing morphology in the bat&nbsp;<em>Carollia perspicillata</em>: biophysical models and energetics. Integrative Zoology. DOI:10.1111/1749-4877.12707</p> <p>File descriptions</p> <p>ArmTA.txt - Temperature and surface areas for arms of <em>C. perspicillata</em> after flight experiment<br> BodyTA.txt - Temperature and surface areas for body of <em>C. perspicillata</em> after flight experiment<br> HeadTA.txt - Temperature and surface areas for head of <em>C. perspicillata</em> after flight experiment<br> WingTA.txt - Temperature and surface areas for wings (patagium) of <em>C. perspicillata</em> after flight experiment<br> WingMorph.txt - Morphological variables measured in the body and wings of <em>C. perspicillata</em><br> HeatLoss.R - Function to estimate heat loss (Qt)<br> PowFlight.R - Function to estimate minimum power required to fly<br> Script-HeatLoss-FlightPerformance.R - R script with set of analyses performed<br> SupportingInformationFile.docx - R notebook with set of analyses performed, word format<br> SupportingInformationFile.nb.html - R notebook with set of analyses performed, html format<br> SupportingInformationFile.Rmd - R notebook with set of analyses performed (R markdown)</p> <p>For the R scripts (Script-HeatLoss-FlightPerformance.R) and notebook (<br> SupportingInformationFile.Rmd) to work and be compiled, all files need to be copied to the same folder.</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Using machine learning to integrate genetic and environmental data to model genotype-by-environment interactions

<p>Files generated from the study described in&nbsp;<a href="https://doi.org/10.1101/2024.02.08.579534">Fernandes et. al (2024)</a> .</p> <p>The file "cvs_h2s.csv" comprises the coefficient of variation and the Cullis heritability for each environment.</p> <p>The file "all_predictions.csv" contains the predictions from all the models evaluated, in different cross-validation (CV) scenarios.</p> <p>The file "coincidence_index.csv" has the Coincidence Index (CI) for each CV and models evaluated in our study.</p> <p>Our study used the multi-environment maize yield trials data from the Genomes to Fields 2022 initiative (<a href="https://doi.org/10.1186/s13104-023-06421-z">Lima et. al 2024</a>).</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

ExcapeDB: An integrated large scale dataset facilitating Big Data analysis in chemogenomics

<p>ExcapeDB: An integrated large scale dataset facilitating Big Data analysis in chemogenomics</p> <p>Supplementary file (full dataset download)</p> <p>- v2 with SMILES errors fixed&nbsp; (19.01.2019)</p>

opencc-by-sa-4.0Nov 2016View details →
zenodo48/100

Data and code related to the paper: "Integrated stretchable pneumatic strain gauges for electronics-free soft robots"

<p>This folder contains the raw data and Matlab scripts to reproduce the plots and supplementary movies for the paper:</p> <p>Anastasia Koivikko, Vilma Lampinen, Mika Pihlajam&auml;ki, Kyriacos Yiannacou, Vipul Sharma &amp; Veikko Sariola, &quot;Integrated Stretchable Pneumatic Strain Gauges for Electronics-Free Soft Robots&quot;, Communications Engineering, 1, 14 (2022).</p> <p><a href="https://doi.org/10.1038/s44172-022-00015-6">Link to the paper</a>.</p> <p>The scripts were tested on Matlab R2021a on Windows.</p> <p>Generally speaking, there is a folder containing the plotting scripts for each figure. In most cases, the folder contains scripts named <strong>plot&lt;...&gt;.m</strong>&nbsp;that recreate the actual plots. Some folders also have a scripts <strong>analyze&lt;...&gt;.m</strong>&nbsp;to analyze the data; these need to be run before the actual plotting.</p> <p>For more details, please see the paper.</p>

opencc-by-4.0Jun 2022View details →
edi48/100

Course Materials for Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730)

In today's world, understanding environmental data and making informed decisions based on it is crucial for addressing complex environmental challenges. Yale School of the Environment's Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730) course serves as an introduction to the integration of environmental data using R programming language, coupled with machine learning techniques. This dataset contains a zip file with all the data files used in this course, along with a README that has the metadata for those files.

openCC (other)Jul 2025View details →
edi48/100

CCE LTER process cruise, in the California Current region, event log records including date, time, position and activity for use in post-cruise data integration based on co-sampling indexes. From 2006 to 2019 CCE LTER used a locally developed event logging system. During P2107, CCE LTER started to utilize the R2R Event Logger on UNOL ships, 2006 - 2024 (ongoing).

The event logger program developed and maintained by the California Cooperative Oceanic Fisheries Investigations, SIO, program is used aboard CCE LTER process cruises to create indexes with temporal, spatial and activity information for post-cruise data integration. The event log is configured aboard the ship for the recording of sampling events by both ship crew personnel on the bridge, and research personnel in the lab. The event log is processed post-cruise to correct for various errors.

openCC0Aug 2025View details →
zenodo44/100

EPA Integrated Planning Model (IPM) National Electric Energy Data System (NEEDS) database

EPA is making the latest power sector modeling platform available, including the associated input data and modeling assumptions, outputs, and documentation.

opencc-zeroFeb 2020View details →
zenodo44/100

RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)

<p>RDF version of the data from Choi, JS., Ha, M.K., Trinh, T.X. et al. Towards a generalized toxicity prediction model for oxide nanomaterials using integrated data from different sources. Sci Rep 8, 6110 (2018)</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Supplementary material 3: World Spider Catalog Bibliographic Data: Treatments from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063

List of journal/publisher by ranked by treatment count exported from the World Spider Catalog 14 October 2014 with total treatments by source, cumulative treatments, and cumulative proportion of treatments.

opencc-by-4.0Feb 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record