Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,634
datasets available to search
ShareScore release 0.9.0
Dataset results
1,634 results for “Data integration”
Extended data integration of motivational aspects in gamification and game-based learning educational designs
<p>This is the extended data for a systematic review article about the integration of motivational aspects in gamification and game-based learning educational designs related to teacher´s training and teacher´s professional development.</p>
Colony forming unit (CFU) data accompanying publication: Bacterial microbiome dynamics in commercial integrated aquaculture systems growing Ulva in abalone effluent water.
<p>Excel sheet with colony forming unit (CFU) data from an abalone farm growing the green seaweed <em>Ulva </em>in abalone effluent water. This dataset characterises the bacterial communities isolated from abalone effluent water entering and leaving the <em>Ulva </em>raceways, as well as from <em>Ulva </em>itself. Dataset is accompanied by two SigmaPlot files showing statistical analyses (statistical outcomes also available in publication).</p>
Characterization Data for the Manuscript: "Unraveling Metal Effects on CO2 Uptake in Pyrene-based Metal-Organic Frameworks through Integrated Lab and Computer Experiments"
<p>This entry contains characterization data for the manuscript "Unraveling Metal Effects on CO2 Uptake in Pyrene-based Metal-Organic Frameworks through Integrated Lab and Computer Experiments".</p>
An Online Integrated Development Environment for Automated Programming Assessment Systems Open Source Data
<p>This dataset accompanies the paper <em>"An Online Integrated Development Environment for Automated Programming Assessment Systems"</em>. It contains data from the usability evaluation of a feature-rich online IDE designed for integration into Automated Programming Assessment Systems (APASs). The dataset includes survey responses from 27 participants based on the Technology Acceptance Model (TAM), performance metrics such as memory usage, and qualitative user feedback. The study highlights challenges in integrating online IDEs with APASs, such as memory efficiency, load balancing, and user experience. The dataset supports further research in developing scalable, effective, and user-friendly programming education tools.<br><br>Here you can find the code changes required for the online IDE in Artemis: <a href="https://github.com/ls1intum/Artemis/pull/6706/files" target="_blank" rel="noopener">Github</a></p>
Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research
<p><strong>This is the dataset of the report: Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research</strong></p> <p>It contains the following information of all the papers from ASE, FSE, and ICSE in 2023:</p> <ul> <li>Paper title</li> <li>Keyword</li> <li>Is the source data available and accessible in the paper?</li> <li>If the source data is not available, do the authors explain why?</li> <li>Hosting platforms</li> <li>Access mode</li> <li>License</li> <li>Is their experiment data reused from previous work, or newly generated specifically for this study, or combination of both? </li> <li>Do the authors change/modify their experiment data before experiment?</li> <li>What modifications do they perform?</li> <li>Does the link provide detailed instructions about how to replicate their paper?</li> <li>Does the link contains their complete experiment data, their source code or other materials that are necessary to replicate their experiments?</li> <li>What's the data format inside the link?</li> <li>What's the content of the link?</li> </ul> <p> </p> <p>We collect the data in a rush.</p> <p>If you want to use this dataset and find any errors, please contact us ;-)</p> <p> </p> <p>Our emails:</p> <ul> <li>echo.xiangchen@gmail.com</li> <li>zhifengyao731@gmail.com</li> </ul>
The ECOLOPES Voxel Model: Multi-domain data integration for ontology-aided generative computational design of ecological building envelopes
<p>The research portrayed in this article is part of the research project ‘ECOlogical building enveLOPES: a game-changing design approach for regenerative ecosystems’ funded by Horizon 2020 Future and Emerging Technologies. The overall research project focuses on developing a multi-domain data-driven computational design framework for the design of ecological building enclosures that addresses humans, plants, animals and microbiota. This article focuses on the development of a key component of the computational workflow in which initial designs are computationally initiated generated and analyzed, namely the ECOLOPES Voxel Model that contains and correlates multi-domain spatialised data for the design process, and its interactions with other components of the ontology-aided generative computational design process for ecological building envelopes.</p> <p>This repository contains all relevant data produced in this paper. Extended technical description is available in the Appendix A to the published paper, containing listing and description of individual voxel data layers. Data were exported from the RDB server (PostgreSQL) in text-based, future-proof format (csv).</p>
Oscillations of Offshore Wind Turbines undergoing Installation II: Filtered and Integrated data - acceleration, velocity, displacement
<p>This is dataset is based on the raw measurement data from <a href="https://zenodo.org/record/5009061">https://zenodo.org/record/5009061</a></p> <p>The data included in the archives are the resampled and high-pass filtered accelerations as well as the velocity and displacement data.</p>
Data from: Integrative ichthyological species delimitation in the Greenthroat Darter complex (Percidae: Etheostomatinae)
<p>Species delimitation is fundamental to deciphering the mechanisms that generate and maintain biodiversity. Alpha taxonomy historically relied on expert knowledge to describe new species using phenotypic and biogeographic evidence, which has the appearance of investigator subjectivity. In contrast, DNA‐based methods using the multispecies coalescent model (MSC) promise a more objective approach to describing biodiversity. However, recent criticisms suggest that under some conditions the MSC may over‐split lineages, identifying species that do not reflect biological reality. Here, we reconcile these approaches using empirical data for the Greenthroat Darter complex (<em>Etheostoma lepidum</em>), a small freshwater fish species with a disjunct distribution in Texas and New Mexico, USA. We demonstrate that MSC methods recognizes all nine sampled populations as distinct species, sometimes splitting specimens from a single locality into multiple species. However, environmental, phenotypic and biogeographic evidence do not corroborate the nine species supported by the MSC. Instead, collective evidence indicates that <em>E. lepidum</em> is comprised of just three species that are consistent with the molecular phylogeny: <em>Etheostoma lepidum</em> (Greenthroat Darter) in rivers draining the eastern Edwards Plateau, <em>Etheostoma</em> cf. <em>lepidum</em> (Texas Darter) in the Concho and San Saba rivers and <em>Etheostoma</em> cf. <em>lepidum</em> (Pecos Darter) in the Pecos River. The Pecos Darter is likely highly imperiled due to its localized distribution and reliance on vanishing spring‐fed stream habitats. The impending biodiversity crisis makes integrative and swift species delimitation more necessary than ever. Our study exemplifies how classic taxonomic expertise combined with molecular phylogenetics can produce a more robust description of threatened biodiversity.</p>
Data for: Efficient geometric integrators for nonadiabatic quantum dynamics. II. The diabatic representation
<p>Data for publication: J. Roulet, S. Choi, J. Vanicek, Efficient geometric integrators for nonadiabatic quantum dynamics. II. The diabatic representation, J. Chem. Phys. <strong>150</strong>, 204113 (2019)</p> <p>Contains the data for reproducing the figures in the abovementioned publication.</p>
Data from: Building on 150 years of knowledge: the freshwater isopod Asellus aquaticus as an integrative eco-evolutionary model system
<p><strong>Introduction</strong></p> <p>This is a literature database with reference information of all papers that use the freshwater isopod <em>Asellus aquaticus</em>; published between the years 1867 and 2020. This database is intended as a starting point for scientists interested in conducting research on and with this organism. The database is currently only available as a single CSV file; future versions may be made available through a more frequently updated SQL database. The database includes specific information about the subject area and content of each paper, as well as bibliographic information. This repository is associated with the paper "Building on 150 years of knowledge: the freshwater isopod<em> Asellus aquaticus</em> as an integrative eco-evolutionary model system", published in Frontiers in Ecology and Evolution.</p> <p><strong>Details on Methods from the electronic supplement:</strong></p> <p>We used the we online search tools of Web of Science (WOS; Clarivate analytics) by searching for the term "asellus aquaticus" in six relevant databases (BIOSIS, CABI, FSTA, Medline, WOS Core Collection and Zoological Records). The database was accessed with a University License (Lund University). We manually downloaded the results and combined them to a single CSV file in Excel (Microsoft). All further processing was done in the statistical programming language R, version 4.0.2 (R Core Team 2020).</p> <p>From the 1238 obtained records we discarded three papers that were published after the year 2020 to work with completed years only. We used the subject areas assigned by WOS to provide an overview of the fields of science in which A. aquaticus has been most studied. Each paper had between one and ten subject areas assigned by WOS (2845 assignments to 1235 papers, meaning 2.3 assignments per paper, on average). To represent these multiple assignments in relation to the actual number of papers per year, we calculated "fractional assignments" by adding up all assignments to a field per year, divided by the total number of assignments in that year, and then multiplied by the number of papers. For example, if there were 12 assignments to "toxicology" in 1993, and 133 assignments in 1993, but only 21 papers published, "toxicology" would get a score of 1.9 papers in 1993 (as calculated by = (12/133)*21). In Figure 1, we represent these "fractional assignments" in the top panel, and the total number of assignments in the lower panel.</p> <p><strong>Caption for figure (1) in publication:</strong></p> <p>FIGURE 1 | Over 150 years of research on and with Asellus aquaticus. The figure summarizes published scientific literature on A. aquaticus. We conducted a quantitative literature survey with the search tools of Web of Science (WOS; Clarivate analytics) by searching for the term "asellus aquaticus" in six databases (i.e., BIOSIS, CABI, FSTA, Medline, WOS Core Collection, and Zoological Records). We found 1235 records, published between 1867 and 2020. (A) The graph shows the number of publications per year within a given subject area, as designated by WOS. (B) The graph shows the total number of publications assigned to a specific subject area. The top 10 fields account for 72.58% of all publications, and are indicated by color coding in A and B (multiple assignments are possible, summing up to 2845 assignments). The inset in B shows a wordcloud with the 100 most used keywords from all A. aquaticus’ publications. Furthermore, we compiled all records with relevant information (e.g., title, keywords, research areas, and abstract) to a single file which is available online. More details can be found in the Supplementary Material.</p>
Making sense of large-scale kinase inhibitor bioactivity data sets: a comparative and integrative analysis
<p>We carried out a systematic evaluation of target selectivity profiles across three recent large-scale biochemical assays of kinase inhibitors and further compared these standardized bioactivity assays with data reported in the widely used databases ChEMBL and STITCH. Our comparative evaluation revealed relative benefits and potential limitations among the bioactivity types, as well as pinpointed biases in the database curation processes. Ignoring such issues in data heterogeneity and representation may lead to biased modeling of drugs' polypharmacological effects as well as to unrealistic evaluation of computational strategies for the prediction of drug-target interaction networks. Toward making use of the complementary information captured by the various bioactivity types, including IC50, K(i), and K(d), we also introduce a model-based integration approach, termed KIBA, and demonstrate here how it can be used to classify kinase inhibitor targets and to pinpoint potential errors in database-reported drug-target interactions. An integrated drug-target bioactivity matrix across 52,498 chemical compounds and 467 kinase targets, including a total of 246,088 KIBA scores, has been made freely available.</p> <p>Please cite: </p> <p>https://pubmed.ncbi.nlm.nih.gov/24521231/ </p> <p>https://pubs.acs.org/doi/10.1021/ci400709d</p>
Data from: Dispersal in a house sparrow metapopulation: an integrative case study of genetic assignment calibrated with ecological data and pedigree information
<p class="western">Dispersal has a crucial role determining eco-evolutionary dynamics through both gene flow and population size regulation. However, to study dispersal and its consequences, one must distinguish immigrants from residents. Dispersers can be identified using telemetry, capture-mark-recapture (CMR) methods, or genetic assignment methods. All of these methods have disadvantages, such as, high costs and substantial field efforts needed for telemetry and CMR surveys, and adequate genetic distance required in genetic assignment. In this study, we used genome-wide 200K Single Nucleotide Polymorphism data and two different genetic assignment approaches (GSI_SIM, Bayesian framework; BONE, network-based estimation) to identify the dispersers in a house sparrow (<i>Passer domesticus</i>) metapopulation sampled over 16 years. Our results showed higher assignment accuracy with BONE. Hence, we proceeded to diagnose potential sources of errors in the assignment results from the BONE method due to variation in levels of inter-population genetic differentiation, intra-population genetic variation and sample size. We show that assignment accuracy is high even at low levels of genetic differentiation and that it increases with the proportion of a population that has been sampled. Finally, we highlight that dispersal studies integrating both ecological and genetic data provide robust assessments of the dispersal patterns in natural populations.</p>
Magnetic-Free Silicon Nitride Integrated Optical Isolator (Original Data)
<p>This is the raw dataset for the paper titled "Magnetic-Free Silicon Nitride Integrated Optical Isolator", which includes the original data and theoretical code. </p>
Data from: Integrating pheromonal and spatial information in the amygdalo-hippocampal network. Villafranca-Faus et al. 2021
<p>The local field potential (LFP) of the dorsal hippocampus (CA1) and cortical amygdala (PMCo) of mice, under head-fix recording and inmersed on a virtual environtmernt.</p> <p><strong>Paper Abstract</strong>: <br> Vomeronasal information is critical in mice for territorial behavior. Consequently, learning the territorial spatial structure should incorporate the vomeronasal signals indicating individual identity into the hippocampal cognitive map. In this work we show in mice that navigating a virtual environment induces synchronic activity, with causality in both directionalities, between the vomeronasal amygdala and the dorsal CA1 of the hippocampus in the theta frequency range. The detection of urine stimuli induces synaptic plasticity in the vomeronasal pathway and the dorsal hippocampus, even in animals with experimentally induced anosmia. In the dorsal hippocampus, this plasticity is associated with the overexpression of pAKT and pGSK3β. An amygdalo-entorhino-hippocampal circuit likely underlies this effect of pheromonal information on hippocampal learning. This circuit likely constitutes the neural substrate of territorial behavior in mice, and it allows the integration of social and spatial information.</p>
Data for: High-order geometric integrators for representation-free Ehrenfest dynamics
<p>Data for publication: S. Choi and J. Vaníček, High-order geometric integrators for representation-free Ehrenfest dynamics,<a href="https://aip.scitation.org/doi/10.1063/5.0061878"> J. Chem. Phys. 155, 124104 (2021)</a>.</p> <p>Contains the data for reproducing the figures in the above-mentioned publication.</p>
Data for Integrated Step Selection Analysis of translocated female greater sage-grouse in the 60 days post-release, North Dakota 2018-2020
<p>The data include used and random available steps at 11-hour resolution generated for 26 female greater sage-grouse in the 60 days post-translocation to North Dakota, with associated environmental predictors and individual information. The code fits individual habitat selection models in an Integrated Step Selection Analysis framework.</p> <p>Data used to fit the models described in:</p> <p>Picardi, S., Ranc, N., Smith, B.J., Coates, P.S., Mathews, S.R., Dahlgren, D.K. <i>Individual variation in temporal dynamics of post-release habitat selection</i>. Frontiers in Conservation Science (in review)</p> <p>Code used to implement the analysis is available on GitHub: https://github.com/picardis/picardi-et-al_2021_sage-grouse_frontiers-in-conservation</p>
Data, code and supplementary material for "A data integration framework for spatial interpolation of temperature observations using climate model data"
<p>Each zipped file contains code and data to reproduce the results in the paper and supplementary material. The Cyprus folder contains also the files to run the model, as well as the associated results. The Morocco folder only contains the results and the code used to manipulate it. </p>
Yield Prediction Through Integration of Genetic, Environment, and Management Data Through Deep Learning: Cleaned Data
<p>The included files and script are to allow for reconstruction of the data directory and cleaned data used in "Yield Prediction Through Integration of Genetic, Environment, and Management Data Through Deep Learning" ( https://doi.org/10.1101/2022.07.29.502051 ). Code used is available at 10.5281/zenodo.7401113 .</p> <table> <tbody> <tr> <th>Filename</th> <th>Description</th> </tr> <tr> <td>interim.tar.gz</td> <td>Contains site grouping dictonary</td> </tr> <tr> <td>processed.tar.gz</td> <td>Processed data</td> </tr> <tr> <td>raw.tar.gz</td> <td>Input data</td> </tr> <tr> <td>SetupInstructions.sh</td> <td>Bash script to prepare folders and unzipped data expected by code in 10.5281/zenodo.7401113</td> </tr> <tr> <td>SetupInstructions.txt</td> <td>Instructions for unzipping the data</td> </tr> <tr> <td>Train_Test_Split_Reference_Phenotypes.csv</td> <td>Reference spreadsheet to allow for easily exploring training and test set groupings</td> </tr> </tbody> </table> <ul> </ul> <p>This work was supported through funding from the USDA Agricultural Research Service, ARS project number 5070-21000-041-000-D. Raw data provided by the [Genomes to Field Initiative](https://www.genomes2fields.org/) and the [Daymet database](https://daymet.ornl.gov/).</p>
Data supplement for the paper "An integrated computational strategy to predict personalized cancer drug combinations by reversing drug resistance signatures"
<p>This dataset contains the the following data created for the paper "An integrated computational strategy to predict personalized cancer drug combinations by reversing drug resistance signatures".</p> <p>Data listing:</p> <p>Cell line-specific drug resistance signatures (CDRSR)</p> <p>Patient-specific drug resistance signatures (CTR-DB)</p>
Data for: Induction of C4 genes during de-etiolation of Gynandropsis gynandra evolved through changes in cis allowing integration into ancestral C3 gene regulatory networks
<p>C4 photosynthesis has evolved repeatedly and in doing so repurposed existing enzymes to drive a carbon pump that limits the oxygenation reaction of RuBisCO. C4 proteins accumulate to levels matching those of the photosynthetic apparatus, and to allow this gene expression must be modified over evolutionary time. To better understand this rewiring of gene expression we undertook RNA-SEQ and <span>DNaseI</span>-SEQ on de-etiolating seedlings of C4 <em>Gynandropsis gynandra</em> which is evolutionarily proximate to C3 <em>A. thaliana</em>. Changes in chloroplast ultrastructure and C4 gene expression in <em>G. gynandra</em> were coordinated and rapid. C3 and C4 photosynthesis genes showed similar induction patterns, but C4 genes from <em>G. gynandra</em> were more strongly induced than orthologs from <em>A. thaliana</em>. The cistrome of <em>G. gynandra</em> was enriched in TGA, TCP and homeodomain binding sites. Furthermore,<em> in vivo</em> binding data in <em>G. gynandra</em> highlighted TGA and homeodomain as well as light responsive elements such as G- and I-box motifs as being associated with the rapid increase in transcripts derived from C4 genes. Although promoters of <em>PPDK</em> and <em>ASP1</em> from <em>G. gynandra</em> contained distinct light responsive elements, promoters from both <em>A. thaliana</em> and <em>G. gynandra</em> allowed high expression. Deletion analysis of the <em>Ppa6</em> gene from <em>G. gynandra</em> showed that regions containing G- and I-boxes were necessary for high expression. The data support a model in which accumulation of transcripts derived from C4 genes in leaves of <em>G. gynandra</em> is enhanced compared with homologs in <em>A. thaliana</em> because a variety of modifications in <em>cis</em> allowed integration into ancestral transcriptional networks.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.