Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “data analysis”
Data from: Process-based modelling of nonharmonic internal tides using adjoint, statistical, and stochastic approaches. Part I: statistical model and analysis of observational data
<p>Meta data updated after publication.</p> <p> </p>
Numerical back-analysis of short-term convergence data of sections within zone A (from chainage 1905 to chainage 2723) in the Fréjus road tunnel
<p>Numerical back-analysis of short-term convergence data of sections within zone A (from chainage 1905 to chainage 2723) in the Fréjus road tunnel</p>
FlyTracker Video Analysis and Data Extraction
<p>Protocol video showing how to use the MATLAB package "FlyTracker" to analyze locomotor behavior in a video and extract the data in a format compatible with any worksheet software.</p> <p>Full uncompressed video made with FinalCut Pro (higher quality than version available at STAR Protocols)</p> <p>MATLAB package: <a href="https://github.com/kristinbranson/FlyTracker/archive/refs/heads/main.zip">FlyTracker</a></p> <p>MATLAB Script: <a href="https://github.com/LaurentSeroude/FlyTrackerExtraction">FlyTracker Extraction</a></p> <p>Peer-reviewed publications:</p> <p>Genome 64,139,2021 <a href="https://github.com/LaurentSeroude/FlyTrackerExtraction/blob/main/Genome%2064%2C139%2C2021.pdf">PDF</a></p> <p><a href="https://star-protocols.cell.com/protocols/2193">STAR Protocols 3,101888,2022</a></p> <p>Peer-reviewed protocol: STAR Protocols in press</p>
Data from: Integrating a UAV-derived DEM in object-based image analysis increases habitat classification accuracy on coral reefs
<p>Very shallow coral reefs (< 5 m deep) are naturally exposed to strong sea surface temperature variations, UV radiation and other stressors exacerbated by climate change, raising great concern over their future. As such, accurate and ecologically informative coral reef maps are fundamental for their management and conservation. Since traditional mapping and monitoring methods fall short in very shallow habitats, shallow reefs are increasingly mapped with Unmanned Aerial Vehicles (UAVs). UAV-imagery is commonly processed with Structure-from-Motion (SfM) to create orthomosaics and Digital Elevation Models (DEMs) spanning several hundred metres. Techniques to convert these SfM products to ecologically relevant habitat maps are still relatively underdeveloped. Here we demonstrate that incorporating geomorphometric variables (the DEM and its derivatives) in addition to spectral information (the orthomosaic) can greatly enhance the accuracy of automatic habitat classification. Therefore, we mapped three very shallow reef areas off KAUST on the Saudi Arabian Red Sea coast with an RTK-ready UAV. Imagery was processed with SfM, and classified through Object-Based Image Analysis (OBIA). Within our OBIA workflow, we observed overall accuracy increases of up to 11% when training a Random Forest classifier on both spectral and geomorphometric variables as opposed to traditional methods that only use spectral information. Our work highlights the potential of incorporating a UAV's DEM in OBIA for benthic habitat mapping, a promising but still scarcely exploited asset.</p>
Supplementary data for: New paromomyids (Mammalia, Primates) from the Paleocene of southwestern Alberta, Canada, and an analysis of paromomyid interrelationships
<p>Paromomyidae are one of several families of plesiadapiforms that flourished during the Paleocene in North America soon after the extinction of non-avian dinosaurs some 66 million years ago. Although they are often among the best-represented plesiadapiforms in mammalian faunas in both North America and Europe, the early history of paromomyids is poorly understood, and their fossil record at higher latitudes is comparatively depauperate. We report here on the discovery of two new species of paromomyids from Paleocene deposits in southwestern Alberta:<strong> <em>Edworthia greggi</em> new species</strong> is the second known species of the basal paromomyid <em>Edworthia</em> Fox, Scott, and Rankin, whereas <strong><em>I</em><em>gnacius glenbowensis</em> new species</strong> is among the most abundantly represented species of <em>Ignacius</em> Matthew and Granger. These new discoveries document for the first time parts of the upper dentition of <em>Edworthia</em>, and the new species of<em> Ignacius</em> represents the first new, pre-Clarkforkian species of the genus to be described in nearly one hundred years. A comprehensive phylogenetic analysis of nearly all known paromomyid taxa (including the new species described herein) recovered both species of <em>Edworthia</em> near the base of the paromomyid tree in a polytomy with <em>Paromomys depressidens</em>, and a paraphyletic <em>Ignacius</em>. The new paromomyids from Alberta not only increase the known taxonomic diversity of <em>Edworthia</em> and <em>Ignacius</em>, but also add significantly to knowledge of the dental anatomy of these poorly known genera, and further add to a uniquely Canadian complement of Paleocene plesiadapiforms.</p>
Cryptocurrency Fraud and Code Sharing Data Set and Analysis Code
<p>This release covers the state of the data and associated analysis code for determining code sharing between cryptocurrency codebases funded through the end of the original NSF CRII award. This material is based on work supported by the National Science Foundation under Grant CNS-1849729.</p>
Data analysis results for: "MoDLE: High-performance stochastic modeling of DNA loop extrusion interactions"
<p>Due to technical issues we are unable to upload the updated version of this dataset on Zenodo.<br> <br> The latest version of this dataset can be found on the NRID research data archive at DOI <a href="https://doi.org/10.11582/2022.00056">10.11582/2022.00056</a>.</p>
Data for Apalachicola Bay oyster demographic analysis
<p>Processed data describing the number and lengths of oysters sampled from long-term fisheries independent sampling in Apalachicola Bay, Florida. Original (raw) data collected by Florida Fish and Wildlife Conservation Commission and Florida Department of Agriculture and Consumer Services.</p>
Data of "A workflow to study the microbiota profile of piglet's umbilical cord blood: from sampling to data analysis".
<p>The present study proposes a workflow – from the sampling method to DNA extraction, bioinformatics and data analysis – that characterises the bacterial profile of umbilical cord blood samples, taking into account the contaminants found throughout the procedure of bacterial DNA extraction and amplification.</p> <p>Ps_umbilical.rds: A phyloseq object file of data containing the amplicon sequences variants (ASVs) of thirteen umbilical cord samples and two negative control samples, created by DADA2.</p> <p>R-script.doc: A word document containing the scripts used to characterize the taxonomical composition of the fifteen umbilical cord samples and two negative control samples before and after the application of Decontam R-package (Davis et al., 2018).</p> <p>metadata.docx: meta data for R-script.doc</p>
Unique or not unique? Comparative genetic analysis of bacterial O-antigens from the Oxalobacteraceae family.Supplementary_data
<p>This is Supplementary files for the paper "Unique or not unique? Comparative genetic analysis of bacterial O-antigens from the Oxalobacteraceae family".</p> <p>The description for each dataset is located inside files.</p>
Data - Retrospective analysis of measures to reduce large whale entanglements in a lucrative commercial fishery
<p>Datasets for the manuscript titled "Retrospective analysis of measures to reduce large whale entanglements in a lucrative commercial fishery" (<a href="https://doi.org/10.1016/j.biocon.2022.109880">https://doi.org/10.1016/j.biocon.2022.109880</a>). </p> <p> </p> <p>Please see the README file for further information on each data file.</p>
dhaw/fluCodeImperial: Using real-time data to guide decision-making during an influenza pandemic: a modelling analysis
<p><strong>All codes and data used for "Using real-time data to guide decision-making during an influenza pandemic: a modelling analysis" are included in this folder. The file "runExamples.m" contains a step-by-step method for reproducing figures and running model fits. Ensure that all data and code files are in the same directory, then a single execution of “runExamples” on the command line will generate all main and supplementary figures in the manuscript. There is one line of code per figure, clearly marked, that can be commented out as desired. In order to run the MCMC adaptive algorithm, a single line of code, also clearly marked, must be commented back in. Instructions to change the single-state example are given at the top of the file “runExamples.m”. The saved state selection of California (“state=1”) is consistent with all results presented in the manuscript. </strong></p> <p><strong> </strong></p> <p><strong>Plots make use of files from the following sources, with some modifications:</strong></p> <p><strong>Holger Hoffmann (2022). Violin Plot (https://www.mathworks.com/matlabcentral/fileexchange/45134-violin-plot);</strong></p> <p><strong>Evan (2022). Plot Groups of Stacked Bars (https://www.mathworks.com/matlabcentral/fileexchange/32884-plot-groups-of-stacked-bars);</strong></p> <p><strong>John Onofrey (2022). Shaded Plots and Statistical Distribution Visualizations (https://www.mathworks.com/matlabcentral/fileexchange/69203-shaded-plots-and-statistical-distribution-visualizations)</strong></p>
Data sets for "Automated cell segmentation for reproducibility in bioimage analysis"
<p>This is the raw data sets used in "Automated cell segmentation for reproducibility in bioimage analysis", published in Synthetic Biology (Oxford Academic)</p>
Data, sample sizes, and R code for analysis of: Variation in mutation (co)variances
<p>Because of pleiotropy, mutations affect the expression and inheritance of multiple traits and, together with selection, are expected to shape standing genetic covariances between traits and eventual phenotypic divergence between populations. It is therefore important to find if the M matrix, describing mutational variances of each trait and covariances between traits, varies between genotypes. We here estimate the M matrix for six locomotion behavior traits in lines of two genotypes of the nematode <em>Caenorhabditis elegans </em>that accumulated mutations in a nearly-neutral manner for 250 generations. We find significant mutational variance along at least one phenotypic dimension of the M matrices, but neither their size nor their orientation had detectable differences between genotypes. The number of generations of mutation accumulation, or the number of MA lines measured, was likely insufficient to sample enough mutations and detect potentially small differences between the two M matrices. We then tested if the M matrices were similar to one G matrix describing the standing genetic (co)variances of a population derived by the hybridization of several genotypes, including the two measured for M, and domesticated to a lab-defined environment for 140 generations. We found that the M and G were different because the genetic covariances caused by mutational pleiotropy in the two genotypes are smaller than those caused by linkage disequilibrium in the lab population. We further show that M matrices differed in their alignment with the lab population G matrix. If generalized to other founder genotypes of the lab population, these observations indicate that selection does not shape the evolution of the M matrix for locomotion behavior in the short-term of a few tens to hundreds of generations and suggests that the hybridization of <em>C. elegans </em>genotypes allows selection on new phenotypic dimensions of locomotion behavior.</p>
Archive of the microtremor data observed at rock/stiff-soil sites and the analysis results
<p>This archive includes the microtremor data observed at 15 rock/stiff-soil sites and the analysis results, which were fully described in a paper "Spatial autocorrelation method for a simple microtremor array survey at rock/stiff-soil sites" by Ikuo Cho (2023, Geophysical Journal International, in press).</p>
Replication files for "The role of actors' issue and sector specialization for policy integration in the parliamentary arena: An analysis of Swiss biodiversity policy using text as data"
<p>The ZIP file contains all data and code to replicate the analyses reported in the following paper.</p> <p>Reber, U., Ingold, K., & Fischer, M. (2023). The role of actors' issue and sector specialization for policy integration in the parliamentary arena: An analysis of Swiss biodiversity policy using text as data. <em>Policy Sciences</em>. <a href="https://doi.org/10.1007/s11077-022-09490-2">https://doi.org/10.1007/s11077-022-09490-2</a></p> <p>If you use any of the material included in this repository, please refer to the paper.</p>
Data and Analysis from "Analysis of context-specific KRAS-effectors (sub)complexes in Caco-2 cells"
<p>Data, data processing and data analysis for manuscript "Analysis of context-specific KRAS-effectors (sub)complexes in Caco-2 cells". (Preprint available <a href="https://doi.org/10.1101/2022.08.15.503960">here</a>)</p> <p><strong>Analysis of AP-MS data</strong>: analysis.zip</p> <p>Contains the following scripts as well as their outputs:</p> <ul> <li>01_preparation.R R script for filtering and processing our mass spec data.</li> <li>02_diffbinding.R R script for differential analysis followed by gene set enrichment.</li> <li>03_funcstats.R R script for statistical analysis over different ontology terms.</li> <li>04_semantic_analysis.R R script for the GO semantic analysis for the output of 02 and 03.</li> <li>05_1_random_walks.py Python script for performing random walks for specific functional terms.</li> <li>05_2_random_walks_analysis.R R script for the analysis and visualization of the output of 05_1.</li> </ul> <p>The required input data is deposited in the "data" sub-folder, taken directly from the linked PRoteomics IDEntification database (PRIDE) <a href="https://www.ebi.ac.uk/pride/archive/projects/PXD035399">entry</a>.</p> <p>Interactive visualization of the results of most of this analysis is available on <a href="https://github.com/PhilippJunk/kras_apms_vis">GitHub </a>as a Shiny app.</p> <p> </p> <p><strong>Analysis of whole cell lysate</strong>: analysis_wholecelllysate.zip</p> <p>Contains the following script, as well as its output:</p> <ul> <li>01_analysis.R R script for loading the data and extracting/visualizing KRAS and effector abundances.</li> </ul> <p>The required data is deposited in the "data" sub-folder, taken directly from the linked PRoteomics IDEntification database (PRIDE) <a href="https://www.ebi.ac.uk/pride/archive/projects/PXD039404">entry</a>.</p>
Exploring Holocene temperature trends and a potential summer bias in simulations and reconstructions: TransEBM1.2 simulation data and analysis
<p>The TransEBM1.2 model code, transient climate simulation data of the last 26 ka and Python scripts to reproduce the analysis and Figures. </p>
FLOC App Prototype - Version 2 - Data Analysis
<p>This database gathers the feedback from 20 users regarding the second version of FLOC (Floating Companies App) - a prototype of an AR application for tablet and Oculus Quest 2. The data was collected anonymously using the 'thinking aloud' method.</p> <p>Floating Companies (FLOC) visualizes the design companies registered in Portugal in 2019, under the form of a set of floating spheres. This project was framed by the Project Design OBS. - For a Design Observatory in Portugal: Models, Instruments, Representation and Strategies, which collects and interprets data on the Portuguese design ecosystem.</p>
Causal Analysis of Google Code Jam Contest Data
<p>This archive is the replication package for the paper:</p> <p>Carlo A. Furia, Richard Torkar, Robert Feldt: <em>Towards Causal Analysis of Empirical Software Engineering Data — The Impact of Programming Languages on Coding Competitions</em>. <a href="https://arxiv.org/abs/2301.07524">arXiv:2301.07524</a>. January 2023.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.