Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Figure data sets for the paper "Non-classical correlations over 1250 modes between telecom photons and 979-nm photons stored 171Yb3+:Y2SiO5"
<p>Processed datasets corresponding to the Figures published in the article.</p>
Data Set_Curantz et al
<p>Data Set for Curantz et al, PLOS Biology, "Cell shape anisotropy contributes to self-organized feather pattern fidelity in birds"</p>
Evaluation data-set for APICarv
<p>Data-set to replicate results reported in the APICarv submission</p>
Figure data sets for the paper "Coherent optical-microwave interface for manipulation of low-field electronic clock transitions in 171Yb3+:Y2SiO5"
<p>Figure data sets.</p>
Data set to Remote sensing-supported mapping of the activity of a subterranean landscape engineer across an afro-alpine ecosystem
<p>This data set is part of the article Wraase et al. (2022): Remote sensing -supported mapping of the activity of a subterranean landscape engineer across an afro-alpine ecosystem. Remote sensing in Ecology and Conservation. (https://doi.org/10.1002/RSE2.303)</p> <p>The repository contains a Readme file ("readme.txt") and two additional folders labeled: “input data” and “script”.<br> <br> The first folder contains 13 data files further divided into three subfolders “cca_analysis”, “main_modelling_prc_texture_idx” and “vectors”. Data formats are .csv format for all tables, .rds files for model objects from R and .shp format for all vector data.</p> <p>The second folder contains all 31 R-scripts necessary to do the analysis, as described in the article. Additionally, the folder is further categorized into five subfolders equivalent to the main analysis operations: “cca_analysis”, “landsat_temp_modelling”, “main_modelling_prc”, “maxent” and “texture_idx”.</p>
Data set
<p>Data set from <strong>Conventional and sustainable agricultural management effect of contrasting crops on carbon and water fluxes in a Mexican semi-arid region</strong></p>
Data set: Native forest conversion alters soil macroinvertabrate diversity and soil quality in tropical mountain landscapes of northern Ecuador
<p>Data set an related material of the article entitled: Native forest conversion alters soil macroinvertabrate diversity and soil quality in tropical mountain landscapes of northern Ecuador</p>
Data set containing the energy landscapes for GPO and GPP tropocollagen models under pulling forces
<p>Energy landscapes (databases of minima and transition states) for GPO and GPP repeat collagen models under constant pulling forces as explored with OPTIM and PATHSAMPLE with an AMBER force field.</p> <p>The systems are seven GPO or GPP per chain capped with ACE and NME.</p> <p> The forces applied are 0 pN (F0), 10 pN (F1), 50 pN (F2), 100 pN (F3), 250 pN (F4), 500 pN (F5) and 750 pN (F6).</p> <p>The folders contains numerous analysis scripts and graphs. Most of these assume python with numpy and pandas, as well as cpptraj from AMBERTools.</p>
Data set for energy landscapes for interrupted sequences
<p>Databases of minima and transition states produced with PATHSAMPLE for interrupted tropocollagen segments.</p> <p>The two databases are for a GPO and a GPP repeat model (7 repeats per chain, with ACE and NME caps).</p> <p>One strand contains a deletion of the form GPO-GP-GPO and GPP-GP-GPP.</p>
Boundary Data set : Land Cover Classification with Gaussian Processes using spatio-spectro-temporal features
<p>Boundary data set used to evaluate the continuity predictions for different models in boundary zones. It is composed of labelled and unlabelled pixels for a boundary size of 100m and 200m.</p> <p>For further details see section VI-A-1 of the pre-print article "Land Cover Classification with Gaussian Processes using spatio-spectro-temporal features ". This article is available <a href="https://hal.archives-ouvertes.fr/hal-03781332">here</a>.</p> <p>To compute the predictions in the boundary zones with different models (GP, RF, MLP, LTAE), the code is available in the <a href="https://gitlab.cesbio.omp.eu/belletv/land_cover_southfrance_gp">open source repository</a>.</p>
Scalable mixed model approaches for set-based association studies on large-scale categorical data analysis and its application to 450k exome sequencing data in UK Biobank
<p>The ongoing release of large-scale sequencing data in the UK Biobank allows for identifying associations between rare variants and complex traits. SAIGE-GENE+ is a valid approach to conducting set-based association tests for quantitative and binary traits. However, for ordinal categorical phenotypes, applying SAIGE-GENE+ with treating the trait as quantitative or binarizing the trait can cause inflated type I error rates or power loss. In this study, we propose a novel method for rare-variant association tests, POLMM-GENE, in which a proportional odds logistic mixed model was used to characterize ordinal categorical phenotypes while adjusting for sample relatedness. POLMM-GENE fully utilizes the categorical nature of phenotypes and thus can well control type I error rates while remaining powerful. In the analyses of UK Biobank 450k whole exome-sequencing data for 5 ordinal categorical traits, POLMM-GENE identified 54 gene-phenotype associations.</p>
Blinded Predictions and Post-hoc Analysis of the Second Solubility Challenge Data: Exploring Training Data and Feature Set Selection for Machine and Deep Learning Models
<p>Training and test datasets and scripts for training models.</p>
Data set for "The effect of pulling and twisting forces on chameleon sequence peptides"
<p>Data set for the publication "The effect of pulling and twisting forces on chameleon sequence peptides"<br> <br> The data contains all energy landscapes explored, structures extracted and analysed, and the necessary utility scripts, written by James Meadows. The README details which data is in each directory. The image directory contains the disconnectivity graphs associated with this data.</p>
Data set responden pengguna Go-Pay pada Go-Jek
<p>Responden hasil survey persepsi layanan Go-Pay pada aplikasi Go-Jek</p>
Analysed data set for rapidly intensifying tropical cyclones in the western North Pacific
<p>These are the extracted data set from (1) GPM IMERG; (2) Himawati-8 brightness temperature in infrared bands; (3) Himawari-8 cloud properties for the rapidly intensifying tropical cyclones in the western North Pacific.</p>
Coupling Silicon Lithography with Metal Casting-Data Set
<p>Dataset corresponding to the article Coupling Silicon Lithography with Metal Casting <a href="https://doi.org/10.1016/j.apmt.2022.101647">https://doi.org/10.1016/j.apmt.2022.101647</a> </p>
Data set for Ligand additivity relationships enable efficient exploration of transition metal chemical space
<p>Dataset of transition metal complexes curated in pickle files and comma delimited format, scripts for CSD curation, and computed DFT properties for associated manuscript.</p>
Data set for the manuscript "Uncertainty quantification and physics-informed forecasting for improved urban flood modeling"
<p>This is a data set for the manuscript "Uncertainty quantification and physics-informed forecasting for improved urban flood modeling."</p>
S1 Minimal Data Set for Selective Autonomic Stimulation of the AV Node Fat Pad to Control Rapid Post-Operative Atrial Arrhythmias
<p>Raw data files for figures on paper.</p>
Compound data sets for differential evolution calculation and a descriptor list
<p>Nine compound activity classes were assembled from ChEMBL version 22 as described in reference 1. In addition, the descriptor list reported in reference 1 is provided.</p> <p>Reference:</p> <p>1 Miyao T, Funatsu K and Bajorath J. Exploring differential evolution for inverse QSAR analysis. F1000Research 2017, <strong>6</strong>(Chemical Inf Sci): 1285</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.