Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
119
datasets available to search
ShareScore release 0.9.0
Dataset results
119 results for “Missing data”
Data from: Out-of-sample predictions from plant–insect food webs: robustness to missing and erroneous trophic interaction records
Open the record for dataset details and reuse information.
Data from: Regional heritability mapping method helps explain missing heritability of blood lipid traits in isolated populations
Open the record for dataset details and reuse information.
Data from: Effects of growth rate, size, and light availability on tree survival across life stages: a demographic analysis accounting for missing values and small sample sizes
Open the record for dataset details and reuse information.
Data from: HIV testing in a South African emergency department: a missed opportunity
Open the record for dataset details and reuse information.
Data from: Genetic differentiation of Alaska Chinook salmon: the missing link for migratory studies
Open the record for dataset details and reuse information.
Data from: Nest inheritance is the missing source of direct fitness in a primitively eusocial insect
Open the record for dataset details and reuse information.
Data from: A hierarchical Bayesian approach for handling missing classification data
Open the record for dataset details and reuse information.
Data from: Missed opportunities for HIV testing among patients newly presenting for HIV care at a Swiss university hospital: a retrospective analysis
Open the record for dataset details and reuse information.
Data from: Resolving the mesoscopic missing link: biophysical modeling of EEG from cortical columns in primates
Open the record for dataset details and reuse information.
Data from: Rewilded mammal assemblages reveal the missing ecological functions of granivores
Open the record for dataset details and reuse information.
Data from: Phenotypic selection favors missing trait combinations in coexisting annual plants
Open the record for dataset details and reuse information.
Repositories for taxonomic data: Where we are and what is missing
Open the record for dataset details and reuse information.
Data from: Misconceptions on missing data in RAD-seq phylogenetics with a deep-scale example from flowering plants
Open the record for dataset details and reuse information.
Data from: Correcting for missing and irregular data in home-range estimation
Open the record for dataset details and reuse information.
Data from: Missing the people for the trees: identifying coupled natural-human system feedbacks driving the ecology of Lyme disease
Open the record for dataset details and reuse information.
Missing data in sea turtle population monitoring: a Bayesian statistical framework accounting for incomplete sampling
Open the record for dataset details and reuse information.
Data from: RADcap: sequence capture of dual-digest RADseq libraries with identifiable duplicates and reduced missing data
Open the record for dataset details and reuse information.
Nonrandom missing data can bias PCA inference of population genetic structure
Open the record for dataset details and reuse information.
GECCO Industrial Challenge 2015 Dataset: A heating system dataset for the 'Recovering missing information in heating system operating data' competition at the Genetic and Evolutionary Computation Conference 2015, Madrid, Spain
<p>Dataset of the 'Industrial Challenge: Recovering missing information in heating system operating data' competition hosted at The Genetic and Evolutionary Computation Conference (GECCO) July 11th-15th 2015, Madrid, Spain</p> <p> </p> <p>The task of the competition was to recover (impute) missing information in heating system operation time series'.</p> <p> </p> <p>Included in zenodo: </p> <p>- dataset of heating system operational time series with missing values</p> <p>- additional material and descriptions provided for the competition</p> <p> </p> <p>The competition was organized by:</p> <p>M. Friese, A. Fischbach, C. Schlitt, T. Bartz-Beielstein (TH Köln)</p> <p> </p> <p>The dataset was provided by:</p> <p>Major German heating systems supplier (S. Moritz)</p> <p> </p> <p> </p> <p>Industrial Challenge: Recovering missing information in heating system operating data</p> <p> </p> <p>The Industrial Challenge will be held in the competition session at the Genetic and Evolutionary Computation Conference. It poses difficult real-world problems provided by industry partners from various fields. Highlights of the Industrial Challenge include interesting problem domains, real-world data and realistic quality measurement</p> <p>Overview</p> <p>In times of accelerating climate change and rising energy costs, increasing energy efficiency and reducing expenses becomes a high priority goal for businesses and private households alike. Modern heating systems record detailed operating data and report this data to a central system. Here, the operating data can be correlated and analyzed to detect potential optimization opportunities or anomalies like unusually high energy consumption. Due to various difficulties this data might be incomplete which makes accurate forecasting even harder.</p> <p>Goal of the GECCO 2015 Industrial Challenge is to develop capable procedures to recover missing information in heating system operating data. Adequate recovery of the missing data enables more accurate forecastings which allow for intelligent control of the heating systems, and therefore contributes to a positive energy balance and reduced expenses.</p> <p> </p> <p><strong>Submission deadline:</strong><br> June 22, 2015</p> <p><strong>Official Webpage:</strong><br> <a href="http://www.spotseven.de/gecco-challenge/gecco-challenge-2015/">www.spotseven.de/gecco-challenge/gecco-challenge-2015/</a></p> <p> </p>
Pleistocene persistence and expansion in tarantulas on the Colorado Plateau and the effects of missing data on phylogeographical inferences from RADseq
Few phylogeographical studies exist for taxa inhabiting the Colorado Plateau province. We combined mitochondrial and genomic data with species distribution modeling to test Pleistocene hypotheses for <i>Aphonopelma marxi</i>, a large tarantula endemic to the plateau region. Mitochondrial and genomic analyses revealed that the species comprises at least three main clades that diverged in the Pleistocene. A clade distributed along the Mogollon Rim appears to have persisted in place during the last glacial maximum, whereas the other two clades probably colonized the central and northeastern portion of the species' range from small refugial areas along river-carved canyons. Climate models support this hypothesis for the Mogollon Rim, but late glacial climate data appear too coarse to detect suitable areas in canyons. Locations of canyon refugia could not be inferred from genomic analyses due to missing data, encouraging us to explore the effect of missing loci in phylogeographical inferences using RADseq. In phylogenetic analyses, node support for major clades decreased with the addition of samples with significant amounts of missing data (more than 30%). Population genomic structure was greatly influenced by missing data, with the group membership of many taxa changing as samples with missing loci were added. Results from DAPC, a distance-based method, did not change as samples with significant amounts missing data were added. We conclude that the specific loci that are missing matters more than the number of missing loci, and that samples with missing data can still add information to RADseq-based analyses as long as results are interpreted cautiously.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.