Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
119
datasets available to search
ShareScore release 0.9.0
Dataset results
119 results for “Missing data”
Data from: How should genes and taxa be sampled for phylogenomic analyses with missing data? An empirical study in iguanian lizards
Targeted sequence capture is becoming a widespread tool for generating large phylogenomic data sets to address difficult phylogenetic problems. However, this methodology often generates data sets in which increasing the number of taxa and loci increases amounts of missing data. Thus, a fundamental (but still unresolved) question is whether sampling should be designed to maximize sampling of taxa or genes, or to minimize the inclusion of missing data cells. Here, we explore this question for an ancient, rapid radiation of lizards, the pleurodont iguanians. Pleurodonts include many well-known clades (e.g., anoles, basilisks, iguanas, and spiny lizards) but relationships among families have proven difficult to resolve strongly and consistently using traditional sequencing approaches. We generated up to 4921 ultraconserved elements with sampling strategies including 16, 29, and 44 taxa, from 1179 to approximately 2.4 million characters per matrix and approximately 30% to 60% total missing data. We then compared mean branch support for interfamilial relationships under these 15 different sampling strategies for both concatenated (maximum likelihood) and species tree (NJst) approaches (after showing that mean branch support appears to be related to accuracy). We found that both approaches had the highest support when including loci with up to 50% missing taxa (matrices with ∼40–55% missing data overall). Thus, our results show that simply excluding all missing data may be highly problematic as the primary guiding principle for the inclusion or exclusion of taxa and genes. The optimal strategy was somewhat different for each approach, a pattern that has not been shown previously. For concatenated analyses, branch support was maximized when including many taxa (44) but fewer characters (1.1 million). For species-tree analyses, branch support was maximized with minimal taxon sampling (16) but many loci (4789 of 4921). We also show that the choice of these sampling strategies can be critically important for phylogenomic analyses, since some strategies lead to demonstrably incorrect inferences (using the same method) that have strong statistical support. Our preferred estimate provides strong support for most interfamilial relationships in this important but phylogenetically challenging group.
Data from: Joined at the hip: linked characters and the problem of missing data in studies of disparity
Paleontological investigations into morphological diversity, or disparity, are often confronted with large amounts of missing data. We illustrate how missing discrete data effects disparity using a novel simulation for removing data based on parameters from published datasets that contain both extinct and extant taxa. We develop an algorithm that assesses the distribution of missing characters in extinct taxa, and simulates data loss by applying that distribution to extant taxa. We term this technique 'linkage'. We compare differences in disparity metrics and ordination spaces produced by linkage and random character removal. When we incorporated linkage among characters, disparity metrics declined and ordination spaces shrank at a slower rate with increasing missing data, indicating that correlations among characters govern the sensitivity of disparity analysis. We also present and test a new disparity method that uses the linkage algorithm to correct for the bias caused by missing data. We equalized proportions of missing data among time bins before calculating disparity, and found that estimates of disparity changed when missing data were taken into account. By removing the bias of missing data, we can gain new insights into the morphological evolution of organisms and highlight the detrimental effects of missing data on disparity analysis.
Data from: Estimation of individual growth trajectories when repeated measures are missing
Individuals in a population vary in their growth due to hidden and observed factors such as age, genetics, environment, disease, and carryover effects from past environments. Because size affects fitness, growth trajectories scale up to affect population dynamics. However, it can be difficult to estimate growth in data from wild populations with missing observations and observation error. Previous work has shown that linear mixed models (LMMs) underestimate hidden individual heterogeneity when over 25% of repeated measures are missing. Here we demonstrate a flexible and robust way to model growth trajectories. We show that state-space models (SSMs), fit using R package growmod, are far less biased than LMMs when fit to simulated datasets with missing repeated measures and observation error. This method is much faster than MCMC methods, allowing more models to be tested in a shorter time. For the scenarios we simulated, SSMs gave estimates with little bias when up to 87.5 % of repeated measures were missing. We use this method to quantify growth of Soay sheep, using data from a long-term mark-recapture study, and demonstrate that growth decreased with age, population density, weather conditions, and when individuals are reproductive. The method improves our ability to quantify how growth varies among individuals in response to their attributes and the environments they experience, with particular relevance for wild populations.
Data from: The case of the missing ancient fungal polyploids
Polyploidy—the increase in the number of whole chromosome sets—is an important evolutionary force in eukaryotes. Polyploidy is well recognized throughout the evolutionary history of plants and animals, where several ancient events have been hypothesized to be drivers of major evolutionary radiations. However, fungi provide a striking contrast: while numerous recent polyploids have been documented, ancient fungal polyploidy is virtually unknown. We present a survey of known fungal polyploids that confirms the absence of ancient fungal polyploidy events. Three hypotheses may explain this finding. First, ancient fungal polyploids are indeed rare, with unique aspects of fungal biology providing similar benefits without genome duplication. Second, fungal polyploids are not successful in the long term, leading to few extant species derived from ancient polyploidy events. Third, ancient fungal polyploids are difficult to detect, causing the real contribution of polyploidy to fungal evolution to be underappreciated. We consider each of these hypotheses in turn and propose that failure to detect ancient events is the most likely reason for the lack of observed ancient fungal polyploids. We examine whether existing data can provide evidence for previously unrecognized ancient fungal polyploidy events but discover that current resources are too limited. We contend that establishing whether unrecognized ancient fungal polyploidy events exist is important to ascertain whether polyploidy has played a key role in the evolution of the extensive complexity and diversity observed in fungi today and, thus, whether polyploidy is a driver of evolutionary diversifications across eukaryotes. Therefore, we conclude by suggesting ways to test the hypothesis that there are unrecognized polyploidy events in the deep evolutionary history of the fungi.
Data from: Missed opportunities for earlier diagnosis of HIV in patients that presented with advanced HIV disease: a retrospective cohort study
OBJECTIVE: To quantify and characterize missed opportunities for earlier HIV diagnosis in patients diagnosed with advanced HIV. DESIGN: A retrospective observational cohort study. SETTING: A central tertiary medical center in Israel. MEASURES: The proportion of patients with advanced HIV, the proportion of missed opportunities to diagnose them earlier, and the rate of clinical indicator diseases (CIDs) in those patients RESULTS: Between 2010-2015, 356 patients were diagnosed with HIV, 118 (33.4 %) were diagnosed late, 57 (16%) with advanced HIV disease. Old age (OR=1.45 [95% CI 1.16-1.74]) and being heterosexual (OR=2.65 [95% CI 1.21-5.78]) were significant risk factors for being diagnosed late. All patients with advanced disease had at least one CID that did not lead to an HIV test in the 5 years prior to AIDS diagnosis. The median time between CID and AIDS diagnosis was 24 month (IQR 10-30). 60% of CIDs were missed by a general practitioner and 40% by a specialist. CONCLUSIONS: Missed opportunities to early diagnosis of HIV occur both in primary and secondary care. Lack of national guidelines, lack of knowledge regarding CIDs and communication barriers with patients may contribute to HIV late diagnosis. 'Strengths and limitations of this study' This study shows for the first time rate and reasons for missed opportunities to diagnose HIV in a low prevalence country like Israel. This study may shed light on the reasons why primary care physicians or specialists are missing to diagnose HIV earlier. Nonexistence of clear national guidelines for HIV testing and ignoring HIV clinically indicator diseases are major reasons for missed diagnosis of HIV. This study was carried out in one center and may not reflect the picture in the all country; Also, the total number of patients is low and this may limit generability of the study
Data from: Unforeseen consequences of excluding missing data from next-generation sequences: simulation study of RAD sequences
There is a lack of consensus on how next-generation sequence data should be considered for phylogenetic and phylogeographic estimates, with some studies excluding loci with missing data, while others include them, even when sequences are missing from a large number of individuals. Here we use simulations, focusing specifically on RAD sequences, to highlight some of the unforeseen consequence of excluding missing data from next-generation sequencing. Specifically, we show that in addition to the obvious effects associated with reducing the amount of data used to make historical inferences, the decisions we make about missing data (such as the minimum number of individuals with a sequence for a locus to be included in the study) also impact the types of loci sampled for a study. In particular, as the tolerance for missing data becomes more stringent, the mutational spectrum represented in the sampled loci becomes truncated such that loci with the highest mutation rates are disproportionately excluded. This effect is exacerbated further by factors involved in the preparation of the genomic library (i.e., the use of reduced representation libraries, as well as the coverage) and the taxonomic diversity represented in the library (i.e., the level of divergence among the individuals). We demonstrate that the intuitive appeals about being conservative by removing loci may be misguided.
Data from: Using multiple imputation to estimate missing data in meta-regression
1. There is a growing need for scientific synthesis in ecology and evolution. In many cases, meta-analytic techniques can be used to complement such synthesis. However, missing data is a serious problem for any synthetic efforts and can compromise the integrity of meta-analyses in these and other disciplines. Currently, the prevalence of missing data in meta-analytic datasets in ecology and the efficacy of different remedies for this problem have not been adequately quantified. 2. We generated meta-analytic datasets based on literature reviews of experimental and observational data and found that missing data were prevalent in meta-analytic ecological datasets. We then tested the performance of complete case removal (a widely used method when data are missing) and multiple imputation (an alternative method for data recovery) and assessed model bias, precision, and multi-model rankings under a variety of simulated conditions using published meta-regression datasets. 3. We found that complete case removal led to biased and imprecise coefficient estimates and yielded poorly specified models. In contrast, multiple imputation provided unbiased parameter estimates with only a small loss in precision. The performance of multiple imputation, however, was dependent on the type of data missing. It performed best when missing values were weighting variables, but performance was mixed when missing values were predictor variables. Multiple imputation performed poorly when imputing raw data which was then used to calculate effect size and the weighting variable. 4. We conclude that complete case removal should not be used in meta-regression, and that multiple imputation has the potential to be an indispensable tool for meta-regression in ecology and evolution. However, we recommend that users assess the performance of multiple imputation by simulating missing data on a subset of their data before implementing it to recover actual missing data.
Data from: Bias and sensitivity in the placement of fossil taxa resulting from interpretations of missing data
The utility of fossils in evolutionary contexts is dependent on their accurate placement in phylogenetic frameworks, yet intrinsic and widespread missing data make this problematic. The complex taphonomic processes occurring during fossilization can make it difficult to distinguish absence from non-preservation, especially in the case of exceptionally preserved soft-tissue fossils: is a particular morphological character (e.g. appendage, tentacle or nerve) missing from a fossil because it was never there (phylogenetic absence), or just happened to not be preserved (taphonomic loss)? Missing data has not been tested in the context of interpretation of non-present anatomy nor in the context of directional shifts and biases in affinity. Here, complete taxa, both simulated and empirical, are subjected to data loss through the replacement of present entries (1s) with either missing (?s) or absent (0s) entries. Both cause taxa to drift down trees, from their original position, toward the root. Absolute thresholds at which downshift is significant are extremely low for introduced absences (2 entries replaced, 6 % of present characters). The opposite threshold in empirical fossil taxa is also found to be low; two absent entries replaced with presences causes fossil taxa to drift up trees. As such, only a few instances of non-preserved characters interpreted as absences will cause fossil organisms to be erroneously interpreted as more primitive than they were in life. This observed sensitivity to coding non-present morphology presents a problem for all evolutionary studies that attempt to use fossils to reconstruct rates of evolution or unlock sequences of morphological change. Stem-ward slippage, whereby fossilization processes cause organisms to appear artificially primitive, appears to be a ubiquitous and problematic phenomenon inherent to missing data, even when no decay biases exist. Absent characters therefore require explicit justification and taphonomic frameworks to support their interpretation.
Data from: Missing data estimation in morphometrics: how much is too much?
Fossil-based estimates of diversity and evolutionary dynamics mainly rely on the study of morphological variation. Unfortunately, organism remains are often altered by post-mortem taphonomic processes such as weathering or distortion. Such a loss of information often prevents quantitative multivariate description and statistically controlled comparisons of extinct species based on morphometric data. A common way to deal with missing data involves imputation methods that directly fill the missing cases with model estimates. Over the last several years, several empirically determined thresholds for the maximum acceptable proportion of missing values have been proposed in the literature, whereas other studies showed that this limit actually depends on several properties of the study dataset and of the selected imputation method, and is by no way generalizable. We evaluate the relative performances of seven multiple imputation techniques through a simulation-based analysis under three distinct patterns of missing data distribution. Overall, Fully Conditional Specification and Expectation-Maximization algorithms provide the best compromises between imputation accuracy and coverage probability. Multiple imputation (MI) techniques appear remarkably robust to the violation of basic assumptions such as the occurrence of taxonomically or anatomically biased patterns of missing data distribution, making differences in simulation results between the three patterns of missing data distribution much smaller than differences between the individual MI techniques. Based on these results, rather than proposing a new (set of) threshold value(s), we develop an approach combining the use of multiple imputations with procrustean superimposition of principal component analysis results, in order to directly visualize the effect of individual missing data imputation on an ordinated space. We provide an R function for users to implement the proposed procedure.
Data for The missing base molecules in atmospheric acid-base nucleation
<p>Measurement data from Beijing and simulation results for figures in the main text</p>
Data from: The phylogenetic trunk: maximal inclusion of taxa with missing data in an analysis of the Lepospondyli (Vertebrata, Tetrapoda)
The importance of fossils to phylogenetic reconstruction is well established. However, analyses of fossil data sets are confounded by problems related to the less complete nature of the specimens. Taxa that are incompletely known are problematic because of the uncertainty of their placement within a tree, leading to a proliferation of most parsimonious solutions and wild card behavior. Problematic taxa are commonly deleted based on a priori criteria of completeness. Paradoxically, a taxon's problematic behavior is tree dependent, and levels of completeness are not directly associated with problematic behavior. Exclusion of taxa based on completeness eliminates real character conflict and, by not allowing incomplete taxa to determine tree topology, the phylogenetic hypothesis is diminished. The phylogenetic trunk approach is proposed to allow optimization of taxonomic inclusion and tree stability. This method is used in an analysis of the Paleozoic Lepospondyli. A single most parsimonious tree, or trunk, is found after removal of one taxon identified as being problematic. The 38 trees found one additional step from this primary trunk are reduced to two by removal of one additional taxon. These trunks are compared to the trees found by excluding taxa with various degrees of completeness. Effects of incomplete taxa are explored in light of the trunk. Correlated characters associated with limblessness are discussed regarding the assumption of character independence, but inclusion of intermediate taxa is found to be the single best method for breaking down long branches.
Data from: Genetic analysis identifies the missing parchment of New Zealand's founding document, The Treaty of Waitangi.
Genetic analyses provide a powerful tool with which to identify the biological components of historical objects. Te Tiriti o Waitangi | The Treaty of Waitangi is New Zealand's founding document, intended to be a partnership between the indigenous Māori and the British Crown. Here we focus on an archived piece of blank parchment that has been proposed to be the missing portion of the lower parchment of the Waitangi Sheet of the Treaty. However, its physical dimensions and characteristics are not consistent with this hypothesis. We perform genetic analyses on the parchment membranes of the Treaty, plus the blank piece of parchment. We find that all three parchments were made from ewes and that the blank parchment is highly likely to be a portion cut from the lower membrane of the Waitangi Sheet because they share identical whole mitochondrial genomes, including an unusual heteroplasmic site. We suggest that the differences in size and characteristics between the two pieces of parchment may have resulted from the Treaty's exposure to water in the early 20th century and the subsequent repair work, light exposure during exhibition or the later conservation treatments in the 1970s and 80s. The blank piece of parchment will be valuable for comparison tests to study the effects of earlier treatments and to monitor the effects of long-term display on the Treaty.
Data from: How should genes and taxa be sampled for phylogenomic analyses with missing data? An empirical study in iguanian lizards
Open the record for dataset details and reuse information.
Data from: Biases with the Generalized Euclidean Distance in disparity analyses with high levels of missing data
Open the record for dataset details and reuse information.
Data from: Genetic analysis identifies the missing parchment of New Zealand’s founding document, The Treaty of Waitangi.
Open the record for dataset details and reuse information.
Data from: The phylogenetic trunk: maximal inclusion of taxa with missing data in an analysis of the Lepospondyli (Vertebrata, Tetrapoda)
Open the record for dataset details and reuse information.
Data from: Missed opportunities for earlier diagnosis of HIV in patients that presented with advanced HIV disease: a retrospective cohort study
Open the record for dataset details and reuse information.
Data from: The case of the missing ancient fungal polyploids
Open the record for dataset details and reuse information.
Data from: Bias and sensitivity in the placement of fossil taxa resulting from interpretations of missing data
Open the record for dataset details and reuse information.
Pleistocene persistence and expansion in tarantulas on the Colorado Plateau and the effects of missing data on phylogeographical inferences from RADseq
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.