Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
119
datasets available to search
ShareScore release 0.9.0
Dataset results
119 results for “Missing data”
Conflict over the eukaryote root resides in strong outliers, mosaics and missing data sensitivity of site-specific (CAT) mixture models
Open the record for dataset details and reuse information.
Data from: What need for speed? Lizards from islands missing predators sprint slower
Open the record for dataset details and reuse information.
Data from: Accounting for missing ticks: Use (or lack thereof) of hierarchical models in tick ecology studies
Open the record for dataset details and reuse information.
Data from: Uneven missing data skew phylogenomic relationships within the lories and lorikeets
<p>Inlcuded is the supplementary data for Smith, B. T., Mauck, W. M., Benz, B., & Andersen, M. J. (2018). Uneven missing data skews phylogenomic relationships within the lories and lorikeets. <em>BioRxiv</em>, 398297. </p> <p>The resolution of the Tree of Life has accelerated with advances in DNA sequencing technology. To achieve dense taxon sampling, it is often necessary to obtain DNA from historical museum specimens to supplement modern genetic samples. However, DNA from historical material is generally degraded, which presents various challenges. In this study, we evaluated how the coverage at variant sites and missing data among historical and modern samples impacts phylogenomic inference. We explored these patterns in the brush-tongued parrots (lories and lorikeets) of Australasia by sampling ultraconserved elements in 105 taxa. Trees estimated with low coverage characters had several clades where relationships appeared to be influenced by whether the sample came from historical or modern specimens, which were not observed when more stringent filtering was applied. To assess if the topologies were affected by missing data, we performed an outlier analysis of sites and loci, and a data reduction approach where we excluded sites based on data completeness. Depending on the outlier test, 0.15% of total sites or 38% of loci were driving the topological differences among trees, and at these sites, historical samples had 10.9x more missing data than modern ones. In contrast, 70% data completeness was necessary to avoid spurious relationships. Predictive modeling found that outlier analysis scores were correlated with parsimony informative sites in the clades whose topologies changed the most by filtering. After accounting for biased loci and understanding the stability of relationships, we inferred a more robust phylogenetic hypothesis for lories and lorikeets.</p>
Data from: Quantifying the impacts of management and herbicide resistance on regional plant population dynamics in the face of missing data
<p>A key challenge in the management of populations is to quantify the impact of interven-tions in the face of environmental and phenotypic variability. However, accurate estima-tion of the effects of management and environment, in large-scale ecological research is often limited by the expense of data collection, the inherent trade-off between quality and quantity, and missing data.</p> <p>In this paper we develop a novel modelling framework, and demographically informed imputation scheme, to comprehensively account for the uncertainty generated by miss-ing population, management, and herbicide resistance data. Using this framework and a large dataset (178 sites over 3 years) on the densities of a destructive arable weed (Alo-pecurus myosuroides) we investigate the effects of environment, management, and evolved herbicide resistance, on weed population dynamics.</p> <p>In this study we quantify the marginal effects of a suite of common management prac-tices, including cropping, cultivation, and herbicide pressure, and evolved herbicide re-sistance, on weed population dynamics.</p> <p>Using this framework, we provide the first empirically backed demonstration that herbi-cide resistance is a key driver of population dynamics in arable weeds at regional scales. Whilst cultivation type had minimal impact on weed density, crop rotation, and earlier cultivation and drill dates consistently reduced infestation severity.</p> <p>Synthesis and applications: As we demonstrate that high herbicide resistance levels can produce extremely severe weed infestations, monitoring of herbicide resistance is a pri-ority for famers across western Europe. Furthermore, developing non chemical control methods is essential to control current weed populations, and prevent further resistance evolution. We recommend that planning interventions that center on crop rotation and incorporate spring sewing and cultivation to provide the best reductions in weed densi-ties. More generally, by directly accounting for missing data our framework permits the analysis of management practices with data that would otherwise be severely compro-mised.</p>
Data for The genetics monopolistic industry, as expected, is missing information two inches beyond their nose: in this case, transcripts
<p><strong>De novo transcriptome assembly is one of the many fundamental pieces of new research in genomics. It is, for example, the preferred method for studying non-model organisms, since it is easier and cheaper than building a genome, and reference methods are not possible without an existing genome. The transcriptomes of these organisms can thus reveal novel proteins and their isoforms that are implicated in such unique biological phenomena. This technique is also useful in cancer research as it makes possible to detect potentially significant chimeric transcripts in cancer and normal somatic tissues.</strong></p> <p><strong>Given that the genetics industry is organized in the form of a monopoly controlled by hidden lobbies who also control Academia, all the available software for de novo transcriptome assembly is being developed by academic researchers under the open-source paradigm. We report here, that as anyone could have very easily deduced from past experiences in other industries, this unethical form of organization in the industry is resulting in incompetence whose effects include missing a significant portion of the available information that could be obtained from some genomics studies. In this case, missing transcripts in transcriptome studies. We won't deep in on the consequences, but these could include overpricing, over costs, and failing to achieve the goals of some studies.</strong></p>
Data and example reconstruction code for "Investigating the missing wedge problem in small-angle x-ray scattering tensor tomography across real and reciprocal space"
<p>Data sets and example code for the paper "Investigating the missing wedge problem in small-angle x-ray scattering tensor tomography across real and reciprocal space".</p> <p> </p> <p>Requires the software Mumott, see: https://doi.org/10.5281/zenodo.7798530</p>
Data for "How to use scale invariant properties of imperviousness in urban areas to handle missing data ?"
<p>The data set corresponds the data presented in the data paper : “How to use scale invariant properties of imperviousness in urban areas to handle missing data ?“ which has been submitted to Water Resources Research ” (https://agupubs.onlinelibrary.wiley.com/journal/19447973).</p> <p>It corresponds to :</p> <p>- the rainfall data collected on 2019-06-02 with 5 min and 30 s time steps by a disdrometer installed on the roof of Ecole des Ponts ParisTech building.</p> <p>- land use distribution for the Jouy-en-Josas catchment (1 = forest, 2= road, 3=Grass, 4=building, 5=Gully, 6=missing data), with pixel size of 10 m and 2 m.</p> <p>More details can be found in the file and in the paper.</p>
Males miss and females forgo: auditory masking from vessel noise impairs foraging efficiency and success in killer whales - CALIBRATED MOVEMENT DATA AND VARIABLES SUPPORTING ANALYSES
<p><strong>Description of the data and file structure<br></strong>This record contains data from animal-borne biologging instruments (Dtags) temporarily affixed to fish-eating killer whales, supporting the analyses presented in the following article:</p> <p> Tennessen. J.B., Holt, M.M., Wright, B.M., Hanson, M.B., Emmons, C.K., Giles, D.A., Hogan, J.T., Thornton, S.J., Deecke, V.B. 2024. Males miss and females forgo: auditory masking from vessel noise impairs foraging efficiency and success in killer whales. <em>Global Change Biology</em>.<strong> </strong>In press.</p> <p>The data include the following: (1) calibrated movement data from analyzed Dtag deployments, and (2) a spreadsheet containing the variables included in the fully-saturated and final models listed in Table 2 in the article cited above. All methodological details necessary to contextualize analysis procedures are provided in the methods section of the article. The following data files are available under separate DOIs: 10.5281/zenodo.13333019 - all 2009 & 2010 audio data; 10.5281/zenodo.13328931 - all 2011 & 2014 audio data.</p> <p>These data are provided by NOAA Fisheries' Northwest Fisheries Science Center, and Fisheries and Oceans Canada, to support reproducibility of all statistical analyses presented in the article. Please cite your usage of our data. For inquiries about data use, or for general questions, please contact Dr. Jennifer B. Tennessen, at jtenness@uw.edu.</p> <p> </p> <p><strong>Description of the movement data files<br></strong>The movement files have been calibrated from the raw data and are ready to use. The files contain the .mat extension, and need to be opened using Matlab and the tagtools tool kit available at https://github.com/animaltags . Tutorials for working with the toolkit are available at animaltags.org . These files contain several vector and matrix variables. We define those used in our analyses below. For questions about how to work with these files, please contact Dr. Jennifer B. Tennessen, at jtenness@uw.edu.</p> <p>Aw: calibrated triaxial accelerometer data (converted from tag frame to whale frame)</p> <p>fs: sample rate (50 Hz)</p> <p>head: animal's circular heading (rotation about the dorsal-ventral axis, in radians)</p> <p>Mw: calibrated triaxial magnetometer data (converted from tag frame to whale frame)</p> <p>p: depth (in meters)</p> <p>pitch: animal's pitch (rotation about the left-right axis, in radians)</p> <p>roll: animal's roll (rotation about the anterior-posterior axis, in radians)</p> <p>tempr: temperature recorded on tag (in Celsius)</p> <p>TT: time cues for the start and end of every analyzed dive within a deployment. This matrix contains 6 columns:<br>-col 1: start cue (in sec)<br>-col 2: end cue (in sec)<br>-col 3: maximum depth of dive (m)<br>-col 4: time cue at max depth (in sec)<br>-col 5: mean depth (m)<br>-col 6: mean compression</p> <p> </p> <p><strong>Description of the analyzed variables<br></strong>The data are provided column-wise in a spreadsheet, whereby each column contains one of several variables used to build the corresponding models listed in Table 2 in the above article. Model details are provided in the above article, including the statistical packages needed to run the models. </p> <p><em>The following is a list of variable names (column headers) and their corresponding definitions:<br></em><strong>bzsounds:</strong> binary presence (1)/absence (0) of buzz bouts within a dive. Buzzing is defined as the occurrence of echolocation clicks with an inter-click interval < 11 ms<br><strong>code:</strong> categorical identifier of the numerical week of year in which the tag was deployed (e.g., week 33 of 2009 is different than week 33 of 2011)<br><strong>deployment:</strong> the event whereby a tag was affixed to an individual killer whale and data were collected via tag sensors; each deployment was assigned a unique deployment ID, consisting of the first letter of the Genus and species names (“oo” for Orcinus orca), followed by two digits corresponding to the year (“09” = 2009), followed by the Julian day of the year (e.g. “234”), followed by a letter indicating the deployment order of the day. NRKW deployments were assigned a through l, and SRKW deployments were assigned m through z (e.g. “a” = first deployment of the day for NRKW, “m” = first deployment of the day for SRKW)<br><strong>durwho: </strong>duration of a whole dive, in seconds. Dives were defined as all departures from the surface, to at least 1 m or deeper, followed by a return to within 0.5 m of the surface<br><strong>divenum: c</strong>hronological identifier for dive position within a deployment (e.g., for the 10<sup>th</sup> dive within a deployment, divenum = 10)<br><strong>kindet: </strong>binary presence (1)/absence (0) of a prey capture event within a dive. Prey capture was informed by the occurrence of stereotyped movement signatures in sensor data indicative of prey capture, following an established method validated with visual and acoustic confirmation of predation events. Prey capture is defined as the occurrence of three movement variables indicative of prey capture (peak jerk, roll and heading variance) each exceeding pre-determined thresholds (see Tennessen et al. 2019b in above article for details)<br><strong>maxdep:</strong> maximum depth of a dive, in meters<br><strong>NLmax: </strong>the maximum noise level received during a dive, measured as the root-mean-square sound pressure level (dB re 1 mPa) within one second bins over the 15-45 kHz frequency band<br><strong>population:</strong> population to which the tagged whale belongs (NRKW = Northern Resident killer whale; SRKW = Southern Resident killer whale)<br><strong>sex:</strong> sex of tagged whale (F = female, M = male, NA = unknown)<br><strong>sc:</strong> binary presence (1)/absence (0) of slow-click sounds within a dive. Slow-clicking is defined as the occurrence of echolocation clicks with an inter-click interval >100 ms<br><strong>tagID:</strong> identifier for the individual tag used for each deployment<br><strong>year:</strong> year of deployment</p>
Data for: Approaches for handling missing values and their impacts on biological inferences: a molecular rate case study
<p>These data files are associated with the manuscript entitled: "Approaches for handling missing values and their impacts on biological inferences: a molecular rate case study" by Jacqueline A. May, Zeny Feng, and Sarah J. Adamowicz. This project entailed an evaluation of missing data handling approach on inferences using a molecular evolution case study. A target mixed-type dataset was first imputed using a real data-driven strategy for imputation method selection. Both trait-only (non-phylogenetic) and phylogenetic imputation methods were used to impute the dataset. Phylogenetic generalized least squares (PGLS) analyses were then applied to the complete-case and imputed datasets, specifying the traits as predictors and molecular evolutionary rates as the response variable. Those traits that associate significantly with molecular rates were identified and PGLS models compared to determine how the approach for handling missing data impacts biological inferences and conclusions.</p> <p>The files stored here are the trees built for phylogenetic imputation (RAxML tree and ultrametric tree versions) and the corresponding GenBank accession numbers.</p>
Data from: RADseq dataset with 90% missing data fully resolves recent radiation of Petalidium (Acanthaceae) in the ultra-arid deserts of Namibia
Open the record for dataset details and reuse information.
Data from: Character evolution and missing (morphological) data across Asteridae
Open the record for dataset details and reuse information.
Data from: Quantifying the impacts of management and herbicide resistance on regional plant population dynamics in the face of missing data
Open the record for dataset details and reuse information.
Data from: Uneven missing data skew phylogenomic relationships within the lories and lorikeets
Open the record for dataset details and reuse information.
Data from: The missing link in grassland restoration: arbuscular mycorrhizal fungi inoculation increases plant diversity and accelerates succession
Open the record for dataset details and reuse information.
Data from: Estimated missing interactions change the structure and alter species roles in one of the world’s largest seed-dispersal networks
Open the record for dataset details and reuse information.
FIGURE5. Maximum likelihood tree based on the Kimura 2-parameter model of the COI sequences from the Siphamia species with P. kauderni as the outgroup. Tree shown here has the highest log likelihood following 10 000 replications. The percentage of trees in which the associated taxa clustered together is shown next to the branches, branch lengths are measured in the number of substitutions per site and all positions containing gaps and missing data have been eliminated. in Redescription and distributional range extension of the Speckled Siphonfish, Siphamia guttulata (Pisces: Apogonidae)
FIGURE5. Maximum likelihood tree based on the Kimura 2-parameter model of the COI sequences from the Siphamia species with P. kauderni as the outgroup. Tree shown here has the highest log likelihood following 10 000 replications. The percentage of trees in which the associated taxa clustered together is shown next to the branches, branch lengths are measured in the number of substitutions per site and all positions containing gaps and missing data have been eliminated.
Dataset for "Recursive Input and State Estimation: A General Framework for Learning from Time Series with Missing Data"
<p>Dataset for "Recursive Input and State Estimation: A General Framework for Learning from Time Series with Missing Data"</p> <p> </p> <p>Missing values in the blood glucose datasets are represented with -2.</p>
Data from: HIV testing in a South African emergency department: a missed opportunity
Background: South Africa (SA) has the highest burden of HIV infection in the world. Even though HIV testing is mandated in all hospital-based facilities in SA, it is rarely implemented in the emergency department (ED). EDs are episodic care centers that provide care to large volumes of undifferentiated patients for short periods of time, and may treat undiagnosed HIV-infected patients not captured through standard clinic based screenings. Methods and Findings: In this prospective study, we implemented the National South African HIV testing guidelines, including 24-hours a day counselor initiated HIV Counseling and Testing (HCT), at Frere Hospital in the Eastern Cape from September 1 to November 30, 2016. All patients that presented for care in the ED during the study period, and who were clinically stable and fully conscious, were eligible to be approached by HCT staff to receive a rapid point-of-care HIV test. A total of 2355 of the 9583 (24.6%) patients that presented to the ED for care during the study period were approached by the HCT staff, of whom 1852 were enrolled in the study. There was high uptake of HIV testing (78.6%) among a predominantly male (58%) patient group that mostly presented with traumatic injuries (70.8%). Four hundred (21.6%) of the enrolled patients were HIV positive, including 115 (6.2%) with previously undiagnosed HIV infection. The overall prevalence of HIV infection in females (29.8%) was twice that compared to males (15.4%), despite the burden of undiagnosed infection being similar (6.0% for all females and 6.4% for all males). Conclusions: Overall there was a high HIV testing uptake that revealed a significant burden of undiagnosed HIV infection in this setting, especially in young males. Future research should focus on testing optimization, and efficient linkage to care from the ED for antiretroviral therapy initiation.
Data from: Effects of growth rate, size, and light availability on tree survival across life stages: a demographic analysis accounting for missing values and small sample sizes
Background: Plant survival is a key factor in forest dynamics and survival probabilities often vary across life stages. Studies specifically aimed at assessing tree survival are unusual and so data initially designed for other purposes often need to be used; such data are more likely to contain errors than data collected for this specific purpose. Results: We investigate the survival rates of ten tree species in a dataset designed to monitor growth rates. As some individuals were not included in the census at some time points we use capture-mark-recapture methods both to allow us to account for missing individuals, and to estimate relocation probabilities. Growth rates, size, and light availability were included as covariates in the model predicting survival rates. The study demonstrates that tree mortality is best described as constant between years and size-dependent at early life stages and size independent at later life stages for most species of UK hardwood. We have demonstrated that even with a twenty-year dataset it is possible to discern variability both between individuals and between species. Conclusions: Our work illustrates the potential utility of the method applied here for calculating plant population dynamics parameters in time replicated datasets with small sample sizes and missing individuals without any loss of sample size, and including explanatory covariates.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.