Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
119
datasets available to search
ShareScore release 0.9.0
Dataset results
119 results for “Missing data”
Solar and interplanetary magnetic field data analyzed in "Optimal frequency-domain analysis for spacecraft time series: Introducing the missing-data multitaper power spectrum estimator"
<p>This dataset contains simultaneous measurements of the interplanetary magnetic field magnitude <B> and the sun's radio flux at 10.7 cm <F10.7>. <B> measurements come from a series of spacecraft located at the L1 point, while <F10.7> was measured by the ongoing monitoring program by Canada's Dominion Radio Astrophysical Observatory. Bartels rotation-averaged data were downloaded from NASA's OMNIWeb, https://omniweb.gsfc.nasa.gov/html/ow_data.html. The file contains other solar wind plasma parameters that were not used in the analysis.</p>
Missing data in the analysis of multilevel and dependent data (Examples)
<p>Example data sets and computer code for the book chapter titled "Missing Data in the Analysis of Multilevel and Dependent Data" submitted for publication in the second edition of "Dependent Data in Social Science Research" (Stemmler et al., 2015). This repository includes the computer code (".R") and the data sets from both example analyses (Examples 1 and 2). The data sets are available in two file formats (binary ".rda" for use in R; plain-text ".dat").</p> <p>The data sets contain simulated data from 23,376 (Example 1) and 23,072 (Example 2) individuals from 2,000 groups on four variables:</p> <p><code>ID</code> = group identifier (1-2000)<br> <code>x</code> = numeric (Level 1)<br> <code>y</code> = numeric (Level 1)<br> <code>w</code> = binary (Level 2)</p> <p>In all data sets, missing values are coded as "NA".</p>
Data from: Accounting for missing ticks: Use (or lack thereof) of hierarchical models in tick ecology studies
<p>Ixodid (hard) ticks play important ecosystem roles and have significant impacts on animal and human health via tick-borne diseases and physiological stress from parasitism. Tick occurrence, abundance, behavior, and key life-history traits are highly influenced by host availability, weather, microclimate, and landscape features. As such, changes in the environment can have profound impacts on ticks, their hosts, and the spread of diseases. Researchers interested in enumerating questing ticks attempt to integrate this heterogeneity by conducting replicate sampling bouts spread over the tick questing period as common field methods notoriously underestimate ticks. However, it is unclear how (or if) tick studies account for this heterogeneity in the modeling process. This step is critical as unaccounted variance in detection can lead to biased estimates of occurrence and abundance. We performed a descriptive review to evaluate the extent to which studies account for the detection process while modeling tick data. We also categorized the types of analyses that are commonly used to model tick data. We used hierarchical models (HMs) that account for imperfect detection to analyze simulated and empirical tick data, demonstrating that inference is muddled when detection probability is not accounted for in the modeling process. Our review indicates that only 5 of 412 (1%) papers explicitly accounted for imperfect detection while modeling ticks. By comparing HMs with the most common approaches used for modeling tick data (e.g., ANOVA), we show that population estimates are biased low for simulated and empirical data when using non-HMs, and that confounding occurs due to not explicitly modeling factors that influenced both detection and abundance. Our review and analysis of simulated and empirical data shows that it is important to account for our ability to detect ticks using field methods with imperfect detection. Not doing so leads to biased estimates of occurrence and abundance which could complicate our understanding of parasite-host relationships and the spread of tick-borne diseases. We highlight the resources available for learning HM approaches and applying them to analyzing tick data.</p>
Phylogenomics and biogeography of Torreya (Taxaceae) – Integrating data from three organelle genomes, morphology, and fossils and a practical method for reducing missing data from RAD-seq
<p><span>Restriction site-associated DNA sequencing (RAD-seq) enables obtaining thousands of genetic markers for phylogenomic studies. However, RAD-seq data are subject to allele dropout (ADO) due to polymorphisms at enzyme cutting sites. We developed a new pipeline, RADADOR, to mitigate the ADO in outgroups by recovering missing loci from previously published transcriptomes in our study of a gymnosperm genus </span><em>Torreya</em><span>. Using the supplemented RAD-seq data in combination with plastome and mitochondrial gene sequences, morphology, and fossil records, we reconstructed the phylogenetic and biogeographic histories of the genus and test hypotheses on diversity anomaly in eastern Asian-North American floristic disjunction. Our results showed that our pipeline recovered many loci missing from the outgroup, and the improved data yielded a more robust phylogeny for </span><em>Torreya</em><span>. Using the fossilized-birth-death model and divergence-extinction-cladogenesis method we resolved detailed biogeographic history of </span><em>Torreya</em><span> that suggested a Jurassic origin in the Laurasia and differential speciation and extinction among continents accounting for the modern diversity anomaly biased toward Eastern Asia (EA). The history also supported a vicariance origin of the modern </span><em>Torreya</em><span> from a widespread ancestor in EA and NA in the mid-Eocene, cross-Beringia exchange in the early Paleogene before the vicariant isolation, in contrast to the "Out of NA" pattern common to gymnosperms and in contrast to the "Out of EA" hypothesis previously proposed for the genus. Furthermore, we observed phylogenetic discordance between the nuclear and plastid phylogenies on </span><em>T. jackii</em><span>, suggesting differential lineage sorting of plastid genomes among </span><em>Torreya</em><span> species or plastid genome capture in </span><em>T. jackii</em><span>.</span></p>
Conflict over the eukaryote root resides in strong outliers, mosaics and missing data sensitivity of site-specific (CAT) mixture models
Abstract Phylogenetic reconstruction using concatenated loci ("phylogenomics" or "supermatrix phylogeny") is a powerful tool for solving evolutionary splits that are poorly resolved in single gene/protein trees (SGTs). However, recent phylogenomic attempts to resolve the eukaryote root have yielded conflicting results, along with claims of various artefacts hidden in the data. We have investigated these conflicts using two new methods for assessing phylogenetic conflict. ConJak uses whole marker (gene or protein) jackknifing to assess deviation from a central mean for each individual sequence, while ConWin uses a sliding window to screen for incongruent protein fragments (mosaics). Both methods allow selective masking of individual sequences or sequence fragments in order to minimize missing data, an important consideration for resolving deep splits with limited data. Analyses focused on a set of 76 eukaryotic proteins of bacterial-ancestry previously used in various combinations to assess the branching order among the three major divisions of eukaryotes: Amorphea (mainly animals, fungi and Amoebozoa), Diaphoretickes (most other well-known eukaryotes and nearly all algae) and Excavata, represented here by Discoba (Jakobida, Heterolobosea, and Euglenozoa). ConJak analyses found strong outliers to be concentrated in under-sampled lineages, while ConWin analyses of Discoba, the most under-sampled of the major lineages, detected potentially incongruent fragments scattered throughout. Phylogenetic analyses of the full data using an LG-gamma model support a Discoba sister scenario (neozoan-excavate root), which rises to 99-100% bootstrap support with data masked according to either protocol. However, analyses with two site-specific (CAT) mixture models yielded widely inconsistent results and a striking sensitivity to missing data. The neozoan-excavate root places Amorphea and Diaphoretickes as more closely related to each other than either is to Discoba, a fundamental relationship that should remain unaffected by additional taxa.
Imputation of missing land carbon sequestration data in the AR6 Scenarios Database
<p>This repository is linked to the following research paper:</p> <ul> <li>Prütz, R., Fuss, S., and Rogelj, J.: Imputation of missing land carbon sequestration data in the AR6 Scenarios Database, Earth Syst. Sci. Data, 2025. <a href="https://doi.org/10.5194/essd-17-221-2025">https://doi.org/10.5194/essd-17-221-2025</a> </li> </ul> <p>This repository includes: </p> <ul> <li>An imputation dataset for missing land carbon sequestation data of the AR6 Scenarios Database for global scenarios and R10 scenario variants</li> <li>Code to test, compare and visualize the performance of regression models to predict missing land removal data</li> <li>Code to compare and visualize available AR6 land removal data and existing AR6 data reanalyses</li> </ul> <p>The following two datasets are required to replicate the analysis:</p> <ul> <li>Byers, E., Krey, V., Kriegler, E., Riahi, K., Schaeffer, R., Kikstra, J., Lamboll, R., Nicholls, Z., Sandstad, M., Smith, C., van der Wijst, K., Al -Khourdajie, A., Lecocq, F., Portugal-Pereira, J., Saheb, Y., Stromman, A., Winkler, H., Auer, C., Brutschin, E., … van Vuuren, D. (2022). AR6 Scenarios Database [Data set]. In Climate Change 2022: Mitigation of Climate Change (1.1). Intergovernmental Panel on Climate Change. <a href="https://doi.org/10.5281/zenodo.7197970">https://doi.org/10.5281/zenodo.7197970</a></li> <li>Gidden, M., Gasser, T., Grassi, G., Forsell, N., Janssens, I., Lamb, W. F., Minx, J., Nicholls, Z., Steinhauser, J., & Riahi, K. (2023). Dataset for Gidden et.al. 2023 Updated AR6 Mitigation Benchmarks using National Emissions Inventories (Version v2) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.10158920">https://doi.org/10.5281/zenodo.10158920</a></li> </ul> <p>The variable imputation is based on the dataset by Byers et al. (2022). The dataset by Gidden et al. (2023) is used for variable comparison. </p>
Data from: A threat to loyalty: Fear of missing out (FOMO) leads to reluctance to repeat current experiences
<p>We investigate a popular but underresearched concept, the fear of missing out (FOMO), on desirable experiences of which an individual is aware, but in which they do not partake. Through laboratory and field studies, we establish FOMO's pervasiveness as a psychological phenomenon, present real-life contexts wherein FOMO may be experienced, and explore its behavioral consequences. Specifically, we show that FOMO poses a threat to loyalty by decreasing one's intentions to repeat a current experience and may decrease the valuation of the current experience.</p>
Figure 6 in Molecules, morphology, missing data and the phylogenetic position of a recently extinct madtom catfish (Actinopterygii: Ictaluridae)
Figure 6. Results of phylogenetic analyses of the morphology data set. A, majority rule consensus topology of four equally parsimonious trees. Numbers above branches represent non-parametric bootstrap support, numbers below are Bremer (1988) decay indices. Species groups and subgenera recognized by Taylor (1969) are indicated as follows: E = Elegans group, FB = Funebris group, FR = Furiosus group, H = Hildebrandi group, M = Miurus group, N = Noturus, R = Rabida, S = Schilbeodes. B, Bayesian consensus topology; asterisks above branches indicate posterior probabilities ± 0.95.
Figure 8 in Molecules, morphology, missing data and the phylogenetic position of a recently extinct madtom catfish (Actinopterygii: Ictaluridae)
Figure 8. Results of phylogenetic analyses of the combined (morphology + cyt b + RAG2) data set. A, single most parsimonious tree resulting from parsimony analyses including (offset box at left) and excluding Noturus trautmani. Numbers above branches indicate non-parametric bootstrap support, numbers below branches correspond to node numbers in Table 2 which outlines partitioned Bremer (1988) support for each node. Nodes recovered in the morphologyonly analysis are indicated by an open circle, molecular analysis by a closed circle and both analyses by an open square. Clades recognized in this study are indicated as follows: a = albater, e = elegans, fb = funebris, fr = furiosus, g = gyrinus, h = hildebrandi, r = rabida. B, Bayesian consensus topology from analyses including (offset box at left) and excluding N. trautmani. Asterisks above branches indicate posterior probabilities ± 0.95. Nodes recovered with posterior probabilities ± 0.95 in the morphology-only analysis are indicated by an open circle, molecular analysis by a closed circle and both analyses by an open square.
Figure 7 in Molecules, morphology, missing data and the phylogenetic position of a recently extinct madtom catfish (Actinopterygii: Ictaluridae)
Figure 7. Results of phylogenetic analyses of the molecular (cyt b + RAG2) data set. A, single most parsimonious tree resulting from parsimony analysis. Numbers above branches indicate non-parametric bootstrap support, numbers below are Bremer (1988) decay indices. B, Bayesian consensus topology; asterisks above branches indicate posterior probabilities ± 0.95.
Figure 5. A, B in Molecules, morphology, missing data and the phylogenetic position of a recently extinct madtom catfish (Actinopterygii: Ictaluridae)
Figure 5. A, B, tripus and os suspensor, ventral view, anterior at top. A, Noturus maydeni, JFBM 40815; B, N. funebris, JFBM 43002. C–E, caudal vertebrae, lateral view, anterior at left. C, N. funebris, JFBM 43002; D, N. gladiator, JFBM 40771; E, N. elegans, JFBM 41047. c1 = vertebral centrum 1; CC1 = complex centrum – Weberian apparatus; osus = os suspensor; tr = tripus. Numbers refer to characters and states listed in Appendix S2. Scale bar = 1 mm.
Figure 4 in Molecules, morphology, missing data and the phylogenetic position of a recently extinct madtom catfish (Actinopterygii: Ictaluridae)
Figure 4. Weberian complex and associated dorsal fin structures, dorsal view, anterior at bottom. A, Noturus munitus, JFBM 43107; B, N. maydeni, JFBM 39170; C, N. gilberti, JFBM 42496; D, N. funebris, JFBM 43002. dsp1 = dorsal fin spine 1; dsp2 = dorsal fin spine 2; np1 = nuchal plate 1; np2 = nuchal plate 2; ns4 = neural spine 4; pr1 = proximal radial 1; tp4 = transverse process fourth vertebra; tp5 = transverse process fifth vertebra. Numbers refer to characters and states listed in Appendix S2. Scale bar = 1 mm.
Figure 2 in Molecules, morphology, missing data and the phylogenetic position of a recently extinct madtom catfish (Actinopterygii: Ictaluridae)
Figure 2. Suspensorium, lateral view, anterior at left. A, Noturus elegans, JFBM 41047; B, N. gladiator, JFBM 40771; C, N. insignis, JFBM 43239; D, N. funebris, JFBM 41501. enp = endopterygoid; h = hyomandibula; iop = interopercle; mpt = metapterygoid; op = opercle; pop = preopercle; q = quadrate. Numbers refer to characters and states listed in Appendix S2. Scale bar = 1 mm.
Figure 3. A, B in Molecules, morphology, missing data and the phylogenetic position of a recently extinct madtom catfish (Actinopterygii: Ictaluridae)
Figure 3. A, B, pectoral girdle, ventral view, anterior at top. A, Noturus munitus, JFBM 43107; B, N. funebris, JFBM 43002. C–E, posttemporo-supracleithrum, posterior view, anterior into page. C, N. munitus, JFBM 43107; D, N. eleutherus, JFBM 43055; E, N. funebris, JFBM 43002. cl = cleithrum; pr = pectoral radial; ps = pectoral fin spine; sco = scapulo-coracoid. Numbers refer to characters and states listed in Appendix S2. Cartilage shown by grey shading. Scale bar = 1 mm.
Figure 1 in Molecules, morphology, missing data and the phylogenetic position of a recently extinct madtom catfish (Actinopterygii: Ictaluridae)
Figure 1. Neurocranium and associated laterosensory canals, dorsal view, anterior at top. A, Noturus albater, JFBM 39350; B, N. eleutherus, JFBM 43055; C, N. insignis, JFBM 41247; D, N. funebris, JFBM 41501. ep = epioccipital; f = frontal; le = lateral ethmoid; me = mesethmoid; n = nasal; pa-so = parieto-supraoccipital; pto = pterotic; spo = sphenotic. Numbers refer to characters and states listed in Appendix S2. Cartilage shown by grey shading. Scale bar = 1 mm.
Uptime data for 'Coherent fiber links operated for years: effect of missing data'
<p>Data of the total uptime of the comparison between two clocks, used in figure 13, in the article "Coherent fiber links operated for years: effect of missing data", doi: 10.1088/1681-7575/ac938e,.</p>
Missing data in amortized simulation-based neural posterior estimation
<p>Supplementary Data to the publication "Missing data in amortized simulation-based neural posterior estimation", Wang et al. 2023.</p>
Data set for figure 2-4 from publication "Missed Evaporation from Atmospherically Relevant Inorganic Mixtures Confounds Experimental Aerosol Studies",
<p>Data set for figure 2-4 from publication "Missed Evaporation from Atmospherically Relevant Inorganic Mixtures Confounds Experimental Aerosol Studies".</p>
Phylogenomics and biogeography of Torreya (Taxaceae) – Integrating data from three organelle genomes, morphology, and fossils and a practical method for reducing missing data from RAD-seq
Open the record for dataset details and reuse information.
Data from: A threat to loyalty: Fear of missing out (FOMO) leads to reluctance to repeat current experiences
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.