Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19
datasets available to search
ShareScore release 0.9.0
Dataset results
19 results for “Systematic error”
Data supporting 'Empirical correction of systematic orthorectification error in Sentinel-2 velocity fields for Greenlandic outlet glaciers'
<p><strong>Note: An updated dataset covering the majority of Greenland's marine-terminating glaciers is available as part of the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) project through the National Snow and Ice Data Center (NSIDC) at <a href="https://doi.org/10.5067/B28FM2QVVYWY">https://doi.org/10.5067/B28FM2QVVYWY</a>. </strong></p> <p>Data supporting the paper:</p> <blockquote> <p>Chudley, T. R., Howat, I. M., Yadav, B. N., & Noh, M. J. (2022). Empirical correction of systematic orthorectification error in Sentinel-2 velocity fields for Greenlandic outlet glaciers. <em>The Cryosphere. </em>16, 2629–2642, https://doi.org/10.5194/tc-16-2629-2022</p> </blockquote> <p>Dataset consists of four netCDF files containing stacked Sentinel-2 velocity data of four Greenlandic outlet glaciers (Helheim Glacier, Jakobshavn Isbræ, Store Glacier, and Kangerlussuaq) between 2017 and 2021. Velocity data are derived and corrected following the methods outlined in Chudley <em>et al.</em> (2022). </p> <p>NetCDF files are created by, and tested to be readable by, Python's xarray package.</p> <p>The dimensions of the netCDF file are as follows:</p> <ul> <li><strong>X</strong> - <em>x </em>coordinates in NSDIC Sea Ice Polar Stereographic North (EPSG:3413).</li> <li><strong>Y</strong> - <em>y</em> coordinates in NSDIC Sea Ice Polar Stereographic North (EPSG:3413).</li> <li><strong>time</strong> - temporal midpoint of velocity field.</li> </ul> <p>The variables of the netCDF file are as follows:</p> <ul> <li><strong>dmag</strong> - the absolute magnitude of the velocity, in metres per day.</li> <li><strong>dx</strong> - the velocity in the <em>x</em> direction, in metres per day.</li> <li><strong>dy</strong> - the velocity in the <em>y</em> direction, in metres per day.</li> <li><strong>date1</strong> - the date and time of the first scene acquisition.</li> <li><strong>date2</strong> - the date and time of the second scene acquisition.</li> <li><strong>baseline</strong> - the temporal baseline, in days, between scene acquisitions.</li> <li><strong>orbit_pair</strong> - the combination of orbital pathways in the string format 'RXXX_RYYY', where XXX is relative orbit number of the first scene and YYY the relative orbit number of the second scene.</li> <li><strong>mag_rmse</strong> - the root mean square error of the absolute velocity of the off-ice area. </li> <li><strong>dx_mean</strong> - the mean velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dx_sd</strong> - the standard deviation of the velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dy_mean</strong> - the mean velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dy_sd</strong> - the standard deviation of the velocity of the off-ice area in the <em>y</em> direction.</li> </ul>
Data from: Improved robustness to gene tree incompleteness, estimation errors, and systematic homology errors with weighted TREE-QMC
Open the record for dataset details and reuse information.
Confronting sources of systematic error to resolve historically contentious relationships: a case study using gadiform fishes (Teleostei, Paracanthopterygii, Gadiformes)
<p>Reliable estimation of phylogeny is central to avoid inaccuracy in downstream macroevolutionary inferences. However, limitations exist in the implementation of concatenated and summary coalescent approaches, and Bayesian and full coalescent inference methods may not yet be feasible for computation of phylogeny using complicated models and large datasets. Here, we explored methodological (e.g., optimality criteria, character sampling, model selection) and biological (e.g., heterotachy, branch length heterogeneity) sources of systematic error that can result in biased or incorrect parameter estimates when reconstructing phylogeny by using the gadiform fishes as a model clade. Gadiformes include some of the most economically important fishes in the world (e.g., Cods, Hakes, and Rattails). Despite many attempts, a robust higher-level phylogenetic framework was lacking due to limited character and taxonomic sampling, particularly from several species-poor families that have been recalcitrant to phylogenetic placement. We compiled the first phylogenomic dataset, including <span>14,208 loci (>2.8 M bp) from 58 species representing all recognized gadiform families, to infer a time-calibrated phylogeny for the group. </span>Data were generated with a gene-capture approach targeting coding DNA sequences from single-copy protein-coding genes. Species-tree and concatenated maximum-likelihood analyses resolved all family-level relationships within Gadiformes. While there were a few differences between topologies produced by the DNA and the amino acid datasets, most of the historically unresolved relationships among gadiform lineages were consistently well resolved with high support in our analyses regardless of the methodological and biological approaches used. However, at deeper levels, we observed inconsistency in branch support estimates between bootstrap and gene and site coefficient factors (gCF, sCF). Despite numerous short internodes, all relationships received unequivocal bootstrap support while gCF and sCF had very little support, reflecting hidden conflict across loci. Most of the gene-tree and species-tree discordance in our study is a result of short divergence times, and consequent lack of informative characters at deep levels, rather than incomplete lineage sorting (ILS). We use this phylogeny to establish a new<span> higher-level classification of Gadiformes as a way of clarifying the evolutionary diversification of the order.</span> We recognize 17 families in five suborders: Bregmacerotoidei, Gadoidei, Ranicipitoidei, Merluccioidei, and Macrouroidei (including two subclades). A time-calibrated analysis using 15 fossil taxa suggests that Gadiformes evolved ~79.5 million years ago (Ma) in the late Cretaceous, but that most extant lineages diverged after the Cretaceous-Paleogene (K-Pg) mass extinction (66 Ma)<span>. </span>Our results reiterate the importance of examining phylogenomic analyses for evidence of systematic error that can emerge as a result of unsuitable modeling of biological factors and/or methodological issues, even when datasets are large and yield high support for phylogenetic relationships.</p>
Synthetic automotive LiDAR with non-systematic error and automotive LiDAR based on active stereo dataset
<p>The synthetic dataset was generated by transforming the original dataset using several methods. Each of these transformations occurs from a use case:</p><ul><li>UC1 is the original dataset obtained from [1] and represents a point cloud dataset captured by an ideal LiDAR</li><li>UC2 is a realistic point cloud dataset obtained by simulating the non-systematic error of a Velodyne HDL-64E and applying to UC1</li><li>UC3 is our approach to replace the LiDAR with an active stereo setup. Where the point cloud are captured using two cameras, operating stereoscopically, and a dot projector. The cameras are perfectly calibrated and the triangulation is always correct.</li><li>UC4 also obtains the point clouds through triangulation. However we introduced a calibration error on the right camera. The error has the value of 1 pixel and is added to every dimension of the rotation matrix of the right camera.</li><li>UC5 performs an ideal triangulation, same as UC3. However, in this use case, we introduce camera noise to the point clouds.</li><li>UC6 is a combination of the triangulation from UC4 and the camera noise from UC5.</li></ul><p>Additionally, the labels, images and calibration file are also in [1]. For further details, please check the dataset generation source code [2]. </p><p>[1] https://zenodo.org/records/7184990</p><p>[2] https://github.com/RobertoGraca/Active_Stereo_Based_LiDAR</p>
Confronting sources of systematic error to resolve historically contentious relationships: a case study using gadiform fishes (Teleostei, Paracanthopterygii, Gadiformes)
Open the record for dataset details and reuse information.
Data from: Variation across mitochondrial gene trees provides evidence for systematic error: how much gene tree variation is biological?
The use of large genomic datasets in phylogenetics has highlighted extensive topological variation across genes. Much of this discordance is assumed to result from biological processes. However, variation among gene trees can also be a consequence of systematic error driven by poor model fit, and the relative importance of biological versus methodological factors in explaining gene tree variation is a major unresolved question. Using mitochondrial genomes to control for biological causes of gene tree variation, we estimate the extent of gene tree discordance driven by systematic error and employ posterior prediction to highlight the role of model fit in producing this discordance. We find that the amount of discordance among mitochondrial gene trees is similar to the amount of discordance found in other studies that assume only biological causes of variation. This similarity suggests that the role of systematic error in generating gene tree variation is underappreciated and critical evaluation of fit between assumed models and the data used for inference is important for the resolution of unresolved phylogenetic questions.
Data from: Phylogenomics of Lophotrochozoa with consideration of systematic error
Open the record for dataset details and reuse information.
Data from: Adaptation to random and systematic errors: Comparison of amputee and non-amputee control interfaces with varying levels of process noise
Open the record for dataset details and reuse information.
Data from: Variation across mitochondrial gene trees provides evidence for systematic error: how much gene tree variation is biological?
Open the record for dataset details and reuse information.
Ordered phylogenomic subsampling enables diagnosis of systematic errors in the placement of the enigmatic arachnid order Palpigradi
<p><span><span><span><span><span><span><span><span><span><span><span>The miniaturized arachnid order Palpigradi has ambiguous phylogenetic affinities, due to its odd combination of plesiomorphic and derived morphological traits. This lineage has never been sampled in phylogenomic datasets because of its small body size and fragility of most species, a sampling gap of immediate concern to recent disputes over arachnid monophyly. To redress this gap, we sampled a population of the cave-inhabiting species <i>Eukoenenia spelaea</i> from Slovakia and inferred its placement in the phylogeny of Chelicerata using dense phylogenomic matrices of up to 1450 loci, drawn from high-quality transcriptomic libraries and complete genomes. The complete matrix included exemplars of all extant orders of Chelicerata. Analyses of the complete matrix recovered palpigrades as the sister group of the long-branch order Parasitiformes (ticks) with high support. However, sequential deletion of long-branch taxa revealed that the position of palpigrades is prone to topological instability. Phylogenomic subsampling approaches that maximized taxon or dataset completeness recovered palpigrades as the sister group of camel spiders (Solifugae), with modest support. While this relationship is congruent with the location and architecture of the coxal glands, a long-forgotten character system that opens in the pedipalpal segments only in palpigrades and solifuges, we show that nodal support values in concatenated supermatrices can mask high levels of underlying topological conflict in the placement of the enigmatic Palpigradi. </span></span></span></span></span></span></span></span></span></span></span></p>
Figure 3 from: Welter-Schultes F, Görlich A, Lutze A (2016) Sherborn's Index Animalium: New names, systematic errors and availability of names in the light of modern nomenclature. In: Michel E (Ed.) Anchoring Biodiversity Information: From Sherborn to the 21st century and beyond. ZooKeys 550: 173–187. https://doi.org/10.3897/zookeys.550.10041
Figure 3 - Example of a new name established in a chaotically arranged early zoological work. Detail of Hartmann (1821: p. 231), with the original description of Helix ruderata Hartmann, 1821 (currently Discus ruderatus). For the non-insider it is very difficult to see that a new name was established here. It is necessary to understand precisely the content of the German text: "s. meine Tab. II. f. 11. Der vorigen ähnlich, aber aufgeblasener, rauher, weniger Umgänge" (= see my plate II, figure 11. Similar to the previous one, but more inflated, rougher, fewer whorls).
Figure 1 from: Welter-Schultes F, Görlich A, Lutze A (2016) Sherborn's Index Animalium: New names, systematic errors and availability of names in the light of modern nomenclature. In: Michel E (Ed.) Anchoring Biodiversity Information: From Sherborn to the 21st century and beyond. ZooKeys 550: 173–187. https://doi.org/10.3897/zookeys.550.10041
Figure 1 - Number of names. Total number of taxonomic names listed in Sherborn's Index Animalium and in AnimalBase (2011) plotted in 5-year intervals. It must be taken into account that Sherborn's numbers included a proportion of 30% of names that were not new, while in AnimalBase this proportion was much lower (less than 5 %).
Figure 2 from: Welter-Schultes F, Görlich A, Lutze A (2016) Sherborn's Index Animalium: New names, systematic errors and availability of names in the light of modern nomenclature. In: Michel E (Ed.) Anchoring Biodiversity Information: From Sherborn to the 21st century and beyond. ZooKeys 550: 173–187. https://doi.org/10.3897/zookeys.550.10041
Figure 2 - Languages used in early zoological literature. Language analysis of 2100 arbitrarily selected binominal zoological works published between 1758 and 1850.
Data from: Bacterial cooperation causes systematic errors in pathogen risk assessment due to the failure of the independent action hypothesis
Open the record for dataset details and reuse information.
Ordered phylogenomic subsampling enables diagnosis of systematic errors in the placement of the enigmatic arachnid order Palpigradi
Open the record for dataset details and reuse information.
Resolving systematic errors in widely-used enhancer activity assays in human cells enables genome-wide functional enhancer characterization.
GEO Series GSE100432. Homo sapiens. 29 samples. Type: Other.
Comparison of systematic sequencing errors using spike-in standards
GEO Series GSE36217. Homo sapiens. 6 samples. Type: Other.
Validation data: Multi-objective support vector regression reduces systematic error in moderate resolution maps of tree species abundance
<p>Validation data used in the analysis presented by Legaard et al. (in review). Validation data are sufficient to replicate model comparisons presented in this paper. Note that model training data were provided by the USDA Forest Service, Forest Inventory and Analysis Program through a collaborative agreement, are maintained by the USDA Forest Service as confidential, and cannot be shared.</p> <p>Legaard, K., Simons-Legaard, E., Weiskittel, A., Multi-objective support vector regression reduces systematic error in moderate resolution maps of tree species abundance, Remote Sensing, in review.</p>
Systematic detection of amino acid substitutions in proteome reveals a mechanistic basis of ribosome errors and selection for translation fidelity
GEO Series GSE128812. Escherichia coli; Escherichia coli BW25113. 9 samples. Type: Non-coding RNA profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.