Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,577
datasets available to search
ShareScore release 0.9.0
Dataset results
2,577 results for “Inference”
Datasets for Training and Inference of DeepRLI
<p>This repository contains datasets used in the development of the <a href="https://github.com/fairydance/DeepRLI" target="_blank" rel="noopener">DeepRLI</a> model for protein–ligand interaction prediction, which includes the training dataset for the model and data related to the PLK1 kinase involved in the case study.</p>
Human pan-body age- and sex-specific molecular phenomena inferred from public transcriptome data using machine learning - Data
<p>Expression data used in manuscript <i>Human pan-body age- and sex-specific molecular phenomena inferred from public transcriptome data using machine learning</i></p>
Fig. 11. Bayesian inference trees. A. 16S rRNA dataset. B. Cytochrome oxidase I in Designation of a neotype for Myxicola infundibulum (Montagu, 1808) (Annelida: Sabellidae) and a new species from the UK
Fig. 11. Bayesian inference trees. A. 16S rRNA dataset. B. Cytochrome oxidase I gene dataset. The first value at each node represents maximum likelihood bootstrap support, the second the Bayesian posterior probabilities and the third the maximum parsimony bootstrap support.
Fig. 2 in Phylogenetic Relationships Of Malayan And Malagasy Pygmy Shrews Of The Genus Suncus (Soricomorpha: Soricidae) Inferred From Mitochondrial Cytochrome B Gene Sequences
Fig. 2. The neighbour-joining (A) and Bayesian (B) trees for Suncus inferred from 1140 base-pairs of cytochrome b gene sequence. Bootstrap and posterior probability values are given above branches.
Fig. 1 in Phylogenetic Relationships Of Malayan And Malagasy Pygmy Shrews Of The Genus Suncus (Soricomorpha: Soricidae) Inferred From Mitochondrial Cytochrome B Gene Sequences
Fig. 1. Male Malayan pygmy shrew (Suncus malayanus) captured in the Cameron Highlands, Pahang, Peninsular Malaysia, in a pitfall trap set on the forest floor. Notice the characteristic large ears and dark fine pelage.
Fig. 3 in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data
Fig. 3. Bayesian consensus tree for the anabantoids, channids and catfishes (silurids, bagrids, clariids) obtained using partial Cytochrome b sequences with cyprinids as outgroup. The heteronchocleidids genera present on the anabantoids and channids are shown with their geographical areas. Values shown at each node refer to Bayesian posterior probabilities. (*refer to Table 3 for names used in GenBank).
Fig. 2. Bayesian consensus tree generated from partial 28S in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data
Fig. 2. Bayesian consensus tree generated from partial 28S rDNA sequences (D1 domain) with Diplectanum spp. and Gyrodactylus spp. as outgroups. Values shown at each node refer to Bayesian (BI) posterior probabilities/maximum likelihood (ML) percentages of the bootstrap values with 100 replicates. Bootstrap values lower than 50 are given as dashes (-).
Fig. 1 in Relationships Of The Heteronchocleidids (Heteronchocleidus, Eutrianchoratus And Trianchoratus) As Inferred From Ribosomal Dna Nucleotide Sequence Data
Fig. 1. Neighbour joining (NJ) tree constructed by PAUP* using partial 28S rDNA sequences (D1 domain) with Diplectanum spp. and Gyrodactylus spp. as outgroups. Percentages of the bootstrap values for neighbour joining (NJ)/maximum parsimony (MP) (NJ & MP=1,000 replicates) are shown along the branches. Bootstrap values lower than 50 are given as dashes (-).
Germline CpG methylation signatures in the human population inferred from genetic polymorphism
<p>This repository contains data released accompanying the manuscript "Germline CpG methylation signatures in the human population inferred from genetic polymorphism". </p>
Training and test data, plus saved models for the upcoming paper `Top-down perceptual inference shaping the activity of early visual cortex'
<p>Each .pkl file contains a training or test dataset in the form of a Python dictionary (generated with Python 3.8.5) with the following fields:</p><ul><li>'train_images': 640,000 float32 images used for model training. These are 40px images that contain 1600 pixel intensities each.</li><li>'train_labels': float32 labels for each image in 'train_images'. All natural images are labeled with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0, according to their texture family.</li><li>'test_images': 64,000 float32 images used for model testing. These are 40px images that contain 1600 pixel intensities each.</li><li>'test_labels': float32 labels for each image in 'test_images'. All natural images are labeled with 0.0. Texture images are labeled with 0.0, 1,0, 2.0, 3.0, or 4.0, according to their texture family.</li></ul><p>The .zip file contains a saved model snapshot and various intermediate evaluative data. Details on these are coming soon.</p>
Dataset: Evolution of large Venusian coronae inferred from structural analyses and the presence of low-angle faults in chasmata
<p>The vector datasets mentioned in this article are available. You can find the Magellan SAR data and the Global Topography Data Records (GTDR) on NASA's Planetary Data System (PDS) website (specific links can be found on https://pdsgeosciences.wustl.edu/missions/magellan/index.htm). In addition, the stereo-derived topography dataset of Herrick et al. 2012 is also available on their personal website https://sites.google.com/alaska.edu/robertherrick/resources/stereo-derived-topography-for-venus.</p>
Data from: Soil incubation methods lead to large differences in inferred methane production temperature sensitivity
<p>Quantifying the temperature sensitivity of methane (CH4) production is crucial for predicting how wetland ecosystems will respond to climate warming. Typically, the temperature sensitivity (often quantified as a Q10 value) is derived from laboratory incubation studies and then used in biogeochemical models. However, studies report wide variation in incubation-inferred Q10 values, with a large portion of this variation remaining unexplained. Here we applied observations in Stordalen Mire, a thawing permafrost peatland, and a well-tested process-rich model, ecosys, to interpret incubation observations and investigate controls on inferred CH4 production temperature sensitivity. We developed a Field-Storage-Incubation (FSI) modeling approach to mimic the full incubation sequence, including field sampling at a particular time in the growing season,refrigerated storage, and the laboratory incubation process, followed by model evaluation. We found that CH4 production rates during incubation are regulated by seasonally-dependent substrate availability and active microbial biomass of key microbial functional groups. Applying a model sensitivity analysis, we found that storage duration, storage temperature, and field sampling time significantly affect CH4 production during incubation. Shorter storage duration and lower storage temperature led to larger CH4 production during incubation. Our findings revealed a wide range of inferred Q10 values (1.2 to 3.5), which we attribute to incubation temperatures, incubation duration, storage duration, and sampling time. Q10 of CH4 production is controlled by many interacting biological, biochemical, and physical processes, which cause the aggregated Q10 values to differ from those of the component processes. Terrestrial ecosystem models that use a constant Q10 value to represent temperature responses may therefore predict biased soil carbon cycling under future climate scenarios.</p> <p>This dataset includes all the data used to plot figures in the manuscript, including Fig.2-6 and Fig.S2-S11. Each sheet in the aggregated spreadsheet corresponds to one figure in the manuscript. The simulation experiment setup and analyses are thoroughly described in the manuscript. Here we provide a brief summary. The data includes field greenhouse gas observations and laboratory incubation measurements of CH4 production in Stordalen Mire. These datasets were already published and references were provided in the manuscript and spreadsheet. The data also includes simulation data, including modeled cumulative CH4 production, CH4 production rates, substrate concentrations, and active microbial biomass under different incubation temperature, sampling time and storage conditions. This data also includes inferred temperature sensitivity of CH4 production as Q10 values under different scenarios. Please refer to the manuscript for more detailed information.</p> <p>Please see "Related works" at the bottom of this page and the "References" tab in the spreadsheet for a full list of source datasets and associated publications.</p> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute, funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council’s grant 4.3-2021-00164. This research used resources of the National Energy Research Scientific Computing Center (NERSC) which is a U.S. Department of Energy Office of Science user facility. This research used the Lawrencium computational cluster resource provided by the IT Division at the Lawrence Berkeley National Laboratory (Supported by the Director, Office of Science, Office of Basic Energy Sciences, of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231). Incubation and field observation data were collected under the IsoGenie Project, which was funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632, DE-SC0010580, and DE-SC0016440.</p>
Clockor2: Inferring global and local strict molecular clocks using root-to-tip regression
<p>Molecular sequence data from rapidly evolving organisms are often sampled at different points in time. Sampling times can then be used for molecular clock calibration. The root-to-tip (RTT) regression is an essential tool to assess the degree to which the data behave in a clock-like fashion. Here, we introduce Clockor2, a client-side web application for conducting RTT regression. Clockor2 uniquely allows users to quickly fit local and global molecular clocks, thus handling the increasing complexity of genomic datasets that sample beyond the assumption homogeneous host populations. Clockor2 is efficient, handling trees of up to the order of 10^4 tips, with significant speed increases compared to other RTT regression applications. Although clockor2 is written as a web application, all data processing happens on the client-side, meaning that data never leaves the user's computer. Clockor2 is freely available at https://clockor2.github.io/</p>
Population size differences can lead to biases in phylogenetic inference and introgression detection in the presence of purifying selection
<p>Phylogenetic reconstruction and introgression detection rely on an assumption about the probability distribution of gene tree topologies. Recently, evidence has emerged that population size differences can affect the probability distribution of gene tree topologies in the presence of purifying selection. Here, using the population genetic simulator SLiM, we provide evidence that in the presence of purifying selection, population size differences can lead to biases in phylogenetic inference. We also provide evidence that in the presence of purifying selection, population size differences can cause statistics used for introgression detection to exhibit patterns resembling those caused by introgression. In addition, we present a theoretical analysis showing that the occurrence of population size–dependent gene tree distributions is an inherent consequence of purifying selection. Our work underscores the importance of considering the potential confounding effect of purifying selection on phylogenetic inference and introgression detection.</p>
Different currencies for calculating resource phenology result in opposite inferences about trophic mismatches
<p>Shifts in phenology are among the key responses of organisms to climate change. When rates of phenological change differ between interacting species they may result in phenological asynchrony. Studies have found conflicting patterns concerning the direction and magnitude of changes in synchrony, which have been attributed to biological factors. A hitherto overlooked additional explanation is differences in the currency used to quantify resource phenology, such as abundance and biomass. Studying an insectivorous bird, Sanderling, and its prey, we show that the median date of cumulative arthropod biomass occurred, on average, 6.9 days after the median date of cumulative arthropod abundance. In some years this difference could be as large as 21 days. For 23 years, hatch dates of Sanderlings became less synchronized with the median date of arthropod abundance, but more synchronized with the median date of arthropod biomass. The currency-specific trends can be explained by our finding that mean biomass per arthropod specimen increased with date. Using a conceptual simulation, we show that estimated rates of phenological change for abundance and biomass can differ depending on temporal shifts in the size distribution of resources. We conclude that studies of trophic mismatch based on different currencies for resource phenology can be incompatible with each other.</p>
FIGURE 1 in Evolutionary history of species of the fireFly subgenus Hotaria (Coleoptera, Lampyridae, Luciolinae, Luciola) inferred from DNA barcoding data
FIGURE 1 Neighbor-joining (NJ) tree of 128 samples of 14 morphospecies based on COI barcode sequences. The percentages at terminal taxa indicate intraspecific genetic divergence. The percentages at each node indicate genetic divergence for the split. Weakly supported nodes (bootstrap values below 70%) are Downloaded from Brill.com 12/12/2023 03:05:57PM shown in red. via Open Access. This is an open access article distributed under the terms of the CC-BY 4.0 License. https://creativecommons.org/licenses/by/4.0/
FIGURE 4 in Evolutionary history of species of the fireFly subgenus Hotaria (Coleoptera, Lampyridae, Luciolinae, Luciola) inferred from DNA barcoding data
FIGURE 4 Time-calibrated phylogram calculated using BEAST based on the COI dataset for 128 samples of 14 morphospecies. Blue numbers below nodes are estimated diversification dates with confidence intervals (blue bars). Posterior probabilities (PP) are marked on nodes with an asterisk (PP = 1.00). Pli = Pliocene, Ple = Pleistocene
FIGURE 5 in Evolutionary history of species of the fireFly subgenus Hotaria (Coleoptera, Lampyridae, Luciolinae, Luciola) inferred from DNA barcoding data
FIGURE 5 Distributional patterns (collection sites) for terminal taxa of Luciola unmunsana (LU), L. papariensis (LP), and L. tsushimana (LT) constructed using BEAST. Red dotted line indicates the approximate location of the "Bekdudaegan" mountains. Yellow dotted line indicates the approximate location of the "Hannam-Geumbuk Jeongmaeck" mountains. Green dotted line denotes the approximate location of the "Nakdong Jeongmaeck" mountains. The map was extracted from Google Earth.
FIGURE 3 Majority-rule consensus tree from a in Evolutionary history of species of the fireFly subgenus Hotaria (Coleoptera, Lampyridae, Luciolinae, Luciola) inferred from DNA barcoding data
FIGURE 3 Majority-rule consensus tree from a Bayesian analysis (BI) of 128 samples of 14 morphospecies based on COI barcode sequences. The numbers at each node indicate Downloadedposteriorfrom Brill. probabilities com 12. /12/ Weakly 2023 03 sup-:05:57PM ported nodes (posterior via probabilityOpen below Access. 0.95) Thisareis an shownopenin red. access article distributed under the terms of the CC-BY 4.0 License. https://creativecommons.org/licenses/by/4.0/
FIGURE 2 in Evolutionary history of species of the fireFly subgenus Hotaria (Coleoptera, Lampyridae, Luciolinae, Luciola) inferred from DNA barcoding data
FIGURE 2 Maximum likelihood (ML) tree of 128 samples of 14 morphospecies based on COI barcode sequences. The numbers at each node indicate support (%). Weakly supported nodes (below 70%) are shown in red. Downloaded from Brill.com 12/12/2023 03:05:57PM via Open Access. This is an open access article distributed under the terms of the CC-BY 4.0 License. https://creativecommons.org/licenses/by/4.0/
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.