Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
56
datasets available to search
ShareScore release 0.9.0
Dataset results
56 results for “variational inference”
Fitness consequences of structural variation inferred from a House Finch pangenome
Open the record for dataset details and reuse information.
Data from: Major histocompatibility complex class II variation in bottlenose dolphin from Adriatic Sea: inferences about the extent of balancing selection
The bottlenose dolphin (Tursiops truncatus) is the most common cetacean species worldwide and the only marine mammal species resident in the Croatian part of the Adriatic Sea. To gain insight into genetic diversity of bottlenose dolphins at adaptively important loci relevant to conservation, we analysed the polymorphism of major histocompatibility complex (MHC) genes, which play a key role in pathogen confrontation and clearance. Specifically, we examined the diversity of MHC class II DRA, DQA and DQB alleles in 50 bottlenose dolphins from the Adriatic Sea collected between 1997 and 2011 and in 12 animals from other Mediterranean locations. Notable variation in DQA, DQB and three-locus haplotypes was found, with all 10 DQA and 12 DQB alleles encoding unique protein products. Analysis of the ratio of non-synonymous to synonymous substitution rates suggests that positive selection acts at both highly variable loci. Phylogenetic analyses revealed trans-species polymorphism at the DQB locus, strongly indicating the influence of balancing selection in the long term. In fact, the balancing selection observed in bottlenose dolphins is higher than that reported for most other cetaceans and comparable to that seen in terrestrial mammals.
Data from: The hidden history of the snowshoe hare, Lepus americanus: extensive mitochondrial DNA introgression inferred from multilocus genetic variation
Hybridization drives the evolutionary trajectory of many species or local populations, and assessing the geographic extent and genetic impact of interspecific gene flow may provide invaluable clues to understand population divergence or the adaptive relevance of admixture. In North America, hares (Lepus spp.) are key species for ecosystem dynamics and their evolutionary history may have been affected by hybridization. Here we reconstructed the speciation history of the three most widespread hares in North America - the snowshoe hare (Lepus americanus), the white-tailed jackrabbit (L. townsendii) and the black-tailed jackrabbit (L. californicus) - by analyzing sequence variation at eight nuclear markers and one mitochondrial DNA (mtDNA) locus (6 240 bp; 94 specimens). A multilocus-multispecies coalescent-based phylogeny suggests that L. americanus diverged ~2.7 Mya and that L. californicus and L. townsendii split more recently (~1.2 Mya). Within L. americanus a deep history of cryptic divergence (~2.0 Mya) was inferred, which coincides with major speciation events in other North American species. While the isolation-with-migration model suggested that nuclear gene flow was generally rare or absent among species or major genetic groups, coalescent simulations of mtDNA divergence revealed historical mtDNA introgression from L. californicus into the Pacific Northwest populations of L. americanus. This finding marks a history of past reticulation between these species, which may have affected other parts of the genome and influence the adaptive potential of hares during climate change.
Data from: Major histocompatability complex variation in insular populations of the Egyptian vulture: inferences about the roles of genetic drift and selection
Insular populations have attracted the attention of evolutionary biologists because of their morphological and ecological peculiarities with respect to their mainland counterparts. Founder effects and genetic drift are known to distribute neutral genetic variability in these demes. However, elucidating whether these evolutionary forces have also shaped adaptive variation is crucial to evaluate the real impact of reduced genetic variation in small populations. Genes of the Major Histocompatibility Complex (MHC) are classical examples of evolutionarily relevant loci because of their well-known role in pathogen confrontation and clearance. In this study, we aim to disentangle the partial roles of genetic drift and natural selection in the spatial distribution of MHC variation in insular populations. To this end, we integrate the study of neutral (22 microsatellites and one mtDNA locus) and MHC class II variation in one mainland (Iberia) and two insular populations (Fuerteventura and Menorca) of the endangered Egyptian vulture (Neophron percnopterus). Overall, the distribution of the frequencies of individual MHC alleles (N=17 alleles from two class II B loci) does not significantly depart from neutral expectations, which indicates a prominent role for genetic drift over selection. However, our results point towards an interesting co-evolution of gene duplicates that maintains different pairs of divergent alleles in strong linkage disequilibrium on islands. We hypothesize that the co-evolution of genes may counteract the loss of genetic diversity in insular demes, maximize antigen recognition capabilities when gene diversity is reduced, and promote the co-segregation of the most efficient allele combinations to cope with local pathogen communities.
Data from: Using genetic variation to infer associations with climate in the common frog, Rana temporaria
Recent and historical species' associations with climate can be inferred using molecular markers. This knowledge of population and species-level responses to climatic variables can then be used to predict the potential consequences of ongoing climate change. The aim of this study was to predict responses of Rana temporaria to environmental change in Scotland by inferring historical and contemporary patterns of gene flow in relation to current variation in local thermal conditions. We first inferred colonization patterns within Europe following the last glacial maximum by combining new and previously published mitochondrial DNA sequences. We found that sequences from our Scottish samples were identical to (92%), or clustered with, the common haplotype previously identified from Western Europe. This clade showed very low mitochondrial variation, which did not allow inference of historical colonization routes but did allow interpretation of patterns of current fine-scale population structure without consideration of confounding historical variation. Second, we assessed fine-scale microsatellite-based patterns of genetic variation in relation to current altitudinal temperature gradients. No population structure was found within altitudinal gradients (average FST = 0.02), despite a mean annual temperature difference of 4.5 °C between low- and high-altitude sites. Levels of genetic diversity were considerable and did not vary between sites. The panmictic population structure observed, even along temperature gradients, is a potentially positive sign for R. temporaria persistence in Scotland in the face of a changing climate. This study demonstrates that within taxonomic groups, thought to be at high risk from environmental change, levels of vulnerability can vary, even within species.
Data from: Low major histocompatibility complex class II variation in the endangered Indo-Pacific humpback dolphin (Sousa chinensis): inferences about the role of balancing selection
It has been widely reported that the major histocompatibility complex (MHC) is under balancing selection due to its immune function across terrestrial and aquatic mammals. The comprehensive studies at MHC and other neutral loci could give us a synthetic evaluation about the major force determining genetic diversity of species. Previously, a low level of genetic diversity has been reported among the Indo-Pacific humpback dolphin (Sousa chinensis) in the Pearl River Estuary (PRE) using both mitochondrial marker and microsatellite loci. Here, the expression and sequence polymorphism of 2 MHC class II genes (DQB and DRB) in 32 S. chinensis from PRE collected between 2003 and 2011 were investigated. High ratios of non-synonymous to synonymous substitution rates, codon-based selection analysis, and trans-species polymorphism (TSP) support the hypothesis that balancing selection acted on S. chinensis MHC sequences. However, only 2 haplotypes were detected at either DQB or DRB loci. Moreover, the lack of deviation from the Hardy–Weinberg expectation at DRB locus combined with the relatively low heterozygosity at both DQB locus and microsatellite loci suggested that balancing selection might not be sufficient, which further suggested that genetic drift associated with historical bottlenecks was not mitigated by balancing selection in terms of the loss of MHC and neutral variation in S. chinensis. The combined results highlighted the importance of maintaining the genetic diversity of the endangered S. chinensis.
Models and data for "Amortized reparametrization: Efficient and Scalable Variational Inference for Latent SDEs"
Open the record for dataset details and reuse information.
Fig. 3 in Genetic variation in the spotted seal (Phoca largha Pallas, 1811) from the Rimsky-Korsakov Archipelago (Peter the Great Bay, western sea of Japan) as inferred from mitochondrial DNA control region sequences
Fig. 3. Consensus maximum likelihood tree demonstrating the matrilineal genealogy of some Phocidae species generated from the 460 bp mtDNA control region sequences. The numbers at branch nodes represent bootstrap support of 1000 replications for the maximum likelihood trees and posterior probabilities of 2,000,000 generations for the Bayesian trees with the same topology, respectively. The accession numbers of Phoca largha from Liaodong Bay are highlighted in bold. Scale bar indicates the relative branch lengths.
Fig. 2 in Genetic variation in the spotted seal (Phoca largha Pallas, 1811) from the Rimsky-Korsakov Archipelago (Peter the Great Bay, western sea of Japan) as inferred from mitochondrial DNA control region sequences
Fig. 2. Minimum spanning network showing the mutational relationships among Phoca largha haplotypes detected in a sample of 32 spotted seal underyearlings. The tick marks on the branches indicate mutational changes. The circle sizes correspond to the number of haplotypes. The haplotypes of groups (A) and (B) are shown in gray and white circles, respectively. The median vectors are indicated by dark dots.
Figure 2 in Complete mitochondrial genome of Tetraophasis szechenyii Madarász, 1885 (Aves: Galliformes: Phasianidae), and its genetic variation as inferred from the mitochondrial DNA Control Region
Figure 2. Median-joining network of all the control region haplotypes found in Tetraophasis szechenyii. Notes: Missing haplotypes in the network are represented by black dots; circle sizes are proportional to the number of individuals sharing the same haplotypes (n); each mutation step is shown as a short line connecting neighbouring haplotypes; numbers of mutations between haplotypes are indicated near branches if greater than 1.
Figure 1 in Complete mitochondrial genome of Tetraophasis szechenyii Madarász, 1885 (Aves: Galliformes: Phasianidae), and its genetic variation as inferred from the mitochondrial DNA Control Region
Figure 1. Molecular phylogenetic tree derived from the complete DNA sequences of 12 mitochondrial protein-coding genes using Bayesian inference and maximum likelihood analyses. Notes:The numbers beside the nodes are Bayesian posterior probabilities (≥ 0.9 retained) and bootstrap proportions of maximum likelihood analyses calculated with 100 replicates (≥ 50% retained); Anas platyrhynchos and Alectura lathami were set as outgroups; *clades not supported by Bayesian inference.
Data from: Low major histocompatibility complex class II variation in the endangered Indo-Pacific humpback dolphin (Sousa chinensis): inferences about the role of balancing selection
Open the record for dataset details and reuse information.
Data from: Major histocompatibility complex class II variation in bottlenose dolphin from Adriatic Sea: inferences about the extent of balancing selection
Open the record for dataset details and reuse information.
Data from: The hidden history of the snowshoe hare, Lepus americanus: extensive mitochondrial DNA introgression inferred from multilocus genetic variation
Open the record for dataset details and reuse information.
Data from: Major histocompatability complex variation in insular populations of the Egyptian vulture: inferences about the roles of genetic drift and selection
Open the record for dataset details and reuse information.
Data from: Using genetic variation to infer associations with climate in the common frog, Rana temporaria
Open the record for dataset details and reuse information.
Data from: Mate choice and the genetic basis for color variation in a polymorphic dart frog: inferences from a wild pedigree
Open the record for dataset details and reuse information.
Data from: Heritable variation in host tolerance and resistance inferred from a wild host– parasite system
Open the record for dataset details and reuse information.
Data and Code for Publication "Testing the Utility of Dental Morphological Trait Combinations for Inferring Human Neutral Genetic Variation"
<p>Data and code for publication: H. Rathmann, H. Reyes-Centeno, Testing the utility of dental morphological trait combinations for inferring human neutral genetic variation. <em>Proc. Natl. Acad. Sci. U.S.A.</em> 117, 10769-10777 (2020). DOI: 10.1073/pnas.1914330117</p> <p>The repository contains:</p> <ul> <li>“R-code.txt”: R code for an exhaustive search algorithm testing the utility of dental morphological traits and trait combinations for inferring human neutral genetic variation.</li> <li>“dental trait frequencies.csv”: Data set with 27 dental morphological trait frequencies for 20 modern human populations worldwide used for analysis. Data from G. R. Scott, C. G. Turner, G. C. Townsend, M. Martinón-Torres, <em>The Anthropology of Modern Human Teeth</em> (Cambridge University Press, 2018). DOI: 10.1017/ 9781316795859</li> <li>“microsatellite loci mean sizes.csv”: Data set with 645 microsatellite mean allele sizes for 20 modern human populations worldwide used for analysis. Data from T. J. Pemberton, M. DeGiorgio, N. A. Rosenberg, Population structure in a comprehensive genomic data set on human microsatellite variation. <em>G3: Genes Genom. Genet.</em> 3, 891–907 (2013). DOI: 10.1534/g3.113.005728</li> <li>“utility estimates for 134217727 trait combinations.txt”: A large table with utility estimates for 27 dental morphological traits and all 134,217,700 possible trait combinations.</li> </ul> <p>Abbreviations for the 20 population names (rows) in “dental trait frequencies.csv” and “microsatellite loci mean sizes.csv” as follows:</p> <ul> <li>AUS = Australia</li> <li>CAS = Central Asia</li> <li>EAF = Eastern Africa</li> <li>EAS = East Asia</li> <li>EEU = Eastern Europe</li> <li>IND = India</li> <li>MAM = Mesoamerica</li> <li>MEL = Melanesia</li> <li>MIC = Micronesia</li> <li>NAF = North Africa</li> <li>NAM = North America</li> <li>NESI = Northeast Siberia</li> <li>NGU = New Guinea</li> <li>NWAM = Na-Dene</li> <li>POL = Polynesia</li> <li>SAM = South America</li> <li>SAN = San</li> <li>SEAS = Southeast Asia</li> <li>WEU = Western Europe</li> <li>WSAF = Sub-Saharan Africa</li> </ul> <p>Abbreviations for the 27 dental morphological trait names (columns) in “dental trait frequencies.csv” as follows:</p> <ul> <li>T1 = Winging (UI1)</li> <li>T2 = Shoveling (UI1)</li> <li>T3 = Double-Shoveling (UI1)</li> <li>T4 = Interruption Grooves (UI2)</li> <li>T5 = Tuberculum Dentale (UI2)</li> <li>T6 = Mesial Ridge (UC)</li> <li>T7 = Distal Accessory Ridge (UC)</li> <li>T8 = Hypocone (UM2)</li> <li>T9 = Carabelli Trait (UM1)</li> <li>T10 = Cusp 5 (UM1)</li> <li>T11 = Enamel Extensions (UM1)</li> <li>T12 = Peg-Reduced-Missing (UM3)</li> <li>T13 = Lingual Cusp Number (LP2)</li> <li>T14 = Groove Pattern (LM2)</li> <li>T15 = Cusp 6 (LM1)</li> <li>T16 = Cusp Number (LM2)</li> <li>T17 = Deflecting Wrinkle (LM1)</li> <li>T18 = Distal Trigonid Crest (LM1)</li> <li>T19 = Protostylid (LM1)</li> <li>T20 = Cusp 7 (LM1)</li> <li>T21 = Odontomes (UP-LP)</li> <li>T22 = Root Number (UP1)</li> <li>T23 = Root Number (UM2)</li> <li>T24 = Root Number (LC)</li> <li>T25 = Tomes’ Root (LP1)</li> <li>T26 = Root Number (LM1)</li> <li>T27 = Root Number (LM2)</li> </ul> <p>Abbreviations for the 645 microsatellite allele locus names (columns) in “microsatellite loci mean sizes.csv” as in T. J. Pemberton, M. DeGiorgio, N. A. Rosenberg, Population structure in a comprehensive genomic data set on human microsatellite variation. <em>G3: Genes Genom. Genet.</em> 3, 891–907 (2013). DOI: 10.1534/g3.113.005728</p>
Data from: Inferring diversification rate variation from phylogenies with fossils
Time-calibrated phylogenies of living species have been widely used to study the tempo and mode of species diversification. However, it is increasingly clear that inferences about species diversification — extinction rates in particular — can be unreliable in the absence of paleontological data. We introduce a general framework based on the fossilized birth-death process for studying speciation-extinction dynamics on phylogenies of extant and extinct species. Our model assumes that phylogenies can be modeled as a mixture of distinct evolutionary rate regimes and that a hierarchical Poisson process governs the number of such rate regimes across a tree. We implemented the model in BAMM, a computational framework that uses reversible jump Markov chain Monte Carlo to simulate a posterior distribution of macroevolutionary rate regimes conditional on the branching times and topology of a phylogeny. The implementation we describe can be applied to paleontological phylogenies, neontological phylogenies, and to phylogenies that include both extant and extinct taxa. We evaluate performance of the model on datasets simulated under a range of diversification scenarios. We find that speciation rates are reliably inferred in the absence of paleontological data. However, the inclusion of fossil observations substantially increases the accuracy of extinction rate estimates. We demonstrate that the inferences are relatively robust to at least some violations of model assumptions, including heterogeneity in preservation rates and misspecification of the number of occurrences in paleontological datasets.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.