Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
56
datasets available to search
ShareScore release 0.9.0
Dataset results
56 results for “variational inference”
Variational Inference for Learning Representations of Natural Language Edits
<p>Performance Evaluation of Edit Representations (PEER), the dataset we use in the paper <a href="https://arxiv.org/abs/2004.09143">"Variational Inference for Learning Representations of Natural Language Edits"</a>.</p>
Constructing a high-density linkage map to infer the genomic landscape of recombination rate variation in European Aspen (Populus tremula)
<p>Data sets and files for linkage map construction and for inferring recombination rate variation in <em>Populus tremula</em>. Associated scripts for analyses can be found at <a href="https://github.com/parkingvarsson/Recombination_rate_variation">https://github.com/parkingvarsson/Recombination_rate_variation</a> </p>
Figure 3 in Determination of genetic variations between Apodemus mystacinus populations distributed in Turkey inferred from mtDNA PCR-RFLP
Figure 3. Restriction patterns of HinfI inferred from D-loop digestion (M: Marker–100bp DNA Ladder, 1. Ordu, 2. Trabzon, 3. Rize, 4. Artvin, 5–6. Erzincan, 7–8. Kahramanmaraş, 9. Adıyaman, 10–11. Adana, 12. Muğla, 13. Burdur, 14. Konya, 15. Antalya, 16. Mersin, 17. Kastamonu, 18. Zonguldak, 19. Düzce, 20. Balıkesir, 21. İzmir, 22. Aydın, 23. A. uralensis, 24. A. witherbyi, 25. D-loop PCR products).
Figure 2 in Determination of genetic variations between Apodemus mystacinus populations distributed in Turkey inferred from mtDNA PCR-RFLP
Figure 2. Restriction patterns of MboI, HaeIII, and RsaI inferred from cytb digestion (M: Marker–100bp DNA Ladder, 1. Ordu, 2. Trabzon, 3. Rize, 4. Artvin, 5. Erzincan, 6. Kahramanmaraş, 7. Adıyaman, 8. Adana, 9. Muğla, 10. Burdur, 11. Konya, 12. Antalya, 13. Mersin, 14. Kastamonu, 15. Zonguldak, 16. Düzce, 17. Balıkesir, 18. İzmir, 19. Aydın, 20. A. uralensis, 21. A. witherbyi, 22. Cytb PCR product).
Figure 5 in Determination of genetic variations between Apodemus mystacinus populations distributed in Turkey inferred from mtDNA PCR-RFLP
Figure 5. PCoA analysis of A. mystacinus clades. The scatter plot is of the scores of three principal eigenvalues inferred from NTSYS software. Each scatter point represents a specimen of A. mystacinus.
Figure 1 in Determination of genetic variations between Apodemus mystacinus populations distributed in Turkey inferred from mtDNA PCR-RFLP
Figure 1. Sampling localities of A. mystacinus specimens. Table 2. Restriction enzymes and their digestion sites with reaction procedures.
Figure 2. A in Genetic differentiation of the Meriones tristrami (Mammalia: Rodentia) subpopulations in Turkey - inferring allozyme variations
Figure 2. A dendrogram summarizing the genetic relationships of M. tristrami subpopulations (NTSYSpc options: Coefficient: SM (SimQual), clustering method: UPGMA) (see Nei, 1978) (Mt1 = Gaziantep, Adana, Mt2 = Central Anatolia (Cihanbeyli/Konya, Sivrihisar/Eskişehir), Mt3 = Denizli, Mt4 = Şanlıurfa, Mt5 = Iğdır, Mt6 = Tosya/Kastamonu, Mt7 = Karadağ/Karaman, Mt8 = Turgutlu/Manisa).
Figure 1 in Genetic differentiation of the Meriones tristrami (Mammalia: Rodentia) subpopulations in Turkey - inferring allozyme variations
Figure 1. Map of the locations of the M. tristrami specimens in Turkey (Gaziantep, Adana: Mt1; Central Anatolia (Cihanbeyli/Konya, Sivrihisar/Eskişehir): Mt2; Denizli: Mt3; Şanlıurfa: Mt4; Iğdır: Mt5; Tosya/Kastamonu: Mt6; Karadağ/Karaman: Mt7; Turgutlu/ Manisa: Mt8).
Data and Code for Publication "Inferring human neutral genetic variation from craniodental phenotypes"
<p>Data and code for publication: H. Rathmann et al., Inferring human neutral genetic variation from craniodental phenotypes. PNAS Nexus.</p> <p>The repository contains:</p> <ul> <li>“<em>R code for DP-DG analysis.txt</em>”: R code for testing levels of neutral evolutionary signals preserved in five craniodental data types: cranial metrics, dental metrics, cranial non-metric traits, dental non-metric traits, and craniodental metrics and non-metric traits combined.</li> </ul> <ul> <li>“<em>Cranial metric data.csv</em>”: Dataset consisting of 37 cranial metric variables for 26 worldwide modern populations, provided in a comma-separated values file format. The data were collected by T. Hanihara and originally presented in the publication titled: T. Hanihara, Comparison of craniofacial features of major human groups. <em>Am. J. Phys. Anthropol.</em> 99, 389–412 (1996) (<a href="https://doi.org/10.1002/(SICI)1096-8644(199603)99:3%3c389::AID-AJPA3%3e3.0.CO;2-S">https://doi.org/10.1002/(SICI)1096-8644(199603)99:3<389::AID-AJPA3>3.0.CO;2-S</a>).</li> </ul> <ul> <li>“<em>Dental metric data.csv</em>”: Dataset comprising 28 dental metric variables for 26 worldwide modern populations, provided in a comma-separated values file format. The data were collected by T. Hanihara and originally presented in the publication titled: T. Hanihara, H. Ishida, Metric dental variation of major human populations. <em>Am. J. Phys. Anthropol.</em> 128, 287–298 (2005) (<a href="https://doi.org/10.1002/ajpa.20080">https://doi.org/10.1002/ajpa.20080</a>).</li> </ul> <ul> <li>“<em>Cranial non-metric trait data.csv</em>”: Dataset consisting of 24 cranial non-metric trait variables for 26 worldwide modern populations, provided in a comma-separated values file format. The data were collected for the most part by T. Hanihara and presented in the publication titled: T. Hanihara, H. Ishida, Y. Dodo, Characterization of biological diversity through analysis of discrete cranial traits. <em>Am. J. Phys. Anthropol.</em> 121, 241–251 (2003) (<a href="https://doi.org/10.1002/ajpa.10233">https://doi.org/10.1002/ajpa.10233</a>).</li> </ul> <ul> <li>“<em>Dental non-metric trait data.csv</em>”: Dataset comprising 25 dental non-metric trait variables for 26 worldwide modern populations, provided in a comma-separated values file format. The data were collected by C. G. Turner II, G. R. Scott, and J. D. Irish. This individual-level dataset was artificially created from population-level trait frequency information presented in the publications: G. R. Scott, J. D. Irish, <em>Human Tooth Crown and Root Morphology </em>(Cambridge University Press, 2017) (<a href="https://doi.org/10.1017/9781316156629">https://doi.org/10.1017/9781316156629</a>); and: J. D. Irish, A. Morez, L. Girdland Flink, E. L. W. Phillips, G. R. Scott, Do dental nonmetric traits actually work as proxies for neutral genomic data? Some answers from continental- and global-level analyses. <em>Am. J. Phys. Anthropol. </em>172, 347–375 (2020) (<a href="https://doi.org/10.1002/ajpa.24052">https://doi.org/10.1002/ajpa.24052</a>).</li> </ul> <ul> <li>“<em>SNP data.txt</em>”: Dataset comprising 8,821 SNP markers for 26 worldwide modern populations, provided in a genepop file format. The data were obtained from various published sources: I. Lazaridis et al., Ancient human genomes suggest three ancestral populations for present-day Europeans. <em>Nature </em>513, 409–413 (2014) (<a href="https://doi.org/10.1038/nature13673">https://doi.org/10.1038/nature13673</a>); P. Qin, M. Stoneking, Denisovan ancestry in east Eurasian and native American populations. <em>Mol. Biol. Evol. </em>32, 2665–2674 (2015) (<a href="https://doi.org/10.1093/molbev/msv141">https://doi.org/10.1093/molbev/msv141</a>); P. Skoglund et al., Genomic insights into the peopling of the Southwest Pacific. <em>Nature </em>538, 510–513 (2016) (<a href="https://doi.org/10.1038/nature19844">https://doi.org/10.1038/nature19844</a>); M. R. Nelson et al., The Population Reference Sample, POPRES: a resource for population, disease, and pharmacological genetics research. <em>Am. J. Hum. Genet. </em>83, 347–358 (2008) (<a href="https://doi.org/10.1016/j.ajhg.2008.08.005">https://doi.org/10.1016/j.ajhg.2008.08.005</a>); J. K. Pickrell, J. K. Pritchard, Inference of population splits and mixtures from genome-wide allele frequency data. <em>PLoS Genet. </em>8, e1002967 (2012) (<a href="https://doi.org/10.1371/journal.pgen.1002967">https://doi.org/10.1371/journal.pgen.1002967</a>); A. Bergström et al., Insights into human genetic variation and population history from 929 diverse genomes. <em>Science </em>367 (2020) (<a href="https://doi.org/10.1126/science.aay5012">https://doi.org/10.1126/science.aay5012</a>); B. M. Henn et al., Genomic ancestry of North Africans supports back-to-Africa migrations. <em>PLoS Genet. </em>8, e1002397 (2012) (<a href="https://doi.org/10.1371/journal.pgen.1002397">https://doi.org/10.1371/journal.pgen.1002397</a>); S. Mallick et al., The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. <em>Nature </em>538, 201–206 (2016) (<a href="https://doi.org/10.1038/nature18964">https://doi.org/10.1038/nature18964</a>); Lao et al., Correlation between genetic and geographic structure in Europe. <em>Curr. Biol. </em>18, 1241–1248 (2008) (<a href="https://doi.org/10.1016/j.cub.2008.07.049">https://doi.org/10.1016/j.cub.2008.07.049</a>); and M. Lipson et al., Population Turnover in Remote Oceania Shortly after Initial Settlement. <em>Curr. Biol. </em>28, 1157-1165.e7 (2018) (<a href="https://doi.org/10.1016/j.cub.2018.02.051">https://doi.org/10.1016/j.cub.2018.02.051</a>).</li> </ul> <p>For population and variable names and abbreviations, see Supplementary Information in: H. Rathmann et al., Inferring human neutral genetic variation from craniodental phenotypes. PNAS Nexus.</p>
Experimental Results for "A Unified Perspective on Natural Gradient Variational Inference with Gaussian Mixture Models"
<p>This package contains the raw data / logs (fetched from WandB) for the experiments of the following publication:</p> <p>O. Arenz, P. Dahlinger, Z. Ye, M. Volpp, and G. Neumann. A unified perspective on natural gradient variational inference with gaussian mixture models. Transactions on Machine Learning Research, 2023. URL: <a href="https://openreview.net/forum?id=tLBjsX4tjs">https://openreview.net/forum?id=tLBjsX4tjs</a>.</p> <p> </p>
Migration trajectories of the diamondback moth Plutella xylostella in China inferred from population genomic variation
<p><span class="fontstyle01"><span><b>BACKGROUND</b></span></span><span class="fontstyle01"><span><b>:</b></span></span></p> <p><span class="fontstyle01"><span>The diamondback moth (DBM),</span></span><span class="fontstyle01"><span><i> Plutella xylostella</i></span></span><span class="fontstyle01"><span> (Lepidoptera: Plutellidae),</span></span><span class="fontstyle01"><span><i> </i></span></span><span class="fontstyle01"><span>is a notorious pest of cruciferous plants. In temperate areas, annual populations of DBM originate from adult migrants. However, the source populations and migration trajectories of immigrants remain unclear. Here, we investigated migration trajectories of DBM in China with genome-wide single nucleotide polymorphisms (SNPs) genotyped using double-digest RAD (ddRAD) sequencing. We first analyzed patterns of spatial and temporal genetic structure among southern source and northern recipient populations, then inferred migration trajectories into northern regions using discriminant analysis of principal components (DAPC), assignment tests and spatial kinship patterns.</span></span></p> <p><span class="fontstyle01"><span><b>RESULTS:</b></span></span></p> <p><span class="fontstyle01"><span>Temporal</span></span><span class="fontstyle01"><span><b> </b></span></span><span class="fontstyle01"><span>genetic differentiation among populations was low, indicating sources of </span></span><span class="fontstyle01"><span>recipient </span></span><span class="fontstyle01"><span>populations and migration trajectories are stable.</span></span><span class="fontstyle01"><span> Spatial genetic structure indicated three genetic clusters in the southern source populations. Assignment tests linked northern populations to the Sichuan cluster, and central-eastern populations to the South and Yunnan clusters, indicating that Sichuan populations are sources of northern immigrants and South and Yunnan populations are sources of central-eastern populations. First-order (full-sib) and second-order (half-sib) kin pairs were always found within populations, but about 35-40% of third-order (cousin) pairs were found in different populations. Closely related individuals in different populations were in about 35-40% of cases found at distances of 900 to 1500 km, while some were separated by over 2000 km.</span></span></p> <p><span class="fontstyle01"><span><b>CONCLUSION:</b></span></span></p> <p><span class="fontstyle01"><span>This study unravels seasonal migration patterns in the DBM. We demonstrate how careful sampling and population genomic analyses can be combined to help understand cryptic migration patterns in insects.</span></span></p>
Data from: A Bayesian approach for inferring the impact of a discrete character on rates of continuous-character evolution in the presence of background-rate variation
Understanding how and why rates of character evolution vary across the Tree of Life is central to many evolutionary questions; e.g., does the trophic apparatus (a set of continuous characters) evolve at a higher rate in fish lineages that dwell in reef versus non-reef habitats (a discrete character)? Existing approaches for inferring the relationship between a discrete character and rates of continuous-character evolution rely on comparing a null model (in which rates of continuous-character evolution are constant across lineages) to an alternative model (in which rates of continuous-character evolution depend on the state of the discrete character under consideration). However, these approaches are susceptible to a "straw-man" effect: the influence of the discrete character is inflated because the null model is extremely unrealistic. Here, we describe MuSSCRat, a Bayesian approach for inferring the impact of a discrete trait on rates of continuous-character evolution in the presence of alternative sources of rate variation ("background-rate variation"). We demonstrate by simulation that our method is able to reliably infer the degree of state-dependent rate variation, and show that ignoring background-rate variation leads to biased inferences regarding the degree of state-dependent rate variation in grunts (the fish group Haemulidae).
Sex-biased admixture and assortative mating shape genetic variation and influence demographic inference in admixed Cabo Verdeans
<p>Inferred ROH and IBD calls from Korunes et al (2022). bioRxiv DOI: https://doi.org/10.1101/2020.12.14.422766</p> <p>Samples originally collected and analyzed in Beleza et al. 2013, PLoS Genetics. Inferred local ancestry information can be found at <a href="https://doi.org/10.5281/zenodo.4021277">https://doi.org/10.5281/zenodo.4021277</a></p> <p>See README.txt in upload for more detailed information.</p>
Data from: A century of genetic variation inferred from a persistent soil-stored seed bank
Stratigraphic accretion of dormant propagules in soil can result in natural archives useful for studying ecological and evolutionary responses to environmental change. Few attempts have been made, however, to use soil-stored seed banks as natural archives, in part because of concerns over non-random attrition and mixed stratification. Here we examine the persistent seed bank of Schoenoplectus americanus, a foundational brackish marsh sedge, to determine whether it can serve as a resource for reconstructing historical records of demographic and population genetic variation. After assembling profiles of the seed bank from radionuclide dated soil cores, we germinated seeds to 'resurrect' cohorts spanning the 20th century. Using microsatellite markers, we assessed genetic diversity and differentiation among depth cohorts, drawing comparisons to extant plants at the study site and in nearby and more distant marshes. We found that seed density peaked at intermediate soil depths. We also detected genotypic differences among cohorts as well as between cohorts and extant plants. Genetic diversity did not decline with depth, indicating that the observed pattern of differentiation is not due to attrition. Patterns of differentiation within and among extant marshes also suggest that local populations persist as aggregates of small clones, likely reflecting repeated seedling recruitment and low immigration from admixed regional gene pools. These findings indicate that persistent and stratified soil-stored seed banks merit further consideration as resources for reconstructing decadal-to-century long records that can lend insight into the tempo and nature of ecological and evolutionary processes that shape populations over time.
Figure 4 in Determination of genetic variations between Apodemus mystacinus populations distributed in Turkey inferred from mtDNA PCR-RFLP
Figure 4. UPGMA dendrogram of the composite data by combining cytb and D-loop regions.
Dataset accompaning "Variational inference for correlated gravitational wave detector network noise"
<p>Dataset accompaning the paper "<strong>Variational inference for correlated gravitational wave detector network noise</strong>"</p> <p> </p> <h3>Raw data files</h3> <ul> <li>ET_caseA_noise.h5 (correlated noise)</li> <li>ET_caseB_noise.h5 (uncorrelated noise)</li> </ul> <p>These contain:</p> <ul> <li>raw_XYZ (the XYZ channels of ET noise)</li> <li>time (time in seconds, corresponding to raw_XYZ)</li> <li>periodogram: <ul> <li>pdgrm (of the above channels, truncated to 5-128 Hz)</li> <li>freq (in Hz)</li> </ul> </li> <li>true_psd <ul> <li>psd </li> <li>freq</li> </ul> </li> </ul> <p><em>Note</em>: case C from the manuscript utilised the case B dataset, (but with a model that does not account for the cross-spectrum). It does not have a separate dataset. </p> <p> </p> <h3><strong>Result file</strong></h3> <ul> <li>ET-CaseA-SGVB-PSD.h5</li> <li>ET-CaseB-SGVB-PSD.h5</li> <li>ET-CaseC-SGVB-PSD.h5</li> </ul> <p>These contain:</p> <ul> <li>psd_quantiles (the lower 0.05, median 0.50, upper 0.95 quantiles of 500 PSD samples)</li> <li>freq (in Hz, associated with the psd_quantiles)</li> </ul>
Code from: When and where do waterbirds need water? Inferring candidate restoration areas from spatio-temporal variation in surface water availability
Open the record for dataset details and reuse information.
Data from: A Bayesian approach for inferring the impact of a discrete character on rates of continuous-character evolution in the presence of background-rate variation
Open the record for dataset details and reuse information.
Migration trajectories of the diamondback moth Plutella xylostella in China inferred from population genomic variation
Open the record for dataset details and reuse information.
Data from: A century of genetic variation inferred from a persistent soil-stored seed bank
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.