Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
131
datasets available to search
ShareScore release 0.7.1
Dataset results
131 results for “genetic loci”
FIGURE 1 in Population genetics of the endangered catfish Pseudoplatystoma magdaleniatum (Siluriformes: Pimelodidae) based on species-specific microsatellite loci
FIGURE 1 | Sampling sites of Pseudoplatystoma magdaleniatum in the Magdalena-Cauca basin.
Integrative multi-ancestry genetic analysis of gene regulation in coronary arteries prioritizes disease risk loci
<p>All full-sample files contain results generated in coronary artery tissue from 138 American adults. Subset analyses utilized 80 individuals selected from the original 138. Scripts accompanying some of these data in downstream analyses can be viewed on our Github, which also contains a link to the current version of our accompanying manuscript: https://github.com/MillerLab-CPHG/CAD_QTL</p> <p>Full summary statistics for eQTL associations using mixQTL (https://github.com/hakyimlab/mixqtl/wiki) by chromosome are located in UVA_coronary_mixQTL_sumstats_by_chromosome.zip</p> <p>Full summary statistics for eQTL associations using mixQTL in the subset of 100% European-ancestry study sample members by chromosome are located in Hodonsky_mixQTL_Euro_sumstats.zip</p> <p>Full summary statistics for eQTL associations using mixQTL in the genetically diverse downsampled subset by chromosome are located in Hodonsky_mixQTL_downsample_sumstats.zip</p> <p>Full summary statistics for nominal pass for all genes identified as significant in the permutation pass using QTLtools (https://qtltools.github.io/qtltools/) adjusting for local ancestry by gene by chromosome are located in Local_ancestry_UVA_coronary_QTLtools_nominal_sumstats.zip</p> <p>Full summary statistics for sQTL associations with splice junctions using QTLtools by gene are located in sQTL_results_UVA_coronary_full_sumstats.zip</p>
Genetic and functional analysis of Raynaud’s syndrome implicates loci in vasculature and immunity
Open the record for dataset details and reuse information.
Population genetic structure of Nephrops norvegicus from the Adriatic Sea inferred using microsatellite loci
Open the record for dataset details and reuse information.
Modelling the genetic aetiology of complex disease: human-mouse conservation of noncoding features and disease-associated loci
Open the record for dataset details and reuse information.
Patterns of recent natural selection on genetic loci associated with sexually differentiated human body size and shape phenotypes
Open the record for dataset details and reuse information.
Genome-wide association mapping to identify genetic loci for cold tolerance and cold recovery during germination in rice
Open the record for dataset details and reuse information.
Data from: Genetic population structure and variation at phenology-related loci in anadromous Arctic char (Salvelinus alpinus)
The Arctic will be especially affected by climate change, resulting in altered seasonal timing. Anadromous Arctic char (Salvelinus alpinus) is strongly influenced by sea surface temperature (SST) delimiting time periods available for foraging in the sea. Recent studies of salmonid species have shown variation at phenology-related loci associated with timing of migration and spawning. We contrasted genetic population structure at 53 SNPs versus four phenology-related loci among 15 anadromous Arctic char populations from Western Greenland and three outgroup populations. Among anadromous populations, the time period available for foraging at sea (> 2oC) ranges from a few weeks to several months, motivating two research questions: 1) Is population structure compatible with possibilities for evolutionary rescue of anadromous populations during climate change? 2) Does selection associated with latitude or SST regimes act on phenology-related loci? In Western Greenland, strong isolation-by-distance at SNPs was observed and spatial autocorrelation analysis showed genetic patch size up to 450 km, documenting contingency and gene flow among populations. Outlier tests provided no evidence for selection at phenology-related loci. However, in Western Greenland, mean allele length at OtsClock1b was positively associated with the time of year when SST first exceeded 2oC and negatively associated with duration of the period where SST exceeded 2oC. This is consistent with local adaptation for making full use of the time period available for foraging in the sea. Current adaptation may become maladaptive under climate change, but long-distance connectivity of anadromous populations could redistribute adaptive variation across populations and lead to evolutionary rescue.
Data from: Sex without sex chromosomes: genetic architecture of multiple loci independently segregating to determine sex ratios in the copepod Tigriopus californicus
Sex determining systems are remarkably diverse and may evolve rapidly. Polygenic sex determination systems are predicted to be transient and evolutionarily unstable yet examples have been reported across a range of taxa. Here we provide the first direct evidence of polygenic sex determination in Tigriopus californicus, a harpacticoid copepod with no heteromorphic sex chromosomes. Using genetically distinct inbred lines selected for male- and female-biased clutches, we generated a genetic map with 39 SNPs across 12 chromosomes. Quantitative trait locus mapping of sex ratio phenotype (the proportion of male offspring produced by an F2 female) in four F2 families revealed six independently segregating quantitative trait loci on five separate chromosomes, explaining 19% of the variation in sex ratios. The sex ratio phenotype varied among loci across chromosomes in both direction and magnitude, with the strongest phenotypic effects on chromosome 10 moderated to some degree by loci on four other chromosomes. For a given locus, sex ratio phenotype varied in magnitude for individuals derived from different dam lines. These data, together with the environmental factors known to contribute to sex determination, characterize the underlying complexity and potential lability of sex determination, and confirm the polygenic architecture of sex determination in T. californicus.
Data from: Squamate Conserved Loci (SqCL): a unified set of conserved loci for phylogenomics and population genetics of squamate reptiles
The identification of conserved loci across genomes, along with advances in target capture methods and high-throughput sequencing, has helped spur a phylogenomics revolution by enabling researchers to gather large numbers of homologous loci across clades of interest with minimal upfront investment in locus design. Target capture for vertebrate animals is currently dominated by two approaches – anchored hybrid enrichment (AHE) and ultraconserved elements (UCE) – and both approaches have proven useful for addressing questions in phylogenomics, phylogeography, and population genomics. However, these two sets of loci have minimal overlap with each other; moreover, they do not include many traditional loci that that have been used for phylogenetics. Here, we combine across UCE, AHE, and traditional phylogenetic gene locus sets to generate the Squamate Conserved Loci (SqCL) set, a single integrated probe set that can generate high-quality and highly complete data across all three loci types. We use these probes to generate data for 44 phylogenetically-disparate taxa that collectively span approximately 33% of terrestrial vertebrate diversity. Our results generated an average of 4.29 Mb across 4709 loci per individual, of which an average of 2.99 Mb was sequenced to high enough coverage (≥10×) to use for population genetic analyses. We validate the utility of these loci for both phylogenomic and population genomic questions, provide a comparison among these locus sets of their relative usefulness, and suggest areas for future improvement.
Data from: Population genetic analyses using 10 new polymorphic microsatellite loci confirms genetic subdivision within the olm, Proteus anguinus
We provide a comparative population genetic study of the elusive amphibian, Proteus anguinus, by comparing the genetic diversity and divergence among four cave populations (96 individuals) sampled in the Dinaric Karst of Croatia. We developed 10 variable microsatellite markers using pyrosequencing and applied them to the four selected populations belonging to four different cave systems. The results showed strong genetic differentiation between the four caves corroborating with previous findings suggesting that Proteus might comprises several unrecognized taxa. Our results confirmed that gene flow should be high within the caves, whereas it is low between hydrographic systems since geological periods. Finally, we conclude that the high genetic subdivision suggests the necessity of treating the four studied Proteus populations as evolutionary significant units.
Data from: SNPs selected by information content outperform randomly selected microsatellite loci for delineating genetic identification and introgression in the endangered dark European honeybee (Apis mellifera mellifera)
The honeybee (Apis mellifera) has been threatened by multiple factors, including pests and pathogens, pesticides, and loss of locally adapted gene complexes due to replacement and introgression. In western Europe, the genetic integrity of the native A.m. mellifera (M-lineage) is endangered due to trading and intensive queen breeding with commercial subspecies of eastern European ancestry (C-lineage). Effective conservation actions require reliable molecular tools to identify purebred A.m. mellifera colonies. Microsatellites have been preferred for identification of A.m. mellifera stocks across conservation centers. However, owing to high-throughput, easy transferability between laboratories and low genotyping error, SNPs promise to become popular. Here, we compared the resolving power of a widely utilized microsatellite dataset to detect structure and introgression with that of different datasets that combine a variable number of SNPs selected for their information content and genomic proximity to the microsatellites. Contrary to every SNP dataset, microsatellites were unable to clearly separate the two European lineages in the PCA space. Mean introgression proportions were identical across the two marker types, although at the individual level microsatellites' performance was relatively poor at the upper range of introgression, a result reflected by their lower precision. Although mean accuracy was relatively high across datasets (>91%), microsatellites were the least accurate and the top-ranked informative 144 SNPs were the most accurate. Comparisons amongst the SNP datasets showed that those combining SNPs flanking microsatellites performed worst. Our results suggest that SNPs are more powerful for identification of A.m. mellifera colonies, especially when they are selected by information content.
Data from: The first set of universal nuclear protein-coding loci markers for avian phylogenetic and population genetic studies
Multiple nuclear markers provide genetic polymorphism data for molecular systematics and population genetic studies. They are especially required for the coalescent-based analyses that can be used to accurately estimate species trees and infer population demographic histories. However, in avian evolutionary studies, these powerful coalescent-based methods are hindered by the lack of a sufficient number of markers. In this study, we designed PCR primers to amplify 136 nuclear protein-coding loci (NPCLs) by scanning the published Red Junglefowl (Gallus gallus) and Zebra Finch (Taeniopygia guttata) genomes. To test their utility, we amplified these loci in 41 bird species representing 23 Aves orders. The sixty-three best-performing NPCLs, based on high PCR success rates, were selected which had various mutation rates and were evenly distributed across 17 avian autosomal chromosomes and the Z chromosome. To test phylogenetic resolving power of these markers, we conducted a Neoavian phylogenies analysis using 63 concatenated NPCL markers derived from 48 whole genomes of birds. The resulting phylogenetic topology, to a large extent, is congruence with results resolved by previous whole genome data. To test the level of intraspecific polymorphism in these makers, we examined the genetic diversity in four populations of the Kentish Plover (Charadrius alexandrinus) at 17 of NPCL markers chosen at random. Our results showed that these NPCL markers exhibited a level of polymorphism comparable with mitochondrial loci. Therefore, this set of pan-avian nuclear protein-coding loci has great potential to facilitate studies in avian phylogenetics and population genetics.
Data from: Genetic variation at aryl hydrocarbon receptor (AHR) loci in populations of Atlantic killifish (Fundulus heteroclitus) inhabiting polluted and reference habitats
Background: The non-migratory killifish Fundulus heteroclitus inhabits clean and polluted environments interspersed throughout its range along the Atlantic coast of North America. Several populations of this species have successfully adapted to environments contaminated with toxic aromatic hydrocarbon pollutants such as polychlorinated biphenyls (PCBs). Previous studies suggest that the mechanism of resistance to these and other "dioxin-like compounds" (DLCs) may involve reduced signaling through the aryl hydrocarbon receptor (AHR) pathway. Here we investigated gene diversity and evidence for positive selection at three AHR-related loci (AHR1, AHR2, AHRR) in F. heteroclitus by comparing alleles from seven locations ranging over 600 km along the northeastern US, including extremely polluted and reference estuaries, with a focus on New Bedford Harbor (MA, USA), a PCB Superfund site, and nearby reference sites. Results: We identified 98 single nucleotide polymorphisms within three AHR-related loci among all populations, including synonymous and nonsynonymous substitutions. Haplotype distributions were spatially segregated and F-statistics suggested strong population genetic structure at these loci, consistent with previous studies showing strong population genetic structure at other F. heteroclitus loci. Genetic diversity at these three loci was not significantly different in contaminated sites as compared to reference sites. However, for AHR2 the New Bedford Harbor population had significant FST values in comparison to the nearest reference populations. Tests for positive selection revealed ten nonsynonymous polymorphisms in AHR1 and four in AHR2. Four nonsynonymous SNPs in AHR1 and three in AHR2 showed large differences in base frequency between New Bedford Harbor and its reference site. Tests for isolation-by-distance revealed evidence for non-neutral change at the AHR2 locus. Conclusion: Together, these data suggest that F. heteroclitus populations in reference and polluted sites have similar genetic diversity, providing no evidence for strong genetic bottlenecks for populations in polluted locations. However, the data provide evidence for genetic differentiation among sites, selection at specific nucleotides in AHR1 and AHR2, and specific AHR2 SNPs and haplotypes that are associated with the PCB-resistant phenotype in the New Bedford Harbor population. The results suggest that AHRs, and especially AHR2, may be important, recurring targets for selection in local adaptation to dioxin-like aromatic hydrocarbon contaminants.
Data from: Heterogeneity in genetic diversity among non-coding loci fails to fit neutral coalescent models of population history
Inferring aspects of the population histories of species using coalescent analyses of non-coding nuclear DNA has grown in popularity. These inferences, such as divergence, gene flow, and changes in population size, assume that genetic data reflect simple population histories and neutral evolutionary processes. However, violating model assumptions can result in a poor fit between empirical data and the models. We sampled 22 nuclear intron sequences from at least 19 different chromosomes (a genomic transect) to test for deviations from selective neutrality in the gadwall (Anas strepera), a Holarctic duck. Nucleotide diversity among these loci varied by nearly two orders of magnitude (from 0.0004 to 0.029), and this heterogeneity could not be explained by differences in substitution rates. Using two different coalescent methods to infer models of population history and then simulating neutral genetic diversity under these models, we found that the among-locus heterogeneity in nucleotide diversity was significantly higher than expected for these simple models. Defining more complex models of population history demonstrated that a pre-divergence bottleneck was also unlikely to explain this heterogeneity. However, both selection and interspecific hybridization could account for the heterogeneity observed among loci. Regardless of the cause of the deviation, our results illustrate that violating key assumptions of coalescent models can mislead inferences of population history.
Data from: Comparative population genetic analysis of bocaccio rockfish Sebastes paucispinis using anonymous and gene-associated simple sequence repeat loci
Comparative population genetic analyses of traditional and emergent molecular markers aid in determining appropriate use of new technologies. The bocaccio rockfish Sebastes paucispinis is a high-gene-flow marine species off the west coast of North America that experienced strong population decline over the past three decades. We used 18 anonymous and 13 gene associated simple sequence repeat loci (EST-SSRs) to characterize range-wide population structure with temporal replicates. No FST-outliers were detected using the LOSITAN program, suggesting that neither balancing nor divergent selection affected the loci surveyed. Consistent hierarchical structuring of populations by geography or year class was not detected regardless of marker class. The EST-SSRs were less variable than the anonymous SSRs, but no correlation between FST and variation or marker class was observed. General Linear Model analysis showed that low EST-SSR variation was attributable to low mean repeat number. Comparative genomic analysis with Gasterosteus aculeatus, Takifugu rubripes, and Oryzias latipes showed consistently lower repeat number in EST-SSRs than SSR loci that were not in ESTs. Purifying selection likely imposed functional constraints on EST-SSRs resulting in low repeat numbers that affected diversity estimates, but did not affect the observed pattern of population structure.
Data from: Multilocus approaches for the measurement of selection on correlated genetic loci
The study of ecological speciation is inherently linked to the study of selection. Methods for estimating phenotypic selection within a generation based on associations between trait values and fitness (e.g. survival) of individuals are established. These methods attempt to disentangle selection acting directly on a trait from indirect selection caused by correlations with other traits via multivariate statistical approaches (i.e. inference of selection gradients). The estimation of selection on genotypic or genomic variation could also benefit from disentangling direct and indirect selection on genetic loci. However, achieving this goal is difficult with genomic data because the number of potentially correlated genetic loci (p) is very large relative to the number of individuals sampled (n). In other words, the number of model parameters exceeds the number of observations (p ≫ n). We present simulations examining the utility of whole-genome regression approaches (i.e. Bayesian sparse linear mixed models) for quantifying direct selection in cases where p ≫ n. Such models have been used for genome-wide association mapping and are common in artificial breeding. Our results show they hold promise for studies of natural selection in the wild and thus of ecological speciation. But we also demonstrate important limitations to the approach and discuss study designs required for more robust inferences.
FIGURE 4 in A molecular phylogenetic study on South Korean Tettigonia species (Orthoptera: Tettigoniidae) using five genetic loci: The possibility of multiple allopatric speciation
FIGURE 4. Inter- (gray) and intraspecific (open) genetic differences in Tettigonia species for CO1 calculated using the pdistance method and treatment of pairwise deletion for gaps with the range of genetic difference within clusters. The box plot displays the median (internal transverse thick line) and interquartile range (box). Short lines indicate maximum and minimum genetic differences. Asterisk denotes a sequence from NCBI; T. viridissima, JN609414–JN609420; T. hispania, EF515121; T. chinensis, HQ609468–HQ609470. (JJ-TU = Jeju Island population of T. ussuriana; JS-TU = Jeongseon population of T. ussuriana; PC-TU = Pyeongchang population of T. ussuriana; MJ-TU = Muju population of T. ussuriana; MG-TU = Mungyeong population of T. ussuriana)
FIGURE 3 in A molecular phylogenetic study on South Korean Tettigonia species (Orthoptera: Tettigoniidae) using five genetic loci: The possibility of multiple allopatric speciation
FIGURE 3. Neighbor-joining tree inferred from the concatenated dataset of all five genetic loci: CO1, CO2, ND1, TA1, and ITS2. Neighbor-joining (left) and parsimony (right) bootstrap values are indicated above internodes; Bayesian posterior probabilities are shown below internodes.
FIGURE 2 in A molecular phylogenetic study on South Korean Tettigonia species (Orthoptera: Tettigoniidae) using five genetic loci: The possibility of multiple allopatric speciation
FIGURE 2. Neighbor-joining (A), parsimony (B), and Bayesian inference (C) trees inferred from the combined dataset of three mtDNA loci (CO1 + CO2 + ND1). Numbers next to nodes are bootstrap or posterior probability values.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.