Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,445
datasets available to search
ShareScore release 0.9.0
Dataset results
2,445 results for “Genetics: population”
Data from: Population genetic structure of Picea engelmannii, P. glauca and their previously unrecognized hybrids in the central Rocky Mountains
Areas of geographic overlap between potentially hybridizing species provide the opportunity to study interspecific gene flow and reproductive barriers. Here we identified hybrids between Picea engelmannii and P. glauca by their genetic composition at 17 microsatellite markers, and determined the broad-scale geographic distribution of hybrids in the central Rocky Mountains of North America, a geographic region where hybrids and isolation between species had not previously been studied. Parameter estimates from admixture models revealed considerable variation in ancestry within and among collection sites, suggesting that within this area of geographic overlap, the interaction of the two species varies extensively. The results document a previously unrecognized patchy distribution of hybrids between P. engelmannii and P. glauca, including locations where hybrids were not known or expected to exist. Further, the ancestry of many hybrids was consistent with multiple generations of hybridization, with probable directional backcrossing to P. engelmannii, suggesting a relatively porous species boundary. The identification and characterization of hybridization between these spruce in this region raises the question of what factors maintain barriers to gene flow in these long-lived forest trees. The current research lays the groundwork for future study of the ecological and evolutionary contexts of their hybridization, as well as of differential introgression and permeability of species boundaries.
Data from: Seven common mistakes in population genetics and how to avoid them
As the data resulting from modern genotyping tools are astoundingly complex, genotyping studies require great care in the sampling design, genotyping, data analysis and interpretation. Such care is necessary because, with data sets containing thousands of loci, small biases can easily become strongly significant patterns. Such biases may already be present in routine tasks that are present in almost every genotyping study. Here, I discuss seven common mistakes that can be frequently encountered in the genotyping literature: (i) giving more attention to genotyping than to sampling, (ii) failing to perform or report experimental randomization in the laboratory, (iii) equating geopolitical borders with biological borders, (iv) testing significance of clustering output, (v) misinterpreting Mantel's r statistic, (vi) only interpreting a single value of k and (vii) forgetting that only a small portion of the genome will be associated with climate. For every of those issues, I give some suggestions how to avoid the mistake. Overall, I argue that genotyping studies would benefit from establishing a more rigorous experimental design, involving proper sampling design, randomization and better distinction of a priori hypotheses and exploratory analyses.
Data from: Artificial barriers prevent genetic recovery of small isolated populations of a low-mobility freshwater fish
Habitat loss and fragmentation often result in small, isolated populations vulnerable to environmental disturbance and loss of genetic diversity. Low genetic diversity can increase extinction risk of small populations by elevating inbreeding and inbreeding depression, and reducing adaptive potential. Due to their linear nature and extensive use by humans, freshwater ecosystems are especially vulnerable to habitat loss and fragmentation. Although the effects of fragmentation on genetic structure have been extensively studied in migratory fish, they are less understood in low-mobility species. We estimated impacts of instream barriers on genetic structure and diversity of the low-mobility river blackfish (Gadopsis marmoratus) within five streams separated by weirs or dams constructed 45–120 years ago. We found evidence of small-scale (<13 km) genetic structure within reaches unimpeded by barriers, as expected for a fish with low mobility. Genetic diversity was lower above barriers in small streams only, regardless of barrier age. In particular, one isolated population showed evidence of a recent bottleneck and inbreeding. Differentiation above and below the barrier was greatest (FST =0.13) in this stream, but in other streams did not differ from background levels. Spatially explicit simulations suggest that short-term barrier effects would not be detected with our dataset unless effective population sizes were very small (<100). Our study highlights that, in structured populations, the ability to detect short-term genetic effects from barriers is reduced and requires more genetic markers compared to panmictic populations. We also demonstrate the importance of accounting for natural population genetic structure in fragmentation studies.
Data from: Genetic diversity and population structure of Urochloa grass accessions from Tanzania using simple sequence repeat (SSR) markers
Urochloa (syn.—Brachiaria s.s.) is one of the most important tropical forages that transformed livestock industries in Australia and South America. Farmers in Africa are increasingly interested in growing Urochloa to support the burgeoning livestock business, but the lack of cultivars adapted to African environments has been a major challenge. Therefore, this study examines genetic diversity of Tanzanian Urochloa accessions to provide essential information for establishing a Urochloa breeding program in Africa. A total of 36 historical Urochloa accessions initially collected from Tanzania in 1985 were analyzed for genetic variation using 24 SSR markers along with six South American commercial cultivars. These markers detected 407 alleles in the 36 Tanzania accessions and 6 commercial cultivars. Markers were highly informative with an average polymorphic information content of 0.79. The analysis of molecular variance revealed high genetic variation within individual accessions in a species (92%), fixation index of 0.05 and gene flow estimate of 4.77 showed a low genetic differentiation and a high level of gene flow among populations. An unweighted neighbor-joining tree grouped the 36 accessions and six commercial cultivars into three main clusters. The clustering of test accessions did not follow geographical origin. Similarly, population structure analysis grouped the 42 tested genotypes into three major gene pools. The results showed the Urochloa brizantha (A. Rich.) Stapf population has the highest genetic diversity (I = 0.94) with high utility in the Urochloa breeding and conservation program. As the Urochloa accessions analyzed in this study represented only 3 of 31 regions of Tanzania, further collection and characterization of materials from wider geographical areas are necessary to comprehend the whole Urochloa diversity in Tanzania.
Data from: Population-specific genetic modification of Huntington's disease in Venezuela
Modifiers of Mendelian disorders can provide insights into disease mechanisms and guide therapeutic strategies. A recent genome-wide association (GWA) study discovered genetic modifiers of Huntington's disease (HD) onset in Europeans. Here, we performed whole genome sequencing and GWA analysis of a Venezuelan HD cluster whose families were crucial for the original mapping of the HD gene defect. The Venezuelan HD subjects develop motor symptoms earlier than their European counterparts, implying the potential for population-specific modifiers. The main Venezuelan HD family inherits HTT haplotype hap.03, which differs subtly at the sequence level from European HD hap.03, suggesting a different ancestral origin but not explaining the earlier age at onset in these Venezuelans. GWA analysis of the Venezuelan HD cluster suggests both population-specific and population-shared genetic modifiers. Genome-wide significant signals at 7p21.2-21.1 and suggestive association signals at 4p14 and 17q21.2 are evident only in Venezuelan HD, but genome-wide significant association signals at the established European chromosome 15 modifier locus are improved when Venezuelan HD data are included in the meta-analysis. Venezuelan-specific association signals on chromosome 7 center on SOSTDC1, which encodes a bone morphogenetic protein antagonist. The corresponding SNPs are associated with reduced expression of SOSTDC1 in non-Venezuelan tissue samples, suggesting that interaction of reduced SOSTDC1 expression with a population-specific genetic or environmental factor may be responsible for modification of HD onset in Venezuela. Detection of population-specific modification in Venezuelan HD supports the value of distinct disease populations in revealing novel aspects of a disease and population-relevant therapeutic strategies.
Data from: A genome-wide assessment of genetic diversity and population structure of Korean native cattle breeds
Background: The native cattle breeds are an important genetic resource for meat and milk production throughout Asia. In Asia cattle were domesticated around 10,000 years ago and in Korea cattle are being raised since 2000 B.C. There are three native breeds of cattle in Korea viz. Brown Hanwoo, Brindle Hanwoo and Jeju Black. While one of these breeds, Brown Hanwoo, is a part of a Food and Agricultural Organization and national genetic evaluation plans, others get little attention. This study is an effort to understand and provide a detailed insight into the population structure and genetic variability of the Korean cattle breeds along with other Asian breeds using various methods. In this study we report the genetic variation and structure of the Korean cattle breeds and their comparison with five other Asian cattle breeds along with a panel of animals from European taurine, African taurine and indicine cattle breeds. Results: Asian cattle were found to be least differentiated which reflects their recent history. Amongst the Asian breeds Hainan, which is an indicine breed, had the lowest gene diversity while Yanbian had the highest followed by Mongolian and Korean cattle. Amongst the Korean breeds Brown Hanwoo had the highest diversity followed by Brindle Hanwoo and Jeju Black. The genetic diversity in Asian cattle breeds was found comparable to the European taurines and more than the African taurines and Zebu cattle. Korean cattle breed, Brown Hanwoo was consistently found to be closer to Yanbian, a Chinese cattle breed. We found low divergence and moderate levels of genetic diversity among the native Korean breeds. Indicine introgression from Hainan was seen in other Asian breeds. From Europe, Limousin, Holstein and Hereford introgression was found in Asian breeds. Conclusions: In this study we provide a genome-wide insight into the genetic history of the native cattle breeds of Korea. The outcomes of this study will help in prioritization and designing of the conservation plans.
Data from: The devil is in the details: genetic variation in introduced populations and its contributions to invasion
The influence of genetic variation on invasion success has captivated researchers since the start of the field of invasion genetics 50 years ago. We review the history of work on this question and conclude that genetic variation—as surveyed with molecular markers—appears to shape invasion rarely. Instead, there is a significant disconnect between marker assays and ecologically relevant genetic variation in introductions. We argue that the potential for adaptation to facilitate invasion will be shaped by the details of genotypes affecting phenotypes, and we highlight three areas in which we see opportunities to make powerful new insights. (i) The genetic architecture of adaptive variation. Traits shaped by large-effect alleles may be strongly impacted by founder events yet more likely to respond to selection when genetic drift is strong. Large-effect loci may be especially relevant for traits involved in biotic interactions. (ii) Cryptic genetic variation exposed during invasion. Introductions have strong potential to uncover masked variation due to alterations in genetic and ecological environments. (iii) Genetic interactions during admixture of multiple source populations. As divergence among sources increases, positive followed by increasingly negative effects of admixture should be expected. Although generally hypothesized to be beneficial during invasion, admixture is most often reported among sources of intermediate divergence, supporting the possibility that incompatibilities among divergent source populations might be limiting their introgression. Finally, we note that these details of invasion genetics can be coupled with comparative demographic analyses to link genetic changes to the evolution of invasiveness itself.
Data from: Effects of sampling close relatives on some elementary population genetics analyses
Many molecular ecology analyses assume the genotyped individuals are sampled at random from a population and thus are representative of the population. Realistically, however, a sample may contain excessive close relatives (ECR) because, for example, localized juveniles are drawn from fecund species. Our knowledge is limited about how ECR affect the routinely conducted elementary genetics analyses, and how ECR are best dealt with to yield unbiased and accurate parameter estimates. This study quantifies the effects of ECR on some popular population genetics analyses of marker data, including the estimation of allele frequencies, F-statistics, expected heterozygosity (He), effective and observed numbers of alleles, and the tests of Hardy-Weinberg equilibrium (HWE) and linkage equilibrium (LE). It also investigates several strategies for handling ECR to mitigate their impact and to yield accurate parameter estimates. My analytical work, assisted by simulations, shows that ECR have large and global effects on all of the above marker analyses. The naïve approach of simply ignoring ECR could yield low-precision and often biased parameter estimates, and could cause too many false rejections of HWE and LE. The bold approach, which simply identifies and removes ECR, and the cautious approach, which estimates target parameters (e.g. He) by accounting for ECR and using naïve allele frequency estimates, eliminate the bias and the false HWE and LE rejections, but could reduce estimation precision substantially. The likelihood approach, which accounts for ECR in estimating allele frequencies and thus target parameters relying on allele frequencies, usually yields unbiased and the most accurate parameter estimates. Which of the four approaches is the most effective and efficient may depend on the particular marker analysis to be conducted. The results are discussed in the context of using marker data for understanding population properties and marker properties.
Data from: Conflict bear translocation: Investigating population genetics and fate of bear translocation in Dachigam National Park, Jammu and Kashmir, India
The Asiatic black bear population in Dachigam landscape, Jammu and Kashmir is well recognized as one of the highest density bear populations in India. Increasing incidences of bear-human interactions and the resultant retaliatory killings by locals have become a serious threat to the survivorship of black bears in the Dachigam landscape. The Department of Wildlife Protection in Jammu and Kashmir has been translocating bears involved in conflicts, henceforth 'conflict bears' from different sites in Dachigam landscape to Dachigam National Park as a flagship activity to mitigate conflicts. We undertook this study to investigate the population genetics and the fate of bear translocation in Dachigam National Park. We identified 109 unique genotypes in an area of ca. 650 km2 and observed bear population under panmixia that showed sound genetic variability. Molecular tracking of translocated bears revealed that mostly bears (7 out of 11 bears) returned to their capture sites, possibly due to homing instincts or habituation to the high quality food available in agricultural croplands and orchards, while only four bears remained in Dachigam National Park after translocation. Results indicated that translocation success was most likely to be season dependent as bears translocated during spring and late autumn returned to their capture sites, perhaps due to the scarcity of food inside Dachigam National Park while bears translocated in summer remained in Dachigam National Park due to availability of surplus food resources. Thus, the current management practices of translocating conflict bears, without taking into account spatio-temporal variability of food resources in Dachigam landscape seemed to be ineffective in mitigating conflicts on a long-term basis. However, the study highlighted the importance of molecular tracking of bears to understand their movement patterns and socio-biology in tough terrains like Dachigam landscape.
Data from: Prediction of genetic values of quantitative traits with epistatic effects in plant breeding populations
Though epistasis has long been postulated to play a critical role in genetic regulation of important pathways as well as provide a major source of variation in the process of speciation, the importance of epistasis for genomic selection in the context of plant breeding is still being debated. In this paper, we report the results on the prediction of genetic values with epistatic effects for 280 accessions in the Nebraska Wheat Breeding Program using adaptive mixed LASSO. The development of adaptive mixed LASSO, originally designed for association mapping, for the context of genomic selection is reported. The results show that adaptive mixed LASSO can be successfully applied to the prediction of genetic values while incorporating both marker main effects and epistatic effects. Especially, the prediction accuracy is substantially improved by the inclusion of two-locus epistatic effects (more than one fold in some cases as measured by cross validation correlation coefficient), which is observed for multiple traits and planting locations. This points to significant potential in using non-additive genetic effects for genomic selection in crop breeding practices.
Data from: Population genetic structure of the Pocillopora damicornis morphospecies along Ningaloo Reef, Western Australia
The effective management of a coral reef system relies on a detailed understanding of the population structure of dominant habitat-forming species. For some corals, however, high levels of phenotypic plasticity have made species delineation based on morphological characteristics alone unreliable, suggesting that previous studies of population genetic structure may have been influenced by the inclusion of multiple genetic lineages in the analyses. We examined the population structure of the Pocillopora damicornis morphospecies along the World Heritage Ningaloo Coast, Western Australia, and recovered 2 mitochondrial haplotypes from sympatrically occurring colonies possessing morphological characteristics consistent with taxonomic classification of P. damicornis. Despite a high degree of genetic differentiation between these lineages, we detected low levels of unidirectional admixture between them, suggesting that reproductive barriers are not fully developed. We found dual modes of reproduction for both lineages with considerable variation in the contribution of sexual reproduction among sample sites. Lastly, we identified a high dispersal potential of sexually produced propagules in the most common lineage with positive spatial autocorrelation detected over distances up to 60 km. Based on these results, it appears that populations of P. damicornis have a high capacity to recover from environmental perturbations as long as the effects of disturbances are patchy across Ningaloo Reef.
Data from: Optimizing the genetic composition of a translocation population: incorporating constraints and conflicting objectives
Translocations of threatened species can reduce the risk of extinction from a catastrophic event. For plants, translocation consists of moving individuals, seeds, or cuttings from a native (source) population to a new site. Ideally a translocation population would be genetically diverse and consist of fit founding individuals. In practice, there are challenges to designing such a population, including constraints on the availability of material, and tradeoffs between different goals. We present an approach for designing a translocation population that identifies sets of founders that are optimized according to multiple criteria (e.g., genetic diversity), while also conforming to constraints on the representation of different founders (e.g., propagation success). It uses flexible inputs, including SNP genotypes, matrices of similarity between individuals, and vectors of phenotype data. We apply the approach to a critically endangered plant, Hibbertia puberula subsp. glabrescens (Dilleniaceae), which was genotyped at thousands of SNP loci. The goals of minimizing genetic similarity among the founding individuals and maximizing genetic diversity were largely complementary – populations optimized for one of these criteria were near-optimal for the other. We also performed analyses in which we minimized genetic similarity among founding individuals while imposing selection (against hypothetical deleterious alleles, and against undesirable phenotypes, respectively), and here characterized sharp tradeoffs. This is useful in allowing the benefits of selection to be weighed against 'costs' in terms of genetic similarity. In sum, we present an approach for designing a translocation population that allows flexible inputs, the imposition of realistic constraints, and examination of conflicting goals.
Data from: The role of ecological factors in determining phylogeographic and population genetic structure of two sympatric island skinks (Plestiodon kishinouyei and P. stimpsonii)
We conducted comparative phylogeographic and population genetic analyses of Plestiodon kishinouyei and P. stimpsonii, two sympatric skinks endemic to islands in the southern Ryukyus, to explore different factors that have influenced population structure. Previous phylogenetic studies using partial mitochondrial DNA (mtDNA) indicate similar divergence times from their respective closest relatives, suggesting that differences in population structure are driven by intrinsic attributes of either species rather than the common set of extrinsic factors that both presumably have been exposed to throughout their history. In this study, analysis of mtDNA sequences and microsatellite polymorphism demonstrate contrasting patterns of phylogeography and population structure: P. kishinouyei exhibits a lower genetic variability and lower genetic differentiation among islands than P. stimpsonii, consistent with recent population expansion. However, historical demographic analyses indicate that the relatively high genetic uniformity in P. kishinouyei is not attributable to recent expansion. We detected significant isolation-by-distance patterns among P. kishinouyei populations on the land bridge islands, but not among P. stimpsonii populations occurring on those same islands. Our results suggest that P. kishinouyei populations have maintained gene flows across islands until recently, probably via ephemeral Quaternary land bridges. The lower genetic variability in P. kishinouyei may also indicate smaller effective population sizes on average than that of P. stimpsonii. We interpret these differences as a consequence of ecological divergence between the two species, primarily in trophic level and habitat preference.
Data from: Genetics of plasminogen activator inhibitor-1 (PAI-1) in a Ghanaian population
Plasminogen activator inhibitor 1 (PAI-1), a major modulator of the fibrinolytic system, is an important factor in cardiovascular disease (CVD) susceptibility and severity. PAI-1 is highly heritable, but the few genes associated with it explain only a small portion of its variation. Studies of PAI-1 typically employ linear regression to estimate the effects of genetic variants on PAI-1 levels, but PAI-1 is not normally distributed, even after transformation. Therefore, alternative statistical methods may provide greater power to identify important genetic variants. Additionally, most genetic studies of PAI-1 have been performed on populations of European descent, limiting the generalizability of their results. We analyzed >30,000 variants for association with PAI-1 in a Ghanaian population, using median regression, a non-parametric alternative to linear regression. Three variants associated with median PAI-1, the most significant of which was in the gene arylsulfatase B (ARSB) (p = 1.09 x 10−7). We also analyzed the upper quartile of PAI-1, the most clinically relevant part of the distribution, and found 19 SNPs significantly associated in this quartile. Of note an association was found in period circadian clock 3 (PER3). Our results reveal novel associations with median and elevated PAI-1 in an understudied population. The lack of overlap between the two analyses indicates that the genetic effects on PAI-1 are not uniform across its distribution. They also provide evidence of the generalizability of the circadian pathway's effect on PAI-1, as a recent meta-analysis performed in Caucasian populations identified another circadian clock gene (ARNTL).
Data from: Evaluation of genetic diversity and population structure of five Chinese indigenous donkey breeds using microsatellite markers
China had the largest population of raising donkeys in the world, however the number of Chinese indigenous donkey decreased dramatically due to the increase of agriculture mechanization in the last century. The species has still been important in China because of its edible and medical value, therefore the survey on its genetic diversity in China is necessary for its conservation and utilization. In this study, 15 microsatellite markers were used to evaluate genetic diversity and population structure of five Chinese indigenous donkey breeds. The mean values of expected heterozygosity, allelic richness, and total number of alleles for all the tested Chinese donkeys were 0.70, 6.04, and 6.28 respectively, suggesting that the genetic diversity of Chinese indigenous donkeys is rich. The Bayesian analysis and principal component analysis plot yielded the same clustering result, which revealed that Guanzhong donkey was the most differentiated breed in all detected samples, and Jinnan (JN) and Guangling (GL) were genetically closed together. Additionally, our results indicated that the heterozygote deficit was severe in two Chinese indigenous donkey breeds (GL and JN), and it warned us that animal conservation activities on this species should be considered carefully in near future.
Data from: Genetic diversity and drivers of dwarfism in extinct island emu populations
Australia's iconic emu (Dromaius novaehollandiae novaehollandiae) is the only living representative of its genus, but fossil evidence and reports from early European explorers suggest that three island forms (at least two of which were dwarfs) became extinct during the 19th century. While one of these - the King Island emu - has been found to be conspecific with Australian mainland emus, little is known about how the other two forms - Kangaroo Island and Tasmanian emus - relate to the others, or even the size of Tasmanian emus. We present a comprehensive genetic and morphological analysis of Dromaius diversity, including data from one of the few definitively genuine Tasmanian emu specimens known. Our genetic analyses suggest that all the island populations represent sub-populations of mainland D. novaehollandiae. Further, the size of island emus and those on the mainland appears to scale linearly with island size but not time since isolation, suggesting that island size—and presumably concomitant limitations on resource availability—may be a more important driver of dwarfism in island emus, though its precise contribution to emu dwarfism remains to be confirmed.
Genetic Heterogeneity of Early-Onset Colorectal Cancer in Asian Populations: A Comprehensive Systematic Review
<p>Genetic Heterogeneity of Early-Onset Colorectal Cancer in Asian Populations: A Comprehensive Systematic Review</p>
Figure 3 from: Patterson BD, Webala PW, Lavery TH, Agwanda BR, Goodman SM, Kerbis Peterhans JC, Demos TC (2020) Evolutionary relationships and population genetics of the Afrotropical leaf-nosed bats (Chiroptera, Hipposideridae). ZooKeys 929: 117-161. https://doi.org/10.3897/zookeys.929.50240
Figure 3 Substitution network plots for Afrotropical hipposiderids ADoryrhinaBMacronycteris.
Figure 3 from: Zhao L, Wang S, Qu F, Liu Z, Gao T (2022) A genetic assessment of the population structure and demographic history of Odontamblyopus lacepedii (Perciformes, Amblyopinae) from the northwestern Pacific. ZooKeys 1088: 1-15. https://doi.org/10.3897/zookeys.1088.70860
Figure 3 Phylogenetic network of all haplotypes.
Figure 1 from: Zhao L, Wang S, Qu F, Liu Z, Gao T (2022) A genetic assessment of the population structure and demographic history of Odontamblyopus lacepedii (Perciformes, Amblyopinae) from the northwestern Pacific. ZooKeys 1088: 1-15. https://doi.org/10.3897/zookeys.1088.70860
Figure 1 Sampling locations of Odontamblyopus lacepedii in the present study.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.