Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,444
datasets available to search
ShareScore release 0.7.1
Dataset results
2,444 results for “population genetics”
State Water Project, Genetic Determination of Population of Origin 2011-2024
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Central Valley Project, Genetic Determination of Population of Origin 2011-2024
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Sacramento trawl – Genetic Determination of Population of Origin 2017-2023
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Chipps Island trawl – Genetic Determination of Population of Origin 2017-2023
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Genetic assignments for Spring Evolutionary Significant Unit reanalysis, Central Valley Chinook Salmon populations, CA, 2011-2024
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is difficult to visually distinguish individuals from the different Evolutionarily Significant Units (ESU). As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to genetic lineage; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program compliance monitoring programs. The genetic lineage was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central V
The genetic population structure of Lake Tanganyika's Lates species flock, an endemic radiation of pelagic top predators
<p>Data associated with the manuscript "The genetic population structure of Lake Tanganyika’s Lates species flock, an endemic radiation of pelagic top predators," where we investigate the genetic population structure of the four endemic <em>Lates </em>species in Lake Tanganyika.</p> <p><strong>Abstract</strong>: Life history traits are important in shaping gene flow within species and can thus determine whether a species exhibits genetic homogeneity or population structure across its range. Understanding genetic connectivity plays a crucial role in species conservation decisions, and genetic connectivity is an important component of modern fisheries management in fishes exploited for human consumption. In this study, we investigated the population genetics of four endemic <em>Lates</em> species of Lake Tanganyika (<em>Lates stappersii</em>, <em>L. microlepis</em>, <em>L. mariae</em> and <em>L. angustifrons</em>), using reduced-representation genomic sequencing methods. We find the four species to be strongly differentiated from one another, with no evidence for contemporary admixture. We also find evidence for high levels of genetic structure within <em>L. mariae</em>, with the majority of individuals from the most southern sampling site forming a genetic group distinct from the individuals at other sampling sites<em>.</em> We find evidence for much weaker structure within the other three species, <em>L. stappersii,</em> <em>L. microlepis</em>, and <em>L. angustifrons</em>, although small and unbalanced sample sizes and imprecise geographic sampling locations may hinder our ability to detect weak population structure. We call for further research into the origins of the genetic differentiation that we observe in these four species, particularly that of <em>L. mariae</em>, which may be important for the conservation and management of this species.</p> <p>Code associated with the analysis of these data can be found on GitHub at <a href="https://github.com/jessicarick/lates-popgen">https://github.com/jessicarick/lates-popgen</a>.</p>
Data from Neutral genetic structuring of pathogen populations during rapid adaptation
<p><strong>Datasets and temporary dataframes relating to the article "Neutral genetic structuring of pathogen populations during rapid adaptation".</strong></p> <p>These datasets and temporary dataframes are necessary to run the scripts from the public GitLab repository: <a href="https://gitlab.com/saubin.meline/neutral-genetic-structuring-adaptation">https://gitlab.com/saubin.meline/neutral-genetic-structuring-adaptation</a>. Please refer to this public GitLab repository for the latest version of the codes and to perform all analyses presented in the article.</p> <p>Original datasets from the demogenetic model:</p> <ul> <li>Output_RandomDesign.txt</li> <li>Output_RegularDesign_With_host_alternation.txt</li> <li>Output_RegularDesign_Without_host_alternation.txt</li> <li>Output_RandomDesign_Mnull_Medoid_With_host_alternation.txt</li> <li>Output_RandomDesign_Mnull_Medoid_Without_host_alternation.txt</li> </ul> <p>All remaining files correspond to temporary dataframes generated by the scripts in the GitLab repository, provided here for reproducibility of the results and to save time at certain time-consuming scripts.</p>
Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat
<p>SNPs obtained by UNEAK pipeline for <em>Habromys schmidlyi </em>and <em>Reithrodontomys microdon</em>. </p> <p>Pleae cite as: </p> <p>Colunga-Salas P., T Marines-Macías, G Hernández-Canchola, S Barbosa, C Ramírez, JB Searle, L León-Paniagua. 2022. <strong>Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat</strong>. Mammalian Reasearch. Doi: 10.1007/s13364-022-00667-x</p>
Central Valley Project, Genetic Determination of Population of Origin 2011-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Knights Landing, California Department of Fish and Wildlife, Genetic Determination of Population of Origin 2017 through 2019
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Sacramento trawl, Delta Juvenile Fish Monitoring Program, Genetic Determination of Population of Origin 2017-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Chipps Island trawl, Delta Juvenile Fish Monitoring Program, Genetic Determination of Population of Origin 2017-2021
Central Valley Chinook Salmon populations differ in their Endangered Species Act listing status. It is often difficult to distinguish individuals from the different Evolutionarily Significant Units. As such, many of the salmon monitoring and evaluation efforts in the Central Valley and San Francisco Bay-Delta are hampered by uncertainty about population (stock) identification and proportional effects of management actions (Dekar et al. 2013; IEP 2019). Studies have identified that the current identification method (length-at-date models) of juvenile Chinook salmon (Fisher 1992) captured in the watershed vary in their accuracy, particularly for spring-run (NMFS 2013; Harvey et al. 2014; Merz et al. 2014). The inaccuracy of the size-based methods is likely due to differences in fish distribution during early rearing, habitat-specific growth rates, and inter-annual variability in temperatures and food availability that lead to overlap in size ranges among stocks. The primary objective of this project was the genetic classification (to race; Evolutionary Significant Unit) of Chinook Salmon captured from State Water Project and Central Valley Project fish protection facilities and Interagency Ecological Program monitoring programs. The population-of-origin was determined for sampled fish by comparing their genotypes to reference genetic baselines. Genetic methods, having less statistical uncertainty that size-based models for population identification, were intended to directly target (and reduce) one source of uncertainty in the estimation of loss (take) from water diversions (operations) and develop the information necessary for understanding stock-specific distribution, habitat utilization, abundance, and life history variation. This project supports recommendations from the Interagency Ecological Program’s Salmon and Sturgeon Assessment of Indicators by Life Stage and Interagency Ecological Program Science Agenda efforts to improve Central Valley salmonid monitoring
Highly parallel genomic selection response in replicated Drosophila melanogaster populations with reduced genetic variation
<p>Many adaptive traits are polygenic and frequently more loci contributing to the phenotype are segregating than needed to express the phenotypic optimum. Experimental evolution with replicated populations adapting to a new controlled environment provides a powerful approach to study polygenic adaptation. Since genetic redundancy often results in non-parallel selection responses among replicates, we propose a modified Evolve and Resequence (E&R) design that maximizes the similarity among replicates. Rather than starting from many founders, we only use two inbred <em>Drosophila melanogaster</em>strains and expose them to a very extreme, hot temperature environment (29°C). After 20 generations, we detect many genomic regions with a strong, highly parallel selection response in 10 evolved replicates. The X chromosome has a more pronounced selection response than the autosomes, which may be attributed to dominance effects. Furthermore, we find that the median selection coefficient for all chromosomes is higher in our two-genotype experiment than in classic E&R studies. Since two random genomes harbor sufficient variation for adaptive responses, we propose that this approach is particularly well-suited for the analysis of polygenic adaptation.</p> <p>See the README.txt file to get a description of the uploaded files. Scripts.zip contains annotated command lines and scripts for the project (see internal README.txt file).</p>
Corpus and list of keywords from Improving sustainable crop protection using population genetics concepts
<p>Corpus extracted in April 2021 from the ISI Web of Science portal (https://www.webofscience.com) with the following request: ‘Plant AND Resistan* AND Durab*’. A first corpus of 2522 articles was built considering all publication years for this extraction. This collection was then refined by categories to remove articles outwith the scope of our search (e.g. related to durable resistant materials for constructions). We also kept only articles cited at least once. The final corpus was composed of 1783 articles from 1979 to 2021:</p> <ul> <li>CORPUS_plant_resistance_durability.zip</li> </ul> <p>List of keywords used for the network presented in the article:</p> <ul> <li>keywords_list.csv</li> </ul>
Retrotransposon-based genetic variation of Poa annua populations from contrasting climate conditions
<p>Raw photographs of agarose electrophoresis. Material: six Poa annua populations. Method: inter-Primer Binding Site (iPBS) markers This is the documentation of studies described in the manuscript entitled "Retrotransposon-based genetic variation of Poa annua populations from contrasting climate conditions" accepted for publication in PeerJ journal (decision received on 02.04.2019)</p>
Example Dataset for npstat: Population genetics from Pooled NGS data NPStat v1: User guide
<p>Example Dataset for npstat to test the program and the different options.</p> <p>The example dataset contains a pileup file with sequences of of the 2L chromosome from fifteen pooled inbreed individuals of <em>Drosophila melanogaster </em>(<span>doi: 10.1038/nature10811</span>). The dataset also contains the sequence reference of the 2L chromosome in fasta format, an outgroup sequence in fasta format of <em>D. yakuba</em> (SRR26246471), a GFF3 annotation file and a file with a brief list of selected SNPs to be analyzed.</p>
Attack of the clones: population genetics reveals clonality of Colletotrichum lupini, the causal agent of lupin anthracnose
<p><em>Colletotrichum lupini</em>, causing lupin anthracnose, is one of the worst pathogens to lupin cultivation worldwide. Understanding its population structure and evolutionary potential is crucial to design successful disease management strategies. The objective of this study was to employ population genetics to investigate the diversity, evolutionary dynamics and molecular basis of host interaction of this notorious lupin pathogen. A collection of globally representative <em>C. lupini </em>isolates was genotyped through triple digest restriction-site associated DNA sequencing (3D-RADseq), resulting in a dataset of unparalleled resolution. Phylogenetic and structural analysis could distinguish four (I – IV) independent lineages. The strong population structure, low recombination and slow linkage decay strongly suggests that <em>C. lupini</em> reproduces clonally. Different morphologies and virulence patterns on white (<em>Lupinus albus</em>) and Andean lupin (<em>L. mutabilis</em>) were observed between and within clonal lineages. Isolates belonging to lineage II were shown to have a mini-chromosome which was also partly present in lineage III and IV, but not in lineage I isolates. Variation in the presence of this mini-chromosome could indicate a role in host interaction. All four lineages were present in the South American Andes region, which is concluded to be the center of origin of this species. Only members of lineage II have been found outside South America since the 1990s, indicating it as the current pandemic population. As a seed-borne pathogen, <em>C. lupini</em> has mainly spread through infected but symptomless seeds, stressing the importance of phytosanitary measures to prevent future outbreaks of strains that are yet confined to South America.</p>
Lepidoptera genomics based on 88 chromosomal reference sequences informs population genetic parameters for conservation
<p>This repository contains (1) germline mutations called by the DeepVariant (v1.1.0) pipeline in VCF format; (2) rejected substitution scores calculated by the Genomic Evolutionary Rate Profiling (GERP++) software on each species and chromosome; and (3) the phylogenetic tree used as guide tree in the Cactus alignment.</p>
A scalable, accurate, and universal analysis framework using individual-level allele frequency for large-scale genetic association studies in an admixed population
<p>Inclusion of individuals with diverse or admixed genetic ancestries is crucial to discover novel findings that may be missed by genomics analyses rooted solely in Caucasian population. Here, we present an analysis framework, SPAmix, which is scalable to a large-scale biobank data analysis including hundreds of thousands of admixed individuals and is universally applicable to various types of complex traits including binary trait, quantitative trait, time-to-event trait, longitudinal traits, etc. For each genetic variant, SPAmix uses genotype data and genetic principal components (PCs) to estimate individual-level allele frequency, which is subsequently used to calibrate p values via a retrospective analysis. A hybrid strategy including saddlepoint approximation (SPA) can greatly increase the accuracy to analyze rare genetic variants, especially if the phenotypic distribution is unbalanced or extremely unbalanced. Compared to Tractor, SPAmix does not require local ancestry information and can be straightforwardly applicable to a multi-way admixed population. Meanwhile, SPAmix can also be extended to SPAmix<sub>local</sub> in which the local ancestry can be incorporated if available. In addition, we propose SPAmix<sub>CCT</sub> to combine the p values of SPAmix and SPAmix<sub>local</sub> via Cauchy combination (CCT). SPAmix<sub>local</sub> performs close to Tractor when analyzing quantitative traits and is more accurate when analyzing binary traits with an unbalanced case-control ratio. And SPAmix<sub>CCT </sub>is an optimal unified approach for various cross-ancestry genetic architectures. Extensive simulation studies and real data analyses of 369,314 UK Biobank individuals from multiple ancestries demonstrated that SPAmix is scalable and can discover novel hits while controlling type I error rates well.</p>
Genetic diversity, population structure, and linkage disequilibrium among tropical quality protein maize (QPM) lines assessed with high-density SNP markers
<p>The study of genetic diversity (GD), population structure, and linkage disequilibrium (LD) provides a better understanding of the genetic relationships between individuals in a population which can be utilized in crop research and improvement. Genotyping-by-sequencing (GBS) was used to detect and genotype single nucleotide polymorphisms (SNPs) in a collection of 74 quality protein maize (QPM) lines and further to characterize their genetic diversity, population structure, and linkage disequilibrium. A total of 235,214 high-quality SNPs were used for different genetic analyses except for structure analysis where 11,950 SNPs were used. Analysis of molecular variance (AMOVA) based on these SNPs revealed high genetic heterozygosity among the five populations with 1% of the total genetic variation present among the subpopulations and 99% of the variation among individuals within the populations. Population structure analysis using Bayesian-based clustering revealed that the 74 lines could be clustered into four groups. However, neighbor-joining trees indicate the lines are grouped into three major clusters. Further analysis using principal component analyses (PCA) clustered the genotypes into five groups which are concordant with the groups based on pedigree information. Higher genetic diversity was detected in population 1 with a GD value of 0.484 and the lowest in population 5 (0.396) and overall, with a mean of 0.434. The LD pattern in the quality protein maize was investigated and we observed a relatively rapid LD decay of 3.53kb and 10.66kb at r<sup>2</sup> =0.2 and r<sup>2</sup>= 0.1, respectively. Our findings provide important information for future Linkage mapping studies, genome-wide association analyses, and marker-assisted selective breeding of maize as well as genomic prediction-based selection in tropical germplasm.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.