Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

289

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

289 results for “Genomic prediction”

Learn how ShareScore rates datasets ↗
dryad32/100

Benchmarking parametric and machine learning models for genomic prediction of complex traits

<p>The usefulness of genomic prediction in crop and livestock breeding programs has prompted efforts to develop new and improved genomic prediction algorithms, such as artificial neural networks and gradient tree boosting. However, the performance of these algorithms has not been compared in a systematic manner using a wide range of datasets and models. Using data of 18 traits across six plant species with different marker densities and training population sizes, we compared the performance of six linear and six non-linear algorithms. First, we found that hyperparameter selection was necessary for all non-linear algorithms and that feature selection prior to model training was critical for artificial neural networks when the markers greatly outnumbered the number of training lines. Across all species and trait combinations, no one algorithm performed best, however predictions based on a combination of results from multiple algorithms (i.e. ensemble predictions) performed consistently well. While linear and non-linear algorithms performed best for a similar number of traits, the performance of non-linear algorithms vary more between traits. Although artificial neural networks did not perform best for any trait, we identified strategies (i.e. feature selection, seeded starting weights) that boosted their performance to near the level of other algorithms. Our results highlight the importance of algorithm selection for the prediction of trait values.</p>

opencc-zeroOct 2019View details →
dryad32/100

Genomic prediction applied to multiple traits and environments in second season maize hybrids

<p>Genomic selection has become a reality in plant breeding programs with the reduction in genotyping costs. Especially in maize breeding programs, it emerges as a promising tool for predicting hybrid performance. The dynamics of a commercial breeding program involve the evaluation of several traits simultaneously in a large set of target environments. Therefore, multi-trait multi-environment (MTME) genomic prediction models can leverage these data sets by exploring the correlation between traits and Genotype-by-Environment (G×E) interaction. Herein, we assess predictive abilities of univariate and multivariate genomic prediction models in a maize breeding program. To this end, we used data from 415 maize hybrids evaluated in four years of second season field trials for the traits grain yield, number of ears and grain moisture. Genotypes of these hybrids were inferred <i>in silico</i> based on their parental inbred lines using Single Nucleotide Polymorphisms (SNPs) markers obtained via genotyping-by-sequencing (GBS). Because genotypic information was available for only 257 hybrids, we used the genomic and pedigree relationship matrices to obtain the <b>H</b> matrix for all 415 hybrids. Our results demonstrated that in the single-environment context the use of multi-trait models was always superior in comparison to their univariate counterparts. Besides that, although MTME models were not particularly successful in predicting hybrid performance in untested years, they improved the ability to predict the performance of hybrids that had not been evaluated in any environment. However, the computational requirements of this kind of model could represent a limitation to its practical implementation and further investigation is necessary.</p>

opencc-zeroMay 2020View details →
dryad32/100

Data from: Genomic prediction accuracies in space and time for height and wood density of Douglas-fir using exome capture as the genotyping platform

Background Genomic selection (GS) can offer unprecedented gains, in terms of cost efficiency and generation turnover, to forest tree selective breeding; especially for late expressing and low heritability traits. Here, we used: 1) exome capture as a genotyping platform for 1372 Douglas-fir trees representing 37 full-sib families growing on three sites in British Columbia, Canada and 2) height growth and wood density (EBVs), and deregressed estimated breeding values (DEBVs) as phenotypes. Representing models with (EBVs) and without (DEBVs) pedigree structure. Ridge regression best linear unbiased predictor (RR-BLUP) and generalized ridge regression (GRR) were used to assess their predictive accuracies over space (within site, cross-sites, multi-site, and multi-site to single site) and time (age-age/ trait-trait). Results The RR-BLUP and GRR models produced similar predictive accuracies across the studied traits. Within-site GS prediction accuracies with models trained on EBVs were high (RR-BLUP: 0.79–0.91 and GRR: 0.80–0.91), and were generally similar to the multi-site (RR-BLUP: 0.83–0.91, GRR: 0.83–0.91) and multi-site to single-site predictive accuracies (RR-BLUP: 0.79–0.92, GRR: 0.79–0.92). Cross-site predictions were surprisingly high, with predictive accuracies within a similar range (RR-BLUP: 0.79–0.92, GRR: 0.78–0.91). Height at 12 years was deemed the earliest acceptable age at which accurate predictions can be made concerning future height (age-age) and wood density (trait-trait). Using DEBVs reduced the accuracies of all cross-validation procedures dramatically, indicating that the models were tracking pedigree (family means), rather than marker-QTL LD. Conclusions While GS models' prediction accuracies were high, the main driving force was the pedigree tracking rather than LD. It is likely that many more markers are needed to increase the chance of capturing the LD between causal genes and markers.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Accurate genomic predictions for chronic wasting disease in U.S. white-tailed deer

<p>The geographic expansion of chronic wasting disease (CWD) in U.S. white-tailed deer (<em>Odocoileus virginianus</em>) has been largely unabated by best management practices, diagnostic surveillance, and depopulation of positive herds. Using a custom Affymetrix Axiom® single nucleotide polymorphism (SNP) array, we demonstrate that both differential susceptibility to CWD, and natural variation in disease progression, are moderately to highly heritable ( among farmed U.S. white-tailed deer, and that loci other than <em>PRNP</em> are involved. Genome-wide association analyses using 123,987 quality filtered SNPs for a geographically diverse cohort of 807 farmed U.S. white-tailed deer (n = 284 CWD positive; n = 523 CWD non-detect) confirmed the prion gene (<em>PRNP</em>; G96S) as a large-effect risk locus (<em>P</em>-value &lt; 6.3E-11), as evidenced by the estimated proportion of phenotypic variance explained (PVE ≥ 0.05), but also demonstrated that more phenotypic variance was collectively explained by loci other than <em>PRNP</em>.<em> </em>Genomic best linear unbiased prediction (GBLUP; n = 123,987 SNPs) with <em>k</em>-fold cross validation (<em>k</em> = 3; <em>k</em> = 5) and random sampling (n = 50 iterations) for the same cohort of 807 farmed U.S. white-tailed deer produced mean genomic prediction accuracies ≥ 0.81; thereby providing the necessary foundation for exploring a genomically-estimated CWD eradication program.</p>

opencc-zeroMar 2020View details →
dryad32/100

Data from: The predictability of genomic changes underlying a recent host shift in Melissa blue butterflies

Despite accumulating evidence that evolution can be predictable, studies quantifying the predictability of evolution remain rare. Here, we measured the predictability of genome-wide evolutionary changes associated with a recent host shift in the Melissa blue butterfly (Lycaeides melissa). We asked whether and to what extent genome-wide patterns of evolutionary change in nature could be predicted (1) by comparisons among instances of repeated evolution, and (2) from SNP $\times$ performance associations in a lab experiment. We delineated the genetic loci (SNPs) most strongly associated with host use in two L. melissa lineages that colonized alfalfa. Whereas most SNPs were strongly associated with host use in none or one of these lineages, we detected a ~two-fold excess of SNPs associated with host use in both lineages. Similarly, we found that host-associated SNPs in nature could also be partially predicted from SNP $\times$ performance (survival and weight) associations in a lab rearing experiment. But the extent of overlap, and thus degree of predictability, was somewhat reduced. Although we were able to predict (to a modest extent) the SNPs most strongly associated with host use in nature (in terms of parallelism and from the experiment), we had little to no ability to predict the direction of evolutionary change during the colonization of alfalfa. Our results show that different aspects of evolution associated with recent adaptation can be more or less predictable, and highlight how stochastic and deterministic processes interact to drive patterns of genome-wide evolutionary change

opencc-zeroDec 2017View details →
dryad32/100

Data from: Efficiency of genomic prediction across two Eucalyptus nitens seed orchards with different selection histories

Genomic selection is expected to enhance the genetic improvement of forest tree species by providing more accurate estimates of breeding values through marker-based relationship matrices compared with pedigree-based methodologies. When adequately robust genomic prediction models are available, an additional increase in genetic gains can be made possible with the shortening of the breeding cycle through elimination of the progeny testing phase and early selection of parental candidates. The potential of genomic selection was investigated in an advanced Eucalyptus nitens breeding population focused on improvement for solid wood production. A high-density SNP chip (EUChip60K) was used to genotype 691 individuals in the breeding population, which represented two seed orchards with different selection histories. Phenotypic records for growth and form traits at age six, and for wood quality traits at age seven were available to build genomic prediction models using GBLUP which were compared to the traditional pedigree-based alternative using BLUP. GBLUP demonstrated that breeding value accuracy would be improved and substantial increases in genetic gains towards solid wood production would be achieved. Cross-validation within and across two different seed orchards indicated that genomic predictions would likely benefit in terms of higher predictive accuracy from increasing the size of the training data sets through higher relatedness and better utilization of LD

opencc-zeroDec 2017View details →
dryad32/100

Data from: Genomic predictions and genome-wide association study of resistance against Piscirickettsia salmonis in coho salmon (Oncorhynchus kisutch) using ddRAD sequencing

Piscirickettsia salmonis is one of the main infectious diseases affecting coho salmon (Oncorhynchus kisutch) farming, and current treatments have been ineffective for the control of this disease. Genetic improvement for P. salmonis resistance has been proposed as a feasible alternative for the control of this infectious disease in farmed fish. Genotyping by sequencing (GBS) strategies allow genotyping of hundreds of individuals with thousands of single nucleotide polymorphisms (SNPs), which can be used to perform genome wide association studies (GWAS) and predict genetic values using genome-wide information. We used double-digest restriction-site associated DNA (ddRAD) sequencing to dissect the genetic architecture of resistance against P. salmonis in a farmed coho salmon population and to identify molecular markers associated with the trait. We also evaluated genomic selection (GS) models in order to determine the potential to accelerate the genetic improvement of this trait by means of using genome-wide molecular information. A total of 764 individuals from 33 full-sib families (17 highly resistant and 16 highly susceptible) were experimentally challenged against P. salmonis and their genotypes were assayed using ddRAD sequencing. A total of 9,389 SNPs markers were identified in the population. These markers were used to test genomic selection models and compare different GWAS methodologies for resistance measured as day of death (DD) and binary survival (BIN). Genomic selection models showed higher accuracies than the traditional pedigree-based best linear unbiased prediction (PBLUP) method, for both DD and BIN. The models showed an improvement of up to 95% and 155% respectively over PBLUP. One SNP related with B-cell development was identified as a potential functional candidate associated with resistance to P. salmonis defined as DD.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Combining high-throughput phenotyping and genomic information to increase prediction and selection accuracy in wheat breeding

Genomics and phenomics have promised to revolutionize the field of plant breeding. The integration of these two fields has just begun and is being driven through big data by advances in next-generation sequencing and developments of field-based high-throughput phenotyping (HTP) platforms. Each year the International Maize and Wheat Improvement Center (CIMMYT) evaluates tens-of-thousands of advanced lines for grain yield across multiple environments. To evaluate how CIMMYT may utilize dynamic HTP data for genomic selection (GS), we evaluated 1170 of these advanced lines in two environments, drought (2014, 2015) and heat (2015). A portable phenotyping system called 'Phenocart' was used to measure normalized difference vegetation index and canopy temperature simultaneously while tagging each data point with precise GPS coordinates. For genomic profiling, genotyping-by-sequencing (GBS) was used for marker discovery and genotyping. Several GS models were evaluated utilizing the 2254 GBS markers along with over 1.1 million phenotypic observations. The physiological measurements collected by HTP, whether used as a response in multivariate models or as a covariate in univariate models, resulted in a range of 33% below to 7% above the standard univariate model. Continued advances in yield prediction models as well as increasing data generating capabilities for both genomic and phenomic data will make these selection strategies tractable for plant breeders to implement increasing the rate of genetic gain.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Accurate genomic prediction of Coffea canephora in multiple environments using whole-genome statistical models

Genomic selection have been proposed as the standard method to predict breeding values in animal and plant breeding. Although some crops have benefited from this methodology, studies in Coffea are still emerging. To date, there have been no studies of how well genomic prediction models work across populations and environments for different complex traits in coffee. Considering that predictive models are based on biological and statistical assumptions, it is expected that their performance vary depending on how well these assumptions align with the true genetic architecture of the phenotype. To investigate this, we used data from two recurrent selection populations of Coffea canephora, evaluated in two locations, and single nucleotide polymorphisms identified by Genotyping-by-Sequencing. In particular, we evaluated the performance of 13 statistical approaches to predict three important traits in the coffee — production of coffee beans, leaf rust incidence and yield of green beans. Analyses were performed for predictions within-environment, across locations and across populations to assess the reliability of genomic selection. Overall, differences in the prediction accuracy of the competing models were small, although the Bayesian methods showed a modest improvement over other methods, at the cost of more computation time. As expected, predictive accuracy for within-environment analysis, on average, were higher than predictions across locations and across populations. Our results support the potential of genomic selection to reshape traditional plant breeding schemes. In practice, we expect to increase the genetic gain per unit of time by reducing the length cycle of recurrent selection in coffee.

opencc-zeroDec 2017View details →
zenodo32/100

Genome size predicts diatom abundance in the polar ocean

<p>This repository contains the datasets, code, and results for:</p> <p>Roberts et al. 2024. Genome size predicts diatom abundance in the polar ocean.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
dryad32/100

Predictability and parallelism in the contemporary evolution of hybrid genomes

<p>Hybridization between species is widespread across the tree of life. As a result, many species, including our own, harbor regions of their genome derived from hybridization. Despite the recognition that this process is widespread, we understand little about how the genome stabilizes following hybridization, and whether the mechanisms driving this stabilization tend to be shared across species. Here, we dissect the drivers of variation in local ancestry across the genome in replicated hybridization events between two species pairs of swordtail fish: <em>Xiphophorus birchmanni </em>× <em>X. cortezi</em> and <em>X. birchmanni </em>× <em>X. malinche</em> . We find surprisingly high levels of repeatability in local ancestry across the two types of hybrid populations. This repeatability is attributable in part to the fact that the recombination landscape and locations of functionally important elements play a major role in driving variation in local ancestry in both types of hybrid populations. Beyond these broad scale patterns, we identify dozens of regions of the genome where minor parent ancestry is unusually low or high across species pairs. Analysis of these regions points to shared sites under selection across species pairs, and in some cases, shared mechanisms of selection. We show that one such region is a previously unknown hybrid incompatibility that is shared across <em>X. birchmanni</em> × <em>X. cortezi</em> and <em>X. birchmanni</em> × <em>X. malinche</em> hybrid populations. </p>

opencc-zeroOct 2021View details →
dryad32/100

Data from: Improvement of genomic predictions in small breeds by construction of genomic relationship matrix through variable selection

<p>Genomic selection has been increasingly implemented in the animal breeding industry, and it is becoming a routine method in many livestock breeding contexts. However, its use is still limited in several small-population local breeds, which are, nonetheless, an important source of genetic variability of great economic value. A major roadblock for their genomic selection is accuracy when population size is limited: to improve breeding value accuracy, variable selection models that assume heterogenous variance have been proposed over the last few years. However, while these models might outperform traditional and genomic predictions in terms of accuracy, they also carry a proportional increase of breeding value bias and dispersion. These mutual increases are especially striking when genomic selection is performed with a low number of phenotypes and high shrinkage value—which is precisely the situation that happens with small local breeds. In our study, we tested several alternative methods to improve the accuracy of genomic selection in a small population. First, we investigated the impact of using only a subset of informative markers regarding prediction accuracy, bias, and dispersion. We used different algorithms to select them, such as recursive feature eliminations, penalized regression, and XGBoost. We compared our results with the predictions of pedigree-based BLUP, single-step genomic BLUP, and weighted single-step genomic BLUP in different simulated populations obtained by combining various parameters in terms of number of QTLs and effective population size. We also investigated these approaches on a real data set belonging to the small local Rendena breed. Our results show that the accuracy of GBLUP in small-sized populations increased when performed with SNPs selected via variable selection methods both in simulated and real data sets. In addition, the use of variable selection models—especially those using XGBoost—in our real data set did not impact bias and the dispersion of estimated breeding values. We have discussed possible explanations for our results and how our study can help estimate breeding values for future genomic selection in small breeds.</p>

opencc-zeroAug 2022View details →
zenodo32/100

Data from: Genome-wide Polygenic Risk Scores Predict Risk of Glioma and Molecular Subtypes

<div> <div> <div> <p><strong>Background</strong>: Polygenic risk scores (PRS) aggregate the contribution of many risk variants to provide a personalized genetic susceptibility profile. Since sample sizes of glioma genome-wide association studies (GWAS) remain modest, there is a need to efficiently capture genetic risk using available data.</p> <p><strong>Methods</strong>: We applied a method based on continuous shrinkage priors (PRS-CS) to model the joint effects of over 1 million common variants on disease risk and compared this to an approach (PRS-CT) that only selects a limited set of independent variants that reach genome-wide significance (P&lt;5&times;10-8). PRS models were trained using GWAS stratified by histological (10,346 cases, 14,687 controls) and molecular subtype (2,632 cases, 2,445 controls), and validated in two independent cohorts.</p> <p><strong>Results</strong>: PRS-CS was generally more predictive than PRS-CT with a median increase in explained variance (R2) of 24% (interquartile range=11-30%) across glioma subtypes. Improvements were pronounced for glioblastoma (GBM), with PRS-CS yielding larger odds ratios (OR) per standard deviation (OR=1.93, P=2.0&times;10-54 vs. OR=1.83, P=9.4&times;10-50) and higher explained variance (R2=2.82% vs. R2=2.56%). Individuals in the 80th percentile of the PRS- CS distribution had significantly higher risk of GBM (0.107%) at age 60 compared to those with average PRS (0.046%, P=2.4&times;10-12). Lifetime absolute risk reached 1.18% for glioma and 0.76% for IDH wildtype tumors for individuals in the 95th PRS percentile. PRS-CS augmented the classification of IDH mutation status in cases when added to demographic factors (AUC=0.839 vs. AUC=0.895, P=6.8&times;10-9).</p> <p><strong>Conclusions</strong>: Genome-wide PRS has potential to enhance the detection of high-risk individuals and help distinguish between prognostic glioma subtypes.</p> <p><strong>Citation</strong>: Nakase T, Guerra GA, Ostrom QT, et al. Genome-wide Polygenic Risk Scores Predict Risk of Glioma and Molecular Subtypes. <em>Neuro-Oncology</em>. Published online June 25, 2024:noae112. doi:10.1093/neuonc/noae112</p> </div> </div> </div>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Gene predictions in GFF format for tuatara genome assembly QEPC00000000.1

<p><strong>Maker gene predictions for Sphenodon punctatus (tuatara) isolate: mauimua-1&nbsp;</strong>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;<br> <br> These GFF files correspond to the genome assembly in GenBank ID QEPC00000000.1<br> https://www.ncbi.nlm.nih.gov/nuccore/QEPC00000000.1</p> <p>Files:</p> <p>&nbsp;<strong>MASKED_Tuatara.DEVO.annot15102.gff.gz</strong><br> &nbsp;&nbsp; - gene and transcript predictions only<br> <br> &nbsp;<strong>20160427.tuatara.maker.raw.gff3.gz</strong><br> &nbsp;&nbsp; - raw Maker output including gene and transcript predictions, alignment results and output of individual gene prediction algorithms (SNAP and Augustus).</p> <p><br> <br> &nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo32/100

FIGURE 2. Predicted secondary structures for 22 in The complete mitochondrial genome of the jumping grasshopper Sinopodisma pieli (Orthoptera: Acrididae) and the phylogenetic analysis of Melanoplinae

FIGURE 2. Predicted secondary structures for 22 tRNA genes of the S. pieli mitogenome. The tRNAs are labeled with the abbreviations of their corresponding amino acids. The minus sign (-) indicates Watson-Crick base pairing and plus sign (.) indicates G-U base pairing.

opennotspecifiedDec 2017View details →
zenodo32/100

Accurate genome-wide predictions of spatio-temporal gene expression during embryonic development

<p>This upload contains the expression prediction dataset discussed in the manuscript &quot;Accurate genome-wide predictions of spatio-temporal gene expression during embryonic development&quot; and used by&nbsp;the webserver&nbsp;https://find.princeton.edu.</p> <p>Abstract:</p> <p>Comprehensive information on the timing and location of gene expression is fundamental to our understanding of embryonic development and tissue formation.&nbsp; While high-throughput&nbsp;<em>in situ</em>&nbsp;hybridization projects provide invaluable information about developmental gene expression patterns for model organisms like&nbsp;<em>Drosophila</em>, the output of these experiments is primarily qualitative, and a high proportion of protein coding genes and most non-coding genes lack any annotation.&nbsp; Accurate data-centric predictions of spatio-temporal gene expression will therefore complement current <em>in situ</em>&nbsp;hybridization efforts.&nbsp; Here, we applied a machine learning approach by training models on all public gene expression and chromatin data, even from whole-organism experiments, to provide genome-wide, quantitative spatio-temporal predictions for all genes.&nbsp; We developed structured&nbsp;in silico&nbsp;nano-dissection, a computational approach that predicts gene expression in &gt;200 tissue-developmental stages. The algorithm integrates expression signals from a compendium of 6,378 genome-wide expression and chromatin profiling experiments in a cell lineage-aware fashion.&nbsp; We systematically evaluated our performance via cross-validation and experimentally confirmed 22 new predictions for four different embryonic tissues.&nbsp; The model also predicts complex, multi-tissue expression and developmental regulation with high accuracy.&nbsp; We further show the potential of applying these genome-wide predictions to extract tissue specificity signals from non-tissue-dissected experiments, and to prioritize tissues and stages for disease modeling.&nbsp; This resource, together with the exploratory tools are freely available at our webserver&nbsp;<a href="http://find.princeton.edu/">http://find.princeton.edu</a>, &nbsp;which provides a valuable tool for a range of applications, from predicting spatio-temporal expression patterns to recognizing tissue signatures from differential gene expression profiles.</p>

opencc-by-4.0Sep 2019View details →
zenodo32/100

varCADD: large sets of standing genetic variation enable genome-wide pathogenicity prediction

<p>Data and trained models for the manuscript <em>varCADD: large sets of standing genetic variation enable genome-wide pathogenicity prediction.</em></p>

openmit-licenseSep 2024View details →
dryad32/100

Data from: Improving genomic prediction for two Yorkshire populations with a limited size using single-step method

In this study, we conducted genomic prediction for two Yorkshire purebred populations (Yichun and Chifeng) from two different provinces of China that both had a limited population size. Two growth traits (age adjusted to 100 kg weight, AGE; back‐fat thickness adjusted to 100 kg weight, BF) and one reproduction trait (total number of piglets born, TNB) were analyzed with four prediction strategies: one‐population BLUP, joint two‐population BLUP, one‐population single‐step BLUP (SSBLUP) and joint two‐population SSBLUP. Our results illustrate that accuracies of genomic estimated breeding values were improved for BF and TNB for the Yichun population and for BF for the Chifeng population by genomic prediction (one‐population SSBLUP and joint two‐population SSBLUP). The accuracy of TNB for the Yichun population was increased two fold when comparing the one‐population SSBLUP to the one‐population BLUP prediction. Meanwhile, prediction biases were dramatically reduced for AGE for the Yichun population and for TNB for the Chifeng population. The conclusions of this study are as follows: first, genomic prediction is useful for improving prediction accuracy for purebred pig breeding farms with a limited population size; second, joint genomic prediction for different populations of the same breed with certain genetic links has the trend to further improve prediction accuracy.

opencc-zeroDec 2018View details →
dryad32/100

Competitiveness prediction for nodule colonization in Sinorhizobium meliloti through combined in vitro tagged strain characterization and genome-wide association analysis

<p>Associations between leguminous plants and symbiotic nitrogen-fixing rhizobia are a classic example of mutualism between a eukaryotic host and a specific group of prokaryotic microbes. Although this symbiosis is in part species-specific, different rhizobial strains may colonise the same nodule. Some rhizobial strains are commonly known as better competitors than others, but detailed analyses that aim to predict rhizobial competitive abilities based on genomes are still scarce. Here, we performed a bacterial <em>genome-wide association (GWAS) analysis to define the </em>genomic determinants related to the competitive capabilities in the model rhizobial species <em>Sinorhizobium meliloti.</em> For this, 13 tester strains were GFP-tagged and assayed <i>vs.</i> 3 RFP-tagged reference competitor strains (<em>Rm1021, AK83, and BL225C) in a</em> <i>Medicago sativa</i> nodule occupancy test. Competition data and strain genomic sequences were employed to build a model for GWAS based on <i>k</i>-mers. Among the <i>k</i>-mers with the highest scores, 51 <i>k</i>-mers mapped on the genomes of four strains showing the highest competition phenotypes (&gt; 60% single strain nodule occupancy; GR4, KH35c, KH46 and SM11) <i>vs.</i> BL225C. These <i>k</i>-mers were mainly located on the symbiosis-related megaplasmid pSymA, specifically on genes coding for transporters, proteins involved in the biosynthesis of cofactors and proteins related to metabolism (e.g., fatty acids). The same analysis was performed considering the sum of single and mixed nodules obtained in the competition assays <em>vs. </em>BL225C, retrieving <i>k</i>-mers mapped on the genes previously found and on <i>vir</i> genes. Therefore, the competition abilities seem to be linked to multiple genetic determinants and comprise several cellular components.</p>

opencc-zeroJul 2021View details →
dryad32/100

Haplotype-based genome-wide association increases the predictability of leaf rust (Puccinia triticina) resistance in wheat

<p></p><p>Resistance breeding is crucial for a sustainable control of wheat leaf rust and SNP-based genome-wide association studies (GWAS) are widely used to dissect leaf rust resistance. Unfortunately, GWAS based on SNPs explained often only a small proportion of the genetic variation. We compared SNP-based GWAS with a method based on functional haplotypes (FH) considering epistasis in a comprehensive hybrid wheat mapping population composed of 133 parents plus their 1,574 hybrids and characterized with 626,245 high-quality SNPs. In total, 2,408 and 1,139,828 significant associations were detected in the mapping population by using SNP-based and FH-GWAS, respectively. These associations mapped to 25 and 69 candidate regions, correspondingly. SNP-based GWAS highlighted two already-known resistance genes, i.e. Lr22a and Lr34-B, while FH-GWAS not only detected associations on these genes but also on two additional genes, i.e. Lr10 and Lr1. As revealed by a second hybrid wheat population for independent validation, using detected associations from SNP-based and FH-GWAS reached predictabilities of 11.72% and 22.86%, respectively. Therefore, FH-GWAS is not only more powerful to detect associations, but also improves the accuracy of marker-assisted selection as compared to the SNP-based approach.</p><p></p>

opencc-zeroAug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record