Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
289
datasets available to search
ShareScore release 0.9.0
Dataset results
289 results for “Genomic prediction”
Predicted genes from the Amblyomma americanum draft genome assembly
<p>Data for pub "Predicted genes from the <em>Amblyomma americanum </em>draft genome assembly."</p> <ul> <li>Amblyomma_americanum_filtered_assembly.fasta: Decontaminated A. americanum genome with bacterial contigs removed</li> <li>Amblyomma_americanum_bacterial_contigs_info.tsv: Information about contigs classified as bacteria that were removed</li> <li>Amblyomma_americanum_annotation_data.tar.gz: Directory of annotation data produced by EvidenceModeler as part of the nf-core/genomeannotator workflow. Includes files in FASTA format (predicted genes and proteins), set of proteins clustered at 99% identity in FASTA format, and annotations in both GFF3 and GTF formats. GTF file produced from the GFF3 file with AGAT.</li> <li>Amblyomma_americanum_transcriptome_assembly_data.tar.gz: Directory of data generated for the transcriptome assembly that was used for gene prediction</li> </ul>
Prediction of nucleosome dyads for the K562 cell line in the hg19 genome assembly
<p><strong><a href="https://andre-rendeiro.com/2015/05/12/predicting_dyads_from_mnase">Predicting dyads from MNase-seq data</a></strong></p> <p>I needed the location of nucleosomal dyads in the K562 cell line (ENCODE tier 1 line). Surprisingly, although plenty of MNase-seq data for that cell line is available, no nucleosome and dyad prediction exists.</p> <p><strong>Running NuMap</strong></p> <p>I found the <a href="http://www-hsc.usc.edu/~valouev/NuMap/NuMap.html">NuMap</a> software by Anton Valouev to do exactly what I intended.</p> <p>Since it is in a somewhat obscure page and this seemed to be the only place where this software was, I have <a href="https://github.com/afrendeiro/NuMap">uploaded it into a Github repository</a> for the sake of preservation (<a href="https://github.com/orphancode/NuMap">https://github.com/orphancode/NuMap</a>).</p> <p>Predicting dyads from MNase-seq data with NuMap seemed trivial: I downloaded the <a href="http://hgdownload.cse.ucsc.edu/goldenPath/hg19/encodeDCC/wgEncodeSydhNsome/">K562 MNase-seq data set</a> (11 replicates ~85Gb!!), combined all replicates and ran NuMap on the data(instructions on the Github README).</p> <p>From NuMap output there are <a href="https://www.dropbox.com/s/asmp7bi40lrvtjb/K562_dyads.bed?dl=0">dyad positions in bed format</a> and you can also produce several metrics to evaluate how good the prediction was.</p> <p><strong>Distograms & phasograms</strong></p> <p>Valouev describes two measurements of the frequencies of distances between MNase-seq reads. The frequency of distances between reads mapping to opposite strands can be used to build a “distogram”, which ilustrates the expected nucleosome length (147 bp) - this is consistent across most eukaryotic cells. The frequency of distances between reads mapping to the same strand gives a measurement of the distance between nucleosomes, as they’re separated by some linker DNA - (Valouev calls this plot a “phasogram”). This measurement, on the other hand tends to be species and cell-type specific.</p> <p><strong>K562 predictions:</strong></p> <p>The expected 147 bp nucleosome length in K562 cells.</p> <p>The average distance between dyads in K562 cells seems to be 185 bp.</p> <p><strong>References:</strong></p> <p>Valouev, A., Johnson, S. M., Boyd, S. D., Smith, C. L., Fire, A. Z., Sidow, A. (2011). Determinants of nucleosome organization in primary human cells. Nature, 474(7352), 516–520. <a href="http://doi.org/10.1038/nature10002">http://doi.org/10.1038/nature10002</a></p>
Data from: Can the genomics of ecological speciation be predicted across the divergence continuum from host races to species? A case study in Rhagoletis
<p>Studies assessing the predictability of evolution typically focus on short-term adaptation within populations or the repeatability of change among lineages. A missing consideration in speciation research is to determine whether natural selection predictably transforms standing genetic variation within populations into differences between species. Here, we test whether host-related selection on diapause timing anticipates genome-wide differentiation during ecological speciation by comparing ancestral hawthorn and newly formed apple-infesting host races of <i>Rhagoletis pomonella </i>to their sibling species <i>R. mendax</i> that attacks blueberries. The responses of 57,857 single nucleotide polymorphisms in a diapause study on the hawthorn race strongly predicted the direction and magnitude of genomic divergence among the three flies at a field site in Fennville, Michigan, USA. As anticipated, the apple race and <i>R. mendax</i> show parallel changes in the frequencies of putative inversions on three chromosomes associated with the earlier fruiting times of apples and blueberries compared to hawthorns. A diapause experiment on <i>R. mendax</i> revealed compensatory mutations throughout the genome accounting for the earlier eclosion of blueberry, but not apple flies. Thus, a degree of predictability, although not complete, exists in the genomics of diapause across the ecological speciation continuum in <i>Rhagoletis</i>. The generality of this result is placed in the context of other similar systems.</p>
Xpresso: Predicting gene expression levels from genomic sequences
<p>Xpresso: Predicting gene expression levels from genomic sequences<br> <br> More info at:<br> Publication: https://doi.org/10.1016/j.celrep.2020.107663<br> Website: https://xpresso.gs.washington.edu/<br> Github: https://github.com/vagarwal87/Xpresso</p>
Growth traits of a tropical timber species at Southeast Asia, Shorea macrophylla, and scripts for genome wide association study and genomic prediction
<p><em><span>Shorea macrophylla</span></em><span> is a commercially important tropical tree species grown for timber and oil. It is amenable to plantation forestry due to its fast initial growth. Genomic selection (GS) has been used in tree breeding studies to shorten long breeding cycles but has not previously been applied to <em>S. macrophylla</em>. To build genomic prediction models for GS, leaves and growth trait data were collected from a half-sib progeny population of <em>S. macrophylla</em> in Sari Bumi Kusuma forest concession, central Kalimantan, Indonesia. 18037 SNP markers were identified in two ddRAD-seq libraries. Genomic prediction models based on these SNPs were then generated for breast height and total height in the 7th year from planting (D7 and H7). These traits were chosen because of their relatively high narrow-sense genomic heritability and because seven years was considered long enough to assess initial growth. Genomic prediction models were built using 12 methods with the full set of identified SNPs and subsets of 48, 96, and 192 SNPs selected based on the results of a genome-wide association study (GWAS). The GBLUP and RKHS methods gave the highest predictive ability (PA) for D7 and H7 and showed that D7 has an additive genetic architecture while H7 has an epistatic genetic architecture. LightGBM and CNN1D also achieved high PA for D7 with 48 and 96 selected SNPs, and for H7 with 96 and 192 selected SNPs, showing that gradient boosting decision trees and deep learning can be useful in genomic prediction. For almost all methods and both traits, PA was higher when SNPs were selected based on their GWAS P-values than when using the full set of SNPs. These results suggest that GS with GWAS-based SNP selection could be used in <em>S. macrophylla </em>breeding to improve initial growth and reduce genotyping costs for next generation seedlings.</span></p>
Plaque assay images from the article "Predicting phage-bacteria interactions at the strain level from genomes"
<p>This dataset contains the plaque assay images for (i) the construction of the 403 bacteria * 96 phages interaction matrix and (ii) the "cocktails" experiment (100 E. coli strains challenged with recommended cocktails vs. a baseline). </p>
Datasets for chromatin hub prediction in six cell lines based on multiple genomic features
<p>Tables with features and classes for machine learning prediction of chromatin hubs. Genomic features include CTCF, EP300, H3K27me3, H3K36me3, H3K4me1, H3K4me2, H3K4me3, H3K9ac, H3K9me3, RAD21, RNAPol2, and RNA.Seq, while the classes are Hubs and Non-Hubs.</p> <p>The cell lines featured here are A549, H1ESC, HeLa, IMR90, K562, and MCF7. They happen to be the 6 cell lines out of 8 existing in our integrative database, GREG (https://doi.org/10.1093/database/baz162). The normalized read-coverages from features (variables) are mapped through genomic intervals of 2 Kbs, genome-wide. Such genomic intervals (bins), are classified as Hubs or Non-Hubs. Hubs are those bins with multiple chromatin interactions, including at least one long-range interaction (larger than 1Mb) or an inter-chromosomal interaction (tagged as Inf).</p> <p>Columns per table:<br> chr start end CTCF EP300 H3K27me3 H3K36me3 H3K4me1 H3K4me2 H3K4me3 H3K9ac H3K9me3 RAD21 RNA.Seq RNAPol2 Class</p> <p>Note that features may be inconsistent across different cell types, due to the availability of data. The BAM files have been sourced from ENCODE and NCBI repositories.</p> <p>The analysis following this data can be found at https://github.com/mora-lab/GREG-Hubs.</p>
Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure
<p>Predicting how tree populations will respond to climate change is an urgent societal concern. An increasingly popular way to make such predictions is the genomic offset (GO) approach, which aims to use genomic and climate data to identify populations that may experience climate maladaptation in the near future. More precisely, GO tries to represent the change in allele frequencies required to maintain the current gene-climate relationships under climate change. However, the GO approach has major limitations and, despite promising validation of its predictions using height data from common gardens, it still lacks broad empirical testing. In the present study, we evaluated the consistency and empirical validity of GO predictions in maritime pine (<em>Pinus pinaster</em> Ait.), a tree species from southwestern Europe and North Africa with a marked population genetic structure. First, gene-climate relationships were estimated using 9,817 SNPs genotyped in 454 trees from 34 populations; and candidate SNPs potentially involved in climate adaptation were identified. Second, GO was predicted using four methods, namely Gradient Forest (GF), Redundancy Analysis (RDA), latent factor mixed model (LFMM) and Generalised Dissimilarity Modeling (GDM), two sets of SNPs (candidate and control SNPs) and five climate general circulation models (GCMs) to account for uncertainty in future climate predictions. Last, the empirical validity of GO predictions was evaluated within a Bayesian framework by estimating the associations between GO predictions and two independent data sources: mortality data from National Forest Inventories (NFI), and mortality and height data from five common gardens in contrasting environments. We found high variability in GO predictions across methods, SNP sets and GCMs. Regarding validation, GO predictions with GDM and GF (and to a lesser extent RDA) based on the candidate SNPs showed the strongest and most consistent associations with mortality rates in common gardens and NFI plots. We found almost no association between GO predictions and tree height in common gardens, most likely due to the overwhelming effect of population genetic structure on tree height in this species. Our study demonstrates the imperative to validate GO predictions with a range of independent data sources before they can be used as informative and reliable metrics in conservation or management strategies.</p>
Swordtail fish hybrids reveal that genome evolution is surprisingly predictable after initial hybridization
<p>Over the past two decades, biologists have come to appreciate that hybridization, or genetic exchange between distinct lineages, is remarkably common – not just in particular lineages but in taxonomic groups across the tree of life. As a result, the genomes of many modern species harbor regions inherited from related species. This observation has raised fundamental questions about the degree to which the genomic outcomes of hybridization are repeatable and the degree to which natural selection drives such repeatability. However, a lack of appropriate systems to answer these questions has limited empirical progress in this area. Here, we leverage independently formed hybrid populations between the swordtail fish <em>Xiphophorus birchmanni </em>and <em>X. cortezi </em>to address this fundamental question. We find that local ancestry in one hybrid population is remarkably predictive of local ancestry in another, demographically independent hybrid population. Applying newly developed methods, we can attribute much of this repeatability to strong selection in the earliest generations after initial hybridization. We complement these analyses with time-series data that demonstrates that ancestry at regions under selection has remained stable over the past ~40 generations of evolution. Finally, we compare our results to the well-studied <em>X. birchmanni×X. malinche </em>hybrid populations and conclude that deeper evolutionary divergence has resulted in stronger selection and higher repeatability in patterns of local ancestry in hybrids between <em>X. birchmanni </em>and <em>X. cortezi</em>.</p>
Data from: Integrating genomic data and simulations to evaluate alternative species distribution models and improve predictions of glacial refugia and future responses to climate change
<p>Climate change poses a threat to biodiversity, and it is unclear whether species can adapt to or tolerate new conditions, or migrate to areas with suitable habitats. Reconstructions of range shifts that occurred in response to environmental changes since the last glacial maximum from species distribution models (SDMs) can provide useful data to inform conservation efforts. However, different SDM algorithms and climate reconstructions often produce contrasting patterns, and validation methods typically focus on accuracy in recreating current distributions, limiting their relevance for assessing predictions to the past or future. We modeled historically suitable habitat for the threatened North American tree green ash (<em>Fraxinus pennsylvanica</em>) using 24 SDMs built using two climate models, three calibration regions, and four modeling algorithms. We evaluated the SDMs using contemporary data with spatial block cross-validation and compared the relative support for alternative models using a novel integrative method based on coupled demographic-genetic simulations. We simulated genomic datasets using habitat suitability of each of the 24 SDMs in a spatially-explicit model. Approximate Bayesian Computation (ABC) was then used to evaluate the support for alternative SDMs through comparisons to an empirical population genomic dataset. Models had very similar performance when assessed with contemporary occurrences using spatial cross-validation, but ABC model selection analyses consistently supported SDMs based on the CCSM climate model, an intermediate calibration extent, and the generalized linear modeling algorithm. Finally, we projected the future range of green ash under four climate change scenarios. Future projections using the SDMs selected via ABC suggest only minor shifts in suitable habitat for this species, while some of those that were rejected predicted dramatic changes. Our results highlight the different inferences that may result from the application of alternative distribution modeling algorithms and provide a novel approach for selecting among a set of competing SDMs with independent data.</p>
Supporting data for RCANE: A Deep Learning Algorithm for Whole-genome Pan-Cancer Somatic Copy Number Aberration Prediction using RNA-seq Data.
<p>This is the data repository for <em>RCANE: A Deep Learning Algorithm for Whole-genome Pan-Cancer Somatic Copy Number Aberration Prediction using RNA-seq Data</em>. To use this dataset, please refer to <a href="https://github.com/HowardGech/RCANE" target="_blank" rel="noopener">https://github.com/HowardGech/RCANE</a>.</p>
Pseudomonas aeruginosa predicted prophages from publicly available genomes
<p>Through the Genome Information by Organism section of the NCBI Genome database, <em>P. aeruginosa</em> bacterial genomic assemblies were downloaded (September 2020). Genome quality was assessed totaling 5,383 genomes total. All 5,383 genomes were then entered into VirSorter v.1 (https://github.com/simroux/VirSorter). The data set provided here includes all category 1 and category 4 predicted prophage sequences.</p>
50000 SNPs for genomic prediction of ash dieback susceptibility in European Ash
<p>This file is <a href="https://www.nature.com/articles/s41559-019-1036-6">Stocks et al (2019)</a> Supplementary Table 7j with major allele (MAA) and minor allele (MIA) identities added.<br> Estimated effect sizes (EES) from genomic prediction model trained on the pool-seq data using the top 50000 SNPs from the pool-seq GWAS<br> Contig = Contig in BATG0.5 assembly <br> Pos = SNP location in contig <br> EES.MIA = Estimated effect size of minor allele <br> EES.MIA.SE = Standard Error of Estimated effect size of minor allele<br> EES.MAA = Estimated effect size of major allele<br> EES.MAA.SE = Standard Error of Estimated effect size of major allele <br> MIA = identity of minor allele <br> MAA = identity of major allele</p>
Supplementary Datasets for "Genome-wide CRISPR off-target prediction and optimization using RNA-DNA interaction fingerprints"
<p>Supplementary Datasets for "Genome-wide CRISPR off-target prediction and optimization using RNA-DNA interaction fingerprints". The deposition contains training/testing datasets used in the article.</p>
Development of whole-genome prediction models to increase the rate of genetic gain in intermediate wheatgrass (Thinopyrum intermedium) breeding
Open the record for dataset details and reuse information.
Habitat association predicts population connectivity and persistence in flightless beetles: a population genomics approach within a dynamic archipelago
Open the record for dataset details and reuse information.
Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure
Open the record for dataset details and reuse information.
Temperature-specific repeatability of evolution and its implications for genomic predictions of adaptation to warming
Open the record for dataset details and reuse information.
Growth traits of a tropical timber species at Southeast Asia, Shorea macrophylla, and scripts for genome wide association study and genomic prediction
Open the record for dataset details and reuse information.
Data from: Can the genomics of ecological speciation be predicted across the divergence continuum from host races to species? A case study in Rhagoletis
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.