Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
289
datasets available to search
ShareScore release 0.9.0
Dataset results
289 results for “Genomic prediction”
Machine learning classification of archaea and bacteria identifies novel predictive genomic features
<p>Dataset used for the classification analysis of archaea and bacteria based on 77 genomic features calculated using GBRAP (GenBank Retrieving, Analyzing and Parsing) tool (Vischioni, C. et al. Gbrap: a tool to retrieve, parse and analyze genbank files of viral and bacterial species. bioRxiv 2021–09 (2021)).</p>
Machine learning reveals the diversity of human 3D chromatin contact patterns (example predictions genome wide)
<p>Example data for the paper: Machine learning reveals the diversity of human 3D chromatin contact patterns</p> <p>GitHub: https://github.com/erin-n-gilbertson/3DGenome-diversity/tree/main</p> <p>biorXiv: https://www.biorxiv.org/content/10.1101/2023.12.22.573104v1.full</p> <p>Manuscript accepted at Molecular Biology and Evolution</p> <p>Of primary interest will be the example predictions genome wide for hg38 reference, human-archaic hominin ancestor and most divergent 1KG individual per genome along with the Jupyter notebook tutorial for making your own Akita predictions given any input 1MB sequence.</p> <div> <ul> <li>bin: contains python script for and qsub array shell script for generating example predictions. These scripts can be modified to take in any fasta files as input.</li> <li>akita_predictions: contains both Akita prediction output arrays and SVG files with predicted contact maps for the hg38 reference, human-archaic hominin ancestor and most divergent 1KG individual in each of 4,873 1MB windows</li> <li>anc_window_spearman.csv: spearman correlation between each 1KG individual and the ancestor for each 1MB window. To calculate 3D divergence subtract these values from 1.</li> <li>basenji: basenji dir from their github, necessary in the directory to run predictions - https://github.com/calico/basenji/tree/master</li> <li>genomes: fasta genomes for hg38 reference and human-archaic hominin ancestor used to make akita predictions</li> <li>divergent_windows: variants and expected divergence distributions for 392 more divergent than expected windows. Defined in the manuscript as windows where 3D divergence between 1KG indiivudals and the ancestor is greater than what would be expected based on sequence divergence. See manuscript Fig. S9 for more details. </li> <li>windows.txt: 4,873 1MB genomic windows with 100% coverage in hg38 used for Akita predictions</li> <li>making_examples.ipynb: jupyter notebook with tutorial instructions for making Akita predictions on any human genome sequence.</li> </ul> <br><br></div>
Mode-of-inheritance predictions for all possible missense variants in the human genome (hg38)
<p><span>Ensemble and consensus approaches to prediction of recessive inheritance for missense variants in human disease.</span></p>
Combined tumor and immune signals from genomes or transcriptomes predict outcomes of checkpoint inhibition in melanoma
<p>Code and data for manuscript "Combined tumor and immune signals from genomes or transcriptomes predict outcomes of checkpoint inhibition in melanoma"</p>
Data from: A mix of old British and modern European breeds: Genomic prediction of breed composition of smallholder pigs in Uganda
<p>Pig herds in Africa comprise genotypes ranging from local ecotypes to commercial breeds. Many animals are composites of these two types and the best levels of crossbreeding for particular production systems are largely unknown. These pigs are managed without structured breeding programs and inbreeding is potentially limiting. The objective of this study was to quantify ancestry contributions and inbreeding levels in a population of smallholder pigs in Uganda. The study was set in the districts of Hoima and Kamuli in Uganda and involved 422 pigs. Pig hair samples were taken from adult and growing pigs in the framework of a longitudinal study investigating productivity and profitability of smallholder pig production. The samples were genotyped using the porcine GeneSeek Genomic Profiler (GGP) 50K SNP Chip. The SNP data was analyzed to infer breed ancestry and autozygosity of the Uganda pigs. The results showed that exotic breeds (modern European and old British) contributed an average of 22.8% with a range of 2–50% while "local" blood contributed 69.2% (36.9–95.2%) to the ancestry of the pigs. Runs of homozygosity (ROH) greater than 2 megabase (Mb) quantified the average genomic inbreeding coefficient of the pigs as 0.043. The scarcity of long ROH indicated low recent inbreeding. We conclude that the genomic background of the pig population in the study is a mix of old British and modern pig ancestries. Best levels of admixture for smallholder pigs are yet to be determined, by linking genotypes and phenotypic records.</p>
Genomically predicted theoretical protein mass database for mass spectrometry (GPMsDB) evaluation datasets
<p>These are datasets obtained for the evaluation of GPMsDB (genomically predicted protein mass database) and its toolkits (GPMsDB-tk/GPMsDB-dbtk). The following datasets are deposited.</p> <ul> <li>The genome sequences of the strains newly sequenced and added using GPMsDB-dbtk (genomes_added.zip)</li> <li>MALDI-TOF-MS peak lists obtained from reference bacterial and archaeal strains (MALDI_peaklists.zip)</li> <li>16S rRNA gene sequences of the faecal isolates (mice_isolates_16S_nanopore.zip)</li> <li>Metagenome-assembled genomes from mouse faeces (mice_MAGs.zip)</li> </ul>
MAC genome assembly and gene prediction of Tetrahymena thermophila SB210
<p>Corrected genome assembly and gene prediction of the MAC genome of T. thermophila SB210. These data were generated and analysed in the manuscript "Single-nucleotide polymorphism landscape of the macronuclear genome of <em>Tetrahymena thermophila".</em></p> <p>Please see the Material & methods and Supplementary data files of this manuscript for more details about these files.</p>
Genome annotation file containing predicted genome features of Phytophthora agathidicida (Strain: 3770, Assembly:ASM2572299v1)
<p>This is the genome annotation file (gff3) containing predicted genome features of the <em>Phytophthora agathidicida </em>(Strain: 3770) genome published in Cox et al (2022). This annotation file is associated with the following entries at Genbank:</p> <p>Assembly: ASM2572299v1<br> Biosample: SAMN19597867<br> BioProject: PRJNA734652</p> <p>Included in the file are predicted functional annotations from Blastp search of all predicted proteins sequences against the Swiss-Prot sequence database (Release 23/02).</p>
Data from: Phenotypic drought stress prediction of European beech (Fagus sylvatica) by genomic prediction and remote sensing
<p><span>Current climate change species response models usually do not include evolution. We integrated remote sensing with population genomics to improve phenotypic response prediction to drought stress in the key forest tree species European beech (<em>Fagus sylvatica</em> L.). We used whole-genome sequencing of pooled DNA from natural stands along an ecological gradient from humid-cold to warm-dry climate. We phenotyped stands for leaf area index (LAI) and moisture stress index (MSI) for the period 2016–2022. We predicted this data with matching meteorological data and a newly developed genomic population prediction score in a Generalised Linear Model. Model selection showed that the addition of genomic prediction decisively increased the explanatory power. We then predicted the response of beech to future climate change under evolutionary adaptation scenarios. A moderate climate change scenario would allow persistence of adapted beech forests, but not worst-case scenarios. Our approach can thus guide mitigation measures, such as allowing natural selection or proactive evolutionary management.</span></p>
Gene prediction for: A reference genome for ecological restoration of the sunflower sea star, Pycnopodia helianthoides
<div> <div> <div> <div>Wildlife diseases, such as the sea star wasting (SSW) epizootic that outbroke in the mid-2010s, appear to be associated with acute and/or chronic abiotic environmental change; dissociating the effects of different drivers can be difficult. The sunflower sea star,<em> Pycnopodia helianthoides</em>, was the species most severely impacted during the SSW outbreak, which overlapped with periods of anomalous atmospheric and oceanographic conditions, and there is not yet a consensus on the cause(s). Genomic data may reveal underlying molecular signatures that implicate a subset of factors and, thus, clarify past events while also setting the scene for effective restoration efforts. To advance this goal, we used Pacific Biosciences HiFi long sequencing reads and Dovetail Omni-C proximity reads to generate a highly contiguous genome assembly that was then annotated using RNA-seq-informed gene prediction. The genome assembly is 484 Mb long, with contig N50 of 1.9 Mb, scaffold N50 of 21.8 Mb, BUSCO completeness score 96.1%, and 22 major scaffolds consistent with prior evidence that sea star genomes comprise 22 autosomes. These statistics generally fall between those of other recently assembled chromosome-scale assemblies for two species in the distantly related asteroid genus <em>Pisaster</em>. These novel genomic resources for <em>Pycnopodia helianthoides</em> will underwrite population genomic, comparative genomic, and phylogenomic analyses — as well as their integration across scales — of SSW and environmental stressors. This data resource contains the files associated with gene prediction.</div> </div> </div> </div>
Data from Genome scale metabolic network modelling for metabolic profile predictions
<p>Data used to produce figures 4, 5 and 6 in the paper Genome scale metabolic network modelling for metabolic profile predictions.</p>
A Study Evaluating Targeted Therapies in Participants Who Have Advanced Solid Tumors With Genomic Alterations or Protein Expression Patterns Predictive of Response
ClinicalTrials.gov study NCT04632992. IPD Sharing: YES. Countries: 1. Publications: 1.
Harnessing underutilized gene bank diversity and genomic prediction of cross usefulness to enhance resistance to Phytophthora cactorum in strawberry
Open the record for dataset details and reuse information.
Data from: Genome assembly of the ragweed leaf beetle, a step forward to better predict rapid evolution of a weed biocontrol agent to environmental novelties
Open the record for dataset details and reuse information.
Phenomic data-driven biological prediction of maize through field-based high throughput phenotyping integration with genomic data
Open the record for dataset details and reuse information.
Switchgrass flowering time measurements for genomic prediction
Open the record for dataset details and reuse information.
Data from: Phenotypic drought stress prediction of European beech (Fagus sylvatica) by genomic prediction and remote sensing
Open the record for dataset details and reuse information.
Gene prediction for: A reference genome for ecological restoration of the sunflower sea star, Pycnopodia helianthoides
Open the record for dataset details and reuse information.
Data from: Wheat genotypic and phenotypic data for multivariate genomic prediction
Open the record for dataset details and reuse information.
Combining climatic and genomic data improves range-wide tree height growth prediction in a forest tree
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.