Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
528
datasets available to search
ShareScore release 0.9.0
Dataset results
528 results for “gene prediction”
On the cross-population generalizability of gene expression prediction models
Open the record for dataset details and reuse information.
Within and cross species predictions of plant specialized metabolism genes using transfer learning
<p>Datasets for <em>Within and cross species predictions of plant specialized metabolism genes using transfer learning.</em> </p> <ol> <li>Dataset 1: All gene features used in the full feature machine learning models including expression, co-expression, evolutionary, duplication, and protein domain features.</li> <li>Dataset 2: All gene features used in shared feature machine learning models including expression, evolutionary, duplication, and protein domain features.</li> <li>Table S1: All gene annotations from TomatoCyc or manual annotation.</li> <li>Table S2: All model scores.</li> <li>Table S3: All gene scores and predictions from each model.</li> <li>Table S4: Feature importance for 5 models.</li> <li>Table S5: Statistical analysis between classes for binary and continuous feature data.</li> <li>Table S6: RNAseq datasets used in analysis.</li> </ol>
FunMap prediction scores for all gene pairs
<p>This gzipped TSV (Tab-Separated Values) file contains prediction scores generated by a predictive model trained on different types of feature set. The file structure is as follows:</p><p><strong>Column 1 (Gene Pair):</strong> This column represents pairs of gene ids. Each entry in this column consists of two gene ids, potentially denoting associations between genes.</p><p><strong>Column 2 (RNA Data Prediction):</strong> The values in this column represent prediction scores generated by the model using RNA features as input. </p><p><strong>Column 3 (Protein Data Prediction):</strong> This column contains prediction scores produced by the model when trained on protein features. </p><p><strong>Column 4 (Combined RNA and Protein Data Prediction):</strong> The values in this column represent prediction scores resulting from the model trained with both RNA and protein features.</p>
Convergent evolution and predictability of gene copy numbers associated with diets in mammals
<p>Convergent evolution, the evolution of the same or similar phenotypes in phylogenetically independent lineages, is a widespread phenomenon in nature. If the genetic basis for convergent evolution is predictable to some extent, it may be possible to infer organismic phenotypes and adaptability based on genome sequence data. While repeated amino acid changes have been studied in association with convergent evolution, relatively little is known about the potential contribution of repeated gene copy number changes. In this study, we explore whether certain gene copy number changes are linked to diet shifts in mammals and assess if trophic ecology can be inferred from the copy numbers of a specific set of genes. Using 86 mammalian genome sequences, we identified several genes with higher copy numbers in herbivores, carnivores, and omnivores, even after phylogenetic corrections. We were able to confirm previous findings on genes such as amylase, olfactory receptor, and xenobiotic metabolism genes, and identify novel genes whose copy numbers correlate with dietary patterns. For example, omnivores exhibited higher copy numbers of genes encoding gene expression regulators. We also established a discriminant function based on the copy numbers of 13 genes that can help predict trophic ecology based on genome sequence data. These findings highlight a possible association between convergent evolution and repeated copy number changes in specific genes, suggesting the potential to develop a method for predicting animal ecology and adaptability from genome sequence data.</p>
Data from: Does the number of functional olfactory receptor genes predict olfactory sensitivity and discrimination performance in mammals?
<p>The number of functional genes coding for olfactory receptors differs markedly between species and has repeatedly been suggested to be predictive of a species' olfactory capabilities. To test this assumption, we compiled a database of all published olfactory detection threshold values in mammals and used three sets of data on olfactory discrimination performance that employed the same structurally related monomolecular odor pairs with different mammal species. We extracted the number of functional olfactory receptor genes of the 20 mammal species for which we found data on olfactory sensitivity and/or olfactory discrimination performance from the Chordata Olfactory Receptor Database. We found that the overall olfactory detection thresholds significantly correlates with the number of functional olfactory receptor genes. Similarly, the overall proportion of successfully discriminated monomolecular odor pairs significantly correlates with the number of functional olfactory receptor genes. These results provide the first statistically robust evidence for the relation between olfactory capabilities and their genomics correlates. However, when analysed individually, of the 44 monomolecular odorants for which data on olfactory sensitivity from at least five mammal species are available, only five yielded a significant correlation between olfactory detection thresholds and the number of functional olfactory receptors genes. Also, for the olfactory discrimination performance, no significant correlation was found for any of the 74 relationships between the proportion of successfully discriminated monomolecular odor pairs and the number of functional olfactory receptor genes. While only a rather limited amount of data on olfactory detection thresholds and olfactory discrimination scores in a rather limited number of mammal species is available so far, we conclude that the number of functional olfactory receptor genes may be a predictor of olfactory sensitivity and discrimination performance in mammals.</p>
Supplementary materials for "Integration of public DNA methylation and expression networks via eQTMs improves prediction of functional gene–gene associations"
<p>This repository contained supplementary materials in the study named: "<strong>Integration of public DNA methylation and expression networks via eQTMs improves prediction of functional gene–gene associations.</strong>"</p> <p>For extracting all files from the downloaded tar.gz file, the following commanda could be used:</p> <pre><code>tar -xf supplementary_materials.tar.gz </code></pre> <p>The supplementary_materials/data diretcory contains the following sections:</p> <p></p> <ol> <li>eqtm_predictions: This contains the training and testing datasets for the eQTM prediction procedures</li> <li>public_methylation_data_and_pca: This contains the harmonized public DNA methylation dataset and its first 100 PCA components</li> <li>cca_data: This contains the CCA components for the public DNA methylation and gene expression datasets for the negative eQTMs, and the input datasets for the functional gene pair prediction analylsis.</li> <li>gene_enrichment_results_for_cca: This contains the gene enrichment results for the CCA components for negative and positive eQTMs</li> </ol> <p>The supplementary_materials/model directory contains the following sections:</p> <ol> <li>disease_tissue_predictions: This contains the models trained for tissue prediction and disease prediction based on the PCA components from the public DNA methylation data</li> <li>eqtm_predictions: This contains models trained for eQTM prediction</li> <li>cca_transformations: This contains CCA transformation models and models for STRING gene pair predictions </li> </ol>
Mimulus cardinalis plasticity analyses and R scripts for: Spatial variation in high temperature-regulated gene expression predicts evolution of plasticity with climate change in the scarlet monkeyflower
<p>A major way that organisms can adapt to changing environmental conditions is by evolving increased or decreased phenotypic plasticity. In the face of current global warming, more attention is being paid to the role of plasticity in maintaining fitness as abiotic conditions change over time. However, given that temporal data can be challenging to acquire, a major question is whether evolution in plasticity across space can predict adaptive plasticity across time. In growth chambers simulating two thermal regimes, we generated transcriptome data for western North American scarlet monkeyflowers (<i>Mimulus cardinalis</i>) collected from different latitudes and years (2010 and 2017) to test hypotheses about how plasticity in gene expression is responding to increases in temperature, and if this pattern is consistent across time and space. Supporting the genetic compensation hypothesis, individuals whose progenitors were collected from the warmer-origin northern 2017 descendant cohort showed lower thermal plasticity in gene expression than their cooler-origin northern 2010 ancestors. This was largely due to a change in response at the warmer (40ºC) rather than cooler (20ºC) treatment. A similar pattern of reduced plasticity, largely due to a change in response at 40ºC, was also found for the cooler-origin northern versus the warmer-origin southern population from 2017. Our results demonstrate that reduced phenotypic plasticity can evolve with warming and that spatial and temporal changes in plasticity predict one another.</p>
A benchmark study of ab initio gene prediction methods in diverse eukaryotic organisms
<p>G3PO (Gene and Protein Prediction PrOgrams) Benchmark was designed to represent many of the typical challenges faced by current genome annotation projects. The benchmark is based on a carefully validated and curated set of real eukaryotic genes from 147 phylogenetically disperse organisms (from human to protists). <br> </p>
Data for: Mating strategy predicts gene presence/absence patterns in a genus of simultaneously hermaphroditic flatworms
<p>This repository contains a record of analysis scripts and similarity score data used for the analyses presented in the manuscript.</p> <p>Some of the R scripts depend on supplementary tables associated with the manuscript.</p> <p>A preprint of the manuscript is available at: <a href="https://www.biorxiv.org/content/10.1101/2022.04.25.489193v2">https://www.biorxiv.org/content/10.1101/2022.04.25.489193v2</a></p>
Towards future directions in data-integrative supervised prediction of human aging-related genes
<p><strong>Corresponding author:</strong> <a href="http://www.nd.edu/~tmilenko">Prof. Tijana Milenković</a>, tmilenko AT nd DOT edu.</p> <p><strong>Detailed descriptions of the data and program</strong>: check <a href="https://nd.edu/~cone/DirectionsAging/">our paper's website</a>.</p> <p><strong>Reference</strong>: Qi Li, Khalique Newaz, and Tijana Milenković (2022). <strong>Towards future directions in data-integrative supervised prediction of human aging-related genes.</strong><em> under review</em></p>
Predicted genes annotated with Prokka for the 623 archaeal UBA genomes
<p>Gene prediction and annotation for the 623 archaeal UBA metagenome-assembled genomes using Prokka v1.12 with Pfam v31 and UniProt databases created on April 17, 2017 according to the Prokka instructions.</p> <p>These genomes are described in:</p> <p>Parks DH, et al. 2017. Recovery of nearly 8,000 metagenome-assembled genomes substantially expands the tree of life. Nat Microbiol, doi:10.1038/s41564-017-0012-7/ .</p> <p>https://www.nature.com/articles/s41564-017-0012-7</p>
Gene and protein sequence features augment HLA class I ligand predictions
<p>Dataset and analyses supporting the manuscript "Gene and protein sequence features augment HLA class I ligand predictions".</p> <p>The "peptides" files contain the mass-spec detected peptides obtained from HLA ligandomics performed on the indicated tumor lines. </p> <p>The "protein data" files contain the RNAseq data (TPM) and Ribosome profiling data (ribosome occupancy) per protein, for each tumor line. </p> <p>The "source data" zip archive contains the source data underlying the figures of the manuscript.</p> <p>The "HLA ligandome analyses" zip archive contains the R scripts used for all data analysis in the manuscript, including all data and output files. These analyses can also be found at https://github.com/kasbress/HLA_Ligandome_Analyses/</p> <p> </p> <p> </p>
Gene Enhancer Predictions in C2C12 (Mouse Myoblast)
<p>This file contains the complete enhancer prediction output of the analysis done in this study published on Nature Communications:</p> <p><a href="https://www.nature.com/articles/s41467-025-57758-x" target="_blank" rel="noopener">https://www.nature.com/articles/s41467-025-57758-x</a></p> <p>using the Activity by Contact Model following the pipeline described here: <a href="https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction">https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction</a></p> <p>To be used in this analysis, a genomewide Hi-C interaction matrix was generated with <a>Juicer </a> v1.6, expression counts for each gene were generated with STAR v2.7.10b and <a>Rsubread</a> v2.8.2, and the sequence alignment maps of ATAC-seq and H3K27ac ChIP-seq were generated with NextGenMap v0.5.5, all using raw sequencing reads downloaded from <a>SRA</a>: Hi-C (SRR16220088), RNA-seq (SRR074113 and SRR074114), <a>ATAC-seq (SRR2999996) and H3K27ac ChIP-seq (SRR358589, SRR358590, and SRR358591).</a></p> <p><a>The chromosome coordinates are of the GRCm38 - mm10 assembly, and the gene IDs are from ENSEMBL annotation.</a></p> <p>Please use the Version 3, and for more information please refer to our publication that used these enhancer predictions.</p>
The telomere regulatory gene POT1 responds to stress and predicts performance in nature: implications for telomeres and life history evolution
<p>Telomeres are emerging as correlates of fitness-related traits and may be important mediators of ecologically relevant variation in life history strategies. Growing evidence suggests that telomere dynamics can be more predictive of performance than length itself, but very little work considers how telomere regulatory mechanisms respond to environmental challenges or influence performance in nature. Here, we combine observational and experimental datasets from free-living tree swallows (<i>Tachycineta bicolor</i>) to assess how performance is predicted by the telomere regulatory gene POT1, which encodes a shelterin protein that sterically blocks telomerase from repairing the telomere. First, we show that lower POT1 gene expression was associated with higher female quality, <i>i.e.</i> earlier breeding and heavier body mass. We next challenged mothers with an immune stressor (lipopolysaccharide injection) that led to 'sickness' in mothers and 24h of food restriction in their offspring. While POT1 did not respond to maternal injection, females with lower constitutive POT1 gene expression were better able to maintain feeding rates following treatment. Maternal injection also generated a one-day stressor for chicks, which responded with lower POT1 gene expression and elongated telomeres. Other putatively stress-responsive mechanisms (i.e. glucocorticoids, antioxidants) showed marginal responses in stress-exposed chicks. Model comparisons indicated that POT1 mRNA abundance was a largely better predictor of performance than telomere dynamics, indicating that telomere regulators may be powerful modulators of variation in life history strategies.</p>
Saved model and preprocessed data for "CRMnet:a deep learning model for predicting gene expression from large regulatory sequence datasets"
<p>Saved TUNet model and preprocessed training data for "CRMnet: a deep learning model for predicting gene expression from large regulatory sequence datasets"</p> <p>To load the trained model:</p> <pre><code class="language-python">import tensorflow as tf tf.keras.models.load_model("path to the model folder")</code></pre> <p>for more information please find our repository: https://github.com/jiayuwen/CRMnet</p>
Speos: An ensemble graph representation framework to predict core genes for complex diseases (Datasets)
<p>the "data.tar.gz" tarball contains the unprocessed or minimally processed data used by Speos. If you intend to use the framework or want to inspect the data, download this part of the dataset.</p> <p>The "final_datasets.tar.gz" tarball contains tsv-formatted, processed data matrices which are directly used as input features for the ensemble models. </p> <p>There are two tsv-files per disease, one labeled "normal", which contains the p input features alongside the gene identifiers and a column which indicates if the gene is labeled as Mendelian or not, and another file labeled "with_n2v_vectors", which also contains the 100-dimensional vectors produced by Node2Vec so the N2V+MLP method can be reproduced with the exact same parameters. </p> <p>All files contain a header row which describes the column and no index column.</p> <p>The "model_parameters.tar.gz" tarball contains the model parameters for all ensemble models used to produce the candidate genes.</p>
MAC genome assembly and gene prediction of Tetrahymena thermophila SB210
<p>Corrected genome assembly and gene prediction of the MAC genome of T. thermophila SB210. These data were generated and analysed in the manuscript "Single-nucleotide polymorphism landscape of the macronuclear genome of <em>Tetrahymena thermophila".</em></p> <p>Please see the Material & methods and Supplementary data files of this manuscript for more details about these files.</p>
Deep model predictive control of gene expression in thousands of single cells
<p>Experimental data, training datasets, and trained models for our study on deep model predictive control of gene expression in bacteria. This data can be used to reproduce our results and figures.</p> <p>See our preprint here: <a href="https://www.biorxiv.org/content/10.1101/2022.10.28.514305">biorxiv.org/content/10.1101/2022.10.28.514305</a></p> <p>And our code repository here: <a href="https://gitlab.com/dunloplab/deepcellcontrol">gitlab.com/dunloplab/deepcellcontrol</a></p> <p><strong>Contents:</strong></p> <ul> <li><em>datasets.zip</em>: Formatted experimental data used to train and validate fluorescence forecasting models.</li> <li><em>experiments.zip:</em> Processed data for all control experiments.</li> <li><em>misc.zip</em>: Files necessary to reproduce some figures or to exactly reproduce some of our results.</li> <li><em>models.zip</em>: Trained neural network models and associated files.</li> </ul>
Models and Data associated with: Single-cell gene expression prediction from DNA sequence at large contexts
<p>This archive holds trained models and associated data for the <a href="https://www.biorxiv.org/content/10.1101/2023.07.26.550634v1">manuscript</a>:<br> "Single-cell gene expression prediction from DNA sequence at large contexts"</p> <p>Structure:</p> <ul> <li>configs - example configs for the workflows to produce publication data </li> <li>data_* - pre-processed single cell data used for publication</li> <li>models_* - model checkpoints, hyperparameters and training progress in tensorboard logs</li> <li>preprocessing - additional data required to reproduce the pre-processing workflow</li> </ul> <p> </p> <p>"Copyright 2023 GlaxoSmithKline Research & Development Limited. All rights reserved."</p>
Gene prediction for: A reference genome for ecological restoration of the sunflower sea star, Pycnopodia helianthoides
<div> <div> <div> <div>Wildlife diseases, such as the sea star wasting (SSW) epizootic that outbroke in the mid-2010s, appear to be associated with acute and/or chronic abiotic environmental change; dissociating the effects of different drivers can be difficult. The sunflower sea star,<em> Pycnopodia helianthoides</em>, was the species most severely impacted during the SSW outbreak, which overlapped with periods of anomalous atmospheric and oceanographic conditions, and there is not yet a consensus on the cause(s). Genomic data may reveal underlying molecular signatures that implicate a subset of factors and, thus, clarify past events while also setting the scene for effective restoration efforts. To advance this goal, we used Pacific Biosciences HiFi long sequencing reads and Dovetail Omni-C proximity reads to generate a highly contiguous genome assembly that was then annotated using RNA-seq-informed gene prediction. The genome assembly is 484 Mb long, with contig N50 of 1.9 Mb, scaffold N50 of 21.8 Mb, BUSCO completeness score 96.1%, and 22 major scaffolds consistent with prior evidence that sea star genomes comprise 22 autosomes. These statistics generally fall between those of other recently assembled chromosome-scale assemblies for two species in the distantly related asteroid genus <em>Pisaster</em>. These novel genomic resources for <em>Pycnopodia helianthoides</em> will underwrite population genomic, comparative genomic, and phylogenomic analyses — as well as their integration across scales — of SSW and environmental stressors. This data resource contains the files associated with gene prediction.</div> </div> </div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.