Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
528
datasets available to search
ShareScore release 0.9.0
Dataset results
528 results for “gene prediction”
Comorbid-phenome prediction and phenotype risk scores enhance gene discovery for generalized anxiety disorder and posttraumatic stress disorder
<p>GWAS summary statistics. See ReadMe for column descriptions.</p>
A new gene set identifies senescent cells and predicts senescence-associated pathways across tissues
<p>Although cellular senescence drives multiple age-related co-morbidities through the senescence-associated secretory phenotype (SASP), <em>in vivo</em> senescent cell identification remains challenging. Here, we generated a gene set (SenMayo) and validated its enrichment in bone biopsies from two aged human cohorts. We further demonstrated reductions in SenMayo in bone following genetic clearance of senescent cells in mice and in adipose tissue from humans following pharmacological senescent cell clearance. We next used SenMayo to identify senescent hematopoietic or mesenchymal cells at the single cell level from human and murine bone marrow/bone scRNA-seq data. Thus, SenMayo identifies senescent cells across tissues and species with high fidelity. Using this senescence panel, we were able to characterize senescent cells at the single-cell level and identify key intercellular signaling pathways. SenMayo also represents a potentially clinically applicable panel for monitoring senescent cell burden with aging and other conditions as well as in studies of senolytic drugs.</p>
Predicting the secondary metabolic potential of microbiomes from marker genes using PSMPA
<p>Supplementary data for the paper <strong>Predicting the secondary metabolic potential of microbiomes </strong><strong>from marker genes using PSMPA</strong>.</p>
Kroll et al., 2024. Behavioural pharmacology predicts disrupted signalling pathways and candidate therapeutics from zebrafish mutants of Alzheimer's disease risk genes
<p>Data repository for</p> <p>François Kroll, Joshua Donnelly, Joshua Donnelly, Güliz Gürel Özcan, Eirinn Mackay, Jason Rihel</p> <div> <div><strong>Behavioural pharmacology predicts disrupted signalling pathways and candidate therapeutics from zebrafish mutants of Alzheimer’s disease risk genes</strong></div> <br> <div>eLife, 2024.</div> <br> <div><a href="https://doi.org/10.7554/eLife.96839.1">https://doi.org/10.7554/eLife.96839.1</a></div> <div> </div> <div>Code is found in the <a href="https://github.com/francoiskroll/ZFAD">GitHub repository</a>. Please see notes there.</div> </div> <p>___</p> <p>Contact:</p> <p>Twitter: @francois_kroll</p> <p>Email: francois@kroll.be</p>
A Riskscore Model for Predicting Survival, Tumor Microenvironment, Immunotherapy and drug sensitivity of Lung Squamous Cell Carcinoma Based on PI3K/AKT/MTOR Pathway-Related Genes
<p>Firstly, the data we provide is the raw data downloaded from the TCGA database. Secondly, we provide the following explanations for the raw data, taking Figure 1 as an example:</p> <p>First, we selected a LUSC dataset from the TCGA database that includes RNA-seq data (FPKM values) and clear clinical information, consisting of 51 normal samples and 501 LUSC samples. The raw data <span>can be found in</span> the "mRNA" document in the "Fig.1" file<span>, and t</span>he data includes TCGA <span>id</span> information.</p> <p>Next, the data was imported into "mRNA_edgeR" to analyze whether the RNA-seq data from these 552 samples meet the criteria of |logFC| > 0.585 and FDR < 0.05, <span>and the results are shown in</span> the "diffSig" table. Subsequently, the genes in the "diffSig" table were intersected with 105 PAGs, <span>and</span> 44 key PAGs <span>were identified, </span>as shown in Figure 1.</p> <p>The <span>other</span> figures can <span>also </span>be reproduced sequentially based on the methods described <span>above</span>.</p> <p>Furthermore, due to the limited number of LUSC patients, the clinical information in the database is relatively incomplete. Hence, the clinical information that we provide constitutes the entire content available in the database.</p>
Multimodal contrastive learning for spatial gene expression prediction using histology images
<p>we employed two human breast cancer datasets and one human cutaneous squamous cell carcinoma (cSCC) dataset.</p>
Gene predictions in GFF format for tuatara genome assembly QEPC00000000.1
<p><strong>Maker gene predictions for Sphenodon punctatus (tuatara) isolate: mauimua-1 </strong> <br> <br> These GFF files correspond to the genome assembly in GenBank ID QEPC00000000.1<br> https://www.ncbi.nlm.nih.gov/nuccore/QEPC00000000.1</p> <p>Files:</p> <p> <strong>MASKED_Tuatara.DEVO.annot15102.gff.gz</strong><br> - gene and transcript predictions only<br> <br> <strong>20160427.tuatara.maker.raw.gff3.gz</strong><br> - raw Maker output including gene and transcript predictions, alignment results and output of individual gene prediction algorithms (SNAP and Augustus).</p> <p><br> <br> </p>
Accurate genome-wide predictions of spatio-temporal gene expression during embryonic development
<p>This upload contains the expression prediction dataset discussed in the manuscript "Accurate genome-wide predictions of spatio-temporal gene expression during embryonic development" and used by the webserver https://find.princeton.edu.</p> <p>Abstract:</p> <p>Comprehensive information on the timing and location of gene expression is fundamental to our understanding of embryonic development and tissue formation. While high-throughput <em>in situ</em> hybridization projects provide invaluable information about developmental gene expression patterns for model organisms like <em>Drosophila</em>, the output of these experiments is primarily qualitative, and a high proportion of protein coding genes and most non-coding genes lack any annotation. Accurate data-centric predictions of spatio-temporal gene expression will therefore complement current <em>in situ</em> hybridization efforts. Here, we applied a machine learning approach by training models on all public gene expression and chromatin data, even from whole-organism experiments, to provide genome-wide, quantitative spatio-temporal predictions for all genes. We developed structured in silico nano-dissection, a computational approach that predicts gene expression in >200 tissue-developmental stages. The algorithm integrates expression signals from a compendium of 6,378 genome-wide expression and chromatin profiling experiments in a cell lineage-aware fashion. We systematically evaluated our performance via cross-validation and experimentally confirmed 22 new predictions for four different embryonic tissues. The model also predicts complex, multi-tissue expression and developmental regulation with high accuracy. We further show the potential of applying these genome-wide predictions to extract tissue specificity signals from non-tissue-dissected experiments, and to prioritize tissues and stages for disease modeling. This resource, together with the exploratory tools are freely available at our webserver <a href="http://find.princeton.edu/">http://find.princeton.edu</a>, which provides a valuable tool for a range of applications, from predicting spatio-temporal expression patterns to recognizing tissue signatures from differential gene expression profiles.</p>
Analysis of opposing histone modifications H3K4me3 and H3K27me3 reveals candidate diagnostic biomarkers for triple negative breast cancer and gene set prediction combinations
<p>Supplementary data of manuscript <strong>'Analysis of opposing histone modifications H3K4me3 and H3K27me3 reveals candidate diagnostic biomarkers for triple negative breast cancer and gene set prediction combinations </strong><strong>Analysis of opposing histone modifications H3K4me3 and H3K27me3 reveals candidate diagnostic biomarkers for triple negative breast cancer and gene set prediction combinations'</strong></p>
Gene Enhancer Predictions in C2C12 (Mouse Myoblast)
<p>This file contains the complete enhancer prediction output of the analysis done in this study published on Nature Communications:</p> <p><a href="https://www.nature.com/articles/s41467-025-57758-x" target="_blank" rel="noopener">https://www.nature.com/articles/s41467-025-57758-x</a></p> <p>using the Activity by Contact Model following the pipeline described here: <a href="https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction">https://github.com/broadinstitute/ABC-Enhancer-Gene-Prediction</a></p> <p>To be used in this analysis, a genomewide Hi-C interaction matrix was generated with <a>Juicer </a> v1.6, expression counts for each gene were generated with STAR v2.7.10b and <a>Rsubread</a> v2.8.2, and the sequence alignment maps of ATAC-seq and H3K27ac ChIP-seq were generated with NextGenMap v0.5.5, all using raw sequencing reads downloaded from <a>SRA</a>: Hi-C (SRR16220088), RNA-seq (SRR074113 and SRR074114), <a>ATAC-seq (SRR2999996) and H3K27ac ChIP-seq (SRR358589, SRR358590, and SRR358591).</a></p> <p><a>The chromosome coordinates are of the GRCm38 - mm10 assembly, and the gene IDs are from ENSEMBL annotation.</a></p> <p> </p> <p>This research was funded in whole or in part by the Austrian Science Fund (FWF) [P29713-B28, P32512-B and P36503-B] to Roland Foisner and a doctorate program funded by the Austrian Science Fund (FWF) [W1261-B28].</p> <p> </p> <p>For more information please refer to our publication that used these enhancer predictions titled:</p> <p>MyoD1 localization at the nuclear periphery is mediated by association of WFS1 with active enhancers</p>
Predicting the pro-longevity or anti-longevity effect of model organism genes with enhanced Gaussian noise augmentation-based contrastive learning on protein-protein interaction networks
<p>The datasets used to evaluate Enhanced Gaussian noise augmentation-based contrastive learning (EGsCL) against predicting the pro-longevity or anti-longevity effect of model organism gene. This repo also includes the pretrained encoders that obtained the best predictive performance for each organism (see Table 2).</p>
Data from: Immune response genes and pathogen presence predict migration survival in wild salmon smolts
We present the first data to link physiological responses and pathogen presence with subsequent fate during migration of wild salmonid smolts. We tagged and non-lethally sampled gill tissue from sockeye salmon (Oncorhynchus nerka) smolts as they left their nursery lake (Chilko Lake, BC, Canada) to compare gene expression profiles and freshwater pathogen loads with migration success over the first ~1150 km of their migration to the North Pacific Ocean using acoustic telemetry. Fifteen percent of smolts were never detected again after release and these fish had gene expression profiles consistent with an immune response to one or more viral pathogens compared with fish that survived their freshwater migration. Among the significantly up-regulated genes of the fish that were never detected post-release were MX (Interferon-induced GTP-binding Protein Mx) and STAT1 (Signal transducer and activator of transcription 1-alpha/beta), which are characteristic of a type I interferon response to viral pathogens. The most commonly detected pathogen in the smolts leaving the nursery lake was infectious hematopoietic necrosis virus (IHNV). Collectively, these data show that some of the fish assumed to have died after leaving the nursery lake appeared to be responding to one or more viral pathogens and had elevated stress levels that could have contributed to some of the mortality shortly after release. We present the first evidence that changes in gene expression may be predictive of some of the fresh water migration mortality in wild salmonid smolts.
Genes with transcripts predicted to be miR-195 or miR-26b targets by all 5 predictive algorithms included in starBase, or experimentally identified as targets by pulldown assay
<p>Genes with transcripts predicted to be miR-195 or miR-26b targets by all 5 predictive algorithms included in starBase, or experimentally identified as targets by pulldown assay</p>
Gene structure prediction results and execution commands and options by GINGER and similar tools
<p>Gene structure prediction results, execution commands, and options by GINGER and similar tools</p> <ul> <li>genome.tar.gz ... genome sequence data used for benchmark test in the paper. </li> <li>input_of_EVM.tar.gz ... input dataset for EVM</li> <li>input_of_GINGER.tar.gz ... input dataset for GINGER</li> <li>input_of_MAKER.tar.gz ... input dataset for MAKER</li> <li>reference_annotation.tar.gz ... gff data used for benchmark test in the paper.</li> <li>result_of_EVM.tar.gz .... results and stats information of EVM (including intermediate result file of <em>C.elegans</em>, <em>D. melanogaster</em>, <em>O. sativa</em>, <em>D. rerio</em>, and <em>H. sapiens</em>)</li> <li>result_of_GINGER.tar.gz .... results and stats information of GINGER (including intermediate result files of <em>C.elegans</em>)</li> <li>result_of_MAKER.tar.gz .... results and stats information of MAKER (including intermediate result files of <em>C.elegans</em>)</li> </ul>
DeepARG: a deep learning approach for predicting antibiotic resistance genes from metagenomic data
<p>Database and models for deepARG: </p> <p>see this link for details: <a href="https://bitbucket.org/gusphdproj/deeparg-ss/src/fbe063e24cf79d83a88499353aa15a85b58a300e/?at=master">gusphdproj / deeparg-ss — Bitbucket</a></p>
Nicotiana model for augustus gene prediction
<p>Nicotiana model for augustus gene prediction.</p>
Gene Expression Profiles in Predicting Chemotherapy Response in Breast Cancer
ClinicalTrials.gov study NCT00212082. IPD Sharing: Not stated. Countries: 1. Publications: 6.
Analyze the Predictive Value of Gene TMPRSS2-ETS in Response to Enzalutamide in Patients With Prostate Cancer
ClinicalTrials.gov study NCT02288936. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Validate Gene Expression and Proteomic Signatures Predictive of Treatment for Response for Breast Cancer Patient
ClinicalTrials.gov study NCT00669773. IPD Sharing: Not stated. Countries: 1. Publications: 3.
Gene Expression in Cumulus Cells to Predict Pregnancy
ClinicalTrials.gov study NCT01732900. IPD Sharing: Not stated. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.