Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

528

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

528 results for “gene prediction”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Genes predict long distance migration and large body size in a migratory fish, Pacific lamprey

Open the record for dataset details and reuse information.

publicAug 2014View details →
dryad28/100

Data from: Repertoire-wide gene structure analyses: a case study comparing automatically predicted and manually annotated gene models

The location and modular structure of eukaryotic protein-coding genes in genomic sequences can be automatically predicted by gene annotation algorithms. These predictions are often used for comparative studies on gene structure, gene repertoires, and genome evolution. However, automatic annotation algorithms do not yet correctly identify all genes within a genome, and manual annotation is often necessary to obtain accurate gene models and gene sets. As manual annotation is time-consuming, only a fraction of the gene models in a genome is typically manually annotated, and this fraction often differs between species. To assess the impact of manual annotation efforts on genome-wide analyses of gene structural properties, we compared the structural properties of protein-coding genes in seven diverse insect species sequenced by the i5k initiative. Our results show that the subset of genes chosen for manual annotation by a research community (3.5-7% of gene models) may have structural properties (e.g., lengths and exon counts) that are not necessarily representative for a species' gene set as a whole. Nonetheless, the structural properties of automatically generated gene models are only altered marginally (if at all) through manual annotation. Major correlative trends, for example a negative correlation between genome size and exonic proportion, can be inferred from either the automatically predicted or manually annotated gene models alike. Vice versa, some previously reported trends did not appear in either the automatic or manually annotated gene sets, pointing towards insect-specific gene structural peculiarities. In our analysis of gene structural properties, automatically predicted gene models proved to be sufficiently reliable to recover the same gene-repertoire-wide correlative trends that we found when focusing on manually annotated gene models only. We acknowledge that analyses on the individual gene level clearly benefit from manual curation. However, as genome sequencing and annotation projects often differ in the extent of their manual annotation and curation efforts, our results indicate that comparative studies analyzing gene structural properties in these genomes can nonetheless be justifiable and informative.

opencc-zeroAug 2020View details →
dryad28/100

Data from: Integration of genomics and transcriptomics predicts diabetic retinopathy susceptibility genes

<p class="Normal1">We determined differential gene expression in response to high glucose in lymphoblastoid cell lines derived from matched individuals with type 1 diabetes with and without retinopathy. Those genes exhibiting the largest difference in glucose response were assessed for association to diabetic retinopathy in a genome-wide association study meta-analysis. Expression Quantitative Trait Loci (eQTLs) of the glucose response genes were tested for association with diabetic retinopathy. We detected an enrichment of the eQTLs from the glucose response genes among small association p-values and identified <i>FLCN</i> as a susceptibility gene for diabetic retinopathy. Expression of <i>FLCN </i>in response to glucose was greater in individuals with diabetic retinopathy. Independent cohorts of individuals with diabetes revealed an association of <i>FLCN</i> eQTLs to diabetic retinopathy. Mendelian randomization confirmed a direct positive effect of increased <i>FLCN</i> expression on retinopathy. Integrating genetic association with gene expression implicated <i>FLCN </i>as a disease gene for diabetic retinopathy.</p>

opencc-zeroNov 2020View details →
dryad28/100

Data from: Variant at serotonin transporter gene predicts increased imitation in toddlers: relevance to the human capacity for cumulative culture

Cumulative culture ostensibly arises from a set of sociocognitive processes which includes high-fidelity production imitation, prosociality and group identification. The latter processes are facilitated by unconscious imitation or social mimicry. The proximate mechanisms of individual variation in imitation may thus shed light on the evolutionary history of the human capacity for cumulative culture. In humans, a genetic component to variation in the propensity for imitation is likely. A functional length polymorphism in the serotonin transporter gene, the short allele at 5HTTLPR, is associated with heightened responsiveness to the social environment as well as anatomical and activational differences in the brain's imitation circuity. Here, we evaluate whether this polymorphism contributes to variation in production imitation and social mimicry. Toddlers with the short allele at 5HTTLPR exhibit increased social mimicry and increased fidelity of demonstrated novel object manipulations. Thus, the short allele is associated with two forms of imitation that may underlie the human capacity for cumulative culture. The short allele spread relatively recently, possibly due to selection, and its frequency varies dramatically on a global scale. Diverse observations can be unified via conceptualization of 5HTTLPR as influencing the propensity to experience others' emotions, actions and sensations, potentially through the mirror mechanism.

opencc-zeroDec 2015View details →
dryad28/100

The predicted haploid gene set of the genome of Nitzschia putrida

<p>Secondary loss of photosynthesis is observed across almost all plastid-bearing branches of the eukaryotic tree of life. However, genome-based insights into the transition from a phototroph into a secondary heterotroph have so far only been revealed for parasitic species. Free-living organisms can yield unique insights into the evolutionary consequence of the loss of photosynthesis, as the parasitic lifestyle requires specific adaptations to host environments. Here we report on the diploid genome of the free-living diatom <i>Nitzschia putrida </i>(35 Mbp), a non-photosynthetic osmotroph whose photosynthetic relatives contribute ca. 40% of net oceanic primary production. Comparative analyses with photosynthetic diatoms and heterotrophic algae with parasitic lifestyle revealed that a combination of gene loss, the accumulation of genes involved in organic carbon degradation, a unique secretome and the rapid divergence of conserved gene families involved in cell wall and extracellular metabolism appear to have facilitated the lifestyle of a free-living secondary heterotroph.</p>

opencc-zeroMar 2022View details →
zenodo28/100

Methylation haplotypes of the insulin gene promoter in children and adolescents with type 1 diabetes: could a dimensionality reduction approach predict the disease?

<p>The aim of the present study was to identify insulin gene promoter (IGP) methyl-haplotypes among children and adolescents with T1D and suggest a predictive model for the discrimination of cases and controls according to methyl-haplotypes. Fourty individuals (20 T1D) participated. IGP-region from peripheral whole blood DNA of 40 participants (20 T1D) was sequenced by next generation sequencing, sequences were read using FASTQ files, and methylation status was calculated by python-based pipeline for targeted deep bisulfite sequenced amplicons (ampliMethProfiler). Methylation profile at 10 CpG sites proximal to transcription start site of the IGP was recorded and coded as 0 for unmethylation or 1 for methylation. A single read could result in &ldquo;1111111111&rdquo; methyl-haplotype (all methylated), &ldquo;000000000&rdquo; methyl-haplotype (all unmethylated) or any other combination.</p>

opencc-by-4.0Jun 2023View details →
ClinicalTrials.gov28/100

Identification of Genes That Predict Local Recurrence in Samples From Patients With Breast Cancer Treated on NSABP-B-28

ClinicalTrials.gov study NCT01420185. IPD Sharing: Not stated. Countries: 0. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov28/100

Gene Expression Levels in Predicting Treatment Response in Patients With Stage IV Non-small Cell Lung Cancer

ClinicalTrials.gov study NCT02145078. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov28/100

Genes in Predicting Outcome of Patients With DLBCL Treated With Rituximab and Combination Chemotherapy (R-CHOP)

ClinicalTrials.gov study NCT00450385. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov28/100

Identifying Genes That Predict Recurrence in Women With Breast Cancer Treated With Chemotherapy

ClinicalTrials.gov study NCT00897299. IPD Sharing: Not stated. Countries: 0. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad28/100

Data from: Predicting gene function from uncontrolled expression variation among individual wild-type Arabidopsis plants

Open the record for dataset details and reuse information.

publicMar 2016View details →
dryad28/100

Data from: Genes and group membership predict gidgee skink (Egernia stokesii) reproductive pairs

Open the record for dataset details and reuse information.

publicMar 2017View details →
dryad28/100

Data from: Variant at serotonin transporter gene predicts increased imitation in toddlers: relevance to the human capacity for cumulative culture

Open the record for dataset details and reuse information.

publicMar 2016View details →
dryad28/100

Data from: Gene prediction and annotation in Penstemon (Plantaginaceae): a workflow for marker development from extremely low-coverage genome sequencing

Open the record for dataset details and reuse information.

publicOct 2015View details →
dryad28/100

The predicted haploid gene set of the genome of Nitzschia putrida

Open the record for dataset details and reuse information.

publicMar 2022View details →
dryad28/100

Data from: Repertoire-wide gene structure analyses: a case study comparing automatically predicted and manually annotated gene models

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad28/100

Data from: Integration of genomics and transcriptomics predicts diabetic retinopathy susceptibility genes

Open the record for dataset details and reuse information.

publicNov 2020View details →
geo24/100

Dicer deficiency reveals microRNAs predicted to control gene expression in developing adrenal cortex [RT-PCR]

GEO Series GSE45810. Mus musculus. 16 samples. Type: Expression profiling by RT-PCR.

openGEO-OpenMay 2013View details →
geo24/100

Contradictory of gene expression signatures for clinical prognostic prediction caused by update of reference genome and annotation

GEO Series GSE143486. Homo sapiens. 30 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2020View details →
geo24/100

Immune gene signatures for predicting durable clinical benefit of anti-PD-1 immunotherapy in patients with non-small cell lung cancer

GEO Series GSE136961. Homo sapiens. 21 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record