Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

665

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

665 results for “exon”

Learn how ShareScore rates datasets ↗
zenodo36/100

Intronization signatures in coding exons reveal the evolutionary fluidity of eukaryotic gene architecture

<p>This dataset accompanies the manuscript &quot;Intronization signatures in coding exons reveal the evolutionary fluidity of eukaryotic gene architecture&quot;. It contains the relevant tables, scripts and figures used for and created during data analysis.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Huntingtin Exon 1 Q16 and Q46 pET32a thioredoxin tagged expression and purification 2018/04/09

<p>Huntingtin structure-function open lab notebook project.&nbsp;Huntingtin Exon 1 Q16 and Q46 pET32a thioredoxin tagged expression and purification 2018/04/09.</p>

opencc-by-4.0Apr 2018View details →
zenodo36/100

Huntingtin Exon 1 Q16 and Q46 pET32a thioredoxin tagged expression and purification 2018/04/02

<p>Huntingtin structure-function open lab notebook project. Huntingtin Exon 1 Q16 and Q46 pET32a thioredoxin tagged, attempted&nbsp;expression and purification 2018/04/02</p>

opencc-by-4.0Apr 2018View details →
zenodo36/100

Electrophysiology Data for "Two functional epithelial sodium channel isoforms are present in rodents despite pronounced evolutionary pseudogenization and exon fusion"

<p>Here we provide&nbsp;the electrophysiology data for&nbsp;the manuscript &quot;Two functional epithelial sodium channel isoforms are present in rodents despite pronounced evolutionary pseudogenization and exon fusion&quot;, published in Molecular Biology and Evolution (2021):&nbsp;msab271 (doi: 10.1093/molbev/msab271).&nbsp;Data are reported as current values in Excel format, sorted according to the appearance in Figures and supplemented by explanatory text on the procedures/data presentation.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

The impact of genetically controlled splicing on exon inclusion and protein structure

<p><strong>This repository contains raw and processed files used in Einson et. al 2022. </strong></p> <p>Code used to generate these files can be found here:&nbsp;https://github.com/jeinson/sqtl_manuscript</p> <p><strong><em>Descriptions of files contained within each sub directory</em></strong></p> <p><strong>01_raw_psi</strong></p> <ul> <li><em>{GTEx_tissue_id}_v8.psi.tsv.gz:&nbsp;</em>Unfiltered PSI output from IPSA-nf, per tissue. See methods for details about how files were created.&nbsp;</li> <li><em>gtex_v8_exon_id_map.tsv:&nbsp;</em>Mapping file between exon coordinates and Ensembl gene IDs, with suffix used in GTEx v8 gencode annotation.&nbsp;</li> </ul> <p><strong>02_qtl_results</strong></p> <ul> <li><strong>cross_tissue</strong> <ul> <li><em>top_sQTLs_MAF05.tsv:&nbsp;</em>List of top GTEx v8 sQTLs across tissues, with one exon and top variant per tissue. See methods for details. See matching file for column descriptions.&nbsp;</li> <li><em>top_sQTLs_median_psi.tsv:&nbsp;</em>The median, mean, and standard deviation of PSI of each significant exon from the previous file, taken across all individuals from GTEx with data available.</li> <li><em>top_sQTLs_MAF05_w_anc_allele.tsv:&nbsp;</em>List of top sQTLs across tissues, with additional columns for the top &psi;QTL ancestral and derived alleles, where available.&nbsp;</li> </ul> </li> <li><strong>per_tissue</strong> <ul> <li><em>{GTEx_tissue_id}_combined_sQTLs.tsv.gz:&nbsp;</em>Raw output of &psi;QTL calling using&nbsp;QTLtools in grouped permutational mode per&nbsp;tissue, with groups specified by gene. See methods for more details, and&nbsp;https://qtltools.github.io/qtltools/ for column descriptions.&nbsp;</li> </ul> </li> </ul> <p><strong>03_qtl_credible_sets</strong></p> <ul> <li>&nbsp;<em>GTEx_psi_{GTEx_tissue_id}.collapsed.txt.gz:&nbsp;</em>Output of the QTL catalog fine mapping pipeline (https://github.com/eQTL-Catalogue/qtlmap), run on all exons and tissues, and collapsed using the procedure described in Methods.&nbsp;</li> </ul> <p><strong>04_qtl_coloc</strong></p> <ul> <li><em>combined_coloc_results_full.tsv.gz:&nbsp;</em>Combined output of running coloc on &psi;QTLs from the 18 GTEx tissues against 87 sets of GWAS summary statistics. This file contains all results, including non-significant associations. A nominal QTLtools pass was used as input. We do not include these files in this repository due to size limitations, but contact the authors if you need access to nominal QTL calls.&nbsp;</li> <li><em>top_sQTLs_with_top_coloc_event.tsv:&nbsp;</em>The QTLs in&nbsp;<em>top_sQTLs_MAF05.tsv</em>&nbsp;with additional columns for the GWAS with the highest posterior probability of a colocalization event. Importantly, the tissue and top variant may not match the main&nbsp;<em>top_sQTLs_MAF05.tsv&nbsp;</em>file for every gene.&nbsp;</li> </ul> <p><strong>05_exon_features:&nbsp;</strong>See matching files for description of each column.&nbsp;</p> <ul> <li><em>cross_tissue_constitutive_exons_with_AF.tsv:&nbsp;</em>Detailed features of cross tissue constitutive exons. See methods for definition of constitutive exons.&nbsp;</li> <li><em>cross_tissue_nonsignificant_genes_with_AF.tsv:&nbsp;</em>Detailed features of sufficiently variable exons with no significant variant across tissues. See methods for more details.&nbsp;</li> <li><em>top_sQTLs_MAF05_with_AF.tsv:&nbsp;</em>Detailed features of top sQTLs.&nbsp;</li> <li><em>top_sQTLs_with_top_coloc_with_AF.tsv:&nbsp;</em>Detailed features of sQTLs that colocalize with at least one GWAS trait. Contains columns for Euclidean distances between PAE matrices and RMSD between isoforms, among genes with a significant GWAS colocalization event.&nbsp;</li> </ul> <p><strong>06_predicted_structures:&nbsp;</strong>Each prediction was run 5 times, and we report the best model in the manuscript.&nbsp;</p> <ul> <li><strong>{protein.id}[_mutant].result</strong> <ul> <li><em>{protein.id}[_mutant]{_run.id}_coverage.png.gz:&nbsp;</em>Plot of the number of sequences per position in MSA</li> <li><em>{protein.id}[_mutant]{_run.id}_PAE.png.gz:&nbsp;</em>PAE matrix plots for each model</li> <li><em>{protein.id}[_mutant]{_run.id}_plddt.png.gz:&nbsp;</em>pLDDT plots for each model</li> <li><em>{protein.id}[_mutant]{_run.id}_predicted_aligned_error_v1.json.gz</em>: A PAE matrix for the best model using <a href="https://alphafold.ebi.ac.uk/faq#faq-7">AlphaFold-DB&#39;s format</a></li> <li><em>{protein.id}[_mutant]{_run.id}_unrelaxed_rank_{rank.num}_model_{model.num}_scores.json.gz</em>: Per model array (list of lists) with PAE, a list of the average pLDDT and the pTM score.&nbsp;</li> <li><em>{protein.id}[_mutant]{_run.id}_unrelaxed_rank_{rank.num}_model_{model.num}_pdb.gz:&nbsp;</em>Per model predicted structure in pd format</li> <li><em>{protein.id}[_mutant]{_run.id}.a3m.gz</em>: A3M formatted input MSA</li> <li><em>cite.bibtex:&nbsp;</em>BibTex file with citations for all used tools and databases</li> <li><em>config.json</em>: Model input parameters</li> </ul> </li> </ul> <p><strong>07_other_data</strong></p> <ul> <li><em>cross_tissue_constitutive_exons.tsv:&nbsp;</em>List of exons that are constitutively spliced across multiple tissues. See methods for details.&nbsp;</li> <li><em>cross_tissue_nonsignificant_genes.tsv</em>: List variably spliced exons with no significant sVariant in any tissue. See methods for details.&nbsp;</li> <li><em>gtex_v8_exon_id_map.rds:&nbsp;</em>rds representation of a map between exon IDs, as used in the modified version of gencode v26, and exon hg38 coordinates.&nbsp;</li> <li><em>gtex_v8_n_exons_per_gene.tsv:&nbsp;</em>Number of exons per gene, as annotated in the modified version of gencode v26 used in GTEx v8.&nbsp;</li> </ul> <p><strong>08_geuvadis</strong></p> <ul> <li><em>geuvadis_psi.tsv.gz:&nbsp;</em>Unfiltered PSI output from IPSA-nf, run on Geuvadis BAM files. See methods for details. (Raw data was downloaded from ftp://ftp.ebi.ac.uk/pub/databases/microarray/data/experiment/GEUV)</li> <li><em>geuvadis_sQTLs.tsv.gz:&nbsp;</em>Raw output of &psi;QTL calling using&nbsp;QTLtools in grouped permutational mode for geuvadis data, with groups specified by gene. See methods for more details, and&nbsp;https://qtltools.github.io/qtltools/ for column descriptions.&nbsp;</li> <li><em>remapped_gencode.v26.GRCh37.GTEx_v8.nochr.genes.gtf.gz:</em>&nbsp;Lifted over version of the gencode v26 gtf file, used to define exons for PSI and qtl mapping in the geuvadis analysis. The original version that was used in the GTEx analysis is based on GRCh38, and is available here:&nbsp;https://storage.googleapis.com/gtex_analysis_v8/reference/gencode.v26.GRCh38.genes.gtf</li> </ul>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Sequences of target genes and probes for construction of exon-targeted libraries for Prunus spp.

<p>(Files)</p> <p>&quot;baits-selection-15171.fas&quot;: Sequences of 15,171 baits used for hybridization reaction.</p> <p>&quot;Targets_mume.15171.xlsx&quot;: Sequences of target genes used for design baits.</p> <p>&nbsp;</p> <p>(Protocol for bait design)</p> <p>To selectively retrieve libraries with exons, a myBaits Custom design kit was used to design 1&ndash;20-K probes (Arbor Biosciences, Ann Arbor, MI, USA), which uses biotinylated RNA probes to concentrate fragments carrying sequences of interest (Gnirke et al., 2009), based on the published genomic and coding sequences of <em>P. mume</em> (Zhang et al., 2012). We subjected 29,621 non-redundant coding loci to search for single hits with BLAST+ (MEGABLAST with -p 70) against the <em>P. mume</em> genome, for the subsequent bait design. A 120-mer bait with 25&ndash;55 GC% per locus was randomly designed for each locus, and finally we obtained a bait set targeting 15,171 coding loci.</p>

opencc-by-4.0Jul 2023View details →
ClinicalTrials.gov36/100

Phase 2 Study of Poziotinib in Participants With NSCLC Having EGFR or HER2 Exon 20 Insertion Mutation

ClinicalTrials.gov study NCT03318939. IPD Sharing: NO. Countries: 8. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

AAV9 U7snRNA Gene Therapy to Treat Boys With DMD Exon 2 Duplications.

ClinicalTrials.gov study NCT04240314. IPD Sharing: Not stated. Countries: 1. Publications: 5.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

TAK-788 as First-Line Treatment Versus Platinum-Based Chemotherapy for Non-Small Cell Lung Cancer (NSCLC) With EGFR Exon 20 Insertion Mutations

ClinicalTrials.gov study NCT04129502. IPD Sharing: YES. Countries: 24. Publications: 1.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

Study of Eteplirsen in Young Participants With Duchenne Muscular Dystrophy (DMD) Amenable to Exon 51 Skipping

ClinicalTrials.gov study NCT03218995. IPD Sharing: NO. Countries: 4. Publications: 1.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

Phase 2 Study of AUY922 in NSCLC Patients With Exon 20 Insertion Mutations in EGFR

ClinicalTrials.gov study NCT01854034. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

Osimertinib for NSCLC With EGFR Exon 20 Insertion Mutation

ClinicalTrials.gov study NCT03414814. IPD Sharing: Not stated. Countries: 1. Publications: 1.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad36/100

Novel DMD mouse model carrying a multi-exonic Dmd deletion exhibit progressive muscular dystrophy and early-onset cardiomyopathy

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad36/100

Data from: Natural variation in the zinc-finger-encoding exon of Prdm9 affects hybrid sterility phenotypes in mice

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad36/100

Assembled exon data for Gobioid phylogenetic study

Open the record for dataset details and reuse information.

publicJul 2025View details →
zenodo32/100

Training data for "Differential exon usage analysis"

<p>In RNA-Seq, we usually want to know the differentially expressed genes, as explained in several Galaxy Training Material tutorials. Sometimes,the question is more &quot;which exons are differentially expressed&quot;. The process to identify differentially expressed exons is really similar to the one for differentially expressed genes.</p> <p>In this tutorial, we identify exons regulated by the <em>Pasilla</em> gene using RNA-Seq data from <a href="https://training.galaxyproject.org/training-material/topics/transcriptomics/tutorials/ref-based/tutorial.html#brooks2011conservation">Brooks <em>et al.</em> 2011</a>.</p>

opencc-by-4.0Dec 2019View details →
dryad32/100

Data from: Evolution of the MHC-DQB exon 2 in marine and terrestrial mammals

On the basis of a general low polymorphism, several studies suggest that balancing selection in the class II major histocompatibility complex (MHC) is weaker in marine mammals as compared with terrestrial mammals. We investigated such differential selection among Cetacea, Artiodactyla, and Primates at exon 2 of MHC-DQB gene by contrasting indicators of molecular evolution such as occurrence of transpecific polymorphisms, patterns of phylogenetic branch lengths by codon position, rates of nonsynonymous and synonymous substitutions as well as accumulation of variable sites on the sampling of alleles. These indicators were compared between the DQB and the mitochondrial cytochrome b gene (cytb) as a reference of neutral expectations and differences between molecular clocks resulting from life history and historical demography. All indicators showed that the influence of balancing selection on the DQB is more variable and overall weaker for cetaceans. In our sampling, ziphiids, the sperm whale, monodontids and the finless porpoise formed a group with lower DQB polymorphism, while mysticetes exhibited a higher DQB variation similar to that of terrestrial mammals as well as higher occurrence of transpecific polymorphisms. Different dolphins appeared in the two groups. Larger variation of selection on the cetacean DQB could be related to greater stochasticity in their historical demography and thus, to a greater complexity of the general ecology and disease processes of these animals.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Resolving the phylogenetic position of Darwin's extinct ground sloth (Mylodon darwinii) using mitogenomic and nuclear exon data

Mylodon darwinii is the extinct giant ground sloth named after Charles Darwin, who first discovered its remains in South America. We have successfully obtained a high-quality mitochondrial genome at 99-fold coverage using an Illumina shotgun sequencing of a 12,880 year-old bone fragment from Mylodon Cave in Chile. Low level of DNA damage showed that this sample was exceptionally well preserved for an ancient sub-fossil, likely the result of the dry and cold conditions prevailing within the cave. Accordingly, taxonomic assessment of our shotgun metagenomic data showed a very high percentage of endogenous DNA with 22% of the assembled metagenomic contigs assigned to Xenarthra. Additionally, we enriched over 15 kilobases of sequence data from seven nuclear exons, using target sequence capture designed against a wide xenarthran dataset. Phylogenetic and dating analyses of the mitogenomic dataset including all extant species of xenarthrans and the assembled nuclear supermatrix unambiguously place Mylodon darwinii as the sister-group of modern two-fingered sloths from which it diverged around 22 million years ago. These congruent results from both the mitochondrial and nuclear data support the diphyly of the two modern sloths lineages, implying the convergent evolution of their unique suspensory behaviour as an adaption to arboreality. Our results offer promising perspectives for whole genome sequencing of this emblematic extinct taxon.

opencc-zeroDec 2017View details →
dryad32/100

Development of novel, Exon-Primed Intron-Crossing (EPIC) markers from EST databases and evaluation of their phylogenetic utility in Commiphora (Burseraceae)

Premise of the study: Novel nuclear exon-primed intron-crossing (EPIC) markers were developed to increase phylogenetic resolution among recently diverged lineages in the frankincense and myrrh family, Burseraceae, using Citrus, Arabidopsis, and Oryza genome resources. Methods and Results: Primer pairs for 48 nuclear introns were developed using the genome resource IntrEST and were screened using species of Commiphora and other Burseraceae taxa. Four putative intron regions (RPT6A, BXL2, mtATP Synthase D, and Rab6) sequenced successfully for multiple taxa and recovered phylogenies consistent with those of existing studies. In some cases, these regions yielded informative sequence variation on par with that of the nrDNA internal transcribed spacer. Conclusions: The combination of freely available genome resources and our design criteria have uncovered four, single-copy nuclear intron regions that are useful for phylogenetic reconstruction of Burseraceae taxa. Because our EPIC primers also amplify Arabidopsis, we recommend their trial in other rosid and eudicot lineages.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Conserved non-exonic elements: a novel class of marker for phylogenomics

Noncoding markers have a particular appeal as tools for phylogenomic analysis because, at least in vertebrates, they appear less subject to strong variation in GC content among lineages. Thus far, ultraconserved elements (UCEs) and introns have been the most widely used noncoding markers. Here we analyze and study the evolutionary properties of a new type of noncoding marker, conserved non-exonic elements (CNEEs), which consists of noncoding elements that are estimated to evolve slower than the neutral rate across a set of species. Although they often include UCEs, CNEEs are distinct from UCEs because they are not ultraconserved, and, most importantly, the core region alone is analyzed, rather than both the core and its flanking regions. Using a data set of 16 birds plus an alligator outgroup, and ∼3600 - ∼3800 loci per marker type, we found that although CNEEs were less variable than bioinformatically-derived UCEs or introns and in some cases exhibited a slower approach to branch resolution as determined by phylogenomic subsampling, the quality of CNEE alignments was superior to those of the other markers, with fewer gaps and missing species. Phylogenetic resolution using coalescent approaches was comparable among the three marker types, with most nodes being fully and congruently resolved. Comparison of phylogenetic results across the three marker types indicated that one branch, the sister group to the passerine+falcon clade, was resolved differently and with moderate (&gt; 70%) bootstrap support between CNEEs and UCEs or introns. Overall, CNEEs appear to be promising as phylogenomic markers, yielding phylogenetic resolution as high as for UCEs and introns but with fewer gaps, less ambiguity in alignments and with patterns of nucleotide substitution more consistent with the assumptions of commonly used methods of phylogenetic analysis.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record