Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
665
datasets available to search
ShareScore release 0.9.0
Dataset results
665 results for “exon”
Intronization signatures in coding exons reveal the evolutionary fluidity of eukaryotic gene architecture
<p>This dataset accompanies the manuscript "Intronization signatures in coding exons reveal the evolutionary fluidity of eukaryotic gene architecture". It contains the relevant tables, scripts and figures used for and created during data analysis.</p>
Huntingtin Exon 1 Q16 and Q46 pET32a thioredoxin tagged expression and purification 2018/04/09
<p>Huntingtin structure-function open lab notebook project. Huntingtin Exon 1 Q16 and Q46 pET32a thioredoxin tagged expression and purification 2018/04/09.</p>
Huntingtin Exon 1 Q16 and Q46 pET32a thioredoxin tagged expression and purification 2018/04/02
<p>Huntingtin structure-function open lab notebook project. Huntingtin Exon 1 Q16 and Q46 pET32a thioredoxin tagged, attempted expression and purification 2018/04/02</p>
Electrophysiology Data for "Two functional epithelial sodium channel isoforms are present in rodents despite pronounced evolutionary pseudogenization and exon fusion"
<p>Here we provide the electrophysiology data for the manuscript "Two functional epithelial sodium channel isoforms are present in rodents despite pronounced evolutionary pseudogenization and exon fusion", published in Molecular Biology and Evolution (2021): msab271 (doi: 10.1093/molbev/msab271). Data are reported as current values in Excel format, sorted according to the appearance in Figures and supplemented by explanatory text on the procedures/data presentation.</p>
The impact of genetically controlled splicing on exon inclusion and protein structure
<p><strong>This repository contains raw and processed files used in Einson et. al 2022. </strong></p> <p>Code used to generate these files can be found here: https://github.com/jeinson/sqtl_manuscript</p> <p><strong><em>Descriptions of files contained within each sub directory</em></strong></p> <p><strong>01_raw_psi</strong></p> <ul> <li><em>{GTEx_tissue_id}_v8.psi.tsv.gz: </em>Unfiltered PSI output from IPSA-nf, per tissue. See methods for details about how files were created. </li> <li><em>gtex_v8_exon_id_map.tsv: </em>Mapping file between exon coordinates and Ensembl gene IDs, with suffix used in GTEx v8 gencode annotation. </li> </ul> <p><strong>02_qtl_results</strong></p> <ul> <li><strong>cross_tissue</strong> <ul> <li><em>top_sQTLs_MAF05.tsv: </em>List of top GTEx v8 sQTLs across tissues, with one exon and top variant per tissue. See methods for details. See matching file for column descriptions. </li> <li><em>top_sQTLs_median_psi.tsv: </em>The median, mean, and standard deviation of PSI of each significant exon from the previous file, taken across all individuals from GTEx with data available.</li> <li><em>top_sQTLs_MAF05_w_anc_allele.tsv: </em>List of top sQTLs across tissues, with additional columns for the top ψQTL ancestral and derived alleles, where available. </li> </ul> </li> <li><strong>per_tissue</strong> <ul> <li><em>{GTEx_tissue_id}_combined_sQTLs.tsv.gz: </em>Raw output of ψQTL calling using QTLtools in grouped permutational mode per tissue, with groups specified by gene. See methods for more details, and https://qtltools.github.io/qtltools/ for column descriptions. </li> </ul> </li> </ul> <p><strong>03_qtl_credible_sets</strong></p> <ul> <li> <em>GTEx_psi_{GTEx_tissue_id}.collapsed.txt.gz: </em>Output of the QTL catalog fine mapping pipeline (https://github.com/eQTL-Catalogue/qtlmap), run on all exons and tissues, and collapsed using the procedure described in Methods. </li> </ul> <p><strong>04_qtl_coloc</strong></p> <ul> <li><em>combined_coloc_results_full.tsv.gz: </em>Combined output of running coloc on ψQTLs from the 18 GTEx tissues against 87 sets of GWAS summary statistics. This file contains all results, including non-significant associations. A nominal QTLtools pass was used as input. We do not include these files in this repository due to size limitations, but contact the authors if you need access to nominal QTL calls. </li> <li><em>top_sQTLs_with_top_coloc_event.tsv: </em>The QTLs in <em>top_sQTLs_MAF05.tsv</em> with additional columns for the GWAS with the highest posterior probability of a colocalization event. Importantly, the tissue and top variant may not match the main <em>top_sQTLs_MAF05.tsv </em>file for every gene. </li> </ul> <p><strong>05_exon_features: </strong>See matching files for description of each column. </p> <ul> <li><em>cross_tissue_constitutive_exons_with_AF.tsv: </em>Detailed features of cross tissue constitutive exons. See methods for definition of constitutive exons. </li> <li><em>cross_tissue_nonsignificant_genes_with_AF.tsv: </em>Detailed features of sufficiently variable exons with no significant variant across tissues. See methods for more details. </li> <li><em>top_sQTLs_MAF05_with_AF.tsv: </em>Detailed features of top sQTLs. </li> <li><em>top_sQTLs_with_top_coloc_with_AF.tsv: </em>Detailed features of sQTLs that colocalize with at least one GWAS trait. Contains columns for Euclidean distances between PAE matrices and RMSD between isoforms, among genes with a significant GWAS colocalization event. </li> </ul> <p><strong>06_predicted_structures: </strong>Each prediction was run 5 times, and we report the best model in the manuscript. </p> <ul> <li><strong>{protein.id}[_mutant].result</strong> <ul> <li><em>{protein.id}[_mutant]{_run.id}_coverage.png.gz: </em>Plot of the number of sequences per position in MSA</li> <li><em>{protein.id}[_mutant]{_run.id}_PAE.png.gz: </em>PAE matrix plots for each model</li> <li><em>{protein.id}[_mutant]{_run.id}_plddt.png.gz: </em>pLDDT plots for each model</li> <li><em>{protein.id}[_mutant]{_run.id}_predicted_aligned_error_v1.json.gz</em>: A PAE matrix for the best model using <a href="https://alphafold.ebi.ac.uk/faq#faq-7">AlphaFold-DB's format</a></li> <li><em>{protein.id}[_mutant]{_run.id}_unrelaxed_rank_{rank.num}_model_{model.num}_scores.json.gz</em>: Per model array (list of lists) with PAE, a list of the average pLDDT and the pTM score. </li> <li><em>{protein.id}[_mutant]{_run.id}_unrelaxed_rank_{rank.num}_model_{model.num}_pdb.gz: </em>Per model predicted structure in pd format</li> <li><em>{protein.id}[_mutant]{_run.id}.a3m.gz</em>: A3M formatted input MSA</li> <li><em>cite.bibtex: </em>BibTex file with citations for all used tools and databases</li> <li><em>config.json</em>: Model input parameters</li> </ul> </li> </ul> <p><strong>07_other_data</strong></p> <ul> <li><em>cross_tissue_constitutive_exons.tsv: </em>List of exons that are constitutively spliced across multiple tissues. See methods for details. </li> <li><em>cross_tissue_nonsignificant_genes.tsv</em>: List variably spliced exons with no significant sVariant in any tissue. See methods for details. </li> <li><em>gtex_v8_exon_id_map.rds: </em>rds representation of a map between exon IDs, as used in the modified version of gencode v26, and exon hg38 coordinates. </li> <li><em>gtex_v8_n_exons_per_gene.tsv: </em>Number of exons per gene, as annotated in the modified version of gencode v26 used in GTEx v8. </li> </ul> <p><strong>08_geuvadis</strong></p> <ul> <li><em>geuvadis_psi.tsv.gz: </em>Unfiltered PSI output from IPSA-nf, run on Geuvadis BAM files. See methods for details. (Raw data was downloaded from ftp://ftp.ebi.ac.uk/pub/databases/microarray/data/experiment/GEUV)</li> <li><em>geuvadis_sQTLs.tsv.gz: </em>Raw output of ψQTL calling using QTLtools in grouped permutational mode for geuvadis data, with groups specified by gene. See methods for more details, and https://qtltools.github.io/qtltools/ for column descriptions. </li> <li><em>remapped_gencode.v26.GRCh37.GTEx_v8.nochr.genes.gtf.gz:</em> Lifted over version of the gencode v26 gtf file, used to define exons for PSI and qtl mapping in the geuvadis analysis. The original version that was used in the GTEx analysis is based on GRCh38, and is available here: https://storage.googleapis.com/gtex_analysis_v8/reference/gencode.v26.GRCh38.genes.gtf</li> </ul>
Sequences of target genes and probes for construction of exon-targeted libraries for Prunus spp.
<p>(Files)</p> <p>"baits-selection-15171.fas": Sequences of 15,171 baits used for hybridization reaction.</p> <p>"Targets_mume.15171.xlsx": Sequences of target genes used for design baits.</p> <p> </p> <p>(Protocol for bait design)</p> <p>To selectively retrieve libraries with exons, a myBaits Custom design kit was used to design 1–20-K probes (Arbor Biosciences, Ann Arbor, MI, USA), which uses biotinylated RNA probes to concentrate fragments carrying sequences of interest (Gnirke et al., 2009), based on the published genomic and coding sequences of <em>P. mume</em> (Zhang et al., 2012). We subjected 29,621 non-redundant coding loci to search for single hits with BLAST+ (MEGABLAST with -p 70) against the <em>P. mume</em> genome, for the subsequent bait design. A 120-mer bait with 25–55 GC% per locus was randomly designed for each locus, and finally we obtained a bait set targeting 15,171 coding loci.</p>
Phase 2 Study of Poziotinib in Participants With NSCLC Having EGFR or HER2 Exon 20 Insertion Mutation
ClinicalTrials.gov study NCT03318939. IPD Sharing: NO. Countries: 8. Publications: 1.
AAV9 U7snRNA Gene Therapy to Treat Boys With DMD Exon 2 Duplications.
ClinicalTrials.gov study NCT04240314. IPD Sharing: Not stated. Countries: 1. Publications: 5.
TAK-788 as First-Line Treatment Versus Platinum-Based Chemotherapy for Non-Small Cell Lung Cancer (NSCLC) With EGFR Exon 20 Insertion Mutations
ClinicalTrials.gov study NCT04129502. IPD Sharing: YES. Countries: 24. Publications: 1.
Study of Eteplirsen in Young Participants With Duchenne Muscular Dystrophy (DMD) Amenable to Exon 51 Skipping
ClinicalTrials.gov study NCT03218995. IPD Sharing: NO. Countries: 4. Publications: 1.
Phase 2 Study of AUY922 in NSCLC Patients With Exon 20 Insertion Mutations in EGFR
ClinicalTrials.gov study NCT01854034. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Osimertinib for NSCLC With EGFR Exon 20 Insertion Mutation
ClinicalTrials.gov study NCT03414814. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Novel DMD mouse model carrying a multi-exonic Dmd deletion exhibit progressive muscular dystrophy and early-onset cardiomyopathy
Open the record for dataset details and reuse information.
Data from: Natural variation in the zinc-finger-encoding exon of Prdm9 affects hybrid sterility phenotypes in mice
Open the record for dataset details and reuse information.
Assembled exon data for Gobioid phylogenetic study
Open the record for dataset details and reuse information.
Training data for "Differential exon usage analysis"
<p>In RNA-Seq, we usually want to know the differentially expressed genes, as explained in several Galaxy Training Material tutorials. Sometimes,the question is more "which exons are differentially expressed". The process to identify differentially expressed exons is really similar to the one for differentially expressed genes.</p> <p>In this tutorial, we identify exons regulated by the <em>Pasilla</em> gene using RNA-Seq data from <a href="https://training.galaxyproject.org/training-material/topics/transcriptomics/tutorials/ref-based/tutorial.html#brooks2011conservation">Brooks <em>et al.</em> 2011</a>.</p>
Data from: Evolution of the MHC-DQB exon 2 in marine and terrestrial mammals
On the basis of a general low polymorphism, several studies suggest that balancing selection in the class II major histocompatibility complex (MHC) is weaker in marine mammals as compared with terrestrial mammals. We investigated such differential selection among Cetacea, Artiodactyla, and Primates at exon 2 of MHC-DQB gene by contrasting indicators of molecular evolution such as occurrence of transpecific polymorphisms, patterns of phylogenetic branch lengths by codon position, rates of nonsynonymous and synonymous substitutions as well as accumulation of variable sites on the sampling of alleles. These indicators were compared between the DQB and the mitochondrial cytochrome b gene (cytb) as a reference of neutral expectations and differences between molecular clocks resulting from life history and historical demography. All indicators showed that the influence of balancing selection on the DQB is more variable and overall weaker for cetaceans. In our sampling, ziphiids, the sperm whale, monodontids and the finless porpoise formed a group with lower DQB polymorphism, while mysticetes exhibited a higher DQB variation similar to that of terrestrial mammals as well as higher occurrence of transpecific polymorphisms. Different dolphins appeared in the two groups. Larger variation of selection on the cetacean DQB could be related to greater stochasticity in their historical demography and thus, to a greater complexity of the general ecology and disease processes of these animals.
Data from: Resolving the phylogenetic position of Darwin's extinct ground sloth (Mylodon darwinii) using mitogenomic and nuclear exon data
Mylodon darwinii is the extinct giant ground sloth named after Charles Darwin, who first discovered its remains in South America. We have successfully obtained a high-quality mitochondrial genome at 99-fold coverage using an Illumina shotgun sequencing of a 12,880 year-old bone fragment from Mylodon Cave in Chile. Low level of DNA damage showed that this sample was exceptionally well preserved for an ancient sub-fossil, likely the result of the dry and cold conditions prevailing within the cave. Accordingly, taxonomic assessment of our shotgun metagenomic data showed a very high percentage of endogenous DNA with 22% of the assembled metagenomic contigs assigned to Xenarthra. Additionally, we enriched over 15 kilobases of sequence data from seven nuclear exons, using target sequence capture designed against a wide xenarthran dataset. Phylogenetic and dating analyses of the mitogenomic dataset including all extant species of xenarthrans and the assembled nuclear supermatrix unambiguously place Mylodon darwinii as the sister-group of modern two-fingered sloths from which it diverged around 22 million years ago. These congruent results from both the mitochondrial and nuclear data support the diphyly of the two modern sloths lineages, implying the convergent evolution of their unique suspensory behaviour as an adaption to arboreality. Our results offer promising perspectives for whole genome sequencing of this emblematic extinct taxon.
Development of novel, Exon-Primed Intron-Crossing (EPIC) markers from EST databases and evaluation of their phylogenetic utility in Commiphora (Burseraceae)
Premise of the study: Novel nuclear exon-primed intron-crossing (EPIC) markers were developed to increase phylogenetic resolution among recently diverged lineages in the frankincense and myrrh family, Burseraceae, using Citrus, Arabidopsis, and Oryza genome resources. Methods and Results: Primer pairs for 48 nuclear introns were developed using the genome resource IntrEST and were screened using species of Commiphora and other Burseraceae taxa. Four putative intron regions (RPT6A, BXL2, mtATP Synthase D, and Rab6) sequenced successfully for multiple taxa and recovered phylogenies consistent with those of existing studies. In some cases, these regions yielded informative sequence variation on par with that of the nrDNA internal transcribed spacer. Conclusions: The combination of freely available genome resources and our design criteria have uncovered four, single-copy nuclear intron regions that are useful for phylogenetic reconstruction of Burseraceae taxa. Because our EPIC primers also amplify Arabidopsis, we recommend their trial in other rosid and eudicot lineages.
Data from: Conserved non-exonic elements: a novel class of marker for phylogenomics
Noncoding markers have a particular appeal as tools for phylogenomic analysis because, at least in vertebrates, they appear less subject to strong variation in GC content among lineages. Thus far, ultraconserved elements (UCEs) and introns have been the most widely used noncoding markers. Here we analyze and study the evolutionary properties of a new type of noncoding marker, conserved non-exonic elements (CNEEs), which consists of noncoding elements that are estimated to evolve slower than the neutral rate across a set of species. Although they often include UCEs, CNEEs are distinct from UCEs because they are not ultraconserved, and, most importantly, the core region alone is analyzed, rather than both the core and its flanking regions. Using a data set of 16 birds plus an alligator outgroup, and ∼3600 - ∼3800 loci per marker type, we found that although CNEEs were less variable than bioinformatically-derived UCEs or introns and in some cases exhibited a slower approach to branch resolution as determined by phylogenomic subsampling, the quality of CNEE alignments was superior to those of the other markers, with fewer gaps and missing species. Phylogenetic resolution using coalescent approaches was comparable among the three marker types, with most nodes being fully and congruently resolved. Comparison of phylogenetic results across the three marker types indicated that one branch, the sister group to the passerine+falcon clade, was resolved differently and with moderate (> 70%) bootstrap support between CNEEs and UCEs or introns. Overall, CNEEs appear to be promising as phylogenomic markers, yielding phylogenetic resolution as high as for UCEs and introns but with fewer gaps, less ambiguity in alignments and with patterns of nucleotide substitution more consistent with the assumptions of commonly used methods of phylogenetic analysis.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.