Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
144
datasets available to search
ShareScore release 0.9.0
Dataset results
144 results for “splice variant”
Nanopore deep sequencing as a tool to characterize and quantify aberrant splicing caused by variants in inherited retinal dystrophy genes
Open the record for dataset details and reuse information.
Identification and characterization of novel splice variants of human farnesoid X receptor
<p>Dataset related to publication:</p> <blockquote> <p>Mustonen E-K, Lee SML, Nieß H, Schwab M, Pantsar T, Burk O: "Identification and characterization of novel splice variants of human farnesoid X receptor". <em>Archives of Biochemistry and Biophysics</em> <a href="https://doi.org/10.1016/j.abb.2021.108893">https://doi.org/10.1016/j.abb.2021.108893</a></p> </blockquote> <p> </p> <p>Including:</p> <p>I. Full-length raw-trajectories of the 1 microsecond Desmond simulations (-out.cms, trj)</p>
Oncogenic role of a developmentally regulated NTRK2 splice variant
<p>Abstract</p> <p>Temporally-regulated alternative splicing choices are vital for proper development yet the wrong splice choice may be detrimental. Here we highlight a novel role for the neurotrophin receptor splice variant TrkB.T1 in neurodevelopment, embryogenesis, transformation, and oncogenesis across multiple tumor types in both humans and mice. TrkB.T1 is the predominant NTRK2 isoform across embryonic organogenesis and forced over-expression of this embryonic pattern causes multiple solid and nonsolid tumors in mice in the context of tumor suppressor loss. TrkB.T1 also emerges the predominant NTRK isoform expressed in a wide range of adult and pediatric tumors, including those harboring TRK fusions. Affinity purification-mass spectrometry (AP-MS) proteomic analysis reveals TrkB.T1 has distinct interactors with known developmental and oncogenic signaling pathways such as Wnt, TGF-ß, Hedgehog, and Ras. From alterations in splicing factors to changes in gene expression, the discovery of isoform specific oncogenes with embryonic ancestry has the potential to shape the way we think about developmental systems and oncology.</p>
Data for Cell-type-specific alternative splicing in the cerebral cortex of a Schinzel-Giedion Syndrome patient variant mouse model
<p><span><strong>data.tar.gz </strong>contains all files from the data directory (except for sam outputs from STAR) associated with the 230926_EJ_Setbp1_AlternativeSplicing GitHub project and includes the following files:</span></p> <p> </p> <p><span><strong>./marvel: </strong>- </span><span>This directory contains rds and Rdata objects that were created using the MARVEL R package</span></p> <p><span>cell_type_goresults.rds - This is the go results split by cell type</span></p> <p><span>marvel_04_split_counts.Rdata - This R data includes all environment objects from MARVEL script 04, and is used for downstream plotting</span></p> <p><span>normalized_sj_expression.Rds - This object is the normalized splice junction expression</span></p> <p><span>Setbp1_marvel_aligned.rds - Final prepared MARVEL object before any SJU analyses have been run</span></p> <p><span>significant_tables.RData - For those who do not want to load multiple massive files, this includes all significant SJU results for each cell type</span></p> <p><span>sj_usage_cell_type.rds - This data object has splice junction usage calculated for each cell type</span></p> <p><span>sj_usage_condition.rds - This data object has splice junction usage calculated for each cell type and also split by condition</span></p> <p> </p> <p><strong><span>./seurat: </span></strong><span>- This directory contains all intermediate and final Seurat single-cell gene expression objects</span></p> <p><span>annotated_brain_samples.rds - This is the final iteration of the processing in Seurat for a final annotated object. Please use this object for any Seurat or single-cell gene expression analyses.</span></p> <p><span>clustered_brain_samples.rds - This is the clustered Seurat object, before cell type annotation based on canonical markers.</span></p> <p><span>filtered_brain_samples_pca.rds - This is the filtered Seurat object, before clustering but after PCA.</span></p> <p><span>filtered_brain_samples.rds - This is the filtered Seurat object, before PCA.</span></p> <p><span>integrated_brain_samples.rds - This the integrated Seurat object, before other steps.</span></p> <p> </p> <p><span><strong>./star: </strong>- </span><span>All files in the STAR directory are outputs from STARsolo, as described in our methods. Each output directory contains the same files, so only one example is included here for brevity. Intermediate SAM files were removed to optimize space.</span></p> <p><span>J1/ - This directory contains outputs for brain sample J1</span></p> <p><span>J13/ - This directory contains outputs for brain sample J13</span></p> <p><span>J15/ - This directory contains outputs for brain sample J15</span></p> <p><span>J2/ - This directory contains outputs for brain sample J2</span></p> <p><span>J3/ - This directory contains outputs for brain sample J3</span></p> <p><span>J4/ - This directory contains outputs for brain sample J4</span></p> <p><span>K1/ - This directory contains outputs for kidney sample K1</span></p> <p><span>K2/ - This directory contains outputs for kidney sample K2</span></p> <p><span>K3/ - This directory contains outputs for kidney sample K3</span></p> <p><span>K4/ - This directory contains outputs for kidney sample K4</span></p> <p><span>K5/ - This directory contains outputs for kidney sample K5</span></p> <p><span>K6/ - This directory contains outputs for kidney sample K6</span></p> <p> </p> <p><span><strong>./star/genome:</strong> - This directory contains outputs from running STAR genomeGenerate. Detailed file descriptions available from</span><a href="https://github.com/alexdobin/STAR/blob/master/doc/STARmanual.pdf"><span> </span><span>https://github.com/alexdobin/STAR/blob/master/doc/STARmanual.pdf</span></a><span> </span></p> <p><span>chrLength.txt</span></p> <p><span>chrNameLength.txt</span></p> <p><span>chrName.txt</span></p> <p><span>chrStart.txt</span></p> <p><span>exonGeTrInfo.tab</span></p> <p><span>exonInfo.tab</span></p> <p><span>geneInfo.tab</span></p> <p><span>Genome</span></p> <p><span>genomeParameters.txt</span></p> <p><span>Log.out</span></p> <p><span>SA</span></p> <p><span>SAindex</span></p> <p><span>sjdbInfo.txt</span></p> <p><span>sjdbList.fromGTF.out.tab</span></p> <p><span>sjdbList.out.tab</span></p> <p><span>transcriptInfo.tab</span></p> <p> </p> <p><span><strong>./star/J1:</strong> - This is the head STAR directory for sample J1. It contains logs, basic QC, and gene and splice junction counts. For more information about the STAR pipeline and its outputs, please refer to the STAR documentation</span><a href="https://github.com/alexdobin/STAR/blob/master/doc/STARmanual.pdf"><span> </span><span>https://github.com/alexdobin/STAR/blob/master/doc/STARmanual.pdf</span></a><span> </span></p> <p><span>Log.final.out</span></p> <p><span>Log.out</span></p> <p><span>Log.progress.out</span></p> <p><span>SJ.out.tab</span></p> <p><span>Solo.out/</span></p> <p><span>STARgenome/</span></p> <p> </p> <p><span><strong>./star/J1/Solo.out:</strong>- This directory contains the outputs used for downstream analysis</span></p> <p><span>Barcodes.stats</span></p> <p><span>GeneFull_Ex50pAS/</span></p> <p><span>SJ/</span></p> <p> </p> <p><span><strong>./star/J1/Solo.out/GeneFull_Ex50pAS: </strong>- This directory contains the filtered and raw barcodes, features, and matrix files for gene expression (including introns)</span></p> <p><span>Features.stats</span></p> <p><span>filtered/</span></p> <p><span>raw/</span></p> <p><span>Summary.csv</span></p> <p><span>UMIperCellSorted.txt</span></p> <p> </p> <p><span><strong>./star/J1/Solo.out/GeneFull_Ex50pAS/filtered: </strong>- This directory contains the filtered tsv and mtx gene expression files required for creating a Seurat object (or other single cell packages)</span></p> <p><span>barcodes.tsv.gz - This file contains filtered cell barcodes</span></p> <p><span>features.tsv.gz - This file contains filtered features (genes)</span></p> <p><span>matrix.mtx.gz - This file contains the filtered cell by gene expression count matrix</span></p> <p> </p> <p><span><strong>./star/J1/Solo.out/GeneFull_Ex50pAS/raw: </strong>- This directory contains the unfiltered tsv and mtx gene expression files required for creating a Seurat object (or other single cell packages). Files are the same as previously described for filtered.</span></p> <p><span>barcodes.tsv</span></p> <p><span>features.tsv</span></p> <p><span>matrix.mtx</span></p> <p> </p> <p><span><strong>./star/J1/Solo.out/SJ: </strong>- This directory contains the QC and raw barcodes, features, and matrix files for splice junction expression</span></p> <p><span>Features.stats</span></p> <p><span>raw/</span></p> <p><span>Summary.csv</span></p> <p> </p> <p><span><strong>./star/J1/Solo.out/SJ/raw:</strong> - This directory contains the raw barcodes, features, and matrix files for splice junction expression</span></p> <p><span>barcodes.tsv - This file contains filtered cell barcodes</span></p> <p><span>features.tsv - This file contains filtered features (splice junctions)</span></p> <p><span>matrix.mtx - This file contains the filtered cell by gene expression count matrix</span></p> <p> </p> <p><span><strong>./star/J1/_STARgenome:</strong> - This directory contains the STARgenome created and used by STAR for this sample. Detailed file descriptions available from</span><a href="https://github.com/alexdobin/STAR/blob/master/doc/STARmanual.pdf"><span> </span><span>https://github.com/alexdobin/STAR/blob/master/doc/STARmanual.pdf</span></a><span> </span></p> <p><span>exonGeTrInfo.tab</span></p> <p><span>exonInfo.tab</span></p> <p><span>geneInfo.tab</span></p> <p><span>sjdbInfo.txt</span></p> <p><span>sjdbList.fromGTF.out.tab</span></p> <p><span>sjdbList.out.tab</span></p> <p><span>transcriptInfo.tab</span></p>
Massively parallel screens to identify splice disruptive variants in human disease genes
<p>Splicing is a critical step in mRNA maturation with roles in gene regulation and proteome diversification. Splice disruptive variants (SDVs) are implicated in diverse human diseases, and 10-33% of exonic variants may disrupt splicing. However, identifying SDVs remains challenging due to the degeneracy and redundancy of the underlying sequence code. Experimental splicing measurements from patient cells or mini-gene assays can detect SDVs but have traditionally been low-throughput. </p> <p>Massively parallel reporter assays (MPSAs) systematically measure splicing impacts at scale and could clarify variant pathogenicity and inform models of splicing regulation. In this assay, complex barcoded libraries of mutant exons are synthesized, cloned into minigene constructs, and transfected into human cells. Splicing outcomes of each mutation are quantified en masse using targeted RNA-seq of minigene-derived transcripts, and I developed custom python package to process the resulting data.</p> <p>In Chapter 2, I apply this assay to the pituitary transcription factor gene <em>POU1F1</em> (in collaboration Dr. Sally Camper’s lab). Mutations in <em>POU1F1</em> cause combined pituitary hormone deficiency (CPHD), a clinically and genetically heterogenous disorder with prevalence ~1:4000. We targeted exon 2, which has two alterative isoforms (alpha and beta) using competing splice acceptors that encode mutually antagonistic proteins. We measured the splicing effects of 1,070 SNVs across the exon and surrounding introns and identified 96 SDVs - 14 of which were synonymous substitutions. Our measurements were concordant with six nearby heterozygous missense and synonymous variants seen in unrelated hypopituitarism patients. This map identifies a putative splice silencer motif that represses the use of the normally lowly expressed beta isoform.</p> <p>In Chapter 3, I apply a MPSA to a critical developmental renal transcription factor gene, <em>WT1</em> (in collaboration with clinical nephrologist Dr. Jen Lai Yee). Mutations in <em>WT1 </em>are implicated in nephrotic syndrome and sexual differentiation phenotypes. I focus on exon 9 which is alternatively spliced at competing donor sites resulting in two isoforms (KTS+ and KTS-). KTS+ and KTS- are normally expressed in ~2:1 ratio, but perturbation of the ratio can lead to Frasier’s syndrome – a rare nephrotic syndrome. We tested 518 SNVs for splicing defects and identified 8 known Frasier’s Syndrome variants as well as 16 additional variants that similarly lowered the KTS ratio. We also detected 19 variants increasing the KTS ratio - two of which have been observed in patients with sexual differentiation phenotypes.</p> <p>Although MPSAs can measure splicing effects of hundreds of variants simultaneously, the current scale of variant discovery via exome and genome sequencing demands efficient and accurate computational approaches to identify splice disruptive variants genome-wide. To evaluate the state of the art within contemporary splice prediction algorithms, in Chapter 4 I employed the results of five high throughput splicing assays and one literature curated variant set. A unique advantage of MPSAs over typical training and validation datasets is that they avoid bias towards essential splice site variants. I found the latest deep learning tools, SpliceAI and Pangolin, were most concordant with the measured splicing effects. However, all tools showed less agreement with exonic splicing outcomes compared to intronic. Some tools’ predictions, like SpliceAI’s, were sensitive to specified annotation files. Thus, there is still room for improvement within the next generation of splice prediction algorithms which future MPSA studies may facilitate.</p>
Splice altering variant predictions in four archaic hominin genomes
Open the record for dataset details and reuse information.
Screening for functional transcriptional and splicing regulatory variants with GenIE
<p>Here we provide data used in the analysis for our paper introducing the GenIE method (Genome engineering-based Interrogation of Enhancers).</p> <p><strong>Abstract</strong></p> <p>Genome-wide association studies (GWAS) have identified numerous genetic loci underlying human diseases, but a fundamental challenge remains to accurately identify the underlying causal genes and variants. Here we describe an arrayed CRISPR screening method, Genome engineering-based Interrogation of Enhancers (GenIE), which assesses the effects of defined alleles on transcription or splicing when introduced in their endogenous genomic location. We use this sensitive assay to validate the activity of transcriptional enhancers and splice regulatory elements in human induced pluripotent stem cells (hiPSCs), and develop a software package (rgenie) to analyse the data. We screen the 99% credible set of Alzheimer’s disease (AD) GWAS variants identified at the clusterin (CLU) locus to identify a subset of likely causal variants, and employ GenIE to understand the impact of specific mutations on splicing efficiency. We thus establish GenIE as an efficient tool to rapidly screen for the role of transcribed variants on gene expression.</p>
Benchmarking splice variant prediction algorithms using massively parallel splicing assays
<p>Dataset, jupyter notebooks, and support python modules for "Benchmarking splice variant prediction algorithms using massively parallel splicing assays" (Smith and Kitzman, 2023)</p>
Calcium Channel Splice Variant Expression in Cardiovascular Disease and Aging
ClinicalTrials.gov study NCT00251615. IPD Sharing: Not stated. Countries: 1. Publications: 3.
A Study of Galeterone Compared to Enzalutamide In Men Expressing Androgen Receptor Splice Variant-7 mRNA (AR-V7) Metastatic CRPC
ClinicalTrials.gov study NCT02438007. IPD Sharing: Not stated. Countries: 8. Publications: 1.
Isoform specific activities of androgen receptor and its splice variants in prostate cancer cells
<p>Androgen receptor (AR) signaling continues to drive castration resistant prostate cancer (CRPC) in spite of androgen deprivation therapy (ADT). Constitutively active shorter variants of AR, lacking the ligand binding domain, are frequently expressed in CRPC and have emerged as a potential mechanism for prostate cancer to escape ADT. ARv7 and AR<sup>v567es </sup>are two of the most commonly detected variants of AR in clinical samples of advanced, metastatic prostate cancer. It is not clear if variants of AR merely act as weaker substitutes for AR or can mediate unique isoform specific activities different from AR. In this study, we employed LNCaP prostate cancer cell lines with inducible expression of ARv7 or AR<sup>v567es </sup>to delineate similarities and differences in transcriptomics, metabolomics and lipidomics resulting from the activation of AR, ARv7 or AR<sup>v567es</sup>. While the majority of target genes were similarly regulated by the action of all three isoforms, we found a clear difference in transcriptomic activities of AR versus the variants, and a few differences between ARv7 and AR<sup>v567es</sup>. Some of the target gene regulation by AR isoforms was similar in the VCaP background as well. Differences in downstream activities of AR isoforms were also evident from comparison of the metabolome and lipidome in an LNCaP model. Overall our study implies that shorter variants of AR are capable of mediating unique downstream activities different from AR and some of these are isoform specific.</p>
Figure 1 from: Abbey M (2016) Functional characterization of the several splice variants of Fmr1. Research Ideas and Outcomes 2: e10593. https://doi.org/10.3897/rio.2.e10593
Figure 1 - Diagrammatic representation of the exon structure of Fmr1 and its corresponding functional domains. The splice acceptor sites are marked as 1 and 2.
Data from: Nonsense-mediated decay of alternative pre-mRNA splicing variants is a major determinant of the Arabidopsis steady state transcriptome
Open the record for dataset details and reuse information.
Isoform specific activities of androgen receptor and its splice variants in prostate cancer cells
Open the record for dataset details and reuse information.
MATR3 pathogenic variants differentially impair its cryptic splicing repression function
GEO Series GSE205343. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.
Expression profiling by RNA-seq of LNCaP cells expressing wild-type androgen receptor (AR-WT), AR-V7 splice variant or mutant AR-Q641X
GEO Series GSE158557. Homo sapiens. 24 samples. Type: Expression profiling by high throughput sequencing.
Androgen Receptor splice variant V7 (AR-V7) mediates AR signalling in castration resistant prostate cancer (CRPC) [RNA-seq]
GEO Series GSE143905. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing.
Whole-transcriptome analysis identifies re-expression of fetal splice variants in cardiac hypertrophy
GEO Series GSE42411. Rattus norvegicus. 9 samples. Type: Expression profiling by high throughput sequencing.
Androgen receptor splicing variant 7 (ARv7) promotes DNA damage response in prostate cancer cells
GEO Series GSE277204. Homo sapiens. 12 samples. Type: Expression profiling by high throughput sequencing.
Generation 2.5 Antisense Oligonucleotides Targeting the Androgen Receptor and its Splice Variants Suppress Enzalutamide-Resistant Prostate Cancer Cell Growth
GEO Series GSE55345. Homo sapiens. 6 samples. Type: Genome variation profiling by genome tiling array.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.