Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

725

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

725 results for “Isoforms”

Learn how ShareScore rates datasets ↗
zenodo44/100

Identification and Functional Characterization of an Alternative Cancer-derived PD-L1 Isoform (supplemental data)

<p>The enclosed files contain all of the supplemental data from: Identification and Functional Characterization of an Alternative Cancer-derived PD-L1 Isoform. The files include the complete tables in CSV-formatted files.</p>

opencc-by-4.0Sep 2018View details →
zenodo44/100

The spatial landscape of gene expression isoforms in tissue sections

<p><strong>This upload&nbsp;provides raw in situ sequencing (ISS) data used to validate Spatial Isoform&nbsp;Transcriptomics (SiT), as well as&nbsp;R scripts required for SiT analysis.</strong></p> <p><strong>GenePlots.zip and Reads.zip are ISS data </strong><strong>generated and collected by the CARTANA ISS service</strong>.&nbsp;<strong>The following data description is cited from the&nbsp;report provided by CARTANA ISS service:</strong></p> <p><em>&quot;Folder &quot;Reads&quot; contains coordinates and gene information of segmented spots.<br> The coordinates are in pixel unit. Scaling factor is 0.32 um/pixel. (0,0) is at northwest (top-left corner).<br> With Low/High Threshold, we refer to the quality thresholding. Our technology is fluorescence based, i.e. with the thresholding one can balance how certain the signals are.</em></p> <p><em>Files ending with _LowThreshold: reads not matching with any known barcode were already discarded.</em></p> <p><em>Files ending with _HighThreshold: has information only about spots that passed additional quality check.</em></p> <p><em>Folder &quot;GenePlots&quot; has plotted images in static .png format, fully zoomed out. LowThreshold and HighThreshold follow the same thresholding strategy as in reads files.&quot;</em></p> <p>&nbsp;</p> <p><strong>SiT-master.zip is a download of the&nbsp;GitHub repository </strong><a href="https://github.com/ucagenomix/SiT">https://github.com/ucagenomix/SiT</a>,&nbsp;<strong>providing figures and analysis scripts for SiT.</strong></p> <p>&nbsp;</p> <p><strong>Related SiT data are deposited through&nbsp;GEO, accession number&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE153859">GSE153859</a></strong></p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Amyloid-motif-dependent tau self-assembly is modulated by isoform sequence context

<p><span>The microtubule-associated protein tau is implicated in neurodegenerative diseases characterized by amyloid formation. Mutations associated with frontotemporal dementia increase tau aggregation propensity and disrupt its endogenous microtubule-binding activity. However, the structural relationship between aggregation propensity and biological activity remains unclear. We employed a multi-disciplinary approach, including computational modeling, NMR, cross-linking mass spectrometry, and cell models to engineer tau sequences that modulate its structural ensemble. Our findings show that substitutions near the conserved 'PGGG' </span><span>&beta;</span><span>-turn motif informed by tau isoform context reduce tau aggregation in vitro and cells and can even counteract aggregation from disease-associated proline-to-serine mutations. Engineered tau sequences maintain microtubule binding and explain why 3R isoforms exhibit reduced pathogenesis compared to 4R. We propose a simple mechanism to reduce the formation of pathogenic tau species while preserving biological function, thus offering insights for therapeutic strategies aimed at reducing tau protein misfolding in neurodegenerative diseases.</span></p> <p><strong>Description of Source Data and Supplementary Data</strong>: All MD, NMR (peptide and tauRD), ThT, XL-MS, MT stabilization, MT:tau modeling, and cell-based aggregation data are available in the Source_Data directory as Data S1, Data S2, Data S3, Data S4, Data S5, Data S6, Data S7, and Data S8, respectively. Supplementary Data for raw MD trajectory files, structure files for MSM modeling and validation file, and tau:MT modeling are available as "Supplementary_Data_MD_Trajectories_Structures", "Supplementary_Data_MSM_models_Structures" and "Supplementary_Data_MT-tau_complex_models", respectively."</p>

opencc-by-4.0Sep 2024View details →
dryad40/100

Developmental isoform diversity in the human neocortex informs neuropsychiatric risk mechanisms

<p>RNA splicing is highly prevalent in the brain and has strong links to neuropsychiatric disorders, yet the role of cell-type-specific splicing or transcript-isoform diversity during human brain development has not been systematically investigated. Here, we leveraged single-molecule long-read sequencing to deeply profile the full-length transcriptome of the germinal zone (GZ) and cortical plate (CP) regions of the developing human neocortex at tissue and single-cell resolution. We identified 214,516 unique isoforms, of which 72.6% are novel (unannotated in Gencode-v33), and uncovered a substantial contribution of transcript-isoform diversity, regulated by RNA binding proteins, in defining cellular identity in the developing neocortex. We leveraged this comprehensive isoform-centric gene annotation to re-prioritize thousands of rare de novo risk variants and elucidate genetic risk mechanisms for neuropsychiatric disorders.</p>

opencc-zeroMar 2024View details →
zenodo40/100

Enhanced Protein Isoform Characterization Through Long-Read Proteogenomics - Workflow Results

<pre>&nbsp;</pre> <p>The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Long-Read-Proteogenomics Workflow Sample and Reference Data</a></li> <li><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></li> </ol> <p>This Repository contains the complete output from the execution of the&nbsp;<a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow</a>, using the input from&nbsp;<a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a>.&nbsp; &nbsp;</p> <p>The file&nbsp;<em>jurkat.flnc.bam&nbsp;</em>was 6.5 GB had to be split into 13 separate files and for use should be rejoined -- here are the steps that were used to split the file up.&nbsp; &nbsp;</p> <p>1. Convert&nbsp;<em>jurkat.flnc.bam</em>&nbsp;(binary format) to sam file (text format) without header:&nbsp;&nbsp;<em>samtools view jurkat.flnc.bam &gt; jurkat.flnc.sam</em></p> <p>2. Capture the header:&nbsp;<em>samtools view -H jurkat.flnc.bam &gt; jurkat.flnc.header.sam</em></p> <p>3. Split&nbsp;<em>jurkat.flnc.sam</em>&nbsp;into smaller files (aim to get final size under 2GB):&nbsp;<em>split -l 400000 jurkat.flnc.sam jurkat.flnc.chunk.</em></p> <p>4. Convert each of these files back to bam for uploading:&nbsp;<em>samtools view -b jurkat.flnc.chunk.a* -o jurkat.flnc.chunk.a*.bam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>After downloading, reverse this process including using the header file which is found in the&nbsp;LRPG-Manuscript-Results-results-results-jurkat-isoseq3-companion-files.tar.gz file&gt;</p> <p>1. Convert the bam files back to sam files:&nbsp;<em>samtools view jurkat.flnc.chunk.a*.bam &gt; jurkat.flnc.chunk.a*.sam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>2. Combine the header together with the sam files:&nbsp;<em>cat jurkat.flnc.chunk.a*sam &gt; jurkcat.flnc.sam (</em>verified the same number of lines of the sam files is identical to the number of lines of the original without header: 4,956,761.&nbsp; Header file is 13 lines.</p> <p>3. Convert to bam files if desired:&nbsp;<em>samtools view -b jurkat.flnc.sam -o jurkat.flnc.bam</em></p> <p>4. Rehead with the header file:&nbsp;<em>samtools reheader -P -i jurkat.flnc.header.sam jurkat.flnc.bam</em></p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Underlying data for IsoAligner: dynamic mapping of amino acidpositions across protein isoforms

<p>The human isoform library (list_of_gene_objects_25th_july_final.txt) for the IsoAligner webtool&nbsp;is generated from these resources.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Long read proteogenomics to characterize protein isoform diversity in human umbilical vein endothelial cells (HUVECs)

<p>Endothelial cells (ECs) comprise the lumenal lining of all blood vessels and are critical for the functioning of the cardiovascular system and their phenotypes can be modulated by protein isoforms. To characterize the isoform landscape within EC, we applied a long read proteogenomics approach to analyze human umbilical vein endothelial cells (HUVECs). Transcripts delineated from PacBio sequencing serve as the basis for a sample-specific protein database used for downstream MS analysis to infer protein isoform expression. We detected 53,836 transcript isoforms from 10,426 genes, with 22,195 of those transcripts being novel. Furthermore, the predominant isoform in HUVECs does not correspond with the accepted &ldquo;reference isoform&rdquo; 25% of the time, with vascular pathway-related genes among this group. We found 2,597 protein isoforms supported through unique peptides, with an additional 2,280 isoforms nominated upon incorporation of long-read transcript evidence. We characterized a novel alternative acceptor for endothelial-related gene <em>CDH5</em>, suggesting potential changes in its associated signaling pathways. Finally, we identified novel protein isoforms arising from a diversity of splicing mechanisms supported by uniquely mapped novel peptides. Our results represent a high resolution atlas of known and novel isoforms of potential relevance to endothelial phenotypes and function.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Dyregulated miRNA isoforms across TCGA and TARGET cohorts

<p>The table reports the complete list of dysregulated miRNA isoform molecules across cohorts/cancer tissues, retained according to a |<em>linear fold change</em>| &gt;1.5 and an <em>FDR adjusted p-value</em> &lt;0.05.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Figure 1 in The Combined Expression Patterns of Ikaros Isoforms Characterize Different Hematological Tumor Subtypes

Figure 1. - Location of sites cited in table I, where the new records were gathered. 1) San Jorge, 2) La Poma, 3) Punta del Diablo, 4) Patos Island, 5) El Chivero, 6) La Reina, 7) La Tordilla, 8) San Pedro Mártir.

opencc-by-4.0Dec 2013View details →
zenodo40/100

TEST DATA for Enhanced protein isoform characterization through long-read proteogenomics

<p>Test data for&nbsp;The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a></li> <li><a href="http://10.5281/zenodo.5920920">Long-Read-Proteogenomics Workflow Results using Jurkat Sample data</a></li> </ol> <p>This Repository contains the test data, specifically:</p> <p><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></p>

opencc-by-4.0Jul 2021View details →
dryad40/100

Developmental isoform diversity in the human neocortex informs neuropsychiatric risk mechanisms

Open the record for dataset details and reuse information.

publicMar 2024View details →
dryad36/100

Data from: Cis-regulatory differences in isoform expression associate with life history strategy variation in Atlantic salmon

<p><span><span><span><span><span><span><span><span><span><span><span><b>A major goal in biology is to understand how evolution shapes variation in individual life histories. Genome-wide association studies have been successful in uncovering genome regions linked with traits underlying life history variation in a range of species. However, lack of functional studies of the discovered genotype-phenotype associations severely restrains our understanding how alternative life history traits evolved and are mediated at the molecular level. Here, we report a <i>cis</i>-regulatory mechanism whereby expression of alternative isoforms of the transcription co-factor <i>vestigial-like 3</i> (<i>vgll3</i>) associate with variation in a key life history trait, age at maturity, in Atlantic salmon (<i>Salmo salar</i>). Using a common-garden experiment, we first show that <i>vgll3 </i>genotype associates with puberty timing in one-year-old salmon males. By way of temporal sampling of <i>vgll3 </i>expression in ten tissues across the first year of salmon development, we identify a pubertal transition in <i>vgll3</i> expression where maturation coincided with a 66% reduction in testicular <i>vgll3</i> expression. The <i>late </i>maturation allele was not only associated with a tendency to delay puberty, but also with expression of a rare transcript isoform of <i>vgll3</i> pre-puberty. By comparing absolute <i>vgll3 </i>mRNA copies in heterozygotes we show that the expression difference between the <i>early</i>and <i>late</i> maturity alleles is largely <i>cis</i>-regulatory. We propose a model whereby expression of a rare isoform from the <i>late </i>allele shifts the liability of its carriers towards delaying puberty. These results exemplify the potential importance of regulatory differences as a mechanism for the evolution of life history traits.</b></span></span></span></span></span></span></span></span></span></span></span></p>

opencc-zeroAug 2020View details →
zenodo36/100

Dateset of in silico investigations on protein_protein Interaction of GST isoforms with ASK1 and JNK1

<p>Please refer to the description of the files document for details on the files in the zip folder.</p> <p>The dataset contains molecular dynamics simulation files, and other in silico investigation&nbsp;output files used to describe the protein-protein interactions of seven GST isoforms with that of MAPK8 (JNK1) and MAP3K5 (ASK1).&nbsp;</p>

opencc-by-4.0Jun 2020View details →
dryad36/100

Molecular Dynamics Simulations and associated data for: Mechanistic and evolutionary insights into isoform-specific 'supercharging' in DCLK family kinases

<p>Catalytic signaling outputs of protein kinases are dynamically regulated by an array of structural mechanisms, including allosteric interactions mediated by intrinsically disordered segments flanking the conserved catalytic domain. The Doublecortin Like Kinases (DCLKs) are a family of microtubule-associated proteins characterized by a flexible C-terminal autoregulatory 'tail' segment that varies in length across the various human DCLK isoforms. However, the mechanism whereby these isoform-specific variations contribute to unique modes of autoregulation is not well understood. Here, we employ a combination of statistical sequence analysis, molecular dynamics simulations and in vitro mutational analysis to define hallmarks of DCLK family evolutionary divergence, including analysis of splice variants within the DCLK1 sub-family, which arise through alternative codon usage and serve to 'supercharge' the inhibitory potential of the DCLK1 C-tail. We identify co-conserved motifs that readily distinguish DCLKs from all other Calcium Calmodulin Kinases (CAMKs), and a 'Swiss-army' assembly of distinct motifs that tether the C-terminal tail to conserved ATP and substrate-binding regions of the catalytic domain to generate a scaffold for auto-regulation through C-tail dynamics. Consistently, deletions and mutations that alter C-terminal tail length or interfere with co-conserved interactions within the catalytic domain alter intrinsic protein stability, nucleotide/inhibitor-binding, and catalytic activity, suggesting isoform-specific regulation of activity through alternative splicing. Our studies provide a detailed framework for investigating kinome–wide regulation of catalytic output through cis-regulatory events mediated by intrinsically disordered segments, opening new avenues for the design of mechanistically-divergent DCLK1 modulators, stabilizers or degraders.</p>

opencc-zeroOct 2023View details →
dryad36/100

A distinct isoform of lymphoid enhancer binding factor 1 (LEF1) epigenetically restricts EBV reactivation to maintain viral latency

<p>As a human tumor virus, EBV is present as a latent infection in its associated malignancies where genetic and epigenetic changes have been shown to impede cellular differentiation and viral reactivation. We reported previously that levels of the Wnt signaling effector, lymphoid enhancer binding factor 1 (LEF1) increased following EBV epithelial infection and an epigenetic reprogramming event was maintained even after loss of the viral genome. Elevated LEF1 levels are also observed in nasopharyngeal carcinoma and Burkitt lymphoma. To determine the role played by LEF1 in the EBV life cycle, we used in silico analysis of EBV type 1 and 2 genomes to identify over 20 Wnt-response elements, which suggests that LEF1 may bind directly to the EBV genome and regulate the viral life cycle. Using CUT&amp;RUN-seq, LEF1 was shown to bind the latent EBV genome at various sites encoding viral lytic products that included the immediate early transactivator BZLF1 and viral primase BSLF1 genes. The LEF1 gene encodes various long and short protein isoforms. siRNA depletion of specific LEF1 isoforms revealed that the alternative-promoter derived isoform with an N-terminal truncation (∆N LEF1) transcriptionally repressed lytic genes associated with LEF1 binding. In addition, forced expression of the ∆N LEF1 isoform antagonized EBV reactivation. As LEF1 repression requires histone deacetylase activity through either recruitment of or direct intrinsic histone deacetylase activity, siRNA depletion of LEF1 resulted in increased histone 3 lysine 9 and lysine 27 acetylation at LEF1 binding sites and across the EBV genome. Taken together, these results indicate a novel role for LEF1 in maintaining EBV latency and restriction viral reactivation via repressive chromatin remodeling of critical lytic cycle factors.</p>

opencc-zeroDec 2023View details →
zenodo36/100

Supplementary data: The APOE isoforms differentially shape the transcriptomic and epigenomic landscapes of human microglia in a xenotransplantation model of Alzheimer's disease

<p>Supplementary data for: The APOE isoforms differentially shape the transcriptomic and epigenomic landscapes of human microglia in a xenotransplantation model of Alzheimer&rsquo;s disease.&nbsp;</p> <p>Supplementary_Table1_QC: Excel sheet containing QC metrics for the RNA-seq data and the other containing QC metrics for the ATAC-seq data.&nbsp;</p> <p>Supplementary_Table2_DEGs: Excel sheet containing DeSeq2 differential expression analysis results for the following comparisons: APOE2 vs APOE3, APOE4 vs APOE3, APOE4 vs APOE2, APOE-KO vs APOE3.&nbsp;</p> <p>Supplementary_Table3_MAGMA_geneset_analysis_res: CSV file containing MAGMA gene set analysis results using the differentially expressed genes (FDR &lt; 0.05) for the comparisons outlined in Supplementary_Table2_DEGs and three independent AD GWAS.&nbsp;</p> <p>Supplementary_Table4_DARs: Excel sheet containing DeSeq2 differential accessibility analysis results for the following comparisons: APOE2 vs APOE3, APOE4 vs APOE3, APOE4 vs APOE2, APOE-KO vs APOE3.&nbsp;</p> <p>Supplementary_Table5_sLDSC_res.csv: CSV file containing s-LDSC results using the consensus set of ATAC-seq peaks with three brain disorder GWAS (Alzheimer's disease, autism spectrum disorder, and amyotrophic lateral sclerosis).&nbsp;</p> <p>Supplementary_Table6_WGCNA_clusterProfiler_pathway_enrichment.csv: CSV file containing pathway enrichment results using two WGCNA-identified modules that were significantly upregulated in APOE2-expressing microglia.&nbsp;</p> <p>Supplementary_Table7_homer_motifEnrichment_res.xlsx: Excel sheet containing Homer motif enrichment analysis results using top 100 peaks with increased and decreased chromatin accessibility for APOE2 vs APOE3, APOE4 vs APOE3, and APOE4 vs APOE2.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Simulated data used in "AIDE: annotation-assisted isoform discovery with high precision"

<p>This repository contains the eight simulated RNA-seq samples (sample1, sample2, ..., sample8 with coverages10x, 20x, ...,&nbsp;80x) simulated from the annotation file &quot;truth.gtf&quot; by R package polyester. The simulation procedure is described in the Methods section &quot;Simulation for comparing isoform reconstruction methods&quot;. For example, sample 1 has mapped and sorted&nbsp;reads in &quot;sample1.sorted.bam&quot;.</p> <p>When we generate the benchmarking results in the Results section &quot;AIDE outperforms state-of-the-art methods on simulated data&quot;,&nbsp;we supply each isoform identification/quantification method (AIDE, Cufflinks, StringTie, or SLIDE) with one&nbsp;synthetic annotation file (in &quot;aGTF.zip&quot;) and one bam file (e.g., sample1.sorted.bam).</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Full-Length Spatial Transcriptomics Reveals the Unexplored Isoform Diversity of the Myocardium Post-MI

<p>We introduce Single-cell Nanopore Spatial Transcriptomics (scNaST), a software suite to facilitate the analysis of spatial gene expression from second- and third-generation sequencing, allowing to generate a full-length near-single-cell transcriptional landscape of the tissue microenvironment. Taking advantage of the Visium Spatial platform, we adapted a strategy recently developed to assign barcodes to long-read single-cell sequencing data for spatial capture technology. Here, we demonstrate our workflow using four short axis sections of the mouse heart following myocardial infarction. We constructed a <em>de novo</em> transcriptome using long-read data, and successfully assigned 19,794 transcript isoforms in total, including clinically-relevant, but yet uncharacterized modes of transcription, such as intron retention or antisense overlapping transcription. We showed a higher transcriptome complexity in the healthy regions, and identified intron retention as a mode of transcription associated with the infarct area. Our data revealed a clear regional isoform switching among differentially used transcripts for genes involved in cardiac muscle contraction and tissue morphogenesis. Molecular signatures involved in cardiac remodeling integrated with morphological context may support the development of new therapeutics towards the treatment of heart failure and the reduction of cardiac complications.</p>

opencc-by-4.0May 2022View details →
dryad36/100

Data from: Long-term severe hypoxia adaptation induces non-canonical EMT and a novel Wilms Tumor 1 (WT1) isoform

<p>The majority of cancer deaths are caused by solid tumors, where the four most prevalent cancers (breast, lung, colorectal and prostate) account for more than 60% of all cases (1). Tumor cell heterogeneity driven by variable cancer microenvironments, such as hypoxia, is a key determinant of therapeutic outcome. We developed a novel culture protocol, termed the Long-Term Hypoxia (LTHY) time course, to recapitulate the gradual development of severe hypoxia seen in vivo to mimic conditions observed in primary tumors. Cells subjected to LTHY underwent a non-canonical epithelial to mesenchymal transition (EMT) based on miRNA and mRNA signatures as well as displayed EMT-like morphological changes. Concomitant to this, we report production of a novel truncated isoform of WT1 transcription factor (tWt1), a non-canonical EMT driver, with expression driven by a yet undescribed intronic promoter through hypoxia-responsive elements (HREs). We further demonstrated that tWt1 initiates translation from an intron-derived start codon, retains proper subcellular localization and DNA binding. A similar tWt1 is also expressed in LTHY-cultured human cancer cell lines as well as primary cancers and predicts long-term patient survival. Our study not only demonstrates the importance of culture conditions that better mimic those observed in primary cancers, especially with regards to hypoxia, but also identifies a novel isoform of WT1 which correlates with poor long-term survival in ovarian cancer.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Ranger models for predicting isoform abundance from UTR sequence features

<p>Each RDS file contains a ranger object, trained on transcripts after removing those associated with the held-out genes in one of the five cross-validation folds. The day and replicate number in the file name corresponds to the neuronal differentiation sample on which the model was trained. The file gene_folds.txt indicates the fold from which each gene was excluded during model training. The file transcript_gene_associations.txt contains transcript-gene associations. The file predictors.RDS contains the matrix of predictor variables.</p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record