Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
725
datasets available to search
ShareScore release 0.9.0
Dataset results
725 results for “Isoforms”
Biosurfer for systematic tracking of regulatory mechanisms leading to protein isoform diversity
<p>This Zenodo repository contains data used for running <a href="https://github.com/sheynkman-lab/biosurfer_analysis" target="_blank" rel="noopener">Biosurfer_analysis</a>, a tool to surf the biological network, from genome to transcriptome to proteome and back to gain insights into human disease biology.<br><br>Repository content:</p> <ol> <li><strong>biosurfer_gencode_toy_data.zip:</strong> is a small subset of the GENCODE version 38 files (GTF, transcript FASTA, and protein FASTA) for the purpose of trial run and testing Biosurfer scripts.<br><br></li> <li><strong>biosurfer_gencode_toy_output.zip: </strong>Biosurfer generated output files for the toy data. <br><br></li> <li><strong>biosurfer_gencode_v42_data.zip: </strong>contains GENCODE version 42 (basic) files (GTF, transcript FASTA, and protein FASTA).<br><br></li> <li><strong>biosurfer_gencode_v42_output.zip: </strong>Biosurfer generated output files for GENCODE 42 input files.<br><br></li> <li><strong>biosurfer_wtc11_data.zip: </strong>contains karyotypically normal human stem cell line data (WTC11) file (GTF, transcript FASTA, and protein FASTA).<br><strong><br></strong></li> <li><strong>biosurfer_wtc11_output.zip: </strong>Biosurfer generated output files for WTC11 files.<br><br></li> <li><strong>APPRIS analysis.zip: </strong>Intermediate files (CSV) detailing the APPRIS isoforms information utilized in the associated manuscript.</li> <li><strong>biosurfer_mouse.zip: </strong>contains GENCODE Mouse version M35 (basic) files (GTF, transcript FASTA, and protein FASTA).</li> <li><strong>biosurfer_mouse_output.zip: </strong>Biosurfer generated output files for GENCODE M35 input files.</li> </ol>
Isoform analysis Galaxy training with EWS-FLI1 example
<p>The fastqs are subsampled from the following SRA:</p> <p><br>SRR8561337<br>SRR8561334<br>SRR8561335<br>SRR8561332<br>SRR8561333<br>SRR8561347</p> <p>The gtf file was reduced from https://ftp.ebi.ac.uk/pub/databases/gencode/Gencode_human/release_19/gencode.v19.annotation.gtf.gz and only lines with 'chr10' in first column were kept.</p> <p>The fasta file comes from https://hgdownload.soe.ucsc.edu/goldenPath/hg19/chromosomes/chr10.fa.gz</p>
Dataset of EPR-monitored-redox potentiometries performed with distinct CrFdx isoforms and variants
<p>The zip-folder contains the raw EPR data and data necessary for construction of the Nernst fits that appear in:</p> <p>Melanie Heghmanns, Alexander Günzel, Dörte Brandis, Yury Kutin, Vera Engelbrecht, Martin Winkler, Thomas Happe, Müge Kasanmascheff, "Fine-tuning of FeS Proteins Monitored via Pulsed EPR Redox Potentiometry at Q-band", Biophysical Reports, 2021, https://doi.org/10.1016/j.bpr.2021.100016.</p> <p>A brief explanation of the files and samples can be found in “Readme.docx”. <br> The relevant data obtained are compiled in “Summary_data.xlsx”. <br> Other experimental details can be found in the paper.</p>
Electrophysiology Data for "Two functional epithelial sodium channel isoforms are present in rodents despite pronounced evolutionary pseudogenization and exon fusion"
<p>Here we provide the electrophysiology data for the manuscript "Two functional epithelial sodium channel isoforms are present in rodents despite pronounced evolutionary pseudogenization and exon fusion", published in Molecular Biology and Evolution (2021): msab271 (doi: 10.1093/molbev/msab271). Data are reported as current values in Excel format, sorted according to the appearance in Figures and supplemented by explanatory text on the procedures/data presentation.</p>
Spliced isoforms of the cardiac Nav1.5 channel modify channel activation by distinct structural mechanisms: Molecular dynamics coordinate and trajectory files
<p>Molecular dynamics files associated with the publication: Spliced isoforms of the cardiac Nav1.5 channel modify channel activation by distinct structural mechanisms</p>
ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data (repository for simulated ONT RNA-seq data)
<p>Simulated ONT direct RNA and 1D cDNA sequencing data of varying sequencing depths (0.5 million, 1 million, 3 million, and 5 million simulated reads) used for benchmark evaluations of transcript discovery and quantification in our paper "ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data". All details can be found in the <strong>Materials and Methods</strong> section of the paper. </p> <p><em>HEK293T_DirectRNA.transcriptome_quantification.tsv</em> and <em>HEK293T_DirectRNA.transcriptome_quantification.tsv </em>are tab-separated files containing estimated raw read counts and normalized abundance values (in TPM) of transcripts annotated in GENCODE v34lift37. Transcript quantification was done using NanoSim (version 3.1.0). </p> <p><em>HEK293T_DirectRNA.NanoSim_500k.fastq.gz</em>,<em> </em><em>HEK293T_DirectRNA.NanoSim_1M.fastq.gz</em>, <em>HEK293T_DirectRNA.NanoSim_3M.fastq.gz</em>, and<em> HEK293T_DirectRNA.NanoSim_5M.fastq.gz </em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT direct RNA sequencing reads respectively. </p> <p><em>HEK293T_1DcDNA.NanoSim_500k.fastq.gz</em>,<em> HEK293T_1DcDNA.NanoSim_1M.fastq.gz</em>, <em>HEK293T_1DcDNA.NanoSim_3M.fastq.gz</em>, and<em> HEK293T_1DcDNA.NanoSim_5M.fastq.gz </em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT 1D cDNA sequencing reads respectively. </p>
Long read isoforms with A-to-I edits for H1975 cell line
<p>FLAIR2 was used to determine isoform models sequenced from H1975 cells using R2C2 cDNA nanopore sequencing. RNA sequence variants observed in the data were called with longshot and integrated into FLAIR2 isoform models.</p>
Phase 1/2 Study of Enasidenib (AG-221) in Adults With Advanced Hematologic Malignancies With an Isocitrate Dehydrogenase Isoform 2 (IDH2) Mutation
ClinicalTrials.gov study NCT01915498. IPD Sharing: Not stated. Countries: 2. Publications: 12.
Molecular Dynamics Simulations and associated data for: Mechanistic and evolutionary insights into isoform-specific 'supercharging' in DCLK family kinases
Open the record for dataset details and reuse information.
A distinct isoform of lymphoid enhancer binding factor 1 (LEF1) epigenetically restricts EBV reactivation to maintain viral latency
Open the record for dataset details and reuse information.
Data from: Long-term severe hypoxia adaptation induces non-canonical EMT and a novel Wilms Tumor 1 (WT1) isoform
Open the record for dataset details and reuse information.
An alternative cytoplasmic SFPQ isoform with reduced phase separation potential is upregulated in ALS
Open the record for dataset details and reuse information.
Data from: Cis-regulatory differences in isoform expression associate with life history strategy variation in Atlantic salmon
Open the record for dataset details and reuse information.
Data from: O-GlcNAc modification differentially regulates microtubule binding and pathological conformations of tau isoforms in vitro
Open the record for dataset details and reuse information.
Ancient medicinal plant rosemary contains a highly efficacious and isoform-selective KCNQ potassium channel opener
Open the record for dataset details and reuse information.
Drosophila SWR1 and NuA4 complexes are defined by DOMINO isoforms
<p>Histone acetylation and deposition of H2A.Z variant are integral aspects of active transcription. In <i>Drosophila</i>, the single DOMINO chromatin regulator complex is thought to combine both activities <i>via</i> an unknown mechanism. Here we show that alternative isoforms of the DOMINO<i> </i>nucleosome remodeling ATPase, DOM-A and DOM-B, directly specify two distinct multi-subunit complexes. Both complexes are necessary for transcriptional regulation but through different mechanisms. The DOM-B complex incorporates H2A.V (the fly ortholog of H2A.Z) genome-wide in an ATP-dependent manner, like the yeast SWR1 complex. The DOM-A complex, instead, functions as an ATP-independent histone acetyltransferase complex similar to the yeast NuA4, targeting lysine 12 of histone H4. Our work provides an instructive example of how different evolutionary strategies lead to similar functional separation. In yeast and humans, nucleosome remodeling and histone acetyltransferase complexes originate from gene duplication and paralog specification. <i>Drosophila</i> generates the same diversity by alternative splicing of a single gene.</p>
MD data for the article "Structure comparison of beta amyloid peptide Aβ 1-42 isoforms. Molecular dynamics modeling" by Anna P. Tolstova, Alexander A. Makarov, Alexei A. Adzhubei.
<p>There are CMD and REMD trajectories for Aβ isoforms discussed in the paper together with final coordinate files for these trajectories. The resulting dataset of modeled structures includes wild type Aβ42, isoD7, pS8, D7H and H6R-Aβ42, and wild type Aβ16, isoD7, pS8, D7H and H6R-Aβ16.</p>
aSyn isoforms
Open the record for dataset details and reuse information.
TPM quantification (StringTie) of long-read isoforms in TCGA breast tumors and GTEx
<p>StringTie was applied to quantify the expression of long-read isoforms sequenced in TCGA and GTEx samples.</p> <p>Inputs:</p> <ul> <li>RNA-seq samples (bam files) from breast cancer samples from TCGA (tumors) and GTEx tissues (healthy controls).</li> <li>GTF: full-length isoforms sequenced using PacBio long-read sequencing in breast tumors. See <a href="https://www.science.org/doi/10.1126/sciadv.abg6711">A comprehensive long-read isoform analysis platform and sequencing resource for breast cancer</a>.</li> </ul> <p>Output: tab-delimited files with TPM values for isoforms parsed from StringTie results.</p> <p>More details: <a href="https://www.science.org/doi/10.1126/sciadv.abg6711">A comprehensive long-read isoform analysis platform and sequencing resource for breast cancer</a></p> <p> </p> <p> </p> <p> </p> <p> </p>
MD simulation trajectories for paper "Distinct Structures and Dynamics of Chromatosomes with Different Human Linker Histone Isoforms"
<div> <div>The dataset includes the MD simulation trajectories for the paper "Distinct Structures and Dynamics of Chromatosomes with Different Human Linker Histone Isoforms". </div> <div>Citation: Zhou BR, Feng H, Kale S, Fox T, Khant H, de Val N, Ghirlando R, Panchenko AR, Bai Y. Distinct Structures and Dynamics of Chromatosomes with Different Human Linker Histone Isoforms. Mol Cell. 2021 Jan 7;81(1):166-182.e6. PMID: 33238161; PMCID: PMC7796963.</div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.