Skip to main content
zenodoopen

Enhanced Protein Isoform Characterization Through Long-Read Proteogenomics - Workflow Results

<pre>&nbsp;</pre> <p>The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Long-Read-Proteogenomics Workflow Sample and Reference Data</a></li> <li><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></li> </ol> <p>This Repository contains the complete output from the execution of the&nbsp;<a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow</a>, using the input from&nbsp;<a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a>.&nbsp; &nbsp;</p> <p>The file&nbsp;<em>jurkat.flnc.bam&nbsp;</em>was 6.5 GB had to be split into 13 separate files and for use should be rejoined -- here are the steps that were used to split the file up.&nbsp; &nbsp;</p> <p>1. Convert&nbsp;<em>jurkat.flnc.bam</em>&nbsp;(binary format) to sam file (text format) without header:&nbsp;&nbsp;<em>samtools view jurkat.flnc.bam &gt; jurkat.flnc.sam</em></p> <p>2. Capture the header:&nbsp;<em>samtools view -H jurkat.flnc.bam &gt; jurkat.flnc.header.sam</em></p> <p>3. Split&nbsp;<em>jurkat.flnc.sam</em>&nbsp;into smaller files (aim to get final size under 2GB):&nbsp;<em>split -l 400000 jurkat.flnc.sam jurkat.flnc.chunk.</em></p> <p>4. Convert each of these files back to bam for uploading:&nbsp;<em>samtools view -b jurkat.flnc.chunk.a* -o jurkat.flnc.chunk.a*.bam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>After downloading, reverse this process including using the header file which is found in the&nbsp;LRPG-Manuscript-Results-results-results-jurkat-isoseq3-companion-files.tar.gz file&gt;</p> <p>1. Convert the bam files back to sam files:&nbsp;<em>samtools view jurkat.flnc.chunk.a*.bam &gt; jurkat.flnc.chunk.a*.sam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>2. Combine the header together with the sam files:&nbsp;<em>cat jurkat.flnc.chunk.a*sam &gt; jurkcat.flnc.sam (</em>verified the same number of lines of the sam files is identical to the number of lines of the original without header: 4,956,761.&nbsp; Header file is 13 lines.</p> <p>3. Convert to bam files if desired:&nbsp;<em>samtools view -b jurkat.flnc.sam -o jurkat.flnc.bam</em></p> <p>4. Rehead with the header file:&nbsp;<em>samtools reheader -P -i jurkat.flnc.header.sam jurkat.flnc.bam</em></p>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
20
Reuse readiness
8
Engagement
0

Topics