Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
75
datasets available to search
ShareScore release 0.9.0
Dataset results
75 results for “Proteogenomics”
Development of a spectral library for the discovery of altered genomic events in Mycobacterium avium associated with virulence using mass spectrometry-based proteogenomic analysis
<p><em>Mycobacterium avium</em> is one of the prominent disease-causing bacteria in humans. It causes lymphadenitis, chronic and extrapulmonary, and disseminated infections in adults, children, and immunocompromised patients. <em>M. avium</em> has ~4,500 predicted protein-coding regions on an average, which can be helpful in discovering several variants at the proteome level. Many of them are potentially associated with virulence, thus identifying such proteins can be a helpful feature in the development of panel-based theranostics. In line with such a long-term goal, we carried out an in-depth proteomic analysis of <em>M. avium</em> with both data-dependent and data-independent acquisition methods. Further, a set of proteogenomic investigations were carried out using the protein database for <em>Mycobacterium tuberculosis,</em> and a genome six-frame translated database and a variant protein database of <em>M. avium</em>. A search of mass spectrometry data analysis against <em>M. avium</em> protein database resulted in the identification of 2,954 proteins. Further, proteogenomic analyses aided in the identification of 1,301 novel peptide sequences and correction of translation start sites for 15 proteins. At the end, we created a spectral library of <em>M. avium</em> proteins including novel genome search-specific peptides and variant peptides detected in this study. We validated the spectral library by a data-independent acquisition of the <em>M. avium</em> proteome. Thus, we present a <em>M. avium </em>spectral library of 29,033 peptide precursors supported by 0.4 million fragment ions for further use by the biomedical community.</p>
Enhanced Protein Isoform Characterization Through Long-Read Proteogenomics - Workflow Results
<pre> </pre> <p>The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Long-Read-Proteogenomics Workflow Sample and Reference Data</a></li> <li><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></li> </ol> <p>This Repository contains the complete output from the execution of the <a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow</a>, using the input from <a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a>. </p> <p>The file <em>jurkat.flnc.bam </em>was 6.5 GB had to be split into 13 separate files and for use should be rejoined -- here are the steps that were used to split the file up. </p> <p>1. Convert <em>jurkat.flnc.bam</em> (binary format) to sam file (text format) without header: <em>samtools view jurkat.flnc.bam > jurkat.flnc.sam</em></p> <p>2. Capture the header: <em>samtools view -H jurkat.flnc.bam > jurkat.flnc.header.sam</em></p> <p>3. Split <em>jurkat.flnc.sam</em> into smaller files (aim to get final size under 2GB): <em>split -l 400000 jurkat.flnc.sam jurkat.flnc.chunk.</em></p> <p>4. Convert each of these files back to bam for uploading: <em>samtools view -b jurkat.flnc.chunk.a* -o jurkat.flnc.chunk.a*.bam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>After downloading, reverse this process including using the header file which is found in the LRPG-Manuscript-Results-results-results-jurkat-isoseq3-companion-files.tar.gz file></p> <p>1. Convert the bam files back to sam files: <em>samtools view jurkat.flnc.chunk.a*.bam > jurkat.flnc.chunk.a*.sam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>2. Combine the header together with the sam files: <em>cat jurkat.flnc.chunk.a*sam > jurkcat.flnc.sam (</em>verified the same number of lines of the sam files is identical to the number of lines of the original without header: 4,956,761. Header file is 13 lines.</p> <p>3. Convert to bam files if desired: <em>samtools view -b jurkat.flnc.sam -o jurkat.flnc.bam</em></p> <p>4. Rehead with the header file: <em>samtools reheader -P -i jurkat.flnc.header.sam jurkat.flnc.bam</em></p>
Long read proteogenomics to characterize protein isoform diversity in human umbilical vein endothelial cells (HUVECs)
<p>Endothelial cells (ECs) comprise the lumenal lining of all blood vessels and are critical for the functioning of the cardiovascular system and their phenotypes can be modulated by protein isoforms. To characterize the isoform landscape within EC, we applied a long read proteogenomics approach to analyze human umbilical vein endothelial cells (HUVECs). Transcripts delineated from PacBio sequencing serve as the basis for a sample-specific protein database used for downstream MS analysis to infer protein isoform expression. We detected 53,836 transcript isoforms from 10,426 genes, with 22,195 of those transcripts being novel. Furthermore, the predominant isoform in HUVECs does not correspond with the accepted “reference isoform” 25% of the time, with vascular pathway-related genes among this group. We found 2,597 protein isoforms supported through unique peptides, with an additional 2,280 isoforms nominated upon incorporation of long-read transcript evidence. We characterized a novel alternative acceptor for endothelial-related gene <em>CDH5</em>, suggesting potential changes in its associated signaling pathways. Finally, we identified novel protein isoforms arising from a diversity of splicing mechanisms supported by uniquely mapped novel peptides. Our results represent a high resolution atlas of known and novel isoforms of potential relevance to endothelial phenotypes and function.</p>
TEST DATA for Enhanced protein isoform characterization through long-read proteogenomics
<p>Test data for The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a></li> <li><a href="http://10.5281/zenodo.5920920">Long-Read-Proteogenomics Workflow Results using Jurkat Sample data</a></li> </ol> <p>This Repository contains the test data, specifically:</p> <p><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></p>
A meta-proteogenomic approach to peptide identification incorporating assembly uncertainty and genomic variation
<p>Supplementary data to "A meta-proteogenomic approach to peptide identification incorporating assembly uncertainty and genomic variation"</p>
The central nervous system's proteogenomic and spatial imprint upon systemic viral infection, like SARS-CoV-2
<p>Data set including image files of histological stainings, immunohistochemistry, MELC, and spatial transcriptomics associated with the study mentioned above.</p>
Proteogenomics_input_files
<p>Proteogenomics tutorial input datasets</p>
Bridging the Chromosome-Centric and Biology and Disease Human Proteome Projects: Accessible and automated tools for interpreting biological and pathological impact of protein sequence variants detected via proteogenomics
<p>Bridging the Chromosome-Centric and Biology and Disease Human Proteome Projects: Accessible and automated tools for interpreting biological and pathological impact of protein sequence variants detected via proteogenomics</p>
Proteogenomic Characterization of Cholangiocarcinoma
<p>supplemental table of Integrated Proteogenomic Characterization of Cholangiocarcinoma Associated with Clinical Outcomes</p>
A massive proteogenomic screen identifies thousands of novel human protein coding sequences
<p>Accurate annotation of genes in the human genome is fundamental for biomedical research and genomic data interpretation. The Ensembl, RefSeq, and GENCODE consortiums continuously update the human genome annotations based on new computational and experimental evidence, and new proteins were identified constantly. The Genotype-Tissue Expression (GTEx) project has generated more than 15,000 RNA sequencing dataset from multiple-tissues of more than 800 donors which allows to model almost all transcripts and proteins in the human genome. Using proteins translated from the GTEx transcript model, more than 21 million in-silico trypsin-digested peptides were generated. To identify high-confidence novel proteins with proteomic support, we screened more than 2,000 proteomic projects in the PRIDE database and selected more than 50,000 mass spectrometry (MS) runs from 923 projects. These MS data were used to validate the predicted novel peptides. With a stringent standard, we identified almost 20,000 novel peptides. </p> <p>This dataset include files used in the the above analysis. More details can be found in the GitHub page (https://github.com/ATPs/human_novo_protein_2022). </p>
Proteogenomic Monitoring and Assessment of Kidney Transplant Recipients
ClinicalTrials.gov study NCT01531257. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Proteogenomic Biomarker Panels in a Serial Blood & Urine Monitoring Study of Kidney Transplant Recipients
ClinicalTrials.gov study NCT01289717. IPD Sharing: Not stated. Countries: 1. Publications: 6.
Gene validation and remodelling using proteogenomics of Phytophthora cinnamomi, the causal agent of Dieback
<p>This spectral data is in support for the manuscript 'Gene validation and remodelling using proteogenomics of Phytophthora cinnamomi, the causal agent of Dieback'. This data was used to detect errors in the draft genome and curate previously undescribed genes in the <em>Phytophthora cinnamomi</em> genome. </p>
Data from: Genome annotation improvements from cross-phyla proteogenomics and time-of-day differences in malaria mosquito proteins using untargeted quantitative proteomics
The malaria mosquito, Anopheles stephensi, and other mosquitoes modulate their biology to match the time-of-day. In the present work, we used a non-hypothesis driven approach (untargeted proteomics) to identify proteins in mosquito tissue, and then quantified the relative abundance of the identified proteins from An. stephensi bodies. Using these quantified protein levels, we then analyzed the data for proteins that were only detectable at certain times-of-the day, highlighting the need to consider time-of-day in experimental design. Further, we extended our time-of-day analysis to look for proteins which cycle in a rhythmic 24-hour ("circadian") manner, identifying 31 rhythmic proteins. Finally, to maximize the utility of our data, we performed a proteogenomic analysis to improve the genome annotation of An. stephensi. We compare peptides that were detected using mass spectrometry but are 'missing' from the An. stephensi predicted proteome, to reference proteomes from 38 other primarily human disease vector species. We found 239 such peptide matches and reveal that genome annotation can be improved using proteogenomic analysis from taxonomically diverse reference proteomes. Examination of 'missing' peptides revealed reading frame errors, errors in gene-calling, overlapping gene models, and suspected gaps in the genome assembly.
Gene validation and remodelling using proteogenomics of Phytophthora cinnamomi, the causal agent of Dieback
Open the record for dataset details and reuse information.
Data from: Genome annotation improvements from cross-phyla proteogenomics and time-of-day differences in malaria mosquito proteins using untargeted quantitative proteomics
Open the record for dataset details and reuse information.
Proteogenomic analysis of psoriasis reveals discordant and concordant changes in mRNA and protein abundance
GEO Series GSE67785. Homo sapiens. 28 samples. Type: Expression profiling by high throughput sequencing.
Spatial proteogenomics reveals distinct and evolutionarily-conserved hepatic macrophage niches (spatial)
GEO Series GSE192741. Homo sapiens; Mus musculus. 15 samples. Type: Expression profiling by high throughput sequencing.
Long read proteogenomics to connect disease-associated sQTLs to the protein isoform effectors in disease
GEO Series GSE224588. Homo sapiens. 11 samples. Type: Expression profiling by high throughput sequencing.
Orthogonal proteogenomic approaches identify the druggable PA2G4-MYC axis in 3q26 AML [scRNA-Seq]
GEO Series GSE256130. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.