Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

75

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

75 results for “Proteogenomics”

Learn how ShareScore rates datasets ↗
zenodo44/100

Development of a spectral library for the discovery of altered genomic events in Mycobacterium avium associated with virulence using mass spectrometry-based proteogenomic analysis

<p><em>Mycobacterium avium</em> is one of the prominent disease-causing bacteria in humans. It causes lymphadenitis, chronic and extrapulmonary, and disseminated infections in adults, children, and immunocompromised patients. <em>M. avium</em> has ~4,500 predicted protein-coding regions on an average, which can be helpful in discovering several variants at the proteome level. Many of them are potentially associated with virulence, thus identifying such proteins can be a helpful feature in the development of panel-based theranostics. In line with such a long-term goal, we carried out an in-depth proteomic analysis of <em>M. avium</em> with both data-dependent and data-independent acquisition methods. Further, a set of proteogenomic investigations were carried out using the protein database for <em>Mycobacterium tuberculosis,</em> and a genome six-frame translated database and a variant protein database of <em>M. avium</em>. A search of mass spectrometry data analysis against <em>M. avium</em> protein database resulted in the identification of 2,954 proteins. Further, proteogenomic analyses aided in the identification of 1,301 novel peptide sequences and correction of translation start sites for 15 proteins. At the end, we created a spectral library of <em>M. avium</em> proteins including novel genome search-specific peptides and variant peptides detected in this study. We validated the spectral library by a data-independent acquisition of the <em>M. avium</em> proteome. Thus, we present a <em>M. avium </em>spectral library of 29,033 peptide precursors supported by 0.4 million fragment ions for further use by the biomedical community.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Enhanced Protein Isoform Characterization Through Long-Read Proteogenomics - Workflow Results

<pre>&nbsp;</pre> <p>The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Long-Read-Proteogenomics Workflow Sample and Reference Data</a></li> <li><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></li> </ol> <p>This Repository contains the complete output from the execution of the&nbsp;<a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow</a>, using the input from&nbsp;<a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a>.&nbsp; &nbsp;</p> <p>The file&nbsp;<em>jurkat.flnc.bam&nbsp;</em>was 6.5 GB had to be split into 13 separate files and for use should be rejoined -- here are the steps that were used to split the file up.&nbsp; &nbsp;</p> <p>1. Convert&nbsp;<em>jurkat.flnc.bam</em>&nbsp;(binary format) to sam file (text format) without header:&nbsp;&nbsp;<em>samtools view jurkat.flnc.bam &gt; jurkat.flnc.sam</em></p> <p>2. Capture the header:&nbsp;<em>samtools view -H jurkat.flnc.bam &gt; jurkat.flnc.header.sam</em></p> <p>3. Split&nbsp;<em>jurkat.flnc.sam</em>&nbsp;into smaller files (aim to get final size under 2GB):&nbsp;<em>split -l 400000 jurkat.flnc.sam jurkat.flnc.chunk.</em></p> <p>4. Convert each of these files back to bam for uploading:&nbsp;<em>samtools view -b jurkat.flnc.chunk.a* -o jurkat.flnc.chunk.a*.bam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>After downloading, reverse this process including using the header file which is found in the&nbsp;LRPG-Manuscript-Results-results-results-jurkat-isoseq3-companion-files.tar.gz file&gt;</p> <p>1. Convert the bam files back to sam files:&nbsp;<em>samtools view jurkat.flnc.chunk.a*.bam &gt; jurkat.flnc.chunk.a*.sam (*=a,b,c,d,e,f,g,h,i,j,k,l,m)</em></p> <p>2. Combine the header together with the sam files:&nbsp;<em>cat jurkat.flnc.chunk.a*sam &gt; jurkcat.flnc.sam (</em>verified the same number of lines of the sam files is identical to the number of lines of the original without header: 4,956,761.&nbsp; Header file is 13 lines.</p> <p>3. Convert to bam files if desired:&nbsp;<em>samtools view -b jurkat.flnc.sam -o jurkat.flnc.bam</em></p> <p>4. Rehead with the header file:&nbsp;<em>samtools reheader -P -i jurkat.flnc.header.sam jurkat.flnc.bam</em></p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Long read proteogenomics to characterize protein isoform diversity in human umbilical vein endothelial cells (HUVECs)

<p>Endothelial cells (ECs) comprise the lumenal lining of all blood vessels and are critical for the functioning of the cardiovascular system and their phenotypes can be modulated by protein isoforms. To characterize the isoform landscape within EC, we applied a long read proteogenomics approach to analyze human umbilical vein endothelial cells (HUVECs). Transcripts delineated from PacBio sequencing serve as the basis for a sample-specific protein database used for downstream MS analysis to infer protein isoform expression. We detected 53,836 transcript isoforms from 10,426 genes, with 22,195 of those transcripts being novel. Furthermore, the predominant isoform in HUVECs does not correspond with the accepted &ldquo;reference isoform&rdquo; 25% of the time, with vascular pathway-related genes among this group. We found 2,597 protein isoforms supported through unique peptides, with an additional 2,280 isoforms nominated upon incorporation of long-read transcript evidence. We characterized a novel alternative acceptor for endothelial-related gene <em>CDH5</em>, suggesting potential changes in its associated signaling pathways. Finally, we identified novel protein isoforms arising from a diversity of splicing mechanisms supported by uniquely mapped novel peptides. Our results represent a high resolution atlas of known and novel isoforms of potential relevance to endothelial phenotypes and function.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

TEST DATA for Enhanced protein isoform characterization through long-read proteogenomics

<p>Test data for&nbsp;The detection of physiologically relevant protein isoforms encoded by the human genome is critical to biomedicine. Mass spectrometry (MS)-based proteomics is the preeminent method for protein detection, but isoform-resolved proteomic analysis relies on accurate reference databases that match the sample; neither a subset nor a superset database is ideal. Long-read RNA sequencing (e.g. PacBio, Oxford Nanopore) provides full-length transcript sequencing, which can be used to predict full-length proteins. Here, we describe a long-read proteogenomics approach for integrating matched long-read RNA-seq and MS-based proteomics data to enhance isoform characterization. We introduce a classification scheme for protein isoforms, discover novel protein isoforms, and present the first protein inference algorithm for the direct incorporation of long-read transcriptome data in protein inference to enable detection of protein isoforms that are intractable to MS detection. We have released an open-source Nextflow pipeline that integrates long-read sequencing in a proteomic workflow for isoform-resolved analysis.</p> <p>Companion Repositories:</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.5920817">Long-Read-Proteogenomics Workflow GitHub Repository Release</a></li> <li><a href="https://doi.org/10.5281/zenodo.5920847">Long-Read-Proteogenomics Analysis GitHub Repository Release</a></li> </ol> <p>Companion Datasets</p> <ol> <li><a href="https://zenodo.org/deposit/5703754">Jurkat Samples and Reference Data</a></li> <li><a href="http://10.5281/zenodo.5920920">Long-Read-Proteogenomics Workflow Results using Jurkat Sample data</a></li> </ol> <p>This Repository contains the test data, specifically:</p> <p><a href="https://doi.org/10.5281/zenodo.5234651">TEST Data for Long-Read-Proteogenomics Workflow GitHub Actions</a></p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

A meta-proteogenomic approach to peptide identification incorporating assembly uncertainty and genomic variation

<p>Supplementary data to &quot;A meta-proteogenomic approach to peptide identification incorporating assembly uncertainty and genomic variation&quot;</p>

opencc-by-4.0May 2019View details →
zenodo40/100

The central nervous system's proteogenomic and spatial imprint upon systemic viral infection, like SARS-CoV-2

<p>Data set including&nbsp;image files of histological stainings, immunohistochemistry, MELC, and spatial transcriptomics associated with the study mentioned above.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Proteogenomics_input_files

<p>Proteogenomics tutorial input datasets</p>

opencc-by-sa-4.0Jun 2018View details →
zenodo36/100

Bridging the Chromosome-Centric and Biology and Disease Human Proteome Projects: Accessible and automated tools for interpreting biological and pathological impact of protein sequence variants detected via proteogenomics

<p>Bridging the Chromosome-Centric and Biology and Disease Human Proteome Projects: Accessible and automated tools for interpreting biological and pathological impact of protein sequence variants detected via proteogenomics</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Proteogenomic Characterization of Cholangiocarcinoma

<p>supplemental table of&nbsp;Integrated Proteogenomic Characterization of Cholangiocarcinoma Associated with Clinical Outcomes</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

A massive proteogenomic screen identifies thousands of novel human protein coding sequences

<p>Accurate annotation of genes in the human genome is fundamental for biomedical research and genomic data interpretation. The Ensembl, RefSeq, and GENCODE consortiums continuously update the human genome annotations based on new computational and experimental evidence, and new proteins were identified constantly. The Genotype-Tissue Expression (GTEx) project has generated more than 15,000 RNA sequencing dataset from multiple-tissues of more than 800 donors which allows to model almost all transcripts and proteins in the human genome. Using proteins translated from the GTEx transcript model, more than 21 million in-silico trypsin-digested peptides were generated. To identify high-confidence novel proteins with proteomic support, we screened more than 2,000 proteomic projects in the PRIDE database and selected more than 50,000 mass spectrometry (MS) runs from 923 projects. These MS data were used to validate the predicted novel peptides. With a stringent standard, we identified almost 20,000 novel peptides.&nbsp;</p> <p>This dataset include files used in the the above analysis. More details can be found in the GitHub page (https://github.com/ATPs/human_novo_protein_2022).&nbsp;</p>

opencc-by-4.0Jul 2022View details →
ClinicalTrials.gov32/100

Proteogenomic Monitoring and Assessment of Kidney Transplant Recipients

ClinicalTrials.gov study NCT01531257. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Proteogenomic Biomarker Panels in a Serial Blood & Urine Monitoring Study of Kidney Transplant Recipients

ClinicalTrials.gov study NCT01289717. IPD Sharing: Not stated. Countries: 1. Publications: 6.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad28/100

Gene validation and remodelling using proteogenomics of Phytophthora cinnamomi, the causal agent of Dieback

<p>This spectral data is in support for the manuscript 'Gene validation and remodelling using proteogenomics of Phytophthora cinnamomi, the causal agent of Dieback'. This data was used to detect errors in the draft genome and curate previously undescribed genes in the <em>Phytophthora cinnamomi</em> genome. </p>

opencc-zeroSep 2020View details →
dryad28/100

Data from: Genome annotation improvements from cross-phyla proteogenomics and time-of-day differences in malaria mosquito proteins using untargeted quantitative proteomics

The malaria mosquito, Anopheles stephensi, and other mosquitoes modulate their biology to match the time-of-day. In the present work, we used a non-hypothesis driven approach (untargeted proteomics) to identify proteins in mosquito tissue, and then quantified the relative abundance of the identified proteins from An. stephensi bodies. Using these quantified protein levels, we then analyzed the data for proteins that were only detectable at certain times-of-the day, highlighting the need to consider time-of-day in experimental design. Further, we extended our time-of-day analysis to look for proteins which cycle in a rhythmic 24-hour ("circadian") manner, identifying 31 rhythmic proteins. Finally, to maximize the utility of our data, we performed a proteogenomic analysis to improve the genome annotation of An. stephensi. We compare peptides that were detected using mass spectrometry but are 'missing' from the An. stephensi predicted proteome, to reference proteomes from 38 other primarily human disease vector species. We found 239 such peptide matches and reveal that genome annotation can be improved using proteogenomic analysis from taxonomically diverse reference proteomes. Examination of 'missing' peptides revealed reading frame errors, errors in gene-calling, overlapping gene models, and suspected gaps in the genome assembly.

opencc-zeroAug 2019View details →
dryad28/100

Gene validation and remodelling using proteogenomics of Phytophthora cinnamomi, the causal agent of Dieback

Open the record for dataset details and reuse information.

publicSep 2020View details →
dryad28/100

Data from: Genome annotation improvements from cross-phyla proteogenomics and time-of-day differences in malaria mosquito proteins using untargeted quantitative proteomics

Open the record for dataset details and reuse information.

publicAug 2019View details →
geo24/100

Proteogenomic analysis of psoriasis reveals discordant and concordant changes in mRNA and protein abundance

GEO Series GSE67785. Homo sapiens. 28 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenAug 2015View details →
geo24/100

Spatial proteogenomics reveals distinct and evolutionarily-conserved hepatic macrophage niches (spatial)

GEO Series GSE192741. Homo sapiens; Mus musculus. 15 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2022View details →
geo24/100

Long read proteogenomics to connect disease-associated sQTLs to the protein isoform effectors in disease

GEO Series GSE224588. Homo sapiens. 11 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMar 2023View details →
geo24/100

Orthogonal proteogenomic approaches identify the druggable PA2G4-MYC axis in 3q26 AML [scRNA-Seq]

GEO Series GSE256130. Homo sapiens. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenApr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record