Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
35
datasets available to search
ShareScore release 0.7.1
Dataset results
35 results for “Peptide identification”
vPro-MS peptide spectral library for the identification of human-pathogenic viruses by untargeted proteomics
<p>The viral proteomics workflow (vPro-MS) enables identification of human-pathogenic viruses from patient samples by untargeted proteomics. vPro-MS is based on an in-silico derived peptide library covering the human virome in <a href="https://www.uniprot.org/" rel="nofollow">UniProtKB</a> (331 viruses, 20,386 genomes, 121,977 peptides). vPro-MS is intended to identify human-pathogenic viruses from DiaNN (<a href="https://github.com/vdemichev/DiaNN">https://github.com/vdemichev/DiaNN</a>) outputs of either DIA or diaPASEF data. A scoring algorithm (vProID) assesses the confidence of virus identification and the results are finally summarized in a report table. </p> <p>The vPro Peptide Library folder contains 3 peptide FASTA files (Contaminants.fasta, Human.fasta, vPro.Virus.fasta), which were used to predict the spectral library (vPro-lib.predicted.speclib). Please note, that the additional commands “--cut” and “--duplicate-proteins” are needed to reprocess the prediction in DiaNN. This spectral library should be used to identify peptide sequences from samples of human origin using DiaNN. Furthermore, the folder contains the metadata file of the viral peptide sequences (vPro.Peptide.Library.txt) and a summary file of the virus taxonomy covered by the library (Taxonomy.Summary.txt). The metadata file is used by the vPro script to identify viruses from the DiaNN main report.</p>
Identification of significant KLH-derived peptide - Fig. S2
<p>Supplemental figure 2 (S2) of the manuscript “Surface LAMP-2 is an endocytic receptor that diverts antigen internalized by human dendritic cells into highly immunogenic exosomes” by Dario A. Leone et al. (Journal of Immunology). </p> <p>ProPresent® antigen presentation assay from ProImmune was used here to identify potential differences in the presentation of immunogenic regions in Keyhole Limpet Hemocyanin (KLH), compared to KLH conjugated to anti-LAMP2 antibody (H4B4*KLH). The peptide eluted from HLA-DR of MoDC isolated from four different donors were identified using mass spectrometry (LCMSMS)-based analysis in order to identify the putative immunogenic peptides from KLH and two endogenous protein – myeloperoxidase(MPO) and heat shock protein 70 (HSc70). </p>
A meta-proteogenomic approach to peptide identification incorporating assembly uncertainty and genomic variation
<p>Supplementary data to "A meta-proteogenomic approach to peptide identification incorporating assembly uncertainty and genomic variation"</p>
Supporting data for "Quality control for the target decoy approach for peptide identification"
<p>Supporting data for the manuscript "Quality control for the target decoy approach for peptide identification". Input files are raw mass spectrometry runs, to be downloaded from the PRIDE Archive (https://www.ebi.ac.uk/pride/), which can be processed with the parameter files, Nextflow workflow, and Python scripts in "workflow-scripts-parameters.zip". The resulting output files that were used for the manuscript are provided in search-results.zip.</p>
MS data set: Identification of Microorganisms by Liquid Chromatography-Mass Spectrometry (LC-MS1) and in silico Peptide Mass Data
<p>Data set consisting of raw LC-MS2 data, LC-MS1 peak data and a description</p> <p>For unreviewed publication preprint: <strong>Identification of Microorganisms by Liquid Chromatography-Mass Spectrometry (LC-MS<sup>1</sup>) and <em>in silico </em>Peptide Mass Data</strong></p> <p>ABSTRACT</p> <p>Over the past decade, modern methods of mass spectrometry (MS) have emerged that allow reliable, fast and cost-effective identification of pathogenic microorganisms. While MALDI-TOF MS has already revolutionized the way microorganisms are identified, recent years have witnessed also substantial progress in the development of liquid chromatography (LC)-MS based proteomics for microbiological applications. For example, LC-tandem mass spectrometry (LC-MS<sup>2</sup>) has been proposed for microbial characterization by means of multiple discriminative peptides that enable identification at the species, or sometimes at the strain level. However, such investigations can be very time-consuming, especially if the experimental LC-MS<sup>2</sup> data are tested against sequence databases covering a broad panel of different microbiological taxa.</p> <p>In this proof of concept study, we present an alternative bottom-up proteomics method for microbial identification. The proposed approach involves efficient extraction of proteins from cultivated microbial cells, digestion by trypsin and LC-MS measurements. MS<sup>1</sup> data are then extracted and systematically tested against an in silico library of peptide mass data compiled in house. The library has been computed from the UniProt Knowledgebase Swiss-Prot and TrEMBL databases and comprises more than 12,000 strain-specific in silico profiles, each containing tens of thousands of peptide mass entries. Identification analysis involves computation of score values derived from spectral distances between experimental and in silico peptide mass data and compilation of score ranking lists. The taxonomic positions of the microbial samples are then determined by using the best-matching database entries. The suggested method is computationally efficient – less than two minutes per sample - and has been successfully tested by a set of 19 different microbial pathogens. The approach is rapid, accurate and automatable and holds great potential for future microbiological applications.</p> <p><em>For details see the following preprint: Lasch, P. Schneider, A. Blumenscheit, C. and Doellinger, J. “Identification of Microorganisms by Liquid Chromatography-Mass Spectrometry (LC-MS1) and in silico Peptide Mass Data”. bioRxiv preprint, http://dx.doi.org/10.1101/870089</em></p> <p> </p>
NGS data produced in 'Rapid selection and identification of functional CD8+ T-cell epitopes from large peptide-coding libraries'; Nature Communications (2019)
<p>Sharma, G et al. Rapid selection and identification of functional CD8+ T-cell epitopes from large peptide-coding libraries. <em>Nature Communications</em>. Accepted (August 2019)</p> <p><strong>Abstract:</strong></p> <p>Cytotoxic CD8+ T-cells recognize and eliminate infected or malignant cells that present, at their cell surfaces, short peptide epitopes derived from intracellularly processed antigens. However, broadly searching for specific major histocompatibility complex (MHC)-bound peptide epitopes that are naturally processed and capable of eliciting a functional T-cell response has been challenging. Here, we report a method for deep and unbiased T-cell epitope profiling, which is done by using <em>in vitro</em> co-culture of CD8+ T-cells and target cells transduced with high-complexity epitope-encoding minigene libraries. Target cells that are subject to cytotoxic attack from T-cells in co-culture are isolated, before they are lost to apoptosis, by fluorescence-activated cell sorting and characterized by sequencing the minigenes encoded within. In the present study, we validate this highly parallelized method using known murine T-cell receptor/peptide-MHC pairs and diverse minigene-encoded epitope libraries to identify naturally processed and MHC-presented peptide epitopes unambiguously and with high sensitivity.</p>
Protein Identification by Nanopore Peptide Profiling
<p>This dataset belongs to “Protein Identification by Nanopore Peptide Profiling” and describes the raw data and analysis of tryptic digested peptides translocating through a mutant Fragaceatoxin C nanopore. A jupyter notebook describing the analysis and structure is added to this dataset.</p> <p> </p> <p><strong>Data description:</strong></p> <p><strong>Protein Identification by Nanopore Peptide Profiling.ipynb</strong></p> <p> Jupyter notebook contained data analysis of data contained in data_0.zip and data_1.zip (Python 3.7)</p> <p><strong>python_scripts.zip</strong></p> <p> Supplementary scripts belonging to “Protein Identification by Nanopore Peptide Profiling.ipynb”. See explanation of custom classes in the jupyter notebook.</p> <p><strong>data_0.zip</strong> - Folder containing raw electrophysiology data and result after analysis with “Protein Identification by Nanopore Peptide Profiling.ipynb”, with each folder containing the following:</p> <p> Alpha casein: Tryptic digest of alpha casein</p> <p> Beta casein: Tryptic digest of beta casein</p> <p> BSA: Tryptic digest of bovine serum albumin</p> <p> Control: Tryptic digest of water (no protein, control measurement)</p> <p> Cytochrome c: Tryptic digest of cytochrome c</p> <p> DHFR_His6: Tryptic digest of dihydropholate reductase (His6 tagged)</p> <p> EFP: Tryptic digest of elongation factor P</p> <p> HMW1Act: Tryptic digest of high molecular weight adhesin protein</p> <p><strong>data_1.zip</strong> - Folder containing raw electrophysiology data, comma-separated MS peptide masses, and result after analysis with “Protein Identification by Nanopore Peptide Profiling.ipynb”, with each folder containing the following:</p> <p> Lysozyme: Tryptic digest of lysozyme </p> <p> PAN: Tryptic digest of proteasome-activating nucleotidase</p> <p> TbpA_Y27A: Tryptic digest of periplasmic binding protein</p> <p> Trypsin: Tryptic digest of bovine trypsin</p> <p> Mass_spec: csv files containing measured ESI-MS peptides</p> <p> Lysozyme synthetic peptides: Synthetic peptides:</p> <p> Lys1: TPGSR</p> <p> Lys2alk: C(+57.02)ELAAAMK</p> <p> Lys3: HGLDNYR</p> <p> Lys4alk: WWC(+57.02)NDGR</p> <p> Lys5: GTDVQAWIR</p> <p> Lys6alk: GYSLGNWVC(+57.02)AAK</p> <p> Lys7: FESNFNTQATNR</p> <p>The structure of the data files is registered data_1.zip in <strong>'index.csv'</strong> (digested proteins) and <strong>'index</strong><strong>_peptides</strong><strong>.csv'</strong> (synthetic peptides) contained in the data folder. In this file, we describe the protein that was measured as well as the folder location and the expected baseline / standard deviation.<br> <br> <strong>Structure of <em>./data/index.csv</em></strong></p> <p><strong>Protein (string) | Folder (string) | Baseline (pA) (float) | Baseline Error (pA) (float)</strong></p> <p> </p> <p>In each <strong>Folder</strong>, there is another <strong>'index.csv'</strong>, explaining which files are with protein and which are without (blank).<br> <br> <strong>Structure of <em>./data/[protein]/[repeat]/index.csv</em></strong></p> <p><strong>blank (boolean) | fname (string)</strong></p> <p> </p> <p>Each folder in data_0.zip and data_1.zip contains a folder for each measure protein, which contains a folder for each repeat. The repeats contain raw axon binary files (.abf), each file contains measurement conditions as follows:</p> <p> [Date of measurement]_[Pore type]_[Buffer conditions]_[added analyte(s)]_[operator initials]</p> <p> <em>e.g</em>: 20200312_1M_KCl_50mM_Citricacid_50mM_BTP_pH_38_FraC_G13F_neg70mV_20ul_CytC_TrypsinGold_FL_0000</p> <p> Measured on 12-03-2020, in 1M KCl buffered with Citricacid (50 mM) adjusted using bis-tris-propane to pH 3.8, using Fragaceatoxin C mutant G13F at a negatively applied potential of 70 mV. 20 µL cytochrome c was added to the cis compartment.</p> <p>The total volume of the container used for all electrophysiology experiments was 400 µL, all samples were prepared at a 1 g/L concentration. A prefix “perf” before analyte description indicates that the chamber was flushed with approximately 2 mL fresh buffer prior to analysis. The buffer condition "BTP" means bis-tris-propane, which is used to titrate to the exact pH of 3.8.</p> <p>Each analysed folder contains<strong> results.pkl</strong> file, containing the analysis result as provided by “Protein Identification by Nanopore Peptide Profiling.ipynb” - see the jupyter notebook</p> <p>Each analysed folder contains <strong>results_analysis.xlsx</strong>, which contains sheets with excluded currents, standard deviations, dwell time and beta value for the pore without analyte added “Blank” and results from the analyte added in “Results”. Parameters used for fitting are contained in “Parameters”. The “Histograms” tab shows the raw data of the excluded current spectra.</p> <p> </p> <p><strong>mass_spec_peaks.zip</strong> – Folder containing mass spectrometry files as analysed by PEAKS Studio</p> <p>The folder contains an subfolder for each protein measured using electrospray ionisation mass spectrometry (ESI-MS).</p> <p> acasein: alpha casein protein</p> <p> b_casein: beta casein protein</p> <p> BSA: bovine serum albumin</p> <p> CytC: cytochrome C digested</p> <p> DHFR: dihydropholate reductase</p> <p> HMW1_Act: high molecular weight adhesin protein</p> <p> PAN: proteasome-activating nucleotidase</p> <p> ThBP: periplasmic thiamine binding protein</p> <p> Trypsin: bovine trypsin</p>
Supplementarty Data: Retention time and fragmentation predictors increase confidence in identification of common variant peptides
<p>Supplementary data related to the paper "Retention time and fragmentation predictors increase confidence in identification of common variant peptides"</p> <p>The database directory contains the FASTA files of the four protein sequence databases used in the analysis of the paper.</p> <p>The data directory contains exports with lists of all peptide-to-spectrum matches obtained from the analysis.</p> <p>More information and scripts to reproduce the post-processing steps are available at: https://github.com/ProGenNo/VariantPeptideIdentification</p>
Test-time training for deep MS/MS spectrum prediction improves peptide identification
<p>In bottom-up proteomics, peptide-spectrum matching is critical for peptide and protein identification. Recently, deep learning models have been used to predict tandem mass spectra of peptides, with the similarity scores of predicted and experimental spectra being integrated into peptide-spectrum matching. These models follow the supervised learning paradigm, which trains a general model using paired peptides and spectra from standard datasets and uses the model for prediction on experimental data. However, this approach can lead to inaccurate predictions due to differences between the training data and the experimental data, such as sample types, enzyme specificity, and instrument calibration. To address this issue, we proposed a Test-Time Training paradigm that adapts the pre-trained model to experimental data-specific models, namely PepT3. PepT3 results in a 10-40\% increase in peptide identification, depending on the distinctness of training and experimental data. Intriguingly, PepT3 improves the identification of tumor-specific neo-epitopes when applied to a complex patient-derived immunopeptidomic sample, with two-thirds of these neo-epitopes predicted to bind to the patient's human leukocyte antigen isoforms</p>
Identification of reactive Borrelia burgdorferi peptides associated with Lyme disease
Open the record for dataset details and reuse information.
Data for "MS²Rescore 3.0 is a modular, flexible, and user-friendly platform to boost peptide identifications, as showcased with MS Amanda 3.0"
<p>This data set contains reanalyzed data from the "MS2 500 ms" and "MS2 750 ms" workflow <a href="https://doi.org/10.1016/j.mcpro.2022.100219">originally published by Furtwängler et al. 2022</a>. </p><p>The full data set is available on ProteomeXchange via the PRIDE partner repository using ID <a href="https://www.ebi.ac.uk/pride/archive/projects/PXD029320">PXD029320</a>. <br> </p><p>This data set contains, for both workflows </p><ul><li>mzML files</li><li>Output files from analysis with MS Amanda + settings file</li><li>Output files from analysis with MS Amanda in combination with Percolator + settings file</li><li>Percolator pin files</li><li>Output files from analysis with MS²Rescore</li></ul>
Genome sequences of Rhizopogon roseolus, Mariannaea elegans, Myrothecium verrucaria, and Sphaerostilbella broomeana and the identification of biosynthetic gene clusters for fungal peptide natural products
<p>Data to accompany the paper, detailed in manifest.txt</p>
Identification of Potential Biomarkers of Peptide Immunotherapy. Part 2 - Gene Array Analysis
ClinicalTrials.gov study NCT01383603. IPD Sharing: Not stated. Countries: 1. Publications: 1.
64Cu-SAR-BBN and 67CU SAR-BBN for Identification and Treatment of Gastrin Releasing Peptide Receptor (GRPR)-Expressing Metastatic Castrate Resistant Prostate Cancer in Patients Who Are Ineligible for
ClinicalTrials.gov study NCT05633160. IPD Sharing: NO. Countries: 1. Publications: 1.
Identification of Potential Biomarkers of Peptide Immunotherapy. Part 1 - Proteomics Analysis
ClinicalTrials.gov study NCT01383590. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Data from: Identification and mapping of linear antibody epitopes in human serum albumin using high-density peptide arrays
We have recently developed a high-density photolithographic, peptide array technology with a theoretical upper limit of 2 million different peptides per array of 2 cm2. Here, we have used this to perform complete and exhaustive analyses of linear B cell epitopes of a medium sized protein target using human serum albumin (HSA) as an example. All possible overlapping 15-mers from HSA were synthesized and probed with a commercially available polyclonal rabbit anti-HSA antibody preparation. To allow for identification of even the weakest epitopes and at the same time perform a detailed characterization of key residues involved in antibody binding, the array also included complete single substitution scans (i.e. including each of the 20 common amino acids) at each position of each 15-mer peptide. As specificity controls, all possible 15-mer peptides from bovine serum albumin (BSA) and from rabbit serum albumin (RSA) were included as well. The resulting layout contained more than 200.000 peptide fields and could be synthesized in a single array on a microscope slide. More than 20 linear epitope candidates were identified and characterized at high resolution i.e. identifying which amino acids in which positions were needed, or not needed, for antibody interaction. As expected, moderate cross-reaction with some peptides in BSA was identified whereas no cross-reaction was observed with peptides from RSA. We conclude that high-density peptide microarrays are a very powerful methodology to identify and characterize linear antibody epitopes, and should advance detailed description of individual specificities at the single antibody level as well as serologic analysis at the proteome-wide level.
Data from: Identification and mapping of linear antibody epitopes in human serum albumin using high-density peptide arrays
Open the record for dataset details and reuse information.
Mining the genome of Arabidopsis thaliana as a basis for the identification of novel bioactive peptides involved in oxidative stress tolerance
GEO Series GSE49001. Arabidopsis thaliana. 4 samples. Type: Expression profiling by genome tiling array.
sRNA and cis-antisense sRNA identification in Staphylococcus aureus highlights an unusual sRNA gene cluster with one encoding a secreted peptide
GEO Series GSE89487. Staphylococcus aureus. 10 samples. Type: Expression profiling by high throughput sequencing.
Peptide Driven Identification of TCRs (PDI-TCR) reveals dynamics and phenotypes of CD4 T cells in tuberculosis [TCR-Seq]
GEO Series GSE301489. Homo sapiens. 124 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.