Spike-in and real-world proteomics data sets used in publication of PRONE
<p>Spike-in and real-world data sets used in the evaluation study by Arend et al. (see reference), and some utilized in the vignettes of PRONE, an R package designed for preprocessing, normalization, and performance evaluation of normalization methods of proteomics data. </p> <h3>Overview of the Data Sets</h3> <p>Due to the unavailability of proteomics quantification data from all original publications, data were extracted from alternative sources, which are also listed in the table below. Please refer to the paper's supplementary material and GitHub repository (https://github.com/lisiarend/PRONE.Evaluation) for more comprehensive information on the data sets. A processed metadata file and the protein quantification data file are provided for all data sets. The original quantification data used to generate these two files for each data set are consistently provided in the `original_data` directory of each data set.</p> <p> </p> <div> <table> <tbody> <tr> <td> <p>Data Set</p> </td> <td> <p>Type</p> </td> <td> <p>Quantification Type</p> </td> <td> <p>Raw Data (ID)</p> </td> <td> <p>Quantification Data</p> </td> </tr> <tr> <td> <p>dS1</p> </td> <td> <p>UPS1 spike-in </p> <p>(4 levels)</p> </td> <td> <p>LFQ</p> </td> <td> <p>Tabb et al. <a href="https://www.zotero.org/google-docs/?gkmf6j">[1]</a></p> </td> <td> <p>Välikangas et al. <a href="https://www.zotero.org/google-docs/?Tt81W8">[10]</a> </p> </td> </tr> <tr> <td> <p>dS2</p> </td> <td> <p>UPS1 spike-in </p> <p>(6 levels)</p> </td> <td> <p>LFQ</p> </td> <td> <p>Ramus et al. <a href="https://www.zotero.org/google-docs/?HP8b9X">[2]</a> (PXD001819)</p> </td> <td> <p>Graw et al. <a href="https://www.zotero.org/google-docs/?aiEVfL">[11]</a></p> </td> </tr> <tr> <td> <p>dS3</p> </td> <td> <p>E.coli spike-in </p> <p>(5 levels)</p> </td> <td> <p>LFQ</p> </td> <td> <p>Shen et al. <a href="https://www.zotero.org/google-docs/?zHIDZy">[3]</a></p> <p>(PXD003881)</p> </td> <td> <p>Sticker et al. <a href="https://www.zotero.org/google-docs/?aZ7oZu">[12]</a></p> </td> </tr> <tr> <td> <p>dS4</p> </td> <td> <p>E.coli spike-in </p> <p>(2 levels)</p> </td> <td> <p>LFQ</p> </td> <td> <p>Cox et al. <a href="https://www.zotero.org/google-docs/?scM19M">[4]</a></p> <p>(PXD00279)</p> </td> <td> </td> </tr> <tr> <td> <p>dS5</p> </td> <td> <p>E.coli spike-in </p> <p>(3 levels)</p> </td> <td> <p>TMT 10-plex (1)</p> </td> <td> <p>Zhu et al. <a href="https://www.zotero.org/google-docs/?qAz08E">[5]</a></p> <p>(PXD013277)</p> </td> <td> <p>Phil Wilmarth <a href="https://www.zotero.org/google-docs/?i7xmgP">[13]</a></p> </td> </tr> <tr> <td> <p>dS6</p> </td> <td> <p>yeast spike-in </p> <p>(3 levels)</p> </td> <td> <p>TMT 11-plex (1)</p> </td> <td> <p>O’Connell et al. <a href="https://www.zotero.org/google-docs/?C2XoaJ">[6]</a></p> <p>(PXD007683)</p> </td> <td> <p>Ammar et al. <a href="https://www.zotero.org/google-docs/?UWcyly">[14]</a></p> </td> </tr> <tr> <td> <p>dR1</p> </td> <td> <p>Osteogenic differentiation of hPCLSCs (4 time points)</p> </td> <td> <p>TMT 6-plex (3)</p> </td> <td> <p>Li et al. <a href="https://www.zotero.org/google-docs/?BouGSs">[8]</a> (PXD020908)</p> </td> <td> <p>MaxQuant executed in-house</p> </td> </tr> <tr> <td> <p>dR2</p> </td> <td> <p>Prospective Ovarian JHU Proteome</p> </td> <td> <p>TMT 10-plex (13)</p> </td> <td> <p>Hu et al. <a href="https://www.zotero.org/google-docs/?kVn4pf">[9]</a></p> <p>(PDC000110)</p> </td> <td> <p>MaxQuant executed in-house</p> </td> </tr> <tr> <td> <p>dR3</p> </td> <td> <p>AROM+ transgenic vs. wild-type mice</p> </td> <td> <p>LFQ</p> </td> <td> <p>Vehmas et al. <a href="https://www.zotero.org/google-docs/?qwYM2A">[7]</a></p> <p>(PXD002025)</p> </td> <td> </td> </tr> <tr> <td> <p>dR4</p> </td> <td> <p>Mycobacterium tuberculosis</p> <p>(healthy, disease vs. treated)</p> </td> <td> <p>TMT 10-plex (2)</p> </td> <td> <p>Schmidt et al. <a href="https://www.zotero.org/google-docs/?m1VHIG">[15]</a></p> <p>(PXD030883)</p> </td> <td> </td> </tr> </tbody> </table> </div> <h3>References</h3> <p>[1] D. L. Tabb et al., ‘Repeatability and Reproducibility in Proteomic Identifications by Liquid Chromatography−Tandem Mass Spectrometry’, J. Proteome Res., vol. 9, no. 2, pp. 761–776, Feb. 2010, doi: 10.1021/pr9006365.<br>[2] C. Ramus et al., ‘Spiked proteomic standard dataset for testing label-free quantitative software and statistical methods’, Data Brief, vol. 6, pp. 286–294, Mar. 2016, doi: 10.1016/j.dib.2015.11.063.<br>[3] X. Shen et al., ‘IonStar enables high-precision, low-missing-data proteomics quantification in large biological cohorts’, Proc. Natl. Acad. Sci., vol. 115, no. 21, pp. E4767–E4776, May 2018, doi: 10.1073/pnas.1800541115.<br>[4] J. Cox, M. Y. Hein, C. A. Luber, I. Paron, N. Nagaraj, and M. Mann, ‘Accurate Proteome-wide Label-free Quantification by Delayed Normalization and Maximal Peptide Ratio Extraction, Termed MaxLFQ *’, Mol. Cell. Proteomics, vol. 13, no. 9, pp. 2513–2526, Sep. 2014, doi: 10.1074/mcp.M113.031591.<br>[5] Y. Zhu et al., ‘DEqMS: A Method for Accurate Variance Estimation in Differential Protein Expression Analysis *’, Mol. Cell. Proteomics, vol. 19, no. 6, pp. 1047–1057, Jun. 2020, doi: 10.1074/mcp.TIR119.001646.<br>[6] J. D. O’Connell, J. A. Paulo, J. J. O’Brien, and S. P. Gygi, ‘Proteome-Wide Evaluation of Two Common Protein Quantification Methods’, J. Proteome Res., vol. 17, no. 5, pp. 1934–1942, May 2018, doi: 10.1021/acs.jproteome.8b00016.<br>[7] A. P. Vehmas et al., ‘Liver lipid metabolism is altered by increased circulating estrogen to androgen ratio in male mouse’, J. Proteomics, vol. 133, pp. 66–75, Feb. 2016, doi: 10.1016/j.jprot.2015.12.009.<br>[8] J. Li et al., ‘Dynamic proteomic profiling of human periodontal ligament stem cells during osteogenic differentiation’, Stem Cell Res. Ther., vol. 12, no. 1, p. 98, Feb. 2021, doi: 10.1186/s13287-020-02123-6.<br>[9] Y. Hu et al., ‘Integrated Proteomic and Glycoproteomic Characterization of Human High-Grade Serous Ovarian Carcinoma’, Cell Rep., vol. 33, no. 3, p. 108276, Oct. 2020, doi: 10.1016/j.celrep.2020.108276.<br>[10] T. Välikangas, T. Suomi, and L. L. Elo, ‘A systematic evaluation of normalization methods in quantitative label-free proteomics’, Brief. Bioinform., vol. 19, no. 1, pp. 1–11, Jan. 2018, doi: 10.1093/bib/bbw095.<br>[11] S. Graw et al., ‘proteiNorm – A User-Friendly Tool for Normalization and Analysis of TMT and Label-Free Protein Quantification’, ACS Omega, vol. 5, no. 40, pp. 25625–25633, Oct. 2020, doi: 10.1021/acsomega.0c02564.<br>[12] A. Sticker, L. Goeminne, L. Martens, and L. Clement, ‘Robust Summarization and Inference in Proteome-wide Label-free Quantification’, Mol. Cell. Proteomics, vol. 19, no. 7, pp. 1209–1219, Jul. 2020, doi: 10.1074/mcp.RA119.001624.<br>[13] ‘understanding_IRS’. Accessed: Mar. 07, 2024. [Online]. Available: https://pwilmart.github.io/IRS_normalization/understanding_IRS.html<br>[14] C. Ammar, M. Gruber, G. Csaba, and R. Zimmer, ‘MS-EmpiRe Utilizes Peptide-level Noise Distributions for Ultra-sensitive Detection of Differentially Expressed Proteins[S]’, Mol. Cell. Proteomics, vol. 18, no. 9, pp. 1880–1892, Sep. 2019, doi: 10.1074/mcp.RA119.001509.<br>[15] F. Biadglegne et al., ‘Mycobacterium tuberculosis Affects Protein and Lipid Content of Circulating Exosomes in Infected Patients Depending on Tuberculosis Disease State’, Biomedicines, vol. 10, no. 4, p. 783, Mar. 2022, doi: 10.3390/biomedicines10040783.</p>
ShareScore
24/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 0