Skip to main content
zenodoopen

Data for "A learned score function improves the power of mass spectrometry database search"

<div> <h1>DATA for "A learned score function improves the power of mass spectrometry database search"</h1> <br> <div>These data files are associated with the following publication:</div> <br> <div> <ul> <li>Varun Ananth, Justin Sanders, Melih Yilmaz, Sewoong Oh and William Stafford Noble. "<a title="biorXiv Preprint Link" href="https://www.biorxiv.org/content/10.1101/2024.01.26.577425v2" target="_blank" rel="noopener">A learned score function improves the power of mass spectrometry database search</a>". Bioinformatics (Proceedings of the ISMB). &nbsp;2024.</li> </ul> </div> <br> <div>For the benchmarking data, we used a dataset that is publicly available on ProteomeXchange (PXD028735). The paper that introduced this dataset is:</div> <br> <div> <ul> <li>Van Puyvelde, B., Daled, S., Willems, S., Gabriels, R., Gonzalez de Peredo, A., Chaoui, K., Mouton-Barbosa, E., Bouyssi&eacute;, D., Boonen, K., Hughes, C. J., Gethings, L. A., Perez-Riverol, Y., Bloomfield, N., Tate, S., Schiltz, O., Martens, L., Deforce, D., &amp; Dhaenens, M. (2022). A comprehensive LFQ benchmark dataset on modern day acquisition strategies in proteomics. In Scientific Data (Vol. 9, Issue 1). Springer Science and Business Media LLC. https://doi.org/10.1038/s41597-022-01216-6</li> </ul> </div> <br> <div>More specifically, the following `.raw` files were downloaded:</div> <br> <ul> <li><code>LFQ_Orbitrap_DDA_Ecoli_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Human_01.raw</code></li> <li><code>LFQ_Orbitrap_DDA_Yeast_01.raw</code></li> </ul> <br> <div>Those files can be accessed via FTP&nbsp;<a title="Link to ProteomeXchange: PXD028735" href="https://ftp.pride.ebi.ac.uk/pride/data/archive/2022/02/PXD028735/" target="_blank" rel="noopener">here</a>.</div> <br> <div>We upload here the annotated <code>.mgf</code> files created from these <code>.raw</code> files, as described in our paper.</div> <br> <div>The human, yeast, and E. coli .fasta files used in all database searches were downloaded from UniProt on 11/6/23, 4:30 PM.</div> <br> <div> <ul> <li>Bateman, A., Martin, M.-J., Orchard, S., Magrane, M., Ahmad, S., Alpi, E., Bowler-Barnett, E. H., Britto, R., Bye-A-Jee, H., Cukura, A., Denny, P., Dogan, T., Ebenezer, T., Fan, J., Garmiri, P., da Costa Gonzales, L. J., Hatton-Ellis, E., Hussein, A., &hellip; Zhang, J. (2022). UniProt: the Universal Protein Knowledgebase in 2023. In Nucleic Acids Research (Vol. 51, Issue D1, pp. D523&ndash;D531). Oxford University Press (OUP). https://doi.org/10.1093/nar/gkac1052</li> </ul> </div> <br> <div>We include these files here, with only minor modifications to replace `U` amino acids with `X` so that all amino acids fall into Casanovo-DB's vocabulary.</div> </div>

ShareScore

40/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
12
Harmonization
4
Access
16
Reuse readiness
0
Engagement
8

Topics