Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “Open modification search”
Enhancing Open Modification Searches via a Combined Approach Facilitated by Ursgal
<p>The identification of peptide sequences and their post-translational modifications (PTMs) is a crucial step in the analysis of bottom-up proteomics data. The recent development of open modification search (OMS) engines allows virtually all PTMs to be searched for. This not only increases the number of spectra that can be matched to peptides but also greatly advances the understanding of biological roles of PTMs through the identification, and thereby facilitated quantification, of peptidoforms (peptide sequences and their potential PTMs). While the benefits of combining results from multiple protein database search engines has been established previously, similar approaches for OMS results are missing so far. Here, we compare and combine results from three different OMS engines, demonstrating an increase in peptide spectrum matches of 8-18%. The unification of search results furthermore allows for the combined downstream processing of search results, including the mapping to potential PTMs. Finally, we test for the ability of OMS engines to identify glycosylated peptides. The implementation of these engines in the Python framework Ursgal facilitates the straightforward application of OMS with unified parameters and results files, thereby enabling yet unmatched high-throughput, large-scale data analysis.</p> <p>This dataset includes all relevant results files, databases, and scripts that correspond to the accompanying journal article. Specifically, the following files are deposited:</p> <ul> <li>Homo_sapiens_PXD004452_results.zip: result files from OMS and CS for the dataset PXD004452</li> <li>Homo_sapiens_PXD013715_results.zip: result files from OMS and CS for the dataset PXD013715</li> <li>Haloferax_volcanii_PXD021874_results.zip: result files from OMS and CS for the dataset PXD021874</li> <li>Escherichia_coli_PXD000498_results.zip: result files from OMS and CS for the dataset PXD000498</li> <li>databases.zip: target-decoy databases for <em>Homo sapiens</em>, <em>Escherichia coli </em>and <em>Haloferax volcanii</em> as well as a glycan database for <em>Homo sapiens</em></li> <li>scripts.zip: example scripts for all relevant steps of the analysis</li> <li>mzml_files.zip: mzML files for all included datasets</li> <li>ursgal.zip: current version of Ursgal (0.6.7) that has been used to generate the results (for most recent versions see https://github.com/ursgal/ursgal)</li> </ul>
Semi-supervised learning for sensitive open modification spectral library searching - dataset PXD009476
<p>This is the analysis results of ANN-SoLo + Rescoring as an integrated module. ANN-SoLo spectral library search engine is a tool for efficient open modification searching. ANN-SoLo uses a cascade search strategy to optimally identify both unmodified and modified peptides: in the first stage a standard search is performed to identify unmodified peptides, after which the remaining unidentified spectra are submitted to second stage during which an open search is performed to additionally identify modified peptides. We have augmented this approach by natively integrating PSM rescoring into ANN-SoLo using the mokapot Python framework for semi-supervised learning for peptide detection.</p> <p>The repository includes the results and search database used to analyze a human glycoproteomics dataset that was acquired from human kidney tissue, serum, and T cells to study O-linked glycosylation. Samples were trypsin-digested and separated into 24 fractions after enrichment of intact glycopeptides and release of O-linked glycopeptides. Next, the data was acquired on a Fusion Lumos mass spectrometer with an Easy-nLC 1200 system or a Q-Exactive HF mass spectrometer with a Waters NanoAcquity UPLC. From this dataset, four raw files derived from kidney tissue samples were retrieved from PRIDE (project PXD009476) and converted to MGF files using ThermoRawFileParser (version 1.7.2).</p> <p>The four files are available both in RAW and MGF format below. </p> <p> </p> <ul> <li>Arab, Issar, William E. Fondrie, Kris Laukens, and Wout Bittremieux. "<strong>Semisupervised Machine Learning for Sensitive Open Modification Spectral Library Searching</strong>." Journal of Proteome Research 22, no. 2 (2023): 585-593. <a title="DOI URL" href="https://doi.org/10.1021/acs.jproteome.2c00616">doi.org/10.1021/acs.jproteome.2c00616</a></li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.