Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “mass spectral library”
Liquid Chromatography - Tandem Mass Spectrometry (LC-MS/MS) and Gas Chromatography - Mass Spectrometry (GC-MS) Reference Libraries from Global Natural Products Social Molecular Networking (GNPS) and National Institute of Standards and Technology (NIST) WebBook Processed for Spectral Library Matching
<div>In order to obtain a high-quality LC-MS/MS reference database for spectral library matching, we selected 22 high-quality GNPS tandem mass spectrometry databases generated under the positive ion mode. Further preprocessing similar to Huber et al involving mass-to-charge (m/z) and intensity filtering yields the database found in the file LCMS_GNPS_reference_library.csv which contains 14,705 electrospray ionization (ESI) mass spectra, each of which corresponds to a unique compound. The NIST WebBook database was used to construct GC-MS database contained in the file GCMS_NIST_WebBook.csv. This database contains 23,721 electron ionization (EI) mass spectra, each of which corresponds to a unique non-hyphenated Chemical Abstract Service (CAS) Registry Number.</div> <div> </div> <div>Both LC-MS/MS and GC-MS databases are organized into three columns: one for the identifier, one for the m/z values, and one for the intensity values. For example, if spectrum A has 20 ion fragments, then there will be 20 rows corresponding to spectrum A in the corresponding database with the identifier A repeated 20 times with the corresponding m/z and intensity values.</div>
Development of a spectral library for the discovery of altered genomic events in Mycobacterium avium associated with virulence using mass spectrometry-based proteogenomic analysis
<p><em>Mycobacterium avium</em> is one of the prominent disease-causing bacteria in humans. It causes lymphadenitis, chronic and extrapulmonary, and disseminated infections in adults, children, and immunocompromised patients. <em>M. avium</em> has ~4,500 predicted protein-coding regions on an average, which can be helpful in discovering several variants at the proteome level. Many of them are potentially associated with virulence, thus identifying such proteins can be a helpful feature in the development of panel-based theranostics. In line with such a long-term goal, we carried out an in-depth proteomic analysis of <em>M. avium</em> with both data-dependent and data-independent acquisition methods. Further, a set of proteogenomic investigations were carried out using the protein database for <em>Mycobacterium tuberculosis,</em> and a genome six-frame translated database and a variant protein database of <em>M. avium</em>. A search of mass spectrometry data analysis against <em>M. avium</em> protein database resulted in the identification of 2,954 proteins. Further, proteogenomic analyses aided in the identification of 1,301 novel peptide sequences and correction of translation start sites for 15 proteins. At the end, we created a spectral library of <em>M. avium</em> proteins including novel genome search-specific peptides and variant peptides detected in this study. We validated the spectral library by a data-independent acquisition of the <em>M. avium</em> proteome. Thus, we present a <em>M. avium </em>spectral library of 29,033 peptide precursors supported by 0.4 million fragment ions for further use by the biomedical community.</p>
Training dataset: Generation of a spectral library from HEK-Ecoli Spike-in mass spectrometry data
<p>The five raw files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>HEK and E.coli cell pellets were lysed with 5 % SDS, 50 mM triethylammonium bicarbonate (TEAB), pH 7.55. The obtained protein extracts were reduced by adding f.c. 5 mM TCEP and alkylated by the addition of f.c. 10 mM iodacetamide. Protein digestion and purification was performed on S-Trap columns. To ensure protein binding to the S-Trap columns, samples were acidified to a final concentration of 1.2 % phosphoric acid (~ pH 2). Six times the sample volume S-Trap buffer (90% aqueous methanol containing a final concentration of 100 mM TEAB, pH 7.1) was added to the samples which were then loaded on the columns and washed with S-Trap buffer. Protein digestion was performed with trypsin and LysC for one hour at 47 °C. Peptides were eluted in three steps with (1) 50 mM TEAB, (2) 0.2 % aqueous formic acid and (3) 50 % acetonitrile containing 0.2 % formic acid. Eluted peptides of HEK and E.coli were mixed in the following ratios (amount in µg):</p> <p>Sample HEK E.coli MS method<br> Sample1 2.5 0.00 DDA<br> Sample2 2.5 0.05 DDA<br> Sample3 2.5 0.15 DDA<br> Sample4 2.5 0.40 DDA<br> Sample5 2.5 0.80 DDA</p> <p>Additionally, iRT peptides were added and 1µg of each samples was measured with a Q-Exactive Plus mass spectrometer. Besides the five raw files, we uploaded two fasta files that serve as human and ecoli protein sequence databases, an transition list for the iRT peptides as well as an experimental design for the MaxQuant search.<br> Additionally, we uploaded the Galaxy MaxQuant training result files: protein groups, peptides, mqpar, msms, evidence and PTXQC.</p>
An 8-(Diazomethyl) Quinoline Derivatized Acyl-CoA In Silico Mass Spectral Library Reveals the Landscape of Acyl-CoA in the Aging Mouse Organs (Data Supplement)
<p>Data supplement for publication "An 8-(Diazomethyl) Quinoline Derivatized Acyl-CoA In Silico Mass Spectral Library Reveals the Landscape of Acyl-CoA in the Aging Mouse Organs (Data Supplement)" (2024)</p> <p>Jinhui Yu<sup>1†</sup>, Menghao Guo<sup>1,3†</sup>, Sha Li<sup>5</sup>, Jian Ni<sup>2</sup>, Yu-Qi Feng<sup>4,5</sup>*, Jun Ding<sup>1,2</sup>*</p> <p>1. CAS Key Laboratory of Plant Germplasm Enhancement and Specialty Agriculture, Wuhan Botanical Garden, Chinese Academy of Sciences, Wuhan, 430074, PR China.</p> <p>2. Renmin Hospital of Wuhan University, Wuhan University, 430072, Wuhan, P. R. China.</p> <p>3. College of Life Sciences, Wuhan University, Wuhan 430072, China.</p> <p>4. School of Bioengineering and Health, Wuhan Textile University, Wuhan 430200, China.</p> <p>5. Frontier Science Center for Immunology and Metabolism, Wuhan University, Wuhan, 430071, China.</p> <p> </p> <p>†The authors contribute equally to this work.</p> <p>* Corresponding author Email: <a href="mailto:dingjun@wbgcas.cn">dingjun@wbgcas.cn</a>, <a href="mailto:yqfeng@whu.edu.cn">yqfeng@whu.edu.cn</a></p> <p>----<br>Content:<br>1) Developement: XLS template sheet for development (can be used to adjust or create new library)<br>2) MSP library: spectra in NIST MSP format (use with MS-Dial or NIST MS Search)<br>3) NIST library: NIST23 compatible 8-DMQ-acyl-CoA library (use with NIST MS Search)<br>4) Reference spectra msp: 8-DMQ-acyl-CoA authentic MS/MS spectrum in NIST MSP format (generated by Thermo QE HF-X MS (HCD), for searching NIST MS-Search)</p> <p>Version 1.0<br>April 22 2024</p>
Data for "On optimizing mass spectral library search and spectral alignment by weighting low-intensity peaks and m/z frequency"
<p>Data for <a href="https://github.com/enveda/weighting-spectral-similarity#on-optimizing-mass-spectral-library-search-and-spectral-alignment-by-weighting-low-intensity-peaks-and-mz-frequency">On optimizing mass spectral library search and spectral alignment by weighting low-intensity peaks and m/z frequency</a>.</p>
MsnLib multi-stage fragmentation mass spectral libraries - mzml positive and negative
<p>The data for <a href="https://doi.org/10.26434/chemrxiv-2024-l1tqh-v2">MSnLib</a> are divided into several Zenodo records due to size constraints. </p> <p>raw positive: <a href="https://doi.org/10.5281/zenodo.10966404">10966404</a><br>raw negative: <a href="https://doi.org/10.5281/zenodo.10967081">10967081</a><br>mzml positive and negative: <a href="https://doi.org/10.5281/zenodo.10966280">10966280</a><br>spectral libraries: <a href="https://doi.org/10.5281/zenodo.11163380">11163380</a></p> <p>This record includes zipped files containing the converted mzml positive and negative data, acquired using a flow injection method on an Orbitrap ID-X instrument, for all compound libraries. The different polartities are kept separated. </p> <p>9 Compound Libraries:</p> <ul> <li>Short Name: Full name, Provider (Catalog number), total compounds (not all detected during library building)</li> <li>MCEBIO: Bioactive Compound Library, MedChemExpress (HY-L001), 10,315 compounds</li> <li>MCESAF: 5k Scaffold Library, MedChemExpress, (HY-L902), 4998 compounds</li> <li>NIHNP: NIH NPAC ACONN collection of NP, NIH/NCATS, 3988 compounds</li> <li>OTAVAPEP: Alpha-helix Peptiomimetic Library, OTAVAchemicals (a-helix-Peptido), 1298 compounds</li> <li>ENAMDISC: Discovery Diversity Set -10, Enamine (DDS-10), 10,240 compounds</li> <li>ENAMMOL: Carboxylic Acid Fragment Library + Random, Enamine and Molport, 4378 compounds</li> <li>MCEDRUG: FDA-Approved Drug Library, MedChemExpress (HY-L022), 2610 compounds</li> <li>MCEDIV_50k_Sub: Subset of 50K Diversity Library, MedChemExpress (HY-L901), 20000 compounds</li> <li>TargetMolNPHTS: Subset of Natural Product Library for HTS, TargetMol (L6000), 2175 compounds</li> </ul>
Matchms cleaned mass spectral library positive mode
<p>The library consists of the GNPS library, MassBank, MoNA and Brungs et al. 's library. The library was cleaned using matchms filtering. The settings for filtering and logging can be found here as well. </p>
Machms cleaned mass spectral library negative ionisation mode
<p>The library consists of the GNPS library, MassBank, MoNA and Brungs et al. 's library. The library was cleaned using matchms filtering. This library contains all negative ionisation mode ms2 spectra. The settings for filtering and logging can be found here as well. </p>
MsnLib multi-stage fragmentation mass spectral libraries - raw negative
<p>The data for <a href="https://doi.org/10.26434/chemrxiv-2024-l1tqh-v2">MSnLib</a> are divided into several Zenodo records due to size constraints. </p> <p>raw positive: <a href="https://doi.org/10.5281/zenodo.10966404">10966404</a><br>raw negative: <a href="https://doi.org/10.5281/zenodo.10967081">10967081</a><br>mzml positive and negative: <a href="https://doi.org/10.5281/zenodo.10966280">10966280</a><br>spectral libraries: <a href="https://doi.org/10.5281/zenodo.11163380">11163380</a></p> <p>This record includes zipped files containing the raw negative data, acquired using a flow injection method on an Orbitrap ID-X instrument, for all compound libraries.</p> <p>7 Compound Libraries:</p> <ul> <li>Short Name: Full name, Provider (Catalog number), total compounds (not all detected during library building)</li> <li>mce_bioactive: Bioactive Compound Library, MedChemExpress (HY-L001), 10,315 compounds</li> <li>mce_scaffold: 5k Scaffold Library, MedChemExpress, (HY-L902), 4998 compounds</li> <li>nih: NIH NPAC ACONN collection of NP, NIH/NCATS, 3988 compounds</li> <li>otavapep: Alpha-helix Peptiomimetic Library, OTAVAchemicals (a-helix-Peptido), 1298 compounds</li> <li>enamdisc: Discovery Diversity Set -10, Enamine (DDS-10), 10,240 compounds</li> <li>enammol: Carboxylic Acid Fragment Library + Random, Enamine and Molport, 4378 compounds</li> <li>mcedrug: FDA-Approved Drug Library, MedChemExpress (HY-L022), 2610 compounds</li> </ul>
MsnLib multi-stage fragmentation mass spectral libraries - raw positive
<p>The data for <a href="https://doi.org/10.26434/chemrxiv-2024-l1tqh-v2">MSnLib</a> are divided into several Zenodo records due to size constraints. </p> <p>raw positive: <a href="https://doi.org/10.5281/zenodo.10966404">10966404</a><br>raw negative: <a href="https://doi.org/10.5281/zenodo.10967081">10967081</a><br>mzml positive and negative: <a href="https://doi.org/10.5281/zenodo.10966280">10966280</a><br>spectral libraries: <a href="https://doi.org/10.5281/zenodo.11163380">11163380</a></p> <p>This record includes zipped files containing the raw positive data, acquired using a flow injection method on an Orbitrap ID-X instrument, for all compound libraries.</p> <p>7 Compound Libraries:</p> <ul> <li>Short Name: Full name, Provider (Catalog number), total compounds (not all detected during library building)</li> <li>mce_bioactive: Bioactive Compound Library, MedChemExpress (HY-L001), 10,315 compounds</li> <li>mce_scaffold: 5k Scaffold Library, MedChemExpress, (HY-L902), 4998 compounds</li> <li>nih: NIH NPAC ACONN collection of NP, NIH/NCATS, 3988 compounds</li> <li>otavapep: Alpha-helix Peptiomimetic Library, OTAVAchemicals (a-helix-Peptido), 1298 compounds</li> <li>enamdisc: Discovery Diversity Set -10, Enamine (DDS-10), 10,240 compounds</li> <li>enammol: Carboxylic Acid Fragment Library + Random, Enamine and Molport, 4378 compounds</li> <li>mcedrug: FDA-Approved Drug Library, MedChemExpress (HY-L022), 2610 compounds</li> </ul>
MSnLib Mass spectral libraries (.mgf and .json)
<p>The data for <a href="https://doi.org/10.26434/chemrxiv-2024-l1tqh-v2">MSnLib</a> are divided into several Zenodo records due to size constraints. </p> <p>raw positive: <a href="https://doi.org/10.5281/zenodo.10966404">10966404</a><br>raw negative: <a href="https://doi.org/10.5281/zenodo.10967081">10967081</a><br>mzml positive and negative: <a href="https://doi.org/10.5281/zenodo.10966280">10966280</a><br>spectral libraries: <a href="https://doi.org/10.5281/zenodo.11163380">11163380</a></p> <p>This record includes the automatically generated spectral libraries (MSnLib) within mzmine, acquired using a flow injection method on an Orbitrap ID-X instrument, for all compound libraries. There are multiple files for each compound library containing MS2 only or MSn in two data formats (.mgf or .json) for both polarities. </p> <p>MS2 contains next to all MS2 spectra all pseudo MS2 spectra (a full MSn tree merged into one spectrum per compound ion). MSn contains all individual MSn stages additionally. The first number for each file highlights the library building date.</p> <p>9 Compound Libraries:</p> <ul> <li>Short Name: Full name, Provider (Catalog number), total compounds (not all detected during library building)</li> <li>MCEBIO: Bioactive Compound Library, MedChemExpress (HY-L001), 10,315 compounds</li> <li>MCESAF: 5k Scaffold Library, MedChemExpress, (HY-L902), 4998 compounds</li> <li>NIHNP: NIH NPAC ACONN collection of NP, NIH/NCATS, 3988 compounds</li> <li>OTAVAPEP: Alpha-helix Peptiomimetic Library, OTAVAchemicals (a-helix-Peptido), 1298 compounds</li> <li>ENAMDISC: Discovery Diversity Set -10, Enamine (DDS-10), 10,240 compounds</li> <li>ENAMMOL: Carboxylic Acid Fragment Library + Random, Enamine and Molport, 4378 compounds</li> <li>MCEDRUG: FDA-Approved Drug Library, MedChemExpress (HY-L022), 2610 compounds</li> <li>MCEDIV_50k_Sub: Subset of 50K Diversity Library, MedChemExpress (HY-L901), 20000 compounds</li> <li>TargetMolHTSNP: Subset of Natural Product Library for HTS, TargetMol (L6000), 2175 compounds</li> </ul> <p>Information regarding the SPECTYPE</p> <ul> <li>no SPECTYPE or SINGLE_BEST_SCAN: Best spectrum for each precursor and energy (highest TIC)</li> <li>'SAME_ENERGY' = Additionally, if a spectrum was acquired multiple times for a precursor with the same energy, they are merged into one spectrum only with the same energy (max. signal height used for each fragment signal).</li> <li>'ALL_ENERGIES' = merged spectrum of all used energies (in our case 3 for each precursor, using the merged (same energy) if available).</li> <li>'ALL_MSN_TO_PSEUDO_MS2' = mzmine merges all MSn into one pseudo MS2.</li> </ul> <p> </p> <p>MCEDIV and TargetmolNPHTS </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.