Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,448

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,448 results for “proteomic”

Learn how ShareScore rates datasets ↗
zenodo40/100

Proteomic data set of the analysis of black poplar (Populus nigra L.) seed storability

<p>Proteomic data set&nbsp;containing&nbsp;protein identification parameters (ESI MS/MS) and GO&nbsp;annotation functional classification (UniProt and QuickGO). Identification parameters of differentially abundant proteins of black poplar (<em>Populus nigra</em> L.) seeds stored in different temperature (3, -3, -20 and -196&deg;C) and time (12 and 24 months) conditions. Proteins were extracted and separated according to their isoelectric point (pI) and mass using 2-dimensional electrophoresis. Proteins that varied in abundance for temperature and time of storage were identified by mass spectrometry (ESI MS/MS). The mascot search algorithm (http://www.matrixscience.com) was used for protein identification against the NCBInr (http://www.ncbi.nig.gov) databases.Identified proteins were grouped due to biological process, molecular function and subcellular localization according to the gene ontology (GO) annotation using UniProt database and QuickGO search (https://www.ebi.ac.uk/QuickGO/).</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Proteomic characterization of human exhaled breath condensate.

<p>datasets from 3 studies, for&nbsp;In-depth proteomics characterization of exhaled breath condensate (EBC).</p> <p>1) Lacombe M. et al, 2018</p> <p>2) Muccilli V. et al, 2015</p> <p>3) Bredberg A.&nbsp;et al, 2012</p>

opencc-by-4.0Feb 2018View details →
zenodo40/100

Proteome of the Ceratopteris richardii fern (strain Hn-n)

<p>Proteome derived from de novo transcriptome assembly from fronds, mature gametophytes and spores of Ceratopteris richardii Hn-n strain. Transcriptome is deposited at&nbsp;&nbsp;<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB33372">https://www.ebi.ac.uk/ena/data/view/PRJEB33372</a></p> <p>After removing low-quality reads (reads lacking all four nucleotides or with a no-call), we assembled transcripts with Velvet (version 1.2.06) and Oases (version 0.2.06) using each of five k-mer values (k=45, 55, 65, 75, 85). Also, we converted the .fastq file to a non-redundant .fasta file and performed separate de novo transcriptome assembly with k=35, 45, 55, 65, and 75. Assembled transcripts were combined for each tissue, then redundant or fragmented sequences were removed based on BLASTN analysis. We determined the translational reading frame and corresponding peptide sequences from each assembled transcript based on BLASTP mapping results (after 6-frame translation in silico) to four plant reference proteome databases (Creinhardtii_169, Osativa_193_pep, Smoellendorffii_91_pep, TAIR10). Sequences lacking significant BLASTP scores to the reference proteomes were considered to be non-coding and omitted from the resulting fern proteome database. The resulting protein sequences derived from the three tissues were combined and a non-redundant protein sequence set computed based on clustering with UCLUST (version 4.2.66), requiring &gt;97% amino acid identity. The supporting code is available from the NuevoTx repository (https://github.com/taejoonlab/NuevoTx).&nbsp;</p> <p>This proteome was assembled for &quot;A pan-plant protein complex map reveals deep conservation and novel assemblies&quot;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Fig. 4 in Proteomic analysis of the venom of the social wasp Apoica pallens (Hymenoptera: Vespidae)

Fig. 4. Identified proteins of social wasp Apoica pallens venom and their functions according to the literature.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Fig. 2 in Proteomic analysis of the venom of the social wasp Apoica pallens (Hymenoptera: Vespidae)

Fig. 2. Two-dimensional reference gel (14%) of the venom of the social wasp Apoica pallens, showing the proteins identified by MALDI-TOF/TOF analysis.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Fig. 3 in Proteomic analysis of the venom of the social wasp Apoica pallens (Hymenoptera: Vespidae)

Fig. 3. Classification of proteins according to the representativity in number of proteins identified in the Apoica pallens venom.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Fig. 1 in Proteomic analysis of the venom of the social wasp Apoica pallens (Hymenoptera: Vespidae)

Fig. 1. Two-dimensional gel (14%) in triplicate (three extracts from a colony), with the three gels (A, B, and C) of the Apoica pallens wasp venom.

opencc-by-4.0Oct 2019View details →
zenodo40/100

Fig. 3. A in Differential proteomic analysis of date palm leaves infested with the red palm weevil (Coleoptera: Curculionidae)

Fig. 3. A pie chart presenting the classification of identified proteins according to their biological functions, expressed in percentage.

opencc-by-4.0Jun 2018View details →
zenodo40/100

Fig. 1 in Differential proteomic analysis of date palm leaves infested with the red palm weevil (Coleoptera: Curculionidae)

Fig. 1. Two-dimensional differential gel electrophoresis representative images of date palm proteins. The protein sample of control, wounded, infested, and internal standard (pooled of all the samples) are individually labeled with Cy dyes, mixed together and separated by two-dimensional differential gel electrophoresis followed by image scanning. (A) image of date palm control sample and labeled with cy3 dye; (B) image of date palm artificially wounded sample labeled with cy5 dye; (C) image of date palm sample infested with red palm weevil and labeled with cy3 dye; (D) image of date palm sample pooled from all and labeled with cy2 dye; (E) overlay gel of control, infested, and wounded along with internal standard.

opencc-by-4.0Jun 2018View details →
zenodo40/100

Fig. 2 in Differential proteomic analysis of date palm leaves infested with the red palm weevil (Coleoptera: Curculionidae)

Fig. 2. Venn diagram for the relative distribution of proteins spots in control, mechanically wounded, and red palm weevil infested date palm samples. The non-overlapping segment of diagram represent the number of proteins which were significantly up-regulated (&gt; 1.5-fold) in the corresponding group when compared with the other two groups. The overlapping region between any two groups represents the number of protein spots significantly up-regulated (&gt; 1.5-fold) compared to the third one. The central overlapping region depicts the protein spots where no statistically significant change in up- or down-regulation was observed.

opencc-by-4.0Jun 2018View details →
zenodo40/100

Long COVID IRIS Study Olink Proteomics Dataset

<p>Plasma samples collected from Stanford University's &ldquo;Infection Recovery in SARS-CoV-2&rdquo; (IRIS) study participants during acute SARS-CoV-2 infection, approximately 3 months post infection, and approximately 12 months post infection were analyzed using the Olink&reg; Target 96 Inflammation panel and Olink&reg; Target 96 Immune Response panel.&nbsp;</p> <p>Proteomic data were obtained from the Olink biomarker platform and presented as "NPX" (i.e. Normalized Protein eXpression) for each protein assay. The dataset features de-identified patient metadata, including participant study number (i.e. IRIS_number), sex, long COVID status (0 = recovered; 1 = Long COVID), and timepoint of the plasma sample (i.e. acute infection sample, approximately 3 months post infection, or approximately 12 months post infection).&nbsp;</p> <p>Notes:</p> <ul> <li>Consistent with the World Health Organization (WHO) definition, long COVID was defined in this study as the continuation or development of symptoms three months after SARS-CoV-2 infection, which were not readily attributable to other etiologies.</li> <li>For acute infection samples, long COVID status (0 = recovered; 1 = Long COVID) refers to whether the patient will have fully recovered or will have long COVID at 3 months post infection.&nbsp;</li> <li>For 3 month samples, long COVID status (0 = recovered; 1 = Long COVID) refers to whether the patient has fully recovered or has long COVID at 3 months post infection.</li> <li>For 12 month samples, long COVID status (0 = recovered; 1 = Long COVID) refers to whether the patient has fully recovered or has ongoing long COVID at 12 months post infection, but these patients all had long COVID at 3 months post infection.</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo40/100

UniSpec: Deep Learning for Predicting the Full Range of Peptide Fragment Ion Series to Enhance the Proteomics Data Analysis Workflow

<p>UniSpec is a comprehensive DL spectrum predictor that can predict the intensity of the entire HCD MS/MS fragment ion series, going beyond existing tools limited to b/y ion series.&nbsp;</p> <p>All datasets developed for UniSpec model are shared on Zenodo as part of the UniSpec publication, "UniSpec: Deep Learning for Predicting Comprehensive Peptide Fragment Ion Series to Improve Peptide-Spectrum Matches from Shotgun Proteomics Experiments".</p> <p>This includes UniSpec datasets, downstream evaluation and analysis, and application case studies.</p> <p>1. pre-processed training, evaluation and testing data for machine learning;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp;UniSpec-Datasets.7z, Readme_UniSpecDatasets.txt</p> <p>2. Streamlined &nbsp;input datasets based on the fragmentation dictionary;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; Streamlined_inputdatasets.7z, Readme_Streamlined_inputdatasets.txt</p> <p>3. Predictions on the validation and test sets;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;UniSpecPred_Validation-Test.7z, Readme_Predictons_ValidationTest.txt</p> <p>4. Evaluation by comparison with Prosit;</p> <p>&nbsp; &nbsp; &nbsp; a. Predictions: prosit_and_unispec_predictions.7z, Readme_prosit_and_unispec_predictions.txt</p> <p>&nbsp; &nbsp; &nbsp; b. Cosine similarity scores: prosit_vs_unispec_CS.7z, Readme_prosit_vs_unispec_CS.txt</p> <p>5. CSS for Different HCD Fragment Ion Series;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp;CS_for_ion_splits.tsv</p> <p>6. Application 1: PSM rescoring;</p> <p>&nbsp; &nbsp; &nbsp; PSM rescoring_zipfiles.7z, &nbsp;PSM rescoring_readme.txt</p> <p>7. Application 2: In-silico spectral library search &nbsp;</p> <p>&nbsp; &nbsp; &nbsp; in-silico_librarysearch.7z, in-silico_librarysearch_readme.txt</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Comprehensive Single Point Mutational Landscape Analysis of the Monkeypox Virus Proteome

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo40/100

Proteome Data for A. thaliana

<p>Datafile with calculated metrics and associated data for A. thaliana proteome</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Proteome Data for S. cerevisiae

<p>Datafile with calculated metrics and associated data for S. cerevisiae proteome</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

A3DyDB: exploring structural aggregation propensities in the yeast proteome.

<h3>The data from the paper: Garcia-Pardo, J., Badaczewska-Dawid, A.E., Pintado-Grima, C. <em>et al.</em>&nbsp;A3DyDB: exploring structural aggregation propensities in the yeast proteome.&nbsp;<em>Microb Cell Fact</em>&nbsp;<strong>22</strong>, 186 (2023). https://doi.org/10.1186/s12934-023-02182-3</h3> <h4>The unified and integrated metadata accompanied by referencing identifiers in the A3D database is available for download in CSV format.</h4> <h4>The unified and integrated metadata accompanied by referencing identifiers in the A3D database is available for download in CSV format.</h4>

opencc-by-4.0Sep 2024View details →
zenodo40/100

A3D Model Organism Database (A3D-MODB): a database for proteome aggregation predictions in model organisms

<p>The unified and integrated metadata accompanied by referencing identifiers from the A3D database is available for download in CSV format.</p> <p>Aleksandra E Badaczewska-Dawid, Aleksander Kuriata, Carlos Pintado-Grima, Javier Garcia-Pardo, Michał Burdukiewicz, Valent&iacute;n Iglesias, Sebastian Kmiecik, Salvador Ventura, A3D Model Organism Database (A3D-MODB): a database for proteome aggregation predictions in model organisms,&nbsp;<em>Nucleic Acids Research</em>, Volume 52, Issue D1, 5 January 2024, Pages D360&ndash;D367,&nbsp;<a href="https://doi.org/10.1093/nar/gkad942">https://doi.org/10.1093/nar/gkad942</a></p>

opencc-by-4.0Sep 2024View details →
dryad40/100

Regression models generated by APRANK (computational prioritization of antigenic proteins and peptides from complete pathogen proteomes)

<p>Availability of highly parallelized immunoassays has renewed interest in the discovery of serology-based biomarkers for infectious diseases. Protein and peptide microarrays now provide a high-throughput platform for immunological screening of potential antigens and B-cell epitopes. However, there is still a need to prioritize relevant probes when designing these arrays. In this work we describe a computational method called APRANK (Antigenic Protein and Peptide Ranker) which integrates multiple molecular features to prioritize antigenic targets starting from a given pathogen proteome. These features include subcellular localization, presence of repetitive motifs, natively disordered regions, secondary structure, transmembrane spans and predicted interaction with the immune system. We applied this method to the prioritization of potential diagnostic antigens and peptides in a number of pathogen proteomes and human diseases: Borrelia burgdorferi (Lyme disease), Brucella melitensis (Brucellosis), Coxiella burnetii (Q fever), Escherichia coli (Gastroenteritis), Francisella tularensis (Tularemia), Leishmania braziliensis (Leishmaniasis), Leptospira interrogans (Leptospirosis), Mycobacterium leprae (Leprae), Mycobacterium tuberculosis (Tuberculosis), Plasmodium falciparum (Malaria), Porphyromonas gingivalis (Periodontal disease), Staphylococcus aureus (Bacteremia), Streptococcus pyogenes (Group A Streptococcal infections), Toxoplasma gondii (Toxoplasmosis) and Trypanosoma cruzi (Chagas Disease). After training a linear regression model the method achieves good to excellent performance on most species, measured by the enrichment of validated antigens at the top of the ranking. An unbiased validation using independent data sets shows APRANK is successful in predicting antigenicity for all pathogen species tested. We make APRANK available to facilitate the identification of novel diagnostic antigens in infectious diseases.</p>

opencc-zeroJun 2021View details →
zenodo40/100

Protein language model embeddings and predictions of the human proteome

<p>Residue and sequence embeddings of the human proteome (SwissProt for organism Human, downloaded on&nbsp;2021.06.09)&nbsp;computed using bio_embeddings (bioembeddings.com) using the ProtT5 embedder at full precision (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3).</p> <p>Additionally:</p> <p>- Sequence-level&nbsp;predictions of subcellular localization in 10 classes using LA (https://www.biorxiv.org/content/10.1101/2021.04.25.441334v1)</p> <p>- Residue-level three state secondary structure prediction (alpha, sheet or other) using models reported&nbsp;in the ProtTrans paper (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3)</p> <p>&nbsp;</p> <p>Files included:</p> <p>- human.fasta --&gt; FASTA-formatted sequences of human from SwissProt</p> <p>-&nbsp;DSSP3_human_ProtT5Sec.fasta --&gt; Secondary structure predictions in three states for each residue of each protein&nbsp;in human.fasta. &quot;H&quot; stands for Helix; &quot;E&quot; stands for Sheet; &quot;C&quot; stands for Other.</p> <p>-&nbsp;subcell_human_LA_ProtT5.csv --&gt; Subcellular location (10 states) and memrane-boundness (2 states)&nbsp;for each protein in human.fasta</p> <p>-&nbsp;embeddings_file.h5 --&gt; per-residue embeddings of sequences in human.fasta. Each dataset&nbsp;in the .h5 file represents a protein sequence and contains a matrix of length Lx1024, with L being the length of the protein sequence. Datasets are indexed using integers. The original sequence identifier (from the FASTA header) can be accessed through the &quot;original_id&quot; attribute. See&nbsp;https://docs.bioembeddings.com/v0.2.0/notebooks/open_embedding_file.html for information on how to open the file</p> <p>-&nbsp;reduced_embeddings_file.h5 --&gt; per-sequence embeddings of sequences in human.fasta (obtained by mean-pooling the residue-embeddings along the length dimension of the protein sequence). Each dataset&nbsp;in the .h5 file represents a protein sequence and contains a vector of size 1024 (meaning, each sequence has the same dimension).</p>

openafl-3.0Jun 2021View details →
zenodo40/100

Proteomic data (SWATH-MS) of mouse uterine horns treated with different types of plasma

<p>This dataset contains the proteomic data (SWATH-MS) from 48 mouse uterine horns corresponding to a murine model of Asherman&#39;Syndrome (presence of intrauterine adhesions). These 48 uterine horns correspond to 26 NOD-SCID mice&nbsp;(mouse uterus are bicornuate - 2 uterine horns per mouse) distributed in 4&nbsp;groups (n = 6 /group), attending to the treatment received:&nbsp;Control (n = 6; milliQ H2O was injected), non-activated umbilical cord plasma &nbsp;(n = 6), activated umbilical cord plasma (n =&nbsp;6), and activated platelet-rich plasma from adult blood (n = 6). To simulate Asherman&#39;s Syndrome, we induced endometrial damage (using a needle) inside the lumen of left uterine horns from all animals, while&nbsp;right horns were left undamaged.</p>

opencc-by-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record