Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,448
datasets available to search
ShareScore release 0.7.1
Dataset results
1,448 results for “Proteomics”
CPT-1 whole-proteome feature matrices (EVE set)
<p><strong>Cross-protein transfer learning for variant effect prediction</strong></p> <p>This repository contains the feature matrices for CPT-1 to make variant effect prediction on 3,045 human proteins within the EVE set (<a href="https://www.nature.com/articles/s41586-021-04043-8">Frazer et al., 2021</a>), initially released with the manuscript "Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects".</p> <p> </p> <p><strong>Citation</strong></p> <p>Jagota, M.*, Ye, C.*, Albors, C., Rastogi, R., Koehl, A., Ioannidis, N., and Song, Y.S.†<br> "Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects", bioRxiv (2022)</p> <p>*These authors contributed equally to this work.<br> †To whom correspondence should be addressed: <a href="mailto:yss@berkeley.edu">yss@berkeley.edu</a></p> <p>DOI: <a href="https://doi.org/10.1101/2022.11.15.516532">https://doi.org/10.1101/2022.11.15.516532</a></p>
ECOD Classification of AFDB 48 Proteomes
<p>ECOD domains classified for the 48 whole proteomes (model_v4) from AFDB. Domains were classified using Domain Parser for Alphafold Models. </p>
Proteomic analysis reveals different molecular mechanisms to face water deficit in mycorrhizal and nonmycorrhizal sorghum plants
<p>Differential accumulated proteins in response to water deficit in mycorrhizal and nonmycorrhizal sorghum plants were recovered from 2D gels and identified by HPLC-MSMS. MS analysis was performed by a Nano acquity nanoflow LC system (Waters, Milford, MA, USA) coupled to a linear ion trap (LTQ) velos mass spectrometer (Thermo Fisher Scientific, Bremen, Germany) equipped with a nanoelectrospray ion source.</p>
Traditional biochemistry versus proteomics
<p>"Traditional Biochemistry versus Proteomics" . Creator: Esteban Núñez, Chile. New version of a classic image (Tao Cartoon Image, https://www.chem.purdue.edu/people/profile/taow).<br> </p>
Tooth enamel proteome of Early Medieval non-adult individuals
<p>This dataset includes .raw LC-MS/MS files, from a proteomic study of deciduous and permanent tooth enamel samples. It includes data of 30 different non-adult individuals (Early Middle Ages, Valdaro, Italy), whose sex has been estimated through amelogenin peptides. This dataset is linked to a submitted publication (Lugli et al., <em>Journal of Archaeological Science: Reports</em>). <br> Please, refer to Lugli et al. (2019, <em>Scientific Reports</em>; doi: 10.1038/s41598-019-49562-7) for methodology. </p>
Proteomes in 3D - Correlation Analyses - Fluxes vs LiP Peptides
<p>Analysis of metabolic fluxes that correlate with protein structural changes across 8 metabolic conditions in E. coli, associated to the manuscript by Cappelletti et al., currently under consideration. Results are expressed as data from a given metabolic condition relative to growth in glucose.</p> <p>Plot title: Protein name_Peptide sequence</p> <p>X-axis: Log<sub>2</sub> FC [LiP peptide (condition/glucose)]</p> <p>Y-axis: Log<sub>2</sub> FC [Flux (condition/glucose)]</p> <p>The following files contain the statistical parameters of the analysis (p and q-values and R-square values) for all the uploaded plots:</p> <p>- Statistics_p_and_q_values.txt</p> <p>- Statistics_R2.txt</p>
Training dataset: Mass spectrometry based proteomics of healthy human serum samples
<p>The two raw files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>Serum of a healthy person was obtained by centrifugation of full blood in a serum-gelmonovette. One serum sample was depleted for high abundant proteins, the other not.<br> For the non-depleted sample: 5µl of serum was diluted with 0.1% Rapigest, resulting in a concentration of 1mg/ml.<br> Depletion was performed with the Seppro IgY14 Spin columns which are able to deplete 14 high abundent blood proteins by immunoaffinity. For the depleted sample 9µl of serum was diluted with TBS/HCl/NaCl buffer and added to the Seppro IgY14 spin column. After depletion the sample was buffered with Hepes pH 8.0 and Rapigest was added to a final 0.1% Rapigest concentration. From here on, both samples were reduced by adding TCEP, alkylated by IAA and quenched with DTT in solution. Digestion was performed by adding trypsin in a ratio of 1:50 to the samples. After incubation at 37°C, 600rpm, over night, the sample clean-up was performed with the PreOmics desalting columns. iRT peptides were added and the sample was measured with a Q-Exactive Plus mass spectrometer. Besides the two raw files, we uploaded a fasta file that serves as human protein sequence database and the Galaxy MaxQuant training result files: protein groups, peptides, mqpar and PTXQC.</p>
Supplementary Material: Conformational Ensemble of the Poliovirus 3CD Precursor Observed by MD Simulations and Confirmed by SAXS: A Strategy to Expand the Viral Proteome?
<p>Supplementary video for <em>Viruses</em> <strong>2015</strong>, <em>7</em>(11), 5962-5986; doi:10.3390/v7112919; http://www.mdpi.com/1999-4915/7/11/2919.</p> <p><strong>Movie S1.</strong> Dynamic interface between 3C and 3D domains revealed by accelerated MD. The 3C and 3D domains are colored cyan and blue, respectively. The active-site residues of the protease (His-40, Glu-71, Cys-147) and the polymerase (Asp-416, Asp-511, Asp-512) domains are represented by red spheres to help identifying the relative orientations of two domains.</p>
Proteomic analysis of silenced cathepsin B expression suggests non-proteolytic cathepsin B functionality.
<p>A list of human protein uniprot IDs. The proteins were identified by LC-MS/MS in the cellular supernatant of MDA-MB-231 cells, originally published in:</p> <p>F.C. Sigloch, J.D. Knopf, J. Weißer, A. Gomez-Auli, M.L. Biniossek, A. Petrera, et al., Proteomic analysis of silenced cathepsin B expression suggests non-proteolytic cathepsin B functionality, Biochim. Biophys. Acta - Mol. Cell Res. 1863 (2016) 2700–2709. doi:10.1016/j.bbamcr.2016.08.005. https://www.ncbi.nlm.nih.gov/pubmed/27526672</p>
SUPPLEMENTARY (For MD) An integrative pan-genome and subtractive proteomics approach for the identification of potential novel therapeutic drug target against antibiotic resistant honeybee pathogen Paenibacillus larvae
<p><strong>Parameters</strong></p><p>Force field: AMBER ff19SB</p><p>Water type: TIP3P</p><p>Ions: NaCl </p><p>Ligand topology force field: GAFF2</p><p>Temperature: 298k</p><p>Pressure: 1 bar</p><p>minimization step: 20000 on 5 nanoseconds</p><p>initial velocity is changed by changing "ntx" and "ig"</p><p>C2: ntx = 5 , ig = 8</p><p>C3: ntx = 2 , ig = 5</p><p> </p><p><strong>Uploads</strong>- </p><p>1. Zip file of all 3 main files</p><p>2. Unzip file of C1 (Trajectory, PDB complex after each 10 ns run, and Mp4 video of Complex)</p><p>3. Zip file of C1</p><p>4. Unzip file of C2 (Trajectory, PDB complex after each 10 ns run, and Mp4 video of Complex)</p><p>5. Zip file of C2</p><p>6. Unzip file of C3 (Trajectory, PDB complex after each 10 ns run, and Mp4 video of Complex)</p><p>7. Zip file of C3</p><p>8. Zip and unzip file of <strong>Initial</strong> PDB of complex prior to MD simulation with <strong>Post</strong> MD PDB (C1, C2, C3)</p><p>9. Zip file of <strong>topology</strong> files for C1, C2, and C3</p>
Supplementary Data for Stabilizing the Proteomes of Acute Myeloid Leukemia Cells: Implications for Cancer Proteomics
<p>Supplementary data for: <br>Stabilizing the Proteomes of Acute Myeloid Leukemia Cells: Implications for Cancer Proteomics<br>Authors: Robert Sprung, Qiang Zhang, Michael H. Kramer, Matthew C. Christopher, Petra Erdmann-Gilmore, Yiling Mi, James P. Malone, Timothy J. Ley, and R. Reid Townsend.</p><p>Table S1 - AML Case descriptors and LC-MS data files<br>Table S2 - All Peptides by Case -LFQ<br>Table S3 - Identification of tryptic and non-tryptic peptides from five AML cases with high and low expression of ELANE<br>Table S4 - Number of proteins identified by LFQ proteomics with a minimum of 2 tryptic peptides<br>Table S5 - DFP Adduct Database Search Tryptic Peptides<br>Table S6 - Protein quantification from TMT 11-plex tryptic peptides with and without DFP<br>Table S7 - Tryptic peptides used for protein quantification from TMT 11-plex with and without DFP<br>Table S8 - Changes in TMT relative abund. with DFP treatment<br>Table S9 - Protein quantification from LFQ tryptic peptides with and without DFP<br>Table S10 - Proteins with significant change in abundance with DFP treatment using Label-Free Quantitation</p>
iGEMME Missense Mutational Effect Predictions for Entire Human Proteome
<p>This dataset contains iGEMME single point mutation predictions of about ~19000 human proteins. In iGEMME predictions, only evolutionary data coming from multiple sequence alignment files is used. </p> <h2>Description of the data and file structure</h2> <p>This dataset contains iGEMME predictions for all human proteins. </p> <p>Data of each human protein is in a folder named after its uniprotID. Inside uniprotID folder, there is a subfolder called results that contain all input and output. An example results folder for uniprotID A0A0B4J245 will contain the following files:</p> <ol> <li> <p><strong>Raw igemme predictions (output file):</strong> A0A0B4J245_normPred_evolCombi_igemme.txt</p> </li> <li> <p><strong>Ranksorted (between 0-1) igemme predictions in csv format (output file):</strong> A0A0B4J245_normPred_evolCombiTransposedRanksorted_igemme.csv</p> </li> <li> <p><strong>Colabfold MSA file (input file):</strong> aliA0A0B4J245.fasta</p> </li> <li> <p><strong>JET2 file containing JET scores for each amino acid (output file) :</strong> A0A0B4J245_jet_igemme.res</p> </li> <li> <p><strong>Configuration file containing default parameters (output file):</strong> default.conf</p> </li> <li> <p><strong>Log file (output file):</strong> igemme.log</p> </li> </ol>
The Q-TOF proteomics data for the identification of mammalian L-fucose dehydrogenase (EC 1.1.1.122)
<p>The enclosed zip file contains data files (RAW format) from MS^E experiment. The experiment was performed with the use of Acquity nanoUPLC coupled with a Synapt G2 HDMS Q-TOF mass spectrometer (Waters) fitted with a nanospray source. It aimed at the identification of proteins present in the gel bands S1-S9 and the gel band Z1. The bands have come from SDS-PAGE and zymography analyses, respectively, of the most active enzyme fraction from the Reactive Red Agarose 120 purification step.</p>
Proteome analysis of Corynebacterium diphtheriae - macrophage interaction
<p>Contact of <em>Corynebacterium diphtheriae</em> with macrophages induce adaptations on both bacterial and cellular sides. Using an experimental design involving gentamicin protection and liquid chromagraphy followed by mass-spectrometry, a multi-species proteomic dataset was analyzed at different time points of the infection assay. Several previously undescribed <em>Corynebacterium </em>proteins were differentially regulated, as well as key macrophage components of the phagolysosome. Overall, Bacteria responded to phagocytosis by changes in DNA repair, transcription and cell wall synthesis proteins, while macrophages showed changes in components of the innate immune system.</p> <p>This dataset consists of:</p> <ul> <li> Raw protein abundance data for Macrophage THP-1 cells (M0) obtained by LC-MS/MS followed by peptide sequencing using Proteome Discoverer (ThermoFisher) -see .zip folder.</li> <li>Raw protein abundance data for Macrophage <em>C. diphtheriae </em>ISS3319 (CD) obtained by LC-MS/MS followed by peptide sequencing using Proteome Discoverer (ThermoFisher) - see .zip folder.</li> <li>Differential protein abundance analysis for CD and M0 using LIMMA. </li> <li>Data underlying the growth curves observed in CD in RPMI + 10% FBS conditions (infection assay conditions).</li> <li>BLASTP searches, PFAM clans, and InterPro annotations for CD.</li> <li>STRING-based PPI network (baseline) and APSPs between differentially abundant proteins in CD. </li> <li>Related R scripts.</li> </ul>
Code to generate figures 3 and 4 of: "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."
<p>Code to generate figures 3 and 4 of the manuscript titled "A comprehensive LFQ benchmark dataset to validate data analysis pipelines on modern day acquisition strategies in proteomics."</p> <p> </p>
Immuno-proteomic profiling reveals aberrant immune cell regulation in the airways of individuals with ongoing post-COVID-19 respiratory disease
<p><span><span><span><span><span><span><span><span><span><span><span>Some patients hospitalized with acute COVID-19 suffer respiratory symptoms that persist for many months. We delineated the immune-proteomic landscape in the airway and peripheral blood of healthy controls and post-COVID-19 patients 3 to 6 months after hospital discharge. Post-COVID-19 patients showed abnormal airway (but not plasma) proteomes, with elevated concentration of proteins associated with apoptosis, tissue repair and epithelial injury versus healthy individuals. Increased numbers of cytotoxic lymphocytes were observed in individuals with greater airway dysfunction, while increased B cell numbers and altered monocyte subsets were associated with more widespread lung abnormalities. 1 year follow-up of some post-COVID-19 patients indicated that these abnormalities resolved over time. In summary, COVID-19 causes a prolonged change to the airway immune landscape in those with persistent lung disease, with evidence of cell death and tissue repair linked to ongoing activation of cytotoxic T cells. </span></span></span></span></span></span></span></span></span></span></span></p>
Mining folded proteomes in the era of accurate structure prediction
<p>Supplementary data to accompany the manuscript “Mining folded proteomes in the era of accurate structure prediction”. Contains three zip files with fold matching search results to support results in the main text.</p>
Protein language model embeddings and predictions for the fly proteome (FlyBase)
<p>Residue and sequence embeddings of the fly (drosophila melanogaster) proteome (FlyBase for organism drosophila melanogaster, downloaded on 2022.03.01) computed using bio_embeddings (bioembeddings.com) using the ProtT5 embedder at full precision (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3). To open the embeddings file, please see <a href="https://github.com/sacdallago/bio_embeddings/blob/develop/notebooks/open_embedding_file.ipynb">this notebook</a>. The embeddings will be indexed by numbers according to the mapping file (mapping_file.csv) in this dataset. All following results will share the same mapping (for instance, to access the variation prediction results, by accessing index "0", you will query results for the sequence "FBpp0304622").</p> <p>Additionally:</p> <p>- Sequence-level predictions of subcellular localization in 10 classes using LA (https://www.biorxiv.org/content/10.1101/2021.04.25.441334v1)</p> <p>- Residue-level three state secondary structure prediction (alpha, sheet or other) using models reported in the ProtTrans paper (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3)</p> <p>- Residue-level prediction of conservation (in 9 states) and of variation effect (from 0 [no-effect] to 1 [effect]) using VESPAl (https://doi.org/10.1007/s00439-021-02411-y)</p> <p> </p> <p>Files included:</p> <p>- dmel-all-translation-r6.44.fasta --> FASTA-formatted sequences of drosophila melanogaster from FlyBase</p> <p>- mapping_file.csv --> A CSV file mapping the identifiers used in the following files (from 0 to 30737) to the identifiers in the FlyBase fasta file (dmel-all-translation-r6.44.fasta).</p> <p>- DSSP3_fly_ProtT5Sec.fasta --> Secondary structure predictions in three states for each residue of each protein in dmel-all-translation-r6.44.fasta. "H" stands for Helix; "E" stands for Sheet; "C" stands for Other.</p> <p>- subcell_fly_LA_ProtT5.csv --> Subcellular location (10 states) and memrane-boundness (2 states) for each protein in dmel-all-translation-r6.44.fasta</p> <p>- embeddings_file.h5 --> per-residue embeddings of sequences in dmel-all-translation-r6.44.fasta. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length Lx1024, with L being the length of the protein sequence. Datasets are indexed using integers. The original sequence identifier (from the FASTA header) can be accessed through the "original_id" attribute. See https://docs.bioembeddings.com/v0.2.0/notebooks/open_embedding_file.html for information on how to open the file.</p> <p>- reduced_embeddings_file.h5 --> per-sequence embeddings of sequences in dmel-all-translation-r6.44.fasta (obtained by mean-pooling the residue-embeddings along the length dimension of the protein sequence). Each dataset in the .h5 file represents a protein sequence and contains a vector of size 1024 (meaning, each sequence has the same dimension).</p> <p>- conspred_probs.h5 --> per-sequence conservation probability (softmax) prediction of sequences in dmel-all-translation-r6.44.fasta in 9 classes. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length 9xL, with L being the length of the protein sequence, and 9 being the predicted conservation class (index 0 = very variable; index 8 = very conserved)</p> <p>- vespal_SAVeffect_fly.zip --> zipped .h5 file of per-sequence variation predictions of sequences in dmel-all-translation-r6.44.fasta on a scale from 0 (neutral) to 1 (effect). -1 indicates WT substitution. Each dataset in the .h5 file represents a protein sequence and contains a matrix of length 20xL, with L being the length of the protein sequence, and 20 being the predicted variation score for each residue substitution (AAs in the following order: "<strong>ALGVSREDTIPKFQNYMHWC</strong>" . Meaning that index 0 = substitution of the residue to "A", index = 1 substitution to residue "L", aso.)</p>
Proteomics data of mitochondrial fraction of CRL-2097 cancer cell line model
<p>The cancer cell line model developed using human dermal fibroblasts CRL-2097 was used in these experiments:</p> <p>Sample 1 - CRL2097 + hTERT</p> <p>Sample 2 - CRL2097 + hTERT + LT</p> <p>Sample 2 - CRL2097 + hTERT + LT + Ras</p> <p>The mitochondrial fraction was prepared from each of these cell lines and analysed via mass spec for their proteomics. The experiment was done in duplicates. </p>
Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity"
<p>Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity".</p> <p>Data and further information at GitHub repository https://github.com/kreutz-lab/dia-benchmarking (DOI: 10.5281/zenodo.6371925)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.