Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,448
datasets available to search
ShareScore release 0.7.1
Dataset results
1,448 results for “proteomic”
Data from: A universal tool for marine metazoan species identification – Towards best practices in proteomic fingerprinting
<p><span>Proteomic fingerprinting using MALDI-TOF mass spectrometry is a well-established tool for identifying microorganisms and has shown promising results for identification of animal species, particularly disease vectors and marine organisms. However, few studies have tested species identification across different orders and classes. In this study, we collected data from 1,246 specimens and 198 species to test species identification in a diverse dataset. We also evaluated different specimen preparation and data processing approaches for machine learning and developed a workflow to optimize classification using random forest. Our results showed high success rates of over 90%, but we also found that the size of the reference library affects classification error. Additionally, we demonstrated the ability of the method to differentiate marine cryptic-species complexes and to distinguish sexes within species.</span></p>
Proteome of peripheral mononuclear cells (PBMCs) from asymptomatic malaria and uninfected individuals and the ensuing malaria episodes
<p>Cumulative malaria parasite exposure in endemic regions often results in the acquisition of partial immunity and asymptomatic infections. There is limited information on how host-parasite interactions mediate maintenance of chronic symptomless infections that sustain malaria transmission. Here, we have determined the gene expression profiles of the parasite population and the corresponding host peripheral blood mononuclear cells (PBMCs) from 21 children (<15 years). We compared children who were defined as uninfected, asymptomatic and those with febrile malaria. Children with asymptomatic infections had a parasite transcriptional profile characterized by a bias toward trophozoite stage (~12 hours-post invasion) parasites and low parasite levels, while earlier ring stage parasites were characteristic of febrile malaria. The host response of asymptomatic children was characterized by downregulated transcription of genes associated with inflammatory responses, compared to children with febrile malaria. Interestingly, the host responses during febrile infections that followed an asymptomatic infection featured stronger inflammatory responses, whereas the febrile host responses from previously uninfected children featured increased humoral immune responses. The priming effect of prior asymptomatic infection may explain the blunted acquisition of antibody responses seen to malaria antigens following natural exposure or vaccination in malaria endemic areas.</p>
Data from: Pre- and post- treatment fold change plasma proteomics in metastatic NSCLC patients
<p>Blood plasma samples were collected from advanced-stage NSCLC patients as part of a clinical study (PROPHETIC; NCT04056247). All clinical sites received IRB approval for the study protocol. Patient blood samples were drawn at baseline (referred to as T0) and, on average, 4 weeks after the treatment commenced, prior to the second dose of treatment (referred to as T1). Blood samples were drawn into tubes containing EDTA as an anticoagulant, and plasma was separated from the whole blood. The protocol adheres to the Clinical and Laboratory Standards Institute (CLSI) guidelines.</p>
DATASET - Mass Spectrometry - Snake venom proteomics of seven taxa of the genera Vipera, Montivipera, Macrovipera and Daboia across Türkiye
<p><strong>Publication: Damm <em>et al.</em> 2024 - <a title="DOI URL" href="https://doi.org/10.1021/acs.jproteome.4c00171">https://doi.org/10.1021/acs.jproteome.4c00171</a></strong></p> <p> </p> <p><strong>This DATASET collection includes the mass spectrometry files for proteomics venom investigation of seven taxa of the genera <em>Vipera</em>, <em>Montivipera</em>, <em>Macrovipera</em> and <em>Daboia </em>across Türkiye.</strong></p> <p><strong>Species list:</strong></p> <ol> <li>Vipera berus barani</li> <li>Vipera darevskii</li> <li>Montivipera bulgardaghica bulgardaghica </li> <li>Montivipera bulgardaghica albizona</li> <li>Montivipera xanthina</li> <li>Macrovipera lebetinus obtusa</li> <li>Daboia palaestinae</li> </ol> <p><strong>Folders 01-07 - BOTTOM-UP PROTEOMICS</strong>: The venom pools were investigated by the bottom-up "snake venomics" (labled as SVX) approach and in short: separated by RP-HPLC, followed by SDS-PAGE separation and the single bands were in-gel processed by DTT, IAC and finally o/n tryptic digested. Samples submitted to HPLC-MS/MS. Early peptidic fractions of the first HPLC run were directly submitted to HPLC-MS/MS analytic w/o further gel procession. Folders 01 to 07 include the MS and MS/MS spectra of the snake species 1-7, respectively. Files are included as RAW and MZML format.</p> <p>Used instrument: LTQ Orbitrap XL mass spectrometer (Thermo, Bremen, Germany) with an Agilent 1260 HPLC system (Agilent Technologies, Waldbronn, Germany) using a reversed-phase Grace Vydac 218MS C18 (2.1 × 150 mm; 5 μm particle size) column.</p> <p>Modifications: UNIMOD:4 - \"Iodoacetamide derivative.\"</p> <p>Used protein database: Uniprot_8570_serpentes_reviewed_canonical_2640_entries_cRAP_210408.fasta</p> <p><strong>Folders 10-11 - TOP-DOWN PROTEOMICS</strong>: The venom pools were investigated by the non-reduced and TCEP reduced top-down (labled as TD) approach and in short: untreated or TCEP reduced samples submitted to HPLC-MS/MS. Folders 10 and 11 include the MS and MS/MS spectra of the snake species 1-7 as labled. Files are included as RAW and MZML format.</p> <p>Used instrument: Q Exactive HF mass spectrometer (Thermo, Bremen, Germany) with a Vanquish ultra-high performance liquid chromatography (UHPLC) system (Agilent Technologies, Waldbronn, Germany) using a reversed-phase Supelco Discovery BIO wide C18 (2.0 × 150 mm; 3 μm particle size; 300 Å pore size).</p> <p>Modifications: none (either red. or non-red. disulfide bridges)</p> <p>Used protein database for TopPIC analysis: Uniprot_8570_serpentes_reviewed_ISOandCAN_2749_entries_NOcRAP_231011.fasta</p> <p> </p>
Data repository associated with 'A Functional Map of the Human Intrinsically Disordered Proteome'
<p><strong>ES_MAP.zip</strong></p> <ul> <li>a hierarchically clustered map of the human IDR-ome</li> <li>.cdt and .gtr files - outputs of Cluster3.0 software</li> <li>can be visualized using JavaTreeView (see Tutorial_ES.pdf)</li> </ul> <p><strong>TUTORIAL.zip</strong>, information on:</p> <ul> <li>visualization and analysis of the human IDR-ome map</li> <li>search for proteins of interest and exploratory analyses of clusters</li> <li>automatic export and analysis of exported clusters (code available at https://github.com/IPritisanac/ES_PW)</li> </ul> <p><strong>IDROME_SEQUENCES.zip</strong></p> <ul> <li>human proteome fasta file</li> <li>IDRome fasta file</li> <li>SPOT-Disorder v1.0 disorder boundaries <ul> <li>13 044 unique protein sequences with at least one IDR (>=30 amino acids)</li> <li>21 252 total unique human IDRs</li> </ul> </li> </ul> <p><strong>IDR_ALN.zip</strong></p> <ul> <li>alignments of IDR sequences across ENSEMBL orthologs</li> <li>19 459 IDR alignments</li> <li>UniProt ID and IDR boundaries for the human sequence are indicated in the name of the file</li> </ul> <p><strong>FAIDR_TSTATS.zip</strong></p> <ul> <li>hierarchical clustering of FAIDR t-statistics for 148 GO terms<br> <ul> <li>.cdt, .gtr files from Cluster3.0</li> <li>can be visualized using JavaTreeView</li> <li>reveals the most predictive molecular features for the top performing 148 models</li> </ul> </li> </ul> <p><strong>CLUSTERS_EXPLORE.zip</strong></p> <ul> <li>clusters obtained through exploratory analysis of the map provided in ES_MAP.zip</li> <li>93 exported clusters in .cdt file format</li> </ul> <p><strong>CLUSTERS_AUTO.zip</strong></p> <ul> <li>clusters extracted from the hierarchically clustered IDR-ome map at a range of distance thresholds (0.4 - 0.8) in .cdt file format</li> <li>distance refers to the uncentered correlation distance between vectors of Z-scores representing human IDRs</li> <li>clusters extracted at different distance thresholds are split into separate archives</li> <li>AUTO_GO_FEATS.xlsx - summary of GO-term overrepresentation and feature enrichment analyses; each distance threshold is in a separate sheet</li> </ul> <p><strong>FAIDR_HIGH_AUC_PPV_GO.zip</strong></p> <ul> <li>target files with annotations of 148 GO terms for which good quality FAIDR models could be obtained (AUC >= 0.7, PPV >= 0.4)</li> <li>file format: three columns; 1st: IDR ID (includes IDR boundaries); 2nd: protein UniProt ID; 3rd: annotation of the protein to a GO term (1 if known to be associated with the GO term, 0 if not)</li> </ul> <p> </p> <p> </p>
Processed OLINK serum proteomics data MIS-C patients versus healthy controls
<p>This dataset contains processed OLINK serum proteomics data MIS-C patients versus healthy controls. Data was generated by Diorio et al. (Diorio, C., Shraim, R., Vella, L.A. <em>et al.</em> Proteomic profiling of MIS-C patients indicates heterogeneity relating to interferon gamma dysregulation and vascular endothelial dysfunction. <em>Nat Commun</em> <strong>12</strong>, 7222 (2021). https://doi.org/10.1038/s41467-021-27544-6). Processing in format provided here was done by dr. Levi Hoste. This table is used in the MultiNicheNet package (https://github.com/saeyslab/multinichenetr) and mentioned in the updated corresponding manuscript. </p>
Peripheral priming induces plastic transcriptomic and proteomic responses in circulating neutrophils required for pathogen containment
<p><strong>When using any of this data, please cite the corresponding manuscript</strong></p> <div> <div> <p><a title="Rainer Kaiser et al., Peripheral priming induces plastic transcriptomic and proteomic responses in circulating neutrophils required for pathogen containment.Sci. Adv.10,eadl1710(2024).DOI:10.1126/sciadv.adl1710" href="https://doi.org/10.1126/sciadv.adl1710">Rainer Kaiser et al., Peripheral priming induces plastic transcriptomic and proteomic responses in circulating neutrophils required for pathogen containment. Sci. Adv. 10, eadl1710 (2024). DOI:10.1126/sciadv.adl1710</a></p> </div> </div> <p><strong>Original Data</strong></p> <p>The following files contain all count data for the original data of this manuscript:</p> <p>sepsis1_raw_feature_bc_matrix.h5 -> raw feature barcode matrix for sepsis1 sequencing<br>sepsis1_velocyto.loom -> velocyto matrices for sepsis1<br>sepsis2_raw_feature_bc_matrix.h5 -> raw feature barcode matrix for sepsis2 sequencing<br>sepsis2_velocyto.loom -> velocyto matrices for sepsis2</p> <p>sepsis_seurat.rds -> processed Seurat object containing original data cells. (Upd.: the meta.data-column "cellnames" contains the cell type annotation given in Figure 1B)</p> <p><strong>Original Data Scripts</strong></p> <p>process.R -> main analysis script<br>functions.R -> helper functions for main analysis script<br>enrichmentAnalysis.R -> script running the enrichment analysis<br>process_wgcna.R -> script performing the wgcna analysis<br>velocities_step1.R -> script performing velocity analysis (from seurat to data matrices)<br>velocities_step2.py -> actual velocity analysis<br><br><strong>GSE137539 Re-Analysis</strong></p> <p>gse137539_processed.Rds -> processed seurat object<br>gse137539_process.R -> analysis script</p> <p><strong>Bulk Analysis</strong></p> <p>MOUSE_SEPTIC_SEPTICACT.inex.DirectDESeq2.xlsx-> Raw UMI counts (intronic+exonic from zUMIs) and DE genes<br>MOUSE_SEPTIC_SEPTICACT.inex.DirectDESeq2.tsv.GeneOntology.BP.up.gsea.tsv -> Gene Set Enrichment Analysis on up-regulated genes using Gene Ontology Biological Process</p>
The Cellular and Extra-Cellular Proteomic Signature of Human Dopaminergic Neurons Carrying the LRRK2 G2019S Mutation
<p>The provided data sets were generated during the work described in "The Cellular and Extra-Cellular Proteomic Signature of Human Dopaminergic Neurons Carrying the LRRK2 G2019S Mutation" published in Frontiers in Neuroscience, 2024.</p> <p>They contain the results from our differential expression analysis of our raw DIA Protemic Data as well as a list of input data for GO analyses.</p>
Spatial-DC: a robust deep learning-based method for deconvolution of spatial proteomics
<p>The processed reference and spatial proteomics datasets, along with the processed mIHC imaging data of mouse PDAC tissue are available in the repository.</p> <p>Also, the source code for pre-processing, data analysis, and generating figure and tables has been deposited in both GitHub [<a href="https://github.com/TencentAILabHealthcare/Spatial-DC">https://github.com/TencentAILabHealthcare/Spatial-DC</a>] and Zenodo [<a href="https://doi.org/10.5281/zenodo.14386585">https://doi.org/10.5281/zenodo.14386585</a>].</p> <p> </p>
SWAT proteomics study
<p>CSV file: Prosessed de-identified proteomics data of a sub-cohort of pateints from the SWAT observational study</p> <p>zip folder: SomaLogic adat file and measurment infomration on the same study</p>
Proteome-wide Prediction of the Functional Impact of Missense Variants with ProteoCast
<p><strong>Proteome-wide Prediction of the Functional Impact of Missense Variants with ProteoCast</strong></p> <p>This dataset contains mutation effect predictions for 22,169 <em>Drosophila melanogaster</em> protein isoforms, classifying over <strong>293 million </strong>amino acid substitutions as neutral, uncertain, or impactful. The predictions were generated using the evolution-based <strong>GEMME </strong>model (E.Laine et <em>al.</em> MBE 2019) with multiple sequence alignments (MSAs) from the highly efficient <strong>ColabFold</strong> protocol (M.Mirdita et<em> al.</em> NatMet 2022, Abakarova et al. <em>GBE</em> 2023). To ensure reliability, we provide<strong> </strong>global (per-protein) and local (per-residue) confidence metrics, since the predictions are sensitive to the input MSA quality.</p> <p>Predictions were validated using natural polymorphisms from the Drosophila Genetic Reference Panel (DGRP) and Drosophila Evolution over Space and Time (DEST2) datasets, as well as FlyBase’s developmentally lethal and hypomorphic mutations. Additionally, the dataset includes sensitivity data for post-translational modifications (PTMs) and short linear motifs (SLiMs), aiding functional site identification.</p> <p>All this data can be visualized at <a href="https://proteocast.ijm.fr/drosophiladb/">proteocast.ijm.fr</a>. </p> <p>Readme:</p> <ol> <li><strong>Drosophila_ProteoCast.tar.gz </strong>- this archive contains ProteoCast predictions and analysis for each unique proteoform, with folder names corresponding to IDs listed in the mapping_database.csv file. A detailed description of the folder structure and contents can be found in ReadMe.txt. </li> <li><strong>data.tar.gz </strong>- this archive contains the data used in this study, sourced from FlyBase, DGRP2, and DEST2.</li> <li><strong>Dmel6.44PredictionsRecap.csv</strong> - this summary file provides detailed information for each proteoform, all FlyBase protein IDs (<em>FBpp_ID</em>) included. It contains the following data: <ul> <li><strong>Identifiers</strong>: FlyBase protein ID (<em>FBpp_ID</em>), protein symbol (Protein_symbol), gene ID (<em>FBgn_ID</em>), transcript ID (<em>FBtr_ID</em>), and UniProt ID if available (<em>UniProt_ID</em>).</li> <li><strong>Protein Characteristics</strong>: Sequence length (<em>Length</em>) and whether the proteoform is representative (<em>Representative_FBpp</em>).</li> <li><strong>MSA and GEMME Predictions</strong>: Fraction of observed mutations (<em>F_obs</em>) and number of sequences (<em>Nb_seq_MSA</em>) in the ColabFold MSA, presence or absence of GEMME predictions (<em>GEMME_prediction</em>), and global confidence score (<em>GlobalConfidence</em>).</li> <li><strong>Mutation Classification</strong>: Thresholds for defining mutations as neutral, uncertain, or impactful (<em>GMM3_uncertain</em>, <em>GMM3_impactful</em>).</li> <li><strong>Genomic Information</strong>: DNA strand (<em>Strand</em>) and exon coordinates (<em>Exons_coordinates</em>).</li> <li><strong>Structural Data</strong>: 3D structure file name if available (<em>Structure_3D_file, Structure_3D</em>).</li> <li><strong>Mutation Counts</strong>: Number of analyzed mutations and affected residues, labelled as lethal, hypomorphic on FlyBase, or from the DEST2 and DGRP datasets (<em>n_Lethal, n_Lethal_res, n_Hypomorphic, n_Hypomorphic_res, n_DEST2, n_DEST2_res, n_DGRP, n_DGRP_res, n_DEST_DGRP_union, n_DEST_DGRP_union_res</em>).</li> </ul> </li> <li><strong>csv.tar.gz </strong>-<strong> </strong>this archive contains the files generated in this study. A detailed description of the folder structure and contents can be found in ReadMe.txt.</li> </ol>
Blood Cell Transcriptomics and Proteomics of Axial Spondyloarthritis patients undergoing adalimumab treatment
<p>This study aims at identifying molecular biomarkers differentiating good responders and non-responders to treatment with TNF inhibitors (TNFi), among patients with axial spondyloarthritis (axSpA). Publication: <em>Biomolecules</em> <strong>2024</strong>, <em>14</em>(3), 382; <a href="https://doi.org/10.3390/biom14030382">https://doi.org/10.3390/biom14030382</a></p>
Whole-cell modeling in yeast predicts compartment-specific proteome constraints that drive metabolic strategies
<p>The pcYeast7.6 model and files, required to reproduce the figures, provided in the publication "Whole-cell modeling in yeast predicts compartment-specific proteome constraints that drive metabolic strategies", accepted in <em>Nat Commun</em>. The Zenodo upload was created by Pranas Grigaitis, p.grigaitis [at] vu.nl.</p> <p>Abstract</p> <p>When conditions change, unicellular organisms rewire their metabolism to sustain cell maintenance and cellular growth. Such rewiring may be understood as resource re-allocation under cellular constraints. Eukaryal cells contain metabolically active organelles such as mitochondria, competing for cytosolic space and resources, and the nature of the relevant cellular constraints remain to be determined for such cells. Here we present a comprehensive metabolic model of the yeast cell, based on its full metabolic reaction network extended with protein synthesis and degradation reactions. The model predicts metabolic fluxes and corresponding protein expression by constraining compartment-specific protein pools and maximising growth rate. Comparing model predictions with quantitative experimental data suggests that under glucose limitation, a mitochondrial constraint limits growth at the onset of ethanol formation - known as the Crabtree effect. Under sugar excess, however, a constraint on total cytosolic volume dictates overflow metabolism. Our comprehensive model thus identifies condition-dependent and compartment-specific constraints that can explain metabolic strategies and protein expression profiles from growth rate optimization, providing a framework to understand metabolic adaptation in eukaryal cells.</p>
Data from: Combined analysis of micro RNA and proteomic profiles and interactions in patients with primary lung adenocarcinoma and lung adenocarcinoma brain metastases
<p>We carried out an analysis of miRNAs expression profiles and protein spectrums of non-metastatic primary lung adenocarcinoma (LP) and patients with brain metastases (BM) to better explore the molecular basis of BM. Files containing raw data of miRNA expression and proteomic profiles in the manuscript "Combined analysis of micro RNA and proteomic profiles and interactions in patients with primary lung adenocarcinoma and lung adenocarcinoma brain metastases".</p>
Mass spectrometry proteomics data obtained from analysis of the secretome of Anisakis simplex (sensu stricto) L3 larvae.
<p>Mass spectrometry proteomics data obtained from analysis of the secretome of <em>Anisakis simplex</em> (sensu stricto) L3 larvae.</p>
VESPAl Prediction of the human proteome (downloaded 22/01/17)
<p>VESPAl Prediction of the human proteome (downloaded 22/01/17) generated with https://github.com/Rostlab/VESPA.</p> <p>For details on VESPAl see</p> <blockquote> <p>Marquet, C., Heinzinger, M., Olenyi, T. <em>et al.</em> Embeddings from protein language models predict conservation and variant effects. <em>Hum Genet</em> (2021). https://doi.org/10.1007/s00439-021-02411-y</p> </blockquote> <p>3 proteins (Q8WZ42,Q8WXI7,Q8NF91) of the human proteome were too long to generate ProtT5 embeddings (Elnaggar et al. 2021). VESPAl predictions are therefore available for 20357 proteins.</p>
Recent speciation and hybridization in Icelandic deep-sea isopods: An integrative approach using genomics and proteomics.
<p>The crustacean marine isopod species <i>Haploniscus bicuspis</i> (G.O. Sars, 1877) shows circum-Icelandic distribution in a wide range of environmental conditions and along well-known geographic barriers, such as the Greenland-Iceland-Faroe (GIF) Ridge. We wanted to explore population genetics, phylogeography and cryptic speciation as well as to investigate whether previously described, but unaccepted subspecies have any merit. Using the same set of specimens, we combined mitochondrial COI sequences, thousands of nuclear loci (ddRAD), and proteomic profiles, plus selected morphological characters using Confocal Laser Scanning Microscopy (CLSM). Five divergent genetic lineages were identified by COI and ddRAD, two south and three north of the GIF Ridge. Assignment of populations to the three northern lineages varied and detailed analyses revealed hybridization and gene flow between them, suggesting a single northern species with a complex phylogeographic history. No apparent hybridization was observed among lineages south of the Ridge, inferring the existence of two more species. Differences in proteomic profiles between the three putative species were minimal, implying an ongoing or recent speciation process. Population differentiation was high, even among closely associated populations, and higher in mitochondrial COI than nuclear ddRAD loci. Gene flow is apparently male-biased, leading to hybrid zones and instances of complete exchange of the local nuclear genome through immigrating males. This study did not confirm the existence of subspecies defined by male characters, which probably characterize different male developmental stages.</p>
Comparative proteomic and metabolomic analyses of plasma reveal the novel biomarker panels for thyroid dysfunction
<p><strong>Abstract</strong><strong>:</strong></p> <p><em>Objectives: </em>Thyroid dysfunction such as hypothyroidism (THO) and hyperthyroidism (THE) are the disease caused by pathological processes in the thyroid. The current diagnosis of thyroid dysfunction is variable because of ages and genders. The aim of this study was to explore the novel candidate biomarker panels for hypothyroidism and hyperthyroidism screening with mass spectrometry and bioinformatics.</p> <p><em>Methods:</em> Plasma samples were collected from 15 THE patients, 9 THO patients, and 15 healthy controls. DIA-based proteomic and untargeted metabolomic analyses were performed to identify the novel biomarker panels for THO and THE. Finally, three candidate biomarkers were verified by ELISA in 34 samples.</p> <p><em>Results:</em> A total of 2738 proteins and 6103 metabolites were identified, and 173 proteins and 2487 metabolites were found to be differentially expressed among THE, THO and control groups. The results of the ensemble feature selection, K-means clustering and the least absolute shrinkage and selection operator (LASSO) regression model showed that four proteins (C4A, C3/C5 convertase, APOL1, and ITIH4) and four metabolites (L-arginine, L-proline, cortisol, and cortisone) identified by plasma proteomics and metabolomics could help distinguish THO and THE patients from healthy controls.</p> <p><em>Conclusions:</em> This study identified and verified two pairs of biomarker panels that can distinguish the THE and THO patients regardless of ages and genders. Consequently, our findings represent a comprehensive analyses of thyroid dysfunction plasma, which is significant for the clinical diagnosis.</p> <p> </p>
Input features and benchmark data sets for protein complex prediction and E. coli proteome application by AF2Complex
<p>Benchmark data sets of AF2Complex, input features for application to E. coli proteome, and predicted structural models of E. coli Ccm I as described in</p> <p><strong>Predicting direct physical interactions in multimeric proteins with deep learning</strong></p> <p><em>Mu Gao, Davi Nakajima An, Jerry M. Parks, Jeffrey Skolnick</em></p> <ol> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/af2complex_bench.tar.gz">af2complex_bench.tar.gz</a>: Benchmark data sets CP17, Dimer1193 and Oligomer562, including input features for AF2Complex/AF-Multimer, both paired and unpaired MSAs, as well as sequences, experimental structures, and results presented in the AF2Complex work (~90GB de-compressed size)</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_Ccm_I.tar.gz">ecoli_Ccm_I.tar.gz</a>: Computational models of the<em> E. coli</em> Ccm I system</li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_set.tar.gz">ecoli_sets.tar.gz</a>: Lists of benchmark sets of positive and negative PPIs from <em>E. coli</em></li> <li><a href="https://zenodo.org/api/files/087ae188-6f7a-4a71-a586-bbc7bdcc6843/ecoli_af_fea.tar.gz">ecoli_af_fea.tar.gz: </a>Pre-generated input features of <em>E. coli</em> proteome for protein complex prediction and modeling by AF2Complex (4,429 proteins, ~800 GB de-compressed size). This data set can be used with AF2Complex to probe the interactions of any combinations among the 4,429 proteins of E. coli.</li> </ol> <p> </p>
LC-MS raw data for proteomic elucidation of the targets and primary functions of picornavirus 2A protease
<p>This dataset contains LC-MS raw data for pulldowns from the project "Proteomic elucidation of the targets and primary functions of picornavirus 2A protease." The descriptions of the raw data files are in Summary_Table_MS_RawData.pdf.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.