Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,448
datasets available to search
ShareScore release 0.7.1
Dataset results
1,448 results for “proteomic”
MALDI-TOF MS data: Species delimitation of Hexacorallia and Octocorallia around Iceland using nuclear and mitochondrial DNA and proteome fingerprinting
<p>Cold-water corals build up reef structures or coral gardens and play an important role for many organisms in the deep sea. Climate change, deep-sea mining, and bottom trawling are severely compromising these ecosystems, making it all the more important to document the diversity, distribution, and impacts on corals. This goes hand in hand with species identification, which is morphologically and genetically challenging for Hexa- and Octocorallia. Morphological variation and slowly evolving molecular markers both contribute to the difficulty of species identification. In this study, a fast and cheap species delimitation tool for Octocorallia and Scleractinia of the Northeast Atlantic was tested based on 49 specimens. Two nuclear markers (ITS2 and 28S rDNA) and two mitochondrial markers (COI and mtMutS) were sequenced. The sequences formed the basis of a reference library for comparison to the results of species delimitation based on proteomic analysis using the MALDI-TOF MS method. The genetic methods were able to distinguish 17 of 18 presumed species. The MALDI-TOF MS method was able to distinguish 7 species. Species that could not be distinguished from one another still achieved good signals but were not represented by enough specimens for comparison. Therefore, it is predicted that with an extensive reference library of proteome spectra for Scleractinia and Octocorallia, MALDI-TOF MS may provide a rapid and cost-effective alternative for species discrimination in corals.</p>
Genomic, transcriptomic and proteomic comparison of MRSA CC398 isolates collected from human and wild animal samples (Genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) for the following methicillin-resistant <em>Staphylococcus aureus</em> (MRSA) strains: MRSA CC398 isolates recovered from humans, namely C5621 and C9017, and from a wild boar, namely OR418.</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB35102).</p>
The proteomics data of Paris polyphylla var. yunnanensis
<p>The mass spectrometry proteomics data and the protein/peptide identifications of seed coat, mature seed and germianting seeds from<strong><em> </em></strong><em>Paris polyphylla</em> var. <em>yunnanensis.</em></p>
Proteomic analysis of serum markers in patients maintained on Antipsychotics
<p><strong>Background:</strong> Schizophrenia (SZ) and bipolar disorder (BD) share many features: overlap in mood and psychotic symptoms, common genetic predisposition, treatment with antipsychotics (APs), and similar metabolic comorbidities. The pathophysiology of both is still not well defined, and no biomarkers can be used clinically for diagnosis and management. This study aimed to assess the plasma proteomics profile of patients with SZ and BD maintained on APs compared to those who had been off APs for six months and to healthy controls (HCs).</p> <p><strong>Methods:</strong> We analyzed the data using functional enrichment, random forest modeling to identify potential biomarkers, and multivariate regression for the associations with metabolic abnormalities. </p> <p><strong>Results:</strong> We identified several proteins known to play roles in the differentiation of the nervous system like NTRK2, CNTN1, ROBO2, and PLXNC1, which were downregulated in AP-free SZ and BD patients but were "normalized" in those on APs. Other proteins (like NCAM1 and TNFRSF17) were "normal" in AP-free patients but downregulated in patients on APs, suggesting that these changes are related to medications' effects. We found significant enrichment of proteins involved in neuronal plasticity, mainly in SZ patients on APs. Most of the proteins associated with metabolic abnormalities were more related to APs use than having SZ or BD. The biomarkers identification showed specific and sensitive results for schizophrenia, where two proteins (PRL and MRC2) produced adequate results. </p> <p><strong>Conclusions:</strong> Our results confirmed the utility of blood samples to identify protein signatures and mechanisms involved in the pathophysiology and treatment of SZ and BD.</p>
Supplementary code and data for: Inferring differential subcellular localisation in comparative spatial proteomics using BANDLE
<p>This repository contains code and data to reproduce the figures in the manuscript: Inferring differential subcellular localisation in comparative spatial proteomics using BANDLE.</p> <p>Please refer to the readme in the repository. </p>
Supplementary table of PRIDE datasets analyzed for "FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data"
<p>Our proteomics dataset comes from The PRoteomics IDEntifications (PRIDE) database, the world’s largest data repository of mass spectrometry-based proteomics data. Specifically, we used 633 human proteomics project experiments with a total of 32,546 runs and reanalyzed them using ionbot with an FDR threshold of 0.01 [16], resulting in a total of 154,885,151 peptide spectrum matches for 18,846 proteins. Here is the full list of projects, runs, and general statistics.</p>
Integrated methylome and phenome study of the circulating proteome reveals markers pertinent to brain health
<p>This repository houses fully-adjusted methylome-wide association study (MWAS) summary statistics for 4,231 SomaScan protein measurements. These were generated as part of the study titled ‘Integrated methylome and phenome study of the circulating proteome reveals markers pertinent to brain health’ by Gadd <em>et al</em>. The Stratifying Resilience and Depression Longitudinally (STRADL) cohort used in this study is a subset of individuals from Generation Scotland: The Scottish Family Health Study. There were 744 individuals with complete protein and DNA methylation measurements available at 772,619 CpG probes. MWAS were performed with protein residuals as the outcome and DNA methylation as the exposure, using the Omics-data-based complex trait analysis (OSCA) software.</p> <p>Fully-adjusted models were run using M-values that were adjusted for age, sex, DNA methylation-derived immune cell estimates, depression status, DNA methylation batch and set, body mass index and a DNA methylation-derived smoking score. Protein levels were rank-based inverse normalised and scaled to have a mean of 0 and standard deviation of 1. Protein levels were residualised by age, sex, available pQTLs, technical covariates and 20 genetic principal components.</p> <p>Four of the 4,235 protein MWAS models did not converge (15509-2 - NAGLU, 15584-9 - CFHR2, 4407-10 - MST1 and 6402-8 - PILRA). Therefore, summary statistics are provided for 4,231 protein levels.</p> <p>Each protein MWAS summary statistics file has been saved with the following naming system: "MWAS_SeqId_Protein_gene.csv". For example, the protein with gene name CRYBB2 and SeqId 10000-28 has the following file name: "MWAS_10000-28_CRYBB2.csv".</p> <p>The SeqIds, UniProt codes, gene names and full UniProt names can be found in "annotation_formatted_for_paper.csv" and the full summary statistics are found within "compressed-protein-ewas.tar.gz".</p> <p>Please contact either <a href="mailto:riccardo.marioni@ed.ac.uk">riccardo.marioni@ed.ac.uk</a> or <a href="mailto:danni.gadd@ed.ac.uk">danni.gadd@ed.ac.uk</a> for any queries. All code is available at the following Github repository: <a href="https://github.com/DanniGadd/Epigenome-and-phenome-wide-study-of-brain-health-outcomes">https://github.com/DanniGadd/Epigenome-and-phenome-wide-study-of-brain-health-outcomes</a>.</p>
Supplementary data for: Effects of thermal acclimation on the proteome of the planarian Crenobia alpina from an alpine freshwater spring
<p>Species' acclimation capacities and their ability to maintain molecular homeostasis outside of ideal temperature ranges will partly predict their success following climate-change induced thermal regime shifts. Theory predicts that ectothermic organisms from thermally stable environments have muted plasticities, and that these species <span>may be</span> particularly vulnerable to temperature increase. Whether such species retained or lost acclimation capacities remains largely unknown. We studied proteome changes in the planarian <em>Crenobia alpina</em>, a prominent member of cold-stable alpine habitats that is considered to be cold-adapted stenotherm. We found that the species' CT<sub>max</sub> is above its experienced habitat temperatures and that different populations exhibit differential CTmax acclimation capacities, whereby an alpine population showed reduced plasticity. In a separate experiment, we acclimated <em>C. alpina</em> individuals from the alpine population to 8, 11, 14, or 17°C over the course of 168 h and compared a comprehensively annotated species-specific proteome. Network analyses of 3399 proteins and protein set enrichment show that while the species' proteome is overall stable across these temperatures, protein sets functioning in oxidative stress response, mitochondria, protein synthesis and turnover are lower abundant following warm acclimation. Proteins associated with an unfolded protein response, ciliogenesis, tissue damage repair, development, and the innate immune system were higher abundant following warm acclimation. Our findings suggest that this species has not suffered DNA decay (e.g., loss of heat-shock proteins) during evolution in a cold-stable environment and retained plasticity in response to elevated temperatures, challenging the notion that stable environments necessarily result in muted plasticity.</p>
Integrated plasma proteomic and single-cell immune signaling network signatures demarcate mild, moderate, and severe COVID-19
<p>The biological determinants underlying the range of COVID-19 clinical manifestations are not fully understood. Here, over 1400 plasma proteins and 2600 single-cell immune features comprising cell phenotype, endogenous signaling activity, and signaling responses to inflammatory ligands are cross-sectionally assessed in peripheral blood from 97 patients with mild, moderate, and severe COVID-19 and 40 uninfected patients. Using an integrated computational approach to analyze the combined plasma and single-cell proteomic data, we identify and independently validate a multivariate model classifying COVID-19 severity (multi-class AUC<sub>training</sub> = 0.799, p-value = 4.2e-6; multi-class AUC<sub>validation</sub> = 0.773, p-value = 7.7e-6). Examination of informative model features reveals novel biological signatures of COVID-19 severity, including the dysregulation of JAK/STAT, MAPK/mTOR, and NF-κB immune signaling networks in addition to recapitulating known hallmarks of COVID-19. These results provide a set of early determinants of COVID-19 severity that may point to therapeutic targets for prevention and/or treatment of COVID-19 progression.</p>
Proteomic characterization of Toxoplasma gondii ME49 derived strains resistant to the artemisinin derivatives artemiside and artemisone
<p>Full datasets of proteomes of artemisone (GC003) and artemiside (GC008) resistant T. gondii Me49 strains.</p> <p>Paper submitted to International Journal of Parasitology-Drugs and Drug Resistance</p> <p>Differential proteomic analysis of the artemisone (GC003<sup>R</sup>) and artemiside (GC008<sup>R</sup>) resistant strains versus their corresponding <em>T. gondii</em> ME49 wildtype yielded 3977 unique peptides matching to 733 <em>T. gondii</em> proteins. The complete dataset is available as supplemental Table S1. A more detailed analysis revealed that 215 proteins were significantly downregulated in GC003R and 8 proteins in GC008R as compared to their wildtype ME49. Two proteins were downregulated in both strains. No proteins were upregulated in the resistant strains as compared to their corresponding wildtype. The complete list of the differentials is given as supplemental Table S2. </p>
Data from: Proteomic fingerprinting enables quantitative biodiversity assessments of species and ontogenetic stages in Calanus congeners (Copepoda, Crustacea) from the Arctic Ocean
<p><span>Species identification is pivotal in biodiversity assessments, and proteomic fingerprinting by MALDI-TOF mass spectrometry has already been shown to reliably identify calanoid copepods to species level. However, MALDI-TOF data may contain more information beyond mere species identification. In this study, we investigated different ontogenetic stages (copepodids C1-C6 females) of three co-occurring <em>Calanus</em> species from the Arctic Fram Strait, which cannot be identified to species level based on morphological characters alone. Differentiation of the three species based on mass spectrometry data was without any error. In addition, a clear stage-specific signal was detected in all species, supported by clustering approaches as well as machine learning using Random Forest. More complex mass spectra in later ontogenetic stages as well as relative intensities of certain mass peaks were found as the main drivers of stage distinction in these species. Through a dilution series, we were able to show that this did not result from the higher amount of biomass that was used in tissue processing of the larger stages. Finally, the data were tested in a simulation for application in a real biodiversity assessment by using Random Forest for stage classification of specimens absent from the training data. This resulted in a successful stage-identification rate of almost 90%, making proteomic fingerprinting a promising tool to investigate polewards shifts of Atlantic <em>Calanus</em> species and, in general, to assess stage compositions in biodiversity assessments of Calanoida, which can be notoriously difficult using conventional identification methods.</span></p>
Data from: Evaluating species richness using proteomic fingerprinting and DNA-barcoding – a case study on meiobenthic copepods from the Clarion Clipperton Fracture Zone
<p><span>The Clarion Clipperton Fracture Zone (CCZ) is a vast deep-sea region harboring a highly diverse benthic fauna, which will be affected by potential future deep-sea mining of metal-rich polymetallic nodules. Despite the need for conservation plans and monitoring strategies in this context, the majority of taxonomic groups remains scientifically undescribed. However, molecular rapid assessment methods such as DNA-barcoding and Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) provide the potential to accelerate specimen identification and biodiversity assessment significantly in the deep-sea areas. In this study, we successfully applied both methods to investigate the diversity of meiobenthic copepods in the eastern CCZ, including the first application of MALDI-TOF MS for the identification of these deep-sea organisms. Comparing several different species delimitation tools for both datasets, we found that biodiversity values were very similar, with Pielou's Evenness varying between 0.97 and 0.99 in all datasets. Still, direct comparisons of species clusters revealed differences between all techniques and methods, which are likely caused by the high number of rare species being represented by only one specimen, despite our extensive dataset of more than 2000 specimens. Hence, we regard our study as a first approach toward setting up a reference library for mass spectrometry data of the CCZ in combination with DNA-barcodes. We conclude that proteome fingerprinting, as well as the more established DNA-barcoding, can be seen as a valuable tool for rapid biodiversity assessments in the future, even when no reference information is available.</span></p>
Proteomic Profiling for Identification of Animal Skin Species in Ancient Egyptian Archaeological Leather using Liquid Chromatography Coupled with Tandem Mass Spectrometry (Nano LC-MS/MS)
<p><strong>Proteomic Profiling Dataset for Identification of Animal Skin Species in Ancient Egyptian Archaeological Leather using Liquid Chromatography Coupled with Tandem Mass Spectrometry (Nano LC-MS/MS)</strong></p>
Mining Proteomic Databases of a Model Plant Medicago truncatula for Mosaic Proteins and Other Unconventional Translation Products
<p><strong>How many different proteins can be produced from a single spliced transcript? Genome annotation projects do not consider the coding potential of reading frames other than that of the reference open reading frames (refORFs). Recently, alternative open reading frames (altORFs) and their translational products, alternative proteins (altProts), have been shown to carry out important functions in various organisms. Overlapping altORFs may be involved in one fundamental mechanism so far overlooked. A few years ago, it was proposed that altORFs may act as building blocks for chimeric (mosaic) polypeptides, which are produced via multiple ribosomal frameshifting events from a single mature transcript. We adopt terminology from that earlier discussion and call this mechanism mosaic translation. This way of extracting and combining genetic information may significantly increase proteome diversity. Thus, we hypothesize that this mechanism may have contributed to the flexibility and adaptability of organisms to a variety of environmental conditions. The idea of mosaic translation is a testable hypothesis, although its direct demonstration is technically very challenging. If confirmed, this concept will revolutionize modern genetics. In this project, we would like to follow a unique strategy for the detection of mosaic proteins in proteomic databases publicly available for a very important model plant <em>Medicago truncatula</em>. The proposed analysis will be based on our own preliminary data already generated in the course of an ongoing TÜBİTAK1002 project. Regardless of whether the evidence for mosaic translation is found in this study, this effort will help identify such proteins later when more proteomic data become available. Finally, our approach can reveal unconventional frameshifting products that derive from the omission of several nucleotides by ribosomes (for example, +2 to +16 frameshifts). Regardless of whether such frameshifted products are parts of mosaic proteins, the potential for their detection makes this project very novel, because frameshifts longer than one nucleotide in the forward direction have not been described so far.</strong></p>
Proteomics LC-MS/MS test dataset for protein quantitation via stable isotope labelling
<p>The provided mzML file can be used as a test dataset for protein identification and quantitation software. It was generated from human embryonic kidney (HEK) cells that were either unlabelled or labelled with heavy SILAC (K6R6, unimod accession 188, PSI-MS Name: "Label:13C(6)"). Apart from different labelling, the HEK cells were kept in exactly the same conditions and harvested simultaneously. Light and heavy labelled proteins from HEK cell lysate were mixed in a certain ratio, digested with Trypsin and measured on a ThermoFisher QExactive mass spectrometer. A more detailed description on the generation of the dataset will soon be accessible at PRIDE.</p> <p>The provided mzML file has been converted from Thermo RAW and slightly modified via msConvert (ProteoWizard). To reduce the filesize and to speed up analysis, it has further been filtered to contain only the data measured between 2,000 sec and 3,000 sec of the original LC-MS/MS run.</p>
Unveiling the autoreactome: Proteome-wide immunological fingerprints reveal the promise of plasma cell depleting therapy
<p>The prevalence and burden of autoimmune and autoantibody mediated disease continues to rise, yet the etiologies of many of these diseases remain unclear. Despite numerous new targeted immunomodulatory therapies, comprehensive approaches to apply and evaluate the effects of these treatments longitudinally are lacking. Here, we leverage advances in PhIPseq methodology to explore the modulation, or lack thereof, for autoreactive antibodies proteome-wide in both health and disease. We demonstrate that each individual, regardless of disease state, possesses a distinct set of autoreactivities constituting a unique immunological fingerprint, or "autoreactome", that is remarkably stable over years. In addition to uncovering important new biology, the autoreactome can be used to better evaluate the relative effectiveness of various therapies in altering autoantibody repertoires. We find that therapies targeting B-Cell Maturation Antigen (BCMA) profoundly alter an individual's autoreactome, while anti-CD19 and CD-20 therapies have minimal effects, strongly suggesting a rationale for BCMA or other plasma cell targeted therapies in autoantibody mediated diseases.</p>
HDCA fetal lung spatial proteomics example datasets
<p>The data provided in this repository is published alongside <a href="https://doi.org/10.1101/2024.01.25.577163" target="_blank" rel="noopener noreferrer">this preprint</a> titled as 'High-parametric protein maps reveal the spatial organization in early-developing human lung' and <a href="https://github.com/CellProfiling/HDCA-FetalLung-SpatialProteomics" target="_blank" rel="noopener noreferrer">this code repository</a>. The preprint and GitHub repository provide further metadata and analysis information. When using the data in this repository, please cite the preprint under DOI: <a href="https://doi.org/10.1101/2024.01.25.577163" target="_blank" rel="noopener noreferrer">https://doi.org/10.1101/2024.01.25.577163</a>.</p>
Resolving single-cell expression profiles by pseudo-temporal integration of transcriptomic and proteomic datasets.
<p>Raw and processed single cell proteomics (scp-MS) and scRNA-Seq data of HEK293-PIP-FUCCI cells which were challanged with hypoxia. The repository contains data for recreating the pseudo-temporal alignment analysis of transcription-translation profiles.</p>
Proteomic analysis of Entamoeba histolytica in vivo assembled pre-mRNA splicing complexes
<p>Dataset for publication "Proteomic analysis of <em>Entamoeba histolytica</em> in vivo assembled pre-mRNA splicing complexes"</p>
Bridging the Chromosome-Centric and Biology and Disease Human Proteome Projects: Accessible and automated tools for interpreting biological and pathological impact of protein sequence variants detected via proteogenomics
<p>Bridging the Chromosome-Centric and Biology and Disease Human Proteome Projects: Accessible and automated tools for interpreting biological and pathological impact of protein sequence variants detected via proteogenomics</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.