Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,448

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,448 results for “Proteomics”

Learn how ShareScore rates datasets ↗
zenodo52/100

Dataset - Decrypting lysine deacetylase inhibitor action and protein modifications by dose-resolved proteomics

<h4><strong>Dataset Summary</strong></h4> <p>Lysine deacetylase inhibitors (KDACis) are approved for cutaneous T-cell lymphoma (CTCL), peripheral T-cell lymphoma (PTCL), and multiple myeloma. Despite the mechanism of action(s) (MoA) remains elusive, these inhibitors lead to increasing acetylation levels of histones and other proteins, altered gene expression and cell death. To characterize the MoA of these drugs in more detail, we systematically measured dose-dependent changes in protein expression, acetylation, and phosphorylation in response to 21 clinical and pre-clinical KDACis. MV4-11 cells were treated for 6 h with 1 vehicle control and 10 increasing doses of the respective drug (from 100 pM to 30 mM). Proteins were digested with trypsin, and the resulting 11 peptide preparations corresponding to one drug dose each were encoded by stable isotopes (tandem mass tags, TMT-11plex) and combined. Acetylated peptides were subsequently enriched by immunoprecipitation and phosphopeptides by immobilized metal affinity chromatography (IMAC). PTM-carrying and unmodified peptides were analyzed separately by liquid chromatography tandem mass spectrometry (LC-MS/MS) for peptide and protein identification and quantification. Additionally, Vorinostat and Panobinostat were also recorded as time-dependent experiments at their pEC50 concentration, respectively.&nbsp;</p> <h4><strong>Dataset structure</strong></h4> <p>Here, we provide all curve data processed with CurveCurator v0.4.0 (<a href="https://github.com/kusterlab/curve_curator">https://github.com/kusterlab/curve_curator</a>). Each drug is a zip folder containing acetylome, phosphoproteome, and fullproteome data. Next to each data set is the toml parameter file used to generate the curves.txt and dashboard.html files. Time-dependent data is indicated by "td" and dose-dependent data is indicated by "dd".</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Haloferax volcanii theoretical proteome exported from HaloLex

<p>The theoretical proteome is derived from the annotated Haloferax volcanii DS2 genome (Hartmann et al, 2010).</p> <p>This proteome represents a Gold Standard Protein based annotation (Pfeiffer and Oesterhelt, 2015).</p> <p>Proteome data from Haloferax volcanii were comprehensively analyzed in a community-based effort (Schulze et al, 2020)</p> <p>Various proteome studies are based on versions of this theoretical proteome. Different versions, which are cited in proteome papers, are made available in this series of Zenodo uploads.</p>

opencc-by-4.0Jan 2017View details →
zenodo52/100

Reprocessing of the dataset "Plasma Proteome Profiling Reveals the Effects of Weight Loss on the Apolipoprotein Family and Systemic Inflammation Status"

<p>Reprocessing of the MassIVE repository MSV000080596, originally generated to investigate the dynamic changes in the plasma proteomes of a cohort of individuals with obesity following weight loss and maintenance. The reprocessing included all samples from 52 individuals&nbsp;taken right after the weight-loss process and during the weight maintenance phase of the study (Weeks 0, 4, 13, 26, 39, and 52).</p> <p>We used the sequence database generated by ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) representing all populations from the 1000 Genomes Project (doi.org/10.5281/zenodo.10149277). For the search, SearchGUI version 4.3.1 and PeptideShaker version 3.0.0 were used with the X!Tandem and Tide search engines. The modification settings specified were carbamidomethylation of C as fixed and oxidation of M, deamidation of N and Q, Pyrrolidone of E and Q, and acetylation of protein N-terminus as variable modifications. The maximum peptide length was set to 40 amino acids and the precursor and fragment ion tolerances were set to 7 and 20 ppm, respectively. Resulting PSMs were processed as described in (doi.org/10.1021/acs.jproteome.3c00243) using Percolator version 3.5 provided with features based on peptide retention time (DeepLC version 1.1.2) and fragmentation predictors (MS2PIP version 3.9.0), and filtered at a 1% estimated FDR.</p> <p>The attached file contains all the peptide-spectrum matches identified at 1% FDR. The peptides have been annotated with transcripts, genes, and alleles using the ProHap Peptide Annotator v1.1 (<a href="https://github.com/ProGenNo/ProHap_PeptideAnnotator">https://github.com/ProGenNo/ProHap_PeptideAnnotator</a>).</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Natrialba magadii theoretical proteome exported from HaloLex

<p>The theoretical proteome is derived from the annotated Natrialba magadii&nbsp;ATCC 43099 genome (Siddaramappa et al, 2012).</p> <p>This proteome represents a Gold Standard Protein based annotation (Pfeiffer and Oesterhelt, 2015).</p> <p>This proteome was exported from HaloLex (Pfeiffer et al, 2008)</p> <p>A proteome study is based on this version of the theoretical proteome (Cerletti et al, 2018).</p>

opencc-by-4.0Dec 2019View details →
zenodo48/100

Darwin: an amino acid sequence collection of complete proteomes from eukaryotes with different phylogenetic affinities (v. 03_2020_137)

<p><strong>Background</strong></p> <p>Every time we find an interesting gene in an organism of interest, the first question is often &ldquo;how widely is this gene distributed in the eukaryotic kingdom?&rdquo;. Naturally, one could use NCBI BLAST search against the non-redundant sequence database provided by GenBank to answer this question. However, it can be cumbersome to parse the results and assign them to taxonomic units. It is also not straightforward to get an overview of which eukaryotic groups are represented in the results. Top BLAST hits can be crowded with sequences from closely-related organisms making it difficult gain an overview of the overall distribution across eukaryotes. To streamline this process, we developed an in-house database of complete eukaryotic proteomes. We tagged each sequence with a eukaryotic group handle (two-character symbol) and combined them into a single data set searchable by standalone BLAST on one&rsquo;s own computer. We named this data set &ldquo;Darwin&rdquo; to reflect the diverse nature of the sequences it contains.&nbsp;</p> <p><strong>Methods</strong></p> <p>We downloaded predicted proteomes in FASTA format from different sources such as GenBank, Joint Genome Institute (Depart of Energy, USA), Broad Institute (Massachusetts Institute of Technology, USA), Phytozome and a number of other specialized websites catering for a specific organism such as the Arabidopsis Information Resource (TAIR), or the Saccharomyces Genome Database (SGD). All the organisms we included in Darwin are listed in Table 1. To reduce redundancy, we took care not to include the same species more than once unless subspecies were known to show wide diversity. Each sequence header was tagged with a eukaryotic group handle composed of two-character symbols (based on Keeling&nbsp;<em>et al</em>., 2005). These handles clearly appear in BLAST output and can be parsed easily. We combined sequences from all proteomes into a single data set and named it &ldquo;Darwin&rdquo;.</p> <p><strong>Results</strong></p> <p>The current version of Darwin (v. 03_2020_137) contains 2,601,132 amino acid sequences from 137 eukaryotes (Table 1, Data file 1). The sizes of the proteomes were diverse, ranging from ~4000 sequences in some alveolates to 60,000-76,000 in plants. Darwin represents most of the supergroups of eukaryotic kingdom described in Keeling&nbsp;<em>et al.,</em>&nbsp;(2005) except those in Rhizaria whose genomes were not available at the time of data set construction. The data set contains larger numbers of proteomes from fungi and plants reflecting areas of interest in our group.&nbsp;</p> <p><strong>Conclusions</strong></p> <p>Darwin is provided as a text fasta file that can be formatted for BLAST searches on standalone computers. The results from the BLAST searches can be parsed to determine how widely a gene of interest is distributed among different eukaryotes. Simple counting of the eukaryotic group handles would also yield an overview of the distribution across taxa. Darwin is also useful for rapidly finding out whether a gene is missing in particular taxa.</p> <p><strong>Reference</strong></p> <p>Keeling PJ, Burger G, Durnford DG, Lang BF, Lee RW, Pearlman RE, Roger AJ, Gray MW (2005) The tree of eukaryotes.&nbsp;<em>Trends Ecol. Evol.</em>&nbsp;<strong>20:</strong>&nbsp;670-676</p>

opencc-by-4.0Mar 2020View details →
zenodo48/100

Comparative proteomics analysis of whole-cell catalyst of K. rhizophila strain SA117 catabolism of SMX

<p><span>Sulfamethoxazole (SMX), an oral sulfonamide antibiotic, presents significant environmental challenges due to its persistence and potential role in promoting antibiotic resistance. The bacterial strain <em><span>Kocuria rhizophila</span></em> SA117, isolated from polluted soils, has demonstrated a remarkable capability to metabolize SMX. Proteomic analysis revealed the presence of various enzymes and metabolic pathways that may contribute to SMX degradation, including those involved in para-aminobenzoate condensation and protocatechuate metabolism. Notably, the genome of SA117 harbors eight monooxygenase genes, including those related to antibiotic biosynthesis and flavin family monooxygenases. Additionally, several cytochrome c-encoding genes, known for their role in respiratory versatility and potential application in bioremediation, were identified. Genes associated with sulfur metabolism, including an iron-sulfur cluster gene cluster (SufB, C, D, R, E) linked to oxidative stress response, were also found. A comparative proteomic study under SMX exposure highlighted significant upregulation of stress-related proteins. These findings underscore the metabolic adaptability of <em><span>Kocuria rhizophila</span></em> SA117 and its potential application in the bioremediation of SMX-contaminated environments.</span></p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Unified Human Gastrointestinal Proteome clustering results by DPCfam

<p>This dataset contains the result of clustering the Unified Human Gastrointestinal Proteome (UHGP) using the DPCfam algorithm.&nbsp;</p> <p>More details on the DPCfam clustering algorithm can be found in the original publication:</p> <p>Russo, Elena Tea, et al. "DPCfam: Unsupervised protein family classification by Density Peak Clustering of large sequence datasets."&nbsp;<em>PLOS Computational Biology</em>&nbsp;18.10 (2022): e1010610. <a href="https://doi.org/10.1371/journal.pcbi.1010610">https://doi.org/10.1371/journal.pcbi.1010610</a></p> <p>All of the putative protein families obtained through DPCfam (including previous results) can be browsed online at our dedicated webserver:&nbsp;<a href="https://dpcfam.areasciencepark.it/uhgp">https://dpcfam.areasciencepark.it/uhgp</a></p> <p>The original protein dataset is version 1.0 of the UHGP-50 dataset, available for download from MGnify&nbsp;at&nbsp;<a href="https://www.ebi.ac.uk/metagenomics/.">https://www.ebi.ac.uk/metagenomics/</a>.</p> <p><strong>FILES DESCRIPTION:</strong></p> <p>Only MCs with seeds with 1) more than 50 elements and 2) average length larger than 50 aminoacids are reported.</p> <p><strong>metaclusters_xml.tar.gz:</strong></p> <ul> <li><strong>dpcfam_uhgp_metaclusters.xml</strong>:&nbsp;Metaclusters' seeds.&nbsp;Metaclusters entries include also some statistical information about each MC (such as size, average length, low complexity fraction, etc.) and Pfam comparison (Dominant Architecture).</li> <li><strong>dpcfam_metaclusters.xsd</strong>: XML schema file for the data.&nbsp;</li> <li><strong>MCxml_to_tables.awk:</strong> Awk script to convert from XML to tabular text files. Use through the parse.sh script.</li> <li><strong>parse.sh</strong>: XML parser.&nbsp;</li> <li><strong>README.md</strong></li> </ul> <p><strong>uhgp_xml.tar.gz:&nbsp;</strong></p> <ul> <li><strong>uhgp_seed_match.xml</strong>: XML file containing all of UHGP-50 proteins and its corresponding sequences, annotated with Pfam and DPCfam metacluster data.&nbsp; Annotations comprise the membership of a protein as a seed or matches found though the profile-hmms of the DPCfam-UHGP and the DPCfam-Uniref clusterings.&nbsp;</li> <li><strong>uhgp_matches.xsd</strong>: XML schema file for the data.&nbsp;</li> <li><strong>xml_to_list.awk:</strong> Awk script to convert from XML to tabular text files. Use through the parse.sh script.</li> <li><strong>xml_to_list_mcfiles.awk:</strong> Awk script to convert from XML to tabular text files (including individual files for metaclusters' seeds). Use through the parse.sh script.</li> <li><strong>parse.sh</strong>: XML parser.&nbsp;</li> <li><strong>README.md</strong></li> </ul> <p><strong>Metacluster Files:</strong></p> <ul> <li><strong>seeds.zip: </strong>Metaclusters' seed sequences. A fasta file for each metacluster before filtering.</li> <li><strong>filtered_seeds.zip:&nbsp;</strong>Metaclusters'&nbsp;seed sequences after clustering at 60 percent identity.&nbsp;</li> <li><strong>metaclusters_hmms.tar.gz:&nbsp;</strong>Metaclusters' profile-hmms.&nbsp;A&nbsp;".hmm" file for each metacluser.&nbsp;</li> <li><strong>metaclusters_msas.tar.gz:&nbsp;</strong>Metaclusters' multiple sequence alignments, in fasta format.&nbsp;</li> </ul> <p><strong>uhgp_protein_mapping.txt:</strong></p> <ul> <li>Contains a mapping between the identifiers of versions 1.0 and 2.0.2 of UHGP. The first column corresponds to the ID in UHGP-50 1.0 (representatives for the clustering at 50% protein identity), the second column to the ID in version 2.0.2 and the third column to the ID of the representative of the protein for clustering at 100% sequence identity, for which the protein sequence can be found in UHGP-100.&nbsp;&nbsp;</li> </ul>

opencc-by-4.0Nov 2022View details →
zenodo48/100

Mass spectrometry raw data for "Proteomics reveals substantial differences between in vitro matured abattoir-derived and in vivo matured oocytes in cattle"

<p><em><span>In vitro</span></em><span> production (IVP) of bovine embryos still has its limitations such as low blastocyst rate and lower embryo quality, resulting in lower pregnancy rates following the transfer of IVP embryos compared to <em>in vivo</em> produced embryos. </span><span>Given these differences in developmental competence, RNA sequencing and microarray technology have been applied to describe the differences in transcriptional activity between <em>in vitro</em> and <em>in vivo</em> produced embryos. All but one of these studies solely utilized oocytes obtained from slaughterhouse material for the <em>in vitro</em> production of embryos, thereby introducing the possibility, that differences between IVP and <em>in vivo</em> embryos are in part attributable to differing sources of oocytes. The aim of the present study was therefore to compare the proteome of oocytes retrieved from slaughterhouse material, with and without a period of <em>in vitro</em> maturation and <em>in vivo</em> matured oocytes obtained from donor cattle following superovulation. <span>For each group the protein pattern of four biological replicates containing ten oocytes each were analyzed via SWATH<sup>TM</sup>-MS.</span></span></p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

R script and data files for Oakley et al (2017) Journal of Proteome Research. DOI: 10.1021/acs.jproteome.6b00797

<p>This R script and data&nbsp;replicates the analysis&nbsp;of Oakley et&nbsp;al&nbsp;(2017) Thermal shock induces host proteostasis disruption and endoplasmic reticulum stress in the model symbiotic Cnidarian <em>Aiptasia</em>. <em>Journal of Proteome Research</em>. 16:2121-2134. DOI: 10.1021/acs.jproteome.6b00797.&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo48/100

Microbial Proteomes with/without experimental optimal growth temperature

<p>This repository contains proteomes of microorganisms&nbsp;with/without experimentally determined&nbsp;optimal growth temperature (OGT), used in&nbsp;the paper &#39;<strong>Li G, Rabe KS, Nielsen J &amp; Engqvist MKM (2019) Machine learning applied to predicting microorganism growth temperatures and enzyme catalytic optima. </strong><em>ACS Synth. Biol.</em><strong> 8: 1411&ndash;1420</strong>&#39;. There are two .tar.gz files:</p> <p>(1) classified.tar.gz. It contains 5761 proteomes with experimental OGT. The name format of each proteome is &#39;{ogt}_{organism_name}_{organism domain}.fasta&#39;. For example, &#39;36_escherichia_coli_bacteria.fasta&#39; for <em>Escherichia coli.</em></p> <p>(2)&nbsp;not_classified.tar.gz. It contains 1803 proteomes without experimental OGT. The name format is similar as in&nbsp;classified.tar.gz. The only different is to use&nbsp;&nbsp;&#39;tt&#39; to represent the unknown OGT value. For example, &#39;tt_candidatus_azobacteroides_bacteria.fasta&#39;.</p> <p>All proteomes are in fasta format.&nbsp;</p> <p>If you used the dataset, please kindly cite the paper mentioned above.</p>

opengpl-2.0Jul 2019View details →
zenodo48/100

Ergolide Mediates Anti-Cancer Effects on Metastatic Uveal Melanoma Cells and Modulates their Cellular and Extracellular Vesicle proteomes

<p>Underlying dataset and extended dataset&nbsp;of the results described in the article &quot;Ergolide Mediates Anti-Cancer Effects on Metastatic Uveal Melanoma Cells and Modulates their Cellular and Extracellular Vesicle proteomes&quot;.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

MATEdb2, a Collection of High-Quality Metazoan Proteomes across the Animal Tree of Life to Speed Up Phylogenomic Studies

<p>Recent advances in high-throughput sequencing have exponentially increased the number of genomic data available for animals (Metazoa) in the last decades, with high-quality chromosome-level genomes being published almost daily. Nevertheless, generating a new genome is not an easy task due to the high cost of genome sequencing, the high complexity of assembly, and the lack of standardized protocols for genome annotation. The lack of consensus in the annotation and publication of genome files hinders research by making researchers lose time in reformatting the files for their purposes but can also reduce the quality of the genetic repertoire for an evolutionary study. Thus, the use of transcriptomes obtained using the same pipeline as a proxy for the genetic content of species remains a valuable resource that is easier to obtain, cheaper, and more comparable than genomes. In a previous study, we presented the Metazoan Assemblies from Transcriptomic Ensembles database (MATEdb), a repository of high-quality transcriptomic and genomic data for the two most diverse animal phyla, Arthropoda and Mollusca. Here, we present the newest version of MATEdb (MATEdb2) that overcomes some of the previous limitations of our database: (i) we include data from all animal phyla where public data are available, and (ii) we provide gene annotations extracted from the original GFF genome files using the same pipeline. In total, we provide proteomes inferred from high-quality transcriptomic or genomic data for almost 1,000 animal species, including the longest isoforms, all isoforms, and functional annotation based on sequence homology and protein language models, as well as the embedding representations of the sequences. We believe this new version of MATEdb will accelerate research on animal phylogenomics while saving thousands of hours of computational work in a plea for open, greener, and collaborative science.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Predicted proteome of Paratrimastix pyriformis

<p>The upload contains a predicted proteome from a genomic assembly of flagellate <em>Paratrimastix pyriformis</em> (Metamonada, Excavata). The publication describing the genomic study in detail is in progress. The final&nbsp;<em>P.&nbsp;pyriformis</em> genome was assembled into 650 scaffolds spanning 56,722,987 bp, with an N50&nbsp;=&nbsp;268,802 bp and a GC content of 60.92%. Manual and automatic gene prediction resulted in 13,532&nbsp;predicted protein-coding genes, which are the subject of this upload. Below we briefly describe, how the data were generated.&nbsp;&nbsp;</p> <p><strong>DNA isolation: </strong>Monoeukaryotic, xenic culture of <em>P. pyriformis</em> (strain RCP-MX, ATCC 50935) was maintained in the Sonneborn&#39;s Paramecium medium ATCC 802 at room temperature. The DNA was isolated from 15 litres of culture using two different kits. The gDNA samples for PacBio, Illumina HiSeq, and Illumina MiSeq sequencing were each isolated using the Qiagen DNeasy Blood &amp; Tissue Kit (Qiagen). The isolated gDNA was further ethanol-precipitated to increase the concentration and remove any contaminants. For nanopore sequencing, the DNA was isolated using Qiagen MagAttract HMW DNA Kit (Qiagen) according to the manufacturer&rsquo;s protocol.</p> <p><strong>Sequencing:</strong>&nbsp;We used three platforms to generate the sequence data - PacBio (RSII sequencer), Illumina (HiSeq and MiSeq) and Oxford Nanopore (two flow cells, MinION Mk1B).&nbsp;</p> <p><strong>Assembling:</strong> Sequencing quality was assessed with FastQC (Andrew 2010).&nbsp; For the Illumina data, adapter and quality trimming was performed using Trimmomatic 0.36 (Bolger et al. 2014), with a quality threshold of 15. For the nanopore data, trimming and removal of chimeric reads was performed using Porechop v0.2.3 (https://github.com/rrwick/Porechop).&nbsp;The initial assembly of the genomes was made only with the Nanopore and PacBio generated reads using Canu v1.7.1 assembler (Koren et al. 2017), with the corMinCoverage and corOutCoverage set to 0 and 100000 respectively. After assembly, the data were binned using tetraESOM (Haddad et al. 2009). The resulting eukaryotic bins were also checked using a combination of BLASTn and BLASTp and a scoring strategy based on the identity and coverage of the scaffold as described in (Treitli et al. 2019). After binning, the resulted genomic bins were polished in two phases. In the first phase, the scaffolds were polished using the raw reads generated by nanopore with Nanopolish (Loman et al. 2015). In the second phase, the resulting scaffolds generated by Nanopolish were further corrected using Illumina short reads with Pilon v1.21 (Walker et al. 2014). Finally, the genome assembly of <em>P. pyriformis</em> was further scaffolded with raw RNA-seq reads using Rascaf (Song et al. 2016).&nbsp;</p> <p><strong>Gene prediction:</strong>&nbsp;For <em>de novo</em> prediction of genes, first, we manually re-trained Augustus using a manually curated set of gene models. After re-training of Augustus, intron hints were generated from the RNAseq data and gene prediction was performed on repeat masked genomes using Augustus 3.2.3 (Stanke and Waack 2003). For polishing of the predicted genes, we mapped the transcriptome assemblies to the genome using PASA (Haas et al. 2003) and used the assembled transcripts by PASA as evidence for gene model polishing with EVM (Haas et al. 2008).&nbsp;</p> <p><strong>Protein annotation: </strong>Automatic annotation of the proteins was performed using KEGG Automatic Annotation Server (Moriya et al. 2007), as well as similarity searches using BLAST against NCBI nr protein database. Manual search and annotation were performed by searching the predicted proteome using BLAST and HMMER (Finn et al. 2011). Proteins of interest were manually investigated and if possible, the gene models were manually corrected.</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Input features of E. coli proteome for predicting and modeling protein-protein interactions with AF2Complex

<p>Input features to be used with AF2Complex for predicting protein-protein interactions among ~4400 E. coli proteins. A pickled feature file was generated by the feature data pipeline of AF2Complex for each E. coli protein. To reduce storage size, we limited up to 10,000 MSA sequences and up to 10 structural templates from the Protein Data Bank. The cutoff date for sequence libraries and the Protein Data Bank releases used for feature generation is no later than 11-30-2021.</p> <ul> <li>ecoli_af2c_fea.txt -- A list of all E coli protein with pre-generated input features</li> <li>af2c_fea_ecoli_220331_msa10ktem10.tar&nbsp;-- Input features named after the UniProt ID of each proteins. Note that after untar the tarball, you may use the gzipped feature pickle files directly with AF2Complex w/o gunzip.</li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Metabolome and proteome dataset from yeast kinase knock-outs

<p>The dataset comprised of processed data produced in Zelezniak at al, Cell Systems 2018 study, please see README.txt for the detailed description of files.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2018View details →
zenodo44/100

A proteome-wide quantitative platform for nanoscale spatially resolved extraction of membrane proteins into native nanodiscs

<p><strong>EM Quantitation:</strong></p> <p>Raw data gathered from EM images taken to determine nanodisc population size distribution.</p> <p>&nbsp;</p> <p><strong>NNB TGN46 analysis:</strong></p> <p>Data analysis of the Native Nanobleach experiments of TGN46 in native nanodiscs to determine population distribution of oligomeric organizations.</p> <p>&nbsp;</p> <p><strong>Polymer conditions:</strong></p> <p>Physiochemical characteristic and extraction conditions for all polymers in the library both commercially available and in-house.</p> <p>&nbsp;</p> <p><strong>Protein groups polymer screen original file:</strong></p> <p>Original output of MaxQuant data processing of polymer screen data.</p> <p>&nbsp;</p> <p><strong>Organelle matching:</strong></p> <p>Code used for mathcing proteins identified in the proteomics output to organelle or residence for all organellar annotations.</p> <p>&nbsp;</p> <p><strong>Polymer code:</strong></p> <p>Code used to process and normalize the MaxQuant output and calulate extraction efficiency across all detected proteins.</p> <p>&nbsp;</p> <p><strong>MAP Library Details:</strong></p> <p>Graphic and table explaining chemical details of all polymer used in the screen, both commerically available and in-house synthesized.</p> <p>&nbsp;</p> <p><strong>NNB TGN46:</strong></p> <p>Raw scope files for the TIRF microscopy single molecule step photobleaching experiment with TGN46.</p> <p>&nbsp;</p> <p><strong>Organellar Breakdown Database:</strong></p> <p>Proteins detected in the polymer screen through proteomics experiments stratified into organelle of residence.</p> <p>&nbsp;</p> <p><strong>Human Proteome FASTA:</strong></p> <p>The FASTA file used for proteome searching in processing the proteomics data to build the screening database.</p> <p>&nbsp;</p> <p><strong>Hand Curated Organellar Proteomes:</strong></p> <p>Organellar proteomes used for organellar sorting and identification of proteins detected in the screen.</p> <p>&nbsp;</p> <p><strong>Polymer SEC Superdex75:</strong></p> <p>Size exculsion chromatography traces for chloroSMA series of polymers. Was used to characterize length and population polydispersity.</p> <p>&nbsp;</p> <p><strong>Negative Stain Raw:</strong></p> <p>RAW TEM scope images of purified synaptophysin-vamp2 containing nanodiscs. Populatoin size distribution was determined.</p> <p>&nbsp;</p> <p><strong>FSEC Polymer CS80:</strong></p> <p>Fluoresence size exclusion chromatogram for purified synaptophysin-vamp2 containing nanodiscs to ensure population homogeneity and purity.</p> <p><strong>NMR Raw data:</strong></p> <p>NMR raw files for characterizing the in-house synthesized Chloro-SMA series and AASTY series.</p> <p>&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

A Bioconductor workflow for processing, evaluating and interpreting expression proteomics data

<p>Files for users of the workflow "A Bioconductor workflow for processing, evaluating and interpreting expression proteomics data". Files include Proteome Discoverer (v2.5) processing and consensus workflows for both TMT and LFQ expression proteomics data. Also provided are the output .txt files of a corresponding Proteome Discoverer identification search, as required for users to follow the workflow themselves. For raw data please refer to PRIDE. Appendix is provided as a PDF.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

vPro-MS peptide spectral library for the identification of human-pathogenic viruses by untargeted proteomics

<p>The viral proteomics workflow (vPro-MS) enables identification of human-pathogenic viruses from patient samples by untargeted proteomics. vPro-MS is based on an in-silico derived peptide library covering the human virome in <a href="https://www.uniprot.org/" rel="nofollow">UniProtKB</a> (331 viruses, 20,386 genomes, 121,977 peptides). vPro-MS is intended to identify human-pathogenic viruses from DiaNN (<a href="https://github.com/vdemichev/DiaNN">https://github.com/vdemichev/DiaNN</a>) outputs of either DIA or diaPASEF data. A scoring algorithm (vProID) assesses the confidence of virus identification and the results are finally summarized in a report table.&nbsp;</p> <p>The vPro Peptide Library folder contains 3 peptide FASTA files (Contaminants.fasta, Human.fasta, vPro.Virus.fasta), which were used to predict the spectral library (vPro-lib.predicted.speclib). Please note, that the additional commands &ldquo;--cut&rdquo; and &ldquo;--duplicate-proteins&rdquo; are needed to reprocess the prediction in DiaNN. This spectral library should be used to identify peptide sequences from samples of human origin using DiaNN. Furthermore, the folder contains the metadata file of the viral peptide sequences (vPro.Peptide.Library.txt) and a summary file of the virus taxonomy covered by the library (Taxonomy.Summary.txt). The metadata file is used by the vPro script to identify viruses from the DiaNN main report.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Supplementary Files for "Proteomic analysis of the sponge Aggregation Factor implicates an ancient toolkit for allorecognition and adhesion in animals"

<p>This repository hosts supplemental files for the Manuscript "Proteomic analysis of the sponge Aggregation Factor implicates an ancient toolkit for allorecognition and adhesion in animals" by Ruperti, et al., 2024.</p> <ul> <li><strong>Suppl_File_wreath_domain_model.pdb</strong>: AlphaFold3 model for the <em>C. prolifera</em> MAFp3 wreath domain (aa 33 - 317)</li> <li><strong>Suppl_File_MAFAP1_Cterm_model.cif</strong>:&nbsp;AlphaFold3 model for the <em>C. prolifera</em> MAFAP1 C-terminal domain, region 1 and 2</li> <li><strong>Suppl_File_AFInteracting_hmm.hmm</strong>: HMM sequence profile of AF-interacting region of C. prolifera proteins</li> <li><strong>XXX_Foldseek.zip</strong>: Foldseek raw search results, separated by target databases (Swissprot, AFDB, CATH50)</li> </ul>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Processed proteomic and phosphoproteomic timeseries from Ostreococcus tauri, with Gene Ontology enrichment, from "A phospho-dawn of protein modification anticipates light onset in the picoeukaryote O. tauri"

<p>Diel regulation of protein levels and protein modification had been less studied than transcript rhythms. These data tables in .XLSX format report partial proteome (Table_S1)&nbsp;and phosphoproteome data (Table_S2), assayed using shotgun mass-spectrometry, from cultures of the alga <em>Ostreococcus tauri&nbsp;</em>under light-dark cycles, sampled at Zeitgeber times (ZT, hours) 0, 4, 8, 12, 16 and 20.&nbsp;10% of quantified proteins but two-thirds of phosphoproteins were rhythmic. Gene Ontology enrichment analysis was applied to infer the functional enrichment of the proteins or phosphoproteins, grouped by their loadings in PCA analysis (Table_S3), by hierarchical clustering (Table_S4) or&nbsp;by the peak time of their rhythmic profile (Table_S5).Prompted by night-peaking and apparently dark-stable proteins, we also tested the proteome of cultures transferred to prolonged darkness for 24, 48, 72 or 96h (Table_S6), where the proteome changed less than under the diel cycle. The raw data are available from ProteomeXchange, with identifiers PXD001734, PXD001735 and PXD002909.</p>

opencc-by-4.0Mar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record