Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,448

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,448 results for “proteomic”

Learn how ShareScore rates datasets ↗
zenodo40/100

Proteomic data (LC-MS/MS) of human plasma samples

<p>LC-MS/MS analysis of 8 different samples of plasma: 4 samples correspond to the activated platelet-rich plasma (PRP) fractions from 4 different patients with infertility due to Asherman&#39;s syndrome and/or endometrial atrophy; 2 samples correspond to the activated and not-activated, respectively, PRP fractions from a control fertile patient; 2 samples&nbsp;correspond to the activated and not-activated, respectively,&nbsp;fractions from a commercial umbilical cord plasma.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Early 3D Evolution of the SARS-CoV-2 proteome -- Supplementary Tables and Models

<p><strong>Evolution of the SARS-CoV-2 proteome in three dimensions (3D) during the first six months of the COVID-19 pandemic</strong></p> <p><a href="https://iqb.rutgers.edu/covid-19_proteome_evolution">https://iqb.rutgers.edu/covid-19_proteome_evolution</a></p> <p>&nbsp;</p> <p><strong>Legends for Supplementary Figures for 29 </strong><strong>SARS-CoV-2 Study Proteins</strong></p> <p><strong>Separate analysis of protein changes was performed for each study protein and complex. Description below applies to all figures.</strong></p> <p><strong>A</strong>: Observed frequencies for all USV substitutions of Native Residue (i.e., amino acid type in the reference protein sequence) changing to Substituted Residue for a given protein/complex. Red boxes enclose conservative substitutions for hydrophobic, uncharged polar, positively charged, and negatively charged amino acids, respectively in order from upper left to lower right. Cysteine, Glycine and Proline are excluded from these groupings.</p> <p><strong>B-D</strong>: Normalized Frequency histograms for &Delta;&Delta;G<sup>App</sup> calculated for all USVs for a given protein/complex. These were calculated using three methods, which we refer to as hard-hard (B), soft-hard (C), and soft-soft (D), based on the scoring functions used for sidechain rotamer optimization and gradient-based energy minimization respectively (see methods). All energy values described in the text were obtained using the soft-hard method. Overlay of energy histogram with fitted bi-Gaussian curve (solid red line) and fitted single Gaussian curves for subsets of USVs with surface (green), boundary layer (yellow), or core (blue) substitutions. USVs with multiple substitutions were included in single Gaussian fitting when all substitutions mapped to the same region of the study protein. The data used for fitting includes the energies of all unique protein models produced by a given method, excluding extreme outliers with energy values greater than 3 standard deviations away from the central mean.</p> <p><strong>E-G</strong>: USV Count histograms indicate the number of USVs among the full set for a given protein in which each site included a substitution. Sites are separated by burial layer. Substitutions at sites that are absent from the available crystal structures are excluded from the histograms. In most cases, only a single protein is analyzed, and only panel E is included. In the case of complexes, a separate histogram is provided for each protein in the complex: for methyltransferase nsp10-nsp16, E is nsp10 and F is nsp16; for RDRP nsp12-nsp7-nsp8, E is nsp7, F is nsp8, and G is nsp12.</p> <p>&nbsp;</p> <p><strong>Legends for Supplementary Tables for 29 </strong><strong>SARS-CoV-2 Study Proteins</strong></p> <p><strong>Table: USVs</strong>: All identified USVs for a protein/complex. Columns are:</p> <ul> <li>date: Date of first collection of a strain with the USV reported to GISAID</li> <li>gisaid_count: The number of sequences in the GISAID database that include the USV</li> <li>id: The GISAID strain identification for the first collected instance of the USV</li> <li>location: The country in which the first strain including the USV was collected</li> <li>substitutions: All substitutions in the USV, in the form [chain]_[sequence][site][substitution], with multiple substitutions separated by semicolons</li> <li>is_in_PDB: whether a substitution is present in the PDB model used to generate the USV structure, with multiple substitutions separated by semicolons</li> <li>multiple: whether more than one amino acid substitution is present in the USV</li> <li>conservative: whether a substitution is conservative, with multiple substitutions separated by semicolons</li> <li>layer: Identification of the burial layer (surface, boundary, or core) of a substitution in the reference structure, with multiple substitutions separated by semicolons and substitutions absent from the PDB excluded</li> <li>sh_rmsd: The RMSD of the USV to the reference structure when modeled using the soft-hard method</li> <li>sh_ddG: The &Delta;&Delta;G<sup>App</sup> of the USV when modeled using the soft-hard method</li> <li>hh_rmsd: The RMSD of the USV to the reference structure when modeled using the hard-hard method</li> <li>hh_ddG: The &Delta;&Delta;G<sup>App</sup> of the USV when modeled using the hard-hard method</li> <li>ss_rmsd: The RMSD of the USV to the reference structure when modeled using the soft-soft method</li> <li>ss_ddG: The &Delta;&Delta;G<sup>App</sup> of the USV when modeled using the soft-soft method</li> </ul> <p>&nbsp;</p> <p><strong>Table: Substitutions</strong>: All substitutions identified for a protein/complex</p> <ul> <li>chain: The chain identifier of the protein in the PDB file in which the substitution is present</li> <li>site: The residue number at which the substitution is present</li> <li>reference: The one-letter amino acid name of the residue in the reference sequence</li> <li>mutant: The one-letter amino acid name of the residue in a USV</li> <li>conservative: Indication of whether a substitution is conservative</li> <li>in_pdb: whether the substitution site is present in the PDB model used to generate the USV structure</li> <li>layer: Identification of the burial layer (surface, boundary, or core) of a substitution in the reference structure</li> <li>date: date: Date of first collection of a strain with the substitution reported to GISAID</li> <li>location: The country in which the first strain including the substitution was collected</li> <li>gisaid_count: The number of sequences in the GISAID database including the substitution</li> <li>usv_count: The number of identified USVs including the substitution</li> <li>ddG: The soft-hard &Delta;&Delta;G<sup>App</sup> of the USV that includes only the substitution, left empty if no single-substitution USV was identified with the substitution</li> <li>single: Indication of whether the substitution was present in a single-substitution USV</li> <li>multiple: Indication of whether the substitution was present in a USV with multiple substitutions</li> <li>associates: List of all other substitutions that were identified in a USV that included the substitution</li> <li>strains: List of all USV-representative GISAID strains that included the substitution, with the single-substitution USV strain listed first if one was available</li> </ul> <p>&nbsp;</p> <p><strong>Table: Gaussian Fit Statistics</strong>: Fitted models for the energies of all USVs either together (ALL) or by study protein.</p> <ul> <li>fit: The number of Gaussian curves in the fitted energy model&nbsp;</li> <li>protein: The protein/complex name</li> <li>method: The modeling method used to calculate energy values</li> <li>layer: The subset burial layer (surface, boundary, or core) of USVs for which the energy model was fitted, excluding all USVs with substitutions not in that layer</li> <li>&mu;<sub>1</sub>: Mean of the first Gaussian in the fitted model</li> <li>&sigma;<sub>1</sub>: Variance of the first Gaussian in the fitted model</li> <li>wt<sub>1</sub>: Weight of the first Gaussian in the fitted model</li> <li>&mu;<sub>2</sub>: Mean of the second Gaussian in the fitted model</li> <li>&sigma;<sub>2</sub>: Variance of the second Gaussian in the fitted model</li> <li>wt<sub>2</sub>: Weight of the second Gaussian in the fitted model</li> <li>R<sup>2</sup>: R-squared value indicating the goodness of fit</li> </ul> <p>&nbsp;</p> <p><strong>Description of Computed Structural Models </strong><strong>for Unique Sequence Variants for 29 </strong><strong>SARS-CoV-2 Study Proteins.</strong></p> <p><strong>USV Computed Structural Models</strong>. Computed structural models for all amino acid substituted USVs. We are providing the structural models of all study proteins modeled using the soft-hard modeling method (see Methods). Structural models are named according to the GISAID strain identification of the first strain in which the USV was identified, followed by an underscore-separated list of substitutions in the form [chain]_[sequence][site][substitution]. Atomic coordinates for each computed structural model are provided in the legacy Protein Data Bank format used by most molecular graphics software tools (see <a href="https://www.wwpdb.org/documentation/file-format-content/format33/v3.3.html">https://www.wwpdb.org/documentation/file-format-content/format33/v3.3.html</a> for detailed description).</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Proteomics and metabolomics data associated with the end-of-life phenotype Smurf

<p>Data obtained from whole bodies of mated females of genotype Drs-GFP at 20 and 40 days for proteomics and 30 days for metabolomics.</p> <p>The data are cited in the following preprint :</p> <p>Smurfness-based two-phase model of ageing helps deconvolve the ageing transcriptional signature</p> <p>Flaminia&nbsp;Zane,&nbsp;Hayet&nbsp;Bouzid,&nbsp;Sofia Sosa&nbsp;Marmol,&nbsp;<a href="http://orcid.org/0000-0003-2448-4022">&nbsp;View ORCID Profile</a>Savandara&nbsp;Besse,&nbsp;Julia Lisa&nbsp;Molina,&nbsp;<a href="http://orcid.org/0000-0002-9579-5250">&nbsp;View ORCID Profile</a>C&eacute;line&nbsp;Cansell,&nbsp;Fanny&nbsp;Aprahamian,&nbsp;<a href="http://orcid.org/0000-0001-6356-1006">&nbsp;View ORCID Profile</a>Sylv&egrave;re&nbsp;Durand,&nbsp;Jessica&nbsp;Ayache,&nbsp;<a href="http://orcid.org/0000-0001-7709-2116">&nbsp;View ORCID Profile</a>Christophe&nbsp;Antoniewski,&nbsp;<a href="http://orcid.org/0000-0002-6574-6511">&nbsp;View ORCID Profile</a>Michael&nbsp;Rera</p> <p>doi:&nbsp;https://doi.org/10.1101/2022.11.22.517330</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

pSCoPE: Prioritized Single-Cell Proteomics (data for generating publication figures)

<p>Major aims of single-cell proteomics include increasing the consistency, sensitivity, and depth of protein quantification, especially for proteins and modifications of biological interest. To simultaneously advance all these aims, we developed prioritized Single Cell ProtEomics (pSCoPE). pSCoPE consistently analyzes thousands of prioritized peptides across all single cells (thus increasing data completeness) while analyzing identifiable peptides at full duty-cycle, thus increasing proteome depth. These strategies increased the sensitivity, data completeness, and proteome coverage over 2-fold. The gains enabled quantifying protein variation in untreated and lipopolysaccharide-treated primary macrophages. Within each condition, proteins covaried within functional sets, including phagosome maturation and proton transport. This protein covariation within a treatment condition was similar across the treatment conditions and coupled to phenotypic variability in endocytic activity. pSCoPE also enabled quantifying proteolytic products, suggesting a gradient of cathepsin activities within a treatment condition. pSCoPE is freely available and widely applicable, especially for analyzing proteins of interest without sacrificing proteome coverage. Support for pSCoPE is available at: <a href="http://scp.slavovlab.net/pSCoPE">scp.slavovlab.net/pSCoPE</a></p> <p>&nbsp;</p> <p>The files contained in this .zip directory are necessary for replicating the analysis and figures associated with the pSCoPE manuscript.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
dryad40/100

Proteomic and metabolomic analysis of COVID-19 nasal swabs

<p>The epithelial barrier's primary role is to protect against entry of foreign and pathogenic elements. Global and targeted approaches were applied to nasal swabs from healthy and COVID-19-confirmed cases within 24 hours post-positive-confirmation and at 3 weeks post-infection to observe changes in proteome and metabolome.</p> <p>We found that the tryptophan/kynurenine metabolism pathway is a pinch-point regulator of canonical and non-canonical transcription activation, macrophage release of cytokines and significant changes in the immune and metabolic status with increasing severity and disease course.</p>

opencc-zeroFeb 2023View details →
zenodo40/100

DEcancer: Proteomics Datasets of Liquid Biopsy Samples

<p>Two datasets used to build the DEcancer pipelines from open access data. The tables in the cancerseek folder are uploaded for convenience and can also be found from the supplementary data of&nbsp;<a href="https://doi.org/10.1016/j.isci.2019.04.035">Early Cancer Detection from Multianalyte Blood Test Results</a>&nbsp;paper. The nanoparticles dataset is processed with the Perseus software on the&nbsp;<a href="http://proteomecentral.proteomexchange.org/cgi/GetDataset?ID=PXD017052">original data</a>&nbsp;from Blume&nbsp;<em>et al&#39;s</em>&nbsp;Rapid, deep and precise profiling of the plasma proteome with multi-nanoparticle protein corona. These are distinct datasets in that the nanoparticles is a cohort of patients from a high dimensional, low sample size proteomics dataset.</p>

opencc-by-4.0Feb 2023View details →
dryad40/100

RNA-Seq and proteomics of Crohn's disease: Terminal ileum of inflamed and non inflamed paired tissue biopsy

<p>The cause of chronic inflammation manifestations such as inflammatory bowel disease (IBD) is still not yet fully understood. Evidence however points to the convergence of environmental factors, genetic factors, intestinal microbiota and the immune response at the intestinal epithelial surface that consequently leads to a breakdown of barrier function in IBD. An increasing number of theories implicate a dysfunctional intestinal epithelial surface in the pathogenesis of IBD and progression to complicated disease, which contributes significantly to the IBD burden. As yet, the significance of the intestinal epithelial surface in IBD manifestation and progression has not been explored using a systems biology approach, that compares the mucosa of healthy patients to those with IBD. In the context of dyregulation of the epithelial barrier, there is a scarcity of publications that explore the contribution of proteins- the effectors of the cells and RNA message in the pathways affected by CD and in particular the localization of inflammatory response. We aim to evaluate the functional pathways of chronic inflammation at the intestinal epithelial surface using a combination of transcriptomic and proteomic analyses on intestinal epithelial tissue samples from paired inflamed and non-inflammed Crohn's disease (CD) and healthy control subjects. Concordance of several biological pathways from both data sets was found to be altered in CD patients' epithelia when compared to healthy controls. This information could be helpful in identifying novel therapeutic targets that aim to restore barrier function at the intestinal epithelial surface and to guide therapy.</p>

opencc-zeroFeb 2023View details →
zenodo40/100

Data for manuscript, "An optimized workflow for MS-based quantitative proteomics of challenging clinical bronchoalveolar lavage fluid (BALF) samples"

<p>Clinical BALF samples are rich in biomolecules, including proteins, and useful for molecular studies of lung health and disease.&nbsp; However, MS based proteomic analysis of BALF is impeded by the dynamic range of protein abundance, and potential for interfering contaminants.&nbsp; We have developed a workflow that eliminates these challenges.&nbsp; By combining high abundance protein depletion, protein trapping, clean-up, and in-situ tryptic digestion, our workflow is compatible with both qualitative and quantitative MS-based proteomic analysis.&nbsp; The workflow includes collection of endogenous peptides for peptidomic analysis of BALF, if desired, as well as amenability to offline semi-preparative or microscale fractionation of peptide mixtures prior to LC-MS/MS analysis, for increased depth of analysis.&nbsp; We show the effectiveness of this workflow on BALF samples from COPD patients.&nbsp; Overall, our workflow should allow MS-based proteomics to be applied to a wide variety of studies focused on BALF clinical samples.&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;</p> <p>Note:&nbsp; Due to the nature of some of the files, file&nbsp;<em>wendt005_ostr0103_18260_20210831_BALF_FAIMS_MS2_TMT16.msf, wendt005_ostr0103_18976_20230202_quantReport.msf, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_1R.raw,&nbsp;cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_2R.raw, cmsptc_higgi022_18988_20230203_18976DW_EnF_hcdlT_3R.raw and cmsptc_higgi022_18988_20230203_18976DW_Eclipse_noFAIMS_quantReport.msf</em>&nbsp;were&nbsp;zipped into&nbsp;compressed folders before uploading.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Supplementary files to An integrated RNA-proteomic landscape of drug induced senescence in a cancer cell line.

<p>Senescent cells are characterized by an arrest in proliferation. In addition to replicative senescence resulting from telomere exhaustion, sub-lethal genotoxic stress resulting from DNA damage, oncogene activation, mitochondrial dysfunction or reactive metabolites also elicits a senescence phenotype. Senescence is a controlled programme affecting a wide variety of biological processes with some core hallmarks of senescence as well as tissue specific changes. This study presents an integrative multi-omic analysis of proteomic and RNA-seq from proliferating and senescent osteosarcoma cells. This study demonstrates senescence induction in a widely used cell line which can be used as a model system for characterising cancer cell responses to sub-lethal doses of chemotherapeutic agents, and makes available both RNA-seq and proteomic data from proliferating and senescent cells in open access repositories to aid reuse by the community.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Phylogenetic profile of 100 annotated low complexity proteins against the Uniprot Reference Proteome dataset

<p>Phylogenetic profile of&nbsp;100 human proteins with characteristic compositional bias,&nbsp;previously recorded by&nbsp;Mier et al (2020)&nbsp;against the Uniprot&nbsp;Reference Proteome,&nbsp;containing a total of 11297 proteomes, excluding viruses. The counts for each protein correspond to homologs found in each proteome.&nbsp;</p> <p>Detailed description of included columns:</p> <p><strong>ref_proteome_identifier</strong>: The<strong>&nbsp;</strong>Uniprot&nbsp;Reference Proteome&nbsp;identifier</p> <p><strong>ncbi_taxid</strong>: The NCBI taxonomy ID</p> <p><strong>species_name</strong>: NCBI common name corresponding to taxonomy ID</p> <p><strong>species_code</strong>: internal species code composed of 9 characters</p> <p><strong>taxonomic domain</strong>: E/B/A for Eukaryota/Bacteria/Archaea classification of proteome</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

UniProt Human Proteome Benchmarking data for "aaHash: recursive amino acid hashing"

<p>aaHash is a rolling hash algorithm tailed for amino acids.&nbsp;Here, we provide the human proteome benchmarking data used in the aaHash&nbsp;paper &quot;aaHash: recursive amino acid sequence hashing&quot;.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Data for 'Deriving spatial features from in situ proteomics imaging to enhance cancer survival analysis'

<p>Additional data for &#39;Deriving spatial features from in situ proteomics imaging to enhance cancer survival analysis&#39;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Revealing the lipidome and proteome of Arabidopsis thaliana plasma membrane

<p>This table contains peaks aera values from GC-MS, TLC-GC-MS and LC-MS for characterization of Arabidopsis thaliana plasma membrane. These data were used for Fig. 6, 7, 8, 9 and S1, S2, S3 and S4 of Bahammou et al. 2023: Revealing the lipidome and proteome of Arabidopsis thaliana plasma membrane</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Pan-cancer Proteomics Analysis to Identify Tumor-Enriched and Highly Expressed Cell Surface Antigens as Potential Targets for Cancer Therapeutics

<p>CPTAC PAN-cancer Data Repository</p> <p>Welcome to the CPTAC PAN-cancer Data Repository! This repository serves as a data repository for the CPTAC PAN-cancer effort, which focuses on cancer target discovery. It contains various data sets related to protein abundance estimation, derived TMT-TPA, iBAQ, iBAQ-derived copy number, and differential protein expression for CPTAC ten indications.</p> <p>## Contents</p> <p>The repository includes the following data:</p> <p>- FragPipe Output: Protein abundance estimation data generated using the FragPipe software.<br> - Derived TMT-TPA: Data derived from Tandem Mass Tag (TMT) based Total Protein Approach (TPA).<br> - iBAQ: Data representing intensity-based absolute quantification (iBAQ) of proteins.<br> - iBAQ-derived Copy Number: Data derived from iBAQ analysis for copy number estimation.<br> - Differential Protein Expression: Data indicating differential expression of proteins between tumor and NAT.</p> <p>## Data Organization</p> <p>The data in this repository is organized in a structured manner to facilitate easy access and navigation. The repository structure is as follows:</p> <p>FragPipe/<br> [fragpipe_data_files]<br> Derived_TMT_TPA/<br> [derived_tmt_tpa_data_files]<br> iBAQ/<br> [ibaq_data_files]<br> iBAQ-derived_copy_number/<br> [ibaq_copy_number_data_files]<br> Differential_protein_expression/<br> [differential_expression_data_files]</p>

opencc-by-4.0May 2023View details →
zenodo40/100

CPT-1 whole-proteome feature matrices (no-EVE set)

<p><strong>Cross-protein transfer learning for variant effect prediction</strong></p> <p>This repository contains the feature matrices for&nbsp;CPT-1 to make variant effect prediction on&nbsp;15,557 human proteins NOT&nbsp;in&nbsp;the EVE set (<a href="https://www.nature.com/articles/s41586-021-04043-8">Frazer et al., 2021</a>), initially released with the manuscript &quot;Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects&quot;.</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>Jagota, M.*, Ye, C.*, Albors, C., Rastogi, R., Koehl, A., Ioannidis, N., and Song, Y.S.&dagger;<br> &quot;Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects&quot;, bioRxiv (2022)</p> <p>*These authors contributed equally to this work.<br> &dagger;To whom correspondence should be addressed:&nbsp;<a href="mailto:yss@berkeley.edu">yss@berkeley.edu</a></p> <p>DOI:&nbsp;<a href="https://doi.org/10.1101/2022.11.15.516532">https://doi.org/10.1101/2022.11.15.516532</a></p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

CPT-1 pre-computed whole-proteome variant effect predictions and model source code

<p><strong>Cross-protein transfer learning for variant effect prediction</strong></p><p>This repository contains the variant effect predictions of CPT-1 for 18,602 human proteins, initially released with the manuscript "Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects". The proteins are split into three files.</p><p><i>CPT1_score_EVE_set.zip</i>: Proteins in the EVE set (<a href="https://www.nature.com/articles/s41586-021-04043-8">Frazer et al., 2021</a>)</p><p><i>CPT1_score_no_EVE_set_1.zip</i> &amp; <i>CPT1_score_no_EVE_set_2.zip</i>: Proteins not in the EVE set. Predictions for these proteins use imputed values for features depending on the EVE MSA.</p><p>The protein names are UniProt gene names.</p><p>We also provide source code to train CPT-1 model and reproduce results in the manuscript :</p><p><i>source_code.zip </i>(corresponds to GitHub repository&nbsp;songlab-cal/CPT version as of Jul 12, 2023)</p><p>&nbsp;</p><p><strong>Citation</strong></p><p>Jagota, M.*, Ye, C.*, Albors, C., Rastogi, R., Koehl, A., Ioannidis, N., and Song, Y.S.†<br>"Cross-protein transfer learning substantially improves zero-shot prediction of disease variant effects", bioRxiv (2022)</p><p>*These authors contributed equally to this work.<br>†To whom correspondence should be addressed:&nbsp;<a href="mailto:yss@berkeley.edu">yss@berkeley.edu</a></p><p>DOI:&nbsp;<a href="https://doi.org/10.1101/2022.11.15.516532">https://doi.org/10.1101/2022.11.15.516532</a></p><p>&nbsp;</p>

opencc-by-4.0May 2023View details →
dryad40/100

Data from: Genotype-by-environment interactions influence the composition of the Drosophila seminal proteome

<p>Ejaculate proteins are key mediators of post-mating sexual selection and sexual conflict, as they can influence both male fertilization success and female reproductive physiology. However, the extent and sources of genetic variation and condition dependence of the ejaculate proteome are largely unknown. Such knowledge could reveal the targets and mechanisms of post-mating selection and inform about the relative costs and allocation of different ejaculate components, each with its own potential fitness consequences. Here, we used liquid chromatography coupled with tandem mass spectrometry to characterize the whole-ejaculate protein composition across twelve isogenic lines of Drosophila melanogaster that were reared on a high- or low-quality diet. We discovered new proteins in the transferred ejaculate and inferred their origin in the male reproductive system. We further found that the ejaculate composition was mainly determined by genotype identity and genotype-specific responses to larval diet, with no clear overall diet effect. Nutrient restriction increased proteolytic protein activity and shifted the balance between reproductive function and RNA metabolism. Our results open new avenues for exploring the intricate role of genotypes and their environment in shaping ejaculate composition, or for studying the functional dynamics and evolutionary potential of the ejaculate in its multivariate complexity.</p>

opencc-zeroAug 2023View details →
zenodo40/100

Copy number variations and their effect on the plasma proteome | Associations

<p>Results and supplementary data from a GWAS investigation the relationship of CNVs and blood protein measurements.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

The salivary proteome of the green peach aphid/peach-potato aphid (Myzus persicae) (Sulzer, 1776) (Hemiptera, Aphididae).

<p><strong>*for correspondence:&nbsp;</strong>saskia.hogenhout@jic.ac.uk</p> <p>&nbsp;</p> <p><strong>Introduction</strong><strong>&nbsp;</strong></p> <p>The green peach aphid/peach-potato aphid <em>Myzus persicae</em> colonizes hundreds of plant species, an ability that is in part due to the delivery of saliva proteins &ndash; often referred to as effectors &ndash;&nbsp; into the host plant that suppress plant defence. As a generalist herbivore with a remarkable ability to colonize new host plants (Dedryver et al., 2010), <em>M. persicae</em> represents an outstanding model system for studying the molecular mechanisms underlying plant-insect interactions.</p> <p>Recent advancements in mass spectrometry instrumentation (Yu et al. 2020) and database search software (Frejno et al., 2024), along with a new high-quality reference genome assembly for <em>M. persicae</em> (Mathers et al., 2021) and a simplified method for improved aphid saliva recovery that we describe here, collectively enhance the detection of saliva proteins with unprecedented sensitivity and specificity.</p> <p>Here, we present a high-quality salivary proteome of <em>M. persicae </em>generated using a nanoLC-MS/MS analysis in combination with an updated annotation of the <em>Myzus persicae </em>clone O genome.</p> <p>Saliva from over 10,000 <em>M. persicae</em> aphids was collected and analysed using nanoLC-MS/MS (Figure 1). We identified 1557 peptide sequences mapped to the <em>M. persicae</em> clone O genome v2.1 annotation (Table 1); 210 of those peptides additionally appeared as modified by oxidation (M) and/or carbamidomethylation (C) so that a total of 1767 peptide forms mapping to <em>M. persicae</em> were detected with high confidence (combined from both search engines). Of those peptides, 120 were only identified by the CHIMERYS search engine, and 52 were only identified by the Mascot search engine (all with high confidence). Most peptides were detected in the concentrated sample; of the 1767 peptide forms, only 411 were detected in the unconcentrated sample and only 10 of those were exclusively identified in the unconcentrated sample.</p> <p>Those peptides were assigned to a total of 423 <em>M. persicae</em> proteins with high confidence and at least 1 unique peptide including all proteins with shared peptides. The software generated 169 protein groups each represented by one master protein. The master protein is the largest protein with the most peptide matches in the group. Of the 169 protein groups, 126 groups were identified with at least 2 unique peptides and 43 groups were identified with only 1 unique peptide. Protein groups that included products of a single gene model were identified from the peptide output of Proteome Discoverer. The 169 protein groups included proteins encoded by 219 gene models which are listed in Table 3, of which 155 had peptides that did not match to any other gene model.</p> <p>Of note, 8 &nbsp;Cathepsin B cysteine peptidases are listed (highlighted in Table 3) which have previously been shown to play an important role in colonisation of host plants (Mathers et al. 2017). This is a recently diversified gene family including 28 members in the <em>Myzus persicae </em>genome. The detected proteins belong to several protein groups including highly similar or identical proteins with shared peptides detected, and also include CathB17 (MYZPE13164_O_EIv2.1_0124700) which was excluded from the database because it is identical to CathB18 (MYZPE13164_O_EIv2.1_0124690)</p> <p>Together, these findings suggest that aphid saliva contains enzymes that likely alter plant physiology by interacting with both plant proteins and small molecules. Future mechanistic studies will be able to precisely characterize the role of these salivary proteins to understand the pathways, hormones, and chemical defences that may be suppressed by aphids.</p> <p>&nbsp;</p> <p><strong>Materials and Methods</strong></p> <p><strong><em>Aphid rearing</em></strong></p> <p><em>M. pers</em>icae Clone O colonies were reared on <em>Arabidopsis thaliana</em> Col-0 plants in a growth chamber maintained at 20&deg;C with a 14-hour light/10-hour dark cycle and 75% relative humidity. The <em>A. thaliana</em> plants were grown at short-day conditions (10 hours light/14 hours dark) at 22&deg;C and 70% relative humidity. For aphid rearing, 4-week-old <em>A. thaliana</em> plants were used, and plants were replaced every two weeks to ensure optimal conditions for aphid growth.&nbsp;</p> <p><strong><em>Saliva collection</em></strong></p> <p>Figure 1 illustrates the schematic overview of our sample collection and preparation workflow. <em>M. persicae</em> Clone O aphids were transferred to 4-week-old <em>A. thaliana</em> plants and maintained for 2-3 weeks to enable reproduction. Approximately 100 aphids were then placed into a 50 mm Petri dish, which was sealed with parafilm and had a hole in the base for introducing the aphids. The hole was subsequently covered with mesh to allow ventilation while preventing aphid escape. A 300 &mu;L aliquot of artificial diet (15% w/v sucrose in Milli-Q water, sterilized via 0.22 &mu;m filtration) was added to the inverted lid of the Petri dish. The dish, with the aphid chamber positioned over the lid, was set up so that the parafilm made contact with the artificial diet. This setup was kept under short-day conditions (14 hours light/10 hours dark) at 20&deg;C and 75% relative humidity for 24 hours.</p> <p>After 24 hours, the artificial diet, now containing aphid saliva, was collected and pooled to produce an unconcentrated saliva sample. A portion of this sample was then concentrated using a Vivaspin concentrator with a 3 kDa molecular weight cut-off (MWCO) at 4&deg;C. The concentrated saliva was snap-frozen in liquid nitrogen and stored at -80&deg;C until further analysis.</p> <p>This procedure was repeated until saliva was collected from approximately 10,000 aphids.</p> <p>&nbsp;</p> <p><strong><em>Saliva preparation and nanoLC-MS/MS analysis</em></strong></p> <p>Saliva samples were precipitated by adding 4 volumes of methanol and 1 volume of chloroform, followed by centrifugation at maximum speed for 10 minutes (Wessel and Fl&uuml;gge, 1984). After removing the supernatant, the pellet was washed once with acetone before proceeding to trypsin digestion. The protein pellet was resuspended in 50 &micro;L of 1.5% sodium deoxycholate (SDC; Merck) in 0.2 M EPPS buffer (Merck), pH 8.5, and vortexed under heating. Cysteine residues were reduced with dithiothreitol, alkylated with iodoacetamide, and proteins were digested with trypsin in the SDC buffer following standard protocols.</p> <p>After digestion, SDC was precipitated by adjusting the solution to 0.2% trifluoroacetic acid (TFA). The clear supernatant was then subjected to C18 solid-phase extraction (SPE) using OMIX 10-100 &mu;L C18 pipette tips (Agilent).</p> <p>The samples were analysed using nanoLC-MS/MS on an Orbitrap Eclipse&trade; Tribrid&trade; mass spectrometer, coupled with an UltiMate&reg; 3000 RSLCnano LC system (Thermo Fisher Scientific, Hemel Hempstead, UK). Samples were loaded onto a trap column (nanoEase M/Z Symmetry C18 Trap Column, Waters) with 0.1% TFA at a flow rate of 15 &micro;L/min for 3 minutes. The trap column was then switched in-line with the analytical column (nanoEase M/Z HSS C18 T3, 1.8 &micro;m, 100 &Aring;, 250 mm x 0.75 &micro;m, Waters) for separation at 40&deg;C. The gradient used for separation was as follows: solvent A (water with 0.1% formic acid) and solvent B (80% acetonitrile with 0.1% formic acid) at a flow rate of 0.2 &micro;L/min: 0-3 minutes at 3% B (parallel to trapping); 3-10 minutes with B increasing to 8%; 10-130 minutes with B linearly increasing to 45%; 130-145 minutes with B linearly increasing to 55%; followed by a ramp to 99% B and re-equilibration to 0% B, for a total runtime of 180 minutes.</p> <p>Mass spectrometry data were acquired in positive ion mode with the following settings: Orbitrap resolution at 120K, profile mode, mass range m/z 300-1800, normalized AGC target at 100%, and a maximum injection time of 50 ms. For MS2 analysis in IT Turbo mode, parameters included quadrupole isolation window of 1.2 Da, charge states 2-5, threshold at 1.9e4, HCD CE and CID CE both set to 33 in parallel, AGC target at 1e4, maximum injection time of 35 ms, and dynamic exclusion set to 1 count for 15 seconds with a mass tolerance of &plusmn;7 ppm.</p> <p>&nbsp;</p> <p><em><strong>M. persicae v2.1 annotation</strong></em></p> <p>The chromosome scale genome assembly of <em>M. persicae </em>clone O (Mathers et al. 2021) was re-annotated for accurate gene prediction as follows. Illumina short read RNAseq data from <em>M. persicae</em> used for previous annotation (Mathers et al. 2017; EBI ENA SAMEA4469192) was used in addition to stranded RNAseq reads of <em>M. persicae </em>clone O<em> </em>feeding from 9 different host plant species (Chen et al. 2020 ; NCBI GEO GSE129669); from males, alate asexual females and winged asexual females, and nymphs (Mathers et al. 2019; NCBI SRA PRJNA437622); dissected organs from winged and alate asexual female <em>M. persicae</em> (EBI ENA PRJEB79119<a>)</a><em>, </em>and PacBio Isoseq RNAseq data from asexual female <em>M. persicae</em> (EBI ENA PRJEB79119<a>).</a></p> <p>Candidate transcript sequences were assembled from RNA-seq reads with Scallop (Shao and Kingsford 2017) and StringTie (Pertea et al. 2015) using a genome guided approach. A filtered set of non-redundant transcripts are derived using Mikado (Venturini et al. 2018) for the final transcript set for annotation. Mikado models together with aligned proteins and repeat annotation are provided as hints to Augustus (<a href="http://bioinf.uni-greifswald.de/augustus/">http://bioinf.uni-greifswald.de/augustus/</a>). Multiple Augustus gene builds were created from alternative evidence inputs or weightings. These were supplemented with gene models derived directly from protein alignments and high confidence models from the Mikado transcript selection stage. Metrics were generated to assess how well supported each gene model is by available evidence and an integrated set of models produced by Mikado.&nbsp;</p> <p>Long non-coding (lnc) RNAs were identified from the assembled RNAseq. Transcripts with open reading frames (Transdecoder, <a href="https://transdecoder.github.io">https://transdecoder.github.io</a>) showing similarity to arthropod protein coding genes (BlastP e&lt;1e-5), or with HAMMER hits against the Pfam database were excluded. Remaining transcripts with coding potential&nbsp; &gt;0.5 (CPC2, Kang et al. 2017) or that were shorter than 200bp were also excluded. Transcripts mapping to rRNA, tRNA, miRNA or transposon loci were excluded using Mikado (Venturini et al. 2018).&nbsp;</p> <p>In total we identified 37,720 total genes (with 58,609 total splice variant isoforms), including 22,796 (47,508 total isoforms encoding 39,681 unique proteins) protein coding and 7,990 (11,101 total isoforms) non-coding genes. (Further details of the annotation process and statistics can be found in CloneO_v2.1_annotation_summary_stats.txt and Myzus_persicae_O_annotation_readme.doc).</p> <p>&nbsp;</p> <p><strong><em>Mass spectrometry</em> <em>data processing</em></strong></p> <p>The mass spectrometry raw data were processed and quantified in Proteome Discoverer 3.1 (Thermo), all mentioned tools of the following workflows are nodes of the proprietary Proteome Discoverer (PD) software. The database search was performed using the search engines CHIMERYS (MSAID, Munich, Germany) and Mascot Server 2.8.3 (Matrixscience, London) in parallel on the following databases: MYZPE13164_O_EIv2.1.annotation.gff3.pep.fasta (39,681 entries after removal of duplicate protein sequences) and common contaminants (MaxQuant.org, 20240812, 246 entries).&nbsp; The databases were imported into PD adding a reversed sequence database for decoy searches. The processing workflow started with spectrum recalibration on the <em>Myzus</em> protein database, Minora Feature Detection with min. trace length 7, S/N 2.5, PSM confidence high, and Top N Peak Filter with 20 peaks per 100 Da. For CHIMERYS, the inferys_3.0.0_fragmentation prediction model with FDR targets 0.01 (strict) and 0.05 (relaxed), a fragment tolerance of 0.3 Da, enzyme trypsin with 2 missed cleavages, variable modification oxidation (M), fixed modification carbamidomethyl (C) were used. For Mascot, the same parameters were used including a precursor tolerance of 5 ppm and a fragment tolerance of 0.5 Da; validation was performed using Percolator based on q-values and FDR targets 0.01 (strict) and 0.05 (relaxed).</p> <p>The consensus workflow in the PD software was used to evaluate the peptide identifications and to measure the abundances of the peptides based on the LC-peak intensities. For identification, an FDR of 0.01 was used as strict threshold. Protein abundance was calculated using the Top3 most abundant peptides. The results were exported into Microsoft Excel including data for protein abundances, number of peptides, protein coverage, the search identification score and other important values (Tables 1 and 2). Identification of protein groups with members encoded by a single gene model was performed by first identifying peptides that mapped to a single gene model, then counting the number of peptides that were specific to each gene model.</p> <p>&nbsp;&nbsp;</p> <p><strong>Data availability statement</strong></p> <p>The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the&nbsp;PRIDE&nbsp;(<a href="https://www.ebi.ac.uk/pride/">https://www.ebi.ac.uk/pride/</a>) partner repository with the dataset identifier PXD055051 and 10.6019/PXD055051.</p> <p><strong>Acknowledgements</strong></p> <p>We would like to thank the Informatics, and Entomology technology platforms at the John Innes Centre for technical support.</p> <p><strong>Conflicts of Interest</strong></p> <p><strong>T</strong>he authors declare that no conflicts of interest exist.</p> <p><strong>Funding Information</strong></p> <p>This work was funded by UK Research and Innovation (UKRI) Biotechnology and Biological Sciences Research Council (BBSRC) grants to S.A.H. (BB/V008544/1 and BB/R009481/1). Additional Support as provided by the BBSRC Institute Strategy Programmes (BBS/E/J/000PR9797 and BBS/E/JI/230001B) awarded to the John Innes Centre (JIC). The JIC is grant-aided by the John Innes Foundation.</p> <p>&nbsp;</p> <p><strong>Figures and legends</strong></p> <p><strong>Fig. 1.</strong> Experimental procedure for detecting <em>M. persicae</em> secretome. (A) The workflow of <em>M. persicae</em> saliva collection and concentration for nanoLC-MS/MS. (B) The diagrammatic representation of mass spectrometry data analysis of <em>M. persicae </em>saliva.</p> <p><strong>Table 1. Full list of detected peptides.</strong></p> <p>List of peptides detected in <em>M. persicae </em>saliva from Mascot and Chimerys searches of the <em>M. persicae </em>clone O v2.1 database.</p> <p><strong>Table 2. Full list of proteins detected</strong></p> <p>List of proteins detected in <em>M. persicae </em>saliva. Closely related proteins with shared peptides are grouped into protein groups by Proteome Discoverer, the highest confidence of these is the Master protein of the group. Master proteins of the 169 protein groups are indicated in the &lsquo;Master&rsquo; column as &lsquo;Master protein&rsquo;, other proteins belonging to these groups have shared peptides. Proteins belonging to groups containing only different isoforms (i.e. splice variants) from the same gene model are indicating by the number of peptides that specifically match that gene model</p> <p><strong>Table 3. Annotated list of unique gene models representing detected proteins.</strong></p> <p>Two hundred and nineteen (219) gene models that encode at least one protein detected in the <em>M. persicae saliva. </em>For each gene model, the highest confidence detected protein is shown: either a master protein of the protein group, or else the longest isoform of that protein. Annotation is derived from Interproscan, including descriptions from pfam, GO and BlastP hits, and include &nbsp;the identities of candidate effector proteins (e.g. Mp1, Mp2) identified in Bos et al. (2010). Cathepsin B proteins are highlighted in green.</p> <p>&nbsp;<strong>&nbsp;</strong></p> <p><strong>Supplementary files:</strong></p> <p><strong>Genome annotation files:</strong></p> <p>Myzus_persicae_O_v2.0.scaffolds.fa (as described in Mathers et al. 2021)</p> <p>CloneO_v2.1_annotation_summary_stats.txt</p> <p>MYZPE13164_O_EIv2.1.annotation.gff3</p> <p>MYZPE13164_O_EIv2.1.annotation_w.functions.gff3</p> <p>MYZPE13164_O_EIv2.1.annotation.gff3.cdna.fasta</p> <p>MYZPE13164_O_EIv2.1.annotation.gff3.cds.fasta</p> <p>MYZPE13164_O_EIv2.1.annotation.gff3.metrics.txt</p> <p>MYZPE13164_O_EIv2.1.annotation.gff3.pep.fasta</p> <p>MYZPE13164_O_EIv2.1.annotation.gff3.cdna.LTPG.fasta</p> <p>MYZPE13164_O_EIv2.1.annotation.gff3.cds.LTPG.fasta</p> <p>MYZPE13164_O_EIv2.1.annotation.gff3.pep.LTPG.fasta</p> <p>Myzus_persicae_O_annotation_readme.doc</p> <p><strong>&nbsp;</strong></p> <p><strong>Literature Cited</strong></p> <p><strong>Chen Y, Singh A, Kaithakottil GG, Mathers TC, Gravino M, Mugford ST, van Oosterhout C, Swarbreck D, Hogenhout SA.</strong> (2020). An aphid RNA transcript migrates systemically within plants and is a virulence factor. Proc Natl Acad Sci U S A. 117(23):12763-12771. doi: 10.1073/pnas.1918410117.</p> <p><strong>Dedryver, C.A., Le Ralec, A., and Fabre, F.</strong> (2010). The conflicting relationships between aphids and men: a review of aphid damage and control strategies. C R Biol <strong>333, </strong>539-553.</p> <p><strong>Frejno, M., Berger, M.T., T&uuml;shaus, J., &hellip; and Wilhelm, M.</strong> (2024). Unifying the analysis of bottom-up proteomics data with CHIMERYS. bioRxiv 2024.05.27.596040; doi: https://doi.org/10.1101/2024.05.27.596040</p> <p><strong>Kang, Y. J., Yang, D. C., Kong, L., Hou, M., Meng, Y. Q., Wei L., and Gao, G.</strong> (2017). CPC2: a fast and accurate coding potential calculator based on sequence intrinsic features. Nucleic Acids Research 45, W12&ndash;W16.</p> <p><strong>Mathers, T.C., Chen, Y., Kaithakottil, G</strong>.<em>, &hellip;</em> <strong>Hogenhout, S.A. </strong>(2017)<em>.</em> Rapid transcriptional plasticity of duplicated gene clusters enables a clonally reproducing aphid to colonise diverse plant species. <em>Genome Biol</em> <strong>18</strong>, 27. https://doi.org/10.1186/s13059-016-1145-3</p> <p><strong>Mathers TC, Mugford ST, Percival-Alwyn L, Chen Y, Kaithakottil G, Swarbreck D, Hogenhout SA, van Oosterhout C.</strong> (2020) Sex-specific changes in the aphid DNA methylation landscape. Mol Ecol.&nbsp; 28(18):4228-4241. doi: 10.1111/mec.15216.</p> <p><strong>Mathers, T.C., Wouters, R.H.M., Mugford, S.T., Swarbreck, D., van Oosterhout, C., Hogenhout, S.A.</strong> (2021). Chromosome-Scale Genome Assemblies of Aphids Reveal Extensively Rearranged Autosomes and Long-Term Conservation of the X Chromosome, <em>Molecular Biology and Evolution</em> <strong>38</strong>, 856&ndash;875.</p> <p><strong>Perez-Riverol Y, Bai J, Bandla C, &hellip;Vizca&iacute;no JA</strong> (2022). The PRIDE database resources in 2022: A Hub for mass spectrometry-based proteomics evidences. Nucleic Acids Res 50(D1):D543-D552 (PubMed ID: 34723319).</p> <p><strong>Pertea, M., Pertea, G.M., Antonescu, C.M., Chang, T.C., Mendell, J.T. and&nbsp; Salzberg, S.L.</strong> (2015). StringTie enables improved reconstruction of a transcriptome from RNA-seq reads Nature Biotechnology 33,290-295. doi:10.1038/nbt.3122</p> <p><strong>Shao, M., Kingsford, C</strong>. (2017).&nbsp; Accurate assembly of transcripts through phase-preserving graph decomposition. Nat Biotechnol 35, 1167&ndash;1169 (2017). https://doi.org/10.1038/nbt.4020</p> <p><strong>Venturini, L., Caim, S.,&nbsp; Kaithakottil, G.G.,&nbsp; Mapleson, D.L., and Swarbreck, D.</strong> (2018). Leveraging multiple transcriptome assembly methods for improved gene structure annotation, GigaScience7giy093 https://doi.org/10.1093/gigascience/giy093</p> <p><strong>Wessel, D., and Fl&uuml;gge, U.I.</strong> (1984). A method for the quantitative recovery of protein in dilute solution in the presence of detergents and lipids. Analytical Biochemistry <strong>138, </strong>141-143.</p> <p><strong>Yu, Q., Paulo, J.A., Naverrete-Perea, J., McAlister, G.C., &hellip; and Gygi, S.P., and Schweppe, D.K. </strong>(2020). Benchmarking the Orbitrap Tribrid Eclipse for Next Generation Multiplexed Proteomics. <em>Analytical Chemistry</em> <strong>2020</strong> <em>92</em>, 6478-6485. DOI: 10.1021/acs.analchem.9b05685</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Illuminating oncogenic KRAS signaling by multi-dimensional chemical proteomics

<p><span>Mutated KRAS is among the most frequent activating genetic alterations in cancer. Drug discovery efforts have led to inhibitors that block mutant KRAS activity. To better understand the molecular basis of their cytostatic rather than cytotoxic effects, w</span>e performed comprehensive dose-dependent proteome-wide target deconvolution, pathway engagement, and protein expression characterization in response to KRAS, MEK, ERK, SHP2, and SOS1 inhibitors in pancreatic (KRAS G12C, G12D) and lung cancer (KRAS G12C) cell lines. Analysis of the dose-response curves available online revealed common and cell line-specific signaling networks dominated by KRAS activity. Time-dose experiments separated early ERK-driven effects from those that result from cell cycle arrest. The transition occurred without substantial proteome re-modelling but extensive changes in phosphorylation and ubiquitinylation. Our resource highlights the complexity of KRAS signaling in cancer and places a large number of new proteins and their modifications into this functional context for further exploration.</p> <p>We provide all dose-response curve data processed using internal pipelines or CurveCurator v0.5.0 (<a href="https://github.com/kusterlab/curve_curator" target="_blank" rel="noopener">https://github.com/kusterlab/curve_curator</a>). A README file is included with details about each file and a Meta table describing the experimental conditions. Each CurveCurator folder contains both the input data (including the TOML parameter file used for curve generation) and the output, which includes interactive dashboards (<strong>dashboard.html</strong>) and processed curve data (<strong>curves.txt</strong>).</p> <p>Phospho-proteome, whole proteome, ubiquitinome, Kinobead pulldown, and cysteine profiling data are provided in separate ZIP folders. Additionally, we include all aggregated supplementary tables and analysis output tables used for figure generation in the manuscript.</p>

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record