Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11,855

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

11,855 results for “Proteins”

Learn how ShareScore rates datasets ↗
zenodo56/100

Homologous membrane protein structures (HOMEP) version 1

<p><strong>Table 1</strong> = List of membrane protein structures in the <strong>HOMEP</strong>&nbsp;data set (version 1).<br> From Forrest, Tang&nbsp;&amp; Honig&nbsp;2006 Biophysical Journal (Supplementary Table 1)<br> <a href="https://www.ncbi.nlm.nih.gov/pubmed/16648166">https://www.ncbi.nlm.nih.gov/pubmed/16648166</a></p> <p>Contains the following columns:<br> PDB-Code Protein-Name &nbsp; &nbsp;Source &nbsp;Res-(&Aring;) Length (Num-TM) Number-of-TM-domains &nbsp; &nbsp;Family</p> <p><strong>Table 2</strong> =&nbsp;List of pairs of membrane protein structures in the <strong>HOMEP</strong>&nbsp;data set (version 1).<br> From Forrest, Tang&nbsp;&amp; Honig&nbsp;2006 Biophysical Journal (Supplementary Table 2)</p> <p>Contains the following columns:<br> Model &nbsp; Family &nbsp;Query &nbsp; Template &nbsp; ID(%) &nbsp;RMS(&Aring;) GDT_TS(%) &nbsp;TM-ID(%) &nbsp; TM-RMS(&Aring;) &nbsp;TM GDT_TS(%)</p> <p><strong>Table 3 </strong>= Manually-defined transmembrane regions in the <strong>HOMEP</strong>&nbsp;data set (version 1),&nbsp;listed for each family by transmembrane segment number. From Forrest, Tang&nbsp;&amp; Honig&nbsp;2006 Biophysical Journal (Supplementary Table 3).</p> <p>Contains&nbsp;the columns defined as follows:<br> Protein chain identifier, start (-s) and end (-e) residues for each PDB structure in the family</p>

opencc-by-4.0Apr 2006View details →
zenodo56/100

Cross-phyla protein annotation by structural prediction and alignment

<p><strong>Background:</strong> Protein annotation is a major goal in molecular biology, yet experimentally determined knowledge is&nbsp;typically limited to a few model organisms. In non-model species, the sequence-based prediction of&nbsp;gene orthology can be used to infer protein identity, however this approach loses predictive power&nbsp;at longer evolutionary distances. Here we propose a workflow for protein annotation using structural&nbsp;similarity, exploiting the fact that similar protein structures often reflect homology and are more&nbsp;conserved than protein sequences.</p> <p><strong>Results:</strong>&nbsp;&nbsp;We propose a workflow of openly available tools for the functional annotation of proteins via&nbsp;structural similarity (MorF: <strong>Mor</strong>pholog<strong>F</strong>inder) and use it to annotate the complete&nbsp;proteome of a sponge. Sponges are highly relevant for inferring the early history of animals, yet&nbsp;their proteomes remain sparsely annotated. MorF accurately predicts the functions of proteins with&nbsp;known homology in &gt;90%&nbsp;cases, and annotates an additional 50%&nbsp;of the proteome beyond&nbsp;standard sequence-based methods. We uncover new functions for sponge cell types, including extensive&nbsp;FGF, TGF and Ephrin signalling in sponge epithelia, and redox metabolism and control in&nbsp;myopeptidocytes. Notably, we also annotate genes specific to the enigmatic sponge mesocytes,&nbsp;proposing they function to digest cell walls.</p> <p><strong>Conclusions:</strong> Our work demonstrates that structural similarity is a powerful approach that complements and extends sequence similarity searches to identify homologous proteins over long evolutionary distances. We anticipate this to be a powerful approach that boosts discovery in numerous -omics datasets, especially for non-model organisms.</p>

opencc-by-4.0Mar 2023View details →
zenodo52/100

Dataset - Decrypting lysine deacetylase inhibitor action and protein modifications by dose-resolved proteomics

<h4><strong>Dataset Summary</strong></h4> <p>Lysine deacetylase inhibitors (KDACis) are approved for cutaneous T-cell lymphoma (CTCL), peripheral T-cell lymphoma (PTCL), and multiple myeloma. Despite the mechanism of action(s) (MoA) remains elusive, these inhibitors lead to increasing acetylation levels of histones and other proteins, altered gene expression and cell death. To characterize the MoA of these drugs in more detail, we systematically measured dose-dependent changes in protein expression, acetylation, and phosphorylation in response to 21 clinical and pre-clinical KDACis. MV4-11 cells were treated for 6 h with 1 vehicle control and 10 increasing doses of the respective drug (from 100 pM to 30 mM). Proteins were digested with trypsin, and the resulting 11 peptide preparations corresponding to one drug dose each were encoded by stable isotopes (tandem mass tags, TMT-11plex) and combined. Acetylated peptides were subsequently enriched by immunoprecipitation and phosphopeptides by immobilized metal affinity chromatography (IMAC). PTM-carrying and unmodified peptides were analyzed separately by liquid chromatography tandem mass spectrometry (LC-MS/MS) for peptide and protein identification and quantification. Additionally, Vorinostat and Panobinostat were also recorded as time-dependent experiments at their pEC50 concentration, respectively.&nbsp;</p> <h4><strong>Dataset structure</strong></h4> <p>Here, we provide all curve data processed with CurveCurator v0.4.0 (<a href="https://github.com/kusterlab/curve_curator">https://github.com/kusterlab/curve_curator</a>). Each drug is a zip folder containing acetylome, phosphoproteome, and fullproteome data. Next to each data set is the toml parameter file used to generate the curves.txt and dashboard.html files. Time-dependent data is indicated by "td" and dose-dependent data is indicated by "dd".</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Dataset / Code: Targeted protein degradation in mycobacteria uncovers antibacterial effects and potentiates antibiotic efficacy

<p><strong>Targeted protein degradation in mycobacteria uncovers antibacterial effects and potentiates antibiotic efficacy</strong></p> <p><strong>&nbsp;</strong></p> <p>Harim I. Won<sup>1,#</sup>, Samuel Zinga<sup>1,#</sup>, Olga Kandror<sup>1</sup>, Tatos Akopian<sup>1</sup>, Ian D. Wolf<sup>1</sup>, Jessica T.P. Schweber<sup>1</sup>, Ernst W. Schmid<sup>2</sup>, Michael C. Chao<sup>1</sup>, Maya Waldor<sup>1</sup>, Eric J. Rubin<sup>1,*</sup>, Junhao Zhu<sup>1,3,*</sup></p> <p><strong>&nbsp;</strong></p> <p><sup>1</sup>Department of Immunology and Infectious Diseases, Harvard T.H. Chan School of Public Health, Boston, Massachusetts 02115, USA.</p> <p><sup>2</sup>Department of Biological Chemistry and Molecular Pharmacology, Harvard Medical School, Blavatnik Institute, Boston, Massachusetts 02115, USA.</p> <p><sup>3</sup>CAS Key Laboratory of Pathogen Microbiology and Immunology, Institute of Microbiology, Chinese Academy of Sciences, Beijing, China.</p> <p><sup>#</sup>These authors contributed equally to this work.</p> <p>*Corresponding authors: <a href="mailto:zhujh@im.ac.cn">zhujh@im.ac.cn</a> (J.Z.), <a href="mailto:erubin@hsph.harvard.edu">erubin@hsph.harvard.edu</a> (E. J. R.)</p> <p><strong>&nbsp;</strong></p> <p><strong>Abstract</strong></p> <p>Proteolysis-targeting chimeras (PROTACs) represent a new therapeutic modality involving selectively directing disease-causing proteins for degradation through proteolytic systems. Our ability to exploit targeted protein degradation (TPD) for antibiotic development remains nascent due to our limited understanding of which bacterial proteins are amenable to a TPD strategy. Here, we use a genetic system to model chemically-induced proximity and degradation to screen essential proteins in <em>Mycobacterium smegmatis </em>(<em>Msm</em>)<em>, </em>a model for the human pathogen <em>M. tuberculosis </em>(<em>Mtb</em>). By integrating experimental screening of 72 protein candidates and machine learning, we find that drug-induced proximity to the bacterial ClpC1P1P2 proteolytic complex leads to the degradation of many endogenous proteins, especially those with disordered termini. Additionally, TPD of essential <em>Msm </em>proteins inhibits bacterial growth and potentiates the effects of existing antimicrobial compounds. Together, our results provide biological principles to select and evaluate attractive targets for future <em>Mtb</em> PROTAC development, as both standalone antibiotics and potentiators of existing antibiotic efficacy.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Full-length and split homologs of human proteins in the gut microbiome

<p>These files were generated as part of the manuscript "Human xenobiotic metabolism proteins have full-length and split homologs in the gut microbiome" (submitted).</p> <p>The .tar file contains .ipc files that are tables of full-length (full_humcover3.ipc) and split homologs (part_humcover3.ipc) of human proteins in the gut microbiome, organized by alignment coverage threshold. For example, the directory `HumanUPR_0.67_src_20000_70` contains results obtained at a 67% alignment coverage threshold for the bacterial protein, and 70% for the human protein. Note that our pipeline collapses full-length alignments to the same UHGP-90 protein family into a single entry per species, with the number of genomes reported in the column nGenomes. Split homologs are not collapsed because genomic context is used to define them, and this context may differ across individual genomes.</p> <p>These files are in Arrow <a href="https://arrow.apache.org/docs/python/ipc.html#ipc">IPC</a> format, which provides compression and fast I/O for large tables. We recommend reading them using <a href="https://pola.rs/">pola.rs</a> or the <a href="https://arrow.apache.org/docs/r/">R Arrow</a> package. In particular, because the full-length homolog table is large, you may wish to work with it without loading it into memory, which can be accomplished using&nbsp;<a href="https://docs.pola.rs/api/python/dev/reference/api/polars.scan_ipc.html">scan_ipc</a> in pola.rs or <a href="https://arrow.apache.org/docs/r/reference/open_dataset.html">open_dataset</a> in R Arrow.</p> <p>We also provide gzipped .csv format datasets of full-length (pgkb_FH_drugs.csv.gz) and split (pgkb_SH_drugs.csv.gz) homologs, at the default 67% alignment coverage threshold for bacterial and 70% for human proteins, organized by their&nbsp;<a href="https://www.pharmgkb.org/">PharmGKB</a> annotations. For each drug annotated in PharmGKB as being metabolized by a human protein with full-length or split homologs, we provide the human protein(s) responsible, its xenobiotic enzyme class, the bacterial protein homolog(s), length and percent identity of the alignment, and either the specific genome (g, split homologs only) or the number of genomes (nGenomes, full homologs only). Xenobiotic enzyme classes are defined as in Figure 4 of the manuscript, with the additional classes "nucl" (nucleobase-containing metabolic proteins not annotated to any other class), "redox" (oxidoreductases not annotated to any other class), and "other" (all remaining proteins).</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Protein haplotype sequences obtained by ProHap from the Haplotype Reference Consortium Release 1.1 dataset

<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Haplotype Reference Consortium, Release 1.1 (<a href="https://ega-archive.org/datasets/EGAD00001002729" target="_blank" rel="noopener">https://ega-archive.org/datasets/EGAD00001002729</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>Release 1.1 of the HRC is provided aligned with the GRCh37 reference genome. We have performed a liftover to the GRCh38 reference using GeneBe (https://genebe.net/tools/liftover). Variants for which the reported alternative allele is considered as reference in GRCh38 were removed. A threshold of 1% minor allele frequency was applied to filter the remaining variants. After translation, a frequency threshold of 0.5% was applied to filter the resulting unique non-canonical sequences. The complete configuration file for the ProHap run is attached to this repository.</p> <p>This dataset contains one compressed directory, contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file.&nbsp;</li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to&nbsp;<a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Va&scaron;&iacute;ček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>

opencc-by-4.0Jul 2024View details →
zenodo52/100

Image data of co-localization of IgG and HEV ORF2 protein in a case of hepatitis E-associated kidney disease

<p><span>Image data for a co-localization study of IgG with HEV ORF2 protein in a </span><span>de novo immune complex-mediated glomerulonephritis (GN) case in</span><span> a kidney transplant recipient </span><span>with chronic hepatitis E (Leblond and Helmchen, et al. 2024).<span>&nbsp; </span>Immunofluorescence images are provided for 25 glomeruli at low magnification (20x, 0.227 micron/pixel) and for 16 glomeruli at high magnification (100x, 0.0454 micron/pixel). For each example glomeruli the green channel represents IgG antibody staining with FITC, and the magenta channel represent anti-HEV ORF2 staining using Alexa Fluor 546.</span></p> <p><span>Methods:&nbsp;</span></p> <p><span>Mouse monoclonal antibody clone 1E6 against the HEV ORF2 protein was incubated for 1h at a dilution of 1:125 followed by a mix of Alexa Fluor 546-conjugated goat anti-mouse antibody (Invitrogen BV, A11018) and FITC-conjugated Rabbit anti-Human IgG (Gamma chain, Diagnostic Biosystem, F008) for 1hat a dilution of 1:50. Following automated staining, the slides were hand -washed in distilled H<sub>2</sub>O. Tissue was covered with Vectashield&reg; Antifade Mounting Medium with DAPI (VectorLaboratories, H-1200), covered with a coverslip and stored at 4&deg;C until evaluation.</span></p> <p><span>Immunofluorescence images were acquired with an upright fluorescence microscope (AxioImager.Z2 controlled by ZEN Blue software; 89 North Photofluor LM-75 light source, and Axiocam 503 mono camera; Zeiss, Jena, Germany), equipped with the following objectives: 20x (NA 0.5, Plan-NEOFLUAR), 40x (NA 1.4 oil, Plan-APOCHROMAT), and 100x (NA 1.45 oil, Plan-APOCHROMAT) objectives. This setup provides an excellent spatial resolution (nominally about 200 nm lateral resolution in our study; pixel size was 45.4 nm for 100x objective). High resolution images were taken with the 100x objective using the ApoTome.2 module with deconvolution (grid 5 lp/mm; section thickness 0.7 &micro;m). We used Vysis Abbott Chroma filter sets (Blue: excitation (ex) 335-383 nm, emission (em) 420-470; green: ex 481-507 nm; em 521-551 nm; red: ex 534-556 nm, em 574- 606 nm). Co-localization of IgG and HEV ORF2 staining was quantified using Fiji software (Schindelin et al., 2012) and the JACoP ImageJ plug-in. </span></p>

opencc-by-4.0Sep 2024View details →
zenodo52/100

Conserved mechanism of Xrn1 regulation by glycolytic flux and protein aggregation

<p>This dataset contains the raw and processed microscopy data that form the basis of our research article titled <em><a href="https://www.cell.com/heliyon/fulltext/S2405-8440(24)14817-7" target="_blank" rel="noopener">Conserved mechanism of Xrn1 regulation by glycolytic flux and protein aggregation</a></em> (DOI: <a href="https://kwnsfk27.r.eu-west-1.awstrack.me/L0/https:%2F%2Fdoi.org%2F10.1016%2Fj.heliyon.2024.e38786/1/010201924b2cde3a-4e9b4379-d805-432d-b94e-11d8b32354dc-000000/-WTssLbV_1JADo7ZwuFgWEmy9i0=394" target="_blank" rel="noopener noreferrer">doi.org/10.1016/j.heliyon.2024.e38786</a>) publlished in the <a href="https://www.cell.com/">CellPress</a> journal <a href="https://www.cell.com/heliyon/home">Heliyon</a> (<a title="Go to table of contents for this volume/issue" href="https://www.sciencedirect.com/journal/heliyon/vol/10/issue/19"><span><span>Volume 10, Issue 19</span></span></a>, 15 October 2024, e38786).<em>&nbsp;</em>The paper describes the mechanism underlying the binding of <a href="https://www.yeastgenome.org/locus/S000003141">yeast Xrn1</a> to the plasma membrane microdomain stabiliser <a href="https://doi.org/10.1016/j.cub.2017.11.073">eisosome</a> in a glucose-dependent manner. The images stored in the dataset were acquired with a <a href="https://www.iem.cas.cz/en/devices/zeiss-lsm-880-airyscan-en/">Zeiss LSM 880 confocal microscope</a> performed at the <a href="https://www.iem.cas.cz/en/department/microscopy-unit/">Microscopy Service Centre</a> of the <a href="https://www.iem.cas.cz/en/home-en/">Institute of Experimental Medicine CAS</a> supported by the MEYS CR (LM2023050 <a href="https://www.czech-bioimaging.cz/">Czech-Bioimaging</a>). Detailed step-by-step instructions for live microscopy sample preparation that we follow can be found at protocols.io: <a href="https://www.protocols.io/view/live-cell-microscopy-sample-preparation-yeast-cult-8epv5r23dg1b">https://www.protocols.io/view/live-cell-microscopy-sample-preparation-yeast-cult-8epv5r23dg1b</a>. The quantification of microscopy data was performed using our custom-developed Fiji and R&nbsp;scripts that can be found at&nbsp;<a href="https://github.com/jakubzahumensky/microscopy_analysis">https://github.com/jakubzahumensky/microscopy_analysis</a>. Their use is described in detail in the&nbsp;<a href="https://doi.org/10.1101/2024.03.28.587214">https://doi.org/10.1101/2024.03.28.587214</a>. For further information, please refer to the README file attached to the dataset.</p> <p>Note: This final version of datasets supplements the previous datasets of version 1 (DOI: <a href="https://doi.org/10.5281/zenodo.12748899">10.5281/zenodo.12748899</a>) and version 2 (DOI: <a href="https://doi.org/10.5281/zenodo.13772845">10.5281/zenodo.13772845</a>).</p>

opencc-by-4.0Jul 2024View details →
zenodo52/100

A novel approach to the detection of unusual mitochondrial protein change suggests hypometabolism of ancestral simians: Supplemental Files

<p><strong>Supplementary Fig. S1</strong>: &theta;<sub>evo</sub> calculated for each analyzed edge for specific OXPHOS complexes. Analyses were performed as in fig. 1F, except that SPCSs calculated from mtDNA-encoded protein positions in Complex I, Complex III, Complex IV, or Complex V were used to generate &theta;evo values.</p> <p><strong>Supplementary Fig. S2</strong>: Mammalian orders differ in their propensity for potentially efficacious mitochondrial protein substitutions within specific OXPHOS complexes (median calculations). Analysis was performed as in fig. 2A, except that &theta;<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V polypeptides.</p> <p><strong>Supplementary Fig. S3</strong>: Mammalian orders differ in their propensity for potentially efficacious mitochondrial protein substitutions within specific OXPHOS complexes (median confidence intervals). Analysis was performed as in (<em>A</em>) fig. 2B or (<em>B</em>) fig. 2C, except that &theta;<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V proteins.</p> <p><strong>Supplementary Fig. S4</strong>: Mammalian families differ in their propensity for potentially efficacious mitochondrial protein substitutions at specific OXPHOS complexes (median calculations). Analysis was performed as in fig. 3A, except that &theta;<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V subunits.</p> <p><strong>Supplementary Fig. S5</strong>: Mammalian families differ in their propensity for potentially efficacious mitochondrial protein substitutions at specific OXPHOS complexes (median confidence intervals ordered by lower 90% median confidence limit). Analysis was performed as in fig. 3B, except that &theta;<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V proteins.</p> <p><strong>Supplementary Fig. S6</strong>: Mammalian families differ in their propensity for potentially efficacious mitochondrial protein substitutions at specific OXPHOS complexes (median confidence intervals ordered by upper 90% median confidence limit). Analysis was performed as in fig. 3C, except that &theta;<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V polypeptides.</p> <p>---</p> <p><strong>Supplementary File 1</strong>: All predicted protein substitutions along all edges at positions containing less than 2% gaps across input and ancestral sequences are listed, along with associated taxonomy information, TSS, and branch length. All alignment positions refer to Bos taurus reference sequences.</p> <p><strong>Supplementary File 2</strong>: The TSS calculated for each mitochondrial protein alignment position. All alignment positions refer to Bos taurus reference sequences.</p> <p><strong>Supplementary File 3</strong>: SPCS and &theta;evo outputs are provided for analyses across all mitochondria-encoded positions, as well as for focused analyses of specific OXPHOS complexes and individual proteins.</p> <p><strong>Supplementary File 4</strong>: A GenBank flat file containing RefSeq entries for mammalian mtDNAs, as well as the entry for the reptile Anolis punctatus.</p> <p><strong>Supplementary File 5</strong>: A maximum likelihood inferred tree generated by a RAxML-NG analysis of concatenated and aligned protein coding sequences from mammalian and Anolis punctatusmtDNAs.</p> <p><strong>Supplementary File 6</strong>: Bootstrap replicates were generated from the alignment of concatenated protein coding sequences. Felsenstein&rsquo;s Bootstrap Proportions (Felsenstein 1985) were calculated and used to label the maximum likelihood inferred tree of mammalian mtDNAs.</p> <p><strong>Supplementary File 7</strong>: Bootstrap replicates were generated using concatenated mammalian mtDNA coding sequences. Transfer Bootstrap Expectations (Lemoine 2018) were calculated and used to label the maximum likelihood inferred tree of mammalian mtDNAs.</p> <p><strong>Supplementary File 8</strong>: PAGAN tree output produced using aligned amino acid sequences and the rooted maximum likelihood inferred tree as input.</p>

opencc-by-4.0Aug 2021View details →
zenodo52/100

Structures of S-protein in complex with ligands deposited in the PDB between the 1st January 2021 and the 13th May 2021

<p>All 174 structures of SARS-CoV-2 S-protein in complex with a ligand released on the PDB between the 1<sup>st</sup> January 2021 and the 13<sup>th</sup> May 2021 were categorised by ligand type: hACE2, antibody Fab fragments, VHH antibody fragments or <em>de novo</em> designed peptide scaffolds. Information concerning the method by which the structures were determined and their resolution were retrieved from the PDB. The categorisation of ligands by S-protein binding site were achieved by visual analysis of all the structures using molecular visualisation software PyMOL, in which no new binding sites were found beyond those already categorised for the structures released on the PDB until the 1<sup>st</sup> January 2021 (10.5281/zenodo.5503855).</p> <p>The Pure project is funded by the European Union&rsquo;s Horizon 2020 program under grant agreement No. 899732.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo52/100

List of the structures of S-protein in complex with ligands deposited in the Protein Data Bank until the 1st January 2021.

<p>All 131 structures of SARS-CoV-2 S-protein in complex with a ligand released on the PDB until the 1<sup>st</sup> January 2021 were categorised by ligand type: hACE2, antibody Fab fragments, VHH antibody fragments or <em>de novo</em> designed peptide scaffolds. The ligands&rsquo; amino acid sequences, the method by which the structures were determined and their resolution were retrieved from the PDB. Information regarding the ligands&#39; production method, dissociation constants (K<sub>D</sub>), S-protein segment against which the K<sub>D</sub> were measured and the determination methods were retrieved from the respective references. The categorisation of ligands by S-protein binding site and listing of S-protein conformation in each structure were achieved by visual analysis of all the structures using molecular visualisation software PyMOL.</p>

opencc-by-4.0Sep 2021View details →
edi52/100

Examination of protein-like fluorophores in chromophoric dissolved organic matter (CDOM) in a wetland and coastal environment for the wet and dry seasons of the years 2002 and 2003 (FCE)

Water samples are collected at the end of the dry and the wet season from all LTER sites and stored on ice until return to the lab. They are pre-filtered through pre-combusted GF/F filters and ultrafiltered and concentrated with a Pellicon 2 Mini tangential flow ultrafiltration system.Concentrated samples were then analyzed using fluorescence and SEC-HPLC. This CDOM optical study revealed the presence of two classes of compounds associated with the protein-like peak (peak T; excitation/emission (Ex/Em) maxima at around 280 nm/325 nm), which have very different chemical structures and ecological roles. In addition to proteins, we propose phenolic compounds as possible origins of peak T in coastal and wetland environments. In this study, natural water samples were obtained from subtropical rivers and estuarine environments within the Florida Coastal Everglades (FCE) ecosystem. The samples were ultra-filtered and excitation-emission fluorescence matrices (EEMs) were obtained. The EEMs showed the presence of four peaks with Ex/Em maxima at around 280 nm/325 nm (T), less than 260 nm/460 nm (A), 300 nm/412nm (M), and 350 nm/470 nm (C). To better understand the nature of peak T, the components originating this peak were separated using size exclusion chromatography (SEC) and detected by fluorescence emission at Ex/Em = 280 nm/325 nm. The elution curves revealed the presence of two elution peaks at a molecular weight of greater than 50K (void volume; T1) and around 7.6K (T2). This result suggested the need of cautious interpretation in the use of peak T as a proxy for the detection of proteinaceous materials in wetland and estuarine environments, since significant amounts of potentially interfering phenolic compounds are leached from senescent biomass in wetland and coastal ecosystems. As such EEM spectra of gallic acid an important component of hydrolysable tannins, and condensed tannins extracted from red mangroves (Rhizophora mangle) showed the presence of a peak maxima

openCC (other)Feb 2024View details →
zenodo48/100

Raw data to accompany the manuscript 'Data for Engineering Lipid Metabolism of Chinese Hamster Ovary (CHO) Cells for Enhanced Recombinant Protein Production' published in the Journal Data in Brief

<p>This repository consists of the raw western blot, microscopy and mass spectrometry data to accompany the manuscript &#39;Data for Engineering Lipid Metabolism of Chinese Hamster Ovary (CHO) Cells for Enhanced Recombinant Protein Production&#39; published in the Journal Data in Brief and associated with the article &#39;<a href="https://www.ncbi.nlm.nih.gov/pubmed/31805379">Engineering of Chinese hamster ovary cell lipid metabolism results in an expanded ER and enhanced recombinant biotherapeutic protein production</a>&#39; published in the journal Metabolic Engineering (see DOI:&nbsp;10.1016/j.ymben.2019.11.007).&nbsp;</p> <p>The western blot raw file is associated with Figure 1a and 1b of the Data in Brief manuscript.</p> <p>The confocal microscopy raw image files (x3) are associated with Figure 1c&nbsp;of the Data in Brief manuscript.</p> <p>The mass spectrometry files are the raw data that refers to the samples presented in Figure 5 of the Data in Brief manuscript. Files are labelled as in the Data in Brief and Metabolic Engineering manuscripts. The file name structures is as follows;</p> <p>CHO-Controlpoolai</p> <p>Where &#39;a&#39; represents replicate &#39;a&#39; of three biological replicates and &#39;i&#39; refers to mass spectrometry technical analysis 1 of 3 technical analyses of each replicate (thus for each cell pool or line there are three biological replicates that are each analysed in triplicate such that there are 9 raw mass spectrometry files for each cell pool or line).</p> <p>All the mass spectrometry files are found in the compressed (zip) file named mass_spectrometry_raw_files_archive.zip</p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

Murine norovirus virulence factor 1 (VF1) protein contributes to viral fitness during persistent infection [Primary data]

<p>Primary data underlying journal article titled &quot;<strong>Murine norovirus virulence factor 1 (VF1) protein contributes to viral fitness during persistent infection</strong>&quot;</p>

opencc-by-4.0Mar 2020View details →
zenodo48/100

Data From: Exploring Gelatin-A and Mouse Proline-Rich Protein 5 as Probes for Wine Polyphenols analysis by Quartz Crystal Microbalance with Dissipation Monitoring

<p>Polyphenols are essential in winemaking, affecting the wine's quality, color, astringency, bitterness, and chemical stability. Conventional methods for assessing polyphenolic content are both expensive and time-intensive, underscoring the need for new, efficient techniques.</p> <p>The Quartz Crystal Microbalance with Dissipation Monitoring (QCM-D) sensor is recognized for its speed and reliability as a label-free detection tool. This study applies QCM-D to evaluate Gelatin Type A (Gel-A) from porcine skin and Mouse Proline-Rich Protein 5 (MP5) for polyphenol analysis in red wines without pre-treatment. MP5 notably exhibited a linear dissipation signal response with both total polyphenol and hydroxybenzoic acid concentrations. These findings highlight the potential for creating a stand-alone sensor platform for real-time polyphenol monitoring in winemaking.</p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Bioactivity deep learning for structure-free compound-protein interaction

<p>CPI2M data for "<strong>Bioactivity deep learning for structure-free compound-protein interaction</strong>".</p> <p>CPI2M_main_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_Kd.csv: Bioactivity data with <strong>pKd</strong> activity type. Used for model training and internal validation.</p> <p>CPI2M_main_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_main_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for model training and internal validation.</p> <p>CPI2M_few_Ki.csv: Bioactivity data with <strong>pKi </strong>activity type. Used for external validation.</p> <p>CPI2M_few_Kd.csv: Bioactivity data with <strong>pKd </strong>activity type. Used for external validation.</p> <p>CPI2M_few_EC50.csv: Bioactivity data with <strong>pEC50 </strong>activity type. Used for external validation.</p> <p>CPI2M_few_IC50.csv: Bioactivity data with <strong>pIC50 </strong>activity type. Used for external validation.</p> <p>potency.csv: BIoactivity data with <strong>pPotency </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>percentage.csv: BIoactivity data with <strong>Percentage Inhibition </strong>activity type. Not used currently but can be potentially adopted as classification data for customized use.</p> <p>Protein_pretrained_feat.zip: pre-calculated protein feature files with UniProt ID naming. <strong>Should be unzipped</strong> before start model training with CPI2M data.</p> <p>&nbsp;</p> <p>For each .csv data, columns include "<strong>smiles</strong>" (ligand SMILES), "<strong>exp_mean</strong>" (nM bioactivity), "<strong>y</strong>" (neg.log nM, final label), "<strong>cliff_mol</strong>" (whether activity cliff or not), "<strong>split</strong>" (splitting label by activity cliff), "<strong>Uniprot_id</strong>" (UniProt ID for protein), "<strong>Sequence</strong>" (wildtype sequence for protein), and "type_id" (bioactivity type token, pKi =0, pKd=1, pEC50=2, pIC50=3).</p> <p>&nbsp;</p> <p>Please find the project code at https://github.com/gu-yaowen/GGAP-CPI</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

CPTAC TMT protein quantifications imputed with Lupine

<p>TMT proteomics data collected from clinical patient samples as part of the Clinical Tumor Atlas Consortium (CPTAC) project were imputed with Lupine, a deep matrix factorization-based proteomics imputation method. All quantifications have been summarized at the protein level.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Overcoming Limitation of AlphaFold2 by Deep-mutational Scanning and Stability-Selection of Protein Sequences

<p>This repository contains the processed datasets and corresponding code used in our study. While AlphaFold2 revolutionizes protein structure prediction, its accuracy critically depends on evolutionary information from natural homologs&mdash;limiting applications for proteins with sparse sequence families. Here, we bypass this bottleneck by employing deep mutational scanning and stability-guided selection to generate artificial homologs. Fed into AlphaFold2, these synthetic sequences match the accuracy achieved on well-predicted proteins with rich natural homology, while providing highly accurate predictions for difficult targets&mdash;including orphan proteins previously deemed "unpredictable." Our approach achieves high accuracy (&lt;3 &Aring; RMSD for 5/8 and &lt;2 &Aring; RMSD for 7/8 targets after excluding intrinsically flexible regions). Thus, integrating simple, scalable molecular biology (mutagenesis/selection) with high-throughput sequencing can deliver the accuracy similar to but at a fraction of the cost and time of traditional experimental structure-determination methods. This hybrid framework could democratize high-resolution structural biology, opening avenues to determine structures of protein complexes, modified proteins, and condition-dependent conformations.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo48/100

Characterization of a loss-offunction NSF attachment protein beta mutation in monozygotic triplets affected with epilepsy and autism using cortical neurons from proband-derived and CRISPR-corrected induced pluripotent stem cell lines

<p>RNA-seq data of matured cortical neurons (8-weeks old) derived from the induced pluripoent stem cells (iPSC) of control parents (CtrlF and CtrlM) and corrected proband. There are three replicates (Rep1, Rep2, Rep3) for each sample&nbsp; with Forwad read (R1_001.fastq.gz)</p> <p>CtrlF:&nbsp; Control Father sample</p> <p>CtrlM: Control mother sample</p> <p>NDD_01_Corr_Het: Heterozygous correction of NAPB mutation (c.354+2T&gt;G) in NDD_01 proband</p> <p>NDD_05_Corr_Hom: Homozygous correction of NAPB mutation (c.354+2T&gt;G) in NDD_05 proband</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Data for manuscript: Functional Protein Dynamics in a Crystal

<p>The data is provided as a part of the manuscript&nbsp;&quot;<strong>Functional Protein Dynamics in a Crystal</strong>&quot;.&nbsp; This repository includes an archive with folders:<br> <br> <strong>md_data</strong></p> <ul> <li>contains various simulation systems (crystal supercell, apo and ligand-bound solution) built from the crystal structure of the PDZ domain (PDB ID: 5E11) and carried out using three force fields: Amber ff14SB, CHARMM36m, Amber ff94.&nbsp;<em>The details of the simulations are provided in the Methods and Supplementary methods sections of&nbsp;the&nbsp;manuscript.&nbsp;</em></li> </ul> <p><strong>fig_data</strong></p> <ul> <li>contains the data sets underlying Figures 1-5 of the manuscript&#39;s main text.&nbsp;</li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record