Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

347

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

347 results for “Protein structures”

Learn how ShareScore rates datasets ↗
zenodo44/100

Data for "SeaMoon: from protein language models to continuous structural heterogeneity"

<p>Datasets used for development of SeaMoon:&nbsp;<br><a href="https://github.com/PhyloSofS-Team/seamoon">https://github.com/PhyloSofS-Team/seamoon</a>.</p> <p>This upload contains the following data:</p> <ul> <li><strong>precomputed_emb.tar.gz</strong> is a compressed archive containing the precomputed data used for training and testing the models of the SeaMoon method, in Torch <strong>.pt </strong>format.&nbsp;<br>The file prefixes consist of two IDs, "ID1_ID2_", identifying the <a href="https://github.com/PhyloSofS-Team/DANCE">DANCE</a> [1] protein conformational collection used for its generation. "ID1" represents the first member of the collection in alphabetical order, while "ID2" is the reference conformation for the structural alignment. The "ESM_data" or "ProstT5_data" suffixes designate the type of embeddings, generated by either ESM2 [2] or ProstT5 [3].<br>The dictionnary contains the following keys: <ul> <li><strong>emb:</strong> The per-residue embedding.</li> <li><strong>data: </strong>A tuple containing "ID2" (the reference), the amino acid sequence, and the coverage of the positions in the original DANCE collection.</li> <li><strong>eigvect:</strong> The eigenvectors of the covariance matrix of the "ID1_ID2" collection, centered on reference conformaton "D2".</li> <li><strong>eigval:&nbsp;</strong>The associated eigenvalues.</li> <li><strong>ref:</strong> The coordinates of the C-alpha atoms of the reference conformaton "ID2".</li> </ul> </li> <li><strong>train_list.txt, train_list_5ref.txt, val_list.txt </strong>and<strong> test_list.txt</strong> contain the identifiers of the samples used for training and evaluating the SeaMoon models. In the "5ref" setting, we used up to 5 reference conformations per collection.&nbsp;</li> </ul> <p>For details on SeaMoon see:</p> <div> <div>SeaMoon: Prediction of molecular motions based on language models</div> </div> <div>Valentin Lombard, Dan Timsit, Sergei Grudinin, Elodie Laine</div> <div>bioRxiv 2024.09.23.614585; doi: https://doi.org/10.1101/2024.09.23.614585</div> <div>&nbsp;</div> <div>For more information on data usage and generation please see <a href="https://github.com/PhyloSofS-Team/seamoon">https://github.com/PhyloSofS-Team/seamoon</a>.</div> <div>&nbsp;</div> <div>Abstract:</div> <p>How protein move and deform determines their interactions with the environment and is thus of utmost importance for cellular functioning. Following the revolution in single protein 3D structure prediction, researchers have focused on repurposing or developing deep learning models for sampling alternative protein conformations. In this work, we explored whether continuous compact representations of protein motions could be predicted directly from protein sequences, without exploiting nor sampling protein structures. Our approach, called SeaMoon, leverages protein Language Model (pLM) embeddings as input to a lightweight (~1M trainable parameters) convolutional neural network. SeaMoon achieves a success rate of up to 40% when assessed against ~1,000 collections of experimental conformations exhibiting a wide range of motions. SeaMoon capture motions not accessible to the normal mode analysis, an unsupervised physics-based method relying solely on a protein structure's 3D geometry, and generalises to proteins that do not have any detectable sequence similarity to the training set. SeaMoon is easily retrainable with novel or updated pLMs.&nbsp;</p> <p>&nbsp;</p> <p>[1] Lombard, V.; Grudinin, S.; Laine, E. Explaining Conformational Diversity in Protein Families through Molecular Motions. Scientific Data 2024, 11, 752.</p> <p>[2] Lin, Z.; Akin, H.; Rao, R.; Hie, B.; Zhu, Z.; Lu, W.; Smetanin, N.; Verkuil, R.; Kabeli, O.; Shmueli, Y.; Dos Santos Costa, A.; Fazel-Zarandi, M.; Sercu, T.; Candido, S.; Rives, A. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023, 379, 1123&ndash;1130.</p> <p>[3] Heinzinger, M.; Weissenow, K.; Sanchez, J. G.; Henkel, A.; Steinegger, M.; Rost, B. ProstT5: Bilingual language model for protein sequence and structure. bioRxiv 2023, 2023&ndash;07.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

PhasAGE Training School 1 - Structure and protein interactions of repeated and low complexity regions - LECTURE

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction

<p>This dataset contains replication data for the paper titled &quot;DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction&quot;. The dataset consists of pickled Pandas DataFrame files, along with training, validation, and (for DB5-Plus) test&nbsp;filename lists for cross-validation, that can be used to develop and evaluate&nbsp;protein interface prediction models. This dataset also contains the externally generated residue-level PSAIA and HH-suite3 features for users&#39; convenience (e.g. raw MSAs and profile HMMs for each protein complex).&nbsp;Our GitHub repository linked in the &quot;Additional notes&quot; metadata section below provides more details on how we parsed through these files to create our cross-validation&nbsp;datasets. The GitHub repository&nbsp;for DIPS-Plus&nbsp;also includes scripts that can be used&nbsp;to impute missing feature values and convert the&nbsp;final &quot;raw&quot; complexes into DGL-compatible graph objects. Since our final DGL graph representation for each complex uses PyTorch tensors in its construction of&nbsp;residue&nbsp;embeddings, the final representation of each complex can easily be adapted to fit the users&#39; needs (e.g. feeding a complex&#39;s&nbsp;2D residue feature tensors into a convolutional neural network).</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Data for: The structure of evolutionary model space for proteins across the tree of life

<p>Supporting data for &quot;The structure of evolutionary model space for proteins across the tree of life,&quot;&nbsp;submitted by GE Scolaro&nbsp;and EL Braun. The data files correspond to three gzipped tarballs including protein multiple sequence alignments, PAML format models of protein evolution, and model fit data; see included README for details.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

ESM Atlas v0 representative random sample of predicted protein structures

<p>A representative random sample of the ESM Atlas v0 dataset introduced in &quot;Evolutionary-scale prediction of atomic level protein structure with a language model.&quot;.<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> Sample size: 997,405.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

ESM Atlas v0 random sample of high confidence predicted protein structures

<p>A random sample out of the 225M high confidence predictions in the ESM Atlas v0 dataset introduced in &quot;Evolutionary-scale prediction of atomic level protein structure with a language model.&quot;.<br> All predictions can be accessed in the ESM Metagenomic Atlas (<a href="https://esmatlas.com/">https://esmatlas.com</a>) open science resource, released on 2022-11-01.<br> High confidence is defined as mean pLDDT &gt; 0.7 and pTM &gt; 0.7 and corresponds to &sim;36% of the total 617M proteins folded.<br> This is the random sample used for analysis in the paper as well as visualization on the&nbsp;<a href="http://esmatlas.com/">esmatlas.com</a>&nbsp;Explore page.<br> Sample size: 999,520 based on 999,996 unique randomly sampled IDs and 0.05% missing data in the processing pipeline.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction (Supplementary Data)

<p>This dataset contains supplementary replication data for the paper titled &quot;DIPS-Plus: The Enhanced Database of Interacting Protein Structures for Interface Prediction&quot;. In particular, it contains a new version of our `final_raw_dips.tar.gz` protein pair representations which now contain (1) residue-level&nbsp;annotations for intrinsic disorder regions (IDRs) as well as (2) a copy of each protein pair representation in the HDF5 file format for programming language-agnostic read capabilities. In addition, this record also contains (3) raw MSAs (in HDF5 file format)&nbsp;generated for each protein pair using Jackhmmer and AlphaFold&#39;s small version of the Big Fantastic Database (BFD). Lastly, this record contains (4) PDB metadata derived for each DIPS-Plus complex using Graphein&#39;s PDBManager&nbsp;API&nbsp;as well as (5) structure-based (i.e., FoldSeek-based) training and validation splits of the dataset&#39;s complexes in the form of respective text files containing the file paths of complexes assigned to each split.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Genetic diversity, population structure, and linkage disequilibrium among tropical quality protein maize (QPM) lines assessed with high-density SNP markers

<p>The study of genetic diversity (GD), population structure, and linkage disequilibrium (LD) provides a better understanding of the genetic relationships between individuals in a population which can be utilized in crop research and improvement. Genotyping-by-sequencing (GBS) was used to detect and genotype single nucleotide polymorphisms (SNPs) in a collection of 74 quality protein maize (QPM) lines and further to characterize their genetic diversity, population structure, and linkage disequilibrium. A total of 235,214 high-quality SNPs were used for different genetic analyses except for structure analysis where 11,950 SNPs were used. Analysis of molecular variance (AMOVA) based on these SNPs revealed high genetic heterozygosity among the five populations with 1% of the total genetic variation present among the subpopulations and 99% of the variation among individuals within the populations. &nbsp;Population structure analysis using Bayesian-based clustering revealed that the 74 lines could be clustered into four groups. However, neighbor-joining trees indicate the lines are grouped into three major clusters.&nbsp; Further analysis using principal component analyses (PCA) clustered the genotypes into five groups which are concordant with the groups based on pedigree information. Higher genetic diversity was detected in population 1 with a GD value of 0.484 and the lowest in population 5 (0.396) and overall, with a mean of 0.434. The LD pattern in the quality protein maize was investigated and we observed a relatively rapid LD decay of 3.53kb and 10.66kb at r<sup>2</sup> =0.2 and r<sup>2</sup>= 0.1, respectively. Our findings provide important information for future Linkage mapping studies, genome-wide association analyses, and marker-assisted selective breeding of maize as well as genomic prediction-based selection in tropical germplasm.</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Protein secondary-structure description with a coarse-grained model: code and datasets in ActivePapers format

<p>This file contains the supplementary material for the publication</p> <p><em>Protein secondary-structure description with a coarse-grained model</em><br /> by Gerald R. Kneller and K. Hinsen<br /> http://dx.doi.org/10.1107/S1399004715007191<br /> Acta Cryst. (2015). D<strong>71</strong>, 1411-1422</p> <p><strong>Datasets in this file</strong></p> <p>1) ScrewFit and ScrewFrame parameters for ideal secondary-structure elements</p> <p>&nbsp;&nbsp; Scripts:<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/import_ideal_structures<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/analyze_ideal_structures</p> <p>1.1) The PDB files generated with Chimera</p> <p>&nbsp;&nbsp; /data/ideal_structures/3-10.pdb<br /> &nbsp;&nbsp; /data/ideal_structures/alpha.pdb<br /> &nbsp;&nbsp; /data/ideal_structures/beta-antiparallel.pdb<br /> &nbsp;&nbsp; /data/ideal_structures/beta-parallel.pdb<br /> &nbsp;&nbsp; /data/ideal_structures/pi.pdb</p> <p>1.2) The corresponding MOSAIC datasets</p> <p>&nbsp;&nbsp; /data/ideal_structures/3-10<br /> &nbsp;&nbsp; /data/ideal_structures/alpha<br /> &nbsp;&nbsp; /data/ideal_structures/beta-antiparallel<br /> &nbsp;&nbsp; /data/ideal_structures/beta-parallel<br /> &nbsp;&nbsp; /data/ideal_structures/pi</p> <p>1.3) The ScrewFit parameters</p> <p>&nbsp;&nbsp; /data/ideal_structures/screwfit/3-10<br /> &nbsp;&nbsp; /data/ideal_structures/screwfit/alpha<br /> &nbsp;&nbsp; /data/ideal_structures/screwfit/beta-antiparallel<br /> &nbsp;&nbsp; /data/ideal_structures/screwfit/beta-parallel<br /> &nbsp;&nbsp; /data/ideal_structures/screwfit/pi</p> <p>1.4) The ScrewFrame parameters</p> <p>&nbsp;&nbsp; /data/ideal_structures/screwframe/3-10<br /> &nbsp;&nbsp; /data/ideal_structures/screwframe/alpha<br /> &nbsp;&nbsp; /data/ideal_structures/screwframe/beta-antiparallel<br /> &nbsp;&nbsp; /data/ideal_structures/screwframe/beta-parallel<br /> &nbsp;&nbsp; /data/ideal_structures/screwframe/pi</p> <p><br /> 2) Statistics for ScrewFit and ScrewFrame parameters computed<br /> &nbsp;&nbsp; for the ASTRAL SCOPe subset with less than 40% sequence identity.</p> <p>&nbsp;&nbsp; Scripts:<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/astral_analysis<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/fit_rho_distributions<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/plot_histograms</p> <p>2.1) The ASTRAL database (link to published ActivePaper)</p> <p>&nbsp;&nbsp; /data/astral_2.04</p> <p>2.2) The histograms for the ScrewFit and ScrewFrame parameters<br /> &nbsp;&nbsp;&nbsp;&nbsp; for the all-alpha and all-beta subsets</p> <p>&nbsp;&nbsp; /data/histograms/astral_alpha/screwfit<br /> &nbsp;&nbsp; /data/histograms/astral_alpha/screwframe</p> <p>&nbsp;&nbsp; /data/histograms/astral_beta/screwfit<br /> &nbsp;&nbsp; /data/histograms/astral_beta/screwframe</p> <p>2.3) The Gaussians fitted to the peaks in the distributions for rho</p> <p>&nbsp;&nbsp; /data/fitted_rho_distributions/screwfit<br /> &nbsp;&nbsp; /data/fitted_rho_distributions/screwframe</p> <p>2.4) Plots</p> <p>&nbsp;&nbsp; /documentation/delta.pdf<br /> &nbsp;&nbsp; /documentation/delta_q.pdf<br /> &nbsp;&nbsp; /documentation/delta_r.pdf<br /> &nbsp;&nbsp; /documentation/p.pdf<br /> &nbsp;&nbsp; /documentation/rho-detail.pdf<br /> &nbsp;&nbsp; /documentation/rho.pdf<br /> &nbsp;&nbsp; /documentation/sigma.pdf<br /> &nbsp;&nbsp; /documentation/tau.pdf</p> <p><br /> 3) Comparison of secondary-structure identification between ScrewFrame<br /> &nbsp;&nbsp; and DSSP.</p> <p>&nbsp;&nbsp; Script:<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/compare_secondary_structure_assignments<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/plot_histograms</p> <p>3.1) The histograms of the lengths of secondary-structure elements</p> <p>&nbsp;&nbsp; /data/histograms/secondary_structure/length-alpha-dssp<br /> &nbsp;&nbsp; /data/histograms/secondary_structure/length-alpha-screwframe<br /> &nbsp;&nbsp; /data/histograms/secondary_structure/length-beta-dssp<br /> &nbsp;&nbsp; /data/histograms/secondary_structure/length-beta-screwframe</p> <p>3.2) The 2D histograms of the number of residues inside identified<br /> &nbsp;&nbsp;&nbsp;&nbsp; secondary-structure elements</p> <p>&nbsp;&nbsp; /data/histograms/secondary_structure/n-alpha<br /> &nbsp;&nbsp; /data/histograms/secondary_structure/n-beta</p> <p>3.3) The distribution of rho inside alpha helices</p> <p>&nbsp;&nbsp; /data/histograms/secondary_structure/rho-alpha-dssp</p> <p>3.3) Plots</p> <p>&nbsp;&nbsp; /documentation/lengths-alpha.pdf<br /> &nbsp;&nbsp; /documentation/lengths-beta.pdf<br /> &nbsp;&nbsp; /documentation/n-alpha.pdf<br /> &nbsp;&nbsp; /documentation/n-beta.pdf<br /> &nbsp;&nbsp; /documentation/rho-alpha-dssp.pdf</p> <p><br /> 4) Illustration for myoglobin and VADC-1</p> <p>&nbsp;&nbsp; Scripts:<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/import_myoglobin_vdac<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/analyze_myoglobin<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/analyze_vdac<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/perturbation_analysis</p> <p>4.1) Imported structures in MOSAIC format:<br /> &nbsp;&nbsp;&nbsp;&nbsp; PDB code 1A6G for myoglobin<br /> &nbsp;&nbsp;&nbsp;&nbsp; PDB code 2K4T for VDAC-1</p> <p>&nbsp;&nbsp; /data/myoglobin<br /> &nbsp;&nbsp; /data/VDAC-1</p> <p>4.2) Plots showing rho and delta</p> <p>&nbsp;&nbsp; /documentation/rho-myoglobin.pdf<br /> &nbsp;&nbsp; /documentation/delta-myoglobin.pdf</p> <p>4.3) Tube models for visualization with Chimera</p> <p>&nbsp;&nbsp; /documentation/myoglobin-tube.bld<br /> &nbsp;&nbsp; /documentation/VDAC-1-tube.bld</p> <p>4.4) Sensitivity to perturbations in the coordinates</p> <p>&nbsp;&nbsp; /documentation/rho-perturbed-myoglobin.pdf<br /> &nbsp;&nbsp; /documentation/delta-perturbed-VDAC-1.pdf<br /> &nbsp;&nbsp; /documentation/rho-perturbed-myoglobin.pdf<br /> &nbsp;&nbsp; /documentation/delta-perturbed-VDAC-1.pdf<br /> &nbsp;&nbsp; /documentation/myoglobin-perturbation.pdf<br /> &nbsp;&nbsp; /documentation/VDAC-1-perturbation.pdf</p> <p>5) Analysis of CA-only structures in the PDB</p> <p>&nbsp;&nbsp; Scripts:<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/ca_analysis<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/import_calpha_structures<br /> &nbsp;&nbsp;&nbsp;&nbsp; /code/plot_histograms</p> <p>5.1) Imported CA-only structures in MOSAIC format</p> <p>&nbsp;&nbsp; /data/pdb_ca_only_structures</p> <p>5.2) Histograms for ScrewFrame parameters</p> <p>&nbsp;&nbsp; /data/histograms/ca_only_structures</p> <p>5.3) Plots</p> <p>&nbsp;&nbsp; /documentation/delta_ca.pdf<br /> &nbsp;&nbsp; /documentation/delta_q_ca.pdf<br /> &nbsp;&nbsp; /documentation/delta_r_ca.pdf<br /> &nbsp;&nbsp; /documentation/p_ca.pdf<br /> &nbsp;&nbsp; /documentation/rho_ca.pdf<br /> &nbsp;&nbsp; /documentation/sigma_ca.pdf<br /> &nbsp;&nbsp; /documentation/tau_ca.pdf</p> <p>&nbsp;</p>

opencc-zeroJul 2015View details →
zenodo40/100

Protein Structure Initiative Publications, 2000-2016

<p>These files contain the full list of the 2313 publications and book chapters written as part of the Protein Structure Initiative, from 2000-2016.  This data was collected from PubMed and by manual entry by the PSI Structural Biology Knowledgebase's Publication Portal, managed by Wladek Minor at University of Virginia.  These files were created at the end of the PSI project on June 30, 2017.</p> <p>The references are provided as lists in two formats:</p> <ul> <li>in CSV (comma-separated variables) format that can be read in Excel or other spreadsheet application, or</li> <li>an Endnote Library Import file (store-endnote-pubs).  To import this library into Endnote, select File --&gt; Import... and then under Options, select the Import Option "Endnote Library Import".  Then this text file will be processed and loaded into the library.</li> </ul> <p>--created by the Structural Biology Knowledgebase, July 5, 2017.  (sbkb.org)</p>

opencc-by-sa-4.0Jun 2017View details →
zenodo40/100

III PhasAGE International Conference - PED in 2024: improving the community deposition of structural ensembles for intrinsically disordered proteins - Lecture

<p>The&nbsp;III PhasAGE International Conference&nbsp;"Multiscale understanding of protein aggregation and biomolecular condensates in aging and disease" brought together members of the PhasAGE consortium as well as outstanding international speakers from multidisciplinary fields dedicated to unraveling the intricacies of protein aggregation and biomolecular condensates in the context of aging and disease. For details on the conference program please see&nbsp;https://phasage.eu/iii-phasage-international-conference/.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Dataset for "Computational prediction of structure, function and interaction of Myzus persicae (green peach aphid) salivary effector proteins "

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Small molecules targeting the structural dynamics of AR-V7 partially disordered protein using deep learning and physics based models.

<p>Partially disordered proteins can contain both stable and unstable secondary structure segments and &nbsp;are involved in various (mis)functions in the cell. The extensive conformational dynamics of partially disordered proteins scaling with extent of disorder and length of the protein hampers the efficiency of traditional experimental and in-silico structure-based drug discovery approaches. Therefore new efficient paradigms in drug discovery taking into account conformational ensembles of proteins need to emerge. In this study, using as a test case the AR-V7 transcription factor splicing variant related to prostate cancer, we present an automated &nbsp;methodology that can accelerate the screening of small molecule binders targeting partially disordered proteins. By swiftly identifying the conformational ensemble of AR-V7, and reducing the dimension of binding-sites by a factor of 90 by applying appropriate physicochemical filters, &nbsp;we combine physics based molecular docking and multi-objective classification machine learning models that speed up the screening of thousands of compounds targeting AR-V7 multiple binding sites. Our method not only identifies previously known binding sites of AR-V7, but also discovers new ones, as well as increases the multi-binding site hit-rate of small molecules by a factor of 17 compared to naive physics-based molecular docking.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

(U)SAXS data (ID02 beamline, ESRF): Effects of pH on the fibrous structure formation of plant proteins during high-moisture extrusion

<p>Due to health and environmental factors, the food industry is looking for ways to introduce meat replacers made from plant-based proteins to consumer markets. The presence of structural anisotropy in the form of fibre is a prerequisite for meat analogues. Structure formation ability depends on the protein ingredients used, which leads to plant protein products with varying texture hardness and extent of fibre alignment. In the current study, we will test if it is possible to tune these properties based on the hypothesis that plant proteins have different structure formation abilities under varying pH conditions.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Data supporting "Slowest-first translation scheme: Structural asymmetry along protein sequences and co-translational folding"

<p>Contains data for a set of 16,200 non-redundant protein structures taken from the Protein Data Bank. Associated code can be found at https://github.com/jomimc/FoldAsymCode.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Double-stranded RNA structural elements holding the key to translational regulation in cancer: the case of editing in RNA Binding Motif Protein 8A

<p>Raw data supporting the manuscript</p> <p>Abukar, A.;Wipplinger, M.;<br> Hariharan, A.; Sun, S.; Ronner, M.;<br> Sculco, M.; Okonska, A.;<br> Kresoja-Rakic, J.; Rehrauer, H.; Qi, W.;<br> et al. Double-Stranded RNA<br> Structural Elements Holding the Key<br> to Translational Regulation in Cancer:<br> The Case of Editing in RNA-Binding<br> Motif Protein 8A. Cells 2021, 10, 3543.<br> https://doi.org/10.3390/<br> cells10123543</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Predictions of the SARS-CoV-2 B.1.1.529 Variant Spike Protein Receptor Binding Domain Structure and Neutralizing Antibody Interactions

<p>Using AlphaFold2 and HADDOCK, we have generated a predicted&nbsp;structure for the SARS-CoV-2 B.1.1.529 variant&#39;s Spike receptor binding domain and then predicted the binding interaction with neutralizing antibodies. This was performed to understand the potential structural changes in&nbsp;the receptor binding domain&nbsp;of&nbsp;B.1.1.529 and how this may affect vaccine efficacy through antibody interaction.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

SHAPE datasets from manuscript: Modulation of pre-mRNA structure by hnRNP proteins regulates alternative splicing of MALT1

<p>Normalized SHAPE reactivity for the MALT1 M1 minigene RNA constructs (wildtype, variant 1, and variant 2) reported in the manuscript titled &#39;Modulation of pre-mRNA structure by hnRNP proteins regulates alternative splicing of <em>MALT1&#39;</em> .</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Data for the "Discovery of Dehydroamino Acid Residues in the Capsid and Matrix Structural Proteins of HIV-1"

<p>Bottom-up mass spectrometry-based proteomic analysis (trypsin) was performed on four biological replicates of HIV-1 virions. These virions were isolated from HEK293T cells transfected with a HIV-1 proviral plasmid derived from the pNL4-3 molecular clone, rendered biosafe due to inactivating point mutations in both the env and vpr reading frames. There are 8 total spectra, 4 are from unlabeled aliquots of sample, and 4 are from aliquots of sample treated with glutathione to label dehydroamino acids (Spectra can be accessed on MassIVE&nbsp;(MSV000088220). All data was analyzed using MetaMorpheus version 0.0.319 (https://github.com/smith-chem-wisc/MetaMorpheus). Provided here are the results of this analysis.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Screening routine for integrative dynamic structural biology using SAXS and intramolecular FRET and DEER-EPR on hGBP1 (human guanalyte binding protein 1)

<p>Initial and selected ensemble for major and minor species of the human guanalyte binding protein 1 with scripts for the reading routine to combine and analyse jointly SAXS, EPR and FRET data.</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record