Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

71

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

71 results for “spike protein”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dynamics of SARS-CoV-2 spike protein in open and closed states and identification of key structural perturbations upon mutations

<p>The SARS-Cov-2 spike protein resides on the exterior surface of the coronavirus, and therefore, acts as the first point of contact that mediates cell attachment and fusion. &nbsp;During this process, it undergoes dramatic conformational changes upon host receptor binding. We are leveraging high-performance computing to identify these structural perturbations in wildtype and mutant spike protein models. The files contain structures from molecular dynamics simulations of closed SARS-Cov-2 spike protein embedded in POPC membrane.</p>

opencc-by-4.0May 2020View details →
Figshare44/100

UnityMol COVID19 Spike protein 360 video

<p>A 360 degree video with a camera path through the COVID19 spike protein-ACE2 complex (example 1 of our paper on biorxiv), illustrating the use of specific cameras to export enriched media.</p> <p>NB: not all video players allow you to experience the 360 degree navigability. The youtube link should work fine.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Dataset for: Pre-pandemic artificial MERS analog of polyfunctional SARS-CoV-2 S1/S2 furin cleavage site domain is unique among spike proteins of genus Betacoronavirus

<table> <tbody> <tr> <th>&nbsp;</th> <td> <div> <h3><strong>Data File Descriptions and Methods</strong></h3> <ol> <li><strong>Data file 1 [betacov_matching_IPR042578.fasta]</strong>: Representative set of 2,465 betacoronavirus S protein overlapping homologous superfamily sequences retrieved in fasta format on 4 December 2022 from the InterPro repository at https://www.ebi.ac.uk/interpro/entry/InterPro/IPR042578/.<br><br></li> <li><strong>Data File 2 [betacov_matching_IPR042578_motif.fasta]</strong>: With Data File 1 as input, extracted 98,122 furin cleavage site (FCS) output motifs of 20 amino acids length, including overlapping and redundant sequences, produced with the FindFur algorithm with preset parameters as described by (Gu, 2020). FindFur as used was deposited on 15 December 2020 at the GitHub software repository at https://github.com/chwisteeng/FindFur.<br><br></li> <li><strong>Data File 3 [table_s1s2_hits_betacov_polyf.pdf]</strong>: Compiled summary table of sequence hits (PDF) of spike S1/S2 domains across genus&nbsp;<em>Betacoronavirus. </em>The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences.<em> </em>Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS.<br><br></li> <li> <p><strong>Data File 4 [table_s1s2_hits_betacov_polyf.xlsx]</strong>: Compiled summary table of sequence hits (MS Excel) of spike S1/S2 domains across genus&nbsp;<em>Betacoronavirus. </em>The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences.<em> </em>Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS.<br><br></p> </li> <li> <p><strong>Data File 5 [betacov_s1s2_nls_pat7_furin_psort.txt]:&nbsp;</strong>Nuclear localization signal (NLS) detection output for 5 representative betacoronavirus spike sequence domains, including the positive hits for pat7 in SARS-CoV-2 and for MERS-MA30 CoV. NLS predictions used the PSORT algorithm available as a webservice at https://wolfpsort.hgc.jp/ which is based on the work of Nakai and Horton (Nakai and Horton, 1999). Numbering refers to Data File 3 and Data File 4.<br><br></p> </li> <li> <p><strong>Data File 6 [betacov_s1s2_oglyc_netogly.txt]:&nbsp;</strong>Detection output for 5 representative betacoronavirus spike sequence domains tested for Thr/Ser O-glycosite residue pairs with the standard prediction software NetOGlyc4.0 (Steentoft et al., 2013) as available at https://services.healthtech.dtu.dk/services/NetOGlyc-4.0/. Positive hits have scores above 0.5. Numbering refers to Data File 3 and Data File 4.<br><br></p> </li> <li> <p><strong>Data File 7 [betacov_s1s2_nls_pat7_furin_blastp.txt]</strong>: Comprehensive sequence database searches using were performed using the NCBI protein BLAST (blastp) algorithm with webservice available at https://blast.ncbi.nlm.nih.gov/Blast.cgi?PAGE=Proteins. The following blastp search parameters and settings were used: Word size=2; Expect value=200000; Hitlist size=500; Gapcosts=9,1; Matrix=PAM30; Filter string=F; Genetic Code=1;Window Size=40; Threshold=11; Composition-based stats=0; Database Posted date=Jan 19, 2023 2:59 AM; Number of letters=17,117,563; Number of sequences=10,766; Entrez query: Includes: Betacoronavirus (taxid:694002); Excludes: SARS-CoV-2 (taxid:2697049). The six polyfunctional input query consensus motif sequences were TXXPR(K/H/R)XRSX and TXXPRX(K/H/R)RSX.</p> </li> </ol> <h3><strong>References</strong></h3> <p>Gu, C., 2020. FindFur: A Tool for Predicting Furin Cleavage Sites of Viral Envelope Substrates. Master&rsquo;s Thesis, San Jose State University, CA, USA. doi: <a href="https://doi.org/10.31979/etd.4ahv-9jya">10.31979/etd.4ahv-9jya</a>&nbsp;</p> <p>Gangavarapu K, Latif AA, Mullen JL, Alkuzweny M, Hufbauer E, Tsueng G, Haag E, Zeller M, Aceves CM, Zaiets K, Cano M, Zhou X, Qian Z, Sattler R, Matteson NL, Levy JI, Lee RTC, Freitas L, Maurer-Stroh S; GISAID Core and Curation Team; Suchard MA, Wu C, Su AI, Andersen KG, Hughes LD. Outbreak.info genomic reports: scalable and dynamic surveillance of SARS-CoV-2 variants and mutations. Nat Methods. 2023. 20(4):512-522. doi: <a href="https://doi.org/10.1038/s41592-023-01769-3">10.1038/s41592-023-01769-3</a>.</p> <p>Nakai, K., Horton, P., 1999. PSORT: a program for detecting sorting signals in proteins and predicting their subcellular localization. Trends Biochem Sci 24, 34&ndash;36. doi: <a href="https://doi.org/10.1016/s0968-0004(98)01336-x">10.1016/s0968-0004(98)01336-x</a></p> <p>Steentoft, C., Vakhrushev, S.Y., Joshi, H.J., Kong, Y., Vester-Christensen, M.B., Schjoldager, K.T.-B.G., Lavrsen, K., Dabelsteen, S., Pedersen, N.B., Marcos-Silva, L., Gupta, R., Bennett, E.P., Mandel, U., Brunak, S., Wandall, H.H., Levery, S.B., Clausen, H., 2013. Precision mapping of the human O-GalNAc glycoproteome through SimpleCell technology. EMBO J 32, 1478&ndash;1488.&nbsp;doi: <a href="https://doi.org/10.1038/emboj.2013.79">10.1038/emboj.2013.79</a></p> </div> </td> </tr> </tbody> </table>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Analysis of the interacting residues between wild type SARS-CoV-2 spike protein and natural ligand hACE2, as well as three engineered alternative ligands

<p>The analysis of residue interactions between the SARS-CoV-2 spike protein and its natural (hACE2 <sup>1</sup>) and engineered binders P17 Fab <sup>2</sup>, Ty1 VHH <sup>3</sup> and LCB1 peptide <sup>4</sup> reveals that glutamine, serine and especially tyrosine residues on the ligand side are more frequent and influence spike binding efficiency, and that spike residues Glu484, Phe486, Tyr489 and Gln493 are more recurrent targets for interactions with ligands. The list of residues establishing contacts between the wild type structure of the SARS-CoV-2 spike protein and the binders defined above are described in Table 1. In Figure 1, the frequency and type of amino acids that interact with each spike residue is illustrated.</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Identifying and profiling structural similarities between Spike of SARS-CoV-2 and other viral or host proteins with Machaon - Pre-computed features for replication

<p>Machaon&#39;s computed features that were used in the structural comparisons with Spike protein.</p> <p>DATA_PDBS_vir_whole_1-3.zip files are parts of a single folder.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Molecular Dynamics Simulation of SARS-CoV-2 Spike Protein

<p>Trajectory data corresponding to the manuscript, tentatively titled &quot;Distant Residues Modulate the Conformational Opening in SARS-CoV-2 Spike Protein&quot;</p> <p>Authors: Dhiman Ray, Ly Le, Ioan Andricioaei</p> <p>Affiliation: University of California Irvine, USA</p> <p>Description: Multiple unbiased simulations of 40 ns were performed for the SARS-CoV-2 spike protein. Frames are saved at 50 ps interval. The initial structures were generated from umbrella sampling simulation starting from PDB ID: 6VSB and 6VXX. The index at the end of filename stands for the umbrella sampling window from which the trajectory was initiated. The indices are not continuous as not all the umbrella sampling windows were used to start trajectories. Additionally 3 trajectories, each of length 80 ns, are included for the closed, partially open and fully open state. The topology is provided as a PDB file (&quot;spike_dry.pdb&quot;).</p> <p>The trajectories are for the spike head only structure obtained from the CHARMM-GUI Covid-19 archive. No solvent or ions are included in the trajectory or the topology.</p> <p>Update: Additional trajectories and PDB files for D614G mutant added. Each trajectory is 40 ns long. The PDB files are named 6VXX_mutant_dry.pdb and 6VSB_mutant_dry.pdb for the closed and partially open state.</p> <p>Pre-print available: https://doi.org/10.1101/2020.12.07.415596</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

MD simulations of SARS-CoV-2 Spike Protein under static electric fields

<p>This dataset contains trajectories corresponding to all-atom MD simulations of segments of the SARS-CoV-2 Spike Protein, and in-silico mutations, under the influence of moderate external electric fields. The final structures of some of the simulations were used to perform in-silico docking with ACE2 receptor to evaluate the effect of comformational changes (docking was perform with PyDOCK).</p> <p>The file trajectories_6vsb_dt1ns.zip contains trajectories of simulations that were performed on a segment of the Protein Data Bank ID 6VSB comprising RBD, SD1 and SD2. The file trajectories_6m0j_dt1ns.zip correspond to the RBD in Protein Data Bank ID 6M0J. The file trajectories_in-silico_mutations_dt1ns.zip correspond to simulations performed on in-silico generated mutations following the mutations corresponding to WHO Variants of Concern UK, South Africa and Brazil. In all cases, simulations were performed at different electric field intensities ranging between 10<sup>4</sup> V/m and 10<sup>7</sup> V/m, with an extra short simulation under very high intensity (10<sup>9</sup> V/m). The file docked_structures_6m0j.zip contains the 100 best scored docked structures for each case as the output of PyDOCK.</p> <p>Trajectories are stored in GROMACS compressed trajectory file format (.xtc), downsampled to a 1ns timestep. Individual trajectories length are between 300 nanoseconds and 1 microsecond. In-silico docked structures are in PDB format. See linked preprint for more details.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

All atom simulations snapshots and contact maps analysis scripts for SARS-CoV-2002 and SARS-CoV-2 spike proteins with and without ACE2 enzyme

<p><strong>The dataset contains a total of 40&nbsp;snapshots of the four&nbsp;trajectories (10&nbsp;snapshots each&nbsp;system =&nbsp;two per replica&nbsp;x 5 replicas/system):</strong></p> <ol> <li>SARS-CoV-2002 spike protein without ACE2</li> <li>SARS-CoV-2&nbsp;spike protein&nbsp;without ACE2</li> <li>SARS-CoV-2002 spike protein with&nbsp;ACE2</li> <li>SARS-CoV-2&nbsp;spike protein with&nbsp;ACE2</li> </ol> <p>Molecular dynamics simulation trajectories (320ns each) have been performed using the Amber&nbsp;ff14SB&nbsp;force field running with the Amber18 package at the&nbsp;the&nbsp;NSF-funded (OAC-1826915, OAC-1828163) ELSA high performance computing cluster at The College of New Jersey.&nbsp;Under the following simulation methodology:</p> <p><em>All-atom simulations were carried out with Amber18 (<a href="https://slack-redir.net/link?url=http%3A%2F%2Fambermd.org">ambermd.org</a>), and system components (protein, ions, water) were modeled with the included FF14SB and TIP3P parameter sets. Energy minimization used CPU pmemd, while later simulation stages used GPU pmemd. CoV2 and CoV1 systems with one RBD up (with/without ACE2) were solvated in 12 angstrom water shells. Cysteine residues identified in the initial models as having a disulfide bond (DB) were bonded using tLeap. All simulations used 0.150 M NaCl. Hydrogen mass repartitioning was applied only to the protein to enable a 4 fs timestep (<a href="https://slack-redir.net/link?url=https%3A%2F%2Fpubs.acs.org%2Fdoi%2Fabs%2F10.1021%2Fct5010406">https://pubs.acs.org/doi/abs/10.1021/ct5010406</a>). The SHAKE algorithm was applied to hydrogens, and a real-space cutoff of 8 angstroms was used. Periodic boundary conditions were applied and PME was used for long-range electrostatics. Minimization was by steepest descent (2000 steps) followed by conjugate gradient (3000 steps). Heating used two stages: (1) NVT heating from 0 K to 100 K (50 ps), and (2) NPT heating from 100 K to 300 K (100 ps). Restraints of 10 kcal mol<sup>-1</sup>&nbsp;angstrom<sup>-2</sup>&nbsp;were applied during minimization and heating to C-alpha atoms. During 6 ns of equilibration at 300 K C-alpha restraints were gradually reduced from 10 kcal mol<sup>-1</sup>&nbsp;angstrom<sup>-2</sup>&nbsp;to 0.1 kcal mol<sup>-1</sup>&nbsp;angstrom<sup>-2</sup>. Finally, restraints were released and 320 ns unrestrained production simulations were carried out for CoV2 and CoV1 systems. Production simulations began from the final equilibrated snapshots, and five copies of each system were simulated. As unrestrained systems can freely rotate we monitored simulations for any close contacts and found that in one copy of the CoV1 simulation without ACE2 and one RBD up that a few contacts close to 8 angstrom occur near the end of the 320 ns between the RBD and a different subdomain of the spike complex in a periodic image. However this did not influence analyzed structural properties which is verified by comparing results across simulations. The Monte Carlo barostat was used to maintain pressure (1 atm), and the Langevin thermostat was used to maintain 300 K temperature (collision frequency 1 ps<sup>-1</sup>), as implemented in Amber18. In aggregate, nearly 7 microseconds of simulation of systems ranging from 396,147 to 879,100 atoms was carried out for this work.</em><br> For further details on the trajectories, please contact&nbsp;Joseph Baker (bakerj@tcnj.edu).</p> <p><strong>Regarding the contact map analysis scripts&nbsp;(contactMaps_Analysis.tar.gz), they contain the following workflow:</strong></p> <p>contactmap &nbsp; &nbsp; &nbsp;--&gt; source files from contact_map executable<br> process_nc.sh &nbsp; --&gt; convert raw data from all-atom simulation to numbered PDB files and get the contact maps<br> frequency.lua &nbsp; --&gt; read a set of PDB files and output the frequency count for each contact<br> consensus.fasta --&gt; align sequence of Covid19 and SARS from Chimera<br> consensus.lua &nbsp; --&gt; read data previously generated and compute the frequency per residue, among other things.<br> consensus.sh &nbsp; &nbsp;--&gt; input information to consensus.lua<br> consensus.gp &nbsp; &nbsp;--&gt; gnuplot script to plot figures</p> <p>This dataset and the code is part of tripartite collaboration between:</p> <ul> <li>The Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland (supported by the National Science Centre, Poland, under grant No. 2017/26/D/NZ1/0046)</li> <li>Department of Chemistry, The College of New Jersey, New Jersey, United States (supported by National Science Foundation under grant numbers OAC-1826915 and OAC-1828163).</li> <li>Jozef Stefan Institute, Ljubljana, Slovenia (supported by the Slovenian Research Agency (Funding No. P1-0055)).</li> </ul>

opencc-by-4.0May 2020View details →
zenodo40/100

Data for the article: "Molecular Modelling Reveals Eight Novel Druggable Binding Sites in SARS-CoV-2's Spike Protein" by Ilke Ugur and Antoine Marion

<p>This upload contains data related to the article<br> published as a preprint on ChemRxiv with DOI<br> https://doi.org/10.26434/chemrxiv.13292768</p> <p>&quot;Molecular Modelling Reveals Eight Novel Druggable Binding Sites in SARS-CoV-2&#39;s Spike Protein&quot;<br> by Ilke Ugur and Antoine Marion (2020)<br> Department of Chemistry, Middle East Technical University, Ankara, Turkey.</p> <p>For further information, please contact:<br> ilkeugur@metu.edu.tr ; amarion@metu.edu.tr</p> <p>The manuscript is currently under peer-review.</p> <p>Content:</p> <p>Library of molecules derived from DrugBank v 5.1.5:<br> - DrugBank_2020_5.1.5/ &nbsp; &nbsp; &nbsp; &nbsp;# All necessary files for the docking and refinement of the library of molecules.<br> -- DB_5.1.5_pH7.4_pdbqt/ &nbsp; &nbsp; &nbsp;## PDBQT readily usable for docking with AutoDock Vina.<br> -- DB_5.1.5_pH7.4_mol2amber/ &nbsp;## mol2 files containing assigned GAFF atom types and Gasteiger atomic charges.<br> -- DB_5.1.5_pH7.4_frcmod/ &nbsp; &nbsp; ## frcmod files containing missing molecular mechanics parameters<br> -- dbID_name.dat &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;## DrugBank ID to generic name dictionary</p> <p>Note: The files were prepared automatically via a series of operations handling openbabel and antechamber.<br> &nbsp; &nbsp; &nbsp; The protonation state of ionizable groups as well as Gasteiger atomic charges were assigned by openbabel for a pH of 7.4<br> &nbsp; &nbsp; &nbsp; mol2 and frcmod files can be used readily via the tleap module of AmberTools to produce topology files.</p> <p><br> Receptor structures:<br> - receptors/ &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# PDB files for the four structures of the spike protein considered in this work<br> -- CS00ns.pdb &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ## Closed state after the remodelling of missing loops (PDB ID 6vxx)<br> -- OS00ns.pdb &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ## Open &nbsp; state after the remodelling of missing loops (PDB ID 6vyb)<br> -- CS25ns.pdb &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ## Closed state after 25 ns of molecular dynamics in explicit water<br> -- OS25ns.pdb &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ## Open &nbsp; state after 25 ns of molecular dynamics in explicit water</p> <p>Note: All structures are aligned to CS00ns.pdb and can be converted to pdbqt for docking with AutoDock Vina</p> <p><br> Docking grid centers:<br> - dockingCenters/ &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # XYZ files containing the coordinates of each docking grid center considered in this work</p> <p>Note: The coordinates are given in the same frame as that of the four structures of the receptor.</p> <p><br> Binding sites:<br> - bindingSites/ &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # XYZ files with the coordinates of the representative atomic centres<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # of each binding site identified in this work (A-H).</p> <p>Note: These files can be used to get a clearer picture of the binding sites within the structures<br> &nbsp; &nbsp; &nbsp; of the spike protein shared in the receptors directory.</p> <p><br> Final modelling results:<br> - allData.txt &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # data for all molecules in the set (approved and investigational)<br> - appData.txt &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # data for approved molecules only<br> - data.xlsx &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # data for all molecules in the set (approved and investigational)<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # as a formatted excel spreadsheet</p> <p>Note: The columns are delimited with semi-colons &quot;;&quot;.<br> &nbsp; &nbsp; &nbsp; The files contain the results for the best pose of all approved molecules for which<br> &nbsp; &nbsp; &nbsp; molecular mechanics-based geometry optimization succeeded, regardless of their score.<br> &nbsp; &nbsp; &nbsp; For other molecules, the result of their best pose is reported only for those complexes<br> &nbsp; &nbsp; &nbsp; having MM interaction energy lower or equal to -22.00 kcal/mol.</p> <p><br> Visualization:<br> - bs.pse &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# pymol session representing the binding sites within the<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # closed state structure of the spike protein (CS00ns)<br> - pt.pse &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;# pymol session representing the docking grid centres within<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # closed statestructure of the spike protein (CS00ns)</p> <p>Note: the PSE files should be compatible with version 7.0 of pymol and later</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Predictions of the SARS-CoV-2 B.1.1.529 Variant Spike Protein Receptor Binding Domain Structure and Neutralizing Antibody Interactions

<p>Using AlphaFold2 and HADDOCK, we have generated a predicted&nbsp;structure for the SARS-CoV-2 B.1.1.529 variant&#39;s Spike receptor binding domain and then predicted the binding interaction with neutralizing antibodies. This was performed to understand the potential structural changes in&nbsp;the receptor binding domain&nbsp;of&nbsp;B.1.1.529 and how this may affect vaccine efficacy through antibody interaction.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Simulation of Receptor Binding Domain of SARS-CoV-2 spike protein (WT and variants) in complex with neutralizing antibodies.

<p>This repository contains the molecular dynamics trajectories of the SARS-CoV-2 Spike RBD bound to BD23 and B38 monoclonal antibodies. The simulations for the RBD only systems are also provided. The trajectories are available for the WT spike protein as well as for four different variants (alpha, beta, kappa and delta). The simulations of the RBD only system are propagated for 300 ns and for the RBD-Antibody complex for 500 ns. The trajectories are saved at 100 ps interval. The Steered MD simulation trajectories&nbsp;(WT_RBD_B38_SMD_1.dcd etc.) and collective variables files are also included (WT_RBD_B38_SMD_1.colvars.traj etc.). There are 5 SMD trajectories for each RBD antibody pair. The details of the simulation can be obtained from the preprint: https://doi.org/10.1101/2021.08.13.456317</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Coronavirus Spike Protein Sequences from NCBI Virus

<p>For the study of taxonomic classification of coronaviruses across all genera (Alpha-, Beta-, Gamma-, and Deltacoronavirus), we can use spike protein sequences downloaded from the NCBI Virus website, \url{https://www.ncbi.nlm.nih.gov/labs/virus/vssi/}. The spike protein sequences used in this study were downloaded on November 21 and 27, 2021 using search terms such as ``spike&#39;&#39;, ``S1 protein&#39;&#39;, ``S2 protein&#39;&#39;, and ``S protein&#39;&#39;, in order to download as many spike protein sequences as possible.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Synthetic single particle cryo-EM dataset of the SARS-CoV-2 spike protein

<p>PDBs were generated using molecular dynamics.<br> See DESRES_README.txt for more details on molecular dynamics simulation.<br> PDBs were converted to volumetric data using EMAN2.<br> The image stack contains 100 000 projection images each&nbsp;<br> of the 10 states (see PDBs), at an SNR of 1/10 in the following order:</p> <p>state00 (closed)<br> state01 (closed)<br> state02 (closed)<br> state10 (intermediate)<br> state11 (intermediate)<br> state12 (intermediate)<br> state13 (intermediate)<br> state20 (open)<br> state21 (open)<br> state22 (open)</p> <p>Projections were made using relion_project.&nbsp;<br> &nbsp;&nbsp;White gaussian noise with standard deviation 1.0<br> &nbsp;&nbsp;CTF multiplied signal<br> &nbsp;&nbsp;High signal-to-noise ratio<br> &nbsp;&nbsp;Image size 96x96x96<br> &nbsp;&nbsp;<br> MRC-files used for the projections not included, but can be generated using the PDB files.<br> Final RELION reconstruction resolution is 5.33334 Angstrom (Nyqvist is at 5.33334).</p> <p>Command line for RELION reconstruction:<br> relion_refine_mpi --o refine3d/run --auto_refine --split_random_halves --i rot_trans_ctf_noise/stack.star --ref pdb2mrc/state21.mrc --ini_high 20 --dont_combine_weights_via_disc --preread_images --pool 30 --pad 2 --ctf --particle_diameter 130 --flatten_solvent --zero_mask --oversampling 1 --healpix_order 2 --auto_local_healpix_order 4 --offset_range 5 --offset_step 2 --low_resol_join_halves 40 --norm --scale --j 2 --gpu --fristiter_cc --grad&nbsp;</p> <p>This dataset is generated as a testbed for cryo-EM heterogeneity analysis.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Analysis of hMSC-derived chondrocytes in response to stimulation with SARS-CoV-2 spike proteins

<p>Previous studies have detected the presence of SARS-CoV-2 viral proteins in bronchial cartilage chondrocytes. We hypothesised that the leakage of viral proteins to local joint tissue was due to virus-induced endothelial dysfunction. Studies have also shown upregulation of endothelin-1 (ET-1), the most potent vasoconstrictor, in COVID patients. We are investigating the direct effect of SARS-CoV-2 spike protein (SP) and the host response to chondrocytes.</p> <p>Human mesenchymal stem cells (hMSCs)-differentiated chondrocytes were treated with either full-length SARS-CoV-2 spike protein (SP) only or a combination of spike protein, neutralising antibody to S1 and endothelin-1 (SAE) to mimic viral insult and host response respectively. RNA sequencing was performed to compared the change in transcriptome in control, SP, and SAE. All samples were processed in the same batch. Default quality control parameters were used.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Production and purification of receptor-binding domain (RBD) of the Spike protein from a transiently transfected mammalian cell

<p>&nbsp;</p> <p>Production and purification of recombinant receptor-binding domain (RBD) of the Spike protein from a transiently transfected EXPI293&nbsp;mammalian cell .</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

Production and purification protocols for Ectodomain of SARS-CoV-2 Spike (S) protein

<p>Production and purification protocols for Ectodomain of SARS-CoV-2 Spike (S) protein &nbsp; &nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

Snapshots, frequency contact maps analysis, Poisson Boltzmann calculations, and data scripts for characterization of structural and energetic differences between conformations of the SARS-CoV-2 spike protein

<p><strong>Molecular dynamics simulation</strong> trajectories, which have been performed using the Amber&nbsp;ff14SB&nbsp;force field running with the Amber18 package at the NSF-funded (OAC-1826915, OAC-1828163) ELSA high performance computing cluster at The College of New Jersey. Simulation methodology and further details are described in [1] and [2]. For further details on the trajectories, please contact&nbsp;Joseph Baker (bakerj@tcnj.edu).</p> <p>The <strong>Poisson Boltzmann </strong>energy calculations have been achieved by using the input_files.tar.xz found here and solving the Poisson Boltzmann equation with pygbe. A more detailed example and tutorial can be found at [4]. For further details contact Horacio V Guzman.</p> <p><strong>The dataset contains </strong></p> <ul> <li><strong>A total of 30&nbsp;snapshots of the three trajectories (10&nbsp;snapshots each&nbsp;system =&nbsp;two per replica&nbsp;x 5 replicas/system):</strong></li> </ul> <ol> <li>SARS-CoV-2002 spike protein with three RBD in the down positions: &quot;COV2-DDD/PDB/&quot; .</li> <li>SARS-CoV-2002 spike protein with one RBD in the up and two RBD in the down positions: &quot;COV2-UDD/PDB/&quot;.</li> <li>SARS-CoV-2002&nbsp;spike protein with two RBD in the up and one RBD in the down positions: &quot;COV2-DUU/PDB/&quot;.</li> </ol> <ul> <li><strong>Input files for Poisson-Boltzmann analysis</strong>:</li> </ul> <ol> <li>PoissonBoltzmann/input_files.tar.xz</li> </ol> <ul> <li><strong>Data for the frequency contact map and processing scripts</strong>:</li> </ul> <ol> <li>cov2-ddd.pdb, cov2-udd.pdb, cov2-duu.pdb reference PDB files.</li> <li>Contact maps [3] at&nbsp; &quot;COV2-DDD/CONTACT_MAP/&quot;,&nbsp; &quot;COV2-UDD/CONTACT_MAP/&quot;,&nbsp; &quot;COV2-DUU/CONTACT_MAP/&quot;.</li> <li>frequency.lua: get frequency of contacts from a set of contacts map files.</li> <li>diff_frequency.lua: get differential frequency of contacts from a set of frequency files.</li> <li>Frequency of contacts listed in frequency.data files at &quot;COV2-DDD/&quot;, &quot;COV2-UDD/&quot; and &quot;COV2-DUU/&quot; directories.</li> </ol> <p>Read the &quot;INFO&quot; files for further informations.</p> <p>This dataset and the code is part of a collaboration between:</p> <ul> <li>The Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland (supported by the National Science Centre, Poland, under grant No. 2017/26/D/NZ1/0046)</li> <li>Department of Chemistry, The College of New Jersey, New Jersey, United States (supported by National Science Foundation under grant numbers OAC-1826915 and OAC-1828163).</li> <li>Jozef Stefan Institute, Ljubljana, Slovenia (supported by the Slovenian Research Agency (Funding No. P1-0055)).</li> <li>School of engineering in bioinformatics, University of Talca, Talca, Chile.</li> </ul> <p>[1] Rodrigo A. Moreira, Mateusz Chwastyk, Joseph L. Baker, Horacio V Guzman, &amp; Adolfo B. Poma. (2020). All-atom simulations snapshots and contact maps analysis scripts for SARS-CoV-2002 and SARS-CoV-2 spike proteins with and without ACE2 enzyme (Version 0.1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3817447</p> <p>[2] Chad W. Hopkins, Scott Le Grand, Ross C. Walker, and Adrian E. Roitberg. Long-Time-Step Molecular Dynamics through Hydrogen Mass Repartitioning. Journal of Chemical Theory and Computation 2015 11 (4), 1864-1874. http://doi.org/10.1021/ct5010406</p> <p>[3] Rodrigo A. Moreira, Mateusz Chwastyk, Joseph L. Baker, Horacio V Guzman, &amp; Adolfo B. Poma. Quantitative determination of mechanical stability in the novel coronavirus spike protein. Nanoscale, 2020,12, 16409-16413. <a href="https://doi.org/10.1039/D0NR03969A">https://doi.org/10.1039/D0NR03969A</a></p> <p>[4] https://github.com/pyF4all</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

Docking data for "The evolution of the SARS-CoV-2 spike protein for differential usage of the host transmembrane serine proteases entry pathway"

<p><br>The dataset includes predicted complexes of the SARS-CoV-2 Spike protein (specifically at the S2' cleavage site) with Hepsin and TMPRSS2 proteins. It contains data on three variants: Wuhan, Delta, and Omicron BA.1.</p> <p><strong>Compressed folders:</strong></p> <p>-357596-DeltaHepsin.tgz</p> <p>-357597-DeltaTMPRSS2.tgz</p> <p>-360039-WuhanHepsin.tgz</p> <p>-360042-TMPRSSWuhan.tgz</p> <p>-392981-TMPRSS-BA1_all.tgz</p> <p>-392982-Hepsin-BA-all.tgz</p> <p><strong>Each compressed folder contains the following:</strong></p> <p>-Initial structures in pdb format</p> <p>-Output complexes in pdb format</p> <p>-Clusters in pdb format</p> <p>-Protocols</p> <p>-Parameters</p> <p>-Scoring files</p> <p>&nbsp;</p> <p><strong>Protein-protein docking&nbsp;</strong><br>Molecular docking between the SARS-CoV-2 S protein of Wuhan, Delta (PDB: 7W92, [DOI: 10.1038/s41467-022-28528-w]), and BA.1 (PDB: 7XO5, [DOI: 10.1038/s41422-022-00672-4]) and the human proteases TMPRSS2 (PDB: 8HD8, [DOI: 10.1038/s41467-023-42527-5]) and Hepsin (PDB: 1Z8G, [DOI: 10.1042/BJ20041955]) was performed using the HADDOCK v2.5-2024.03 webserver ([DOI: 10.1021/ja026939x], [DOI: 10.1016/j.jmb.2015.09.014]). Missing loops in the protein structures were reconstructed using Modeller v10.5 ([DOI: 10.1006/jmbi.1993.1626]). Every heteroatom was removed from the reference structures. The relaxed atomistic coordinates for each S protein variant were derived via all-atom molecular dynamics (MD) simulations. These simulations were performed using AMBER22 with the FF19SB force fields and the pmemd.cuda module for enhanced performance ([DOI: 10.1021/acs.jcim.3c01153], [DOI: 10.1021/jz501780a], [DOI:10.1021/ct400314y]). For the Wuhan variant the S protein was retrieved from our previous modeling study [DOI: 10.1039/D0NR03969A] where for Delta and BA.1, ecah S protein was placed in a dodecahedral box, extending 20 &Aring; beyond the solute in every cartesian direction, and solvated with the four-site OPC water model ([DOI: 10.1021/jz501780a]). The systems were neutralized with counterions, specifically one Cl&minus; ion for the Delta variant and three Cl- ions for the BA.1 variant. To remove local clashes, a geometric optimization was performed using the steepest descent algorithm for 5000 cycles. The MD equilibration process consisted of several stages. First, temperature equilibration in the NVT ensemble was performed by gradually increasing the temperature through steps of 150, 200, 250, 300, and finally 310 K, each lasting 200 ps. During this phase, position restraints were applied to the heavy atoms of the proteins, with progressively decreasing spring constants of 5.0, 4.0, 3.0, and 1.0 kcal mol&minus;1 &Aring;&minus;2, facilitating gradual relaxation. This was followed by a 1 ns equilibration at 310 K in the NPT ensemble without restraints. For production MD, the simulations were run in the NPT ensemble with periodic boundary conditions and Particle Mesh Ewald (PME) method ([DOI: 10.1063/5.0040966], [DOI: 10.1021/ct9001015]) using a grid spacing of 1.0 &Aring; for long-range electrostatics. Non-bonded interactions were modeled with a Lennard-Jones potential using a 9&Aring; cutoff. Temperature control was maintained using Langevin dynamics ([DOI: 10.1021/ct800573m]) with a collision frequency of 4.0 ps&minus;1, and pressure control was managed by the Monte Carlo barostat ([DOI: 10.1016/j.cplett.2003.12.039]) with a 2.0 ps relaxation time at 1 bar. Bond constraints on hydrogen atoms were applied using the SHAKE algorithm ([DOI: 10.1016/0021-9991(77)90098-5]), and the hydrogen mass repartitioning scheme was applied via ParmEd ([DOI: 10.1371/journal.pcbi.1005659]), enabling a 4 fs integration time step ([DOI: 10.1021/ct5010406]). Each protein complex was simulated for a total of 20 ns. For the Wuhan variant, the 3D coordinates were retrieved from [DOI: 10.5281/zenodo.3817446].<br>The active interaction region on the spike protein was defined as the cleavage site (residues P809-R815). For TMPRSS2 and Hepsin, the active sites were defined based on their catalytic residues: H296, D345, D435, S441, S460, and G462 for TMPRSS2, and H203, D257, D347, A348, and S353 for Hepsin. These specific regions were selected to guide the docking process and maximize biologically relevant interactions. Docking clusters were analyzed by selecting those with the lowest interaction energies for further structural analysis. To evaluate binding accuracy, native contacts between the S protein and proteases were computed using the contact map analysis based on the OV+rCSU method ([DOI: 10.12693/APhysPolA.145.S9, 10.1021/acs.jctc.6b00986]), which allows for a precise identification of critical stabilizing interactions, both specific and non-specifics. High-frequency contacts, defined as those appearing in over 70% of the generated models, were highlighted as key determinants of protein-protein recognition, providing insight into the most stable and consistent interactions across docking configurations.</p>

opencc-by-4.0Nov 2024View details →
dryad36/100

The use of nanobodies in a sensitive ELISA test for SARS-CoV-2 Spike 1 protein

<p>A rapid detection method for SARS-CoV-2 spike protein is essential for control of COVID19. We investigated various combinations of engineered nanobodies in a sandwich ELISA to detect the Spike protein of SARS-CoV-2. We have identified an optimal combination of nanobodies. These were selectively functionalised to further improve antigen capture. This dataset contains data from ELISA experiments described in the manuscript.<span>                                                                                                                                           </span></p> <p><span>Plate coating of nanobodies for ELISA by passive adsorption vs biotinylation was compared. A series of nanobody pairings (two cluster 2 ACE2-binding epitope and two cluster 1 CR3022 epitope) were screened for optimum sensitivity. The optimal pair were then tested against a series of SARS-COV-2 antigens: recombinant spike 1 protein; recombinant receceptor binding domain (RBD); pseudotyped HIV-1 and heat-empigen inactivated SARS-CoV-2 virus. X-ray irradiated SARS-CoV-2 was also tested. Sensitivity to these antigens was compared with nanobodies biotinylated a) site-selectively and b) in a non-specific stochastic manner. Batch-to-batch viral variation and effects of inactivating agents were investigated. Limit of detection was compared against delta and beta viral mutants. Combining optimal nanobody pairing and site-selective biotinylation, we observed a limit of detection of 147 pg/mL for Spike protein; 33 pg/mL for RBD; 16 TCID50/mL of pseudovirus and 15 ffu/mL of heat-Empigen inactivated SARS-CoV-2. The pairing also showed sensitivity towards delta variant. We have demonstrated the use and sensitivity of nanobodies in ELISA by detection of recombinant and viral SARS-CoV-2 antigens.</span></p>

opencc-zeroJun 2021View details →
zenodo36/100

SARS-CoV-2 spike protein complexes and their contacts

<p>This archive contains protein complex structures containing SARS-CoV-2 spike protein as well as derived data.</p>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record