Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,489
datasets available to search
ShareScore release 0.7.1
Dataset results
2,489 results for “SARS-CoV-2”
Tagged Twitter timelines for users reporting SARS-CoV-2 infections and related data
<p>Twitter data was collected through the Twitter API v2.0, specifically through the timeline endpoint. Details about the inference of SARS-CoV-2 self-reports and the tagging of the full timeline of each user can be found in the <a href="https://github.com/digitalepidemiologylab/content_changes_paper">GitHub repository</a>. The larger dataset ("preprocessed_data.csv") consists of a total of 8,534,171 tweets posted by 30,856 users from January 1, 2020 to October to September 30, 2021.</p> <p>The raw data from Twitter, including tweet and user IDs, has been removed or anonymized in order to comply with the EPFL guidelines for data sharing.</p> <p>In particular, the date of the tweets was removed, the text of the tweets, URLs and URL domains have been substituted with the "text", "<URL>" and "<URL_DOMAIN>" token respectively.</p> <p>User IDs have been substitued with new IDs in the [0, number of users] range (e.g. U0, U1, ...) .</p> <p>Tweet IDs have been substitued with new IDs in the [0, number of tweets] range (e.g. T0, T1, ...) .</p> <p>In addition to self-explanatory columns about topics, emotions, URL classification and symptoms we tagged, we also share the columns:</p> <ul> <li>pdate: date of the SARS-CoV-2 infection self-report for that user (adjusted with SUTime)</li> <li>effective_date: date of the tweet adjusted with SUTime, when the SUTime columns is available.</li> <li>rel_effective_day(week, month): days (weeks, months) computed with respect to the positivity date (i.e. "pdate" column). Negative numbers refer to tweets posted before the user reported a COVID-19 infection on Twitter.</li> </ul>
Dataset of "Exposure to airborne SARS-CoV-2 in four hospital wards and ICUs of Cyprus. A detailed study accounting for day-to-day operations and aerosol generating procedures."
<p>The authors highly appreciate being contacted if the data is to be used for any purpose.</p> <p>The following data set was used in the study entitled "<strong>Exposure to airborne </strong><strong>SARS-CoV-2 in four hospital wards and ICUs of Cyprus. A detailed study accounting for day-to-day operations and aerosol generating procedures.</strong>" and published in <em>Heliyon</em> Journal.</p> <p>This study characterized the transmission dynamics of airborne SARS-CoV-2 in normal and intensive care units. The data were collected over the period of 2020. In total, 165 and 62 air and environmental samples, respectively, were collected in four COVID-19 wards and ICUs in Cyprus and analyzed by RT-PCR. The comparison between RT-PCR and an alternative method for SARS-CoV-2 detection in air that provides comparable results but is less cumbersome and time demanding, is also given in the tab "Comparison with BELD".</p> <p>The data from sampling airborne SARS-CoV-2 using a MOUDI impactor are not included in this document but can be found in the supplement of the relevant publication.</p> <p>Please refer to the manuscript and its supplementary material for more information about how the data was collected. </p> <p> </p>
A evolução do vírus SARS-COV-2 no Nordeste brasileiro.
<p>Os dados estão distribuídos da seguinte forma:</p> <p><em><strong>Região</strong></em> – Nome da região, no caso Nordeste;</p> <p><em><strong>Estado</strong></em> – Nome da UF;</p> <p><em><strong>Município</strong></em> – Nome do município;</p> <p><em><strong>Coduf</strong></em> – Código da UF;</p> <p><em><strong>Codmun</strong></em> – Código do município;</p> <p><em><strong>CodRegiaoSaude</strong></em> - Código da Região de Saúde;</p> <p><em><strong>NomeRegiaoSaude</strong></em> – Nome da Respectiva região de saúde;</p> <p><em><strong>Data</strong></em> – Data da coleta dos dados;</p> <p><em><strong>SemanaEpi</strong></em> – Número de semanas epidemológica;</p> <p><em><strong>PopulacaoTCU2019</strong></em> – População do Município;</p> <p><em><strong>CasosAcumulado</strong></em> – Casos acumulados no Município;</p> <p><em><strong>CasosNovos</strong></em> – Novos casos no Município;</p> <p><em><strong>ÓbitosAcumulado</strong></em> – Total de óbitos acumulados no município;</p> <p><em><strong>ÓbitosNovos</strong></em> - Total de novos óbitos no município;</p>
Glycosylated models for: The diversity of the glycan shield of sarbecoviruses closely related to SARS-CoV-2
<p>Glycosylated models (as PDB files) of the sarbecovirus spike proteins used in the study: The diversity of the glycan shield of sarbecoviruses closely related to SARS-CoV-2.</p>
Supporting data for "CoVEffect: Interactive System for Mining the Effects of SARS-CoV-2 Mutations and Variants Based on Deep Learning"
<p>This repository contains the datasets created and extracted for the paper:</p> <p>Giuseppe Serna García, Ruba Al Khalaf, Francesco Invernici, Stefano Ceri, and Anna Bernasconi. 2022.<br> "<strong>CoVEffect</strong>: Interactive System for Mining the <strong>Effects of SARS-CoV-2 Mutations and Variants</strong> Based on Deep Learning". (Available online at http://gmql.eu/coveffect)</p> <p>--------------------------------------------------------------------------------<br> LIST OF FILES WITH DESCRIPTION:<br> --------------------------------------------------------------------------------</p> <p>AdditionalFile1-effects-taxonomy:<br> Descriptions of legal values for the 'Effect' field, based on a categorized taxonomy.</p> <p>AdditionalFile2-levels-taxonomy:<br> Descriptions of legal values for the 'Level' field.</p> <p>AdditionalFile3-training_dataset_target:<br> List of target tuples (manually annotated) of 221 abstracts considered for training the model. For each abstract, target tuples follow the schema ID, DOI, title, entity, effect, level, type (mutation or variant), tuples_count (>1 when an effect/level is shared by multiple entities, #abstracts containing the same effect described in the tuple).</p> <p>AdditionalFile4-validation_dataset_target:<br> List of target tuples (manually annotated) of 50 abstracts considered for validating the prepared prediction model.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3.</p> <p>AdditionalFile5-validation_dataset_highlighted:<br> Textual abstracts of the 50 manuscripts considered for validation; the text used to support the manual target annotations has been highlighted in yellow.</p> <p>AdditionalFile6-validation_dataset_prediction:<br> List of predicted annotations of 50 abstracts considered for validating the prepared prediction model. The file is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).</p> <p>AdditionalFile7-keywords_query_list:<br> Keyword-based search run on the CORD-19 dataset to extract a relevant subset of abstracts regarding the scope of interest of CoVEffect. The Boolean logic used to combine keywords is explained in the section 'Annotations of the biology-related CORD-19 cluster'.</p> <p>AdditionalFile8-CORD-19_batch_dataset_metadata:<br> Metadata of the 7,230 papers extracted by the keyword-based query in AdditionalFile7.<br> These abstracts have been annotated by the prediction framework.</p> <p>AdditionalFile9-CORD-19_batch_dataset_prediction:<br> List of predicted annotations of 7,230 abstracts extracted from the biology-related cluster of CORD-19.</p> <p>AdditionalFile10-test_dataset_target:<br> List of target tuples (manually annotated) of 100 abstracts randomly selected from the 7,230 extracted as in AdditionalFile8.<br> For each abstract, target tuples follow the schema defined for AdditionalFile3.</p> <p>AdditionalFile11-test_dataset_prediction:<br> List of predicted annotations of 100 abstracts considered for testing the prediction model on a subset of the CORD-19 biology-related cluster. As AdditionalFile6, it is split in 4 TSV, respectively for entity (a), effect (b), level (c), and whole tuple predictions (d).</p>
SARS-CoV-2 Infection and Clinical Signs in Cats and Dogs from Confirmed Positive Households in Germany
<p>Supplemental material and raw data referring to specified publication</p>
Free energy simulations of receptor-binding domain opening in the SARS-CoV-2 spike indicate a barrierless transition with slow conformational motions
<p>This online data set accompanies the manuscript entitled "Free energy<br> simulations of receptor-binding domain opening in the SARS-CoV-2 spike<br> indicate a barrierless transition with slow conformational motions."</p> <p>The dataset is composed of the following files:</p> <p>* pmf0-now.dcd -- pmf63-now.dcd : molecular dynamics trajectory frames in<br> each of the 64 umbrella sampling windows, from which water has been<br> removed to save space</p> <p>* s1am_0-now.pdb -- s1am_63-now.pdb : initial coordinates in each of the 64<br> umbrella sampling windows, from which water has been removed,<br> corresponding to the trajectory data above</p> <p>* view -- Visual Molecular Dynamics command script to load a trajectory, <br> e.g., in Linux, use "vmd -e view"</p> <p>* s1am_0-cg.dcd -- s1am_63-cg.dcd : molecular dynamics<br> trajectory frames in each of the 64 umbrella sampling windows, coarse-grained to<br> 1 bead per residue.</p> <p>* s1am_0-cg.pdb -- s1am_63-cg.pdb : initial coordinates in each of the 64<br> umbrella sampling windows, corresponding to the coarse-grained trajectory<br> data above.</p> <p>* viewcg -- Visual Molecular Dynamics command script to load a<br> coarse-grained trajectory, e.g., in Linux, use "vmd -e viewcg"</p> <p>* 0readme -- brief instructions on how to view the trajectories</p> <p>* colors.vmd -- utility script for VMD</p> <p>* covmacros.vmd -- VMD script to define coronavirus spike subdomains</p> <p>* fe.zip -- ZIP archive that contains data and Matlab analysis files to<br> reproduce the free energy profiles</p> <p>* diff.zip -- ZIP archive that contains data and Matlab analysis files to<br> reproduce the diffusion and mean first passage times calculations</p> <p>* pca-qha.zip -- ZIP archive that contains the data and Matlab analysis files<br> to compute the autocorrelation functions of trajectory displacements<br> along principal/quasiharmonic modes</p> <p>Each ZIP archive contains a "0readme" file with brief instructions, and also the <br> results of the calculations<br> </p>
All atom simulations snapshots and contact maps analysis scripts for SARS-CoV-2002 and SARS-CoV-2 spike proteins with and without ACE2 enzyme
<p><strong>The dataset contains a total of 40 snapshots of the four trajectories (10 snapshots each system = two per replica x 5 replicas/system):</strong></p> <ol> <li>SARS-CoV-2002 spike protein without ACE2</li> <li>SARS-CoV-2 spike protein without ACE2</li> <li>SARS-CoV-2002 spike protein with ACE2</li> <li>SARS-CoV-2 spike protein with ACE2</li> </ol> <p>Molecular dynamics simulation trajectories (320ns each) have been performed using the Amber ff14SB force field running with the Amber18 package at the the NSF-funded (OAC-1826915, OAC-1828163) ELSA high performance computing cluster at The College of New Jersey. Under the following simulation methodology:</p> <p><em>All-atom simulations were carried out with Amber18 (<a href="https://slack-redir.net/link?url=http%3A%2F%2Fambermd.org">ambermd.org</a>), and system components (protein, ions, water) were modeled with the included FF14SB and TIP3P parameter sets. Energy minimization used CPU pmemd, while later simulation stages used GPU pmemd. CoV2 and CoV1 systems with one RBD up (with/without ACE2) were solvated in 12 angstrom water shells. Cysteine residues identified in the initial models as having a disulfide bond (DB) were bonded using tLeap. All simulations used 0.150 M NaCl. Hydrogen mass repartitioning was applied only to the protein to enable a 4 fs timestep (<a href="https://slack-redir.net/link?url=https%3A%2F%2Fpubs.acs.org%2Fdoi%2Fabs%2F10.1021%2Fct5010406">https://pubs.acs.org/doi/abs/10.1021/ct5010406</a>). The SHAKE algorithm was applied to hydrogens, and a real-space cutoff of 8 angstroms was used. Periodic boundary conditions were applied and PME was used for long-range electrostatics. Minimization was by steepest descent (2000 steps) followed by conjugate gradient (3000 steps). Heating used two stages: (1) NVT heating from 0 K to 100 K (50 ps), and (2) NPT heating from 100 K to 300 K (100 ps). Restraints of 10 kcal mol<sup>-1</sup> angstrom<sup>-2</sup> were applied during minimization and heating to C-alpha atoms. During 6 ns of equilibration at 300 K C-alpha restraints were gradually reduced from 10 kcal mol<sup>-1</sup> angstrom<sup>-2</sup> to 0.1 kcal mol<sup>-1</sup> angstrom<sup>-2</sup>. Finally, restraints were released and 320 ns unrestrained production simulations were carried out for CoV2 and CoV1 systems. Production simulations began from the final equilibrated snapshots, and five copies of each system were simulated. As unrestrained systems can freely rotate we monitored simulations for any close contacts and found that in one copy of the CoV1 simulation without ACE2 and one RBD up that a few contacts close to 8 angstrom occur near the end of the 320 ns between the RBD and a different subdomain of the spike complex in a periodic image. However this did not influence analyzed structural properties which is verified by comparing results across simulations. The Monte Carlo barostat was used to maintain pressure (1 atm), and the Langevin thermostat was used to maintain 300 K temperature (collision frequency 1 ps<sup>-1</sup>), as implemented in Amber18. In aggregate, nearly 7 microseconds of simulation of systems ranging from 396,147 to 879,100 atoms was carried out for this work.</em><br> For further details on the trajectories, please contact Joseph Baker (bakerj@tcnj.edu).</p> <p><strong>Regarding the contact map analysis scripts (contactMaps_Analysis.tar.gz), they contain the following workflow:</strong></p> <p>contactmap --> source files from contact_map executable<br> process_nc.sh --> convert raw data from all-atom simulation to numbered PDB files and get the contact maps<br> frequency.lua --> read a set of PDB files and output the frequency count for each contact<br> consensus.fasta --> align sequence of Covid19 and SARS from Chimera<br> consensus.lua --> read data previously generated and compute the frequency per residue, among other things.<br> consensus.sh --> input information to consensus.lua<br> consensus.gp --> gnuplot script to plot figures</p> <p>This dataset and the code is part of tripartite collaboration between:</p> <ul> <li>The Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland (supported by the National Science Centre, Poland, under grant No. 2017/26/D/NZ1/0046)</li> <li>Department of Chemistry, The College of New Jersey, New Jersey, United States (supported by National Science Foundation under grant numbers OAC-1826915 and OAC-1828163).</li> <li>Jozef Stefan Institute, Ljubljana, Slovenia (supported by the Slovenian Research Agency (Funding No. P1-0055)).</li> </ul>
Raw diffraction data (CBF) for a structure of SARS-CoV-2 Main Protease bound to 2-Methyl-1-tetralone
<p>Data collected at beamline P11/PETRAIII Deutsches Elektronen Synchrotron DESY</p> <p>Info:</p> <p>run type: regular<br> run name: l6p17_09_001<br> start angle: 0.000000deg<br> frames: 1000<br> degrees/frame: 0.200000deg<br> exposure time: 40.000000ms<br> energy: 11.999832keV<br> wavelength: 1.033214A<br> detector distance: 200.000000mm<br> resolution: 1.304257A<br> aperture: 100um<br> filter transmission: 71.798748%<br> filter thickness: 75um<br> ring current: 119.222664mA</p> <p>Crystal-info:</p> <p>Co-crystallization of Sars-CoV-2 MPro with the compound was achieved by equlibrating a 6.25 mg/ml protein solution in 20 mM HEPES buffer (pH 7.8) containing 1 mM DTT, 1mM EDTA, and 150 mM NaCl against a reservoir solution of 100 mM MIB buffer (2:3:3 molar ratio of malonic acid, imidazole, and boric acid), pH 7.5, containing 25% v/v PEG 1500 and 5% v/v DMSO. Prior to crystallization compound solutions in DMSO were dried onto the wells of SwissCI 96-well plates. To achieve reproducible crystal growth seeding was used. Crystals appeared within a few hours and reached their final size after 2 -3 days. Crystals were manually harvested and flash cooled in liquid nitrogen for subsequent X-ray diffraction data collection.</p>
Genomic determinants of pathogenicity in SARS-CoV-2 and other human coronaviruses
<p><strong>Dataset S1.</strong>Complete nucleotide sequence alignment of all human CoV used for region identification. </p> <p><strong>Dataset S2.</strong>Complete nucleotide sequence alignment of all CoV (of human and non-human hosts).</p> <p><strong>Dataset S3.</strong>Distances between leaves (each CoV strain in Dataset S2 was considered), from every reference genome of each of the seven human CoV.</p> <p><strong>Dataset S4.</strong>Alignment of strains used for zoonotic jump analysis.</p>
Molecular dynamics trajectories for SARS-CoV-2 Mpro with 7 HIV inhibitors
<p>Raw trajectory data (GROMACS format) of all atom molecular dynamics simulation of COVID-19 related SARS-CoV-2 dimeric main protease (based on PDB 6LU7) with 7 kinds of HIV inhibitors (darunavir, indinavir, lopinavir, nelfinavir, ritonavir, saquinavir, and tipranavir) were calculated on massively parallel supercomputer HOKUSAI Big Waterfall at RIKEN ISC, and a special-purpose computer, MDGRAPE-4A, at RIKEN BDR, JAPAN. For each ligand, 200ns length 28 trajectories were calculated. Some of these trajectories were calculated further longer. We can observe formation of encounter complex and investigate potential binding sites on the surface of the dimeric protease. We hope that these raw data are valuable for further drug repurposing/development research targeting the SARS-CoV-2 main protease. We will submit analysis of these data to refereed journal.</p> <p>Molecular dynamics simulations were performed under NVT at 310K, with the time step 2.5fs. The starting structure was prepared based on PDB 6LU7, with amber14sb force field in about 10nm cubic box with periodic boundary conditions. The ligands were initially placed apart from the active sites of the dimeric main protease.</p> <p>We have also already deposited 10 microseconds trajectories of the dimeric protease without ligand (with amber99sb-ildn force field) in the repository https://data.mendeley.com/datasets/vpps4vhryg/1 (DOI:10.17632/vpps4vhryg.1).</p> <p>Files:</p> <ul> <li><strong><em>LIG</em></strong>_28traj200ns_every200ps.zip (28trajectories for each ligand) <ul> <li>traj200ns_every200ps/<strong><em>LIG</em></strong>/<strong><em>a</em></strong>/traj200ns_every200ps_<em><strong>LIG</strong>-<strong>a</strong>-<strong>n</strong></em>.xtc <ul> <li>(trajectory in GROMACS XTC)</li> </ul> </li> <li>traj200ns_every200ps/<strong><em>LIG</em></strong>/<strong><em>a</em></strong>/conf.gro <ul> <li>(initial condition in GROMACS GRO)</li> </ul> </li> <li>traj200ns_every200ps/<em><strong>LIG</strong></em>/topology/ <ul> <li>(contains topology files)</li> </ul> </li> <li>traj200ns_every200ps/<strong><em>LIG</em></strong>/mdp/ <ul> <li>(contains run paramter files)</li> </ul> </li> </ul> </li> <li>ZZZ_20traj1us_every200ps.zip (20trajectories extended to 1microsecond) <ul> <li>traj1us_every200ps/traj1us_every200ps_<em><strong>LIG</strong>-<strong>a</strong>-<strong>n</strong></em>.xtc <ul> <li>DAR-C-06, DAR-D-07</li> <li>IND-C-05, IND-C-06, IND-D-06</li> <li>LOP-A-02, LOP-D-03</li> <li>NEL-B-01, NEL-C-07, NEL-D-02</li> <li>RIT-B-07, RIT-C-07</li> <li>SAQ-B-01, SAQ-C-04, SAQ-D-03</li> <li>TPR-A-07, TPR-B-04, TPR-B-06, TPR-C-05, TPR-D-02</li> </ul> </li> </ul> </li> <li>ZZZ_3traj6us_every1ns.zip (3trajectories extended to 6microseconds or more) <ul> <li>traj6us_every1ns/traj6us_every1ns_<em><strong>LIG</strong>-<strong>a</strong>-<strong>n</strong></em>.xtc <ul> <li>IND-D-06, NEL-B-01, TPR-B-04</li> </ul> </li> </ul> </li> </ul> <p> </p> <ul> <li>ZZZ_LigandBindingPosePDBs.zip (pickup 3 snapshots for each ligand) <ul> <li>LigandBindingPosePDBs/<em><strong>LIG</strong>-<strong>a</strong>-<strong>n</strong></em>_frame.pdb</li> </ul> </li> </ul> <p> </p> <ul> <li>movies_overlooking_28traj200ns.zip (7x2movies) <ul> <li>movies_28traj200ns/movie_overlooking_<strong><em>LIG</em></strong>_28traj200ns-viewA.mp4 <ul> <li>inspecting 28traj at once</li> </ul> </li> <li>movies_28traj200ns/movie_overlooking_<strong><em>LIG</em></strong>_28traj200ns-viewB.mp4 <ul> <li>from the opposite angle</li> </ul> </li> </ul> </li> <li>movies_1us.zip (17movies) <ul> <li>movies_1us/movie_<em><strong>LIG</strong>-<strong>a</strong>-<strong>n</strong></em>_1us.mp4</li> </ul> </li> <li>movies_6us.zip (3movies) <ul> <li>movies_6us/movie_<em><strong>LIG</strong>-<strong>a</strong>-<strong>n</strong></em>_6us.mp4</li> </ul> </li> </ul> <p> where</p> <p> <em><strong>LIG</strong></em>={DAR,IND,LOP,NEL,RIT,SAQ,TPR}<br> DAR:darunavir<br> IND:indinavir<br> NEL:nelfinavir<br> RIT:ritonavir<br> SAQ:saquinavir<br> TPR:tipranavir<br> <em><strong>a</strong></em>={A,B,C,D}<br> <em><strong>n</strong></em>={01,02,03,04,05,06,07}</p> <p> </p> <ul> <li>ZZZz_3traj1us_every200ps_unbinding.zip (3trajectories extended to 1microsecond exhibiting unbinding)</li> <li>ZZZz_56traj200ns_every200ps_negative_control.zip (56trajectories as a negative control)</li> <li>ZZZz_LigandBindingPosePDBsRevised.zip (pickuped 3 snapshots for each ligand)</li> </ul> <p> </p>
Electron microscopy of SARS-CoV-2 particles - Dataset 05
<p>The dataset contains transmission electron microscopy image stacks (tomograms) of ultrathin sections through extracellular SARS-CoV-2 particles in Vero cell cultures. The dataset contains 17 image stacks of slightly variable pixel dimensions, which were recorded at either 1.17 or 0.96 nm pixel size (12 bit). Image stacks were size calibrated and stored in 16 bit TIF format. Visualization can be done using ImageJ or Fiji. Each image stack in TIF format is supplemented by a file containing the corresponding raw image tilt series (MRC format; plus meta data files) generated by the tomography acquisition software and by a file with the aligned tiltseries. A PDF document describes the methods used for generation of the image files. The dataset was generated as dataset 05 for a comparative morphometric analysis of SARS-CoV and SARS-CoV-2. Further datasets which were used for the analysis are available in this repository (see dataset description document).</p> <p>Related publication: Laue M, Kauter A, Hoffmann T, Möller L, Michel J, Nitsche A. Morphometry of SARS-CoV and SARS-CoV-2 particles in ultrathin plastic sections of infected Vero cell cultures. Sci Rep. 2021 Feb 10;11(1):3515. doi: 10.1038/s41598-021-82852-7. PMID: 33568700; PMCID: PMC7876034.</p> <p> </p>
Incidence of SARs-CoV-2 in Gütersloh county, Germany, after the outbreak in the slaughterhouse and meat packing plant Tönnies
<p>Figure </p> <p>Seven day incidence of SARS-CoV-2 per 100,000 people from March 15 to September 3, 2020 in Gütersloh, North Rhine-Westphalia, Germany</p> <p>Table</p> <p>Pandemic control measures in Gütersloh county</p>
Studied disinfectant substances against SARS-CoV-2 and other coronaviruses
<p>This data-sheet covers those disinfectants tested against SARS-CoV-2 or other coronaviruses. Data were extracted from several research articles indicated in the reference row. The data-sheet comprises a total of 11 fields with info regarding the virus (virus and strain/isolate names), formulation (substance(s) and its concentration in percentage) and test characteristics (suspension or surface tested, kind of surface, use dilution before testing, disinfectant and inoculum volumes, organic load type and concentrations and contact time) as well as their results, normalized in terms of Log<sub>10 </sub>viral infectivity reduction. Data is included and comented in the following journal article: <a href="https://doi.org/10.3390/foods10020283">https://doi.org/10.3390/foods10020283</a> Please reference also to this publication if using the data-sheet.</p>
Data for the article: "Molecular Modelling Reveals Eight Novel Druggable Binding Sites in SARS-CoV-2's Spike Protein" by Ilke Ugur and Antoine Marion
<p>This upload contains data related to the article<br> published as a preprint on ChemRxiv with DOI<br> https://doi.org/10.26434/chemrxiv.13292768</p> <p>"Molecular Modelling Reveals Eight Novel Druggable Binding Sites in SARS-CoV-2's Spike Protein"<br> by Ilke Ugur and Antoine Marion (2020)<br> Department of Chemistry, Middle East Technical University, Ankara, Turkey.</p> <p>For further information, please contact:<br> ilkeugur@metu.edu.tr ; amarion@metu.edu.tr</p> <p>The manuscript is currently under peer-review.</p> <p>Content:</p> <p>Library of molecules derived from DrugBank v 5.1.5:<br> - DrugBank_2020_5.1.5/ # All necessary files for the docking and refinement of the library of molecules.<br> -- DB_5.1.5_pH7.4_pdbqt/ ## PDBQT readily usable for docking with AutoDock Vina.<br> -- DB_5.1.5_pH7.4_mol2amber/ ## mol2 files containing assigned GAFF atom types and Gasteiger atomic charges.<br> -- DB_5.1.5_pH7.4_frcmod/ ## frcmod files containing missing molecular mechanics parameters<br> -- dbID_name.dat ## DrugBank ID to generic name dictionary</p> <p>Note: The files were prepared automatically via a series of operations handling openbabel and antechamber.<br> The protonation state of ionizable groups as well as Gasteiger atomic charges were assigned by openbabel for a pH of 7.4<br> mol2 and frcmod files can be used readily via the tleap module of AmberTools to produce topology files.</p> <p><br> Receptor structures:<br> - receptors/ # PDB files for the four structures of the spike protein considered in this work<br> -- CS00ns.pdb ## Closed state after the remodelling of missing loops (PDB ID 6vxx)<br> -- OS00ns.pdb ## Open state after the remodelling of missing loops (PDB ID 6vyb)<br> -- CS25ns.pdb ## Closed state after 25 ns of molecular dynamics in explicit water<br> -- OS25ns.pdb ## Open state after 25 ns of molecular dynamics in explicit water</p> <p>Note: All structures are aligned to CS00ns.pdb and can be converted to pdbqt for docking with AutoDock Vina</p> <p><br> Docking grid centers:<br> - dockingCenters/ # XYZ files containing the coordinates of each docking grid center considered in this work</p> <p>Note: The coordinates are given in the same frame as that of the four structures of the receptor.</p> <p><br> Binding sites:<br> - bindingSites/ # XYZ files with the coordinates of the representative atomic centres<br> # of each binding site identified in this work (A-H).</p> <p>Note: These files can be used to get a clearer picture of the binding sites within the structures<br> of the spike protein shared in the receptors directory.</p> <p><br> Final modelling results:<br> - allData.txt # data for all molecules in the set (approved and investigational)<br> - appData.txt # data for approved molecules only<br> - data.xlsx # data for all molecules in the set (approved and investigational)<br> # as a formatted excel spreadsheet</p> <p>Note: The columns are delimited with semi-colons ";".<br> The files contain the results for the best pose of all approved molecules for which<br> molecular mechanics-based geometry optimization succeeded, regardless of their score.<br> For other molecules, the result of their best pose is reported only for those complexes<br> having MM interaction energy lower or equal to -22.00 kcal/mol.</p> <p><br> Visualization:<br> - bs.pse # pymol session representing the binding sites within the<br> # closed state structure of the spike protein (CS00ns)<br> - pt.pse # pymol session representing the docking grid centres within<br> # closed statestructure of the spike protein (CS00ns)</p> <p>Note: the PSE files should be compatible with version 7.0 of pymol and later</p>
Supplementary Data -A STUDY ON SOME STRUCTURAL FEATURES RESPONSIBLE FOR SARS-COV-2 INFECTION FATALITY
<p>A correlation between hydrodynamic properties like radius of gyration ( Rg ) vs Molecular weight of spike protein of SARS - COV-2 biopolymers .</p>
Vector sequences in early WIV SRA sequencing data of SARS-CoV-2 inform on a potential large-scale security breach at the beginning of the COVID-19 pandemic
<p>DESCRIPTION</p> <p>Sequences identified as Influenza A virus, Spodoptera frugiperda rhabdovirus and Nipah henipavirus have been previously identified within the early HiSeq 1000 and HiSeq 3000 sequencing data of SARS-CoV-2, SRR11092059,SRR11092060,SRR11092061 and SRR11092062, and were being used to support the hypothesis that a "simultaneous outbreak of multiple zoonotic viruses" have happened in the Huanan Seafood market. https://doi.org/10.31219/osf.io/s4td6</p> <p>However, a closer examination of these sequences revealed that they were not sequences of actual wild viruses, but were in stead fragments left behind from PCR products and cloning vectors harboring both cDNA clones and infectious clones of such viruses, with evidence of viral sequences being joined directly to DNA sequences of vector and non-human origin within the same short reads.</p> <p>Here are the vector sequences and PCR product-like sequences recovered from the earliest WIV SRA sequencing data of Human SARS-CoV-2 from dataset SRR11092059,SRR11092060,SRR11092061,SRR11092062.</p> <p>Sequences associated with Vectors and PCR products from 3 distinct viral species have been obtained: The 3'-end of a Nipah Henipahvirus with fusion to a Hepatitis D virus Ribozyme, a T7 terminator and a Tetracycline resistance gene, The 5'-end of the same Nipah Henipahvirus with fusion to sequences found in diverse vectors, A complete vector genome encoding the HA gene of Influenza A virus subtype H7N9 under a CMV promoter and a bgH polyA terminator, and 221 Contiguous sequences corresponding to the Spodoptera frugiperda rhabdovirus reference genome fused to sequences that were homologous to multiple Plastid sequences and Notably Mitochondrial sequences of Rodents.</p> <p>As sequences corresponding to a rescued infectious clone of a BSL-4 organism (Nipah Henipahvirus) were found in sample sequences that supposedy represents patient samples that were obtained from Hospital ICU and sequenced in a pathogen diagnosis laboratory (which is separate from the Virology Research laboratory which is implied by the context of an Infectious Clone of such an organism, evident by the 3'-HDV ribozyme and T7 terminator fused directly to the 3'-terminus of the Nipah Henipahvirus reads), The discovery of artifact-containing sequences of at least 3 different pathogen species that are phylogenetically and methodologically distinct from each other in samples that were supposedly submitted by a laboratory that is Separate from the virological research laboratories that could have hosted such clone sequences imply extensive crosstalk and cross-contamination between the various laboratories within the Wuhan Institute of Virology, which includes at least one BSL-4 laboratory with evidence of containment breach of a BSL-4 organism and it's subsequent introduction into RNA-seq samples that were processed by a laboratory of distinct and separate purposes than the basic virological research evidenced by the Infectious Clone of the Hipah Henipahvirus.</p> <p>Such a discovery therefore likely imply a major security breach happening within the Wuhan institute of Virology at the time when the first sequences of SARS-CoV-2 was sampled and sequenced, which have important implications on the origins of the SARS-CoV-2 virus itself.</p> <p>METHODS</p> <p>The metagenomic sequencing datasets, SRR11092059,SRR11092060,SRR11092061 and SRR11092062 were first analyzed using the NCBI phylogenetic analysis tool, which identified viral sequences that is not related to SARS-CoV-2 itself. These include Influenza A virus (IAV, subtype H7N9), Spodoptera frugiperda rhabdovirus and Nipah Henipahvirus.</p> <p>The datasets were then subjected to BLAST search using MEGABLAST against the reference sequences of such viruses to verify the existence of the viral sequences and determine the exact sybtype of such viruses and the closest sequences on GenBank that corresponds to the reads. There seuqences are MH926031.1 for the Spodoptera frugiperda rhabdovirus, KY199425.1 for the Influenza A virus and AY988601.1 for the Nipah Henipahvirus.</p> <p>A second round BLAST analysis with these identified sequences were then performed, which unexpectedly revealed numerous reads corresponding to Cloning vectors and non-human Mitochondrial and Plastid sequences being fused directly to the sequences of the identified viral species. Reads were then downloaded and subjected to assembly using the CAP3 sequence assembly program and the EGASSEMBLER tool. Contig sequences were then queried against the NCBI nr/nt database which unanimously identified the original sample sequences as viral sequences inserted into cloning vectors.</p> <p>The complete sequence of the Influenza A virus Haemagluttinin (HA) gene clone was obtained from SRR11092061,SRR11092062 using multiple rounds of BLAST search and sequence assembly expansion on the existing vector-virus junction contigs, and a partial sequence corresponding the 3'-end of Nipah Henipahvirus AY988601.1 fused to a 3'-HDV ribozyme, T7 terminator and a Tet resistance gene was obtained from SRR11092059. In addition, 221 Contig sequences corresponding to the Rhabdovirus MH926031.1 fused to Chloroplast sequence MN524635.1 and Rodent Mitochondrial sequence MT241668.1 have been recovered from SRR11092061.</p> <p>We then performed a BLAST search using the identified vector sequences on SRR11092059,SRR11092060,SRR11092061 and SRR11092062, which confirms the existence of these two vetor sequences in all 4 datasets.</p>
Epidemiology, risk factors and clinical course of SARS-CoV-2 infected patients in a Swiss university hospital: an observational retrospective study
<p>This is the dataset of the study called "Epidemiology, risk factors and clinical course of SARS-CoV-2 infected patients in a Swiss university hospital: an observational retrospective study". <br> <br> <strong>Abstract: </strong></p> <p>Background<br> Coronavirus disease 2019 (COVID-19) is now a global pandemic with Europe and the USA at its epicenter. Little is known about risk factors for progression to severe disease in Europe. This study aims to describe the epidemiology of COVID-19 patients in a Swiss university hospital.</p> <p>Methods<br> This retrospective observational study included all adult patients hospitalized with a laboratory confirmed SARS-CoV-2 infection from March 1 to March 25, 2020. We extracted data from electronic health records. The primary outcome was the need to mechanical ventilation at day 14. We used multivariate logistic regression to identify risk factors for mechanical ventilation. Follow-up was of at least 14 days. <br> <br> Results<br> 200 patients were included, of whom 37 (18·5%) needed mechanical ventilation at 14 days. The median time from symptoms onset to mechanical ventilation was 9·5 days (IQR 7.00, 12.75). Multivariable regression showed increased odds of mechanical ventilation in males (3.26, 1.21-9.8; p=0.025), in patients who presented with a qSOFA score ≥2 (6.02, 2.09-18.82; p=0.001), with bilateral infiltrate (5.75, 1.91-21.06; p=0.004) or with a CRP of 40 mg/l or greater (4.73, 1.51-18.58; p=0.013). <br> <br> Conclusions<br> This study gives some insight in the epidemiology and clinical course of patients admitted in a European tertiary hospital with SARS-CoV-2 infection. Male sex, high qSOFA score, CRP of 40 mg/l or greater and a bilateral radiological infiltrate could help clinicians identify patients at high risk for mechanical ventilation.</p>
All-atom Molecular Dynamics Simulations of SARS-CoV-2 Spike Receptor-binding Domain bound with ACE2
<p>Data includes all of the trajectories (1000) of classical all-atom molecular dynamics (MD) simulations of of SARS-CoV2 Spike Protein/ACE2 complex (PDB ID: 6M0J). In order to decrease the size of the file only protein rajectories were provided. Simulation has been performed with Desmond. Protein was placed in the cubic boxes with explicit TIP3P water models that have 10.0 Å thickness from surfaces of protein. The system is neutralized by adding counter ions, and salt solution of 0.15M NaCl was also used to adjust the concentration of the systems. The long-range electrostatic interactions were calculated by the particle mesh Ewald method. A cutoff radius of 9.0 Å was used for both van der Waals and Coulombic interactions. The temperature was set as 310K initially, and Nose–Hoover thermostat was used for adjustment. Martyna–Tobias–Klein protocol was employed to control the pressure, which was set at 1.01325 bar. The time-step was assigned as 2.0 fs. The default values were used for minimization and equilibration steps, and finally 100 ns production run was performed for the simulation.</p>
Files and code for English dictionaries, gold and silver standard corpora for biomedical natural language processing related to SARS-CoV-2 and COVID-19
<p><span lang="EN-GB">Automated information extraction with natural language processing (NLP) tools is required to gain systematic insights from the large number of COVID-19 publications, reports and social media posts, which far exceed human processing capabilities. </span></p> <p><span lang="EN-GB">Here we present an NLP toolbox comprising COVID-19-related dictionaries and annotated corpora in English as well as useful code and workflows for their update and use. The dictionaries contain terms referring to the COVID-19 disease, the SARS-CoV-2 virus, its variants and common mutations, respectively. They were used together with the EasyNER NLP tool to extract and annotate all 764 398 abstracts in the CORD-19 dataset, creating a very large silver standard corpus (named Lund-Annotated-CORD-19 corpus). This was complemented with a small gold standard corpus consisting of PubMed abstracts manually annotated for key entity classes such as disease, virus, symptom, protein/gene, cell type, chemical and species terms. </span></p> <p><span lang="EN-GB">The toolbox can support various text analysis tasks related to COVID-19 such as named entity recognition and co-mention analysis. A preliminary version of the toolbox, which was released early in the pandemic, was</span><span lang="EN-GB"> for example already used to create a COVID-19 knowledge graph and study the evolution and variation of COVID-19-related terminology. In addition, the toolbox can be applied in the development of other NLP tools, for example to train and evaluate large language models.</span></p> <p><span lang="EN-GB">When using the toolbox, please cite this record and the associated article.</span></p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.