Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,489
datasets available to search
ShareScore release 0.9.0
Dataset results
2,489 results for “Sars-CoV-2”
Oral microbiota composition and its relationship with epidemiological, clinical and microbiological variables of SARS-CoV-2 infection in children and adults under strict home confinement in Barcelona, Spain
Open the record for dataset details and reuse information.
Figure 3 in Potential histopathological and immunological effects of SARS-CoV-2 on the liver
Figure 3. Immune-mediated liver injury adapted from Spearman et al. (2021).
Figure 4 in Potential histopathological and immunological effects of SARS-CoV-2 on the liver
Figure 4. The major liver histological features adapted from Díaz et al. (2020).
Figure 1 in Potential histopathological and immunological effects of SARS-CoV-2 on the liver
Figure 1. SARS-CoV-2 structure and life cycle adapted from Zhong et al. (2020).
The early evolution of the BA.2.86 VOI sheds light on the origins of highly divergent SARS-CoV-2 lineages
<p><span>Alignments corresponding to the D1-D4.1 datasets. Five different datasets (named here D1, D2, D3, D4 and D4.4) comprising complete SARS-CoV-2<strong> </strong>genomes (low-coverage excluded, Ns ≤ 5%) were compiled from assemblies retrieved from GISAID. Genomes within these datasets were collated at different time points, aligned and sequentially subsampled under a phylogenetically-informed approach. </span></p>
Data set for "SARS-CoV-2 introductions to the island of Ireland: a phylogenetic and geospatiotemporal study of infection dynamics"
<p>Please see README.txt for detailed information about the contents of this data repository.</p>
Intra-host single-genome sequences of SARS-CoV-2 spike from people without HIV and people living with HIV
<p>Supporting data and code for the study "Rapid Intra-host Diversification and Evolution of SARS-CoV-2 in Advanced HIV Infection" (2024) by Ko & Radecki et al.</p> <p>See <a href="https://github.com/niaid/UMI-pacbio-pipeline/releases/tag/SC2-HIV-demo">the GitHub</a> for demo of SGS and haplotype generation using example data from the study.</p> <p>Raw sequencing data are deposited at <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1055920">PRJNA1055920</a>.</p> <p>READMEs are available within individual folders and releases.</p>
Bioinformática para el análisis de SARS-CoV-2 para principiantes / Bioinformatics for SARS-CoV-2 analysis for beginners
<ul> <li>Data for the e-learning tutorial <a href="https://github.com/cabana-online/Course_SARS_CoV-2">Bioinformatics for SARS-CoV-2 analysis for beginners</a></li> <li>Datos para el tutorial e-learning <a href="Bioinformática para el análisis de SARS-CoV-2 para principiantes" target="_blank" rel="noopener">Bioinformática para el análisis de SARS-CoV-2 para principiantes</a></li> </ul>
Random samples of SARS-CoV-2 sequences for Germany
<p>Mutation profiles and meta data for SARS-CoV-2 lineages in Germany submitted via DESH to the Robert Koch Institute (<a href="https://github.com/robert-koch-institut/SARS-CoV-2-Sequenzdaten_aus_Deutschland">https://github.com/robert-koch-institut/SARS-CoV-2-Sequenzdaten_aus_Deutschland</a>), containing a subset from https://doi.org/10.5281/zenodo.12806236.<br>For this CSV file, the raw data (fasta) was processed with covSonar (https://github.com/rki-mf1/covsonar) and the resulting database queried for random samples (RKI_STICHPROBE, DESH_STICHPROBE) and with dates ranging from 2021-01-01 until 2024-07-24.</p>
Evolutionary trends in the compositional structure of the SARS-CoV-2 genome during the pandemic
<p>By computing a measure of compositional genome structure in random datasets of coronavirus genomes, we observed a long-term decreasing trend in this measure, accompanied by an increasing evolutionary rate.</p> <p> </p>
Docking data for "The evolution of the SARS-CoV-2 spike protein for differential usage of the host transmembrane serine proteases entry pathway"
<p><br>The dataset includes predicted complexes of the SARS-CoV-2 Spike protein (specifically at the S2' cleavage site) with Hepsin and TMPRSS2 proteins. It contains data on three variants: Wuhan, Delta, and Omicron BA.1.</p> <p><strong>Compressed folders:</strong></p> <p>-357596-DeltaHepsin.tgz</p> <p>-357597-DeltaTMPRSS2.tgz</p> <p>-360039-WuhanHepsin.tgz</p> <p>-360042-TMPRSSWuhan.tgz</p> <p>-392981-TMPRSS-BA1_all.tgz</p> <p>-392982-Hepsin-BA-all.tgz</p> <p><strong>Each compressed folder contains the following:</strong></p> <p>-Initial structures in pdb format</p> <p>-Output complexes in pdb format</p> <p>-Clusters in pdb format</p> <p>-Protocols</p> <p>-Parameters</p> <p>-Scoring files</p> <p> </p> <p><strong>Protein-protein docking </strong><br>Molecular docking between the SARS-CoV-2 S protein of Wuhan, Delta (PDB: 7W92, [DOI: 10.1038/s41467-022-28528-w]), and BA.1 (PDB: 7XO5, [DOI: 10.1038/s41422-022-00672-4]) and the human proteases TMPRSS2 (PDB: 8HD8, [DOI: 10.1038/s41467-023-42527-5]) and Hepsin (PDB: 1Z8G, [DOI: 10.1042/BJ20041955]) was performed using the HADDOCK v2.5-2024.03 webserver ([DOI: 10.1021/ja026939x], [DOI: 10.1016/j.jmb.2015.09.014]). Missing loops in the protein structures were reconstructed using Modeller v10.5 ([DOI: 10.1006/jmbi.1993.1626]). Every heteroatom was removed from the reference structures. The relaxed atomistic coordinates for each S protein variant were derived via all-atom molecular dynamics (MD) simulations. These simulations were performed using AMBER22 with the FF19SB force fields and the pmemd.cuda module for enhanced performance ([DOI: 10.1021/acs.jcim.3c01153], [DOI: 10.1021/jz501780a], [DOI:10.1021/ct400314y]). For the Wuhan variant the S protein was retrieved from our previous modeling study [DOI: 10.1039/D0NR03969A] where for Delta and BA.1, ecah S protein was placed in a dodecahedral box, extending 20 Å beyond the solute in every cartesian direction, and solvated with the four-site OPC water model ([DOI: 10.1021/jz501780a]). The systems were neutralized with counterions, specifically one Cl− ion for the Delta variant and three Cl- ions for the BA.1 variant. To remove local clashes, a geometric optimization was performed using the steepest descent algorithm for 5000 cycles. The MD equilibration process consisted of several stages. First, temperature equilibration in the NVT ensemble was performed by gradually increasing the temperature through steps of 150, 200, 250, 300, and finally 310 K, each lasting 200 ps. During this phase, position restraints were applied to the heavy atoms of the proteins, with progressively decreasing spring constants of 5.0, 4.0, 3.0, and 1.0 kcal mol−1 Å−2, facilitating gradual relaxation. This was followed by a 1 ns equilibration at 310 K in the NPT ensemble without restraints. For production MD, the simulations were run in the NPT ensemble with periodic boundary conditions and Particle Mesh Ewald (PME) method ([DOI: 10.1063/5.0040966], [DOI: 10.1021/ct9001015]) using a grid spacing of 1.0 Å for long-range electrostatics. Non-bonded interactions were modeled with a Lennard-Jones potential using a 9Å cutoff. Temperature control was maintained using Langevin dynamics ([DOI: 10.1021/ct800573m]) with a collision frequency of 4.0 ps−1, and pressure control was managed by the Monte Carlo barostat ([DOI: 10.1016/j.cplett.2003.12.039]) with a 2.0 ps relaxation time at 1 bar. Bond constraints on hydrogen atoms were applied using the SHAKE algorithm ([DOI: 10.1016/0021-9991(77)90098-5]), and the hydrogen mass repartitioning scheme was applied via ParmEd ([DOI: 10.1371/journal.pcbi.1005659]), enabling a 4 fs integration time step ([DOI: 10.1021/ct5010406]). Each protein complex was simulated for a total of 20 ns. For the Wuhan variant, the 3D coordinates were retrieved from [DOI: 10.5281/zenodo.3817446].<br>The active interaction region on the spike protein was defined as the cleavage site (residues P809-R815). For TMPRSS2 and Hepsin, the active sites were defined based on their catalytic residues: H296, D345, D435, S441, S460, and G462 for TMPRSS2, and H203, D257, D347, A348, and S353 for Hepsin. These specific regions were selected to guide the docking process and maximize biologically relevant interactions. Docking clusters were analyzed by selecting those with the lowest interaction energies for further structural analysis. To evaluate binding accuracy, native contacts between the S protein and proteases were computed using the contact map analysis based on the OV+rCSU method ([DOI: 10.12693/APhysPolA.145.S9, 10.1021/acs.jctc.6b00986]), which allows for a precise identification of critical stabilizing interactions, both specific and non-specifics. High-frequency contacts, defined as those appearing in over 70% of the generated models, were highlighted as key determinants of protein-protein recognition, providing insight into the most stable and consistent interactions across docking configurations.</p>
Supplementary Data for Ongoing Global and Regional Adaptive Evolution of SARS-CoV-2
<p>This repository includes:<br> The global tree in Newick format: global_2_13_21.main.tre<br> The global, ultrametric tree in Newick format: global_2_13_21.ultra.tre<br> All metadata used in this study including the GISAID acknowledgements: metadata.tgz<br> Tables S1-S4 in plain text format.</p>
The use of nanobodies in a sensitive ELISA test for SARS-CoV-2 Spike 1 protein
<p>A rapid detection method for SARS-CoV-2 spike protein is essential for control of COVID19. We investigated various combinations of engineered nanobodies in a sandwich ELISA to detect the Spike protein of SARS-CoV-2. We have identified an optimal combination of nanobodies. These were selectively functionalised to further improve antigen capture. This dataset contains data from ELISA experiments described in the manuscript.<span> </span></p> <p><span>Plate coating of nanobodies for ELISA by passive adsorption vs biotinylation was compared. A series of nanobody pairings (two cluster 2 ACE2-binding epitope and two cluster 1 CR3022 epitope) were screened for optimum sensitivity. The optimal pair were then tested against a series of SARS-COV-2 antigens: recombinant spike 1 protein; recombinant receceptor binding domain (RBD); pseudotyped HIV-1 and heat-empigen inactivated SARS-CoV-2 virus. X-ray irradiated SARS-CoV-2 was also tested. Sensitivity to these antigens was compared with nanobodies biotinylated a) site-selectively and b) in a non-specific stochastic manner. Batch-to-batch viral variation and effects of inactivating agents were investigated. Limit of detection was compared against delta and beta viral mutants. Combining optimal nanobody pairing and site-selective biotinylation, we observed a limit of detection of 147 pg/mL for Spike protein; 33 pg/mL for RBD; 16 TCID50/mL of pseudovirus and 15 ffu/mL of heat-Empigen inactivated SARS-CoV-2. The pairing also showed sensitivity towards delta variant. We have demonstrated the use and sensitivity of nanobodies in ELISA by detection of recombinant and viral SARS-CoV-2 antigens.</span></p>
Training data for "Identification of allelic variants in SARS-CoV-2 from deep sequencing reads"
<p>Effectively monitoring global infectious disease crises, such as the COVID-19 pandemic, requires capacity to generate and analyze large volumes of sequencing data in near real time. These data have proven essential for monitoring the emergence and spread of new variants, and for understanding the evolutionary dynamics of the virus.</p> <p>Two sequencing platforms in combination with several established library preparation strategies are predominantly used to generate SARS-CoV-2 sequence data. However, data alone do not equal knowledge: they need to be analyzed. The Galaxy community developed analysis workflows to support the <strong>identification of allelic variants (AVs) in SARS-CoV-2 from deep sequencing reads</strong>.</p> <p>These workflows allow one to identify AVs and lineages in SARS-CoV-2 genomes with variant allele frequencies ranging from 5% to 100% (i.e., they detect variants with intermediate frequencies as well.</p> <p>In this tutorial we will see how to run these workflows for the different types of input data:</p> <ul> <li>Single end data derived from Illumina-based RNAseq experiments</li> <li>Paired end data derived from Illumina-based RNAseq experiments</li> <li>Paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols</li> <li>ONT fastq files generated with Oxford nanopore (ONT)-based Ampliconic (ARTIC) protocols</li> </ul> <p>To illustrate the tutorial, we took some example datasets (paired-end data generated with Illumina-based Ampliconic (ARTIC) protocols) from COG-UK, the COVID-19 Genomics UK Consortium.</p>
Molecular Dynamics of SARS-CoV-2 Delta Variant Receptor Binding Domain in Complex with ACE2 Receptor
<p>Molecular dynamics simulation for 10 ns at 37 C degrees of SARS-CoV-2 delta variant. Performed with NAMD and visualized/analyzed in ChimeraX software using Frontera supercomputer from Texas Advanced Computing Center. By Victor Padilla-Sanchez, PhD.</p> <p>https://www.youtube.com/watch?v=8N_MjWwxbMQ</p>
CSV-format data for: Increased mortality in community-tested cases of SARS-CoV-2 lineage B.1.1.7
<p>This is a supplementary upload to <a href="https://zenodo.org/record/4579857">https://zenodo.org/record/4579857</a>.</p> <p>This upload provides the same anonymised individual-level SARS-CoV-2 testing data for England as that earlier upload provided, but provides it in CSV format (comma-separated values) instead of in the previous QS format. The QS format requires specialised software (e.g. the <strong>qs</strong> package for R) to read, so I am providing the data in CSV format to facilitate access to and re-use of the data.</p> <p>Please see the original data upload, <a href="https://zenodo.org/record/4579857">https://zenodo.org/record/4579857</a>, the associated journal article, <a href="https://www.nature.com/articles/s41586-021-03426-1">Increased mortality in community-tested cases of SARS-CoV-2 lineage B.1.1.7</a>, and the project's Github repository, <a href="https://github.com/nicholasdavies/cfrvoc">https://github.com/nicholasdavies/cfrvoc</a>, for full details.</p>
Evidence summary on activities or settings associated with a higher risk of SARS-CoV-2 transmission: Summary of included evidence syntheses
<p>This is a data extraction table associated with the HIQA report entitled: Evidence summary on activities or settings associated with a higher risk of SARS-CoV-2 transmission</p>
Pandemic, epidemic, endemic: B cell repertoire analysis reveals unique anti-viral responses to SARS-CoV-2, Ebola and Respiratory Syncytial Virus
<p>VDJ gene usage, and associated amino acid sequences and properties, from healthy controls as well as patients with COVID-19, RSV or Ebola and Yellow fever vaccine recipients.</p>
SARS-CoV-2 raw fastq files
<p>Illumina paired end FASTQ files from MiSeq. Sequenced at the African Centre of Excellence for Genomics of Infectious Diseases (ACEGID), Redeemer's University, Ede, Osun State.</p> <p>Training data</p>
SARS-CoV-2 viral loads across the upper and lower respiratory tract, sex, disease severity and age groups for adult and pediatric COVID-19
<p>This dataset shows SARS-CoV-2 respiratory viral loads (viral RNA concentration in the respiratory tract) in the upper and lower respiratory tract for age, sex and COVID-19 severity groups. The data were obtained from a systematic review. The model outputs show the Weibull distributions, case percentiles, and sensitivity & specificity when using SARS-CoV-2 viral load (URT or LRT) as a prognostic indicator. See our paper for more information.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.