Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,562
datasets available to search
ShareScore release 0.9.0
Dataset results
2,562 results for “SARS CoV 2”
SARS-CoV-2 variants
<p>This dataset is a temporary upload and will be removed after double-blind peer review process is finished.</p>
Data and scripts from Epidemiological and clinical insights from SARS-CoV-2 RT-PCR crossing threshold values, France, January to November 2020
<p>Raw data and scripts used in the publication "<em>Epidemiological and clinical insights from SARS-CoV-2 RT-PCR crossing threshold values, France, January to November 2020 separator commenting unavailable</em>" in Eurosurveillance in 2022.</p> <p> </p> <p>https://www.eurosurveillance.org/content/10.2807/1560-7917.ES.2022.27.6.2100406</p>
Automated and manual pooled sample testing with panther fusion and aptima SARS-CoV-2 assays
<p>Combining diagnostic specimens into pools has been considered as a strategy to augment throughput, decrease turnaround time, and leverage resources. This study utilized a multi-parametric approach to assess optimum pool size, impact of automation, and effect of nucleic acid amplification chemistries on the detection of SARS-CoV-2 RNA in pooled samples for surveillance testing on the Hologic Panther Fusion® System. Dorfman pooled testing was conducted with previously tested SARS-CoV-2 nasopharyngeal samples using Hologic's Aptima® and Panther Fusion® SARS-CoV-2 Emergency Use Authorization assays. A manual workflow was used to generate pool sizes of 5:1 (five samples: one positive, four negative) and 10:1. An automated workflow was used to generate pool sizes of 3:1, 4:1, 5:1, 8:1 and 10:1. The impact of pool size, pooling method, and assay chemistry on sensitivity, specificity, and lower limit of detection (LLOD) was evaluated. Both the Hologic Aptima® and Panther Fusion® SARS-CoV-2 assays demonstrated >85% positive percent agreement between neat testing and pool sizes ≤5:1, satisfying FDA recommendation. Discordant results between neat and pooled testing were more frequent for positive samples with CT>35. Fusion® CT (cycle threshold) values for pooled samples increased as expected for pool sizes of 5:1 (CT increase of 1.92 - 2.41) and 10:1 (CT increase of 3.03 - 3.29). The Fusion® assay demonstrated lower LLOD than the Aptima® assay for pooled testing (956 vs 1503 cp/mL, pool size of 5:1). Lowering the cut-off threshold of the Aptima® assay from 560 kRLU (manufacturer's setting) to 350 kRLU improved the assay sensitivity to that of the Fusion® assay for pooled testing. Both Hologic's SARS-CoV-2 assays met the FDA recommended guidelines for percent positive agreement (>85%) for pool sizes ≤5:1. Automated pooling increased test throughput and enabled automated sample tracking while requiring less labor. The Fusion® SARS-CoV-2 assay, which demonstrated a lower LLOD, may be more appropriate for surveillance testing.</p>
Etiology of upper respiratory tract infection in outpatients before and during the SARS-CoV-2 pandemic
<p>Illumina MiSeq Viral reads after quality passing and detection in zipped FASTQ format. Files are named by Patient's codes and whether DNA or RNA workflow is used for sample preparation.</p>
Novel trehalose-based excipients for stabilizing nebulized anti-SARS-CoV-2 antibody
<p>Raw data associated to the study entitled "Novel trehalose-based excipients for stabilizing nebulized anti-SARS-CoV-2 antibody", including:</p> <p>- Synthesis development: experimental procedures and characterization</p> <p>- DLS data for CxxTreSuc excipients</p> <p>- Cytotoxicity</p>
SARS-CoV-2 RBD data along with ESM embeddings
<p>This data is published along with the paper "Biophysical principles predict fitness of SARS-CoV-2 variants" and the <a href="https://github.com/Dianzhuo-Wang/COVID19-Biophysical-Model">code</a>. </p> <p>Description:</p> <p><strong>rbd_df.csv</strong> : RBD sequences filtered from GISAID data, along with occurence time, up untill May 2023.</p> <p><strong>unique_mutant_sequence_emb_esm1v_650m.pkl</strong> : esm1v embeddings for unqiue RBDs in rbd_df.csv</p> <p><strong>df_Desai_15loci_complete.csv</strong>: esm1v embeddings for the Desai combinatoric dataset</p> <p>If you use the code or predictions please consider citing:</p> <pre><code>@article{ doi:10.1073/pnas.2314518121, author = {Dianzhuo Wang and Marian Huot and Vaibhav Mohanty and Eugene I. Shakhnovich }, title = {Biophysical principles predict fitness of SARS-CoV-2 variants}, journal = {Proceedings of the National Academy of Sciences}, volume = {121}, number = {23}, pages = {e2314518121}, year = {2024}, doi = {10.1073/pnas.2314518121}, URL = {https://www.pnas.org/doi/abs/10.1073/pnas.2314518121}, eprint = {https://www.pnas.org/doi/pdf/10.1073/pnas.2314518121}, }</code></pre>
Supplemental Figures - Detection of SARS-CoV-2-Specific Secretory IgA and Neutralizing Antibodies in the Nasal Secretions of Exposed Seronegative Individuals
<p>Figure S1: Flow diagram of exposed seronegative cohort.</p> <p>Figure S2: SARS-CoV-2-specific neutralization activity at Days 1 and 8 relative to enrollment in exposed seronegative nasal SIgA positive and infected participants. NPS SARS-CoV-2-specific neutralization activity is shown for exposed seronegative and infected participants with normalized OD490 at Day 1 and Day 8, respectively. Sample sizes (N) are indicated in parentheses. Wilcoxon signed-rank tests were used to determine if the median SARS-CoV-2 nasal SIgA neutralization activity differed significantly. A two-tailed p < 0.05 was considered significant.</p>
Long-term wastewater monitoring of SARS-CoV-2 viral loads and variants at the major international passenger hub Amsterdam Schiphol Airport: a valuable addition to COVID-19 surveillance
<p>Datasets used for the manuscript: <em>Long-term wastewater monitoring of SARS-CoV-2 viral loads and variants at the major international passenger hub Amsterdam Schiphol Airport: a valuable addition to COVID-19 surveillance</em></p> <p><em>pandemic_daily_passenger_counts.tsv</em>: An overview of daily passenger arrival counts at Amsterdam Schiphol Airport per continent of origin during the study period 16-02-2020 - 04-09-2022</p> <p><em>pre-pandemic_daily_passenger_averages.tsv: </em>An overview of mean daily passenger arrival counts at Amsterdam Schiphol Airport in the pre-pandemic period 2017-2019.</p> <p><em>viral_load_data.tsv: </em>Sample metadata (sample identifier, sampling date, flow, average # particles per ml, and flow-corrected viral-load) for samples taken at the wastewater treatment plant of Amsterdam Schiphol Airport.</p> <p><em>wastewater_variant_frequencies.tsv: </em>SARS-CoV-2 lineage estimates in samples taken at the wastewater treatment plant of Amsterdam Schiphol Airport, analyzed using whole-genome tiled amplicon sequencing.</p> <p> </p> <p> </p> <p> </p>
Oral microbiota composition and its relationship with epidemiological, clinical and microbiological variables of SARS-CoV-2 infection in children and adults under strict home confinement in Barcelona, Spain
Open the record for dataset details and reuse information.
Figure 3 in Potential histopathological and immunological effects of SARS-CoV-2 on the liver
Figure 3. Immune-mediated liver injury adapted from Spearman et al. (2021).
Figure 4 in Potential histopathological and immunological effects of SARS-CoV-2 on the liver
Figure 4. The major liver histological features adapted from Díaz et al. (2020).
Figure 1 in Potential histopathological and immunological effects of SARS-CoV-2 on the liver
Figure 1. SARS-CoV-2 structure and life cycle adapted from Zhong et al. (2020).
The early evolution of the BA.2.86 VOI sheds light on the origins of highly divergent SARS-CoV-2 lineages
<p><span>Alignments corresponding to the D1-D4.1 datasets. Five different datasets (named here D1, D2, D3, D4 and D4.4) comprising complete SARS-CoV-2<strong> </strong>genomes (low-coverage excluded, Ns ≤ 5%) were compiled from assemblies retrieved from GISAID. Genomes within these datasets were collated at different time points, aligned and sequentially subsampled under a phylogenetically-informed approach. </span></p>
Data set for "SARS-CoV-2 introductions to the island of Ireland: a phylogenetic and geospatiotemporal study of infection dynamics"
<p>Please see README.txt for detailed information about the contents of this data repository.</p>
Intra-host single-genome sequences of SARS-CoV-2 spike from people without HIV and people living with HIV
<p>Supporting data and code for the study "Rapid Intra-host Diversification and Evolution of SARS-CoV-2 in Advanced HIV Infection" (2024) by Ko & Radecki et al.</p> <p>See <a href="https://github.com/niaid/UMI-pacbio-pipeline/releases/tag/SC2-HIV-demo">the GitHub</a> for demo of SGS and haplotype generation using example data from the study.</p> <p>Raw sequencing data are deposited at <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1055920">PRJNA1055920</a>.</p> <p>READMEs are available within individual folders and releases.</p>
Bioinformática para el análisis de SARS-CoV-2 para principiantes / Bioinformatics for SARS-CoV-2 analysis for beginners
<ul> <li>Data for the e-learning tutorial <a href="https://github.com/cabana-online/Course_SARS_CoV-2">Bioinformatics for SARS-CoV-2 analysis for beginners</a></li> <li>Datos para el tutorial e-learning <a href="Bioinformática para el análisis de SARS-CoV-2 para principiantes" target="_blank" rel="noopener">Bioinformática para el análisis de SARS-CoV-2 para principiantes</a></li> </ul>
Random samples of SARS-CoV-2 sequences for Germany
<p>Mutation profiles and meta data for SARS-CoV-2 lineages in Germany submitted via DESH to the Robert Koch Institute (<a href="https://github.com/robert-koch-institut/SARS-CoV-2-Sequenzdaten_aus_Deutschland">https://github.com/robert-koch-institut/SARS-CoV-2-Sequenzdaten_aus_Deutschland</a>), containing a subset from https://doi.org/10.5281/zenodo.12806236.<br>For this CSV file, the raw data (fasta) was processed with covSonar (https://github.com/rki-mf1/covsonar) and the resulting database queried for random samples (RKI_STICHPROBE, DESH_STICHPROBE) and with dates ranging from 2021-01-01 until 2024-07-24.</p>
Evolutionary trends in the compositional structure of the SARS-CoV-2 genome during the pandemic
<p>By computing a measure of compositional genome structure in random datasets of coronavirus genomes, we observed a long-term decreasing trend in this measure, accompanied by an increasing evolutionary rate.</p> <p> </p>
Docking data for "The evolution of the SARS-CoV-2 spike protein for differential usage of the host transmembrane serine proteases entry pathway"
<p><br>The dataset includes predicted complexes of the SARS-CoV-2 Spike protein (specifically at the S2' cleavage site) with Hepsin and TMPRSS2 proteins. It contains data on three variants: Wuhan, Delta, and Omicron BA.1.</p> <p><strong>Compressed folders:</strong></p> <p>-357596-DeltaHepsin.tgz</p> <p>-357597-DeltaTMPRSS2.tgz</p> <p>-360039-WuhanHepsin.tgz</p> <p>-360042-TMPRSSWuhan.tgz</p> <p>-392981-TMPRSS-BA1_all.tgz</p> <p>-392982-Hepsin-BA-all.tgz</p> <p><strong>Each compressed folder contains the following:</strong></p> <p>-Initial structures in pdb format</p> <p>-Output complexes in pdb format</p> <p>-Clusters in pdb format</p> <p>-Protocols</p> <p>-Parameters</p> <p>-Scoring files</p> <p> </p> <p><strong>Protein-protein docking </strong><br>Molecular docking between the SARS-CoV-2 S protein of Wuhan, Delta (PDB: 7W92, [DOI: 10.1038/s41467-022-28528-w]), and BA.1 (PDB: 7XO5, [DOI: 10.1038/s41422-022-00672-4]) and the human proteases TMPRSS2 (PDB: 8HD8, [DOI: 10.1038/s41467-023-42527-5]) and Hepsin (PDB: 1Z8G, [DOI: 10.1042/BJ20041955]) was performed using the HADDOCK v2.5-2024.03 webserver ([DOI: 10.1021/ja026939x], [DOI: 10.1016/j.jmb.2015.09.014]). Missing loops in the protein structures were reconstructed using Modeller v10.5 ([DOI: 10.1006/jmbi.1993.1626]). Every heteroatom was removed from the reference structures. The relaxed atomistic coordinates for each S protein variant were derived via all-atom molecular dynamics (MD) simulations. These simulations were performed using AMBER22 with the FF19SB force fields and the pmemd.cuda module for enhanced performance ([DOI: 10.1021/acs.jcim.3c01153], [DOI: 10.1021/jz501780a], [DOI:10.1021/ct400314y]). For the Wuhan variant the S protein was retrieved from our previous modeling study [DOI: 10.1039/D0NR03969A] where for Delta and BA.1, ecah S protein was placed in a dodecahedral box, extending 20 Å beyond the solute in every cartesian direction, and solvated with the four-site OPC water model ([DOI: 10.1021/jz501780a]). The systems were neutralized with counterions, specifically one Cl− ion for the Delta variant and three Cl- ions for the BA.1 variant. To remove local clashes, a geometric optimization was performed using the steepest descent algorithm for 5000 cycles. The MD equilibration process consisted of several stages. First, temperature equilibration in the NVT ensemble was performed by gradually increasing the temperature through steps of 150, 200, 250, 300, and finally 310 K, each lasting 200 ps. During this phase, position restraints were applied to the heavy atoms of the proteins, with progressively decreasing spring constants of 5.0, 4.0, 3.0, and 1.0 kcal mol−1 Å−2, facilitating gradual relaxation. This was followed by a 1 ns equilibration at 310 K in the NPT ensemble without restraints. For production MD, the simulations were run in the NPT ensemble with periodic boundary conditions and Particle Mesh Ewald (PME) method ([DOI: 10.1063/5.0040966], [DOI: 10.1021/ct9001015]) using a grid spacing of 1.0 Å for long-range electrostatics. Non-bonded interactions were modeled with a Lennard-Jones potential using a 9Å cutoff. Temperature control was maintained using Langevin dynamics ([DOI: 10.1021/ct800573m]) with a collision frequency of 4.0 ps−1, and pressure control was managed by the Monte Carlo barostat ([DOI: 10.1016/j.cplett.2003.12.039]) with a 2.0 ps relaxation time at 1 bar. Bond constraints on hydrogen atoms were applied using the SHAKE algorithm ([DOI: 10.1016/0021-9991(77)90098-5]), and the hydrogen mass repartitioning scheme was applied via ParmEd ([DOI: 10.1371/journal.pcbi.1005659]), enabling a 4 fs integration time step ([DOI: 10.1021/ct5010406]). Each protein complex was simulated for a total of 20 ns. For the Wuhan variant, the 3D coordinates were retrieved from [DOI: 10.5281/zenodo.3817446].<br>The active interaction region on the spike protein was defined as the cleavage site (residues P809-R815). For TMPRSS2 and Hepsin, the active sites were defined based on their catalytic residues: H296, D345, D435, S441, S460, and G462 for TMPRSS2, and H203, D257, D347, A348, and S353 for Hepsin. These specific regions were selected to guide the docking process and maximize biologically relevant interactions. Docking clusters were analyzed by selecting those with the lowest interaction energies for further structural analysis. To evaluate binding accuracy, native contacts between the S protein and proteases were computed using the contact map analysis based on the OV+rCSU method ([DOI: 10.12693/APhysPolA.145.S9, 10.1021/acs.jctc.6b00986]), which allows for a precise identification of critical stabilizing interactions, both specific and non-specifics. High-frequency contacts, defined as those appearing in over 70% of the generated models, were highlighted as key determinants of protein-protein recognition, providing insight into the most stable and consistent interactions across docking configurations.</p>
Supplementary Data for Ongoing Global and Regional Adaptive Evolution of SARS-CoV-2
<p>This repository includes:<br> The global tree in Newick format: global_2_13_21.main.tre<br> The global, ultrametric tree in Newick format: global_2_13_21.ultra.tre<br> All metadata used in this study including the GISAID acknowledgements: metadata.tgz<br> Tables S1-S4 in plain text format.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.