Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,489

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,489 results for “SARS-CoV-2”

Learn how ShareScore rates datasets ↗
zenodo52/100

Transcriptomic response of human cells to SARS-CoV-2, RSV and H1N1 (STAR + StringTie)

<p>These data represent results from:</p> <ol> <li>Processing reads from 20 experiments (part of GSE147507) by following a standard approach, which includes using STAR to align the reads to GRCh38 and StringTie to calculate the (raw) counts per experiment. These results depict the transcriptomic response&nbsp;of human cells to SARS-CoV-2, RSV and H1N1, and enrichment analyses based on genes differentially expressed in&nbsp;SARS-CoV-2 but not in RSV or H1N1. (Authors: V.A.-P., M.G.F. and A.G.)</li> <li>Aligning to SARS-CoV-2 and quantifying reads&nbsp;by using HISAT2 and StringTie. (Author: C.R.-A.)</li> </ol> <p>Disclaimer: These results were obtained during the virtual BioHackathon 2020. As such, they&nbsp;are subject to ongoing research and have thus NOT yet undergone any scientific peer-review. That is, none of the contents can be considered to be free of errors and must be taken with caution!</p>

opencc-zeroApr 2020View details →
zenodo52/100

Data for At-home testing to characterize SARS-CoV-2 seroprevalence among children and adolescents

<div> <div>This repository contains the data used to reproduce *At-home testing to characterize SARS-CoV-2 seroprevalence among children and adolescents* by Ahmed et al.</div> </div>

opencc-by-4.0May 2024View details →
zenodo48/100

Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention

<p>Single cell RNA seq datasets used for analysis in the&nbsp;Bulk and single-cell gene expression profiling of SARS-CoV-2 infected human cell lines identifies molecular targets for therapeutic intervention</p>

opencc-by-4.0Sep 2020View details →
zenodo48/100

UShER performance statistics, SARS-CoV-2 daily builds 2021-2023

<p>For each day from 2021-01-07 through 2023-08-01 on which the daily build update of the UShER tree of SARS-CoV-2 genomes completed, the number of new sequences added to the tree, the number of sequences in the updated tree, the number of parallel usher jobs (original usher through 2022-04-27, usher-sampled starting 2022-04-29), the number of CPU cores per usher job, and approximate runtime of the usher batch in hours are listed.&nbsp; The number of sequences in the updated tree is "n/a" for most dates prior to 2021-03-11 because before that point, daily updates were for the public-sequence-only tree and the comprehensive GISAID and public sequence tree was updated only occasionally.&nbsp; On and after 2021-03-11, the comprehensive tree and public tree were updated daily.&nbsp; The runtime figures are approximate because they are calculated by subtracting the file modification date of the VCF input to usher from the file modification date of the MAT output of usher.&nbsp; On most days, that was a good proxy for usher runtime, but occasionally there was a crash that required debugging and/or restart, and the "runtime" includes those delays.</p>

opencc-by-sa-4.0Nov 2023View details →
zenodo48/100

Public sequence accessions from INSDC, COG-UK and CNCB and EPI_SET from GISAID for SARS-CoV-2 genome sequences in 2023-08-01 UShER tree

<p>Genome sequences and metadata for the accessions in the .tsv.gz (gzip-compressed tab-separated text) files are freely available from their corresponding sources:</p><ul><li>insdc.accessionNameDate.tsv.gz: INSDC (GenBank, ENA, DDBJ) sequences and metadata may be downloaded using NCBI Datasets: https://www.ncbi.nlm.nih.gov/datasets/taxonomy/2697049/ (7,361,734 accessions used on 2023-08-01)</li><li>cog.accessionNameDate.tsv.gz: COG-UK sequences and metadata may be downloaded from https://cog-uk.s3.climb.ac.uk/phylogenetics/latest (as of publication); most COG-UK sequences have been submitted to ENA and are available from INSDC/NCBI Datasets as well. &nbsp;(724,978 accessions used on 2023-08-01)</li><li>cncb.accessionNameDate.tsv.gz: Sequences and metadata from several databases at the China National Center for Bioinformation (CNCB) may be downloaded from GenBase: https://ngdc.cncb.ac.cn/genbase/ (26,604 accessions used on 2023-08-01)</li></ul><p>GISAID data are subject to restrictions on sharing described in https://gisaid.org/terms-of-use/. &nbsp;Genome sequences and metadata are available to registered GISAID users as part of EPI_SET_231106ax at https://doi.org/10.55876/gis8.231106ax (7,718,061 accessions used on 2023-08-01).</p>

opencc-by-sa-4.0Nov 2023View details →
zenodo48/100

Supplementary Datasets for the publication "Increased Susceptibility of Rousettus aegyptiacus Bats to Respiratory SARS-CoV-2 Challenge Despite Its Distinct Tropism for Gut Epithelia in Bats"

<p>Increasing evidence suggests bats are the ancestral hosts of the majority of coronaviruses. In gen-eral, coronaviruses primarily target the gastrointestinal system, while some strains, especially Be-tacoronaviruses with the most relevant representatives SARS-CoV, MERS-CoV, and SARS-CoV-2, also cause severe respiratory disease in humans and other mammals. We previously reported the susceptibility of Rousettus aegyptiacus (Egyptian fruit bats) to intranasal SARS-CoV-2 infection. Here, we compared their permissiveness to an oral infection versus respiratory challenge (in-tranasal or orotracheal) by assessing virus shedding, host immune responses, tissue-specific pa-thology, and physiological parameters. While respiratory challenge with a moderate infection dose of 1 &times; 104 TCID50 caused a systemic infection with oral and nasal shedding of replica-tion-competent virus, the oral challenge only induced nasal shedding of low levels of viral RNA. Even after a challenge with a higher infection dose of 1 &times; 106 TCID50, no replication-competent vi-rus was detectable in any of the samples of the orally challenged bats. We postulate that SARS-CoV-2 is inactivated by HCl and digested by pepsin in the stomach of R. aegyptiacus, thereby decreasing the efficiency of an oral infection. Therefore, fecal shedding of RNA seems to depend on systemic dissemination upon respiratory infection. These findings may influence our general understanding of the pathophysiology of coronavirus infections in bats.</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Observatorium serologischer Studien zu SARS-CoV-2 in Deutschland

<p>Die seit 2019 auftretende Infektionskrankheit COVID-19, hervorgerufen durch das neuartige SARS-CoV-2-Virus, f&uuml;hrte zu gesundheitspolitischen und gesamtgesellschaftlichen Herausforderungen. Um geeignete Ma&szlig;nahmen zur Eind&auml;mmung der Pandemie ergreifen zu k&ouml;nnen und neue Erkenntnisse &uuml;ber die Pandemie zu gewinnen, gibt es vermehrt Forschungsbedarfe zu COVID-19. Ein Ansatzpunkt hierf&uuml;r sind die gewonnenen Blutproben von infizierten sowie von nicht infizierten Personen, die in Laboren auf Antik&ouml;rper gegen das SARS-CoV-2-Virus getestet und analysiert werden. Sie geben Aufschluss &uuml;ber den Anteil der Bev&ouml;lkerung, der bereits eine Infektion mit SARS-CoV-2 durchgemacht hat, und schlie&szlig;en dabei nicht erkannte Infektionen (Untererfassung) ein.<br>Das Projekt 'Observatorium serologischer Studien zu SARS-CoV-2 in Deutschland' (SERO-OBS Corona) gibt eine &Uuml;bersicht zu Antik&ouml;rper-Studien (sogenannte seroepidemiologische Studien) in Deutschland. Die seroepidemiologischen Studien basieren auf Blutproben von B&uuml;rgerinnen und B&uuml;rgern, die zu unterschiedlichen Zeitpunkten der Pandemie auf Antik&ouml;rper gegen das SARS-CoV-2-Virus getestet wurden. Dabei sollen z. B. folgende Fragen beantwortet werden: Wie ist die H&auml;ufigkeit von SARS-CoV-2-Infektionen in verschiedenen Bev&ouml;lkerungsgruppen? Wie hoch ist der Untererfassungsfaktor, der zeigt, wie viel Mal mehr Infektionen im Vergleich zu den bislang bekannten (gemeldeten) F&auml;llen aufgetreten sind? In dem vorliegenden Projekt werden in Deutschland durchgef&uuml;hrte seroepidemiologische Studien zu SARS-CoV-2 seit dem Fr&uuml;hjahr 2020 &uuml;ber systematische Recherchen in Studienregistern, Literaturdatenbanken einschlie&szlig;lich Vorver&ouml;ffentlichungen sowie Medienberichten fortlaufend identifizier und Studieninformationen sowie Ergebnis&uuml;bersichten verf&uuml;gbar gemacht.</p> <p>Die Ergebnisse des Projektes SERO-OBS-Corona werden auf der Webseite <a href="http://www.rki.de/covid-19-ak-studien">www.rki.de/covid-19-ak-studien</a>, auf Deutsch, sowie der Webseite <a href="http://www.rki.de/covid-19-serostudies-germany">www.rki.de/covid-19-serostudies-germany</a>, auf Englisch, bereitgestellt und regelm&auml;&szlig;ig aktualisiert.</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

SMDP: SARS-CoV-2 Mutation Distribution Profiler for rapid estimation of mutational histories of unusual lineages

<p>Supplementary information relating to the manuscript titled "SMDP: SARS-CoV-2 Mutation Distribution Profiler for rapid estimation of mutational histories of unusual lineages" that has been published on the preprint server arXiv.</p> <ul> <li>PersistentInfectionScore.nb: Mathematica code used to process the data and generate Figure 2</li> <li>PersistentInfectionScore.pdf: pdf version of the above file</li> <li>Supplementary_tables_Harari_et_al_2022.xlsx: raw data from (<a href="https://www.nature.com/articles/s41591-022-01882-4#Sec19">Harari et al. 2022</a>) that was used to generate mutation distributions</li> </ul>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Genome-wide structure and function modeling of SARS-COV-2

<p>Homology models and function annotation for all proteins in the SARS-CoV-2 genome. For a description of each file, follow <a href="https://zhanglab.ccmb.med.umich.edu/COVID-19/">this link</a>.&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

All-atom 500-nano seconds Molecular Dynamics Simulations of SARS-CoV-2 Spike Receptor-binding Domain bound with ACE2

<p>Data includes all of the trajectories (1000) of classical all-atom molecular dynamics (MD) simulations of of SARS-CoV2 Spike Protein/ACE2 complex (PDB ID: 6M0J). In order to decrease the size of the file only protein rajectories were provided.&nbsp;&nbsp;Simulation has been performed with Desmond.&nbsp; Protein was placed in the cubic boxes with explicit TIP3P water models that have 10.0 &Aring; thickness from surfaces of protein. The system is&nbsp;neutralized by adding counter ions, and salt solution of 0.15M NaCl was also used to adjust the concentration of the systems. The long-range electrostatic interactions were calculated by the particle mesh Ewald method. A cutoff radius of 9.0 &Aring; was used for both van der Waals and Coulombic interactions. The temperature was set as 310K initially, and Nose&ndash;Hoover thermostat was used for adjustment. Martyna&ndash;Tobias&ndash;Klein protocol was employed to control the pressure, which was set at 1.01325 bar. The time-step was assigned as 2.0 fs. The default values were used for minimization and equilibration steps, and finally 500 nano-seconds (ns) production run was performed for the simulation.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Dynamics of SARS-CoV-2 spike protein in open and closed states and identification of key structural perturbations upon mutations

<p>The SARS-Cov-2 spike protein resides on the exterior surface of the coronavirus, and therefore, acts as the first point of contact that mediates cell attachment and fusion. &nbsp;During this process, it undergoes dramatic conformational changes upon host receptor binding. We are leveraging high-performance computing to identify these structural perturbations in wildtype and mutant spike protein models. The files contain structures from molecular dynamics simulations of closed SARS-Cov-2 spike protein embedded in POPC membrane.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Data and code for the analysis in "Assessing the impact of non-pharmaceutical interventions on SARS-CoV-2 transmission in Switzerland"

<p>Data and code used for the analysis in <em>Assessing the impact of non-pharmaceutical interventions on SARS-CoV-2 transmission in Switzerland</em> (Lemaitre et al., Swiss Medial Weekly 2020).</p>

opencc-by-4.0May 2020View details →
Figshare44/100

SARS-CoV-2 main protease 3D print model

<p>A 3D model for printing&nbsp;SARS-CoV-2 main protease from our paper on FAIR sharing molecular visualization experiences.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

SARS-CoV-2 NSP13; A Target Enabling Package

<p>To contribute towards the development of novel anti-viral therapeutics targeting the current and future emerging coronavirus threats, the Gileadi lab at the University of Oxford, together with the XChem team at Diamond Light Source, have teamed up to perform a crystallographic fragment screen against SARS-CoV-2 NSP13 helicase. &nbsp;NSP13 is believed to act in concert with the replication-transcription complex (NSP7/NSP8<sub>2</sub>/NSP12), possibly being involved in either disrupting downstream RNA secondary structures or template switching, and plays an essential role in the life cycle of SARS-CoV-2.</p> <p>This TEP includes expression clones and methods for producing the full length NSP13, and fluorescence-based activity assays suitable for compound screening. We provide a crystallisation system that produces reproducible crystals that diffract to high resolution, and have performed a crystallographic fragment screen revealing 63 fragment hits across 51 datasets. The fragment hits include several hits in pockets predicted to be of functional importance, including the nucleotide and nucleic acid binding sites, opening the way to development of novel antiviral agents.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

SARS-CoV-2 Nidoviral RNA Uridylate‐Specific Endoribonuclease (NSP15); A Target Enabling Package

<p>The non-structural protein 15 (NSP15, NendoU<sup>SARS-CoV-2</sup>) from severe acute respiratory syndrome 2 virus (SARS-CoV-2) is an uridylate-specific endoribonuclease, likely responsible in the viral immune evasion mechanism. This TEP provides a set of reagents for further interrogation of the molecular function of NSP15. We have established a purification protocol for the active protein for biochemical and structural studies. Moreover, we have crystallised the protein and performed a crystallographic fragment screen which yielded several hits. Data generated here will be used for the development of enzyme inhibitors that would illuminate the biological role of the gene product, and eventually point the way to new antiviral therapies.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

DATA ANALYSIS - SARS-COV-2 ( Del69-70 VARIANT ) – NEW UK MUTANTS

<p>The data for S - genome sequence analysis known as Del69-70 is under variant of concern ( VOC ) . It is also termed as variant of investigation ( VUI ) . The data for VUI is statistically analysed by datewise and regionwise . The software used for data analysis is CURVE FINDER V.1.4 . The reproducibility of correlation and standard error is reported here for analysis of scattered data an attempt to study the Rational Fit and Harris Fit .</p>

opencc-by-4.0Feb 2021View details →
zenodo44/100

Identifying the interplay between protective measures and settings on the SARS-CoV-2 transmission using a Bayesian network [Dataset]

<p>data07B.csv: dataset for the study of the SARS-CoV-2 transmission.</p> <p>CPTNetica.txt: conditional probabilities tables of each variable given through Netica once the BN obtained in R code is loaded.</p> <p>code01.R: code to learn structure and parameters of the SARS-CoV-2 BN model.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Dataset for: Pre-pandemic artificial MERS analog of polyfunctional SARS-CoV-2 S1/S2 furin cleavage site domain is unique among spike proteins of genus Betacoronavirus

<table> <tbody> <tr> <th>&nbsp;</th> <td> <div> <h3><strong>Data File Descriptions and Methods</strong></h3> <ol> <li><strong>Data file 1 [betacov_matching_IPR042578.fasta]</strong>: Representative set of 2,465 betacoronavirus S protein overlapping homologous superfamily sequences retrieved in fasta format on 4 December 2022 from the InterPro repository at https://www.ebi.ac.uk/interpro/entry/InterPro/IPR042578/.<br><br></li> <li><strong>Data File 2 [betacov_matching_IPR042578_motif.fasta]</strong>: With Data File 1 as input, extracted 98,122 furin cleavage site (FCS) output motifs of 20 amino acids length, including overlapping and redundant sequences, produced with the FindFur algorithm with preset parameters as described by (Gu, 2020). FindFur as used was deposited on 15 December 2020 at the GitHub software repository at https://github.com/chwisteeng/FindFur.<br><br></li> <li><strong>Data File 3 [table_s1s2_hits_betacov_polyf.pdf]</strong>: Compiled summary table of sequence hits (PDF) of spike S1/S2 domains across genus&nbsp;<em>Betacoronavirus. </em>The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences.<em> </em>Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS.<br><br></li> <li> <p><strong>Data File 4 [table_s1s2_hits_betacov_polyf.xlsx]</strong>: Compiled summary table of sequence hits (MS Excel) of spike S1/S2 domains across genus&nbsp;<em>Betacoronavirus. </em>The compiled table of hits removed from Data File 2 sequences corresponding to spike protein fragments (incomplete length spike proteins as deposited at GenBank) and duplicates (redundant parts identically overlapping within the 20 amino acids motif windows), and then selected one sequence representative for multiple but identical sequences.<em> </em>Collection dates and geographical locations were retrieved from the NCBI Genbank protein database at https://www.ncbi.nlm.nih.gov/protein/. For SARS-CoV-2 spike variants, these data were also cross-validated with the SARS-CoV-2 lineage mutation tracker (Gangavarapu, 2023) available at https://outbreak.info which was based on extensive sequencing data from the global GISAID initiative (https://gisaid.org/). Solid lines (-) depict pat7 NLS, asterisks (*) O-glycosites, and circumflex (^) symbols FCS.<br><br></p> </li> <li> <p><strong>Data File 5 [betacov_s1s2_nls_pat7_furin_psort.txt]:&nbsp;</strong>Nuclear localization signal (NLS) detection output for 5 representative betacoronavirus spike sequence domains, including the positive hits for pat7 in SARS-CoV-2 and for MERS-MA30 CoV. NLS predictions used the PSORT algorithm available as a webservice at https://wolfpsort.hgc.jp/ which is based on the work of Nakai and Horton (Nakai and Horton, 1999). Numbering refers to Data File 3 and Data File 4.<br><br></p> </li> <li> <p><strong>Data File 6 [betacov_s1s2_oglyc_netogly.txt]:&nbsp;</strong>Detection output for 5 representative betacoronavirus spike sequence domains tested for Thr/Ser O-glycosite residue pairs with the standard prediction software NetOGlyc4.0 (Steentoft et al., 2013) as available at https://services.healthtech.dtu.dk/services/NetOGlyc-4.0/. Positive hits have scores above 0.5. Numbering refers to Data File 3 and Data File 4.<br><br></p> </li> <li> <p><strong>Data File 7 [betacov_s1s2_nls_pat7_furin_blastp.txt]</strong>: Comprehensive sequence database searches using were performed using the NCBI protein BLAST (blastp) algorithm with webservice available at https://blast.ncbi.nlm.nih.gov/Blast.cgi?PAGE=Proteins. The following blastp search parameters and settings were used: Word size=2; Expect value=200000; Hitlist size=500; Gapcosts=9,1; Matrix=PAM30; Filter string=F; Genetic Code=1;Window Size=40; Threshold=11; Composition-based stats=0; Database Posted date=Jan 19, 2023 2:59 AM; Number of letters=17,117,563; Number of sequences=10,766; Entrez query: Includes: Betacoronavirus (taxid:694002); Excludes: SARS-CoV-2 (taxid:2697049). The six polyfunctional input query consensus motif sequences were TXXPR(K/H/R)XRSX and TXXPRX(K/H/R)RSX.</p> </li> </ol> <h3><strong>References</strong></h3> <p>Gu, C., 2020. FindFur: A Tool for Predicting Furin Cleavage Sites of Viral Envelope Substrates. Master&rsquo;s Thesis, San Jose State University, CA, USA. doi: <a href="https://doi.org/10.31979/etd.4ahv-9jya">10.31979/etd.4ahv-9jya</a>&nbsp;</p> <p>Gangavarapu K, Latif AA, Mullen JL, Alkuzweny M, Hufbauer E, Tsueng G, Haag E, Zeller M, Aceves CM, Zaiets K, Cano M, Zhou X, Qian Z, Sattler R, Matteson NL, Levy JI, Lee RTC, Freitas L, Maurer-Stroh S; GISAID Core and Curation Team; Suchard MA, Wu C, Su AI, Andersen KG, Hughes LD. Outbreak.info genomic reports: scalable and dynamic surveillance of SARS-CoV-2 variants and mutations. Nat Methods. 2023. 20(4):512-522. doi: <a href="https://doi.org/10.1038/s41592-023-01769-3">10.1038/s41592-023-01769-3</a>.</p> <p>Nakai, K., Horton, P., 1999. PSORT: a program for detecting sorting signals in proteins and predicting their subcellular localization. Trends Biochem Sci 24, 34&ndash;36. doi: <a href="https://doi.org/10.1016/s0968-0004(98)01336-x">10.1016/s0968-0004(98)01336-x</a></p> <p>Steentoft, C., Vakhrushev, S.Y., Joshi, H.J., Kong, Y., Vester-Christensen, M.B., Schjoldager, K.T.-B.G., Lavrsen, K., Dabelsteen, S., Pedersen, N.B., Marcos-Silva, L., Gupta, R., Bennett, E.P., Mandel, U., Brunak, S., Wandall, H.H., Levery, S.B., Clausen, H., 2013. Precision mapping of the human O-GalNAc glycoproteome through SimpleCell technology. EMBO J 32, 1478&ndash;1488.&nbsp;doi: <a href="https://doi.org/10.1038/emboj.2013.79">10.1038/emboj.2013.79</a></p> </div> </td> </tr> </tbody> </table>

opencc-by-4.0Jul 2024View details →
zenodo44/100

SARS-CoV-2 wastewater surveillance data and metadata in the Open Data Model format. Part 1: Québec City

<p>SARS-CoV-2 wastewater surveillance data and metadata in the Open Data Model format. Part 1: Qu&eacute;bec City Authors</p> <ul> <li>Therrien, J-D<sup>1</sup></li> <li>Maere, T.<sup>1</sup></li> <li>Sanchez-Quete, F.<sup>2</sup></li> <li>Tsitouras, A.<sup>2</sup></li> <li>Goitom, E.<sup>3</sup></li> <li>Cloutier, F.<sup>4</sup></li> <li>Dufour, D.<sup>4</sup></li> <li>Proulx, F. <sup>4</sup></li> <li>Nicola&iuml;, N.<sup>1</sup></li> <li>Philippe, R.<sup>1</sup></li> <li>Tohidi, M.<sup>1</sup></li> <li>Dorner, S.<sup>3</sup></li> <li>Frigon, D.<sup>2</sup></li> <li>Vanrolleghem, P.A.<sup>1</sup></li> </ul> <p>Affiliations</p> <ul> <li><sup>1</sup> model<em>EAU</em>, D&eacute;partement de g&eacute;nie civil et de g&eacute;nie des eaux, Universit&eacute; Laval</li> <li><sup>2</sup> Microbial Community Engineering Lab (MiCEL), Department of Civil Engineering, McGill University</li> <li><sup>3</sup> Polytechnique Montr&eacute;al</li> <li><sup>4</sup> Ville de Qu&eacute;bec</li> </ul> <p>General Remarks</p> <p>Wastewater-based surveillance of SARS-CoV-2 virus can detect between 1 and 30 infected individuals per 100,000 (including asymptomatic ones) by analyzing the population&#39;s sewage. As such, this method is very attractive since it costs only a fraction of clinical testing (as low as 1%). Human faeces may contain the virus a few days before a person becomes ill. Thus, this approach allows for detection of outbreaks 2-7 days before the increase in reported cases stemming from clinical screening tests (Bibby et al., 2021). Wastewater-based surveillance complements clinical testing by geolocating outbreaks, which may help targeting intensive screening programs. Moreover, it provides a quick indication of whether new public health measures (e.g., masks, social distancing, confinement, and curfew) are effective.</p> <p>Sampling</p> <p>The reported dataset contains open data collected in the province of Qu&eacute;bec as part of the SARS-CoV-2 wastewater-based surveillance program <a href="https://www.centreau.ulaval.ca/en/covid/">CentrEau</a>-COVID. Four of the largest cities in the province (Montr&eacute;al, Laval, Qu&eacute;bec City, and Trois-Rivi&egrave;res), as well as the municipalities of four rural regions (Mauricie, Centre-du-Qu&eacute;bec, Bas-St-Laurent, and Gasp&eacute;sie) participated in the program. The entire dataset includes 31 sampling sites covering approximately half the population of the province of Qu&eacute;bec (population size of 8.5 million). The timeframe covered by the dataset varies for each site. The earliest surveillance program was launched in March 2020, others followed soon after. Samples were collected using various methods, such as 24h composite samples, grab samples, and passive sampling using variations on the Moore swab method (Schang et al., 2020)</p> <p>Analysis</p> <p>Prior to the analysis of the samples for SARS-CoV-2, physiochemical parameters such as total suspended solids (TSS), turbidity, conductivity, ammonium concentration, and pH were measured. The samples were subsequently concentred by filtration using a MEC filter (0.45 um), followed by total RNA extraction using the Qiagen AllPrep PowerViral DNA/RNA Kit (Qiagen, USA) with some modifications (beta-mercaptoethanol concentration raised to 10% and lysis performed at 55 &deg;C for 30 minutes) (Ahmed et al., 2020). SARS-CoV-2 viral RNA was detected by a one-step RT-qPCR. To assess the RNA recovery rate of the procedure, samples were spiked before extraction with a known concentration of Bovine Respiratory Syncytial Virus (BRSV) using the Zoetis INFORCE 3 vaccine (Zoetis, USA). In addition to SARS-CoV-2, samples were assessed for Pepper Mild Mottle Virus (PMMoV), the daily load of which is hypothesized to represent the fecal load contributions to the samples at a given site and time. PCR conditions and primer used to collect viral data are described in the files <code>primers.md</code> and <code>PCR conditions.md</code>.</p> <p>Compilation</p> <p>The measurements on wastewater samples carried out by the participating laboratories of this study are found in the <code>WWMeasure</code> table. The values provided by municipalities come from laboratories accredited by the Centre d&#39;expertise en analyse environnementale du Qu&eacute;bec (CEAEQ), in compliance with the latter&#39;s quality assurance protocols. The COVID-19-related public health data found in the <code>CPHD</code> table were collected from the Institut National de Sant&eacute; Publique du Qu&eacute;bec (INSPQ)&#39;s public reports. Wastewater data taken in-situ at the sampling sites (e.g., the flow at pumping stations or water resource recovery facilities (WRRFs)) are found in the <code>SiteMeasure</code> table and were taken by the institutions responsible for managing the sites. All of the data, stemming from multiple sources, were combined into the <a href="https://github.com/Big-Life-Lab/ODM">Open Data Model (ODM)</a> standard format using the <a href="https://github.com/modelEAU/ODM-Import">ODM-Import python package</a> (see also Structure).</p> <p>Validation</p> <p>Wastewater and sample data were manually assessed for quality by our research collaborators. Data points for which the quality appeared to be uncertain were tagged with the value <code>True</code> in the <code>qualityFlag</code> column. Conversely, data deemed of good quality have a quality flag of <code>False</code>. Data that were not checked have a quality flag of <code>NA</code>. Textual comments describing the issues with the data points in more detail are also included in the dataset using the <code>notes</code> column of the relevant tables. Note that data validation was carried out by the data custodians responsible for each city in the dataset according to available resources. As the project continues and data validation is undertaken on more sections of the dataset, data may be re-analyzed, flagged, or commented as needed. Revisions to the dataset will be reported to the best of our ability.</p> <p>Structure</p> <p>The data contained in this dataset has been structured according to the <a href="https://github.com/Big-Life-Lab/ODM">Open Data Model (ODM) for Wastewater-Based Surveillance</a>. This model provides a standardized dictionary to collect and share data and metadata stemming from wastewater-based surveillance programs. By convention, it splits all data into 10+ thematic tables with each record representing a unique measurement, i.e., long format. For convenience, the <code>wide</code> folder presents the data found in all the other tables in a wide format, i.e., multiple measurements are aligned by <code>timestamp</code>, with each column representing a different parameter.</p> <p>Acknowledgements</p> <p>The authors would like to acknowledge that this dataset was collected thanks to the financial support of the Fonds de Recherche du Qu&eacute;bec, the Molson Foundation, the Trottier Family Foundation, CentrEau and NSERC. The authors would also like to acknowledge the efforts of Douglas Manuel (Ottawa Hospital) and Howard Swerdfeger (Public Health Agency of Canada) for their original idea for the Open Data Model and continued development.</p> <p>References</p> <ol> <li> <p>Ahmed, W., Bertsch, P.M., Bivins, A., Bibby, K., Farkas, K., Gathercole, A., Haramoto, E., Gyawali, P., Korajkic, A., McMinn, B.R., Mueller, J.F., Simpson, S.L., Smith, W.J.M., Symonds, E.M., Thomas, K. v., Verhagen, R., Kitajima, M., 2020. Comparison of virus concentration methods for the RT-qPCR-based recovery of murine hepatitis virus, a surrogate for SARS-CoV-2 from untreated wastewater. Science of the Total Environment 739. <a href="https://doi.org/10.1016/j.scitotenv.2020.139960">https://doi.org/10.1016/j.scitotenv.2020.139960</a></p> </li> <li> <p>Bibby, K., Bivins, A., Wu, Z., North, D., 2021. Making waves: Plausible lead time for wastewater based epidemiology as an early warning system for COVID-19. Water Research 202, 117438. <a href="https://doi.org/10.1016/j.watres.2021.117438">https://doi.org/10.1016/j.watres.2021.117438</a></p> </li> <li> <p>Schang, C., Crosbie, N., Nolan, M., Poon, R., Wang, M., Jex, A., Scales, P., Schmidt, J., Thorley, B.R., Henry, R., Kolotelo, P., Langeveld, J., Schilperoort, R., Shi, B., Einsiedel, S., Thomas, M., Black, J., Wilson, S., McCarthy, D.T., 2020. Passive sampling of viruses for wastewater-based epidemiology: a case-study of SARS-CoV-2 [WWW Document]. URL <a href="https://www.researchgate.net/publication/347103410\_Passive\_sampling\_of\_viruses\_for\_wastewater-based\_epidemiology\_a\_case-study\_of\_SARS-CoV-2?channel=doi&amp;linkId=5fd800f392851c13fe892393&amp;showFulltext=true">https://www.researchgate.net/publication/347103410\_Passive\_sampling\_of\_viruses\_for\_wastewater-based\_epidemiology\_a\_case-study\_of\_SARS-CoV-2?channel=doi&amp;linkId=5fd800f392851c13fe892393&amp;showFulltext=true</a> (accessed 1.18.21).</p> </li> </ol>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Dry trajectories of SARS-CoV-2 RBD from accelerated molecular dynamics simulation

<p>These are supplementary files to the preprint/paper &quot;SARS-CoV-2 spike protein unlikely to bind to integrins via the Arg-Gly-Asp (RGD) motif of the Receptor Binding Domain: evidence from structural analysis and microscale accelerated molecular dynamics&quot; (http://dx.doi.org/10.1101/2021.05.24.445335).</p> <p>The attached code in Jupyter notebook can be run after installing the virtual environment using the `environment.yml `</p> <p>The file `data.zip` needs to be extracted to the same path where the notebook is run from</p>

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record