Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,489
datasets available to search
ShareScore release 0.7.1
Dataset results
2,489 results for “SARS-CoV-2”
Figure 2 in Potential histopathological and immunological effects of SARS-CoV-2 on the liver
Figure 2. Potential mechanisms of hepatic injury with SARS-CoV-2 infection adapted from Yang et al. (2020).
Electron microscopy images and morphometric data of SARS-CoV-2 variants in ultrathin plastic sections - Dataset 06 (SARS-CoV-2 Omicron B.1.1.529; BA.2)
<p>Dataset 06 comprises 164 transmission electron microscopy images of extracellular SARS-CoV-2 (isolate Omicron B.1.1.529; BA.2) particles in ultrathin plastic sections (45 nm) through Vero cell cultures. The images were recorded with dimensions of 4112 x 3008 pixels at a pixel size of 0.1641 nm and stored in 16-bit TIF format. It is recommended that an image viewer capable of reading 16-bit images, such as IrfanView, be used to visualize the images. The image files have been size calibrated and can be opened with the correct size calibration using ImageJ or Fiji with the Bioformats importer. A PDF document is provided with the image files, which describes the methods used for the generation of the images. Additionally, an XLSX file is included, offering morphometric particle measurements and the calculated statistical values for their distribution. The dataset was produced as dataset 06 for a comparative morphometric analysis of evolving SARS-CoV-2 variants. Further datasets used for the analysis are available in this repository (see dataset description document).</p>
Electron microscopy images and morphometric data of SARS-CoV-2 variants in ultrathin plastic sections - Dataset 04 (SARS-CoV-2 Beta B.1.351)
<p>Dataset 04 comprises 132 transmission electron microscopy images of extracellular SARS-CoV-2 (isolate Beta B.1.351) particles in ultrathin plastic sections (45 nm) through Vero cell cultures. The images were recorded with dimensions of 4112 x 3008 pixels at a pixel size of 0.1641 nm and stored in 16-bit TIF format. It is recommended that an image viewer capable of reading 16-bit images, such as IrfanView, be used to visualize the images. The image files have been size calibrated and can be opened with the correct size calibration using ImageJ or Fiji with the Bioformats importer. A PDF document is provided with the image files, which describes the methods used for the generation of the images. Additionally, an XLSX file is included, offering morphometric particle measurements and the calculated statistical values for their distribution. The dataset was produced as dataset 04 for a comparative morphometric analysis of evolving SARS-CoV-2 variants. Further datasets used for the analysis are available in this repository (see dataset description document).</p>
Electron microscopy images and morphometric data of SARS-CoV-2 variants in ultrathin plastic sections - Dataset 02 (SARS-CoV-2 Italy-INMI1)
<p>Dataset 02 comprises 154 transmission electron microscopy images of extracellular SARS-CoV-2 (isolate Italy-INMI1) particles in ultrathin plastic sections (45 nm) through Vero cell cultures. The images were recorded with dimensions of 4112 x 3008 pixels at a pixel size of 0.1641 nm and stored in 16-bit TIF format. It is recommended that an image viewer capable of reading 16-bit images, such as IrfanView, be used to visualize the images. The image files have been size calibrated and can be opened with the correct size calibration using ImageJ or Fiji with the Bioformats importer. A PDF document is provided with the image files, which describes the methods used for the generation of the images. Additionally, an XLSX file is included, offering morphometric particle measurements and the calculated statistical values for their distribution. The dataset was produced as dataset 02 for a comparative morphometric analysis of evolving SARS-CoV-2 variants. Further datasets used for the analysis are available in this repository (see dataset description document).</p>
Supplemental data for: Evaluation of SARS-CoV-2 response at the University of North Carolina (UNC) at Charlotte using percent positivity data and viral genomic sequence data.
<p>Supplemental data for:</p> <p>Evaluation of SARS-CoV-2 response at the<br>University of North Carolina (UNC) at Charlotte<br>using percent positivity data and viral genomic<br>sequence data.</p> <p>Submitted to Biocarla 2024</p> <p>https://carla2024.org/portfolios/biocarla/</p> <p>Authors:</p> <p>Daniel Janies 1,2,3 [0000−0002−7890−9906], Shirish Yasa 1,2,3 [0000−0003−3217−4921],<br>Colby T. Ford 1,3,4 [0000−0002−7859−3622] Jannatul Ferdous 2,3 [0000−0003−3053−9616],<br>William Taylor 2,3 [0009−0000−6204−1172], April Harris 2,3 [0009−0009−2557−7926],<br>Sam Kunkleman 2,3 [0000−0002−2309−6418], Juan Bolanos 2,3, Kevin Lambirth 2,3 [0000-0002-6568-543X], Denis<br>Jacob Machado 1,2,3 [0000−0001−9858−4515], Cynthia Gibas 1,2,3 [0000−0002−1288−9543],<br>and Jessica Schlueter 1,2,3 [0000−0002−6490−0580]</p> <p>Affiliations:</p> <p>1) Center for Computational Intelligence to Predict Health and Environmental Risks<br>(CIPHER), University of North Carolina at Charlotte 28223, USA<br>Correspondence to: djanies@charlotte.edu<br>https://cipher.charlotte.edu<br>2) Department of Bioinformatics and Genomics, University of North Carolina at<br>Charlotte 28223, USA https://cci.charlotte.edu/departments/<br>department-of-bioinformatics-and-genomics/<br>3) College of Computing and Informatics, University of North Carolina at Charlotte<br>28223, USA https://cci.charlotte.edu<br>4) School of Data Science, University of North Carolina at Charlotte 28223, USA<br>https://sds.charlotte.edu<br>5) Division of Research, University of North Carolina at Charlotte 28223, USA<br>https://research.charlotte.edu/</p> <p> </p>
Data for manuscript: The Conformational Space of the SARS-CoV-2 Main Protease Active Site Loops is Determined by Ligand Binding and Interprotomer Allostery
<div>The data is provided as a part of the manuscript "<strong>The Conformational Space of the SARS-CoV-2 Main Protease Active Site Loops is Determined by Ligand Binding and Interprotomer Allostery</strong>". This repository includes an archive with folders:</div> <div> </div> <div><strong>md_data </strong></div> <div> <ul> <li>a directory with MD data for all simulation systems considered in the manuscript. Initial and final conformations are provided.</li> </ul> </div> <div> </div> <div><strong>fig_data</strong></div> <div> <ul> <li>a directory with the data underlying all the main text in the manuscript. </li> </ul> </div> <div> </div> <div>Videos S1-S3 are also included.</div>
Metaproteomics reveals age-specific alterations of gut microbiome in hamsters with SARS-CoV-2 infection
<p><span>The gut microbiome's pivotal role in health and disease is well-established. SARS-CoV-2 infection often causes gastrointestinal symptoms and is associated with changes of the microbiome in both human and animal studies. While hamsters serve as important animal models for coronavirus research, there exists a notable void in functional characterization of their microbiomes with metaproteomics. In this study, we present a workflow for analyzing the hamster gut microbiome, including a metagenomics-derived hamster gut microbial protein database and a data-independent acquisition metaproteomics method. Using this workflow, we identified 32419 protein groups from the fecal microbiomes of young and old hamsters infected with SARS-CoV-2 . We showed age-specific changes in the expressions of microbiome functions and host proteins associated with microbiomes, providing further functional insight into the dysbiosis and aberrant cross-talks between the microbiome and host in SARS-CoV-2 infection. Altogether this study established and demonstrated the capability of metaproteomics for the study of hamster microbiomes.<span> </span></span></p>
Linked collectors and determiners for: Detección de patógenos emergentes (SARS-CoV-2) y re-emergentes (Rickettsias) en caninos procedentes del área metropolitana de Bucaramanga.
Natural history specimen data linked to collectors and determiners held within, "Detección de patógenos emergentes (SARS-CoV-2) y re-emergentes (Rickettsias) en caninos procedentes del área metropolitana de Bucaramanga". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/224f87f5-79f8-431b-9fb7-a5bc78e2e666">https://bionomia.net/dataset/224f87f5-79f8-431b-9fb7-a5bc78e2e666</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/224f87f5-79f8-431b-9fb7-a5bc78e2e666">https://gbif.org/dataset/224f87f5-79f8-431b-9fb7-a5bc78e2e666</a>. Formatted as a Frictionless Data package.
Linked collectors and determiners for: Identificación de ectoparásitos para la detección de patógenos emergentes (SARS-CoV-2) y re-emergentes (Rickettsias) en caninos procedentes del área metropolitana de Bucaramanga.
Natural history specimen data linked to collectors and determiners held within, "Identificación de ectoparásitos para la detección de patógenos emergentes (SARS-CoV-2) y re-emergentes (Rickettsias) en caninos procedentes del área metropolitana de Bucaramanga". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/17e160c4-11ea-432e-95af-d65e84ec5953">https://bionomia.net/dataset/17e160c4-11ea-432e-95af-d65e84ec5953</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/17e160c4-11ea-432e-95af-d65e84ec5953">https://gbif.org/dataset/17e160c4-11ea-432e-95af-d65e84ec5953</a>. Formatted as a Frictionless Data package.
Risk of bias assessments for the Cochrane review 'SARS-CoV-2-neutralising monoclonal antibodies for treatment of COVID-19'
<p>Risk of bias assessments and support for judgement with ROB 2 tool for the Cochrane Review: SARS-CoV-2-neutralising monoclonal antibodies for treatment of COVID-19.</p>
Early 3D Evolution of the SARS-CoV-2 proteome -- Supplementary Tables and Models
<p><strong>Evolution of the SARS-CoV-2 proteome in three dimensions (3D) during the first six months of the COVID-19 pandemic</strong></p> <p><a href="https://iqb.rutgers.edu/covid-19_proteome_evolution">https://iqb.rutgers.edu/covid-19_proteome_evolution</a></p> <p> </p> <p><strong>Legends for Supplementary Figures for 29 </strong><strong>SARS-CoV-2 Study Proteins</strong></p> <p><strong>Separate analysis of protein changes was performed for each study protein and complex. Description below applies to all figures.</strong></p> <p><strong>A</strong>: Observed frequencies for all USV substitutions of Native Residue (i.e., amino acid type in the reference protein sequence) changing to Substituted Residue for a given protein/complex. Red boxes enclose conservative substitutions for hydrophobic, uncharged polar, positively charged, and negatively charged amino acids, respectively in order from upper left to lower right. Cysteine, Glycine and Proline are excluded from these groupings.</p> <p><strong>B-D</strong>: Normalized Frequency histograms for ΔΔG<sup>App</sup> calculated for all USVs for a given protein/complex. These were calculated using three methods, which we refer to as hard-hard (B), soft-hard (C), and soft-soft (D), based on the scoring functions used for sidechain rotamer optimization and gradient-based energy minimization respectively (see methods). All energy values described in the text were obtained using the soft-hard method. Overlay of energy histogram with fitted bi-Gaussian curve (solid red line) and fitted single Gaussian curves for subsets of USVs with surface (green), boundary layer (yellow), or core (blue) substitutions. USVs with multiple substitutions were included in single Gaussian fitting when all substitutions mapped to the same region of the study protein. The data used for fitting includes the energies of all unique protein models produced by a given method, excluding extreme outliers with energy values greater than 3 standard deviations away from the central mean.</p> <p><strong>E-G</strong>: USV Count histograms indicate the number of USVs among the full set for a given protein in which each site included a substitution. Sites are separated by burial layer. Substitutions at sites that are absent from the available crystal structures are excluded from the histograms. In most cases, only a single protein is analyzed, and only panel E is included. In the case of complexes, a separate histogram is provided for each protein in the complex: for methyltransferase nsp10-nsp16, E is nsp10 and F is nsp16; for RDRP nsp12-nsp7-nsp8, E is nsp7, F is nsp8, and G is nsp12.</p> <p> </p> <p><strong>Legends for Supplementary Tables for 29 </strong><strong>SARS-CoV-2 Study Proteins</strong></p> <p><strong>Table: USVs</strong>: All identified USVs for a protein/complex. Columns are:</p> <ul> <li>date: Date of first collection of a strain with the USV reported to GISAID</li> <li>gisaid_count: The number of sequences in the GISAID database that include the USV</li> <li>id: The GISAID strain identification for the first collected instance of the USV</li> <li>location: The country in which the first strain including the USV was collected</li> <li>substitutions: All substitutions in the USV, in the form [chain]_[sequence][site][substitution], with multiple substitutions separated by semicolons</li> <li>is_in_PDB: whether a substitution is present in the PDB model used to generate the USV structure, with multiple substitutions separated by semicolons</li> <li>multiple: whether more than one amino acid substitution is present in the USV</li> <li>conservative: whether a substitution is conservative, with multiple substitutions separated by semicolons</li> <li>layer: Identification of the burial layer (surface, boundary, or core) of a substitution in the reference structure, with multiple substitutions separated by semicolons and substitutions absent from the PDB excluded</li> <li>sh_rmsd: The RMSD of the USV to the reference structure when modeled using the soft-hard method</li> <li>sh_ddG: The ΔΔG<sup>App</sup> of the USV when modeled using the soft-hard method</li> <li>hh_rmsd: The RMSD of the USV to the reference structure when modeled using the hard-hard method</li> <li>hh_ddG: The ΔΔG<sup>App</sup> of the USV when modeled using the hard-hard method</li> <li>ss_rmsd: The RMSD of the USV to the reference structure when modeled using the soft-soft method</li> <li>ss_ddG: The ΔΔG<sup>App</sup> of the USV when modeled using the soft-soft method</li> </ul> <p> </p> <p><strong>Table: Substitutions</strong>: All substitutions identified for a protein/complex</p> <ul> <li>chain: The chain identifier of the protein in the PDB file in which the substitution is present</li> <li>site: The residue number at which the substitution is present</li> <li>reference: The one-letter amino acid name of the residue in the reference sequence</li> <li>mutant: The one-letter amino acid name of the residue in a USV</li> <li>conservative: Indication of whether a substitution is conservative</li> <li>in_pdb: whether the substitution site is present in the PDB model used to generate the USV structure</li> <li>layer: Identification of the burial layer (surface, boundary, or core) of a substitution in the reference structure</li> <li>date: date: Date of first collection of a strain with the substitution reported to GISAID</li> <li>location: The country in which the first strain including the substitution was collected</li> <li>gisaid_count: The number of sequences in the GISAID database including the substitution</li> <li>usv_count: The number of identified USVs including the substitution</li> <li>ddG: The soft-hard ΔΔG<sup>App</sup> of the USV that includes only the substitution, left empty if no single-substitution USV was identified with the substitution</li> <li>single: Indication of whether the substitution was present in a single-substitution USV</li> <li>multiple: Indication of whether the substitution was present in a USV with multiple substitutions</li> <li>associates: List of all other substitutions that were identified in a USV that included the substitution</li> <li>strains: List of all USV-representative GISAID strains that included the substitution, with the single-substitution USV strain listed first if one was available</li> </ul> <p> </p> <p><strong>Table: Gaussian Fit Statistics</strong>: Fitted models for the energies of all USVs either together (ALL) or by study protein.</p> <ul> <li>fit: The number of Gaussian curves in the fitted energy model </li> <li>protein: The protein/complex name</li> <li>method: The modeling method used to calculate energy values</li> <li>layer: The subset burial layer (surface, boundary, or core) of USVs for which the energy model was fitted, excluding all USVs with substitutions not in that layer</li> <li>μ<sub>1</sub>: Mean of the first Gaussian in the fitted model</li> <li>σ<sub>1</sub>: Variance of the first Gaussian in the fitted model</li> <li>wt<sub>1</sub>: Weight of the first Gaussian in the fitted model</li> <li>μ<sub>2</sub>: Mean of the second Gaussian in the fitted model</li> <li>σ<sub>2</sub>: Variance of the second Gaussian in the fitted model</li> <li>wt<sub>2</sub>: Weight of the second Gaussian in the fitted model</li> <li>R<sup>2</sup>: R-squared value indicating the goodness of fit</li> </ul> <p> </p> <p><strong>Description of Computed Structural Models </strong><strong>for Unique Sequence Variants for 29 </strong><strong>SARS-CoV-2 Study Proteins.</strong></p> <p><strong>USV Computed Structural Models</strong>. Computed structural models for all amino acid substituted USVs. We are providing the structural models of all study proteins modeled using the soft-hard modeling method (see Methods). Structural models are named according to the GISAID strain identification of the first strain in which the USV was identified, followed by an underscore-separated list of substitutions in the form [chain]_[sequence][site][substitution]. Atomic coordinates for each computed structural model are provided in the legacy Protein Data Bank format used by most molecular graphics software tools (see <a href="https://www.wwpdb.org/documentation/file-format-content/format33/v3.3.html">https://www.wwpdb.org/documentation/file-format-content/format33/v3.3.html</a> for detailed description).</p>
Wastewater monitoring of SARS-CoV-2 from wastewater in Bratislava, Slovakia and other Slovak cities
<p>The statistics for Bratislava and other selected Slovak cities - Košice, Prešov, Žilina, Banská Bystrica, Trenčín, Trnava, Piešťany, Nováky, Poprad (Petržalka is a district of Bratislava) shows the correlation between the positive PCR tests, deaths caused by COVID-19 and detected virus particles from wastewater (based on a wastewater model developed by the Slovak University of Technology, Faculty of Chemical and Food Technology) in a certain period of time in 2020-2021.</p> <p>The statistics were developed as a result of the Project VIR-SCAN – Wastewater Monitoring Data as an Early Warning Tool to alert COVID-19 in the Population, supported by the EOSCsecretariat.eu</p> <p>Project VIR-SCAN – Wastewater Monitoring Data as an Early Warning Tool to alert COVID-19 in the Population“ has received funding from the EOSC Secretariat project. EOSCsecretariat.eu has received funding from the European Union’s Horizon Programme call H2020-INFRAEOSC-05-2018-2019, grant Agreement number 831644.</p>
The CHASING COVID Cohort Study: A national, community-based prospective cohort study of SARS-CoV-2 pandemic outcomes in the USA
<p>The Communities, Households and SARS-CoV-2 Epidemiology (CHASING) COVID Cohort Study is a community-based prospective cohort study launched during the upswing of the USA COVID-19 epidemic. The objectives of the cohort study are to: (1) estimate and evaluate determinants of the incidence of SARS-CoV-2 infection, disease and deaths; (2) assess the impact of the pandemic on psychosocial and economic outcomes and (3) assess the uptake of pandemic mitigation strategies. 6740 people are enrolled in the cohort, including participants from all 50 US states, the District of Columbia, Puerto Rico and Guam. Participants are contacted regularly to complete study assessments, including interviews and dried blood spot specimen collection for serologic testing.</p> <p>Datasets are provided in CSV and sas7bdat (with formatting script) file formats.</p>
The central nervous system's proteogenomic and spatial imprint upon systemic viral infection, like SARS-CoV-2
<p>Data set including image files of histological stainings, immunohistochemistry, MELC, and spatial transcriptomics associated with the study mentioned above.</p>
Relative role of border restrictions, case finding and contact tracing in controlling SARS-CoV-2 in the presence of undetected transmission: a mathematical modelling study
<p>Data and code for publication on <em>Relative role of border restrictions, case finding and contact tracing in controlling SARS-CoV-2 in the presence of undetected transmission: a mathematical modelling study</em></p>
The benefit of augmenting open data with clinical data-warehouse EHR for forecasting SARS-CoV-2 hospitalizations in Bordeaux area, France
<p><strong>Objective</strong></p> <p>The aim of this study was to develop an accurate regional forecast algorithm to predict the number of hospitalized patients and to assess the benefit of the Electronic Health Records (EHR) information to perform those predictions. Materials and Methods Aggregated data from SARS-CoV-2 and weather public database and data warehouse of the Bordeaux hospital were extracted from May 16, 2020, to January 17, 2022. The outcomes were the number of hospitalized patients in the Bordeaux Hospital at 7 and 14 days. We compared the performance of different data sources, feature engineering, and machine learning models.</p> <p><strong>Results </strong></p> <p>During the period of 88 weeks, 2561 hospitalizations due to COVID-19 were recorded at the Bordeaux Hospital. The model achieving the best performance was an elastic-net penalized linear regression using all available data with a median relative error at 7 and 14 days of 0.136 [0.063; 0.223] and 0.198 [0.105; 0.302] hospitalizations, respectively. Electronic health records (EHRs) from the hospital data warehouse improved median relative error at 7 and 14 days by 10.9% and 19.8%, respectively. Graphical evaluation showed remaining forecast error was mainly due to delay in slope shift detection.</p> <p><strong>Discussion </strong></p> <p>Forecast models showed overall good performance both at 7 and 14 days which was improved by the addition of the data from Bordeaux Hospital data warehouse.</p> <p><strong>Conclusions </strong></p> <p>The development of hospital data warehouses might help to get more specific and faster information than traditional surveillance systems, which in turn will help to improve epidemic forecasting at a larger and finer scale.</p>
FASTA consensus sequences obtained using amplicon-based genome sequencing of SARS-CoV-2
<p>Set of 22 FASTA consensus sequences that were produced during routine SARS-CoV-2 sequencing obtained using amplicon-based sequencing (ARTIC protocol). Those sequences were compared to those generated in NASCarD applications.</p>
SARS-CoV-2 Omicron Boosting Induces De Novo B Cell Response in Humans
<p>These are the<strong> processed</strong> BCR repertoire and transcriptomics data described in <a href="https://doi.org/10.1038/s41586-023-06025-4">Alsoussi & Malladi et al., <em>Nature</em>, 2023</a>. The <strong>raw</strong> sequencing data new to this study are available on SRA under BioProject <a href="https://www.ncbi.nlm.nih.gov/sra/?term=PRJNA800176">PRJNA800176</a>. This study also used BCR repertoire data from <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., <em>Nature</em>, 2021</a> (<a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA731610">PRJNA731610</a>), <a href="https://doi.org/10.1016/j.immuni.2021.08.013">Schmitz, Turner & Liu et al., <em>Immunity</em>, 2021</a> (<a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA741267">PRJNA741267</a>), and <a href="https://doi.org/10.1038/s41586-022-04527-1">Kim & Zhou et al., <em>Nature</em>, 2022</a> (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA777934/">PRJNA777934</a>).</p> <p> </p> <p><strong>Code</strong></p> <p>Code along with Docker containers for reproducing the NGS data-based figures and analyses in the published paper can be <a href="https://github.com/julianqz/wustl_published/tree/main/nature_2023">found on GitHub</a>.</p> <p> </p> <p><strong>Metadata</strong></p> <p>File: WU382_alsoussi_et_al_nature_2023_meta.tsv.gz</p> <p>Notes:</p> <ul> <li>181 samples in total, including: <ul> <li>78 new</li> <li>90 from <a href="https://doi.org/10.1038/s41586-022-04527-1">Kim & Zhou et al., <em>Nature</em>, 2022</a></li> <li>8 from <a href="https://doi.org/10.1016/j.immuni.2021.08.013">Schmitz, Turner & Liu et al., <em>Immunity</em>, 2021</a></li> <li>5 from <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., <em>Nature</em>, 2021</a></li> </ul> </li> <li>Participant IDs: 6 participants who were in previous studies and who continued in the new study were referenced by new participant IDs. Correspondence with previous participant IDs is as follows: <ul> <li>382-01 = 368-22</li> <li>382-02 = 368-20</li> <li>382-07 = 368-02a</li> <li>382-08 = 368-04</li> <li>382-13 = 368-01a</li> <li>382-15 = 368-10</li> </ul> </li> <li>Sample collection time was originally recorded in days in the `timepoint` column. Values in parentheses indicate variations in which the BCR data was coded. Timepoints were mainly referenced in weeks in the manuscript, as shown in the `timepoint_ms` column. </li> <li>Pre-3rd dose ("pre-boost") samples were coded `b0` in the `booster_num` column; post-3rd dose ("post-boost") samples were coded `b1`.</li> <li>The `booster_type` column records the 3rd dose ("booster") variant. <ul> <li>`regular` = mRNA-1273 (WA1/2020)</li> <li>`beta_delta` = mRNA-1273.213 (Beta & Delta)</li> <li>`v1.1.529` = mRNA-1273.529 (Omicron)</li> </ul> </li> <li>382-02/07/08 received mRNA-1273; 382-01/13/15 received mRNA-1273.213; 382-53/54/55 received mRNA-1273.529.</li> <li>The `seq_type` column indicates the platform from which sequences originated. <ul> <li>`bulk` = bulk BCR sequencing</li> <li>`tgx` = 10x Genomics single-cell VDJ + 5' gene expression</li> <li>`mab`, `mab_1`, `mab_2`: single-cell sorted mAb synthesis. The suffixes were purely for the convenience of distinguishing originating studies.</li> </ul> </li> </ul> <p>Abbreviations:</p> <ul> <li>LN = lymph node</li> <li>BM = bone marrow</li> <li>PB = plasmablast</li> <li>GC = germinal centre</li> <li>LLPC = long-lived plasma cell</li> <li>NS = no sorting</li> <li>mAb = monoclonal antibody</li> </ul> <p> </p> <p><strong>[Beta & Delta booster] Processed BCR data - heavy chains</strong></p> <p>File: WU382_alsoussi_et_al_nature_2023_betaDelta_bcr_heavy.tsv.gz</p> <p><em>Analysis was based on heavy chain-based clonal inference.</em></p> <p>Notes on columns:</p> <p>The columns largely follow the <a href="https://changeo.readthedocs.io/en/stable/standard.html">AIRR-C Rearrangement format</a>. The main deviation is that CDR3s were used, as opposed to IMGT-defined "junctions". Nonetheless, junction-related columns are included here as some repositories such as <a href="https://gateway.ireceptor.org/login"><em>iReceptor</em></a> use these. Non-standard columns are noted below.</p> <ul> <li>`cell_id`: Only sequences from single-cell samples and synthesized mAbs have cell IDs. 10x sequences follow the format `[donor]_[sample]@[id]`. `NA` for bulk sequences.</li> <li>`sequence_id`: Sequence IDs follow the format `[donor]_[sample]@[id]`.</li> <li>`v_call_genotyped`: V gene annotation reassigned after individualized genotyping by <a href="https://tigger.readthedocs.io/en/stable/">TIgGER</a>.</li> <li>`germline_[vdj]_call`: Clonal consensus germline calls after corresponding clonal consensus sequences were reconstructed via <a href="https://changeo.readthedocs.io/en/stable/methods/germlines.html">`CreateGermlines.py --cloned` from Change-O</a>.</li> <li>`collapse_count`: Number of duplicate IMGT-aligned V(D)J sequences that were collapsed by <a href="https://alakazam.readthedocs.io/en/stable/topics/collapseDuplicates/">`alakazam::collapseDuplicates`</a>.</li> <li>`timepoint`: Timepoints follow the format `b[01]_d*`, where `b0` and `b1` correspond to pre-3rd dose ("pre-boost") and post-3rd dose ("post-boost") respectively, and `d*` indicates the timepoint in days. There's one exception: `b0_m6or9` for pre-3rd dose d201 or d280 (m6or9 = 6 or 9 months).</li> <li>`gex_anno`: Cell type identity annotation based on transcriptomic profiles. Mapped from `anno_leiden_0.35` from WU382_alsoussi_et_al_nature_2023_betaDelta_gex_b_cells.h5ad.</li> <li>`compartment`: B cell compartment</li> <li>`clone_id`: B cell clonal lineage IDs follow the format `[donor]@[id]`.</li> <li>`s_pos_clone`: `TRUE` if a sequence belonged to a B cell clone that was designated as S-binding by virtue of containing one of the recombinant mAbs that tested positive via ELISA.</li> <li>`expressed_id`: mAb IDs of mAbs from <a href="https://doi.org/10.1038/s41586-021-03738-2">Turner & O'Halloran et al., <em>Nature</em>, 2021</a> and the current study; and of recombinant mAbs generated based on 10x BCRs from <a href="https://doi.org/10.1038/s41586-022-04527-1">Kim & Zhou et al., <em>Nature</em>, 2022</a>. `NA` for everything else.</li> <li>`elisa`: ELISA results for binding of recombinant mAbs to SARS-CoV-2 S. `TRUE` if positive (WA1+); `FALSE` if negaive; `NA` if not tested or test failed.</li> <li>`nuc_RS_19_312`: number of replacement and silent mutations between IMGT-numbered nucleotide positions 19-312 along IGHV sequences, calculated by <a href="https://shazam.readthedocs.io/en/stable/topics/calcObservedMutations/">`shazam::calcObservedMutations`</a>.</li> <li>`nuc_denom_19_312`: number of informative nucleotide positions for counting mutations, excluding non-A/T/G/C positions (such as "N", "-", ".").</li> <li>`nuc_RS_freq_19_312`: nucleotide-level mutation frequency (= nuc_RS_19_312 / nuc_denom_19_312).</li> </ul> <p> </p> <p><strong>[Beta & Delta booster] Processed BCR data - light chains</strong></p> <p>File: WU382_alsoussi_et_al_nature_2023_betaDelta_bcr_light.tsv.gz</p> <p><em>Light chains were not used for heavy chain-based clonal inference or analysis.</em></p> <p> </p> <p><strong>[Beta & Delta booster] Processed transcriptomics data</strong></p> <p>Files: </p> <ul> <li>WU382_alsoussi_et_al_nature_2023_betaDelta_gex_all_cells.h5ad</li> <li>WU382_alsoussi_et_al_nature_2023_betaDelta_gex_b_cells.h5ad</li> <li>WU382_alsoussi_et_al_nature_2023_betaDelta_gex_b_cell_umap.tsv.gz</li> </ul> <p>Notes on the `h5ad` files:</p> <ul> <li>These files can be imported into <a href="https://scanpy.readthedocs.io/en/stable/index.html">Scanpy</a> as an <a href="https://scanpy.readthedocs.io/en/stable/usage-principles.html#anndata">AnnData object</a>.</li> <li>Each `AnnData` object has 3 `.layers`, each representing a version of the count matrix. <ul> <li>`raw_counts`: Imported from `<a href="https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/6.0/using/aggregate">cellranger aggr</a>` output by `scanpy.read_10x_mtx`.</li> <li>`log_norm`: Log-noramlized expression values outputted by `scanpy.pp.normalize_total` followed by `scanpy.pp.log1p`.</li> <li>`scaled`: The `log_norm` layer scaled to unit variance and zero mean by `scanpy.pp.scale`. </li> </ul> </li> <li>The `gene_name` and `biotype` columns in `.var` were extracted from GENCODE v32 GTF.</li> <li>Columns in `.obs` (each row corresponds to a cell) <ul> <li>`n_feature`: The `n_genes_by_counts` column produced by `scanpy.pp.calculate_qc_metrics`, renamed. The number of genes expressed. This is before subsetting the genes.</li> <li>`n_umi`: The `total_counts` column produced by `scanpy.pp.calculate_qc_metrics`, renamed. The total UMI counts in a cell.</li> <li>`pct_mt`: The `pct_counts_mt` column produced by `scanpy.pp.calculate_qc_metrics`, renamed. The percentage of counts in mitochondrial genes.</li> <li>`n_hkg`: The number of housekeeping genes for which expression was detected.</li> <li>`n_gene_expressed`: The total number of genes for which expression was detected. This is after subsetting the genes.</li> <li>`pre_qc_bcr`: `TRUE` if a cell also had paired BCR data available. Produced by cross-referencing the cellular barcodes in `cell_barcodes.json` outputted by `cellranger vdj`. At this point the BCR data had not gone through the QC process in the BCR processing pipeline (hence `pre_qc`). </li> <li>`leiden_[resolution]`: Cluster assignment by `scanpy.tl.leiden`.</li> <li>`anno_leiden_[resolution]`: Cell type identity annotations based on transcriptomic profiles. This was mapped onto the `gex_anno` column in the processed heavy chain BCR data.</li> </ul> </li> <li>UMAP coordinates can be found in `.obsm["X_umap"]`.</li> <li>`.X` has been set to `None` in order to reduce file size.</li> </ul> <p>Note on the `tsv.gz` file: This file was derived from WU382_alsoussi_et_al_nature_2023_betaDelta_gex_b_cells.h5ad. It contains UMAP coordinates and select attributes of the cells, including their log-normalized expression values of XBP1 (`ln_XBP1`). For analysis and visualization in conjunction with BCR data.</p> <p>In addition, the preprocessed count matrix outputted by `<a href="https://support.10xgenomics.com/single-cell-gene-expression/software/pipelines/6.0/using/aggregate">cellranger aggr</a>` is available from <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE227562">GEO under BioProject PRJNA800176</a>.</p> <p> </p> <p><strong>[Omicron booster] Processed BCR data - heavy chains</strong></p> <p>File: WU382_alsoussi_et_al_nature_2023_omicron_bcr_heavy.tsv.gz</p> <p>Notes on columns:</p> <ul> <li>`elisa`: ELISA results for mAbs, with values being one of `WA1+`, `BA1+WA1-`, or `negative`. `NA` for bulk sequences.</li> <li>`clone_type`: If a sequence was in an S-binding B cell clone (`TRUE` for `s_pos_clone`), its `clone_type` was based on the `elisa` value of the S-binding mAb in that clone -- either `WA1+` or `BA1+WA1-`; otherwise `NA`.</li> </ul>
Dataset of the Article "Reconstruction of the unbinding pathways of new inhibitors of the SARS-CoV-2 Papain-like protease using molecular dynamics simulation"
<p>This dataset contains concatenated trajectory files of the SuMD simulation of the unbinding pathways of the new inhibitors for SARS-CoV-2 papain-like protease. This data will be published in an article titled: "<strong>Reconstruction of the unbinding pathways of new inhibitors of the SARS-CoV-2 Papain-like protease using molecular dynamics simulation".</strong></p>
COVFlow: performing virus phylodynamics analyses from selected SARS-CoV-2 genome sequences
<p>This upload contains pipeline configuration files, output data, scripts and data identifiers (GISAID EPI_ISL_ID) required to reproduce the results of the article entitled "COVFlow: performing virus phylodynamics analyses from selected SARS-CoV-2 genome sequences".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.