Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
95
datasets available to search
ShareScore release 0.9.0
Dataset results
95 results for “Viral genome”
COG-UK Viral Genome Sequences
<p>COG-UK Consortium has published dataset contains over 10K SARS-CoV-2 viral genome sequences available as open access. The current COVID-19 pandemic, caused by the SARS-CoV-2 virus, represents a major threat to health in the UK and globally. To fully understand the transmission and evolution of the virus requires sequencing and analysing viral genomes at scale and speed. The numbers of samples calls for a rapid increase in the UK’s pathogen genome sequencing capacity rapidly and robustly. To provide this increased capacity to collect, sequence and analyse the whole genomes of virus samples in the UK, the COVID-19 Genomics UK (COG-UK) consortium is pooling the world-leading knowledge and expertise in genomics of the four UK Public Health Agencies, multiple regional University hubs, and large sequencing centres such as the Wellcome Sanger Institute.</p> <ul> <li>Protocols: https://www.cogconsortium.uk/protocols/</li> </ul>
Adaptive changes in the genomes of wild rabbits after 16 years of viral epidemics
<p>Since its introduction to control overabundant alien rabbits (Oryctolagus cuniculus), the highly virulent Rabbit Haemorrhagic Disease Virus (RHDV) has caused regular annual disease outbreaks in Australian rabbit populations. Although initially reducing rabbit abundance by 60%, continent-wide, experimental evidence has since indicated increased genetic resistance in wild rabbits that have experienced RHDV-driven selection. To identify genetic adaptations, which explain the increased resistance to this biocontrol virus, we investigated genome-wide SNP (single nucleotide polymorphism) allele frequency changes in a South Australian rabbit population that was sampled in 1996 (pre-RHD genomes) and after 16 years of RHDV outbreaks. We identified several SNPs with changed allele frequencies within or in proximity of genes that have roles potentially important for increased RHD resistance. Many of the identified genes are known to be involved in virus infections or immunity, or had previously been identified as being differentially expressed in healthy vs. acutely RHDV-infected rabbits. Furthermore, we show in a simulation study that the allele/genotype frequency changes cannot be explained by drift alone, and that several candidate genes had also been identified as being associated with surviving RHD in a different Australian rabbit population. Our unique dataset allowed us to identify candidate genes for RHDV resistance that have evolved under natural conditions, and over a time span that would not have been feasible to study in an experimental setting. Moreover, it provides a rare example of host genetic adaptations to virus-driven selection in response to a suddenly emerging infectious disease.</p>
"Genome binning of viral entities from bulk metagenomics data" - CAMISIM simulated datasets and genomes
<p><strong>Genome binning of viral entities from bulk metagenomics data</strong></p> <p> </p> <p><strong>Authors</strong></p> <p><strong>Joachim Johansen1,2, Damian R. Plichta2, Jakob Nybo Nissen1,3, Marie Louise Jespersen1,4, Shiraz A. Shah5, Ling Deng6, Jakob Stokholm5,6, Hans Bisgaard5, Dennis Sandris Nielsen6, Søren Sørensen7, Simon Rasmussen1</strong></p> <p> </p> <p><strong>Affiliations</strong></p> <p>1 Novo Nordisk Foundation Center for Protein Research, Faculty of Health and Medical Sciences, University of Copenhagen, Copenhagen N, Denmark</p> <p>2 Infectious Disease and Microbiome Program, Broad Institute of MIT and Harvard, Cambridge, MA, USA</p> <p>3 Statens Serum Institut, Viral & Microbial Special diagnostics, Copenhagen, Denmark</p> <p>4 National Food Institute, Technical University of Denmark, Kongens Lyngby, Denmark</p> <p>5 Copenhagen Prospective Studies on Asthma in Childhood (COPSAC), Herlev and Gentofte Hospital, University of Copenhagen, Copenhagen, Denmark</p> <p>6 Section of Food Microbiology and Fermentation, Department of Food Science, Faculty of Science, University of Copenhagen, Copenhagen, Denmark</p> <p>7 Section of Microbiology, Department of Biology, University of Copenhagen, Copenhagen, Denmark</p> <p><strong>Methods description</strong></p> <p>We compared the viral binning performance of VAMB and MetaBAT2 using the official CAMI consortium method to create assemblies and metagenome profiles. To this end we generated 3 different metagenome compositions with up to 308 reference genomes; one mixed with bacteria, plasmids and viruses to test binning in complex samples i.e. high diversity (1), one with only crass-like viruses to test binning with highly similar viruses i.e. high relatedness (2) and a set of small-viruses (<6,000 bp) including members of the Microviridae family to address the bias of size (3). Bacterial genomes were gathered from NCBIs refseq genome repository 2021, plasmids from the PLSDB database (v. 2021_06_23) and viral genomes from the recent MGV database. </p> <p>Dataset A contained a mixture of bacteria (N=8), plasmids (N=20) and viruses (N=280) to test binning in complex samples, i.e. high diversity. Dataset B contained only crass-like viruses (N=80) to test binning with highly similar viruses i.e. high relatedness. Dataset C contained small-viruses (N=50, <6,000 bp) of the Microviridae family to address the bias of size. Bacterial genomes were sampled from the Refseq genome repository 2021, plasmids from the PLSDB database and viral genomes from the recent MGV database (Nayfach, et al. Nature Microbiology 2021).</p> <p> </p> <p> </p>
Virus classification for viral genomic fragments using PhaGCN2
<p>Viruses are the most ubiquitous and diverse entities in the biome. Due to the rapid growth of newly identified viruses, there is an urgent need for accurate and comprehensive virus classification, particularly for novel viruses. Here, we present PhaGCN2, which can rapidly classify the taxonomy of viral sequences at family level and supports the visualization of the associations of all families. We evaluate the performance of PhaGCN2 and compare it with the state-of-the-art virus classification tools, such as vConTACT2, CAT, and VPF-Class, using the widely accepted metrics. The results show that PhaGCN2 largely improves the precision and recall of virus classification, increases the number of classifiable virus sequences in the Global Ocean Virome dataset (v2.0) by 4 times, and classifies more than 90% of the Gut Phage Database. PhaGCN2 makes it possible to conduct high-throughput and automatic expansion of the database of the International Committee on Taxonomy of Viruses.</p>
Deep-sequencing of viral genomes from treatment-naive HIV-infected persons shows positive association between intrahost genetic diversity and viral load
<p><strong><span>Background:</span></strong><span> Infection with human immunodeficiency virus type 1 (HIV) typically results from transmission of a small and genetically uniform viral population. Following transmission, the virus population becomes more diverse because of recombination and acquired mutations through genetic drift and selection. Viral intrahost genetic diversity remains a major obstacle to the cure of HIV; however, there is a disagreement whether intrahost viral genetic diversification associates positively or negatively with disease progression and progression markers. Viral load is a key progression marker and understanding its relationship to viral intrahost genetic diversity could help design future strategies for HIV monitoring and treatment.</span></p> <p><span><strong>Methods:</strong> </span><span>We analyzed deep-sequenced viral genomes from 2,650 treatment-naive HIV-infected persons to measure the intrahost genetic diversity of 2,447 genomic codon positions as calculated by Shannon entropy. We tested for associations between viral load (VL) and amino acid (AA) entropy accounting for sex, age, race, duration of infection, and HIV population structure.</span></p> <p><strong><span>Results:</span></strong><span><strong> </strong>We confirmed that the intrahost genetic diversity is highest in the <em>env</em> gene. Furthermore, we showed that mean Shannon entropy is significantly associated with VL, especially in infections of >24 months duration. We identified 16 significant associations between VL (p-value<2.0x10<sup>-5</sup>) and Shannon entropy at AA positions which in our association analysis explained 13% of the variance in VL.</span></p> <p><strong><span>Conclusions: </span></strong><span>Our results elucidate that viral intrahost genetic diversity is associated with VL and could be used as a better disease progression marker than HIV consensus sequence variants, especially in infections of longer duration. We emphasize that viral intrahost diversity should be considered when studying viral genomes and infection outcomes.</span></p>
Data Pack fro VirSorter: mining viral signal from microbial genomic data
<p>This is the data pack for VirSorter, the publication of which by Roux et al. titled "<strong>VirSorter: mining viral signal from microbial genomic data</strong>" appeared in PeerJ on 2015-05-28 (<a href="https://doi.org/10.7717/peerj.985">doi:10.7717/peerj.985</a>).</p> <p>Most up-to-date tutorials and the code for VirSorter can be found at the GitHub repository <a href="https://github.com/simroux/VirSorter">https://github.com/simroux/VirSorter</a>.</p> <p>The original source of this data pack was here (last accessed 2018-02-03): <a href="http://datacommons.cyverse.org/browse/iplant/home/shared/imicrobe/VirSorter/virsorter-data.tar.gz">http://datacommons.cyverse.org/browse/iplant/home/shared/imicrobe/VirSorter/virsorter-data.tar.gz</a></p>
Supplemental data for: Evaluation of SARS-CoV-2 response at the University of North Carolina (UNC) at Charlotte using percent positivity data and viral genomic sequence data.
<p>Supplemental data for:</p> <p>Evaluation of SARS-CoV-2 response at the<br>University of North Carolina (UNC) at Charlotte<br>using percent positivity data and viral genomic<br>sequence data.</p> <p>Submitted to Biocarla 2024</p> <p>https://carla2024.org/portfolios/biocarla/</p> <p>Authors:</p> <p>Daniel Janies 1,2,3 [0000−0002−7890−9906], Shirish Yasa 1,2,3 [0000−0003−3217−4921],<br>Colby T. Ford 1,3,4 [0000−0002−7859−3622] Jannatul Ferdous 2,3 [0000−0003−3053−9616],<br>William Taylor 2,3 [0009−0000−6204−1172], April Harris 2,3 [0009−0009−2557−7926],<br>Sam Kunkleman 2,3 [0000−0002−2309−6418], Juan Bolanos 2,3, Kevin Lambirth 2,3 [0000-0002-6568-543X], Denis<br>Jacob Machado 1,2,3 [0000−0001−9858−4515], Cynthia Gibas 1,2,3 [0000−0002−1288−9543],<br>and Jessica Schlueter 1,2,3 [0000−0002−6490−0580]</p> <p>Affiliations:</p> <p>1) Center for Computational Intelligence to Predict Health and Environmental Risks<br>(CIPHER), University of North Carolina at Charlotte 28223, USA<br>Correspondence to: djanies@charlotte.edu<br>https://cipher.charlotte.edu<br>2) Department of Bioinformatics and Genomics, University of North Carolina at<br>Charlotte 28223, USA https://cci.charlotte.edu/departments/<br>department-of-bioinformatics-and-genomics/<br>3) College of Computing and Informatics, University of North Carolina at Charlotte<br>28223, USA https://cci.charlotte.edu<br>4) School of Data Science, University of North Carolina at Charlotte 28223, USA<br>https://sds.charlotte.edu<br>5) Division of Research, University of North Carolina at Charlotte 28223, USA<br>https://research.charlotte.edu/</p> <p> </p>
Correlation of the kinetics of viral antigen and genomic RNA with restoration of normal cell homeostasis
<p>The main objective of the studies within COCID Work Package 6 is to understand basic mechanisms by which viral replication machineries are removed from cells after pharmacological interruption of viral replication using functional as well as imaging techniques, including soft X-ray tomography.</p> <p>In this document, we highlight the progress in establishing the HCV replication models (replicons), the antiviral treatment chosen to eliminate viral structures from the host cells and the correlation between elimination of the viral replication machinery and restoration of a normal host cell homeostasis, using markers of HCV-induced stress. These data are essential to frame the experimental setup chosen for the imaging process required in subsequent stps of the project.</p>
Data corresponding to: Evaluation of sequencing and PCR-based methods for the quantification of the viral genome formula
<p>Viruses show great diversity in their genome organisation. Multipartite viruses package their genome segments into separate particles, most or all of which are required to initiate infection in the host cell. The benefits of such seemingly inefficient genome organization are not well understood. One hypothesised benefit of multipartition is that it allows for flexible changes in gene expression by altering the frequency of each genome segment in different environments, such as encountering different host species. The ratio of the frequency of segments is termed the genome formula (GF). Thus far, formal studies quantifying the GF have been performed for well-characterised virus-host systems in experimental settings using RT-qPCR. However, to understand GF variation in natural populations or novel virus-host systems, a comparison of several methods for GF estimation including high-throughput sequencing (HTS) based methods is needed. Currently, it is unclear how HTS-methods compare a golden standard, such as RT-qPCR. Here we show a comparison of multiple GF quantification methods (RT-qPCR, RT-digital PCR, Illumina RNAseq and Nanopore direct RNA sequencing) using three host plants (<em>Nicotiana tabacum</em>, <em>Nicotiana benthamiana</em>, and <em>Chenopodium quinoa</em>) infected with cucumber mosaic virus (CMV), a tripartite RNA virus. Our results show that all methods give roughly similar results, though there is a significant method effect on genome formula estimates. While the RT-qPCR and RT-dPCR GF estimates are congruent, the GF estimates from HTS methods deviate from those found with PCR. Our findings emphasise the need to tailor the GF quantification method to the experimental aim, and highlight that it may not be possible to compare HTS and PCR-based methods directly. The difference in results between PCR-based methods and HTS highlights that the choice of quantification technique is not trivial.</p>
Adaptive changes in the genomes of wild rabbits after 16 years of viral epidemics
Open the record for dataset details and reuse information.
Deep-sequencing of viral genomes from treatment-naive HIV-infected persons shows positive association between intrahost genetic diversity and viral load
Open the record for dataset details and reuse information.
Data corresponding to: Evaluation of sequencing and PCR-based methods for the quantification of the viral genome formula
Open the record for dataset details and reuse information.
Virus classification for viral genomic fragments using PhaGCN2
Open the record for dataset details and reuse information.
Seagrass associated viral genomes
<p>High-quality viral sequences associated with:</p> <p>A genomic resource for exploring bacterial-viral dynamics in seagrass ecosystems</p> <p>Analysis, code, intermediate and supporting files are archived here: <a href="https://doi.org/10.5281/zenodo.14226514">10.5281/zenodo.14226514</a></p> <p>Bacterial metagenome-assembled genomes from this work are archived here: <a href="https://doi.org/10.5281/zenodo.14225974" target="_blank" rel="noopener">10.5281/zenodo.14225974</a><br><br>This archive contains:<br>(i) Fasta file representing the 354 viral sequences in the final catalog described in the above titled work<br>(ii) Metadata file describing the viral catalog (i.e., Table S2 from the above work)</p>
Freshwater viral metagenome assembled genomes (vMAGs) used for vContact2 analysis in publication Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments
<p>This dataset contains all freshwater viruses that were mined from publicly available data in an effort to provide biogeographical context to viral communities identified from the Columbia River. These two files include data from:</p> <p>1) East River, CO (PRJNA579838)</p> <p>2) A previous study from the Columbia River, WA (PRJNA375338)</p> <p>3) Prairie Potholes, ND (PRJNA365086)</p> <p>4) Amazon River (PRJNA237344)</p> <p> </p> <p>Manuscript title Genome-resolved metaproteomics decodes the microbial and viral contributions to coupled carbon and nitrogen cycling in river sediments</p>
Exploiting functional regions in the viral RNA genome as druggable entities
<p><span>XML files with normalized SHAPE reactivities are provided in SHAPE_react_rep1/2.react.xml; </span></p> <p><span>WIG files with Shannon entropies are provided in SHAPE_react_rep1/2.react.xml,</span></p> <p><span>and the full secondary structure are provided in PEDV_incell_secondary structure.ct</span></p>
Viral metagenome assembled genomes (vMAGs) from Columbia River hyporheic sediments
<p>Fasta file containing 111 viral metagenome assembled genomes (vMAGs) from publication to be submitted titled "<strong>Microbial genome-resolved metaproteomic analyses frame intertwined carbon and nitrogen cycles in river hyporheic sediments". </strong></p>
Viral reference genomes to disentangle the recombinant phylogenetic history of the potyviruses
<p>Potyviruses are a large genus of plant-infecting RNA viruses in the family Potyviridae. Due to their rapid diversification and frequent recombination, reconstructing the phylogenetic history of the potyviruses has proven difficult. Phylogenies reconstructed from different protein-coding regions of the viral genome often reveal conflicing or discordant relationships. But the extent to which discordance is due to interspecific recombination versus phylogenetic noise or errors in reconstruction is unclear. </p> <p>To explore the recombinant history of the potyviruses, we assembled a dataset containing referece genomes for 131 species of potyviruses. High-quality, full-length reference genomes for all species were obtained form NCBI GenBank. Viral genomes were carefully aligned at the codon-level and screened for recombination. The full alignment was then partitioned into several sub-alignments between each detected recombination event, such that each sub-aligment corresponds to a non-recombinant block (NRB) free of detected recombination events. Local phylogenetic trees for each NRB were then reconstructed to explore how phylogenetic relationships varied across different regions of the potyvirus genome. </p> <p>We then used our program Espalier to disentangle the phylogenetic history of the potyviruses. Espalier reconciles and removes discordances between phylogenetic trees that are likely attributable to phylogenetic error while retaining recombination events that are strongly supported by the sequence data. Applying Espalier to the potyviruses revealed that most phylogenetic discordace between local trees is likely attributable to phylogenetic noise. Removing the discordance attributable to phylogenetic error allows us to much more clearly visualize the phylogenetic history of the potyviruses.</p>
SARS-CoV-2 Viral Samples and Reference Genome for Galaxy Training Network SARS with Galaxy on AnVIL Tutorial
<p>In the lab activity, we'll see if there are genomic differences in the collected sample compared to the original SARS-CoV-2 genome (the reference). We need three files to do this:</p> <ul> <li><strong>SARS-CoV-2_reference_genome.fasta</strong> : the reference genome</li> <li><strong>VA_sample_forward_reads.fastq.gz</strong>: 1 of 2 raw read data files</li> <li><strong>VA_sample_reverse_reads.fastq.gz</strong>: 2 of 2 raw read data files</li> </ul> <p>The sample for this activity was derived from data collected at Virginia Commonwealth University in Richmond, VA. The researchers collected the sample with the goal of being able to track the spread and evolution of this virus state-wide, nationally, and internationally. You can download the original data <a href="https://www.ncbi.nlm.nih.gov/sra/?term=XGTK449087">here</a>.</p> <p>You can download the reference genome <a href="https://www.ncbi.nlm.nih.gov/nuccore/1798174254">here</a>.</p>
Data for "Mixed viral infection constrains the genome formula of multipartite cucumber mosaic virus"
<p>We performed a study to assess the effect of mixed infection on the genome formula of multipartite virus CMV. We used qPCR to determine genome formulas and titer. Additionally, simulation models were developed to describe mechanisms for genome formula change under mixed infection. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.