Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

62

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

62 results for “virome”

Learn how ShareScore rates datasets ↗
zenodo32/100

Urban wastewater virome by viral metagenomics and target enrichment sequencing

<p>Metagenomic analysis of virus in raw sewage.</p>

opencc-by-4.0Jan 2021View details →
zenodo32/100

Dereplicated vOTUs recovered from a viromics extraction buffer chemistry experiment

<p>Database of de-replicated viral operational taxonomic units recovered from a viromics (viral-size fraction metagenomics) buffer chemistry experiment.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Birth of new protein folds and functions in the virome - Structure Database

<p>This is the database of protein structures described in the manuscript "Birth of new protein folds and functions in the virome", by authors Jason Nomburg, Nate Price, and Jennifer A. Doudna.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

Virome of Helicoverpa armigera, Tuta absoluta, Nezara viridula

<p>Virome of Helicoverpa armigera, Tuta absoluta, Nezara viridula</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Characterisation Of The Excreted Virome In The Population: NGS Tools And Epidemiologically Significant Pathogens

<p>Metagenomic analysis of virus in raw sewage.</p>

opencc-by-4.0Nov 2019View details →
zenodo32/100

Dataset for: Tameness selection pressure affects gut virome diversity in mice

<p>This dataset comprises a non-redundant set of wild hetrogenouse stock mice virome (WHS-MV) cataleogue of 6,078 vOTUs.&nbsp; All 80 raw shotgun metagenomic sequencing datasets (150 bp reads) used to generate WHS-MV catalogues are available in the NCBI under BioProject PRJDB15857 with BioSample accession numbers from SAMD00614304 to SAMD00614383 and SRA accession numbers from DRR480456 to DRR480535. Raw shotgun metagenomic sequencing datasets of four samples (250 bp reads) are available under BioProject PRJDB18588, with BioSample accession numbers SAMD00805331 to SAMD00805334 and SRA accession numbers DRR585793 to DRR585796.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Database of clustered vOTUs recovered from a viromics prescribed burn study of forest soil

<p>Database of dereplicated viral operational taxonomic units (vOTUs) recovered from a viromics (viral-size fraction metagenomics) prescribed burn study of forest soil</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

Virus sequences associated with the rodent virome database for the years 2014 and 2016/2017.

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

Hybrid Illumina/PacBio assembly of a virome from the Rhode River at the Smithsonian Environmental Research Center (SERC).

<p>Hybrid Illumina/PacBio assembly of a virome from the Rhode River at the Smithsonian Environmental Research Center (SERC). This dataset was used in the following publication:&nbsp;https://www.frontiersin.org/articles/10.3389/fmicb.2018.03053/full.</p>

opencc-by-4.0Dec 2018View details →
dryad32/100

The chronic wound virome: phage diversity and associations with wounds and healing outcomes

Open the record for dataset details and reuse information.

publicApr 2022View details →
dryad32/100

Data from: Demography, life history trade-offs, and the gastrointestinal virome of wild chimpanzees

Open the record for dataset details and reuse information.

publicSep 2020View details →
zenodo28/100

Limited host-specificity of eukaryotic virome in Hymenoptera

<p>Additional datafiles for Bee_Euvir repository.</p>

opencc-by-4.0Aug 2020View details →
dryad28/100

Data from: Evolution and diversity of the Microviridae viral family through a collection of 81 new complete genomes assembled from virome reads.

Recent studies suggest that members of the Microviridae (a family of ssDNA bacteriophages) might play an important role in a broad spectrum of environments, as they were found dominant among the viral fraction from seawater and human gut samples. 24 completely sequenced Microviridae have been described so far, divided into three distinct groups named Microvirus, Gokushovirinae and Alpavirinae, this last group being only composed of prophages. In this study, we present the analysis of 81 new complete Microviridae genomes, assembled from viral metagenomes originating from various ecosystems. The phylogenetic analysis of the core genes concludes to the existence of four groups, confirming the three sub-families described so far and exhibiting a new group, named Pichovirinae. The genomic organizations of these viruses are strikingly coherent with their phylogeny, the Pichovirinae being the only group of this family with a different organization of the three core genes. Analysis of the structure of the major capsid protein reveals the presence of mushroom-like insertions conserved within all the groups except for the Microvirus. In addition, a peptidase gene was found in 11 Microviridae and its analysis concludes to a horizontal gene transfer that occurred several times between these viruses and their bacterial hosts. This is the first report of such gene transfer in microviruses. Finally, searches against viral metagenomes revealed the presence of highly similar sequences in a variety of biomes indicating that Microviridae probably have both an important role in these ecosystems and an ancient origin.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Using viromes to predict novel immune proteins in non-model organisms

Immunity is mostly studied in a few model organisms, leaving the majority of immune systems on the planet unexplored. To characterize the immune systems of non-model organisms alternative approaches are required. Viruses manipulate host cell biology through the expression of proteins that modulate the immune response. We hypothesized that metagenomic sequencing of viral communities would be useful to identify both known and unknown host immune proteins. To test this hypothesis, a mock human virome was generated and compared to the human proteome using tBLASTn, resulting in 36 proteins known to be involved in immunity. This same pipeline was then applied to reef-building coral, a non-model organism that currently lacks traditional molecular tools like transgenic animals, gene-editing capabilities, and in vitro cell cultures. Viromes isolated from corals and compared with the predicted coral proteome resulted in 2503 coral proteins, including many proteins involved with pathogen sensing and apoptosis. There were also 159 coral proteins predicted to be involved with coral immunity but currently lacking any functional annotation. The pipeline described here provides a novel method to rapidly predict host immune components that can be applied to virtually any system with the potential to discover novel immune proteins.

opencc-zeroDec 2015View details →
zenodo28/100

Data from: Identification of prokaryotic and eukaryotic virus-derived sequences in virome using deep learning

<h4>This repository contains the data and Docker image to reproduce the results of our paper: <strong>identification of prokaryotic and eukaryotic virus-derived sequences in virome using deep learning</strong></h4> <p>Authors: Hengchuang Yin, Shufang Wu, Jie Tan, Qian Guo, Mo Li, Jinyuan Guo, Yaqi Wang, Xiaoqing Jiang, and Huaiqiu Zhu*</p> <p><strong>This work has been accepted by GigaScience.&nbsp;</strong></p> <p><strong>Hengchuang Yin, Shufang Wu, Jie Tan, Qian Guo, Mo Li, Jinyuan Guo, Yaqi Wang, Xiaoqing Jiang, and Huaiqiu Zhu. "IPEV: Identification of Prokaryotic and Eukaryotic Virus-Derived Sequences in Virome Using Deep Learning." GigaScience 13 (2024): giae018.&nbsp;<a href="https://doi.org/10.1093/gigascience/giae018" rel="nofollow">https://doi.org/10.1093/gigascience/giae018</a>.</strong></p> <div>&nbsp;</div> <p>&nbsp;</p> <p><strong>Background: </strong>The virome obtained through virus-like particle enrichment contains a mixture of prokaryotic and eukaryotic virus-derived fragments. Accurate identification and classification of these elements are crucial to understanding their roles and functions in microbial communities. However, the rapid mutation rates of viral genomes pose challenges in developing high-performance tools for classification, potentially limiting downstream analyses.</p> <p><strong>Findings: </strong>We present IPEV, a novel method to distinguish prokaryotic and eukaryotic viruses in viromes, with a 2D convolutional neural network combining trinucleotide pair relative distance and frequency. Cross-validation assessments of IPEV demonstrate its state-of-the-art precision, significantly improving the F1-score by approximately 22% on an independent test set compared to existing methods when query viruses share less than 30% sequence similarity with known viruses.&nbsp;Furthermore, IPEV outperforms other methods in accuracy on marine and gut virome samples based on annotations by sequence alignments. IPEV reduces runtime by at most 1,225 times compared to existing methods under the same computing configuration. We also utilized IPEV to analyze longitudinal samples and found that the gut virome exhibits a higher degree of temporal stability than previously observed in persistent personal viromes, providing novel insights into the resilience of the gut virome in individuals.&nbsp;</p> <p><strong>Conclusions:&nbsp;</strong>IPEV is a high-performance, user-friendly tool that assists biologists in identifying and classifying prokaryotic and eukaryotic viruses within viromes. The tool is available at&nbsp;https://github.com/basehc/IPEV.</p> <p>&nbsp;</p> <p><strong>5_fold_cross_validation.zip:</strong> Dataset of cross-validation of IPEV</p> <p><strong>Eukaryotic_virus_CV_Dataset-1.csv:</strong> GI, and accession ID for the cross-validation Dataset-1 (eukaryotic virus)</p> <p><strong>Prokaryotic_virus_CV_Dataset-1.csv:</strong> GI, and accession ID for the cross-validation Dataset-1 (prokaryotic virus)</p> <p><strong>Test_Prokaryotic_virus_Dataset-1.fasta: </strong>An independent test set of IPEV (prokaryotic&nbsp;virus)</p> <p><strong>Test_Eukaryotic_virus_Dataset-1.fasta: </strong>An independent test set of IPEV (eukaryotic&nbsp;virus)</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Dataset_sequencing_error.zip: </strong>Simulated dataset with sequencing errors</p> <p><strong>Cap_enzyme_sequence.fasta: </strong>Accession IDs of Receptor Binding Proteins (RBPs) in phages collected by our article</p> <p><strong>Dataset_runtime_evaluation.zip: </strong>Dataset for evaluating the runtime of IPEV</p> <p><strong>Receptor_binding_protein_accession_id:</strong>&nbsp;Accession IDs of Receptor Binding Proteins (RBPs) in phages collected by our article</p> <p>&nbsp;</p> <p><strong>archaea_ID.txt</strong> Accession ID information for the reference archaea dataset</p> <p><strong>bacteria_ID.txt</strong> Accession ID information for the reference bacterial dataset</p> <p><strong>marine_virome_id.csv:</strong> Ocean virome data information used in our paper</p> <p><strong>gut_virome.csv</strong>:Gur virome data information used in our paper</p> <p><strong>fungi.txt: </strong>Negative sequence information used to train, validate, and test the model in the decontamination function</p> <p><strong>bacteria.txt: </strong>Negative sequence information used to train, validate, and test the model in the decontamination function</p> <p>&nbsp;</p> <p>&nbsp;</p> <h4><strong>Reproduce the results of our paper from a Docker image.</strong></h4> <p>&nbsp;</p> <p>We also provide a Docker image file that does not require any environment configuration. You can reproduce the results of our paper (e.g., train and test our IPEV model) in a Docker image.</p> <p>Pull the<a href="https://hub.docker.com/r/dryinhc/ipev_v1"> <em>dryinhc/ipev_v1</em></a> image from Docker Hub. Open a terminal window and run the following command:</p> <p><em>docker pull dryinhc/ipev_v1</em></p> <p>This will download the image to your local machine.</p> <p>Run the <em>dryinhc/ipev_v1</em> image. In the same terminal window, run the following command:</p> <p><em>docker run -it --rm dryinhc/ipev_v1</em></p> <p>This will start a container based on the image and run the IPEV tool.</p> <p>And you can run cd train or cd other file folders in the container.</p> <p>To exit the container, press <em>Ctrl+D</em> or type&nbsp;<em>exit</em>.</p> <p>It contains 4 directories, namely 5 fold cross validation, independent set, marine virome, and gut virome. The 5-fold cross-validation directory holds the scripts required for implementing the 5-fold cross-validation method. The independent set directory contains scripts necessary for working with an independent set. Lastly, the marine virome and gut virome directories store scripts for analyzing real datasets.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>We hereby confirm that the dataset associated with the research described in this work is made available to the public under the Creative Commons Zero (CC0) license.</p> <p>&nbsp;</p> <p>&nbsp;</p> <h4><strong>Contact&nbsp;</strong></h4> <p>&nbsp;</p> <p>If you have any questions, please don't hesitate to ask me: yinhengchuang@pku.edu.cn or hqzhu@pku.edu.cn</p>

openother-pdNov 2023View details →
zenodo28/100

Data to search Lake Baikal viromes against GOV 2.0

<p>Includes reads from &quot;Metagenomic Analysis of Virioplankton from the Pelagic Zone of Lake Baikal&quot; (https://doi.org/10.3390/v11110991).&nbsp; MG-RAST IDs:&nbsp;BVP1 --&nbsp;mgm4814173.3, BVP2 --&nbsp;mgm4816981.3.</p> <p>Also, GOV 2.0 viral populations over 10kb&nbsp;from&nbsp;http://datacommons.cyverse.org/browse/iplant/home/shared/iVirus/GOV2.0/GOV2_viral_populations_larger_than_10KB_or_circular.zip.&nbsp; (From&nbsp;Marine DNA Viral Macro- and Microdiversity from Pole to Pole,&nbsp;https://doi.org/10.1016/j.cell.2019.03.040).</p> <p>The data is grouped together here for ease of access.</p>

opencc-by-4.0Nov 2019View details →
zenodo28/100

Data files for "Comprehensive Wastewater Sequencing Reveals Community and Variant Dynamics of the Collective Human Virome"

<p>Files required to run analyses and generate charts for R notebooks at:&nbsp;<a href="https://github.com/cmmr/TX_wastewater_virome">https://github.com/cmmr/TX_wastewater_virome</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
dryad28/100

Data from: Evolution and diversity of the Microviridae viral family through a collection of 81 new complete genomes assembled from virome reads.

Open the record for dataset details and reuse information.

publicJul 2012View details →
dryad28/100

Data from: Using viromes to predict novel immune proteins in non-model organisms

Open the record for dataset details and reuse information.

publicAug 2016View details →
dryad28/100

CRISPR-Cas system of a prevalent human gut bacterium reveals hyper-targeting against phages in a human virome catalog

Open the record for dataset details and reuse information.

publicFeb 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record