Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,212

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,212 results for “virality”

Learn how ShareScore rates datasets ↗
zenodo40/100

Database for EVEREST (pipEline for Viral assEmbly and chaRactEriSaTion).

<p>Database for EVEREST (pipEline for Viral assEmbly and chaRactEriSaTion).</p> <p>Funding Statement:</p> <p>This work was supported by the National Health and Medical Research Council of Australia, Synergy grant (APP1183640).</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

Viral Communication: Longitudinal Survey Data on the Social Dimensions of the COVID-19 Pandemic

<p>This dataset represents the anonymised data collected as part of the Viral Communication (Understand-ELSED) project, which focussed on the social and ethical dimensions of the COVID-19 pandemic in Germany. It includes the three measurements; Phase I (30 October 2020 and 14 December 2020), Phase II (2 March 2021 and 22 March 2021) and Phase III.</p> <p>The first phase built the foundation for the wider suite of data collection approaches and research methods used in the Viral Communication project by allowing respondents to opt-in to multiple research pathways.</p> <p>Overall sample frame (Phase I): <em>N </em>= 1480</p> <p>Phase II sample frame: <em>N </em>= 482</p> <p>Phase III sample frame: <em>N </em>= 426</p> <p>Computed variables such as weights, groupings (experimental set-ups), and composite scores are included in the dataset.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

"Genome binning of viral entities from bulk metagenomics data" - CAMISIM simulated datasets and genomes

<p><strong>Genome binning of viral entities from bulk metagenomics data</strong></p> <p>&nbsp;</p> <p><strong>Authors</strong></p> <p><strong>Joachim Johansen1,2, Damian R. Plichta2, Jakob Nybo Nissen1,3, Marie Louise Jespersen1,4, Shiraz A. Shah5, Ling Deng6, Jakob Stokholm5,6, Hans Bisgaard5, Dennis Sandris Nielsen6, S&oslash;ren S&oslash;rensen7, Simon Rasmussen1</strong></p> <p>&nbsp;</p> <p><strong>Affiliations</strong></p> <p>1 Novo Nordisk Foundation Center for Protein Research, Faculty of Health and Medical Sciences, University of Copenhagen, Copenhagen N, Denmark</p> <p>2 Infectious Disease and Microbiome Program, Broad Institute of MIT and Harvard, Cambridge, MA, USA</p> <p>3 Statens Serum Institut, Viral &amp; Microbial Special diagnostics, Copenhagen, Denmark</p> <p>4 National Food Institute, Technical University of Denmark, Kongens Lyngby, Denmark</p> <p>5 Copenhagen Prospective Studies on Asthma in Childhood (COPSAC), Herlev and Gentofte Hospital, University of Copenhagen, Copenhagen, Denmark</p> <p>6 Section of Food Microbiology and Fermentation, Department of Food Science, Faculty of Science, University of Copenhagen, Copenhagen, Denmark</p> <p>7 Section of Microbiology, Department of Biology, University of Copenhagen, Copenhagen, Denmark</p> <p><strong>Methods description</strong></p> <p>We compared the viral binning performance of VAMB and MetaBAT2 using the official CAMI consortium method to create assemblies and metagenome profiles. To this end we generated 3 different metagenome compositions with up to 308 reference genomes; one mixed with bacteria, plasmids and viruses to test binning in complex samples i.e. high diversity (1), one with only crass-like viruses to test binning with highly similar viruses i.e. high relatedness (2) and a set of small-viruses (&lt;6,000 bp) including members of the Microviridae family to address the bias of size (3). Bacterial genomes were gathered from NCBIs refseq genome repository 2021, plasmids from the PLSDB database (v. 2021_06_23)&nbsp;and viral genomes from the recent MGV database.&nbsp;&nbsp;</p> <p>Dataset A contained a mixture of bacteria (N=8), plasmids (N=20) and viruses (N=280) to test binning in complex samples, i.e. high diversity. Dataset B contained only crass-like viruses (N=80) to test binning with highly similar viruses i.e. high relatedness. Dataset C contained small-viruses (N=50, &lt;6,000 bp) of the Microviridae family to address the bias of size. Bacterial genomes were sampled from the Refseq genome repository 2021, plasmids from the PLSDB database&nbsp; and viral genomes from the recent MGV database (Nayfach, et al. Nature Microbiology 2021).</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

FaCov Dataset: COVID-19 Viral News and Rumors Fact-Check Articles Dataset

<p>The data were collected by web-scraping pages from the websites collected earlier, using the <a href="https://webscraper.io/">Web Scraper browser extension</a>.</p> <p>More specifically, the sections of these websites that dealt exclusively with COVID-19 related content were scraped. In cases where the website did not have such a specified section, the search functionality within the website was used to query terms related to COVID-19 and the articles in the search results were scraped. Also in some cases, all articles were scraped and those unrelated to COVID-19 were filtered out in the pre-processing stage. All the samples collected were then put together into one CSV</p> <p>The following information was extracted along with the articles:</p> <p>Title of the fact check article</p> <p>URL of the fact check article</p> <p>Claim being discussed in the article (if available)</p> <p>Summary of the fact check article (if available)</p> <p>Content of the fact check article&bull; Label assigned by the article to the claim</p> <p>Author of the fact check article (if available)</p> <p>Date of publication of the article (if available)</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Extended Data Figure 7 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 7 | Global distribution of viral and host species richness for primates (order Primates). a, Observed total viral richness (for n = 71 host spp.); b, predicted total viral richness given maximum research effort; c, missing viruses or predicted minus observed total viral richness; d, observed zoonotic viral richness (n = 73);e, predicted zoonotic viral richness given maximum research effort; f, missing zoonoses or predicted minus observed zoonotic viral richness (same as included in Fig. 3e); g, global host species richness for Primates (n = 400); h, host species richness for Primates in our database (n = 98);i, primate species with no described viruses in the literature. Warmer colours (larger values) in c and f highlight areas predicted to be of greatest value for discovering novel viruses or novel viral zoonoses, respectively, in primates. Red/pink colours in panel i highlight areas with poor viral surveillance in primate species to date. Hatched regions represent areas where model predictions deviate systematically for the collection of species in that grid cell (see Methods).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Extended Data Figure 4 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 4 | Global distribution of viral and host species richness for wild carnivores (order Carnivora). a, Observed total viral richness (for n = 55 host spp.); b, predicted total viral richness given maximum research effort; c, missing viruses or predicted minus observed total viral richness; d, observed zoonotic viral richness (n = 55); e, predicted zoonotic viral richness given maximum research effort; f, missing zoonoses or predicted minus observed zoonotic viral richness (same as included in Fig. 3b); g, global host species richness for Carnivora (n = 276);h, host species richness for Carnivora in our database (n = 79); i, species of the order Carnivora with no described viruses in the literature. Warmer colours (larger values) in c and f highlight areas predicted to be of greatest value for discovering novel viruses or novel viral zoonoses, respectively, in carnivores. Red/pink colours in panel i highlight areas with poor viral surveillance in carnivore species to date. Hatched regions represent areas where model predictions deviate systematically for the collection of species in that grid cell (see Methods).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Extended Data Figure 3 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 3 | Global distribution of viral and host species richness for all wild mammals. a, Observed total viral richness (for n = 576 host spp.); b, predicted total viral richness given maximum research effort; c, missing viruses or predicted minus observed total viral richness; d, observed zoonotic viral richness (n = 584);e, predicted zoonotic viral richness given maximum research effort; f, missing zoonoses or predicted minus observed zoonotic viral richness (same as included in Fig. 3a); g, global mammal species richness (n = 5,290); h, mammal richness for species in our database (n = 753);i, mammal species with no described viruses in the literature. Warmer colours (larger values) in panels c and f highlight areas predicted to be of greatest value for discovering novel viruses or novel viral zoonoses, respectively, in mammals. Red/pink colours in panel i highlight areas with poor viral surveillance in mammal species to date. Hatched regions represent areas where model predictions deviate systematically for the collection of species in that grid cell (see Methods).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 4 in Host and viral traits predict zoonotic spillover from mammals

Figure 4 | Traits that predict zoonotic potential of a virus. a, Box plot of maximum phylogenetic host breadth per virus (PHB, see methods) for each of 586 mammalian viruses, aggregated by 28 viral families. Individual points represent viral species, colour-coded by zoonotic status. Box plots coloured and sorted by the proportion of zoonoses in each viral family. b–d, Partial effect plots for the best-fit GAM to predict the zoonotic potential of a virus. b, Maximum PHB. Viruses that infect a phylogenetically broader range of hosts are more likely to be zoonotic. c, Research effort (log, number of PubMed citations per viral species). d, Whether or not a virus replicates in the cytoplasm or is vector-borne. Viral genome length and whether or not a virus is enveloped improved the overall predictive power but were non-significant and are not shown (see Extended Data Table 1).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Extended Data Figure 2 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 2 | Heat map of observed total viral richness by mammalian order and viral family. Dataset includes 754 mammalian species and 586 unique ICTV recognized viral species. Heat map aggregated by rows and columns to group taxa with similar levels of observed viral richness.

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 3 in Host and viral traits predict zoonotic spillover from mammals

Figure 3 | Global distribution of the predicted number of 'missing zoonoses' by order. Warmer colours highlight areas predicted to be of greatest value for discovering novel zoonotic viruses. a, All wild mammals (n = 584 spp. included in the best-fit model). b, Carnivores (order Carnivora, n = 55).c, Even-toed ungulates (order Cetartiodactyla, n = 70). d, Bats (order Chiroptera, n = 157).e, Primates (order Primates, n = 73). f, Rodents (order Rodentia, n = 183). Hatched regions represent areas where model predictions deviate systematically for the assemblage of species in that grid cell (approximately 18 km × 18 km, see Methods). Animal silhouettes from PhyloPic.

opencc-by-4.0Jun 2017View details →
zenodo40/100

Extended Data Figure 6 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 6 | Global distribution of viral and host species richness for bats (order Chiroptera). a, Observed total viral richness (for n = 156 host spp.); b, predicted total viral richness given maximum research effort; c, missing viruses or predicted minus observed total viral richness; d, observed zoonotic viral richness (n = 157);e, predicted zoonotic viral richness given maximum research effort; f, missing zoonoses or predicted minus observed zoonotic viral richness (same as included in Fig. 3d); g, global host species richness for Chiroptera (n = 1117);h, host species richness for Chiroptera in our database (n = 192);i, species of the order Chiroptera with no described viruses in the literature. Warmer colours (larger values) in c and f highlight areas predicted to be of greatest value for discovering novel viruses or novel viral zoonoses, respectively, in bats. Red/pink colours in panel i highlight areas with poor viral surveillance in bat species to date. Hatched regions represent areas where model predictions deviate systematically for the collection of species in that grid cell (see Methods).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Extended Data Figure 1 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 1 | Conceptual model of zoonotic spillover, viral richness, and summary of models. a, Conceptual model of zoonotic spillover showing primary risk factors examined, colour-coded according to generalized additive models used. b, Conceptual model of observed, predicted, and actual viral richness in mammals. c, GAMs used in our study to address specific components of a and b, colour-coded by model. Variables listed with 'or' under each GAM covaried and were provided as competing terms in model selection, and those in bold were included in the best-fit model using all host–virus associations. Significant variables from each best-fit GAM are noted with an asterisk. Zoonotic viral spillover first depends on the underlying total viral richness in mammal populations and the ecological, taxonomic, and life-history traits that govern this diversity (GAM 1). Second, host- and virus-specific factors may facilitate viral spillover. We examine the relative importance of host phylogenetic distance to humans, ecological opportunity for contact, or other species-specific life-history and taxonomic traits (GAM 2), and identify viral traits associated with a higher likelihood of an observed virus being zoonotic (GAM 3). We estimate the total and zoonotic viral richness per host species using GAMs 1 and 2, and calculate the missing viruses and missing zoonoses under a scenario of increased research effort (b, Methods). Owing to imperfect surveillance in both humans and wildlife and biases in viral detection, there may be uncertainty in the exact proportion of viruses that are zoonotic (b, light grey), and also between the actual, or true, viral richness (dotted lines) and the predicted maximum viral richness per host (dashed line).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Extended Data Figure 9 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 9 | Order-level phylogenies showing residuals from zoonoses model. a–e, Subtrees from cytochrome b maximum likelihood phylogeny for 558 mammal species (constrained to order-level topology of mammal supertree) for bats (a), carnivores (b), even-toed ungulates (c), rodents (d) and primates (e). Species included have at least one described virus association and available genetic data. Wildlife species names and terminal branches are colour-coded by the residuals (predicted minus observed) from the best-fit GAM to predict the number of zoonotic viruses using all data. Species with residual values between −1 and 1 (black) are accurately predicted within one virus. Warm colours represent species with positive residuals (orange&gt;1 to 3; red&gt;3). Cool colours represent species with negative residuals (green &lt;−1 to−3; blue&lt;−3). Marine mammals, domestic animals, and species with missing data and not included in the best-fit models are shown in grey.

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 2 in Host and viral traits predict zoonotic spillover from mammals

Figure 2 | Host traits that predict total viral richness (top row) and proportion of zoonotic viruses (bottom row) per wild mammal species. Partial effect plots show the relative effect of each variable included in the best-fit GAM, given the effect of the other variables. Shaded circles represent partial residuals; shaded areas, 95% confidence intervals around mean partial effect. a–e, Best model for total viral richness includes: a, number of disease-related citations per host species (research effort, log); b, phylogenetic eigenvector regression (PVR) of body mass (log); c, geographic range area of each species (log km2); d, number of sympatric mammal species overlapping with at least 20% area of target species range; and e, mammalian orders. f–i, Best model for proportion of zoonoses includes: f, research effort (log); g, phylogenetic distance from humans (cytochrome b tree constrained to the topology of the mammal supertree28); h, ratio of urban to rural human population within species range; and i, three mammalian orders. Bats are the only order with a significantly larger proportion of zoonotic viruses than would be predicted by the other variables in the all-data model. Three additional mammalian orders, and whether or not a species is hunted, improved the overall predictive power of the best zoonotic virus model but were non-significant and are not shown (see Extended Data Table 1).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Extended Data Figure 8 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 8 | Global distribution of viral and host species richness for rodents (order Rodentia). a, Observed total viral richness (for n = 178 host spp.); b, predicted total viral richness given maximum research effort; c, missing viruses or predicted minus observed total viral richness; d, observed zoonotic viral richness (n = 183);e, predicted zoonotic viral richness given maximum research effort; f, missing zoonoses or predicted minus observed zoonotic viral richness (same as included in Fig. 3f); g, global host species richness for Rodentia (n = 2206);h, host species richness for Rodentia in our database (n = 221); i, rodent species with no described viruses in the literature. Warmer colours (larger values) in c and f highlight areas predicted to be of greatest value for discovering novel viruses or novel viral zoonoses, respectively, in wild rodents. Red/pink colours in panel i highlight areas with poor viral surveillance in rodent species to date. Hatched regions represent areas where model predictions deviate systematically for the collection of species in that grid cell (see Methods).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Extended Data Figure 5 in Host and viral traits predict zoonotic spillover from mammals

Extended Data Figure 5 | Global distribution of viral and host species richness for wild even-toed ungulates (order Cetartiodactyla). a, Observed total viral richness (for n = 70 host spp.); b, predicted total viral richness given maximum research effort; c, missing viruses or predicted minus observed total viral richness; d, observed zoonotic viral richness (n = 70);e, predicted zoonotic viral richness given maximum research effort; f, missing zoonoses or predicted minus observed zoonotic viral richness (same as included in Fig. 3c); g, global host species richness for Cetartiodactyla (n = 229);h, host species richness for Cetartiodactyla in our database (n = 105);i, species of the order Cetartiodactyla with no described viruses in the literature. Warmer colours (larger values) in c and f highlight areas predicted to be of greatest value for discovering novel viruses or novel viral zoonoses, respectively, in even-toed ungulates. Red/pink colours in panel i highlight areas with poor viral surveillance in even-toed ungulates species to date. Hatched regions represent areas where model predictions deviate systematically for the collection of species in that grid cell (see Methods).

opencc-by-4.0Jun 2017View details →
zenodo40/100

Combining Policies to Reduce the Spread of Viral Misinformation Online

<p>Data were collected as part of the Election Integrity Partnership. Instances of potential misinformation were flagged as tickets. These were reviewed and categorized as misinformation if they made false claims related to election integrity. Details on the collection methods of the EIP can be found in our Report. These tickets were grouped together into qualitatively similar incidents. For example, tickets regarding false narratives about the use of Benford&#39;s law to detect fraud in Wisconsin became an incident. For each incident, search terms and appropriate data-ranges were determined to query our database.&nbsp;</p> <p>Our full database consisted of all tweets matching an evolving set of keywords, collected in real time, using the Twitter API. To maintain user privacy, we are providing data segmented into events and aggregated into 5-minute blocks of time. This should be sufficient for replicating our findings (predicated on the aggregation and segmentation). In order to permit analysis under various user-removal conditions, we have provided multiple versions of this dataset with users removed according to the conditions evaluated in the manuscript. We encourage anyone with the need for more granular data or alternate conditions to reach out to the University of Washington Center for an Informed Public.</p>

opencc-by-4.0Apr 2022View details →
dryad40/100

Virus classification for viral genomic fragments using PhaGCN2

<p>Viruses are the most ubiquitous and diverse entities in the biome. Due to the rapid growth of newly identified viruses, there is an urgent need for accurate and comprehensive virus classification, particularly for novel viruses. Here, we present PhaGCN2, which can rapidly classify the taxonomy of viral sequences at family level and supports the visualization of the associations of all families. We evaluate the performance of PhaGCN2 and compare it with the state-of-the-art virus classification tools, such as vConTACT2, CAT, and VPF-Class, using the widely accepted metrics. The results show that PhaGCN2 largely improves the precision and recall of virus classification, increases the number of classifiable virus sequences in the Global Ocean Virome dataset (v2.0) by 4 times, and classifies more than 90% of the Gut Phage Database. PhaGCN2 makes it possible to conduct high-throughput and automatic expansion of the database of the International Committee on Taxonomy of Viruses.</p>

opencc-zeroApr 2022View details →
dryad40/100

Deep-sequencing of viral genomes from treatment-naive HIV-infected persons shows positive association between intrahost genetic diversity and viral load

<p><strong><span>Background:</span></strong><span> Infection with human immunodeficiency virus type 1 (HIV) typically results from transmission of a small and genetically uniform viral population. Following transmission, the virus population becomes more diverse because of recombination and acquired mutations through genetic drift and selection. Viral intrahost genetic diversity remains a major obstacle to the cure of HIV; however, there is a disagreement whether intrahost viral genetic diversification associates positively or negatively with disease progression and progression markers. Viral load is a key progression marker and understanding its relationship to viral intrahost genetic diversity could help design future strategies for HIV monitoring and treatment.</span></p> <p><span><strong>Methods:</strong> </span><span>We analyzed deep-sequenced viral genomes from 2,650 treatment-naive HIV-infected persons to measure the intrahost genetic diversity of 2,447 genomic codon positions as calculated by Shannon entropy. We tested for associations between viral load (VL) and amino acid (AA) entropy accounting for sex, age, race, duration of infection, and HIV population structure.</span></p> <p><strong><span>Results:</span></strong><span><strong> </strong>We confirmed that the intrahost genetic diversity is highest in the <em>env</em> gene. Furthermore, we showed that mean Shannon entropy is significantly associated with VL, especially in infections of &gt;24 months duration. We identified 16 significant associations between VL (p-value&lt;2.0x10<sup>-5</sup>) and Shannon entropy at AA positions which in our association analysis explained 13% of the variance in VL.</span></p> <p><strong><span>Conclusions: </span></strong><span>Our results elucidate that viral intrahost genetic diversity is associated with VL and could be used as a better disease progression marker than HIV consensus sequence variants, especially in infections of longer duration. We emphasize that viral intrahost diversity should be considered when studying viral genomes and infection outcomes.</span></p>

opencc-zeroJun 2022View details →
zenodo40/100

Viral Culture in Early Nineteenth-Century Europe newspaper dataset

<p>Dataset produced during the project Viral Culture in Early Nineteenth-Century Europe.</p> <p>The project traced text reuse by analysing large OCR&#39;d newspaper collections using a BLAST based algorithm. This algorithm produces text clusters.</p> <p>This dataset contains two produced cluster datasets based on two different data collections.</p> <p>For the first dataset, the Austrian ANNO newspaper collection, this dataset contains metadata describing the used newspapers.</p> <p>For the second dataset, German-language newspapers in the Europeana collection, this dataset contains project produced metadata describing the newspapers used by the project, as well as the OCR&#39;s content for these newspaper issues. The OCR is produced with Tesseract OCR from digital page images downloaded from the Europeana services.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record