Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
320
datasets available to search
ShareScore release 0.7.1
Dataset results
320 results for “NGS”
Simulated NGS read datasets for bacterial pathogenic potential prediction
<p>## Predicting pathogenic potentials from NGS reads: novel bacterial species</p> <p>This repository contains simulated Illumina read datasets for bacterial pathogenic potential prediction and associated metadata extracted from the IMG Database (https://img.jgi.doe.gov/). The reads are 250bp long and were simulated with Mason (https://www.seqan.de/apps/mason/) from genomes downloaded from NCBI. The training-validation-test split was done on the species level to ensure "novelty" of validation and test species. The training sets contain 10 million reads per class, validation sets - 1.25 million reads per class, and test sets - 1.25 million paired reads per class. Additional, imbalanced training sets contain 2.5 million "nonpathogenic" and 17.5 million "pathogenic" reads, keeping the mean covarage constant for all species. The temporal benchmark test set contains reads from 3 additional pathogenic species in the Pantoea genus.</p> <p>## Predicting pathogenic potentials from NGS reads: novel strains of known species</p> <p>The BacPaCS datasets contain reads simulated from the dataset compiled by Barash et al. (https://doi.org/10.1093/bioinformatics/bty928). It this case, the training-validation-test split was done on the strain level (so different strains of the same species may be present in all three sets).</p>
Data set for publication: Determination of Virulence-Associated Genes and Antimicrobial Resistance Profiles in Brucella Isolates Recovered from Humans and Animals in Iran Using NGS Technology
<p>This dataset includes information on resistance profiling, as well as antimicrobial resistance (AMR) genes and virulence-related factors that were identified in <em>Brucella</em> isolates recovered from humans and animals in different regions of Iran using classical phenotyping and next-generation sequencing (NGS) technology.</p>
Simulated NGS datasets for real-time detection of novel pathogens
<p>Datasets based on the <a href="https://doi.org/10.5281/zenodo.3678563">bacterial</a> and <a href="https://doi.org/10.5281/zenodo.4312525">viral</a> simulated NGS datasets. Fastq files correspond to tests sets of those datasets. Basecall files were generated based on the fastq files with an 8nt simulated barcode between the mates of a read pair. The "rn" datasets containg random length subreads (25-250bp) of the original validation and training reads.</p> <p>The Nanopore datasets were resimulated with <a href="https://github.com/liyu95/DeepSimulator">DeepSimulator 1.5</a> (Li et al., 2020) based on the original datasets (i.e. using the same species composition as the original data). The test Nanopore dataset contains full reads (target average length: 8kb) and the training and validation datasets - 250bp subreads.</p>
Simulated NGS read datasets for prediction of novel fungal pathogens and multiple pathogen classes
<p>This repository contains simulated Illumina read datasets for novel fungal pathogen prediction and real-time detection of multiple pathogen classes. They were used to train the models hosted at <a href="https://zenodo.org/record/5711877">https://zenodo.org/record/5711877</a>.<br> The reads were simulated with Mason (<a href="https://www.seqan.de/apps/mason/">https://www.seqan.de/apps/mason/</a>) from genomes downloaded from NCBI, based on metadata stored in a manually curated database (<a href="https://zenodo.org/record/5846345">https://zenodo.org/record/5846345</a>).</p> <p>We provide the following:</p> <p>1) An rds file describing assignment of fungal species from the database (<a href="https://zenodo.org/record/5846345">https://zenodo.org/record/5846345</a>) to training, validation and test sets (TrainValTest_fungi.rds). A second rds file (TrainValTest_temporal.rds) includes species added within 12 weeks after the original datasets were compiled. Those species were used for a temporal benchmark.</p> <p>2) Fungal validation and test sets. Each contains 1.25 million, 250bp-long reads simulated from non-overlapping sets of human ("pathogenic") or non-human ("nonpathogenic") pathogens. The test set contains paired reads ("_1" and "_2" for the first and second mate). The number of reads per species is proportional to the respective genome length. An additional, temporal test set (*temporal*fasta.gz) includes 15 species added after 12 weeks from the consturction of the original datasets.</p> <p>3) Fungal training sets. They contain 250bp-long reads simulated from species not present in the validation or test sets. There are four variants:<br> 3a) "low-coverage, linear" - 20 million reads, number of reads per species proportional to genome length<br> 3b) "low-coverage, logarithmic" - 20 million reads, number of reads per species proportional to the logarithm of genome length ("log")<br> 3c) "high-coverage, linear" - 240 million reads, number of reads per species proportional to genome length ("24")<br> 3d) "high-coverage, logarithmic" - 240 million reads, number of reads per species proportional to the logarithm of genome length ("24log")</p> <p>4) Training, validation and test sets for the multiclass models. They should be used together with the "pathogenic" read sets hosted at <a href="https://zenodo.org/record/4456857">https://zenodo.org/record/4456857</a>. Here, we share sets for two of the four total classes:<br> 4a) The 'non-pathogen' class is a mixture of "nonpathogenic" biacterial and viral read sets, concatenated and downsampled to the original read number (20M for training, 1.25M for validation and test). The training and validation sets contain mixed-length (25-20bp) simulated subreads (original sets hosted here: <a href="https://zenodo.org/record/4456857">https://zenodo.org/record/4456857</a>). The test set contains 250bp long reads based on the test sets from here: <a href="https://zenodo.org/record/3678563">https://zenodo.org/record/3678563</a> and here: <a href="https://zenodo.org/record/4312525">https://zenodo.org/record/4312525</a>; it was also sorted by species.<br> 4b) Mixed-length versions of the "pathogenic" fungal training and validation sets, prepared by random shortening of the "low-coverage" read sets in the "linear" (_rn_) and "logarithmic" (_rn_*log_) flavours.</p> <p>See also the preprint: <a href="https://www.biorxiv.org/content/10.1101/2021.11.30.470625">https://www.biorxiv.org/content/10.1101/2021.11.30.470625</a></p>
NGS datasets for "Paired nicking-mediated COL17A1 reframing for junctional epidermolysis bullosa"
<p>Illumina MiSeq PE250 Nano on-target data for treated immortalized and primary keratinocytes (n=3 each) and PPN-treated single cell clones as well as off-target data for sgRNA 5' and sgRNA 3'.</p>
Example Dataset for npstat: Population genetics from Pooled NGS data NPStat v1: User guide
<p>Example Dataset for npstat to test the program and the different options.</p> <p>The example dataset contains a pileup file with sequences of of the 2L chromosome from fifteen pooled inbreed individuals of <em>Drosophila melanogaster </em>(<span>doi: 10.1038/nature10811</span>). The dataset also contains the sequence reference of the 2L chromosome in fasta format, an outgroup sequence in fasta format of <em>D. yakuba</em> (SRR26246471), a GFF3 annotation file and a file with a brief list of selected SNPs to be analyzed.</p>
Simulated NGS read datasets for novel human virus prediction
<p>This repository contains simulated Illumina read datasets for novel human virus prediction and associated metadata extracted from the Virus Host Database (<a href="https://www.genome.jp/virushostdb/">https://www.genome.jp/virushostdb/</a>). The reads are 250bp long and were simulated with Mason (<a href="https://www.seqan.de/apps/mason/">https://www.seqan.de/apps/mason/</a>) from genomes downloaded from NCBI. The training-validation-test split was done on whole viral sequences to ensure "novelty" of validation and test viruses. The training sets contain 10 million reads per class, validation sets - 1.25 million reads per class, and test sets - 1.25 million paired reads per class. The negative class sets contain reads simulated from chordate-infecting ("cho"), metazoan-infecting ("met"), eukariote-infecting ("euk") and all-nonhuman viruses. The positive class contains human-infecting viruses. The stratified dataset ("strat") contains an equal number of reads from "cho", "met but not cho", "euk but not met" and "all but not euk". </p> <p>Species-level datasets ("humspec", "allspec" and "chospec", with the corresponding fasta and *_species.rds files) are constructed analogously, but ensuring that all viruses of a given species were assigned to either training, val or test set. This is a stricter setting modelling a "novel viral species" scenario while reflecting within-species phenotype diversity.</p> <p>blast_hits.gz contains blast hits of human virome reads form Moustafa et al., 2017 (https://doi.org/10.1371/journal.ppat.1006292) blasted against our training database (see paper for details). In the second column you can find the matched label and the accession number of the matched reference. blast_labels_complete.gz contains extracted labels for all virome reads, including those without any matches. Note: one of the read headers (>3c8ac47039d32b11c8fe23f588e444e9) from Moustafa et al. is slightly corrupted with null characters. You can remove them with sed 's/\x0//g' or equivalent.</p> <p> </p>
Raw NGS data for the study 'Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing'
<p>This file contains the original next-generation sequencing data (raw sequences in fastq format) which were analyzed in the study titled: "Spouse-to-spouse Transmission and Evolution of Hypervariable Region 1 and 5’ Untraslated Region of Hepatitis C Virus Analyzed by Next-generation Sequencing".</p> <p> </p> <p> </p>
NGS in the clinical microbiology settings
<p>This dataset contains bam files from 16s rRNA targeted NGS sequencing from a variety of clinical specimens. The study included 55 specimens from 51 patients. We compared targeted NGS to Sanger sequencing in clinical samples submitted for Sanger sequencing. </p> <p>NGS: We sequenced the amplified V1-2 or V4 regions of 16s rRNA gene with Ion S5™ System (Thermo Fisher Scientific). Data were analyzed using the Ion Reporter bioinformatics Software pipeline (Thermo Fisher Scientific), using a threshold of 1000 mapped reads for designating significant pathogens. BAM files uploaded to the Ion Reporter were mapped to the Silva 138 SSU database.</p>
Implementation of Genomic Variant Calling Using GATK4, SPARK, WDL, CROMWELL and DOCKER Over Simulated Ebola NGS Dataset.
<p>Ebola genome is manually mutated to contain non-structural as well as structural variants. One ebola genome contains non-structural variants - 10 SNPs, 10 INDELs, 05 TRANSLOCATIONs, 05 INSERSIONs and their reverse complements. Similarly, other two set of mutated genomes contain structural variants. Each set contains seven mutated ebola genome each one for large deletion, insertion, duplication, translocation, inversion, complex variant1 (consecutive three mutations - insertion, duplication and deletion) and complex variants2 (consecutive three mutations - deletion, duplication and deletion). All insertions are novel sequence insertion.</p> <p> </p>
Fig. 5. Specimen NGS−F−V−258 in New evidence on the taphonomic context of the Ediacaran Pteridinium
Fig. 5. Specimen NGS−F−V−258 collected on Farm Aar. Multiple membrane−like surfaces are visible. These are interpreted as flexible organic structures, see discussion for more information.
Fig. 11 in NGS-barcodes, haplotype networks combined to external morphology help to identify new species in the mangrove genus Ngirhaphium Evenhuis & Grootaert, 2002 (Diptera: Dolichopodidae: Rhaphiinae) in Southeast Asia
Fig. 11. Haplotype network of Ngirhaphium sivasothii. Thailand: Krabi, Phangnga, Satun; Singapore: Sarimbun, Pulau Tekong, Pulau Ubin, Labrador, Semakau Island, and Sungei Buloh.
Fig. 10 in NGS-barcodes, haplotype networks combined to external morphology help to identify new species in the mangrove genus Ngirhaphium Evenhuis & Grootaert, 2002 (Diptera: Dolichopodidae: Rhaphiinae) in Southeast Asia
Fig. 10. Haplotype network of Ngirhaphium murphyi. Singapore: Pulau Ubin, Mandai Sungei Buloh; Thailand: Satun and Krabi.
Fig. 12 in NGS-barcodes, haplotype networks combined to external morphology help to identify new species in the mangrove genus Ngirhaphium Evenhuis & Grootaert, 2002 (Diptera: Dolichopodidae: Rhaphiinae) in Southeast Asia
Fig. 12. Compilation of male terminalia of Ngirhaphium in lateral view. A, Ngirhaphium sivasothii left side with ventral surstylus removed; cerci dorsally; B, Ngirhaphium murphyi left side with ventral surstylus removed; cerci and dorsal surstyli dorsally; C, Ngirhaphium caeruleum left side with ventral surstylus removed; cerci and dorsal surstyli dorsally; D, Ngirhaphium meieri, new species left side with ventral surstylus removed; E, Ngirhaphium chutamasae left side with ventral surstylus removed; F, Ngirhaphium thaicum, new species right side. Scale = 0.1 mm.
Fig. 8 in NGS-barcodes, haplotype networks combined to external morphology help to identify new species in the mangrove genus Ngirhaphium Evenhuis & Grootaert, 2002 (Diptera: Dolichopodidae: Rhaphiinae) in Southeast Asia
Fig. 8. Ngirhaphium meieri, new species, holotype male terminalia: A, Lateral view of genital capsule with left ventral surstylus removed; B, Cerci dorsally; C. Left ventral surstylus; D, Ventral view of genital capsule. Abbreviations: ae = aedeagus; c = cercus; ds = dorsal surstylus; hy = hypandrium; sp = sperm pump; vs = ventral surstylus. Scale = 0.1 mm.
Fig. 5 in NGS-barcodes, haplotype networks combined to external morphology help to identify new species in the mangrove genus Ngirhaphium Evenhuis & Grootaert, 2002 (Diptera: Dolichopodidae: Rhaphiinae) in Southeast Asia
Fig. 5. Ngirhaphium thaicum, new species male terminalia (27_009) A, epandrium left side; B, cerci dorsal view; C, left surstylus inside view. Scale = 0.1 mm.
Fig. 2 in NGS-barcodes, haplotype networks combined to external morphology help to identify new species in the mangrove genus Ngirhaphium Evenhuis & Grootaert, 2002 (Diptera: Dolichopodidae: Rhaphiinae) in Southeast Asia
Fig. 2. Ngirhaphium caeruleum-complex. View of inside of the dorsal surstylus of the left side and the fused dorsal and ventral surstyli of the right side. Scale = 0.1mm. A, green caeruleum from Tutong, Brunei (ZRCBDP0066395) at 1.6 % from the Semakau population; B, green caeruleum from Pulau Tekong, Singapore (ZRCBDP0001462) at 0.6% from the Semakau population; C, green caeruleum from Pulau Ubin, Singapore (ZRC_BDP_0084430) at 0.6% from the Semakau population; D, blue caeruleum from the type locality on Semakau Island, Singapore (ZRCBDP0118762); E, green thaicum, new species from Surat Thani, Thailand (24-018) at 4.2% from the Semakau population; F, green thaicum, new species from Cambodia (JP3C_Ngi-cambodiensis_Misc002) within the variability of the southern Thailand populations.
Fig. 1 in NGS-barcodes, haplotype networks combined to external morphology help to identify new species in the mangrove genus Ngirhaphium Evenhuis & Grootaert, 2002 (Diptera: Dolichopodidae: Rhaphiinae) in Southeast Asia
Fig. 1. Haplotype network of the Ngirhaphium caeruleum-complex. Brunei: Berambang, Tutong, Labu; Singapore: Pulau Ubin, Pulau Tekong, Semakau Island; Thailand: Chumphon, Surat Thani, and Cambodia.
dataset related to article "Multiple Genetic Rare Variants in Autism Spectrum Disorders: A Single-Center Targeted NGS Study"
<p>dataset contains: NGS data (.vcf; .bam; .bam.bai) of all 40 ASD patients analysed in the study at title</p>
NGS Data Accompanying "Deep Learning Enables Design of Multifunctional Synthetic Human Gut Microbiome Dynamics"
<p>NGS Data Accompanying "Deep Learning Enables Design of Multifunctional Synthetic Human Gut Microbiome Dynamics", currently in review.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.