Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
595
datasets available to search
ShareScore release 0.9.0
Dataset results
595 results for “High-throughput sequencing”
Fig. 1 in Do we similarly assess diversity with microscopy and high-throughput sequencing? Case of microalgae in lakes
Fig. 1 Location of the sampled lakes in the French Northern Alps
Data from: On the optimal trimming of high-throughput mRNA sequence data
The widespread and rapid adoption of high-throughput sequencing technologies has afforded researchers the opportunity to gain a deep understanding of genome level processes that underlie evolutionary change, and perhaps more importantly, the links between genotype and phenotype. In particular, researchers interested in functional biology and adaptation have used these technologies to sequence mRNA transcriptomes of specific tissues, which in turn are often compared to other tissues, or other individuals with different phenotypes. While these techniques are extremely powerful, careful attention to data quality is required. In particular, because high-throughput sequencing is more error-prone than traditional Sanger sequencing, quality trimming of sequence reads should be an important step in all data processing pipelines. While several software packages for quality trimming exist, no general guidelines for the specifics of trimming have been developed. Here, using empirically derived sequence data, I provide general recommendations regarding the optimal strength of trimming, specifically in mRNA-Seq studies. Although very aggressive quality trimming is common, this study suggests that a more gentle trimming, specifically of those nucleotides whose Phred score < 2 or < 5, is optimal for most studies across a wide variety of metrics.
A Streamlined and High-Throughput Error-Corrected Next-Generation Sequencing Method for Low Variant Allele Frequency Quantitation
<p></p><p>Quantifying mutant or variable allele frequencies (VAFs) of ≤10−3 using next-generation sequencing (NGS) has utility in both clinical and nonclinical settings. Two common approaches for quantifying VAFs using NGS are tagged single-strand sequencing and duplex sequencing. While duplex sequencing is reported to have sensitivity up to 10−8 VAF, it is not a quick, easy, or inexpensive method. We report a method for quantifying VAFs that are ≥10−4 that is as easy and quick for processing samples as standard sequencing kits, yet less expensive than the kits. The method was developed using PCR fragment-based VAFs of Kras codon 12 in log10 increments from 10−5 to 10−1, then applied and tested on native genomic DNA. For both sources of DNA, there is a proportional increase in the observed VAF to input VAF from 10−4 to 100% mutant samples. Variability of quantitation was evaluated within experimental replicates and shown to be consistent across sample preparations. The error at each successive base read was evaluated to determine if there is a limit of read length for quantitation of ≥10−4, and it was determined that read lengths up to 70 bases are reliable for quantitation. The method described here is adaptable to various oncogene or tumor suppressor gene targets, with the potential to implement multiplexing at the initial tagging step. While easy to perform manually, it is also suited for robotic handling and batch processing of samples, facilitating detection and quantitation of genetic carcinogenic biomarkers before tumor formation or in normal-appearing tissue.</p><p></p>
Figure 5. The polymorphism sites among the 25 in Biased heteroplasmy within the mitogenomic sequences of Gigantometra gigas revealed by sanger and high-throughput methods
Figure 5. The polymorphism sites among the 25 different cloning sequences of cox1. Weblogo 3.0 was used to show the nucleotide content of 25 cloning sequences of cox1 (Crooks et al., 2004). The abscissa stands for the number of the bases, while the ordinate stands for the proportion of nucleotide content provided by the 25 different cloning sequences in the same position. The sequence length between the two arrows stands for the barcode fragment size of cox1. The black triangles show the polymorphism positions in the 25 different cloning sequences, the red circles show the positions exhibited obvious second-peak in the results of direct Sanger sequencing without cloning, and the yellow stars show the different sites between the results of Sanger and HTS.
Figure 2 in Biased heteroplasmy within the mitogenomic sequences of Gigantometra gigas revealed by sanger and high-throughput methods
Figure 2. The different nucleotides of all 13 PCGs in mitogenomes obtained by Sanger and HTS sequencing. The different sequences of HTS sequencing are separately compared with the consequence of Sanger method. The horizontal axis stands for the nucleotide position, of which the sequences of 13 PCGs are ordered according to the circular mitochondrial DNA from nad2 to nad1 in the clockwise direction. The vertical axis stands for the number of the different sites in the sequences of 13 PCGs, and the different nucleotide in each site of all three HTS sequences compared with Sanger is shown in the corresponding panel.
Supplementary material 1 from: Röder N, Schwenk K (2023) Direct PCR meets high-throughput sequencing – metabarcoding of chironomid communities without DNA extraction. Metabarcoding and Metagenomics 7: e102455. https://doi.org/10.3897/mbmg.7.102455
Overview of chironomid size classes
Supplementary material 2 from: Röder N, Schwenk K (2023) Direct PCR meets high-throughput sequencing – metabarcoding of chironomid communities without DNA extraction. Metabarcoding and Metagenomics 7: e102455. https://doi.org/10.3897/mbmg.7.102455
Composition of the two artificial chironomid communities
Data from: Universal and blocking primer mismatches limit the use of high-throughput DNA sequencing for the quantitative metabarcoding of arthropods
Open the record for dataset details and reuse information.
Data from: Defining the alloreactive T cell repertoire using high-throughput sequencing of mixed lymphocyte reaction culture
Open the record for dataset details and reuse information.
Data from: More affordable and effective noninvasive SNP genotyping using high-throughput amplicon sequencing
Open the record for dataset details and reuse information.
Data from: On the optimal trimming of high-throughput mRNA sequence data
Open the record for dataset details and reuse information.
Data from: Prevention, diagnosis, and treatment of high-throughput sequencing data pathologies
Open the record for dataset details and reuse information.
A Streamlined and High-Throughput Error-Corrected Next-Generation Sequencing Method for Low Variant Allele Frequency Quantitation
Open the record for dataset details and reuse information.
Data from: Improving transcriptome assembly through error correction of high-throughput sequence reads
Open the record for dataset details and reuse information.
Data from: SSR_pipeline: a bioinformatic infrastructure for identifying microsatellites from paired-end Illumina high-throughput DNA sequencing data
Open the record for dataset details and reuse information.
Data from: Robust DNA isolation and high-throughput sequencing library construction for herbarium specimens
Open the record for dataset details and reuse information.
Data from: Sequence Capture using PCR-generated Probes (SCPP): a cost-effective method of targeted high-throughput sequencing for non-model organisms
Open the record for dataset details and reuse information.
High-throughput discovery of phage receptors using transposon insertion sequencing of bacteria
Open the record for dataset details and reuse information.
Data from: Large-scale biomonitoring of remote and threatened ecosystems via high-throughput sequencing
Open the record for dataset details and reuse information.
A Multiplexed, Next-Generation Sequencing Platform for High-Throughput Detection of SARS-CoV-2 [pilot cohort]
GEO Series GSE160033. synthetic RNA; synthetic construct; Homo sapiens. 862 samples. Type: Other.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.