Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
767
datasets available to search
ShareScore release 0.7.1
Dataset results
767 results for “high throughput sequencing”
Flow diagram for analysis of high-throughput sequencing data
<p>Tex code and resulting pdf image, summarising the data processing pipeline of high-throughput sequencing data (fastq format files), through mapping the data to a reference genome, and then discovery and genotyping of sequence variants. The latter stage uses both 'GATK Haplotype Caller' for smaller variants, such as single-nucleotide polymorphisms and insertion-deletion polymorphisms, and Genomestrip for variants such as deletions and duplications greater than 1000 nucelotide bases in length. Note that the flow diagram is intended to represent what steps were take in the study, and does not necessarily represent the current optimum methods.</p> <p>The manuscript for which this image is a part of can be found open-access at F1000 Research "Whole genome resequencing of a laboratory-adapted <em>Drosophila melanogaster </em>population sample" https://f1000research.com/articles/5-2644/v1 doi: 10.12688/f1000research.9912.1</p> <p> </p>
Identification of grapevine clones via high-throughput amplicon sequencing: a proof-of-concept study VCF files
<p>VCF files used and cited in the article: Identification of grapevine clones via high-throughput amplicon sequencing: a proof-of-concept study</p>
High-Throughput Sequencing of Human Immunoglobulin Variable Regions with Subtype Identification
<p>Raw Illumina MiSeq data in zipped FASTQ format. The data set includes demultiplexed samples from three different time points (_wk*_) of patient ZA159 (159_*), samples from four different preps of a healthy donor (HD1_*), and samples from IgG subtype sorted cells of a healthy donor (HD3_*). Every sample consists of forward (_R1_), reverse (_R2_) and index read 1 (_I1_). </p>
Supplementary Data for "Sequencing the Pandemic: Rapid and High-Throughput Processing and Analysis of COVID-19 Clinical Samples for 21st Century Public Health"
<p>Supplementary material for F1000 methods manuscript. Includes raw sequencing metrics for two COVID sequencing methodologies, as well as a complete cost breakdown for each methodology.</p>
Simultaneous genotyping of snails and infecting trematode parasites using high-throughput amplicon sequencing.
<p>Several methodological issues currently hamper the study of entire trematode communities within populations of their intermediate snail hosts. Here we develop a new workflow using high-throughput amplicon sequencing to simultaneously genotype snail hosts and their infecting trematode parasites. We designed primers to amplify 4 snail and 5 trematode markers in a single multiplex PCR. While also applicable to other genera, we focused on medically and economically important snail genera within the Superorder Hygrophila and targeted a broad taxonomic range of parasites within the Class Trematoda. We tested the workflow using 417 <i>Biomphalaria glabrata </i>specimens experimentally infected with <i>Schistosoma rodhaini</i>, two strains of<i> Schistosoma mansoni</i>,<i> </i>and combinations thereof. We evaluated the reliability of infection diagnostics, the robustness of the workflow, its specificity related to host and parasite identification, and the sensitivity to detect co-infections, immature infections, and changes of parasite biomass during the infection process. Finally, we investigated its applicability in wild-caught snails of other genera naturally infected with diverse trematode assemblages. After stringent quality control the workflow allows the identification of snails to species level, and of trematodes to taxonomic levels ranging from family to strain. It is sensitive to detect immature infections and changes in parasite biomass described in previous experimental studies. Co-infections were successfully identified, opening the possibility to examine parasite-parasite interactions such as interspecific competition. Altogether, these results demonstrate that our workflow provides a powerful tool to analyze the processes shaping trematode communities within natural snail populations.</p>
Figure S1 in Biased heteroplasmy within the mitogenomic sequences of Gigantometra gigas revealed by sanger and high-throughput methods
Figure S1. Map of the Gigantometra gigas mitogenome using Sanger method (GenBank accession number: MF177288). Genes in the outer circle indicate the direction of transcription of the majority strand (J-strand), and those in the inner circle indicate that of the minority strand (N-strand). The GC content, GC skew+, and GC skew- are separately shown in the circle.
Figure 6 in Biased heteroplasmy within the mitogenomic sequences of Gigantometra gigas revealed by sanger and high-throughput methods
Figure 6. Two examples of the heteroplasmic sites in Sanger sequencing which correspond to the differently sequenced sites. Panels A and B indicate the sites at which the second-peak is obviously higher than the third-peak and fourth-peak, and the base state of the second-peak can be obtained by at least one result of HTS. The different fluorescence densities of base situated at np 1923 in the cox1 are shown in the panel A, and the panel B shows the nucleotides with amino acids at np 1923 in the results of Sanger and HTS methods. The nucleotides are C in the results of HTS sequencing, while the corresponding nucleotides are T in the results of Sanger method in both positions, and the different nucleotides lead not to the amino acids changed. Panels C and D indicate the site at the unobvious second-peak, which is slightly higher than the third-peak and fourth-peak, and the base state of the second-peak can also be obtained by at least one result of HTS. Panel C shows the unobvious second-peak at np 7125, and the nucleotide and amino acid of the site in the results of Sanger and HTS methods are shown in panel D. The amino acids are listed using single-letter amino acid abbreviations.
Figure 4. Intraspecific pairwise K2P in Biased heteroplasmy within the mitogenomic sequences of Gigantometra gigas revealed by sanger and high-throughput methods
Figure 4. Intraspecific pairwise K2P distance of G. gigas based on barcode fragment size of cox1 (Sanger). The red boxplot shows the genetic distances of individuals in all three collecting sites, and the boxplots (blue, green, and yellow) separately show the distances of individuals within each place (HNYG, HNDL, and VIET). The pink boxplot shows the distances of the corresponding cox1 sequences obtained by the two sequencing methods. Abbreviation: HNYG—Yinggeling Nature Reserve, Hainan; HNDL— Diaoluoshan Nature Reserve, Hainan; VIET—northern Vietnam.
Figure S5 in Biased heteroplasmy within the mitogenomic sequences of Gigantometra gigas revealed by sanger and high-throughput methods
Figure S5. The coverage of short fragments at each position in the assembly results of HTS. The three results of HTS method were separately used as reference sequences to be mapped back onto the corresponding HTS scaffolds, and the mitochondrial genes were shown below the corresponding coverage. The scale bar had an indicator at the mean coverage level and the coverage for each nucleotide position was indicated by the height of the blue line.
Figure 3 in Biased heteroplasmy within the mitogenomic sequences of Gigantometra gigas revealed by sanger and high-throughput methods
Figure 3. The different nucleotides in the ITS-1 and ITS-2 regions are shown. The result shows the different nucleotides at nucleotide position np 1897 (G nucleotide and T nucleotide) and np 2790 (C nucleotide and T nucleotide) obtained by Sanger and HTS methods.
Figure 1. Gigantometra gigas. A. Female, dorsal view. B. Male, dorsal view. C in Biased heteroplasmy within the mitogenomic sequences of Gigantometra gigas revealed by sanger and high-throughput methods
Figure 1. Gigantometra gigas. A. Female, dorsal view. B. Male, dorsal view. C. The narrow distribution of G. gigas.
Data for: High-throughput profiling of sequence recognition by tyrosine kinases and SH2 domains using bacterial peptide display
<p>Tyrosine kinases and SH2 (phosphotyrosine recognition) domains have binding specificities that depend on the amino acid sequence surrounding the target (phospho)tyrosine residue. Although the preferred recognition motifs of many kinases and SH2 domains are known, we lack a quantitative description of sequence specificity that could guide predictions about signaling pathways or be used to design sequences for biomedical applications. Here, we present a platform that combines genetically-encoded peptide libraries and deep sequencing to profile sequence recognition by tyrosine kinases and SH2 domains. We screened several tyrosine kinases against a million-peptide random library and used the resulting profiles to design high-activity sequences. We also screened several kinases against a library containing thousands of human proteome-derived peptides and their naturally-occurring variants. These screens recapitulated independently measured phosphorylation rates and revealed hundreds of phosphosite-proximal mutations that impact phosphosite recognition by tyrosine kinases. We extended this platform to the analysis of SH2 domains and showed that screens could predict relative binding affinities. Finally, we expanded our method to assess the impact of non-canonical and post-translationally modified amino acids on sequence recognition. This specificity profiling platform will shed new light on phosphotyrosine signaling and could readily be adapted to other protein modification/recognition domains.</p>
Data for: High-throughput profiling of sequence recognition by tyrosine kinases and SH2 domains using bacterial peptide display
Open the record for dataset details and reuse information.
Simultaneous genotyping of snails and infecting trematode parasites using high-throughput amplicon sequencing.
Open the record for dataset details and reuse information.
Antarctic eukaryotic soil diversity of the Prince Charles Mountains revealed by high-throughput sequencing
<p>Analysis scripts and raw data for manuscript titled: "<em>Antarctic eukaryotic soil diversity of the Prince Charles Mountains revealed by high-throughput sequencing</em>" </p>
Application of high-throughput sequencing (HTS) metabarcoding to diatom biomonitoring: Do DNA extraction methods matter?
<p>The 8 benthic samples from Mainland France (stream Edian, stream Aire, lake Geneva), Sweden (stream Dåmman, Agricultural stream, lake Båtkåjaure) and Mayotte (stream Dapani, stream Majimbini) were collected by scraping material from the surface of stones, following the French standard (AFNOR 2007) used in routine biomonitoring programs.DNA was extracted from each sample (2 replicates) using five DNA extraction methods, followed by the amplification of a short rbcL DNA barcode (312bp) specific to diatoms. PCR products were then sequenced in one random direction using the Ion Torrent™ Personal Genome Machine® (PGM) System according to the manufacturer’s instructions. The data file contains one fastq file per library sequenced with the raw DNA reads, as provided by the sequencing platform (demultiplexing performed by the sequencing platform). An excel file is also provided to make the link between the fastq file number and the sample information (sample origin, DNA extraction method used, number of raw reads).</p>
Shark-dust: Application of high-throughput DNA sequencing of processing residues for trade monitoring of threatened sharks and rays
<p>Data repository accompanying manuscript titled of "Shark-dust: Application of high-throughput DNA sequencing of processing residues for trade monitoring of threatened sharks and rays."</p> <p>Prasetyo, A. P., Murray, J. M., Kurniawan, M. F. A. K., Sales, N. G., McDevitt, A. D., & Mariani, S. (2023). Shark-dust: Application of high-throughput DNA sequencing of processing residues for trade monitoring of threatened sharks and rays. Conservation Letters, 16, e12971. https://doi.org/10.1111/conl.12971</p>
A Framework for High-throughput Sequence Alignment using Real Processing-in-Memory Systems
<p>Sequence alignment is a fundamentally memory bound computation whose performance in modern systems is limited by the memory bandwidth bottleneck. Processing-in-memory architectures alleviate this bottleneck by providing the memory with computing competencies. We propose Alignment-in-Memory (AIM), a framework for high-throughput sequence alignment using processing-in-memory, and evaluate it on UPMEM, the first publicly-available general-purpose programmable processing-in-memory system.</p>
Do we similarly assess diversity with microscopy and High Throughput Sequencing? Case of microalgae in lakes.
<p>These are the repository files of the paper "Do we similarly assess diversity with microscopy and High Throughput Sequencing? Case of microalgae in lakes" published in Organisms Diversity and Evolution</p> <p>In these files are given:</p> <p>- the sampling sites descrptions (coordinates, lake names)</p> <p>- species relative abundances obtained with light microscopy for each sampling site</p> <p>- OTUs amounts and relative abundances obtained High-Throughput Sequencing for each sampling site</p> <p>- code correspondence between lake codes and FastQ files codes</p> <p>- FastQ files of the sampling sites</p>
Discovery of tandem and interspersed segmental duplications using high throughput sequencing
<p>We developed novel algorithms to accurately characterize tandem, direct and inverted interspersed segmental duplications using short read whole genome sequencing data sets. We integrated these methods to our TARDIS tool, which is now capable of detecting various types of SVs using multiple sequence signatures such as read pair, read depth and split read. We evaluated the prediction performance of our algorithms through several experiments using both simulated and real data sets. In the simulation experiments, using a 30x coverage TARDIS achieved 96% sensitivity with only 4% false discovery rate. For experiments that involve real data, we used two haploid genomes (CHM1 and CHM13) and one human genome (NA12878) from the Illumina Platinum Genomes set. Comparison of our results with orthogonal PacBio call sets from the same genomes revealed higher accuracy for TARDIS than state of the art methods. Furthermore, we showed a surprisingly low false discovery rate of our approach for discovery of tandem, direct and inverted interspersed segmental duplications prediction on CHM1 (less than 5\% for the top 50 predictions). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.