Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
63
datasets available to search
ShareScore release 0.9.0
Dataset results
63 results for “massively parallel sequencing”
Robustness of massively parallel sequencing platforms
<p>The improvements in high throughput sequencing technologies (HTS) made clinical sequencing projects such as ClinSeq and Genomics England feasible. Although there are significant improvements in accuracy and reproducibility of HTS based analyses, the usability of these types of data for diagnostic and prognostic applications necessitates a near perfect data generation. To assess the usability of a widely used HTS platform for accurate and reproducible clinical applications in terms of robustness, we generated whole genome shotgun (WGS) sequence data from the genomes of two human individuals in two different genome sequencing centers. After analyzing the data to characterize SNPs and indels using the same tools (BWA, SAMtools, and GATK), we observed significant number of discrepancies in the call sets. As expected, the most of the disagreements between the call sets were found within genomic regions containing common repeats and segmental duplications, albeit only a small fraction of the discordant variants were within the exons and other functionally relevant regions such as promoters. We conclude that although HTS platforms are sufficiently powerful for providing data for first-pass clinical tests, the variant predictions still need to be confirmed using orthogonal methods before using in clinical applications. </p>
Kinetic sequencing (k-Seq) as a massively parallel assay for ribozyme kinetics: utility and critical parameters
Open the record for dataset details and reuse information.
High frequency of X4/DM-tropic viruses in PBMC samples from HIV-1 recently infected blood donors by massively parallel sequencing: the REDS II Study
<p>Here is a sub-library of the <em>env</em> V3 massively parallel sequencing proviral data generated (by Illumina MiSeq platform) during the early phase of HIV-1 infection in a group of first-time blood donors. Only paired-end reads that encompass the complete V3 region from each dataset were extracted, uploaded and considered for the analysis to avoid artificial generation of <em>in silico</em> chimeras through assembly and to evade inflating the diversity estimates of the V3 region</p>
Massively parallel sequencing data of the HIV-1 pol region generated from the plasma of therapy-naïve chronically infected Brazilian blood donors
<p>The submitted massively parallel sequencing (MPS) data were partial data from the pol region of HIV-1 plasma viruses. Samples were obtained from 18 therapy-naive HIV-1 Brazilian blood donors with longstanding infection. Illumina ultra-deep sequencing technology (MiSeq platform) was used to generate the sequences. </p>
Data and processing scripts for PRISM barcode sequencing data used in "Massively parallel pooled screening reveals genomic determinants of nanoparticle-cell interactions"
<p>Sequencing data for the PRISM barcodes generated after nano-particle treatment is presented in this repository alongside the code to process the sequencing counts to generate the binning probabilities and weighted scores. <br> <br> For the details please see the original publication or the bioarxiv preprint: https://doi.org/10.1101/2021.04.05.438521<br> <br> The raw data is provided in PILOT_DATA_COUNTS.csv and EXPERIMENT_DATA_COUNTS.csv files, for the pilot and the actual experiment. <br> <br> For each of these files an R script is provided to process them, along with the output of the scripts (PILOT_DATA_PROBABILITIES.csv and EXPERIMENT_DATA_PROBABILITIES.csv)</p>
Data from: Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform
Genetic information is a valuable component of biosystematics, especially specimen identification through the use of species-specific DNA barcodes. Although many genomics applications have shifted to High-Throughput Sequencing (HTS) or Next-Generation Sequencing (NGS) technologies, sample identification (e.g., via DNA barcoding) is still most often done with Sanger sequencing. Here, we present a scalable double dual-indexing approach using an Illumina Miseq platform to sequence DNA barcode markers. We achieved 97.3% success by using half of an Illumina Miseq flowcell to obtain 658 base pairs of the cytochrome c oxidase I DNA barcode in 1,010 specimens from eleven orders of arthropods. Our approach recovers a greater proportion of DNA barcode sequences from individuals than does conventional Sanger sequencing, while at the same time reducing both per specimen costs and labor time by nearly 80%. In addition, the use of HTS allows the recovery of multiple sequences per specimen, for deeper analysis of genetic variation in target gene regions.
MADDD-seq, a novel massively parallel sequencing tool for simultaneous detection of DNA damage and mutations
<p>The file "data.tar" contains the output of the MADDD-seq pipepline. There is one sub-folder per sample. For each sample, the most important files are:</p> <ul> <li>max_variants_2.adduct.gtf : A GTF file with the location (and details) about each adduct called by the pipepline</li> <li>max_variants_2.DSC.vcf.gz : A VCF (Variant Call File) with information about mutations called.</li> <li>coverage.rds : pre-computed coverage information in binary format to be loaded in R.</li> </ul> <p>To analyze this data, use the following R files: adducts.R, mutations.R and jason-function-2022-04.R</p> <p> </p> <p>The file "kallisto-h5.tar" contains the output of running Kallisto on the regular RNAseq data (for expression level analysis). To analyze this data, use the following R files: Yeast-MNNG-MGT.Rmd and myDESeq2.R</p> <p> </p> <p>The source code of the R files will need to be modified to point at the location of files on the computer being used. These modifications are pointed by comments in the code and are located towards the start of each file.</p>
Data from: Comparative study of the validity of three regions of 18S-rRNA gene for massively parallel sequencing-based monitoring of the planktonic eukaryote community
Open the record for dataset details and reuse information.
Data from: Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform
Open the record for dataset details and reuse information.
Data from: Within-host competition between Borrelia afzelii ospC strains in wild hosts as revealed by massively parallel amplicon sequencing
Open the record for dataset details and reuse information.
Data from: ASSET: analysis of sequences of synchronous events in massively parallel spike trains
With the ability to observe the activity from large numbers of neurons simultaneously using modern recording technologies, the chance to identify sub-networks involved in coordinated processing increases. Sequences of synchronous spike events (SSEs) constitute one type of such coordinated spiking that propagates activity in a temporally precise manner. The synfire chain was proposed as one potential model for such network processing. Previous work introduced a method for visualization of SSEs in massively parallel spike trains, based on an intersection matrix that contains in each entry the degree of overlap of active neurons in two corresponding time bins. Repeated SSEs are reflected in the matrix as diagonal structures of high overlap values. The method as such, however, leaves the task of identifying these diagonal structures to visual inspection rather than to a quantitative analysis. Here we present ASSET (Analysis of Sequences of Synchronous EvenTs), an improved, fully automated method which determines diagonal structures in the intersection matrix by a robust mathematical procedure. The method consists of a sequence of steps that i) assess which entries in the matrix potentially belong to a diagonal structure, ii) cluster these entries into individual diagonal structures and iii) determine the neurons composing the associated SSEs. We employ parallel point processes generated by stochastic simulations as test data to demonstrate the performance of the method under a wide range of realistic scenarios, including different types of non-stationarity of the spiking activity and different correlation structures. Finally, the ability of the method to discover SSEs is demonstrated on complex data from large network simulations with embedded synfire chains. Thus, ASSET represents an effective and efficient tool to analyze massively parallel spike data for temporal sequences of synchronous activity.
Data from: "You are not what you eat: massive parallel sequencing reveals that gut microbiome is not diet-related in larval Dilophus febrilis (Diptera: Bibionidae)" in Genomic Resources Notes Accepted 1 June 2015 to 31 July 2015
This article documents the public availability of metagenome sequence data from 454 amplicon sequencing of larval dipteran gut (Dilophus febrilis) and their potential food sources dwarf shrub litter (Vaccinium gaultheroides), grass litter (Dactylis glomerata), and cow dung (Bos primigenius taurus).
Data from: Target capture and massively parallel sequencing of ultraconserved elements for comparative studies at shallow evolutionary time scales
Comparative genetic studies of non-model organisms are transforming rapidly due to major advances in sequencing technology. A limiting factor in these studies has been the identification and screening of orthologous loci across an evolutionarily distant set of taxa. Here, we evaluate the efficacy of genomic markers targeting ultraconserved DNA elements (UCEs) for analyses at shallow evolutionary timescales. Using sequence capture and massively parallel sequencing to generate UCE data for five co-distributed Neotropical rainforest bird species, we recovered 776–1516 UCE loci across the five species. Across species, 53–77% of the loci were polymorphic, containing between 2.0 and 3.2 variable sites per polymorphic locus, on average. We performed species tree construction, coalescent modeling, and species delimitation, and we found that the five co-distributed species exhibited discordant phylogeographic histories. We also found that species trees and divergence times estimated from UCEs were similar to the parameters obtained from mtDNA. The species that inhabit the understory had older divergence times across barriers, contained a higher number of cryptic species, and exhibited larger effective population sizes relative to the species inhabiting the canopy. Because orthologous UCEs can be obtained from a wide array of taxa, are polymorphic at shallow evolutionary timescales, and can be generated rapidly at low cost, they are an effective genetic marker for studies investigating evolutionary patterns and processes at shallow timescales.
Rapid discovery of high-affinity antibodies via massively parallel sequencing, ribosome display and affinity screening
<p>Deep screening datasets for experiments conducted in Porebski et al., (2023) Rapid discovery of high-affinity antibodies via massively parallel sequencing, ribosome display and affinity screening. <em>Nat. Biol. Eng., doi: 10.1038/s41551-023-01093-3</em>.</p> <p>Datasets are made available under a CC BY-NC-ND 4.0 licence.</p>
Data from: Phylogenetic affiliation of SSU rRNA genes generated by massively parallel sequencing: new insights into the freshwater protist diversity
Open the record for dataset details and reuse information.
Data from: Sex matters in massive parallel sequencing: Evidence for biases in genetic parameter estimation and investigation of sex determination systems
Open the record for dataset details and reuse information.
Data from: ASSET: analysis of sequences of synchronous events in massively parallel spike trains
Open the record for dataset details and reuse information.
Data from: EPA-ng: massively parallel evolutionary placement of genetic sequences
Open the record for dataset details and reuse information.
Data from: Target capture and massively parallel sequencing of ultraconserved elements for comparative studies at shallow evolutionary time scales
Open the record for dataset details and reuse information.
Data from: Target capture and massively parallel sequencing of ultraconserved elements for comparative studies at shallow evolutionary time scales
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.