Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
72
datasets available to search
ShareScore release 0.7.1
Dataset results
72 results for “PacBio”
Genome assemblies of four MDR B. fragilis isolates using PacBio data - supporting the PhD Thesis
<p>Unicycler and Canu assemblies using PacBio data of four MDR B. fragilis isolates.</p> <p>Data supporting the PhD Thesis <em>Epidemiology and Genomics of antimicrobial resistance in the Bacteroides fragilis group </em>by Thomas Vognbjerg Sydenham, The research unit of Clinical Microbiology, Department of Clinical Research, Faculty of Health Sciences, University of Southern Denmark September 2019.</p> <p> </p>
PacBio HiFi de-novo assembled genome and mitochondrial genome for Orbicella faveolata
<p>Final assembly using Funannotate of <i>Orbicella faveolata</i> from PacBio HiFi reads. For full methods please see the publication. </p>
SNCA targeted gDNA: PacBio raw data
<p>The landscape of SNCA transcripts across synucleinopathies: New insights from long reads sequencing analysis</p> <p> </p> <p><strong>gDNA capture using IDT xGen</strong><strong>® Lockdown</strong><strong>® Probes and single-molecule sequencing</strong></p> <p>2µg of each gDNA sample was sheared to 6kb using the Covaris g-TUBE and ligated with barcoded adapters. An equimolar pool of 12-plex barcoded gDNA library (2µg total) was input into the probe based capture with a custom designed SNCA gene panel.</p> <p>A SMRTBell library was constructed using 626ng of captured and re-amplified gDNA (https://www.pacb.com/wp-content/uploads/Procedure-Checklist-%E2%80%93-Multiplex-Genomic-DNA-Target-Capture-Using-IDT-xGen-Lockdown-Probes.pdf). A total of 3 SMRT Cells (6 hour movie) were sequenced on the PacBio Sequel platform using 2.0 chemistry. </p> <p> </p> <p>Barcodes used for gDNA demultiplexing</p> <p>bc1001: gcagtcgaacatgtagctgactcaggtcacCACATATCAGAGTGCG<br> bc1002: gcagtcgaacatgtagctgactcaggtcacACACACAGACTGTGAG<br> bc1003: gcagtcgaacatgtagctgactcaggtcacACACATCTCGTGAGAG<br> bc1004: gcagtcgaacatgtagctgactcaggtcacCACGCACACACGCGCG<br> bc1005: gcagtcgaacatgtagctgactcaggtcacCACTCGACTCTCGCGT<br> bc1006: gcagtcgaacatgtagctgactcaggtcacCATATATATCAGCTGT<br> bc1007: gcagtcgaacatgtagctgactcaggtcacTCTGTATCTCTATGTG<br> bc1008: gcagtcgaacatgtagctgactcaggtcacACAGTCGAGCGCTGCG<br> bc1009: gcagtcgaacatgtagctgactcaggtcacACACACGCGAGACAGA<br> bc1010: gcagtcgaacatgtagctgactcaggtcacACGCGCTATCTCAGAG<br> bc1011: gcagtcgaacatgtagctgactcaggtcacCTATACGTATATCTAT<br> bc1012: gcagtcgaacatgtagctgactcaggtcacACACTAGATCGCGTGT</p> <p> </p>
HiPR-FISH PacBio Sequencing Data
<p>This dataset contains raw fastq files from PacBio sequencing for HiPR-FISH experiments.</p>
SMRT PacBio Sequel sequencing of 12 Silene herbarium specimens
<p>The folder includes reads sequenced with SMRT PacBio Sequel from twelve Silene herbarium specimens. Both demultiplexed and non demultiplexed reads are included. The demultiplexed reads are in folder called "demultiplexed.zip", in files starting with "Barcodenumber_speciesnamesamplingyear.fastq". The barcode number and sequence can be found below, in the description. The non demultiplexed reads are in file called "ps_248_001.ccsreads.fastq.gz".</p> <p>#NebNext Barcodes<br> #Adapter('Barcode 1 (forward)',<br> # start_sequence=('BC01', 'CGTGAT'),<br> # end_sequence=('BC01_rev', 'ATCACG')),<br> #Adapter('Barcode 2 (forward)',<br> # start_sequence=('BC02', 'ACATCG'),<br> # end_sequence=('BC02_rev', 'CGATGT')),<br> #Adapter('Barcode 3 (forward)',<br> # start_sequence=('BC03', 'GCCTAA'),<br> # end_sequence=('BC03_rev', 'TTAGGC')),<br> #Adapter('Barcode 4 (forward)',<br> # start_sequence=('BC04', 'TGGTCA'),<br> # end_sequence=('BC04_rev', 'TGACCA')),<br> #Adapter('Barcode 5 (forward)',<br> # start_sequence=('BC05', 'CACTGT'),<br> # end_sequence=('BC05_rev', 'ACAGTG')),<br> #Adapter('Barcode 6 (forward)',<br> # start_sequence=('BC06', 'ATTGGC'),<br> # end_sequence=('BC06_rev', 'GCCAAT')),<br> #Adapter('Barcode 7 (forward)',<br> # start_sequence=('BC07', 'GATCTG'),<br> # end_sequence=('BC07_rev', 'CAGATC')),<br> #Adapter('Barcode 8 (forward)',<br> # start_sequence=('BC08', 'TCAAGT'),<br> # end_sequence=('BC08_rev', 'ACTTGA')),<br> #Adapter('Barcode 9 (forward)',<br> # start_sequence=('BC09', 'CTGATC'),<br> # end_sequence=('BC09_rev', 'GATCAG')),<br> #Adapter('Barcode 10 (forward)',<br> # start_sequence=('BC10', 'AAGCTA'),<br> # end_sequence=('BC10_rev', 'TAGCTT')),<br> #Adapter('Barcode 11 (forward)',<br> # start_sequence=('BC11', 'GTAGCC'),<br> # end_sequence=('BC11_rev', 'GGCTAC')),<br> #Adapter('Barcode 12 (forward)',<br> # start_sequence=('BC12', 'TACAAG'),<br> # end_sequence=('BC12_rev', 'CTTGTA')),]</p> <p> </p>
Comparing bacterial microbiome composition of Xylocopa species across populations using PacBio 16S rRNA gene sequencing
Open the record for dataset details and reuse information.
Pacbio of Broiler chicken: cecum metagenome
<p>The goal is to sequence and assemble broiler chicken's intestine microbial meta-genome. The data we show is sequenced from cecum microbial by PacBio.</p>
Control panel created from 30-40 Nanopore or PacBio HiFi sequencing data from the Human Pangenome Reference Consortium
<p>This is control panel for <a href="https://github.com/friend1ws/nanomonsv">nanomonsv</a> software, which is expected to exclude many false positives as well as improve computational cost. This is made by aligning 30-40 Nanopore or PacBio HiFi sequencing data from Human Pangenome Reference Consortium (HPRC) to the GRCh38 or CHM13 reference genomes with <a href="https://github.com/lh3/minimap2">minimap2</a> version 2.24.</p> <p><strong>When you use these control panels and publish, do not forget to credit to <a href="https://humanpangenome.org/data-use-protocol/">HPRC</a>!</strong></p> <div> <div> <div> <p>Reference genomes:</p> <ul> <li>GRCh38: <a href="https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/000/001/405/GCA_000001405.15_GRCh38/seqs_for_alignment_pipelines.ucsc_ids/GCA_000001405.15_GRCh38_no_alt_analysis_set.fna.gz">Download GRCh38</a></li> <li>CHM13: <a href="https://s3-us-west-2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/analysis_set/chm13v2.0_maskedY_rCRS.fa.gz">Download CHM13</a> <div> <div> <div> <div> </div> </div> </div> </div> </li> </ul> </div> </div> </div>
Combined high-depth Illumina+PacBio Sequencing of several samples from FDA-ARGOS
<p>The (real) sequencing data is compiled from a concatenation of sequencing runs from Database for Reference Grade Microbial Sequences (FDA-ARGOS). Specifically, the following samples were sequenced with both Illumina and PacBio. The sample accessions are shown below.</p> <pre><code> BioSample Run Platform Organism bases source <chr> <chr> <chr> <chr> <dbl> <chr> 1 SAMN06173354 SRR5409204 ILLUMINA Bacillus anthracis 1778000000 Colorado Serum Co., Anthrax Spore Vaccine 2 SAMN06173354 SRR5409205 PACBIO_SMRT Bacillus anthracis 2420000000 Colorado Serum Co., Anthrax Spore Vaccine 3 SAMN06173356 SRR5448657 ILLUMINA Bacillus circulans 2311000000 swab with brown-gray powder 4 SAMN06173356 SRR5448656 PACBIO_SMRT Bacillus circulans 242000000 swab with brown-gray powder 5 SAMN04875535 SRR4123920 ILLUMINA Elizabethkingia anophelis 1357000000 blood 6 SAMN04875535 SRR4123919 PACBIO_SMRT Elizabethkingia anophelis 2173000000 blood 7 SAMN06173306 SRR5413253 ILLUMINA Escherichia coli O157 3051000000 clinical isolate 8 SAMN06173306 SRR5413252 PACBIO_SMRT Escherichia coli O157 915000000 clinical isolate 9 SAMN06173318 SRR5413272 ILLUMINA Mycobacterium avium subsp. paratuberculosis 1032000000 feces 10 SAMN06173318 SRR5413271 PACBIO_SMRT Mycobacterium avium subsp. paratuberculosis 462000000 feces 11 SAMN07312468 SRR5879398 ILLUMINA Mycobacterium tuberculosis 1054000000 human 12 SAMN07312468 SRR5879396 PACBIO_SMRT Mycobacterium tuberculosis 1854000000 human 13 SAMN04875542 SRR4123931 ILLUMINA Neisseria gonorrhoeae 1053000000 ATCC strain 14 SAMN04875542 SRR4123930 PACBIO_SMRT Neisseria gonorrhoeae 1117000000 ATCC strain </code></pre> <p>Samples were selected with the SRA Run selector. The SraRunTable.txt file was exported containing the metadata for each sample, and fastq-dump from the SRA toolkit was used to write out fastq files for each run, with paired Illumina data being split into separate _1.fastq.gz and _2.fastq.gz files. </p>
IsoQuant graphs from PacBio and ONT Mouse sequencing data
<p>Graph files and simulated ONT BAM files.</p>
PacBio IsoSeq reference transcriptomes for Pinus taeda L.
<p>Fusiform rust disease, caused by the endemic fungus <i>Cronartium quercuum</i> f. sp. <i>fusiforme</i>, is the most damaging disease affecting economically important pine species in the southeast United States. In this report, we detail the genomic localization and sequence-level discovery of candidate race-nonspecific broad-spectrum fusiform rust resistance genes in <i>Pinus taeda </i>L. Two full-sib families, each with ~1000 progeny, were challenged with a complex inoculum consisting of over 150 pathogen isolates. High-density linkage mapping revealed three QTL distributed on two linkage groups. The two QTL on linkage group 2 were additive with respect to their effects on the probability of disease outcome. All three QTL were validated using a population of 2057 cloned pine genotypes in a six-year-old multi-environmental field trial. As a complement to the QTL mapping approach, bulked segregant RNAseq analysis revealed a small number of candidate nucleotide binding leucine rich repeat genes harboring SNP significantly associated with disease resistance. The results of this study demonstrate that single qualitative resistance genes can confer effective resistance against genetically diverse mixtures of an endemic pathogen.</p>
Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data
<p class="MsoNormal">Whole genome sequencing enables us to ask fundamental questions about the genetic basis of adaptation, population structure, and epigenetic mechanisms, but usually requires a suitable reference genome for making sense of the sequence data. While the availability of reference genomes has significantly improvement in both taxonomic coverage and overall quality, this poses a challenge for researchers in determining which reference genome best suits their data. Here we compare the use of two different reference genomes for the three-spined stickleback (<em>Gasterosteus aculeatus</em>), one novel genome from a European individual and the published reference genome of a North American individual. Specifically, we investigate the impact of using a local reference versus one generated from a differentiated population on several commonly used metrics in population genomics. Through mapping genome resequencing data of 60 sticklebacks from across Europe and North America, we confirmed genome quality is an important factor in choosing a reference genome. A local reference genome did offer increased mapping efficiency and genotyping accuracy, likely stemming from the higher similarity in genome sequence and synteny. Despite comparable distributions of the metrics generated across the genome using SNP data (i.e., π, Tajima's D, and FST), window-based statistics using different references resulted in different outlier genes and enriched gene functions. In contrast, the marker-based analysis utilising DNA methylation distributions had a considerably higher overlap in outlier genes and functions when using different reference genomes. Overall, our results highlight how using a local reference genome can increase the resolution of genome scans when multiple similar-quality reference genomes are available. Such results have implications in the detection of signatures of selection.</p>
Pacbio single molecule sequencing
<p>As a leading genomics company, CD Genomics provides next-generation sequencing and bioinformatics services to pharmaceutical and biotech companies, as well as academic and government agencies around the world, relying on advanced sequencing instruments and rich project experience. We are offering long-read sequencing services from PacBio and Oxford Nanopore, giving researchers a wide range of cutting-edge sequencing services to suit the particular needs of any project. Here, we offer long-read sequencing services which are based on <a href="https://longseq.cd-genomics.com/pacbio-smrt-sequencing-technology.html">PacBio single molecule sequencing</a> technology, the Sequel II System. This technology can produce highly accurate (> 99.9% average concordance accuracy) long sequences up to 30 kb long, helping to identify biologically important structural variants, RNA splice site isoforms of cDNA, and provide insights into hard-to-sequence regions of the genome, among others.</p>
PacBio whole-genome sequencing and draft assembly of Stentor coeruleus
Open the record for dataset details and reuse information.
HG002 DNA (PacBio and ONT) and UHRR RNA (ONT) base modification data for minimod
Open the record for dataset details and reuse information.
PacBio IsoSeq reference transcriptomes for Pinus taeda L.
Open the record for dataset details and reuse information.
Gasterosteus aculeatus gynogenetic reference genome and functional annotations version 1 and raw PacBio and Illumina data
Open the record for dataset details and reuse information.
PacBio sequencing output increased through uniform and directional 5-fold concatenation
Open the record for dataset details and reuse information.
PacBio Simulation CRAM files for "Characterization of large-scale structural variants using Linked-Reads" (Part 2 of 2)
<p>Here we propose novel algorithms to characterize large (>40 Kbp) interspersed segmental duplications, (> 80 Kbp) inversions, (> 100 Kbp) deletions, and (> 100 Kbp) translocations using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.</p> <p><br> Our methods rely on split molecule sequence signature that we have previously described. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.</p> <p><br> We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large <br> structural variation. We implement our new algorithms in a new software package, called VALOR2. </p>
PacBio Simulation CRAM files for "Characterization of large-scale structural variants using Linked-Reads" (Part 1 of 2)
<p>Here we propose novel algorithms to characterize large (>40 Kbp) interspersed segmental duplications, (> 80 Kbp) inversions, (> 100 Kbp) deletions, and (> 100 Kbp) translocations using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules.</p> <p><br> Our methods rely on split molecule sequence signature that we have previously described. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous.</p> <p><br> We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large <br> structural variation. We implement our new algorithms in a new software package, called VALOR2. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.