Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

72

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

72 results for “PacBio”

Learn how ShareScore rates datasets ↗
dryad32/100

Hooded crow genome assembly (PacBio long reads + BioNano optical maps + DoveTail HiC maps)

<p>Structural variation (SV) is an important component of mutations providing the raw material for evolution. Here, we uncover the genome-wide spectrum of intra- and interspecific SV segregating in natural populations of seven songbird species in the genus Corvus. Combining short-read (N = 127) and long-read re-sequencing (N = 31), as well as optical mapping (N = 16), we apply both assembly- and read mapping approaches to detect SV and characterize a total of 220,452 insertions, deletions and inversions. We exploit sampling across wide phylogenetic timescales to validate SV genotypes and assess the contribution of SV to evolutionary processes in an avian model of incipient speciation. We reveal an evolutionary young (~530,000 years) cis-acting 2.25-kb retrotransposon insertion reducing expression of the NDP gene with consequences for premating isolation. Our results attest to the wealth and evolutionary significance of SV segregating in natural populations and highlight the need for reliable SV genotyping.</p>

opencc-zeroJun 2020View details →
dryad32/100

Data from: Rapid allopolyploid radiation of moonwort ferns (Botrychium ; Ophioglossaceae) revealed by PacBio sequencing of homologous and homeologous nuclear regions

Polyploidy is a major speciation process in vascular plants, and is postulated to be particularly important in shaping the diversity of extant ferns. However, limitations in the availability of bi-parental markers for ferns have greatly limited phylogenetic investigation of polyploidy in this group. With a large number of allopolyploid species, the genus Botrychium is a classic example in ferns where recurrent polyploidy is postulated to have driven frequent speciation events. Here, we use PacBio sequencing and the PURC bioinformatics pipeline to capture all homeologous or allelic copies of four long (∼1kb) low-copy nuclear regions from a sample of 45 specimens (25 diploids and 20 polyploids) representing 37 Botrychium taxa, and three outgroups. This sample includes most currently recognized Botrychium species in Europe and North America, and the majority of our specimens were genotyped with co-dominant nuclear allozymes to ensure species identification. We analyzed the sequence data using maximum likelihood (ML) and Bayesian inference (BI) concatenated-data ("gene tree") approaches to explore the relationships among Botrychium species. Finally, we estimated divergence times among Botrychium lineages and inferred the multi-labeled polyploid species tree showing the origins of the polyploid taxa, and their relationships to each other and to their diploid progenitors. We found strong support for the monophyly of the major lineages within Botrychium and identified most of the parental donors of the polyploids; these results largely corroborate earlier morphological and allozyme-based investigations. Each polyploid had at least two distinct homeologs, indicating that all sampled polyploids are likely allopolyploids (rather than autopolyploids). Our divergence-time analyses revealed that these allopolyploid lineages originated recently—within the last two million years—and thus that the genus has undergone a recent radiation, correlated with multiple independent allopolyploidizations across the phylogeny. Also, we found strong parental biases in the formation of allopolyploids, with individual diploid species participating multiple times as either the maternal or paternal donor (but not both). Finally, we discuss the role of polyploidy in the evolutionary history of Botrychium and the interspecific reproductive barriers possibly involved in these parental biases.

opencc-zeroDec 2016View details →
zenodo32/100

A near telomere-to-telomere phased reference assembly for the male mountain gorilla (Gorilla beringei beringei) - Pacbio HIFI reads

<p>The critically endangered mountain gorilla Gorilla beringei beringei faces numerous threats to its survival, highlighting the urgent need for genomic resources to aid conservation efforts. Here, we present a near telomere-to-telomere, haplotype-phased reference genome assembly for a male mountain gorilla generated using Pacbio HiFi and Oxford Nanopore Ultralong data. The resulting assembly exhibits exceptional contiguity, with contig N50 of ~ 95 Mbps for the combined pseudohaplotype (3,540,458,497 bps, and 56.5 Mbps (3.1 Gbps) and 51.0 Mbps (3.2 Gbps) for the maternal and paternal haplotypes and an average QV of 65.15 (error rate = 3.1 x 10-7) and 0% switch errors detected. These represent substantial improvements over most other available primate genomes. This high-quality reference genome provides an invaluable resource for future studies on gorilla evolution, adaptation, and conservation, ultimately contributing to the long-term survival of this iconic species.</p> <p>This read set is comprised of fastqs from Pacbio HIFI (3 runs).</p> <p>A preprint for this work is available at bioRXiv, doi: https://doi.org/10.1101/2024.10.28.620258</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

WGS of E. coli K12 with PacBio HiFi

<p>WGS of E. coli K12 with PacBio HiFi reads for genome assembly.</p>

opencc-by-4.0Feb 2020View details →
dryad32/100

Evaluating Illumina-, Nanopore-, and PacBio-based genome assembly strategies with the bald notothen, Trematomus borchgrevinki

<p>For any genome-based research, a robust genome assembly is required. <em>De novo</em> assembly strategies have evolved with changes in DNA sequencing technologies and have been through at least three phases: i) short-read only, ii) short- and long-read hybrid, and iii) long-read only assemblies. Each of the phases has their own error model. We hypothesized that hidden scaffolding errors in short-read assembly and erroneous long-read contigs degrade the quality of short- and long-read hybrid assemblies. We assembled the genome of <em>T. borchgrevinki</em> from data generated during each of the three phases and assessed the quality problems we encountered. We developed strategies such as k-mer-assembled region replacement, parameter optimization, and long-read sampling to address the error models. We demonstrated that a k-mer-based strategy improved short-read assemblies as measured by BUSCO while mate-pair libraries introduced hidden scaffolding errors and perturbed BUSCO scores. Further, we found that although hybrid assemblies can generate higher contiguity, they tend to suffer from lower quality. In addition, we found long-read-only assemblies can be optimized for contiguity by sub-sampling length-restricted raw reads. Our results indicate that long-read contig assembly is the current best choice and that assemblies from phase I and phase II were of lower quality.</p>

opencc-zeroAug 2022View details →
dryad32/100

Data from: Rapid allopolyploid radiation of moonwort ferns (Botrychium ; Ophioglossaceae) revealed by PacBio sequencing of homologous and homeologous nuclear regions

Open the record for dataset details and reuse information.

publicDec 2018View details →
dryad32/100

Evaluating Illumina-, Nanopore-, and PacBio-based genome assembly strategies with the bald notothen, Trematomus borchgrevinki

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad32/100

Hooded crow genome assembly (PacBio long reads + BioNano optical maps + DoveTail HiC maps)

Open the record for dataset details and reuse information.

publicJun 2020View details →
dryad32/100

Data from: The effects of read length, quality and quantity on microsatellite discovery and primer development: from Illumina to PacBio

Open the record for dataset details and reuse information.

publicFeb 2014View details →
dryad32/100

NBS-LRR PacBio sequencing from a Glycine max diversity panel

Open the record for dataset details and reuse information.

publicOct 2024View details →
zenodo28/100

PacBio amplicon re-sequencing of 62 P. tricornutum genomic loci – processed datasets

<p>Processed datasets from&nbsp;PacBio amplicon re-sequencing of 62 P. tricornutum genomic loci.&nbsp;Raw data are available at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA658511.&nbsp;</p> <p>Available datasets:&nbsp;</p> <p>- reference file for regions selected for amplicon sequencing: <em>Phaeodactylum_tricornutum_amplicon_sequencing_loci.fa</em></p> <p>- final .bam files containing processed PacBio sequencing reads&nbsp;aligned to the reference:</p> <p><em>PacBio_amplicon_seq_T1.bam</em> &nbsp;and&nbsp;<em>PacBio_amplicon_seq_T6.bam</em>&nbsp;</p> <p>- . table files with the position, reference and alternative allele for reliable biallelic SNPs selected in ILLUMINA sequencing of the culture at T1:&nbsp;</p> <p><em>P_tricornutum_PacBio_amplicon_sequencing_T1_SNPs.table</em> and&nbsp;<em>P_tricornutum_PacBio_amplicon_sequencing_T6_SNPs.table</em></p> <p>Re-sequencing of 62 endogenous P. tricornutum loci selected in a genome-wide analysis of haplotype diversity. The goal was to determine the number of haplotypes per locus and the appearance of new haplotypes over time. The length of the sequenced loci was 2kb (+/- 5%).&nbsp; Loci were amplified by emulsion PCR on the same culture harvested in two time points: five loci were amplified one month (T1) and all loci were amplified 6 months (T6) after the start of the culture from a single cell. Plasmids containing cloned GFP or YFP were amplified separately as a control for random errors. Control reactions for artificial haplotypes detection consisted of mixed CFP with YFP or CFP with GFP. Amplicons were pooled together into two samples. Sample PacBio_AS_T1 contained five P. tricornutum endogenous amplicons from DNA harvested at T1 time point, GFP amplified separately and CFP+YFP amplified in one reaction. Sample PacBio_AS_T6 contained 63 P. tricornutum endogenous amplicons from DNA harvested at T6 time point, YFP amplified separately and CFP+GFP amplified in one reaction. Samples were mixed in 1:9 PacBio_AS_T1: PacBio_AS_T6 ratio before sequencing on one PacBio Sequel SMRT cell.</p>

opencc-by-4.0Aug 2020View details →
zenodo28/100

Metagenomic Std ATCC MSA-1003: PacBio HiFi Reads FASTQ

<p>WGS of mock metagenomic community ATCC MSA-1003 using PacBio HiFi (CCS) Sequencing</p>

opencc-by-4.0Apr 2020View details →
zenodo28/100

supplemental files for Assembly and analysis of sequence from a spring and winter type Camelina sativa by whole genome PacBio HiFi technologies

<p><span>Supplemental files for Assembly and analysis of sequence from a spring and winter type <em>Camelina sativa</em> by whole genome PacBio HiFi technologies</span></p>

opencc-by-4.0Jan 2024View details →
zenodo28/100

SNCA targeted cDNA: PacBio Iso-Seq raw data

<p><strong>The landscape of <em>SNCA</em> transcripts across </strong><strong>synucleinopathies</strong><strong>: New insights from long reads sequencing analysis</strong></p> <p><strong>cDNA capture using IDT xGen</strong><strong>&reg; Lockdown</strong><strong>&reg; Probes and single-molecule Isoform-Sequencing (Iso-Seq)</strong></p> <p>100-150ng of total RNA per reaction was reverse transcribed using the Clontech SMARTer cDNA synthesis kit and 12 sample specific barcoded oligo dT (with PacBio 16mer barcode sequences, see Supplementary Methods).&nbsp; Three reverse transcription (RT) reactions were processed in parallel for each sample. PCR optimization was used to determine the optimal amplification cycle number for the downstream large-scale PCR reactions. A single primer (primer IIA from the Clontech SMARTer kit 5&rsquo; AAG CAG TGG TAT CAA CGC AGA GTA C 3&rsquo;) was used for all PCR reactions post-RT. Large scale PCR products were purified separately with 1X AMPure PB beads and the bioanalyzer was used for QC. An equimolar pool of 12-plex barcoded cDNA library (1&micro;g total) was input into the probe based capture with a custom designed SNCA gene panel.</p> <p>A SMRTBell library was constructed using 874ng of captured and re-amplified cDNA (https://www.pacb.com/wp-content/uploads/Procedure-Checklist-%E2%80%93-cDNA-Capture-Using-IDT-xGen-Lockdown-Probes.pdf).&nbsp; One SMRT Cell (6 hour movie) was sequenced on the PacBio Sequel platform using 2.0 chemistry.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>The 5&rsquo; primer is identical for all patient samples: AAGCAGTGGTATCAACGCAGAGTACATGGG</p> <p>The 3&rsquo; primer for the 12 patient samples are below:</p> <p>PD-1: CGCACTCTGATATGTGGTACTCTGCGTTGATACCACTGCTT</p> <p>PD-2: CTCACAGTCTGTGTGTGTACTCTGCGTTGATACCACTGCTT</p> <p>PD-3: CTCTCACGAGATGTGTGTACTCTGCGTTGATACCACTGCTT</p> <p>PD-4: CGCGCGTGTGTGCGTGGTACTCTGCGTTGATACCACTGCT</p> <p>N-1: ACGCGAGAGTCGAGTGGTACTCTGCGTTGATACCACTGCTT</p> <p>N-2: ACAGCTGATATATATGGTACTCTGCGTTGATACCACTGCTT</p> <p>N-3: CACATAGAGATACAGAGTACTCTGCGTTGATACCACTGCTT</p> <p>N-4: CGCAGCGCTCGACTGTGTACTCTGCGTTGATACCACTGCTT</p> <p>DLB-1: TCTGTCTCGCGTGTGTGTACTCTGCGTTGATACCACTGCTT</p> <p>DLB-2: CTCTGAGATAGCGCGTGTACTCTGCGTTGATACCACTGCTT</p> <p>DLB-3: ATAGATATACGTATAGGTACTCTGCGTTGATACCACTGCTT</p> <p>DLB-4: ACACGCGATCTAGTGTGTACTCTGCGTTGATACCACTGCTT</p>

opencc-by-4.0Nov 2018View details →
zenodo28/100

Supplementary material 1 from: Gueidan C, Elix JA, McCarthy PM, Roux C, Mallen-Cooper M, Kantvilas G (2019) PacBio amplicon sequencing for metabarcoding of mixed DNA samples from lichen herbarium specimens. MycoKeys 53: 73-91. https://doi.org/10.3897/mycokeys.53.34761

: Data type: measurement

opencc-zeroJun 2019View details →
zenodo28/100

PacBio sequencing E. coli doi: 10.1099/mgen.0.000352

<p>PacBio sequencing E. coli doi: 10.1099/mgen.0.000352</p>

openSep 2024View details →
dryad28/100

Data from: Next-generation polyploid phylogenetics: rapid resolution of hybrid polyploid complexes using PacBio single-molecule sequencing

Difficulties in generating nuclear data for polyploids have impeded phylogenetic study of these groups. We describe a high-throughput protocol and an associated bioinformatics pipeline (PURC: "Pipeline for Untangling Reticulate Complexes") that is able to generate these data quickly and conveniently, and demonstrate its efficacy on accessions from the fern family Cystopteridaceae. We conclude with a demonstration of the downstream utility of these data by inferring a multilabeled species tree for a subset of our accessions. We amplified four ~1kb-long nuclear loci and sequenced them in a parallel-tagged amplicon sequencing approach using the PacBio platform. PURC infers the final sequences from the raw reads via an iterative approach that corrects PCR and sequencing errors and removes PCR-mediated recombinant sequences (chimeras). We generated data for all gene copies (homeologs, paralogs, and segregating alleles) present in each of three sets of 50 mostly-polyploid accessions, for four loci, in three PacBio runs (one run per set). From the raw sequencing reads PURC was able to accurately infer the underlying sequences. This approach makes it easy and economical to study the phylogenetics of polyploids, and in conjunction with recent analytical advances, facilitates investigation of broad patterns of polyploid evolution.

opencc-zeroDec 2015View details →
zenodo28/100

Raw data from Pacbio, assemblies and annotation of Xen31 and Xen36 strains

<p>This deposit contains:</p> <ol> <li>raw pacbio data related to two samples of&nbsp;Staphylococcus aureus subsp. aureus MRSA252 (strains Xen31 and&nbsp;Xen36).</li> <li>assemblies in forms of fasta files</li> <li>annotations of the two main chromosomes</li> </ol> <p>The sequencing was performed on Pacbio Sequel on different so-called smartcells. Full details about library prepration, sequencing runs, etc can be found on array express under number&nbsp;E-MTAB-12210 (https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-12210).&nbsp;</p> <p>Because there were two&nbsp;smartcells, there are two files per samples. The two files xen31_A.subreads.bam and xen31_B.subreads.bam are therefore the raw data of the strain Xen31. similarly&nbsp;&nbsp;the two files xen36_A.subreads.bam and xen36_B.subreads.bam are the raw data of the strain Xen36. These 2 files should be merged before analysis. Be aware that the data is not hifi data nor CCS (corrected). CCS files are provided on the array express link here above.</p> <p>Extra information about the quality assessments (in terms of coverage) of the assemblies can be found here :&nbsp;https://github.com/biomics-pasteur-fr/manuscript_B4410/edit/main/README.md and in&nbsp;a paper to be provided.</p> <p>In this deposit you can find the&nbsp;</p> <ol> <li>the four raw files named e.g. xen31_A.subreads.bam for the first smrtcell of strain 31</li> <li>the results of the assemblies using Canu1.6 . In both strains a main contig corresponding to the entire S. aureus chromosome was obtained e.g. xen31_chromosome.fa Extra contigs are stored in xen31_other_contigs.fa</li> <li>annotations of the main chromosome were performed with prokka 1.14.5 and genbank and GFF format are provided here.</li> </ol>

opencc-by-4.0Jan 2023View details →
dryad28/100

PacBio HiFi based haplotype-aware assemblies of tomato hybrid varieties Funtelle and Maxeza

<p>Modern commercial varieties of tomato (<em>Solanum lycopersicum</em>) are typically F1 hybrids that are genetically heterozygous. Here we generated haplotype-aware assemblies of two different tomato commercial hybrids (Funtelle and Maxeza) using PacBio HiFi reads. The HiFi data was assembled using the Hifiasm assembler allowing for the generation of contigs that distinguish the two parental haplotypes (haplotype-aware assembly). Reference based scaffolding was used to generate the chromosome-scale assemblies available here. It should be noted that although the raw assembly manages to fully distinguish haplotypes we did not test whether the working reference sequence we make available here is fully phased at the chromosome level.</p>

opencc-zeroOct 2023View details →
dryad28/100

PacBio HiFi based haplotype-aware assemblies of tomato hybrid varieties Funtelle and Maxeza

Open the record for dataset details and reuse information.

publicOct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record