Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

233

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

233 results for “long-read”

Learn how ShareScore rates datasets ↗
zenodo32/100

Expanded dynamic methylome and quantitative trait detection by long-read epigenome profiling of personal DNA

<p>Scripts and methylation frequencies for "Expanded dynamic methylome and quantitative trait detection by long-read epigenome profiling of personal DNA"&nbsp;manuscript.</p>

opencc-by-4.0Mar 2024View details →
dryad32/100

Data from: Structural variation and its potential impact on genome instability: novel discoveries in the EGFR landscape by long-read sequencing

<p>Studies of structural variation (SV) have been challenging due to technological contraints. With the advent of third generation (long-read) sequencing technology, exploration of longer stretches of DNA not easily examined previously has been made possible. In the present study, we utilized third generation (long-read) sequencing techniques to examime SV in the <em>EGFR </em>landscape of four haplotypes derived from two human samples. We analyzed the <em>EGFR</em> gene and its landscape (+/- 500,000 base pairs) using this sequencing approach and were able to identify regions of non-coding DNA which had relatively high similarity to the most common activating <em>EGFR</em> mutation in non-small cell lung cancer. We discovered that reverse complements to the exon 19 deletion mutation which had at least 60% homology to the <em>EGFR</em> exon 19 canonical deletion and were within ± 421,000 bp of the deletion varied across the five haploid genomes examined (4 patient landscapes and hg38). Although the sample size is limited in this study, the estimated variation observed in genomic stability between the five <em>EGFR</em> haplotypes examined is novel and encourages further work to examine structural variation in larger cohorts.</p>

opencc-zeroAug 2021View details →
zenodo32/100

Long-read genome sequencing accelerated the cloning of Pm69 by resolving the complexity of a rapidly evolving resistance gene cluster in wheat

<p>Oxford Nanopore assembly of&nbsp;<em>Triticum turgidum</em>&nbsp;ssp.&nbsp;<em>dicoccoides, </em>cv. G305-3M.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Long-read transcriptome data (Iso-Seq) of the human and mouse brain

<p>Datasets from&nbsp;&quot;<strong>Full-length transcript sequencing of human and mouse cerebral cortex identifies widespread isoform diversity and alternative splicing</strong>&quot;, SK.Leung, A.Jeffries. et al. (2021)</p> <p>Deposited files are generated from running Cupcake,&nbsp;SQANTI2 (v7.4) and filter.&nbsp;<br> Note, only the files generated from SQANTI2 filtering are deposited.&nbsp;<br> To re-run SQANTI, use the files in cupcake_collapse folder as input.&nbsp;</p> <p>Please refer to code (https://github.com/SziKayLeung/Whole_Transcriptome_Paper) for more information.&nbsp;</p> <p>Datasets:&nbsp;<br> - AdultCTX: Adult human prefrontal cortex tissue (n = 4) merged dataset<br> - FetalCTX: Fetal human prefrontal cortex tissue (n = 3) merged dataset<br> - FetalHIP: Fetal human hippocampus tissue (n = 2, subset of FetalCTX) merged dataset<br> - FetalSTR: Fetal human striatum tissue (n = 2, subset of FetalCTX) merged dataset<br> - MouseCTX: Mouse entorhinal cortex tissue (n = 12) merged dataset</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

The data used in correspondence on the article entitled "Reply: Correspondence on NanoVar's performance outlined by Jiang T. et al. in 'Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluation'."

<p>The data used in correspondence on the article entitled &ldquo;Reply: Correspondence on NanoVar&rsquo;s performance outlined by Jiang T. et al. in &lsquo;Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluation&rsquo;.&rdquo;. It contains the benchmarking results evaluated by&nbsp;Truvari. The running log files of NanoVar are also included in this repository.</p>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Purification of High Molecular Weight DNA for Long-Read Sequencing Using a High-Salt Gel Electroelution Trap

<p><strong>Figure 3. Yield and purity of HMW DNA obtained from difficult samples using the method proposed in this study</strong>.</p> <p>(A) The proposed method extracts more HMW DNA from a complex soil sample than a commercial column purification kit. Shown is a negative image of an ethidium bromide-stained agarose gel. Lane 1: 1/10 aliquot of ~90 ng of HMW DNA isolated using the E.Z.N.A. soil DNA extraction kit from a soil sample containing ~1,5 &mu;g of total DNA (HMW DNA yield around 6%). Lane 2: ~100 ng of CTAB-extracted DNA from the same soil sample. ~10 &mu;g of this crude DNA preparation was used as input for HMW DNA purification using the proposed method. Lane 3: 1/100 aliquot of ~3 &mu;g of HMW DNA isolated using the proposed method from ~10 &mu;g of the CTAB-extracted DNA (HMW DNA yield around 30%).</p> <p>(B) The proposed method yields high-purity HMW DNA from a complex plant sample, as determined by agarose gel electrophoresis. Lane 1: molecular weight marker (GeneRuler DNA ladder, Thermo Fisher Scientific). Lane 2: crude nucleic acid preparation extracted with SDS/Proteinase K from <em>Zingeria trichopoda</em> leaves, which served as an input for HMW DNA purification using the proposed method. Lane 3: purified HMW DNA. Note the absence of low-molecular-weight nucleic acids and heavy covalent complexes in the purified sample.</p> <p>(C) Same as (B) except that crude, CTAB-extracted DNA from a complex soil sample was used as an input for HMW DNA purification. Note the absence of a continuous smear of fragmented DNA as well as heavy covalent complexes in the purified sample (lane 3). Molecular weight marker sizes are indicated in base pairs to the left of each panel.</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Tradeoffs in alignment and assembly-based methods for structural variant detection with long-read sequencing data

<p>Source data for&nbsp;the paper &quot;Tradeoffs in alignment and assembly-based methods for structural variant detection with long-read sequencing data&quot;</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

isoMiGA: Long-read isoforms discovered in human microglia

<p>https://github.com/RajLabMSSM/isoMiGA&nbsp;&nbsp;</p> <p>Novel isoform discovery from long-read RNA-seq of 30 human microglia samples with PacBio CCS.</p> <p><strong>Isoform sets used</strong></p> <p>isomiga_full: all isoforms discovered from hybrid assembly of long-read and short-read data</p> <p>isomiga_novel: just the novel isoforms</p> <p>isomiga_gencode: the novel isoforms combined with all isoforms in GENCODE v38.</p> <p>Not included: GENCODE v38 - available at https://www.gencodegenes.org/human/release_38.html</p> <p><strong>Isoform coordinates (GTF)</strong></p> <p>All files bgzipped and tabixed for easy random access by IGV. {file}.gtf.gz is the Isoform list, {file}.gtf.gz.tbi is the tabix index.</p> <p><strong>Isoform sequences (FASTA)</strong></p> <p>All files gzipped.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
dryad32/100

Construction of a chromosome-scale long-read reference genome assembly for potato

Open the record for dataset details and reuse information.

publicAug 2020View details →
dryad32/100

Data from: Structural variation and its potential impact on genome instability: novel discoveries in the EGFR landscape by long-read sequencing

Open the record for dataset details and reuse information.

publicAug 2021View details →
dryad32/100

Long-read genome sequencing of bread wheat facilitates disease resistance gene cloning

Open the record for dataset details and reuse information.

publicDec 2021View details →
dryad32/100

Data from: Haplotype sequence collection of ABO blood group alleles by long-read sequencing reveals putative A1-diagnostic variants

Open the record for dataset details and reuse information.

publicOct 2022View details →
dryad32/100

Improved contiguity of the threespine stickleback genome using long-read sequencing

Open the record for dataset details and reuse information.

publicFeb 2021View details →
zenodo28/100

Supporting data for the manuscript "metaFlye: scalable long-read metagenome assembly using repeat graphs"

<p>Genome assemblies, simulated datasets and&nbsp;evaluations described in the&nbsp;manuscript&nbsp;&quot;metaFlye: scalable long-read metagenome assembly using repeat graphs&quot;.</p>

opencc-by-4.0Aug 2020View details →
zenodo28/100

Figure 1 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431

Figure 1 Examples of lichen herbarium specimens used for this study AParmotrema perlatum, specimen J.A. Elix 43686 (CANB790817) BEndocarpon pusillum, specimen H. Streiman 45100 (CBG9011273) CBuellia albula, specimen J.A. Elix 45138 (CANB810791) DCatillaria sp., specimen J.A Elix 37142 (CANB872684). Scale bar: 1 cm. Photos C. Gueidan.

opencc-by-4.0Feb 2022View details →
zenodo28/100

Figure 2 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431

Figure 2 Sequencing success for different morphological groups of taxa included in this study. Specimens were grouped into three main morphological categories: 1Buellia, Catillaria and other crustose saxicolous taxa 2Endocarpon and other squamulose terricolous taxa 3 the foliose corticolous genus Parmotrema. In the graph, stalked columns show successful samples (sequence generated for the target species) in dark grey and unsuccessful samples (no sequence generated or generated sequences not from the target species) in light grey. The total number of samples (N) is indicated below each corresponding column.

opencc-by-4.0Feb 2022View details →
zenodo28/100

Figure 3 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431

Figure 3 Sequencing success for different ages of specimens included in this study. Specimens were grouped in five categories: 1966–1980, 1981–1990, 1991–2000, 2001–2010, 2011–2020. In the graph, stalked columns show successful samples (sequence generated for the target species) in dark grey and unsuccessful samples (no sequence generated or generated sequences not from the target species) in light grey. The total number of samples (N) is indicated below each corresponding column.

opencc-by-4.0Feb 2022View details →
zenodo28/100

Haplotype-phasing of long-read HiFi data to enhance structural variant detection through a Skip-Gram model

<p>Example dataset&nbsp;for DipPAV</p>

opencc-by-4.0May 2022View details →
zenodo28/100

Healthy and cancer individuals from Long-read sequencing reveals aberrant fragmentation patterns and origins of circulating DNA in cancer

<p>Oxford Nanopore sequencing data from 61 samples described in manuscript "Long-read sequencing reveals aberrant fragmentation patterns and origins of circulating DNA in cancer"<br><br></p> <p>All data are aligned to UCSC analysisSet hg38 (https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/analysisSet/), and use 0-based coordinates.</p> <p><strong>Level 2:</strong> In order to provide combined fragmentation and methylation information, Biscuit was used to create &ldquo;epiBED&rdquo; files (see Methods). epiBED files contain each read on a separate line, with each DNA methylation call on the read. epiBED files can be used for combined methylation/fragmentomic analysis, and are compatible with the CelFiE-ISH software that was used for cell of origin deconvolution.</p> <p>Creation of Biscuit EpiBED files:<br>Biscuit (https://huishenlab.github.io/biscuit/) v. 1.4.1-dev &nbsp;was used with the command &ldquo;biscuit epiread -M -b 0 -m 0 -a 0 -5 0 -3 0 -y 0.9 -L 1000000 hg38.analysisSet.fa&rdquo;, where the reference is the same UCSC reference genome used for alignment. Files are available in the Zenodo repository listed in Data Availability.</p> <p><strong>Level 3</strong>: DNA methylation BED files. BED files created by modkit (see Methods) provide one line for each CpG covered, and can be used for basic DNA methylation analysis.<br><br>Creation of modkit BED files:<br>BED files were created using modkit (https://github.com/nanoporetech/modkit) v. 0.1.5 with the command &ldquo;modkit pileup --cpg --combine-strands --ignore h --filter-threshold 0.9 --bedgraph".&nbsp;</p> <p>&nbsp;</p>

openMay 2024View details →
zenodo28/100

Leaf: an ultrafast filter for population-scale long-read SV detection: dataset 2

<p>Simulated PacBio Simuluated ONT reads for assessment</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record