Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
233
datasets available to search
ShareScore release 0.9.0
Dataset results
233 results for “long-read”
Expanded dynamic methylome and quantitative trait detection by long-read epigenome profiling of personal DNA
<p>Scripts and methylation frequencies for "Expanded dynamic methylome and quantitative trait detection by long-read epigenome profiling of personal DNA" manuscript.</p>
Data from: Structural variation and its potential impact on genome instability: novel discoveries in the EGFR landscape by long-read sequencing
<p>Studies of structural variation (SV) have been challenging due to technological contraints. With the advent of third generation (long-read) sequencing technology, exploration of longer stretches of DNA not easily examined previously has been made possible. In the present study, we utilized third generation (long-read) sequencing techniques to examime SV in the <em>EGFR </em>landscape of four haplotypes derived from two human samples. We analyzed the <em>EGFR</em> gene and its landscape (+/- 500,000 base pairs) using this sequencing approach and were able to identify regions of non-coding DNA which had relatively high similarity to the most common activating <em>EGFR</em> mutation in non-small cell lung cancer. We discovered that reverse complements to the exon 19 deletion mutation which had at least 60% homology to the <em>EGFR</em> exon 19 canonical deletion and were within ± 421,000 bp of the deletion varied across the five haploid genomes examined (4 patient landscapes and hg38). Although the sample size is limited in this study, the estimated variation observed in genomic stability between the five <em>EGFR</em> haplotypes examined is novel and encourages further work to examine structural variation in larger cohorts.</p>
Long-read genome sequencing accelerated the cloning of Pm69 by resolving the complexity of a rapidly evolving resistance gene cluster in wheat
<p>Oxford Nanopore assembly of <em>Triticum turgidum</em> ssp. <em>dicoccoides, </em>cv. G305-3M.</p>
Long-read transcriptome data (Iso-Seq) of the human and mouse brain
<p>Datasets from "<strong>Full-length transcript sequencing of human and mouse cerebral cortex identifies widespread isoform diversity and alternative splicing</strong>", SK.Leung, A.Jeffries. et al. (2021)</p> <p>Deposited files are generated from running Cupcake, SQANTI2 (v7.4) and filter. <br> Note, only the files generated from SQANTI2 filtering are deposited. <br> To re-run SQANTI, use the files in cupcake_collapse folder as input. </p> <p>Please refer to code (https://github.com/SziKayLeung/Whole_Transcriptome_Paper) for more information. </p> <p>Datasets: <br> - AdultCTX: Adult human prefrontal cortex tissue (n = 4) merged dataset<br> - FetalCTX: Fetal human prefrontal cortex tissue (n = 3) merged dataset<br> - FetalHIP: Fetal human hippocampus tissue (n = 2, subset of FetalCTX) merged dataset<br> - FetalSTR: Fetal human striatum tissue (n = 2, subset of FetalCTX) merged dataset<br> - MouseCTX: Mouse entorhinal cortex tissue (n = 12) merged dataset</p>
The data used in correspondence on the article entitled "Reply: Correspondence on NanoVar's performance outlined by Jiang T. et al. in 'Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluation'."
<p>The data used in correspondence on the article entitled “Reply: Correspondence on NanoVar’s performance outlined by Jiang T. et al. in ‘Long-read sequencing settings for efficient structural variation detection based on comprehensive evaluation’.”. It contains the benchmarking results evaluated by Truvari. The running log files of NanoVar are also included in this repository.</p>
Purification of High Molecular Weight DNA for Long-Read Sequencing Using a High-Salt Gel Electroelution Trap
<p><strong>Figure 3. Yield and purity of HMW DNA obtained from difficult samples using the method proposed in this study</strong>.</p> <p>(A) The proposed method extracts more HMW DNA from a complex soil sample than a commercial column purification kit. Shown is a negative image of an ethidium bromide-stained agarose gel. Lane 1: 1/10 aliquot of ~90 ng of HMW DNA isolated using the E.Z.N.A. soil DNA extraction kit from a soil sample containing ~1,5 μg of total DNA (HMW DNA yield around 6%). Lane 2: ~100 ng of CTAB-extracted DNA from the same soil sample. ~10 μg of this crude DNA preparation was used as input for HMW DNA purification using the proposed method. Lane 3: 1/100 aliquot of ~3 μg of HMW DNA isolated using the proposed method from ~10 μg of the CTAB-extracted DNA (HMW DNA yield around 30%).</p> <p>(B) The proposed method yields high-purity HMW DNA from a complex plant sample, as determined by agarose gel electrophoresis. Lane 1: molecular weight marker (GeneRuler DNA ladder, Thermo Fisher Scientific). Lane 2: crude nucleic acid preparation extracted with SDS/Proteinase K from <em>Zingeria trichopoda</em> leaves, which served as an input for HMW DNA purification using the proposed method. Lane 3: purified HMW DNA. Note the absence of low-molecular-weight nucleic acids and heavy covalent complexes in the purified sample.</p> <p>(C) Same as (B) except that crude, CTAB-extracted DNA from a complex soil sample was used as an input for HMW DNA purification. Note the absence of a continuous smear of fragmented DNA as well as heavy covalent complexes in the purified sample (lane 3). Molecular weight marker sizes are indicated in base pairs to the left of each panel.</p>
Tradeoffs in alignment and assembly-based methods for structural variant detection with long-read sequencing data
<p>Source data for the paper "Tradeoffs in alignment and assembly-based methods for structural variant detection with long-read sequencing data"</p>
isoMiGA: Long-read isoforms discovered in human microglia
<p>https://github.com/RajLabMSSM/isoMiGA </p> <p>Novel isoform discovery from long-read RNA-seq of 30 human microglia samples with PacBio CCS.</p> <p><strong>Isoform sets used</strong></p> <p>isomiga_full: all isoforms discovered from hybrid assembly of long-read and short-read data</p> <p>isomiga_novel: just the novel isoforms</p> <p>isomiga_gencode: the novel isoforms combined with all isoforms in GENCODE v38.</p> <p>Not included: GENCODE v38 - available at https://www.gencodegenes.org/human/release_38.html</p> <p><strong>Isoform coordinates (GTF)</strong></p> <p>All files bgzipped and tabixed for easy random access by IGV. {file}.gtf.gz is the Isoform list, {file}.gtf.gz.tbi is the tabix index.</p> <p><strong>Isoform sequences (FASTA)</strong></p> <p>All files gzipped. </p> <p> </p>
Construction of a chromosome-scale long-read reference genome assembly for potato
Open the record for dataset details and reuse information.
Data from: Structural variation and its potential impact on genome instability: novel discoveries in the EGFR landscape by long-read sequencing
Open the record for dataset details and reuse information.
Long-read genome sequencing of bread wheat facilitates disease resistance gene cloning
Open the record for dataset details and reuse information.
Data from: Haplotype sequence collection of ABO blood group alleles by long-read sequencing reveals putative A1-diagnostic variants
Open the record for dataset details and reuse information.
Improved contiguity of the threespine stickleback genome using long-read sequencing
Open the record for dataset details and reuse information.
Supporting data for the manuscript "metaFlye: scalable long-read metagenome assembly using repeat graphs"
<p>Genome assemblies, simulated datasets and evaluations described in the manuscript "metaFlye: scalable long-read metagenome assembly using repeat graphs".</p>
Figure 1 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431
Figure 1 Examples of lichen herbarium specimens used for this study AParmotrema perlatum, specimen J.A. Elix 43686 (CANB790817) BEndocarpon pusillum, specimen H. Streiman 45100 (CBG9011273) CBuellia albula, specimen J.A. Elix 45138 (CANB810791) DCatillaria sp., specimen J.A Elix 37142 (CANB872684). Scale bar: 1 cm. Photos C. Gueidan.
Figure 2 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431
Figure 2 Sequencing success for different morphological groups of taxa included in this study. Specimens were grouped into three main morphological categories: 1Buellia, Catillaria and other crustose saxicolous taxa 2Endocarpon and other squamulose terricolous taxa 3 the foliose corticolous genus Parmotrema. In the graph, stalked columns show successful samples (sequence generated for the target species) in dark grey and unsuccessful samples (no sequence generated or generated sequences not from the target species) in light grey. The total number of samples (N) is indicated below each corresponding column.
Figure 3 from: Gueidan C, Li L (2022) A long-read amplicon approach to scaling up the metabarcoding of lichen herbarium specimens. MycoKeys 86: 195-212. https://doi.org/10.3897/mycokeys.86.77431
Figure 3 Sequencing success for different ages of specimens included in this study. Specimens were grouped in five categories: 1966–1980, 1981–1990, 1991–2000, 2001–2010, 2011–2020. In the graph, stalked columns show successful samples (sequence generated for the target species) in dark grey and unsuccessful samples (no sequence generated or generated sequences not from the target species) in light grey. The total number of samples (N) is indicated below each corresponding column.
Haplotype-phasing of long-read HiFi data to enhance structural variant detection through a Skip-Gram model
<p>Example dataset for DipPAV</p>
Healthy and cancer individuals from Long-read sequencing reveals aberrant fragmentation patterns and origins of circulating DNA in cancer
<p>Oxford Nanopore sequencing data from 61 samples described in manuscript "Long-read sequencing reveals aberrant fragmentation patterns and origins of circulating DNA in cancer"<br><br></p> <p>All data are aligned to UCSC analysisSet hg38 (https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/analysisSet/), and use 0-based coordinates.</p> <p><strong>Level 2:</strong> In order to provide combined fragmentation and methylation information, Biscuit was used to create “epiBED” files (see Methods). epiBED files contain each read on a separate line, with each DNA methylation call on the read. epiBED files can be used for combined methylation/fragmentomic analysis, and are compatible with the CelFiE-ISH software that was used for cell of origin deconvolution.</p> <p>Creation of Biscuit EpiBED files:<br>Biscuit (https://huishenlab.github.io/biscuit/) v. 1.4.1-dev was used with the command “biscuit epiread -M -b 0 -m 0 -a 0 -5 0 -3 0 -y 0.9 -L 1000000 hg38.analysisSet.fa”, where the reference is the same UCSC reference genome used for alignment. Files are available in the Zenodo repository listed in Data Availability.</p> <p><strong>Level 3</strong>: DNA methylation BED files. BED files created by modkit (see Methods) provide one line for each CpG covered, and can be used for basic DNA methylation analysis.<br><br>Creation of modkit BED files:<br>BED files were created using modkit (https://github.com/nanoporetech/modkit) v. 0.1.5 with the command “modkit pileup --cpg --combine-strands --ignore h --filter-threshold 0.9 --bedgraph". </p> <p> </p>
Leaf: an ultrafast filter for population-scale long-read SV detection: dataset 2
<p>Simulated PacBio Simuluated ONT reads for assessment</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.