Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18
datasets available to search
ShareScore release 0.7.1
Dataset results
18 results for “haplotype-resolved”
Large structural variations in the haplotype-resolved African cassava genome
<p>Cassava TME7 haplotype resolved assemblies and annotation</p> <p> </p> <p>ABSTRACT:</p> <p>Cassava (<em>Manihot esculenta</em> Crantz, 2n=36) is a global food security crop. Cassava has a highly heterozygous genome, high genetic load, and genotype-dependent asynchronous flowering. It is typically propagated by stem cuttings and any genetic variation between haplotypes, including large structural variations, is preserved by such clonal propagation. Traditional genome assembly approaches generate a collapsed haplotype representation of the genome. In highly heterozygous plants, this results in artifacts and an oversimplification of heterozygous regions. We used a combination of Pacific Biosciences (PacBio), Illumina, and Hi-C to resolve each haplotype of the genome of a farmer-preferred cassava line, TME7 (Oko-iyawo). PacBio reads were assembled using the FALCON suite. Phase switch errors were corrected using FALCON-Phase and Hi-C read data. The ultra-long-range information from Hi-C sequencing was also used for scaffolding. Comparison of the two phases revealed more than 5,000 large haplotype-specific structural variants affecting over 8 Mb, including insertions and deletions spanning thousands of base pairs. The potential of these variants to affect allele specific expression was further explored. RNA-seq data from 11 different tissue types were mapped against the scaffolded haploid assembly and gene expression data are incorporated into our existing easy-to-use web-based interface to facilitate use by the broader plant science community. These two assemblies provide an excellent means to study the effects of heterozygosity, haplotype-specific structural variation, gene hemizygosity, and allele specific gene expression contributing to important agricultural traits and further our understanding of the genetics and domestication of cassava.</p>
Haplotype-Resolved and Gap-Free Genome of a Floating Aquatic Plant from the Oryzeae Tribe, Hygroryza aristata
<p><em><span>Hygroryza aristata</span></em><span> (Retz.) Nees ex Wight & Arn.</span><span> </span><span>is a floating aquatic plant. <span>Genomic DNA and RNA samples of </span><em><span>H. aristata</span></em><span> were extracted from plants clonally propagated from a single individual. Long-read sequencing of PacBio HiFi and ultra-long (UL) ONT (read lengths > 100 kb), and short-read sequencing of Hi-C, WGS, and RNA-seq, were performed. </span></span></p> <p><span>For genome assembly, 31.91 Gb of PacBio HiFi and 22.36 Gb of UL-ONT sequencing data sets were utilized. Assembly was conducted using HiFiAsm (v0.20.0-r639) under HiFi + UL-ONT mode with the following parameters: -l 3 -r 5 -a 6 -n 10 --ctg-n 10 -w 63 -k 63. Chromosome IDs and strand directions were determined by aligning the assemblies to the rice (</span><em><span>O. sativa</span></em><span>) genome. The resulting assemblies of unphased two haplotypes, designated as hap1 and hap2, were obtained. </span></p> <p><span><span>Both hap1 and hap2 are complete genomes,<span> </span>each comprising 12 chromosomes with genome sizes of 349.74 Mb and 347.98 Mb, respectively. Notably, both haplotypes are gap-free. Telomere detection using Seqtk telo (v1.4-r122) revealed that each haplotype contains 23 telomeres. In conclusion, this study presents a haplotype-resolved and gap-free genome assembly. </span></span></p>
Chromosome-scale, haplotype-resolved genome assembly of Suaeda glauca
<p><em>Suaeda glauca</em>is an annual herb of Suaeda and an important saline-alkali plant resource, which is widespread on beaches and saline lands around the world. It is also a good candidate for food, feed, and drug development. There has been no publication of the <em>Suaeda glauca</em>genome assembly, limiting the evolutionary study of Amaranthaceae and the bioavailability of <em>Suaeda glauca</em>.</p> <p>Using PacBio HiFi and Hi-C sequencing data, we successfully generated chromosome-scale, haplotype-resolved assemblies of the <em>Suaeda glauca</em>genome. The size of the final primary assembly was 622.95 Mb, and the contig N50 was 19.42 Mb, which was successfully anchored to 9 chromosomes, accounting for 96.79% of the total assembly size. The repeat content and genome size of <em>Suaeda glauca</em>are much higher than those of the same genus <em>Suaeda aralocaspica</em>, presumably due to a recent burst of LTR insertions. Using HiFi reads, we assembled the complete circular chloroplast genome of <em>Suaeda glauca</em>. Through gene family and phylogenetic tree analysis, it was shown that <em>Suaeda glauca</em>and <em>Suaeda aralocaspica</em>differentiated at ~26.36 million years ago (MYA), and Amaranthaceae species began to differentiate at ~52.00 MYA.</p>
Linked-read sequencing enables haplotype-resolved resequencing at population scale
The feasibility to sequence entire genomes of virtually any organism provides unprecedented insights into the evolutionary history of populations and species. Nevertheless, many population genomic inferences – including the quantification and dating of admixture, introgression and demographic events, and inference of selective sweeps – are still limited by the lack of high-quality haplotype information. The newest generation of sequencing technology now promises significant progress. To establish the feasibility of haplotype-resolved genome resequencing at population scale, we investigated properties of linked-read sequencing data of songbirds of the genus Oenanthe across a range of sequencing depths. Our results based on the comparison of downsampled (25x, 20x, 15x, 10x, 7x, and 5x) with high-coverage data (46-68x) of seven bird genomes mapped to a reference suggest that phasing contiguities and accuracies adequate for most population genomic analyses can be reached already with moderate sequencing effort. At 15x coverage, phased haplotypes span about 90% of the genome assembly, with 50 and 90 percent of phased sequences located in phase blocks longer than 1.25-4.6 Mb (N50) and 0.27-0.72 Mb (N90). Phasing accuracy reaches beyond 99% starting from 15x coverage. Higher coverages yielded higher contiguities (up to about 7 Mb/1Mb (N50/N90) at 25x coverage), but only marginally improved phasing accuracy. Phase block contiguity improved with input DNA molecule length; thus, higher-quality DNA may help keeping sequencing costs at bay. In conclusion, even for organisms with gigabase-sized genomes like birds, linked-read sequencing at moderate depth opens an affordable avenue towards haplotype-resolved genome resequencing at population scale.
Haplotype-resolved and near-T2T assembly of the African catfish (Clarias gariepinus)
<p>Airbreathing catfishes are a group of stenohaline freshwater fish that can withstand various environmental conditions and farming practices, including the ability to breathe atmospheric oxygen. This unique ability has allowed them to thrive in semi-terrestrial habitats. However, the genomic mechanisms underlying their adaptation to adverse ecological conditions remain to fully investigate, due to the absence of gold standard reference genomes. The present study aimed to sequence and characterize the genome of the African catfish (<em>Clarias gariepinus</em>), a representative air-breathing catfish, to elucidate the genomic underpinnings of its remarkable adaptability. By generating a near telomere-to-telomere (T2T) assembly with high-resolution haplotypes, we sought to identify genomic and evolutionary features that may have contributed to its ability to withstand adverse conditions and transition to semi-terrestrial life. \textbf{Methods:} We conducted a comprehensive genomic analysis of the African catfish using a multi-platform sequencing approach, integrating Oxford Nanopore, PacBio HiFi, Illumina, and Hi-C technologies to achieve a haplotype-resolved chromosome-scale genome assembly. Functional annotations and comparative genomic analyses, including gene family evolution and positive selection studies, were performed to identify the genomic mechanisms underlying the species' resilience and adaptation to diverse environments.<strong> Results:</strong> This multifaceted approach has provided novel insights into the African catfish's complex genomic architecture and adaptive strategies. The near-T2T diploid assembly yielded 48 contigs spanning 969.62 Mb with a contig N50 of 33.71 Mb. We report 25,655 predicted protein-coding genes and 43.94\% repetitive elements in the African catfish genome. Several gene families involved in ion transport, osmoregulation, oxidative stress response, and muscle metabolism were expanded and positively selected in clariids, suggesting a potential role in their transition and adaptation to semi-terrestrial habitats. <strong>Conclusion</strong>: Our study provides a comprehensive genomic resource for \textit{Clarias gariepinus}, shedding light on the genetic and genomic mechanisms of clariids' adaptation to adverse ecological environments. The findings enhance our understanding of resilience in <em>C. gariepinu</em>s and offer valuable insights for improving aquaculture and studying related teleosts.</p>
Linked-read sequencing enables haplotype-resolved resequencing at population scale
Open the record for dataset details and reuse information.
Three haplotype-resolved pentaploid Rosa assemblies with assembled and extracted single copy orthologue (SCO) sequences from Rosa canina genome, diploid Rosa species, and sect. Caninae pollen
Open the record for dataset details and reuse information.
Supplementary tables for chapter 3: "Haplotype-resolved transcriptomics defines the inheritance and genetic architecture of response to Citrus Greening Disease"
<p>Dissertation chapter 3: "<span>Haplotype-resolved transcriptomics defines the inheritance and genetic architecture of response to Citrus Greening Disease"</span></p>
Haplotype-resolved genome analyses of a heterozygous diploid potato
<p>Potato (<i>Solanum tuberosum</i> L.) is the most important tuber crop worldwide. An effort is underway to transform the crop from a clonally propagated tetraploid into a diploid seed-propagated, inbred line-based hybrid, which requires a better understanding of its highly heterozygous genome of potato. Here, we report the 1.67 Gb haplotype-resolved assembly of a diploid potato, RH89-039-16, using the combination of multiple sequencing and mapping strategies, including circular consensus sequencing. Comparison of the two haplotypes revealed ~2.1% intra-genome diversity, including 22,134 predicted deleterious mutations in 10,642 annotated genes. In a total of 20,583 pairs of allelic genes, 16.6% and 30.8% exhibited differential expression and methylation between alleles, respectively. Deleterious mutations and differentially expressed alleles were dispersed throughout both haplotypes, complicating strategies to eradicate deleterious alleles or stacking of beneficial alleles, via meiotic recombination. Further cataloguing of functional haplotypes, in diploid potato, could enable exploitation of heterosis using genotypes with complementary haplotypes. This study offers a holistic view of the genome organization of a clonally propagated diploid species, as well as provides insights into technological evolution in resolving complex genomes.</p>
Haplotype-resolved genome analyses of a heterozygous diploid potato
Open the record for dataset details and reuse information.
CRISPR-based targeted haplotype-resolved assembly of a megabase region [WGBS]
GEO Series GSE192499. Homo sapiens. 3 samples. Type: Methylation profiling by high throughput sequencing.
CRISPR-based targeted haplotype-resolved assembly of a megabase region
GEO Series GSE192502. Homo sapiens. 4 samples. Type: Methylation profiling by array; Methylation profiling by high throughput sequencing.
Integrative analysis of haplotype-resolved epigenomes across human tissues
GEO Series GSE58752. Homo sapiens. 4 samples. Type: Other.
CRISPR-based targeted haplotype-resolved assembly of a megabase region [EPIC]
GEO Series GSE192501. Homo sapiens. 1 samples. Type: Methylation profiling by array.
The haplotype-resolved T2T genome of teinturier cultivar Yan73 reveals the genetic basis of anthocyanin biosynthesis in grapes
<p>The cultivar "Yan73' was used for assembling the first T2T genome of teinturier grapes by applying the PacBio Sequel Ⅱ platform, Hi-C technology and ultralong Oxford Nanoore Technologies (ONT) . Two haplotypes genomes were assembled, with the sizes of 501Mb and 493.38 Mb. The Yan73' sequencing on the PacBig Sequel Ⅱ platform generated a total of 29.41 Gb HiFi reads and two haplotypes were finally assembled.</p><p>We employed K-mers to assess genomic heterozygosity, estimating it at 1.35%. BUSCO was used to evaluate genomic completeness, with approximately 98.4% completeness for Yan73 haplotype 1 and 98.1% for Yan73 haplotype 2 in terms of core conserved plant genes within the genome assembly. The first genome of the teinturier grape Yan73 was successfully assembled, identifying 334,930 and 34,919 genes in haplotype 1 and haplotype 2 genomes, respectively.</p><p>The Yan73hap1 genome assembly: Yan73hap1.fa</p><p>The Yan73hap2 genome assembly: Yan73hap2.fa</p><p>The Yan73hap1 gene annotation: Yan73hap1.gff3</p><p>The Yan73hap2 gene annotation: Yan73hap2.gff3</p><p>The Yan73hap1 TE annotation : Yan73hap1.TE.gff</p><p>The Yan73hap2 TE annotation : Yan73hap2.TE.gff</p>
Intraspecific sequence variation and complete haplotype-resolved assemblies refine the identification of rapidly evolving regions in humans
GEO Series GSE311407. Mus musculus; Homo sapiens; Escherichia coli. 7 samples. Type: Expression profiling by high throughput sequencing.
Haplotype-Resolved Analysis of the Filaggrin Gene Elucidates its Complex Role in Human Adaptation and Disease
GEO Series GSE316840. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.
The haplotype-resolved genome assemblies of the hybrid Pinot Noir
<p>The haplotype-resolved genome assemblies of the hybrid Pinot Noir, PN1(495.185MB) and PN2(489.609 MB).<br> To validate the quality of our assembly, K-mer and BUSO0 were conducted. We used K-mer to evaluate genomic heterozygosity, estimated 1.43%. BUSCO to evaluate genomic completeness about 98.3% in PN1 of thecore conserved plant genes were found complete in the genome assembly. For genome annotation, the number of genesidentified by the genome is similar, more than 33,803 genes were found for Pinot Noir, PN1<br> The PN1 genome assembly: PNhap1.v1. fa<br> The PN1gene annotation: PNhap1.v1. gff3<br> The PN1 TE annotation: PNhap1_TE.v1.gff<br> The PN2 genome assembly: PNhap2.v1. fa</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.