Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

18

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

18 results for “haplotype-resolved”

Learn how ShareScore rates datasets ↗
zenodo44/100

Large structural variations in the haplotype-resolved African cassava genome

<p>Cassava TME7 haplotype resolved assemblies and annotation</p> <p>&nbsp;</p> <p>ABSTRACT:</p> <p>Cassava (<em>Manihot esculenta</em> Crantz, 2n=36) is a global food security crop. Cassava has a highly heterozygous genome, high genetic load, and genotype-dependent asynchronous flowering. It is typically propagated by stem cuttings and any genetic variation between haplotypes, including large structural variations, is preserved by such clonal propagation. Traditional genome assembly approaches generate a collapsed haplotype representation of the genome. In highly heterozygous plants, this results in artifacts and an oversimplification of heterozygous regions. We used a combination of Pacific Biosciences (PacBio), Illumina, and Hi-C to resolve each haplotype of the genome of a farmer-preferred cassava line, TME7 (Oko-iyawo). PacBio reads were assembled using the FALCON suite. Phase switch errors were corrected using FALCON-Phase and Hi-C read data. The ultra-long-range information from Hi-C sequencing was also used for scaffolding. Comparison of the two phases revealed more than 5,000 large haplotype-specific structural variants affecting over 8 Mb, including insertions and deletions spanning thousands of base pairs. The potential of these variants to affect allele specific expression was further explored. RNA-seq data from 11 different tissue types were mapped against the scaffolded haploid assembly and gene expression data are incorporated into our existing easy-to-use web-based interface to facilitate use by the broader plant science community. These two assemblies provide an excellent means to study the effects of heterozygosity, haplotype-specific structural variation, gene hemizygosity, and allele specific gene expression contributing to important agricultural traits and further our understanding of the genetics and domestication of cassava.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Haplotype-Resolved and Gap-Free Genome of a Floating Aquatic Plant from the Oryzeae Tribe, Hygroryza aristata

<p><em><span>Hygroryza aristata</span></em><span> (Retz.) Nees ex Wight &amp; Arn.</span><span> </span><span>is a floating aquatic plant. <span>Genomic DNA and RNA samples of </span><em><span>H. aristata</span></em><span> were extracted from plants clonally propagated from a single individual. Long-read sequencing of PacBio HiFi and ultra-long (UL) ONT (read lengths &gt; 100 kb), and short-read sequencing of Hi-C, WGS, and RNA-seq, were performed.&nbsp;</span></span></p> <p><span>For genome assembly, 31.91 Gb of PacBio HiFi and 22.36 Gb of UL-ONT sequencing data sets were utilized. Assembly was conducted using HiFiAsm (v0.20.0-r639)&nbsp;under HiFi + UL-ONT mode with the following parameters: -l 3 -r 5 -a 6 -n 10 --ctg-n 10 -w 63 -k 63. Chromosome IDs and strand directions were determined by aligning the assemblies to the rice (</span><em><span>O. sativa</span></em><span>) genome. The resulting assemblies of unphased two haplotypes, designated as hap1 and hap2, were obtained. </span></p> <p><span><span>Both hap1 and hap2 are complete genomes,<span> </span>each comprising 12 chromosomes with genome sizes of 349.74 Mb and 347.98 Mb, respectively. Notably, both haplotypes are gap-free. Telomere detection using Seqtk telo (v1.4-r122) revealed that each haplotype contains 23 telomeres. In conclusion, this study presents a haplotype-resolved and gap-free genome assembly.&nbsp;</span></span></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Chromosome-scale, haplotype-resolved genome assembly of Suaeda glauca

<p><em>Suaeda glauca</em>is an annual herb of Suaeda and an important saline-alkali plant resource, which is widespread on beaches and saline lands around the world. It is also a good candidate for food, feed, and drug development. There has been no publication of the&nbsp;<em>Suaeda glauca</em>genome assembly, limiting the evolutionary study of Amaranthaceae and the bioavailability of&nbsp;<em>Suaeda glauca</em>.</p> <p>Using PacBio HiFi and Hi-C sequencing data, we successfully generated chromosome-scale, haplotype-resolved assemblies of the&nbsp;<em>Suaeda glauca</em>genome. The size of the final primary assembly was 622.95 Mb, and the contig N50 was 19.42 Mb, which was successfully anchored to 9 chromosomes, accounting for 96.79% of the total assembly size. The repeat content and genome size of&nbsp;<em>Suaeda glauca</em>are much higher than those of the same genus&nbsp;<em>Suaeda aralocaspica</em>, presumably due to a recent burst of LTR insertions. Using HiFi reads, we assembled the complete circular chloroplast genome of&nbsp;<em>Suaeda glauca</em>. Through gene family and phylogenetic tree analysis, it was shown that&nbsp;<em>Suaeda glauca</em>and&nbsp;<em>Suaeda aralocaspica</em>differentiated at ~26.36 million years ago (MYA), and Amaranthaceae species began to differentiate at ~52.00 MYA.</p>

opencc-by-4.0Feb 2022View details →
dryad36/100

Linked-read sequencing enables haplotype-resolved resequencing at population scale

The feasibility to sequence entire genomes of virtually any organism provides unprecedented insights into the evolutionary history of populations and species. Nevertheless, many population genomic inferences – including the quantification and dating of admixture, introgression and demographic events, and inference of selective sweeps – are still limited by the lack of high-quality haplotype information. The newest generation of sequencing technology now promises significant progress. To establish the feasibility of haplotype-resolved genome resequencing at population scale, we investigated properties of linked-read sequencing data of songbirds of the genus Oenanthe across a range of sequencing depths. Our results based on the comparison of downsampled (25x, 20x, 15x, 10x, 7x, and 5x) with high-coverage data (46-68x) of seven bird genomes mapped to a reference suggest that phasing contiguities and accuracies adequate for most population genomic analyses can be reached already with moderate sequencing effort. At 15x coverage, phased haplotypes span about 90% of the genome assembly, with 50 and 90 percent of phased sequences located in phase blocks longer than 1.25-4.6 Mb (N50) and 0.27-0.72 Mb (N90). Phasing accuracy reaches beyond 99% starting from 15x coverage. Higher coverages yielded higher contiguities (up to about 7 Mb/1Mb (N50/N90) at 25x coverage), but only marginally improved phasing accuracy. Phase block contiguity improved with input DNA molecule length; thus, higher-quality DNA may help keeping sequencing costs at bay. In conclusion, even for organisms with gigabase-sized genomes like birds, linked-read sequencing at moderate depth opens an affordable avenue towards haplotype-resolved genome resequencing at population scale.

opencc-zeroMay 2020View details →
zenodo36/100

Haplotype-resolved and near-T2T assembly of the African catfish (Clarias gariepinus)

<p>Airbreathing catfishes are a group of stenohaline freshwater fish that can withstand various environmental conditions and farming practices, including the ability to breathe atmospheric oxygen. This unique ability has allowed them to thrive in semi-terrestrial habitats. However, the genomic mechanisms underlying their adaptation to adverse ecological conditions remain to fully investigate, due to the absence of gold standard reference genomes. The present study aimed to sequence and characterize the genome of the African catfish (<em>Clarias gariepinus</em>), a representative air-breathing catfish, to elucidate the genomic underpinnings of its remarkable adaptability. By generating a near telomere-to-telomere (T2T) assembly with high-resolution haplotypes, we sought to identify genomic and evolutionary features that may have contributed to its ability to withstand adverse conditions and transition to semi-terrestrial life. \textbf{Methods:} We conducted a comprehensive genomic analysis of the African catfish using a multi-platform sequencing approach, integrating Oxford Nanopore, PacBio HiFi, Illumina, and Hi-C technologies to achieve a haplotype-resolved chromosome-scale genome assembly. Functional annotations and comparative genomic analyses, including gene family evolution and positive selection studies, were performed to identify the genomic mechanisms underlying the species' resilience and adaptation to diverse environments.<strong> Results:</strong> This multifaceted approach has provided novel insights into the African catfish's complex genomic architecture and adaptive strategies. The near-T2T diploid assembly yielded 48 contigs spanning 969.62 Mb with a contig N50 of 33.71 Mb. We report 25,655 predicted protein-coding genes and 43.94\% repetitive elements in the African catfish genome. Several gene families involved in ion transport, osmoregulation, oxidative stress response, and muscle metabolism were expanded and positively selected in clariids, suggesting a potential role in their transition and adaptation to semi-terrestrial habitats. <strong>Conclusion</strong>: Our study provides a comprehensive genomic resource for \textit{Clarias gariepinus}, shedding light on the genetic and genomic mechanisms of clariids' adaptation to adverse ecological environments. The findings enhance our understanding of resilience in <em>C. gariepinu</em>s and offer valuable insights for improving aquaculture and studying related teleosts.</p>

opencc-by-4.0Jun 2024View details →
dryad36/100

Linked-read sequencing enables haplotype-resolved resequencing at population scale

Open the record for dataset details and reuse information.

publicMay 2020View details →
dryad36/100

Three haplotype-resolved pentaploid Rosa assemblies with assembled and extracted single copy orthologue (SCO) sequences from Rosa canina genome, diploid Rosa species, and sect. Caninae pollen

Open the record for dataset details and reuse information.

publicOct 2025View details →
zenodo32/100

Supplementary tables for chapter 3: "Haplotype-resolved transcriptomics defines the inheritance and genetic architecture of response to Citrus Greening Disease"

<p>Dissertation chapter 3: "<span>Haplotype-resolved transcriptomics defines the inheritance and genetic architecture of response to Citrus Greening Disease"</span></p>

opencc-by-4.0Nov 2024View details →
dryad28/100

Haplotype-resolved genome analyses of a heterozygous diploid potato

<p>Potato (<i>Solanum tuberosum</i> L.) is the most important tuber crop worldwide. An effort is underway to transform the crop from a clonally propagated tetraploid into a diploid seed-propagated, inbred line-based hybrid, which requires a better understanding of its highly heterozygous genome of potato. Here, we report the 1.67 Gb haplotype-resolved assembly of a diploid potato, RH89-039-16, using the combination of multiple sequencing and mapping strategies, including circular consensus sequencing. Comparison of the two haplotypes revealed ~2.1% intra-genome diversity, including 22,134 predicted deleterious mutations in 10,642 annotated genes. In a total of 20,583 pairs of allelic genes, 16.6% and 30.8% exhibited differential expression and methylation between alleles, respectively. Deleterious mutations and differentially expressed alleles were dispersed throughout both haplotypes, complicating strategies to eradicate deleterious alleles or stacking of beneficial alleles, via meiotic recombination. Further cataloguing of functional haplotypes, in diploid potato, could enable exploitation of heterosis using genotypes with complementary haplotypes. This study offers a holistic view of the genome organization of a clonally propagated diploid species, as well as provides insights into technological evolution in resolving complex genomes.</p>

opencc-zeroDec 2019View details →
dryad28/100

Haplotype-resolved genome analyses of a heterozygous diploid potato

Open the record for dataset details and reuse information.

publicJul 2020View details →
geo24/100

CRISPR-based targeted haplotype-resolved assembly of a megabase region [WGBS]

GEO Series GSE192499. Homo sapiens. 3 samples. Type: Methylation profiling by high throughput sequencing.

openGEO-OpenNov 2022View details →
geo24/100

CRISPR-based targeted haplotype-resolved assembly of a megabase region

GEO Series GSE192502. Homo sapiens. 4 samples. Type: Methylation profiling by array; Methylation profiling by high throughput sequencing.

openGEO-OpenNov 2022View details →
geo24/100

Integrative analysis of haplotype-resolved epigenomes across human tissues

GEO Series GSE58752. Homo sapiens. 4 samples. Type: Other.

openGEO-OpenFeb 2015View details →
geo24/100

CRISPR-based targeted haplotype-resolved assembly of a megabase region [EPIC]

GEO Series GSE192501. Homo sapiens. 1 samples. Type: Methylation profiling by array.

openGEO-OpenNov 2022View details →
zenodo20/100

The haplotype-resolved T2T genome of teinturier cultivar Yan73 reveals the genetic basis of anthocyanin biosynthesis in grapes

<p>The cultivar "Yan73' was used for assembling the first T2T genome of teinturier grapes by applying the PacBio Sequel Ⅱ platform, Hi-C technology and ultralong Oxford Nanoore Technologies (ONT) . Two haplotypes genomes were assembled, with the sizes of 501Mb and 493.38 Mb. The Yan73' sequencing on the PacBig Sequel Ⅱ platform generated a total of 29.41 Gb HiFi reads and two haplotypes were finally assembled.</p><p>We employed K-mers to assess genomic heterozygosity, estimating it at 1.35%. BUSCO was used to evaluate genomic completeness, with approximately 98.4% completeness for Yan73 haplotype 1 and 98.1% for Yan73 haplotype 2 in terms of core conserved plant genes within the genome assembly. The first genome of the teinturier grape Yan73 was successfully assembled, identifying 334,930 and 34,919 genes in haplotype 1 and haplotype 2 genomes, respectively.</p><p>The&nbsp;Yan73hap1 genome assembly: Yan73hap1.fa</p><p>The&nbsp;Yan73hap2 genome assembly: Yan73hap2.fa</p><p>The Yan73hap1 gene annotation: Yan73hap1.gff3</p><p>The Yan73hap2 gene annotation: Yan73hap2.gff3</p><p>The&nbsp;Yan73hap1&nbsp;TE annotation : Yan73hap1.TE.gff</p><p>The&nbsp;Yan73hap2&nbsp;TE annotation : Yan73hap2.TE.gff</p>

restrictedcc-by-4.0Oct 2024View details →
geo16/100

Intraspecific sequence variation and complete haplotype-resolved assemblies refine the identification of rapidly evolving regions in humans

GEO Series GSE311407. Mus musculus; Homo sapiens; Escherichia coli. 7 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2025View details →
geo16/100

Haplotype-Resolved Analysis of the Filaggrin Gene Elucidates its Complex Role in Human Adaptation and Disease

GEO Series GSE316840. Homo sapiens. 6 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenFeb 2026View details →
zenodo12/100

The haplotype-resolved genome assemblies of the hybrid Pinot Noir

<p>The haplotype-resolved genome assemblies of the hybrid Pinot Noir, PN1(495.185MB) and PN2(489.609 MB).<br> To validate the quality of our assembly, K-mer and BUSO0 were conducted. We used K-mer to evaluate genomic heterozygosity, estimated 1.43%. BUSCO to evaluate genomic completeness about 98.3% in PN1 of thecore conserved plant genes were found complete in the genome assembly. For genome annotation, the number of genesidentified by the genome is similar, more than 33,803 genes were found for Pinot Noir, PN1<br> The PN1 genome assembly: PNhap1.v1. fa<br> The PN1gene annotation: PNhap1.v1. gff3<br> The PN1 TE annotation: PNhap1_TE.v1.gff<br> The PN2 genome assembly: PNhap2.v1. fa</p>

restrictedDec 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record