Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

423

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

423 results for “Haplotypes”

Learn how ShareScore rates datasets ↗
dryad36/100

Detecting selected haplotype blocks in evolve and resequence experiments

Shifting from the analysis of single nucleotide polymorphisms to the reconstruction of selected haplotypes greatly facilitates the interpretation of Evolve and Resequence (E&R) experiments. Merging highly correlated hitchhiker SNPs into haplotype blocks reduces thousands of candidates to few selected regions. Current methods of haplotype reconstruction from Pool-Seq data need a variety of data-specific parameters that are typically defined ad hoc and require haplotype sequences for validation. Here, we introduce haplovalidate, a tool which detects selected haplotypes in Pool-seq time series data without the need for sequenced haplotypes. Haplovalidate makes data-driven choices of two key parameters for the clustering procedure, the minimum correlation between SNPs constituting a cluster and the window size. Applying haplovalidate to simulated and experimental E&R data reliably detects selected haplotype blocks with low false discovery rates. Importantly, our analyses identified a restriction of the haplotype block-based approach to describe the genomic architecture of adaptation. We detected a substantial fraction of haplotypes containing multiple selection targets. These blocks were considered as one region of selection and therefore led to under-estimation of the number of selection targets. We demonstrate that the separate analysis of earlier time points can significantly increase the separation of selection targets into individual haplotype blocks. We conclude that the analysis of selected haplotype blocks has great potential for the characterisation of the adaptive architecture with E&R experiments.

opencc-zeroJan 2021View details →
zenodo36/100

NMDP 7-locus haplotype frequencies trimmed

<p>This is a gzipped tar file of csv files; one per population with HLA haplotype frequency estimates for the US population across the 7 loci:</p> <p>A~C~B~DRB3/4/5~DRB1~DQB1~DPB1</p>

opencc-by-nc-nd-4.0Jan 2021View details →
dryad36/100

Linked-read sequencing enables haplotype-resolved resequencing at population scale

The feasibility to sequence entire genomes of virtually any organism provides unprecedented insights into the evolutionary history of populations and species. Nevertheless, many population genomic inferences – including the quantification and dating of admixture, introgression and demographic events, and inference of selective sweeps – are still limited by the lack of high-quality haplotype information. The newest generation of sequencing technology now promises significant progress. To establish the feasibility of haplotype-resolved genome resequencing at population scale, we investigated properties of linked-read sequencing data of songbirds of the genus Oenanthe across a range of sequencing depths. Our results based on the comparison of downsampled (25x, 20x, 15x, 10x, 7x, and 5x) with high-coverage data (46-68x) of seven bird genomes mapped to a reference suggest that phasing contiguities and accuracies adequate for most population genomic analyses can be reached already with moderate sequencing effort. At 15x coverage, phased haplotypes span about 90% of the genome assembly, with 50 and 90 percent of phased sequences located in phase blocks longer than 1.25-4.6 Mb (N50) and 0.27-0.72 Mb (N90). Phasing accuracy reaches beyond 99% starting from 15x coverage. Higher coverages yielded higher contiguities (up to about 7 Mb/1Mb (N50/N90) at 25x coverage), but only marginally improved phasing accuracy. Phase block contiguity improved with input DNA molecule length; thus, higher-quality DNA may help keeping sequencing costs at bay. In conclusion, even for organisms with gigabase-sized genomes like birds, linked-read sequencing at moderate depth opens an affordable avenue towards haplotype-resolved genome resequencing at population scale.

opencc-zeroMay 2020View details →
zenodo36/100

HLA and KIR allele genotyping for HPRC-frz2 haplotype assemblies

<p>HLA/KIR annotation for available long reads assemblies, &nbsp;constructed by the pipeline : https://github.com/YingZhou001/Immuannot</p> <p>IPD-KIR version: V2.13.0, IPD-IMGT/HLA version: V3.59.0</p> <p>472 haploid assemblies included</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Dataset for paper 'McAN: a novel computational algorithm and platform for constructing and visualizing haplotype networks'

<p>The .zip file includes four datasets for testing the performance of McAN (doi: https://doi.org/10.1093/bib/bbad174).</p>

opencc-by-4.0Oct 2023View details →
dryad36/100

Mitogenome alignment of 159 unique haplotypes representing 455 individual killer whales

<p><span>Genome sequences can reveal the extent of inbreeding in small populations. Here we present the first genomic characterization of type D killer whales, a distinctive eco/morphotype with a circumpolar, subantarctic distribution. Effective population size is the lowest estimated from any killer whale genome and indicates a severe population bottleneck. Consequently, type D genomes show among the highest level of inbreeding reported for any mammalian species (F<sub>ROH</sub> </span><span></span><span> 0.65). Detected recombination events of different haplotypes are up to an order of magnitude rarer than in other killer whale genomes studied to date. Comparison of genomic data from a museum specimen of a type D killer whale that stranded in New Zealand in 1955, with three modern genomes from the Cape Horn area, reveals high covariance and identity-by-state of alleles, suggesting these genomic characteristics and demographic history are shared among different social groups within this morphotype. Limitations to the insights gained in this study stem from the </span><span>non-independence of the three closely related modern genomes, the short coalescence time of most variation within the genomes, and the nonequilibrium population history which violates the assumptions of many model-based methods. Long-range linkage disequilibrium and extensive runs of homozygosity found in type D genomes provide the potential basis for coupling of genetic barriers to gene flow with other killer whale populations, and the distinctive morphology. </span></p>

opencc-zeroDec 2023View details →
dryad36/100

Data from: Introgression of non-native mitochondrial haplotypes from farmed to wild Atlantic salmon

<p>Farmed salmon escape and interbreed with wild Atlantic salmon on a large scale. We studied introgression of mitochondrial haplotypes from farmed Atlantic salmon originating from the Eastern Atlantic phylogenetic group to wild salmon of the Barents-White Sea phylogenetic group. We find that farmed genetic introgression introduced novel, non-native haplotypes into the Barents-White Sea phylogenetic group. The mitochondrial genome has important functional effects and is inherited as a haploid from the mother. Hence, the observed introgression across natural genetic barriers is expected to cause long-lasting functional maladaptation of the hybrids in the maternal line. As the use of farmed Atlantic salmon from non-native phylogenetic groups is widespread in aquaculture, the impact on wild Atlantic salmon may be more severe than previously recognized. Our results highlight the ecological risks of releasing non-native wild and domesticated animals.</p>

opencc-zeroMar 2024View details →
zenodo36/100

Resolution, validation and divergence of heterozygous haplotypes from pooled long read sequencing of the diamondback moth (Lepidoptera: Plutellidae)

<p>This data sets includes the full set of intermediate genome assemblies produced during our analyses. We make these available for researchers who may be interested in the variation between assembly results and variation within the study organism (<em>Plutella xylostella</em>) prior to removal during subsequent genome processing.</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

A protective HLA extended haplotype outweighs the major COVID-19 risk factor inherited from Neanderthals in the Sardinian population

<p>Sardinia has one of the lowest incidences of hospitalization and related mortality in Europe. In this dataset we reported 358 patients with COVID-19, in which we evaluated the frequency of the Neanderthal risk locus variant on chromosome 3 (rs35044562), considered to be a major risk factor for a severe SARS-CoV-2 disease course.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Data for "Haplotype-aware pantranscriptome analyses using spliced pangenome graphs"

<p>This archive contains mapping benchmark&nbsp;tables and haplotype-specific expression estimates from the paper&nbsp;&quot;Haplotype-aware pantranscriptome analyses using spliced pangenome graphs&quot;.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
dryad36/100

Determining haploblocks and haplotypes in the MAGIC winter wheat population WM-800 based on the wheat 15k Infinium and the 135k Affymetrix SNP arrays

<p><span>Haplotypes are derived from single nucleotide polymorphisms (SNPs). They are beneficial (i) to remove redundant sequence information in genetic populations and, more important, (ii) to distinguish more than two variants/alleles at a genomic locus. A haploblock locus, made of multiple haplotypes, is very useful in multiparent-advanced-generation-intercross (MAGIC) populations, where, ideally, multiple founder alleles need to be distinguished at each locus to subsequently carry out efficient genome-wide association analysis studies (GWAS). </span></p> <p><span>In this regard, the dataset contains genotype matrices (made of SNP, haploblock and haplotype data) for 800 lines of the MAGIC WHEAT population WM-800 (Sannemann et al. 2018). The datasets are based on genotyping the lines with both the already published wheat 15k Infinium SNP array (Sannemann et al. 2018) and the new wheat 135k Affymetrix SNP array.</span></p>

opencc-zeroOct 2022View details →
zenodo36/100

Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Real and mock datasets

<p>This repository contains the reads, assemblies, and references required to replicate the <strong>real and mock</strong>&nbsp;results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Strainy: phasing and assembly of strain haplotypes from long-read metagenome sequencing - Simulated datasets

<p>This repository contains the reads, assemblies, and references required to replicate the <strong>simulated</strong>&nbsp;results presented in the paper: https://doi.org/10.1101/2023.01.31.526521</p>

opencc-by-4.0May 2024View details →
dryad36/100

Data from: Symbiont infection and psyllid haplotype influence phenotypic plasticity during host switching events

<p>Many herbivorous insect species exhibit phenotypic plasticity when using multiple hosts, which facilitates survival in heterogeneous host environments. Physiological host acclimation is an important part of it, yet the effects of host acclimation on insect feeding behavior are not well studied, particularly for insect vectors of plant pathogens. We studied the combined effects of host acclimation and infection with a plant pathogenic symbiont on feeding behavior of <em>Bactericera cockerelli</em>,<em> </em>an oligophagous psyllid widespread in both crop and natural habitats that feeds primarily on Solanaceae and transmits an economically important plant pathogen, <em>Candidatus</em> Liberibacter solanacearum (<em>C</em>Lso). We used a factorial design and the electrical penetration graphing technique to disentangle the effects of host acclimation, <em>C</em>Lso infection, and psyllid haplotype on the within-plant feeding behavior of <em>B. cockerelli</em> during conspecific and heterospecific host switches. This approach allows to connect phenotypic plasticity with the role of <em>B. cockerelli </em>as a vector by quantifying the frequency and duration of behaviors involved in <em>C</em>Lso transmission. We found significant reductions in multiple metrics of <em>B. cockerelli</em> feeding efficiency, exacerbated by infection with <em>C</em>Lso, which could lead to reduced transmission of this pathogen. Psyllid genotype was also important; the Central haplotype exhibited less dramatic changes in feeding efficiency than the Western haplotype during heterospecific host switches. Our study shows that host acclimation and heterospecific host switching directly alter feeding behaviors underlying pathogen transmission, and that the magnitude of feeding efficiency reductions depends on both host genotype and infection status.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Haplotype-resolved and near-T2T assembly of the African catfish (Clarias gariepinus)

<p>Airbreathing catfishes are a group of stenohaline freshwater fish that can withstand various environmental conditions and farming practices, including the ability to breathe atmospheric oxygen. This unique ability has allowed them to thrive in semi-terrestrial habitats. However, the genomic mechanisms underlying their adaptation to adverse ecological conditions remain to fully investigate, due to the absence of gold standard reference genomes. The present study aimed to sequence and characterize the genome of the African catfish (<em>Clarias gariepinus</em>), a representative air-breathing catfish, to elucidate the genomic underpinnings of its remarkable adaptability. By generating a near telomere-to-telomere (T2T) assembly with high-resolution haplotypes, we sought to identify genomic and evolutionary features that may have contributed to its ability to withstand adverse conditions and transition to semi-terrestrial life. \textbf{Methods:} We conducted a comprehensive genomic analysis of the African catfish using a multi-platform sequencing approach, integrating Oxford Nanopore, PacBio HiFi, Illumina, and Hi-C technologies to achieve a haplotype-resolved chromosome-scale genome assembly. Functional annotations and comparative genomic analyses, including gene family evolution and positive selection studies, were performed to identify the genomic mechanisms underlying the species' resilience and adaptation to diverse environments.<strong> Results:</strong> This multifaceted approach has provided novel insights into the African catfish's complex genomic architecture and adaptive strategies. The near-T2T diploid assembly yielded 48 contigs spanning 969.62 Mb with a contig N50 of 33.71 Mb. We report 25,655 predicted protein-coding genes and 43.94\% repetitive elements in the African catfish genome. Several gene families involved in ion transport, osmoregulation, oxidative stress response, and muscle metabolism were expanded and positively selected in clariids, suggesting a potential role in their transition and adaptation to semi-terrestrial habitats. <strong>Conclusion</strong>: Our study provides a comprehensive genomic resource for \textit{Clarias gariepinus}, shedding light on the genetic and genomic mechanisms of clariids' adaptation to adverse ecological environments. The findings enhance our understanding of resilience in <em>C. gariepinu</em>s and offer valuable insights for improving aquaculture and studying related teleosts.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Fig. 5 in Variation In Cone And Seed Morphology Traits Among The Mitochondrial Dna Haplotypes Of Scots Pine (Pinus Sylvestris L.)

Fig. 5. Distribution of the seed wing shape (%, units) between and within Scots pine mitotypes.

opencc-by-4.0Dec 2017View details →
zenodo36/100

The alignment of 163 plastome haplotypes of East Asian Cerris oaks and 29 plastomes of related oak species

<p>This dataset includes the alignment of 163 plastome haplotypes of East Asian Cerris oaks and 29 plastomes of related oak species. The alignment was generated through four steps: (1) We used PhyloSuite v.1.2.2 to extract protein-coding genes (PCGs), tRNA genes, rRNA genes, introns, and intergenic spacers (IGSs) from the plastomes of 761 East Asian Cerris oak trees and 29 accessions of 22 related oak species. (2) The extracted regions were aligned individually with MAFFT v.7.313 and adjusted manually using BioEdit v.7.2.5. Specifically, inversions and length variations in simple sequence repeats were excluded because of their tendency for homoplasy. Eight ambiguously aligned regions in rps16-trnQ, psbM-trnD, ndhF-rpl32, rpl32-trnL, ndhD-psaC, and psaC-ndhE IGSs, and ndhF and ycf1 PCGs were also discarded to reduce phylogenetic noise. (3) The individual alignments were concatenated according to their respective positions in the plastome to generate a whole-plastome alignment with only one inverted repeat (IR) retained. (4) Unique plastome haplotypes were determined by DnaSP v.5.10.01.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Figure 3 in Increased haplotype diversity of Emys orbicularis (Linnaeus, 1758) (Reptilia: Emydidae) in northern Iran

Figure 3. Haplotype network of all studied samples based on the Cytochrome b gene fragment.

opencc-by-4.0Sep 2021View details →
dryad36/100

Borrelia infection in bank voles Myodes glareolus is associated with specific DQB haplotypes which affect allelic divergence within individuals

<p>The high polymorphism of Major Histocompatibility Complex (MHC) genes is generally considered to be a result of pathogen-mediated balancing selection. Such selection may operate in the form of heterozygote advantage, and/or through specific MHC allele–pathogen interactions. Specific MHC allele–pathogen interactions may promote polymorphism via negative frequency-dependent selection (NFDS), or selection that varies in time and/or space because of variability in the composition of the pathogen community (fluctuating selection; FS). In addition, divergent allele advantage (DAA) may act on top of these forms of balancing selection, explaining the high sequence divergence between MHC alleles. DAA has primarily been thought of as an extension of heterozygote advantage. However, DAA could also work in concert with NFDS though this is yet to be tested explicitly. To evaluate the importance of DAA in pathogen-mediated balancing selection, we surveyed allelic polymorphism of MHC class II DQB genes in wild bank voles (<i>Myodes glareolus</i>) and tested for associations between DQB haplotypes and infection by <i>Borrelia afzelii</i>, a tick-transmitted bacterium causing Lyme disease in humans. We found two significant associations between DQB haplotypes and infection status: one haplotype was associated with lower risk of infection (resistance), while another was associated with higher risk of infection (susceptibility). Interestingly, allelic divergence within individuals was higher for voles with the resistance haplotype compared to other voles. In contrast, allelic divergence was lower for voles with the susceptibility haplotype than other voles. The pattern of higher allelic divergence in individuals with the resistance haplotype is consistent with NFDS favouring divergent alleles in a natural population, hence selection where DAA works in concert with NFDS. </p>

opencc-zeroJul 2021View details →
zenodo36/100

Fig. 13 in NGS-barcodes, haplotype networks combined to external morphology help to identify new species in the mangrove genus Ngirhaphium Evenhuis & Grootaert, 2002 (Diptera: Dolichopodidae: Rhaphiinae) in Southeast Asia

Fig. 13. Distribution map of Ngirhaphium in Southeast Asia

opencc-by-4.0Nov 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record