Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
423
datasets available to search
ShareScore release 0.9.0
Dataset results
423 results for “Haplotypes”
Microsatellite (13 loci) and plastid DNA haplotypes in a population of Antirrhinum charidemi
<p>Genotype matrix of 182 Antirrhinum charidemi individuals sampled in 2007-2009 in the Barranco del Dragoncillo Blanco population in Cabo de Gata, Almería, Spain. Genotypes are given for 13 microsatellite loci and also include 3 plastid DNA haplotypes. Details on loci and genotyping conditions can be found in Forrest et al. 2017, https://doi.org/10.1093/botlinnean/bow002. Each individual is geolocated. Additional information include its corolla colour, its ancestry score in four gene pools obtained in Bayesian genetic cluster analysis (STRUCTURE), and its assignment to geo-genetic subpopulations. Metadata are available in a separate tab in the submitted spreadsheet. The data are analysed in a paper expected to be published in AoB Plants in 2025, titled: "Fine-scale genetic differentiation in the bee-specialized Antirrhinum charidemi covaries more strongly with microenvironment than with corolla colour"</p>
Influence of voltine ecotype and geographic distance on genetic and haplotype variation in the Asian corn borer
<p>Diapause is an adaptive dormancy strategy by which arthropods endure extended periods of adverse climatic conditions. Seasonal variation in larval diapause initiation and duration in the Asian corn borer, <i>Ostrinia furnacalis</i>, influences adult mating generation number (voltinism) across local environmental conditions. Degree of mating period overlap between sympatric voltinism ecotypes influence hybridization level, but impact on <i>O. furnacalis</i> population genetic structure and evolution of divergent adaptive phenotypes remains uncertain. Genetic differentiation was estimated between voltinism ecotypes collected from 8 locations in Jilin Province, China [3 single generation (univoltine), 3 two generation (bivoltine), and 2 sympatric locations] in 2014. Bayesian and phylogenetic clustering partitioned mitochondrial cytochrome <i>c</i> oxidase subunit I (COI) haplotypes mostly into groups corresponding to historically uni- or bivoltine population origins, whereas samples from sympatric locations were interspersed between voltinism-specific clusters. Additionally, analyses of single nucleotide polymorphism (SNP) genotype data implicate voltinism, as opposed to geographic distance, as a factor contributing to differentiation among sample site. Temporal analysis of SNP genotypes from a sympatric location showed significant variation between adult moths collected within non-overlapping periods corresponding to bivoltine and univoltine flights. Regardless, only 11 of 257 SNP loci were predicted to be under selection, suggesting population genetic homogenization except at loci in proximity to factors responsible for locally adaptive or voltinism-specific traits. These findings provide evidence that divergent voltinism ecotype-specific traits and mitochondrial haplotypes may be maintained in allopatric as well as sympatric areas despite relatively high rates of nuclear gene flow.</p>
Data from: Genomic region associated with run-timing has similar haplotypes and phenotypic effects across three lineages of Chinook salmon
<p><span><span><span><span><span><span><span><span><span><span><span>Conserving life history variation is a stated goal of many management programs, but the most effective means by which to accomplish this are often far from clear. Early, premature migrating and late, mature migrating forms of Chinook salmon face unequal pressure from natural and anthropogenic forces. These forces may result in the diminishment of migration variation in some stocks because migration timing is known to be highly heritable. Genomic regions of chromosome 28 are known to be strongly associated with migration variation in adult Chinook salmon, but it remains unclear whether there is consistent association among the diverse populations of Chinook salmon. Therefore, application of this association for management may be premature. We examined the association of genetic variation in 28 markers on chromosome 28 surveyed with high-throughout genotyping (GT-seq) with individual run timing characteristics gleaned from passive integrated transponder recordings of over 5,000 Chinook salmon from the three phylogeographic lineages that inhabit the Columbia River Basin. Despite the strong genetic differences among them, the three lineages exhibited very similar genetic variants in the chromosome 28 region and moderate to strong association of these variants and run timing phenotypes. This is particularly notable for the interior stream-type lineage, which exhibits an earlier and more constrained migration of fish than the other lineages and which are exclusively premature when they enter freshwater. In both interior stream-type and interior-ocean type Chinook salmon, heterozygotes of the most strongly associated linkage groups are largely intermediate to homozygotes in migration timing, and while we make no robust conclusions about dominance, results indicate codominance or marginal partial dominance of the early migrating allele. Our results lend support to the cautious utilization of chromosome 28 variation in tracking and predicting run timing in these Chinook salmon.</span></span></span></span></span></span></span></span></span></span></span></p>
Drakaea glyptodon nuclear microsatellite and chloroplast haplotype data
<p class="CxSpFirst">Many orchids are characterized by small, patchily distributed populations. Resolving how they persist is important for understanding the ecology of this hyper-diverse family, many members of which are of conservation concern. <span>Ten</span> populations of the common terrestrial orchid <i>Drakaea glyptodon</i> from Southwest Australia were genotyped with ten nuclear and five chloroplast SSR markers. Levels and partitioning of genetic variation, and effective population sizes (<i>N</i><sub>e</sub>), were estimated. Spatial genetic structure of nuclear diversity, together with chloroplast data, are used to infer the effective number of seed parents per population. We found high genetic diversity, <i>N</i><sub>e</sub> values that generally exceed predictions based on the number of flowering individuals, and moderate levels of gene flow. Two populations were founded by < 5 colonists suggesting some populations are colonized by few seeds, with growth largely resulting from <i>in situ</i> recruitment. A value of 3.65 for <i>m</i><sub>p </sub>/<i>m</i><sub>s</sub> indicates that pollinators play a greater role than seed in introducing genetic diversity to populations via gene flow. Our results highlight that <i>D. glyptodon</i> is highly effective at persisting in patchily distributed populations. However, it is important to examine how insights from this common, widespread species transfer to species that are rare and/or occur in fragmented landscapes.</p>
Detecting selection using extended haplotype homozygosity (EHH)-based statistics in unphased or unpolarized data
<p>Analysis of population genetic data often includes the search for genomic regions with signs of recent positive selection. One of the approaches involves the concept of Extended Haplotype Homozygosity (EHH) and its associated statistics. These statistics typically need phased haplotypes and, some of them, polarized variants.<br> Here, we unify and extend previously proposed modifications to loosen these requirements. We compare the modified versions with the original ones by measuring the False Discovery Rate in simulated whole-genome scans and quantifying the overlap of inferred candidate regions in empirical data. We find that phasing information is indispensable for the accurate estimation of within-population statistics for all but very large samples and of cross-population statistics for small samples. Ancestry information, in contrast, is of lesser importance for both.<br> Our publicly available R package rehh incorporates the modified statistics presented here.</p>
Example of KRAS oncogene mutational regression and HLA haplotype characterization in oligo-metastatic colorectal cancer patient (ID: PAT2)
<p>KRAS oncogene mutational regression from primary tumour to lung metastasis (evaluated through TSO500 panel, Illumina Novaseq 6000 platform) and HLA haplotype characterization (through PCR) in a representative oligo-metastatic colorectal cancer patient (ID: PAT2).</p>
Data from: Community assembly and metaphylogeography of soil biodiversity: insights from haplotype-level community DNA metabarcoding within an oceanic island
<p>Most of our understanding of island diversity comes from the study of aboveground systems, while the patterns and processes of diversification and community assembly for belowground biotas remain poorly understood. Here we take advantage of a relatively young and dynamic oceanic island to advance our understanding of eco-evolutionary processes driving community assembly within soil mesofauna. Using whole organism community DNA (wocDNA) metabarcoding and the recently developed metaMATE pipeline, we have generated spatially explicit and reliable haplotype-level DNA sequence data for soil mesofauna assemblages sampled across the four main habitats within the island of Tenerife. Community ecological and metaphylogeographic analyses have been performed at multiple levels of genetic similarity, from haplotypes to species and supraspecific groupings. Broadly consistent patterns of local-scale species richness across different insular habitats have been found, whereas local insular richness is lower than in continental settings. Our results reveal an important role for niche conservatism as a driver of insular community assembly of soil mesofauna, with only limited evidence for habitat shifts promoting diversification. Furthermore, support is found for a fundamental role of habitat in the assembly of soil mesofauna, where habitat specialism is mainly due to colonisation and the establishment of preadapted species. Hierarchical patterns of distance decay at the community level and metaphylogeographical analyses support a pattern of geographic structuring over limited spatial scales, from the level of haplotypes through to species and lineages, as expected for taxa with strong dispersal limitations. Our results demonstrate the potential for wocDNA metabarcoding to advance our understanding of biodiversity.</p>
Data from: Within-trio tests provide little support for post-copulatory selection on MHC haplotypes in a free-living population
<p>Sexual selection has been proposed as a force that could maintain the diversity of major histocompatibility complex (MHC) genes in vertebrates. Potential selective mechanisms can be divided into pre-copulatory and post-copulatory, and in both cases the evidence for occurrence is mixed, especially in natural populations. In this study, we used a large number of parent-offspring trios that were diplotyped for MHC class II genes in a wild population of Soay sheep (<i>Ovis aries</i>) to examine whether there was within-trio post-copulatory selection on MHC genes at both the haplotype and diplotype levels. We found there was transmission ratio distortion of one the eight MHC class II haplotype (E) which was transmitted less than expected by fathers, and transmission ratio distortion of another haplotype (A) which was transmitted more than expected by chance to male offspring. However, in both cases these deviations were not significant after correction for multiple tests. In addition, we did not find any evidence of post-copulatory selection on diplotype level. These results imply given known parents, there is no strong post-copulatory selection on MHC genes in this population.</p>
Data from: Haplotype sequence collection of ABO blood group alleles by long-read sequencing reveals putative A1-diagnostic variants
<p>In the era of blood group genomics, reference collections of complete and fully-resolved blood group gene alleles have gained high importance. For most blood groups, however, such collections are currently lacking, as resolving full-length gene sequences as haplotypes (i.e. separated maternal/paternal origin) remains exceedingly difficult with both Sanger and short-read next-generation sequencing. Using the latest third-generation long-read sequencing, we generated a collection of fully-resolved sequences for all six main <em>ABO</em> allele groups: <em>ABO</em>*<em>A1</em>/<em>A2</em>/<em>B</em>/<em>O.01.01</em>/<em>O.01.02</em>/<em>O.02</em>. We selected 77 samples from an <em>ABO</em> genotype dataset (n=25,200) of serologically-typed Swiss blood donors. The entire <em>ABO</em> gene was amplified in two overlapping long-range PCRs (covering ~23.6 kb) and sequenced by long-read Oxford Nanopore sequencing. For quality validation, two samples per <em>ABO</em> group were re-sequenced using Illumina and PacBio technology. All 154 full-length <em>ABO</em> sequences were resolved as haplotypes. We observed novel, distinct sequence patterns for each <em>ABO</em> group. Most genetic diversity was found between, not within, <em>ABO</em> groups. Phylogenetic tree and haplotype network analyses highlighted distinct clades of each <em>ABO</em> group. Strikingly, our data uncovered four genetic variants putatively specific for <em>ABO</em>*<em>A1</em>, for which direct diagnostic targets are currently lacking. We validated <em>A1</em>-diagnostic potential using whole-genome data (n=4,872) of a multi-ethnic cohort. Overall, our sequencing strategy proved powerful for producing high-quality <em>ABO</em> haplotypes and holds promise for generating similar collections for other blood groups. The publicly available collection of 154 haplotypes will serve as a valuable resource for molecular analyses of <em>ABO</em>, as well as studies about the function and evolutionary history of <em>ABO</em>.</p>
Snakemake report for manuscript "Orthanq: transparent and uncertainty-aware haplotype quantification with application in HLA-typing"
<p>For viewing the report, unzip the file and open index.html in your browser.</p>
Figure 2. Geographic distribution and haplotype networks for 12S in Effects of Quaternary climatic oscillations over the Chacoan fauna: phylogeographic patterns in the southern three-banded armadillo Tolypeutes matacus (Cingulata: Chlamyphoridae)
Figure 2. Geographic distribution and haplotype networks for 12S (top panels) and control region (bottom panels). The panels on the left plot the geographical distribution and frequency of haplotypes in the different localities analysed. Localities were labelled according to their ID (see Table 1). The right panels show the haplotype networks, where the dashes on the lines represent mutations, and the black circles represent intermediate variants not found. Principal Chacoan rivers are shown in light-blue labels. Capitalized labels indicate names of Argentinean provinces, and labels with all letters in uppercase refer to neighbouring countries.
Haplotype analysis of the mitochondrial DNA d-loop region reveals the maternal origin and historical dynamics among the indigenous goat populations in east and west of the Democratic Republic of Congo (DRC)
<p><span>This study aimed at assessing haplotype diversity and population dynamics of three Congolese indigenous goat populations that included Kasai goat (KG), small goat (SG), and dwarf goat (DG) of the Democratic Republic of Congo (DRC). The 1,169 bp <em>d-loop</em> region of mitochondrial DNA (mtDNA) was sequenced for 339 Congolese indigenous goats. The total length of sequences was used to generate the haplotypes and evaluate their diversities, whereas the hypervariable region (HVI, 453 bp) was analyzed to define the maternal variation and the demographic dynamic. A total of 568 segregating sites that generated 192 haplotypes were observed from the entire <em>d-loop</em> region (1,169 bp <em>d-loop</em>). Phylogenetic analyses using reference haplotypes from the six globally defined goat mtDNA haplogroups showed that all the three Congolese indigenous goat populations studied clustered into the dominant haplogroup A, as revealed by the Neighbor-joining (NJ) tree and median-joining (MJ) network. Nine haplotypes were shared between the studied goats and goat populations from Pakistan (1 haplotype), Kenya, Ethiopia and Algeria (1 haplotype), Zimbabwe (1 haplotype), Cameroon (3 haplotypes), and Mozambique (3 haplotypes). The population pairwise analysis (<em>F<sub>ST</sub></em>) indicated a weak differentiation between the Congolese indigenous goat populations. Negative and significant (<em>p</em>-value < 0.05) values for <em>F</em>u's <em>F</em>s (-20.418) and Tajima's (-2.189) tests showed the expansion in the history of the three Congolese indigenous goat populations. These results suggest a weak differentiation and a single maternal origin for the studied goats. This information will contribute to the improvement of the management strategies and long-term conservation of indigenous goats in DRC</span><span>.</span></p>
Replication data for: HairSplitter: separating haplotypes with long reads
<p>Replication data for the paper "HairSplitter: separating haplotypes with long reads".</p> <p>Contains 1) Reads used to assemble when not available on public repositories, 2) The assemblies obtained using Flye or hifiasm, 3) The separated assemblies obtained by running the benchmarked software (HairSplitter, Strainberry, stRainy, HaploDMF, iGDA, Strainline) and 4) command lines used to launch the assemblies.</p>
Fig. 5 Haplotype networks for datasets IV in Dugesia hepta and Dugesia benazzii (Platyhelminthes: Tricladida): two sympatric species with occasional sex?
Fig. 5 Haplotype networks for datasets IV (a Dunuc12) and III (b Cox1). Haplotypes are depicted as individual circles which are proportional to their abundancy (number of sequences), highlighted in a white square. Mutations are either depicted with black bars or black triangles when the number of mutations between linked haplotypes is equal to or exceeds a certain
Fig. 3 Median-joining haplotype network obtained for 867 in Diversification and evolutionary history of brush-tailed mice, Calomyscidae (Rodentia), in southwestern Asia
Fig. 3 Median-joining haplotype network obtained for 867 bp of mitochondrial Cyt b of the genus Calomyscus. Circle size is relative to haplotype frequency; black circles represent extinct or unsampled
FIGURE 2. Haplotype network calculated from a 360 in A new species of Uroplatus (Gekkonidae) from Ankarana National Park Madagascar, of remarkably high genetic divergence
FIGURE 2. Haplotype network calculated from a 360 bp segment of the nuclear gene CMOS for the species in the Uroplatus ebenaui group.
Fig. 2 Mitochondrial haplotype network using the 590 in Differentiation of North African foxes and population genetic dynamics in the desert-insights into the evolutionary history of two sister taxa, Vulpes rueppellii and Vulpes vulpes
Fig. 2 Mitochondrial haplotype network using the 590-bp concatenated sequences from Cyt-b and D-loop and a total of 46 sequences (same as in Fig. 1, except for C. lupus not being used as an outgroup in the TCS network). a Neighbour-Net network based on uncorrected patristic distances as implemented in SPLITSTREE. Canis lupus (DQ480504) was used as an outgroup. Numbers indicate bootstrap values. Scale bar represents 0.01 sequence divergence. Highlighoed are the four species, the three V. vulpes clades and the location within the network of the V. vulpes sample from Egypt. Colour patterns are concordant with Fig. 1 and b. b Statistical parsimony network assuming a 95 % parsimony threshold, as constructed by TCS. Symbol size and branch lengths are proportional to the number of shared individuals per haplotype and the number of mutational steps amongst haplotypes, respectively. Numbers in black background also refer to the number of mutation steps between species and V. vulpes clades. Symbols and colours are concordant with Fig. 1 and a. Haplotype codes, sample origin and corresponding accession numbers are available in Online Resource Table S1
Fig. 3 Haplotype-networks for a in Species status and population structure of mussels (Mollusca: Bivalvia: Mytilus spp.) in the Wadden Sea of Lower Saxony (Germany)
Fig. 3 Haplotype-networks for a COI (n haplotypes 0 15; n sequences 0 111), b VD1 (n haplotypes 0 17; n sequences 0 81), and c the combined data set (n haplotypes 0 16; n sequences 0 64). The sizes of the symbols are proportional to the number of individuals sharing that haplotypes (unique haplotypes are not included), with the rectangular haplotype having had the largest outgroup weight. Each node corresponds to one mutation step. The patterns used for the symbols match those used in the geographical distribution maps (Fig. 2)
Fig. 2 in Using haplotype networks, estimation of gene flow and phenotypic characters to understand species delimitation in fungi of a predominantly Antarctic Usnea group (Ascomycota, Parmeliaceae)
Fig. 2 Enlargement of the Usnea aurantiaco-atra group of the Bayesian inference (Fig. 1) depicting 101 taxa. Posterior probabilities ≥ 0.95 are visualized by bold branches. Colors match the sampling site colours in Fig. 4a. A! fertile specimen with apothecia, S! vegetative reproduction via soralia. Most important clades of the nested clade analysis (Fig. 4a,b) are plotted on the phylogenetic tree
Fig. 8 Haplotype network for 44 in New insights into the phylogeny and taxonomy of Chinese species of Gagea (Liliaceae)-speciation through hybridization
Fig. 8 Haplotype network for 44 cpDNA haplotypes (psbA-trnH IGS+trnL-trnF IGS) including 38 sequences of representatives of Gagea sect. Gagea: G. aipetriensis (aip), G. ancestralis, G. angelae (ang), G. artemczukii (art), G. capusii (cap), G. erubescens (eru), G. helenae (hel), G. huochengensis (huo), G. lutea (lut), G. nakaiana (nak), G. paczoskii (pac), G. podolica (pod), G. pomeranica (pom), G. pratensis (pra), G. pusilla (pus), G. rubicunda (rub), G. shmakoviana (shm), G. terraccianoana (ter), G. tisoniana (tis), G.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.