Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
423
datasets available to search
ShareScore release 0.7.1
Dataset results
423 results for “haplotypes”
Protein haplotype sequences obtained by ProHap from the Haplotype Reference Consortium Release 1.1 dataset
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Haplotype Reference Consortium, Release 1.1 (<a href="https://ega-archive.org/datasets/EGAD00001002729" target="_blank" rel="noopener">https://ega-archive.org/datasets/EGAD00001002729</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>Release 1.1 of the HRC is provided aligned with the GRCh37 reference genome. We have performed a liftover to the GRCh38 reference using GeneBe (https://genebe.net/tools/liftover). Variants for which the reported alternative allele is considered as reference in GRCh38 were removed. A threshold of 1% minor allele frequency was applied to filter the remaining variants. After translation, a frequency threshold of 0.5% was applied to filter the resulting unique non-canonical sequences. The complete configuration file for the ProHap run is attached to this repository.</p> <p>This dataset contains one compressed directory, contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
Protein haplotype sequences obtained by ProHap from the Human Pangenome Reference Consotruim dataset
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the Human Pangenome Reference Consotruim (HPRC), first release (<a href="https://github.com/human-pangenomics/hpp_pangenome_resources">https://github.com/human-pangenomics/hpp_pangenome_resources</a>), 44 samples. We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts.</p> <p>This repository contains one database created using all 43 samples of the HPRC release (the haplotypes of the sample NA21309 did not encode any non-canonical sequences), and then a database for each of the 43 samples separately. No filtering on allele frequency or haplotype frequency was applied in any of the databases. The complete configuration file for the ProHap run is attached to this repository.</p> <p>There is one compressed directory for each of the databases, containing the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>For this dataset, only the simplified format is provided. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
Data from: Complex population structure and haplotype patterns in Western Europe honey bee from sequencing a large panel of haploid drones
<p>This vcf file contains 7.023.689 SNPs and 870 honey bee samples, as described in the paper "Complex population structure and haplotype patterns in Western Europe honey bee from sequencing a large panel of haploid drones" by Wragg et al., available at https://doi.org/10.1101/2021.09.20.460798 as preprint.</p> <p>Eight hundred and seventy haploid drone samples from several honey bee subspecies hybrids were sequenced and aligned to the HAv3.1 reference genome. Sequence read alignment and genotyping quality filters were used to obtain a selection of 7.023.689 high-quality SNPs. The file Diversity_Study_629_Samples.txt corresponds to the 629 unique samples that were used for the diversity study described in the paper and can be used to recreate the restricted diversity dataset using bcftools or an equivalent software.</p> <p>Having sequenced haploid drones, heterozygous SNPs resulting from duplicated regions could be filtered out and the data is phased.</p>
Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S1 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"
<p>This dataset contains the raw read counts and phased SNP counts for every single cell in the sequencing datasets of breast cancer patient S1 from “Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL” [Zaccaria & Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S1. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S1 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz </em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz </em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>
Raw read counts and phased SNP counts for every single cell in the sequencing datasets of the breast cancer patient S0 from "Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL"
<p>This dataset contains the raw read counts and phased SNP counts for every single cell in the sequencing datasets of breast cancer patient S0 from “Characterizing allele- and haplotype-specific copy numbers in single cells with CHISEL” [Zaccaria & Raphael, 2020]. These data enable the full reproduction of all the results in the related manuscript for breast cancer patient S0. Specifically, the data are provided in two files for every dataset <em>DAT</em> of patient S0 with the following format:</p> <ol> <li><em>DAT.raw</em>_<em>read</em>_<em>counts.bed.gz </em>is a multi-cell BED file containing the raw read counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>START: the starting genomic position of a genomic bin in the chromosome</li> <li>END: the ending genomic position of the genomic bin in the chromosome</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>NORMAL: the raw read count for the specified bin from a matched-normal sample</li> <li>COUNT: the raw read count for the specified bin in the specified cell</li> <li>RDR: the estimated read-depth ratio for the specified bin in the specified cell</li> </ul> </li> <li><em>DAT.phased</em>_<em>snps</em>_<em>counts.pos.gz </em>is a multi-cell POS file containing the phased SNP counts in the following fields: <ul> <li>CHROMOSOME: the name of a human chromosome</li> <li>POS: the genomic position in the chromosome of a germline SNP</li> <li>CELL: the cell barcode that uniquely identifies a cell</li> <li>COUNT_HAPLOTYPE_A: the count of reads that cover the SNP and that belong to haplotype A in the specified cell</li> <li>COUNT_HAPLOTYPE_B: the count of reads that cover the SNP and that belong to haplotype B in the specified cell</li> </ul> </li> </ol> <p>All the files have been compressed using standard <em>gzip</em>.</p>
Large structural variations in the haplotype-resolved African cassava genome
<p>Cassava TME7 haplotype resolved assemblies and annotation</p> <p> </p> <p>ABSTRACT:</p> <p>Cassava (<em>Manihot esculenta</em> Crantz, 2n=36) is a global food security crop. Cassava has a highly heterozygous genome, high genetic load, and genotype-dependent asynchronous flowering. It is typically propagated by stem cuttings and any genetic variation between haplotypes, including large structural variations, is preserved by such clonal propagation. Traditional genome assembly approaches generate a collapsed haplotype representation of the genome. In highly heterozygous plants, this results in artifacts and an oversimplification of heterozygous regions. We used a combination of Pacific Biosciences (PacBio), Illumina, and Hi-C to resolve each haplotype of the genome of a farmer-preferred cassava line, TME7 (Oko-iyawo). PacBio reads were assembled using the FALCON suite. Phase switch errors were corrected using FALCON-Phase and Hi-C read data. The ultra-long-range information from Hi-C sequencing was also used for scaffolding. Comparison of the two phases revealed more than 5,000 large haplotype-specific structural variants affecting over 8 Mb, including insertions and deletions spanning thousands of base pairs. The potential of these variants to affect allele specific expression was further explored. RNA-seq data from 11 different tissue types were mapped against the scaffolded haploid assembly and gene expression data are incorporated into our existing easy-to-use web-based interface to facilitate use by the broader plant science community. These two assemblies provide an excellent means to study the effects of heterozygosity, haplotype-specific structural variation, gene hemizygosity, and allele specific gene expression contributing to important agricultural traits and further our understanding of the genetics and domestication of cassava.</p>
Appendix List of samples of deep frozen frog legs with purchase date, collection number, haplotype number, taxonomic identification, tibia length (TL) and estimated snout vent length (SVL). in Which frog's legs do froggies eat? The use of DNA barcoding for identification of deep frozen frog legs (Dicroglossidae, Amphibia) commercialized in France
Appendix List of samples of deep frozen frog legs with purchase date, collection number, haplotype number, taxonomic identification, tibia length (TL) and estimated snout vent length (SVL).
Direct chromosome-length haplotyping by single-cell sequencing.
<p>Selected Strand-seq libraries from PMID:27646535 study. Data were originally shared on the European Nucleotide Archive (http://www.ebi.ac.uk/ena) under the accession number: PRJEB14185</p>
Multi-cell type deconvolution using a probabilistic model for single-molecule DNA methylation haplotypes
<p>Files required to run deconvolution with CelFIE-ISH and Epistate, in U250 regions from Loyfer et al. 2023, in both "pat" and "epiread" formats. </p>
Fig. 2. Haplotype network calculated from the E in Molecular assessment of commercial and laboratory stocks of Eisenia spp. (Oligochaeta: Lumbricidae) from South Africa
Fig. 2. Haplotype network calculated from the E. andrei COI haplotypes found in the South African earthworm groups investigated. The size of the circles is proportional to the number of earthworms sharing the same haplotype. The numbers on the branches indicate the positions of mutations on the COI sequences, mv1 represents a median vector (intermediate haplotypes, not found in this study).
Figure 5. Unrooted haplotype network. Each circle represents a in Distribution and molecular differentiation of Culex pipiens complex species in the Middle and Eastern Black Sea Regions of Turkey
Figure 5. Unrooted haplotype network. Each circle represents a haplotype, and the lines above each link indicate one mutation. Small black dots indicate intermediate, missing, or unsampled haplotypes.
Haplotype-Resolved and Gap-Free Genome of a Floating Aquatic Plant from the Oryzeae Tribe, Hygroryza aristata
<p><em><span>Hygroryza aristata</span></em><span> (Retz.) Nees ex Wight & Arn.</span><span> </span><span>is a floating aquatic plant. <span>Genomic DNA and RNA samples of </span><em><span>H. aristata</span></em><span> were extracted from plants clonally propagated from a single individual. Long-read sequencing of PacBio HiFi and ultra-long (UL) ONT (read lengths > 100 kb), and short-read sequencing of Hi-C, WGS, and RNA-seq, were performed. </span></span></p> <p><span>For genome assembly, 31.91 Gb of PacBio HiFi and 22.36 Gb of UL-ONT sequencing data sets were utilized. Assembly was conducted using HiFiAsm (v0.20.0-r639) under HiFi + UL-ONT mode with the following parameters: -l 3 -r 5 -a 6 -n 10 --ctg-n 10 -w 63 -k 63. Chromosome IDs and strand directions were determined by aligning the assemblies to the rice (</span><em><span>O. sativa</span></em><span>) genome. The resulting assemblies of unphased two haplotypes, designated as hap1 and hap2, were obtained. </span></p> <p><span><span>Both hap1 and hap2 are complete genomes,<span> </span>each comprising 12 chromosomes with genome sizes of 349.74 Mb and 347.98 Mb, respectively. Notably, both haplotypes are gap-free. Telomere detection using Seqtk telo (v1.4-r122) revealed that each haplotype contains 23 telomeres. In conclusion, this study presents a haplotype-resolved and gap-free genome assembly. </span></span></p>
Chromosome-scale, haplotype-resolved genome assembly of Suaeda glauca
<p><em>Suaeda glauca</em>is an annual herb of Suaeda and an important saline-alkali plant resource, which is widespread on beaches and saline lands around the world. It is also a good candidate for food, feed, and drug development. There has been no publication of the <em>Suaeda glauca</em>genome assembly, limiting the evolutionary study of Amaranthaceae and the bioavailability of <em>Suaeda glauca</em>.</p> <p>Using PacBio HiFi and Hi-C sequencing data, we successfully generated chromosome-scale, haplotype-resolved assemblies of the <em>Suaeda glauca</em>genome. The size of the final primary assembly was 622.95 Mb, and the contig N50 was 19.42 Mb, which was successfully anchored to 9 chromosomes, accounting for 96.79% of the total assembly size. The repeat content and genome size of <em>Suaeda glauca</em>are much higher than those of the same genus <em>Suaeda aralocaspica</em>, presumably due to a recent burst of LTR insertions. Using HiFi reads, we assembled the complete circular chloroplast genome of <em>Suaeda glauca</em>. Through gene family and phylogenetic tree analysis, it was shown that <em>Suaeda glauca</em>and <em>Suaeda aralocaspica</em>differentiated at ~26.36 million years ago (MYA), and Amaranthaceae species began to differentiate at ~52.00 MYA.</p>
Parent-of-origin detection and chromosome-scale haplotyping using long-read DNA methylation sequencing and Strand-seq
<p>Hundreds of loci in human genomes have alleles that are methylated differentially according to their parent of origin. These imprinted loci generally show little variation across tissues, individuals, and populations. We show that such loci can be used to distinguish the maternal and paternal homologs for all autosomes, without the need for the parental DNA. We integrate methylation-detecting nanopore sequencing with the long-range phase information in Strand-seq data to determine the parent of origin of chromosome-length haplotypes for both DNA sequence and DNA methylation in five trios with diverse genetic backgrounds.</p>
Vamos: VNTR annotation using efficient motif sets, 64-haplotype motifs
<p>This is the efficient motif set for 382,492 VNTR loci in 64 haplotype-resolved assemblies (10.1126/science.abf7117). The initial release is with compression parameter (delta) = 0.1</p>
Data for paper MethPhaser: methylation-based haplotype phasing of human genomes
<p>Data for paper MethPhaser: methylation-based haplotype phasing of human genomes. </p> <p>Files start with R9 or R10 are HG002 sample data.</p> <p>Zip files are block connection intermediate for MethPhaser. </p> <p>GTFs are block assignments. </p>
Fig. 3 in Haplotype variation in the Physa acuta group (Basommatophora): genetic diversity and distribution in Serbia Abstract
Fig. 3: Haplotype networks from 43 Physa acuta group specimens, obtained using statistical parsimony (TCS). Circles represent specific haplotypes; the size of the circles reflects the number of individuals with a particular haplotype (not to scale); the dots between the circles represent mutational steps.
Fig. 2 in Haplotype variation in the Physa acuta group (Basommatophora): genetic diversity and distribution in Serbia Abstract
Fig. 2: Phylogenetic trees based on mt16S rDNA, obtained using the Maximum Likelihood (ML) method. Bootstrap values are indicated below the branches. Scale bar indicates the number of substitutions per site.
Fig. 3 in Variation In Cone And Seed Morphology Traits Among The Mitochondrial Dna Haplotypes Of Scots Pine (Pinus Sylvestris L.)
Fig. 3. Dependence of seed number per cone on cone length for the type A and type B mitotypes of Scots pine. Individual cone values are shown.
Protein haplotype sequences obtained by ProHap from the 1000 Genomes Project data set
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the 1000 Genomes Project, aligned with the GRCh38 genome build (<a href="https://www.internationalgenome.org/data-portal/data-collection/grch38">https://www.internationalgenome.org/data-portal/data-collection/grch38</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts. The complete configuration file for each ProHap run is attached to this repository.</p> <p>This data set contains six compressed directories, five representing the superpopulations included in the 1000 Genomes Project (<a href="https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project">https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project</a>), and one created using all the samples included in the 1000 Genomes data set:</p> <ul> <li>AFR - African</li> <li>AMR - American</li> <li>EUR - European</li> <li>SAS - South Asian</li> <li>EAS - East Asian</li> <li>ALL - all participants in the 1000 Genomes Project</li> </ul> <p>Each of the directories contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap, using alleles with at least 1 % frequency within the selected population</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the simplified fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.