Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

756

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

756 results for “Hi-C”

Learn how ShareScore rates datasets ↗
zenodo48/100

Test datasets for Hi-C scaffolding

<p>We provided two datasets for testing Hi-C scaffolding tools. For the CHM13 test dataset, we randomly chunked the first 10Mb of chr1, chr2 and chr3 of the T2T-CHM13v1.1 human genome assembly (Nurk et&nbsp;al. 2022) into 57 contigs. The Hi-C data downloaded from the telomere-to-telomere consortium GitHub repository (https://github.com/marbl/CHM13) were mapped to the reference genome and the reads mapped to these regions were extracted to generate&nbsp;Hi-C alignment files. For the LYZE01 test dataset, the Saccharomyces cerevisiae strain W303 genome assembly (Matheson et al. 2017) was split at positions with gaps (&lsquo;N&rsquo;) to get the original contigs. An&nbsp;independent Hi-C data library was&nbsp;downloaded from the NCBI repository (GEO Accession GSM2417297) and&nbsp;downsampled to approximately 20X. The downsampled Hi-C data were mapped to the contigs to generate Hi-C alignment files.</p> <p>We provided five files for each test dataset:&nbsp;the contig file in FASTA format, the FASTA index file generated with SAMtools faidx command, and the Hi-C alignment file in BAM format sorted by coordinate,&nbsp;in BAM format sorted by query names (with the identifier &#39;qn&#39; in the file name), and in BED format.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Coolpup.py – a versatile tool to perform pile-up analysis of Hi-C data

<p>Data used for the analysis and to generate the figures. After un-tar-ing, the data is present in three folders. /coolers contains Hi-C data in the .cool format, and associated text files. /beds contains .bed and .bedpe files used in the analysis. /enrichment_jsons contains .json files with results of the &quot;loop-ability&quot; analysis. Code for the data analysis is available here: https://github.com/Phlya/coolpuppy_paper</p>

opencc-by-4.0Mar 2019View details →
zenodo40/100

IGM Population of HFF structures using Hi-C, laminB1 DamID, 3D HIPMAp FISH and single cell SPRITE data

<p>This repository accompanies the manuscript &quot;<strong>Integrative Genome Modeling Platform reveals essentiality of rare contact events in 3D genome organizations</strong>&quot;, to appear in Nat. Methods (2022), see also&nbsp;https://www.biorxiv.org/content/10.1101/2021.08.22.457288v1.</p> <p>It&nbsp;contains the preprocessed input data files (Hi-C, laminB1 DamID, 3D&nbsp;HIPMAp FISH and single cell SPRITE) for the HFF fibroblast cell line to be used in the Integrative Genome Modeling platform (IGM) developed in the Alber lab at UCLA (https://github.com/alberlab/igm).</p> <p>Also, we provide the configuration file to run IGM with those datasets, as we did in generating the HDSF population discussed in the accompanying manuscript. Such population is also provided as an &quot;hss&quot; file. Documentation and a simple demo/tutorial on how IGM can be run is given on the Alber lab Github @&nbsp;https://github.com/alberlab/igm.</p> <p>All files can be read in using the&nbsp;<em>h5py</em> and <em>alabtools</em> (available @https://github.com/alberlab/alabtools) Python packages. More detailed information is provided in the manuscript and associated Supplementary Information file.&nbsp;&nbsp;</p> <p>For any inquiry/suggestions/doubts please reach out to Lorenzo Boninsegna (bonimba@g.ucla.edu) or Dr. Frank Alber (falber@g.ucla.edu).</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Processed Hi-C contact matrices for "Three invariant Hi-C interaction patterns: applications to genome assembly"

<p>Processed Hi-C interaction matrices, saved in numpy npz format.</p> <p>Matrices were processed using Dekker lab cMapping pipeline.</p> <p>Raw sequence data was taken from:</p> <p>Hap1: Haarhuis et al 10.1016/j.cell.2017.04.013</p> <p>IMR90, H1ESC, MESC, MCORTEX: Dixon et al 10.1038/nature11082</p> <p>Worm: Crane et al 10.1038/nature14450</p> <p>Caulobacter: Le et al 10.1126/science.1242059</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2018View details →
zenodo40/100

Processed Hi-C contact matrices for "Single-cell DNA replication profiling identifies spatiotemporal developmental dynamics of chromosome organization"

<p>Processed Hi-C interaction matrices (iterative correction) saved in .hic format (40kb bins).</p> <p>.hic files were generated by juicer pipeline using processed Hi-C interaction matrices.</p> <p>Only <em>cis&nbsp;</em>interactions were available.</p> <p>To extract the data, please see&nbsp;</p> <p>https://github.com/aidenlab/juicer/wiki/Data-Extraction</p>

opencc-by-4.0Aug 2019View details →
dryad40/100

Hi-reComb: Constructing recombination maps from bulk gamete Hi-C sequencing

Open the record for dataset details and reuse information.

publicJul 2025View details →
zenodo36/100

FISH datasets used in Zou et al. integrating multi-track Hi-C data for genome-scale reconstruction of 3D chromatin structure

<p>This upload contains the FISH datasets used in Zou et al. integrating multi-track Hi-C data for genome-scale reconstruction of 3D chromatin structure.</p> <p>If you use the datasets, we would be grateful if you cited the following paper:</p> <p>Zou, C., Zhang, Y., Ouyang, Z. (2016) HSA: integrating multi-track Hi-C data for genome-scale reconstruction of 3D chromatin structure. Genome Biology, 17: 40.</p>

opengpl-2.0Feb 2016View details →
zenodo36/100

A dataset for testing a single-cell Hi-C softwares

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
dryad36/100

Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C

<p><i>Telopea speciosissima, </i>the New South Wales waratah, is an Australian endemic woody shrub in the family Proteaceae. Waratahs have great potential as a model clade to better understand processes of speciation, introgression and adaptation, and are significant from a horticultural perspective. Here, we report the first chromosome-level genome for <i>T. speciosissima</i>. Combining Oxford Nanopore long-reads, 10x Genomics Chromium linked-reads and Hi-C data, the assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 97.8 % of Embryophyta BUSCOs 'Complete'. We present a new method in Diploidocus (<a href="https://github.com/slimsuite/diploidocus">https://github.com/slimsuite/diploidocus</a>) for classifying, curating and QC-filtering scaffolds, which combines read depths, <i>k</i>-mer frequencies and BUSCO predictions. We also present a new tool, DepthSizer (<a href="https://github.com/slimsuite/depthsizer">https://github.com/slimsuite/depthsizer</a>), for genome size estimation from the read depth of single-copy orthologues and estimate the genome size to be approximately 900 Mb. The largest 11 scaffolds contained 94.1 % of the assembly, conforming to the expected number of chromosomes (2<i>n</i> = 22). Genome annotation predicted 40,158<code> </code>protein-coding genes, 351 rRNAs and 728 tRNAs. We investigated <i>CYCLOIDEA </i>(<i>CYC</i>)<i> </i>genes, which have a role in determination of floral symmetry, and confirm the presence of two copies in the genome. Read depth analysis of 180 'Duplicated' BUSCO genes using a new tool, DepthKopy (<a href="https://github.com/slimsuite/depthkopy">https://github.com/slimsuite/depthkopy</a>), suggests almost all are real duplications, increasing confidence in the annotation and highlighting a possible need to revise the BUSCO set for this lineage. The chromosome-level <i>T. speciosissima</i> reference genome (Tspe_v1) provides an important new genomic resource of Proteaceae to support the conservation of flora in Australia and further afield.</p>

opencc-zeroDec 2021View details →
zenodo36/100

Human assemblies evaluated in the hifiasm (Hi-C) paper

<p>This repository contains all human assemblies evaluated in the paper titled:&nbsp;<strong>Haplotype-resolved assembly of diploid genomes without parental data</strong>.&nbsp;Non-human assemblies are available at&nbsp;<a href="https://zenodo.org/record/5953248">10.5281/zenodo.5953248</a>. &quot;*.DipAsm*&quot; by&nbsp;<a href="https://www.nature.com/articles/s41587-020-0711-0">Garg et al (2020)</a>&nbsp;and &quot;*.PAGS*&quot; by&nbsp;<a href="https://www.nature.com/articles/s41587-020-0719-5">Porubsky et al (2020)</a>.&nbsp;The rest of the assemblies were generated by Cheng et al.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Contigs scaffolding with Hi-C for plant genomes

<p>Example contigs fasta file use in the protocol:&nbsp;Contigs scaffolding with Hi-C for plant genomes</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Galaxy Hi-C Training material dm3

<p>Hi-C data for Galaxy training, dm3 cells.</p>

opencc-by-4.0Feb 2018View details →
zenodo36/100

Supplementary Data for "FreeHi-C: high fidelity Hi-C data simulation for benchmarking and data augmentation"

<p>Simulated and processed data utilized in the paper &quot;FreeHi-C: high fidelity Hi-C data simulation for benchmarking and data augmentation&quot; for visualization and further analysis.</p>

opencc-by-4.0Jul 2019View details →
zenodo36/100

In silico prediction of high-resolution Hi-C interaction matrices (part III)

<p>The uploaded files are source datasets for the HiC-Reg approach. HiC-Reg is a regression based method that predict contact counts from one-dimensional regulatory signals such as epigenetic marks and regulatory protein binding. See more details here (<a href="https://github.com/Roy-lab/HiC-Reg">https://github.com/Roy-lab/HiC-Reg</a>).&nbsp;There are a total of six files in this dataset:&nbsp;Data.tgz, Gm12878.tgz, Hmec.tgz, K562.tgz, Huvec.tgz&nbsp;and Nhek.tgz.&nbsp;The Data.tgz&nbsp;include&nbsp;predictions and other downstream analysis such as feature importance analysis, significant interaction calling, and data files for select figures. The Gm12878.tgz, K562.tgz, Huvec.tgz, Hmec.tgz and Nhek.tgz&nbsp;contain trained models, predictions, feature files for two chromosomes for in each&nbsp;cell line.</p> <p>This is part III of the dataset which contains Nhek.tgz and Data.tgz.</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

In silico prediction of high-resolution Hi-C interaction matrices (part I)

<p>The uploaded files are source datasets for the HiC-Reg approach. HiC-Reg is a regression based method that predict contact counts from one-dimensional regulatory signals such as epigenetic marks and regulatory protein binding. See more details here (<a href="https://github.com/Roy-lab/HiC-Reg">https://github.com/Roy-lab/HiC-Reg</a>).&nbsp;There are a total of six files in this dataset:&nbsp;Data.tgz, Gm12878.tgz, Hmec.tgz, K562.tgz, Huvec.tgz&nbsp;and Nhek.tgz.&nbsp;The Data.tgz&nbsp;include&nbsp;predictions and other downstream analysis such as feature importance analysis, significant interaction calling, and data files for select figures. The Gm12878.tgz, K562.tgz, Huvec.tgz, Hmec.tgz and Nhek.tgz&nbsp;contain trained models, predictions, feature files for two chromosomes for in each&nbsp;cell line.</p> <p>This is part I of the dataset which contains Gm12878.tgz and Hmec.tgz.</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

In silico prediction of high-resolution Hi-C interaction matrices (part II)

<p>The uploaded files are source datasets for the HiC-Reg approach. HiC-Reg is a regression based method that predict contact counts from one-dimensional regulatory signals such as epigenetic marks and regulatory protein binding. See more details here (<a href="https://github.com/Roy-lab/HiC-Reg">https://github.com/Roy-lab/HiC-Reg</a>).&nbsp;There are a total of six files in this dataset:&nbsp;Data.tgz, Gm12878.tgz, Hmec.tgz, K562.tgz, Huvec.tgz&nbsp;and Nhek.tgz.&nbsp;The Data.tgz&nbsp;include&nbsp;predictions and other downstream analysis such as feature importance analysis, significant interaction calling, and data files for select figures. The Gm12878.tgz, K562.tgz, Huvec.tgz, Hmec.tgz and Nhek.tgz&nbsp;contain trained models, predictions, feature files for two chromosomes for in each&nbsp;cell line.</p> <p>This is part II of the dataset which contains K562.tgz and&nbsp;Huvec.tgz.</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

Single-cell Hi-C matrix Nagano 2017

<p>Single-cell Hi-C interaction matrix in mcool format, provided in 10kb and 1Mb. Data based on Nagano 2017:&nbsp;Cell-cycle dynamics of chromosomal organization at single-cell resolution</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

DiCARN-DNase: Enhancing Cell-to-Cell Hi-C Resolution Using Dilated Cascading ResNet with Self-Attention and DNase-seq Chromatin Accessibility Data

<p>This repository contains the training, validation, and testing data for the DiCARN-DNase project. The GitHub repository is available via https://github.com/OluwadareLab/DiCARN_DNase</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Novel canine high-quality metagenome-assembled genomes by long-read metagenomics together with Hi-C proximity ligation

<p>We characterized a canine fecal sample of a healthy dog by combining a long-read metagenomics assembly (Nanopore sequencing) with Hi-C cross-linking data, and further correction of the frameshift errors. We retrieved and characterized 27 HQ MAGs and seven MQ MAGs considering MIMAG criteria.</p> <p>Find in this repository the final Hi-C genomics bins (CanMAG_XX-HiCbin.fa), including both the genome and the extra-chromosomal elements within the bin.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Long reads and Hi-C sequencing illuminate the two-compartment genome of the model arbuscular mycorrhizal symbiont Rhizophagus irregularis

<p>This repository contains annotations for the strains of <em>R. irregularis</em> chromosome assemblies.</p>

opencc-by-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record