Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

15,745

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

15,745 results for “ChIP-seq”

Learn how ShareScore rates datasets ↗
zenodo44/100

DCSsim (simulated) and DCSsub (sub-sampled) ChIP-seq data for benchmarking DCS tools.

<p>These data are the results from five independent runs of DCSsim and DCSsub for TF, sharp and broad mark signals in 50:50 and 100:0 regulation scenarios.</p> <p>&nbsp;</p> <p>Simulated data from DCSsim: simulated_ChIP-seq_data.zip</p> <ul> <li>Set1: TF 50:50</li> <li>Set2: TF 100:0</li> <li>Set3: Sharp mark 50:50</li> <li>Set4: Sharp mark 100:0</li> <li>Set5: Broad mark 50:50</li> <li>Set6: Broad mark 100:0</li> </ul> <p>&nbsp;</p> <p>Sub-sampled data from DCSsub: sub-sampled_ChIP-seq_data.zip</p> <ul> <li>Set1: Cebpa-ChIP-seq 50:50</li> <li>Set2: Cebpa-ChIP-seq 100:0</li> <li>Set3: H3K27ac-ChIP-seq 50:50</li> <li>Set4: H3K27ac-ChIP-seq 100:0</li> <li>Set5: H3K36me3-ChIP-seq 50:50</li> <li>Set6: H3K36me3-ChIP-seq 100:0</li> </ul>

opencc-by-4.0May 2022View details →
zenodo44/100

Training material for ChIP-seq analysis

<p>The data provided here are part of a Galaxy tutorial that analyzes ChIP-seq data from a study published by Wu et al., 2014 (DOI:10.1101/gr.164830.113). The goal of this study was to investigate "the dynamics of occupancy and the role in gene regulation of the transcription factor Tal1, a critical regulator of hematopoiesis, at multiple stages of hematopoietic differentiation." To this end, ChIP-seq experiments were performed in multiple mouse cell types including a G1E cell line and megakaryocytes, the two cell types represented here. The dataset contains biological replicate Tal1 ChIP-seq and input control experiments (*.fastqsanger files). Because of the long processing time for the large original files, we have downsampled the original raw data files to include only reads that align to chromosome 19 and a subset of interesting genomic loci (ChIPseq_regions_of_interest_v4.bed) pulled from the Wu et al. publication. Also included is a gene annotation file (RefSeq_gene_annotations_mm10.bed) with gene names added for viewing in a genome browser.</p>

opencc-by-4.0Dec 2016View details →
zenodo44/100

S1Data: ChIP-seq Data from Ferrie et. al. "p300 Is an Obligate Integrator of Combinatorial Transcription Factors Inputs"

<p>ChIP data from Ferrie et. al. "p300 Is an Obligate Integrator of Combinatorial Transcription Factors Inputs"</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

DCSsim (simulated) and DCSsub (sub-sampled) ChIP-seq data from different chromosomes.

<p>These data are the results from three independent runs of DCSsim and DCSsub for TF, sharp and broad mark signals in 50:50 regulation scenarios for mm10 chr1, chr8, chr11, chr19 and chrX.</p> <p>Simulated data from DCSsim: simulated_ChIP-seq_data.zip</p> <p>set1: TF 50:50 chr11<br> set4: TF 50:50 chr8<br> set7: TF 50:50 chrX<br> set10: TF 50:50 chr1<br> set22: TF 50:50 chr19</p> <p>set2: Sharp mark 50:50 chr11<br> set5: Sharp mark 50:50 chr8<br> set8: Sharp mark 50:50 chrX<br> set11: Sharp mark 50:50 chr1<br> set23: Sharp mark 50:50 chr19</p> <p>set3: Broad mark 50:50 chr11<br> set6: Broad mark 50:50 chr8<br> set9: Broad mark 50:50 chrX<br> set12: Broad mark 50:50 chr1<br> set24: Broad mark 50:50 chr19</p> <p><br> Sub-sampled data from DCSsub: sub-sampled_ChIP-seq_data.zip</p> <p>Set1: C/EBPa-ChIP-seq 50:50 chr11<br> Set2: C/EBPa-ChIP-seq 50:50 chr8<br> Set3: C/EBPa-ChIP-seq 50:50 chrX<br> Set4: C/EBPa-ChIP-seq 50:50 chr1</p> <p>Set5: H3K27ac-ChIP-seq 50:50 chr11<br> Set6: H3K27ac-ChIP-seq 50:50 chr8<br> Set7: H3K27ac-ChIP-seq 50:50 chrX<br> Set8: H3K27ac-ChIP-seq 50:50 chr1</p> <p>Set9: H3K36me3-ChIP-seq 50:50 chr11<br> Set10: H3K36me3-ChIP-seq 50:50 chr8<br> Set11: H3K36me3-ChIP-seq 50:50 chrX<br> Set12: H3K36me3-ChIP-seq 50:50 chr1</p> <p><br> C/EBPa-ChIP-seq 50:50 chr19 can be found in sub-sampled_ChIP-seq_data.zip of the FRIP data set (DOI: 10.5281/zenodo.6042902 set8)<br> H3K27ac-ChIP-seq 50:50 chr19 can be found in sub-sampled_ChIP-seq_data.zip of the FRIP data set (DOI: 10.5281/zenodo.6042902 set9)<br> H3K36me3-ChIP-seq 50:50 chr19 can be found in sub-sampled_ChIP-seq_data.zip of the FRIP data set (DOI: 10.5281/zenodo.6042902 set10)</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Datasets for predicting TF binding using Virtual ChIP-seq

<p>This repository contains datasets necessary for using the Virtual ChIP-seq software.</p> <p>Virtual ChIP-seq requires the following datasets to predict transcription factor binding:</p> <ul> <li> <p>chipExpDir_AtoH_V1.0.0.tar.gz: Reference matrices of correlation between TF binding and gene expression for TFs starting with letters A-H.</p> </li> <li> <p>chipExpDir_ItoZ_V1.0.0.tar.gz: Reference matrices of correlation between TF binding and gene expression for TFs starting with letters I-Z.</p> </li> <li> <p>refTables_V1.1.0.tar.gz: PhastCons genomic conservation, FIMO PWM scores for JASPAR motifs, and ChIP-seq data of ENCODE and Cistrome database.</p> </li> <li> <p>hg38_chrsize.tsv: Length of chromosomes in hg38</p> </li> <li> <p>trainedModels_V1.0.0.tar.gz: Virtual ChIP-seq scikit-learn trained models saved in joblib format</p> </li> <li> <p>&lt;CellType&gt;.tar.gz: Pre-calculated matrices suitable for training with other algorithms or re-training with Virtual ChIP-seq.</p> </li> </ul> <p>Some predictive features of TF binding&nbsp;are the same in each cell type and are stored together for simplicity in refTables_V1.0.0.tar.gz. You can use datasets from other cell types (named&nbsp;here as&nbsp; &lt;CellType&gt;.tar.gz) for the purpose of re-training the model. The &lt;CellType&gt;.tar.gz files contain pre-calculated predictive features of transcription factor binding in 4 chromosomes (5, 10, 15, 20).</p> <p>These features include:</p> <ul> <li> <p>PhastCons genomic conservation</p> </li> <li> <p>FIMO score for sequence motifs of TF in the JASPAR database</p> </li> <li> <p>Chromatin accessibility</p> </li> <li> <p>TF binding in ENCODE + Cistrome DB datasets</p> </li> <li> <p>Virtual ChIP-seq expression score</p> </li> </ul> <p>&nbsp;</p>

opencc-zeroFeb 2018View details →
zenodo44/100

Virtual ChIP-seq predictions of binding of 36 transcription factor in Roadmap Epigenomics Project tissues

<p>This dataset contains predictions of Virtual ChIP-seq for binding of 36&nbsp;transcription factors in Roadmap Epigenomics dataset tissues with matched DNase-seq and RNA-seq data.</p> <p>Tarball contains subfolders for each of the 36&nbsp;TFs where Virtual ChIP-seq median MCC&nbsp;in validation cell types was &gt; 0.3.</p> <p>Each subfolder contains gzipped BED files. Each file is named as &lt;Tissue&gt;_&lt;Age&gt;_&lt;TF&gt;_&lt;Accession&gt;_Predictions.bed.gz. Columns correspond to Chromosome, Start, End,&nbsp;&lt;Tissue&gt;_&lt;Age&gt;_&lt;TF&gt;_&lt;Accession&gt;, Posterior probability</p> <p>You can use the posterior probabilities provided in Virchip_PosteriorCutoffs_V3.0.0.tsv. These are posterior probability cutoffs which maximized MCC in H1-hESC cell type, or are set to 0.4 if there was no ChIP-seq data of that TF in H1-hESC (0.4 is the mode of all optimal posterior probability cutoffs in H1-hESC).</p>

opencc-zeroOct 2018View details →
zenodo44/100

Collection of Schistosoma mansoni ChIP-Seq input fastq files

<p>These are fastq files of ChIP-Seq input files for different life cycle stages of <em>Schistosoma mansoni</em>.</p> <ul> <li>adult female worms</li> <li>pairs of adults</li> <li>female cercariae</li> <li>miracidia</li> <li>primary sporocysts (sp1)</li> </ul> <p>Produced at IHPE (http://ihpe.univ-perp.fr/)</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Convolutional Neural Net (CNN) models for ENCODE-Roadmap DNase-seq peaks and Transcription Factor ChIP-seq peaks - Basset architecture

<p>Deep learning models trained on epigenomic landscapes from ENCODE and Roadmap Epigenomics. The models are Basset convolutional neural networks (Kelley, et al 2016). The dataset used to train these models can be found at https://doi.org/10.5281/zenodo.4059038. The file `nn.encode-roadmap.models.basset.clf.tar.gz` contains 10 cross-validated models in Tensorflow framework files as well as details on the architecture, cross-validation scheme, and training of these models. The file `nn.encode-roadmap.models.basset.clf.np_weights.tar.gz` contains the 10 cross-validated models&#39; weights extracted to numpy array files (.npz).</p>

openmit-licenseSep 2020View details →
zenodo40/100

Training data for ChIP-seq data analysis (Galaxy Training Material): Identification of the binding sites of the Estrogen receptor

<p>The data provided here are part of a Galaxy Training Network tutorial that analyzes ChIP-seq data from a study published by Ross-Inness et al., 2012 (DOI:10.1038/nature10730) to identify the binding sites of the Estrogen receptor, a transcription factor known to be associated with different types of breast cancer.</p>

opencc-by-4.0Sep 2017View details →
zenodo40/100

The raw data of Chip-seq in ESM-DBP

<p>The raw sequencing data in fastq format of Chip-seq using GTF2E2 and SUPT6H antibodies on Hale cell.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Virtual ChIP-seq predictions for TF binding in Cistrome and ENCODE-DREAM datasets

<p>Each gzipped tarball contains BED files of Virtual ChIP-seq posterior probabilities.</p> <p>Each BED file corresponds to binding of a TF in one chromosome (one of chr5, chr10, chr15, and chr20) in a validation cell type of Cistrome (virchipCistromePredictions_V2.0.0.tar.gz) or ENCODE-DREAM Challenge datasets (virchipDreamPredictions_V1.0.0.tar.gz).</p> <p>&nbsp;</p> <p>V2.0.0 update:</p> <ul> <li>We provided our predictions on ENCODE-DREAM Challenge validation chromosomes (chr1, chr8, and chr21) in virchipDreamPredictions_V2.0.0.tar.gz</li> <li>FOXA1 predictions for binding in ch5 and chr10 of liver, and CTCF predictions for binding in chr10 of&nbsp;MCF-7&nbsp;were missing from virchipCistromePredictions_V1.0.0.tar.gz</li> </ul> <p>&nbsp;</p>

opencc-zeroMar 2018View details →
zenodo40/100

Bamfiles ChIP-seq

<p>Bamfiles resulting from mapping reads deposited in GEO (Accession GSE121283) to the genome of <em>Fusarium oxysporum</em> f. sp. <em>lycopersici </em>4287 (Fol4287).</p> <p>Adapter sequences were removed&nbsp;and quality scores were converted to Sanger format with the MAQ sol2sanger command&nbsp;if needed. Quality was checked manually using FastQC.</p> <p>Reads were aligned to the genome of Fol4287 using `bwa aln`&nbsp;(bwa version 0.7.12)<br> Duplicate reads were removed with `Picard tools` (version 1.134) (<a href="http://broadinstitute.github.io/picard,">http://broadinstitute.github.io/picard,</a>&nbsp;MarkDuplicates)</p> <p>See&nbsp;<a href="https://doi.org/10.1016/j.fgb.2015.03.006">10.1016/j.fgb.2015.03.006</a>&nbsp;for details on ChIP-seq experiment protocols. This dataset is described in&nbsp;10.1101/465070.&nbsp;&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

Sub-sampled Fastq Files for ChIP-seq datasets from Click-Seq Science Paper (Science 2017, 10.1126/science.aal2066)

<p>Data were downloaded from SRA. We then randomly&nbsp;sub-sampled&nbsp;20% of the reads using&nbsp;seqtk_sample v 1.2 in galaxy (seed 4).</p> <p>This dataset is a support dataset for the&nbsp;<a href="https://www.embl.de/training/events/2019/EPI19-01/index.html">EMBL Course: Chromatin Signatures During Differentiation: Integrated Omics</a></p>

opencc-by-4.0Aug 2019View details →
zenodo40/100

LFY ChIP-SEQ analysis Galaxy Training Material

<p>Datasets for Galaxy Training on ChIP-SEQ analysis. Raw files can be downloaded from SRA project&nbsp;SRP051214</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Nature Communications 2020 Kopp et al JunD ChIP-seq SeqDatas

<p>This record represents SeqData objects saved as Zarr files (https://github.com/ML4GLand/SeqData) derived from ENCODE consortium ChIP-seq experiments with the JunD transcription. This data was used in one of the use cases in the EUGENe publication (https://github.com/ML4GLand/EUGENe_paper), and includes objects used in various tutorials available in the ML4GLand GitHub organization.</p> <p>These files are primarily accessed via the SeqDatasets package (https://github.com/ML4GLand/SeqDatasets).</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

ChIP-seq of plasma cell-free nucleosomes identifies cell-of-origin gene expression programs

<p>Genomic DNA is packed by histone proteins that carry a multitude of post-translational modifications&nbsp; that reflect cellular transcriptional state. Cell-free DNA (cfDNA) is derived from fragmented chromatin in dying cells, and as such it retains the histones markings present in the cells of origin. Here, we pioneer chromatin immunoprecipitation followed by sequencing of cell-free nucleosomes (cfChIP-seq) carrying active chromatin marks. Our results show that cfChIP-seq provides multidimensional epigenetic information that recapitulates the epigenetic and transcriptional landscape in the cells of origin. We applied cfChIP-seq to 268 samples including samples from patients with heart and liver pathologies, and 135 samples from 56 metastatic CRC patients. We show that cfChIP-seq can detect pathology-related transcriptional changes at the site of the disease, beyond the information on tissue of origin. In CRC patients we detect clinically-relevant, and patient-specific information, including transcriptionally active HER2 amplifications. cfChIP-seq provides genome-wide information and requires low sequencing depth. Altogether, we establish cell-free chromatin immunoprecipitation as an exciting modality with potential for diagnosis and interrogation of physiological and pathological processes using a simple blood test.</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

Xrp1 ChIP-seq database (data from Xrp1 is a transcription factor required for cell competition-driven elimination of loser cells)

<p>Xrp1 ChIP-seq data from Baillon et al. 2018, "Xrp1 is a transcription factor required for cell competition-driven elimination of loser cells".</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

DCSsim (simulated) and DCSsub (sub-sampled) ChIP-seq data with different FRIP.

<p>These data are the results from three independent runs of DCSsim and DCSsub for TF, sharp and broad mark signals in 50:50 regulation scenarios for four (sim) and three (sub) different FRIP ranges.</p> <p>Simulated data from DCSsim: simulated_ChIP-seq_data.zip</p> <p>Set13: TF 50:50 x0.5 background<br> Set14: TF 50:50 x2 background<br> Set15: TF 50:50 x3 background<br> Set25: TF 50:50 x1 background</p> <p>Set16: Sharp mark 50:50 x0.5 background<br> Set17: Sharp mark 50:50 x2 background<br> Set18: Sharp mark 50:50 x3 background<br> Set26: Sharp mark 50:50 x1 background</p> <p>Set19: Broad mark 50:50 x0.5 background<br> Set20: Broad mark 50:50 x2 background<br> Set21: Broad mark 50:50 x3 background<br> Set27: Broad mark 50:50 x1 background</p> <p><br> Sub-sampled data from DCSsub: sub-sampled_ChIP-seq_data.zip</p> <p>Set1: PU1-ChIP-seq 50:50<br> Set2: STAT6-ChIP-seq 50:50<br> Set8: C/EBPa-ChIP-seq 50:50</p> <p>Set4: H3K4me3-ChIP-seq 50:50<br> Set9: H3K27ac-ChIP-seq 50:50<br> Set15: H3K9ac-ChIP-seq 50:50</p> <p>Set6: H3K27me3-ChIP-seq 50:50<br> Set10: H3K36me3-ChIP-seq 50:50<br> Set11: H3K79me2-ChIP-seq 50:50</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Transferrin gene ChIP-seq data

<p>ChIP-seq Peak-calling data to support the supplementary information to:</p> <p><strong>In severe alcoholic hepatitis, serum transferrin constitutes an independent mortality predictor indicating an impaired HNF4</strong><strong>&alpha;signaling.&nbsp;</strong>Stephen R. Atkinson, MD<sup>*</sup>; Karim Hamesch, MD<sup>*</sup>; Igor Spivak, MD; Nurdan Guldiken, PhD; Joaqu&iacute;n Cabezas, PhD<sup>3</sup>; Josepmaria Argemi, PhD; Igor Theurl, MD, Heinz Zoller, MD; Sheng Cao, MD; Philippe Mathurin, MD; Vijay H. Shah, MD; Christian Trautwein, MD; Ramon Bataller, MD; Mark R. Thursz, MD<sup>+</sup>; Pavel Strnad, MD<sup>+</sup></p>

opencc-by-4.0Jun 2019View details →
zenodo36/100

ChIP-seq data for training purposes

<p>Data for the EBAII training and concerning reads extracted from GSE40129 serie (<a href="https://www.ncbi.nlm.nih.gov/geo">NCBI GEO dataset</a>) by human chr11 mapping.</p> <p>siNT_ER_E2_r3_chr11.fastq.gz (from <a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM986063">GSM986063</a> sample) : ChIP-seq</p> <p>MCF_input_r3_chr21.fastq.gz (from&nbsp;<a href="https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSM986091">GSM986091</a> sample) : input</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record