Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
124
datasets available to search
ShareScore release 0.9.0
Dataset results
124 results for “Transcription Factor Binding Sites”
DoubleChEC program to identify transcription factor binding sites from mapped ChEC-seq data
<p>ChIP-seq (chromatin immunoprecipitation followed by sequencing) is commonly used to identify genome-wide protein-DNA interactions. However, ChIP-seq often gives a low yield, which is not ideal for quantitative outcomes. An alternative method to ChIP-seq is ChEC-seq (Chromatin endogenous cleavage with high-throughput sequencing). In this method, the endogenous TF (transcription factor) of interest is fused with MNase (micrococcal nuclease) that non-specifically cleaves DNA near binding sites. Compared to the <a href="https://www.nature.com/articles/ncomms9733" rel="nofollow">original ChEC-seq method</a>, the <a href="https://sites.northwestern.edu/bricknerlab/" rel="nofollow">modified version</a> requires far less amplification. Since <a href="https://github.com/macs3-project/MACS/tree/master#introduction">MACS3</a> failed to identify peaks in data generated from the modified ChEC-seq method, a new peak finder has been developed specifically for it.</p> <p>There are three functions in the <em><code>peak_finder/</code></em>. <code>callpeaks()</code> is used to identify peaks from BAM files. <code>goanalysis()</code> is used to make GO (Gene Ontology) term plots from peaks. <code>bedtomeme()</code> is a wrapper function to perform <a href="https://meme-suite.org/meme/tools/meme" rel="nofollow">MEME analysis</a> in R <strong>after <a href="https://meme-suite.org/meme/doc/download.html" rel="nofollow">MEME Suite</a> is installed locally</strong>.</p>
Mammalian Evolution of Human cis-regulatory Elements and Transcription Factor Binding Sites
<p>Code and data associated with the manuscript entitled "Mammalian Evolution of Human cis-regulatory Elements and Transcription Factor Binding Sites "</p>
DoubleChEC program to identify transcription factor binding sites from mapped ChEC-seq data
Open the record for dataset details and reuse information.
Quantitative modulation of a spatial enhancer through the biophysical properties of a transcription factor binding site
Open the record for dataset details and reuse information.
Pleiotropic Expression Quantitative Trait Loci Are Enriched in Enhancers and Transcription Factor Binding Sites and Impact More Genes
<p>This dataset comprises two files that accompany the article (link to be added upon publication).</p> <h2>1. gwas2eqtl_colocalization_full.tar.gz</h2> <p>This file contains the complete colocalization dataset generated using the code from the following GitHub repository: gwas2eqtl. This dataset is used as input for the pleiotropic eQTL analysis available at gwas2eqtl_pleiotropy, which produces the figures in the article.</p> <p><strong>File structure:</strong></p> <blockquote> <p>.<br>└── gwas417<br> └── coloc<br> ├── ebi-a-GCST000998<br> │ └── pval_5e-08<br> │ └── r2_0.1<br> │ └── kb_1000<br> │ └── window_1000000<br> │ ├── Alasoo_2018_ge_macrophage_IFNg+Salmonella.tsv<br> │ ├── Alasoo_2018_ge_macrophage_IFNg.tsv<br>...</p> </blockquote> <p>Each TSV file contains the following columns:</p> <blockquote> <p>chrom pos rsid ref alt eqtl_gene_id gwas_beta gwas_pval gwas_id eqtl_beta eqtl_pval eqtl_id PP.H4.abf SNP.PP.H4 nsnps PP.H3.abf PP.H2.abf PP.H1.abf PP.H0.abf coloc_variant_id coloc_region<br>1 109272258 rs4970834 C T ENSG00000168765 -0.12874.25001047052626e-09 ebi-a-GCST000998 -0.250697 0.0893351 Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.0520205502224409 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>1 109274968 rs12740374 G T ENSG00000168765 -0.103341 1.63998546891446e-09 ebi-a-GCST000998 -0.197397 0.172673Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.0857585178966856 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>1 109275216 rs660240 T C ENSG00000168765 0.1044492.78997299740827e-09 ebi-a-GCST000998 0.214165 0.139318 Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.0557749486050279 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>1 109275684 rs629301 G T ENSG00000168765 0.1054716.129993302249e-10 ebi-a-GCST000998 0.197397 0.172673 Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.22229240584331 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>1 109278889 rs602633 T G ENSG00000168765 0.1034352.15998134341707e-09 ebi-a-GCST000998 0.226329 0.102673 Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.0782482718431504 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>...</p> </blockquote> <p> </p> <p>The dataset provides colocalization statistics for GWAS-eQTL pairs, including posterior probabilities and variant annotations.</p> <h2>2. gwas2eqtl0.1.3.tsv.gz</h2> <p>This file is a filtered version of the colocalization dataset, refined based on cutoffs of PP.H4.abf ≥ 0.75 and SNP.PP.H4 ≥ 0. This subset is utilized in the gwas2eqtl web application for data visualization.</p> <p>Sample Columns:</p> <blockquote> <p>chrom pos19 pos38 cytoband rsid ref alt gwas_trait gwas_class gwas_beta eqtl_gene_symbol eqtl_beta eqtl_id eqtl_gene_id gwas_id gwas_pval eqtl_pval pp_h4_abf snp_pp_h4 tophits_variant_id nsnps<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.175816 BrainSeq_ge_brain ENSG00000235098 ebi-a-GCST003043 2.85292266979231e-10 0.000841046 0.978425254116226 6.18454060493069e-12 1_1312114_T_C 3<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.175816 BrainSeq_ge_brain ENSG00000235098 ieu-a-294 2.85292266979231e-10 0.000841046 0.974019788384412 7.52286530905187e-12 1_1312114_T_C 4<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.293529 CommonMind_ge_DLPFC_naive ENSG00000235098 ebi-a-GCST003043 2.85292266979231e-10 2.30452e-06 0.953333690803618 6.9758581004380506e-15 1_1312114_T_C 6<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.293529 CommonMind_ge_DLPFC_naive ENSG00000235098 ieu-a-294 2.85292266979231e-10 2.30452e-06 0.951092499048109 7.0876793540547e-15 1_1312114_T_C 7<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.510549 FUSION_ge_adipose_naive ENSG00000235098 ebi-a-GCST003043 2.85292266979231e-10 1.2212e-06 0.974412793352836 1.70246427912903e-11 1_1312114_T_C 6</p> </blockquote> <p> </p> <p>This filtered dataset focuses on high-confidence colocalization events for functional exploration of genetic associations and regulatory mechanisms.</p>
Data from: Discovery and information-theoretic characterization of transcription factor binding sites that act cooperatively
Transcription factor binding to the surface of DNA regulatory regions is one of the primary causes of regulating gene expression levels. A probabilistic approach to model protein–DNA interactions at the sequence level is through position weight matrices (PWMs) that estimate the joint probability of a DNA binding site sequence by assuming positional independence within the DNA sequence. Here we construct conditional PWMs that depend on the motif signatures in the flanking DNA sequence, by conditioning known binding site loci on the presence or absence of additional binding sites in the flanking sequence of each siteʼs locus. Pooling known sites with similar flanking sequence patterns allows for the estimation of the conditional distribution function over the binding site sequences. We apply our model to the Dorsal transcription factor binding sites active in patterning the Dorsal–Ventral axis of Drosophila development. We find that those binding sites that cooperate with nearby Twist sites on average contain about 0.5 bits of information about the presence of Twist transcription factor binding sites in the flanking sequence. We also find that Dorsal binding site detectors conditioned on flanking sequence information make better predictions about what is a Dorsal site relative to background DNA than detection without information about flanking sequence features.
The developmental and evolutionary characteristics of transcription factor binding site clustered regions based on an explainable machine learning model
<p>## Identification of transcription factor binding sites clustered regions</p> <p>First, the TFBSs were identified from ATAC-seq peaks by FIMO. The position-specific weight matrices (PWMs) of transcription factors were downloaded from CIS-BP databases. The genomic sequences under the open chromatin regions were used as inputs for FIMO with a custom library of all motifs for each species to scan for motif instances at a p-value threshold of 1e-5. </p> <p>Then, an established method was used to identify TFCRs by performing the Gaussian kernel density estimations across the genome (with a bandwidth of 300bp centered on each TFBS). Each peak in density profile was considered a TFCR. To determine the complexity of each TFCR, the Gaussian kernelized distances from each peak that contributed at least 0.1 to its strength were determined. The complexity of each TFCR was determined by the quantity and proximity of the contributing TFBS. We combined motif instances based on the TF family information from CIS-BP to calculate the complexity of TFCR. The window for each TFCR was determined by finding the maximum distance (in bp) from the TFCR to a contributing TF and then adding 150 bp (one-half of the bandwidth). Each window was centered on the TFCR. The identified TFCR was grouped into 10 groups based on their complexity from low to high. </p> <p>usage: <br>indir="Human_fimo" # the directory where you put the output files of FIMO <br>motifMap="Homo_sapiens_2020_0920/TF_Information_all_motifs_plus.txt" # the mapping relationship of TF and its TF family from CIS-BP <br>cd Codes/TFCR_embryo <br>perl d-motif_combine.pl $indir TFfamily $motifMap <br>perl e-tfpos_combine.pl TFfamily <br>perl f1-tf_bed-new-c.pl TFfamily <br>perl 0-merge-TFCR.pl $indir TFfamily </p>
Data from: Discovery and information-theoretic characterization of transcription factor binding sites that act cooperatively
Open the record for dataset details and reuse information.
DNA features beyond the transcription factor binding sites specify target recognition by plant bHLHs
GEO Series GSE155321. Marchantia polymorpha; Arabidopsis thaliana; Solanum lycopersicum. 35 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Genome binding/occupancy profiling by array.
Identification of transcription factor MAB-5 binding sites
GEO Series GSE15625. Caenorhabditis elegans. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Identification of binding sites of the Six1 transcription factor in mouse primary myoblasts and myotubes
GEO Series GSE175999. Mus musculus. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Quantifying Transcription Factor Specificity with Advanced DNA Universal Microarrays Featuring Long and Modified Binding Sites [48k]
GEO Series GSE299470. Homo sapiens; synthetic construct. 2 samples. Type: Other.
Genome-wide maps of Egr2 transcription factor binding sites in NKT and anti TCRb injected thymocytes.
GEO Series GSE34254. Mus musculus. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
DNA features beyond the transcription factor binding sites specify target recognition by plant bHLHs [GenomeData]
GEO Series GSE155320. Arabidopsis thaliana. 4 samples. Type: Genome binding/occupancy profiling by array.
Transcription factor binding in human cells occurs in dense clusters formed around cohesin anchor sites
GEO Series GSE49402. Homo sapiens. 225 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Transcription Factor Binding Sites by Motifs Scan
GEO Series GSE53962. Homo sapiens. 0 samples. Type: Third-party reanalysis; Genome binding/occupancy profiling by high throughput sequencing.
Distinct properties of cell type-specific and shared transcription factor binding sites
GEO Series GSE49993. Mus musculus. 4 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Genome-wide mapping of the Arabidopsis thaliana heat shock transcription factor A1b binding sites under non-stress and heat stress conditions [ChIP-seq]
GEO Series GSE85651. Arabidopsis thaliana. 16 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
Transcription Factor Binding Sites clustered Regions of 133 Cell Lines
GEO Series GSE59016. Homo sapiens. 0 samples. Type: Third-party reanalysis; Genome binding/occupancy profiling by high throughput sequencing.
ENCODE Transcription Factor Binding Sites by ChIP-seq from Stanford/Yale/USC/Harvard
GEO Series GSE31477. Homo sapiens. 426 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.