Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,276
datasets available to search
ShareScore release 0.9.0
Dataset results
4,276 results for “transcription factors”
An activity-specificity trade-off encoded in human transcription factors
<p>Data repository for the publication <strong>"An activity-specificity trade-off encoded in human transcription factors"</strong>. </p> <p>Imaging datasets with larger sizes can be downloaded here: https://owww.molgen.mpg.de/~TFsuboptimization/</p>
Monosaccharide transporter OsMST6 is activated by transcription factor OsERF120 to enhance chilling tolerance in rice seedlings
<p>Chilling stress caused by extreme weather is threatening global rice (<em>Oryza sativa</em> L.) production. Identifying components of the signal transduction pathways underlying chilling tolerance in rice would advance molecular breeding. Here, we report that <em>OsMST6</em>, which<em> </em>encodes<em> </em>a monosaccharide transporter, positively regulates the chilling tolerance of rice seedlings. The <em>mst6</em> mutants showed hypersensitivity to chilling, while the <em>OsMST6</em> overexpression lines were tolerant. During chilling stress, OsMST6 transported more glucose into cells to modulate sugar and ABA signal pathways. We showed that the transcription factor OsERF120 could bind to the DRE/CRT element of the <em>OsMST6</em> promoter and activate the expression of <em>OsMST6</em> to positively regulate chilling tolerance. Genetically, OsERF120 was functionally dependent on OsMST6 when promoting chilling tolerance. In summary, OsERF120 and OsMST6 form a new downstream chilling regulatory pathway in rice in response to chilling stress, providing valuable findings for molecular breeding aimed at achieving global food security.</p> <p><span>The stored data is the analysis results of transcriptome data from ZH11 and <em>mst6-1</em> after 0 hours and 4 hours of chilling treatment. These data form the basis for the transcriptome data visualization in the article.</span></p>
Identification and expression analysis of transcription factors in the Carallia brachiata genome
<p>Rhizophoraceae has 2 terrestrial genera and 4 marine genera. The intertidal zone in which marine mangroves are located is known for its low oxygen and high salinity. Marine and terrestrial genera have evolved distinct adaptive characteristics, among which viviparous reproduction is the most unique. To investigate the genetic foundations difference underlying the adaptive mechanisms of marine–terrestrial genera, we selected two species from Rhizophoraceae. <em>Kandelia obovata</em> is marine and viviparous, and <em>Carallia brachiata</em> is terrestrial and non-viviparous.<em> </em>We compared their transcriptome of 8 tissues (root, stem, leaf, flower, ovule, fruit, seed, embryo) and found that the mature reproductive organs (fruit, seed, embryo) of <em>K. obovata </em>did not reduce metabolic activity compared to <em>C. brachiata</em>. The reproductive organs of <em>K. obovata</em> were regulated by the same gene set as vegetative organs. This contrasted with <em>C. brachiata</em>. Eight kinds of hormone transduction genes were up-regulated in the seed of <em>K. obovata</em>. Finally, and most importantly, the transcriptional factors AP2 and ARF families were significantly more expressed in the reproductive organs of <em>K. obovata</em> than in those of <em>C. brachiata</em>. At the same time, the ERF family was more expressed in its roots. The findings suggested that the hormone transduction may contribute to viviparous initiation. Transcriptional factors were quite crucial for mangroves' adaptation to wetlands.</p>
Pleiotropic Expression Quantitative Trait Loci Are Enriched in Enhancers and Transcription Factor Binding Sites and Impact More Genes
<p>This dataset comprises two files that accompany the article (link to be added upon publication).</p> <h2>1. gwas2eqtl_colocalization_full.tar.gz</h2> <p>This file contains the complete colocalization dataset generated using the code from the following GitHub repository: gwas2eqtl. This dataset is used as input for the pleiotropic eQTL analysis available at gwas2eqtl_pleiotropy, which produces the figures in the article.</p> <p><strong>File structure:</strong></p> <blockquote> <p>.<br>└── gwas417<br> └── coloc<br> ├── ebi-a-GCST000998<br> │ └── pval_5e-08<br> │ └── r2_0.1<br> │ └── kb_1000<br> │ └── window_1000000<br> │ ├── Alasoo_2018_ge_macrophage_IFNg+Salmonella.tsv<br> │ ├── Alasoo_2018_ge_macrophage_IFNg.tsv<br>...</p> </blockquote> <p>Each TSV file contains the following columns:</p> <blockquote> <p>chrom pos rsid ref alt eqtl_gene_id gwas_beta gwas_pval gwas_id eqtl_beta eqtl_pval eqtl_id PP.H4.abf SNP.PP.H4 nsnps PP.H3.abf PP.H2.abf PP.H1.abf PP.H0.abf coloc_variant_id coloc_region<br>1 109272258 rs4970834 C T ENSG00000168765 -0.12874.25001047052626e-09 ebi-a-GCST000998 -0.250697 0.0893351 Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.0520205502224409 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>1 109274968 rs12740374 G T ENSG00000168765 -0.103341 1.63998546891446e-09 ebi-a-GCST000998 -0.197397 0.172673Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.0857585178966856 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>1 109275216 rs660240 T C ENSG00000168765 0.1044492.78997299740827e-09 ebi-a-GCST000998 0.214165 0.139318 Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.0557749486050279 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>1 109275684 rs629301 G T ENSG00000168765 0.1054716.129993302249e-10 ebi-a-GCST000998 0.197397 0.172673 Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.22229240584331 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>1 109278889 rs602633 T G ENSG00000168765 0.1034352.15998134341707e-09 ebi-a-GCST000998 0.226329 0.102673 Alasoo_2018_ge_macrophage_IFNg+Salmonella 0.108283426725895 0.0782482718431504 6 0.000552728832012655 2.09397277427793e-07 0.890881419183881 0.000282215860934432 1_109279544_G_A 1:108779544-109779543<br>...</p> </blockquote> <p> </p> <p>The dataset provides colocalization statistics for GWAS-eQTL pairs, including posterior probabilities and variant annotations.</p> <h2>2. gwas2eqtl0.1.3.tsv.gz</h2> <p>This file is a filtered version of the colocalization dataset, refined based on cutoffs of PP.H4.abf ≥ 0.75 and SNP.PP.H4 ≥ 0. This subset is utilized in the gwas2eqtl web application for data visualization.</p> <p>Sample Columns:</p> <blockquote> <p>chrom pos19 pos38 cytoband rsid ref alt gwas_trait gwas_class gwas_beta eqtl_gene_symbol eqtl_beta eqtl_id eqtl_gene_id gwas_id gwas_pval eqtl_pval pp_h4_abf snp_pp_h4 tophits_variant_id nsnps<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.175816 BrainSeq_ge_brain ENSG00000235098 ebi-a-GCST003043 2.85292266979231e-10 0.000841046 0.978425254116226 6.18454060493069e-12 1_1312114_T_C 3<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.175816 BrainSeq_ge_brain ENSG00000235098 ieu-a-294 2.85292266979231e-10 0.000841046 0.974019788384412 7.52286530905187e-12 1_1312114_T_C 4<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.293529 CommonMind_ge_DLPFC_naive ENSG00000235098 ebi-a-GCST003043 2.85292266979231e-10 2.30452e-06 0.953333690803618 6.9758581004380506e-15 1_1312114_T_C 6<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.293529 CommonMind_ge_DLPFC_naive ENSG00000235098 ieu-a-294 2.85292266979231e-10 2.30452e-06 0.951092499048109 7.0876793540547e-15 1_1312114_T_C 7<br>1 1163804 1228424 1p36.33 rs7515488 C T Inflammatory bowel disease Autoimmune dis. 0.0874308 ANKRD65 -0.510549 FUSION_ge_adipose_naive ENSG00000235098 ebi-a-GCST003043 2.85292266979231e-10 1.2212e-06 0.974412793352836 1.70246427912903e-11 1_1312114_T_C 6</p> </blockquote> <p> </p> <p>This filtered dataset focuses on high-confidence colocalization events for functional exploration of genetic associations and regulatory mechanisms.</p>
Transfer learning reveals sequence determinants of the quantitative response to transcription factor dosage
<p>Processed data and code for "Transfer learning reveals sequence determinants of the quantitative response to transcription factor dosage," Naqvi et al 2025.</p> <p>Directory is organized into the following subfolders, each tar'ed and gzipped:</p> <p><strong>data_analysis.tar.gz - Processed data for modulation of TWIST1 levels and calculation of RE responsiveness to TWIST1 dosage</strong></p> <ul> <li>atac_design.txt - design matrix for ATAC-seq TWIST1 titration samples</li> <li>all.sub.150bpclust.greater2.500bp.merge.TWIST1.titr.ATAC.counts.txt - ATAC-seq counts from all samples over all reproducible ATAC-seq peak regions, as defined in Naqvi et al 2023</li> <li>atac_deseq_fitmodels_moded50.R - R code for calculating new version of ED50 and response to full depletion from TWIST1 titration data (note, uses drm.R function from <a href="https://doi.org/10.5281/zenodo.7689948">10.5281/zenodo.7689948</a>, install drc() with this version to avoid errors)</li> </ul> <p><strong>baseline_models.tar.gz - Code and data for training baseline models to predict RE responsiveness to SOX9/TWIST1 dosage</strong></p> <ul> <li>{sox9|twist1}.{0v100|ed50}.{train|valid|test}.txt - Training/testing/validation data (ED50 or full TF depletion effect for SOX9 or TWIST1), split into train/test/validation folds</li> <li>HOCOMOCOv11_core_HUMAN_mono_jaspar_format.all.sub.150bpclust.greater2.500bp.merge.minus300bp.p01.maxscore.mat.cpg.gc.basemean.txt.gz - matrix of predictors for all REs. Quantitative encoding of PWM match for all HOCOMOCO motifs + CpG + GC content, plus unperturbed ATAC-seq signal</li> <li>train_baseline.R - R code to train baseline (LASSO regression or random forest) models using predictor matrix and the provided training data. <ul> <li>Note: training the random forest to predict full TF depletion is computationally intensive because it is across all REs, if doing this run on CPU for ~6 hrs. </li> </ul> </li> </ul> <p><strong>chrombpnet_models.tar.gz - Remainder of code, data, and models for fine-tuning and interpreting ChromBPNet mdoels to </strong><strong>predict RE responsiveness to SOX9/TWIST1 dosage</strong></p> <ul> <li>Fine-tuning code, data, models <ul> <li>{all|sox9.direct|twist1.bound.down}.{train|valid|test}.{ed50|0v100.log2fc}.txt - Training/testing/validation data (ED50 or full TF depletion effect for SOX9 or TWIST1), split into train/test/validation folds</li> <li>pretrained.unperturbed.chrombpnet.h5 - Pretrained model of unperturbed ATAC-seq signal in CNCCs, obtained by running ChromBPNet (https://github.com/kundajelab/chrombpnet) on DMSO-treated SOX9/TWIST1-tagged ATAC-seq data</li> <li>finetune_chrombpnet.py - code for fine-tuning the pretrained model for any of the relevant prediction tasks (ED50/ effect of full TF depletion for SOX9/TWIST1)</li> <li>best.model.chrombpnet.{0v100|ed50}.{sox9|twist1}.h5 - output of finetune_chrombpnet.py, best model after 10 training epochs for the indicated task</li> <li>chrombpnet.{0v100|ed50}.{sox9|twist1}.contrib.{h5|bw} - contribution scores for the indicated predictive model, obtained by running chrombpnet contribs_bw on the corresponding model h5 file.</li> <li>chrombpnet.{0v100|ed50}.{sox9|twist1}.contrib.modisco.{h5|bw} - TF-MoDIsCo output from the corresponding contribution score file</li> </ul> </li> <li>Interpretation code, data, models <ul> <li>contrib_h5_to_projshap_npy.py - code to convert contrib .h5 files into .npy files containing projected SHAP scores (required because the CWM matching code takes this format of contribution scores)</li> <li>sox9.direct.10col.bed, twist1.bound.down.10col.uniq.bed - regions over which CWMs will be matched (likely direct targets of each TF)</li> <li>match_cwms.py - Python code to match individual CWM instances. Takes as input: modisco .h5 file, SHAP .npy file, bed file of regions to be matched. Output is a bed file of all CWM matches (not pruned, contains many redundant matches).</li> <li>chrombpnet.ed50.{sox9|twist1}.contrib.perc05.matchperc10.allmatch.bed - output of match_cwms.py </li> <li>take_max_overlap.py - code to merge output of match_cwms.py into clusters, and then take the maximum (length-normalized) match score in each cluster as the representative CWM match of that cluster. Requires upstream bedtools commands to be piped in, see example usage in file. </li> <li>chrombpnet.ed50.{sox9|twist1}.contrib.perc05.matchperc10.allmatch.maxoverlap.bed - output of take_max_overlap.py. These CWM instances are the ones used throughout the paper.</li> </ul> </li> </ul> <p><strong>modisco_reports.zip -</strong><strong> TF-MoDIsCo reports from running on the fine-tuned ChromBPNet models</strong></p> <ul> <li>modisco_report_{sox9|twist1}_{0v100|ed50}: folders containing images of discovered CWMs and HTMLs/PDFs of summarized reports from running TF-MoDisCo on the indicated fine-tuned ChromBPNet model</li> </ul> <p><strong>chrombpnet_models_supp.tar.gz - Alternative ChromBPNet mdoels to </strong><strong>predict SOX9/TWIST1 ED50 using varying definitons of direct targets</strong></p> <ul> <li> <p>best.model.chrombpnet.ed50.twist1.3hdn.h5 - TWIST1 direct targets defined using response to full 3h depletion (as was done for SOX9 throughout the rest of the paper)</p> </li> <li> <p>best.model.chrombpnet.ed50.sox9.v5chip.h5 - SOX9 direct targets defined using V5 ChIP-seq from SOX9-tagged lines (as was done for TWIST1 throughout the paper)</p> </li> </ul> <p><strong>mirny_model.tar.gz - Code and data for analyzing and fitting Mirny model of TF-nucleosome competition to observed RE dosage response curves</strong></p> <ul> <li>twist1.strong.multi.only.ed50.cutoff.true.hill.txt - ED50 and signed hill coefficients for all TWIST1-dependent REs with only buffering Coordinators (mostly one or two) and no other TFs' binding sites. "ed50_new" is the ED50 calculation used in this paper. </li> <li>twist1.strong.weak{1|2|3}.ed50.cutoff.true.hill.txt - ED50 and signed hill coefficients for all TWIST1-dependent REs with only buffering Coordinators (mostly one or two) and the indicated number of sensitizing (weak) Coordinators and no other TFs' binding sites. "ed50_new" is the ED50 calculation used in this paper. </li> <li>MirnyModelAnalysis.py - Python code for analysis of Mirny model of TF-nucleosome competition. Contains implementations of analytic solutions, as well as code to fit model to observed ED50 and hill coefficients in the provided data files.</li> </ul> <p><strong>nucleoatac.tar.gz - Output files from running NucleoATAC on merged ATAC-seq from each of 5 TWIST1 dosages</strong></p> <ul> <li>TWIST1_{dosage}_merge.nucmap_combined.bed.gz - see NucleoATAC docs for output format</li> </ul>
Distinct and essential roles of bZIP transcription factors in stress response and pathogenesis in Alternaria alternata
<p>The ability to cope with environmental abiotic stress and biotic stress is crucial for the survival of plants and microorganisms, which enable them to occupy multiple niches in the environment. Previous studies have shown that transcription factors play crucial roles in regulating various biological processes including multiple stress tolerance and response in eukaryotes. This work identified multiple critical transcription factor genes, metabolic pathways and gene ontology (GO) terms related to abiotic stress response were broadly activated by analyzing the transcriptome of phytopathogenic fungus Alternaria alternata un- der metal ions stresses, oxidative stress, salt stresses, and host-pathogen interaction. We determined the biological functions and regulatory roles of the bZIP transcriptional factor (TF) genes in the phytopathogenic fungus A. alternata by analyzing targeted gene deletion mutants. Morphological analysis provides evidence that bZIPs including Gcn4, MeaB, Atf1, Hac1 and Ada1 are required for morphogenesis as the colony morphology of these gene deletion mutants was significantly different from that of the wild-type. In addition, bZIPs are involved in the resistance to multiple stresses such as oxidative stress (Ada1, Yap1, MetR) and virulence (Hac1, MetR, Yap1, Ada1) at varying degrees. Transcriptome data demonstrated that the inactivation of bZIPs (Hac1, Atf1, Ada1 and Yap1) significantly affected many genes in multiple critical metabolism pathways and gene ontology (GO) terms. Moreover, the ΔHac1 mutants displayed reduced aerial hypha and are hypersensitivity to endoplasmic reticulum disruptors such as tunicamycin and dithiothreitol. Transcriptome analysis showed that inactivation of Hac1 significantly affected the proteasome process and its downstream unfolded protein binding, indicating that Hac1 participates in the endoplasmic reticulum stress response through the conserved unfolded protein response. Taken together, our findings identified many crucial transcription factor genes and pathways related to cell development, abiotic stress response and pathogenesis, and expand our understanding of how microbial pathogens utilize these genes to deal with environmental stresses and achieve successful infection in the host plant.</p>
MeDeMo - a dependency model for DNA methylation-aware transcription factor binding predictions
<p>The uploaded <em>fasta </em>files contain extended reference genomes for three cell lines HepG2, GM12878, K562 (ENCODE) and two primary liver hepatocyte samples from the german epigenomics consortium (DEEP). The extended reference genomes contain information on DNA methylation in a CpG context. They can be used as input for <em>MeDeMo</em>, a tool to infer transcription factor binding sites incorporating not only sequence specificity but also DNA methylation. <em>MeDeMo </em>is available online at: <a href="http://www.jstacs.de/index.php/MeDeMo">http://www.jstacs.de/index.php/MeDeMo</a>.</p> <p>We considered the files ENCFF279HCL and ENCFF835NTC for GM12878, ENCFF867JRG and ENCFF721JMB for K562 as well as ENCFF064GJQ and ENCFF369YQW for HepG2. From DEEP, we considered samples 41_Hf01 and 41_Hf03 which are available through the International Human Epigenomics Consortium (IHEC).</p> <p>In addition, we provide all models trained using the mentioned data sets as well models for and motifs from genome wide predictions.</p>
The transcription factor network of E. coli steers global responses to shifts in RNAP concentration
<p><span>The robustness and sensitivity of gene networks to environmental changes</span><span> is critical for cell survival. How gene networks produce specific, chronologically ordered responses to genome-wide perturbations, while robustly maintaining homeostasis, remains an open question. We analysed if short- and mid-term genome-wide responses to shifts in RNA polymerase (RNAP) concentration are influenced by the <em>known</em> topology and logic of the transcription factor network (TFN) of <em>Escherichia coli</em>. We found that, at the gene cohort level, the magnitude of the single-gene, mid-term transcriptional responses to changes in RNAP concentration can be explained by the absolute difference between the gene's numbers of activating and repressing input transcription factors (TFs)</span><span>. Interestingly, this difference is strongly positively correlated with the number of input TFs of the gene. Meanwhile, short-term responses showed only weak influence from the TFN. </span><span>Our results suggest that the global topological traits of the TFN of <em>E. coli</em> shape which gene cohorts respond to genome-wide stresses.</span></p>
A dual selection system for directed evolution to identify allosteric transcription factor PobR variants responsive to different aromatic compounds
<p>This dataset includes all the raw data of our characterization experiments during the work titled “A dual selection system for directed evolution to identify allosteric transcription factor PobR variants responsive to different aromatic compounds”.</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Figure5e_2
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019 <a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p>for Figure 5e (second part)</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Figure5f_1
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019 <a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p>for Figure 5f (first part)</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Figure5f_2
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019 <a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p>for Figure 5f (second part)</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_FigureS6
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019 <a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p>for FigureS6</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Figure5e_1
<p>Raw microscopy movies for smFRET experiments with various chromatin templates for Mivelaz M., et al, 2019</p> <p><a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a></p> <p> </p> <p>for Figure 5e (first part)</p> <p>see attached documentation for more details</p>
Chromatin fiber invasion and nucleosome displacement by the Rap1 transcription factor_Fig.2,3,S3,S4,S7
<p>Raw microscopy movies for colocalization TIRF experiments (Rap1 binding) with various chromatin templates for Mivelaz M., et al, 2019 (<a href="https://doi.org/10.1016/j.molcel.2019.10.025">https://doi.org/10.1016/j.molcel.2019.10.025</a>)</p> <p>for Figures Fig.2,3,S3,S4,S7</p> <p>see attached documentation for more details</p>
Bigwig files for paper "STK19 is a transcription-coupled repair factor that participates in UVSSA ubiquitination and TFIIH loading"
<p>Bigwig files for paper "STK19 is a transcription-coupled repair factor that participates in UVSSA ubiquitination and TFIIH loading". </p>
GRAS Family Transcription Factor Binding Behaviors in Sorghum bicolor, Oryza, and Maize
<p>Supplemental Data Files for the manuscript entitled "GRAS Family Transcription Factor Binding Behaviors in Sorghum bicolor, Oyrza, and Maize".</p>
Specifying cellular context of transcription factor regulons for exploring context-specific gene regulation programs
<p>This repository contains the raw and processed files used in Minaeva et al. 2024.</p> <p>In this version, we have revised the regulon construction pipeline and expanded the dataset to cover 40 common cell lines.</p> <p>The code used to generate these files is available at <a href="https://github.com/LappalainenLab/chip_seq_regulons" target="_new" rel="noreferrer">GitHub - LappalainenLab/chip_seq_regulons</a>.</p> <p>The descriptions of the files contained within each subdirectory are as follows:</p> <h3>1-dataset_stats</h3> <ul> <li><code>per_gene_stats_{approach}_{cell_line}.tsv</code>: Number of TFs regulating a gene according to the respective approach (S2Mb, M2Kb, or S2Kb) in a given cell line.</li> <li><code>per_tf_stats_{approach}_{cell_line}.tsv</code>: Number of target genes regulated by a TF according to the respective approach (S2Mb, M2Kb, or S2Kb) in a given cell line.</li> </ul> <h3>1-network_enrichment</h3> <ul> <li><code>enrich_scores_remap_all_tfs_K562.tsv</code>: Results of fitting logistic regression for testing the enrichment of the K562 regulon in other biological networks (PPI, coexpression, experimental trans-networks).</li> </ul> <h3>2-plot_decoupler_comparison_benchmark_across_cells</h3> <ul> <li><code>{cell_line}_comparison_benchmark.tsv</code>: Results of benchmarking S2Mb, M2Kb, CollecTri, Dorothea, ChIP-Atlas, RegNet, and TRRUST regulons using the decoupler package and the KnockTF database. Cell lines considered are K562, HepG2, and MCF7 (see Methods for benchmarking pipeline details).</li> </ul> <h3>2-plot_decoupler_filter_benchmark_across_methods</h3> <ul> <li><code>{cell_line}_filtering_benchmark.tsv</code>: Results of benchmarking S2Mb, M2Kb, and S2Kb regulons with different filters applied using the decoupler package and the KnockTF database. Cell lines considered are K562, HepG2, and MCF7 (see Methods for benchmarking pipeline details).</li> </ul> <h3>3-tf_activity</h3> <ul> <li><code>aml_k562_activity_{regulon}_sc.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy hematopoietic stem cells (HSCs) and abnormal AML progenitor cells following the decoupler pipeline. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_activity_estimates_hsc_sc.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy HSCs and abnormal AML progenitor cells across regulons.</li> <li><code>aml_dhsc_ahsc_activity_{regulon}_sc.tsv</code>: Results of TF activity analysis based on a respective regulon between leukemic activated and dormant HSCs following the decoupler pipeline. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_activity_estimates_dhsc_ahsc_sc.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between leukemic activated and dormant HSCs across regulons.</li> <li><code>bc_bas_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer following the decoupler pipeline. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_activity_estimates_bas.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer across regulons.</li> <li><code>bc_lum_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer following the decoupler pipeline. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_activity_estimates_lum.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer across regulons.</li> <li><code>hep_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between neoplastic and healthy liver cells following the decoupler pipeline. Regulons considered are HepG2-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>hep_activity_estimates.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between neoplastic and healthy liver cells across regulons.</li> </ul> <h3>3-tf_disease_enrichment</h3> <ul> <li><code>aml_{database}_enrich_{regulon}_hsc_sc.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy HSCs and abnormal AML progenitor cells following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_{database}_enrich_{regulon}_dhsc_ahsc_sc.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between leukemic activated and dormant HSCs following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_{database}_enrich_{regulon}_bas.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, and OMIM. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_{database}_enrich_{regulon}_lum.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, and OMIM. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>hep_{database}_enrich_{regulon}.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between neoplastic and healthy liver cells following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are HepG2-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> </ul> <h3>regulons</h3> <ul> <li><code>{cell_line}_regulon.tsv</code>: S2Mb, M2Kb, and S2Kb regulons generated in this study with all acquired annotations (see Methods for details).</li> </ul> <p>External regulons used for comparison. Cell lines considered are K562, HepG2, MCF7, and GM12878:</p> <ul> <li><code>ChIP-Atlas_target_genes_{cell_line}.tsv</code>: Customized ChIP-Atlas regulons (see Methods for details).</li> <li><code>Revised_Supplemental_Table_S3_Normal.csv</code>: Dorothea regulon collected from supplementary materials of Garcia-Alonso et al. (2019).</li> </ul> <h3>s3-network_enrichment</h3> <ul> <li><code>enrich_scores_remap_all_tfs_{cell_line}.tsv</code>: Results of fitting logistic regression for testing the enrichment of cell-line-specific regulons in PPI networks (see Methods and corresponding GitHub repository for details). Cell lines considered are K562, HepG2, MCF7, and GM12878.</li> </ul> <p> </p>
Expression and Purification of DNA-binding domain of T-box Transcription Factor TBXTA-c021
<p>A detailed protocol for the expression and purification of G177D variant of TBXT DNA-binding domain.</p>
Expression and Purification of Full-Length T-box Transcription Factor TBXTA-c027
<p>A detailed protocol for the expression and purification of G177D variant of biotinylated, full-length TBXT</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.