Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,291
datasets available to search
ShareScore release 0.9.0
Dataset results
3,291 results for “transcription regulation”
Supplementary Data to "Disparate regulation of Smad3 phosphorylation and collagen gene transcription by full-length IL-33"
<p>These are Supplementary Figures for the article "Disparate regulation of Smad3 phosphorylation and collagen transcription by full-length IL-33"</p>
Data and code for the publication "DNA methylation underpins the epigenomic landscape regulating genome transcription in Arabidopsis"
<p>The zipped file of this repository contains code and data to reproduce the results of the publication:</p> <p>Zhao et al, DNA methylation underpins the epigenomic landscape regulating genome transcription in Arabidopsis. Genome Biology (2022). </p> <p>All sequence data have been deposited in NCBI GEO accession codes GSE183987 and GSE169497.</p> <p> </p> <p>Please see the README document for detailed:</p> <p>- Descriptions of the code and data provided</p> <p>- Lists of the required dependencies</p>
Direct molecular evidence for an ancient, conserved developmental toolkit controlling post-transcriptional gene regulation in land plants
<p>In plants, miRNA production is orchestrated by a suite of proteins that control transcription of the pri-miRNA gene, post-transcriptional processing and nuclear export of the mature miRNA. Post-transcriptional processing of miRNAs is controlled by a pair of physically-interacting proteins, HYL1 and DCL1. However, the evolutionary history and structural basis of the HYL1-DCL1 interaction is unknown. Here we use ancestral sequence reconstruction and functional characterization of ancestral HYL1 <em>in vitro</em> and in <em>Arabidopsis thaliana </em>to better understand the origin and evolution of the HYL1-DCL1 interaction and its impact on miRNA production and plant development. We found the ancestral plant HYL1 evolved high affinity for both double-stranded RNA (dsRNA) and its DCL1 partner before the divergence of mosses from seed plants (~500 Ma), and these high-affinity interactions remained largely conserved throughout plant evolutionary history. Structural modeling and molecular binding experiments suggest that the second of two double-stranded RNA-binding motifs (DSRMs) in HYL1 may interact tightly with the first of two C-terminal DCL1 DSRMs to mediate the HYL1-DCL1 physical interaction necessary for efficient miRNA production. Transgenic expression of the nearly 200 Ma-old ancestral flowering-plant HYL1 in <em>A. thaliana</em> was sufficient to rescue many key aspects of plant development disrupted by HYL1<sup>-</sup> knockout and restored near-native miRNA production, suggesting that the functional partnership of HYL1-DCL1 originated very early in and was strongly conserved throughout the evolutionary history of terrestrial plants. Overall, our results are consistent with a model in which miRNA-based gene regulation evolved as part of a conserved plant ‘developmental toolkit’.</p>
Diel-regulated transcriptional cascades of microbial eukaryotes in the North Pacific Subtropical Gyre
<p>Trinity <em>de novo </em>assemblies of 24 poly-A+ selected, combined-replicate metatranscriptomes from HOE-Legacy 2 cruise KM1513 (Jul 24 - Aug 6, 2015). KM1513 cruise information, plots, and associated environmental data for the HOE Legacy II cruise can be found online at <a href="http://hahana.soest.hawaii.edu/hoelegacy/hoelegacy.html">http://hahana.soest.hawaii.edu/hoelegacy/hoelegacy.html</a>. Raw metatranscriptome short-read sequence data is available in the NCBI Sequence Read Archive under BioProject ID PRJNA492142. Code associated with this project is available on Github (<a href="https://github.com/armbrustlab/diel_eukaryotes">https://github.com/armbrustlab/diel_eukaryotes</a>).</p> <p> </p> <p> </p>
Analysis accompanying "Dynamically regulated transcription factors are encoded by highly unstable mRNAs in the Drosophila larval brain"
<p>This repository documents the raw data processing and figure generation for the article “Dynamically regulated transcription factors are encoded by highly unstable mRNAs in the <em>Drosophila </em>larval brain”, doi: 10.1261/rna.079552.122.</p>
Transcriptional regulation underlying the temperature response of embryonic development rate in the winter moth
<p>Climate change will strongly affect the developmental timing of insects, as their development rate largely depends on ambient temperature. However, we know little about the genetic mechanisms underlying the temperature sensitivity of embryonic development in insects. We investigated embryonic development rate in the winter moth (<em>Operophtera brumata</em>), a species with egg dormancy that has been under selection due to climate change. We used RNAseq to investigate which genes are involved in the regulation of winter moth embryonic development rate in response to temperature. Over the course of development, we sampled eggs before and after an experimental change in ambient temperature, including two early development weeks when the temperature sensitivity of eggs is low and two late development weeks when temperature sensitivity is high. We found temperature-responsive genes that responded in a similar way across development, as well as genes with a temperature response specific to a particular development week. Moreover, we identified genes whose temperature effect size changed around the switch in temperature sensitivity of development rate. Interesting candidate genes for regulating the temperature sensitivity of egg development rate included genes involved in histone modification, hormonal signalling, nervous system development, and circadian clock genes. In conclusion, the diverse sets of temperature-responsive genes we found here indicate that there are many potential targets of selection to change the temperature sensitivity of embryonic development rate. Identifying for which of these genes there is genetic variation in wild insect populations will give insight into their adaptive potential in the face of climate change.</p>
Architecture of Pol II(G) and molecular mechanism of transcription regulation by Gdown1
<p>This repository contains the modeling files and the analysis related to the article <a href="https://www.ncbi.nlm.nih.gov/pubmed/30190596">"Architecture of Pol II(G) and molecular mechanism of transcription regulation by Gdown1"</a> by Jishage et al. in Nat Struct Mol Biol 2018.</p> <p><strong>For more information</strong> about how to reproduce this modeling, see the <a href="https://salilab.org/pol_ii_g/">Sali lab website</a> or the README file.</p>
Analysis Products: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains analysis products for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>. Please refer to the READMEs in the directories, which are summarized below.</p> <p>The record contains the following files:<br> <br> `clusters.tsv`: <strong> </strong>contains the cluster id, name and colour of clusters in the paper</p> <p><strong>scATAC.zip</strong></p> <p>Analysis products for the single-cell ATAC-seq data. Contains:</p> <p>- `cells.tsv`: list of barcodes that pass QC. Columns include:<br> - `barcode`<br> - `sample`: (time point)<br> - `umap1`<br> - `umap2`<br> - `cluster`<br> - `dpt_pseudotime_fibr_root`: pseudotime values treating a fibroblast cell as root<br> - `dpt_pseudotime_xOSK_root`: pseudotime values treating xOSK cell as root<br> - `peaks.bed`: list of peaks of 500bp across all cell states. 4th column contains the peak set label. Note that ~5000 peaks are not assigned to any peak set and are marked as NA.<br> - `features.tsv`: 50 dimensional representation of each cell <br> - `cell_x_peak.mtx.gz`: sparse matrix of fragment counts within peaks. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (combine sample + barcode). Rows correspond to peaks in `peaks.bed` </p> <p><strong>scATAC_clusters.zip</strong></p> <p>Analysis products corresponding to cluster pseudo-bulks of the single-cell ATAC-seq data. </p> <p>- `clusters.tsv`: contains the cluster id, name and colour used in the paper<br> - `peaks`: contains `overlap_reproducibilty/overlap.optimal_peak` peaks called using ENCODE bulk ATAC-seq pipeline in the narrowPeak format.<br> - `fragments`: contains per cluster fragment files </p> <p><strong>scATAC_scRNA_integration.zip</strong></p> <p>Analysis products from the integration of scATAC with scRNA. Contains:</p> <p>- `peak_gene_links_fdr1e-4.tsv`: file with peak gene links passing FDR 1e-4. For analyses in the paper, we filter to peaks with absolute correlation >0.45.<br> - `harmony.cca.30.feat.tsv`: 30 dimensional co-embedding for scATAC and scRNA cells obtained by CCA followed by applying Harmony over assay type.<br> - `harmony.cca.metadata.tsv`: UMAP coordinates for scATAC and scRNA cells derived from the Harmony CCA embedding. First column contains barcode.</p> <p><strong>scRNA.zip</strong></p> <p>Analysis products for the single-cell RNA-seq data. Contains:</p> <p>- `seurat.rds`: seurat object that contains expression data (raw counts, normalized, and scaled), reductions (umap, pca), knn graphs, all associated metadata. Note that barcode suffix (1-9 corresponds to samples D0, D2, ..., D14, iPSC)<br> - `genes.txt`: list of all genes<br> - `cells.tsv`: list of barcodes that pass QC across samples. Contains:<br> - `barcode_sample`: barcode with index of sample (1-9 corresponding to D0, D2, ..., D14, iPSC) <br> - `sample`: sample name (D0, D2, .., D14, iPSC)<br> - `umap1`<br> - `umap2`<br> - `nCount_RNA`<br> - `nFeature_RNA`<br> - `cluster`<br> - `percent.mt`: percent of mitochondrial transcripts in cell<br> - `percent.oskm`: percent of OSKM transcripts in cell<br> - `gene_x_cell.mtx.gz`: sparse matrix of gene counts. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (barcode suffix contains sample information). Rows correspond to genes in `genes.txt` <br> - `pca.tsv`: first 50 PC of each cell<br> - `oskm_endo_sendai.tsv`: estimated raw counts (cts, may not be integers) and log(1+ tp10k) normalized expression (norm) for endogenous and exogenous (Sendai derived) counts of POU5F1 (OCT4), SOX2, KLF4 and MYC genes. Rows are consistent with `seurat.rds` and `cells.tsv`</p> <p><strong>multiome.zip</strong></p> <p><em>multiome/snATAC:</em></p> <p>These files are derived from the integration of nuclei from multiome (D1M and D2M), with cells from day 2 of scATAC-seq (labeled D2). </p> <p>- `cells.tsv`: This is the list of nuclei barcodes that pass QC from multiome AND also cell barcodes from D2 of scATAC-seq. Includes:<br> - `barcode`<br> - `umap1`: These are the coordinates used for the figures involving multiome in the paper.<br> - `umap2`: ^^^ <br> - `sample`: D1M and D2M correspond to multiome, D2 corresponds to day 2 of scATAC-seq<br> - `cluster`: For multiome barcodes, these are labels transfered from scATAC-seq. For D2 scATAC-seq, it is the original cluster labels. <br> - `peaks.bed`: This is the same file as scATAC/peaks.bed. List of peaks of 500bp. 4th column contains the peak set label. Note that ~5000 peaks are not assigned to any peak set and are marked as NA.<br> - `cell_x_peak.mtx.gz`: sparse matrix of fragment counts within peaks. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (combine sample + barcode). Rows correspond to peaks in `peaks.bed`.<br> - `features.no.harmony.50d.tsv`: 50 dimensional representation of each cell prior to running Harmony (to correct for batch effect between D2 scATAC and D1M,D2M snMultiome). Rows correspond to cells from `cells.tsv`.<br> - `features.harmony.10d.tsv`: 10 dimensional representation of each cell after running Harmony. Rows correspond to cells from `cells.tsv`.</p> <p><em>multiome/snRNA:</em></p> <p>- `seurat.rds`: seurat object that contains expression data (raw counts, normalized, and scaled), reductions (umap, pca),associated metadata. Note that barcode suffix (1,2 corresponds to samples D1M, D2M). Please use the UMAP/features from snATAC/ for consistency.<br> - `genes.txt`: list of all genes (this is different from the list in scRNA analysis)<br> - `cells.tsv`: list of barcodes that pass QC across samples. Contains:<br> - `barcode_sample`: barcode with index of sample (1,2 corresponding to D1M, D2M respectively) <br> - `sample`: sample name (D1M, D2M)<br> - `nCount_RNA`<br> - `nFeature_RNA`<br> - `percent.oskm`: percent of OSKM genes in cell<br> - `gene_x_cell.mtx.gz`: sparse matrix of gene counts. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (barcode suffix contains sample information). Rows correspond to genes in `genes.txt` </p>
Transcriptional regulation underlying the temperature response of embryonic development rate in the winter moth
Open the record for dataset details and reuse information.
Data from: FSHB transcription is regulated by a novel 5’ distal enhancer containing a fertility-associated single nucleotide polymorphism
Open the record for dataset details and reuse information.
Transient intracellular acidification regulates the core transcriptional heat shock response
<p>Heat shock induces a conserved transcriptional program regulated by heat shock factor 1 (Hsf1) in eukaryotic cells. Activation of this heat-shock response is triggered by heat-induced misfolding of newly synthesized polypeptides, and so has been thought to depend on ongoing protein synthesis. Here, using the budding yeast <em>Saccharomyces cerevisiae</em>, we report the discovery that Hsf1 can be robustly activated when protein synthesis is inhibited, so long as cells undergo cytosolic acidification. Heat shock has long been known to cause transient intracellular acidification which, for reasons which have remained unclear, is associated with increased stress resistance in eukaryotes. We demonstrate that acidification is required for heat shock response induction in translationally inhibited cells, and specifically affects Hsf1 activation. Physiological heat-triggered acidification also increases population fitness and promotes cell cycle reentry following heat shock. Our results uncover a previously unknown adaptive dimension of the well-studied eukaryotic heat shock response. </p>
Transcription regulator ACTR contributes pathogenicity through mediating ACT toxin synthesis gene ACTS4 in Alternaria alternata
<p>Host-selective ACT toxin are critical for the pathogenesis of the citrus fungal pathogen <em>Alternaria alternata</em>. The biosynthesis of ACT toxin is mainly regulated by multiple ACT toxin genes located in the secondary metabolite gene cluster. However, the regulatory hierarchy of ACT toxin synthesis by these ACT genes have not been explored. In this study, we reported a transcription regulator <em>ACTR</em> contributes ACT toxin biosynthesis through mediating ACT toxin synthesis gene ACTS4 in <em>Alternaria alternata.</em> We generated <em>ACTR</em>-disrupted and -silenced mutants in the tangerine pathotype of <em>A. alternata.</em> Phenotype analysis showed that the <em>ACTR</em> mutants displayed a significant loss of ACT toxin production and a decreased virulence on citrus leaves whereas the vegetative growth and sporulation were not affected, indicating an essential role of <em>ACTR</em> in both ACT toxin biosynthesis and pathogenicity. To elucidate the transcription network of ACTR, we performed RNA-Seq experiments on wild-type and <em>ACTR</em> null mutant and identified genes that were differentially expressed between two genotypes. Transcriptome profiling and RT-qPCR analysis demonstrated that the ACT toxin biosynthetic gene <em>ACTS4</em> is down-regulated in<em> ACTR </em>mutant<em>.</em> We generated <em>ACTS4 </em>knock-down mutant and found that the pathogenicity of <em>ACTS4</em> mutant was severely impaired. Interestingly, both <em>ACTR</em> and <em>ACTS4</em> are not involved in the response to different abiotic stresses including oxidative stress, salt stress, cell-wall disrupting regents, and Cu<sup>2+</sup>, indicating the function of these two genes is highly specific. In conclusion, our results highlight the important regulatory role of <em>ACTR</em> in ACT toxin biosynthesis through mediating ACT toxin synthesis gene ACTS4 and underline the essential role of in the tangerine pathotype of <em>A. alternata</em>.</p>
Dataset of results for a copper switch for inducing CRISPR/Cas9-based transcriptional activation tightly regulates gene expression in Nicotiana benthamiana.
<p>CRISPR-based programmable transcriptional activators (PTAs) are used in plants for rewiring gene networks. Better tuning of their activity in a time and dose-dependent manner should allow precise control of gene expression. Here, we report the optimization of a Copper Inducible system called CI-switch for conditional gene activation in Nicotiana benthamiana. In the presence of copper, the copper-responsive factor CUP2 undergoes a conformational change and binds a DNA motif named copper-binding site (CBS). In this study, we tested several activation domains fused to CUP2 and found that the non-viral Gal4 domain results in strong activation of a reporter gene equipped with a minimal promoter, offering advantages over previous designs. To connect copper regulation with downstream programmable elements, several copper-dependent configurations of the strong dCasEV2.1 PTA were assayed, aiming at maximizing activation range, while minimizing undesired background expression. The best configuration involved a dual copper regulation of the two protein components of the PTA, namely dCas9:EDLL and MS2:VPR, and a constitutive RNA pol III-driven expression of the third component, a guide RNA with anchoring sites for the MS2 RNA-binding domain. With these optimizations, the CI/dCasEV2.1 system resulted in copper-dependent activation rates of 2,600-fold and 245-fold for the endogenous N. benthamiana DFR and PAL2 genes, respectively, with negligible expression in the absence of the trigger. The tight regulation of copper over CI/dCasEV2.1 makes this system ideal for the conditional production of plant-derived metabolites and recombinant proteins in the field.</p>
Data from: Transcriptional profiling of lung macrophages following ozone exposure in mice identifies signaling pathways regulating immunometabolic activation
<p>Macrophages play a key role in ozone-induced lung injury by regulating both the initiation and resolution of inflammation. These distinct activities are mediated by pro-inflammatory and anti-inflammatory/pro-resolution macrophages which sequentially accumulate in injured tissues. Macrophage activation is dependent, in part, on intracellular metabolism. Herein, we used RNA-sequencing (seq) to identify signaling pathways regulating macrophage immunometabolic activity following exposure of mice to ozone (0.8 ppm, 3 hr) or air control. Analysis of lung macrophages using an Agilent Seahorse showed that inhalation of ozone increased macrophage glycolytic activity and oxidative phosphorylation at 24 and 72 hr post exposure. An increase in the percentage of macrophages in the S phase of the cell cycle was observed 24 hr post ozone. RNA-seq revealed significant enrichment of pathways involved in innate immune signaling and cytokine production among differentially expressed genes at both 24 and 72 hr after ozone, while pathways involved in cell cycle regulation were upregulated at 24 hr and intracellular metabolism at 72 hr. An interaction network analysis identified tumor suppressor 53 (TP53), E2F family of transcription factors (E2Fs), Cyclin Dependent Kinase Inhibitor 1A (CDKN1a/p21), and Cyclin D1 (CCND1) as upstream regulators of cell cycle pathways at 24 hr and TP53, nuclear receptor subfamily 4 group a member 1 (NR4A1/Nur77), and estrogen receptor alpha (ESR1/ERα) as central upstream regulators of mitochondrial respiration pathways at 72 hr. These results highlight the complex interaction between cell cycle, intracellular metabolism, and macrophage activation which may be important in the initiation and resolution of inflammation following ozone exposure.</p>
Post-transcriptional regulation supports the homeostatic expression of mature RNA
<h3>Overview</h3> <p>This dataset consists of RNA sequencing data stored in HDF5 files. The data has been processed to quantify gene expression changes at the precursor RNA (preRNA) and mature RNA (matureRNA) levels across various biological conditions, including normal tissues and diseases. The dataset supports the study "Post-transcriptional regulation supports the homeostatic expression of mature RNA," providing insight into the general influence of post-transcriptional regulation (PTR) on gene expression homeostasis.</p> <h3>File Contents</h3> <p>Each HDF5 file contains the following datasets:</p> <ul> <li><strong>logFC_MatureRNA</strong>: Represents the log fold change of mature RNA expression levels between different conditions or tissues.</li> <li><strong>logCPM_MatureRNA</strong>: Represents the log counts per million for mature RNA.</li> <li><strong>PValue_MatureRNA</strong>: Represents the p-value for the statistical test performed on mature RNA expression levels.</li> <li><strong>FDR_MatureRNA</strong>: Represents the false discovery rate for the mature RNA expression levels.</li> <li><strong>logFC_preRNA</strong>: Represents the log fold change of precursor RNA expression levels between different conditions or tissues.</li> <li><strong>logCPM_preRNA</strong>: Represents the log counts per million for precursor RNA.</li> <li><strong>PValue_preRNA</strong>: Represents the p-value for the statistical test performed on precursor RNA expression levels.</li> <li><strong>FDR_preRNA</strong>: Represents the false discovery rate for the precursor RNA expression levels.</li> <li><strong>genes</strong>: A list of gene identifiers used in the study.</li> <li><strong>studies</strong>: A list of study or condition names included in the dataset.</li> </ul> <h3>File Naming Convention</h3> <p>The HDF5 files are named to reflect the filtering applied to exclude lowly expressed genes (logCPM > 5), and DESeq2 and edgeR refer to the tools used for differential expression analysis:</p> <ul> <li><code>DESeq2_filtered_exon_intron_abnormal_conditions_all_reads.h5</code></li> <li><code>edgeR_filtered_exon_intron_abnormal_conditions_all_reads.h5</code></li> <li><code>DESeq2_filtered_exon_intron_gtex_tissue_all_reads.h5</code></li> <li><code>edgeR_filtered_exon_intron_gtex_tissue_all_reads.h5</code></li> </ul> <h3>Example Usage</h3> <p>To query the HDF5 files, you can use the provided `<a href="https://github.com/suzheng/PTR_RNA_seq/blob/main/notebooks/parse_h5_file.ipynb">parse_h5_file.ipynb</a>` example notebook to parse the HDF5 files.</p>
Interaction of N-3-oxododecanoyl homoserine lactone with transcriptional regulator LasR of Pseudomonas aeruginosa: Insights from molecular docking and dynamics simulations
<p>Dataset and supplementary files of the research: Interaction of N-3-oxododecanoyl homoserine lactone with transcriptional regulator LasR of Pseudomonas aeruginosa: Insights from molecular docking and dynamics simulations (https://doi.org/10.1101/121681)</p> <p>- Supporting Information</p> <p>- Input: Parameters and initial structures</p> <p>- Output: Trajectories, Docking poses</p> <p>Gromacs (multi-core with CUDA) was used for the simulations.</p> <p>Autodock Vina, FlexAid and rDock were used for molecular docking.</p>
Transcriptional stochasticity reveals multiple mechanisms of long non-coding RNA regulation at the Xist – Tsix locus
<p>Data and codes for Figure reprodicbility, image processing and demo.</p>
Specifying cellular context of transcription factor regulons for exploring context-specific gene regulation programs
<p>This repository contains the raw and processed files used in Minaeva et al. 2024.</p> <p>In this version, we have revised the regulon construction pipeline and expanded the dataset to cover 40 common cell lines.</p> <p>The code used to generate these files is available at <a href="https://github.com/LappalainenLab/chip_seq_regulons" target="_new" rel="noreferrer">GitHub - LappalainenLab/chip_seq_regulons</a>.</p> <p>The descriptions of the files contained within each subdirectory are as follows:</p> <h3>1-dataset_stats</h3> <ul> <li><code>per_gene_stats_{approach}_{cell_line}.tsv</code>: Number of TFs regulating a gene according to the respective approach (S2Mb, M2Kb, or S2Kb) in a given cell line.</li> <li><code>per_tf_stats_{approach}_{cell_line}.tsv</code>: Number of target genes regulated by a TF according to the respective approach (S2Mb, M2Kb, or S2Kb) in a given cell line.</li> </ul> <h3>1-network_enrichment</h3> <ul> <li><code>enrich_scores_remap_all_tfs_K562.tsv</code>: Results of fitting logistic regression for testing the enrichment of the K562 regulon in other biological networks (PPI, coexpression, experimental trans-networks).</li> </ul> <h3>2-plot_decoupler_comparison_benchmark_across_cells</h3> <ul> <li><code>{cell_line}_comparison_benchmark.tsv</code>: Results of benchmarking S2Mb, M2Kb, CollecTri, Dorothea, ChIP-Atlas, RegNet, and TRRUST regulons using the decoupler package and the KnockTF database. Cell lines considered are K562, HepG2, and MCF7 (see Methods for benchmarking pipeline details).</li> </ul> <h3>2-plot_decoupler_filter_benchmark_across_methods</h3> <ul> <li><code>{cell_line}_filtering_benchmark.tsv</code>: Results of benchmarking S2Mb, M2Kb, and S2Kb regulons with different filters applied using the decoupler package and the KnockTF database. Cell lines considered are K562, HepG2, and MCF7 (see Methods for benchmarking pipeline details).</li> </ul> <h3>3-tf_activity</h3> <ul> <li><code>aml_k562_activity_{regulon}_sc.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy hematopoietic stem cells (HSCs) and abnormal AML progenitor cells following the decoupler pipeline. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_activity_estimates_hsc_sc.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy HSCs and abnormal AML progenitor cells across regulons.</li> <li><code>aml_dhsc_ahsc_activity_{regulon}_sc.tsv</code>: Results of TF activity analysis based on a respective regulon between leukemic activated and dormant HSCs following the decoupler pipeline. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_activity_estimates_dhsc_ahsc_sc.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between leukemic activated and dormant HSCs across regulons.</li> <li><code>bc_bas_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer following the decoupler pipeline. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_activity_estimates_bas.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer across regulons.</li> <li><code>bc_lum_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer following the decoupler pipeline. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_activity_estimates_lum.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer across regulons.</li> <li><code>hep_activity_{regulon}.tsv</code>: Results of TF activity analysis based on a respective regulon between neoplastic and healthy liver cells following the decoupler pipeline. Regulons considered are HepG2-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>hep_activity_estimates.tsv</code>: Summary of the TF activity analysis for statistically significantly dysregulated TFs between neoplastic and healthy liver cells across regulons.</li> </ul> <h3>3-tf_disease_enrichment</h3> <ul> <li><code>aml_{database}_enrich_{regulon}_hsc_sc.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy HSCs and abnormal AML progenitor cells following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>aml_{database}_enrich_{regulon}_dhsc_ahsc_sc.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between leukemic activated and dormant HSCs following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are K562-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_{database}_enrich_{regulon}_bas.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from basal breast cancer following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, and OMIM. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>bc_{database}_enrich_{regulon}_lum.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between healthy epithelial breast cells and malignant epithelial cells from luminal type A breast cancer following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, and OMIM. Regulons considered are MCF7-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> <li><code>hep_{database}_enrich_{regulon}.tsv</code>: Results of enrichment analysis of dysregulated TFs identified based on a respective regulon between neoplastic and healthy liver cells following the decoupler pipeline. Databases considered are COSMIC, DisGeNet, OMIM, and KEGG. Regulons considered are HepG2-specific ChIP-Atlas and M2Kb regulons, and generalized CollecTri regulon.</li> </ul> <h3>regulons</h3> <ul> <li><code>{cell_line}_regulon.tsv</code>: S2Mb, M2Kb, and S2Kb regulons generated in this study with all acquired annotations (see Methods for details).</li> </ul> <p>External regulons used for comparison. Cell lines considered are K562, HepG2, MCF7, and GM12878:</p> <ul> <li><code>ChIP-Atlas_target_genes_{cell_line}.tsv</code>: Customized ChIP-Atlas regulons (see Methods for details).</li> <li><code>Revised_Supplemental_Table_S3_Normal.csv</code>: Dorothea regulon collected from supplementary materials of Garcia-Alonso et al. (2019).</li> </ul> <h3>s3-network_enrichment</h3> <ul> <li><code>enrich_scores_remap_all_tfs_{cell_line}.tsv</code>: Results of fitting logistic regression for testing the enrichment of cell-line-specific regulons in PPI networks (see Methods and corresponding GitHub repository for details). Cell lines considered are K562, HepG2, MCF7, and GM12878.</li> </ul> <p> </p>
Data for "Dynamic switching of transcriptional regulators between two distinct low-mobility chromatin states"
<p>This deposit contains all the single-molecule trajectories reported in "Dynamic switching of transcriptional regulators between two distinct low-mobility chromatin states", Science Advances, 2023.</p> <p>To access the tracks, open the mat file in MATLAB. This contains a MATLAB table with the following fields:</p> <p><strong>summary_table.cell_protein{i}</strong> identifies the i<sup>th</sup> dataset i.e. cell line + protein + treatment.</p> <p><strong>summary_table.X{i}{j}</strong> is an Nx2 array of x and y coordinates (in microns) for track j in condition i. N is the number of localizations in that track.</p> <p>Time interval between localizations is 200 ms.</p> <p>Details on data acquisition and tracking parameters can be found in the associated manuscript.</p>
Transcription Factor Binding Regulates Chromatin Architecture
<p>Matlab scripts to calculate fiber packing ratio, sedimentation coefficient, volume, and radius of gyration.</p> <p>Representative structures for chromatin fibers in pdb format. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.