Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,276
datasets available to search
ShareScore release 0.9.0
Dataset results
4,276 results for “Transcription Factors”
The transcription factor network of E. coli steers global responses to shifts in RNAP concentration
Open the record for dataset details and reuse information.
Crystal structure of the coild-coil oligomerisation domain of the transcription factor PHOSPHATE STARVATION RESPONSE 1 (PHR1) from Arabidopsis - native dataset form 3
<p>gzipped tar archive contains the raw diffraction images (Pilatus 2M-F detector, SLS beamline PXIII), the data processing directory (xds_nat1) and the scaled and converted structure factors (xdsconv_nat1).</p>
Crystal structure of the coild-coil oligomerisation domain of the transcription factor PHOSPHATE STARVATION RESPONSE 1 (PHR1) from Arabidopsis - native dataset form 2
<p>the gzipped tar archive contains the raw diffraction images (Swiss Light Source beamline PXIII, Pilatus 2M-F detector), the data processing directory (xds_nat1) and the scaled and converted structure factors (xdsconv_nat1)</p>
Crystal structure of the coild-coil oligomerisation domain of the transcription factor PHOSPHATE STARVATION RESPONSE 1 (PHR1) from Arabidopsis - native dataset form 1
<p>gz tar archive contains the diffraction images (Pilatus 2M-F detector, SLS beamline PXIII), the data processing files (xds_nat1) and the scaled and converted data in mtz format (xdsconv_nat1)</p>
Classification dataset for ENCODE-Roadmap DNase-seq peaks and Transcription Factor ChIP-seq peaks
<p>Classification dataset for machine learning on epigenomic landscapes from ENCODE and Roadmap Epigenomics. The dataset includes the original peak files (BED format) for the DNase-seq and transcription factor (TF) ChIP-seq experiments that were used to label genomic regions as positives or negatives. The DNase-seq peak files along with processing details are in the file "encode-roadmap.DNase-seq.peaks.tar.gz". The TF ChIP-seq peak files along with processing details are in the file "encode.ChIP-seq.peaks.tar.gz". The processed dataset stored in hdf5 format files along with processing details are in the file "nn.encode-roadmap.hdf5_files.tar.gz".</p>
DNA methylation regulates transcription factor specific neurodevelopmental but not sexually dimorphic gene expression dynamics in zebra finch telencephalon
<p>Supplementary Information for the manuscript: "DNA methylation regulates transcription factor specific neurodevelopmental but not sexually dimorphic gene expression dynamics in zebra finch telencephalon"</p>
Data from: Discovery and information-theoretic characterization of transcription factor binding sites that act cooperatively
Transcription factor binding to the surface of DNA regulatory regions is one of the primary causes of regulating gene expression levels. A probabilistic approach to model protein–DNA interactions at the sequence level is through position weight matrices (PWMs) that estimate the joint probability of a DNA binding site sequence by assuming positional independence within the DNA sequence. Here we construct conditional PWMs that depend on the motif signatures in the flanking DNA sequence, by conditioning known binding site loci on the presence or absence of additional binding sites in the flanking sequence of each siteʼs locus. Pooling known sites with similar flanking sequence patterns allows for the estimation of the conditional distribution function over the binding site sequences. We apply our model to the Dorsal transcription factor binding sites active in patterning the Dorsal–Ventral axis of Drosophila development. We find that those binding sites that cooperate with nearby Twist sites on average contain about 0.5 bits of information about the presence of Twist transcription factor binding sites in the flanking sequence. We also find that Dorsal binding site detectors conditioned on flanking sequence information make better predictions about what is a Dorsal site relative to background DNA than detection without information about flanking sequence features.
Data from: Adaptive evolution and divergent expression of heat stress transcription factors in grasses
Background: Heat stress transcription factors (Hsfs) regulate gene expression in response to heat and many other environmental stresses in plants. Understanding the adaptive evolution of Hsf genes in the grass family will provide potentially useful information for the genetic improvement of modern crops to handle increasing global temperatures. Results: In this work, we performed a genome-wide survey of Hsf genes in 5 grass species, including rice, maize, sorghum, Setaria, and Brachypodium, by describing their phylogenetic relationships, adaptive evolution, and expression patterns under abiotic stresses. The Hsf genes in grasses were divided into 24 orthologous gene clusters (OGCs) based on phylogeneitc relationship and synteny, suggesting that 24 Hsf genes were present in the ancestral grass genome. However, 9 duplication and 4 gene-loss events were identified in the tested genomes. A maximum-likelihood analysis revealed the effects of positive selection in the evolution of 11 OGCs and suggested that OGCs with duplicated or lost genes were more readily influenced by positive selection than other OGCs. Further investigation revealed that positive selection acted on only one of the duplicated genes in 8 of 9 paralogous pairs, suggesting that neofunctionalization contributed to the evolution of these duplicated pairs. We also investigated the expression patterns of rice and maize Hsf genes under heat, salt, drought, and cold stresses. The results revealed divergent expression patterns between the duplicated genes. Conclusions: This study demonstrates that neofunctionalization by changes in expression pattern and function following gene duplication has been an important factor in the maintenance and divergence of grass Hsf genes.
Identify key transcription factors in cell fate determination using ANANSE
<p>The example data used in ANANSE protocol.</p>
Data from: Limits on information transduction through amplitude and frequency regulation of transcription factor activity
Signaling pathways often transmit multiple signals through a single shared transcription factor (TF) and encode signal information by differentially regulating TF dynamics. However, signal information will be lost unless it can be reliably decoded by downstream genes. To understand the limits on dynamic information transduction, we apply information theory to quantify how much gene expression information the yeast TF Msn2 can transduce to target genes in the amplitude or frequency of its activation dynamics. We find that although the amount of information transmitted by Msn2 to single target genes is limited, information transduction can be increased by modulating promoter cis-elements or by integrating information from multiple genes. By correcting for extrinsic noise, we estimate an upper bound on information transduction. Overall, we find that information transduction through amplitude and frequency regulation of Msn2 is limited to error-free transduction of signal identity, but not signal intensity information.
Transcription Factor Co-Expression Mediates Lineage Priming for Embryonic and Extra-Embryonic Differentiation
<p>In early mammalian development, cleavage stage blastomeres and inner cell mass (ICM) cells co-express embryonic and extra-embryonic transcriptional determinants. Using a double protein-based reporter we identify an embryonic stem cell (ESC)population that co-expresses the extra-embryonic factor GATA6 alongside the embryonic factor SOX2. Based on single cell transcriptomics, we find this population resembles the unsegregated ICM, exhibiting enhanced differentiation potential for endoderm while maintaining epiblast competence. To relate transcription factor binding in these to future fate, we describe a complete enhancer set in both ESCs and naïve extra-embryonic endoderm stem cells and assess SOX2 and GATA6 binding at these elements in the ICM-like ESC sub-population. Both factors support cooperative recognition in these lineages, with GATA6 bound alongside SOX2 on a fraction of pluripotency enhancers and SOX2 alongside GATA6 more extensively on endoderm enhancers, suggesting that cooperative binding between these antagonistic factors both supports self-renewal and prepares progenitor cells for later differentiation.</p>
Supplementary File to "Both binding strength and evolutionary accessibility affect the population frequency of transcription factor binding sequences in Arabidopsis thaliana" (Genome Biology and Evolution)
<p>This data is supplementary file 1 of the following publication:</p> <p>Schweizer G, Wagner A. "Both binding strength and evolutionary accessibility affect the population frequency of transcription factor binding sequences in Arabidopsis thaliana" (Genome Biology and Evolution)</p>
The underground life of homeodomain-leucine zipper transcription factors
<p class="western"><span>Roots are the anchorage organs of plants, responsible for water and nutrient uptake, exhibiting high plasticity. Root architecture is driven by the interactions of biomolecules, including transcription factors (TFs) and hormones that are crucial players regulating root plasticity. Multiple TF families are involved in root development; some, such as ARFs and LBDs, have been well characterized, whereas others remain less investigated. In this review, we synthesize the current knowledge about the involvement of the large family of homeodomain-leucine zipper (HD-Zip) TFs in root development. This family is divided into four subfamilies (I to IV), mainly according to structural features, such as additional motifs aside from HD-Zip, as well as their size, gene structure, and expression patterns. We explored and analyzed public databases and the scientific literature regarding HD-Zip TFs in Arabidopsis and other species. Most members of the four HD-Zip subfamilies are expressed in specific cell types and several ones from each group have assigned functions in root development. Notably, a high proportion of the studied proteins are part of intricate regulation pathways involved in primary and lateral root growth and development.</span></p>
Transcription factor expression is the main determinant of variability in gene co-activity
<p><strong>Summary</strong></p> <p>Co-activity scores for 343 GEUVADIS LCLs and ABC scores for 68 LCLs, of which 30 are contained in both.</p> <p><strong>Project abstract</strong></p> <p>Many genes are co-regulated and, when proximal, form domains of coordinated gene activity. However, the regulatory determinants of domain co-activity remain unclear. Here, we leverage human individual variation in gene expression to characterize the regulatory processes underlying the activities of such domains and systematically quantify their effect sizes. We employ transcriptional decomposition to extract from RNA expression data an expression component related to co-activity revealed by genomic positioning. This strategy reveals close to 1,500 domains of co-activity, covering most expressed genes, of which the large majority are invariable across individuals. Focusing specifically on domains with high variation in co-activity reveals that neighboring genes contained within variable co-activity domains have a higher sharing of eQTLs, a higher variability in enhancer interactions, and a specific enrichment of binding by variably expressed transcription factors. Through careful quantification of the relative contributions of regulatory activities underlying co-activity, we find transcription factor expression levels to be the main determinant of gene co-activity, indicating that distal <em>trans</em> effects contribute more than local genetic variation to individual variation in co-activity domains. </p> <p><strong>Included files</strong></p> <p>Co_activity_scores_343_individuals.tsv.zip - Contains co-activity scores for included individuals. Columns include chromosome, bin, start, end and one column per individual containing the co-activity score.</p> <p>ABC_scores_68_individuals.tsv.zip - Contains ABC scores for included individuals. Columns include chromosome, start of putative enhancer region, end of putative enhancer region, name of putative enhancer region, target gene, TSS of target gene, LCL identifier, ABC score</p>
The developmental and evolutionary characteristics of transcription factor binding site clustered regions based on an explainable machine learning model
<p>## Identification of transcription factor binding sites clustered regions</p> <p>First, the TFBSs were identified from ATAC-seq peaks by FIMO. The position-specific weight matrices (PWMs) of transcription factors were downloaded from CIS-BP databases. The genomic sequences under the open chromatin regions were used as inputs for FIMO with a custom library of all motifs for each species to scan for motif instances at a p-value threshold of 1e-5. </p> <p>Then, an established method was used to identify TFCRs by performing the Gaussian kernel density estimations across the genome (with a bandwidth of 300bp centered on each TFBS). Each peak in density profile was considered a TFCR. To determine the complexity of each TFCR, the Gaussian kernelized distances from each peak that contributed at least 0.1 to its strength were determined. The complexity of each TFCR was determined by the quantity and proximity of the contributing TFBS. We combined motif instances based on the TF family information from CIS-BP to calculate the complexity of TFCR. The window for each TFCR was determined by finding the maximum distance (in bp) from the TFCR to a contributing TF and then adding 150 bp (one-half of the bandwidth). Each window was centered on the TFCR. The identified TFCR was grouped into 10 groups based on their complexity from low to high. </p> <p>usage: <br>indir="Human_fimo" # the directory where you put the output files of FIMO <br>motifMap="Homo_sapiens_2020_0920/TF_Information_all_motifs_plus.txt" # the mapping relationship of TF and its TF family from CIS-BP <br>cd Codes/TFCR_embryo <br>perl d-motif_combine.pl $indir TFfamily $motifMap <br>perl e-tfpos_combine.pl TFfamily <br>perl f1-tf_bed-new-c.pl TFfamily <br>perl 0-merge-TFCR.pl $indir TFfamily </p>
Dissociation rate compensation mechanism for budding yeast pioneer transcription factors
<p>Single molecule data sets related to "Dissociation rate compensation mechanism for budding yeast pioneer transcription factors."</p>
Benchmark and integration of resources for the estimation of human transcription factor activities
<p>Data used to benchmark human TF-target datasets via TF activities in 3 benchmark datasets. Described in <a href="https://www.biorxiv.org/content/biorxiv/early/2018/06/18/337915.full.pdf">Garcia-Alonso et al 2019</a></p> <p>Check <a href="https://github.com/saezlab/TFbenchmark">https://github.com/saezlab/TFbenchmark</a> to access the corresponding code.</p> <p> </p> <p> </p> <p><strong>Study abstract</strong></p> <p>Prediction of transcription factor (TF) activities from the gene expression of their targets (i.e. TF regulon) is becoming a widely-used approach to characterize the functional status of transcriptional regulatory circuits. Several strategies and datasets have been proposed to link the target genes likely regulated by a TF, each one providing a different level of evidence. The most established ones are: (i) manually curated repositories, (ii) interactions derived from ChIP-seq binding data, (iii) <em>in silico</em> prediction of TF binding on gene promoters, and (iv) reverse-engineered regulons from large gene expression datasets. However, it is not known how these different sources of regulons affect the TF activity estimations, and thereby downstream analysis and interpretation. Here we compared the accuracy and biases of these strategies to define human TF regulons by means of their ability to predict changes in TF activities in three reference benchmark datasets. We assembled a collection of TF-target interactions among 1,541 TFs and evaluated how the different molecular and regulatory properties of the TFs, such as the DNA-binding domain, specificities or mode of interaction with the chromatin, affect the predictions of TF activity changes. We assessed their coverage and found little overlap on the regulons derived from each strategy and better performance by literature-curated information followed by ChIP-seq data. We provide an integrated resource of all TF-target interactions derived through these strategies with a confidence score, as a resource for enhanced prediction of TF activities.</p>
Chlamydomonas pacifica Lipid Transcription Factors
<p>Contains the plasmid sequence for <span>pJPCHx1_CpaDpWRI1, pJPCHx1_CpaAtWRI1, pJPCHx1_CpaMYB6, pJPCHx1_CpabZIP1, pJPCHx1_CpaSPL12, pJPCHx1_CpaPSR1, pJPCHx1_CpaCHT7, pJPCHx1_CpaNRR1, and pJPCHx1_CpaLRL1.</span></p>
Transcription factor clusters in fruit fly embryos
<p>These files represent the dataset used in generating plots shown in <a href="https://arxiv.org/abs/2403.02943">[2403.02943] Transcription factor clusters as information transfer agents (arxiv.org)</a></p> <p>The MATLAB programs to generate the plots using this data can be found in:</p> <p>https://github.com/ancientman/clusters-plots</p> <p>Please refer to the instructions there.</p> <p> </p> <p> </p>
Fig 5 in Docosahexaenoic Acid (DHA) Reduces LPS-Induced Inflammatory Response Via ATF3 Transcription Factor and Stimulates Src/Syk Signaling-Dependent Phagocytosis in Microglia
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.