Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,276
datasets available to search
ShareScore release 0.9.0
Dataset results
4,276 results for “Transcription factors”
A dominant-negative SOX18 mutant disrupts multiple regulatory layers essential to transcription factor activity
<p>Few genetically dominant mutations involved in human disease have been fully explained at the molecular level. In cases where the mutant gene encodes a transcription factor, the dominant-negative mode of action of the mutant protein is particularly poorly understood. Here, we studied the genome-wide mechanism underlying a dominant-negative form of the SOX18 transcription factor (SOX18<sup>RaOp</sup>) responsible for both the classical mouse mutant <u>Ra</u>gged <u>Op</u>ossum and the human genetic disorder Hypotrichosis-Lymphedema-Telangiectasia-Renal Syndrome. Combining three single-molecule imaging assays in living cells together with genomics and proteomics analysis, we found that SOX18<sup>RaOp</sup> disrupts the system through an accumulation of molecular interferences which impair several functional properties of the wild-type SOX18 protein, including its target gene selection process. The dominant-negative effect is further amplified by poisoning the interactome of its wild-type counterpart, which perturbs regulatory nodes such as SOX7 and MEF2C. Our findings explain in unprecedented detail the multi-layered process that underpins the molecular aetiology of dominant-negative transcription factor function.</p>
Raw Data supporting Promoters adopt distinct dynamic manifestations depending on transcription factor context
<p>All source data is provided as MAT-files, which can be opened in Matlab. In total, the source data contain 270 core data files and 540 processed data files. The data for each individual promoter is stored in a different directory and the 9 promoters are:</p> <ul> <li> <p><em>ALD3 </em></p> </li> <li> <p><em>DCS2 </em></p> </li> <li> <p><em>DDR2 </em></p> </li> <li> <p><em>HXK1 </em></p> </li> <li> <p><em>RTN2 </em></p> </li> <li> <p><em>TKL2 </em></p> </li> <li> <p><em>SIP18 </em></p> </li> <li> <p><em>pSIP18_mut6</em> (also referred to as mutant A4)</p> </li> <li> <p><em>pSIP18_mut21</em> (also referred to as mutant D6)</p> </li> </ul> <p>The data for the first 7 promoters was previously reported in (Hansen and O’Shea, 2013), though in an unnormalized form. That is, it was previously reported as a concentration per cell in arbitrary fluorescence units (AU). In the present manuscript, we have calibrated the data to obtain absolute abundances, such that the MAT-files now contain both the old AU concentration as well as absolute abundances, i.e. number of YFP molecules per cell. The calibration was performed as described in (Huang et al., 2016). Similarly, the data for the last 2 promoters (A4 and D6) was previously reported in its unnormalized form in (Hansen and O’Shea, 2015) and it is here also reported in the form of absolute abundances.</p> <p> </p> <p>The MAT-files containing the raw data have the suffix “_size.mat”. The name of the MAT files describes the experiment. If the file name contains “DM”, then it is a single pulse. Thus, “SIP18_DM_40min_275nM_size.mat” refers to a single 40 min pulse with 275 nM 1-NM-PP1 for the <em>SIP18</em> promoter. Similarly if the file name contains “FM”, e.g. “RTN2_FM_8_5min_690nM_size.mat” then it refers to eight 5 min pulses separated by 5 min intervals at 690 nM for the <em>RTN2</em> promoter. Finally, if the file name contains “FM4”, e.g. “TKL2_FM4_15minINT_690nM_size.mat” then the experiment was four 5 min pulses separated by 15 min intervals at 690 nM for the <em>TKL2</em> promoter. The concentration is the concentration of 1-NM-PP1 that was used and 100 nM, 275 nM, 690 nM and 3mM refers to approximately, 25%, 50%, 75% and 100% Msn2 activation. For full experimental details please see (Hansen and O’Shea, 2013; Hansen et al., 2015).</p> <p> </p> <p>The “_size.mat” MAT-files contain the following variables:</p> <ul> <li> cell_size_pixels</li> <li> CFP</li> <li> CFP_molecules</li> <li> CFP_raw</li> <li> inhibitor_conc</li> <li> MSN2_raw</li> <li> MSN2_RFP</li> <li> pulse_parameters • time</li> <li> YFP</li> <li> YFP_molecules</li> <li> YFP_raw</li> </ul> <p>CFP, CFP_molecules, CFP_raw and YFP, YFP_molecules, YFP_raw are Nx64 matrices, where each row N correspond to a different cell and the 64 columns correspond to the 64 experimentally measured timepoints corresponding to the “time” vector running from -5 min to 152.5 min in increments of 2.5 min and the 1NM-PP1 inhibitor was added at time 0. “CFP_raw” and “YFP_raw” contains raw, uncorrected data, so without photobleaching correction and background subtraction. “CFP” and “YFP” contain corrected data in arbitrary fluorescence units (AU) and report on the concentration (i.e. size normalized). Finally, “CFP_molecules” and “YFP_molecules” contains the total number of CFP and YFP molecules per cell (i.e. this is not a concentration, but the absolute abundance). The area of each cell at each timepoint can be found in the matrix “cell_size_pixels”. Since the cells are live and growing, this will tend to increase during the experiments. Occasionally large fluctuations can occur due to errors in cell segmentation or due to division. For full details on the image analysis and cell segmentation, please see (Hansen and O’Shea, 2013; Hansen et al., 2015).</p> <p> </p> <p>The variables “inhibitor_conc” and “pulse_parameters” refer to the type of experiment and is also given by the name. “inhibitor_conc” gives the 1NMPP1 concentration: 100 nM, 275 nM, 690 nM or 3000 nM. “pulse_parameters” contains either 2 or 3 elements and given the dynamical pulse sequence parameters. Column 1 contains the number of pulses and column 2 the duration of the pulses. Column 3 gives the interval between the pulses if more than one pulse is used – otherwise column 3 is zero.</p> <p> </p> <p>Moreover, on a more technical note it should be noted that the signal-to-noise of the CFP reporter is worse than the YFP reporter. Therefore, we always use the YFP reporter for quantitative analysis. Furthermore, the two other MAT-files “…MSN2.mat” and “…YFP.mat” contain processed data. Please see the ReadMe file on the code for a full description and how these were derived.</p> <p> </p> <p>Finally, Supplementary Table 1 contains the model-inferred parameters for each promoter and condition.</p> <p> </p> <p> </p> <p><strong>References</strong></p> <p>Hansen, A.S., and O’Shea, E.K. (2013). Promoter decoding of transcription factor dynamics involves a trade-off between noise and control of gene expression. Mol. Syst. Biol.</p> <p>Hansen, A.S., and O’Shea, E.K. (2015). Cis Determinants of Promoter Threshold and Activation Timescale. Cell Rep.</p> <p>Hansen, A.S., Hao, N., and OShea, E.K. (2015). High-throughput microfluidics to control and measure signaling dynamics in single yeast cells. Nat. Protoc.</p> <p>Huang, L., Pauleve, L., Zechner, C., Unger, M., Hansen, A.S., and Koeppl, H. (2016). Reconstructing dynamic molecular states from single-cell time series. J. R. Soc. Interface.</p>
Fluorescence complementation enables quantitative imaging of cell penetrating peptide-mediated protein delivery in plants including WUSCHEL transcription factor
<p>These are data related to the manuscript titled "Fluorescence complementation enables quantitative imaging of cell penetrating peptide-mediated protein delivery in plants including WUSCHEL transcription factor" whose preprint can be found here: https://doi.org/10.1101/2022.05.03.490515</p>
Reproductive transcriptome of Nicotiana tabacum male gametophyte unveils decreasing number of new transcription factors during pollen ontogeny
<p>Plants with highly reduced male gametophytes represent a successful adaptation to sexual reproduction, which plays an important role in the colonization and radiation of terrestrial ecosystems. During pollen maturation, microsporocytes to mature pollen grain cells switch from mitosis to meiosis and ultimately form a haploid male gametophyte, a widely used model to study plant development. We performed RNASeq analysis of <em>Nicotiana tabacum</em> at six developmental stages from microspores to mature pollen grain to characterize in detail key transcription factor (TF) genes involved in pollen ontogeny. Our results provide the most complete transcriptomic data during pollen development in the important model plant <em>Nicotiana tabacum. </em>We have identified DEGs associated with six ontogenetic stages of the male gametophyte, providing insights into the molecular regulation of reproductive development by TFs that have a high potential for agronomic research investigating molecular networks associated with pollen sterility and related issues.</p>
Precise modulation of transcription factor levels identifies features underlying dosage sensitivity
<p><strong>Processed data and code for "Precise modulation of transcription factor levels reveals drivers of dosage sensitivity," Naqvi et al 2022.</strong></p> <p><strong>Count/expression data</strong></p> <ul> <li>all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz - ATAC-seq counts from all samples (SOX9 titration and depletion) over all reproducible ATAC-seq peak regions</li> <li>all.sub.150bpclust.greater2.500bp.merge.k27ac.txt.gz - H3K27ac ChIP-seq counts from SOX9 depletion samples over all reproducible peak regions</li> <li>all.sub.150bpclust.greater2.500bp.merge.SOX9titr.V5.counts.txt.gz - V5 (SOX9) ChIP-seq counts from partial SOX9 titration (100%, 60%, 30%, 0%) over all reproducible peak regions</li> <li>all.sub.150bpclust.greater2.500bp.merge.SOX9titr.TWIST1.in.counts.tab.txt.gz - TWIST1 and input ChIP-seq counts from partial SOX9 titration (100%, 60%, 30%, 0%) over all reproducible peak regions</li> <li>rna.salmon.7rep.txi.counts.txt.gz - RNA-seq counts from SOX9 titration samples </li> <li>rna.salmon.7rep.txi.abundance.txt.gz - RNA-seq TPM values from SOX9 titration samples </li> <li>slam.tcreadcount.txt.gz - SLAM-seq T-C conversion-containing read counts (representing newly transcribed mRNAs) from SOX9 depletion samples</li> <li>slam.readcount.txt.gz - SLAM-seq read counts (representing all mRNAs) from SOX9 depletion samples</li> </ul> <p><strong>Metadata</strong></p> <ul> <li>all.protcod.gene.features.txt.gz - Features of interest for all protein-coding genes</li> <li>all.sub.150bpclust.greater2.500bp.merge.features.txt.gz - Features of interest for all reproducible peak regions</li> <li>atac_depletion_3h_24h_design.txt - design matrix for ATAC-seq SOX9 depletion samples</li> <li>atac_titration_48h_design.txt - design matrix for ATAC-seq SOX9 titration samples</li> <li>Homo_sapiens.GRCh38.cdna.all.txid2gene.id.symbol.type.txt.gz - Ensembl transcript types (for filtering to protein-coding genes in various analyses)</li> <li>k27_depletion_3h_24h_design.txt - design matrix for H3K27ac ChIP-seq SOX9 depletion samples</li> <li>v5_sox9titration_design.txt - design matrix for V5 (SOX9) ChIP-seq SOX9 partial titration samples</li> <li>twist1_sox9titration_design.txt - design matrix for TWIST1 ChIP-seq SOX9 partial titration samples</li> <li>rna_titration_48h_design.txt - design matrix for RNA-seq SOX9 titration samples</li> <li>slam_depletion_3h_24h_design.txt - design matrix for SLAM-seq SOX9 depletion samples</li> <li>facialgwas_snpia_ld0.5.hg38.bed - SNPs in LD (r2 > 0.5) with any of the facial GWAS lead SNPs in Supplementary Table 2 of Naqvi, Hoskens, et al, Annu Rev. Hum Genet. Genom. 2022. </li> <li>facialgwas_prsendo_7e5_either_snpia_ld0.5.hg38.bed - SNPs in LD (r2 > 0.5) with the subset of the same facial GWAS SNPs that show significant (p-value < 7e-05, ~corresponding to Bonferonni-corrected p-value of 0.01) association with the PRS endophenotype GWAS in either US or UK cohort. </li> </ul> <p><strong>Scripts</strong></p> <ul> <li>atac_deseq_fitmodels_bs_parallel.R - R code for fitting bootstrapped Hill equations to all SOX9-dependent REs (computationally intensive, so has been coded for parallelization over multiple cores) <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz, atac_titration_48h_design.txt</li> <li>Output: enh_linear_sig_aic_bsmat.txt, enh_linear_sig_aic_bsmat_enhind.txt</li> </ul> </li> <li>atac_deseq_fitmodels.R - R code for fitting Hill equations (no bootstrap) to all SOX9-dependent REs <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz, atac_titration_48h_design.txt</li> <li>Output: enh_linear_sig_aic.rds</li> </ul> </li> <li>atac_k27_depletion_deseq.R - R code for DESeq2 analysis of ATAC and H3K27ac ChIP SOX9 depletion (3h and 24h) <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz, all.sub.150bpclust.greater2.500bp.merge.k27ac.txt.gz, atac_depletion_3h_24h_design.txt</li> <li>Output: atac_depletion_3h_24h_deseq.txt, k27_depletion_3h_24h_deseq.txt </li> </ul> </li> <li>v5_twist1_sox9titration_deseq.R - R code for DESeq2 analysis of V5 (SOX9) and TWIST1 ChIP in partial SOX9 titration (100%, 60%, 30%, 0%) <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.SOX9titr.V5.counts.txt.gz, all.sub.150bpclust.greater2.500bp.merge.SOX9titr.TWIST1.in.counts.tab.txt.gz, v5_sox9titration_design.txt, twist1_sox9titration_design.txt</li> <li>Output: v5_sox9titration_deseq.txt, twist1_sox9titration_deseq.txt </li> </ul> </li> <li>drm.R - Modified version of code from drc() package to prevent errors, install drc() with this version to avoid errors</li> <li>group_comparisons.Rmd - R code to compare computed parameters (i.e. ED50, Hill) between sets of REs/genes utilizing bootstrap information <ul> <li>Input: enh_linear_sig_aic_bsmat_enhind.txt, enh_linear_sig_aic_bsmat.txt.gz, enh_linear_sig_aic.rds, gene_linear_sig_aic_bsmat_enhind.txt, gene_linear_sig_aic_bsmat.txt.gz, gene_linear_sig_aic.rds, all.sub.150bpclust.greater2.500bp.merge.features.txt.gz, all.protcod.gene.features.txt.gz</li> <li>Uses: summarize_bs_helper.R</li> </ul> </li> <li>plot_re_gene_fits.Rmd - R code for plotting individual RE/gene counts and Hill/linear fits <ul> <li>Input: all.sub.150bpclust.greater2.500bp.merge.ATAC.counts.fulldep.3h.24h.txt.gz, atac_titration_48h_design.txt, rna.salmon.7rep.txi.counts.txt.gz, rna.salmon.7rep.txi.abundance.txt.gz, rna_titration_48h_design.txt</li> </ul> </li> <li>rna_deseq_fitmodels_bs_parallel.R - R code for fitting bootstrapped Hill equations to all SOX9-dependent genes (computationally intensive, so has been coded for parallelization over multiple cores) <ul> <li>Input: rna.salmon.7rep.txi.counts.txt.gz, rna.salmon.7rep.txi.abundance.txt.gz, rna_titration_48h_design.txt</li> <li>Output: gene_linear_sig_aic_bsmat.txt, gene_linear_sig_aic_bsmat_enhind.txt</li> </ul> </li> <li>rna_deseq_fitmodels.R - R code for fitting Hill equations (no bootstrap) to all SOX9-dependent genes <ul> <li>Input: rna.salmon.7rep.txi.counts.txt.gz, rna.salmon.7rep.txi.abundance.txt.gz, rna_titration_48h_design.txt</li> <li>Output: gene_linear_sig_aic.rds</li> </ul> </li> <li>slam_depletion_deseq.R - R code for DESeq2/sva analysis of SLAM-seq SOX9 depletion (3h and 24h) <ul> <li>Input: slam.tcreadcount.txt.gz, slam_depletion_3h_24h_design.txt</li> <li>Output: slam_depletion_3h_24h_deseq.txt</li> </ul> </li> <li>summarize_bs_helper.R - Helper functions for group_comparisons.Rmd</li> </ul> <p><strong>Intermediate/output files (some files are gzipped to save space, the Rscripts that output them won't gzip but they expect gzipped input when indicated)</strong></p> <ul> <li>atac_depletion_3h_24h_deseq.txt.gz - DESeq2 output of ATAC SOX9 depletion (3h and 24h)</li> <li>enh_linear_sig_aic_bsmat_enhind.txt - index to name file for SOX9-dependent RE bootstrap output</li> <li>enh_linear_sig_aic_bsmat.txt.gz - SOX9-dependent RE bootstrap output</li> <li>enh_linear_sig_aic.rds - Parameters from Hill equation fit on all SOX9-dependent REs (no bootstrap) (RDS file)</li> <li>gene_linear_sig_aic_bsmat_enhind.txt - index to name file for SOX9-dependent gene bootstrap output</li> <li>gene_linear_sig_aic_bsmat.txt.gz - SOX9-dependent gene bootstrap output</li> <li>gene_linear_sig_aic.rds - Parameters from Hill equation fit on all SOX9-dependent gene (no bootstrap) (RDS file)</li> <li>k27_depletion_3h_24h_deseq.txt.gz - DESeq2 output of H3K27ac ChIP-seq SOX9 depletion (3h and 24h)</li> <li>slam_depletion_3h_24h_deseq.txt.gz - DESeq2 output of SLAM-seq SOX9 depletion (3h and 24h)</li> <li>v5_sox9titration_deseq.txt.gz - DESeq2 output of V5 (SOX9) ChIP-seq in partial SOX9 titration (100%, 60%, 30%, 0%)</li> <li>twist1_sox9titration_deseq.txt.gz - DESeq2 output of TWIS1 ChIP-seq in partial SOX9 titration (100%, 60%, 30%, 0%)</li> </ul> <p><strong>chromatin_predictions.tar.gz (self-contained folder for chromatin-based predictions of gene expression change) contains:</strong></p> <ul> <li>ABC_6conc.sh - Bash script to calculate predicted gene expression change based on ATAC-seq fold-change at each of five SOX9 concentrations (warning: creates a number of very large additional intermediate output files). Requires as input all files in this folder except for all.sub.150bpclust.greater2.500bp.merge.ABC.5Mb.power-0.7.norm.6conc.all.total.txt</li> <li><strong> </strong>all.sub.150bpclust.greater2.500bp.merge.ABC.5Mb.power-0.7.norm.6conc.all.total.txt - ATAC-based predicted fold-change of all genes each of five SOX9 concentrations (78, 52, 25, 8, 0, in that order), relative to 100% SOX9</li> <li>all.sub.150bpclust.greater2.500bp.merge.ATAC.DMSO.counts.txt - ATAC-seq counts over all reproducible peak regions in updepleted samples</li> <li>all.sub.150bpclust.greater2.500bp.merge.bed - bed file of all reproducible peak regions</li> <li>all.sub.150bpclust.greater2.500bp.merge.deseq.allconc.lfc.txt - DESeq2 output from ATAC SOX9 titration, comparing each lowered SOX9 concentration to 100% SOX9</li> <li>all.sub.150bpclust.greater2.500bp.merge.k27ac.dmso.counts.txt - H3K27ac ChIP-seq counts over all reproducible peak regions in updepleted samples</li> <li>hg38_refGene_TSS_collapsed.bed - collapsed TSSs for all genes</li> <li>hg38.genome - genome file</li> </ul> <p> </p>
Transcription factor dataset in the pea genome
<p><strong>Table 1:</strong> TF dataset identified in the pea genome. The 1807 TF dataset is shown with TF family prediction (from PlantTFDB) and the gene description available from the <em>Pisum sativum</em> genome (v1a). The best BlastP hit against <em>M. truncatula </em>(v4) and <em>A. thaliana</em> (TAIR10) proteomes is presented with both coverage (%) and e-value.</p> <p><strong>Table 2:</strong> Pea, Medicago and Arabidopsis TF datasets with unique gene identifier, predicted TF family (from PlantTFDB) and gene description (retrieved from respective genome annotation) is displayed on a second sheet.</p> <p> </p>
Transcription Factor Binding Regulates Chromatin Architecture
<p>Matlab scripts to calculate fiber packing ratio, sedimentation coefficient, volume, and radius of gyration.</p> <p>Representative structures for chromatin fibers in pdb format. </p>
Mapped data: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains mapped sequencing data for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>. It contains single-cell RNA-seq (scRNA) and single-cell ATAC-seq (scATAC) data from a time course of human dermal fibroblasts induced with Yamanaka factors OSKM using a Sendai virus based delivery system. The scRNA and scATAC data is performed at days 0, 2, 4, 6, 8, 10, 12, 14 and the final iPSCs. The experiment was re-performed and single-nucleus multiome (ATAC+RNA) was collected on days 1 and 2. </p> <p>The data is as follows:</p> <p><strong>scATAC</strong>: We used Chromap (commit <a href="https://github.com/haowenz/chromap/tree/6e97125b9">https://github.com/haowenz/chromap/tree/6e97125b9</a>, <a href="https://doi.org/10.1038/s41467-021-26865-w">https://doi.org/10.1038/s41467-021-26865-w</a>) to perform barcode correction, alignment and filtering for each of our samples. The corresponding fragment files (tab separated file containing mapped fragments with columns: chr, start, end, barcode, number of reads) and their tabix indices are available for each sample.</p> <p><strong>scRNA</strong>: We used cellranger v6.0.2 for read mapping and quantification to obtain the counts matrix. We used the GRCh38 2020-A reference. For each sample, the raw and filtered counts matrices are provided. E.g. `D0/raw_feature_bc_matrix.h5` contains an HDF5 object containing gene counts for each barcode and associated metadata for the Day 0 sample. Similarly, the files in `D0/raw_feature_bc_matrix/` contain the same gene x barcode matrix, with the counts matrix in Matrix Market format (`matrix.mtx.gz`), and gene (`features.tsv.gz`) and barcode names (`barcodes.tsv.gz`). </p> <p><strong>multiome</strong>: The ATAC and RNA components are separately processed using the same tools as mentioned above for scATAC and scRNA. Outputs are in the `snATAC` and `snRNA` subdirectories respectively. In addition, the `ATAC.RNA.bc.map.tsv` file contains a map to link snATAC barcodes to snRNA barcodes. </p>
ChromBPNet models and data: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains ChromBPNet models and data used to train the models for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>.</p> <p>`data` contains bigwigs and regions (peaks + non-peaks) used for training each of the models. See `data/README.txt` for more details.</p> <p><strong>Models:</strong></p> <p><em>Loading the model:</em></p> <p>The models were trained using tf1.14. The models are provided in h5 format for tf1.14 (py3.7) and SavedModel format for tf2.X. tf2.X tested only for py3.8-11, tf2.8-13.</p> <p>To load the models in tf1.14:</p> <pre><code class="language-python">model = tf.keras.models.load_model("path/to/model.h5")</code></pre> <p>In tf2:</p> <pre><code class="language-python">model = tf.keras.models.load_model("path/to/model_dir")</code></pre> <p>If all fails, you can load the architecture as provided in `model_arch.py` with default parameters (`bpnet_seq` for bias model and `chrombpnet` for chrombpnet model), and then load the weights using `model.load_weights` from the weights provided in the `weights` directory.</p> <p> </p> <p><em>Usage:</em></p> <p>The bias models take as input one-hot sequence of length 2000. It has 2 outputs, a vector of logits of length 2000, and 1 logcounts scalar:</p> <pre><code class="language-python"># seq_one_hot of length B x 2000 x 4 out_bias_logits, out_bias_logcounts = bias_model.predict(seq_one_hot) # out_bias_logits: B x 2000 # out_bias_logcounts: B x 1</code></pre> <p>The ChromBPNet model takes as input a one-hot sequence of length 2000, bias logits of length 2000 and bias log-counts scalar. It has the same output types as the bias model. To run the chrombpnet model to obtain predictions:</p> <pre><code class="language-python">pred_profile, pred_logcounts = chrombpnet_model.predict([seq_one_hot, out_bias_logits, out_bias_logcounts]) # pred_profile: B x 2000 # pred_logcounts: B x 1 </code></pre> <p>If you wish to obtain the "de-biased" predictions (see Methods), simply pass in zeros instead of the bias model predictions as:</p> <pre><code class="language-python">pred_profile_debiased, pred_logcounts_debiased = chrombpnet_model.predict([seq_one_hot, np.zeros((seq_one_hot.shape[0], 2000)), np.zeros((seq_one_hot.shape[0], 1))])</code></pre> <p>To obtain predicted per-base predicted counts (with or without bias):</p> <pre><code class="language-python">pred_per_base_counts = scipy.special.softmax(pred_profile, axis=-1) * (np.exp(pred_logcounts)-1) # pred_per_base_counts: B x 2000 </code></pre> <p>Note that in general predicted counts can't be compared across models as they are not corrected for sequencing depth.</p> <p> </p> <p><em>Note:</em></p> <p>All bias models used across folds are identical, except for the final intercept term in the counts output (see Methods), that is specific to each cell state, fold combination.</p> <p> </p> <p><em>Folds:</em></p> <p>The splits used for training the different folds are as below:</p> Fold Test Chromosomes Validation Chromosomes 0 chr1 chr8, chr10 1 chr2, chr19 chr1 2 chr3, chr20 chr2, chr19 3 chr6, chr13, chr22 chr3, chr20 4 chr5, chr16, chrY chr6, chr13, chr22 5 chr4, chr15, chr21 chr5, chr16, chrY 6 chr7, chr18, chr14 chr4, chr15, chr21 7 chr11, chr17, chrX chr7, chr18, chr14 8 chr9, chr12 chr11, chr17, chrX 9 chr8, chr10 chr9, chr12 <p>Remaining chromosomes were used as the training chromosome for each fold.</p>
Trihelix transcription factor SlGT31 regulates fruit ripening mediated by ethylene in tomato
<p class="MsoNormal"><span>Trihelix proteins are plant-specific transcription factors that play crucial roles in plant development and stress responses. However, the involvement of trihelix proteins in fruit ripening and transcriptional regulatory mechanisms remains largely unclear. In this study, we cloned a trihelix <span>gene<em> SlGT31</em>, whose relative expression was significantly induced by the application of exogenous ethylene but repressed by 1-methylcyclopropene (1-MCP). Suppression of <em>SlGT31</em> resulted in delayed fruit ripening, decreased accumulation of total carotenoids and ethylene content, and inhibition of relative expression of genes related to ethylene and fruit ripening. Conversely, the opposite results were observed in <em>SlGT31</em>-overexpression lines. Yeast one-hybrid and dual-luciferase assays suggested that SlGT31 could bind to the promoters of two key ethylene biosynthesis genes <em>ACO1 </em>and <em>ACS4.</em> These results indicate that SlGT31 may act as a positive modulator during fruit ripening.</span></span></p>
Monosaccharide transporter OsMST6 is activated by transcription factor OsERF120 to enhance chilling tolerance in rice seedlings
Open the record for dataset details and reuse information.
Data from: The transcription factor Pou3f1 promotes neural fate commitment via activation of neural lineage genes and inhibition of external signaling pathways
Open the record for dataset details and reuse information.
Data from: Short activation domains control chromatin association of transcription factors
Open the record for dataset details and reuse information.
Data from: Decoupling transcription factor expression and activity enables dimmer switch gene regulation
Open the record for dataset details and reuse information.
Data from: A clinal polymorphism in the insulin signaling transcription factor foxo contributes to life-history adaptation in Drosophila
Open the record for dataset details and reuse information.
Data from: Protein kinase FaSnRK2.6 phosphorylates transcription factor FabHLH3 to regulate anthocyanin homeostasis during strawberry fruit ripening
Open the record for dataset details and reuse information.
A dominant-negative SOX18 mutant disrupts multiple regulatory layers essential to transcription factor activity
Open the record for dataset details and reuse information.
Natural variation in a cortex/epidermis-specific transcription factor bZIP89 determines lateral root development and drought resilience in maize
Open the record for dataset details and reuse information.
Data from: Essential function of transmembrane transcription factor MYRF in promoting transcription of miRNA lin-4 during C. elegans development
Open the record for dataset details and reuse information.
Trihelix transcription factor SlGT31 regulates fruit ripening mediated by ethylene in tomato
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.