Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,276
datasets available to search
ShareScore release 0.9.0
Dataset results
4,276 results for “transcription factors”
Fig 1 in Docosahexaenoic Acid (DHA) Reduces LPSInduced Inflammatory Response Via ATF3 Transcription Factor and Stimulates Src/ Syk Signaling-Dependent Phagocytosis in Microglia
<p>Viability of microglia incubated with various concentration of DHA for 12 h (A) and LPS for 2.5 h (B). Viability of microglia incubated with 20 μM DHA followed by 10 ng/ ml LPS treatment (C).</p>
Transfer learning and DNA language models enhance transcription factor binding predictions
<p>This is the dataset for replicating the results of the paper called "Transfer learning and DNA language models enhance transcription factor binding predictions" by Ekin Deniz Aksu and Martin Vingron.</p> <p>See https://github.com/ekinda/tfbs_prediction_paper</p>
Flow Cytometry data from: "The EMT transcription factor Zeb1 is essential for HSPC differentiation that acts synergistically with Zeb2 in fine-tuning hematopoietic lineage fidelity"
<p>Abstract:</p> <p>The Zeb2 transcription factor has been demonstrated to play important roles in hematopoiesis and leukemic transformation. Zeb1 is a close family member of Zeb2 but has remained more enigmatic concerning its roles in hematopoiesis. Here we show using conditional loss of function approaches and bone marrow reconstitution experiments that Zeb1 plays cell autonomous role in hematopoietic lineage differentiation, particularly as a positive regulator of monocyte development in addition to its previously reported important role in T-cell differentiation. Analysis of existing single cell RNAseq data of early hematopoiesis has revealed distinctive expression differences between Zeb1 and Zeb2 in HSPC differentiation with Zeb2 being more highly and broadly expressed that Zeb1 except at a key transition point (ST-HSCàMPP1) whereby Zeb1 appears to be the dominantly expressed family member. Inducible deletion of both Zeb1 and Zeb2 using a tamoxifen inducible Cre-mediated approach leads to acute bone marrow failure at this transition point with increased long-term and shortterm hematopoietic stem cell numbers and an accompanying decrease in all hematopoietic lineage differentiation. Bioinformatics analysis of RNAseq data has revealed that Zeb2 acts predominantly as a transcriptional repressor involved in restraining mature hematopoietic lineage gene expression programs from being expressed too early in hematopoietic stem and progenitor cells (HSPCs). Zeb1 appears to fine tune this repressive role during hematopoiesis to ensure hematopoietic lineage fidelity. Analysis of ROSA26 locus based transgenic models has revealed that Zeb1 as well as Zeb2 overexpression within the hematopoietic system can drive extramedullary hematopoiesis/splenomegaly and enhanced monocyte development. Finally, deletion of Zeb2 alone or Zeb1/2 together was found to enhance survival in secondary MLL-AF9 AML models attesting to the oncogenic role of Zeb1/2 in AML.</p> <p> </p> <p>Flow cytometric and Hematocrit analysis methods: </p> <p><br> Cells were stained with antibodies listed in the provided Supplemental Table (Antibodies.xlsx) according to the manufacturer guidelines. Flow cytometric analyses were performed on the LSRII and Fortessa X-20 cytometer (BD Biosciences) and the results were analysed by FACSDiva or FlowJo software (BD Biosciences). Cells for MLL-AF9 experiments and RNA-seq were stained and sorted on Influx or FACSAria Fusion sorters (BD Biosciences) at AMREP Flow Cytometry Core Facility and FlowCore, Monash University. <br> Submandibular blood samples were collected into EDTA-coated tubes, and hematology parameters were measured using a HemaVet 950FS automated blood analysis machine (Drew Scientific).</p>
Supplementary material for: RSAT variation-tools: An accessible and flexible framework to predict the impact of regulatory variants on transcription factor binding
<p>Supplementary Material for the Article</p> <p>Santana-Garcia, W., Rocha-Acevedo, M., Ramirez-Navarro, L., Mbouamboua, Y., Thieffry, D., Thomas-Chollier, M., Contreras-Moreira, B., van Helden, J., Medina-Rivera, A., 2019. RSAT variation-tools: An accessible and flexible framework to predict the impact of regulatory variants on transcription factor binding. Comput. Struct. Biotechnol. J. 17, 1415–1428.</p> <p> </p>
Cooperative action of separate interaction domains promotes high-affinity DNA binding of Arabidopsis thaliana ARF transcription factors
<p>The repository contains the smFRET and SAXS data presented in the preprint https://doi.org/10.1101/2022.11.16.516730 (BioRxiv)</p>
Mammalian Evolution of Human cis-regulatory Elements and Transcription Factor Binding Sites
<p>Code and data associated with the manuscript entitled "Mammalian Evolution of Human cis-regulatory Elements and Transcription Factor Binding Sites "</p>
Dataset supporting the paper "Expanding the coverage of regulons from high-confidence prior knowledge for accurate estimation of transcription factor activities"
<p>Datasets involved in the construction and benchmarking of the CollecTRI-derived regulons as presented in the paper "Expanding the coverage of regulons from high-confidence prior knowledge for accurate estimation of transcription factor activities".</p>
Analysis Products: Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency
<p>This record contains analysis products for the paper "Transcription factor stoichiometry, motif affinity and syntax regulate single cell chromatin dynamics during fibroblast reprogramming to pluripotency" by Nair, Ameen <em>et al</em>. Please refer to the READMEs in the directories, which are summarized below.</p> <p>The record contains the following files:<br> <br> `clusters.tsv`: <strong> </strong>contains the cluster id, name and colour of clusters in the paper</p> <p><strong>scATAC.zip</strong></p> <p>Analysis products for the single-cell ATAC-seq data. Contains:</p> <p>- `cells.tsv`: list of barcodes that pass QC. Columns include:<br> - `barcode`<br> - `sample`: (time point)<br> - `umap1`<br> - `umap2`<br> - `cluster`<br> - `dpt_pseudotime_fibr_root`: pseudotime values treating a fibroblast cell as root<br> - `dpt_pseudotime_xOSK_root`: pseudotime values treating xOSK cell as root<br> - `peaks.bed`: list of peaks of 500bp across all cell states. 4th column contains the peak set label. Note that ~5000 peaks are not assigned to any peak set and are marked as NA.<br> - `features.tsv`: 50 dimensional representation of each cell <br> - `cell_x_peak.mtx.gz`: sparse matrix of fragment counts within peaks. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (combine sample + barcode). Rows correspond to peaks in `peaks.bed` </p> <p><strong>scATAC_clusters.zip</strong></p> <p>Analysis products corresponding to cluster pseudo-bulks of the single-cell ATAC-seq data. </p> <p>- `clusters.tsv`: contains the cluster id, name and colour used in the paper<br> - `peaks`: contains `overlap_reproducibilty/overlap.optimal_peak` peaks called using ENCODE bulk ATAC-seq pipeline in the narrowPeak format.<br> - `fragments`: contains per cluster fragment files </p> <p><strong>scATAC_scRNA_integration.zip</strong></p> <p>Analysis products from the integration of scATAC with scRNA. Contains:</p> <p>- `peak_gene_links_fdr1e-4.tsv`: file with peak gene links passing FDR 1e-4. For analyses in the paper, we filter to peaks with absolute correlation >0.45.<br> - `harmony.cca.30.feat.tsv`: 30 dimensional co-embedding for scATAC and scRNA cells obtained by CCA followed by applying Harmony over assay type.<br> - `harmony.cca.metadata.tsv`: UMAP coordinates for scATAC and scRNA cells derived from the Harmony CCA embedding. First column contains barcode.</p> <p><strong>scRNA.zip</strong></p> <p>Analysis products for the single-cell RNA-seq data. Contains:</p> <p>- `seurat.rds`: seurat object that contains expression data (raw counts, normalized, and scaled), reductions (umap, pca), knn graphs, all associated metadata. Note that barcode suffix (1-9 corresponds to samples D0, D2, ..., D14, iPSC)<br> - `genes.txt`: list of all genes<br> - `cells.tsv`: list of barcodes that pass QC across samples. Contains:<br> - `barcode_sample`: barcode with index of sample (1-9 corresponding to D0, D2, ..., D14, iPSC) <br> - `sample`: sample name (D0, D2, .., D14, iPSC)<br> - `umap1`<br> - `umap2`<br> - `nCount_RNA`<br> - `nFeature_RNA`<br> - `cluster`<br> - `percent.mt`: percent of mitochondrial transcripts in cell<br> - `percent.oskm`: percent of OSKM transcripts in cell<br> - `gene_x_cell.mtx.gz`: sparse matrix of gene counts. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (barcode suffix contains sample information). Rows correspond to genes in `genes.txt` <br> - `pca.tsv`: first 50 PC of each cell<br> - `oskm_endo_sendai.tsv`: estimated raw counts (cts, may not be integers) and log(1+ tp10k) normalized expression (norm) for endogenous and exogenous (Sendai derived) counts of POU5F1 (OCT4), SOX2, KLF4 and MYC genes. Rows are consistent with `seurat.rds` and `cells.tsv`</p> <p><strong>multiome.zip</strong></p> <p><em>multiome/snATAC:</em></p> <p>These files are derived from the integration of nuclei from multiome (D1M and D2M), with cells from day 2 of scATAC-seq (labeled D2). </p> <p>- `cells.tsv`: This is the list of nuclei barcodes that pass QC from multiome AND also cell barcodes from D2 of scATAC-seq. Includes:<br> - `barcode`<br> - `umap1`: These are the coordinates used for the figures involving multiome in the paper.<br> - `umap2`: ^^^ <br> - `sample`: D1M and D2M correspond to multiome, D2 corresponds to day 2 of scATAC-seq<br> - `cluster`: For multiome barcodes, these are labels transfered from scATAC-seq. For D2 scATAC-seq, it is the original cluster labels. <br> - `peaks.bed`: This is the same file as scATAC/peaks.bed. List of peaks of 500bp. 4th column contains the peak set label. Note that ~5000 peaks are not assigned to any peak set and are marked as NA.<br> - `cell_x_peak.mtx.gz`: sparse matrix of fragment counts within peaks. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (combine sample + barcode). Rows correspond to peaks in `peaks.bed`.<br> - `features.no.harmony.50d.tsv`: 50 dimensional representation of each cell prior to running Harmony (to correct for batch effect between D2 scATAC and D1M,D2M snMultiome). Rows correspond to cells from `cells.tsv`.<br> - `features.harmony.10d.tsv`: 10 dimensional representation of each cell after running Harmony. Rows correspond to cells from `cells.tsv`.</p> <p><em>multiome/snRNA:</em></p> <p>- `seurat.rds`: seurat object that contains expression data (raw counts, normalized, and scaled), reductions (umap, pca),associated metadata. Note that barcode suffix (1,2 corresponds to samples D1M, D2M). Please use the UMAP/features from snATAC/ for consistency.<br> - `genes.txt`: list of all genes (this is different from the list in scRNA analysis)<br> - `cells.tsv`: list of barcodes that pass QC across samples. Contains:<br> - `barcode_sample`: barcode with index of sample (1,2 corresponding to D1M, D2M respectively) <br> - `sample`: sample name (D1M, D2M)<br> - `nCount_RNA`<br> - `nFeature_RNA`<br> - `percent.oskm`: percent of OSKM genes in cell<br> - `gene_x_cell.mtx.gz`: sparse matrix of gene counts. Load using scipy.io.mmread in python or readMM in R. Columns correspond to cells from `cells.tsv` (barcode suffix contains sample information). Rows correspond to genes in `genes.txt` </p>
Development and validation of a novel plasmid chassis system for screening of metabolite-responsive transcription factors
<p>This dataset contains the raw data that lie at the basis of the results discussed in <strong>Chapter 3: Development and validation of a novel plasmid chassis system for screening of metabolite-responsive transcription factors </strong>of the PhD thesis of Amber Bernauw. The README.txt file provides more information on the different data files.</p>
In vivo screening of Lrp-type transcription factors in Escherichia coli
<p>This dataset contains the raw data that lie at the basis of the results discussed in <strong>Chapter 4: <em>In vivo</em> screening of Lrp-type transcription factors in <em>Escherichia coli</em></strong><strong> </strong>of the PhD thesis of Amber Bernauw. The README.txt file provides more information on the different data files.</p>
Image stacks for full-body transcription factor expression atlas with completely resolved cell identities in C. elegans
<p>Each image stack presented as zip file. Once decompressed, each folder contain '.ano' linker file, straightening C. elegans L1 images file, the segmentation mask image file and the cell annotation file. The image files are stored in Peng Hanchuan RAW/TIFF format, and the cell annotation file is stored in simple comma separated values format. To visualize the image stack data, drag the '.ano' linker file to VANO interface. </p> <p>vano_win32_1.741.zip contains VANO for worm visualization.</p>
Supplementary data accompanying HOCOMOCO v12 collection of transcription factor binding motifs
<p><strong>*** SUMMARY ***</strong></p><p>This dataset contains supplementary data accompanying HOCOMOCO v12 collection</p><p>of DNA binding motifs for human and mouse transcription factors, https://hocomoco.autosome.org</p><p> </p><p>The contents include:</p><p>- the complete initial set of motifs discovered from ChIP-Seq and HT-SELEX data;</p><p>- the curated subset of motifs associated with distinct motif subtypes that were used in benchmarking;</p><p>- the benchmarking results and the resulting final motif collections, including motif logos;</p><p>- accompanying metadata.</p><p> </p><p>Please refer to the README and the HOCOMOCO website for further details.</p><p> </p>
DoubleChEC program to identify transcription factor binding sites from mapped ChEC-seq data
Open the record for dataset details and reuse information.
Quantitative modulation of a spatial enhancer through the biophysical properties of a transcription factor binding site
Open the record for dataset details and reuse information.
Data for the paper "Insights gained from a comprehensive all-against-all transcription factor binding motif benchmarking study".
<p>Data for the paper "Insights gained from a comprehensive all-against-all transcription factor binding motif benchmarking study".</p>
Supporting data for "Dynamics of RNA polymerase II and elongation factor Spt4/5 recruitment during activator-dependent transcription"
<p>Supporting data for</p> <p><strong>Dynamics of RNA polymerase II and elongation factor Spt4/5 recruitment</strong></p> <p><strong>during activator-dependent transcription </strong></p> <p>Grace A. Rosen<sup>a,1</sup>, Inwha Baek<sup>b,1</sup>, Larry J. Friedman<sup>a</sup>, Yoo Jin Joo<sup>b</sup>, Stephen Buratowski<sup>b,2</sup>, Jeff Gelles<sup>a,2</sup></p> <p><sup>a</sup>Department of Biochemistry, Brandeis University, Waltham, Massachusetts 02454, USA.</p> <p><sup>b</sup>Department of Biological Chemistry and Molecular Pharmacology, Harvard Medical School, Boston, Massachusetts 02115, USA.</p> <p><sup>1</sup>Equal contributions</p> <p><sup>2</sup>Corresponding authors: <a href="mailto:steveb@hms.harvard.edu">steveb@hms.harvard.edu</a>; +1 (617) 432-0696 (S.B.) and <a href="mailto:gelles@brandeis.edu">gelles@brandeis.edu</a>; +1 (781) 736-2377 (J.G.)</p> <p>See <strong>Source data index.pdf</strong> for description of files.</p>
LERC: Linked Extended Regulatory Circuits dataset on interactions between transcription factors and genes
<p>Turtles files of the project: <a href="https://regulatorycircuits-lod.genouest.org/">https://regulatorycircuits-lod.genouest.org/</a><br> 808 samples-specific files<br> 394 tissues-specific files<br> 2 mapping (FMA and Uberon)<br> 1 graph of experimental context<br> 1 graph of metadata</p> <p> </p>
Data from: A clinal polymorphism in the insulin signaling transcription factor foxo contributes to life-history adaptation in Drosophila
A fundamental aim of adaptation genomics is to identify polymorphisms that underpin variation in fitness traits. In D. melanogaster latitudinal life-history clines exist on multiple continents and make an excellent system for dissecting the genetics of adaptation. We have previously identified numerous clinal SNPs in insulin/insulin-like growth factor signaling (IIS), a pathway known from mutant studies to affect life history. However, the effects of natural variants in this pathway remain poorly understood. Here we investigate how two clinal alternative alleles at foxo, a transcriptional effector of IIS, affect fitness components (viability, size, starvation resistance, fat content). We assessed this polymorphism from the North American cline by reconstituting outbred populations, fixed for either the low- or high-latitude allele, from inbred DGRP lines. Since diet and temperature modulate IIS, we phenotyped alleles across two temperatures (18°C, 25°C) and two diets differing in sugar source and content. Consistent with clinal expectations, the high-latitude allele conferred larger body size and reduced wing loading. Alleles also differed in starvation resistance and expression of InR, a transcriptional target of FOXO. Allelic reaction norms were mostly parallel, with few GxE interactions. Together, our results suggest that variation in IIS makes a major contribution to clinal life-history adaptation.
Molecular dynamics trajectories, GROMACS input files, and analysis code from "Rational optimization of a transcription factor activation domain inhibitor" by Basu et. al, Nature Structural & Molecular Biology, 2023
<p>Molecular dynamics trajectories, GROMACS input files, and analysis code from "Rational optimization of a transcription factor activation domain inhibitor" by Basu et. al, Nature Structural & Molecular Biology, 2023</p> <p> </p> <p> </p>
Xrp1 ChIP-seq database (data from Xrp1 is a transcription factor required for cell competition-driven elimination of loser cells)
<p>Xrp1 ChIP-seq data from Baillon et al. 2018, "Xrp1 is a transcription factor required for cell competition-driven elimination of loser cells".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.