Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
145
datasets available to search
ShareScore release 0.9.0
Dataset results
145 results for “DNA-binding”
Predictive modeling of moonlighting DNA-binding proteins
<p>This repository contains the codes used for the prediction of moonlighting proteins in the paper "Predictive modeling of moonlighting DNA binding proteins".</p> <p>The repository is organized as the following:</p> <p>1. The DNA binding protein identifiers and their features that were used to train the models for the prediction of DNA binding Moonlighting proteins.</p> <p>2. Five feature sets were used to create Catboost models that make predictions. The source code for generating predictions based on all the features and predictions based on particular features is supplied. In addition, the source code for generating maximum and average ensemble predictions has been made available. A detailed explanation is given in README file.</p>
RNAseq sequences of the study "Transactive response DNA-binding Protein (TARDBP/TDP-43) regulates early HIV-1 entry and infection" (1/2)
<p>Each pair of FASTQ files corresponds to a specific sample condition:</p> <table> <thead> <tr> <th scope="col">Condition</th> <th scope="col">Sample</th> <th scope="col">FASTQ name R1</th> <th scope="col">FASTQ name R2</th> </tr> </thead> <tbody> <tr> <td>Cneg</td> <td>RNASEQ-AVF1</td> <td>RNASEQ-AVF1_S1_R1_001.fastq.gz</td> <td>RNASEQ-AVF1_S1_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-wt-TDP-43</td> <td>RNASEQ-AVF2</td> <td>RNASEQ-AVF2_S2_R1_001.fastq.gz</td> <td>RNASEQ-AVF2_S2_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-NLS-mut-TDP-43</td> <td>RNASEQ-AVF3</td> <td>RNASEQ-AVF3_S3_R1_001.fastq.gz</td> <td>RNASEQ-AVF3_S3_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF4</td> <td>RNASEQ-AVF4_S4_R1_001.fastq.gz</td> <td>RNASEQ-AVF4_S4_R2_001.fastq.gz</td> </tr> <tr> <td>Scramble</td> <td>RNASEQ-AVF5</td> <td>RNASEQ-AVF5_S5_R1_001.fastq.gz</td> <td>RNASEQ-AVF5_S5_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA A</td> <td>RNASEQ-AVF6</td> <td>RNASEQ-AVF6_S6_R1_001.fastq.gz</td> <td>RNASEQ-AVF6_S6_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA B</td> <td>RNASEQ-AVF7</td> <td>RNASEQ-AVF7_S7_R1_001.fastq.gz</td> <td>RNASEQ-AVF7_S7_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA C</td> <td>RNASEQ-AVF8</td> <td>RNASEQ-AVF8_S8_R1_001.fastq.gz</td> <td>RNASEQ-AVF8_S8_R2_001.fastq.gz</td> </tr> </tbody> </table> <p> </p>
RNAseq sequences of the study "Transactive response DNA-binding Protein (TARDBP/TDP-43) regulates early HIV-1 entry and infection" (2/2)
<p>Each pair of FASTQ files corresponds to a specific sample condition:</p> <table> <thead> <tr> <th scope="col">Condition</th> <th scope="col">Sample</th> <th scope="col">FASTQ name R1</th> <th scope="col">FASTQ name R2</th> </tr> </thead> <tbody> <tr> <td>TDP-43 siRNA D</td> <td>RNASEQ-AVF9</td> <td>RNASEQ-AVF9_S1_R1_001.fastq.gz</td> <td>RNASEQ-AVF9_S1_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF10</td> <td>RNASEQ-AVF10_S2_R1_001.fastq.gz</td> <td>RNASEQ-AVF10_S2_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-wt-TDP-43</td> <td>RNASEQ-AVF11</td> <td>RNASEQ-AVF11_S3_R1_001.fastq.gz</td> <td>RNASEQ-AVF11_S3_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-NLS-mut-TDP-43</td> <td>RNASEQ-AVF12</td> <td>RNASEQ-AVF12_S4_R1_001.fastq.gz</td> <td>RNASEQ-AVF12_S4_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF13</td> <td>RNASEQ-AVF13_S5_R1_001.fastq.gz</td> <td>RNASEQ-AVF13_S5_R2_001.fastq.gz</td> </tr> <tr> <td>Scramble</td> <td>RNASEQ-AVF14</td> <td>RNASEQ-AVF14_S6_R1_001.fastq.gz</td> <td>RNASEQ-AVF14_S6_R2_001.fastq.gz</td> </tr> <tr> <td>Oligos B+C</td> <td>RNASEQ-AVF15</td> <td>RNASEQ-AVF15_S7_R1_001.fastq.gz</td> <td>RNASEQ-AVF15_S7_R2_001.fastq.gz</td> </tr> <tr> <td>Oligos A+B+C</td> <td>RNASEQ-AVF16</td> <td>RNASEQ-AVF16_S8_R1_001.fastq.gz</td> <td>RNASEQ-AVF16_S8_R2_001.fastq.gz</td> </tr> </tbody> </table> <p> </p>
Data Sets and Results for "Improved data sets and evaluation methods for the automatic prediction of DNA-binding proteins"
<p>Data sets and results for "Improved data sets and evaluation methods for the automatic prediction of DNA-binding proteins"<br> <br> The file "dna_binding_protein_sequences.zip" has the training and testing sets from the paper:<br> <br> RLL - "random_<train/test>_full_1000.csv"<br> RSL - "random_<train/test>_40.csv"<br> RS&LL - "random_<train/test>_40_1000.csv"<br> RLL where included positive examples have verified DNA binding activity - "random_<train/test>_hq_1000.csv"<br> The 10 RS&LL data sets - "random_<train/test>_40_1000.csv" + "random_<train/test>_40_1000_cv_<0-8>.csv"</p> <p>The results files are named similarly. See "see_results.ipynb" in the <a href="https://doi.org/10.5281/zenodo.5153683">codebase</a> that supplement these data sets<br> The species data sets are derived from "uniprot_data_bac.tab" and "uniprot_data_not_bac.tab." See <a href="https://doi.org/10.5281/zenodo.5153683">code</a>. <br> <br> The ESM embeddings used by the XGBoost model are in "dna_binding_protein_esm.zip"</p>
A Conserved Inhibitory Interdomain Interaction Regulates DNA-binding Activities of Hybrid Two-component Systems in Bacteroides
<p>The study reveals a highly conserved inhibitory mechanism to regulate the activities of hybrid two-component systems (HTCSs) in <em>Bacteroides</em>. HTCSs comprise a major class of transcription regulators of polysaccharide utilization genes in <em>Bacteroides</em>. A conserved sequence motif has been discovered to correlate with the interdomain arrangement of HTCS domains. Presence or absence of this motif is likely predictive of the regulatory mechanism evolved for utilization of different glycans.</p> <p> </p> <p>This dataset includes sequence analyses and structure predictions of HTCSs.</p> <p>List of files:</p> <p>· AlphaFold-HTCS-RR.zip AlphaFold results of all HTCS-RR fragments in <em>B. theta</em></p> <p>· AlphaFold-HTCS-cyto-dimer.zip AlphaFold results of all HTCS-cyto dimers in <em>B. theta</em></p> <p>· HTCSbacteroides-MAFFT-fasta Sequence alignment of 6908 HTCSs from <em>Bacteroides</em></p> <p>· HMM-AllHTCS.hmm HMM of HTCSs generated from the MAFFT alignment</p> <p>· HMM-DBD-PF12833 HMM of HTH18 (Pfam: PF12833) from Interpro</p> <p>· HMM-REC-PF00072 HMM of REC (Pfam: PF00072) from Interpro</p> <p>· B_theta_RR_fasta Sequence alignment of 32 HTCS-RRs in <em>B. theta</em></p> <p>· B_theta_RR-tree A neighbor-joining phylogenetic tree of 32 HTCS-RRs in <em>B. theta</em> </p>
Expression and Purification of DNA-binding domain of T-box Transcription Factor TBXTA-c021
<p>A detailed protocol for the expression and purification of G177D variant of TBXT DNA-binding domain.</p>
X-ray scattering datasets and simulations associated with the publication "Bio-SAXS of Single-Stranded DNA-Binding Proteins: Radiation Protection by the Compatible Solute Ectoine"
<p>This dataset contains the processed and analysed small-angle X-ray scattering data associated with all samples from the publications "Bio-SAXS of Single-Stranded DNA-Binding Proteins: Radiation Protection by the Compatible Solute Ectoine" (<a href="https://doi.org/10.1039/D2CP05053F">https://doi.org/10.1039/D2CP05053F</a>)</p> <p>Files associated with McSAS3 analyses are included, alongside the relevant SAXS data, with datasets labelled in accordance to the protein (G5P), its concentration (1, 2 or 4 mg/mL), and if Ectoine is present (Ect) or absent (Pure). PEPSIsaxs simulations of the GVP monomer (PDB structure: 1GV5 ) and dimer are also included.[1]</p> <p>TOPAS-bioSAXS-dosimetry extension for TOPAS-nBio based particle scattering simulations can be obtained from <a href="https://github.com/MarcBHahn/TOPAS-bioSAXS-dosimetry">https://github.com/MarcBHahn/TOPAS-bioSAXS-dosimetry</a> which is further described in <a href="https://doi.org/10.26272/opus4-55751">https://doi.org/10.26272/opus4-55751</a>.</p> <p>This work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under grant number 442240902 (HA 8528/2-1 and SE 2999/2-1). We acknowledge Diamond Light Source for time on Beamline B21 under Proposal SM29806. This work has been supported by iNEXT-Discovery, grant number 871037, funded by the Horizon 2020 program of the European Commission.</p> <p>[1] S. Su, Y.-G. Gao, H. Zhang, T. C. Terwilliger and A. H.-J. Wang, Protein Science, 1997, 6, 771–780.</p> <p> </p> <p><br> </p>
TIRF Microscopy Data Files. "Multiple RNA- and DNA-binding proteins exhibit direct transfer of polynucleotides: Implications for target site search"
<p>TIRF-microscopy images and analysis data from single-molecule experiments assessing the direct transfer phenomenon in the TREX1 exonuclease. </p>
Data from: Topological DNA-binding of SMC-like RecN promotes RecA-mediated DNA double-strand break repair
<p>Bacterial RecN, closely related to the structural maintenance of chromosomes (SMC) family of proteins, functions in the repair of DNA double-strand breaks (DSBs) by homologous recombination. However, the understanding of how RecN acts in concert with the RecA recombinase to promote DSB repair remains limited. Here, we demonstrated that the purified Escherichia coli RecN protein topologically loads onto both single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA) that has a preference for ssDNA. RecN topologically bound to dsDNA slides off the end of linear dsDNA, but this is prevented by RecA nucleoprotein filaments on ssDNA, thereby allowing RecN to translocate to DSBs. Furthermore, we found that, once RecN is recruited onto ssDNA, it can topologically capture a second dsDNA substrate in an ATP-dependent manner, suggesting a role in synapsis. Indeed, RecN stimulates RecA-mediated D-loop formation and subsequent strand exchange activities. Our findings provide mechanistic insights into the recruitment of RecN to DSBs and sister chromatid interactions by RecN, both of which function in RecA-mediated DSB repair.</p>
Structural adaptation of the single-stranded DNA-binding protein C-terminal to DNA metabolizing partners guides in-hibitor design
<p> </p> <ul> <li>Fluorescence anisotropy metadata and polarization values </li> <li>Isotermal titration calorimetry raw files (protein-ligand and background titrations)</li> <li>Input files for PRALINE calculation in FASTA format</li> <li>CSV file containing the SSB sequences in SMILES format and binding data</li> <li>The input files for the ExoI-bound wtSSB-Ct, E1-sSSB-Ct, and E2-sSSB-Ct and RecO-bound wtSSB-Ct, R1-sSSB-Ct, and R2-sSSB-Ct together with the corresponding trajectory files </li> </ul>
Dataset associated with the study titled "Systematic discovery of regulatory motifs associated with human insulator sites" It includes data used for predicting insulator-associated DNA-binding proteins.
<p><strong>This dataset was used in the study on the prediction of insulator-associated DNA-binding proteins.</strong></p> <p>The three files included here are intended for use with the code available in the following GitHub repository:<br><a href="https://github.com/reposit2/insulator" target="_new" rel="noopener">https://github.com/reposit2/insulator</a></p> <p>To use the data, extract all three files and place them in the same directory as the code files.</p>
Data from: Topological DNA-binding of SMC-like RecN promotes RecA-mediated DNA double-strand break repair
Open the record for dataset details and reuse information.
Genome-Wide Specificity of DNA-Binding, Gene Regulation, and Chromatin Remodeling by TALE- and CRISPR/Cas9-Based Transcription Factors
GEO Series GSE68341. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.
Locus-specific paramutation in Zea mays is maintained by a chromodomain helicase DNA-binding 3 protein controlling development and male gametophyte function
GEO Series GSE158990. Zea mays. 5 samples. Type: Non-coding RNA profiling by high throughput sequencing.
A framework to validate fluorescently labeled DNA-binding proteins for single-molecule experiments
GEO Series GSE212751. Bacillus subtilis PY79. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
High-throughput data and modeling reveal insights into the mechanisms of cooperative DNA-binding by transcription factor proteins
GEO Series GSE171735. Homo sapiens; synthetic construct. 5 samples. Type: Other.
The Compendium of DNA-Binding Specificities of Transcription Factors in a Pathogenic Bacterium
GEO Series GSE146697. Pseudomonas savastanoi pv. phaseolicola 1448A. 513 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Other.
ETHYLENE RESPONSE DNA-BINDING FACTORs are transcriptional repressors responsible for hormone cross-regulation during the ethylene response
GEO Series GSE182617. Arabidopsis thaliana. 116 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.
In vivo probing of the DNA-binding Architecture by bacterial arginine repressor
GEO Series GSE60546. Escherichia coli str. K-12 substr. MG1655. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
High-resolution Brachyury DNA-binding Sites Mapping by ChIP-exo
GEO Series GSE54963. Mus musculus. 1 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.