Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

145

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

145 results for “DNA-binding”

Learn how ShareScore rates datasets ↗
zenodo44/100

Predictive modeling of moonlighting DNA-binding proteins

<p>This repository contains the codes used for the prediction of moonlighting proteins&nbsp;in the paper &quot;Predictive modeling of moonlighting DNA binding proteins&quot;.</p> <p>The repository is organized as the following:</p> <p>1. The DNA binding protein identifiers&nbsp;and their features that were used to train the models for the prediction of DNA binding Moonlighting proteins.</p> <p>2. Five feature sets were used to create Catboost models that make predictions. The source code for generating predictions based on all the features and predictions based on particular features is supplied. In addition, the source code for generating maximum and average ensemble predictions has been made available. A detailed explanation is given in README file.</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

RNAseq sequences of the study "Transactive response DNA-binding Protein (TARDBP/TDP-43) regulates early HIV-1 entry and infection" (1/2)

<p>Each pair of FASTQ files corresponds to a specific sample condition:</p> <table> <thead> <tr> <th scope="col">Condition</th> <th scope="col">Sample</th> <th scope="col">FASTQ name R1</th> <th scope="col">FASTQ name R2</th> </tr> </thead> <tbody> <tr> <td>Cneg</td> <td>RNASEQ-AVF1</td> <td>RNASEQ-AVF1_S1_R1_001.fastq.gz</td> <td>RNASEQ-AVF1_S1_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-wt-TDP-43</td> <td>RNASEQ-AVF2</td> <td>RNASEQ-AVF2_S2_R1_001.fastq.gz</td> <td>RNASEQ-AVF2_S2_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-NLS-mut-TDP-43</td> <td>RNASEQ-AVF3</td> <td>RNASEQ-AVF3_S3_R1_001.fastq.gz</td> <td>RNASEQ-AVF3_S3_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF4</td> <td>RNASEQ-AVF4_S4_R1_001.fastq.gz</td> <td>RNASEQ-AVF4_S4_R2_001.fastq.gz</td> </tr> <tr> <td>Scramble</td> <td>RNASEQ-AVF5</td> <td>RNASEQ-AVF5_S5_R1_001.fastq.gz</td> <td>RNASEQ-AVF5_S5_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA A</td> <td>RNASEQ-AVF6</td> <td>RNASEQ-AVF6_S6_R1_001.fastq.gz</td> <td>RNASEQ-AVF6_S6_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA B</td> <td>RNASEQ-AVF7</td> <td>RNASEQ-AVF7_S7_R1_001.fastq.gz</td> <td>RNASEQ-AVF7_S7_R2_001.fastq.gz</td> </tr> <tr> <td>TDP-43 siRNA C</td> <td>RNASEQ-AVF8</td> <td>RNASEQ-AVF8_S8_R1_001.fastq.gz</td> <td>RNASEQ-AVF8_S8_R2_001.fastq.gz</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

RNAseq sequences of the study "Transactive response DNA-binding Protein (TARDBP/TDP-43) regulates early HIV-1 entry and infection" (2/2)

<p>Each pair of FASTQ files corresponds to a specific sample condition:</p> <table> <thead> <tr> <th scope="col">Condition</th> <th scope="col">Sample</th> <th scope="col">FASTQ name R1</th> <th scope="col">FASTQ name R2</th> </tr> </thead> <tbody> <tr> <td>TDP-43 siRNA D</td> <td>RNASEQ-AVF9</td> <td>RNASEQ-AVF9_S1_R1_001.fastq.gz</td> <td>RNASEQ-AVF9_S1_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF10</td> <td>RNASEQ-AVF10_S2_R1_001.fastq.gz</td> <td>RNASEQ-AVF10_S2_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-wt-TDP-43</td> <td>RNASEQ-AVF11</td> <td>RNASEQ-AVF11_S3_R1_001.fastq.gz</td> <td>RNASEQ-AVF11_S3_R2_001.fastq.gz</td> </tr> <tr> <td>Flag-NLS-mut-TDP-43</td> <td>RNASEQ-AVF12</td> <td>RNASEQ-AVF12_S4_R1_001.fastq.gz</td> <td>RNASEQ-AVF12_S4_R2_001.fastq.gz</td> </tr> <tr> <td>Cneg</td> <td>RNASEQ-AVF13</td> <td>RNASEQ-AVF13_S5_R1_001.fastq.gz</td> <td>RNASEQ-AVF13_S5_R2_001.fastq.gz</td> </tr> <tr> <td>Scramble</td> <td>RNASEQ-AVF14</td> <td>RNASEQ-AVF14_S6_R1_001.fastq.gz</td> <td>RNASEQ-AVF14_S6_R2_001.fastq.gz</td> </tr> <tr> <td>Oligos B+C</td> <td>RNASEQ-AVF15</td> <td>RNASEQ-AVF15_S7_R1_001.fastq.gz</td> <td>RNASEQ-AVF15_S7_R2_001.fastq.gz</td> </tr> <tr> <td>Oligos A+B+C</td> <td>RNASEQ-AVF16</td> <td>RNASEQ-AVF16_S8_R1_001.fastq.gz</td> <td>RNASEQ-AVF16_S8_R2_001.fastq.gz</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Data Sets and Results for "Improved data sets and evaluation methods for the automatic prediction of DNA-binding proteins"

<p>Data sets and results for &quot;Improved data sets and evaluation methods for the automatic prediction of DNA-binding proteins&quot;<br> <br> The file &quot;dna_binding_protein_sequences.zip&quot; has the training and testing sets from the paper:<br> <br> RLL - &quot;random_&lt;train/test&gt;_full_1000.csv&quot;<br> RSL - &quot;random_&lt;train/test&gt;_40.csv&quot;<br> RS&amp;LL - &quot;random_&lt;train/test&gt;_40_1000.csv&quot;<br> RLL where included positive examples have verified DNA binding activity -&nbsp;&quot;random_&lt;train/test&gt;_hq_1000.csv&quot;<br> The 10 RS&amp;LL data sets - &quot;random_&lt;train/test&gt;_40_1000.csv&quot; +&nbsp;&quot;random_&lt;train/test&gt;_40_1000_cv_&lt;0-8&gt;.csv&quot;</p> <p>The results files are&nbsp;named similarly. See &quot;see_results.ipynb&quot; in the <a href="https://doi.org/10.5281/zenodo.5153683">codebase</a> that supplement these&nbsp;data sets<br> The species data sets are derived from &quot;uniprot_data_bac.tab&quot; and &quot;uniprot_data_not_bac.tab.&quot;&nbsp; See <a href="https://doi.org/10.5281/zenodo.5153683">code</a>.&nbsp;<br> <br> The ESM embeddings used by the XGBoost model are in &quot;dna_binding_protein_esm.zip&quot;</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

A Conserved Inhibitory Interdomain Interaction Regulates DNA-binding Activities of Hybrid Two-component Systems in Bacteroides

<p>The study reveals a highly conserved inhibitory mechanism to regulate the activities of hybrid two-component systems (HTCSs) in <em>Bacteroides</em>. HTCSs comprise a major class of transcription regulators of polysaccharide utilization genes in <em>Bacteroides</em>. A conserved sequence motif has been discovered to correlate with the interdomain arrangement of HTCS domains. Presence or absence of this motif is likely predictive of the regulatory mechanism evolved for utilization of different glycans.</p> <p>&nbsp;</p> <p>This dataset includes sequence analyses and structure predictions of HTCSs.</p> <p>List of files:</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; AlphaFold-HTCS-RR.zip&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp;&nbsp; AlphaFold results of all HTCS-RR fragments in <em>B. theta</em></p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; AlphaFold-HTCS-cyto-dimer.zip&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; AlphaFold results of all HTCS-cyto dimers in <em>B. theta</em></p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; HTCSbacteroides-MAFFT-fasta&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Sequence alignment of 6908 HTCSs from <em>Bacteroides</em></p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; HMM-AllHTCS.hmm&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; HMM of HTCSs generated from the MAFFT alignment</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; HMM-DBD-PF12833&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; HMM of HTH18 (Pfam: PF12833) from Interpro</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; HMM-REC-PF00072&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; HMM of REC (Pfam: PF00072) from Interpro</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; B_theta_RR_fasta&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Sequence alignment of 32 HTCS-RRs in <em>B. theta</em></p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; B_theta_RR-tree&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A neighbor-joining phylogenetic tree of 32 HTCS-RRs in <em>B. theta</em>&nbsp;&nbsp;&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Expression and Purification of DNA-binding domain of T-box Transcription Factor TBXTA-c021

<p>A detailed protocol for the expression and purification of G177D variant of TBXT DNA-binding domain.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

X-ray scattering datasets and simulations associated with the publication "Bio-SAXS of Single-Stranded DNA-Binding Proteins: Radiation Protection by the Compatible Solute Ectoine"

<p>This dataset contains the processed and analysed small-angle X-ray scattering data associated with all samples from the publications &quot;Bio-SAXS of Single-Stranded DNA-Binding Proteins: Radiation Protection by the Compatible Solute Ectoine&quot; (<a href="https://doi.org/10.1039/D2CP05053F">https://doi.org/10.1039/D2CP05053F</a>)</p> <p>Files associated with McSAS3 analyses are included, alongside the relevant SAXS data,&nbsp;with datasets labelled in accordance to the protein (G5P), its concentration (1, 2 or 4 mg/mL), and if Ectoine is present (Ect) or absent&nbsp;(Pure). PEPSIsaxs simulations of the GVP monomer&nbsp;(PDB structure: 1GV5&nbsp;) and dimer are also included.[1]</p> <p>TOPAS-bioSAXS-dosimetry extension for TOPAS-nBio based particle scattering simulations can be obtained from&nbsp;<a href="https://github.com/MarcBHahn/TOPAS-bioSAXS-dosimetry">https://github.com/MarcBHahn/TOPAS-bioSAXS-dosimetry</a>&nbsp;which is further described in&nbsp;<a href="https://doi.org/10.26272/opus4-55751">https://doi.org/10.26272/opus4-55751</a>.</p> <p>This work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under grant number 442240902 (HA 8528/2-1 and SE 2999/2-1). We acknowledge Diamond Light Source for time on Beamline B21 under Proposal SM29806. This work has been supported by iNEXT-Discovery, grant number 871037, funded by the Horizon 2020 program of the European Commission.</p> <p>[1]&nbsp;S. Su, Y.-G. Gao, H. Zhang, T. C. Terwilliger and A. H.-J. Wang, Protein Science, 1997, 6, 771&ndash;780.</p> <p>&nbsp;</p> <p><br> &nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

TIRF Microscopy Data Files. "Multiple RNA- and DNA-binding proteins exhibit direct transfer of polynucleotides: Implications for target site search"

<p>TIRF-microscopy images and analysis data from single-molecule experiments assessing the direct transfer phenomenon in the TREX1 exonuclease.&nbsp;</p>

opencc-by-4.0Dec 2021View details →
dryad32/100

Data from: Topological DNA-binding of SMC-like RecN promotes RecA-mediated DNA double-strand break repair

<p>Bacterial RecN, closely related to the structural maintenance of chromosomes (SMC) family of proteins, functions in the repair of DNA double-strand breaks (DSBs) by homologous recombination. However, the understanding of how RecN acts in concert with the RecA recombinase to promote DSB repair remains limited. Here, we demonstrated that the purified Escherichia coli RecN protein topologically loads onto both single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA) that has a preference for ssDNA. RecN topologically bound to dsDNA slides off the end of linear dsDNA, but this is prevented by RecA nucleoprotein filaments on ssDNA, thereby allowing RecN to translocate to DSBs. Furthermore, we found that, once RecN is recruited onto ssDNA, it can topologically capture a second dsDNA substrate in an ATP-dependent manner, suggesting a role in synapsis. Indeed, RecN stimulates RecA-mediated D-loop formation and subsequent strand exchange activities. Our findings provide mechanistic insights into the recruitment of RecN to DSBs and sister chromatid interactions by RecN, both of which function in RecA-mediated DSB repair.</p>

opencc-zeroOct 2020View details →
zenodo32/100

Structural adaptation of the single-stranded DNA-binding protein C-terminal to DNA metabolizing partners guides in-hibitor design

<p>&nbsp;</p> <ul> <li>Fluorescence anisotropy metadata and polarization values&nbsp;</li> <li>Isotermal titration calorimetry raw files (protein-ligand and background titrations)</li> <li>Input files for PRALINE calculation in FASTA format</li> <li>CSV file containing the SSB sequences in SMILES format and binding data</li> <li>The input files for the ExoI-bound wtSSB-Ct, E1-sSSB-Ct, and E2-sSSB-Ct and RecO-bound wtSSB-Ct, R1-sSSB-Ct, and R2-sSSB-Ct together with the corresponding trajectory files&nbsp;</li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo32/100

Dataset associated with the study titled "Systematic discovery of regulatory motifs associated with human insulator sites" It includes data used for predicting insulator-associated DNA-binding proteins.

<p><strong>This dataset was used in the study on the prediction of insulator-associated DNA-binding proteins.</strong></p> <p>The three files included here are intended for use with the code available in the following GitHub repository:<br><a href="https://github.com/reposit2/insulator" target="_new" rel="noopener">https://github.com/reposit2/insulator</a></p> <p>To use the data, extract all three files and place them in the same directory as the code files.</p>

opencc-by-4.0Dec 2023View details →
dryad32/100

Data from: Topological DNA-binding of SMC-like RecN promotes RecA-mediated DNA double-strand break repair

Open the record for dataset details and reuse information.

publicSep 2021View details →
geo24/100

Genome-Wide Specificity of DNA-Binding, Gene Regulation, and Chromatin Remodeling by TALE- and CRISPR/Cas9-Based Transcription Factors

GEO Series GSE68341. Homo sapiens. 9 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenMay 2015View details →
geo24/100

Locus-specific paramutation in Zea mays is maintained by a chromodomain helicase DNA-binding 3 protein controlling development and male gametophyte function

GEO Series GSE158990. Zea mays. 5 samples. Type: Non-coding RNA profiling by high throughput sequencing.

openGEO-OpenNov 2020View details →
geo24/100

A framework to validate fluorescently labeled DNA-binding proteins for single-molecule experiments

GEO Series GSE212751. Bacillus subtilis PY79. 8 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenSep 2023View details →
geo24/100

High-throughput data and modeling reveal insights into the mechanisms of cooperative DNA-binding by transcription factor proteins

GEO Series GSE171735. Homo sapiens; synthetic construct. 5 samples. Type: Other.

openGEO-OpenJul 2023View details →
geo24/100

The Compendium of DNA-Binding Specificities of Transcription Factors in a Pathogenic Bacterium

GEO Series GSE146697. Pseudomonas savastanoi pv. phaseolicola 1448A. 513 samples. Type: Genome binding/occupancy profiling by high throughput sequencing; Other.

openGEO-OpenAug 2020View details →
geo24/100

ETHYLENE RESPONSE DNA-BINDING FACTORs are transcriptional repressors responsible for hormone cross-regulation during the ethylene response

GEO Series GSE182617. Arabidopsis thaliana. 116 samples. Type: Expression profiling by high throughput sequencing; Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenApr 2023View details →
geo24/100

In vivo probing of the DNA-binding Architecture by bacterial arginine repressor

GEO Series GSE60546. Escherichia coli str. K-12 substr. MG1655. 2 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenApr 2015View details →
geo24/100

High-resolution Brachyury DNA-binding Sites Mapping by ChIP-exo

GEO Series GSE54963. Mus musculus. 1 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.

openGEO-OpenMar 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record