Benchmarking Recent Computational Tools Available for DNA Binding Protein Identification
<p><strong>Train_Test_PSSM.pickle:</strong></p> <p>We have provided a train and a test fasta file containing DNA binding and non-binding protein sequences. The PSSM matrices for these sequences are available inside this .pickle file. You will need to provide the path of this .pickle file when running predictions on the provided train or test fasta file sequences using the two provided tools (LocalDPP and LSTM-CNN_Fusion)</p> <p><strong>db.zip:</strong></p> <p>You will need to unzip this database file and provide the unzipped folder as path to the PSSM generation method. The PSSM creation method provided in github uses this database to generate the .pickle file (full of PSSM matrices) from user provided fasta file sequences.</p>
ShareScore
24/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 0