Skip to main content
zenodoopen

DBLP

<p>DBLPdataset is a set of research papers. This dataset is composed of 38,12$ papers in the computer science field. Each paper is classified into one of the following knowledge subareas: computer vision, computational linguistics, biomedical engineering, software engineering, graphics, data mining, security and cryptography, signal processing, robotics, and theory. We had removed the venue for prevent lack of information about the subarea class.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_&lt;k&gt;.pkl:&nbsp;&nbsp;pandas DataFrame with k-cross validation partition.</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
4
Harmonization
8
Access
20
Reuse readiness
8
Engagement
4