Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

11 results for “DBLP”

Learn how ShareScore rates datasets ↗
zenodo44/100

DBLP

<p>DBLPdataset is a set of research papers. This dataset is composed of 38,12$ papers in the computer science field. Each paper is classified into one of the following knowledge subareas: computer vision, computational linguistics, biomedical engineering, software engineering, graphics, data mining, security and cryptography, signal processing, robotics, and theory. We had removed the venue for prevent lack of information about the subarea class.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_&lt;k&gt;.pkl:&nbsp;&nbsp;pandas DataFrame with k-cross validation partition.</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

coauth-DBLP

<h3><strong>Overview</strong></h3> <p>This is a temporal higher-order network dataset, which here means a sequence of timestamped hyperedges where each hyperedge is a set of nodes. In this dataset, nodes are authors, and a hyperedge is a publication recorded on DBLP. Timestamps are the year of publication. Some basic statistics of this dataset are:</p> <ul> <li>number of nodes: 1,924,991</li> <li>number of timestamped hyperedges: 3,700,067</li> <li>number of unique hyperedges: 2,599,087</li> </ul> <p><strong>Changelog</strong></p> <ul> <li>v0.1: fixed year format with PR #32 https://github.com/xgi-org/xgi-data/pull/32</li> <li>v0: initial version</li> </ul> <p><strong>Source of original data</strong></p> <ul> <li> <p><a href="https://www.cs.cornell.edu/~arb/data/coauth-DBLP/">https://www.cs.cornell.edu/~arb/data/coauth-DBLP/</a></p> </li> </ul> <h4><strong>References</strong></h4> <p>If you use this data, please cite the following paper:</p> <ul> <li><a href="https://doi.org/10.1073/pnas.1800683115">Simplicial closure and higher-order link prediction</a>. Austin R. Benson, Rediet Abebe, Michael T. Schaub, Ali Jadbabaie, and Jon Kleinberg. <em>Proceedings of the National Academy of Sciences (PNAS)</em>, 2018.</li> </ul>

opencc-by-4.0Nov 2023View details →
zenodo40/100

OpenCSMap dataset: DBLP articles with affiliations

<p>Dataset contains DBLP articles (journal and conference papers) along with one affiliation found for some of them.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

DBLP-QuAD

<p>In this work we create a question answering dataset over the DBLP scholarly knowledge graph (KG). DBLP is an on-line reference for bibliographic information on major computer science publications that indexes over 4.4 million publications, published by more than 2.2 million authors. Our dataset consists of 10,000 question answer pairs with the corresponding SPARQL queries which can be executed over the DBLP KG to fetch the correct answer. To the best of our knowledge, this is the first QA dataset for scholarly KGs.</p> <p>The DBLP KG dump used to create this dataset can be found on this link <a href="https://zenodo.org/record/7638511">https://zenodo.org/record/7638511</a></p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

DBLP-QuAD DBLP Dump

<p>This is the RDF dump of DBLP released on August 1, 2022. The DBLP RDF dump is published to allow fair and replicable evaluation of KGQA systems with the <a href="https://zenodo.org/record/7554379">DBLP-QuAD dataset</a>.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

DBLP Publications Network

<p>This dataset contains information about academic articles, their authors and venues of publication. The dataset has the form of a graph. It has been produced by the SmartDataLake project (<a href="https://smartdatalake.eu">https://smartdatalake.eu</a>), using data collected from Aminer (<a href="https://aminer.org">https://aminer.org</a>).</p>

opencc-by-4.0Jun 2019View details →
zenodo32/100

DBLP-coauthorship

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

Open citations involving Computer Science publications listed in DBLP

<p>Data used in the presentation &quot;Open citations in Informatics&quot; held during ECSS 2021. It includes six different files obtained using the software available at <a href="https://github.com/essepuntato/ecss-2021">https://github.com/essepuntato/ecss-2021</a>.</p>

opencc-zeroOct 2021View details →
zenodo28/100

Two Test Collections for the Author Name Disambiguation Problem based on DBLP

<p>This data set contains&nbsp;two test collections for the author name disambiguation problem. Both collections harness the manual correction work which has been invested into the DBLP (https://dblp.org)&nbsp;collection for computer science literature. The collections are published under the Open Data Commons Attribution License (ODC-By) v1.0.</p> <p>Version 2.0 contains more recent historical data (June 1999 - March 2018) than the previous version (June 1999 - October 2015).&nbsp;</p>

openodc-byMar 2018View details →
zenodo24/100

kgbench: dblp

<p>Graph neural networks and other machine learning models offer a promising direction for interpretable machine learning on relational and multimodal data. Until now, however, progress in this area is difficult to gauge. This is primarily due to a limited number of datasets with (a) a high enough number of labeled nodes in the test set for precise measurement of performance, and (b) a rich enough variety of of multimodal information to learn from. Here, we introduce a set of new benchmark tasks for node classification on knowledge graphs. We focus primarily on node classification, since this setting cannot be solved purely by node embedding models, instead requiring the model to pool information from several steps away in the graph. However, the datasets may also be used for link prediction. For each dataset, we provide test and validation sets of at least 1000 instances, with some containing more than 10\;000 instances. Each task can be performed in a purely relational manner, to evaluate the performance of a relational graph model in isolation, or with multimodal information, to evaluate the performance of multimodal relational graph models. All datasets are packaged in a CSV format that is easily consumable in any machine learning environment, together with the original source data in RDF and pre-processing code for full provenance. We provide code for loading the data into \texttt{numpy} and \texttt{pytorch}. We compute performance for several baseline models.</p>

opencc-pddcDec 2020View details →
zenodo20/100

hdblp: historical data of the dblp collection

<p>This data set contains historical data of the dblp collection (https://dblp.org). I.e., for each metadata record, the collection contains all known revisions that existed in dblp.&nbsp;</p> <p>With hdblp, the state of dblp can be restored for each day between June 2 1999 and December 17 2018. hdblp can be used to study the development of the dblp collection.</p> <p>Correction of version 2 which omitted the data set description.&nbsp;</p>

openodc-byApr 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record