Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “DBLP”
DBLP
<p>DBLPdataset is a set of research papers. This dataset is composed of 38,12$ papers in the computer science field. Each paper is classified into one of the following knowledge subareas: computer vision, computational linguistics, biomedical engineering, software engineering, graphics, data mining, security and cryptography, signal processing, robotics, and theory. We had removed the venue for prevent lack of information about the subarea class.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_<k>.pkl: pandas DataFrame with k-cross validation partition.</p>
coauth-DBLP
<h3><strong>Overview</strong></h3> <p>This is a temporal higher-order network dataset, which here means a sequence of timestamped hyperedges where each hyperedge is a set of nodes. In this dataset, nodes are authors, and a hyperedge is a publication recorded on DBLP. Timestamps are the year of publication. Some basic statistics of this dataset are:</p> <ul> <li>number of nodes: 1,924,991</li> <li>number of timestamped hyperedges: 3,700,067</li> <li>number of unique hyperedges: 2,599,087</li> </ul> <p><strong>Changelog</strong></p> <ul> <li>v0.1: fixed year format with PR #32 https://github.com/xgi-org/xgi-data/pull/32</li> <li>v0: initial version</li> </ul> <p><strong>Source of original data</strong></p> <ul> <li> <p><a href="https://www.cs.cornell.edu/~arb/data/coauth-DBLP/">https://www.cs.cornell.edu/~arb/data/coauth-DBLP/</a></p> </li> </ul> <h4><strong>References</strong></h4> <p>If you use this data, please cite the following paper:</p> <ul> <li><a href="https://doi.org/10.1073/pnas.1800683115">Simplicial closure and higher-order link prediction</a>. Austin R. Benson, Rediet Abebe, Michael T. Schaub, Ali Jadbabaie, and Jon Kleinberg. <em>Proceedings of the National Academy of Sciences (PNAS)</em>, 2018.</li> </ul>
OpenCSMap dataset: DBLP articles with affiliations
<p>Dataset contains DBLP articles (journal and conference papers) along with one affiliation found for some of them.</p>
DBLP-QuAD
<p>In this work we create a question answering dataset over the DBLP scholarly knowledge graph (KG). DBLP is an on-line reference for bibliographic information on major computer science publications that indexes over 4.4 million publications, published by more than 2.2 million authors. Our dataset consists of 10,000 question answer pairs with the corresponding SPARQL queries which can be executed over the DBLP KG to fetch the correct answer. To the best of our knowledge, this is the first QA dataset for scholarly KGs.</p> <p>The DBLP KG dump used to create this dataset can be found on this link <a href="https://zenodo.org/record/7638511">https://zenodo.org/record/7638511</a></p>
DBLP-QuAD DBLP Dump
<p>This is the RDF dump of DBLP released on August 1, 2022. The DBLP RDF dump is published to allow fair and replicable evaluation of KGQA systems with the <a href="https://zenodo.org/record/7554379">DBLP-QuAD dataset</a>.</p>
DBLP Publications Network
<p>This dataset contains information about academic articles, their authors and venues of publication. The dataset has the form of a graph. It has been produced by the SmartDataLake project (<a href="https://smartdatalake.eu">https://smartdatalake.eu</a>), using data collected from Aminer (<a href="https://aminer.org">https://aminer.org</a>).</p>
DBLP-coauthorship
Open the record for dataset details and reuse information.
Open citations involving Computer Science publications listed in DBLP
<p>Data used in the presentation "Open citations in Informatics" held during ECSS 2021. It includes six different files obtained using the software available at <a href="https://github.com/essepuntato/ecss-2021">https://github.com/essepuntato/ecss-2021</a>.</p>
Two Test Collections for the Author Name Disambiguation Problem based on DBLP
<p>This data set contains two test collections for the author name disambiguation problem. Both collections harness the manual correction work which has been invested into the DBLP (https://dblp.org) collection for computer science literature. The collections are published under the Open Data Commons Attribution License (ODC-By) v1.0.</p> <p>Version 2.0 contains more recent historical data (June 1999 - March 2018) than the previous version (June 1999 - October 2015). </p>
kgbench: dblp
<p>Graph neural networks and other machine learning models offer a promising direction for interpretable machine learning on relational and multimodal data. Until now, however, progress in this area is difficult to gauge. This is primarily due to a limited number of datasets with (a) a high enough number of labeled nodes in the test set for precise measurement of performance, and (b) a rich enough variety of of multimodal information to learn from. Here, we introduce a set of new benchmark tasks for node classification on knowledge graphs. We focus primarily on node classification, since this setting cannot be solved purely by node embedding models, instead requiring the model to pool information from several steps away in the graph. However, the datasets may also be used for link prediction. For each dataset, we provide test and validation sets of at least 1000 instances, with some containing more than 10\;000 instances. Each task can be performed in a purely relational manner, to evaluate the performance of a relational graph model in isolation, or with multimodal information, to evaluate the performance of multimodal relational graph models. All datasets are packaged in a CSV format that is easily consumable in any machine learning environment, together with the original source data in RDF and pre-processing code for full provenance. We provide code for loading the data into \texttt{numpy} and \texttt{pytorch}. We compute performance for several baseline models.</p>
hdblp: historical data of the dblp collection
<p>This data set contains historical data of the dblp collection (https://dblp.org). I.e., for each metadata record, the collection contains all known revisions that existed in dblp. </p> <p>With hdblp, the state of dblp can be restored for each day between June 2 1999 and December 17 2018. hdblp can be used to study the development of the dblp collection.</p> <p>Correction of version 2 which omitted the data set description. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.