Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

6 results for “Dataset, Author Name Disambiguation”

Learn how ShareScore rates datasets ↗
zenodo44/100

LSPO: A Large-Scale Physics ORCiD-Linked Dataset for Author Name Disambiguation

<p>The LSPO dataset, a Large-Scale Physics ORCiD-Linked Dataset for Author Name Disambiguation is comprised of 554,962 NASA/ADS publications linked to 125,486 unique researchers through ORCiD identifiers. The available meta-data fields are: ORCiD identifier, author name, affiliation, title, asbtract, and name block. The dataset can be utilized to make pairs or triplets for training a author name disambiguation model.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

CHEN-AND: A Labeled Dataset for Chinese and English Joint Author Name Disambiguation

<p>Abstract: Author name disambiguation (AND) is an important problem in literature databases and is even more prominent in the cross-language (database) context. Extensive research has been conducted in the academic community to eliminate such ambiguity. However, the existing research mainly focuses on monolingual literature, with less attention paid to author disambiguation in cross-language environments. In this regard, this study focuses on a typical cross-language author disambiguation task - Chinese and English joint author disambiguation. We first propose an automated dataset construction method for Chinese-English literature joint AND using online open resources, with this method, we create a dataset named CHEN-AND for joint Chinese and English author disambiguation. Then we propose a merging-first-then-disambiguation (MFTD)--based disambiguation framework and evaluate several variants of this method on the test dataset.&nbsp;For details on building this dataset, please refer to <a href="https://github.com/carmanzhang/CHEN-AND">this repository</a> on my GitHub page.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

LAGOS-AND: A Large Gold Standard Dataset for MAG/OpenAlex Author Name Disambiguation

<p>We present a large gold standard dataset for author name disambiguation (AND) research (LAGOS-AND), which contains two sub-datasets, LAGOS-AND-BLOCK and LAGOS-AND-PAIRWISE. The datasets were automatically built by using the two authoritative sources, ORCID and DOI, based on the ORCID open database and an open literature database (MAG or OpenAlex).</p> <p>The currently available versions of the LAGOS-AND datasets are:</p> <ol> <li><strong>Version 1.0 (MAG+ORCID)</strong>: This is the initial version of the LAGOS-AND dataset, the evaluation results and quality control measures of this version dataset can be found in this paper <a href="https://arxiv.org/abs/2104.01821">https://arxiv.org/abs/2104.01821</a>.</li> <li><strong>Version 2.0 (OpenAlex+ORCID)</strong>: This version builds on OpenAlex instead of MAG because MAG was discontinued on 31 December 2021, and OpenAlex not only positions itself as a drop-in replacement for MAG but also keeps evolving by aggregating academic resources from other repositories. In addition, the pairwise-based sub-dataset (LAGOS-AND-PAIRWISE v2.0) improves the accuracy of labeled authorship (class label) as compared to LAGOS-AND-PAIRWISE v1.0.</li> </ol> <p>Note that there are other versions of the dataset, which we call &quot;pre-release&quot;. We recommend users to use the normal version of the dataset. These pre-release versions were originally intended to be released as the normal versions. However, during the preparation of the research paper, we found an issue with the dataset and the reviewers also made some reasonable requests for the dataset. This led us to update the dataset. Unfortunately, the Zenodo platform does not allow updates to the same version of dataset, so we had to create new versions. The created pre-release datasets are as follows:</p> <ol> <li><strong>Version 2.0-alpha</strong>: For few samples in LAGOS-AND-PAIRWISE, the class labels are incorrect. We improve the accuracy of the class label in the normal Version 2.0 dataset by using a better random sampling approach.</li> <li><strong>Version 1.0-beta</strong>: We created this version because a sub-dataset of this version LAGOS-AND-PAIRWISE contains only ~500K author pairs, while it should contain ~1M author pairs, as described in our paper. We fixed the problem in the Version 1.0 dataset.</li> <li><strong>Version 1.0-alpha</strong>: The earliest dataset uploaded to Zenodo, corresponding to the original dataset before addressing the reviewers&#39; comments and suggestions. In contrast, the dataset in Version 1.0 is the dataset after the reviewers&#39; comments and suggestions have been addressed.</li> </ol>

opencc-by-4.0Feb 2021View details →
zenodo36/100

Dataset for paper: " Knowledge Graph Embeddings based Approach for Author Name Disambiguation using Literals"

<p>This dataset consists in two distinct scholarly&nbsp;knowledge graph created from two publicly available bibliographic datasets: 1) a triplestore covering information about the journal <em>Scientometrics</em>&nbsp;provided by&nbsp;<em>OpenCitations</em> (available <a href="https://doi.org/10.5281/zenodo.5151264">here</a>), and 2) the <em>AMiner </em>AND benchmark from 2018&nbsp;available <a href="https://static.aminer.cn/misc/na-data-kdd18.zip">here</a>. This KG was extracted&nbsp;for a research project on knowledge graph embeddings (KGEs)&nbsp;for author disambiguation. Structural triples of the knowledge graphs are split into training, testing and validation for applying representation learning methods. Textual literals and numeric literals were stored separately in order to implement multimodal approaches for KGEs (see&nbsp;<a href="https://arxiv.org/abs/1802.00934">arXiv:1802.00934</a>). For the same reason, textual literals and numeric literals are already stored into sentence embeddings and a&nbsp;numeric matrix&nbsp;respectively in the files&nbsp;<em>textual_literals.npy&nbsp;</em>and&nbsp;<em>numeric_literals.npy </em>in order to simplify the representation learning task. The file <em>and_eval.json</em> of each KG&nbsp;contains the evaluation dataset used for evaluating our AND architecture. For the script used to gather this dataset see <a href="https://github.com/sntcristian/and-kge/tree/main/src/AMiner-534K">https://github.com/sntcristian/and-kge/tree/main/src/AMiner-534K</a> and&nbsp;<a href="https://github.com/sntcristian/and-kge/tree/main/src/OC-782K">https://github.com/sntcristian/and-kge/tree/main/src/OC-782K</a>.</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Sampled data for varifying the correctness of CHEN-AND (Labeled Dataset for Chinese and English Joint Author Name Disambiguation)

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo28/100

LAGOS-AND-PM: A Large Gold Standard Dataset for PubMed Author Name Disambiguation

<p>LAGOS-AND-PM is an author name disambiguation (AND) dataset for the PubMed database, containing several versions, and they are built based on the ORCID database and the PubMed literature database.</p> <p>Note that we have previously created another dataset named <a href="https://zenodo.org/record/7313380">LAGOS-AND</a>, which refers to a series of AND datasets created for the MAG/OpenAlex database.</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record