Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “Dataset, Author Name Disambiguation”
LSPO: A Large-Scale Physics ORCiD-Linked Dataset for Author Name Disambiguation
<p>The LSPO dataset, a Large-Scale Physics ORCiD-Linked Dataset for Author Name Disambiguation is comprised of 554,962 NASA/ADS publications linked to 125,486 unique researchers through ORCiD identifiers. The available meta-data fields are: ORCiD identifier, author name, affiliation, title, asbtract, and name block. The dataset can be utilized to make pairs or triplets for training a author name disambiguation model. </p>
CHEN-AND: A Labeled Dataset for Chinese and English Joint Author Name Disambiguation
<p>Abstract: Author name disambiguation (AND) is an important problem in literature databases and is even more prominent in the cross-language (database) context. Extensive research has been conducted in the academic community to eliminate such ambiguity. However, the existing research mainly focuses on monolingual literature, with less attention paid to author disambiguation in cross-language environments. In this regard, this study focuses on a typical cross-language author disambiguation task - Chinese and English joint author disambiguation. We first propose an automated dataset construction method for Chinese-English literature joint AND using online open resources, with this method, we create a dataset named CHEN-AND for joint Chinese and English author disambiguation. Then we propose a merging-first-then-disambiguation (MFTD)--based disambiguation framework and evaluate several variants of this method on the test dataset. For details on building this dataset, please refer to <a href="https://github.com/carmanzhang/CHEN-AND">this repository</a> on my GitHub page.</p>
LAGOS-AND: A Large Gold Standard Dataset for MAG/OpenAlex Author Name Disambiguation
<p>We present a large gold standard dataset for author name disambiguation (AND) research (LAGOS-AND), which contains two sub-datasets, LAGOS-AND-BLOCK and LAGOS-AND-PAIRWISE. The datasets were automatically built by using the two authoritative sources, ORCID and DOI, based on the ORCID open database and an open literature database (MAG or OpenAlex).</p> <p>The currently available versions of the LAGOS-AND datasets are:</p> <ol> <li><strong>Version 1.0 (MAG+ORCID)</strong>: This is the initial version of the LAGOS-AND dataset, the evaluation results and quality control measures of this version dataset can be found in this paper <a href="https://arxiv.org/abs/2104.01821">https://arxiv.org/abs/2104.01821</a>.</li> <li><strong>Version 2.0 (OpenAlex+ORCID)</strong>: This version builds on OpenAlex instead of MAG because MAG was discontinued on 31 December 2021, and OpenAlex not only positions itself as a drop-in replacement for MAG but also keeps evolving by aggregating academic resources from other repositories. In addition, the pairwise-based sub-dataset (LAGOS-AND-PAIRWISE v2.0) improves the accuracy of labeled authorship (class label) as compared to LAGOS-AND-PAIRWISE v1.0.</li> </ol> <p>Note that there are other versions of the dataset, which we call "pre-release". We recommend users to use the normal version of the dataset. These pre-release versions were originally intended to be released as the normal versions. However, during the preparation of the research paper, we found an issue with the dataset and the reviewers also made some reasonable requests for the dataset. This led us to update the dataset. Unfortunately, the Zenodo platform does not allow updates to the same version of dataset, so we had to create new versions. The created pre-release datasets are as follows:</p> <ol> <li><strong>Version 2.0-alpha</strong>: For few samples in LAGOS-AND-PAIRWISE, the class labels are incorrect. We improve the accuracy of the class label in the normal Version 2.0 dataset by using a better random sampling approach.</li> <li><strong>Version 1.0-beta</strong>: We created this version because a sub-dataset of this version LAGOS-AND-PAIRWISE contains only ~500K author pairs, while it should contain ~1M author pairs, as described in our paper. We fixed the problem in the Version 1.0 dataset.</li> <li><strong>Version 1.0-alpha</strong>: The earliest dataset uploaded to Zenodo, corresponding to the original dataset before addressing the reviewers' comments and suggestions. In contrast, the dataset in Version 1.0 is the dataset after the reviewers' comments and suggestions have been addressed.</li> </ol>
Dataset for paper: " Knowledge Graph Embeddings based Approach for Author Name Disambiguation using Literals"
<p>This dataset consists in two distinct scholarly knowledge graph created from two publicly available bibliographic datasets: 1) a triplestore covering information about the journal <em>Scientometrics</em> provided by <em>OpenCitations</em> (available <a href="https://doi.org/10.5281/zenodo.5151264">here</a>), and 2) the <em>AMiner </em>AND benchmark from 2018 available <a href="https://static.aminer.cn/misc/na-data-kdd18.zip">here</a>. This KG was extracted for a research project on knowledge graph embeddings (KGEs) for author disambiguation. Structural triples of the knowledge graphs are split into training, testing and validation for applying representation learning methods. Textual literals and numeric literals were stored separately in order to implement multimodal approaches for KGEs (see <a href="https://arxiv.org/abs/1802.00934">arXiv:1802.00934</a>). For the same reason, textual literals and numeric literals are already stored into sentence embeddings and a numeric matrix respectively in the files <em>textual_literals.npy </em>and <em>numeric_literals.npy </em>in order to simplify the representation learning task. The file <em>and_eval.json</em> of each KG contains the evaluation dataset used for evaluating our AND architecture. For the script used to gather this dataset see <a href="https://github.com/sntcristian/and-kge/tree/main/src/AMiner-534K">https://github.com/sntcristian/and-kge/tree/main/src/AMiner-534K</a> and <a href="https://github.com/sntcristian/and-kge/tree/main/src/OC-782K">https://github.com/sntcristian/and-kge/tree/main/src/OC-782K</a>.</p>
Sampled data for varifying the correctness of CHEN-AND (Labeled Dataset for Chinese and English Joint Author Name Disambiguation)
Open the record for dataset details and reuse information.
LAGOS-AND-PM: A Large Gold Standard Dataset for PubMed Author Name Disambiguation
<p>LAGOS-AND-PM is an author name disambiguation (AND) dataset for the PubMed database, containing several versions, and they are built based on the ORCID database and the PubMed literature database.</p> <p>Note that we have previously created another dataset named <a href="https://zenodo.org/record/7313380">LAGOS-AND</a>, which refers to a series of AND datasets created for the MAG/OpenAlex database.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.