Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.9.0
Dataset results
13 results for “citation impact”
Citation data of arXiv eprints and the associated quantitatively-and-temporally normalised impact metrics
<p><strong>Data collection</strong></p> <p>This dataset contains information on the eprints posted on arXiv from its launch in 1991 until the end of 2019 (1,589,006 unique eprints), plus the data on their citations and the associated impact metrics. Here, eprints include preprints, conference proceedings, book chapters, data sets and commentary, i.e. every electronic material that has been posted on arXiv. </p> <p>The content and metadata of the arXiv eprints were retrieved from the arXiv API (https://arxiv.org/help/api/) as of 21st January 2020, where the metadata included data of the eprint’s title, author, abstract, subject category and the arXiv ID (the arXiv’s original eprint identifier). In addition, the associated citation data were derived from the Semantic Scholar API (https://api.semanticscholar.org/) from 24th January 2020 to 7th February 2020, containing the citation information in and out of the arXiv eprints and their published versions (if applicable). Here, whether an eprint has been published in a journal or other means is assumed to be inferrable, albeit indirectly, from the status of the digital object identifier (DOI) assignment. It is also assumed that if an arXiv eprint received <em>c</em><sub>pre</sub> and <em>c</em><sub>pub</sub> citations until the data retrieval date (7th February 2020) before and after it is assigned a DOI, respectively, then the citation count of this eprint is recorded in the Semantic Scholar dataset as <em>c</em><sub>pre</sub> + <em>c</em><sub>pub</sub>. Both the arXiv API and the Semantic Scholar datasets contained the arXiv ID as metadata, which served as a key variable to merge the two datasets.</p> <p>The classification of research disciplines is based on that described in the arXiv.org website (https://arxiv.org/help/stats/2020_by_area/). There, the arXiv subject categories are aggregated into several disciplines, of which we restrict our attention to the following six disciplines: Astrophysics (‘astro-ph’), Computer Science (‘comp-sci’), Condensed Matter Physics (‘cond-mat’), High Energy Physics (‘hep’), Mathematics (‘math’) and Other Physics (‘oth-phys’), which collectively accounted for 98% of all the eprints. Those eprints tagged to multiple arXiv disciplines were counted independently for each discipline. Due to this overlapping feature, the current dataset contains a cumulative total of 2,011,216 eprints. </p> <p>Some general statistics and visualisations per research discipline are provided in the original article (Okamura, 2022), where the validity and limitations associated with the dataset are also discussed.</p> <p> </p> <p><strong>Description of columns (variables)</strong></p> <ul> <li><strong>arxiv_id</strong> : arXiv ID</li> <li><strong>category</strong> : Research discipline</li> <li><strong>pre_year</strong> : Year of posting v1 on arXiv</li> <li><strong>pub_year</strong> : Year of DOI acquisition</li> <li><strong>c_tot</strong> : No. of citations acquired during 1991–2019</li> <li><strong>c_pre</strong> : No. of citations acquired before and including the year of DOI acquisition</li> <li><strong>c_pub</strong> : No. of citations acquired after the year of DOI acquisition</li> <li><strong>c_<em>yyyy</em></strong> (<em>yyyy</em> = 1991, …, 2019) : No. of citations acquired in the year <em>yyyy</em> (with ‘<em>yyyy</em>’ running from 1991 to 2019)</li> <li><strong>gamma</strong> : The quantitatively-and-temporally normalised citation index</li> <li><strong>gamma_star</strong> : The quantitatively-and-temporally standardised citation index</li> </ul> <p><em>Note:</em> The definition of the quantitatively-and-temporally normalised citation index (γ; ‘gamma’) and that of the standardised citation index (γ*; ‘gamma_star’) are provided in the original article (Okamura, 2022). Both indices can be used to compare the citational impact of papers/eprints published in different research disciplines at different times. </p> <p> </p> <p><strong>Data files</strong></p> <p>A comma-separated values file (‘<strong>arXiv_impact.csv</strong>’) and a Stata file (‘<strong>arXiv_impact.dta</strong>’) are provided, both containing the same information.</p> <p> </p>
Triangle of Biomedicine Framework to Analyze the Citations' Impact on Categories Dissemination in the PubMed Database
<p>This is the data and the most relevant script of the paper 'Triangle of Biomedicine Framework to Analyze the Citations’ Impact on Categories Dissemination in the PubMed Database'.</p>
Exploring the Impact of Neuroscience Preprints: A Citation Analysis
<p>1. Neuroscience_Records_Contain_Reference_to_Preprints.Scopus.V3.xlsx</p> <p>This Excel file contains the titles, DOIs, references, and EIDs of those Neuroscience publications (journal articles, books/book chapters, conference papers, notes, etc.) from 2004 to 2022 that have at least one reference to a preprint. For example, if a Neuroscience journal article has 40 references and one of these references is a preprint, then it's included in this Excel file. These records are retrieved from Scopus through the following query:</p> <p>REFSRCTITLE ( "OSF Preprints" OR "open science foundation Preprints" OR *africarxiv* OR *agrixiv* OR *arabixiv* OR *arxiv* OR *biohackrxiv* OR *biorxiv* OR *bodoarxiv* OR *cogprints* OR *eartharxiv* OR *ecoevorxiv* OR *ecsarxiv* OR *edarxiv* OR *engrxiv* OR *frenxiv* OR "INA-Rxiv" OR *indiarxiv* OR *lawarxiv* OR "LIS Scholarship Archive" OR *marxiv* OR *mediarxiv* OR *metaarxiv* OR mindrxiv OR *nutrixiv* OR paleorxiv OR "Preprints.org" OR psyarxiv OR *repec* OR *socarxiv* OR *sportrxiv* OR "Thesis Commons" OR "CoP preprint" OR "FocUS Archive preprint" OR "PeerJ preprint" OR "Law Archive preprint" OR *medrxiv* ) AND SUBJAREA ( neur ) AND PUBYEAR < 2023</p> <p> </p> <p>2. ReferencesToPreprints.V3.txt</p> <p>References of the publications are split through a Python code (SplitReferences.py) and organized into separate lines in a text file. For example, if a publication has 40 references, all of these 40 references are split into 40 separate lines. After splitting references, those lines containing one of these words/terms ("OSF Preprints" OR "open science foundation preprints" OR africarxiv OR agrixiv OR arabixiv OR arxiv OR biohackrxiv OR biorxiv OR bodoarxiv OR cogprints OR eartharxiv OR ecoevorxiv OR ecsarxiv OR edarxiv OR engrxiv OR frenxiv OR "INA-Rxiv" OR indiarxiv OR lawarxiv OR "LIS Scholarship Archive" OR marxiv OR mediarxiv OR metaarxiv OR mindrxiv OR nutrixiv OR paleorxiv OR "Preprints.org" OR psyarxiv OR repec OR socarxiv OR sportrxiv OR "Thesis Commons" OR "CoP preprint" OR "FocUS Archive preprint" OR "PeerJ preprint" OR "Law Archive preprint" OR medrxiv) are selected (through RetrieveLinesContainingSpeceficString.py) and organized into this text file (ReferencesToPreprints.V3.txt). Each reference contains an EID (separated by ";") in order to specify which publication contains this specific reference.</p> <p>After this step, through a Python code (AddPreprintServerToEndOfLines.py) the name of a certain preprint was added to the end of each line. For example, if a line (or a reference) contains "biorxiv", the word "biorxiv" will be added to the end of this line after the "@" sign.</p>
Data for "Open Access impact on citations: a case study"
<p>This dataset is a list of 347 papers published in 2010 and retrieved from the Web of Science, Scopus and Google Scholar. For each paper, the number of citations and the citation date(s) have been collected. If the full-text is available online, the date of "liberation" and the URL of the file have been retrieved as well. The objective was to assess the impact of Open access on citation rate and more particularly the impact before and after full-text "liberation".</p> <p> </p>
Data for "Measuring Back: Bibliodiversity and the Journal Impact Factor brand. A Case study of IF-journals included in the 2021 Journal Citations Report."
<p>This is the open data for the preprint "Measuring Back: Bibliodiversity and the Journal Impact Factor brand. A Case study of IF-journals included in the 2021 Journal Citations Report."</p>
Exploring the Impact of Negative Sampling on Patent Citation Recommendation
<ul> <li> <p><strong>pcr_patents.csv </strong>is the dataset which is generated by collecting samples randomly from Google Patents by exploiting a <a href="https://pypi.org/project/google-patent-scraper/">Python library</a>. The dataset comprises around 250,000 US patents and their titles, abstracts, and citations. Each patent has roughly on average 27 citations.</p> </li> </ul> <p>The zip file contains 3 different datasets for training and testing patent citation recommendation systems. These datasets were generated by utilizing the main dataset. They consist of around 1 million instances which are positive as well as negative samples. </p> <ul> <li> <p><strong>pcr_cpc_negative_sample_data.csv</strong> consists of negative samples that were generated based on CPC subclass codes. </p> </li> <li> <p><strong>pcr_random_negative_sample_data.csv</strong> consists of negative samples that were generated randomly. </p> </li> <li> <p><strong>pcr_sem_sim_negative_sample_data_2.csv</strong> consists of negative samples that were generated based on nearest neighbor relation.</p> </li> </ul>
Dataset for Spence et al., "Availability of study protocols for randomized trials published in high-impact medical journals: cross-sectional analysis" (CITATION)
<p>Contains our extraction sheets (as SAS data files), code to calculate the values in the tables in our manuscript, and a supplemental file with additional notes on methods used in our study.</p>
Citation reasons and their impact on knowledge
Open the record for dataset details and reuse information.
Hidden citations obscure true impact in science
<p>The repo consists of the intermediate results (subGPhy10000/) for "Hidden citations obscure true impact in science" (https://academic.oup.com/pnasnexus/article/3/5/pgae155/7664049). The corresponding code is available at <a href="https://github.com/Barabasi-Lab/hidden-citation" target="_blank" rel="noopener">https://github.com/Barabasi-Lab/hidden-citation</a>.</p>
Analysis of the Emerging Source Citation Index (coverage and impact) in social science and humanities (2005-2018)
<p>The information presented is supplementary material to the paper " <strong>Is the Emerging Source Citation Index an aid to assess the citation impact in social science and humanities? </strong> "</p>
A multidimensional framework for characterizing the citation impact of scientific publications
<p>This data set pertains to the following research article: Bu, Y., Waltman, L., & Huang, Y. (2020). <em>A multidimensional framework for characterizing the citation impact of scientific publications</em>. arXiv:1901.09663.</p>
Data from: The assessment of science: the relative merits of post-publication review, the impact factor and the number of citations
Background: The assessment of scientific publications is an integral part of the scientific process. Here we investigate three methods of assessing the merit of a scientific paper: subjective post-publication peer review, the number of citations gained by a paper and the impact factor of the journal in which the article was published. Methodology/principle findings: We investigate these methods using two datasets in which subjective post-publication assessments of scientific publications have been made by experts. We find that there are moderate, but statistically significant, correlations between assessor scores, when two assessors have rated the same paper, and between assessor score and the number of citations a paper accrues. However, we show that assessor score depends strongly on the journal in which the paper is published, and that assessors tend to over-rate papers published in journals with high impact factors. If we control for this bias, we find that the correlation between assessor scores and between assessor score and the number of citations is weak, suggesting that scientists have little ability to judge either the intrinsic merit of a paper or its likely impact. We also show that the number of citations a paper receives is an extremely error-prone measure of scientific merit. Finally, we argue that the impact factor is likely to be a poor measure of merit, since it depends on subjective assessment. Conclusions: We conclude that the three measures of scientific merit considered here are poor; in particular subjective assessments are an error-prone, biased and expensive method by which to assess merit. We argue that the impact factor may be the most satisfactory of the methods we have considered, since it is a form of pre-publication review. However, we emphasise that it is likely to be a very error-prone measure of merit that is qualitative, not quantitative.
Data from: The assessment of science: the relative merits of post-publication review, the impact factor and the number of citations
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.