Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
23
datasets available to search
ShareScore release 0.9.0
Dataset results
23 results for “citation analysis”
Foundation Species Revisited: Citation Analysis of Ellison et al. 2005
Ecologists and environmental scientists often prioritize research efforts with conservation importance. Dominant, widespread, or locally abundant species at low risk of extinction receive relatively little attention unless they are invasive. Native foundation species create habitats and environmental conditions that support many associated species and modulate local-scale ecosystem processes, but the generally high local or regional abundance of foundation species may lead to less research about them. We used citation analysis (2005-2014) to examine research following from a suggestion to identify and study foundation species while they were still common and not threatened. We explored the use and expanding definition of the foundation species concept, as well as the trajectory and ecological focus of research on foundation species throughout the world in 378 papers published in this nine-year span. Contemporary authors who cite key papers defining a foundation species pay little attention to its actual definition and species studied in this context rarely were identified as foundation species. Although functions and roles of foundation species, such as creating unique microclimates or supporting dependent species, are being studied, less research is focused on identifying them before they are threatened or lost from the ecosystem that they otherwise define. Invasive species were identified as the most common threat to foundation species. Our citation analysis and synthesis provides a new conceptual framework linking identification of and research about foundation species with their functional roles and our ability to manage emerging threats to them.
Data on a citation context analysis focusing on natural sciences and social sciences and humanities
<p>This dataset contains data on citation context analysis between natural sciences (NS) and social sciences and humanities (SSH). In particular, the data were created through manual coding of each citation between papers related to SDG7 (renewable energy) and SDG13 (climate change) and papers cited by them. This dataset consists of 9 files, associated with the article: Nishikawa, K. How and why are citations between disciplines made? A citation context analysis focusing on natural sciences and social sciences and humanities. Scientometrics (2023). <a href="https://doi.org/10.1007/s11192-023-04664-y">https://doi.org/10.1007/s11192-023-04664-y</a></p> <p> </p> <p>The files are numbered as follows:</p> <ul> <li>00 – README</li> <li>01 – Data by citation pair for SDG7 (original)</li> <li>02 – Data by citation pair for SDG13 (original)</li> <li>03 – Data by mention location for SDG7 (original)</li> <li>04 – Data by mention location for SDG13 (original)</li> <li>05 – Data by citation pair for SDG7 (additional)</li> <li>06 – Data by citation pair for SDG13 (additional)</li> <li>07 – Data by mention location for SDG7 (additional)</li> <li>08 – Data by mention location for SDG13 (additional)</li> </ul> <p>See README for more information.</p>
Citation analysis of Brown 1988 and Charnov 1976, for Figure 1 of manuscript "Halloween Charnov", Calcagno et al. 2023
<p>This contains the R script (bibliom.txt) and the ciatation data files (three .csv files) needed to generate Figure 1 in manuscript "Taking fear back into the Marginal Value Theorem: the risk-MVT and optimal boldness", by Calcagno, Gorgnard, Hamelin and Mailleret, 2023.</p>
Datasets and results of the paper titled "Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification"
<p>These are the <strong>input datasets</strong> and the <strong>results of the analyses</strong> reported on the paper titled <strong>"Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification"</strong>.</p> <p><strong>Abstract:</strong> </p> <p>The aim of this paper is to study the role of citation network measures in the assessment of scientific maturity. Referring to the case of the Italian national scientific qualification (ASN), we investigate if there is a relationship between citation network indices and the results of the researchers’ evaluation procedures. In particular, we want to understand if network measures can enhance the prediction accuracy of the results of the evaluation procedures beyond basic performance indices. Moreover, we want to highlight which citation network indices prove to be more relevant in explaining the ASN results, and if quantitative indices used in the citation-based disciplines assessment can replace the citation network measures in non-citation-based disciplines. Data concerning Statistics and Computer Science disciplines are collected from different sources (ASN, Italian Ministry of University and Research, and Scopus) and processed in order to calculate the citation-based measures used in this study. Following, we apply classification models to estimate the effects of network variables. We find that network measures are strongly related to the results of the ASN and significantly improve the explanatory power of the models, especially for the research fields of Statistics. Additionally, citation networks in the specific sub-disciplines are far more relevant than those in the general disciplines. Finally, results show that the citation network measures are not a substitute of the citation-based bibliometric indices.</p> <p><strong>Code</strong></p> <p>The code to collect and process the data used in this paper is available on GitHub at <a href="https://github.com/DigitalDataLab/ASN16-18_CitationNetwork">https://github.com/DigitalDataLab/ASN16-18_CitationNetwork</a><strong>.</strong> </p> <p><strong>Dataset description</strong></p> <p>The files <strong>AdjacencyMatrix_01B1.csv</strong>, <strong>AdjacencyMatrix_09H1.csv</strong>, <strong>AdjacencyMatrix_13D1.csv</strong>, <strong>AdjacencyMatrix_13D2.csv</strong> and <strong>AdjacencyMatrix_13D3.csv</strong> are the citation matrices for Italian academics (i.e. ASN candidates and permanent positions in the Italian academic system) in the Recruitment Fields (RFs) 01/B1, 09/H1, 13/D1, 13/D2 and 13/D3, respectively.</p> <p>The files <strong>AdjacencyMatrix_CS.csv</strong> and <strong>AdjacencyMatrix_ST.csv</strong> are the citation matrices for the Italian academics in the Computer Science disciplines (i.e. RFs 01/B1 and 09/H1) and the Statistical disciplines (i.e. RFs 13/D1, 13/D2 and 13/D3), respectively.</p> <p>The files <strong>CS_01B1_1.csv, CS_09H1_1.csv, ST_13D1_1.csv, ST_13D2_1.csv</strong> and <strong>ST_13D3_1.csv</strong> contain the data used to build the logistic regression models presented in the paper for the Italian academics at the Full Professor (FP) level.</p> <p>The files <strong>CS_01B1_2.csv, CS_09H1_2.csv, ST_13D1_2.csv, ST_13D2_2.csv</strong> and <strong>ST_13D3_2.csv</strong> contain the data used to build the logistic regression models presented in the paper for the Italian academics at the Associate Professor (AP) level.</p> <p>The file <strong>Codebook.pdf</strong> is the codebook of the previous ten files.</p> <p>The file <strong>Appendix.pdf</strong> contains the final results of the stepwise logistic regressions computed for each level (i.e. Full Professor and Associate Professor) and Recruitment Field in the Computer Science and Statistics disciplines.</p> <p>The file <strong>NormalityAssessment.pdf</strong> contains the normality assessment of citation network indices. </p>
Computer Applications in Archaeology conference proceedings citation analysis
<p>Comma-separated values (csv) files to accompany the paper: </p> <p>Huggett, J. (2024). 'Changing Theory and Practice? CAA and Archaeology's Digital Turn'. <em>Journal of Computer Applications in Archaeology</em> 7(1), pp. 316–331. DOI: <a href="https://doi.org/10.5334/jcaa.144" target="_blank" rel="noopener">10.5334/jcaa.144</a></p> <p>The data are based on the Scopus citation database (see <a href="https://www.elsevier.com/products/scopus" target="_blank" rel="noopener">https://www.elsevier.com/products/scopus</a>). The copyright for the database itself is held by Elsevier and requires a subscription to access. These files are derived data only, consisting of counts and associated statistics.</p>
Exploring the Impact of Neuroscience Preprints: A Citation Analysis
<p>1. Neuroscience_Records_Contain_Reference_to_Preprints.Scopus.V3.xlsx</p> <p>This Excel file contains the titles, DOIs, references, and EIDs of those Neuroscience publications (journal articles, books/book chapters, conference papers, notes, etc.) from 2004 to 2022 that have at least one reference to a preprint. For example, if a Neuroscience journal article has 40 references and one of these references is a preprint, then it's included in this Excel file. These records are retrieved from Scopus through the following query:</p> <p>REFSRCTITLE ( "OSF Preprints" OR "open science foundation Preprints" OR *africarxiv* OR *agrixiv* OR *arabixiv* OR *arxiv* OR *biohackrxiv* OR *biorxiv* OR *bodoarxiv* OR *cogprints* OR *eartharxiv* OR *ecoevorxiv* OR *ecsarxiv* OR *edarxiv* OR *engrxiv* OR *frenxiv* OR "INA-Rxiv" OR *indiarxiv* OR *lawarxiv* OR "LIS Scholarship Archive" OR *marxiv* OR *mediarxiv* OR *metaarxiv* OR mindrxiv OR *nutrixiv* OR paleorxiv OR "Preprints.org" OR psyarxiv OR *repec* OR *socarxiv* OR *sportrxiv* OR "Thesis Commons" OR "CoP preprint" OR "FocUS Archive preprint" OR "PeerJ preprint" OR "Law Archive preprint" OR *medrxiv* ) AND SUBJAREA ( neur ) AND PUBYEAR < 2023</p> <p> </p> <p>2. ReferencesToPreprints.V3.txt</p> <p>References of the publications are split through a Python code (SplitReferences.py) and organized into separate lines in a text file. For example, if a publication has 40 references, all of these 40 references are split into 40 separate lines. After splitting references, those lines containing one of these words/terms ("OSF Preprints" OR "open science foundation preprints" OR africarxiv OR agrixiv OR arabixiv OR arxiv OR biohackrxiv OR biorxiv OR bodoarxiv OR cogprints OR eartharxiv OR ecoevorxiv OR ecsarxiv OR edarxiv OR engrxiv OR frenxiv OR "INA-Rxiv" OR indiarxiv OR lawarxiv OR "LIS Scholarship Archive" OR marxiv OR mediarxiv OR metaarxiv OR mindrxiv OR nutrixiv OR paleorxiv OR "Preprints.org" OR psyarxiv OR repec OR socarxiv OR sportrxiv OR "Thesis Commons" OR "CoP preprint" OR "FocUS Archive preprint" OR "PeerJ preprint" OR "Law Archive preprint" OR medrxiv) are selected (through RetrieveLinesContainingSpeceficString.py) and organized into this text file (ReferencesToPreprints.V3.txt). Each reference contains an EID (separated by ";") in order to specify which publication contains this specific reference.</p> <p>After this step, through a Python code (AddPreprintServerToEndOfLines.py) the name of a certain preprint was added to the end of each line. For example, if a line (or a reference) contains "biorxiv", the word "biorxiv" will be added to the end of this line after the "@" sign.</p>
PLOS ONE – a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)
<p>This is a dataset used in and produced by research described in article "PLOS ONE - a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)" that is translation of the original Polish text "PLOS ONE – studium przypadku analizy cytowań prac naukowych na podstawie danych otwartego indeksu cytowań (OpenCitations Corpus)" published by EBiB bulletin (2017, No 176).</p> <p>Data were extracted, as nodes (PLOS_cited_nodes.csv) and edges (PLOS_edges.csv) files from the OpenCitations Corpus (http://opencitations.net/download) on 2017.07.25 and describe all cited papers published by PLOS ONE (nodes), and all citing relations (edges). The research was conducted using Gephi (https://gephi.org/) platform so the same source data are also avaiable as GEXF file (for "one-click" import capabilities). In addition, the same data are published in NET format (but be warned that due to this format limitations, information about the publication year of papers has been lost) used by PAJEK platform, as it is very popular tool for analysis of network data.</p> <p>Published figures have prefix names corresponding to figures captions in the original paper, where they have been thoroughly discussed. This data set contains also the additional figure not published in the article, showing most cited paper with citing chains of articles of lenght not greater than 3.<br> These pictures have much better quality than those published in the article, which allows for "drill down"/zoom-in analysis and large format printing.</p>
Diversity in citations to a single study: Supplementary data set for citation context network analysis
<p><strong>Introduction</strong></p> <p>This document describes the data set used for all analyses in 'Diversity in citations to a single study: A citation context network analysis of how evidence from a prospective cohort study was cited' accepted for publication in <em>Quantitative Science Studies</em> [1].</p> <p><strong>Data Collection</strong></p> <p>The data collection procedure has been fully described [1]. Concisely, the data set contains bibliometric data collected from Web of Science Core Collection via the University of Edinburgh’s Library subscription concerning all papers that cited a cohort study, Paul <em>et al.</em> [2], in the period <1985. This includes a full list of citing papers, and the citations between these papers. Additionally, it includes textual passages (citation contexts) from 343 citing papers, which were manually recovered from the full-text documents accessible via the University of Edinburgh’s Library subscription. These data have been cleaned, converted into network readable datasets, and are coded into particular classifications reflecting content, which are described fully in the supplied code book and within the manuscript [1]. </p> <p><strong>Data description</strong></p> <p>All relevant data can be found in the attached file 'Supplementary_material_Leng_QSS_2021.xlsx', which contains the following five workbooks:</p> <ul> <li><strong>“Overview”</strong> includes a list of the content of the workbooks.</li> <li><strong>“Code Book”</strong> contains the coding rules and definitions used for the classification of findings and paper titles.</li> <li><strong>“Node attribute list”</strong> includes a workbook containing all node attributes for the citation network, which includes Paul et al. [2] and its citing papers as of 1984. Highlighted in yellow at the bottom of this workbook is two papers that were discarded due to duplication - remove these if analysing this dataset in a network analysis. The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Label</em>, the formal citation of the paper to which data within this row corresponds. Citation is in the following format: last name of first author, year of publication, journal of publication, volume number, start page, and DOI (if available). </li> <li><em>Title</em>, the paper title for the paper in question.</li> <li><em>Publication_year</em>, the year of publication.</li> <li><em>Document_type, </em>the document type (e.g. review, article)</li> <li><em>WoS_ID</em>, the paper’s unique Web of Science accession number.</li> <li><em>Citation_context</em>, a column specifying whether citation context data is available from that paper</li> <li><em>Explanans</em>, the title explanans terms for that paper;</li> <li><em>Explanandum</em>, the explanandum terms for that paper.</li> <li><em>Combined_Title_Classification</em>, the combined terms used for fig 2 of the published manuscript.</li> <li><em>Serum_cholesterol_(SC)</em>, a column identifying papers that cited the serum cholesterol findings.</li> <li><em>Blood_Pressure_(BP), </em>a column identifying papers that cited the blood pressure findings.</li> <li><em>Coffee_(C),</em> a column identifying papers that cited the coffee findings.</li> <li><em>Diet_(D), </em>a column identifying papers that cited the dietary findings.</li> <li><em>Smoking_(S), </em>a column identifying papers that cited the smoking findings.</li> <li><em>Alcohol_(A), </em>a column identifying papers that cited the alcohol findings.</li> <li><em>Physical_Activity_(PA),</em> a column identifying papers that cited the physical activity findings.</li> <li><em>Body_Fatness (BF), </em>a column identifying papers that cited the body fatness findings.</li> <li><em>Indegree,</em> the number of within network citations to that paper, calculated for the network shown in Fig 4 of the manuscript.</li> <li><em>Outdegree</em>, the number of within network references of that paper as calculated for the network in Fig 4.</li> <li><em>Main_component</em>, a column specifying whether a node is contained in the largest weakly connect component as shown in Fig 4 of the manuscript.</li> <li><em>Cluster</em>, provides the cluster membership number as discussed within the manuscript (Fig 5).</li> </ol> <ul> <li><strong>“Edge list”</strong> includes a workbook including the edges for the network. The columns refer to:</li> </ul> <ol> <li><em>Source</em>, contains the node identifier of the citing paper.</li> <li><em>Target,</em> contains the node identifier of the cited paper.</li> </ol> <ul> <li><strong>“Citation context classification</strong>” includes a workbook containing the WoS accession number for the paper analysed, and any finding category discussed in that paper established via context analysis (see the code book for definitions). The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Finding_Class, </em>the findings discussed from Paul et al. within the body of the citing paper. </li> </ol> <ul> <li><strong> “Citation context data”</strong> includes a workbook containing the WoS accession number for papers in which citation context data was available, the citation context passages, the reference number or format of Paul et al. within the citing paper, and the finding categories discussed in those contexts (see code book for definitions). The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Citation_context</em>, the passage copied from the full text of the citing paper containing discussion of the findings of Paul et al.</li> <li><em>Reference_in_citing_article</em>, the reference number or format of Paul et al. within the citing paper.</li> <li><em>Finding_class, </em>the findings discussed from Paul et al. within the body of the citing paper. </li> </ol> <p><strong>Software recommended for analysis</strong></p> <p>For the analyses performed within the manuscript, Gephi version 0.9.2 was used [3], and both the edge and node lists are in a format that is easily read into this software. The Sci2 tool was used to parse data initially [4].</p> <p><strong>Notes</strong></p> <ol> <li>Leng, R. I. (Forthcoming). Diversity in citations to a single study: A citation context network analysis of how evidence from a prospective cohort study was cited. Quantitative Science Studies.</li> <li>Paul, O., Lepper, M. H., Phelan, W. H., Dupertuis, G. W., Macmillan, A., McKean, H., <em>et al.</em> (1963). A longitudinal study of coronary heart disease. <em>Circulation, </em><strong>28</strong>, 20-31. <a href="https://doi.org/10.1161/01.cir.28.1.20">https://doi.org/10.1161/01.cir.28.1.20</a>.</li> <li>Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media.</li> <li>Sci2 Team. (2009). Science of Science (Sci2) Tool. Indiana University and SciTech Strategies. Stable URL: <a href="https://sci2.cns.iu.edu">https://sci2.cns.iu.edu</a></li> </ol>
COVID-19++: A Citation-Aware Covid-19 Dataset for the Analysis of Research Dynamics
<p>COVID-19++ is a citation-aware COVID-19 dataset for the analysis of research dynamics. In addition to primary COVID-19 related articles and preprints from 2020, it includes citations and the metadata of first-order cited work. All publications are annotated with MeSH terms, either from the ground truth, or via ConceptMapper, if no ground truth was available. </p> <p>The data is organized in CSV files</p> <p>- Paper metadata (paper_id, publdate, title, data_source): paper.csv</p> <p>- Annotation data, mapping paper_id to MeSH terms: annotation.csv </p> <p>- Authorship data, mapping paper_id to author, optionally with ORCID: authorship.csv<br> - Paired DOIs of citing and cited papers: references.csv</p> <p>The column data source within the paper metadata has the value KE (for metadata from ZB MED KE), PP (for preprints) or CR (for cited resources from CrossRef)<br> </p> <p>This work was supported by BMBF within the programme ``Quantitative Wissenschaftsforschung'' under grant numbers 01PU17013A, 01PU17013B, 01PU17013C.<br> </p>
Statistical Analysis of the Effect of Equations on Citations
<p>Statistical analysis of a data set of number of equations and number of citations of papers published in volumes 94 and 104 of the journal <em>Physical Review Letters</em>. This analysis is referred to by the paper <strong>Equation-dense papers receive fewer citations—in physics as well as biology</strong> in the <em>New Journal of Physics </em>(vol. 18, article 118003) by Andrew D Higginson and Tim W Fawcett. http://iopscience.iop.org/article/10.1088/1367-2630/18/11/118003</p>
Reference Manager Data Citation Analysis
<p>DESCRIPTION:</p> <p>This package contains data used to analyze citation metadata completeness and correctness for several common reference managers used in scholarly research and several common repositories in the Earth, space, and environmental sciences.</p> <p>METHODS:</p> <p>Metadata fields for import and export methods and for 8 metadata fields (authors/creators, publisher, DOI, dataset title, version, access date, publication date, and resource type) were collected from reference managers via all import methods available (app or wizard and plugin) during summer 2024 from most recent software versions of all. To encode data, citation information for each dataset as imported by Reference Manager was compared to that registered for the DOI with DataCite. Correct metadata for each of 8 fields for both import and export was encoded as 0, incorrect as 1, and missing as '' or nan. See publication and software package for more information.</p> <p>FILES:</p> <p>FOLDER 'coded-data' contains files that include information (DOIs) about the data examined in this study, preserved copies of exported data citations used in the data interpretation and processing, and the processed data itself encoded in columns.</p> <p>FOLDER 'datacite-metadata-profiles' includes the raw metadata from each dataset DOI at the time of analysis, included for reproducibility purposes. </p> <p>FOLDER 'bibtex-files' includes the downloaded .bib files, where available, for each dataset DOI examined.</p> <p>See README file for more information.</p>
Inputs and results of "A qualitative and quantitative analysis of open citations to retracted articles: the Wakefield 1998 et al.'s case"
<p>This repository contains the datasets and visualizations generated in our work: <strong>"A qualitative and quantitative analysis of open citations to retracted articles: the Wakefield 1998 et al.’s case"</strong>.</p> <p><strong>Note:</strong> the data are all contained inside the <strong><em>data.zip</em> </strong>file. You need to unzip the container to get access to all the files and directories listed below.</p> <p>The data (citations) gathered accompanied by their annotated characteristics are stored in <strong><em>data/</em>:</strong></p> <ul> <li><em><strong>"cits_features.csv": </strong></em>a dataset containing all the entities (rows in the CSV) which have cited the Wakefield et al. retracted article, and a set of features characterizing each citing entity (columns in the CSV). The features included are: DOI ("doi"), year of publication ("year"), the title ("title"), the venue identifier ("source_id"), the title of the venue ("source_title"), yes/no value in case the entity is retracted as well ("retracted"), the subject area ("area"), the subject category ("category"), the sections of the in-text citations ("intext_citation.section"), the value of the reference pointer ("intext_citation.pointer"), the in-text citation function ("intext_citation.intent"), the in-text citation perceived sentiment ("intext_citation.sentiment"), and a yes/no value to denote whether the in-text citation context mentions the retraction of the cited entity ("intext_citation.section.ret_mention").<br> <strong>Note: </strong>this dataset is licensed under a <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">Creative Commons public domain dedication (CC0)</a>.</li> <li><em><strong>"cits_text.csv": </strong>this dataset stores the abstract ("abstract") and the in-text citations context ("intext_citation.context") </em>for each citing entity identified using the DOI value ("doi").<br> <strong>Note: </strong>the data keep their original license (the one provided by their publisher). This dataset is provided in order to favor the reproducibility of the results obtained in our work.</li> </ul> <p><strong>Topic modeling</strong></p> <p>We run a topic modeling analysis on the textual features gathered (i.e. abstracts and citation contexts). The results are stored inside the <em><strong>topic_modeling/</strong></em> directory. The topic modeling has been done using MITAO, a tool for mashing up automatic text analysis tools and creating a completely customizable visual workflow [1]. The topic modeling results for each textual feature are separated into two different folders, <em><strong>abstract/</strong></em> for the abstracts, and <em><strong>intext_cit/</strong></em> for the in-text citation contexts. Both the directories contain the datasets and visualizations generated using MITAO. </p> <p> </p> <p><strong>References</strong></p> <p>[1] Ferri, P., Heibi, I., Pareschi, L., & Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135–149. <a href="https://doi.org/10.19245/25.05.pij.5.2.3">https://doi.org/10.19245/25.05.pij.5.2.3</a></p>
Inputs and results of "A quantitative and qualitative citation analysis to retracted articles in the humanities domain"
<p>This repository contains the datasets and visualizations generated in our work: <strong>"A quantitative and qualitative citation analysis to retracted articles in the humanities domain"</strong>.</p> <p><strong>Note:</strong> the data are all contained inside the <strong><em>data.zip</em> </strong>file. You need to unzip the container to get access to all the files and directories listed below.</p> <p>The data (citations) gathered accompanied by their annotated characteristics are stored in <strong><em>data/</em>:</strong></p> <ul> <li><em>cits.csv: </em>a dataset containing all the entities (rows in the CSV) which have cited a retracted article in the humanities domain. Each citing entity (row) is accompanied by a set of features (columns) that characterizes it.<br> <strong>Note: </strong>this dataset is licensed under a <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">Creative Commons public domain dedication (CC0)</a>.</li> <li><em>content.csv: </em>a dataset containing the abstracts and the in-text citation contexts of all the citing entities gathered.<br> <strong>Note: </strong>the data keep their original license (the one provided by their publisher). This dataset is provided in order to favor the reproducibility of the results obtained in our work.</li> <li><em>excluded_hum_retractions.csv: </em>a list of the 12 humanities retracted articles with a humanities affinity score < 2, therefore excluded from the analysis. </li> </ul> <p> </p> <p><strong>Topic modeling</strong></p> <p>We run a topic modeling analysis on the textual features gathered (i.e. abstracts and citation contexts). The results are stored inside the <em><strong>topic_model/</strong></em> directory. The topic modeling has been done using MITAO, a tool for mashing up automatic text analysis tools and creating a completely customizable visual workflow [1]. The directory <em><strong>workflow/ </strong></em>contains the workflows used in MITAO. The topic modeling results for each textual feature are separated into two different folders, <em><strong>abstract/</strong></em> for the abstracts, and <em><strong>cits_context/</strong></em> for the in-text citation contexts. Both the directories contain the following directories/files: </p> <ul> <li> <p><em><strong>datasets_and_views/: </strong></em>the datasets and visualizations generated using MITAO. </p> </li> <li> <p><em><strong>ldamodel_corpus_dict/: </strong></em>it contains the dictionary, the LDA topic model, and the tokenized and vectorized corpus.</p> </li> <li><em><strong>rawdata/: </strong></em>the textual collection, metadata, and stopwords used as input in the workflow of MITAO</li> </ul> <p> </p> <p><strong>References</strong></p> <p>[1] Ferri, P., Heibi, I., Pareschi, L., & Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135–149. <a href="https://doi.org/10.19245/25.05.pij.5.2.3">https://doi.org/10.19245/25.05.pij.5.2.3</a></p> <p> </p> <ol> </ol>
Methodology data of "A qualitative and quantitative citation analysis toward retracted articles: a case of study"
<p>This document contains the datasets and visualizations generated after the application of the methodology defined in our work: <em>"A qualitative and quantitative citation analysis toward retracted articles: a case of study"</em>. The methodology defines a citation analysis of the Wakefield et al. [1] retracted article from a quantitative and qualitative point of view. The data contained in this repository are based on the first two steps of the methodology. The first step of the methodology (i.e. “Data gathering”) builds an annotated dataset of the citing entities, this step is largely discussed also in [2]. The second step (i.e. "Topic Modelling") runs a topic modeling analysis on the textual features contained in the dataset generated by the first step. </p> <p><strong>Note:</strong> the data are all contained inside the "<strong><em>method_data.zip"</em> </strong>file. You need to unzip the file to get access to all the files and directories listed below.</p> <p> </p> <p><strong>Data gathering</strong></p> <p>The data generated by this step are stored in <strong>"<em>data/</em>"</strong>:</p> <ol> <li><em><strong>"cits_features.csv": </strong></em>a dataset containing all the entities (rows in the CSV) which have cited the Wakefield et al. retracted article, and a set of features characterizing each citing entity (columns in the CSV). The features included are: DOI ("doi"), year of publication ("year"), the title ("title"), the venue identifier ("source_id"), the title of the venue ("source_title"), yes/no value in case the entity is retracted as well ("retracted"), the subject area ("area"), the subject category ("category"), the sections of the in-text citations ("intext_citation.section"), the value of the reference pointer ("intext_citation.pointer"), the in-text citation function ("intext_citation.intent"), the in-text citation perceived sentiment ("intext_citation.sentiment"), and a yes/no value to denote whether the in-text citation context mentions the retraction of the cited entity ("intext_citation.section.ret_mention").<br> <strong>Note: </strong>this dataset is licensed under a <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">Creative Commons public domain dedication (CC0)</a>.<br> </li> <li><em><strong>"cits_text.csv": </strong>this dataset stores the abstract ("abstract") and the in-text citations context ("intext_citation.context") </em>for each citing entity identified using the DOI value ("doi").<br> <strong>Note: </strong>the data keep their original license (the one provided by their publisher). This dataset is provided in order to favor the reproducibility of the results obtained in our work.</li> </ol> <p> </p> <p><strong>Topic modeling</strong><br> We run a topic modeling analysis on the textual features gathered (i.e. abstracts and citation contexts). The results are stored inside the <em><strong>"topic_modeling/"</strong></em> directory. The topic modeling has been done using MITAO, a tool for mashing up automatic text analysis tools, and creating a completely customizable visual workflow [3]. The topic modeling results for each textual feature are separated into two different folders, <em><strong>"abstracts/"</strong></em> for the abstracts, and <em><strong>"intext_cit/"</strong></em> for the in-text citation contexts. Both the directories contain the following directories/files: <strong> </strong></p> <ol> <li> <p><em><strong>"mitao_workflows/"</strong></em>: the workflows of MITAO. These are JSON files that could be reloaded in MITAO to reproduce the results following the same workflows.</p> </li> <li> <p><em><strong>"corpus_and_dictionary/": </strong></em>it contains the dictionary and the vectorized corpus given as inputs for the LDA topic modeling.</p> </li> <li> <p><em><strong>"coherence/coherence.csv":</strong></em> the coherence score of several topic models trained on a number of topics from 1 - 40.</p> </li> <li> <p><em><strong>"datasets_and_views/": </strong></em>the datasets and visualizations generated using MITAO. </p> </li> </ol> <p> </p> <p><strong>References</strong></p> <ol> <li>Wakefield, A., Murch, S., Anthony, A., Linnell, J., Casson, D., Malik, M., Berelowitz, M., Dhillon, A., Thomson, M., Harvey, P., Valentine, A., Davies, S., & Walker-Smith, J. (1998). RETRACTED: Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. <em>The Lancet</em>, <em>351</em>(9103), 637–641. <a href="https://doi.org/10.1016/S0140-6736(97)11096-0">https://doi.org/10.1016/S0140-6736(97)11096-0</a></li> <li> <p>Heibi, I., & Peroni, S. (2020). A methodology for gathering and annotating the raw-data/characteristics of the documents citing a retracted article v1 (protocols.io.bdc4i2yw) [Data set]. In protocols.io. ZappyLab, Inc. <a href="https://doi.org/10.17504/protocols.io.bdc4i2yw">https://doi.org/10.17504/protocols.io.bdc4i2yw</a></p> </li> <li> <p> </p> Ferri, P., Heibi, I., Pareschi, L., & Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135–149. <a href="https://doi.org/10.19245/25.05.pij.5.2.3">https://doi.org/10.19245/25.05.pij.5.2.3</a> <p> </p> </li> </ol>
Dataset for Spence et al., "Availability of study protocols for randomized trials published in high-impact medical journals: cross-sectional analysis" (CITATION)
<p>Contains our extraction sheets (as SAS data files), code to calculate the values in the tables in our manuscript, and a supplemental file with additional notes on methods used in our study.</p>
Dataset for Citation Network Analysis on Quality Cues for Meat Purchases
<p>Data are retrieved from two databases permissible to the Vosviewer® software: Scopus and Web of Science. Three sets of Boolean search strings are used to find articles in the databases:</p> <p>(i) [<em>(“certification” AND “meat”) OR (“certification” AND “dairy”)</em>]</p> <p>(ii) [<em>(“credence” AND meat) OR (“experience” AND meat)</em>]<em> </em></p> <p><em>(iii) </em>[<em>(“meat” AND “intrinsic”) OR (“meat” AND “extrinsic”) OR (“meat” AND “quality cue”)</em>].</p> <p>The searches targeted the title, abstract and keywords sections for the Scopus database, and the topic section for the Web of Science database. The search was limited to peer-reviewed articles published in the English language; no timeline restrictions were set for the initial search. </p>
Dataset for "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations"
<p>This is the dataset for the paper "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations" submitted to iConference 2025.</p>
Dataset for: Fifty years of research on questionable research practices in science: Quantitative analysis of co-citation patterns
<p>Questionable research practices (QRPs) have been the focus of the scientific community amid greater scrutiny and evidence highlighting issues with replicability across many fields of science. To capture the most impactful publications and the main thematic domains in the literature on QRPs, this study uses a document co-citation analysis. The analysis was conducted on a sample of 341 documents that covered the past 50 years of research in QRPs. Nine major thematic clusters emerged. Statistical reporting and statistical power emerged as key areas of research, where systemic-level factors in how research is conducted are consistently raised as the precipitating factors for QRPs. There is also an encouraging shift in the focus of research into open science practices designed to address engagement in QRPs. Such a shift is indicative of the growing momentum of the open science movement, and more research can be conducted on how these practices are employed on the ground and how their uptake by researchers can be further promoted. However, the results suggest that, while pre-registration and registered reports receive the most research interest, less attention has been paid to other open science practices (e.g., data and methods sharing).</p>
Dataset for: Fifty years of research on questionable research practices in science: Quantitative analysis of co-citation patterns
Open the record for dataset details and reuse information.
Agrivoltaic grazing systems for a sustainable future: Citation database for a multi-disciplinary review & gap analysis
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.