Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

214

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

214 results for “citations”

Learn how ShareScore rates datasets ↗
zenodo44/100

Preprint Citations in PLOS Dataset

<p>Preprints are research articles that have been published online before undergoing peer review. The role of preprints in the scientific production has been growing in recent years. Our objective is to study these practices and evaluate the differences that exist between citations to preprints and citations to peer-reviewed articles.</p> <p>This dataset contains citation contexts to preprints extracted from the PLOS dataset. We have processed all PLOS articles published up to January 2021. Preprint citations were identified by matching cited source metadata against a list of existing preprint databases. For each citation we have extracted the sentence and its position in the IMRaD structure of the article.</p> <p>The data is presented in a tsv file that contains the following columns :</p> <ul> <li>id: identifier.</li> <li>source_name: name of the preprint database where the preprint is published. In some cases source_name is &quot;preprint kw&quot; which means that it has been identified by the presence of the &quot;preprint&quot; keyword in the source metadata, but could not be linked to a known preprint database.</li> <li>jtitle: title of the PLOS journal from which the citation context is extracted.</li> <li>imrad_code: one of &quot;I&quot;, &quot;M&quot;, &quot;R&quot;, &quot;D&quot;, indicating the name of the section of the citation context in the IMRaD (Introduction, Methods, Results and Discussion) structure.</li> <li>perc: a number between 0 and 100, indicating the position of the citation context in terms of percentage of the text progression of the section in which it appears. This position has been calculated by dividing the number of the sentence of the citation context by the total number of sentences in the section.</li> <li>pub_year: publication year of the article</li> <li>sentence_text: sentence containing the citation to the preprint.</li> </ul> <p>The full description of the dataset and the processing steps to obtain it are described in:</p> <p>Bertin, Marc and Atanassova, Iana (2022). &quot;Preprint Citation Praxis in PLOS&quot;. Scientometrics.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Datasets and results of the paper titled "Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification"

<p>These&nbsp;are&nbsp;the <strong>input&nbsp;datasets</strong> and the <strong>results of the analyses</strong>&nbsp;reported on&nbsp;the paper titled <strong>&quot;Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification&quot;</strong>.</p> <p><strong>Abstract:</strong>&nbsp;</p> <p>The aim of this paper is to study the role of citation network measures in the assessment of scientific maturity. Referring to the case of the Italian national scientific qualification (ASN), we investigate if there is a relationship between citation network indices and the results of the researchers&rsquo; evaluation procedures. In particular, we want to understand if network measures can enhance the prediction accuracy of the results of the evaluation procedures beyond basic performance indices. Moreover, we want to highlight which citation network indices prove to be more relevant in explaining the ASN results, and if quantitative indices used in the citation-based disciplines assessment can replace the citation network measures in non-citation-based disciplines. Data concerning Statistics and Computer Science disciplines are collected from different sources (ASN, Italian Ministry of University and Research, and Scopus) and processed in order to calculate the citation-based measures used in this study. Following, we apply classification models to estimate the effects of network variables. We find that network measures are strongly related to the results of the ASN and significantly improve the explanatory power of the models, especially for the research fields of Statistics. Additionally, citation networks in the specific sub-disciplines are far more relevant than those in the general disciplines. Finally, results show that the citation network measures are not a substitute of the citation-based bibliometric indices.</p> <p><strong>Code</strong></p> <p>The code to collect&nbsp;and process the data used in this paper is available on GitHub at <a href="https://github.com/DigitalDataLab/ASN16-18_CitationNetwork">https://github.com/DigitalDataLab/ASN16-18_CitationNetwork</a><strong>.</strong>&nbsp;</p> <p><strong>Dataset description</strong></p> <p>The files&nbsp;<strong>AdjacencyMatrix_01B1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_09H1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_13D1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_13D2.csv</strong> and&nbsp;<strong>AdjacencyMatrix_13D3.csv</strong> are the&nbsp;citation matrices for Italian academics (i.e. ASN candidates and permanent positions in the Italian academic system) in the Recruitment Fields (RFs) 01/B1, 09/H1, 13/D1, 13/D2 and&nbsp;13/D3, respectively.</p> <p>The files&nbsp;<strong>AdjacencyMatrix_CS.csv</strong>&nbsp;and&nbsp;<strong>AdjacencyMatrix_ST.csv</strong> are the citation matrices for the Italian academics in the Computer Science disciplines (i.e. RFs 01/B1 and 09/H1) and the Statistical disciplines (i.e. RFs 13/D1,&nbsp;13/D2 and&nbsp;13/D3), respectively.</p> <p>The files&nbsp;<strong>CS_01B1_1.csv,&nbsp;CS_09H1_1.csv, ST_13D1_1.csv,&nbsp;ST_13D2_1.csv</strong> and&nbsp;<strong>ST_13D3_1.csv</strong>&nbsp;contain the data used to build the&nbsp;logistic regression models presented in the paper for the Italian academics at the Full Professor (FP) level.</p> <p>The files&nbsp;<strong>CS_01B1_2.csv,&nbsp;CS_09H1_2.csv, ST_13D1_2.csv,&nbsp;ST_13D2_2.csv</strong> and&nbsp;<strong>ST_13D3_2.csv</strong>&nbsp;contain the data used to build the&nbsp;logistic regression models presented in the paper for the Italian academics at the Associate Professor (AP) level.</p> <p>The file&nbsp;<strong>Codebook.pdf</strong>&nbsp;is the codebook of the previous ten files.</p> <p>The file <strong>Appendix.pdf</strong> contains the final results of the stepwise logistic regressions computed for each level (i.e. Full Professor and Associate Professor) and Recruitment Field in the Computer Science and Statistics disciplines.</p> <p>The file&nbsp;<strong>NormalityAssessment.pdf</strong>&nbsp;contains the&nbsp;normality assessment of citation network indices.&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Libraries & Recommended Citations for using PLAsTiCC Models

<p>Text file libraries for transient and variable source models used in the &quot;Photometric LSST Astronomical Time-Series Classification Challenge&quot; (PLAsTiCC). &nbsp;The original challenge (Sep 28, 2018 - Dec 17, 2018) was hosted at&nbsp;&nbsp;https://www.kaggle.com/c/PLAsTiCC-2018. See AAA_README.pdf for more information.</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Reliquary of contacts for: A pragmatic approach to complex citations, closing the provenance gap between IPCC AR6 figures and CMIP6 simulations

<p>Photos and metadadata pannels of a "Reliquary of contacts for: A pragmatic approach to complex citations, closing the provenance gap between IPCC AR6 figures and CMIP6 simulations" produced to support the "A pragmatic approach to complex citations, closing the provenance gap between IPCC AR6 figures and CMIP6 simulations" presentation given at EGU 2024.</p> <p>------</p> <p>With ever growing abilities to process greater volumes of data the abiiity to sustain the citability and tracability of the underluing source data within outputs such as publications is becoming increasingly challenging. With a range of use-cases, work on how to handle complex citations from the perspective of those producing outputs, journals and those handling the knowledge graph and associated services, is exmaning a how to handle these situations in a sustainable and manageable fashion.<br><br>At the European Geophysical Union (EGU) General Assembly in Vienna, 2024, a pragmatic solution using Zenodo to store 'reliquary' objects was presented. The poster presentation demonstrated the use of existing strucutres within a Zenodo object to address the complex citation use-case around figure, the related data and the source datasets related to the IPCC's AR5 figure data. I.e. how to utulise the existing constructs of a Zenodo item and the range of available metadata fields to give an off-the-shelf solution to allow tracability to the specific datasets used (via their Handle identifiers) and citability of the higher level, DOI-ed dataset collections within which the specific Handle-ed datasets were selected from. Additionally, the connectivity between these two levels of PID objects was also captured within the stored files around which the rich metata was captured.<br><br>The concept of a complex citation 'reliquary' as a metadtata rich object, acting as a referencable nexus in the knowledge graph has been put forth as a solution to the complex citation challenge. It borrows the concept from its historical use, denoting a container or shrine, often richly embellished, for sacred relics (e.g. saints bones, artefacts etc). In the same way here we have both the rich metadata 'container' around the specific details (the 'bones in the box', with their preserved connectivity).<br><br>However, the term 'reliquary' is often a hard one to convey, being somewhat of an obscure term (likewise the term 'nexus' may also be one lacking wider recogniton). Thus, to aid the discussions around the presentation by Pascoe et al. (2024) at the EGU 2023 General Assembly, a physical representation of a metadata reliquary object was produced.<br><br>The purpose of this object was two fold:<br><br>&nbsp;- The first was to show how the reliquary container itself is metadata rich, detailing through the use of ORCIDS, RORs and a DOI, references to external items, complemented by further metadata concerning the specifics of the reliquary's own metadata (its title and the credit for the artist that created it). Futher more, the relationship between the reliquary and those referenced parties/objects was also captured. The contents were also used to demonstrate the importance of making the contents useful for onward users (in this case contact details on business cards). <br>&nbsp;- The second, and for the funder of this piece, arguably the most important aspect was a degree of outreach this provided, both to engage the audience of Pascoe et al (2024), and directly to the artist to demonstrate the importance of this work to the international research data management community and overall to aid engagemeng with the funder's work.<br><br>This resource is provided here as a repository of images of the reliquary itself and in context at the EGU 2024 event as a potential resource others may use to aid further discussions around the use of reliquaries with regards to complex citations. The slides provided of the reliquary box labels are also provided with some annotation to further expand on the metadata aspects of their content.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

DOI-to-DOI citations of articles '10.1016/j.websem.2012.08.001' and '10.1016/j.websem.2017.06.001'

<p>This dataset contains the citations of two articles I&#39;ve coauthored, i.e.&nbsp;&#39;10.1016/j.websem.2012.08.001&#39; and &#39;10.1016/j.websem.2017.06.001&#39; published on Elsevier&#39;s Web Semantics. The citation data are compliant with the CSV format for CROCI, http://opencitations.net/index/croci.</p>

opencc-zeroMar 2019View details →
zenodo44/100

Citations to software and data in Zenodo via open sources

<p>In January 2019, the Asclepias Broker harvested citation links to Zenodo objects from three discovery systems: the NASA Astrophysics Datasystem (ADS), Crossref Event Data and Europe PMC. Each row of our dataset represents one unique link between a citing publication and a Zenodo DOI. Both endpoints are described by basic metadata. The second dataset contains usage metrics for every cited Zenodo DOI of our data sample.&nbsp;</p> <p>&nbsp;</p>

opencc-zeroOct 2019View details →
zenodo44/100

Computer Applications in Archaeology conference proceedings citation analysis

<p>Comma-separated values (csv) files to accompany the paper:&nbsp;</p> <p>Huggett, J. (2024). 'Changing Theory and Practice? CAA and Archaeology's Digital Turn'. <em>Journal of Computer Applications in Archaeology</em> 7(1), pp. 316&ndash;331. DOI: <a href="https://doi.org/10.5334/jcaa.144" target="_blank" rel="noopener">10.5334/jcaa.144</a></p> <p>The data are based on the Scopus citation database (see <a href="https://www.elsevier.com/products/scopus" target="_blank" rel="noopener">https://www.elsevier.com/products/scopus</a>). The copyright for the database itself is held by Elsevier and requires a subscription to access. These files are derived data only, consisting of counts and associated statistics.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Source Data for Manuscript: Identifying genomic data use with the Data Citation Explorer

<p>This page contains the source data for the manuscript describing the Data Citation Explorer, currently in review for publication. The preprint version can be found on this page.</p> <p>Files:</p> <p><strong>DCE_manual_eval_sample.xlsx:</strong></p> <p>This file was used to manually evaluate hits generated by the Data Citation Explorer. There are two separate sheets: one with publications returned by searches in PubMed and PubMed Central and another with publications returned by searches in Dimensions. Column descriptions can be found in the file itself. Each row in each evaluation sheet refers to a pair between a JAMO record and a linked publication.</p> <p><strong>DCE_citation_report.csv</strong></p> <p>Contains JAMO record IDs and PubMed IDs from the initial 2020 DCE trial run. There are 238,994 unique JAMO IDs and 30,641 unique PubMed IDs. 78,104 JAMO records are linked with publications.</p> <p>Columns:</p> <ul> <li>jamo_id - unique JAMO record ID</li> <li>sample_group - Sample strata from which manually evaluated records were pulled</li> <li>citation_count - Number of citations associated with each record</li> <li>citations - comma-delimited PubMed IDs for linked publications</li> <li>sampled - True/False, denoting which records were included in the initial evaluation sample</li> <li>notes - descriptions for why certain sampled records were excluded from manual evaluation</li> <li>unprocessed - True/False. These 7,890 records contained anomalous fields that caused them to be rejected for processing. They are represented as zero-length files in the archive.</li> </ul> <p><strong>DCE_source_files.zip:</strong></p> <p>This folder contains 3 files for each JAMO record in DCE_citation_report.tsv. For each JAMO record listed in the citation report, three files are provided:</p> <ol> <li>JAMO_ID_source.yaml - The fields extracted from the JAMO record that were relevant to the citation search, including any previously known PMIDs (manually curated).</li> <li>JAMO_ID_expand.yaml - The source record augmented with additional metadata discovered in other resources, including the citations that were discovered based on querying PubMed Central for the values in those metadata fields.</li> <li>JAMO_ID_audit.json - The audit path as a directed acyclic graph, in JSON.</li> </ol>

openJul 2024View details →
zenodo44/100

BIP! NDR (NoDoiRefs): a dataset of citations from papers without DOIs in computer science conferences and workshops

<h2>Overview</h2> <p>In the field of Computer Science, conference and workshop papers serve as important contributions, carrying substantial weight in research assessment processes, compared to other disciplines. However, a considerable number of these papers are not assigned a Digital Object Identifier (DOI), hence their citations are not reported in widely used citation datasets like OpenCitations and Crossref, raising limitations to citation analysis. While the Microsoft Academic Graph (MAG) previously addressed this issue by providing substantial coverage, its discontinuation&nbsp; has created a void in available data.</p> <p>BIP! NDR aims to alleviate this issue and enhance the research assessment processes within the field of Computer Science. To accomplish this, it leverages a workflow that identifies and retrieves Open Science papers lacking DOIs from the DBLP Corpus, and by performing text analysis, it extracts citation information directly from their full text.</p> <p>The current version of the dataset contains&nbsp;<em>~4.3M citations</em> made by approximately <em>211K open access Computer Science conference or workshop papers</em> that, according to DBLP, do not have a DOI. The DBLP snapshot used for this version was the one released on <em>September 2025</em>.&nbsp;</p> <h2>Dataset files</h2> <h3>1. Core Non-DOI Citation Dataset - bip_ndr_{version}.tar.gz</h3> <p>The dataset is formatted as a JSON Lines (JSONL) file (one JSON Object per line) to facilitate file splitting and streaming.&nbsp;</p> <p>Each JSON object has three main fields:</p> <ul> <li> <p>&ldquo;_id&rdquo;: a unique identifier,</p> </li> <li> <p>&ldquo;citing_paper&rdquo;, the &ldquo;dblp_id&rdquo; of the citing paper,</p> </li> <li> <p>&ldquo;cited_papers&rdquo;: array containing the objects that correspond to each reference found in the text of the &ldquo;citing_paper&rdquo;; each object may contain the following fields:</p> <ul> <li> <p>&ldquo;dblp_id&rdquo;: the &ldquo;dblp_id&rdquo; of the cited paper. Optional - this field is required if a &ldquo;doi&rdquo; is not present.</p> </li> <li> <p>&ldquo;doi&rdquo;: the doi of the cited paper. Optional - this field is required if a &ldquo;dblp_id&rdquo; is not present.</p> </li> <li> <p>&ldquo;bibliographic_reference&rdquo;: the raw citation string as it appears in the citing paper.</p> </li> </ul> </li> </ul> <p>Changes from previous version:</p> <ul> <li>Added more papers from DBLP.</li> </ul> <h3>2. Citation Intents Dataset - bip_ndr_ci_{version}.tar.gz</h3> <p>This file enriches the BIP! NDR dataset with citation-level intent classification.<br>It preserves the same base structure of the previous file, while adding a nested array of "citations" with each element of "cited_papers".</p> <p>Each "citation" provides the local textual context, section, and intent of the citation in the following format:</p> <ul> <li>"citation_id": Unique identifier in the format {citing_id}&gt;{cited_id}_CIT{index} linking the citing and cited entities.</li> <li>"section": The section of the citing paper where the citation occurs (e.g., Introduction, Methods, Results).</li> <li>"intent": Inferred purpose of the citation based on textual context (see classification schema below).</li> </ul> <p>The "intent" field follows the SciCite classification schema, which categorizes citations into three high-level functional types:</p> <ol> <li>background information: The citation states, mentions, or points to the background information giving more context about a problem, concept, approach, topic, or importance of the problem in the field.</li> <li>method: Making use of a method, tool, approach or dataset.</li> <li>results comparison: Comparison of the paper's results/findings with the results/findings of other work.</li> </ol> <p>The classification is done with the <a href="https://huggingface.co/sknow-lab/Qwen2.5-14B-CIC-SciCite">Qwen2.5-14B-CIC-SciCite fine-tuned Large Language Model, published by Athena RC</a>.&nbsp;</p> <p>Changes from previous version:&nbsp;</p> <ul> <li>Added more papers with intent</li> </ul>

opencc-zeroMay 2023View details →
zenodo44/100

Co-citations Map: Literature review on Learning Ecologies, 1991-2018

<p>The co-citation map was adopted as&nbsp;type of analysis over 85 papers sampled from five scientific databases (pls. cfr: <a href="https://zenodo.org/record/1503775#.W_rUn-hKg2w">Systematic Review on the Research Topic &quot;Learning Ecologies&quot; - Dataset and Analysi</a>s&quot;. This was adopted as method to further understand the relationships and advancement of work relating the concept of LE.</p> <p>The software tool to carry out this phase of the study was <a href="http://www.citnetexplorer.nl/ ">Citenet Explorer</a>&nbsp;for the analysis and visualization of co-citations along the period 1999-2018.</p> <p>The total number of citations and the relationships between most cited authors and all authors were extracted from the corpus analyzed (<a href="https://zenodo.org/record/1503775#.W_rUn-hKg2w">Systematic Review on the Research Topic &quot;Learning Ecologies&quot; - Dataset and Analysi</a>s&quot;; a specific dataset was created and the outputs were processed by the specialized software that delivers the bibliometric maps as output.</p> <p>All the instructions to use the two files that compose the dataset requested by the software, as well as the software specifications can be found at CiteNet Explore page:&nbsp;http://www.citnetexplorer.nl/&nbsp;</p> <p>Finally, the Fig. 7 with the analysis we performed shows the co-citations map, where at a first sight it is possible to see two main groups of nodes or authors cited (X axis) across a timespan (Y axis), and sparse elements at the center and at the beginning of the period (1991). The main and more compact group (also in terms of clusterisation of nodes, which are showed in green) related the publications that cite the seminal work of Barron (2004). These publications are mostly classified as using socio-constructivist theories and are placed in the area of social sciences, while for other categories (such as the methodological approach, research applications and the alignment are more fragmented). &nbsp;There are four seminal works (Abd-El-Khalick &amp; Akerson, 2004; Barron, 2004; T ; Okamoto, Kayama, Inoue, &amp; Cristea, 2002; Toshio Okamoto &amp; Kayama, 2004) to which other papers can be connected. Beyond the mentioned work of Barron, the other three works can be placed in the area of technology (development of eLearning environments) and STEM education, supporting the idea of disciplinary fragmentation.</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

Coverage of DOAJ journals' citations through OpenCitations - Result DataSet

<p>The dataset contains:&nbsp;</p> <ul> <li><strong>by_journal.json</strong>: a file containing all information extracted by Open Citations about DOAJ journals divide by year and journal name. Inside the file, the metadata about the journal are:&nbsp; <ul> <li>ISSN</li> <li>EISSN</li> <li>number of articles overall in the journal</li> <li>subject(s)&nbsp;</li> <li>number of citations received</li> <li>number of citations done</li> <li>ratio between citations done and received</li> <li>number of citations received from DOAJ journals</li> <li>number of citations done to DOAJ journals</li> <li>ratio between citations done to and received from DOAJ journals.</li> </ul> </li> </ul> <ul> <li><strong>normal.json</strong>: a file containing all information extracted from Open Citations about DOAJ journals divided only by year. Inside the file, the data by year are: <ul> <li>number of citations received.</li> <li>number of citations done.</li> <li>ratio between citations done and received.</li> <li>number of self-citations made by DOAJ inside Open Citations.</li> <li>ratio between the self-citation and the total citations received and done by DOAJ.</li> </ul> </li> </ul> <ul> <li><strong>errors.json</strong>: a file containing the count of all errors obtained from computations. Inside the file: <ul> <li>errors about records that don&#39;t have any specified date (null dates).</li> <li>errors about records that have impossible dates (wrong dates).</li> <li>errors about articles that don&#39;t have any specified Dois.</li> <li>errors about Open Citations records that don&#39;t have any Dois in the citing or cited fields.</li> </ul> </li> <li><strong>DOAJ_metrics.json</strong>: a file containing metrics about DOAJ and Open Citations, obtained by computations. Inside the file are these fields: <ul> <li>number of journals with dois.</li> <li>number of articles which have been processed during computations.</li> <li>number of used Dois. All dois (with no repetition) which are used for the adding journal operation.</li> <li>number of repeated Dois. All dois which are repeated inside the same or in another journal.</li> <li>number of accepted Dois. All articles (with repetition) which have both a defined journal and a defined doi.</li> </ul> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Dataset Citation and Re-use Data

<p>This dataset includes processed citation data for datasets recorded in OpenAlex as of May 2022. It identifies self-citations to these datasets at the individual, institutional, and country level, and includes domain classifications of the citing works using the Science-Metrix classifications.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Exploring the Impact of Neuroscience Preprints: A Citation Analysis

<p>1.&nbsp;Neuroscience_Records_Contain_Reference_to_Preprints.Scopus.V3.xlsx</p> <p>This Excel file contains the titles, DOIs, references, and EIDs of those Neuroscience publications (journal articles, books/book chapters, conference papers, notes, etc.) from 2004 to 2022 that have at least one reference to a preprint. For example, if a&nbsp;Neuroscience journal article has 40 references and one of these references is a preprint, then it&#39;s included in this Excel file. These records are retrieved from Scopus through the following query:</p> <p>REFSRCTITLE ( &quot;OSF Preprints&quot; OR &quot;open science foundation Preprints&quot; OR *africarxiv* OR *agrixiv* OR *arabixiv* OR *arxiv* OR *biohackrxiv* OR *biorxiv* OR *bodoarxiv* OR *cogprints* OR *eartharxiv* OR *ecoevorxiv* OR *ecsarxiv* OR *edarxiv* OR *engrxiv* OR *frenxiv* OR &quot;INA-Rxiv&quot; OR *indiarxiv* OR *lawarxiv* OR &quot;LIS Scholarship Archive&quot; OR *marxiv* OR *mediarxiv* OR *metaarxiv* OR mindrxiv OR *nutrixiv* OR paleorxiv OR &quot;Preprints.org&quot; OR psyarxiv OR *repec* OR *socarxiv* OR *sportrxiv* OR &quot;Thesis Commons&quot; OR &quot;CoP preprint&quot; OR &quot;FocUS Archive preprint&quot; OR &quot;PeerJ preprint&quot; OR &quot;Law Archive preprint&quot; OR *medrxiv* ) AND SUBJAREA ( neur ) AND PUBYEAR &lt; 2023</p> <p>&nbsp;</p> <p>2.&nbsp;ReferencesToPreprints.V3.txt</p> <p>References of the publications are split through a Python code (SplitReferences.py) and organized into separate lines in a text file. For example, if a publication has 40 references, all of these 40 references are split into 40 separate lines. After splitting references, those lines containing one of these words/terms (&quot;OSF Preprints&quot; OR &quot;open science foundation preprints&quot; OR africarxiv OR agrixiv OR arabixiv OR arxiv OR biohackrxiv OR biorxiv OR bodoarxiv OR cogprints OR eartharxiv OR ecoevorxiv OR ecsarxiv OR edarxiv OR engrxiv OR frenxiv OR &quot;INA-Rxiv&quot; OR indiarxiv OR lawarxiv OR &quot;LIS Scholarship Archive&quot; OR marxiv OR mediarxiv OR metaarxiv OR mindrxiv OR nutrixiv OR paleorxiv OR &quot;Preprints.org&quot; OR psyarxiv OR repec OR socarxiv OR sportrxiv OR &quot;Thesis Commons&quot; OR &quot;CoP preprint&quot; OR &quot;FocUS Archive preprint&quot; OR &quot;PeerJ preprint&quot; OR &quot;Law Archive preprint&quot; OR medrxiv) are selected (through RetrieveLinesContainingSpeceficString.py) and organized into this text file (ReferencesToPreprints.V3.txt). Each reference contains an EID (separated by &quot;;&quot;) in order to specify which publication contains this specific reference.</p> <p>After this step, through a Python code (AddPreprintServerToEndOfLines.py) the name of a certain preprint was added to the end of each line. For example, if a line (or a reference) contains &quot;biorxiv&quot;, the word &quot;biorxiv&quot; will be added to the end of this line after the &quot;@&quot; sign.</p>

opencc-by-4.0Sep 2023View details →
edi44/100

Data and code for EDI overview paper, data collection characteristics, FAIR evaluation, downloads, and citations

The Environmental Data Initiative (EDI) is a trustworthy, stable data repository and data management support organization for the environmental scientist. EDI provides tools and support that allow the environmental researcher to easily integrate data publishing into the research workflow. Almost ten years since going into production, these data and code were used to provide a general description of EDI’s collection of data and its data management philosophy and placement in the repository landscape. They show how comprehensive metadata and the repository infrastructure lead to highly findable, accessible, interoperable, and reusable (FAIR) data by evaluating compliance with specific community proposed FAIR criteria. Finally, they provide measures and patterns of data (re)use, assuring that EDI is fulfilling its stated premise.

openCC0Aug 2022View details →
zenodo40/100

Pre- and post-publication citations to published arXiv preprints

<p>This dataset contains citations to published preprints, both before they are published and after they are published. Details of the data are provided in the <code>README.md</code>.</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

PatCit: A Comprehensive Dataset of Patent Citations

<p><em><strong>patCit:&nbsp;A Comprehensive Dataset of Patent Citations</strong></em>&nbsp;[<a href="https://tinyletter.com/patcit">Newsletter</a>,&nbsp;<a href="https://github.com/cverluise/PatCit">GitHub</a>]</p> <p>Patents are at the crossroads of many innovation nodes: science, industry, products, competition, etc. Such interactions can be identified through citations&nbsp;<em>in a broad sense</em>.</p> <p>It is now common to use front-page patent citations to study some aspects of the innovation system. However, <strong>there is much more buried in the Non Patent Literature (NPL) citations and in the patent text itself</strong>.&nbsp;<strong>patCit extracts and structures these citations.</strong></p> <blockquote> <p>Want to know more? Read patCit&nbsp;<a href="https://docs.google.com/presentation/d/11COlz64EZn8PipXvnDBBZI_bnDD0fpm6tyx1_EqD6lU/edit?usp=sharing">academic presentation</a>&nbsp;or dive into usage and technical guides on patCit&nbsp;<a href="https://cverluise.github.io/PatCit/">documentation website</a>.</p> </blockquote> <p><strong><em>IN PRACTICE</em></strong></p> <p>At patCit, we are building a&nbsp;<em>comprehensive</em>&nbsp;dataset of patent citations to help the community explore this&nbsp;<em>terra incognita</em>. patCit has the following features:</p> <ul> <li>global coverage</li> <li>front-page and in-text citations</li> <li>all categories&nbsp;of NPL documents</li> </ul> <p><strong><em>Front-page</em></strong></p> <p>patCit builds on&nbsp;<a href="https://www.epo.org/searching-for-patents/data/bulk-data-sets/docdb.html#tab-1">DOCDB</a>, the largest database of Non Patent Literature (NPL) citations. First, we deduplicate this corpus and organize it into 10 categories (bibliographical reference, database, norm &amp; standard, etc). Then, we design and apply category specific information extraction models using&nbsp;<a href="https://github.com/explosion/spaCy">spaCy</a>. Eventually, when possible, we enrich the data using external domain specific high quality databases (e.g. Crossref for bibliographical references).</p> <p><strong><em>In-text</em></strong></p> <p>patCit builds on Google Patents corpus of&nbsp;<a href="https://console.cloud.google.com/bigquery?project=patcit-public-data&amp;p=patents-public-data&amp;d=patents&amp;t=publications&amp;page=table">USPTO full-text patents</a>. First, we extract patent and bibliographical reference citations. Then, we parse detected in-text citations into a series of category dependent attributes using&nbsp;<a href="https://github.com/kermitt2/grobid">grobid</a>. Patent citations are matched with a standard publication number using the Google Patents&nbsp;<a href="https://patents.google.com/api/match">matching API</a>&nbsp;and bibliographical references are matched with a DOI using&nbsp;<a href="https://github.com/kermitt2/biblio-glutton">biblio-glutton</a>. Eventually, when possible, we enrich the data using external domain specific high quality databases (e.g. Crossref for bibliographical references).</p> <p>&nbsp;</p> <p><strong>FAIR</strong></p> <p><strong>Find</strong>&nbsp;- The patCit dataset is available on&nbsp;<a href="https://console.cloud.google.com/bigquery?project=patcit-public-data&amp;p=patcit-public-data&amp;page=project">BigQuery</a>&nbsp;in an interactive environment. For those who have a smattering of SQL, this is the perfect place to explore the data. It can also be downloaded on&nbsp;<a href="https://zenodo.org/record/3710994">Zenodo</a>.</p> <p><strong>Interoperate</strong>&nbsp;- Interoperability is at the core of patCit ambition. We take care to extract unique identifiers whenever it is possible to enable data enrichment for domain specific high quality databases. This includes the DOI, PMID and PMCID for bibliographical references, the Technical Doc Number for standards, the Accession Number for Genetic databases, the publication number for PATSTAT and Claims, etc. See specific table for more details.</p> <p><strong>Reproduce</strong>&nbsp;- Our <a href="https://github.com/cverluise/PatCit">gitHub</a> repository is the project factory. You can learn more about data recipes and models on the patCit&nbsp;<a href="https://cverluise.github.io/PatCit/">documentation website</a>.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Harvard Citation - Faculty of Medicine UNS

<p>Harvard Citation, Faculty of Medicine UNS version, is a dataset format for Mendeley referencing, which can be used for student&#39;s thesis.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Data for "Open Access impact on citations: a case study"

<p>This dataset is a list of 347 papers published in 2010 and retrieved from the Web of Science, Scopus and Google Scholar. For each paper, the number of citations and the citation date(s) have been collected. If the full-text is available online, the date of &quot;liberation&quot; and the URL of the file have been retrieved as well. The objective was to assess the impact of Open access on citation rate and more particularly the impact before and after&nbsp;full-text&nbsp;&quot;liberation&quot;.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2016View details →
zenodo40/100

What students answer when discussing about citation practices

<p>This document explain how data were generated and how to interpret them.   </p> <p> </p> <p>LICENSE: CC0<br> But if you want to combine data with other datasets, feel free to use them as if they were published under CC0 license.   <br> Data were published in February 2017. At that time, Zenodo only provided CC BY, CC BY-SA, CC BY-NC, CC BY-ND and CC BY-NC-ND. No CC0 option was available.</p> <p> </p> <p>HOW DATA WERE COLLECTED<br> The 21 recorded sessions took place between February 2013 and December 2016.   <br> Data were collected using Turning Technologies' remote controls (called *clickers*) and TurningPoint software.</p> <p>The 4 versions of the quiz used during these 4 years are provided in the 'quizzes' folder for information purpose (in PDF and Powerpoint formats).</p> <p>Turning Technologies records data in a closed format (.tpzx) that can be exported and converted them into 3 formats provided here (these 3 files contain the same data):</p> <p>* Excel (.xslx)<br> * Comma-spearated values (.csv)<br> * SQLite (.sqlite)</p> <p>The first one was directly exported from TurningPoint and is provided for Excel users who can't read CSV correctly.   <br> CSV was converted from Excel and is provided for non-Excel users.   <br> Finally, SQLite is provided in order to apply different sorting and filters to the data. It can be read using SQLite manager for Firefox ([https://addons.mozilla.org/en-US/firefox/addon/sqlite-manager/](https://addons.mozilla.org/en-US/firefox/addon/sqlite-manager/)).</p> <p> </p> <p>CODEBOOK<br> Here is the name, the meaning and the possible values of the columns (name - meaning [possible values]). If students didn't answer the question, the value is '-'.   </p> <p>Session - session number (chronological) [1 to 21]<br> AcademicYear - academic year [12-13, 13-14, 14-15, 15-16, 16-17]<br> Year - calendar year [2013, 2014, 2015, 2016]<br> Month - month (number) [1 to 12]<br> Day - day (number) [1 to 31]<br> Section - section abbreviation [CH, ESC, GM, IF, SIE, SV]<br> Level - students' level [BA2, BA3, MA]<br> Language - course's language [FR or EN]<br> DeviceID - clicker's ID [(unique ID within a session)]<br> Q1 - answers to question 1 [A, B, C, D, E]<br> Q2 - answers to question 2 [A, B, C, D]<br> Q3 - answers to question 3 [A or B]<br> Q4 - answers to question 4 [A or B]<br> Q5 - answers to question 5 [A or B]<br> Q6 - answers to question 6 [A or B]<br> Q7 - answers to question 7 [A or B]<br> Q8 - answers to question 8 [A or B]<br> Q9 - answers to question 9 [A or B]<br> Q8-9 - answers to the question 8-9 (merge) [A or B]<br> Q10 - answers to question 10 [1, 2]<br> Q11 - answers to question 11 [A or B]<br> Q12 - answers to question 12 [A, B]</p> <p>Section abbreviation meaning<br> * CH: chemistry<br> * ESC: school of criminal justice (Unil)<br> * GM: mechanical engineering<br> * IF: financial engineering<br> * SIE: environmental engineering<br> * SV: life sciences</p> <p>Level meaning  <br> * BA2: 2nd year of Bachelor<br> * BA3: 3rd year of Bachelor<br> * MA: Master level</p> <p>Question types<br> For some questions, multiple answers were allowed: Q1, Q2, Q10 &amp; Q12.   <br> Half of the questions have only one correct answer, true or false: Q3, Q5, Q6, Q7, Q8, Q9 &amp; Q8-9.   <br> Finally, for 2 questions only one answer was accepted, but there is not only one correct answer: Q4 &amp; Q11.</p> <p> </p> <p>INFORMATION ABOUT THE SESSIONS<br> Except otherwise stated below, all sessions were conducted like the original one: Q1 to Q12 (no Q8-9).<br> The original French version of the quiz has been translated into English for a few sessions with Master students.<br> For sessions 14 and 20, Q5 was removed and Q8 &amp; Q9 were merged in Q8-9.   <br> Session 18 was a short one with only 7 sevens questions: Q1, Q2, Q3, Q4, Q6, Q7 &amp; Q9.   </p> <p> </p> <p>CONTACT INFORMATION<br> If you have any question about these data, contact formations.bib@epfl.ch.</p> <p> </p>

opencc-by-4.0Feb 2017View details →
zenodo40/100

Eco-hydrology Cikapundung Project: citation connections and research profile building

<p>This image is uploaded as an integrated part of Eco-Hydrology Cikapundung project. This image will be cited across all future publications related to this project as CC-BY image. Therefore it should not be treated as prior publication of any kind.</p> <p>We used www.Draw.io and the source code is available on GIthub (https://github.com/dasaptaerwin/CikapundungProject/blob/master/citationConnection.xml).</p> <p>---<br> Dokumen ini disusun sebagai pelengkap riset untuk menggambarkan kaitan sitasi antar dokumen dari hulu ke hilir. Setiap dokumen diupayakan ber-DOI agar dapat <em>autosync</em> dengan profil riset yang tersedia: Google Scholar, Sinta, ORCID. Dengan dibuatnya dokumen hubungan sitasi ini, maka diharapkan dapat menjelaskan bahwa tidak terjadi duplikasi dalam publikasi.</p>

opencc-by-4.0Jun 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record