Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
612
datasets available to search
ShareScore release 0.9.0
Dataset results
612 results for “citing”
Wilcoxon Rank Sum Test and Keyphrase Extraction Data Cited in "What Everyone Says: Public Perceptions of the Humanities in the Media"
<p>This repository contains Wilcoxon rank sum test and keyphrase extraction data cited in the WhatEvery1Says (WE1S) Project's article "What Everyone Says: Public Perceptions of the Humanities in the Media". The organization of the materials is discussed below.</p> <p><strong>Wilcoxon Rank Sum Test</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>wilcoxon-tests</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. The Wilcoxon rank sum test identifies specific words that appear significantly more in one group of documents as compared to another, thus providing researchers with an understanding of what words are “distinctive” to each group. Further information on WE1S's use of Wilcoxon rank sum testing can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf</a>.</p> <p>Each subdirectory in the <code>wilcoxon-test</code> folder contains the data and results of a particular comparison experiment based on a metadata category such as whether the data contained articles published by public or private institutions. Each data file is a <code>.txt</code> file representing a sample of the overall data from the collection. The <code>README</code> file provides information on the collection used, the sample size, and the nature of the comparison. The results for the test are in a file called <code>results.csv</code>.</p> <p>The <code>results.csv</code> file for each test includes a row for each term included in the test. Each row displays the term, the term's raw count in each category compared (count 1 and count 2), the difference between the 2 counts (count 1 minus count 2), the percentage change in the counts, the Wilcoxon statistic, and the Wilcoxon p-value. Sorting the csv by the Wilcoxon stat from greatest to least will cause the terms most strongly associated with category 1 to come to the top (category 1 is the category listed first in the title field of the README.md file for each test), while sorting it by the Wilcoxon stat from least to greatest will cause the terms most strongly associated with category 2 to come to the top (category 2 is the category listed second). The p-value column provides you with information about how confident you can be about each comparison's significance.</p> <p><strong>Keyphrase Extraction</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>keyphrase-extraction</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. Keyphrase extraction generates a list of the most significant words or phrases (1-6 words long) within individual documents. WE1S takes the top ten keyphrases in each document and ranks them according to their frequency across the collection. WE1S uses the SGRank algorithm for keyphrase extraction, and because this algorithm is computationally intensive, WE1S limits keyphrases to lemmatized nouns and proper nouns within a window of 70 words to either side of candidate keyphrases. Further information on WE1S's use of keyphrase extaction can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf</a>.</p> <p>Each subdirectory in the <code>keyphrase-extraction</code> folder contains the data and results of keyphrase extraction on a particular collection. Details of the collection and resulting files can be found in each subdirectory. Each list of keyphrases is in a file called <code>SGRank.csv</code>, which lists the keyphrases and their number of occurrences in the collection. The article additionally cites keyphrases that are shared with the terms in the public topic model produced by Andrew Goldstone and Ted Underwood, “The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us,” <em>New Literary History</em> 45, no. 3 (2014): 359–84, <a href="https://doi.org/10.1353/nlh.2014.0025">https://doi.org/10.1353/nlh.2014.0025</a>. The list of terms is derived from the public visualization at <a href="https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words">https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words</a>. Keyphrases extracted from WE1S data were split into single-word terms and compared with the list of vocabulary in Goldstone and Underwood's word list (<code>quiet_transformations_wordlist.txt</code>) to compile lists of shared vocabulary. These lists are given in files called <code>shared_terms.txt</code>.</p> <p>Note that keyphrases were extracted for corpora produced using the Python <a href="https://textacy.readthedocs.io/en/latest/index.html">Textacy</a> library. Because these corpora contain the full text of articles with intellectual property restrictions they cannot be reproduced here.</p>
Motivations for citing research in comments to YouTube videos
<p><strong>This dataset includes 300 YouTube comments that link to research publications or preprints. Each comment has been assigned into one category that define why individuals mention scholarly publications in comments to YouTube videos. </strong></p> <p>Each row of the file "Categories_300random.xlsx" represents one distinct comment to YouTube video.<br> The file includes the following columns:</p> <ul> <li><strong>Comment_text </strong>- the full text of a comment,</li> <li><strong>YouTube_link </strong>- URL to the video where the comment was left,</li> <li><strong>Comment_link </strong>- URL to the thread with the comment,</li> <li><strong>Category </strong>- the final category that was assigned to the comment,</li> <li>Columns <strong>Category_old_schema_SS</strong>, <strong>Category_old_schema_IP</strong>, <strong>Category_old_schema_OZ </strong>- categories assigned by different researchers according to old categorization schema and were used for validation reasons,</li> <li>Columns <strong>Category_OZ</strong>, and <strong>Caregory_LB </strong>- categories assigned by different researchers according to the final schema and were used for validation.<br> <br> </li> </ul>
Complementary dataset of Overton metadata on citing policy-related documents for the study "From intent to impact: Investigating the effects of open sharing commitments"
<p>This document provides the underlying dataset for the bibliometric component for the 2022 study "From intent to impact: Investigating the effects of open sharing commitments" by Research Consulting and Science-Metrix.</p> <p>Before reproducing the study findings or re-using the underlying datasets for other purposes, please cautiously review their limitations in the study's technical annex and main report, available at: https://zenodo.org/communities/data-sharing-in-public-health-emergencies/ </p> <p>Special thanks from the Science-Metrix / Elsevier teams to Euan Adie and Overton for this exceptional public release of Overton metadata, and for conducting extraordinary data collection to retrieve citations towards arXiv preprints.</p> <p> </p> <p>Scope: note that this file combines cited journal publications and preprints from the Covid19, HVRD, Zika and HVVD thematic sets.</p> <p>Data treatment: this data is intend foremost to provide manual validation or qualitative triangulation of our findings. No special efforts have been made to process and clean the data for its eventual re-use in secondary analysis or text mining approaches.</p> <p>Definitions used in this table:</p> <table> <tbody> <tr> <td>Column name </td> <td>Definition</td> </tr> <tr> <td>document_type</td> <td>preprint or journal publication</td> </tr> <tr> <td>doi</td> <td>digital object identifier</td> </tr> <tr> <td>arxiv_id</td> <td>arXiv preprint server's unique identifier for its preprints</td> </tr> <tr> <td>ssrn_id</td> <td>SSRN preprint server's unique identifier for its preprints. Note that some of these IDs are contained within the DOIs also assigned to some (but not all) SSRN preprints , in the form of "10.2139/ssrn." + 'ssrn_id'</td> </tr> <tr> <td>coalesce_id</td> <td>coalesce function applied to the DOI, arxiv_id and ssrn_id. Redundant for journal publications.</td> </tr> <tr> <td>policy_source_title</td> <td>name of the policy-related organization</td> </tr> <tr> <td>published_on</td> <td>publication date of the citing policy-related document</td> </tr> <tr> <td>title</td> <td>title of the citing policy-related document</td> </tr> <tr> <td>pdf_url</td> <td>URL for the online version of the policy-related document</td> </tr> <tr> <td>snippet</td> <td>Where available, excerpt of the text immediatly before and after the citation to a journal publication or preprint found in the citing policy-related document</td> </tr> </tbody> </table> <p> </p>
Citation network dataset covering the work of RP Millar and its citing literature
<p>The following describes the citation network datasets that underpins the manuscript “A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research” [1].</p> <p><strong>Data collection</strong></p> <p>We retrieved data from the Web of Science Core Collection under the University of Edinburgh’s subscription in January 2024. We sought to retrieve all indexed papers of Professor Robert P. Millar (RPM). We searched the following <em>AU = (Millar, R)</em>, and then retrieved records that corresponded to his WoS profile (n=428 records) and an additional 49 paper that were authored by Robert but had not been included in his WoS record – validating the records against a CV of his published works.</p> <p>We retrieved the full citation history as record by Web of Science to these papers from other indexed records. The 477 RPM papers had been cited 21,677 times by 11,138 documents by date of retrieval, and removing self-citations left 19,256 citations by 10,719 documents. We then retrieved all metadata from WoS concerning the 477 RPM papers and the 10,719 citation papers, resulting in a dataset covering 11,196 documents.</p> <p><strong>Citation network dataset</strong></p> <p>We constructed a citation network dataset by parsing data from each paper’s full bibliography consisting of:</p> <p>i. ‘Edge-list’ that records citation links from a citing to a cited document. This is constructed by assigning unique IDs to each retrieved paper and to every unique reference string contained in their bibliographies. The edge list is composed of a ‘Source’ column that contains the ID of the <em>citing</em> document and a ‘Target’ column containing the IDs of its citations, with one record per row. Given that we were only interested in citations between the WoS retrieved documents, we discarded any reference string that represented a document outwith our search.</p> <p>ii. ‘Node-attribute list’ that contains the ID, with relevant metadata contained in adjacent columns to identify documents, including authors, title of publication, journal, year of publication. We also parsed into this dataset the WoS full citation count for each paper and the total number of references in the bibliographies of each paper.</p> <p>This results in a dataset containing 11,196 nodes and 115,834 edges between nodes. We removed a total of 67 papers for which metadata was incomplete and/or corrupted. We further focussed on the largest interconnected component, removing nodes with no connections (isolates) or smaller components that were detached from the main network. We excluded papers <10 references to remove meeting abstracts and other minor journal items, and papers not published in English. This resulted in a final dataset containing 10,901 nodes and 113,742 edges, and it is this dataset that we share as it is the basis for the analyses within the paper.</p> <p><strong>Description of dataset variables</strong></p> <p><strong>‘RPM_Edgelist.csv’</strong> is a comma-separate values file that consists of all 113,742 citations between the 10,901 documents of the citation network analysed in the manuscript. The columns refer to:</p> <ul> <li>‘<em>Source</em>’, the unique identifier for the <em>citing </em>document</li> <li>‘<em>Target’</em>, the unique identifier for the <em>cited</em> document</li> <li>‘<em>Syr</em>’, the year of publication of the <em>citing</em> document</li> <li>‘<em>Tyr</em>’, the year of publication of the <em>cited</em> document</li> <li>‘<em>SC</em>’, the cluster ID of the <em>citing</em> document</li> <li>‘<em>TC</em>’, the cluster ID of the <em>cited </em>document</li> </ul> <p><strong>‘RPM_Nodelist.csv’</strong> is a comma-separate values file that consists of the 10,901 documents of the citation network analysed in the manuscript. The columns refer to:</p> <ul> <li>‘<em>Id</em>’, the unique ID assigned to a document that corresponds with the edgelist</li> <li>‘<em>Reference string</em>’, the reference string of the document</li> <li>‘<em>WoS ID</em>’, the unique accession number assigned to a document by the Web of Science. These can be used to query WoS to find further data on all papers via the ‘UT= ’ field tag.</li> <li>‘<em>Authors</em>’, all authors formatted by full last name and initials</li> <li>‘<em># of authors’</em>, number of authors</li> <li>‘<em>Title</em>’, title of document</li> <li>‘<em>Publication year</em>’, publication year of document</li> <li>‘<em>Document type</em>’, document type defined by WoS (e.g. article, review, etc.)</li> <li>‘<em>Total references</em>’, total number of references within a documents bibliography as recorded by WoS</li> <li>‘<em>Total WoS citations</em>’, total number of citations recorded to a document from other documents indexed in the Web of Science</li> <li>‘<em>Indegree</em>’, total number of within network citations (i.e. counting only citations from other papers retrieved by our query)</li> <li>‘<em>Outdegree</em>’, total number of within network references (i.e. counting only reference to other papers retrieved by our query)</li> <li>‘<em>Degree</em>’, total number of node connections (i.e. indegree + outdegree)</li> <li>‘<em>Class</em>’, variable used to distinguish between RPM’s publications (‘RPM’) and the citing documents (‘CITE’)</li> <li>‘<em>Cluster</em>’, provides the cluster membership number as discussed within the manuscript. This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.67 | 25 clusters).</li> </ul> <p><strong>References</strong></p> <p>[1] Leng, R. I., Leng. G. (Under review). A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research. <em>J. Neuroendocrinol</em></p> <p>All bibliographic data included in this study are derived originally from Clarivate™ (Web of Science™) and downloaded in January 2024. © Clarivate 2024. All rights reserved. </p>
Speadsheet data cited in Annex 2 of Deliverable 4.1 of project ASTRail
<p>This repository contains the spreadsheet for the "Ranking Matrix" computations,<br> as mentioned in the Annex 2 of Deliverable 4.1 "Report on Analysis and on Ranking of Formal Methods"<br> of the project ASTRail (http://www.astrail.eu).</p>
Dataset for "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations"
<p>This is the dataset for the paper "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations" submitted to iConference 2025.</p>
CITES Trade Shipments Database
Reprocessed records database from CITES in parquet form. Originally provided at trade.cites.org.
HTML Status Codes of Publications citing ICPSR
<p>This dataset contains the HTML status codes of ICPSR citing literature, which were tested in 2020 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>
PBMC CITE-seq reference
<p>This PBMC CITE-seq reference object was constructed using Seurat v5.</p>
Dataset for: Do you Cite What you Tweet? Investigating the relationship between tweeting and citing research articles
<p>This dataset was used for the work "Do you cite what you tweet? Investigating the relationship between tweeting and citing research articles", submitted to the Quantitative Science Studies Journal. The accompanying R script was used for the logistic regression model in the paper.</p> <p> </p> <p> </p>
For Articles Published in 1995: Evolution Of The Top 50 Most Cited
<p>ascii data files (cleaned, calibrated, processed) and journal-level figure figure (in vector graphics format) shown in the the youtube video "The Top 50 Most Cited Articles - The Short and Long of It" at <a href="https://www.youtube.com/watch?v=4wyy80QZ0lg">https://www.youtube.com/watch?v=4wyy80QZ0lg</a></p>
Data file for paper: Highly-cited papers in software engineering: the top-100
<p>Data file for paper: Highly-cited papers in software engineering: the top-100</p> <p>https://doi.org/10.1016/j.infsof.2015.11.003</p>
Data from: Persistence of distinctive morphotypes in the native range of the CITES-listed Aldabra giant tortoise
Understanding the extent of morphological variation in the wild population of Aldabra giant tortoises is important for conservation, as morphological variation in captive populations has been interpreted as evidence for lingering genes from extinct tortoise lineages. If true, this could impact reintroduction programmes in the region. The population of giant tortoises on Aldabra Atoll is subdivided and distributed around several islands. Although pronounced morphological variation was recorded in the late 1960s, it was thought to be a temporary phenomenon. Early researchers also raised concerns over the future of the population, which was perceived to have exceeded its carrying capacity. We analyzed monthly monitoring data from 12 transects spanning a recent 15-year period (1998–2012) during which animals from four subpopulations were counted, measured, and sexed. In addition, we analyzed survival data from individuals first tagged during the early 1970s. The population is stable with no sign of significant decline. Subpopulations differ in density, but these differences are mostly due to differences in the prevailing vegetation type. However, subpopulations differ greatly in both the size of animals and the degree of sexual dimorphism. Comparisons with historical data reveal that phenotypic differences among the subpopulations of tortoises on Aldabra have been apparent for the last 50 years with no sign of diminishing. We conclude that the giant tortoise population on Aldabra is subject to varying ecological selection pressures, giving rise to stable morphotypes in discrete subpopulations. We suggest therefore that (1) the presence of morphological differences among captive Aldabra tortoises does not alone provide convincing evidence of genes from other extinct species; and (2) Aldabra serves as an important example of how conservation and management in situ can add to the scientific value of populations and perhaps enable them to better adapt to future ecological pressures.
Anatomy of top 1% most highly-cited publications. An empirical comparison of two approaches. Dataset
<p>Supplementary material containing tables with the main pieces of data used in the publication entitled Anatomy of top 1% most highly-cited publications. An empirical comparison of two approaches. This pieces of data were downloaded from the November 2022 snapshot of OpenAlex.</p>
Cites received by Brazilian Journal of Information Science: research trends (BRAJIS)
<p>Dataset with the data of the paper "BRAJIS en WOS: impacto observado y visibilidad global"</p>
FIGURES 37–40 in The Hemiptera-Sternorrhyncha (Insecta) of Hong Kong, China-an annotated inventory citing voucher specimens and published records
FIGURES 37–40, Rhachisphora takahashii sp. n. (Aleyrodidae, Aleyrodinae), holotype puparium. (37) submedial dorsum of metathorax and abdominal segments I-IV to show chaetotaxy and geminate pore / porettes. (38) vasiform orifice and eighth abdominal setae. (39) thoracic tracheal opening at margin. (40) caudal setae, caudal tracheal opening at margin and posterior part of caudal furrow.
FIGURES 41–43 in The Hemiptera-Sternorrhyncha (Insecta) of Hong Kong, China-an annotated inventory citing voucher specimens and published records
FIGURES 41–43, Rhachisphora spp. (Aleyrodidae, Aleyrodinae). (41) R. takahashii sp. n., post-emergence male pupal case, entire puparium, particularly to show rhachis with 5 pairs of abdominal lateral arms, and submarginal geminate pore / porettes. (42) R. takahashii sp. n., detail of basal parts of three lateral abdominal rhachis arms, to show dentate anterior edges. (43) R. maesae Takahashi, original drawing after Takahashi (1932).
FIGURES 25–30 in The Hemiptera-Sternorrhyncha (Insecta) of Hong Kong, China-an annotated inventory citing voucher specimens and published records
FIGURES 25–30, Coccoidea. (25) Cribropulvinaria tailungensis (Coccidae), adult females and nymphs on Aporusa dioica. (26) Coccus formicarii (Coccidae), spherical mature females on bark of Schefflera heptaphylla (originally covered over with debris by ants). (27) Fistulococcus pokfulamensis (Coccidae), microscope slide preparation of adult female, showing glandular structures responsible for secretion of waxy material (28) Fistulococcus pokfulamensis, adult females and nymphs almost invisible beneath a layer of secreted white meal, under leaf of Gnetum luofuense. 29) Neoparlatoria formosana (Diaspididae, Leucaspidinae), adult females on Cyclobalanopsis sp. (30) Pseudaulacaspis cockerelli (Diaspididae, Diaspidinae), adult female on Michelia figo.
FIGURES 31–36 in The Hemiptera-Sternorrhyncha (Insecta) of Hong Kong, China-an annotated inventory citing voucher specimens and published records
FIGURES 31–36, Coccoidea and Psylloidea. (31) Drosicha undetermined sp. (Coccoidea, Monophlebidae), adult female on undetermined host. (32) Ferrisia virgata (Pseudococcidae), adult females on Citrus sp. (33) Paurocephala bifasciata (Psylloidea, Psyllidae), adult and nymph on Ficus hispida. (34) Homotoma?yunnanica (Psylloidea, Homotomidae), adult female from Ficus tinctoria gibbosa. (35, 36)?Cacopsylla sp., adult and nymph on Rhaphiolepis indica.
FIGURES 19–24 in The Hemiptera-Sternorrhyncha (Insecta) of Hong Kong, China-an annotated inventory citing voucher specimens and published records
FIGURES 19–24, Aphididae. (19) Neohormaphis undetermined sp., alatoid nymph on Cyclobalanopsis championii. (20) Dermaphis undetermined sp., mature aptera on Cyclobalanopsis championii. (21) Cerataphis brasiliensis, aptera on Archontophoenix alexandrae. (22) Capitophorus sp., aptera and first-instar nymph on Polygonun sinense. (23) Phyllaphoides bambusicola, alata and nymphs on bamboo leaf. (24) Toxoptera odinae, apterae and nymphs on undetermined host.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.