Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2 results for “Keyphrases”

Learn how ShareScore rates datasets ↗
zenodo40/100

Learning Rich Representation of Keyphrases from Text

<p>In this work, we explore how to learn task-specific language models aimed towards learning rich representation of keyphrases from text documents. We experiment with different masking strategies for pre-training transformer language models (LMs) in discriminative as well as generative settings. In the discriminative setting, we introduce a new pre-training objective - Keyphrase Boundary Infilling with Replacement (KBIR), showing large gains in performance (up to 9.26 points in F1) over SOTA, when LM pre-trained using KBIR is fine-tuned for the task of keyphrase extraction. In the generative setting, we introduce a new pre-training setup for BART - KeyBART, that reproduces the keyphrases related to the input text in the CatSeq format, instead of the denoised original input. This also led to gains in performance (up to 4.33 points in F1@M) over SOTA for keyphrase generation. Additionally, we also fine-tune the pre-trained language models on named entity recognition (NER), question answering (QA), relation extraction (RE), abstractive summarization and achieve comparable performance with that of the SOTA, showing that learning rich representation of keyphrases is indeed beneficial for many other fundamental NLP tasks.</p> <p>As a part of this zip file we release the KBIR model which is continually pre-trained on RoBERTa-Large and also the KeyBART model which is continually pre-trained on BART-Large. Both these models can be used in place of a RoBERTa-Large or BART-Large model in PyTorch codebases and also with HuggingFace.</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Wilcoxon Rank Sum Test and Keyphrase Extraction Data Cited in "What Everyone Says: Public Perceptions of the Humanities in the Media"

<p>This repository contains Wilcoxon rank sum test and keyphrase extraction data cited in the WhatEvery1Says (WE1S) Project&#39;s article &nbsp;&quot;What Everyone Says: Public Perceptions of the Humanities in the Media&quot;. The organization of the materials is discussed below.</p> <p><strong>Wilcoxon Rank Sum Test</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>wilcoxon-tests</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. The Wilcoxon rank sum test identifies specific words that appear significantly more in one group of documents as compared to another, thus providing researchers with an understanding of what words are &ldquo;distinctive&rdquo; to each group. Further information on WE1S&#39;s use of Wilcoxon rank sum testing can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf</a>.</p> <p>Each subdirectory in the <code>wilcoxon-test</code> folder contains the data and results of a particular comparison experiment based on a metadata category such as whether the data contained articles published by public or private institutions. Each data file is a <code>.txt</code> file representing a sample of the overall data from the collection. The <code>README</code> file provides information on the collection used, the sample size, and the nature of the comparison. The results for the test are in a file called <code>results.csv</code>.</p> <p>The <code>results.csv</code> file for each test includes a row for each term included in the test. Each row displays the term, the term&#39;s raw count in each category compared (count 1 and count 2), the difference between the 2 counts (count 1 minus count 2), the percentage change in the counts, the Wilcoxon statistic, and the Wilcoxon p-value. Sorting the csv by the Wilcoxon stat from greatest to least will cause the terms most strongly associated with category 1 to come to the top (category 1 is the category listed first in the title field of the README.md file for each test), while sorting it by the Wilcoxon stat from least to greatest will cause the terms most strongly associated with category 2 to come to the top (category 2 is the category listed second). The p-value column provides you with information about how confident you can be about each comparison&#39;s significance.</p> <p><strong>Keyphrase Extraction</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>keyphrase-extraction</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. Keyphrase extraction generates a list of the most significant words or phrases (1-6 words long) within individual documents. WE1S takes the top ten keyphrases in each document and ranks them according to their frequency across the collection. WE1S uses the SGRank algorithm for keyphrase extraction, and because this algorithm is computationally intensive, WE1S limits keyphrases to lemmatized nouns and proper nouns within a window of 70 words to either side of candidate keyphrases. Further information on WE1S&#39;s use of keyphrase extaction can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf</a>.</p> <p>Each subdirectory in the <code>keyphrase-extraction</code> folder contains the data and results of keyphrase extraction on a particular collection. Details of the collection and resulting files can be found in each subdirectory. Each list of keyphrases is in a file called <code>SGRank.csv</code>, which lists the keyphrases and their number of occurrences in the collection. The article additionally cites keyphrases that are shared with the terms in the public topic model produced by Andrew Goldstone and Ted Underwood, &ldquo;The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us,&rdquo; <em>New Literary History</em> 45, no. 3 (2014): 359&ndash;84, <a href="https://doi.org/10.1353/nlh.2014.0025">https://doi.org/10.1353/nlh.2014.0025</a>. The list of terms is derived from the public visualization at <a href="https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words">https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words</a>. Keyphrases extracted from WE1S data were split into single-word terms and compared with the list of vocabulary in Goldstone and Underwood&#39;s word list (<code>quiet_transformations_wordlist.txt</code>) to compile lists of shared vocabulary. These lists are given in files called <code>shared_terms.txt</code>.</p> <p>Note that keyphrases were extracted for corpora produced using the Python <a href="https://textacy.readthedocs.io/en/latest/index.html">Textacy</a> library. Because these corpora contain the full text of articles with intellectual property restrictions they cannot be reproduced here.</p>

opencc-by-sa-4.0Jul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record