Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “Keyphrases”
Learning Rich Representation of Keyphrases from Text
<p>In this work, we explore how to learn task-specific language models aimed towards learning rich representation of keyphrases from text documents. We experiment with different masking strategies for pre-training transformer language models (LMs) in discriminative as well as generative settings. In the discriminative setting, we introduce a new pre-training objective - Keyphrase Boundary Infilling with Replacement (KBIR), showing large gains in performance (up to 9.26 points in F1) over SOTA, when LM pre-trained using KBIR is fine-tuned for the task of keyphrase extraction. In the generative setting, we introduce a new pre-training setup for BART - KeyBART, that reproduces the keyphrases related to the input text in the CatSeq format, instead of the denoised original input. This also led to gains in performance (up to 4.33 points in F1@M) over SOTA for keyphrase generation. Additionally, we also fine-tune the pre-trained language models on named entity recognition (NER), question answering (QA), relation extraction (RE), abstractive summarization and achieve comparable performance with that of the SOTA, showing that learning rich representation of keyphrases is indeed beneficial for many other fundamental NLP tasks.</p> <p>As a part of this zip file we release the KBIR model which is continually pre-trained on RoBERTa-Large and also the KeyBART model which is continually pre-trained on BART-Large. Both these models can be used in place of a RoBERTa-Large or BART-Large model in PyTorch codebases and also with HuggingFace.</p>
Wilcoxon Rank Sum Test and Keyphrase Extraction Data Cited in "What Everyone Says: Public Perceptions of the Humanities in the Media"
<p>This repository contains Wilcoxon rank sum test and keyphrase extraction data cited in the WhatEvery1Says (WE1S) Project's article "What Everyone Says: Public Perceptions of the Humanities in the Media". The organization of the materials is discussed below.</p> <p><strong>Wilcoxon Rank Sum Test</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>wilcoxon-tests</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. The Wilcoxon rank sum test identifies specific words that appear significantly more in one group of documents as compared to another, thus providing researchers with an understanding of what words are “distinctive” to each group. Further information on WE1S's use of Wilcoxon rank sum testing can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-15-Wilcoxon-Test.pdf</a>.</p> <p>Each subdirectory in the <code>wilcoxon-test</code> folder contains the data and results of a particular comparison experiment based on a metadata category such as whether the data contained articles published by public or private institutions. Each data file is a <code>.txt</code> file representing a sample of the overall data from the collection. The <code>README</code> file provides information on the collection used, the sample size, and the nature of the comparison. The results for the test are in a file called <code>results.csv</code>.</p> <p>The <code>results.csv</code> file for each test includes a row for each term included in the test. Each row displays the term, the term's raw count in each category compared (count 1 and count 2), the difference between the 2 counts (count 1 minus count 2), the percentage change in the counts, the Wilcoxon statistic, and the Wilcoxon p-value. Sorting the csv by the Wilcoxon stat from greatest to least will cause the terms most strongly associated with category 1 to come to the top (category 1 is the category listed first in the title field of the README.md file for each test), while sorting it by the Wilcoxon stat from least to greatest will cause the terms most strongly associated with category 2 to come to the top (category 2 is the category listed second). The p-value column provides you with information about how confident you can be about each comparison's significance.</p> <p><strong>Keyphrase Extraction</strong></p> <p>All data and results from Wilcoxon rank sum testing can be found in the <code>keyphrase-extraction</code> folder of the extracted zip file <code>we1s_about_the_humanities.zip</code>. Keyphrase extraction generates a list of the most significant words or phrases (1-6 words long) within individual documents. WE1S takes the top ten keyphrases in each document and ranks them according to their frequency across the collection. WE1S uses the SGRank algorithm for keyphrase extraction, and because this algorithm is computationally intensive, WE1S limits keyphrases to lemmatized nouns and proper nouns within a window of 70 words to either side of candidate keyphrases. Further information on WE1S's use of keyphrase extaction can be found at <a href="https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf">https://we1s.ucsb.edu/wp-content/uploads/M-14-Keyphrase-Extraction.pdf</a>.</p> <p>Each subdirectory in the <code>keyphrase-extraction</code> folder contains the data and results of keyphrase extraction on a particular collection. Details of the collection and resulting files can be found in each subdirectory. Each list of keyphrases is in a file called <code>SGRank.csv</code>, which lists the keyphrases and their number of occurrences in the collection. The article additionally cites keyphrases that are shared with the terms in the public topic model produced by Andrew Goldstone and Ted Underwood, “The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us,” <em>New Literary History</em> 45, no. 3 (2014): 359–84, <a href="https://doi.org/10.1353/nlh.2014.0025">https://doi.org/10.1353/nlh.2014.0025</a>. The list of terms is derived from the public visualization at <a href="https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words">https://www.sas.rutgers.edu/virtual/ag978/quiet/#/words</a>. Keyphrases extracted from WE1S data were split into single-word terms and compared with the list of vocabulary in Goldstone and Underwood's word list (<code>quiet_transformations_wordlist.txt</code>) to compile lists of shared vocabulary. These lists are given in files called <code>shared_terms.txt</code>.</p> <p>Note that keyphrases were extracted for corpora produced using the Python <a href="https://textacy.readthedocs.io/en/latest/index.html">Textacy</a> library. Because these corpora contain the full text of articles with intellectual property restrictions they cannot be reproduced here.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.