Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.7.1
Dataset results
6 results for “word sense disambiguation”
RUSSE'2018: Human-Annotated Sense-Disambiguated Word Contexts for Russian
<p>This dataset contains human-annotated sense identifiers for 2562 contexts of 20 words used in the <a href="https://russe.nlpub.org/2018/wsi/">RUSSE'2018</a> shared task on Word Sense Induction and Disambiguation for the Russian language; part of the <em>bts-rnc</em> evaluation dataset. These sense identifiers are disambiguated as according to the sense inventory of the <a href="http://gramota.ru/slovari/info/bts/">Large Explanatory Dictionary of Russian</a>.</p> <p>The annotation is done on December 1, 2017, on the <a href="https://tolokanyandex.com/">Yandex.Toloka</a> crowdsourcing platform. In particular, 80 pre-annotated contexts are used for training the human annotators, 2562 contexts are annotated by humans such that each context was annotated by 9 different annotators. The annotation reliability is indicated by a high value of Krippendorff's α = 0.83. After the annotation, every context was additionally inspected (“curated”) by the organizers of the shared task.</p> <p>The following words are represented: <em>акция</em> (action / stock), <em>байка</em> (yarn / tale), <em>гвоздика</em> (carnation / nail), <em>гипербола</em> (hyperbole), <em>град</em> (avalanche), <em>гусеница</em> (grub), <em>домино</em> (domino), <em>кабачок</em> (marrow / pub), <em>капот</em> (hood), <em>карьер</em> (mine / career), <em>кок</em> (cook), <em>крона</em> (top / crown), <em>круп</em> (croup), <em>мандарин</em> (mandarine), <em>рок</em> (fate / rock), <em>слог</em> (syllable), <em>стопка</em> (glass, stack), <em>таз</em> (bowl), <em>такса</em> (rate / badger-dog), <em>шах</em> (shah / check).</p> <p>The following files are included in this dataset:</p> <ul> <li>Toloka assignments (training: <em>tasks-train.tsv</em>, annotation: <em>tasks-test.tsv</em>)</li> <li>Toloka output (non-aggregated: <em>assignments_01-12-2017.tsv.xz</em>, aggregated: <em>aggregated_results_pool_1036853__2017_12_01.tsv</em>)</li> <li>annotator agreement report (<em>agreement.txt</em>)</li> <li>curated report (<em>report-curated.tsv.xz</em> and a supplementary file <em>tasks-eval.tsv.xz</em>)</li> <li>the final aggregated dataset (<em>bts-rnc-crowd.tsv</em>)</li> </ul> <p>The <em>bts-rnc-crowd.tsv</em> file has the following format: <em>id</em>, <em>lemma</em>, <em>sense_id</em>, <em>left</em> hand side context, <em>word</em> form, <em>right</em> hand side context, list of <em>senses</em>. The encoding is UTF-8 and the line breaks are LF (UNIX).</p>
French Word Sense Disambiguation with Princeton WordNet Identifiers
<p>This is a dataset for the Word Sense Disambiguation of French using Princeton WordNet identifiers. It contains two training corpora : the SemCor and the WordNet Gloss Corpus, both automatically translated from their original English version, and with sense tags automatically aligned. It contains also a test corpus : the task 12 of SemEval 2013, originally sense annotated with BabelNet identifiers, converted into Princeton WordNet 3.0.</p>
Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation
<p>This dataset contains the models for interpretable Word Sense Disambiguation (WSD) that were employed in Panchenko et al. (2017; the paper can be accessed at https://www.lt.informatik.tu-darmstadt.de/fileadmin/user_upload/Group_LangTech/publications/EACL_Interpretability___FINAL__1_.pdf).</p> <p>The files were computed on a 2015 dump from the English Wikipedia. Their contents:</p> <ul> <li>Induced Sense Inventories: <strong>wp_stanford_sense_inventories.tar.gz</strong><br> This file contains 3 inventories (coarse, medium fine)</li> <li>Language Model (3-gram): <strong>wiki_text.3.arpa.gz</strong><br> This file contains all n-grams up to n=3 and can be loaded into an index</li> <li>Weighted Dependency Features: <strong>wp_stanford_lemma_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000.gz</strong><br> This file contains weighted word--context-feature combinations and includes their count and an LMI significance score</li> <li>Distributional Thesaurus (DT) of Dependency Features: <strong>wp_stanford_lemma_BIM_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000_simsortlimit200_feature expansion.gz</strong><br> This file contains a DT of context features. The context feature similarities can be used for context expansion</li> </ul> <p>For further information, consult the paper and the companion page: http://jobimtext.org/wsd/</p> <p>Panchenko A., Ruppert E., Faralli S., Ponzetto S. P., and Biemann C. (2017): Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL'2017). Valencia, Spain. Association for Computational Linguistics.</p> <p> </p> <p> </p>
SemEval-2010 Task 14: Word Sense Induction & Disambiguation
<p>Data originally provided by the SemEval-2010 Task 14 organizers and which was available through the following links:</p> <p>1. https://www.cs.york.ac.uk/semeval2010_WSI/files/evaluation.zip<br> 2. https://www.cs.york.ac.uk/semeval2010_WSI/files/training_data.tar.gz<br> 3. https://www.cs.york.ac.uk/semeval2010_WSI/files/test_data.tar.gz</p>
Sense-Tagged corpora for Word Sense Disambiguation in several languages
<p>Copies of Sense-Tagged corpora for Word Sense Disambiguation in several languages</p> <p>English, French, Spanish, Russian</p> <p> </p>
Amharic WSD Dataset: Advancing Word Sense Disambiguation in Amharic
<p>This dataset is specifically designed for the Word Sense Disambiguation (WSD) task in the Amharic language, consisting of 50,415 annotated sentences. Each sentence includes the correct sense for one of 200 ambiguous words chosen based on homonymy relations, where a single word may have multiple meanings depending on its context.</p> <p>The ambiguous words were selected to capture the nuances of Amharic vocabulary, drawing from diverse textual sources such as news articles, literature, and social media. This ensures a broad and representative range of usage across various contexts, making the dataset particularly valuable for advancing Amharic NLP research. Potential applications include improvements in machine translation, sentiment analysis, and other semantic processing tasks in Amharic.</p> <p>The dataset is organized in a structured format, with each entry containing fields for sentence, ambiguous word, sense, gloss, and sense label, facilitating ease of use for machine learning models.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.