Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6 results for “word sense disambiguation”

Learn how ShareScore rates datasets ↗
zenodo44/100

RUSSE'2018: Human-Annotated Sense-Disambiguated Word Contexts for Russian

<p>This dataset contains human-annotated sense identifiers for 2562 contexts of 20 words used in the <a href="https://russe.nlpub.org/2018/wsi/">RUSSE&#39;2018</a> shared task on Word Sense Induction and Disambiguation for the Russian language; part of the&nbsp;<em>bts-rnc</em>&nbsp;evaluation dataset. These sense identifiers are disambiguated as according to the sense inventory of the <a href="http://gramota.ru/slovari/info/bts/">Large Explanatory Dictionary of Russian</a>.</p> <p>The annotation is done on December 1, 2017, on the&nbsp;<a href="https://tolokanyandex.com/">Yandex.Toloka</a>&nbsp;crowdsourcing platform. In particular, 80 pre-annotated contexts are used for&nbsp;training the human annotators, 2562 contexts are annotated by humans such that each&nbsp;context was annotated by 9 different annotators. The annotation reliability&nbsp;is indicated by a high value of Krippendorff&#39;s&nbsp;&alpha; = 0.83. After the annotation, every context was additionally inspected (&ldquo;curated&rdquo;) by the organizers of the shared task.</p> <p>The following words are represented:&nbsp;<em>акция</em> (action / stock), <em>байка</em> (yarn / tale), <em>гвоздика</em> (carnation / nail), <em>гипербола</em> (hyperbole), <em>град</em> (avalanche), <em>гусеница</em> (grub), <em>домино</em> (domino), <em>кабачок</em> (marrow / pub), <em>капот</em> (hood), <em>карьер</em> (mine / career), <em>кок</em> (cook), <em>крона</em> (top / crown), <em>круп</em> (croup), <em>мандарин</em> (mandarine), <em>рок</em> (fate / rock), <em>слог</em> (syllable), <em>стопка</em> (glass, stack), <em>таз</em> (bowl), <em>такса</em> (rate / badger-dog), <em>шах</em> (shah / check).</p> <p>The following files are included in this dataset:</p> <ul> <li>Toloka assignments (training:&nbsp;<em>tasks-train.tsv</em>, annotation: <em>tasks-test.tsv</em>)</li> <li>Toloka output (non-aggregated: <em>assignments_01-12-2017.tsv.xz</em>, aggregated:&nbsp;<em>aggregated_results_pool_1036853__2017_12_01.tsv</em>)</li> <li>annotator agreement report (<em>agreement.txt</em>)</li> <li>curated report (<em>report-curated.tsv.xz</em> and a supplementary file&nbsp;<em>tasks-eval.tsv.xz</em>)</li> <li>the final aggregated dataset (<em>bts-rnc-crowd.tsv</em>)</li> </ul> <p>The <em>bts-rnc-crowd.tsv</em>&nbsp;file has the following format: <em>id</em>, <em>lemma</em>, <em>sense_id</em>, <em>left</em> hand side context, <em>word</em> form, <em>right</em> hand side context, list of&nbsp;<em>senses</em>. The encoding is UTF-8 and the line breaks are LF (UNIX).</p>

opencc-by-sa-4.0Jan 2018View details →
zenodo44/100

French Word Sense Disambiguation with Princeton WordNet Identifiers

<p>This is a dataset for the Word Sense Disambiguation of French using Princeton WordNet identifiers. It contains two training corpora : the SemCor and the WordNet Gloss Corpus, both automatically translated from their original English version, and with sense tags automatically aligned. It contains also a test corpus : the task 12 of SemEval 2013, originally sense annotated with BabelNet identifiers, converted into Princeton WordNet 3.0.</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation

<p>This dataset contains the models for interpretable Word Sense Disambiguation (WSD) that were employed in Panchenko et al. (2017; the paper can be accessed at https://www.lt.informatik.tu-darmstadt.de/fileadmin/user_upload/Group_LangTech/publications/EACL_Interpretability___FINAL__1_.pdf).</p> <p>The files were computed on a 2015 dump from the English Wikipedia. Their contents:</p> <ul> <li>Induced Sense Inventories: <strong>wp_stanford_sense_inventories.tar.gz</strong><br> This file contains 3 inventories (coarse, medium fine)</li> <li>Language Model (3-gram): <strong>wiki_text.3.arpa.gz</strong><br> This file contains all n-grams up to n=3 and can be loaded into an index</li> <li>Weighted Dependency Features: <strong>wp_stanford_lemma_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000.gz</strong><br> This file contains weighted word--context-feature combinations and includes their count and an LMI significance score</li> <li>Distributional Thesaurus (DT) of Dependency Features: <strong>wp_stanford_lemma_BIM_LMI_s0.0_w2_f2_wf2_wpfmax1000_wpfmin2_p1000_simsortlimit200_feature expansion.gz</strong><br> This file contains a DT of context features. The context feature similarities can be used for context expansion</li> </ul> <p>For further information, consult the paper and the companion page: http://jobimtext.org/wsd/</p> <p>Panchenko A., Ruppert E., Faralli S., Ponzetto S. P., and Biemann C. (2017): Unsupervised Does Not Mean Uninterpretable: The Case for Word Sense Induction and Disambiguation. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL'2017). Valencia, Spain. Association for Computational Linguistics.</p> <p> </p> <p> </p>

opencc-by-4.0Mar 2017View details →
zenodo32/100

SemEval-2010 Task 14: Word Sense Induction & Disambiguation

<p>Data originally provided by the SemEval-2010 Task 14 organizers and which was available through the following links:</p> <p>1.&nbsp;https://www.cs.york.ac.uk/semeval2010_WSI/files/evaluation.zip<br> 2.&nbsp;https://www.cs.york.ac.uk/semeval2010_WSI/files/training_data.tar.gz<br> 3.&nbsp;https://www.cs.york.ac.uk/semeval2010_WSI/files/test_data.tar.gz</p>

opencc-by-4.0Jul 2010View details →
zenodo28/100

Sense-Tagged corpora for Word Sense Disambiguation in several languages

<p>Copies of Sense-Tagged corpora for Word Sense Disambiguation in several languages</p> <p>English, French, Spanish, Russian</p> <p>&nbsp;</p>

opencc-byOct 2019View details →
zenodo16/100

Amharic WSD Dataset: Advancing Word Sense Disambiguation in Amharic

<p>This dataset is specifically designed for the Word Sense Disambiguation (WSD) task in the Amharic language, consisting of 50,415 annotated sentences. Each sentence includes the correct sense for one of 200 ambiguous words chosen based on homonymy relations, where a single word may have multiple meanings depending on its context.</p> <p>The ambiguous words were selected to capture the nuances of Amharic vocabulary, drawing from diverse textual sources such as news articles, literature, and social media. This ensures a broad and representative range of usage across various contexts, making the dataset particularly valuable for advancing Amharic NLP research. Potential applications include improvements in machine translation, sentiment analysis, and other semantic processing tasks in Amharic.</p> <p>The dataset is organized in a structured format, with each entry containing fields for sentence, ambiguous word, sense, gloss, and sense label, facilitating ease of use for machine learning models.</p>

restrictedcc-by-4.0Oct 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record