Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6
datasets available to search
ShareScore release 0.9.0
Dataset results
6 results for “WordNet”
Open Dutch Wordnet v1.3
<p>Dutch version of <a href="https://wordnet.princeton.edu/">WordNet</a></p><p>More information about the project: <a href="http://www.cltl.nl/projects/current-projects/opensourcewordnet/">http://www.cltl.nl/projects/current-projects/opensourcewordnet/</a></p><p>Authors and main associated publication:</p><p>@InProceedings{Postma:Miltenburg:Segers:Schoen:Vossen:2016, author = "Marten Postma and Emiel van Miltenburg and Roxane Segers and Anneleen Schoen and Piek Vossen", title = "Open {Dutch} {WordNet}", booktitle = "Proceedings of the Eight Global Wordnet Conference", year = 2016, address = "Bucharest, Romania", }</p>
French Word Sense Disambiguation with Princeton WordNet Identifiers
<p>This is a dataset for the Word Sense Disambiguation of French using Princeton WordNet identifiers. It contains two training corpora : the SemCor and the WordNet Gloss Corpus, both automatically translated from their original English version, and with sense tags automatically aligned. It contains also a test corpus : the task 12 of SemEval 2013, originally sense annotated with BabelNet identifiers, converted into Princeton WordNet 3.0.</p>
WordNet–Wikipedia–Wiktionary alignment
<p>This distribution contains the three-way alignments between WordNet 3.0, the English edition of Wikipedia, and the English edition of Wiktionary, as described in the LREC 2014 paper by Tristan Miller and Iryna Gurevych (see below).</p> <p><strong>Format</strong></p> <p>Here you will find two tab-delimited text files, <code>alignment_3way.tsv</code> and <code>alignment_3way_conjoint.tsv</code>. The first of these contains the full alignment of WordNet, Wikipedia, and Wiktionary, except for the unaligned singleton senses. The second file contains the conjoint alignment of WordNet, Wikipedia, and Wiktionary.</p> <p>The format of both files is the same: each line consists of a tab-delimited list of “sense” identifiers which refer to the same concept. Identifiers for Wiktionary are prefixed with a <code>#</code> character, and take the form of the unique sense identifier generated by the <a href="https://dkpro.github.io/dkpro-jwktl/">JWKTL library</a> for a 3 April 2010 dump of the English edition of Wiktionary. Identifiers for Wikipedia are prefixed with a <code>%</code> character, and take the form of the article title (with underscores replacing spaces) as found in a 22 August 2009 snapshot of the English edition of Wikipedia. Identifiers for WordNet are prefixed with a <code>=</code> character and take the form of a synset offset, followed by a hyphen (<code>-</code>), followed by a part of speech label (<code>a</code>, <code>n</code>, <code>r</code>, or <code>v</code>, for adjectives, nouns, adverbs, and verbs, respectively).</p> <p><strong>Citing this resource</strong></p> <p>If you use this resource in your own work, please cite the following paper:</p> <p>Tristan Miller and Iryna Gurevych. <a href="http://www.lrec-conf.org/proceedings/lrec2014/pdf/4_Paper.pdf">WordNet–Wikipedia–Wiktionary: Construction of a three-way alignment</a>. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Hrafn Loftsson, Bente Maegaard, Joseph Mariani, Asunción Moreno, Jan Odijk, and Stelios Piperidis, editors, <em>Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC 2014)</em>, pages 2094–2100. European Language Resources Association, May 2014. ISBN 978-2-9517408-8-4.</p> <p>You can use the following BibTeX entry:</p> <pre>@inproceedings{miller2014wordnet, author = {Tristan Miller and Iryna Gurevych}, title = {{WordNet}--{Wikipedia}--{Wiktionary}: Construction of a Three-way Alignment}, booktitle = {Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC 2014)}, year = 2014, editor = {Nicoletta Calzolari and Khalid Choukri and Thierry Declerck and Hrafn Loftsson and Bente Maegaard and Joseph Mariani and Asunci{\'{o}}n Moreno and Jan Odijk and Stelios Piperidis}, pages = {2094--2100}, month = may, publisher = {European Language Resources Association}, pdf = {http://www.lrec-conf.org/proceedings/lrec2014/pdf/4_Paper.pdf}, isbn = {978-2-9517408-8-4}, }</pre>
Embeddings of the UMBC Corpus over WordNet using Shallow Connectivity Disambiguation
<p>Embeddings of the UMBC Corpus over WordNet using Shallow Connectivity Disambiguation. Also includes HolE style word embeddings generated from WordNet.</p>
LMMS Wordnet Embeddings for SemCor corpus
<p>This dataset contains word vectors generated after training the LMMS <code>Language Modelling Makes Sense (ACL 2019)</code> model with the whole train set of SemCor, adapted by rdenaux.</p> <p>The main modifications include:</p> <ul> <li>support for <a href="https://github.com/huggingface/transformers">transformers</a> backend ** this makes it possible to experiment with other transformer architectures besides BERT, e.g. XLNet, XLM, RoBERTa ** optimised training since we no longer have to pad sequences to 512 wordpiece tokens</li> <li>Introduced <code>SentenceEncoder</code> which is an experimental generalisation of bert-as-service like encoding services using the transformers backend ** allows to extract various types of embeddings from a single execution of a batch of sequences</li> <li>rolling cosine similarity metrics during training phase</li> </ul> <p>The original repository includes the code to replicate the experiments in the <a href="https://arxiv.org/abs/1906.10007">"Language Modelling Makes Sense (ACL 2019)"</a> paper.</p> <p>This project is designed to be modular so that others can easily modify or reuse the portions that are relevant for them. Its composed of a series of scripts that when run in sequence produce most of the work described in the paper (for simplicity, we've focused this release on BERT, let us know if you need ELMo).</p> <p>The code is available <a href="https://github.com/rdenaux/LMMS">here</a>.</p> <p> </p>
Studying Taxonomy Enrichment on Diachronic WordNet Versions
<p>We choose two versions of WordNet and then select words which appear only in a newer version. For each word, we get its hypernyms from the newer WordNet version and consider them as gold standard hypernyms. We add words to the dataset if only their hypernyms appear in both snippets. We do not consider adjectives and adverbs, because they often introduce abstract concepts and are difficult to interpret by context.</p> <p>Previous dataset (RUSSE'2020) does not include short words (<4 symbols), diminutives, named entities and other constraints described in the shared task paper. We remove those constraints and present a non-restricted Russian dataset and a symmetrical English dataset from WordNet database.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.