Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “Wiktionary”
Parallel Translations from the English Wiktionary
<p>Parallel translations from English words and expressions as extracted from the English Wiktionary (database version of 2018-06-01). The data includes 2,169,063 different entries from the translation of 149,530 English words and expressions in 2,358 languages (with much variation in vocabulary size among languages: 931 languages have only one entry and German, the largest language after English, 97,091 entries). Data is offered in a tabular textual format, and all entries include (a) a unique ID, (b) a concept ID referring to the source English word, (c) a description string with the English source and a short definition (such as “dictionary/publication that explains the meanings of an ordered list of words“), (d) a language ID from the Glottolog catalog, (e) the text of the translation as given in the Wiktionary, and (f) an extra field holding complementary information, when available (such as phonetic transcription of the text, noun gender, etc.). Data is also offer in a set of files (tabular textual files, bibtex sources, and JSON metadata) following the Cross-Linguistic Data Formats (CLDF), a specification designed to allow the exchange of cross-linguistic data.</p> <p>Code for the extraction is available at http://github.com/tresoldi/wiktionary_parser and a longer description in a blog post of our group, at https://calc.hypotheses.org/?p=32</p>
WordNet–Wikipedia–Wiktionary alignment
<p>This distribution contains the three-way alignments between WordNet 3.0, the English edition of Wikipedia, and the English edition of Wiktionary, as described in the LREC 2014 paper by Tristan Miller and Iryna Gurevych (see below).</p> <p><strong>Format</strong></p> <p>Here you will find two tab-delimited text files, <code>alignment_3way.tsv</code> and <code>alignment_3way_conjoint.tsv</code>. The first of these contains the full alignment of WordNet, Wikipedia, and Wiktionary, except for the unaligned singleton senses. The second file contains the conjoint alignment of WordNet, Wikipedia, and Wiktionary.</p> <p>The format of both files is the same: each line consists of a tab-delimited list of “sense” identifiers which refer to the same concept. Identifiers for Wiktionary are prefixed with a <code>#</code> character, and take the form of the unique sense identifier generated by the <a href="https://dkpro.github.io/dkpro-jwktl/">JWKTL library</a> for a 3 April 2010 dump of the English edition of Wiktionary. Identifiers for Wikipedia are prefixed with a <code>%</code> character, and take the form of the article title (with underscores replacing spaces) as found in a 22 August 2009 snapshot of the English edition of Wikipedia. Identifiers for WordNet are prefixed with a <code>=</code> character and take the form of a synset offset, followed by a hyphen (<code>-</code>), followed by a part of speech label (<code>a</code>, <code>n</code>, <code>r</code>, or <code>v</code>, for adjectives, nouns, adverbs, and verbs, respectively).</p> <p><strong>Citing this resource</strong></p> <p>If you use this resource in your own work, please cite the following paper:</p> <p>Tristan Miller and Iryna Gurevych. <a href="http://www.lrec-conf.org/proceedings/lrec2014/pdf/4_Paper.pdf">WordNet–Wikipedia–Wiktionary: Construction of a three-way alignment</a>. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Hrafn Loftsson, Bente Maegaard, Joseph Mariani, Asunción Moreno, Jan Odijk, and Stelios Piperidis, editors, <em>Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC 2014)</em>, pages 2094–2100. European Language Resources Association, May 2014. ISBN 978-2-9517408-8-4.</p> <p>You can use the following BibTeX entry:</p> <pre>@inproceedings{miller2014wordnet, author = {Tristan Miller and Iryna Gurevych}, title = {{WordNet}--{Wikipedia}--{Wiktionary}: Construction of a Three-way Alignment}, booktitle = {Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC 2014)}, year = 2014, editor = {Nicoletta Calzolari and Khalid Choukri and Thierry Declerck and Hrafn Loftsson and Bente Maegaard and Joseph Mariani and Asunci{\'{o}}n Moreno and Jan Odijk and Stelios Piperidis}, pages = {2094--2100}, month = may, publisher = {European Language Resources Association}, pdf = {http://www.lrec-conf.org/proceedings/lrec2014/pdf/4_Paper.pdf}, isbn = {978-2-9517408-8-4}, }</pre>
Wiktionary taxonomic data (ru and en)
<p>Wiktionary taxonomic data (ru and en) about synonyms, hypernyms and meanings</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.