Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3 results for “Wiktionary”

Learn how ShareScore rates datasets ↗
zenodo44/100

Parallel Translations from the English Wiktionary

<p>Parallel translations from English words and expressions as extracted from the English Wiktionary (database version of 2018-06-01). The data includes 2,169,063 different entries from the translation of 149,530 English words and expressions in 2,358 languages (with much variation in vocabulary size among languages: 931 languages have only one entry and German, the largest language after English, 97,091 entries). Data is offered in a tabular textual format, and all entries include (a) a unique ID, (b) a concept ID referring to the source English word, (c) a description string with the English source and a short definition (such as &ldquo;dictionary/publication that explains the meanings of an ordered list of words&ldquo;), (d) a language ID from the Glottolog catalog, (e) the text of the translation as given in the Wiktionary, and (f) an extra field holding complementary information, when available (such as phonetic transcription of the text, noun gender, etc.). Data is also offer in a set of files (tabular textual files, bibtex sources, and JSON metadata) following the Cross-Linguistic Data Formats (CLDF), a specification designed to allow the exchange of cross-linguistic data.</p> <p>Code for the extraction is available at http://github.com/tresoldi/wiktionary_parser and a longer description in a blog post of our group, at https://calc.hypotheses.org/?p=32</p>

opencc-by-sa-4.0Jun 2018View details →
zenodo40/100

WordNet–Wikipedia–Wiktionary alignment

<p>This distribution contains the three-way alignments between WordNet 3.0, the English edition of Wikipedia, and the English edition of Wiktionary, as described in the LREC 2014 paper by Tristan Miller and Iryna Gurevych (see below).</p> <p><strong>Format</strong></p> <p>Here you will find two tab-delimited text files, <code>alignment_3way.tsv</code> and <code>alignment_3way_conjoint.tsv</code>. The first of these contains the full alignment of WordNet, Wikipedia, and Wiktionary, except for the unaligned singleton senses. The second file contains the conjoint alignment of WordNet, Wikipedia, and Wiktionary.</p> <p>The format of both files is the same: each line consists of a tab-delimited list of &ldquo;sense&rdquo; identifiers which refer to the same concept. Identifiers for Wiktionary are prefixed with a <code>#</code> character, and take the form of the unique sense identifier generated by the <a href="https://dkpro.github.io/dkpro-jwktl/">JWKTL library</a> for a 3 April 2010 dump of the English edition of Wiktionary. Identifiers for Wikipedia are prefixed with a <code>%</code> character, and take the form of the article title (with underscores replacing spaces) as found in a 22 August 2009 snapshot of the English edition of Wikipedia. Identifiers for WordNet are prefixed with a <code>=</code> character and take the form of a synset offset, followed by a hyphen (<code>-</code>), followed by a part of speech label (<code>a</code>, <code>n</code>, <code>r</code>, or <code>v</code>, for adjectives, nouns, adverbs, and verbs, respectively).</p> <p><strong>Citing this resource</strong></p> <p>If you use this resource in your own work, please cite the following paper:</p> <p>Tristan Miller and Iryna Gurevych. <a href="http://www.lrec-conf.org/proceedings/lrec2014/pdf/4_Paper.pdf">WordNet&ndash;Wikipedia&ndash;Wiktionary: Construction of a three-way alignment</a>. In Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Hrafn Loftsson, Bente Maegaard, Joseph Mariani, Asunci&oacute;n Moreno, Jan Odijk, and Stelios Piperidis, editors, <em>Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC 2014)</em>, pages 2094&ndash;2100. European Language Resources Association, May 2014. ISBN 978-2-9517408-8-4.</p> <p>You can use the following BibTeX entry:</p> <pre>@inproceedings{miller2014wordnet, author = {Tristan Miller and Iryna Gurevych}, title = {{WordNet}--{Wikipedia}--{Wiktionary}: Construction of a Three-way Alignment}, booktitle = {Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC 2014)}, year = 2014, editor = {Nicoletta Calzolari and Khalid Choukri and Thierry Declerck and Hrafn Loftsson and Bente Maegaard and Joseph Mariani and Asunci{\&#39;{o}}n Moreno and Jan Odijk and Stelios Piperidis}, pages = {2094--2100}, month = may, publisher = {European Language Resources Association}, pdf = {http://www.lrec-conf.org/proceedings/lrec2014/pdf/4_Paper.pdf}, isbn = {978-2-9517408-8-4}, }</pre>

opencc-by-4.0Mar 2014View details →
zenodo28/100

Wiktionary taxonomic data (ru and en)

<p>Wiktionary taxonomic data (ru and en) about synonyms, hypernyms and meanings</p>

opencc-by-4.0Nov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record