Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

277

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

277 results for “Dictionary”

Learn how ShareScore rates datasets ↗
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Macedonian-Bulgarian dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2010-2013, Kevin Brubeck Unhammer 2010-2014, Francis M. Tyers 2010, Jim O'Regan 2010, Tihomir Rangelov 2011, Pminervini 2011, Trond Trosterud 2012, Anthony J. Bentley

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Icelandic-Swedish dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2010-2013, Tihomir Rangelov 2011-2013, Francis M. Tyers 2013, Kevin Brubeck Unhammer 2013, Trond Trosterud 2014, Tino Didriksen

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Occitan (post 1500)-Catalan dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2005 Universitat d'Alacant (Transducens group) (c) 2005--2008 Generalitat de Catalunya (c) 2005 Universitat Pompeu Fabra (IULA) (c) 2007--2008 Prompsit Language Engineering S.L.

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Spanish-Asturian dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2008-2010, Universidá d'Uviéu (Equipu d'Investigación Eslema / Dptos. de Filoloxía Española ya Informática, http://di098.edv.uniovi.es/apertium/comun/nos.php) (c) 2007-2009, Prompsit Language Engineering (Parts of build system, documentation, and lexical data) (c) 2005-2006, Universidade de Vigo (Seminario de Lingüística Informática, http://sli.uvigo.es) (parts of Spanish monolingual data) (c) 2005-2006, Universitat d'Alacant (Transducens group, http://transducens.dlsi.ua.es) (Spanish monolingual data) (c) 2005-2006, Universitat Politècnica de Catalunya (TALP group, http://www.talp.upc.edu) (parts of Spanish monolingual data)

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Serbo-Croatian-Macedonian dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2007-2013, Ivica Dimitrijev 2007-2013, Dejan Čabrilo 2007-2008, Jim O'Regan 2007-2008, Proleter 2007-2014, Francis M. Tyers 2007, Jernej Vicic 2007, Troy 2008, Misoss 2009, Pminervini 2011-2012, Ljubisha Rus 2011-2013, Hrvoje Peradin 2011, Trond Trosterud 2012-2013, Filip Petkovski 2012-2014, Kevin Brubeck Unhammer 2012, Anthony J. Bentley 2012, Cromulen7

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Swedish-Norwegian dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2013-2019, Kevin Brubeck Unhammer 2016-2018, Trond Trosterud 2013-2016, Francis M. Tyers 2013, Sjur Nørstebø Moshagen

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Swedish-Danish dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> 2015-2019, Kevin Brubeck Unhammer 2017, Trond Trosterud 2008-2016, Francis M. Tyers 2012-2015, Per Tunedal 2009-2010, Jacob Nordfalk 2008-2009, Michael Kristensen 2005-2009, Universitat d'Alacant (Transducens group)

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Spanish-Italian dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> (c) 2007--2008 Prompsit Language Engineering S.L. (http://www.prompsit.com) (c) 2008 Universitat d'Alacant, Grup Transducens (http://transducens.dlsi.ua.es)

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Ontolex-lemon and TIAD versions of Apertium Northern Sami-Norwegian Bokmål dictionary

<p>OntoLex-lemon and TSV&nbsp;conversion of Apertium Bidix. For more details, see&nbsp;<a href="https://www.aclweb.org/anthology/2020.lrec-1.401/">https://www.aclweb.org/anthology/2020.lrec-1.401/</a></p> <p>Authors of the original data:</p> Copyright (C) 2010--2015 Francis Tyers Copyright (C) 2010--2015 Kevin Brubeck Unhammer Copyright (C) 2010--2015 Trond Trosterud

opengpl-2.0-or-laterMar 2020View details →
zenodo44/100

Relations in the Biographical Dictionary of Republican China - Raw NLP Output

<p>This file provides the data on relations as extracted from the Biographical Dictionary of Republican China (BDRC) with CoreNLP. This is the raw output before any form of processing and cleaning.</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

Education Data in the Biographical Dictionary of Republican China

<p>This dataset contains records of education of the 589 historical figures in the Biographical Dictionary of Republican China.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Relations in the Biographical Dictionary of Republican China - Node & Edge lists and Attribute File

<p>This dataset contains the node and edge lists of relations in the BDRC. They are based on the &quot;Relations in the Biographical Dictionary of Republican China - Standardized output&quot; file in this repository. The attributes of the nodes refer only to the main 589 figures in the BDRC.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Position Data in the Biographical Dictionary of Republican China

<p>This dataset contains records of positions of the 589 historical figures in the Biographical Dictionary of Republican China.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Austronesian Comparative Dictionary

<p>Cite the source of the dataset as:</p> <blockquote> <p>Robert Blust and Stephen Trussel 2020. The Austronesian Comparative Dictionary. Web edition</p> </blockquote>

opencc-by-4.0Mar 2023View details →
zenodo44/100

The Great Ape Dictionary Video Data Ark

<p>We study the behaviour and cognition of wild apes and other species (elephants, corvids, dogs). Our video archive is called the Great Ape Dictionary, you can find out more here <a href="https://greatapedictionary.ac.uk/">www.greatapedictionary.com</a> or about our lab group here <a href="https://www.wildminds.ac.uk/">www.wildminds.ac.uk</a>&nbsp;We consider these videos to be a data ark&nbsp;that we would like to make as accessible as possible.&nbsp;While we are unable to make the original video files open access at the present time you can search this database to explore what is available, and then request access for collaborations of different kinds by contacting us directly or <a href="https://greatapedictionary.ac.uk/video-resources/request-access/">through our website</a>.</p> <p>We label all videos in the Great Ape Dictionary video archive with basic meta-data on the location, date, duration, individuals present, and behaviour present. Version 1.0.0 contains current data&nbsp;from the Budongo East African chimpanzee population (n=13806 videos). These datasets are being updated regularly&nbsp;and new data will be incorporated here with versioning. As well as the database there is a second read.me file which contains the ethograms used for each variable coded, and a short summary of other datasets that are in preparation for subsequent&nbsp;version(s). If you are interested in these data please contact us.&nbsp;Please note that not all variables are labeled for all videos, the detailed Ethogram categories are only available for a subset of data. All videos are labelled with up to 5 Contexts (at least one, rarely 5). If you are interested in finding a good example video for a particular behaviour, search for 'Library' = Y, this indicates that this clip contains a very clear example of the behaviour.<br><br>March 17th 2025: Version 1.1.0 contains added data from the Bwindi mountain gorilla chimpanzee population (n=3537 videos).</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

dictionaria/teanu: Teanu dictionary (Solomon Islands)

François, Alexandre. 2021. Teanu dictionary (Solomon Islands). Dictionaria 15. 1-1877. (Available online at https://dictionaria.clld.org/contributions/teanu)

opencc-by-4.0Nov 2021View details →
zenodo44/100

Olonets-Karelian-to-X XML Dictionary

<p>This .zip file includes XML dictionary document, divided according to part of speech, for Olonets-Karelian (LIvvi, ISO-693: olo) to Finnish and Russian. The Livvi-Finnish glosses originate from the Kone Foundation&#39;s &quot;Language Programme&quot; and represent work done by Timo Rantakaulio (Creation of morphological parsers for minority Finno-Ugric languages [Morfologisten j&auml;sentimien luominen suomalais-ugrilaisille v&auml;hemmist&ouml;kielille] 2013&ndash;2014) in close collaboration with Jack Rueter, who was able to formulate the finite-state morphological description of the Olonets-Karelian language on the basis of Timo&#39;s expertise. Subsequently, the Olonets-Karelian and Russian language pair has also been aligned in tandum with the Livvi-Finnish pair, but there may be some descrepencies, as no two languages follow identical lexical systems, much less three.</p> <p>The XMLs contain both lemma and stem forms of Livvi with an indication of distinct paradigm information.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

The ALPIN Sentiment Dictionary: Austrian Language Polarity in Newspapers

<p>These datasets are part of the submitted paper for the LREC2022 conference entitled: &quot;The ALPIN Sentiment Dictionary: Austrian Language Polarity in Newspapers&quot;</p> <p>The various data sources, as well as the methodology, are explained in detail in the research paper which will be available soon.</p> <p>ALPIN stands for Austrian Language Polarity in Newspapers. The dictionary consists of three different parts which were merged together:</p> <ul> <li>Austrian Media Corpus: AMC (AMC_v1.0.csv)</li> <li>STANDARD posts: STP (STP_v1.0.csv)</li> <li>Austriacisms: AUT (AUT_v1.0.csv)</li> </ul> <p>Austrian Media Corpus (AMC) (Ransmayr et al., 2017) &amp; STANDARD posts (STP) (Schabus et al., 2017) rely on the SPLM algorithm as used in SentiDraw (Sharma &amp; Dutta 2021). Austriacisms (AUT) was generated by using the Best-Worst scaling (BWS) (Kiritchenko and Mohammad, 2017b). The AUT list was collected from the &ldquo;Variantenw&ouml;rterbuch des Deutschen&rdquo; (Ammon et al., 2016) (thereby only selecting those words that only surface in Austrian German and in no other variety of German) and an austriacism list of Wikipedia (https://de.wikipedia.org/wiki/Liste_von_Austriazismen).</p> <p>The scores are scaled to the interval [-1, 1] using the min-max-abs scaling, ranging from negative to positive.</p> <p>References:<br> Sharma, S. S., &amp; Dutta, G. (2021). SentiDraw: Using star ratings of reviews to develop domain specific sentiment lexicon for polarity determination. Information Processing &amp; Management, 58(1), 102412.<br> Kiritchenko, S. and Mohammad, S. M. (2017b). Capturing reliable fine-grained sentiment associations by crowdsourcing and best-worst scaling.<br> Schabus, D., Skowron, M., &amp; Trapp, M. (2017). One Million Posts: A Data Set of German Online Discussions. Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, 1241&ndash;1244. https://doi.org/10.1145/3077136.3080711<br> Ransmayr, J., M&ouml;rth, K., &amp; Ďurčo, M. (2017). AMC (Austrian Media Corpus). In Korpusbasierte Forschungen zum &ouml;sterreichischen Deutsch. In Digitale Methoden der Korpusforschung in &Ouml;sterreich (= Ver&ouml;ffentlichungen zur Linguistik und Kommunikationsforschung Nr. 30) (pp. 27&ndash;38). Verlag der &Ouml;sterreichischen Akademie der Wissenschaften.<br> Ammon, U., Bickel, H., &amp; Ebner, J. (2016). Variantenw&ouml;rterbuch des Deutschen : die Standardsprache in &Ouml;sterreich, der Schweiz, Deutschland, Liechtenstein, Luxemburg, Ostbelgien und S&uuml;dtirol sowie Rum&auml;nien, Namibia und Mennonitensiedlungen. Walter de Gruyter.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

PanLex, OntoLex-Lemon edition, Batch 1/8 (dictionaries 0..999)

<p>PanLex (panlex.org) is a large-scale collection of dictionaries. Here, we provide a machine-readable edition in RDF, using the OntoLex-Lemon vocabulary and under resolvable URIs. We achieve resolvability by means of a redirection service, at the moment, purl.org. Feel free to explore this data with http://purl.org/acoli/dicts/panlex/000/0.ttl, for example.</p> <p>For details of the conversion process, see (and refer to):</p> <p>Chiarcos, C., F&auml;th, C., &amp; Ionov, M. (2020, May). The ACoLi dictionary graph. In <em>Proceedings of The 12th Language Resources and Evaluation Conference</em> (pp. 3281-3290), https://aclanthology.org/2020.lrec-1.401.pdf</p> <p>Note that Zenodo does not allow us to publish this data under an RDF-compliant media type. We therefore deviate from general naming recommendations and include the file type in the URIs. Some SPARQL clients (e.g., Apache Jena) use the file extension to guess the proper media type. Also note that these heuristics will fail if the standard Zenodo flags are attached to files, so please use bare files, only.</p> <p>Licensing follows the original data. See PanLex batch on language metadata.</p>

opencc-zeroJun 2022View details →
zenodo44/100

Pronunciation Dictionaries for the Alsatian Dialects

<p>This dataset contains a collection of pronunciation dictionaries which were manually transcribed using the X-SAMPA transcription system. The transcriptions were performed based on audio recordings available on the following websites :</p> <ul> <li>OLCA: <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Els&auml;ssich Web diktionnair: <a href="http://www.ami-hebdo.com/elsadico/index.php">http://www.ami-hebdo.com/elsadico/index.php</a></li> </ul> <p>The dataset was produced in the context of the RESTAURE project, funded by the French ANR. The transcription process is described in the following research report : 10.5281/zenodo.1174219. It is also detailed in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01704814">http://hal.archives-ouvertes.fr/hal-01704814</a></p> <p>Three pronunciation dictionaries are available :</p> <ul> <li>elsassich_dico.csv : 702 transcriptions from the &quot;Els&auml;ssich Web diktionnair&quot;</li> <li>olca67.csv : 1,458 transcriptions from the &quot;OLCA&quot; lexicons for the northern part of the Alsace region (Bas-Rhin)</li> <li>olca68.csv : 1,401 transcriptions from the &quot;OLCA&quot; lexicons for the southern part of the Alsace region (Haut-Rhin)</li> </ul>

opencc-by-sa-4.0Feb 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record