Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
277
datasets available to search
ShareScore release 0.9.0
Dataset results
277 results for “dictionaries”
Automatic TEI encoding of manuscripts catalogues with GROBID-Dictionaries
<p>Manuscript Sales Catalogues (MSC) are highly important for authenticating documents and studying the reception of authors. Their regular publication throughout Europe since the beginning of the 19th c. has consequently raised the interest around scaling up the means for automatically structuring their contents. </p> <p>Following successful first encoding tests with <em>GROBID-Dictionaries</em> on a single MSC collection, we aim in this paper to present the results of more advanced tests of the system’s capacity to handle a larger corpus with MSC of different dealers, and therefore multiple layouts. Four different types of catalogues published between the middle of the 19th c. and the beginning of the 20th c. have been tested.</p>
CLDF dataset on Panoan Languages in Standardized Transcription derived from Key and Comrie's "Intercontinental Dictionary Series" from 2023
<p>Cite the source of the dataset as:</p> <blockquote> <p>Miller, J. and List, J.-M. (2024): Providing standardized phonetic transcriptions for the Panoan languages in the Intercontinental Dictionary Series. Computer-Assisted Language Comparison in Practice 7.2. URL: https://calc.hypotheses.org/7503.</p> </blockquote>
HELP Study Data Dictionary / Catalog of Items
<p>The <a href="https://doi.org/10.1136/bmjopen-2019-033391">HELP study</a> was a clinical trial in the form of a multicenter interventional randomized controlled trial (RCT) at five German university hospitals, conducted from 2020 to 2022. It aimed to enhance the clinical management of Staphylococcus bacteremia and used data both from Electronic Health Records (EHR), provided by German university hospital's Data Integration Centers in the HL7 FHIR format (German profiles of the Medical Informatics Initiative Core Data Set – <a href="https://www.medizininformatik-initiative.de/en/medical-informatics-initiatives-core-data-set">MII CDS</a>), as well as data from Electronic Case Report Forms (eCRF) used in the study for data which was too unstructered or not availabe in the EHR (hybrid data collection approach).<br>This dataset is a tabular listing and description of the data items used in the study and <em>serves (primarily) as a template for information and data modeling.</em></p>
CLDF dataset accompanying List's "Inference of Partial Colexifications" from 2023 building on Key and Comrie's "Intercontinental Dictionary Series" from 2021
<p>Cite the source of the dataset as:</p> <blockquote> <p>List, J.-M. (2023): Inference of partial colexifications from multilingual wordlists. Frontiers in Psychology 14.1156540. 1-10.</p> </blockquote>
dictionaria/hdi: Hdi Dictionary
Frajzyngier, Zygmunt and Eguchi, Paul and Prafé, Roger and Schwabauer, Megan (with Erin Shay, Henry Tourneux). 2017. Hdi dictionary. Dictionaria 2. 1-1681. DOI: 10.5281/zenodo.3067648 (Available online at https://dictionaria.clld.org/contributions/hdi)
dictionaria/guarayu: Guarayu. A revised dictionary by Alfred Hoeller
Danielsen, Swintha and Sell, Lena and Terhart, Lena. 2019. Guarayu. A revised dictionary by Alfred Hoeller. Dictionaria 7. 1-3590 (Available online at https://dictionaria.clld.org/contributions/guarayu)
dictionaria/daakaka: Daakaka Dictionary
von Prince, Kilu. 2017. Daakaka dictionary. Dictionaria 1. 1-2167 (Available online at https://dictionaria.clld.org/contributions/daakaka)
dictionaria/wersing: Wersing dictionary
Schapper, Antoinette. 2020. Wersing dictionary. Dictionaria 11. 1-1430 (Available online at https://dictionaria.clld.org/contributions/wersing)
dictionaria/teop: A multifunctional Teop-English dictionary
Mosel, Ulrike. 2019. A multifunctional Teop-English dictionary. Dictionaria 4. 1-6488. (Available online at https://dictionaria.clld.org/contributions/teop, Accessed on 2021-02-08.)
dictionaria/palula: Palula dictionary
Liljegren, Henrik. 2019. Palula dictionary. Dictionaria 3. 1-2700 (Available online at https://dictionaria.clld.org/contributions/palula)
dictionaria/sanzhi: Sanzhi Dargwa dictionary
Forker, Diana. 2019. Sanzhi Dargwa dictionary. Dictionaria 5. 1-5533. (Available online at https://dictionaria.clld.org/contributions/sanzhi, Accessed on 2021-02-08.)
dictionaria/medialengua: Media Lengua dictionary
Visser, Eline. 2020. Kalamang dictionary. Dictionaria 13. 1-2737 (Available online at http://dictionaria.clld.org/contributions/kalamang)
dictionaria/kalamang: Kalamang dictionary
Visser, Eline. 2020. Kalamang dictionary. Dictionaria 13. 1-2737 (Available online at http://dictionaria.clld.org/contributions/kalamang)
CLDF dataset derived from Key and Comrie's "Intercontinental Dictionary Series" from 2023
<p>Cite the source of the dataset as:</p> <blockquote> <p>Key, Mary Ritchie & Comrie, Bernard (eds.) 2023. The Intercontinental Dictionary Series. Leipzig: Max Planck Institute for Evolutionary Anthropology.</p> </blockquote>
CLDF dataset derived from a preprint of 'Új magyar etimológiai szótár' [New Hungarian Etymological Dictionary] by Károly Gerstner (ed.)
<p>Cite the source of the dataset as:</p> <blockquote> <p>Gerstner, Károly (ed.) (2011-2023). Új magyar Etimológiai Szótár. Hungarian Academy of Sciences, Budapest. http://uesz.nytud.hu/.</p> </blockquote>
Uzbek School corpus (dictionary)
<p><strong>Uzbek School Literature Corpus: A Comprehensive Collection for Language Research and Education</strong></p> <p>This project presents a comprehensive language corpus comprising various forms of school literature in Uzbek, spanning grades 1 to 9. The corpus includes diverse texts such as storybooks, textbooks, poems, and essays, providing a representative sample of linguistic and literary contexts within the educational system. It serves as a valuable resource for linguistic research, educational development, and language-related applications. The corpus can be accessed through a repository, fostering collaboration and advancing linguistic research, education, and language technology.</p>
Ministerratsprotokolle: Dictionary files
<p>This is a collection of Habsburg Monarchy dictionary files.</p> <p>The .dic files contain UTF-8 encoded wordlists that may serve as helper files for spell-checking, training models for Named entity recognition etc. In general, this is german-language material, with the exception of proper names. Manual curation has been done to our best knowledge. Overlap in placenames is on purpose as several places fulfil multiple roles.</p> <ul> <li>gerichtsbezirke.dic - jurisdiction related placenames from all parts (Kronländer) of Cisleithania (the Habsburg Monarchy’s Austrian part) | 1702 names</li> <li>institutionen.dic - wordlist extracted from institution names of Austro-Hungarian Monarchy administration | 356 words</li> <li>politische_bezirke.dic - administrative districts placenames from all parts (Kronländer)of Cisleithania (the Habsburg Monarchy’s Austrian part) | 315 names</li> <li>wahlkreise.dic - election districts placenames from all parts (Kronländer) of of Cisleithania (the Habsburg Monarchy’s Austrian part) (data from 1907) | 1940 names</li> <li>wordlist_register-serie1.dic - list of unique index entries from 28 vols of Ministerratsprotokolle 1848-1867 not contained in standard hunspell `german.dic` | 1064 words</li> <li>wordlist_rgbl_titel.dic - wordlist created from all legal regulations named in the "Reichsgesetzblatt" 1848-1918, based on publicly available alex.onb.ac.at data; duplicates from the above wordlists have been removed | 5195 words</li> </ul> <p>This publication aims to provide words that have disappeared from normal discourse, in an effort to save some linguistic heritage as a byproduct to the Austrian Academy of Scienes’ Ministerratsprotokolle digital edition efforts...</p>
Duhumbi Dictionary - Illustrations
<p>This repository contains a Zip file with in it all the illustrations belonging to the publication:</p> <p><br> Title: Duhumbi Dictionary - Duhumbi Tshikze - དུ་ཧུམ་བི་ཚིག་མཛོད།</p> <p>Author: Timotheus Adrianus Bodt</p> <p>Year of publication: 2020</p> <p>Publisher: Monpasang Publications, Arnhem, the Netherlands. </p> <p>ISBN/EAN: 978-90-818610-9-0</p> <p>All illustration used in this dictionary were taken by the author, except those noted here with the original copyright holders, permission for usage of the illustrations is available on file. Blyth’s tragopan: Rofikul Islam, Oriental Bird Club, <a href="https://www.orientalbirdclub.org/">https://www.orientalbirdclub.org/</a>; Monal pheasant: Subhajit Chaudhuri, Oriental Bird Club, <a href="https://www.orientalbirdclub.org/">https://www.orientalbirdclub.org/</a>; Temminck's tragopan: James Eaton, Oriental Bird Club, <a href="https://www.orientalbirdclub.org/">https://www.orientalbirdclub.org/</a>; Kalij pheasant and Himalayan swiftlet: Ayuwat Jearwattanakanok, Oriental Bird Club, <a href="https://www.orientalbirdclub.org/">https://www.orientalbirdclub.org/</a>; Tibetan-eared pheasant: Ian Merrill, Surfbirds online photo gallery, <a href="http://www.surfbirds.com/">http://www.surfbirds.com/</a>. If anyone thinks his / her copyright was infringed upon, please contact the author at <a href="mailto:monpasang@gmail.com">monpasang@gmail.com</a> for amendments. The drawings in the first section of the book, with the Duhumbi alphabet, were made by Debbie Patterson and earlier published in the Duhumbi Storybook (Bodt, 2018).</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collector of the material. By downloading this material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: monpasang (at) gmail (dot) com</p>
Data accompanying the book review "A new general dictionary of Ume Saami"
<p>This is a set of tables illustrating the differences between two dictionaries of Ume Saami (Schlachter 1958 and Barruk 2018). Headwords with the initial <em>v</em> from both dictionaries have been aligned to show the common and differing shares of vocabulary. A further analysis is published in the book review (Kuokkala 2020) in Finnisch-Ugrische Forschungen (<a href="https://doi.org/10.33339/fuf.99934">doi.org/10.33339/fuf.99934</a>). The data from Schlachter 1958 has been converted into the modern orthography with a script found at <a href="https://doi.org/10.5281/zenodo.4162535">doi.org/10.5281/zenodo.4162535</a> and published in its entirety at <a href="https://doi.org/10.5281/zenodo.4163676">doi.org/10.5281/zenodo.4163676</a>.</p>
Files and code for English dictionaries, gold and silver standard corpora for biomedical natural language processing related to SARS-CoV-2 and COVID-19
<p><span lang="EN-GB">Automated information extraction with natural language processing (NLP) tools is required to gain systematic insights from the large number of COVID-19 publications, reports and social media posts, which far exceed human processing capabilities. </span></p> <p><span lang="EN-GB">Here we present an NLP toolbox comprising COVID-19-related dictionaries and annotated corpora in English as well as useful code and workflows for their update and use. The dictionaries contain terms referring to the COVID-19 disease, the SARS-CoV-2 virus, its variants and common mutations, respectively. They were used together with the EasyNER NLP tool to extract and annotate all 764 398 abstracts in the CORD-19 dataset, creating a very large silver standard corpus (named Lund-Annotated-CORD-19 corpus). This was complemented with a small gold standard corpus consisting of PubMed abstracts manually annotated for key entity classes such as disease, virus, symptom, protein/gene, cell type, chemical and species terms. </span></p> <p><span lang="EN-GB">The toolbox can support various text analysis tasks related to COVID-19 such as named entity recognition and co-mention analysis. A preliminary version of the toolbox, which was released early in the pandemic, was</span><span lang="EN-GB"> for example already used to create a COVID-19 knowledge graph and study the evolution and variation of COVID-19-related terminology. In addition, the toolbox can be applied in the development of other NLP tools, for example to train and evaluate large language models.</span></p> <p><span lang="EN-GB">When using the toolbox, please cite this record and the associated article.</span></p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.