Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

277

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

277 results for “dictionaries”

Learn how ShareScore rates datasets ↗
zenodo40/100

Semantic Enrichment of the Laboratory Data Dictionary of the Study of Health in Pomerania (SHIP-START-4) with LOINC; Detailed Mapping Results

<p>Unlike West Germany, high morbidity and mortality have been observed in East Germany over the last century. The regional population-based Study of Health in Pomerania (SHIP) therefore investigates the long-term progression of sub-clinical findings, their determinants and prognostic values, to acquire knowledge that facilitates early diagnosis and thus helps prevent the progression of disease. &nbsp;The SHIP covers various areas of patient health. Each SHIP data set is accompanied by a data dictionary (DD) which provides descriptions of variables and definitions.</p> <p>This work shows the detailed mapping results of the semantic enrichment of the SHIP-START-4 medical laboratory data dictionary with LOINC codes. This work also provides detailed descriptions of the concepts applied in the semnatic enrichment. The results of this work serve as a critical step towards improving its interoperability and hence FAIRness for the SHIP laboratory-related measurements. &nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Metabolomics WorkBench Compound Dictionary

<p>mwTAB file for each study in the Metabolomics Workbench (https://www.metabolomicsworkbench.org/) was processed through an automated data merging pipeline. It yielded -</p> <p>&nbsp;</p> <ul> <li>~237,333 unique chemical names</li> <li>~13,810 PubChem CIDs</li> <li>~14,298 KEGG IDs</li> <li>~5,130 CAS Numbers</li> <li>- 147 unique species</li> <li>- 175 unique sample type</li> </ul> <p>&nbsp;</p> <p>Note:&nbsp;workbench_curated_compound_list.csv contains some of the curated compound names. Basic curation has happened to connect metabolites names to PubChem identifiers.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Metabolomics Data Dictionary (Metabolon Inc. ) - from PMC articles (OA)

<p>A collection of metabolite&nbsp;and chemical names reported by Metabolon&nbsp;Inc. in the&nbsp;supplementary tables of open-access full-text articles from the PMC database.</p> <p><strong>Many entries can be duplicate because of chemical name variants reported in different articles.&nbsp;</strong></p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Data and source code for Automatic generation of a large dictionary with concreteness/abstractness ratings based on a small human dictionary

<p>We present a method for automatic ranking concreteness of words and propose an approach to significantly decrease amount of expert assessment. The method has been evaluated on a large test set for English. The quality of the constructed dictionaries is comparable to the expert ones. The correlation between predicted and expert ratings is higher comparing to the state-of-the-art methods.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Machine-readable Northern Karelian Proper-Livvi bilingual translation dictionary

<p>This machine readable bilingual translation dictionary of Northern Karelian Proper (ISO-639: krl) to Livvi aka Olonets-Karelian (ISO-639 olo) was created by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd&#39; with translation suggestions generated by Khalid Alnajjar and Mika H&auml;m&auml;l&auml;inen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Machine-readable Finnish-Livvi bilingual translation dictionary

<p>This machine readable bilingual translation dictionary of Finnish to Livvi aka Olonets-Karelian (ISO-396: olo) was proofread and extended by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd&#39; with translation suggestions generated by Khalid Alnajjar and Mika H&auml;m&auml;l&auml;inen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Machine-readable Finnish-Karelian bilingual translation dictionary

<p>This machine readable bilingual translation dictionary of Finnish to Northern Karelian Proper (ISO-396: krl) was proofread and extended by Timo Rantakaulio during Google Summer of Code 2021 at Apertium. His work was facilitated through an online back-end dictionary editing tool `https://akusanat.com/verdd&#39; with translation suggestions generated by Khalid Alnajjar and Mika H&auml;m&auml;l&auml;inen developers of the Verdd editor. Workflow, coordination, instruction and advice were given by Flammie A. Pirinen and Jack Rueter.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

A Dive into the State of the Practice of Brazilian Game Software Ecosystems - Codes and Relationships Dictionary (PT-BR)

<p>This spreadsheet is related to Grounded Theory open and axial coding processes regarding the answers captured by the survey about the identification of positives, negatives, and opportunities&nbsp;in Brazilian Game Software Ecosystems (https://zenodo.org/record/6632485#.YqV2XajMKW). It is essential to note that the responses are in the native language of the target audience (Brazilian Portuguese).</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Figure 3. Using the ADX agent for obtaining definitions, synonyms and antonyms, for a given word, during MS Word editing-ADX – Agent for Morphologic Analysis of Lexical Entries in a Dictionary

<p>In figure 3 we present a capture screen of using the ADX1 agent in editing the text in<br> Microsoft Word. It displays the definition of the current word, but it also generates the synonyms,<br> used in the application of some web searching rules, used by another intelligent agent, called ASR.<br> The ASR agent will automatically compose some search strings to use for a web search engine, like<br> Google, Yahoo, or other. For example, if the user will search the word zăpadă (snow) on Yahoo<br> search engine, then one can also generate a search for the word nea (synonym of zăpadă) by using<br> the following ASR rule:<br> # Yahoo search<br> IF<br> http://search.yahoo.com/search?p=^X^&amp;fr=yfp-t-309&amp;toggle=1&amp;cop=mss&amp;ei=UTF-8<br> THEN<br> http://search.yahoo.com/search?p=^Clasa(X)^&amp;fr=yfp-t-309&amp;toggle=1&amp;cop=mss&amp;ei=UTF-8</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

Figure 1. Linguistic analysis - ADX – Agent for Morphologic Analysis of Lexical Entries in a Dictionary

<p>The first three analyses constitute an important phase, having as a result the obtaining of a<br> phrase that we cannot say is incorrect in that language, e. g. &ldquo;Copilul a m&acirc;ncat bătaie&rdquo;. It is hard to<br> suppose, of course, that a child can have a &ldquo;bataie&rdquo; (&ldquo;beating&rdquo;) for his meal, but the problem is<br> solvable from a semantic point of view, because &ldquo;a m&acirc;nca bătaie&rdquo; is a phrasal verb having the<br> meaning of &ldquo;a fi bătut&rdquo; (&ldquo;to be beaten&rdquo;). (The correct translation of the &ldquo;Copilul a m&acirc;ncat bătaie&rdquo; is<br> &ldquo;The child has been beaten.&rdquo;).</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

Figure 2. The structures of the two tables from the Dex Online database-ADX – Agent for Morphologic Analysis of Lexical Entries in a Dictionary

<p>Dex Online is a project initiated and coordinated by Catalin Francu [3]. He intended to<br> realise an online database for all the words in the Romanian language, using the main explanatory<br> dictionaries, dictionaries of synonyms, neologisms, published by the Romanian Academy and other<br> scientific forums.<br> The database was completed by volunteers, similarly to the Wikipedia system. They actually<br> transcribed the information from different important dictionaries, but many words have been<br> electronically entered by two companies (Siveco and Litera International Publishing House).</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

Table 1. The declination of feminine nouns (after [4])-ADX — Agent for Morphologic Analysis of Lexical Entries in a Dictionary

<p>An example of such searching trees for nouns of all genders (masculine, feminine, neuter)<br> together with the entire analysis made by the ADX agent for the word &ldquo;cepelor&rdquo; is presented in<br> figure 2. The pairs of numbers in brackets form lists of lines and columns from the inflections&rsquo;<br> charts where endings from the top of the list are found. We have noted the zero ending with &ndash; (that<br> is always the root of the tree) and the void list with * (Null/Nil).</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

Concept list of Vietnamese based on the Intercontinental Dictionary Series (IDS)

<div> <p>The list includes 1310 concepts for Vietnamese. The concepts are based on the items provided in the Intercontinental Dictionary Series (<a href="https://ids.clld.org/">IDS</a>, Key &amp; Comrie 2023).&nbsp;</p> </div>

opencc-by-4.0Jun 2024View details →
zenodo40/100

English-Portuguese Dictionary of Verbal Collocations

<p>This is the json-LD version of the data of Tagnin's <strong>English-Portuguese Dictionary of Verbal Collocation</strong>s, which is available as a proof-of-concept interactive resource at&nbsp;<a href="https://mangalamresearch.shinyapps.io/EnglishPortugueseVerbalCollocations/" target="_blank" rel="nofollow noopener noreferrer">https://mangalamresearch.shinyapps.io/EnglishPortugueseVerbalCollocations/</a> .</p> <p>This is still work in progress; comments and suggestions are most welcome. Contact me at seotagni@gmail.com.</p> <p>The conversion of the dictionary data and its display in the form of a digital dictionary were made possible thanks Ligeia Lugli, Tilak Balavijayan and the NEH-funded project&nbsp;<em>Democratizing Digital Lexicography</em> (HAA-290402-23).</p> <p>For more information about the criteria adopted in the dictionary, please go to the Word file included in this repository.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Database of Hokkien Dictionaries and Textbooks

<p>Hokkien (a.k.a. Minnan 閩南話, Southern Min, Taiwanese) is a variety of Chinese spoken in the southern part of Fujian province (China), Taiwan and by a large number of overseas Chinese all over Southeast Asia. Hokkien dictionaries and textbooks have been published since the 16th century in a number of locations and in a variety of languages (Classical Chinese, Dutch, English, Hokkien, Japanese, Latin, Mandarin, Spanish) for purposes such as education for locals, Christian missionary work and colonial administration.</p> <p>In this post I would like to share a database project I created for the &quot;Working with Digital Data for Historians&quot; class taught by Prof. Tara Andrews at Uni Vienna in the autumn semester of 2018/19. The project work is based on my 3-years MA studies at Xiamen University (Fujian prov., China) and is connected to an article of mine (<em>Xiamen at the Crossroads of Sino-Foreign Interaction During the Late Qing and Republican Periods: The Issue of Hokkien Phoneticization)&nbsp;</em>already peer-reviewed and accepted for publication in the journal&nbsp;<em>Crossroads - Studies on the History of Exchange Relations in the East Asian World</em>. The database is an Excel spreadsheet-based SQL database containing 121 titles of Hokkien linguistic works, connected to relevant information (author, place/date of publishing, publisher, language, transcription method, etc.).</p> <p>Hereby, I submit the dump version of the database (a .sql file containing the SQL scheme and the data as well) and a zip collection containing further relevant materials (project description, original spreadsheets, SQL command scheme, .ipynb documentation of converting the Excel spreadsheets into the SQL database using Jupyter Notebook).</p> <p>Since during my research work I focused on the pre-WWII period, data regarding the post-WWII is highly insufficient. Therefore, the database is up for further expansion. (My e-mail address: sebestyen.hompot@outlook.com)</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

Coarse-Grained Sense Inventories Based on Semantic Matching between English Dictionaries

<p><strong>Abstract</strong> (our paper)</p> <p>WordNet is one of the largest handcrafted concept dictionaries visualizing word connections through semantic relationships. It is widely used as a word sense inventory in natural language processing tasks. However, WordNet's fine-grained senses have been criticized for limiting its usability. In this paper, we semantically match sense definitions from Cambridge dictionaries and WordNet and develop new coarse-grained sense inventories. We verify the effectiveness of our inventories by comparing their semantic coherences with that of Coarse Sense Inventory. The advantages of the proposed inventories include their low dependency on large-scale resources, better aggregation of closely related senses, CEFR-level assignments, and ease of expansion and improvement. Our inventories are publicly available for free use.</p> <p><strong>Publication</strong></p> <p>These datasets are part of our research results. If you make use of our datasets, please cite:</p> <ul> <li>Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono. Coarse-Grained Sense Inventories Based on Semantic Matching between English Dictionaries. In <em>Proceedings of the 11th International Conference on Advanced Informatics: Concepts, Theory and Applications (ICAICTA 2024)</em>. 6 pages, 2024.</li> </ul>

opencc-zeroSep 2024View details →
zenodo40/100

Dataset and Data Dictionary for "Enhancing Consumer Satisfaction in Live Commerce: A Study of Middle-Aged Women's Cosmetics Purchases Using TAM, PVT, and SIT Models

<p>This dataset is part of a study investigating the underexplored factors driving middle-aged Chinese women&rsquo;s purchasing behavior in live commerce, particularly in the context of their decision-making amidst the rapid expansion of e-commerce. The study employs a comprehensive theoretical framework based on the Technology Acceptance Model (TAM), Perceived Value Theory (PVT), and Social Influence Theory (SIT).</p> <p>Data were collected through a structured survey administered to 653 women aged 40 to 59. The dataset captures key variables including ease of use, pricing, consumer trust, platform interactivity, and purchase satisfaction. These variables are essential for understanding the complex relationships that influence purchasing decisions in live-stream shopping environments.</p> <p>The dataset has been analyzed using Structural Equation Modeling (SEM), revealing that factors such as ease of use, perceived value from competitive pricing, consumer trust, and real-time platform interactivity significantly enhance purchase satisfaction. Moreover, the results demonstrate that perceived value moderates these relationships, amplifying their effects under conditions of high perceived value.</p> <p>This dataset provides valuable insights into the psychological and social factors that shape e-commerce behavior, offering implications for optimizing platform design to promote consumer trust and long-term engagement.</p> <p>Keywords: Live commerce, Middle-aged women, Purchasing behavior, Technology Acceptance Model (TAM), Perceived Value Theory (PVT), Social Influence Theory (SIT), Structural Equation Modeling (SEM), Consumer trust, E-commerce.</p>

opencc-by-4.0Sep 2024View details →
zenodo40/100

Ivercori Dataset and dictionary. Ivermectin impact in COVID-19 pneumonia mortality and need of respiratory support

<p>Dataset of IVERCORI and variables dictionary. IVERCORI is a propensity matched score retrospective study that analyses the impact of ivermectin in COVID-19 pneumonia in-hospital mortality and need of respiratory support:</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Dictionary of Emotional and Nonemotional Frames

<p>This resource is a&nbsp;dictionary of frames-to-emotion associations. It&nbsp;contains&nbsp;multiple FrameNet frames&nbsp;and the&nbsp;strength of their association&nbsp;to &quot;emotionality&quot;,i.e., the degree to which a frame expresses&nbsp;an emotion, irrespective of what specific emotion that is.</p>

opencc-by-4.0Aug 2023View details →
edi40/100

The ecocomDP Annotation Dictionary

This data package contains a comprehensive set of semantic annotations (URIs and labels) from datasets in the ecocomDP format published in EDI. The table of annotations, referred to as the ecocomDP Annotation Dictionary, can be viewed in RStudio using the view_annotation_dictionary function of the ecocomDP R package.

openCC0Sep 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record