Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
14
datasets available to search
ShareScore release 0.9.0
Dataset results
14 results for “lexical database”
Cariban Lexical Database (CaLeD)
<p>This dataset contains a comprehensive collection of lexical items from various languages within the Carib linguistic family. It is structured to facilitate computational historical linguistics analysis, offering detailed information on language characteristics, word forms, and cognacy judgments. The data is curated to support research in linguistic typology, historical linguistics, and related fields.</p> <h4>Data Structure</h4> <p>The dataset is presented in a TSV (Tab-Separated Values) format, ensuring easy integration with common data analysis tools. Each lexical item in the dataset is detailed with multiple linguistic attributes, including phonological transcriptions, morphological analysis, and cognacy information. The following table summarizes the fields included in the dataset:</p> <table> <tbody> <tr> <th>Field Name</th> <th>Data Type</th> <th>Description</th> </tr> </tbody> <tbody> <tr> <td>ID</td> <td>string</td> <td>Unique identifier for each dataset entry.</td> </tr> <tr> <td>ID_lang</td> <td>string</td> <td>Unique identifier for the language within the dataset.</td> </tr> <tr> <td>Glottocode</td> <td>string</td> <td>Code uniquely identifying the language in the Glottolog database.</td> </tr> <tr> <td>Glottolog_Name</td> <td>string</td> <td>Name of the language as recorded in the Glottolog database.</td> </tr> <tr> <td>ISO639P3code</td> <td>string</td> <td>ISO 639-3 code for the language.</td> </tr> <tr> <td>ID_param</td> <td>string</td> <td>Unique identifier for the linguistic parameter or concept within the dataset.</td> </tr> <tr> <td>Concepticon_ID</td> <td>integer</td> <td>Identifier for the concept in the Concepticon database.</td> </tr> <tr> <td>Concepticon_Gloss</td> <td>string</td> <td>Gloss or definition of the concept from the Concepticon database.</td> </tr> <tr> <td>Value</td> <td>string</td> <td>Value of the linguistic data point, typically a word or phrase in the language.</td> </tr> <tr> <td>Form</td> <td>string</td> <td>Phonetic or phonological transcription of the linguistic data point.</td> </tr> <tr> <td>Segments</td> <td>string</td> <td>Further phonetic or phonological breakdown of the form.</td> </tr> <tr> <td>Source</td> <td>string</td> <td>Reference to the source or citation where the data was obtained.</td> </tr> <tr> <td>Morphemes</td> <td>string</td> <td>Morphological breakdown of the form.</td> </tr> <tr> <td>SimpleCognate</td> <td>integer</td> <td>Cognacy judgment, indicating whether the form is cognate with forms of the same meaning in related languages.</td> </tr> <tr> <td>PartialCognates</td> <td>string</td> <td>Partial cognacy coding, detailing the cognacy of individual segments or morphemes.</td> </tr> </tbody> </table> <h4>Intended Use</h4> <p>This dataset is intended for researchers and linguists specializing in the Carib linguistic family. It provides valuable insights into the lexical similarities and differences across the languages within this family, supporting studies on language evolution, relationships, and structure.</p> <h4>Additional Resources</h4> <ol> <li> <p><strong>Metadata for Validation</strong>: This dataset comes with comprehensive metadata following the Frictionless Data standard, ensuring that the data structure and types are accurately described for validation purposes. This metadata aids in maintaining the integrity and usability of the data across various computational platforms and research projects.</p> </li> <li> <p><strong>CLDF Version Available</strong>: For researchers utilizing the Cross-Linguistic Data Formats (CLDF), a version of this dataset is available in CLDF specifications. This version is provided as a zipped file, facilitating easier distribution and handling.</p> </li> </ol>
CLDF dataset derived from Carroll et al. "Yamfinder: The Southern New Guinea Lexical Database"
<p>Cite the source of the dataset as:</p> <blockquote> <p>Carroll, Matthew J., Barth, Wolfgang, Nicholas Evans, I Wayan Arka, Christian Döhler, Eri Kashima, Volker Gast, Tina Gregor, Kate L. Lindsey, Julia Miller, Emil Mittag, Bruno Olsson, Dineke Schokkin, Jeff Siegel, Charlotte van Tongeren, Kyla Quinn. [DATE ACCESSED]. Yamfinder: Southern New Guinea Lexical Database. Available online at: http://www.yamfinder.com</p> </blockquote>
TuLeD. Tupían Lexical Database
<p>Cite the source of the dataset as:</p> <blockquote> <p>Fabrício Ferraz Gerardi, Stanislav Reichert, Carolina Aragon, Johann-Mattis List, & Tim Wientzek. (2021). TuLeD: Tupían lexical database. Max Planck Institute for Evolutionary Anthropology: Leipzig</p> </blockquote>
CLDF dataset derived from Sagart et al.'s "Sino-Tibetan Database of Lexical Cognates" from 2019
<p>Cite the source of the dataset as:</p> <blockquote> <p>Laurent Sagart, Jacques, Guillaume, Yunfan Lai, and Johann-Mattis List (2019): Sino-Tibetan Database of Lexical Cognates. Jena: Max Planck Institute for the Science of Human History.</p> </blockquote>
Ume Saami lexical data from Álgu database with modernized spelling
<p>The files in this archive contain several processed versions of Ume Saami dictionary data originating from (Schlachter 1958). The data has been retrieved from the Álgu database <<a href="http://kaino.kotus.fi/algu/">http://kaino.kotus.fi/algu/</a>> and checked against the original dictionary for headword variants. The final versions have the headword mechanically converted to (approximate) the current Ume Saami orthography (as in Barruk 2018). In addition to an alphabetized list, a reverse-alphabetized (a tergo) file is provided.</p> <p>For convenience, .xlsx versions of the final files are provided in addition to the plain-text (.tsv) versions (in zip file).<br> The orthography conversion program can be found at <<a href="https://doi.org/10.5281/zenodo.4162535">https://doi.org/10.5281/zenodo.4162535</a>>.</p> <p><strong>NOTE: The mechanically converted lemmas do not always fully correspond to the modern orthography and should be checked before use in further applications!</strong></p> <p>Authors: Juha Kuokkala (data processing and programming), Wolfgang Schlachter (original dictionary data).</p>
Figure 2. The structures of the two tables from the Dex Online database-ADX – Agent for Morphologic Analysis of Lexical Entries in a Dictionary
<p>Dex Online is a project initiated and coordinated by Catalin Francu [3]. He intended to<br> realise an online database for all the words in the Romanian language, using the main explanatory<br> dictionaries, dictionaries of synonyms, neologisms, published by the Romanian Academy and other<br> scientific forums.<br> The database was completed by volunteers, similarly to the Wikipedia system. They actually<br> transcribed the information from different important dictionaries, but many words have been<br> electronically entered by two companies (Siveco and Litera International Publishing House).</p>
CLDF dataset derived from MaJeLeD (Macro-Je Lexical Database)
<p>Cite the source of the dataset as:</p> <blockquote> <p>Gerardi, F. F., Tresoldi, T. (2024)</p> </blockquote>
Linkeast: A lexical database of Eastern Polynesian Languages
<p>This dataset is an extraction of the data contained in <em>LinkEast</em>, a lexical database of Eastern Polynesian languages. <em>LinkEast </em>is housed at the Université de la Polynésie française, within <em>Anareo</em>, a digital infrastructure dedicated to research on and documentation of the languages of French Polynesia. This data (v1.0) is based primarily on the Linguistic Atlas of French Polynesia (Charpentier & François, 2015). </p> <p> </p> <p>References:<br> Charpentier, Jean-Michel, & François, Alexandre, 2015, <em>Atlas linguistique de la Polynésie française</em>, Berlin et Papeete, Mouton de Gruyter et Université de la Polynésie française.</p> <p> </p>
DravLex: A Dravidian lexical database.
<p>First release of DravLex</p>
Lexical Database of Bororoan
No description provided.
Mondzish lexical database
<p>Mondzish (Mangish) lexical database, including transcriptions of my audio recordings collected in China in from 2012-2015.</p>
Lolo-Burmese lexical database
<p>Comparative vocabulary of Lolo-Burmese languages, with glosses and word lists merged from Shintani (2001) and Lama (2012). This spreadsheet is a work in progress.</p>
Central Loloish lexical isogloss database
<p>Database of potential lexical isoglosses in Central Loloish (Ngwi) languages</p>
Mienic lexical isogloss database
<p>Database of identified or potential lexical isoglosses in Mienic lects</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.