Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

132

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

132 results for “vocabularies”

Learn how ShareScore rates datasets ↗
zenodo36/100

CLDF dataset derived from Greenhill et al.'s "Austronesian Basic Vocabulary Database" from 2020 focusing on Oceanic languages

<p>Cite the source of the dataset as:</p> <blockquote> <p>Greenhill, S.J., Blust. R, &amp; Gray, R.D. (2008). The Austronesian Basic Vocabulary Database: From Bioinformatics to Lexomics. Evolutionary Bioinformatics, 4:271-283.</p> </blockquote>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Deccani-English Vocabulary

<p>Robert Blair Munro Binning, <em>Deccani-English Vocabulary</em> (Madras, circa 1844).</p>

opencc-by-4.0Feb 2019View details →
zenodo36/100

Kokinwakashu Hyoshaku by Motoomi Kaneko translation sentence vocabulary dataset

<p>"Kokin wakashu hyoshaku" by Kaneko Motoomi.</p> <p>古今和歌集評釈 金子元臣著 1927年</p> <p>It was published in 1927.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

NCAS General Data Standard Controlled Vocabularies

Controlled Vocabularies for the NCAS General Data Standard

opencc-by-4.0Nov 2024View details →
zenodo36/100

Vocabulary list from the 2022 Vocabulary Workshop

<p>Across the workshop, participants were asked to review and add to a list of vocabularies used within their domain, and across domains. To support post-workshop discussion of questions such as</p> <ul> <li>Which vocabularies are required across domains?</li> <li>Which vocabularies are domain-specific but are needed by other domains?</li> <li>How might guidance on vocabulary selection be provided?</li> <li>What to do about multiple copies of similar vocabularies?</li> </ul>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Vocabulary tools list. Referred to within the 2022 Vocabulary workshop

<p><strong>Initially created during 2021.&nbsp; Discussed during the 2022 Vocabulary Workshop</strong></p> <p>&quot;This catalogue is a copy of a draft output of a 2021 Vocabulary Workshop. Participants were interested in putting together a catalogue of vocabulary tools. To address questions like...<br> &quot;If you want to do something with vocabularies, and there are lots of tools to choose from, how do you decide which one to use?&quot;</p> <p>The first iteration of this resource was created by Nick Car (SurroundAustralia), Edmond Chuc (TERN), Kheeran Dharmawardena (DAWE), Michael Lawley(CSIRO) and Simon Cox (CSIRO) &quot;</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

VOYAGE: A Large Collection of Vocabulary Usage in Open RDF Datasets

<p><strong>List of files:</strong></p> <ul> <li>odps.json: for each of the accessed ODPs, its name, URL, API type, API URL, and the IDs of RDF datasets collected from it <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;name&#39; (string), &#39;URL&#39; (string), &#39;API type&#39; (string), &#39;API URL&#39; (string), and &#39;collected datasets IDs&#39; (list of integers)</p> </li> </ul> </li> <li>datasets.json: for each of the crawled RDF datasets, its ID, title, description, author, license, dump file URLs, and PLDs <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;ID&#39; (integer), &#39;title&#39; (string), &#39;description&#39; (string), &#39;author&#39; (string), &#39;license&#39; (string), &#39;dump file URLs&#39; (list of strings), and &#39;PLDs&#39; (list of strings)</p> </li> </ul> </li> <li>deduplicated_datasets.json: the IDs of the deduplicated RDF datasets and whether they are in the LOD Cloud <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;ID&#39; (integer) and &#39;in LOD Cloud&#39; (boolean)</p> </li> </ul> </li> <li>terms.json: the extracted classes, properties, and the IDs of RDF datasets using each term <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;term&#39; (string), &#39;is class&#39; (boolean), &#39;is property&#39; (boolean), and &#39;used in dataset IDs&#39; (list of integers)</p> </li> </ul> </li> <li>vocabularies.json: the extracted vocabularies, the classes and properties in each vocabulary, and the IDs of RDF datasets using each vocabulary <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;vocabulary&#39; (string), &#39;classes&#39; (list of strings), &#39;properties&#39; (list of strings), and &#39;used in dataset IDs&#39; (list of integers).</p> </li> </ul> </li> <li>edps.json: the extracted distinct EDPs and the IDs of RDF datasets using each EDP <ul> <li> <p>JSON structure: a list of objects, where each object contains the following attributes - &#39;classes&#39; (list of strings), &#39;forward properties&#39; (list of strings), &#39;backward properties&#39; (list of strings), and &#39;used in dataset IDs&#39; (list of integers)</p> </li> </ul> </li> <li>clusters.json: the clusters of vocabularies generated by MV-ITCC and LDA <ul> <li> <p>JSON&nbsp;structure:&nbsp;{&quot;LDA&quot;:&nbsp;{&quot;vocabularies&quot;:&nbsp;{VOCABULARY_CLUSTER_ID_1:&nbsp;[LIST_OF_VOCABULARIES],&nbsp;VOCABULARY_CLUSTER_ID_2:&nbsp;[LIST_OF_VOCABULARIES],&nbsp;...}},&nbsp;&quot;MV-ITCC&quot;:&nbsp;{&quot;vocabularies&quot;:&nbsp;{VOCABULARY_CLUSTER_ID_1:&nbsp;[LIST_OF_VOCABULARIES],&nbsp;VOCABULARY_CLUSTER_ID_2:&nbsp;[LIST_OF_VOCABULARIES],&nbsp;...},&nbsp;&quot;dataset&nbsp;IDs&quot;:&nbsp;{DATASET_CLUSTER_ID_1:&nbsp;[LIST_OF_DATASET_IDS],&nbsp;DATASET_CLUSTER_ID_2:&nbsp;[LIST_OF_DATASET_IDS],&nbsp;...}}}</p> </li> </ul> </li> </ul>

opencc-by-4.0May 2023View details →
zenodo36/100

A controlled vocabulary for research and innovation in the field of Cultural Heritage & Heritage Sciences

<p>This controlled vocabulary of keywords related to the field of Cultural Heritage and Heritage Sciences was built by SIRIS Academic in collaboration with IRPET (the Regional Institute for Economic Planning of Tuscany) and the ISPC (Institute of Heritage Science of CNR), in order to identify Cultural related research, development, and innovation activities. The work was carried out by consulting domain experts&#39; advice, and it was ultimately applied to inform regional strategies on Cultural Heritage and research and innovation policy.</p> <p>The aim of this vocabulary is to enable one to retrieve texts (e.g. R&amp;D projects and scientific publications) featuring the concepts included in the present vocabulary in their titles and abstracts, assuming that these records have a certain contribution of applications, techniques and issues, in the domain of Cultural Heritage and Heritage Sciences.</p> <p>The aim of this classification is to identify research products in the domain of Cultural Heritage, ranging from documents in some of its &ldquo;traditional&rdquo; disciplines, but also from documents emerging from interdisciplinary projects that apply novel areas and technologies in the domain of Cultural Heritage. The identification of texts in the domain of Cultural Heritage requires a task of text classification. Developing a method that could be applied to decide if a text can be relevant or have some relation to the domain of Cultural Heritage is a challenging task. The definition of what Cultural Heritage is and what it includes is a complex activity, even for domain experts. This is in particular because Cultural Heritage is quite a broad field of knowledge, and there is no full agreement on where the borders of the domain are. To define the scope of the perimeter, in this project, many of the available definitions were taken into account.</p> <p>Because of the high number of resources available in the domain, among thesauruses and taxonomies, the construction of a weakly-supervised controlled vocabulary was considered as the best way of retrieving documents in the domain. Since there is no annotated corpus/dataset of research texts in the domain capable of generalising the diversity of publications that can be related to the cultural domain, but stemming from different disciplines, we have opted for a text classification technique based on rules &ndash; specifically, a weakly-supervised controlled vocabulary.</p> <p>As defined by the Getty Institute, a controlled vocabulary is an organized arrangement of words and phrases used to index content and/or to retrieve content through browsing or searching. It typically includes preferred and variant terms and has a defined scope or describes a specific domain. The purpose of controlled vocabularies is to organize information and to provide terminology to catalogue and retrieve information. While capturing the richness of variant terms, controlled vocabularies also promote consistency in preferred terms and the assignment of the same terms to similar content (Harping, 2010).</p> <p>In short, Cultural Heritage is a rather abstractly-defined field, and Heritage Science is a particularly &ldquo;fuzzy&rdquo; field within Cultural Heritage. One of the main limitations of the approach we used is that the controlled vocabularies never capture all the lexical and linguistic variants of a term, and we may miss relevant texts if we cannot find the correct pattern to match during the search. But on the other hand, the controlled vocabulary is built from available vocabularies and thesauruses in the domain of Cultural Heritage, which are large resources. All the concepts in these resources are not included directly in the controlled vocabularies, because they would add noise to the classification. Therefore, the automatic weak supervision and a human curation of the final controlled vocabulary is fundamental for achieving correct results.</p> <p>The controlled vocabulary is built taking advantage of these four resources:</p> <ul> <li>The <a href="https://www.getty.edu/research/tools/vocabularies/aat/"><strong>Art and Architecture Thesaurus (AAT)</strong></a>: this is a structured vocabulary with approximately 34,000 concepts, including 131,000 words, descriptions and other information related to art, architecture, decorative arts, archival material and material culture, commonly used for cataloguing and for information retrieval.</li> <li> <p>Some cultural heritage categories in <strong><a href="https://en.wikipedia.org/wiki/Category:Cultural_heritage">Wikipedia</a> </strong>and <strong><a href="https://dbpedia.org/page/Cultural_heritage">DBpedia</a></strong>: these categories have been used to collect all related articles and subcategories, in order to obtain relevant, similar and specific instances of concepts linked to the domain.&nbsp;</p> </li> <li> <p>The <strong><a href="https://www.riches-project.eu/riches-taxonomy.html">RICHES Taxonomy</a></strong>: this taxonomy is a theoretical framework of related terms and their definitions, referring to the new concepts in the digital era, with the aim of defining the scope of some digital technologies applied to cultural heritage.</p> </li> <li> <p><strong><a href="https://www.heritagedata.org/blog/">Heritage Data - Linked Data Vocabularies for Cultural Heritage</a></strong>: a dataset which includes several cultural heritage thesauruses and vocabularies and is recognised as a reference point in the United Kingdom in the domain of cultural heritage.</p> </li> </ul> <p>The collection of concepts extracted from these four resources was composed of more than 60,000 terms, which have been refined as described in the next section.</p> <p>&nbsp;</p> <p><strong>## Automatic validation of the controlled vocabulary</strong></p> <p>In order to refine the collection of concepts to have a final set of relevant concepts and terms in the domain of Cultural Heritage, a semi-automatic validation has been applied to remove the irrelevant, too general, and ambiguous terms.</p> <p>To keep the relevant ones, the <a href="https://ncses.nsf.gov/pubs/nsb20206/specialization-and-impact-analysis-combined#:~:text=The%20specialization%20index%20(SI)%20is,the%20total%20output%20across%20all">specialization index (SI) </a>metric has been calculated for each of the keywords in the collection. In this case, the SI can be obtained measuring the fraction of publications with a keyword in a set of publications in the domain of Cultural Heritage and normalizing over the fraction of publications in the open domain with that keyword.</p> <p>After the calculation of the SI, all the keywords below a certain threshold are removed, and a manual supervision step is applied in order to remove non-pertinent keywords. An example of this automatic validation can be observed in the next table:</p> <table> <tbody> <tr> <td> <p><strong>Keyword</strong></p> </td> <td> <p><strong>Specialization Index</strong></p> </td> <td> <p><strong>Automatic threshold</strong></p> </td> <td> <p><strong>Manual supervision</strong></p> </td> </tr> <tr> <td> <p>male</p> </td> <td> <p>0.27</p> </td> <td> <p>Removed</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>3-d laser scanning</p> </td> <td> <p>0.7</p> </td> <td> <p>Removed</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>78 rpm records</p> </td> <td> <p>20.7</p> </td> <td> <p>Accepted</p> </td> <td> <p>Removed</p> </td> </tr> <tr> <td> <p>vienna</p> </td> <td> <p>3.48</p> </td> <td> <p>Accepted</p> </td> <td> <p>Removed</p> </td> </tr> <tr> <td> <p>radiocarbon dating</p> </td> <td> <p>13.6</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>graffiti</p> </td> <td> <p>25</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>bark painting</p> </td> <td> <p>20.7</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>pompeii</p> </td> <td> <p>16.23</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> </tbody> </table> <p>The SI of the final keywords can be used as a probabilistic metric for each keyword.</p> <p>The final list of keywords was manually curated by domain experts.</p> <p>&nbsp;</p> <p><strong>## Evaluation of the controlled vocabulary</strong></p> <p>The final controlled vocabulary was evaluated with an external dataset with the aim of calculating its degree of precision. The evaluation dataset was composed of a collection of articles in 4 journals unequivocally considered to fall within the domain of Cultural Heritage. These four journals were: <em>(1) Journal Of Cultural Heritage, (2) Journal On Computing And Cultural Heritage, (3) Journal Of Cultural Heritage Management And Sustainable Development and (4) Digital Applications In Archaeology And Cultural Heritage.</em> This collection was composed of 5,000 articles, considered as the positive set, and another collection of randomly selected 5,000 articles outside of the Cultural Heritage domain, considered as the false set.</p> <p>The Cultural Heritage vocabulary was applied to the evaluation data set, obtaining a 95% of precision. After a set of improvements on the vocabulary, based on the exploration of publications not identified in the first test and the false positive results, we obtained a 98% of precision. The application of the vocabulary taking advantage of the probability of each keyword as its weight of being in the domain did not improve the results, and for this reason the probabilistic approach was discarded.</p> <p>&nbsp;</p> <p>##&nbsp;<strong>Using the vocabulary to classify publications concerning Cultural Heritage</strong></p> <p>The definition of the vocabulary does not, per se, allow to identify research contributions in Cultural Heritage: this is performed by actually matching the terms in the controlled vocabulary to the content of the gathered research textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance.</p> <p>In the following table we present some examples of the tagging process on some abstracts:</p> <table> <tbody> <tr> <td> <p><strong>Publication title</strong></p> </td> <td> <p><strong>Publication abstract</strong></p> </td> </tr> <tr> <td> <p>Egocentric visitor localization and artwork detection in cultural sites using synthetic data</p> </td> <td> <p>Computer vision and machine learning can be used in <strong>cultural heritage to augment the experience of visitors during the exploration of the cultural site</strong>, as well as to assist its management. To achieve such goals, two fundamental tasks should be addressed, i.e., localizing <strong>visitors and recognizing the observed artworks</strong>. Wearable cameras offer a convenient setting to address both tasks through the analysis of images acquired from the visitors&rsquo; points of view. However, the engineering of approaches to address such tasks generally requires large amounts of labeled data. We propose a tool which can be used to collect and automatically label synthetic visual data suitable to study image-based localization and artwork detection. The tool simulates a virtual agent navigating the <strong>3D model of a real cultural site</strong> and automatically captures video frames along with the related ground truth camera poses and semantic masks indicating the position of artworks. We generate a dataset of synthetic images starting from the 3D model of a <strong>museum located in Siracusa</strong>, Italy. The experiments suggest that the proposed tool allows to drastically reduce the effort needed to collect and label data, providing a means to generate large-scale datasets suitable to study localization and <strong>artwork detection in cultural sites</strong>.</p> </td> </tr> <tr> <td> <p>Discovering Leonardo with artificial intelligence and holograms: A user study</p> </td> <td> <p>Cutting-edge visualization and interaction technologies are increasingly used in<strong> museum exhibitions</strong>, providing novel ways to engage visitors and enhance their <strong>cultural experience</strong>. Existing applications are commonly built upon a single technology, focusing on visualization, motion or verbal interaction (e.g., high-resolution projections, gesture interfaces, chatbots). This aspect limits their potential, since museums are highly heterogeneous in terms of visitors profiles and interests, requiring multi-channel, customizable interaction modalities. To this aim, this work describes and evaluates an artificial intelligence powered, interactive holographic stand aimed at describing <strong>Leonardo Da Vinci&#39;s art</strong>. This system provides the users with accurate<strong> 3D representations of Leonardo&#39;s machines</strong>, which can be interactively manipulated through a touchless user interface. It is also able to dialog with the users in natural language about Leonardo&#39;s art, while keeping the context of conversation and interactions. Furthermore, the results of a large user study, carried out during art and tech exhibitions, are presented and discussed. The goal was to assess how users of different ages and interests perceive, understand and explore <strong>cultural objects </strong>when holograms and artificial intelligence are used as instruments of knowledge and analysis.</p> </td> </tr> <tr> <td> <p>Hybrid query expansion using lexical resources and word embeddings for sentence retrieval in question answering</p> </td> <td> <p>Question Answering (QA) systems based on Information Retrieval return precise answers to natural language questions, extracting relevant sentences from document collections. However, questions and sentences cannot be aligned terminologically, generating errors in the sentence retrieval. In order to augment the effectiveness in retrieving relevant sentences from documents, this paper proposes a hybrid Query Expansion (QE) approach, based on lexical resources and word embeddings, for QA systems. In detail, synonyms and hypernyms of relevant terms occurring in the question are first extracted from MultiWordNet and, then, contextualized to the document collection used in the QA system. Finally, the resulting set is ranked and filtered on the basis of wording and sense of the question, by employing a semantic similarity metric built on the top of a Word2Vec model. This latter is locally trained on an extended corpus pertaining the same topic of the documents used in the QA system. This QE approach is implemented into an existing QA system and experimentally evaluated, with respect to different possible configurations and selected baselines, for the <strong>Italian language and in the Cultural Heritage domain</strong>, assessing its effectiveness in retrieving sentences containing proper answers to questions belonging to four different categories.</p> </td> </tr> <tr> <td> <p>&quot;3D reconstruction and validation of historical background for immersive VR applications and games: The case study of the Forum of Augustus in Rome&quot;</p> </td> <td> <p>&quot;In the last decades, thanks to the success of the video games industry, the sector of technologies applied to cultural heritage has begun to envisage, in this domain, new possibilities for the <strong>dissemination of heritage and the study of the past </strong>through edutainment models. More recently, experimentation in the field of<strong> virtual archaeology </strong>has led to the development of virtual museums and interactive applications. Among these, the &ldquo;serious game&rdquo; segment &ndash; the<strong> application of interactive technologies to the cultural heritage domain</strong> &ndash; is rapidly growing, also including immersive VR technologies. Applied VR games and applications are characterized by a thorough <strong>historical background and a validated 3D reconstruction</strong>. Indeed, producing such products requires a tailored workflow and large effort in terms of time and professionals involved to guarantee such faithfulness. Drawing on our previous work in the<strong> field of virtual archaeology</strong> and referring to recent experiences related to the deployment of applied VR games on PlayStation VR, we describe and assess a workflow for the production of <strong>historically accurate 3D assets</strong>, targeting interactive, immersive VR products. The workflow is supported by the case study of the <strong>Forum of Augustus </strong>and different output applications, highlighting peculiarities and issues emerging from a multi and interdisciplinary approach.</p> </td> </tr> </tbody> </table> <p>Through this classification process, we identified projects and publications related to heritage, with different levels of relationship and relevance, but mostly relevant to understanding the research competencies in the domain. The resulting research records were reviewed by experts in the domain, given the occurrence of some false positives.</p> <p>Among the main strengths of this step, it&rsquo;s worth mentioning the fact that the vocabulary is broad and not restricted to the field of Heritage Science (that is, to STEM applications in Cultural Heritage), as it takes advantage of a variety of available resources. Moreover, by looking directly at the textual data, instead of using the assigned bibliometric areas, we can better capture interdisciplinary research. The limitations of this approach were presented at the beginning of this document: for example, relevant texts could be missed if the correct pattern to match during the search is not found.</p> <p>&nbsp;</p> <p><strong>## Vocabulary of concepts related to Key Enabling Technologies in the domain of Cultural Heritage and Culture</strong></p> <p>For the development of this vocabulary, the definition of key enabling technologies in the domain of Cultural Heritage and Culture, was based on reference of the <a href="http://www.irpet.it/archives/53165">report &#39;Technologies, Cultural Heritage and Culture&#39; published on&nbsp; March 2019</a> by IRPET.</p> <p>A vocabulary for each Key Enabling Technology (hereafter, KET) was prepared by extracting the relevant concepts, words, technologies and examples from the Platform Report document &#39;Technologies, Cultural Heritage and Culture, within APPENDIX A. DESCRIPTION OF MAIN TECHNOLOGIES FOR ROADMAP (p. 45-61). Each vocabulary contains a set of terms divided into subdomains.</p> <p>The KETs have been divided into the following six groups:</p> <ul> <li> <p>ICT</p> </li> <li> <p>PHOTONICS, MICRO- AND NANO-ELECTRONICS</p> </li> <li> <p>PLATFORMS</p> </li> <li> <p>NANO AND BIOTECHNOLOGY, ADVANCED MATERIALS</p> </li> <li> <p>PARTICLE ANALYTICAL SYSTEMS</p> </li> </ul> <p>The initial keywords extracted from the document were enriched following the approach based on semantic keyword enrichment based on combination of concurrent keywords and word embeddings&nbsp; (Duran-Silva et al., 2019; Duran-Silva et al., 2021).</p> <p>This second vocabulary has to be used in combination with the Cultural Heritage vocabulary to capture KETs&nbsp;within the domain of cultural heritage.</p> <p>&nbsp;</p> <p>##&nbsp;<strong>Use of the controlled vocabulary</strong></p> <p>The definition of the vocabulary does not, per se, allow identifying STI contributions to the domain: this activity in fact boils down to actually matching the terms in the controlled vocabulary to the content of the gathered STI textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance. Some relatively ambiguous keywords (which may match unwanted pieces of text), have a set of associated &ldquo;extra&rdquo; terms. These &ldquo;extra&rdquo; terms are defined as further terms that must co-appear, in the same sentence, together with their associated ambiguous keywords. SIRIS Academic has developed the&nbsp;<a href="https://github.com/sirisacademic/VocTagger">voc_tagger tool</a>, a multiprocess information extraction system able to identify hidden knowledge in textual documents using &ldquo;controlled vocabularies&rdquo;, openly available at GitHub and compatible with these controlled vocabularies.</p> <p>&nbsp;</p> <p><strong>## Bibliography</strong></p> <p>Harpring, P. (2010). Introduction to controlled vocabularies: terminology for art, architecture, and other cultural works. Getty Publications.</p> <p>Nicolau Duran-Silva, Enric Fuster, Francesco Alessandro Massucci, C&eacute;sar Parra-Rojas, Arnau Quinquill&agrave;, Fernando Roda, Bernardo Rondelli, Nicandro Bovenzi, &amp; Chiara Toietta. (2021). A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI) (Version 2) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5591987</p> <p>Duran-Silva, Nicolau, Fuster, Enric, Massucci, Francesco Alessandro, &amp; Quinquill&agrave;, Arnau. (2019). A controlled vocabulary defining the semantic perimeter of Sustainable Development Goals (1.2) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3567769</p>

opencc-by-sa-4.0Nov 2021View details →
zenodo36/100

SatTerm experience: vocabulary control and facet analysis help improve the software requirement solicitation process

<p>One of the most difficult steps in the software development process is moving from&nbsp;requirements written in natural, uncontrolled language, to the formalisms required by the design&nbsp;modelling languages. To solve this issue, practitioners should pay attention to the possibility of&nbsp;applying vocabulary control and knowledge representation techniques to produce better&nbsp;specifications. The use of controlled vocabularies and the modelling of the conceptual relationships&nbsp;between concepts in a specific domain are expected to improve the quality of the specifications.<br> Vocabulary control and semantic modelling are promising tools to avoid the most frequent problems&nbsp;in the requirements specification process: lack of consistency and ambiguity.&nbsp;This paper provides a detailed description of the development process of an ontology used for&nbsp;requirements modelling in the area of satellite control systems. The process applied is based on well-established practices and guidelines applied for the construction of controlled vocabularies and&nbsp;faceted classifications schemas. Engineers can use the ontology when writing system specifications&nbsp;using predefined templates. The use of this ontology ensures the consistency of the specifications&nbsp;written by different engineers improves the communication with other parties involved in the system&nbsp;construction activities and sets the foundations for a semi-automated generation of models for&nbsp;subsequent design activities.</p>

opencc-by-4.0Jul 2013View details →
dryad36/100

Supplementary methods and data for: Dogs with a vocabulary of object label remember labels for at ‎least two years

Open the record for dataset details and reuse information.

publicApr 2024View details →
zenodo32/100

A controlled vocabulary defining the semantic perimeter of Sustainable Development Goals

<p>A set of controlled terms that define the scope and breadth of <a href="https://sustainabledevelopment.un.org/">Sustainable Development Goals (SDGs) as defined by the United Nations</a>.&nbsp; These terms may be used to tag and index textual records in accordance with SDGs.</p> <p>The vocabulary is constructed by means of the following steps:</p> <ol> <li>An initial set of terms per SDG target is built by extracting key terms from the UN official list of Goals, Targets and Indicators</li> <li>The list is manually enriched by performing a review of the literature produced around SDGs and by compiling lists of pertinent words per Target mentioned by the reviewed documents</li> <li>A reference textual corpus is downloaded by searching for the initial set terms defined at step 1. and 2. The corpus is used to train a Word2Vec word embedding model (a machine learning model based on neural networks).</li> <li>The terms&rsquo; list is then enriched by means of automatic methods, which are run in parallel: <ul> <li>The trained Word2Vec model is used to select, among the indexed keywords of the reference corpus, all terms &ldquo;semantically close&rdquo; to the initial set of words. This step is carried out to select terms that might not appear in the texts themselves, but that were deemed pertinent to label the textual records.</li> <li>Further terms that are mentioned in the texts of the reference corpus and that are valued by the trained Word2Vec model as &ldquo;semantically close&rdquo; to the initial set of words are also retained. This step is performed to include in the controlled vocabulary a series of terms that are related to the focus of the SDGs and which are used by practitioners.</li> <li>An automated algorithm is used to retrieve, from the APIs of WikiPedia a series of terms that have some categorical relationships (i.e. those that are indexed as &ldquo;a broader concept of&rdquo;, or &ldquo;equivalent to&rdquo; in DBpedia) with the initial set of words.</li> </ul> </li> <li>The final list produced by steps 1-4 s finally manually revised</li> </ol>

opencc-by-sa-4.0Dec 2019View details →
zenodo32/100

A collection of North Tujia (Bifzivsar 北部土家语) vocabulary and textual passages for use in NLP

<p>This is a collection of North Tujia (Bifzivsar 北部土家语) vocabulary and textual passages for use in NLP.</p>

openother-pdJan 2021View details →
zenodo32/100

Figure 2 The mean scores of vocabulary and morphemes, reading, listening, speaking, and writing.

<p><strong>Figure 2 for &#39;Mind Mapping improved the Effect of Medical English Learning for non-Native English Speakers in non Target Language Environment:A Randomized Controlled Trial&#39;</strong></p>

opencc-by-sa-4.0Dec 2015View details →
zenodo32/100

A Metadata Application Profile for KOS Vocabulary Registries

<p>This paper reports on the end-products of a Dublin Core Application Profile for KOS Resources (KOS-AP) developed by a DCMI Task Group, which builds on work done by the NKOS group during the last decade. Included are user scenarios, a FRBR-based domain model, the attributes of KOS <i>work, expression,&nbsp;</i>and <i>manifestation</i> and the associated metadata elements defined in the context of user tasks. With the testing of real KOS systems that involve multiple editions, languages, delivery formats, and derivations, (such as the Dewey Decimal Classification 22nd edition and of the ASIS Thesaurus 3rd edition), the authors identified some unique issues related to KOS descriptions, for example the designation of a '<i>work'</i> in theory and practice, and the shift of management level from the whole KOS expression to concept- and label-based electronic systems. &nbsp;This KOS-AP effort is becoming more meaningful with the fast development of Linked Data, the success of which depends heavily on the sharing of standardized value vocabularies. The theoretical exploration, the conceptual model, and the core elements to be used for describing and accessing KOS products can be implemented by KOS registries and used in microdata of KOS Websites, as well as be adopted for use by other types of specifications that share common characteristics of KOS, such as frequently updated, translated, and derived handbooks, technical manuals, and schemas. &nbsp;</p>

openJul 2013View details →
zenodo32/100

Analysis of Questions in Vocabulary Instruction in Grade 2 Mother Tongue, Filipino, and English Self-Learning Modules

<p>This is a master's thesis that analyzed the questions on vocabulary instruction in Grade 2 Mother Tongue, Filipino, and English self-learning modules.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

SCENT for GLAM: a tool for giving meaning to professional controlled vocabularies

<p><span>This paper discusses SCENT for GLAM, which stands for Semantic and Collaborative Environment for a Network of Terminology for Galleries, Libraries, Archives and Museums. SCENT for GLAM offers a number of features that allow an institution to create or import terminology, to convert it into SKOS format, to edit it, make alignments with other terminologies from the same institution or another one and finally publish and share this terminology and alignments. The idea is to provide a repository of terminologies that will include all of the cultural sector concepts.</span></p>

openJul 2015View details →
zenodo32/100

Ciência Vitae controlled vocabulary - Current status

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

Vietic and Early Chinese Grammatical Vocabulary in Vietnamese: Native Vietic and Austroasiatic Etyma versus Early Chinese Loanwords

<p>In this talk, I present Vietnamese, Vietic, and Old and Middle Chinese lexical data with a focus on function words. The data consists of (a) about 60 Vietic reconstructions and (b) about 60 early Chinese loanwords in Vietnamese (i.e. those borrowed prior to Late Middle Chinese and the speciation of Viet-Muong). Core Vietic function words have been retained in Vietnamese (e.g., pronouns, numbers, location words, etc.), while Sinitic contributed words with other structural functions (comparative words, conjunctions, measures, etc.). This data has ethnohistorical linguistic implications for Vietic before and after language contact with Sinitic.</p>

opencc-by-4.0May 2021View details →
zenodo32/100

Currungulla vocabulary

<p>Currungulla tribal dialect by W.H.W. [<a href="https://en.wikipedia.org/wiki/Currawilla">William Henry Watson</a>] of Currawilla Queensland. <em>Science of Man</em> 13(10):211, 13(11):231, 13(12):251</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

Ontology of time metrology vocabulary

<p>The advent of the digital era has put forward an urgent need for the digitization of metrology, and the digitization of metrology vocabularies is one of the fundamental and critical steps to achieve the digital transformation of metrology. Metrology vocabulary ontology can facilitate the exchange and sharing of data, which is an important way to achieve the digitization of metrology vocabulary. Time metrology vocabulary is a special and important part of the whole metrology vocabulary, and constructing its ontology can reduce the problems caused by semantic confusion, help the smooth progress of metrological work, and promote the digital transformation of metrology. Currently, the existing ontology for metrology vocabulary is primarily the MetrOnto ontology, but it lacks a systematic description of the vocabulary of time metrology. Based on Web Ontology Language (OWL), a bilingual time metrology vocabulary ontology OTMV is designed and constructed using a seven-step ontology development methodology.OTMV provides a standardized, interoperable, and unified architecture, which realizes the bilingual digital representation of time metrology vocabulary in the national standards of China, metrological technical specifications JJF1180-2007 "Glossary and Definition of Time and Frequency Metrology", JJG 2007-2015 "Time and Frequency Measuring Instruments", and relevant materials published by the China National Institute of Metrology.</p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record