Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
42
datasets available to search
ShareScore release 0.7.1
Dataset results
42 results for “Controlled vocabularies”
A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI)
<p><strong>A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI)</strong></p> <p>This controlled vocabulary of keywords related to the field of Artificial Intelligence (AI) was built by SIRIS Academic in collaboration with ART-ER (the R&I and sustainable development in-house agency of the Emilia-Romagna region in Italy) and the Generalitat de Catalunya (the regional government of Catalonia, Spain), in order to identify AI research, development and innovation activities. The work was carried out by consulting domain experts' advice and it was ultimately applied to inform regional strategies on AI and research and innovation policy.</p> <p>The aim of this vocabulary is to enable one to retrieve texts (e.g. R&D projects and scientific publications) featuring the concepts included in the present vocabulary in their titles and abstracts, assuming that these records have a certain contribution of applications, techniques and issues, in the domain of AI.</p> <p>The present effort was carried out because, despite the high number of contributions and technological developments in the field of AI, there is no closed or static vocabulary of concepts that allows to unequivocally define the boundaries of what should be considered “an Artificial Intelligence intellectual product” (or what should not). Indeed, the literature presents different definitions of the domain, with visions that could be contradictory. AI encompasses today a wide variety of subdomains, ranging from general purpose areas such as learning and perception to more specific ones such as autonomous vehicle driving, theorem proving, or industrial process monitoring. AI synthesises and automates intellectual tasks, and is therefore potentially relevant to any area of human intellectual activity. In this sense, it is a genuinely universal and multidisciplinary field. AI draws upon disciplines as diverse as cybernetics, mathematics, philosophy, sociology and economics.</p> <p>As a ground for the construction of the AI controlled vocabulary, an initial set of concepts was taken from different subdomains of the <em>ACM Computing Classification System 2012, </em> to define the boundaries of the AI domain. Notably, although some relevant AI subdomains have an independent category in the ACM taxonomy outside of AI, they have been included in the list of subdomains. In order to align the ACM taxonomical definition with the Catalan Strategy of AI, <em>CATALONIA.AI</em>, in <em>version 1 </em>of this resource the emerging area of AI Ethics was included in the vocabulary, while some other categories which are not relevant for the objectives were removed from the subdomains list. In the current <em>version 2</em>, the classification and the labels of the subdomains have been revised because of the evolution of the field. Some fields have been grouped in order to reduce the overlap between subdomains and to provide a taxonomy that makes more sense for the analysis of R&I ecosystems. </p> <p>The different subdomains in the versions are presented in the following table:</p> <table> <tbody> <tr> <td><strong>Version </strong></td> <td><strong>Subdomains</strong></td> </tr> <tr> <td> <p><em>Version 2</em></p> </td> <td> <p>(1) Machine learning and deep learning; (2) Computer Vision; (3) Natural Language Processing and speech recognition; (4) Intelligent agents, planning, scheduling, problem-solving, control methods, and search; (5) Expert Systems, Knowledge representation and reasoning; (6) AI Ethics.</p> </td> </tr> <tr> <td><em>Version 1</em></td> <td>(1) General, (2) Machine Learning, (3) Computer Vision, (4) Natural Language Processing, (5) Knowledge Representation and Reasoning, (6) Distributed Artificial Intelligence, (7) Expert Systems, Problem-Solving, Control Methods and Search and (8) AI Ethics.</td> </tr> </tbody> </table> <p>Although a keyword rule-based approach suffers from the major shortcomings of not capturing all the lexical and linguistic variants of specific concepts nor the context of the words - namely, keyword-based approaches would miss relevant texts if the specific pattern is not matched during the search - the present vocabulary allowed us to obtain fairly good results, due to the specificity of the concepts describing the AI domain. Furthermore, an understandable and transparent controlled vocabulary allows a better control of the final results and the final definition of the domain borders. Also, a plain list of terms allows a much easier and interactive engagement of interested stakeholders with different degrees of knowledge (such as, for instance, domain experts, policy-makers and potential users) who can make use of vocabulary to retrieve pertinent literature or to enrich the resource itself.</p> <p>The vocabulary has been built taking advantage of advanced language models and resources from knowledge datasets such as arXiv, DBpedia and Wikipedia. The resulting vocabulary comprises 833 keywords, and has been validated by experts from several universities in Emilia-Romagna and Catalonia.</p> <p>The <em>version 0.5</em> of this resource was developed by the SIRIS Academic in 2019 in collaboration with ART-ER, Emilia-Romagna (Quinquillá et <em>al.</em>, 2020), the <em>version 1 </em>was the result of an update done in 2020 in collaboration with the Generalitat de Catalunya, and the current version (<em>version 2</em>) has resulted in 2021 from the collaboration with ART-ER and the integration of an additional set of keywords provided by the <em>Artificial Intelligence and Intelligence Systems (AIIS)</em> Laboratory of the CINI (<em>Consorzio interuniversitario nazionale per l’informatica </em>based in Rome, Italy).</p> <p>The methodology for the construction of the controlled vocabulary is presented in the following steps:</p> <ol> <li> <p>An initial set of scientific publications was collected by retrieving the following records as a weakly-supervised (in the sense that records are linked to AI by their taxonomy and not by a manual label) dataset in the domain of Artificial Intelligence :</p> <ol> <li> <p>Publications from Scopus with the keyword “Artificial Intelligence”</p> </li> <li> <p>Publications from arXiv in the category “Artificial Intelligence”</p> </li> <li> <p>Publications in relevant journals in the scientific domain of “Artificial Intelligence”</p> </li> </ol> </li> <li> <p>An automated algorithm was used to retrieve, from the APIs of DBpedia, a series of terms that have some categorical relationships (i.e. those that are indexed as “sub-categories of”, “equivalent to”, among other relations in DBpedia) with the Artificial Intelligence concept and with the AI categories in the ACM taxonomy. The DBpedia tree has been exploited down to the level 3, and the relevant categories have been manually selected (for instance: <em>Classification algorithms</em>,<em> Machine learning</em> or <em>Evolutionary computation</em>) and others were ignored (for instance: <em>Artificial intelligence in fiction</em>, <em>Robots</em> or <em>History of artificial intelligence</em>) because they were not relevant, or not specifically in the domain.</p> </li> <li> <p>The keywords in publications in the dataset were extracted from the keyword sections and from the abstracts. The keywords with a higher <em>TF-IDF</em>, using an <em>IDF</em> matrix in the open domain, have been selected. The co-occurrence of keywords with categories in specific AI subdomain and a clusterization of the main keywords has been used for a categorization of the keywords at the thematic level.</p> </li> <li> <p>This list of keywords tagged by thematic category has been manually revised, removing the non-pertinent keywords and changing the wrong categorizations by fields.</p> </li> <li> <p>The weak-supervised dataset in the domain of Artificial Intelligence is used to train a Word2Vec (Mikolov <em>et al.</em>, 2013) word embedding model (a machine learning model based on neural networks).</p> </li> <li> <p>The terms’ list is then enriched by means of automatic methods, which are run in parallel: </p> <ol> <li> <p>The trained Word2Vec model is used to select, among the indexed keywords of the reference corpus, all terms “semantically close” to the initial set of words. This step is carried out to select terms that might not appear in the texts themselves, but that were deemed pertinent to label the textual records.</p> </li> <li> <p>Further, terms that are mentioned in the texts of the reference corpus and that are valued by the trained Word2Vec model as “semantically close” to the initial set of words are also retained. This step is performed to include in the controlled vocabulary a series of terms that are related to the focus of the SDGs and which are used by practitioners.</p> </li> </ol> </li> <li> <p>The final list produced by steps 2-6 is manually revised.</p> </li> </ol> <p> </p> <p>The definition of the vocabulary does not, per se, allow to identify STI contributions to AI: this activity in fact boils down to actually matching the terms in the controlled vocabulary to the content of the gathered STI textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance. Some relatively ambiguous keywords (which may match unwanted pieces of text), have a set of associated “extra” terms. These “extra” terms are defined as further terms that must co-appear, in the same sentence, together with their associated ambiguous keywords.</p> <p>Finally, each keyword in the vocabulary was assigned one or more AI subdomains, so that the vocabulary can also be used to tag collections of texts within narrower AI sub-domains. In order to complement the alignment between keywords and subdomains, a set of subdomain-specific keywords have been defined to better capture the scope of the subdomains. These allow better characterization of subdomains that are more difficult to define only by means of unambiguous specific concepts, or that overlap with the wide “machine learning” subdomain (example: machine learning applied to object recognition or text translation). The alignment between keywords and subdomains, and these keyword lists of each subdomain, have been applied to capture AI subdomains in research outputs. Through this classification process, we have identified projects and publications related to AI, with a focus on mapping the research competencies in the AI domain in Emilia-Romagna. The resulting research records have been reviewed by experts in the domain, given the occurrence of some false positives, which have been used to improve the approach.</p> <p>The final controlled vocabulary has been evaluated with an external test set, proposed by (Dunham <em>et al.,</em> 2020). The test set consists of the abstract of 10,606 papers published in the arXiv repository, of which 1,076 within the Artificial Intelligence subcategories and 9,530 in arXiv categories other than Artificial Intelligence. Evaluating the controlled vocabulary on this data set, we observe accuracy of .94. However, because the pertinence of these publications to the field of AI is based solely on their taxonomic classification (i.e., on whether they are classified in the arXiv within Artificial Intelligence and not on a manual labelling), this evaluation can only yield an orientative performance assessment.</p> <p>The version 2 includes new keywords extracted from the (1) re-training of the enrichment pipeline (steps 5-6 in the methodology) considering as initial set of terms the version 1 of the vocabulary on a reference corpus of new publications, and (2) from the flat keywords list provided by the <em>Artificial Intelligence and Intelligence Systems</em> <em>(AIIS)</em> Lab of CINI (Consorzio interuniversitario nazionale per l’informatica). The keywords in (2) have been cleaned by calculating precision and f-measure on the dataset (Dunham et al., 2020), selecting those keywords with the highest scores, and being manually validated a posteriori.</p> <p>The AI controlled vocabulary has been applied in two practical cases, which have the purpose of identifying skills, stakeholders and capabilities, of a specific research ecosystem at the regional level. See the following references:</p> <ul> <li> <p>Quinquillá, Arnau, Duran-Silva, Nicolau, Massucci, Francesco Alessandro, Fuster, Enric, Rondelli, Bernardo, Bologni, Leda, … Moretti, Giorgio. (2020). Text mining to identify skills, stakeholders and capabilities: the case of Artificial Intelligence in Emilia-Romagna. Zenodo. <a href="http://doi.org/10.5281/zenodo.3606342">http://doi.org/10.5281/zenodo.3606342</a>. Poster presented at: World Open Innovation Conference 2019 (WOIC); 11th december 2019, Rome, Italy.</p> </li> <li> <p>Bigas, E., Duran, N., Fuster, E., Parra, C., Fernández, T. (2021): “Anàlisi de l’especialització en intel·ligència artificial”. Col·lecció Monitoratge de la RIS3CAT, Generalitat de Catalunya <a href="http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf">http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf</a></p> </li> </ul> <p> </p> <p><strong>Acknowledgements</strong></p> <ul> <li> <p>Tatiana Fernández (Direcció General de Promoció Econòmica, Competència i Regulació, de la Generalitat de Catalunya), </p> </li> <li> <p>Daniel Marco, Daniel Santanach and Eduard Balbuena (Departament de Polítiques Digitals i Administració Pública, de la Generalitat de Catalunya) </p> </li> <li> <p>Albert Sabater (Observatori d’Ètica en Intel·ligència Artificial i Universitat de Girona)</p> </li> <li> <p>Leda Bologni, Lucia Mazzoni and Giorgio Moretti (Art-ER)</p> </li> <li> <p>Prof. RIta Cucchiara and Dr. Lorenzo Baraldi (Università degli Studi di Modena e Reggio Emilia)</p> </li> <li> <p>Artificial Intelligence and Intelligence Systems (AIIS) Lab of CINI (Consorzio interuniversitario nazionale per l’informatica)</p> </li> </ul> <p> </p> <p><strong>Bibliography</strong></p> <p>Bigas, E., Duran, N., Fuster, E., Parra, C., Fernández, T. (2021): “Anàlisi de l’especialització en intel·ligència artificial”. Col·lecció Monitoratge de la RIS3CAT, Generalitat de Catalunya <a href="http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf">http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf</a></p> <p>Dunham, J.W., Melot, J., & Murdick, D. (2020). Identifying the Development and Application of Artificial Intelligence in Scientific Text. ArXiv, abs/2002.07143. Available at: <a href="https://arxiv.org/abs/2002.07143">https://arxiv.org/abs/2002.07143</a></p> <p>Mikolov, Tomas & Corrado, G.s & Chen, Kai & Dean, Jeffrey. (2013). Efficient Estimation of Word Representations in Vector Space. 1-12.</p> <p>Quinquillá, Arnau, Duran-Silva, Nicolau, Massucci, Francesco Alessandro, Fuster, Enric, Rondelli, Bernardo, Bologni, Leda, … Moretti, Giorgio. (2020). Text mining to identify skills, stakeholders and capabilities: the case of Artificial Intelligence in Emilia-Romagna. Zenodo. <a href="http://doi.org/10.5281/zenodo.3606342">http://doi.org/10.5281/zenodo.3606342</a>. Poster presented at: World Open Innovation Conference 2019 (WOIC); 11th december 2019, Rome, Italy.</p>
Figure 5 in Masner, a new genus of Ceraphronidae (Hymenoptera, Ceraphronoidea) described using controlled vocabularies
Figure 5. Male antenna of Ceraphronoidea A–B Masner lubomirus Deans and Mikó, sp. n. C Aphanogmus sp. D Ceraphron sp. E Dendrocerus sp. F Megaspilus sp. Unlabelled arrows indicate antennal sensilla located on sensillar patch. Scale bars in micrometer.
Figure 7. Ceraphronoidea A–B Metasoma, ventral view A Trichosteresis glabra B in Masner, a new genus of Ceraphronidae (Hymenoptera, Ceraphronoidea) described using controlled vocabularies
Figure 7. Ceraphronoidea A–B Metasoma, ventral view A Trichosteresis glabra B Lagynodes sp. C Masner lubomirus Deans and Mikó, sp. n., lateral view. Scale bars in micrometer.
Figure 6. Ceraphronoidea A in Masner, a new genus of Ceraphronidae (Hymenoptera, Ceraphronoidea) described using controlled vocabularies
Figure 6. Ceraphronoidea A Trichosteresis glabra, head, anterior view B Aphanogmus sp., mesosoma, lateral view, anterior to the left C Aphanogmus sp., head, posterior view D Aphanogmus sp., head, anterior view E Megaspilus sp., male genitalia, ventral view F Conostigmus sp., apex of mesotibia, ventral view. Scale bars in micrometer.
Figure 4 in Masner, a new genus of Ceraphronidae (Hymenoptera, Ceraphronoidea) described using controlled vocabularies
Figure 4. Male genitalia and Waterston's evaporatorium of Ceraphronidae A–B Masner lubomirus Deans and Mikó, sp. n., male genitalia A ventral view8 B lateral view9 C–F Waterston's evaporatorium C Masner lubomirus Deans and Mikó, sp. n., dorsal view D Aphanogmus sp. dorsal view E Masner lubomirus Deans and Mikó, sp. n. F Aphanogmus sp., anterior view. Scale bars in micrometer.
A controlled vocabulary for research and innovation in the field of Circular Bioeconomy
<p>We live in a world of limited resources. Facing global challenges such as climate change and degradation of natural capital, in addition to the increasing rate of resource consumption, we are compelled to look for new ways of producing and consuming that respect the ecological limits of our planet. A model based on the Circular Bioeconomy (<em>CBE</em>) is key to successfully tackling the complexity of this paradigm shift and addressing these challenges: CBE is a relatively new and fast evolving concept, which is still in the conceptualisation phase. It stems from the concepts of “<em>bioeconomy</em>” and “<em>circular economy</em>”, which have become progressively interlinked in recent years.</p> <p>Although the wider community of stakeholders has not reached yet a consensus on a single definition for CBE, we try to go beyond this limitation, by proposing a controlled vocabulary of keywords related to the field, built by eliciting domain knowledge from both experts and policy-makers. This effort responds to an explorative project launched by the Generalitat de Catalunya (the regional government of Catalonia, Spain - specifically, the Ministry of the Economy and Finance, Secretary for Economic Affairs and European Funds) in order to identify CBE research, development and innovation activities. The work, coordinated by SIRIS Academic, was carried out by consulting the advice of experts in the domain (see Acknowledgements, below), and it was ultimately applied to inform regional strategies on bioeconomy and the Research and Innovation Strategy for Smart Specialisation.</p> <p>After a qualitative analysis of numerous examples, the keywords extracted have been classified into three categories:</p> <ul> <li> <p><strong>Unequivocal</strong>: keywords positively associated with the CBE domain (e.g. <em>bioplastic conversion, biorefinery, or biomass</em>),</p> </li> <li> <p><strong>Bioeconomy</strong>: keywords in the bioeconomy domain but not necessarily within the circular paradigm (e.g. <em>agriculture, algae, or organic waste</em>),</p> </li> <li> <p><strong>Technologies and processes</strong>: keywords concerning circular processes and technologies (e.g. <em>valorisation, or bioconversion</em>).</p> </li> </ul> <p>This distinction allows one to apply the vocabulary to link a given text to the CBE domain. Indeed, a text may be considered in the CBE perimeter if it mentions an unequivocal keyword, or if it contains a concept concerning bioeconomy co-appearing with a concept concerning circular technologies and processes (e.g. food waste with composting, or vegetation with biosynthesis).</p> <p>The circular bioeconomy vocabulary, in its current form, contains 393 keywords (49 unequivocal keywords, 229 bioeconomy keywords, and 115 concerning circular processes and technologies). <br> </p> <p><strong>Datasets Provided</strong></p> <p><br> In this publication, we provide:</p> <p> </p> <ul> <li><strong>A document outlining the context that lead to the creation of the vocabulary</strong>, the methodology used to build it and the strategy to apply it</li> <li><strong>The full vocabulary of terms as a csv file</strong>. For each keyword of the vocabulary, the property “type” specifies whether the term is either an “Unequivocal” CBE keyword, a wider “Bioeconomy” term or a specification of “Technologies and processes”.</li> <li>In addition to the controlled vocabulary, we also provide <strong>a labelled dataset to foster the development of new text mining initiatives</strong> on the results obtained. The dataset consists of the description of 2000 R&D projects funded by the Horizon 2020 framework of the European Commission. Of these, 1000 have been associated with the field of CBE via the vocabulary and further manual validation, while further 1000 descriptions are unrelated to CBE. Each project entry has the following attributes: title, abstract, ecRef (project identifier in the EU information service, CORDIS), and label (where True are projects in the CBE domain)</li> </ul> <p><strong>Acknowledgements</strong></p> <p>The work described above, leading to the production of the datasets provided, was developed in collaboration with Generalitat de Catalunya officials, as well as CBE research and innovation experts and stakeholders that contributed essential input to the conceptual definition of the CBE, the review and improvement of the controlled vocabulary and the quality of the classification, as well as the the revision of the analytical results described in the section Applicability, in particular:</p> <ul> <li> <p>Tatiana Fernández (Direcció General de Promoció Econòmica, Competència i Regulació, de la Generalitat de Catalunya)</p> </li> <li> <p>Mercè Balcells i Ramón Canela (Universitat de Lleida)</p> </li> <li> <p>Teresa Botargues (Diputació de Lleida)</p> </li> <li> <p>Sergio Ponsà (Centre Tecnològic BETA)</p> </li> <li> <p>Ignacio Rodríguez and Jaume Sió (Departament d’Agricultura, Ramaderia, Pesca i Alimentació, de la Generalitat de Catalunya)</p> </li> </ul>
Controlled vocabularies and knowledge organisation for Digital Humanities - Part I
<p>The workshop "Controlled vocabularies and knowledge organisation for the digital humanities" aimed to exchange experiences and applications of controlled vocabularies for research and projects in the field of digital humanities, including research infrastructures, libraries and other cultural institutions linked to social sciences, arts and humanities (SSAH).</p> <p>This is part I, presented by: Helen Goulis (DARIAH Thesaurus Maintenance WG); José Moreiro González (University Carlos III of Madrid); Bruno Almeida (ROSSIO Infrastructure / NOVA CLUNL); Filipa Medeiros (Art Library and Archives of the Calouste Gulbenkian Foundation).</p> <p>This online workshop was organized by ROSSIO Infrastructure, Department of Linguistics of NOVA FCSH, NOVA CLUNL and Art Library and Archives of the Calouste Gulbenkian Foundation (FCG), and took place on the 12th of July 2021.</p>
Controlled vocabularies and knowledge organisation for Digital Humanities - Part II
<p>The workshop "Controlled vocabularies and knowledge organisation for the digital humanities" aimed to exchange experiences and applications of controlled vocabularies for research and projects in the field of digital humanities, including research infrastructures, libraries and other cultural institutions linked to social sciences, arts and humanities (SSAH).</p> <p>This is part II, presented by: Sébastien Durost (BIBRACTE EPCC), Guillaume Reich (MSHE Ledoux) & Jean Pierre Girard (UMR Archéorient); Susana Medina (FEUP); Ana Paula Figueiredo (Directorate-General Cultural Heritage of Portugal); Teresa Borges (Portuguese Film Archives).</p> <p>This online workshop was organized by ROSSIO Infrastructure, Department of Linguistics of NOVA FCSH, NOVA CLUNL and Art Library and Archives of the Calouste Gulbenkian Foundation (FCG), and took place on the 12th of July 2021.</p>
NCAS General Data Standard Controlled Vocabularies
Controlled Vocabularies for the NCAS General Data Standard
A controlled vocabulary for research and innovation in the field of Cultural Heritage & Heritage Sciences
<p>This controlled vocabulary of keywords related to the field of Cultural Heritage and Heritage Sciences was built by SIRIS Academic in collaboration with IRPET (the Regional Institute for Economic Planning of Tuscany) and the ISPC (Institute of Heritage Science of CNR), in order to identify Cultural related research, development, and innovation activities. The work was carried out by consulting domain experts' advice, and it was ultimately applied to inform regional strategies on Cultural Heritage and research and innovation policy.</p> <p>The aim of this vocabulary is to enable one to retrieve texts (e.g. R&D projects and scientific publications) featuring the concepts included in the present vocabulary in their titles and abstracts, assuming that these records have a certain contribution of applications, techniques and issues, in the domain of Cultural Heritage and Heritage Sciences.</p> <p>The aim of this classification is to identify research products in the domain of Cultural Heritage, ranging from documents in some of its “traditional” disciplines, but also from documents emerging from interdisciplinary projects that apply novel areas and technologies in the domain of Cultural Heritage. The identification of texts in the domain of Cultural Heritage requires a task of text classification. Developing a method that could be applied to decide if a text can be relevant or have some relation to the domain of Cultural Heritage is a challenging task. The definition of what Cultural Heritage is and what it includes is a complex activity, even for domain experts. This is in particular because Cultural Heritage is quite a broad field of knowledge, and there is no full agreement on where the borders of the domain are. To define the scope of the perimeter, in this project, many of the available definitions were taken into account.</p> <p>Because of the high number of resources available in the domain, among thesauruses and taxonomies, the construction of a weakly-supervised controlled vocabulary was considered as the best way of retrieving documents in the domain. Since there is no annotated corpus/dataset of research texts in the domain capable of generalising the diversity of publications that can be related to the cultural domain, but stemming from different disciplines, we have opted for a text classification technique based on rules – specifically, a weakly-supervised controlled vocabulary.</p> <p>As defined by the Getty Institute, a controlled vocabulary is an organized arrangement of words and phrases used to index content and/or to retrieve content through browsing or searching. It typically includes preferred and variant terms and has a defined scope or describes a specific domain. The purpose of controlled vocabularies is to organize information and to provide terminology to catalogue and retrieve information. While capturing the richness of variant terms, controlled vocabularies also promote consistency in preferred terms and the assignment of the same terms to similar content (Harping, 2010).</p> <p>In short, Cultural Heritage is a rather abstractly-defined field, and Heritage Science is a particularly “fuzzy” field within Cultural Heritage. One of the main limitations of the approach we used is that the controlled vocabularies never capture all the lexical and linguistic variants of a term, and we may miss relevant texts if we cannot find the correct pattern to match during the search. But on the other hand, the controlled vocabulary is built from available vocabularies and thesauruses in the domain of Cultural Heritage, which are large resources. All the concepts in these resources are not included directly in the controlled vocabularies, because they would add noise to the classification. Therefore, the automatic weak supervision and a human curation of the final controlled vocabulary is fundamental for achieving correct results.</p> <p>The controlled vocabulary is built taking advantage of these four resources:</p> <ul> <li>The <a href="https://www.getty.edu/research/tools/vocabularies/aat/"><strong>Art and Architecture Thesaurus (AAT)</strong></a>: this is a structured vocabulary with approximately 34,000 concepts, including 131,000 words, descriptions and other information related to art, architecture, decorative arts, archival material and material culture, commonly used for cataloguing and for information retrieval.</li> <li> <p>Some cultural heritage categories in <strong><a href="https://en.wikipedia.org/wiki/Category:Cultural_heritage">Wikipedia</a> </strong>and <strong><a href="https://dbpedia.org/page/Cultural_heritage">DBpedia</a></strong>: these categories have been used to collect all related articles and subcategories, in order to obtain relevant, similar and specific instances of concepts linked to the domain. </p> </li> <li> <p>The <strong><a href="https://www.riches-project.eu/riches-taxonomy.html">RICHES Taxonomy</a></strong>: this taxonomy is a theoretical framework of related terms and their definitions, referring to the new concepts in the digital era, with the aim of defining the scope of some digital technologies applied to cultural heritage.</p> </li> <li> <p><strong><a href="https://www.heritagedata.org/blog/">Heritage Data - Linked Data Vocabularies for Cultural Heritage</a></strong>: a dataset which includes several cultural heritage thesauruses and vocabularies and is recognised as a reference point in the United Kingdom in the domain of cultural heritage.</p> </li> </ul> <p>The collection of concepts extracted from these four resources was composed of more than 60,000 terms, which have been refined as described in the next section.</p> <p> </p> <p><strong>## Automatic validation of the controlled vocabulary</strong></p> <p>In order to refine the collection of concepts to have a final set of relevant concepts and terms in the domain of Cultural Heritage, a semi-automatic validation has been applied to remove the irrelevant, too general, and ambiguous terms.</p> <p>To keep the relevant ones, the <a href="https://ncses.nsf.gov/pubs/nsb20206/specialization-and-impact-analysis-combined#:~:text=The%20specialization%20index%20(SI)%20is,the%20total%20output%20across%20all">specialization index (SI) </a>metric has been calculated for each of the keywords in the collection. In this case, the SI can be obtained measuring the fraction of publications with a keyword in a set of publications in the domain of Cultural Heritage and normalizing over the fraction of publications in the open domain with that keyword.</p> <p>After the calculation of the SI, all the keywords below a certain threshold are removed, and a manual supervision step is applied in order to remove non-pertinent keywords. An example of this automatic validation can be observed in the next table:</p> <table> <tbody> <tr> <td> <p><strong>Keyword</strong></p> </td> <td> <p><strong>Specialization Index</strong></p> </td> <td> <p><strong>Automatic threshold</strong></p> </td> <td> <p><strong>Manual supervision</strong></p> </td> </tr> <tr> <td> <p>male</p> </td> <td> <p>0.27</p> </td> <td> <p>Removed</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>3-d laser scanning</p> </td> <td> <p>0.7</p> </td> <td> <p>Removed</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>78 rpm records</p> </td> <td> <p>20.7</p> </td> <td> <p>Accepted</p> </td> <td> <p>Removed</p> </td> </tr> <tr> <td> <p>vienna</p> </td> <td> <p>3.48</p> </td> <td> <p>Accepted</p> </td> <td> <p>Removed</p> </td> </tr> <tr> <td> <p>radiocarbon dating</p> </td> <td> <p>13.6</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>graffiti</p> </td> <td> <p>25</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>bark painting</p> </td> <td> <p>20.7</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>pompeii</p> </td> <td> <p>16.23</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> </tbody> </table> <p>The SI of the final keywords can be used as a probabilistic metric for each keyword.</p> <p>The final list of keywords was manually curated by domain experts.</p> <p> </p> <p><strong>## Evaluation of the controlled vocabulary</strong></p> <p>The final controlled vocabulary was evaluated with an external dataset with the aim of calculating its degree of precision. The evaluation dataset was composed of a collection of articles in 4 journals unequivocally considered to fall within the domain of Cultural Heritage. These four journals were: <em>(1) Journal Of Cultural Heritage, (2) Journal On Computing And Cultural Heritage, (3) Journal Of Cultural Heritage Management And Sustainable Development and (4) Digital Applications In Archaeology And Cultural Heritage.</em> This collection was composed of 5,000 articles, considered as the positive set, and another collection of randomly selected 5,000 articles outside of the Cultural Heritage domain, considered as the false set.</p> <p>The Cultural Heritage vocabulary was applied to the evaluation data set, obtaining a 95% of precision. After a set of improvements on the vocabulary, based on the exploration of publications not identified in the first test and the false positive results, we obtained a 98% of precision. The application of the vocabulary taking advantage of the probability of each keyword as its weight of being in the domain did not improve the results, and for this reason the probabilistic approach was discarded.</p> <p> </p> <p>## <strong>Using the vocabulary to classify publications concerning Cultural Heritage</strong></p> <p>The definition of the vocabulary does not, per se, allow to identify research contributions in Cultural Heritage: this is performed by actually matching the terms in the controlled vocabulary to the content of the gathered research textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance.</p> <p>In the following table we present some examples of the tagging process on some abstracts:</p> <table> <tbody> <tr> <td> <p><strong>Publication title</strong></p> </td> <td> <p><strong>Publication abstract</strong></p> </td> </tr> <tr> <td> <p>Egocentric visitor localization and artwork detection in cultural sites using synthetic data</p> </td> <td> <p>Computer vision and machine learning can be used in <strong>cultural heritage to augment the experience of visitors during the exploration of the cultural site</strong>, as well as to assist its management. To achieve such goals, two fundamental tasks should be addressed, i.e., localizing <strong>visitors and recognizing the observed artworks</strong>. Wearable cameras offer a convenient setting to address both tasks through the analysis of images acquired from the visitors’ points of view. However, the engineering of approaches to address such tasks generally requires large amounts of labeled data. We propose a tool which can be used to collect and automatically label synthetic visual data suitable to study image-based localization and artwork detection. The tool simulates a virtual agent navigating the <strong>3D model of a real cultural site</strong> and automatically captures video frames along with the related ground truth camera poses and semantic masks indicating the position of artworks. We generate a dataset of synthetic images starting from the 3D model of a <strong>museum located in Siracusa</strong>, Italy. The experiments suggest that the proposed tool allows to drastically reduce the effort needed to collect and label data, providing a means to generate large-scale datasets suitable to study localization and <strong>artwork detection in cultural sites</strong>.</p> </td> </tr> <tr> <td> <p>Discovering Leonardo with artificial intelligence and holograms: A user study</p> </td> <td> <p>Cutting-edge visualization and interaction technologies are increasingly used in<strong> museum exhibitions</strong>, providing novel ways to engage visitors and enhance their <strong>cultural experience</strong>. Existing applications are commonly built upon a single technology, focusing on visualization, motion or verbal interaction (e.g., high-resolution projections, gesture interfaces, chatbots). This aspect limits their potential, since museums are highly heterogeneous in terms of visitors profiles and interests, requiring multi-channel, customizable interaction modalities. To this aim, this work describes and evaluates an artificial intelligence powered, interactive holographic stand aimed at describing <strong>Leonardo Da Vinci's art</strong>. This system provides the users with accurate<strong> 3D representations of Leonardo's machines</strong>, which can be interactively manipulated through a touchless user interface. It is also able to dialog with the users in natural language about Leonardo's art, while keeping the context of conversation and interactions. Furthermore, the results of a large user study, carried out during art and tech exhibitions, are presented and discussed. The goal was to assess how users of different ages and interests perceive, understand and explore <strong>cultural objects </strong>when holograms and artificial intelligence are used as instruments of knowledge and analysis.</p> </td> </tr> <tr> <td> <p>Hybrid query expansion using lexical resources and word embeddings for sentence retrieval in question answering</p> </td> <td> <p>Question Answering (QA) systems based on Information Retrieval return precise answers to natural language questions, extracting relevant sentences from document collections. However, questions and sentences cannot be aligned terminologically, generating errors in the sentence retrieval. In order to augment the effectiveness in retrieving relevant sentences from documents, this paper proposes a hybrid Query Expansion (QE) approach, based on lexical resources and word embeddings, for QA systems. In detail, synonyms and hypernyms of relevant terms occurring in the question are first extracted from MultiWordNet and, then, contextualized to the document collection used in the QA system. Finally, the resulting set is ranked and filtered on the basis of wording and sense of the question, by employing a semantic similarity metric built on the top of a Word2Vec model. This latter is locally trained on an extended corpus pertaining the same topic of the documents used in the QA system. This QE approach is implemented into an existing QA system and experimentally evaluated, with respect to different possible configurations and selected baselines, for the <strong>Italian language and in the Cultural Heritage domain</strong>, assessing its effectiveness in retrieving sentences containing proper answers to questions belonging to four different categories.</p> </td> </tr> <tr> <td> <p>"3D reconstruction and validation of historical background for immersive VR applications and games: The case study of the Forum of Augustus in Rome"</p> </td> <td> <p>"In the last decades, thanks to the success of the video games industry, the sector of technologies applied to cultural heritage has begun to envisage, in this domain, new possibilities for the <strong>dissemination of heritage and the study of the past </strong>through edutainment models. More recently, experimentation in the field of<strong> virtual archaeology </strong>has led to the development of virtual museums and interactive applications. Among these, the “serious game” segment – the<strong> application of interactive technologies to the cultural heritage domain</strong> – is rapidly growing, also including immersive VR technologies. Applied VR games and applications are characterized by a thorough <strong>historical background and a validated 3D reconstruction</strong>. Indeed, producing such products requires a tailored workflow and large effort in terms of time and professionals involved to guarantee such faithfulness. Drawing on our previous work in the<strong> field of virtual archaeology</strong> and referring to recent experiences related to the deployment of applied VR games on PlayStation VR, we describe and assess a workflow for the production of <strong>historically accurate 3D assets</strong>, targeting interactive, immersive VR products. The workflow is supported by the case study of the <strong>Forum of Augustus </strong>and different output applications, highlighting peculiarities and issues emerging from a multi and interdisciplinary approach.</p> </td> </tr> </tbody> </table> <p>Through this classification process, we identified projects and publications related to heritage, with different levels of relationship and relevance, but mostly relevant to understanding the research competencies in the domain. The resulting research records were reviewed by experts in the domain, given the occurrence of some false positives.</p> <p>Among the main strengths of this step, it’s worth mentioning the fact that the vocabulary is broad and not restricted to the field of Heritage Science (that is, to STEM applications in Cultural Heritage), as it takes advantage of a variety of available resources. Moreover, by looking directly at the textual data, instead of using the assigned bibliometric areas, we can better capture interdisciplinary research. The limitations of this approach were presented at the beginning of this document: for example, relevant texts could be missed if the correct pattern to match during the search is not found.</p> <p> </p> <p><strong>## Vocabulary of concepts related to Key Enabling Technologies in the domain of Cultural Heritage and Culture</strong></p> <p>For the development of this vocabulary, the definition of key enabling technologies in the domain of Cultural Heritage and Culture, was based on reference of the <a href="http://www.irpet.it/archives/53165">report 'Technologies, Cultural Heritage and Culture' published on March 2019</a> by IRPET.</p> <p>A vocabulary for each Key Enabling Technology (hereafter, KET) was prepared by extracting the relevant concepts, words, technologies and examples from the Platform Report document 'Technologies, Cultural Heritage and Culture, within APPENDIX A. DESCRIPTION OF MAIN TECHNOLOGIES FOR ROADMAP (p. 45-61). Each vocabulary contains a set of terms divided into subdomains.</p> <p>The KETs have been divided into the following six groups:</p> <ul> <li> <p>ICT</p> </li> <li> <p>PHOTONICS, MICRO- AND NANO-ELECTRONICS</p> </li> <li> <p>PLATFORMS</p> </li> <li> <p>NANO AND BIOTECHNOLOGY, ADVANCED MATERIALS</p> </li> <li> <p>PARTICLE ANALYTICAL SYSTEMS</p> </li> </ul> <p>The initial keywords extracted from the document were enriched following the approach based on semantic keyword enrichment based on combination of concurrent keywords and word embeddings (Duran-Silva et al., 2019; Duran-Silva et al., 2021).</p> <p>This second vocabulary has to be used in combination with the Cultural Heritage vocabulary to capture KETs within the domain of cultural heritage.</p> <p> </p> <p>## <strong>Use of the controlled vocabulary</strong></p> <p>The definition of the vocabulary does not, per se, allow identifying STI contributions to the domain: this activity in fact boils down to actually matching the terms in the controlled vocabulary to the content of the gathered STI textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance. Some relatively ambiguous keywords (which may match unwanted pieces of text), have a set of associated “extra” terms. These “extra” terms are defined as further terms that must co-appear, in the same sentence, together with their associated ambiguous keywords. SIRIS Academic has developed the <a href="https://github.com/sirisacademic/VocTagger">voc_tagger tool</a>, a multiprocess information extraction system able to identify hidden knowledge in textual documents using “controlled vocabularies”, openly available at GitHub and compatible with these controlled vocabularies.</p> <p> </p> <p><strong>## Bibliography</strong></p> <p>Harpring, P. (2010). Introduction to controlled vocabularies: terminology for art, architecture, and other cultural works. Getty Publications.</p> <p>Nicolau Duran-Silva, Enric Fuster, Francesco Alessandro Massucci, César Parra-Rojas, Arnau Quinquillà, Fernando Roda, Bernardo Rondelli, Nicandro Bovenzi, & Chiara Toietta. (2021). A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI) (Version 2) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5591987</p> <p>Duran-Silva, Nicolau, Fuster, Enric, Massucci, Francesco Alessandro, & Quinquillà, Arnau. (2019). A controlled vocabulary defining the semantic perimeter of Sustainable Development Goals (1.2) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3567769</p>
SatTerm experience: vocabulary control and facet analysis help improve the software requirement solicitation process
<p>One of the most difficult steps in the software development process is moving from requirements written in natural, uncontrolled language, to the formalisms required by the design modelling languages. To solve this issue, practitioners should pay attention to the possibility of applying vocabulary control and knowledge representation techniques to produce better specifications. The use of controlled vocabularies and the modelling of the conceptual relationships between concepts in a specific domain are expected to improve the quality of the specifications.<br> Vocabulary control and semantic modelling are promising tools to avoid the most frequent problems in the requirements specification process: lack of consistency and ambiguity. This paper provides a detailed description of the development process of an ontology used for requirements modelling in the area of satellite control systems. The process applied is based on well-established practices and guidelines applied for the construction of controlled vocabularies and faceted classifications schemas. Engineers can use the ontology when writing system specifications using predefined templates. The use of this ontology ensures the consistency of the specifications written by different engineers improves the communication with other parties involved in the system construction activities and sets the foundations for a semi-automated generation of models for subsequent design activities.</p>
A controlled vocabulary defining the semantic perimeter of Sustainable Development Goals
<p>A set of controlled terms that define the scope and breadth of <a href="https://sustainabledevelopment.un.org/">Sustainable Development Goals (SDGs) as defined by the United Nations</a>. These terms may be used to tag and index textual records in accordance with SDGs.</p> <p>The vocabulary is constructed by means of the following steps:</p> <ol> <li>An initial set of terms per SDG target is built by extracting key terms from the UN official list of Goals, Targets and Indicators</li> <li>The list is manually enriched by performing a review of the literature produced around SDGs and by compiling lists of pertinent words per Target mentioned by the reviewed documents</li> <li>A reference textual corpus is downloaded by searching for the initial set terms defined at step 1. and 2. The corpus is used to train a Word2Vec word embedding model (a machine learning model based on neural networks).</li> <li>The terms’ list is then enriched by means of automatic methods, which are run in parallel: <ul> <li>The trained Word2Vec model is used to select, among the indexed keywords of the reference corpus, all terms “semantically close” to the initial set of words. This step is carried out to select terms that might not appear in the texts themselves, but that were deemed pertinent to label the textual records.</li> <li>Further terms that are mentioned in the texts of the reference corpus and that are valued by the trained Word2Vec model as “semantically close” to the initial set of words are also retained. This step is performed to include in the controlled vocabulary a series of terms that are related to the focus of the SDGs and which are used by practitioners.</li> <li>An automated algorithm is used to retrieve, from the APIs of WikiPedia a series of terms that have some categorical relationships (i.e. those that are indexed as “a broader concept of”, or “equivalent to” in DBpedia) with the initial set of words.</li> </ul> </li> <li>The final list produced by steps 1-4 s finally manually revised</li> </ol>
SCENT for GLAM: a tool for giving meaning to professional controlled vocabularies
<p><span>This paper discusses SCENT for GLAM, which stands for Semantic and Collaborative Environment for a Network of Terminology for Galleries, Libraries, Archives and Museums. SCENT for GLAM offers a number of features that allow an institution to create or import terminology, to convert it into SKOS format, to edit it, make alignments with other terminologies from the same institution or another one and finally publish and share this terminology and alignments. The idea is to provide a repository of terminologies that will include all of the cultural sector concepts.</span></p>
Ciência Vitae controlled vocabulary - Current status
Open the record for dataset details and reuse information.
From one place to a different place: problems of vocabulary control of geographic names and concepts
<p>Although apparently easily understood and straightforward to manage, the representation of place as a concept in indexing descriptions can be subject to subtle nuances of meaning and interpretation. A current re-classification project at the British Museum Anthropology Library reveals that a place can be understood in different ways, and defined using a variety of criteria. This is in addition to the conventional problems of vocabulary control of geographic names caused by the occurrence of official, local, and vernacular names, and the form of name in different natural languages.<br>The use of well formulated design principles for the construction of faceted knowledge organization systems can provide a solution for both organization and representation of place concepts. When applied, such principles create classification data that is readily comprehended and managed by machines, and software which supports the automatic generation of knowledge organization tools.</p>
RiboSeqOrg Controlled Vocabularies For Metadata Cleaning - 02/10/2024
<p>This set of sheets describe the terms used to clean Ribo-Seq metadata after they have been fetched from SRA and aggregated. <br><br>The first sheet 'Content' describes the terms within cells (i.e. metadata content) where anything that matches All Names will become Main Name. <br><br>A similar principle of All Names -> Main Names is applied to columns which are handled across numerous sheets based on column type. <br><br>First columns are standardised with the various column sheets and then the values within those columns are standardised using the Content sheet. </p>
Ciência Vitae controlled vocabulary - Type of event
Open the record for dataset details and reuse information.
Ciência Vitae controlled vocabulary - Type of participation
Open the record for dataset details and reuse information.
Ciência Vitae controlled vocabulary - Productions
Open the record for dataset details and reuse information.
Ciência Vitae controlled vocabulary - Function Performed
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.