Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7
datasets available to search
ShareScore release 0.9.0
Dataset results
7 results for “Heritage Science”
Hyperspectral Imaging dataset for use in Heritage Science
<p>The following data sets were collected to support the potential uses of opensource data in the context of digital humanities and heritage sciences. </p> <p>This proposed experiment is conducted by the UCL Institute for Sustainable Heritage in collaboration with the Centre for Digital Humanities. Imaging methods including Photography, Multispectral Imaging, Hyperspectral Imaging and Xray Fluorescence Mapping have been collected along with the complete readout metadata of the instrumentation. </p> <p>We hope that you find the data helpful, and we welcome you to use the data in any way you wish, for all and any analysis development purposes. For us to build upon this research, we ask that in return you would be willing to share in some regard your experiences in using open-source data, using our data, successes and issues. </p> <p>If you would be willing to engage with us in this endeavor, please feel free to contact us so that we may be able to follow up with you. </p> <p>Other Data sets available <a href="https://zenodo.org/record/7319696#.Y3NuOXbP2Uk">Here</a></p> <p> </p> <p>E: <a href="mailto:molly.fort.21@ucl.ac.uk">molly.fort.21@ucl.ac.uk</a> </p> <p>Object Paradata; </p> <ul> <li><strong>Postcard – c. Early 1900's </strong></li> <li><strong>Language – Eng. </strong></li> <li><strong>Materials – colour print on card, metallic leafing. </strong></li> <li><strong>Front transcription - </strong></li> <li><strong> ‘Greetings’ </strong></li> <li><strong> ‘May your Birthday bring you Peace & perfect Happiness, Golden hopes & Love of Friends, And every Happiness this world can send.’ </strong></li> <li><strong>Object Dimensions – 138mm X 88mm </strong></li> </ul> <p>The postcard is an item of ephemera donated to the UCLDH Digitisation Suite by Prof Melissa Terras, for teaching and training purposes in 2015.</p> <p>This folder contains:</p> <p>Hyperspectral Image data collected using a <a href="https://www.clydehsi.com/hyperspectral-cameras">ClydeHSI VNIR-HR+ Hyperspectral Imaging System</a>.</p> <p>Images captured : 477x484 pixel, 304 spectral band images, 4*4 pixel binning</p> <ul> <li>*.hdr - Header file read out from the ClydeHSI systems instructions for reading the subsequent .raw spectral database. </li> <li>*.raw - Hyperspectral image data cube information. Combination with hdr file creates a ENVI file format, this can be read into a variety of image analysis software packages. </li> <li>postcardhsi.ini - Metadata collected and read out from ClydeHSI system.</li> <li>Dark/White.corr - Correction files taken from camera for processing and minimalising system noise and illumination variences.</li> <li>*_refl.* - Pre - Corrected hyperspectral image data, using provided ClydeHSI software.</li> <li>Truecolour RGB reference image </li> </ul> <p>Each raw and header file set makes-up a single data set in ENVI file format. </p> <p>ENVI reading support exists in Python, R, Matlab, and other common image analysis packages.</p>
Image Data sets for use in Heritage Science
<p>The following data sets were collected to support the potential uses of opensource data in the context of digital humanities and heritage sciences. </p> <p><a href="https://doi.org/10.5281/zenodo.7292917">Photographs</a> </p> <p><a href="https://doi.org/10.5281/zenodo.7292961">X-Ray Fluorescence</a> </p> <p><a href="https://doi.org/10.5281/zenodo.7322908">Hyperspectral Imaging</a> </p> <p><a href="https://doi.org/10.5281/zenodo.7292714">Multispectral Imaging</a></p> <p> </p> <p>This proposed experiment is conducted by the UCL Institute for Sustainable Heritage in collaboration with the Centre for Digital Humanities. Imaging methods including Photography, Multispectral Imaging, Hyperspectral Imaging and Xray Fluorescence Mapping have been collected along with the complete readout metadata of the instrumentation.</p> <p>We have collected this as an example of typical, unprocessed imaging datasets that would be found in standard image conditions. This data is not optimized, nor do we claim it to be perfect quality, our aim is to provide users with access to a range of imaging data sets. We have included the data with minimum processing, as it is read straight from our systems, with the accompanying metadata provided from capture alone.</p> <p>We hope that you find the data helpful, and we welcome you to use the data in any way you wish, for all and any analysis development purposes. For us to build upon this research, we ask that in return you would be willing to share in some regard your experiences in using open-source data, using our data, successes and issues. </p> <p>If you would be willing to engage with us in this endeavor, please feel free to contact us so that we may be able to follow up with you. </p> <p>E: <a href="mailto:molly.fort.21@ucl.ac.uk">molly.fort.21@ucl.ac.uk</a> </p> <p>Object Paradata; </p> <ul> <li><strong>Postcard – c. Early 1900's </strong></li> <li><strong>Language – Eng. </strong></li> <li><strong>Materials – colour print on card, metallic leafing. </strong></li> <li><strong>Front transcription - </strong></li> <li><strong> ‘Greetings’ </strong></li> <li><strong> ‘May your Birthday bring you Peace & perfect Happiness, Golden hopes & Love of Friends, And every Happiness this world can send.’ </strong></li> <li><strong>Object Dimensions – 138mm X 88mm </strong></li> </ul> <p>The postcard is an item of ephemera donated to the UCLDH Digitisation Suite by Prof Melissa Terras, for teaching and training purposes in 2015.</p>
Multispectral Spectral Imaging dataset for use in Heritage Science
<p>The following data sets were collected to support the potential uses of opensource data in the context of digital humanities and heritage sciences. </p> <p>This proposed experiment is conducted by the UCL Institute for Sustainable Heritage in collaboration with the Centre for Digital Humanities. Imaging methods including Photography, Multispectral Imaging, Hyperspectral Imaging and Xray Fluorescence Mapping have been collected along with the complete readout metadata of the instrumentation. </p> <p>We hope that you find the data helpful, and we welcome you to use the data in any way you wish, for all and any analysis development purposes. For us to build upon this research, we ask that in return you would be willing to share in some regard your experiences in using open-source data, using our data, successes and issues. </p> <p>If you would be willing to engage with us in this endeavor, please feel free to contact us so that we may be able to follow up with you. </p> <p>Other Data sets available <a href="https://zenodo.org/record/7319696#.Y3NuOXbP2Uk">Here</a></p> <p>E: <a href="mailto:molly.fort.21@ucl.ac.uk">molly.fort.21@ucl.ac.uk</a> </p> <p>Object Paradata; </p> <ul> <li><strong>Postcard – c. Early 1900's </strong></li> <li><strong>Language – Eng. </strong></li> <li><strong>Materials – colour print on card, metallic leafing. </strong></li> <li><strong>Front transcription - </strong></li> <li><strong> ‘Greetings’ </strong></li> <li><strong> ‘May your Birthday bring you Peace & perfect Happiness, Golden hopes & Love of Friends, And every Happiness this world can send.’ </strong></li> <li><strong>Object Dimensions – 138mm X 88mm </strong></li> </ul> <p>The postcard is an item of ephemera donated to the UCLDH Digitisation Suite by Prof Melissa Terras, for teaching and training purposes in 2015.</p> <p>This folder contains:</p> <ul> <li>Images captured using a <a href="https://photography.phaseone.com/xf-camera-system/">PhaseOne XF Multispectral Camera System.</a> <ul> <li>Image filenames are arranged as postacards_postcardmsi-<strong>Postcard</strong>- <strong>(Wavelength No.)(Filter)</strong>_*sequence order number*_R.tif where wavelength number is the nominal central illumination wavelength in nm (365, 385, 410, 420, 450, 480, 510, 550, 600, 630, 640, 660, 740, 850, 940), Filter is the colour of the long-pass filter (N - no filter, I - Infrared filter, G - Green filter, R - Red filter) and sequence order number is a count from 0001 denoting the order in which the image was acquired)</li> <li>Complementary flats for each of the object images, used typically to process even illumination distribution, captured of white, flat, smooth, non-chemically processed imaging standard flat paper with the same naming convention as above. </li> </ul> </li> <li>postcard_postcardmsi-Postcard.json - Metadata read out collected from MS camera system</li> <li>Truecolour RGB reference image</li> </ul>
Imaging dataset for use in Heritage Science
<p>The following data sets were collected to support the potential uses of opensource data in the context of digital humanities and heritage sciences. </p> <p>This proposed experiment is conducted by the UCL Institute for Sustainable Heritage in collaboration with the Centre for Digital Humanities. Imaging methods including Photography, Multispectral Imaging, Hyperspectral Imaging and Xray Fluorescence Mapping have been collected along with the complete readout metadata of the instrumentation. </p> <p>We hope that you find the data helpful, and we welcome you to use the data in any way you wish, for all and any analysis development purposes. For us to build upon this research, we ask that in return you would be willing to share in some regard your experiences in using open-source data, using our data, successes and issues. </p> <p>If you would be willing to engage with us in this endeavor, please feel free to contact us so that we may be able to follow up with you. </p> <p>Other Data sets available <a href="https://zenodo.org/record/7319696#.Y3NuOXbP2Uk">Here</a></p> <p>E: <a href="mailto:molly.fort.21@ucl.ac.uk">molly.fort.21@ucl.ac.uk</a> </p> <p>Object Paradata; </p> <ul> <li><strong>Postcard – c. Early 1900's </strong></li> <li><strong>Language – Eng. </strong></li> <li><strong>Materials – colour print on card, metallic leafing. </strong></li> <li><strong>Front transcription - </strong></li> <li><strong> ‘Greetings’ </strong></li> <li><strong> ‘May your Birthday bring you Peace & perfect Happiness, Golden hopes & Love of Friends, And every Happiness this world can send.’ </strong></li> <li><strong>Object Dimensions – 138mm X 88mm </strong></li> </ul> <p>The postcard is an item of ephemera donated to the UCLDH Digitisation Suite by Prof Melissa Terras, for teaching and training purposes in 2015.</p> <p>This folder contains:</p> <ul> <li>Photographs at various conditions</li> <li>Flats at matching conditions</li> <li>Metadata stored within image file</li> </ul> <p>Reading file names: </p> <p>(Flat or postcard)<strong>-exp</strong>(exposure value). format</p> <p>*.Jpg - compressed JPEG readout, saved straight from camera</p> <p>*.tiff - non-compressed TIFF readout- saved straight from camera</p> <p>Exposure values (time {sec}): 1/<strong>30</strong>, 1/<strong>60</strong>, 1/<strong>125</strong>, 1/<strong>200</strong>, 1/<strong>400</strong></p>
X-ray Fluorescence Mapping dataset for use in Heritage Science
<p>The following data sets were collected to support the potential uses of opensource data in the context of digital humanities and heritage sciences. </p> <p>This proposed experiment is conducted by the UCL Institute for Sustainable Heritage in collaboration with the Centre for Digital Humanities. Imaging methods including Photography, Multispectral Imaging, Hyperspectral Imaging and Xray Fluorescence Mapping have been collected along with the complete readout metadata of the instrumentation. </p> <p>We hope that you find the data helpful, and we welcome you to use the data in any way you wish, for all and any analysis development purposes. For us to build upon this research, we ask that in return you would be willing to share in some regard your experiences in using open-source data, using our data, successes and issues. </p> <p>If you would be willing to engage with us in this endeavor, please feel free to contact us so that we may be able to follow up with you. </p> <p>Other Data sets available <a href="https://zenodo.org/record/7319696#.Y3NuOXbP2Uk">Here</a></p> <p>E: <a href="mailto:molly.fort.21@ucl.ac.uk">molly.fort.21@ucl.ac.uk</a> </p> <p>Object Paradata; </p> <ul> <li><strong>Postcard – c. Early 1900's </strong></li> <li><strong>Language – Eng. </strong></li> <li><strong>Materials – colour print on card, metallic leafing. </strong></li> <li><strong>Front transcription - </strong></li> <li><strong> ‘Greetings’ </strong></li> <li><strong> ‘May your Birthday bring you Peace & perfect Happiness, Golden hopes & Love of Friends, And every Happiness this world can send.’ </strong></li> <li><strong>Object Dimensions – 138mm X 88mm</strong></li> </ul> <p>The postcard is an item of ephemera donated to the UCLDH Digitisation Suite by Prof Melissa Terras, for teaching and training purposes in 2015.</p> <p>This folder contains:</p> <p>X-ray fluorescence imaging map collected from a <a href="https://www.bruker.com/en/products-and-solutions/elemental-analyzers/micro-xrf-spectrometers/m4-tornado-plus.html">Bruker M4+ Tornado Micro-XRF System</a></p> <ul> <li>Postcardmap1.bcf - Bruker composite file containing full fluorescence spectral and mapping data with some other information. Can be read by Bruker software or using freely dowloadable <a href="https://hyperspy.org/">Hyperspy</a>.</li> <li>EDX.hdf5 - full fluorescence spectral and mapping data in open format, created and readable via. Hyperspy.</li> <li>Postcard_data.txt - metadata saved in ASCII format by Bruker system.</li> <li>postcardmap1_*.png - individual png images of mapped elements, listed in filename. </li> </ul> <p> </p>
A controlled vocabulary for research and innovation in the field of Cultural Heritage & Heritage Sciences
<p>This controlled vocabulary of keywords related to the field of Cultural Heritage and Heritage Sciences was built by SIRIS Academic in collaboration with IRPET (the Regional Institute for Economic Planning of Tuscany) and the ISPC (Institute of Heritage Science of CNR), in order to identify Cultural related research, development, and innovation activities. The work was carried out by consulting domain experts' advice, and it was ultimately applied to inform regional strategies on Cultural Heritage and research and innovation policy.</p> <p>The aim of this vocabulary is to enable one to retrieve texts (e.g. R&D projects and scientific publications) featuring the concepts included in the present vocabulary in their titles and abstracts, assuming that these records have a certain contribution of applications, techniques and issues, in the domain of Cultural Heritage and Heritage Sciences.</p> <p>The aim of this classification is to identify research products in the domain of Cultural Heritage, ranging from documents in some of its “traditional” disciplines, but also from documents emerging from interdisciplinary projects that apply novel areas and technologies in the domain of Cultural Heritage. The identification of texts in the domain of Cultural Heritage requires a task of text classification. Developing a method that could be applied to decide if a text can be relevant or have some relation to the domain of Cultural Heritage is a challenging task. The definition of what Cultural Heritage is and what it includes is a complex activity, even for domain experts. This is in particular because Cultural Heritage is quite a broad field of knowledge, and there is no full agreement on where the borders of the domain are. To define the scope of the perimeter, in this project, many of the available definitions were taken into account.</p> <p>Because of the high number of resources available in the domain, among thesauruses and taxonomies, the construction of a weakly-supervised controlled vocabulary was considered as the best way of retrieving documents in the domain. Since there is no annotated corpus/dataset of research texts in the domain capable of generalising the diversity of publications that can be related to the cultural domain, but stemming from different disciplines, we have opted for a text classification technique based on rules – specifically, a weakly-supervised controlled vocabulary.</p> <p>As defined by the Getty Institute, a controlled vocabulary is an organized arrangement of words and phrases used to index content and/or to retrieve content through browsing or searching. It typically includes preferred and variant terms and has a defined scope or describes a specific domain. The purpose of controlled vocabularies is to organize information and to provide terminology to catalogue and retrieve information. While capturing the richness of variant terms, controlled vocabularies also promote consistency in preferred terms and the assignment of the same terms to similar content (Harping, 2010).</p> <p>In short, Cultural Heritage is a rather abstractly-defined field, and Heritage Science is a particularly “fuzzy” field within Cultural Heritage. One of the main limitations of the approach we used is that the controlled vocabularies never capture all the lexical and linguistic variants of a term, and we may miss relevant texts if we cannot find the correct pattern to match during the search. But on the other hand, the controlled vocabulary is built from available vocabularies and thesauruses in the domain of Cultural Heritage, which are large resources. All the concepts in these resources are not included directly in the controlled vocabularies, because they would add noise to the classification. Therefore, the automatic weak supervision and a human curation of the final controlled vocabulary is fundamental for achieving correct results.</p> <p>The controlled vocabulary is built taking advantage of these four resources:</p> <ul> <li>The <a href="https://www.getty.edu/research/tools/vocabularies/aat/"><strong>Art and Architecture Thesaurus (AAT)</strong></a>: this is a structured vocabulary with approximately 34,000 concepts, including 131,000 words, descriptions and other information related to art, architecture, decorative arts, archival material and material culture, commonly used for cataloguing and for information retrieval.</li> <li> <p>Some cultural heritage categories in <strong><a href="https://en.wikipedia.org/wiki/Category:Cultural_heritage">Wikipedia</a> </strong>and <strong><a href="https://dbpedia.org/page/Cultural_heritage">DBpedia</a></strong>: these categories have been used to collect all related articles and subcategories, in order to obtain relevant, similar and specific instances of concepts linked to the domain. </p> </li> <li> <p>The <strong><a href="https://www.riches-project.eu/riches-taxonomy.html">RICHES Taxonomy</a></strong>: this taxonomy is a theoretical framework of related terms and their definitions, referring to the new concepts in the digital era, with the aim of defining the scope of some digital technologies applied to cultural heritage.</p> </li> <li> <p><strong><a href="https://www.heritagedata.org/blog/">Heritage Data - Linked Data Vocabularies for Cultural Heritage</a></strong>: a dataset which includes several cultural heritage thesauruses and vocabularies and is recognised as a reference point in the United Kingdom in the domain of cultural heritage.</p> </li> </ul> <p>The collection of concepts extracted from these four resources was composed of more than 60,000 terms, which have been refined as described in the next section.</p> <p> </p> <p><strong>## Automatic validation of the controlled vocabulary</strong></p> <p>In order to refine the collection of concepts to have a final set of relevant concepts and terms in the domain of Cultural Heritage, a semi-automatic validation has been applied to remove the irrelevant, too general, and ambiguous terms.</p> <p>To keep the relevant ones, the <a href="https://ncses.nsf.gov/pubs/nsb20206/specialization-and-impact-analysis-combined#:~:text=The%20specialization%20index%20(SI)%20is,the%20total%20output%20across%20all">specialization index (SI) </a>metric has been calculated for each of the keywords in the collection. In this case, the SI can be obtained measuring the fraction of publications with a keyword in a set of publications in the domain of Cultural Heritage and normalizing over the fraction of publications in the open domain with that keyword.</p> <p>After the calculation of the SI, all the keywords below a certain threshold are removed, and a manual supervision step is applied in order to remove non-pertinent keywords. An example of this automatic validation can be observed in the next table:</p> <table> <tbody> <tr> <td> <p><strong>Keyword</strong></p> </td> <td> <p><strong>Specialization Index</strong></p> </td> <td> <p><strong>Automatic threshold</strong></p> </td> <td> <p><strong>Manual supervision</strong></p> </td> </tr> <tr> <td> <p>male</p> </td> <td> <p>0.27</p> </td> <td> <p>Removed</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>3-d laser scanning</p> </td> <td> <p>0.7</p> </td> <td> <p>Removed</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>78 rpm records</p> </td> <td> <p>20.7</p> </td> <td> <p>Accepted</p> </td> <td> <p>Removed</p> </td> </tr> <tr> <td> <p>vienna</p> </td> <td> <p>3.48</p> </td> <td> <p>Accepted</p> </td> <td> <p>Removed</p> </td> </tr> <tr> <td> <p>radiocarbon dating</p> </td> <td> <p>13.6</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>graffiti</p> </td> <td> <p>25</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>bark painting</p> </td> <td> <p>20.7</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> <tr> <td> <p>pompeii</p> </td> <td> <p>16.23</p> </td> <td> <p>Accepted</p> </td> <td> <p>Accepted</p> </td> </tr> </tbody> </table> <p>The SI of the final keywords can be used as a probabilistic metric for each keyword.</p> <p>The final list of keywords was manually curated by domain experts.</p> <p> </p> <p><strong>## Evaluation of the controlled vocabulary</strong></p> <p>The final controlled vocabulary was evaluated with an external dataset with the aim of calculating its degree of precision. The evaluation dataset was composed of a collection of articles in 4 journals unequivocally considered to fall within the domain of Cultural Heritage. These four journals were: <em>(1) Journal Of Cultural Heritage, (2) Journal On Computing And Cultural Heritage, (3) Journal Of Cultural Heritage Management And Sustainable Development and (4) Digital Applications In Archaeology And Cultural Heritage.</em> This collection was composed of 5,000 articles, considered as the positive set, and another collection of randomly selected 5,000 articles outside of the Cultural Heritage domain, considered as the false set.</p> <p>The Cultural Heritage vocabulary was applied to the evaluation data set, obtaining a 95% of precision. After a set of improvements on the vocabulary, based on the exploration of publications not identified in the first test and the false positive results, we obtained a 98% of precision. The application of the vocabulary taking advantage of the probability of each keyword as its weight of being in the domain did not improve the results, and for this reason the probabilistic approach was discarded.</p> <p> </p> <p>## <strong>Using the vocabulary to classify publications concerning Cultural Heritage</strong></p> <p>The definition of the vocabulary does not, per se, allow to identify research contributions in Cultural Heritage: this is performed by actually matching the terms in the controlled vocabulary to the content of the gathered research textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance.</p> <p>In the following table we present some examples of the tagging process on some abstracts:</p> <table> <tbody> <tr> <td> <p><strong>Publication title</strong></p> </td> <td> <p><strong>Publication abstract</strong></p> </td> </tr> <tr> <td> <p>Egocentric visitor localization and artwork detection in cultural sites using synthetic data</p> </td> <td> <p>Computer vision and machine learning can be used in <strong>cultural heritage to augment the experience of visitors during the exploration of the cultural site</strong>, as well as to assist its management. To achieve such goals, two fundamental tasks should be addressed, i.e., localizing <strong>visitors and recognizing the observed artworks</strong>. Wearable cameras offer a convenient setting to address both tasks through the analysis of images acquired from the visitors’ points of view. However, the engineering of approaches to address such tasks generally requires large amounts of labeled data. We propose a tool which can be used to collect and automatically label synthetic visual data suitable to study image-based localization and artwork detection. The tool simulates a virtual agent navigating the <strong>3D model of a real cultural site</strong> and automatically captures video frames along with the related ground truth camera poses and semantic masks indicating the position of artworks. We generate a dataset of synthetic images starting from the 3D model of a <strong>museum located in Siracusa</strong>, Italy. The experiments suggest that the proposed tool allows to drastically reduce the effort needed to collect and label data, providing a means to generate large-scale datasets suitable to study localization and <strong>artwork detection in cultural sites</strong>.</p> </td> </tr> <tr> <td> <p>Discovering Leonardo with artificial intelligence and holograms: A user study</p> </td> <td> <p>Cutting-edge visualization and interaction technologies are increasingly used in<strong> museum exhibitions</strong>, providing novel ways to engage visitors and enhance their <strong>cultural experience</strong>. Existing applications are commonly built upon a single technology, focusing on visualization, motion or verbal interaction (e.g., high-resolution projections, gesture interfaces, chatbots). This aspect limits their potential, since museums are highly heterogeneous in terms of visitors profiles and interests, requiring multi-channel, customizable interaction modalities. To this aim, this work describes and evaluates an artificial intelligence powered, interactive holographic stand aimed at describing <strong>Leonardo Da Vinci's art</strong>. This system provides the users with accurate<strong> 3D representations of Leonardo's machines</strong>, which can be interactively manipulated through a touchless user interface. It is also able to dialog with the users in natural language about Leonardo's art, while keeping the context of conversation and interactions. Furthermore, the results of a large user study, carried out during art and tech exhibitions, are presented and discussed. The goal was to assess how users of different ages and interests perceive, understand and explore <strong>cultural objects </strong>when holograms and artificial intelligence are used as instruments of knowledge and analysis.</p> </td> </tr> <tr> <td> <p>Hybrid query expansion using lexical resources and word embeddings for sentence retrieval in question answering</p> </td> <td> <p>Question Answering (QA) systems based on Information Retrieval return precise answers to natural language questions, extracting relevant sentences from document collections. However, questions and sentences cannot be aligned terminologically, generating errors in the sentence retrieval. In order to augment the effectiveness in retrieving relevant sentences from documents, this paper proposes a hybrid Query Expansion (QE) approach, based on lexical resources and word embeddings, for QA systems. In detail, synonyms and hypernyms of relevant terms occurring in the question are first extracted from MultiWordNet and, then, contextualized to the document collection used in the QA system. Finally, the resulting set is ranked and filtered on the basis of wording and sense of the question, by employing a semantic similarity metric built on the top of a Word2Vec model. This latter is locally trained on an extended corpus pertaining the same topic of the documents used in the QA system. This QE approach is implemented into an existing QA system and experimentally evaluated, with respect to different possible configurations and selected baselines, for the <strong>Italian language and in the Cultural Heritage domain</strong>, assessing its effectiveness in retrieving sentences containing proper answers to questions belonging to four different categories.</p> </td> </tr> <tr> <td> <p>"3D reconstruction and validation of historical background for immersive VR applications and games: The case study of the Forum of Augustus in Rome"</p> </td> <td> <p>"In the last decades, thanks to the success of the video games industry, the sector of technologies applied to cultural heritage has begun to envisage, in this domain, new possibilities for the <strong>dissemination of heritage and the study of the past </strong>through edutainment models. More recently, experimentation in the field of<strong> virtual archaeology </strong>has led to the development of virtual museums and interactive applications. Among these, the “serious game” segment – the<strong> application of interactive technologies to the cultural heritage domain</strong> – is rapidly growing, also including immersive VR technologies. Applied VR games and applications are characterized by a thorough <strong>historical background and a validated 3D reconstruction</strong>. Indeed, producing such products requires a tailored workflow and large effort in terms of time and professionals involved to guarantee such faithfulness. Drawing on our previous work in the<strong> field of virtual archaeology</strong> and referring to recent experiences related to the deployment of applied VR games on PlayStation VR, we describe and assess a workflow for the production of <strong>historically accurate 3D assets</strong>, targeting interactive, immersive VR products. The workflow is supported by the case study of the <strong>Forum of Augustus </strong>and different output applications, highlighting peculiarities and issues emerging from a multi and interdisciplinary approach.</p> </td> </tr> </tbody> </table> <p>Through this classification process, we identified projects and publications related to heritage, with different levels of relationship and relevance, but mostly relevant to understanding the research competencies in the domain. The resulting research records were reviewed by experts in the domain, given the occurrence of some false positives.</p> <p>Among the main strengths of this step, it’s worth mentioning the fact that the vocabulary is broad and not restricted to the field of Heritage Science (that is, to STEM applications in Cultural Heritage), as it takes advantage of a variety of available resources. Moreover, by looking directly at the textual data, instead of using the assigned bibliometric areas, we can better capture interdisciplinary research. The limitations of this approach were presented at the beginning of this document: for example, relevant texts could be missed if the correct pattern to match during the search is not found.</p> <p> </p> <p><strong>## Vocabulary of concepts related to Key Enabling Technologies in the domain of Cultural Heritage and Culture</strong></p> <p>For the development of this vocabulary, the definition of key enabling technologies in the domain of Cultural Heritage and Culture, was based on reference of the <a href="http://www.irpet.it/archives/53165">report 'Technologies, Cultural Heritage and Culture' published on March 2019</a> by IRPET.</p> <p>A vocabulary for each Key Enabling Technology (hereafter, KET) was prepared by extracting the relevant concepts, words, technologies and examples from the Platform Report document 'Technologies, Cultural Heritage and Culture, within APPENDIX A. DESCRIPTION OF MAIN TECHNOLOGIES FOR ROADMAP (p. 45-61). Each vocabulary contains a set of terms divided into subdomains.</p> <p>The KETs have been divided into the following six groups:</p> <ul> <li> <p>ICT</p> </li> <li> <p>PHOTONICS, MICRO- AND NANO-ELECTRONICS</p> </li> <li> <p>PLATFORMS</p> </li> <li> <p>NANO AND BIOTECHNOLOGY, ADVANCED MATERIALS</p> </li> <li> <p>PARTICLE ANALYTICAL SYSTEMS</p> </li> </ul> <p>The initial keywords extracted from the document were enriched following the approach based on semantic keyword enrichment based on combination of concurrent keywords and word embeddings (Duran-Silva et al., 2019; Duran-Silva et al., 2021).</p> <p>This second vocabulary has to be used in combination with the Cultural Heritage vocabulary to capture KETs within the domain of cultural heritage.</p> <p> </p> <p>## <strong>Use of the controlled vocabulary</strong></p> <p>The definition of the vocabulary does not, per se, allow identifying STI contributions to the domain: this activity in fact boils down to actually matching the terms in the controlled vocabulary to the content of the gathered STI textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance. Some relatively ambiguous keywords (which may match unwanted pieces of text), have a set of associated “extra” terms. These “extra” terms are defined as further terms that must co-appear, in the same sentence, together with their associated ambiguous keywords. SIRIS Academic has developed the <a href="https://github.com/sirisacademic/VocTagger">voc_tagger tool</a>, a multiprocess information extraction system able to identify hidden knowledge in textual documents using “controlled vocabularies”, openly available at GitHub and compatible with these controlled vocabularies.</p> <p> </p> <p><strong>## Bibliography</strong></p> <p>Harpring, P. (2010). Introduction to controlled vocabularies: terminology for art, architecture, and other cultural works. Getty Publications.</p> <p>Nicolau Duran-Silva, Enric Fuster, Francesco Alessandro Massucci, César Parra-Rojas, Arnau Quinquillà, Fernando Roda, Bernardo Rondelli, Nicandro Bovenzi, & Chiara Toietta. (2021). A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI) (Version 2) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5591987</p> <p>Duran-Silva, Nicolau, Fuster, Enric, Massucci, Francesco Alessandro, & Quinquillà, Arnau. (2019). A controlled vocabulary defining the semantic perimeter of Sustainable Development Goals (1.2) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3567769</p>
Book of Abstracts from Conference for the Society for the Preservation of Natural History Collections (SPNHC), International Partner – BHL (Biodiversity Heritage Library) and National Partner – NatSCA (Natural Sciences Collections Association) 2022
<p>The complete book of abstracts (Posters & Oral Presentations). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.