Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
86
datasets available to search
ShareScore release 0.7.1
Dataset results
86 results for “semantic data”
Raw data for manuscript Semantic context can mask intelligibility declines at above-conversational speech levels in normal-hearing listeners
<p>Raw data for the manuscript in doc file. <br> Copied from the Matlab .m file. used for the analysis.</p> <p>To be updated.</p> <p>For details, contact me at mfer@health.sdu.dk</p>
TFG Systematization process of generating semantic data and ontology
<pre>Set of TALIS files, the main source of data for the investigation, in CSV format.General ontology manually and by new software. Semantic data, as well as the DSL code. Finally, the web design of the new tool.</pre>
Source data of the study of semantic and syntactic relations
<p>The data contain detailed answers of test participants to questions asked in accordance with the scenario of examining semantic and syntactic relations between tactile cartographic signs, planned to be placed on tactile maps of historic gardens in various garden design styles. The study was conducted as part of testing the legibility of cartographic tactile signs, realized within the research project No. Rzeczy są dla ludzi/0005/2020-00, titled “Technology for the development of tactile maps of historic parks”, financed by the National Centre for Research and Development, for the years 2021–2024, and realized at the Military University of Technology, Faculty of Civil Engineering, and Geodesy (Warsaw, Poland).</p>
Data from: Two complementary AI approaches for predicting UMLS semantic group assignment: heuristic reasoning and deep learning
<p><strong>Objective</strong>: Use heuristic, deep learning (DL), and hybrid AI methods to predict semantic group (SG) assignments for new UMLS Metathesaurus atoms, with target accuracy ≥ 95%.</p> <p><strong>Materials and Methods</strong>: We used train-test datasets from successive 2020AA-2022AB UMLS Metathesaurus releases. Our heuristic "waterfall" approach employed a sequence of seven different SG prediction methods. Atoms not qualifying for a method were passed on to the next method. The DL approach generated BioWordVec and SapBERT embeddings for atom names, BioWordVec embeddings for source vocabulary names, and BioWordVec embeddings for atom names of the second-to-top nodes of an atom's source hierarchy. We fed a concatenation of the four embeddings into a fully connected multi-layer neural network with an output layer of 15 nodes (one for each SG). Both methods were capable of estimating the probability that their predicted SG for an atom would be correct. We developed two hybrid SG prediction methods combining the strengths of heuristic and DL methods.</p> <p><strong>Results</strong>: The heuristic waterfall approach accurately predicted 94.3% of SGs for 1,563,692 new unseen atoms. The DL accuracy on the same dataset was also 94.3%. The hybrid approaches achieved an average accuracy of 96.5%.</p> <p><strong>Conclusion</strong>: Our study demonstrated that AI methods can predict SG assignments for new UMLS atoms with sufficient accuracy to be potentially useful as an intermediate step in the time-consuming task of assigning new atoms to UMLS concepts (CUIs). We showed that for SG prediction, combining heuristic methods and DL methods can produce better results than either alone.</p>
Mapping data files to semantic data models using the CaosDB crawler
<p>Data from data acquisition can lead to a high variety of data files on file systems. The figure illustrates that these files can be mapped to semantic data models in the research data management system CaosDB using a customizable crawler.</p>
Data from: Intermediate acoustic-to-semantic representations link behavioural and neural responses to natural sounds
Open the record for dataset details and reuse information.
Data from: Early detection of encroaching woody Juniperus virginiana and its classification in multi-species forest using UAS imagery and semantic segmentation algorithms
Open the record for dataset details and reuse information.
Data from: Two complementary AI approaches for predicting UMLS semantic group assignment: heuristic reasoning and deep learning
Open the record for dataset details and reuse information.
Piveau: A Large-scale Open Data Management Platform based on Semantic Web Technologie
<p>This file contains the sources that were used to create the feature comparison in "Piveau: A Large-scale Open Data Management Platform based on Semantic Web Technologies".</p>
Semantic Web und Linked Data: Generierung von Interoperabilität in archäologischen Fachdaten am Beispiel römischer Töpferstempel - Datasets
<p><strong>Datasets</strong></p> <p>Gegenstand der Masterarbeit ist die Verwendung aktueller Technologien interoperabler Datenhaltung, insbesondere das Konzept der Linked Open Data (LOD) und der semantischen Modellierung, zur Verdeutlichung ihres Potentials in archäologischen Informationen am Beispiel von Terra Sigillata-Fundorten, -Töpfern und -Keramikfragmenten. Die Arbeit zeigt eine Migration von Daten, sowie die Möglichkeiten und die Problematik der Modellierung der Attribute und Beziehungen mit Hilfe bestehender LOD-Konzepte und kontrollierter Vokabularien, sowie eigene Ansätze zur Lösung. Diese Daten werden mittels REST-Schnittstelle zur Verfügung gestellt. Ein Schwerpunkt wird auf die Verlinkung zu anderen bereits bestehenden Projekten gelegt, wodurch eine Vielzahl weiterer archäologischer und historischer Informationen z.B. über das Pelagios Projekt eingebunden werden. Zudem wird das Potential der Verlinkung und Abfrage von heterogenen Informationen zwischen Töpfern, Fragmenten und Orten deren relativ chronologische Beziehungen über LOD mit einer webbasierten Schnittstelle aufgezeigt.</p> <p>The subject matter of this master thesis is using current technologies in interoperable data management, in particular the illustration of the potential of Linked Open Data (LOD) and semantic modelling in archaeological information, as used on samian ware places and their corresponding potters and ceramic fragments. The thesis demonstrates a migration of data as well as possibilities and problems of modelling attributes and relationships using existing LOD concepts and controlled vocabularies as well as novel self-developed approaches to the solution. These data are provided by a ReST interface. One focus is linking to other existing projects, creating associations to other archaeological and historical information, for example the Pelagios project. Moreover, a web-based interface shows the potential of linking and retrieval of heterogeneous information among pottery, fragments and places and their relative chronological relationships via LOD.</p>
A process proposal on how to move from data requirements to semantic data in the context of the MDA
<p>This work proposes a process for the identification of the requirements associated with the data, along with and a set of transformations in Model Driven Architecture (MDA) context, in order to obtain a semantically annotated dataset, as a result of the unification and alignment of the data in the context of its initial domain. Our proposal identifies four phases (from CIM to code), in which is describe the artifacts and transformations required to progress to the next phase: a target domain model is first obtained from the data requirements, after which the ontological schema and the ontology is generated from the previous model. A domain-specific language (DSL), also proposed in this work, is then used to obtain the semantic data model (the DSL code), which generates the final semantic dataset. We have validated the proposal by studying two cases: one with data from the public transport domain and the other with data concerning those affected by the COVID-19 pandemic.</p>
Processed data for the "Deriving Semantics-Aware Fuzzers from Web API Schemas" paper
<p>Processed data for the "Deriving Semantics-Aware Fuzzers from Web API Schemas" paper. Each directory in the archive consists of:</p> <p>- metadata.json. Metadata about a test run - tested fuzzer name, run duration, etc</p> <p>- fuzzer.json - Structured fuzzer output</p> <p>- deduplicated_cases.json - Deduplicated reported failures, when fuzzers provide it</p> <p>- sentry.json - Cleaned Sentry events for this run</p> <p>- target.json - Parsed stdout for Gitlab & Disease.sh targets that were tested without Sentry integration</p>
Semantic object-scene inconsistencies affect eye movements, but not in the way predicted by contextualized meaning maps - data
<p>Data from the article<strong><em> Semantic object-scene inconsistencies affect eye movements, but not in the way predicted by contextualized meaning maps</em></strong> published in Journal of Vision.</p> <p>code: https://zenodo.org/record/5999215<br> data: https://zenodo.org/record/5999046</p> <p><br> Marek A. Pedziwiatr<br> marek.pedziwi@gmail.com<br> February 2022</p>
Aesthetic Trends and Semantic Web Adoption of Media Outlets Identified through Automated Archival Data Extraction
<p>This dataset includes a variety of structured data gathered via various Web data extraction techniques which were employed in order to collect current and archival data from almost a thousand news websites that are popular in Greece, for the purpose of monitoring and recording their progress through time. The collected information, that took the form of a website’s source code and an impression of their homepage in different time instances of the last decade, has been used to identify trends concerning Semantic Web integration, DOM structure complexity, number of graphics, color usage and more. In total more than ten thousands impressions (including screenshots and source code) were analyzed which resulted to conclusions regarding the evolution of aesthetics and the adoption of new technologies.</p>
Association Rules and Semantic Relatedness (ARSR) - Evaluation Data
<p><em><strong>ARSR</strong></em> stands for "<strong><em>Association Rules and Semantic Relatedness</em></strong>",<br> a recommender system that combines association rule mining and semantic relatedness to create recommendations for form fields.</p> <p>ARSR was evaluated with focus on recommending values for fields of metadata forms.<br> The evaluation was performed with two sets of association rules (R1 and R2).<br> <strong>R1</strong> was primarily used to assess the recommendation performance.<br> <strong>R2</strong> served as an alternative and led to similar results.</p> <p>This dataset contains the raw data ("<strong>collected</strong>", JSON format), generated during the evaluation.<br> It includes i.a. input (populated fields and target field), expected output, and the top 40 generated recommendations for each test combination.</p> <p>In addition, the processed data ("<strong>analysed</strong>", CSV format) is provided.<br> It is based on the raw data and is used to calculate metrics and plot the results.</p> <p>The <strong>source code</strong> for ARSR and the evaluation is available on <a href="https://gitlab.com/dlr-dw/arsr">GitLab (gitlab.com)</a>.</p> <p> </p> <p><em>Note:</em><br> To perform the analysis on the raw data (collected) yourself, make sure to follow the setup instruction in the evaluation repository first.<br> More specifically, install dependencies and unzip "data/cedar/test-instances.zip" so that URI mappings can be accessed.<br> Then follow the instructions provided by the README file in the raw data archives (e.g. collected-R1.zip).</p>
Figures - Semantic analysis of web archive historical data 1983 "Marche pour l'égalité et contre le racisme"
Open the record for dataset details and reuse information.
Data for "Improving semantic video retrieval models by training with a relevance-aware online mining strategy"
<p>This repository contains all the data available for the publication:</p> <p><a href="https://doi.org/10.1016/j.cviu.2024.104035">Alex Falcon, Giuseppe Serra, and Oswald Lanz. <em>Improving semantic video retrieval models by training with a relevance-aware online mining strategy</em>. <strong>Computer Vision and Image Understanding</strong>. 2024.</a></p> <p>Code is available at: <a href="https://github.com/aranciokov/ranp/">https://github.com/aranciokov/ranp/</a></p> <p>The data includes:</p> <ul> <li>pre-extracted features (ordered_feature_*.zip files)</li> <li>annotations, such as pre-extracted semantic graphs, glove checkpoints, class annotations, etc (annotations_*.zip files)</li> <li>train/val/test, when available, split information (public_split_*.zip) files</li> <li>pretrained models for HGR and EAO (details in the github repo)</li> </ul>
Dataset for "Learning Scene Semantics from Vehicle-centric Data for City-scale Digital Twins", Fürntratt et al.
<p>Dataset for "Learning Scene Semantics from Vehicle-centric Data for City-scale Digital Twins", Fürntratt et al.</p> <p>Data are anonymized and provided with segmentation mask ground truths. </p>
Unsupervised detection of semantic correlations in big data
<p>simulated spin configurations from statistical mechanics models at thermal equilibrium. Further details are in our manuscript with the same title. </p>
Data from: Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narrative speech
People routinely hear and understand speech at rates of 120–200 words per minute [1, 2]. Thus, speech comprehension must involve rapid, online neural mechanisms that process words' meanings in an approximately time-locked fashion. However, in the context of continuous speech, electrophysiological evidence for such time-locked processing has been lacking. Whilst valuable insights into the semantic processing of speech have been provided by the "N400 component" of the event-related potential [3-6], this literature has been dominated by paradigms using incongruous words within specially constructed sentences, and may not accurately reflect natural, narrative speech comprehension. Building on the discovery that cortical activity "tracks" the dynamics of running speech [7-9], and psycholinguistic work both demonstrating [10-12] and modeling [13-15] how context rapidly impacts on word processing, we describe a new approach for deriving an electrophysiological correlate of natural speech comprehension. We used a computational model [16] to quantify the meaning carried by each word based on how semantically dissimilar it was to its preceding context and then regressed this quantity against electroencephalographic (EEG) data recorded from subjects as they listened to narrative speech. This produced a prominent negativity at a time-lag of 200–600 ms on centro-parietal EEG channels, characteristics common to the N400. Applying this approach to EEG datasets involving time-reversed speech, cocktail party attention and audiovisual speech-in-noise demonstrated that this response was very sensitive to whether or not subjects understood the speech they heard. These findings demonstrate that, when successfully comprehending natural speech, the human brain responds to the contextual semantic content of each word in a relatively time-locked fashion.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.