Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
483
datasets available to search
ShareScore release 0.9.0
Dataset results
483 results for “SEMANTICS”
Semantic Web resources and Machine Learning systems - Knowledge Graph (SWeMLS-KG)
<p>This resource is part of our submission to ESWC 2023 resource track, which includes:</p> <p>Datasets:<br> - Folder "pattern" - a set of SWeMLS patterns represented based on OPMW and P-Plan ontology,<br> - Folder "shapes" - a set of SHACL constraints to check the conformance of SWeML Systems against SWeMLS patterns as well as a set of SHACL-AF rules to generate links between system components,<br> - File "swemls-ontology.ttl" - an ontology to represent Semantic Web resources and Machine Learning systems (SWeMLS),<br> - File "swemls-instances.ttl" - a set of triples representing the extracted metadata from 476 SWeML systems and papers,<br> - File "swemls-kg.ttl" - an integrated and validated KG containing all above files, including enrichment from SHACL-AF rules using "swemls-toolkit" [2].</p> <p>These resources are produced based on the result of the Systematic Mapping Study (SMS) reported in [1]. The latest SNAPSHOT-version of the resource can be accessed through our resource landing page: <a href="https://w3id.org/semsys/sites/swemls-kg/">https://w3id.org/semsys/sites/swemls-kg/</a></p> <p>[1] Breit, A., Waltersdorfer, L., Ekaputra, J.F., Sabou, M., Ekelhart, A., Iana, A., Paulheim, H., Portisch, J., Revenko, A., Ten Teije, A., van Harmelen, F.: Combining Machine Learning and Semantic Web -A Systematic Mapping Study (under review). ACM CSUR (2022)<br> [2] Source code of swemls-toolkit is available at: https://github.com/semanticsystems/swemls-toolkit</p>
Unified MOOCs Semantic Search Engine
<p>Full dataset of Unified MOOCs</p>
How doctors apply semantic components to specify search in work-related information retrieval
<p>Workplace searching is often context-specific and targets a ‘right answer’ within some<br> domain-specific aspect of the search topic. We have developed the semantic component<br> (SC) model that allows searchers to specify a search within context-specific aspects of the<br> main topic of documents. The goal of our study was to gain insight into how family practice<br> physicians at sundhed.dk, a national healthcare portal in Denmark, applied the SC model<br> to formulate queries to solve work-related search tasks. The results showed that doctors<br> used the model purposively when choosing search facets and search concepts. They were<br> relatively consistent in their use. The findings provide promising evidence of the model’s<br> potential usefulness.</p>
Background-Foreground-Segmentation Labels for "Self-improving Semantic Perception on a Construction Robot"
<p>Background-foreground-segmentation labels the paper "Self-improving Semantic Perception on a Construction Robot" of CoRL 2021</p>
TFG Systematization process of generating semantic data and ontology
<pre>Set of TALIS files, the main source of data for the investigation, in CSV format.General ontology manually and by new software. Semantic data, as well as the DSL code. Finally, the web design of the new tool.</pre>
Source data of the study of semantic and syntactic relations
<p>The data contain detailed answers of test participants to questions asked in accordance with the scenario of examining semantic and syntactic relations between tactile cartographic signs, planned to be placed on tactile maps of historic gardens in various garden design styles. The study was conducted as part of testing the legibility of cartographic tactile signs, realized within the research project No. Rzeczy są dla ludzi/0005/2020-00, titled “Technology for the development of tactile maps of historic parks”, financed by the National Centre for Research and Development, for the years 2021–2024, and realized at the Military University of Technology, Faculty of Civil Engineering, and Geodesy (Warsaw, Poland).</p>
Data from: Two complementary AI approaches for predicting UMLS semantic group assignment: heuristic reasoning and deep learning
<p><strong>Objective</strong>: Use heuristic, deep learning (DL), and hybrid AI methods to predict semantic group (SG) assignments for new UMLS Metathesaurus atoms, with target accuracy ≥ 95%.</p> <p><strong>Materials and Methods</strong>: We used train-test datasets from successive 2020AA-2022AB UMLS Metathesaurus releases. Our heuristic "waterfall" approach employed a sequence of seven different SG prediction methods. Atoms not qualifying for a method were passed on to the next method. The DL approach generated BioWordVec and SapBERT embeddings for atom names, BioWordVec embeddings for source vocabulary names, and BioWordVec embeddings for atom names of the second-to-top nodes of an atom's source hierarchy. We fed a concatenation of the four embeddings into a fully connected multi-layer neural network with an output layer of 15 nodes (one for each SG). Both methods were capable of estimating the probability that their predicted SG for an atom would be correct. We developed two hybrid SG prediction methods combining the strengths of heuristic and DL methods.</p> <p><strong>Results</strong>: The heuristic waterfall approach accurately predicted 94.3% of SGs for 1,563,692 new unseen atoms. The DL accuracy on the same dataset was also 94.3%. The hybrid approaches achieved an average accuracy of 96.5%.</p> <p><strong>Conclusion</strong>: Our study demonstrated that AI methods can predict SG assignments for new UMLS atoms with sufficient accuracy to be potentially useful as an intermediate step in the time-consuming task of assigning new atoms to UMLS concepts (CUIs). We showed that for SG prediction, combining heuristic methods and DL methods can produce better results than either alone.</p>
Mapping data files to semantic data models using the CaosDB crawler
<p>Data from data acquisition can lead to a high variety of data files on file systems. The figure illustrates that these files can be mapped to semantic data models in the research data management system CaosDB using a customizable crawler.</p>
Wikibio: a Semantic Resource for the Intersectional Analysis of Biographical Events
<p>If you use this resource please cite</p> <p> </p> <pre>@inproceedings{stranisci2023wikibio, title={WikiBio: a Semantic Resource for the Intersectional Analysis of Biographical Events}, author={Stranisci, Marco Antonio and Damiano, Rossana and Mensa, Enrico and Patti, Viviana and Radicioni, Daniele and Caselli, Tommaso and others}, booktitle={Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, volume={1}, pages={12370--12384}, year={2023}, organization={Association for Computational Linguistics} }</pre>
Natural Language Annotations for Reasoning about Program Semantics
<p>Natural language annotations about Python statements</p> <p>The dataset is made of pairs of files sharing the prefix of the filename</p> <p>* Annotations are in JSONL format (filenames ending with '_annot'), i.e. JSON objects separated by newline ('\n') characters</p> <p>* Reference source code files are in JSON format. (filenames ending with '_code')</p> <p> </p> <p>Source dataset : Programming Puzzles (Schuster et al. 2021, NeurIPS Dataset and benchmarks track) - MIT License - https://github.com/microsoft/PythonProgrammingPuzzles</p>
Word-Retrieval Treatment for Aphasia: Semantic Feature Analysis
ClinicalTrials.gov study NCT00125242. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Semantic Memory, Financial Capacity, and Brain Perfusion in Mild Cognitive Impairment (MCI) (CASL)
ClinicalTrials.gov study NCT00880555. IPD Sharing: NO. Countries: 1. Publications: 3.
Data from: Intermediate acoustic-to-semantic representations link behavioural and neural responses to natural sounds
Open the record for dataset details and reuse information.
Data from: Early detection of encroaching woody Juniperus virginiana and its classification in multi-species forest using UAS imagery and semantic segmentation algorithms
Open the record for dataset details and reuse information.
Data from: Two complementary AI approaches for predicting UMLS semantic group assignment: heuristic reasoning and deep learning
Open the record for dataset details and reuse information.
Decrypting cryptic crosswords: Semantically complex wordplay puzzles as a target for NLP
Open the record for dataset details and reuse information.
U-net for automated thoracic CT semantic segmentation
Open the record for dataset details and reuse information.
Christchurch Aerial Semantic Dataset
<p> </p> <p>The <strong>Christchurch Aerial Semantic Dataset (CASD)</strong> comprises aerial imagery at very high resolution over Christchurch, New-Zealand, and reference semantic data for urban objects such as buildings, vegetation and vehicles.</p> <p>It aims to foster developments of new methods for Earth observation including automatic image classification, object detection, semantic segmentation, semi-supervised learning, etc...</p> <p><strong>Aerial imagery</strong> consists in the Christchurch Earthquake Imagery dataset, released by Land Information New Zealand:</p> <p><a href="https://www.linz.govt.nz/land/maps/linz-topographic-maps/imagery-orthophotos/christchurch-earthquake-imagery">https://www.linz.govt.nz/land/maps/linz-topographic-maps/imagery-orthophotos/christchurch-earthquake-imagery</a></p> <p><strong>Annotations</strong> were produced by ONERA/DTIS on 4 images. Three classes were tagged: buildings (797 objects), cars (2357 objects), and vegetation (938 objects). All objects are given as polygonal bounding boxes (shapefiles) and image semantic masks (rasters).</p> <p> </p> <p><strong>License</strong></p> <p>This imagery is licensed under a Creative Commons Attribution 4.0 International licence (<a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a>) with Crown copyright reserved. This means anyone is free to copy, distribute, and adapt the imagery so long as it is attributed to the Crown, eg ”Crown Copyright Reserved.”</p> <p>The annotations are licensed also under a Creative Commons Attribution 4.0 International licence (<a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a>) with authors’ and ONERA copyright reserved. This means anyone is free to copy, distribute, and adapt the annotations as long as they are attributed to the authors and ONERA.</p> <p>For research papers, acknowledgments can be done by citing the authors’ works:</p> <p><em>H. Randrianarivo, B. Le Saux, and M. Ferecatu. Man-made structure detection with deformable part-based models. In IEEE Int. Geoscience and Remote Sensing Symposium (IGARSS), Melbourne, Australia, July 2013</em></p> <p><em>N. Audebert, B. Le Saux, and S. Lefèvre. Segment-before-Detect: Vehicle Detection and Classification through Semantic Segmentation of Aerial Images. Remote Sensing, 9(4):1–18, April 2017</em></p> <p> </p>
Piveau: A Large-scale Open Data Management Platform based on Semantic Web Technologie
<p>This file contains the sources that were used to create the feature comparison in "Piveau: A Large-scale Open Data Management Platform based on Semantic Web Technologies".</p>
Datasets for: Semantic Robustness of Models of Source Code
<p>Datasets for Semantic Robustness of Models of Source Code.</p> <p>Includes the c2s/java-small, csn/java, csn/python, and sri/py150 in the following representations:</p> <ol> <li>Raw [in raw.tar.gz]</li> <li>Normalized [in normalized.tar.gz]</li> <li>Pre-processed (<em>ast-paths and tokens</em>) [in preprocessed.tar.gz]</li> <li>Transformed [in transformed.tar.gz] <ol> <li>Normalized <ol> <li>transforms.All</li> <li>transforms.ShuffleLocalVariables</li> <li>transforms.ShuffleParameters</li> <li>transforms.RenameLocalVariables</li> <li>transforms.RenameFields</li> <li>transforms.RenameParameters</li> <li>transforms.ReplaceTrueFalse</li> <li>transforms.InsertPrintStatements</li> <li>transforms.Identity</li> </ol> </li> <li>Pre-processed (<em>ast-paths and tokens</em>) <ol> <li>transforms.Identity</li> <li>transforms.InsertPrintStatements</li> <li>transforms.ReplaceTrueFalse</li> <li>transforms.RenameParameters</li> <li>transforms.RenameFields</li> <li>transforms.RenameLocalVariables</li> <li>transforms.ShuffleParameters</li> <li>transforms.ShuffleLocalVariables</li> <li>transforms.All</li> </ol> </li> </ol> </li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.