Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
483
datasets available to search
ShareScore release 0.9.0
Dataset results
483 results for “SEMANTICS”
Measuring semantic memory using associative and dissociative retrieval tasks
<p>Recent theoretical advances highlighted the need for novel means of assessing semantic cognition. Here, we introduce the Associative-Dissociative Retrieval Task (ADT), positing a novel way to test inhibitory control over semantic memory retrieval by contrasting the efficacy of associative (automatic) and dissociative (controlled) retrieval on standard set of verbal stimuli. All ADT measures achieved excellent reliability, homogeneity, and short-term temporal stability. Moreover, in-depth stimulus level analyses showed that associating is easier for words evoking few but strong associates, yet such propensity hampers the inhibition. Finally, we provided critical support for the construct validity of the ADT measures, demonstrating reliable correlations with domain-specific measures of semantic memory functioning (semantic fluency and associative combination) but negligible correlations with domain-general capacities (processing speed and working memory). Together, we show that ADT provides simple yet potent and psychometrically sound measures of semantic memory retrieval and offers noteworthy advantages over the currently available assessment methods.</p>
Supplementary Data for "Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"
<p>This dataset contains quality assessment results for 26 vocabularies. The assessment was conducted using the <a href="https://skos-play.sparna.fr/skos-testing-tool/">qSKOS vocabulary quality assessment tool</a>.</p> <p>The 26 assessed vocabularies were converted from their original formats into the Simple Knowledge Organization System (SKOS) data model using the approach described in our paper titled <a href="https://doi.org/10.1007/978-3-031-62362-2_9">"Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"</a>, presented at the <a href="https://doi.org/10.1007/978-3-031-62362-2">24th International Conference on Web Engineering (ICWE 2024)</a>.</p> <p>The dataset contains a quality assessment for the following vocabularies:</p> <ol> <li>A Taxonomy of Evaluation Towards Standards</li> <li>Cross-Device Taxonomy</li> <li>What Makes a Data-driven Business Model? A Consolidated Taxonomy</li> <li>DDI Aggregation Method</li> <li>DDI Mode of Collection</li> <li>Building a New Taxonomy for Data Discretization Techniques</li> <li>Demopaedia</li> <li>Data Science Glossary</li> <li>A Taxonomy of Evaluation Approaches in Software Engineering</li> <li>Evaluation Thesaurus</li> <li>The Glossary of Human Computer Interaction</li> <li>Human-Factors Taxonomy</li> <li>A Taxonomy to Structure and Analyze Human–Robot Interaction</li> <li>A Taxonomy of Interaction for Instructional Multimedia</li> <li>A Taxonomy of Interrogation Methods</li> <li>Design Vocabulary for Human–IoT Systems Communication</li> <li>Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors</li> <li>Thesaurus Mass Communication</li> <li>Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey</li> <li>A Taxonomy of Quality of Service and Quality of Experience of Multimodal Human-Machine Interaction</li> <li>A Human-Centered Taxonomy of Interaction Modalities and Devices</li> <li>A Taxonomy of Spatial Interaction Patterns and Techniques</li> <li>A Taxonomy of Social Errors in Human-Robot Interaction</li> <li>Taxonomy of Digital Research Activities in the Humanities</li> <li>Virtual Reality and the CAVE: Taxonomy, Interaction Challenges and Research Directions </li> <li>Cross-Device Interaction</li> </ol>
dataset for "basic setting", "+ binary semantic loss", "+ class weights", "+ height weights", "+ region weights", "+ elastic distortion and subsampling", "+ TreeMix" in paper Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learning
<p>dataset for "basic setting", "+ binary semantic loss", "+ class weights", "+ height weights", "+ region weights", "+ elastic distortion and subsampling", "+ TreeMix" in paper Automated forest inventory: analysis of high-density airborne LiDAR point clouds with 3D deep learning</p>
PSemQE: Disambiguating Short Queries Through Personalised Semantic Query Expansion
<p>This repository contains research data associated to the publication "PSemQE: Disambiguating Short Queries Through Personalised Semantic Query Expansion". Included are the experimental results, as well as a knowledge graph of Wikipedia articles and categories.</p> <p>For further details, please see the README.md</p>
Supplementary Material for "To Do or Not to Do: Semantics and Patterns for Do Activities in UML PSSM State Machines"
<p>This dataset provides artifacts about the semantics of doActivity in the Precise Semantics of UML State Machines (PSSM) specification. It collects:</p> <ul> <li>execution traces and screenshots from two simulators (Cameo, Papyrus Moka),</li> <li>analysis about doActivity features present in PSSM test suite, focusing on concurrency,</li> <li>collection of state machine models and doActivity patterns used in the Thirty Meter Telescope (TMT) SysML model.</li> </ul> <p>Find the related paper at <a href="https://arxiv.org/abs/2309.14884" target="_blank" rel="noopener">arXiv:2309.14884</a>.</p>
tBiomed: Semantic Table Annotations Benchmark for Biomedical Domain
<p><strong>tBiomed </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiomed </strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p><strong>tBiomed </strong>contains <strong>26,778</strong> entity and horizontal tables, while this repository contains only a <strong>validation fold</strong> of the original data representing <strong>20%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>1</strong> <strong>GB</strong>.</p> <p>We included the full version of the dataset. We will update this repository ground truth data of the test set in the Future.</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>
SemTab 2024: Semantic Web Challenge on Tabular Data to Knowledge Graph Matching Data Sets - WikidataTables2024R1 and WikidataTables2024R2
<p>Data Sets from the ISWC 2024 Semantic Web Challenge on Tabular Data to Knowledge Graph Matching, Round 1, Wikidata Tables. Links to other datasets can be found on the challenge website: https://sem-tab-challenge.github.io/2024/ as well as the proceedings of the challenge published on CEUR.</p> <p>For details about the challenge, see: http://www.cs.ox.ac.uk/isg/challenges/sem-tab/</p> <p>For 2024 edition, see: https://sem-tab-challenge.github.io/2024/</p> <p>Note on License: This data includes data from the following sources. Refer to each source for license details:<br>- Wikidata https://www.wikidata.org/</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>
Evaluation Results - Semantic Zoom With Immersive Detail View for ExplorViz
<p>This archive contains the evaluation results of the master thesis 'Semantic Zoom With Immersive Detail View for ExplorViz'.</p> <p>The evaluation is divided into a user evaluation of usability and user performance and a rendering performance evaluation.</p> <p>The evaluation compares the version of ExplorViz with Semantic Zoom and without Semantic Zoom.</p> <p>The complete user survey can be viewed in the PDF: 'Printed version of the survey - ExplorViz with Semantic Zoom.pdf'.</p> <p><br>- The file 'survey_archive_277626.lsa' is exported from LimeSurvey and contains the survey and the responses.<br>- results-survey277626.csv' contains the results in csv format.<br>- The file 'results-statistics.pdf' is a pdf that contains statistics about the survey results.<br>- The file 'results-all-answers-ExplorViz with Semantic Zoom.pdf' lists all the participants' answers in text format.<br>- The file 'allChartImages.zip' displays the results data in graphs.</p> <p><br>As part of a performance evaluation of the frontend, a Python script using Selenium was used.<br>The results can be found in the csv files:<br>- 'performance_RendertimeTracegen - XXXL world with high communication2024-11-19--22-13-36-SZLongTerm'<br>- 'performance_RendertimeTracegen - XXXL world with high communication2024-11-19--22-09-24-NoSZLongTerm'</p> <p>The Python script is split into two files:<br>- 'selenium_test.py'<br>- 'helpers.py'</p>
ForestSemantic: A Dataset for Semantic Learning of Forest from Close-Range Sensing
<p><strong>ForestSemantic</strong> is a dataset for forest semantic studies at both tree- and plot-levels. The dataset supports both instance and semantic segmentation, such as the tree detection and segmentation and the classification of ground, trunk, branches, and foliage components at both tree- and plot-levels. Also, the instance of each first-order branch is provided,</p> <p>For each plot, three files are provided, i.e., "Plot_x.las", "Plot_x_Tree_Reference.xlsx" and "Plot_x_Branch_Reference.txt", where x means the x-th plot.<br>1) "Plot_x.las" is the data file, which includes the point coordinates and intensity, as well as tree-, classification-, and First-order branch IDs. The tree-, classification-, and First-order branch IDs are stored in the field of "Point Source ID", "Classification" and "GPS Time", respectively.</p> <p>2) "Plot_x_Tree_Reference.xlsx" includes the reference of the tree structure traits for each tree in the plot. The reference of each tree takes up one row. The tree-ID, position_x, position_y, tree height (m), DBH (m), First-order branch (m), Crown Projection area (m<sup>2</sup>), Crown Surface area (m<sup>2</sup>), Crown Volume (m<sup>3</sup>) are in the column 1 to 9, respectively.</p> <p>3) "Plot_x_Branch_Reference.txt" includes the reference of the First-order branch in the plot, including the tree ID, branch ID, the start and end point positions of each branch. The record of each individual First-order branch takes up one row, and the column 1 to 9 are tree-ID, First-order Branch ID, Start_x, Start_y, Start_z, End_x, End_y, End_z, and Length.</p> <p>4) The calculation of the reference of the tree structure traits can be found in <a href="https://doi.org/10.1080/10095020.2024.2313325">https://doi.org/10.1080/10095020.2024.2313325.</a></p> <p>5) For more details about the data, readers are referred to "Read me.pdf".</p> <p>If you used this dataset, please cite the following paper:</p> <p>Liang, Xinlian, Hanwen Qi, Xuejie Deng, Jianchang Chen, Shangshu Cai, Qingjun Zhang, Yunsheng Wang, Antero Kukko, and Juha Hyyppä. 2024. “ForestSemantic: A Dataset for Semantic Learning of Forest from Close-Range Sensing.” Geo-Spatial Information Science, March, 1–27. doi:10.1080/10095020.2024.2313325.</p>
Data for the Article: Cross-validation of a semantic segmentation network for natural history collection specimens
<p>This deposit contains six datasets which were used for testing and validating a semantic segmentation network. The purpose was to evaluate the suitability of the segmentation network for use in the processing of images from Natural History Collections.</p>
LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation
<p>The benchmark code is available at: <a href="https://github.com/Junjue-Wang/LoveDA">https://github.com/Junjue-Wang/LoveDA</a></p> <p><strong>Highlights: </strong></p> <ol> <li>5987 high spatial resolution (0.3 m) remote sensing images from Nanjing, Changzhou, and Wuhan</li> <li>Focus on different geographical environments between Urban and Rural</li> <li>Advance both semantic segmentation and domain adaptation tasks</li> <li>Three considerable challenges: multi-scale objects, complex background samples, and inconsistent class distributions</li> </ol> <p><strong>Reference:</strong></p> <pre><code>@inproceedings{wang2021loveda, title={Love{DA}: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation}, author={Junjue Wang and Zhuo Zheng and Ailong Ma and Xiaoyan Lu and Yanfei Zhong}, booktitle={Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks}, editor = {J. Vanschoren and S. Yeung}, year={2021}, volume = {1}, pages = {}, url={https://datasets-benchmarks proceedings.neurips.cc/paper/2021/file/4e732ced3463d06de0ca9a15b6153677-Paper-round2.pdf} }</code></pre> <p><strong>License:</strong></p> <p>The owners of the data and of the copyright on the data are RSIDEA, Wuhan University. Use of the Google Earth images must respect the "Google Earth" terms of use. All images and their associated annotations in LoveDA can be used for academic purposes only, <strong>but any commercial use is prohibited. (CC BY-NC-SA 4.0)</strong></p>
OMOP2OBO: Semantic Integration of Standardized Clinical Terminologies to Power Translational Digital Medicine Across Health Systems (Recorded Introduction)
<p>This entry contains the recorded introduction that was presented at the 2020 Observational Health Data Science Initiative Symposium (<a href="https://www.ohdsi.org/events/2020-ohdsi-symposium/">https://www.ohdsi.org/events/2020-ohdsi-symposium/</a>).</p>
Towards a reproducible interactome: semantic-based detection of redundancies to unify protein-protein interaction databases
<p>Protein-protein interactions (PPIs) play an ubiquitous and fundamental role in all biological processes. Information on PPIs described in the literature is annotated and made available by several protein-interaction databases. Because most databases have their own curation rules and priorities, they often annotate overlapping sets of publications, which leads to redundancies. We developed a semantic-based approach which enables to accurately detect redundancies within PPI datasets from multiple databases. We applied this approach to assemble a "reproducible interactome", with PPIs supported by at least two methods or publications.</p>
VC-SLAM Versatile Corpus for Semantic Labeling And Modeling
<p>Benchmark Corpus for semantic labeling and modeling.</p> <p>This corpus contains 101 data sets from different open data portals.<br> Each data set consists of the following data:</p> <ul> <li>Raw csv data [rawdata_csv]</li> <li>Large json data sample [json_sample_large]</li> <li>Small json data sample [json_sample_small]</li> <li>Raw data in csv format [rawdata_csv]</li> <li>Raw data samples in csv format [rawdata_csv_samples]</li> <li>Mappings to translate between csv and json files [csv_json_mappings]</li> <li>Textual description / Metadata [descriptions]</li> <li>Semantic model as rdf/ttl [semantic_models]</li> <li>Mappings describing mapping between raw data attributes and concepts from the ontology [mappings]</li> <li>List of attributes that have been ignored during modeling [ignored_attributes]</li> </ul> <p>Additionally the corpus contains a target ontology as rdf/ttl [ontology].</p> <p>The individual data sets are licensed by the licenses specified in the attached Excel sheet (DataSetOverview.xlsx)</p> <p> </p> <p>These data are provided "as is", without any warranties of any kind. The data are provided under the Creative Commons Attribution 4.0 International license.</p>
Data used to evaluate ORBITS: Optimal Repair-Based Inconsistency-Tolerant Semantics
<p>This dataset provides the input files that were used in the evaluation of the ORBITS system (Optimal Repair-Based Inconsistency-Tolerant Semantics, <a href="https://github.com/bourgaux/orbits">https://github.com/bourgaux/orbits</a>). A detailed description is available in a technical report on arXiv (<a href="https://arxiv.org/abs/2202.07980">https://arxiv.org/abs/2202.07980</a>).</p> <p><strong>Content:</strong></p> <p>Folders <em>cqapri_benchmark</em>, <em>food_inspection_benchmark</em>, and <em>physicians_benchmark</em> contain JSON files of conflict graphs and candidate queries and their causes.<br> These files are named using the following pattern: files of candidate answers and their causes are named <database>_<query>_answers_causes.json, and conflict graphs are named <database>_conflictGraph_<priority relation>.json where <priority relation> says whether the priority relation is score-structured (prio_score) or not (prio_non_score) and the probability (p<proba>) or number of scores (n<number>) used to build the priority relation.</p> <p>Folder <em>original_datasets_and_queries</em> contains the Food Inspection and Physicians datasets used to generate files from <em>food_inspection_benchmark</em> and <em>physicians_benchmark</em>.<br> Files from <em>cqapri_benchmark</em> have been generated from the CQAPri benchmark available at <a href="https://lahdak.lri.fr/CQAPri/CQAPri.php">https://lahdak.lri.fr/CQAPri/CQAPri.php</a>.<br> In all cases, we use ProvSQL (<a href="https://github.com/PierreSenellart/provsql">https://github.com/PierreSenellart/provsql</a>) to build conflict graphs and causes from the datasets.</p>
Metadata and metadata semantics
<p>The dataset has a study of international scientific panorama on term "semantic metadata" using bibliometric indicators and the positioning of Information Science in relation to this topic. Bibliographic research was used for the theoretical construction of the subject and as a method, bibliometric studies, in order to identify scientific production indicators on semantic metadata and their conceptual basis, as a way of showing how this theme is seen in the international scenario.</p>
SROADEX: Dataset for binary recognition and semantic segmentation of road surface areas from high resolution Aerial Orthoimages Covering Approximately 8,650 km2 of the Spanish Territory Tagged with Road Information
<p>The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography representing the axes of the different types of roads (urban, interurban and rural). This cartography has been obtained from different Spanish official sources (National Geographic Institute and autonomic cartographic agencies) that we have revised and edited in a meticulous and systematic way to verify that the roads are represented on the cartography according to the orthoimages, available on January 1, 2021 in the download center of the National Center of Geographic Information (CNIG), on 16 rectangular areas (28,5 km * 18,5 km) of the Spanish territory (insular and peninsular).</p> <p>The dataset consists of 777599 images in png format of 256x256 pixels, organized in folders for the different trainings, separating those corresponding to training, testing and validation.</p> <p>The structure of the data is as follows:<br> 1-Road-Ortho and 1-Road-Mask contain the images and ground true for training the semantic segmentation networks.<br> 1-Road-Ortho and 2-NoRoad-Ortho contain aerial images containing or not containing vials, for the training of binary tessellation networks identifying tessellations with vials.<br> Moreover, in each folder the structure is the same: train, test, validation containing 90%, 5% and 5% of the total images and masks of each type.</p> <p>1-Road-Ortho</p> <p> |----Train</p> <p> |----Test</p> <p> -----Validation</p> <p>1-Road-Mask</p> <p> |----Train</p> <p> |----Test</p> <p> -----Validation</p> <p>2-NoRoad-Ortho</p> <p> |----Train</p> <p> |----Test</p> <p> -----Validation</p> <p> </p>
[Stimulus Set] Evoking the N400 event-related potential (ERP) component using a publicly available novel set of sentences with semantically incongruent or congruent eggplants (endings)
<p>During speech comprehension, the ongoing context of a sentence is used to predict sentence outcome by limiting subsequent word likelihood. Neurophysiologically, violations of context-dependent predictions result in amplitude modulations of the N400 event-related potential (ERP) component. While N400 is widely used to measure semantic processing and integration, no publicly-available auditory stimulus set is available to standardize approaches across the field. Here, we developed an auditory stimulus set of 442 sentences that utilized the semantic anomaly paradigm, provided cloze probability for all stimuli, and was developed for both children and adults. With 20 neurotypical adults, we validated that this set elicits robust N400's, as well as two additional semantically-related ERP components: the recognition potential (~250 ms) and the late positivity component (~600 ms). This stimulus set (<a href="https://doi.org/10.5061/dryad.9ghx3ffkg">https://doi.org/10.5061/dryad.9ghx3ffkg</a>) and the 20 high-density (128-channel) electrophysiological datasets (<a href="https://doi.org/10.5061/dryad.6wwpzgmx4">https://doi.org/10.5061/dryad.6wwpzgmx4</a>) are made publicly available to promote data sharing and reuse. Future studies that use this stimulus set to investigate sentential semantic comprehension in both control and clinical populations may benefit from the increased comparability and reproducibility within this field of research.</p>
Dataset for generating LOD3 building models from structure-from-motion and semantic segmentation
<p>This repository contains the codes for computing geometrical digital twins as LOD3 models for buildings, using a structure from motion and semantic segmentation. The methodology hereby implements was presented in the paper [Generating LOD3 building models from structure-from-motion and semantic segmentation" by Pantoja-Rosero et., al. (2022)] (<a href="https://doi.org/10.1016/j.autcon.2022.104430">https://doi.org/10.1016/j.autcon.2022.104430</a>)</p>
A Dataset of Synthetic Images of Outdoor Scenes Taken from Sidewalks, for Temporal Semantic Segmentation Applications
<p>This dataset has been generated using the CARLA simulator (release 0.9.11), an open-source 3D simulator for experiments in autonomous vehicle, based on the Unreal Engine game engine. It comes with pre-made city environment maps. CARLA is distributed with several integrated maps as well as parameters to increase the variety in the dataset. In the release that we have used, there are 13 semantic segmentation classes: None, Building, Fence, Other, Pedestrian, Pole, Lane-marking, Road, Sidewalk, Vegetation, Vehicle, Wall, and Traffic sign. The "None" category corresponds to textures that are not part of an object, such as lawns which are not part of "Vegetation", or sky. In the “Other” category are found objects that are not included in the other classes like plant and flower pots. For smart mobility applications, the “Sidewalks” and “Road” classes are of particular importance to find the way forward, as well as “Buildings” and “Poles” for obstacle avoidance. Sequences are made of 4 images. The dataset is composed of 46436 frames (11609 sequences) partitioned in 41024 frames (10256 sequences) for train, 2696 frames (674 sequences) for validation, and 2716 for test (679 sequences). The size of the images is 800 x 600 (resp. width x height).</p> <p>Additionaly, we have generated another smaller dataset with images taken from 2 different viewpoints: one located on the road and the other located on the sidewalk. The number of frames for train/validation/test is respectively 7288 (1822 sequences) partitioned in 6344 (1687 sequences) for train, 416 frames (104 sequences) for validation, and 424 for test (106 sequences). This smaller dataset is aimed at showing the importance of the viewpoint in the result of semantic segmentation. This can be done by cross-validation: learning on images taken from a viewpoint located on the road and test on images with a viewpoint located on the sidewalk, and vice versa.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.