Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

483

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

483 results for “SEMANTICS”

Learn how ShareScore rates datasets ↗
zenodo40/100

Experimental Datasets and Processing Codes for the Semantic PHD Filter

<p>The water bottle detection dataset and measurement model dataset for the paper titled &quot;The Semantic PHD Filter for Multi-class Target Tracking: From Theory to Practice&quot; by Jun Chen, Zhanteng Xie and Philip Dames, and the paper titled &quot;Experimental Datasets and Processing Codes for the Semantic PHD Filter&quot; by&nbsp;Zhanteng Xie,&nbsp;Jun Chen and Philip Dames</p> <p><strong>1. Detection dataset:&nbsp;&nbsp;</strong></p> <p>Size:&nbsp;<br> Total: 4870 images<br> Training: 4000 images<br> Validation: 870 images</p> <p>Bottle Classes: Aquafina, Deer, Kirkland, Nestle</p> <p>Format: PASCAL VOC, Darknet</p> <p>Folder Structure:<br> &nbsp;- Annotations: containing the xml label files in PASCAL VOC format<br> &nbsp;- ImageSets: containing the training index files&nbsp;<br> &nbsp;- JPEGImages: containing the image data in jpg format<br> &nbsp;- Labels: containing the txt label files in Darknet format</p> <p><strong>2. Measurement model dataset:</strong></p> <p>Format: ROSBAG</p> <p>Duration: 19:59s (1199s)</p> <p>Topics:<br> /darknet_ros/detection_image 3543 msgs : sensor_msgs/Image<br> /map 1 msg : nav_msgs/OccupancyGrid<br> /sphd_measurements 3585 msgs : sphd_msgs/SPHDMeasurements<br> /tf 142727 msgs : tf2_msgs/TFMessage<br> /tf_static 1 msg : tf2_msgs/TFMessage</p> <p>Message Types:<br> nav_msgs/OccupancyGrid<br> sensor_msgs/Image<br> sphd_msgs/SPHDMeasurements<br> tf2_msgs/TFMessage</p> <p>&nbsp;</p> <p><strong>3. Processing codes:</strong></p> <p>Detection processing:<br> Zenodo:&nbsp;https://doi.org/10.5281/zenodo.7066045<br> GitHub: https://github.com/TempleRAIL/yolov3_bottle_detector<br> <br> Measurement model processing:<br> Zenodo:&nbsp;&nbsp;https://doi.org/10.5281/zenodo.7066050<br> GitHub: https://github.com/TempleRAIL/sphd_sensor_models</p>

opencc-by-4.0Sep 2022View details →
dryad40/100

Evoking the N400 Event-Related Potential (ERP) component using a publicly available novel set of sentences with semantically incongruent or congruent eggplants (endings)

<p class="MsoNoSpacing"><span>During speech comprehension, the ongoing context of a sentence is used to predict sentence outcome by limiting subsequent word likelihood. Neurophysiologically, violations of context-dependent predictions result in amplitude modulations of the N400 event-related potential (ERP) component. While N400 is widely used to measure semantic processing and integration, </span><span>no publicly-available auditory stimulus set is available to standardize approaches across the field. Here, we developed an auditory stimulus set of 442 sentences that utilized the semantic anomaly paradigm, provided cloze probability for all stimuli, and was developed for both children and adults. With 20 neurotypical adults, we validated that this set elicits robust N400's, as well as two additional semantically-related ERP components: the recognition potential (~250 ms) and the late positivity component (~600 ms). This stimulus set (<a href="https://doi.org/10.5061/dryad.9ghx3ffkg">https://doi.org/10.5061/dryad.9ghx3ffkg</a>) and the 20 high-density (128-channel) electrophysiological datasets (<a href="https://doi.org/10.5061/dryad.6wwpzgmx4">https://doi.org/10.5061/dryad.6wwpzgmx4</a>) </span><span>are made publicly available to promote data sharing and reuse. Future studies that use this stimulus set to investigate sentential semantic comprehension in both control and clinical populations may benefit from the increased comparability and reproducibility within this field of research.</span></p>

opencc-zeroOct 2022View details →
zenodo40/100

Figure 1: The ECG model-MAPPING BETWEEN SEMANTIC GRAPHS AND SENTENCES IN GRAMMAR INDUCTION SYSTEM

<p>The following Figure 1 shows a sample semantic graph that describes a<br> simple test world.<br> During the processing of the ECG, the base units of the graph are the ECG<br> atoms. An ECG atom corresponds to a primitive statements related to one<br> predicate. It has a structure of one-level deep tree, where the root of the tree<br> is the predicate and the concepts linked to it are the leaves. The child concept<br> of the root predicate may be not only a single concept but it can be another<br> ECG atom.</p>

opencc-by-4.0Jun 2010View details →
zenodo40/100

Figure 11. Comparison of MultiNet and MST representations-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>In the sentence (7) there are two clauses describing two hypothetical situations. These two<br> situations are connected by the relation COND in MultiNet representation (Figure 11).</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 10. Comparison of MultiNet and MST representations.-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>Helbig [18, p. 507] defines the expression (s MCONT c) as &ldquo;a specification of the<br> informational or mental content c of a mental or informational process s&hellip; By default, the second<br> argument c is assumed to be a hypothetical object or situation.&rdquo;<br> MCONT relation properties are roughly equivalent to those of subject-verb complexes<br> (SVSBs) of MST. MCONT relation can generally be treated as capturing the idea behind what<br> philosophers call propositional attitudes in representational theories of mind [12] or opaque<br> contexts [10] in semantics. For example, consider the sentence in (3) rewritten below as (6) and<br> semantic representation of which is given in Figure 10.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 8. Comparison of MultiNet and MST representations-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>Figure 8 shows the representation of the<br> first reading of the sentence (5) in MWR. The MultiNet is shown at top of the figure, and the corresponding mental space representation is shown at the bottom side. Blue circles mark<br> hypothetical or nonreal objects and situations and red circles mark real objects and situations.<br> Facticity value for John is real, for unicorn it is nonreal, and for the process of riding, marked with<br> blue broken circle, it is non-real as well.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 7. The mental spaces set up by the sentence If John buys the car, he will drive to Berlin.-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>In the sentence (4) there are three proper names which constitute elements of the base space<br> or reality space (R). The CNSB if sets up the hypothetical space (H) with elements identical to those<br> of reality space (Figure 7). Every space&rsquo;s internal structure is presented in the boxes next to them.<br> In the next part, we will see how MultiNet&rsquo;s built-in meaning representation mechanisms are<br> capable of representing the basic principles of mental space building outlined above.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 6. The mental spaces set up by the sentence Mary thinks that John smokes.-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>The proper nouns Mary and John setup a base space (B). By the help of background<br> knowledge and activated frames we know that they are names of female and male humans. Not<br> having access to the previous discourse, we also consider their existence presupposed. The SVSB<br> Mary thinks that sets up a belief space (L) relative to space B (Figure 6). The identity connector<br> maintains the referential link between elements a and a ׳ both referring to the same person.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 9. Comparison of MultiNet and MST representations-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>Second interpretation of the sentence (5) can be represented with changing John&rsquo;s Facticity<br> attribute-value from [FACT = real] to [FACT = nonreal] making the whole situation and its<br> elements non-real (Figure 9).</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 5. The mental spaces set up by the second interpretation of the sentence in the film, John is riding a unicorn.-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>Second interpretation: The proper name John exists only in film space without having a<br> counterpart in base space. Therefore, the second interpretation of the sentence is: John is a<br> film character who is riding a unicorn in the film (Figure 5).</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 4. The mental spaces set up by the first interpretation of the sentence in the film, John is riding a unicorn-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>First interpretation: The reality space (let&rsquo;s call it R) contains an element a associated with<br> the proper name John. The noun phrase a unicorn introduces an element b ׳ to the film space<br> (call it F). I is the connector linking a in the space B to a ׳ in the space F (Figure 4). Since the<br> elements of both mental spaces are co-referential, this connector is an identity connector.<br> The rectangles represent the internal structure of the spaces next to them. The dashed line<br> indicates that the space F is set up in relation to R and that it is subordinate to R in<br> discourse.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 1. The upper ontology of sorts in MultiNet (after Helbig [17])-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>One of the distinguishing features of MultiNet is its commitment to the Cognitive Adequacy<br> requirement. According this requirement [9], semantic representations and knowledge<br> representations should be centered around concepts. Concepts2 are represented by nodes in the<br> graphical representation of the network. Every node belongs to a specific sort defined by the<br> MultiNet&rsquo;s ontology of sorts (Figure 2).</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 1. Mental space representation of "In the play, Mary is excited"-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>Thus an entity can have a variable reality status depending on the mental<br> space to which it belongs. The mental space constructed by the sentence Poirot is a Belgian<br> detective is a non-real imaginary story space of which Poirot is an element. But when we say in<br> reality, Poirot is not Belgian the constructed space is reality space in which Poirot (the actor, not<br> the character) does not have the fictional nationality.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 2. Semantic representation of the sentence Peter finished the discussion in MultiNet after Helbig [8, p. 447].-Representing Mental Spaces and Dynamics of Natural Language Semantics

<p>In Figure 3, semantic frame of the concept Finish realized in the form of the verb finish<br> requires two C-roles: An agent represented by the relation AGT, and an affected entity represented<br> by the relation AFF. Here agent is Peter and the affected entity is an abstract object ([SORT = ad]<br> means the concept is a dynamic abstraction).</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

SemTab 24: Semantic Table Annotations Benchmark for LLM-based approaches

<p><strong>SuperSemtab24 </strong>is a dataset for tabular data to knowledge graph matching.</p> <p>The dataset is divided into training and validation sets. The dataset includes general-purpose tables and intentionally misspelled entities to evaluate the model's robustness. Participants must annotate the entity mentions in the validation set and submit their annotations (following a target file).</p> <p>The repository contains the full version of the dataset; the ground truth (GT) of the test set will be uploaded in the future.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Train and Evaluation Code, Road Classification Models and Test set of the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography"

<p>This repository contains the Python scripts built for training and evaluation of the implementation, together with the test data and the resulting road segmentation models corresponding to the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography". The scripts make use of the Tensorflow with Keras framework and their additional required dependencies.</p> <p>The training and validation set is based on the binary SROADEX dataset (<a href="../records/6482346">https://zenodo.org/records/6482346</a>) that was re-split into tiles that feature the image resolutions (256 x 256, 512 x 512, and 1024 x 1024 pixels) and image overlaps (0% and 12.5%) considered in this study. The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography that represents the axes of the different types of roads (urban, interurban and rural). This binary road data contains information from 16 full orthoimages (28.5 km * 18.5 km) with spatial resolution of 0.5 m/pixel from the insular and peninsular Spanish territory. Due to the size on disk of approximately 492 gigabytes, this training and validation data is only available upon request from the corresponding author. The test set has been generated from a novel area from Palencia (Spain) and features 18 million pixels labelled with the positive "Road" class. The test sets are provided in the repository for each resolution (with no overlap), so that additional DL models can be evaluated on the same data and compared with the results achieved in this study.</p> <p>The structure of the information shared in this repository is as follows:<br>The scripts have been grouped by tile resolution (256, 512 and 1024). First, the test set and the evaluation script can be found. For each tile resolution, there are two subfolders (corresponding to the "no overlap" and "12.5% overlap"). In each case, the Python scripts for training the models in the three repetitions are shared, and the trained models (H5 format) are shared in compressed form. Finally, for each resolution we also share the testing dataset which consists of two folders.</p> <p>The material is distributed under a CC-BY 4.0 license.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Lexical Semantic Change Cause-Type-Definitions Benchmark

<p>The Lexical Semantic Change Cause-Type-Definitions (LSC-CTD) Benchmark is a digitised dataset that builds on and extends the Blank's seminal 1997 taxonomy of semantic change. This collection categorises 657 instances of linguistic evolution across the vocabulary of the Romance languages, with additional entries of German and English instances. Each entry is accompanied by a new pair (Old and New Meaning) of english definitions, manually curated by a historical linguist.</p> <p>The dataset includes a detailed classification of causes of change such as semantic wear, lexical gap, orphaned word, lexical complexity, atypical actant structure, frame, socio-cultural change, abstract concept, atypical part of speech, new concept, taboo, expressivity and prototype. It also includes types of semantic shift as classified by Blank, i.e. specialisation, generalisation, co-hyponymous transfer, auto-antonym, metaphor, antiphrasis, metonymy, auto-converse, ellipsis, folk etymology, analogy, meaning dilution, meaning reinforcement and doubtful cases.&nbsp;</p> <p><br><strong>Reference</strong></p> <p>The accompanying paper where this resource is described in detail will be published at ACL 2024.<br><br><span>Pierluigi Cassotti, Stefano De Pascale, and Nina Tahmasebi. 2024. <a href="https://aclanthology.org/2024.acl-long.249">Using Synchronic Definitions and Semantic Relations to Classify Semantic Change Types</a>. In <em>Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</em>, pages 4539&ndash;4553, Bangkok, Thailand. Association for Computational Linguistics.</span></p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Dataset of KO journal paper: Semantic analysis of archival concepts in CIDOC-CRM and in RiC-CM and RiC-O

<p>Semantic analysis of archival concepts (class, relations, atributtes, and relation attributes) presents in Records in Context family (conceptual model and ontology) and its possible equivalents in CIDOC-CRM.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

M4.4 (associated data) - FAIR-IMPACT Review of Semantic Artefact Catalogues and technologies

<p>This dataset (version 1) takes the form of a spreadsheet corresponding to the associated data described by&nbsp;<a href="../records/12799796" target="_blank" rel="noopener"><strong>M4.4 - Review of Semantic Artefact Catalogues and guidelines for serving FAIR semantic artefacts in EOSC</strong></a></p> <p>The spreadsheept (available here in ODS format) contains the listing of Semantic Artefact Catalogues (SACs) done within FAIR-IMPACT's WP4, their classifications (by status, type, discipline and technology) and the evaluation of their FAIR-enabling dimensions.&nbsp;</p> <p>A "live" version of this spreadsheet is available as an open Google Sheet open for comments and suggestions. We will take external contributions and comments into consideration when producing new verion of this dataset. Contributions could be of several types:&nbsp;</p> <ul> <li>New SAC or modification of the ones currently identified;</li> <li>New SAC technology or modification of the ones currently identified;</li> <li>New FAIR-enabling assessment or modification of the ones currently available.</li> </ul> <p>For more information, interested parties may contact the authors.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Figure 4. After merging, overview is more transparent. Tens of persons were merged together into clusters in order to clarify the visualization. Firms and persons are recognized based on their icons.-Browsing Semantic Data in Slovakia

<p>The usefulness of such visualization has its key points regarding connections. Thanks to SBR browsing module, we were able to get 22 firm records for &ldquo;V&aacute;hostav&rdquo; query. Between any 2 companies, connections may be (and often are) not bidirectional, so, in order to navigate through connections, we have refined all 22 records. Although, even being filtered, graph is still complex. And it is possible to further navigate and search for outgoing connections, for example firm &ldquo;MERLIN TRADE, a.s.&rdquo; on Fig.4 contains item on &ldquo;J&aacute;n Kato&rdquo;, which is already included in our graph and connected to &ldquo;V&Aacute;HOSTAV&amp;SK&amp;DEVELOPEMENT&rdquo; on bottom left side and &ldquo;V&Aacute;HOSTAV&amp;SK, a.s.&rdquo; in the center. Edge coloring and drawing is helpful with overlapped edges. For methods of visualization, including coloring, we refer to studies of H. Omote and K. Sugiyama (2006), and I. Herman, G. Melanon, and M. S. Marshall (2000) &nbsp;or our study on graph clutter filtering and connectivity distance (Mojzis &amp; Laclavik, 2014).</p>

opencc-by-4.0Nov 2015View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record