Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
483
datasets available to search
ShareScore release 0.9.0
Dataset results
483 results for “SEMANTICS”
Figure 3. Graph created directly from the results.-Browsing Semantic Data in Slovakia
<p>Resulting graph is rather complex. There are 175 vertices and 201 edges, which were created directly, containing 158 persons and 17 companies. As we see on Figure 3, some filtering methods are required in order to create suitable overview of relations. Figure 4 thus shows the visualization of the same graph, but with tens of vertices merged. Now it contains 22 persons and 17 companies. A merging was performed for clarification and is built inside the visualization module. A simple condition says that a merging is performed if persons are unique, thus if a person is connected only to 1 firm. More formally, person vertices are merged, if their vertex degree equals 1 (each).</p>
Figure 2. Results for firm name "Váhostav" are in table. Each row defines a firm with its name, identification number and address. Then a connection is specified (whether it be a person or another firm).-Browsing Semantic Data in Slovakia
<p>We have searched for firm “Váhostav”, which is a rather big firm in Slovakia, with many press articles published about31. On Figure 2 there is a browsing window, for SBR data results, displaying tabular structure, which was refined from SBR dataset by continuous querying.</p>
Figure 1. Schema of clientside application for semantic browsing.-Browsing Semantic Data in Slovakia
<p>With the aim primary on unstructured information extraction and refining, relationship discovery and visualization, we propose our solution for SBR in the first place. The reason for this is, primary, that HTML formatted results of SBR are very jerky and uncertainty regarding the structure of information is very high. Readers can also be pointed by J. Suchal and P. Vojtek (2009), that care should be taken towards type errors. We discuss that later. In this work, we try to fill&up the gap of visualization and, somehow limited data access offered by SBR, adapting to the problems disclaimed above. We suggest a new client& side paradigm, which does not depend on a particular website like foaf.sk. Figure 1 describes the schema briefly and the key elements are parsers with other tools on the top and structured formats, for datastore, on the bottom.</p>
Datasets of the paper "Retrieval-Mediated Directed Forgetting in the Item-Method Paradigm: The Effect of Semantic Cues"
<p>Data sets and analyses scripts of the paper "Retrieval-Mediated Directed Forgetting in the Item-Method Paradigm: The Effect of Semantic Cues".</p>
FN-RE: A Corpus of Requirements Documents Enriched with Semantic Frame Annotations
<p>FN-RE is a human-labelled dataset using FrameNet scheme. The dataset is distributed and can be viewed using a web-index page. For further details about the annotation procedures, please refer to the annotation guidelines included in the folder.</p>
Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - model weights
<p>In the related <a href="https://github.com/HybridNLP2018/tutorial">notebook </a>we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains the model weights trained on such large corpora.</p>
Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - images
<p>In this notebook we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains high quality versions of the images used in the analysis.</p>
Semantic Frame Embeddings for Detecting Relations between Software Requirements
<p><strong>FN-RE frame embeddings-is semantic resource built based on embedding-based representations of semantic frames in FrameNet, which was developed to support the detection of relations between software requirements. Our embeddings, which encapsulate contextual information at the semantic frame level, were trained on a large corpus of requirements (i.e., a collection of more than three million mobile application reviews). </strong></p>
Don't count, predict! Semantic vectors
<p>Semantic vectors associated with the paper "<a href="https://www.aclweb.org/anthology/P14-1023">Don't count, predict! A systematic comparison of context-counting vs context-predicting semantics vectors</a>"</p> <p><strong>Abstract:</strong> context-predicting models (more commonly known as embeddings or neural language models) are the new kids on the distributional semantics block. Despite the buzz surrounding these models, the literature is still lacking a systematic comparison of the predictive models with classic, count-vector-based distributional semantic approaches. In this paper, we perform such an extensive evaluation, on a wide range of lexical semantics tasks and across many parameter settings. The results, to our own surprise, show that the buzz is fully justified, as the context-predicting models obtain a thorough and resounding victory against their count-based counterparts.</p>
Responses to "Semantic Web: Perspectives" Questionnaire
<p>The dataset provides material relating to a questionnaire entitled "Semantic Web: Perspectives". This questionnaire was addressed to the W3C Semantic Web mailing list (semantic-web@w3.org) and was open to responses from May 12th to May 25th, 2019. A total of 113 responses were collected in this time. The following files are provided:</p> <ul> <li><strong>public-comments.txt:</strong> provides the public comments of respondents in plain text;</li> <li><strong>questionnaire-form.pdf:</strong> illustrates the design of the questionnaire, including questions, types of responses permitted, etc.;</li> <li><strong>questionnaire-responses.tsv:</strong> lists the individual responses (without private comments) as a tab-separated values file;</li> <li><strong>success-keywords.xlsx:</strong> provides a spreadsheet mapping success story responses to a list of keywords, further providing statistics on these keywords;</li> <li><strong>wordcloud-bw.svg:</strong> provides a word-cloud of success-story keywords in black & white;</li> <li><strong>wordcloud-colour.svg:</strong> provides a word-cloud of success-story keywords in colour.</li> </ul> <p>The word-clouds were produced using <a href="https://www.jasondavies.com/wordcloud/">Jason Davies' online service</a>, copying and pasting the keywords from the <strong>success-keywords.xlsx</strong> spreadsheet (e.g., Column A, Sheet Statistics) into the text field; the following settings were selected: Orientations from 0° to 0°, Spiral: Rectangular; Scale: n; Number of words: 400; One word per line: ticked; Font: Patua One (must be installed locally beforehand). The resulting SVG files were later modified in a text editor to add a link to the font used, to tighten the bounding box, and to produce a black & white version.</p> <p>We thank the respondents for providing their input.</p>
FAIRsFAIR Data of Survey on Semantics and interoperability solutions
<p>As part of the EOSC project family the FAIRsFAIR - Fostering Fair Data Practices in Europe - project aims to supply practical solutions for the use of the FAIR data principles throughout the research data life cycle. The work package "WP2 FAIR Practices: Semantics, Interoperability, and Services" will produce three reports on FAIR requirements for persistence and interoperability to identify domain-specific standards and practices in use. These will review and document commonalities and possible gaps regarding semantic interoperability, and the use of metadata and persistent identifiers across infrastructures. They will also look into differences in terms of standards, vocabularies and ontologies. The collected information will be updated during the course of the project in cooperation with other tasks and EOSC projects.</p> <p>This survey was done to complement and validate the information from desk research for the first of these reports. It was aimed at data managers and data support experts. We hoped to get information about tools and services we might have missed, but also some reflections on the thinking around identifiers and ontologies and other semantic artefacts. The information was also collected to support preparing workshops on semantics and interoperability that are forthcoming in the project, as well as the work on software and services. The survey covers questions about metadata, use of persistent identifiers, use of semantic artefacts and handling research software.</p> <p>The survey was conducted as a joint effort with WP3, FAIR Policy and Practice and its open consultation, and was disseminated on the fairsfair.eu web pages, social media channels and via email lists. We received 66 answers during the period the survey was open, that is between 15 July to 2 October 2019.</p>
Coarse-Grained Sense Inventories Based on Semantic Matching between English Dictionaries
<p><strong>Abstract</strong> (our paper)</p> <p>WordNet is one of the largest handcrafted concept dictionaries visualizing word connections through semantic relationships. It is widely used as a word sense inventory in natural language processing tasks. However, WordNet's fine-grained senses have been criticized for limiting its usability. In this paper, we semantically match sense definitions from Cambridge dictionaries and WordNet and develop new coarse-grained sense inventories. We verify the effectiveness of our inventories by comparing their semantic coherences with that of Coarse Sense Inventory. The advantages of the proposed inventories include their low dependency on large-scale resources, better aggregation of closely related senses, CEFR-level assignments, and ease of expansion and improvement. Our inventories are publicly available for free use.</p> <p><strong>Publication</strong></p> <p>These datasets are part of our research results. If you make use of our datasets, please cite:</p> <ul> <li>Masato Kikuchi, Masatsugu Ono, Toshioki Soga, Tetsu Tanabe, Tadachika Ozono. Coarse-Grained Sense Inventories Based on Semantic Matching between English Dictionaries. In <em>Proceedings of the 11th International Conference on Advanced Informatics: Concepts, Theory and Applications (ICAICTA 2024)</em>. 6 pages, 2024.</li> </ul>
A semantic segmentation dataset of Arctic sea ice from Operation IceBridge data
<p>This dataset is a semantic segmentation dataset of Arctic sea ice based on deep learning method from Operation IceBridge images. It contains 29,372 labeled images, each of which corresponds to an image of Operation IceBridge and are stored in TIFF format. The dataset can be accessed using ArcGIS, ENVI, and the GDAL library in Python easily. Where label 1 represents melt ponds, label 2 represents sea ice/snow, label 3 represents submerged ice, and label 4 represents open ocean water, respectively.</p>
Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms Dataset
<p>This is the official data repository for the paper: "Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms", available at:<a href="https://www.arxiv.org/abs/2409.19371"> https://www.arxiv.org/abs/2409.19371</a>. The corresponding code is available at: <a href="https://github.com/david-stojanovski/EDMLX">https://github.com/david-stojanovski/echo_from_noise</a></p> <p> </p> <p>The synthetic data is produced using a variety of generative architectures, including the <strong>Elucidating Diffusion Model (EDM), Variance Exploding (VE), Variance Preserving (VP)</strong>, and our novel models, <strong>EDM-L64</strong> and <strong>EDM-L128</strong>, which employ <strong>latent diffusion</strong> strategies to significantly reduce computational cost. By incorporating <strong>spatially adaptive normalization (SPADE) blocks</strong> and <strong>Γ-distribution-based Variational Autoencoders (Γ-VAE)</strong>, these datasets ensure that the generated images preserve the essential semantic features required for training deep learning models.</p> <p> </p> <p>All pretrained classification and segmentation models can be found within the <strong>trained_models </strong>file.</p> <p>All generated images can be found within the <strong>generated_data </strong>file. Included is the <strong>CAMUS</strong> and original <strong>Semantic Diffusion Model (SDM) </strong>data, as well as a folder labelled <strong>easy_inference</strong> designed to contain all relevant labelmaps in a convenient folder for generating replicas of the dataset (detailed at codebase).</p>
Semantic Sea-Ice Classification for Belgica Bank in Greenland
<p>Each Sentinel-1 image is tiled into patches of 256x256 pixels. The size of the images is different and we reduced to the smallest one. In total for each image are 6,400 patches [1-2, 4]. See the excel file for the 24 Sentinel-1 ids.</p> <p>The semantic classes are:</p> <ol> <li><em>Black border</em></li> <li><em>Old ice</em></li> <li><em>First-Year ice </em></li> <li><em>Glaciers </em></li> <li><em>Icebergs</em></li> <li><em>Mountains </em></li> <li><em>Young ice</em></li> <li><em>Water group</em></li> </ol> <p>The last class combines the <em>Floating ice, Water body, Water ice current and melted snow</em> defined in [3] because they have very similar physical properties.</p> <p> </p> <p><strong>References:</strong></p> <p>1. C.O. Dumitru, G. Schwarz, C. Karmakar, and M. Datcu, “Machine Learning-Based Paradigm for Boosting the Semantic Annotation of EO Images”, IGARSS, Belgium, July 2021, pp. 1-4.</p> <p>2. C. Karmakar, C.O. Dumitru, and M. Datcu, “Explainable AI for SAR Image Time Series: Knowledge Extraction for Polar Areas”, MDPI Remote Sensing Journal, 2021, pp. 1-21 (under review).</p> <p>3. C.O. Dumitru, V. Andrei, G. Schwarz, and M. Datcu, “Machine Learning for Sea Ice Monitoring from Satellites”, Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sci., XLII-2/W16, pp. 83-89, 2019.</p> <p>4. C. Karmakar, C.O. Dumitru, G. Schwarz, and M. Datcu, “<em>Feature-Free Explainable Data Mining in SAR Images Using Latent Dirichlet Allocation</em>”, IEEE JSTARS, vol. 14, pp. 676-689, 2021.</p>
SemEval-2021 Task 10: Source-Free Domain Adaptation for Semantic Processing
<p>Data sharing restrictions are common in NLP datasets. For example, Twitter policies do not allow sharing of tweet text, though tweet IDs may be shared. The situation is even more common in clinical NLP, where patient health information must be protected, and annotations over health text, when released at all, often require the signing of complex data use agreements. The SemEval-2021 Task 10 framework asks participants to develop semantic annotation systems in the face of data sharing constraints. A participant's goal is to develop an accurate system for a target domain when annotations exist for a related domain but cannot be distributed. Instead of annotated training data, participants are given a model trained on the annotations. Then, given unlabeled target domain data, they are asked to make predictions.</p> <p>Website: <a href="https://machine-learning-for-medical-language.github.io/source-free-domain-adaptation/">https://machine-learning-for-medical-language.github.io/source-free-domain-adaptation/</a></p> <p>CodaLab site: <a href="https://competitions.codalab.org/competitions/26152">https://competitions.codalab.org/competitions/26152</a></p> <p>Github repository: <a href="https://github.com/Machine-Learning-for-Medical-Language/source-free-domain-adaptation">https://github.com/Machine-Learning-for-Medical-Language/source-free-domain-adaptation</a></p>
Supplementary materials accompanying "Baring the bones: the lexico-semantic association of bone with strength in Melanesia and the study of colexification"
<p>Supplementary materials accompanying "Baring the bones: the lexico-semantic association of bone with strength in Melanesia and the study of colexification". Two appendices. Appendix I: Melanesian languages with associations of ‘bone’ with strength. Appendix II: Colexifications of ‘bone’ in databases.</p>
Dataset for semantic segmentation in NDT with step-heating thermography for CFRP laminates
<p>Dataset composed of 36 images (640x480 pixels) of 30 bands or channels from step-heating for the same Carbon Fiber Reinforced Polymer (CFRP) laminate.</p> <p>In each image the specimen is rotated 10º to generate new data, with different illumination, background, and heating/cooling sequences.</p> <p>Images are generated from a video composed of the heating process, which takes 10 seconds, and the cooling process, which takes another 10 seconds, to a total of 1000 frames per video. This video is then reduced to 30 images using different processing methods explained below.</p> <p>The 30 bands or channels consist of the following bands repeated for the heating and cooling processes:</p> <ol> <li>First PCT component</li> <li>Second PCT component</li> <li>Third PCT component</li> <li>Fourth PCT component</li> <li>PPT</li> <li>Kurtosis</li> <li>Skewness</li> <li>TSR Coefficient 7</li> <li>TSR Coefficient 6</li> <li>TSR Coefficient 5</li> <li>TSR Coefficient 4</li> <li>TSR Coefficient 3</li> <li>TSR Coefficient 2</li> <li>TSR Coefficient 1</li> <li>TSR Coefficient 0</li> </ol> <p>Please cite the original paper:</p> <p>LINK: TODO</p> <p>BibTex:</p> <p>TODO</p> <p> </p> <p> </p>
Synset-based dataset built by using semantic-based feature reduction techniques.
<p>Dataset was built by applying an enhanced semantic-based feature selection method over to publicly available datasets: SpamAssassin (SA.zip) and Youtube Comments (YT.zip).</p>
NIVA Common Semantic Model
<p>NIVA has developed under Enterprise Architect a UML data model for some pieces of IACS data or processes, namely core geographic data, EO monitoring and Farm Registry. The NIVA model may be found under: EU Common Agricultural Model /Conceptual model. More detailed explanations about content of this model may be found in NIVA deliverable D3.2 Common Semantic Model, available on the NIVA web site : <a href="https://www.niva4cap.eu/deliverables/">Deliverables – Niva4cap</a> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.