Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
483
datasets available to search
ShareScore release 0.9.0
Dataset results
483 results for “SEMANTICS”
Semantic object-scene inconsistencies affect eye movements, but not in the way predicted by contextualized meaning maps - data
<p>Data from the article<strong><em> Semantic object-scene inconsistencies affect eye movements, but not in the way predicted by contextualized meaning maps</em></strong> published in Journal of Vision.</p> <p>code: https://zenodo.org/record/5999215<br> data: https://zenodo.org/record/5999046</p> <p><br> Marek A. Pedziwiatr<br> marek.pedziwi@gmail.com<br> February 2022</p>
Evaluating a Semantic-based Automated Approach to Task-Relevant Text Identification: Supporting Material
<p>This is the supplementary material for the evaluation of a semantic-based automated approach to task-relevant text identification.</p> <p> </p> <p><strong>Where:</strong></p> <ul> <li><strong>ds-python:</strong> contains the dataset produced as part of this study</li> <li><strong>output:</strong> contains the results from our data analysis</li> <li><strong>responses:</strong> contains the participants' responses to each of the major questions we asked them during our experiment</li> </ul>
Aesthetic Trends and Semantic Web Adoption of Media Outlets Identified through Automated Archival Data Extraction
<p>This dataset includes a variety of structured data gathered via various Web data extraction techniques which were employed in order to collect current and archival data from almost a thousand news websites that are popular in Greece, for the purpose of monitoring and recording their progress through time. The collected information, that took the form of a website’s source code and an impression of their homepage in different time instances of the last decade, has been used to identify trends concerning Semantic Web integration, DOM structure complexity, number of graphics, color usage and more. In total more than ten thousands impressions (including screenshots and source code) were analyzed which resulted to conclusions regarding the evolution of aesthetics and the adoption of new technologies.</p>
Association Rules and Semantic Relatedness (ARSR) - Evaluation Data
<p><em><strong>ARSR</strong></em> stands for "<strong><em>Association Rules and Semantic Relatedness</em></strong>",<br> a recommender system that combines association rule mining and semantic relatedness to create recommendations for form fields.</p> <p>ARSR was evaluated with focus on recommending values for fields of metadata forms.<br> The evaluation was performed with two sets of association rules (R1 and R2).<br> <strong>R1</strong> was primarily used to assess the recommendation performance.<br> <strong>R2</strong> served as an alternative and led to similar results.</p> <p>This dataset contains the raw data ("<strong>collected</strong>", JSON format), generated during the evaluation.<br> It includes i.a. input (populated fields and target field), expected output, and the top 40 generated recommendations for each test combination.</p> <p>In addition, the processed data ("<strong>analysed</strong>", CSV format) is provided.<br> It is based on the raw data and is used to calculate metrics and plot the results.</p> <p>The <strong>source code</strong> for ARSR and the evaluation is available on <a href="https://gitlab.com/dlr-dw/arsr">GitLab (gitlab.com)</a>.</p> <p> </p> <p><em>Note:</em><br> To perform the analysis on the raw data (collected) yourself, make sure to follow the setup instruction in the evaluation repository first.<br> More specifically, install dependencies and unzip "data/cedar/test-instances.zip" so that URI mappings can be accessed.<br> Then follow the instructions provided by the README file in the raw data archives (e.g. collected-R1.zip).</p>
Semantic Breaking Issue Detector
<p>Semantic Breaking Issue Detector</p>
Dataset for paper "Automated Static Warning Identification via Path-based Semantic Representation" submitted to JOS
<p>The project includes the datasets and running example used in the submitted JOS paper titled "Automated Static Warning Identification via Path-based Semantic Representation"</p>
Figure 5 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 5. The Wallich Catalogue. Screenshot of Wallich Catalogue hosted by Royal Botanic Garden Edinburgh showing popup for stable URI containing information hosted at Botanic Garden and Botanical Museum Berlin-Dahlem.
Figure 4 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 4. The CETAF Specimen URI Tester provides for any given Specimen URI an overview of the redirection process as well as a preview of machine-readable and human-readable data associated with the URI.
Figure 3 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 3. CETAF stable HTTP URIs in the GBIF data portal. The Global Biodiversity Information Facility (GBIF) publishes CETAF stable HTTP URIs via their data portal.
Figure 2 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects
Figure 2. Basic redirection mechanisms. Human users are redirected to a human-readable web-representation of the specimens. Software systems are re-directed to a machine-readable metadata record.
Overfitting in semantics-based program repair
<p>Data for our emse journal paper</p>
Choice-dependent delta-band neural trajectory during semantic category decision making in the human brain
<p>The dataset and code provided correspond to Manuscript Number: ISCIENCE-D-23-09408R1, titled "Choice-dependent delta-band neural trajectory during semantic category decision making in the human brain."</p> <p>This dataset comprises delta and alpha filtered data from a total of 19 participants. Each .mat file contains 800 cells, representing the number of trials. Each cell contains a matrix of size 128 x 1750. Here, 128 denotes the number of EEG channels, and 1750 represents the number of time points. Time points are sampled from -1 s to 2.5 s relative to the onset of the first stimulus, with a 2ms interval.<br><br></p> <p>The participants' behavioral data are stored in separate .mat files for each run (or block). Upon loading these files, a struct named "data" is loaded, containing six variables, each representing a 1 x 200 vector:</p> <ol> <li> <p>Cat1: Represents the category of the first stimulus. It takes a value of 1 for animate words and 2 for inanimate words.</p> </li> <li> <p>Cat2: Denotes the category of the second stimulus. It is assigned 1 for animate words and 2 for inanimate words.</p> </li> <li> <p>Same: Indicates the correct response for each trial. A value of 1 signifies a match between the categories of Stimulus 1 and Stimulus 2, while 2 indicates a non-match.</p> </li> <li> <p>Resp: Records the participant's decision for each trial. A value of 1 denotes a match between the categories of Stimulus 1 and Stimulus 2, whereas 2 represents a non-match.</p> </li> <li> <p>RT: Represents the participant's response time in seconds for each trial.</p> </li> <li> <p>Corr: Indicates the correctness of the participant's response. A value of 1 signifies a correct response, while 0 indicates an incorrect response.</p> </li> </ol> <p> </p>
Figures - Semantic analysis of web archive historical data 1983 "Marche pour l'égalité et contre le racisme"
Open the record for dataset details and reuse information.
SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages
<p>This is the GitHub repository hosting data and code for SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages. For additional information about the task, please consult the official website.</p>
QSage: Structural and Semantic Metric Analysis for Quantum Code Smell Detection
Open the record for dataset details and reuse information.
Data for "Improving semantic video retrieval models by training with a relevance-aware online mining strategy"
<p>This repository contains all the data available for the publication:</p> <p><a href="https://doi.org/10.1016/j.cviu.2024.104035">Alex Falcon, Giuseppe Serra, and Oswald Lanz. <em>Improving semantic video retrieval models by training with a relevance-aware online mining strategy</em>. <strong>Computer Vision and Image Understanding</strong>. 2024.</a></p> <p>Code is available at: <a href="https://github.com/aranciokov/ranp/">https://github.com/aranciokov/ranp/</a></p> <p>The data includes:</p> <ul> <li>pre-extracted features (ordered_feature_*.zip files)</li> <li>annotations, such as pre-extracted semantic graphs, glove checkpoints, class annotations, etc (annotations_*.zip files)</li> <li>train/val/test, when available, split information (public_split_*.zip) files</li> <li>pretrained models for HGR and EAO (details in the github repo)</li> </ul>
Portal frame bridge video with semantic predictions
<p>This is a video representing how a 3D point cloud acquisition is made for a portal frame bridge. The predictions of a semantic segmentation model that are inferred during the point cloud acquisition are shown on the side. When the scanning of the bridge is complete, the 3D point cloud automatically includes the semantic label per 3D point, thanks to the semantic segmentation model. </p>
Semantic fields and Castilianization in Galician: A comparative study with the Loanword Typology project (dataset)
<p>Dataset containing the data analyzed in "Semantic fields and Castilianization in Galician: A comparative study with the Loanword Typology project" (Fernández Rei, Elisa & Regueira Fernández, Xosé Luís (eds.), <em>New Developments in Galician Linguistics, </em>special issue of Languages) (forthcoming) https://www.mdpi.com/journal/languages/special_issues/8109M79HP8</p> <p> </p>
Embeddings for the paper "The Zapatista Semantic Struggle: Analysing the Linguistic Innovation of the EZLN with Semantic Difference Keywords (SDKs)"
<p>Embeddings from word2vec models described in "The Zapatista Semantic Struggle: Analysing the Linguistic Innovation of the EZLN with Semantic Difference Keywords (SDKs)". Full reference TBD.</p>
Dataset for "Learning Scene Semantics from Vehicle-centric Data for City-scale Digital Twins", Fürntratt et al.
<p>Dataset for "Learning Scene Semantics from Vehicle-centric Data for City-scale Digital Twins", Fürntratt et al.</p> <p>Data are anonymized and provided with segmentation mask ground truths. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.