Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
45
datasets available to search
ShareScore release 0.9.0
Dataset results
45 results for “Semantic annotations”
Langmark: annotations for scenes with semantic inconsistencies connecting distributional semantic models to vision science – data and code
<p>Data (including object annotations) and code from the following manuscript:</p> <p><em>Langmark: annotations for scenes with semantic inconsistencies connecting distributional semantic models to vision science</em>.</p>
Annotated dataset for the semantic segmentation of radishes
<p>This folder contains pictures of radishes collected on the PMF experimental field during Spring 2017. There are two kinds of labeled images in the following folders:</p> <p><br> - human annotations: human annotators draw polygons around each plant and those were then refined using an active contours algorithm.<br> - machine annotations: An SVM trained on the human annotations was used to produce labeled images. Images with bad segmentation were manually discarded.</p> <p>Each of these folder contains an images folder containing original pictures and a labels folder containing binary segmentation masks.</p>
Natural Language Annotations for Reasoning about Program Semantics
<p>Natural language annotations about Python statements</p> <p>The dataset is made of pairs of files sharing the prefix of the filename</p> <p>* Annotations are in JSONL format (filenames ending with '_annot'), i.e. JSON objects separated by newline ('\n') characters</p> <p>* Reference source code files are in JSON format. (filenames ending with '_code')</p> <p> </p> <p>Source dataset : Programming Puzzles (Schuster et al. 2021, NeurIPS Dataset and benchmarks track) - MIT License - https://github.com/microsoft/PythonProgrammingPuzzles</p>
SiriusGeoOnto: an ontology tailored to the semantic annotation of geological images
<p>We have designed and implemented the <strong>SiriusGeoOnto</strong> ontology to cover the information embedded in the geological images. <strong>SiriusGeoOnto</strong> has been modelled in the OWL 2 ontology language using the ontology editor Protégé. <strong>SiriusGeoOnto </strong>is currently integrated with the <strong>SiriusGeoAnnotator. </strong></p> <p><strong>SiriusGeoAnnotator</strong> is a system that generates annotations in the form of a knowledge graph and drives the annotation process according to <strong>SiriusGeoOnto</strong> and the previously generated annotations.</p> <p><strong>SiriusGeoAnnotator: </strong><a href="https://sws.ifi.uio.no/project/sirius-geo-annotator/">https://sws.ifi.uio.no/project/sirius-geo-annotator/</a></p> <p> </p> <p> </p> <p> </p> <p> </p>
SAUCE (Semantic Annotated University Campus Environment)
<p>The dataset is composed of a set of 30 manually-annotated images with a 640x480 resolution. Frames were acquired by a Bumblebee2 stereo rig sensor, which is mounted on the forepart of the iCab research platform.</p> <p>Four different classes are used for labelling the dataset. The labelled categories correspond to the most popular instances found around the university campus and compose the minimum set required for in-campus navigation. The categories are:</p> <ul> <li>Traversable area (blue: RGB(0, 0, 255))</li> <li>Garden (green: RGB(0, 255, 0))</li> <li>Obstacles (red: RGB(255, 0, 0))</li> <li>Pedestrian (yellow: RGB(255, 246, 0))</li> </ul> <p> </p>
Materials in Vessels Dataset, Annotated images of materials in transparent vessels for semantic segmentation
<p> Data set of materials in vessels<br> The handling of materials in glassware vessels is the main task in chemistry laboratory research as well as a large number of other activities. Visual recognition of the physical phase of the<br> materials is essential for many methods ranging from a simple task such as fill-level evaluation to the<br> identification of more complex properties such as solvation, precipitation, crystallization and phase<br> separation. To help train neural nets for this task, a new data set was created. The data set contains a<br> thousand images of materials, in different phases and involved in different chemical processes, in a<br> laboratory setting. Each pixel in each image is labeled according to several layers of classification, as<br> given below:</p> <p>a. Vessel/Background: For each pixel assign value of one if it is part of the vessel and zero otherwise.<br> This annotation was used as the ROI map for the valve filter method.</p> <p>b. Filled/Empty: This is similar to the above, but also distinguishes between the filled and empty<br> regions of the vessel. For each pixel, one of the following three values is assigned:0 (background); 1<br> (empty vessel); or 2 (filled vessel).</p> <p>c. Phase type: This is similar to the above but distinguishes between liquid and solid regions of the<br> filled vessel. For each pixel, one of the following four values: 0 (background); 1 (empty vessel); 2<br> (liquid); or 3 (solid).</p> <p>d. Fine-grained physical phase type: This is similar to the above but distinguishes between specific<br> classes of physical phase. For each pixel, one of 15 values is assigned: 1 (background); 2 (empty<br> vessel); 3 (liquid); 4 (liquid phase two, in the case where more than one phase of the liquid appears in<br> the vessel); 5 (suspension); 6 (emulsion); 7 (foam); 8 (solid); 9 (gel); 10 (powder); 11 (granular); 12<br> (bulk); 13 (solid-liquid mixture); 14 (solid phase two, in the case where more than one phase of solid<br> exists in the vessel): and 15 (vapor).<br> The annotations are given as images of the size of the original image, where the pixel value is the<br> class number. The annotation of the vessel region (a) is used in the ROI input for the valve filter net .</p> <p>4.1. Validation/testing set<br> The data set is divided into training and testing sets. The testing set is itself divided into two subsets;<br> one contains images extracted from the same YouTube channels as the training set, and therefore was<br> taken under similar conditions as the training images. The second subset contains images extracted<br> from YouTube channels not included in the training set, and hence contains images taken under<br> different conditions from those used to train the net.</p> <p>4.2. Creating the data set<br> The creation of a large number of images with a variety of chemical processes and settings could have<br> been a daunting task. Luckily, several YouTube channels dedicated to chemical experiments exist<br> which offer high-quality footage of chemistry experiments. Thanks to these channels, including<br> NurdRage, NileRed, ChemPlayer, it was possible to collect a large number of high-quality images in a<br> short time. Pixel-wise annotation of these images was another challenging task, and was performed by<br> Alexandra Emanuel and Mor Bismuth.</p> <p>For more details see: <a href="https://arxiv.org/pdf/1708.08711.pdf">Setting attention region for convolutional neural networks using region selective features, for recognition of materials within glass vessels</a></p> <p>This dataset was first published in 2017.8</p> <p>For newer and Bigger datasets see</p> <p>https://zenodo.org/record/4736111#.YbG-RrtyZH4</p> <p>https://zenodo.org/record/3697452#.YbG-TLtyZH4</p> <p> </p>
Subset of AmazonQA annotated with answerability and semantic & syntactic embedding
<p>3755 Instances taken from the review-question dataset AmazonQA. The file data_answerability contain the 3755 instances annotated with answerability, answer tag (aligned with the passage) and answer text. The file data_annotated contains 1818 answerable instances, annotated with embedding constructions in the answer text. Inventory_phrases and inventory_simple_words contain the expressions of logical operator, implicative and factive predicates that can be used in embedding annotation. annotate.py is the script for embedding annotation. </p>
Graphical and Collaborative Annotation Support for Semantic Web Services
<p><strong>Graphical and Collaborative Annotation Support for Semantic Web Services presentation</strong></p>
Data from: A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation
The Neotropical evaniid genus Evaniscus Szépligeti currently includes six species. Two new species are described, Evaniscus lansdownei Mullins, sp. n. from Colombia and Brazil and Evaniscus rafaeli Kawada, sp. n. from Brazil. Evaniscus sulcigenis Roman, syn. n., is synonymized under Evaniscus rufithorax Enderlein. An identification key to species of Evaniscus is provided. Thirty-five parsimony informative morphological characters are analyzed for six ingroup and four outgroup taxa. A topology resulting in a monophyletic Evaniscus is presented with Evaniscus tibialis and Evaniscus rafaeli as sister to the remaining Evaniscus species. The Hymenoptera Anatomy Ontology and other relevant biomedical ontologies are employed to create semantic phenotype statements in Entity-Quality (EQ) format for species descriptions. This approach is an early effort to formalize species descriptions and to make descriptive data available to other domains.
Semantic metadata annotation: tagging medline abstracts for enhanced information access
<p>The object of this study is to develop methods for automatically annotating the argumentative role of sentences in scientific abstracts. Working from Medline abstracts, we classified sentences into four major argumentative roles: objective, method, result, conclusion. The idea is that if the role of each sentence can be marked up, then this metadata can be used during information retrieval to seek for particular types of information such as novelty, conclusions, methodologies, aims/goals of a scientific piece of work.</p> <p> </p>
Data from: Evaluating active learning methods for annotating semantic predications
Objectives: This study evaluated and compared a variety of active learning strategies, including a novel strategy we proposed, as applied to the task of filtering incorrect SemRep semantic predications. Materials and Methods: We evaluated three types of active learning strategies – uncertainty, representative, and combined– on two datasets of semantic predications from SemMedDB covering the domains of substance interactions and clinical medicine, respectively. We also designed a novel combined strategy with dynamic β without hand-tuned hyperparameters. Each strategy was assessed by the Area under the Learning Curve (ALC) and the number of training examples required to achieve a target Area Under the ROC curve (AUC). We also visualized and compared the query patterns of the query strategies. Results: Combined strategies outperformed all other methods in terms of ALC, outperforming the baseline by over 0.05 ALC for both datasets and reducing 58% annotation efforts in the best case. While representative strategies performed well, their performance was matched or outperformed by the combined methods. All the uncertainty sampling methods beat the baseline but they were the worst performing methods overall. Our proposed AL method with dynamic β shows promising ability to achieve near-optimal performance across two datasets. Discussion: Our visual analysis of query patterns indicates that strategies which efficiently obtain a representative subsample perform better on this task. Conclusion: Active learning is shown to be effective at reducing annotation costs for filtering incorrect semantic predications from SemRep. Our proposed AL method demonstrated promising performance.
Figure 31 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figure 31 - Most parsimonious tree from exhaustive search in PAUP*. Numbers above nodes show bootstrap support and numbers below nodes show jackknife support from the maximum parsimony analysis.
Figures 13-18 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 13-18 - Brightfield images of Evaniscus rufithorax Enderlein. 13, 14 Lateral habitus 15, 16 Dorsal habitus 17 Anterior oblique 18 Anterior face.
Figures 7-12 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 7-12 - Brightfield images of Evaniscus rafaeli Kawada sp. n. 7, 8 Lateral habitus 9, 10 Dorsal habitus 11 Anterior oblique 12 Anterior face.
Figures 1-6 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 1-6 - Brightfield images of Evaniscus lansdownei Mullins sp. n. 1, 2 Lateral habitus 3, 4 Dorsal habitus 5 Anterior oblique 6 Anterior face.
Figures 25-30 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 25-30 - Brightfield images of Evaniscus tibialis Szépligeti. 25, 26 Lateral habitus 27, 28 Dorsal habitus 29 Anterior oblique 30 Anterior face.
Figures 32-33 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 32-33 - Brightfield images of Evaniscus rufithorax . 32 Male specimen; arrow points to visible lower tubular sclerite 33 Female specimen; arrow shows where lower tubular sclerite is not visible.
Figures 19-24 from: Mullins P, Kawada R, Balhoff J, Deans A (2012) A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation. ZooKeys 223: 1-38. https://doi.org/10.3897/zookeys.223.3572
Figures 19-24 - Brightfield images of Evaniscus marginatus Cameron. 19, 20 Lateral habitus 21, 22 Dorsal habitus 23 Anterior oblique 24 Anterior face.
Data from: A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation
Open the record for dataset details and reuse information.
Data from: Evaluating active learning methods for annotating semantic predications
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.