Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
483
datasets available to search
ShareScore release 0.9.0
Dataset results
483 results for “SEMANTICS”
Pairwise Multi-Class Document Classification for Semantic Relations between Wikipedia Articles (Dataset, Models & Code)
<p>Many digital libraries recommend literature to their users considering the similarity between a query document and their repository. However, they often fail to distinguish what is the relationship that makes two documents alike. In this paper, we model the problem of finding the relationship between two documents as a pairwise document classification task. To find the semantic relation between documents, we apply a series of techniques, such as GloVe, Paragraph-Vectors, BERT, and XLNet under different configurations (e.g., sequence length, vector concatenation scheme), including a Siamese architecture for the Transformer-based systems. We perform our experiments on a newly proposed dataset of 32,168 Wikipedia article pairs and Wikidata properties that define the semantic document relations. Our results show vanilla BERT as the best performing system with an F1-score of 0.93,<br> which we manually examine to better understand its applicability to other domains. Our findings suggest that classifying semantic relations between documents is a solvable task and motivates the development of recommender systems based on the evaluated techniques. The discussions in this paper serve as first steps in the exploration of documents through SPARQL-like queries such that one could find documents that are similar in one aspect but dissimilar in another.</p> <p>Additional information can be found on <a href="https://github.com/malteos/semantic-document-relations/">GitHub</a>.</p> <p>The following data is supplemental to the experiments described in our research paper. The data consists of:</p> <ul> <li>Datasets (articles, class labels, cross-validation splits)</li> <li>Pretrained models (Transformers, GloVe, Doc2vec)</li> <li>Model output (prediction) for the best performing models</li> </ul> <p><strong>Dataset</strong></p> <p>The Wikipedia article corpus is available in <code>enwiki-20191101-pages-articles.weighted.10k.jsonl.bz2</code>. The original data have been downloaded as <a href="https://dumps.wikimedia.org/enwiki/">XML dump</a>, and the corresponding articles were extracted as plain-text with <a href="https://radimrehurek.com/gensim/scripts/segment_wiki.html">gensim.scripts.segment_wiki</a>. The archive contains only articles that are available in training or test data.</p> <p>The actual dataset is provided as used in the stratified k-fold with <code>k=4</code> in <code>train_testdata__4folds.tar.gz</code>.</p> <pre><code>├── 1 │ ├── test.csv │ └── train.csv ├── 2 │ ├── test.csv │ └── train.csv ├── 3 │ ├── test.csv │ └── train.csv └── 4 ├── test.csv └── train.csv 4 directories, 8 files </code></pre> <p>Pretrained models</p> <p>PyTorch: vanilla and Siamese BERT + XLNet</p> <p>Pretrained model for each fold is available in the corresponding model archives:</p> <pre><code># Vanilla model_wiki.bert_base__joint__seq512.tar.gz model_wiki.xlnet_base__joint__seq512.tar.gz # Siamese model_wiki.bert_base__siamese__seq512__4d.tar.gz model_wiki.xlnet_base__siamese__seq512__4d.tar.gz </code></pre>
A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains
<p><strong>Content</strong></p> <p>This repository contains pre-trained computer vision models, data labels, and images used in the pre-print publication "A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains":</p> <ol> <li><em>ADPdevkit</em>: a folder containing the 50 validation ("tuning") set and 50 evaluation ("segtest") set of images from the Atlas of Digital Pathology database formatted in the VOC2012 style--the full database of 17,668 images is available for download from the original website</li> <li><em>VOCdevkit</em>: a folder containing the relevant files for the PASCAL VOC2012 Segmentation dataset, with both the trainaug and test sets</li> <li><em>DGdevkit</em>: a folder containing the 803 test images of the DeepGlobe Land Cover challenge dataset formatted in the VOC2012 style</li> <li><em>cues</em>: a folder containing the pre-generated weak cues for ADP, VOC2012, and DeepGlobe datasets, as required for the SEC and DSRG methods</li> <li><em>models_cnn</em>: a folder containing the pre-trained CNN models</li> <li><em>models_wsss</em>: a folder containing the pre-trained SEC, DSRG, and IRNet models, along with dense CRF settings</li> </ol> <p><strong>More information</strong></p> <p>For more information, please refer to the following article. <strong>Please cite this article when using the data set.</strong></p> <p>@misc{chan2019comprehensive,<br> title={A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains},<br> author={Lyndon Chan and Mahdi S. Hosseini and Konstantinos N. Plataniotis},<br> year={2019},<br> eprint={1912.11186},<br> archivePrefix={arXiv},<br> primaryClass={cs.CV}<br> }</p> <p>For the full code released on GitHub, please visit the repository at: <a href="https://github.com/lyndonchan/wsss-analysis">https://github.com/lyndonchan/wsss-analysis</a></p> <p><strong>Contact</strong></p> <p>For questions, please contact:<br> Lyndon Chan<br> lyndon.chan@mail.utoronto.ca<br> http://orcid.org/0000-0002-1185-7961</p>
Semantic Web und Linked Data: Generierung von Interoperabilität in archäologischen Fachdaten am Beispiel römischer Töpferstempel - Datasets
<p><strong>Datasets</strong></p> <p>Gegenstand der Masterarbeit ist die Verwendung aktueller Technologien interoperabler Datenhaltung, insbesondere das Konzept der Linked Open Data (LOD) und der semantischen Modellierung, zur Verdeutlichung ihres Potentials in archäologischen Informationen am Beispiel von Terra Sigillata-Fundorten, -Töpfern und -Keramikfragmenten. Die Arbeit zeigt eine Migration von Daten, sowie die Möglichkeiten und die Problematik der Modellierung der Attribute und Beziehungen mit Hilfe bestehender LOD-Konzepte und kontrollierter Vokabularien, sowie eigene Ansätze zur Lösung. Diese Daten werden mittels REST-Schnittstelle zur Verfügung gestellt. Ein Schwerpunkt wird auf die Verlinkung zu anderen bereits bestehenden Projekten gelegt, wodurch eine Vielzahl weiterer archäologischer und historischer Informationen z.B. über das Pelagios Projekt eingebunden werden. Zudem wird das Potential der Verlinkung und Abfrage von heterogenen Informationen zwischen Töpfern, Fragmenten und Orten deren relativ chronologische Beziehungen über LOD mit einer webbasierten Schnittstelle aufgezeigt.</p> <p>The subject matter of this master thesis is using current technologies in interoperable data management, in particular the illustration of the potential of Linked Open Data (LOD) and semantic modelling in archaeological information, as used on samian ware places and their corresponding potters and ceramic fragments. The thesis demonstrates a migration of data as well as possibilities and problems of modelling attributes and relationships using existing LOD concepts and controlled vocabularies as well as novel self-developed approaches to the solution. These data are provided by a ReST interface. One focus is linking to other existing projects, creating associations to other archaeological and historical information, for example the Pelagios project. Moreover, a web-based interface shows the potential of linking and retrieval of heterogeneous information among pottery, fragments and places and their relative chronological relationships via LOD.</p>
SiriusGeoOnto: an ontology tailored to the semantic annotation of geological images
<p>We have designed and implemented the <strong>SiriusGeoOnto</strong> ontology to cover the information embedded in the geological images. <strong>SiriusGeoOnto</strong> has been modelled in the OWL 2 ontology language using the ontology editor Protégé. <strong>SiriusGeoOnto </strong>is currently integrated with the <strong>SiriusGeoAnnotator. </strong></p> <p><strong>SiriusGeoAnnotator</strong> is a system that generates annotations in the form of a knowledge graph and drives the annotation process according to <strong>SiriusGeoOnto</strong> and the previously generated annotations.</p> <p><strong>SiriusGeoAnnotator: </strong><a href="https://sws.ifi.uio.no/project/sirius-geo-annotator/">https://sws.ifi.uio.no/project/sirius-geo-annotator/</a></p> <p> </p> <p> </p> <p> </p> <p> </p>
Do Synthesis Centers Synthesize? A Semantic Analysis of Topical Diversity in Research
<p>Synthesis centers are a form of scientific organization that catalyzes and supports research that integrates diverse theories, methods and data across spatial or temporal scales to increase the generality, parsimony, applicability, or empirical soundness of scientific explanations. Synthesis working groups are a distinctive form of scientific collaboration that produce consequential, high-impact publications. But no one has asked if synthesis working groups synthesize: are their publications substantially more diverse than others, and if so, in what ways and with what effect? We investigate these questions by using Latent Dirichlet Analysis to compare the topical diversity of papers published by synthesis center collaborations with that of papers in a reference corpus. Topical diversity was operationalized and measured in several ways, both to reflect aggregate diversity and to emphasize particular aspects of diversity (such as variety, evenness, and balance). Synthesis center publications have greater topical variety and evenness, but less disparity, than do papers in the reference corpus. The influence of synthesis center origins on aspects of diversity is only partly mediated by the size and heterogeneity of collaborations: when taking into account the numbers of authors, distinct institutions, and references, synthesis center origins retain a significant direct effect on diversity measures. Controlling for the size and heterogeneity of collaborative groups, synthesis center origins and diversity measures significantly influence the visibility of publications, as indicated by citation measures. We conclude by suggesting social processes within collaborations that might account for the observed effects, by inviting further exploration of what this novel textual analysis approach might reveal about interdisciplinary research, and by offering some practical implications of our results.</p>
Semantic Representation of the Joconde Database
<p>Semantic Representation of the Joconde Database, based on CIDOC-CRM.</p> <p>Near from 600000 creative works described. More than 11M triples.</p> <p>Source for Joconde Database:</p> <p><a href="https://www.data.gouv.fr/fr/datasets/collections-des-musees-de-france-extrait-de-la-base-joconde/">https://www.data.gouv.fr/fr/datasets/collections-des-musees-de-france-extrait-de-la-base-joconde/</a></p> <p>(json format)</p> <p>More information: <a href="https://github.com/datamusee/semjoconde">https://github.com/datamusee/semjoconde</a></p> <p> </p>
A controlled vocabulary defining the semantic perimeter of Sustainable Development Goals
<p>A set of controlled terms that define the scope and breadth of <a href="https://sustainabledevelopment.un.org/">Sustainable Development Goals (SDGs) as defined by the United Nations</a>. These terms may be used to tag and index textual records in accordance with SDGs.</p> <p>The vocabulary is constructed by means of the following steps:</p> <ol> <li>An initial set of terms per SDG target is built by extracting key terms from the UN official list of Goals, Targets and Indicators</li> <li>The list is manually enriched by performing a review of the literature produced around SDGs and by compiling lists of pertinent words per Target mentioned by the reviewed documents</li> <li>A reference textual corpus is downloaded by searching for the initial set terms defined at step 1. and 2. The corpus is used to train a Word2Vec word embedding model (a machine learning model based on neural networks).</li> <li>The terms’ list is then enriched by means of automatic methods, which are run in parallel: <ul> <li>The trained Word2Vec model is used to select, among the indexed keywords of the reference corpus, all terms “semantically close” to the initial set of words. This step is carried out to select terms that might not appear in the texts themselves, but that were deemed pertinent to label the textual records.</li> <li>Further terms that are mentioned in the texts of the reference corpus and that are valued by the trained Word2Vec model as “semantically close” to the initial set of words are also retained. This step is performed to include in the controlled vocabulary a series of terms that are related to the focus of the SDGs and which are used by practitioners.</li> <li>An automated algorithm is used to retrieve, from the APIs of WikiPedia a series of terms that have some categorical relationships (i.e. those that are indexed as “a broader concept of”, or “equivalent to” in DBpedia) with the initial set of words.</li> </ul> </li> <li>The final list produced by steps 1-4 s finally manually revised</li> </ol>
Paper-author bipartite graph from Semantic Scholar
<p>Paper-author bipartite graph created from <a href="https://api.semanticscholar.org/corpus">Semantic Scholar's Open Research Corpus</a>, version 2018-05-03. Vertices are papers (39,219,709 of them) in one part and authors (12,862,455 of them) in the other. A paper is connected to all its co-authors, and an author is connected to all the papers they wrote, leading to 139,268,795 edges. A citation count (the number of times the paper was cited) is available for each paper (from 0 to 37,230 citations per paper).</p>
A process proposal on how to move from data requirements to semantic data in the context of the MDA
<p>This work proposes a process for the identification of the requirements associated with the data, along with and a set of transformations in Model Driven Architecture (MDA) context, in order to obtain a semantically annotated dataset, as a result of the unification and alignment of the data in the context of its initial domain. Our proposal identifies four phases (from CIM to code), in which is describe the artifacts and transformations required to progress to the next phase: a target domain model is first obtained from the data requirements, after which the ontological schema and the ontology is generated from the previous model. A domain-specific language (DSL), also proposed in this work, is then used to obtain the semantic data model (the DSL code), which generates the final semantic dataset. We have validated the proposal by studying two cases: one with data from the public transport domain and the other with data concerning those affected by the COVID-19 pandemic.</p>
The role of semantics in the perceptual organization of shape
<p>Dataset relative to the following publication:</p> <p>Schmidt, F., Kleis, J., Morgenstern, Y., & Fleming, R. W. (2020). The role of semantics in the perceptual organization of shape. <em>Scientific Reports, 10</em>, 22141. <a href="https://doi.org/10.1038/s41598-020-79072-w">https://doi.org/10.1038/s41598-020-79072-w</a></p> <p>Each experiment folder contains the data relative to one experiment and a Matlab script for plotting the data, together with a text file with comments. The stimuli folder contains image files of all experimental stimuli and corresponding copyright information.</p>
Swiss3DCities: Aerial Photogrammetric 3D Pointcloud Dataset with Semantic Labels
<p>We introduce a new outdoor urban 3D pointcloud dataset, covering a total area of 2.7 km<sup>2</sup>, sampled from three Swiss cities with different characteristics. The dataset is manually annotated for semantic segmentation with per-point labels, and is built using photogrammetry from images acquired by multirotors equipped with high-resolution cameras. In contrast to datasets acquired with ground LiDAR sensors, the resulting point clouds are uniformly dense and complete, and are useful to disparate applications, including autonomous driving, gaming, smart city planning, and robotics.</p>
Reproducibility and robustness of graph measures of the Associative-Semantic Network
<p>Matfiles and matlab scripts used to study the reproducibility and robustness of graph measures of the Associative-Semantic Network.</p>
SAUCE (Semantic Annotated University Campus Environment)
<p>The dataset is composed of a set of 30 manually-annotated images with a 640x480 resolution. Frames were acquired by a Bumblebee2 stereo rig sensor, which is mounted on the forepart of the iCab research platform.</p> <p>Four different classes are used for labelling the dataset. The labelled categories correspond to the most popular instances found around the university campus and compose the minimum set required for in-campus navigation. The categories are:</p> <ul> <li>Traversable area (blue: RGB(0, 0, 255))</li> <li>Garden (green: RGB(0, 255, 0))</li> <li>Obstacles (red: RGB(255, 0, 0))</li> <li>Pedestrian (yellow: RGB(255, 246, 0))</li> </ul> <p> </p>
Large-scale semantic indexing of Spanish biomedical literature using contrastive transfer learning
Open the record for dataset details and reuse information.
Effect of Semantically Equivalent Embeddings on Generalizations
<p>Reproducibility package for the paper </p><blockquote><p>Effect of Semantically Equivalent Embeddings on Generalizations<br>Francesco Bertolotti and Walter Cazzola </p></blockquote><p>currently submitted to IEEE Transactions on Neural Networks and Learning Systems</p>
DepthMars Dataset for Semantic Segmentation of the Martian Surface from Rover Images
<p>This dataset is resulted from a research article, "DepthFormer: Depth-Enhanced Transformer Network for Semantic Segmentation of the Martian Surface from Rover Images", which includes surface images on Mars collected by the Zhurong rover along its traverse, depth images generated from stereo images, and corresponding manually labeled images.</p>
Appendix 5 Semantic prosody of ser+PP
<p>Appendix 5 of the paper "<span>Between source language constructions and target language expectations. </span>An analysis of passive constructions in translated and non-translated Spanish", published in <em>Review of Cognitive Linguistics.</em></p> <p>It contains an additional analysis of the semantic prosody of the verbs used with Spanish ser+PP. </p>
Decoding Knowledge Claims: the Evaluation of Scientific Publication Contributions through Semantic Analysis
<p>This data were used to compute the RWMD distance as described in the study submitted for the STI 2024 conference, Berlin.</p>
Processed data for the "Deriving Semantics-Aware Fuzzers from Web API Schemas" paper
<p>Processed data for the "Deriving Semantics-Aware Fuzzers from Web API Schemas" paper. Each directory in the archive consists of:</p> <p>- metadata.json. Metadata about a test run - tested fuzzer name, run duration, etc</p> <p>- fuzzer.json - Structured fuzzer output</p> <p>- deduplicated_cases.json - Deduplicated reported failures, when fuzzers provide it</p> <p>- sentry.json - Cleaned Sentry events for this run</p> <p>- target.json - Parsed stdout for Gitlab & Disease.sh targets that were tested without Sentry integration</p>
Materials in Vessels Dataset, Annotated images of materials in transparent vessels for semantic segmentation
<p> Data set of materials in vessels<br> The handling of materials in glassware vessels is the main task in chemistry laboratory research as well as a large number of other activities. Visual recognition of the physical phase of the<br> materials is essential for many methods ranging from a simple task such as fill-level evaluation to the<br> identification of more complex properties such as solvation, precipitation, crystallization and phase<br> separation. To help train neural nets for this task, a new data set was created. The data set contains a<br> thousand images of materials, in different phases and involved in different chemical processes, in a<br> laboratory setting. Each pixel in each image is labeled according to several layers of classification, as<br> given below:</p> <p>a. Vessel/Background: For each pixel assign value of one if it is part of the vessel and zero otherwise.<br> This annotation was used as the ROI map for the valve filter method.</p> <p>b. Filled/Empty: This is similar to the above, but also distinguishes between the filled and empty<br> regions of the vessel. For each pixel, one of the following three values is assigned:0 (background); 1<br> (empty vessel); or 2 (filled vessel).</p> <p>c. Phase type: This is similar to the above but distinguishes between liquid and solid regions of the<br> filled vessel. For each pixel, one of the following four values: 0 (background); 1 (empty vessel); 2<br> (liquid); or 3 (solid).</p> <p>d. Fine-grained physical phase type: This is similar to the above but distinguishes between specific<br> classes of physical phase. For each pixel, one of 15 values is assigned: 1 (background); 2 (empty<br> vessel); 3 (liquid); 4 (liquid phase two, in the case where more than one phase of the liquid appears in<br> the vessel); 5 (suspension); 6 (emulsion); 7 (foam); 8 (solid); 9 (gel); 10 (powder); 11 (granular); 12<br> (bulk); 13 (solid-liquid mixture); 14 (solid phase two, in the case where more than one phase of solid<br> exists in the vessel): and 15 (vapor).<br> The annotations are given as images of the size of the original image, where the pixel value is the<br> class number. The annotation of the vessel region (a) is used in the ROI input for the valve filter net .</p> <p>4.1. Validation/testing set<br> The data set is divided into training and testing sets. The testing set is itself divided into two subsets;<br> one contains images extracted from the same YouTube channels as the training set, and therefore was<br> taken under similar conditions as the training images. The second subset contains images extracted<br> from YouTube channels not included in the training set, and hence contains images taken under<br> different conditions from those used to train the net.</p> <p>4.2. Creating the data set<br> The creation of a large number of images with a variety of chemical processes and settings could have<br> been a daunting task. Luckily, several YouTube channels dedicated to chemical experiments exist<br> which offer high-quality footage of chemistry experiments. Thanks to these channels, including<br> NurdRage, NileRed, ChemPlayer, it was possible to collect a large number of high-quality images in a<br> short time. Pixel-wise annotation of these images was another challenging task, and was performed by<br> Alexandra Emanuel and Mor Bismuth.</p> <p>For more details see: <a href="https://arxiv.org/pdf/1708.08711.pdf">Setting attention region for convolutional neural networks using region selective features, for recognition of materials within glass vessels</a></p> <p>This dataset was first published in 2017.8</p> <p>For newer and Bigger datasets see</p> <p>https://zenodo.org/record/4736111#.YbG-RrtyZH4</p> <p>https://zenodo.org/record/3697452#.YbG-TLtyZH4</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.