Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
483
datasets available to search
ShareScore release 0.9.0
Dataset results
483 results for “SEMANTICS”
Japan landslide dataset for semantic segmentation
<p>This database contains images used for the semantic segmentation of landslide scars from a fully convolutional neural network U-Net.</p> <p>1. <strong>Training dataset: </strong>it contains 125 GeoTIFF 8 bits images and associated PNG masks (scars indicated in white and background in black color).</p> <p>2. <strong>Validation dataset</strong>: it contains 10 GeoTIFF 8 bits images and associated PNG masks used for U-Net validation step.</p> <p>3. <strong>Test dataset: </strong>it contains 10 GeoTIFF 8 bits images and associated PNG masks for testing.</p> <p>Also, the "SHAPEFILES_LANDSLIDES.rar" file contains the vector layers of the masked images in .shp format.</p>
Investigating Distributional Robustness: Semantic Perturbations Using Generative Models (ImageNet Examples)
<p>This dataset contains examples of semantically-perturbed images, for NeurIPS 2020 submission #4915.</p> <p>There are four top-level folders, each containing results for semantic perturbations restricted to adjust the activation values at only certain layers of the BigGAN generative network: the first six layers, the middle six layers, the last six layers, and all layers.</p> <p>Within each top-level folder, there are a further four folders, each corresponding to a classifier neural network whose evaluation is being evaluated. These are EfficientNet-B4 with NoisyStudent training [1], the standard ResNet50 [2], a pixel-perturbation-robust ResNet50 trained by Engstrom et al. [3] and another trained by Wong et al. [4], using their "Fast is better than free" technique.</p> <p>Within each of these, there are many folders, named 'version_$N'. Each one of these contains three images: the unperturbed generated image, named unpert_generated_x_grid_0.png; the semantically-perturbed generated image, named generated_x_grid_0.png; and an image named semantic_pert_diffs_grid_0.png showing the pixel-space effect of the semantic perturbation, that is, the diff between the perturbed and unperturbed images. Note that if the perturbed and unperturbed images are identical, and the classifier misclassifies the unperturbed images, and so we skip this example.</p> <p>Along with the 'version_$N' folders containing the images, there exists a file for each classifier named results.json. Each top-level item in this JSON file corresponds to one 'version_$N' example. There are 5 attributes: 'label', indicating the target label of the unperturbed image; 'magnitude', which gives the magnitude of the semantic perturbation found; 'skipped_cla', which is 1 if the example is skipped because the classifier did not correctly classify the unperturbed image; 'skipped_judge', which is 1 if the human judged that the unperturbed image did not match its label, so this example is skipped; and 'pert_judgement', which is 1 if the semantically-perturbed image is judged by the human to be of the same class as the unperturbed image. These judgements on these images were used to construct the main graphs in the paper.</p> <p> </p> <p>[1] Qizhe Xie, Eduard H. Hovy, Minh-Thang Luong, and Quoc V. Le. Self-training with Noisy Student improves ImageNet classification. CoRR, abs/1911.04252, 2019. URL http://arxiv.org/abs/1911.04252.</p> <p>[2] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 770–778. IEEE Computer Society, 2016. doi: 10.1109/CVPR.2016.90. URL https://doi.org/10.1109/CVPR.2016.90.</p> <p>[3] Logan Engstrom, Andrew Ilyas, Shibani Santurkar, and Dimitris Tsipras. Robustness (Python library), 2019. URL ttps://github.com/MadryLab/robustness.</p> <p>[4] Eric Wong, Leslie Rice, and J. Zico Kolter. Fast is better than free: Revisiting adversarial training. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL<br> https://openreview.net/forum?id=BJx040EFvH.</p>
Semantic text mining in early drug discovery for type 2 diabetes
<p>BACKGROUND: Surveying the scientific literature is an important part of early drug discovery; and with the ever-increasing amount of biomedical publications it is imperative to focus on the most interesting articles. Here we present a project that highlights new understanding (e.g.\ recently discovered modes of action) and identifies potential novel drug target, via a novel, data-driven text mining approach to score type 2 diabetes (T2D) relevance. We focused on monitoring trends and jumps in T2D relevance to help us be timely informed of important breakthroughs.<br> <br> METHODS: We extracted over 7 million <em>n</em>-grams from PubMed and then clustered around 240,000 linked to T2D into almost 50,000 T2D relevant `semantic concepts'. To score papers, these concepts were weighted depending on co-mentioning with core T2D proteins. A protein's current T2D relevance was determined by combining the scores of the papers mentioning it in the preceeding five years. The significance of a jump in a protein's rank was assessed by comparing it to previously observed jumps.<br> <br> RESULTS: We show that T2D relevant papers, also those not mentioning T2D explicitly, got assigned high scores by mentioning semantic concepts often used in connection with T2D, as shown by the enrichment of well known T2D proteins among the top scoring proteins. Our `high jumpers' identified important past developments in the apprehension of how certain key proteins relate to T2D, indicating that our method will make us aware of future breakthroughs. In summary, this project facilitated keeping up with current T2D research by repeatedly providing short lists of potential novel targets into our early drug discovery pipeline.</p>
Amazon Rainforest dataset for semantic segmentation V2
<p>This database contains images used for training a fully convolutional neural network for the semantic segmentation of forested areas in images from the Sentinel-2 Level 2A Satellite.</p> <p>The images refer to the RGB composition (bands 4, 3 and 2). The histogram for each image was selected as follows:</p> <ul> <li>Band 4 (red): values ranging from 103 to 2724;</li> <li>Band 3 (green): values ranging from 194 to 2888;</li> <li>Band 2 (blue): values ranging from 99 to 2798.</li> </ul> <p>After that, each band was converted to a byte type (0-255).</p> <p>The images are still divided into three main sets: training, validation and testing:</p> <ol> <li><strong>Training dataset: </strong>it contains 1.123 GeoTIFF images with 512x512 pixels and associated PNG masks (clouds indicated in white and background in black color).</li> <li><strong>Validation dataset</strong>: it contains 100 GeoTIFF images with 512x512 pixels and associated PNG masks used for validation step.</li> <li><strong>Test dataset: </strong>it contains 100 GeoTIFF images 512x512 pixels for testing.</li> </ol>
Impact of Combining Syntactic and Semantic Similarities on Patch Prioritization
<p>This dataset contains 246 bugs, their fixes and corresponding buggy projects from historical bug fixes dataset (https://github.com/xuanbachle/data-bugfixes) that fulfill the following criteria:</p> <ol> <li>Unique</li> <li>Satisfy redundancy assumption at file level</li> <li>Fixed by applying replacement mutation</li> <li>Require fixing at expression level</li> <li>Having available project and dependency files</li> </ol> <p>For details, please view https://www.scitepress.org/Link.aspx?doi=10.5220/0009411301700180</p>
Graphical and Collaborative Annotation Support for Semantic Web Services
<p><strong>Graphical and Collaborative Annotation Support for Semantic Web Services presentation</strong></p>
Data from: A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation
The Neotropical evaniid genus Evaniscus Szépligeti currently includes six species. Two new species are described, Evaniscus lansdownei Mullins, sp. n. from Colombia and Brazil and Evaniscus rafaeli Kawada, sp. n. from Brazil. Evaniscus sulcigenis Roman, syn. n., is synonymized under Evaniscus rufithorax Enderlein. An identification key to species of Evaniscus is provided. Thirty-five parsimony informative morphological characters are analyzed for six ingroup and four outgroup taxa. A topology resulting in a monophyletic Evaniscus is presented with Evaniscus tibialis and Evaniscus rafaeli as sister to the remaining Evaniscus species. The Hymenoptera Anatomy Ontology and other relevant biomedical ontologies are employed to create semantic phenotype statements in Entity-Quality (EQ) format for species descriptions. This approach is an early effort to formalize species descriptions and to make descriptive data available to other domains.
Replication Package for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"
<p>Replication package for our analysis of Semantic Versioning in Ansible Galaxy role repositories.</p> <p>This replication package consists of three parts:</p> <ul> <li> <p>Classification Model: Contains Jupyter notebooks used to train and evaluate a Random Forest classification model based on structural features. Training and evaluation data is included.</p> </li> <li> <p>Quantitative Notebooks: Contains Jupyter notebooks used to perform quantitative analyses of versions and changes.</p> </li> <li> <p>data: CSV files of the data used in the Quantitative Notebooks, and the source data for the classification model. Should be downloaded separately fromthe classification model. Should be downloaded separately from <a href="https://doi.org/10.5281/zenodo.4991955">https://doi.org/10.5281/zenodo.4991955</a>.</p> </li> </ul> <p>The data is under the Creative Commons Attribution Share-Alike 4.0 license. The source code is under the GNU General Public License.</p>
Semantic Concordance rate data used to generate S2_Table10
<p>Semantic Concordance rate data used to generate S2_Table10 described in Appendix S2.</p>
THE PROBLEM OF SEMANTIC TAGGING OF PHILOSOPHICAL TERMS IN THE UZBEK LANGUAGE IN THE CORPUS
Open the record for dataset details and reuse information.
SOME CHARACTERISTICS OF WORD SEMANTIC SPECIFICATIONS IN FRENCH AND UZBEKI LANGUAGES
Open the record for dataset details and reuse information.
STRUCTURAL-SEMANTIC TYPES AND STYLISTIC FEATURES OF THE COMPOUND SENTENCE WITH MEASURE-DEGREE CLAUSE
Open the record for dataset details and reuse information.
GPTCloneBench: A comprehensive benchmark of semantic clones and cross-language clones using GPT-3 model and SemanticCloneBench
<p>This is the full dataset of GPTCloneBench (version 2)</p>
Diachronic semantics: changes of meaning of words over time and the consequences for keeping classification systems up to date
<p>Meanings of words in a natural language are changing over time under the influences of different factors. Words adapt to new meanings, lose old meanings, rearrange current meanings and change some parts of previous meanings, etc. The language as a living organism needs to be able to accept and adapt to those changes. Like natural languages, artificial languages such as classification systems or subject indexing systems have to adjust to those changes too. Linguist F. de Saussure defines language as a 'system of signs'. If we transfer this definition into the artificial language such as a classification system (e.g., Universal Decimal Classification (UDC)) and define it also as a 'system of signs' which has a vocabulary, a grammar and a syntax, we could draw parallels between phenomena which occur in natural and in artificial languages. The young linguistic discipline which deals with changes in meanings over time is called diachronic semantics and this paper explores how its mechanisms can be used to analyze the changes that occur in classification systems over time. Diachronic semantics uses various mechanisms to describe adapting and changing meanings of words. These mechanisms include: metaphor, metonymy, specialization, generalization, analogy and splitting. This paper also aims to explain the borrowing and adjusting of the theory from the field of linguistics into the field of information sciences.</p>
SEMANTIC STRUCTURE OF THE WORD IN WORLD LINGUISTICS
Open the record for dataset details and reuse information.
SEMANTIC AND STRUCTURAL ANALYSIS OF THE TOURISTIC TERMS IN THE ENGLISH AND UZBEK LANGUAGES.
Open the record for dataset details and reuse information.
LEXICAL SEMANTIC FEATURES OF SPECIFIC GENDER AFFILIATION IN THE LANGUAGE COMMUNITY
Open the record for dataset details and reuse information.
SEMANTICS WITHIN THE FRAMEWORK OF THE ANALYTIC LANGUAGE IN TERMS OF PHRASEOLOGY
Open the record for dataset details and reuse information.
CONCEPT AS A FORM OF EXPRESSION OF KNOWLEDGE ABOUT THE WORLD OF COGNITIVE SEMANTICS
Open the record for dataset details and reuse information.
STYLISTIC AND SEMANTIC FEATURES OF ENGLISH TAX TERMS
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.