Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

483

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

483 results for “SEMANTICS”

Learn how ShareScore rates datasets ↗
zenodo28/100

Japan landslide dataset for semantic segmentation

<p>This database contains images used for the semantic segmentation of landslide scars from a fully convolutional neural network U-Net.</p> <p>1. <strong>Training dataset: </strong>it contains 125 GeoTIFF 8 bits images and associated PNG masks (scars indicated in white and background in black color).</p> <p>2. <strong>Validation dataset</strong>: it contains 10 GeoTIFF 8 bits images and associated PNG masks used for U-Net validation step.</p> <p>3. <strong>Test dataset:&nbsp;</strong>it contains 10 GeoTIFF 8 bits images and associated PNG masks for testing.</p> <p>Also, the &quot;SHAPEFILES_LANDSLIDES.rar&quot; file contains the vector layers of the masked images in .shp format.</p>

opencc-by-4.0Apr 2020View details →
zenodo28/100

Investigating Distributional Robustness: Semantic Perturbations Using Generative Models (ImageNet Examples)

<p>This dataset contains examples of semantically-perturbed images, for NeurIPS 2020 submission #4915.</p> <p>There are four top-level folders, each containing results for semantic perturbations restricted to adjust the activation values at only certain layers of the BigGAN generative network: the first six layers, the middle six layers, the last six layers, and all layers.</p> <p>Within each top-level folder, there are a further four folders, each corresponding to a classifier neural network whose evaluation is being evaluated. These are&nbsp;EfficientNet-B4 with NoisyStudent training [1],&nbsp;the standard ResNet50 [2], a pixel-perturbation-robust ResNet50 trained by&nbsp;Engstrom et al. [3]&nbsp;and another trained by&nbsp;Wong et al. [4], using their &quot;Fast is better than free&quot; technique.</p> <p>Within each of these, there are many folders, named &#39;version_$N&#39;. Each one of these contains three images: the unperturbed generated image, named&nbsp;unpert_generated_x_grid_0.png; the semantically-perturbed generated image, named&nbsp;generated_x_grid_0.png; and an image named semantic_pert_diffs_grid_0.png showing the pixel-space effect of the semantic perturbation, that is, the diff between the perturbed and unperturbed images. Note that if the&nbsp;perturbed and unperturbed images are identical, and the classifier misclassifies the unperturbed images, and so we skip this example.</p> <p>Along with the &#39;version_$N&#39; folders containing the images, there exists a file for each classifier named results.json. Each top-level item in this JSON file corresponds to one &#39;version_$N&#39; example. There are 5 attributes: &#39;label&#39;, indicating the target label of the unperturbed image; &#39;magnitude&#39;, which gives the magnitude of the semantic perturbation found; &#39;skipped_cla&#39;, which is 1 if the example is skipped because the classifier did not correctly classify the unperturbed image; &#39;skipped_judge&#39;, which is 1 if the human judged that the unperturbed image did not match its label, so this example is skipped; and &#39;pert_judgement&#39;, which is 1 if the semantically-perturbed image is judged by the human to be of the same class as the unperturbed image. These judgements on these images were used to construct the main graphs in the paper.</p> <p>&nbsp;</p> <p>[1]&nbsp;Qizhe Xie, Eduard H. Hovy, Minh-Thang Luong, and Quoc V. Le. Self-training with Noisy Student improves ImageNet classification. CoRR, abs/1911.04252, 2019. URL http://arxiv.org/abs/1911.04252.</p> <p>[2] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 770&ndash;778. IEEE Computer Society, 2016. doi: 10.1109/CVPR.2016.90. URL https://doi.org/10.1109/CVPR.2016.90.</p> <p>[3] Logan Engstrom, Andrew Ilyas, Shibani Santurkar, and Dimitris Tsipras. Robustness (Python library), 2019. URL ttps://github.com/MadryLab/robustness.</p> <p>[4] Eric Wong, Leslie Rice, and J. Zico Kolter. Fast is better than free: Revisiting adversarial training. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL<br> https://openreview.net/forum?id=BJx040EFvH.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Semantic text mining in early drug discovery for type 2 diabetes

<p>BACKGROUND: Surveying the scientific literature is an important part of early drug discovery; and with the ever-increasing amount of biomedical publications it is imperative to focus on the most interesting articles. Here we present a project that highlights new understanding (e.g.\ recently discovered modes of action) and identifies potential novel drug target, via a novel, data-driven text mining approach to score type 2 diabetes (T2D) relevance. We focused on monitoring trends and jumps in T2D relevance to help us be timely informed of important breakthroughs.<br> &nbsp;&nbsp; &nbsp;<br> METHODS: We extracted over 7 million <em>n</em>-grams from PubMed and then clustered around 240,000 linked to T2D into almost 50,000 T2D relevant `semantic concepts&#39;. To score papers, these concepts were weighted depending on co-mentioning with core T2D proteins. A protein&#39;s current T2D relevance was determined by combining the scores of the papers mentioning it in the preceeding five years. The significance of a jump in a protein&#39;s rank was assessed by comparing it to previously observed jumps.<br> &nbsp;&nbsp; &nbsp;<br> RESULTS: We show that T2D relevant papers, also those not mentioning T2D explicitly, got assigned high scores by mentioning semantic concepts often used in connection with T2D, as shown by the enrichment of well known T2D proteins among the top scoring proteins. Our `high jumpers&#39; identified important past developments in the apprehension of how certain key proteins relate to T2D, indicating that our method will make us aware of future breakthroughs. In summary, this project facilitated keeping up with current T2D research by repeatedly providing short lists of potential novel targets into our early drug discovery pipeline.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

Amazon Rainforest dataset for semantic segmentation V2

<p>This database contains images used for training a fully convolutional neural network for the semantic segmentation of forested areas in images from the Sentinel-2 Level 2A Satellite.</p> <p>The images refer to the RGB composition (bands 4, 3 and 2). The histogram for each image was selected as follows:</p> <ul> <li>Band 4 (red): values ranging from 103 to 2724;</li> <li>Band 3 (green): values ranging from 194 to 2888;</li> <li>Band 2 (blue): values ranging from 99 to 2798.</li> </ul> <p>After that, each band was converted to a byte type (0-255).</p> <p>The images are still divided into three main sets: training, validation and testing:</p> <ol> <li><strong>Training dataset: </strong>it contains 1.123 GeoTIFF images with 512x512 pixels and associated PNG masks (clouds indicated in white and background in black color).</li> <li><strong>Validation dataset</strong>: it contains 100 GeoTIFF images with 512x512 pixels and associated PNG masks used for validation step.</li> <li><strong>Test dataset:&nbsp;</strong>it contains 100 GeoTIFF images 512x512 pixels for testing.</li> </ol>

opencc-by-4.0Aug 2020View details →
zenodo28/100

Impact of Combining Syntactic and Semantic Similarities on Patch Prioritization

<p>This dataset contains 246 bugs,&nbsp;their fixes&nbsp;and corresponding buggy projects from historical bug fixes dataset (https://github.com/xuanbachle/data-bugfixes) that fulfill the following criteria:</p> <ol> <li>Unique</li> <li>Satisfy redundancy assumption at file level</li> <li>Fixed by applying replacement mutation</li> <li>Require fixing at expression level</li> <li>Having available project and dependency files</li> </ol> <p>For details, please view https://www.scitepress.org/Link.aspx?doi=10.5220/0009411301700180</p>

opencc-by-4.0May 2020View details →
zenodo28/100

Graphical and Collaborative Annotation Support for Semantic Web Services

<p><strong>Graphical and Collaborative Annotation Support for Semantic Web Services presentation</strong></p>

opencc-by-4.0Oct 2020View details →
dryad28/100

Data from: A revision of Evaniscus (Hymenoptera, Evaniidae) using ontology-based semantic phenotype annotation

The Neotropical evaniid genus Evaniscus Szépligeti currently includes six species. Two new species are described, Evaniscus lansdownei Mullins, sp. n. from Colombia and Brazil and Evaniscus rafaeli Kawada, sp. n. from Brazil. Evaniscus sulcigenis Roman, syn. n., is synonymized under Evaniscus rufithorax Enderlein. An identification key to species of Evaniscus is provided. Thirty-five parsimony informative morphological characters are analyzed for six ingroup and four outgroup taxa. A topology resulting in a monophyletic Evaniscus is presented with Evaniscus tibialis and Evaniscus rafaeli as sister to the remaining Evaniscus species. The Hymenoptera Anatomy Ontology and other relevant biomedical ontologies are employed to create semantic phenotype statements in Entity-Quality (EQ) format for species descriptions. This approach is an early effort to formalize species descriptions and to make descriptive data available to other domains.

opencc-zeroDec 2011View details →
zenodo28/100

Replication Package for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"

<p>Replication package for our analysis of Semantic Versioning in Ansible Galaxy role repositories.</p> <p>This replication package consists of three parts:</p> <ul> <li> <p>Classification Model: Contains Jupyter notebooks used to train and evaluate a Random Forest classification model based on structural features. Training and evaluation data is included.</p> </li> <li> <p>Quantitative Notebooks: Contains Jupyter notebooks used to perform quantitative analyses of versions and changes.</p> </li> <li> <p>data: CSV files of the data used in the Quantitative Notebooks, and the source data for the classification model. Should be downloaded separately fromthe classification model. Should be downloaded separately from <a href="https://doi.org/10.5281/zenodo.4991955">https://doi.org/10.5281/zenodo.4991955</a>.</p> </li> </ul> <p>The data is under the Creative Commons Attribution Share-Alike 4.0 license. The source code is under the GNU General Public License.</p>

openother-openJun 2021View details →
zenodo28/100

Semantic Concordance rate data used to generate S2_Table10

<p>Semantic Concordance rate data used to generate S2_Table10 described in Appendix S2.</p>

opencc-by-4.0Nov 2016View details →
zenodo28/100

THE PROBLEM OF SEMANTIC TAGGING OF PHILOSOPHICAL TERMS IN THE UZBEK LANGUAGE IN THE CORPUS

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo28/100

SOME CHARACTERISTICS OF WORD SEMANTIC SPECIFICATIONS IN FRENCH AND UZBEKI LANGUAGES

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo28/100

STRUCTURAL-SEMANTIC TYPES AND STYLISTIC FEATURES OF THE COMPOUND SENTENCE WITH MEASURE-DEGREE CLAUSE

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo28/100

GPTCloneBench: A comprehensive benchmark of semantic clones and cross-language clones using GPT-3 model and SemanticCloneBench

<p>This is the full dataset of GPTCloneBench (version 2)</p>

opencc-by-4.0Aug 2023View details →
zenodo28/100

Diachronic semantics: changes of meaning of words over time and the consequences for keeping classification systems up to date

<p>Meanings of words in a natural language are changing over time under the influences of different factors. Words adapt to new meanings, lose old meanings, rearrange current meanings and change some parts of previous meanings, etc. The language as a living organism needs to be able to accept and adapt to those changes. Like natural languages, artificial languages such as classification systems or subject indexing systems have to adjust to those changes too. Linguist F. de Saussure defines language as a 'system of signs'. If we transfer this definition into the artificial language such as a classification system (e.g., Universal Decimal Classification (UDC)) and define it also as a 'system of signs' which has a vocabulary, a grammar and a syntax, we could draw parallels between phenomena which occur in natural and in artificial languages. The young linguistic discipline which deals with changes in meanings over time is called diachronic semantics and this paper explores how its mechanisms can be used to analyze the changes that occur in classification systems over time. Diachronic semantics uses various mechanisms to describe adapting and changing meanings of words. These mechanisms include: metaphor, metonymy, specialization, generalization, analogy and splitting. This paper also aims to explain the borrowing and adjusting of the theory from the field of linguistics into the field of information sciences.</p>

openJul 2015View details →
zenodo28/100

SEMANTIC STRUCTURE OF THE WORD IN WORLD LINGUISTICS

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo28/100

SEMANTIC AND STRUCTURAL ANALYSIS OF THE TOURISTIC TERMS IN THE ENGLISH AND UZBEK LANGUAGES.

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

LEXICAL SEMANTIC FEATURES OF SPECIFIC GENDER AFFILIATION IN THE LANGUAGE COMMUNITY

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

SEMANTICS WITHIN THE FRAMEWORK OF THE ANALYTIC LANGUAGE IN TERMS OF PHRASEOLOGY

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo28/100

CONCEPT AS A FORM OF EXPRESSION OF KNOWLEDGE ABOUT THE WORLD OF COGNITIVE SEMANTICS

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo28/100

STYLISTIC AND SEMANTIC FEATURES OF ENGLISH TAX TERMS

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record