Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

38

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

38 results for “semantic model”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"

<p>This dataset accompanies a replication package provided for a study on Semantic Versioning for Ansible Galaxy roles.</p> <p>The replication package is available at https://github.com/ROpdebee/ansible_semver_ext_replication</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

Semantic segmentation model of construction waste landfill based on high-resolution satellite images

<p>CWLD_model project shows scripts and instructions on how to use this dataset (<a href="../records/10686118">https://zenodo.org/records/10686118</a>) to train a segmentation model. requirements.txt files provide the libraries you need to run your project. The README.md document details the deployment process and features of each module.</p> <p>You can also visit the GitHub page for scripts and instructions on how to use this dataset for visualizing and plotting basic statistics. The models and the code to execute them are released on&nbsp;<a href="https://github.com/huangleinxidimejd/CWLD_Model">https://github.com/huangleinxidimejd/CWLD_Model</a>.</p> <h2>Training details</h2> <p>The model was trained with two GPUs, an Nvidia GeForce RTX 2080Ti, and the following parameters:</p> <ul> <li>'train_batch_size': 4,</li> <li>'val_batch_size': 4,</li> <li>'train_crop_size': 512,</li> <li>'val_crop_size': 512,</li> <li>'lr': 0.001, # the learning rate used during training. It determines how quickly the model learns from the data</li> <li>'Epoch Times': 200,</li> <li>'gpu': correct,</li> <li>'weight_decay': 5E-4,</li> <li>'Momentum': 0.9,</li> <li>'print_freq': 100,</li> <li>'predict_step': 5,</li> </ul> <h2>usage</h2> <ul> <li>After downloading the dataset from Zenodo, place the train and val files from the Deep Learning Datasets file into the data folder of the CWLD semantic segmentation model.</li> <li>Open: CWLD_ Open the root directory in CWLD_model/dataset/ and start training with the WasteSeg_Train.py file. The modelss module provides five convolutional networks, Improved_DeeplabV3_plus, PSPNet, ResNet, SegNet, and UNet, which can be selected and modified accordingly.</li> <li>The utils package provides a large number of data processing tools to use.</li> <li>The trained model can be predicted from a EvalSeg.py file.</li> </ul>

opencc-by-4.0Apr 2024View details →
zenodo44/100

JeSemE models for lexical semantic change

<p>Models for diachronic lexical semantics used by&nbsp;the&nbsp;<a href="http://jeseme.org">Jena Semantic Explorer (JeSemE)</a> web site described in our&nbsp;<a href="http://aclweb.org/anthology/C18-2003">COLING 2018 paper &quot;JeSemE: A Website for Exploring Diachronic Changes in Word Meaning and Emotion&quot;</a>.</p> <p>Also described and applied in Johannes Hellrich&#39;s Ph.D. thesis &quot;Word Embeddings: Reliability &amp; Semantic Change&quot; who was funded by the&nbsp;Deutsche Forschungsgemeinschaft (DFG) within the graduate school &quot;The Romantic Model&quot;&nbsp;(GRK 2041/1).</p> <p>One ZIP&nbsp;file per corpus, each containing several CSV files:</p> <ul> <li>CHI.csv with&nbsp;&chi;<sup>2&nbsp;</sup>word association values (structure: word-id, word-id, time, value)</li> <li>EMBEDDING.csv with SVD-PPMI word embeddings (aligned;&nbsp;structure: word-id, time, values)</li> <li>EMOTION.csv with VAD&nbsp;word emotion values (structure: word-id, time, values)</li> <li>FREQUENCY.csv with relative word&nbsp;frequency values (structure: word-id, time, value)</li> <li>PPMI.csv with PPMI<sup>&nbsp;</sup>word association values (structure: word-id, word-id, time, value)</li> <li>SIMILARITY.csv with word embedding derived word similarity&nbsp;values (structure: word-id, word-id, time, value)</li> <li>WORDIDS.csv mapping words to their corpus specific&nbsp;IDs</li> </ul> <p>Corpora are:</p> <ul> <li> <p>coha:&nbsp;Corpus of Historical American English</p> </li> <li> <p>dta:&nbsp;Deutsches Textarchiv &#39;German Text Archive&#39;</p> </li> <li> <p>google_fiction:&nbsp;Google Books N-Gram corpus, English fiction subcorpus</p> </li> <li> <p>google_german:&nbsp;Google Books N-Gram corpus, German subcorpus</p> </li> <li> <p>rsc: Royal Society Corpus&nbsp;</p> </li> </ul>

opencc-by-4.0Mar 2018View details →
zenodo44/100

Semantic 3D Tree Model Dresden 2017

<p>The semantic 3D tree model contains reconstructed tree crowns within the City of Dresden (Germany). Area-wide availability of such models and their integration into semantic 3D city models facilitates enriched visualizations of urban areas as well as 3D spatial modeling that simulates the interaction of trees with buildings and the built environment.</p> <p>The individual tree crowns were modeled using geometric primitives and correspond to the CityGML Level of Detail (LoD) 2. Individual modeling parameters were determined for each tree aiming for a realistic volume replication. LiDAR data from a survey in the year 2017 were used to parameterize the tree crowns. Tree crowns were modeled via ellipsoids fitted to crown extent, cylinders were used for trunk representation. The framework for segmenting individual trees in the LiDAR point cloud and for modeling individual tree crowns via geometric primitives is described in <a href="https://doi.org/10.1016/j.ufug.2022.127637">this article</a>.</p> <p>The tree models are available as CityGML files in the coordinate system ETRS89/UTM zone 33 (EPSG: 25833). The dataset was divided into tiles. The tile number results from the coordinate of the lower left corner in the coordinate reference system.</p> <p>The source data used was made freely available by the &ldquo;Landesamt f&uuml;r Geobasisinformation Sachsen&rdquo; (GeoSN) under the license &quot;Data license Germany - attribution - Version 2.0&quot; and can be downloaded under the following links:<br> LiDAR: <a href="https://www.geodaten.sachsen.de/downloadbereich-digitale-hoehenmodelle-4851.html">https://www.geodaten.sachsen.de/downloadbereich-digitale-hoehenmodelle-4851.html</a><br> 3D Building Model: <a href="https://www.geodaten.sachsen.de/downloadbereich-digitale-3d-stadtmodelle-4875.html">https://www.geodaten.sachsen.de/downloadbereich-digitale-3d-stadtmodelle-4875.html</a><br> Aerial Imagery: <a href="https://www.geodaten.sachsen.de/downloadbereich-dop-4826.html">https://www.geodaten.sachsen.de/downloadbereich-dop-4826.html</a></p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

smashHit semantic model

<p>The smashHit core semantic model defines the main entities that are important for the smashHit modules (e.g. consent, metadata, contract, etc.) and investigated the several existing ontologies seeking to find the ones that better matches the smashHit needs.&nbsp;</p> <p>The version v0.2 shows the final development stage of the smashHit semantic model (smashHit core ontology) within the smashHit project.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

SEMAFORA Semantic Reference Data Models

<p>To support the aim of the Semafora project, a series of Semantic Reference Data Models were created to provide a target semantic structure for the integration of standard archaeological survey data.&nbsp;</p> <p>&nbsp;</p> <p>The following models constitute the Semafora SRDM package:</p> <p>&nbsp;</p> <p>Place: This model is used to document any places associated with the archaeological survey.</p> <p>&nbsp;</p> <p>Institution: This model is used to document any institution associated with the survey.</p> <p>&nbsp;</p> <p>Period: This model is used to document the generic historical period assigned to the production of artefacts, existence of sites or other observable archaeological and historical events.</p> <p>&nbsp;</p> <p>Feature: This model is used to document any physical features, such as walls and other human-made structures observable on the field.</p> <p>&nbsp;</p> <p>Project: This model is used to document the overarching project, a part of which is the archaeological survey. Some projects may involve surveys, excavations, and other archaeological activities.</p> <p>&nbsp;</p> <p>Site: This model is used to document a site declared as archaeological as a result of the survey process.</p> <p>&nbsp;</p> <p>Digital Object: This model is used to document any type of digital asset associated with the survey.</p> <p>&nbsp;</p> <p>Survey Unit: This model is used to document a defined survey unit where the survey activity happens. It has both the properties of a place with dimensions and coordinates and of a physical thing from which samples can be collected.</p> <p>&nbsp;</p> <p>Collection: This model is used to document a collection of physical things, usually artefacts, collected while surveying.</p> <p>&nbsp;</p> <p>Artefact: This model is used to document individual artifacts collected from while surveying as a part of a larger collection of material things or as a singular artefact collection or documentation.</p> <p>&nbsp;</p> <p>Image: This model is used to document any image representing components of the archaeological survey, such as artefacts, features, places, people, etc.</p> <p>&nbsp;</p> <p>Observation: This model is used to document the act of observation usually associated with archaeological sites or survey units and the properties assigned to those as a result of the observation.</p> <p>&nbsp;</p> <p>Bibliography: This model is used to document any textual object associated with the survey or any components of it.</p> <p>&nbsp;</p> <p>Sample: This model is used to document a material sample of the survey unit. It partially overlaps with collection but acts as a parent sample that may contain other physical things besides human-made objects.</p> <p>&nbsp;</p> <p>Person: This model is used to document an individual person (alive or dead) involved in some way in the survey process.</p> <p>&nbsp;</p> <p>These models are intended to be used in order to guide semantic data mapping processes as well as to provide instructions for the creation of a target data semantic data management system.</p> <p>&nbsp;</p> <p>Each model&rsquo;s semantic reference data model description is stored here as a csv. The ongoing curation and updating of these SRDMs is undertaken using the Zellij system and can be accessed here:</p> <p>&nbsp;</p> <p><a href="https://zellij.pythonanywhere.com/docs/list/appaCfYmUmH78z85c">https://zellij.pythonanywhere.com/docs/list/appaCfYmUmH78z85c</a></p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

A core ontology for modeling life cycle sustainability assessment on the Semantic Web with Accompanying Database

<p>To enable and support the uptake of semantic ontologies, we present a core ontology developed specifically to capture the data relevant for life cycle sustainability assessment. We further demonstrate the utility of the ontology by using it to integrate data relevant to sustainability assessments, such as EXIOBASE and the Yale Stocks and Flow Database to the Semantic Web. These datasets can be accessed by the machine-readable endpoint using SPARQL, a semantic query language.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Dataset for semantic segmentation of the laboratory model of manufacturing environment

<p>This dataset includes images and labels used for semantic segmentation of the laboratory model of the manufacturing environment, at the University of Belgrade - Faculty of Mechanical Engineering. The dataset is gathered by using mobile robot RAICO (Robot with Artificial Intelligence based COgnition) and its stereo visual system made from two Basler acA1920-25uc&nbsp;cameras with Fujinon&nbsp;lens DF6HA-1B. The dataset includes close to 430 images with a&nbsp;resolution of 640x360. Images are acquired by both cameras at different mobile robot poses in the laboratory model of a manufacturing environment. Five classes are introduced in the dataset, machines 1 to 4, and a background class.&nbsp;The exact names of the classes are:</p> <p>classNames = [&quot;Machine_1&quot;, &quot;Machine_2&quot;, &quot;Machine_3&quot;, &quot;Machine_4&quot;, &quot;Background&quot;];</p> <p>while the labels of the classes (RGB values of label images) are:</p> <p>labelIDs = [ ...<br> &nbsp; &nbsp; 000 000 255; ... % &quot;Machine 1&quot;<br> &nbsp; &nbsp; 000 255 255; ... % &quot;Machine 2&quot;<br> &nbsp; &nbsp; 255 255 000; ... % &quot;Machine 3&quot;<br> &nbsp; &nbsp; 255 000 000; ... % &quot;Machine 4&quot;<br> &nbsp; &nbsp; 255 255 255; ... &nbsp; % &quot;Background&quot;<br> &nbsp; &nbsp; ];</p> <p>Image and label pairs are entitled 1 to 430, and e.g. label 5 corresponds to images 5.</p> <p>This dataset was developed with the support&nbsp;of the Science Fund of the Republic of Serbia, Grant No. 6523109, AI - MISSION4.0, 2020-2022.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

VC-SLAM Versatile Corpus for Semantic Labeling And Modeling

<p>Benchmark Corpus for semantic labeling and modeling.</p> <p>This corpus contains 101 data sets from different open data portals.<br> Each data set consists of the following data:</p> <ul> <li>Raw csv data [rawdata_csv]</li> <li>Large json data sample [json_sample_large]</li> <li>Small json data sample [json_sample_small]</li> <li>Raw data in csv format [rawdata_csv]</li> <li>Raw data samples in csv format [rawdata_csv_samples]</li> <li>Mappings to translate between csv and json files [csv_json_mappings]</li> <li>Textual description / Metadata [descriptions]</li> <li>Semantic model as rdf/ttl [semantic_models]</li> <li>Mappings describing mapping between raw data attributes and concepts from the ontology [mappings]</li> <li>List of attributes that have been ignored during modeling [ignored_attributes]</li> </ul> <p>Additionally the corpus contains a target ontology as rdf/ttl [ontology].</p> <p>The individual data sets are licensed by the licenses specified in the attached Excel sheet (DataSetOverview.xlsx)</p> <p>&nbsp;</p> <p>These data are provided &quot;as is&quot;, without any warranties of any kind. The data are provided under the Creative Commons Attribution 4.0 International license.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Dataset for generating LOD3 building models from structure-from-motion and semantic segmentation

<p>This repository contains the codes for computing geometrical digital twins as LOD3 models for buildings, using a structure from motion and semantic segmentation. The methodology hereby implements was presented in the paper [Generating LOD3 building models from structure-from-motion and semantic segmentation&quot; by Pantoja-Rosero et., al. (2022)] (<a href="https://doi.org/10.1016/j.autcon.2022.104430">https://doi.org/10.1016/j.autcon.2022.104430</a>)</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Figure 1: The ECG model-MAPPING BETWEEN SEMANTIC GRAPHS AND SENTENCES IN GRAMMAR INDUCTION SYSTEM

<p>The following Figure 1 shows a sample semantic graph that describes a<br> simple test world.<br> During the processing of the ECG, the base units of the graph are the ECG<br> atoms. An ECG atom corresponds to a primitive statements related to one<br> predicate. It has a structure of one-level deep tree, where the root of the tree<br> is the predicate and the concepts linked to it are the leaves. The child concept<br> of the root predicate may be not only a single concept but it can be another<br> ECG atom.</p>

opencc-by-4.0Jun 2010View details →
zenodo40/100

Train and Evaluation Code, Road Classification Models and Test set of the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography"

<p>This repository contains the Python scripts built for training and evaluation of the implementation, together with the test data and the resulting road segmentation models corresponding to the paper "Insights into the Effects of Image Overlap and Image Size on Semantic Segmentation Models Trained for Road Surface Area Extraction from Aerial Orthophotography". The scripts make use of the Tensorflow with Keras framework and their additional required dependencies.</p> <p>The training and validation set is based on the binary SROADEX dataset (<a href="../records/6482346">https://zenodo.org/records/6482346</a>) that was re-split into tiles that feature the image resolutions (256 x 256, 512 x 512, and 1024 x 1024 pixels) and image overlaps (0% and 12.5%) considered in this study. The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography that represents the axes of the different types of roads (urban, interurban and rural). This binary road data contains information from 16 full orthoimages (28.5 km * 18.5 km) with spatial resolution of 0.5 m/pixel from the insular and peninsular Spanish territory. Due to the size on disk of approximately 492 gigabytes, this training and validation data is only available upon request from the corresponding author. The test set has been generated from a novel area from Palencia (Spain) and features 18 million pixels labelled with the positive "Road" class. The test sets are provided in the repository for each resolution (with no overlap), so that additional DL models can be evaluated on the same data and compared with the results achieved in this study.</p> <p>The structure of the information shared in this repository is as follows:<br>The scripts have been grouped by tile resolution (256, 512 and 1024). First, the test set and the evaluation script can be found. For each tile resolution, there are two subfolders (corresponding to the "no overlap" and "12.5% overlap"). In each case, the Python scripts for training the models in the three repetitions are shared, and the trained models (H5 format) are shared in compressed form. Finally, for each resolution we also share the testing dataset which consists of two folders.</p> <p>The material is distributed under a CC-BY 4.0 license.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - model weights

<p>In the related <a href="https://github.com/HybridNLP2018/tutorial">notebook&nbsp;</a>we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains the model weights trained on such large corpora.</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms Dataset

<p>This is the official data repository for the paper: "Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms", available at:<a href="https://www.arxiv.org/abs/2409.19371"> https://www.arxiv.org/abs/2409.19371</a>. The corresponding code is available at: <a href="https://github.com/david-stojanovski/EDMLX">https://github.com/david-stojanovski/echo_from_noise</a></p> <p>&nbsp;</p> <p>The synthetic data is produced using a variety of generative architectures, including the <strong>Elucidating Diffusion Model (EDM), Variance Exploding (VE), Variance Preserving (VP)</strong>, and our novel models, <strong>EDM-L64</strong> and <strong>EDM-L128</strong>, which employ <strong>latent diffusion</strong> strategies to significantly reduce computational cost. By incorporating&nbsp;<strong>spatially adaptive normalization (SPADE) blocks</strong> and <strong>&Gamma;-distribution-based Variational Autoencoders (&Gamma;-VAE)</strong>, these datasets ensure that the generated images preserve the essential semantic features required for training deep learning models.</p> <p>&nbsp;</p> <p>All pretrained classification and segmentation models can be found within the <strong>trained_models </strong>file.</p> <p>All generated images can be found within the&nbsp;<strong>generated_data&nbsp;</strong>file. Included is the <strong>CAMUS</strong> and original <strong>Semantic Diffusion Model (SDM)&nbsp;</strong>data, as well as a folder labelled&nbsp;<strong>easy_inference</strong> designed to contain all relevant labelmaps in a convenient folder for generating replicas of the dataset (detailed at codebase).</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

NIVA Common Semantic Model

<p>NIVA has developed under Enterprise Architect a UML data model for some pieces of IACS data or processes, namely core geographic data, EO monitoring and Farm Registry. The NIVA model may be found under: EU Common Agricultural Model /Conceptual model.&nbsp;More detailed explanations about content of this model may be found in NIVA deliverable D3.2 Common Semantic Model, available on the NIVA web site : <a href="https://www.niva4cap.eu/deliverables/">Deliverables &ndash; Niva4cap</a>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Echo from noise: synthetically generated cardiac ultrasound data using semantic diffusion models

<p>This is the data repository for the paper: &quot;Echo from noise: synthetic ultrasound image generation using diffusion models for real image segmentation&quot;, available at: https://arxiv.org/abs/2305.05424. The corresponding code is available at:&nbsp;https://github.com/david-stojanovski/echo_from_noise</p> <p>&nbsp;</p> <p>This is the first work to utilize Denoising Diffusion Probabilistic Models (DDPMs)&nbsp;for generating medical images using semantic label maps as a source image for conditioning the generated image.</p> <p>Each of the 400+50 CAMUS patients contributes with 4 labelled frames (ED and ES for 2 chamber and 4 chamber), totalling 1800 initial semantic maps, to which we added the sector label. These semantic maps then had five random deformations applied (a combination of random affine and elastic deformation) to produce, 9000 transformed semantic maps (8000 for training and 1000 for validation).&nbsp;</p> <p>Affine transformation ranges for rotation degrees, translate, scale and shear were: (-5, 5), (0, 0.05), (0.8, 1.05)&nbsp;and 5&nbsp;respectively. This was implemented using the torchvision python package. Elastic deformation was implemented using the TorchIO package. The settings for number of control points and max displacement were (10, 10, 4)&nbsp;and (0, 30, 30)&nbsp;respectively.</p> <p>Using these 9000 semantic maps as input to the generative models, we produced 9000 synthetic ultrasound images.</p> <p>Each echo view folder contains 3 folders:</p> <p>1) annotations: augmented labels, with no sector label and no clipping due to sector</p> <p>2) images: semantic diffusion model inferenced images</p> <p>3) sector_annotations: label maps which contain ultrasound cone sector, which were used to generate corresponding semantic diffusion model images</p> <p>ema_0.9999_050000_2ch_ed_256.pt and&nbsp;ema_0.9999_050000_4ch_ed_256.pt are the saved checkpoints for the 2 and 4 chamber diffusion models respectively.</p> <p>The pretrained segmentation networks are provided within the&nbsp;<a href="https://zenodo.org/api/files/0af4e6a3-234d-40a3-8351-c91261628982/final_models.zip">final_models.zip</a>&nbsp;file.</p> <p>A diagram of image numbers is shown in&nbsp;<a href="https://zenodo.org/api/files/0af4e6a3-234d-40a3-8351-c91261628982/Data%20diagram.png">Data diagram.png</a></p>

opencc-by-4.0May 2023View details →
zenodo40/100

AGREE: a New Benchmark for the Evaluation of Semantic Models of Ancient Greek

<p>AGREE (Ancient Greek Relatedness Embeddings Evaluation) is a benchmark for the evaluation of semantic models of Ancient Greek created at the University of Groningen (The Netherlands). More information about it can be found in the following publication:</p> <p>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim, AGREE: a new benchmark for the evaluation of distributional semantic models of ancient Greek,&nbsp;<em>Digital Scholarship in the Humanities</em>, Volume 39, Issue 1, April 2024, Pages 373&ndash;392,&nbsp;<a href="https://doi.org/10.1093/llc/fqad087">https://doi.org/10.1093/llc/fqad087</a></p> <p>&nbsp;</p> <p><strong>1. Overview of the repository</strong></p> <p>This benchmark was created from a mix of expert judgements about relatedness between Ancient Greek words and model outputs validated by human experts. The evaluation items are pairs of Ancient Greek lemmas with a high semantic relatedness.</p> <p>The human judgements were collected via two questionnaires, proposing two different tasks to the experts. The evaluation items included in the AGREE benchmark are a selection of the most strictly related pairs of lemmas obtained from the two tasks. Here an overview of the contents of the repository:</p> <ul> <li><strong>1_agree_task1.json</strong>&nbsp;includes all the data collected with the first task. The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'frequency': the number of times that the pair was suggested as related by an expert;</li> <li>'POS1': part-of-speech of the first lemma;</li> <li>'POS2': part-of-speech of the second lemma;</li> <li>'benchmark': inclusion of the pair in the AGREE benchmark ('yes'/'no').</li> </ul> </li> <li><strong>2_agree_task2.json&nbsp;</strong>includes all the data collected with the second task.&nbsp;The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'origin':&nbsp; <ul> <li>'common_pair' = one of the two pairs proposed to all participants in the second task;</li> <li>'task1' = pairs proposed by experts in the first task;</li> <li>'models_easy_rel' = output of word2vec models, pair considered as strictly related;</li> <li>'models_task1' = pairs proposed by experts in the first task and also output by word2vec models;</li> <li>'models' = output of word2vec language models;</li> <li>'unrelated' = made up pairs of unrelated lemmas (control pairs);</li> </ul> </li> <li>'respondents': number of experts evaluating a pair;</li> <li>'score': average relatedness score given by the experts on a 0-100 scale;</li> <li>'agreement': inter-annotated agreement between all experts who evaluated the&nbsp;block of pairs to which the current&nbsp;pair belongs&nbsp;(when available, i.e. when the block of pairs was presented to more than one participant);</li> <li>'benchmark': inclusion of the pair in the AGREE benchmark ('yes'/'no').</li> </ul> </li> <li><strong>3_agree_final_benchmark.json </strong>includes the&nbsp;final selection of items that constitutes AGREE. The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'origin': <ul> <li>'task1': pair either proposed more than once in the first task or proposed only once, but scored&nbsp;&gt;=&nbsp;70 in the second task;</li> <li>'task2': pair scored by more than one respondent in the second task and with average score &gt;= 70.</li> </ul> </li> </ul> </li> </ul> <p>&nbsp;</p> <p>This updated version of the repository includes the individual answers to the two questionnaires (see files 'answers_Task1_postprocessed.xlsx' and 'raw_answers_Task2.xlsx').</p> <p>&nbsp;</p> <p><strong>2. Acknowledgements</strong></p> <div>This work was partially supported by the Young Academy Groningen through the PhD scholarship of Silvia Stopponi.<br>&nbsp;<br>We acknowledge the financial support of Anchoring Innovation. Anchoring Innovation is the Gravitation Grant research agenda of the Dutch National Research School in Classical Studies, OIKOS. It is financially supported by the Dutch ministry of Education, Culture and Science (NWO project number 024.003.012). For more information about the research programme and its results, see the website <a href="https://www.anchoringinnovation.nl">www.anchoringinnovation.nl</a>.<br>&nbsp;<br>We want to thank the experts of Ancient Greek around the world who shared their knowledge of Ancient Greek semantics and donated some of their precious time. Without them the creation of this benchmark would not have been possible.<br>&nbsp;<br>We also want to thank the many colleagues from the University of Groningen, the National Research School OIKOS, and other Universities abroad who contributed to this work with discussion and advice.</div> <div>&nbsp;</div> <div>&nbsp;<br><strong>3. Citation</strong><br>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim, AGREE: a new benchmark for the evaluation of distributional semantic models of ancient Greek, <em>Digital Scholarship in the Humanities</em>, Volume 39, Issue 1, April 2024, Pages 373&ndash;392, <a href="https://doi.org/10.1093/llc/fqad087">https://doi.org/10.1093/llc/fqad087</a></div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0Feb 2023View details →
zenodo36/100

On the Effect of Semantically Enriched Context Models on Software Modularization

<p>The dataset used for evaluating the approaches outlined in this paper, comprising of 10 open source Java projects. The algorithms employed can be found at https://github.com/amirms/GeLaToLab, </p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

Raw data for the creation of a maturity model for Catalogues of Semantic Artefacts

<p>This dataset includes two data collections (in two different formats, i.e. CSV and XLSX) with the raw data used for creating the <a href="https://doi.org/10.5281/zenodo.10618105">Maturity Dimensions and Sub-Criteria for Catalogues of Semantic Artefacts</a>. In particular:</p> <p>1. <em>Dimension identification in literature</em> includes the list of relevant materials gathered involving all the members of the EOSC Task Force on Semantic Interoperability that include (1) definitions of semantic artefact catalogues and (2) dimensions that can be used to measure the maturity of such catalogues;</p> <p>2. <em>Catalogue assessment</em> is the result of the analysis of 26 different catalogues of semantic artefacts against the dimensions and sub-criteria described in the maturity model.</p>

opencc-zeroMay 2023View details →
zenodo36/100

Data, code, models for "Weakly Supervised Semantic Segmentation for Joint Key Local Structure Localization and Classification of Aurora Image"

<p>Data, code and models for https://ieeexplore.ieee.org/document/8410588/</p>

opencc-by-4.0Jul 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record