Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
483
datasets available to search
ShareScore release 0.9.0
Dataset results
483 results for “SEMANTICS”
A Simple Semantic-based Data Storage Layout for Querying Point Clouds
<p>Dataset contains three datasets: SMALL, MEDIUM and LARGE point cloud.</p> <p>Contains Python code which is made up of two files pc_new_semantics.py and paperutils.py</p> <p> </p>
NCI Semantic Competency Query Review
<p><strong>Overview</strong></p> <p>NCI held a Workshop on Semantics to support the NCI Cancer Research Data Commons (CRDC) in May 2018 at the National Cancer Institute in Rockville, MD. This workshop brought together experts in various areas of semantics, data integration and harmonization, Natural Language Processing (NLP) and other relevant areas to discuss and gather recommendations on semantic support for the CRDC.</p> <p>The workshop goals were to:<br> 1) Identify high-level requirements to address semantic needs and potential approaches for evaluation testing of the Cancer Data Aggregator (CDA)</p> <p>2) A set of options for using and/or extending current methods and resources (e.g. NLP) to</p> <ul> <li> <p>support semantic query capabilities</p> </li> </ul> <ul> <li> <p>facilitate metadata annotation</p> </li> <li> <p>minimize efforts for data validation and submission</p> </li> </ul> <p>3) Develop recommendations to support ongoing engagement with the community to ensure the semantics underlying the CDA improve and evolve as people contribute to and use the CDA</p> <p><strong>Participants</strong></p> <p>In total, 33 participants attended the meeting, coming from various backgrounds including clinicians, ontologists, bioinformaticians, data scientists, and project managers. Participants had expertise in semantic technologies, software and infrastructure development, data standards, data integration, clinical research, open source tool development. </p> <p><strong>Competency Queries</strong></p> <p>At the workshop, participants were asked to brainstorm ‘competency queries’, potential queries or questions that they would ask the future Cancer Data Aggregator (CDA) in order to retrieve data from across the CRDC. At the workshop, the breakout groups documented 237 queries for the CDA. </p> <p>Competency queries are often used to inform requirements to build a data model and/or ontology. They can help inform the scope of the model:what queries should the model support; 2) the content of the model, in terms of what types of entity types and attributes are needed to answer these queries; 3) the structure of the model in terms of what types of relationships between entities are needed to efficiently answer queries; 4) the semantics of the data, meaning which terminologies/ontologies would be useful for representing data to support query needs; 5) how to test and improve a completed model to ensure it can efficiently support queries determined to be in scope.</p> <p>After the workshop, a small subgroup assessed, organized and summarized the 237 queries that were noted at the workshop. A spreadsheet was created containing the queries and the keywords in each query were highlighted. From the highlighted keywords, a column was added to capture the core search parameters or classifications for each query. In evaluating the queries it was observed that some were really not a query, but expressed various observations about the data that one might hope to make. For example “Patients with a certain temporal pattern of diagnoses, both cancer and comorbidities”. This submission indicates that the data returned would need to include diagnosis and other conditions and can help to inform the requirements for CRDC data models. </p> <p>Using the information from the keyword analysis, the queries were initially categorized across various classifications, such as queries that included exposure information, diagnosis or cancer types, anatomical location of tumor, etc. In total, the queries were classified amongst 25 different parameters, or an ‘other’ category, where the query did not fit the classification scheme, or was out of scope. To further refine this list into a more manageable list, 82 representative queries were pulled out, with at least 2 examples from every classification parameter. The goal was to identify a minimal or at least smaller subset that was still representative. This list of 82 queries was then reviewed with a larger group of experts and further refined. Additional classification parameters were added, for a total of 31 parameters and an ‘other’ category. Some classifications were subdivided into more granular classifications, such as treatment was subdivided into surgery/radiation and protocols/regimens.</p> <p>In classifying each query, the exact words from each query that fit the classification scheme were noted. For example, consider the query, “What environmental exposures are typically associated with the development of salivary gland cancer?”; this query is classified as an exposure (environmental exposure), a diagnosis or specific cancer type (salivary gland cancer), and a tumor location (salivary gland). </p> <p>For a central query to be effective, we felt that the use of preferred terms for each classification parameter would be useful, so each was mapped to a relevant terminology or ontology. For example, exposure data is represented in the Environmental and Exposures Ontology (ECTO), as well as NCIt. The Uber Anatomy Ontology (Uberon) contains classifications of anatomical structures, which can be used to classify tumor locations. Many of the parameters covered by specialized terminologies are also covered by the NCIt. In some cases, the parameters were covered by multiple ontologies. At some point a preferred terminology will need to be selected for each parameter, perhaps informed by an assessment of what is being used in the CRDC data. It is likely that all terminologies and ontologies will need to be extended to cover all the terminology needed. Mappings between these terminologies and those used in the CRDC data will need to be developed. </p> <p>Finally, categories were prioritized based on how relevant and feasible they were for the CRDC as either high priority, nice to have or low priority.</p>
Semantic links between selected CSV datasets harvested by the European Data Portal and the DBpedia knowledge graph
<p>These dataset contains the results of the interlinking process between selected csv datasets harvested by the European DAta Portal and the DBpedia knowledge graph. </p> <p>We aim at answering the following questions:<br> What are the more popular column types? This will provide hindsight about what the datasets hold and how they can be joined. It will also provide hindsight on what specific linking schemes could be applied in future elements.<br> What datasets have columns of the same type? This will suggest datasets that may be similar or related.<br> What entities appear in most datasets (co-referent entities)? This will suggest entities for which more data is published.<br> What datasets share a particular entity? This will suggest datasets that may be joined, or are related through that particular entity</p> <p>Results are provided as augmented tables, that contain the columns of the original csv, plus a metadata file in JSON-LD format. The metadata files can be loaded in an RDF-store and queried.</p> <p>Refer to the accompanying report of activities for more details on the methodolog and how to query the dataset.</p> <p><br> </p>
Towards Green Cartography & Visualization: An automated, semantically-enriched method of generating energy-aware color schemes for digital maps and visualizations
<p>Towards Green Cartography & Visualization: An automated, semantically-enriched method of generating energy-aware color schemes for digital maps and visualizations</p>
Learning Unsupervised Knowledge-Enhanced Representations to Reduce the Semantic Gap in Information Retrieval (Evaluation datasets)
<p>This dataset contains all the runs, pools, plots and analyses to reproduce the results presented in the paper: "Learning Unsupervised Knowledge-Enhanced Representations to Reduce the Semantic Gap in Information Retrieval ", 2020.</p>
SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection
<p><strong>Authors</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky, and Nina Tahmasebi</p> <p><strong>Description</strong></p> <p>This data collection contains the <strong>post-evaluation</strong> data for <a href="https://languagechange.org/semeval">SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>:</p> <ul> <li>the starting kit to download data, and examples for competing in the CodaLab challenge including baselines</li> <li>the true binary change scores of the targets for Subtask 1, and their true graded change scores for Subtask 2 (<code>test_data_truth/</code>),</li> <li>the scoring program used to score submissions against the true test data in the evaluation and post-evaluation phase (<code>scoring_program/</code>),</li> <li>the results of the evaluation phase including <ul> <li>the final rankings of the participating teams by their best submission (<code>results/rankings_teams.csv</code>),</li> <li>the submitted files of each team (<code>results/submissions/</code>),</li> <li>an overview of the results for each submission ordered by team (<code>results/submissions_results.csv</code>),</li> <li>analysis plots (<code>plots/</code>) displaying the results: <ul> <li>under <code>per_target/</code> we provide the gold change scores and the normalized prediction error of target words plotted against their frequency and polysemy statistics,</li> <li>under <code>per_team/</code> we provide the model predictions from the best submission per team (per subtask) plotted against frequency/polysemy statistics and performance on gold data (gray lines give the correlation with the respective variable in the gold data); we also provide plots of visualizing the teams' prediction similarities.</li> </ul> </li> </ul> </li> </ul> <p>Some remarks:</p> <ul> <li>the paper referenced below remains the only source for the rankings between teams,</li> <li>some teams were disqualified, and are thus removed from the analyses and the rankings present in the paper,</li> <li>some teams have changed names, resulting in a discrepancy between team names under <code>results/</code> and team names in the paper. The paper contains a key to match old names with new names.</li> </ul> <p><strong>Test Data </strong>for SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection can be found using the links below:</p> <ul> <li><a href="https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd-eng/">English</a></li> <li><a href="https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd-ger/">German</a></li> <li><a href="https://zenodo.org/record/3734089">Latin</a></li> <li><a href="https://zenodo.org/record/3730550">Swedish</a></li> </ul> <p>Please find more information on the provided data in the paper referenced below.</p> <p><strong>Reference</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi. 2020. <a href="https://languagechange.org/semeval">SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. SemEval@COLING2020.</p> <p>The resources are freely available for education, research and other non-commercial purposes.</p> <pre><code>@inproceedings{schlechtweg2020semeval, title = "{S}em{E}val-2020 {T}ask 1: {U}nsupervised {L}exical {S}emantic {C}hange {D}etection", author = "Schlechtweg, Dominik and McGillivray, Barbara and Hengchen, Simon and Dubossarsky, Haim and Tahmasebi, Nina", booktitle = "To appear in Proceedings of the 14th International Workshop on Semantic Evaluation", year = "2020", address = "Barcelona, Spain", publisher = "Association for Computational Linguistics"}</code></pre> <p> </p>
Clouds dataset for semantic segmentation
<p>This database contains images used for training a fully convolutional neural network for the semantic segmentation of clouds in images from the Sentinel-2 Level 2A Satellite.</p> <p>The images refer to the RGB composition (bands 4, 3, and 2).</p> <p>After that, each band was converted to a byte type (0-255).</p> <p>The images are still divided into three main sets: training, validation and testing:</p> <ol> <li><strong>Training dataset: </strong>it contains 1.160 GeoTIFF images with 512x512 pixels and associated PNG masks (clouds indicated in white and background in black color).</li> <li><strong>Validation dataset</strong>: it contains 100 GeoTIFF images with 512x512 pixels and associated PNG masks used for validation step.</li> <li><strong>Test dataset: </strong>it contains 100 GeoTIFF images 512x512 pixels for testing.</li> </ol>
Brazil's South Region landslide dataset for semantic segmentation
<p>This database contains images used for the semantic segmentation of landslide scars from a fully convolutional neural network U-Net of the three States of Brazil's South Region: Rio Grande do Sul, Santa Catarina and Paraná. Each .rar file contains 3 folders:</p> <p><strong>Image: </strong>GeoTIFF 8 bits images of locations with landslide scars.</p> <p><strong>Masks: </strong>PNG masks (scars indicated in white and background in black color).</p> <p><strong>Slope: </strong>GeoTIFF float rasters of slope (in degrees).</p>
From father Busa to Linked Data. What does Thomas Aquinas have to do with the Semantic Web
<p>5<sup>th</sup> Lecture</p>
Amazon and Atlantic Forest image datasets for semantic segmentation
<p>This database contains images from<strong> Amazon </strong>and <strong>Atlantic Forest </strong>brazilian biomes used for training a fully convolutional neural network for the semantic segmentation of forested areas in images from the Sentinel-2 Level 2A Satellite.</p> <p>The images refer to the composition of bands 4, 3, 2 and 8. Each band was converted to a byte type (0-255).</p> <p>The images are still divided into three main sets: training, validation and testing:</p> <ol> <li><strong>Training dataset: </strong>it contains 499 and 485 GeoTIFF images (Amazon and Atlantic Forest, respectively) with 512x512 pixels and associated PNG masks (forest indicated in white and background in black color).</li> <li><strong>Validation dataset</strong>: it contains 100 GeoTIFF images for each biome with 512x512 pixels and associated PNG masks used for validation step.</li> <li><strong>Test dataset: </strong>it contains 20 GeoTIFF images for each biome with 512x512 pixels for testing.</li> </ol>
An appendix to the article: "'Completive' semantic component and a lexeme-based approach to actional classification in Russian" (Приложение к статье: «Семантический компонент 'комплетив' и «полексемный» подход к акциональной классификации в русском языке»)
<p>An appendix to the article “‘Completive’ semantic component and a lexeme-based approach to actional classification in Russian” by Maksim (L.) Fedotov [in Russian] (<em>Voprosy Jazykoznanija</em>, 2019, 3: 7–44. URL: <a href="http://vja.ruslang.ru/en/archive/2019-3/7-44">http://vja.ruslang.ru/en/archive/2019-3/7-44</a>. DOI: <a href="https://doi.org/10.31857/S0373658X0004896-4">10.31857/S0373658X0004896-4</a>). The appendix contains a table, in which 45 Russian perfective and imperfective verbs are analyzed by applying to them a system of tests descibed in the article and thus assigning to them an actional (aspectual) characteristic.</p> <p>Приложение к статье М. Л. Федотова «Семантический компонент ‘комплетив’ и «полексемный» подход к акциональной классификации в русском языке» (<em>Вопросы языкознания</em>, 2019, 3: 7–44. URL: <a href="http://vja.ruslang.ru/ru/archive/2019-3/7-44">http://vja.ruslang.ru/ru/archive/2019-3/7-44</a>. DOI: <a href="https://doi.org/10.31857/S0373658X0004896-4">10.31857/S0373658X0004896-4</a>). Приложение содержит таблицу с анализом 45 глаголов русского языка совершенного и несовершенного вида, к которым применяется описанная в статье система тестов, и которым на её основании присваиваются те или иные акциональные ярлыки.</p> <p>* * *</p> <p>Annotations for the article itself:</p> <p>Maksim (L.) Fedotov. <em>‘‘Completive’ semantic component and a lexeme-based approach to actional classification in Russian</em> [in Russian].</p> <p>The paper discusses two related aspectological topics. First section examines the ‘completive’ — i. e. ‘attainment of the internal limit’ — meaning (together with its counterpart ‘incompletive’, i. e. ‘non-attainment of the internal limit’). Its localization in the semantic structure of the utterance is determined: between aspect proper and actionality proper. Also, ‘completive’ can be included under the semantic scope of an iterative operator. It is argued that ‘completive’ is contained as a fixed component in the semantics of some Russian Imperfective verbs such as <em>sgorat’ </em>‘burn (down)’ and <em>pročityvat’ </em>‘read (through)’.</p> <p>Second section demonstrates practical possibility and the advantages of a single-verb approach to actional classification in Russian, an approach which is not based on the notion of aspectual pairs. Actional properties are ascribed separately to single Perfective and Imperfective verbs on the basis of uniform tests. The efficiency of the approach is demonstrated on a pilot sample of Perfective and Imperfective verbs.</p> <p>М. Л. Федотов. <em>Семантический компонент ‘комплетив’ и «полексемный» подход к акциональной классификации в русском языке</em>.</p> <p>В статье обсуждаются два связанных аспектологических сюжета. Во-первых, рассматривается значение ‘комплетив’, т. е. ‘достижение предела’ (а также парное ‘инкомплетив’, т. е. ‘недостижение предела’). Определяется его локализация в семантической структуре высказывания: между собственно аспектом и собственно акциональностью. Также ‘комплетив’ может включаться в сферу действия итеративного оператора. Аргументируется гипотеза о наличии фиксированного компонента ‘комплетив’ в семантике русских глаголов НСВ типа <em>сгорать </em>и <em>прочитывать</em>.</p> <p>Во-вторых, демонстрируется практическая возможность и преимущества «полексемного» — без апелляции к видовым парам — подхода к акциональной классификации в русском языке. Акциональные характеристики приписываются отдельным глаголам совершенного вида (СВ) и несовершенного вида (НСВ) на основании однотипных тестов. Работоспособность метода демонстрируется на тестовой выборке глаголов СВ и НСВ.</p>
GERBIL evaluation data of Robust and Collective Entity Disambiguation through Semantic Embeddings
<p>A table containing the SIGIR 2016 experiments performed with GERBIL in context of the SIGIR 2016 work "Robust and Collective Entity Disambiguation through Semantic Embeddings" by Stefan Zwicklbauer, Christin Seifert and Michael Granitzer</p> <p>It also contains the original URL to the GERBIL website</p> <p>Corresponding GitHub Repository:</p> <p>https://github.com/quhfus/</p> <p> </p>
SBML Test Suite Semantic Test Cases 3.2.0
<p>The SBML Test Suite is a conformance testing system. It allows developers and users to test the degree and correctness of the SBML support provided in a software package. A core part of the SBML Test Suite is the collection of test cases. There are 3 sets of tests: <strong>semantic</strong> (for deterministic simulation behavior), <strong>stochastic</strong>(for stochastic simulation behavior), and <strong>syntactic</strong> (for basic parsing).</p> <p>This is the version 3.2.0 release of the <strong>semantic</strong> test cases archive.</p> <p>For more information about SBML and the SBML Test Suite, please visit http://sbml.org.</p>
Modulating the assessment of semantic speech–gesture relatedness via transcranial direct current stimulation of the left frontal cortex
<p>Raw data related to the publication:</p> <p>Schülke, R., & <strong>Straube, B.</strong> (accepted). Modulating the assessment of semantic speech-gesture relatedness via transcranial direct current stimulation of the left frontal cortex. Brain Stimulation. DOI: 10.1016/j.brs.2016.10.012.</p> <p> </p> <p>Statistical software: SPSS</p> <p>Variables:</p> <p>Subject<br> Stimulus<br> SessionNr<br> Stimulation<br> Localisation - frontal/parietal/frontoparietal<br> Polarisation - anode left/right<br> Relatedness - related/unrelated<br> Gesture_type - iconic/metaphoric<br> Reaction_time - in milliseconds<br> Rating - on a scale from 1-7</p>
On the Effect of Semantically Enriched Context Models on Software Modularization
<p>The dataset used for evaluating the approaches outlined in this paper, comprising of 10 open source Java projects. The algorithms employed can be found at https://github.com/amirms/GeLaToLab, </p>
tFood: Semantic Table Annotations Benchmark for Food Domain
<p>tFood is a dataset for tabular data to knowledge graph matching. It is derived for the Food domain and has two types of tables. On the one hand, <strong>Horizontal Relational Tables</strong> are where each table represents a collection of entities. On the other hand, <strong>Entity Tables </strong>are where each of which represents a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p>The supported tasks for semantic table annotations are: </p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> </ol> <p>This dataset version will be used during SemTab 2023 - Round 1. So, the ground truth data for the test set is currently hidden. We will add such ground truth after the conclusion of the challenge. </p> <p> </p> <p> </p>
Test Dataset for 3D semantic image segmentation of the various organs from CT and MR scans
<p>These test cases are for the <a href="https://github.com/MIC-DKFZ/nnUNet/releases/tag/v1.7.1">nnUnet v1</a> models trained on the following datasets:<br><br></p> <table> <tbody> <tr> <td>Dataset </td> <td>Task</td> <td>Model Details on Zenodo</td> </tr> <tr> <td> <a href="../record/6802614">TotalSegmentator</a> and <a href="../record/5903672">FLARE21</a> datasets</td> <td>Segment Liver from CT scans</td> <td>https://zenodo.org/record/8274976</td> </tr> <tr> <td><a href="https://kits-challenge.org/kits23/">KiTS23</a> datasets and a subset of the<a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=5800386#5800386566e265abf95408aa64c4917f0cbe5d9"> TCGA-KIRC </a>dataset</td> <td>Segment Kidney, Cyst, and Tumors from CT Scans</td> <td>https://zenodo.org/records/8277846</td> </tr> <tr> <td><a href="http://ji%20yuanfeng.%20(2022).%20amos%20a%20large-scale%20abdominal%20multi-organ%20benchmark%20for%20versatile%20medical%20image%20segmentation%20[data%20set].%20zenodo.%20https">AMOS</a> and <a href="http://macdonald,%20jacob%20a.,%20zhu,%20zhe,%20konkel,%20brandon,%20mazurowski,%20maciej,%20wiggins,%20walter,%20&%20bashir,%20mustafa.%20(2020).%20duke%20liver%20dataset%20(mri)%20v2%20(2.0.0)%20[data%20set].%20zenodo.%20https//doi.org/10.5281/zenodo.7774566">DUKE Liver</a> datasets</td> <td>Segment Liver from the MR scans</td> <td>https://zenodo.org/record/8290124</td> </tr> <tr> <td>Data from m <a href="../record/6624726">pi-cai</a></td> <td>Segment Prostate region from MR scans</td> <td>https://zenodo.org/record/8290093</td> </tr> </tbody> </table>
ANR REaDY-SPOK-Semantic Representations
<p>The study investigated changes in preexisting representations stored in semantic memory after a dialogue. Researchers can find the pictures used in the set of three experiments, the data and R-scripts of the experiments, an example of confederate's script, and a new picture database. This research was supported by the French National Research Agency (ANR-19-CE28–0006).</p>
Raw data for the creation of a maturity model for Catalogues of Semantic Artefacts
<p>This dataset includes two data collections (in two different formats, i.e. CSV and XLSX) with the raw data used for creating the <a href="https://doi.org/10.5281/zenodo.10618105">Maturity Dimensions and Sub-Criteria for Catalogues of Semantic Artefacts</a>. In particular:</p> <p>1. <em>Dimension identification in literature</em> includes the list of relevant materials gathered involving all the members of the EOSC Task Force on Semantic Interoperability that include (1) definitions of semantic artefact catalogues and (2) dimensions that can be used to measure the maturity of such catalogues;</p> <p>2. <em>Catalogue assessment</em> is the result of the analysis of 26 different catalogues of semantic artefacts against the dimensions and sub-criteria described in the maturity model.</p>
Semantic Event Extraction Datasets
<p>Two datasets for Semantic Event Extraction that have been created using a distant labelling approach.</p> <p>They contain sentences from Wikipedia and events identified within them, their relations and attributes regarding the classes, properties and entities in the Wikidata and DBpedia knowledge graphs.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.