Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

29

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

29 results for “Semantic Web”

Learn how ShareScore rates datasets ↗
zenodo44/100

WikiMuTe: A web-sourced dataset of semantic descriptions for music audio

<p>This upload contains the supplementary material for our <a href="https://arxiv.org/abs/2312.09207" target="_blank" rel="noopener">paper</a> presented at the <a href="https://mmm2024.org/" target="_blank" rel="noopener">MMM2024 conference</a>.</p> <h2>Dataset</h2> <p>The dataset contains rich text descriptions for music audio files collected from Wikipedia articles.</p> <p>The audio files are freely accessible and available for download through the URLs provided in the dataset.</p> <h3>Example</h3> <p>A few hand-picked, simplified examples of the dataset.&nbsp;</p> <table> <tbody> <tr> <td> <p><strong>file</strong></p> </td> <td> <p><strong>aspects</strong></p> </td> <td> <p><strong>sentences</strong></p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/7/7a/Bongo_sound.wav" target="_blank" rel="noopener"><strong>🔈 Bongo sound.wav</strong></a></p> </td> <td> <p>['bongoes', 'percussion instrument', 'cumbia', 'drums']</p> </td> <td> <p>['a loop of bongoes playing a cumbia beat at 99 bpm']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/4/46/Example_of_double_tracking_in_a_pop-rock_song_%283_guitar_tracks%29.ogg" target="_blank" rel="noopener"><strong>🔈 Example of double tracking in a pop-rock song (3 guitar tracks).ogg</strong></a></p> </td> <td> <p>['bass', 'rock', 'guitar music', 'guitar', 'pop', 'drums']</p> </td> <td> <p>['a pop-rock song']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/6/62/OriginalDixielandJassBand-JazzMeBlues.ogg" target="_blank" rel="noopener"><strong>🔈 OriginalDixielandJassBand-JazzMeBlues.ogg</strong></a></p> </td> <td> <p>['jazz standard', 'instrumental', 'jazz music', 'jazz']</p> </td> <td> <p>['Considered to be a jazz standard', 'is an jazz composition']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/5/58/Colin_Ross_-_Etherea.ogg" target="_blank" rel="noopener"><strong>🔈 Colin Ross - Etherea.ogg</strong></a></p> </td> <td> <p>['chirping birds', 'ambient percussion', 'new-age', 'flute', 'recorder', 'single instrument', 'woodwind']</p> </td> <td> <p>['features a single instrument with delayed echo, as well as ambient percussion and chirping birds', 'a new-age composition for recorder']</p> </td> </tr> <tr> <td> <p><a href="https://upload.wikimedia.org/wikipedia/commons/8/8b/Belau_rekid_%28instrumental%29.oga" target="_blank" rel="noopener"><strong>🔈 Belau rekid (instrumental).oga</strong></a></p> </td> <td> <p>['instrumental', 'brass band']</p> </td> <td> <p>['an instrumental brass band performance']</p> </td> </tr> <tr> <td> <p><strong>...</strong></p> </td> <td> <p>...</p> </td> <td> <p>...</p> </td> </tr> </tbody> </table> <h3>Dataset structure</h3> <p>We provide three variants of the dataset in the&nbsp;<code>data</code>&nbsp;folder.</p> <p>All are described in the paper.</p> <ol> <li><code>all.csv</code> contains all the data we collected, without any filtering.</li> <li><code>filtered_sf.csv</code> contains the data obtained using the&nbsp;<em>self-filtering</em> method.</li> <li><code>filtered_mc.csv</code> contains the data obtained using the <em>MusicCaps</em>&nbsp;dataset method.</li> </ol> <h3>File structure</h3> <p>Each CSV file contains the following columns:</p> <ul> <li><code>file</code>: the name of the audio file</li> <li><code>pageid</code>: the ID of the Wikipedia article where the text was collected from</li> <li><code>aspects</code>: the short-form (tag) description texts collected from the Wikipedia articles</li> <li><code>sentences</code>: the long-form (caption) description texts collected from the Wikipedia articles</li> <li><code>audio_url</code>: the URL of the audio file</li> <li><code>url</code>: the URL of the Wikipedia article where the text was collected from</li> </ul> <h3>Citation</h3> <div> <p>If you use this dataset in your research, please cite the following paper:</p> <div> <pre><code>@inproceedings{wikimute,</code><br><code> title = {WikiMuTe: {A} Web-Sourced Dataset of Semantic Descriptions for Music Audio},</code><br><code> author = {Weck, Benno and Kirchhoff, Holger and Grosche, Peter and Serra, Xavier},</code><br><code> booktitle = "MultiMedia Modeling",</code><br><code> year = "2024",</code><br><code> publisher = "Springer Nature Switzerland",</code><br><code> address = "Cham",</code><br><code> pages = "42--56",</code><br><code> doi = {10.1007/978-3-031-56435-2_4},</code><br><code> url = {https://doi.org/10.1007/978-3-031-56435-2_4},</code><br><code>}</code></pre> </div> </div> <h3>License</h3> <p>The data is available under the&nbsp;<a href="https://creativecommons.org/licenses/by-sa/3.0/" target="_blank" rel="noopener">Creative Commons Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) license</a>.</p> <p>Each entry in the dataset contains a URL linking to the article, where the text data was collected from.</p>

opencc-by-sa-3.0Dec 2023View details →
zenodo44/100

Examining LGBTQ+-related Concepts in the Semantic Web: Link Discovery, Concept Drift, Ambiguity, and Multilingual Information Reuse

<div> <h1>Examining LGBTQ+-related Concepts in the Semantic Web</h1> </div> <div> <h2>Introduction</h2> </div> <p>Welcome to the project. We study the links between LGBTQ+ ontologies and structured vocabularies. More specifically, we focus on GSSO, Homosaurus, QLIT, and Wikidata. The code is free for use with the license GPL 3,0. You can resue/extend the code for free as long as you give credits to us in your publication/data. Citation information will be added after the corresponding paper gets accepted. The paper is under submission and will be included soon.&nbsp;</p> <p>If you would like to extend this work, you may want to contact the experts in the acknowledgement before releasing your data/code about legal and ethical issues. The DOI for this version is 10.5281/zenodo.12684870. The latest code can be found at https://github.com/Multilingual-LGBTQIA-Vocabularies/Examing_LGBTQ_Concepts.&nbsp;</p> <p>To reproduce the results or extend our work, you need to take the following steps.</p> <div> <h2>Step 1: Preparing the data</h2> </div> <p>In this project, the following datasets were used:</p> <ul> <li>QLIT: version 1.0</li> <li>Homosaurus: version 3.5 and version 2.3</li> <li>Wikidata: retrieved from the SPARQL Endpoint (<a href="https://query.wikidata.org/sparql" rel="nofollow">https://query.wikidata.org/sparql</a>) and processed between 5th May and 8th May, 2024.</li> <li>GSSO: we used gsso.owl (version 2.0.10) obtained from its Github (<a href="https://github.com/Superraptor/GSSO">https://github.com/Superraptor/GSSO</a>).</li> <li>LCSH was obtained from the official website:&nbsp;<a href="https://id.loc.gov/authorities/subjects.html" rel="nofollow">https://id.loc.gov/authorities/subjects.html</a>&nbsp;on 9th May, 2024. The LCSH data was converted to its HDT format.</li> </ul> <p>Please put the corresponding files in the following folders (and change its names where necessary) to make sure that the Python scripts can find your code.</p> <ul> <li>./data/GSSO/gsso.owl</li> <li>./data/Homosaurus/v2.ttl and ./data/Homosaurus/v3.ttl</li> <li>./data/LCSH/lcsh.hdt (we used its HDT format for fast query and analysis). The original file is also attached: subjects.skosrdf.nt.</li> <li>./data/QLIT/Qlit-v1.ttl</li> </ul> <p>The case of Wikidata is more complicated. The following scripts were used for the retrival of data. These scripts are all in the folder ./data/wikidata/</p> <ul> <li>We used the Wikidata SPARQL endpoint:&nbsp;<a href="https://query.wikidata.org/" rel="nofollow">https://query.wikidata.org/</a></li> </ul> <p>The following relations from Wikidata were used while extracting triples.</p> <ul> <li>Wikidata - GSSO:&nbsp;<a href="http://www.wikidata.org/prop/direct/P9827" rel="nofollow">http://www.wikidata.org/prop/direct/P9827</a></li> <li>Wikidata - Homosaurus 2:&nbsp;<a href="http://www.wikidata.org/prop/direct/P6417" rel="nofollow">http://www.wikidata.org/prop/direct/P6417</a></li> <li>Wikidata - Homosaurus 3:&nbsp;<a href="http://www.wikidata.org/prop/direct/P10192" rel="nofollow">http://www.wikidata.org/prop/direct/P10192</a></li> <li>Wikidata - LCSH:&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a></li> </ul> <p>The generated files are:</p> <ul> <li>'wikidata-homosaurus-v2-links.nt'</li> <li>'wikidata-homosaurus-v3-links.nt'</li> <li>'wikidata-gsso-links.nt'</li> <li>'wikidata-qlit-links.nt'</li> <li>'wikidata-lcsh-links-all.nt'</li> </ul> <p>Please note that the case of Wikdiata-LCSH is more complicated: there are so many links that are nothing to do with the entities in our scope. We restrict it to only entities in the scope of this paper. See below for more details.</p> <p>You can find all the scripts in the corresponding folder in the data folder.</p> <p>All the SPARQL queries used can be found in the folder ./SPARQL/</p> <p>Note! For GSSO, the following two mistakes were corrected while preprocessing:</p> <ul> <li><a href="https://www.wikidata.org/wiki/Q1823134" rel="nofollow">https://www.wikidata.org/wiki/Q1823134</a>&nbsp;should not be used as a relation. We have replaced it with&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a>.</li> <li>Instead of referring to the page, we refer to the entity. We use&nbsp;<a href="http://www.wikidata.org/entity/" rel="nofollow">http://www.wikidata.org/entity/</a>* instead of&nbsp;<a href="https://www.wikidata.org/wiki/" rel="nofollow">https://www.wikidata.org/wiki/</a>*</li> </ul> <p>The redirection test was conducted on 30th April, 2024, between 6PM and 8PM. The files can be found in the folder of ./data/Homosaurus/redirect/.</p> <div> <h2>Integrating the data</h2> </div> <p>In the folder ./integrated_data/, you can find all the scripts related to the integrated data. Unfortunately, due to the CC-BY-NC-ND license of GSSO and Homosaurus, the integrated data will not be made available. But you can generate it with the instructions above and by using the following scripts.</p> <p>The script ./integrated_data/integrate.py takes advantage of the data generated. It first integrates a list of files of links. Then we go through the links between Wikidata and LCSH. Only those that are in the scope of the study are included.</p> <ul> <li>If your steps are correct and using the same version as we did, you should be able to get four files:</li> <li>a) the integrated file as integrated.nt</li> <li>b) the links that are relevant for this study: wikidata-lcsh-links-selected.nt.</li> <li>c) a plot of the distribution of the size of WCCs</li> <li>d) a mapping of entities and their corresponding ID of WCCs.</li> </ul> <div> <h2>Weakly Connected Components</h2> </div> <p>The weakly connected components (WCCs) were computed for the following three purposes:</p> <p>a) Discovering missing links. See the section below for details.</p> <p>b) The WCCs can be used for manual examination. These are entities that form clusters about related concepts. The intuition is that the larger they are, the more likely there is concept drift/change, ambiguity, and mistakes.</p> <p>c) Multilingual information reuse. Smaller WCCs with exactly one entity from each dataset (e.g. Homosaurus and Wikidata) can then be used to suggest labels for the one with fewer labels for some given languages. See below for more details.</p> <p>As mentioned above, the distribution has been plotted. You can find this plot here: ./integrated_data/frequency.png</p> <p>In the folder ./integrated_data/weakly_connected_components/, you can find all the WCCs and their links.</p> <p>Two examples were given in the folder. The largest WCC about sex, gender, fucking, etc. The other is about BDSM and fetish.</p> <div> <h2>Discovering missing and outdated links</h2> </div> <p>Taking advantage of WCCs, we can further find missing and outdated links. The scripts are in the folder ./discover_missing_links.</p> <p>Three examples were given. The first two is about discovering missing links. The last one is about finding outdated links.</p> <ul> <li> <p>The script ./discover_missing_links/discover_H3_LCSH.py and ./discover_missing_links/discover_QLIT_LCSH.py are scripts that outputs links that could be missing in Homosaurus and QLIT respectively. This was computed by looking at the WCCs. If two entities are both involved in the same WCC, there could be a link between them. The csv files in the same folder are the corresponding links found.</p> </li> <li> <p>The script ./discover_missing_links/find_qlit_outdated_links/ is used to discover the outdated links between QLIT and Homosaurus v3. There was only one link found.</p> </li> <li> <p>The 105 potentially missing links were taken for further review by Swedish-speaking experts from the QLIT team, which showed that 78 (72.38%) suggested links should be included: 38 (36.19%) can be included using skos:exactMatch and another 38 (36.19%) using skos:closeMatch. 28 (26.67%) suggested links are incorrect. The manual annotation are included in the file ./discover_missing_links/Annotated_found_new_links_qlit-lcsh.xlsx.</p> </li> </ul> <div> <h2>Multilingual Information Reuse</h2> </div> <p>You can find two attempts in the folders about the use of GSSO and Wikidata for Homosaurus respectively.</p> <ul> <li>./WCC-based-gsso-multilingual_info_reuse/</li> <li>./WCC-based-wikidata-multilingual_info_reuse/</li> </ul> <p>Additionally, we provide also some code for the reuse of Wikidata multilingual info for QLIT. It's in the folder</p> <ul> <li>./WCC-based-QLIT-info-reuse-from-Wikidata/</li> </ul> <p>They follow very similar steps:</p> <ol> <li> <p>Compute the one-to-one mapping using the WCCs. The script is named compute-one-to-one-mapping.py</p> </li> <li> <p>Extract the multilingual labels from sources. The corresponding file is extract_multilingual_labels_from_one_to_one_mappings.py</p> </li> <li> <p>Provide the extracted multilingual as suggestions for targeting entities. The name of the corresponding files are like "*suggesting-labels.py", where the * is replaced by the actual source/target.</p> </li> </ol> <p>For GSSO, we use the following relations:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasExactSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasExactSynonym</a></li> <li><a href="http://purl.org/dc/terms/replaces" rel="nofollow">http://purl.org/dc/terms/replaces</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P5191" rel="nofollow">https://www.wikidata.org/wiki/Property:P5191</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P1813" rel="nofollow">https://www.wikidata.org/wiki/Property:P1813</a></li> <li><a href="https://schema.org/alternateName" rel="nofollow">https://schema.org/alternateName</a></li> <li><a href="http://www.w3.org/2002/07/owl#annotatedTarget" rel="nofollow">http://www.w3.org/2002/07/owl#annotatedTarget</a></li> </ul> <p>Additioinally, we found the relation to be studied in the future:&nbsp;<a href="http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym</a></p> <p>For Wikidata, there are only two:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.w3.org/2004/02/skos/core#altLabel" rel="nofollow">http://www.w3.org/2004/02/skos/core#altLabel</a></li> </ul> <div> <h2>Additional analysis</h2> </div> <p>Additionally, we perform an analysis using only redirection and replacement for GSSO and Homosaurus. The scripts are in the folder ./additional_test_gsso_multilingual_info_reuse. We consider also Homosaurus v2. This additional analysis shows the following:</p> <ul> <li> <p>For the Turkish language, in total there are 103 triples about labels about 23 entities. The average suggested labels per entity is 3.0.</p> </li> <li> <p>For the Spanish language, in total there are 205 triples about labels about 43 entities. The average suggested labels per entity is 2.12.</p> </li> <li> <p>For the French language, in total there are 277 triples about labels about 47 entities. The average suggested labels per entity is 2.19.</p> </li> <li> <p>For the Danish language, in total there are 115 triples about labels about 47 entities. The average suggested labels per entity is 2.70.</p> </li> </ul> <p>Some analysis about the replacement relations of Homosaurus is in the folder ./data/Homosaurus/replace_relations_homosaurus/.</p> <p>Finally, some additional analysis is included in the folder ./analysis_integrated_graph. Currently, there is only one that is about outdated entities in Homosaurus v3. Some more analysis will be added in the future.</p> <div> <h2>Acknowledgement</h2> </div> <p>The authors appreciate the help of the following researchers:</p> <ul> <li>Siska Humlesj&ouml;, QLIT, G&ouml;teborgs Universitet (<a href="mailto:siska.humlesjo@lir.gu.se">siska.humlesjo@lir.gu.se</a>)</li> <li>Olov Kristr&ouml;m, former member of QLIT</li> <li>Jack van der Wel, IHLIA (<a href="mailto:jack@ihlia.nl">jack@ihlia.nl</a>)</li> <li>Clair Kronk, GSSO (<a href="mailto:clair.kronk@mountsinai.org">clair.kronk@mountsinai.org</a>)</li> </ul> <div> <p>If you would like to extend this work, you may want to contact them before releasing your data/code about legal and ethical issues.</p> <h2>Contact</h2> </div> <ul> <li>Shuai Wang, Vrije Universiteit Amsterdam (<a href="mailto:shuai.wang@vu.nl">shuai.wang@vu.nl</a>)</li> <li>Maria Adamidou, Vrije Universiteit Amsterdam (<a href="mailto:m.adamidou@student.vu.nl">m.adamidou@student.vu.nl</a>)</li> </ul> <p>&nbsp;</p> <p>Thank you very much for your interest in our project!</p>

opengpl-3.0-or-laterJul 2024View details →
zenodo44/100

A decade of Semantic Web research through the lenses of a mixed methods approach (Resources)

<p>This work has been submitted to&nbsp;<a href="http://www.semantic-web-journal.net/content/decade-semantic-web-research-through-lenses-mixed-methods-approach">Semantic Web Journal</a>. We provide here resources to reproduce our approach.</p> <p>In this paper, we aim to provide a broader and more complete picture of Semantic Web topics and trends by adopting a mixed methods methodology, which allows a combined use of both qualitative and quantitative approaches. Concretely, we build on a qualitative analysis of the main seminal papers, which adopt a top-down approach, and on quantitative results derived with three bottom-up data-driven approaches (<a href="https://technologies.kmi.open.ac.uk/Rexplore/">Rexplore</a>, <a href="http://saffron.insight-centre.org/">Saffron</a>, <a href="https://www.poolparty.biz/">PoolParty</a>), on a corpus of Semantic Web papers published in the last decade. In this process, we both use the latter for &ldquo;fact-checking&rdquo; on the former and also to derive key findings in relation to the strengths and weaknesses of top-down and bottom-up approaches to research topic identification.</p> <p>Please access the full set of resources at:&nbsp;<a href="https://aic.ai.wu.ac.at/qadlod/SW/">https://aic.ai.wu.ac.at/qadlod/SW/</a></p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

Planet Microbe Functional and Taxonomic annotation of Illumina WGS Prokaryotic Fraction for Semantic Web Analysis

<p>Functional and Taxonomic annotations computed from a subset of Illumina Whole-Genome Sequencing samples from the prokaryotic fraction of the <a href="https://www.planetmicrobe.org/">Planet Microbe</a> database. Data was computed using the pipeline available from https://github.com/hurwitzlab/planet-microbe-functional-annotation/, and post processing scripts from https://github.com/hurwitzlab/planet-microbe-semantic-web-analysis. Files contain total annotation counts of Interpro, GO and NCBITaxon annotations, as well as additional sample metadata. See readme.txt file for more information.</p>

opencc-zeroJan 2022View details →
zenodo44/100

RDF dataset produced in the work "Exploring Adverse Outcome Pathways for Nanomaterials with semantic web technologies"

<p>Adverse Outcome Pathways (AOPs) have been proposed to facilitate mechanistic understanding of interactions of chemicals/materials with biological systems. Each AOP starts with a molecular initiating event (MIE) and possibly ends with adverse outcome(s) (AOs) via a series of key events (KEs). So far, the interaction of engineered nanomaterials (ENMs) with biomolecules, biomembranes, cells, and biological structures, in general, is not yet fully elucidated. There is also a huge lack of information on which AOPs are ENMs-relevant or -specific, despite numerous published data on toxicological endpoints they trigger, such as oxidative stress and inflammation. We propose to integrate related data and knowledge recently collected. Our approach combines the annotation of nanomaterials and their MIEs with ontology annotation to demonstrate how we can then query AOPs and biological pathway information for these materials. We conclude that a FAIR (Findable, Accessible, Interoperable, Reusable) representation of the ENM-MIE knowledge simplifies integration with other knowledge.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

A core ontology for modeling life cycle sustainability assessment on the Semantic Web with Accompanying Database

<p>To enable and support the uptake of semantic ontologies, we present a core ontology developed specifically to capture the data relevant for life cycle sustainability assessment. We further demonstrate the utility of the ontology by using it to integrate data relevant to sustainability assessments, such as EXIOBASE and the Yale Stocks and Flow Database to the Semantic Web. These datasets can be accessed by the machine-readable endpoint using SPARQL, a semantic query language.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Accompanying Dataset migr_asyappctzm for Efficient Analytical Queries on Semantic Web Data Cubes

<p>This dataset&nbsp; shows how the Eurostat data cube in the orginal publicatin is modelled in QB4OLAP.</p> <p>This data is based on statistical data about asylum applications to the European Union, provided by Eurostat on</p> <p><a href="http://ec.europa.eu/eurostat/web/products-datasets/-/migr_asyappctzm">http://ec.europa.eu/eurostat/web/products-datasets/-/migr_asyappctzm</a></p> <p>Further data has been integrated from: https://github.com/lorenae/qb4olap/tree/master/examples</p>

opencc-by-4.0Oct 2017View details →
zenodo40/100

Supplementary Data for "Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"

<p>This dataset contains quality assessment results for 26 vocabularies. The assessment was conducted using the <a href="https://skos-play.sparna.fr/skos-testing-tool/">qSKOS vocabulary quality assessment tool</a>.</p> <p>The 26 assessed vocabularies were converted from their original formats into the Simple Knowledge Organization System (SKOS) data model using the approach described in our paper titled <a href="https://doi.org/10.1007/978-3-031-62362-2_9">"Streamlining Vocabulary Conversion to SKOS: A YAML-based Approach to Facilitate Participation in the Semantic Web"</a>, presented at the <a href="https://doi.org/10.1007/978-3-031-62362-2">24th International Conference on Web Engineering (ICWE 2024)</a>.</p> <p>The dataset contains a quality assessment for the following vocabularies:</p> <ol> <li>A Taxonomy of Evaluation Towards Standards</li> <li>Cross-Device Taxonomy</li> <li>What Makes a Data-driven Business Model? A Consolidated Taxonomy</li> <li>DDI Aggregation Method</li> <li>DDI Mode of Collection</li> <li>Building a New Taxonomy for Data Discretization Techniques</li> <li>Demopaedia</li> <li>Data Science Glossary</li> <li>A Taxonomy of Evaluation Approaches in Software Engineering</li> <li>Evaluation Thesaurus</li> <li>The Glossary of Human Computer Interaction</li> <li>Human-Factors Taxonomy</li> <li>A Taxonomy to Structure and Analyze Human&ndash;Robot Interaction</li> <li>A Taxonomy of Interaction for Instructional Multimedia</li> <li>A Taxonomy of Interrogation Methods</li> <li>Design Vocabulary for Human&ndash;IoT Systems Communication</li> <li>Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors</li> <li>Thesaurus Mass Communication</li> <li>Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey</li> <li>A Taxonomy of Quality of Service and Quality of Experience of Multimodal Human-Machine Interaction</li> <li>A Human-Centered Taxonomy of Interaction Modalities and Devices</li> <li>A Taxonomy of Spatial Interaction Patterns and Techniques</li> <li>A Taxonomy of Social Errors in Human-Robot Interaction</li> <li>Taxonomy of Digital Research Activities in the Humanities</li> <li>Virtual Reality and the CAVE: Taxonomy, Interaction Challenges and Research Directions&nbsp;</li> <li>Cross-Device Interaction</li> </ol>

opencc-by-4.0Feb 2024View details →
zenodo40/100

SemTab 2024: Semantic Web Challenge on Tabular Data to Knowledge Graph Matching Data Sets - WikidataTables2024R1 and WikidataTables2024R2

<p>Data Sets from the ISWC 2024 Semantic Web Challenge on Tabular Data to Knowledge Graph Matching, Round 1, Wikidata Tables. Links to other datasets can be found on the challenge website: https://sem-tab-challenge.github.io/2024/ as well as the proceedings of the challenge published on CEUR.</p> <p>For details about the challenge, see: http://www.cs.ox.ac.uk/isg/challenges/sem-tab/</p> <p>For 2024 edition, see: https://sem-tab-challenge.github.io/2024/</p> <p>Note on License: This data includes data from the following sources. Refer to each source for license details:<br>- Wikidata&nbsp;https://www.wikidata.org/</p> <p>THIS DATA IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Responses to "Semantic Web: Perspectives" Questionnaire

<p>The dataset provides material relating to a questionnaire entitled &quot;Semantic Web: Perspectives&quot;. This questionnaire was addressed to the W3C Semantic Web mailing list (semantic-web@w3.org) and was open to responses from May 12th to May 25th, 2019. A total of 113 responses were collected in this time. The following files are provided:</p> <ul> <li><strong>public-comments.txt:</strong>&nbsp;provides the public comments of respondents in plain text;</li> <li><strong>questionnaire-form.pdf:</strong>&nbsp;illustrates the design of the questionnaire, including questions, types of responses permitted, etc.;</li> <li><strong>questionnaire-responses.tsv:</strong>&nbsp;lists the individual responses (without private comments) as a tab-separated values file;</li> <li><strong>success-keywords.xlsx:</strong>&nbsp;provides a spreadsheet mapping success story responses to a list of keywords, further providing statistics on these keywords;</li> <li><strong>wordcloud-bw.svg:</strong>&nbsp;provides a word-cloud of success-story keywords in black &amp; white;</li> <li><strong>wordcloud-colour.svg:</strong>&nbsp;provides a word-cloud of success-story keywords in colour.</li> </ul> <p>The word-clouds were produced using <a href="https://www.jasondavies.com/wordcloud/">Jason Davies&#39; online service</a>, copying and pasting the keywords from the&nbsp;&nbsp;<strong>success-keywords.xlsx</strong>&nbsp;spreadsheet (e.g., Column A, Sheet Statistics) into the text field; the following settings were selected: Orientations from 0&deg; to 0&deg;, Spiral: Rectangular;&nbsp;Scale: n;&nbsp;Number of words: 400;&nbsp;One word per line: ticked; Font: Patua One (must be installed locally beforehand). The resulting SVG files were later modified in a text editor to add a link to the font used, to tighten the bounding box, and to produce a black &amp; white version.</p> <p>We thank the respondents for providing their input.</p>

opencc-by-4.0May 2019View details →
zenodo36/100

From father Busa to Linked Data. What does Thomas Aquinas have to do with the Semantic Web

<p>5<sup>th</sup> Lecture</p>

opencc-by-4.0Jul 2018View details →
zenodo36/100

Mining the UK Web Archive for Semantic Change Detection (Dataset)

<p>The dataset that was used and released with the RANLP 2019 paper, titled &quot;Mining the UK Web Archive for Semantic Change Detection&quot; (see&nbsp;<a href="https://github.com/adtsakal/Semantic_Change">https://github.com/adtsakal/Semantic_Change</a>).&nbsp;It contains annual word2vec representations of more than 47K words over the period 2000-2013, along with a list of 65 words with known semantic change over the same time period.&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

Semantic Web resources and Machine Learning systems - Knowledge Graph (SWeMLS-KG)

<p>This resource is part of our submission to ESWC 2023 resource track, which includes:</p> <p>Datasets:<br> - Folder &quot;pattern&quot; - a set of SWeMLS patterns represented based on OPMW and P-Plan ontology,<br> - Folder &quot;shapes&quot; - a set of SHACL constraints to check the conformance of SWeML Systems against SWeMLS patterns as well as a set of SHACL-AF rules to generate links between system components,<br> - File &quot;swemls-ontology.ttl&quot; - an ontology to represent Semantic Web resources and Machine Learning systems (SWeMLS),<br> - File &quot;swemls-instances.ttl&quot; - a set of triples representing the extracted metadata from 476 SWeML systems and papers,<br> - File &quot;swemls-kg.ttl&quot; - an integrated and validated KG containing all above files, including enrichment from SHACL-AF rules using &quot;swemls-toolkit&quot; [2].</p> <p>These resources are produced based on the result of the Systematic Mapping Study (SMS) reported in [1]. The latest SNAPSHOT-version of the resource can be accessed through our resource landing page: <a href="https://w3id.org/semsys/sites/swemls-kg/">https://w3id.org/semsys/sites/swemls-kg/</a></p> <p>[1] Breit, A., Waltersdorfer, L., Ekaputra, J.F., Sabou, M., Ekelhart, A., Iana, A., Paulheim, H., Portisch, J., Revenko, A., Ten Teije, A., van Harmelen, F.: Combining Machine Learning and Semantic Web -A Systematic Mapping Study (under review). ACM CSUR (2022)<br> [2] Source code of swemls-toolkit is available at: https://github.com/semanticsystems/swemls-toolkit</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Piveau: A Large-scale Open Data Management Platform based on Semantic Web Technologie

<p>This file contains the sources that were used to create the feature comparison in &quot;Piveau: A Large-scale Open Data Management Platform based on Semantic Web Technologies&quot;.</p>

opencc-by-4.0Dec 2019View details →
zenodo32/100

Semantic Web und Linked Data: Generierung von Interoperabilität in archäologischen Fachdaten am Beispiel römischer Töpferstempel - Datasets

<p><strong>Datasets</strong></p> <p>Gegenstand der Masterarbeit ist die Verwendung aktueller Technologien interoperabler Datenhaltung, insbesondere das Konzept der Linked Open Data (LOD) und der semantischen Modellierung, zur Verdeutlichung ihres Potentials in arch&auml;ologischen Informationen am Beispiel von Terra Sigillata-Fundorten, -T&ouml;pfern und -Keramikfragmenten. Die Arbeit zeigt eine Migration von Daten, sowie die M&ouml;glichkeiten und die Problematik der Modellierung der Attribute und Beziehungen mit Hilfe bestehender LOD-Konzepte und kontrollierter Vokabularien, sowie eigene Ans&auml;tze zur L&ouml;sung. Diese Daten werden mittels REST-Schnittstelle zur Verf&uuml;gung gestellt. Ein Schwerpunkt wird auf die Verlinkung zu anderen bereits bestehenden Projekten gelegt, wodurch eine Vielzahl weiterer arch&auml;ologischer und historischer Informationen z.B. &uuml;ber das Pelagios Projekt eingebunden werden. Zudem wird das Potential der Verlinkung und Abfrage von heterogenen Informationen zwischen T&ouml;pfern, Fragmenten und Orten deren relativ chronologische Beziehungen &uuml;ber LOD mit einer webbasierten Schnittstelle aufgezeigt.</p> <p>The subject matter of this master thesis is using current technologies in interoperable data management, in particular the illustration of the potential of Linked Open Data (LOD) and semantic modelling in archaeological information, as used on samian ware places and their corresponding potters and ceramic fragments. The thesis demonstrates a migration of data as well as possibilities and problems of modelling attributes and relationships using existing LOD concepts and controlled vocabularies as well as novel self-developed approaches to the solution. These data are provided by a ReST interface. One focus is linking to other existing projects, creating associations to other archaeological and historical information, for example the Pelagios project. Moreover, a web-based interface shows the potential of linking and retrieval of heterogeneous information among pottery, fragments and places and their relative chronological relationships via LOD.</p>

opencc-by-4.0Dec 2013View details →
zenodo32/100

Processed data for the "Deriving Semantics-Aware Fuzzers from Web API Schemas" paper

<p>Processed data for the &quot;Deriving Semantics-Aware Fuzzers from Web API Schemas&quot; paper.&nbsp; Each directory in the archive consists of:</p> <p>- metadata.json. Metadata about a test run - tested fuzzer name, run duration, etc</p> <p>- fuzzer.json&nbsp;- Structured fuzzer output</p> <p>-&nbsp;deduplicated_cases.json - Deduplicated reported failures, when fuzzers provide it</p> <p>- sentry.json&nbsp;- Cleaned Sentry events for this run</p> <p>- target.json&nbsp;- Parsed stdout for Gitlab &amp; Disease.sh targets that were tested without Sentry integration</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Aesthetic Trends and Semantic Web Adoption of Media Outlets Identified through Automated Archival Data Extraction

<p>This dataset includes a variety of structured data gathered via various Web data extraction techniques which were employed in order to collect current and archival data from almost a thousand news websites that are popular in Greece, for the purpose of monitoring and recording their progress through time. The collected information, that took the form of a website&rsquo;s source code and an impression of their homepage in different time instances of the last decade, has been used to identify trends concerning Semantic Web integration, DOM structure complexity, number of graphics, color usage and more. In total more than ten thousands impressions (including screenshots and source code) were analyzed which resulted to conclusions regarding the evolution of aesthetics and the adoption of new technologies.</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

Figure 5 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects

Figure 5. The Wallich Catalogue. Screenshot of Wallich Catalogue hosted by Royal Botanic Garden Edinburgh showing popup for stable URI containing information hosted at Botanic Garden and Botanical Museum Berlin-Dahlem.

opennotspecifiedOct 2017View details →
zenodo32/100

Figure 4 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects

Figure 4. The CETAF Specimen URI Tester provides for any given Specimen URI an overview of the redirection process as well as a preview of machine-readable and human-readable data associated with the URI.

opennotspecifiedOct 2017View details →
zenodo32/100

Figure 3 in Actionable, long-term stable and semantic web compatible identifiers for access to biological collection objects

Figure 3. CETAF stable HTTP URIs in the GBIF data portal. The Global Biodiversity Information Facility (GBIF) publishes CETAF stable HTTP URIs via their data portal.

opennotspecifiedOct 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record