Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

483

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

483 results for “SEMANTICS”

Learn how ShareScore rates datasets ↗
zenodo40/100

Semantic Similarity of IT Support Tickets

<p>Collection of 300 support tickets manually labeled for semantic similarity, obtained from a IT support company in the Florian&oacute;polis (Brazil) region. Each ticket is represented by an unstructured text field, which is typed by the user that opened the call. The labeling process was performed in 2022 by three IT support professionals. The corpus contains tickets in many languages, mainly English, German, Portuguese and Spanish.</p> <p>All Personal Identifiable Information (PII) and sensitive information were removed (substituted by a tag indicating the original content, for instance: the sentence &quot;this text was written by Leonardo&quot; is converted to &quot;this text was written by [NAME]&quot;). The removal was performed in three steps: first, the automated machine learning-based tool AWS Comprehend PII Removal was used; then, a sequence of custom regular expressions was applied; last, the entire corpus was manually verified.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

PDF4 - Semantics and Vocabularies Hackathon SWOT Visualization

<p>This SWOT visualization outlines the community&#39;s feedback regarding the overall polar RDM strengths, weaknesses, opportunities, and threats. This visualization is valuable to convey the overarching themes/ topics of interest, and allows for stakeholders to determine where resources should be allocated. The content for this visualization was compiled through community input during the &#39;Semantics and Vocabularies&#39; hackathon at the 4th Polar Data Forum in September 2021.&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

WE3DS: An RGB-D image dataset for semantic segmentation in agriculture

<p>Here, we introduce a novel RGB-D image database (WE3DS) for semantic segmentation in crop farming. It contains 2,568 RGB-D images (color image and distance map) and hand-annotated ground-truth masks for semantic segmentation and is the first RGB-D image dataset for multi-class plant species semantic segmentation task. Images were taken under natural light conditions using an RGB-D sensor consisting of two RGB cameras in a stereo setup.</p> <p>&nbsp;</p> <p><strong>Please cite the original source when using this dataset.</strong></p> <p>Kitzler, F.; Barta, N.; Neugschwandtner, R.W.; Gronauer, A.; Motsch, V. WE3DS: An RGB-D Image Dataset for Semantic Segmentation in Agriculture. <em>Sensors</em> <strong>2023</strong>, <em>23</em>, 2713. <a href="https://doi.org/10.3390/s23052713">https://doi.org/10.3390/s23052713 </a></p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

The evolution of lexical semantics dynamics, directionality, and drift: S4

<p>Supplementary Material (S4) for the study &quot;The evolution of lexical semantics dynamics, directionality, and drift&quot;, for Frontiers in Communication, Special Issue &quot;<a href="https://www.frontiersin.org/research-topics/38650/the-evolution-of-meaning-challenges-in-quantitative-lexical-typology?fbclid=IwAR3AeXx_11P-8CZG0UavlOgqvbVo6MlMxX8AejtPiREAKmeLyy3LIB3Ux24">The Evolution of Meaning: Challenges in Quantitative Lexical Typology</a>&quot;, ed. Gerd Carling &amp; Annemarie Verkerk</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

SPVPANELEX: Dataset containing aerial orthoimages (covering 257.93 km2 of the Spanish territory, with a spatial resolution of 0.5 m) labelled with photovoltaic panel information for binary recognition and semantic segmentation

<p>The data have been generated using scripts developed in Python with Open-Source libraries (GDAL/OGR and MapScript) to rasterize of vector cartography representing the photovoltaic (PV) panels instalations in urban, industrial, and rural areas. This PV panels cartography has been generated by manual digitalizing the PV panels found latest aerial orthofotographs available on June 1, 2021 from Plano Nacional de Ortofotograf&iacute;a A&eacute;rea (PNOA), produced by the National Geographic Institute of Spain, using the Web Map Service PNOA-MA.<br> <br> The dataset consists of 239,680 images of 256 &times; 256 pixels in size, in png format, labelled with Class_1: &ldquo;Contains PV panel&rdquo; and Class_2: &ldquo;Does not contain PV panel&rdquo;, that were pre-divided with a split criterion of 70:10:20%. in train, validation and test folders, respectively.<br> <br> The structure of the data is as follows:<br> 1-Panels-Ortho and 1-Panels-Mask contain the images featuring PV panels and their corresponding ground truth mask for training the semantic segmentation networks.<br> 1-Panels-Ortho and 2-NoPanels-Ortho contain images containing and not containing PV panels, for the training of binary recognition models of PV panels.<br> <br> Moreover, in each folder the structure is the same: train, test, validation containing 70%, 10% and 20% of the total images and masks of each type.<br> <br> 1-Panels-Ortho<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation<br> <br> 1-Panels-Mask<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation<br> <br> 2-NoPanels-Ortho<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

A Semantically Annotated 15-Class Ground Truth Dataset for Substation Equipment

<p>This dataset contains 1660 images of electric substations with 50705 annotated objects. The images were obtained using different cameras, including cameras mounted on Autonomous Guided Vehicles (AGVs), fixed location cameras and those captured by humans using a variety of cameras. A total of 15 classes of objects were identified in this dataset, and the number of instances for each class is provided in the following table:</p> <table align="center"> <caption>Object classes and how many times they appear in the dataset.</caption> <thead> <tr> <th scope="col">Class</th> <th scope="col">Instances</th> </tr> </thead> <tbody> <tr> <td>Open blade disconnect</td> <td>310</td> </tr> <tr> <td>Closed blade disconnect switch</td> <td>5243</td> </tr> <tr> <td>Open tandem disconnect switch</td> <td>1599</td> </tr> <tr> <td>Closed tandem disconnect switch</td> <td>966</td> </tr> <tr> <td>Breaker</td> <td>980</td> </tr> <tr> <td>Fuse disconnect switch</td> <td>355</td> </tr> <tr> <td>Glass disc insulator</td> <td>3185</td> </tr> <tr> <td>Porcelain pin insulator</td> <td>26499</td> </tr> <tr> <td>Muffle</td> <td>1354</td> </tr> <tr> <td>Lightning arrester</td> <td>1976</td> </tr> <tr> <td>Recloser</td> <td>2331</td> </tr> <tr> <td>Power transformer</td> <td>768</td> </tr> <tr> <td>Current transformer</td> <td>2136</td> </tr> <tr> <td>Potential transformer</td> <td>654</td> </tr> <tr> <td>Tripolar disconnect switch</td> <td>2349</td> </tr> </tbody> </table> <p>All images in this dataset were collected from a single electrical distribution substation in Brazil over a period of two years. The images were captured at various times of the day and under different weather and seasonal conditions, ensuring a diverse range of lighting conditions for the depicted objects. A team of experts in Electrical Engineering curated all the images to ensure that the angles and distances depicted in the images are suitable for automating inspections in an electrical substation.</p> <p>The file structure of this dataset contains the following directories and files:</p> <p>&nbsp;images: This directory contains 1660 electrical substation images in JPEG format.</p> <p>images: This directory contains 1660 electrical substation images in JPEG format.</p> <ul> <li><strong>labels_json: </strong>This directory contains JSON files annotated in the VOC-style polygonal format. Each file shares the same filename as its respective image in the images directory.</li> <li><strong>15_masks:</strong> This directory contains PNG segmentation masks for all 15 classes, including the porcelain pin insulator class. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>14_masks:</strong> This directory contains PNG segmentation masks for all classes except the porcelain pin insulator. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>porcelain_masks:</strong> This directory contains PNG segmentation masks for the porcelain pin insulator class. Each file shares the same name as its corresponding image in the images directory.</li> <li><strong>classes.txt:</strong> This text file lists the 15 classes plus the background class used in LabelMe.</li> <li><strong>json2png.py:</strong> This Python script can be used to generate segmentation masks using the VOC-style polygonal JSON annotations.</li> </ul> <p>The dataset aims to support the development of computer vision techniques and deep learning algorithms for automating the inspection process of electrical substations. The dataset is expected to be useful for researchers, practitioners, and engineers interested in developing and testing object detection and segmentation models for automating inspection and maintenance activities in electrical substations.</p> <p>The authors would like to thank UTFPR for the support and infrastructure made available for the development of this research and COPEL-DIS for the support through project PD-2866-0528/2020&mdash;Development of a Methodology for Automatic Analysis of Thermal Images. We also would like to express our deepest appreciation to the team of annotators who worked diligently to produce the semantic labels for our dataset. Their hard work, dedication and attention to detail were critical to the success of this project.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Taxonomies for Semantic Research Data Annotation

<p>This dataset contains 35 of 39 taxonomies that were the result of a systematic review. The systematic review was conducted with the goal of identifying taxonomies suitable for semantically annotating research data. A special focus was set on research data from the hybrid societies domain.</p> <p>The following taxonomies were identified as part of the systematic review:</p> <table> <tbody> <tr> <td> <p><strong>Filename</strong></p> </td> <td> <p><strong>Taxonomy Title</strong></p> </td> </tr> <tr> <td> <p>acm_ccs</p> </td> <td> <p>ACM Computing Classification System [1]</p> </td> </tr> <tr> <td> <p>amec</p> </td> <td> <p>A Taxonomy of Evaluation Towards Standards [2]</p> </td> </tr> <tr> <td> <p>bibo</p> </td> <td> <p>A BIBO Ontology Extension for Evaluation of Scientific Research Results [3]</p> </td> </tr> <tr> <td> <p>cdt</p> </td> <td> <p>Cross-Device Taxonomy [4]</p> </td> </tr> <tr> <td> <p>cso</p> </td> <td> <p>Computer Science Ontology [5]</p> </td> </tr> <tr> <td> <p>ddbm</p> </td> <td> <p>What Makes a Data-driven Business Model? A Consolidated Taxonomy [6]</p> </td> </tr> <tr> <td> <p>ddi_am</p> </td> <td> <p>DDI Aggregation Method [7]</p> </td> </tr> <tr> <td> <p>ddi_moc</p> </td> <td> <p>DDI Mode of Collection [8]</p> </td> </tr> <tr> <td> <p>n/a</p> </td> <td> <p>DemoVoc [9]</p> </td> </tr> <tr> <td> <p>discretization</p> </td> <td> <p>Building a New Taxonomy for Data Discretization Techniques [10]</p> </td> </tr> <tr> <td> <p>dp</p> </td> <td> <p>Demopaedia [11]</p> </td> </tr> <tr> <td> <p>dsg</p> </td> <td> <p>Data Science Glossary [12]</p> </td> </tr> <tr> <td> <p>ease</p> </td> <td> <p>A Taxonomy of Evaluation Approaches in Software Engineering [13]</p> </td> </tr> <tr> <td> <p>eco</p> </td> <td> <p>Evidence &amp; Conclusion Ontology [14]</p> </td> </tr> <tr> <td> <p>edam</p> </td> <td> <p>EDAM: The Bioscientific Data Analysis Ontology [15]</p> </td> </tr> <tr> <td> <p>n/a</p> </td> <td> <p>European Language Social Science Thesaurus [16]</p> </td> </tr> <tr> <td> <p>et</p> </td> <td> <p>Evaluation Thesaurus [17]</p> </td> </tr> <tr> <td> <p>glos_hci</p> </td> <td> <p>The Glossary of Human Computer Interaction [18]</p> </td> </tr> <tr> <td> <p>n/a</p> </td> <td> <p>Humanities and Social Science Electronic Thesaurus [19]</p> </td> </tr> <tr> <td> <p>hcio</p> </td> <td> <p>A Core Ontology on the Human-Computer Interaction Phenomenon [20]</p> </td> </tr> <tr> <td> <p>hft</p> </td> <td> <p>Human-Factors Taxonomy [21]</p> </td> </tr> <tr> <td> <p>hri</p> </td> <td> <p>A Taxonomy to Structure and Analyze Human&ndash;Robot Interaction [22]</p> </td> </tr> <tr> <td> <p>iim</p> </td> <td> <p>A Taxonomy of Interaction for Instructional Multimedia [23]</p> </td> </tr> <tr> <td> <p>interrogation</p> </td> <td> <p>A Taxonomy of Interrogation Methods [24]</p> </td> </tr> <tr> <td> <p>iot</p> </td> <td> <p>Design Vocabulary for Human&ndash;IoT Systems Communication [25]</p> </td> </tr> <tr> <td> <p>kinect</p> </td> <td> <p>Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors [26]</p> </td> </tr> <tr> <td> <p>maco</p> </td> <td> <p>Thesaurus Mass Communication [27]</p> </td> </tr> <tr> <td> <p>n/a</p> </td> <td> <p>Thesaurus Cognitive Psychology of Human Memory [28]</p> </td> </tr> <tr> <td> <p>mixed_initiative</p> </td> <td> <p>Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey [29]</p> </td> </tr> <tr> <td> <p>qos_qoe</p> </td> <td> <p>A Taxonomy of Quality of Service and Quality of Experience of Multimodal Human-Machine Interaction [30]</p> </td> </tr> <tr> <td> <p>ro</p> </td> <td> <p>The Research Object Ontology [31]</p> </td> </tr> <tr> <td> <p>senses_sensors</p> </td> <td> <p>A Human-Centered Taxonomy of Interaction Modalities and Devices [32]</p> </td> </tr> <tr> <td> <p>sipat</p> </td> <td> <p>A Taxonomy of Spatial Interaction Patterns and Techniques [33]</p> </td> </tr> <tr> <td> <p>social_errors</p> </td> <td> <p>A Taxonomy of Social Errors in Human-Robot Interaction [34]</p> </td> </tr> <tr> <td> <p>sosa</p> </td> <td> <p>Semantic Sensor Network Ontology [35]</p> </td> </tr> <tr> <td> <p>swo</p> </td> <td> <p>The Software Ontology [36]</p> </td> </tr> <tr> <td> <p>tadirah</p> </td> <td> <p>Taxonomy of Digital Research Activities in the Humanities [37]</p> </td> </tr> <tr> <td> <p>vrs</p> </td> <td> <p>Virtual Reality and the CAVE: Taxonomy, Interaction Challenges and Research Directions&nbsp; [38]</p> </td> </tr> <tr> <td> <p>xdi</p> </td> <td> <p>Cross-Device Interaction [39]</p> </td> </tr> </tbody> </table> <p><br> We converted the taxonomies into SKOS&nbsp;(Simple Knowledge Organisation System) representation. The following 4 taxonomies were not converted as they were already available in SKOS and were for this reason excluded from this dataset:</p> <p>1) DemoVoc, cf. <a href="http://thesaurus.web.ined.fr/navigateur/">http://thesaurus.web.ined.fr/navigateur/</a><br> available at <a href="https://thesaurus.web.ined.fr/exports/demovoc/demovoc.rdf">https://thesaurus.web.ined.fr/exports/demovoc/demovoc.rdf</a></p> <p>2) European Language Social Science Thesaurus, cf. <a href="https://thesauri.cessda.eu/elsst/en/">https://thesauri.cessda.eu/elsst/en/</a><br> available at <a href="https://zenodo.org/record/5506929">https://zenodo.org/record/5506929</a></p> <p>3) Humanities and Social Science Electronic Thesaurus, cf. <a href="https://hasset.ukdataservice.ac.uk/hasset/en/">https://hasset.ukdataservice.ac.uk/hasset/en/</a><br> available at <a href="https://zenodo.org/record/7568355">https://zenodo.org/record/7568355</a></p> <p>4) Thesaurus Cognitive Psychology of Human Memory, cf. <a href="https://www.loterre.fr/presentation/">https://www.loterre.fr/presentation/</a><br> available at <a href="https://skosmos.loterre.fr/P66/en/">https://skosmos.loterre.fr/P66/en/</a></p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>[1]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &ldquo;The 2012 ACM Computing Classification System,&rdquo; <em>ACM Digital Library</em>, 2012. <a href="https://dl.acm.org/ccs">https://dl.acm.org/ccs</a> (accessed May 08, 2023).</p> <p>[2]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; AMEC, &ldquo;A Taxonomy of Evaluation Towards Standards.&rdquo; Aug. 31, 2016. Accessed: May 08, 2023. [Online]. Available: <a href="https://amecorg.com/amecframework/home/supporting-material/taxonomy/">https://amecorg.com/amecframework/home/supporting-material/taxonomy/</a></p> <p>[3]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; B. Dimić Surla, M. Segedinac, and D. Ivanović, &ldquo;A BIBO ontology extension for evaluation of scientific research results,&rdquo; in <em>Proceedings of the Fifth Balkan Conference in Informatics</em>, in BCI &rsquo;12. New York, NY, USA: Association for Computing Machinery, Sep. 2012, pp. 275&ndash;278. doi: <a href="https://doi.org/10.1145/2371316.2371376">10.1145/2371316.2371376</a>.</p> <p>[4]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; F. Brudy <em>et al.</em>, &ldquo;Cross-Device Taxonomy: Survey, Opportunities and Challenges of Interactions Spanning Across Multiple Devices,&rdquo; in <em>Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems</em>, in CHI &rsquo;19. New York, NY, USA: Association for Computing Machinery, Mai 2019, pp. 1&ndash;28. doi: <a href="https://doi.org/10.1145/3290605.3300792">10.1145/3290605.3300792</a>.</p> <p>[5]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A. A. Salatino, T. Thanapalasingam, A. Mannocci, F. Osborne, and E. Motta, &ldquo;The Computer Science Ontology: A Large-Scale Taxonomy of Research Areas,&rdquo; in <em>Lecture Notes in Computer Science 1137</em>, D. Vrandečić, K. Bontcheva, M. C. Su&aacute;rez-Figueroa, V. Presutti, I. Celino, M. Sabou, L.-A. Kaffee, and E. Simperl, Eds., Monterey, California, USA: Springer, Oct. 2018, pp. 187&ndash;205. Accessed: May 08, 2023. [Online]. Available: <a href="http://oro.open.ac.uk/55484/">http://oro.open.ac.uk/55484/</a></p> <p>[6]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; M. Dehnert, A. Gleiss, and F. Reiss, &ldquo;What makes a data-driven business model? A consolidated taxonomy,&rdquo; presented at the European Conference on Information Systems, 2021.</p> <p>[7]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; DDI Alliance, &ldquo;DDI Controlled Vocabulary for Aggregation Method,&rdquo; 2014. <a href="https://ddialliance.org/Specification/DDI-CV/AggregationMethod_1.0.html">https://ddialliance.org/Specification/DDI-CV/AggregationMethod_1.0.html</a> (accessed May 08, 2023).</p> <p>[8]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; DDI Alliance, &ldquo;DDI Controlled Vocabulary for Mode Of Collection,&rdquo; 2015. <a href="https://ddialliance.org/Specification/DDI-CV/ModeOfCollection_2.0.html">https://ddialliance.org/Specification/DDI-CV/ModeOfCollection_2.0.html</a> (accessed May 08, 2023).</p> <p>[9]&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; INED - French Institute for Demographic Studies, &ldquo;Th&eacute;saurus DemoVoc,&rdquo; Feb. 26, 2020. <a href="https://thesaurus.web.ined.fr/navigateur/en/about">https://thesaurus.web.ined.fr/navigateur/en/about</a> (accessed May 08, 2023).</p> <p>[10]&nbsp;&nbsp;&nbsp; A. A. Bakar, Z. A. Othman, and N. L. M. Shuib, &ldquo;Building a new taxonomy for data discretization techniques,&rdquo; in <em>2009 2nd Conference on Data Mining and Optimization</em>, Oct. 2009, pp. 132&ndash;140. doi: <a href="https://doi.org/10.1109/DMO.2009.5341896">10.1109/DMO.2009.5341896</a>.</p> <p>[11]&nbsp;&nbsp;&nbsp; N. Brouard and C. Giudici, &ldquo;Unified second edition of the Multilingual Demographic Dictionary (Demopaedia.org project),&rdquo; presented at the 2017 International Population Conference, IUSSP, Oct. 2017. Accessed: May 08, 2023. [Online]. Available: <a href="https://iussp.confex.com/iussp/ipc2017/meetingapp.cgi/Paper/5713">https://iussp.confex.com/iussp/ipc2017/meetingapp.cgi/Paper/5713</a></p> <p>[12]&nbsp;&nbsp;&nbsp; DuCharme, Bob, &ldquo;Data Science Glossary.&rdquo; https://www.datascienceglossary.org/ (accessed May 08, 2023).</p> <p>[13]&nbsp;&nbsp;&nbsp; A. Chatzigeorgiou, T. Chaikalis, G. Paschalidou, N. Vesyropoulos, C. K. Georgiadis, and E. Stiakakis, &ldquo;A Taxonomy of Evaluation Approaches in Software Engineering,&rdquo; in <em>Proceedings of the 7th Balkan Conference on Informatics Conference</em>, in BCI &rsquo;15. New York, NY, USA: Association for Computing Machinery, Sep. 2015, pp. 1&ndash;8. doi: <a href="https://doi.org/10.1145/2801081.2801084">10.1145/2801081.2801084</a>.</p> <p>[14]&nbsp;&nbsp;&nbsp; M. C. Chibucos, D. A. Siegele, J. C. Hu, and M. Giglio, &ldquo;The Evidence and Conclusion Ontology (ECO): Supporting GO Annotations,&rdquo; in <em>The Gene Ontology Handbook</em>, C. Dessimoz and N. &Scaron;kunca, Eds., in Methods in Molecular Biology. New York, NY: Springer, 2017, pp. 245&ndash;259. doi: <a href="https://doi.org/10.1007/978-1-4939-3743-1_18">10.1007/978-1-4939-3743-1_18</a>.</p> <p>[15]&nbsp;&nbsp;&nbsp; M. Black <em>et al.</em>, &ldquo;EDAM: the bioscientific data analysis ontology,&rdquo; <em>F1000Research</em>, vol. 11, Jan. 2021, doi: <a href="https://doi.org/10.7490/f1000research.1118900.1">10.7490/f1000research.1118900.1</a>.</p> <p>[16]&nbsp;&nbsp;&nbsp; Council of European Social Science Data Archives (CESSDA), &ldquo;European Language Social Science Thesaurus ELSST,&rdquo; 2021. <a href="https://thesauri.cessda.eu/en/">https://thesauri.cessda.eu/en/</a> (accessed May 08, 2023).</p> <p>[17]&nbsp;&nbsp;&nbsp; M. Scriven, <em>Evaluation Thesaurus</em>, 3rd Edition. Edgepress, 1981. Accessed: May 08, 2023. [Online]. Available: <a href="https://us.sagepub.com/en-us/nam/evaluation-thesaurus/book3562">https://us.sagepub.com/en-us/nam/evaluation-thesaurus/book3562</a></p> <p>[18]&nbsp;&nbsp;&nbsp; Papantoniou, Bill <em>et al.</em>, <em>The Glossary of Human Computer Interaction</em>. Interaction Design Foundation. Accessed: May 08, 2023. [Online]. Available: <a href="https://www.interaction-design.org/literature/book/the-glossary-of-human-computer-interaction">https://www.interaction-design.org/literature/book/the-glossary-of-human-computer-interaction</a></p> <p>[19]&nbsp;&nbsp;&nbsp; &ldquo;UK Data Service Vocabularies: HASSET Thesaurus.&rdquo; <a href="https://hasset.ukdataservice.ac.uk/hasset/en/">https://hasset.ukdataservice.ac.uk/hasset/en/</a> (accessed May 08, 2023).</p> <p>[20]&nbsp;&nbsp;&nbsp; S. D. Costa, M. P. Barcellos, R. de A. Falbo, T. Conte, and K. M. de Oliveira, &ldquo;A core ontology on the Human&ndash;Computer Interaction phenomenon,&rdquo; <em>Data Knowl. Eng.</em>, vol. 138, p. 101977, Mar. 2022, doi: <a href="https://doi.org/10.1016/j.datak.2021.101977">10.1016/j.datak.2021.101977</a>.</p> <p>[21]&nbsp;&nbsp;&nbsp; V. J. Gawron <em>et al.</em>, &ldquo;Human Factors Taxonomy,&rdquo; <em>Proc. Hum. Factors Soc. Annu. Meet.</em>, vol. 35, no. 18, pp. 1284&ndash;1287, Sep. 1991, doi: <a href="https://doi.org/10.1177/154193129103501807">10.1177/154193129103501807</a>.</p> <p>[22]&nbsp;&nbsp;&nbsp; L. Onnasch and E. Roesler, &ldquo;A Taxonomy to Structure and Analyze Human&ndash;Robot Interaction,&rdquo; <em>Int. J. Soc. Robot.</em>, vol. 13, no. 4, pp. 833&ndash;849, Jul. 2021, doi: <a href="https://doi.org/10.1007/s12369-020-00666-5">10.1007/s12369-020-00666-5</a>.</p> <p>[23]&nbsp;&nbsp;&nbsp; R. A. Schwier, &ldquo;A Taxonomy of Interaction for Instructional Multimedia.&rdquo; Sep. 28, 1992. Accessed: May 09, 2023. [Online]. Available: <a href="https://eric.ed.gov/?id=ED352044">https://eric.ed.gov/?id=ED352044</a></p> <p>[24]&nbsp;&nbsp;&nbsp; C. Kelly, J. Miller, A. Redlich, and S. Kleinman, &ldquo;A Taxonomy of Interrogation Methods,&rdquo; <em>Psychol. Public Policy Law</em>, vol. 19, p. 165, May 2013, doi: <a href="https://doi.org/10.1037/a0030310">10.1037/a0030310</a>.</p> <p>[25]&nbsp;&nbsp;&nbsp; Y. Chuang, L.-L. Chen, and Y. Liu, &ldquo;Design Vocabulary for Human-IoT Systems Communication,&rdquo; in <em>Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems</em>, Montreal QC Canada: ACM, Apr. 2018, pp. 1&ndash;11. doi: <a href="https://doi.org/10.1145/3173574.3173848">10.1145/3173574.3173848</a>.</p> <p>[26]&nbsp;&nbsp;&nbsp; N. D&iacute;az Rodr&iacute;guez, R. Wikstr&ouml;m, J. Lilius, M. P. Cu&eacute;llar, and M. Delgado Calvo Flores, &ldquo;Understanding Movement and Interaction: An Ontology for Kinect-Based 3D Depth Sensors,&rdquo; in <em>Ubiquitous Computing and Ambient Intelligence. Context-Awareness and Context-Driven Interaction</em>, G. Urzaiz, S. F. Ochoa, J. Bravo, L. L. Chen, and J. Oliveira, Eds., in Lecture Notes in Computer Science, vol. 8276. Cham: Springer International Publishing, 2013, pp. 254&ndash;261. doi: <a href="https://doi.org/10.1007/978-3-319-03176-7_33">10.1007/978-3-319-03176-7_33</a>.</p> <p>[27]&nbsp;&nbsp;&nbsp; &ldquo;Thesaurus: mass communication - UNESCO Digital Library.&rdquo; <a href="https://unesdoc.unesco.org/ark:/48223/pf0000015031">https://unesdoc.unesco.org/ark:/48223/pf0000015031</a> (accessed May 08, 2023).</p> <p>[28]&nbsp;&nbsp;&nbsp; Institute for Scientific and Technical Information, <em>Thesaurus Cognitive Psychology of Human Memory</em>, Version 2.0. 2021. Accessed: May 08, 2023. [Online]. Available: <a href="https://fairsharing.org/FAIRsharing.LcyXdU">https://fairsharing.org/FAIRsharing.LcyXdU</a></p> <p>[29]&nbsp;&nbsp;&nbsp; S. Jiang and R. C. Arkin, &ldquo;Mixed-Initiative Human-Robot Interaction: Definition, Taxonomy, and Survey,&rdquo; in <em>2015 IEEE International Conference on Systems, Man, and Cybernetics</em>, Oct. 2015, pp. 954&ndash;961. doi: <a href="https://doi.org/10.1109/SMC.2015.174">10.1109/SMC.2015.174</a>.</p> <p>[30]&nbsp;&nbsp;&nbsp; S. Moller, K.-P. Engelbrecht, C. Kuhnel, I. Wechsung, and B. Weiss, &ldquo;A taxonomy of quality of service and Quality of Experience of multimodal human-machine interaction,&rdquo; in <em>2009 International Workshop on Quality of Multimedia Experience</em>, Jul. 2009, pp. 7&ndash;12. doi: <a href="https://doi.org/10.1109/QOMEX.2009.5246986">10.1109/QOMEX.2009.5246986</a>.</p> <p>[31]&nbsp;&nbsp;&nbsp; K. Belhajjame <em>et al.</em>, &ldquo;Using a suite of ontologies for preserving workflow-centric research objects,&rdquo; <em>J. Web Semant.</em>, vol. 32, pp. 16&ndash;42, May 2015, doi: <a href="https://doi.org/10.1016/j.websem.2015.01.003">10.1016/j.websem.2015.01.003</a>.</p> <p>[32]&nbsp;&nbsp;&nbsp; M. Augstein and T. Neumayr, &ldquo;A Human-Centered Taxonomy of Interaction Modalities and Devices,&rdquo; <em>Interact. Comput.</em>, vol. 31, no. 1, pp. 27&ndash;58, Jan. 2019, doi: <a href="https://doi.org/10.1093/iwc/iwz003">10.1093/iwc/iwz003</a>.</p> <p>[33]&nbsp;&nbsp;&nbsp; J. Jerald, &ldquo;A Taxonomy of Spatial Interaction Patterns and Techniques,&rdquo; <em>IEEE Comput. Graph. Appl.</em>, vol. 38, no. 1, pp. 11&ndash;19, Jan. 2018, doi: <a href="https://doi.org/10.1109/MCG.2018.011461524">10.1109/MCG.2018.011461524</a>.</p> <p>[34]&nbsp;&nbsp;&nbsp; L. Tian and S. Oviatt, &ldquo;A Taxonomy of Social Errors in Human-Robot Interaction,&rdquo; <em>ACM Trans. Hum.-Robot Interact.</em>, vol. 10, no. 2, pp. 1&ndash;32, Jun. 2021, doi: <a href="https://doi.org/10.1145/3439720">10.1145/3439720</a>.</p> <p>[35]&nbsp;&nbsp;&nbsp; A. Haller, K. Janowicz, S. Cox, D. Phuoc, K. Taylor, and M. Lefran&ccedil;ois, <em>Semantic Sensor Network Ontology</em>. 2017.</p> <p>[36]&nbsp;&nbsp;&nbsp; J. Malone <em>et al.</em>, &ldquo;The Software Ontology (SWO): a resource for reproducibility in biomedical data analysis, curation and digital preservation,&rdquo; <em>J. Biomed. Semant.</em>, vol. 5, no. 1, p. 25, Jun. 2014, doi: <a href="https://doi.org/10.1186/2041-1480-5-25">10.1186/2041-1480-5-25</a>.</p> <p>[37]&nbsp;&nbsp;&nbsp; L. Borek, Q. Dombrowski, J. Perkins, and C. Sch&ouml;ch, &ldquo;TaDiRAH: a Case Study in Pragmatic Classification,&rdquo; <em>Digit. Humanit. Q.</em>, vol. 010, no. 1, Feb. 2016.</p> <p>[38]&nbsp;&nbsp;&nbsp; M. A. Muhanna, &ldquo;Virtual reality and the CAVE: Taxonomy, interaction challenges and research directions,&rdquo; <em>J. King Saud Univ. - Comput. Inf. Sci.</em>, vol. 27, no. 3, pp. 344&ndash;361, Jul. 2015, doi: <a href="https://doi.org/10.1016/j.jksuci.2014.03.023">10.1016/j.jksuci.2014.03.023</a>.</p> <p>[39]&nbsp;&nbsp;&nbsp; F. Scharf, C. Wolters, M. Herczeg, and J. Cassens, &ldquo;Cross-Device Interaction: Definition, Taxonomy and Application,&rdquo; presented at the AMBIENT 2013 : The Third International Conference on Ambient Computing, Applications, Services and Technologies, Porto, Portugal: IARIA, 2013, pp. 35&ndash;41. Accessed: May 08, 2023. [Online]. Available: <a href="https://www.imis.uni-luebeck.de/de/forschung/publikationen/6380">https://www.imis.uni-luebeck.de/de/forschung/publikationen/6380</a></p>

opencc-by-4.0May 2023View details →
zenodo40/100

Echo from noise: synthetically generated cardiac ultrasound data using semantic diffusion models

<p>This is the data repository for the paper: &quot;Echo from noise: synthetic ultrasound image generation using diffusion models for real image segmentation&quot;, available at: https://arxiv.org/abs/2305.05424. The corresponding code is available at:&nbsp;https://github.com/david-stojanovski/echo_from_noise</p> <p>&nbsp;</p> <p>This is the first work to utilize Denoising Diffusion Probabilistic Models (DDPMs)&nbsp;for generating medical images using semantic label maps as a source image for conditioning the generated image.</p> <p>Each of the 400+50 CAMUS patients contributes with 4 labelled frames (ED and ES for 2 chamber and 4 chamber), totalling 1800 initial semantic maps, to which we added the sector label. These semantic maps then had five random deformations applied (a combination of random affine and elastic deformation) to produce, 9000 transformed semantic maps (8000 for training and 1000 for validation).&nbsp;</p> <p>Affine transformation ranges for rotation degrees, translate, scale and shear were: (-5, 5), (0, 0.05), (0.8, 1.05)&nbsp;and 5&nbsp;respectively. This was implemented using the torchvision python package. Elastic deformation was implemented using the TorchIO package. The settings for number of control points and max displacement were (10, 10, 4)&nbsp;and (0, 30, 30)&nbsp;respectively.</p> <p>Using these 9000 semantic maps as input to the generative models, we produced 9000 synthetic ultrasound images.</p> <p>Each echo view folder contains 3 folders:</p> <p>1) annotations: augmented labels, with no sector label and no clipping due to sector</p> <p>2) images: semantic diffusion model inferenced images</p> <p>3) sector_annotations: label maps which contain ultrasound cone sector, which were used to generate corresponding semantic diffusion model images</p> <p>ema_0.9999_050000_2ch_ed_256.pt and&nbsp;ema_0.9999_050000_4ch_ed_256.pt are the saved checkpoints for the 2 and 4 chamber diffusion models respectively.</p> <p>The pretrained segmentation networks are provided within the&nbsp;<a href="https://zenodo.org/api/files/0af4e6a3-234d-40a3-8351-c91261628982/final_models.zip">final_models.zip</a>&nbsp;file.</p> <p>A diagram of image numbers is shown in&nbsp;<a href="https://zenodo.org/api/files/0af4e6a3-234d-40a3-8351-c91261628982/Data%20diagram.png">Data diagram.png</a></p>

opencc-by-4.0May 2023View details →
zenodo40/100

Supplementary code and data for the paper: 'The fall of genres that did not happen: formalising history of the "universal" semantics of Russian iambic tetrameter'

<p>The dataset provides preprocessed data and the full code used in the paper 'The fall of genres that did not happen: formalising history of the "universal" semantics of Russian iambic tetrameter'. The code can be also be accessed as rendered notebooks on Github.</p><p>The dataset is structured as follows:</p><ul><li><i>data/ </i>: This folder contains preprocessed data,including a sampled corpus of periodicals and a document-term matrix used for topic modelling;&nbsp;</li><li><i>scr/</i> : The code used for the analysis, with separate scripts for figures;&nbsp;</li><li><i>plots/</i> : The figures used in the paper, which correspond to the aforementioned code.</li></ul>

openother-openMay 2023View details →
zenodo40/100

Main and extended tables for the 207-word Swadesh list of Early Sranan and Modern Sranan with parts of speech, semantic categories, source languages and semantic and lexical changes

<p>The dataset was made for the purposes of the author&#39;s master thesis, titled <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&amp;lang=slv">&quot;Socio-Cultural Motivations for the Acquisition of Lexical Items in Sranan Tongo&rsquo;s Core Vocabulary&quot;</a>. The dataset includes two worksheets. The first is titled &quot;Main table&quot;, and it includes all the data, where each Swadesh gloss (1 to 207) is assigned one ID (No., first column), even if there are multiple Modern Sranan (MSr) equivalents. The second worksheet, titled &quot;Extended table&quot;, includes additional IDs (No., first column) by hyphenating, so that each MSr equivalent has its separate ID number (e. g. gloss numbered 2 has 3 MSr equivalents, so these are now numbered 2-1, 2-2, and 2-3, respectively).&nbsp;<br> This allowed the author to also make a clearer distinction according to source languages, as the MSr equivalents for the same gloss sometimes come from different source languages. More about the methodology of the tables and their importance for the research is available in the thesis, available <a href="https://repozitorij.uni-lj.si/Dokument.php?id=170462&amp;lang=slv">at&nbsp;this link</a>.&nbsp;&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

AGREE: a New Benchmark for the Evaluation of Semantic Models of Ancient Greek

<p>AGREE (Ancient Greek Relatedness Embeddings Evaluation) is a benchmark for the evaluation of semantic models of Ancient Greek created at the University of Groningen (The Netherlands). More information about it can be found in the following publication:</p> <p>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim, AGREE: a new benchmark for the evaluation of distributional semantic models of ancient Greek,&nbsp;<em>Digital Scholarship in the Humanities</em>, Volume 39, Issue 1, April 2024, Pages 373&ndash;392,&nbsp;<a href="https://doi.org/10.1093/llc/fqad087">https://doi.org/10.1093/llc/fqad087</a></p> <p>&nbsp;</p> <p><strong>1. Overview of the repository</strong></p> <p>This benchmark was created from a mix of expert judgements about relatedness between Ancient Greek words and model outputs validated by human experts. The evaluation items are pairs of Ancient Greek lemmas with a high semantic relatedness.</p> <p>The human judgements were collected via two questionnaires, proposing two different tasks to the experts. The evaluation items included in the AGREE benchmark are a selection of the most strictly related pairs of lemmas obtained from the two tasks. Here an overview of the contents of the repository:</p> <ul> <li><strong>1_agree_task1.json</strong>&nbsp;includes all the data collected with the first task. The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'frequency': the number of times that the pair was suggested as related by an expert;</li> <li>'POS1': part-of-speech of the first lemma;</li> <li>'POS2': part-of-speech of the second lemma;</li> <li>'benchmark': inclusion of the pair in the AGREE benchmark ('yes'/'no').</li> </ul> </li> <li><strong>2_agree_task2.json&nbsp;</strong>includes all the data collected with the second task.&nbsp;The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'origin':&nbsp; <ul> <li>'common_pair' = one of the two pairs proposed to all participants in the second task;</li> <li>'task1' = pairs proposed by experts in the first task;</li> <li>'models_easy_rel' = output of word2vec models, pair considered as strictly related;</li> <li>'models_task1' = pairs proposed by experts in the first task and also output by word2vec models;</li> <li>'models' = output of word2vec language models;</li> <li>'unrelated' = made up pairs of unrelated lemmas (control pairs);</li> </ul> </li> <li>'respondents': number of experts evaluating a pair;</li> <li>'score': average relatedness score given by the experts on a 0-100 scale;</li> <li>'agreement': inter-annotated agreement between all experts who evaluated the&nbsp;block of pairs to which the current&nbsp;pair belongs&nbsp;(when available, i.e. when the block of pairs was presented to more than one participant);</li> <li>'benchmark': inclusion of the pair in the AGREE benchmark ('yes'/'no').</li> </ul> </li> <li><strong>3_agree_final_benchmark.json </strong>includes the&nbsp;final selection of items that constitutes AGREE. The following labels are used: <ul> <li>'pair': two Ancient Greek lemmas;</li> <li>'origin': <ul> <li>'task1': pair either proposed more than once in the first task or proposed only once, but scored&nbsp;&gt;=&nbsp;70 in the second task;</li> <li>'task2': pair scored by more than one respondent in the second task and with average score &gt;= 70.</li> </ul> </li> </ul> </li> </ul> <p>&nbsp;</p> <p>This updated version of the repository includes the individual answers to the two questionnaires (see files 'answers_Task1_postprocessed.xlsx' and 'raw_answers_Task2.xlsx').</p> <p>&nbsp;</p> <p><strong>2. Acknowledgements</strong></p> <div>This work was partially supported by the Young Academy Groningen through the PhD scholarship of Silvia Stopponi.<br>&nbsp;<br>We acknowledge the financial support of Anchoring Innovation. Anchoring Innovation is the Gravitation Grant research agenda of the Dutch National Research School in Classical Studies, OIKOS. It is financially supported by the Dutch ministry of Education, Culture and Science (NWO project number 024.003.012). For more information about the research programme and its results, see the website <a href="https://www.anchoringinnovation.nl">www.anchoringinnovation.nl</a>.<br>&nbsp;<br>We want to thank the experts of Ancient Greek around the world who shared their knowledge of Ancient Greek semantics and donated some of their precious time. Without them the creation of this benchmark would not have been possible.<br>&nbsp;<br>We also want to thank the many colleagues from the University of Groningen, the National Research School OIKOS, and other Universities abroad who contributed to this work with discussion and advice.</div> <div>&nbsp;</div> <div>&nbsp;<br><strong>3. Citation</strong><br>Silvia Stopponi, Saskia Peels-Matthey, Malvina Nissim, AGREE: a new benchmark for the evaluation of distributional semantic models of ancient Greek, <em>Digital Scholarship in the Humanities</em>, Volume 39, Issue 1, April 2024, Pages 373&ndash;392, <a href="https://doi.org/10.1093/llc/fqad087">https://doi.org/10.1093/llc/fqad087</a></div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div> <div>&nbsp;</div>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Immediate neural impact and incomplete compensation after semantic hub disconnection

<p>Data repository for &quot;Immediate neural impact and incomplete compensation after semantic hub disconnection&quot;</p> <p>Abstract: The human brain extracts meaning using an extensive neural system for semantic knowledge.&nbsp;<br> Whether broadly distributed systems depend on or can compensate after losing a highly interconnected hub is controversial.&nbsp;<br> We report rare intracranial recordings from two patients during a speech prediction task,&nbsp;<br> obtained minutes before and after neurosurgical treatment requiring disconnection of the left anterior temporal lobe (ATL),&nbsp;<br> a candidate semantic knowledge hub. Informed by modern diaschisis and predictive coding frameworks,&nbsp;<br> we tested hypotheses ranging from solely neural network disruption to complete compensation by&nbsp;<br> the indirectly affected language-related and speech processing sites. Immediately after ATL disconnection,&nbsp;<br> we observed substantial neurophysiological alterations in the recorded frontal and auditory sites,&nbsp;<br> providing direct evidence for the importance of the ATL as a semantic hub. We also obtained evidence for rapid, albeit incomplete,&nbsp;<br> attempts at neural network compensation, with neural impact largely in the forms stipulated by the predictive coding framework,&nbsp;<br> in specificity, and the modern diaschisis framework, more generally. The overall results validate these frameworks and reveal&nbsp;<br> a remarkable immediate impact and capability of the human brain to adjust after losing a brain hub.</p> <p>In this dataset,&nbsp;you will be able to access the intracranial recordings from 2 subjects,&nbsp;<br> with the description of recording channels, along with event codes and times in ms. Both the raw data,&nbsp;<br> and the preprocessed (DBT denoised and SVD applied; matrix in times x channels format, 1000 Hz sampling rate) data are available. For further information,&nbsp;please consult the paper (currently in press, but will provide a DOI&nbsp;when available),&nbsp;<br> or email zsuzsanna-kocsis@uiowa.edu or zkocsis@andrew.cmu.edu<br> &nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

NCCD-PF - A pre-failure narrow concrete cracks dataset for engineering structures damage classification and semantic segmentation

<p>The&nbsp;NCCD-PF dataset was developed for the classification and semantic segmentation of narrow concrete cracks in engineering structures elements at the pre-failure state. It only includes cracks whose width is narrower than 0.3 mm, i.e. the limit value specified in EC 1992-1-1 for typical elements of engineering structures and environmental conditions.</p> <p>This dataset is dedicated to the early crack detection at a stage when the serviceability limit state has not yet been exceeded and the failure of a structural element has not occurred. By implementing the early crack detection approach, it is possible to protect cracks in order to stop or slow down their propagation and thus to extend the structure's lifespan.</p> <p>This dataset contains images of cracks appearing on various elements of engineering structures (bridges, viaducts, tunnels) made of reinforced concrete (including abutments, tunnel walls, concrete barriers, pillars). The images were captured on construction sites and during inspections of engineering structures, at different stages of the reinforced concrete structure's working conditions - from the construction stage (when the elements are loaded only by their own weight) to the structure's use stage (when the elements are loaded by most of the design loads). The images are also differentiated by the cause of the cracking (ex., thermal and shrinkage stresses in young concrete, excessive stresses). The images were acquired using fixed-focus cameras without prior conditioning in order to represent the real working conditions of a bridge engineer during structural inspections. The images are characterised by a high degree of complexity due to the quality of the concrete surface finish (e.g. presence of formwork marks, concrete trowel marks), which could potentially be recognised&nbsp;as cracks.</p> <p>This dataset is dedicated to researchers working in the fields of computer vision, machine learning and deep learning. In particular, it contains domain knowledge in structural health monitoring, so that it can support the work of engineers in detecting cracks of concrete elements in a pre-failure state.</p> <p>A detailed description of the dataset is presented in <a href="https://www.nature.com/articles/s41597-023-02839-z" target="_blank" rel="noopener">A pre-failure narrow concrete cracks dataset for engineering structures damage classification and segmentation</a> (DOI: 10.1038/s41597-023-02839-z).</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Semantics 2023 Knowledge Graph

<p>Input data sources, mappings and RDF results used to generate the Semantics 2023 KG.</p> <p>We follow a simple data model, reusing the [Schema.org](https://schema.org/) vocabulary as much as possible. All papers are schema:ScholarlyArticles, that are part of tracks (schema:Event) and are authored by schema:Persons. Each article has either datasets (schema:Dataset), code repositories (schema:SoftwareSourceCode), ontologies (owl:Ontology) or Demos (schema:SoftwareApplication).</p> <p>As a result, our KG has 2521 triples, which describe 50 papers, 22 code repositories, 11 datasets, 7 demos and 5 ontologies.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Benchmark for Pairs of Papers in Semantic Scholar: 1 hop vs. 2-4 hops version 0.0

<p><strong>Benchmark for Pairs of Papers in Semantic Scholar: 1 hop vs.&nbsp; 2-4 hops (version 0.0)</strong></p> <p>There are two files: valid.txt and test.txt; both files use the same format.</p> <p>Columns 2 and 3 are corpus ids from Semantic Scholar.</p> <p>Column 1 is the distance between the two papers in the citation index.</p> <p>Columns 4 and 5 are the bins of the two paper, respectively.&nbsp; The bin is a number between 0 and 100.&nbsp; Papers are sorted by publication date.&nbsp; There are about 2M papers per bin, with the oldest papers in bin 0, and the newest papers in bin 99.</p> <p>Bin 100 is a catch-all for papers with unknown publication dates.</p> <p>head valid.txt</p> <p>1 &nbsp; &nbsp; &nbsp; 248518397 &nbsp; &nbsp; &nbsp; 1041744 97&nbsp; &nbsp; &nbsp; 51</p> <p>2 &nbsp; &nbsp; &nbsp; 248518397 &nbsp; &nbsp; &nbsp; 23848439&nbsp; &nbsp; &nbsp; &nbsp; 97&nbsp; &nbsp; &nbsp; 21</p> <p>3 &nbsp; &nbsp; &nbsp; 248518397 &nbsp; &nbsp; &nbsp; 4235810 97&nbsp; &nbsp; &nbsp; 12</p> <p>4 &nbsp; &nbsp; &nbsp; 248518397 &nbsp; &nbsp; &nbsp; 82079949&nbsp; &nbsp; &nbsp; &nbsp; 97&nbsp; &nbsp; &nbsp; 11</p> <p>1 &nbsp; &nbsp; &nbsp; 3374228 140728989 &nbsp; &nbsp; &nbsp; 79&nbsp; &nbsp; &nbsp; 0</p> <p>1 &nbsp; &nbsp; &nbsp; 68334187&nbsp; &nbsp; &nbsp; &nbsp; 36144275&nbsp; &nbsp; &nbsp; &nbsp; 58&nbsp; &nbsp; &nbsp; 34</p> <p>2 &nbsp; &nbsp; &nbsp; 68334187&nbsp; &nbsp; &nbsp; &nbsp; 7008060 58&nbsp; &nbsp; &nbsp; 4</p> <p>1 &nbsp; &nbsp; &nbsp; 205881482 &nbsp; &nbsp; &nbsp; 94036919&nbsp; &nbsp; &nbsp; &nbsp; 77&nbsp; &nbsp; &nbsp; 72</p> <p>2 &nbsp; &nbsp; &nbsp; 205881482 &nbsp; &nbsp; &nbsp; 95069173&nbsp; &nbsp; &nbsp; &nbsp; 77&nbsp; &nbsp; &nbsp; 53</p> <p>3 &nbsp; &nbsp; &nbsp; 205881482 &nbsp; &nbsp; &nbsp; 53480264&nbsp; &nbsp; &nbsp; &nbsp; 77&nbsp; &nbsp; &nbsp; 52</p> <p>Each row is assigned to a bin, B, where B = max(col4, col5).</p> <p>&nbsp;</p> <p><strong>Task</strong>: the task is to distinguish pairs of papers with distance == 1 from pairs of papers with distance &gt; 1.</p> <p><strong>Test/Train splits</strong>: For all thresholds, 0 &lt;= T_{train} &lt;= 99, train a model on rows in bins between 0 and T_{train} (inclusively).&nbsp; Test these models on rows in all bins 0 &lt;= T_{test} &lt;= 99.&nbsp; Report average accuracy for all combinations of T_{train} and T_{test}.</p> <p>&nbsp;</p> <p>Average Accuracy is defined as: mean(Predict(row) == 1, Gold(row) == 1)</p> <p>The means are computed over rows in a test bin.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

FOR-instance: a UAV laser scanning benchmark dataset for semantic and instance segmentation of individual trees

<p>The challenge of accurately segmenting individual trees from laser scanning data hinders the assessment of crucial tree parameters necessary for effective forest management, impacting many downstream applications. While dense laser scanning offers detailed 3D representations, automating the segmentation of trees and their structures from point clouds remains difficult. The lack of suitable benchmark datasets and reliance on small datasets have limited method development. The emergence of deep learning models exacerbates the need for standardized benchmarks.&nbsp;Addressing these gaps, the FOR-instance data represent a novel benchmarking dataset to enhance forest measurement using dense airborne laser scanning data, aiding researchers in advancing segmentation methods for forested 3D scenes.</p> <p>In this repository, users will&nbsp;find forest laser scanning point clouds from unamnned aerial vehicle (using Riegl sensors) that are manually segmented according to the individual trees (1130 trees) and semantic classes. The point clouds are subdivided into five data collections representing different forests in Norway, the Czech Republic, Austria, New Zealand, and Australia.&nbsp;</p> <p>These data are meant to be used either for developement of new methods (using the dev data) or for testing of exisitng methods (test data). The data splits are provided in the&nbsp;data_split_metadata.csv file.</p> <p>A full description of the FOR-instance data can be found at&nbsp;<a href="http://arxiv.org/abs/2309.01279">http://arxiv.org/abs/2309.01279</a>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

WhiteRoadLines: Dataset of 27,025 images (256x256 pixeles at 0.15 m/ pixel) containing representative road lines and markings labelled for multi-class semantic segmentation

<p>The dataset consists of 27,025 PNG images (256x256 pixels) of high resolution aerial orthoimages at 0,15 m/pixel of resolution. The images contain information related to representative road lines and markings found on highway pavement and is labelled for multi-class semantic segmentation with tree classes of white road<br>lines and markings: (1) continuous line (black color), (2) dashed line (dark gray color) and (3) separation of entry and exit lanes (light gray color), together with (4) the background (white color).&nbsp;<br>&nbsp;</p><p>The dataset has been created in the framework of the SROADEX project to train a multiclass semantic segmentation process based on Deep Learning.<br>The data have been generated using scripts developed in Python using Open Source libraries (GDAL/OGR and MapScript) for rasterization of vector cartography representing the three different types of white road lines. This cartography has been obtained from Spanish official sources (National Geographic Institute) that we have<br>revised and edited in a meticulous and systematic way to verify that the road lines are represented on the cartography according to the orthoimages, available on January 1, 2022 in the download center of the National Center of Geographic Information (CNIG).&nbsp;</p><p>In the digitisation process, 46 homogeneously distributed areas of Spain have been selected. The orthoimages used have been resampled from the original resolution of 0,25m/pixel to 0,15m/pixel, as this is closer to the width of two of the three classes of white lines in the dataset. It resulted in 80% of the images for training (21622), 10% for validation (2702) and 10% for testing (2701). The following table summarises the number of pixels of each category included in each of the three sub-datasets</p><p>&nbsp;</p><p>Set&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Nº images &nbsp; Class_1 (continuous line) &nbsp; Class_2 (discontinuous line) Class_3 (line defining highway entrance or exit) &nbsp;Class_4 (background)</p><p>Train &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 21,622 &nbsp; &nbsp; &nbsp;27,633,537 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 4,543,552 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 3,284,380 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 1,381,557,923</p><p>Validation &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2,702 &nbsp; &nbsp; &nbsp; &nbsp;3,433,103 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;570,741 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;395,646 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 172,678,782</p><p>Test &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;2,701 &nbsp; &nbsp; &nbsp; &nbsp;3,435,072 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;536,838 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;429,527 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 172,611,299</p><p>Total &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 27,025 &nbsp; &nbsp; &nbsp;34,501,712 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 5,651,131 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 4,109,553 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;1,726,848,004</p>

opencc-by-4.0Oct 2023View details →
dryad40/100

Evoking the N400 Event-Related Potential (ERP) component using a publicly available novel set of sentences with semantically incongruent or congruent eggplants (endings)

Open the record for dataset details and reuse information.

publicOct 2024View details →
dryad40/100

[Stimulus Set] Evoking the N400 event-related potential (ERP) component using a publicly available novel set of sentences with semantically incongruent or congruent eggplants (endings)

Open the record for dataset details and reuse information.

publicSep 2024View details →
dryad40/100

Measuring semantic memory using associative and dissociative retrieval tasks

Open the record for dataset details and reuse information.

publicJan 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record