Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

483

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

483 results for “semantics”

Learn how ShareScore rates datasets ↗
zenodo44/100

Historical City Maps Semantic Segmentation Dataset

<p>This dataset includes a total of 635 annotated image patches from historical city maps. It is designed for the semantic segmentation of the maps into 5 semantic classes (building blocks, non-built, water, road network, background frame). 330 patches are taken from maps of the city of Paris, while the 305 others are taken from a balanced corpus of city maps from 90 countries all around the world.</p> <p>Please read the detailed informations about data collection methodology, associated metadata and annotation ontology in README.md hereunder :</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other)

<p><em><strong>Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other)</strong></em></p> <p><strong>Description</strong></p> <p>579 images and 579 associated labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts. The 4 classes are 0=water, 1=whitewater, 2=sediment, 3=other</p> <p>These images and labels have been made using the Doodleverse software package, Doodler*. These images and labels could be used within numerous Machine Learning frameworks for image segmentation, but have specifically been made for use with the Doodleverse software package, Segmentation Gym**.</p> <p>Some (422) of these images and labels were originally included in the Coast Train*** data release, and have been modified from their original by reclassifying from the original classes to the present 4 classes.</p> <p>The label images are a subset of the following data release**** <a href="https://doi.org/10.5281/zenodo.7335647">https://doi.org/10.5281/zenodo.7335647</a></p> <p>Imagery comes from the following 10 sand beach sites:</p> <ol> <li>Duck, NC, Hatteras NC, USA</li> <li>Santa Cruz CA, USA</li> <li>Galveston TX, USA</li> <li>Truc Vert,France</li> <li>Sunset State Beach CA, USA</li> <li>Torrey Pines CA, USA</li> <li>Narrabeen, NSW, Australia</li> <li>Elwha WA, USA</li> <li>Ventura region, CA, USA</li> <li>Klamath region, CA USA</li> </ol> <p>Imagery are a mixture of 10-m Sentinel-2 and 15-m pansharpened Landsat 7, 8, and 9 visible-band imagery of various sizes. Red, Green, Blue, NIR, and SWIR bands only</p> <p><strong>File descriptions</strong></p> <ol> <li>classes.txt, a file containing the class names</li> <li>images.zip, a zipped folder containing the 3-band RGB images of varying sizes and extents</li> <li>nir.zip, a zipped folder containing the corresponding near-infrared (NIR) imagery</li> <li>swir.zip, a zipped folder containing the corresponding shortwave-infrared (SWIR) imagery</li> <li>labels.zip, a zipped folder containing the 1-band label images</li> <li>overlays.zip, a zipped folder containing a semi-transparent overlay of the color-coded label on the image (blue=0=water, red=1=whitewater, yellow=2=sediment, green=3=other)</li> <li>resized_images.zip, RGB images resized to 512x512x3 pixels</li> <li>resized_nir.zip, NIR images resized to 512x512x3 pixels</li> <li>resized_swir.zip, SWIR images resized to 512x512x3 pixels</li> <li>resized_labels.zip, label images resized to 512x512 pixels</li> </ol> <p><strong>References</strong></p> <p>*Doodler: Buscombe, D., Goldstein, E.B., Sherwood, C.R., Bodine, C., Brown, J.A., Favela, J., Fitzpatrick, S., Kranenburg, C.J., Over, J.R., Ritchie, A.C. and Warrick, J.A., 2021. Human‐in‐the‐Loop Segmentation of Earth Surface Imagery. Earth and Space Science, p.e2021EA002085<a href="https://doi.org/10.1029/2021EA002085">https://doi.org/10.1029/2021EA002085</a>. See <a href="https://github.com/Doodleverse/dash_doodler">https://github.com/Doodleverse/dash_doodler.</a></p> <p>**Segmentation Gym: Buscombe, D., &amp; Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>***Coast Train data release: Wernette, P.A., Buscombe, D.D., Favela, J., Fitzpatrick, S., and Goldstein E., 2022, Coast Train--Labeled imagery for training and evaluation of data-driven models for image segmentation: U.S. Geological Survey data release, <a href="https://doi.org/10.5066/P91NP87I">https://doi.org/10.5066/P91NP87I</a>. See <a href="https://coasttrain.github.io/CoastTrain/">https://coasttrain.github.io/CoastTrain/ </a>for more information</p> <p>**** Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jes&uacute;s Gonz&aacute;lez Guill&eacute;n, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, &amp; Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.7335647">https://doi.org/10.5281/zenodo.7335647</a></p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other)

<p><strong>Description</strong></p> <p>1018 images and 1018 associated labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts. The 4 classes are 0=water, 1=whitewater, 2=sediment, 3=other</p> <p>These images and labels have been made using the Doodleverse software package, Doodler*. These images and labels could be used within numerous Machine Learning frameworks for image segmentation, but have specifically been made for use with the Doodleverse software package, Segmentation Gym**.</p> <p>Some (473) of these images and labels were originally included in the Coast Train*** data release, and have been modified from their original by reclassifying from the original classes to the present 4 classes.</p> <p>Imagery comes from the following 10 sand beach sites:</p> <ol> <li>Duck, NC, Hatteras NC, USA</li> <li>Santa Cruz CA, USA</li> <li>Galveston TX, USA</li> <li>Truc Vert,France</li> <li>Sunset State Beach CA, USA</li> <li>Torrey Pines CA, USA</li> <li>Narrabeen, NSW, Australia</li> <li>Elwha WA, USA</li> <li>Ventura region, CA, USA</li> <li>Klamath region, CA USA</li> </ol> <p>Imagery are a mixture of 10-m Sentinel-2 and 15-m pansharpened Landsat 7, 8, and 9 visible-band imagery of various sizes. Red, Green, and Blue bands only</p> <p><strong>File descriptions</strong></p> <ol> <li>classes.txt, a file containing the class names</li> <li>images.zip, a zipped folder containing the 3-band images of varying sizes and extents</li> <li>labels.zip, a zipped folder containing the 1-band label images</li> <li>overlays.zip, a zipped folder containing a semi-transparent overlay of the color-coded label on the image (blue=0=water, red=1=whitewater, yellow=2=sediment, green=3=other)</li> <li>resized_images.zip, RGB images resized to 512x512x3 pixels</li> <li>resized_labels.zip, label images resized to 512x512 pixels</li> </ol> <p><strong>References</strong></p> <p>*Doodler: Buscombe, D., Goldstein, E.B., Sherwood, C.R., Bodine, C., Brown, J.A., Favela, J., Fitzpatrick, S., Kranenburg, C.J., Over, J.R., Ritchie, A.C. and Warrick, J.A., 2021. Human‐in‐the‐Loop Segmentation of Earth Surface Imagery. Earth and Space Science, p.e2021EA002085<a href="https://doi.org/10.1029/2021EA002085">https://doi.org/10.1029/2021EA002085</a>. See <a href="https://github.com/Doodleverse/dash_doodler">https://github.com/Doodleverse/dash_doodler.</a></p> <p>**Segmentation Gym: Buscombe, D., &amp; Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>***Coast Train data release: Wernette, P.A., Buscombe, D.D., Favela, J., Fitzpatrick, S., and Goldstein E., 2022, Coast Train--Labeled imagery for training and evaluation of data-driven models for image segmentation: U.S. Geological Survey data release, <a href="https://doi.org/10.5066/P91NP87I">https://doi.org/10.5066/P91NP87I</a>. See <a href="https://coasttrain.github.io/CoastTrain/">https://coasttrain.github.io/CoastTrain/ </a>for more information</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Dataset: Comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research.

<p>Supplementary material for a comparative evaluation of a keyword based search and semantic search in a data portal for biodiversity research. We conducted a relevance evaluation with 6 users over 19 search questions in two search interfaces.</p> <p>The users provided up to five search questions and relevant keywords from their research background. We setup a dataset search over a corpus of ~92,000 randomly selected metadata files from GFBio (<a href="https://www.gfbio.org">https://www.gfbio.org</a>). For each of their own search queries, the users got two result sets presented. The first one displayed results obtained from a keyword search. The second panel contained dataset results from a prototypical semantic search. Instead of results with exact mentions of the query terms, the semantic search also presented related results with synonyms and more specific terms or terms obtained from concept nodes of a higher hierarchy level.</p> <p>Each user rated the relevance of his/her own search queries on a 7-point Likert scale for both search results.<br> In addition, users also assessed the expanded keywords for each question.</p> <p>More information can be found in our publication:</p> <p>L&ouml;ffler, F. and Klan, F. (2016): Does Term Expansion Matter for the Retrieval of Biodiversity Data? in Joint Proceedings of the Posters and Demos Track of the 12th International Conference on Semantic Systems - SEMANTiCS2016 and the 1st International Workshop on Semantic Change &amp; Evolving Semantics (SuCCESS&#39;16), co-located with the 12th International Conference on Semantic Systems (SEMANTiCS 2016),2016, <a href="http://ceur-ws.org/Vol-1695/paper2.pdf">http://ceur-ws.org/Vol-1695/paper2.pdf</a></p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, other)

<p><em><strong>Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, other)</strong></em></p> <p>Images and 2-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, other)</p> <p><strong>Description</strong></p> <p>4088 images and 4088 associated labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts. The 2 classes are 1=water, 0=other. Imagery are a mixture of 10-m Sentinel-2 and 15-m pansharpened Landsat 7, 8, and 9 visible-band imagery of various sizes. Red, Green, Blue bands only</p> <p>These images and labels could be used within numerous Machine Learning frameworks for image segmentation, but have specifically been made for use with the Doodleverse software package, Segmentation Gym**.</p> <p>Two data sources have been combined</p> <p><strong>Dataset 1</strong></p> <ul> <li>1018 image-label pairs from the following data release**** https://doi.org/10.5281/zenodo.7335647</li> <li>Labels have been reclassified from 4 classes to 2 classes.</li> <li>Some (422) of these images and labels were originally included in the Coast Train*** data release, and have been modified from their original by reclassifying from the original classes to the present 2 classes.</li> <li>These images and labels have been made using the Doodleverse software package, Doodler*.</li> </ul> <p><strong>Dataset 2</strong></p> <ul> <li>3070 image-label pairs from the Sentinel-2 Water Edges Dataset (SWED)***** dataset, https://openmldata.ukho.gov.uk/, described by Seale et al. (2022)******</li> <li>A subset of the original SWED imagery (256 x 256 x 12) and labels (256 x 256 x 1) have been chosen, based on the criteria of more than 2.5% of the pixels represent water</li> </ul> <p><strong>File descriptions</strong></p> <ul> <li>&nbsp;&nbsp;&nbsp; classes.txt, a file containing the class names</li> <li>&nbsp;&nbsp;&nbsp; images.zip, a zipped folder containing the 3-band RGB images of varying sizes and extents</li> <li>&nbsp;&nbsp;&nbsp; labels.zip, a zipped folder containing the 1-band label images</li> <li>&nbsp;&nbsp;&nbsp; overlays.zip, a zipped folder containing a semi-transparent overlay of the color-coded label on the image (red=1=water, bllue=0=other)</li> <li>&nbsp;&nbsp;&nbsp; resized_images.zip, RGB images resized to 512x512x3 pixels</li> <li>&nbsp;&nbsp;&nbsp; resized_labels.zip, label images resized to 512x512x1 pixels</li> </ul> <p><strong>References</strong></p> <p>*Doodler: Buscombe, D., Goldstein, E.B., Sherwood, C.R., Bodine, C., Brown, J.A., Favela, J., Fitzpatrick, S., Kranenburg, C.J., Over, J.R., Ritchie, A.C. and Warrick, J.A., 2021. Human‐in‐the‐Loop Segmentation of Earth Surface Imagery. Earth and Space Science, p.e2021EA002085https://doi.org/10.1029/2021EA002085. See https://github.com/Doodleverse/dash_doodler.</p> <p>**Segmentation Gym: Buscombe, D., &amp; Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. https://doi.org/10.1029/2022EA002332 See: https://github.com/Doodleverse/segmentation_gym</p> <p>***Coast Train data release: Wernette, P.A., Buscombe, D.D., Favela, J., Fitzpatrick, S., and Goldstein E., 2022, Coast Train--Labeled imagery for training and evaluation of data-driven models for image segmentation: U.S. Geological Survey data release, https://doi.org/10.5066/P91NP87I. See https://coasttrain.github.io/CoastTrain/ for more information</p> <p>****Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jes&uacute;s Gonz&aacute;lez Guill&eacute;n, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, &amp; Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7335647</p> <p>*****Seale, C., Redfern, T., Chatfield, P. 2022. Sentinel-2 Water Edges Dataset (SWED) https://openmldata.ukho.gov.uk/</p> <p>******Seale, C., Redfern, T., Chatfield, P., Luo, C. and Dempsey, K., 2022. Coastline detection in satellite imagery: A deep learning approach on new benchmark data. Remote Sensing of Environment, 278, p.113044.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Semantic 3D Tree Model Dresden 2017

<p>The semantic 3D tree model contains reconstructed tree crowns within the City of Dresden (Germany). Area-wide availability of such models and their integration into semantic 3D city models facilitates enriched visualizations of urban areas as well as 3D spatial modeling that simulates the interaction of trees with buildings and the built environment.</p> <p>The individual tree crowns were modeled using geometric primitives and correspond to the CityGML Level of Detail (LoD) 2. Individual modeling parameters were determined for each tree aiming for a realistic volume replication. LiDAR data from a survey in the year 2017 were used to parameterize the tree crowns. Tree crowns were modeled via ellipsoids fitted to crown extent, cylinders were used for trunk representation. The framework for segmenting individual trees in the LiDAR point cloud and for modeling individual tree crowns via geometric primitives is described in <a href="https://doi.org/10.1016/j.ufug.2022.127637">this article</a>.</p> <p>The tree models are available as CityGML files in the coordinate system ETRS89/UTM zone 33 (EPSG: 25833). The dataset was divided into tiles. The tile number results from the coordinate of the lower left corner in the coordinate reference system.</p> <p>The source data used was made freely available by the &ldquo;Landesamt f&uuml;r Geobasisinformation Sachsen&rdquo; (GeoSN) under the license &quot;Data license Germany - attribution - Version 2.0&quot; and can be downloaded under the following links:<br> LiDAR: <a href="https://www.geodaten.sachsen.de/downloadbereich-digitale-hoehenmodelle-4851.html">https://www.geodaten.sachsen.de/downloadbereich-digitale-hoehenmodelle-4851.html</a><br> 3D Building Model: <a href="https://www.geodaten.sachsen.de/downloadbereich-digitale-3d-stadtmodelle-4875.html">https://www.geodaten.sachsen.de/downloadbereich-digitale-3d-stadtmodelle-4875.html</a><br> Aerial Imagery: <a href="https://www.geodaten.sachsen.de/downloadbereich-dop-4826.html">https://www.geodaten.sachsen.de/downloadbereich-dop-4826.html</a></p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

smashHit semantic model

<p>The smashHit core semantic model defines the main entities that are important for the smashHit modules (e.g. consent, metadata, contract, etc.) and investigated the several existing ontologies seeking to find the ones that better matches the smashHit needs.&nbsp;</p> <p>The version v0.2 shows the final development stage of the smashHit semantic model (smashHit core ontology) within the smashHit project.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Spanish semantic fields

<p>A database with 73.000 frequent spanish words classified in semantic fields in a three level hierarchy</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

SEMAFORA Semantic Reference Data Models

<p>To support the aim of the Semafora project, a series of Semantic Reference Data Models were created to provide a target semantic structure for the integration of standard archaeological survey data.&nbsp;</p> <p>&nbsp;</p> <p>The following models constitute the Semafora SRDM package:</p> <p>&nbsp;</p> <p>Place: This model is used to document any places associated with the archaeological survey.</p> <p>&nbsp;</p> <p>Institution: This model is used to document any institution associated with the survey.</p> <p>&nbsp;</p> <p>Period: This model is used to document the generic historical period assigned to the production of artefacts, existence of sites or other observable archaeological and historical events.</p> <p>&nbsp;</p> <p>Feature: This model is used to document any physical features, such as walls and other human-made structures observable on the field.</p> <p>&nbsp;</p> <p>Project: This model is used to document the overarching project, a part of which is the archaeological survey. Some projects may involve surveys, excavations, and other archaeological activities.</p> <p>&nbsp;</p> <p>Site: This model is used to document a site declared as archaeological as a result of the survey process.</p> <p>&nbsp;</p> <p>Digital Object: This model is used to document any type of digital asset associated with the survey.</p> <p>&nbsp;</p> <p>Survey Unit: This model is used to document a defined survey unit where the survey activity happens. It has both the properties of a place with dimensions and coordinates and of a physical thing from which samples can be collected.</p> <p>&nbsp;</p> <p>Collection: This model is used to document a collection of physical things, usually artefacts, collected while surveying.</p> <p>&nbsp;</p> <p>Artefact: This model is used to document individual artifacts collected from while surveying as a part of a larger collection of material things or as a singular artefact collection or documentation.</p> <p>&nbsp;</p> <p>Image: This model is used to document any image representing components of the archaeological survey, such as artefacts, features, places, people, etc.</p> <p>&nbsp;</p> <p>Observation: This model is used to document the act of observation usually associated with archaeological sites or survey units and the properties assigned to those as a result of the observation.</p> <p>&nbsp;</p> <p>Bibliography: This model is used to document any textual object associated with the survey or any components of it.</p> <p>&nbsp;</p> <p>Sample: This model is used to document a material sample of the survey unit. It partially overlaps with collection but acts as a parent sample that may contain other physical things besides human-made objects.</p> <p>&nbsp;</p> <p>Person: This model is used to document an individual person (alive or dead) involved in some way in the survey process.</p> <p>&nbsp;</p> <p>These models are intended to be used in order to guide semantic data mapping processes as well as to provide instructions for the creation of a target data semantic data management system.</p> <p>&nbsp;</p> <p>Each model&rsquo;s semantic reference data model description is stored here as a csv. The ongoing curation and updating of these SRDMs is undertaken using the Zellij system and can be accessed here:</p> <p>&nbsp;</p> <p><a href="https://zellij.pythonanywhere.com/docs/list/appaCfYmUmH78z85c">https://zellij.pythonanywhere.com/docs/list/appaCfYmUmH78z85c</a></p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Planet Microbe Functional and Taxonomic annotation of Illumina WGS Prokaryotic Fraction for Semantic Web Analysis

<p>Functional and Taxonomic annotations computed from a subset of Illumina Whole-Genome Sequencing samples from the prokaryotic fraction of the <a href="https://www.planetmicrobe.org/">Planet Microbe</a> database. Data was computed using the pipeline available from https://github.com/hurwitzlab/planet-microbe-functional-annotation/, and post processing scripts from https://github.com/hurwitzlab/planet-microbe-semantic-web-analysis. Files contain total annotation counts of Interpro, GO and NCBITaxon annotations, as well as additional sample metadata. See readme.txt file for more information.</p>

opencc-zeroJan 2022View details →
zenodo44/100

MammoTab 22: a giant and comprehensive dataset for Semantic Table Interpretation

<p>MammoTab is a dataset designed to evaluate semantic table annotation approaches.</p> <p>It includes two types of annotation:</p> <ol> <li>cell/mentions to Knowledge Graph (KG) entity matching (CEA task) and;</li> <li>column to KG&nbsp;class matching (CTA task).</li> </ol> <p>It is composed of 980254 tables extracted from 21149260 Wikipedia pages and annotated through Wikidata v. 20220708. The dataset is compliant with the data format used in&nbsp;<a href="https://www.cs.ox.ac.uk/isg/challenges/sem-tab/2019/index.html">SemTab2019</a>.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

June 2023 Supplement Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other)

<p><strong>June 2023 Supplement of Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB, NIR, and SWIR satellite images of coasts (water, whitewater, sediment, other)</strong></p> <p><strong>Description</strong></p> <p>Supplementary dataset to:</p> <p>Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jes&uacute;s Gonz&aacute;lez Guill&eacute;n, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, &amp; Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7335647</p> <p>This supplemental dataset consists of 283 RGB images and 283 associated labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts. Of these, 77 images-label pairs also have a corresponding NIR and SWIR satellite image. The 4 classes are 0=water, 1=whitewater, 2=sediment, 3=other</p> <p>These images and labels have been made using the Doodleverse software package, Doodler*. These images and labels could be used within numerous Machine Learning frameworks for image segmentation, but have specifically been made for use with the Doodleverse software package, Segmentation Gym**.</p> <p>Imagery are a mixture of 10-m Sentinel-2 and 15-m pansharpened Landsat 7, 8, and 9 visible-band imagery of various sizes. NIR, SWIR, Red, Green, and Blue bands only</p> <p><strong>File descriptions</strong></p> <ol> <li>classes.txt, a file containing the class names</li> <li>images.zip, a zipped folder containing the 3-band images of varying sizes and extents</li> <li>labels.zip, a zipped folder containing the 1-band label images</li> <li>overlays.zip, a zipped folder containing a semi-transparent overlay of the color-coded label on the image (blue=0=water, red=1=whitewater, yellow=2=sediment, green=3=other)</li> <li>nir.zip</li> <li>swir.zip</li> </ol> <p><strong>References</strong></p> <p>Buscombe, Daniel, Goldstein, Evan, Bernier, Julie, Bosse, Stephen, Colacicco, Rosa, Corak, Nick, Fitzpatrick, Sharon, del Jes&uacute;s Gonz&aacute;lez Guill&eacute;n, Anais, Ku, Venus, Paprocki, Julie, Platt, Lindsay, Steele, Bethel, Wright, Kyle, &amp; Yasin, Brandon. (2022). Images and 4-class labels for semantic segmentation of Sentinel-2 and Landsat RGB satellite images of coasts (water, whitewater, sediment, other) (v1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7335647</p> <p>*Doodler: Buscombe, D., Goldstein, E.B., Sherwood, C.R., Bodine, C., Brown, J.A., Favela, J., Fitzpatrick, S., Kranenburg, C.J., Over, J.R., Ritchie, A.C. and Warrick, J.A., 2021. Human‐in‐the‐Loop Segmentation of Earth Surface Imagery. Earth and Space Science, p.e2021EA002085<a href="https://doi.org/10.1029/2021EA002085">https://doi.org/10.1029/2021EA002085</a>. See <a href="https://github.com/Doodleverse/dash_doodler">https://github.com/Doodleverse/dash_doodler.</a></p> <p>**Segmentation Gym: Buscombe, D., &amp; Goldstein, E. B. (2022). A reproducible and reusable pipeline for segmentation of geoscientific imagery. Earth and Space Science, 9, e2022EA002332. <a href="https://doi.org/10.1029/2022EA002332">https://doi.org/10.1029/2022EA002332</a> See: <a href="https://github.com/Doodleverse/segmentation_gym">https://github.com/Doodleverse/segmentation_gym</a></p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

RDF dataset produced in the work "Exploring Adverse Outcome Pathways for Nanomaterials with semantic web technologies"

<p>Adverse Outcome Pathways (AOPs) have been proposed to facilitate mechanistic understanding of interactions of chemicals/materials with biological systems. Each AOP starts with a molecular initiating event (MIE) and possibly ends with adverse outcome(s) (AOs) via a series of key events (KEs). So far, the interaction of engineered nanomaterials (ENMs) with biomolecules, biomembranes, cells, and biological structures, in general, is not yet fully elucidated. There is also a huge lack of information on which AOPs are ENMs-relevant or -specific, despite numerous published data on toxicological endpoints they trigger, such as oxidative stress and inflammation. We propose to integrate related data and knowledge recently collected. Our approach combines the annotation of nanomaterials and their MIEs with ontology annotation to demonstrate how we can then query AOPs and biological pathway information for these materials. We conclude that a FAIR (Findable, Accessible, Interoperable, Reusable) representation of the ENM-MIE knowledge simplifies integration with other knowledge.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

A core ontology for modeling life cycle sustainability assessment on the Semantic Web with Accompanying Database

<p>To enable and support the uptake of semantic ontologies, we present a core ontology developed specifically to capture the data relevant for life cycle sustainability assessment. We further demonstrate the utility of the ontology by using it to integrate data relevant to sustainability assessments, such as EXIOBASE and the Yale Stocks and Flow Database to the Semantic Web. These datasets can be accessed by the machine-readable endpoint using SPARQL, a semantic query language.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

Accompanying Dataset migr_asyappctzm for Efficient Analytical Queries on Semantic Web Data Cubes

<p>This dataset&nbsp; shows how the Eurostat data cube in the orginal publicatin is modelled in QB4OLAP.</p> <p>This data is based on statistical data about asylum applications to the European Union, provided by Eurostat on</p> <p><a href="http://ec.europa.eu/eurostat/web/products-datasets/-/migr_asyappctzm">http://ec.europa.eu/eurostat/web/products-datasets/-/migr_asyappctzm</a></p> <p>Further data has been integrated from: https://github.com/lorenae/qb4olap/tree/master/examples</p>

opencc-by-4.0Oct 2017View details →
zenodo44/100

Semantic annotation of a part of the Italian Copyright Legislation

<p>The dataset is a structured JSONL file focusing on copyright law. Each entry contains key fields that annotate legal texts, mainly in Italian. These fields include:</p> <p>1. ID: A unique numerical identifier.</p> <p>2. Text: Contains the actual legal provisions.</p> <p>3. Chapter ID &amp; Heading: Identifiers and titles for chapters, categorizing the legal text.</p> <p>4. Article and Paragraph ID: Further break down of the text into articles and paragraphs.</p> <p>5. Insertions: Highlights inserted text fragments in the legal text.</p> <p>6. References: Cites external references with URLs and descriptions.</p> <p>7. Entities: Labels sections of the text, identifying their beginning and ending offsets.</p> <p>8. Relations: Intended to describe relationships between entities, although this field is empty in the sample.</p> <p>9. Comments: A field for comments, also empty in the sample.</p> <p>&nbsp;</p>

openmit-licenseSep 2023View details →
zenodo44/100

Semantic-Discrepant Outliers on CIFAR-10 Dataset

<p>We provide synthetic Out-of-distibution (OOD) dataset, which is called Semantic-Discrepant (SD)&nbsp;outliers, on&nbsp;CIFAR-10 dataset. SD outliers&nbsp;can be utilized for&nbsp;boosting OOD detection model performance. For the details, SD outliers are&nbsp;realistic OOD samples that contains incoherent semantic shift while preserving nuisances with in-distribution (ID). SD-outliers are generated&nbsp;from ID training samples using semantic-discrepant sampling in the diffusion model.&nbsp;&nbsp;so SD-outliers on CIFAR-10 contains 50000 32X32 images which is same as CIFAR-10 training dataset size. The dataset has a capacity of 768MB.</p>

opencc-by-4.0Sep 2023View details →
edi44/100

Results of semantic queries for "carbon cycling" for datasets in the DataONE catalog

DataONE (https://www.dataone.org) is a federation of institutions involved with the earth and environmental sciences that share data through common cyberinfrastructure. In 2016, the DataONE project carried out a quantification of the utility of semantic query, by measuring the precision and recall of relevant datasets available through that catalog. Precision is defined as the proportion of relevant data in the retrieved results, and recall is the proportion of relevant data retrieved, compared to all relevant data present in the repository (see Methods). This dataset contains the queries and results of that study. Four data tables are included. First, a table of the 10 queries, which were formatted in several ways, including natural language and text strings (for plain text searches of various parts of metadata), and URIs for measurements in the EcoSystem Ontology (ECSO). A second table contains 994 relevant datasets in the DataONE catalog, with a column for each of the ten queries and boolean value indicating whether the dataset is a match for that query. Two query results tables are included, for the raw and summarized results of the query tests. A fifth entity contains the zipped code (R language) used to perform the queries in the DataONE system. When run against approximately 1000 datasets (in October, 2016), results for the ten queries ranged from 0-50% (precision) and 0-100% (recall), indicating that traditional searches may sometimes be adequate to return all relevant data in a corpus, but results can be erratic and inconsistent, with potentially large returns of irrelevant data in the result set. When querying through semantic classes, precision and recall were much higher and more consistent (90-100% and 75-100%, respectively).

openCC0Mar 2023View details →
zenodo40/100

BreXLiMe: A Semantically Enriched Dataset With News Articles, Micro-Posts, and TV Shows Related to the Brexit

<p>We provide a <strong>large data set of media content metadata</strong> from various media sources (including online news sites, social media, and live-TV) in three languages (<strong>English, German, and Spanish</strong>). Overall, the data set contains rich metadata for about <strong>240 thousand news articles, 12 million micro-posts, and 900 TV shows</strong>. All media content information has been semantically enriched with annotations of both entities and categories from DBpedia.</p> <p>The data can be used as a valuable data basis for applications and studies of various disciplines (e.g., social studies, political science, and humanities) on the case of Brexit, particularly on the <strong>media landscape before the Brexit referendum held on June 23, 2016</strong>.</p> <p>We provide the data set in the RDF serialization format Turtle (.ttl) as well as in XML.</p> <p>If you use our data set, please <strong>cite</strong> it as follows:</p> <pre><code>Lei Zhang, Maribel Acosta, Michael Färber, Steffen Thoma and Achim Rettinger. "BreXearch: Exploring Brexit Data Using Cross-Lingual and Cross-Media Semantic Search". In: Proceedings of the ISWC 2017 Posters &amp; Demonstrations Track within the 16th International Semantic Web Conference (ISWC 2017). Vienna, Austria, 2017.</code></pre> <p>&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo40/100

Training dataset for semantic segmentation (U-Net) of structural conservation practices

<p>In this research, the best management practices include vegetative/structural conservation practices (SCP) across crop fields, such as grassed waterways&nbsp;and terraces. This reference dataset includes 500,000 pair patches (false-color image (B1: NIR, B2: Red, B3: Green)&nbsp;and binary label (SCP: yes[1] or no[0]).&nbsp;These training samples were randomly extracted from Iowa BMP project (<a href="https://www.gis.iastate.edu/gisf/projects/conservation-practices">https://www.gis.iastate.edu/gisf/projects/conservation-practices</a>) and present 90% of patches with SCP areas and 10% of patches non-SCP area. The patch dimension is 256 x&nbsp; 256 pixels at 2-m resolution. Due to the file size, the images were upload in different *.rar files (imagem_0_200k.rar, imagem_200_400k.rar, imagem_400_500k.rar), and the user should download all and merge them in the same folder. The corresponding labels are all in &quot;class_bin.rar&quot; file.</p> <p>Application: These pair images are useful for conservation practitioners interested in the classification of vegetative/structural SCPs using deep-learning semantic segmentation methods.</p> <p>Further information will be available in future.</p>

opencc-by-4.0May 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record