Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

7,523

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

7,523 results for “Annotation”

Learn how ShareScore rates datasets ↗
zenodo48/100

Sentinel2GlobalLULC: A dataset of Sentinel-2 georeferenced RGB imagery annotated for global land use/land cover mapping with deep learning (License CC BY 4.0)

<p>Sentinel2GlobalLULC is a deep learning-ready dataset of RGB images from the Sentinel-2 satellites designed for global land use and land cover (LULC) mapping. Sentinel2GlobalLULC v2.1&nbsp;contains 194,877 images in GeoTiff and JPEG format corresponding to 29 broad LULC classes. Each image has 224 x 224 pixels at 10 m spatial resolution and was produced by assigning the 25th percentile of all available observations in the Sentinel-2 collection between June 2015 and October 2020 in order to remove atmospheric effects (i.e., clouds, aerosols, shadows, snow, etc.). A spatial purity value was assigned to each image based on the consensus across 15 different global LULC products available in Google Earth Engine (GEE).&nbsp;</p> <p>&nbsp;</p> <p>Our dataset is structured into 3 main zip-compressed folders, an Excel file with a dictionary for class names and descriptive statistics per LULC class, and a python script to convert RGB GeoTiff images into JPEG format. The first folder called &quot;Sentinel2LULC_GeoTiff.zip&quot;&nbsp;contains 29 zip-compressed subfolders where each one corresponds to a specific LULC class with hundreds to thousands of GeoTiff Sentinel-2 RGB images. The second folder called &quot;Sentinel2LULC_JPEG.zip&quot; contains 29 zip-compressed subfolders with a JPEG formatted version of the same images provided in the first main folder. The third folder called &quot;Sentinel2LULC_CSV.zip&quot; includes 29 zip-compressed CSV files with as many rows as provided images and with 12&nbsp;columns containing the following metadata (this same metadata is provided in the image filenames):&nbsp;</p> <ul> <li>Land Cover Class ID: is the identification number of each LULC class</li> <li>Land Cover Class Short Name: is the short name of each LULC class</li> <li>Image ID: is the identification number of each image within its corresponding LULC class&nbsp;</li> <li>Pixel purity Value: is the spatial purity of each pixel for its corresponding LULC class calculated as the spatial consensus across up to 15 land-cover products&nbsp;</li> <li>GHM Value: is the spatial average of the Global Human Modification index (gHM) for each image</li> <li>Latitude: is the latitude of the center point of each image</li> <li>Longitude: is the longitude of the center point of each image</li> <li>Country Code: is the Alpha-2 country code of each image as described in the ISO 3166 international standard. To understand the country codes, we recommend the user to visit the following website where they present the Alpha-2 code for each country as described in the ISO 3166 international standard:https: //www.iban.com/country-codes</li> <li>Administrative Department Level1: is the administrative level 1 name to which each image belongs</li> <li>Administrative Department Level2: is the administrative level 2 name to which each image belongs</li> <li>Locality: is the name of the locality to which each image belongs</li> <li>Number of S2 images : is&nbsp;the number of found instances in the corresponding Sentinel-2 image collection between June 2015 and October 2020, when compositing&nbsp;and exporting&nbsp;its corresponding&nbsp;image tile</li> </ul> <p>For seven LULC classes, we could not export from GEE all images that fulfilled a spatial purity of 100% since there were millions of them. In this case, we exported a stratified random sample of 14,000 images and provided an additional CSV file with the images actually contained in our dataset. That is, for these seven LULC classes, we provide these 2 CSV files:</p> <ul> <li>A CSV file that contains all exported images for this class&nbsp;</li> <li>A CSV file that contains all images available for this class at spatial purity of 100%, both the ones exported and the ones not exported, in case the user wants to export them. These CSV filenames end with &quot;including_non_downloaded_images&quot;.</li> </ul> <p>To clearly state the geographical coverage of images available in this dataset,&nbsp; we&nbsp;included in the version v2.1, &nbsp;a compressed folder called &quot;Geographic_Representativeness.zip&quot;. This zip-compressed folder&nbsp;contains a csv file&nbsp;for each LULC class that provides the complete list of countries represented in that class. Each csv file has two columns, the first one gives the country code and the second one gives the number of images provided in that country for that LULC class. In addition to these 29 csv files, we provided another csv file that maps each ISO Alpha-2 country code to its original full country name.</p> <p>&copy;&nbsp;<a href="https://doi.org/10.5281/zenodo.5055632">Sentinel2GlobalLULC Dataset&nbsp;</a>by&nbsp;&nbsp;Yassir Benhammou, Domingo Alcaraz-Segura, Emilio Guirado, Rohaifa Khaldi, Boujem&acirc;a Achchab, Francisco Herrera &amp; Siham Tabik&nbsp;is marked with Attribution 4.0 International&nbsp;(CC-BY 4.0)</p>

opencc-by-4.0Jul 2022View details →
zenodo48/100

Thermal Bridges on Building Rooftops - Hyperspectral (RGB + Thermal + Height) drone images of Karlsruhe, Germany, with thermal bridge annotations

<p><strong>Overview:</strong></p> <p>The dataset of <strong>Thermal Bridges on Building Rooftops (TBBR dataset)</strong> consists of annotated combined RGB and thermal drone images with a height map. All images were converted to a uniform format of 3000x4000 pixels, aligned, and cropped to <strong>2680x3370</strong>&nbsp;to remove empty borders. See the &quot;Usage&quot; section below for details about the stored&nbsp;formats made available here.</p> <p>The raw images for our dataset were recorded with a normal (RGB) and a FLIR-XT2 (thermal) camera on a DJI M600 drone. They show six large building blocks of around 20 buildings per block recorded in the city centre of the German city Karlsruhe east of the market square. Because of a high overlap rate of the images, the same buildings are on average recorded from different angles in different images about 20 times.</p> <p>All images were recorded during a drone flight on March 19, 2019 from 7 a.m. to 8 a.m. At this time, temperatures were between 3.78 &deg; C and 4.97 &deg; C, humidity between 80% and 98%. There was no rain on the day of the flight, but there was 2.3mm/m&sup2; 48 hours beforehand. For recording the thermographic images an emissivity of 1.0 was set. The global radiation during this period was between 38.59 W / m&sup2; and 120.86 W / m&sup2;. No direct sunlight can be seen visually on any of the recordings.</p> <p>The dataset contains <strong>926&nbsp;images</strong> with a total of <strong>6,927&nbsp;annotations</strong> of thermal bridges on rooftops, split into train and test subsets with 723&nbsp;(5,614) and 203&nbsp;(1,313) images (annotations), respectively. The annotations only include thermal bridges that are visually identifiable with the human eye. Because of the aforementioned&nbsp;image overlap, each thermal bridge is annotated multiple times from different angles.</p> <p>For the annotation of the thermal images the image processing program <em>VGG Image Annotator </em>from the Visual Geometry Group, version 2.0.10, was used. The thermal bridge annotations are outlined with polygon shapes. These polygon lines were placed as close as possible but outside the area of significant temperature increase. If a detected thermal bridge was partially covered by another building component located in the foreground, the thermal bridge was also marked across the covering in case of minor coverings. Adjacent thermal bridges, which affect different rooftop components, were annotated separately. For example, a window with poor insulation of the window reveal located in the area of a poorly insulated roof is annotated individually. There is no overlap between annotated areas. While each image contains annotations, they&nbsp;also include&nbsp;thermal bridges present that are not annotated.</p> <p><strong>Usage:</strong></p> <p>Each compressed archive file represents one of the six flight paths.&nbsp;For the related publication the final path (Flug1_105Media) was used as a hold-out test sample. The archives contain Numpy files (one per image) of shape (2680, 3370, 5), where the final dimension is the colour channel&nbsp;in the format [B, G, R, Thermal, Height].</p> <p>Archives were compressed using&nbsp;<a href="https://facebook.github.io/zstd/">ZStandard</a> compression. They can be decompressed in a terminal by running e.g.</p> <pre><code class="language-bash">tar -I zstd -xvf Flug1_105Media.tar.zst</code></pre> <p>these will be decompressed into the file structure:</p> <pre><code>images/ └── Flug1_105Media/ └── DJI_0004_R.npy └── DJI_0006_R.npy └── ...</code></pre> <p>Corresponding annotations are provided in the COCO JSON format. There is one file for training (Flug1_100Media - Flug1_104Media blocks) and one for test (Flug1_105Media block). They contain a single class (thermal bridge) and expect the folder structure shown below.</p> <p>Note: The annotation files contain&nbsp;<em>relative</em>&nbsp;paths to numpy files, in case of problems please convert to <em>absolute</em> paths (i.e. insert the containing directory before each file path in the JSON annotation files).</p> <p>We provide the <a href="https://github.com/Helmholtz-AI-Energy/TBBRDet"><strong>TBBRDet software</strong></a>&nbsp;which includes a dataloader and dataset inspection tools which make use of the <a href="https://github.com/facebookresearch/detectron2">Detectron2</a> and&nbsp;<a href="https://github.com/open-mmlab/mmdetection">MMDetection</a> libraries.</p> <p>We recommend the following folder structure for use:</p> <pre><code>├── train/ │ ├── Flug1_100-104Media_coco.json │ └── images/ │ ├── Flug1_100Media/ │ │ ├── DJI_XXXX_R.npy │ │ └── ... │ ├── ... │ └── Flug1_104Media/ │ ├── DJI_XXXX_R.npy │ └── ... └── test/ ├── Flug1_105Media_coco.json └── images/ └── Flug1_105Media/ ├── DJI_XXXX_R.npy └── ...</code></pre> <p><strong>Metadata:</strong></p> <p>The experimental metadata was structured with the <strong>Spatio Temporal Asset Catalog (STAC)</strong> specification family.&nbsp;This specification provides a standardized way for describing geospatial assets. It defines related JSON object types of Item, Catalog, and Catalog, extending on Collection as the basis.</p> <p>One STAC Collection JSON object provides information about the recorded images and the environmental conditions during recordings. It also contains information about the overall bounding box of the entire area in which images were recorded.</p> <p>This object links to related STAC Item JSON objects containing information about the recorded city blocks and the cameras. The objects for the city blocks contain the GeoJSON geometry of the respective block and the<br> corresponding bounding box. The objects containing the camera information are based on an existing STAC extension for camera related metadata.</p> <p>Metadata of the archived NumPy files for each image was structured using the <strong>Data Package</strong> schema from the <strong>Frictionless Standards</strong>. This standard describes a collection of data files. Therefore, metadata about all containerized NumPy files of the six flight paths (Flug1_100Media - Flug1_104Media blocks and Flug1_105Media block) is provided within a JSON-based file.</p> <p>Note that <strong>camera1</strong> corresponds to the <strong>RGB camera</strong> and&nbsp;<strong>camera2</strong> the <strong>thermal</strong>.</p> <p><strong>FAIR Digital Objects:</strong></p> <p>All files are represented in a standardized way as <strong>FAIR Digital Objects<br> (FAIR DOs)</strong> to enable machine actionable decisions on the data in spirit of<br> the FAIR principles.</p> <p><strong>Persistent Identifier (PID):</strong></p> <p>Persistent Identifiers (PIDs) are&nbsp;resolvable with the <a href="https://hdl.handle.net/">Handle.Net Registry (HNR)</a>.</p> <table> <thead> <tr> <th scope="col">File</th> <th scope="col">Persistent Identifier (PID)</th> </tr> </thead> <tbody> <tr> <td>Flug1_100-104Media_coco.json</td> <td>21.11152/6ea60288-d895-414e-80c0-26c9fdd662b2</td> </tr> <tr> <td>Flug1_105Media_coco.json</td> <td>21.11152/58d43ddc-5e29-4980-8675-ae579b50a1e2</td> </tr> <tr> <td>Flug1_100.tar.zst</td> <td>21.11152/6858a0b5-cc60-40e9-afef-8c2dd8b35e8e</td> </tr> <tr> <td>Flug1_101.tar.zst</td> <td>21.11152/e670f510-7e00-4d3a-9b90-3bac7a7c069e</td> </tr> <tr> <td>Flug1_102.tar.zst</td> <td>21.11152/3ab9f444-05f6-445e-a691-62fae4021bea</td> </tr> <tr> <td>Flug1_103.tar.zst</td> <td>21.11152/365fd8cf-8e86-41b8-9d0e-b816fdd01d29</td> </tr> <tr> <td>Flug1_104.tar.zst</td> <td>21.11152/041a6111-644a-4617-afb3-3c421a88e8e3</td> </tr> <tr> <td>Flug1_105.tar.zst</td> <td>21.11152/f48bf4e7-3879-4216-8f64-45a060b8f658</td> </tr> <tr> <td>Flug1_100-105_frictionless_standards.json</td> <td>21.11152/7b58b3b5-75eb-4417-ac4d-abe025e159f6</td> </tr> <tr> <td>Flug1_collection_stac_spec.json</td> <td>21.11152/ba370aa3-6422-428c-9ff7-c2ef429df603</td> </tr> <tr> <td>Flug1_100_stac_spec.json</td> <td>21.11152/09cb76fc-b8cb-4116-a22a-68c5bdfa77b0</td> </tr> <tr> <td>Flug1_101_stac_spec.json</td> <td>21.11152/24a55398-b96b-43dd-b0fb-cd8ce302c7ce</td> </tr> <tr> <td>Flug1_102_stac_spec.json</td> <td>21.11152/721234ac-4b5a-4d02-9944-82a08ef2db35</td> </tr> <tr> <td>Flug1_103_stac_spec.json</td> <td>21.11152/ebaeb5bc-0514-47c9-bcd2-98f0253843d8</td> </tr> <tr> <td>Flug1_104_stac_spec.json</td> <td>21.11152/9854677c-77c5-4a0b-916b-57dd9ec20198</td> </tr> <tr> <td>Flug1_105_stac_spec.json</td> <td>21.11152/cfd0fc0e-f5ea-464e-a57f-28e882924860</td> </tr> <tr> <td>Flug1_camera1_stac-spec.json</td> <td>21.11152/976fcf28-f924-4a21-b53d-5d054ad8198d</td> </tr> <tr> <td>Flug1_camera2_stac-spec.json</td> <td>21.11152/37833c54-1d36-42e4-858d-831447122863</td> </tr> </tbody> </table>

opencc-by-4.0Aug 2022View details →
zenodo48/100

MicroCT scans of a hybrid poplar leaf dehydrating, with annotated slices for model training

<p>Dataset of a leaf segment of a hybrid poplar (<em>P. maximowiczii x P. nigra</em> &lsquo;Max3&rsquo;) leaf scanned using microcomputed tomography (microCT) over time as it dehydrates.</p> <p>&nbsp;</p> <p><strong>Data acquisition methodology</strong></p> <p>Plants were brought to the TOMCAT tomographic beamline of the Swiss Light Source at the Paul Scherrer Institute (Villigen, Switzerland). Before microCT scanning, a young fully expanded leaf was detached from the plant and a short strip (0.4 x 1.5 cm) was cut between second-order veins. The base of the strip was wrapped in polyimide tape and inserted into a styrofoam block fixed on a sample holder. The strip was immediately scanned by imaging 1801 projections of 100 ms under a beam energy of 21 keV and a magnification of 40x, yielding a final voxel size of 0.1625 &micro;m (field of view: ~416x416x312 &micro;m). The leaf was left to dehydrate in the holder and additional scans were taken 10, 20, 25, and 30 minutes after the initial scan. Scanned projections were reconstructed to a transverse view using both absorption (gridrec; Marone <em>et al.</em> 2012) and phase contrast enhancement (Paganin <em>et al.</em> 2002) reconstruction.</p> <p>&nbsp;</p> <p><strong>Dataset description</strong></p> <p>On the reconstructed images a region of interest was identified using a paradermal view (i.e. top to bottom of the leaf) and used to manually align the scans of each time step. Thereafter, all images were cropped to that ROI, ensuring that the same region of the leaf was present in all image stacks.</p> <p>For all stacks, files start with:<br> <em>DEHYDRATION_small_Leaf4_time_N_</em><br> where N is the time point, with values from 1 to 5 equaling 0, 10, 20, 25, and 30 minutes.</p> <p>Following this prefix is either GRID (gridrec reconstruction), PAGANIN (phase contrast enhancement reconstruction), or LABELLED (hand labelled slices or ground truth). For GRID and PAGANIN, 8-bit grayscale stacks are provided. The AOI suffix indicates the region of interest.</p> <p>Stacks have been hand labelled over three orientations (for visual examples of the orientations see <a href="https://zenodo.org/api/files/6f06d15b-3ee9-412d-82ca-a20336c4bffa/Labeled_Sections_order_time1.png?versionId=26fc15aa-702e-4052-b162-702cc567634c">Labeled_Sections_order_time1.png</a> and <a href="https://zenodo.org/api/files/6f06d15b-3ee9-412d-82ca-a20336c4bffa/Labeled_Sections_order_time2.png">Labeled_Sections_order_time2.png</a>):</p> <ol> <li>CROSS (cross sectional, or transverse, view)</li> <li>LONGI (longitudinal view: similar to cross sectional view but starting normal to it, i.e. along the depth of the stack starting from the left of the cross-sectional view)</li> <li>PARADERMAL (top to bottom view: starting at the upper epidermis)</li> </ol> <p>A general idea of the slice range within one LABELLED stack is presented after the orientation, as:<br> <em>STARTtoENDbyRANGE</em><br> The exact position of the labelled slices for each time point can be found in the <a href="https://zenodo.org/api/files/6f06d15b-3ee9-412d-82ca-a20336c4bffa/Labeled_slices_positions.txt?versionId=93d7e22f-9f07-4f98-8c49-d93e9a2e1ce5">Labeled_slices_positions.txt </a>file. <strong>Note that one-based indexing is used (as in ImageJ), not zero-based indexing (as in e.g. Python).</strong></p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>Marone F, Stampanoni M. 2012. Regridding reconstruction algorithm for realtime tomographic imaging. Journal of Synchrotron Radiation 19: 1029&ndash;1037.</p> <p>Paganin D, Mayo SC, Gureyev TE, Miller PR, Wilkins SW. 2002. Simultaneous phase and amplitude extraction from a single defocused image of a homogeneous object. Journal of Microscopy 206: 33&ndash;40.</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

StreetSurfaceVis: a dataset of street-level imagery with annotations of road surface type and quality

<h1>StreetSurfaceVis</h1> <p><em>StreetSurfaceVis</em> is an image dataset containing <strong>9,122 street-level images from Germany</strong> with labels on <strong>road surface type and quality.</strong> The CSV file <code>streetSurfaceVis_v1_0.csv</code> contains all image metadata and four folders contain the image files.&nbsp;All images are available in four different sizes, based on the image width, in 256px, 1024px, 2048px and the original size.<br>Folders containing the images are named according to the respective image size. Image files are named based on the <code>mapillary_image_id</code>.</p> <p>You can find the corresponding publication here: &nbsp;<a href="https://www.nature.com/articles/s41597-024-04295-9#citeas">StreetSurfaceVis: a dataset of crowdsourced street-level imagery with semi-automated annotations of road surface type and quality</a></p> <p>&nbsp;</p> <h3>Image metadata</h3> <p>Each CSV record contains information about one street-level image with the following attributes:</p> <ul> <li><code>mapillary_image_id</code>: ID provided by Mapillary (see information below on Mapillary)</li> <li><code>user_id</code>: Mapillary user ID of contributor</li> <li><code>user_name</code>: Mapillary user name of contributor</li> <li><code>captured_at</code>: timestamp, capture time of image</li> <li><code>longitude</code>, <code>latitude</code>: location the image was taken at</li> <li><code>train</code>: Suggestion to split train and test data. `True` for train data and `False` for test data. Test data contains data from 5 cities which are excluded in the training data.</li> <li><code>surface_type</code>: Surface type of the road in the focal area (the center of the lower image half) of the image. Possible values: asphalt, concrete, paving_stones, sett, unpaved</li> <li><code>surface_quality</code>: Surface quality of the road in the focal area of the image. Possible values: (1) excellent, (2) good, (3) intermediate, (4) bad, (5) very bad (see the attached <strong>Labeling Guide document</strong> for details)</li> </ul> <p>&nbsp;</p> <h3>Image source</h3> <p>Images are obtained from <a href="https://www.mapillary.com/">Mapillary</a>, a crowd-sourcing plattform for street-level imagery.&nbsp;More metadata about each image can be obtained via the <a href="https://www.mapillary.com/developer/api-documentation">Mapillary API . </a>User-generated images are shared by Mapillary under the <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA</a> License.</p> <p>For each image, the dataset contains the <code>mapillary_image_id</code> and <code>user_name</code>.&nbsp;<br>You can access user information on the Mapillary website by <code>https://www.mapillary.com/app/user/&lt;USER_NAME&gt;&nbsp;</code><br>and image information by <code>https://www.mapillary.com/app/?focus=photo&amp;pKey=&lt;MAPILLARY_IMAGE_ID&gt;</code></p> <p>If you use the provided images, please adhere to the <a href="https://www.mapillary.com/terms">terms of use of Mapillary.</a></p> <p>&nbsp;</p> <h3>Instances per class</h3> <p>Total number of images: 9,122</p> <table> <tbody> <tr> <td>&nbsp;</td> <td><strong>excellent</strong></td> <td><strong>good</strong></td> <td><strong>intermediate</strong></td> <td><strong>bad</strong></td> <td><strong>very bad</strong></td> </tr> <tr> <td><strong>asphalt</strong></td> <td>971</td> <td>1697</td> <td>821</td> <td>246</td> <td>-</td> </tr> <tr> <td><strong>concrete</strong></td> <td>314</td> <td>350</td> <td>250</td> <td>58</td> <td>-</td> </tr> <tr> <td><strong>paving stones</strong></td> <td>385</td> <td>1063</td> <td>519</td> <td>70</td> <td>-</td> </tr> <tr> <td><strong>sett</strong></td> <td>-</td> <td>129</td> <td>694</td> <td>540</td> <td>-</td> </tr> <tr> <td><strong>unpaved</strong></td> <td>-</td> <td>-</td> <td>326</td> <td>387</td> <td>303</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>For modeling, we recommend using a train-test split where the test data includes geospatially distinct areas, thereby ensuring the model's ability to generalize to unseen regions is tested. We propose five cities varying in population size and from different regions in Germany for testing - images are tagged accordingly.</p> <p>Number of test images (train-test split): 776</p> <h3>Inter-rater-reliablility</h3> <p>Three annotators labeled the dataset, such that each image was annotated by one person. Annotators were encouraged to consult each other for a second opinion when uncertain.<br>1,800 images were annotated by all three annotators, resulting in a <em>Krippendorff's alpha</em> of 0.96 for surface type and 0.74 for surface quality.</p> <h3>Recommended image preprocessing</h3> <p>As the focal road located in the bottom center of the street-level image is labeled, it is recommended to crop images to their lower and middle half prior using for classification tasks.</p> <p>This is an exemplary code for recommended image preprocessing in <strong>Python</strong>:</p> <pre><code>from PIL import Image<br></code><code>img = Image.open(image_path)</code><br><code>width, height = img.size</code><br><code>img_cropped = img.crop((0.25 * width, 0.5 * height, 0.75 * width, height))</code></pre> <h3><br><strong>License</strong></h3> <p><a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA</a></p> <p>&nbsp;</p> <h3><strong>Citation</strong></h3> <p>If you use this dataset, please cite as:&nbsp;</p> <p>&nbsp;</p> <p>Kapp, A., Hoffmann, E., Weigmann, E. <em>et al.</em> StreetSurfaceVis: a dataset of crowdsourced street-level imagery annotated by road surface type and quality. <em>Sci Data</em> <strong>12</strong>, 92 (2025). https://doi.org/10.1038/s41597-024-04295-9</p> <p>&nbsp;</p> <p><code>@article{kapp_streetsurfacevis_2025,<br>&nbsp; &nbsp; title = {{StreetSurfaceVis}: a dataset of crowdsourced street-level imagery annotated by road surface type and quality},<br>&nbsp; &nbsp; volume = {12},<br>&nbsp; &nbsp; issn = {2052-4463},<br>&nbsp; &nbsp; url = {https://doi.org/10.1038/s41597-024-04295-9},<br>&nbsp; &nbsp; doi = {10.1038/s41597-024-04295-9},<br>&nbsp; &nbsp; pages = {92},<br>&nbsp; &nbsp; number = {1},<br>&nbsp; &nbsp; journaltitle = {Scientific Data},<br>&nbsp; &nbsp; shortjournal = {Scientific Data},<br>&nbsp; &nbsp; author = {Kapp, Alexandra and Hoffmann, Edith and Weigmann, Esther and Mihaljević, Helena},<br>&nbsp; &nbsp; date = {2025-01-16},<br>}</code></p> <p>&nbsp;</p> <p>-----------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p>This is part of the SurfaceAI project at the University of Applied Sciences, HTW Berlin.</p> <p><br>- Prof. Dr. Helena Mihajlević<br>- Alexandra Kapp<br>- Edith Hoffmann<br>- Esther Weigmann</p> <p>Contact: surface-ai@htw-berlin.de</p> <p>https://surfaceai.github.io/surfaceai/</p> <p><strong>Funding</strong>: SurfaceAI is a mFund project funded by the Federal Ministry for Digital and Transportation Germany.</p> <p>&nbsp;</p>

opencc-by-sa-4.0Jun 2024View details →
zenodo48/100

An annotated high-content fluorescence microscopy dataset with EGFP-Galectin-3-stained cells and manually labelled outlines

<p>Here we present a benchmarking dataset of fluorescence microscopy images with EGFP-Galectin-3-stained cells together with annotations of their outlines. Images were randomly selected from an RNA interference screen with a modified U2OS osteosarcoma cell line, acquired on a Thermo Fischer CX7 high-content imaging system at 20x magnification.&nbsp;</p> <p>The dataset contains 60 images showing over 2000 labelled nuclear objects in total, which is sufficiently large to train well-performing neural networks for instance or semantic segmentation. It is pre-split into training, development and test set, each in a zip file. The dataset should be referred to as Aitslab_bioimaging2.</p> <p>For most of the images, nuclear staining and annotations have been published previously in the dataset Aitslab_bioimaging1 (https://doi.org/10.5281/zenodo.6657260). The conversion script to produce the png images from the C01 images was published together with this dataset.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Corpus of Occitan Written Traditional Folktales Annotated with Part-Of-Speech (OWT-Tag)

<p>This resource contains 5 extracts of texts in Occitan which were manually annotated with lemmas and parts-of-speech, following the Grace standard. It was produced during the ExpressioNarration project, funded by a Marie Curie Individual Fellowship, in order to evaluate the performance of an Occitan Part-Of-Speech tagger, Talismane, to the specifities of the corpus of the project called Oral Occitan (OcOr), also available on https://zenodo.org/record/1451753#.W78FJWOYSpo.<br> Each extract contains around 1500 words. They are extracted from &#39;Contes et proverbes populaires recueillis en armagnac et Contes populaires recueillis en agenais&#39; de J.-F. Blad&eacute;, &#39;Coundes biarn&eacute;s, cou&eacute;ilhuts a&uuml;s pars&agrave;as mi&eacute;ytad&egrave;s dou p&eacute;ys d&eacute; Biarn&#39; de J.-V. Lalanne, &#39;Contes populaires du Languedoc&#39; de L. Lambert and &#39;Contes populaires recueillis dans la Grande-Lande&#39; de F. Arnaudin.<br> The annotation process is described in the following article available on https://www.openscience.fr/IMG/pdf/iste_modocv1n1_2.pdf.</p>

opencc-by-sa-4.0Oct 2018View details →
zenodo48/100

Emergency management/Natural Hazards: annotated tweets

<p>A set of annotated tweets related to natural hazards and emergency management.</p> <p>Information available for each tweet:</p> <p>- tweet id</p> <p>- boolean flags about its content: floods;storms;landslides;snow;infrastructures;affected individuals;caution advice;donations &amp; volunteering;emotional support;other info;panic</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo48/100

MiRoR11 - P2 - Annotated corpus for semantic similarity of clinical trial outcomes

<p>Outcome similarity corpus</p> <p>This dataset contains annotations of semantic similarity for pairs of primary and reported outcomes.<br> Tab-separated format is used. The files contain the following columns:<br> filename, sentence pair ID, sentence pair text, primary outcome, primary outcome start position, primary outcome end position, reported outcome, reported outcome start position, reported outcome end position, label</p> <p>The folder out_relations_split contains the dataset splits for 10-fold cross-validation.</p>

opencc-by-4.0May 2019View details →
zenodo48/100

MiRoR11 - P2 - Annotated corpus for the relation between reported outcomes and their significance levels

<p>Corpus of relations between outcomes and significance levels</p> <p>This dataset contains annotations of the relations between reported outcomes and their significance levels.<br> Tab-separated format is used. The file contains the following comumns:<br> filename, sentence text, outcome, primary outcome start position, primary outcome end position, reported outcome, reported outcome start position, reported outcome end position, label</p> <p>The folder out_sig_rel contains the dataset splits for 10-fold cross-validation.</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Image segmentations produced by BAMF under the AIMI Annotations initiative

<p>The Imaging Data Commons (IDC)(<a href="https://imaging.datacommons.cancer.gov/">https://imaging.datacommons.cancer.gov/</a>) [1] connects researchers with publicly available cancer imaging data, often linked with other types of cancer data. Many of the collections have limited annotations due to the expense and effort required to create these manually. The increased capabilities of AI analysis of radiology images provide an opportunity to augment existing IDC collections with new annotation data. To further this goal, we trained several nnUNet [2] based models for a variety of radiology segmentation tasks from public datasets and used them to generate segmentations for IDC collections.</p> <p>To validate the model's performance, roughly 10% of the AI predictions were assigned to a validation set. For this set, a board-certified radiologist graded the quality of AI predictions on a Likert scale. If they did not 'strongly agree' with the AI output, the reviewer corrected the segmentation.&nbsp;</p> <p>This record provides the AI segmentations, Manually corrected segmentations, and Manual scores for the inspected IDC Collection images.</p> <p><em>Only 10% of the AI-derived annotations provided in this dataset are verified by expert radiologists . More details, on model training and annotations are provided within the associated manuscript to ensure transparency and reproducibility.</em></p> <p>&nbsp;</p> <p>This work was done in two stages. Versions 1.x of this record were from the first stage. Versions 2.x added additional records. In the Version 1.x collections, a medical student (non-expert) reviewed all the AI predictions and rated them on a 5-point Likert Scale, for any AI predictions in the validation set that they did not 'strongly agree' with, the non-expert provided corrected segmentations. This non-expert was not utilized for the Version 2.x additional records.</p> <p>&nbsp;</p> <h3>Likert Score Definition:</h3> <p>Guidelines for reviewers to grade the quality of AI segmentations.</p> <ul> <li>5 Strongly Agree - Use-as-is (i.e., clinically acceptable, and could be used for treatment without change)</li> <li>4 Agree - Minor edits that are not necessary. Stylistic differences, but not clinically important. The current segmentation is acceptable</li> <li>3 Neither agree nor disagree - Minor edits that are necessary. Minor edits are those that the review judges can be made in less time than starting from scratch or are expected to have minimal effect on treatment outcome</li> <li>2 Disagree - Major edits. This category indicates that the necessary edit is required to ensure correctness, and sufficiently significant that user would prefer to start from the scratch</li> <li>1 Strongly disagree - Unusable. This category indicates that the quality of the automatic annotations is so bad that they are unusable.</li> </ul> <p>&nbsp;</p> <h3>Zip File Folder Structure</h3> <p>Each zip file in the collection correlates to a specific segmentation task. The common folder structure is</p> <ul> <li><em>ai-segmentations-dcm </em>This directory contains the AI model predictions in DICOM-SEG format for all analyzed IDC collection files</li> <li><em>qa-segmentations-dcm </em>This directory contains manual corrected segmentation files, based on the AI prediction, in DICOM-SEG format. Only a fraction, ~10%, of the AI predictions were corrected. Corrections were performed by radiologist (rad*) and non-experts (ne*)</li> <li><em>qa-results.csv</em> CSV file linking the study/series UIDs with the ai segmentation file, radiologist corrected segmentation file, radiologist ratings of AI performance.</li> </ul> <p>&nbsp;</p> <h3><strong><em>qa-results.csv Columns</em></strong></h3> <p>The qa-results.csv file contains metadata about the segmentations, their related IDC case image, as well as the Likert ratings and comments by the reviewers.</p> <div> <table> <tbody> <tr> <td> <p><strong>Column</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p><em>Collection</em></p> </td> <td> <p>The name of the IDC collection for this case</p> </td> </tr> <tr> <td> <p><em>PatientID</em></p> </td> <td> <p>PatientID in DICOM metadata of scan. Also called Case ID in the IDC</p> </td> </tr> <tr> <td> <p><em>StudyInstanceUID</em></p> </td> <td> <p>StudyInstanceUID in the DICOM metadata of the scan</p> </td> </tr> <tr> <td> <p><em>SeriesInstanceUID</em></p> </td> <td> <p>SeriesInstanceUID in the DICOM metadata of the scan</p> </td> </tr> <tr> <td> <p><em>Validation</em></p> </td> <td> <p>true/false if this scan was manually reviewed</p> </td> </tr> <tr> <td> <p><em>Reviewer</em></p> </td> <td> <p>Coded ID of the reviewer. Radiologist IDs start with &lsquo;rad&rsquo; non-expect IDs start with &lsquo;ne&rsquo;</p> </td> </tr> <tr> <td> <p><em>AimiProjectYear</em></p> </td> <td> <p>2023 or 2024, This work was split over two years. The main methodology difference between the two is that in 2023, a non-expert also reviewed the AI output, but a non-expert was not utilized in 2024.</p> </td> </tr> <tr> <td> <p><em>AISegmentation</em></p> </td> <td> <p>The filename of the AI prediction file in DICOM-seg format. This file is in the ai-segmentations-dcm folder.</p> </td> </tr> <tr> <td> <p><em>CorrectedSegmentation</em></p> </td> <td> <p>The filename of the reviewer-corrected prediction file in DICOM-seg format. This file is in the qa-segmentations-dcm folder. If the reviewer strongly agreed with the AI for all segments, they did not provide any correction file.</p> </td> </tr> <tr> <td> <p><em>Was the AI predicted ROIs accurate?</em></p> </td> <td> <p>This column appears one for each segment in the task for images from AimiProjectYear 2023. The reviewer rates segmentation quality on a Likert scale. In tasks that have multiple labels in the output, there is only one rating to cover them all.</p> </td> </tr> <tr> <td> <p><em>Was the AI predicted {SEGMENT_NAME} label accurate?</em></p> <p><em><strong>&nbsp;</strong></em></p> </td> <td> <p>This column appears one for each segment in the task for images from AimiProjectYear 2024. The reviewer rates each segment for its quality on a Likert scale.</p> </td> </tr> <tr> <td> <p><em>Do you have any comments about the AI predicted ROIs?</em></p> <p><em><strong>&nbsp;</strong></em></p> </td> <td> <p>Open ended question for the reviewer</p> </td> </tr> <tr> <td> <p><em>Do you have any comments about the findings from the study scans?</em></p> </td> <td> <p>Open ended question for the reviewer</p> </td> </tr> </tbody> </table> </div> <p>&nbsp;</p> <h3>File Overview</h3> <h4>brain-mr.zip</h4> <ul> <li>Segment Description: brain tumor regions: necrosis, edema, enhancing</li> <li>IDC Collection: <a href="https://www.cancerimagingarchive.net/collection/upenn-gbm/">UPENN-GBM</a></li> <li>Links: <a href="../records/11582627">model weights</a>, <a href="https://github.com/bamf-health/aimi-brain-mr">github</a></li> </ul> <h4>breast-fdg-pet-ct.zip</h4> <ul> <li>Segment Description: FDG-avid lesions in breast from FDG PET/CT scans QIN-Breast</li> <li>IDC Collection: <a href="https://www.cancerimagingarchive.net/collection/qin-breast/">QIN-Breast</a></li> <li>Links: <a href="https://doi.org/10.5281/zenodo.8290054">model weights, </a><a href="https://github.com/bamf-health/aimi-breast-pet-ct">github</a></li> </ul> <h4>breast-mr.zip</h4> <ul> <li>Segment Description: Breast, Fibroglandular tissue, structural tumor</li> <li>IDC Collection: <a href="https://www.cancerimagingarchive.net/collection/duke-breast-cancer-mri/">duke-breast-cancer-mri</a></li> <li>Links: <a href="../records/11998679">model weights</a>, <a href="https://github.com/bamf-health/aimi-breast-mr">github</a></li> </ul> <h4>kidney-ct.zip</h4> <ul> <li>Segment Description: Kidney, Tumor, and Cysts from contrast enhanced CT scans</li> <li>IDS Collection: <a href="https://www.cancerimagingarchive.net/collection/tcga-kirc/">TCGA-KIRC,</a> <a href="https://www.cancerimagingarchive.net/collection/tcga-kirp/">TCGA-KIRP</a>, <a href="https://www.cancerimagingarchive.net/collection/tcga-kich/">TCGA-KICH</a>, <a href="https://www.cancerimagingarchive.net/collection/cptac-ccrcc/">CPTAC-CCRCC</a></li> <li>Links: <a href="https://doi.org/10.5281/zenodo.8277845">model weights, </a><a href="https://github.com/bamf-health/aimi-kidney-ct">github</a></li> </ul> <h4>liver-ct.zip</h4> <ul> <li>Segment Description: Liver from CT scans</li> <li>IDC Collection: <a href="https://www.cancerimagingarchive.net/collection/TCGA-LIHC/">TCGA-LIHC</a></li> <li>Links: <a href="https://doi.org/10.5281/zenodo.8270230">model weights, </a><a href="https://github.com/bamf-health/aimi-liver-ct">github</a></li> </ul> <h4>liver2-ct.zip</h4> <ul> <li>Segment Description: Liver and Lesions from CT scans</li> <li>IDC Collection: <a href="https://www.cancerimagingarchive.net/collection/hcc-tace-seg/">HCC-TACE-SEG</a>, <a href="https://www.cancerimagingarchive.net/collection/colorectal-liver-metastases/">COLORECTAL-LIVER-METASTASES</a></li> <li>Links: <a href="../records/11582728">model weights</a>, <a href="https://github.com/bamf-health/aimi-liver-tumor-ct">github</a></li> </ul> <h4>liver-mr.zip</h4> <ul> <li>Segment Description: Liver from T1 MRI scans</li> <li>IDC Collection: <a href="https://www.cancerimagingarchive.net/collection/TCGA-LIHC/">TCGA-LIHC</a></li> <li>Links: <a href="https://doi.org/10.5281/zenodo.8290123">model weights, </a><a href="https://github.com/bamf-health/aimi-liver-mr">github</a></li> </ul> <h4>lung-ct.zip</h4> <ul> <li>Segment Description: Lung and Nodules (3mm-30mm) from CT scans</li> <li>IDC Collections:<br> <ul> <li><a href="https://www.cancerimagingarchive.net/collection/anti-pd-1_lung/">Anti-PD-1-Lung</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/lung-pet-ct-dx/">LUNG-PET-CT-Dx</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/nsclc-radiogenomics/">NSCLC Radiogenomics</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/rider-lung-pet-ct/">RIDER Lung PET-CT</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/TCGA-LUAD/">TCGA-LUAD</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/TCGA-LUSC/">TCGA-LUSC</a></li> </ul> </li> <li>Links: <a href="https://doi.org/10.5281/zenodo.8290146">model weights 1, </a><a href="../record/8290169">model weights 2, </a><a href="https://github.com/bamf-health/aimi-lung-ct">github</a></li> </ul> <h4>lung2-ct.zip</h4> <ul> <li>Improved model version</li> <li>Segment Description: Lung and Nodules (3mm-30mm) from CT scans</li> <li>IDC Collections:<br> <ul> <li><a href="https://www.cancerimagingarchive.net/collection/QIN-LUNG-CT">QIN-LUNG-CT</a>,&nbsp;<a href="https://www.cancerimagingarchive.net/collection/spie-aapm-lung-ct-challenge/">SPIE-AAPM Lung CT Challenge</a></li> </ul> </li> <li>Links: <a href="../records/11582738">model weights</a>, <a href="https://github.com/bamf-health/aimi-lung2-ct">github</a></li> </ul> <h4>lung-fdg-pet-ct.zip</h4> <ul> <li>Segment Description: Lungs and FDG-avid lesions in the lung from FDG PET/CT scans</li> <li>IDC Collections: <ul> <li><a href="https://www.cancerimagingarchive.net/collection/acrin-nsclc-fdg-pet/">ACRIN-NSCLC-FDG-PET</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/anti-pd-1_lung/">Anti-PD-1-Lung</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/lung-pet-ct-dx/">LUNG-PET-CT-Dx</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/nsclc-radiogenomics/">NSCLC Radiogenomics</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/rider-lung-pet-ct/">RIDER Lung PET-CT</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/TCGA-LUAD/">TCGA-LUAD</a></li> <li><a href="https://www.cancerimagingarchive.net/collection/TCGA-LUSC/">TCGA-LUSC</a></li> </ul> </li> <li>Links: <a href="https://doi.org/10.5281/zenodo.8290054">model weights, </a><a href="https://github.com/bamf-health/aimi-lung-pet-ct">github</a></li> </ul> <h4>prostate-mr.zip</h4> <ul> <li>Segment Description: Prostate from T2 MRI scans</li> <li>IDC Collection: <a href="https://www.cancerimagingarchive.net/collection/ProstateX/">ProstateX,</a> <a href="https://www.cancerimagingarchive.net/collection/prostate-mri-us-biopsy/">Prostate-MRI-US-Biopsy</a></li> <li>Links: <a href="https://doi.org/10.5281/zenodo.8290092">model weights, </a><a href="https://github.com/bamf-health/aimi-prostate-mr">github</a></li> </ul> <p>&nbsp;</p> <p><strong>Changelog</strong></p> <ul> <li>2.0.2 - Fix the brain-mr segmentations to be transformed correctly</li> <li>2.0.1 - added AIMI 2024 radiologist comments to qa-results.csv</li> <li>2.0.0 - added AIMI 2024 segmentations</li> <li>1.X - AIMI 2023 segmentations and reviewer scores</li> </ul>

opencc-by-4.0Nov 2023View details →
zenodo48/100

GitHub Profiles (users/organisations) and Repositories (research/non-research) of Potsdam Researchers and Research Organisations: An annotated dataset of with howfairis and software quality variables.

<p>This dataset accompanies the paper <em>"Software FAIRness, Documentation and Development Practices in Potsdam Researchers' GitHub Repositories"</em> It includes 3 CSV files that contain data related to github profiles of users/organisations, their repositories annotated as research/non-research repositories and followed by FAIRness and other software qualtiy variables. The data were collected using <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP">SWORDS-template-UP</a> (v1.0.0) methods (collect_users, collect_repositories, collect_variables) which is extended version of&nbsp;<a href="https://github.com/UtrechtUniversity/SWORDS-template">SWORS-template</a> adopted according our needs and detailed in the paper.</p> <p><strong>GitHub (research) user/organisation profiles. ( <em>github_profiles.csv )</em></strong></p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description&nbsp;</strong></td> </tr> <tr> <td>user_id</td> <td>GitHub username &nbsp;</td> </tr> <tr> <td>html_url &nbsp;</td> <td>URL of the GitHub profile &nbsp;</td> </tr> <tr> <td>type &nbsp; &nbsp;</td> <td>Type of profile (user or organization)</td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the organization &nbsp; &nbsp;</td> </tr> </tbody> </table> <p><strong>GitHub repositories&nbsp;<em>(github_repositories.csv)</em></strong></p> <p>This file contains the repositories scraped from the GitHub profiles of research users and organizations.</p> <table> <tbody> <tr> <td><strong>Column name&nbsp;</strong></td> <td><strong>Description&nbsp;</strong></td> </tr> <tr> <td>html_url &nbsp;</td> <td>URL link to the repository &nbsp;</td> </tr> <tr> <td>description</td> <td>GitHub project description &nbsp;</td> </tr> <tr> <td>project</td> <td>Specifies if the project is research or non-research</td> </tr> <tr> <td>language</td> <td>Programming language used in the project &nbsp;</td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the university, institution, or research organization</td> </tr> <tr> <td>research_group</td> <td>Acronym or name of the research group the repository belongs to</td> </tr> </tbody> </table> <p><strong>Research repositories filtered and annotated&nbsp;<em>(github_research_repositories_filtered_annotated.csv)</em></strong></p> <p>This file contains filtered and annotated information about research repositories.</p> <table> <tbody> <tr> <td><strong>Column Name&nbsp;</strong></td> <td><strong>Description&nbsp;</strong></td> <td><strong>Collection Method&nbsp;</strong></td> </tr> <tr> <td>html_url &nbsp;</td> <td>Repository URL &nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>howfairis_repository</td> <td>Indicates if the repository is public or private (True/False) &nbsp;</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_license &nbsp;</td> <td>Indicates if the repository has a license (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_registry</td> <td>Indicates if the repository has implemented community registry (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_citation</td> <td>Indicates if the repository has a .cff file (True/False) &nbsp;</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_checklist</td> <td>Indicates if the repository has implemented OpenSSF best practices badge (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>fair_score</td> <td>Score based on howfairis variables (0-5) &nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>dlr_soft_class</td> <td>Name of the university, company, research institute, or research organization</td> <td>(Manual) Annotated the repository based on <a href="https://core.ac.uk/reader/211557820">DLR software engineering guideline.</a> There are no specific definitions on metrics how to categorise them (github repositories) into application classes. Which were needed to do a comparitive analysis.&nbsp;</td> </tr> <tr> <td>installation_instruction</td> <td>Presence of installation instruction (True/False) &nbsp;</td> <td>(Manual) Checked the presense of Installation Instruction in the readme or in the project wiki pages.&nbsp;</td> </tr> <tr> <td>project_information &nbsp;</td> <td>Presence of basic project information in README (True/False) &nbsp;</td> <td>(Manual) Checked if the readme have basic information about the project.&nbsp;</td> </tr> <tr> <td>usage_guide</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td>(Manual) Checked the presense of Usage Guide in the readme or in the project wiki pages. For command line tools checked if they have help command which guides how to use the tool. &nbsp;</td> </tr> <tr> <td>test_folder</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td> <p>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/docs/collect_variables/scripts/soft_dev_pract/test_folder.py">test_folder.py</a>) Checks the folder names test/tests in the root directory of the repository.</p> </td> </tr> <tr> <td>requirements_explicit &nbsp;</td> <td>Explicit requirements for Python, R, C++ repositories (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/requirement_explicit.py">requirement_explicit.py</a>) Checks the files (requirements.txt, DESCRIPTION, CMakeLists.txt) in the root directory.&nbsp;</td> </tr> <tr> <td>continuous_integration</td> <td>Indicates if the repository uses continuous integration (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>ci_tool &nbsp;</td> <td>Name of the continuous integration tool used</td> <td>(Script-&nbsp;<a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>add_lint_rule &nbsp;</td> <td>Indicates if additional linting rules are present (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the&nbsp;<br>.github/workflows directory to detect the presence of (linters)&nbsp;Python, R, and C++.</td> </tr> <tr> <td>add_test_rule</td> <td>Indicates if additional testing rules are present (True/False) &nbsp;</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the&nbsp;<br>.github/workflows directory to detect the presence of (testing libraries) Python, R, and C++.</td> </tr> <tr> <td>comment_at_start</td> <td>Indicates the level of comments at the start of the program (most, more, some, less)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/comment_at_start.py">comment_at_start.py</a>) Checks the presence of brief comments at the start at source code files in GitHub repositories.</td> </tr> <tr> <td>language &nbsp;</td> <td>Programming language used in the repository &nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>type &nbsp;</td> <td>Specifies if the profile is a user or organization &nbsp;</td> <td>Github organisation or user profiles.</td> </tr> <tr> <td>organisation &nbsp;</td> <td>Name of the university, company, research institute, or research organization</td> <td>Oraganisation name (from where the user was found)</td> </tr> <tr> <td>research_group</td> <td>Name or acronym of the research group &nbsp;</td> <td>&nbsp;</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Data for publication - https://github.com/Software-Engineering-Group-UP/potsdam-research-repos</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Hong Kong Annotated Airborne LiDAR Point Clouds

<p>The annotated point clouds were generated to train the weakly supervised semantic segmentation algorithm Semantic Query Network (SQN) to classify point clouds <sup>[1]</sup>. The dataset covers 16 tiles of airborne LiDAR data in an area of 7.2 km2&nbsp; in Shatin, Hong Kong, China. 11 tiles were used for training, while 5 tiles were used for validation. There are multiple types of construction in the dataset including high-rise residential buildings, low-rise village houses, and large public buildings. Green spaces are mainly composed of wood areas in open spaces (e.g., in parks and hills) and planted trees in residential gardens and nearby roads. Point clouds are classified in ground, buildings, and trees.</p> <p>The LiDAR data is owned by the Hong Kong government. Please visit the Spatial Data Portal, Survey Division, CEDD (https://sdportal.cedd.gov.hk/#/en/) for more details.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Three Annotated Anomaly Detection Datasets for Line-Scan Algorithms

<h1>Summary</h1> <p>This dataset contains two hyperspectral and one multispectral anomaly detection images, and their corresponding binary pixel masks. They were initially used for real-time anomaly detection in line-scanning, but they can be used for any anomaly detection task.</p> <p>They are in .npy file format (will add tiff or geotiff variants in the future), with the image datasets being in the order of (height, width, channels). The SNP dataset was collected using sentinelhub, and the Synthetic dataset was collected from AVIRIS. The Python code used to analyse these datasets can be found at: https://github.com/WiseGamgee/HyperAD</p> <h1>How to Get Started</h1> <p>All that is needed to load these datasets is Python (preferably 3.8+) and the NumPy package. Example code for loading the Beach Dataset if you put it in a folder called "data" with the python script is:</p> <pre><code>import numpy as np # Load image file hsi_array = np.load("data/beach_hsi.npy") n_pixels, n_lines, n_bands = hsi_array.shape print(f"This dataset has {n_pixels} pixels, {n_lines} lines, and {n_bands}.") # Load image mask mask_array = np.load("data/beach_mask.npy") m_pixels, m_lines = mask_array.shape print(f"The corresponding anomaly mask is {m_pixels} pixels by {m_lines} lines.")</code></pre> <h1>Citing the Datasets</h1> <p>If you use any of these datasets, please cite the following paper:</p> <pre><code>@article{garske2024erx,</code><br><code>&nbsp; title={ERX - a Fast Real-Time Anomaly Detection Algorithm for Hyperspectral Line-Scanning},</code><br><code>&nbsp; author={Garske, Samuel and Evans, Bradley and Artlett, Christopher and Wong, KC},</code><br><code>&nbsp; journal={arXiv preprint arXiv:2408.14947},</code><br><code>&nbsp; year={2024},</code><br><code>}</code></pre> <div> <pre>If you use the beach dataset please cite the following paper as well (original source):</pre> </div> <pre><code>@article{mao2022openhsi, title={OpenHSI: A complete open-source hyperspectral imaging solution for everyone}, author={Mao, Yiwei and Betters, Christopher H and Evans, Bradley and Artlett, Christopher P and Leon-Saval, Sergio G and Garske, Samuel and Cairns, Iver H and Cocks, Terry and Winter, Robert and Dell, Timothy}, journal={Remote Sensing}, volume={14}, number={9}, pages={2244}, year={2022}, publisher={MDPI} }</code></pre>

opencc-by-4.0Aug 2024View details →
zenodo48/100

Metagenome quality metrics and taxonomical annotation visualization through the integration of MAGFlow and BIgMAG (Sup. Material)

<p>Dataset encompassing:</p> <ul> <li>The recovered MAGs by 6 different metagenomics pipelines (ATLAS, DATMA, MetaWRAP, MUFFIN, nf-core/mag and SnakeMAGs) using a mock community as input (SRR8359173 and SRR9328980), complemented with the output from MAGFlow (v1.0.0) using these MAGs as input for their quality assessment and taxonomical annotation.&nbsp;</li> <li>The MAGs produced by nf-core/mag using rice/rhizosphere sequenced libraries (PRJNA663614, PRJNA448773 and PRJNA645385) in either single assembly/single binning or co-assembly/co-binning mode, complemented with the output from MAGFlow (v1.0.0) using these MAGs as input for their quality assessment and taxonomical annotation.</li> <li>Scripts, commands and configuration files to run the different pipelines (ATLAS, DATMA, MetaWRAP, MUFFIN, nf-core/mag and SnakeMAGs) and reproduce the experimental conditions.</li> <li>Outputs, commands and scripts to run Metabinner and Semibin in their default configuration using the rice soil samples co-assembly, along with the MAGFlow (v1.1.0) output to compare these binners against MetaBAT2.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo48/100

Klf14 mouse white adipose tissue histology DeepZoom files and AIDA annotations for visualisation of DeepCytometer white adipocyte segmentations

<p>Latest description of this data set:&nbsp;<a href="https://github.com/MRC-Harwell/cytometer/blob/main/DATA.md">Data.md at cytometer project</a></p> <pre># Publications related to the data The data associated to the DeepCytometer project (https://github.com/MRC-Harwell/cytometer) is available from Zenodo (doi: 10.5281/zenodo.5137433 and 10.5281/zenodo.5149005). The histology and mouse measures were generated as part of the Small et al. 2018 study: &gt; Small et al. &quot;Regulatory variants at KLF14 influence type 2 diabetes risk via a female-specific effect on adipocyte size and body composition&quot;. Nature Genetics, 50:572&ndash;580, 2018. The hand traced data set, colour maps, and automatic segmentations were generated for the Casero et al. 2021 paper: &gt; Casero et al. &quot;Phenotyping of Klf14 mouse white adipose tissue enabled by whole slide segmentation with deep neural networks&quot;. bioRxiv, 2021. doi: [10.1101/2021.06.03.444997](https://www.biorxiv.org/content/10.1101/2021.06.03.444997v1.full). # Data protocols ## Histology and laboratory measures To develop and evaluate our methods we used Klf14tm1(KOMP)Vlcg C57BL/6NTac (B6NTac) mice tissue samples and additional data generated as part of the Small et al. 2018 study(Small et al. 2018). It should be noted that the single exon Klf14 gene is imprinted and only expressed from the maternally inherited allele(Parker-Katiraee et al. 2007). This was taken into account by (Small et al. 2018) by crossing a Het parent with a WT parent, so that each offspring inherited a WT allele from the WT parent, and the Klf14 gene knockout or a WT allele from the other parent (from the father, PAT, or the mother, MAT). We also take Klf14 imprinting into account by using as controls the PAT mice and comparing them to the MAT WT and MAT Het (or functional KO, FKO) mice.&nbsp; We used a total of 76 Klf14-B6NTac mice (nfemale=nmale=38), of which 20 mice from the Control and FKO groups were used for training and testing the DeepCytometer pipeline, as well as the hand traced population experiment (summary in Table MICE). The histopathology screen involved fixing, processing and embedding in wax, sectioning and staining with Hematoxylin and Eosin (H&amp;E) both inguinal subcutaneous and gonadal adipose depots. For paraffin-embedded sections, all samples were fixed in 10% neutral buffered formalin (Surgipath) for at least 48 hours at RT and processed using an Excelsior&trade; AS Tissue Processor (Thermo Scientific). Samples were embedded in molten paraffin wax and 8 &mu;m sections were cut through the respective depots using a Finesse&trade; ME+ microtome (Thermo Scientific). Sampling was conducted at 2sxns per slide, 3 slides per depot block onto simultaneous charged slides, stained with haematoxylin Gill 3 and eosin (Thermo scientific) and scanned using an NDP NanoZoomer Digital pathology scanner (RS C10730 Series; Hamamatsu).&nbsp;Body weight (BW) and depot weight (DW) were measured with Satorius BAL7000 scales. ## White adipose tissue segmentation For cell area quantification, we applied DeepCytometer v8 to 75 inguinal subcutaneous and 72 gonadal whole histology slides with DeepCytometer (with the Corrected method), including the 20 slides sampled for the hand-traced data set, corresponding to 73 females and 74 males, to produce 2,560,067 subcutaneous and 2,467,686 gonadal cells (on average, 34,134 and 34,273 cells per slide, respectively). Full segmentation of all whole slides was performed with script [klf14_b6ntac_exp_0106_full_slide_pipeline_v8.py](https://github.com/MRC-Harwell/cytometer/blob/39358ed1d79df07d1d522b98728c7efd745513f7/scripts/klf14_b6ntac_exp_0106_full_slide_pipeline_v8.py). In this case, the segmentation contours were grouped by tiles in the output AIDA annotation `.json` file (one contour per cell, one file per slide). Non-white adipocyte contours were filtered out, and white adipocyte contours were aggregated into an AIDA annotation `.json` file with a single tile with script [klf14_b6ntac_exp_0106_annotations_postprocessing_v8.py](https://github.com/MRC-Harwell/cytometer/blob/39358ed1d79df07d1d522b98728c7efd745513f7/scripts/klf14_b6ntac_exp_0106_annotations_postprocessing_v8.py) (one contour per cell, one file per slide). # List of directories and files ## Casero et al. (2021) &quot;DeepCytometer pipeline parameter files, Klf14 mouse white adipose tissue histology and hand-traced training contours&quot; (doi: 10.5281/zenodo.5137433) ### `deepcytometer_pipeline_v8.zip` (60.6 MB) Weights, colourmaps, etc. necessary to run the pipeline (v8, with mode colour correction). This is the version of the pipeline described in the paper. There are 10 weight files per convolutional neural network (CNN), corresponding to 10-fold cross-validation * `klf14_b6ntac_exp_0086_cnn_dmap_model_fold_[0..9].h5`: Keras weights for the **EDT CNN** (Histology to Euclidean Distance Transform regression) * `klf14_b6ntac_exp_0089_cnn_segmentation_correction_overlapping_scaled_contours_model_fold_[0..9].h5`: Keras weights for the **Correction CNN** (Segmentation Correction regression) * `klf14_b6ntac_exp_0091_cnn_contour_after_dmap_model_fold_[0..9].h5`: Keras weights for the **Contour CNN** (EDT to Contour detection) * `klf14_b6ntac_exp_0095_cnn_tissue_classifier_fcn_model_fold_[0..9].h5`: Keras weights for the **Tissue CNN** (Pixel-wise tissue classifier) * `klf14_b6ntac_exp_0094_generate_extra_training_images.pickle`: training dataset description * **&#39;file_list&#39;**: list of SVG files with hand-traced contours for network training. Each SVG file has a corresponding TIFF file with the histology used for segmentation * **&#39;idx_test&#39;**: 10 lists with file indices for testing in 10-fold cross-validation * **&#39;idx_train&#39;**: 10 lists with file indices for training in 10-fold cross-validation * **&#39;fold_seed&#39;**: seed number used for the random number generator to assign file indices to folds * `klf14_b6ntac_exp_0098_filename_area2quantile.npz`: quantile colour maps calculated in `klf14_b6ntac_exp_0098_full_slide_size_analysis_v7.py` using the whole Klf14 data set with v7 of the pipeline, and used in earlier experiments, including some where v8 of the pipeline was used for segmentation. * `klf14_b6ntac_exp_0106_filename_area2quantile_v8.npz`: quantile colour maps calculated in `klf14_b6ntac_exp_0106_full_slide_pipeline_v8.py` using the whole Klf14 data set with v8 of the pipeline, and used in later experiments. * `klf14_training_colour_histogram.npz`: statistics from Klf14 histology images to be used in colour correction * **&#39;xbins_edge&#39;**, **&#39;xbins&#39;**: edges and centres of the bins used for histogram calculations * **&#39;hist_r_q1&#39;**, **&#39;hist_r_q2&#39;**, **&#39;hist_r_q3&#39;** * **&#39;hist_g_q1&#39;**, **&#39;hist_g_q2&#39;**, **&#39;hist_g_q3&#39;** * **&#39;hist_b_q1&#39;**, **&#39;hist_b_q2&#39;**, **&#39;hist_b_q3&#39;**: density quartiles (Q1, Q2, Q3) for RGB channels for each bin the histogram * **&#39;mode_r&#39;**, **&#39;mode_g&#39;**, **&#39;mode_b&#39;**: modes for RGB channels (this corresponds to the most typical background colour in the histology images) * **&#39;mean_l&#39;**, **&#39;mean_a&#39;**, **&#39;mean_b&#39;**: mean intensity for L*a*b channels of the image * **&#39;std_l&#39;**, **&#39;std_a&#39;**, **&#39;std_b&#39;**: intensity standard deviations for L*a*b channels of the image * `klf14_exp_0112_training_colour_histogram.npz`: other statistics from Klf14 histology images to be used in colour correction * **&#39;p&#39;**: vector of quantile values used in ECDF calculations * **&#39;val_r_klf14&#39;**, **&#39;val_g_klf14&#39;**, **&#39;val_b_klf14&#39;**: all intensity values for the RGB channels of Klf14 training images that contain at least a white adipocyte * **&#39;f_ecdf_to_val_r_klf14&#39;**, **&#39;f_ecdf_to_val_g_klf14&#39;**, **&#39;f_ecdf_to_val_b_klf14&#39;**: linear interpolation function that maps ECDF quantiles to intensity values in the Klf14 training data set. These functions can be used together with intensity-&gt;quantile interpolation functions calculated for a new histology image to perform histogram matching colour correction * **&#39;mean_klf14&#39;**, **&#39;std_klf14&#39;**: mean and standard deviation of the **&#39;val_r_klf14&#39;**, **&#39;val_g_klf14&#39;**, **&#39;val_b_klf14&#39;** vectors There are also weight files for the pipeline trained with all the data, instead of the 10-fold cross-validation partition. These were not used for the paper, but could be useful for future experiments * `klf14_b6ntac_exp_0101_cnn_dmap_model.h5`: Keras weights for the **EDT CNN** (Histology to Euclidean Distance Transform regression) * `klf14_b6ntac_exp_0104_cnn_segmentation_correction_overlapping_scaled_contours_model.h5`: Keras weights for the **Correction CNN** (Segmentation Correction regression) * `klf14_b6ntac_exp_0102_cnn_contour_after_dmap_model.h5`: Keras weights for the **Contour CNN** (EDT to Contour detection) * `klf14_b6ntac_exp_0103_cnn_tissue_classifier_fcn_model.h5`: Keras weights for the **Tissue CNN** (Pixel-wise tissue classifier) ### `histology.7z` (29.1 GB) 165 H&amp;E histology whole slides from Hamamatsu scanner (`.ndpi`). ### `klf14.7z` (2.3 GB) Mice metadata, training/testing data sets for the pipeline, intermediate files created during training, and neural network weights for multiple experiments. * `klf14_b6ntac_meta_info.csv`: Klf14 mice metadata * **Animal Identifier**, **id:** unique ID for each mouse * **ko_parent:** heterozygous parent of origin for the KO allele (father, PAT or mother, MAT) * **sex:** female or male * **genotype:** wild type (KLF14-KO:WT) or heterozygous (KLF14-KO:Het) * **BW:** body weight (g) * **SC:** subcutaneous depot weight (g) * **gWAT:** gonadal depot weight (g) * **Liver:** livel weight (g) * **cull_age:** age at time of culling (days) * **BW_alive:** body weight measured before culling * **BW_alive_date:** age at time of BW_alive measure * **mother:** unique ID for mouse&#39;s mother * **mother_genotype:** mouse&#39;s mother genotype * `klf14_b6ntac_training`: Directory with hand-traced segmentations of training histology windows. 131 windows sampled from 20 whole slides, plus hand-traced contours that were used for training DeepCytometer and compute population distributions. These segmentations were used for CNN training, but note that there&#39;s a cleaned-up version of these data below, and it was the cleaned-up version that was used for the paper experiments * `ndpifile_row_YYYYYY_col_XXXXXX[.tif/.xcf/.svg]`: * **ndpifile:** name of the whole slide file (e.g. `KLF14-B6NTAC 36.1c PAT 98-16 C1 - 2016-02-11 10.45.00`) * **row_YYYYYY:** Y-coordinate of the top-left corner of the sampling window, in pixels * **col_XXXXXX:** X-coordinate of the top-left corner of the sampling window, in pixels * **.tif:** TIFF file with the histology sampling window * **.xcf:** Gimp file with the histology and hand-traced contours (the contours were drawn in Gimp) * **.svg:** SVG (Scalable Vector Graphics) that contains the hand-traced contours in the XCF file * `klf14_b6ntac_training_v2`: Same as `klf14_b6ntac_training`, but the hand-traced data set was cleaned up to remove small contours of dubious cells, or cells that are fully overlapped by others * `klf14_b6ntac_training_non_overlap`: Directory with intermediate images to train the networks. These images are generated by script [`klf14_b6ntac_training_non_overlap`](https://github.com/MRC-Harwell/cytometer/blob/main/scripts/klf14_b6ntac_exp_0077_generate_non_overlap_training_images.py) * `klf14_b6ntac_training_augmented`: Directory with intermediate images used to train the networks (using augmentation to reduce overfitting). These images are generated by script [`klf14_b6ntac_exp_0078_generate_augmented_training_images.py`](https://github.com/MRC-Harwell/cytometer/blob/main/scripts/klf14_b6ntac_exp_0078_generate_augmented_training_images.py) * `klf14_b6ntac_seg`: Deprecated. Directory to store whole slide coarse segmentations in old experiments (e.g. `klf14_b6ntac_exp_0076_generate_training_images.py`). Of little interest for most users * `klf14_b6ntac_results`: Deprecated. Directory to store miscellanea output from some experiments. Of little interest for most users ## Casero et al. (2021). &quot;Klf14 mouse white adipose tissue histology DeepZoom files and AIDA annotations for visualisation of DeepCytometer white adipocyte segmentations&quot; (doi: 10.5281/zenodo.5149005) ### `aida_data_Klf14_v8_images.7z` (16.9 GB) Histology images converted to DeepZoom so that they can be visualised with [AIDA](https://github.com/alanaberdeen/AIDA). To use this, decompress this file and put the resulting `images` directory in your `AIDA/dist/data/` directory. ### `aida_data_Klf14_v8_annotations.7z` (18 GB) White adipocyte segmentations in AIDA annotation `.json` files (one contour per cell, one file per whole slide). Each slide has the following files: * `SLIDENAME.json`: Soft link to the annotations file that we want to associate to slide `SLIDENAME.ndpi`, e.g. `SLIDENAME` = `KLF14-B6NTAC-PAT-39.2d 454-16 B1 - 2016-03-17 12.16.06` * `SLIDENAME.lock`: Empty file used to tell the pipeline that `SLIDENAME.ndpi` has already been processed or is being currently processed * `SLIDENAME_coarse_mask.npz`: File with the coarse tissue segmentation of `SLIDENAME.ndpi` and the internal state of the pipeline (execution times, steps, etc) * `SLIDENAME_exp_0106_auto.json`: Annotations (all segmentations without filtering from the Auto algorithm, i.e. segmentation without object overlap). Contours are grouped by the tile they were processed in * `SLIDENAME_exp_0106_auto_aggregated.json`: Filtered annotations (non-white adipocytes removed) of the Auto algorithm. All contours aggregated into a single tile * `SLIDENAME_exp_0106_corrected.json`: Annotations (all segmentations without filtering from the Corrected algorithm, i.e. segmentation with object overlap). Contours are grouped by the tile they were processed in * `SLIDENAME_exp_0106_corrected_aggregated.json`: Filtered annotations (non-white adipocytes removed) of the Corrected algorithm. All contours aggregated into a single tile To use this, decompress this file and put the resulting `annotations` directory in your `AIDA/dist/data/` directory.</pre>

opencc-by-4.0Jul 2021View details →
zenodo48/100

French Entity-Linking dataset between annotated tweets collected during major crises in France and French Wikipedia corpus

<p>Most of the available datasets are not particularly adapted to our target application: geolocate natural disasters from social networks. First, social media posts are largely underrepresented in these datasets, and the only Twitter dataset lacks Entity-Linking annotations. Second, none of the datasets focuses on a crisis or natural disaster event.</p> <p>To mitigate these issues, we extracted a collection of French tweets written during earthquakes and major floods that have occurred in France in recent years. We set up Label-Studio in order to annotate these tweets. A total of 4617 tweets were annotated, including 1678 tweets posted during earthquakes and 2939 during floods. For each annotated tweet, mentions were annotated using the set of labels described earlier in the paper as well as, when possible, the target Wikipedia title.</p> <p>Named &ldquo;R&eacute;SoCIO&rdquo; in reference to the research project in which it was carried out, the dataset resulting from this work contains a total of 12 828 annotated mentions and 1 513 distinct Wikipedia entities. 85% of mentions were associated with a Wikipedia page and 94 % if we ignore the RISKNAT and DAMAGES labels, which are often difficult to map to an existing entity.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entities</strong></td> </tr> <tr> <td>PERSON</td> <td>315</td> <td>263</td> <td>136</td> </tr> <tr> <td>ORG</td> <td>863</td> <td>790</td> <td>281</td> </tr> <tr> <td>GEOLOC</td> <td>4375</td> <td>4234</td> <td>701</td> </tr> <tr> <td>TRANSPORT</td> <td>250</td> <td>203</td> <td>101</td> </tr> <tr> <td>EVENT</td> <td>35</td> <td>21</td> <td>16</td> </tr> <tr> <td>FACILITY</td> <td>129</td> <td>94</td> <td>49</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>128</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>223</td> <td>200</td> <td>46</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>12828</strong></td> <td><strong>1322</strong></td> <td><strong>1513</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the Twitter dataset. #Mentions&nbsp;shows the total number of mentions per label, #Linked the number of mentions linked&nbsp;to an entity and #Entities the number of distinct entities per label present in the&nbsp;dataset.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entitie</strong>s</td> </tr> <tr> <td>PERSON</td> <td>1100102</td> <td>1098406</td> <td>557697</td> </tr> <tr> <td>ORG</td> <td>750925</td> <td>749504</td> <td>130394</td> </tr> <tr> <td>GEOLOC</td> <td>2729702</td> <td>2728296</td> <td>215924</td> </tr> <tr> <td>TRANSPORT</td> <td>161539</td> <td>160487</td> <td>53405</td> </tr> <tr> <td>EVENT</td> <td>798433</td> <td>798251</td> <td>86471</td> </tr> <tr> <td>FACILITY</td> <td>258835</td> <td>258513</td> <td>109867</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>127</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>4340621</td> <td>4339658</td> <td>682458</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>10146795</strong></td> <td><strong>10138230</strong></td> <td><strong>1836399</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the full dataset. #Mentions shows&nbsp;the total number of mentions per label, #Linked the number of mentions linked to an&nbsp;entity and #Entities the number of distinct entities per label present in the dataset.</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Tweets containing "climate change" with topic annotations

<p>This dataset contains the Twitter IDs of all ~20M tweets containing the phrase "climate change" 2018-2021. Additionally, it contains the topical annotations and 2D semantic representation of our thematic analysis based on ~980 topic clusters that are grouped by hand into seven themes (COVID-19, Politics, Contrarian, Movements, Solutions, Impacts, Causes) as well as "non-relevant/spam", "others", and highlighting of potentially interesting topics.</p> <p>Code and additional notes are available on GitHub:&nbsp;https://github.com/TimRepke/twitter-climate</p> <p>The topics, including statistics and the annotator labels for broader themes (aka "super topics") are contained in the spreadsheet. This data is extrapolated to the tweets contained in the share.jsonl file containing&nbsp;one json object per line with the following fields:</p> <ul> <li><strong>'rel':</strong>&nbsp;true iff Tweet is contained in analysis</li> <li><strong>'filters':</strong> null if Tweet is not included, otherwise contains an object with "reasons" why this tweet was excluded <ul> <li><strong>'dup'</strong>:&nbsp;1 iff this is a duplicate (excl first)</li> <li><strong>'lan':</strong>&nbsp;&nbsp;1 iff language is English (and not None)</li> <li><strong>'txt': </strong>1 iff status text is not None</li> <li><strong>'mit':&nbsp;</strong>1 iff text has minimum number of tokens (&gt;=4)</li> <li><strong>'mah'</strong>:&nbsp;1 iff text has less than maximum number of hashtags (&lt;=5),</li> <li><strong>'pfd'</strong>:&nbsp;1 iff tweet was posted after 01.01.2018</li> <li><strong>'ptd'</strong>: 1 iff tweet was posted before 31.12.2021</li> <li><strong>'cli':</strong>&nbsp;1 iff tweet actually contains "climate change" (API matches some false positives)</li> </ul> </li> <li><strong>'ann'</strong>: null if Tweet is not included, otherwise contains an object with topic annotations <ul> <li><strong>'t_km': </strong>&nbsp;topic (based on "keep &amp; majority vote" strategy)</li> <li><strong>'t_kp': </strong>&nbsp;topic (based on "keep &amp; closest topic centroid [proximity]" strategy)</li> <li><strong>'t_fm': </strong>&nbsp;topic (based on "drop sample topic [fresh] &amp; majority vote" strategy)</li> <li><strong>'t_fp':</strong> &nbsp;topic (based on "drop sample topic [fresh] &amp; closest topic centroid [proximity]")</li> <li><strong>'st_int':</strong> &nbsp;theme annotation "Interesting"</li> <li><strong>'st_nr': </strong>&nbsp;theme annotation "Non-relevant / spam"</li> <li><strong>'st_cov':</strong> &nbsp;theme annotation "COVID"</li> <li><strong>'st_pol': </strong>&nbsp;theme annotation "Politics"</li> <li><strong>'st_mov': </strong>&nbsp;theme annotation "Movements"</li> <li><strong>'st_imp':</strong> &nbsp;theme annotation "Impacts"</li> <li><strong>'st_cau': </strong>&nbsp;theme annotation "Causes"</li> <li><strong>'st_sol': </strong>&nbsp;theme annotation "Solutions"</li> <li><strong>'st_con': &nbsp;</strong>theme annotation "Contrarian"</li> <li><strong>'st_oth': </strong>&nbsp;theme annotation "Other"</li> <li><strong>'x': </strong>&nbsp;x position in 2D representation</li> <li><strong>'y': </strong>&nbsp;x position in 2D representation</li> <li><strong>'sample':</strong> &nbsp;true iff this tweet was in the original topic model sample</li> </ul> </li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo48/100

OcWikiAnnot: Annotated Wikipedia Corpus of Occitan

<p>OcWikiAnnot is a corpus of Wikipedia content in Occitan that is tokenized, PoS-tagged and lemmatized. The corpus contains 100 000 sentences for a total of 2&nbsp;037&nbsp;723 tokens. It is based on the Wikipedia corpus in Occitan that is part of the <a href="https://corpora.uni-leipzig.de/en?corpusId=oci_wikipedia_2021">Leipzig Corpora Collection</a>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

TG-CSR Annotations

<p>Individual raw and normalized label data for the TG-CSR (Theoretically-Grounded Commonsense Reasoning) benchmark.</p>

opencc-by-4.0May 2023View details →
zenodo48/100

A Curated Gene and Biological System Annotation of Adverse Outcome Pathways Related to Human Health

<p>Adverse Outcome Pathways (AOPs) are multi-scale models of biological mechanisms connecting molecular initiating events to adverse outcomes through measurable key events.&nbsp;AOPs can guide the use and development of new approach methodologies (NAMs) aimed at reducing animal experimentation in chemical safety assessment. Here, we present a comprehensive molecular annotation of AOPs relevant to human health to embed the AOP framework into molecular data interpretation, which supports the development and application of novel AOP-based approaches in biomedical research.</p> <p>Please cite the following publication alongside this Zenodo entry when using the data:</p> <p>Saarim&auml;ki, L.A., Fratello, M., Pavel, A.&nbsp;<em>et al.</em>&nbsp;A curated gene and biological system annotation of adverse outcome pathways related to human health.&nbsp;<em>Sci Data</em>&nbsp;<strong>10</strong>, 409 (2023). https://doi.org/10.1038/s41597-023-02321-w</p>

opencc-by-4.0Oct 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record