Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
271
datasets available to search
ShareScore release 0.7.1
Dataset results
271 results for “annotated dataset”
Occurrence Record Dataset from "Annotated checklist of the bees of Bonaire, with a focus on host plants"
<p>This is the occurrence dataset created for the publication "Annotated checklist of the bees of Bonaire, with a focus on host plants" (<a href="https://natuurtijdschriften.nl/pub/1026875" target="_blank" rel="noopener">https://natuurtijdschriften.nl/pub/1026875</a>).</p> <p>Observation and specimen data were assembled for this dataset, with the majority of records obtained during the Bonaire Estafette Expeditie (BEE). All citizen science records from Observation.org and iNaturalist.org up to December 2023 have been critically reviewed.<br>A project was created (<a href="https://www.inaturalist.org/projects/flower-visitors-and-pollinators-of-the-caribbean" target="_blank" rel="noopener">Flower visitors and pollinators of the Caribbean</a>) to improve standardized data collecting of plant-pollinator interactions and on <a href="https://observation.org/">observation.org</a> the standardized fields for interactions were used.<br>Records from passive trapping methods are not included. All bees were either observed or collected by hand or insect net. The majority of specimens will be accessible in the collection of Naturalis Biodiversity Center (RMNH), Leiden (the Netherlands). A synoptic collection is retained at the University of Tartu Zoological Collections in Tartu, Estonia (TUZ).</p> <p>The occurrence dataset (Version 1.4 and later) is:</p> <ul> <li>conform Darwin Core (DwC): <a href="https://dwc.tdwg.org/terms/">https://dwc.tdwg.org/terms</a></li> <li>in the data format CSV (tab delimited values) and UTF-8 encoded</li> </ul> <p> </p> <p><strong>DwC terms (Column labels) used in the dataset with their description:</strong></p> <table> <tbody> <tr> <td><strong>Column label</strong></td> <td><strong>Column description</strong></td> </tr> <tr> <td>occurrenceID</td> <td>Unique identifier or URI (GUID) for each record, mainly unique URLs generated by the web-based data holder.</td> </tr> <tr> <td>catalogNumber</td> <td>Unique code derived from URI in occurrenceID. Each specimen bears a label with this identifier and multimedia are tagged with this identifier.</td> </tr> <tr> <td>recordNumber</td> <td>Sample field ID used to manage data of preserved specimen occurrence records.</td> </tr> <tr> <td>otherCatalogNumbers</td> <td>Other unique identifiers used on specimen labels, but not derived from an URI.</td> </tr> <tr> <td>scientificName</td> <td>The scientific name of the lowest taxonomic rank to which the individual(s) was identified.</td> </tr> <tr> <td>scientificNameAuthorship</td> <td>The author name and year of publication in accordance with ICZN rules.</td> </tr> <tr> <td>verbatimIdentification</td> <td>The original identification, including qualifiers if needed.</td> </tr> <tr> <td>individualCount</td> <td>The number of individuals present at the time of the occurrence.</td> </tr> <tr> <td>sex</td> <td>The sex of the individual(s). The values female, male or unknown are used, if a mixed group is observed multiple values are listed.</td> </tr> <tr> <td>lifeStage</td> <td>The life stage of the individual(s).</td> </tr> <tr> <td>basisOfRecord</td> <td>The specific nature of the data record at the time of the identification (e.g. PreservedSpecimen).</td> </tr> <tr> <td>identifiedBy</td> <td>The name of the person who made the identification in the field or based on collected evidence (e.g. specimen or photo).</td> </tr> <tr> <td>identificationQualifier</td> <td>In case the identification could be given only to a species group 'cf.' is recorded.</td> </tr> <tr> <td>dateIdentified</td> <td>The year when the identification was made.</td> </tr> <tr> <td>previousIdentifications</td> <td>The scientific name originally given to the observed or collected individual(s).</td> </tr> <tr> <td>order</td> <td>The name of the order (e.g. Hymenoptera).</td> </tr> <tr> <td>family</td> <td>The name of the family (e.g. Apidae).</td> </tr> <tr> <td>genus</td> <td>The name of the genus (e.g. Apis).</td> </tr> <tr> <td>subgenus</td> <td>The name of the subgenus (e.g. Apis).</td> </tr> <tr> <td>specificEpithet</td> <td>The name of the species, epithet as given in dwc:scientificName.</td> </tr> <tr> <td>taxonRank</td> <td>The taxonomic rank of the most specific name in dwc:scientificName.</td> </tr> <tr> <td>eventDate</td> <td>The date-time when the event was observed and recorded. The event date uses the ISO 8601-1:2019 standard, with the following formatting being used: format YYYY-MM-DD, or YYYY if only the year is known. If time of capture is known, then format is YYYY-MM-DDTHH:MM, with HH:MM the local time.</td> </tr> <tr> <td>year</td> <td>The year in which the event was observed and recorded.</td> </tr> <tr> <td>month</td> <td>The month in which the event was observed and recorded.</td> </tr> <tr> <td>day</td> <td>The day in which the event was observed and recorded.</td> </tr> <tr> <td>eventTime</td> <td>The time or interval during which the event occurred.</td> </tr> <tr> <td>samplingProtocol</td> <td>The name or description of the collecting or recording method used.</td> </tr> <tr> <td>behavior</td> <td>A description of the behavior shown by the individual(s) recorded in this occurrence.</td> </tr> <tr> <td>decimalLatitude</td> <td>The geographic latitude in decimal degrees recorded by a GPS device (WGS84) when observing and recording the occurrence.</td> </tr> <tr> <td>decimalLongitude</td> <td>The geographic longitude in decimal degrees recorded by a GPS device (WGS84) when observing and recording the occurrence.</td> </tr> <tr> <td>geodeticDatum</td> <td>The ellipsoid, geodetic datum, or spatial reference system (SRS) upon which the geographic coordinates given in dwc:decimalLatitude and dwc:decimalLongitude is based.</td> </tr> <tr> <td>verbatimLocality</td> <td>The original textual description of the place.</td> </tr> <tr> <td>island</td> <td>The name of the island.</td> </tr> <tr> <td>countryCode</td> <td>The standard ISO 3166-1 alpha-2 country code for the country.</td> </tr> <tr> <td>coordinateUncertaintyInMeters</td> <td> <p>The horizontal distance (in meters) from the given dwc:decimalLatitude and dwc:decimalLongitude describing the smallest circle containing the actual location, usually the EPE (Estimated Position Error) from the GPS device. The EPE is here measured as the horizontal position error in meters.</p> </td> </tr> <tr> <td>recordedBy</td> <td>A person, group, or organization observing and recording the occurrence.</td> </tr> <tr> <td>associatedTaxa</td> <td>The type of association and the scientific name of the host taxon is recorded that is associated/has relationship with the taxon in dwc:scientificName. The association/relationship is recorded using the format as in the following example: "floral host":"Lantana sp."</td> </tr> <tr> <td>occurrenceRemarks</td> <td>Comments or notes about the dwc:Occurrence.</td> </tr> <tr> <td>associatedSequences</td> <td>A list (concatenated and separated) of identifiers (publication, global unique identifier, URI) of genetic sequence information.</td> </tr> <tr> <td>typeStatus</td> <td>A list (concatenated and separated) of nomenclatural types (type status, typified scientific name, publication) applied to the subject.</td> </tr> <tr> <td>collectionCode</td> <td>The name, acronym, coden, or initialism identifying the collection or data set from which the record was derived.</td> </tr> <tr> <td>identificationRemarks</td> <td>Comments or notes about the identification.</td> </tr> <tr> <td>identificationReferences</td> <td>A reference or list of references (publication, global unique identifier, URI) used for the identification.</td> </tr> <tr> <td>nameAccordingTo</td> <td>A reference to the checklist or publication that was followed to record the name in dwc:scientificName.</td> </tr> <tr> <td>samplingEffort</td> <td>The amount of effort, expressed in minutes or hours, to obtain and record the occurrences.</td> </tr> <tr> <td>occurrenceStatus</td> <td>A statement about the presence or absence of a taxon during the time of an event.</td> </tr> <tr> <td>disposition</td> <td>The current state of a specimen with respect to a collection.</td> </tr> <tr> <td>language</td> <td>The language of the record using ISO 639-1 codes, e.g. en</td> </tr> </tbody> </table>
Mappings for "Developing a Scalable Annotation Method for Large Datasets That Enhances Alarms With Actionability Data to Increase Informativeness: Mixed Methods Approach"
<p>Studies identified false and non-actionnable alarms as a factor for alarm fatigue in intensive care units.</p> <p>To annotate patient alarms, and analyse the alarm situation in intensive care units, we conceptualized and performed data mappings related to airway management and medication interventions. The mappings were based on information retrieved from the patient data management system (PDMS) and clinical expertise. For the airway management mappings, we used additional resources such as ISO 19223:2019 or ventilator instruction manuals. The mappings do not include patient data.</p> <p>As the mappings are generic, they could be used in other contexts than alarm annotation and research.</p> <p><strong>1. Respiratory Management Mappings:</strong></p> <ul> <li>General tables summarizing the 1) categories based on ISO 19223:2019 to describe respiratory support therapies (RSTs), 2) defining the invasiveness level of a RST and 3) listing the abbreviations used in the mappings</li> <li> <p>Tables including PDMS entries for airway devices (ADs), ventilation devices (VDs), and ventilation modes (VMs)</p> </li> <li> <p>Mapping of AD entries (from the PDMS) to defined categories</p> </li> <li> <p>Mapping of VDs, VMs, and ADs to defined RSTs, including information on invasiveness</p> </li> <li> <p>Table specifying suitable ventilation parameters in the context of each RST</p> </li> </ul> <p><strong>2. Medication Mappings:</strong></p> <ul> <li> <p>General tables providing information on physiological alarm conditions (PACs), interventions, routes, and techniques of administration of interest</p> </li> <li> <p>Mapping of routes of administration to techniques of administration including PDMS entries</p> </li> <li> <p>Mapping of active ingredients (including SNOMED CT Fully Specified Names and Identifiers), related PDMS information, and routes and techniques of administration to defined PAC and interventions</p> </li> </ul>
DATASET: De novo assembly and functional annotation of the heart + hemolymph transcriptome in the Caribbean spiny lobster Panulirus argus
<p>The spiny lobster <em>Panulirus argus</em> is an ecologically relevant species in shallow water coral reefs and target of the most lucrative fishery in the greater Caribbean region. This study reports, for the first time, the heart + hemolymph transcriptome of the Caribbean spiny lobster<em> Panulirus argus</em> assembled from short Illumina 150 bp PE raw reads. A total 80,152,094 raw reads were assembled using the Oyster River Protocol pipeline that aspires to become the standard protocol for <em>de novo</em> transcriptome assembly. The assembly resulted in a total of 254,773 transcripts. Functional gene annotation was conducted using the software package 'dammit' that also aspires to become the standard protocol for <em>de novo</em> transcriptome annotation. Lastly, gene enrichment analyses were conducted using the Gene Ontology (GO), KEGG pathway analyses (Kaas), and KOG (WebMGA) databases. This resource will be of utmost importance in future research aiming at exploring the effect of local and regional anthropogenic disturbances as well as global climate change on the molecular physiology of this overexploited species.</p>
CODE-test: An annotated 12-lead ECG dataset
<pre># Annotated 12 lead ECG dataset Contain 827 ECG tracings from different patients, annotated by several cardiologists, residents and medical students. It is used as test set on the paper: "Automatic diagnosis of the 12-lead ECG using a deep neural network". https://www.nature.com/articles/s41467-020-15432-4. It contain annotations about 6 different ECGs abnormalities: - 1st degree AV block (1dAVb); - right bundle branch block (RBBB); - left bundle branch block (LBBB); - sinus bradycardia (SB); - atrial fibrillation (AF); and, - sinus tachycardia (ST). Companion python scripts are available in: https://github.com/antonior92/automatic-ecg-diagnosis -------- Citation ``` Ribeiro, A.H., Ribeiro, M.H., Paixão, G.M.M. et al. Automatic diagnosis of the 12-lead ECG using a deep neural network. Nat Commun 11, 1760 (2020). https://doi.org/10.1038/s41467-020-15432-4 ``` Bibtex: ``` @article{ribeiro_automatic_2020, title = {Automatic Diagnosis of the 12-Lead {{ECG}} Using a Deep Neural Network}, author = {Ribeiro, Ant{\^o}nio H. and Ribeiro, Manoel Horta and Paix{\~a}o, Gabriela M. M. and Oliveira, Derick M. and Gomes, Paulo R. and Canazart, J{\'e}ssica A. and Ferreira, Milton P. S. and Andersson, Carl R. and Macfarlane, Peter W. and Meira Jr., Wagner and Sch{\"o}n, Thomas B. and Ribeiro, Antonio Luiz P.}, year = {2020}, volume = {11}, pages = {1760}, doi = {https://doi.org/10.1038/s41467-020-15432-4}, journal = {Nature Communications}, number = {1} } ``` ----- ## Folder content: - `ecg_tracings.hdf5`: The HDF5 file containing a single dataset named `tracings`. This dataset is a `(827, 4096, 12)` tensor. The first dimension correspond to the 827 different exams from different patients; the second dimension correspond to the 4096 signal samples; the third dimension to the 12 different leads of the ECG exams in the following order: `{DI, DII, DIII, AVR, AVL, AVF, V1, V2, V3, V4, V5, V6}`. The signals are sampled at 400 Hz. Some signals originally have a duration of 10 seconds (10 * 400 = 4000 samples) and others of 7 seconds (7 * 400 = 2800 samples). In order to make them all have the same size (4096 samples) we fill them with zeros on both sizes. For instance, for a 7 seconds ECG signal with 2800 samples we include 648 samples at the beginning and 648 samples at the end, yielding 4096 samples that are them saved in the hdf5 dataset. All signal are represented as floating point numbers at the scale 1e-4V: so it should be multiplied by 1000 in order to obtain the signals in V. In python, one can read this file using the following sequence: ```python import h5py with h5py.File(args.tracings, "r") as f: x = np.array(f['tracings']) ``` - The file `attributes.csv` contain basic patient attributes: sex (M or F) and age. It contain 827 lines (plus the header). The i-th tracing in `ecg_tracings.hdf5` correspond to the i-th line. - `annotations/`: folder containing annotations csv format. Each csv file contain 827 lines (plus the header). The i-th line correspond to the i-th tracing in `ecg_tracings.hdf5` correspond to the in all csv files. The csv files all have 6 columns `1dAVb, RBBB, LBBB, SB, AF, ST` corresponding to weather the annotator have detect the abnormality in the ECG (`=1`) or not (`=0`). 1. `cardiologist[1,2].csv` contain annotations from two different cardiologist. 2. `gold_standard.csv` gold standard annotation for this test dataset. When the cardiologist 1 and cardiologist 2 agree, the common diagnosis was considered as gold standard. In cases where there was any disagreement, a third senior specialist, aware of the annotations from the other two, decided the diagnosis. 3. `dnn.csv` prediction from the deep neural network described in the paper. THe threshold is set in such way it maximizes the F1 score. 4. `cardiology_residents.csv` annotations from two 4th year cardiology residents (each annotated half of the dataset). 5. `emergency_residents.csv` annotations from two 3rd year emergency residents (each annotated half of the dataset). 6. `medical_students.csv` annotations from two 5th year medical students (each annotated half of the dataset). </pre>
Public metagenome datasets annotated using SingleM
<p>These data underlie the community profiles shown at <a href="https://sandpiper.qut.edu.au">https://sandpiper.qut.edu.au</a></p> <p> </p> <h2>Changelog</h2> <p>version 1.0.0</p> <ul> <li>Public metagenomes published before Feb 20, 2025 were analysed using SingleM pipe v0.18.3 (the default R220 metapackage), and then renewed using an R226 metapackage (S5.4.0.GTDB_r226.metapackage_20250331).</li> </ul> <p>version 0.3.0</p> <ul> <li>Update profiles to use GTDB R220, generated using SingleM renew v0.17.0.</li> </ul> <p>version 0.2.0</p> <ul> <li>Initial version. Created using a GTDB R214-based reference SingleM metapackage S3.2.1.GTDB_r214.metapackage_20231006 based on public datasets available Dec 15, 2021.</li> </ul>
scRNA-seq atlases for 3 Caenorhabditis species - annotated cell datasets
<p>Annotated datasets (monocle3 objects) of scRNA-seq data for <em>C. elegans</em>, <em>C. briggsae</em> and <em>C. tropicalis</em> L2 nematodes. The datasets are published together with the manuscript "Divergence in neuronal signaling pathways despite conserved neuronal identity among <em>Caenorhabditis</em> species".</p> <p><a href="https://doi.org/10.1016/j.cub.2025.05.036" target="_blank" rel="noopener">https://doi.org/10.1016/j.cub.2025.05.036</a></p> <p>Files deposited include cell datasets for all sequenced cells ("all_cds") and datasets for all cells annotated as neurons ("neu_cds"). </p> <p><em>C. elegans</em> strain - N2.</p> <p><em>C. briggsae</em> strain - AF16.</p> <p><em>C. tropicalis</em> strain - NIC203.</p>
Curlie Enhanced with LLM Annotations: Two Datasets for Advancing Homepage2Vec's Multilingual Website Classification
<h3>Advancing Homepage2Vec with LLM-Generated Datasets for Multilingual Website Classification</h3> <p>This dataset contains two subsets of labeled website data, specifically created to enhance the performance of Homepage2Vec, a multi-label model for website classification. The datasets were generated using Large Language Models (LLMs) to provide more accurate and diverse topic annotations for websites, addressing a limitation of existing Homepage2Vec training data.</p> <p><strong>Key Features:</strong></p> <ul> <li><strong>LLM-generated annotations:</strong> Both datasets feature website topic labels generated using LLMs, a novel approach to creating high-quality training data for website classification models.</li> <li><strong>Improved multi-label classification:</strong> Fine-tuning Homepage2Vec with these datasets has been shown to improve its macro F1 score from 38% to 43% evaluated on a human-labeled dataset, demonstrating their effectiveness in capturing a broader range of website topics.</li> <li><strong>Multilingual applicability:</strong> The datasets facilitate classification of websites in multiple languages, reflecting the inherent multilingual nature of Homepage2Vec.</li> </ul> <p><strong>Dataset Composition:</strong></p> <ul> <li><strong>curlie-gpt3.5-10k:</strong> 10,000 websites labeled using GPT-3.5, context 2 and 1-shot</li> <li><strong>curlie-gpt4-10k:</strong> 10,000 websites labeled using GPT-4, context 2 and zero-shot</li> </ul> <p><strong>Intended Use:</strong></p> <ul> <li>Fine-tuning and advancing Homepage2Vec or similar website classification models</li> <li>Research on LLM-generated datasets for text classification tasks</li> <li>Exploration of multilingual website classification</li> </ul> <p><strong>Additional Information:</strong></p> <ul> <li><strong>Project and report repository:</strong> https://github.com/CS-433/ml-project-2-mlp</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>This dataset was created as part of a project at EPFL's Data Science Lab (DLab) in collaboration with <a href="https://people.epfl.ch/robert.west">Prof. Robert West</a> and <a href="https://tizianopiccardi.github.io/" rel="nofollow">Tiziano Piccardi.</a></p>
Sentinel2GlobalLULC: A dataset of Sentinel-2 georeferenced RGB imagery annotated for global land use/land cover mapping with deep learning (License CC BY 4.0)
<p>Sentinel2GlobalLULC is a deep learning-ready dataset of RGB images from the Sentinel-2 satellites designed for global land use and land cover (LULC) mapping. Sentinel2GlobalLULC v2.1 contains 194,877 images in GeoTiff and JPEG format corresponding to 29 broad LULC classes. Each image has 224 x 224 pixels at 10 m spatial resolution and was produced by assigning the 25th percentile of all available observations in the Sentinel-2 collection between June 2015 and October 2020 in order to remove atmospheric effects (i.e., clouds, aerosols, shadows, snow, etc.). A spatial purity value was assigned to each image based on the consensus across 15 different global LULC products available in Google Earth Engine (GEE). </p> <p> </p> <p>Our dataset is structured into 3 main zip-compressed folders, an Excel file with a dictionary for class names and descriptive statistics per LULC class, and a python script to convert RGB GeoTiff images into JPEG format. The first folder called "Sentinel2LULC_GeoTiff.zip" contains 29 zip-compressed subfolders where each one corresponds to a specific LULC class with hundreds to thousands of GeoTiff Sentinel-2 RGB images. The second folder called "Sentinel2LULC_JPEG.zip" contains 29 zip-compressed subfolders with a JPEG formatted version of the same images provided in the first main folder. The third folder called "Sentinel2LULC_CSV.zip" includes 29 zip-compressed CSV files with as many rows as provided images and with 12 columns containing the following metadata (this same metadata is provided in the image filenames): </p> <ul> <li>Land Cover Class ID: is the identification number of each LULC class</li> <li>Land Cover Class Short Name: is the short name of each LULC class</li> <li>Image ID: is the identification number of each image within its corresponding LULC class </li> <li>Pixel purity Value: is the spatial purity of each pixel for its corresponding LULC class calculated as the spatial consensus across up to 15 land-cover products </li> <li>GHM Value: is the spatial average of the Global Human Modification index (gHM) for each image</li> <li>Latitude: is the latitude of the center point of each image</li> <li>Longitude: is the longitude of the center point of each image</li> <li>Country Code: is the Alpha-2 country code of each image as described in the ISO 3166 international standard. To understand the country codes, we recommend the user to visit the following website where they present the Alpha-2 code for each country as described in the ISO 3166 international standard:https: //www.iban.com/country-codes</li> <li>Administrative Department Level1: is the administrative level 1 name to which each image belongs</li> <li>Administrative Department Level2: is the administrative level 2 name to which each image belongs</li> <li>Locality: is the name of the locality to which each image belongs</li> <li>Number of S2 images : is the number of found instances in the corresponding Sentinel-2 image collection between June 2015 and October 2020, when compositing and exporting its corresponding image tile</li> </ul> <p>For seven LULC classes, we could not export from GEE all images that fulfilled a spatial purity of 100% since there were millions of them. In this case, we exported a stratified random sample of 14,000 images and provided an additional CSV file with the images actually contained in our dataset. That is, for these seven LULC classes, we provide these 2 CSV files:</p> <ul> <li>A CSV file that contains all exported images for this class </li> <li>A CSV file that contains all images available for this class at spatial purity of 100%, both the ones exported and the ones not exported, in case the user wants to export them. These CSV filenames end with "including_non_downloaded_images".</li> </ul> <p>To clearly state the geographical coverage of images available in this dataset, we included in the version v2.1, a compressed folder called "Geographic_Representativeness.zip". This zip-compressed folder contains a csv file for each LULC class that provides the complete list of countries represented in that class. Each csv file has two columns, the first one gives the country code and the second one gives the number of images provided in that country for that LULC class. In addition to these 29 csv files, we provided another csv file that maps each ISO Alpha-2 country code to its original full country name.</p> <p>© <a href="https://doi.org/10.5281/zenodo.5055632">Sentinel2GlobalLULC Dataset </a>by Yassir Benhammou, Domingo Alcaraz-Segura, Emilio Guirado, Rohaifa Khaldi, Boujemâa Achchab, Francisco Herrera & Siham Tabik is marked with Attribution 4.0 International (CC-BY 4.0)</p>
StreetSurfaceVis: a dataset of street-level imagery with annotations of road surface type and quality
<h1>StreetSurfaceVis</h1> <p><em>StreetSurfaceVis</em> is an image dataset containing <strong>9,122 street-level images from Germany</strong> with labels on <strong>road surface type and quality.</strong> The CSV file <code>streetSurfaceVis_v1_0.csv</code> contains all image metadata and four folders contain the image files. All images are available in four different sizes, based on the image width, in 256px, 1024px, 2048px and the original size.<br>Folders containing the images are named according to the respective image size. Image files are named based on the <code>mapillary_image_id</code>.</p> <p>You can find the corresponding publication here: <a href="https://www.nature.com/articles/s41597-024-04295-9#citeas">StreetSurfaceVis: a dataset of crowdsourced street-level imagery with semi-automated annotations of road surface type and quality</a></p> <p> </p> <h3>Image metadata</h3> <p>Each CSV record contains information about one street-level image with the following attributes:</p> <ul> <li><code>mapillary_image_id</code>: ID provided by Mapillary (see information below on Mapillary)</li> <li><code>user_id</code>: Mapillary user ID of contributor</li> <li><code>user_name</code>: Mapillary user name of contributor</li> <li><code>captured_at</code>: timestamp, capture time of image</li> <li><code>longitude</code>, <code>latitude</code>: location the image was taken at</li> <li><code>train</code>: Suggestion to split train and test data. `True` for train data and `False` for test data. Test data contains data from 5 cities which are excluded in the training data.</li> <li><code>surface_type</code>: Surface type of the road in the focal area (the center of the lower image half) of the image. Possible values: asphalt, concrete, paving_stones, sett, unpaved</li> <li><code>surface_quality</code>: Surface quality of the road in the focal area of the image. Possible values: (1) excellent, (2) good, (3) intermediate, (4) bad, (5) very bad (see the attached <strong>Labeling Guide document</strong> for details)</li> </ul> <p> </p> <h3>Image source</h3> <p>Images are obtained from <a href="https://www.mapillary.com/">Mapillary</a>, a crowd-sourcing plattform for street-level imagery. More metadata about each image can be obtained via the <a href="https://www.mapillary.com/developer/api-documentation">Mapillary API . </a>User-generated images are shared by Mapillary under the <a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA</a> License.</p> <p>For each image, the dataset contains the <code>mapillary_image_id</code> and <code>user_name</code>. <br>You can access user information on the Mapillary website by <code>https://www.mapillary.com/app/user/<USER_NAME> </code><br>and image information by <code>https://www.mapillary.com/app/?focus=photo&pKey=<MAPILLARY_IMAGE_ID></code></p> <p>If you use the provided images, please adhere to the <a href="https://www.mapillary.com/terms">terms of use of Mapillary.</a></p> <p> </p> <h3>Instances per class</h3> <p>Total number of images: 9,122</p> <table> <tbody> <tr> <td> </td> <td><strong>excellent</strong></td> <td><strong>good</strong></td> <td><strong>intermediate</strong></td> <td><strong>bad</strong></td> <td><strong>very bad</strong></td> </tr> <tr> <td><strong>asphalt</strong></td> <td>971</td> <td>1697</td> <td>821</td> <td>246</td> <td>-</td> </tr> <tr> <td><strong>concrete</strong></td> <td>314</td> <td>350</td> <td>250</td> <td>58</td> <td>-</td> </tr> <tr> <td><strong>paving stones</strong></td> <td>385</td> <td>1063</td> <td>519</td> <td>70</td> <td>-</td> </tr> <tr> <td><strong>sett</strong></td> <td>-</td> <td>129</td> <td>694</td> <td>540</td> <td>-</td> </tr> <tr> <td><strong>unpaved</strong></td> <td>-</td> <td>-</td> <td>326</td> <td>387</td> <td>303</td> </tr> </tbody> </table> <p> </p> <p>For modeling, we recommend using a train-test split where the test data includes geospatially distinct areas, thereby ensuring the model's ability to generalize to unseen regions is tested. We propose five cities varying in population size and from different regions in Germany for testing - images are tagged accordingly.</p> <p>Number of test images (train-test split): 776</p> <h3>Inter-rater-reliablility</h3> <p>Three annotators labeled the dataset, such that each image was annotated by one person. Annotators were encouraged to consult each other for a second opinion when uncertain.<br>1,800 images were annotated by all three annotators, resulting in a <em>Krippendorff's alpha</em> of 0.96 for surface type and 0.74 for surface quality.</p> <h3>Recommended image preprocessing</h3> <p>As the focal road located in the bottom center of the street-level image is labeled, it is recommended to crop images to their lower and middle half prior using for classification tasks.</p> <p>This is an exemplary code for recommended image preprocessing in <strong>Python</strong>:</p> <pre><code>from PIL import Image<br></code><code>img = Image.open(image_path)</code><br><code>width, height = img.size</code><br><code>img_cropped = img.crop((0.25 * width, 0.5 * height, 0.75 * width, height))</code></pre> <h3><br><strong>License</strong></h3> <p><a href="https://creativecommons.org/licenses/by-sa/4.0/">CC-BY-SA</a></p> <p> </p> <h3><strong>Citation</strong></h3> <p>If you use this dataset, please cite as: </p> <p> </p> <p>Kapp, A., Hoffmann, E., Weigmann, E. <em>et al.</em> StreetSurfaceVis: a dataset of crowdsourced street-level imagery annotated by road surface type and quality. <em>Sci Data</em> <strong>12</strong>, 92 (2025). https://doi.org/10.1038/s41597-024-04295-9</p> <p> </p> <p><code>@article{kapp_streetsurfacevis_2025,<br> title = {{StreetSurfaceVis}: a dataset of crowdsourced street-level imagery annotated by road surface type and quality},<br> volume = {12},<br> issn = {2052-4463},<br> url = {https://doi.org/10.1038/s41597-024-04295-9},<br> doi = {10.1038/s41597-024-04295-9},<br> pages = {92},<br> number = {1},<br> journaltitle = {Scientific Data},<br> shortjournal = {Scientific Data},<br> author = {Kapp, Alexandra and Hoffmann, Edith and Weigmann, Esther and Mihaljević, Helena},<br> date = {2025-01-16},<br>}</code></p> <p> </p> <p>-----------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p>This is part of the SurfaceAI project at the University of Applied Sciences, HTW Berlin.</p> <p><br>- Prof. Dr. Helena Mihajlević<br>- Alexandra Kapp<br>- Edith Hoffmann<br>- Esther Weigmann</p> <p>Contact: surface-ai@htw-berlin.de</p> <p>https://surfaceai.github.io/surfaceai/</p> <p><strong>Funding</strong>: SurfaceAI is a mFund project funded by the Federal Ministry for Digital and Transportation Germany.</p> <p> </p>
An annotated high-content fluorescence microscopy dataset with EGFP-Galectin-3-stained cells and manually labelled outlines
<p>Here we present a benchmarking dataset of fluorescence microscopy images with EGFP-Galectin-3-stained cells together with annotations of their outlines. Images were randomly selected from an RNA interference screen with a modified U2OS osteosarcoma cell line, acquired on a Thermo Fischer CX7 high-content imaging system at 20x magnification. </p> <p>The dataset contains 60 images showing over 2000 labelled nuclear objects in total, which is sufficiently large to train well-performing neural networks for instance or semantic segmentation. It is pre-split into training, development and test set, each in a zip file. The dataset should be referred to as Aitslab_bioimaging2.</p> <p>For most of the images, nuclear staining and annotations have been published previously in the dataset Aitslab_bioimaging1 (https://doi.org/10.5281/zenodo.6657260). The conversion script to produce the png images from the C01 images was published together with this dataset.</p>
GitHub Profiles (users/organisations) and Repositories (research/non-research) of Potsdam Researchers and Research Organisations: An annotated dataset of with howfairis and software quality variables.
<p>This dataset accompanies the paper <em>"Software FAIRness, Documentation and Development Practices in Potsdam Researchers' GitHub Repositories"</em> It includes 3 CSV files that contain data related to github profiles of users/organisations, their repositories annotated as research/non-research repositories and followed by FAIRness and other software qualtiy variables. The data were collected using <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP">SWORDS-template-UP</a> (v1.0.0) methods (collect_users, collect_repositories, collect_variables) which is extended version of <a href="https://github.com/UtrechtUniversity/SWORDS-template">SWORS-template</a> adopted according our needs and detailed in the paper.</p> <p><strong>GitHub (research) user/organisation profiles. ( <em>github_profiles.csv )</em></strong></p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description </strong></td> </tr> <tr> <td>user_id</td> <td>GitHub username </td> </tr> <tr> <td>html_url </td> <td>URL of the GitHub profile </td> </tr> <tr> <td>type </td> <td>Type of profile (user or organization)</td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the organization </td> </tr> </tbody> </table> <p><strong>GitHub repositories <em>(github_repositories.csv)</em></strong></p> <p>This file contains the repositories scraped from the GitHub profiles of research users and organizations.</p> <table> <tbody> <tr> <td><strong>Column name </strong></td> <td><strong>Description </strong></td> </tr> <tr> <td>html_url </td> <td>URL link to the repository </td> </tr> <tr> <td>description</td> <td>GitHub project description </td> </tr> <tr> <td>project</td> <td>Specifies if the project is research or non-research</td> </tr> <tr> <td>language</td> <td>Programming language used in the project </td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the university, institution, or research organization</td> </tr> <tr> <td>research_group</td> <td>Acronym or name of the research group the repository belongs to</td> </tr> </tbody> </table> <p><strong>Research repositories filtered and annotated <em>(github_research_repositories_filtered_annotated.csv)</em></strong></p> <p>This file contains filtered and annotated information about research repositories.</p> <table> <tbody> <tr> <td><strong>Column Name </strong></td> <td><strong>Description </strong></td> <td><strong>Collection Method </strong></td> </tr> <tr> <td>html_url </td> <td>Repository URL </td> <td> </td> </tr> <tr> <td>howfairis_repository</td> <td>Indicates if the repository is public or private (True/False) </td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_license </td> <td>Indicates if the repository has a license (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_registry</td> <td>Indicates if the repository has implemented community registry (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_citation</td> <td>Indicates if the repository has a .cff file (True/False) </td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_checklist</td> <td>Indicates if the repository has implemented OpenSSF best practices badge (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>fair_score</td> <td>Score based on howfairis variables (0-5) </td> <td> </td> </tr> <tr> <td>dlr_soft_class</td> <td>Name of the university, company, research institute, or research organization</td> <td>(Manual) Annotated the repository based on <a href="https://core.ac.uk/reader/211557820">DLR software engineering guideline.</a> There are no specific definitions on metrics how to categorise them (github repositories) into application classes. Which were needed to do a comparitive analysis. </td> </tr> <tr> <td>installation_instruction</td> <td>Presence of installation instruction (True/False) </td> <td>(Manual) Checked the presense of Installation Instruction in the readme or in the project wiki pages. </td> </tr> <tr> <td>project_information </td> <td>Presence of basic project information in README (True/False) </td> <td>(Manual) Checked if the readme have basic information about the project. </td> </tr> <tr> <td>usage_guide</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td>(Manual) Checked the presense of Usage Guide in the readme or in the project wiki pages. For command line tools checked if they have help command which guides how to use the tool. </td> </tr> <tr> <td>test_folder</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td> <p>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/docs/collect_variables/scripts/soft_dev_pract/test_folder.py">test_folder.py</a>) Checks the folder names test/tests in the root directory of the repository.</p> </td> </tr> <tr> <td>requirements_explicit </td> <td>Explicit requirements for Python, R, C++ repositories (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/requirement_explicit.py">requirement_explicit.py</a>) Checks the files (requirements.txt, DESCRIPTION, CMakeLists.txt) in the root directory. </td> </tr> <tr> <td>continuous_integration</td> <td>Indicates if the repository uses continuous integration (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>ci_tool </td> <td>Name of the continuous integration tool used</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>add_lint_rule </td> <td>Indicates if additional linting rules are present (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the <br>.github/workflows directory to detect the presence of (linters) Python, R, and C++.</td> </tr> <tr> <td>add_test_rule</td> <td>Indicates if additional testing rules are present (True/False) </td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the <br>.github/workflows directory to detect the presence of (testing libraries) Python, R, and C++.</td> </tr> <tr> <td>comment_at_start</td> <td>Indicates the level of comments at the start of the program (most, more, some, less)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/comment_at_start.py">comment_at_start.py</a>) Checks the presence of brief comments at the start at source code files in GitHub repositories.</td> </tr> <tr> <td>language </td> <td>Programming language used in the repository </td> <td> </td> </tr> <tr> <td>type </td> <td>Specifies if the profile is a user or organization </td> <td>Github organisation or user profiles.</td> </tr> <tr> <td>organisation </td> <td>Name of the university, company, research institute, or research organization</td> <td>Oraganisation name (from where the user was found)</td> </tr> <tr> <td>research_group</td> <td>Name or acronym of the research group </td> <td> </td> </tr> </tbody> </table> <p> </p> <p>Data for publication - https://github.com/Software-Engineering-Group-UP/potsdam-research-repos</p>
Three Annotated Anomaly Detection Datasets for Line-Scan Algorithms
<h1>Summary</h1> <p>This dataset contains two hyperspectral and one multispectral anomaly detection images, and their corresponding binary pixel masks. They were initially used for real-time anomaly detection in line-scanning, but they can be used for any anomaly detection task.</p> <p>They are in .npy file format (will add tiff or geotiff variants in the future), with the image datasets being in the order of (height, width, channels). The SNP dataset was collected using sentinelhub, and the Synthetic dataset was collected from AVIRIS. The Python code used to analyse these datasets can be found at: https://github.com/WiseGamgee/HyperAD</p> <h1>How to Get Started</h1> <p>All that is needed to load these datasets is Python (preferably 3.8+) and the NumPy package. Example code for loading the Beach Dataset if you put it in a folder called "data" with the python script is:</p> <pre><code>import numpy as np # Load image file hsi_array = np.load("data/beach_hsi.npy") n_pixels, n_lines, n_bands = hsi_array.shape print(f"This dataset has {n_pixels} pixels, {n_lines} lines, and {n_bands}.") # Load image mask mask_array = np.load("data/beach_mask.npy") m_pixels, m_lines = mask_array.shape print(f"The corresponding anomaly mask is {m_pixels} pixels by {m_lines} lines.")</code></pre> <h1>Citing the Datasets</h1> <p>If you use any of these datasets, please cite the following paper:</p> <pre><code>@article{garske2024erx,</code><br><code> title={ERX - a Fast Real-Time Anomaly Detection Algorithm for Hyperspectral Line-Scanning},</code><br><code> author={Garske, Samuel and Evans, Bradley and Artlett, Christopher and Wong, KC},</code><br><code> journal={arXiv preprint arXiv:2408.14947},</code><br><code> year={2024},</code><br><code>}</code></pre> <div> <pre>If you use the beach dataset please cite the following paper as well (original source):</pre> </div> <pre><code>@article{mao2022openhsi, title={OpenHSI: A complete open-source hyperspectral imaging solution for everyone}, author={Mao, Yiwei and Betters, Christopher H and Evans, Bradley and Artlett, Christopher P and Leon-Saval, Sergio G and Garske, Samuel and Cairns, Iver H and Cocks, Terry and Winter, Robert and Dell, Timothy}, journal={Remote Sensing}, volume={14}, number={9}, pages={2244}, year={2022}, publisher={MDPI} }</code></pre>
French Entity-Linking dataset between annotated tweets collected during major crises in France and French Wikipedia corpus
<p>Most of the available datasets are not particularly adapted to our target application: geolocate natural disasters from social networks. First, social media posts are largely underrepresented in these datasets, and the only Twitter dataset lacks Entity-Linking annotations. Second, none of the datasets focuses on a crisis or natural disaster event.</p> <p>To mitigate these issues, we extracted a collection of French tweets written during earthquakes and major floods that have occurred in France in recent years. We set up Label-Studio in order to annotate these tweets. A total of 4617 tweets were annotated, including 1678 tweets posted during earthquakes and 2939 during floods. For each annotated tweet, mentions were annotated using the set of labels described earlier in the paper as well as, when possible, the target Wikipedia title.</p> <p>Named “RéSoCIO” in reference to the research project in which it was carried out, the dataset resulting from this work contains a total of 12 828 annotated mentions and 1 513 distinct Wikipedia entities. 85% of mentions were associated with a Wikipedia page and 94 % if we ignore the RISKNAT and DAMAGES labels, which are often difficult to map to an existing entity.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entities</strong></td> </tr> <tr> <td>PERSON</td> <td>315</td> <td>263</td> <td>136</td> </tr> <tr> <td>ORG</td> <td>863</td> <td>790</td> <td>281</td> </tr> <tr> <td>GEOLOC</td> <td>4375</td> <td>4234</td> <td>701</td> </tr> <tr> <td>TRANSPORT</td> <td>250</td> <td>203</td> <td>101</td> </tr> <tr> <td>EVENT</td> <td>35</td> <td>21</td> <td>16</td> </tr> <tr> <td>FACILITY</td> <td>129</td> <td>94</td> <td>49</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>128</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>223</td> <td>200</td> <td>46</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>12828</strong></td> <td><strong>1322</strong></td> <td><strong>1513</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the Twitter dataset. #Mentions shows the total number of mentions per label, #Linked the number of mentions linked to an entity and #Entities the number of distinct entities per label present in the dataset.</p> <table> <tbody> <tr> <td><strong>Labels</strong></td> <td><strong>#Mentions</strong></td> <td><strong>#Linked</strong></td> <td><strong>#Entitie</strong>s</td> </tr> <tr> <td>PERSON</td> <td>1100102</td> <td>1098406</td> <td>557697</td> </tr> <tr> <td>ORG</td> <td>750925</td> <td>749504</td> <td>130394</td> </tr> <tr> <td>GEOLOC</td> <td>2729702</td> <td>2728296</td> <td>215924</td> </tr> <tr> <td>TRANSPORT</td> <td>161539</td> <td>160487</td> <td>53405</td> </tr> <tr> <td>EVENT</td> <td>798433</td> <td>798251</td> <td>86471</td> </tr> <tr> <td>FACILITY</td> <td>258835</td> <td>258513</td> <td>109867</td> </tr> <tr> <td>RISKNAT</td> <td>5502</td> <td>4994</td> <td>127</td> </tr> <tr> <td>DAMAGES</td> <td>1136</td> <td>121</td> <td>56</td> </tr> <tr> <td>OTHER</td> <td>4340621</td> <td>4339658</td> <td>682458</td> </tr> <tr> <td><strong>Total</strong></td> <td><strong>10146795</strong></td> <td><strong>10138230</strong></td> <td><strong>1836399</strong></td> </tr> </tbody> </table> <p>Overview of the mentions annotated in the full dataset. #Mentions shows the total number of mentions per label, #Linked the number of mentions linked to an entity and #Entities the number of distinct entities per label present in the dataset.</p>
MiRoR11 - P2 - Annotated dataset for spin-related types of statements (statements of similarity and within-group comparisons)
<p>180 abstracts / 2401 sentences annotated for 2 types of spin-related statements: statements of similarity and within-group comparisons.</p>
Annotated Dataset for Uncertainty Mining : Gold Standard
<p> </p> <h1>Description of the dataset</h1> <p>In order to study the expression of uncertainty in scientific articles, we have put together an interdisciplinary corpus of journals in the fields of Science, Technology and Medicine (STM) and the Humanities and Social Sciences (SHS). The selection of journals in our corpus is based on the Scimago Journal and Country Rank (SJR) classification, which is based on Scopus, the largest academic database available online. We have selected journals covering various disciplines, such as medicine, biochemistry, genetics and molecular biology, computer science, social sciences, environmental sciences, psychology, arts and humanities. For each discipline, we selected the five highest-ranked journals. In addition, we have included the journals PLoS ONE and Nature, both of which are interdisciplinary and highly ranked.</p> <p>Based on the corpus of articles from different disciplines described above, we created a set of annotated sentences as follows:</p> <ul> <li>593 were pre-selected automatically, by studying the occurrences of the lists of uncertainty indices proposed by Bongelli et al. (2019), Chen et al. (2018) and Hyland (1996).</li> <li>The remaining sentences were extracted from a subset of articles, consisting of two randomly selected articles per journal. These articles were examined by two human annotators to identify sentences containing uncertainty and to annotate them.</li> <li>600 sentences not expressing scientific uncertainty were manually identified and reviewed by two annotators<br><br></li> </ul> <p>The sentences were annotated by two independent annotators following the annotation guide proposed by Ningrum and Atanassova (2024). The annotators were trained on the basis of an annotation guide and previously annotated sentences in order to guarantee the consistency of the annotations. <br>Each sentence was annotated as expressing or not expressing uncertainty (<strong>Uncertainty</strong> and <strong>No Uncertainty)</strong>.<br>Sentences expressing uncertainty were then annotated along five dimensions: Reference , Nature, Context , Timeline and Expression. <br>The annotators reached an average agreement score of 0.414 according to Cohen's Kappa test, which shows the difficulty of the task of annotating scientific uncertainty.<br>Finally, conflicting annotations were resolved by a third independent annotator.</p> <p><br>Our final corpus thus consists of a total of 1 840 sentences from 496 articles in 21 English-language journals from 8 different disciplines.<br>The columns of the table are as follows:</p> <ol> <li><strong>journal</strong>: name of the journal from where the article originates</li> <li><strong>article_title</strong>: title of the article from where the sentence is extracted</li> <li><strong>publication_year</strong>: year of publication of the article</li> <li><strong>sentence_text</strong>: text of the sentence expressing or not expressing uncertainty</li> <li><strong>uncertainty</strong>: 1 if the sentence expresses uncertainty and 0 otherwise;</li> <li><strong>ref, nature, context, timeline, expression</strong>: annotations of the type of uncertainty according to the annotation framework proposed by Ningrum and Atanassova (2023). The annotation of each dimension in this dataset are in numeric format rather than textual. The mapping betwen textual and numeric labels is presented in the Table below.</li> </ol> <table> <tbody> <tr> <td>Dimension</td> <td>1</td> <td>2</td> <td>3</td> <td>4</td> <td>5</td> </tr> <tr> <td>Reference</td> <td>Author</td> <td>Former</td> <td>Both</td> <td> </td> <td> </td> </tr> <tr> <td>Nature</td> <td>Epistemic</td> <td>Aleatory</td> <td>Both</td> <td> </td> <td> </td> </tr> <tr> <td>Context</td> <td>Background</td> <td>Methods</td> <td>Res&Disc</td> <td>Conclusion</td> <td>Others</td> </tr> <tr> <td>Timeline</td> <td>Past</td> <td>Present</td> <td>Future</td> <td> </td> <td> </td> </tr> <tr> <td>Expression</td> <td>Quantified</td> <td>Unquantified</td> <td> </td> <td> </td> <td> </td> </tr> </tbody> </table> <p><br>This gold standard has been produced as part of the <a href="https://project-inscim.github.io/">ANR InSciM (Modelling Uncertainty in Science) project.</a> </p> <h1>References</h1> <p><br>Bongelli, R., Riccioni, I., Burro, R., & Zuczkowski, A. (2019). Writers’ uncertainty in scientific and popular biomedical articles. A comparative analysis of the British Medical Journal and Discover Magazine [Publisher: Public Library of Science]. PLoS ONE, 14 (9). <a href="https://doi.org/10.1371/journal.pone.0221933">https://doi.org/10.1371/journal.pone.0221933</a></p> <p>Chen, C., Song, M., & Heo, G. E. (2018). A scalable and adaptive method for finding semantically equivalent cue words of uncertainty. Journal of Informetrics, 12 (1), 158–180. <a href="https://doi.org/10.1016/j.joi.2017.12.004">https://doi.org/10.1016/j.joi.2017.12.004</a></p> <p><br>Hyland, K. E. (1996). Talking to the academy forms of hedging in science research articles [Publisher: SAGE Publications Inc.]. Written Communication, 13 (2), 251–281. <a href="https://doi.org/10.1177/0741088396013002004">https://doi.org/10.1177/0741088396013002004</a></p> <p>Ningrum, P. K., & Atanassova, I. (2023). Scientific Uncertainty: An Annotation Framework and Corpus Study in Different Disciplines. 19th International Conference of the International Society for Scientometrics and Informetrics (ISSI 2023). <a href="https://doi.org/10.5281/zenodo.8306035">https://doi.org/10.5281/zenodo.8306035</a></p> <p>Ningrum, P. K., & Atanassova, I. (2024). Annotation of scientific uncertainty using linguistic patterns. Scientometrics. <a href="https://doi.org/10.1007/s11192-024-05009-z">https://doi.org/10.1007/s11192-024-05009-z</a></p>
AckSent: Human Annotated Dataset of Support and Sentiments in Dissertation Acknowledgments
<p>This data is supplementary to the paper:</p> <blockquote> <p><em>Manika Lamba, You Peng, Sophie Nikolov, and J. Stephen Downie. 2024. <strong>AckSent: Human Annotated Dataset of Support and Sentiments in Dissertation Acknowledgments</strong>. In The 2024 ACM/IEEE Joint Conference on Digital Libraries (JCDL ’24), December 2024, Hong Kong, China. ACM, New York, NY, USA, 5 pages. https://doi.org/10.1145/3677389.3702594</em></p> </blockquote>
Tree Annotation Vocabulary (TAV) - Knowledge Graph and Annotated Dataset
<p>This dataset contains all the files used in developing the Tree-KG, the knowledge graph to capture the tree annotations in the works of Vladimir Nabokov. </p> <p>In the Annotated Dataset folder, 6 spreadsheets in excel (.xlsx) format are provided. They are numbered. Note that annotated data are all in English as the consulted works are the English translations of the literary works of Nabokov.</p> <p>(1) contains the tree annotations from the novels originally written in Russian by Vladimir Nabokov.</p> <p>(2) contains the tree annotations from the novels originally written in English by Vladimir Nabokov.</p> <p>(3) contains the tree annotations from the short stories originally written in Russian and English by Vladimir Nabokov.</p> <p>(4) is the knowledge base (KB) developed to link the annotated trees to Wikidata and DBPedia.</p> <p>(5) is the benchmarking results of some entity recognition tools. It includes the relevant passages from Nabokov's novels that were used in the experiments as well as the prompts used in getting the results.</p> <p>(6) represents the complete bibliographic details of the works of Vladimir Nabokov (https://thenabokovian.org/abbreviations).</p> <p>In the Ontology Versions folder, four ontology (TAV) files in turtle (.ttl) format are provided. They are all numbered and dated to represent their different versions. Some sample SPARQL queries are provided in a .txt file. The KG was developed on Protégé. </p> <p>(1) contains the essential schema for the TAV vocabulary.</p> <p>(2) contains the schema for TAV vocabulary with links to external vocabularies (Schema.Org; Open Annotation, etc.). </p> <p>(3) contains the Tree-KG in so far it reflects data from three novels (Mary; King, Queen, Knave; Glory).</p> <p>(4) contains the entire Tree-KG based on all the works mentioned in the excel sheets (20 books).</p> <p>(5) contains some sample SPARQL queries (.txt) file.</p>
Breast Micro-Calcifications Dataset with Precisely Annotated Sequential Mammograms
<p><strong>Dataset Version 3 Update</strong></p> <p><strong>The ground truth images (.jpg) match the dimensions of the corresponding original images (.dcm), ensuring consistency across the dataset.</strong><br><br></p> <p><strong>Breast Micro-Calcifications Dataset with Precisely Annotated Sequential Mammograms</strong></p> <p><strong>Citing the Dataset</strong></p> <p>The dataset is released under a Creative Commons Attribution license, so please cite the dataset if it is used in your work in any form. Published academic papers should use the academic paper citation for our paper. Personal works, such as projects or blog posts, should provide a URL to this Zenodo page, though a reference to our paper would also be appreciated.</p> <p><em>Academic paper citation</em></p> <p>Loizidou, K., Skouroumouni, G., Pitris, C. <em>et al.</em> Digital subtraction of temporally sequential mammograms for improved detection and classification of microcalcifications. <em>Eur Radiol Exp</em> <strong>5, </strong>40 (2021). https://doi.org/10.1186/s41747-021-00238-w</p> <p><em>Personal use citation</em></p> <p>Include a link to this Zenodo page - 10.5281/zenodo.14859694</p> <p><strong>ACKNOWLEDGMENT</strong></p> <p>This research is funded by the European Union’s Horizon 2020 research and innovation program under grant agreement No. 739551 (KIOS CoE) and from the Republic of Cyprus through the Directorate General for European Programs, Coordination and Development.</p> <p><strong>Contact Information</strong></p> <p>If you would like further information about the dataset, or if you experience any issues downloading files, please contact us at cloizi01@ucy.ac.cy.</p> <p><strong>General Information</strong></p> <p>This dataset consists of 100 pairs of mammograms, from two temporally sequential rounds. Specifically, this dataset includes the prior and recent mammograms of CC and MLO view of each patient. This is a complete dataset for the detection and BI-RADS classification of breast micro-calcifications, using digital mammograms. It contains normal (BI-RADS 1), benign (BI-RADS 2), and suspicious (BI-RADS 4-5) cases, and for each mammogram, an image with precise annotation of each individual micro-calcification, by two expert radiologists, is provided. In 32 suspicious cases, the biopsy results are also available.</p> <p><strong>More details are available in the README.txt</strong></p>
Manually Annotated Drone Imagery (RGB) Dataset for automatic coastline delineation of Southern Baltic Sea, Poland with polyline annotations (0.1.1)
<p><strong>Overview:</strong></p> <p>The Manually Annotated Drone Imagery Dataset (MADRID) consists of hand annotated high resolution RGB images taken in two different types of coasts in Poland, Miedzyzdroje - cliff coast and in Mrzezyno - dune coast in 2022-2023. All images were converted into a uniform format of 1440x2560 pixels, polyline annotated and set into file structure format suited for semantic segmentation tasks (See "Usage" notes below for more details).</p> <p>The raw images of our dataset were captured Zenmuse L1 Sensor (RGB) mounted on a DJI Matrice 300 RTK Drone. Total of 4895 images were captured, however the dataset contains 3876 images with each image annotated with coastline. The dataset only include images with coastlines that are visually identifiable with the human eye. For the annotations of the images, CVAT v2.13 open-source software was utilized.</p> <p><strong>Usage:</strong></p> <p>The compressed RAR file contains two folders train and test. Each folder contains the file that represents the date at which the image was captured in the format of (year, month, day), number of the image and the name of the drone utilized to capture the image. For example, DJI_20220111140051_0051_Zenmuse-L1-mission and DJI_20220111140105_0053_Zenmuse-L1-mission. Additionally, the test folder contains annotations (one per image) which are extracted from the original XML annotation file provided in the CVAT 1.1 image format.</p> <p>Archives were compressed using RAR compression. They can be decompressed in a terminal by opening and extracting Madrid_v0.1_Data.zip.</p> <p>The subset of the data with the name Madrid_subset_data.zip has been added which contains a small portion of train and test images for purpose of inspecting the dataset without downloading the entire dataset.</p> <p>The training images for both training data and testing data are structured as follows.</p> <pre><code>Train/ └── images/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.JPG └── DJI_20220111140105_0053_Zenmuse-L1-mission.JPG └── ...<br>└── masks/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.PNG └── DJI_20220111140105_0053_Zenmuse-L1-mission.PNG └── ...<br><br>Test/ └── images/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.JPG └── DJI_20220111140105_0053_Zenmuse-L1-mission.JPG └── ...<br>└── masks/ └── DJI_20220111140051_0051_Zenmuse-L1-mission.PNG └── DJI_20220111140105_0053_Zenmuse-L1-mission.PNG └── ...</code></pre> <p> </p>
AUTH-OpenDR Mixed Image Annotated Dataset for Human-centric Perception Tasks
<p>The dataset was generated through a mixed (real and synthetic) image data generation method which utilizes real background images and DL-generated human models. It contains 50000 real images depicting urban scenes, populated by synthetic human models in various positions and poses and is suitable for training/evaluating (a) pose estimation, (b) person detection, (c) identity recognition methods. Annotations for 2D bounding boxes of the depicted humans, their IDs and 2D keypoints etc are provided. The 133 3D human models, required by the method, were generated using the Pixel-aligned Implicit Function (PIFu) and full-body images of people from the Clothing Co-Parsing (CCP) dataset. As background images, a subset of the Cityscapes dataset was used. The Cityscapes license prohibits the distribution of any modified versions of itself. Thus, we provide code that can re-generate the exact same dataset, given that the Cityscapes dataset is downloaded by the website of its authors.</p> <p>Code and instructions for re-generating the dataset are provided <a href="https://github.com/opendr-eu/opendr/tree/master/projects/python/simulation/human_dataset_generation">here</a>.</p> <p>The dataset was developed by Aristotle University of Thessaloniki (AUTH) within the H2020 OpenDR Project.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.