Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4 results for “crowdsourced annotations”

Learn how ShareScore rates datasets ↗
zenodo40/100

A crowdsourced dataset of aerial images with annotated solar photovoltaic arrays and installation metadata

<p><strong>Summary</strong></p> <p>Photovoltaic (PV) energy generation plays a crucial role in the energy transition. Small-scale, residential PV installations are deployed at an unprecedented pace, and their safe integration into the grid necessitates up-to-date, high-quality information. Overhead imagery is increasingly used to improve the knowledge of residential PV installations with machine learning models capable of automatically mapping these installations. However, these models cannot be reliably transferred from one region or imagery source to another without incurring a decrease in accuracy. To address this issue, known as distribution shift, and foster the development of PV array mapping pipelines, we propose a dataset containing aerial images, segmentation masks, and installation metadata. We provide installation metadata for more than 28000 installations. We provide ground truth segmentation masks for 13000 installations, including 7000 with annotations for two different image providers. Finally, we provide installation metadata that matches the annotation for more than 8000 installations. Dataset applications include end-to-end PV registry construction, robust PV installations mapping, and analysis of crowdsourced datasets.</p> <p>This dataset contains the complete records&nbsp;associated with the article &quot;A crowdsourced dataset of aerial images of solar panels, their segmentation masks, and characteristics&quot;, published in Scientific data. The article is accessible here :&nbsp;<a href="https://www.nature.com/articles/s41597-023-01951-4">https://www.nature.com/articles/s41597-023-01951-4</a> These complete records consist of:</p> <ol> <li>The complete training dataset containing RGB overhead imagery, segmentation masks and metadata of PV installations (folder <strong>bdappv</strong>),</li> <li>The raw crowdsourcing data, and the postprocessed data for replication and validation (folder <strong>data</strong>).</li> </ol> <p><strong>Data records</strong></p> <p>Folders are organized as follows:</p> <ul> <li><strong>bdappv/</strong> Root data folder <ul> <li><strong>google / ign:</strong>&nbsp; One folder for each campaign <ul> <li><strong>img/</strong>: Folder containing all the images presented to the users. This folder contains 28807 images for Google and 17325 images for IGN.</li> <li><strong>mask/</strong>: Folder containing all segmentations masks generated from the polygon annotations of the users. This folder contains 13303 masks for Google and&nbsp;7686 masks for IGN.</li> </ul> </li> <li><em>metadata.csv</em> The <code>.csv</code>&nbsp; file with the installations&#39; metadata.</li> </ul> </li> </ul> <p>&nbsp;</p> <ul> <li><strong>data/ </strong>Root data folder <ul> <li><strong>raw/</strong> Folder containing the raw crowdsourcing data and raw metadata; <ul> <li><em>input-google.json</em>: <code>.json </code>input data data containing all information on images and raw annotators&rsquo; contributions for both phases (clicks and polygons) during the first annotation campaign;</li> <li><em>input-ign.json</em>:<em> </em><code>.json </code>input data containing all information on images and raw annotators&rsquo; contributions for both phases (clicks and polygons) during the second annotation campaign;</li> <li><em>raw-metadata.json</em>: <code>.json </code>output containing the PV systems&rsquo; metadata extracted from the BDPV database before filtering. It can be used to replicate the association between the installations and the segmentation masks, as done in the notebook metadata.</li> </ul> </li> <li><strong>replication/</strong> Folder containing the compiled data used to generate the segmentation masks; <ul> <li><strong>campaign-google/campaign-ign</strong>: One folder for each campaign <ul> <li><em>click-analysis.json</em>: <code>.json </code>output on the click analysis, compiling raw input into a few best-guess locations for the PV arrays. This dataset enables the replication of our annotations,</li> <li><em>polygon-analysis.json</em>: <code>.json </code>output of polygon analysis, compiling raw input into a best-guess polygon for the PV arrays.</li> </ul> </li> </ul> </li> <li><strong>validation/</strong> Folder containing the compiled data used for technical validation. <ul> <li><strong>campaign-google/campaign-ign</strong>: One folder for each campaign <ul> <li><em>click-analysis-thres=1.0.json</em>: <code>.json </code>output of the click analysis with a lowered threshold to analyze the effect of the threshold on image classification, as done in the notebook annotation;</li> <li><em>polygon-analysis-thres=1.0.json</em>: <code>.json </code>output of polygon analysis, with a lowered threshold to analyze the effect of the threshold on polygon annotation, as done in the notebook annotations.</li> </ul> </li> <li><em>metadata.csv</em>: the <code>.csv </code>file of filtered installations&#39; metadata.</li> </ul> </li> </ul> </li> </ul> <p><strong>License</strong></p> <p>We extracted the thumbnails contained in the <strong>google/img/</strong> folder using Google Earth Engine API and we generated the thumbnails contained in the <strong>ign/img</strong><strong>/</strong> folder from high resolution tiles downloaded from the online IGN portal accessible here: <a href="https://geoservices.ign.fr/bdortho">https://geoservices.ign.fr/bdortho</a>. Images provided by Google are subjet to Google&#39;s terms and conditions. Images provided by the IGN are subject to an open license 2.0.</p> <p>Access the terms and conditions of Google images at this URL: <a href="https://www.google.com/intl/en/help/legalnotices_maps/">https://www.google.com/intl/en/help/legalnotices_maps/</a></p> <p>Access the terms and conditions of IGN images at this URL: <a href="https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf">https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf</a></p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

New crowdsourced annotations for the GiantSteps Tempo dataset

<p>Raw beat and derived tempo annotations&nbsp;from&nbsp;<a href="https://doi.org/10.5281/zenodo.1492437">A Crowdsourced Experiment for Tempo Estimation of Electronic Dance Music</a>.</p>

opencc-by-4.0Sep 2018View details →
zenodo36/100

The Belfort dataset: Handwritten Text Recognition from Crowdsourced Annotations

<p>This dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946. Documents include deliberations, lists of councillors, convocations, and agendas.</p> <p>The dataset includes 24,105 text-line images that were automatically detected from pages. Up to 4 transcriptions are available for each line image: two from humans, and two from automatic models.</p> <p>We would like to thank the <em>Archives municipales de la ville de Belfort, France</em> for giving us access to these documents.</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

Crowdsourced LibriTTS Speech Prominence Annotations

<p>Dataset corresponding to the ICASSP 2024 paper "Crowdsourced and Automatic Speech Prominence Estimation" <a href="https://arxiv.org/abs/2310.08464">[link]</a></p> <p>This dataset is useful for training machine learning models to perform automatic emphasis annotaiton, as well as downstream tasks such as&nbsp;emphasis-controlled TTS, emotion recognition, and text summarization. The dataset is described in Section 3 (Emphasis Annotation Dataset). The contents of this section are copied below for convenience.</p> <p>We used our crowdsourced annotation system to perform human annotation on one eighth of the train-clean-100 partition of the LibriTTS [1] dataset. Specifically, participants annotated 3,626 utterances with a total length of 6.42 hours and 69,809 words from 18 speakers (9 male and 9 female). We collected at least one annotation of all 3,626 utterances, at least two annotations of 2,259 of those utterances, at least four annotations of 974 utterances, and at least eight annotations of 453 utterances. We did this in order to explore (in Section 6) whether it is more cost-effective to train a system on multiple annotations of fewer utterances or fewer annotations of more utterances. We paid 298 annotators to annotate batches of 20 utterances, where each batch takes approximately 15 minutes. We paid $3.34 for each completed batch (estimated $13.35 per hour). Annotators each annotated between one and six batches. We recruited on MTurk US residents with an approval rating of at least 99 and at least 1000 approved tasks. Today, microlabor platforms like MTurk are plagued by automated task-completion software agents (bots) that randomly fill out surveys. We filtered out bots by excluding annotations from an additional 107 annotators that marked more than 2/3 of words as emphasized in eight or more utterances of the 20 utterances in a batch. Annotators who fail the bot filter are blocked from performing further annotation. We also recorded participants' native country and language, but note these may be unreliable as many MTurk workers use VPNs to subvert IP region filters on MTurk [2].</p> <p>The average Cohen Kappa score for annotators with at least one overlapping utterance is 0.226 (i.e., ``Fair'' agreement)---but not all annotators annotate the same utterances, and this overemphasizes pairs of annotators with low overlap. Therefore, we use a one-parameter logistic model (i.e., a Rasch model) computed via py-irt [3], which predicts heldout annotations from scores of overlapping annotators with 77.7% accuracy (50% is random).</p> <p>The structure of this dataset is a single JSON file of word-aligned emphasis annotations. The JSON references file stems of the LibriTTS dataset, which can be found <a href="https://www.openslr.org/60/">here</a>. All code used in the creation of the dataset can be found <a href="https://github.com/interactiveaudiolab/emphases">here</a>. The format of the JSON file is as follows.</p> <p>&nbsp;</p> <pre><code>{ &lt;anonymized_participant_id_0&gt;: { "annotations": [ { "score": [ &lt;word_0_prominence&gt;,<br> &lt;word_1_prominence&gt;, &nbsp; &nbsp; &nbsp;... ], "stem": &lt;libritts_file_stem&gt;, "words": [ [ &lt;word_0&gt;, &lt;word_0_start_time&gt;, &lt;word_0_end_time&gt; ], [ &lt;word_1&gt;, &lt;word_1_start_time&gt;, &lt;word_1_end_time&gt; ],<br> ... &nbsp; &nbsp; ] },<br> ... ], "country": &lt;participant_0_country&gt;, "language": &lt;participant_0_language&gt; }, ... }</code></pre> <p><br>[1] Zen et al., &ldquo;LibriTTS: A corpus derived from LibriSpeech for text-to-speech,&rdquo; in Interspeech, 2019.<br>[2] Moss et al., &ldquo;Bots or inattentive humans? Identifying sources of low-quality data in online platforms,&rdquo; PsyArXiv preprint PsyArXiv:wr8ds, 2021.<br>[3] John Patrick Lalor and Pedro Rodriguez, &ldquo;py-irt: A scalable item response theory library for Python,&rdquo; INFORMS Journal on Computing, 2023.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record