Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
125
datasets available to search
ShareScore release 0.9.0
Dataset results
125 results for “Crowdsourcing”
Description of the Features of the Open Research Knowledge Graph as a Crowdsourcing Platform based on the 4 Pillars of Crowdsourcing
<p>This dataset provides the description of the features of the <a href="https://www.orkg.org/orkg/">Open Research Knowledge Graph</a> (ORKG) based on the 4 pillars of crowdsourcing according to the reference model for crowdsourcing by Hosseini et al. [1]. This overview represents the features of the current implementation status of ORKG as a crowdsourcing platform.</p> <p>[1] M. Hosseini, K. Phalp, J. Taylor, and R. Ali, "<a href="https://ieeexplore.ieee.org/abstract/document/6861072?casa_token=zWTHNBHeH6kAAAAA:1n3EJjajqgSkSk154g4DlNFAmJs_7o3KY7LnobvP_W7AUIr-FvM5OyGP1FaRn68zUX-2oYHh7A">The Four Pillars of Crowdsourcing: A Reference Model</a>", in 2014 IEEE 8th International Conference on Research Challenges in Information Science (RCIS). IEEE, 2014, pp. 1–12.</p>
A crowdsourced dataset of aerial images with annotated solar photovoltaic arrays and installation metadata
<p><strong>Summary</strong></p> <p>Photovoltaic (PV) energy generation plays a crucial role in the energy transition. Small-scale, residential PV installations are deployed at an unprecedented pace, and their safe integration into the grid necessitates up-to-date, high-quality information. Overhead imagery is increasingly used to improve the knowledge of residential PV installations with machine learning models capable of automatically mapping these installations. However, these models cannot be reliably transferred from one region or imagery source to another without incurring a decrease in accuracy. To address this issue, known as distribution shift, and foster the development of PV array mapping pipelines, we propose a dataset containing aerial images, segmentation masks, and installation metadata. We provide installation metadata for more than 28000 installations. We provide ground truth segmentation masks for 13000 installations, including 7000 with annotations for two different image providers. Finally, we provide installation metadata that matches the annotation for more than 8000 installations. Dataset applications include end-to-end PV registry construction, robust PV installations mapping, and analysis of crowdsourced datasets.</p> <p>This dataset contains the complete records associated with the article "A crowdsourced dataset of aerial images of solar panels, their segmentation masks, and characteristics", published in Scientific data. The article is accessible here : <a href="https://www.nature.com/articles/s41597-023-01951-4">https://www.nature.com/articles/s41597-023-01951-4</a> These complete records consist of:</p> <ol> <li>The complete training dataset containing RGB overhead imagery, segmentation masks and metadata of PV installations (folder <strong>bdappv</strong>),</li> <li>The raw crowdsourcing data, and the postprocessed data for replication and validation (folder <strong>data</strong>).</li> </ol> <p><strong>Data records</strong></p> <p>Folders are organized as follows:</p> <ul> <li><strong>bdappv/</strong> Root data folder <ul> <li><strong>google / ign:</strong> One folder for each campaign <ul> <li><strong>img/</strong>: Folder containing all the images presented to the users. This folder contains 28807 images for Google and 17325 images for IGN.</li> <li><strong>mask/</strong>: Folder containing all segmentations masks generated from the polygon annotations of the users. This folder contains 13303 masks for Google and 7686 masks for IGN.</li> </ul> </li> <li><em>metadata.csv</em> The <code>.csv</code> file with the installations' metadata.</li> </ul> </li> </ul> <p> </p> <ul> <li><strong>data/ </strong>Root data folder <ul> <li><strong>raw/</strong> Folder containing the raw crowdsourcing data and raw metadata; <ul> <li><em>input-google.json</em>: <code>.json </code>input data data containing all information on images and raw annotators’ contributions for both phases (clicks and polygons) during the first annotation campaign;</li> <li><em>input-ign.json</em>:<em> </em><code>.json </code>input data containing all information on images and raw annotators’ contributions for both phases (clicks and polygons) during the second annotation campaign;</li> <li><em>raw-metadata.json</em>: <code>.json </code>output containing the PV systems’ metadata extracted from the BDPV database before filtering. It can be used to replicate the association between the installations and the segmentation masks, as done in the notebook metadata.</li> </ul> </li> <li><strong>replication/</strong> Folder containing the compiled data used to generate the segmentation masks; <ul> <li><strong>campaign-google/campaign-ign</strong>: One folder for each campaign <ul> <li><em>click-analysis.json</em>: <code>.json </code>output on the click analysis, compiling raw input into a few best-guess locations for the PV arrays. This dataset enables the replication of our annotations,</li> <li><em>polygon-analysis.json</em>: <code>.json </code>output of polygon analysis, compiling raw input into a best-guess polygon for the PV arrays.</li> </ul> </li> </ul> </li> <li><strong>validation/</strong> Folder containing the compiled data used for technical validation. <ul> <li><strong>campaign-google/campaign-ign</strong>: One folder for each campaign <ul> <li><em>click-analysis-thres=1.0.json</em>: <code>.json </code>output of the click analysis with a lowered threshold to analyze the effect of the threshold on image classification, as done in the notebook annotation;</li> <li><em>polygon-analysis-thres=1.0.json</em>: <code>.json </code>output of polygon analysis, with a lowered threshold to analyze the effect of the threshold on polygon annotation, as done in the notebook annotations.</li> </ul> </li> <li><em>metadata.csv</em>: the <code>.csv </code>file of filtered installations' metadata.</li> </ul> </li> </ul> </li> </ul> <p><strong>License</strong></p> <p>We extracted the thumbnails contained in the <strong>google/img/</strong> folder using Google Earth Engine API and we generated the thumbnails contained in the <strong>ign/img</strong><strong>/</strong> folder from high resolution tiles downloaded from the online IGN portal accessible here: <a href="https://geoservices.ign.fr/bdortho">https://geoservices.ign.fr/bdortho</a>. Images provided by Google are subjet to Google's terms and conditions. Images provided by the IGN are subject to an open license 2.0.</p> <p>Access the terms and conditions of Google images at this URL: <a href="https://www.google.com/intl/en/help/legalnotices_maps/">https://www.google.com/intl/en/help/legalnotices_maps/</a></p> <p>Access the terms and conditions of IGN images at this URL: <a href="https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf">https://www.etalab.gouv.fr/wp-content/uploads/2018/11/open-licence.pdf</a></p>
Crowdsourced COVID-19 Cases and Outbreaks across Canadian Schools 2020-21: COVID Schools Canada
<p>This archive contains the final data freeze for COVID Schools Canada, and the software used to compile, clean, and plot the data. </p> <p>The <strong>Canada COVID-19 School Tracker</strong> was a 100% volunteer-led project tracking COVID-19 cases and outbreaks in schools across Canada from September 2020 to June 2021. The goal of this project was to highlight the impact of COVID-19 on schools and families; to advocate for safer schools; and to advocate for transparency in our educational system.<br> This project is an initiative of grassroots advocacy group <a href="https://masks4canada.org/">Masks4Canada</a>.</p> <p>To learn more about project and see the interactive map of COVID-19 cases and outbreaks across Canadian schools as compiled by this project, visit <a href="https://covidschoolscanada.org/">https://covidschoolscanada.org/ </a></p> <p>For questions about these data, please contact <a href="mailto:shraddha.pai@utoronto.ca?subject=COVID%20Schools%20Canada%20Zenodo%20archive">Shraddha Pai</a>.</p>
Crowdsourced dataset of firefly trajectories obtained by automated stereo calibration of 360-degree cameras
<p>Advancements in animal tracking techniques, spanning from migrating mammals to swarming insects, have resulted in remarkable progress in the fields of behavioral ecology and conservation science. Recently, we have devised a method for tracking luminous fireflies in their natural habitat using stereoscopic pairs of 360-degree cameras. This method offers affordability, versatility, and ease of setup; however, the process of camera calibration has remained tedious and time-consuming. Now, we have introduced an enhanced algorithm that achieves spatial and temporal stereo calibration directly from the data, eliminating the need for manual procedures both in the field and during video processing. The algorithm relies on cross-correlation of flashing patterns and numerical estimation of camera pose. Utilizing this improved protocol and processing software, we have compiled an extensive dataset comprising over 100 reconstructed firefly swarms of various species. This data was gathered throughout the United States by numerous contributors following a straightforward protocol. The dataset holds significant potential for advancing our comprehension of firefly collective behavior, facilitating population monitoring, and expanding citizen science initiatives.</p>
Commodifying infrastructure spatial dynamics with crowdsourced smartphone data
Open the record for dataset details and reuse information.
Crowdsourced dataset of firefly trajectories obtained by automated stereo calibration of 360-degree cameras
Open the record for dataset details and reuse information.
New insights into the patterns and drivers of avian altitudinal migration from a growing crowdsourcing data source
Altitudinal migration is a common and important but understudied behavior in birds. Difficulty in characterizing avian altitudinal migration has prevented a comprehensive understanding of this behavior. To address this, we investigated the altitudinal migration patterns and explored potential drivers for a major proportion (~70%) of the entire resident bird community along an almost 4,000 m elevational gradient on the main island of Taiwan. Based on the occurrence records collected by citizen scientists, we examined the seasonal shifts in the center and the upper and lower boundaries of elevational distributions for 104 individual species. We then built phylogeny-controlled regression models to investigate the associations between the birds' seasonal distribution shifts and seven of their traits, and examined whether the observed shifts can be explained by three main hypotheses on potential drivers. Results showed that at least 60 species (58%) seasonally changed their distributions along elevations. While most of them (42 species) tended to move downhill in winter, a considerable number of species (14) tended to move uphill. While the species breeding at high or low elevations tended to move downhill in winter, those breeding at medium-low elevations tended to move or extend their distributions to higher elevations. Our regression models suggested that seasonal variations in climates and food availability could be major drivers of the behavior. However, the three hypotheses can only partially explain the observed downhill migration patterns and none of them can well explain the uphill patterns, indicating an important knowledge gap. This study investigated avian altitudinal migration from a new perspective with a novel and generalizable approach, and revealed interesting patterns that could be difficult to identify with conventional approaches. It demonstrated the power of citizen science data to provide new insights into this behavior by characterizing the general patterns and mechanisms across a large number of species.
Replication data for: Consumer preference testing of boiled sweetpotato (Ipomoea batatas (L.) Lam.) using crowdsourced citizen science in Ghana and Uganda
<p>Crowdsourced citizen science is an emerging approach in plant sciences. The triadic comparison of technologies (tricot) approach has been successfully utilised by demand-led breeding programmes to identify varieties for dissemination suited to specific geographic and climatic regions. An important feature of this approach is the independent way in which farmers individually evaluate the varieties on their own farms as ‘citizen scientists’. In this study, we adapted this approach to evaluate consumer preferences to boiled sweetpotato (<em>Ipomoea batatas</em> (L.) Lam) roots of 21 advanced breeding materials and varieties in Ghana and 6 released varieties in Uganda. We were specifically interested in evaluating if a more independent style of evaluation (home tasting) would produce results comparable to an approach that involves control over preparation (centralised tasting). We compiled data from 1,433 participants who individually contributed to a home tasting (de-centralised) and a centralised tasting trial in Ghana and Uganda, evaluating overall acceptability, and indicating the reasons for their preferences. Geographic factors showed important contribution to define consumers’ preference to boiled sweetpotato genotypes. Home and centralised tasting approaches gave similar rankings for overall acceptability, which was strongly correlated to taste. In both Ghana and Uganda, it was possible to robustly identify superior sweetpotato genotypes from consumers’ perspectives. Our results indicate that the tricot approach can be successfully applied to consumer preference studies.</p>
Library of Celsus - Crowdsourced photogrammetry
Library of Celsus - Crowdsourced model made from 200 photos from several sources public avaible. The Library of Celsus is an ancient Roman building in Ephesus, Anatolia, now part of Selçuk, Turkey. It was built in honour of the Roman Senator Tiberius Julius Celsus Polemaeanus completed between circa 114–117 A.Dby Celsus' son, Gaius Julius Aquila (consul, 110 AD). The library was "one of the most impressive buildings in the Roman Empire" and built to store 12,000 scrolls and to serve as a mausoleum for Celsus, who is buried in a crypt beneath the library in a decorated marble sarcophagus The Library of Celsus was the "third-largest library in the ancient world" behind both Alexandria and Pergamum. Source: Objaverse 1.0 / Sketchfab
AD-DAYR (Petra) Crowdsourced photogrammetry scan
El Monasterio o Al-Dayr es uno de los monumentos más grandes de la ciudad de Petra, mide 47 metros de ancho por 48.3 de alto. Fue construido con Khazna (El Tesoro de Petra) como modelo pero en este caso los bajo relieves fueron sustituidos por espacios para albergar esculturas. Tiene un pórtico columnado que se extiende por todo el frontal de la fachada. http://www.pasaporteblog.com/monasterio-ad-dayr-ciudad-de-petra/# Source: Objaverse 1.0 / Sketchfab
Crowdsourcing thesaurus
<p>Thesaurus used for text mining analysis on a corpus about digital libraries and crowdsourcing. It contains ideologies, taxonomies and motivations.</p> <p>Analysis were used for a French article : Andro, M. (2016). Bibliothèques numériques et crowdsourcing : analyses bibliométriques et text mining</p>
CrowdTruth/Crowdsourcing-NamedEntities-GoldStandard: Data release for crowdsourcing named entity gold standards
<p>This repository contains the experimental results of identifying and typing named entities in English Wikipedia sentences by using a hybrid Multi-NER - crowd-enhanced approach.</p>
Ermita de San Bernabé (Crowdsourced photos)
Reconstructed using crowdsourced photos (about 100) from different public sources. Está situada junto a la entrada principal del Complejo de Ojo Guareña y es parte de las cuevas. Se desconoce la fecha de su construcción, unos la sitúan en los siglos VIII-IX, pero Gómez Grinda cree que es del siglo XIII. Comenzó estando dedicada a San Tirso sólo. En el siglo XVIII reúne las dos advocaciones. Las reformas de acondicionamiento comienzan a mediados del siglo XVII. Las bóvedas poseen pinturas, algunas deterioradas por las filtraciones del agua de las corrientes de la gran cueva. También hay frescos y un retablo. El conjunto fue declarado Monumento Histórico Artístico Nacional el 23 de abril de 1970 Source: Objaverse 1.0 / Sketchfab
Paper is not enough: Crowdsourcing the T<sub>1</sub> mapping common ground via the ISMRM reproducibility challenge
Dataset provided for NeuroLibre preprint. Author repo: https://www.github.com/rrsg2020/note NeuroLibre fork:https://github.com/roboneurolibre/note <p>For details, please visit the corresponding <a href="https://github.com/neurolibre/neurolibre-reviews/issues/23">NeuroLibre technical screening.</a></p> <p><strong><a href="https://neurolibre.org" target="NeuroLibre">https://neurolibre.org</a></strong></p>
WhichDog: A crowdsourced dataset including candidate set-based labelling
<p>A dataset with crowdsourced labels for aggregation and supervised classification.<br> It contains 400 images of dogs from the Stanford Dogs dataset (http://vision.stanford.edu/aditya86/ImageNetDogs/). Images of dogs that belong to 32 different breeds (classes) are included. Annotators were asked to provide two types of labelling: full labelling (each labeler is allowed to provide a single label for each image) and candidate labelling (each labeler is allowed to provide a set of candidate labels for each image). It includes a total of 61227 annotations (30628 full and 30599 candidate) obtained from a set of 1028 different labelers.</p> <p>The labels were collected through the online crowdsourcing platform Amazon mTurk thanks to funds provided by the Basque Government through the Elkartek program (KK-2018/00071). The assignments were designed as sequences of 64 images that were given to the annotators. Each image in the sequence was provided together with a specific subset of possible labels (with the number of options ranging from 4 to 32), and a instruction for the annotator to perform a specific type of labelling (full or candidate). Each labeler performed at least one assignment. Not all the labelers completed the 64 annotations in their assignments.<br> <br> The file 'whichdog.zip' contains a folder ('images') with the 400 images of dogs, a text file ('breed_names.txt') that indicates the names of the different breeds and their assigned label (a number in the interval from 0 to 31) and a CSV file ('whichdog_all_annots.csv') that contains the information about the annotations. Each row of the CSV file represents a single annotation, and each column shows:<br> - image_id: ID number of the image.<br> - is_candidate: indicates whether the requested labelling is full (0) or candidate (1).<br> - labeler_id: ID number of the labeler.<br> - time: time employed by the labeler to perform the annotation.<br> - answer: label or set of labels provided by the labeler as annotation.<br> - options: subset of possible labels shown to the labeler.<br> - assignment_id: ID number of the assignment<br> - sequence_point: number that indicates the point of the sequence of images of the assignment in which the annotation was provided.<br> - class: ground truth label of the image.</p>
Crowdsourced WiFi database and benchmark software for indoor positioning
<p>This dataset contains two Wi-Fi databases (one for training and one for test/estimation purposes in indoor positioning applications), collected in a crowdsourced mode (i.e., via 21 different devices and different users), together with a benchmarking utility software (in Matlab and Python) to illustrate various algorithms of indoor positioning based solely on WiFi information (MAC addresses and RSS values). </p> <p>The data was collected in a 4-floor university building in Tampere, Finland, during Jan-Aug 2017 and it comprises 687 training fingerprints and 3951 test or estimation fingerprints.</p> <p>13.10.2017: Version 2 uploaded; the revised version contains improved readme files and improved Python SW.</p> <p>The dataset and/or the associated software are to be cited as follows:</p> <p>E.S. Lohan, J. Torres-Sospedra, P. Richter, H. Leppäkoski, J. Huerta, A. Cramariuc, “Crowdsourced WiFi-fingerprinting database and benchmark software for indoor positioning”, Zenodo repository, DOI 10.5281/zenodo.889798</p>
EUCROWD - Citizen's crowdsourcing in politics and policy-making
<p>TRILLION project and platform presentation as a case presentation on citizen's crowdsourcing in politics and policy-making.</p>
Crowdsourcing StoryLines with CrowdTruth
<p>Crowdsourced ground truth dataset for 1,204 sentences and 7,778 event pairs covering 22 news topics. The corpus was created by using the CrowdTruth methodology, as described in the following paper:</p> <ul> <li>Tommaso Caselli and Oana Inel: Crowdsourcing StoryLines: Harnessing the Crowd for Causal Relation Annotation. Events and Stories in the News Workshops, COLING 2018</li> </ul> <p>If you find this data useful in your research, please consider citing:</p> <pre>@inproceedings{caselli2018crowdsourcing, title={Crowdsourcing StoryLines: Harnessing the Crowd for Causal Relation Annotation}, author={Caselli, Tommaso and Inel, Oana}, booktitle={Proceedings of the Workshop Events and Stories in the News 2018}, pages={44--54}, year={2018} } </pre> <p>Crowdsourcing results and evaluation against expert data are available in folder: <code>|--data/results/</code></p> <p>Expert ground truth data is available in folder: <code>|--data/ground_truth/</code></p> <p>Aggregated raw crowdsourcig data is available in folder: <code>|--data/aggregated_input/</code></p> <p>Raw crowdsourcig data is available in folder: <code>|--data/input/</code></p> <p> </p> <p><strong>Running the notebooks</strong></p> <p>To run and regenerate the results, you need to install the stable version of the <em><strong>crowdtruth==2.0</strong></em> package from PyPI using:</p> <p>pip install crowdtruth==2.0</p> <pre> </pre>
Crowdsourcing Topical Relevance with CrowdTruth
<p>This repository contains the crowdsourcing annotations for topical relevance referenced in the following paper:</p> <ul> <li>Oana Inel, Giannis Haralabopoulos, Dan Li, Christophe Van Gysel, Zoltán Szlávik, Elena Simperl, Evangelos Kanoulas and Lora Aroyo: Studying Topical Relevance with Evidence-based Crowdsourcing. CIKM 2018.</li> </ul> <p> </p> <p>If you find this data useful in your research, please consider citing:</p> <pre>@inproceedings{inel2018studying, title={Studying Topical Relevance with Evidence-based Crowdsourcing}, author={Inel, Oana and Haralabopoulos, Giannis and Li, Dan and Van Gysel, Christophe and Szl{\'a}vik, Zolt{\'a}n and Simperl, Elena and Kanoulas, Evangelos and Aroyo, Lora}, booktitle={Proceedings of the 27th ACM International Conference on Information and Knowledge Management}, pages={1253--1262}, year={2018}, organization={ACM} }</pre> <p> </p> <p><strong>Running the notebooks</strong></p> <p>To run and regenerate the results, you need to install the stable version of the <em><strong>crowdtruth==2.0</strong></em> package from PyPI using:<br> pip install crowdtruth==2.0<br> </p>
New crowdsourced annotations for the GiantSteps Tempo dataset
<p>Raw beat and derived tempo annotations from <a href="https://doi.org/10.5281/zenodo.1492437">A Crowdsourced Experiment for Tempo Estimation of Electronic Dance Music</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.