Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

125

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

125 results for “Crowdsourcing”

Learn how ShareScore rates datasets ↗
zenodo36/100

Estonian Historical Newspaper Crowdsourced OCR Corrections

<div> <div>This dataset consists of newspaper articles from the National Library of Estonia's DIGAR archive and their respective crowdsourced corrections.</div> </div>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Results from Crowdsourcing Evaluation for Anirudh Prabhu's PhD Dissertation

<p>The dataset for Appendix D of Anirudh Prabhu&#39;s PhD dissertation. Contains the results of the human evaluation conducted on the crowdsourcing data. The &quot;Input&quot; folder contains the original results from the crowdsourcing experiment, and the &quot;Output&quot; and &quot;OutputPdfs&quot; folder contain the final outputs submitted by the participants of the crowdsourcing evaluation.&nbsp;</p>

opencc-by-nc-nd-3.0Aug 2021View details →
zenodo36/100

Crowdsourcing historical text and data with the Chinese Text Project

<p>Paper presented on Friday 11 June 2021 at the Digital Medievalist Global Symposium <em>The past, present, and future of Digital Medieval Studies</em> for the Asia &amp; Oceania Panel, in the session Engaging in Chinese Literature.</p> <p>The Chinese Text Project (<a href="https://ctext.org">https://ctext.org</a>) is a crowdsourced digital library of premodern Chinese writing, containing over 35 million pages of scanned primary source material and billions of words of transcribed text. In this talk I describe the implementation of a crowdsourced semantic annotation system for these texts, as well as the joint construction of a crowdsourced knowledge graph recording data covering close to 3000 years of Chinese history.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

FPCA - From mobile app-based crowdsourcing to crowd-trusted food price estimates in Nigeria: pre-processing and post-sampling strategy for optimal statistical inference

<p>Timely and reliable monitoring of commodity food prices is an essential requirement for the assessment of market and food security risks and the establishment of early warning systems, especially in developing economies. However, data from regional or national systems for tracking changes of food prices in sub-Saharan Africa lacks the temporal or spatial richness and is often insufficient to inform targeted interventions. In addition to limited opportunity for [near-]real-time assessment of food prices, various stages in the commodity supply chain are mostly unrepresented, thereby limiting insights on stage-related price evolution. Yet, governments and market stakeholders rely on commodity price data to make decisions on appropriate interventions or commodity-focused investments. Recent rapid technological development indicates that digital devices and connectivity services are becoming affordable for many, including in remote areas of developing economies. This offers a great opportunity both for the harvesting of price data (via new data collection methodologies, such as crowdsourcing/crowdsensing &mdash; i.e. citizen-generated data &mdash; using mobile apps/devices), and for disseminating it (via web dashboards or other means) to provide real-time data that can support decisions at various levels and related policy-making processes. However, market information that aims at improving the functioning of markets and supply chains requires a continuous data flow as well as quality, accessibility and trust. More data does not necessarily translate into better information. Citizen-based data-generation systems are often confronted by challenges related to data quality and citizen participation, which may be further complicated by the volume of data generated compared to traditional approaches. Following the food price hikes during the first noughties of the 21st century, the European Commission&#39;s Joint Research Centre (JRC) started working on innovative methodologies for real-time food price data collection and analysis in developing countries. The work carried out so far includes a pilot initiative to crowdsource data from selected markets across several African countries, two workshops (with relevant stakeholders and experts), and the development of a spatial statistical quality methodology to facilitate the best possible exploitation of geo-located data. Based on the latter, the JRC designed the Food Price Crowdsourcing Africa (FPCA) project and implemented it within two states in Northern Nigeria. The FPCA is a credible methodology, based on the voluntary provision of data by a crowd (people living in urban, suburban, and rural areas) using a mobile app, leveraging monetary and non-monetary incentives to enhance contribution, which makes it possible to collect, analyse and validate, and disseminate staple food price data in real time across market segments. The granularity and high frequency of the crowdsourcing data open the door to real-time space-time analysis, which can be essential for policy and decision making and rapid response on specific geographic regions.&nbsp;<a href="https://datam.jrc.ec.europa.eu/datam/perm/news/870?rdr=1666109837893">Link to the project</a></p>

opencc-by-4.0Oct 2022View details →
dryad36/100

Source Data for Crowdsourcing Bridge Dynamic Monitoring with Smartphone Vehicle Trips

<p>This data accompanies the study "Crowdsourcing Bridge Dynamic Monitoring with Smartphone Vehicle Trips" published in (Nature) Communications Engineering. This paper focuses on using large and inexpsensive datasets for obtaining information on the dynamics of bridges. In this study, data is collected by smartphones in moving vehicles as the cross over a bridge, in three distinct applications. Smartphone data was collected in controlled field experiments and uncontrolled Uber rides on a long-span suspension bridge in the USA (The Golden Gate Bridge) and an analytical method was developed to accurately recover modal properties. The method was also successfully applied to partially-controlled crowdsourced data collected on a short-span highway bridge in Italy. The results suggest that larve and inexpensive datasets collected by smartphones could play a role in monitoring the health of existing transportation infrastructure.</p> <p>The data provided includes the source data for the figures in the publication as well as the "controlled data" referenced in the study.</p>

opencc-zeroNov 2022View details →
zenodo36/100

Integration of the Drug–Gene Interaction Database (DGIdb 4.0) with open crowdsource efforts

<p>The Drug-Gene Interaction Database (DGIdb,&nbsp;<a href="http://www.dgidb.org/">www.dgidb.org</a>) is a web resource that provides information on drug-gene interactions and druggable genes from publications, databases, and other web-based sources. Drug, gene, and interaction data are normalized and merged into conceptual groups. The information contained in this resource is available to users through a straightforward search interface, an application programming interface (API), and TSV data downloads. DGIdb 4.0 is the latest major version release of this database. A primary focus of this update was integration with crowdsourced efforts, leveraging the Drug Target Commons for community-contributed interaction data, Wikidata to facilitate term normalization, and export to NDEx for drug-gene interaction network representations. Seven new sources have been added since the last major version release, bringing the total number of sources included to 41. Of the previously aggregated sources, 15 have been updated. DGIdb 4.0 also includes improvements to the process of drug normalization and grouping of imported sources. Other notable updates include the introduction of a more sophisticated Query Score for interaction search results, an updated Interaction Score, the inclusion of interaction directionality, and several additional improvements to search features, data releases, licensing documentation&nbsp;and the application framework.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

The Belfort dataset: Handwritten Text Recognition from Crowdsourced Annotations

<p>This dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946. Documents include deliberations, lists of councillors, convocations, and agendas.</p> <p>The dataset includes 24,105 text-line images that were automatically detected from pages. Up to 4 transcriptions are available for each line image: two from humans, and two from automatic models.</p> <p>We would like to thank the <em>Archives municipales de la ville de Belfort, France</em> for giving us access to these documents.</p>

opencc-by-4.0Jun 2023View details →
ClinicalTrials.gov36/100

Spurring Innovation to Promote HIV Testing: An RCT Evaluating Crowdsourcing

ClinicalTrials.gov study NCT02248558. IPD Sharing: Not stated. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov36/100

A Crowdsourced Social Media Portal for Parents of Very Young Children With Type 1 Diabetes

ClinicalTrials.gov study NCT03222180. IPD Sharing: YES. Countries: 1. Publications: 4.

controlledIPD-YESFeb 2026View details →
dryad36/100

Crowdsourced data reveal shortcomings in precipitation phase products for rain and snow partitioning

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad36/100

New insights into the patterns and drivers of avian altitudinal migration from a growing crowdsourcing data source

Open the record for dataset details and reuse information.

publicSep 2020View details →
dryad36/100

Source Data for Crowdsourcing Bridge Dynamic Monitoring with Smartphone Vehicle Trips

Open the record for dataset details and reuse information.

publicNov 2022View details →
zenodo32/100

Crowdsourcing public perceptions of urban green space quality: A case study of Rembrandt park in Amsterdam

<p><strong><em>City-dwellers are realizing the benefits of green spaces and are flocking to urban parks. City planners face the challenge of ensuring that urban green spaces are functional for all citizens. To make informed choices they need the right information and that is where the Mijn Park app can help. </em></strong></p> <p>Research shows that when considering the social functions of urban green spaces, quality is just as important as quantity. It is easy enough to map how much green spaces there are, but how do we measure their quality? How do city planners ensure that the city&rsquo;s green areas are attractive, accessible and inclusive &ndash; for everyone? The Vrije Universiteit Amsterdam in collaboration with the International Institute for Applied Systems Analysis developed Mijn Park, a mobile application that will help city planners do just that. As part of the LandSense Citizen Observatory, the &lsquo;Mijn Park&rsquo; (My Park) app asks respondents to go to several locations in a park and give subjective responses to those locations. They are then further questioned about how they use the whole park and how much they would like to see certain changes made in the park. This information provides information that can help to inform decisions about any renovations or improvements to the park.</p> <p>A pilot campaign was conducted in the summer of 2018 in Rembrandt park in Amsterdam and insights from the citizen-driven observations were shared with the Department of Planning and Sustainability of Amsterdam.</p> <p>This dataset includes responses and photographs collected by 129 unique volunteers providing 377 observations in select locations across Rembrandt Park. The following files are available:</p> <ul> <li>_Preview MijnPark-Amsterdam-LandSense.png</li> <li>Attributes-MijnPark-Amsterdam-LandSense.csv</li> <li>MijnPark-Amsterdam-LandSense.csv</li> <li>MijnPark-Amsterdam-LandSense.geoJSON</li> <li>README.txt</li> </ul> <p>&nbsp;</p> <p>This dataset is licensed under a Creative Commons Attribution 4.0 International. It is attributed to the <a href="https://landsense.eu/">LandSense Citizen Observatory</a>, <a href="https://www.vu.nl/">Vrije Universiteit</a> (VU), Amsterdam and the <a href="https://iiasa.ac.at/">International Institute for Applied Systems Analysis</a> (IIASA).</p> <p>This project has received funding from the European Union&rsquo;s Horizon 2020 research and innovation programme under grant agreement no 689812.</p>

opencc-by-4.0Jan 2020View details →
zenodo32/100

Crowdsourced air traffic data from The OpenSky Network 2020 [CC-BY]

<p><strong>Motivation</strong></p> <p>The data in this dataset is derived and cleaned from the full OpenSky dataset to illustrate the development of air traffic during the COVID-19 pandemic. It spans all flights seen by the network&#39;s more than 2500 members since 1 January 2020. More data will be periodically included in the dataset until the end of the COVID-19 pandemic.</p> <p><strong>License</strong></p> <p>Creative Commons CC-BY</p> <p>The only difference with the <a href="https://zenodo.org/record/3928550">original dataset</a> comes from anonymised aircraft information.</p> <p><strong>Disclaimer</strong></p> <p>The data provided in the files is provided as is. Despite our best efforts at filtering out potential issues, some information could be erroneous.</p> <ul> <li>Origin and destination airports are computed online based on the ADS-B trajectories on approach/takeoff: no crosschecking with external sources of data has been conducted.<br> Fields <strong>origin</strong> or <strong>destination</strong> are empty when no airport could be found.</li> <li>Aircraft information come from the OpenSky aircraft database. Fields <strong>typecode</strong> and <strong>registration</strong> are empty when the aircraft is not present in the database.</li> </ul> <p><strong>Description of the dataset</strong></p> <p>One file per month is provided as a csv file with the following features:</p> <ul> <li><strong>callsign</strong>: the identifier of the flight displayed on ATC screens (usually the first three letters are reserved for an airline: AFR for Air France, DLH for Lufthansa, etc.)</li> <li><strong>number</strong>: the commercial number of the flight, when available (the matching with the callsign comes from public open API)</li> <li><strong>aircraft_uid</strong>: a unique anonymised identifier for aircraft;</li> <li><strong>typecode</strong>: the aircraft model type (when available);</li> <li><strong>origin</strong>: a four letter code for the origin airport of the flight (when available);</li> <li><strong>destination</strong>: a four letter code for the destination airport of the flight (when available);</li> <li><strong>firstseen</strong>: the UTC timestamp of the first message received by the OpenSky Network;</li> <li><strong>lastseen</strong>: the UTC timestamp of the last message received by the OpenSky Network;</li> <li><strong>day</strong>: the UTC day of the last message received by the OpenSky Network;</li> <li><strong>latitude_1</strong>, <strong>longitude_1</strong>, <strong>altitude_1</strong>: the first detected position of the aircraft;</li> <li><strong>latitude_2</strong>, <strong>longitude_2</strong>, <strong>altitude_2</strong>: the last detected position of the aircraft.</li> </ul> <p><strong>Examples</strong></p> <p>Possible visualisations and a more detailed description of the data are available at the following page:<br> &lt;<a href="https://traffic-viz.github.io/scenarios/covid19.html">https://traffic-viz.github.io/scenarios/covid19.html</a>&gt;</p> <p><strong>Credit</strong></p> <p>If you use this dataset, please cite the original OpenSky paper:</p> <p>Matthias Sch&auml;fer, Martin Strohmeier, Vincent Lenders, Ivan Martinovic and Matthias Wilhelm.<br> &quot;Bringing Up OpenSky: A Large-scale ADS-B Sensor Network for Research&quot;.<br> In<em> Proceedings of the 13th IEEE/ACM International Symposium on Information Processing in Sensor Networks (IPSN)</em>, pages 83-94, April 2014.</p> <p>and the traffic library used to derive the data:</p> <p>Xavier Olive.<br> &quot;traffic, a toolbox for processing and analysing air traffic data.&quot;<br> <em>Journal of Open Source Software</em> 4(39), July 2019.</p>

opencc-by-4.0Jul 2020View details →
zenodo32/100

Evolutionary "crowdsourcing": alignment of fitness landscapes allows for cross-species adaptation of a horizontally transferred gene

<p>This repository accompanies the publication of <i><strong>Evolutionary "crowdsourcing": alignment of fitness landscapes allows for cross-species adaptation of a horizontally transferred gene</strong></i> by Kosterlitz et. al. This research project explores the cross-species adaptation of a horizontally transferred gene through evolutionary "crowdsourcing." The repository provides all relevant data, code, and figures associated with the publication, enabling users to replicate the results and explore the findings in-depth.</p>

opencc-by-4.0Oct 2023View details →
zenodo32/100

Crowdsourced LibriTTS Speech Prominence Annotations

<p>Dataset corresponding to the ICASSP 2024 paper "Crowdsourced and Automatic Speech Prominence Estimation" <a href="https://arxiv.org/abs/2310.08464">[link]</a></p> <p>This dataset is useful for training machine learning models to perform automatic emphasis annotaiton, as well as downstream tasks such as&nbsp;emphasis-controlled TTS, emotion recognition, and text summarization. The dataset is described in Section 3 (Emphasis Annotation Dataset). The contents of this section are copied below for convenience.</p> <p>We used our crowdsourced annotation system to perform human annotation on one eighth of the train-clean-100 partition of the LibriTTS [1] dataset. Specifically, participants annotated 3,626 utterances with a total length of 6.42 hours and 69,809 words from 18 speakers (9 male and 9 female). We collected at least one annotation of all 3,626 utterances, at least two annotations of 2,259 of those utterances, at least four annotations of 974 utterances, and at least eight annotations of 453 utterances. We did this in order to explore (in Section 6) whether it is more cost-effective to train a system on multiple annotations of fewer utterances or fewer annotations of more utterances. We paid 298 annotators to annotate batches of 20 utterances, where each batch takes approximately 15 minutes. We paid $3.34 for each completed batch (estimated $13.35 per hour). Annotators each annotated between one and six batches. We recruited on MTurk US residents with an approval rating of at least 99 and at least 1000 approved tasks. Today, microlabor platforms like MTurk are plagued by automated task-completion software agents (bots) that randomly fill out surveys. We filtered out bots by excluding annotations from an additional 107 annotators that marked more than 2/3 of words as emphasized in eight or more utterances of the 20 utterances in a batch. Annotators who fail the bot filter are blocked from performing further annotation. We also recorded participants' native country and language, but note these may be unreliable as many MTurk workers use VPNs to subvert IP region filters on MTurk [2].</p> <p>The average Cohen Kappa score for annotators with at least one overlapping utterance is 0.226 (i.e., ``Fair'' agreement)---but not all annotators annotate the same utterances, and this overemphasizes pairs of annotators with low overlap. Therefore, we use a one-parameter logistic model (i.e., a Rasch model) computed via py-irt [3], which predicts heldout annotations from scores of overlapping annotators with 77.7% accuracy (50% is random).</p> <p>The structure of this dataset is a single JSON file of word-aligned emphasis annotations. The JSON references file stems of the LibriTTS dataset, which can be found <a href="https://www.openslr.org/60/">here</a>. All code used in the creation of the dataset can be found <a href="https://github.com/interactiveaudiolab/emphases">here</a>. The format of the JSON file is as follows.</p> <p>&nbsp;</p> <pre><code>{ &lt;anonymized_participant_id_0&gt;: { "annotations": [ { "score": [ &lt;word_0_prominence&gt;,<br> &lt;word_1_prominence&gt;, &nbsp; &nbsp; &nbsp;... ], "stem": &lt;libritts_file_stem&gt;, "words": [ [ &lt;word_0&gt;, &lt;word_0_start_time&gt;, &lt;word_0_end_time&gt; ], [ &lt;word_1&gt;, &lt;word_1_start_time&gt;, &lt;word_1_end_time&gt; ],<br> ... &nbsp; &nbsp; ] },<br> ... ], "country": &lt;participant_0_country&gt;, "language": &lt;participant_0_language&gt; }, ... }</code></pre> <p><br>[1] Zen et al., &ldquo;LibriTTS: A corpus derived from LibriSpeech for text-to-speech,&rdquo; in Interspeech, 2019.<br>[2] Moss et al., &ldquo;Bots or inattentive humans? Identifying sources of low-quality data in online platforms,&rdquo; PsyArXiv preprint PsyArXiv:wr8ds, 2021.<br>[3] John Patrick Lalor and Pedro Rodriguez, &ldquo;py-irt: A scalable item response theory library for Python,&rdquo; INFORMS Journal on Computing, 2023.</p>

opencc-by-4.0Dec 2023View details →
zenodo32/100

First insights into the scale of invasions in African marine protected areas: leveraging global databases and crowdsourced data

<p>This dataset stems from the paper "First insights into the scale of invasions in African marine protected areas: leveraging global databases and crowdsourced data". Additional data is provided as supplementary material and is associated with the manuscript.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Test.wiki: Plataforma Crowdsourcing para Testes Funcionais

<p>Ent&atilde;o coloca na descri&ccedil;&atilde;o que &eacute; o v&iacute;deo referente ao trabalho Test.wiki: Plataforma Crowdsourcing para Testes Funcionais, apresenta&ccedil;&atilde;o ERES</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Qualitative research: Behavioral antecedents of crowdsourcing in science: academic teachers' perspective

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Focus research group: Behavioral antecedents of crowdsourcing in science: academic teachers' perspective

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record