Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
125
datasets available to search
ShareScore release 0.9.0
Dataset results
125 results for “Crowdsourcing”
Crowd4SDG - Crowdsourced image classification and damage assessment
<p>This data set contains crowdsourced classification and damage assessment of images of an earthquake extracted from social media. <br> <br> A data set of 907 images posted on Twitter related to the 2019 Albanian Earthquake, that are filtered and pre-classified using an automated technique is cross-validated for accuracy by two different crowds. One, digital humanitarian volunteers using the crowdsourcing platform <a href="http://www.crowd4ems.org">CROWD4EMS</a> and another, paid micro-taskers of the Amazon Mechanical Turk. In order to compare and evaluate the efficiency and accuracy of the volunteers and the paid micro taskers, ground truth is established with the help of a team of experts, who validated the same set of data. <br> <br> <strong>Parameters considered for volunteer contributions:</strong> The dataset was imported to the Crowd4EMS platform for Crowd contribution. In the forum, each volunteer will see the image to be validated along with the tweet text and the link to the original tweet. The user has to validate whether the given image is <em>relevant or</em> <em>irrelevant</em> to the disaster. In case of doubt, the user can refer to the tutorial explaining the relevance or skip the task. Once the image's relevance is validated, the user will be asked to label the <em>severity</em> of the impact, as seen in the image.</p> <p>The Automated algorithm has pre-classified the images as <em>severe </em>and <em>minimal </em>damage. The Crowd4EMS platform lets the volunteer label them as '<em>severe damage</em>,' <em>moderate damage'</em>,' <em>minimal damage', </em>and' <em>no damage'.</em> Each task has to be answered <em>at least three times</em>, and the final consensus is taken as per the<em> inter-rater agreement. </em><br> <br> <strong>Parameters considered for micro-taskers contribution:</strong> The dataset was imported to the <em>Amazon Mechanical Turk</em> platform for Crowd contribution. In the platform, each worker will see only the image that is to be categorised as follows: The user has to validate whether the given image depicts <em>severe damage, moderate damage, minimal damage, no damage </em>or <em>irrelevant</em> to the disaster. Each task has to be answered <em>at least ten times</em>, and the final consensus is taken as per the<em> inter-rater agreement. </em><br> <br> <strong>Acknowledgements:</strong> We want to thank Muhammad Imran of Qatar Computing Research Institute for sharing their pre-filtered social media imagery dataset on the Albanian earthquake from the Artificial Intelligence for Disaster Response (AIDR) Platform. We would also like to extend our gratitude to the volunteers for their contribution on the Crowd4EMS Platform.<br> </p>
Crowdsourcing vibration data stemming from different transportation usages
<p> </p> <p>Crowdsourcing vibration data stemming from different activities and transportation usages (by trains, by buses, by bicycles by walking). We present a comprehensive dataset that provides the pattern of five activities walking, cycling, taking a train, a bus or a taxi. The measurements are carried out by embedded sensor accelerometer in smartphones. The dataset offers dynamic responses of subjects carrying smartphones in varied styles as they performing the five activities through vibrations acquired by accelerometers. The dataset contains corresponding time stamps and vibrations in three directions longitudinal, horizontal, and vertical stored in an Excel Macro-enabled Workbook (xlsm) format can be used to train an AI model in a smartphone which has potentials to collect people’s vibration data and decides what movement is being conducted. Besides, with more data are received, the database can be updated and it can be fed to train the model with a larger dataset. The prevalent of the smartphone opens the door of crowdsensing which leads to the pattern of people talking public transports can be understood. Furthermore, the time consumed in each activity is available in the dataset. Therefore, with a better understanding of people using public transports, the service and schedule can be planned perceptively. Activities to obtain the dataset are jointly funded by H2020 and Hitachi Europe.</p>
Integration of the Drug-Gene Interaction Database (DGIdb 4.0) with open crowdsource efforts.
<p><strong>ABSTRACT </strong></p> <p>It contains the data of drug targets (gene names), uniprot identifiers, secondary linked data sources (e.g., PharmGKB), market drug name, chembl identifier, and pubchem compound identifier obtained from DGIdb.</p> <p><strong>Instructions: </strong></p> <p>Data were cleaned and duplicates were removed. Data were all categorical features.</p> <p><strong>Inspiration:</strong></p> <p>This dataset uploaded to U-BRITE for "DRG_DEPOT" summer 2023 team project. It is used for constructing R2G dataset, which will map drugs to their drug targets (gene -> protein = drug target)</p> <p><strong>Acknowledgements</strong></p> <p>Freshour SL, Kiwala S, Cotto KC, Coffman AC, McMichael JF, Song JJ, Griffith M, Griffith OL, Wagner AH. Integration of the Drug-Gene Interaction Database (DGIdb 4.0) with open crowdsource efforts. Nucleic Acids Res. 2021 Jan 8;49(D1):D1144-D1151. doi: 10.1093/nar/gkaa1084. PMID: 33237278; PMCID: PMC7778926.</p> <p>https://www.dgidb.org/</p> <p><strong>U-BRITE last update date:</strong> 06/09/2023</p>
Public Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior"
<p>Dataset for the "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior" paper, published in ICWSM 2018. The full text of the paper can be found <a href="https://arxiv.org/pdf/1802.00393.pdf">here</a>. </p> <p>The dataset provided here includes an updated version of the original dataset, with ~100k tweets annotated using the CrowdFlower platform: </p> <ul> <li> <p>hatespeech_id_label_PUBLIC_100K.csv: contains ~100K rows, where every row consists of a unique Tweet ID. </p> </li> <li> <p>hatespeech_text_label_vote_RESTRICTED_100K.csv: contains ~100K rows, where every row consists of the tweet text, its label according to majority annotation and the number of majority annotators. Available only <a href="https://zenodo.org/record/3706866#.Xmkh6i97FQI">here</a>.</p> </li> <li> <p>retweets.csv: contains ~2K rows, where every row consists of the row number in the hatespeech_text_label_vote_RESTRICTED_100K.csv file which is the first occurrence of a Tweet text followed by comma-separated row numbers of all other occurrences of the same Tweet text in the same file. There are ~8K other occurrences due to retweets. Available only <a href="https://zenodo.org/record/3706866#.Xmkh6i97FQI">here</a>.</p> </li> </ul> <p> </p> <p>UPDATE: It has come to our understanding that a number of the tweets are not available anymore for download on Twitter. Therefore, we provide <a href="https://zenodo.org/record/3706866#.YYLG6S8RqjQ">here </a>the hatespeech_text_label_vote_RESTRICTED_100K file with the full ~100K tweet texts, their associated majority label, and the number of votes for the majority label. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be in line with the T&C of Twitter). </p> <p>Please cite the paper in any published work that uses any of these resources. </p> <p>@inproceedings{founta2018large, <br> title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior}, <br> author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas}, <br> booktitle={11th International Conference on Web and Social Media, ICWSM 2018}, <br> year={2018}, <br> organization={AAAI Press} <br> } </p> <p>For any further questions contact a.m.founta at gmail dot com AND markos.charalambous at eecei dot cut dot ac dot cy </p>
Crowdsourcing Document Similarity Judgements
<p>This is the data obtained from crowdsourcing tasks which ask workers to provide similarity metrics between pairs of documents. Each document, as well as each pair, has a unique ID. We provide crowd workers with the pairs through three different task variations:</p> <ul> <li>Variation 1: We showed workers 5 pairs of documents and, for each, asked them to rate their similarity in a 4-level Likert scale (None, Low, Medium, High), tell us a confidence level of how sure they were (from 0 to 4) and a written reason as to why they chose that similarity level. For quality reasons, two of the 5 pairs were golden-standards, which means we knew their ratings already and checked the workers' responses. They had to give the golden pair with the higher similarity a higher score than the other golden pair, otherwise, their answer would be rejected.</li> <li>Variation 2: We repeated variation 1 but with a slight alteration: instead of a Likert scale for the similarity score, we asked for a Magnitude Estimation, which is any number above 0. It could be 1, 0.0001, 1000, 42, as long as it was coherent, as in a more similar pair had a higher score than a less similar pair and vice-versa;</li> <li>Variation 3: We showed workers 5 rankings. Each ranking had a main document and 3 auxiliary documents to be compared against the main one. They also had to report a confidence score and give a short written reason, just like variation 1. The first ranking is a golden-standard, and we knew the values for the 3 pairs in it (the pairs were the main document paired with each of the 3 auxiliary documents), and they had to give the golden pair with the highest similarity a higher rank than the one with the lower similarity.</li> </ul> <p>The raw results from the tasks are recorded in the JSON file CrowdResults.json. For a description of its contents, please read the file CrowdResults_README.md.</p> <p>These raw annotations from the crowd were then parsed into the three CSVs you see, each corresponding to the aggregated results from one of the task variations.</p> <ul> <li><em>final_scores_likert.csv</em> is the resulting scores for each pair using the variation 1 tasks; <ul> <li><em>pair_id </em>is a unique identifier for each pair;</li> <li><em>similarity_alg </em>is the similarity assigned to the pair of documents from an automated similarity algorithm;</li> <li><em>relation </em>is the type of relationship shown by the pair, where smaller values indicate more similar pairs;</li> <li><em>similarity_crowd_simple_maj </em>stores the simple majority result from the crowd's annotations;</li> <li><em>similarity_crowd_simple_mean </em>stores the mean of the crowd's annotations;</li> <li><em>similarity_crowd_simple_median </em>stores the median of the crowd's annotations;</li> </ul> </li> <li><em>final_scores_magnitude.csv</em> is the resulting scores for each pair using the variation 2 tasks; <ul> <li><em>pair_id </em>is a unique identifier for each pair;</li> <li><em>similarity_alg </em>is the similarity assigned to the pair of documents from an automated similarity algorithm;</li> <li><em>relation </em>is the type of relationship shown by the pair, where smaller values indicate more similar pairs;</li> <li><em>scaled_similarity_worker</em> is the magnitude score scaled based on worker's behaviours</li> <li><em>scaled_similarity_worker_docset </em>is the magnitude score scaled based both on the worker's behaviour and on the pair</li> </ul> </li> <li><em>final_scores_ranking.csv</em> is the resulting scores for each pair using the variation 3 tasks; <ul> <li><em>pair_id </em>is a unique identifier for each pair;</li> <li><em>similarity_alg </em>is the similarity assigned to the pair of documents from an automated similarity algorithm;</li> <li><em>relation </em>is the type of relationship shown by the pair, where smaller values indicate more similar pairs;</li> <li><em>mean_similarity</em> is the mean ranking from that value</li> </ul> </li> </ul> <p>This dataset was built and used as part of the <a href="https://theybuyforyou.eu/">TheyBuyForYou </a>project.</p>
Crowdsourced air traffic data from The OpenSky Network 2020 [CC-BY]
<p><strong>WARNING! </strong>This dataset is no longer updated after January 2022. Refer to the <a href="https://doi.org/10.5281/zenodo.3737101">original dataset</a> with different license terms for an up to date version.</p> <p><strong>Motivation</strong></p> <p>The data in this dataset is derived and cleaned from the full OpenSky dataset to illustrate the development of air traffic during the COVID-19 pandemic. It spans all flights seen by the network's more than 2500 members since 1 January 2019. More data will be periodically included in the dataset until the end of the COVID-19 pandemic.</p> <p><strong>License</strong></p> <p>Creative Commons CC-BY</p> <p>The only difference with the <a href="https://doi.org/10.5281/zenodo.3737101">original dataset</a> comes from anonymised aircraft information.</p> <p><strong>WARNING:</strong>This dataset is now longer updated after January 2022. The original dataset is still updated.</p> <p><strong>Disclaimer</strong></p> <p>The data provided in the files is provided as is. Despite our best efforts at filtering out potential issues, some information could be erroneous.</p> <ul> <li>Origin and destination airports are computed online based on the ADS-B trajectories on approach/takeoff: no crosschecking with external sources of data has been conducted.<br> Fields <strong>origin</strong> or <strong>destination</strong> are empty when no airport could be found.</li> <li>Aircraft information come from the OpenSky aircraft database. Fields <strong>typecode</strong> and <strong>registration</strong> are empty when the aircraft is not present in the database.</li> </ul> <p><strong>Description of the dataset</strong></p> <p>One file per month is provided as a csv file with the following features:</p> <ul> <li><strong>callsign</strong>: the identifier of the flight displayed on ATC screens (usually the first three letters are reserved for an airline: AFR for Air France, DLH for Lufthansa, etc.)</li> <li><strong>number</strong>: the commercial number of the flight, when available (the matching with the callsign comes from public open API)</li> <li><strong>aircraft_uid</strong>: a unique anonymised identifier for aircraft;</li> <li><strong>typecode</strong>: the aircraft model type (when available);</li> <li><strong>origin</strong>: a four letter code for the origin airport of the flight (when available);</li> <li><strong>destination</strong>: a four letter code for the destination airport of the flight (when available);</li> <li><strong>firstseen</strong>: the UTC timestamp of the first message received by the OpenSky Network;</li> <li><strong>lastseen</strong>: the UTC timestamp of the last message received by the OpenSky Network;</li> <li><strong>day</strong>: the UTC day of the last message received by the OpenSky Network;</li> <li><strong>latitude_1</strong>, <strong>longitude_1</strong>, <strong>altitude_1</strong>: the first detected position of the aircraft;</li> <li><strong>latitude_2</strong>, <strong>longitude_2</strong>, <strong>altitude_2</strong>: the last detected position of the aircraft.</li> </ul> <p><strong>Examples</strong></p> <p>Possible visualisations and a more detailed description of the data are available at the following page:<br> <<a href="https://traffic-viz.github.io/scenarios/covid19.html">https://traffic-viz.github.io/scenarios/covid19.html</a>></p> <p><strong>Credit</strong></p> <p>Martin Strohmeier, Xavier Olive, Jannis Lübbe, Matthias Schäfer, and Vincent Lenders<br> <strong>"</strong>Crowdsourced air traffic data from the OpenSky Network 2019–2020<strong>"</strong><br> <em>Earth System Science Data</em> 13(2), 2021<br> <a href="https://doi.org/10.5194/essd-13-357-2021">https://doi.org/10.5194/essd-13-357-2021</a></p> <p> </p>
Estimating the Global Distribution of Field Size using Crowdsourcing
<p>There is increasing evidence that smallholder farms contribute substantially to food production globally yet spatially explicit data on agricultural field sizes are currently lacking. Automated field size delineation using remote sensing or the estimation of average farm size at subnational level using census data are two approaches that have been used but both have limitations, e.g. limited geographical coverage by remote sensing or coarse spatial resolution when using census data. This paper demonstrates another approach to quantifying and mapping field size globally using crowdsourcing. A campaign was run in June 2017 where participants were asked to visually interpret very high resolution satellite imagery from Google Maps and Bing using the Geo-Wiki application. During the campaign, participants collected field size data for 130K unique locations around the globe. Using this sample, we have produced an improved global field size map (over the previous version) and estimated the percentage of different field sizes, ranging from very small to very large, in agricultural areas at global, continental and national levels. The results show that smallholder farms occupy no more than 40% of agricultural areas, which means that, potentially, there are much more smallholder farms in comparison with the current global estimate of 12%. The global field size map and the crowdsourced data set are openly available and can be used for integrated assessment modelling, comparative studies of agricultural dynamics across different contexts and contribute to SDG 2, among many others.</p> <p> </p> <p>The dataset (global field sizes.zip) contains:<br> - map of dominant field sizes (dominant_field_size_categories.tif) and description of legend items (legend_items.txt)<br> - table with all submissions by the participant (those who completed more than 10 classifications) and table description<br> - table with quality score of all the participants and table description<br> - table with estimated dominant field sizes at each location and table description</p>
The COUGHVID crowdsourcing dataset: A corpus for the study of large-scale cough analysis algorithms
<p><strong>Overview</strong></p> <p>Cough audio signal classification has been successfully used to diagnose a variety of respiratory conditions, and there has been significant interest in leveraging Machine Learning (ML) to provide widespread COVID-19 screening. The COUGHVID dataset provides over 30,000 crowdsourced cough recordings representing a wide range of subject ages, genders, geographic locations, and COVID-19 statuses. Furthermore, experienced pulmonologists labeled more than 2,000 recordings to diagnose medical abnormalities present in the coughs, thereby contributing one of the largest expert-labeled cough datasets in existence that can be used for a plethora of cough audio classification tasks. As a result, the COUGHVID dataset contributes a wealth of cough recordings for training ML models to address the world’s most urgent health crises.</p> <p><strong>Private Set and Testing Protocol</strong></p> <p>Researchers interested in testing their models on the private test dataset should contact us at coughvid@epfl.ch, briefly explaining the type of validation they wish to make, and their obtained results obtained through cross-validation with the public data. Then, access to the unlabeled recordings will be provided, and the researchers should send the predictions of their models on these recordings. Finally, the performance metrics of the predictions will be sent to the researchers. The private testing data is not included in any file within our Zenodo record, and it can only be accessed by contacting the COUGHVID team at the aforementioned e-mail address.</p> <p><strong>New Semi-Supervised Labeling</strong></p> <p>The third version of the COUGHVID dataset contains thousands of additional recordings obtained through October 2021. Additionally, the recordings containing coughs were re-labeled according to a semi-supervised learning algorithm that combined the user labels with those of the expert physicians, which were modeled using ML and expanded on the previously unlabeled data. These labels can be found in the "status_SSL" column of the "metadata_compiled.csv" file.</p>
Ransomwhere: A Crowdsourced Ransomware Payment Dataset
<p>Ransomwhere is the largest dataset of ransomware payment addresses, comprising over a billion dollars in payments. The dataset contains payment addresses, transactions, and the associated ransomware family. Anyone — whether a victim, a firm, or a security researcher — can help grow the data by submitting addresses of ransomware actors. For more information and to submit data, see <a href="https://ransomwhe.re/">ransomwhe.re</a>.</p>
crowdsourced body parameters of workshop attendants at the Helmholtz MT ARD ST3 meeting
<p>This data set was crowdsourced at the 2021 Helmholtz MT ARD ST3 meeting from attendants of the Machine Learning Tutorial on Sep 30, 2021. For more details on the event, see<br> https://indico.desy.de/event/28823/</p> <p>The CSV contains 4 columns:</p> <p>- is_female : fill with 1 if participant is female, filled with 0 if not female</p> <p>- shoesize_europe : your european shoesize</p> <p>- weight_kg : your weight in kilograms</p> <p>- height_cm : your height in centimeters</p> <p>Participants were encouraged to +1 or -1 to individual body properties in case they do not feel confident providing their true numbers.</p>
CrowdSpeech and Vox DIY: Benchmark Dataset for Crowdsourced Audio Transcription
<p>We collect and release CrowdSpeech — the first publicly available large-scale dataset of crowdsourced audio transcriptions. e show its applicability on an under-resourced language by constructing VoxDIY — a counterpart of CrowdSpeech for the Russian language.</p>
Image descriptions for 7471 of the DEArt images, obtained via the Zooniverse crowdsourcing platform
<p>This dataset contains crowdsourced image descriptions for 7471 images from the DEArt dataset. There are typically 4-5 descriptions per image, provided by volunteers of the Zooniverse crowdsourcing platform. The guidelines we used, together with some examples of possible captions, are also published in Zooniverse. We are in debt with all the Zooniverse volunteers and coordinators, and in particular with Samantha Blickhan, for making this possible. Some image descriptions were filtered out whenever they did not contain any textual data, nevertheless the dataset may still contain some descriptions that are not according to the guidelines we provided to the annotators. The dataset is a data export from Zooniverse in CSV format, in which the user information has been anonymized. The CSV also includes a field with the filename of the described image in the DEArt dataset.</p>
Extensive crowdsourced dataset of in-situ evaluated binaural soundscapes of private dwellings containing subjective sound-related and situational ratings along with person factors to study time-varying influences on sound perception — research data
<p><strong>Abstract:</strong></p> <p>The soundscape approach highlights the role of situational factors in sound evaluations; however, only a few studies have applied a multi‐domain approach including sound‐related, person‐related, and time‐varying situational variables. Therefore, we conducted a study based on the Experience Sampling Method to measure the relative contribution of a broad range of potentially relevant acoustic and non‐auditory variables in predicting indoor soundscape evaluations. Here we present the comprehensive dataset for which 105 participants reported temporally (rather) stable trait variables such as noise sensitivity, trait affect, and quality of life. They rated 6.594 situations regarding the soundscape standard dimensions, perceived loudness, and the saliency of its sound components and evaluated situational variables such as state affect, perceived control, activity, and location. To complement these subject‐centered data, we additionally crowdsourced object‐centered data by having participants make binaural measurements of each indoor soundscape at their homes using a low‐(self‐)noise recorder. These recordings were used to compute (psycho‐)acoustical indices such as the energetically averaged loudness level, the A‐weighted energetically averaged equivalent continuous sound pressure level, and the A‐weighted five‐percent exceedance level. This complex hierarchical data can be used to investigate time‐varying non‐auditory influences on sound perception and to develop soundscape indicators based on the binaural recordings to predict soundscape evaluations.</p> <p><strong>Content:</strong></p> <ul> <li><a href="https://zenodo.org/record/7858848/files/01%20StudyDescription.pdf">01 StudyDescription.pdf </a> <ul> <li>Description of the field study.</li> <li>Information about the methods and materials used.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/02%20Dataset.csv">02 Dataset.csv</a> <ul> <li>The dataset, consisting of 93 variables describing 6594 observations taken by 105 participants.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/03%20VariableDescriptions_EnglishPersonQuestionnaire.pdf">03 VariableDescriptions_EnglishPersonQuestionnaire.pdf</a> <ul> <li>Descriptions of all variables, their measurement scale, scale ranges and levels.</li> <li>Questions and task descriptions of the Experience Sampling Method questionnaire in German language with an English translation.</li> <li>English translations of questions asked in the person questionnaire.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/04%20ESM-Questionnaire.pdf">04 ESM-Questionnaire.pdf</a> <ul> <li>Screenshots of the original Experience Sampling Method questionnaire with English translations.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/05%20PersonQuestionnaire_OriginalGermanVersion.pdf">05 PersonQuestionnaire_OriginalGermanVersion.pdf</a> <ul> <li>Original version of the person questionnaire in German language.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/06%20HelpTexts.pdf">06 HelpTexts.pdf</a> <ul> <li>Descriptions of the study task.</li> <li>Explanations of the scales used in the questionnaire.</li> <li>Explanations of the sound categories and the soundscape composition.</li> <li>Explanation of the operation of the recording device.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_README.md">AcousticFeatures_README.md</a> <a href="https://zenodo.org/api/files/3d784540-c0f4-412f-8742-df1db6f5401d/TimeSeries_and_Spectrograms_README.md?versionId=9291496c-d2c6-4151-96f1-a2ad99e1a540"> </a> <ul> <li>Descriptions of the structure of the AcousticFeatures_xxx.csv and .zip files.</li> <li>Analyis settings used in Artemis Suite to generate the acoustic features.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_SingleValues.csv">AcousticFeatures_SingleValues.csv</a> <ul> <li>All acoustic features, aggregated to single values per feature, recording, and channel.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_Spectra.csv">AcousticFeatures_Spectra.csv</a> <ul> <li>Time-averaged 1/3 octave spectra of each channel of each recording, A-weichted and un-weighted.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_Spectrograms.zip">AcousticFeatures_Spectrograms.zip</a> <ul> <li>13188 .csv files with un-weighted spetrograms of each channel of each recording.</li> </ul> </li> <li><a href="https://zenodo.org/record/7858848/files/AcousticFeatures_TimeSeries.zip">AcousticFeatures_TimeSeries.zip</a> <ul> <li>A .csv file containing LAeq and LZeq time series of each channel of each recording.</li> </ul> </li> </ul> <p><strong>Publications refering to this dataset:</strong></p> <p>Versümer, Siegbert; Steffens, Jochen; Weinzierl, Stefan (currently under review): "The role of loudness predictions, personal and situational factors in day-to-day loudness assessments of indoor soundscapes."</p> <p><strong>Funding:</strong></p> <p>This study was sponsored by the German Federal Ministry of Education and Research. “FHprofUnt” funding code: 13FH729IX6. </p> <p><strong>License: </strong></p> <p>CC 4.0 BY, <a href="https://creativecommons.org/licenses/by/4.0/legalcode">https://creativecommons.org/licenses/by/4.0/legalcode</a></p> <p><strong>Version history:</strong></p> <p>Details can be found in the <a href="https://zenodo.org/api/files/a15d6a91-1a35-4b5e-a7ec-da8a9bcbee2b/Changelog.md">Changelog.md</a> file.</p> <ul> <li> V.01.0. March 7, 2023: Initial publication. <a href="https://doi.org/10.5281/zenodo.7193938">https://doi.org/10.5281/zenodo.7193938</a></li> <li> V.01.1. April 25, 2023. <a href="https://doi.org/10.5281/zenodo.7858848">https://doi.org/10.5281/zenodo.7858848</a></li> </ul>
A crowdsourced chemical-induced disease relation corpus
<p>A sentence-bound chemical-induced disease relation corpus was produced with crowdsourcing as part of the BioCreative V challenge.</p>
A crowdsourced sentence-bound chemical-induced disease relationship corpus
<p>A set of 3000 abstracts from PubMed were annotated for sentence-bound chemical-induced disease relationships in order to train a machine learning algorithm for the BioCreative V challenge.</p>
Dataset from Remote analysis of Sputum Smears for Mycobacterium Tuberculosis Quantification using Digital Crowdsourcing
<p>Worldwide, TB is one of the top 10 causes of death and the leading cause from a single infectious agent. Although the development and roll out of Xpert MTB/RIF has recently become a major breakthrough in the field of TB diagnosis, smear microscopy remains the most widely used method for TB diagnosis, especially in low- and middle-income countries.</p> <p>This is a minimal dataset to reproduce our research that tests the feasibility of a crowdsourced approach to tuberculosis image analysis. In particular, we investigated whether anonymous volunteers with no prior experience would be able to count acid-fast bacilli in digitized images of sputum smears by playing an online game. Following this approach 1790 people identified the acid-fast bacilli present in 60 digitized images, the best overall performance was obtained with a specific number of combined analysis from different players and the performance was evaluated with the F1 score, sensitivity and positive predictive value, reaching values of 0.933, 0.968 and 0.91, respectively.</p> <p>The dataset includes 24 digitized images of sputum smears and the corresponding gameplays clicks. </p>
Data from: Crowdsourcing training material for automated bird sound classification – a pilot study
<p>Data from the manuscript "Crowdsourcing training material for automated bird sound classification – a pilot study" by Petteri Lehikoinen, Meeri Rannisto, Ulisses Camargo, Aki Aintila, Patrik Lauha, Esko Piirainen, Panu Somervuo & Otso Ovaskainen</p>
Emozionalmente: a crowdsourced Italian speech emotional corpus
<p>This repository contains Emozionalmente: an extensive simulated speech emotional corpus in Italian. The dataset comprises 6,902 labeled samples acted out by 431 amateur actors, each verbalizing 18 different sentences to express the Big Six emotions (anger, disgust, fear, joy, sadness, surprise) plus neutrality. The labels represent the emotional communicative intention of the actors (i.e., the seven emotional states).</p> <p>Key details about the dataset:<br>- **Recording specifications**: The recordings were generally obtained with non-professional equipment. They are .wav files, mono-channel, with a sample size of 16 bits and a sample rate of 16,000 Hz. Each audio recording lasts 3.81 seconds on average (SD = 0.99 seconds).<br>- **Validation**: To validate the emotional content of the clips, 829 humans evaluated each audio recording, providing five evaluations per audio. The general Unweighted Average Recall (UAR) achieved by the evaluators was 66%, which is comparable to previous literature in the field.</p> <p>The repository includes the following additional resources:<br>1. **Demographic information**: Three .csv files describing the demographics of the actors and evaluators, as well as the emotions they expressed and recognized for each audio sample.<br>2. **Data splits**: A speaker-independent train-dev-test split, stratified by emotion, gender, and age.</p> <p>If you use this dataset, please cite the following paper:</p> <blockquote> <p>F. Catania, J. W. Wilke and F. Garzotto,<br><em>"Emozionalmente: A Crowdsourced Corpus of Simulated Emotional Speech in Italian,"</em><br>IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 1142–1155, 2025.<br>doi: <a target="_new" rel="noopener">10.1109/TASLPRO.2025.3540662</a></p> </blockquote>
Video of Using LOD to crowdsource Dutch WW2 underground newspapers on Wikipedia, SWIB2016, Bonn, 29-11-2016
<p><strong> </strong> Video registration of presentation about <a title="File:Using LOD to crowdsource Dutch WW2 underground newspapers on Wikipedia, SWIB2016, 29-11-2016.pdf" href="https://commons.wikimedia.org/wiki/File:Using_LOD_to_crowdsource_Dutch_WW2_underground_newspapers_on_Wikipedia,_SWIB2016,_29-11-2016.pdf">Using Linked Open Data to crowdsource Dutch WW2 underground newspapers on Wikipedia</a>, <a href="https://swib.org/swib16/programme.html" rel="nofollow">SWIB2016</a>, 29-11-2016 in Bonn.</p> <p>The video describes a project <a title="nl:Wikipedia:Wikiproject/Verzetskranten" href="https://nl.wikipedia.org/wiki/Wikipedia:Wikiproject/Verzetskranten">to systematically describe and interlink 1,300 Dutch underground newspapers from World War 2</a> on Wikipedia using linked open data.</p> <p>The project extracts contextual information about the newspapers from a book, converts it to structured data, and generates Wikipedia stubs linked to metadata, full texts, and each other.</p> <p>Volunteers are expanding the stubs into full articles, improving access to information about this historical period.</p> <div> <h2>Abstract</h2> </div> <p>During the second World War some 1.300 illegal newspapers were issued by the Dutch resistance. Right after the war as many of these newspapers as possible were physically preserved by Dutch memory institutions. They were described in formal library catalogues that were digitized and brought online in the 1990s. In 2010 the national collection of underground newspapers - some 200.000 pages - was full-text digitized in Delpher, the national aggregator for historical full-texts. Having created online metadata and full-texts for these publications, the third pillar <em>context</em> was still missing, making it hard for people to understand the historic background of the newspapers. We are currently running a project to tackle this contextual problem. We started by extracting contextual entries from a hard-copy standard work on Dutch illegal press and combined these with data from the library catalogue and Delpher into a central LOD triple store. We then created links between historically related newspapers and used Named Entity Recognition to find persons, organisations and places related to the newspapers. We further semantically enriched the data using DBPedia. Next, using an article template to ensure uniformity and consistency, we generated 1.300 Wikipedia article stubs from the database. Finally, we sought collaboration with the Dutch Wikipedia volunteer community to extend these stubs into full encyclopedic articles. In this way we can give every newspaper its own Wikipedia article, making these WW2 materials much more visible to the Dutch public, over 80% of whom uses Wikipedia. At the same time the triple store can serve as a source for alternative applications, like data visualizations. This will enable us to visualize connections and networks between underground newspapers, as they developed over time between 1940 and 1945.</p> <div> <h2>Presentation slides</h2> </div> <ul> <li><a title="File:Using LOD to crowdsource Dutch WW2 underground newspapers on Wikipedia, SWIB2016, 29-11-2016.pdf" href="https://commons.wikimedia.org/wiki/File:Using_LOD_to_crowdsource_Dutch_WW2_underground_newspapers_on_Wikipedia,_SWIB2016,_29-11-2016.pdf">Using LOD to crowdsource Dutch WW2 underground newspapers on Wikipedia</a></li> <li><a href="https://swib.org/swib16/slides/janssen_using_lod.pdf" rel="nofollow">https://swib.org/swib16/slides/janssen_using_lod.pdf</a></li> <li><a href="https://doi.org/10.5281/zenodo.13132987" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13132987</a></li> </ul>
Deprecated Dataset for "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior"
<p>This dataset is deprecated. <strong>The updated version of this Dataset is here:</strong> <a href="https://zenodo.org/record/3678559#.Xl9-Ji97FhE">https://zenodo.org/record/3678559#.Xl9-Ji97FhE</a></p> <p>Dataset for the publication "Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior". Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos and Nicolas Kourtellis. International AAAI Conference on Web and Social Media (ICWSM), 2018.</p> <p>The dataset provided here includes an updated version of the original dataset, with ~100k tweets annotated using the CrowdFlower platform:</p> <ul> <li>hatespeech_labels.csv: contains ~100k rows, where every row consists of a unique Tweet ID and its associated majority annotation</li> </ul> <p><em>UPDATE</em>: It has come to our understanding that a number of the tweets are not available anymore for download on Twitter. Therefore, <strong>upon request</strong>, we can provide one more file with the full ~100k tweet text and their associated majority labels. The tweets are shuffled so that there is no connection between tweet IDs and texts (in order to be aligned with the T&C of Twitter).</p> <p>To obtain the file contact a.m.founta at gmail dot com <strong>AND </strong>antonis26papa at gmail dot com.</p> <p><em>Please cite the paper in any published work that uses any of these resources.</em></p> <blockquote> <p>@inproceedings{founta2018large,<br> title={Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior},<br> author={Founta, Antigoni-Maria and Djouvas, Constantinos and Chatzakou, Despoina and Leontiadis, Ilias and Blackburn, Jeremy and Stringhini, Gianluca and Vakali, Athena and Sirivianos, Michael and Kourtellis, Nicolas},<br> booktitle={11th International Conference on Web and Social Media, ICWSM 2018},<br> year={2018},<br> organization={AAAI Press}<br> }</p> </blockquote> <p>For any further questions contact a.m.founta at gmail dot com.</p> <p> </p> <p>Publication DOI: <a href="https://doi.org/10.5281/zenodo.1443348">https://doi.org/10.5281/zenodo.1443348</a></p> <p>Github: <a href="https://github.com/ENCASEH2020/hatespeech-twitter">https://github.com/ENCASEH2020/hatespeech-twitter</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.