Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
227
datasets available to search
ShareScore release 0.7.1
Dataset results
227 results for “Crowding out”
Lexical Relations from the Wisdom of the Crowd 1.0
<p>A set of 300 most frequent nouns has been extracted from the Russian National Corpus. Then, each method or resource, including RuThes, produced at most five hypernyms, if possible. In case it is not possible, missing answers treated as empty results. This resulted in 9 322 unique non-empty subsumption pairs that have been passed for crowdsourcing annotation on the Yandex.Toloka microtask platform. Each pair has been annotated by seven different annotators whose mother tongue is Russian and the age is at least 20 by February 1, 2017.</p> <p>The layout of the human intelligence task (HIT) design assumes the direct answer to a simple question: does the given pair of words represent a meaningful <em>is-a</em> relation? Since the crowd workers are not expert lexicographers and this question might be difficult for them, it has been rephrased as “Is it correct that a <em>kitten</em> is a kind of <em>mammal</em>?” (in Russian).</p> <p>The answers have been aggregated using the Yandex.Toloka proprietary answer aggregation mechanism. As the result, 3 940 out of 9 322 pairs have been annotated as positive while the rest 5 382 have been annotated as negative.</p> <p>Interestingly, the workers were more confident in negative answers rather than in the positive ones. These negative answers are extremely useful for both training and testing different relation extraction methods. To the best of our knowledge, this is the first dataset of this kind made for the Russian language using microtask-based crowdsourcing.</p>
Large-Scale Dataset for Radio Frequency based Device-Free Crowd Estimation
<p>This dataset serves to estimate the status, in particular the size, of a crowd given the impact on radio frequency communication links within a wireless sensor network. To quantify this relation, signal strengths across sub-GHz communication links are collected at the premises of the Tomorrowland music festival. The communication links are formed between the network nodes of wireless sensor networks deployed in three of the festival's stage environments. </p> <p>The table below lists the eighteen dataset files. They are collected at the music festival's 2017 and 2018 editions. There are three environments, labeled: ‘Freedom Stage 2017’, ‘Freedom Stage 2018’, and ‘Main Comfort 2018’. Each environment has both 433 MHz and 868 MHz data. The measurements at each environment were collected over a period of three festival days. The dataset files are formatted as Comma-Separated Values (CSV).</p> <pre><code class="language-markdown">| Dataset file | Reference file | Number of messages | |-------------------- |------------------------- |-------------------- | | free17_433_fri.csv | None | 393 852 | | free17_868_fri.csv | None | 472 202 | | free17_433_sat.csv | free17_transactions.csv | 996 033 | | free17_868_sat.csv | free17_transactions.csv | 1 023 059 | | free17_433_sun.csv | free17_transactions.csv | 1 007 066 | | free17_868_sun.csv | free17_transactions.csv | 1 036 456 | | free18_433_fri.csv | None | 765 024 | | free18_868_fri.csv | None | 757 657 | | free18_433_sat.csv | free18_transactions.csv | 711 438 | | free18_868_sat.csv | free18_transactions.csv | 714 390 | | free18_433_sun.csv | free18_transactions.csv | 648 329 | | free18_868_sun.csv | free18_transactions.csv | 656 290 | | main18_433_fri.csv | None | 791 462 | | main18_868_fri.csv | None | 908 407 | | main18_433_sat.csv | main18_counts.csv | 863 666 | | main18_868_sat.csv | main18_counts.csv | 884 682 | | main18_433_sun.csv | main18_counts.csv | 903 862 | | main18_868_sun.csv | main18_counts.csv | 894 496 |</code></pre> <p>In addition to the datasets and reference files, a software example is provided to illustrate the data use and visualise the initial findings and relation between crowd size and network signal strength impact.</p> <p>In order to use the software, please retain the following file structure: </p> <pre><code class="language-markdown">. ├── data ├── data_reference ├── graphs └── software</code></pre> <p>The peer-reviewed data descriptor for this dataset has now been published in MDPI Data - an open access journal aiming at enhancing data transparency and reusability, and can be accessed here: <a href="https://doi.org/10.3390/data5020052">https://doi.org/10.3390/data5020052</a>.<br> Please cite this when using the dataset.</p>
Lexical Relations from the Wisdom of the Crowd 1.1
<p>A set of 300 most frequent nouns has been extracted from the Russian National Corpus. Then, each method or resource, including RuThes and RuWordNet, produced at most five hypernyms, if possible. In case it is not possible, missing answers treated as empty results. This resulted in 10,600 unique non-empty subsumption pairs that have been passed for crowdsourcing annotation on the Yandex.Toloka microtask platform. Each pair has been annotated by seven different annotators whose mother tongue is Russian and the age is at least 20 by February 1, 2017.</p> <p>The layout of the human intelligence task (HIT) design assumes the direct answer to a simple question: does the given pair of words represent a meaningful <em>is-a</em> relation? Since the crowd workers are not expert lexicographers and this question might be difficult for them, it has been rephrased as “Is it correct that a <em>kitten</em> is a kind of <em>mammal</em>?” (in Russian).</p> <p>The answers have been aggregated using the Yandex.Toloka proprietary answer aggregation mechanism. As the result, 4,576 out of 10,600 pairs have been annotated as positive while the rest 6,024 have been annotated as negative.</p> <p>Interestingly, the workers were more confident in negative answers rather than in the positive ones. These negative answers are extremely useful for both training and testing different relation extraction methods. To the best of our knowledge, this is the first dataset of this kind made for the Russian language using microtask-based crowdsourcing.</p>
Supplementary Material for "Using Unstructured Crowd-sourced Data to Evaluate Urban Tolerance of Terrestrial Native Animal Species within a California Mega-City"
<p>This data repository is for the publication "Using Unstructured Crowd-sourced Data to Evaluate Urban Tolerance of Terrestrial Native Animal Species within a California Mega-City" and contains all R scripts and data files to reproduce results as well as all supplementary tables and figures.</p>
The Unfolding Journey of Superoxide Dismutase 1 Barrels Under Crowding: Atomistic Simulations Shed Light on Intermediate States and Their Interactions With Crowders
<p>This data accompanies the article entitled <em>The Unfolding Journey of Superoxide Dismutase 1 Barrels Under Crowding: Atomistic Simulations Shed Light on Intermediate States and Their Interactions With Crowders</em>, published in J. Phys. Chem. Lett. (<a href="https://doi.org/10.1021/acs.jpclett.0c00699">https://doi.org/10.1021/acs.jpclett.0c00699</a>).</p> <p><strong>01_SOD1bar_unfolding_REST2.zip: </strong>The zip archive includes REST2 trajectories for the three systems investigated in the paper: 1:1 packing, 2:1 packing, and the dilute case. The trajectories are saved in the GROMACS XTC file format, separately for each temperature (i=0,...,23). Given the large trajectory sizes, only protein coordinates (SOD1bar + crowders) are reported, and the output frequency is reduced to 100 ps. A starting geometry (in the Gromos87 GRO format) after equilibration of the initial packing is provided for each REST2 simulation (conf_prot.gro). Moreover, for each REST2 simulation, an xarray (http://xarray.pydata.org) dataset, saved in the netCDF file format, is included with computed fraction of native contacts, secondary structure content, and the Calpha RMSD of the barrel core (beta sheets beta1 - beta8) with respect to the crystal structure.</p> <p><strong>02_SOD1bar_geometries_representative_unfolding.zip:</strong> Representative SOD1bar geometries along the unfolding pathway (presented in Figure 3 of the paper).</p> <p><strong>03_SOD1bar_geometries_loopVII.zip: </strong>SOD1bar geometries with varying loop VII conformation which were isolated from dilute REST2 and which are presented in Figure S9 of the paper.</p>
Data and R-Scripts for "Quality and timing of crowd-based water level class observations"
<p>This are the data and the R-scripts used for the manuscript "Quality and timing of crowd-based water level class observations" accepted for publication in the journal Hydrological Processes in July 2020 as a Scientific Briefing. To run the code, just run the R-script with the name "RunThisForResults.R". Results will be written to the "Figures" and the "Results" folder.</p>
Enhancing Crowd Creativity as Innovation via Teamwork: Dataset
<p>These data files correspond to the results produced in the following paper currently accepted for publication.</p> <p>Pradeep K. Murukannaiah, Nirav Ajmeri, and Munindar P. Singh. 2022. Enhancing Creativity as Innovation via Asynchronous Crowdwork. In Proceedings of the 14th ACM Web Science Conference. Pages 1--9. To Appear.</p> <p>----------<br> Data files<br> ----------</p> <p>* all_scenarios.csv: 1,823 scenarios produced by MTurk workers</p> <p>* rated_scenarios.csv: 639 scenarios rated for creativity by three authors (700 scenarios were randomly selected for rating. 61 scenarios of these 700 scenarios were unclear or irrelevant and thus were discarded)</p> <p>* creativity.csv: data for RQ1 (Creativity)</p> <p>* personality-creativty.csv and team-composition-creativity.csv: data for RQ2 (Personality)</p> <p>* efficiency.csv: data for RQ3 (Efficiency)</p> <p>* emotions.csv: data for RQ4 (Emotions)</p>
Stabilizing or Destabilizing: Simulations of Chymotrypsin Inhibitor 2 under Crowding Reveal Existence of a Crossover Temperature
<p>This data accompanies the paper entitled <em>Stabilizing or Destabilizing: Simulations of Chymotrypsin Inhibitor 2 under Crowding Reveal Existence of a Crossover Temperature</em> (https://dx.doi.org/10.1021/acs.jpclett.0c03626).</p> <p>CI2_REST2.zip: The zip archive includes REST2 trajectories for the three systems investigated in the paper: dilute conditions, crowding by BSA, and crowding by lysozyme. The trajectories are saved in the GROMACS XTC file format, separately for each temperature (i=0,...,23). Given the large trajectory sizes, only protein coordinates (CI2 + crowder(s)) are reported, and the output frequency is reduced to 100 ps. A starting geometry (in the Gromos87 GRO format) after a short relaxation is provided for each REST2 simulation (conf_prot.gro). Moreover, for each REST2 simulation, an xarray (http://xarray.pydata.org) dataset, saved in the netCDF file format, is included with the following observables computed for CI2: fraction of native contacts relative to crystal structure, radius of gyration, secondary-structure content, fraction of native contacts evaluated separately for the alpha helix and the two beta strands.</p>
List of past crowd crushes
<p>List of past crowd crushes (until 2022) based on six existing sources: <a href="https://en.wikipedia.org/w/index.php?title=List_of_fatal_crowd_crushes&oldid=1136597272">Wikipedia</a>, <a href="https://github.com/Mind-the-Cap/Observatory/tree/main/Wikidata">Wikidata</a>, <a href="https://www.gkstill.com/ExpertWitness/CrowdDisasters.html">Keith Still's list</a>, the <a href="https://yorku.maps.arcgis.com/apps/webappviewer/index.html?id=e7c52856187642e19bd227865393432c">World Crowd Disasters Web App</a>, the <a href="https://www.workingwithcrowds.com/crowd-disasters-and-incidents/">Working with Crowds website</a> and <a href="https://link.springer.com/book/10.1007/978-3-030-90012-0">Introduction to Crowd Management</a></p>
SCoRe-LFC: Platform data on crowd collaboration in higher education
<p>SCoRe (short for Student Crowd Research) was a joint research project between the Universities of Bremen (UB), Hamburg (UHH) and Kiel (CAU), the Macromedia University of Applied Sciences (HMM) and the Ghostthinker GmbH (GT). The overall aim of the project was to develop a digital learning and research environment as well as didactic scenarios that foster collaborative processes of research-based learning in large groups of students (crowd). The main subject area was research for sustainable development. Towards this end, the project consortium drew on the partners’ expertise on advanced video-technologies (HHM), virtual collaboration in interdisciplinary and largescale groups (CAU), research-based learning (UHH) and education and research for sustainable development (UB). To achieve its goals, the project adopted a design-based research approach. The Project started in Oct. 2018 and was funded for 3.5 years by the Federal Ministry of Education and Research (BMBF) in a funding scheme on digital higher education.</p> <p>The work in the department of media-pedagogy and educational computer sciences at Kiel University was focused on the sub-project „SCoRe - learning and researching in the crowd“. The sub-project was aimed at the development, implementation and evaluation of pedagogical and organizational measures for the seeding, coordination and orchestration of collaborative research and learning processes in crowd scenarios. Particular emphasis was placed on crowd-specific characteristics of productive knowledge work in large and interdisciplinary groups.</p> <p>This dataset contains interaction data as well as textual content data. As ongoing development of the software platform led to a continuous integration of new features into the platform itself as well as changes to the data collection functions, making this an evolving dataset. Some inconsistencies exist due to software bugs.</p> <p><strong><a href="https://scorelfc.github.io/gestaltungsbericht3/img/datastructure.png">Platform data structure diagram</a> </strong></p> <p>Further Readings to gain an understanding of the platform and its interaction posbilities (in german):</p> <p><a href="https://scorelfc.github.io/gestaltungsbericht2">Design Report Prototype 2</a></p> <p><a href="https://scorelfc.github.io/gestaltungsbericht3">Design Report Prototype 3</a></p> <p> </p> <p><strong>Contained Files</strong></p> <table> <tbody> <tr> <td> <p><strong>Filename</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>annotations.csv</p> </td> <td> <p>Annotations (comment and/or drawings on the video) of video files</p> </td> </tr> <tr> <td> <p>content.csv</p> </td> <td> <p>Content of <a href="https://scorelfc.github.io/gestaltungsbericht3/umsetzungen/u03/">sections</a></p> </td> </tr> <tr> <td> <p>events.csv</p> </td> <td> <p>All events triggered by user interaction</p> </td> </tr> <tr> <td> <p>media.csv</p> </td> <td> <p>Uploaded <a href="https://scorelfc.github.io/gestaltungsbericht3/umsetzungen/u03/">images and videos</a></p> </td> </tr> <tr> <td> <p>messages.csv</p> </td> <td> <p><a href="https://scorelfc.github.io/gestaltungsbericht3/umsetzungen/u08/">Chat messages</a></p> </td> </tr> <tr> <td> <p>sequences.csv</p> </td> <td> <p><a href="https://scorelfc.github.io/gestaltungsbericht3/umsetzungen/u21/">Sequences</a> of video files</p> </td> </tr> </tbody> </table> <p> </p> <p><strong>Columns</strong></p> <p>(not all are present in each file. 0, “null” or “none” might mean not applicable)</p> <table> <tbody> <tr> <td> <p><strong>Column name</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Format</strong></p> </td> </tr> <tr> <td> <p>Index (empty column name) </p> </td> <td> <p>unique identifier of the corresponding event in the original dataset</p> </td> <td> <p>UUID (int on rare occasions)</p> </td> </tr> <tr> <td> <p>Actor-Name</p> </td> <td> <p>Unique identifier of an actor – “MA” identifies project staff </p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Annotation-ID</p> </td> <td> <p>Unique identifier of an annotation</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Annotation-Text</p> </td> <td> <p>Label of an annotation</p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Version-ID</p> </td> <td> <p>Unique identifier of a version of an auditable object (e.g. a section)</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Version-Changelog</p> </td> <td> <p>Changelog message on saving a new version of a section</p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Case-ID</p> </td> <td> <p>Unique identifier of a case (if applicable, coded by research team)</p> </td> <td> <p>String </p> </td> </tr> <tr> <td> <p>Media-Caption</p> </td> <td> <p>Title of a media file (image, video)</p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Media-ID</p> </td> <td> <p>Unique identifier of a media file (image, video)</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Media-Timestamp</p> </td> <td> <p>Timestamp in a video</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Message-ID</p> </td> <td> <p>Unique identifier of a chat message</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Message-Text</p> </td> <td> <p>Content of a chat-message</p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Object-Type</p> </td> <td> <p>Type of an object an action refers to</p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Project-ID</p> </td> <td> <p>Unique identifier of a project</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Research-Task-Type</p> </td> <td> <p>Type of research task (if applicable, coded by research team, see table below)</p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Section-Content</p> </td> <td> <p>Content of a section (in a specific version) </p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Section-Outline-Level</p> </td> <td> <p>Outline level of a section (in a specific version)</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Section-ID</p> </td> <td> <p>Unique identifier of a section</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Section-Index</p> </td> <td> <p>Position of a section in the project (in a specific version)</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Section-Status</p> </td> <td> <p>Status of a section</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Section-Title</p> </td> <td> <p>Title of a section (in a specific version)</p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Sequence-Description</p> </td> <td> <p>Description of a video sequence</p> </td> <td> <p>string</p> </td> </tr> <tr> <td> <p>Sequence-Duration</p> </td> <td> <p>Length of a video sequence</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Sequence-ID</p> </td> <td> <p>Unique identifier of a sequence</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>Sequence-Timestamp</p> </td> <td> <p>Timestamp of the start of a sequence in a video</p> </td> <td> <p>int</p> </td> </tr> <tr> <td> <p>timestamp</p> </td> <td> <p>timestamp of an event</p> </td> <td> <p>datetime</p> </td> </tr> <tr> <td> <p>Verb</p> </td> <td> <p>Action type of an event (see table below)</p> </td> <td> <p>string</p> </td> </tr> </tbody> </table> <p> </p> <p><strong>Verbs</strong></p> <table> <tbody> <tr> <td> <p><strong>Value</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>canceled editing of</p> </td> <td> <p>Actor canceled editing of a section</p> </td> </tr> <tr> <td> <p>clicked</p> </td> <td> <p>Actor clicked a link</p> </td> </tr> <tr> <td> <p>collapsed</p> </td> <td> <p>Actor collapsed a section (hides its content form being viewed)</p> </td> </tr> <tr> <td> <p>compared versions of</p> </td> <td> <p>Actor compared two versions of a section</p> </td> </tr> <tr> <td> <p>created</p> </td> <td> <p>Actor created a new section, video sequence, video annotation, video playback command, project or news</p> </td> </tr> <tr> <td> <p>deleted</p> </td> <td> <p>Actor deleted a section, video sequence, video annotation, video playback command, project or news</p> </td> </tr> <tr> <td> <p>ended</p> </td> <td> <p>Actor played a video hitting its end</p> </td> </tr> <tr> <td> <p>expanded</p> </td> <td> <p>Actor expanded a collapsed section</p> </td> </tr> <tr> <td> <p>inserted</p> </td> <td> <p>Actor inserted a video comment (on occasions instead of created)</p> </td> </tr> <tr> <td> <p>left</p> </td> <td> <p>Actor left a context (e.g. a project, a chat window) by e.g. closing it using platform functions, changing a browser tab, etc.</p> </td> </tr> <tr> <td> <p>mentioned</p> </td> <td> <p>Actor mentioned another actor in a chat message</p> </td> </tr> <tr> <td> <p>opened</p> </td> <td> <p>Actor opened a context (e.g. a project, a chat window) by e.g. accessing it using platform functions or changing a browser tab</p> </td> </tr> <tr> <td> <p>paused</p> </td> <td> <p>Actor paused a video</p> </td> </tr> <tr> <td> <p>played</p> </td> <td> <p>Actor played a video</p> </td> </tr> <tr> <td> <p>read</p> </td> <td> <p>Actor read an activity message or news</p> </td> </tr> <tr> <td> <p>read all messages and activities of</p> </td> <td> <p>Actor used switch to mark all chat and activity messages read</p> </td> </tr> <tr> <td> <p>restored</p> </td> <td> <p>Actor restored a deleted section</p> </td> </tr> <tr> <td> <p>reverted</p> </td> <td> <p>Actor restored a deleted section</p> </td> </tr> <tr> <td> <p>reverted version of</p> </td> <td> <p>Actor reverted a section to an earlier version</p> </td> </tr> <tr> <td> <p>seeked</p> </td> <td> <p>Actor seeked on a video timeline</p> </td> </tr> <tr> <td> <p>sent</p> </td> <td> <p>Actor sent a chat message</p> </td> </tr> <tr> <td> <p>started editing of</p> </td> <td> <p>Actor started editing of a section</p> </td> </tr> <tr> <td> <p>switched</p> </td> <td> <p>Actor switched chat focus between project and section chat</p> </td> </tr> <tr> <td> <p>typed</p> </td> <td> <p>Actor typed into the chat</p> </td> </tr> <tr> <td> <p>updated</p> </td> <td> <p>Actor updated an existing section (changing content, heading, heading-depth or status), video sequence, video annotation, video playback command, project or news</p> </td> </tr> <tr> <td> <p>uploaded</p> </td> <td> <p>Actor uploaded an image or video</p> </td> </tr> <tr> <td> <p>viewed</p> </td> <td> <p>Actor viewed an entity (had it on screen for 5 seconds), e.g. a section or video comment</p> </td> </tr> <tr> <td> <p>viewed history of</p> </td> <td> <p>Actor viewed history of a section</p> </td> </tr> </tbody> </table> <p> </p> <p><strong>Project-ID</strong></p> <table> <tbody> <tr> <td> <p><strong>Project-ID</strong></p> </td> <td> <p><strong>Case-IDs</strong></p> </td> <td> <p><strong>Title</strong></p> </td> <td> <p> </p> <p><strong>Type</strong></p> </td> <td> <p><strong>Period</strong></p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>a1-a*</p> </td> <td> <p>Urbane Grünflächen</p> </td> <td> <p>Research project</p> </td> <td> <p>1.11.20-31.3.21</p> </td> </tr> <tr> <td> <p>4</p> </td> <td> <p>b1-b*</p> </td> <td> <p>Nachhaltiger Verkehr</p> </td> <td> <p>Research project</p> </td> <td> <p>1.11.20-31.3.21</p> </td> </tr> <tr> <td> <p>166</p> </td> <td> <p>c1-c*</p> </td> <td> <p>UGF - Urbane Grünflächen</p> </td> <td> <p>Research project</p> </td> <td> <p>1.4.21-30.09.21</p> </td> </tr> <tr> <td> <p>168</p> </td> <td> <p> </p> </td> <td> <p>LGS - Foyer</p> </td> <td> <p>Onboarding of students in LGS Projects</p> </td> <td> <p>1.4.21-30.09.21</p> </td> </tr> <tr> <td> <p>188</p> </td> <td> <p> </p> </td> <td> <p>LGS - Reflexionsraum</p> </td> <td> <p>Reflection project for students in LGS Projects</p> </td> <td> <p>1.4.21-30.09.21</p> </td> </tr> <tr> <td> <p>207</p> </td> <td> <p> </p> </td> <td> <p>LGS - Nachhaltiger Konsum</p> </td> <td> <p>Research project</p> </td> <td> <p>1.4.21-30.09.21</p> </td> </tr> <tr> <td> <p>210</p> </td> <td> <p> </p> </td> <td> <p>LGS - Bildungsangebote für nachhaltige Entwicklung</p> </td> <td> <p>Research project</p> </td> <td> <p>1.4.21-30.09.21</p> </td> </tr> <tr> <td> <p>264</p> </td> <td> <p> </p> </td> <td> <p>LGS - Fahrradmobilität in Städten</p> </td> <td> <p>Research project</p> </td> <td> <p>1.4.21-30.09.21</p> </td> </tr> <tr> <td> <p>271</p> </td> <td> <p> </p> </td> <td> <p>Fahrradmobilität in Städten</p> </td> <td> <p>Research project</p> </td> <td> <p>1.10.21-30.11.21</p> </td> </tr> <tr> <td> <p>269</p> </td> <td> <p>e1-e*</p> </td> <td> <p>Kaufentscheidung vs. Nachhaltigkeit</p> </td> <td> <p>Research project</p> </td> <td> <p>1.10.21-30.11.21</p> </td> </tr> <tr> <td> <p>262</p> </td> <td> <p>d1-d*</p> </td> <td> <p>Urbane Grünflächen</p> </td> <td> <p>Research project</p> </td> <td> <p>1.10.21-30.11.21</p> </td> </tr> <tr> <td> <p>199</p> </td> <td> <p> </p> </td> <td> <p>Basiskurs</p> </td> <td> <p>Basic course for onboarding of students on the platform</p> </td> <td> <p>1.10.21-30.11.21</p> </td> </tr> <tr> <td> <p>5</p> </td> <td> <p> </p> </td> <td> <p>Glossar</p> </td> <td> <p>Glossar of definitions</p> </td> <td> <p>persistent</p> </td> </tr> <tr> <td> <p>6</p> </td> <td> <p> </p> </td> <td> <p>Erste Schritte</p> </td> <td> <p>How to start using score-docs platform</p> </td> <td> <p>persistent</p> </td> </tr> <tr> <td> <p>8</p> </td> <td> <p> </p> </td> <td> <p>Testbereich</p> </td> <td> <p>Area for testing score-docs functionalities</p> </td> <td> <p>persistent</p> </td> </tr> <tr> <td> <p>7</p> </td> <td> <p> </p> </td> <td> <p>Hilfestellungen</p> </td> <td> <p>Helpful links </p> </td> <td> <p>persistent</p> <p> </p> </td> </tr> </tbody> </table> <p>Other Project-IDs refer to personal assessment documents of individual students which are not included in content.csv</p> <p><strong>Research-Task-Type</strong></p> <table> <tbody> <tr> <td> <p><strong>Value</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>Erheben</p> </td> <td> <p>Data collection</p> </td> </tr> <tr> <td> <p>Analysieren</p> </td> <td> <p>Case-specific data analysis and sensemaking</p> </td> </tr> <tr> <td> <p>Synthetisieren</p> </td> <td> <p>Cross-case data analysis and sensemaking</p> </td> </tr> <tr> <td> <p>Sonstiges</p> </td> <td> <p>Other, e.g. communication between participants</p> </td> </tr> </tbody> </table>
early developmental milestones from a crowd-based application
<p>Dataset for a research paper on developmental milestone data from a crowd-authored tracking application. </p>
OpenChart-SE: A corpus of artificial Swedish electronic health records for imagined emergency care patients written by physicians in a crowd-sourcing project
<p>Electronic health records (EHRs) are a rich source of information for medical research and public health monitoring. Information systems based on EHR data could also assist in patient care and hospital management. However, much of the data in EHRs is in the form of unstructured text, which is difficult to process for analysis. Natural language processing (NLP), a form of artificial intelligence, has the potential to enable automatic extraction of information from EHRs and several NLP tools adapted to the style of clinical writing have been developed for English and other major languages. In contrast, the development of NLP tools for less widely spoken languages such as Swedish has lagged behind. A major bottleneck in the development of NLP tools is the restricted access to EHRs due to legitimate patient privacy concerns. To overcome this issue we have generated a citizen science platform for collecting artificial Swedish EHRs with the help of Swedish physicians and medical students. These artificial EHRs describe imagined but plausible emergency care patients in a style that closely resembles EHRs used in emergency departments in Sweden. In the pilot phase, we collected a first batch of 50 artificial EHRs, which has passed review by an experienced Swedish emergency care physician. We make this dataset publicly available as OpenChart-SE corpus (version 1) under an open-source license for the NLP research community. The project is now open for general participation and Swedish physicians and medical students are invited to submit EHRs on the project website (<a href="https://github.com/Aitslab/openchart-se">https://github.com/Aitslab/openchart-se</a>), where additional batches of quality-controlled EHRs will be released periodically. </p> <p> </p> <p><strong>Dataset content</strong></p> <p><em>OpenChart-SE, version 1 corpus (txt files and and dataset.csv)</em></p> <p>The OpenChart-SE corpus, version 1, contains 50 artificial EHRs (note that the numbering starts with 5 as 1-4 were test cases that were not suitable for publication). The EHRs are available in two formats, structured as a .csv file and as separate textfiles for annotation. Note that flaws in the data were not cleaned up so that it simulates what could be encountered when working with data from different EHR systems. All charts have been checked for medical validity by a resident in Emergency Medicine at a Swedish hospital before publication.</p> <p> </p> <p><em>Codebook.xlsx</em></p> <p>The codebook contain information about each variable used. It is in XLSForm-format, which can be re-used in several different applications for data collection.</p> <p> </p> <p><em>suppl_data_1_openchart-se_form.pdf</em></p> <p>OpenChart-SE mock emergency care EHR form.</p> <p> </p> <p><em>suppl_data_3_openchart-se_dataexploration.ipynb</em></p> <p>This jupyter notebook contains the code and results from the analysis of the OpenChart-SE corpus.</p> <p> </p> <p>More details about the project and information on the upcoming preprint accompanying the dataset can be found on the project website (<a href="https://github.com/Aitslab/openchart-se">https://github.com/Aitslab/openchart-se</a>).</p>
Toloker Graph: Interaction of Crowd Annotators
<p>The graph contains 11,758 nodes and 519,000 edges representing interactions between crowd annotators on a project labeled on the <a href="https://toloka.ai/">Toloka</a> crowdsourcing platform (see the <a href="https://toloka.ai/en/docs/guide/concepts/overview">Toloka overview</a> for the details on the used terminology).</p> <p>Each node represents an individual annotator; nodes are provided with four numerical and three categorical features. An edge is drawn between a pair of annotators if they annotated the same task. Also, each node is provided with a label showing whether the annotator was banned on this project, or not.</p> <p><strong>Nodes</strong> are stored in the <a href="https://github.com/Toloka/TolokerGraph/blob/main/nodes.tsv">nodes.tsv</a> file in the TSV format of the following structure:</p> <ul> <li><code>id</code>: unique identifier of the annotator</li> <li><code>approved_rate</code>: percentage of the approved labels of this annotator</li> <li><code>skipped_rate</code>: percentage of the skipped tasks of this annotator</li> <li><code>expired_rate</code>: percentage of the expired tasks of this annotator</li> <li><code>rejected_rate</code>: percentage of the rejected labels of this annotator</li> <li><code>education</code>: level of education as self-reported by this annotator (<code>none</code>, <code>basic</code>, <code>middle</code>, <code>high</code>)</li> <li><code>english_profile</code>: knowledge of English as self-reported by this annotator (<code>0</code> for no, <code>1</code> for yes)</li> <li><code>english_tested</code>: whether the annotator passed the Toloka language test for English (<code>0</code> for no, <code>1</code> for yes)</li> <li><code>banned</code>: whether the annotator was banned on this project (<code>0</code> for no, <code>1</code> for yes)</li> </ul> <p>The <code>*_rate</code> attributes should sum up to 1.</p> <p><strong>Edges</strong> are stored in the <a href="https://github.com/Toloka/TolokerGraph/blob/main/edges.tsv">edges.tsv</a> file in the TSV format of the following structure:</p> <ul> <li><code>source</code>: source identifier of the annotator</li> <li><code>target</code>: target identifier of the annotator</li> </ul> <p>As the graph is undirected, <code>source</code> and <code>target</code> can be interchanged for the given pair of nodes.</p>
Results of the crowd-mapping action within the project TeRRIFICA [Dataset No. 1 dated 2022-09-19]
<p>The dataset includes the results of the crowd-mapping action within the project "Territorial RRI fostering innovative climate action" - TeRRIFICA (Horizon 2020 under GA 824489) dated 2022-09-19. The data are points added to the map by the users (volunteers) and represent locations where climate change-related issues occur regarding air temperature, air quality, water, soil, and wind (SPOTS). The second part of the dataset is related to the crowd-mapping users and their anonymized characteristics (USERS). More details are available at https://terrifica.eu/.</p>
Audiovisual crowd counting dataset
<p>This dataset contains 1,935 annotated images, each image has one-second audio and a density map. For more details, please refer to our paper <a href="https://arxiv.org/abs/2005.07097">Ambient Sound Helps: Audiovisual Crowd Counting in Extreme Conditions</a> and <a href="https://github.com/qingzwang/AudioVisualCrowdCounting">code</a>.</p>
The Collection Management System Collection - Crowd-sourcing a list of digital repository options
<p><strong>The Collection Management System Collection - Crowd-sourcing a list of digital repository options</strong></p> <p>This dataset contains a list of digital repository options for collection management systems. It has been started and complited by Ashley Blewer.<br> The data set contains:</p> <ul> <li>a PDF capture of the blog describing motivation and background, columns of the spreadsheet and further resources; originally published at https://bits.ashleyblewer.com/blog/2017/08/09/collection-management-system-collection/</li> <li>The dataset / spreadsheet of The Collection Management System Collection, originally published at https://docs.google.com/spreadsheets/d/1cXOug3qM0pNNeD_wssiVEv9c0W1Y5I1VDTnSPTk7fb4/<br> The data was exported from the google spreadsheet on November 14th 2020 into the following formats: <ul> <li>PDF</li> <li>XLSX</li> <li>CSV</li> <li>TSV</li> </ul> </li> </ul> <p>The list contains basic information, administration considerations, interface considerations, technical considerations and social considerations for 70 different repository systems.</p>
Dataset - The influence of the crowding assumptions in biofilm simulations
<p><strong>Dataset simulated for the manuscript "The influence of the crowding assumptions in biofilm simulations" by Angeles-Martinez and Hatzimanikatis.</strong></p>
Dataset - Spatio-temporal modeling of the crowding conditions and metabolic variability in microbial communities
<p><strong>Dataset simulated for the manuscript "Spatio-temporal modeling of the crowding conditions and metabolic variability in microbial communities" by Angeles-Martinez and Hatzimanikatis.</strong></p>
Representation of crowd accidents in popular media
<p>This repository contains results related to the analysis of a corpus of news reports covering the topic of crowd accidents. To facilitate online visualization and offline analysis, the files are organized by assigning a number to each. The number system and the details of each set of files are described as follows:</p> <ul> <li><strong>Class 0</strong> – This contains the same files provided in this repository, but they are organized into folders to make analysis easier. If you intend to analyze the data from our lexical analysis, we suggest using this file since it is better organized and can be directly downloaded.</li> <li><strong>Class 1</strong> – This contains the sources and relevant information for people who are interested in replicating our dataset or accessing the news reports used in our analysis. Please note that due to copyright regulations, the texts cannot be shared. However, you can refer to the links provided in these files to access the news articles and Wikipedia pages. Some links have stopped working during the time we were working on this study, and others may be unreachable in the future.</li> <li><strong>Class 2</strong> – This contains the results from a lexical analysis of the corpus. The HTML page allows you to visualize each result interactively through the online VOSviewer app (you need to download the file and open it using a browser since Zenodo does not recognize this as a link). It is possible that this service (VOSviewer app) may be discontinued at some point in the future. PNG images of lexical maps are, therefore, available for download through the ZIP archive, although they do not allow interactive access. If you plan to read our results using the offline VOSviewer software or perform a more systematic analysis, JSON files are available for each category (time period, geographical area of the reporting institution, and purpose of gathering). The same files can be also find in the ZIP archive in class 0.</li> <li><strong>Class 3</strong> – These are the results of the sentiment analysis. For each report, a single result is generated for the title. However, for the body, the text is divided into parts, which are analyzed independently.</li> <li><strong>Class 4</strong> – These two files contains the corpus of Wikipedia relative to 68 crowd accidents which occurred between 1990 and 2019. The text for all accidents were scraped on October 15th, 2022 (<em>before</em> the tragedy in Itaewon) and on May 25th, 2023 (<em>after</em> the tragedy). Sources relative to the content in Wikipedia are listed in the file contained in Class 1 ("1_list_wiki_report.csv"). More generally, accidents listed on dedicated Wikipedia pages on <a href="https://en.wikipedia.org/wiki/List_of_fatal_crowd_crushes" target="_blank" rel="noopener">https://en.wikipedia.org/wiki/List_of_fatal_crowd_crushes</a> are reported in the corpus provided here (the period 1900-2019 is considered here).</li> </ul> <p>The format of CSV and JSON files should be self-explanatory after reading our publication. For specific questions or queries, please contact one of the authors, and we will try to assist you.</p>
GRN_MARVEL_AUDIO_VISUAL_CROWD_COUNTING
<p>The raw audio-video data was collected from Mgarr, a rural town on the western coast from IP cameras. Data has been manually annotated for pedestrians.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.