Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

130

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

130 results for “categorization”

Learn how ShareScore rates datasets ↗
OpenNeuro48/100

Go-nogo categorization and detection task

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
edi48/100

CoRRE Trait Data: A collection of 17 categorical and continuous traits for more than 4000 grassland species worldwide

In our changing world, it is critical to understand and predict plant community responses to global change drivers. Plant functional traits promise to be a key predictive tool for many ecosystems, including grasslands, however their use requires both complete plant community and functional trait data. Yet, representation of these data in global databases is incredibly sparse, particularly beyond a handful of most used traits and common species. Here we present the CoRRE Trait Database, spanning 17 traits (9 categorical, 8 continuous) anticipated to predict species’ responses to global change for 4,079 vascular plant species across 173 plant families present in 390 grassland experiments from around the world. The database contains complete categorical trait records for all 4,079 plant species, obtained from a comprehensive literature search. Additionally, the database contains nearly complete coverage (99.97%) of species mean values for continuous traits for a subset of 2,927 plant species, predicted from observed trait data drawn from TRY and a variety of other plant trait databases using Bayesian Probabilistic Matrix Factorization (BHPMF) and multivariate imputation using chained equations (MICE). These data will shed light on mechanisms underlying population, community, and ecosystem responses to global change in grasslands worldwide.

openCC BYMay 2024View details →
edi48/100

Categorical traits for macroalgae species of the Santa Barbara Channel

These data describe 32 categorical traits for 50 species of macroalgae found in the Santa Barbara Channel. Data are contained in one table, including a list of species, trait states for any given trait, and citations for where this information was found.

openCC (other)Aug 2025View details →
zenodo44/100

Investigating the Effects of Embodiment on Emotional Categorization of Faces and Words in Children and Adults

<p>The three data files uploaded here contain the data used for the analyses in experiments 1a, 1b, and 2 as described in the article carrying the same title as this dataset, published in the journal Frontiers in Psychology. All analyses were carried out in SPSS version 22 as described in the published article.</p> <p>Article Abstract:</p> <p>The facial feedback hypothesis (FFH) indicates that besides being involved in the production of facial expressions, the musculature of the face also influences one&rsquo;s perception of emotional stimuli. Recently, this effect has been the focus of increased scrutiny as efforts to replicate a key study with adult participants supporting this hypothesis, using the so-called &ldquo;pen-in-the-mouth&rdquo; task, have not been successful at several labs. Our series of experiments attempted to investigate whether the assumed embodiment effect can be reproduced in a simplified emotional categorization task for emotional faces and words. We also wanted to test whether the embodiment effect can be detected in children because it is assumed that their bodily processes are especially closely linked with their sensory and cognitive processes. Our experiments involved child and adult participants categorizing faces and words as positive or negative as quickly as possible, while inducing a positive or negative facial or bodily state (holding a straw in the mouth such that a smile or a frown was generated, or creating a positive or negative body posture). The positive or negative facial and bodily states could therefore be either congruent or incongruent with the valence of the target face and word stimuli. Our results did not show any significant differences between the congruent and incongruent conditions in either children or adults. This suggests that embodiment effects either do not significantly impact valence-based categorization or are not strong enough to be detected by our approach considering the sample size in the present study.</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

162 Human Error Descriptions and Categorizations from a User Study

<p><i><strong>Software Engineers' Human Errors</strong></i></p><p>This dataset contains descriptions of 162 human errors experienced by software engineering students during a user study described in the following publication:</p><ul><li>Benjamin S. Meyers and Andrew Meneely. Taxonomy-Based Human Error Assessment for Senior Software Engineering Students. Special Interest Group on Computer Science Education (SIGCSE) Technical Symposium. Forthcoming in 2024.</li></ul><p><i><strong>Included Files</strong></i></p><p>The "experienced_human_errors.csv" file contains a dataset of 162 human errors experienced during our user study. Participants documented their human errors in a Google Form with 8 questions.</p><p><i><strong>CSV Fields</strong></i></p><ul><li><strong>PARTICIPANT</strong>: Anonymous participant ID.</li><li><strong>INTERVIEW_DATE</strong>: Date of interview discussing human error.</li><li><strong>ID</strong>: Unique ID for experienced human error. Prefixed with "P1" for Phase 1 or "P2" for Phase 2.</li><li><strong>FINAL_CATEGORIZATION</strong>: Agreed upon T.H.E.S.E. categorization following discussion with interview facilitator.</li><li><strong>QUESTION_1</strong>: Anonymized participant answer to Question 1.</li><li><strong>QUESTION_2</strong>: Anonymized participant answer to Question 2.</li><li><strong>QUESTION_3</strong>: Anonymized participant answer to Question 3.</li><li><strong>QUESTION_4</strong>: Anonymized participant answer to Question 4.</li><li><strong>QUESTION_5</strong>: Anonymized participant answer to Question 5.</li><li><strong>QUESTION_6</strong>: Anonymized participant answer to Question 6.</li><li><strong>QUESTION_7</strong>: Anonymized participant answer to Question 7.</li><li><strong>QUESTION_8</strong>: Anonymized participant answer to Question 8.</li></ul><p><i><strong>Interview Questions</strong></i></p><ol><li>Please briefly describe the human error that you experienced.</li><li>If the human error you experienced resulted in a defect that was committed, please provide a link (or Git commit hash) to the commit below.</li><li>Is your human error a slip, lapse, or mistake?</li><li>Now, please examine the Taxonomy of Human Errors in Software Engineering (T.H.E.S.E.) and choose the specific human error that most accurately describes the human error you experienced. If you experienced multiple human errors, please submit this form once for each human error.</li><li>If there are other categories of human error that also describe the human error that you experienced, please note them here.</li><li>If you chose a 'General' or 'Other' category in Question (4), this question is required. Do you believe there is a missing human error category that better describes the human error that you experienced? If yes, please describe it below.</li><li>On a scale of 1 (not at all confident) to 5 (completely confident), how confident are you in your classification in the previous question?</li><li>Do you have any additional comments about this human error?</li></ol><p><i><strong>Anonymity</strong></i></p><p>Institutional Review Board approval for this research involving human subjects was granted by the Human Subjects Research Office at RIT on March 18, 2022. Participants signed an informed consent form acknowledging that (1) their participation was entirely voluntary and had no impact on their grades, and (2) their survey responses would be published in an anonymized format. All data released with this publication has been anonymized by replacing any personally identifiable information with participant identifiers.</p><p><i><strong>Contact</strong></i></p><p>Please contact Benjamin S. Meyers (<a href="mailto:bsm9339@rit.edu">email</a>) with questions about this data and its collection.</p><p><i><strong>Acknowledgments</strong></i></p><p>Collection of this data has been sponsored in part by the National Science Foundation (grant 1922169), by the NSA Science of Security Lablet program (grant H98230-17-D-0080/2018-0438-02), and by a Department of Defense DARPA SBIR program (grant 140D63-19-C-0018).</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

ERA5 dataset for the categorization of Stratospheric Final Warming

<p>This file in HDF5 format includes daily values of the following quantities derived from ERA5 data:</p> <ul> <li>Zonal wind at 60&deg;N &ndash; 10 hPa</li> <li>Polar temperature averaged over 80-90&deg;N and 50-10 hPa</li> <li>Amplitude of geopotential wave 1 at 60&deg;N &ndash; 10 hPa</li> <li>Zonal-mean meridional heat flux averaged over 45-75&deg;N at 10 hPa</li> </ul> <p>Each data set includes 25933 daily values from January 1n, 1950 to December 31, 2020.</p> <p>They are computed from ERA5 data extracted at 12UT each day and a resolution 2.5 x 2.5 degrees in latitude and longitude.</p> <p>The ERA5 data are provided by ECMWF Copernicus Climate Change Service from their data server <a href="https://confluence.ecmwf.int/display/CKB/How+to+download+ERA5">https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-pressure-levels?tab=form</a>.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Stack Exchange Open Source site questions categorization

<p>This dataset contains the posts of Open Source Stack Exchange site, collected at the end of 2020, along with the categorization of the posts. For each post a category, and potentially a second one is indicated, along with the cluster (generic group) each category belongs to. The coding task of assigning each question to a category was performed by two independent coders for each question (the categorization of each coder is also provided in the dataset). The dataset contains also (in a separate file) a dictionary of the most correlated unigrams and bigrams per category.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

From social categorization to implicit citizenship theories: Advancing the socio-cognitive foundations of state–citizen interactions

<p>Data for the following article: Vogel, R., Vogel, D., Liegat, M. C., &amp; Hensel, D. (2024). From social categorization to implicit citizenship theories: Advancing the socio‐cognitive foundations of state&ndash;citizen interactions. Public Administration Review, Article puar.13844. Advance online publication. https://doi.org/10.1111/puar.13844</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Raw data of the study: Categorizing urban avoiders, utilizers, and dwellers for identifying bird conservation priorities in a northern Andean city

<p>This datasheet contains raw data on bird count records made from 2016 and 2019. Data were taken in urban and adjacent non-urban areas of Medell&iacute;n, Colombia. It was part of a collaborative sampling effort during environmental assessments and personal research, summarizing systematic information on 139 sampling points (124 within the city and 15 in adjacent non-urban areas). All points were sampled under the same protocol in order to facilited data for research; in all cases, sampling was in charge of ornithologist with at least 4 years of previous experience in bird surveys. This protocol consisted in sampling during 10 minutes, four times per point (i.e., repetitions), using a fixed radius of 25 m.&nbsp;</p> <p>Information on bird surveys (Count_Data within the corresponding datasheet tab) contains the ID of each site; whether corresponded to a urban or non-urban site; in what category of urban development the site was located, based on 1000, 500 and 200 m buffers (from the observer during bird counts: moderate, low or high); the taxonomic information of each species (order, family, scientific name); the number of recorded individuals; &nbsp;the repetition or number of the visit (1, 2, 3, or 4); the name of the project; the name of the observer, and the date of sampling.&nbsp;</p> <p>Information on categorization of bird species (Categorization within the corresponding datasheet tab) represents additional information on altitudinal ranges, trophic guilds, distribution, and others. In addition, information on frequency for each bird species is given, according to the location of each sampling site and the way it was grouped. This information was the base for categorizing bird species as urban avoider, utilizer, or dweller, under the calculations and decision rules that are also given within the corresponding cells of the datasheet.</p> <p>Any further information or questions about this data could be ask directly, writing to the e-mails: jgarizabal@unal.edu.co or njmacer@unal.edu.co.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

H.I.D.R.A.: A Hierarchical, Interactive and Dynamic Recognition Architecture for Product Categorization

<p>The Hierarchical, Interactive and Dynamic Recognition Architecture (H.I.D.R.A.) for Product Categorization is a new intelligent system architecture developed by Elo7 to easily evolve its category tree and automatically classify millions of products, thus improving the page ranking of our marketplace.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

Categorical perception for red and brown

<p>This data supplements the article:</p> <p>Witzel, C., &amp; Gegenfurtner, K. R. (2016). Categorical perception for red and brown. Journal of Experimental Psychology: Human Perception &amp; Performance, 42(4), 540-570. doi:10.1037/xhp0000154</p> <p>The Excell-file with the data includes 3 sheets:</p> <p><strong>Sheet 1 (jnd): </strong>JND data from Figure 4.a of the above article.</p> <p>- columns = 20 test colours.</p> <p>- rows = 14 observers.</p> <p><strong>Sheet 2 (rt): </strong>Response time data from Figure 6.a of the above article.</p> <p>- columns = three kinds of colour pairs (AB, BC, &amp; CD) and the location of the target (left vs. right).</p> <p>- rows = 15 observers.</p> <p><strong>Sheet 3 (er): </strong>Error rates from Figure 6.b of the above article.</p> <p>- columns and rows as in sheet 2.</p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

Categorical facilitation with equally discriminable colors

<p>This data supplements the study of:</p> <p>Witzel, C., &amp; Gegenfurtner, K. R. (2015). Categorical facilitation with equally discriminable colors. Journal of Vision, 15(8), 22. doi:10.1167/15.8.22, http://jov.arvojournals.org/article.aspx?articleid=2381517</p> <p>The Excell-file provides the data shown in Figure 4 of the above article, which shows the main results. The first sheet (trained) provides the data for the first, experienced group of participants, the second sheet (naive) the data for the naive, untrained group of participants.</p> <p>Rows refer to the 20 stimulus pairs.</p> <p>Columns:</p> <p>sti_ctg = category membership of each colour in a pair.</p> <p>sti_type = type of colour pair: 1 = centre pair, 2 = boundary, 3 &amp; 4 = transitional pairs</p> <p>sti_azi = Hue (azimuth) in DKL-space</p> <p>rt = response times, one column for each observer</p> <p>er = error rates, one column for each observer</p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

Figure 1. The ecoregions are categorized within 14 in Terrestrial Ecoregions of the World: A New Map of Life on Earth

Figure 1. The ecoregions are categorized within 14 biomes and eight biogeographic realms to facilitate representation analyses.

opencc-by-4.0Oct 2001View details →
zenodo40/100

GeoEDdA: A Gold Standard Dataset for Named Entity Recognition and Span Categorization Annotations of Diderot & d'Alembert's Encyclopédie

<p>This repository contains a gold standard dataset for named entity recognition and span categorization annotations from Diderot &amp; d&rsquo;Alembert&rsquo;s Encyclop&eacute;die entries.</p> <p>The dataset is available in the following formats:</p> <ul> <li>JSONL format provided by <a href="https://prodi.gy/" rel="nofollow">Prodigy</a></li> <li>binary spaCy format (ready to use with the spaCy train pipeline)</li> </ul> <p>The Gold Standard dataset is composed of 2,200 paragraphs out of 2,001 Encyclop&eacute;die's entries randomly selected. All paragraphs were written in 19th-century French.</p> <p>The spans/entities were labeled by the project team along with using pre-labelling with early machine learning models to speed up the labelling process. A train/val/test split was used. Validation and test sets are composed of 200 paragraphs each: 100 classified under 'G&eacute;ographie' and 100 from another knowledge domain. The datasets have the following breakdown of tokens and spans/entities.</p> <h2>Tagset</h2> <ul> <li><strong>NC-Spatial</strong>: a common noun that identifies a spatial entity (nominal spatial entity) including natural features, e.g. <code>ville</code>,&nbsp;<code>la rivi&egrave;re</code>, <code>royaume</code>.</li> <li><strong>NP-Spatial</strong>: a proper noun identifying the name of a place (spatial named entities), e.g. <code>France</code>, <code>Paris</code>, <code>la Chine</code>.</li> <li><strong>ENE-Spatial</strong>: nested spatial entity , e.g. <code>ville de France</code> , <code>royaume de Naples</code>, <code>la mer Baltique</code>.</li> <li><strong>Relation</strong>: spatial relation, e.g. <code>dans</code>, <code>sur</code>, <code>&agrave; 10 lieues de</code>.</li> <li><strong>Latlong</strong>: geographic coordinates, e.g. <code>Long. 19. 49. lat. 43. 55. 44.</code></li> <li><strong>NC-Person</strong>: a common noun that identifies a person (nominal spatial entity), e.g. <code>roi</code>, <code>l'empereur</code>, <code>les auteurs</code>.</li> <li><strong>NP-Person</strong>: a proper noun identifying the name of a person (person named entities), e.g. <code>Louis XIV</code>, <code>Pline</code>.</li> <li><strong>ENE-Person</strong>: nested people entity, e.g. <code>le czar Pierre</code>, <code>roi de Mac&eacute;doine</code>.</li> <li><strong>NP-Misc</strong>: a proper noun identifying entities not classified as spatial or person, e.g. <code>l'Eglise</code>, <code>1702</code>, <code>P&eacute;lasgique</code></li> <li><strong>ENE-Misc</strong>: nested named entity not classified as spatial or person, e.g. <code>l'ordre de S. Jacques</code>, <code>la d&eacute;claration du 21 Mars 1671</code>.</li> <li><strong>Head</strong>: entry name</li> <li><strong>Domain-Mark</strong>: words indicating the knowledge domain (usually after the head and between parenthesis), e.g. <code>G&eacute;ographie</code>, <code>Geog.</code>, <code>en Anatomie</code>.</li> </ul> <h2>HuggingFace</h2> <p>The GeoEDdA dataset is available on the HuggingFace Hub: <a href="https://huggingface.co/datasets/GEODE/GeoEDdA">https://huggingface.co/datasets/GEODE/GeoEDdA</a></p> <h2>spaCy Custom Spancat trained on Diderot &amp; d&rsquo;Alembert&rsquo;s Encyclop&eacute;die entries</h2> <p>This dataset was used to train and evaluate a custom spancat model for French using <a href="https://spacy.io/" rel="nofollow">spaCy</a>. The model is available on HuggingFace's model hub: <a href="https://huggingface.co/GEODE/fr_spacy_custom_spancat_edda" rel="nofollow">https://huggingface.co/GEODE/fr_spacy_custom_spancat_edda</a>.</p> <h2>Acknowledgement</h2> <p>The authors are grateful to the <a href="https://aslan.universite-lyon.fr/" rel="nofollow">ASLAN project</a> (ANR-10-LABX-0081) of the Universit&eacute; de Lyon, for its financial support within the French program "Investments for the Future" operated by the National Research Agency (ANR). Data courtesy the <a href="https://artfl-project.uchicago.edu/" rel="nofollow">ARTFL Encyclop&eacute;die Project</a>, University of Chicago.</p>

opencc-by-sa-4.0Jan 2024View details →
zenodo40/100

Categorization of countries into Global South and Global North

<p>Structured data categorizing countries into two blocks: Global South and Global North. This categorization took into account factors such as HDI, colonization, and dependency. In the specific case of this categorization, countries of the Global North were understood as those with high economic income.</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Categorical variables based on cross country household survey on energy consumption

<p>The data used in this file was collected via two large-scale surveys conducted in Italy, Switzerland and the Netherlands. A total of 6,138 responses were recorded, containing information on socio-demographic and socio-psychological characteristics, dwelling and household characteristics, technologies and energy services used, and their metered electricity consumption. There were a large number of missing responses for metered electricity consumption in the Netherlands, leading to an under-representation of data from this country. The survey responses were used to construct newly defined energy efficiency indicators, and energy service indicators. This allows two distinct factors to be separated: service consumption, and energy efficiency relative to the demanded service. Firstly, dwelling characteristics and survey responses related to energy services (e.g. floorspace, ownership of specific appliances and number of lightbulbs), were regressed to the collected metered electricity data. For each household, this allowed us to calculate the expected lighting and appliance electricity demand based on the level service that the household demanded, which is referred to as&nbsp;lighting and appliance service demand indicators. The idea is that a larger house, or a house with more appliances for example is expected to use more electricity. Relative to this expected electricity demand energy efficiency can be calculated. All variables are&nbsp;categorised in categorical variables deducted based on the questions asked in the two surveys.&nbsp;The survey responses were clustered based on the lighting service demand, appliance service demand and the efficiency gap (k-means clustering with Jaccard dissimilarity measure) which is described in Edelenbosch, Miu et al (2022).&nbsp;Translating observed household energy behaviour to agent-based technology choices in an integrated modelling framework. <em>Iscience</em> (accepted).</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Accompanying simulated data for "Go multivariate: a Monte Carlo study of a multilevel hidden Markov model with categorical data of varying complexity"

<p>The multilevel hidden Markov model (MHMM) is a promising vehicle to investigate latent dynamics over time in social and behavioral processes. By including continuous individual random effects, the model accommodates variability between individuals, providing individual-specific trajectories and facilitating the study of individual differences. However, the performance of the MHMM has not been sufficiently explored. Currently, there are no practical guidelines on the sample size needed to obtain reliable estimates related to categorical data characteristics We performed an extensive simulation to assess the effect of the number of dependent variables (1-4), the number of individuals (5-90), and the number of observations per individual (100-1600) on the estimation performance of group-level parameters and between-individual variability on a Bayesian MHMM with categorical data of various levels of complexity. We found that using multivariate data generally alleviates the sample size needed and improves the stability of the results. Regarding the estimation of group-level parameters, the number of individuals and observations largely compensate for each other. Meanwhile, only the former drives the estimation of between-individual variability. We conclude with guidelines on the sample size necessary based on the complexity of the data and the study objectives of the practitioners.</p> <p>This repository contains data generated&nbsp;for the manuscript: &quot;Go multivariate: a Monte Carlo study of a multilevel hidden Markov model&nbsp;with categorical data of varying complexity&quot;. It comprehends: (1) model outputs (maximum a posteriori estimates) for&nbsp;each repetition (n=100) of&nbsp;each scenario (n=324) of the main simulation, (2) complete model outputs (including estimates for&nbsp;4000 MCMC iterations) for two chains of each&nbsp;repetition (n=3)&nbsp;of&nbsp;each scenario (n=324). Please note that the empirical data used in the manuscript&nbsp;is not available as part of this repository.&nbsp;A subsample of the data used in the empirical example are openly available as an example data set in the R package <a href="https://cran.r-project.org/web/packages/mHMMbayes/index.html">mHMMbayes on CRAN</a>. The full data set&nbsp;is available on request from the authors.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Online Repository of the Study "I want to RIDE my e-bicycle!": Supporting Developers Categorizing User Issues of a Mobility-as-a-Service Platform

<p><strong>Online Repository of the Study </strong><em>&ldquo;I want to RIDE my e-bicycle!&quot;: Supporting Developers Categorizing User Issues of a Mobility-as-a-Service Platform</em></p> <p><strong>Introduction</strong></p> <p>In the Mobility-as-a-Service (MaaS) context, e-bikes are important and environmental-friendly transportation resources providing flexibility, time and cost savings, and reducing traffic congestion. Additional to user satisfaction and marketing advantages, the resolution of user-reported issues is regulated in many cities. In order to efficiently solve the issues, it is essential to quickly identify their types (e.g., software- or hardware-related?) to assign them to the responsible team. But for popular e-mobility services, the manual analysis of the reports is inefficient because of its tediousness, high time requirements, and error-proneness.&nbsp;</p> <p>Our empirical study, carried out in the context of a <em>Mobility as a Service </em>start-up company, proposes an approach for the automated identification of relevant concerns reported by users of e-bike services. The company has more than 20,000 private customers across seven different countries and dedicates considerable effort in analyzing user behavior. However, the current manual process of analyzing and triaging user-reported issues hinders MaaS-company&rsquo;s ability to grow and expand its services.&nbsp;</p> <p>To help MaaS providers identify relevant user-reported issues, In the study, we (i) manually inspect about 3,000 user-reported issues received by the MaaS company; (ii) design a taxonomy modeling the types of relevant issues reported by users; and (iii) propose MaaS-RIDE, an approach to automatically classify the user-reported issues according to the categories of the devised taxonomy.&nbsp;</p> <p>Our results demonstrate that MaaS-RIDE is able to accurately (F-measure &ge; 93%) identify software and hardware user-reported issues. This result is critical for e-bike sharing companies to address such issues in an agile way and achieve the required user satisfaction.</p> <p><strong>Dataset Overview</strong></p> <p>The dataset is composed of the following different sorts of data:&nbsp;</p> <ul> <li>&nbsp;&ldquo;<em>Data_and_preprocessing</em>&rdquo; folder&nbsp; <ul> <li>o the user-reported issues data</li> <li>o the user-reported issues data processed as Bag of Words for Machine Learning training.&nbsp; <ul> <li>For this look at the sub-folder &ldquo;<em>input_data_for_ML</em>&rdquo; and the following matrices: <ul> <li><em>tf-idf-matrix-of-comment_finals_with_oracle_info_low_level.csv</em></li> <li><em>tf-idf-matrix-of-comment_finals_with_oracle_info.csv</em></li> </ul> </li> <li>Moreover, a sample of selected issues was reported in the replication package: <ul> <li>see file &ldquo;<em>randomSamples.csv</em>&rdquo; (due to a non-disclosure agreement with our industrial partner, we are unauthorized to share the whole raw user reports used in our experiments)</li> <li>&nbsp;&ldquo;RQ1&rdquo; folder: Types of E-bikes User-reported Issues</li> </ul> </li> </ul> </li> <li>&nbsp;the resulting taxonomy after the analysis of the issues</li> <li>&nbsp;&ldquo;RQ2&rdquo; folder: Classifying E-bikes Issue types</li> <li>&nbsp;the trained models&nbsp;</li> <li>&nbsp;the results of the models</li> </ul> </li> </ul> <p>The following sections describe more in detail what each of those folders and files contain.</p> <p><strong>&ldquo;Data_and_preprocessing&rdquo; folder</strong></p> <ul> <li><strong>User-reported issues subset.</strong></li> </ul> <p>In an industrial setting, due to privacy reasons, we disclose only an example subset of the user-reported issues, this information is in the file <em>randomSamples.csv</em>.</p> <p>The <em>randomSamples.csv </em>a subset that was generated randomly adding 20 examples using a stratified sampling from the High-level categories and 20 from the Low-level categories. This subset is not exhaustive but serves the purpose of showing the reviewers the kind of issues that this particular industrial set is confronted with. The file contains:</p> <ul> <li> <ul> <li>&nbsp;the Id of the user report;&nbsp;</li> <li>&nbsp;the column &quot;comment_final&quot;<strong> </strong>contains the issue text after the replacement of information that needed anonymization (e.g., vehicle-plates, personal names, addresses and timestamps);&nbsp;</li> <li>&nbsp;the column &quot;High_level_category&quot; contains the selected category from the 5 first level categories of the presented <em>Three-level taxonomy of e-bike user reported issues</em>;&nbsp;</li> <li>&bull; the columns &lsquo;Low_level_category&quot; and &quot;Fine_grained_topic&quot; contain the assigned, if existing, respective category.&nbsp;</li> </ul> </li> <li><strong>Bag of Words Term by Document matrix.</strong></li> </ul> <p>An important input for training the ML models is the Bag of Words representation generated after processing the&nbsp; 2,989 manually-labeled user issues. The result of this process is a Term-by-Document matrix. We share this matrix in the files in the sub-folder <em>input_data_for_ML </em>where they are labeled for High- and Low-level categories.&nbsp;</p> <p>In the <em>tf-idf-matrix-of-comment_finals_with_oracle_info.csv</em> and <em>tf-idf-matrix-of-comment_finals_with_oracle_info_low_level.csv</em> files, the first column refers to the issue &ldquo;Id&rdquo;, the last column &ldquo;oracle&rdquo; is the labeled category, the rest of the columns represent the terms contained in the 2,989 user-reported issues and in each row the weight of the i&minus;𝑡ℎ term contained in the j&minus;𝑡ℎ user issue by using the tf-idf score.</p> <p><strong>&ldquo;RQ1&rdquo; folder</strong></p> <ul> <li><strong>&ldquo;Three-level taxonomy of e-bike user-reported issues.pdf<em>&rdquo; file</em></strong></li> </ul> <p>The taxonomy derives from the manual analysis of the 2,989 user issues. We found that a three-level taxonomy provides significant granularity to the MaaS-company. The taxonomy encompasses 5 High-level categories, 16 Low-level categories, and 15 Low-level subcategories of e-bike user-reported issues. The file <em>Three-level taxonomy of e-bike user-reported issues.pdf</em> &nbsp;presents the taxonomy categories and in the columns &ldquo;Nr.&rdquo; and &ldquo;%&rdquo; it shows the number of occurrences within the analyzed dataset, and the corresponding percentages.</p> <p><strong>&ldquo;RQ2&rdquo; folder</strong></p> <ul> <li><strong>&ldquo;Trained Models&rdquo; folder</strong></li> </ul> <p>We provide the trained machine and deep learning models in the sub-folder <em>ML_DL_models</em>. Our approach experimented with classic machine learning models based on the Bag-of-Words approach using SVM, on Word Embeddings using FastText, and Language models leveraging BERT. The SVM and BERT models were trained using the open source low-code data analytics platform KNIME and were used to classify issues corresponding to the first and second levels of the taxonomy from the &ldquo;RQ1&rdquo; folder. A 10-fold cross validation strategy was used to assess the classification performance.&nbsp;&nbsp;</p> <p>The fastText model was trained by using default values of parameters (https://fasttext.cc/docs/en/options.html) and a 10-fold cross-validation strategy. With fastText, we classified issues corresponding only to the first level of the taxonomy from &ldquo;RQ1&rdquo; folder, since fastText is more effective when more data points are available in the training set (i.e., lower levels in the taxonomy have fewer well-represented issue types).</p> <ul> <li><strong>&ldquo;Model results&rdquo; folder</strong></li> </ul> <p>In the sub-folder model_results we provide the tables summarizing the results of using the proposed MaaS-RIDE approach, with which we automatically identify and categorize user-reported issues according to the High-level and Low-level categories of the taxonomy devised in RQ1, which are relevant for the MaaS-company.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Superordinate Categorization Based on the Perceptual Organization of Parts - Dataset

<p>This is the dataset for:<br> Tiedemann, H.; Schmidt, F.; Fleming, R.W. Superordinate Categorization Based on the Perceptual Organization of Parts. Brain Sci. 2022, 12, 667. https://doi.org/10.3390/brainsci1205066</p> <p>Available at:<br> https://www.mdpi.com/2076-3425/12/5/667</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Scripts and analysis files for categorization of PKZILLA matching proteomic peptides into protein-unique, protein-multimatch & exon-unique, exon-multimatch categories.

<p>A .zip file containing the source data files &amp; Jupyter notebook for analysis of the <em>Prymnesium parvum</em> 12B1 PKZILLA-detecting proteomic results (<a href="https://doi.org/10.5281/zenodo.10023441">https://doi.org/10.5281/zenodo.10023441</a>), and the resulting files from the workflow. See "Analysis of proteomic results" section of the manuscript Materials and Methods for further detail.&nbsp;</p> <p><strong>Key files:</strong></p> <ul> <li>'PKZILLA-1_classify_peptides.txt' - A plaintext report of the # of classified peptides for PKZILLA-1</li> <li>'PKZILLA-2_classify_peptides.txt' - A plaintext report of the # of classified peptides for PKZILLA-2</li> <li>'./hierarchical_classified_xlsx/' - Excel spreadsheets with the classified peptides for PKZILLA-1 and PKZILLA-2</li> <li>'./Process_into_polypeptide_coordinates/' - Workflow, results, and plots for back-alignment of peptides back to PKZILLA-1 and PKZILLA-2 genomic loci</li> </ul>

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record