Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,359

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,359 results for “Online”

Learn how ShareScore rates datasets ↗
zenodo40/100

The Reddit Politosphere: A Large-Scale Text and Network Resource of Online Political Discourse

<p>The Reddit Politosphere is a large-scale resource of online political discourse covering more than 600 political discussion groups over a period of 12 years. Based on the <a href="https://doi.org/10.5281/zenodo.3608135">Pushshift Reddit Dataset</a>, it is to the best of our knowledge the largest and ideologically most comprehensive dataset of its type now available. One key feature of the Reddit Politosphere is that it consists of both text and network data. We also release annotated metadata for subreddits and users.</p> <p>Documentation and scripts for easy data access are provided in an associated&nbsp;<a href="https://github.com/valentinhofmann/politosphere">repository</a>&nbsp;on GitHub.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Dataset: Sentiment Analysis annotation of News headlines covering the Olympic legacy of Rio 2016 and London 2012 published by the Brazilian and British online media

<p>Dataset of 464 news headlines with sentiment manually annotated by a domain expert using the labels positive, negative and neutral. Data contains URLs for news articles published between 2004-2020 by the British and Brazilian media in English and Brazilian Portuguese covering the Olympic legacies of London 2012 and Rio 2016. Articles were collected from the news outlets&rsquo; websites using Google search engine.</p> <p>News outlets:</p> <ul> <li>The Guardian</li> <li>Daily Mail</li> <li>Globo</li> <li>Estadao</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Replication files for "Comparing Online and Offline Political Support".

<p>The Zip-folder contains replication materials for the article&nbsp;&quot;Comparing Online and Offline Political Support&quot;. The folder contains a readme with information and details on the replication data and scripts.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Combining Policies to Reduce the Spread of Viral Misinformation Online

<p>Data were collected as part of the Election Integrity Partnership. Instances of potential misinformation were flagged as tickets. These were reviewed and categorized as misinformation if they made false claims related to election integrity. Details on the collection methods of the EIP can be found in our Report. These tickets were grouped together into qualitatively similar incidents. For example, tickets regarding false narratives about the use of Benford&#39;s law to detect fraud in Wisconsin became an incident. For each incident, search terms and appropriate data-ranges were determined to query our database.&nbsp;</p> <p>Our full database consisted of all tweets matching an evolving set of keywords, collected in real time, using the Twitter API. To maintain user privacy, we are providing data segmented into events and aggregated into 5-minute blocks of time. This should be sufficient for replicating our findings (predicated on the aggregation and segmentation). In order to permit analysis under various user-removal conditions, we have provided multiple versions of this dataset with users removed according to the conditions evaluated in the manuscript. We encourage anyone with the need for more granular data or alternate conditions to reach out to the University of Washington Center for an Informed Public.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Using a Hybrid Kano-Importance Questionnaire in the Acquisition of Data Related to Students' Expectations from Online Educational Platforms

<p>This dataset contains the data collected for the assessment of the quality attributes of a new online educational platform. The questionnaire used for data collection the Kano methodology and was designed as a hybrid Kano-importance questionnaire. The purpose of this data collection consists of the analysis of the students&rsquo; expectations regarding the features proposed for a new online educational platform. This analysis facilitates the identification of student needs during times of COVID-19 pandemic and post-pandemic times, while a transition to an online educational system was used throughout the world.&nbsp;</p>

opencc-by-4.0May 2022View details →
dryad40/100

Online electronic material for: Macroevolutionary dynamics of climatic niche space

<p><span>How and why lineages evolve along niche space as they diversify and adapt to different environments is fundamental to evolution. Progress has been hampered by the difficulties of linking a robust empirical characterization of species niches with flexible evolutionary models that describe their evolution. Consequently, the relative influence of abiotic and biotic factors remains poorly understood. Here we characterize species' two-dimensional temperature and precipitation niche space occupied (i.e., species niche envelope) as complex geometries and assess their evolution across all Aves using a model that captures heterogeneous evolutionary rates on time-calibrated phylogenies. We find that extant birds coevolved from warm, mesic climatic niches into colder and drier environments and responded to the K-Pg boundary with a dramatic increase in disparity. Contrary to expectations of subsiding rates of niche evolution, our results show that overall rates have increased steadily, with some lineages experiencing exceptionally high evolutionary rates, associated with colonization of novel niche spaces, and others showing niche stasis. Both competition- and environmental change-driven niche evolution transpire and result in highly heterogeneous rates near the present. Our findings highlight the growing ecological and conservation insights arising from model-based integration of comprehensive environmental and phylogenetic information.</span></p>

opencc-zeroMay 2022View details →
zenodo40/100

Online Repository of the Study "I want to RIDE my e-bicycle!": Supporting Developers Categorizing User Issues of a Mobility-as-a-Service Platform

<p><strong>Online Repository of the Study </strong><em>&ldquo;I want to RIDE my e-bicycle!&quot;: Supporting Developers Categorizing User Issues of a Mobility-as-a-Service Platform</em></p> <p><strong>Introduction</strong></p> <p>In the Mobility-as-a-Service (MaaS) context, e-bikes are important and environmental-friendly transportation resources providing flexibility, time and cost savings, and reducing traffic congestion. Additional to user satisfaction and marketing advantages, the resolution of user-reported issues is regulated in many cities. In order to efficiently solve the issues, it is essential to quickly identify their types (e.g., software- or hardware-related?) to assign them to the responsible team. But for popular e-mobility services, the manual analysis of the reports is inefficient because of its tediousness, high time requirements, and error-proneness.&nbsp;</p> <p>Our empirical study, carried out in the context of a <em>Mobility as a Service </em>start-up company, proposes an approach for the automated identification of relevant concerns reported by users of e-bike services. The company has more than 20,000 private customers across seven different countries and dedicates considerable effort in analyzing user behavior. However, the current manual process of analyzing and triaging user-reported issues hinders MaaS-company&rsquo;s ability to grow and expand its services.&nbsp;</p> <p>To help MaaS providers identify relevant user-reported issues, In the study, we (i) manually inspect about 3,000 user-reported issues received by the MaaS company; (ii) design a taxonomy modeling the types of relevant issues reported by users; and (iii) propose MaaS-RIDE, an approach to automatically classify the user-reported issues according to the categories of the devised taxonomy.&nbsp;</p> <p>Our results demonstrate that MaaS-RIDE is able to accurately (F-measure &ge; 93%) identify software and hardware user-reported issues. This result is critical for e-bike sharing companies to address such issues in an agile way and achieve the required user satisfaction.</p> <p><strong>Dataset Overview</strong></p> <p>The dataset is composed of the following different sorts of data:&nbsp;</p> <ul> <li>&nbsp;&ldquo;<em>Data_and_preprocessing</em>&rdquo; folder&nbsp; <ul> <li>o the user-reported issues data</li> <li>o the user-reported issues data processed as Bag of Words for Machine Learning training.&nbsp; <ul> <li>For this look at the sub-folder &ldquo;<em>input_data_for_ML</em>&rdquo; and the following matrices: <ul> <li><em>tf-idf-matrix-of-comment_finals_with_oracle_info_low_level.csv</em></li> <li><em>tf-idf-matrix-of-comment_finals_with_oracle_info.csv</em></li> </ul> </li> <li>Moreover, a sample of selected issues was reported in the replication package: <ul> <li>see file &ldquo;<em>randomSamples.csv</em>&rdquo; (due to a non-disclosure agreement with our industrial partner, we are unauthorized to share the whole raw user reports used in our experiments)</li> <li>&nbsp;&ldquo;RQ1&rdquo; folder: Types of E-bikes User-reported Issues</li> </ul> </li> </ul> </li> <li>&nbsp;the resulting taxonomy after the analysis of the issues</li> <li>&nbsp;&ldquo;RQ2&rdquo; folder: Classifying E-bikes Issue types</li> <li>&nbsp;the trained models&nbsp;</li> <li>&nbsp;the results of the models</li> </ul> </li> </ul> <p>The following sections describe more in detail what each of those folders and files contain.</p> <p><strong>&ldquo;Data_and_preprocessing&rdquo; folder</strong></p> <ul> <li><strong>User-reported issues subset.</strong></li> </ul> <p>In an industrial setting, due to privacy reasons, we disclose only an example subset of the user-reported issues, this information is in the file <em>randomSamples.csv</em>.</p> <p>The <em>randomSamples.csv </em>a subset that was generated randomly adding 20 examples using a stratified sampling from the High-level categories and 20 from the Low-level categories. This subset is not exhaustive but serves the purpose of showing the reviewers the kind of issues that this particular industrial set is confronted with. The file contains:</p> <ul> <li> <ul> <li>&nbsp;the Id of the user report;&nbsp;</li> <li>&nbsp;the column &quot;comment_final&quot;<strong> </strong>contains the issue text after the replacement of information that needed anonymization (e.g., vehicle-plates, personal names, addresses and timestamps);&nbsp;</li> <li>&nbsp;the column &quot;High_level_category&quot; contains the selected category from the 5 first level categories of the presented <em>Three-level taxonomy of e-bike user reported issues</em>;&nbsp;</li> <li>&bull; the columns &lsquo;Low_level_category&quot; and &quot;Fine_grained_topic&quot; contain the assigned, if existing, respective category.&nbsp;</li> </ul> </li> <li><strong>Bag of Words Term by Document matrix.</strong></li> </ul> <p>An important input for training the ML models is the Bag of Words representation generated after processing the&nbsp; 2,989 manually-labeled user issues. The result of this process is a Term-by-Document matrix. We share this matrix in the files in the sub-folder <em>input_data_for_ML </em>where they are labeled for High- and Low-level categories.&nbsp;</p> <p>In the <em>tf-idf-matrix-of-comment_finals_with_oracle_info.csv</em> and <em>tf-idf-matrix-of-comment_finals_with_oracle_info_low_level.csv</em> files, the first column refers to the issue &ldquo;Id&rdquo;, the last column &ldquo;oracle&rdquo; is the labeled category, the rest of the columns represent the terms contained in the 2,989 user-reported issues and in each row the weight of the i&minus;𝑡ℎ term contained in the j&minus;𝑡ℎ user issue by using the tf-idf score.</p> <p><strong>&ldquo;RQ1&rdquo; folder</strong></p> <ul> <li><strong>&ldquo;Three-level taxonomy of e-bike user-reported issues.pdf<em>&rdquo; file</em></strong></li> </ul> <p>The taxonomy derives from the manual analysis of the 2,989 user issues. We found that a three-level taxonomy provides significant granularity to the MaaS-company. The taxonomy encompasses 5 High-level categories, 16 Low-level categories, and 15 Low-level subcategories of e-bike user-reported issues. The file <em>Three-level taxonomy of e-bike user-reported issues.pdf</em> &nbsp;presents the taxonomy categories and in the columns &ldquo;Nr.&rdquo; and &ldquo;%&rdquo; it shows the number of occurrences within the analyzed dataset, and the corresponding percentages.</p> <p><strong>&ldquo;RQ2&rdquo; folder</strong></p> <ul> <li><strong>&ldquo;Trained Models&rdquo; folder</strong></li> </ul> <p>We provide the trained machine and deep learning models in the sub-folder <em>ML_DL_models</em>. Our approach experimented with classic machine learning models based on the Bag-of-Words approach using SVM, on Word Embeddings using FastText, and Language models leveraging BERT. The SVM and BERT models were trained using the open source low-code data analytics platform KNIME and were used to classify issues corresponding to the first and second levels of the taxonomy from the &ldquo;RQ1&rdquo; folder. A 10-fold cross validation strategy was used to assess the classification performance.&nbsp;&nbsp;</p> <p>The fastText model was trained by using default values of parameters (https://fasttext.cc/docs/en/options.html) and a 10-fold cross-validation strategy. With fastText, we classified issues corresponding only to the first level of the taxonomy from &ldquo;RQ1&rdquo; folder, since fastText is more effective when more data points are available in the training set (i.e., lower levels in the taxonomy have fewer well-represented issue types).</p> <ul> <li><strong>&ldquo;Model results&rdquo; folder</strong></li> </ul> <p>In the sub-folder model_results we provide the tables summarizing the results of using the proposed MaaS-RIDE approach, with which we automatically identify and categorize user-reported issues according to the High-level and Low-level categories of the taxonomy devised in RQ1, which are relevant for the MaaS-company.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Online Supplement for An Arabic Version of The Visual Aesthetics of Websites Inventory (AR-VisAWI): Translation and Psychometric Properties

<p>This is the online supplement for a translation of the Visual Aesthetics of Websites Inventory into Arabic (AR-VisAWI).&nbsp;</p> <p>In the field of human-computer interaction, the concept of visual aesthetics gained popularity after researchers started to recognize its merits and effects on user experience. Yet, no proper instrument exists to assess the visual aesthetics of websites that are intended for users who speak Arabic as their native language. As such, the aim of this study was to develop and evaluate an Arabic version of the Visual Aesthetics of Websites Inventory (VisAWI, Moshagen &amp; Thielsch, 2010) and its short version (VisAWI-S, Moshagen &amp; Thielsch, 2013). For this purpose, participants were asked to evaluate a randomly assigned website with the AR-VisAWI and with different validating instruments. A final sample of 223 participants was included in the analyses.</p> <p>This online supplement includes</p> <ul> <li>a codebook describing all instructions and items</li> <li>raw data (anonymised) and analysis script (Note: The raw data contains only the information of persons who have agreed to be included in the analysis. Some demographic information was deleted to ensure anonymity.)</li> <li>Questionnaire template and scoring instructions</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Continuous Integration and Delivery Practices for Cyber-Physical Systems: An Interview-Based Study - Online Dataset

<p>This package contains the online dataset of the manuscript:</p> <p>Continuous Integration and Delivery Practices for Cyber-Physical Systems: An Interview-Based Study</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

WP2 DESIRA_Online survey_Data

<p>The online survey complements the key-informants interviews and participatory workshop (or focus group discussions) organized in the living labs (LL) to provide a thorough and holistic socio-economic impact assessment to co-create knowledge on possible impact and shared transition pathways toward digitalization. The survey was designed jointly by UNIPI and KIT-ITAS and addresses all stakeholders involved in the LL.</p>

opencc-by-4.0Dec 2021View details →
dryad40/100

Paid and hypothetical time preferences are the same: Lab, field and online evidence

<p class="MsoNormal"><span>The use of real decision-making incentives remains under debate after decades of economic experiments. In time preferences experiments involving future payments, real incentives are particularly problematic due to between-options differences in transaction costs, among other issues. What if hypothetical payments provide accurate data which, moreover, avoid transaction cost problems? In this paper, we test whether the use of hypothetical or one-out-of-ten-participants probabilistic—versus real—payments affects the elicitation of short-term and long-term discounting in a standard multiple price list task. We analyze data from a lab experiment in Spain and well-powered field and online experiments in Nigeria and the UK, respectively (N = 2,043). Our results indicate that the preferences elicited using the three payment methods are mostly the same: we can reject that either hypothetical or one-out-of-ten payments change any of the four preference measures considered by more than 0.18 SD with respect to real payments.</span></p>

opencc-zeroAug 2022View details →
zenodo40/100

Figure 1 in The Tydeoidea (Ereynetidae, Iolinidae, Triophtydeidae and Tydeidae) - An online database in the Wikispecies platform

Figure 1 Diachronic classification of Tydeoidea. Abbreviations: pcp = post-cunliffean period, Pseudot. = Pseudotydeinae, R = Riccardoellinae, s = synonymy, subfam. = subfamilies, Trioph. = Triophtydeidae.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 4 A-B in The Tydeoidea (Ereynetidae, Iolinidae, Triophtydeidae and Tydeidae) - An online database in the Wikispecies platform

Figure 4 A-B – Tetranychus urticae; C – T. viburni; D – Tydeus goetzi; A – Dissecting microscope view; B-C – Facsimile of Koch's figures (same magnification) with some dorsal setae notation added; D – Compound microscope view, Agroscope Changins [Switzerland], routine black chlorazol coloration by Marc Baillod, scale bar = 100 µm. Koch's "Schulterborsten" correspond to scapular setaesc(1 andsc2) plus the subhumeral seta (c3). A – photoghraphy by Gilles San Martin. CC-BY.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Figure 3 in The Tydeoidea (Ereynetidae, Iolinidae, Triophtydeidae and Tydeidae) - An online database in the Wikispecies platform

Figure 3 The number of ereynetid mites described by Fain and by other acarologists (data grouped by decade).

opencc-by-4.0Dec 2021View details →
zenodo40/100

Data for: Image-based evaluation of beers at an online Pint of Science festival using Projective Mapping, Check-All-That-Apply and Acceptability

<p>Data obtained from&nbsp;n=67 untrained attendants at an outreach Pint of Science festival, online because of the COVID-19 pandemic but usually held at bars. The participants&nbsp;used images of brand logos to evaluate eight beers among the most commonly consumed in Spain. Three sensory analysis techniques were used: Projective Mapping, Acceptability and Check-All-That-Apply (CATA).</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Figure 2. The structures of the two tables from the Dex Online database-ADX – Agent for Morphologic Analysis of Lexical Entries in a Dictionary

<p>Dex Online is a project initiated and coordinated by Catalin Francu [3]. He intended to<br> realise an online database for all the words in the Romanian language, using the main explanatory<br> dictionaries, dictionaries of synonyms, neologisms, published by the Romanian Academy and other<br> scientific forums.<br> The database was completed by volunteers, similarly to the Wikipedia system. They actually<br> transcribed the information from different important dictionaries, but many words have been<br> electronically entered by two companies (Siveco and Litera International Publishing House).</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

Survey for online registered users of HistoricGraves platform

<div>This survey is being conducted by Eachtra Archaeological Projects as part of INCULTUM (2021-2024), a tourism-oriented HORIZON2020 funded project. The main goal of this survey was to better understand users and&nbsp;usage of the Historic Graves website and how to improve the visitor experience.</div>

opencc-by-sa-4.0Apr 2024View details →
zenodo40/100

Updates applied to Flemish online news and their associated change types

<p>This dataset contains 291,666 news articles produced by six different Flemish online news outlets (VRT NWS, Knack, Het Laatste Nieuws, Het Nieuwsblad, De Morgen, De Standaard), together with (in total 197,979)&nbsp;updates applied to these news articles in the first 24 hours after publication. The respective article versions (one row per article version that has been put online over time) can be found in the &#39;article_versions_vrt.csv&#39;,&nbsp;&#39;article_versions_knack.csv&#39;,&nbsp;&#39;article_versions_hln.csv&#39;,&nbsp;&#39;article_versions_nieuwsblad.csv&#39;,&nbsp;&#39;article_versions_demorgen.csv&#39; and&nbsp;&#39;article_versions_standaard.csv&#39;. Documentation regarding the meaning of the attributes in these files is provided in &#39;article_versions_README.txt&#39;.</p> <p>Next to this dataset, we also provide a coded set of changes made during a subset of the news updates in &#39;coded_article_change_types.csv&#39;. The file contains all text extracts that are&nbsp;part of a specific change, together with the type of the change to which the text extract belongs. Corresponding documentation is provided in &#39;coded_article_change_types_README.txt&#39;.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Supplementary online material for KIC 4150611: A quadruply eclipsing heptuple star system with a g-mode period-spacing pattern. Eclipse modelling of the triple and spectroscopic analysis

<p>Additional figures and data supplementary to the published (or soon-to-be-published) paper KIC 4150611: A quadruply eclipsing heptuple star system with a g-mode period-spacing pattern Eclipse modelling of the triple and spectroscopic analysis.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Medical Students use Online Study Materials more than School-Provided Resources when preparing for USMLE Step 1

<p>In this excel file contains the raw data collected from a survery sent out to students attend ULSOM and UNRSOM. The raw data was processed and analyzed in the sheet titled "graphs".&nbsp;</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record