Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
174
datasets available to search
ShareScore release 0.9.0
Dataset results
174 results for “online data”
Figure 4 from: Senderov V, Georgiev T, Penev L (2016) Online direct import of specimen records into manuscripts and automatic creation of data papers from biological databases . Research Ideas and Outcomes 2: e10617. https://doi.org/10.3897/rio.2.e10617
Figure 4 - Download of an EML from the GBIF Integarted Publishuing Toolkit (IPT)
Figure 1 from: Senderov V, Georgiev T, Penev L (2016) Online direct import of specimen records into manuscripts and automatic creation of data papers from biological databases . Research Ideas and Outcomes 2: e10617. https://doi.org/10.3897/rio.2.e10617
Figure 1 - Poll results about composition of audience during live participation.
CS#2 Data from online sensors on DWSP (drinking water service public – houses of water)
<div> <p>Data gathered from online sensor installed on DWSPs used to obtain infographics for monitoring water quality</p> </div>
Advanced Planning for Online Accounts and Data
ClinicalTrials.gov study NCT05222308. IPD Sharing: NO. Countries: 0. Publications: 0.
Gadgetron 2020 online class data - day 4
<p>MRI data in the MRD format for the Gadgetron Online class.</p>
Data from: Classification of mathematics ability and use of online tools
<p><span>Enrolment in secondary school science and mathematics subjects has been declining for some time. This decrease is particularly evident in the final years of secondary education </span><span>and occurs in parallel with the delivery of advanced mathematics. </span><span>This decreasing trend in science, technology, engineering and mathematics (STEM) enrolment at the secondary level are by no means restricted to Australia, with reports from the UK and Europe, the Middle East and Asia all echoing the same decline. Alongside this decline in mathematics enrolment comes an increased interest in, and reliance upon, online resources to bolster learning. </span><span>The educational challenge arising from such an environment, while notable in the high school context, has profound follow-on effects within the higher education sector</span><span>.</span></p> <p><span>The data in this submission is twofold and addresses both the decay in learning, and the increase in tool reliance. For the former we developed a tool to classify mathematics capacity in order to provide the educator with an easily deployed and robust manner. For the second, we analysed responses as to whether students (a) made use of online tools at all, and (b) did they use tools provided by us or another offering. Further to this, we analysed responses relating to student perception of the value of online tools.</span></p> <p><span>Building on previous works around self-efficacy, and including level of previous learning, we found three distinct groups (low, medium and high) which reflected pre-delivery testing scores which might allow an educator to design interventions targeting a students mathematical capacity. We also found no difference in perceived value of the use of online tools where students used either supplied or other programs.</span></p>
Data from: Classification of mathematics ability and use of online tools
Open the record for dataset details and reuse information.
Data from: Deep Sea Spy: an online citizen science annotation platform for science and ocean literacy
<p>Data sets of buccinid <em>Buccinum thermophilum </em>and crab <em>Segonzacia mesatlantica </em>after identifying unique groups (i.e. individuals) — using the updated version (v0.0.3) of the <a title="Deep Sea Spy - deeptools" href="https://github.com/DeepSeaSpy/deeptools">deeptools</a> package — among Deep Sea Spy citizen participants and expert.</p> <p>Data cleaning : annotated buccinids in background removed in <span><a href="https://zenodo.org/api/records/14203506/draft/files/CS_groups.csv/content" target="_blank" rel="noopener noreferrer">CS_groups.csv</a></span></p> <p>R script used to generate and analyse the dataset.</p> <p><strong>Please, use this version (v5).<br></strong></p>
Studying the Impact of Noises in Build Breakage Data [Online Appendix]
<p>This document represents an Online Appendix of our article, which has been accepted for publication in the IEEE Transactions on Software Engineering (TSE) in August 2019. The Online Appendix includes a replication package of the scripts and data we used in our work, in addition to more detailed results of the findings reported in the article.</p> <p> </p> <p><strong>Abstract.</strong></p> <p>Much research has investigated the common reasons for build breakages. However, prior research has paid little attention to builds that may break due to reasons that are unlikely to be related to development activities. For example, Continuous Integration (CI) builds may break due to timeout or connection errors while generating the build. Such kinds of build breakages potentially introduce noises to build breakage data. Not considering such noises may lead to misleading results when studying CI builds. In this paper, we propose three criteria to identify build breakages that can potentially introduce noises to build breakage data. We apply these criteria to a dataset of 350,246 builds from 153 GitHub projects that are linked with Travis CI. Our results reveal that 33% of the build breakages are due to environmental factors (e.g., errors in CI servers), 29% are due to (unfixed) errors in previous builds, and 9% are due to build jobs that were later deemed by developers as noisy (there is an overlap of 17% between these three types of breakages). We measure the impact of noises in build breakage data on modeling build breakages. We observe that models that use uncleaned build breakage data can lead to misleading associations between build breakages and development activities (e.g., the role of developer). However, such associations could not be observed after eliminating noisy build breakages. Moreover, we replicate a prior study that investigates the association between build breakages and development activities using data from 14 GitHub projects. We observe that some observations reported by the prior study (e.g., pull requests cause more breakages) do not hold after eliminating the noises from build breakage data.</p> <p> </p> <p><strong>Citing this research.</strong></p> <p>If you intend to use any materials (i.e., scripts, data, approach, or findings) of this work, please cite our article as follows:</p> <p>@article{ghaleb2019studying,<br> title={Studying the Impact of Noises in Build Breakage Data},<br> author={Ghaleb, Taher Ahmed and da Costa, Daniel Alencar and Zou, Ying and Hassan, Ahmed E.},<br> journal={IEEE Transactions on Software Engineering},<br> pages={1--14},<br> year={2019},<br> publisher={IEEE},<br> DOI={10.1109/TSE.2019.2941880}<br> }</p> <p>The official publication of this research can be found at <a href="https://dx.doi.org/10.1109/TSE.2019.2941880">https://dx.doi.org/10.1109/TSE.2019.2941880</a></p> <p> </p> <p><strong>Replication Package.</strong></p> <p>You may download the whole replication package (i.e., scripts, data, approach, or findings) using the zip file shown below. Please note that the size of the raw build logs used in this work is approximately 107 GB. You can download raw build logs using from Travis CI (using Travis API) or its AWS S3 backend (using Amazon S3 API).</p> <p> </p>
Data Set of an online controlled experiment to study adaptive learning
<p>Online-controlled experiment evaluation - Data Set</p> <p>Digital learning platforms are more and more used in blended classroom scenarios in Germany. However, as learning processes are different among students, adaptive learning platforms can offer personalized learning, e.g. by individual feedback and corrections, task sequencing, or recommendations. As digital learning platforms are already used in classroom settings, we propose the transformation of these plat-forms into adaptive learning environments. To measure the effectiveness and improvements achieved through the adaptions an online-controlled experiment design is created. In our experiment, we therefore investigate the effectiveness of different inter-ventions on a large user group in a four-month online-controlled experiment. For this purpose, the highly frequented German learning platform Orthografietrainer.net was transformed into an adaptive learning platform and users were randomly assigned to different interventions.</p> <p>The experimental design is published here: N. Rzepka, K. Simbeck, H.-G. Müller, and N. Pinkwart An Online Controlled Experiment Design to Support the Transformation of Digital Learning towards Adaptive Learning Platforms Proceedings of the 14th International Conference on Computer Supported Education - Volume 2: CSEDU,, SciTePress, 2022, ISBN 978-989-758-562-3 </p> <p>The architectural concept is published here: Rzepka, N., Simbeck, K., Müller, H.-G. & Pinkwart, N., (2022). Adaptive Learning as a Service – A concept to extend digital learning platforms?. In: Henning, P. A., Striewe, M.-0. 0. & Wölfel, M.-0. 0. (Hrsg.), 20. Fachtagung Bildungstechnologien (DELFI). Bonn: Gesellschaft für Informatik e.V.. (S. 237-238). DOI: 10.18420/delfi2022-049 </p> <p>The findings of this experiment are published here: tba</p> <p>The code to this evaluation can be found on Zenodo: <a href="https://doi.org/10.5281/zenodo.7755546">10.5281/zenodo.7755546</a></p>
Data Set: Solution Probability in Online Learning Environments
<pre>Solution Probability Model and Fairness Evaluation This in-session prediction model seeks to predict the users’ performance on the Orthografietrainer.net platform. The target variable is binary and predicts if the user will do the following sentence correctly or not. For fairness evaluations the best models (MLP and DTE), and the worst model (SVM) are considered. A random state is not set, thus, results might differ marginally. A detailed description of the solution probability model and the fairness evaluation can be found here: tba</pre>
Data for Reported user-generated online hate speech: The 'ecosystem', frames, and ideologies
<p>This is the dataset for the article entitled Reported user-generated online hate speech: The 'ecosystem', frames, and ideologies. The same dataset is provided in two different formats: comma-separated values (.csv) and Excel format (.xlsx). A basic legend to the data is provided separately in the corresponding PDF document. An extended legend to the data is available at: <strong><a href="https://doi.org/10.5281/zenodo.6656185">https://doi.org/10.5281/zenodo.6656185</a></strong>.</p>
Data and materials for "Exploring the spatial segmentation of housing markets from online listings"
<div> <div>This folder contains the materials from the publication: "Exploring the spatial segmentation of housing markets from online listings".</div> <br> <div>It includes the weighted networks of spatial units, where the link between two spatial units account not only for the presence of common real estate agencies operating in both units but also its influence, accounting for the relative share of those agencies in that spatial units. Additionally, it also includes the necessary code to perform the stochastic aggregative method from generic census data (using as example the IRIS codes).</div> <br> <div>We construct the networks using 3 different spatial units: 1000 m square grid cells, municipalities and census tracts.</div> </div>
Adaptation and evolution of teaching method for university programming subject to the online learning environment - Changelog Data
<p>This dataset contains raw data from changelogs of students studying Operating Systems class at the Technical University of Košice in the year 2020/2021.</p> <p>All of the data is anonymized and all names are replaced with the string *Anonymized name*. All of the content is in the Slovak language.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.