Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,481
datasets available to search
ShareScore release 0.9.0
Dataset results
3,481 results for “data set”
Data Set: In-session dropout prediction model
<p>In-session dropout prediction model</p> <p>This project describes an in-session prediction model that predicts student early dropout from online learning exercises.<br> Dropout prediction models for Massive Open Online Courses (MOOCs) have shown high accuracy rates in<br> the past and make personalized interventions possible. While MOOCs have traditionally high dropout rates,<br> school homework and assignments are supposed to be completed by all learners. In the pandemic, online<br> learning platforms were used to support school teaching. In this setting, dropout predictions have to be designed differently as a simple dropout from the (mandatory) class is not possible. The aim of our work is to<br> transfer traditional temporal dropout prediction models to in-session dropout prediction for school-supporting<br> learning platforms. For this purpose, we used data from more than 164,000 sessions by 52,000 users of the<br> online language learning platform orthografietrainer.net. We calculated time-progressive machine learning<br> models that predict dropout after each step (completed sentence) in the assignment using learning process<br> data. The multilayer perceptron is outperforming the baseline algorithms with up to 87% accuracy. By extending the binary prediction with dropout probabilities, we were able to design a personalized intervention<br> strategy that distinguishes between motivational and subject-specific interventions. <br> A random state is not set, thus, results might differ marginally.</p> <p>Whole project described in: <br> N. Rzepka, K. Simbeck, H.-G. Müller, and N. Pinkwart<br> Keep It Up: In-session Dropout Prediction to Support Blended Classroom Scenarios<br> Proceedings of the 14th International Conference on Computer Supported Education - Volume 2: CSEDU,<br> SciTePress, 2022, ISBN 978-989-758-562-3 </p> <p> </p> <p> </p>
Data Set of an online controlled experiment to study adaptive learning
<p>Online-controlled experiment evaluation - Data Set</p> <p>Digital learning platforms are more and more used in blended classroom scenarios in Germany. However, as learning processes are different among students, adaptive learning platforms can offer personalized learning, e.g. by individual feedback and corrections, task sequencing, or recommendations. As digital learning platforms are already used in classroom settings, we propose the transformation of these plat-forms into adaptive learning environments. To measure the effectiveness and improvements achieved through the adaptions an online-controlled experiment design is created. In our experiment, we therefore investigate the effectiveness of different inter-ventions on a large user group in a four-month online-controlled experiment. For this purpose, the highly frequented German learning platform Orthografietrainer.net was transformed into an adaptive learning platform and users were randomly assigned to different interventions.</p> <p>The experimental design is published here: N. Rzepka, K. Simbeck, H.-G. Müller, and N. Pinkwart An Online Controlled Experiment Design to Support the Transformation of Digital Learning towards Adaptive Learning Platforms Proceedings of the 14th International Conference on Computer Supported Education - Volume 2: CSEDU,, SciTePress, 2022, ISBN 978-989-758-562-3 </p> <p>The architectural concept is published here: Rzepka, N., Simbeck, K., Müller, H.-G. & Pinkwart, N., (2022). Adaptive Learning as a Service – A concept to extend digital learning platforms?. In: Henning, P. A., Striewe, M.-0. 0. & Wölfel, M.-0. 0. (Hrsg.), 20. Fachtagung Bildungstechnologien (DELFI). Bonn: Gesellschaft für Informatik e.V.. (S. 237-238). DOI: 10.18420/delfi2022-049 </p> <p>The findings of this experiment are published here: tba</p> <p>The code to this evaluation can be found on Zenodo: <a href="https://doi.org/10.5281/zenodo.7755546">10.5281/zenodo.7755546</a></p>
Data Set: Solution Probability in Online Learning Environments
<pre>Solution Probability Model and Fairness Evaluation This in-session prediction model seeks to predict the users’ performance on the Orthografietrainer.net platform. The target variable is binary and predicts if the user will do the following sentence correctly or not. For fairness evaluations the best models (MLP and DTE), and the worst model (SVM) are considered. A random state is not set, thus, results might differ marginally. A detailed description of the solution probability model and the fairness evaluation can be found here: tba</pre>
A Pseudonymized Rehydrated Tweets Data Set collected before and after the Onset of the War between Russia and Ukraine in 2022
<p>Related to our other dataset publication, which can be found <a href="https://zenodo.org/record/6381899#.Yj1mCy8w2J8">here</a>, we publish the pseudonymized texts of the tweets no longer available after the rehydration process via Twitter's historic search API. Please note that we restrict access to these datasets to specific parties under certain conditions, as can be seen below. </p> <p>The datasets contain the timestamps of the tweets, the pseudonymized full text, as well as the time of rehydration. </p> <p>If you gain access to this dataset, please cite our paper: </p> <blockquote> <p>Pohl, Janina Susanne and Seiler, Moritz Vinzent and Assenmacher, Dennis and Grimme, Christian, A Twitter Streaming Dataset collected before and after the Onset of the War between Russia and Ukraine in 2022 (March 25, 2022). Available at SSRN: <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4066543">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4066543</a></p> </blockquote>
Hateful Messages: A Conversational Data Set of Hate Speech produced by Adolescents on Discord
<p>With the rise of social media, a rise of hateful content can be observed. Even though the understanding and definitions of hate speech varies, platforms, communities, and legislature all acknowledge the problem. Therefore, adolescents are a new and active group of social media users. The majority of adolescents experience or witness online hate speech. Research in the field of automated hate speech classification has been on the rise and focuses on aspects such as bias, generalizability, and performance. To increase generalizability and performance, it is important to understand biases within the data. This research addresses the bias of youth language within hate speech classification and contributes by providing a modern and anonymized hate speech youth language data set consisting of 88.395 annotated chat messages. The data set consists of publicly available online messages from the chat platform Discord. ~6,42\% of the messages were classified by a self-developed annotation schema as hate speech. For 35.553 messages, the user profiles provided age annotations setting the average author age to under 20 years old.</p>
Data Sets of Timing and Kinematics of Late Cenozoic Strike-slip Faults in Western Tibetan Plateau: New Constraints from Provenance Analysis and Detrital Thermochronology
<p>Data Sets of Wei et al. (2023). The manuscript is titled Timing and Kinematics of Late Cenozoic Strike-slip Faults in Western Tibetan Plateau: New Constraints from Provenance Analysis and Detrital Thermochronology</p>
Korean Colorectal Cancer Cohort RNA-seq data set from Asan Medical Center
<p>This study presents a comprehensive molecular analysis of colorectal cancer (CRC) in Korean patients using RNA-sequencing data and clinical information. Key findings include significant differences in gene expression between tumor and normal tissues and identification of dysregulated pathways related to tumor progression. CRC samples were stratified into molecular subtypes, confirming their similarity to established CRC classifications. Immunological analysis highlighted distinguishing features between tumor and normal tissues, and some patients were found to be responsive to immunotherapy, especially those with microsatellite stable (MSS) tumors, who demonstrated a better prognosis. These insights enhance the understanding of CRC in Korean patients and support the use of RNA-seq data to inform cancer treatment and improve patient care.</p>
Korean Colorectal Cancer Cohort RNA-seq data set from Seoul National University Bundang Hospital and Uijeongbu St. Mary's Hospital
<p>Colorectal cancer (CRC) is a leading cause of cancer-related mortality worldwide, and understanding the molecular mechanisms underlying CRC development and progression is critical for developing effective treatments. In this study, we performed RNA sequencing (RNA-seq) analysis of CRC tissue samples and adjacent normal tissue samples from 384 cohort to investigate differential gene expression between the two groups. We identified a total of 19,635 expressed genes in CRC tissue and normal tissue, with 3,816 differentially expressed genes (DEGs) identified between the two groups. Functional annotation analysis revealed that upregulated DEGs were significantly enriched in pathways related to cell cycle, DNA replication, and IL-17, while downregulated DEGs were enriched in metabolic pathways. We also analyzed relationships between the clinical information and subtypes using Consensus Molecular Subtype (CMS). Our findings provide valuable insights into the molecular mechanisms underlying Korean CRC patients and suggest potential targets for developing new treatments for this deadly disease. The raw data and processed results of this study have been deposited in a public repository to enable further analysis and exploration.</p> <p>Please cite the original article:</p>
Data set for Persistent LBP study
<p>Raw data for the study 'Prevalence and characteristics of women with persistent LBP postpartum'</p>
WRF-Chem configurations and input data sets for sensitivity tests of emission inventories
<p>WRF-Chem and WPS v4.4 source codes and their configurations with namelist files.</p> <p>Emission inventory data sets (EDGAR-HTAP v2 and v3) for 'anthro_emis' input are included.</p> <p>The KORUS v5 emission data are provided with 'wrfchemi' format.</p> <p>The 'namelist.input' contains physics and chemistry options that are used for WRF-Chem model.</p> <p>The model grid information is available in 'namelist.wps'.</p> <p> </p> <p>Kim, K.-M., Kim, S.-W., Seo, S., Blake, D. R., Cho, S., Crawford, J. H., Emmons, L., Fried, A., Herman, J. R., Hong, J., Jung, J., Pfister, G., Weinheimer, A. J., Woo, J.-H., and Zhang, Q.: Sensitivity of the WRF-Chem v4.4 ozone, formaldehyde, and precursor simulations to multiple bottom-up emission inventories over East Asia during the KORUS-AQ 2016 field campaign, Geosci. Model Dev. Discuss. [preprint], https://doi.org/10.5194/gmd-2023-132, in review, 2023.</p>
SCANS Data Set - A Reference Data Set for Handwritten Text Detection
<p># SCANS Data Set<br> <br> A scientific paper with handwritten annotations as reference data set for handwritten text detection.<br> <br> The paper is</p> <pre><code>Sharon Fogel, Hadar Averbuch-Elor, Sarel Cohen, Shai Mazor, Roee Litman; Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti on (CVPR), 2020, pp. 4324-4333</code></pre> <p><br> The annotations were made by two people with different pens. The text is randomly taken from "The Fellowship Of The Ring" by JRR Tolkien.<br> <br> The documents in folder `images` were scanned with a Canon Pixima TR4550 with 300DPI. The `labels-` folders contain the labels for paragraph, line, and word level in following JSON format:<br> </p> <pre><code>{ "image_file": "paper-0001.png", # The corresponding image in folder `images` "shape": [ # The size of the image 3495, 2473 ], "properties": [ # A list of all line bounding boxes (for paragraphs, lines, or words) { "type": "HWL", # HWA: handwritten paragraph, HWL: handwritten line, HWW: handwritten word "id": "18", "points": [ # The coordinates of the bounding boxes [ 1243.87096093748, 1515.21674083916 ], [ 2149.01588303892, 1515.21674083916 ], ... ] }, ... ] }</code></pre> <p> </p>
Dynamic rewiring of transcription factor networks during smooth muscle cell phenotypic modulation (RNA-Seq data sets)
GEO Series GSE111714. Rattus norvegicus. 14 samples. Type: Expression profiling by high throughput sequencing.
CASSINI RSS RAW DATA SET - GWE2 V1.0
not applicable
CASSINI RSS RAW DATA SET - GWE3 V1.0
not applicable
Phoenix Robotic Arm Data Set
The Phoenix Robotic Arm derived and test data set migrated from PDS3
GALILEO PROBE ASI RAW DATA SET
This data set is most accurately described in [SEIFFETAL1996]:
CASSINI RSS RAW DATA SET - GWE1 V1.0
not applicable
CASSINI RSS RAW DATA SET - SCE1 V1.0
not applicable
Affymetrix U133Plus2.0 data sets of Human Myeloma Cellline JJN3 with non-specific, scrambled shRNA or CKS1B shRNA
GEO Series GSE3369. Homo sapiens. 6 samples. Type: Expression profiling by array.
Dynamic rewiring of transcription factor networks during smooth muscle cell phenotypic modulation (ChIP-Seq data sets)
GEO Series GSE111712. Rattus norvegicus. 22 samples. Type: Genome binding/occupancy profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.