Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,688
datasets available to search
ShareScore release 0.7.1
Dataset results
3,688 results for “computer”
The Globus Compute Dataset
<p>We present a unique function-as-a-service (FaaS) dataset capturing the use of the Globus Compute (previously funcX) platform. Globus Compute implements a federated model via which users may deploy endpoints on arbitrary remote computers, from the edge to high-performance computing (HPC) cluster, and they may then invoke Python functions on those endpoints via a reliable cloud-hosted service. The dataset covers 31 weeks and includes 2,121,472 task submissions from 252 users executed on 580<br>remote computing endpoints. It includes 277,386 registered functions. </p>
Dataset for Investigating Anomalies in Compute Clusters
<p><strong>Abstract</strong></p><p>The dataset was collected for 332 compute nodes throughout May 19 - 23, 2023. May 19 - 22 characterizes normal compute cluster behavior, while May 23 includes an anomalous event. The dataset includes eight CPU, 11 disk, 47 memory, and 22 Slurm metrics. It represents five distinct hardware configurations and contains over one million records, totaling more than 180GB of raw data.</p><p><strong>Background</strong></p><p>Motivated by the goal to develop a digital twin of a compute cluster, the dataset was collected using a Prometheus server (1) scraping the Thomas Jefferson National Accelerator Facility (JLab) batch cluster used to run an assortment of physics analysis and simulation jobs, where analysis workloads leverage data generated from the laboratory's electron accelerator, and simulation workloads generate large amounts of flat data that is then carved to verify amplitudes. Metrics were scraped from the cluster throughout May 19 - 23, 2023. Data from May 19 to May 22 primarily reflected normal system behavior, while May 23, 2023, recorded a notable anomaly. This anomaly was severe enough to necessitate intervention by JLab IT Operations staff.</p><p>The metrics were collected from CPU, disk, memory, and Slurm. Metrics related to CPU, disk, and memory provide insights into the status of individual compute nodes. Furthermore, Slurm metrics collected from the network have the capability to detect anomalies that may propagate to compute nodes executing the same job.</p><p><strong>Usage Notes</strong></p><p>While the data from May 19 - 22 characterizes normal compute cluster behavior, and May 23 includes anomalous observations, the dataset cannot be considered labeled data. The set of nodes and the exact start and end time affected nodes demonstrate abnormal effects are unclear. Thus, the dataset could be used to develop unsupervised machine-learning algorithms to detect anomalous events in a batch cluster.</p><p><a href="https://doi.org/10.48550/arXiv.2311.16129">https://doi.org/10.48550/arXiv.2311.16129</a></p>
IonSolv-Aq Dataset for: Experimental Compilation and Computation of Hydration Free Energies for Ionic Solutes
<p>This repository includes datasets and supplementary materials for the manuscript "Experimental Compilation and Computation of Hydration Free Energies for Ionic Solutes" by Jonathan W. Zheng and William H. Green. <strong>Citations should refer directly to the manuscript:</strong></p> <blockquote> <p>Zheng, J. W., & Green, W. H. (2023). Experimental Compilation and Computation of Hydration Free Energies for Ionic Solutes. <em>The Journal of Physical Chemistry A</em>, <em>127</em>(48), 10268-10281.</p> </blockquote> <p>This compilation includes experimental and computed solvation free energies for the compounds in the IonSolv-Aq dataset, as well as .xyz files for all conformers used in the corresponding work. The lower-quality set of data described in the manuscript is also available in the "extra-anion-data.zip" archive file.</p>
Phlorest phylogeny derived from Bowern & Atkinson 2012 'Computational phylogenetics and the internal structure of Pama-Nyungan'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Bowern C & Atkinson QD. 2012. Computational phylogenetics and the internal structure of Pama-Nyungan. Language, 88(4), 817-845.</p> </blockquote>
Phlorest phylogeny derived from Chacon & List 2015 'Improved computational models of sound change shed light on the history of the Tukanoan languages'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Chacon TC, List J-M (2015) Improved computational models of sound change shed light on the history of the Tukanoan languages. Journal of Language Relationship, 3:177–203.</p> </blockquote>
Discrete Dataset - A dataset for computing with operators defined on a finite chain
<p>Discrete Dataset is a collection of the main operators defined on the finite chain L_n={0,1,...,n} up to n=11. These operators have been computationally generated, with the aim of being used to study properties of the operators. The easiest way to use these operators is with the Python package "DiscreteFuzzyOperators", published at https://zenodo.org/doi/10.5281/zenodo.5031268. It currently contains the following operators:</p><ul><li>For n=1,2,3,4, it contains:<ul><li>Discrete aggregation functions.</li><li>Smooth discrete aggregation functions.</li><li>Smooth and commutative discrete aggregation functions.</li><li>Discrete conjunctions.</li><li>Discrete disjunctions.</li><li>Discrete t-norms.</li><li>Discrete t-conorms.</li><li>Discrete uninorms.</li><li>Discrete negations.</li></ul></li><li>For n=5,6, it contains:<ul><li>Smooth and commutative discrete aggregation functions.</li><li>Discrete conjunctions.</li><li>Discrete disjunctions.</li><li>Discrete t-norms.</li><li>Discrete t-conorms.</li><li>Discrete uninorms.</li><li>Discrete negations.</li></ul></li><li>For n=7,8, it contains:<ul><li>Discrete t-norms.</li><li>Discrete t-conorms.</li><li>Discrete uninorms.</li><li>Discrete negations.</li></ul></li><li>For n=9,10,11, it contains:<ul><li>Discrete t-norms.</li><li>Discrete t-conorms.</li><li>Discrete negations.</li></ul></li></ul>
Synchrotron X-ray Computed Tomography scan of a wasp
<h4>Contents:</h4><ul><li><i>bee_yazeed-20231001T170032.h5</i> - SXCT scan of a wasp performed at beamline <a href="https://www.sesame.org.jo/beamlines/beats">ID10-BEATS</a> of SESAME.</li><li><i>SESAME_wasp_yazeed.avi -</i> 3D video rendering of phase-contrast CT reconstruction of <i>bee_yazeed-20231001T170032</i>. The dataset was reconstructed using <a href="https://github.com/gianthk/alrecon/tree/master">alrecon</a>. The video was created using ORS Dragonfly.</li></ul><h4>H5 dataset information:</h4><ul><li>Raw experimental data (sinogram, flat fields and dark fields) and metadata are stored in a common .H5 file.</li><li>The HDF5 file is organized hierarchically following the <a href="https://dxfile.readthedocs.io/en/latest/">Scientific Data Exchange (DXfile)</a> community standard.</li></ul><h4>How to reconstruct:</h4><ul><li>You can use <a href="http://www.silx.org/">Silx</a> to read and explore the .H5 dataset.</li><li>The file can be read within Python using the <a href="https://dxchange.readthedocs.io/en/latest/">DXChange</a> package.</li><li>See the <a href="https://beats.readthedocs.io/reconstruction.html">ID10-BEATS beamline user guide</a> for a detailed description on how to process and reconstruct the scan.</li></ul>
4D cone beam computed tomography phantom data set
<p>This data set accompanies the following Medical Physics publication: <a href="https://doi.org/10.1002/mp.14441"><i>Madesta, F., Sentker, T., Gauer, T., & Werner, R. (2020). Self‐contained deep learning‐based boosting of 4D cone‐beam CT reconstruction. Medical Physics, 47(11), 5619-5631</i></a><i>.</i></p><p>It comprises 6 time-resolved (4D) cone-beam computed tomography scans with the following scan configurations:</p><ul><li>4D CBCT Scanner: Varian TrueBeam (the detailed scan geometry and further details can be found in Scan.xml included in each scan)</li><li>Phantom: <a href="https://www.cirsinc.com/products/radiation-therapy/dynamic-thorax-motion-phantom/">Dynamic Thorax Phantom: Model 008A</a></li><li>The following motion patterns are included:<ol><li>SI amplitude of insert: ±10mm, pattern: sin, period: 5.0s</li><li>SI amplitude of insert: ±10mm, pattern: cos**4, period: 5.0s</li><li>SI amplitude of insert: ±10mm, pattern: sin, period: 2.5s</li><li>SI amplitude of insert: ±10mm, pattern: cos**4, period: 2.5s</li><li>SI amplitude of insert: ±10mm, pattern: sin, period: 7.5s</li><li>SI amplitude of insert: ±10mm, pattern: cos**4, period: 7.5s</li></ol></li></ul>
Resistive switching in benzylammonium-based Ruddlesden–Popper layered hybrid perovskites for non-volatile memory and neuromorphic computing
<p><span>Structural, optoelectronic, and supplementary characterisation data for “</span><span>Resistive Switching in Benzylammonium-Based Ruddlesden-Popper Layered Hybrid Perovskites for Non-Volatile Memory and Neuromorphic Computing ”</span><span>, DOI:</span><span>10.1039/d3ma00618b</span><span>.</span></p>
Fast Symbolic Computation of Bottom SCCs - TACAS 2024 artifact
<p>This is the artifact for the paper "<em>Fast Symbolic Computation of Bottom SCCs</em>", by Anna B. Jakobsen, Rasmus S. M. Jørgensen, Jaco van de Pol and Andreas Pavlogiannis, appearing in TACAS 2024.</p> <p>The artifact contains the LTSmin toolset, extended with an implementation of the algorithms from the paper, to compute the Bottom Strongly Connected Components of a directed graph, provided symbolically by BDDs (Binary Decision Diagrams).</p> <p>The artifact also contains the data set, consisting of directed graphs (state spaces) specified in DVE (Divine), PNML (Petri Nets) and BN (Boolean Networks).</p> <p>The file README.md contains the instructions how to setup the artifact on Ubuntu and how to run the experiment scripts.</p>
A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting
<p><span><span><span><span>We demonstrate a resilient workflow enabled by the LEXIS Platform, running a time- and safety-critical biomedical simulation of virtual stent placement in intracranial arteries using the HemoFlow application. The workflow, as captured on the video, gracefully handles failures of single computing steps or entire computing systems and thus lends itself to urgent computing applications. <br><br><span><span>The concept of this workflow has potential for realising ab-initio computational biomedical simulations which can provide live, targeted guidance to surgeons.</span></span></span></span></span></span></p>
Raw data for PIP2 interaction with TRPC3, explored through computation and electrophysiology
<p>The transient receptor potential canonical type 3 (TRPC3) channel plays a pivotal role in regulating neuronal excitability within the brain via its constitutive activity. The channel is intricately regulated by lipids and has previously been demonstrated to be positively modulated by PIP2. Using molecular dynamics simulations and patch clamp techniques, we reveal that PIP2 predominantly interacts with TRPC3 at the L3 lipid binding site, located at the intersection of pre-S1 and S1 helices. We propose a novel signal transduction pathway from the L3 through the re-entrant loop to a salt bridge between the TRP helix and S4-S5 linker. Notably, we find that both stimulated and constitutive TRPC3 activity require PIP2. These structural insights into the function of TRPC3 are invaluable for understanding the role of the TRPC subfamily in health and disease in native tissue.</p>
Additional Artifacts - Supplements to: A Resilient Workflow to Control a Biomedical HPC Simulation in an Urgent Computing Setting
<p>In this dataset, we have collected supplementary artifacts to support an understanding of the workflow presented in the submission cited (see related identifiers).</p> <p>These artifacts are (cf. README.md in the main folder of the tar.gz archive):</p> <p>A1: modified HemoFlow code (cf. https://github.com/gzavo/hemoflow) for our workflow experiments (subfolder "hemoflowcfd");<br>A2: workflow descriptions in python for Apache Airflow (subfolder "workflow");<br>A3: inputs (.xml/.npz) and output (.txt) for the example (subfolder "case").</p> <p> </p>
Bio-logger Ethogram Benchmark: A benchmark for computational analysis of animal behavior, using animal-borne tags
<p>This repository contains the datasets and experiment results presented in our <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a>:</p> <blockquote> <p>B. Hoffman, M. Cusimano, V. Baglione, D. Canestrari, D. Chevallier, D. DeSantis, L. Jeantet, M. Ladds, T. Maekawa, V. Mata-Silva, V. Moreno-González, A. Pagano, E. Trapote, O. Vainio, A. Vehkaoja, K. Yoda, K. Zacarian, A. Friedlaender, "A benchmark for computational analysis of animal behavior, using animal-borne tags," 2023.</p> </blockquote> <p>Standardized code to implement, train, and evaluate models can be found at <a href="https://github.com/earthspecies/BEBE/">https://github.com/earthspecies/BEBE/</a>. </p> <p>Please note the licenses in each dataset folder.</p> <p><strong>Zip folders beginning with "formatted":</strong> These are the datasets we used to run the experiments reported in the benchmark paper. </p> <p><strong>Zip folders beginning with "raw": </strong>These are the unprocessed datasets used in BEBE. Code to process these raw datasets into the formatted ones used by BEBE can be found at <a href="https://github.com/earthspecies/BEBE-datasets/">https://github.com/earthspecies/BEBE-datasets/</a>.</p> <p><strong>Zip folders beginning with "experiments": </strong>Results of the cross-validation experiments reported in the paper, as well as hyperparameter optimization. Confusion matrices for all experiments can also be found here. Note that dt, rf, and svm refer to the feature set from Nathan et al., 2012.</p> <p><em>Results used in Fig. 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (deep neural networks vs. classical models)</em><br>{dataset}_ harnet_nogyr<br>{dataset}_CRNN<br>{dataset}_CNN<br>{dataset}_dt<br>{dataset}_rf<br>{dataset}_svm<br>{dataset}_wavelet_dt<br>{dataset}_wavelet_rf<br>{dataset}_wavelet_svm</p> <p><em>Results used in Fig. 5D of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (full data setting)<br></em>If dataset contains gyroscope (HAR, jeantet_turtles, vehkaoja_dogs):<br>{dataset}_harnet_nogyr<br>{dataset}_harnet_random_nogyr<br>{dataset}_harnet_unfrozen_nogyr<br>{dataset}_RNN_nogyr<br>{dataset}_CRNN_nogyr<br>{dataset}_rf_nogyr<br><br>Otherwise:<br>{dataset}_harnet_nogyr<br>{dataset}_harnet_unfrozen_nogyr<br>{dataset}_harnet_random_nogyr<br>{dataset}_RNN_nogyr<br>{dataset}_CRNN<br>{dataset}_rf</p> <p><em>Results used in Fig. 5E of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a> (reduced data setting)<br></em>If dataset contains gyroscope (HAR, jeantet_turtles, vehkaoja_dogs):<br>{dataset}_harnet_low_data_nogyr<br>{dataset}_harnet_random_low_data_nogyr<br>{dataset}_harnet_unfrozen_low_data_nogyr<br>{dataset}_RNN_low_data_nogyr<br>{dataset}_wavelet_RNN_low_data_nogyr<br>{dataset}_CRNN_low_data_nogyr<br>{dataset}_rf_low_data_nogyr</p> <p>Otherwise:<br>{dataset}_harnet_low_data_nogyr<br>{dataset}_harnet_random_low_data_nogyr<br>{dataset}_harnet_unfrozen_low_data_nogyr<br>{dataset}_RNN_low_data_nogyr<br>{dataset}_wavelet_RNN_low_data_nogyr<br>{dataset}_CRNN_low_data<br>{dataset}_rf_low_data<br><br></p> <p><strong>CSV files</strong>: we also include summaries of the experimental results in experiments_summary.csv, experiments_by_fold_individual.csv, experiments_by_fold_behavior.csv. </p> <p><em>experiments_summary.csv - results averaged over individuals and behavior classes<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting <br>fig4 (bool): True if dataset+experiment was used in figure 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>f1_mean (float): mean of macro-averaged F1 score, averaged over individuals in test folds<br>f1_std (float): standard deviation of macro-averaged F1 score, computed over individuals in test folds<br>prec_mean, prec_std (float): analogous for precision<br>rec_mean, rec_std (float): analogous for recall<em><br><br>experiments_by_fold_individual.csv - results per individual in the test folds<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting <br>fig4 (bool): True if dataset+experiment was used in figure 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fold (int): test fold index<br>individual (int): individuals are numbered zero-indexed, starting from fold 1<br>f1 (float): macro-averaged f1 score for this individual<br>precision (float): macro-averaged precision for this individual<br>recall (float): macro-averaged recall for this individual<em><br></em></p> <p><em>experiments_by_fold_behavior.csv - results per behavior class, for each test fold<br></em>dataset (str): name of dataset<br>experiment (str): name of model with experiment setting <br>fig4 (bool): True if dataset+experiment was used in figure 4 of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5d (bool): True if dataset+experiment was used in figure 5d of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fig5e (bool): True if dataset+experiment was used in figure 5e of <a href="https://arxiv.org/abs/2305.10740">arxiv paper</a><br>fold (int): test fold index<br>behavior_class (str): name of behavior class<br>f1 (float): f1 score for this behavior, averaged over individuals in the test fold<br>precision (float): precision for this behavior, averaged over individuals in the test fold<br>recall (float): recall for this behavior, averaged over individuals in the test fold<br>train_ground_truth_label_counts (int): number of timepoints labeled with this behavior class, in the training set<em><br></em></p>
Chronic Ethanol Exposure Produces Sex-Dependent Impairments in Value Computations in the Striatum
<div> <div>These datasets and scripts are organized by figures. All data are stored as .mat format and can be open and manipulated using MATLAB. Scripts are all written in MATLAB and can be ran in MATLAB.</div> <div>There are two ways to run the code to reproduce each figures and statistics.</div> <div>1. Run RUN_ME.m. In this case, the file will automatically excute scripts to load corresponding data and figures.</div> <div>2. Open individual script to load corresponding data and generate statistics and figures.</div> <br> <div>All scripts here have been validated and tested. The system and coding environment is:</div> <div>- Windows 11 24H2</div> <div>- MATLAB 2023a</div> <br> <div>Matlab dependent package (not all are required but those are installed in my environment):</div> <div>- Bioinformatics Toolbox v4.17</div> <div>- Communications Toolbox v8.0</div> <div>- Computer Vision Toolbox v10.4</div> <div>- Curve Fitting Toolbox v3.9</div> <div>- Data Acquisition Toolbox v4.7</div> <div>- Database Toolbox v11.0</div> <div>- Deep Learning HDL Toolbox v1.5</div> <div>- Deep Learning Toolbox v14.6</div> <div>- DSP HDL Toolbox v1.2</div> <div>- Econometrics Toolbox v6.2</div> <div>- Financial Toolbox v6.5</div> <div>- Fixed-point Designer v7.6</div> <div>- Image Processing Toolbox v11.7</div> <div>- MATLAB Coder v5.6</div> <div>- MATLAB Compiler v8.6</div> <div>- MATLAB Compiler SDK v7.2</div> <div>- MATLAB Report Generator v5.14</div> <div>- MATLAB Support for MinGW-w64 C/C++ Compiler v23.1.0</div> <div>- Optimization Toolbox v9.5</div> <div>- Parallel Computing Toolbox v9.5</div> <div>- FR Toolbox v4.5</div> <div>- Signal Integrity Toolbox v1.3</div> <div>- Simulink v10.7</div> <div>- Statistics and Machine Learning Toolbox v12.5</div> <div>- Symbolic Math Toolbox v9.3</div> <div>- Text Analytics Toolbox v1.10</div> <div>- Wavelet Toolbox v6.3</div> </div>
Evaluation datasets and results of the paper "Efficient Online Computation of Business Process State From Trace Prefixes via N-Gram Indexing"
<p>Event logs, process models, and results corresponding to the paper "Efficient Online Computation of Business Process State From Trace Prefixes via N-Gram Indexing".</p> <p><em><strong>Inputs</strong></em>: preprocessed event logs and discovered process models (and their characteristics) used in the evaluation.</p> <ul> <li><em><strong>Real-life</strong></em>: preprocessed event logs (<em>xes</em> and <em>csv</em>) corresponding to the real-life processes used in the evaluation. Process models (<em>pnml</em>) discovered with the Inductive Miner infrequent for thresholds of 10%, 20%, and 50%. Characteristics (<em>txt</em>) of the event logs and process models. Ongoing cases result from splitting each case in the preprocessed event logs (under folder <em>split</em>).</li> <li><em><strong>Synthetic</strong></em>: simulated event logs (<em>csv</em>) corresponding to the synthetic processes used in the evaluation. Designed process models (<em>bpmn</em> and <em>pnml</em>). Ongoing cases result from splitting each case in the preprocessed event logs (under folder <em>split</em>). Ongoing cases with injected noise as described in the publication (under folders <em>noise_1</em>, <em>noise_2</em>, and <em>noise_3</em>).</li> </ul>
High-throughput robotic titration using computer vision
<ul> <li> <p>An automated HTE robotic titration using a liquid-handling robot Opentrons(OT-2) and a standard webcam enables in-situ, affordable titration analyses.</p> </li> <li>Its modular design allows adaptability for materials chemsitry and integration into automated workflows, enhancing efficiency in chemical search.</li> </ul>
Extended dataset for the validation the competent Computational Thinking test in grades 3-6
<p>Extended dataset for the validation the competent Computational Thinking test in grades 3-6<br>=======================================================</p> <p>• If you publish material based on this dataset, please cite the following :</p> <p> • The Zenodo repository : Laila El-Hamamsy, Barbara Bruno, Jessica Dehler Zufferey, & Francesco Mondada (2023). Extended dataset for the validation of the competent Computational Thinking test in grades 3-6 [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7983525 </p> <p> • The article on the validation of the computational thinking test for grades 3-6 : El-Hamamsy, L., Zapata-Cáceres, M., Martín-Barroso, E., Mondada, F., Zufferey, J. D., Bruno, B., & Román-González, M. (2025). The competent Computational Thinking test (cCTt): A valid, reliable and gender-fair test for longitudinal CT studies in grades 3–6. <em>Technology, Knowledge and Learning</em>, 1-55. https://doi.org/10.1007/s10758-024-09777-8 </p> <p>• License : This work is licensed under a Creative Commons Attribution 4.0 International license (CC-BY-4.0)</p> <p>• Creators : El-Hamamsy, L., Bruno, B., Dehler Zufferey, J., and Mondada, F.</p> <p>• Date May 30th 2023</p> <p>• Subject : Computational Thinking (CT), Assessment, Primary education, Psychometric validation</p> <p>• Dataset format : CSV. The dataset contains four files (one per grade, see detailed description below). Please note that the spreadsheets may contain missing values due to students not being present for a part of the data collection. To have access to the specific cCTt questions please refer to the original publication [1] and Zenodo repository [2] which provide the full set of questions and correct responses.</p> <p>• Dataset size < 500 kB</p> <p>• Data collection period : January and November 2021</p> <p>• Abbreviations :<br> - CT : Computational Thinking<br> - cCTt: competent CT test</p> <p>• Funding : This work was funded by the the NCCR Robotics, a National Centre of Competence in Research, funded by the Swiss National Science Foundation (grant number 51NF40_185543)</p> <p># References</p> <p>[1] El-Hamamsy, L., Zapata-Cáceres, M., Barroso, E. M., Mondada, F., Zufferey, J. D., & Bruno, B. (2022). The Competent Computational Thinking Test: Development and Validation of an Unplugged Computational Thinking Test for Upper Primary School. Journal of Educational Computing Research, 60(7), 1818–1866. https://doi.org/10.1177/07356331221081753 </p> <p>[2] El-Hamamsy, L., Zapata-Cáceres, M., Marcelino, P., Dehler Zufferey, J., Bruno, B., Martín Barroso, E., & Román-González, M. (2022). Dataset for the comparison of two Computational Thinking (CT) test for upper primary school (grades 3-4) : the Beginners' CT test (BCTt) and the competent CT test (cCTt) (Version 1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5885034 </p> <p>[3] El-Hamamsy, L., Zapata-Cáceres, M., Martín-Barroso, E. <em>et al.</em> The Competent Computational Thinking Test (cCTt): A Valid, Reliable and Gender-Fair Test for Longitudinal CT Studies in Grades 3–6. <em>Tech Know Learn</em> (2025). https://doi.org/10.1007/s10758-024-09777-8</p> <p>[4] Brennan, K. and Resnick, M. (2012). New frameworks for studying and assessing the development of computational thinking. page 25</p> <p>[5] El-Hamamsy, L., Zapata-Cáceres, M., Marcelino, P., Bruno, B., Dehler Zufferey, J., Martín-Barroso, E., & Román-González, M. (2022). Comparing the psychometric properties of two primary school Computational Thinking (CT) assessments for grades 3 and 4: The Beginners’ CT test (BCTt) and the competent CT test (cCTt). Frontiers in Psychology, 13. https://www.frontiersin.org/articles/10.3389/fpsyg.2022.1082659</p>
Music Data Sharing Platform for Computational Musicology Research (CCMUSIC DATASET)
<p>This platform is a multi-functional music data sharing platform for Computational Musicology research. It contains many music datas such as the sound information of Chinese traditional musical instruments and the labeling information of Chinese pop music, which is available for free use by computational musicology researchers.</p> <p>This platform is also a large-scale music data sharing platform specially used for Computational Musicology research in China, including 3 music databases: Chinese Traditional Instrument Sound Database (CTIS), Midi-wav Bi-directional Database of Pop Music and Multi-functional Music Database for MIR Research (CCMusic). All 3 databases are available for free use by computational musicology researchers. For the contents contained in the database, we will provide audio files recorded by the professional team of the conservatory of music, as well as corresponding labelled files, which have no commodity copyright problem and facilitate large-scale promotion. We hope that this music data sharing platform can meet the one-stop data needs of users and contribute to the research in the field of Computational Musicology.</p> <p> </p> <p>If you want to know more information or obtain complete files, please go to the official website of this platform:</p> <p><a href="https://ccmusic-database.github.io/en/">Music Data Sharing Platform for Academic Research</a></p> <p> </p> <ul> <li> <p><strong>Chinese Traditional Instrument Sound Database (CTIS)</strong></p> </li> </ul> <p>This database is developed by Prof. Han Baoqiang's team for many years, which collects sound information about Chinese traditional musical instruments. The database includes 287 Chinese national musical instruments, including traditional musical instruments, improved musical instruments and ethnic minority musical instruments.</p> <ul> <li> <p><strong>Multi-functional Music Database for MIR Research</strong></p> </li> </ul> <p>This database collects sound materials of pop music, folk music and hundreds of national musical instruments, and makes comprehensive annotation to form a multi-purpose music database for MIR researchers.</p> <ul> <li><strong>Midi-wav Bi-directional Database of Pop Music</strong></li> </ul> <p>This database contains hundreds of Chinese pop songs, and each song contains the corresponding midi-audio-lyric information. Among them, recording the vocal part and accompaniment part of audio independently is helpful to study the MIR task under the ideal situation. In addition, the information of singing techniques consistent with vocal part (such as breath sound, falsetto, breathing, vibrato, mute, slide, etc.) is marked in MuseScore, which constitutes a Midi-Wav bi-direction corresponding pop music database.</p>
Dataset of limericks for computational poetics
<p>Herein is a data set comprising 98k limericks scraped from the <a href="https://nam12.safelinks.protection.outlook.com/?url=http%3A%2F%2Fwww.oedilf.com%2Fdb%2FLim.php&data=04%7C01%7CAlmas.Abdibayev.GR%40dartmouth.edu%7Cbeec80009bb9495ff7d508d97dbaeb72%7C995b093648d640e5a31ebf689ec9446f%7C0%7C0%7C637679064035896909%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C1000&sdata=w4m26CuFXTpPdOl455RGF%2Bt5naG8Jb7Nt2mStXHvIms%3D&reserved=0">The Omnificent English Dictionary In Limerick Form - OEDILF</a>. It is a subset of the full data set, filtered to pass a basic test of standard limerick form (i.e., ensuring five lines, no emojis, no symbols). Each limerick was written by a human contributor whose work has passed through a rigorous moderation. This dataset is released alongside two companion papers: "BPoMP: The Benchmark of Poetic Minimal Pairs – Limericks, Rhyme, and Narrative Coherence" (Abdibayev, Riddell, Rockmore, RANLP 2021) and "Automating the Detection of Poetic Features: The Limerick as Model Organism" (Abdibayev, Riddell, Igarashi, Rockmore, SIGHUM 2021). The dataset is primarily released for use by NLP researchers interested in studying formal structure of poetry and more generally, interested in computational poetics. Each limerick is accompanied by metadata: author information, id within the website and "is_limerick" field, which denotes if limerick was recognized by our custom filter that was built to check for formal limerick properties (this tagging was a goal of the SIGHUM paper and reflects the results reported there - see the paper for details). Thus, if "is_limerick"=True this is a true positive, "is_limerick"=False is (almost surely) a false negative. We identify 70% of these as limericks and provide the tagging as a benchmark for the community to improve upon. With these considerations in mind we hope that NLP community will use this dataset to study poetical knowledge of language models trained on large corpora as many of their properties still remain a mystery to the community at large. We are excited for the possibilities ahead!</p> <p><strong>UPDATE</strong>: we released a new version of our dataset that contains all of the limericks that we planned to publish. Previous version (v2) was created using code that contained a bug which in turn lowered the number of available limericks.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.