Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12
datasets available to search
ShareScore release 0.9.0
Dataset results
12 results for “symbolic datasets”
Dataset supporting the paper: Symbolic Versus Numerical Computation and Visualization of Parameter Regions for Multistationarity of Biological Networks
<p>Dataset supporting the paper:</p> <p>Matthew England, Hassan Errami, Dima Grigoriev, Ovidiu Radulescu, Thomas Sturm, and Andreas Weber. Symbolic Versus Numerical Computation and Visualization of Parameter Regions for Multistationarity of Biological Networks. In Proceedings of CASC ’17, Beijing, China, September 18-22 2017, 15 pages. Springer, 2017.</p> <p>The files whose name starts with "SamplePoints" are text files containing the data that produced the plots in the paper.</p> <p>The files whose name starts with "Sys" show the Maple computations used to produce the data. The mw files are to be run with the Maple Computer Algebra System (https://www.maplesoft.com/products/maple/). Pdf printouts of these have also been included for those who do not have access to Maple.</p> <p> </p>
CLDF dataset derived from the Johansson et al.'s "The typology of sound symbolism" from 2020
<p>Cite the source of the dataset as:</p> <blockquote> <p>Erben Johansson, N., Anikin, A., Carling, G., & Holmer, A. (2020). The typology of sound symbolism: Defining macro-concepts via their semantic and phonetic features, Linguistic Typology , 24(2), 253-310. doi: https://doi.org/10.1515/lingty-2020-2034</p> </blockquote>
OTMM Symbolic Section Dataset
<p>otmm_symbolic_section_dataset</p> <p>The section test dataset of music scores of Ottoman-Turkish makam music</p> <p>This repository contains the audio section annotations and the scores used in the paper:</p> <blockquote> <p>Şentürk, S., & Serra X. (2016). A method for structural analysis of Ottoman-Turkish makam music scores. In Proceedings of 6th International Workshop on Folk Music Analysis (FMA 2016), (pp. 39–46)., Dublin, Ireland.</p> </blockquote> <p>Please cite the publication above in any work using this dataset.</p> <p>The repository contains a test dataset of the SymbTr scores of 23 vocal compositions in the şarkı form and 42 instrumental compositions in peşrev and sazsemaisi forms in the txt and pdf format. The scores are selected from the release version 2.4.2 and the section annotations done by the first author of the paper. In the sections folder, the annotated sections for each score are stored in a csv file, which has the same name as the SymbTr-name (makam--form--usul--name--composer) of the annotated score. The fields are:</p> <p>start_note: The starting note index in the SymbTr-txt score end_note: The ending note index in the SymbTr-txt score name: The basic semantic name of the section. Right now, it is the name (TESLİM, ARANAĞME...) annotated in the score for instumental sections or "VOCAL_SECTION" for vocal sections. melodic_structure: Melodic semiotic label lyric_structure: Lyrical semiotic label lyrics: Lyrics of the section slug: The processed version of "name" field with the Turkish characters and special characters handled.</p> <p>For the details of the dataset, please refer to the paper. For any further information please contact the authors.</p>
Turkish Makam Symbolic Phrase Segmentation Dataset
<p>makam-symbolic-phrase-segmentation-dataset</p> <p><strong>Data sets containing pieces in SymbTr2 format segmented into phrases</strong></p> <p>This study presents a large machine-readable dataset of Turkish makam music scores segmented into phrases by experts of this music. The segmentation facilitates computational research on melodic similarity between phrases, and relation between melodic phrasing and meter, rarely studied topics due to unavailability of data resources. It consists of 31362 phrases on a set of 480 scores of different compositions annotated by 3 experts.</p> <p>Please refer to the following publication if you use this data in your research:</p> <blockquote> <p>M. K. Karaosmanoglu, B. Bozkurt, A. Holzapfel, N. D. Disiacik, A symbolic dataset of Turkish makam music phrases, Folk Music Analysis Workshop (FMA), Istanbul, 2014.</p> </blockquote> <p>The refactored code for automatic phrase segmentation can be found here.</p> <p>For other deliverables of the paper please visit: http://www.rhythmos.org/shareddata/turkishphrases.html</p>
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
<p>We introduce <strong>PDMX</strong>: a <strong>P</strong>ublic <strong>D</strong>omain <strong>M</strong>usic<strong>X</strong>ML dataset for symbolic music processing. Refer to our <a title="PDMX Paper" href="https://arxiv.org/abs/2409.10831" target="_blank" rel="noopener">paper</a> for more information, and our <a title="PDMX GitHub Repository" href="https://github.com/pnlong/PDMX/" target="_blank" rel="noopener">GitHub repository</a> for any code-related details. Please cite both our paper and <a href="https://arxiv.org/abs/2410.02084" target="_blank" rel="noopener">our collaborators' paper</a> if you use this dataset (see our GitHub for more information).</p> <p>Upon further use of the PDMX dataset, we discovered a discrepancy between the public-facing copyright metadata on the <a href="https://musescore.com/">MuseScore website</a> and the internal copyright data of the MuseScore files themselves, which affected 31,221 (12.29% of) songs. We have decided to proceed with the former given its public visibility on Musescore (i.e. this is what the MuseScore website presents its users with). We have noted files with conflicting internal licenses in the <em><strong>license_conflict</strong></em> column of PDMX. We recommend using the <em><strong>no_license_conflict</strong></em> subset of PDMX (which still includes 222,856 songs) moving forward.</p> <p>Additionally, for each song in PDMX, we not only provide the <em>MusicRender</em> and metadata JSON files, but we also try to include the associated compressed MusicXML (MXL), sheet music (PDF), and MIDI (MID) files when available. Due to the corruption of 42 of the original MuseScore files, these songs lack those associated files (since they could not be converted to those formats) and only include the <em>MusicRender</em> and metadata JSON files. The <em><strong>all_valid</strong></em> subset of PDMX describes the songs where all associated files are valid.</p>
Datasets for "Machine-Learning-Enhanced Symbolic Regression for Methane Storage Prediction in Covalent Organic Frameworks"
<p>This collection contains the datasets and associated files used in the research presented in the manuscript titled "Machine Learning-Enhanced Symbolic Regression for Methane Storage Prediction in Covalent Organic Frameworks". The datasets are critical for the development and validation of machine learning and symbolic regression models aiming to predict methane storage capacities in covalent organic frameworks (COFs).</p> <p><strong>Included Datasets:</strong></p> <ol> <li><code>COF_Data_for_ML.csv</code>: This dataset was utilized for the development of machine learning models.</li> <li><code>COF_Data_for_SISSO.csv</code>: This dataset was employed for the development of SISSO-based symbolic regression models.</li> <li><code>ML_vs_GCMC.xlsx</code>: This comparative dataset features GCMC-calculated results alongside machine learning predictions.</li> <li><code>Feature_Combination.xlsx</code>: This file contains data detailing all the feature combinations explored in the study.</li> <li><code>ML_SISSO_GCMC.xlsx</code>: This comparative dataset includes GCMC calculations, SISSO-based symbolic regression model predictions, and ML predictions.</li> <li><code>Crystallographic_Properties_of_535k_COFs.xlsx</code>: This consolidated dataset presents the crystallographic properties of 535,293 COFs.</li> </ol> <p><strong>Software Used:</strong></p> <ul> <li>Machine Learning Computations: Scikit-Learn (<a href="https://scikit-learn.org/stable/" target="_new">https://scikit-learn.org/stable/</a>)</li> <li>GCMC Simulations: RASPA2 (<a href="https://github.com/iRASPA/RASPA2" target="_new">https://github.com/iRASPA/RASPA2</a>)</li> <li>SISSO Calculations: SISSO toolkit (<a href="https://github.com/rouyang2017/SISSO" target="_new">https://github.com/rouyang2017/SISSO</a>)</li> <li>Crystallographic property calculations: Zeo++ (<a href="https://www.zeoplusplus.org/" target="_new">https://www.zeoplusplus.org/</a>)</li> </ul> <p>The datasets are provided to enable replication of the study's findings, encourage further research in the field, and facilitate the development of advanced predictive models by the scientific community. Researchers who use these datasets are requested to cite this Zenodo entry as well as the associated paper upon its publication.</p>
GPAM: Genetic Programming with Associative Memory - datasets for symbolic regression
<p>This collection contains five data sets for symbolic regression generated using functions known as Koza-1, Nguyen-7, Nguyen-10, Korns-1, and Korns-4. In each function, we replaced some data points with randomly generated values from interval [−10, 10].</p>
Emotion4MIDI: A Lyrics-Based Emotion-Labeled Symbolic Music Dataset
<p>This dataset includes emotion labels for the publicly available MIDI dataset, namely Lakh MIDI Dataset and Reddit MIDI dataset. The values represent the probability of containing a particular emotion. For a single song, more than one emotion can be present, hence the values don't add up to 1.</p>
Dataset of A User-driven Hybrid Neuro-symbolic Approach for Knowledge Graph Creation from Relational Data
<p>This dataset contains the following:</p> <p>1. achieved percentage values of the generated RML rules using LXS and manually</p> <p>2. basic information about the example used and with which creation type users started</p> <p>3. all answers of users to the User Experience Questionnaire</p> <p>4. Answers to the structured part of the user interview</p>
NEON Neuro-Symbolic Dataset
<p>This is the dataset used in the experiments from the manuscript</p> <blockquote> <p>Harmon, I., Weinstein, B., Bohlman, S., White, E., & Wang, D. Z. (2024). A Neuro-Symbolic Framework for Tree Crown Delineation and Tree Species Classification. Remote Sensing, 16(23), 4365. <a href="https://doi.org/10.3390/rs16234365">https://doi.org/10.3390/rs16234365</a></p> </blockquote> <p> This dataset is the RGB images combined with canopy height models (CHM) for the NEON sites NIWO and TEAK. The TEAK data is from the following publication:</p> <blockquote> <p>Geoffrey A Fricker, Jonathan Daniel Ventura, Jeffrey Wolf, Malcolm P. North, Frank W. Davis, & Janet Franklin. (2019). A Convolutional Neural Network classifier identifies tree species in mixed-conifer forest from hyperspectral imagery [Data set]. In Remote Sensing. Zenodo. <a href="https://doi.org/10.5281/zenodo.3463589" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.3463589</a>.</p> </blockquote>
Dataset for Transformers to Predict the Applicability of Symbolic Integration Routines
<p>This is the version related to the publication. </p>
UMD-350MB: Refined MIDI Dataset for Symbolic Music Generation
<p><strong>UMD-350MB</strong></p> <p>The Universal MIDI Dataset 350MB (UMD-350MB) is a proprietary collection of 85,618 MIDI files curated for research and development within our organization. This collection is a subset sampled from a larger dataset developed for pretraining symbolic music models.</p> <p>The field of symbolic music generation is constrained by limited data compared to language models. Publicly available datasets, such as the Lakh MIDI Dataset, offer large collections of MIDI files sourced from the web. While the sheer volume of musical data might appear beneficial, the actual amount of valuable data is less than anticipated, as many songs contain less desirable melodies with erratic and repetitive events.</p> <p>The UMD-350MB employs an attention-based approach to achieve more desirable output generations by focusing on human-reviewed training examples of single-track melodies, chord progressions, leads and arpeggios with an average duration of 8 bars. This was achieved by refining the dataset over 24 months, ensuring consistent quality and tempo alignment. Moreover, the dataset is normalized by setting the timing information to 120 BPM with a tick resolution (PPQ) of 96 and transposing the musical scales to C major and A minor (natural scales).</p> <p><strong>Melody Styles</strong></p> <p>A major portion of the dataset is composed of newly produced private data to represent modern musical styles.</p> <ul> <li>Pop: 1970s to 2020s Pop music</li> <li>EDM: Trance, House, Synthwave, Dance, Arcade</li> <li>Jazz: Bebop, Ballad, Latin-Jazz, Bossa-Jazz, Ragtime</li> <li>Soul: 80s Classic, Neo-Soul, Latin-Soul</li> <li>Urban: Pop, Hip-Hop, Trap, R&B, Afrobeat</li> <li>World: Latin, Bossa Nova, European</li> <li>Other: Film, Cinematic, Game music and piano references</li> </ul> <p><em>Actual MIDI files are unlabeled for unsupervised training.</em></p> <p><strong>Dataset Access</strong></p> <p>Please note that this is a closed-source dataset with very limited access. Considerations for access include proposals for data augmentation, chord extraction and other enhancement methods, whether through scripts, algorithmic techniques, manual editing in a DAW or additional processing methods.</p> <p>For inquiries about this dataset, please email us.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.