Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
174
datasets available to search
ShareScore release 0.9.0
Dataset results
174 results for “symbol”
Dataset supporting the paper: Symbolic Versus Numerical Computation and Visualization of Parameter Regions for Multistationarity of Biological Networks
<p>Dataset supporting the paper:</p> <p>Matthew England, Hassan Errami, Dima Grigoriev, Ovidiu Radulescu, Thomas Sturm, and Andreas Weber. Symbolic Versus Numerical Computation and Visualization of Parameter Regions for Multistationarity of Biological Networks. In Proceedings of CASC ’17, Beijing, China, September 18-22 2017, 15 pages. Springer, 2017.</p> <p>The files whose name starts with "SamplePoints" are text files containing the data that produced the plots in the paper.</p> <p>The files whose name starts with "Sys" show the Maple computations used to produce the data. The mw files are to be run with the Maple Computer Algebra System (https://www.maplesoft.com/products/maple/). Pdf printouts of these have also been included for those who do not have access to Maple.</p> <p> </p>
Bayesian Symbolic Learning to Build Analytical Correlations from Rigorous Process Simulations: Application to CO2 Capture Technologies
<p>Dataset of process simulations results of the natural gas sweetening and flue gas treatment (first and second sheet, respectively as indicated by the sheet name in the .xlsx file). The dataset refers to the publication <em>Bayesian Symbolic Learning to Build Analytical Correlations from Rigorous Process Simulations: Application to CO<sub>2</sub> Capture Technologies </em>by V. Negri, Vàzquey D., Sales-Pardo, Marta, Guimerà, R. and Guillén-Gosàlbez, G. The training and testing dataset are used to generate the figures in the main manuscript and supplementary information. </p> <p> </p>
Fast Symbolic Computation of Bottom SCCs - TACAS 2024 artifact
<p>This is the artifact for the paper "<em>Fast Symbolic Computation of Bottom SCCs</em>", by Anna B. Jakobsen, Rasmus S. M. Jørgensen, Jaco van de Pol and Andreas Pavlogiannis, appearing in TACAS 2024.</p> <p>The artifact contains the LTSmin toolset, extended with an implementation of the algorithms from the paper, to compute the Bottom Strongly Connected Components of a directed graph, provided symbolically by BDDs (Binary Decision Diagrams).</p> <p>The artifact also contains the data set, consisting of directed graphs (state spaces) specified in DVE (Divine), PNML (Petri Nets) and BN (Boolean Networks).</p> <p>The file README.md contains the instructions how to setup the artifact on Ubuntu and how to run the experiment scripts.</p>
Linguistic-symbolic classification of occupations
<p>As a result, it has been considered the following occupational categories, based on the degree of symbolic analysis and language intensity: (A) high symbolic analysts, (B) low symbolic analysts, (C) high intensity oral interaction with public (high public service), (D) low intensity oral interaction with public (low public service), and (E) manual labor occupations with limited symbolic and oral demands. Further, and within (B) category -low symbolic analysts-, it is distinguished between (B1) those whose work is generally inside the organization (such as file clerks) and (B2) those whose work includes interacting with the public (such as receptionists). In a similar way, (C) category is also divided into (C1) category of nurses and (C2-C5) which group the remainder of the high public service occupations.</p> <p>The result of the categorization of occupations by language use is summarized in the table below:</p> <table> <thead> <tr> <th>Major occupational classification</th> <th>Linguistic characteristics of occupation</th> <th>Sub classification</th> <th>Example of occupation</th> </tr> </thead> <tbody> <tr> <td>A: High symbolic analysts</td> <td>Produce/consume long or complex written communications, with variable but often important oral communication</td> <td>A1. Upper management</td> <td>Chief executive, human resource executive</td> </tr> <tr> <td> </td> <td> </td> <td>A2. Professionals</td> <td>Lawyer, doctor</td> </tr> <tr> <td> </td> <td> </td> <td>A3. Lower management</td> <td>First line manager/supervisor</td> </tr> <tr> <td> </td> <td> </td> <td>A4. High symbolic analysts, not managers</td> <td>Public relations specialists, computer systems specialists</td> </tr> <tr> <td>B: Low symbolic analysts</td> <td>Produce/consume short or simple written communications, with variable but often important oral communication</td> <td>B1. Low symbolic analysts with low likelihood of public interaction</td> <td>File clerks</td> </tr> <tr> <td> </td> <td> </td> <td>B2. Low symbolic analysts with high likelihood of public interaction</td> <td>Receptionists, billing/appointment clerks</td> </tr> <tr> <td>C: In-person service workers with high communicative demands</td> <td>Important oral communication, limited but present written skills, and high public interaction</td> <td>C1. Nurses</td> <td>Nurses</td> </tr> <tr> <td> </td> <td> </td> <td>C2. Assistants and technicians in public service settings</td> <td>Medical technicians</td> </tr> <tr> <td> </td> <td> </td> <td>C3. Police, etc.</td> <td>Police, detectives, investigators</td> </tr> <tr> <td> </td> <td> </td> <td>C4. Firefighters, emergency medical technicians</td> <td>Firefighters, emergency medical technicians</td> </tr> <tr> <td> </td> <td> </td> <td>C5. Miscellaneous</td> <td>Counselors, dispatchers</td> </tr> <tr> <td>D. In-person service workers with low communicative demands</td> <td>Simple oral communication and public interaction, very limited or no writing</td> <td>(no subcategories in our study)</td> <td>home health care aides, security guards</td> </tr> <tr> <td>E. Manual work</td> <td>Limited oral and written consumption and production</td> <td>E1. Skilled manual work</td> <td>Plumber</td> </tr> <tr> <td> </td> <td> </td> <td>E2. Unskilled manual work</td> <td>Janitor</td> </tr> </tbody> </table> <p> </p>
Symbol Representation of the Three Gluon Form Factor in N=4 Planar Super Yang-Mills Theory
<p>Datasets describing the symbol of the three-gluon form factor in N=4 planar super Yang-Mills theory, generated using the amplitude bootstrap approach. The file "EZ_symb_new_norm" contains the symbol form of this quantity at 1 through 5 loops of precision, while the file "EZ6_symb_new_norm" contains the symbol at 6 loops. The file "EZ_symb_quad_new_norm" contains the symbol at 1 through 6 loops in compressed "quad" form, where the final-entry conditions described in (https://arxiv.org/pdf/2204.11901) are used to dramatically reduce the total number of terms in the symbol. The file "EZ7_symb_quad_new_norm" contains the symbol at 7 loops in the "quad" form. </p> <p>The tag “new_norm” refers to the fact that in the symbols given here, the letters a,b,c are defined by a = sqrt(u/(v*w)), b = sqrt(v/(w*u)), c = sqrt(w/(u*v)), as in arXiv:2405.06107, in order to make all coefficients integers. In contrast, in arXiv:2204.11901, the letters a,b,c were defined by a = u/(v*w), b = v/(w*u), c = w/(u*v).</p> <p>In addition to the funding sources listed, MW was supported by research grant 00025445 from Villum Fonden.</p>
CLDF dataset derived from the Johansson et al.'s "The typology of sound symbolism" from 2020
<p>Cite the source of the dataset as:</p> <blockquote> <p>Erben Johansson, N., Anikin, A., Carling, G., & Holmer, A. (2020). The typology of sound symbolism: Defining macro-concepts via their semantic and phonetic features, Linguistic Typology , 24(2), 253-310. doi: https://doi.org/10.1515/lingty-2020-2034</p> </blockquote>
Thinking about Vector Symbolic Architectures (video recording)
<p>Video recording of the keynote presentation "Thinking about Vector Symbolic Architectures" given on 2023-06-15 at the <a href="https://sites.google.com/ltu.se/midnightvsa/home?authuser=0">Midnight Sun Workshop on Vector Symbolic Architectures</a> in Luleå, Sweden.</p> <p><strong>Abstract</strong></p> <p>Vector Symbolic Architectures are defined in terms of a very small set of operators acting on a vector space. The task of the VSA researcher is to discover the implications that follow from the definition in terms of the systems that can be implemented with VSAs. The VSA definitions are the researcher’s raw materials, but they also need tools to transform those raw materials into useful hypotheses and system designs. One important tool for a researcher is a conceptual framework, which specifies how the researcher thinks about VSAs and relates them to the other things they know. It is the researcher’s mental model of how VSAs work. The primary requirement for a conceptual framework is that it is productive; it should make it easy for the researcher to generate interesting hypotheses and designs. These hypotheses and designs don’t have to be correct, just plausible. Beating them into shape is a different part of the research process. Most VSA research papers contain a statement of the VSA definition. Very few mention the researcher’s conceptual framework. In this talk I will sketch out my conceptual framework - how I think about Vector Symbolic Architectures - in the hope that it might be interesting and useful to other researchers.</p>
Handedness and Symbolic Number Representation
Open the record for dataset details and reuse information.
FIG. 6 in Plate f of the Gundestrup "cauldron": symbols of spring and fertility
FIG. 6. — Bird helmet from Ciumeşti, Romania, second century BC, with zygo- dactylic feet (from Rusu 1971).
FIG. 1 in Plate f of the Gundestrup "cauldron": symbols of spring and fertility
FIG. 1. — The Gundestrup "cauldron", plate f is facing (© National Museum of Denmark, photo Kim Bach).
OTMM Symbolic Section Dataset
<p>otmm_symbolic_section_dataset</p> <p>The section test dataset of music scores of Ottoman-Turkish makam music</p> <p>This repository contains the audio section annotations and the scores used in the paper:</p> <blockquote> <p>Şentürk, S., & Serra X. (2016). A method for structural analysis of Ottoman-Turkish makam music scores. In Proceedings of 6th International Workshop on Folk Music Analysis (FMA 2016), (pp. 39–46)., Dublin, Ireland.</p> </blockquote> <p>Please cite the publication above in any work using this dataset.</p> <p>The repository contains a test dataset of the SymbTr scores of 23 vocal compositions in the şarkı form and 42 instrumental compositions in peşrev and sazsemaisi forms in the txt and pdf format. The scores are selected from the release version 2.4.2 and the section annotations done by the first author of the paper. In the sections folder, the annotated sections for each score are stored in a csv file, which has the same name as the SymbTr-name (makam--form--usul--name--composer) of the annotated score. The fields are:</p> <p>start_note: The starting note index in the SymbTr-txt score end_note: The ending note index in the SymbTr-txt score name: The basic semantic name of the section. Right now, it is the name (TESLİM, ARANAĞME...) annotated in the score for instumental sections or "VOCAL_SECTION" for vocal sections. melodic_structure: Melodic semiotic label lyric_structure: Lyrical semiotic label lyrics: Lyrics of the section slug: The processed version of "name" field with the Turkish characters and special characters handled.</p> <p>For the details of the dataset, please refer to the paper. For any further information please contact the authors.</p>
Turkish Makam Symbolic Phrase Segmentation Dataset
<p>makam-symbolic-phrase-segmentation-dataset</p> <p><strong>Data sets containing pieces in SymbTr2 format segmented into phrases</strong></p> <p>This study presents a large machine-readable dataset of Turkish makam music scores segmented into phrases by experts of this music. The segmentation facilitates computational research on melodic similarity between phrases, and relation between melodic phrasing and meter, rarely studied topics due to unavailability of data resources. It consists of 31362 phrases on a set of 480 scores of different compositions annotated by 3 experts.</p> <p>Please refer to the following publication if you use this data in your research:</p> <blockquote> <p>M. K. Karaosmanoglu, B. Bozkurt, A. Holzapfel, N. D. Disiacik, A symbolic dataset of Turkish makam music phrases, Folk Music Analysis Workshop (FMA), Istanbul, 2014.</p> </blockquote> <p>The refactored code for automatic phrase segmentation can be found here.</p> <p>For other deliverables of the paper please visit: http://www.rhythmos.org/shareddata/turkishphrases.html</p>
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
<p>We introduce <strong>PDMX</strong>: a <strong>P</strong>ublic <strong>D</strong>omain <strong>M</strong>usic<strong>X</strong>ML dataset for symbolic music processing. Refer to our <a title="PDMX Paper" href="https://arxiv.org/abs/2409.10831" target="_blank" rel="noopener">paper</a> for more information, and our <a title="PDMX GitHub Repository" href="https://github.com/pnlong/PDMX/" target="_blank" rel="noopener">GitHub repository</a> for any code-related details. Please cite both our paper and <a href="https://arxiv.org/abs/2410.02084" target="_blank" rel="noopener">our collaborators' paper</a> if you use this dataset (see our GitHub for more information).</p> <p>Upon further use of the PDMX dataset, we discovered a discrepancy between the public-facing copyright metadata on the <a href="https://musescore.com/">MuseScore website</a> and the internal copyright data of the MuseScore files themselves, which affected 31,221 (12.29% of) songs. We have decided to proceed with the former given its public visibility on Musescore (i.e. this is what the MuseScore website presents its users with). We have noted files with conflicting internal licenses in the <em><strong>license_conflict</strong></em> column of PDMX. We recommend using the <em><strong>no_license_conflict</strong></em> subset of PDMX (which still includes 222,856 songs) moving forward.</p> <p>Additionally, for each song in PDMX, we not only provide the <em>MusicRender</em> and metadata JSON files, but we also try to include the associated compressed MusicXML (MXL), sheet music (PDF), and MIDI (MID) files when available. Due to the corruption of 42 of the original MuseScore files, these songs lack those associated files (since they could not be converted to those formats) and only include the <em>MusicRender</em> and metadata JSON files. The <em><strong>all_valid</strong></em> subset of PDMX describes the songs where all associated files are valid.</p>
Fig.ç3.R elationships of (A) barbel length and (B) pectoral- n length to standard length in Upeneus guttatus (closed symbols: stars from Japan, squares from Indo–West Paci c) and U. japonicus (open circles). in First Records of the Two-tone Goatfish, Upeneus guttatus, from Japan, and Comparisons with U. japonicus (Perciformes: Mullidae)
Fig.ç3.R elationships of (A) barbel length and (B) pectoral- n length to standard length in Upeneus guttatus (closed symbols: stars from Japan, squares from Indo–West Paci c) and U. japonicus (open circles).
Fig. 10. Distribution records for Epicharis cockerelli Friese, 1900, E. duckei Friese, 1901 and E. iheringi Friese, 1899. Symbols with a in Taxonomic revision of the oil-collecting bee subgenus Epicharis (Epicharitides) Moure, 1945 (Hymenoptera: Apidae), with the description of two new species
Fig. 10. Distribution records for Epicharis cockerelli Friese, 1900, E. duckei Friese, 1901 and E. iheringi Friese, 1899. Symbols with a cross represent records from literature.
Datasets for "Machine-Learning-Enhanced Symbolic Regression for Methane Storage Prediction in Covalent Organic Frameworks"
<p>This collection contains the datasets and associated files used in the research presented in the manuscript titled "Machine Learning-Enhanced Symbolic Regression for Methane Storage Prediction in Covalent Organic Frameworks". The datasets are critical for the development and validation of machine learning and symbolic regression models aiming to predict methane storage capacities in covalent organic frameworks (COFs).</p> <p><strong>Included Datasets:</strong></p> <ol> <li><code>COF_Data_for_ML.csv</code>: This dataset was utilized for the development of machine learning models.</li> <li><code>COF_Data_for_SISSO.csv</code>: This dataset was employed for the development of SISSO-based symbolic regression models.</li> <li><code>ML_vs_GCMC.xlsx</code>: This comparative dataset features GCMC-calculated results alongside machine learning predictions.</li> <li><code>Feature_Combination.xlsx</code>: This file contains data detailing all the feature combinations explored in the study.</li> <li><code>ML_SISSO_GCMC.xlsx</code>: This comparative dataset includes GCMC calculations, SISSO-based symbolic regression model predictions, and ML predictions.</li> <li><code>Crystallographic_Properties_of_535k_COFs.xlsx</code>: This consolidated dataset presents the crystallographic properties of 535,293 COFs.</li> </ol> <p><strong>Software Used:</strong></p> <ul> <li>Machine Learning Computations: Scikit-Learn (<a href="https://scikit-learn.org/stable/" target="_new">https://scikit-learn.org/stable/</a>)</li> <li>GCMC Simulations: RASPA2 (<a href="https://github.com/iRASPA/RASPA2" target="_new">https://github.com/iRASPA/RASPA2</a>)</li> <li>SISSO Calculations: SISSO toolkit (<a href="https://github.com/rouyang2017/SISSO" target="_new">https://github.com/rouyang2017/SISSO</a>)</li> <li>Crystallographic property calculations: Zeo++ (<a href="https://www.zeoplusplus.org/" target="_new">https://www.zeoplusplus.org/</a>)</li> </ul> <p>The datasets are provided to enable replication of the study's findings, encourage further research in the field, and facilitate the development of advanced predictive models by the scientific community. Researchers who use these datasets are requested to cite this Zenodo entry as well as the associated paper upon its publication.</p>
Derby database for mapping secondary to primary HGNC gene symbols
<p>The datasets (hgnc_complete_set and withdrawn) used to create this ID mapping database were downloaded from HGNC (<em>HUGO Gene Nomenclature Committee at the European Bioinformatics Institute, </em>website URL: https://www.genenames.org/) on 09/05/2022. </p> <p>This database was used for the <a href="https://github.com/tabbassidaloii/BridgeDbDemoBioSB2022">BridgeDb demo at BioSB 2022</a> conference.</p> <p>The scripts used to create this database based on HGNC: https://github.com/tabbassidaloii/create-bridgedb-secondary2primary</p> <p>This work was funded by the <a href="https://fairplus-project.eu/">FAIRplus project</a> (grant agreement no 802750) and <a href="https://www.nwo.nl/en/researchprogrammes/open-science/open-science-fund/open-science-fund-2021-awarded-grants">NWO Open Science Fund</a> (grant no <a href="https://www.nwo.nl/en/projects/203001121">203.001.121</a>).</p>
Text-fig. 3. Geology of the Cheringoma Plateau, Mozambique. Sections and geological map adapted from Tinley (1977). The star symbols close to Mhengere Hill represent fossil wood and stem sites. Note that the fault relationships proposed in the northernmost Inhaminga section require re-examination. The Nguere Hills were called Gadjiua by Tinley (1977). in Stratigraphy, Chronology And Palaeontology Of The Tertiary Rocks Of The Cheringoma Plateau, Mozambique
Text-fig. 3. Geology of the Cheringoma Plateau, Mozambique. Sections and geological map adapted from Tinley (1977). The star symbols close to Mhengere Hill represent fossil wood and stem sites. Note that the fault relationships proposed in the northernmost Inhaminga section require re-examination. The Nguere Hills were called Gadjiua by Tinley (1977).
Text-fig. 1. a: Po Plain and foothills of the Northern Apennine in Northern Italy (inset) with the location of Oriolo (black star) and other Early and Middle Pleistocene plant localities, Enza and Stirone. Red lines indicate the frontal thrust arcs (modified from Martinetto et al. 2015). b: The "La Salita" section, Oriolo and chronology of the two "Sabbie gialle" cycles based on large mammals and palaeomagnetic correlation (modified from Toniato et al. 2017; IMMS 2020* [Italian Mediterranean Marine Stages] updated from Cohen and Gibbars 2020; GTS 2021* [Global Time Scale] updated from Head et al. 2021). c: Quarry "La Salita", Oriolo, in 1987. Main unconformities (U) separating the two "Sabbie gialle" cycles and terrestrial deposits on top are shown. Leaf symbols indicate the positions of some of the layers rich in fossil leaves (photo by G. B. Vai, modified). d: Surroundings of Faenza with the location of Oriolo and adjacent coeval sites yielding plant macrofossils. in The Late Early Pleistocene Flora Of Oriolo, Faenza (Italy): Assembly Of The Modern Forest Biome
Text-fig. 1. a: Po Plain and foothills of the Northern Apennine in Northern Italy (inset) with the location of Oriolo (black star) and other Early and Middle Pleistocene plant localities, Enza and Stirone. Red lines indicate the frontal thrust arcs (modified from Martinetto et al. 2015). b: The "La Salita" section, Oriolo and chronology of the two "Sabbie gialle" cycles based on large mammals and palaeomagnetic correlation (modified from Toniato et al. 2017; IMMS 2020* [Italian Mediterranean Marine Stages] updated from Cohen and Gibbars 2020; GTS 2021* [Global Time Scale] updated from Head et al. 2021). c: Quarry "La Salita", Oriolo, in 1987. Main unconformities (U) separating the two "Sabbie gialle" cycles and terrestrial deposits on top are shown. Leaf symbols indicate the positions of some of the layers rich in fossil leaves (photo by G. B. Vai, modified). d: Surroundings of Faenza with the location of Oriolo and adjacent coeval sites yielding plant macrofossils.
Text-fig. 1. Locality map with the approximate extent of Clarkia Lake during Miocene times in what is today northern Idaho, USA. Black dots mark three of the localities yielding the Miocene Clarkia flora; the fossil leaf of Nymphaea sp. described here comes from locality P-33. Other symbols: Dashed lines for county boundaries; a thin dotted line for Idaho State Hwy 3; a triangle for the local peak of Bechtel Butte; and a star for the town of Clarkia. Inset: Location of the map in northern Idaho. Abbreviations: WA – Washington state, OR – Oregon, ID – Idaho, MT – Montana. Map redrawn from Ladderud et al. (2015). in First Water Lily, A Leaf Of Nymphaea Sp., From The Miocene Clarkia Flora, Northern Idaho, Usa: Occurrence, Taphonomic Observations, Floristic Implications
Text-fig. 1. Locality map with the approximate extent of Clarkia Lake during Miocene times in what is today northern Idaho, USA. Black dots mark three of the localities yielding the Miocene Clarkia flora; the fossil leaf of Nymphaea sp. described here comes from locality P-33. Other symbols: Dashed lines for county boundaries; a thin dotted line for Idaho State Hwy 3; a triangle for the local peak of Bechtel Butte; and a star for the town of Clarkia. Inset: Location of the map in northern Idaho. Abbreviations: WA – Washington state, OR – Oregon, ID – Idaho, MT – Montana. Map redrawn from Ladderud et al. (2015).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.