Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
130
datasets available to search
ShareScore release 0.7.1
Dataset results
130 results for “categorization”
APRSuite: A Suite of Components and Use Cases Based on Categorical Decomposition of Automatic Program Repair Techniques and Tools
<p><strong>During the last decade, we are witnessing the advent of a proliferation of techniques and associated tools for automatic program repair (APR). The current techniques and tools provide rich sources of knowledge that should be taken into consideration for future research. An overview of the current APR techniques and tools can serve the research community as a knowledge accumulator. However, APR techniques and tools differ in many aspects making knowledge accumulation challenging. To overcome this challenge, in this paper, we propose to leverage common components that constitute the APR techniques and tools. To achieve this objective, we surveyed current APR techniques and tools to identify the APR Suite of common constituent components, namely as APRSuite. Repair source and defect class are examples of identified components. We grouped these components into several categories such as patch evaluation and target defects. We have also identified some of the possible use cases per component as well as different lessons learned in studies for each component and for each use case. In addition, we developed a principled way for application of the components. The <em>APRSuite</em> and the <em>principled way</em> to apply it comprise a <em>framework</em> for knowledge accumulation, evaluation, and comparison of APR techniques and tools. The novelty of our work lies in its original viewpoint to the process of literature review in the APR research field. To demonstrate the applicability of the framework, we mapped out several concrete APR techniques, as a first instantiation of the framework. We observed that the framework brings discipline into the evaluation and/or comparison of APR techniques and tools. The framework offers these benefits objectively and systematically. We concluded that knowledge accumulation and characterization through literature reviews can be therefore facilitated through the identified suite of components while at the same time the existing component suite can be modified, augmented, or improved.</strong></p>
Categorization of articles 2017 with authorship of Pontificia Universidad Católica de Chile, through the SDGs
<p>The dataset comprises a single list of publications exported from Web of Science (Clarivate Analytics) and Scopus (Elsevier) databases, to which a process was applied that eliminated duplicate records. 2,379 scientific publications in English or Spanish of the "Article" type from the year 2017, with authorship associated with the Pontificia Universidad Católica de Chile, were considered.</p> <p>In addition, the classification process carried out by the team of specialists that considered three consecutive milestones is included: establishment of the reading level applied to each publication record; assignment of one of the 18 categories identified in the information analysis, which include the 17 SDGs and the option "Unclassified" and one of the 169 subcategories corresponding to the goals; and, finally, the status of the review process carried out.</p>
Expectation-Maximization enables phylogenetic dating under a Categorical Rate Model
<div> <div> <div> <p>Dating phylogenetic trees to obtain branch lengths in the unit of time is essential for many downstream applications but has remained challenging. Dating requires inferring mutation rates that can change across the tree. While we can assume to have information about a small subset of nodes from the fossil record or sampling times (for fast-evolving organisms), inferring the ages of the other nodes essentially requires extrapolation and interpolation. Assuming a clock model that defines a distribution over rates, we can formulate dating as a constrained maximum likelihood (ML) estimation problem. While ML dating methods exist, their accuracy degrades in the face of model misspecification where the assumed parametric statistical clock model vastly differs from the true distribution. Notably, existing methods tend to assume rigid, often unimodal rate distributions. A second challenge is that the likelihood function involves an integral over the continuous domain of the rates and often leads to difficult non-convex optimization problems. To tackle these two challenges, we propose a new method called Molecular Dating using Categorical-models (MD-Cat). MD-Cat uses a categorical model of rates inspired by non-parametric statistics and can approximate a large family of models by discretizing the rate distribution into k categories. Under this model, we can use the Expectation-Maximization (EM) algorithm to co-estimate rate categories and branch lengths in the time unit. Our model has fewer assumptions about the true clock model than parametric models such as Gamma or LogNormal distribution. Our results on two simulated and real datasets of Angiosperms and HIV and a wide selection of rate distributions show that MD-Cat is often more accurate than the alternatives, especially on datasets with nonmodal or multimodal clock models.</p> </div> </div> </div>
Dataset: Invariant categorical color regions across illuminant change coincide with focal colors
<p>Experimental data for a paper titled "Invariant categorical color regions across illuminant change coincide with focal colors“ published in Journal of Vision.</p> <p>This repository contains following data. Please refer to the article for details about the dataset.</p> <p> </p> <p><strong>1. IlluminantsANDReflectances.xlsx</strong></p> <p>This file stores spectral radiance for each illuminant and OSA coordinates (jgL) and reflectance values for 424 OSA color samples.</p> <p> </p> <p><strong>2. Exp1_1_RawData</strong></p> <p>Raw data for Experiment 1-1.</p> <p>This folder has xlsx files, and filename indicates the illuminant condiion.</p> <p>In each xlsx file, different sheets contain a different observer’s data.</p> <p>The number (1-424) corresponds to the number of OSA sample (shown in IlluminantsANDReflectances.xlsx).</p> <p> </p> <p><strong>3. Exp1_2_RawData</strong></p> <p>Raw data for Experiment 1-2. The same format as Exp1_1_RawData.</p> <p> </p> <p><strong>4. Exp2_RawData.xlsx</strong></p> <p>Raw matching data for Experiment 2.</p> <p>The name of each sheet shows the observer & illuminant condition.</p> <p>Each row shows Ljg coordinates of the OSA sample and associated matching result (in RGB and xyL).</p>
Defining Categorical Reasoning of Numerical Feature Models with Feature-Wise and Variant-Wise Quality Attributes
<p><strong>To watch it in Youtube:</strong></p> <p><a href="https://youtu.be/Uq2qtb4_K2U">https://youtu.be/Uq2qtb4_K2U</a></p> <p><strong>This is a pre-print, please access and cite the published version:</strong></p> <p><a href="https://doi.org/10.1145/3503229.3547057">https://doi.org/10.1145/3503229.3547057</a></p> <p>Automatic analysis of variability is an important stage of <em>Software Product Line</em> (SPL) engineering. Incorporating quality information into this stage poses a significant challenge. However, quality-aware automated analysis tools are rare, mainly because in existing solutions variability and quality information are not unified under the same model.</p> <p>In this paper, we make use of the <em>Quality Variability Model</em> (QVM), based on <em>Category Theory</em> (CT), to redefine reasoning operations. We start defining and composing the six most common operations in SPL, but now as quality-based queries, which tend to be unavailable in other approaches. Consequently, QVM supports interactions between variant-wise and feature-wise quality attributes. As a proof of concept, we present, implement and execute the operations as lambda reasoning for CQL IDE -- the state-of-the-art CT tool.</p>
Ontology metadata categorization
<p>This file summarizes an analysis aimed at finding the support for each MOD metadata category by existing ontologies and vocabularies. For each property, we map it against the most relevant MOD metadata category to find their support.</p>
Overcoming the pitfalls of categorizing continuous variables in ecology and evolutionary biology
<ol> <li><span>Many metrics in biological research – from body size to life history timing to environmental metrics – are measured continuously (e.g., body size in grams) but analyzed as categories (e.g., large versus small). The pitfalls of categorization are well-recognized in statistics, but many scientists in the fields of ecology, evolution, and behavior may not be aware of this literature. These fields lack a review of common examples and feasible solutions to avoid the hazards of categorizing continuous data. </span></li> <li><span>Our goal was to summarize current practices of categorizing continuous predictors in ecology and evolutionary biology and provide guidance for overcoming those pitfalls. We conducted a mini-review of 72 recent publications in six popular journals to quantify the prevalence of categorization. We then summarized commonly categorized metrics and simulated a dataset to demonstrate the drawbacks of categorization using common metrics and realistic examples from ecology and evolutionary biology. </span></li> <li><span>We show that categorizing continuous variables is common (31% of publications reviewed), especially in the animal behavior field, and underscore that predictor variables – including abiotic, morphological, physiological, behavioral, and demographic metrics – can and should be collected and analyzed continuously. Our analysis of the simulated field dataset demonstrates how categorizing continuous variables can lower statistical power and change interpretation, especially when arbitrary breakpoints are used. Finally, we provide recommendations on how to keep variables continuous throughout the entire scientific process. </span></li> <li><span>Together, these pieces comprise an actionable guide to increasing statistical power and facilitating large synthesis studies by simply leaving continuous variables alone. Overcoming the pitfalls of categorizing continuous variables will allow ecologists and evolutionary biologists to continue making trustworthy conclusions about natural processes, along with predictions about their responses to climate change and other environmental contexts. We hope that this manuscript and its associated code will provide a useful lab practical for students and teachers to develop programming skills including data simulation, plotting, and model comparisons, as well as research skills including reporting and interpretation. </span></li> </ol>
Overcoming the pitfalls of categorizing continuous variables in ecology, evolution, and behavior
Open the record for dataset details and reuse information.
Expectation-Maximization enables phylogenetic dating under a Categorical Rate Model
Open the record for dataset details and reuse information.
The color communication game: how categorical understanding of colors can be shown without considering color naming data
Open the record for dataset details and reuse information.
A dataset for evaluating one-shot categorization of novel object classes
<p>From just a single example, we can derive quite precise intuitions about what other class members look like. This stands in stark contrast to machine learning algorithms, which typically require tens or even hundreds of thousands of examples to learn a new category. One of the most important open questions in our field is: How do humans achieve this? The stimuli and data provided here (in MATLAB format) are from thousands of crowd-sourced human responses to novel objects. The data can be used to test machine learning generalization as compared to human and also can be used as a test bed for various kinds of category learning models. </p>
Categorical versus geometric morphometric approaches to characterising the evolution of morphological disparity in Osteostraci (Vertebrata, stem-Gnathostomata)
Morphological variation (disparity) is almost invariably characterised by two non-mutually exclusive approaches: (i) quantitatively, through geometric morphometrics, and (ii) in terms of discrete, 'cladistic', or categorical characters. Uncertainty over the comparability of these approaches diminishes the potential to obtain nomothetic insights into the evolution of morphological disparity and the few benchmarking studies conducted so far show contrasting results. Here, we apply both approaches to characterising morphology in the stem-gnathostome clade Osteostraci in order to assess congruence between these alternative methods as well as to explore the evolutionary patterns of the group in terms of temporal disparity and the influence of phylogenetic relationships and habitat on morphospace occupation. Our results suggest that both approaches yield similar results in morphospace occupation and clustering, but also some differences indicating that these metrics may capture different aspects of morphology. Phylomorphospaces reveal convergence towards a generalised 'horseshoe'-shaped cranial morphology and two strong trends involving major groups of osteostracans (benneviaspidids and thyestiids), which probably reflect adaptations to different lifestyles. Temporal patterns of disparity obtained from categorical and morphometric approaches appear congruent, however disparity maxima are recorded at very different times in the evolutionary history of the group when increasing the number of taxa and characters in the categorical dataset. The results of our analyses indicate that categorical and continuous datasets may characterize different patterns of morphological disparity and that discrepancies could reflect preservational limitations of morphometric data and differences in the potential of each data type for characterizing more or less inclusive aspects of overall phenotype.
Branch Policies in CI/CD: Categorization, Adoption, and Usage - Supplemental for replicability
Open the record for dataset details and reuse information.
Behavioral data associated with "Passive exposure to task-relevant stimuli enhances categorization learning"
<p>Behavioral data associated with Schmid et al. (2023) "<i>Passive exposure to task-relevant stimuli enhances categorization learning</i>", and example code for loading these data. See README.md for details.</p><p> </p>
Young and older adult vowel categorization responses
<p>Age-related changes in auditory processing may reduce physiological coding of acoustic cues, contributing to older adults' difficulty perceiving speech in background noise. This study investigated whether older adults differed from young adults in patterns of acoustic cue weighting for categorizing vowels in quiet and in noise. All participants relied primarily on spectral quality to categorize /Ꜫ/ and /æ/ sounds in both listening conditions. However, relative to young adults, older adults exhibited greater reliance on duration and less reliance on spectral quality. These results suggest that aging alters patterns of perceptual cue weights that may influence speech recognition abilities.</p>
European Starling categorical perception chronic ephys and behavior dataset
<p>This dataset corresponds to the currently in-press paper "Expectation-driven sensory adaptations support enhanced acuity during categorical perception" in Nature Neuroscience. </p> <p> </p> <p>This dataset corresponds to the code at <a href="https://github.com/timsainb/cdcp_paper">https://github.com/timsainb/cdcp_paper</a></p> <p>Please refer to the readme for this GitHub repo, which contains all the necessary information for reproducing our analyses or using this data for additional analyses. </p> <p> </p> <p> </p> <pre> </pre>
Cooperative cortical network for categorical processing of Chinese lexical tone
<p><strong>This dataset contains the ECoG, CT and MRI data for the six subjects associated with the manuscript, "Cooperative cortical network for categorical processing of Chinese lexical tone", as well as stimulus sound files used in the study.</strong></p> <p>data_MRI.rar - MRI scan before the ECoG electrode implantation</p> <p>data_CT.rar - CT images with ECoG electrode implanted</p> <p>data_ECoG.rar - ECoG recording of six patients reported in the manuscript</p> <p>stimuli_BehaviorContinumm.rar - Chinese tone continuum stimuli (Token 1-13 as in Fig 1 of the manuscript)</p> <p>stimuli_ECoG_MMN.rar - Chinese tone stimuli used for oddball paradigm in ECoG experiment (Token 2, 5 and 8, see Fig 1 of the manuscript)</p> <p>Upon email request (hongbo@tsinghua.edu.cn), the authors may provide analysis code. </p> <p> </p> <p><strong>Original manuscript:</strong></p> <p>Si, Zhou and Hong. <em>PNAS</em>, 2017</p> <p><strong>Cooperative cortical network for categorical processing of Chinese lexical tone</strong></p> <p><strong>Abstract:</strong> In tonal languages such as Chinese, lexical tone with varying pitch contours serves as a key feature to provide contrast in word meaning. Similar to phoneme processing, behavioral studies have suggested that Chinese tone is categorically perceived. However, its underlying neural mechanism remains poorly understood. By conducting cortical surface recordings in surgical patients, we revealed a cooperative cortical network along with its dynamics responsible for this categorical perception. Based on an oddball paradigm, we found amplified neural dissimilarity between cross-category tone pairs, rather than between within-category tone pairs, over cortical sites covering both the ventral and dorsal streams of speech processing. The bilateral superior temporal gyrus (STG) and the middle temporal gyrus (MTG) exhibited increased response latencies and enlarged neural dissimilarity, suggesting a ventral hierarchy that gradually differentiates the acoustic features of lexical tones. In addition, the bilateral motor cortices were also found to be involved in categorical processing, interacting with both the STG and the MTG and exhibiting a response latency in between. Moreover, the motor cortex received enhanced Granger causal influence from the semantic hub, the anterior temporal lobe, in the right hemisphere. These unique data suggest that there exists a distributed cooperative cortical network supporting the categorical processing of lexical tone in tonal language speakers, not only encompassing a bilateral temporal hierarchy that is shared by categorical processing of phonemes but also involving intensive speech-motor interactions over the right hemisphere, which might be the unique machinery responsible for the reliable discrimination of tone identities.</p>
InCLosure Code for Real Automatic Differentiation: Supplementary Material for Article "A Consistent and Categorical Axiomatization of Differentiation Arithmetic Applicable to First and Higher Order Derivatives"
<p>InCLosure Code for Real Automatic Differentiation: Supplementary Material for Article "A Consistent and Categorical Axiomatization of Differentiation Arithmetic Applicable to First and Higher Order Derivatives", Punjab University Journal of Mathematics, October 2019. Download latest release of InCLosure via <a href="https://doi.org/10.5281/zenodo.2702404">https://doi.org/10.5281/zenodo.2702404</a></p>
Dataset for Categorizing Unicode and Non-Unicode Transmitted Via H.323 Protocol
<p>To secure data transmitted through the H.323 protocol, a new data security algorithm has been introduced. This algorithm combines hybrid cryptography and real-time steganography. The hybrid cryptography includes a multilayer encryption algorithm and the RSA algorithm, while steganography involves embedding encrypted data in the 3rd LSBs of video frame pixels. To ensure data integrity and assess the reliability of the proposed algorithm using a Machine Learning model, a dataset containing 50,523 rows and 10 columns was generated. These rows and columns contain information about 50,523 Unicode and Non-Unicode characters (including all special characters) from around 159 languages. Columns were created by following the steps of the multi-layer symmetric cryptography algorithm. This means that the dataset was created by encrypting each of the 50,523 characters separately using the multi-layer cryptography algorithm and then filling up the 10 columns of each row with the encrypted result. The dataset features include Unicode_Numbers, Decimal_Values, Hex_Addition_Result, Decimal_of_Hex_Operation, XORed_Values, Linear_Arithmetic_Result_1, Linear_Arithmetic_Result_2, Type, Character_Type_in_Numerals, and Class, with Class being the dependent variable. The 'Type' column indicates the character type based on Unicode and Non-Unicode distinctions, while the 'Character_Type_in_Numerals' column contains a numeric float value based on each character's type. Class 0 is assigned to Non-Unicode characters, and Class 1 is assigned to Unicode characters. To obtain the encrypted result, a key generation algorithm creates 7 corresponding keys from a 256-bit randomly generated hexadecimal string for use in multi-layer encryption.</p>
Categorization of Decentralized Autonomous Organizations in the Aragon platform
<p>Dataset result of the research paper "A Categorization of Decentralized Autonomous Organizations: The Case of the Aragon Platform" published in <em>IEEE Transactions on Computational Social Systems</em>. The paper proposes an empirically grounded categorization of DAOs in terms of their operative domain, purpose, scope, voting processes, and use of crypto-tokens. The categorization was applied to 40 DAO communities hosted in the Aragon platform, analyzing 15 dimensions in each of them. Further details on the variables and the annotation process are described in the paper.</p> <p>Recommended citation for the article: </p> <p>Peña-Calvin, A., Saldivar, J., Arroyo, J., & Hassan, S. (2023). A Categorization of Decentralized Autonomous Organizations: The Case of the Aragon Platform. <em>IEEE Transactions on Computational Social Systems</em>, doi: 10.1109/TCSS.2023.3299254</p> <p>WHAT'S NEW IN THIS VERSION</p> <p>- Readme file. It includes detailed information on the content of the rest of the files.</p> <p>- Erratum corrected in file Dataset - Tokens.csv. Token NU was labelled as TOK-4 and now is labelled as TOK-3.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.