Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

28

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

28 results for “Categorical data”

Learn how ShareScore rates datasets ↗
edi48/100

CoRRE Trait Data: A collection of 17 categorical and continuous traits for more than 4000 grassland species worldwide

In our changing world, it is critical to understand and predict plant community responses to global change drivers. Plant functional traits promise to be a key predictive tool for many ecosystems, including grasslands, however their use requires both complete plant community and functional trait data. Yet, representation of these data in global databases is incredibly sparse, particularly beyond a handful of most used traits and common species. Here we present the CoRRE Trait Database, spanning 17 traits (9 categorical, 8 continuous) anticipated to predict species’ responses to global change for 4,079 vascular plant species across 173 plant families present in 390 grassland experiments from around the world. The database contains complete categorical trait records for all 4,079 plant species, obtained from a comprehensive literature search. Additionally, the database contains nearly complete coverage (99.97%) of species mean values for continuous traits for a subset of 2,927 plant species, predicted from observed trait data drawn from TRY and a variety of other plant trait databases using Bayesian Probabilistic Matrix Factorization (BHPMF) and multivariate imputation using chained equations (MICE). These data will shed light on mechanisms underlying population, community, and ecosystem responses to global change in grasslands worldwide.

openCC BYMay 2024View details →
zenodo44/100

Raw data of the study: Categorizing urban avoiders, utilizers, and dwellers for identifying bird conservation priorities in a northern Andean city

<p>This datasheet contains raw data on bird count records made from 2016 and 2019. Data were taken in urban and adjacent non-urban areas of Medell&iacute;n, Colombia. It was part of a collaborative sampling effort during environmental assessments and personal research, summarizing systematic information on 139 sampling points (124 within the city and 15 in adjacent non-urban areas). All points were sampled under the same protocol in order to facilited data for research; in all cases, sampling was in charge of ornithologist with at least 4 years of previous experience in bird surveys. This protocol consisted in sampling during 10 minutes, four times per point (i.e., repetitions), using a fixed radius of 25 m.&nbsp;</p> <p>Information on bird surveys (Count_Data within the corresponding datasheet tab) contains the ID of each site; whether corresponded to a urban or non-urban site; in what category of urban development the site was located, based on 1000, 500 and 200 m buffers (from the observer during bird counts: moderate, low or high); the taxonomic information of each species (order, family, scientific name); the number of recorded individuals; &nbsp;the repetition or number of the visit (1, 2, 3, or 4); the name of the project; the name of the observer, and the date of sampling.&nbsp;</p> <p>Information on categorization of bird species (Categorization within the corresponding datasheet tab) represents additional information on altitudinal ranges, trophic guilds, distribution, and others. In addition, information on frequency for each bird species is given, according to the location of each sampling site and the way it was grouped. This information was the base for categorizing bird species as urban avoider, utilizer, or dweller, under the calculations and decision rules that are also given within the corresponding cells of the datasheet.</p> <p>Any further information or questions about this data could be ask directly, writing to the e-mails: jgarizabal@unal.edu.co or njmacer@unal.edu.co.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Accompanying simulated data for "Go multivariate: a Monte Carlo study of a multilevel hidden Markov model with categorical data of varying complexity"

<p>The multilevel hidden Markov model (MHMM) is a promising vehicle to investigate latent dynamics over time in social and behavioral processes. By including continuous individual random effects, the model accommodates variability between individuals, providing individual-specific trajectories and facilitating the study of individual differences. However, the performance of the MHMM has not been sufficiently explored. Currently, there are no practical guidelines on the sample size needed to obtain reliable estimates related to categorical data characteristics We performed an extensive simulation to assess the effect of the number of dependent variables (1-4), the number of individuals (5-90), and the number of observations per individual (100-1600) on the estimation performance of group-level parameters and between-individual variability on a Bayesian MHMM with categorical data of various levels of complexity. We found that using multivariate data generally alleviates the sample size needed and improves the stability of the results. Regarding the estimation of group-level parameters, the number of individuals and observations largely compensate for each other. Meanwhile, only the former drives the estimation of between-individual variability. We conclude with guidelines on the sample size necessary based on the complexity of the data and the study objectives of the practitioners.</p> <p>This repository contains data generated&nbsp;for the manuscript: &quot;Go multivariate: a Monte Carlo study of a multilevel hidden Markov model&nbsp;with categorical data of varying complexity&quot;. It comprehends: (1) model outputs (maximum a posteriori estimates) for&nbsp;each repetition (n=100) of&nbsp;each scenario (n=324) of the main simulation, (2) complete model outputs (including estimates for&nbsp;4000 MCMC iterations) for two chains of each&nbsp;repetition (n=3)&nbsp;of&nbsp;each scenario (n=324). Please note that the empirical data used in the manuscript&nbsp;is not available as part of this repository.&nbsp;A subsample of the data used in the empirical example are openly available as an example data set in the R package <a href="https://cran.r-project.org/web/packages/mHMMbayes/index.html">mHMMbayes on CRAN</a>. The full data set&nbsp;is available on request from the authors.</p>

opencc-by-4.0Mar 2022View details →
dryad40/100

The color communication game: how categorical understanding of colors can be shown without considering color naming data

Open the record for dataset details and reuse information.

publicOct 2024View details →
zenodo36/100

Behavioral data associated with "Passive exposure to task-relevant stimuli enhances categorization learning"

<p>Behavioral data associated with Schmid et al. (2023) "<i>Passive exposure to task-relevant stimuli enhances categorization learning</i>", and example code for loading these data. See README.md for details.</p><p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
dryad32/100

Data from: Control of adaptive action selection by secondary motor cortex during flexible visual categorization

<p>Adaptive action selection during stimulus categorization is an important feature of flexible behavior. To examine neural mechanism underlying this process, we trained mice to categorize the spatial frequencies of visual stimuli according to a boundary that changed between blocks of trials in a session. Using a model with a dynamic decision criterion, we found that sensory history was important for adaptive action selection after the switch of boundary. Bilateral inactivation of the secondary motor cortex (M2) impaired adaptive action selection by reducing the behavioral influence of sensory history. Electrophysiological recordings showed that M2 neurons carried more information about upcoming choice and previous sensory stimuli when sensorimotor association was being remapped than when it was stable. Thus, M2 causally contributes to flexible action selection during stimulus categorization, with the representations of upcoming choice and sensory history regulated by the demand to remap stimulus-action association.</p>

opencc-zeroJul 2020View details →
dryad32/100

Data from: Rapid categorization of natural face images in the infant right hemisphere

Human performance at categorizing natural visual images surpasses automatic algorithms, but how and when this function arises and develops remain unanswered. We recorded scalp electrical brain activity in 4–6 months infants viewing images of objects in their natural background at a rapid rate of 6 images/second (6 Hz). Widely variable face images appearing every 5 stimuli generate an electrophysiological response over the right hemisphere exactly at 1.2 Hz (6 Hz/5). This face-selective response is absent for phase-scrambled images and therefore not due to low-level information. These findings indicate that right lateralized face-selective processes emerge well before reading acquisition in the infant brain, which can perform figure-ground segregation and generalize face-selective responses across changes in size, viewpoint, illumination as well as expression, age and gender. These observations made with a highly sensitive and objective approach open an avenue for clarifying the developmental course of natural image categorization in the human brain.

opencc-zeroDec 2014View details →
zenodo32/100

Accompanying simulated data for "Go multivariate: recommendations on multilevel hidden Markov models with categorical data of varying complexity"

<p>The multilevel hidden Markov model (MHMM) is a promising vehicle to investigate latent dynamics over time in social and behavioral processes. By including continuous individual random effects, the model accommodates variability between individuals, providing individual-specific trajectories and facilitating the study of individual differences. However, the performance of the MHMM has not been sufficiently explored. Currently, there are no practical guidelines on the sample size needed to obtain reliable estimates related to categorical data characteristics We performed an extensive simulation to assess the effect of the number of dependent variables (1-4), the number of individuals (5-90), and the number of observations per individual (100-1600) on the estimation performance of group-level parameters and between-individual variability on a Bayesian MHMM with categorical data of various levels of complexity. We found that using multivariate data generally alleviates the sample size needed and improves the stability of the results. Regarding the estimation of group-level parameters, the number of individuals and observations largely compensate for each other. Meanwhile, only the former drives the estimation of between-individual variability. We conclude with guidelines on the sample size necessary based on the complexity of the data and the study objectives of the practitioners.</p> <p>This repository contains data generated&nbsp;for the manuscript: &quot;Go multivariate: recommendations on multilevel hidden Markov models with categorical data of varying complexity&quot;. It comprehends: (1) model outputs (maximum a posteriori estimates) for&nbsp;each repetition (n=100) of&nbsp;each scenario (n=324) of the main simulation, (2) complete model outputs (including estimates for&nbsp;4000 MCMC iterations) for two chains of each&nbsp;repetition (n=3)&nbsp;of&nbsp;each scenario (n=324). Please note that the empirical data used in the manuscript&nbsp;is not available as part of this repository.&nbsp;A subsample of the data used in the empirical example are openly available as an example data set in the R package&nbsp;<a href="https://cran.r-project.org/web/packages/mHMMbayes/index.html">mHMMbayes on CRAN</a>. The full data set&nbsp;is available on request from the authors.</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Scalable mixed model approaches for set-based association studies on large-scale categorical data analysis and its application to 450k exome sequencing data in UK Biobank

<p>The ongoing release of large-scale sequencing data in the UK Biobank allows for identifying associations between rare variants and complex traits. SAIGE-GENE+ is a valid approach to conducting set-based association tests for quantitative and binary traits. However, for ordinal categorical phenotypes, applying SAIGE-GENE+ with treating the trait as quantitative or binarizing the trait can cause inflated type I error rates or power loss. In this study, we propose a novel method for rare-variant association tests, POLMM-GENE, in which a proportional odds logistic mixed model was used to characterize ordinal categorical phenotypes while adjusting for sample relatedness. POLMM-GENE fully utilizes the categorical nature of phenotypes and thus can well control type I error rates while remaining powerful. In the analyses of UK Biobank 450k whole exome-sequencing data for 5 ordinal categorical traits, POLMM-GENE identified 54 gene-phenotype associations.</p>

opencc-by-4.0Sep 2022View details →
zenodo32/100

FIGURE 3. Species Richness per Ecoregion. The richness data per ecoregion were categorized into classes with equal intervals, using 14 in Revealing the Baja California Peninsula's Hidden Treasures: An Annotated checklist of the native bees (Hymenoptera: Apoidea: Anthophila)

FIGURE 3. Species Richness per Ecoregion. The richness data per ecoregion were categorized into classes with equal intervals, using 14 breaks. However, the map displays only the eight categories where ecoregional richness is concentrated. Ecoregion: Coastal Sage Matorral (CSM); Chaparral (Ch); Baja California Mountains (BCM); Succulent Coastal Matorral (SCM); Lower Colorado Desert (LCD); Central Desert (CD); Vizcaíno Desert (VD); Gulf Coast (GC); La Giganta Ranges (GR); Magdalena Plains (MP); Tropical Dry Forest (TDF); Cape Mountains (CM); Sarcocaulescent Shrubland (SS).

opennotspecifiedOct 2024View details →
zenodo32/100

Data from: Sound categorization by crocodilians

<p>Dataset, acoustic signals and all original statistical codes used in the article &quot;Sound categorization by crocodilians&quot;.</p>

opencc-by-4.0Mar 2023View details →
dryad32/100

Data from: Control of adaptive action selection by secondary motor cortex during flexible visual categorization

Open the record for dataset details and reuse information.

publicJul 2020View details →
dryad32/100

Data for: Correlated evolution of categorical characters under a simple model

Open the record for dataset details and reuse information.

publicNov 2024View details →
dryad32/100

Data from: Rapid categorization of natural face images in the infant right hemisphere

Open the record for dataset details and reuse information.

publicMay 2016View details →
dryad32/100

Data from: Use and categorization of Light Detection and Ranging vegetation metrics in avian diversity and species distribution research

Open the record for dataset details and reuse information.

publicApr 2019View details →
dryad32/100

Data from: SpeciesGeoCoder: fast categorization of species occurrences for analyses of biodiversity, biogeography, ecology and evolution

Open the record for dataset details and reuse information.

publicJul 2016View details →
dryad28/100

Data from: Are categorical spatial relations encoded by shifting visual attention between objects?

Perceiving not just values, but relations between values, is critical to human cognition. We tested the predictions of a proposed mechanism for processing categorical spatial relations between two objects—the shift account of relation processing—which states that relations such as 'above' or 'below' are extracted by shifting visual attention upward or downward in space. If so, then shifts of attention should improve the representation of spatial relations, compared to a control condition of identity memory. Participants viewed a pair of briefly flashed objects and were then tested on either the relative spatial relation or identity of one of those objects. Using eye tracking to reveal participants' voluntary shifts of attention over time, we found that when initial fixation was on neither object, relational memory showed an absolute advantage for the object following an attention shift, while identity memory showed no advantage for either object. This result is consistent with the shift account of relation processing. When initial fixation began on one of the objects, identity memory strongly benefited this fixated object, while relational memory only showed a relative benefit for objects following an attention shift. This result is also consistent, although not as uniquely, with the shift account of relation processing. Taken together, we suggest that the attention shift account provides a mechanistic explanation for the overall results. This account can potentially serve as the common mechanism underlying both linguistic and perceptual representations of spatial relations.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Categorical colour perception occurs in both signalling and non-signalling colour ranges in a songbird

Although perception begins when a stimulus is transduced by a sensory neuron, numerous perceptual mechanisms can modify sensory information as it is processed by an animal's nervous system. One such mechanism is categorical perception, in which 1) continuously-varying stimuli are labelled as belonging to a discrete number of categories and 2) there is enhanced discrimination between stimuli from different categories as compared to equally-different stimuli from within the same category. We have shown previously that female zebra finches (Taeniopygia guttata) categorically perceive colours along an orange-red continuum that aligns with the carotenoid-based colouration of male beaks, a trait that serves as an assessment signal in female mate choice. Here we demonstrate that categorical perception occurs along a blue-green continuum as well, suggesting that categorical colour perception may be a general feature of zebra finch vision. Although we identified two categories in both the blue-green and the orange-red ranges, we also found that individuals could better differentiate colours from within the same category in the blue-green as compared to the orange-red range, indicative of less clear categorization in the blue-green range. We discuss reasons why categorical perception may vary across the visible spectrum, including the possibility that such differences are linked to the behavioural or ecological function of different colour ranges.

opencc-zeroDec 2018View details →
dryad28/100

Data from: Categorizing and assessing comprehensive drivers of provider behavior for optimizing quality of health care

<p>Inadequate quality of care in healthcare facilities is one of the primary causes of patient mortality in low- and middle-income countries, and understanding the behavior of healthcare providers is key to addressing it. Much of the existing research concentrates on improving resource-focused issues, such as staffing or training, but these interventions do not fully close the gaps in quality of care. By contrast, there is a lack of knowledge regarding the full contextual and internal drivers–such as social norms, beliefs, and emotions–that influence the clinical behaviors of healthcare providers. We aimed to provide two conceptual frameworks to identify such drivers, and investigate them in a facility setting where inadequate quality of care is pronounced. Using immersion interviews and a novel decision-making game incorporating concepts from behavioral science, we systematically and qualitatively identified an extensive set of contextual and internal behavioral drivers in staff nurses working in reproductive, maternal, newborn, and child health (RMNCH) in government public health facilities in Uttar Pradesh, India. We found that the nurses operate in an environment of stress, blame, and lack of control, which appears to influence their perception of their role as often significantly different from the RMNCH program's perspective. That context influences their perceptions of risk for themselves and for their patients, as well as self-efficacy beliefs, which could lead to avoidance of responsibility, or incorrect care. A limitation of the study is its use of only qualitative methods, which provide depth, rather than prevalence estimates of findings. This exploratory study identified previously under-researched contextual and internal drivers influencing the care-related behavior of staff nurses in public facilities in Uttar Pradesh. We recommend four types of interventions to close the gap between actual and target behaviors: structural improvements, systemic changes, community-level shifts, and interventions within healthcare facilities.</p>

opencc-zeroDec 2019View details →
zenodo28/100

supporting data for 'categorical colour metric' publication

<p>sampledMT.txt are metric tensors for the &#39;categorical colour metric&#39; describe in the paper of the same name (to be) published in PLOS One.</p> <p>The data is stored as a nested list of dimension 21*21*21*3*3; with the outer dimension being the R coordinate of where the tensor lies running from 0.00, 0.05,...,0.95,1.00; next dimension G; then Bl then the tensors themselves.</p> <p>It can be directly read into Mathematica using &lt;&lt;</p>

opencc-by-4.0Mar 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record