Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
31
datasets available to search
ShareScore release 0.9.0
Dataset results
31 results for “Activity Classification”
Classification of GTP-dependent K-Ras4B active and inactive conformational states
<p>Dataset for the paper: Classification of GTP-dependent K-Ras4B active and inactive conformational states.</p> <p>Cite as: J. Chem. Phys. 158, 000000 (2023); DOI: 10.1063/5.0139181<br> Submitted: 18 December 2022; Accepted: 13 February 2023; Published Online: 13 February 2023.</p> <p>All molecular dynamics and molecular docking data presented, analyzed, and discussed in this paper are available at reasonable<br> requests submitted to the corresponding author. The following data are also available online: (i) file KRas4B_pdbs.zip, a compressed<br> archive including coordinate (.pdb) and structure (.psf) files for KRas-4B WT and D33E proteins, (ii) KRas4B_WT_traj.zip,<br> a compressed archive including trajectory files (.trr) for the WT KRas-4B runs, including 120 "*.trr"-formatted trajectories each corresponding to 40 ns of MD simulation time, and (iii) a sample Python script to generate a free energy plot as shown in Fig. 2.</p>
Human activity classification
<p>Collected dataset comprises body position recordings for various persons of diverse profile while performing six physical activities.</p> <p>The values include: </p> <p>- Cod_pat is the unique code for each person;<br> - Time is the timestamp;<br> - Ax, Ay, Az represent the axis of the acceleration;<br> - Mx, My, Mz represent the axis of the magnetometer;<br> - Ch is the compass heading;<br> - Label is the Class that provide the activity of the person.</p> <p> </p>
Data for paper titled : Comparing Clothing-Mounted Sensors with Wearable Sensors for Movement Analysis and Activity Classification (published in Sensors (MDPI))
<p>Data for paper titled : Comparing Clothing-Mounted Sensors with Wearable Sensors for Movement Analysis and Activity Classification (published in Sensors (MDPI))</p>
Author Classifications of O*NET Individual Work Activities (IWA)
<p>Author Classifications of O*NET Individual Work Activities (IWA) used in "Innovations and Economic Output Scale with Social Interactions in the Workforce"</p>
Improving Open Source Face Detection by Combining an Adapted Cascade Classification Pipeline and Active Learning
<p>The <em><strong>EAVISE Open Source Face Detection Dataset</strong></em> consists of several items that were used to generate the improved frontal face detection model using LBP features and AdaBoost for OpenCV3.2.</p> <ul> <li>The annotations of the FDDB dataset, converted to the OpenCV format for doing a correct evaluation.</li> <li>The final trained model (IterativeHardPositives+ model) which is included in the OpenCV 3.2 framework.</li> </ul>
Venus coronae activity classification
<p>This is a KML file based on the coronae classification that accompanies the manuscript "Corona structures driven by plume-lithosphere interactions and evidence for ongoing plume activity on Venus" by Gülcher et al. (2020) <em>(Nature Geoscience, </em><strong>13</strong>, 547-554, <a href="https://doi.org/10.1038/s41561-020-0606-1">https://doi.org/10.1038/s41561-020-0606-1</a><em>)</em>. The final version relates to the published manuscript, and includes more data points and an improved analysis compared to the first version (review round) than the first version. This KML file includes all of the coronae that were analysed in this study (from Table 3) and can be used in Google Earth or Google Venus. Markers are color-coded depending on the inferred corona activity (Table 2).</p>
Automated classification of avian vocal activity using acoustic indices in regional and heterogeneous datasets
<p>Acoustic indices combined with clustering and classification approaches have been increasingly used to automate identification of the presence of vocalizing taxa or acoustic events of interest. While most studies using this approach standardize data collection and study design parameters at the project or study level, recent trends in ecological research are to investigate patterns at regional or continental scales. Large-scale studies often require collaboration between research groups and integration of data from multiple sources to fulfill objectives, which can lead to variation in recording equipment and data collection protocols.</p> <p>Our objectives were to determine how analytical approaches and variation in data collection and processing that is typical of regional acoustic monitoring programs influences accuracy when identifying vocal activity in migratory breeding birds. We used data from three regional datasets in Northern Alberta, Northern British Columbia, and Southern and Central Yukon, Canada to investigate the effect of analytical framework, sample size, local species richness, and data collection variables on classification accuracy.</p> <p>We found supervised classification approaches to be the most effective, with boosted regression trees identifying vocal activity with a 92.0% accuracy and easily able to accommodate variation in data collection and processing parameters. We also provide recommendations on effectively processing large and heterogeneous datasets including sufficient sample size, accommodating nuisance variables, and selecting suitable model training data.</p> <p>The results presented in this study can help inform decisions in data collection, data processing, and study design and analysis, maximize performance and accuracy during analysis, and efficiently process large, heterogeneous datasets to answer questions at scales previously difficult to investigate.</p>
Plant carbohydrate-active enzymes in bamboo (Neosinocalamus affinis): identification, classification and function in lignocellulose biosynthesis in herbivore defence
<p><i><span>Neosinocalamus affinis</span></i>, a type of cluster bamboo,<i> </i>is a good candidate feedstock for biomass energy. In the study, we found a total of 686 genes were identified as belonging to CAZyme families in the <i><span>N. affinis</span></i> transcriptome, including 222 glycoside hydrolases (GHs), 288 glycosyltransferases (GTs), 64 carbohydrate esterases (CEs), 70 auxiliary activities (AAs), 37 carbohydrate binding modules (CBMs) and five polysaccharide lyases (PLs). Expression profiles revealed that several CAZyme genes were up-regulated after insect infestation, particularly the GT, GH, AA and CE family members. Lignocellulose assays showed that the contents of three components, cellulose, hemicellulose and lignin, increased after insect infestation. Our findings showed that CAZyme genes were abundant in the <i><span>N. affinis</span></i> transcriptome and were involved in the response to herbivory. These findings could be applied to protect bamboo against herbivores, such as the bamboo snout beetle <i><span>Cyrtotrachelus buqueti</span></i>, and develop low-cost chemical feedstock from bamboo.</p>
Semi-Supervised Active Learning for Sound Classification in Hybrid Learning Environments
<p>There are 16,930 sound instances in our database with durations ranging 242 from 1 to 10 seconds, which correspond to (approximately) 15 hours of environmental 243 sounds. All sound files were converted into raw 16 bit encoding, mono-channel, and 16 244 kHz sampling rate, as various formats and rates were used in the original versions 245 retrieved from the web.</p>
Semi-Supervised Active Learning for Sound Classification in Hybrid Learning Environments
<p>There are 16,930 sound instances in our database with durations ranging 242 from 1 to 10 seconds, which correspond to (approximately) 15 hours of environmental 243 sounds. All sound files were converted into raw 16 bit encoding, mono-channel, and 16 244 kHz sampling rate, as various formats and rates were used in the original versions 245 retrieved from the web.</p>
IMBALANCED MACHINE LEARNING CLASSIFICATION MODELS FOR REMOVAL BIOSIMILAR DRUGS AND INCREASED ACTIVITY IN PATIENTS WITH RHEUMATIC DISEASES
<p>Objective: Predict long-term disease worsening and the removal of biosimilar medication in patients with rheumatic diseases.</p><p>Methodology: Observational, retrospective, and descriptive study. Review of a database of patients with immune-mediated inflammatory rheumatic diseases. Disease worsening and removing biosimilars are imbalanced variables, that require using imbalanced machine learning models selected based on their superior f1-scores and great accuracy. Previously, we selected the most important variables using mutual information tests.</p><p>Results: The best imbalanced machine learning models to predict disease worsening and the removal of the biosimilar obtained f1-scores of 0.52 and 0.63, respectively. Both models are decision trees. In the first one, two important factors are switching of biosimilar and age, and in the second, the relevant variables are optimization and the value of the initial CRP. </p><p>Conclusions: Biosimilar drugs do not always work well for rheumatic diseases. We obtained two imbalanced machine learning models to detect those cases, where the drug should be removed or where the activity of the disease increases from low to high. Our decision trees use variables, such as age or switching, not considered in previous studies.</p>
Fink: early supernovae Ia classification using active learning
<p>Data and code used to obtain results presented in <a href="https://arxiv.org/abs/2111.11438">Leoni et al., 2021, Fink: early supernovae Ia classification using active learning</a></p>
GitRanking: A Ranking of GitHub Topics for Software Classification using Active Sampling
<p>Replication package for our paper: GitRanking: A Ranking of GitHub Topics for Software Classification using Active Sampling</p>
NIH BREATHE mHealth pediatric activity classification dataset
<p>The Biomedical REAl-Time Health Evaluation (BREATHE) platform is part of the Los Angeles PRISMS (Pediatric Research with Integrated Monitoring Systems) NIH-funded Center, which focuses on informatics platforms for mHealth (NIH/NIBIB U54 EB022002; PI: Alex Bui). This dataset was developed to better establish models for child activity classification based on smartwatch sensors. Specifically, a Motorola 360 smartwatch was worn by 21 subjects (age ranged 11-18), generating annotated data across different sensors (3-axis accelerometry, gyroscope, heart rate) for 6 different types of physical activity:three static postures (standing, sitting, lying) and three dynamic activities (walking, walking downstairs and walking upstairs). The study was motivated by the fact that current activity classifiers are typically geared towards <em>adult </em>bodies and patterns of motion; children exhibit different motion patterns, particularly as they grow. This dataset is not intended to be comprehensive, but rather a starting point for exploring differences in motion patterns across children (and relative to adults).</p>
World Odonata List (ODO) - active: Odonata species & higher classification (ODO)
Species List from World Odonata List version 20 April 2017. <p></p>https://www.pugetsound.edu/academics/academic-resources/slater-museum/biodiversity-resources/dragonflies/world-odonata-list2/<p></p>Species based on World Odonata List version 10 August 2020 The Encyclopedia of Life (EOL, eol.org) aggregates biodiversity information from more than 400 sources and provides access to the data through taxon pages, visual query and application programming interfaces. Scientific names are essential elements of the data integration infrastructure, but their shortcomings as key identifiers are well documented (Patterson et al., 2016). Complex automated workflows and continuous manual curation are required to address idiosyncrasies of source taxonomies, variation in data quality, and conflicting taxonomic opinions. To achieve a harmonized taxonomic view of EOL content, names from data sources are mapped to a dynamic reference hierarchy ([see current version here](<p></p>https://opendata.eol.org/dataset/tram-807-808-809-810-dh-v1-1/resource/00adb47b-57ed-4f6b-8f66-83bfdb5120e8)) using an algorithm that leverages canonical name strings, hierarchical information (ancestry, descendants), taxonomic ranks, synonym data, and author strings. Names that cannot be associated with a reference taxon are still accessible, but their unmapped status excludes them and any associated content from certain core EOL functions. For more information about the EOL taxonomy, see [EOL Dynamic Hierarchy](<p></p>https://eol.org/docs/eol-dynamic-hierarchy)
Vickery's late ideas on classification by phenomena and activities
<p>Classification was always a major interest for Brian Vickery. In his last years he contributed to the theoretical debate on classification with original ideas that are not well known yet, as some of them were consigned to ephemeral Web pages or private discussion. This paper attempts to report them and to discuss their implications for the current theory of knowledge organization.<br> Vickery’s proposals especially concern: application of the theory of integrative levels to classification as recommended by the Classification Research Group; the various dimensions of knowledge that are involved in the steps ‘from the world to the classifier’; progressive identification by science of phenomena within reality; interplay between the phenomena studied and the human activities providing the context and purpose for studying them; identification of facets of activities as well as facets of phenomena.<br> It is concluded that these ideas offer a substantial contribution to the theory of knowledge organization, and thus should be known, discussed and further developed. industry body. Instead, it requires a bold leap from an informed theoretical base to an implementable strategy.</p>
Plant carbohydrate-active enzymes in bamboo (Neosinocalamus affinis): identification, classification and function in lignocellulose biosynthesis in herbivore defence
Open the record for dataset details and reuse information.
Automated classification of avian vocal activity using acoustic indices in regional and heterogeneous datasets
Open the record for dataset details and reuse information.
Classification of Binding Modes for Kinase-Inhibitor Complex Structures, 3D Activity Cliffs Formed by Kinase Inhibitors, and Structural Analogues of 3D-Cliff Compounds
<p>The classification of crystallographic binding modes is provided for 884 kinase-inhibitor complex structures that were assembled from PDB. In addition, a total of 105 three-dimensional activity cliffs formed by 3D kinase inhibitors are listed. Their corresponding potency information is also given. Furthermore, the 2D structural analogues of 3D cliff-forming inhibitors were identified from ChEMBL database, on the basis of matched molecular pairs. These analogs and their activity information are also provided.</p>
Dynamic Visualization of ResNet Layer Activations for Brain Health Classification
<p>This GIF file provides a dynamic visualization of the internal representations (activations) from the ResNet 18 model layers (2 to 69) during a brain health classification task. The sequence begins by showing the original input image, followed by successive activation maps visualized using the "jet" colormap. Each frame corresponds to the activations extracted from a specific layer in the ResNet, resized to match the input image dimensions for better interpretability. </p> <p>The dataset used for this visualization is from S. Bhuvaji, "Brain Tumor Classification MRI," published on Kaggle in 2023. This dataset contains MRI images of brain tumors and has been utilized to train and evaluate the ResNet model for the classification of brain health states. The input sample displayed in the GIF is one such MRI image from the dataset, highlighting the model's ability to extract and analyze features relevant to brain tumor diagnosis.</p> <p>The activations reveal how the ResNet processes the input image hierarchically. In the <strong>early layers (e.g., Layers 2-10)</strong>, the network preserves much of the spatial structure of the original image, focusing on edges and low-level features. Moving to the <strong>intermediate layers (e.g., Layers 11-40)</strong>, the network begins to emphasize localized patterns while filtering out irrelevant structures such as the skull, concentrating instead on regions associated with tumors or health-related features. Finally, the <strong>deep layers (e.g., Layers 41-69)</strong> extract highly abstract and classification-relevant patterns, concentrating on tumor-related features while discarding most of the background.</p> <p>The progression of the activations in the GIF demonstrates how the network transitions from general image features to highly specialized, diagnostic features that are critical for the classification task. This visualization helps provide an intuitive understanding of the hierarchical processing capabilities of convolutional neural networks (CNNs) in medical image analysis.</p> <h3>Key Features:</h3> <p>The visualization includes an input brain image and its corresponding activation maps, extracted from each ResNet layer. The activation maps are resized to match the original image for consistency, and the "jet" colormap is applied to enhance visual interpretation of activation intensities. Each frame in the GIF dynamically updates to show the activations of the next layer, offering an engaging representation of the network's internal behavior.</p> <h3>Use Cases:</h3> <p>This GIF is a valuable resource for education, research, and presentations. It can be used to illustrate how deep learning models process medical images, providing insights into the hierarchical feature extraction process. Researchers and educators can leverage this visualization to explain the concept of feature abstraction in CNNs. It is also ideal for inclusion in talks, posters, and papers to showcase the dynamic analysis of neural network activations.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.