Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
106
datasets available to search
ShareScore release 0.7.1
Dataset results
106 results for “Classification model”
HETEAC – The Hybrid End-To-End Aerosol Classification model for EarthCARE: Look-Up Table (LUT) for aerosol mixtures
<p>The dataset contains the look-up table (LUT) of EarthCARE’s Hybrid End-To-End Aerosol Classification (HETEAC) model. The LUT contains optical and radiative parameters for four pure aerosol components (fine mode weakly absorbing, fine mode strongly absorbing, coarse mode spherical and coarse mode non-spherical) and their mixtures. In total, 314 aerosol mixtures are considered. The LUT returns the mixing state of an aerosol mixture based on the lidar ratio and the particle linear depolarization ratio at 355 nm. The mixing state is expressed in terms of relative volume contribution of the four pure aerosol components. Additionally, the LUT returns the effective radius, the asymmetry parameter, the single scattering albedo (at 355, 532, 550, 670, 865, 1064, 1650 and 2210 nm) and the Angstrom exponent (at 28 wavelength combinations) of the aerosol mixture. The lidar ratio and the particle linear depolarization ratio is also provided at 532, 550, 670, 865, 1064, 1650 and 2210 nm.</p> <p>The datafile contains two top-level groups: the HeaderData, which contains the header variables, and the ScienceData with the variables. The latter contains two groups, the AerosolComponents, which includes the aerosol-component-related optical and microphysical variables, and the LookUpTable, which contains the HETEAC LUT variables.</p> <p>The variables included in the datafile are listed below. For each variable, a full description is provided in the long_name attribute.</p> <ul> <li>HeaderData <ul> <li>angstrom_exponent_header</li> </ul> </li> <li>ScienceData <ul> <li>AerosolComponents <ul> <li>backscatter</li> <li>effective_radius</li> <li>extinction</li> <li>logarithmic_width</li> <li>mode_radius_number</li> <li>mode_radius_volume</li> <li>particle_linear_depolarization_ratio</li> <li>refractive_index_imaginary</li> <li>refractive_index_real</li> <li>scattering</li> </ul> </li> <li>LookUpTable <ul> <li>angstrom_exponent</li> <li>asymmetry_parameter</li> <li>effective_radius</li> <li>lidar_ratio</li> <li>particle_linear_depolarization_ratio</li> <li>relative_volume_contribution</li> <li>single_scattering_albedo</li> </ul> </li> <li>radiation_wavelength</li> </ul> </li> </ul> <p>Contact</p> <p>For any further clarifications or expression of interest with respect to the EarthCARE LUT, please contact Ulla Wandinger (ulla.wandinger@tropos.de) and/or Athena Augusta Floutsi (floutsi@tropos.de).</p> <ul> </ul>
Small mammal classification model
<p>Classification model for small mammals and images used for training, validating and testing the model. The script used for training and further information are available on https://github.com/hannaboe/camera_trap_workflow.</p> <p> </p>
Modelling heterogeneity in the classification process in multi-species distribution models can improve predictive performance
Open the record for dataset details and reuse information.
Data from: Performance of unmarked abundance models with data from machine-learning classification of passive acoustic recordings
Open the record for dataset details and reuse information.
MVCNN++: CAD model shape classification and retrieval using multi-view convolutional neural networks
<p>Deep neural networks have shown promising success towards the classification and retrieval tasks for images and text data. While there have been several implementations of deep networks in the area of computer graphics, these algorithms do not translate easily across different datasets, especially for shapes used in product design and manufacturing domain. Unlike datasets used in the 3D shape classification and retrieval in the computer graphics domain, engineering level description of 3D models do not yield themselves to neat distinct classes. The current study looks at an improved form of the 3D shape deep learning algorithm for classification and retrieval through the use of techniques such as relaxed classification, use of prime angled camera angles for capturing feature detail and transfer learning for reducing the amount of data and processing time needed to train shape recognition algorithms. The proposed algorithm (MVCNN++) builds on top of multi-view convolutional neural network (MVCNN) algorithm, improving its efficacy for manufacturing part classification by enabling use of part metadata, yielding an improvement of almost 6% over the original version. With the explosive growth of 3D product models available in publicly available repositories, search and discovery of relevant models is critical to democratizing access to design models.</p>
PyTorch model for frame classification
<p>This is a trained PyTorch model for classifying a DNA sequence's (preferably of length 300) frame within an ORF.</p>
IMBALANCED MACHINE LEARNING CLASSIFICATION MODELS FOR REMOVAL BIOSIMILAR DRUGS AND INCREASED ACTIVITY IN PATIENTS WITH RHEUMATIC DISEASES
<p>Objective: Predict long-term disease worsening and the removal of biosimilar medication in patients with rheumatic diseases.</p><p>Methodology: Observational, retrospective, and descriptive study. Review of a database of patients with immune-mediated inflammatory rheumatic diseases. Disease worsening and removing biosimilars are imbalanced variables, that require using imbalanced machine learning models selected based on their superior f1-scores and great accuracy. Previously, we selected the most important variables using mutual information tests.</p><p>Results: The best imbalanced machine learning models to predict disease worsening and the removal of the biosimilar obtained f1-scores of 0.52 and 0.63, respectively. Both models are decision trees. In the first one, two important factors are switching of biosimilar and age, and in the second, the relevant variables are optimization and the value of the initial CRP. </p><p>Conclusions: Biosimilar drugs do not always work well for rheumatic diseases. We obtained two imbalanced machine learning models to detect those cases, where the drug should be removed or where the activity of the disease increases from low to high. Our decision trees use variables, such as age or switching, not considered in previous studies.</p>
The fully executable procedure of the U-Net model combined with the Multi-textRG algorithm to achieve fine ice-water classification ---- another 332 scenes of data-fused SIC labels.
<p>This data source is related to the manuscript titled "Combining the U-Net model and a Multi-textRG algorithm for fine SAR ice-water classification", which will be submitted to the journal---The Cryosphere. </p> <ul> <li>The"ready-to-train-fused_01.zip" to "ready-to-train-fused_10.zip" includes 200 scenes of data-fused SIC labels accessible with doi: 10.5281/zenodo.10973107, https://zenodo.org/records/10973107. </li> <li>The "ready-to-train-fused_11.zip" to "ready-to-train-fused_21.zip" includes another 332 scenes of data-fused SIC labels. </li> </ul>
Classification results extracted by the baseline RFSS+SI+GLCM and U-Net models.
<p>Selected indicative cases demonstrate (A) S2_12-12-20_16PCC_6, (B) S2_22-12-20_18QYF_0, (C) S2_27-1-19_16QED_14 and (D) S2_14-9-18_16PCC_13 patches on test set.</p>
Pre-trained Models for SMP Classification and Segmentation
<p>This dataset provides access to pre-trained models that were used for SnowMicroPen profile classification and segmentation. The models were trained on a part of the MOSAiC SMP dataset, available on <a href="https://doi.pangaea.de/10.1594/PANGAEA.935554">https://doi.pangaea.de/10.1594/PANGAEA.935554</a>. The labeled training data consists mostly of profiles from leg three of the expedition (January - May 2020), some profiles from leg one and two, and no profiles from leg four. Please refer to the snowdragon GitHub repository (<a href="https://github.com/liellnima/snowdragon">https://github.com/liellnima/snowdragon</a>) to access the models' training code and be directed to current publications.</p> <p>The following trained models are available here (alphabetically ordered):</p> <ul> <li>Artificial neural networks <ul> <li>Bi-directional long short-term memory <em>(blstm.hdf5)</em></li> <li>Encoder-decoder <em>(enc_dec.hdf5)</em></li> <li>Long short-term memory <em>(lstm.hdf5)</em></li> </ul> </li> <li>Baseline <ul> <li>Majority vote classifier <em>(baseline.model)</em></li> </ul> </li> <li>Semi-supervised models <ul> <li>Cluster-then-predict models: <ul> <li>Bayesian Gaussian mixture model <em>(gmm.model)</em></li> <li>Bayesian mixture model <em>(bmm.model)</em></li> <li>K-means clustering <em>(kmeans.model)</em></li> </ul> </li> <li>Label propagation <em>(label_spreading.model)</em></li> <li>Self-trained classifier <em>(self_trainer.model)</em></li> </ul> </li> <li>Supervised models <ul> <li>Balanced random forest <em>(rf_bal.model)</em></li> <li>Easy ensemble <em>(easy_ensemble.model)</em></li> <li>K-nearest neighbors <em>(knn.model)</em></li> <li>Random forest <em>(rf.model)</em></li> <li>Support vector machines <em>(svm.model)</em></li> </ul> </li> </ul> <p><br> <em>Loading Instructions:</em><br> The models with the file-ending ".model" are pickeled Python objects and can be loaded with ``pickle.load(your_model.model)``. The random forest must be loaded with ``joblib.load(rf.model)``. All artificial neural networks are h5py.File objects (tf.keras models) and can be loaded with ``tf.keras.models.load_model(your_ann.model)``.</p>
A synthetic dataset for the exploration of survival and classification models: prediction of heart attack or stroke within a 10-year follow-up period
<div> <div></div> </div> <div> <div> <div> <p><span>Machine learning methodologies are increasingly popular in health care research. This shift to integrated data science approaches necessitates professional development of the existing health care data analyst workforce. To enhance a smooth transition, educational resources need to be developed. Barriers to accessing real healthcare datasets, vital for health care data analyses methodologies training purposes, include financial, ethical and patient confidentiality concerns. Synthetic datasets mimicking real-world complexities offer a simpler solution.</span></p> <p>We present a synthetic dataset which mirrors routinely collected primary care data on heart attack and stroke among the adult population. The data incorporates much of the practical challenges encountered in routinely collected primary care systems such as missing data, informative censoring, interactions, variable irrelevance, and noise and can be used for training in methods which handle these difficulties. The intent is for the user to build models of heart/stroke risk using survival-based methodologies.</p> <p>By sharing this synthetic dataset openly, our goal is to contribute a transformative asset for professional training in health and social care data analysis. The dataset covers demographics, lifestyle variables, comorbidities, systolic blood pressure, hypertension treatment, family history of cardiovascular diseases, respiratory functioning, and experience of heart-attack and/or stroke. This initiative aims to bridge the gap in sophisticated healthcare datasets for training, fostering professional development of the health and social care research workforce.</p> <p>This study is funded by the National Institute for Health and Care Research ARC Wessex and the National Centre for Research Methods. The views expressed in this summary are those of the author(s) and not necessarily those of the National Institute for Health and Care Research or the Department of Health and Social Care.</p> <p> </p> </div> </div> </div>
Trained Random Forest Model for PNW Seismic Event Classification Trained on 150s waveforms (P-50, P+100), 50 Hz, and 1-10 Hz BP Filtered
<p>This dataset contains three trained random forest models named as following - </p> <ul> <li>P_10_100_F_1_10_50.joblib - This is a model trained on 110s long waveforms (origin time - 10, origin time +100) in case of earthquakes and explosions and (first arrival pick -10, first arrival pick + 100) in case of surface events, the waveforms are tapered using 10% cosine taper, bandpass filtered between 1-10 Hz using Butterworth four corner filter, normalized and resampled to 50 Hz. </li> <li>P_50_100_F_1_10_50.joblib </li> <li>P_10_30_F_1_15_50.joblib. </li> </ul> <p>And also the standard scaler parameters for each features that will be used to normalize them. </p>
Data, code, models for "Weakly Supervised Semantic Segmentation for Joint Key Local Structure Localization and Classification of Aurora Image"
<p>Data, code and models for https://ieeexplore.ieee.org/document/8410588/</p>
Assessing the peatland hummock-hollow classification framework using high-resolution elevation models: Implications for appropriate complexity ecosystem modelling
<p>The hummock-hollow classification framework used to categorize peatland ecosystem microtopography is pervasive throughout peatland experimental designs and current peatland ecosystem modelling approaches. However, identifying what constitutes a representative hummock-hollow pair within a site and characterizing hummock-hollow variability within or between peatlands remains largely unassessed. Using structure-from-motion (SfM), high resolution digital elevation models (DEM) of hummock-hollow microtopography were used to: 1) examine how much area needs to be sampled to characterize site-level microtopographic variation; and 2) examine the potential role of microtopographic shape/structure on biogeochemical fluxes using data from 9 northern peatlands. This data set is comprised of plot DEMs, supporting data, and the script used to analyze data and produce figures presented in the manuscript submitted to Biogeosciences Discussion "ASSESSING THE PEATLAND HUMMOCK-HOLLOW CLASSIFICATION FRAMEWORK USING HIGH-RESOLUTION ELEVATION MODELS: IMPLICATIONS FOR APPROPRIATE COMPLEXITY ECOSYSTEM MODELLING".</p>
GIS data - Klasifikační model terénního reliéfu ČR | GIS data - Terrain Classification Model of the Czech Republic
<p>GeoTIFF layer (8 x 8 m) based on areial laser scanning data, containing a relief classification model of the Czech Republic processed according to the procedure described in a separately published article (see https://www.researchgate.net/publication/333633314_Wykorzystanie_ALS_do_zautomatyzowanej_analizy_krajobrazom_krajobrazom_krajobrazu_Use_animage_Use_animage_Use_animage_Use_animage_Use_animage_Use_animage_Use_animage_Use).</p>
Convolutional Neural Networks for Classification of Alzheimer's Disease: Overview and Reproducible Evaluation [Models]
<p>This file contains the pretrained models and the evaluation of the pipelines described in the paper <em>Convolutional Neural Networks for Classification of Alzheimer’s Disease: Overview and Reproducible Evaluation</em>.</p> <p>Source code can be downloaded at: <a href="https://github.com/aramis-lab/AD-DL">https://github.com/aramis-lab/AD-DL</a></p> <p>Also, single files can be obtained at: <a href="https://aramislab.paris.inria.fr/clinicadl/files/models/v0.0.1/">https://aramislab.paris.inria.fr/clinicadl/files/models/v0.0.1/</a></p> <p>The structure of the compressed file is as follows:</p> <p>clinicadl_models/<br> ├── 2D_slice<br> │ ├── baseline<br> │ │ ├── AD_CN<br> │ │ │ ├── best_model<br> │ │ │ └── performances<br> │ │ └── AD_CN_dataleakage<br> │ │ ├── best_model<br> │ │ └── performances<br> │ └── longitudinal<br> │ └── AD_CN<br> │ ├── best_model<br> │ └── performances<br> ├── 3D_patch<br> │ ├── baseline<br> │ │ ├── AD_CN<br> │ │ │ ├── best_model<br> │ │ │ └── performances<br> │ │ └── sMCI_pMCI<br> │ │ ├── best_model<br> │ │ └── performances<br> │ └── longitudinal<br> │ ├── AD_CN<br> │ │ ├── best_model<br> │ │ └── performances<br> │ └── sMCI_pMCI<br> │ ├── best_model<br> │ └── performances<br> ├── 3D_ROI_based<br> │ ├── baseline<br> │ │ ├── AD_CN<br> │ │ │ ├── best_model<br> │ │ │ └── performances<br> │ │ └── sMCI_pMCI<br> │ │ ├── best_model<br> │ │ └── performances<br> │ └── longitudinal<br> │ ├── AD_CN<br> │ │ ├── best_model<br> │ │ └── performances<br> │ └── sMCI_pMCI<br> │ ├── best_model<br> │ └── performances<br> ├── 3D_subject<br> │ ├── baseline<br> │ │ ├── AD_CN<br> │ │ │ ├── best_model<br> │ │ │ └── performances<br> │ │ └── sMCI_pMCI<br> │ │ ├── best_model<br> │ │ └── performances<br> │ └── longitudinal<br> │ ├── AD_CN<br> │ │ ├── best_model<br> │ │ └── performances<br> │ └── sMCI_pMCI<br> │ ├── best_model<br> │ └── performances<br> ├── autoencoders<br> │ ├── 3D_patch<br> │ │ ├── baseline<br> │ │ │ └── best_model<br> │ │ └── longitudinal<br> │ │ └── best_model<br> │ ├── 3D_ROI_based<br> │ │ ├── baseline<br> │ │ │ └── best_model<br> │ │ └── longitudinal<br> │ │ └── best_model<br> │ └── 3D_subject<br> │ └── baseline<br> │ ├── extensive<br> │ └── minimal<br> └── svm<br> ├── baseline<br> │ ├── AD_CN<br> │ │ ├── all_subjects.tsv<br> │ │ └── classifier<br> │ └── sMCI_pMCI<br> │ ├── all_subjects.tsv<br> │ └── classifier<br> └── longitudinal<br> ├── AD_CN<br> │ ├── all_subjects.tsv<br> │ └── classifier<br> └── sMCI_pMCI<br> ├── all_subjects.tsv<br> └── classifier</p> <p>We provide the pretrained CNN models for the frameworks 3D subject-level, 3D ROI-based, 3D patch-level and 2D slice-level. This models can be found as a <strong><em>.pth.tar</em> </strong>file (<em>Pytorch</em> format) inside the <em>best_model</em> folder for each framework (and for each fold). We also provide the autoencoders that initialize the training stage of the CNN networks. The <em>performances </em>folder contains the computed metrics for the correponding model (ACC, BA, etc). <em> </em></p> <p>For the svn classification, we provide files with the dual coefficients, the support vector indices and the weights. Also, <em>tsv</em> files with the subject list.</p>
Data for "Entanglement Dynamics in Monitored Kitaev Circuits: Loop Models, Symmetry Classification, and Quantum Lifshitz Scaling"
<p>We provide the data and scripts used to produce the figures shown in our publication "Entanglement Dynamics in Monitored Kitaev Circuits:<br>Loop Models, Symmetry Classification, and Quantum Lifshitz Scaling".</p>
Evaluating Machine Learning Models for Supernova Gravitational Wave Signal Classification
<p>This dataset contains gravitational wave (GW) data used in our research work <a href="https://doi.org/10.1088/2632-2153/ada33a" target="_blank" rel="noopener">Abylkairov et al. (2024)</a>. The first 10,000 columns represent the gravitational wave strain <em>D · h</em> [cm] for the corresponding time values ranging from -993 ms to 6.9 ms, with a step size of 0.1 ms. The zero time refers to the time of core bounce. Each row within these first 10,000 columns corresponds to 864 different gravitational wave signals.</p> <p>Columns 10,001 to 10,005 contain the following additional parameters:</p> <ul> <li><strong>T/|W|</strong>: The rotational parameter.</li> <li><strong>GR_or_GREP</strong>: Binary indicator for the signal type, where 0 denotes GR and 1 denotes GREP.</li> <li><strong>EOS</strong>: The equation of state (EOS) model, where 0 corresponds to SFHo, 1 to LS220, 2 to HSDD2, and 3 to GShenFSU2.1.</li> <li><strong>f_peak</strong>: The peak frequency [Hz].</li> <li><strong>D Delta h</strong>: <em>D · ∆h</em> [cm].</li> </ul> <p>For each row (representing a single gravitational wave signal), these parameters provide information about the signal's rotational parameter, type (GR or GREP), EOS model, peak frequency, and <em>D · ∆h</em>.</p> <p><strong>Note</strong>: In the f_peak calculation procedure, we truncated the GW signal at 4.5 ms after the end of the core bounce (see <a href="https://doi.org/10.1103/PhysRevD.95.063019">Richers et al. (2017)</a> for details).</p>
Summary table outlining key features of different OA business model classifications
<p>A summary table created as a quick reference for the paper published at https://doi.org/10.1629/uksg.667</p>
SPH modelling of companion-perturbed AGB outflows including a new morphology classification scheme
<p><strong>ABSTRACT</strong></p> <p><em>Context.</em> Asymptotic giant branch (AGB) stars are known to lose a significant amount of mass by a stellar wind, which controls the remainder of their stellar lifetime. High angular-resolution observations show that the winds of these cool stars typically exhibit mid- to small-scale density perturbations such as spirals and arcs, believed to be caused by the gravitational interaction with a (sub-)stellar companion.<br> <em>Aims.</em> We aim to explore the effects of the wind-companion interaction on the 3D density and velocity distribution of the wind, as a function of three key parameters: wind velocity, binary separation and companion mass. For the first time, we compare the impact on the outflow of a planetary companion to that of a stellar companion. We intend to devise a morphology classification scheme based on a singular parameter.<br> <em>Methods.</em> We ran a small grid of high-resolution polytropic models with the smoothed particle hydrodynamics (SPH) numerical code Phantom to examine the 3D density structure of the AGB outflow in the orbital and meridional plane and around the poles. By constructing a basic toy model of the gravitational acceleration due to the companion, we analysed the terminal velocity reached by the outflow in the simulations.<br> <em>Results.</em> We find that models with a stellar companion, large binary separation and high wind speed obtain a wind morphology in the orbital plane consisting of a single spiral structure, of which the two edges diverge due to a velocity dispersion caused by the gravitational slingshot mechanism. In the meridional plane the spiral manifests itself as concentric arcs, reaching all latitudes. When lowering the wind velocity and/or the binary separation, the morphology becomes more complex: in the orbital plane a double spiral arises, which is irregular for the closest systems, and the wind material gets focussed towards the orbital plane, with the formation of an equatorial density enhancement (EDE) as a consequence. Lowering the companion mass from a stellar to a planetary mass, reduces<br> the formation of density perturbations significantly.<br> <em>Conclusions.</em> With this grid of models we cover the prominent morphology changes in a companion-perturbed AGB outflow: slow winds with a close, massive binary companion show a more complex morphology. Additionally, we prove that massive planets are able to significantly impact the density structure of an AGB wind. We find that the interaction with a companion affects the terminal velocity of the wind, which can be explained by the gravitational slingshot mechanism. We distinguish between two types of wind focussing to the orbital plane resulting from distinct mechanisms: global flattening of the outflow as a result of the AGB star’s orbital motion and the formation of an EDE as a consequence of the companion’s gravitational pull. We investigate different morphology classification schemes and uncover that the ratio of the gravitational potential energy density of the companion to the kinetic energy density of the AGB outflow yields a robust classification parameter for the models presented in this paper.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.