Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
598
datasets available to search
ShareScore release 0.9.0
Dataset results
598 results for “classifier”
Development of classifier of engagement in occupation with machine learning (CEOML) for quantifying context
<p>These are the raw data, code, and development model from the development of the classifier of engagement in occupation with machine learning (CEOML).</p>
Dataset for the Real/Bogus classifier in the Tomo-e Gozen transient survey
<p>This is a subset of the dataset used in the following paper:<br> Ichiro Takahashi, Ryo Hamasaki, Naonori Ueda, Masaomi Tanaka, Nozomu Tominaga, Shigeyuki Sako, Ryou Ohsawa, Naoki Yoshida, Deep-learning real/bogus classification for the Tomo-e Gozen transient survey, Publications of the Astronomical Society of Japan, 2022;, psac047, <a href="https://doi.org/10.1093/pasj/psac047">https://doi.org/10.1093/pasj/psac047</a></p> <p>This dataset is available for training and testing the Real/Bogus classifier that classifies real transient objects and false detections in the Tomo-e Gozen transient survey. The source code for training is available at <a href="https://github.com/ichiro-takahashi/tomoe-realbogus">https://github.com/ichiro-takahashi/tomoe-realbogus</a>.</p> <p>The dataset consists of images (npy) and meta data (csv). Each sample of the image data includes a set of three images: reference image, observed new image, and subtracted image. Each of them is a cutout image around the transient candidate (29×29 pixels).</p> <p><br> The training data are divided into the following four parts, all of which must be unzipped into the same directory.</p> <ul> <li>tomoerb_train_Q1.tar.gz</li> <li>tomoerb_train_Q2.tar.gz</li> <li>tomoerb_train_Q3.tar.gz</li> <li>tomoerb_train_Q4.tar.gz</li> </ul> <p>The unzipped data are separated according to the detector ID of the Tomo-e Gozen camera (111-444).</p> <p>The meta data contain the following labels:</p> <ul> <li>rand_artificial_real: Real</li> <li>galx_artificial_real: Real</li> <li>artifact: Bogus</li> </ul> <p>The test data are separated into real objects and bogus objects, and the meta data contain the detector ID.</p> <p>The number of samples in the training and test data are as follows:</p> <ul> <li>Training data: <ul> <li>Real: 1224710</li> <li>Bogus: 2031132</li> </ul> </li> <li>Test data: <ul> <li>Real: 292</li> <li>Bogus: 255711</li> </ul> </li> </ul> <p> </p> <p>Acknowledgement:</p> <p>This work has been supported by Japan Science and Technology Agency (JST) AIP Acceleration Research Grant Number JP20317829 and the Japan Society for the Promotion of Science (JSPS) KAKENHI grants 21H04491, 18H05223, and 17H06363.</p> <p>This work is supported in part by the Optical and Near-Infrared Astronomy Inter-University Cooperation Program.</p> <p>The Pan-STARRS1 Surveys (PS1) and the PS1 public science archive have been made possible through contributions by the Institute for Astronomy, the University of Hawaii, the Pan-STARRS Project Office, the Max-Planck Society and its participating institutes, the Max Planck Institute for Astronomy, Heidelberg and the Max Planck Institute for Extraterrestrial Physics, Garching, The Johns Hopkins University, Durham University, the University of Edinburgh, the Queen's University Belfast, the Harvard-Smithsonian Center for Astrophysics, the Las Cumbres Observatory Global Telescope Network Incorporated, the National Central University of Taiwan, the Space Telescope Science Institute, the National Aeronautics and Space Administration under Grant No. NNX08AR22G issued through the Planetary Science Division of the NASA Science Mission Directorate, the National Science Foundation Grant No. AST-1238877, the University of Maryland, Eotvos Lorand University (ELTE), the Los Alamos National Laboratory, and the Gordon and Betty Moore Foundation.</p>
Research data supporting: "Classifying soft self-assembled materials via unsupervised machine learning of defects"
<p>Research data supporting: "Classifying soft self-assembled materials via unsupervised machine learning of defects".</p> <p>The root folder contains 5 folders:</p> <ol> <li>FIBERS</li> <li>MEMBRANES_and_MICELLES</li> <li>NANOPARTICLES</li> <li>COMPARISON</li> <li>paper_images</li> </ol> <p>The folders 1. to 3. contain the data for every soft-matters architecture used to produce the results discussed in the main paper. Each of these folders contain additional sub-fordels: TRAJ, SOAP, PCA, CLUSTERING, containing the files discussed in the main paper.</p> <p>Folder 4. contains the data of the comparison between different classes of materials (SOAP, PCA, and CLUSTERING sub-folders).</p> <p>Folder 5. contains the images that are showed in the main paper and in the Supporting Information.</p>
Deep learning generates custom-made logistic regression models for explaining how breast cancer subtypes are classified
<p>Breast cancer is the most frequently found cancer in women and the one most often subjected to genetic analysis. Nonetheless, it has been causing the largest number of women's cancer-related deaths. PAM50, the intrinsic subtype assay for breast cancer, is beneficial for diagnosis and stratified treatment but does not explain each subtype's mechanism. Nowadays, deep learning can predict the subtypes from genetic information more accurately than conventional statistical methods. However, the previous studies did not directly use deep learning to examine which genes associate with the subtypes. Ours is the first study on a deep-learning approach to reveal the mechanisms embedded in the PAM50-classified subtypes. We developed an explainable deep learning model called a point-wise linear model, which uses a meta-learning approach to generate a custom-made logistic regression model for each sample. Logistic regression is familiar to physicians and medical informatics researchers, and we can use it to analyze which genes are important for subtype prediction. The custom-made logistic regression models generated by the point-wise linear model for each subtype used the specific genes selected in other subtypes compared to the conventional logistic regression model: the overlap ratio is less than twenty percent. And analyzing the point-wise linear model's inner state, we found that the point-wise linear model used genes relevant to the cell cycle-related pathways. The results of this study suggest the potential of our explainable deep learning to play a vital role in cancer treatment.</p>
Ground truth data used to train the synapse classifier used in Lillvis et al., 2022 for ExLLSM circuit reconstruction
<p class="MsoNormal">Brain function is mediated by the physiological coordination of a vast, intricately connected network of molecular and cellular components. The physiological properties of network components can be quantified with high throughput; the ability to assess many animals per study has been key to relating physiological properties to behavior. Conversely, detailed anatomical properties (e.g., the synaptic connectivity of molecularly-defined cell types across an entire circuit) are presently quantifiable only with low throughput; thus we know very little about how network structure, and structural variation, influences behavior. For neuroanatomical reconstruction there is a methodological gulf between electron-microscopic (EM) methods, which yield dense connectomes (but at great expense and low throughput) and light-microscopic methods, which provide molecular and cell-type specificity with high throughput (but without synaptic resolution). We developed a high-throughput analysis pipeline and imaging protocol using tissue expansion and light sheet microscopy (ExLLSM) to rapidly reconstruct selected circuits across many animals with single-synapse resolution and molecular contrast. Using <em>Drosophila </em>to validate this approach, we demonstrate that it yields synaptic counts similar to those obtained by EM, enables synaptic connectivity to be compared across sex and experience, and can be used to correlate structural connectivity, functional connectivity, and behavior. This approach fills a critical methodological gap in studying variability in the structure and function of neural circuits across individuals within and between species.</p> <p class="MsoNormal">Here, we share the data used to train the synapse classifier that was utilized in the analysis pipeline. All additional software, code, and usage examples to train and run the classifier can be found at Github: <a href="https://github.com/JaneliaSciComp/exllsm-circuit-reconstruction">https://github.com/JaneliaSciComp/exllsm-circuit-reconstruction</a></p>
CLDF dataset derived from Tang and Her's "Quantitative typological data on classifiers and plural markers" from 2019
<p><strong>Tang, Marc and One-Soon Her. 2019. Insights on the Greenberg-Sanches-Slobin Generalization: Quantitative typological data on classifiers and plural markers. Folia Linguistica, 53(2): 297-331. https://doi.org/10.1515/flin-2019-2013.</strong></p>
MNBC: a multithreaded Minimizer-based Naïve Bayes Classifier for improved metagenomic sequence classification
<p>These files provide supplementary data underlying the article <a title="https://doi.org/10.1093/bioinformatics/btae601" href="https://doi.org/10.1093/bioinformatics/btae601" target="_blank" rel="noopener noreferrer nofollow">doi.org/10.1093/bioinformatics/btae601</a>:</p> <ul> <li>CAMI2_reference_database.tar.gz_1 to CAMI2_reference_database.tar.gz_6: Merge them into a single file using the <em>cat</em> command. The folder produced by decompressing this file is the reference database for CAMI2 (RefSeq sequence filenames are in the file CAMI2_reference_database_16864_genomes_list.txt at <a href="https://dx.doi.org/10.5281/zenodo.10568965">https://dx.doi.org/10.5281/zenodo.10568965</a>).</li> </ul> <p> </p>
MNBC: a multithreaded Minimizer-based Naïve Bayes Classifier for improved metagenomic sequence classification
<p>These files provide supplementary data underlying the article <a title="https://doi.org/10.1093/bioinformatics/btae601" href="https://doi.org/10.1093/bioinformatics/btae601" target="_blank" rel="noopener noreferrer nofollow">doi.org/10.1093/bioinformatics/btae601</a> (see Figure 1 of the article):</p> <ul> <li>37345_filtered_training_and_test_genomes_sequence_files.tar.gz_1 to 37345_filtered_training_and_test_genomes_sequence_files.tar.gz_6: Merge them into a single file using the <em>cat</em> command. The folder produced by decompressing this file contains the filtered RefSeq sequence files of all 37345 training and test genomes, which were used to build the uniform reference database and generate the test reads (The filenames are in the file 37345_filtered_training_and_test_genomes_list.txt at <a href="https://dx.doi.org/10.5281/zenodo.10568965">https://dx.doi.org/10.5281/zenodo.10568965</a>).</li> </ul>
Neighborhood benthic configuration reveals hidden social diversity: classified benthic data
<p>Ecological interactions among benthic communities are crucial for shaping marine ecosystems. Understanding these interactions is essential for predicting how ecosystems will respond to environmental changes, invasive species, and conservation management. However, determining the prevalence of species interactions at the community scale is challenging. To overcome this challenge, we employ tools from social network analysis, specifically exponential random graph modeling (ERGM). Our approach explores the relationships among animal and plant organisms within their neighborhoods. Inspired by companion planting in agriculture, we use spatiotemporal co-occurrence as a measure of mixed species interaction. In other words, the variety of community interactions based on co-occurrence defines what we call "co-occurrence social diversity." Our objective is to use ERGM to quantify the proportion of interactions at both the simple paired level and the more complex triangle level, enabling us to measure and compare co-occurrence social diversity. Applying our approach to the Spanish coastal zone across 8 sites, 5 depths, and sunlit/shaded aspects, we discover that 80% of sessile communities, consisting of over a hundred species, exhibit co-occurrence social diversity, with 5% of species consistently forming associations with other species. These organism-level interactions likely have a significant impact on the overall character of the site.</p>
Benchmark for classifying camera motion in coral videos
<p>A benchmark of coral vidoes where camera motion which indicates coral structures being viewed from multiple angles have been identified. This may be used for 3D reconstructions of corals. Time stamps are seperated by ";" and time segments are joined by "-". <br><br>This work was partially supported by the Data Science Research Center at the University of Haifa through the Israel PBC grant Advancing Data Science to Serve Humanity and Protect the Global Environment (grant no. 100009443)</p>
Automated Trustworthiness Testing for Machine Learning Classifiers
<p>This repository includes data for the paper <em>Automated Trustworthiness Testing for Machine Learning Classifiers</em>.</p>
Dataset for "Who is behind the Model? Classifying Modelers based on Pragmatic Model Features"
<p>This dataset contains pragmatic features computed after each interaction with the modeling tool.</p> <p>The dataset has been collected in several modeling sessions in 2010. Student data was collected using students from the Technical University of Eindhoven. Practitioner data, in turn, was collected as part of a Dutch BPM roundtable event in Eindhoven as well as in Berlin.</p>
ASSOS Classifier Datasets
<p>This 7zip File contains the in the <a href="https://www.researchgate.net/project/ASSOS-Animal-Sound-Sensor-Observation-Service">ASSOS project</a> used classification data as <a href="https://github.com/cran/monitoR">monitoR </a>corellation templates and there corrosponding wav files. The wav-files are part of the <a href="http://www.tierstimmenarchiv.de/">Tierstimmenarchiv</a>. The data collectors are referenced in the contributor section.</p>
MEASURING EFFICIENCY AND BENCHMARKING CLASSIFIED TWO - FIVE STAR HOTELS IN NAIROBI AND MOMBASA, KENYA
<p>The purpose of this study was to measure the relative efficiency of the hotels in Nairobi and Mombasa using Data Envelopment Analysis. The study was a longitudinal survey in which data are collected for periods; 2007, 2008 and. The study was limited to two-five star hotels. The study sample consisted of 36 hotels. The results revealed that technical inefficiencies of the hotels were mainly due to the pure technical inefficiencies rather than the scale inefficiencies.</p>
Classifying Generated White-Box Tests: An Exploratory Study
<p>The data and analysis scripts for our paper entitled Classifying Generated White-Box Tests: An Exploratory Study</p>
l-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"
<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of large size.</p>
s-sized Training and Evaluation Data for Publication "Using Supervised Learning to Classify Metadata of Research Data by Field of Study"
<p>Automated classification of metadata of research data by their discipline(s) of research can be used in scientometric research, by repository service providers, and in the context of research data aggregation services. Openly available metadata of the DataCite index for research data were used to compile a large training and evaluation set comprised of 609,524 records. This is the cleaned and vectorized version with a feature selection of small size.</p>
PMF-LP: the first 10 m plastic-mulched farmland distribution map (2019-2021) in the Loess Plateau of China generated using training sample generation and classifier transfer method
<p>This dataset provides 10-m resolution plastic-mulched farmland distribution map in the Loess Plateau of China from 2019 to 2021</p> <p>*** The data file is in ".tif" format</p> <p>*** Temporal resolution: Annually</p> <p>*** Temporal coverage: 2019-2021</p> <p>*** Pixel size: 10 m</p> <p>*** Projection information: EPSG: 4326</p> <p>*** Values: 1 denotes plastic-mulched farmland (PMF) and 0 denotes non-plastic-mulched farmland (non-PMF)</p>
Images dataset for Chemical Images Classifier model
<h1>Original paper</h1> <p>The manually curated images dataset is a part of the Supplementary Materials of the paper: <code>A. Krasnov, S. Barnabas, T. Böhme, S. Boyer, L. Weber, Comparing software tools for optical chemical structure recognition, Digital Discovery (2024). <a href="https://doi.org/10.1039/D3DD00228D">https://doi.org/10.1039/D3DD00228D</a></code><br><br></p> <h1>Images dataset description</h1> <p>The dataset was used to generate the image classifier model. The dataset consists of <strong>16,000</strong> images that were collected from different sources:</p> <p>1) Chemical data images extracted from EP, US, and WO patents by OntoChem GmbH.</p> <p>2) Images from the MolScribe datasets <code><a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c01480">https://pubs.acs.org/doi/10.1021/acs.jcim.2c01480</a></code></p> <p>3) DECIMER–hand-drawn molecule images dataset <code>H.O. Brinkhaus, A. Zielesny, C. Steinbeck, K. Rajan, “DECIMER - hand-drawn molecule images dataset”, 2022, Journal of Cheminformatics, 14, 36. <a href="https://doi.org/10.1186/s13321-022-00620-9">https://doi.org/10.1186/s13321-022-00620-9</a></code></p> <p>4) Images from the Rxnscribe training set <code>Y. Qian, J. Guo, Z. Tu, C.W. Coley, R. Barzilay, “RxnScribe: A Sequence Generation Model for Reaction Diagram Parsing”, 2023, arXiv:2305.11845v1, <a href="https://doi.org/10.48550/arXiv.2305.11845">https://doi.org/10.48550/arXiv.2305.11845</a> </code></p> <p>5) Formulas images from the im2latex-100k dataset <code>A prebuilt dataset for OpenAI's task for image-2-latex system, <a href="../record/56198#.YJjuCGZKgox">https://zenodo.org/record/56198#.YJjuCGZKgox</a> (accessed 16 Januar 2024)</code></p> <h1>Structure of dataset</h1> <p>The dataset consists of two directories:</p> <p>The "<strong>classified</strong>" directory contains manually labeled images. These images are divided into four distinct categories, with each category including 4000 images:</p> <p>● one_molecule</p> <p>● several_molecules</p> <p>● reactions</p> <p>● other</p> <p>In the “<strong>for_model</strong>” folder, we have split the images for training, validation, and testing in order to create a Chemical Image Classifier model:</p> <p>● training: 12,804 images</p> <p>● test: 1,604 images</p> <p>● validation: 1,604 images.</p> <p> </p>
Reproduction package for the paper "Classifying Linux commits"
<p>This is the reproduction package of the paper "Classifying Linux commits". This package includes the datasets and the software for methods presented in the paper.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.