Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
598
datasets available to search
ShareScore release 0.9.0
Dataset results
598 results for “classifier”
Data from: Are landscape attributes a useful shortcut for classifying vegetation in the tropics? A case study of La Amistad International Park
Effective vegetation classification schemes identify the processes determining species assemblages and support the management of protected areas. They can also provide a framework for ecological research. In the tropics, elevation-based classifications dominate over alternatives such as river catchments. Given the existence of floristic data for many localities, we ask how useful floristic data are for developing classification schemes in species-rich tropical landscapes and whether floristic data provide support for classification by river catchment. We analyzed the distribution of vascular plant species within 141 plots across an elevation gradient of 130 to 3200 m asl within La Amistad National Park. We tested the hypothesis that river catchment, combined with elevation, explains much of the variation in species composition. We found that annual mean temperature, elevation, and river catchment variables best explained the variation within local species communities. However, only plots in high-elevation oak forest and Páramo were distinct from those in low- and mid-elevation zones. Beta diversity did not significantly differ in plots grouped by elevation zones, except for low-elevation forest, although it did differ between river catchments. None of the analyses identified discrete vegetation assemblages within mid-elevation (700–2600 m asl) plots. Our analysis supports the hypothesis that river catchment can be an alternative means for classifying tropical forest assemblages in conservation settings.
A genetic and computational approach to structurally classify neuronal types
<p>This a missing piece of supplemental information that should accompany the article</p> <p>A genetic and computational approach to structurally classify neuronal types</p> <ul> <li>Uygar Sümbül,</li> <li>Sen Song,</li> <li>Kyle McCulloch,</li> <li>Michael Becker,</li> <li>Bin Lin,</li> <li>Joshua R. Sanes,</li> <li>Richard H. Masland</li> <li>& H. Sebastian Seung</li> </ul> <p>The file was kindly supplied by Uygar Sumbul, but has not yet appeared on the nature communications website.</p> <p> </p>
Data used for training glioblastoma NF1 classifier
<p>All data is publicly available and downloaded from UCSC Xena<br /> https://genome-cancer.ucsc.edu/proj/site/xena/datapages/?cohort=TCGA%20Pan-Cancer</p> <p>Because the database is continously updated and to ensure reproducibility, access data from this cached download.</p> <p>RNAseq and Clincal data were downloaded on 8 March 2016<br /> Mutation data was downloaded on 12 June 2015</p>
Databases needed to run snakemake classifier
<p>Parsed Silva and Unite databases required as input for a snakemake workflow used to classify 18S and ITS1 sequences</p>
Supplementary Data and Videos for "Problems of classifying predator-induced prey immobility: an unexpected case of post-contact freezing"
<p> </p> <p>The data file and five videos below accompany our paper "Problems of classifying predator-induced prey immobility: an unexpected case of post-contact freezing", published in <em>Web Ecology</em> <strong>24,</strong> 35-40 (2024), https://doi.org/10.5194/we-24-35-2024.</p> <p> </p> <p><strong>Video 1.</strong> <em>Pachyoliva semistriata</em> freezes after making contact with <em>Agaronia propatula</em> (I).</p> <p><strong>Video 2.</strong> <em>Pachyoliva semistriata</em> freezes after making contact with <em>Agaronia propatula</em> (II). Surface water flows against the direction of movement of <em>Agaronia</em>.</p> <p><strong>Video 3.</strong> <em>Pachyoliva semistriata</em> freezes after making contact with <em>Agaronia propatula</em> (III). Surface water flows in the direction of movement of <em>Agaronia</em>.</p> <p><strong>Video 4.</strong> Responses of <em>Pachyoliva semistriata</em> when making contact with an obstacle; control experiments.</p> <p><strong>Video 5.</strong> <em>Pachyoliva semistriata</em> freezes when making contact with an obstacle and sensing <em>Agaronia</em> odours.</p> <p><strong>Data File 1.</strong> Numerical data used in creating Figure 2.</p>
Opioid classifiers: dataset sharing for submission Cot 31, 2023
<p>Contains datasets for submission of manuscript "Enhancing Opioid Bioactivity Predictions through Integration of Ligand-Based and Structure-Based Drug Discovery Strategies with Transfer and Deep Learning Techniques" Davide Provasi and Marta Filizola</p>
DeepMRG: a multi-label deep learning classifier for predicting bacterial metal resistance genes
Open the record for dataset details and reuse information.
A Comprehensive Dataset of Classified Citations with Identifiers from English Wikipedia (2024)
<p><strong>2024 (new!)</strong></p> <p>This is a dataset of 44.766.800 (+9.2%) citations extracted from the English Wikipedia February 2024 dump (<a href="https://dumps.wikimedia.org/enwiki/20240220/">https://dumps.wikimedia.org/enwiki/20240220/</a>).</p> <p>The same extraction and template harmonization pipeline was used as the year before. The published dataset fields are like in the previous dataset. A classification label is assigned to each citation (either 'news', 'book', 'journal' or 'other)' by the deterministic rule-based classifier that analyses available identifiers (see code documentation for details), revealing the following citation subgroups:</p> <ol> <li>The total number of news: 10.958.151 (+9.4%)</li> <li>The total number of books:* 3.277.629 (+8.6%)</li> <li>The total number of journals*: 2.248.748 (+8.7%)</li> </ol> <p>* Please note that these numbers do not represent the overall number of book and journal citations, we count only citations with DOI, PMID, PMC and ISBN identifiers assigned by authors (prior to the lookup process that augments citations with missing identifiers). </p> <p>This dataset is not equipped with identifiers located via the lookup process (no 'acquired_ID_list' field). If there is interest in such an augmented version, see the source code for instructions or contact authors for assistance with this task. </p> <p><strong>2023</strong></p> <p>This is a dataset of 40.664.485 citations extracted from the English Wikipedia February 2023 dump (<a href="https://dumps.wikimedia.org/enwiki/20230220/">https://dumps.wikimedia.org/enwiki/20230220/</a>).</p> <p>Version 1: en_citations.zip is a dataset of extracted citations </p> <p>Version 2: en_final.zip is the same dataset with classified citations augmented with identifiers </p> <p>The fields are as follows:</p> <ul> <li>type_of_citation - Wikipedia template type used to define the citation, e.g., 'cite journal', 'cite news', etc.</li> <li>page_title - title of the Wikipedia article from which the citation was extracted.</li> <li>Title - source title, e.g., title of the book, newspaper article, etc.</li> <li>URL - link to the source, e.g., webpage where news article was published, description of the book at the publisher's website, online library webpage, etc.</li> <li>tld - top link domain extracted from the URL, e.g., 'bbc' for https://www.bbc.co.uk/... </li> <li>Authors - list of article or book authors, if available.</li> <li>ID_list - list of publication identifiers mentioned in the citation, e.g., DOI, ISBN, etc.</li> <li>citations - citation text as used in Wikipedia code</li> <li>actual_label - 'book', 'journal', 'news', or 'other' label assigned based on the analysis of citation identifiers or top link domain. </li> <li>acquired_ID_list - identifiers located via Google Books and Crossref APIs for citations which are likely to refer to books or journals, i.e., defined using 'cite book', 'cite journal', 'cite encyclopedia', and 'cite proceedings' templates.</li> </ul> <ol> <li>The total number of news: 9.926.598</li> <li>The total number of books: 2.994.601</li> <li>The total number of journals: 2.052.172</li> <li>Augmented with IDs via lookup 929.601 (out of 2.445.913 book, journal, encyclopedia, and proceedings template citations not classified as books or journals via given identifiers). </li> </ol> <p>The source code to extract citations can be found here: <strong><a href="https://github.com/albatros13/wikicite">https://github.com/albatros13/wikicite</a>. </strong></p> <p>The code is a fork of the earlier project on Wikipedia citation extraction: <a href="https://github.com/Harshdeep1996/cite-classifications-wiki">https://github.com/Harshdeep1996/cite-classifications-wiki</a>.</p> <p> </p>
Light of History: Model weights and classified output
Open the record for dataset details and reuse information.
Algorithm Selection with Probing Trajectories: Benchmarking the Choice of Classifier Model - Data
<p>This repository contains the data and additional information for the paper 'Algorithm Selection with Probing Trajectories: Benchmarking the Choice of Classifier Model'. </p> <p>The following files are included:</p> <ul> <li>accuracy.zip: raw performance files for all models;</li> <li>plots.zip: additional plots;</li> <li>tuning.zip: tuning log files.</li> </ul>
Supporting Dataset For Paper: "Ranks underlie outcome of combining classifiers: quantitative roles for Diversity and Accuracy"
<p>This zipfile contains the DXA/MRI dataset used in the publication: "Ranks<br> underlie outcome of combining classifiers: quantitative roles for Diversity and<br> Accuracy".</p> <p>To be updated with publication citation/DOI when provided.</p>
A Machine Learning-Based Method for Classifying Well Test Responses in Naturally Fractured Reservoirs
<p>Complete dataset (raw data), processing codes (MATLAB v2020b) and numerical simulation model (Petrel v2017, ECLIPSE) for clustering of pressure derivatives in Naturally Fractured Reservoirs</p> <p> </p> <p> </p>
Distribution. Madeira Archipelago (Madeira and Porto Santo) and W Canary Is (La Palma, La Gomera, El Hierro, and Tenerife). Individuals classified as Pipistrellus sp. from the Azores have been suggested to be Madeira Pipistrelles. in Vespertilionidae
Distribution. Madeira Archipelago (Madeira and Porto Santo) and W Canary Is (La Palma, La Gomera, El Hierro, and Tenerife). Individuals classified as Pipistrellus sp. from the Azores have been suggested to be Madeira Pipistrelles.
Experiment data for the 2022 GECCO paper on the Bayesian Learning Classifier System
<p>Data collected during the empirical study for the paper *Pätzel and Hähner. 2022. The Bayesian Learning Classifier System: Implementation, Replicability, Comparison with XCSF* (DOI: https://doi.org/10.1145/3512290.3528736).</p> <p>To evaluate the data, see https://doi.org/10.5281/zenodo.6460994 .</p>
Dataset associated with the manuscript " Locally developed models improve the accuracy of remotely assessed metrics as a rapid tool to classify sandy beach morphodynamics"
<p>Raw dataset associated with the manuscript " Locally developed models improve the accuracy of remotely assessed metrics as a rapid tool to classify sandy beach morphodynamics"</p>
Thwaites CNN Classified
<p>First attempt at Thwaites classification</p>
Planar Graph Classifier - Result Data
<p>CSV file containing loss and accuracy of the Planar Graph Classifier evaluation on the test set.</p>
Planar Graph Classifier - Training History
<p>Training history tracking loss and accuracy during the keras model fitting in numpy file format.</p>
Planar Graph Classifier - Training History Plot
<p>Visualization of the loss and accuracy per epoch during training of the Planar Graph Classifier</p>
Planar Graph Classifier - Generated Graph Images
<p>These images were generated with networkx layout algorithms and matplotlib graph visualization from the graph input datasets of the Planar Graph Classifier. Some randomization in size and color was used for a slight augmentation effect.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.