Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
598
datasets available to search
ShareScore release 0.9.0
Dataset results
598 results for “classifier”
Fig. 1 in Aiteng Ater, New Genus, New Species, An Amphibious And Insectivorous Sea Slug That Is Difficult To Classify [Mollusca: Gastropoda: Opisthobranchia: Sacoglossa(?): Aitengidae, New Family]
Fig. 1. Aiteng ater, new species: A, photo of a wayang (shadow puppet) and of a statue of Ai Theng at a shop near Pak Phanang; B, creeping adult specimen length 9.0 mm in lab; C, creeping young slug length 2.0 mm; D, young slug curled up length 2.1 mm; E, nerve ring, width of photo 170 µm; F, digestive tubes in ventral view in posterior part of body, width of photo 2600 µm; G, posterior lateral side of notum in dorsal view with skin largely removed showing digestive follicles and dorsal vessels; H, prostate dorsal view, width photo 270 µm; I, parasites around salivary follicles, ventral view; J, ventral view of albumen gland, ampulla, and follicles; K, eye; L, penis; M, lateral view of radular teeth. Legend: ag − albumen gland; am − ampulla; dg − digestive follicles; dt − digestive tubes; dv − dorsal vessels; ey − eye; fb − foot border; fo − follicles of ovotestis; nb − notum border; pa − parasite; ph − pharynx; pr − prostate; sg − salivary follicles; sr − seminal receptacle; st − stomach; ve −velum.
FIG. 10 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 10. — Veranichthys ventralis (Agassiz, 1839) n. comb.: A, lectotype, MNHN Bol 275 (10724); B, paralectotype, MNHN Bol 24 (10723) (lectotype of Serranus rugosus Heckel, 1854). Middle Eocene of Bolca, northern Italy. Scale bars: 10 mm.
FIG. 8 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 8. — Goujetia crassispina (Agassiz, 1839) n. comb., holotype MNHN Bol 8 (10811); dentary, medial premaxillary teeth (upper left) and pharyngeal teeth (upper right). Middle Eocene of Bolca, northern Italy. Scale bar: 5 mm.
FIG. 11 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 11. — Veranichthys ventralis (Agassiz, 1839) n. comb., reconstruction of the skeleton based on two syntypes; scales omitted.
FIG. 7 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 7. — Goujetia crassispina (Agassiz, 1839) n. comb., holotype, MNHN Bol 8 (10811). Middle Eocene of Bolca, northern Italy. Scale bar: 10 mm.
FIG. 9 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 9. — Goujetia crassispina (Agassiz, 1839) n. comb., reconstruction of the skeleton based on the holotype (deformations corrected); scales omitted.
FIG. 6 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 6. — Ottaviania leptacanthus (Agassiz, 1839) n. comb., reconstruction of the skeleton based on the holotype (position of supraneurals changed, see text); scales omitted.
FIG. 3 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 3. — Pseudosparnodus microstomus (Agassiz, 1839), MNHN Bol 49 (10752b, holotype of Odonteus sparoides Agassiz, 1839). Middle Eocene of Bolca, northern Italy. Scale bar: 10 mm.
FIG. 5 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 5. — Ottaviania leptacanthus (Agassiz, 1839) n. comb., holotype, MNHN Bol 14b (10809b). Middle Eocene of Bolca, northern Italy. Scale bar: 10 mm.
FIG. 4 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 4. — Sparnodus elongatus Agassiz, 1839, MNHN Bol 69 (10754, holotype of Pristipoma furcatum Agassiz, 1839). Middle Eocene of Bolca, northern Italy. Scale bar: 10 mm.
FIG. 2 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 2. — The urohyal of Pseudosparnodus microstomus (Agassiz, 1839): A, MNHN Bol 252 (10795), c. 118 mm SL; B, MNHN Bol 269 (10730a), 109 mm SL. Middle Eocene of Bolca, northern Italy. Scale bar: 5 mm.
FIG. 1 in Fishes from the Eocene of Bolca, northern Italy, previously classified in the Sparidae, Serranidae and Haemulidae (Perciformes)
FIG. 1. — Pseudosparnodus microstomus (Agassiz, 1839), MNHN Bol 255 (10726, lectotype of Serranus microstomus Agassiz, 1839 as designated by Eastman [1905]). Middle Eocene of Bolca, northern Italy. Scale bar: 10 mm.
Qiime2 classifiers
<p>These files are companion classifiers for QIIME2 snakemake pipeline.</p> <p> </p> <p>If you use these data files please cite:</p> <ul> <li> <p>Michael S Robeson II, Devon R O’Rourke, Benjamin D Kaehler, Michal Ziemski, Matthew R Dillon, Jeffrey T Foster, Nicholas A Bokulich. RESCRIPt: Reproducible sequence taxonomy reference database management for the masses. bioRxiv 2020.10.05.326504; doi: <a href="https://doi.org/10.1101/2020.10.05.326504">https://doi.org/10.1101/2020.10.05.326504</a></p> </li> <li> <p>Bokulich, N.A., Kaehler, B.D., Rideout, J.R. et al. Optimizing taxonomic classification of marker-gene amplicon sequences with QIIME 2’s q2-feature-classifier plugin. Microbiome 6, 90 (2018). <a href="https://doi.org/10.1186/s40168-018-0470-z">https://doi.org/10.1186/s40168-018-0470-z</a></p> </li> <li> <p>See the <a href="https://www.arb-silva.de/">SILVA website</a> and the latest <a href="https://www.nature.com/articles/ismej2011139">Greengenes publication</a> for the latest citation information for these reference databases.</p> </li> <li> <p>We will update this information later.</p> </li> </ul>
Supporting information for: Age-specific sensitivity analysis of stable, stochastic and transient growth for stage-classified populations
<p>The study associated with this dataset proposes a way of performing age-specific sensitivity analysis of stable, stochastic and transient growth for stage-classified populations. Here, you find simulation code in R to produce figures in the manuscript and matrices reporting demographic data upon which code computations are performed.</p>
Image to Classify type of SSW in JRA55 (1958-2005)
<p>Imágenes de altura geopotencial a 10 hPa durante los SSW en el hemisferio Norte de 1958-2005 de JRA55. Usadas para clasificar el tipo de evento (División o desplazamiento de vórtice).</p>
LAGOS-US Race and Ethnicity: Dataset classifying conterminous US lake communities based on racial and ethnic demographic data (v 2.0)
<p>Knowing the consistency of water quality sampling in lakes surrounded by a variety of racial and ethnic communities is important for potential policy uses and to assess community impacts. By using the 2010 US Census race and ethnicity demographic block group data, we analyzed the frequency (i.e., number of years and consistency) of lake water quality sampling according to racial and ethnic demographics in surrounding neighborhoods. Our approach classified human communities near lakes as predominantly White or people of color (POC), and Hispanic or non-Hispanic (NH). Associated R analysis scripts can also be found in this folder. Our data and approach can be used for future studies seeking to analyze environmental monitoring practices in relation to human demographic variables, particularly from US Census data.</p>
Zenodo Open Metadata snapshot - Training dataset for records and communities classifier building
<p>This dataset contains Zenodo's published open access records and communities metadata, including entries marked by the Zenodo staff as spam and deleted.</p> <p>The datasets are gzipped compressed JSON-lines files, where each line is a JSON object representation of a Zenodo record or community.</p> <p><strong>Records dataset</strong></p> <p>Filename:<strong> </strong>zenodo_open_metadata_{ date of export }.jsonl.gz</p> <p>Each object contains the terms: <em>part_of, thesis, description, doi, meeting, imprint, references, recid, alternate_identifiers, resource_type, journal, related_identifiers, title, subjects, notes, creators, communities, access_right, keywords, contributors, publication_date</em></p> <p>which correspond to the fields with the same name available in Zenodo's record JSON Schema at <a href="https://zenodo.org/schemas/records/record-v1.0.0.json">https://zenodo.org/schemas/records/record-v1.0.0.json</a>.</p> <p>In addition, some terms have been altered:</p> <ul> <li>The term <strong>files</strong> contains a list of dictionaries containing <strong>filetype</strong>, <strong>size,</strong> and <strong>filename </strong>only.</li> <li>The term <strong>license</strong> contains a short Zenodo ID of the license (e.g. "cc-by").</li> </ul> <p><strong>Communities dataset</strong></p> <p>Filename:<strong> </strong>zenodo_community_metadata_{ date of export }.jsonl.gz</p> <p>Each object contains the terms: <em>id, title, description, curation_policy, page </em></p> <p>which correspond to the fields with the same name available in Zenodo's community creation form.</p> <p><strong>Notes for all datasets</strong></p> <p>For each object the term <strong>spam</strong> contains a boolean value, determining whether a given record/community was marked as spam content by Zenodo staff.</p> <p>Some values for the top-level terms, which were missing in the metadata may contain a <strong>null</strong> value.</p> <p>A smaller uncompressed random sample of 200 JSON lines is also included for each dataset to test and get familiar with the format without having to download the entire dataset.</p>
Dataset for the paper "Website Fingerprinting: Attacking Popular Privacy Enhancing Technologies with the Multinomial Naïve-Bayes Classifier"
<p>This dataset contains website fingerprints of 775 websites analyzed in the paper "Website Fingerprinting: Attacking Popular Privacy Enhancing Technologies with the Multinomial Naïve-Bayes Classifier" published in the Proceedings of the 2009 ACM workshop on Cloud computing security (CCSW 2009, DOI: 10.1145/1655008.1655013).</p>
Data for: PerchPicker classifier model v7: A catalog of American silver perch (Bairdiella chrysoura) calls for machine learning
<p>This data repository contains labeled passive underwater acoustic data used to train and test the machine-learning model of Bohnenstiehl (in prep - 2023), <em>Automated cataloging of American silver perch (Bairdiella chrysoura) calls using machine learning</em>. The software accompanying this paper is known as PerchPicker (<a href="https://github.com/drbohnen/PerchPicker" rel="noopener">GitHub - drbohnen/PerchPicker)</a>, and the classifier model presented in the paper is v7. It consists of more than 6000 labeled perch and 6000 labeled other signals. Labeled scalogram images are provided, along with pressure-corrected waveforms (micro-Pascals) sampled at 24 kHz. Each waveform sample is 90 ms long. The center 30 ms of these waveform segments represent the portion of the signal used in training and testing the classifier model. Waveform data are provided in multiple formats: 1) MATLAB (.mat) files containing the 'perch' and 'other' waveforms stored in column format, and 2) individual .wav files, each containing a labeled waveform example. Codes are provided to demonstrate how these .wav files can be read into MATLAB and PYTHON. These labeled data can be used to re-train the PerchPicker model or develop alternative classifiers. </p>
Manually labeled Bird song dataset of 22 species from Xeno-canto to enhance deep learning acoustic classifiers with contextual information.
<p>Data accompanying the paper: Jeantet and Dufourq (2023). Empowering Deep Learning Acoustic Classifiers with Human-like Ability to Utilize Contextual Information for Wildlife Monitoring. <em>Ecological Informatics</em>. 77, 15749541, DOI: 10.1016/j.ecoinf.2023.102256</p> <p> </p> <p>Our investigation contributes to the field of deep learning and bioacoustics by highlighting the potential for improved classification performance through the incorporation of contextual information such as time and location.</p> <p>To test if spatial-temporal information can enhance deep learning classifier, we developed a subset dataset derived from Xeno-Canto that included location metadata as input alongside the spectrogram. We used this dataset with the primary purpose of creating a bird song classification task with species carefully selected to share similar vocal characteristics but from distinct geographical distributions. We only considered the recordings of category `A', corresponding to the best quality score in the database.</p> <p>The dataset contains songs of <strong>22 bird species</strong> from 5 families and genera differents. The recordings were downloaded from the Xeno-canto database in .wav format and each recording was <strong>manually annotated </strong>by labelling the start and stop time for every vocalisation occurrence using Sonic Visualiser. In total, database contained 6537 occurrences of bird songs of various length from <strong>967 file recordings</strong>. A precise description of the distribution by species and country can be found in the associated article.</p> <p> </p> <p>The audio files are provided in "Audio.zip" and the manually verified annotation in "Annotations.zip". The name of each file follows the following nomenclature: Family_genus_species_country of recording_date of recording_ID Xenocanto_type of song.wav/svl. The meta-data information of each file can be find in the csv file provided (Xenocanto_metadata_qualityA_selection) based on the number of the ID Xeno-canto. The annotations can be viewed using the Sonic Visualiser software. The python codes to process these files and train neural networks can be found here : github</p> <p>The files were divided into a <strong>training folder</strong> and a<strong> validation folder</strong> to train and evaluate the efficiency of each method. For each species and country, we randomly selected 70% of the downloaded recordings for the training dataset and kept the remaining 30% for validation.</p> <p><strong>Process to select the species</strong> : We selected the ten most recorded families in the Passeriformes order, the most represented order in Xeno-canto database. From each of the ten families, we again sub-samples the ten most recorded genera. For each genus, we observed the countries of the recordings and the number of available recordings per species and countries. From these observations, we made a self-selection of genera containing species with similar songs but recorded in different regions, with enough recordings available by species and country to form a dataset . At the end, 5 genus were selected containing 22 species. We considered only recordings associated with bird songs, specifically, within Xeno-canto we selected the `song' type. To balance the number of recordings between species of the same genus, we reduced the number of recordings for the most represented species. Thus, for each genus we calculated the average of the number of records available per species and per country and limited the number of recordings for the species/country pairs that were in greater number to this value plus two.</p> <p> </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.