Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

598

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

598 results for “classifier”

Learn how ShareScore rates datasets ↗
edi60/100

LAGOS-US RESERVOIR: Data module classifying conterminous U.S. lakes 4 hectares and larger as natural lakes or reservoirs

The LAGOS-US RESERVOIR data module (hereafter, RESERVOIR) classifies all 137,465 lakes > 4 hectares in the conterminous U.S. into one of the following three categories using a machine-learning predictive model based on visual interpretation of lake outlines and a classification rule based on lake shape. Natural Lakes (NLs) are defined as lakes that are likely to be entirely or mostly naturally-formed and that do not have large, flow-altering structures on or near them; Reservoir Class A’s (RSVR_A) are defined as lakes that are likely to be either human-made or highly human-altered by the presence of a relatively large water control structure that appears to significantly change the flow of water; and Reservoir Class B’s (RSVR_Bs) are lakes that are likely to be entirely human-made based on isolation from rivers and a highly angular shape that is rarely, if ever, seen in natural lakes also often. We trained the machine learning models on 12,162 manually-classified lakes to assign probabilities of a lake being in 1 of 2 of the categories (NL or RSVR), then we further classified the RSVR classification into either A or B based on NHD Fcodes, isolation, and angularity. The data module includes a detailed User Guide, metadata tables, and a data table that includes information such as location, lake geometry, surface water connectivity class, and official name. Using our definition, our classification indicates that over 46 % of lakes > 4 ha in the conterminous U.S. are reservoir lakes. These data can be combined with other LAGOS-US data modules and U.S. national databases using unique lake identifiers to study both reservoir lakes and natural lakes at broad scales.

openCC (other)Nov 2022View details →
zenodo48/100

Aurora SDG Research Dashboard and Classifier - Instructions Videos

<p>Instruction videos about the Aurora SDG Research Dashboard, SDG Classifier, Badges and API.</p><p><a href="https://zenodo.org/doi/10.5281/zenodo.10040524">Also read the User Guides</a>.</p>

opencc-by-4.0Oct 2023View details →
zenodo48/100

18S V9 metabarcoding reference databases and naive-bayes classifier

<p>18S metabarcoding databases and naive-bayes classifiers specific to the V9 region. Built&nbsp;from&nbsp;the <a href="https://pr2-database.org/">PR2 database</a> using Qiime2 (version 2023.2)<a href="https://github.com/BenKaehler/q2-clawback">.</a> Includes&nbsp;a naive-bayes classifier for use with Qiime2. Sequences were dereplicated with Rescript --p-mode 'uniq' ,&nbsp;retaining identical sequence records that have differing taxonomies.</p><p>Primers used:</p><p>EMP 18S 1391f:&nbsp;GTACACACCGCCCGTC</p><p>EMP 18S EukBr:&nbsp;TGATCCTTCTGCAGGTTCACCTAC</p><p><strong>Stats</strong></p><p>19,470 unique sequences</p><p>39,170 total sequences</p><p>11,748 unique taxa&nbsp;</p><p>Note: there were 221,085 sequences in the original PR2 database. Many were filtered out due to the in-silico extraction with our V9 primers.</p><h3>File Descriptions</h3><p><strong>Files in bold are recommended for taxonomic classification.</strong></p><p>Create naive-bayes classifier for 18S PR2 database.md: &nbsp;Markdown with code used to generate databases |</p><p><strong>pr2_v5.0.0_SSU_18S-V9_uniq-classifier.qza</strong>: Unweighted naive-bayes classifier for 18S V9 (primers 1391f, EukBr), extracted from PR2 v5.0.1, dereplicated, generated by qiime2-2023.2 |</p><p><strong>pr2_version_5.0.0_SSU_18S-V9_uniq_seqs.qza</strong>: Sequences for 18S V9 (primers 1391f, EukBr), extracted from PR2 v5.0.1, dereplicated, generated by qiime2-2023.2 |</p><p><strong>pr2_version_5.0.0_SSU_18S-V9_uniq_tax.qza</strong>: Taxa for pr2_version_5.0.0_SSU_18S-V9_uniq_seqs.qza (dereplicated) |</p><p>pr2_version_5.0.0_SSU_18S-V9_seqs.qza: Sequences for 18S V9 (primers 1391f, EukBr), extracted from PR2 v5.0.1, NOT dereplicated, generated by qiime2-2023.2 |</p><p>pr2_version_5.0.0_SSU_18S-V9_tax.qza: Taxa for pr2_version_5.0.0_SSU_18S-V9_seqs.qza (NOT dereplicated)&nbsp;</p><p>pr2_version_5.0.0_SSU_mothur.fasta: SSU sequences downloaded from PR2 v 5.0.1 &nbsp;|</p><p>pr2_version_5.0.0_SSU_mothur.tax: SSU taxa downloaded from PR2 v5.0.1 |</p><p>pr2_version_5.0.0_taxonomy.xlsx: Detailed taxonomy downloaded from PR2 v5.0.1 |</p>

opencc-by-4.0Nov 2023View details →
zenodo48/100

S92 | FLUOROPHARMA | List of ~340 ATC classified fluoro-pharmaceuticals

<p>This is the collection associated with list S92 FLUOROPHARMA, List of ~340 ATC classified fluoro-pharmaceuticals on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>&nbsp;</p> <p>A list of ~340 fluoro-pharmaceuticals classified as per WHO'S Anatomical Therapeutic Chemical (ATC) classification based on their medical application, described in Inoue et. al. DOI:&nbsp;<a href="https://pubs.acs.org/doi/10.1021/acsomega.0c00830">10.1021/acsomega.0c00830</a>&nbsp;</p> <p>Structural identifiers and mapping to DTXSID, CASRN provided by ECI. v0.1.1: removed three non-F containing entries (CIDs 1061, 5360545, 56603655) and added information for CID 3039780.&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo48/100

S94 | FLUOROPEST | List of 423 FRAC/HRAC/IRAC classified fluoro-agrochemicals

<p>This is the collection associated with list S94&nbsp;FLUOROPEST,&nbsp;List of 423 FRAC/HRAC/IRAC classified fluoro-agrochemicals on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>A list of 423 Fungicide Resistance Action Committee (FRAC), Herbicide Resistance Action Committee (HRAC) or Insecticide Resistance Action Committee (IRAC) classified fluoro-agrochemicals based on their&nbsp;chemotype and mode of action , described in Ogawa et al DOI:<a href="https://doi.org/10.1016/j.isci.2020.101467">10.1016/j.isci.2020.101467</a>&nbsp;</p> <p>Structural identifiers and mapping to DTXSID, CASRN&nbsp;provided by ECI.</p>

opencc-by-4.0Feb 2022View details →
zenodo48/100

Convolutional Neural Networks for Classifying Combinatorial Metamaterials

<p>This dataset contains the training and test data, as well as the trained neural networks&nbsp;as used for the paper &#39;Machine Learning of Implicit Combinatorial Rules in Mechanical Metamaterials&#39;, as published in Physical Review Letters.</p> <p>In this paper, a neural network is used to classify each&nbsp;<span class="math-tex">\(k \times k\)</span> unit cell design of metamaterial M1 and M2&nbsp;into one of two classes (C or I).&nbsp;Additionally, the performance of the trained networks is analysed in detail. A more detailed description of the contents of the dataset follows below.</p> <p><strong>NeuralNetwork_train_and_test_data.zip</strong></p> <p>This file contains the train and test data used to train the Convolutional Neural Networks (CNNs) of the paper. Each unit cell size has its own file, and is saved in a zipped numpy file type (.npz). It contains data for metamaterial M1 (&quot;smiley_cube&quot;), and metamaterial M2 classification (i) (&quot;prek_xy&quot;) and (ii) (&quot;unimodal_vs_oligomodal_inc_stripmodes&quot;).</p> <p><strong>CNN_saves_kxk.zip</strong></p> <p>This file contains the parameter configurations of the CNNs trained on <span class="math-tex">\(k \times k\)</span>&nbsp;unit cells for metamaterial M2 classification (ii). Classification (i) is denoted by an additional M2ii in the file name. Metamaterial M1 is denoted by an extra M1 in the file name.&nbsp;Every hyperparameter (number of filters<em> nf,</em> number of hidden neurons<em> nh</em>, learning rate<em> lr</em>) combination is saved separately. The neural networks can be loaded using Google&#39;s TensorFlow package in Python, specifically using the &#39;tf.keras.models.load_model&#39; function.&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo48/100

Multi-head CRF classifier for biomedical multi-class Named Entity Recognition on Spanish clinical notes

<p>This contains the merged dataset as described in the work "<strong>Multi-head CRF classifier for biomedical multi-class Named Entity Recognition on Spanish clinical notes"</strong>.</p> <p>This dataset consists of 4 seperate datasets:</p> <ul> <li><a href="../records/8224056" target="_blank" rel="noopener">MedProcNer</a></li> <li><a href="../records/7614764" target="_blank" rel="noopener">DisTEMIST</a></li> <li><a href="../records/4270158" target="_blank" rel="noopener">PharmaCoNER</a></li> <li><a href="../records/10635215" target="_blank" rel="noopener">SympTEMIST</a></li> </ul> <p>The dataset contains two tasks:</p> <p><strong>Task 1:</strong> This task is related to multi-class Named Entity Recognition. This dataset contains 5 possible classes: SYMPTOM, PROCEDURE, DISEASE, CHEMICAL and PROTEIN.</p> <p><strong>Task 2:</strong> This task is related to Named Entity Linking, where each code corresponds to a code within the SNOMED-CT corpus. The exact corpus used can be obtained <a href="https://download.nlm.nih.gov/umls/kss/IHTSDO20190131/SnomedCT_SpanishRelease-es_PRODUCTION_20190430T120000Z.zip" target="_blank" rel="noopener">here</a>. Further for the MedProcNER, SympTEMIST and DisTEMIST datasets, a gazetteer is provided in the original datasets.&nbsp;</p> <p>For more information on the construction of the dataset, aswell as dataloaders, we refer you to our <a href="https://github.com/ieeta-pt/Multi-Head-CRF" target="_blank" rel="noopener">GitHub repository</a>.<br><br>Further this also contains the embeddings from the <a href="https://huggingface.co/cambridgeltl/SapBERT-UMLS-2020AB-all-lang-from-XLMR-large" target="_blank" rel="noopener">SapBERT</a> model.</p> <p><strong>Please, cite:</strong></p> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <blockquote> <div>@article{jonker2024a, title = {Multi-head {{CRF}} classifier for biomedical multi-class named entity recognition on {{Spanish}} clinical notes}, author = {Jonker, Richard A. A. and Almeida, Tiago and Antunes, Rui and Almeida, Jo{\~a}o R. and Matos, S{\'e}rgio}, year = {2024}, journal = {Database}, publisher = {Oxford University Press} }</div> </blockquote> <div>Jonker, R. A. A., Almeida, T., Antunes, R., Almeida, J. R., &amp; Matos, S. (2024). Multi-head CRF classifier for biomedical multi-class named entity recognition on Spanish clinical notes. (Submitted.)&nbsp;</div> <div>&nbsp;</div> <div> <p><strong>License</strong></p> <p>This work is licensed under a&nbsp;<a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div>

opencc-by-4.0May 2024View details →
zenodo48/100

Dataset containing binominal lexemes in Harakmbut (isolate, Peru), for "The derivational use of classifiers in Western Amazonia" and "When the alienability contrast fails to surface in adnominal possession: Bound nouns in Harakmbut"

<p>This is the dataset used, amongst others, in the paper: Van linden, An. Forthcoming. When the alienability contrast fails to surface in adnominal possession: Bound nouns in Harakmbut. Special Issue &ldquo;Re-assessing the explanatory potential of alienability contrasts&rdquo;, guest-edited by Fran&ccedil;oise Rose &amp; An Van linden. <em>Linguistics &ndash; An Interdisciplinary Journal of the Language Sciences</em>. [<a href="https://doi.org/10.1515/ling-2022-0039">https://doi.org/10.1515/ling-2022-0039</a>]</p> <p>For more details, see the ReadMe file.</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Harmonised LUCAS database classified by crop sequence type

<p>Assessing the benefits of crop diversification &ndash; a pillar of the agroecological transition &ndash; on a large scale requires a description of current crop sequences as a baseline, which is lacking at the scale of the European Union (EU). This work is based on the Harmonised LUCAS in-situ land cover and use database for field surveys from 2006 to 2018 in the European Union (doi: <a href="http://doi.org/10.2905/f85907ae-d123-471f-a44a-8cca993485a2">10.2905/f85907ae-d123-471f-a44a-8cca993485a2)</a> to fill this gap, We completed this dataset with a crop sequence type information for each point under non-perennial agricultural land cover in 2012, 2015 and 2018.</p> <p>The dataset lucas_classified.csv includes 31 159 points. Variables &quot;point_id&quot;, &quot;nuts0&quot;, &quot;nuts2&quot;, &quot;th_lat&quot;, &quot;th_long&quot;, &quot;LC1_2012&quot;, &quot;LC1_2015&quot;, &quot;LC1_2018&quot; are inherited from the Harmonised LUCAS databse. Variables &quot;cereals&quot;, &quot;corn&quot;, &quot;rapeseed&quot;, &quot;sunflower&quot;, &quot;pulses&quot;, &quot;rootCrops&quot;, &quot;forageLeg&quot;, &quot;grassland&quot; correspond to the temporal frequencies of respectively cereals, corn, rapeseed, sunflower, pulses, root crops, forage legumes and grassland within the 2012, 2015 and 2018 crop sequence for each point. Variable &quot;crop_sequence_type&quot; is the crop sequence type assigned to each point, among eight options: cereals, corn and cereals, forage legumes and cereals, pulses and cereals, rapeseed and cereals, root crops and cereals, sunflower and cereals, temporary grasslands.</p> <p>This dataset could be used to map current dominant crop sequences in the European Union, as illustrated in the map attached, and to assess the benefits of future crop diversification.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

CurveCurator: A recalibrated F-statistic to assess, classify, and explore significance of dose-response curves - Example Datasets

<p>CurveCurator is an open-source analysis platform for any dose-dependent data. It fits a classical 4-parameter equation to estimate effect potency, effect size, and the statistical significance of the observed response. 2D-thresholding efficiently reduces false positives in high-throughput experiments and separates relevant from irrelevant or insignificant hits in an automated and unbiased manner. An interactive dashboard allows users to quickly explore data locally.</p> <p><br> Here, we store example dose-dependent data, parameter files, and the corresponding CurveCurator pipeline outputs (v.0.2.0). Example data sets include Kinobeads Drug-binding data (1), CTRP Viability data sets (2), and deryptM Proteomics data sets (3). The F-value matrices for developing the CurveCurator tools are deposited as well.</p> <p>Original data sources:</p> <p>(1)<a href="https://doi.org:10.1126/science.aan4368"> https://doi.org:10.1126/science.aan4368</a></p> <p>(2) <a href="https://doi.org:10.1158/2159-8290.CD-15-0235">https://doi.org:10.1158/2159-8290.CD-15-0235</a></p> <p>(3) <a href="https://doi.org:10.1126/science.ade3925">https://doi.org:10.1126/science.ade3925</a></p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

The tpm metabarcoding DNA sequence database for taxonomic allocations using RDP classifier implemented in DADA2.

<p><strong>The </strong><em>tpm</em><strong> metabarcoding DNA sequence database for taxonomic allocations using the Mothur and DADA2 bio-informatic tools</strong></p> <p>A.C.M. Pozzi<sup>1</sup>, R. Bouchali<sup>1</sup>, L. Marjolet<sup>1</sup>, B. Cournoyer<sup>1</sup></p> <p><sup>1 </sup><em>University of Lyon, UMR Ecologie Microbienne Lyon (LEM), CNRS 5557, INRAE 1418, Universit&eacute; Claude Bernard Lyon 1, VetAgro Sup, Research Team &ldquo;Bacterial Opportunistic Pathogens and Environment&rdquo; (BPOE), 69280 Marcy L&rsquo;Etoile, France.</em></p> <p><strong>Corresponding authors: </strong></p> <ul> <li>A.C.M. Pozzi, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L&rsquo;Etoile, France. Tel. (+33) 478 87 39 47. Fax. (+33) 472 43 12 23. Email: <a href="mailto:adrien.meynier_pozzi@vetagro-sup.fr">adrien.meynier_pozzi@vetagro-sup.fr</a></li> <li>B. Cournoyer, UMR Microbial Ecology, CNRS 5557, CNRS 1418, VetAgro Sup, Main building, aisle 3, 1st floor, 69280 Marcy-L&rsquo;Etoile, France. Tel. (+33) 478 87 56 47. Fax. (+33) 472 43 12 23. Email: and <a href="mailto:benoit.cournoyer@vetagro-sup.fr">benoit.cournoyer@vetagro-sup.fr</a></li> </ul> <p><strong>Keywords:</strong></p> <p>BACtpm, Bacteria, <em>tpm</em>, thiopurine-<em>S</em>-methyltransferase EC:2.1.1.67, Nucleotide sequences, PCR products, Next-Generation-Sequencing, OTHU</p> <p><strong>Description:</strong></p> <ul> <li>The <em>tpm</em> gene codes for the thiopurine-<em>S</em>-methyltransferase (TPMT), an enzyme that can detoxify metalloid-containing oxyanions and xenobiotics (Cournoyer et al., 1998). Bacterial TPMTs radiated apart from human and animal TPMTs, and showed a vertical evolution in line with the 16S rRNA gene molecular phylogeny (Favre‐Bont&eacute; et al., 2005).</li> <li>The <em>tpm</em> database, named BACtpm, was designed to apply the <em>tpm</em>-metabarcoding analytical scheme published in Aigle et al. (2021). It includes the full <em>tpm</em> identifiers, GenBank accession numbers, complete taxonomic records (domain down to strain code) of about 215 nucleotide-long <em>tpm</em> sequences of 840 unique taxa belonging to 139 genera.</li> <li>Nucleotide sequences of <em>tpm</em> (range: 190-233 nucleotides) were either retrieved from public repositories (GenBank) or made available by B. Cournoyer&rsquo;s research group. Colin et al. (2020) described the PCR and high throughput Illumina Miseq DNA sequencing procedures used to produce <em>tpm</em> sequences.</li> <li>BACtpm v.2.0.1 (June 2021 release) is made available under the Creative Commons Attribution 4.0 International Licence. It can be used for the taxonomic allocations of <em>tpm </em>sequences down to the species and strain levels. Data is stored in the csv format enabling future user to reformat it to fit their specific needs.</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>We thank the worldwide community of microbiologists who made contributions to public databases in the past decades, and made possible the elaboration of the BACtpm database. We also thank the Field Observatory in Urban Hydrology (OTHU, <a href="http://www.graie.org/othu/">www.graie.org/othu/</a>), Labex IMU (Intelligence des Mondes Urbains), the Greater Lyon Urban Community, the School of Integrated Watershed Sciences H2O&#39;LYON, and the Lyon Urban School for their support in the development of this database. This work was funded by the French national research program for environmental and occupational health of ANSES under the terms of project &ldquo;Iouqmer&rdquo; EST 2016/1/120, l&#39;Agence Nationale de la Recherche through ANR-16-CE32-0006, ANR-17-CE04-0010, ANR-17-EURE-0018 and ANR-17-CONV-0004, by the MITI CNRS project named Urbamic, and the French water agency for the Rh&ocirc;ne, Mediterranean and Corsica areas through the Desir and DOmic projects. We thank former BPOE lab members who contributed to start and expand the BACtpm database: C&eacute;line COLINON, Romain MARTI, Emilie BOURGEOIS, S&eacute;bastien RIBUN and Yannick COLIN.</p> <p><strong>References:</strong></p> <p>Aigle, A., Colin, Y., Bouchali, R., Bourgeois, E., Marti, R., Ribun, S., Marjolet, L., Pozzi, A.C.M., Misery, B., Colinon, C., Bernardin-Souibgui, C., Wiest, L., Blaha, D., Galia, W., Cournoyer, B., 2021. Spatio-temporal variations in chemical pollutants found among urban deposits match changes in thiopurine S-methyltransferase-harboring bacteria tracked by the tpm metabarcoding approach. Sci. Total Environ. 767, 145425. https://doi.org/10.1016/j.scitotenv.2021.145425</p> <p>Colin, Y., Bouchali, R., Marjolet, L., Marti, R., Vautrin, F., Voisin, J., Bourgeois, E., Rodriguez-Nava, V., Blaha, D., Winiarski, T., Mermillod-Blondin, F., Cournoyer, B., 2020. Coalescence of bacterial groups originating from urban runoffs and artificial infiltration systems among aquifer microbiomes. Hydrol. Earth Syst. Sci. 24, 4257&ndash;4273. https://doi.org/10.5194/hess-24-4257-2020</p> <p>Cournoyer, B., Watanabe, S., Vivian, A., 1998. A tellurite-resistance genetic determinant from phytopathogenic pseudomonads encodes a thiopurine methyltransferase: evidence of a widely-conserved family of methyltransferases1The International Collaboration (IC) accession number of the DNA sequence is L49178.1. Biochim. Biophys. Acta BBA - Gene Struct. Expr. 1397, 161&ndash;168. https://doi.org/10.1016/S0167-4781(98)00020-7</p> <p>Favre‐Bont&eacute;, S., Ranjard, L., Colinon, C., Prigent‐Combaret, C., Nazaret, S., Cournoyer, B., 2005. Freshwater selenium-methylating bacterial thiopurine methyltransferases: diversity and molecular phylogeny. Environ. Microbiol. 7, 153&ndash;164. https://doi.org/10.1111/j.1462-2920.2004.00670.x</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

RNAPosers: Machine Learning Classifiers For RNA-Ligand Poses [Data Set]

<ul> <li>This dataset contains the decoys poses used to train and test RNAPosers, a set of RNA-ligand pose classifiers.</li> <li>The folder of&nbsp;each RNA-ligand complex (identified using its PDB ID) contains: <ul> <li>Ligand SMILES: lig.smi</li> <li>Ligand coordinate:&nbsp;lig.sd</li> <li>Receptor coordinate:&nbsp;receptor.mol2</li> <li>Pose&nbsp;coordinates: poses.sd</li> <li>Pose similarity data:&nbsp;rmsd.txt</li> </ul> </li> </ul>

opencc-by-4.0Mar 2020View details →
zenodo44/100

A Data Set of 255,000 Randomly Selected and Manually Classified Extracted Ion Chromatograms for Evaluation of Peak Detection Methods

<p>Non-targeted mass spectrometry (MS) has become an important method over the last years in the fields of metabolomics and environmental research. While more and more algorithms and workflows become available to process a large number of data sets nontargeted, there still exist few manually evaluated universal test data sets for refining and evaluating these methods. The first step of non-targeted screening, peak detection (and refinement of it) is arguably the most important step for non-targeted screening. However, the absence of a model data set makes it harder for researchers to evaluate peak detection methods. In this Data Descriptor, we provide a manually checked data set consisting of 255,000 EICs (5000 peaks randomly sampled from across 51 samples) for the evaluation on peak detection and gap filling algorithms. The data set was created from a previous real-world study, of which a subset was used to extract and manually classify ion chromatograms by three mass spectrometry experts. The data set consists of:</p> <ul> <li>51 converted mass spectral files in mzML format</li> <li>An .RData-file containing the extracted ion chromtograms (EICs)</li> <li>The randomly selected subset and the original output table of MZmine in .csv-format</li> <li>Example .xlsx files for the classification</li> <li>2 central classification tables</li> <li>Several tables with additional information about the sampling, chemical analysis and expert jugdement on EICs</li> </ul> <p>For a full description of the experiment and the data set, please read the related Data Descriptor with the title &quot;A data set of 255000 randomly selected and manually classified extracted ion chromatograms for evaluation of peak detection methods&quot; in Metabolites (https://www.mdpi.com/journal/metabolites; DOI: https://doi.org/10.3390/metabo10040162).</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

SuperWASP Variable Stars: Classifying Light Curves Using Citizen Science

<p>Table of 301 previously unidentified SuperWASP stellar variables and related characteristics, not including rotators and unknown variables. The variable type has been decided by citizen scientists through the SuperWASP Variable Stars Zooniverse project.&nbsp;The types and periods of each object have been assessed by the authors to correct for mis-classifications; whilst they have been corrected as much as possible, some types periods remain best guesses. All periods have an uncertainty of 0.1%.</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

RDP Classifier 2.14 and the RDP bacterial and archaeal taxonomy training set No. 19

<p>RDP Classifier 2.14 (August 2023) Release Note:</p><p>The Bacteria and Archaea hierarchy model used by RDP Classifier has been updated to training set No. 19. The new version has over 600 genera and 2500 species added since last version No. 18 released in July 2020. The information that is used to update the RDP taxonomy to training set version No. 19, and RDP Classifier version 2.14 came from publicly available scientific articles and public sequence repository, mostly from International Journal of Systematic and Evolutionary Microbiology (IJSEM), the All-Species Living Tree Project (LTP) and GenBank. &nbsp;</p><p>It is worth noting that most of the phyla have new names, according to article "</p><p>Oren A, Garrity GM. Valid publication of the names of forty-two phyla of prokaryotes. Int J Syst Evol Microbiol. 2021 Oct;71(10). doi: 10.1099/ijsem.0.005056. PMID: 34694987."</p><p>In addition to the files to train and run the RDP Classifier, new file formats are made available to accommodate the needs of users:</p><p>1. A new file trainset19_072023_speciesrank.fa has been added to the release in RDPClassifier_16S_trainsetNo19_rawtrainingdata.zip. This file is NOT needed to train the classifier. In addition to sequences, it contains genus, species, strain, type status and taxonomy rank, which are useful for closest species identification using third-party tools (e.g. BLAST).</p><p>2. Two new files in RDPClassifier_16S_trainsetNo19_QiimeFormat.zip to retrain the RDP Classifier included in Qiime2 package.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Qiime2 classifiers (rbcl, Mollusc 18s) for testing the validity of using eDNA for carbon origin analysis from sediment cores

<p>Qiime2 formatted classifiers that were created for a Natural England funded project by researchers at the James Hutton Institute. The pilot project aims to test the validity of using eDNA for carbon origin analysis from sediment cores. These classifiers for the rbcl and 18 Mollusc genes were made using RESCRIPt and Qiime2.&nbsp;</p> <p>The scripts used to created these classifiers are available at the James Hutton ICS GitHub <a href="https://github.com/HuttonICS/blue-carbon-db">blue-carbon-db</a> . The files are as follows:</p> <p><a href="../api/records/10046481/draft/files/mollusc-espineira-classifier.qza/content" target="_blank" rel="noopener noreferrer">mollusc-espineira-classifier.qza</a> is a classifer built from ncbi 18s Mollusc sequences, trained on the primer set from Espi&ntilde;eira et al (2009).</p> <div>rbcl-vasselon-zimmerman-F3-R1-classifier.qza is a classifer built from ncbi rbcl sequences, trained on the F3 and R1 primer set fromVasselon et al (2017).</div> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Important: </strong>If you use these classifiers please be aware of the process used to create them and be sure to review the methods. These databases were created by downloading data from the NCBI in October 2023, sequence data available at the NCBI changes over time. To create the most up to date database a fresh download and re-evaluations of the databases would be preferable. All method and scripts can be found at <a href="https://github.com/HuttonICS/blue-carbon-db">blue-carbon-db&nbsp;</a></p> <p>If you use these database please reference this repository along with RESCRIPt and Qiime2&nbsp;</p> <p>&nbsp;</p> <p>Espi&ntilde;eira, M., Gonz&aacute;lez-Lav&iacute;n, N., Vieites, J. M. and Santaclara, F. J. 2009 Development of a method for the genetic identification of commercial bivalve species based on mitochondrial 18S rRNA sequences. J Agric Food Chem, 28, 495-502 https://doi.org/10.1021/jf802787d</p> <p>&nbsp;</p> <p>Vasselon, V., Rimet, F., Tapolczai, K. and Bouchez, A. 2017. Assessing ecological status with diatoms DNA metabarcoding: Scaling-up on a WFD monitoring network (Mayotte island, France). Ecological Indicators, 82, 1-12 <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.ecolind.2017.06.024" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.ecolind.2017.06.024</a></p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Benchmark for classifying presence of coral and camera motion in underwater

<p>Benchmark for classifying presence of coral in underwater videos, and camera motion that would be necessary for 3d reconstruction of coral. Videos are collected from the YouTube-8M dataset.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Geomagnetic Storms - Classified - 1993 - 2025

<p>List of Geomagnetic Storms from 1993 to 2025 classified in main and recovery phase. The requirement for a storm to be identified is that it reaches an SMR index of -50 [nT]</p> <p>More information on the SMR ring current index can be found here : http://supermag.jhuapl.edu&nbsp;</p> <p>This version is based on the work published on Geophysical Research letters (GRL) "Plasma sheet Magnetic Flux Transport During Geomagnetic Storms" : https://agupubs.onlinelibrary.wiley.com/doi/epdf/10.1029/2024GL110839</p> <p>Columns are in order: index, storm number, minimum SMR, start time, end time, phase characterization, and duration.</p> <p>Comapred to version V2 the latest storms require manual verification</p>

opencc-by-4.0May 2024View details →
zenodo44/100

patccat: A classifier for patent claims

<p><strong>Data version: 3.3.0</strong></p> <p>Authors:<br>Bernhard Ganglmair (University of Mannheim, Department of Economics, and ZEW Mannheim)<br>W. Keith Robinson (Wake Forest University, School of Law)<br>Michael Seeligson (Southern Methodist University, Cox School of Business)</p> <p><br>1. Notes on Data Construction<br>2. Citation and Code<br>3. Description of the Data Files<br>3.1. File List<br>3.2. List of Variables for Files with Claim-Level Information<br>3.3. List of Variables for Files with Patent-Level Information<br>4. Coming Soon!</p> <p><br><strong>1. Notes on Data Construction</strong></p> <p>This is version 3.3.0 of the patccat data (patent claim classification by algorithmic text analysis).</p> <p>Patent claims define an invention. A patent application is required to have one or more claims that distinctly claim the subject matter which the patent applicant regards as her invention or discovery. We construct a classifier of patent claims that identifies three distinct claim types: process claims, product claims, and product-by-process claims.</p> <p>For this classification, we combine information obtained from both the preamble and the body of a claim. The preamble is a general description of the invention (e.g., a method, an apparatus, or a device), whereas the body identifies steps and elements (specifying in detail the invention laid out in the preamble) that the applicant is claiming as the invention. The combination of the preamble type and the body type provides us with a more detailed and more accurate classification of claims than other approaches in the literature. This approach also accounts for unconventional drafting approaches. We eventually validate our classification using close to 10,000 manually classified claims.</p> <p>The data files contain the results of our classification. We provide claim-level information for each independent claim of U.S. utility patents granted between 1836 and 2020. We also provide patent-level information, i.e., the counts of different claim types for a given patent.</p> <p>For a detailed description of our classification approach, please take a look at the accompanying paper (Ganglmair, Robinson, and Seeligson 2022).</p> <p><strong>2. Citation</strong></p> <p>Please cite the following paper when using the data in your own work:</p> <p>Ganglmair, Bernhard, W. Keith Robinson, and Michael Seeligson (2022): "The Rise of Process Claims: Evidence from a Century of U.S. Patents," unpublished manuscript available at <a href="https://papers.ssrn.com/abstract=4069994">https://papers.ssrn.com/abstract=4069994</a>.</p> <p>In the paper, we document the use of process claims in the U.S. over the last century, using the patccat data. We show an increase in the annual share of process claims of about 25 percentage points (from below 10% in 1920). This rise in process intensity of patents is not limited to a few patent classes, but we observe it across a broad spectrum of technologies. Process intensity varies by applicant type: companies file more process-intense patents than individuals, and U.S. applicants file more process-intense patents than foreign applicants. We further show that patents with higher process intensity are more valuable but are not necessarily cited more often. Last, process claims are on average shorter than product claims (with the gap narrowing since the 1970s).</p> <p>We would love to see how other researchers use the data and eventually learn from it. If you have a discussion paper or a publication in which you use the data, please send us a copy at patccat.data@gmail.com.</p> <p>We will the R code used to construct the data on Github with the next data version (version 3.4.0). Contact us at b.ganglmair@gmail.com if you would like to take a look at an earlier version of the code.</p> <p><br><strong>3. Description of the Data Files</strong></p> <p>The data files contain claim-level information for independent claims of 10,140,848 U.S. utility patents granted between 1836 and 2020. The files further contain patent-level information for U.S. utility patents.</p> <p><em>3.1. File List</em></p> File list <table><tbody> <tr> <td>claims-patccat-v3-3-sample.csv</td> <td>claim-level information for independent claims of a sample of 1000 patents issued between 1976 and 2020</td> </tr> <tr> <td>claims-patccat-v3-3-1836-1919.csv</td> <td>claim-level information for independent claims of 1,038,041 patents issued between 1836 and 1919</td> </tr> <tr> <td>claims-patccat-v3-3-1920-2020.csv</td> <td>claim-level information for independent claims of 9,102,807 patents issued between 1920 and 2020</td> </tr> <tr> <td>patents-patccat-v3-3-sample.csv</td> <td>patent-level information for a sample of 1000 patents issued between 1976 and 2020</td> </tr> <tr> <td>patents-patccat-v3-3-1836-1919.csv</td> <td>patent-level information for 1,038,041 patents issued between 1836 and 1919</td> </tr> <tr> <td>patents-patccat-v3-3-1920-2020.csv</td> <td>patent-level information for 9,102,807 patents issued between 1920 and 2020</td> </tr> </tbody> </table> <p><br><em>3.2. List of Variables for Files with Claim-Level Information</em></p> <p>For detailed descriptions, see the appendix in Ganglmair, Robinson, and Seeligson (2022).</p> List of Variables (Claim-Level Information) <table><tbody> <tr> <td>PatentClaim</td> <td>patent claim identifier; 8-digit patent number and 4-digit claim number (Ex: 01234567-0001)</td> </tr> <tr> <td>singleLine</td> <td>=1 if claim is published in single-line format</td> </tr> <tr> <td>singleReformat</td> <td>outcome code of reformating of single-line claims</td> </tr> <tr> <td>Jepson</td> <td>=1 if claim is a Jepson claim</td> </tr> <tr> <td>JepsonReformat</td> <td>outcome code of reformating of Jepson claims</td> </tr> <tr> <td>inBegin</td> <td>=1 if claim begins with the word "in"</td> </tr> <tr> <td>wordsPreamble</td> <td>number of words in the claim preamble</td> </tr> <tr> <td>wordsBody</td> <td>number of words in the claim body</td> </tr> <tr> <td>dependentClaims</td> <td>number of dependent claims that refer to this independent claim</td> </tr> <tr> <td>isMeansPreamble</td> <td>=1 if term "means" is used in the preamble</td> </tr> <tr> <td>isMeansBody</td> <td>=1 if term "means" is used in the body</td> </tr> <tr> <td>isMeans</td> <td>=1 if term "means" is used anywhere in the claim (~ means-plus-function claim)</td> </tr> <tr> <td>processPreamble</td> <td>=1 if terms "method" or "process" are used in the preamble</td> </tr> <tr> <td>processBody</td> <td>=1 if terms "method" or "process" are used in the body</td> </tr> <tr> <td>processSimple</td> <td>=1 if terms "method" or "process" are used anywhere in the claim (for simple approach of process claim classification)</td> </tr> <tr> <td>claimType</td> <td>claim type of full classification (1 = process; 2 = product; 3 = product-by-process; 0 = no type)</td> </tr> <tr> <td>preambleType</td> <td>preamble type</td> </tr> <tr> <td>preambleTerm</td> <td>keyword used to classify preamble type</td> </tr> <tr> <td>preambleTermAlt</td> <td>alternative keyword (if preambleTerm were not used)</td> </tr> <tr> <td>preambleTextStub</td> <td>first 15 words of the preamble</td> </tr> <tr> <td>bodyType</td> <td>body type</td> </tr> <tr> <td>bodyLinesStep</td> <td>number of steps in the body</td> </tr> <tr> <td>bodyLinesElement</td> <td>number of elements in the body</td> </tr> <tr> <td>bodyLinesTotal</td> <td>total number of identified lines in the body</td> </tr> <tr> <td>label</td> <td>2-character label of the preamble-body combination; classification table maps label to claim type</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><em>3.3. List of Variables for Files with Patent-Level Information</em></p> <p>For detailed descriptions, see the appendix in Ganglmair, Robinson, and Seeligson (2022).</p> List of Variables (Patent-Level Information) <table><tbody> <tr> <td>patent_id</td> <td>U.S. patent number (8-digit patent number)</td> </tr> <tr> <td>claims</td> <td>number of independent claims (the sum of the four claim types: 0, 1, 2, and 3)</td> </tr> <tr> <td>noCategory</td> <td>number of claims without a classified type</td> </tr> <tr> <td>processClaims</td> <td>number of process claims</td> </tr> <tr> <td>productClaims</td> <td>number of product claims</td> </tr> <tr> <td>prodByProcessClaims</td> <td>number of product-by-process claims</td> </tr> <tr> <td>firstClaim</td> <td>type of the first claim (1 = process; 2 = product; 3 = product-by-process; 0 = no type)</td> </tr> <tr> <td>simpleProcessClaims</td> <td>number of process claims by simple approach (terms "method" or "process" anywhere in the claim)</td> </tr> <tr> <td>simpleProcessPreamble</td> <td>number of process claims by simple approach (terms "method" or "process" in the preamble)</td> </tr> <tr> <td>meansClaims</td> <td>number of means-plus-function claims</td> </tr> <tr> <td>meansFirst</td> <td>=1 if first claim is a means-plus-function claim</td> </tr> <tr> <td>JepsonClaims</td> <td>number of Jepson claims</td> </tr> <tr> <td>JepsonFirst</td> <td>=1 if first claim is a Jepson claim</td> </tr> </tbody> </table> <p><br>Note: The following variables/fields are currently empty (March 30, 2020); we will populate these variables/fields with data version 3.4.0.</p> <p>preambleTerm<br>preambleTermAlt<br>preambleTextStub<br>bodyLinesStep<br>bodyLinesElement<br>bodyLinesTotal</p> <p>Note: We will release the data for patents issued in 2021 with data version 3.4.0.</p> <p><br><strong>4. Coming Soon!</strong></p> <p>We are working on a number of extensions of the patccat data.</p> <p>- With data version 3.4.0, we plan to release data for all published U.S. patent applications (2001 through 2021)<br>- In late spring/early summer 2022, we will release data for patents issued by the European Patent Office (EPO) [<strong>Update: March 28, 2023</strong>: see <a href="https://doi.org/10.5281/zenodo.7776092">https://doi.org/10.5281/zenodo.7776092</a>]<br>- In late spring/early summer 2022, we will release data for patents issued by the Canadian Intellectual Property Office (CIPO)</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo44/100

One Classifier Ignores a Feature

<p>The data sets are used in a controlled experiment, where two classifiers should be compared. train_a.csv and explain.csv are slices from the original data set. train_b.csv contains the same instances as in train_a.csv, but with feature x1 set to 0 to make it unusable to classifier B.</p> <p>The original data set was created and split using this Python code:</p> <pre><code class="language-python">from sklearn.datasets import make_classification from sklearn.model_selection import train_test_split from sklearn.linear_model import LogisticRegression X, y = make_classification(n_samples=300, n_features=2, n_redundant=0, n_informative=2, n_clusters_per_class=1, class_sep=0.75, random_state=0) X *= 100 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.5, random_state=0) lm = LogisticRegression() lm.fit(X_train, y_train) clf_a = lm clf_b = LogisticRegression() X2 = X.copy() X2[:, 0] = 0 X2_train, X2_test, y2_train, y2_test = train_test_split(X2, y, test_size=0.5, random_state=0) clf_b.fit(X2_train, y2_train) X_explain = X_test y_explain = y_test</code></pre>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record