Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

598

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

598 results for “classifier”

Learn how ShareScore rates datasets ↗
zenodo40/100

Xception trained model for classifying large ornithopod dinosaur footprints

<p>PLOS ONE: Classification of large ornithopod dinosaur footprints using Xception transfer learning</p><p>The trained model using Xception transfer learning, provided in https://github.com/CNUGeophysics/Xception_ornithopod.git</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Diversity-aware Fairness Testing of Machine Learning Classifiers through Hashing-based Sampling

<p>The experimental results of the evaluation of VBT-X.</p> <h2>Abstract</h2> <div> <h3>Context:</h3> <p>There are growing concerns about algorithmic fairness, as some machine learning (ML)-based algorithms have been found to exhibit biases against protected attributes such as gender, race, age and so on. Individual fairness requires an ML classifier to produce similar outputs for similar individuals. Verification Based Testing (<span>Vbt</span>) is a state-of-the-art black-box testing algorithm for individual fairness that leverages constraint solving to generate test cases.</p> </div> <div> <h3>Objective:</h3> <p>Generating diverse test cases is expected to facilitate efficient detection of diverse discriminatory data instances (i.&nbsp;e., cases that violate individual fairness). Hashing-based sampling techniques draw a sample approximately uniformly at random from the set of solutions of given Boolean constraints. We propose <span>Vbt</span>-X, which improves <span>Vbt</span> with hashing-based sampling, aiming to improve its testing performance.</p> </div> <div> <h3>Method:</h3> <p>We realize hashing-based sampling for <span>Vbt</span>. The challenge is that the off-the-shelf hashing-based sampling techniques cannot be integrated in a straightforward manner because the constraints in <span>Vbt</span> are generally not Boolean. Moreover, we propose several enhancement techniques to make <span>Vbt</span>-X more efficient.</p> </div> <div> <h3>Results:</h3> <p>To evaluate our method, we conduct experiments, where <span>Vbt</span>-X is compared to <span>Vbt</span>, <span>Sg</span> and ExpGA (other well-known fairness testing algorithms) over a set of configurations consisting of several datasets, protected attributes, and ML classifiers. The results show that, with each configuration, <span>Vbt</span>-X detects more discriminatory data instances with higher diversity than <span>Vbt</span> and <span>Sg</span>. <span>Vbt</span>-X detects discriminatory data instances with higher diversity than ExpGA, though the number of discriminatory data instances detected by <span>Vbt</span>-X is lesser than ExpGA.</p> </div> <div> <h3>Conclusion:</h3> <p>Our proposed method performs better than other state-of-the-art black-box fairness testing algorithms, particularly in terms of diversity. Our method can serve to efficiently identify flaws in ML classifiers with respect to individual fairness for subsequent improvements of an ML classifier. On the other hand, although our method is specific to individual fairness, it could work for testing other aspects of a software system such as security and counterfactual explanations with some technical adaptations, which remains for future work.</p> </div> <p>&nbsp;</p> <div> <h2>Acknowledgments</h2> <p>This paper is partly based on results obtained from a project, JPNP20006, commissioned by the New Energy and Industrial Technology Development Organization (NEDO). This paper is supported by JST SPRING, Grant Number JPMJSP2131.</p> </div>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Data Sets ''Kombucha--Proteinoid Biosynthetic Classifiers of Audio Signals''

<p>These voltage signals represent the electrical response generated by the Kombucha-Proteinoid biosynthetic system when exposed to various audio signals. The analysis of this voltage signal data contributes to the understanding and development of novel audio classification methods based on organic systems.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Lund Bladder Cancer Group - Lund Taxonomy 2023 Classifier - Training Data

<p>Raw and processed training data used for the Lund Taxonomy 2023 gene expression classifier for urothelial carcinoma.</p><p>The dataset contains the raw .CEL files for 3 Affymetrix Gene 1.0 ST cohorts: Lund2017 (n=307, GSE83586), Lund2020 (n=173, GSE128959), Lund2022 (n=310, 117 from GSE169455, 37 from GSE222073, and 156 previously unpublished samples). Code examples for RMA and SCAN.UPC normalization for Affymetrix Gene, Exon, and HTA2 arrays used in the study is included. Hybridization batch information is available for each array cohort.</p><p>The dataset contains Kallisto and Salmon transcript quantification files for 265 RNA-sequenced tumors and the tx2gene file+code for transcript to gene summarization using txImport</p><p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Classifying Transients and Variable Objects with SCONE

<p>In this work, we expand the use of the Supernova Classifier with a Convolutional Neural Network (SCONE) to include classification of an additional two classes of transient objects and five classes of variable objects. SCONE has been adapted for these additional types of objects by updating its data processing pipeline to better handle variable objects, which exhibit repeated changes in brightness over time, as well as introducing class weights to ensure the model picks up on the nuances of the imbalanced training and test sets (both pulled from the PLAsTiCC dataset). Preliminary results show SCONE is capable of classifying transients with 80% accuracy and variable stars with 70% accuracy. To further improve its performance, we are now working on an original method of simulating light curves using templates found in the PLAsTiCC dataset to create a larger and more representative test set.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Dataset: A multi-label classifier for predicting the most appropriate instrumental method for the analysis of contaminants of emerging concern

<p>NORMAN Suspect List Exchange was used for the generation of the dataset. Datasets with clear label (LC or GC) were used. More specifically, we used S3 NORMANCT15, which contains a list of compounds that were detected in surface water from the Danube River in a pan-European collaborative trial employing both GC-HRMS and LC-HRMS. Moreover, the GC and LC target list were used by the following two institutes: National and Kapodistrian University of Athens (NKUA) and Helmholtz Centre for Environmental Research (UFZ). S21 UATHTARGETS is the LC target list of NKUA, S65 UATHTARGETSGC is the GC target list of NKUA and S53 UFZWANATARG contains the LC and GC target list of UFZ. Finally, two GC target lists (S51 WRIGCHRMS and S70 EISUSGCEIMS) were used. These lists contain GC substance lists and were provided by two Slovak institutes, the Water Research Institute (WRI) and Environmental Institute. The aforementioned compound lists were merged together to form a labelled dataset. The SMILES were used to calculate 1446 molecular descriptors. 1446 descriptors were produced by PaDEL-descriptor, logP was produced by JRgui and boiling point by USEPA ECOSAR.</p> <p>The dataset is used in the publication:</p> <p>&quot;A multi-label classifier for predicting the most appropriate instrumental method for the analysis of contaminants of emerging concern&quot; authored by</p> <p>Nikiforos Alygizakis, Vasileios Konstantakos, Grigoris Bouziotopoulos , Evangelos Kormentzas, Jaroslav Slobodnik and Nikolaos S. Thomaidis</p> <p>Github repository:&nbsp;https://github.com/nalygizakis/LCvsGC</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

supplementary materials: Classifying Be star variability with TESS I: the southern ecliptic

<p>The archived files contain ascii light curves (LC_data.tgz) and plots (LC_plots.tgz) from TESS cycle 1 for the Be star sample of &quot;Classifying Be star variability with TESS I: the southern ecliptic&quot;. The light curve files have columns: 1) TESS JD (JD - 2457000), 2) relative flux, 3) photometric error.&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text

<p>Sentiment Analysis outputs based on the combination of three classifiers for news headlines and body text covering the Olympic legacy of Rio 2016 and London 2012. Data was searched via Google search engine. It is composed of sentiment labels assigned to 1271 news articles in total.</p> <p><strong>News outlets:</strong></p> <ul> <li>BBC</li> <li>Daily Mail</li> <li>The Telegraph</li> <li>The Guardian</li> <li>Globo</li> <li>Estadao</li> <li>Folha de S. Paulo</li> </ul> <p><strong>Events covered by the articles:</strong></p> <ul> <li>London 2012 Olympic legacy</li> <li>Rio 2016 Olympic legacy</li> </ul> <p>All classifiers were used in texts in English. Text originally published in Portuguese by the Brazilian media were automatically translated.</p> <p><strong>Sentiment classifiers used:</strong></p> <ul> <li>Vader</li> <li>BERT (Trained on Amazon data)</li> <li>BERT (Trained on twitter data - 140)</li> </ul> <p>Each document (spreadsheet - xlsx) refers to one outlet and one event (London 2012 or Rio 2016).</p> <p><strong>How were labels assigned to the texts?</strong></p> <p>These labels are a combination of the three sentiment classifiers listed above. If two of them agree with the same label, then this label would be considered as right. Otherwise, the label &lsquo;other&rsquo; was assigned.</p> <p>For news article body text: the proportion of sentences of each sentiment type was used to assign labels to the whole article instead of averaging the sentence scores. For example, if the proportion of sentences with negative labels is greater than 50%, then the article is assigned a negative label.</p> <p><strong>The documents are composed of the following columns:</strong></p> <ul> <li>Rank: the position of the article on Google search ranking</li> <li>Date: date of article&#39;s publication (DD/MM/YYYY)</li> <li>Link: article&#39;s link</li> <li>Title: article&#39;s title</li> <li>Sentiment_Title: final sentiment for article headline</li> <li>Sentiment_Text: final sentiment for article&#39;s body text</li> </ul> <p><em>PS: Documents do not include articles&#39; body text. </em></p> <p><strong>Sentiment is presented in labels as follows:</strong></p> <ul> <li>Pos: Positive</li> <li>Neg: Negative</li> <li>Neutral: Neutral</li> <li>other: inconclusive - if each of the 3 classifiers assigned a different label to the article, the label &#39;other&#39; was used. Therefore, &#39;other&#39; identifies contradictory results.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

InpactorDB: A Plant classified lineage-level LTR retrotransposon reference library for free-alignment methods based on Machine Learning

<p>LTR retrotransposons are mobile elements that make up the major part of most plant genomes. Their identification and annotation via bioinformatics approaches represent a major challenge in the era of massive plant genome sequencing. In addition to their involvement in the variation in genome size, these elements are also associated in the function and structure of different chromosomal regions and in the alteration of the function of coding regions, among others. Several plant retrotransposon sequence databases of LTR retrotransposons are available with public access such as PGSB, RepetDB or restricted access such as Repbase. Although they are useful for approaches to identify LTR-RTs in new genomes by similarity, the elements of these databases are not classified down to the lineage/family level. with great depth.&nbsp;</p> <p>Here, we present InpactorDB a semi-curated dataset composed of 130,511 elements from 195 plant genomes (belonging to 108 plant species), classified down to the lineage level. This data set has been used to train two deep neural networks (one fully connected and one convolutional) for fast classification of elements. Used in lineage-level classification approaches, we obtain a score above 98% of F1-score, precision and recall.&nbsp;</p> <p>In order to classify elements of the &lsquo;LTR_STRUC&rsquo; and &lsquo;EDTA&rsquo; datasets, we used the methodology proposed by Inpactor, which uses homology-based strategy with known coding domains belonging to LTR-RTs. We utilized the RexDB &nbsp;domain library as reference. LTR-RTs were classified into superfamilies, Gypsy (RLG) or Copia (RLC) and sub-classified into lineages according to the similarities of five different amino acid reference domains (GAG, AP, RT, RNAseH, and INT domains). In addition, we applied filters to remove keep only intact elements:</p> <p>1) to remove predicted elements with domains from two different superfamilies (i.e. Gypsy and Copia),</p> <p>2) or elements with domains belonging to two or more different lineages,</p> <p>3) to remove elements with lengths different than those reported by Gypsy Database (GyDB) with a tolerance of 20%,</p> <p>4) to delete incomplete elements which has less than three identified domains, and</p> <p>5) to remove elements with insertions of TE class II (reported in Repbase).&nbsp;</p> <p>The final non-redundant version of InpactorDB consists of 67,305 LTR retrotransposons. Both redundant and non-redundant versions of InpactorDB are available in Fasta&nbsp;format in which sequences have identifiers with the following general&nbsp;Identification code:</p> <p>&gt;Superfamily-Lineage-plant_family-specie-source-length-ID,</p> <p>Where Superfamily&nbsp;can&nbsp;is either RLC (for Copia) or RLG (for&nbsp;Gypsy), Lineage/family&nbsp;follows&nbsp;following&nbsp;the RexDB nomenclature, source&nbsp;(can be&nbsp;Repbase, RepetDB, PGSB, LTR_STRUC or EDTA&nbsp;datasets), length, and ID,&nbsp;is&nbsp;a unique number which identify each element inside&nbsp;the&nbsp;InpactorDB.</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Brazilian tweets classified for sentiment analysis

<p>Brazilian tweets classified for sentiment analysis</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

ARIA - Accessible Rich Internet Applications Landmarks classified dataset

<p>ARIA Landmarks classified dataset. ARIA Landmarks are regions of web applications that can be accessed by specific keyboard shortcuts and are specially useful for blind users. They comprise the banner, complementary, contentinfo, form, main, navigation, region and&nbsp;search classes.</p> <p>In the dataset, samples represent DOM elements of web applications with different attributes extracted when these elements are rendered in the browser.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Malaria Stage Classifier dataset

<p>This is the dataset for the Malaria Stage Classifier, which introduces a new method for the stage-specific classification of malaria-infected red blood cells (RBCs) and provides a fast, high-accuracy recognition even with limited training sets by a smart reduction of data dimension. RBCs are extracted from an image, reduced to characteristic one-dimensional cross-sections, and classified by a pretrained neural network. The method is applicable to images recorded by various microscopy techniques. The dataset can be used to retrain the neural network with new data.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Endmember spectra and classified maps derived from CRISM targeted data at the south pole of Mars

<p><strong>Overview</strong></p> <p>Current maps of compositional variation across south polar ice exposures on Mars do not resolve the meter-scales at which erosional processes are most active, ultimately limiting our understanding of how the deposits form and evolve and how they can be used to interpret long-term climate records. In this study, we use&nbsp;<em>k</em>-means clustering and random forest classification to identify and map a set of universal spectral endmembers across 167 high-resolution observations acquired during southern summer by the Compact Reconnaissance Imaging Spectrometer for Mars (CRISM). The 21 endmembers show distinct combinations and strengths of key&nbsp;infrared absorption features reflecting diverse mixtures of CO<sub>2</sub> ice, H<sub>2</sub>O ice, and dust. The resulting compositional framework&nbsp;can be used to characterize the nature of both seasonal CO<sub>2</sub> frost and the residual ices it overlies across a variety of terrains.&nbsp;</p> <p>&nbsp;</p> <p><strong>Contents</strong></p> <p>The repository contains three .zip files, which can be expanded to access the files described below:</p> <ul> <li><strong>classified_maps.zip</strong> <ul> <li>lookup_files <ul> <li><em>SP_CRISM_RF_ColorMap.clr</em> : An ESRI-formatted color map file that can be used to apply the endmember color scheme&nbsp;to random forest-classified maps in ArcGIS (see Symbology settings).</li> <li><em>SP_CRISM_RF_ColorMap.txt</em> : A text file that can be loaded into Python as a numpy array and used to generate a matplotlib color ramp. Row indices correspond to endmember numbers (see below) and columns are red, green, and blue values scaled from 0 to 1.</li> <li><em>SP_CRISM_RF_Endmember_Lookup.csv</em> : A lookup table that can be used to cross-reference between endmember numbers (as stored in GeoTiffs or the indices of the colormaps&nbsp;above) and corresponding endmember names (C1, Dc2, etc.).</li> </ul> </li> <li>morphologic_reference <ul> <li>Contains GeoTiffs of the R1330 spectral parameter (reflectance at 1330 nm) for each processed CRISM observation. Filenames indicate the observation ID, Mars Year, and solar longitude (Ls) of acquisition (&quot;Ls314-07&quot; = Ls 314.07&ordm;). These images can be used as a reference for the surface morphology and albedo of each scene. Tie points used to georeference these images to other datasets can be applied to the random forest classification maps to properly align endmember mapping results.</li> </ul> </li> <li>random_forest_classification <ul> <li>Contains GeoTiffs of the random forest classification results for each processed CRISM observation. Filenames indicate the observation ID, Mars Year, and solar longitude (Ls) of acquisition (&quot;Ls314-07&quot; = Ls 314.07&ordm;). These images are not rendered to display colors consistent with the publication figures, but instead store the endmember classification for each pixel as a number from 0 to 21; use the contents of <em>lookup_files</em>&nbsp;to find the corresponding&nbsp;endmember name or render the image with the color scheme from the publication. To view rendered summary plots of each observation, see the contents of <em>observation_info</em>. To georeference these images to other datasets, use the corresponding morphologic reference (see above) to set tie points.</li> </ul> </li> </ul> </li> <li><strong>observation_info.zip</strong> <ul> <li>footprint_shapefile <ul> <li>Contains the components of an ESRI shapefile that outlines the surface footprint/coverage of each processed CRISM observation. The attributes associated with each observation are the same as those in&nbsp;<em>observation_lookup</em> below. Note that the polygons extend slightly beyond the area shown in maps in&nbsp;<em>random_forest_classification</em> due to the inclusion of border pixels.</li> </ul> </li> <li>observation_lookup <ul> <li><em>SP_CRISM_Classified_Obs_Info.csv</em> : Information on each processed CRISM observation; this is the same file as Table S1 in the publication. Includes the MY and Ls&nbsp;of acquisition and (where applicable) the figure panel where the observation&nbsp;appears in the publication. The location of each observation is indicated with Center Latitude/Longitude and the assigned Spatial Domain (see Figure 1 in the publication). The Observation ID can be used to locate the source Targeted Reduced Data Record (TRDRs, Version 3) on the Geosciences Node of the Planetary Data System. The spatial extent of each observation is provided in the <em>footprint_shapefile</em> described above.&nbsp;</li> </ul> </li> <li>summary_plots <ul> <li>Contains summary plots of the endmember map generated for each processed CRISM observation. Each plot notes the observation ID, Mars Year, and Ls and displays the morphologic reference map, random forest classification map, and a breakdown of the endmembers that are present. To access the maps rendered here, see the contents of <em>classified_maps.zip</em>.</li> <li>Also contains the full-resolution version of Figure S3 from the publication (<em>All_Obs_Unprojected.png</em>), which can be used to lookup observations of interest via small labels above each map.&nbsp;</li> </ul> </li> </ul> </li> <li><strong>spectral_library.zip</strong> <ul> <li>Contains spectral libraries&nbsp;with the median spectrum&nbsp;of each endmember as presented in Figure 3 in the publication. A basic text file listing the wavelength (WVL) and normalized reflectance values for each endmember is included (<em>SP_CRISM_EndmemberMedians.txt</em>) as well as an ENVI-formatted spectral library (S<em>P_CRISM_EndmemberMedians_ENVI.sli</em>). Note that a&nbsp;subset of the 438 wavelengths sampled in the source CRISM data were removed around the longest and shortest wavelengths and the filter boundary to avoid error-prone bands, leaving these spectral libraries with 404 bands.</li> </ul> </li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Figure 2 Bar chart represents the percentage of the correct diagnosis using fuzzy diagnosis, K- nearest neighbor and Naïve Bayes classifiers.-Comparison of Fuzzy Diagnosis with K-Nearest Neighbor and Naïve Bayes Classifiers in Disease Diagnosis

<p>Figure 2 Bar chart represents the percentage of the correct diagnosis using fuzzy<br> diagnosis, K- nearest neighbor and Na&iuml;ve Bayes classifiers.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Figure 1. The comparison of the Area Under the ROC Curve (AUC) for fuzzy Diagnosis, KNN and NB-Comparison of Fuzzy Diagnosis with K-Nearest Neighbor and Naïve Bayes Classifiers in Disease Diagnosis

<p>The area under the receiver operating characteristic (ROC) curve (AUC) is used to measure<br> the performance of fuzzy diagnosis, KNN and NB. In order to show the difference between AUC<br> for the three methods, a single figure which combines the three AUC for the three methods was<br> used for comparison as shown in figure 1 below.</p>

opencc-by-4.0Dec 2010View details →
zenodo40/100

Train and test datasets used for the paper "Neural network time-series classifiers for gravitational-wave searches in single-detector periods"

<p>This repository contains the datasets used for training and testing during the work discussed in the paper "<a href="https://iopscience.iop.org/article/10.1088/1361-6382/ad40f0" target="_blank" rel="noopener">Neural network time-series classifiers for gravitational-wave searches in single-detector periods</a>". Please refer to this paper for more details on how the dataset was produced and cite it if you use these data:</p> <p><em>A. Trovato et al "Neural network time-series classifiers for gravitational-wave searches in single-detector periods", Class. Quant. Grav. 2024 DOI 10.1088/1361-6382/ad40f0.</em></p> <p>In this repository you will find six files in format npz, three of which refer to the test dataset and three to the train dataset. Each file name is of the type {label}_{train or test}.npz where "label" can be "glitch", "noise" or "signal", while the second part of the name indicates whether the file was used for training or testing.</p> <p>Each file is a collection of numpy arrays so it should be read with python. It contains 3 numpy arrays: 'X', 'Y' and 'metadata'. 'X' is a matrix containing 1-second segments of data sampled at 2048 Hz of the LIGO-Livingston detector, so it has shape: (number of samples, 2048). 'Y' contains the label for each segment, which is 0 for noise, 1 for signal and 2 for glitch, so it has shape: (number of samples,). In this case, the information on 'Y' is redundant since it's given directly by the filename. The 'metadata' matrix contains 17 metadata for each sample only for the case of signals, for glitch or noise it contains just 17 zeros for each sample. The shape of 'metadata' is thus: (number of samples, 17). For the signal files, for each sample the metadata is an array with these components:</p> <ol> <li>GPS start of the file from which this segment comes</li> <li>starting GPS time of this segment</li> <li>duration of the segment [s]</li> <li>mass1 [solar masses]</li> <li>mass2 [solar masses]</li> <li>spin1z</li> <li>spin2z</li> <li>inclination [radians]</li> <li>coalescence phase [radians]</li> <li>distance [Mpc]</li> <li>right_ascension [radians]</li> <li>declination [radians]</li> <li>polarization [radians]</li> <li>SNR (signal to noise ratio)</li> <li>shift of the signal w.r.t. the timeseries [s]</li> <li>length of the signal [s]</li> <li>fraction of the signal contained in the time window</li> </ol> <p>Number of samples:</p> <ul> <li>80000 for the file glitch_test.npz</li> <li>69998 for the file glitch_train.npz</li> <li>500000 for the file noise_test.npz</li> <li>250000 for the file noise_train.npz</li> <li>500000 for the file signal_test.npz</li> <li>250000 for the file signal_train.npz</li> </ul> <p>An example of few lines of python code to read each file is:</p> <pre><code>import numpy as np f = np.load("filename.npz") X = f['X'] Y = f['Y'] m = f['metadata'] </code></pre> <p>For the preparation of these data, we acknowledge the use of the following software packages: GWpy [1], PyCBC [2] and LALSuite [3].&nbsp;</p> <p>This research has made use of data or software obtained from the Gravitational Wave Open Science Center (<a href="https://gwosc.org/" target="_blank" rel="noopener">gwosc.org</a>), a service of the LIGO Scientific Collaboration, the Virgo Collaboration, and KAGRA. This material is based upon work supported by NSF's LIGO Laboratory which is a major facility fully funded by the National Science Foundation, as well as the Science and Technology Facilities Council (STFC) of the United Kingdom, the Max-Planck-Society (MPS), and the State of Niedersachsen/Germany for support of the construction of Advanced LIGO and construction and operation of the GEO600 detector. Additional support for Advanced LIGO was provided by the Australian Research Council. Virgo is funded, through the European Gravitational Observatory (EGO), by the French Centre National de Recherche Scientifique (CNRS), the Italian Istituto Nazionale di Fisica Nucleare (INFN) and the Dutch Nikhef, with contributions by institutions from Belgium, Germany, Greece, Hungary, Ireland, Japan, Monaco, Poland, Portugal, Spain. KAGRA is supported by Ministry of Education, Culture, Sports, Science and Technology (MEXT), Japan Society for the Promotion of Science (JSPS) in Japan; National Research Foundation (NRF) and Ministry of Science and ICT (MSIT) in Korea; Academia Sinica (AS) and National Science and Technology Council (NSTC) in Taiwan.</p> <p>[1] https://gwpy.github.io<br>[2] https://pycbc.org<br>[3] https://lscsoft.docs.ligo.org/lalsuite</p>

opencc-by-4.0Apr 2024View details →
dryad40/100

Spatial behavior and diet data for discrete-choice analyses: data observed and classified from GPS video camera collars worn by female members of the Fortymile Caribou Herd across Alaska, USA, and Yukon, Canada

<p>Competition for resources and space can drive forage selection of large herbivores from the bite through the landscape scale. Animal behavior and foraging patterns are also influenced by abiotic and biotic factors. Fine-scale mechanisms of density-dependent foraging at the bite scale are likely consistent with density-dependent behavioral patterns observed at broader scales, but few studies have directly tested this assertion. Here, we tested if space use intensity, a proxy of spatiotemporal density, affects foraging mechanisms at fine spatial scales similarly to density-dependent effects observed at broader scales in caribou. We specifically assessed how behavioral choices are affected by space use intensity and environmental processes using behavioral state and forage selection data from caribou (&lt;i&gt;Rangifer tarandus granti&lt;/i&gt;) observed from GPS video-camera collars using a multivariate discrete-choice modeling framework. We found that the probability of eating shrubs increased with increasing caribou space use intensity and cover of &lt;i&gt;Salix&lt;/i&gt; spp. shrubs, whereas the probability of eating lichen decreased. Insects also affected fine-scale foraging behavior by reducing the overall probability of eating. Strong eastward winds mitigated the negative effects of insects and resulted in higher probabilities of eating lichen. Lastly, caribou exhibited foraging functional responses wherein their probability of selecting each food type increased as the availability (% cover) of that food increased. Space use intensity signals of fine-scale foraging were consistent with density-dependent responses observed at larger scales and with recent evidence suggesting declining reproductive rates in the same caribou population. Our results highlight the potential risks of overgrazing on sensitive forage species such as lichen. Remote investigation of the functional responses of foraging behaviors provides exciting future applications where spatial models can identify high-quality habitats for conservation.</p>

opencc-zeroMay 2024View details →
zenodo40/100

RDP Classifier training files for 16S rRNA sequences from GTDB

<p>16S rRNA gene sequences from the <a href="https://gtdb.ecogenomic.org/">Genome Taxonomy Database</a> (GTDB release 220) were used to retrain the <a href="https://github.com/rdpstaff/classifier">RDP Classifier</a> (version 2.13). Two sets of training files are provided:</p> <ul> <li><code>genus.zip</code> - Genus level</li> <li><code>species.zip</code> - Species level</li> </ul> <p>The code in <code>prepare_files.R</code> was used to prepare the GTDB sequence and taxonomy files for retraining the RDP Classifier. Notes:</p> <ul> <li>Steps to retrain the RDP Classifier are adapted from <a href="https://john-quensen.com/tutorials/training-the-rdp-classifier/">https://john-quensen.com/tutorials/training-the-rdp-classifier/</a></li> <li>Python scripts (lineage2taxTrain.py and addFullLineage.py) are available at <a href="https://github.com/rdpstaff/classifier/issues/18">https://github.com/rdpstaff/classifier/issues/18</a></li> <li>The first 1000 training sequences (<code>train_nodups_1000.fasta</code>) are used for benchmarking the classification accuracy (see results at end of <code>prepare_files.R</code>).</li> </ul>

opencc-by-4.0Jun 2024View details →
zenodo40/100

MedMNIST-C: Comprehensive benchmark and improved classifier robustness by simulating realistic image corruptions

<p><strong>Abstract: </strong>The integration of neural-network-based systems into clinical practice is limited by challenges related to domain generalization and robustness. The computer vision community established benchmarks such as ImageNet-C as a fundamental prerequisite to measure progress towards those challenges. Similar datasets are largely absent in the medical imaging community which lacks a comprehensive benchmark that spans across imaging modalities and applications. To address this gap, we create and open-source MedMNIST-C, a benchmark dataset based on the MedMNIST+ collection, covering 12 datasets and 9 imaging modalities. We simulate task and modality-specific image corruptions of varying severity to comprehensively evaluate the robustness of established algorithms against real-world artifacts and distribution shifts. We further provide quantitative evidence that our simple-to-use artificial corruptions allow for highly performant, lightweight data augmentation to enhance model robustness. Unlike traditional, generic augmentation strategies, our approach leverages domain knowledge, exhibiting significantly higher robustness when compared to widely adopted methods. By introducing MedMNIST-C and open-sourcing the corresponding library allowing for targeted data augmentations, we contribute to the development of increasingly robust methods tailored to the challenges of medical imaging. The code is available at <a href="https://github.com/francescodisalvo05/medmnistc-api">github.com/francescodisalvo05/medmnistc-api</a>.</p> <blockquote> <p>This work has been accepted at the Workshop on Advancing Data Solutions in Medical Imaging AI @ MICCAI 2024 [<a href="https://arxiv.org/pdf/2406.17536" target="_blank" rel="noopener">preprint</a>].</p> </blockquote> <p><strong>Note: </strong>Due to space constraints, we have uploaded all datasets except TissueMNIST-C. However, it can be reproduced via our APIs.&nbsp;</p> <p><strong>Usage:&nbsp;</strong>We recommend using the demo code and tutorials available on our GitHub&nbsp;<a href="https://github.com/francescodisalvo05/medmnistc-api">repository</a>.</p> <p><strong>Citation: </strong>If you find this work useful, please consider citing us:</p> <blockquote> <pre>@article{disalvo2024medmnist, title={MedMNIST-C: Comprehensive benchmark and improved classifier robustness by simulating realistic image corruptions}, author={Di Salvo, Francesco and Doerrich, Sebastian and Ledig, Christian}, journal={arXiv preprint arXiv:2406.17536}, year={2024} }</pre> </blockquote> <p><strong>Disclaimer: </strong>This repository is inspired by MedMNIST APIs and the ImageNet-C repository. Thus, please also consider citing <a href="https://www.nature.com/articles/s41597-022-01721-8" rel="nofollow">MedMNIST</a>, the respective source datasets (described&nbsp;<a href="https://medmnist.com/" rel="nofollow">here</a>), and&nbsp;<a href="https://arxiv.org/abs/1903.12261" rel="nofollow">ImageNet-C</a>.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 2. Correct Classified Instances for different data sets-Intelligent System for Diagnosis of a Three-Phase Separator

<p>The data mining models may be considered a superior</p> <p>technique that may be successful</p> <p>applied in diagnosis and may be develop in the futu</p> <p>re on the base of more training data to increase</p> <p>the accuracy of results.</p> <p>Industrial processes are dynamic processes with ran</p> <p>dom behavior and whose evolution over</p> <p>time cannot be predicted unless it is well known th</p> <p>e process model and use advanced predictive</p> <p>techniques. Consequently, design and implement an a</p> <p>utomated online monitoring and diagnosis</p> <p>three-phase separator remains a future direction of</p> <p>research conducted so far.</p> <p>Conceptually, this system should have permanent acc</p> <p>ess to data collected from field</p> <p>transducers, to be able to identify the type of fau</p> <p>lt occurred, to locate the fault and provide</p> <p>recommendations to remedy abnormal operating condit</p> <p>ion. Also, updating the database defects with</p> <p>new types of defects occurred and the adequate solu</p> <p>tions adopted for eliminating errors in the</p> <p>operating mode is an important feature to be consid</p> <p>ered during the design of the online diagnosis</p> <p>system. This is possible if the system would have s</p> <p>elf-learning capabilities. To acquire this &quot;skill&quot;,</p> <p>the automatic online diagnosis system may contain a</p> <p>diagnosis module based on artificial neural</p> <p>networks.</p>

opencc-by-4.0Jan 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record