Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,250

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,250 results for “Classification”

Learn how ShareScore rates datasets ↗
dryad40/100

Virus classification for viral genomic fragments using PhaGCN2

<p>Viruses are the most ubiquitous and diverse entities in the biome. Due to the rapid growth of newly identified viruses, there is an urgent need for accurate and comprehensive virus classification, particularly for novel viruses. Here, we present PhaGCN2, which can rapidly classify the taxonomy of viral sequences at family level and supports the visualization of the associations of all families. We evaluate the performance of PhaGCN2 and compare it with the state-of-the-art virus classification tools, such as vConTACT2, CAT, and VPF-Class, using the widely accepted metrics. The results show that PhaGCN2 largely improves the precision and recall of virus classification, increases the number of classifiable virus sequences in the Global Ocean Virome dataset (v2.0) by 4 times, and classifies more than 90% of the Gut Phage Database. PhaGCN2 makes it possible to conduct high-throughput and automatic expansion of the database of the International Committee on Taxonomy of Viruses.</p>

opencc-zeroApr 2022View details →
zenodo40/100

Radar-based severe thunderstorm classification in Switzerland from 2016-2021

<p>The data is owned by MeteoSwiss and the Laboratoire de T&eacute;l&eacute;d&eacute;tection Environnementale, EPFL. It is linked to the research article: Feldmann et al., 2022, Hailstorms and rainstorms versus supercells - a regional analysis of convective storm types in the Alpine region, submitted to npj Climate and Atmospheric Science. Whenever using the dataset, please include a reference to this article.</p> <p>This dataset contains thunderstorm track data of thunderstorms severe thunderstorms in Switzerland from 2016-2020. Tracked variables include time, position, size, translational speed and several intensity variables.</p> <p>The README.txt file contains a comprehensive list of all included variables and their units. The data is provided as a&nbsp; .txt file. The methods used to obtain this data are described in the corresponding publication (Feldmann et al., 2022).</p>

opencc-by-4.0May 2022View details →
zenodo40/100

A tempοral Deep Convolutional Neural Network model on Sentinel-1 Image Time Series for pixel-wise Flood Classification (dataset)

<p>This is a dataset which has been designed to be used for flood time series classification. Each time series is annotated as flood or no-flood and represents a pixel-wise time series derived from stack of Sentinel-1 IW GRD images that have been pre-processed according to <a href="http://doi.org/10.5281/zenodo.6510223">https://doi.org/10.5281/zenodo.6510223</a>.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

ReSpa - Towards an automatic requirements classification in a new Spanish dataset

<p>ReSpa&nbsp;(Spanish Dataset for requirements classification) dataset is conformed by requirements collected from final&nbsp;degree projects from one University. It was presented in the paper &#39;Towards an automatic requirements classification in a new Spanish dataset&#39; and used also in the paper &#39;Requirements Classification Using FastText and BETO in Spanish Documents&#39;.</p> <p>Cited as:</p> <p>Limaylla-Lunarejo, M. I., Condori-Fernandez, N., &amp; Luaces, M. R. (2022, August). Towards an automatic requirements classification in a new Spanish dataset. In&nbsp;<em>2022 IEEE 30th International Requirements Engineering Conference (RE)</em>&nbsp;(pp. 270-271). IEEE.&nbsp;https://doi.org/10.1109/RE54965.2022.00039</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Crop Classification Data-set

<p>The dataset consists of crop type training and testing data of more than 10 classes, collected using Ground Truth Surveys, in Harichand region of Khyber Pakhtoonkhwa, Pakistan.&nbsp;</p> <p>The dataset also contains 2 tiff files having Planet-Scope and Sentinel-2 raster data.&nbsp;</p> <p>https://drive.google.com/drive/folders/1SweabTezj78btq9wd3PRZYWrR_4gUuC4</p>

opencc-by-4.0May 2022View details →
zenodo40/100

PhytoNodes for Environmental Monitoring: Stimulus Classification based on Natural Plant Signals in an Interactive Energy-efficient Bio-hybrid System

<p>Cities worldwide are growing, putting bigger populations at risk due to urban pollution. Environmental monitoring is essential and requires a major paradigm shift. We need green and inexpensive means of measuring at high sensor densities and with high user acceptance. We propose using phytosensing: using natural living plants as sensors. In plant experiments we gather electrophysiological data with sensor nodes. We expose the plant <em>Zamioculcas zamiifolia</em> to five different stimuli: wind, temperature, blue light, red light, or no stimulus. Using that data we train ten different types of artificial neural networks to classify measured time series according to the respective stimulus. We achieve good accuracy and succeed in running trained classifying artificial neural networks online on the microcontroller of our small energy-efficient sensor node. To indicate later possible use cases, we showcase the system by sending a notification to a smartphone application once our continuous signal analysis detects a given stimulus.</p> <p>&nbsp;</p> <p>Data repository for our paper &quot;PhytoNodes for Environmental Monitoring: Stimulus Classification based on<br> Natural Plant Signals in an Interactive Energy-efficient Bio-hybrid System&quot;, submitted to the GoodIT conference. Please refer to the paper for more information.</p> <p>&nbsp;</p> <p><strong>Contents of this repository</strong></p> <ul> <li><em>mu_interface:</em> Code for our data collection plant experiments, based on Raspberry Pis and the <a href="http://cybertronica.co/?q=products/phytosensor">Cybertronica phytosensing and phytoactuating system</a>.</li> <li><em>raw_data: </em>The datasets from our plant experiments for the stimuli wind, temperature, red light, blue light, and no stimulus.</li> <li><em>dl-4-tsc:</em> Deep learning framework developed by <a href="https://doi.org/10.1007/s10618-019-00619-1">Fawaz et. al (Deep learning for time series classification: a review)</a> and adapted to our use case. Find the training and testing datasets in the archives folder as well as the trained classifiers in the results folder.</li> <li><em>classification_results.ods: </em>Overview of the results from the deep learning framework (accuracy, precision, recall, training time).</li> <li><em>TFLite_Models: </em>The trained classifiers in TensorFlow Lite Format.</li> <li><em>00_AI_BLE_MeasuringOnlyWind: </em>Source code for classification on STM-based PhytoNodes (using MCDCNN two-class classifier) and Bluetooth communication. The code is written for the STM32WB55 Nucleo board and can be transferred to the dongle.</li> <li><em>zavrsniProjekt_iOS: </em>Source code of the iOS app used to receive data from the STM-based PhytoNodes.</li> <li><em>Watchplant_application_documentation.pdf: </em>Instructions to build and use the iOS app.</li> </ul>

opencc-by-4.0Jun 2022View details →
zenodo40/100

ARIPO patent data for IVD-related patent classifications

<p>ARIPO Dataset_v2.ods dataset was scraped from the ARIPO E-Service database using Octoparse&nbsp;using the IPC classifications below. The exported task is provided as ARIPO Database_Copy.otd:</p> <table> <tbody> <tr> <td> <p><strong>IPC Code</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>B01L 3/00</p> </td> <td> <p>Containers or dishes for laboratory use, e.g. laboratory glassware (bottles B65D; apparatus for enzymology or microbiology C12M 1/00); Droppers (receptacles for volumetric purposes G01F) [2006.01]</p> </td> </tr> <tr> <td> <p>G01N 33/50</p> </td> <td> <p>Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing (measuring or testing processes other than immunological involving enzymes or microorganisms, compositions or test papers therefor, processes of forming such compositions, condition responsive control in microbiological or enzymological processes C12Q) [2006.01]</p> </td> </tr> <tr> <td> <p>C12N 9/00</p> </td> <td> <p>Enzymes, e.g. ligases (6.); Proenzymes; Compositions thereof (preparations containing enzymes for cleaning teeth A61K 8/66, A61Q 11/00; medicinal preparations containing enzymes or proenzymes A61K 38/43; enzyme containing detergent compositions C11D); Processes for preparing, activating, inhibiting, separating, or purifying enzymes [2006.01]</p> </td> </tr> <tr> <td> <p>C12N 11/00</p> </td> <td> <p>Carrier-bound or immobilised enzymes; Carrier-bound or immobilised microbial cells; Preparation thereof [2006.01]</p> </td> </tr> <tr> <td> <p>C12N 15/00<br> &nbsp;</p> </td> <td> <p>Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor (mutants or genetically engineered microorganisms C12N 1/00, C12N 5/00, C12N 7/00; new plants A01H; plant reproduction by tissue culture techniques A01H 4/00; new animals A01K 67/00; use of medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases, gene therapy A61K 48/00; peptides in general C07K) [2006.01]</p> </td> </tr> <tr> <td> <p>C12Q</p> </td> <td> <p>Measuring or testing processes involving enzymes, nucleic acids or microorganisms (immunoassay G01N 33/53); compositions or test papers therefor; processes of preparing such compositions; condition-responsive control in microbiological or enzymological processes [3]</p> </td> </tr> </tbody> </table> <p><strong>Octoparse Method</strong></p> <ol> <li> <p>Go to ARIPO Advanced Search website</p> </li> </ol> <ol> <li> <p>Uncheck Utility Model, Industrial Design, Trademark checkboxes</p> </li> <li> <p>Enter date range 01011960~04222022</p> </li> <li> <p>Enter IPC categories as per Appendix 1 [only one per search]</p> </li> <li> <p>Click Search</p> </li> <li> <p>Adjust items per page from 10 to 40</p> </li> <li> <p>Extract data, looping through each record on the page</p> </li> <li> <p>Click next page button and repeat</p> </li> <li> <p>Click next button to bring up next five pages and repeat until end.</p> </li> </ol> <p>&nbsp;</p> <p><strong>ARIPO_Journal_Dataset_v2.ods</strong></p> <p>This dataset was downloaded from the ARIPO journal publication server using a browser-based scraping software.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-zeroMay 2022View details →
zenodo40/100

Dataset for tumor infiltrating lymphocyte classification (304,097 image patches from TCGA)

<p>This is a dataset of images with or without tumor-infiltrating lymphocytes (TILs). The original images are from Abousamra et al. (2022) and Saltz et al. (2018), and the original whole slide images are from TCGA. This dataset is a subset of the data presented in Abousamra et al. (2022) (with new data partitions).</p> <p>If you use this dataset, please cite the following papers, as well as this Zenodo page.</p> <p>Abousamra, S., Gupta, M. D., Hou, L., Batiste, R., Zhao, T., Shankar, A., Rao, A., Chen, C., Samaras, D., Kurc, T., &amp; Saltz, J. (2022). Deep Learning-Based Mapping of Tumor Infiltrating Lymphocytes in Whole Slide Images of 23 Types of Cancer. <em>Frontiers in Oncology</em>, 5971. https://doi.org/10.3389/fonc.2021.806603</p> <p>Saltz, J., Gupta, R., Hou, L., Kurc, T., Singh, P., Nguyen, V., Samaras, D., Shroyer, K. R., Zhao, T., Batiste, R., &amp; Danilova, L. (2018). Spatial organization and molecular correlation of tumor-infiltrating lymphocytes using deep learning on pathology images. <em>Cell Reports</em>, <em>23</em>(1), 181-193.</p> <p>&nbsp;</p> <p>The acknowledgements from the <em>Frontiers in Oncology</em> and <em>Cell Reports</em> papers are included below:</p> <blockquote> <p>This work was supported by the National Institutes of Health (NIH) and National Cancer Institute (NCI) grants UH3-CA22502103, U24-CA21510904, 1U24CA180924-01A1, 3U24CA215109-02, and 1UG3CA225021-01 as well as generous private support from Bob Beals and Betsy Barton. AR and AS were partially supported by NCI grant R37-CA214955 (to AR), the University of Michigan (U-M) institutional research funds and also supported by ACS grant RSG-16-005-01 (to AR). AS was supported by the Biomedical Informatics &amp; Data Science Training Grant (T32GM141746). This work was enabled by computational resources supported by National Science Foundation grant number ACI-1548562, providing access to the Bridges system, which is supported by NSF award number ACI-1445606, at the Pittsburgh Supercomputing Center, and also a DOE INCITE award joint with the MENNDL team at the Oak Ridge National Laboratory, providing access to Summit high performance computing system. The funders were not involved in the study design, collection, analysis, interpretation of data, the writing of this article or the decision to submit it for publication.</p> </blockquote> <p>&nbsp;</p> <blockquote> <p>We are grateful to all the patients and families who contributed to this study. Funding from the Cancer Research Institute is gratefully acknowledged, as&nbsp;is&nbsp;support from National Cancer Institute (NCI) through U54 HG003273, U54 HG003067, U54 HG003079, U24 CA143799, U24 CA143835, U24 CA143840, U24 CA143843, U24 CA143845,U24 CA143848, U24 CA143858, U24 CA143866, U24 CA143867, U24 CA143882, U24 CA143883, U24 CA144025, P30 CA016672, U24CA180924, U24CA210950, U24CA215109, NCI Contract HHSN261201400007C, and Leidos Biomedical Contract 14X138. A.U.K.R. and P.S were supported by CCSG Bioinformatics Shared Resource P30 CA01667, ITCR U24 Supplement 1U24CA199461-01, a gift from Agilent technologies, CPRIT RP150578, and a Research Scholar Grant from the American Cancer Society (RSG-16-005-01). This work used the Extreme Science and Engineering Discovery Environment (XSEDE), which is supported by National Science Foundation XSEDE Science Gateways program under grant ACI-1548562 allocation TG-ASC130023. The authors would like to thank Stony Brook Research Computing and Cyberinfrastructure and the Institute for Advanced Computational Science at Stony Brook University for access to the high-performance LIred and SeaWulf computing systems, the latter of which was supported by National Science Foundation grant (#1531492).</p> </blockquote> <p>------------------------------------</p> <p>This dataset includes 304,097 image patches. All images are 100 x 100 pixels at 0.5 micrometers per pixel. An image is TIL-positive if there are at least two TILs present.</p> <p>Refer to `images-tcga-tils-metadata.csv` for information about each image. That spreadsheet has the following columns:</p> <pre><code>partition,study,barcode,label,path,md5</code></pre> <p>Partition specifies which partition the image is part of (train, val, test). Study is the TCGA study the image is part of (e.g., acc for TCGA-ACC). Barcode is the TCGA participant barcode. This is used during partitioning, to ensure that images from the same participant are not present in different data partitions. Label is either til-negative or til-positive. An image is til-positive if there are at least two TILs in the image. Path is the path to the PNG image. All images are stored as PNG. Md5 is the md5 hash of the image. This can be used to ensure there are no duplicate images and to verify the integrity of images.</p> <p>There are study-specific directories in the directory `images-tcga-tils`, and there is a directory named `pancancer` that includes images from all the included TCGA studies. That directory uses symlinks to avoid storing duplicate data.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

MEDIC: A Multi-Task Learning Dataset for Disaster Image Classification

<p>Recent research in disaster informatics demonstrates a practical and important use case of artificial intelligence to save human lives and suffering during natural disasters based on social media contents (text and images). While notable progress has been made using texts, research on exploiting the images remains relatively under-explored. To advance image-based approaches, we propose MEDIC\footnote{Available~at: \url{https://crisisnlp.qcri.org/medic/index.html}}, which is the largest social media image classification dataset for humanitarian response consisting of 71,198 images to address four different tasks in a multi-task learning setup. This is the first dataset of its kind: social media images, disaster response, and multi-task learning research. An important property of this dataset is its high potential to facilitate research on \textit{multi-task learning}, which recently receives much interest from the machine learning community and has shown remarkable results in terms of memory, inference speed, performance, and generalization capability. Therefore, the proposed dataset is an important resource for advancing image-based disaster management and multi-task machine learning research.&nbsp;<br> &nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Fig. 3 in Phylogeny and classification of Odonata using targeted genomics

Fig. 3. Current state of odonate phylogeny. Summary of the phylogenetic hypothesis for Odonata from Fig. 2. Support values for each node can be found in Fig. 2. Grey text highlights both the reinstated (Tatocnemididae stat. res., Rhipidolestidae stat. res.) and the proposed new families (Protolestidae fam. nov., Priscagrionidae fam. nov., Mesopodagrionidae fam. nov., Mesagrionidae fam. nov., and Amanipodagrionidae fam. nov.). Discussion of Grps. 1, 2, 3, and 4 is found in the Calopterygoidea section of the manuscript. Almost all major lineages are included and now have a hypothesized phylogenetic placement. The zygopteran genera Rimanella (Rimanellidae) and Sciotropis (Incertae Sedis group 8) were not included in this analysis.

opencc-by-4.0Feb 2021View details →
zenodo40/100

Fig. 2 in Phylogeny and classification of Odonata using targeted genomics

Fig. 2. AHE topology. Results of the phylogenetic reconstruction of Odonata using loci captured from anchored hybridization enrichment. A) ML and Bayesian phylogenetic reconstruction of the Odonata using 478 loci. AZ in the right hand column represents the suborder Anisozygoptera. See "key to nodal support" for a visual guide to nodal support. Support values of 100 bootstrap and 1.0 posterior probability are not shown. Quartet sampling (QS) that shows full support (1/NA/1) is not shown at the node but all other QS is shown at each node along with a measure of taxon rogueness (QF score) at each branch tip. Bolded GF scores represent the lowest values across the topology. Newly established or reestablished families are shown in grey text. B) Arepresentation of the branch lengths are shown with all three suborders designated.

opencc-by-4.0Feb 2021View details →
zenodo40/100

Chess piece dataset for image classification

<p>Chess piece dataset for image classification. Contains 4 different chess sets, 3 used for training and the remainder for validation purposes. Each chess piece from each set has been photograph by a static bird&#39;s eyes camera from each of the 64 squares that form a chess board. This way, each piece is seen from all its different angles.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

A revised phylogenetic classification for Viola (Violaceae)

<p>The genus <em>Viola</em> (Violaceae) is among the 40&ndash;50 largest genera among angiosperms, yet its taxonomy has not been revised for nearly a century. In the most recent revision, by Wilhelm Becker in 1925, the then known 400 species were distributed among 14 sections and numerous unranked groups. Here we provide an updated, comprehensive classification of the genus, based on data from phylogeny, morphology, chromosome counts, and ploidy, and based on modern principles of monophyly. The revision is presented as an annotated global checklist of accepted species of <em>Viola</em>, an updated multigene phylogenetic network and an ITS phylogeny with denser taxon sampling, a brief summary of the taxonomic changes from Becker&rsquo;s classification and their justification, a morphological binary key to the accepted subgenera, sections and subsections, and an account of each infrageneric subdivision with justifications for delimitation and rank including a description, a list of apomorphies, molecular phylogenies where possible or relevant, a distribution map, and a list of included species. We distribute the 664 species accepted by us into 2 subgenera, 31 sections, and 20 subsections. We erect one new subgenus of <em>Viola</em> (subg. <em>Neoandinium</em>, a replacement name for the illegitimate subg. <em>Andinium</em>), six new sections (sect. <em>Abyssinium</em>, sect. <em>Himalayum</em>, sect. <em>Melvio</em>, sect. <em>Nematocaulon</em>, sect. <em>Spathulidium</em>, sect. <em>Xanthidium</em>), and seven new subsections (subsect. <em>Australasiaticae</em>, subsect. <em>Bulbosae</em>, subsect. <em>Clausenianae</em>, subsect. <em>Cleistogamae</em>, subsect. <em>Dispares</em>, subsect. <em>Formosanae</em>, subsect. <em>Pseudorupestres</em>). Evolution within the genus is discussed in light of biogeography, fossil record, morphology, and particular traits. <em>Viola</em> is among very few temperate and widespread genera that originated in South America. The biggest identified knowledge gaps for <em>Viola</em> concern the South American taxa, for which basic knowledge from phylogeny, chromosome counts, and fossil data is virtually absent. <em>Viola</em> has also never been subject to comprehensive anatomical study. Study on seed anatomy and morphology is required to understand the fossil record of the genus.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

CNN for the classification of ICE-CAMERA images of Antarctic ice particles

<p>-The file &#39;ZENODO_FILES.rar&#39; contains the GoogleNet Convolutional Neural Network (CNN) trained to classify pre-processed ICE-CAMERA images (224*224*3) into 14 classes. CNN was developed for (Mathworks) MATLAB&reg; R2020b.</p> <p>-The ICE-CAMERA images used for training, validation and testing the CNN are also contained in specific folders.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Animal Recognition Using Methods Of Fine-Grained Visual Analysis - YOLOv5 Breed Classification Dataset (Oxford-IIIT Pet)

<p>Oxford-IIIT Pet Dataset with ground truth labels for breeds&nbsp;(from https://public.roboflow.com/object-detection/oxford-pets).</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Evaluation of tools for the taxonomical classification of viruses

<p>50G (1942_50GL), 500G (1943_500GL) and 1000G (1752_1000GL) sets are simulated reads previously reported Tangherlini et al. [1], and represent large reads like contigs.</p> <p>FISH-I (FISH_1_Q_sS_dup_HNoM_RbNoM), PB3(PB3_Q_sS_dup_HNoM_RbNoM), I5-8(ISA32_R1_Q_CDhit_HmNoM_RbNoM and ISA32_R1_Q_CDhit_HmNoM_RbNoM) and 121-1 (ISA44_R1_Q_CDhit_HmNoM_RbNoM and ISA44_R1_Q_CDhit_HmNoM_RbNoM) sets are real metagenomics reads previously reported Taboada et al. [2,3].</p> <p>Eukaryotic, Prokaryotic, Unclassified, Bacterial and Human sets are simulated reads, generated by Grinder [4] software and represent short reads like those of Illumina technology.<br> &nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Classification of unstructured text in types of violence against women using text mining and Machine learning techniques

<p>These are the data used for the development of the investigation.</p> <p>This file was extracted from our mongoDB database. The data set contains real news of violence against women, which were organized with their date, the title and the body of the news.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Data Cleaning, Translation & Split of the Dataset for the Automatic Classification of Documents for the Classification System for the Berliner Handreichungen zur Bibliotheks- und Informationswissenschaft

<ul> <li>Cleaned_Dataset.csv &ndash; The combined CSV files of all scraped documents from DABI, e-LiS, o-bib and Springer.</li> <li>Data_Cleaning.ipynb &ndash; The Jupyter Notebook with python code for the analysis and cleaning of the original dataset.</li> <li>ger_train.csv &ndash; The German training set as CSV file.</li> <li>ger_validation.csv &ndash; The German validation set as CSV file.</li> <li>en_test.csv &ndash; The English test set as CSV file.</li> <li>en_train.csv &ndash; The English training set as CSV file.</li> <li>en_validation.csv &ndash; The English validation set as CSV file.</li> <li>splitting.py &ndash; The python code for splitting a dataset into train, test and validation set.</li> <li>DataSetTrans_de.csv &ndash; The final German dataset as a CSV file.</li> <li>DataSetTrans_en.csv &ndash; The final English dataset as a CSV file.</li> <li>translation.py &ndash; The python code for translating the cleaned dataset.</li> </ul>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Classification of hierarchical text using geometric deep learning: the case of clinical trials corpus

<p>We consider the hierarchical representation of documents as graphs and use geometric deep learning to classify them into different categories. While graph neural networks can efficiently handle the variable structure of hierarchical documents using the permutation invariant message passing operations, we show that we can gain extra performance improvements using our proposed selective graph pooling operation that arises from the fact that some parts of the hierarchy are invariable across different documents. We applied our model to classify clinical trial (CT) protocols into completed and terminated categories. We use bag-of-words based as well as pre-trained transformer-based embeddings to featurize the graph nodes, achieving f1-scores $\simeq 0.85$ on a publicly available large scale CT registry of around 360K protocols. We further demonstrate how the selective pooling can add insights into the CT termination status prediction.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Systematic Literature Review about Text Classification

<p>This work represents a Systematic Literature Review about Text Classification with search keys and query results. Also the classification criteria of the examples for the SLR.</p>

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record