Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

23

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

23 results for “text classification”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data set for (binary) text classification, involving spoken utterances and written text

<p>This data set contains sentences belonging to either of two classes: Transcripts of spoken<br> (informal) text (Class 0), and written, formal text (Class 1). Sentences in Class 0 were<br> obtained from publicly available transcripts of radio shows (e.g. NPR),<br> whereas Sentences in Class 1 were obtained from Wikipedia.</p> <p>The data set is divided into&nbsp;three subsets: Training, validation, and test (specified&nbsp;by the file names).<br> Each set contains a large number of sentences, belonging to either of the two classes:</p> <p>In total, there are 13,640,458 sentences, of which 6,374,487 in Class 0 and 7,265,971.<br> The training set contains 9.743,188 sentences (of which 4,553,205 in Class 0 and 5,189,983 in Class 1),&nbsp;<br> the validation set contains 1,948,639 sentences (of which 910,641 in Class 0 and 1,037,998 in Class 1), and the&nbsp;<br> test set contains 1,948,631 sentences (of which 910,641 in Class0 and 1,037,990 in Class1).&nbsp;</p> <p>The data sets are in plain text format. Every row contains (i) the class label (0 or 1) and<br> (ii) the text of the sentence, separated from the class label by a tab character.</p> <p>Note that the&nbsp;sentences contain 5 tokens or more (including punctuation marks).&nbsp;&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

TCAB: Text Classification Attack Benchmark Dataset

<p>TCAB is a large collection of successful adversarial attacks on state-of-the-art&nbsp;text classification models trained on multiple sentiment and abuse&nbsp;domain datasets.</p> <p>The dataset is broken up into 2&nbsp;files: <em>train.csv and</em>&nbsp;<em>val.csv</em>.&nbsp;The training set contains 1,448,751&nbsp;instances (552,364&nbsp;are &quot;clean&quot; unperturbed instances) and&nbsp;the validation set contains 482,914&nbsp;instances (178,607&nbsp;are &quot;clean&quot;). Each instance contains the&nbsp;following attributes:</p> <p><strong>scenario</strong>: Domain, either&nbsp;<em>abuse</em>&nbsp;or&nbsp;<em>sentiment</em>.</p> <p><strong>target_model_dataset</strong>: Dataset being attacked.</p> <p><strong>target_model_train_dataset</strong>: Dataset the target model trained on.</p> <p><strong>target_model</strong>: Type of victim model (e.g.,&nbsp;<em>bert</em>,&nbsp;<em>roberta</em>,&nbsp;<em>xlnet</em>).</p> <p><strong>attack_toolchain</strong>: Open-source attack toolchain, either&nbsp;TextAttack or OpenAttack.</p> <p><strong>attack_name</strong>: Name of the attack method.</p> <p><strong>original_text</strong>: Original input text.</p> <p><strong>original_output</strong>: Prediction probabilities of the target model on the original text.</p> <p><strong>ground_truth</strong>: Encoded label for the original task of the domain dataset. 1 and 0 means toxic and toxic for abuse datasets, respectively. 1 and 0 means positive and negative sentiment for sentiment datasets. If there is a neutral sentiment, then 2, 1, 0 means positive, neutral, and negative sentiment.</p> <p><strong>status</strong>: Unperturbed example if &quot;clean&quot;; successful adversarial attack if &quot;success&quot;.</p> <p><strong>perturbed_text</strong>: Text after it has been perturbed by an attack.</p> <p><strong>perturbed_output</strong>: Prediction probabilities of the target model on the perturbed text.</p> <p><strong>attack_time</strong>: Time taken to execute the attack.</p> <p><strong>num_queries</strong>: Number of queries performed while attacking.</p> <p><strong>frac_words_changed</strong>: Fraction of words changed due to an attack.</p> <p><strong>test_index</strong>: Index of&nbsp;each unique source&nbsp;example (original instance) (LEGACY - necessary for backwards compatibility).</p> <p><strong>original_text_identifier</strong>: Index of&nbsp;each unique source&nbsp;example (original instance).</p> <p><strong>unique_src_instance_identifier</strong>: Primary key to uniquely identify to every source instance; comprised of&nbsp;(<em>target_model_dataset</em>,&nbsp;<em>test_index</em>,&nbsp;<em>original_text_identifier</em>).</p> <p><strong>pk</strong>: Primary key to uniquely identify every attack instance; comprised of&nbsp;(<em>attack_name</em>,&nbsp;<em>attack_toolchain</em>,&nbsp;<em>original_text_identifier</em>,&nbsp;<em>scenario</em>,&nbsp;<em>target_model</em>,&nbsp;<em>target_model_dataset</em>,&nbsp;<em>test_index).</em></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

A recent overview of the state-of-the-art elements of text classification - dataset

<p>The two available datasets were used to conduct the quantitative analysis of the text classification area. The set, such as:</p> <ol> <li>biblio.bib contains all articles that are grouped in categories</li> <li>biblio.csv contains processed records from biblio.bib, based on it were built the statistics presented in the article</li> </ol>

opencc-by-4.0Oct 2017View details →
zenodo44/100

DataCI Continuous Text Classification Example Using Yelp Dataset

<p>We are using the <a href="https://www.yelp.com/dataset">Yelp Review Dataset</a> as the streaming data source for the DataCI example. We have processed the Yelp review dataset into a daily-based dataset by its `date`. In this dataset, we will only use the data from 2020-09-01 to 2020-11-30 to simulate the streaming data scenario. We are downloading two versions of the training and validation datasets:</p> <ul> <li>`yelp_review_train@2020-10`: from 2020-09-01 to 2020-10-15</li> <li>`yelp_review_val@2020-10`: from 2020-10-16 to 2020-10-31</li> <li>`yelp_review_train@2020-11`: from 2020-10-01 to 2020-11-15</li> <li>`yelp_review_val@2020-11`: from 2020-11-16 to 2020-11-30</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Dataset used in the publication "Using of Transformers Models for Text Classification to Mobile Educational Applications"

<p>Dataset used in the publication "Using of Transformers Models for Text Classification to Mobile Educational Applications".</p> <p>More info about the dataset can be found in the published article.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Multi-ideology ISIS/Jihadist White Supremacist (MIWS) Dataset for Multi-class Extremism Text Classification

<p>**************Information on how to use our Multi-ideology ISIS/Jihadist White Supremacist Dataset(MIWS)for Multi-class Extremism Text Classification ************</p> <p>Folder name: Seed_MIWS<br> Sub Folder 1 : Seed_Dataset<br> Inside this folder there are two .csv files.<br> 1) ISIS/Jihadist_Seed_Dataset<br> 2) White_Supremacist_Seed_Dataset</p> <p>These files have common features as:<br> ***************************** Common Features in Seed ******************************************************<br> *********Source :- Contains Author, Article Name or Hyperlink to Article************************************<br> *********Type_of_Source :- Whether Source is Research Article or Report or Website**************************<br> *********Text :- Contains Extremist Text provided in Source*************************************************<br> *********Ideology :- Ideology of Text mentioned in Source i.e. ISIS/Jihadist or White Supremacist***********<br> *********Label :- Labels for Text provided by Source i.e. Propaganda, Radicalization or Recruitment*********<br> *********Geographical_Location :- Location mentioned in Text. Geographical Location is manually identified**<br> *********Author_Country_Affiliation :- Country of origin of Research Article, Report or Website in Source***</p> <p>Sub Folder 2 : MIWS<br> Inside this folder ther is one .csv file:</p> <p>It contains features as:<br> *****************************Features in MIWS********************************************************************************<br> **********Tweet_ID :- Unique Identification for a Tweet provided by Twitter**************************************************<br> **********Created_Date :- Date and Time at whic Tweet was created or posted**************************************************<br> **********Geo_Enabled :- Boolean value. True if location is made public by User**********************************************<br> **********Geographical_Location :- Manually extracted list of Locations within the tweet. &#39;Undefined&#39; if no location present*<br> **********Ideology :- Manually provided during Tweet Collection i.e. ISIS/Jihadist or White Supremacist**********************<br> **********Labels :-&nbsp; Annotated by comparing with Seed, i.e. Propaganda, Radicalization and Recruitment***********************</p> <p>MIWS file can be used to collect tweets and train model for extremism detection.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

Classification of unstructured text in types of violence against women using text mining and Machine learning techniques

<p>These are the data used for the development of the investigation.</p> <p>This file was extracted from our mongoDB database. The data set contains real news of violence against women, which were organized with their date, the title and the body of the news.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Classification of hierarchical text using geometric deep learning: the case of clinical trials corpus

<p>We consider the hierarchical representation of documents as graphs and use geometric deep learning to classify them into different categories. While graph neural networks can efficiently handle the variable structure of hierarchical documents using the permutation invariant message passing operations, we show that we can gain extra performance improvements using our proposed selective graph pooling operation that arises from the fact that some parts of the hierarchy are invariable across different documents. We applied our model to classify clinical trial (CT) protocols into completed and terminated categories. We use bag-of-words based as well as pre-trained transformer-based embeddings to featurize the graph nodes, achieving f1-scores $\simeq 0.85$ on a publicly available large scale CT registry of around 360K protocols. We further demonstrate how the selective pooling can add insights into the CT termination status prediction.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Systematic Literature Review about Text Classification

<p>This work represents a Systematic Literature Review about Text Classification with search keys and query results. Also the classification criteria of the examples for the SLR.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Classification of Text Data on Rare Diseases

<div>A dataset with text and labels for&nbsp;3 categories:&nbsp;</div> <div> <div>&nbsp;</div> <div>- Rare Diseases</div> <div>- Non-Rare Diseases</div> <div>- Other</div> </div> <div>&nbsp;</div> <div>It is a subset of abstracts obtained from PubMed and sorted into the 3 classes on the basis of their MeSH terms.</div> <div>&nbsp;</div> <div>The dataset is provided for demonstration and methodology validation purposes. The original PubMed data was randomly under-sampled.&nbsp;</div> <p>The dataset consists of 3 files in the Tab-Separated Values (TSV) format, corresponding to the 3 splits used in the article:</p> <p>Rei L, Pita Costa J, Zdol&scaron;ek Draksler T. Automatic Classification and Visualization of Text Data on Rare Diseases. _Journal of Personalized Medicine_. 2024; 14(5):545. https://doi.org/10.3390/jpm14050545</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Hierarchical Text Classification corpora

<p>A set of 3 datasets for Hierarchical Text Classification (HTC), with samples divided into training and testing splits. The hierarchies of labels within all datasets have depth 2.</p> <ul> <li>The <strong>Amazon5x5</strong> dataset contains 500,000 user reviews tagged with the reviewed product's categories. There are 5 product categories with 100,000 examples each, and each category has 5 sub-categories.</li> <li>The <strong>Bugs</strong> dataset contains 30,050 bugs of the Linux kernel, labeled with exactly two categories identifying the affected component.</li> <li>Finally, the <strong>Web Of Science</strong> dataset contains 46,960 abstracts of scientific papers, labeled the article's domain (see <a href="https://data.mendeley.com/datasets/9rw3vkcfy4/6">original repo</a> for more details).</li> </ul> <p>Datasets are published in JSONL format, where each line is a string formatted as a JSON, like in the example below.</p> <pre><code>{ "text": &lt;article text&gt;, "labels": [&lt;label1&gt;, &lt;label2&gt;, ...] }</code></pre> <p>The <em>hierarchical structure</em> of labels in each dataset is documented in <a href="https://gitlab.com/distration/dsi-nlp-publib/-/tree/main/htc-survey-24/data/taxonomies">this repository</a>.</p> <p>&nbsp;</p> <p>These datasets have been presented in this paper:</p> <ul> <li>"Hierarchical Text Classification and its Foundations: a Review of Current Research" - DOI: <a href="https://doi.org/10.3390/electronics13071199">10.3390/electronics13071199</a></li> </ul> <p>Some of these datasets have also been used in:</p> <ul> <li>"Ticket Automation: an Insight into Current Research with Applications to Multi-level Classification Scenarios" - DOI: <a href="https://doi.org/10.1016/j.eswa.2023.119984">10.1016/j.eswa.2023.119984</a></li> <li>"A multi-level approach for hierarchical Ticket Classification", accepted at WNUT 2022 - <a href="https://aclanthology.org/2022.wnut-1.22/">link</a></li> </ul> <p>&nbsp;</p> <p>These datasets are partially derived from previous work, namely:</p> <ul> <li>[Amazon] J. Ni, J. Li, J. McAuley, "Justifying Recommendations using Distantly-Labeled Reviews and Fine-Grained Aspects", EMNLP 2019, doi: <a href="http://dx.doi.org/10.18653/v1/D19-1018">10.18653/v1/D19-1018</a></li> <li>[WOS] K. Kowsari, D. E. Brown, M. Heidarysafa, K. Jafari Meimandi, M. S. Gerber and L. E. Barnes, "HDLTex: Hierarchical Deep Learning for Text Classification," 2017 16th IEEE International Conference on Machine Learning and Applications (ICMLA), 2017, pp. 364-371, doi: <a href="http://doi.org/10.1109/ICMLA.2017.0-134">10.1109/ICMLA.2017.0-134</a></li> <li>[Linux Bugs] V. Lyubinets, T. Boiko and D. Nicholas, "Automated Labeling of Bugs and Tickets Using Attention-Based Mechanisms in Recurrent Neural Networks," <em>2018 IEEE Second International Conference on Data Stream Mining &amp; Processing (DSMP)</em>, 2018, pp. 271-275, doi: <a href="http://doi.org/10.1109/DSMP.2018.8478511">10.1109/DSMP.2018.8478511</a></li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo36/100

From Networks of Texts to Networks of Genres? On the Classification of Texts in Compilations with a View towards Manuscript Transmission

<p>At the end of the 14th century, Jakob Twinger von K&ouml;nigshofen, a cleric from Strasbourg, composed&nbsp;a chronicle in the vernacular that spread widely &ndash; up to today nearly 130 manuscripts&nbsp;are known that contain the text, wholly or in parts, and that were produced not only in&nbsp;Strasbourg, but as far as Cologne, Augsburg, or Tyrol. About thirty qualify as true copies,&nbsp;while in the big majority of the witnesses, the text is altered in various ways: abbreviated,&nbsp;augmented, updated, corrected, put in a dierent order, and more often than not combined&nbsp;with other texts, either with distinct text boundaries or resulting in new compositions made&nbsp;of several texts.<br> In several manuscripts, a combination of historiographical texts &ndash; chronicles, annals, lists,&nbsp;etc. &ndash; can be observed, leading to the assumption that Twinger&rsquo;s work was preferably copied&nbsp;for historiographical compilations and that its structure facilitated some historiographical&nbsp;activity of the recipients. But this view puts a big weight on the genre of Twinger&rsquo;s work,&nbsp;making it a filter through which all the other texts in a codex are seen. As I have been&nbsp;working on the chronicle transmission, the attribute in common of the known manuscripts&nbsp;is of course the appearance of at least some lines of the Twinger chronicle; but the attempt&nbsp;to look at a single codex as objectively as possible, without giving preference to a particular&nbsp;text, can not only reveal connections between different manuscripts, but also the fluidity&nbsp;and flexibility of medieval texts. The classification of a work as part of a certain genre does&nbsp;not necessarily hold true for its entire transmission, for the manifestation of a distinct text&nbsp;&ndash; what is, on the one hand, a nuisance, but on the other a chance to better understand&nbsp;medieval text transmission, manuscript production and the transfer of knowledge.<br> The co-occurrence of certain texts in several manuscripts hints towards intentional copying&nbsp;processes that are reflecting particular interests not only of one individual scribe or commissioner;&nbsp;multiple occurrences of particular compilations can reveal networks that go unseen&nbsp;if the content of a codex is not regarded as a whole. While a text-based analysis often faces&nbsp;difficulties that result from insufficient manuscript descriptions, a broader view that would&nbsp;compare less particular texts, but more areas of interest or fields of knowledge, has to deal&nbsp;with problems regarding the classification of the single texts: Apart from an inevitable subjectivity,<br> questions about the criteria and levels of classification have to be addressed, while&nbsp;the danger of over- or under-representation is always lurking. Discussing these issues and the&nbsp;applicability and usefulness of the comparison of codices from a kind-of-genre-perspective&nbsp;could be fruitful to develop a better understanding for the transmission of manuscripts, of&nbsp;texts and of knowledge.</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

Text-fig. 14: Classification of the Gobioninae according to Naseka (1996). in Revision Of The Cyprinids From The Early Oligocene Of The České Středohoří Mountains, And The Phylogenetic Relationships Of Protothymallus Laube, 1901 (Teleostei, Cyprinidae, Gobioninae)

Text-fig. 14: Classification of the Gobioninae according to Naseka (1996).

opencc-by-4.0Dec 2007View details →
zenodo36/100

Multi-label Datasets used in "Adapting Transformers for Multi-Label Text Classification"

<p>The three Multi-Label&nbsp;datasets used in the article &quot;Adapting Transformers for Multi-Label Text Classification&quot;.</p> <p>- AAPD Dataset&nbsp;&nbsp;(ArXiv Academic Paper Dataset) [Yang et al. 2018]<sup>1</sup></p> <p>- Reuters-21578 Dataset:&nbsp;https://archive.ics.uci.edu/ml/datasets/reuters-21578+text+categorization+collection</p> <p>- MFHAD (Multilabel French HAL Abstracts Dataset)</p> <p>&nbsp;</p> <p><sup>1</sup>Pengcheng Yang, Xu Sun, Wei Li, Shuming Ma, Wei Wu, and Houfeng Wang. 2018.<br> SGM: Sequence Generation Model for Multi-label Classification. In Proceedings<br> of the 27th International Conference on Computational Linguistics. Association for<br> Computational Linguistics, Santa Fe, New Mexico, USA, 3915&ndash;3926.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

TeCla: Text Classification Catalan dataset

<p><em>Corpus de not&iacute;cies en catal&agrave; per a classificaci&oacute; textual, extret del web de l&#39;<a href="http://www.acn.cat">Ag&egrave;ncia Catalana de Not&iacute;cies</a> sota llic&egrave;ncia CC-BY-NC-ND</em></p> <p>TeCla (Text Classification) is a Catalan News corpus for thematic multi-class Text Classification tasks. The present version (2.0) contains 113.376 articles classified under a hierarchical class structure consisting of a coarse-grained and a fine-grained class. Each of the 4 coarse-grained classes accept a subset of fine-grained ones, 53 in total.</p> <p>The source data is crawled from the ACN (Catalan News Agency) site: <a href="http://www.acn.cat">http://www.acn.cat</a>, and used under CC-BY-NC-ND 4.0 licence. The dataset is released under the same licence, and is intended exclusively for training Machine Learning models.</p> <p>This dataset was developed by BSC TeMU as part of the AINA project, and intended as part of CLUB (Catalan Language Understanding Benchmark).</p>

opencc-by-nc-nd-4.0Mar 2021View details →
zenodo32/100

AlleNoise - large-scale text classification benchmark dataset with real-world label noise

<div> <div> <div> <div> <p><span>AlleNoise</span><span> is a benchmark dataset for large-scale multi-class text classification with real-world label noise. It consists of e-commerce product titles from Allegro.com with corresponding category labels. The noise distribution comes from actual users of a major e-commerce marketplace, so it realistically reflects the semantics of human mistakes. In addition to the noisy labels, we provide human-verified clean labels and a meaningful, hierarchical taxonomy of categories. Code and data is available at https://github.com/allegro/AlleNoise.<br></span></p> </div> </div> </div> </div>

opencc-by-nc-nd-4.0Jun 2024View details →
zenodo32/100

DBpedia abstracts for text classification

<p>These are the dataset from DBpedia version April 1st 2017 required for testing text classification models.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

DBpedia derived abstracts for text classification

<p>The derived datasets obtained after processing the raw texts from the DBpedia dataset.</p>

opencc-by-4.0Jul 2024View details →
zenodo28/100

CLASSIFICATION OF ADVERTISING TEXTS.

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo24/100

Data associated with Automated Biomedical Text Classification with Research Domain Criteria

<p>Each file contains a set of abstracts for a given RDoC construct.&nbsp; Each file is named after one RDoC category. In each file, each line represents one PubMed abstract that belong to the RDoC category. Each line has a PubMeD ID and an abstract text&nbsp; separated by a tab. The dataset was created on August 2018.&nbsp;</p> <p>Please cite: M. Anani and I. Kahanda, Automated Biomedical Text Classification with Research Domain Criteria, International Conference on Bioinformatics and Computational Biology, Las Vegas, NV, 2018.</p>

opencc-by-4.0Dec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record