Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “Offensive language”
Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying
<p>Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying.</p> <p> </p> <p>* To cite this dataset refer to <a href="http://dx.doi.org/10.12785/ijcds/130177" target="_blank" rel="nofollow noopener">http://dx.doi.org/10.12785/ijcds/130177</a><br>Mazari, A. C., & Kheddar, H. (2023). "Deep Learning-based Analysis of Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying." IJCDS, 13(1).</p> <p> </p> <div> <p>* Due to the nature of this Dataset, comments contain offensiveness and hate speech. This does not reflect author values, however the aim is to providing a resource to help in detecting and preventing spread of such harmful content.</p> </div> <div> <h3>Features</h3> <ul> <li>Algerian Dialect</li> <li>Cyberbullying</li> <li>Hate speech</li> <li>Offensive Language</li> <li>Dialect Dataset</li> </ul> </div>
Offensive content dataset in Urdu language
<p>The archive contains python code and various feature files of offensive language dataset in urdu. The purpose of sharing this archive is to regenerate the results produced by the research article and can extend the findings.</p>
SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)
<p>The task involves three subtasks corresponding to the hierarchical taxonomy of the OLID schema (Zampieri et al., 2019) from OffensEval 2019. The task featured five languages and this upload is for the English language. In addition, English also featured Subtasks B and C. OffensEval 2020 was one of the most popular tasks at SemEval-2020 attracting a large number of participants across all subtasks and also across all languages. A total of 528 teams signed up to participate in the task, 145 teams submitted systems during the evaluation period, and 70 submitted system description papers.</p> <p>This upload includes a test set used in the paper describing the dataset used in the shared task as well as the official test set used in the shared task.</p> <p>The evaluation phase for English is available on Codalab: <a href="https://competitions.codalab.org/competitions/23285">https://competitions.codalab.org/competitions/23285</a></p> <p>The Website for the shared task is <a href="https://sites.google.com/site/offensevalsharedtask/home">https://sites.google.com/site/offensevalsharedtask/home</a></p>
Pashto Offensive Language Dataset
<p>Visit my <strong><a href="https://ijaz.me/">Website</a></strong></p> <p>Visit my <strong><a href="https://github.com/dr-ijaz/nlpashto" target="_blank" rel="noopener">GitHub</a></strong></p> <p>Visit my <strong><a href="https://www.linkedin.com/in/drijazulhaq/">LinkedIn</a></strong></p> <p>Pre-trained Model on <strong><a href="https://huggingface.co/ijazulhaq">HuggingFace</a></strong></p>
Influence of offensive language usage on Anger modulation - Dataset
<p>The dataset aims to find the response elicited by after usage of swear words on anger management in people during distinct patience-demanding situations. A total of 150 participants (75 men and 75 women aged between 18 - 34 years) rated their anger level on a 6-point scale, with respect to the scenarios under three different conditions. baseline anger level during the scenario, anger level on suppressing the reaction, anger level after using offensive language. Valid responses (139) are highlighted in green and invalid responses (11) are highlighted in red.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.