Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “Online comments”
Online Real-Time Delphi Survey for the research project "MENARA" - Compilation of all Comments to Closed and Open Questions
<p><strong>Looking into the Futures: Delphi Survey about the MENA region</strong></p> <p>In order to get a more realistic overview of the situation and trends, of the potentials, problems and potentials of the countries of the MENA region a Real Time Delphi survey was conducted. This is an important tool of modern future research. It was managed by the IZT- Institute for Future Studies in Berlin. A group of 139 experts and researchers from different institutes and organizations were invited to participate at the Online Real-Time Delphi Survey (RTD) about possible and likely futures of the MENA region. The experts were asked to answer questions and provide their opinions on twelve topics such as social unrest, youth unemployment, urbanization, gender equality, security etc. In this dataset all comments to the closed and the open questions are compiled.</p> <p>The output was one of the basic material used for the creation of future regional scenarios for mid-term (2025) and long-term (2050) time horizons. Focus scenarios were produced in order to exemplify selected characteristic and important future options, in terms of chances and risks (e.g. energy futures).</p>
Online supplementary data linked to the publication "Aubenas-les-Alpes (S-E France). Part III – Last and final part of the mammalian assemblage with some comments on the palaeoenvironment and palaeobiogeography" doi:10.1016/j.annpal.2019.03.001
<p>Online supplementary appendix including the list of Oligocene localities and associated faunal lists compared to Aubenas-les-Alpes, and the size estimation of the non-predatory species for the construction of Fig.10.</p>
Bubble reachers and uncivil discourse in polarized online public sphere comments dataset
<p>This dataset contains comments in Portuguese and English gathered from various sources, such as news websites from Brazil and Canada, social media sites like Facebook and Reddit, e-commerce reviews, Wikipedia comments, among others. Each comment is accompanied by a "toxicity" score provided by the Perspective API.</p> <p><strong>Disclaimer</strong>: This file includes words or language that is considered profane, vulgar or offensive by some readers. Due to the topic studied in this article, quoting offensive language is academically justified, but we nor PLOS in no way endorse the use of these words or the content of the quotes. Likewise, the quotes do not represent the opinions of us or that of PLOS, and we condemn online harassment and offensive language.</p> <p>Column information:</p> <ul> <li><strong>preprocessed_text</strong>: the text after undergoing preprocessing steps;</li> <li><strong>dataset</strong>: the given name of the dataset;</li> <li><strong>source</strong>: the dataset's source name;</li> <li><strong>dataset_source</strong>: a combination of the dataset name with its source to facilitate data aggregation tasks;</li> <li><strong>TOXICITY</strong>: a continuous score between 0.0 and 1.0 provided by the Perspective API.</li> </ul> <p>In addition to the comments, there is a spreadsheet containing analyses referenced in the article associated with this dataset.</p>
Dataset: Bengali Online Comments
<p>This dataset contains almost 16,000 Bengali comments collected manually, each annotated with labels indicating different categories of content: troll, sexual, religious, threat, and non-bully. The data is organized in an Excel file. Designed for NLP tasks such as content classification, moderation, and toxicity detection, this dataset is particularly useful for applications in social media monitoring, automated content moderation, and the analysis of user-generated content in Bengali, enabling the detection and filtering of potentially harmful or inappropriate comments.</p>
HOCON34k: A Corpus of Hate speech in Online Comments from German Newspapers
<p>We have compiled a dataset containing 34,223 comments in German, authored by users from online-platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by a group of 29 volunteers, using a binary annotation approach. The inter-rater reliability for hate speech is 0.4428 across all annotators and increases to 0.6078 when considering an optimized subset of 12 annotators, as measured by Fleiss’ Kappa. Additionally, we present a baseline text classification using BERT, achieving an MCC-score up to 0.32 and an F2-score up to 0.64 in our initial experiment on this new corpus. The data set, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.</p>
Historic news comments from the G1 brazilian online newspaper
<p>This dataset contains data of user comments about the news published in the brazilian online newspaper G1 from 2010 and before. It contains 54,634 comments distributed between 4,549 news, which includes news published date, comments texts (in brazilian portuguese) and comments published date. That dataset intentionally ommits the titles, links, usernames or other kind of data that can redirect back to someones identity.</p>
News comments from Folha de São Paulo brazilian online newspaper
<p>This dataset contains data of user comments about the news published in the brazilian online newspaper Folha de São Paulo. It contains 36,437 comments distributed between 2,665 news, which includes news published date and keywords, and comments texts (in brazilian portuguese), amount of likes and published date. That dataset intentionally ommits the titles, links, usernames or other kind of data that can redirect back to someones identity. This data is respective to the first 50 comments of every news published in the online newspaper at the month of jun. 2020.</p> <p>The data was used in the published work "Manifestação do conteúdo de ódio em comentários de notícias na versão online da Folha de São Paulo".</p>
News comments from the G1 brazilian online newspaper
<p>This dataset contains data of user comments about the news published in the brazilian online newspaper G1, between 28 mar. 2020 and 11 set. 2020. It contains 1,059,672 comments distributed between 18,014 news, which includes news published date and cited entities, and comments texts (in brazilian portuguese), amount of likes and published date. That dataset intentionally ommits the titles, links, usernames or other kind of data that can redirect back to someones identity.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.