Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2
datasets available to search
ShareScore release 0.9.0
Dataset results
2 results for “Wikipedia conversations”
The Online Conversation Threads Repository (Slashdot, Barrapunto, Wikipedia talk)
<p>This repository contains datasets with online conversation threads collected and analyzed by different researchers. Currently, you can find datsets from different news aggregators (Slashdot, Barrapunto) and the English Wikipedia talk pages.</p> <p>- Slashdot conversations (Aug 2005 - Aug 2006) Online conversations generated at Slashdot during a year. Posts and comments published between August 26th, 2005 and August 31th, 2006. For each discussion thread: sub-domains, title, topics and hierarchical relations between comments. For each comment: user, date, score and textual content. This dataset is different from the Slashdot Zoo social network (it is not a signed network of users) contained in the SNAP repository and represents the full version of the dataset used in the CAW 2.0 - Content Analysis for the WEB 2.0 workshop for the WWW 2009 conference that can be found in several repositories such as Konect Barrapunto conversations (Jan 2005 - Dec 2008)</p> <p>- Online conversations generated at Barrapunto (Spanish clone of Slashdot) during three years. For each discussion thread: sub-domains, title, topics and hierarchical relations between comments. For each comment: user, date, score and textual content Wikipedia (2001 - Mar 2010)</p> <p>- Data from articles discussions (talk) pages of the English Wikipedia as of March 2010. It contains comments on about 870,000 articles (i.e. all articles which had a corresponding talk page with at least one comment), in total about 9.4 million comments. The oldest comments date back to as early as 2001.</p> <p> </p>
WAC Corpus - Wikipedia Abusive Conversations
<p>This repository contains conversations between Wikipedia editors, which are annotated in terms of various types of abuse, at the level of messages. This corpus is described in the following publication:</p> <p>N. Cécillon, V. Labatut, R. Dufour, and G. Linarès, “WAC: A Corpus of Wikipedia Conversations for Online Abuse Detection,” in <em>12th Language Resources and Evaluation Conference</em>, 2020, pp. 1375–1383. ⟨<a href="https://hal.archives-ouvertes.fr/hal-02497514">hal-02497514</a>⟩</p> <p>The repository also contains the figures shown in this article.</p> <p><strong>Sources. </strong>Our corpus aligns two existing corpora:</p> <ul> <li>Messages and conversation structures of <em>WikiConv</em> (<a href="https://github.com/conversationai/wikidetox/tree/master/wikiconv">https://github.com/conversationai/wikidetox/tree/master/wikiconv</a>)</li> <li>Manual annotations in toxicity of <em>Wikipedia Comment Corpus</em> (WCC -- <a href="https://doi.org/10.6084/m9.figshare.4054689">https://doi.org/10.6084/m9.figshare.4054689</a>)</li> </ul> <p><strong>Citation. </strong>If you use this dataset, please cite the above article.</p> <p><br><code>@InProceedings{Cecillon2020,</code><br><code> author = {Cécillon, Noé and Labatut, Vincent and Dufour, Richard and Linarès, Georges},</code><br><code> title = {{WAC}: A Corpus of {W}ikipedia Conversations for Online Abuse Detection},</code><br><code> booktitle = {12\textsuperscript{th} Language Resources and Evaluation Conference},</code><br><code> year = {2020},</code><br><code> pages = {1375-1383},</code><br><code> address = {Marseille, FR},</code><br><code> url = {http://www.lrec-conf.org/proceedings/lrec2020/pdf/2020.lrec-1.172.pdf},</code><br><code>}</code></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.