Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
39
datasets available to search
ShareScore release 0.7.1
Dataset results
39 results for “reddit”
Hybrid Approaches to Detect Comments Violating Macro Norms on Reddit
<p>[<strong>Content warning: </strong><em>Files may contain instances of highly inflammatory and offensive content.]</em></p> <p><br> This dataset was generated as an extension of our <a href="https://www.cc.gatech.edu/~eshwar3/uploads/3/8/0/4/38043045/eshwar-norms-cscw2018.pdf">CSCW 2018 paper</a>:</p> <p><em>Eshwar Chandrasekharan, Mattia Samory, Shagun Jhaver, Hunter Charvat, Amy Bruckman, Cliff Lampe, Jacob Eisenstein, and Eric Gilbert. 2018. The Internet’s Hidden Rules: An Empirical Study of Reddit Norm Violations at Micro, Meso, and Macro Scales. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 32.</em></p> <p><strong>Description:</strong></p> <p>Working with over 2M removed comments collected from 100 different communities on Reddit (subreddit names listed in data/study-subreddits.csv), we identified <strong>8 macro norms</strong>, i.e., norms that are widely enforced on most parts of Reddit. We extracted these macro norms by employing a hybrid approach—classification, topic modeling, and open-coding—on comments identified to be norm violations within at least 85 out of the 100 study subreddits. Finally, we labelled over 40K Reddit comments removed by moderators according to the specific type of macro norm being violated, and make this dataset publicly available (also available on <a href="https://github.com/ceshwar/reddit-norm-violations">Github</a>).</p> <p>For each of the labeled topics, we identified the top 5000 removed comments that were best fit by the LDA topic model. In this way, we identified over 5000 removed comments that are examples of each type of macro norm violation described in the paper. The removed comments were sorted by their topic fit, stored into respective files based on the type of norm violation they represent, and are made available on this repo.</p> <p>Here we make the following datasets publicly available:</p> <p>* <strong>1 file</strong> containing the log of over 2M removed comments obtained from the top 100 subreddits between May 2016 to March 2017, after filtering out the following comments: 1) comments by u/AutoModerator, 2) replies to removed comments (i.e., children of the poisoned tree - refer to the paper for more information), and 3) non-readable comments (not utf-8 encoded).</p> <p>* <strong>8 files</strong>, each containing 5000+ removed comments obtained from Reddit, are stored in: data/macro-norm-violations/ , and they are split into different files based on the macro norm they violated. Each new line in the files represent a comment that was posted on Reddit between May 2016 to March 2017, and subsequently removed by subreddit moderators for violating community norms. All comments were preprocessed using the script in code/preprocessing-reddit-comments.py , in order to do the following: 1. remove new lines, 2. convert text to lowercase, and 3. strip numbers and punctuations from comments.</p> <p><strong>Description of 1 file</strong> containing over<em> 2M removed comments </em>from <em>100 subreddits.</em></p> <ul> <li>"reddit-removal-log.csv" - all comments that were removed from the 100 study subreddits during the study period described above (post-filtering).</li> </ul> <p><strong>Descriptions of each file</strong> containing <em>5059 comments</em> (that were removed from Reddit, and preprocessed)<strong> violating macro norms </strong>present in data/macro-norm-violations/:</p> <ul> <li>"macro-norm-violations-n10-t0-misogynistic-slurs.csv" - Comments that use misogynistic slurs.</li> <li>"macro-norm-violations-n15-t2-hatespeech-racist-homophobic.csv" - Comments containing hate speech that is racist or homophobic.</li> <li>"macro-norm-violations-n10-t3-opposing-political-views-trump.csv", "macro-norm-violations-n15-t10-opposing-political-views-trump.csv" - Comments with opposing political views around Trump (depends on originating sub).</li> <li>"macro-norm-violations-n10-t4-verbal-attacks-on-Reddit.csv" - Comments containing verbal attacks on Reddit or specific subreddits.</li> <li>"macro-norm-violations-n10-t5-porno-links.csv" - Comments with pornographic links.</li> <li>"macro-norm-violations-n10-t8-personal-attacks.csv", "macro-norm-violations-n10-t9-personal-attacks.csv"- Comments containing personal attacks.</li> <li>"macro-norm-violations-n15-t3-abusing-and-criticisizing-mods.csv" - Comments abusing and criticisizng moderators.</li> <li>"macro-norm-violations-n15-t9-namecalling-claiming-other-too-sensitive.csv" - Comments with name-calling, or claiming that the other person is too sensitive.</li> </ul> <p>More details about the dataset can be found on arXiv: <a href="https://arxiv.org/abs/1904.03596">https://arxiv.org/abs/1904.03596</a></p>
Structure and dynamics of growing networks of Reddit threads
<p>Data used in the paper "<a href="https://doi.org/10.1007/s41109-024-00654-y" target="_blank" rel="noopener">Structure and dynamics of growing networks of Reddit threads</a>".</p> <p>This dataset is made of 6366 threads collected from the r/AmITheAsshole community on Reddit. The dataset contains a total of 6,372,251 comments. The collected threads constitute the “top” submissions — those having the highest score, measured as the difference between upvotes and downvotes of a post. We downloaded them using PRAW, running 10 different queries across various temporal scopes, and then cleaning the obtained dataset by removing duplicated threads. Please refer to the paper, specifically to <a href="https://appliednetsci.springeropen.com/articles/10.1007/s41109-024-00654-y/tables/3" target="_blank" rel="noopener">Table 3</a>, for more details about the dataset.</p> <p><strong>If you use this data, please cite the following source:</strong> Goglia, D., Vega, D. Structure and dynamics of growing networks of Reddit threads. Appl Netw Sci 9, 48 (2024). https://doi.org/10.1007/s41109-024-00654-y</p>
Cross-mentions between 4chan and Reddit
<p>The included datasets are at the basis of the article "No space for Reddit spacing: Mapping the reflexive relationship between groups on 4chan and Reddit", published in Social Media + Society. They include cross-mentions between 4chan and Reddit, as well as various metrics associated to these cross-references.</p><p>The timeframe ranges from the earliest date available (for Reddit: June 2006; for 4chan/b/: April 2006; for 4chan/pol/: December 2013) and ends in January 2023 (except for the 4chan/b/ dataset, which ends in December 2008).</p><p>The datasets specifically entail the following:</p><h3><strong>1. Cross-mentions from Reddit to 4chan</strong></h3><p><i>reddit-mentions-to-4chan.csv</i> </p><p>I used the Pushshift API's search endpoint to fetch Reddit comments (so no opening posts) with the keyword "4chan" (note: this Pushshift functionality is now deprecated). I also used a rudimentary filter to remove posts by bots, specifically by 1) deleting posts from every account that had "bot" or "auto" in the username and 2) removing all posts by authors with 100 or more contributions and which I manually identified as automated accounts.</p><p>I removed URL-only cross-references, i.e. posts that only mentioned "://boards.4chan.org" or "://boards.4channel.org" without another 4chan-reference/</p><p>This resulted in 2,638,621 "4chan" references across Reddit.</p><h3><strong>2. Cross-mentions from 4chan/pol/ to Reddit</strong></h3><p><i>4chan-pol_mentions-of-reddit.csv</i></p><p>With a complete dataset of /pol/ collected through <a href="https://4cat.nl">4CAT</a>, I queried for "reddit" or the common synonym "plebbit", capital-insensitive, with post- and suffixes allowed (e.g. "Redditor").</p><p>I removed URL-only cross-references, i.e. posts that only mentioned "://reddit.com/", "www.reddit.com/", or "i.reddit.com/". without another Reddit-reference/</p><p>This resulted in 1,640,273 "Reddit" references on /pol/.</p><h3><strong>3. Cross-mentions from 4chan/b/ to Reddit</strong></h3><p><i>4chan-b_mentions-to-reddit.csv</i></p><p>I extracted five million posts from <a href="https://archive.org/details/4chan_threads_archive_10_billion">Jason Scott's 4chan/b/ dump</a>. I then queried for "reddit" or the common synonym "plebbit", capital-insensitive, with post- and suffixes allowed (e.g. "Redditor").</p><p>I removed URL-only cross-references, i.e. posts that only mentioned "://reddit.com/", "www.reddit.com/", or "i.reddit.com/". without another Reddit-reference/</p><p>This resulted in 1,287 "Reddit" references on /b/.</p><p>See <a href="https://oilab.eu/an-overview-of-4chan-b-archives-what-is-left-of-the-internets-cesspool">Hagen (2020)</a> for more information on the 4chan/b/ dataset.</p><h3><strong>4. Cross-mention metrics</strong></h3><p><i>cross-mention-metrics.xlsx</i></p><p>I extracted the following metrics from the datasets above:</p><p><i>4.1 The total number of cross-mentions, absolute and relative, per month</i><br>This simply used the monthly counts from datasets 1 and 2.</p><p><i>4.2 The most mentioned subreddits on /pol/, per year</i><br>Using the regular expression:<i> r\/[a-zA-Z_]</i></p><p><i>4.3 Subreddits that mention 4chan most often, per year</i></p><p><i>4.4 4chan boards mentioned across Reddit, per month</i></p><p><i>4.5 4chan boards mentioned by subreddits</i></p><p>I counted every subreddit- or board-mention <i>per post</i> instead of total occurrences.</p><p>For 4.4 and 4.5, I used the following regular expression to extract 4chan board names:</p><p><i>(\s|^|4chan)\/(a|b|c|d|e|f|g|gif|h|hr|k|m|o|p|t|v|vg|vm|vmg|vr|vrpg|vst|w|wg|i|ic|r9k|s4s|vip|qa|cm|hm|lgbt|y|3|aco|adv|an|bant|biz|cgl|ck|co|diy|fa|fit|gd|hc|his|int|jp|lit|mlp|mu|n|news|out|po|pol|pw|qst|sci|soc|sp|tg|toy|trv|tv|vp|vt|wsg|wsr|x|xs|new)\/(\s|$)</i></p><p>I also omitted 4chan's /r/, /u/, and /s/ boards; despite their small scale, they appeared as false positives due to their unrelated vernacular meaning on Reddit (e.g. /u/ as a username prefix).</p><p>4.5 was also transformed and included as a Gephi network file (<i>subreddit-board-mentions.gephi</i>).</p><p>Lastly, I also included:</p><p>4.6 The <i>total amount of posts</i> on 4chan and Reddit</p><p>This was used to calculate 4.1. It uses <a href="https://files.pushshift.io">Pushshift's database statistics</a> (which as of Nov. 2023 requires a login; see <a href="https://pastebin.com/McS2DSNz">this Pastebin</a> for an alternative) and metrics of total 4chan post counts from <a href="https://4stats.io">4stats.io</a>.</p><p>Each of these metrics has their own corresponding tab in the Excel file.</p><h3><strong>5. Co-words of "4chan" and "reddit" in the cross-mentions</strong></h3><p><i>co-words.xslx</i></p><p>Using datasets 1, 2, and 3, I extracted the top ten words appearing directly next to "4chan" on Reddit, and next to "Reddit" on 4chan, <i>per year</i>.</p><p>I first pre-processed the text, which involved tokenisation, filtering of unwanted text elements like URLs, stop word removal (I whitelisted <i>back</i>), and lemmatisation.</p><p>For the co-word extraction I used a window size of two. I excluded a range of semantically uninteresting words or commonly used hate speech terms prevalent throughout 4chan.</p><h3><strong>6. Annotated cross-mentions between Reddit and 4chan/pol/ in September 2014</strong></h3><p><i>annotations_4chanpol-2014.csv</i><br><i>annotations_reddit-2014-kotakuinaction-anonimised.csv</i><br><i>annotations_reddit-2014-tumblrinaction-anonimised.csv</i></p><p>I extracted cross-mentions from /pol/ to Reddit and from Reddit to 4chan in September 2014 for close-reading and annotation.</p><p>__ </p><p>The author names are removed for all datasets.</p>
Toxic Content Detection in online social networks: a new dataset from Brazilian Reddit Communities
<p>This is new dataset of 2,500 manually annotated examples of comments extracted from the top 10 largest Brazilian subreddits on Reddit. The dataset has been annotated by crowd-sourcing efforts with contributions from the departments of computer science (DCC) and the linguistic group @ UFMG. As part of our contribution to the toxicity automatic detection and moderation of online social networks, we're making the dataset public for research.</p> <h3>Dataset</h3> <p>The dataset contains 2,500 manually annotated comments from the most popular brazilian communities on Reddit. The data sampling proccess was a stratified sampling by the number of generated publications by subreddit and the month of publication. The list of communities collected is presented below. The collected data period ranges from January 2022 to December 2022.</p> <p> </p> <table> <tbody> <tr> <td><strong>Subreddit</strong></td> <td><strong>Posts</strong></td> <td><strong>Comments</strong></td> </tr> <tr> <td>r/brasil</td> <td>110,829 </td> <td>2,136,866</td> </tr> <tr> <td>r/desabafos</td> <td>115,876</td> <td>1,211,643</td> </tr> <tr> <td>r/futebol</td> <td>35,826</td> <td>1,214,412</td> </tr> <tr> <td>r/saopaulo</td> <td>7,308</td> <td>81,969</td> </tr> <tr> <td>r/eu_nvr</td> <td>12,631</td> <td>188,620</td> </tr> <tr> <td>r/botecodoreddit</td> <td>7,059</td> <td>57,298</td> </tr> <tr> <td>r/conversas</td> <td>21,967</td> <td>326,061</td> </tr> <tr> <td>r/investimentos</td> <td>9,756</td> <td>141,823</td> </tr> <tr> <td>r/tiodopave</td> <td>2,371</td> <td>11,584</td> </tr> <tr> <td>r/brasilivre</td> <td>67,301</td> <td>1,219265</td> </tr> <tr> <td>Total</td> <td>390,924</td> <td>6,589,541</td> </tr> </tbody> </table> <p> </p> <h3><strong>Annotation proccess</strong></h3> <p>The annotators were divided into groups of raters and each group was assigned a batch of comments to label. The raters were then asked to label a comment as <strong>Toxic</strong>, <strong>Non-toxic</strong>, <strong>I do not know</strong> and <strong>Missing info</strong>. During the annotation process, the raters were encouraged to assign one of the uncertain labels when they're not sure about the toxicity of a comment or the context is missing. </p> <h3>Available data</h3> <p>The dataset is available as csv file and the label was assigned as a majority vote among the raters. The available data are the original collected comment id and body. The label was created from the original classification from the annotators. No data processing has been done on this version of the dataset. The overall schema of the dataset if presented below.</p> <p>- <strong>id</strong>: The unique identifier of the comment on the Reddit platform<br>- <strong>body</strong>: The original comment text publication<br>- <strong>is_toxic</strong>: The final label of a given comment. The label is <strong>0</strong> for non-toxic comments, <strong>1</strong> for toxic comments and <strong>-1</strong> for comments where the raters disagreed about the toxicity.</p>
Reddit WSB Annotated Dataset 2021
<p>This is a WIP randomized sample of comments from the Wallstreetbets community on Reddit during the GameStop event during the 2021 rise. The sample contains 5000 observations of which 3000 were annotated and agreed upon by two annotators. The next 600 were annotated by two authors but due to time constraints, only 1 author corrected them. The remaining 1400 have only been annotated by one author and have not been compared.</p> <p>The second file is the annotation ruleset used to annotate the dataset, we briefly summarize the rules here:</p> <p>Annotations are broken into two main categories, support (the comment indicates some level of support for either the company GameStop, the stock price, or the narrative of 'us' vs 'them'.). Support can be either Y= Yes, N= No, U= Unsure, I= Informative.</p> <p>The second category, 'intent' indicates the individual has expressed intentions or interest in the stock, or has already purchased the stock during the event period. Intent can be either Y= Yes, N= No, M= Maybe, U= Unsure, or I= Informative.</p>
Reddit photo Critique Dataset
<p>The Reddit Photo Critique Dataset (RPCD) contains tuples of image and photo critiques. RPCD consists of 74K images<br> and 220K comments and is collected from a Reddit community used by hobbyists and professional photographers to improve their photography skills by leveraging constructive community feedback. The proposed dataset differs from previous aesthetics datasets mainly in three aspects, namely (i) the large scale of the dataset and the extension of the comments criticizing different aspects of the image, (ii) it contains mostly UltraHD images, and (iii) it can easily be extended to new data as it is collected through an automatic pipeline.</p> <p> </p> <p>More info about the dataset can be found at the Github repo: https://github.com/mediatechnologycenter/aestheval</p>
Reddit Climate Change Debate Dataset
<p> </p> <p>This dataset contains pairwise interactions between Reddit users debating climate change on general-purpose subreddits. Each account is enriched with information about the stance concerning climate change (e.g., whether one denies or believes climate change exists) estimated by a deep neural model. </p> <p>All data is anonymized, and no personally identifiable information is released.</p> <h2><strong>Dataset</strong></h2> <p>Interactions are stored in four files, each encompassing 3 months of interactions in 2022. <span>Each interaction file <strong><em>climatechange-X.csv</em></strong> contains three columns identifying source, target, and weight, respectively.</span> The resulting graphs are directed.</p> <p>The <strong>climatechange-opinions.csv</strong> file contains three columns identifying node, opinion, and time window. Opinions are stored as floats in [-1,1] such that 1 implies maximum adherence with deniers, -1 implies maximum adherence with supporters, and 0 implies neutrality. Thus, a line like 42,0.99,2 should be read as "node 42 is a climate change denier in the second quarter of 2022".</p> <p>Code to reproduce the experiments in the paper is released in a jupyter notebook.</p> <p>For further information on fields and volumes, please refer to the data paper.</p> <h3><strong>Citation</strong></h3> <p>If used for research purposes, please cite the following paper describing the dataset details:</p> <p><em>TBD</em></p> <h3><strong>Acknowledgements</strong></h3> <p>This work is supported by:</p> <ul> <li>the European Union – Horizon 2020 Program under the scheme “INFRAIA-01-2018-2019 – Integrating Activities for Advanced Communities”,<br>Grant Agreement n.871042, “SoBigData++: European Integrated Infrastructure for Social Mining and Big Data Analytics” (http://www.sobigdata.eu); </li> <li>SoBigData.it which receives funding from the European Union – NextGenerationEU – National Recovery and Resilience Plan (Piano Nazionale di Ripresa e Resilienza, PNRR) – Project: “SoBigData.it – Strengthening the Italian RI for Social Mining and Big Data Analytics” – Prot. IR0000013 – Avviso n. 3264 del 28/12/2021;</li> <li>EU NextGenerationEU programme under the funding schemes PNRR-PE-AI FAIR (Future Artificial Intelligence Research). </li> </ul> <p> </p> <p> </p>
Blockchain and Smart Contracts Topics on Reddit
<p>This dataset includes posts related to Ethereum, Stellar, Algorand, Hyperledger, and smart contracts subreddits. The dataset includes URL, title, author, number of comments, and technology (Ethereum, Stellar, etc.). This dataset is used to spot relevant smart contracts topics Reddit's community discusses. Please consider that this dataset will be extended with more information and smart contracts platform-related subreddits.</p>
Analyzing the Human Recommendation Community 'ifyoulikeblank' on Reddit — Auxiliary materials
<p>This repository contains auxiliary materials for the iConference 2024 paper "“If I like BLANK, what else will I like?”: Analyzing<br>a Human Recommendation Community on Reddit" by Thi Binh Minh Cao and Toine Bogers (= corresponding author)</p> <p>Published in: <em>Proceedings of the 2024 iConference</em>, April 15--26, 2024, Changchun, China.</p> <p>The paper presents the results of an analysis of /r/ifyoulikeblank, a Reddit community dedicated to requesting and providing for recommendations. This repository contains the following auxiliary materials:</p> <ul> <li>The annotated sample of threads from the /r/ifyoulikeblank subreddit (<strong>annotated-dataset.xlsx</strong>). The second sheet in the Excel file explains the contents of the file.</li> <li>The R code for performing the analysis described in the paper (<strong>annotation-analysis.R</strong>)</li> <li>CSV file containing the genres attributed to the artists as crawled from the Spotify API (<strong>artist-spotify-genres.csv</strong>)</li> <li>CSV file containing the popularity scores crawled from the Spotify API for the seed items and Spotify recommendations (<strong>recommendation-popularity.reddit-vs-spotify.csv</strong>)</li> <li>CSV file containing the popularity scores crawled from the Spotify API for the Reddit (<strong>recommendation-popularity.reddit.csv</strong>)</li> <li>The stopwords file used in the textual analysis (<strong>stopwords.csv</strong>)</li> <li>Excel file containing the activity data for the /r/ifyoulikeblank subreddit (<strong>subreddit-stats.xlsx</strong>)</li> </ul>
The Reddit Politosphere: A Large-Scale Text and Network Resource of Online Political Discourse
<p>The Reddit Politosphere is a large-scale resource of online political discourse covering more than 600 political discussion groups over a period of 12 years. Based on the <a href="https://doi.org/10.5281/zenodo.3608135">Pushshift Reddit Dataset</a>, it is to the best of our knowledge the largest and ideologically most comprehensive dataset of its type now available. One key feature of the Reddit Politosphere is that it consists of both text and network data. We also release annotated metadata for subreddits and users.</p> <p>Documentation and scripts for easy data access are provided in an associated <a href="https://github.com/valentinhofmann/politosphere">repository</a> on GitHub.</p>
MineDojo Internet Knowledge Base (Reddit)
<p><strong>Project website:</strong> <a href="https://minedojo.org">minedojo.org</a></p> <p><strong>Paper:</strong> <a href="https://arxiv.org/abs/2206.08853">arxiv.org/abs/2206.08853</a></p> <p><strong>GitHub:</strong> <a href="https://github.com/MineDojo/MineDojo">github.com/MineDojo/MineDojo</a></p> <p><strong>We collect 340K+ Reddit posts along with 6.6M comments under the “<a href="https://www.reddit.com/r/minecraft">r/Minecraft</a>” subreddit.</strong> These posts ask questions on how to solve certain tasks, showcase cool architectures and achievements in image/video snippets, and discuss general tips and tricks for players of all expertise levels. Large language models can be finetuned on our Reddit corpus to internalize Minecraft-specific concepts and develop sophisticated strategies.</p> <p>Check out our paper!</p> <p> </p> <pre><code class="language-markdown">@article{fan2022minedojo, title = {MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge}, author = {Linxi Fan and Guanzhi Wang and Yunfan Jiang and Ajay Mandlekar and Yuncong Yang and Haoyi Zhu and Andrew Tang and De-An Huang and Yuke Zhu and Anima Anandkumar}, year = {2022}, journal = {arXiv preprint arXiv: Arxiv-2206.08853} }</code></pre> <p> </p>
Reddit r/cryptocurrency posts and comments January 2021 - December 2022
<p>The dataset comprises data from two primary sources: Bitcoin market data and user activity data from the r/cryptocurrency subreddit. The time period covered by the dataset spans from January 1, 2021, to December 31, 2022.</p> <h1>Bitcoin Market Data</h1> <p>The Bitcoin market data includes the following metrics:</p> <ul> <li>Open: The opening price of Bitcoin recorded daily.</li> <li>High: The highest price of Bitcoin recorded daily.</li> <li>Low: The lowest price of Bitcoin recorded daily.</li> <li>Close: The closing price of Bitcoin recorded daily.</li> <li>Volume: The daily trading volume of Bitcoin.</li> </ul> <p>These metrics were collected from CoinMarketCap (https://coinmarketcap.com/).</p> <h1>Reddit Activity Data</h1> <p>The Reddit activity data consists of posts and comments from the r/cryptocurrency subreddit, focusing on the most popular content. The data includes:</p> <ul> <li>Posts: 770 of the most popular posts from the specified time period, selected based on upvotes and engagement.</li> <li>Comments: 14,886 comments associated with the collected posts, representing the most popular comments in terms of upvotes and responses.</li> </ul> <p>For each post, the following attributes were recorded:</p> <ul> <li>Title: The post title</li> <li>Score: The upvote score of the post or comment, indicating its popularity.</li> <li>URL: The URL of the post</li> <li>Number of comments: The total number of comments received by each post.</li> <li>Body: The text posted with the post submission</li> <li>Date: The date and time of the post submission</li> </ul> <p>For each comment, the following attributes were recorded:</p> <ul> <li>Date: The date the comment was posted</li> <li>Comment: The content of the comment</li> </ul> <p>The combined dataset aims to provide a comprehensive view of both market and social media activity related to Bitcoin, enabling a detailed analysis of the interplay between market dynamics and user sentiment.</p>
Myers Briggs Personality Tags on Reddit Data
<p>This data was pulled on 11/10/2018 from google big query using the following query:</p> <pre><code class="language-sql">SELECT flair_text.author_flair_text as flair_text, comments.body as body, comments.subreddit as subreddit, comments.author as author FROM ( SELECT author,author_flair_text FROM [fh-bigquery:reddit_comments.all] WHERE author_flair_text != 'null' AND REGEXP_MATCH(author_flair_text,r'([IEie][SNsn][TFtf][JPjp]\W)') GROUP BY author,author_flair_text ) AS flair_text INNER JOIN ( SELECT author_flair_text, body, subreddit, author FROM [fh-bigquery:reddit_comments.all] ) AS comments ON comments.author = flair_text.author </code></pre> <p> </p>
Reddit dataset about the Great Ban moderation intervention
<p>Dataset concerning the Reddit Great Ban moderation intervention (2020-06-29), for a total of 48M comments divided into 5 distinct datasets:</p> <ul> <li>Baseline, before the Great Ban, in all Reddit.</li> <li>Baseline, after the Great Ban, in all Reddit.</li> <li>Group of interest, before the Great Ban, into the analyzed subreddits.</li> <li>Group of interest, before the Great Ban, in all Reddit.</li> <li>Group of interest, after the Great Ban, in all Reddit.</li> </ul> <p>More information are on the README file.</p>
Reddit EU language dataset
<p>This dataset has been created for a personal project related to the recognition of the original language of someone writing in english.</p> <p><strong>Origin</strong></p> <p>The dataset has been crawled from the subreddit r/europe and contains around 1.5 milions posts in it's raw form.</p> <p><strong>Structure</strong></p> <p>This repo contains both the <a href="https://github.com/Tsadoq/Reddit_EU_language_dataset/tree/main/data">raw data</a> and the <a href="https://github.com/Tsadoq/Reddit_EU_language_dataset/tree/main/cleaned_data">cleaned data</a>, the latter, purged of deleted comments and of those that were not linked to the provenience of the writer, contains around 450k datapoints and has the following structure:</p> <ul> <li>body: the text content of the comment</li> <li>country_name: extended name of the country</li> <li>permalink: link to the comment</li> <li>author: username of the creator</li> <li>created_utc: utc creation</li> <li>datetime: date and time of creation</li> <li>alpha2: ISO country alpha2 code</li> <li>alpha3: ISO country alpha3 code</li> <li>numeric: ISO country number</li> <li>apolitical_name: apolitical country name</li> </ul>
Patient-Centric Reddit Cancer Dataset
<p>The Patient-Centric Reddit Cancer Dataset (PCRCD) consists of all posts from r/Cancer during the period 01/01/2014 and 04/30/2020. A total of 23,028 unique posts were extracted. These posts were labelled according to three patient demographic traits: patient versus caregiver, sex and age. Each post was independently labelled by three annotators. PCRCD can help with the understanding of the challenges and concerns being reported by the patients and caregivers dealing with cancer. It also exemplifies the potential and limitations associated with the translation of raw posts from online health forums into a common resource for researchers.</p>
Reddit and StackOverflow dataset (Programming languages)
<p>This data set contains anonymized data collected from Reddit (via the <a href="https://api.pushshift.io/redoc">Pushshift API</a>) and StackOverflow (from <a href="https://www.kaggle.com/datasets/stackoverflow/stackoverflow">Kaggle's dataset</a>).</p> <p>Each folder includes the data split by trimester. The schema of StackOverflow and Reddit-related files follows:</p> <ul> <li>Fields from StackOverflow <ul> <li>question_id</li> <li>answer_id</li> <li>creation_date - answer creation_date</li> <li>score - score of the question/answer</li> <li>tags - all tags flagged for a question</li> <li>answer_count - number of answers for a question</li> <li>start_question - question's time of creation</li> <li>last_activity_date - last update on the question</li> <li>new_id - hashed id of the answerer</li> <li>q_new_id - hashed id of the questioner</li> </ul> </li> <li>Fields from Reddit <ul> <li>comment_id</li> <li>submission_id</li> <li>score - score of the question/submission</li> <li>subreddit</li> <li>created_utc - time of creation (unrelated to last modified comments)</li> <li>new_id - hashed id</li> </ul> </li> </ul> <p>The .txt files represent the structure of the corresponding hypergraphs.</p>
Moral Judgments in Narratives on Reddit: Investigating Moral Sparks via Social Commonsense and Linguistic Signals
<ol> <li>The file 'post_instances.jsonl' contains instances extracted from specific posts. In this file, instances that contain moral sparks are labeled as '1,' while others are labeled differently or as '0.'</li> <li>Each instance is scraped by using PushShift API by searching for an unique id. And each instance contains its comment ids that use ">" to quote excerpts in a post. We removed author names and make it left with ids and contexts. The "label" field is computed by using regular expressions to match predefined r/AmItheAsshole verdict codes.</li> <li>The sup_documents.pdf includes full lists of c-event clusters and parameters of linguistic features used in our paper.</li> <li>The regular expressions used to extract the verdicts are as follows:AUTHOR = (0, 'YTA', [<br> r'\m(?i:YWBTA?)\M',<br> r'\m(?i:YTAH?)\M',<br> r"(?e)(?i:"<br> r"you(?:'re| r| are| were| would be| will be) "<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r"(?:an? |the )?"<br> r"(?:huge |big |giant )?"<br> r"(?:asshole|a-?hole)"<br> r"){e<=1}",<br> r"(?e)(?i:"<br> r"you "<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r"(?: r| are| were)? (?:an? |the )?"<br> r"(?:huge |big |giant )?"<br> r"(?:asshole|a-?hole)"<br> r"){e<=1}"<br> ])<br> OTHER = (1, 'NTA', [<br> r'\m(?i:YWNBTA?)\M',<br> r'\m(?i:Y?NTAH?)\M',<br> r'(?e)(?i:'<br> r"you(?:'re| r| are| were| would| will) "<br> r"(?!both)"<br> r'(?:not| not be) '<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'(?:an? |the )?'<br> r"(?:asshole|a-?hole)"<br> r'){e<=1}',<br> r'(?e)(?i:'<br> r"(he|she)(?:'s|s| s| is| was)"<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'(?:an? |the )?'<br> r"(?:asshole|a-?hole)"<br> r'){e<=1}',<br> r'(?e)(?i:'<br> r"they(?:'re|r| r| are| were)"<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'(?:the )?'<br> r"(?:asshole|a-?hole)"<br> r'){e<=1}'<br> ])<br> EVERYBODY = (2, 'ESH', [<br> r'\m(?i:ESH)\M',<br> r'(?e)(?i:every(?:one|body) sucks here){e<=1}',<br> r'(?e)(?i:you both suck){e<=1}',<br> r'(?e)(?i:'<br> r"you(?:'re| r| are| were) "<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'both (?:the )? (?:assholes?|a-?holes?)){e<=1}',<br> r'(?e)(?i:'<br> r"you both"<br> r"(?:'re| r| are| were)? "<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'(?:the )?'<br> r"(?:assholes?|a-?holes?)"<br> r'){e<=1}',<br> r'(?e)(?i:'<br> r"there(?: r| are| were)(?: any| all) "<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'(?:assholes?|a-?holes?)){e<=1}'<br> <br> ])<br> NOBODY = (3, 'NAH', [<br> r'\m(?i:NAH?H)\M',<br> r'(?e)(?i:no (?:assholes|a-?holes|asshole) here)',<br> r'(?e)(?i:no one is the (?:asshole|a-?hole)){e<=1}',<br> r'(?e)(?i:'<br> r"you both"<br> r"(?:'re| r| are| were)? "<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'not (?:an? |the )?'<br> r"(?:assholes?|a-?holes?)"<br> r'){e<=1}',<br> r'(?e)(?i:'<br> r"you(?:'re| r| are| were) "<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'both not (?:an? |the )? (?:assholes?|a-?holes?)){e<=1}',<br> r'(?e)(?i:'<br> r"you "<br> r"both(?: weren't| aren't)? (?:an? |the )? (?:assholes?|a-?holes?)){e<=1}",<br> r'(?e)(?i:'<br> r"there(?: r| are| were)"<br> r' no '<br> r"(?:(?:kind|sort) of |really |indeed |just |definitely |exactly |absolutely |certainly |obviously )?"<br> r'(?:assholes?|a-?holes?)){e<=1}'<br> ])<br> INFO = (4, 'INFO', [<br> r'\m(?i:INFO)\M',<br> r'(?e)(?i:not enough info){e<=1}',<br> r'(?e)(?i:needs? more info){e<=1}',<br> r"(?e)(?i:more info(?:'s| is)? required){e<=1}"<br> ])</li> </ol> <p> </p>
Reddit Entity Linking
<p>An entity linking dataset created from the social media website, Reddit. The dataset contains 619 posts and 1,243 corresponding comments that were selected and given to human annotators. Three different human annotators were used to annotate each grouping of text. The resulting mentions and entities are included with a breakdown of the inter-annotator agreement between the various mention-entity pairs.</p> <p>The mention-entity pairs collected are broken into groups based on the level of inter-annotator agreement. </p> <p>Gold annotations - all three agree</p> <p>Silver annotations - two out of three annotators agree</p> <p>Bronze annotations - an individual annotator's annotation that the other two did not have</p> <p>In total the dataset contains 1,342 gold annotations, 2,723 silver annotations, and 7,038 bronze annotations.</p> <p>A readme file is provided that describes the structure of the files and the information within each one.</p>
Sample reddit posts 5 Oct 2023
<p>Reddit posts from technology related subreddits from 5 Oct 2023</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.