iDRAMA-rumble-2024: A Dataset of Podcasts from Rumble Spanning 2020 to 2022
<h3>ABSTRACT</h3> <p>---------------<br>Rumble has emerged as a prominent platform hosting controversial figures facing restrictions on YouTube. Despite this, the academic community’s engagement with Rumble has been minimal. To help researchers address this gap, we introduce a comprehensive dataset of about 6.7K podcast videos from August 2020 to December 2022, amounting to over 5.6K hours of content. Besides covering metadata of these podcast videos, we provide speech-to-text transcriptions for future analysis. We also provide speaker diarization information, a collection of ~250K unique representative images from podcast videos, and face embeddings of ~400K extracted faces. With the rise of the influence of podcasts and populist figures, this dataset provides a rich resource for identifying challenges in cyber social threats in a relatively underexplored space.</p> <ul> <li>Rumble platform: <a href="http://rumble.com/">http://rumble.com/</a></li> <li>Link to paper: <a href="https://workshop-proceedings.icwsm.org/abstract.php?id=2024_07">https://workshop-proceedings.icwsm.org/abstract.php?id=2024_07</a></li> <li>License: <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en">CC BY-NC-SA 4.0</a></li> </ul> <h3>Dataset Summary</h3> <p><em><strong>iDRAMA-rumble-2024</strong></em> is a large-scale dataset of 6,735 podcast videos from Rumble, an alternative Youtube-like platform. Using state-of-the-art models, we extract information across three modalities: 1) text, 2) audio, and 3) video. We detail the methodology for extracting information from podcast videos in the paper and release a first-of-its-kind dataset including data from different modalities:</p> <ul> <li><strong>Metadata:</strong> Details about podcast videos, e.g., channel name, video name, video description, and more.</li> <li><strong>Text:</strong> Transcription (i.e., speech-to-text) of podcast videos.</li> <li><strong>Audio:</strong> Speaker diarization information providing speaker detection over time for each video.</li> <li><strong>Video:</strong> Sampled representative video frames from each video, totaling 200K images. We also detect ~400K non-unique faces from these images and release face embeddings.</li> </ul> <h3>Repository links</h3> <ul> <li><strong>Zenodo:</strong> On Zenodo, we provide JSON formatted dataset for all modalities and representative images in compressed files.</li> <li><strong>Github:</strong> The main repository of this dataset, where we provide code snippets to get started with this dataset. <ul> <li>Link here: <a href="https://github.com/idramalab/iDRAMA-rumble-2024">https://github.com/idramalab/iDRAMA-rumble-2024</a></li> </ul> </li> <li><strong>Huggingface:</strong> On Huggingface, we provide a dataset that can be accessed through Huggingface APIs in a `parquet` format. <ul> <li>Link here: <a href="https://hf.co/datasets/iDRAMALab/iDRAMA-rumble-2024">https://hf.co/datasets/iDRAMALab/iDRAMA-rumble-2024</a></li> </ul> </li> </ul> <h3>Dataset Info</h3> <p>The dataset is organized by modalities -- transcripts, representative images, speaker diarization, and face embeddings.</p> <table> <tbody> <tr> <td><strong>Config</strong></td> <td><strong>Data-points</strong></td> </tr> <tr> <td>Podcast videos</td> <td>6,735</td> </tr> <tr> <td>Representative images</td> <td>252,387</td> </tr> <tr> <td>Face embeddings</td> <td>399,333</td> </tr> <tr> <td>Transcripts & Speaker diarization</td> <td>6,735</td> </tr> </tbody> </table> <h3>Zenodo Dataset Files Info</h3> <table> <tbody> <tr> <td> </td> <td>#Files</td> <td>File names</td> </tr> <tr> <td>Metadata</td> <td>1</td> <td><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-metadata.ndjson/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-metadata.ndjson</a></td> </tr> <tr> <td>Speaker diarization</td> <td>1</td> <td><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-speaker-dirization.zip/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-speaker-dirization.zip</a></td> </tr> <tr> <td>Face embeddings</td> <td>1</td> <td><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-face-embeddings.ndjson/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-face-embeddings.ndjson</a></td> </tr> <tr> <td>Representation images</td> <td>5</td> <td> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-repr-images-set1.tar.gz/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-repr-images-set1.tar.gz</a></p> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-repr-images-set2.tar.gz/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-repr-images-set2.tar.gz</a></p> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-repr-images-set3.tar.gz/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-repr-images-set3.tar.gz</a></p> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-repr-images-set4.tar.gz/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-repr-images-set4.tar.gz</a></p> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-repr-images-set5.tar.gz/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-repr-images-set5.tar.gz</a></p> </td> </tr> <tr> <td> <p>Transcription Lite</p> <p>(Minimal information)</p> </td> <td>3</td> <td> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-transcription-lite_part_1.ndjson/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-transcription-lite_part_1.ndjson</a></p> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-transcription-lite_part_2.ndjson/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-transcription-lite_part_2.ndjson</a></p> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-transcription-lite_part_3.ndjson/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-transcription-lite_part_3.ndjson</a></p> </td> </tr> <tr> <td>Transcription</td> <td>3</td> <td> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-transcription_part_1.ndjson/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-transcription_part_1.ndjson</a></p> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-transcription_part_2.ndjson/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-transcription_part_2.ndjson</a></p> <p><a href="../api/records/10515991/draft/files/iDRAMA-rumble-2024-transcription_part_3.ndjson/content" target="_blank" rel="noopener noreferrer">iDRAMA-rumble-2024-transcription_part_3.ndjson</a></p> </td> </tr> </tbody> </table> <h3>Authorship</h3> <p>This dataset is published in the "<em><strong>Workshop Proceedings of the 18th International AAAI Conference on Web and Social Media</strong></em>" hosted in Buffalo, NY, USA.</p> <ul> <li>Academic Organization: <a href="https://idrama.science/people/">iDRAMA Lab</a></li> <li>Authors: Utkucan Balci, Jay Patel, Berkan Balci, Jeremy Blackburn</li> <li>Affiliation: Binghamton University, Middle East Technical University</li> </ul> <h3>Licensing</h3> <p>This dataset is available for free to use under terms of the non-commercial license <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en">CC BY-NC-SA 4.0</a>.</p> <h3>Citation</h3> <blockquote> <p>@article{balci2024idrama,<br> title = {iDRAMA-rumble-2024: A Dataset of Podcasts from Rumble Spanning 2020 to 2022},<br> author = {Balci, Utkucan and Patel, Jay and Balci, Berkan and Blackburn, Jeremy},<br> year = {2024},<br> journal = {Workshop Proceedings of the 18th International AAAI Conference on Web and Social Media}<br>}</p> </blockquote>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 4
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 4