Set of obfuscated spam dataset by using LeetSpeak transformations
<p>The usage of LeetSpeak and other text hiding tricks is often used by spammers in the distribution of unsolicited contents. To evaluate deobfuscation techniques and their impact on spam content classification, we preprocessed several popular public datasets to partially obfuscate the text. The datasets transformed are:</p> <ul> <li>YouTube Spam Collection [2, 3] which is available on <a href="https://www.dt.fee.unicamp.br/~tiago/youtubespamcollection/">https://www.dt.fee.unicamp.br/~tiago/youtubespamcollection/</a>.</li> <li>a subset of YouTube Comments [4, 5] which is available on <a href="http://mlg.ucd.ie/yt/">http://mlg.ucd.ie/yt/</a>.</li> <li>CSDMC2010 which is available on <a href="http://csmining.org/index.php/spam-email-datasets-.html">http://csmining.org/index.php/spam-email-datasets-.html</a>.</li> <li>TREC2007 which is available on <a href="https://plg.uwaterloo.ca/~gvcormac/treccorpus07/">https://plg.uwaterloo.ca/~gvcormac/treccorpus07/</a></li> </ul>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0