Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9
datasets available to search
ShareScore release 0.9.0
Dataset results
9 results for “vandalism”
Wikipedia Multilingual Vandalism Detection Dataset
<p>This dataset accompanies a research paper that introduces a novel system designed to support the Wikipedia community in combating vandalism on the platform. The dataset has been prepared to enhance the accuracy and efficiency of Wikipedia patrolling in multiple languages.</p> <p>The release of this comprehensive dataset aims to encourage further research and development in vandalism detection techniques, fostering a safer and more inclusive environment for the Wikipedia community. Researchers and practitioners can utilize this dataset to train and validate their models for vandalism detection and contribute to improving online platforms' content moderation strategies.</p> <p><strong>Dataset Details:</strong></p> <ul> <li><strong>Number of Languages:</strong> 47</li> <li><strong>Observation period: </strong>6 months training, one week hold-out testing</li> <li><strong>Use Case:</strong> The dataset is primarily intended for training and evaluating vandalism detection systems.</li> <li><strong>Features:</strong> Each record characterizes the corresponding revision of the Wikipedia page, including revision metadata, user details, text inserted, removed, or changed, and corresponding MLMs-based features. </li> <li><strong>Data Filtering and Feature Engineering:</strong> Advanced filtering and feature engineering techniques were applied to ensure the dataset's quality and relevance for effectively training the vandalism detection system.</li> <li><strong>Files: </strong>Training and hold-out testing datasets of anonymous and all users. </li> </ul> <p> </p> <p><strong>Related paper citation:</strong></p> <pre><code>@inproceedings{10.1145/3580305.3599823, author = {Trokhymovych, Mykola and Aslam, Muniza and Chou, Ai-Jou and Baeza-Yates, Ricardo and Saez-Trumper, Diego}, title = {Fair Multilingual Vandalism Detection System for Wikipedia}, year = {2023}, isbn = {9798400701030}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3580305.3599823}, doi = {10.1145/3580305.3599823}, abstract = {This paper presents a novel design of the system aimed at supporting the Wikipedia community in addressing vandalism on the platform. To achieve this, we collected a massive dataset of 47 languages, and applied advanced filtering and feature engineering techniques, including multilingual masked language modeling to build the training dataset from human-generated data. The performance of the system was evaluated through comparison with the one used in production in Wikipedia, known as ORES. Our research results in a significant increase in the number of languages covered, making Wikipedia patrolling more efficient to a wider range of communities. Furthermore, our model outperforms ORES, ensuring that the results provided are not only more accurate but also less biased against certain groups of contributors.}, booktitle = {Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining}, pages = {4981–4990}, numpages = {10}, location = {Long Beach, CA, USA}, series = {KDD '23} }</code></pre> <p> </p>
Wikidata Vandalism Corpus 2015 (WDVC-15)
<p>The Wikidata vandalism corpus 2015 (WDVC-15) is a corpus for the evaluation of automatic vandalism detectors for Wikidata. For research purposes the corpus can be used free of charge.</p>
PAN Wikipedia Vandalism Corpus 2010 (PAN-WVC-10)
<p>The PAN Wikipedia Vvandalism Ccorpus 2010 (PAN-WVC-10) is a corpus for the evaluation of automatic vandalism detectors for Wikipedia. For research purposes the corpus can be used free of charge.</p> <p>This corpus is supplemented by the <a href="https://doi.org/10.5281/zenodo.3342157">PAN-WVC-11</a>, which features additional edits in English, Spanish and German. Both corpora should be used to get more representative results.</p> <p>As part of our research on automatic vandalism detection we have compiled a corpus of vandalism cases found in Wikipedia. The corpus compiles 32452 edits on 28468 Wikipedia articles, among which 2391 vandalism edits have been identified. To annotate the corpus we have used Amazon's Mechanical Turk; 753 workers have been recruited who cast more than 150000 votes on the edits, so that each edit was reviewed by at least 3 annotators. The achieved level of agreement was analyzed in order to label an edit as "regular" or "vandalism."</p>
Webis Wikipedia Vandalism Corpus (Webis-WVC-07)
<p>This corpus is outdated. Please use its successors <a href="https://doi.org/10.5281/zenodo.3341488">PAN-WVC-10</a> and <a href="https://doi.org/10.5281/zenodo.3342157">PAN-WVC-11</a>.</p> <p>The Webis Wikipedia Vandalism Corpus (Webis-WVC-07) is a corpus for the evaluation of automatic vandalism detection algorithms for Wikipedia. For research purposes the corpus can be used free of charge.</p> <p>The corpus is the first standardized test collection for the comparison of vandalism detection algorithms. It comprises 940 edits from which 301 are marked as vandalism by human evaluators.</p>
PAN Wikipedia Vandalism Corpus 2011 (PAN-WVC-11)
<p>The PAN Wikipedia Vandalism Corpus 2011 (PAN-WVC-11) is a corpus for the evaluation of automatic vandalism detectors for Wikipedia. For research purposes the corpus can be used free of charge.</p> <p>This corpus supplements the <a href="https://doi.org/10.5281/zenodo.3341488">PAN-WVC-10</a>, which features only English edits. Both corpora should be used to get more representative results.</p> <p>The corpus compiles 29949 edits on 24351 Wikipedia articles, among which 2813 vandalism edits have been identified. The corpus features 9985 English edits, 9990 German edits, and 9974 Spanish edits. To annotate the corpus we have used Amazon's Mechanical Turk; each edit was presented to a number of annotators who were asked to decide whether it is vandalism or regular, and the agreement of the annotators was analyzed in order to label an edit.</p>
Petroglyph Park Vandalism
In April 2021 someone used red spraypaint to graffiti these rocks near prehistoric rock imagery. These petroglyphs are approximately 2,000 years old and created by Late Archaic and Fremont people who lived in the area. This site is still loved by the community, today that community is the nearby town of Eagle Mountain, Utah. The Utah State Historic Preservation Office is working with the City of Eagle Mountain, Utah Cultural Site Stewards, and Logan Simpson design to study these petroglyphs and remove the spraypaint graffiti. Source: Objaverse 1.0 / Sketchfab
Vandalized 3D scan
[Original HD object](https://skfb.ly/KsHn) (292k verts) from MGD Films.  This is my first test to develop an efficient workflow to "vandalize" existing scans, by baking new textures onto the already baked model (~500 verts). The final idea being to seamlessly vandalize famous buildings with famous paintings. This particuliar object was remeshed with mmgs followed by heavy decimate modifiers in blender. I did not care at all about topology. Baking from diffuse and normals only, using a plane with the Banksy texture as an emission shader. Source: Objaverse 1.0 / Sketchfab
Blockage, vandalism, and harassment activities for the cause of climate change mitigation: data deposit
<p>The paper describes a dataset of metadata of 89 blockage, vandalism, and harassment events happening in recent years. The dataset comprises three main categories: 1) Events, 2) Activists, and 3) Consequences. For researchers interested in environmental activism, climate change, and sustainability, the dataset is helpful in studying the effectiveness and appropriateness of strategies to raise public awareness and support. For researchers in the field of security studies and green criminology, the dataset offers resources to study features and impacts of blockage, vandalism, and harassment events. The Bayesian Mindsponge Framework (BMF) analytics was employed to validate the dataset. Consequently, the estimated result aligns with the Mindsponge Theory’s theoretical reasoning.</p>
Wikidata Vandalism Corpus 2016 (WDVC-16)
<p>The Wikidata vandalism corpus 2016 (WDVC-16) is a corpus for the evaluation of automatic vandalism detectors for Wikidata. It was employed as part of the WSDM Cup 2017. For research purposes the corpus can be used free of charge.</p> <p>When using the data, please make sure to refer to it as follows:</p> <pre><code>@inproceedings{heindorf2017overview, author = {Stefan Heindorf and Martin Potthast and Gregor Engels and Benno Stein}, title = {Overview of the Wikidata Vandalism Detection Task at {WSDM} Cup 2017}, booktitle = {{WSDM Cup 2017 Notebook Papers}}, url = {https://arxiv.org/abs/1712.05956}, year = {2017} }</code></pre> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.