Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
21
datasets available to search
ShareScore release 0.7.1
Dataset results
21 results for “plagiarism”
Software Plagiarism Detection on Intermediate Representation Data Set
<p>This data set contains the necessary files used for the bachelor's thesis Software Plagiarism Detection on Intermediate Representation.</p> <p>This includes the data sets for the tests, the implemented code and scripts for the evaluation as well as referenced work.</p> <p> </p>
Identifying Machine-Paraphrased Plagiarism
<p>README.txt</p> <p>Title: <em>Identifying Machine-Paraphrased Plagiarism</em><br> Authors: Jan Philip Wahle, Terry Ruas, Tomas Foltynek, Norman Meuschke, and Bela Gipp<br> contact email: wahle@gipplab.org; ruas@gipplab.org;<br> Venue: iConference<br> Year: 2022<br> ================================================================<br> <strong>Dataset Description:</strong></p> <p><em><strong>Training:</strong></em><br> 200,767 paragraphs (98,282 original, 102,485paraphrased) extracted from 8,024 Wikipedia (English) articles (4,012 original, 4,012 paraphrased using the SpinBot API).</p> <p><em><strong>Testing:</strong></em><br> SpinBot: <br> arXiv - Original - 20,966; Spun - 20,867<br> Theses - Original - 5,226; Spun - 3,463<br> Wikipedia - Original - 39,241; Spun - 40,729<br> <br> SpinnerChief-4W: <br> arXiv - Original - 20,966; Spun - 21,671<br> Theses - Original - 2,379; Spun - 2,941<br> Wikipedia - Original - 39,241; Spun - 39,618<br> <br> SpinnerChief-2W: <br> arXiv - Original - 20,966; Spun - 21,719<br> Theses - Original - 2,379; Spun - 2,941<br> Wikipedia - Original - 39,241; Spun - 39,697</p> <p>================================================================<br> Dataset Structure:</p> <p><strong>[human_evaluation]</strong> folder: human evaluation to identify human-generated text and machine-paraphrased text. It contains the files (original and spun) as for the answer-key for the survey performed with human subjects (all data is anonymous for privacy reasons).</p> <p>NNNNN.txt - whole document from which an extract was taken for human evaluation<br> key.txt.zip - information about each case (ORIG/SPUN)<br> results.xlsx - raw results downloaded from the survey tool (the extracts which humans judged are in the first line)<br> results-corrected.xlsx - at the very beginning, there was a mistake in one question (wrong extract). These results were excluded.</p> <p><br> <strong>[automated_evaluation]: </strong>contains all files used for the automated evaluation considering [spinbot] (https://spinbot.com/API) and [spinnerchief] (http://developer.spinnerchief.com/API_Document.aspx).</p> <ul> <li>Each paraphrase tool folder contains:</li> <li><strong>[corpus] </strong>and<strong> [vectors]</strong> sub-folders.</li> <li>For [spinnerchief], two variations are included, with 4-word-chaging ratio (default) and 2-word-chaging ratio. </li> </ul> <p><strong>[vectors] sub-folder</strong> contains the average of all word vectors for each paragraph. Each line has the number of dimensions of the word embeddings technique used (see paper for more details) followed by its respective class (i.e., label mg or og). Each file belongs to one class, either "mg" or "og". The values are comma-separated (.csv). The extension is .arff can be read as a normal .txt file.</p> <ul> <li>The word embedding technique used is described in the file name with the following structure: <technique>-<type>-mean-<data>.arff . Where</li> </ul> <p><em><technique></em> - d2v - doc2vec<br> google - word2vec<br> fasttextnw - fastText without subwording<br> fasttextsw - fastText with subwording<br> glove - Glove<br> <br> Details for each technique used can be found in the paper.<br> <br> <em><type> - </em> arxivp - arXiv paragraph split<br> thesisp - Theses paragraph split<br> wikip - Wikipedia paragraph split (wikipedia_paragraph_vector_train are the vectors used for training. It follows the same wikip structure) </p> <p>Details for each technique used can be found in the paper referenced at the start of this README file.</p> <p><strong>[corpus] sub-folder:</strong> contains de raw text (No pre-processing) used for train and test at a paragraph level.</p> <ul> <li>The Spun paragraphs used for <strong>training</strong> are only generated using the <strong>SpinBot tool</strong>. For test both SpinBot and SpinnerChief are used. </li> <li>The paragraph split is generated by selecting paragraphs from the original documents with 3 or more sentences. Each folder is divided in mg (i.e., machine-generated through SpinBot and SpinnerChief) and og (i.e., original-generated file). the document split is not avaiable since our experiments only use the paragraph level.</li> <li>Machine Learning models: SVM, Naive Bayes, and Logistic Regression. The grid search for hyperparameter adjustments for the machine learning classifiers is described in the paper.</li> </ul> <p>@incollection{WahleRFM22,<br> title = {Identifying {{Machine-Paraphrased Plagiarism}}},<br> booktitle = {Information for a {{Better World}}: {{Shaping}} the {{Global Future}}},<br> author = {Wahle, Jan Philip and Ruas, Terry and Folt{\’y}nek, Tom{\’a}{\v s} and Meuschke, Norman and Gipp, Bela},<br> editor = {Smits, Malte},<br> year = {2022},<br> volume = {13192},<br> pages = {393--413},<br> publisher = {{Springer International Publishing}},<br> address = {{Cham}},<br> doi = {10.1007/978-3-030-96957-8_34},<br> isbn = {978-3-030-96956-1 978-3-030-96957-8},<br> }</p> <p> </p> <p>For our previous publication using only SpinBot and Wikipedia articles for document and paragraph split, please see the following publication. The dataset used is hosted in <a href="https://deepblue.lib.umich.edu/data/concern/data_sets/2801pg45f?locale=en">DeepBlue</a></p> <p><br> </p>
PAN Arabic Intrinsic Plagiarism Detection Shared Task Corpus
<p>Evaluation corpus for ARAbic INtrinsic plagiarism detection (InAra Corpus) </p> <p> </p> <p>This corpus has been used in AraPlagDet 2015 shared task </p> <p>More details could be found in : <a href="http://araplagdet.misc-lab.org">https://araplagdet.misc-lab.org/</a> or <a href="http://pan.webis.de/fire15/pan15-web/index.html">https://pan.webis.de/fire15/pan15-web/index.html</a> </p> <p> </p> <p><strong>I. SYNOPSIS </strong></p> <p>InAra corpus comprises 2048 documents; 80% of them contain passages borrowed from other documents to simulate documents that contain plagiarized fragments. The corpus involves 2 parts: Training and test.</p> <p> </p> <p><strong>II. DESCRIPTION </strong></p> <p>Each part of the corpus (training and test) consists mainly of 2 datasets: textual files and XML files. The textual files represent the suspicious documents i.e., the documents that contain artificial plagiarism; and the XML files are the plagiarism annotation i.e. they provide for each plagiarized passage its starting offset in the suspicious document and its length (offset and length are both expressed in characters). A suspicious document file and its plagiarism annotation file share the same name.</p> <p> </p> <p><strong>III. PURPOSE </strong></p> <p>The purpose of InAra corpus is to evaluate automatic plagiarism detection methods, notably methods of the intrinsic approach. This approach consists in uncovering the plagiarized passages on the basis of the writing style inconsistency in a given suspicious document. As opposed to the external approach, the intrinsic approach does not necessitate any comparison of the suspicious document against the potential sources of plagiarism. Hence, InAra corpus is not appropriate for the evaluation of the external plagiarism detection because the source of plagiarism are not provided.</p> <p>It should be noted that some documents in InAra corpus contain religious quotations (e.g., Quran and Hadith). These quotations have a peculiar writing style and then a simple intrinsic plagiarism detection software can consider them as plagiarism. However, quotations are not plagiarism, and they are not annotated in the XML files in InAra. Hence, it is an important feature for the plagiarism detection systems evaluated on InAra to not consider religious quotations as plagiarism cases unless they appear as part of a larger plagiarism case.</p> <p> </p> <p><strong>IV. BUILDING METHODS </strong></p> <p>The documents that compose InAra corpus do not contain actual plagiarism cases. They are rather artificial suspicious documents in which plagiarism was created automatically by a software that takes fragments of text from one or more sources documents and inserts them in another one according to a set of parameters, namely the percentage of plagiarism and the plagiarized passages lengths. This building method is the same used to construct PAN 2009-2011 corpora of plagiarism detection (see <a href="http://pan.webis.de">http://pan.webis.de</a> for more information on PAN competition and its corpora). </p> <p> </p> <p><strong>V. LANGUAGE AND ENCODING </strong></p> <p>All the textual documents of this corpus are written in Arabic language and encoded in UTF-8 without BOM.</p> <p> </p> <p><strong>VI. SOURCES OF TEXTS </strong></p> <p>Texts used to build this corpus, either suspicious documents or the inserted passages, are taken mainly from the open library Arabic Wikisource (http://ar.wikisource.org), one of Wikimedia Foundation projects. A few numbers of documents were taken from other websites, namely: </p> <ul> <li>Create your own country blog: http://diycountry.blogspot.com </li> <li>Corpus of Classical Arabic (KSUCCA): http://ksucorpus.ksu.edu.sa </li> <li>Islamic book web site: http://www.islamicbook.ws </li> </ul> <p> </p> <p><strong>VII. COPYRIGHT AND AVAILABILITY </strong></p> <p>We were very careful to build the corpus with copyright-free texts only, to be able to make it publicly available without any sort of problems with texts owners. </p> <p> </p> <p><strong>VIII. HOW TO CITE THE CORPUS ?</strong></p> <p>If you publish a paper about your experimentations using InAra corpus, please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., & Chikhi, S.: Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection. In P. Majumder, M. Mitra, M. Agrawal, & P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India, December 4-6, CEUR proceedings vol. 1587 (pp. 111–122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on InAra with the methods of AraPlagDet competition described in the paper above.</p> <p>Additional information on the corpus building are in the papers:</p> <ul> <li>Bensalem, I., Rosso, P., Chikhi, S.: A New Corpus for the Evaluation of Arabic Intrinsic Plagiarism Detection. In: Forner, P., Müller, H., Paredes, R., Rosso, P., and Stein, B. (eds.) CLEF 2013, LNCS, vol. 8138. pp. 53–58. Springer, Heidelberg (2013).</li> <li>Bensalem, I., Rosso, P., Chikhi, S.: Building Arabic Corpora from Wikisource. 10th ACS/IEEE International Conference on Computer Systems and Applications (AICCSA’13),May 27-30 Fes/Ifran, Morocco (2013).IEEE. </li> </ul> <p> </p> <p>You may wish to compare the results of your experiments with the result of the following papers that used InAra corpus:</p> <ul> <li>Bensalem I, Rosso P, Chikhi S (2019) On the use of character n-grams as the only intrinsic evidence of plagiarism. Language Resources and Evaluation 53:363–396. doi: 10.1007/s10579-019-09444-w</li> <li>Mahgoub AY, Magooda A, Rashwan M, et al (2015) RDI System for Intrinsic Plagiarism Detection (RDI_RID), Working Notes for PAN-AraPlagDet at FIRE 2015. In: Majumder P, Mitra M, Agrawal M, Mehta P (eds) Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India, December 4-6, CEUR proceedings vol. 1587. CEUR-WS.org, pp 129–130</li> </ul> <p> </p> <p><strong>IX. WARNING </strong></p> <p>It should be noted that the Arabic texts may contain quotations from the Quran and the Hadith; and due to the fact that text insertion is automatic and in random positions, it is possible that the plagiarized text is inserted unintentionally between Quranic verses or sentences of a Hadith cited in a document. Hence, the inserted passages may alter the meaning of the original text. For these reasons, this corpus must not be used outside the purpose for which it was built. Examples of the inappropriate use include using the corpus documents as a source of knowledge or distributing them without mentioning that they contain borrowed texts. If you are not interested in plagiarism detection and you are retaining the corpus because it contains books you want to read, then this corpus is not the right source. Please, you should refer to the </p> <p>sources mentioned in Section VI where you can find the original content of the books you are looking for. We emphasize that we are not responsible for the results of any use of this corpus other than the evaluation of the intrinsic plagiarism detection methods. </p> <p> </p> <p><strong>X. CONTACT US</strong></p> <p>We will be happy to hear from you about your experience in using InAra corpus. Please do not hesitate to contact us with the following email address: bens.imene@gmail.com</p> <p> </p> <p>Imene Bensalem¹, Paolo Rosso², Salim Chikhi¹</p> <p>¹MISC Lab. Constantine 2 university, Algeria</p> <p>²PRHLT, Universitat Politècnica de València, Spain </p>
PAN Arabic External Plagiarism Detection Shared Task Corpus
<p>Evaluation Corpus for ARAbic EXternal plagiarism detection (ExAra Corpus) </p> <p> </p> <p>This corpus has been used in AraPlagDet 2015 shared task </p> <p>More details could be found in : <a href="http://araplagdet.misc-lab.org">https://araplagdet.misc-lab.org/</a> or <a href="http://pan.webis.de/fire15/pan15-web/index.html">https://pan.webis.de/fire15/pan15-web/index.html</a></p> <p>If you publish a paper about your experimentations using ExAra corpus, please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., & Chikhi, S.: Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection. In P. Majumder, M. Mitra, M. Agrawal, & P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India, December 4-6, CEUR proceedings vol. 1587 (pp. 111–122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on ExAra with the methods of AraPlagDet competition described in the paper above.</p> <p> </p> <p><strong>I. SYNOPSIS </strong></p> <p>ExAra corpus comprises 2345 documents; almost half of them (suspecious doucuments) contain passages borrowed from the other half (source docucments) to simulate documents that contain plagiarized fragments. The corpus involves 2 parts: Training and test.</p> <p> </p> <p><strong>II. DESCRIPTION </strong></p> <p>Each part of the corpus (training and test) consists mainly of 3 datasets: 2 sets of textual files and 1 set of XML files. The 2 sets of the textual files are the suspicious documents (i.e. the documents that contain artificial plagiarism) and the source documents (i.e., the documents from which the suspicious passages have been plagiarised). The 3rd set of documents contains XML files, which are the plagiarism annotation, i.e., they provide for each plagiarized passage its starting offset and its length in both the suspicious and source documents (offset and length were both expressed in characters). A suspicious document file (.txt) and its plagiarism annotation file (.xml) share the same name.</p> <p> </p> <p><strong>III. PURPOSE </strong></p> <p>The purpose of ExAra corpus is to evaluate automatic plagiarism detection methods, notably methods of the External approach. This approach consists in uncovering the plagiarized passages on the basis of their similarity with passages in the source documents.</p> <p>It should be noted that some suspicious documents in ExAra corpus contain religious quotations (e.g., Quran and Hadith) and common phrases. Some of them appear also in some source documents, and hence a simple plagiarism detection software can consider them as plagiarism. However, quotations and common phrases are legitimate text reuse cases and are not annotated in the XML files in ExAra. Therefore, it is an important feature for the plagiarism detection systems evaluated on ExAra to not consider religious quotations and common phrases as plagiarism cases unless they appear as part of a larger plagiarism case.</p> <p> </p> <p><strong>IV. BUILDING METHODS </strong></p> <p>The documents that compose ExAra corpus do not contain actual plagiarism cases, they are rather artificial suspicious documents in which plagiarism was created automatically by a software that takes fragments of text from one or more sources documents and inserts them in another one according to a set of parameters, namely the percentage of plagiarism and the lengths of the plagiarized passages. Some of the plagiarised fragments are obfuscated manually or automatically before inserting them in the suspicious documents.</p> <p>This building method is the same used to construct PAN 2009-2011 corpora of plagiarism detection (see http://pan.webis.de for more information on PAN competition and its corpora). </p> <p> </p> <p><strong>V. LANGUAGE AND ENCODING </strong></p> <p>All the textual documents of this corpus are written in Arabic language and encoded in UTF-8 without BOM.</p> <p> </p> <p><strong>VI. HOW TO CITE THE CORPUS ?</strong></p> <p>If you publish a paper about your experimentations using ExAra corpus, please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., & Chikhi, S.: Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection. In P. Majumder, M. Mitra, M. Agrawal, & P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India, December 4-6, CEUR proceedings vol. 1587 (pp. 111–122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on ExAra with the methods of AraPlagDet competition described in the paper above.</p> <p> </p> <p><strong>VII. WARNING </strong></p> <p>It should be noted that the Arabic texts may contain quotations from the Quran and the Hadith; and due to the fact that text insertion is automatic and in random positions, it is possible that the plagiarized text is inserted unintentionally between Quranic verses or sentences of a Hadith cited in a document. Hence, the inserted passages may alter the meaning of the original text. For these reasons, this corpus must not be used outside the purpose for which it was built. Examples of the inappropriate use include using the corpus documents as a source of knowledge or distributing them without mentioning that they contain borrowed texts. If you are not interested in plagiarism detection, and you are retaining the corpus because it contains articles you want to read, then this corpus is not the right source. Please, you should refer to the sources mentioned in (Bensalem et al. 2015) (i.e.,the paper above) where you can find the original content of the articles you are looking for.</p> <p>We emphasize that we are not responsible for the results of any use of this corpus other than the evaluation of the external plagiarism detection methods. </p> <p> </p> <p><strong>VIII. CONTACT US</strong></p> <p>We will be happy to hear from you about your experience in using ExAra corpus. Please do not hesitate to contact us with the following email address: bens.imene@gmail.com</p> <p> </p> <p>Imene Bensalem¹, Imene Boukhalfa¹, Paolo Rosso², Salim Chikhi¹</p> <p>¹MISC Lab. Constantine 2 university, Algeria</p> <p>²PRHLT, Universitat Politècnica de València, Spain </p>
PAN Plagiarism Corpus 2010 (PAN-PC-10)
<p>This corpus is outdated. Please use its successor PAN-PC-11: https://doi.org/10.5281/zenodo.3250095</p> <p>The PAN plagiarism corpus 2010 (PAN-PC-10) is a corpus for the evaluation of automatic plagiarism detection algorithms. For research purposes the corpus can be used free of charge.</p> <p>The PAN-PC-10 contains documents in which artificial plagiarism has been inserted automatically as well as documents in which simulated plagiarism has been inserted manually. The former have been constructed using a so-called random plagiarist, a computer program which constructs plagiarism according to a number of parameters, while the latter have been obtained with crowdsourcing via Amazon's Mechanical Turk.</p>
PAN Plagiarism Corpus 2009 (PAN-PC-09)
<p>This corpus is outdated. Please use its successor PAN-PC-11: https://doi.org/10.5281/zenodo.3250095</p> <p>The PAN plagiarism corpus 2009 (PAN-PC-09) is a corpus for the evaluation of automatic plagiarism detection algorithms. For research purposes the corpus can be used free of charge.</p> <p>The PAN-PC-09 contains documents in which artificial plagiarism has been inserted automatically. The plagiarism cases have been constructed using a so-called random plagiarist, a computer program which constructs plagiarism according to a number of random variables. The variables include the percentage of plagiarism in the whole corpus, the percentage of plagiarism per document, the length of a single plagiarized section, and the degree of obfuscation per plagiarized section.</p>
PAN Plagiarism Corpus 2011 (PAN-PC-11)
<p>The PAN plagiarism corpus 2011 (PAN-PC-11) is a corpus for the evaluation of automatic plagiarism detection algorithms. For research purposes the corpus can be used free of charge.</p> <p>The PAN-PC-11 contains documents in which plagiarism has been inserted automatically as well as documents in which plagiarism has been inserted manually. The former have been constructed using a so-called random plagiarist, a computer program which constructs plagiarism according to a number of parameters, while the latter have been obtained with crowdsourcing via Amazon's Mechanical Turk.</p>
Cite & reference to avoid plagiarism
<p>This video explains the importance of citing and referencing, and shows what name/date and footnote citation style formats look like.</p>
Detecting Cross-Language Plagiarism using Open Knowledge Graphs
<p>Corresponding authors: <a href="mailto:meuschke@uni-wuppertal.de?subject=Inquiry%20about%20CL-OSA%20dataset">Norman Meuschke</a>, <a href="mailto:ruas@uni-wuppertal.de?subject=Inquiry%20about%20CL-OSA%20dataset">Terry Ruas</a><br> Venue: 2nd Workshop on Extraction and Evaluation of Knowledge Entities from Scientific Documents (EEKE2021)<br> at the ACM/IEEE Joint Conference on Digital Libraries 2021 (JCDL2021)</p> <p>==========================================================================</p> <p><strong>Source code: <a href="https://github.com/ag-gipp/cl-osa">https://github.com/ag-gipp/cl-osa</a> </strong></p> <p>==========================================================================</p> <p><strong>Dataset Details</strong></p> <p><em><a href="https://jipsti.jst.go.jp/aspec/">ASPEC</a></em>. The Asian Scientific Paper Excerpt Corpus comprises excepts of scientific papers in Japanese that have been manually translated to English and Chinese. We use both subsets of the ASPEC corpus. </p> <ul> <li><em>ASPEC-JC</em><strong> </strong>contains abstracts and paragraphs from the main text of research papers that were translated manually from Japanese to Chinese.</li> <li><em>ASPEC-JE</em> contains abstracts of approx. two million research papers that were translated manually from Japanese to English. </li> </ul> <p><em><a href="https://ec.europa.eu/jrc/en/language-technologies/jrc-acquis">JRC-Acquis</a></em>. The corpus consists of legislative texts in 22 languages, which the European Union's Joint Research Centre (JRC) selected from the cumulative body of EU laws (the so called Acquis communautaire). We sampled our test cases from the 10,000 document pairs in the English-French subset of the corpus.</p> <p><em><a href="https://www.statmt.org/europarl/">Europarl</a></em>. The corpus contains transcripts of European Parliament proceedings in 21 European languages. We exclusively sampled test cases from the 9,443 document pairs in the English-French subset of the corpus.</p> <p><em><a href="https://zenodo.org/record/3250095#.YOr25jqxVH4">PAN-PC-11</a></em>. The corpus contains instances of simulated monolingual and cross-language plagiarism that were used for evaluating plagiarism detection methods as part of the workshop series <a href="https://pan.webis.de/">Plagiarism Analysis, Authorship Identification, and Near-Duplicate Detection (PAN)</a>. Most of the 26,939 documents in the corpus were created by extracting text from openly available books. The documents are partially interspersed with instances of simulated plagiarism that were created and obfuscated automatically or by crowdsourced workers. We exclusively sampled test cases from the 2,921 Spanish-English aligned document pairs in the corpus, for which simulated plagiarism instances were either machine-generated or created manually by crowdsourced workers.</p> <ul> </ul> <p>==========================================================================</p> <p><strong>File Structure</strong></p> <p><strong>[corpus_documents] folder</strong>: Corpora of translation-aligned documents used in our experiments composed of: </p> <ul> <li>aspec: Japanese and English</li> <li>aspecx: Japanese and Chinese </li> <li>jrc: English and French </li> <li>europarl: English and French</li> <li>pan: English and Spanish</li> </ul> <p>Each sub-corpus consists of 4,000 translation-aligned files (2,000 per language); the entire corpus has thus 20,000 files.<br> Each set of translation-aligned documents was randomly selected from the original datasets (details in the paper).<br> The Japanese files in aspec and aspecx do not necessarily overlap even though they are from the same dataset.</p> <p><br> <strong>[vectors_documents] folder</strong>: Average vector representation of the documents in the datasets from two pre-trained models:</p> <ul> <li>Universal Sentence Encoder - Multilingual (USE-ML)</li> <li>ConceptNet Numberbatch</li> </ul> <p> </p> <p><strong>Naming convention</strong>: <model>_<dataset>_<language>;</p> <ul> <li>Example: cn_jrc_es: <ul> <li>model: ConceptNet Numberbatch</li> <li>corpus: JRC-Acquis</li> <li>language: Spanish</li> </ul> </li> </ul> <ul> <li>Labels: <ul> <li><model>:<br> cn - ConceptNet Numberbatch<br> um - USE-ML<br> </li> <li><dataset><br> aspec - ASPEC (Asian Scientific Paper Excerpt Corpus) - English and Japanese<br> aspecx - ASPEC (Asian Scientific Paper Excerpt Corpus) - Japanese and Chinese<br> jrc - JRC-Acquis<br> europarl - Europarl<br> pan - PAN-PC-11<br> </li> <li><language><br> en - English<br> es - Spanish<br> fr - French<br> ja - Japanese<br> zh - Chinese</li> </ul> </li> </ul>
Is plagiarism on the rise in Ethiopian politics? The case of three Master's theses at Addis Ababa University
<p>After an inquiry for plagiarism, the University of Düsseldorf in Germany <a href="https://www.theguardian.com/world/2013/feb/09/german-education-minister-quits-phd-plagiarism">revoked Annette Schavan's doctorate degree in 2013</a>. Schavan (CDU) served as Germany's Federal Minister of Education from 2005 to 2013, and resigned after her PhD was revoked. Karl-Theodor zu Guttenberg (CDU), Germany's defense minister, <a href="https://www.theguardian.com/world/2011/mar/01/german-defence-minister-resigns-plagiarism">resigned in 2011</a> when plagiarism was discovered in his PhD dissertation. As a result, the VroniPlag Wiki was established, where the level of plagiarism in German doctorate theses is investigated and documented via crowdsourcing. As a consequence, dozens of politicians were deprived of their degrees, voluntarily abandoned them, or fled politics. <a href="https://www.faz.net/aktuell/karriere-hochschule/hoersaal/franziska-giffey-verzicht-auf-doktortitel-ist-nicht-moeglich-18746134.html">Franziska Giffey (SPD) and Martin Huber (CSU)</a> are the most recent additions in 2023.</p> <p>The prestige of having an advanced degree is quite high in Ethiopia, as it is in Germany, and as a result, the temptation for shortcuts may be quite strong. We documented three instances of possible MSc thesis plagiarism at Addis Abeba University (AAU): Abraham Belay (2007), Takele Uma (2014), and Dagmawit Moges (2019), all of whom are (or were until recently) ministers in the Ethiopian government.</p> <p>The MSc theses were formally analysed for similarity with earlier works using <a href="https://www.turnitin.com/">Turnitin anti-plagiarism software</a>. The automated output was then further analysed. Similarities with works produced after or concurrently with the publishing date of the MSc theses were omitted from the similarity report; this pertains to a few instances of theses and articles in predatory journals that plagiarized passages from those theses. Also short similarities (eight words or less) were removed from the report. After that, the similarity scores remain high for the three MSc theses, with numerous fully copied paragraphs and sections.</p> <p><strong>Abraham Belay</strong> published an MSc thesis on “<a href="http://etd.aau.edu.et/bitstream/handle/123456789/6598/Abraham%20belay.pdf?sequence=1&isAllowed=y">DSP based vector control of induction motor</a>” at AAU’s Department of Electrical and Computer Engineering in 2007 (<a href="https://web.archive.org/web/20230124144908/http:/etd.aau.edu.et/bitstream/handle/123456789/6598/Abraham%20belay.pdf">archive</a>). In 2020, Abraham Belay (Prosperity Party) became minister of Innovation and Technology and in 2021 Minister of Defence. His MSc thesis presents 62% similarities with earlier work by other authors. See <strong>Annex A</strong>.</p> <p><strong>Takele Uma</strong> published an MSc thesis on “<a href="http://etd.aau.edu.et/handle/123456789/11897">Environmental and Economic Benefits of Bioslurry from Coffee Husk Relative to Chemical Fertilizer</a>” at AAU’s School of Chemical and Bio-Engineering in 2014 (<a href="https://web.archive.org/web/20230409120611/http:/etd.aau.edu.et/bitstream/handle/123456789/11897/Takele%20Uma.pdf">archive</a>). Takele Uma (PP) was mayor of Addis Ababa from 2018 to 2020 and minister of Mines and Petroleum of Ethiopia from 2020 to 2023. His MSc thesis presents 38% similarities with earlier work. See <strong>Annex B</strong>.</p> <p><strong>Dagmawit Moges</strong> (PP) published an MSc thesis on “<a href="http://etd.aau.edu.et/handle/123456789/19432">The Challenges and Prospects of Dynamic Electronic Service in Strengthening Demand-Responsive Transportation System in Addis Ababa, Ethiopia</a>” at AAU’s Department of Public Administration and Development Management in 2019 (<a href="https://web.archive.org/web/20230409120757/http:/etd.aau.edu.et/bitstream/handle/123456789/19432/Dagmawit%20Moges.pdf">archive</a>). Dagmawit was minister of Transport and Communications of Ethiopia from 2018 to 2023. In March 2023, she became director at the Peace Fund Secretariat of the African Union Commission. She <a href="https://smartermobility-africa.com/speakers/h-e-dagmawit-moges-bekele/">chairs also the boards</a> of Woldiya University, the Ethiopian Post, and the Ethiopian Roads Authority. Her MSc thesis presents 33% similarities with earlier work. <strong>See Annex C</strong>.</p> <p>We encourage readers to review the entire Turnitin reports for the three theses. It is not impossible that further paragraphs extracted from grey literature were missed by the Turnitin algorithms.</p> <p>On 28 February 2020, Addis Ababa University sent a <a href="https://archive.today/2023.04.08-084235/https:/t.me/Addisababauniversity/105">message in Amharic on its Telegram</a> channel, stating in substance that AAU has an anti-plagiarism policy, since plagiarism has been a significant hindrance to the University's efforts to provide excellent education. AAU further states that, earlier on, it has revoked a master's degree it had issued to a student due to plagiarism. It is only logical, then, that the three MSc theses annexed be formally scrutinized for plagiarism and the degrees possibly rescinded.</p> <p>Last but not least, examining the legitimacy of the country's politicians' academic degrees would necessitate a collaborative effort, and we urge concerned academics, students, and alumni of Ethiopian universities to initiate a crowdsourcing effort to detect plagiarism in theses, as is done in Germany with <a href="https://en.wikipedia.org/wiki/VroniPlag_Wiki">VroniPlag Wiki</a> and in Russia with <a href="https://en.wikipedia.org/wiki/Dissernet">Dissernet</a>.</p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>The authors acknowledge the owners of Twitter accounts @EgleGeek and @Zeate1 for pointing out that the three MSc theses addressed here reused texts from previous works. We would also like to thank Alex de Waal (Tufts University, Medford, MA, USA) and Boud Roukema (Nicolaus Copernicus University, Toruń, Poland) for exchanges of thoughts.</p> <p> </p> <p><strong>Annexes</strong></p> <p>Annex A. <a href="https://zenodo.org/record/7810624/files/Annex%20A%20-%20Abraham%20belay%20MSc%20thesis%20SimCheck2.pdf">Turnitin similarity report of the MSc thesis by Abraham Belay (2007)</a>.</p> <p>Annex B. <a href="https://zenodo.org/record/7810624/files/Annex%20B%20-%20Takele%20Uma%20MSc%20thesis%20similarity%20check.pdf">Turnitin similarity report of the MSc thesis by Takele Uma (2014).</a></p> <p>Annex C. <a href="https://zenodo.org/record/7810624/files/Annex%20C%20-%20Dagmawit%20Moges%20MSc%20thesis%20similarity%20check.pdf">Turnitin similarity report of the MSc thesis by Dagmawit Moges (2019)</a>.</p>
Supplementary Material for "How Students Plagiarize Modeling Assignments"
<p>Supplementary Material for the Paper "How Students Plagiarize Modeling Assignments" at the Educators Symposium at MODELS'23.</p>
Replication Package for "Mitigating Automated Obfuscation Attacks on Software Plagiarism Detection Systems"
<p>This is the replication package for the doctoral dissertation titled "<em>Mitigating Automated Obfuscation Attacks on Software Plagiarism Detection Systems</em>".</p> <p>The contributions of the dissertation were also integrated into the source code plagiarism detection system <a href="https://github.com/jplag/JPlag/">JPlag</a> to ensure they are widely accessible.</p> <p><strong>Contents Overview:</strong></p> <p>- <strong>Datasets</strong>: the artifacts of the evaluation datasets.<br>- <strong>Raw Results</strong>: the measured results of our evaluation.<br>- <strong>Evaluation Scripts</strong>: the evaluation code for plotting and statistical tests.<br>- <strong>Implementation</strong>: the source code of the JPlag-based implementation and the prebuilt application as a JAR file.<br>- <strong>Other</strong>: additional plots.</p>
Unintentional Plagiarism
<p>This video was produced for an information literacy course at University of Ottawa’s School of Information Studies (ÉSIS) and is shared as an open educational resource to be (re)used to teach viewers about Unintentional Plagiarism.</p> <p>Unintentional plagiarism refers to acts of plagiarism committed unknowingly, due to a lack of awareness rather than malicious intent. The video provides tips for students to avoid plagiarism in the academic work.</p> <p>This video and its contents were produced and recorded using Microsoft PowerPoint.</p> <p>This video was created using inspiration from the following sources:</p> <p>Evering, L. C. & Moorman G. (2012, Sept.). Rethinking plagiarism in the digital age. <em>Journal of Adolescent and Adult Literacy, 56</em>(1), 35-44. https://doi:10.1002/JAAL.00100</p> <p>Lin, Y., & Clark, K. D. (2021). Speech assignments and plagiarism in first year public speaking classes: an investigation of students’ moral attributes in relation to their behavioral intention. <em>Communication Quarterly, 69</em>(1), 23–42. https://doi.org/10.1080/01463373.2020.1864429</p> <p>Zafron, M. L. (2012). Good intentions: providing students with skills to avoid accidental plagiarism. <em>Medical Reference Services Quarterly, 31</em>(2), 225–229. https://doi.org/10.1080/02763869.2012.670605</p>
Reproduction Package for "Preventing Refactoring Attacks on Software Plagiarism Detection through Graph-Based Structural Normalization"
<p>This repository stores all data used in the evaluation of the master's thesis "Preventing Refactoring Attacks on Software Plagiarism Detection through Graph-Based Structural Normalization". It ensures the continuous reproducibility of the results of the thesis.</p> <p>Content:</p> <ul> <li>JPlag v5.1.0 including the Java CPG frontend <ul> <li>Code base</li> <li>Runnable JAR</li> </ul> </li> <li>Data sets used for evaluation</li> <li>Evaluation results</li> <li>R script used to process the results</li> <li>Graphics and tables generated from the results</li> </ul> <p>The data sets were generated by Nils Niehues and Moritz Brödel and were originally published here:</p> <ul> <li><a href="../records/10430322">Supplementary Material for "Detecting Automatic Software Plagiarism via Token Sequence Normalization" (zenodo.org)</a></li> <li><a href="../records/10149536">Reproduction package for: Intelligent Match Merging to Prevent Obfuscation Attacks on Software Plagiarism Detectors (zenodo.org)</a></li> </ul> <p>The data sets are based partly on PROGPedia, available here:</p> <ul> <li><a href="../records/7449056">PROGpedia (zenodo.org)</a></li> </ul> <p>Visit <a title="State-of-the-Art Software Plagiarism & Collusion Detection" href="jplag.github.io/JPlag/" target="_blank" rel="noopener">JPlag</a> on GitHub for the current version.<br>See the thesis document for more information.</p> <p>Read about similar publications about Plagiarism Detection <a title="JPlag" href="https://jplag.github.io/MinimalLandingPage/" target="_blank" rel="noopener">here</a>.</p>
Supplementary Material for "Detecting Automatic Software Plagiarism via Token Sequence Normalization"
<p>This repository contains additional material supporting the paper titled "Detecting Automatic Software Plagiarism via Token Sequence Normalization", presented at ICSE 2024 (research track).</p> <div> <div> <div> <p>The paper presents a defense mechanism against automated plagiarism generators utilizing program dependence graphs and demonstrates its effectiveness in countering insertion-based and reordering-based obfuscation attacks.</p> </div> </div> </div> <p>The defense mechanism was also integrated into the software plagiarism detector <a title="JPlag Repository on GitHub" href="https://github.com/jplag/JPlag">JPlag</a>, thus providing a widely accessible solution.</p> <p><strong>Contents Overview:</strong></p> <ul> <li> <p><strong>Datasets:</strong> Two datasets from the <a title="PROGpedia Repository" href="../record/7449056">PROGpedia</a> collection and two internal datasets. For the latter, only the metadata is available due to the sensitive nature of the data.</p> </li> <li> <p><strong>Plagiarized Submissions:</strong> Generated plagiarism instances illustrating various obfuscation methods such as insertion, reordering, and insert-after-reordering.</p> </li> <li> <p><strong>Evaluation Data:</strong> JSON files detailing calculated similarities and runtime measurements for all datasets.</p> </li> <li> <p><strong>Source Code:</strong> The implementation of our defense mechanism based on the software plagiarism detector <a title="JPlag Repository on GitHub" href="https://github.com/jplag/JPlag">JPlag</a> (v4.0.0). Note that JPlag is licensed under the GPL-3.0 license.</p> </li> <li> <p><strong>Evaluation Code:</strong> Python code for runtime measurements.</p> </li> <li> <p><strong>Interactive Plots:</strong> HTML visualizations of the paper's plots, offering dynamic insights into the research findings. particularly focusing on the detection of automatic software plagiarism through token sequence normalization.</p> </li> <li><strong>Demo:</strong> A packaged JAR of our implementation alongside an instruction on how to execute it.</li> </ul>
Supplementary Material for "Automated Detection of AI-Obfuscated Plagiarism in Modeling Assignments"
<p>This repository contains additional material supporting the paper titled "Automated Detection of AI-Obfuscated Plagiarism in Modeling Assignments", presented at ICSE 2024 (SEET track).</p> <p>The paper presents a token-based approach for detecting modeling plagiarism. It leverages a novel normalization technique to achieve resilience against common obfuscation attacks.</p> <p>The approach was also integrated into the software plagiarism detector <a title="JPlag Repository on GitHub" href="https://github.com/jplag/JPlag">JPlag</a>, thus providing a widely accessible solution.</p> <p><strong>Contents Overview:</strong></p> <ul> <li><strong>Source Code:</strong> The implementation of our approach (contribution 1) based on the software plagiarism detector <a title="JPlag Repository on GitHub" href="https://github.com/jplag/JPlag">JPlag</a> (v4.0.0). Note that JPlag is licensed under the GPL-3.0 license.</li> <li><strong>ChatGPT Study:</strong> An exploration of how ChatGPT can be exploited for cheating in modeling assignments.</li> <li><strong>Datasets:</strong> The three datasets of our evaluation based on EMF modeling assignments.</li> <li><strong>Raw Results:</strong> All raw data of our evaluation results, as used in our plots and tables.</li> <li><strong>Demo:</strong> A packaged JAR of our approach's implementation alongside an instruction on how to use it.</li> </ul>
ConPlag: a Dataset of Programming Contest Plagiarism in Java
<p>This package presents ConPlag, the first dataset of programming contest plagiarism in Java. To find the details about how to use the dataset and how to run tools on it, please refer to the README.</p>
Reproduction package for: Intelligent Match Merging to Prevent Obfuscation Attacks on Software Plagiarism Detectors
<p>This repository serves as the reproduction package for the master's thesis titled 'Intelligent Match Merging to Prevent Obfuscation Attacks on Software Plagiarism Detectors'. It includes datasets, experimental results and the implementation of the proposed approach. For additional details on the motivation, methodology, and analysis, please refer to the corresponding thesis document.</p>
ConPlag: a Dataset of Programming Contest Plagiarism in Java
<p>This package presents ConPlag, the first dataset of programming contest plagiarism in Java. To find the details about how to use the dataset and how to run tools on it, please refer to the README.</p>
Webis Plagiarism Corpus 2008 (Webis-PC-08)
<p>This corpus is outdated. Please use its successor PAN-PC-11: https://doi.org/10.5281/zenodo.3250095</p> <p>The Webis plagiarism corpus 2008 (Webis-PC-08) is a corpus for the evaluation of automatic plagiarism detection algorithms. For research purposes the corpus can be used free of charge, however, since the documents in the corpus are not free of copyrights we need assurance that you have legal access to the ACM digital library.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.