Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

411

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

411 results for “Arabs”

Learn how ShareScore rates datasets ↗
zenodo40/100

Fig. 2 in A new species of the genus Phytocoris (Heteroptera: Miridae) from the United Arab Emirates

Fig. 2. Phytocoris sweihanus sp. nov.: A – male head in lateral view; B – pygophore in lateral view; C – triangular lobe on the left side of genital opening, interior view; D – right paramere; E-G – left paramere in different views.

opencc-by-4.0Nov 2006View details →
zenodo40/100

Figures 19-20 in The first record of Orthochirus glabrifrons (Kraepelin, 1903) (Scorpiones: Buthidae) from the United Arab Emirates

Figures 19-20. Localities of Orthochirus glabrifrons in the United Arab Emirates. Figure 19. Wadi Ghulayyil Khun, near Bidiya, Fujairah. Figure 20. Wadi Saham, Fujairah.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figures 14 in The first record of Orthochirus glabrifrons (Kraepelin, 1903) (Scorpiones: Buthidae) from the United Arab Emirates

Figures 14. Distribution of Orthochirus glabrifrons in Oman and the United Arab Emirates. Red circles: type locality. Blue squair: new locality records.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figures 15-18 in The first record of Orthochirus glabrifrons (Kraepelin, 1903) (Scorpiones: Buthidae) from the United Arab Emirates

Figures 15-18. Localities of Orthochirus glabrifrons in the United Arab Emirates. Figure 15. Al Aqah, Fujairah. Figure 16. Green Mubazzarah, Al Ain. Figure 17. Wadi Madhab, Fujairah. Figure 18. Environs of Masafi, Ras al-Khaimah.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figures 2–3 in The first record of Orthochirus glabrifrons (Kraepelin, 1903) (Scorpiones: Buthidae) from the United Arab Emirates

Figures 2–3. Orthochirus glabrifrons, lectotype female, dorsal (2) and ventral (3) views. Scale bar: 10 mm.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figures 21-22. Orthochirus glabrifrons. Figure 21 in The first record of Orthochirus glabrifrons (Kraepelin, 1903) (Scorpiones: Buthidae) from the United Arab Emirates

Figures 21-22. Orthochirus glabrifrons. Figure 21. United Arab Emirates, Wadi Wurayah NP, Fujairah. Figure 22. Male from the United Arab Emirates, Wadi Saham, Fujairah, in vivo habitus.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 1. Orthochirus glabrifrons, a female from Oman, 30 in The first record of Orthochirus glabrifrons (Kraepelin, 1903) (Scorpiones: Buthidae) from the United Arab Emirates

Figure 1. Orthochirus glabrifrons, a female from Oman, 30 km N of Nizwa, Jabal Akhdar, in vivo habitus.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figures 4–9 in The first record of Orthochirus glabrifrons (Kraepelin, 1903) (Scorpiones: Buthidae) from the United Arab Emirates

Figures 4–9: Orthochirus glabrifrons, female from Oman, 30 km N of Nizwa, Jabal Akhdar. Figure 4. Pedipalp chela dorsal. Figures 5–6. Right legs III–IV, retrolateral aspect. Figures 7–9. Metasoma and telson, dorsal (7), ventral (8) and lateral (9).

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figures 10-13 in The first record of Orthochirus glabrifrons (Kraepelin, 1903) (Scorpiones: Buthidae) from the United Arab Emirates

Figures 10-13. Orthochirus glabrifrons from United Arab Emirates, Fujairah, Dhadna District. Figures 10–11. Male, dorsal (10) and ventral (11) views. Figures 12–13. Female, dorsal (12) and ventral (13) views. Scale bar: 10 mm.

opencc-by-4.0Dec 2023View details →
zenodo40/100

SDADDS-Guelma : A Multi-purpose Dataset for Synthetic Degraded Arabic Documents

<h1><strong>SDADDS-Guelma : A Multi-purpose Dataset for Synthetic Degraded Arabic Documents </strong></h1> <h2><strong>Description:</strong></h2> <p>This is a partial release of the SDADDS-Guelma dataset.</p> <p>SDADDS-Guelma (Synthetic Degraded Arabic Document DataSet of the University of Guelma) is a database of synthetic noisy or degraded Arabic document images. It was created by Dr. Abderrahmane Kefali and his team to support research on preprocessing, analysis, and recognition of degraded Arabic documents, where having a large set of images for training and testing is essential. This dataset is made publicly available to researchers in the field of document analysis and recognition, with the hope that it will be useful and contribute to their research endeavors.</p> <p>In this first release of the dataset, 84 handwritten images and 120 printed images have been used, along with 25 images of historical backgrounds, forming a total of 26316 synthetic images of degraded Arabic documents along with their corresponding ground-truth files.</p> <p>This release is separated into two parts to facilitate upload and use: one for the handwritten documents and the second for the printed documents.</p> <h2><strong>Composition of the dataset:</strong></h2> <p>Each of the parts of the SDADDS-Guelma dataset is organized into directories as follows:</p> <ul> <li>TXT_Files: Contains texts in UTF-8 format. </li> <li>IMG: Contains images of printed and handwritten Arabic text constructed from the text files. </li> <li>Bin_IMG: Contains binary images corresponding to the original images. </li> <li>BG_IMG: Contains images of empty old document backgrounds used for the generation of synthetic historical document images. </li> <li>GT_Files: Contains XML annotation files corresponding to the text images.</li> <li>Degraded_IMG: This directory contains synthetically generated degraded images, separated into sub-directories based on noise types such as Local_Noise, Show_through, Rotation, Curvature, Comb_IMG, etc.</li> </ul> <h2><strong>Ground-truth information:</strong></h2> <p>Ground truth information is essential for a document dataset, as it annotates documents and represents their essential characteristics. Our dataset is designed to be a large-scale and multipurpose dataset. As such, our methodology ensures that ground truth information is provided at three levels: text level (character codes), pixel level (binary and cleaned image), and document physical structure and other annotation information level.</p> <ul> <li>Textual Ground Truth: these are identical to the original texts.&nbsp;</li> <li>Pixel-level ground truth: presented in the form of binary images.</li> <li>Ground truth at the document structure level: the structure of each document image, alongside the textual transcription of the words and PAWs, is recorded in a corresponding XML annotation file. The XML format utilized resembles that employed in similar works with adjustments made according to the specific characteristics of Arabic texts, including the presence of PAWs.&nbsp;</li> </ul> <p>Consequently, each original text image in our dataset is associated to an XML file detailing the entire ground truth and associated metadata.&nbsp;</p> <h3><em><strong>Structure of XML file:</strong></em></h3> <p>Each XML annotation file contains metadata about the document image and text content within the image, including the language, number of lines, and font attributes. It also provides detailed information about each text line, word, and Part of Arabic Words (PAWs), including their bounding boxes and textual transcriptions.</p> <p>Thus, each ground truth file takes the following form:</p> <pre><code>&lt;DOCUMENT imageName="PR1Kufi_bin.png" height="2631" width="1860" nbTextLines="8" language="Arabic" fontName="Kufi" fontSize="34"&gt; &lt;TEXTLINE id="0" nbWords="3" boundingBox="215,481,355,1379"&gt; &lt;WORD id="0" nbPAWs="2" boundingBox="217,1065,341,1379" transcription="خصائص"&gt; &lt;PAW id="0" nbCCs="2" boundingBox="217,1206,322,1379" transcription="خصا"&gt; &lt;CC id="0" nbPixels="4110" pixels="(217,1206,1206);(218,1206,1207);(219,1206,1208);(220,1206,1208);(221,1206,1211);(222,1206,1211);..."&gt; &lt;/CC&gt; &lt;CC id="1" nbPixels="80" pixels="(263,1330,1336);(264,1329,1336);(265,1328,1337);(266,1328,1337);(267,1328,1337);(268,1328,1337);(269,1328,1337);...."&gt; &lt;/CC&gt; &lt;/PAW&gt; .... &lt;/WORD&gt; &lt;WORD id="1" nbPAWs="2" boundingBox="215,817,338,1044" transcription="التفسير"&gt; &lt;PAW id="0" nbCCs="1" boundingBox="215,1030,322,1044" transcription="ا"&gt; &lt;CC id="0" nbPixels="1037" pixels="(215,1030,1030);(216,1030,1030);(217,1030,1031);..."&gt;&lt;/CC&gt; .... &lt;/PAW&gt; .... &lt;/WORD&gt; &lt;/TEXTLINE&gt; .... &lt;/DOCUMENT&gt;</code></pre> <h1><strong>Contact:</strong></h1> <p>Name: &nbsp; &nbsp; &nbsp; &nbsp; Dr. Abderrahmane Kefali<br>Affiliation: &nbsp; &nbsp; University of 8 May 1945-Guelma, Algeria<br>Email: &nbsp; &nbsp; &nbsp; &nbsp; kefali.abderrahmane@univ-guelma.dz</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

National Checklists: United Arab Emirates Species List

Data from: GBIF.org (23 January 2025) GBIF Occurrence Download <a href="https://doi.org/10.15468/dl.vd2ajk" target="_blank" rel="noopener">https://doi.org/10.15468/dl.vd2ajk</a>

opencc-zeroAug 2024View details →
zenodo40/100

A Corpus of Arabic Literature (19-20th centuries) for Stylometric Tests

<p>The dataset contains three collections of mainly literary Arabic texts from the 19th and early 20th centuries.</p> <ul> <li><code>corpus022_JurjiZaydan_Dated</code>&nbsp;is a dated corpus of 22 historical novels by Jurjī Zaydān. It is well established that Jurjī Zaydān was publishing roughly one novel per year and the dates of publication are well known, which makes this corpus a valuable material for testing chronological changes in the style of individual writers.</li> <li><code>corpus065</code>&nbsp;is a corpus of 65 books by 8 authors;</li> <li><code>corpus300</code>&nbsp;contains 300 books by 28 authors;</li> </ul> <p>Texts have been collected from&nbsp;<a href="https://www.hindawi.org/">https://www.hindawi.org/</a>; the original EPUB files have been converted into clean text files (UTF8 encoding) and Arabic orthography has been normalized in the following manner: short vowels removed; the orthography of&nbsp;<em>alif</em>&nbsp;simplified (all&nbsp;<em>alif</em>s converted into bare&nbsp;<em>alif</em>s are used);&nbsp;<em>alif maqṣūraŧ</em>s converted to&nbsp;<em>yāʾ</em>s.</p> <p>For the names of authors and the names of works, the URIs follow the naming conventions used in the OpenITI project (<a href="https://github.com/openiti">https://github.com/openiti</a>, for the most up-to-date description, see&nbsp;<a href="https://kitab-project.org/corpus-and-data">https://kitab-project.org/corpus-and-data</a>). However, for better compatibility with R&nbsp;<code>stylo</code>,&nbsp;<em>underscore</em>&nbsp;(<code>_</code>) is used to connect URIs: AUTHOR_TITLE (in the dated Jurjī Zaydān corpus files are named YEAR_TITLE).</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Supporting data for the AI education publication statistics in "An Experience Report of Executive-Level Artificial Intelligence Education in the United Arab Emirates"

<p>Supporting data for the AI education publication statistics presented in the paper &quot;An Experience Report of Executive-Level Artificial Intelligence Education in the United Arab Emirates&quot; to be published at the Twelfth AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-22). The data was used to plot the figure showing the cumulative number of publications from 1976 to 2020 relating to AI education.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic

<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Figs 11–24 in Soft-winged flower beetles (Coleoptera: Malachiidae) of the United Arab Emirates

Figs 11–24. Tonycolotes kovari (Švihla, 1987) gen. et. comb. nov., ♂ (SCH_ISEA). 11. External appearance, dorsal view. 12. External appearance, lateral view. 13. Head and pronotum, subdorsal view. 14. Head and pronotum, dorsal view. 15. Right antenna. 16–18. Scapus in different positions. 19. Left palp. 20. Left anterior leg. 21. Pygidium. 22. Ultimate abdominal ventrite. 23. Aedeagus, lateral view. 24. Tegmen. Scale bars: 0.5 mm.

opencc-by-4.0Apr 2022View details →
zenodo40/100

Figs 1–10 in Soft-winged flower beetles (Coleoptera: Malachiidae) of the United Arab Emirates

Figs 1–10. Tonyattalus vanharteni gen. et sp. nov., holotype (SCH_ISEA_000133), ♂ (1–8), ♀ (9– 10). 1, 9. External appearance, dorsal view. 2, 10. External appearance, lateral view. 3. Left antenna. 4. Head and pronotum, dorsal view. 5. Left anterior tarsus. 6. Pygidium. 7. Ultimate abdominal ventrite. 8. Aedeagus, and tegmen, lateral view. Scale bars: 0.5 mm.

opencc-by-4.0Apr 2022View details →
zenodo40/100

PAN Arabic Intrinsic Plagiarism Detection Shared Task Corpus

<p>Evaluation corpus for ARAbic INtrinsic plagiarism detection (InAra Corpus)&nbsp;</p> <p>&nbsp;</p> <p>This corpus has been used in AraPlagDet 2015 shared task&nbsp;</p> <p>More details could be found in : <a href="http://araplagdet.misc-lab.org">https://araplagdet.misc-lab.org/</a>&nbsp;or <a href="http://pan.webis.de/fire15/pan15-web/index.html">https://pan.webis.de/fire15/pan15-web/index.html</a>&nbsp;</p> <p>&nbsp;</p> <p><strong>I. SYNOPSIS&nbsp;</strong></p> <p>InAra corpus comprises 2048 documents; 80% of them contain passages&nbsp;borrowed from other documents to simulate documents that contain&nbsp;plagiarized fragments. The corpus involves 2 parts: Training and test.</p> <p>&nbsp;</p> <p><strong>II. DESCRIPTION&nbsp;</strong></p> <p>Each part of the corpus (training and test) consists mainly of 2 datasets:&nbsp;textual files and XML files.&nbsp;The textual files represent the suspicious documents i.e., the documents&nbsp;that contain artificial plagiarism; and the XML files are the plagiarism&nbsp;annotation i.e. they provide for each plagiarized passage its starting&nbsp;offset in the suspicious document and its length (offset and length are both expressed in characters). A suspicious document file and its plagiarism&nbsp;annotation file share the same name.</p> <p>&nbsp;</p> <p><strong>III. PURPOSE&nbsp;</strong></p> <p>The purpose of InAra corpus is to evaluate automatic plagiarism&nbsp;detection methods, notably methods of the intrinsic approach. This&nbsp;approach consists in uncovering the plagiarized passages on the basis of&nbsp;the writing style inconsistency in a given suspicious document. As&nbsp;opposed to the external approach, the intrinsic approach does not&nbsp;necessitate any comparison of the suspicious document against the&nbsp;potential sources of plagiarism. Hence, InAra corpus is not appropriate for the evaluation of the external plagiarism detection because the source of plagiarism are not provided.</p> <p>It should be noted that some documents in InAra corpus contain religious&nbsp;quotations (e.g., Quran and Hadith). These quotations have a peculiar writing style and then a simple intrinsic plagiarism detection software can consider them as plagiarism. However, quotations are not plagiarism, and they are not&nbsp;annotated in the XML files in InAra. Hence, it is an important feature for the plagiarism detection systems evaluated on InAra to not consider religious quotations as plagiarism cases unless they appear as part of a larger&nbsp; plagiarism case.</p> <p>&nbsp;</p> <p><strong>IV. BUILDING METHODS&nbsp;</strong></p> <p>The documents that compose InAra corpus do not contain actual plagiarism&nbsp;cases. They are rather artificial suspicious documents in which&nbsp;plagiarism was created automatically by a software that takes fragments&nbsp;of text from one or more sources documents and inserts them in another&nbsp;one according to a set of parameters, namely the percentage of plagiarism&nbsp;and the plagiarized passages lengths. This building method is the same&nbsp;used to construct PAN 2009-2011 corpora of plagiarism detection (see&nbsp;<a href="http://pan.webis.de">http://pan.webis.de</a> for more information on PAN competition and its&nbsp;corpora).&nbsp;</p> <p>&nbsp;</p> <p><strong>V. LANGUAGE AND ENCODING&nbsp;</strong></p> <p>All the textual documents of this corpus are written in Arabic language&nbsp;and encoded in UTF-8 without BOM.</p> <p>&nbsp;</p> <p><strong>VI. SOURCES OF TEXTS&nbsp;</strong></p> <p>Texts used to build this corpus, either suspicious documents or the&nbsp;inserted passages, are taken mainly from the open library Arabic&nbsp;Wikisource (http://ar.wikisource.org), one of Wikimedia Foundation&nbsp;projects. A few numbers of documents were taken from other websites,&nbsp;namely:&nbsp;</p> <ul> <li>Create your own country blog: http://diycountry.blogspot.com&nbsp;</li> <li>Corpus of Classical Arabic (KSUCCA): http://ksucorpus.ksu.edu.sa&nbsp;</li> <li>Islamic book web site: http://www.islamicbook.ws&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>VII. COPYRIGHT AND AVAILABILITY&nbsp;</strong></p> <p>We were very careful to build the corpus with copyright-free texts only,&nbsp;to be able to make it publicly available without any sort of problems&nbsp;with texts owners.&nbsp;</p> <p>&nbsp;</p> <p><strong>VIII. HOW TO CITE THE CORPUS ?</strong></p> <p>If you publish a paper about your experimentations using InAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on InAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>Additional information on the corpus building are in the papers:</p> <ul> <li>Bensalem, I., Rosso, P., Chikhi, S.: A New Corpus for the Evaluation&nbsp;of Arabic Intrinsic Plagiarism Detection. In: Forner, P., M&uuml;ller, H., Paredes, R., Rosso, P., and Stein, B. (eds.) CLEF 2013, LNCS, vol. 8138. pp. 53&ndash;58. Springer, Heidelberg (2013).</li> <li>Bensalem, I., Rosso, P., Chikhi, S.: Building Arabic Corpora from&nbsp;Wikisource. 10th ACS/IEEE International Conference on Computer Systems&nbsp;and Applications (AICCSA&rsquo;13),May 27-30 Fes/Ifran, Morocco (2013).IEEE.&nbsp;</li> </ul> <p>&nbsp;</p> <p>You may wish to&nbsp;compare the results of your experiments with the result of the following papers that used InAra corpus:</p> <ul> <li>Bensalem I, Rosso P, Chikhi S (2019) On the use of character n-grams&nbsp;as the only intrinsic evidence of plagiarism. Language Resources and&nbsp;Evaluation 53:363&ndash;396. doi: 10.1007/s10579-019-09444-w</li> <li>Mahgoub AY, Magooda A, Rashwan M, et al (2015) RDI System for&nbsp;Intrinsic Plagiarism Detection (RDI_RID), Working Notes for&nbsp;PAN-AraPlagDet at FIRE 2015. In: Majumder P, Mitra M, Agrawal M,&nbsp;&nbsp;Mehta P (eds) Post Proceedings of the Workshops at the 7th Forum for&nbsp;Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587. CEUR-WS.org, pp 129&ndash;130</li> </ul> <p>&nbsp;</p> <p><strong>IX. WARNING&nbsp;</strong></p> <p>It should be noted that the Arabic texts may contain quotations from the&nbsp;Quran and the Hadith; and due to the fact that text insertion is&nbsp;automatic and in random positions, it is possible that the plagiarized&nbsp;text is inserted unintentionally between Quranic verses or sentences of&nbsp;a Hadith cited in a document. Hence, the inserted passages may alter&nbsp;the meaning of the original text. For these reasons, this corpus must&nbsp;not be used outside the purpose for which it was built. Examples of the&nbsp;inappropriate use include using the corpus documents as a source of&nbsp;knowledge or distributing them without mentioning that they contain&nbsp;borrowed texts. If you are not interested in plagiarism detection and&nbsp;you are retaining the corpus because it contains books you want to read,&nbsp;then this corpus is not the right source. Please, you should refer to the&nbsp;</p> <p>sources mentioned in Section VI where you can find the original content of&nbsp;the books you are looking for. We emphasize that we are not responsible&nbsp;for the results of any use of this corpus other than the evaluation of&nbsp;the intrinsic plagiarism detection methods.&nbsp;</p> <p>&nbsp;</p> <p><strong>X. CONTACT US</strong></p> <p>We will be happy to hear from you about your experience in using InAra&nbsp;corpus. Please do not hesitate to contact us with the following email&nbsp;address: bens.imene@gmail.com</p> <p>&nbsp;</p> <p>Imene Bensalem&sup1;, Paolo Rosso&sup2;, Salim Chikhi&sup1;</p> <p>&sup1;MISC Lab. Constantine 2 university, Algeria</p> <p>&sup2;PRHLT, Universitat Polit&egrave;cnica de Val&egrave;ncia, Spain&nbsp;</p>

opencc-by-4.0Jun 2015View details →
zenodo40/100

PAN Arabic External Plagiarism Detection Shared Task Corpus

<p>Evaluation Corpus for ARAbic EXternal plagiarism detection (ExAra Corpus)&nbsp;</p> <p>&nbsp;</p> <p>This corpus has been used in AraPlagDet 2015 shared task&nbsp;</p> <p>More details could be found in : <a href="http://araplagdet.misc-lab.org">https://araplagdet.misc-lab.org/</a>&nbsp;or <a href="http://pan.webis.de/fire15/pan15-web/index.html">https://pan.webis.de/fire15/pan15-web/index.html</a></p> <p>If you publish a paper about your experimentations using ExAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on ExAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>&nbsp;</p> <p><strong>I. SYNOPSIS&nbsp;</strong></p> <p>ExAra corpus comprises 2345 documents; almost half of them (suspecious doucuments) contain passages borrowed from the other half (source docucments) to simulate documents that contain plagiarized fragments. The corpus involves 2 parts: Training and test.</p> <p>&nbsp;</p> <p><strong>II. DESCRIPTION&nbsp;</strong></p> <p>Each part of the corpus (training and test) consists mainly of 3 datasets:&nbsp; 2 sets of textual files and 1 set of XML files. The 2 sets of the textual files are the suspicious documents (i.e. the documents that contain artificial plagiarism) and the source documents (i.e., the documents from which the suspicious passages have been plagiarised). The 3rd set of documents contains XML files, which are the plagiarism annotation, i.e., they provide for each plagiarized passage its starting offset and its length in both the suspicious and source documents (offset and length were both expressed in characters). A suspicious document file (.txt) and its plagiarism annotation file (.xml) share the same name.</p> <p>&nbsp;</p> <p><strong>III. PURPOSE&nbsp;</strong></p> <p>The purpose of ExAra corpus is to evaluate automatic plagiarism detection methods, notably methods of the External approach. This approach consists in uncovering the plagiarized passages on the basis of their similarity with passages in the source documents.</p> <p>It should be noted that some suspicious documents in ExAra corpus contain religious quotations (e.g., Quran and Hadith) and common phrases. Some of them appear also in some source documents, and hence a simple plagiarism detection software can consider them as plagiarism. However, quotations and common phrases are legitimate text reuse cases and are not annotated in the XML files in ExAra. Therefore, it is an important feature for the&nbsp;plagiarism detection systems evaluated on ExAra to not consider religious quotations and common phrases as plagiarism cases unless they appear as part of a larger plagiarism case.</p> <p>&nbsp;</p> <p><strong>IV. BUILDING METHODS&nbsp;</strong></p> <p>The documents that compose ExAra corpus do not contain actual plagiarism&nbsp; cases, they are rather artificial suspicious documents in which&nbsp;plagiarism was created automatically by a software that takes fragments&nbsp;of text from one or more sources documents and inserts them in another&nbsp;one according to a set of parameters, namely the percentage of plagiarism&nbsp;and the lengths of the plagiarized passages. Some of the plagiarised fragments are&nbsp;obfuscated manually or automatically before inserting them in the suspicious documents.</p> <p>This building method is the same used to construct PAN 2009-2011 corpora of plagiarism detection (see http://pan.webis.de for more information on PAN competition and its&nbsp;corpora).&nbsp;</p> <p>&nbsp;</p> <p><strong>V. LANGUAGE AND ENCODING&nbsp;</strong></p> <p>All the textual documents of this corpus are written in Arabic language&nbsp;and encoded in UTF-8 without BOM.</p> <p>&nbsp;</p> <p><strong>VI. HOW TO CITE THE CORPUS ?</strong></p> <p>If you publish a paper about your experimentations using ExAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on ExAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>&nbsp;</p> <p><strong>VII. WARNING&nbsp;</strong></p> <p>It should be noted that the Arabic texts may contain quotations from the&nbsp;Quran and the Hadith; and due to the fact that text insertion is&nbsp;automatic and in random positions, it is possible that the plagiarized&nbsp;text is inserted unintentionally between Quranic verses or sentences of&nbsp;a Hadith cited in a document. Hence, the inserted passages may alter&nbsp;the meaning of the original text. For these reasons, this corpus must&nbsp;not be used outside the purpose for which it was built. Examples of the&nbsp;inappropriate use include using the corpus documents as a source of&nbsp;knowledge or distributing them without mentioning that they contain&nbsp;borrowed texts. If you are not interested in plagiarism detection, and you are retaining the corpus because it contains articles you want to read,&nbsp;then this corpus is not the right source. Please, you should refer to the&nbsp;sources mentioned in (Bensalem et al. 2015) (i.e.,the paper above) where&nbsp;you can find the original content of the articles you are looking for.</p> <p>We emphasize that we are not responsible for the results of any use of this corpus other than the evaluation of the external plagiarism detection methods.&nbsp;</p> <p>&nbsp;</p> <p><strong>VIII. CONTACT US</strong></p> <p>We will be happy to hear from you about your experience in using ExAra&nbsp;corpus. Please do not hesitate to contact us with the following email&nbsp;address: bens.imene@gmail.com</p> <p>&nbsp;</p> <p>Imene Bensalem&sup1;, Imene Boukhalfa&sup1;, Paolo Rosso&sup2;, Salim Chikhi&sup1;</p> <p>&sup1;MISC Lab. Constantine 2 university, Algeria</p> <p>&sup2;PRHLT, Universitat Polit&egrave;cnica de Val&egrave;ncia, Spain&nbsp;</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

Online Supplement for An Arabic Version of The Visual Aesthetics of Websites Inventory (AR-VisAWI): Translation and Psychometric Properties

<p>This is the online supplement for a translation of the Visual Aesthetics of Websites Inventory into Arabic (AR-VisAWI).&nbsp;</p> <p>In the field of human-computer interaction, the concept of visual aesthetics gained popularity after researchers started to recognize its merits and effects on user experience. Yet, no proper instrument exists to assess the visual aesthetics of websites that are intended for users who speak Arabic as their native language. As such, the aim of this study was to develop and evaluate an Arabic version of the Visual Aesthetics of Websites Inventory (VisAWI, Moshagen &amp; Thielsch, 2010) and its short version (VisAWI-S, Moshagen &amp; Thielsch, 2013). For this purpose, participants were asked to evaluate a randomly assigned website with the AR-VisAWI and with different validating instruments. A final sample of 223 participants was included in the analyses.</p> <p>This online supplement includes</p> <ul> <li>a codebook describing all instructions and items</li> <li>raw data (anonymised) and analysis script (Note: The raw data contains only the information of persons who have agreed to be included in the analysis. Some demographic information was deleted to ensure anonymity.)</li> <li>Questionnaire template and scoring instructions</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo40/100

IPBES Invasive Alien Species Assessment: Summary for Policymakers. Figures, tables and captions in Arabic

<p>Arabic translations of the figures, tables and their captions from the Summary for Policymakers of the Thematic Assessment Report on Invasive Alien Species and their Control of the Intergovernmental Science-Policy Platform on Biodiversity and Ecosystem Services.</p>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record