Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
73
datasets available to search
ShareScore release 0.9.0
Dataset results
73 results for “Competition datasets”
"ICDAR2023 Competition on Detection and Recognition of Greek Letters on Papyri" Dataset
<h1><strong>Dataset description of the “ICDAR2023 Competition on Detection and Recognition of Greek Letters on Papyri”</strong></h1> <p>Prof. Dr. Isabelle Marthot-Santaniello, Dr. Olga Serbaeva</p> <p>2024.09.16</p> <h2>Introduction</h2> <p>The present dataset stems from the ICDAR2023 Competition on Detection and Recognition of Greek Letters on Papyri (original links to the competition are provided in the file “1b.CompetitionLinks.”)</p> <p>The aim of this competition was to investigate the performance of glyph detection and recognition in a very challenging type of historical document: Greek papyri. The detection and recognition of Greek letters on papyri is a preliminary step for computational analysis of handwriting that can lead to major steps forward in our understanding of this important source of information on Antiquity. Such detection and recognition can be done manually by trained papyrologists. It is, however, a time-consuming task that would need automatising. </p> <p>We provide here the documents related to two different tasks: localisation and classification. The document images are provided by several institutions and are representative of the diversity of book hands on papyri (a millennium time span, various script styles, provenance, states of preservation, means of digitization and resolution).</p> <h2>How the dataset was constructed</h2> <p>In the frame of <a href="https://d-scribes.philhist.unibas.ch/en/case-studies/iliad-208/" target="_blank" rel="noopener">D-Scribes project</a> lead by Prof. Dr. Isabelle Marthot-Santaniello, 2018-2023, around 150 papyri fragments containing Iliad were manually annotated at a letter-level in <a href="https://github.com/readsoftware/read" target="_blank" rel="noopener">READ</a>.</p> <p>The editions were taken, for the major part, from <a href="papyri.info" target="_blank" rel="noopener">papyri.info</a>, and were simplified, i.e. the accents, editorial marks, and other additional information were removed to be as close as possible to what is to be found on papyri. When the text was not available on papyri.info, the relevant passage was extracted from the <a href="https://github.com/PerseusDL/canonical-greekLit/blob/master/data/tlg0012/tlg001/tlg0012.tlg001.perseus-grc2.xml" target="_blank" rel="noopener">Homer Iliad of Perseus</a>.</p> <p>From those, 150 plus papyri fragments, 185 surfaces (sides of fragments) belonging to 136 different manuscript identified by their Trismegistos numbers, (further TMs) were selected to serve as a material for Competition. These 185 surfaces were separated into the “training set” and the “test set” provided for the competition as a set of images and corresponding data in JSON format.</p> <p>Details on the competition summarised in "ICDAR 2023 Competition on Detection and Recognition of Greek Letters on Papyri", by Mathias Seuret, Isabelle Marthot-Santaniello, Stephen A. White, Olga Serbaeva Saraogi, Selaudin Agolli, Guillaume Carrière, Dalia Rodriguez-Salas, and Vincent Christlein; edited by G. A. Fink et al. (Eds.): <em>ICDAR 2023,</em> LNCS 14188, pp. 498–507, 2023. https://doi.org/10.1007/978-3-031-41679-8_29.</p> <p>After the competition ended, the decision was taken to release manually annotated dataset for the “test set” as well. Please find the description of each included document below.</p> <h2><br>Dataset Structure</h2> <p><br><strong>“1. CompetitionOverview.xlsx”</strong> contains the metadata of the used images in Excel file, state 2024.09.19. Here is the structure of the Excel file:</p> <p> </p> <table> <tbody> <tr> <td> <p><strong>Excel columns</strong></p> </td> <td> <p><strong>Name</strong></p> </td> <td> <p><strong>Content</strong></p> </td> <td> <p><strong>Notes</strong></p> </td> </tr> <tr> <td> <p><strong>A</strong></p> </td> <td> <p><strong>TM</strong></p> </td> <td> <p><strong>Trismegistos number is internationally used for papyri identification</strong></p> </td> <td> <p><strong>With READ item name in ().</strong></p> </td> </tr> <tr> <td> <p><strong>B</strong></p> </td> <td> <p><strong>Papyri.info link</strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p> </p> </td> </tr> <tr> <td> <p><strong>C</strong></p> </td> <td> <p><strong>Fragments' Owning Institution (from <a href="http://papyri.info"><u>papyri.info</u></a>) </strong></p> </td> <td> <p><strong>Institution’s name</strong></p> </td> <td> <p><strong>Institution that physically stores the papyri</strong></p> </td> </tr> <tr> <td> <p><strong>D</strong></p> </td> <td> <p><strong>Availability (of metadata, <a href="http://papyri.info"><u>papyri.info</u></a>) </strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p><strong>Metadata reuse clarification</strong></p> </td> </tr> <tr> <td> <p><strong>E</strong></p> </td> <td> <p><strong>text ID (READ)</strong></p> </td> <td> <p><strong>Number from READ SQL database that was used to link the images and the editions.</strong></p> </td> <td> <p><strong>Serves to locate the attached images and understand the JSON structure.</strong></p> </td> </tr> <tr> <td> <p><strong>F</strong></p> </td> <td> <p><strong>Test/Training</strong></p> </td> <td> <p> </p> </td> <td> <p><strong> I.e. the image was originally included in the training or in the test set of the dataset.</strong></p> </td> </tr> <tr> <td> <p><strong>G</strong></p> </td> <td> <p><strong>Image Name (for orientation)</strong></p> </td> <td> <p> </p> </td> <td> <p><strong>As in READ</strong></p> </td> </tr> <tr> <td> <p><strong>H</strong></p> </td> <td> <p><strong>Cedopal link</strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p><strong>Contains additional metadata and includes the links to all available online images.</strong></p> </td> </tr> <tr> <td> <p><strong>I</strong></p> </td> <td> <p><strong>License from the Institution webpage.</strong></p> </td> <td> <p><strong>Either license or usage summary.</strong></p> </td> <td> <p><strong>If no precise licence has been given, the summary of the reuse rights is provided with a link to the regulations in column K</strong></p> </td> </tr> <tr> <td> <p><strong>J</strong></p> </td> <td> <p><strong>Image URL</strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p><strong>Not all images are available online. Please contact the owning institution directly if the image is not available.</strong></p> </td> </tr> <tr> <td> <p><strong>K</strong></p> </td> <td> <p><strong>Information on the image usage from the institution</strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p><strong>In case of any doubt, please contact the owning institution directly.</strong></p> </td> </tr> <tr> <td> <p><strong>L</strong></p> </td> <td> <p><strong>Notes</strong></p> </td> <td> <p> </p> </td> <td> <p> </p> </td> </tr> </tbody> </table> <p>For the purpose of an easy overview, the items with special problems, i.e. images not online or missing links, have been marked in red.</p> <p><strong>2. There are three data subsets:</strong></p> <p><strong>2a. “Training file” </strong><br>(containing 150 papyri images separated into 108 texts and HomerCompTraining.json). The images are those of papyri containing Iliad of Homer in JPG-format. These were processed in READ, namely, each visible letter on a given papyri was linked to the edition of the Iliad, through this process, each linked letter of the edition was linked to its coordinates in pixels on the HTML-surface of the image. All that information is provided in the JSON-file.</p> <p>The JSON file contains the <strong>“annotations”</strong> (b-boxes of each letter/sign), <strong>“categories”</strong> (Greek letters),<strong> “images”</strong> (Image IDs), and <strong>“licenses”</strong>. The links between image and bboxes is defined via the “id” in the “images” part (for example, "id": 6109). This same id is encoded as “"image_id": 6109” in the “annotations”. Alternatively, “text_id” which can be found in the “images” URL and in the file-names provided here and containing images, can be used for data linking.</p> <p>Let us now describe the content of each part of the JSON file:<br>Each <strong>“annotation”</strong> contains<br>“area" characterised as “bbox" with coordinates, <br>“category_id”, that allows to identify which Greek letter in categories is represented by the number; “id”, which is a unique number of the cliplet, i.e. area; <br>“image_id”, that links cliplet to the surface of the image having the same id; <br>“iscrowd" and “seg_id" are useful to find the information back in READ database; <br>and, finally, “tags”.</p> <p>In tags, “BaseType" was used to annotate quality as described below. “FootMarkType”, ft1, etc., was used for clustering tests, but played no role for the Competition.<br>“BaseType” ot bt-tags were assigned to the letters to mark the quality of preservation: <br>bt-1: well-preserved letter that should allows easy identification for both human eyes and the Computer-vision; <br>bt-2: Partially preserved letter that might also have some background damage (holes, additional ink, etc), but remains readable, and has one interpretation. <br>bt-3: Letters damaged to such an extant that they cannot be identified without reading an edition. These are treated as traces of ink. <br>bt-4: The letters that have some damage, but this damage is of such kind that it makes possible multiple interpretations. For example, missing/defaced horizontal stroke makes alpha indistinguishable from damaged delta or lambda.</p> <p>Each <strong>“category”</strong> contains <br>“id”, this is a number references also in “annotations” and it allows to identify which Greek letter was in the bbox; <br>”name”, for example, “χ”; <br>and “supercategory”, i.e. “Greek”.</p> <p>Each <strong>“image”</strong> contains the following sub fields: <br>“bln_id" is an internal READ number of the html surface; <br>"date_captured": null - is another READ field; <br>"file_name": “./images/homer2/txt1/P.Corn.Inv.MSS.A.101.XIII.jpg", allows to link easy image and text, i.e. for the image in question the JPG will be in the file called “txt1”, it is very similar by structure and function to "img_url": "./images/homer2/txt1/P.Corn.Inv.MSS.A.101.XIII.jpg"; <br>each image has “height" and “width" expressed in pixels. <br>Each image has “id”, and this id is referenced in the “annotations” under “image_id”. <br>Finally, each image contains a link to “license”, expressed as a number. </p> <p>Each <strong>“licence”</strong> lists a license as it was found during the time of competition, i.e. in February 2023.</p> <p><strong>2b. “Test file”</strong> <br>contains 34 papyri image sides separated into 31 TMs and HomerCompTesting.json The JSON file here only allows to connect the images with the “categories”, “images”, “licenses”, but without the “annotations”. The structure and logic is otherwise the same like in “Training” JSON.</p> <p><strong>2c. “Answers file” </strong><br>Containing the “annotations” and other information for the 34 papyri of the “Testing” dataset. The structure and logic is the same like in “Training” JSON.</p> <p><strong>3. “Additional files” </strong><br>Containing lists of duplicate segments id (multiple possible readings or tags), respectively 6 items for “Training”, 17 for “Testing” and 15 for “Answers”.</p> <p><strong>4. “Dataset Description”</strong><br>This same description included for completeness.</p> <h2>References</h2> <p>The Dataset was reused or mentioned in a number of publications (state September 2024)</p> <p>Mohammed, H., Jampour, M. (2024). "From Detection to Modelling: An End-to-End Paleographic System for Analysing Historical Handwriting Styles". In: Sfikas, G., Retsinas, G. (eds) <em>Document Analysis Systems. DAS 2024.</em> Lecture Notes in Computer Science, vol 14994. Springer, Cham, pp. 363–376. https://doi.org/10.1007/978-3-031-70442-0_22</p> <p>De Gregorio, G., Perrin, S., Pena, R.C.G., Marthot-Santaniello, I., Mouchère, H. (2024). "NeuroPapyri: A Deep Attention Embedding Network for Handwritten Papyri Retrieval". In: Mouchère, H., Zhu, A. (eds) <em>Document Analysis and Recognition – ICDAR 2024 Workshops. ICDAR 2024.</em> Lecture Notes in Computer Science, vol 14936. Springer, Cham, pp. 71–86. https://doi.org/10.1007/978-3-031-70642-4_5</p> <div> <p>Vu, M. T., Beurton-Aimar, M. "PapyTwin net: a Twin network for Greek letters detection on ancient Papyri". <em>HIP '23: 7th International Workshop on Historical Document Imaging and Processing, San Jose, CA, USA, August 2023.</em><br>https://doi.org/10.1145/3604951.3605522<br>https://dl.acm.org/doi/fullHtml/10.1145/3604951.3605522</p> <p>Turnbull, R., Mannix, E. "Detecting and recognizing characters in Greek papyri with YOLOv8, DeiT and SimCLR". (Preprint).<br>arXiv:2401.12513<br>https://doi.org/10.48550/arXiv.2401.12513</p> </div>
Factors determining distributions of rainforest Drosophila shift from interspecific competition to high temperature with decreasing elevation (original datasets)
<p>This repository provides the data for the manuscript "Factors determining distributions of rainforest Drosophila shift from interspecific competition to high temperature with decreasing elevation"</p> <p>We investigated thermal tolerances and interspecific competition as causes of species turnover in the nine most abundant species of <em>Drosophila</em> along elevational gradients in the Australian Wet Tropics. Specifically, we 1) analyzed the distribution patterns of the studies <em>Drosophila</em> species; 2) fitted thermal performance curves; 3) tested the correlation between multiple thermal traits and distribution patterns; 4) fitted the Beverton-Holt model to describe the single-generation intra- and inter-specific competition effect; 5) examined the long-term effect of competition and temperature on the population size of a pair of Drosophila species.</p> <p>More details are provided in the README file.</p>
Datasets and Supporting Materials for the IPIN 2021 Competition Track 3 (Smartphone-based, off-site)
<p>This package contains the datasets and supplementary materials used in the IPIN 2021 Competition.</p> <p><strong>Contents:</strong></p> <ul> <li>IPIN2021_Track03_TechnicalAnnex_V1-02.pdf: Technical annex describing the competition</li> <li>01-Logfiles: This folder contains a subfolder with the 105 training logfiles, 80 of them single floor indoors, 10 in outdoor areas, 10 of them in the indoor auditorium with floor-trasitio and 5 of them in floor-transition zones, a subfolder with the 20 validation logfiles, and a subfolder with the 3 blind evaluation logfile as provided to competitors.</li> <li>02-Supplementary_Materials: This folder contains the matlab/octave parser, the raster maps, the files for the matlab tools and the trajectory visualization.</li> <li>03-Evaluation: This folder contains the scripts used to calculate the competition metric, the 75th percentile on the 82 evaluation points. It requires the Matlab Mapping Toolbox. The ground truth is also provided as 3 csv files. Since the results must be provided with a 2Hz freq. starting from apptimestamp 0, the GT files include the closest timestamp matching the timing provided by competitors for the 3 evaluation logfiles. It contains samples of reported estimations and the corresponding results.</li> </ul> <p><strong>Please, cite the following works when using the datasets included in this package:</strong></p> <ul> <li>Torres-Sospedra, J.; et al. Datasets and Supporting Materials for the IPIN 2021 Competition Track 3 (Smartphone-based, off-site). http://dx.doi.org/10.5281/zenodo.5948678</li> </ul>
Dataset for ICFHR2018 Competition on Automated Text Recognition on a READ Dataset
<p>The main idea of this dataset is to analyse the impact of training data. How many training data specific to the document, you are transcribing, is necessary? </p> <p><strong>general data: </strong>This is a collection of heterogeneous documents to train an initial system. For each text line there is an image file of that line, a file with the ground truth text and an information file containing an automatically generated surrounding polygon.</p> <p><strong>specific data: </strong>The specific data contains documents related to the test data. For the specific systems only the images of the train list may be used. The file are of the same type as the general data.</p> <p><strong>test data: </strong>The test data contains only the images and the information files.</p> <p>More Information, some published results and an evaluation procedure at https://scriptnet.iit.demokritos.gr/competitions/10/</p>
Datasets and Supporting Materials for the IPIN 2017 Competition Track 3 (Smartphone-based, off-site)
<p>This package contains the datasets and supplementary materials used in the IPIN 2017 Competition (Sapporo, Japan).</p> <p><strong>Contents:</strong></p> <ol> <li>Track3_LogfileDescription_and_SupplementaryMaterial.pdf: Description of the logfiles and supplemental materials.</li> <li>Track3_TechnicalAnnex.pdf: Technical annex describing the competition </li> <li>01-Logfiles: This folder contains a subfolder with the 25 training logfiles, a subfolder with the 9 validation logfiles, and a subfolder with the 7 blind evaluation logfiles as provided to competitors.</li> <li>02-Supplementary_Materials: This folder contains the Matlab/Octave parser, the raster maps, the visualization of the training routes and the location of the BLE beacon (CAR) and some Wi-Fi APs (UJIUB).</li> <li>03-Evaluation: This folder contains the scripts used to calculate the competition metric, the 75th percentile on the 505 evaluation points. The ground truth is also provided in MatLab format and as a CSV file. Since the results must be provided with a 2Hz freq. starting from apptimestamp 0, the GT includes the closest timestamp matching the timing provided by competitors.</li> </ol> <p><strong>Please, cite the following works when using the datasets included in this package:</strong></p> <ul> <li>Torres-Sospedra, J.; Jiménez, A. R.; Moreira, A.; Lungenstrass, T.; Lu, W.-C.; Knauth, S.; Mendoza-Silva, G.M.; Seco, F.; Perez-Navarro, A.; Nicolau, M.J.; Costa, A.; Meneses, F.; Farina, J.; Morales, J.P.; Lu, W.-C.; Cheng, H.-T.; Yang, S.-S.; Fang, S.-H.; Chien, Y.-R. and Tsao, Y. Off-line evaluation of mobile-centric Indoor Positioning Systems: the experiences from the 2017 IPIN competition Sensors Vol. 18(2), 2018. <a href="http://dx.doi.org/10.3390/s18020487">http://dx.doi.org/10.3390/s18020487</a></li> <li>Jimenez, A.R.; Mendoza-Silva, G.M.; Seco, F.; Torres-Sospedra, J. Datasets and Supporting Materials for the IPIN 2017 Competition Track 3 (Smartphone-based, off-site). <a href="http://dx.doi.org/10.5281/zenodo.2823924">http://dx.doi.org/10.5281/zenodo.2823924</a> </li> </ul> <p><strong>Additional information can be found at:</strong></p> <ul> <li><a href="http://evaal.aaloa.org/2017/2017-competition-home">http://evaal.aaloa.org/2017/2017-competition-home</a></li> <li><a href="http://indoorloc.uji.es/ipin2017track3/">http://indoorloc.uji.es/ipin2017track3/</a></li> </ul> <p><strong>For any further questions about the database and this competition track, please contact: </strong></p> <ul> <li>Joaquín Torres (<a href="mailto:jtorres@uji.es?subject=IPIN%202016%20Competition%20Dataset%20(Zenodo)">jtorres@uji.es</a>) Institute of New Imaging Technologies, Universitat Jaume I, Spain. </li> <li>Antonio R. Jiménez (<a href="mailto:antonio.jimenez@csic.es?subject=IPIN%202016%20Competition%20Dataset%20(Zenodo)">antonio.jimenez@csic.es</a>) Center of Automation and Robotics (CAR)-CSIC/UPM, Spain. </li> </ul> <p><br> </p>
Datasets and Supporting Materials for the IPIN 2018 Competition Track 3 (Smartphone-based, off-site)
<p>This package contains the datasets and supplementary materials used in the IPIN 2018 Competition (Nantes, France).</p> <p><strong>Contents:</strong></p> <ol> <li>IPIN2018_CallForCompetition_v2.1: Call for competition including the technical annex describing the competition </li> <li>01-Logfiles: This folder contains a subfolder with the 22 training logfiles, a subfolder with the 15 (13 + 2) validation logfiles, and a subfolder with the 1 blind evaluation logfile as provided to competitors.</li> <li>02-Supplementary_Materials: This folder contains the Matlab/octave parser, the raster maps, the vector maps and the visualization of the training routes.</li> <li>03-Evaluation: This folder contains the scripts used to calculate the competition metric, the 75th percentile on the 99 evaluation points. The ground truth is also provided in MatLab format and as a CSV file. Since the results must be provided with a 2Hz freq. starting from apptimestamp 0, the GT includes the closest timestamp matching the timing provided by competitors.</li> <li>03-Evaluation_alternative: This folder contains the alternative scripts used to calculate the competition metric, the 75th percentile on the 99 evaluation points. This version is compatible with MatLab and Octave and does not require any toolbox. In some cases, the differences in the reported errors might be around 10 cm with respect to the script used in the competition. The ground truth is also provided in MatLab format and as a CSV file. Since the results must be provided with a 2Hz freq. starting from apptimestamp 0, the GT includes the closest timestamp matching the timing provided by competitors.</li> </ol> <p><strong>Please, cite the following works when using the datasets included in this package:</strong></p> <ul> <li>Jimenez, A.R.; Mendoza-Silva, G.M.; Ortiz, M.; Perez-Navarro, A.; Perul, J.; Seco, F.; Torres-Sospedra, J. Datasets and Supporting Materials for the IPIN 2018 Competition Track 3 (Smartphone-based, off-site). <a href="http://dx.doi.org/10.5281/zenodo.2823964">http://dx.doi.org/10.5281/zenodo.2823964</a></li> <li>Renaudin, V.; Ortiz, M.; Perul, J.; Torres-Sospedra, J.; Ramón Jimenez, A.; Pérez-Navarro, A.; Martín Mendoza-Silva, G.; Seco, F.; Landau, Y.; Marbel, R.; Ben-Moshe, B.; Zheng, X.; Ye, F.; Kuang, J.; Li, Y.; Niu, X.; Landa, V.; Hacohen, S.; Shvalb, N.; Lu, C.; Uchiyama, H.; Thomas, D.; Shimada, A.; Taniguchi, R.; Ding, Z.; Xu, F.; Kronenwett, N.; Vladimirov, B.; Lee, S.; Cho, E.; Jun, S.; Lee, C.; Park, S.; Lee, Y.; Rew, J.; Park, C.; Jeong, H.; Han, J.; Lee, K.; Zhang, W.; Li, X.; Wei, D.; Zhang, Y.; Park, S. Y.; Park, C. G.; Knauth, S.; Pipelidis, G.; Tsiamitros, N.; Lungenstrass, T.; Pablo Morales, J.; Trogh, J.; Plets, D.; Opiela, M.; Shih-Hau Fang Tsao, Y.; Chien, Y.-R.; Yang, S.-S.; Ye, S.-J.; Ali, M. U.; Hur, S.; and Park, Y. Evaluating Indoor Positioning Systems in a Shopping Mall: The Lessons Learned from the IPIN 2018 Competition IEEE Access Vol. 7, pp. 148594-148628, 2019. http://dx.doi.org/10.1109/ACCESS.2019.2944389</li> </ul> <p><strong>Additional information can be found at:</strong></p> <ul> <li><a href="http://evaal.aaloa.org/2018/call-for-competitions">http://evaal.aaloa.org/2018/call-for-competitions</a></li> <li><a href="http://ipin-conference.org/2018/ipincompetition/">http://ipin-conference.org/2018/ipincompetition/</a></li> </ul> <p><strong>For any further questions about the database and this competition track, please contact: </strong></p> <ul> <li>Joaquín Torres (<a href="mailto:jtorres@uji.es?subject=IPIN%202016%20Competition%20Dataset%20(Zenodo)">jtorres@uji.es</a>) Institute of New Imaging Technologies, Universitat Jaume I, Spain. </li> <li>Antonio R. Jiménez (<a href="mailto:antonio.jimenez@csic.es?subject=IPIN%202016%20Competition%20Dataset%20(Zenodo)">antonio.jimenez@csic.es</a>) Center of Automation and Robotics (CAR)-CSIC/UPM, Spain. </li> </ul>
An EEG dataset for cross-session mental workload estimation: Passive BCI competition of the Neuroergonomics Conference 2021
<p>The dataset is part of a new open EEG database designed to answer a need for more publicly available EEG-based dataset to design and benchmark passive brain-computer interface pipelines (as detailed in [Hinss2021]). This database is currently being created and will be fully released before the end of the year. It will include data acquired over 30 participant, 4 tasks and 3 sessions. For this competition, hosted by the Neuroergonomics Conference 2021, only one task and half the participants will be analyzed. Hence, this competition focuses on a renowned task that elicits various levels of mental/cognitive workload: the Multi-Atribute Task Battery-II (MATB-II) developed by NASA (https://matb.larc.nasa.gov/). It is composed of 4 sub-tasks: system monitoring, tracking, resource management and communications. By varying the number and complexity of the sub-tasks, 3 levels of workload were elicited (verified through statistical analyzes of both subjective and objective -behavioral and cardiac- data). Each difficulty level was performed by 15 subjects (6 female; 9 average 25 y.o.) during 5 minutes per session, in a pseudo-randomized order. Each session was separated by 7 days. We used a 62 actiChamp EEG channels device (BrainProducts; electrode placement 10-20 system).</p> <p> </p> <p><strong>For the competition, your goal is to predict the mental workload for a given subject (intra-subject estimation) using the EEG data from another session (inter-session adaptation). More information on the conference website and in the documentation file.</strong></p>
Datasets used for the manuscript: "Sibling competition, dispersal and fitness outcomes in humans"
<p>Datasets used for the manuscript: “Sibling competition, dispersal and fitness outcomes in humans”, 10.1038/s41598-023-33700-3</p>
Datasets and Supporting Materials for the IPIN 2023 Competition Track 3 (Smartphone-based, off-site)
<p>This package contains the datasets and supplementary materials used in the IPIN 2023 Competition.</p><p><strong>Contents</strong></p><ul><li><i>Track-3_TA-2023.pdf: </i>Technical annexe describing the competition (Version 2)</li><li><i>01 Logfiles: </i>This folder contains a subfolder with the 54 training trials, a subfolder with the 4 testing trials (validation), and a subfolder with the 2 blind scoring trials (test) as provided to competitors.</li><li><i>02 Supplementary_Materials: </i>This folder contains the Matlab/octave parser, the raster maps, the files for the Matlab tools and the trajectory visualization.</li><li><i>03 Evaluation: </i>This folder contains the scripts we used to calculate the competition metric, the 75th percentile on the 69 evaluation points. It requires the Matlab Mapping Toolbox. We also provide the ground truth as 2 CSV files. It contains samples of reported estimations and the corresponding results.</li></ul><p>We provide additional information on the competition at: https://evaal.aaloa.org/2023/call-for-competition</p><p><strong>Citation Policy</strong> </p><p>Please cite the following works when using the datasets included in this package:</p><p><i>Torres-Sospedra, J.; et al. Datasets and Supporting Materials for the IPIN 2023</i><br><i>Competition Track 3 (Smartphone-based, off-site), Zenodo 2023</i><br><i>http://dx.doi.org/10.5281/zenodo.8362205</i></p><p>Check the updated citation policy at: http://dx.doi.org/10.5281/zenodo.8362205</p><p><strong>Contact</strong></p><p>For any further questions about the database and this competition track, please contact: </p><p>Joaquín Torres-Sospedra <br>Centro ALGORITMI,<br>Universidade do Minho, Portugal<br>info@jtorr.es - jtorres@algoritmi.uminho.pt<br> <br>Antonio R. Jiménez <br>Centre of Automation and Robotics (CAR)-CSIC/UPM, Spain <br>antonio.jimenez@csic.es</p><p>Antoni Pérez-Navarro<br>Faculty of Computer Sciences, Multimedia and Telecommunication, Universitat Oberta de Catalunya, Barcelona, Spain<br>aperezn@uoc.edu</p><p><strong>Acknowledgements</strong></p><p>We thank Maximilian Stahlke and Christopher Mutschler at Fraunhofer ISS, as well as Miguel Ortiz and Ziyou Li at Université Gustave Eiffel, for their invaluable support in collecting the datasets. And last but certainly not least, Antonino Crivello and Francesco Potortì for their huge effort in georeferencing the competition venue and evaluation points.</p><p>We extend our appreciation to the staff at the Museum for Industrial Culture (Museum Industriekultur) for their unwavering patience and invaluable support throughout our collection days.</p><p>We are also grateful to Francesco Potortì, the ISTI-CNR team (Paolo, Michele & Filippo), and the Fraunhofer IIS team (Chris, Tobi, Max, ...) for their invaluable commitment to organizing and promoting the IPIN competition.</p><p>This work and competition belong to the IPIN 2023 Conference in Nuremberg (Germany). </p><p>Parts of this work received the financial support received from projects and grants: </p><ul><li>ORIENTATE (H2020-MSCA-IF-2020, Grant Agreement 101023072)</li><li>GeoLibero (from CYTED)</li><li>INDRI (MICINN, ref. PID2021-122642OB-C42, PID2021-122642OB-C43, PID2021-122642OB-C44, MCIU/AEI/FEDER UE)</li><li>MICROCEBUS (MICINN, ref. RTI2018-095168-B-C55, MCIU/AEI/FEDER UE)</li><li>TARSIUS (TIN2015-71564-C4-2-R, MINECO/FEDER)</li><li>SmartLoc(CSIC-PIE Ref.201450E011)</li><li>LORIS (TIN2012-38080-C04-04)</li></ul>
GECCO Industrial Challenge 2017 Dataset: A water quality dataset for the 'Monitoring of drinking-water quality' competition at the Genetic and Evolutionary Computation Conference 2017, Berlin, Germany.
<p>Dataset of the 'Industrial Challenge: Monitoring of drinking-water quality' competition hosted at The Genetic and Evolutionary Computation Conference (GECCO) July 15th-19th 2017, Berlin, Germany</p> <p> </p> <p>The task of the competition was to develop an anomaly detection algorithm for a water- and environmental data set.</p> <p> </p> <p>Included in zenodo: </p> <p>- dataset of water quality data</p> <p>- additional material and descriptions provided for the competition</p> <p> </p> <p>The competition was organized by:</p> <p>M. Friese, J. Stork, A. Fischbach, M. Rebolledo, T. Bartz-Beielstein (TH Köln)</p> <p> </p> <p>The dataset was provided and prepared by:</p> <p>Thüringer Fernwasserversorgung,</p> <p>IMProvT research project (S. Moritz)</p> <p><br> </p> <p>Industrial Challenge: Monitoring of drinking-water quality</p> <p> </p> <p>Description:</p> <p>Water covers 71% of the Earth's surface and is vital to all known forms of life. The provision of safe and clean drinking water to protect public health is a natural aim. Performing regular monitoring of the water-quality is essential to achieve this aim.</p> <p>Goal of the GECCO 2017 Industrial Challenge is to analyze drinking-water data and to develop a highly efficient algorithm that most accurately recognizes diverse kinds of changes in the quality of our drinking-water.</p> <p> </p> <p>Submission deadline:</p> <p>June 30, 2017</p> <p>Official webpage:</p> <p><a href="http://www.spotseven.de/gecco-challenge/gecco-challenge-2017/">http://www.spotseven.de/gecco-challenge/gecco-challenge-2017/</a></p>
Datasets and Supporting Materials for the MALIN-ANR 2019 Competition (French national research agency)
<p>This "ZENODO deposit" provides a multiple sensor dataset collected by the CyborgLOC team during the intermediate competition of the Challenge MALIN (<em>MA</em><em>îtrise</em><em> de la </em><em>L</em><em>ocalisation </em><em>IN</em><em>door</em>), which is a competition for indoor/outdoor real-time positioning. The sensors, including a GNSS receiver Ublox NEO-M8N, a Realsense D435i stereo camera, three Xsens MTi-300 and one PERSY (<strong>PE</strong>destrian <strong>R</strong>eference <strong>SY</strong>stem), are mounted on different parts of the subject’s body. The PERSY is a foot-mounted positioning device with a tri-axial accelerometer, a tri-axial gyroscope, a tri-axial magnetometer as well as a GNSS receiver Ublox M8T. The two scenarios are designed in a training center of firefighters CFIS (Fire and Rescue Training Center) in Blois, France to simulate the situation of firefighters during interventions. With total distances around 2 km for each scenario, the travelled trajectories passed through challenging environments including indoor, outdoor, urban canyon. The indoor part contains different stair levels, from the underground up to the 6th floor. The travel modes are vehicles and pedestrians. Several classical activities of firefighters are realized such as walking, running, stair-climbing, side-walking, crawling, passing above/below obstacles, carrying a stretcher, ladder climbing, etc. High accurate ground truth of stationary points and enclosing volumes are provided by the organizers of the competition, i.e., the French Ministry of Defense (DGA: Direction Générale de l’Armement). Provided with raw data, they allow the evaluation of the positioning performances.</p> <p>To facilitate the use of our dataset under Rosbag format, a toolkit of python scripts named <em>MALIN Data Processing Tools</em> is provided on GitHub (<a href="https://github.com/4g-group/malin_data_processing_tools">https://github.com/4g-group/malin_data_processing_tools</a>). It allows merging Rosbags, converting Rosbag files to CSV files as well as republishing camera’s topics as decompressed data. Details about these processing tools could be found in the Readme file on the Github page. </p>
GECCO Industrial Challenge 2019 Dataset: A water quality dataset for the 'Internet of Things: Online Event Detection for Drinking Water Quality Control' competition at the Genetic and Evolutionary Computation Conference 2019, Prague, Czech Republic.
<p>Dataset of the 'Internet of Things: Online Event Detection for Drinking Water Quality Control' competition hosted at The Genetic and Evolutionary Computation Conference (GECCO) July 13th-17th 2019, Prague, Czech Republic</p> <p> </p> <p>The task of the competition was to develop an anomaly detection algorithm for a water- and environmental data set.</p> <p> </p> <p>Included in zenodo: </p> <p>1. Original train dataset of water quality data provided to participants (identical to gecco2019_train_water_quality.csv)</p> <p>2. Call for Participation</p> <p>3. Rules and Description of the Challenge</p> <p>4. Resource Package provided to participants</p> <p>5. The complete dataset, consisting of train, test and validation merged together (gecco2019_all_water_quality.csv)</p> <p>6. The test dataset, which was used for creating the leaderboard on the server (gecco2019_test_water_quality.csv)</p> <p>7. The train dataset, which participants had available for training their models (gecco2019_train_water_quality.csv)</p> <p>8. The validation dataset, which was used for the end results for the challenge (gecco2019_valid_water_quality.csv)</p> <p> </p> <p>The challenge required the participants to submit a program for event detection. A training dataset was available to the participants (gecco2019_train_water_quality.csv). During the challenge the participants were able to upload a version of their program to out online platform, where this version was scored against the testing dataset (gecco2019_test_water_quality.csv), thus an intermediate leaderboard was available. To avoid overfitting against this dataset, at the end of the challenge, the end result was created from scoring with the validation dataset (gecco2019_valid_water_quality.csv). </p> <p>Train, Test, Validation dataset are from the same measuring station and are in chronological order. So the timestamps from the test dataset begin directly after the train timestamps, while the validation timestamps begin directly after the test timestamps. </p> <p> </p> <p>The competition was organized by:</p> <p>F. Rehbach, S. Moritz, T. Bartz-Beielstein (TH Köln)</p> <p> </p> <p>The dataset was provided by:</p> <p>Thüringer Fernwasserversorgung and IMProvT research project</p> <p> </p> <p> </p> <p>Internet of Things: Online Event Detection for Drinking Water Quality Control</p> <p> </p> <p>Description:</p> <p>For the 8th time in GECCO history, the SPOTSeven Lab is hosting an industrial challenge in cooperation with various industry partners. This years challenge, based on the 2018 challenge, is held in cooperation with "Thüringer Fernwasserversorgung" which provides their real-world data set. The task of this years competition is to develop an anomaly detection algorithm for the water- and environmental data set. Early identification of anomalies in water quality data is a challenging task. It is important to identify true undesirable variations in the water quality. At the same time, false alarm rates have to be very low.</p> <p><br> Competition Opens: End of January/Start of February 2019<br> Final Submission: 30 June 2019</p> <p>Official webpage:</p> <p><a href="https://www.th-koeln.de/informatik-und-ingenieurwissenschaften/gecco-challenge-2019_63244.php">https://www.th-koeln.de/informatik-und-ingenieurwissenschaften/gecco-challenge-2019_63244.php</a></p> <p> </p>
Dataset and Jupyter worksheet interpreting the (results from) small- and wide-angle scattering data from a series of boehmite/epoxy nanocomposites. Accompanies the publication "Competition of nanoparticle-induced mobilization and immobilization effects on segmental dynamics of an epoxy-based nanocomposite"
<p>Dataset and Jupyter worksheet interpreting the (results from) small- and wide-angle scattering data from a series of boehmite/epoxy nanocomposites. Accompanies the publication "Competition of nanoparticle-induced mobilization and immobilization effects on segmental dynamics of an epoxy-based nanocomposite", by Paulina Szymoniak, Brian R. Pauw, Xintong Qu, and Andreas Schönhals.</p> <p>Datasets are in three-column ascii (processed and azimuthally averaged data) from a Xenocs NanoInXider SW instrument. Monte-Carlo analyses were performed using McSAS 1.3.1, other analyses are in the Python 3.7 worksheet. Graphics and result tables are output by the worksheet. </p>
ICDAR 2015 Competition HTRtS: Handwritten Text Recognition on the tranScriptorium Dataset
<p>This dataset comprises the dataset used for the ICDAR 2015 Competition on Handwritten Text Recognition on the tranScriptorium Dataset. The handwritten images for this contest were drawn from the English “Bentham collection” dataset used in the TRAN SCRIPTORIUM project. The selected data has been written by several hands and entails significant variabilities and difficulties regarding the quality of text images, writing styles and crossed-out text. This contest is clearly more difficult than the the first edition both for training and for testing. A portion of the training dataset and the full test dataset were provided in the form of carefully segmented line images, along with the corresponding transcripts. Another portion of the training dataset was provided as raw images and their corresponding transcripts at region level.<br> </p> <p>ICDAR 2015 competition HTRtS: handwritten text recognition on the tranScriptorium dataset<br> JA Sánchez, AH Toselli, V Romero, E Vidal. In International Conference on Document Analysis and Recognition (ICDAR), pp. 1166-1170, 2015.</p>
Train-B dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR). Batch 1 and Batch 2.
<p>Train-B Dataset. Dataset of pages without any layout or text line information. The corresponding transcripts are provided at page level with line breaks. It has 10k pages, though for convenience it is divided into two 5k page batches. This information is provided in PAGE format. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439807#.WOIBZ3WLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p> <p> </p>
Train-A dataset for ICDAR2017 Competition on Handwritten Text Recognition on the READ Dataset (ICDAR2017 HTR)
<p>Train-A Dataset of pages with manually revised baselines and the corresponding transcripts associated to them. This batch is small, 50 pages. Please, keep in mind that only the baselines have been manually corrected, The polygons associated to each line have not been manually reviewed. </p> <p>This dataset is complementary to this other dataset:</p> <p>https://zenodo.org/record/439811#.WOIF9HWLSkA</p> <p>More information at:</p> <p>https://scriptnet.iit.demokritos.gr/competitions/~icdar2017htr/</p>
Dataset supplementing Marx, S., Gruenhage, G., Walper, D., Rutishauser, U., Einhäuser, W. (2015). Competition with and without priority control: linking rivalry to attention through winner-take-all networks with memory. Annals of the New York Academy of Sciences. 1339, 138-153.
<p>Data supplementing the paper Marx, S., Gruenhage, G., Walper, D., Rutishauser, U., Einhäuser, W. (2015). Competition with and without priority control: linking rivalry to attention through winner-take-all networks with memory. <em>Annals of the New York Academy of Sciences. 1339, </em>138-153. doi: 10.1111/nyas.12575 The files can be freely used for scientific purposes, provided this reference is appropriately cited.</p> <p>Files contain the behavioral data, the model can be found at https://doi.org/10.5281/zenodo.573026</p> <p> </p> <p>The following files are contained in this folder:</p> <p>dataExp1.mat contains the data of experiment 1</p> <p>The variables durationLeft and durationRight contain 5 x 6 x 6 cell arrays with the dominance durations for the left and right grating, respectively. Dimensions are subject x contrast level left x contrast level right.</p> <p><br> dataExp2.mat contains the data of experiment 2</p> <p>Variables buttonStart, buttonEnd and whichButton contain 3x4x5 (contrast levels x blank duration levels x subjects) cell arrays that contain the start time and end time of each button press, and which button (1/2) was pressed, respectively.</p> <p>Variables presStart and presEnd contain 3x4x5 (contrast levels x blank duration levels x subjects) cell arrays that contain start and end of each blank period. All time stamps refer to the onset of the first blanking trial (end of continuous presentation)</p> <p>Variable prevPerz contains the percept (button) that was pressed at the end of the continuous presentation period.</p> <p><br> figure3_human.m, figure4_human.m and figure6_human.m exemplify the usage of the data by re-plotting the figures containing human data of the aforementioned paper</p>
EmoPairCompete - Physiological Signals Dataset for Emotion and Frustration Assessment under Team and Competitive Behaviours
<p>Please refer to the documentation at: https://github.com/DTUComputeStatisticsAndDataAnalysis/EmoPairCompete</p>
Microsoft Indoor Localization Competition 2018 Dataset
<p>Detailed ground truth measurements and error visualization for each team, as well as the 3D point cloud of the evaluation area related to the Microsoft Indoor Localization Competition 2018.</p> <p>Additional details can be found here:</p> <p>https://www.microsoft.com/en-us/research/event/microsoft-indoor-localization-competition-ipsn-2018/</p>
Datasets and Supporting Materials for the IPIN 2018 Competition Track 4 (Foot-Mounted IMU based Positioning, off-site)
<p>This package contains the datasets and supplementary materials used in the IPIN 2018 Competition (Nantes, France).</p> <p><strong>Contents:</strong></p> <ol> <li>IPIN2018_CallForCompetition_v2.1: Call for competition including the technical annex describing the competition </li> <li>01-Logfiles: This folder contains 2 zip files.<br> - HKB08.zip : for sensors bias estimation.<br> - HKB82.zip : for trajectory estimation.<br> Each archive contains 4 files :<br> - HKBxx_mag.csv : magnetometer data<br> - HKBxx_sti.csv : inertial data<br> - HKBxx_ublox.ubx : GNSS data<br> - HKBxx_INFO.txt : info file<br> see page 16 of IPIN2018_CallForCompetition_v2.1.pdf for more details.</li> <li>02-Supplementary_Materials: This folder contains the datasheet files of the different sensors.</li> <li>03-Evaluation: This folder contains the scripts used to calculate the competition metric, the 75th percentile on all evaluation points. The ground truth is provided csv file.</li> </ol> <p><strong>Please, cite the following works when using the datasets included in this package:</strong></p> <ul> <li>Ortiz, M.; Perul, J.; Torres-Sospedra, J. Renaudin, V. Datasets and Supporting Materials for the IPIN 2018 Competition Track 4 (Foot-Mounted IMU based Positioning, off-site), Zenodo 2018 <a href="http://dx.doi.org/10.5281/zenodo.3228012">http://dx.doi.org/10.5281/zenodo.3228012</a></li> </ul> <p><strong>Additional information can be found at:</strong></p> <ul> <li><a href="http://evaal.aaloa.org/2018/call-for-competitions">http://evaal.aaloa.org/2018/call-for-competitions</a></li> <li><a href="http://ipin-conference.org/2018/ipincompetition/">http://ipin-conference.org/2018/ipincompetition/</a></li> </ul> <p><strong>For any further questions about the database and this competition track, please contact to: </strong></p> <ul> <li> <p>Miguel Ortiz (<a href="mailto:miguel.ortiz@ifsttar.fr">miguel.ortiz@ifsttar.fr</a>) at the French Institute of Science and Technology for Transport, Development and Networks (IFSTTAR) France.</p> <p> </p> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.