Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,549
datasets available to search
ShareScore release 0.9.0
Dataset results
1,549 results for “benchmarks”
PaRoutes: a framework for benchmarking retrosynthesis route predictions
<p>PaRoutes is a framework for benchmarking multi-step retrosynthesis methods, i.e. route predictions.</p> <p>It provides:</p> <ul> <li>A curated reaction dataset for building one-step retrosynthesis models</li> <li>Two sets of 10,000 routes</li> <li>Two sets of stock molecules to use as stop-criterion for the search</li> </ul> <p>Homepage: <a href="https://github.com/MolecularAI/PaRoutes">https://github.com/MolecularAI/PaRoutes</a></p>
HornMT – Machine Translation Benchmark Dataset for Languages in the Horn of Africa
<p>The <strong>HornMT</strong> repository contains data and the associated metadata for the project <a href="https://lesan.ai/benchmark">Machine Translation Benchmark Dataset for Languages in the Horn of Africa</a>. It is a multi-way parallel corpus that will serve as a benchmark to accelerate progress in machine translation research and production systems for languages in the Horn of Africa.</p> <p>Supported Languages</p> <table> <tbody> <tr> <td> <p>Language</p> </td> <td> <p>ISO 639-3 code</p> </td> </tr> </tbody> <tbody> <tr> <td> <p>Afar</p> </td> <td> <p>aaf</p> </td> </tr> <tr> <td> <p>Amharic</p> </td> <td> <p>amh</p> </td> </tr> <tr> <td> <p>English</p> </td> <td> <p>eng</p> </td> </tr> <tr> <td> <p>Oromo</p> </td> <td> <p>orm</p> </td> </tr> <tr> <td> <p>Somali</p> </td> <td> <p>som</p> </td> </tr> <tr> <td> <p>Tigrinya</p> </td> <td> <p>tir</p> </td> </tr> </tbody> </table> <p><strong> </strong></p> <p>data/ contains one text file per language and each file contains news snippets in the same order for each language.</p> <p>data<br> ├── aar.txt<br> ├── amh.txt<br> ├── eng.txt<br> ├── orm.txt<br> ├── som.txt<br> └── tir.txt</p> <p>metadata.tsv contains tab separated data describing each news snippet. The metadata contains the following fields.</p> <ul> <li> <p><strong>Scope</strong> - describes whether the news is global or local. It takes two values: Global news and Local news.</p> </li> <li> <p><strong>Category</strong> - News category covering the following 12 topics</p> <ul> <li> <p>Art and Culture</p> </li> <li> <p>Business and Economy</p> </li> <li> <p>Conflicts and Attacks</p> </li> <li> <p>Disaster and Accidents</p> </li> <li> <p>Entertainment</p> </li> <li> <p>Environment</p> </li> <li> <p>Health</p> </li> <li> <p>International Relations</p> </li> <li> <p>Law and Crime</p> </li> <li> <p>Politics</p> </li> <li> <p>Science and Technology</p> </li> <li> <p>Sport</p> </li> </ul> </li> <li> <p><strong>Source</strong> - List of one or more URLs from which the news content is extracted or based on.</p> </li> <li> <p><strong>Domain</strong> - TLD corresponding to the URL(s) in Source.</p> </li> <li> <p><strong>Date</strong> - The publication date of the source article. The format is yyyy-mm-dd.</p> </li> </ul> <p>Other formats</p> <p>All the data and associated metadata together in one file is also available in other file formats.</p> <p><strong>HornMT.xlsx</strong> - data and associated metadata in xlsx format.</p> <p><strong>HornMT.json</strong> - data and associated metadata in json format.</p> <p>Below is an example row.</p> <pre><code class="language-javascript">{ "data":{ "eng":"The World Meteorological Organisation reports that the ozone layer is damaged to its worst extent ever in the Arctic.", "aaf":"Baad Metrolojih Eglali Areketekeh Addal Ozonih qelu faxe waktik lafetle calat biyakisem xayose.", "amh":"የአለም የአየር ንብረት ድርጅት በአርክቲክ አካባቢ ያለው የኦዞን ምንጣፍ ከፍተኛ ጉዳት እንደደረሰበት አስታወቀ፡፡", "orm":"Dhaabbanni Meetiroolojii Addunyaa baqqaanni oozonii Arkiitik keessatti gara sadarkaa isa hamaa haga ammaatti akka miidhame gabaase.", "som":"Ururka Saadaasha Hawada Adduunka ayaa ku warramaya in lakabka ozoneka ee Ka koreeya dhulka baraflayda uu waxyeelladii abid ugu darnaa soo gaadhay.", "tir":"ውድብ ሜትሮሎጂ ዓለም ኣብ ኣርክቲክ ዝርከብ ናሕሲ ኦዞን ኣዝዩ ብዝኸፍአ ደረጃ ከምዝተጎድአ ሓቢሩ፡፡" }, "metadata":{ "scope":"Global", "category":"Science and Technology", "source":"https://www.independent.co.uk/environment/climate-change/ozone-layer-damaged-by-unusually-harsh-winter-2263653.html", "domain":"www.independent.co.uk", "date":"2011-04-05" } }</code></pre> <p><strong>Team</strong></p> <p>Afar</p> <ul> <li> <p>Mohammed Deresa</p> </li> <li> <p>Yasin Nur</p> </li> </ul> <p>Amharic</p> <ul> <li> <p>Tigist Taye</p> </li> <li> <p>Selamawit Hailemariam</p> </li> <li> <p>Wako Tilahun</p> </li> </ul> <p>Oromo</p> <ul> <li> <p>Gemechis Melkamu</p> </li> <li> <p>Galata Girmaye</p> </li> </ul> <p>Somali</p> <ul> <li> <p>Abdiselam Mohamed</p> </li> <li> <p>Beshir Abdi</p> </li> </ul> <p>Tigrinya</p> <ul> <li> <p>Berhanu Abadi Weldegiorgis</p> </li> <li> <p>Michael Minassie</p> </li> <li> <p>Nureddin Mohammedshiek</p> </li> </ul> <p><strong>Project Leaders</strong></p> <ul> <li> <p>Asmelash Teka Hadgu <a href="mailto:asme@lesan.ai">asme@lesan.ai</a></p> </li> <li> <p>Gebrekirstos G. Gebremeskel <a href="mailto:gebrekirstos.gebremeskel@ru.nl">gebrekirstos.gebremeskel@ru.nl</a></p> </li> <li> <p>Abel Aregawi <a href="mailto:abel@lesan.ai">abel@lesan.ai</a></p> </li> </ul> <p><strong>License</strong></p> <p>Shield: <a href="http://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a></p> <p>This work is licensed under a<br> <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p>
A benchmark dataset of herbarium specimen images with label data: Summary
<p>This landing page contains a CSV file compiling all data associated with herbarium specimens that are part of this dataset, as they could be found on GBIF, JACQ or FinBIF. A CSV file with and without Darwin Core extension data is available, as some CSV readers have trouble with the JSON format that is used for those extensions.</p> <p>In addition, DOI's of the individual specimens uploaded to Zenodo and direct links to the different files (JPEG, TIFF, JSON, PNG) are also included. Index of these added variables:</p> <p>- persistentID: Persistent Identifier of the collection specimen. Data uploaded as part of this dataset will not be kept in sync with changes at the collection's repository. Hence, this URI will always point to the most up to date information known about the herbarium specimen.</p> <p>- jpegURL, tiffURL, jsonURL: URL's pointing straight to the respective image and data files themselves, to facilitate (selective) batch downloads.</p> <p>- pngSegAllURL and pngSegSelURL: Segmented overlays of the herbarium specimens indicating the location of different labels and reference material on the sheet ("All") and their content ("Sel"). More information can be found in the paper (in prep) associated with this data publication and the individual depositions themselves.</p> <p>- DOI: The DOI of the deposition of images and data of these specimens on Zenodo. DOI's point to the most up-to-date version of these depositions at the time of the publication of this CSV file. As a rule, this CSV file will be updated should any changes happen to any of the depositions.</p> <p>- jpegURL2, tiffURL2: A few herbarium sheets had labels on the back and consisted therefore of two scans. As a rule, the label scans are in this category.</p>
Motion Capture Benchmark of Industrial Tasks for Ergonomic Assessment and European Historic Crafts
<p><strong>General Info:</strong></p> <p>This benchmark provides motion capture (MoCap) files in .bvh form. The recordings were done in the span of May 2019 to January 2020 for the needs of the <a href="https://collaborate-project.eu/"><strong>CoLLaboratE</strong></a> and <a href="http://www.mingei-project.eu/"><strong>MINGEI</strong></a> H2020 projects<strong> </strong>funded by the European Commission. The tasks included are:</p> <ul> <li>TV assembling</li> <li>Airplane component manufacturing</li> <li>High ergonomic hazard motions </li> <li>Silk-Weaving</li> <li>Glassblowing</li> <li>Mastic Cultivation</li> </ul> <p>The TV assembly and airplane component manufacturing tasks were recorded in real-world conditions inside the factory during the actual production of the items. The high ergonomic hazard motions were recorded in a controlled lab environment and serve as baseline/prototype motions for ergonomic risk assessment.</p> <p>The silk-weaving, glassblowing, and mastic cultivation data sets were created, corresponding to movements performed by skilled craftsmen and mastic farmers. These data sets were produced in order to extract the expert's gestural knowledge and analyze their dexterity while doing their crafts.</p> <p><strong>Naming Convention:</strong></p> <p>All files in this benchmark follow a strict naming convention to allow for easier parsing by scripts. The names have a total of 12 or 13 characters that convey the following information:</p> <ul> <li>The first three or fours characters label the <strong>recording session </strong>(e.g., LAB, PLN, GBBC, MCSN, etc.)</li> <li>The next three characters label the <strong>subject number </strong>(e.g., S01, S02, S03, etc.)</li> <li>The next three characters label the <strong>posture or gesture number </strong>(e.g., P01, P02, G01, G02, etc.)</li> <li>The final three characters label the <strong>repetition number </strong>(e.g., R01, R02, R03, etc.)</li> </ul> <p>For example, LABS02P03R01 denotes a lab recording of the second subject, performing the third posture for the first time.</p> <p><strong>Recording Sessions:</strong></p> <p>There are six recording sessions in this benchmark, the ergonomic risk motion recorded in the lab (denoted as "<strong>LAB</strong>"), the construction of an airplane component (denoted as "<strong>PLN</strong>"), and the assembling and packaging of TVs (denoted as "<strong>TV*</strong>"), the silk weaving (denoted as "<strong>SW*</strong>"), glassblowing (denoted as "<strong>GB*</strong>"), and mastic cultivation (denoted as "<strong>MC*</strong>").</p> <p>The postures are the following:</p> <p><strong>LAB:</strong></p> <ul> <li><strong>Standing:</strong> <ul> <li><strong>P01</strong>: The subject stays in I-pose</li> <li><strong>P02:</strong> The subject rotates his/her torso to the left as far the person can</li> <li><strong>P03: </strong>The subject will laterally bend his/her torso to the left for 6 seconds</li> <li><strong>P04</strong>: The subject bends more than 20° but less than 60°</li> <li><strong>P05:</strong> The subject bends more than 20° but less than 60° while rotating and laterally bending the torso to the left</li> <li><strong>P06: </strong>The subject stretches his/her arms, and bends forward more than 20° but less than 60° while rotating and laterally bending the torso to the left</li> <li><strong>P07</strong>: The subject bends more than 60°</li> <li><strong>P08:</strong> The subject bends more than 60° while rotating and laterally bending the torso to the left</li> <li><strong>P09: </strong>The subject stretches his/her arms, and bends forward more than 60° while rotating and laterally bending the torso to the left</li> <li><strong>P10:</strong> The subject upright, raises the elbows above the shoulder level with the forearms bent 90° (</li> <li><strong>P11</strong>: The subject raises the elbows above the shoulder level with the forearms bent 90° while rotating and laterally bending the torso to the left</li> <li><strong>P12:</strong> The subject raises the elbows above the shoulder level with the arms stretched while rotating and laterally bending the torso to the left</li> <li><strong>P13:</strong> The subject upright, raises the hands above the head</li> <li><strong>P14: </strong>The subject raises the hands above the head with the arms stretched while rotating and laterally bending the torso to the left</li> </ul> </li> <li><strong>Sitting on a chair:</strong> <ul> <li><strong>P15: </strong>The subject sits upright</li> <li><strong>P16:</strong> The subject bends forward more than 60°</li> <li><strong>P17:</strong> The subject bends forward more than 60° while rotating and laterally bending the torso to the left</li> <li><strong>P18:</strong> The subject stretches the arms, and bends forward more than 60° while rotating and laterally bending the torso to the left</li> <li><strong>P19:</strong> The subject raises the hands above the head with arms stretched</li> <li><strong>P20:</strong> The subject raises the hands above the head with the arms stretched while rotating and laterally bending the torso to the left</li> </ul> </li> <li><strong>Kneeling:</strong> <ul> <li><strong>P21:</strong> The subject stays upright</li> <li><strong>P22:</strong> The subject rotates the torso to the left as far he/she can</li> <li><strong>P23: </strong>The subject will laterally bend the torso to the left for 6 seconds</li> <li><strong>P24:</strong> The subject bends more than 60°</li> <li><strong>P25:</strong> The subject bends more than 60° while rotating and laterally bending the torso to the left</li> <li><strong>P26:</strong> The subject stretches the arms, and bends forward more than 60° while rotating and laterally bending the torso to the left</li> <li><strong>P27: </strong>The subject upright, raises the elbows to the shoulder level with the arms stretched</li> <li><strong>P28:</strong> The subject raises the elbows to the shoulder level with the arms stretched while rotating and laterally bending the torso to the left</li> </ul> </li> </ul> <p>The TV assembling tasks are further divided. The subtasks are: packing the TVs on a stack for shipping (denoted as "<strong>TVP</strong>" for medium-sized TVs and "<strong>TVL</strong>" for larger TVs), placing assembling and placing electronic circuit boards on the chassis (denoted as "<strong>TVB</strong>"), and screwing the boards on the TV chassis (denoted as "<strong>TV_</strong>"). Each task is comprised of a number of postures. </p> <p><strong>TV Assembling:</strong></p> <ul> <li><strong>Assembling the board and placing it on the TV chassis (TVB):</strong> <ul> <li><strong>P01: </strong>Reaching high, above the shoulder level, to pick one component</li> <li><strong>P02: </strong>Reaching low, below the knee level, to pick up the second component</li> <li><strong>P03: </strong>Connecting the components and placing the board on the chassis to be screwed</li> </ul> </li> <li><strong>Screwing an electrical circuit board on the TV chassis (TV_) :</strong> <ul> <li><strong>P01: </strong>A screw is placed on a power tool and it is being screwed on the chassis. The process is repeated four times</li> </ul> </li> <li><strong>Preparing TVs for Shipping (TVP & TVL):</strong> <ul> <li><strong>P01: </strong>Placing TVs on a wooden pallet (bottom level)</li> <li><strong>P02:</strong> Preparing to wrap the bottom level with a membrane</li> <li><strong>P03:</strong> Wrapping the bottom level</li> <li><strong>P04:</strong> Placing TVs on top of the bottom level (second level)</li> <li><strong>P05:</strong> Placing TVs on top of the second level (third level)</li> <li><strong>P06: </strong>Wrapping the second level with a plastic membrane</li> <li><strong>P07:</strong> Wrapping the third level with a plastic membrane</li> <li><strong>P08:</strong> Placing TVs on top of the third level (fourth level)</li> <li><strong>P09:</strong> Wrapping the fourth level with a plastic membrane</li> </ul> </li> </ul> <p><strong>Riveting of an airplane floater (PLN):</strong></p> <ul> <li><strong>P01:</strong> Rivet with the pneumatic hammer.</li> <li><strong>P02:</strong> Prepare the pneumatic hammer and grab rivets. </li> <li><strong>P03:</strong> Place the bucking bar to counteract the incoming rivet.</li> </ul> <p>The tasks recorded for silk weaving, glassblowing, and mastic cultivation data sets were segmented by gestures (e.g., G01, G02, etc.) . The tasks recorded for these three data sets are the following:</p> <p><strong>Silk weaving (SW*):</strong></p> <ul> <li>The creation of the punch cards <strong>(SWPC)</strong>.</li> <li>Preparation of the beam <strong>(SWPB)</strong>.</li> <li>Wrapping of the beam <strong>(SWWB)</strong>.</li> <li>Jacquard weaving with small loom <strong>(SWSL)</strong>.</li> <li>Jacquard weaving with medium size loom <strong>(SWML)</strong>.</li> <li>Jacquard weaving with large loom <strong>(SWLL)</strong>.</li> </ul> <p><strong>Glassblowing (GB*):</strong></p> <ul> <li>Beak cutting <strong>(GBBC)</strong>.</li> <li>Blowing and shaping <strong>(GBBS)</strong>.</li> <li>Cervix refining <strong>(GBCR)</strong>.</li> <li>Cord laying <strong>(GBCL)</strong>.</li> <li>Finish details <strong>(GBFD)</strong>.</li> <li>Handle laying <strong>(GBHL)</strong>.</li> <li>Transfer to punty <strong>(GBTP)</strong>.</li> <li>Leg and foot laying <strong>(GBLF)</strong>.</li> </ul> <p><strong>Mastic Cultivation (MC*):</strong></p> <ul> <li>Scrapping with new tool <strong>(MCSN)</strong>.</li> <li>Scrapping with old tool <strong>(MCSO)</strong>.</li> <li>Sweeping <strong>(MCSW)</strong>.</li> <li>Dusting <strong>(MCDU)</strong>.</li> <li>Embroidery A <strong>(MCEA)</strong>.</li> <li>Embroidery B <strong>(MCEB)</strong>.</li> <li>Embroidery with an axe <strong>(MCEX)</strong>.</li> <li>Gathering <strong>(MCGA)</strong>.</li> <li>Harvesting <strong>(MCHA)</strong>.</li> <li>Wiping <strong>(MCWI)</strong>.</li> <li>Shifting A <strong>(MCSA)</strong>.</li> <li>Shifting B <strong>(MCSB)</strong>.</li> <li>Cleaning with the wind <strong>(MCCW).</strong></li> </ul> <p>The motion capture files were processed and segmented with a 3D character animation software (MotionBuilder, Autodesk Inc., San Rafael, CA. USA) and exported to Biovision Hierarchy (BVH) files.</p>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Annotation V2.0)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is a new version of the Annotation of MELA dataset, in which we add a missing annotation for 'mela_0732'. This file includes the annotations of the whole training set and validation set. </p> <p>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</p> <p> `public_id: anonymous patient ID to match images and annotations.<br> `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>
Transprecision Computing Benchmarks
<p>The datasets have been collected by benchmarking three algorithms for Transprecision Computing (Correlation, Convolution, Saxpy), on three different hardware platforms (pc, vm, g100).</p> <p>Transprecision Computing<sup>1</sup> is a paradigm that allows users to trade the energy associated with computation in exchange for a reduction in the quality of the computation results. In this complex domain, a typical target are Floating-point (FP) operations: transprecision techniques allow to specify the number of bits used to represent FP variables, and using a smaller number of bits decreases the precision, thus saving energy. To analytically calculate the impact of varying the number of bits on the computation results for programs with more than a couple of instructions is a crucial point. However, this relationship can be learned from data.</p> <p>The provided benchmarks have been used for training several machine learning models, to predict the performance (time, error, memory) of a given algorithm, when running with a particular configuration (the precision assigned to each variable) on a certain hardware architecture. Afterward, the produced models have been embedded into HADA, an optimization engine for hardware dimensioning and algorithm configuration, developed by the AI research group at the University of Bologna, as partner of the EU Horizon 2020 Project StairwAI (g.a. 101017142).</p> <p> </p> <p><strong>Bibliography</strong></p> <p>1. Andrea Borghesi, Giuseppe Tagliavini, Michele Lombardi, Luca Benini and Michela Milano. 2020. Combining learning and optimization for transprecision computing. In <em>Proceedings of the 17th ACM International Conference on Computing Frontiers </em>(<em>CF '20</em>). Association for Computing Machinery, New York, NY, USA, 10–18. https://doi.org/10.1145/3387902.3392615</p>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Validation Set and Annotation)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is the Validation Set and Annotation of MELA dataset, including 110 CTs and the annotations of the whole training set and validation set. Files include:</p> <ol> <li>Val.zip: 110 CTs in NII format (nii.gz).</li> <li>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</li> </ol> <p> `public_id: anonymous patient ID to match images and annotations.<br> `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>
Supplemental Material for a Systematic Literature Review on Benchmarks for Evaluating Debugging Approaches
<p>Bug benchmarks are used in development and evaluation of debugging approaches. Quantitative performance comparison of different debugging approaches is only possible when they have been evaluated on the same dataset or benchmark. However, benchmarks are often specialized towards usage for certain debugging approaches in their contained data, metrics, and artifacts. Such benchmarks can not be easily used on debugging approaches outside their scope as such approach may rely on specific data such as bug reports or code metrics not included in the dataset. Furthermore, benchmarks vary in their size w.r.t. the number of subject programs and the size of the individual subject programs. For these reasons, we have performed a systematic literature review where we have identified 73 benchmarks that can be used to evaluate debugging approaches.</p> <p>We compare the different benchmarks with respect to their size and the provided information such as bug reports, contained test cases, and other code metrics. Furthermore, we have investigated how well the benchmarks realize the <a href="https://www.go-fair.org/fair-principles/">FAIR guiding principles</a>. This comparison is intended to help researchers to quickly identify all suitable benchmarks for evaluating their specific debugging approaches. More information can be found in the publication:</p> <blockquote> <p>Thomas Hirsch and Birgit Hofer: "A Systematic Literature Review on Benchmarks for Evaluating Debugging Approaches", Journal of Systems and Software, in press, 2022.</p> </blockquote>
Benchmark map of deforestation for Amazonia
<p><strong>Title: </strong>Benchmark map of deforestation for Amazonia</p> <p><strong>Contact:</strong> Celso H. L. Silva-Junior (celsohlsj@gmail.com)</p> <p><strong>Data:</strong> Old-growth forest deforestation</p> <p><strong>Coverage:</strong> Amazonia</p> <p><strong>Period:</strong> 1986 to onwards (according to new MapBiomas project collections)</p> <p><strong>Spatial resolution:</strong> 30-meters</p> <p><strong>Temporal resolution:</strong> Annual</p> <p><strong>Coordinate reference system:</strong> Geographic Coordinate System with Datum WGS-84</p> <p><strong>File format: </strong>The zip file contains nine tiles in compressed TIFF format. The pixel values represent the year of deforestation.</p> <p><strong>Code:</strong> <a href="https://github.com/celsohlsj/amazonia_deforestation">https://github.com/celsohlsj/amazonia_deforestation</a></p> <p><strong>Dataset usage</strong>: It is free to use, but if you use this dataset in your work, please cite the repository and our paper correctly. We also welcome users to invite us for collaboration.</p> <p><strong>Associated publication: </strong>Silveira, M.V. F., Silva-Junior, C.H.L., Anderson, L.O., Aragão, L.E.O.C. Amazon fires in the 21st century: the year of 2020 in evidence. Global Ecology and Biogeography (2022). https://doi.org/10.1111/geb.13577</p> <p><strong>For this dataset, please use the following references:</strong></p> <ul> <li>Silveira, M.V. F., et al. Amazon fires in the 21st century: the year of 2020 in evidence. <em>Global Ecology and Biogeography</em> (2022). doi: 10.1111/geb.13577</li> <li>Silva-Junior, C. H. L. . (2022). Benchmark maps of deforestation for Amazonia [Data set]. <em>Zenodo</em>. https://doi.org/10.5281/zenodo.6808579</li> </ul>
The benchmark datasets for Multi-class Change Detection (MCD)
<p>Change detection (CD) provides a research basis for environmental monitoring, urban expansion and reconstruction as well as disaster assessment, by identifying the changes of ground objects in different time periods. Traditional CD focused on the binary change detection (BCD), focusing solely on the change and no-change regions. Due to the dynamic progress of earth observation satellite techniques, the spatial resolution of remote sensing images continues to increase, multi-class change detection (MCD) which can reflect more detailed land change has become a hot research direction in the field of CD.<strong> </strong></p> <p>We have collected the current open source benchmark datasets in the MCD of remote sensing imagery , in order to facilitate the sharing of the latest research datasets in the MCD field. Users can access the relevant MCD datasets through the links in the files.</p> <p>Source:</p> <p>Q. Zhu, X. Guo, Ziqi Li, D. Li*, “A review of Multi-class Change Detection for Remote Sensing Imagery” Geo-spatial information science, 2022</p>
TriCera Benchmarks: SMT-LIB Encodings of SV-COMP 2022 Benchmarks by TriCera
<p>This repository contains the SMT-LIB v2.6 encodings of a subset of SV-COMP 2022 benchmarks from the <em>reach-safety</em> and the <em>memsafety</em> categories, encoded by <a href="https://github.com/uuverifiers/tricera">TriCera</a> and <a href="https://github.com/zafer-esen/heap2array">heap2array</a> [1, 2]. The source benchmarks from SV-Comp 2022 (`.c` and `.i` files together with their accompanying `.yml` files) are not provided due to licensing reasons; the files are freely available for download at <a href="https://zenodo.org/record/5831003/export/hx">https://zenodo.org/record/5831003/export/hx</a></p> <p>- `benchmarks/heap` contains all heap benchmarks (encoded by TriCera using the <a href="https://arxiv.org/abs/2104.04224">theory of heaps</a>, used in [1] and [2]).<br> - `benchmarks/heap2array` contains all heap2array encoded benchmarks (encoded from heap benchmarks using heap2array, used in [1] and [2]).<br> - `benchmarks/nonHeapBms` contains all non-heap benchmarks (encoded by TriCera, used in [2]).</p> <p>[1] (to appear) Zafer Esen and Philipp Rümmer, "An SMT-LIB Theory of Heaps", SMT 2022<br> [2] (to appear) Zafer Esen and Philipp Rümmer, "TriCera: Verifying C Programs Using the Theory of Heaps", FMCAD 2022</p>
Benchmark for deterministic traffic simulator - parameter space exploration (Prague, Jun 6 2021)
<p>The benchmark is meant for deterministic traffic simulator for optimising traffic flow within a city. The simulator is one part of a traffic modeling framework for intelligent transportation in smart cities. In contrast to standard navigation systems where the navigation is optimised for drivers, we aim to optimise a distribution of the global traffic flow. We utilise HPC resources for the simulator’s parameters exploration for which EVEREST SDK is used.</p> <p>The traffic simulator is available at: <a href="https://github.com/It4innovations/ruth">github.com/It4innovations/ruth</a><strong>.</strong></p> <p><br> The benchmark contains input data, routing map, and skript to run it with HyperQueue. Simulator v1.0 was used.</p>
Bus Violence: a large-scale benchmark for video violence detection in public transport
<p><strong>Dataset</strong></p> <p>The <em>Bus Violence </em>dataset<em> </em>is a large-scale collection of videos depicting violent and non-violent situations in public transport environments. This benchmark was gathered from multiple cameras located inside a moving bus where several people simulated violent actions, such as stealing an object from another person, fighting between passengers, etc. It contains 1,400 video clips manually annotated as having or not violent scenes, making it one of the biggest benchmarks for video violence detection in the literature.</p> <p>Specifically, videos are recorded from three cameras at 25 Frames Per Second (FPS) --- two cameras located in the corners of the bus (with resolution 960x540 px) and one fisheye in the middle (1280x960 px). The clips have a minimum length of 16 frames and a maximum of 48 frames, capturing a very precise action (either violence or non-violence). The dataset is perfectly balanced, containing 700 videos of violence and 700 videos of non-violence.</p> <p>The <em>Bus Violence</em> dataset is intended as a test data benchmark. However, for researchers interested in using our data also for training purposes, we provide training and test splits.</p> <p>In this repository, we provide</p> <ul> <li> <p>the 1,400 video clips divided into two folders named Violence /NoViolence, containing clips of violent situations and non-violent situations, respectively;</p> </li> <li> <p>two txt files containing the names of the videos belonging to the training and test splits, respectively.</p> </li> </ul> <p> </p> <p><strong>Citing our work</strong></p> <p>If you found this dataset useful, please cite the following paper</p> <blockquote> <pre>@inproceedings{bus_violence_dataset_2022, title = {Bus Violence: An Open Benchmark for Video Violence Detection on Public Transport}, doi = {10.3390/s22218345}, url = {https://doi.org/10.3390%2Fs22218345}, year = 2022, month = {oct}, publisher = {{MDPI} {AG}}, volume = {22}, number = {21}, pages = {8345}, author = {Luca Ciampi and Pawe{\l} Foszner and Nicola Messina and Micha{\l} Staniszewski and Claudio Gennaro and Fabrizio Falchi and Gianluca Serao and Micha{\l} Cogiel and Dominik Golba and Agnieszka Szcz{\k{e}}sna and Giuseppe Amato}, journal = {Sensors} } </pre> </blockquote> <p>and this Zenodo Dataset</p> <blockquote> <pre>@dataset{pawel_bus_violence_zenodo, author = {Paweł Foszner, Michał Staniszewski, Agnieszka Szczęsna, Michał Cogiel, Dominik Golba, Luca Ciampi, Nicola Messina, Claudio Gennaro, Fabrizio Falchi, Giuseppe Amato, Gianluca Serao}, title = {{Bus Violence: a large-scale benchmark for video violence detection in public transport}}, month = sep, year = 2022, publisher = {Zenodo}, version = {1.0.0}, doi = {10.5281/zenodo.7044203}, url = {https://doi.org/10.5281/zenodo.7044203} } </pre> </blockquote> <p> </p> <p><strong>Contact Information</strong></p> <p>Blees Sp. z o.o., Gliwice, Poland<br> mstaniszewski@blees.co</p> <p> </p> <p><strong>Acknowledgments</strong></p> <p>The presented dataset was supported by: European Union funds awarded to Blees Sp. z o.o. under grant POIR.01.01.01-00-0952/20-00 “Development of a system for analysing vision data captured by public transport vehicles interior monitoring, aimed at detecting undesirable situations/behaviours and passenger counting (including their classification by age group) and the objects they carry”); EC H2020 project "AI4media: a Centre of Excellence delivering next generation AI Research and Training at the service of Media, Society and Democracy" under GA 951911; research project INAROS (INtelligenza ARtificiale per il mOnitoraggio e Supporto agli anziani), Tuscany POR FSE CUP B53D21008060008.</p> <p> </p> <p><strong>License</strong></p> <p>The <em>Bus Violence </em>dataset was acquired by Blees Sp. z o.o. and is released under a Creative Commons Attribution license for non-commercial use.</p>
Twist Whole-Exome Sequencing Dataset - High Coverage - WGGC SIG4 Benchmarking
<p>GIAB Reference Genome for Benchmarking Initiatives in the West German Genome Center (WGGC) - SIG4. </p> <p>Twist Whole-Exome Sequencing Dataset - High Coverage - 200M Reads.</p> <p> </p> <p> </p>
GeoWaVe Cytometry Benchmark Data
<p>Contained within this folder are six benchmark datasets (Levine13, Levine32, Samusik, Sepsis, and PD) used for the evaluation of the GeoWaVe ensemble clustering algorithm, part of the cytocluster (https://github.com/burtonrj/CytoCluster) package.</p> <p>The data are compensated, arc-sine transformed, and debris and dead cells removed. See manuscript for details: https://doi.org/10.1101/2022.06.30.496829</p> <p>Each dataset is available as a CSV file and includes two additional columns: UMAP1 and UMAP2. The UMAP columns contain embeddings generated using UMAP (2 components and n_neighbours=30) and were used for visualisation purposes. The column 'population' contains the original population labels generated using manual gating.</p>
Unsupervised New Physics detection at 40 MHz: h+ -> tau nu Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of h+ -> tau nu decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
Unsupervised New Physics detection at 40 MHz: h^0 -> tau tau Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of h^0 -> tau tau decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
Unsupervised New Physics detection at 40 MHz: LQ -> b tau Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of Leptoquarks -> b tau decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
Unsupervised New Physics detection at 40 MHz: A -> 4 leptons Signal Benchmark Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Signal Benchmark Dataset consisting of A -> 4 leptons decays produced in collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
TCAB: Text Classification Attack Benchmark Dataset
<p>TCAB is a large collection of successful adversarial attacks on state-of-the-art text classification models trained on multiple sentiment and abuse domain datasets.</p> <p>The dataset is broken up into 2 files: <em>train.csv and</em> <em>val.csv</em>. The training set contains 1,448,751 instances (552,364 are "clean" unperturbed instances) and the validation set contains 482,914 instances (178,607 are "clean"). Each instance contains the following attributes:</p> <p><strong>scenario</strong>: Domain, either <em>abuse</em> or <em>sentiment</em>.</p> <p><strong>target_model_dataset</strong>: Dataset being attacked.</p> <p><strong>target_model_train_dataset</strong>: Dataset the target model trained on.</p> <p><strong>target_model</strong>: Type of victim model (e.g., <em>bert</em>, <em>roberta</em>, <em>xlnet</em>).</p> <p><strong>attack_toolchain</strong>: Open-source attack toolchain, either TextAttack or OpenAttack.</p> <p><strong>attack_name</strong>: Name of the attack method.</p> <p><strong>original_text</strong>: Original input text.</p> <p><strong>original_output</strong>: Prediction probabilities of the target model on the original text.</p> <p><strong>ground_truth</strong>: Encoded label for the original task of the domain dataset. 1 and 0 means toxic and toxic for abuse datasets, respectively. 1 and 0 means positive and negative sentiment for sentiment datasets. If there is a neutral sentiment, then 2, 1, 0 means positive, neutral, and negative sentiment.</p> <p><strong>status</strong>: Unperturbed example if "clean"; successful adversarial attack if "success".</p> <p><strong>perturbed_text</strong>: Text after it has been perturbed by an attack.</p> <p><strong>perturbed_output</strong>: Prediction probabilities of the target model on the perturbed text.</p> <p><strong>attack_time</strong>: Time taken to execute the attack.</p> <p><strong>num_queries</strong>: Number of queries performed while attacking.</p> <p><strong>frac_words_changed</strong>: Fraction of words changed due to an attack.</p> <p><strong>test_index</strong>: Index of each unique source example (original instance) (LEGACY - necessary for backwards compatibility).</p> <p><strong>original_text_identifier</strong>: Index of each unique source example (original instance).</p> <p><strong>unique_src_instance_identifier</strong>: Primary key to uniquely identify to every source instance; comprised of (<em>target_model_dataset</em>, <em>test_index</em>, <em>original_text_identifier</em>).</p> <p><strong>pk</strong>: Primary key to uniquely identify every attack instance; comprised of (<em>attack_name</em>, <em>attack_toolchain</em>, <em>original_text_identifier</em>, <em>scenario</em>, <em>target_model</em>, <em>target_model_dataset</em>, <em>test_index).</em></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.