Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,549
datasets available to search
ShareScore release 0.9.0
Dataset results
1,549 results for “benchmarks”
FDup deduplication software data benchmark: 10Mi OpenAIRE Publications Dump
<p>This dataset is a random subset of publications extracted from the OpenAIRE Research Graph (<a href="http://doi.org/10.5281/zenodo.4707307">http://doi.org/10.5281/zenodo.4707307</a>). The dataset contains ~10Mi JSON publications records. </p> <p>The file is a zip archive containing gz files, each with one JSON per line. Each JSON is compliant to the schema available at <a href="http://doi.org/10.5281/zenodo.4723403">http://doi.org/10.5281/zenodo.4723403</a>.<br> <br> Learn more about the OpenAIRE Research Graph at <a href="https://graph.openaire.eu/">https://graph.openaire.eu</a>.</p>
QR-Code Optical Covert Channel Benchmark
<p>Benchmark results of the QR-Code Optical Covert Channel existing in the reference implementation of a open-source secure data infrastructure and processes.</p>
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
<p>This benchmark dataset is published with the article: </p> <p><em>Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Martin Katz, and Nikolaos Aletras. 2021. LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. ArXiv.</em></p> <p><strong>Short Description</strong></p> <p>Inspired by the recent widespread use of the GLUE multi-task benchmark NLP dataset (Wang et al., 2018), the subsequent more difficult SuperGLUE (Wang et al., 2019), other previous multi-task NLP benchmarks (Conneau and Kiela,2018; McCann et al., 2018), and similar initiatives in other domains (Peng et al., 2019), we introduce LexGLUE, a benchmark dataset to evaluate the performance of NLP methods in legal tasks. LexGLUE is based on seven existing legal NLP datasets:</p> <ul> <li>ECtHR Task A (Chalkidis et al., 2019)</li> <li>ECtHR Task B (Chalkidis et al., 2021a)</li> <li>SCOTUS (Spaeth et al., 2020)</li> <li>EUR-LEX (Chalkidis et al., 2021b)</li> <li>LEDGAR (Tuggener et al. (2020)</li> <li>UNFAIR-ToS (Lippi et al., 2019)</li> <li>CaseHOLD (Zheng et al., 2021)</li> </ul>
Fatigue Crack Propagation Benchmark, GDR 3651 FATACRACK
<p>This is a data set for fatigue crack propagation following the benchmark defined within the french research network GDR 3651 FATACRACK funded by CNRS <a href="http://www.gdr3651.cnrs.fr/">http://www.gdr3651.cnrs.fr</a>. </p> <p>Only two test configurations are reported in this data set. But DIC allows to provide for the analysis not only of the crack tip state ( tip position, crack growth rate and stress intensity factors) but also of the displacement amplitude along the boundary of the analyzed domain. The data set can thus be used to validate fatigue crack growth models.</p> <p>A detailed description of the data set is given in the pdf file Fatigue_Crack_Propagation_Benchmark.pdf.</p> <p><strong>!!!!! there is unfortunately a mistake in the pdf document: the sample thickness is 4 mm !!!!!</strong></p> <p><br> </p> <p> </p>
MARCONI KLN supercomputer HPCG benchmark
<p>This dataset collects High Performance Conjugate Gradient (HPCG) benchmark results of Marconi KNL (Intel Xeon PHI partition) supercomputer, carried out in May 2017.</p>
Marconi supercomputer KNL partition benchmark (HPL and single node stream)
<p>This dataset collect the data stored during the procedure of evaluation of the Marconi KNL machine at CINECA in order to classify it for Top500 list.</p> <p>Dataset also include a report summarizing the results of benchmarks (STREAM for single node memory assessment and HPL for HPC parallel performance) carried out in Nov. 2016.</p>
Data for - Tracking one-in-a-million: Large-scale benchmark for microbial single-cell tracking with experiment-aware robustness metrics
<p><strong>Large-scale Corynebacterium glutamicum data set with Segmentation and Tracking Annotation</strong></p> <p>We provide five time-lapse sequences with manually corrected segmentation and tracking annotations of growing <strong><em>C. glutamicum</em></strong> cultivations. The dataset contains more than 1.4 million cell observations in 29k cell tracks and 14k cell divisions. We provide videos of the annotations (videos.zip) and the dataset in <a href="http://celltrackingchallenge.net/datasets/">Cell Tracking Challenge</a> format (ctc_format.zip). In the videos, cell contours are rendered in yellow, cell links between frames are colored red and cell divisions, and their links are colored in blue.</p> <p><strong>Data Acquisition</strong></p> <p><strong><em>Corynebacterium glutamicum</em></strong> ATCC 13032 was cultivated in BHI-medium at 30°C in this study. From and overnight preculture, the main culture was inoculated the next day with a starting OD600 of 0.05 and grown at 120 rpm to a OD600 of 0.25. A chip was fabricated, according to <a href="https://doi.org/10.1039/D0LC00711K">(Täuber et al., 2020)</a>, and fixed to the microscope’s holder. The main culture cells were transferred to monolayer growth chambers (height = 720 nm) on the microfluidic chip. Flow through the microfluidic device was mediated by pressure driven pumps with a pressure of 100 mbar on the medium reservoir.</p> <p>The time-lapse phase contrast images of five monolayer growth chambers were taken every minute using an inverted microscope (Nikon Eclipse Ti2) with a 100x oil emersion objective and a DS-QI2 camera (Nikon) at 15 % relative DIA-illumination intensity and 100 ms exposure time. The spatial image resolution is 0.072 μm/px.</p>
NGS competence network - pipeline benchmark data
<p>VCF files generated with the megSAP pipeline for a pipeline benchmark performed by the NGS competence network.</p>
WHU-OHS: A benchmark dataset for large-scale Hyperspectral Image classification
<p>The WHU-OHS dataset is made up of 42 OHS satellite images acquired from more than 40 different locations in China. The imagery has a spatial resolution of 10 m (nadir) and a swath width of 60 km (nadir). There are 32 spectral channels ranging from the visible to near-infrared range, with an average spectral resolution of 15 nm. We cropped each image into 512 × 512 pixels with a stride of 32. There are 4822, 513, and 2460 sub-images in the training, validation, and test sets, respectively.</p> <p>For transferability test, we choose eight pairs of OHS images, and each pair contains one source image (S) and one target image (T):</p> <p>S1: Changchun</p> <p>T1: Jilin</p> <p>S2: Wuxi</p> <p>T2: Shanghai</p> <p>S3: Guangzhou</p> <p>T3: Zhongshan</p> <p>S4: Xining</p> <p>T4: Lanzhou</p> <p>S5: Hetian</p> <p>T5: Kelamayi</p> <p>S6: Anyi</p> <p>T6: Nanchang</p> <p>S7: Changde</p> <p>T7: Changsha</p> <p>S8: Tianjin</p> <p>T8: Tangshan</p> <p>The 26 OHS images except for the eight pairs:</p> <p>O1: Baoding</p> <p>O2: Chongqing</p> <p>O3: Fujin</p> <p>O4: Huainan</p> <p>O5: Huhehaote</p> <p>O6: Jinzhong</p> <p>O7: Luliang</p> <p>O8: Manasi_1</p> <p>O9: Manasi_2</p> <p>O10: Nanmulin</p> <p>O11: Neimenggu</p> <p>O12: Qingdao</p> <p>O13: Qinghuangdao</p> <p>O14: Shawan</p> <p>O15: Shenyang</p> <p>O16: Shuozhou</p> <p>O17: Songpan</p> <p>O18: Taian</p> <p>O19: Tongjiang_1</p> <p>O20: Tongjiang_2</p> <p>O21: Wuzhong</p> <p>O22: Xundian</p> <p>O23: Xuzhou</p> <p>O24: Yidu</p> <p>O25: Zangzu</p> <p>O26: Zhongshan</p> <p>The image patches have been normalized and scaled by 10000 to reduce storage cost. Divide the pixel values by 10000 and then the image patches can be used directly.</p>
A boreal forest model benchmarking dataset for North America: a case study with the Canadian Land Surface Scheme including Biogeochemical Cycles (CLASSIC)
<p>A boreal forest model benchmarking dataset for North America by harmonizing eddy covariance and supporting measurements from black spruce (Picea mariana)-dominated mature forest stands.</p> <p>Dataset glossary and users’ instructions are documented in ‘README.md’. </p>
CrossCodeBench: Benchmarking Cross-Task Generalization of Source Code Models
<p>The raw datasets and tasks of the paper "CrossCodeBench: Benchmarking Cross-Task Generalization of Source Code Models". Source code is available at https://anonymous.4open.science/r/CrossCodeBench-C538/.</p>
Benchmark datasets for detection and identification of insects from camera trap images with deep learning
<p><strong>Insect benchmark datasets for training, validation and test (train1201.zip, val1201.zip and test1201.zip) with time-lapse images as described in paper:</strong></p> <p><a href="https://www.biorxiv.org/content/10.1101/2022.10.25.513484v1">Bjerge K, Alison J, Dyrmann M, Frigaard C.E., Mann H. M. R., Høye T.T., Accurate detection and identification of insects from camera trap images with deep learning, bioRxiv:10.1101/2022.10.25.513484v1</a></p> <p>Labels in <strong>YOLO format: <a href="https://github.com/ultralytics/yolov5/issues/2293">ultralytics/yolov5: label format</a></strong></p> <p>The annotated training and validation datasets contains insects of nine different species as listed below:</p> <table> <tbody> <tr> <td>0 <em>Coccinellidae septempunctata</em></td> </tr> <tr> <td>1 <em>Apis mellifera</em></td> </tr> <tr> <td>2 <em>Bombus lapidarius</em></td> </tr> <tr> <td>3 <em>Bombus terrestris</em></td> </tr> <tr> <td>4 <em>Eupeodes corolla</em></td> </tr> <tr> <td>5 <em>Episyrphus balteatus</em></td> </tr> <tr> <td>6 <em>Aglais urticae</em></td> </tr> <tr> <td>7 <em>Vespula vulgaris</em></td> </tr> <tr> <td>8 <em>Eristalis tenax</em></td> </tr> </tbody> </table> <p>The test dataset contains additional classes of insects.</p> <table> <tbody> <tr> <td>9 Non-Bombus Anthophila</td> </tr> <tr> <td>10 Bombus spp.</td> </tr> <tr> <td>11 Syrphidae</td> </tr> <tr> <td>12 Fly spp.</td> </tr> <tr> <td>13 Unclear insect</td> </tr> <tr> <td>14 Mixed animals:<br> ——————————<br> Rhopalocera<br> Non-Anthophila Hymenoptera<br> Non-Syrphidae Diptera<br> Non-Conccinalidae Coleoptera<br> Concinellidae<br> Other animals</td> </tr> </tbody> </table> <p><strong>There are two naming conventions for image (.jpg) and label (.txt) files.</strong></p> <p><em>Background images without insects are named</em>:<br> “<strong>X_Seq-YYYYMMDDHHMMSS</strong>-snapshot”.<br> E.g.:<br> Background image: 12_13-20190704172200-snapshot.jpg<br> Empty label file: 12_13-20190704172200-snapshot.txt</p> <p><em>Images annotated with insects are named:</em><br> “<strong>SZ_IP-MonthDate_C_Seq-YYYYMMDDHHMMSS</strong>”.<br> E.g.:<br> Image file: S1_146-Aug23_1_156-20190822133230.jpg<br> Label file: S1_146-Aug23_1_156-20190822133230.txt</p> <p><strong>Abbreviations</strong>:</p> <p><strong>YYYYMMDDHHMMSS </strong>– Capture timestamp with year, month, date, hour, minutes, and second<br> <strong>Seq</strong> – Sequence number created by the motion program to separate images<br> <strong>C</strong> – Identification of two cameras with Id=0 or Id=1 in system identified by <strong>SZ_IP</strong><br> <strong>MonthDate </strong>– Folder name for where the original image were stored in the system<br> <strong>SZ_IP</strong> – Identification of five camera systems: S1_123, S2_146, S3_194, S4_199, S5_187 (Two cameras in each system)<br> <strong>X</strong> – An index number related to a specific camera and folder ensuring unique file names of background images from different camera systems.<br> <br> The important information in a filename is system (<strong>SZ_IP</strong>), camera Id (<strong>C</strong>) and timestamp (<strong>YYYYMMDDHHMMSS</strong>).</p> <p><strong>The three best YOLOv5 models (YOLOv5models.zip) from the paper are available in pytorch format.</strong></p> <p>All models are tested with YOLOv5 release v7.0 (22-11-2022): <a href="https://github.com/ultralytics/yolov5">ultralytics/yolov5: YOLOv5 in PyTorch</a></p> <p><strong>insect1201-bestF1-640v5m.pt</strong>: Model no. 6 in Table 2 (F1=0.912)<br> <strong>insect1201-bestF1-1280v5m6.pt</strong>: Model no. 8 in Table 2 (F1=0.925)<br> <strong>insect1201-bestF1-1280v5m6.pt</strong>: Model no. 10 in Table 2 (F1=0.932)</p> <p><strong>insects-1201val.yaml</strong>: YAML file with label names to train YOLOv5</p> <p><strong>trainInsects-1201m.sh</strong>: Linux bash shell script with parameters to train YOLOv5m6<br> <strong>valInsectsF1-1201.sh</strong>: Linux bash shell script with parameters to validated models</p> <p> </p>
IndQNER: Indonesian Benchmark Dataset from the Indonesian Translation of the Quran
<h2>IndQNER</h2> <p>IndQNER is a Named Entity Recognition (NER) benchmark dataset that was created by manually annotating 8 chapters in the Indonesian translation of the Quran. The annotation was performed using a web-based text annotation tool, <a href="https://www.tagtog.com/" target="_blank" rel="noopener">Tagtog</a>, and the BIO (Beginning-Inside-Outside) tagging format. The dataset contains:</p> <ul> <li>3117 sentences</li> <li>62027 tokens</li> <li>2475 named entities</li> <li>18 named entity categories</li> </ul> <h2>Named Entity Classes</h2> <p>The named entity classes were initially defined by analyzing the existing Quran concepts ontology. The initial classes were updated based on the information acquired during the annotation process. Finally, there are 20 classes, as follows:</p> <ol> <li>Allah</li> <li>Allah's Throne</li> <li>Artifact</li> <li>Astronomical body</li> <li>Event</li> <li>False deity</li> <li>Holy book</li> <li>Language</li> <li>Angel</li> <li>Person</li> <li>Messenger</li> <li>Prophet</li> <li>Sentient</li> <li>Afterlife location</li> <li>Geographical location</li> <li>Color</li> <li>Religion</li> <li>Food</li> <li>Fruit</li> <li>The book of Allah</li> </ol> <h2>Annotation Stage</h2> <p>There were eight annotators who contributed to the annotation process. They were informatics engineering students at the State Islamic University Syarif Hidayatullah Jakarta.</p> <ol> <li>Anggita Maharani Gumay Putri</li> <li>Muhammad Destamal Junas</li> <li>Naufaldi Hafidhigbal</li> <li>Nur Kholis Azzam Ubaidillah</li> <li>Puspitasari</li> <li>Septiany Nur Anggita</li> <li>Wilda Nurjannah</li> <li>William Santoso</li> </ol> <h2>Verification Stage</h2> <p>We found many named entity and class candidates during the annotation stage. To verify the candidates, we consulted Quran and Tafseer (content) experts who are lecturers at Quran and Tafseer Department at the State Islamic University Syarif Hidayatullah Jakarta.</p> <ol> <li>Dr. Eva Nugraha, M.Ag.</li> <li>Dr. Jauhar Azizy, MA</li> <li>Dr. Lilik Ummi Kultsum, MA</li> </ol> <h2>Evaluation</h2> <p>We evaluated the annotation quality of IndQNER by performing experiments in two settings: supervised learning (BiLSTM+CRF) and transfer learning (<a href="https://huggingface.co/indobenchmark/indobert-base-p1" target="_blank" rel="noopener">IndoBERT</a> fine-tuning).</p> <h3>Supervised Learning Setting</h3> <p>The implementation of BiLSTM and CRF utilized <a href="https://huggingface.co/indobenchmark/indobert-base-p1" target="_blank" rel="noopener">IndoBERT</a> to provide word embeddings. All experiments used a batch size of 16. These are the results:</p> <table> <tbody> <tr> <td>Maximum sequence length</td> <td>Number of e-poch</td> <td>Precision</td> <td>Recall</td> <td>F1 score</td> </tr> <tr> <td>256</td> <td>10</td> <td>0.94</td> <td>0.92</td> <td>0.93</td> </tr> <tr> <td>256</td> <td>20</td> <td> 0.99</td> <td>0.97</td> <td>0.98</td> </tr> <tr> <td>256</td> <td>40</td> <td>0.96</td> <td>0.96</td> <td>0.96</td> </tr> <tr> <td>256</td> <td>100</td> <td>0.97</td> <td>0.96</td> <td>0.96</td> </tr> <tr> <td>512</td> <td>10</td> <td>0.92</td> <td>0.92</td> <td>0.92</td> </tr> <tr> <td>512</td> <td>20</td> <td>0.96</td> <td>0.95</td> <td>0.96</td> </tr> <tr> <td>512</td> <td>40</td> <td>0.97</td> <td>0.95</td> <td>0.96</td> </tr> <tr> <td>512</td> <td>100</td> <td>0.97</td> <td>0.95</td> <td>0.96</td> </tr> </tbody> </table> <h3>Transfer Learning Setting</h3> <p>We performed several experiments with different parameters in IndoBERT fine-tuning. All experiments used a learning rate of 2e-5 and a batch size of 16. These are the results:</p> <table> <tbody> <tr> <td>Maximum sequence length</td> <td>Number of e-poch</td> <td>Precision</td> <td>Recall</td> <td>F1 score</td> </tr> <tr> <td>256</td> <td>10</td> <td>0.67</td> <td>0.65</td> <td>0.65</td> </tr> <tr> <td>256</td> <td>20</td> <td> 0.60</td> <td>0.59</td> <td>0.59</td> </tr> <tr> <td>256</td> <td>40</td> <td>0.75</td> <td>0.72</td> <td>0.71</td> </tr> <tr> <td>256</td> <td>100</td> <td>0.73</td> <td>0.68</td> <td>0.68</td> </tr> <tr> <td>512</td> <td>10</td> <td>0.72</td> <td>0.62</td> <td>0.64</td> </tr> <tr> <td>512</td> <td>20</td> <td>0.62</td> <td>0.57</td> <td>0.58</td> </tr> <tr> <td>512</td> <td>40</td> <td>0.72</td> <td>0.66</td> <td>0.67</td> </tr> <tr> <td>512</td> <td>100</td> <td>0.68</td> <td>0.68</td> <td>0.67</td> </tr> </tbody> </table> <p> </p> <p>This dataset is also part of the <a href="https://github.com/IndoNLP/nusa-crowd" target="_blank" rel="noopener">NusaCrowd project</a> which aims to collect Natural Language Processing (NLP) datasets for Indonesian and its local languages.</p> <h2>How to Cite</h2> <p>@InProceedings{10.1007/978-3-031-35320-8_12,<br>author="Gusmita, Ria Hari<br>and Firmansyah, Asep Fajar<br>and Moussallem, Diego<br>and Ngonga Ngomo, Axel-Cyrille",<br>editor="M{\'e}tais, Elisabeth<br>and Meziane, Farid<br>and Sugumaran, Vijayan<br>and Manning, Warren<br>and Reiff-Marganiec, Stephan",<br>title="IndQNER: Named Entity Recognition Benchmark Dataset from the Indonesian Translation of the Quran",<br>booktitle="Natural Language Processing and Information Systems",<br>year="2023",<br>publisher="Springer Nature Switzerland",<br>address="Cham",<br>pages="170--185",<br>abstract="Indonesian is classified as underrepresented in the Natural Language Processing (NLP) field, despite being the tenth most spoken language in the world with 198 million speakers. The paucity of datasets is recognized as the main reason for the slow advancements in NLP research for underrepresented languages. Significant attempts were made in 2020 to address this drawback for Indonesian. The Indonesian Natural Language Understanding (IndoNLU) benchmark was introduced alongside IndoBERT pre-trained language model. The second benchmark, Indonesian Language Evaluation Montage (IndoLEM), was presented in the same year. These benchmarks support several tasks, including Named Entity Recognition (NER). However, all NER datasets are in the public domain and do not contain domain-specific datasets. To alleviate this drawback, we introduce IndQNER, a manually annotated NER benchmark dataset in the religious domain that adheres to a meticulously designed annotation guideline. Since Indonesia has the world's largest Muslim population, we build the dataset from the Indonesian translation of the Quran. The dataset includes 2475 named entities representing 18 different classes. To assess the annotation quality of IndQNER, we perform experiments with BiLSTM and CRF-based NER, as well as IndoBERT fine-tuning. The results reveal that the first model outperforms the second model achieving 0.98 F1 points. This outcome indicates that IndQNER may be an acceptable evaluation metric for Indonesian NER tasks in the aforementioned domain, widening the research's domain range.",<br>isbn="978-3-031-35320-8"<br>}</p> <h2>Contact</h2> <p>If you have any questions or feedback, feel free to contact us at ria.hari.gusmita@uni-paderborn.de or ria.gusmita@uinjkt.ac.id</p>
EUPPBench postprocessing benchmark dataset - gridded data - Part III
<p>The EUMETNET EUPPBench postprocessing benchmark gridded data is an analysis-ready dataset to perform benchmarks of different postprocessing methods on a common dataset.</p> <p>This dataset is using the <a href="https://zarr.dev/">Zarr</a> format. Please look at the <a href="https://zarr.readthedocs.io/en/stable/">Zarr documentation</a> to see how to load and access the data.</p> <p>The documentation of the dataset is available on <a href="https://eupp-benchmark.github.io/EUPPBench-doc/">https://eupp-benchmark.github.io/EUPPBench-doc/</a> .</p> <p>The official way to download the dataset is through the <a href="https://github.com/ecmwf/climetlab">climetlab</a> <a href="https://github.com/EUPP-benchmark/climetlab-eumetnet-postprocessing-benchmark">EUMETNET postprocessing benchmark plugin</a>.</p> <p>This Zenodo repository aims to preserve the dataset by providing long-term storage.</p> <p>Please read the LICENSE file for more information on the data licenses.</p> <p><strong>Installation procedure</strong></p> <p>Download the 3 parts of the dataset</p> <ol> <li> <a href="https://doi.org/10.5281/zenodo.7429236">EUPPBench-gridded.part.z01</a></li> <li> <a href="http://Remark You might also be interested by the gridded data part of this dataset also available on Zenodo here: https://doi.org/10.5281/zenodo.7428239">EUPPBench-gridded.part.z02</a></li> <li> <a href="https://doi.org/10.5281/zenodo.7429917">EUPPBench-gridded.part.zip</a></li> </ol> <p>in a given folder, and on a Linux (or mac OS) terminal, and still in this folder, enter the following commands</p> <pre><code class="language-bash">zip -FF EUPPBench-gridded.part.zip --out EUPPBench-gridded.zip rm EUPPBench-gridded.part.* unzip EUPPBench-gridded.zip </code></pre> <p>This will unpack the dataset. You need at least 250Gb of free space on your disk to perform this operation.</p> <p><strong>Citation</strong></p> <p>If you use this dataset for a publication, please cite the dataset article:</p> <ul> <li>Demaeyer, J., Bhend, J., Lerch, S., Primo, C., Van Schaeybroeck, B., Atencia, A., Ben Bouallègue, Z., Chen, J., Dabernig, M., Evans, G., Faganeli Pucer, J., Hooper, B., Horat, N., Jobst, D., Merše, J., Mlakar, P., Möller, A., Mestre, O., Taillardat, M., and Vannitsem, S.: The EUPPBench postprocessing benchmark dataset v1.0, Earth Syst. Sci. Data Discuss. [preprint], <a href="https://doi.org/10.5194/essd-2022-465">https://doi.org/10.5194/essd-2022-465</a>, in review, 2023.</li> </ul> <p><strong>Remark</strong></p> <p>You might also be interested by the station data part of this dataset also available on Zenodo here: <a href="https://doi.org/10.5281/zenodo.7708362">https://doi.org/10.5281/zenodo.7708362</a>.</p>
EUPPBench postprocessing benchmark dataset - gridded data - Part II
<p>The EUMETNET EUPPBench postprocessing benchmark gridded data is an analysis-ready dataset to perform benchmarks of different postprocessing methods on a common dataset.</p> <p>This dataset is using the <a href="https://zarr.dev/">Zarr</a> format. Please look at the <a href="https://zarr.readthedocs.io/en/stable/">Zarr documentation</a> to see how to load and access the data.</p> <p>The documentation of the dataset is available on <a href="https://eupp-benchmark.github.io/EUPPBench-doc/">https://eupp-benchmark.github.io/EUPPBench-doc/</a> .</p> <p>The official way to download the dataset is through the <a href="https://github.com/ecmwf/climetlab">climetlab</a> <a href="https://github.com/EUPP-benchmark/climetlab-eumetnet-postprocessing-benchmark">EUMETNET postprocessing benchmark plugin</a>.</p> <p>This Zenodo repository aims to preserve the dataset by providing long-term storage.</p> <p>Please read the LICENSE file for more information on the data licenses.</p> <p><strong>Installation procedure</strong></p> <p>Download the 3 parts of the dataset</p> <ol> <li> <a href="https://doi.org/10.5281/zenodo.7429236">EUPPBench-gridded.part.z01</a></li> <li> <a href="https://doi.org/10.5281/zenodo.7429420">EUPPBench-gridded.part.z02</a></li> <li> <a href="https://doi.org/10.5281/zenodo.7429917">EUPPBench-gridded.part.zip</a></li> </ol> <p>in a given folder, and on a Linux (or mac OS) terminal, and still in this folder, enter the following commands</p> <pre><code class="language-bash">zip -FF EUPPBench-gridded.part.zip --out EUPPBench-gridded.zip rm EUPPBench-gridded.part.* unzip EUPPBench-gridded.zip </code></pre> <p>This will unpack the dataset. You need at least 250Gb of free space on your disk to perform this operation.</p> <p> </p> <p><strong>Citation</strong></p> <p>If you use this dataset for a publication, please cite the dataset article:</p> <ul> <li>Demaeyer, J., Bhend, J., Lerch, S., Primo, C., Van Schaeybroeck, B., Atencia, A., Ben Bouallègue, Z., Chen, J., Dabernig, M., Evans, G., Faganeli Pucer, J., Hooper, B., Horat, N., Jobst, D., Merše, J., Mlakar, P., Möller, A., Mestre, O., Taillardat, M., and Vannitsem, S.: The EUPPBench postprocessing benchmark dataset v1.0, Earth Syst. Sci. Data Discuss. [preprint], <a href="https://doi.org/10.5194/essd-2022-465">https://doi.org/10.5194/essd-2022-465</a>, in review, 2023.</li> </ul> <p><strong>Remark</strong></p> <p>You might also be interested by the station data part of this dataset also available on Zenodo here: <a href="https://doi.org/10.5281/zenodo.7708362">https://doi.org/10.5281/zenodo.7708362</a>.</p>
EUPPBench postprocessing benchmark dataset - gridded data - Part I
<p>The EUMETNET EUPPBench postprocessing benchmark gridded data is an analysis-ready dataset to perform benchmarks of different postprocessing methods on a common dataset.</p> <p>This dataset is using the <a href="https://zarr.dev/">Zarr</a> format. Please look at the <a href="https://zarr.readthedocs.io/en/stable/">Zarr documentation</a> to see how to load and access the data.</p> <p>The documentation of the dataset is available on <a href="https://eupp-benchmark.github.io/EUPPBench-doc/">https://eupp-benchmark.github.io/EUPPBench-doc/</a> .</p> <p>The official way to download the dataset is through the <a href="https://github.com/ecmwf/climetlab">climetlab</a> <a href="https://github.com/EUPP-benchmark/climetlab-eumetnet-postprocessing-benchmark">EUMETNET postprocessing benchmark plugin</a>.</p> <p>This Zenodo repository aims to preserve the dataset by providing long-term storage.</p> <p>Please read the LICENSE file for more information on the data licenses.</p> <p><strong>Installation procedure</strong></p> <p>Download the 3 parts of the dataset</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.7429236">EUPPBench-gridded.part.z01</a></li> <li><a href="https://doi.org/10.5281/zenodo.7429420">EUPPBench-gridded.part.z02</a></li> <li><a href="https://doi.org/10.5281/zenodo.7429917">EUPPBench-gridded.part.zip</a></li> </ol> <p>in a given folder, and on a Linux (or mac OS) terminal, and still in this folder, enter the following commands</p> <pre><code class="language-bash">zip -FF EUPPBench-gridded.part.zip --out EUPPBench-gridded.zip rm EUPPBench-gridded.part.* unzip EUPPBench-gridded.zip</code></pre> <p>This will unpack the dataset. You need at least 250Gb of free space on your disk to perform this operation.</p> <p><strong>Citation</strong></p> <p>If you use this dataset for a publication, please cite the dataset article:</p> <ul> <li>Demaeyer, J., Bhend, J., Lerch, S., Primo, C., Van Schaeybroeck, B., Atencia, A., Ben Bouallègue, Z., Chen, J., Dabernig, M., Evans, G., Faganeli Pucer, J., Hooper, B., Horat, N., Jobst, D., Merše, J., Mlakar, P., Möller, A., Mestre, O., Taillardat, M., and Vannitsem, S.: The EUPPBench postprocessing benchmark dataset v1.0, Earth Syst. Sci. Data Discuss. [preprint], <a href="https://doi.org/10.5194/essd-2022-465">https://doi.org/10.5194/essd-2022-465</a>, in review, 2023.</li> </ul> <p> </p> <p><strong>Remark</strong></p> <p>You might also be interested by the station data part of this dataset also available on Zenodo here: <a href="https://doi.org/10.5281/zenodo.7708362">https://doi.org/10.5281/zenodo.7708362</a>.</p>
Input data - Wasteaware Cities Benchmark Indicators - WABI 2023 - Global data analytics
<p>This is the input dataset for the research publication "<em>Socio-economic development drives solid waste management performance in cities: A global analysis using machine learning</em>". It features </p> <ul> <li>Metadata info used by R codes</li> <li>Full data set for the WABI, used by the R codes</li> <li>Data required for plotting the map in Figure 1</li> </ul> <p>The independent variables data set refers to specific indicators of the WABI methodology (<a href="https://www.sciencedirect.com/science/article/pii/S0956053X14004905">https://www.sciencedirect.com/science/article/pii/S0956053X14004905</a>) which generates solid waste management and resource recovery profiles for cities. It is applied here for 40 cities around the world. The data set contains also values for a series of explanatory variables, which are measures of the level of socioeconomic development at country level.</p> <p> </p> <p> </p> <p> </p>
Dataset for "Deep reinforcement learning for the olfactory search POMDP: a quantitative benchmark"
<p>Dataset containing the results shown in Fig. 6 of "Deep reinforcement learning for the olfactory search POMDP: a quantitative benchmark".</p>
Benchmarks of the bench/coreutils programs with and without flambda
<p>This dataset represents the benchmark results of Goblint, conducted on all coreutils programs of the goblint/bench repository. It includes the following measurements (note: flambda enabled means that Goblint is compiled with flambda and the provided optimization flags):</p> <ul> <li>Without flambda</li> <li>With flambda's -Oclassic flag</li> <li>With flambda's -O2 flag</li> <li>With flambda's -O3 flag</li> <li>With flambda's -O3 flag, as well as the following additional flags: -inline-toplevel=400 -inline-max-depth=1 -inline-max-unroll=0</li> </ul>
An Agnostic Benchmark for Optical Remote Sensing Image Super-Resolution
<p>In remote sensing, image super-resolution (ISR) is a technique used to create high-resolution (HR) images from low-resolution (R) satellite images, giving a more detailed view of the Earth’s surface. However, with the constant development and introduction of new ISR algorithms, it can be challenging to stay updated on the latest advancements and evaluate their performance objectively. To address this issue, we introduce SRcheck, a Python package that provides an easy-to-use interface for comparing and benchmarking various ISR methods. SRcheck includes a range of datasets that consist of high-resolution and low-resolution image pairs, as well as a set of quantitative metrics for evaluating the performance of SISR algorithms.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.