Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

369

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

369 results for “Datasets, Benchmarking”

Learn how ShareScore rates datasets ↗
zenodo40/100

LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

<p>This benchmark dataset is published with the article:&nbsp;</p> <p><em>Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Martin Katz, and Nikolaos Aletras. 2021.&nbsp;LexGLUE: A Benchmark Dataset for Legal Language Understanding in English. ArXiv.</em></p> <p><strong>Short Description</strong></p> <p>Inspired by the recent widespread use of the GLUE multi-task benchmark NLP dataset (Wang et al., 2018), the subsequent more difficult SuperGLUE (Wang et al., 2019), other previous multi-task NLP benchmarks (Conneau and Kiela,2018; McCann et al., 2018), and similar initiatives in other domains (Peng et al., &nbsp;2019), we introduce LexGLUE, a benchmark dataset to evaluate the performance of NLP methods in legal tasks. LexGLUE is based on seven existing legal NLP datasets:</p> <ul> <li>ECtHR Task A &nbsp;(Chalkidis et al., 2019)</li> <li>ECtHR Task B &nbsp;(Chalkidis et al., 2021a)</li> <li>SCOTUS (Spaeth et al., 2020)</li> <li>EUR-LEX (Chalkidis et al., 2021b)</li> <li>LEDGAR (Tuggener et al. (2020)</li> <li>UNFAIR-ToS (Lippi et al., 2019)</li> <li>CaseHOLD (Zheng et al., 2021)</li> </ul>

opencc-by-4.0Sep 2021View details →
zenodo40/100

WHU-OHS: A benchmark dataset for large-scale Hyperspectral Image classification

<p>The WHU-OHS dataset is made up of 42 OHS satellite images acquired from more than 40 different locations in China. The imagery has a spatial resolution of 10 m (nadir) and a swath width of 60 km (nadir). There are 32 spectral channels ranging from the visible to near-infrared range, with an average spectral resolution of 15 nm. We cropped each image into 512 &times; 512 pixels with a stride of 32. There are 4822, 513, and 2460 sub-images in the training, validation, and test sets, respectively.</p> <p>For transferability test, we choose eight pairs of OHS images, and each pair contains one source image (S) and one target image (T):</p> <p>S1: Changchun</p> <p>T1: Jilin</p> <p>S2: Wuxi</p> <p>T2: Shanghai</p> <p>S3: Guangzhou</p> <p>T3: Zhongshan</p> <p>S4: Xining</p> <p>T4: Lanzhou</p> <p>S5: Hetian</p> <p>T5: Kelamayi</p> <p>S6: Anyi</p> <p>T6: Nanchang</p> <p>S7: Changde</p> <p>T7: Changsha</p> <p>S8: Tianjin</p> <p>T8: Tangshan</p> <p>The 26 OHS images except for the eight pairs:</p> <p>O1: Baoding</p> <p>O2: Chongqing</p> <p>O3: Fujin</p> <p>O4: Huainan</p> <p>O5: Huhehaote</p> <p>O6: Jinzhong</p> <p>O7: Luliang</p> <p>O8: Manasi_1</p> <p>O9: Manasi_2</p> <p>O10: Nanmulin</p> <p>O11: Neimenggu</p> <p>O12: Qingdao</p> <p>O13: Qinghuangdao</p> <p>O14: Shawan</p> <p>O15: Shenyang</p> <p>O16: Shuozhou</p> <p>O17: Songpan</p> <p>O18: Taian</p> <p>O19: Tongjiang_1</p> <p>O20: Tongjiang_2</p> <p>O21: Wuzhong</p> <p>O22: Xundian</p> <p>O23: Xuzhou</p> <p>O24: Yidu</p> <p>O25: Zangzu</p> <p>O26: Zhongshan</p> <p>The image patches have been normalized and scaled by 10000 to reduce storage cost. Divide the pixel values by 10000 and then the image patches can be used directly.</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

A boreal forest model benchmarking dataset for North America: a case study with the Canadian Land Surface Scheme including Biogeochemical Cycles (CLASSIC)

<p>A boreal forest model benchmarking dataset for North America by harmonizing eddy covariance and supporting measurements from black spruce (Picea mariana)-dominated mature forest stands.</p> <p>Dataset glossary and users&rsquo; instructions are documented in &lsquo;README.md&rsquo;.&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Benchmark datasets for detection and identification of insects from camera trap images with deep learning

<p><strong>Insect benchmark datasets for training, validation and test (train1201.zip, val1201.zip and test1201.zip)&nbsp;with time-lapse images as described in paper:</strong></p> <p><a href="https://www.biorxiv.org/content/10.1101/2022.10.25.513484v1">Bjerge K, Alison J, Dyrmann M, Frigaard C.E., Mann H. M. R., H&oslash;ye T.T., Accurate detection and identification of insects from camera trap images with deep learning, bioRxiv:10.1101/2022.10.25.513484v1</a></p> <p>Labels in&nbsp;<strong>YOLO format:&nbsp;<a href="https://github.com/ultralytics/yolov5/issues/2293">ultralytics/yolov5: label format</a></strong></p> <p>The annotated training and validation datasets contains insects of nine different species as listed below:</p> <table> <tbody> <tr> <td>0&nbsp;<em>Coccinellidae septempunctata</em></td> </tr> <tr> <td>1&nbsp;<em>Apis mellifera</em></td> </tr> <tr> <td>2&nbsp;<em>Bombus lapidarius</em></td> </tr> <tr> <td>3&nbsp;<em>Bombus terrestris</em></td> </tr> <tr> <td>4&nbsp;<em>Eupeodes corolla</em></td> </tr> <tr> <td>5&nbsp;<em>Episyrphus balteatus</em></td> </tr> <tr> <td>6&nbsp;<em>Aglais urticae</em></td> </tr> <tr> <td>7&nbsp;<em>Vespula vulgaris</em></td> </tr> <tr> <td>8&nbsp;<em>Eristalis tenax</em></td> </tr> </tbody> </table> <p>The test dataset contains additional classes of insects.</p> <table> <tbody> <tr> <td>9 Non-Bombus Anthophila</td> </tr> <tr> <td>10 Bombus spp.</td> </tr> <tr> <td>11 Syrphidae</td> </tr> <tr> <td>12 Fly spp.</td> </tr> <tr> <td>13 Unclear insect</td> </tr> <tr> <td>14 Mixed animals:<br> &mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;&mdash;<br> Rhopalocera<br> Non-Anthophila Hymenoptera<br> Non-Syrphidae Diptera<br> Non-Conccinalidae Coleoptera<br> Concinellidae<br> Other animals</td> </tr> </tbody> </table> <p><strong>There are two naming conventions for image (.jpg) and label (.txt) files.</strong></p> <p><em>Background images without insects are named</em>:<br> &ldquo;<strong>X_Seq-YYYYMMDDHHMMSS</strong>-snapshot&rdquo;.<br> E.g.:<br> Background image: 12_13-20190704172200-snapshot.jpg<br> Empty label file: 12_13-20190704172200-snapshot.txt</p> <p><em>Images annotated with insects are named:</em><br> &ldquo;<strong>SZ_IP-MonthDate_C_Seq-YYYYMMDDHHMMSS</strong>&rdquo;.<br> E.g.:<br> Image file: S1_146-Aug23_1_156-20190822133230.jpg<br> Label file: S1_146-Aug23_1_156-20190822133230.txt</p> <p><strong>Abbreviations</strong>:</p> <p><strong>YYYYMMDDHHMMSS&nbsp;</strong>&ndash; Capture timestamp with year, month, date, hour, minutes, and second<br> <strong>Seq</strong>&nbsp;&ndash; Sequence number created by the motion program to separate images<br> <strong>C</strong>&nbsp;&ndash; Identification of two cameras with Id=0 or Id=1 in system identified by&nbsp;<strong>SZ_IP</strong><br> <strong>MonthDate&nbsp;</strong>&ndash; Folder name for where the original image were stored in the system<br> <strong>SZ_IP</strong>&nbsp;&ndash; Identification of five camera systems: S1_123, S2_146, S3_194, S4_199, S5_187 (Two cameras in each system)<br> <strong>X</strong>&nbsp;&ndash; An index number related to a specific camera and folder ensuring unique file names of background images from different camera systems.<br> <br> The important information in a filename is system (<strong>SZ_IP</strong>), camera Id (<strong>C</strong>) and timestamp (<strong>YYYYMMDDHHMMSS</strong>).</p> <p><strong>The three best YOLOv5 models (YOLOv5models.zip)&nbsp;from the paper are available in pytorch format.</strong></p> <p>All models are tested with YOLOv5 release v7.0 (22-11-2022):&nbsp;<a href="https://github.com/ultralytics/yolov5">ultralytics/yolov5: YOLOv5&nbsp;&nbsp;in PyTorch</a></p> <p><strong>insect1201-bestF1-640v5m.pt</strong>: Model no. 6 in Table 2 (F1=0.912)<br> <strong>insect1201-bestF1-1280v5m6.pt</strong>: Model no. 8 in Table 2 (F1=0.925)<br> <strong>insect1201-bestF1-1280v5m6.pt</strong>: Model no. 10 in Table 2 (F1=0.932)</p> <p><strong>insects-1201val.yaml</strong>: YAML file with label names to train YOLOv5</p> <p><strong>trainInsects-1201m.sh</strong>: Linux bash shell script with parameters to train YOLOv5m6<br> <strong>valInsectsF1-1201.sh</strong>: Linux bash shell script with parameters to validated models</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

IndQNER: Indonesian Benchmark Dataset from the Indonesian Translation of the Quran

<h2>IndQNER</h2> <p>IndQNER is a Named Entity Recognition (NER) benchmark dataset that was created by manually annotating 8 chapters in the Indonesian translation of the Quran. The annotation was performed using a web-based text annotation tool, <a href="https://www.tagtog.com/" target="_blank" rel="noopener">Tagtog</a>, and the BIO (Beginning-Inside-Outside) tagging format. The dataset contains:</p> <ul> <li>3117 sentences</li> <li>62027 tokens</li> <li>2475 named entities</li> <li>18 named entity categories</li> </ul> <h2>Named Entity Classes</h2> <p>The named entity classes were initially defined by analyzing the existing Quran concepts ontology. The initial classes were updated based on the information acquired during the annotation process. Finally, there are 20 classes, as follows:</p> <ol> <li>Allah</li> <li>Allah's Throne</li> <li>Artifact</li> <li>Astronomical body</li> <li>Event</li> <li>False deity</li> <li>Holy book</li> <li>Language</li> <li>Angel</li> <li>Person</li> <li>Messenger</li> <li>Prophet</li> <li>Sentient</li> <li>Afterlife location</li> <li>Geographical location</li> <li>Color</li> <li>Religion</li> <li>Food</li> <li>Fruit</li> <li>The book of Allah</li> </ol> <h2>Annotation Stage</h2> <p>There were eight annotators who contributed to the annotation process. They were informatics engineering students at the State Islamic University Syarif Hidayatullah Jakarta.</p> <ol> <li>Anggita Maharani Gumay Putri</li> <li>Muhammad Destamal Junas</li> <li>Naufaldi Hafidhigbal</li> <li>Nur Kholis Azzam Ubaidillah</li> <li>Puspitasari</li> <li>Septiany Nur Anggita</li> <li>Wilda Nurjannah</li> <li>William Santoso</li> </ol> <h2>Verification Stage</h2> <p>We found many named entity and class candidates during the annotation stage. To verify the candidates, we consulted Quran and Tafseer (content) experts who are lecturers at Quran and Tafseer Department at the State Islamic University Syarif Hidayatullah Jakarta.</p> <ol> <li>Dr. Eva Nugraha, M.Ag.</li> <li>Dr. Jauhar Azizy, MA</li> <li>Dr. Lilik Ummi Kultsum, MA</li> </ol> <h2>Evaluation</h2> <p>We evaluated the annotation quality of IndQNER by performing experiments in two settings: supervised learning (BiLSTM+CRF) and transfer learning (<a href="https://huggingface.co/indobenchmark/indobert-base-p1" target="_blank" rel="noopener">IndoBERT</a> fine-tuning).</p> <h3>Supervised Learning Setting</h3> <p>The implementation of BiLSTM and CRF utilized <a href="https://huggingface.co/indobenchmark/indobert-base-p1" target="_blank" rel="noopener">IndoBERT</a> to provide word embeddings. All experiments used a batch size of 16. These are the results:</p> <table> <tbody> <tr> <td>Maximum sequence length</td> <td>Number of e-poch</td> <td>Precision</td> <td>Recall</td> <td>F1 score</td> </tr> <tr> <td>256</td> <td>10</td> <td>0.94</td> <td>0.92</td> <td>0.93</td> </tr> <tr> <td>256</td> <td>20</td> <td>&nbsp;0.99</td> <td>0.97</td> <td>0.98</td> </tr> <tr> <td>256</td> <td>40</td> <td>0.96</td> <td>0.96</td> <td>0.96</td> </tr> <tr> <td>256</td> <td>100</td> <td>0.97</td> <td>0.96</td> <td>0.96</td> </tr> <tr> <td>512</td> <td>10</td> <td>0.92</td> <td>0.92</td> <td>0.92</td> </tr> <tr> <td>512</td> <td>20</td> <td>0.96</td> <td>0.95</td> <td>0.96</td> </tr> <tr> <td>512</td> <td>40</td> <td>0.97</td> <td>0.95</td> <td>0.96</td> </tr> <tr> <td>512</td> <td>100</td> <td>0.97</td> <td>0.95</td> <td>0.96</td> </tr> </tbody> </table> <h3>Transfer Learning Setting</h3> <p>We performed several experiments with different parameters in IndoBERT fine-tuning. All experiments used a learning rate of 2e-5 and a batch size of 16. These are the results:</p> <table> <tbody> <tr> <td>Maximum sequence length</td> <td>Number of e-poch</td> <td>Precision</td> <td>Recall</td> <td>F1 score</td> </tr> <tr> <td>256</td> <td>10</td> <td>0.67</td> <td>0.65</td> <td>0.65</td> </tr> <tr> <td>256</td> <td>20</td> <td>&nbsp;0.60</td> <td>0.59</td> <td>0.59</td> </tr> <tr> <td>256</td> <td>40</td> <td>0.75</td> <td>0.72</td> <td>0.71</td> </tr> <tr> <td>256</td> <td>100</td> <td>0.73</td> <td>0.68</td> <td>0.68</td> </tr> <tr> <td>512</td> <td>10</td> <td>0.72</td> <td>0.62</td> <td>0.64</td> </tr> <tr> <td>512</td> <td>20</td> <td>0.62</td> <td>0.57</td> <td>0.58</td> </tr> <tr> <td>512</td> <td>40</td> <td>0.72</td> <td>0.66</td> <td>0.67</td> </tr> <tr> <td>512</td> <td>100</td> <td>0.68</td> <td>0.68</td> <td>0.67</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>This dataset is also part of the <a href="https://github.com/IndoNLP/nusa-crowd" target="_blank" rel="noopener">NusaCrowd project</a> which aims to collect Natural Language Processing (NLP) datasets for Indonesian and its local languages.</p> <h2>How to Cite</h2> <p>@InProceedings{10.1007/978-3-031-35320-8_12,<br>author="Gusmita, Ria Hari<br>and Firmansyah, Asep Fajar<br>and Moussallem, Diego<br>and Ngonga Ngomo, Axel-Cyrille",<br>editor="M{\'e}tais, Elisabeth<br>and Meziane, Farid<br>and Sugumaran, Vijayan<br>and Manning, Warren<br>and Reiff-Marganiec, Stephan",<br>title="IndQNER: Named Entity Recognition Benchmark Dataset from the Indonesian Translation of the Quran",<br>booktitle="Natural Language Processing and Information Systems",<br>year="2023",<br>publisher="Springer Nature Switzerland",<br>address="Cham",<br>pages="170--185",<br>abstract="Indonesian is classified as underrepresented in the Natural Language Processing (NLP) field, despite being the tenth most spoken language in the world with 198 million speakers. The paucity of datasets is recognized as the main reason for the slow advancements in NLP research for underrepresented languages. Significant attempts were made in 2020 to address this drawback for Indonesian. The Indonesian Natural Language Understanding (IndoNLU) benchmark was introduced alongside IndoBERT pre-trained language model. The second benchmark, Indonesian Language Evaluation Montage (IndoLEM), was presented in the same year. These benchmarks support several tasks, including Named Entity Recognition (NER). However, all NER datasets are in the public domain and do not contain domain-specific datasets. To alleviate this drawback, we introduce IndQNER, a manually annotated NER benchmark dataset in the religious domain that adheres to a meticulously designed annotation guideline. Since Indonesia has the world's largest Muslim population, we build the dataset from the Indonesian translation of the Quran. The dataset includes 2475 named entities representing 18 different classes. To assess the annotation quality of IndQNER, we perform experiments with BiLSTM and CRF-based NER, as well as IndoBERT fine-tuning. The results reveal that the first model outperforms the second model achieving 0.98 F1 points. This outcome indicates that IndQNER may be an acceptable evaluation metric for Indonesian NER tasks in the aforementioned domain, widening the research's domain range.",<br>isbn="978-3-031-35320-8"<br>}</p> <h2>Contact</h2> <p>If you have any questions or feedback, feel free to contact us at ria.hari.gusmita@uni-paderborn.de or ria.gusmita@uinjkt.ac.id</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

EUPPBench postprocessing benchmark dataset - gridded data - Part III

<p>The EUMETNET EUPPBench postprocessing benchmark gridded data is an analysis-ready dataset to perform benchmarks of different postprocessing methods on a common dataset.</p> <p>This dataset is using the <a href="https://zarr.dev/">Zarr</a> format. Please look at the <a href="https://zarr.readthedocs.io/en/stable/">Zarr documentation</a> to see how to load and access the data.</p> <p>The documentation of the dataset is available on <a href="https://eupp-benchmark.github.io/EUPPBench-doc/">https://eupp-benchmark.github.io/EUPPBench-doc/</a> .</p> <p>The official way to download the dataset is through the <a href="https://github.com/ecmwf/climetlab">climetlab</a> <a href="https://github.com/EUPP-benchmark/climetlab-eumetnet-postprocessing-benchmark">EUMETNET postprocessing benchmark plugin</a>.</p> <p>This Zenodo repository aims to preserve the dataset by providing long-term storage.</p> <p>Please read the LICENSE file for more information on the data licenses.</p> <p><strong>Installation procedure</strong></p> <p>Download the 3 parts of the dataset</p> <ol> <li>&nbsp;&nbsp;&nbsp; <a href="https://doi.org/10.5281/zenodo.7429236">EUPPBench-gridded.part.z01</a></li> <li>&nbsp;&nbsp;&nbsp; <a href="http://Remark You might also be interested by the gridded data part of this dataset also available on Zenodo here: https://doi.org/10.5281/zenodo.7428239">EUPPBench-gridded.part.z02</a></li> <li>&nbsp;&nbsp;&nbsp; <a href="https://doi.org/10.5281/zenodo.7429917">EUPPBench-gridded.part.zip</a></li> </ol> <p>in a given folder, and on a Linux (or mac OS) terminal, and still in this folder, enter the following commands</p> <pre><code class="language-bash">zip -FF EUPPBench-gridded.part.zip --out EUPPBench-gridded.zip rm EUPPBench-gridded.part.* unzip EUPPBench-gridded.zip </code></pre> <p>This will unpack the dataset. You need at least 250Gb of free space on your disk to perform this operation.</p> <p><strong>Citation</strong></p> <p>If you use this dataset for a publication, please cite the dataset article:</p> <ul> <li>Demaeyer, J., Bhend, J., Lerch, S., Primo, C., Van Schaeybroeck, B., Atencia, A., Ben Bouall&egrave;gue, Z., Chen, J., Dabernig, M., Evans, G., Faganeli Pucer, J., Hooper, B., Horat, N., Jobst, D., Mer&scaron;e, J., Mlakar, P., M&ouml;ller, A., Mestre, O., Taillardat, M., and Vannitsem, S.: The EUPPBench postprocessing benchmark dataset v1.0, Earth Syst. Sci. Data Discuss. [preprint], <a href="https://doi.org/10.5194/essd-2022-465">https://doi.org/10.5194/essd-2022-465</a>, in review, 2023.</li> </ul> <p><strong>Remark</strong></p> <p>You might also be interested by the station data part of this dataset also available on Zenodo here: <a href="https://doi.org/10.5281/zenodo.7708362">https://doi.org/10.5281/zenodo.7708362</a>.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

EUPPBench postprocessing benchmark dataset - gridded data - Part II

<p>The EUMETNET EUPPBench postprocessing benchmark gridded data is an analysis-ready dataset to perform benchmarks of different postprocessing methods on a common dataset.</p> <p>This dataset is using the <a href="https://zarr.dev/">Zarr</a> format. Please look at the <a href="https://zarr.readthedocs.io/en/stable/">Zarr documentation</a> to see how to load and access the data.</p> <p>The documentation of the dataset is available on <a href="https://eupp-benchmark.github.io/EUPPBench-doc/">https://eupp-benchmark.github.io/EUPPBench-doc/</a> .</p> <p>The official way to download the dataset is through the <a href="https://github.com/ecmwf/climetlab">climetlab</a> <a href="https://github.com/EUPP-benchmark/climetlab-eumetnet-postprocessing-benchmark">EUMETNET postprocessing benchmark plugin</a>.</p> <p>This Zenodo repository aims to preserve the dataset by providing long-term storage.</p> <p>Please read the LICENSE file for more information on the data licenses.</p> <p><strong>Installation procedure</strong></p> <p>Download the 3 parts of the dataset</p> <ol> <li>&nbsp;&nbsp;&nbsp; <a href="https://doi.org/10.5281/zenodo.7429236">EUPPBench-gridded.part.z01</a></li> <li>&nbsp;&nbsp;&nbsp; <a href="https://doi.org/10.5281/zenodo.7429420">EUPPBench-gridded.part.z02</a></li> <li>&nbsp;&nbsp;&nbsp; <a href="https://doi.org/10.5281/zenodo.7429917">EUPPBench-gridded.part.zip</a></li> </ol> <p>in a given folder, and on a Linux (or mac OS) terminal, and still in this folder, enter the following commands</p> <pre><code class="language-bash">zip -FF EUPPBench-gridded.part.zip --out EUPPBench-gridded.zip rm EUPPBench-gridded.part.* unzip EUPPBench-gridded.zip </code></pre> <p>This will unpack the dataset. You need at least 250Gb of free space on your disk to perform this operation.</p> <p>&nbsp;</p> <p><strong>Citation</strong></p> <p>If you use this dataset for a publication, please cite the dataset article:</p> <ul> <li>Demaeyer, J., Bhend, J., Lerch, S., Primo, C., Van Schaeybroeck, B., Atencia, A., Ben Bouall&egrave;gue, Z., Chen, J., Dabernig, M., Evans, G., Faganeli Pucer, J., Hooper, B., Horat, N., Jobst, D., Mer&scaron;e, J., Mlakar, P., M&ouml;ller, A., Mestre, O., Taillardat, M., and Vannitsem, S.: The EUPPBench postprocessing benchmark dataset v1.0, Earth Syst. Sci. Data Discuss. [preprint], <a href="https://doi.org/10.5194/essd-2022-465">https://doi.org/10.5194/essd-2022-465</a>, in review, 2023.</li> </ul> <p><strong>Remark</strong></p> <p>You might also be interested by the station data part of this dataset also available on Zenodo here: <a href="https://doi.org/10.5281/zenodo.7708362">https://doi.org/10.5281/zenodo.7708362</a>.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

EUPPBench postprocessing benchmark dataset - gridded data - Part I

<p>The EUMETNET EUPPBench postprocessing benchmark gridded data is an analysis-ready dataset to perform benchmarks of different postprocessing methods on a common dataset.</p> <p>This dataset is using the <a href="https://zarr.dev/">Zarr</a> format. Please look at the <a href="https://zarr.readthedocs.io/en/stable/">Zarr documentation</a> to see how to load and access the data.</p> <p>The documentation of the dataset is available on <a href="https://eupp-benchmark.github.io/EUPPBench-doc/">https://eupp-benchmark.github.io/EUPPBench-doc/</a> .</p> <p>The official way to download the dataset is through the <a href="https://github.com/ecmwf/climetlab">climetlab</a> <a href="https://github.com/EUPP-benchmark/climetlab-eumetnet-postprocessing-benchmark">EUMETNET postprocessing benchmark plugin</a>.</p> <p>This Zenodo repository aims to preserve the dataset by providing long-term storage.</p> <p>Please read the LICENSE file for more information on the data licenses.</p> <p><strong>Installation procedure</strong></p> <p>Download the 3 parts of the dataset</p> <ol> <li><a href="https://doi.org/10.5281/zenodo.7429236">EUPPBench-gridded.part.z01</a></li> <li><a href="https://doi.org/10.5281/zenodo.7429420">EUPPBench-gridded.part.z02</a></li> <li><a href="https://doi.org/10.5281/zenodo.7429917">EUPPBench-gridded.part.zip</a></li> </ol> <p>in a given folder, and on a Linux (or mac OS) terminal, and still in this folder, enter the following commands</p> <pre><code class="language-bash">zip -FF EUPPBench-gridded.part.zip --out EUPPBench-gridded.zip rm EUPPBench-gridded.part.* unzip EUPPBench-gridded.zip</code></pre> <p>This will unpack the dataset. You need at least 250Gb of free space on your disk to perform this operation.</p> <p><strong>Citation</strong></p> <p>If you use this dataset for a publication, please cite the dataset article:</p> <ul> <li>Demaeyer, J., Bhend, J., Lerch, S., Primo, C., Van Schaeybroeck, B., Atencia, A., Ben Bouall&egrave;gue, Z., Chen, J., Dabernig, M., Evans, G., Faganeli Pucer, J., Hooper, B., Horat, N., Jobst, D., Mer&scaron;e, J., Mlakar, P., M&ouml;ller, A., Mestre, O., Taillardat, M., and Vannitsem, S.: The EUPPBench postprocessing benchmark dataset v1.0, Earth Syst. Sci. Data Discuss. [preprint], <a href="https://doi.org/10.5194/essd-2022-465">https://doi.org/10.5194/essd-2022-465</a>, in review, 2023.</li> </ul> <p>&nbsp;</p> <p><strong>Remark</strong></p> <p>You might also be interested by the station data part of this dataset also available on Zenodo here: <a href="https://doi.org/10.5281/zenodo.7708362">https://doi.org/10.5281/zenodo.7708362</a>.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Dataset for "Deep reinforcement learning for the olfactory search POMDP: a quantitative benchmark"

<p>Dataset containing the results shown in Fig. 6 of &quot;Deep reinforcement learning for the olfactory search POMDP: a quantitative benchmark&quot;.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Materials Science Optimization Benchmark Dataset for High-dimensional, Multi-objective, Multi-fidelity Optimization of CrabNet Hyperparameters

Benchmarks are an essential driver of progress in scientific disciplines. Ideal benchmarks mimic real-world tasks as closely as possible, where insufficient difficulty or applicability can stunt growth in the field. Benchmarks should also have sufficiently low computational overhead to promote accessibility and repeatability. The goal is then to win a "Turing test" of sorts by creating a surrogate model that is indistinguishable from the ground truth observation (at least within the dataset bounds that were explored), necessitating a large amount of data. In materials science and chemistry, industry-relevant optimization tasks are often hierarchical, noisy, multi-fidelity, multi-objective, high-dimensional, and non-linearly correlated while exhibiting mixed numerical and categorical variables subject to linear and non-linear constraints. To complicate matters, unexpected, failed simulation or experimental regions may be present in the search space. In this study, 173219 quasi-random hyperparameter combinations were generated across 23 hyperparameters and used to train CrabNet on the Matbench experimental band gap dataset. The results were logged to a free-tier shared MongoDB Atlas dataset. This study resulted in a regression dataset mapping hyperparameter combinations (including repeats) to MAE, RMSE, computational runtime, and model size for CrabNet model trained on the Matbench experimental band gap benchmark task1. This dataset is used to create a surrogate model as close as possible to running the actual simulations by incorporating heteroskedastic noise. Failure cases for bad hyperparameter combinations were excluded via careful construction of the hyperparameter search space, and so were not considered as was done in prior work. For the regression dataset, percentile ranks were computed within each of the groups of identical parameter sets to enable capturing heteroskedastic noise. This contrasts with a more traditional approach that imposes a-priori assumptions such as Gaussian noise, e.g., by providing a mean and standard deviation. A similar approach can be applied to other benchmark datasets to bridge the gap between optimization benchmarks with low computational overhead and realistically complex, real-world optimization scenarios.

opencc-zeroMar 2023View details →
zenodo40/100

Supramolecular cages benchmark datasets

<p>Two benchmark datasets, comprising 22 well-known supramolecular cages, have been selected from the supramolecular chemistry literature to evaluate cavity detection and volume characterization.</p> <p>All supramolecular cage structures&nbsp;were prepared in the following way: the original CIF files were downloaded from the Cambridge Structural Database (CSD)&nbsp;and atoms and molecular fragments,&nbsp;that are not part of the cage-framework, were removed from the structures,&nbsp;using Diamond Crystal and Molecular Structure Visualization software and saved in a PDB format.</p> <ul> <li><strong>Benchmark&nbsp;dataset 1:</strong> host-guest pairs</li> </ul> <p>For benchmark dataset 1, the guest structures were obtained, from the corresponding CIF files, by removing the supramolecular cage and other molecular fragments that are not part of the guest.&nbsp;&nbsp;</p> <table> <tbody> <tr> <td> <p><strong>Supramolecular Cage Identifier&nbsp;</strong></p> </td> <td> <p><strong>Guest&nbsp;</strong></p> </td> <td> <p><strong>CSD Identifier&nbsp;</strong></p> </td> <td> <p><strong>DOI </strong></p> </td> </tr> <tr> <td> <p><strong>B1</strong>&nbsp;</p> </td> <td> <p>Et4N+&nbsp;</p> </td> <td> <p>718468&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1021/ic8012848">10.1021/ic8012848</a></p> </td> </tr> <tr> <td> <p><strong>B2</strong>&nbsp;</p> </td> <td> <p>BnNMe3+&nbsp;</p> </td> <td> <p>718469&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1021/ic8012848">10.1021/ic8012848</a></p> </td> </tr> <tr> <td> <p><strong>B3</strong>&nbsp;</p> </td> <td> <p>(Cp)2Co+&nbsp;</p> </td> <td> <p>718470&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1021/ic8012848">10.1021/ic8012848</a></p> </td> </tr> <tr> <td> <p><strong>B4</strong>&nbsp;</p> </td> <td> <p>(Cp*)2Co+&nbsp;</p> </td> <td> <p>718471&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1021/ic8012848">10.1021/ic8012848</a></p> </td> </tr> <tr> <td> <p><strong>B5</strong>&nbsp;</p> </td> <td> <p>BF4-&nbsp;</p> </td> <td> <p>1862753&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1002/asia.201801262">10.1002/asia.201801262</a></p> </td> </tr> <tr> <td> <p><strong>B6</strong>&nbsp;</p> </td> <td> <p>ClO4-&nbsp;</p> </td> <td> <p>1862752&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1002/asia.201801262">10.1002/asia.201801262</a></p> </td> </tr> <tr> <td> <p><strong>B7</strong>&nbsp;</p> </td> <td> <p>C60&nbsp;</p> </td> <td> <p>942782&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1021/ja4110446">10.1021/ja4110446</a></p> </td> </tr> <tr> <td> <p><strong>B8</strong>&nbsp;</p> </td> <td> <p>C60&nbsp;</p> </td> <td> <p>1872778&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1002/chem.201805353">10.1002/chem.201805353</a></p> </td> </tr> <tr> <td> <p><strong>B9</strong>&nbsp;</p> </td> <td> <p>Adamantane-2,6-dione&nbsp;</p> </td> <td> <p>183906&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1002/1521-3757(20021018)114:20&lt;3947::AID-ANGE3947&gt;3.0.CO;2-X">10.1002/1521-3757(20021018)114:20&lt;3947::AID-ANGE3947&gt;3.0.CO;2-X</a></p> </td> </tr> <tr> <td> <p><strong>B10</strong>&nbsp;</p> </td> <td> <p>Et4N+&nbsp;</p> </td> <td> <p>100947&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1002/(SICI)1521-3773(19980803)37:13/14&lt;1840::AID-ANIE1840&gt;3.0.CO;2-D">10.1002/(SICI)1521-3773(19980803)37:13/14&lt;1840::AID-ANIE1840&gt;3.0.CO;2-D</a></p> </td> </tr> <tr> <td> <p><strong>B11</strong>&nbsp;</p> </td> <td> <p>Corannulene and Cyclohexane&nbsp;</p> </td> <td> <p>2068665&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1038/s41467-021-24344-w">10.1038/s41467-021-24344-w</a></p> </td> </tr> <tr> <td> <p><strong>B12</strong>&nbsp;</p> </td> <td> <p>C60&nbsp;</p> </td> <td> <p>2068666&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1038/s41467-021-24344-w">10.1038/s41467-021-24344-w</a></p> </td> </tr> <tr> <td> <p><strong>B13</strong>&nbsp;</p> </td> <td> <p>C70&nbsp;</p> </td> <td> <p>2068667&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1038/s41467-021-24344-w">10.1038/s41467-021-24344-w</a></p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Benchmark dataset 2: </strong>topologically and morphologically distinct supramolecular cages</p> <p>The single resorcin[4]arene ligand (O2) was prepared from the corresponding resorcin[4]arene-cage (A1) by removing 5 ligands. In the cyclotricatechylene cage (C1), the outward-pointing phenyl groups (that do not affect the cavity calculations) were removed.</p> <table> <tbody> <tr> <td> <p><strong>Supramolecular Cage Identifier&nbsp;</strong></p> </td> <td> <p><strong>CSD Identifier&nbsp;</strong></p> </td> <td> <p><strong>DOI </strong></p> </td> </tr> <tr> <td> <p><strong>A1</strong>&nbsp;</p> </td> <td> <p>1207879&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1038/38985">10.1038/38985</a></p> </td> </tr> <tr> <td> <p><strong>C1</strong>&nbsp;</p> </td> <td> <p>1892128&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1039/C9CC02103E">10.1039/C9CC02103E</a></p> </td> </tr> <tr> <td> <p><strong>F1</strong>&nbsp;</p> </td> <td> <p>293777&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1126/science.1124985">10.1126/science.1124985</a></p> </td> </tr> <tr> <td> <p><strong>F2</strong>&nbsp;</p> </td> <td> <p>1831430&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1038/nature20771">10.1038/nature20771</a></p> </td> </tr> <tr> <td> <p><strong>H1</strong>&nbsp;</p> </td> <td> <p>768969&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1039/C0CC00234H">10.1039/C0CC00234H</a></p> </td> </tr> <tr> <td> <p><strong>N1</strong>&nbsp;</p> </td> <td> <p>1541839&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1021/jacs.7b05202">10.1021/jacs.7b05202</a></p> </td> </tr> <tr> <td> <p><strong>O1</strong>&nbsp;</p> </td> <td> <p>2074472&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1021/acs.joc.1c00794">10.1021/acs.joc.1c00794</a></p> </td> </tr> <tr> <td> <p><strong>O2</strong>&nbsp;</p> </td> <td> <p>1207879&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1038/38985">10.1038/38985</a></p> </td> </tr> <tr> <td> <p><strong>W1</strong>&nbsp;</p> </td> <td> <p>1416694&nbsp;</p> </td> <td> <p><a href="https://doi.org/10.1038/nchem.2452">10.1038/nchem.2452</a></p> </td> </tr> </tbody> </table>

opencc-by-4.0Mar 2023View details →
zenodo40/100

SciQA benchmark: Dataset and RDF dump

<p>SciQA benchmark of questions and queries.</p> <p>The data dump is in NTriples format (RDF NT) taken from the ORKG system on 14.02.2023 at 02:04PM.<br> The dump can be imported into a virtuoso endpoint or any RDF engine so it can be queried.</p> <p>The questions/queries are provided as JSON files, also train and test files are provided for each of the sets.&nbsp;</p> <p><strong>Types</strong> of questions and queries:</p> <ul> <li>Handcrafted set of 100 questions</li> <li>Auto-generated set of 2465 questions</li> </ul> <p><strong>More details</strong> on certain columns:<br> &quot;Classification rationale&quot; It may contain the following values:</p> <ul> <li>Nested facts in the question</li> <li>Sorting, sum, average, minimum, maximum or count calculation required</li> <li>Filter used</li> <li>Mappings of Asking Point in the question to the ORKG ontology</li> </ul> <p><strong>Explanation of Rationale for Non-factoid</strong>:</p> <ul> <li>Nested facts in the question. An entity (e.g., a system or a paper) or predicate is requested that is not explicitly stated in the question text and must be inferred while searching for an answer.&nbsp;</li> <li>Sorting, sum, average, minimum, maximum or count calculation required. To get the answer to the question it is necessary to make an aggregation of the query results.&nbsp;</li> <li>Filter used. To get the answer to the question it is necessary to use filtering of the query results by some conditions.<br> &nbsp;</li> </ul>

opencc-by-4.0Jan 2022View details →
zenodo40/100

CWFETB-China: Gridded dataset of consumptive water footprints, evaporation, transpiration, and associate benchmarks of crop production in China (2000-2018)

<p>The CWFETB-China is a 5-arcmin gridded dataset of monthly green and blue water footprint of crop production (WFCP), evaporation (E), transpiration (Tr), and associated unit WFCP benchmarks for 21 crops grown in China during 2000-2018. As compared to the existing gridded WFCP datasets, the CWFETB-China has four improvements: (i) It evaluated the effects of different water supply modes (irrigated or rain-fed) and irrigation practices (furrow, sprinkler, and micro-irrigation) on water consumption throughout the crop growth period. (ii) It distinguished between monthly blue and green water consumption via soil evaporation and crop transpiration. (iii) The dataset encompassed both the WFCP in m<sup>3 </sup>yr<sup>-1</sup> and the uWFCP in m<sup>3 </sup>ton<sup>-1</sup>. (iv) It identified uWFCP benchmarks that differentiated between various climatic zones and irrigation practices. The dataset is able to support for precise crop water productivity assessments, agricultural water-saving evaluations, the development of sustainable irrigation techniques, cropping structure optimisation, and crop-related interregional virtual water trade analysis.</p> <p>&nbsp;</p> <p>Format: NetCDF-4 (5 arcmin) or .xlsx files (benchmark data).</p> <p>Projected coordinate system: WGS 84</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

MetroPT2: A Benchmark dataset for predictive maintenance

<p><strong>Abstract</strong></p> <p>The MetroPT2 data set is an outcome of a eXplainable Predictive Maintenance (XPM) project with an urban metro public transportation service in Porto, Portugal. The data was collected in 2022 that aimed to evaluate machine learning methods for online anomaly detection and failure prediction. By capturing several analogic sensor signals (pressure, temperature, current consumption), digital signals (control signals, discrete signals), and GPS information (latitude, longitude, and speed), we provide a dataset that can be easily used to evaluate online machine learning methods. This dataset contains some interesting characteristics and can be a good benchmark for predictive maintenance models.</p> <table> <tbody> <tr> <td> <p>Data Set Characteristics:</p> </td> <td> <p>Multivariate Time series</p> </td> <td> <p>Number of Instances:</p> </td> <td> <p>7116940</p> </td> </tr> <tr> <td> <p>Attribute Characteristics:</p> </td> <td> <p>Real</p> </td> <td> <p>Number of Attributes</p> </td> <td> <p>21</p> </td> </tr> <tr> <td> <p>Associated Tracks:</p> </td> <td> <p>Classification, Regression</p> </td> <td> <p>Missing Values</p> </td> <td> <p>N/A</p> </td> </tr> </tbody> </table> <p><strong>Data Set Information:</strong></p> <p>The dataset was collected to support the development of predictive maintenance, anomaly detection, and remaining useful life (RUL) prediction models for compressors using deep learning and machine learning methods.</p> <p>It consists of multivariate time series data obtained from several analogue and digital sensors installed on the compressor of a train. The data span between 2022-04-28 and 2022-07-28 and includes 16 signals, such as pressures, motor current, oil temperature, flowmeter and electrical signals of air intake valves. The monitoring and logging of industrial equipment events, such as temporal behaviour and fault events, were obtained from records generated by the sensors. The data were logged at 1Hz by an onboard embedded device. You can find a schematic diagram of the air production unit of the compressor system in Figure 4 of the accompanying paper [1]. Also, the paper [2] provides a detailed examination of data collection and specifications of various types of potential failures in an air compressor system.&nbsp;</p> <p><strong>Relevant Papers:</strong></p> <p>[1]- Davari, N., Veloso, B., Ribeiro, R.P., Pereira, P.M., Gama, J.: Predictive maintenance based on anomaly detection using deep learning for air production unit in the railway industry. In: 2021 IEEE 8th International Conference on Data Science and Advanced Analytics (DSAA). pp. 1&ndash;10. IEEE (2021) (DOI: <a href="https://doi.org/10.1109/DSAA53316.2021.9564181">10.1109/DSAA53316.2021.9564181</a>)</p> <p>[2] Veloso, B., Ribeiro, R.P., Pereira, P.M., Gama, J.: The MetroPT dataset for predictive maintenance. Scientific Data 9, no. 1 (2022): 764. (DOI: 10.1038/s41597-022-01877-3)</p> <p>[3]-Barros, M., Veloso, B., Pereira, P.M., Ribeiro, R.P., Gama, J.: Failure detection of an air production unit in the operational context. In: IoT Streams for Data-Driven Predictive Maintenance and IoT, Edge, and Mobile for Embedded Machine Learning, pp. 61&ndash;74. Springer (2020) (DOI: 10.1007/978-3-030-66770-2_5)</p> <p><strong>Failure Information:</strong></p> <p>The dataset is unlabeled, but the failure reports provided by the company are available in the following table. This allows for evaluating the effectiveness of anomaly detection, failure prediction, and RUL estimation algorithms.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td> <p>Nr.</p> </td> <td> <p>Start Time</p> </td> <td> <p>End Time</p> </td> <td>Failure</td> </tr> <tr> <td> <p>1</p> </td> <td> <p>2022-06-04 10:19:24.300</p> </td> <td> <p>2022-06-04 14:22:39.188</p> </td> <td>Air Leak</td> </tr> <tr> <td> <p>2</p> </td> <td> <p>2022-07-11 10:10:18.948</p> </td> <td> <p>2022-07-14 10:22:08.046</p> </td> <td>Oil Leak</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

ArabicSL-Net: A Benchmark Video Dataset for Arabic Words Sign Language

<p>The data was captured by mobile camera in four main organization namely Bank , Cafe , Hospital , and Train&nbsp;station. The ArabicSL-Net initially&nbsp;consists of 307&nbsp;words recorded in&nbsp;approximately 30,000 videos. For each organization, we capture the most&nbsp;representative words that are used in those places. For&nbsp;Bank data, we have a total of&nbsp;76 of words, while&nbsp;Cafe data contains 54 words. For&nbsp;&nbsp;Hospital, we collects videos for&nbsp;102 words, and collects videos for&nbsp;71 words in Train station.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

A Benchmark Dataset with Knowledge Graph Generation for Industry 4.0 Production Lines

<p>A benchmark dataset for knowledge graph generation in Industry 4.0 production lines and to show the benefits of using ontologies and semantic annotations of data to showcase how I4.0 industry can benefit from KGs and semantic datasets. This work is a&nbsp; result of collaborations with the production line managers, supervisors, and engineers of a football industry to acquire realistic production line data. Knowledge Graphs (KGs) or a Knowledge Graph (KG) emerged as a significant technology to store the semantics of the domain entities.&nbsp;The data is mapped and populated&nbsp;with RGOM classes and relations using an automated solution based on JenaAPI, producing an I4.0 KG.&nbsp;<br> <br> Usage:<br> <br> Recently, we use this dataset&nbsp; to analyze the performance of the five state-of-the-art KG embedding models, namely ComplEx, DistMult,TransE, ConvKB, and ConvE. We evaluated the models using two key metrics: Mean Reciprocal Rank (MRR), and Hits@N (Hits@10, Hits@3, and Hits@1). We observed that the TransE model outperforms other models, followed by ComplEx and DistMult, with ConvE demonstrating the lowest performance. Similarly, the dataset can be used alternatively in other potential scenarios.</p>

openmit-licenseMar 2023View details →
zenodo40/100

SCA-2023: A two-part dataset for benchmarking the methods of image precompensation for users with refractive errors

<p>The recent practices of demonstrating various static and video images to users by means of digital, processor-controlled, often self-luminous devices (computer monitors, smartphone and tablet screens, etc.) have spurred the development of various methods for improving the perception of such images through their computer processing. In particular, this applies to the task of precompensating images shown to users with various anomalies of refraction of the eyes (e.g. myopia or astigmatism) in situations where they are not equipped with glasses or other corrective devices. Researchers have proposed a considerable number of such precompensation methods, but to this day there has been no way to accurately compare their quality. We propose an original dataset, which we called &ldquo;SCA-2023&rdquo;, of images specially designed for this purpose. Its most important feature is the fact that it includes not only a set of ground-truth images for implementing the precompensation transform, but also a separate set of images characterizing specific types and degrees of manifestation of the refractive errors. The benchmarking procedure itself includes applying the precompensation transformation to a certain image from the first part of the dataset, computer simulation of the so-called retinal image (distribution of light on the retina of an imaginary observer) based on the selection of the &ldquo;distorting eye&rdquo; from the second part of the dataset, and evaluating the similarity of this image to the ground-truth image, using any of the commonly used similarity metrics for this purpose.</p>

openmit-licenseApr 2023View details →
zenodo40/100

Hang-Time HAR: A Benchmark Dataset for Basketball Activity Recognition using Wrist-worn Inertial Sensors

<p>In this paper we present a benchmark dataset for evaluation of physical human activity recognition from wrist-worn sensors, for the specific setting of basketball training, drills, and games.<br> Basketball activities lend themselves well for measurement by wrist-worn inertial sensors, and systems that are able to detect such sport-relevant activities could be used in applications toward game analysis, guided training, and personal physical activity tracking.<br> The dataset was recorded for two teams from separate countries (USA and Germany) with a total of 24 players who wore an inertial sensor on their wrist and spanned both repetitive basketball training sessions and full games.<br> Particular features of this dataset include an inherent variance through cultural differences in game rules and styles as the data was recorded in two countries, as well as different sport skill levels, since the participants were heterogeneous in terms of prior basketball experience.<br> We illustrate the datasets&#39; features in several time-series analyses and report on a baseline classification performance study with a state-of-the-art deep learning architecture.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

MeetingBank: A Benchmark Dataset for Meeting Summarization

<p>MeetingBank, a benchmark dataset created from the city councils of 6 major U.S. cities to supplement existing datasets. It contains 1,366 meetings with over 3,579 hours of video, as well as transcripts, PDF documents of meeting minutes, agenda, and other metadata. On average, a council meeting is 2.6 hours long and its transcript contains over 28k tokens, making it a valuable testbed for meeting summarizers and for extracting structure from meeting videos. The datasets contains 6,892 segment-level summarization instances for training and evaluating of performance.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

DIA benchmarking dataset (based on Gotti et al., 2021)

<p>Benchmarking dataset of DIA acquisition methods, based on the manuscript &quot;Extensive and Accurate Benchmarking of DIA Acquisition Methods and Software Tools Using a Complex Proteomic Standard&quot; by <a href="https://pubs.acs.org/doi/10.1021/acs.jproteome.1c00490">Gotti et al. (2021)</a>, for the purpose of R vignette describing analysis of DIA data using the <a href="https://bioconductor.org/packages/release/bioc/html/QFeatures.html">QFeatures</a> platform. The raw data are publicly available on the ProteomeXchange platform under identifier <a href="http://proteomecentral.proteomexchange.org/cgi/GetDataset?ID=PXD026600">PXD026600</a>. Subset of files, containing &#39;overlapped&#39; in the File Name, were searched using the <a href="https://github.com/vdemichev/DiaNN">DIA-NN </a>software. Here, pipeline settings, as well as FASTA files and resulting report.tsv file (here labelled as benchmarkingDIA.tsv) are provided.</p>

opencc-by-4.0Jun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record