Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
369
datasets available to search
ShareScore release 0.9.0
Dataset results
369 results for “Datasets Benchmarking”
Datasets for a data-centric image classification benchmark for noisy and ambiguous label estimation
<p>This is the official data repository of the Data-Centric Image Classification (DCIC) Benchmark. The goal of this benchmark is to measure the impact of tuning the dataset instead of the model for a variety of image classification datasets. Full details about the collection process, the structure and automatic download at</p> <p>Paper: https://arxiv.org/abs/2207.06214</p> <p>Source Code: https://github.com/Emprime/dcic</p> <p>The license information is given below as download.</p> <p><strong>Citation</strong></p> <p>Please cite as</p> <pre><code>@article{schmarje2022benchmark, author = {Schmarje, Lars and Grossmann, Vasco and Zelenka, Claudius and Dippel, Sabine and Kiko, Rainer and Oszust, Mariusz and Pastell, Matti and Stracke, Jenny and Valros, Anna and Volkmann, Nina and Koch, Reinahrd}, journal = {36th Conference on Neural Information Processing Systems (NeurIPS 2022) Track on Datasets and Benchmarks}, title = {{Is one annotation enough? A data-centric image classification benchmark for noisy and ambiguous label estimation}}, year = {2022} }</code></pre> <p>Please see the full details about the used datasets below, which should also be cited as part of the license.</p> <pre><code>@article{schoening2020Megafauna, author = {Schoening, T and Purser, A and Langenk{\"{a}}mper, D and Suck, I and Taylor, J and Cuvelier, D and Lins, L and Simon-Lled{\'{o}}, E and Marcon, Y and Jones, D O B and Nattkemper, T and K{\"{o}}ser, K and Zurowietz, M and Greinert, J and Gomes-Pereira, J}, doi = {10.5194/bg-17-3115-2020}, journal = {Biogeosciences}, number = {12}, pages = {3115--3133}, title = {{Megafauna community assessment of polymetallic-nodule fields with cameras: platform and methodology comparison}}, volume = {17}, year = {2020} } @article{Langenkamper2020GearStudy, author = {Langenk{\"{a}}mper, Daniel and van Kevelaer, Robin and Purser, Autun and Nattkemper, Tim W}, doi = {10.3389/fmars.2020.00506}, issn = {2296-7745}, journal = {Frontiers in Marine Science}, title = {{Gear-Induced Concept Drift in Marine Images and Its Effect on Deep Learning Classification}}, volume = {7}, year = {2020} } @article{peterson2019cifar10h, author = {Peterson, Joshua and Battleday, Ruairidh and Griffiths, Thomas and Russakovsky, Olga}, doi = {10.1109/ICCV.2019.00971}, issn = {15505499}, journal = {Proceedings of the IEEE International Conference on Computer Vision}, pages = {9616--9625}, title = {{Human uncertainty makes classification more robust}}, volume = {2019-Octob}, year = {2019} } @article{schmarje2019, author = {Schmarje, Lars and Zelenka, Claudius and Geisen, Ulf and Gl{\"{u}}er, Claus-C. and Koch, Reinhard}, doi = {10.1007/978-3-030-33676-9_26}, issn = {23318422}, journal = {DAGM German Conference of Pattern Regocnition}, number = {November}, pages = {374--386}, publisher = {Springer}, title = {{2D and 3D Segmentation of uncertain local collagen fiber orientations in SHG microscopy}}, volume = {11824 LNCS}, year = {2019} } @article{schmarje2021foc, author = {Schmarje, Lars and Br{\"{u}}nger, Johannes and Santarossa, Monty and Schr{\"{o}}der, Simon-Martin and Kiko, Rainer and Koch, Reinhard}, doi = {10.3390/s21196661}, issn = {1424-8220}, journal = {Sensors}, number = {19}, pages = {6661}, title = {{Fuzzy Overclustering: Semi-Supervised Classification of Fuzzy Labels with Overclustering and Inverse Cross-Entropy}}, volume = {21}, year = {2021} } @article{schmarje2022dc3, author = {Schmarje, Lars and Santarossa, Monty and Schr{\"{o}}der, Simon-Martin and Zelenka, Claudius and Kiko, Rainer and Stracke, Jenny and Volkmann, Nina and Koch, Reinhard}, journal = {Proceedings of the European Conference on Computer Vision (ECCV)}, title = {{A data-centric approach for improving ambiguous labels with combined semi-supervised classification and clustering}}, year = {2022} } @article{obuchowicz2020qualityMRI, author = {Obuchowicz, Rafal and Oszust, Mariusz and Piorkowski, Adam}, doi = {10.1186/s12880-020-00505-z}, issn = {1471-2342}, journal = {BMC Medical Imaging}, number = {1}, pages = {109}, title = {{Interobserver variability in quality assessment of magnetic resonance images}}, volume = {20}, year = {2020} } @article{stepien2021cnnQuality, author = {St{\c{e}}pie{\'{n}}, Igor and Obuchowicz, Rafa{\l} and Pi{\'{o}}rkowski, Adam and Oszust, Mariusz}, doi = {10.3390/s21041043}, issn = {1424-8220}, journal = {Sensors}, number = {4}, title = {{Fusion of Deep Convolutional Neural Networks for No-Reference Magnetic Resonance Image Quality Assessment}}, volume = {21}, year = {2021} } @article{volkmann2021turkeys, author = {Volkmann, Nina and Br{\"{u}}nger, Johannes and Stracke, Jenny and Zelenka, Claudius and Koch, Reinhard and Kemper, Nicole and Spindler, Birgit}, doi = {10.3390/ani11092655}, journal = {Animals 2021}, pages = {1--13}, title = {{Learn to train: Improving training data for a neural network to detect pecking injuries in turkeys}}, volume = {11}, year = {2021} } @article{volkmann2022keypoint, author = {Volkmann, Nina and Zelenka, Claudius and Devaraju, Archana Malavalli and Br{\"{u}}nger, Johannes and Stracke, Jenny and Spindler, Birgit and Kemper, Nicole and Koch, Reinhard}, doi = {10.3390/s22145188}, issn = {1424-8220}, journal = {Sensors}, number = {14}, pages = {5188}, title = {{Keypoint Detection for Injury Identification during Turkey Husbandry Using Neural Networks}}, volume = {22}, year = {2022} }</code></pre> <p>Addition: This repository also contains the original data from the paper "Annotating Ambiguous Images" (https://arxiv.org/abs/2306.12189). The data is created based on the original datasets and license from https://osf.io/t98fz/ and https://osf.io/nqjyw/</p>
IoMT-TrafficData: A Dataset for Benchmarking Intrusion Detection in IoMT
<h3><strong>Article Information<br></strong></h3> <p>The work involved in developing the dataset and benchmarking its use of machine learning is set out in the article ‘IoMT-TrafficData: Dataset and Tools for Benchmarking Intrusion Detection in Internet of Medical Things’. DOI: 10.1109/ACCESS.2024.3437214.</p> <p>Please do cite the aforementioned article when using this dataset. </p> <h3><strong>Abstract</strong></h3> <p>The increasing importance of securing the Internet of Medical Things (IoMT) due to its vulnerabilities to cyber-attacks highlights the need for an effective intrusion detection system (IDS). In this study, our main objective was to develop a Machine Learning Model for the IoMT to enhance the security of medical devices and protect patients’ private data. To address this issue, we built a scenario that utilised the Internet of Things (IoT) and IoMT devices to simulate real-world attacks. We collected and cleaned data, pre-processed it, and provided it into our machine-learning model to detect intrusions in the network. Our results revealed significant improvements in all performance metrics, indicating robustness and reproducibility in real-world scenarios. This research has implications in the context of IoMT and cybersecurity, as it helps mitigate vulnerabilities and lowers the number of breaches occurring with the rapid growth of IoMT devices. The use of machine learning algorithms for intrusion detection systems is essential, and our study provides valuable insights and a road map for future research and the deployment of such systems in live environments. By implementing our findings, we can contribute to a safer and more secure IoMT ecosystem, safeguarding patient privacy and ensuring the integrity of medical data.</p> <h3><strong>ZIP Folder Content</strong></h3> <p>The ZIP folder comprises two main components: <strong>Captures</strong> and <strong>Datasets</strong>. Within the captures folder, we have included all the captures used in this project. These captures are organized into separate folders corresponding to the type of network analysis: BLE or IP-Based. Similarly, the datasets folder follows a similar organizational approach. It contains datasets categorized by type: <strong>BLE</strong>, <strong>IP-Based Packet</strong>, and <strong>IP-Based Flows</strong>.</p> <p>To cater to diverse analytical needs, the datasets are provided in two formats: CSV (Comma-Separated Values) and pickle. The CSV format facilitates seamless integration with various data analysis tools, while the pickle format preserves the intricate structures and relationships within the dataset.</p> <p>This organization enables researchers to easily locate and utilize the specific captures and datasets they require, based on their preferred network analysis type or dataset type. The availability of different formats further enhances the flexibility and usability of the provided data.</p> <h3><strong>Datasets' Content</strong></h3> <p>Within this dataset, three sub-datasets are available, namely <strong>BLE, IP-Based Packet, and IP-Based Flows</strong>. Below is a table of the features selected for each dataset and consequently used in the evaluation model within the provided work.</p> <p>Identified Key Features Within Bluetooth Dataset</p> <table> <tbody> <tr> <td><strong>Feature</strong></td> <td><strong>Meaning</strong></td> </tr> <tr> <td>btle.advertising_header</td> <td>BLE Advertising Packet Header</td> </tr> <tr> <td>btle.advertising_header.ch_sel</td> <td>BLE Advertising Channel Selection Algorithm</td> </tr> <tr> <td>btle.advertising_header.length</td> <td>BLE Advertising Length</td> </tr> <tr> <td>btle.advertising_header.pdu_type</td> <td>BLE Advertising PDU Type</td> </tr> <tr> <td>btle.advertising_header.randomized_rx</td> <td>BLE Advertising Rx Address</td> </tr> <tr> <td>btle.advertising_header.randomized_tx</td> <td>BLE Advertising Tx Address</td> </tr> <tr> <td>btle.advertising_header.rfu.1</td> <td>Reserved For Future 1</td> </tr> <tr> <td>btle.advertising_header.rfu.2</td> <td>Reserved For Future 2</td> </tr> <tr> <td>btle.advertising_header.rfu.3</td> <td>Reserved For Future 3</td> </tr> <tr> <td>btle.advertising_header.rfu.4</td> <td>Reserved For Future 4</td> </tr> <tr> <td>btle.control.instant</td> <td>Instant Value Within a BLE Control Packet</td> </tr> <tr> <td>btle.crc.incorrect</td> <td>Incorrect CRC</td> </tr> <tr> <td>btle.extended_advertising</td> <td>Advertiser Data Information</td> </tr> <tr> <td>btle.extended_advertising.did</td> <td>Advertiser Data Identifier</td> </tr> <tr> <td>btle.extended_advertising.sid</td> <td>Advertiser Set Identifier</td> </tr> <tr> <td>btle.length</td> <td>BLE Length</td> </tr> <tr> <td>frame.cap_len</td> <td>Frame Length Stored Into the Capture File</td> </tr> <tr> <td>frame.interface_id</td> <td>Interface ID</td> </tr> <tr> <td>frame.len</td> <td>Frame Length Wire</td> </tr> <tr> <td>nordic_ble.board_id</td> <td>Board ID</td> </tr> <tr> <td>nordic_ble.channel</td> <td>Channel Index</td> </tr> <tr> <td>nordic_ble.crcok</td> <td>Indicates if CRC is Correct</td> </tr> <tr> <td>nordic_ble.flags</td> <td>Flags</td> </tr> <tr> <td>nordic_ble.packet_counter</td> <td>Packet Counter</td> </tr> <tr> <td>nordic_ble.packet_time</td> <td>Packet time (start to end)</td> </tr> <tr> <td>nordic_ble.phy</td> <td>PHY</td> </tr> <tr> <td>nordic_ble.protover</td> <td>Protocol Version</td> </tr> </tbody> </table> <p> </p> <p>Identified Key Features Within IP-Based Packets Dataset</p> <table> <tbody> <tr> <td><strong>Feature</strong></td> <td><strong>Meaning</strong></td> </tr> <tr> <td>http.content_length</td> <td>Length of content in an HTTP response</td> </tr> <tr> <td>http.request</td> <td>HTTP request being made</td> </tr> <tr> <td>http.response.code</td> <td>Sequential number of an HTTP response</td> </tr> <tr> <td>http.response_number</td> <td>Sequential number of an HTTP response</td> </tr> <tr> <td>http.time</td> <td>Time taken for an HTTP transaction</td> </tr> <tr> <td>tcp.analysis.initial_rtt</td> <td>Initial round-trip time for TCP connection</td> </tr> <tr> <td>tcp.connection.fin</td> <td>TCP connection termination with a FIN flag</td> </tr> <tr> <td>tcp.connection.syn</td> <td>TCP connection initiation with SYN flag</td> </tr> <tr> <td>tcp.connection.synack</td> <td>TCP connection establishment with SYN-ACK flags</td> </tr> <tr> <td>tcp.flags.cwr</td> <td>Congestion Window Reduced flag in TCP</td> </tr> <tr> <td>tcp.flags.ecn</td> <td>Explicit Congestion Notification flag in TCP</td> </tr> <tr> <td>tcp.flags.fin</td> <td>FIN flag in TCP</td> </tr> <tr> <td>tcp.flags.ns</td> <td>Nonce Sum flag in TCP</td> </tr> <tr> <td>tcp.flags.res</td> <td>Reserved flags in TCP</td> </tr> <tr> <td>tcp.flags.syn</td> <td>SYN flag in TCP</td> </tr> <tr> <td>tcp.flags.urg</td> <td>Urgent flag in TCP</td> </tr> <tr> <td>tcp.urgent_pointer</td> <td>Pointer to urgent data in TCP</td> </tr> <tr> <td>ip.frag_offset</td> <td>Fragment offset in IP packets</td> </tr> <tr> <td>eth.dst.ig</td> <td>Ethernet destination is in the internal network group</td> </tr> <tr> <td>eth.src.ig</td> <td>Ethernet source is in the internal network group</td> </tr> <tr> <td>eth.src.lg</td> <td>Ethernet source is in the local network group</td> </tr> <tr> <td>eth.src_not_group</td> <td>Ethernet source is not in any network group</td> </tr> <tr> <td>arp.isannouncement</td> <td>Indicates if an ARP message is an announcement</td> </tr> </tbody> </table> <p> </p> <p>Identified Key Features Within IP-Based Flows Dataset</p> <table> <tbody> <tr> <td><strong>Feature</strong></td> <td><strong>Meaning</strong></td> </tr> <tr> <td>proto</td> <td>Transport layer protocol of the connection</td> </tr> <tr> <td>service</td> <td>Identification of an application protocol</td> </tr> <tr> <td>orig_bytes</td> <td>Originator payload bytes</td> </tr> <tr> <td>resp_bytes</td> <td>Responder payload bytes</td> </tr> <tr> <td>history</td> <td>Connection state history</td> </tr> <tr> <td>orig_pkts</td> <td>Originator sent packets</td> </tr> <tr> <td>resp_pkts</td> <td>Responder sent packets</td> </tr> <tr> <td>flow_duration</td> <td>Length of the flow in seconds</td> </tr> <tr> <td>fwd_pkts_tot</td> <td>Forward packets total</td> </tr> <tr> <td>bwd_pkts_tot</td> <td>Backward packets total</td> </tr> <tr> <td>fwd_data_pkts_tot</td> <td>Forward data packets total</td> </tr> <tr> <td>bwd_data_pkts_tot</td> <td>Backward data packets total</td> </tr> <tr> <td>fwd_pkts_per_sec</td> <td>Forward packets per second</td> </tr> <tr> <td>bwd_pkts_per_sec</td> <td>Backward packets per second</td> </tr> <tr> <td>flow_pkts_per_sec</td> <td>Flow packets per second</td> </tr> <tr> <td>fwd_header_size</td> <td>Forward header bytes</td> </tr> <tr> <td>bwd_header_size</td> <td>Backward header bytes</td> </tr> <tr> <td>fwd_pkts_payload</td> <td>Forward payload bytes</td> </tr> <tr> <td>bwd_pkts_payload</td> <td>Backward payload bytes</td> </tr> <tr> <td>flow_pkts_payload</td> <td>Flow payload bytes</td> </tr> <tr> <td>fwd_iat</td> <td>Forward inter-arrival time</td> </tr> <tr> <td>bwd_iat</td> <td>Backward inter-arrival time</td> </tr> <tr> <td>flow_iat</td> <td>Flow inter-arrival time</td> </tr> <tr> <td>active</td> <td>Flow active duration</td> </tr> </tbody> </table>
NabilaKumala/5026211016: Dataset Benchmark Kebijakan Privasi Pengguna Aplikasi Kesehatan
<p><a href="https://github.com/NabilaKumala/5026211016/files/12818874/Dataset.Benchmark.Kebijakan.Aplikasi.Telemedisin.xlsx">Dataset Benchmark Kebijakan Aplikasi Telemedisin.xlsx</a></p>
The benchmark datasets for object tracking in satellite videos
<p>A new type of earth observation satellite uses "gaze" method to continuously observe a certain area, and uses "video recording" method to record dynamic information and analyze its instantaneous characteristics. With the development of dynamic acquisition technology for satellite video data, significant progress in target tracking has been made in recent years, which plays an important role in monitoring rapidly changing events. Different from targets in ordinary videos, targets in satellite videos usually demonstrate the phenomenon of a small size occupation pattern (small) and weak feature capture (dim) due to occlusion, illumination variation, and confusion with the surroundings.</p>
A benchmark dataset for Manipuri Meetei-Mayek handwritten character recognition
Open the record for dataset details and reuse information.
LEN-DB - Local earthquakes detection: a benchmark dataset of 3-component seismograms built on a global scale
<p>In this study ( <a href="http://www.sciencedirect.com/science/article/pii/S2666544120300010">The paper</a> ) we present a large dataset of 1,249,411 3-component seismograms, recorded along the vertical, north, and east components of 1487 broad-band or very broad-band receivers distributed worldwide, including 631,105 3-component seismograms generated by 304,878 local earthquakes and labeled as earthquakes (EQ), and 618,306 ones labeled as noise (AN). The choice of collecting only local earthquake-data is motivated by the fact that small-magnitude events, which generate relatively small amplitudes and are easily attenuated, are often problematic to detect but provide valuable information about earthquake processes. The labeled data are split into HDF5-Groups: <em>EQ</em> and <em>AN</em>. Each of these groups contains as many HDF5-Datasets as the number of 3-component seismograms; these are labeled in accordance to the format <em>net_sta_starttime</em>, where <em>net</em>, <em>sta</em>, and <em>starttime</em> represent the seismic network, station, and start time of the seismograms. Each HDF5-Dataset (i.e. each triplet of seismograms) has an attribute, which allows accessing the respective metadata. In addition, the HDF5-Group <em>Stations</em> allows accessing stations’ metadata through as many HDF5-Datasets (which are labeled in accordance to the format <em>net_sta)</em> as the number of receivers employed for collecting the waveforms.</p> <p>This global dataset is intended to be used for carrying out a multitude of seismological and signal processing tasks on single-station recordings, and its size particularly suits machine learning (ML) applications.. Application of ML to this dataset shows that a simple Convolutional Neural Network of 67,939 parameters allows discriminating between earthquakes and noise single-station recordings with high accuracy (93.2%), even if applied in regions not investigated by the training set. We make the dataset publicly available as a unique file in HDF5 data format, intending to provide the seismological and broader scientific community with a benchmark for time-series to be used as a testing ground in seismology and signal processing.</p>
CoPhy: Counterfactual Learning of Physical Dynamics (Benchmark Dataset)
<p> </p> <p>Benchmark website: https://projet.liris.cnrs.fr/cophy/</p> <p>Understanding causes and effects in mechanical systems is an essential component of reasoning in the physical world. This work poses a new problem of counterfactual learning of object mechanics from visual input. We develop the COPHY benchmark to assess the capacity of the state-of-the-art models for causal physical reasoning in a synthetic 3D environment. Having observed a mechanical experiment that involves, for example, a falling tower of blocks, a set of bouncing balls or colliding objects, we require to learn to predict how its outcome is affected by an arbitrary intervention on its initial conditions, such as displacing one of the objects in the scene.</p> <p>The main objective for the creation of our benchmark is (a) to focus specifically on evaluating capabilities of state of the art models for performing counterfactual reasoning, (b) to be unbiased in terms of distributions of parameters to be estimated and balanced with respect to possible outcomes, and (c) to have sufficient variety in terms of scenarios<br> and latent physical characteristics of the scene that are not visually observed and therefore can act<br> as confounders.</p> <p>If you use this benchmark, you need to cite the following paper:</p> <p>Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, Christian Wolf. COPHY: Counterfactual Learning of Physical Dynamics. pre-print arXiv:1909.12000, 2019.</p>
Dataset for Article Titled Benchmarking Multiphysics Software
<p>Software files containing input and output data for the article titled "Benchmarking multiphysics software for mantle convection."</p>
RibFrac Dataset: A Benchmark for Rib Fracture Detection, Segmentation and Classification (Tuning/Validation Set)
<p>RibFrac dataset is a benchmark for developping algorithms on rib fracture detection, segmentation and classification. We hope this large-scale dataset could facilitate both clinical research for automatic rib fracture detection and diagnoses, and engineering research for 3D detection, segmentation and classification.</p> <p>This is the Tuning Set (a.k.a. Validation Set in machine learning terminology) of RibFrac dataset, including 80 CTs and the corresponding annotations. Files include:</p> <ol> <li>ribfrac-val-images.zip: 80 chest-abdomen CTs in NII format (nii.gz).</li> <li>ribfrac-val-labels.zip: 80 annotations in NII format (nii.gz).</li> <li>ribfrac-val-info.csv: labels in the annotation NIIs. <ul> <li>public_id: anonymous patient ID to match images and annotations.</li> <li>label_id: discrete label value in the NII annotations.</li> <li>label_code: 0, 1, 2, 3, 4, -1 <ul> <li>0: it is background</li> <li>1: it is a displaced rib fracture</li> <li>2: it is a non-displaced rib fracture</li> <li>3: it is a buckle rib fracture</li> <li>4: it is a segmental rib fracture</li> <li>-1: it is a rib fracture, but we could not define its type due to ambiguity, diagnosis difficulty, etc. Ignore it in the classification task. </li> </ul> </li> </ul> </li> </ol> <p> </p> <p>If you find this work useful in your research, please acknowledge the RibFrac project teams in the paper and cite this project as:</p> <p><em>Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, Jiajun Chen, Ming Li. Deep-</em><em>Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet. EBioMedicine (2020). (<a href="https://doi.org/10.1016/j.ebiom.2020.103106">DOI</a>)</em></p> <p>or using bibtex</p> <p><em>@article{ribfrac2020,<br> title={Deep-Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet},<br> author={Jin, Liang and Yang, Jiancheng and Kuang, Kaiming and Ni, Bingbing and Gao, Yiyi and Sun, Yingli and Gao, Pan and Ma, Weiling and Tan, Mingyu and Kang, Hui and Chen, Jiajun and Li, Ming},<br> journal={EBioMedicine},<br> year={2020},<br> publisher={Elsevier}<br> }</em></p> <p> </p> <p>The RibFrac dataset is a research effort of thousands of hours by experienced radiologists, computer scientists and engineers. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p> </p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>
RibFrac Dataset: A Benchmark for Rib Fracture Detection, Segmentation and Classification (Training Set Part 2)
<p>RibFrac dataset is a benchmark for developping algorithms on rib fracture detection, segmentation and classification. We hope this large-scale dataset could facilitate both clinical research for automatic rib fracture detection and diagnoses, and engineering research for 3D detection, segmentation and classification.</p> <p>Due to size limit of zenodo.org, we split the whole RibFrac Training Set into 2 parts; This is the Training Set Part 2 of RibFrac dataset, including 120 CTs and the corresponding annotations. Files include:</p> <ol> <li>ribfrac-train-images-2.zip: 120 chest-abdomen CTs in NII format (nii.gz).</li> <li>ribfrac-train-labels-2.zip: 120 annotations in NII format (nii.gz).</li> <li>ribfrac-train-info-2.csv: labels in the annotation NIIs. <ul> <li>public_id: anonymous patient ID to match images and annotations.</li> <li>label_id: discrete label value in the NII annotations.</li> <li>label_code: 0, 1, 2, 3, 4, -1 <ul> <li>0: it is background</li> <li>1: it is a displaced rib fracture</li> <li>2: it is a non-displaced rib fracture</li> <li>3: it is a buckle rib fracture</li> <li>4: it is a segmental rib fracture</li> <li>-1: it is a rib fracture, but we could not define its type due to ambiguity, diagnosis difficulty, etc. Ignore it in the classification task. </li> </ul> </li> </ul> </li> </ol> <p> </p> <p>If you find this work useful in your research, please acknowledge the RibFrac project teams in the paper and cite this project as:</p> <p><em>Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, Jiajun Chen, Ming Li. Deep-</em><em>Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet. EBioMedicine (2020). (<a href="https://doi.org/10.1016/j.ebiom.2020.103106">DOI</a>)</em></p> <p>or using bibtex</p> <p><em>@article{ribfrac2020,<br> title={Deep-Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet},<br> author={Jin, Liang and Yang, Jiancheng and Kuang, Kaiming and Ni, Bingbing and Gao, Yiyi and Sun, Yingli and Gao, Pan and Ma, Weiling and Tan, Mingyu and Kang, Hui and Chen, Jiajun and Li, Ming},<br> journal={EBioMedicine},<br> year={2020},<br> publisher={Elsevier}<br> }</em></p> <p> </p> <p>The RibFrac dataset is a research effort of thousands of hours by experienced radiologists, computer scientists and engineers. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p> </p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>
RibFrac Dataset: A Benchmark for Rib Fracture Detection, Segmentation and Classification (Training Set Part 1)
<p>RibFrac dataset is a benchmark for developping algorithms on rib fracture detection, segmentation and classification. We hope this large-scale dataset could facilitate both clinical research for automatic rib fracture detection and diagnoses, and engineering research for 3D detection, segmentation and classification.</p> <p>Due to size limit of zenodo.org, we split the whole RibFrac Training Set into 2 parts; This is the Training Set Part 1 of RibFrac dataset, including 300 CTs and the corresponding annotations. Files include:</p> <ol> <li>ribfrac-train-images-1.zip: 300 chest-abdomen CTs in NII format (nii.gz).</li> <li>ribfrac-train-labels-1.zip: 300 annotations in NII format (nii.gz).</li> <li>ribfrac-train-info-1.csv: labels in the annotation NIIs. <ul> <li>public_id: anonymous patient ID to match images and annotations.</li> <li>label_id: discrete label value in the NII annotations.</li> <li>label_code: 0, 1, 2, 3, 4, -1 <ul> <li>0: it is background</li> <li>1: it is a displaced rib fracture</li> <li>2: it is a non-displaced rib fracture</li> <li>3: it is a buckle rib fracture</li> <li>4: it is a segmental rib fracture</li> <li>-1: it is a rib fracture, but we could not define its type due to ambiguity, diagnosis difficulty, etc. Ignore it in the classification task. </li> </ul> </li> </ul> </li> </ol> <p> </p> <p>If you find this work useful in your research, please acknowledge the RibFrac project teams in the paper and cite this project as:</p> <p><em>Liang Jin, Jiancheng Yang, Kaiming Kuang, Bingbing Ni, Yiyi Gao, Yingli Sun, Pan Gao, Weiling Ma, Mingyu Tan, Hui Kang, Jiajun Chen, Ming Li. Deep-</em><em>Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet. EBioMedicine (2020). (<a href="https://doi.org/10.1016/j.ebiom.2020.103106">DOI</a>)</em></p> <p>or using bibtex</p> <p><em>@article{ribfrac2020,<br> title={Deep-Learning-Assisted Detection and Segmentation of Rib Fractures from CT Scans: Development and Validation of FracNet},<br> author={Jin, Liang and Yang, Jiancheng and Kuang, Kaiming and Ni, Bingbing and Gao, Yiyi and Sun, Yingli and Gao, Pan and Ma, Weiling and Tan, Mingyu and Kang, Hui and Chen, Jiajun and Li, Ming},<br> journal={EBioMedicine},<br> year={2020},<br> publisher={Elsevier}<br> }</em></p> <p> </p> <p>The RibFrac dataset is a research effort of thousands of hours by experienced radiologists, computer scientists and engineers. We kindly ask you to respect our effort by appropriate citation and keeping data license.</p> <p> </p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>
Datasets and benchmarks for airport ground movement research
<p>Datasets and benchmarks for airport ground movement research, derived from OpenStreetMap and data supplied by Manchester Airport PLC</p>
Imbalanced dataset for benchmarking
<p>Imbalanced dataset for benchmarking<br /> =======================</p> <p>The different algorithms of the `imbalanced-learn` toolbox are evaluated on a set of common dataset, which are more or less balanced. These benchmark have been proposed in [1]. The following section presents the main characteristics of this benchmark.</p> <p>Characteristics<br /> -------------------</p> <p>|ID |Name |Repository & Target |Ratio |# samples| # features |<br /> |:---:|:----------------------:|--------------------------------------|:------:|:-------------:|:--------------:|<br /> |1 |Ecoli |UCI, target: imU |8.6:1 |336 |7 |<br /> |2 |Optical Digits |UCI, target: 8 |9.1:1 |5,620 |64 |<br /> |3 |SatImage |UCI, target: 4 |9.3:1 |6,435 |36 |<br /> |4 |Pen Digits |UCI, target: 5 |9.4:1 |10,992 |16 |<br /> |5 |Abalone |UCI, target: 7 |9.7:1 |4,177 |8 |<br /> |6 |Sick Euthyroid |UCI, target: sick euthyroid |9.8:1 |3,163 |25 |<br /> |7 |Spectrometer |UCI, target: >=44 |11:1 |531 |93 |<br /> |8 |Car_Eval_34 |UCI, target: good, v good |12:1 |1,728 |6 |<br /> |9 |ISOLET |UCI, target: A, B |12:1 |7,797 |617 |<br /> |10 |US Crime |UCI, target: >0.65 |12:1 |1,994 |122 |<br /> |11 |Yeast_ML8 |LIBSVM, target: 8 |13:1 |2,417 |103 |<br /> |12 |Scene |LIBSVM, target: >one label |13:1 |2,407 |294 |<br /> |13 |Libras Move |UCI, target: 1 |14:1 |360 |90 |<br /> |14 |Thyroid Sick |UCI, target: sick |15:1 |3,772 |28 |<br /> |15 |Coil_2000 |KDD, CoIL, target: minority |16:1 |9,822 |85 |<br /> |16 |Arrhythmia |UCI, target: 06 |17:1 |452 |279 |<br /> |17 |Solar Flare M0 |UCI, target: M->0 |19:1 |1,389 |10 |<br /> |18 |OIL |UCI, target: minority |22:1 |937 |49 |<br /> |19 |Car_Eval_4 |UCI, target: vgood |26:1 |1,728 |6 |<br /> |20 |Wine Quality |UCI, wine, target: <=4 |26:1 |4,898 |11 |<br /> |21 |Letter Img |UCI, target: Z |26:1 |20,000 |16 |<br /> |22 |Yeast _ME2 |UCI, target: ME2 |28:1 |1,484 |8 |<br /> |23 |Webpage |LIBSVM, w7a, target: minority|33:1 |49,749 |300 |<br /> |24 |Ozone Level |UCI, ozone, data |34:1 |2,536 |72 |<br /> |25 |Mammography |UCI, target: minority |42:1 |11,183 |6 |<br /> |26 |Protein homo. |KDD CUP 2004, minority |111:1|145,751 |74 |<br /> |27 |Abalone_19 |UCI, target: 19 |130:1|4,177 |8 |</p> <p>References<br /> ----------<br /> [1] Ding, Zejin, "Diversified Ensemble Classifiers for H<br /> ighly Imbalanced Data Learning and their Application in Bioinformatics." Dissertation, Georgia State University, (2011).</p> <p>[2] Blake, Catherine, and Christopher J. Merz. "UCI Repository of machine learning databases." (1998).</p> <p>[3] Chang, Chih-Chung, and Chih-Jen Lin. "LIBSVM: a library for support vector machines." ACM Transactions on Intelligent Systems and Technology (TIST) 2.3 (2011): 27.</p> <p>[4] Caruana, Rich, Thorsten Joachims, and Lars Backstrom. "KDD-Cup 2004: results and analysis." ACM SIGKDD Explorations Newsletter 6.2 (2004): 95-108.</p> <p> </p>
CodeEval: Pedagogy Based Benchmark Dataset for Evaluation Of Large Language Models Trained On Code
<p>Categories has been modified to 21 from 26.</p>
CimpleG DNAm benchmarking datasets for cell-type classification and deconvolution
<p><strong>Two large DNAm benchmarking datasets</strong> specifically gathered and <strong>curated for cell-type classification and deconvolution problems.</strong></p> <p>It includes <strong>a leukocytes dataset</strong> and <strong>a somatic cells dataset</strong> in the GenomicRatioSet format from the minfi package.</p> <p>These can be easily <strong>loaded into R with the readRDS function:</strong></p> <blockquote> <p>my_data <- readRDS("CimpleG_benchmarking_datasets_2/leukocytes/tidy_leuk_data.rds")</p> </blockquote> <p>Each dataset includes therein sample data like GEO accession numbers, sample name or ID in their original dataset, cell-type label, one-hot encoded data for each cell-type, preferred train/test splits, and others.</p> <p>Alternatively, <strong>you can also load the individual .csv files</strong>. If you choose this option, I recommend using the function fread from the package data.table. Below I briefly describe these (.csv and .txt) files for the leukocytes dataset, the same logic applies to the somatic cells dataset:</p> <ul> <li> <div> <div>tidy_leuk_data_beta-values.csv</div> <div> <ul> <li>Methylation Beta values matrix</li> </ul> </div> </div> </li> <li> <div> <div> <div>tidy_leuk_data_m-values.csv</div> <div> <ul> <li>Methylation M values matrix</li> </ul> </div> </div> </div> </li> <li> <div> <div> <div>tidy_leuk_data_probe-annotation.txt</div> <div> <ul> <li>Note regarding probe annotation</li> </ul> </div> </div> </div> </li> <li> <div> <div> <div>tidy_leuk_data_probe-metadata.csv</div> <div> <ul> <li>Probe metadata matrix (chr and location)</li> </ul> </div> </div> </div> </li> <li> <div> <div> <div>tidy_leuk_data_sample-metadata.csv</div> <div> <ul> <li>Sample metadata matrix (sample ID, cell type labels, one-hot encoded labels, etc.)</li> </ul> </div> </div> </div> </li> </ul>
AeroPath: An airway segmentation benchmark dataset with challenging pathology
Open the record for dataset details and reuse information.
CSS and Benchmark Datasets of GeminiMol
<p><span>The molecular representation model is a neural network that converts molecular representations (SMILES, Graph) into feature vectors, that carries the potential to be applied across a wide scope of drug discovery scenarios. However, current molecular representation models have been limited to 2D or static 3D structures, overlooking the dynamic nature of small molecules in solution and their ability to adopt flexible conformational changes crucial for drug-target interactions. </span></p> <p><span>To address this limitation, we propose a novel strategy that incorporates the conformational space profile into molecular representation learning. By capturing the intricate interplay between molecular structure and conformational space, our strategy enhances the representational capacity of our model named GeminiMol. Consequently, when pre-trained on a miniaturized molecular dataset, the GeminiMol model demonstrates a balanced and superior performance not only on traditional molecular property prediction tasks but also on zero-shot learning tasks, including virtual screening and target identification. By capturing the dynamic behavior of small molecules, our strategy paves the way for rapid exploration of chemical space, facilitating the transformation of drug design paradigms. </span></p> <p>In this study, a diverse collection of <strong>39,290</strong> molecules was employed for conformational searching and shape alignment to generate a comprehensive dataset of molecular conformational space similarity. To assess the model's performance, the <strong>benchmark datasets comprising over millions molecules</strong> was utilized for downstream tasks. Here, we provide all the training and benchmarking data used for this study to facilitate the reproducibility of the work.</p>
CodeUltraFeedback - Dataset and Benchmark
Open the record for dataset details and reuse information.
Benchmark Dataset of Cropland Parcel Boundaries from Multi-Source Remote Sensing Imagery for AI application
<p><span>We provide a standardized dataset of diversified parcel labels based on multi-source remote sensing imagery. The parcels included in this dataset are sourced from four countries: the Netherlands, Denmark, Spain, and China, encompassing over 1</span><span>4</span><span>0,000 </span><span>cropland parcel</span><span>. The average parcel size across different study areas ranges from 0.05 hectares to 10 hectares. The remote sensing imagery is derived from more than 10 data sources, including satellites, UAV (unmanned aerial vehicle), and public mapping platforms, with spatial resolutions varying from 0.05 meters to 10 meters.</span></p>
A Benchmark Dataset for Lightning Nowcasting in Latin America
<p>In 2023, the <a href="https://wmo.int/" target="_blank" rel="nofollow noreferrer noopener">World Meteorological Organization</a> (WMO) commissioned a pilot project for nowcasting convective weather hazards using artificial intelligence (AI) methods. The project focuses on predicting lightning and quantitative precipitation estimation (QPE) over Latin America, Africa, and southeast Asia. This project is called the AI Nowcasting Pilot Project (AINPP), and is part of WMO's Early Warning for All initiative (<a href="https://wmo.int/activities/early-warnings-all/wmo-and-early-warnings-all-initiative" target="_blank" rel="nofollow noreferrer noopener">EW4All</a>).</p> <p>This benchmark dataset is for evaluating the prediction of lightning in Latin America. The dataset currently consists of 30 days data at 10-minute temporal resolution (5 consecutive days for 6 consecutive months). <br><br>Contents:</p> <ul> <li>mexico_targets.tar.gz: The targets or predictand over Mexico, using 1- and 2-hour accumulations of flash-extent density observed by the GOES-16 Geostationary Lightning Mapper.</li> <li>mexico_example_predictions.tar.gz: Example predictions from the <a href="https://cimss.ssec.wisc.edu/severe_conv/ptlg.html" target="_blank" rel="noopener">NOAA/CIMSS LightningCast model</a>, matching the files in mexico_targets.tar.gz.</li> <li>south_america_targets.tar.gz: The targets or predictand over South America, using 1- and 2-hour accumulations of flash-extent density observed by the GOES-16 Geostationary Lightning Mapper.</li> <li>south_america_example_predictions.tar.gz: Example predictions from the <a href="https://cimss.ssec.wisc.edu/severe_conv/ptlg.html" target="_blank" rel="noopener">NOAA/CIMSS LightningCast model</a>, matching the files in south_america_targets.tar.gz.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.