Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
431
datasets available to search
ShareScore release 0.7.1
Dataset results
431 results for “Training Datasets”
NewsEye / READ AS training dataset from French Newspapers (19th, early 20th C.)
<p>The dataset comprises French newspaper pages from 19th and early 20th century with annotated text. The page images were provided by the <a href="https://www.bnf.fr/en">French National Library</a> and comprise 183 pages (training set). The data are formed according to the PAGE format (cf. Cf. <a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a> and the <a href="http://read.transkribus.eu/">READ </a>project. The guidelines with which the AS GT was created are uploaded here as well.</p>
Dataset of behavioral and neurophysiological data of a virtual sailing task published in: "Providing task instructions during motor training enhances performance and modulates attentional brain networks"
<p>Dataset belonging to the behavioral and neurophysiological data of the publication: "Providing task instructions during motor training enhances performance and modulates attentional brain networks". The two uploaded Zip files contain kinematic and electroencephalographic data of 36 participants for the Obstacle and HorizonTask.</p>
The DR-Train dataset: dynamic responses, GPS positions and environmental conditions of two light rail vehicles in Pittsburgh
<p><strong>Note: Downloading the large data file could have a timeout issue. If you cannot directly download it here, please use the following link as a complementary method for getting the data. </strong></p> <p><a href="https://drive.google.com/drive/folders/1oKn7IN7zznQuhwjDCDdjq8r9wHJYBEhj?usp=sharing">https://drive.google.com/drive/folders/1oKn7IN7zznQuhwjDCDdjq8r9wHJYBEhj?usp=sharing</a></p> <p> </p> <p>This dataset contains the dynamic responses (acceleration records) of two passenger trains with corresponding GPS positions, environmental conditions and track maintenance schedules for a light rail network in the city of Pittsburgh, Pennsylvania in the United States of America.</p> <p>In particular, two light rail vehicles were instrumented (identified as LRV4306 and LRV4313): <br> LRV 4306 has 5 acceleration channels, corresponding to the two uni-axial accelerometers inside the train and the three channels of the tri-axial accelerometer on the wheel truck.</p> <p><em>- The last digit of each acceleration file: 1, 2, 3, 4, 5<br> - Corresponding sensor channels: tri-axial x, tri-axial y, tri-axial z, front cabinet uni-axial, back cabinet uni-axial</em></p> <p><br> LRV 4313 has 8 acceleration channels, corresponding to the two uni-axial accelerometer and the two tri-axial accelerometers inside the train.</p> <p><em>- The last digit of each acceleration file: 1, 2, 3, 4, 5, 6, 7, 8<br> - Corresponding sensor channels: front cabinet uni-axial, back cabinet uni-axial, front tri-axial x, front tri-axial y, front tri-axial z, back tri-axial x, back tri-axial y, back tri-axial z.<br> - x longitudinal (vehicle moving direction); y-axis, transverse; z-axis, vertical.</em></p> <p>The dataset contained in this repository is a condensed version of the original raw data. While the accelerometers on the train were sampled continuously, this dataset contains only those measurements for when the train was actually moving along the track (i.e. not idling at a terminal).</p> <p>The data is stored in binary MAT-files (a MATLAB/Octave data format). These files contain MATLAB objects of the class "pass", which is defined in the file pass.m that can be found in the "code" folder. Specifically, two MAT-files named "obj_dic.mat", and found in the "LRV4306" and "LRV4313" folders, contain the "pass" objects of the two trains, respectively.</p> <p>Each category is described in detail. For more detail on the regions of the track, refer to the 'region.fig' file in this folder. The track was divided into distinct regions so that the data over specific sections of track could be compared. These regions were chosen for two reasons: <br> (1) within a region, the train always followed the same track and <br> (2) there are no tunnels in them so the GPS data is relatively consistent. </p> <p>To get started, using MATLAB or Octave try running "main_script.m" in the "code" folder.</p> <p>A data descriptor paper with details of the data collection process was published.</p> <p>Please cite as</p> <p><strong>Liu, J., Chen, S., Lederman, G., Kramer, D. B., Noh, H. Y., Bielak, J., Garrett, J. H., Kovačević, J., & Berges, M. Dynamic responses, GPS positions and environmental conditions of two light rail vehicles in Pittsburgh. Scientific Data, 6, 146. <a href="https://doi.org/10.1038/s41597-019-0148-9">https://doi.org/10.1038/s41597-019-0148-9</a>(2019)</strong></p> <p><strong>Liu, J., Chen, S., Lederman, G., Kramer, D. B., Noh, H. Y., Bielak, J., Garrett, J. H., Kovačević, J., & Berges, M. The DR-Train dataset: dynamic responses, GPS positions and environmental conditions of two light rail vehicles in Pittsburgh. Zenodo, <a href="https://doi.org/10.5281/zenodo.1432702">https://doi.org/10.5281/zenodo.1432702</a>(2018).</strong></p> <p>For questions or suggestions please e-mail Jingxiao Liu <liujx@stanford.edu></p>
Training Deep Learning Models to Estimate Permeability using Geophysical Datasets
<p>This folder contains the dataset for training deep learning models to estimate permeability using hydro-geophysics simulations</p>
Unsupervised New Physics detection at 40 MHz: Training Dataset
<p>Unsupervised New Physics detection at 40 MHz data challenge</p> <p>Training dataset, consisting of a cocktail of Standard Model collision events (simulation of LHC 13 TeV proton-proton collisions) pre-filtered by a requirement of a muon or electron with 23 GeV transverse momentum. Data format description available on the data challenge web page: https://mpp-hep.github.io/ADC2021/</p>
The PANORAMA Challenge: Public Training and Development Dataset (3)
<p>This dataset represents the <strong><a href="https://panorama.grand-challenge.org/" target="_blank" rel="noopener">PANORAMA</a>: Public Training and Development Dataset</strong>. It contains 2238 anonymized contrast-enhanced CT (CECT) scans acquired at two centers (Radboud University Medical Center, University Medical Center Groningen) based in The Netherlands. Additionally, it contains 194 cases from the <strong><a href="http://medicaldecathlon.com/" target="_blank" rel="noopener">Medical Segmentation Decathlon</a> </strong>dataset and 80 cases from<strong> <a href="https://www.cancerimagingarchive.net/collection/pancreas-ct/" target="_blank" rel="noopener">National Institutes of Health</a></strong>. For all updates/fixes regarding this dataset, please join the challenge and check out our <a href="https://grand-challenge.org/forums/forum/panorama-pancreatic-cancer-diagnosis-radiologists-meet-ai-711/topic/public-training-and-development-dataset-updates-and-fixes-2213/" target="_blank" rel="noopener">dedicated forum post</a> on this topic. The corresponding labels of the PANORAMA dataset can be found <a href="https://github.com/DIAGNijmegen/panorama_labels">here</a>. </p> <p>The PANORAMA challenge is an all-new grand challenge that aims to validate the diagnostic performance of artificial intelligence and radiologists at pancreatic ductal adenocarcinoma (PDAC) detection/diagnosis in CECT, with histopathology and follow-up (≥ 3 years) as the reference standard, in a retrospective setting in the hidden testing dataset. The study hypothesizes that state-of-the-art AI algorithms are non-inferior to radiologists reading CECT.</p> <p>Key aspects of the PANORAMA study design have been established in conjunction with an international scientific advisory board of 13 experts in AI and pancreas radiology as well as a patient representative —to unify and standardize present-day guidelines, and to ensure meaningful validation of pancreas AI towards clinical translation (<strong><a href="https://www.sciencedirect.com/science/article/pii/S2405456921001607">Reinke et al., 2021</a></strong>).</p> <p><em>This PANORAMA dataset contains: batch <strong>3</strong> <strong>out of 4</strong></em></p>
The PANORAMA Challenge: Public Training and Development Dataset (4)
<p>This dataset represents the <strong><a href="https://panorama.grand-challenge.org/" target="_blank" rel="noopener">PANORAMA</a>: Public Training and Development Dataset</strong>. It contains 2238 anonymized contrast-enhanced CT (CECT) scans acquired at two centers (Radboud University Medical Center, University Medical Center Groningen) based in The Netherlands. Additionally, it contains 194 cases from the <strong><a href="http://medicaldecathlon.com/" target="_blank" rel="noopener">Medical Segmentation Decathlon</a> </strong>dataset and 80 cases from<strong> <a href="https://www.cancerimagingarchive.net/collection/pancreas-ct/" target="_blank" rel="noopener">National Institutes of Health</a></strong>. For all updates/fixes regarding this dataset, please join the challenge and check out our <a href="https://grand-challenge.org/forums/forum/panorama-pancreatic-cancer-diagnosis-radiologists-meet-ai-711/topic/public-training-and-development-dataset-updates-and-fixes-2213/" target="_blank" rel="noopener">dedicated forum post</a> on this topic. The corresponding labels of the PANORAMA dataset can be found <a href="https://github.com/DIAGNijmegen/panorama_labels">here</a>. </p> <p>The PANORAMA challenge is an all-new grand challenge that aims to validate the diagnostic performance of artificial intelligence and radiologists at pancreatic ductal adenocarcinoma (PDAC) detection/diagnosis in CECT, with histopathology and follow-up (≥ 3 years) as the reference standard, in a retrospective setting in the hidden testing dataset. The study hypothesizes that state-of-the-art AI algorithms are non-inferior to radiologists reading CECT.</p> <p>Key aspects of the PANORAMA study design have been established in conjunction with an international scientific advisory board of 13 experts in AI and pancreas radiology as well as a patient representative —to unify and standardize present-day guidelines, and to ensure meaningful validation of pancreas AI towards clinical translation (<strong><a href="https://www.sciencedirect.com/science/article/pii/S2405456921001607">Reinke et al., 2021</a></strong>).</p> <p><em>This PANORAMA dataset contains: batch <strong>4</strong> <strong>out of 4</strong></em></p>
Trackerless 3D Freehand Ultrasound Reconstruction Challenge 2024 - Train Dataset (Part 2)
<blockquote> <p><strong>This Challenge will be an open-ended challenge, and we welcome your submission. Please register your team via this <a title="https://forms.office.com/e/dPg47ktV7M" href="https://forms.office.com/e/dPg47ktV7M" target="_blank" rel="noopener">form</a>. You can submit the algorithm via this <a title="https://forms.office.com/e/QChhNkLYiu" href="https://forms.office.com/e/QChhNkLYiu" target="_blank" rel="noopener noreferrer">form</a> for TUS-REC2024 Challenge, and we will test your submitted docker on the test set.</strong></p> <p><strong>We are organising TUS-REC2025 at MICCAI2025. More information is available on the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/" target="_blank" rel="noopener">TUS-REC2025 challenge website</a> and <a href="https://github.com/QiLi111/TUS-REC2025-Challenge_baseline" target="_blank" rel="noopener">Baseline code repo</a>.</strong></p> </blockquote> <p><strong>This is the second part of the Challenge dataset. <a href="../doi/10.5281/zenodo.11178509" target="_blank" rel="noopener">Link</a> to first part; <a href="../doi/10.5281/zenodo.11355500" target="_blank" rel="noopener">Link</a> to third part. <a href="../doi/10.5281/zenodo.12979481" target="_blank" rel="noopener">Link</a> to validation dataset.</strong></p> <p>Acquisition devices and config: The 2D US images were acquired using an Ultrasonix machine (BK, Europe) with a curvilinear probe (4DC7-3/40). The associated position information of each frame was recorded by an optical tracker (NDI Polaris Vicra, Northern Digital Inc., Canada). The acquired US frames were recorded at 20 fps, with an image size of 480×640, without speckle reduction. The frequency was set at 6MHz with a dynamic range of 83 dB, an overall gain of 48% and a depth of 9 cm. </p> <div> <p>Scanning protocol: Both left and right forearms of volunteers were scanned. For each forearm, the US probe moves in three different trajectories (straight line shape, "C" shape, and "S" shape), in a distal-to-proximal direction followed by a proximal-to-distal direction, with the US plane perpendicular of and parallel to the scanning direction. The train dataset contains 1200 scans in total, 24 scans associated with each subject.</p> <div> <div> <p>For detailed information please refer to the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/TUS-REC2024/" target="_blank" rel="noopener">Challenge website</a>. Baseline code is also provided, which can be found at this <a href="https://github.com/QiLi111/tus-rec-challenge_baseline" target="_blank" rel="noopener">repo</a>.</p> <p>Dataset structure: </p> </div> <div> <ul> <li> <p>The dataset contains 50 folders (one subject per folder), each with 24 scans. Each .h5 file corresponds to one scan, storing image and transformation of each frame within this scan. Key-value pairs in each .h5 file are explained below.</p> <ul> <li> <p>“frames” - All frames in the scan; with a shape of [N,H,W], where N refers to the number of frames in the scan, H and W denote the height and width of a frame. </p> </li> <li> <p>“tforms” - All transformations in the scan; with a shape of [N,4,4], where N is the number of frames in the scan, and the transformation matrix denotes the transformation from tracker tool space to camera space. </p> </li> <li> <p>Notations in the name of each .h5 file: “RH”: right arm; “LH”: left arm; “Per”: perpendicular; “Par”: parallel; “L”: straight line shape; “C”: C shape; “S”: S shape; “DtP”: distal-to-proximal direction; “PtD”: proximal-to-distal direction; For example, “RH_Per_L_DtP.h5” denotes a scan on the right forearm, with ultrasound probe perpendicular of the forearm sweeping along straight line, in distal-to-proximal direction.</p> </li> </ul> </li> <li> <p>Calibration matrix: The calibration matrix was obtained using a pinhead-based method. The "scaling_from_pixel_to_mm" and "spatial_calibration_from_image_coordinate_system_to_tracking_tool_coordinate_system" are provided in the “calib_matrix.csv”. </p> </li> </ul> <div> <p><strong>Data Usage Policy:</strong></p> <ul> <li>The training and validation data provided may be utilized within the research scope of this challenge and in subsequent research-related publications. However, commercial use of the training and validation data is prohibited. In cases where the intended use is ambiguous, participants accessing the data are requested to abstain from further distribution or use outside the scope of this challenge.</li> <li>If you use our dataset in your publication, please cite the challenge paper and some of the following optional articles: <ul> <li>Challenge paper: <ul> <li><strong>Qi Li et al. "TUS-REC2024: A Challenge to Reconstruct 3D Freehand Ultrasound Without External Tracker." <em>arXiv preprint arXiv:<a title="https://arxiv.org/abs/2506.21765" href="https://doi.org/10.48550/arXiv.2506.21765" target="_blank" rel="noopener">2506.21765</a></em> (2025).</strong></li> </ul> </li> <li>Optional articles: <ul> <li>Qi Li, Ziyi Shen, Qianye Yang, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Nonrigid Reconstruction of Freehand Ultrasound without a Tracker." In <em>International Conference on Medical Image Computing and Computer-Assisted Intervention</em>, pp. 689-699. Cham: Springer Nature Switzerland, 2024. doi: <a href="https://doi.org/10.1007/978-3-031-72083-3_64" target="_blank" rel="noopener">10.1007/978-3-031-72083-3_64.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Long-term Dependency for 3D Reconstruction of Freehand Ultrasound Without External Tracker." IEEE Transactions on Biomedical Engineering, vol. 71, no. 3, pp. 1033-1042, 2024. doi: <a href="https://ieeexplore.ieee.org/abstract/document/10288201" target="_blank" rel="noopener">10.1109/TBME.2023.3325551</a>.</li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Trackerless freehand ultrasound with sequence modelling and auxiliary transformation over past and future frames." In 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), pp. 1-5. IEEE, 2023. doi: <a href="https://doi.org/10.1109/ISBI53787.2023.10230773" target="_blank" rel="noopener">10.1109/ISBI53787.2023.10230773.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Privileged Anatomical and Protocol Discrimination in Trackerless 3D Ultrasound Reconstruction." In International Workshop on Advances in Simplifying Medical Ultrasound, pp. 142-151. Cham: Springer Nature Switzerland, 2023. doi: <a href="https://doi.org/10.1007/978-3-031-44521-7_14" target="_blank" rel="noopener">https://doi.org/10.1007/978-3-031-44521-7_14.</a></li> </ul> </li> </ul> </li> </ul> </div> </div> </div> </div>
Trackerless 3D Freehand Ultrasound Reconstruction Challenge 2024 - Train Dataset (Part 1)
<blockquote> <p><strong>This Challenge will be an open-ended challenge, and we welcome your submission. Please register your team via this <a title="https://forms.office.com/e/dPg47ktV7M" href="https://forms.office.com/e/dPg47ktV7M" target="_blank" rel="noopener">form</a>. You can submit the algorithm via this <a title="https://forms.office.com/e/QChhNkLYiu" href="https://forms.office.com/e/QChhNkLYiu" target="_blank" rel="noopener noreferrer">form</a> for TUS-REC2024 Challenge, and we will test your submitted docker on the test set.</strong></p> <p><strong>We are organising TUS-REC2025 at MICCAI2025. More information is available on the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/" target="_blank" rel="noopener">TUS-REC2025 challenge website</a> and <a href="https://github.com/QiLi111/TUS-REC2025-Challenge_baseline" target="_blank" rel="noopener">Baseline code repo</a>.</strong></p> </blockquote> <p><strong>This is the first part of the Challenge train dataset. <a href="../doi/10.5281/zenodo.11180795" target="_blank" rel="noopener">Link</a> to second part; <a href="../doi/10.5281/zenodo.11355499" target="_blank" rel="noopener">Link</a> to third part. <a href="../doi/10.5281/zenodo.12979481" target="_blank" rel="noopener">Link</a> to validation dataset.</strong></p> <p>Acquisition devices and config: The 2D US images were acquired using an Ultrasonix machine (BK, Europe) with a curvilinear probe (4DC7-3/40). The associated position information of each frame was recorded by an optical tracker (NDI Polaris Vicra, Northern Digital Inc., Canada). The acquired US frames were recorded at 20 fps, with an image size of 480×640, without speckle reduction. The frequency was set at 6MHz with a dynamic range of 83 dB, an overall gain of 48% and a depth of 9 cm. </p> <div> <p>Scanning protocol: Both left and right forearms of volunteers were scanned. For each forearm, the US probe moves in three different trajectories (straight line shape, "C" shape, and "S" shape), in a distal-to-proximal direction followed by a proximal-to-distal direction, with the US plane perpendicular of and parallel to the scanning direction. The train dataset contains 1200 scans in total, 24 scans associated with each subject.</p> <p>For detailed information please refer to the <a href="https://github-pages.ucl.ac.uk/tus-rec-challenge/TUS-REC2024/" target="_blank" rel="noopener">Challenge website</a>. Baseline code is also provided, which can be found at this <a href="https://github.com/QiLi111/tus-rec-challenge_baseline" target="_blank" rel="noopener">repo</a>.</p> <p>Dataset structure: </p> </div> <div> <ul> <li> <p>The dataset contains 50 folders (one subject per folder), each with 24 scans. Each .h5 file corresponds to one scan, storing image and transformation of each frame within this scan. Key-value pairs in each .h5 file are explained below.</p> <ul> <li> <p>“frames” - All frames in the scan; with a shape of [N,H,W], where N refers to the number of frames in the scan, H and W denote the height and width of a frame. </p> </li> <li> <p>“tforms” - All transformations in the scan; with a shape of [N,4,4], where N is the number of frames in the scan, and the transformation matrix denotes the transformation from tracker tool space to camera space. </p> </li> <li> <p>Notations in the name of each .h5 file: “RH”: right arm; “LH”: left arm; “Per”: perpendicular; “Par”: parallel; “L”: straight line shape; “C”: C shape; “S”: S shape; “DtP”: distal-to-proximal direction; “PtD”: proximal-to-distal direction; For example, “RH_Per_L_DtP.h5” denotes a scan on the right forearm, with ultrasound probe perpendicular of the forearm sweeping along straight line, in distal-to-proximal direction.</p> </li> </ul> </li> <li> <p>Calibration matrix: The calibration matrix was obtained using a pinhead-based method. The "scaling_from_pixel_to_mm" and "spatial_calibration_from_image_coordinate_system_to_tracking_tool_coordinate_system" are provided in the “calib_matrix.csv”. </p> </li> </ul> <div> <p><strong>Data Usage Policy:</strong></p> <ul> <li>The training and validation data provided may be utilized within the research scope of this challenge and in subsequent research-related publications. However, commercial use of the training and validation data is prohibited. In cases where the intended use is ambiguous, participants accessing the data are requested to abstain from further distribution or use outside the scope of this challenge.</li> <li>If you use our dataset in your publication, please cite the challenge paper and some of the following optional articles: <ul> <li>Challenge paper: <ul> <li><strong>Qi Li et al. "TUS-REC2024: A Challenge to Reconstruct 3D Freehand Ultrasound Without External Tracker." <em>arXiv preprint arXiv:<a title="https://arxiv.org/abs/2506.21765" href="https://doi.org/10.48550/arXiv.2506.21765" target="_blank" rel="noopener">2506.21765</a></em> (2025).</strong></li> </ul> </li> <li>Optional articles: <ul> <li>Qi Li, Ziyi Shen, Qianye Yang, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Nonrigid Reconstruction of Freehand Ultrasound without a Tracker." In <em>International Conference on Medical Image Computing and Computer-Assisted Intervention</em>, pp. 689-699. Cham: Springer Nature Switzerland, 2024. doi: <a href="https://doi.org/10.1007/978-3-031-72083-3_64" target="_blank" rel="noopener">10.1007/978-3-031-72083-3_64.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Long-term Dependency for 3D Reconstruction of Freehand Ultrasound Without External Tracker." IEEE Transactions on Biomedical Engineering, vol. 71, no. 3, pp. 1033-1042, 2024. doi: <a href="https://ieeexplore.ieee.org/abstract/document/10288201" target="_blank" rel="noopener">10.1109/TBME.2023.3325551</a>.</li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Trackerless freehand ultrasound with sequence modelling and auxiliary transformation over past and future frames." In 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), pp. 1-5. IEEE, 2023. doi: <a href="https://doi.org/10.1109/ISBI53787.2023.10230773" target="_blank" rel="noopener">10.1109/ISBI53787.2023.10230773.</a></li> <li>Qi Li, Ziyi Shen, Qian Li, Dean C. Barratt, Thomas Dowrick, Matthew J. Clarkson, Tom Vercauteren, and Yipeng Hu. "Privileged Anatomical and Protocol Discrimination in Trackerless 3D Ultrasound Reconstruction." In International Workshop on Advances in Simplifying Medical Ultrasound, pp. 142-151. Cham: Springer Nature Switzerland, 2023. doi: <a href="https://doi.org/10.1007/978-3-031-44521-7_14" target="_blank" rel="noopener">https://doi.org/10.1007/978-3-031-44521-7_14.</a></li> </ul> </li> </ul> </li> </ul> </div> </div>
OATH Training Dataset
<p>Data from: http://tid.uio.no/plasma/oath/</p> <p>with regard to paper: <a href="https://doi.org/10.1029/2018JA025274">https://doi.org/10.1029/2018JA025274</a></p> <p> </p>
Dataset: Diamonds from Hadley's ggplot2 for Galaxy training
<p>Sample dataset created from https://doi.org/10.5281/zenodo.3522106 by selecting carat,price,color,clarity and cut columns only. In addition color and clarity are factors with integer values so we can reuse the dataset directly with an existing workflow (taught in Galaxy 101 for everyone).</p>
Fuτure - dataset for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons
<h1> Data description</h1> <h2>MC Simulation</h2> <p><br>The <strong>Fuτure</strong> dataset is intended for studies, development, and training of algorithms for reconstructing and identifying hadronically decaying tau leptons. The dataset is generated with Pythia 8, with the full detector simulation being performed by Geant4 with the CLIC-like detector setup CLICdet (CLIC_o3_v14) setup. Events are reconstructed using the Marlin reconstruction framework and interfaced with Key4HEP. Particle candidates in the reconstructed events are reconstructed using the PandoraPF algorithm.</p> <p>In this version of the dataset no γγ -> hadrons background is included.</p> <h2>Samples</h2> <p><br>This dataset contains e+e- samples with Z->ττ, ZH,H->ττ and Z->qq events, with approximately 2 million events simulated in each category.</p> <p>The following processes e+e- were simulated with Pythia 8 at sqrt(s) = 380 GeV:</p> <ul> <li>p8_ee_qq_ecm380 [Z -> qq events]</li> <li>p8_ee_ZH_Htautau [ZH -> Ztautau]</li> <li>p8_ee_Z_Ztautau_ecm380 [ZH -> Ztautau]</li> </ul> <p>The .root files from the MC simulation chain are eventually processed by the software found in <a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a> in order to create flat ntuples as the final product.</p> <h2><br>Features</h2> <p><br>The basis of the ntuples are the particle flow (PF) candidates from PandoraPF. Each PF candidate has four momenta, charge and particle label (electron / muon / photon / charged hadron / neutral hadron). The PF candidates in a given event are clustered into jets using generalized kt algorithm for ee collisions, with parameters p=-1 and R=0.4. The minimum pT is set to be 0 GeV for both generator level jets and reconstructed jets. The dataset contains the four momenta of the jets, with the PF candidates in the jets with the above listed properties.</p> <p>Additionally, a set of variables describing the tau lifetime are calculated using the software in <a href="https://github.com/HEP-KBFI/ml-tau-en-reg">Github</a>. As tau lifetime is very short, these variables are sensitive to true tau decays. In the calculation of these lifetime variables, we use a linear approximation.</p> <p>In summary, the features found in the flat ntuples are:</p> <p> </p> <table> <tbody> <tr> <td><strong>Name</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>reco_cand_p4s</td> <td>4-momenta per particle in the reco jet.</td> </tr> <tr> <td>reco_cand_charge</td> <td>Charge per particle in the jet.</td> </tr> <tr> <td>reco_cand_pdg</td> <td>PDGid per particle in the jet.</td> </tr> <tr> <td>reco_jet_p4s</td> <td>RecoJet 4-momenta.</td> </tr> <tr> <td>reco_cand_dz</td> <td>Longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dz_err</td> <td>Uncertainty of the longitudinal impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy</td> <td>Transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>reco_cand_dxy_err</td> <td>Uncertainty of the transverse impact parameter per particle in the jet. For future steps. Fill value used for neutral particles as no track parameters can be calculated.</td> </tr> <tr> <td>gen_jet_p4s</td> <td>GenJet 4-momenta. Matched with RecoJet within a cone of radius dR < 0.3.</td> </tr> <tr> <td>gen_jet_tau_decaymode</td> <td>Decay mode of the associated genTau. Jets that have associated leptonically decaying taus are removed, so there are no DM=16 jets. If no GenTau can be matched to GenJet within dR < 0.4, a fill value is used.</td> </tr> <tr> <td>gen_jet_tau_p4s</td> <td>Visible 4-momenta of the genTau. If no GenTau can be matched to GenJet within dR<0.4, a fill value is used.</td> </tr> </tbody> </table> <p>The ground truth is based on stable particles at the generator level, before detector simulation. These particles are clustered into generator-level jets and are matched to generator-level τ leptons as well as reconstructed jets. In order for a generator-level jet to be matched to generator-level τ lepton, the τ lepton needs to be inside a cone of dR = 0.4. The same applies for the reconstructed jet, with the requirement on dR being set to dR = 0.3. For each reconstructed jet, we define three target values related to τ lepton reconstruction:</p> <ul> <li> a binary flag <strong>isTau</strong> if it was matched to a generator-level hadronically decaying τ lepton. <strong>gen_jet_tau_decaymode</strong> of value -1 indicates no match to generator-level hadronically decaying τ.</li> <li> the categorical decay mode of the τ <strong>gen_jet_tau_decaymode</strong> in terms of the number of generator level charged and neutral hadrons. Possible <strong>gen_jet_tau_decaymode</strong> are {0, 1, . . . , 15}.</li> <li> if matched, the visible (neglecting neutrinos), reconstructable pT of the τ lepton. This is inferred from the <strong>gen_jet_tau_p4s</strong></li> </ul> <h2>Contents:</h2> <ul> <li>qq_test.parquet</li> <li>qq_train.parquet</li> <li>zh_test.parquet</li> <li>zh_train.parquet</li> <li>z_test.parquet</li> <li> z_train.parquet</li> <li>data_intro.ipynb</li> </ul> <h2>Dataset characteristics</h2> <p> </p> <table> <tbody> <tr> <td><strong>File</strong></td> <td><strong># Jets</strong></td> <td><strong>Size</strong></td> </tr> <tr> <td>z_test.parquet</td> <td> <pre>870 843</pre> </td> <td>171 MB</td> </tr> <tr> <td>z_train.parquet</td> <td> <pre>3 483 369</pre> </td> <td>681 MB</td> </tr> <tr> <td>zh_test.parquet</td> <td> <pre>1 068 606</pre> </td> <td>213 MB</td> </tr> <tr> <td>zh_train.parquet</td> <td> <pre>4 274 423</pre> </td> <td>851 MB</td> </tr> <tr> <td>qq_test.parquet</td> <td> <pre>6 366 715</pre> </td> <td>1.4 GB</td> </tr> <tr> <td>qq_train.parquet</td> <td> <pre>25 466 858</pre> </td> <td>5.6 GB</td> </tr> </tbody> </table> <p>The dataset consists of 6 files of 8.9 GB in total.</p> <h2>How can you use these data?</h2> <p>The .parquet files can be directly loaded with the Awkward Array Python library.<br>An example how one might use the dataset and the features is given in <strong>data_intro.ipynb</strong></p>
EyeOnWater training dataset for assessing the inclusion of water images
<h1>Training dataset</h1> <p>The EyeOnWater app is designed to assess the ocean's water quality using images captured by regular citizens. In order to have an extra helping hand in determining whether an image meets the criteria for inclusion in the app, the YOLOv8 model for image classification is employed. With the help of this model all uploaded pictures are assessed. If the model deems a water image unsuitable, it is excluded from the app's online database. In order to train this model a training dataset containing a large pool of different images is required. The dataset contains a total of 13,766 images, categorized into three distinct classes: “water_good,” “water_bad,” and “other.” The “water_good” class includes images that meet the requirements of EyeOnWater. The “water_bad” class comprises images of water that do not fulfill these requirements. Finally, the “other” class consists of miscellaneous images that users submitted, which do not depict water. This categorization enables precise filtering and analysis of images relevant to water quality assessment.</p>
FeM dataset – An iron ore labeled images dataset for segmentation training and testing
<p>This dataset is composed of 81 pairs of correlated images. Each pair contains one image of an iron ore sample acquired through reflected light microscopy (RGB, 24-bit), and the corresponding binary reference image (8-bit), in which the pixels are labeled as belonging to one of two classes: ore (0) or embedding resin (255).</p> <p>The sample came from an itabiritic iron ore concentrate from Quadrilátero Ferrífero (Brazil) mainly composed of hematite and quartz, with little magnetite and goethite. It was classified by size and concentrated with a dense liquid. Then, the fraction -149+105 μm with density greater than 3.2 was cold mounted with epoxy resin and subsequently ground and polished.</p> <p>Correlative microscopy was employed for image acquisition. Thus, 81 fields were imaged on a reflected light microscope with a 10× (NA 0.20) objective lens and on a scanning electron microscope (SEM). In sequence, they were registered, resulting in images of 999×756 pixels with a resolution of 1.05 µm/pixel. Finally, the images from SEM were thresholded to generate the reference images.</p> <p>Further description of this sample and its imaging procedure can be found in the work by Gomes and Paciornik (2012).</p> <p>This dataset was created for developing and testing deep learning models on semantic segmentation tasks. The paper of Filippo et al. (2021) presented a variant of the DeepLabv3+ model that reached mean values of 91.43% and 93.13% for overall accuracy and F1 score, respectively, for 5 rounds of experiments (training and testing), each with a different, random initialization of network weights.</p> <p>For further questions and suggestions, please do not hesitate to contact us.</p> <p> </p> <p><strong>Contact email</strong>: ogomes@gmail.com</p> <p> </p> <p>If you use this dataset in your own work, please cite this DOI: 10.5281/zenodo.5014700</p> <p> </p> <p>Please also cite this paper, which provides additional details about the dataset:</p> <p>Michel Pedro Filippo, Otávio da Fonseca Martins Gomes, Gilson Alexandre Ostwald Pedro da Costa, Guilherme Lucio Abelha Mota. <em>Deep learning semantic segmentation of opaque and non-opaque minerals from epoxy resin in reflected light microscopy images</em>. <strong>Minerals Engineering</strong>, Volume 170, 2021, 107007, https://doi.org/10.1016/j.mineng.2021.107007.</p> <p> </p>
Treatment Trains for Water Reclamation (Dataset)
<p>This dataset lists typical treatment trains (combination of unit processes) from global water reuse and reclamation practices - PDF</p>
Human Respiration Dataset - Training purposes
<p>This is a training dataset holding the heart rate (hr), blood oxygenation (ox), pulmonary pressure (pp), and the control variable (ctrl) or pushed per minute that a mechanical ventilator should perform under the given conditions.</p> <p>We include the results on a healthy patient with the respirator configured to update the working frequency every 30 seconds (Capturing 5 minutes total).</p>
GNSS training - dataset - Wrocław - 1st GATHERS Summer School
<p>The files in the collection contain a 3D time-series of displacements in [m]. Data were collected before 1st GATHERS Summer School in Wrocław as a reference data. The observations were taken at the roof of Wrocław University of Environmental and Life Sciences - Institute of Geodesy and Geoinformatics: <a href="https://goo.gl/maps/W94mr2UhJKn7XH1z9">Location in Google Maps</a></p> <p>Please note that each file is containing 3 time-series in North, East and Up direction.</p> <p>There are three headers beginning each of the time-series.</p> <p>The GNSS_PPP.ascii contains the results of kinematic PPP (Precise Point Positioning)</p> <p>The GNSS_VAD.ascii contains the results of a kinematic variometric approach to GNSS data processing (from VADASE software).</p> <p>The ACC_DISPL.ascii contains the displacements calculated from accelerations measured with GeoTiny!AC accelerometer.</p> <p>The GNSS antenna and the accelerometer were co-located with a distance of 10 cm and rigidly connected.</p> <p>If some questions arise, please send a message to:</p> <p><a href="mailto:jan.kaplon@upwr.edu.pl">jan.kaplon@upwr.edu.pl</a>, </p> <p> </p>
replicAnt - Plum2023 - Detection & Tracking Datasets and Trained Networks
<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks used for detection and tracking experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data have been generated with the open-source photgrammetry platform scAnt <a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>. All synthetic data has been generated with the associated replicAnt project available from <a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Two video datasets were curated to quantify detection performance; one in laboratory and one in field conditions. The laboratory dataset consists of top-down recordings of foraging trails of <em>Atta vollenweideri</em> (Forel 1893) leaf-cutter ants. The colony was collected in Uruguay in 2014, and housed in a climate chamber at 25°C and 60% humidity. A recording box was built from clear acrylic, and placed between the colony nest and a box external to the climate chamber, which functioned as feeding site. Bramble leaves were placed in the feeding area prior to each recording session, and ants had access to the recording area at will. The recorded area was 104 mm wide and 200 mm long. An OAK-D camera (OpenCV AI Kit: OAK-D, Luxonis Holding Corporation) was positioned centrally 195 mm above the ground. While keeping the camera position constant, lighting, exposure, and background conditions were varied to create recordings with variable appearance: The “base” case is an evenly lit and well exposed scene with scattered leaf fragments on an otherwise plain white backdrop. A “bright” and “dark” case are characterised by systematic over- or underexposure, respectively, which introduces motion blur, colour-clipped appendages, and extensive flickering and compression artefacts. In a separate well exposed recording, the clear acrylic backdrop was substituted with a printout of a highly textured forest ground to create a “noisy” case. Last, we decreased the camera distance to 100 mm at constant focal distance, effectively doubling the magnification, and yielding a “close” case, distinguished by out-of-focus workers. All recordings were captured at 25 frames per second (fps).<br> <br> The field datasets consists of video recordings of <em>Gnathamitermes</em> sp. desert termites, filmed close to the nest entrance in the desert of Maricopa County, Arizona, using a Nikon D850 and a Nikkor 18-105 mm lens on a tripod at camera distances between 20 cm to 40 cm. All video recordings were well exposed, and captured at 23.976 fps.<br> <br> Each video was trimmed to the first 1000 frames, and contains between 36 and 103 individuals. In total, 5000 and 1000 frames were hand-annotated for the laboratory- and field-dataset, respectively: each visible individual was assigned a constant size bounding box, with a centre coinciding approximately with the geometric centre of the thorax in top-down view. The size of the bounding boxes was chosen such that they were large enough to completely enclose the largest individuals, and was automatically adjusted near the image borders. A custom-written Blender Add-on aided hand-annotation: the Add-on is a semi-automated multi animal tracker, which leverages blender’s internal contrast-based motion tracker, but also include track refinement options, and CSV export functionality. Comprehensive documentation of this tool and Jupyter notebooks for track visualisation and benchmarking is provided on the <a href="https://github.com/evo-biomech/replicAnt"><em>replicAnt</em></a> and <a href="https://github.com/FabianPlum/blenderMotionExport">BlenderMotionExport</a> GitHub repositories.</p> <p><strong>Synthetic data generation</strong></p> <p>Two synthetic datasets, each with a population size of 100, were generated from 3D models of \textit{Atta vollenweideri} leaf-cutter ants. All 3D models were created with the <em>scAnt</em> photogrammetry workflow. A “group” population was based on three distinct 3D models of an ant minor (1.1 mg), a media (9.8 mg), and a major (50.1 mg) (see <a href="https://zenodo.org/record/7849059">10.5281/zenodo.7849059</a>)). To approximately simulate the size distribution of <em>A. vollenweideri </em>colonies, these models make up 20%, 60%, and 20% of the simulated population, respectively. A 33% within-class scale variation, with default hue, contrast, and brightness subject material variation, was used. A “single” population was generated using the major model only, with 90% scale variation, but equal material variation settings.<br> <br> A <em>Gnathamitermes</em> sp. synthetic dataset was generated from two hand-sculpted models; a worker and a soldier made up 80% and 20% of the simulated population of 100 individuals, respectively with default hue, contrast, and brightness subject material variation. Both 3D models were created in Blender v3.1, using reference photographs.<br> <br> Each of the three synthetic datasets contains 10,000 images, rendered at a resolution of 1024 by 1024 px, using the default generator settings as documented in the Generator_example level file (see documentation on <a href="https://github.com/evo-biomech/replicAnt">GitHub</a>). To assess how the training dataset size affects performance, we trained networks on 100 (“small”), 1,000 (“medium”), and 10,000 (“large”) subsets of the “group” dataset. Generating 10,000 samples at the specified resolution took approximately 10 hours per dataset on a consumer-grade laptop (6 Core 4 GHz CPU, 16 GB RAM, RTX 2070 Super).</p> <p><br> Additionally, five datasets which contain both real and synthetic images were curated. These “mixed” datasets combine image samples from the synthetic “group” dataset with image samples from the real “base” case. The ratio between real and synthetic images across the five datasets varied between 10/1 to 1/100.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College’s President’s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>
replicAnt - Plum2023 - Pose-Estimation Datasets and Trained Models
<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks used for semantic and instance segmentation experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data have been generated with the open-source photgrammetry platform scAnt <a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>. All synthetic data has been generated with the associated replicAnt project available from <a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Two pose-estimation datasets were procured. Both datasets used first instar <em>Sungaya nexpectata</em> (Zompro 1996) stick insects as a model species. Recordings from an evenly lit platform served as representative for controlled laboratory conditions; recordings from a hand-held phone camera served as approximate example for serendipitous recordings in the field. </p> <p>For the platform experiments, walking <em>S. inexpectata</em> were recorded using a calibrated array of five FLIR blackfly colour cameras (Blackfly S USB3, Teledyne FLIR LLC, Wilsonville, Oregon, U.S.), each equipped with 8 mm c-mount lenses (M0828-MPW3 8MM 6MP F2.8-16 C-MOUNT, CBC Co., Ltd., Tokyo, Japan). All videos were recorded with 55 fps, and at the sensors’ native resolution of 2048 px by 1536 px. The cameras were synchronised for simultaneous capture from five perspectives (top, front right and left, back right and left), allowing for time-resolved, 3D reconstruction of animal pose.<br> <br> The handheld footage was recorded in landscape orientation with a Huawei P20 (Huawei Technologies Co., Ltd., Shenzhen, China) in stabilised video mode: <em>S. inexpectata </em>were recorded walking across cluttered environments (hands, lab benches, PhD desks etc), resulting in frequent partial occlusions, magnification changes, and uneven lighting, so creating a more varied pose-estimation dataset.<br> <br> Representative frames were extracted from videos using DeepLabCut (DLC)-internal k-means clustering. 46 key points in 805 and 200 frames for the platform and handheld case, respectively, were subsequently hand-annotated using the DLC annotation GUI.</p> <p><strong>Synthetic data</strong></p> <p>We generated a synthetic dataset of 10,000 images at a resolution of 1500 by 1500 px, based on a 3D model of a first instar <em>S. inexpectata </em>specimen, generated with the <a href="https://peerj.com/articles/11155/"><em>scAnt</em> photogrammetry workflow</a>. Generating 10,000 samples took about three hours on a consumer-grade laptop (6 Core 4 GHz CPU, 16 GB RAM, RTX 2070 Super). We applied 70\% scale variation, and enforced hue, brightness, contrast, and saturation shifts, to generate 10 separate sub-datasets containing 1000 samples each, which were combined to form the full dataset.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College’s President’s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>
replicAnt - Plum2023 - Segmentation Datasets and Trained Models
<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks used for semantic and instance segmentation experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data have been generated with the open-source photgrammetry platform scAnt <a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>. All synthetic data has been generated with the associated replicAnt project available from <a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Semantic and instance segmentation is used only rarely in non-human animals, partially due to the laborious process of curating sufficiently large annotated datasets. <em>replicAnt </em>can produce pixel-perfect segmentation maps with minimal manual effort. In order to assess the quality of the segmentations inferred by networks trained with these maps, semi-quantitative verification was conducted using a set of macro-photographs of <em>Leptoglossus zonatus</em> (Dallas, 1852) and <em>Leptoglossus phyllopus</em> (Linnaeus, 1767), provided by Prof. Christine Miller (University of Florida), and Royal Tyler (Bugwood.org. For further qualitative assessment of instance segmentation, we used laboratory footage, and field photographs of <em>Atta vollenweideri</em> provided by Prof. Flavio Roces. More extensive quantitative validation was infeasible, due to the considerable effort involved in hand-annotating larger datasets on a per-pixel basis.</p> <p><strong>Synthetic data</strong></p> <p>We generated two synthetic datasets from a single 3D scanned <em>Leptoglossus zonatus</em> (Dallas, 1852) specimen: one using the default pipeline, and one with additional plant assets, spawned by three dedicated scatterers. The plant assets were taken from the Quixel library and include 20 grass and 11 fern and shrub assets. Two dedicated grass scatterers were configured to spawn between 10,000 and 100,000 instances; the fern and shrub scatterer spawned between 500 to 10,000 instances. A total of 10,000 samples were generated for each sub dataset, leading to a combined dataset comprising 20,000 image render and ID passes. The addition of plant assets was necessary, as many of the macro-photographs also contained truncated plant stems or similar fragments, which networks trained on the default data struggled to distinguish from insect body segments. The ability to simply supplement the asset library underlines one of the main strengths of <em>replicAnt</em>: training data can be tailored to specific use cases with minimal effort.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College’s President’s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.