Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11,687

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

11,687 results for “training”

Learn how ShareScore rates datasets ↗
zenodo44/100

PhasAGE Training School 1 - Computational prediction and databases of protein phase separation - PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Phase separation in diseases - LECTURE

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Computational prediction and databases of protein phase separation Overview- LECTURE

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Phase separation in virus-host interactions- LECTURE

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Structure and protein interactions of repeated and low complexity regions - LECTURE

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1-Protein aggregation prediction-PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Linear motifs identification and prediction - PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Computational prediction of intrinsic disorder in proteins-DisProt - PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

PhasAGE Training School 1 - Computational prediction of intrinsic disorder in proteins-MobiDB - PRACTICAL

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

SAGE Rejected Article Tracker Training Data

<p>The enclosed dataset shows metadata for ArXiv preprints uploaded to ArXiv in 2012.</p> <p>For each preprint, there are 2 rows of search data:</p> <ul> <li>ArXiv preprint metadata plus the CrossRef API data for the&nbsp;<em>correct</em>&nbsp;search result (which is the metadata for the published version of that preprint).</li> <li>ArXiv preprint metadata and the metadata for the top&nbsp;<em>incorrect</em>&nbsp;CrossRef API search result for the title and author-names associated with the preprint.</li> </ul> <p>ArXiv preprints are referred to as &#39;query&#39; documents and CrossRef documents are referred to as &#39;match&#39; documents.</p> <p>This dataset is created using the&nbsp;<a href="https://github.com/sagepublishing/rejected_article_tracker_pkg">SAGE Rejected Article Tracker</a>&nbsp;and is supplementary to that project. Similar custom datasets can be created using the SAGE Rejected Article Tracker with different parameters (e.g. different timeframes).</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Named-Entity Recognition for Modern Tibetan Newspapers: Tagset, Guidelines and Training Data

<p>This dataset, tagset and guidelines were the output&nbsp;of a six-month incubator project on the feasibility of developing Named-Entity Recognition (NER) for modern Tibetan, primarily for use with contemporary Tibetan-language newspapers and media published inside the PRC.&nbsp;The project was carried out by the Mongolian and Inner Asian Studies Unit at Cambridge University&rsquo;s Department of Social Anthropology. It was funded by an incubator grant from Cambridge Language Sciences. The project title was&nbsp;&ldquo;Named-Entity Recognition in Tibetan and Mongolian Newspapers.&rdquo; The Project PI was Dr Hildegard Diemberger (Cambridge),&nbsp;the Coordinator and Lead Author was Dr Robert Barnett (SOAS), and Senior Advisers were Dr Nathan Hill (SOAS), Dr Marieke Meelen (Cambridge), and Dr Thomas White (Cambridge).&nbsp;<br> <br> Although some forms of NER and other NLP procedures have been developed within China for modern Tibetan (see Liu, Nuo <em>et al</em>, 2011), the data underlying those initiatives have not been made publicly available and their findings cannot be tested or reproduced. Significant work on developing NLP for Tibetan has been carried out outside China, but has focused largely on classical Tibetan and religious texts (see Hill &amp; Garrett, Edward, 2017).&nbsp;</p> <p>The Cambridge incubator project therefore produced a tagset, guidelines and training data for developing NER for modern Tibetan, with a focus on historical and political analysis of contemporary newspapers, media and other public documents in Tibetan. We compiled 3.11m syllables of data in Tibetan extracted from articles downloaded from Chinese-language news aggregator sites within China, primarily tibet.cpc.people.com.cn and tibet.people.com.cn. From this data, we selected texts containing 280,000 syllables in Tibetan, grouped in 26,000 utterances/sentences (available on request). Using Lighttag, an online annotation site, we developed a tagset for NER consisting of 17 tags (and one for wrong segmentation if using segmented data). We annotated approximately 186,000 syllables, leading to 9,884 annotations. Of these, after discounting flawed data, we produced training data containing c.6,700&nbsp;annotations.&nbsp; We carried out the secondary, manual review offline (for our method of converting Lighttag&nbsp;data for offline review, see the attached report &ldquo;Using Spreadsheets to Review Annotations Offline.pdf&rdquo;), and found an error rate of 3.6%. The final total of reviewed annotations was 6,624.&nbsp;</p> <p>The dataset, tagset, guidelines and reports were developed and documented by Robert Barnett, with assistance from Tsering Samdrup, Dr Hill and Dr Meelen. Primary annotation was by Tsering Samdrup, assisted by Dr Barnett.<br> <br> The datasets published here include:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> <ol> <li>The <strong>tagseet guidelines and annotation manual</strong>, including the 17-tag tagset, guidelines, and recommendations&nbsp;(&quot;NER for Modern Tibetan-tagset and guidelines.pdf&quot;).</li> <li>The <strong>tagged training data </strong>in .csv format (&quot;Tibetan NER Training Data-tagged, reviewed wth context-v10-UTF-8.csv&quot;) and .xls format (&quot;Tibetan NER Training Data-tagged with context-v10-UTF-8.xlsx&quot;). This includes&nbsp;6,624&nbsp;reveiwed annotations, arranged according to the Tibetan alphabet&nbsp;together&nbsp;with the tags and context (utterance) for each annotation.</li> <li>The <strong>raw annotation results </strong>downloaded&nbsp;from Lighttag as .json files&nbsp;(&quot;Raw Training Data for NER in Modern Tibetan -Jobs2-11-JSON.zip&quot;) and as .xls files&nbsp;(&quot;Training Data for NER in Modern Tibetan -Jobs2-11-XLS.zip&quot;). These include&nbsp;10 &quot;tasks&quot; or datasets of articles scraped from Tibetan-language websites within Tibet.&nbsp;&nbsp;&nbsp;&nbsp;</li> <li>A <strong>guide to preparing Lighttag annotation results for manual review offline </strong>(&ldquo;Using Spreadsheets to Review Annotations Offline.pdf&rdquo;).</li> </ol> <p>The project&#39;s findings regarding the status of NER and NLP for vertical Mongolian are available at DOI: 10.5281/zenodo.5103499.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Ecore Metamodels and EcoreBERT Pre-trained Language Model

<p>This dataset contains ecore metamodels from the MAR dataset&nbsp;transformed into tree representations.&nbsp;The original dataset can be found here:&nbsp;<a href="http://mar-search.org/experiments/models20/">http://mar-search.org/experiments/models20/</a></p> <p>The data contained in this repository were used to conduct the experiments in the paper: <strong>Recommending Metamodel Concepts during Modeling Activities with Pre-Trained Language Models.&nbsp;</strong>Link to the paper:&nbsp;<a href="https://arxiv.org/abs/2104.01642">https://arxiv.org/abs/2104.01642</a></p> <p>The data are organized as follows:</p> <ul> <li>model : our model trained on the tree representations of metamodels with RoBERTa architecture.</li> <li>tokenizers : the byte-level BPE tokenizer we used to train our model.</li> <li>train : the training data separated into a training and validation set.</li> <li>test : the test data of all experiments conducted in the paper.</li> </ul> <p>This data repository is linked with the following Github repository containing our code:&nbsp;<a href="https://github.com/mweyssow/ecore-bert">https://github.com/martiwey/metamodel-concepts-bert</a></p>

opencc-by-4.0Apr 2021View details →
zenodo44/100

Treatment Trains for Water Reclamation (Dataset)

<p>This dataset lists typical treatment trains (combination of unit processes) from global water reuse and reclamation practices - PDF</p>

opencc-by-4.0Dec 2018View details →
zenodo44/100

Human Respiration Dataset - Training purposes

<p>This is a training dataset holding the heart rate (hr), blood oxygenation (ox), pulmonary pressure (pp), and the control variable (ctrl) or pushed per minute that a mechanical ventilator should perform under the given conditions.</p> <p>We include the results on a healthy patient with the respirator configured to update the working frequency every 30 seconds (Capturing 5 minutes total).</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

GNSS training - dataset - Wrocław - 1st GATHERS Summer School

<p>The files in the collection&nbsp;contain a 3D time-series of displacements in [m]. Data were collected before 1st GATHERS Summer School in Wrocław as a reference data. The observations were taken at the roof of Wrocław University of Environmental and Life Sciences - Institute of Geodesy and Geoinformatics:&nbsp;<a href="https://goo.gl/maps/W94mr2UhJKn7XH1z9">Location in Google Maps</a></p> <p>Please note that each file is containing 3 time-series in North, East and Up direction.</p> <p>There are three headers beginning each of the time-series.</p> <p>The GNSS_PPP.ascii contains the results of kinematic PPP (Precise Point Positioning)</p> <p>The GNSS_VAD.ascii contains the results of a kinematic variometric approach to GNSS data processing (from VADASE software).</p> <p>The ACC_DISPL.ascii contains the displacements calculated from accelerations measured with GeoTiny!AC accelerometer.</p> <p>The GNSS antenna and the accelerometer were co-located with a distance of 10 cm and rigidly connected.</p> <p>If some questions arise, please send a message to:</p> <p><a href="mailto:jan.kaplon@upwr.edu.pl">jan.kaplon@upwr.edu.pl</a>,&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Data for Project 'Feasibility, Usability and Acceptance of a Newly Developed Exergame-Based Training Concept for Older Adults with Mild Neurocognitive Disorder - A Pilot Randomized Controlled Trial'

<p>Data for Project &#39;Feasibility, Usability and Acceptance of a Newly Developed Exergame-Based Training Concept for Older Adults with Mild Neurocognitive Disorder - A Pilot Randomized Controlled Trial&#39; (trial&nbsp;registered at clinicaltrials.gov (<a href="https://clinicaltrials.gov/ct2/show/NCT04996654">NCT04996654</a>; date of registration: 11 July 2021), consisting&nbsp;of:</p> <p>(1) the&nbsp;original and complete data set for all primary outcomes (&#39;Data_Primary-Outcomes_Brain-IT-Pilot-Feasibility-RCT_for-publication.xlsx&#39;);</p> <p>(2) the original and complete data set for all secondary outcomes (&#39;Data_Secondary-Outcomes_Brain-IT-Pilot-Feasibility-RCT_for-publication.xlsx&#39;);</p> <p>(3) the&nbsp;original and complete data set for all other outcomes (i.e. baseline factors (demographic data, type of usual care interventions) and training heart rate; &#39;Data_Other-Outcomes_Brain-IT-Pilot-Feasibility-RCT_for-publication.xlsx&#39;);</p> <p>(4) folder including the raw and processed heart rate variability (HRV) and electroencephalography (EEG)&nbsp;data for all participants and measurements (HRV-and-EEG_raw-and-processed-data.zip);</p> <p>(5)&nbsp;a corresponding README file including (a) general information, (b) data and file overview, (c) sharing and access information, (d) methodological information, and (e) data-specific information.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

PAsCAL WP6 Pilot 2 Autonomous Driving Training

<p>This dataset&nbsp;was&nbsp;collected within the context of the PAsCAL research project between January 2022&nbsp;and February 2022&nbsp;at the ACI Vallelunga test circuit and premises&nbsp;in Rome, Italy. Subject of the pilot was a driving training for advanced ADAS systems and test driving of a Level-2+ autonomous vehicle on a test track, performing several different manoeuvres to test the capability of the ADAS systems.</p> <p>Some of the participants were subjected to a driving training for autonomous vehicles before they did the test drive, wherein they had to perform several difficult driving manoeuvres (such as on slippery ground or emergency braking). The purpose of this pilot was to observe whether a driving training improves the driver&#39;s capability to use ADAS systems and therefore operate the vehicle in a safer way. Depending on the pilot, it is recommended to adapt existing driving training for beginners, professionals and experienced drivers.</p> <p>In order to analyse the answers given to the questions, it is&nbsp;recommended to consult also the &quot;PAsCAL WP6 Pilots Surveys&quot; dataset, which contains all questions and possible answers.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

ShiftCrypt training data set for training ConforMine

<p>Data set containing the ShiftCrypt values of the proteins used for training ConforMine. This set is not to be confused the the Molecular Dynamics (MD) data set, also used for training of this model.&nbsp;</p> <p>The data set also contains a python script which recreates the filtering of sequences performed in the training steps of ConforMine, which discards all proteins for which no valid ShiftCrypt predictions were obtained.&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

DUDE competition train - validation - test splits ground truth

<p>This JSON file contains the ground truth annotations for the train and validation set of the DUDE competition (https://rrc.cvc.uab.es/?ch=23&amp;com=tasks) of ICDAR 2023 (https://icdar2023.org/).</p> <p>&nbsp;</p> <p><strong>V1.0.7&nbsp;release</strong>: 41454 annotations for 4974 documents (train-validation-test)</p> <pre>DatasetDict({ &nbsp; &nbsp; train: Dataset({ &nbsp; &nbsp; &nbsp; &nbsp; features: [&#39;docId&#39;, &#39;questionId&#39;, &#39;question&#39;, &#39;answers&#39;, &#39;answers_page_bounding_boxes&#39;, &#39;answers_variants&#39;, &#39;answer_type&#39;, &#39;data_split&#39;, &#39;document&#39;, &#39;OCR&#39;], &nbsp; &nbsp; &nbsp; &nbsp; num_rows: 23728 &nbsp; &nbsp; }) &nbsp; &nbsp; val: Dataset({ &nbsp; &nbsp; &nbsp; &nbsp; features: [&#39;docId&#39;, &#39;questionId&#39;, &#39;question&#39;, &#39;answers&#39;, &#39;answers_page_bounding_boxes&#39;, &#39;answers_variants&#39;, &#39;answer_type&#39;, &#39;data_split&#39;, &#39;document&#39;, &#39;OCR&#39;], &nbsp; &nbsp; &nbsp; &nbsp; num_rows: 6315 &nbsp; &nbsp; }) &nbsp; &nbsp; test: Dataset({ &nbsp; &nbsp; &nbsp; &nbsp; features: [&#39;docId&#39;, &#39;questionId&#39;, &#39;question&#39;, &#39;answers&#39;, &#39;answers_page_bounding_boxes&#39;, &#39;answers_variants&#39;, &#39;answer_type&#39;, &#39;data_split&#39;, &#39;document&#39;, &#39;OCR&#39;], &nbsp; &nbsp; &nbsp; &nbsp; num_rows: 11402 &nbsp; &nbsp; }) }) ++update on answer_type +++formatting change to answers_variants ++++stricter check on answer_variants &amp; rename annotations file <strong>+ blind test set (no ground truth answers provided) </strong>++ removed duplicates from test set:&nbsp; </pre> <blockquote> <p>&nbsp; &nbsp; &quot;92bd5c758bda9bdceb5f67c17009207b_ac6964cbdf483e765b6668e27b3d0bc4&quot;,</p> <p>&nbsp; &nbsp; &quot;6ee71a16d4e4d1dbd7c1f569a92d4e08_549f2a163f8ff3e9f0293cf59fdd98bc&quot;,</p> <p>&nbsp; &nbsp; &quot;e6f3855472231a7ca6aada2f8e85fe5a_827c03a72f2552c722f2c872fd7f74c3&quot;,</p> <p>&nbsp; &nbsp; &quot;e3eecd7cca5de11f1d17cd94ae6a8d77_6300df64e4cf6ba0600ac81278f68de2&quot;,</p> <p>&nbsp; &nbsp; &quot;107b4037df8127a92ee4b6ae9b5df8fb_d7a60e7a9fc0b27487ea39cd7f56f98e&quot;,</p> <p>&nbsp; &nbsp; &quot;300cc3900080064d308983f958141232_6a7cf1aad908d58a75ab8e02ddc856f4&quot;,</p> <p>&nbsp; &nbsp; &quot;fdd3308efacddb88d4aa6e2073f481d4_138cb868ecc804a63cc7a4502c0009b2&quot;,</p> <p>&nbsp; &nbsp; &quot;1f7de256ff1743d329a8402ba0d132e7_95b6e8758533a9817b9f20a958e7b776&quot;,</p> <p>&nbsp; &nbsp; &quot;4f399b8c526ffb6a2fd585a18d4ed5ec_51097231bc327c26c59a4fd8d3ff3069&quot;,</p> </blockquote> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

replicAnt - Plum2023 - Detection & Tracking Datasets and Trained Networks

<p>This dataset contains all recorded and hand-annotated as well as all synthetically generated data as well as representative trained networks used for detection and tracking experiments in the<em> replicAnt - generating annotated images of animals in complex environments using Unreal Engine</em> manuscript. Unless stated otherwise, all 3D animal models used in the synthetically generated data&nbsp;have been generated with the open-source photgrammetry platform scAnt&nbsp;<a href="http://peerj.com/articles/11155/">peerj.com/articles/11155/</a>.&nbsp;&nbsp;All synthetic data has been generated with the associated replicAnt project available from&nbsp;<a href="https://github.com/evo-biomech/replicAnt">https://github.com/evo-biomech/replicAnt</a>.</p> <p><strong>Abstract:</strong></p> <p>Deep learning-based computer vision methods are transforming animal behavioural research. Transfer learning has enabled work in non-model species, but still requires hand-annotation of example footage, and is only performant in well-defined conditions. To overcome these limitations, we created&nbsp;replicAnt, a configurable pipeline implemented in Unreal Engine 5 and Python, designed to generate large and variable training datasets on consumer-grade hardware instead. replicAnt&nbsp;places 3D animal models into complex, procedurally generated environments, from which automatically annotated images can be exported. We demonstrate that synthetic data generated with replicAnt can significantly reduce the hand-annotation required to achieve benchmark performance in common applications such as animal detection, tracking, pose-estimation, and semantic segmentation; and that it increases the subject-specificity and domain-invariance of the trained networks, so conferring robustness. In some applications, replicAnt may even remove the need for hand-annotation altogether. It thus represents a significant step towards porting deep learning-based computer vision tools to the field.</p> <p><strong>Benchmark data</strong></p> <p>Two video datasets were curated to quantify detection performance; one in laboratory and one in field conditions. The laboratory dataset consists of top-down recordings of foraging trails of <em>Atta vollenweideri</em> (Forel 1893) leaf-cutter ants. The colony was collected in Uruguay in 2014, and housed in a climate chamber at 25&deg;C and 60% humidity. A recording box was built from clear acrylic, and placed between the colony nest and a box external to the climate chamber, which functioned as feeding site. Bramble leaves were placed in the feeding area prior to each recording session, and ants had access to the recording area at will. The recorded area was 104 mm wide and 200 mm long. An OAK-D camera (OpenCV AI Kit: OAK-D, Luxonis Holding Corporation) was positioned centrally 195 mm above the ground. While keeping the camera position constant, lighting, exposure, and background conditions were varied to create recordings with variable appearance: The &ldquo;base&rdquo; case is an evenly lit and well exposed scene with scattered leaf fragments on an otherwise plain white backdrop. A &ldquo;bright&rdquo; and &ldquo;dark&rdquo; case are characterised by systematic over- or underexposure, respectively, which introduces motion blur, colour-clipped appendages, and extensive flickering and compression artefacts. In a separate well exposed recording, the clear acrylic backdrop was substituted with a printout of a highly textured forest ground to create a &ldquo;noisy&rdquo; case. Last, we decreased the camera distance to 100 mm at constant focal distance, effectively doubling the magnification, and yielding a &ldquo;close&rdquo; case, distinguished by out-of-focus workers. All recordings were captured at 25 frames per second (fps).<br> <br> The field datasets consists of video recordings of <em>Gnathamitermes</em>&nbsp;sp. desert termites, filmed close to the nest entrance in the desert of Maricopa County, Arizona, using a Nikon D850 and a Nikkor 18-105 mm lens on a tripod at camera distances between 20 cm to 40 cm. All video recordings were well exposed, and captured at 23.976 fps.<br> <br> Each video was trimmed to the first 1000 frames, and contains between 36 and 103 individuals. In total, 5000 and 1000 frames were hand-annotated for the laboratory-&nbsp;and field-dataset, respectively: each visible individual was assigned a constant size bounding box, with a centre coinciding approximately with the geometric centre of the thorax in top-down view. The size of the bounding boxes was chosen such that they were large enough to completely enclose the largest individuals, and was automatically adjusted near the image borders. A custom-written Blender Add-on aided hand-annotation: the Add-on is a semi-automated multi animal tracker, which leverages blender&rsquo;s internal contrast-based motion tracker, but also include track refinement options, and CSV export functionality. Comprehensive documentation of this tool and Jupyter notebooks for track visualisation and benchmarking is provided on the <a href="https://github.com/evo-biomech/replicAnt"><em>replicAnt</em></a>&nbsp;and <a href="https://github.com/FabianPlum/blenderMotionExport">BlenderMotionExport</a>&nbsp;GitHub repositories.</p> <p><strong>Synthetic data generation</strong></p> <p>Two synthetic datasets, each with a population size of 100, were generated from 3D models of \textit{Atta vollenweideri} leaf-cutter ants. All 3D models were created with the <em>scAnt</em>&nbsp;photogrammetry workflow. A &ldquo;group&rdquo; population was based on three distinct 3D models of an ant minor (1.1 mg), a media (9.8 mg), and a major (50.1 mg) (see <a href="https://zenodo.org/record/7849059">10.5281/zenodo.7849059</a>)). To approximately simulate the size distribution of <em>A. vollenweideri&nbsp;</em>colonies, these models make up 20%, 60%, and 20% of the simulated population, respectively. A 33% within-class scale variation, with default hue, contrast, and brightness subject material variation, was used. A &ldquo;single&rdquo; population was generated using the major model only, with 90% scale variation, but equal material variation settings.<br> <br> A <em>Gnathamitermes</em>&nbsp;sp. synthetic dataset was generated from two hand-sculpted models; a worker and a soldier made up 80% and 20% of the simulated population of 100 individuals, respectively with default hue, contrast, and brightness subject material variation. Both 3D models were created in Blender v3.1, using reference photographs.<br> <br> Each of the three synthetic datasets contains 10,000 images, rendered at a resolution of 1024 by 1024 px, using the default generator settings as documented in the Generator_example level file (see documentation on <a href="https://github.com/evo-biomech/replicAnt">GitHub</a>). To assess how the training dataset size affects performance, we trained networks on 100 (&ldquo;small&rdquo;), 1,000 (&ldquo;medium&rdquo;), and 10,000 (&ldquo;large&rdquo;) subsets of the &ldquo;group&rdquo; dataset. Generating 10,000 samples at the specified resolution took approximately 10 hours per dataset on a consumer-grade laptop (6 Core 4 GHz CPU, 16 GB RAM, RTX 2070 Super).</p> <p><br> Additionally, five datasets which contain both real and synthetic images were curated. These &ldquo;mixed&rdquo; datasets combine image samples from the synthetic &ldquo;group&rdquo; dataset with image samples from the real &ldquo;base&rdquo; case. The ratio between real and synthetic images across the five datasets varied between 10/1 to 1/100.</p> <p><strong>Funding</strong></p> <p>This study received funding from Imperial College&rsquo;s President&rsquo;s PhD Scholarship (to Fabian Plum), and is part of a project that has received funding from the European Research Council (ERC) under the European Union&rsquo;s Horizon 2020 research and innovation program (Grant agreement No. 851705, to David Labonte). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record