Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,139

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

2,139 results for “recognition”

Learn how ShareScore rates datasets ↗
zenodo40/100

FIG. 5 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.

FIG. 5. — The holotype of Ptisana soluta (Compton) Murdock & Perrie, comb. nov., stat. nov. (Compton 1674, Ignambi, 1914, BM[BM000787128]) showing how the lamina transitions from 3-pinnate proximally to 2-pinnate distally. CC BY The Trustees of the Natural History Museum, London.

opencc-by-4.0Feb 2023View details →
zenodo40/100

FIG. 3 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.

FIG. 3. — Ptisana soluta (Compton) Murdock & Perrie, comb. nov., stat. nov., field photos: A, frond with lamina 3-pinnate proximally and 2-pinnate distally. The frond at top-right is P. attenuata (Labill.) Murdock; B, stipes are greenish-brown at a distance; C, abaxial surface of costae and fertile lamina, showing transition from 3-pinnate to 2-pinnate; D, stipe greenish-brown and smooth; E, stipules around stipe bases. Photos: A-D, Leon Perrie from near Nouméa; E, Rémy Amice, from near Nouméa.

opencc-by-4.0Feb 2023View details →
zenodo40/100

FIG. 1 in A synopsis of Ptisana Murdock ferns (Marattiaceae) in New Caledonia based on sequence data and morphology with the recognition of a new vulnerable species, P. soluta (Compton) Murdock & Perrie, comb. nov., stat. nov.

FIG. 1. — Ptisana attenuata (Labill.) Murdock, field photos: A, 3-pinnate frond; B, stipes are dark at a distance; C, abaxial surface of costae and lamina, with synangia; D, stipe dark and wrinkled; E, divided stipules around stipe bases. Photos: Leon Perrie, from near Nouméa.

opencc-by-4.0Feb 2023View details →
dryad40/100

Data from: The CellPhe toolkit for cell phenotyping using time-lapse imaging and pattern recognition

<p>With phenotypic heterogeneity in whole cell populations widely recognised, the demand for quantitative and temporal analysis approaches to characterise single cell morphology and dynamics has increased. We present CellPhe, a pattern recognition toolkit for the unbiased characterisation of cellular phenotypes within time-lapse videos. CellPhe imports tracking information from multiple segmentation and tracking algorithms to provide automated cell phenotyping from different imaging modalities, including fluorescence. To maximise data quality for downstream analysis, our toolkit includes automated recognition and removal of erroneous cell boundaries induced by inaccurate tracking and segmentation. We provide an extensive list of features extracted from individual cell time series, with custom feature selection to identify variables that provide the greatest discrimination for the analysis in question. Using ensemble classification for accurate prediction of cellular phenotype and clustering algorithms for the characterisation of heterogeneous subsets, we validate and prove adaptability using different cell types and experimental conditions.</p>

opencc-zeroFeb 2023View details →
zenodo40/100

LivingNER corpus: Named entity recognition, normalization & classification of species, pathogens and food

<p><strong>LivingNER Gold Standard corpus (includes training, validation, test and background&nbsp;sets + MULTILINGUAL RESOURCES</strong>)</p><p>&nbsp;</p><p><strong>Please cite if you use this dataset:</strong></p><p>A. Miranda-Escalada, E. Farré-Maduell, S. Lima-López, D. Estrada, L. Gascó, M. Krallinger, Mention detection, normalization &amp; classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources, <i>Procesamiento del Lenguaje Natural</i> (2022)</p><p>@article{amiranda2022nlp, title={Mention detection, normalization \&amp; classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources}, author={Miranda-Escalada, Antonio and Farr{\'e}-Maduell, Eul{`a}lia and Lima-L{\'o}pez, Salvador and Estrada, Darryl and Gasc{\'o}, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, year={2022} }</p><p>&nbsp;</p><p><i><strong>1. Introduction</strong></i></p><p>The LivingNER Gold Standard corpus is a collection of<strong> 2000 clinical case reports</strong> covering a <strong>broad range of medical specialities</strong>, i.e. infectious diseases (including Covid-19 cases), cardiology, neurology, oncology, dentistry, pediatrics, endocrinology, primary care, allergology, radiology, psychiatry, ophthalmology, urology, internal medicine, emergency and intensive care medicine, tropical medicine, and dermatology <strong>annotated with species</strong> [SPECIES] (including <strong>living organisms</strong> and <strong>microorganisms</strong>) and <strong>infectious diseases</strong> [ENFERMEDAD] mentions. Species mentions include many <strong>pathogens</strong> and infectious agents, but also <strong>food</strong>, allergens, <strong>pets</strong> or other species, taxonomic groups and organisms of clinical relevance.&nbsp;</p><p>The &nbsp;LivingNER corpus has also annotations of mentions of <strong>humans</strong> (tag HUMAN), including the patients itself, <strong>family members</strong>, healhcare professionals or other persons mentioned in the case reports. Thus it can be useful to extract family history information of patients or information about the social and healthcare personal environment and interactions.</p><p>All mentions have been exhaustively manually mapped by experts to their corresponding <a href="https://www.ncbi.nlm.nih.gov/taxonomy"><strong>NCBI Taxonomy</strong></a> identifiers.&nbsp;</p><p>It was used for the&nbsp;<a href="https://temu.bsc.es/livingner/">LivingNER</a>&nbsp;Shared Task on pathogens and living beings detection and normalization in Spanish medical documents, which was celebrated as part of IberLEF 2022.</p><p>&nbsp;</p><p><i><strong>2. Training, validation, test and background sets</strong></i></p><p>The training set is composed of 1000 clinical case reports. The validation set includes 500 clinical case reports with the same characteristics and the test set includes 485. The background set is a collection of around 13k unannotated case reports that were originally added to prevent manual annotations in the test set during the competition and to create a Silver Standard.</p><p><i><strong>2.1 Annotations format</strong></i></p><p>Annotations and text files are distributed separately. The texts are in plain text (.txt in UTF-8) format, while the annotations are are distributed in a tab-separated file&nbsp;(.tsv) file with one row per annotation:</p><p>- For&nbsp;<strong>subtask 1 (LivingNER-Species NER track)</strong>, the .tsv file has the following columns:</p><ul><li>filename: document name</li><li>mark: identifier mention mark</li><li>label: mention type (SPECIES or HUMAN)</li><li>off0: starting&nbsp;position of the mention in the document</li><li>off1: ending position of the mention in the document</li><li>span: textual span</li></ul><p>&nbsp;- For <strong>subtask 2 (LivingNER-Species Norm track)</strong>, the .tsv file has the same columns as the previous one, plus:</p><ul><li>isH: whether the span is narrower than the&nbsp;NCBITax assigned code</li><li>isN: whether the mention corresponds to a nosocomial infection</li><li>iscomplex: whether the span has assigned a combination of NCBITax&nbsp;codes</li><li>NCBITax: mention code in the&nbsp;NCBI Taxonomy</li></ul><p>- For&nbsp;<strong>subtask 3 (LivingNER-Clinical IMPACT track)</strong>,&nbsp;the .tsv file has the following columns:</p><ul><li>filename</li><li>isPet (Yes/No)</li><li>PetIDs (NCBITaxonomy codes of pet &amp; farm animals present in document)</li><li>isAnimalInjury (Yes/No)</li><li>AnimalInjuryIDs (NCBITaxonomy codes of animals causing injuries present in document)</li><li>IsFood (Yes/No)</li><li>FoodIDs (NCBITaxonomy codes of food mentions present in document)</li><li>isNosocomial (Yes/No)</li><li>NosocomialIDs (NCBITaxonomy codes of nosocomial species mentions present in document)</li></ul><p><i><strong>2.2 Important notes about subtask 3 (LivingNER-Clinical IMPACT track):</strong></i></p><ul><li><strong>Less clinical case reports</strong>. Subtask&nbsp;3 (LivingNER-Clinical IMPACT track) contains half of the clinical case reports&nbsp;(500 in the training partition, 250 in the validation partition). The list of valid&nbsp;clinical case reports for task 3 is included in the data&nbsp;(train_files_task3.txt and validation_files_task3.txt)</li><li><strong>Enriched dataset.</strong>&nbsp;The GS format is the one described above (a TSV with one line per clinical case report). However, we believe participants may find useful and&nbsp;<strong>enriched dataset.&nbsp;</strong>Then, we provide an additional dataset, with the mentions of the NER track classified in the 4 Clinical impact categories (food, pet&amp;farm animals, animals causing injuries and nosocomial). It is a TSV file with one row per annotation, and with the following columns:&nbsp;filename,&nbsp;mark,&nbsp;label,&nbsp;off0,&nbsp;off1,&nbsp;span,&nbsp;isPet,&nbsp;isAnimalInjury,&nbsp;isFood,&nbsp;isNosocomial,&nbsp;isH,&nbsp;iscomplex,&nbsp;code</li></ul><p>&nbsp;</p><p><i><strong>3. Multilingual resources</strong></i></p><p>We have generated the annotated training and validation sets in <strong>7 languages</strong>:</p><ul><li><i><strong>English</strong></i></li><li><i><strong>Portuguese</strong></i></li><li><i><strong>Catalan</strong></i></li><li><i><strong>Galician</strong></i></li><li><i><strong>Italian</strong></i></li><li><i><strong>French</strong></i></li><li><i><strong>Romanian</strong></i></li></ul><p>&nbsp;</p><p>The process was:</p><ol><li>The&nbsp;text files were translated with a neural machine translation system.</li><li>The annotations were translated with the same&nbsp;neural machine translation system.</li><li>The translated annotations were transferred to the translated&nbsp;text files using an annotation transfer technology.</li></ol><p>The&nbsp;text files are stored in the&nbsp;multilingual_resources/<strong>training-text-files</strong> and&nbsp;multilingual_resources/<strong>validation-text-files </strong>subfolders.</p><p>The annotated TSV files are stored in the&nbsp;multilingual_resources/<strong>annotation_transfer </strong>subfolder.</p><p>For the sake of comparison, we incorporate as well the annotations that resulted from the <a href="https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-11-85">LINNAEUS tool</a>&nbsp;in the&nbsp;multilingual_resources/<strong>linneaus</strong>&nbsp;subfolder.</p><p>If you want to visualize the multilingual resources, check out this Brat server:&nbsp;<a href="https://temu.bsc.es/mLivingNER/#/translations/">https://temu.bsc.es/mLivingNER/#/translations/</a></p><p>For instance, you can see the parallel annotations in <a href="https://temu.bsc.es/mLivingNER/diff.xhtml#/translations/en/annotation_transfer/train/casos_clinicos_cardiologia34?diff=/translations/fr/annotation_transfer/train/">English vs&nbsp;in French</a>, or <a href="https://temu.bsc.es/mLivingNER/diff.xhtml#/translations/cat/annotation_transfer/train/casos_clinicos_cardiologia35?diff=/gold-standard/train/">in Spanish (the gold standard) vs in Catalan.</a></p><p>&nbsp;</p><p><strong>Resources</strong></p><ul><li><a href="https://temu.bsc.es/livingner/"><strong>Task Web</strong></a></li><li><strong>Citation:&nbsp;</strong>A. Miranda-Escalada, E. Farré-Maduell, S. Lima-López, D. Estrada, L. Gascó, M. Krallinger, Mention detection, normalization &amp; classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources, <i>Procesamiento del Lenguaje Natural</i> (2022)</li><li><a href="https://doi.org/10.5281/zenodo.6385162"><strong>Annotation guidelines</strong></a></li><li><a href="https://github.com/tonifuc3m/livingner-evaluation-library"><strong>Evaluation library</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.6390506">LivingNER terminology</a></li><li><a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6444"><strong>Overview paper</strong></a></li><li><a href="https://ceur-ws.org/Vol-3202/"><strong>Proceedings participant papers</strong></a></li><li><a href="https://www.youtube.com/watch?v=8VcZw8ywyJY&amp;list=PL5uSCzf1azhA_gMLC3DBZe6NvmMJiggTg"><strong>Youtube videos</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/mention-detection-normalization-classification-of-species-pathogens-humans-and-food-in-clinical-documents-overview-of-the-livingner-shared-task-and-resources-talk-at-iberlef-sepln-2022"><strong>LivingNER overview talk sides at IberLEF/SEPLN</strong></a></li></ul><p>&nbsp;</p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in SympTEMIST, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/8413866">SympTEMIST</a> (Corpus of symptoms, sign and findings mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), differentdocument collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, different document collection, some overlapping documents)</li></ul>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Datasets and codes for De Lorm et al. 2023: Optimising the automated recognition of individual animals to support population monitoring

<p>Reliable estimates of population size and demographic rates are central to assessing the status of threatened species. However, obtaining individual-based demographic rates requires long-term data, which is often costly and difficult to collect. Photographic data offer an inexpensive, non-invasive method for individual-based monitoring of species with unique markings, and could therefore increase&nbsp;available demographic data for many species.&nbsp;However, selecting suitable images and identifying individuals from&nbsp;photographic&nbsp;catalogues is prohibitively time-consuming. Automated identification software can significantly speed up this process. Nevertheless, automated methods for selecting suitable images are lacking, as are studies comparing the performance of the most prominent identification software packages.</p> <p>&nbsp;</p> <p>In this study, we develop a framework that automatically selects images suitable for individual identification, and compare the performance of three commonly used identification software packages; Hotspotter, I<sup>3</sup>S-Pattern, and WildID. As a case study, we consider the African wild dog&nbsp;<em>Lycaon pictus</em>, a species whose conservation is limited by a lack&nbsp;of cost-effective large-scale monitoring. To evaluate intra-specific variation in the performance of software packages, we compare&nbsp;identification accuracy&nbsp;between two populations (in Kenya and Zimbabwe) that have markedly different coat colouration patterns.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p>The process of selecting suitable images was automated using Convolutional Neural Nets that crop individuals from images, filter out unsuitable images, separate left and right flanks, and remove image backgrounds. Hotspotter had the highest image-matching accuracy for both populations. However, the accuracy was significantly lower for the Kenyan population (62%), compared to the Zimbabwean population (88%).&nbsp;</p> <p>&nbsp;</p> <p>Our automated image pre-processing has immediate application for expanding monitoring based on image-matching. However, the difference in accuracy between populations highlights that population-specific detection rates are likely and may influence certainty in derived statistics. For species such as the African wild dog, where monitoring is both challenging and expensive, automated individual recognition could greatly expand and expedite conservation efforts.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

An Occlusion and Pose Sensitive Image Dataset for Black Ear Recognition

<p><strong>RESEARCH APPROACH</strong></p> <p>The research approach adopted for the study consists of seven phases which includes as shown in Figure 1:</p> <ol> <li>Pre-acquisition</li> <li>data pre-processing</li> <li>Raw images collection</li> <li>Image pre-processing</li> <li>Naming of images</li> <li>Dataset&nbsp;Repository</li> <li>Performance Evaluation</li> </ol> <p>The different phases in the study are discussed in the sections below.</p> <p>&nbsp;</p> <p><strong>PRE-ACQUISITION</strong></p> <p>The volunteers are given brief orientation on how their data will be managed and used for research purposes only. After the volunteers agrees, a consent form is given to be read and signed. The sample of the consent form filled by the volunteers is shown in Figure 1.</p> <p>The capturing of images was started with the setup of the imaging device. The camera is set up on a tripod stand in stationary position at the height 90 from the floor and distance 20cm from the subject.</p> <p>&nbsp;</p> <p><strong>EAR </strong><strong>IMAGE ACQUISITION</strong></p> <p>Image acquisition is an action of retrieving image from an external source for further processing. The image acquisition is purely a hardware dependent process by capturing unprocessed images of the volunteers using a professional camera. This was acquired through a subject posing in front of the camera. It is also a process through which digital representation of a scene can be obtained. This representation is known as an image and its elements are called pixels (picture elements). The imaging sensor/camera used in this study is a Canon E0S 60D professional camera which is placed at a distance of 3 feet form the subject and 20m from the ground.&nbsp;</p> <p>This is the first step in this project to achieve the project&rsquo;s aim of developing an occlusion and pose sensitive image dataset for black ear recognition. (OPIB ear dataset). To achieve the objectives of this study, a set of black ear images were collected mostly from undergraduate students at a public University in Nigeria.</p> <p>&nbsp;</p> <p>The image dataset required is captured in two scenarios:</p> <p>1. uncontrolled environment with a surveillance camera</p> <ol> </ol> <p>The image dataset captured is purely black ear with partial occlusion in a constrained and unconstrained environment.</p> <p>&nbsp;</p> <p>2. controlled environment with professional cameras</p> <p>The ear images captured were from black subjects in controlled environment. To make the OPIB dataset pose invariant, the volunteers stand on a marked positions on the floor indicating the angles at which the imaging sensor was captured the volunteers&rsquo; ear. The capturing of the images in this category requires that the subject stand and rotates in the following angles 60<sup>o</sup>, 30<sup>o</sup> and 0<sup>o</sup> towards their right side to capture the left ear and then towards the left to capture the right ear (Fernando <em>et al.,</em> 2017) as shown in Figure 4. Six (6) images were captured per subject at angles 60<sup>o</sup>, 30<sup>o</sup> and 0<sup>o</sup> for the left and right ears of 152 volunteers making a total of 907 images <strong><em>(five volunteers had 5 images instead of 6, hence f</em></strong><strong><em>olders 34, 22, 51, 99 and&nbsp;102 contain 5 images).</em></strong></p> <p>To make the OPIB dataset occlusion and pose sensitive, partial occlusion of the subject&rsquo;s ears were simulated using rings, hearing aid, scarf, earphone/ear pods, etc. before the images are captured.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td> <p><strong>CONSENT FORM</strong></p> <p>This form was designed to obtain participant&rsquo;s consent on the project titled: <strong>An Occlusion and Pose Sensitive Image Dataset for Black Ear Recognition</strong><strong> (OPIB)</strong>. The information is purely needed for academic research purposes and the ear images collected will curated anonymously and the identity of the volunteers will not be shared with anyone. The images will be uploaded on online repository to aid research in ear biometrics.</p> <p>The participation is voluntary, and the participant can withdraw from the project any time before the final dataset is curated and warehoused.</p> <p>Kindly sign the form to signify your consent.</p> <p><strong><em>I consent to my image being recorded in form of still images or video surveillance as part of the OPIB ear images project.</em></strong></p> <p><strong>Tick as appropriate:</strong></p> <p><strong>GENDER</strong> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Male &nbsp;&nbsp; Female</p> <p><strong>AGE</strong> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; (18-25)&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; (26-35)&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; (36-50)</p> <p>&nbsp;</p> <p>&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;&hellip;..</p> <p>SIGNED</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Figure 1</strong>: Sample of Subject&rsquo;s Consent Form for the OPIB ear dataset</p> <p>&nbsp;</p> <p><strong>RAW IMAGE COLLECTION</strong></p> <p>The ear images were captured using a digital camera which was set to JPEG because if the camera format is set to raw, no processing will be applied, hence the stored file will contain more tonal and colour data. However, if set to JPEG, the image data will be processed, compressed and stored in the appropriate folders.</p> <p>&nbsp;</p> <p><strong>IMAGE PRE-PROCESSING </strong></p> <p>The aim of pre-processing is to improve the quality of the images with regards to contrast, brightness and other metrics. It also includes operations such as: cropping, resizing, rescaling, etc. which are important aspect of image analysis aimed at dimensionality reduction. The images are downloaded on a laptop for processing using MATLAB.</p> <p>&nbsp;</p> <p><strong>Image Cropping</strong></p> <p>The first step in image pre-processing is image cropping. Some irrelevant parts of the image can be removed, and the image Region of Interest (ROI) is focused. This tool provides a user with the size information of the cropped image. MATLAB function for image cropping realizes this operation interactively by waiting for a user to specify the crop rectangle with the mouse and operate on the current axes. The output images of the cropping process are of the same class as the input image.</p> <p><strong>Naming of OPIB Ear Images</strong></p> <p>The OPIB ear images were labelled based on the naming convention formulated from this study as shown in Figure 5. The images are given unique names that specifies the subject, the side of the ear (left or right) and the angle of capture. The first and second letters (SU) in the image names is block letter simply representing subject for subject 1-to-n in the dataset, while the left and right ears is distinguished using L1, L2, L3 and R1, R2, R3 for angles 60<sup>0</sup>, 30<sup>0</sup> and 0<sup>0</sup><sub>, </sub>respectively as shown in Table 1.</p> <p>&nbsp;</p> <p><strong>Table 1: Naming Convention for OPIB ear images</strong></p> <table align="center"> <tbody> <tr> <td> <p>NAMING CONVENTION</p> </td> </tr> <tr> <td> <p>Label</p> <p>Degrees&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 60<sup>0</sup>&nbsp;&nbsp;&nbsp;&nbsp; 30<sup>0</sup>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 0<sup>0</sup></p> </td> </tr> <tr> <td> <p>No of the degree&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 1&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 2&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 3</p> </td> </tr> <tr> <td> <p>Subject 1&nbsp;&nbsp; indicates&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; (first image in dataset) SU<sub>1</sub></p> </td> </tr> <tr> <td> <p>Subject n&nbsp;&nbsp; indicates&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; (last image in dataset) SU<sub>n</sub></p> </td> </tr> <tr> <td> <p>Left Image 1&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; L 1</p> <p>Left image n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; L n</p> <p>Right Image 1&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; R 1</p> <p>Right Image n&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; R n</p> </td> </tr> <tr> <td> <p>SU1L<sub>1</sub>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; SU1R<sub>I</sub></p> <p>SU1L<sub>2</sub>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; SU1R<sub>2</sub>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> <p>SU1L<sub>3</sub>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; SU1R<sub>3</sub></p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>OPIB EAR DATASET EVALUATION</strong></p> <p>The prominent challenges with the current evaluation practices in the field of ear biometrics are the use of different databases, different evaluation matrices, different classifiers that mask the feature extraction performance and the time spent developing framework (Abaza <em>et al.</em>, 2013; Emer&scaron;ič <em>et al.,</em> 2017).</p> <p>The toolbox provides environment in which the evaluation of methods for person recognition based on ear biometric data is simplified. It executes all the dataset reads and classification based on ear descriptors.</p> <p>&nbsp;</p> <p><strong>DESCRIPTION OF OPIB EAR DATASET</strong></p> <p>OPIB ear dataset was organised into a structure with each folder containing 6 images of the same person. The images were captured with both left and right ear at angle 0, 30 and 60 degrees. The images were occluded with earing, scarves and headphone etc. &nbsp;The collection of the dataset was done both indoor and outdoor.&nbsp; The dataset was gathered through the student at a public university in Nigeria. The percentage of female (40.35%) while Male (59.65%).&nbsp; The ear dataset was captured through a profession camera Nikon D 350. It was set-up with a camera stand where an individual captured in a process order. A total number of 907 images was gathered.</p> <p>The challenges encountered in term of gathering students for capturing, processing of the images and annotations. The volunteers were given a brief orientation on what their ear could be used for before, it was captured, for processing.&nbsp; It was a great task in arranging the ear (dataset) into folders and naming accordingly.</p> <p>&nbsp;</p> <p><strong>Table 2</strong>: Overview of the OPIB Ear Dataset</p> <table align="left"> <tbody> <tr> <td> <p>Location</p> </td> <td> <p>Both Indoor and outdoor environment</p> </td> </tr> <tr> <td> <p>Information about Volunteers</p> </td> <td> <p>Students</p> </td> </tr> <tr> <td> <p>Gender</p> </td> <td> <p>Female (40.35%) and male (59.65%)</p> </td> </tr> <tr> <td> <p>Head Side Left and Right</p> </td> <td> <p>Side Left and Right</p> </td> </tr> <tr> <td> <p>Total number of volunteers</p> </td> <td> <p>152</p> </td> </tr> <tr> <td> <p>Per Subject images</p> </td> <td> <p>3 images of left ear and 3 images of right ear</p> </td> </tr> <tr> <td> <p>Total Images</p> </td> <td> <p>907</p> </td> </tr> <tr> <td> <p>Age group</p> </td> <td> <p>18 to 35 years</p> </td> </tr> <tr> <td> <p>Colour Representation</p> </td> <td> <p>RGB</p> </td> </tr> <tr> <td> <p>Image Resolution</p> </td> <td> <p>224x224</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Appendix: Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition

<p>Appendix tables for the paper &quot;Informative Speech Features based on Emotion Classes and Gender in Explainable Speech Emotion Recognition&quot;.</p> <p>Feature informativeness information was gathered using SHAP values.</p> <p>TABLE VIII: Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE IX: (continue)Table of Statistics and Individual t-Test Results Between Each Emotion vs Neutral</p> <p>TABLE X: Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XI: (continue) Table of Statistics and Individual t-Test Results Between Genders</p> <p>TABLE XII: Table of 5 the most informative feature for each model, according to SHAP values</p> <p>TABLE XIII: (continue)Table of 5 the most informative feature for each model, according to SHAP values</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Brain-inspired multimodal hybrid neural network for robot place recognition

<p>Brain-inspired multimodal hybrid neural network for robot place recognition</p>

openmit-licenseApr 2023View details →
zenodo40/100

Dataset for the paper "An end-to-end deep learning framework for wideband signal recognition"

<p>This dataset comprises a collection of synthetically generated wideband signals, which were used in experiments conducted for the paper &quot;An end-to-end deep learning framework for wideband signal recognition,&quot; submitted for publication to IEEE <em>Access</em>. In this work, the proposed learning-based approach for signal detection, localization, and classification was evaluated on the public wideband signal recognition dataset introduced by West <em>et al.</em> (<a href="https://ieeexplore.ieee.org/document/9593265">https://ieeexplore.ieee.org/document/9593265</a>). We note that the dataset provided here contains only additional auxiliary data that were generated to complement the public benchmark dataset in experiments that employ data mixing and transfer learning techniques. As such, the synthetic wideband signals in our dataset are designed to mimic those in the benchmark dataset.<br> <br> Specifically, in our simulations, a wideband signal is modeled as a superposition of several narrowband emissions and receiver noise. Similar to the benchmark dataset, all wideband signals have a normalized sampling frequency of <span class="math-tex">\(F_\text{s} = 1\)</span>&nbsp;sample per second (sps), duration of <span class="math-tex">\(L = 1 000 000\)</span> samples, and a center frequency of <span class="math-tex">\(F_\text{c} = 0\)</span> Hz. The narrowband signals have varying center frequencies, bandwidths, start times, and durations, and are modulated with a range of modulation classes, including:</p> <ol> <li>Amplitude Modulation - Double Sideband (AM-DSB)</li> <li>Amplitude Modulation - Single Sideband (AM-SSB)</li> <li>Frequency Modulation (FM)</li> <li><span class="math-tex">\(M\)</span>-Frequency-Shift Keying (<span class="math-tex">\(M\)</span>-FSK) for&nbsp;<span class="math-tex">\(M \in \{2, 4\}\)</span></li> <li><span class="math-tex">\(M\)</span>-Continuous Phase Frequency-Shift Keying (<span class="math-tex">\(M\)</span>-CPFSK) for&nbsp;<span class="math-tex">\(M \in \{2, 4\}\)</span></li> <li>Gaussian Minimum Shift Keying (GMSK)</li> <li>On-Off Keying (OOK)</li> <li><span class="math-tex">\(M\)</span>-Phase-Shift Keying (<span class="math-tex">\(M\)</span>-PSK) for <span class="math-tex">\(M \in \{2, 4, 8\}\)</span></li> <li><span class="math-tex">\(M\)</span>-Quadrature Amplitude Modulation (<span class="math-tex">\(M\)</span>-QAM) for <span class="math-tex">\(M \in \{16, 64, 256\}\)</span></li> </ol> <p>The modulation parameters for each narrowband signal are randomly selected from predefined typical ranges. For all wideband signals, the value of the power spectral density of the added white Gaussian noise has been initially set to&nbsp;-174 dBm/Hz. Further noise addition according to the desired&nbsp;signal-to-noise ratio (SNR) level can be applied by the user.&nbsp;<br> <br> The data are stored and documented according to the Signal Metadata Format (SigMF) standard (<a href="https://github.com/sigmf/SigMF">https://github.com/sigmf/SigMF</a>). Each wideband signal in the dataset&nbsp;is represented by a data file (a binary file containing digital I/Q samples of the signal) and a metadata file (in JSON format) that annotates the signal. The annotations include the general properties of the wideband recording (sampling&nbsp;rate, center frequency, signal duration, and noise power spectral density), as well as specific information for each narrowband signal (start time, duration, lowest and highest frequency, modulation class label, and power). For every signal in the dataset, the data file and the annotations file share the same name. For example, the first training wideband signal&nbsp;is stored in the file &quot;train_1.sigmf-data&quot; and is annotated by the file &quot;train_1.sigmf-meta&quot;.</p>

opencc-by-sa-4.0Apr 2023View details →
zenodo40/100

SPVPANELEX: Dataset containing aerial orthoimages (covering 257.93 km2 of the Spanish territory, with a spatial resolution of 0.5 m) labelled with photovoltaic panel information for binary recognition and semantic segmentation

<p>The data have been generated using scripts developed in Python with Open-Source libraries (GDAL/OGR and MapScript) to rasterize of vector cartography representing the photovoltaic (PV) panels instalations in urban, industrial, and rural areas. This PV panels cartography has been generated by manual digitalizing the PV panels found latest aerial orthofotographs available on June 1, 2021 from Plano Nacional de Ortofotograf&iacute;a A&eacute;rea (PNOA), produced by the National Geographic Institute of Spain, using the Web Map Service PNOA-MA.<br> <br> The dataset consists of 239,680 images of 256 &times; 256 pixels in size, in png format, labelled with Class_1: &ldquo;Contains PV panel&rdquo; and Class_2: &ldquo;Does not contain PV panel&rdquo;, that were pre-divided with a split criterion of 70:10:20%. in train, validation and test folders, respectively.<br> <br> The structure of the data is as follows:<br> 1-Panels-Ortho and 1-Panels-Mask contain the images featuring PV panels and their corresponding ground truth mask for training the semantic segmentation networks.<br> 1-Panels-Ortho and 2-NoPanels-Ortho contain images containing and not containing PV panels, for the training of binary recognition models of PV panels.<br> <br> Moreover, in each folder the structure is the same: train, test, validation containing 70%, 10% and 20% of the total images and masks of each type.<br> <br> 1-Panels-Ortho<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation<br> <br> 1-Panels-Mask<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation<br> <br> 2-NoPanels-Ortho<br> &nbsp; &nbsp; |----Train<br> &nbsp; &nbsp; |----Test<br> &nbsp; &nbsp; -----Validation</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Dataset for Paper: A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents

<p>Dataset for the paper: &quot;A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents&quot;, P. Kaddas, K. Palaiologos, B. Gatos, V. Katsouros, K. Christopoulou, 17th&nbsp;International Conference on Document Analysis and Recognition (ICDAR), San Jose, California, USA</p> <p>The dataset consists&nbsp;of 57 pages from the third edition of the Greek New Testament published by Robert Estienne (1503&ndash;1559), who was appointed &ldquo;Royal Typographer&rdquo; by the King of France Fran&ccedil;ois I (1494&ndash;1547). Robert Estienne produced this edition in 1550 using the grecs du roi typeface, produced by Claude Garamont on the basis of the Greek minuscule style of the calligrapher Angelos Vergikios (1505&ndash;1569) from Crete, who active copying Greek manuscripts in Venice and France. The dataset consists of 2045 cropped text line images in .png format with their corresponding OCR in .txt format, where 1431 used for training, 204 for validation and 410 for test. Initial images acquired from: https://bibles-online.net/flippingbook/1550/</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Figs 64, 65 in A revision of Afrotropical Chyromyidae (excluding Gymnochiromyia Hendel) (Diptera: Schizophora), with the recognition of two subfamilies and the description of new genera

Figs 64, 65. (64) Phylogenetic tree (tree 0) of the genera of Chyromyidae with synapomorphies mapped on to the clades, outgroup taxon: Heleomyzinae; (65) Phylogenetic tree (tree 0) of the genera of Chyromyidae showing Standard Bootstrap (GC values, 1000 replicates, cut =1), resampling results (with percentages at each clade) that confirm the parent tree, outgroup taxon: Heleomyzinae.

opencc-by-4.0Dec 2009View details →
zenodo40/100

Fig. 44 in A revision of Afrotropical Chyromyidae (excluding Gymnochiromyia Hendel) (Diptera: Schizophora), with the recognition of two subfamilies and the description of new genera

Fig. 44. Aphaniosoma frequens sp. n., ^postabdomen: (a) posterior, (b) lateral, (c) spermathecae. Scale bar = 0.15 mm.

opencc-by-4.0Dec 2009View details →
zenodo40/100

Figs 62, 63 in A revision of Afrotropical Chyromyidae (excluding Gymnochiromyia Hendel) (Diptera: Schizophora), with the recognition of two subfamilies and the description of new genera

Figs 62, 63. (62) Cladogram of the genera of Chyromyidae showing Strict Concensus of 3 trees (0 taxa excluded), outgroup taxon: Heleomyzinae; (63) Cladogram of the genera of Chyromyidae showing result of Majority Rule (from 3 trees, cut 50), outgroup taxon: Heleomyzinae.

opencc-by-4.0Dec 2009View details →
zenodo40/100

Fig. 60 in A revision of Afrotropical Chyromyidae (excluding Gymnochiromyia Hendel) (Diptera: Schizophora), with the recognition of two subfamilies and the description of new genera

Fig. 60. Tethysimyia deemingi (Ebejer), ơ: (a) postabdomen, lateral; (b, c) hypopygium, ventral (b) and lateral (c). Scale bar = 0.15 mm.

opencc-by-4.0Dec 2009View details →
zenodo40/100

Fig 61 in A revision of Afrotropical Chyromyidae (excluding Gymnochiromyia Hendel) (Diptera: Schizophora), with the recognition of two subfamilies and the description of new genera

Fig 61. Tethysimyia deemingi (Ebejer), ^postabdomen: (a) lateral, (b) ventral, (c) spermatheca. Scale bar = 0.15 mm.

opencc-by-4.0Dec 2009View details →
zenodo40/100

Fig. 42 in A revision of Afrotropical Chyromyidae (excluding Gymnochiromyia Hendel) (Diptera: Schizophora), with the recognition of two subfamilies and the description of new genera

Fig. 42. Aphaniosoma flavescens sp. n., ^postabdomen: (a) posterior, (b) lateral, (c) spermatheca. Scale bar = 0.15 mm.

opencc-by-4.0Dec 2009View details →
zenodo40/100

Fig. 36 in A revision of Afrotropical Chyromyidae (excluding Gymnochiromyia Hendel) (Diptera: Schizophora), with the recognition of two subfamilies and the description of new genera

Fig. 36. Aphaniosoma atriceps sp. n.: (a) ơ hypopygium, lateral, scale bar = 0.15 mm; (b–d) ^postabdomen: (b) posterior, (c) lateral, (d) spermathecae, scale bar = 0.1 mm.

opencc-by-4.0Dec 2009View details →
zenodo40/100

Fig. 28 in A revision of Afrotropical Chyromyidae (excluding Gymnochiromyia Hendel) (Diptera: Schizophora), with the recognition of two subfamilies and the description of new genera

Fig. 28. Somatiosoma nitescens Frey: (a) ơ: hypopygium, lateral; (b) prg, ventral; (c, d) part of ơ postabdomen, lateral (c) and dorsal (d); (e–g) ^postabdomen: (e) dorsal, (f) ventral, (g) spermatheca, enlarged. Scale bars = 0.15 mm.

opencc-by-4.0Dec 2009View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record