Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

14,447

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

14,447 results for “Identification”

Learn how ShareScore rates datasets ↗
zenodo44/100

Artificial fingerprints engraved through block-copolymers as nanoscale physical unclonable functions for authentication and identification - Dataset

<p>This is the dataset of "Artificial fingerprints engraved through block-copolymers as nanoscale physical unclonable functions for authentication and identification" by Irdi Murataj, Chiara Magosso, Stefano Carignano, Matteo Fretto, Federico Ferrarese Lupi, and Gianluca Milano, Nature Communications (2024), DOI: 10.1038/s41467-024-54492-8</p> <p>Part of this was funded by the project MEMQuD, code 20FUN06. The project has received funding from the EMPIR program co-financed by the Participating States and from the European Union's Horizon 2020 research and innovation program.</p> <p>Part of this work was supported by the European project OpMetBat, code 21GRD01. The project has received funding from the European Partnership on Metrology, cofinanced from the the European Union's Horizon Europe Research and Innovation Programme, and by Participating States.</p> <p>Part of this work was supported by the European Union - Next Generation EU under the National Recovery and Resilience Plan (NRRP), Mission 04 Component 2 Investment 3.1 | Project Code: IR0000027 - CUP: B33C22000710006 - iENTRANCE@ENL: Infrastructure for Energy TRAnsition aNd Circular Economy @EuroNanoLab.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Data for "Identification of 4876 Bent-Tail Radio Galaxies in the FIRST Survey using Deep Learning Combined with Visual Inspection"

<p>The data are the full versions of tables that will be published in the manuscript titled "Identification of 4876 Bent-Tail Radio Galaxies in the FIRST Survey using Deep Learning Combined with Visual Inspection" by The Astrophysical Journal Supplement Series.</p> <p>The table file named "FIRST_bt_table1.csv" is the full table for "A catalog of 4876 BTRGs identified from VLA FIRST survey". &nbsp;</p> <p>The table file named "FIRST_bt_table2.csv" is the full table for "Cluster details for BTRGs". &nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo44/100

Version 4.2 (20230306) of the MALDI-ToF Mass Spectrometry Database for Identification and Classification of Highly Pathogenic Microorganisms from the Robert Koch-Institute (RKI)

<p><em>(Version </em>20230306<em>, </em>btmsp files modified May 31, 2023, additional taxonomic information added Dec 27, 2024<em>) </em></p> <p>Version 4.2 (20230306) of the RKI MALDI-ToF mass spectra database represents the third update of the original database (version 20161027,&nbsp;<a href="http://doi.org/10.5281/zenodo.163517">https://doi.org/10.5281/zenodo.163517</a>). The RKI Database v.4.2 now contains a total of 11055 MALDI-ToF mass spectra from 1601 microbial strains of highly pathogenic (i.e. biosafety level 3, BSL-3) bacteria such as <em>Bacillus anthracis</em>, <em>Brucella melitensis</em>, <em>Yersinia pestis</em>, <em>Burkholderia mallei / pseudomallei</em> and <em>Francisella tularensis</em> as well as a selection of spectra of their close and distant relatives. The database can be used as a reference for the diagnosis of BSL-3 bacteria using proprietary and free software packages for MALDI-ToF MS-based microbial identification. The spectral data are provided as a zip archive (<a href="https://zenodo.org/records/14562231/files/zenodo%20db%20230306.zip?download=1&amp;preview=1">zenodo db 230306.zip</a>) containing the original mass spectra in their native data format (Bruker Daltonics). Please refer to the pdf file (<a href="https://zenodo.org/records/14562231/files/230306-ZENODO-Metadata.pdf?download=1&amp;preview=1">230306-ZENODO-Metadata.pdf</a>) for information on cultivation conditions, sample preparation and details of the spectra acquisition. Please do not try to print this document (&gt;1600 pages!).</p> <p>Version 20230306 of the RKI database contains for the first time files in the <em>btmsp</em> format (e.g.&nbsp; <a href="https://zenodo.org/records/14562231/files/2023-May-23-Bacillus-RKI-Database-568.btmsp?download=1&amp;preview=1">2023-May-23-Bacillus-RKI-Database-568.btmsp </a> <a href="https://zenodo.org/api/files/35e90a0c-653d-4ba4-bf93-50b2bd80d073/2023-May-23-Bacillus-RKI-Database-570.btmsp"> </a>and others). These files were generated using the MALDI Biotyper software (Bruker Daltonics) and contain a total of 1601 main spectra (msp) from the BSL-3 database in the proprietary data format of the MALDI Biotyper software. *.<em>btmsp </em>files can be imported and used for identification with this software solution. Please refer to the manufacturer's manual for details on importing <em>btmsp </em>files. Note that the btmsp file available in database version 4 is broken and cannot be imported.</p> <p>The pkf files (<a href="https://zenodo.org/records/14562231/files/230306_ZENODO_30Peaks_0.75.pkf?download=1&amp;preview=1">230306_ZENODO_30Peaks_0.75.pkf</a>, <a href="https://zenodo.org/records/14562231/files/230306_ZENODO_45Peaks_0.75.pkf?download=1&amp;preview=1">230306_ZENODO_45Peaks_0.75.pkf</a>) represent two versions of the MS peak list data in a Matlab compatible format. The latter data can be imported into MicrobeMS, a free Matlab-based software solution developed at the RKI. MicrobeMS can be used for the identification of microorganisms by MALDI-ToF MS and is available at <a href="https://wiki-ms.microbe-ms.com">https://wiki-ms.microbe-ms.com</a>.</p> <p>The Excel file <a href="https://zenodo.org/records/14562231/files/Taxonomy%20information%20-%20RKI%20MALDI-ToF%20MS%20database%20of%20HPB%20at%20ZENODO%20v.4.xlsx?download=1&amp;preview=1">Taxonomy information - RKI MALDI-ToF MS database of HPB at ZENODO v.4.xlsx</a> contains additional taxonomic information such as a detailed list of bacterial MALDI-ToF mass spectra (sheet #1), overviews on the number of spectra per strain, species or bacterial genus (sheet #2), numbers of strains per species, or genus (sheet #3), etc.</p> <p>The RKI mass spectrometry database is updated regularly.</p> <p>The author would like to thank the following individuals for providing microbial strains and species or mass spectra thereof. Without their help, this work would not have been possible.</p> <ul> <li><strong>Wolfgang Beyer</strong> - University of Hohenheim, Faculty of Agricultural Sciences, Stuttgart, Germany</li> <li><strong>Guido Werner</strong> - Robert Koch-Institute, Nosocomial Pathogens and Antibiotic Resistances (FG13), Wernigerode, Germany</li> <li><strong>Alejandra Bosch</strong> - CINDEFI, CONICET-CCT La Plata, Facultad de Ciencias Exactas, Universidad Nacional de La Plata, La Plata, Buenos Aires, Argentina</li> <li><strong>Michal Drevinek</strong> - National Institute for Nuclear, Biological and Chemical Protection, Milin, Czech Republic</li> <li><strong>Roland Grunow, Daniela Jacob, Silke Klee, Susann Dupke </strong>and <strong>Holger Scholz</strong> - Robert Koch-Institute, Highly Pathogenic Microorganisms (ZBS2), Berlin, Germany</li> <li><strong>J&ouml;rg Rau </strong>- Chemisches und Veterin&auml;runtersuchungsamt Stuttgart, Fellbach, Germany</li> <li><strong>Jens Jacob</strong> - Robert Koch-Institute, Hospital Hygiene, Infection Prevention and Control (FG14), Berlin, Germany</li> <li><strong>Martin Mielke</strong> - Robert Koch-Institute, Department 1 - Infectious Diseases, Berlin, Germany</li> <li><strong>Monika Ehling-Schulz</strong> - Functional Microbiology, Institute of Microbiology, University of Veterinary Medicine, Vienna, Austria</li> <li><strong>Armand Paauw</strong> - Department of Medical Microbiology, CBRN protection, Universitair Medisch Centrum Utrecht, TNO, Rijswijk, The Netherlands</li> <li><strong>Herbert Tomaso</strong><strong> </strong>&ndash; Friedrich-L&ouml;ffler-Institut (FLI), Federal Research Institute for Animal Health, Jena, Germany</li> <li><strong>Gabriel Karner</strong><strong> </strong>- Karner D&uuml;ngerproduktion GmbH, Research &amp; Development, Neulengbach, Austria</li> <li><strong>Rainer </strong><strong>Borriss</strong><strong> </strong>- Institute of Marine Biotechnology e.V. (IMaB), Greifswald, Germany</li> <li><strong>Le Thi Thanh Tam</strong><strong> </strong>- Division of Plant Pathology and Phyto-Immunology, Plant Protection Research Institute, Hanoi, Socialist Republic of Vietnam</li> <li><strong>Xuewen</strong><strong> Gao</strong><strong> </strong>- College of Plant Protection, Nanjing Agricultural University, Key Laboratory of Integrated Management of Crop Diseases and Pests, Nanjing, People&rsquo;s Republic of China</li> </ul> <p>For a detailed description of the database see: Lasch, P., Beyer, W., Bosch, A. <em>et al.</em> A MALDI-ToF mass spectrometry database for identification and classification of highly pathogenic bacteria. <em>Sci Data</em> <strong>12</strong>, 187 (2025). <a href="https://doi.org/10.1038/s41597-025-04504-z">https://doi.org/10.1038/s41597-025-04504-z</a></p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Long-term live imaging, cell identification and cell tracking in regenerating crustacean legs

<p>Supplementary data and videos for the manuscript 'Long-term live imaging, cell identification and cell tracking in regenerating crustacean legs', by &Ccedil;evrim,<sup> </sup>Laplace-Builh&eacute;,<sup> </sup>Sugawara, Rusciano, Labert, Brocard, Almaz&aacute;n and Averof.</p> <p>The supplementary data include:</p> <p><strong>Supplementary Data 1 (.csv file);&nbsp; Live imaging of regenerating <em>Parhyale</em> legs: image acquisition settings</strong></p> <p>Table with information on the 22 time lapse recordings presented in Figure 3, including image acquisition settings, temperature and duration of the recordings.</p> <p><strong>Supplementary Data 2 (.zip file);&nbsp; Live imaging of regenerated <em>Parhyale</em> legs: maximum projections</strong></p> <p>Compressed folder including maximum projections for each of the 22 time lapse recordings presented in Figure 3. These files were generated by projecting all or a subset of the z slices acquired at each time point. A 20 micron scale bar was added on the first time point. These files serve as a quick way to examine the 22 time lapse recordings.</p> <p><strong>Supplementary Data 3 (22 .tif files);&nbsp; Live imaging of regenerated <em>Parhyale</em> legs: complete datasets</strong></p> <p>Complete image 3D+T hyperstacks for each of the 22 time lapse recordings presented in Figure 3. These files have been generated by concatenating the original image stacks and correcting any image shifts, as described in the Methods section of the paper.</p> <p><strong>Supplementary Data 4 (.zip file);&nbsp; Analysis of trade-offs of imaging resolution and image quality</strong></p> <p>The data used for the analysis of trade-offs in imaging and the results shown in Table 1 are included in this compressed folder. Folders for the original recording (labelled 00), for each of the subsampled datasets (labelled 01 to 05), and for the denoised and deconvoluted datasets each include the corresponding image data and ground truth cell tracking files (.tif, .h5, .xml and .mastodon files) and three sets of cell track predictions (.mastodon files). There are also separate folders containing the Elephant detection and flow model parameters for each set of predictions.</p> <p dir="ltr"><strong>Supplementary Data 4 (.zip file);&nbsp; Analysis of trade-offs of imaging resolution and image quality</strong></p> <p dir="ltr">The data used for the analysis of trade-offs in imaging and the results shown in Table 1 are included in two folders. The folder named Image_and_tracking_data includes the image data (.tif, .h5, .xml), ground truth cell tracking files (.mastodon files) and three sets of cell track predictions (.mastodon files) for the original recording (labelled 00), for each of the subsampled datasets (labelled 01 to 05), and for the denoised and deconvoluted datasets. It also includes separate folders containing the Elephant detection and flow model parameters for each set of predictions. The folder named CTC_tracking_results includes the ground-truth data along with three sets of predictions for detection and tracking for each dataset, following the Cell Tracking Challenge format. For each dataset we include label image files (.tif) for every time point along with tracking results in .txt format, and each results directory (01_RES_*) also contains the evaluation results from the Cell Tracking Challenge Evaluation Software. For a detailed explanation of the folder structure, please refer to the Cell Tracking Challenge documentation.</p> <p><strong>Supplementary Data 5 (.zip file);&nbsp; Tracking the progenitors of spineless-expressing cells in the distal carpus</strong></p> <p>The data used to generate Figure 7 are included in this compressed folder, including the live imaging and cell tracking files (.h5, .xml and .mastodon files) and the image stack of the spineless and futsch HCR and DAPI stainings (.tif file). Channel 2 shows spineless expression (mostly nascent transcripts in nuclei), as well as background signal in epidermal nuclei (possibly due to photoconversion of DAPI, see Karg &amp; Golic 2018, Chromosoma 127: 235-245) and strong autofluorescence in granular cells (also visible in channel 1, depicting futsch HCR).</p> <p><strong>Supplementary Data 6 (.txt file);&nbsp; Sequences of <em>Parhyale</em> genes targeted by the HCR probes</strong></p> <p>The sequences are provided in FASTA format.</p> <p dir="ltr"><strong>Supplementary Data 7 (.zip file);&nbsp; Apoptosis in legs that have not been subjected to live imaging</strong></p> <p dir="ltr">The data used to generate Figure 2 supplement 2 are contained in this compressed folder, including 9 image stacks of T4 and T5 legs fixed and stained with DAPI 3 days post amputation (with apoptotic nuclei marked) and a .txt file containing the apoptotic cell counts.</p> <p dir="ltr"><strong>Supplementary Data 8 (.zip file);&nbsp; Analysis of tracking performance in relation to imaging depth</strong></p> <p dir="ltr">The data used to generate Figure 5 are contained in this compressed folder, including separate folders for the data extracted from the analysis of datasets #1 to #5. Each folder includes data from three replicates (batches 001 to 003), with .csv files listing the z location of nucleus centroids (in &micro;m) for the nuclei that were incorrectly detected by Elephant &ndash; either as false positives (FP) or as false negatives (FN) &ndash; and the ground truth data (GT). The folder also includes an .xlsx file gathering all the relevant data and the measurements of precision and recall.</p> <p dir="ltr"><strong>Supplementary Data 9 (.zip file);&nbsp; Detecting the temporal pattern of cell divisions in regenerating legs</strong></p> <p dir="ltr">The data used to generate Figure 4 are contained in this compressed folder, including the five image datasets (.tif, .h5, .xml), the detected cell divisions (.mastodon files), and an .xlxs file containing all the cell divisions counts and graphs.</p> <p><strong>Video 1.&nbsp; Time lapse recording of regeneration in a Parhyale T5 leg (dataset li48-t5)</strong></p> <p>Live imaging of nuclei labelled with H2B-mREFruby (maximum projection of z slices 3-10). Proximal parts of the leg are to the left and the amputation site is at the right of the frame. For annotations of different features please refer to Figure 2. Shortly after leg amputation (0 hpa) hemocytes adhere to the wound. By 16 hpa the wound has melanized. Up to ~32 hpa epithelial cells can be seen migrating and accumulating at the wound, below the melanized scab (Figure 2A,B). Around 31 hpa, the leg tissues become detached from the scab (Figure 2C). At 43 hpa, the carpus-propodus boundary first becomes visible, and thereafter many cells can be observed dividing at the distal part of the leg stump (Figure 2D). At 56 hpa, the propodus-dactylus boundary first becomes visible (Figure 2E). At later stages, tissues in more proximal parts of the leg retract, making space for the regenerating leg to grow (Figure 2F,G). After ~90 hpa cell proliferation there is less cell proliferation and cell movements, and the nuclear positions within the tissue become fixed. Scale bars, 20 &micro;m.</p> <p><strong>Video 2.&nbsp; Time lapse recording of regeneration in a Parhyale T5 leg (dataset li36-t5)</strong></p> <p>Live imaging of nuclei labelled with H2B-mREFruby (maximum projection of z slices 3-15). Proximal parts of the leg are to the left and the amputation site is at the right of the frame. The sequence of events is similar to that described in Video 1, but the progression is slower: epithelial migration towards the wound is observed up to 40 hpa, tissues detach from the scab at 65 hpa, and the carpus-propodus and propodus-dactylus boundaries first become visible at 78 and 91 hpa. The tissues making up the carpus and propodus can be seen pulsating from 105 to 145 hpa. Scale bars, 20 &micro;m.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Surrogate-based optimization using an artificial neural network for a parameter identification in a 3D marine ecosystem model

<p><strong>Abstract:</strong></p> <p>Parameter identification for marine ecosystem models is important for the assessment and validation of marine ecosystem models against observational data. The surrogate-based optimization (SBO) is a computationally efficient method to optimize complex models. SBO replaces the computationally expensive (high-fidelity) model by a surrogate constructed from a less accurate but computationally cheaper (low-fidelity) model in combination with an appropriate correction approach, which improves the accuracy of the low-fidelity model. To construct a computationally cheap low-fidelity model, we tested three different approaches to compute an approximation of the annually periodic solution (i.e., a steady annual cycle) of a marine ecosystem model: firstly, a reduced number of spin-up iterations (several decades instead of millennia), secondly, an artificial neural network (ANN) approximating the steady annual cycle and, finally, a combination of the both approaches. Except for the low-fidelity model using only the ANN, the SBO yielded a solution close to the target and reduced the computational effort significantly. If an ANN approximating appropriately a marine ecosystem model is available, the SBO using this ANN as low-fidelity model presents a promising and computational efficient method for the validation.</p> <p>&nbsp;</p> <p><strong>Content:</strong></p> <ul> <li>SQLite database including the data of the different optimization runs</li> <li>Structure and weights of the used artificial neural network</li> <li>Tracer concentrations obtain from the high-fidelity model for the different optimization runs</li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo44/100

CWID-hi: A Dataset for Complex Word Identification in Hindi Text

<p>This dataset was created by conducting a human intelligence test, wherein native and non-native Hindi speakers annotated words they could not understand in Hindi text. They were then asked to rank the complexity of these words along with their synonyms. A word that received an average rank of &lt;=3 (out of 5) is labeled 1 and the word that received an average rank of &gt;3 is labeled 0. 1 indicates complex and 0 indicates simple.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Extended data for Manuscript: Identification of potential biological targets of oxindole scaffolds via in silico repositioning strategies

<p>This is the Extended Data for the manuscript &quot;<strong>Identification of potential biological targets of oxindole scaffolds via <em>in silico</em> repositioning strategies&quot;&nbsp;</strong>submitted to F1000 Research.</p> <p>Extended Data include a list of all the accession codes as mentioned in the text,&nbsp;the results of 2D fingerprint-based similarity analyses and ligand-protein complexes predicted by rigid docking and Induced Fit Docking calculations.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Dataset for the identification of hypertension in school-aged children from Gqeberha, South Africa

<p>Dataset used to evaluate and compare different international references to identify hypertension among South African school-aged children from disadvantaged communities.</p> <p>It encompasses anonymized, unique, identification numbers, anthropometric and blood pressure measures, as well as blood pressure percentiles and the assigned categories derived from four different reference populations (American, German, global and the study population).</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Genome-wide identification of cell-surface and intracellular immune receptors in 350 plant species

<p>Here we identified cell-surface (LRR-RLKs, LRR-RLPs, LysM-RLKs and LysM-RLPs) and intracellular immune receptors (NB-ARCs) from the genomes of 350 plant species.&nbsp;</p> <p>&nbsp;</p> <p>Zip file contains:</p> <p>Folder &#39;Immune_receptor_sequences&#39; - FASTA files of the identified LRR-RLPs, Lys-RLKs, LysM-RLPs and NB-ARCs.</p> <p>Folder &#39;RLK_sequences&#39; -&nbsp;FASTA files of the identified LRR-RLKs (all and 20 individual subgroups).</p> <p>Folder &#39;RLK_trees&#39; - Phylogenetic TREE files of&nbsp;the identified LRR-RLKs (all and 20 individual subgroups); classified according to their kinase domains.</p> <p>238.species -&nbsp;Phylogenetic tree of the 238 plant species used in the analyses (taken from&nbsp;<a href="https://doi.org/10.1093/jpe/rtv047">https://doi.org/10.1093/jpe/rtv047</a>).</p> <p>350.species&nbsp;&nbsp;-&nbsp;Phylogenetic tree of the 350 plant species used in the analyses.</p> <p>simple.to.original.ids-&nbsp;Translator file&nbsp;for the original ID of each gene.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Identification of Thalweg and Ridge Networks as Landmarks for Terrain Partitioning

<p>Grid digital elevation models having resolution of 1 m or less are increasingly available to scientists and engineers interested in describing current state and evolution of Earth and space topography. Significant information loss is, however, clearly observed when existing terrain analysis methods are used in geophysical modeling, especially when coarse meshes are needed for computational efficiency. The present study shows how thalweg and ridge networks can be extracted automatically from any high-resolution grid digital elevation model without the need to alter the observed topographic data, and how these networks can be used as landmarks for terrain partitioning. The slopeline network extracted in grid digital elevation models is used to determine ridge points, related average rejunction lengths of slopelines extending from ridge points on opposite slopes, exorheic and endorheic basins. Exorheic and endorheic basins are connected through the spilling saddles from endorheic basins to form the thalweg network, and the related ridge network is identified. The obtained thalweg and ridge networks are characterized by using the known concept of drainage area and the new concept of spread area to provide physically meaningful&nbsp;landmarks&nbsp;for terrain partitioning at the desired level of detail. Although the developed methods are inspired by the observation of gravity-driven processes, they support any investigation in Earth and space science where thalweg and ridge networks are relevant topographic features. Potential impacts are exemplified by quantifications of preserved depressions over a mountain area and benefits from physically meaningful unstructured terrain partitioning in surface flow propagation over a complex floodplain.</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

InsectSet32: Dataset for automatic acoustic identification of insects (Orthoptera and Cicadidae)

<p>This dataset contains recordings of 32 sound producing insect species with a total 335 files and a length of 57 minutes. The dataset was compiled for training neural networks to automatically identify insect species while comparing adaptive, waveform-based frontends to conventional mel-spectrogram frontends for audio feature extraction. This work was <a href="https://doi.org/10.1371/journal.pcbi.1011541">published</a> in PLOS Computational Biology and this dataset can be used to replicate the results, as well as other uses.&nbsp;The scripts for audio processing and the machine learning implementations are published on <a href="https://github.com/mariusfaiss/InsectSet32-Adaptive-Representations-of-Sound-for-Automatic-Insect-Recognition">Github</a>.</p> <p>The recordings are split into two datasets.&nbsp;Roughly half of the&nbsp;recordings (147) are of nine species belonging to&nbsp;the order Orthoptera. These recordings stem from a dataset that was&nbsp;originally compiled by <a href="https://orcid.org/0000-0002-8929-2737">Baudewijn Od&eacute;</a>&nbsp;(unpublished).&nbsp;</p> <p>The remaining recordings&nbsp;(188)&nbsp;are of 23 species in the family Cicadidae.&nbsp;These recordings were selected&nbsp;from the Global Cicada Sound Collection hosted on&nbsp;<a href="https://bio.acousti.ca/">Bioacoustica</a>&nbsp;(<a href="https://doi.org/10.1093/database/bav054">doi.org/10.1093/database/bav054</a>), including recordings published in&nbsp;<a href="https://doi.org/10.3897/BDJ.3.e5792">doi.org/10.3897/BDJ.3.e5792</a>&nbsp;&amp;&nbsp;<a href="https://doi.org/10.11646/zootaxa.4340.1">doi.org/10.11646/zootaxa.4340.1</a>.&nbsp;Many recordings from this collection included speech annotations in the beginning of the recordings, therefore the last ten seconds of audio were extracted and used in this dataset.&nbsp;</p> <p>All files were manually inspected and files with strong noise interference or with sounds of multiple species were removed. Between species, the number of files ranges from four to 22 files and the length from 40 seconds to almost nine minutes of audio material for a single species. The files range in length from less than one second to several minutes. All original files were available with sample rates of at least&nbsp;44.1 kHz or higher but were resampled to 44.1 kHz mono WAV&nbsp;files for consistency. The annotation files contain information for each recording, including the file name, species name and identifier, as well as the data subset they were included in for training the neural network (training, test, validation).</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Collation and orthology-based identification of hormone-related genes in bread wheat

<p>Plant hormones coordinate a plethora of developmental processes in plants, including responses to abiotic and biotic stressors. Here, we collate the findings of previous studies identifying bread wheat (<em>Triticum aestivum</em>) genes related to hormonal processes (<strong>biosynthesis</strong>, <strong>transport</strong>, <strong>signalling</strong>, and <strong>catabolism</strong>) and collect wheat orthologues from hundreds of additional hormone-related genes utilising the Ensembl Plants Compara database. We have initially conducted this procedure for <strong>abscisic acid</strong>, <strong>auxins </strong>(IAA and IBA), <strong>brassinosteroids</strong>, <strong>cytokinins</strong>, <strong>ethylene</strong>, <strong>gibberellins</strong>, and <strong>strigolactone</strong>, yielding a total of over 1,700 putative wheat orthologues. We aim to provide a community resource to aid gene annotation and subsequent analyses. We warmly welcome feedback from the community.</p> <p>Please refer to the file <strong>README.pdf</strong> for further details, including methods and references.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

The Identification of Extinct Megafauna in Rock art Using Geometric Morphometrics: A Genyornis newtoni Painting in Arnhem Land, Northern Australia?

<p>Raw data files used for the analysis of a contentiously identified rock-art image located in Arnhem Land, Northern Australia. The data were used to test a novel approach to quantifying species identification in rock art images to assess the extent to which an image resembles other rock art of sound identification or anatomical images of visually similar species.</p> <p>Included files are the raw coordinate data files ("[feature] PCA file", .txt format) for use in Morphologika2, and formatted files for use in CVAGen8 ("[feature]" x1y1 file for CVA", .x1y1 format; "[feature] group file", .txt format).</p> <p>Files produced using software by Rohlf (2015) and Sheets (2014)</p>

opencc-by-4.0Aug 2017View details →
zenodo44/100

Polidoc.net CODEBOOK: National and Regional Manifestos and other Political Documents Collected for the Research Projects "Representation in Europe: Congruence between Preferences of Elites and Voters" (REPCONG) and "The Impact of EU Cohesion Policy on European Identification" (COHESIFY)

<p>The Political Documents Archive http://www.polidoc.net/&nbsp;contains election manifestos, coalition agreements, government declarations and various other documents of political actors from developed democracies. Currently, the archive builds on a stock of more than 3000 political documents from 20 European countries. The aim of the repository is to provide political texts in order to facilitate scholarly research in different areas of comparative politics such as party competition, coalition politics, legislative decision-making or electoral behavior.</p> <p>National electoral manifestos have been collected in the course of the REPCONG project (&quot;Representation in Europe: Policy Congruence between Citizens and Elites&quot;), and the archive includes party manifestos for regional elections in several European democracies. Because the process of European integration resulted in a strengthening of regions in EU member states and in countries that want to join the European Union, the relevance of the regional level for political decision-making has increased during the last decades. Therefore, also the policy profiles of regional parties are required to get a full picture of democratic responsiveness in European states across all levels of the political system. The collection of regional manifestos was supported by the COHESIFY project (www.cohesify.eu), funded under the Horizon 2020 Framework Programme for Research and Innovation. The aim of COHESIFY is to study whether the European Structural and Investment Funds affect people&rsquo;s support for and identification with the European project.</p> <p>The archive is freely accessible (after a simple registration) and meant to foster rigorous research in these areas by enabling scholars to produce valid and reliable findings from empirical studies of textual data rather than unnecessarily struggling to obtain and process texts.</p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

A data set on "Utilizing Constant Energy Difference between sp-Peak and C 1s Core Level in Photoelectron Spectra for Unambiguous Identification and Quantification of Diamond Phase in Nanodiamonds"

<p>The data set to paper:&nbsp;</p> <p>Utilizing Constant Energy Difference between sp-Peak and C 1s Core Level in Photoelectron Spectra for Unambiguous Identification and Quantification of Diamond Phase in Nanodiamonds</p> <p>Oleksandr Romanyuk1,*, &Scaron;těp&aacute;n Stehl&iacute;k1,2, Josef Zemek1, Kateřina Aubrechtov&aacute; Dragounov&aacute;1,3 and Alexander Kromka1</p> <p>1 Institute of Physics of the Czech Academy of Sciences, Cukrovarnick&aacute; 10, 162 00 Prague, Czech Republic<br>2 New Technologies&mdash;Research Centre, University of West Bohemia, Univerzitn&iacute; 8, 306 14 Pilsen, Czech Republic<br>3 Faculty of Nuclear Sciences and Physical Engineering, Czech Technical University in Prague, Břehov&aacute; 7, 115 19 Prague, Czech Republic</p> <p>* corresponding author: romanyuk@fzu.cz</p> <p>Data manager: Krist&yacute;na Dost&aacute;lov&aacute;: dostalovak@fzu.cz</p> <p>Date of data collection: 1. 1. 2024 - 15. 03. 2024</p> <p>All the data showed in the pictures are provided in X-Y format with described sample. Always, the respective Figure to which the data belong is provided in high resolution.&nbsp;<br>The data are in the following formats:&nbsp;<br>Figure 1: tiff, csv<br>Figure 2: tiff, csv<br>Figure 3: tiff, csv<br>Figure 4: tiff, csv<br>Figure 5: tiff, csv</p> <p>Data acquistion and processing is provided in the Experimental part in the publication: DOI:10.3390/nano14070590</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Hailstorm Identification and Tracking over Brazil (HIToB): A Storm Polygons Database From GOES ABI Data from 2018 to 2023

<p>This dataset comprises a detailed record of deep convective storm events tracked across South America from 2018 to 2023, utilizing brightness temperature (BT) data from Channel 13 of the GOES-16 Advanced Baseline Imager (ABI) and the TATHU (Tracking and Analysis of Thunderstorms) toolset, that caused hail-fall over Brazil. The database includes storm identification, tracking details, and associated meteorological variables such as brightness temperature statistics inside the storm polygon at each scene and event classifications (e.g., spontaneous generation, continuity, split, merge). The storms were detected and tracked based on brightness temperature threshold of 235 K, with tracking data refined by a 10% overlap criterion between sequential scenes. The tracked convective systems were filtered for intersections in space and time with verified hail reports from Prevots group. The whole family of storm polygons that matched the reports were exported to this database with SpatiaLite enabled dtaa format, in order to make it easier for spatial data queries and analysis. Some example queries using Python library SQLAlchemy are displayed in the code repository as well as the process of creating the tables in the database.<br><br>The data is organized in three tables: "storms", "storm_events" and "intersections". In table "storms" are the records of storm families identifier. Each identifier represents a sequence of storm polygons tracked over subsequent satellite scenes. Table "storm_events" holds the evolution of the storm's geometry through its lifecycle, including BT's mean, minimum and standard deviation inside the storm polygon; as well as storm's pixel count (i.e. storm size). Intersections table stores every instance where a storm event polygon intersects with a hailstorm report's buffer at the corresponding time. In total, there are 9893 intersections belonging to 2172 unique storm families.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Results: Towards Realistic SATD Identification Through Machine Learning Models: Ongoing Research and Preliminary Results

<p>Automated identification of self-admitted technical debt (SATD) has been crucial for advancements in managing such debt.&nbsp;<br>However, state-of-the-arts studies often overlook chronological factors, leading to experiments that do not faithfully replicate the conditions developers face in their daily routines.<br>This study initiates a chronological analysis of SATD identification through machine learning models, emphasizing the significance of temporal factors in automated SATD detection.&nbsp;<br>The research is in its preliminary phase, divided into two stages: evaluating model performance trained on historical data and tested in prospective contexts, and examining model generalization across various projects. Preliminary results reveal that the chronological factor can positively or negatively influence model performance and that some models are not sufficiently general when trained and tested on different projects.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Identification of Southeast Asian Anopheles mosquito species with matrix-assisted laser desorption/ionization time-of-flight mass spectrometry using a cross-correlation approach

<p>This is the dataset used in the analysis "Identification of Southeast Asian <em>Anopheles </em>mosquito species with matrix-assisted laser desorption/ionization time-of-flight mass spectrometry using a cross-correlation approach". It consists in&nbsp;3584 raw mass spectra (mzXML file format) of the head of 359 <em>Anopheles </em>mosquito specimens collected in Karen (Kayin state) in Myanmar between 2020 and 2022 and associated metadata (Rdata file format) including sample information (taxonomy.Rdata) and spectra information (metadata.Rdata).</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Firemaker image collection for benchmarking forensic writer identification using image-based pattern recognition

<p>Disclaimer and terms of use:<br> ============================</p> <p>/*****************************************************************************\<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; This is the Firemaker NFI-images Distribution &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; This distribution contains 1000 images of scanned handwritten text, &nbsp; &nbsp; &nbsp; *<br> * &nbsp; scanned at resolution 300dpi grey scale, containing pages of &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;*<br> * &nbsp; handwritten text by 250 writers, four pages per writer, from four &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; writing conditions, one condition per page. The conditions are: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; p1: copied, natural style, p2: copied, UPPER case, p3: copied and forged, *<br> * &nbsp; i.e.,&quot;try to write in a different style than your natural style&quot;, and p4, *<br> * &nbsp; self generated, i.e., text produced to describe a given cartoon. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; Copyright The International Unipen Foundation, 2000, All rights reserved &nbsp;*<br> *******************************************************************************<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;DISCLAIMER AND COPYRIGHT NOTICE FOR ALL DATA CONTAINED ON THIS CDROM: &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;1) PERMISSION IS HEREBY GRANTED TO USE THE DATA FOR RESEARCH &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; PURPOSES. IT IS NOT ALLOWED TO DISTRIBUTE THIS DATA FOR COMMERCIAL &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; PURPOSES. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;2) PROVIDER GIVES NO EXPRESS OR IMPLIED WARRANTY OF ANY KIND AND ANY &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR PURPOSE ARE &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; DISCLAIMED. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;3) PROVIDER SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, SPECIAL, &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF ANY USE OF THIS &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; DATA. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;4) THE USER SHOULD REFER TO THE FIRST PUBLIC ARTICLE ON THIS DATA SET: &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; M. Bulacu, L. Schomaker &amp; L. Vuurpijl (2003). &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; Writer identification using edge-based directional features. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; ICDAR &#39;03: Proceedings of the 7th International Conference on Document &nbsp;*<br> * &nbsp; &nbsp; Analysis and Recognition, pp. 937-941. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; Piscataway: IEEE Computer, ISBN 0-7695-1960-1 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;5) THE RECIPIENT SHOULD REFRAIN FROM PROLIFERATING THE DATA SET TO THIRD &nbsp; *<br> * &nbsp;PARTIES EXTERNAL TO HIS/HER LOCAL RESEARCH GROUP. PLEASE REFER INTERESTED &nbsp;*<br> * &nbsp;RESEARCHERS TO HTTP://UNIPEN.ORG FOR OBTAINING THEIR OWN COPY. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> \*****************************************************************************/</p> <p>BibTeX entry: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <p>&nbsp; @inproceedings{Firemaker, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<br> &nbsp; &nbsp; author = {Bulacu, M. and Schomaker, L.R.B. and Vuurpijl, L.}, &nbsp; &nbsp;<br> &nbsp; &nbsp; title = {Writer Identification Using Edge-Based Directional Features},<br> &nbsp; &nbsp; booktitle = {ICDAR &#39;03: Proceedings of the 7th International&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Conference on Document Analysis and Recognition},<br> &nbsp; &nbsp; year = {2003},<br> &nbsp; &nbsp; isbn = {0-7695-1960-1},<br> &nbsp; &nbsp; pages = {937-941},<br> &nbsp; &nbsp; publisher = {IEEE Computer Society},<br> &nbsp; &nbsp; address = {Washington, DC, USA},<br> &nbsp; &nbsp;}</p> <p>In the project &quot;Vergelijk&quot;, a grant obtained from the Dutch Forensic Science<br> Institute, two existing professional writer-identification systems have been&nbsp;<br> compared regarding usability studies and in particular recognition&nbsp;<br> performance (Schomaker &amp; Vuurpijl, 2000). The results of this comparison&nbsp;<br> are contained in a confidential report:</p> <p>&nbsp; L.R.B. Schomaker and L.G. Vuurpijl (2000).&nbsp;<br> &nbsp; Forensic writer identification: A benchmark data set&nbsp;<br> &nbsp; and a comparison of two systems. Technical report,&nbsp;<br> &nbsp; Nijmegen Institute for Cognition and Information (NICI),&nbsp;<br> &nbsp; University of Nijmegen, The Netherlands.</p> <p>Informative and non-confidential details from this report are&nbsp;<br> given in the accompanying file: &nbsp;&#39;firemaker-dbase.pdf&#39;</p> <p>To compare both systems, a carefully designed experiment was conducted to<br> record handwritten samples from male and female writers in several conditions:</p> <p>Condition 1: Normal constrained handwriting<br> ==============================================</p> <p>Below, the Dutch text writers had to produce in normal handwriting is given.&nbsp;</p> <p>--- start text ----<br> Zij bezochten veilingen en reisden met de KLM. Voor<br> korte afstanden huurden ze een auto, meestal een VW<br> of een Ford.<br> &lt;EMPTY LINE&gt;<br> De veilingen waren van 7-4-1993 tot 3-5-1993 in New<br> York, Tokyo, Qu&eacute;bec, Rome, Parijs, Z&uuml;rich en Oslo.<br> &lt;EMPTY LINE&gt;<br> Omdat de veilingen steeds begonnen om 12 uur en je<br> gemiddeld 200 tot 300 kilometer moest rijden,<br> stonden zij steeds om 6.30 uur op en vertrokken om<br> 8 uur uit het hotel.<br> &lt;EMPTY LINE&gt;<br> Elke dag hadden ze vijfhonderd (f 500,-) gulden<br> nodig. Daarvoor gebruikten ze elke keer een cheque<br> van tweehonderd (f 200,-) en een cheque van<br> driehonderd (f 300,-) gulden. Aan geschenken gaven<br> ze ongeveer honderd gulden (f 100,-) uit.<br> --- end text ----</p> <p><br> Condition 2: Production of constrained block capital handwriting<br> ================================================================</p> <p>In this condition, the writers had to produce the following text<br> in block-capital handwriting:</p> <p>--- start text ----<br> NADAT ZE IN NEW YORK, TOKYO, QU&Eacute;BEC, PARIJS, Z&Uuml;RICH<br> EN OSLO WAREN GEWEEST, VLOGEN ZE UIT DE USA TERUG<br> MET VLUCHT KL 658 OM 12 UUR.<br> &lt;empty line&gt;<br> ZE KWAMEN AAN IN DUBLIN OM 7 UUR EN IN AMSTERDAM OM<br> 9.40 UUR &#39;S AVONDS. DE FIAT VAN BOB EN DE VW VAN<br> DAVID STONDEN IN R3 VAN HET PARKEERTERREIN.<br> HIERVOOR MOESTEN ZE HONDERD GULDEN (F 100,-)<br> BETALEN.<br> --- end text ----</p> <p><br> Condition 3: Production of free-forged handwriting<br> ==================================================</p> <p>Below, the text writers had to produce in the free-forged handwriting<br> condition is given. No example of handwriting is given which they have to<br> mimick (forge), the condition concerns a self-conceived distorted&nbsp;<br> handwriting style.</p> <p>--- start text ----<br> Nog dezelfde avond reden ze naar hun vrienden<br> Chris, Emile, Jan, Irene en Henk, nadat ze hun<br> vriendinnen Greta en Maria hadden opgehaald.<br> &lt;EMPTY LINE&gt;<br> Samen hadden ze vijfhonderd (500) zeldzame<br> postzegels gekocht, Bob driehonderd (300) en David<br> tweehonderd (200).<br> &lt;EMPTY LINE&gt;<br> De reis was de moeite waard geweest.<br> --- end text ----</p> <p><br> Condition 4: Production of unconstrained handwriting<br> ====================================================</p> <p>The final text writers had to produce is unconstrained handwriting.<br> The cartoon, a series of pictures concerning a &#39;UFO&#39; landing had<br> to be described in their own words, in at least six lines of text.<br> See image file &quot;space.gif&quot;.</p> <p><br> Thruth labels and writer identifications<br> ========================================</p> <p>Each writer has a unique id, specified as:</p> <p>&nbsp; &nbsp;id: &nbsp; {num}{set}<br> &nbsp; num: &nbsp; a three-digit number<br> &nbsp;set: &nbsp; &nbsp;either 01, 02, 03 or 04, identifying one of the 4 experiments</p> <p>The vast majority of the writers producing sets 01, 02 and 03 mimicked the<br> content and layout (empty lines) of the constrained texts they had to copy<br> sufficiently accurately, such that the example texts are a good indication of<br> the contents. However, as set 04 (&quot;describe cartoon story&quot;) &nbsp;contains<br> unconstrained self-generated handwriting, the corresponding thruth &nbsp;labels had<br> to be extracted manually. The resulting label files are contained in &nbsp;the<br> directory ./300dpi/p4-self-natural/labels/</p> <p>Note: no letter, word, line or paragraph segmentation is provided with this<br> data set. The main text can be cropped easily. Since the orientation is<br> horizontal, projection techniques can be used to extract lines, using<br> a line-spacing parameter (~94 pixels line height) as an additional check.&nbsp;</p> <p><br> Overview of directories:</p> <p>300dpi/<br> &nbsp; &nbsp;p1-copy-normal/ &nbsp; &nbsp; &nbsp;Copying task, normal writing style &nbsp;<br> &nbsp; &nbsp;p2-copy-upper/ &nbsp; &nbsp; &nbsp; Copying task, UPPER-case&nbsp;<br> &nbsp; &nbsp;p3-copy-forged/ &nbsp; &nbsp; &nbsp;Copying task, instructed to mimic another script style<br> &nbsp; &nbsp;p4-self-natural/ &nbsp; &nbsp; Self-generated text, natural writing condition</p> <p>Note: the original raw collection contained writer #155, who has been removed<br> from this data set, as his first condition (p1) was started in upper case and<br> the page was not &nbsp;completed. Deleted files were 15501.tif, 15502.tif, 15503.tif<br> and 15504.tif.</p> <p>Note: the name of this data set (Firemaker) is a contraction of the names<br> Vuurpijl and Schomaker.</p> <p>Note b: Example of a cutout of essential handwritten text using NetPBM tools: &nbsp;<br> &nbsp;tifftopnm 15201.tif | pnmcut -left 50 -right 2400 -top 700 -bottom 3250 &gt; handwriting.pgm</p> <p>&nbsp;For an experiment, the upper and lower halves of the resulting image were<br> &nbsp;usually used in the Schomaker &amp; Bulacu studies to obtain two samples of&nbsp;<br> &nbsp;handwriting for a writer.</p> <p>&nbsp;http://www.ai.rug.nl/~lambert<br> &nbsp;http://www.ai.rug.nl/~bulacu</p> <p>Our features for writer identification:</p> <p>Lambert Schomaker<br> &nbsp;http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> &nbsp;L. Schomaker &amp; M. Bulacu (2004).&nbsp;<br> &nbsp;Automatic writer identification using connected-component contours and edge-based features of upper-case Western script.&nbsp;<br> &nbsp;IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.</p> <p>Marius Bulacu<br> &nbsp;http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> &nbsp;Bulacu, M. &amp; Schomaker, L.R.B. (2007).&nbsp;<br> &nbsp;Text-independent Writer Identification and Verification Using Textural and Allographic Features,&nbsp;<br> &nbsp;IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</p> <p>Axel Brink<br> &nbsp;http://www.ai.rug.nl/~axel/ &nbsp;&#39;Quill&#39; feature<br> &nbsp;A.A. Brink, J. Smit, M.L. Bulacu, and L.R.B. Schomaker (2011).&nbsp;<br> &nbsp;Writer identification using directional ink-trace width measurements,&nbsp;<br> &nbsp;Pattern Recognition (July 2011), doi: 10.1016/j.patcog.2011.07.005<br> &nbsp;<br> These three feature groups (hinge, fraglets, quill) have been combined in<br> a single MS Windows application, GIWIS which is available for scientific<br> use upon request (schomaker@ai.rug.nl)</p> <p>Note c.</p> <p>The accompanying file &#39;Firemaker-writer-info.dat&#39; contains some<br> writer information:&nbsp;<br> Column 1: writer identification code<br> Column 2: sex<br> Column 3: handedness,&nbsp;<br> Column 4: age in years<br> Column 5: major Western script group (print,cursive or mixed)<br> &nbsp;</p>

opencc-by-4.0Dec 1999View details →
zenodo44/100

ImUnipen image data set for writer identification (N=208) - vectorial handwriting converted to usable images

<p><br> ==============<br> Terms of Usage<br> ==============</p> <p>The ImUnipen data set is intended for non-commercial, scientific use,<br> and is distributed under auspices of the Unipen Foundation.</p> <p>Please always refer to the following paper in IEEE PAMI when using<br> the ImUnipen data set:</p> <p>&nbsp;Bulacu, M.; Schomaker, L.<br> &nbsp;Text-Independent Writer Identification and Verification<br> &nbsp;Using Textural and Allographic Features<br> &nbsp;Pattern Analysis and Machine Intelligence, IEEE Transactions on<br> &nbsp;Volume 29, Issue 4, April 2007 Page(s):701 - 717</p> <p>The ImUnipen data set is derived from the Unipen (unipen.org)<br> data set of on-line (i.e., vectorial, xy) handwriting.<br> The xy-coordinates and a line-generator algorithm are used<br> to generate a raster image, as if the data were optically scanned.</p> <p>Contents: for 208 writers, there are two PNG images per writer of<br> an artificially constructed table of naturally written words (49MByte).<br> These words are pasted onto a white page. For systematics reasons,<br> we call such a page a Paragraph, see below.</p> <p>The file names are organized as (example):</p> <p>&nbsp;&nbsp; Writ990221.Doc01.Par00.png<br> &nbsp;&nbsp; Writ990221.Doc01.Par01.png</p> <p>&nbsp;&nbsp; meaning: writer number 990221, document 01 (there exists only Doc01)<br> and the image with artificial &quot;paragraph&quot; of isolated words &quot;Par00&quot;<br> and &quot;Par01&quot;.</p> <p>The Par00 and Pa01 images are typically used as the query<br> and best match in a leave-one-out setting for writer identification.<br> For instance, Par00 is the query, and Par01 is added to the total set<br> of all other images as the attractor for an identification search.</p> <p>For these experiments, word labels are not given in this data set,<br> on purpose, as the goal is to test recognition-free writer identification<br> methods.</p> <p>For a description of the regular<br> Unipen data set, please visit http://unipen.org</p> <p>Lambert Schomaker constructed this set in 2005</p>

opencc-by-4.0Sep 2008View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record