Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
38,240
datasets available to search
ShareScore release 0.7.1
Dataset results
38,240 results for “Imaging”
Underwater images collected by an Autonomous Surface Vehicle in Hermitage, Réunion - 2024-06-28
<i>This dataset was collected by an Autonomous Surface Vehicle in Hermitage, Réunion - 2024-06-28.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 29.62 GB of MP4 files, which were trimmed into 9962 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 99.97% of these extracted images are useful and 0.03% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 87.68 %, Q2: 5.45 %, Q5: 6.87 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We only keep the values which have a GPS correction in Q1.<br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 0.2 m and 50.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. Then we apply the local geoid if available.<br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.141 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://github.com/SeatizenDOI/plancha-workflow/releases/tag/v1.0.3" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://github.com/SeatizenDOI/plancha-inference/releases/tag/v1.0.0" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>
Underwater images collected by an Autonomous Surface Vehicle in Tessier, Réunion - 2024-05-15
<i>This dataset was collected by an Autonomous Surface Vehicle in Tessier, Réunion - 2024-05-15.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 19.42 GB of MP4 files, which were trimmed into 6576 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 99.7% of these extracted images are useful and 0.3% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 23.49 %, Q2: 75.4 %, Q5: 1.1 % <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://github.com/SeatizenDOI/plancha-workflow/releases/tag/v1.0.3" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://github.com/SeatizenDOI/plancha-inference/releases/tag/v1.0.0" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>
Underwater images collected by an Autonomous Surface Vehicle in Pointe-Des-Aigrettes, Réunion - 2024-07-08
<i>This dataset was collected by an Autonomous Surface Vehicle in Pointe-Des-Aigrettes, Réunion - 2024-07-08.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 91.26 %, Q2: 2.39 %, Q5: 6.35 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We only keep the values which have a GPS correction in Q1.<br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 3.0 m and 50.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. Then we apply the local geoid if available.<br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.733 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://github.com/SeatizenDOI/plancha-workflow/releases/tag/v1.0.3" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://github.com/SeatizenDOI/plancha-inference/releases/tag/v1.0.0" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>
Underwater images collected by an Autonomous Surface Vehicle in Hermitage, Réunion - 2023-11-20
<i>This dataset was collected by an Autonomous Surface Vehicle in Hermitage, Réunion - 2023-11-20.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 83.45 %, Q2: 16.28 %, Q5: 0.27 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We only keep the values which have a GPS correction in Q1.<br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 0.2 m and 50.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. Then we apply the local geoid if available.<br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.524 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://github.com/SeatizenDOI/plancha-workflow/releases/tag/v1.0.3" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://github.com/SeatizenDOI/plancha-inference/releases/tag/v1.0.0" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>
Image-derived indicators of phytoplankton community responses to Pseudo-nitzschia blooms
<p>Data associated with the manuscript "Image-derived indicators of phytoplankton community responses to <em>Pseudo-nitzschia</em> blooms" submitted to the journal <em>Harmful Algae</em>. There is an additional R script that calculates an interaction metric as described in the paper. </p>
Digital Repository of Ireland Member Digitisation Workflows for 2D Image Files: Survey Questions and Dataset
<p>The Digital Repository of Ireland (DRI) issued a survey to its membership, <strong>DRI Member Digitisation Workflows for 2D Images</strong>, which ran from December 7, 2023–January 31, 2024. The survey was conducted to improve the DRI’s understanding of the technical processes and metadata workflows that our members use to digitise and share images in the Repository, in order to better tailor our support for this work and deliver the most complete information about digital images files available to our users. </p> <p>The survey informed the actions taken in WorldFAIR Project WP13 deliverable <a href="https://doi.org/10.5281/zenodo.10850009" target="_blank" rel="noopener">13.3 Implementing and Testing the Cultural Heritage Image Sharing Recommendations: DRI Case Study Report</a>. The data will inform ongoing work at DRI aimed at improving the transparency of technical information associated with digital assets accessed through the Repository.</p> <p>Read more about the Cultural Heritage Image Sharing Case Study DRI on our website: <a href="https://dri.ie/the-worldfair-project/">https://dri.ie/the-worldfair-project/</a>. </p> <p>Summary: DRI is Ireland's national repository for the arts, humanities, and social sciences data, and operates on a membership scheme. There were 20 respondents to the survey, giving us a response rate of about 35% of DRI's membership. Representation from professional fields of work across the cultural heritage sector was captured in the results (note that some institutions gave multiple responses): 17 Archives, 12 Libraries, 5 Museums and 11 Higher Education Institutions. </p>
IN02005 Trivikrama Image Inscription of Tilaganga. Sanskrit XML file, draft epidoc edition
<p>IN02005 Trivikrama Image Inscription of Tilaganga. Sanskrit XML file (without metadata). Draft epidoc edition to be incorporated into 'Siddham' archive</p>
Bodhgayā, Bihar. Main image in the Mahābodhi temple.
<p>Bodhgayā, Bihar. Main image in the Mahābodhi temple, as placed in the sanctum by J. D. Beglar with the permission of the Mahant of Bodhgayā after 1876. British Museum, Department of Asia, Cunningham archive.</p>
MPS Data set with images of medieval charters for handwriting-style based dating of manuscripts
<pre>The MPS benchmark data set for handwritten manuscript dating ____________________________________________________________ This data set is collected for the Dutch NWO project: Medieval Paleographical Scale (MPS) by Petros Samara Project website: http://application02.target.rug.nl/monk/Projects/MPS/ Copyright (c) Huygensinstituut, Den Haag, 2016 University of Groningen, 2016. All rights reserved. Organisation of the data: Each .tar.gz file contains a number of NetPBM images. The format is chosen because of its simplicity. Also, there is no doubt about lossy compression in the processing chain. The file names are of the format 'MPS<year>_<seqnr>.ppm', for example, 'MPS1300_0056.ppm'. Note: the files are not in a separate directory, they will be extracted in place. However, due to the unique naming, there is no problem extracting them in one single current (destination) directory. The actual type of the image can be gray scale (.pgm) or color (.ppm), in '8-bit DirectClass' according to ImageMagick's 'identify' tool. The images were cropped out of larger photographs because of irrelevant elements such as a Kodak color calibrator and non-text content such as supporting surface (table) backgrounds, seals (emblems), ribbons, etc. No effort has been made to obtain a balanced set of samples over years: the given frequencies of occurrence in archives are used. There is evidently less data in years before 1375 A.D. while some periods provides us with ample data for historical reasons (e.g, 1450 A.D.). It would have been a pity if the scarce years had determined and limited the size of this data set. Selection criteria for data reduction, whether random or systematic, would have been arbitrary. In any case, these images were used in our publications, such that the performance results of future attempts on manuscript dating can be compared with earlier results. The performances that have been reached using our algorithms are in the order of an MAE (mean average error) of 10 years. If you have any questions, please contact us: Sheng He (heshengxgd@gmail.com) Petros Samara (petros.samara@huygens.knaw.nl) Jan Burgers (jan.burgers@huygens.knaw.nl) Lambert Schomaker (L.Schomaker@ai.rug.nl) Please cite our papers if you use this data set: [1] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Image-based historical manuscript dating using contour and stroke fragments. Pattern Recognition(PR), Vol. 59, pp. 159-171, 2016 [2] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Towards style-based dating of historical documents. International Conference on Frontiers in Handwriting Recognition(ICFHR), Crete, Greece, 2014 [3] Sheng He, Petros Samara, Jan Burgers, Lambert Schomaker. Multiple-Label Guided Clustering Algorithm for Historical Document Dating and Localization IEEE Trans. on Image Processing, Vol. 25(11), Nov. 2016. http://ieeexplore.ieee.org/document/7551181/</pre> <p>Data are collected thanks to Dutch NWO grant project 380-50-006</p>
Firemaker image collection for benchmarking forensic writer identification using image-based pattern recognition
<p>Disclaimer and terms of use:<br> ============================</p> <p>/*****************************************************************************\<br> * *<br> * *<br> * This is the Firemaker NFI-images Distribution *<br> * *<br> * This distribution contains 1000 images of scanned handwritten text, *<br> * scanned at resolution 300dpi grey scale, containing pages of *<br> * handwritten text by 250 writers, four pages per writer, from four *<br> * writing conditions, one condition per page. The conditions are: *<br> * p1: copied, natural style, p2: copied, UPPER case, p3: copied and forged, *<br> * i.e.,"try to write in a different style than your natural style", and p4, *<br> * self generated, i.e., text produced to describe a given cartoon. *<br> * *<br> * *<br> * *<br> * Copyright The International Unipen Foundation, 2000, All rights reserved *<br> *******************************************************************************<br> * *<br> * *<br> * DISCLAIMER AND COPYRIGHT NOTICE FOR ALL DATA CONTAINED ON THIS CDROM: *<br> * *<br> * *<br> * 1) PERMISSION IS HEREBY GRANTED TO USE THE DATA FOR RESEARCH *<br> * PURPOSES. IT IS NOT ALLOWED TO DISTRIBUTE THIS DATA FOR COMMERCIAL *<br> * PURPOSES. *<br> * *<br> * *<br> * 2) PROVIDER GIVES NO EXPRESS OR IMPLIED WARRANTY OF ANY KIND AND ANY *<br> * IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR PURPOSE ARE *<br> * DISCLAIMED. *<br> * *<br> * 3) PROVIDER SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, SPECIAL, *<br> * INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF ANY USE OF THIS *<br> * DATA. *<br> * *<br> * 4) THE USER SHOULD REFER TO THE FIRST PUBLIC ARTICLE ON THIS DATA SET: *<br> * *<br> * M. Bulacu, L. Schomaker & L. Vuurpijl (2003). *<br> * Writer identification using edge-based directional features. *<br> * ICDAR '03: Proceedings of the 7th International Conference on Document *<br> * Analysis and Recognition, pp. 937-941. *<br> * Piscataway: IEEE Computer, ISBN 0-7695-1960-1 *<br> * *<br> * 5) THE RECIPIENT SHOULD REFRAIN FROM PROLIFERATING THE DATA SET TO THIRD *<br> * PARTIES EXTERNAL TO HIS/HER LOCAL RESEARCH GROUP. PLEASE REFER INTERESTED *<br> * RESEARCHERS TO HTTP://UNIPEN.ORG FOR OBTAINING THEIR OWN COPY. *<br> \*****************************************************************************/</p> <p>BibTeX entry: </p> <p> @inproceedings{Firemaker, <br> author = {Bulacu, M. and Schomaker, L.R.B. and Vuurpijl, L.}, <br> title = {Writer Identification Using Edge-Based Directional Features},<br> booktitle = {ICDAR '03: Proceedings of the 7th International <br> Conference on Document Analysis and Recognition},<br> year = {2003},<br> isbn = {0-7695-1960-1},<br> pages = {937-941},<br> publisher = {IEEE Computer Society},<br> address = {Washington, DC, USA},<br> }</p> <p>In the project "Vergelijk", a grant obtained from the Dutch Forensic Science<br> Institute, two existing professional writer-identification systems have been <br> compared regarding usability studies and in particular recognition <br> performance (Schomaker & Vuurpijl, 2000). The results of this comparison <br> are contained in a confidential report:</p> <p> L.R.B. Schomaker and L.G. Vuurpijl (2000). <br> Forensic writer identification: A benchmark data set <br> and a comparison of two systems. Technical report, <br> Nijmegen Institute for Cognition and Information (NICI), <br> University of Nijmegen, The Netherlands.</p> <p>Informative and non-confidential details from this report are <br> given in the accompanying file: 'firemaker-dbase.pdf'</p> <p>To compare both systems, a carefully designed experiment was conducted to<br> record handwritten samples from male and female writers in several conditions:</p> <p>Condition 1: Normal constrained handwriting<br> ==============================================</p> <p>Below, the Dutch text writers had to produce in normal handwriting is given. </p> <p>--- start text ----<br> Zij bezochten veilingen en reisden met de KLM. Voor<br> korte afstanden huurden ze een auto, meestal een VW<br> of een Ford.<br> <EMPTY LINE><br> De veilingen waren van 7-4-1993 tot 3-5-1993 in New<br> York, Tokyo, Québec, Rome, Parijs, Zürich en Oslo.<br> <EMPTY LINE><br> Omdat de veilingen steeds begonnen om 12 uur en je<br> gemiddeld 200 tot 300 kilometer moest rijden,<br> stonden zij steeds om 6.30 uur op en vertrokken om<br> 8 uur uit het hotel.<br> <EMPTY LINE><br> Elke dag hadden ze vijfhonderd (f 500,-) gulden<br> nodig. Daarvoor gebruikten ze elke keer een cheque<br> van tweehonderd (f 200,-) en een cheque van<br> driehonderd (f 300,-) gulden. Aan geschenken gaven<br> ze ongeveer honderd gulden (f 100,-) uit.<br> --- end text ----</p> <p><br> Condition 2: Production of constrained block capital handwriting<br> ================================================================</p> <p>In this condition, the writers had to produce the following text<br> in block-capital handwriting:</p> <p>--- start text ----<br> NADAT ZE IN NEW YORK, TOKYO, QUÉBEC, PARIJS, ZÜRICH<br> EN OSLO WAREN GEWEEST, VLOGEN ZE UIT DE USA TERUG<br> MET VLUCHT KL 658 OM 12 UUR.<br> <empty line><br> ZE KWAMEN AAN IN DUBLIN OM 7 UUR EN IN AMSTERDAM OM<br> 9.40 UUR 'S AVONDS. DE FIAT VAN BOB EN DE VW VAN<br> DAVID STONDEN IN R3 VAN HET PARKEERTERREIN.<br> HIERVOOR MOESTEN ZE HONDERD GULDEN (F 100,-)<br> BETALEN.<br> --- end text ----</p> <p><br> Condition 3: Production of free-forged handwriting<br> ==================================================</p> <p>Below, the text writers had to produce in the free-forged handwriting<br> condition is given. No example of handwriting is given which they have to<br> mimick (forge), the condition concerns a self-conceived distorted <br> handwriting style.</p> <p>--- start text ----<br> Nog dezelfde avond reden ze naar hun vrienden<br> Chris, Emile, Jan, Irene en Henk, nadat ze hun<br> vriendinnen Greta en Maria hadden opgehaald.<br> <EMPTY LINE><br> Samen hadden ze vijfhonderd (500) zeldzame<br> postzegels gekocht, Bob driehonderd (300) en David<br> tweehonderd (200).<br> <EMPTY LINE><br> De reis was de moeite waard geweest.<br> --- end text ----</p> <p><br> Condition 4: Production of unconstrained handwriting<br> ====================================================</p> <p>The final text writers had to produce is unconstrained handwriting.<br> The cartoon, a series of pictures concerning a 'UFO' landing had<br> to be described in their own words, in at least six lines of text.<br> See image file "space.gif".</p> <p><br> Thruth labels and writer identifications<br> ========================================</p> <p>Each writer has a unique id, specified as:</p> <p> id: {num}{set}<br> num: a three-digit number<br> set: either 01, 02, 03 or 04, identifying one of the 4 experiments</p> <p>The vast majority of the writers producing sets 01, 02 and 03 mimicked the<br> content and layout (empty lines) of the constrained texts they had to copy<br> sufficiently accurately, such that the example texts are a good indication of<br> the contents. However, as set 04 ("describe cartoon story") contains<br> unconstrained self-generated handwriting, the corresponding thruth labels had<br> to be extracted manually. The resulting label files are contained in the<br> directory ./300dpi/p4-self-natural/labels/</p> <p>Note: no letter, word, line or paragraph segmentation is provided with this<br> data set. The main text can be cropped easily. Since the orientation is<br> horizontal, projection techniques can be used to extract lines, using<br> a line-spacing parameter (~94 pixels line height) as an additional check. </p> <p><br> Overview of directories:</p> <p>300dpi/<br> p1-copy-normal/ Copying task, normal writing style <br> p2-copy-upper/ Copying task, UPPER-case <br> p3-copy-forged/ Copying task, instructed to mimic another script style<br> p4-self-natural/ Self-generated text, natural writing condition</p> <p>Note: the original raw collection contained writer #155, who has been removed<br> from this data set, as his first condition (p1) was started in upper case and<br> the page was not completed. Deleted files were 15501.tif, 15502.tif, 15503.tif<br> and 15504.tif.</p> <p>Note: the name of this data set (Firemaker) is a contraction of the names<br> Vuurpijl and Schomaker.</p> <p>Note b: Example of a cutout of essential handwritten text using NetPBM tools: <br> tifftopnm 15201.tif | pnmcut -left 50 -right 2400 -top 700 -bottom 3250 > handwriting.pgm</p> <p> For an experiment, the upper and lower halves of the resulting image were<br> usually used in the Schomaker & Bulacu studies to obtain two samples of <br> handwriting for a writer.</p> <p> http://www.ai.rug.nl/~lambert<br> http://www.ai.rug.nl/~bulacu</p> <p>Our features for writer identification:</p> <p>Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> L. Schomaker & M. Bulacu (2004). <br> Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. <br> IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.</p> <p>Marius Bulacu<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> Bulacu, M. & Schomaker, L.R.B. (2007). <br> Text-independent Writer Identification and Verification Using Textural and Allographic Features, <br> IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</p> <p>Axel Brink<br> http://www.ai.rug.nl/~axel/ 'Quill' feature<br> A.A. Brink, J. Smit, M.L. Bulacu, and L.R.B. Schomaker (2011). <br> Writer identification using directional ink-trace width measurements, <br> Pattern Recognition (July 2011), doi: 10.1016/j.patcog.2011.07.005<br> <br> These three feature groups (hinge, fraglets, quill) have been combined in<br> a single MS Windows application, GIWIS which is available for scientific<br> use upon request (schomaker@ai.rug.nl)</p> <p>Note c.</p> <p>The accompanying file 'Firemaker-writer-info.dat' contains some<br> writer information: <br> Column 1: writer identification code<br> Column 2: sex<br> Column 3: handedness, <br> Column 4: age in years<br> Column 5: major Western script group (print,cursive or mixed)<br> </p>
ImUnipen image data set for writer identification (N=208) - vectorial handwriting converted to usable images
<p><br> ==============<br> Terms of Usage<br> ==============</p> <p>The ImUnipen data set is intended for non-commercial, scientific use,<br> and is distributed under auspices of the Unipen Foundation.</p> <p>Please always refer to the following paper in IEEE PAMI when using<br> the ImUnipen data set:</p> <p> Bulacu, M.; Schomaker, L.<br> Text-Independent Writer Identification and Verification<br> Using Textural and Allographic Features<br> Pattern Analysis and Machine Intelligence, IEEE Transactions on<br> Volume 29, Issue 4, April 2007 Page(s):701 - 717</p> <p>The ImUnipen data set is derived from the Unipen (unipen.org)<br> data set of on-line (i.e., vectorial, xy) handwriting.<br> The xy-coordinates and a line-generator algorithm are used<br> to generate a raster image, as if the data were optically scanned.</p> <p>Contents: for 208 writers, there are two PNG images per writer of<br> an artificially constructed table of naturally written words (49MByte).<br> These words are pasted onto a white page. For systematics reasons,<br> we call such a page a Paragraph, see below.</p> <p>The file names are organized as (example):</p> <p> Writ990221.Doc01.Par00.png<br> Writ990221.Doc01.Par01.png</p> <p> meaning: writer number 990221, document 01 (there exists only Doc01)<br> and the image with artificial "paragraph" of isolated words "Par00"<br> and "Par01".</p> <p>The Par00 and Pa01 images are typically used as the query<br> and best match in a leave-one-out setting for writer identification.<br> For instance, Par00 is the query, and Par01 is added to the total set<br> of all other images as the attractor for an identification search.</p> <p>For these experiments, word labels are not given in this data set,<br> on purpose, as the goal is to test recognition-free writer identification<br> methods.</p> <p>For a description of the regular<br> Unipen data set, please visit http://unipen.org</p> <p>Lambert Schomaker constructed this set in 2005</p>
Raw images (photographs) of urban text scenes for camera-based Thai text recognition
<p>Raw image collection of city scenes in Thailand with text content.<br> Text is photographed from diffeent angles. Also morning and evening<br> photographs were taken in order to capture different lighting<br> conditions. The material, 309 images, was photographed in 2013 by<br> Bowornrat Sriman and volunteers.</p> <p>Example EXIF:<br> JPEG image data, Exif standard: [TIFF image data, little-endian,<br> direntries=13, height=2448, manufacturer=SAMSUNG, model=GT-I9300,<br> orientation=upper-right, xresolution=220, yresolution=228,<br> resolutionunit=2, software=I9300XXEMA2, datetime=2013:03:14 18:17:49,<br> GPS-Data, width=3264], baseline, precision 8, 3264x2448, frames 3</p> <p>The images are not labeled. The orientation (landscape/portrait) is<br> not corrected yet. This material was used in preparation of the publication:</p> <p>Sriman, B. & Schomaker, L. (2015).<br> Object Attention Patches for Text Detection and Recognition in Scene Images using SIFT,<br> Proceedings of the International Conference on Pattern Recognition Applications and<br> Methods: ICPRAM 2015. De Marsico, M., Figueiredo, M. & Fred, A.<br> (Eds.). Lisbon, Portugal: SciTePress, Vol. 1, p. 304-311 8 p.</p> <p>Please cite this publication when using these data.</p>
PlantVillage Disease Classification Challenge - Color Images
<p><br> This work is licensed under a <a href="http://creativecommons.org/licenses/by-sa/3.0/us/">Creative Commons Attribution-ShareAlike 3.0 United States License</a>.<br> <br> # Data origins<br> The dataset is originally hosted at <a href="https://www.crowdai.org/challenges/plantvillage-disease-classification-challenge">PlantVillage Disease Classification Challenge</a>.<br> We use the modified version in <a href="https://github.com/salathegroup/plantvillage_deeplearning_paper_dataset">this github repository</a> to do controlled experiments.<br> We only use the raw color images dataset and delete the unconventional characters in the classes directory name and `.csv` filenames.<br> <br> # Directory explanation<br> The `80-20` direcotry has multiple `.txt` files which contain the training (~80%), validation(~10%) and testing (~10%) datasets instances filenames and the corresponding label indexes. The validation dataset quantity is `5430` in all data separation. In our experiment code (not included in this archive), the validation and testing dataset are merged together.<br> <br> # Data usage<br> ## Replicate our experiments<br> We have used this dataset in writing our paper. The reference information can be seen at https://<a href="https://gitlab.com/huix/leaf-disease-plant-village">gitlab.com/huix/leaf-disease-plant-village</a>.<br> <br> ### Steps<br> 1. `cd` to the direcotry (e.g. `/home/usrname/plantvillage_deeplearning_paper_dataset`) that contains the `color` directory.<br> 2. run `python change_filename_prefix.py --prefix /home/usrname/plantvillage_deeplearning_paper_dataset` to modify the prefix path (which is `/home/h/plantvillage_deeplearning_paper_dataset` in our former generated datasets).<br> 3. Fin. You can use our <a href="https://gitlab.com/huix/leaf-disease-plant-village">opens ource codes repository</a> to do the later experiments.<br> <br> ## Generate your own training/validation/testing datasets<br> This data separation generating code isn't included in the dataset archive, it is in our open source code. Please see our <a href="https://gitlab.com/huix/leaf-disease-plant-village">open source code repository</a> for the detailed information.<br> If you have any questions, you can contact the author through email.<br> The email address is a QR code in the archive.</p>
Aegean v2.0 Simulated Test Image
<p>A radio sky image simulated for the purposes of testing v2.0 of the Aegean source finding algorithm.</p>
Data sets used for: Urban runoff velocity measurement with consumer-grade surveillance cameras and surface structure image velocimetry
<p>Original videos and reference bulk velocity and water depth data sets used to develop the study: <em>Urban runoff velocity measurement with consumer-grade surveillance cameras and surface structure image velocimetry.</em></p> <p>The reference bulk velocity and water depth data sets were obtained with the Nivus OFR Radar and Nivus NivuCompact sensors, respectively.</p>
Dataset of B-mode fatty liver ultrasound images
<p>The dataset used and described in: M. Byra, G. Styczynski, C. Szmigielski, P. Kalinowski. Ł. Michałowski4. R. Paluszkiewicz. B. Ziarkiewicz-Wróblewska, K. Zieniewicz. P. Sobieraj, A. Nowicki. Transfer learning with deep convolutional neural network for liver steatosis assessment in ultrasound images. International Journal of Computer Assisted Radiology and Surgery, 2018. DOI: 10.1007/s11548-018-1843-2. </p> <p>Please refer to the above work if you use the dataset in your research. </p> <p>Contact:<br> Michal Byra<br> Department of Ultrasound<br> Institute of Fundamental Technological Research<br> Polish Academy of Sciences, Warsaw, Poland<br> mbyra@ippt.pan.pl<br> byra.michal@gmail.com</p>
Blood Vessels Dataset obtained from Retina Images of Healthy and Diabetic Retinopathy Individual
<p>This dataset contains blood vessels image files extracted from publicly available fundus retina images</p>
Dataset of High Resolution Mammographic Images
<p>This dataset contains 138 high resolution mamographic images. Contrast Limited Adjustment Histogram Equalization (CLAHE) was used to enhance selected raw mamographic images.available in mammographic image analysis society (MIAS) database. </p>
In silico 2D photoacoustic imaging data
<p>Here you find the data that was used for the experiments in the paper <strong>Confidence estimation for machine learning-based quantitative photoacoustics</strong> by <em>Janek Gröhl</em>, <em>Thomas Kirchner</em>, <em>Tim Adler</em>, and <em>Lena Maier-Hei</em>n.</p>
Herbarium specimen image of Barleria cuspidata Champl., part of the collection of Botanic Garden and Botanical Museum Berlin
Part of a training dataset of scanned herbarium specimens. The data paper and a summary landing page will be published on Zenodo as it gets published.<br><br>Content of this deposition:<br><br>- A JSON-LD datafile listing the label data associated with this herbarium specimen. The Darwin and Dublin Core data standards are used for most values.<br>- A JPEG image file of the scanned herbarium sheet.<br>- A lossless TIFF image from which the JPEG image has been derived.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.