Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
105
datasets available to search
ShareScore release 0.9.0
Dataset results
105 results for “Ground Truth”
REPUBLIC PageXML ground truth handwritten resolutions States General
<p><strong>Using annotation software provided through the Transkribus Platform we annotated scans, concerning mostly 17th century handwritten documents from the National Archive of the Netherlands, with their textual transcriptions. The resulting ground truth was used to train a machine learning model yielding very accurate results. The ground truth was made available as an open access dataset. </strong></p>
ViF-GTAD: A new Automotive Data Set with Ground Truth for ADAS/AD Development, Testing and Validation
<p>A new dataset for automated driving, which is the subject matter of this paper, identifies and addresses a gap in existing similar perception data sets. While the most state-of-the-art perception data sets primarily focus on provision of various on-board sensor measurements along with the semantic information under various driving conditions, the provided information is often insufficient since the object list and position data provided include unknown and time-varying errors. The current paper and the associated data-set describes the first publicly available perception measurement data that include not only the on-board sensor information from camera, Lidar and radar with semantically classified objects, but also the high precision ground-truth position measurements enabled by the accurate RTK assisted GPS localization systems available on both the ego vehicle and the dynamic target objects. This paper provides insight on the capturing of the data, explicitly explaining the meta data structure and the content, as well as the potential application examples where it has been, and can potentially be, applied and implemented in relation to automated driving and environmental perception systems development, testing and validation.</p>
HawaiiCoast_GT: Curated AIS for Hawaii's coast correlated with ground truth incidents
<p>Because of the high-risk nature of emergencies and illegal activities at sea, it is critical that algorithms designed to detect anomalies from maritime traffic data be robust. However, there exist no publicly available maritime traffic datasets with real-world labelled anomalies. As a result, most anomaly detection algorithms for maritime traffic are validated without ground truth. We introduce the HawaiiCoast_GT dataset, the first ever publicly available automatic identification system dataset with a large corresponding set of true anomalous incidents. This dataset—cleaned and curated from Bureau of Ocean Energy Management (BOEM) and National Oceanic and Atmospheric Administration (NOAA) automatic identification system (AIS) data--covers Hawaii’s coastal waters for four years (2017-2020) and contains 88,749,176 AIS points for a total of 2,622 unique vessels. 208 tracks are labelled corresponding to 154 labelled real-world incidents. The codebase used to curate the original AIS data is being made openly available on GitHub.</p>
Ground-Truthing Satellite Imagery with Phenological Observations: Visual Observations from Grasslands at the Sevilleta National Wildlife Refuge, New Mexico
Phenology is the study of recurring natural phenomena. The seasonal "greening-up" and "greening-down" of dominant vegetation can be used as a predictor for a variety of processes and variables at local to global scales. The use of satellites to monitor land surface phenology is important for understanding local and regional ecosystem variability, identifying change over time, and potentially predicting ecosystem response to short and long-term changes in climate. However, the relationship between how phenology is expressed on the ground and how it is interpreted from satellites is poorly understood because phenological stages do not always correspond well to changes in spectral reflectance. In this study, we explored the relationship between greenness as measured by digital camera, the human eye, and ASTER imagery in two perennial grasslands at the Sevilleta National Wildlife Refuge in central New Mexico.
Comparative Study of Data-driven Solar Coronal Field Models Using a Flux Emergence Simulation as a Ground-truth Data Set
<p>For a better understanding of magnetic field in the solar corona and dynamic activities such as flares and coronal mass ejections, it is crucial to measure the time-evolving coronal field and accurately estimate the magnetic energy. Recently, a new modeling technique called the data-driven coronal field model, in which the time evolution of magnetic field is driven by a sequence of photospheric magnetic and velocity field maps, has been developed and revealed the dynamics of flare-productive active regions. Here we report on the first qualitative and quantitative assessment of different data-driven models using a magnetic flux emergence simulation as a ground-truth (GT) data set. We compare the GT field with those reconstructed from the GT photospheric field by four data-driven algorithms. It is found that, at least, the flux rope structure is reproduced in all coronal field models. Quantitatively, however, the results show a certain degree of model dependence. In most cases, the magnetic energies and relative magnetic helicity are comparable to or at most twice of the GT values. The reproduced flux ropes have a sigmoidal shape (consistent with GT) of various sizes, a vertically-standing magnetic torus, or a packed structure with curled field lines. The observed discrepancies can be attributed to the highly non-force-free input photospheric field, from which the coronal field is reconstructed, and to the modeling constraints such as the treatment of background atmosphere, the bottom boundary setting, and the spatial resolution.</p>
A ground-truth dataset to identify bots in GitHub
<p>This dataset is a ground truth dataset we used to identify bots. Each account in this dataset is rated by at least 3 raters with high interrater agreement.</p> <p>===</p> <p>This dataset is outdated (it was created in 2020) and therefore no longer recommended for use. Many of the classified GitHub bot accounts are no longer active or even available today, and some may even have changed their status from bot to human (or conversely) since. If you want to use a ground-truth dataset of bot accounts for academic (or other) purposes, we therefore recommend to use a more recent and more complete dataset of GitHub bot accounts. Such a dataset can be found here:</p> <p><a href="https://doi.org/10.5281/zenodo.7740520">https://doi.org/10.5281/zenodo.7740520</a></p> <p>===</p>
Evaluating registrations of serial sections with distortions of the ground truths. Supplemental data
<p><strong>Evaluating Registrations of Serial Sections With Distortions of the Ground Truths</strong></p> <p>This is the supplemental data for our paper on how to benchmark registrations of serial sections with ground truths. The files are named as follows:</p> <ul> <li>*_challenge.7z: local distortions and global rigid transformations applied, the input for the benchmark we used. Use this to test your rigid and non-rigid methods.</li> <li>*_local-only.7z: only local distortions applied.</li> <li>*_local-DIST.7z: the distortion maps for local distortions.</li> <li>*_SURF-rigid.7z: local distortions and global rigid transformations applied, rigid transformations undone with SURF-based rigid-only method. Local distortions remain. Use this if your method does not cope well with large rigid transformations.</li> <li>_*vis.7z: visualizations of distortions.</li> <li>_rigid_ground.7z: the real rigid transformations used in the global phase.</li> <li>*_ground.7z: the ground truth. All data fit each other, no distortions. Use this to compare your registration result to it.</li> </ul> <p>There are three main modalities and one further, as a reference:</p> <ul> <li>CT_*: µCT data, a rabbit lung, 600 images. (In ground truth, and local distortions, and global transformations we supply more images that went into the benchmark, 50 more from both beginning and end.)</li> <li>EM_*: an EM serial block-face (SBF-SEM) data set of adult mouse lung, 1000 images. (EM ground truth is individually normalized, see paper.)</li> <li>LS_*: a lung from the light sheet microscopy from a male 24 week-old rat, 300 images. (LS ground truth is individually normalized, too.)</li> <li>REAL_*: a region from real serial sections from a rabbit lung, 2 images.</li> </ul> <p>We also supply elastix parameter files.</p> <p>A preprint has been uploaded to <a href="https://arxiv.org/abs/2011.11060">arXiv</a>. The definite version is available from <a href="https://ieeexplore.ieee.org/abstract/document/9594850/media#media">IEEE</a>. The source code of the distorter is available from <a href="https://github.com/olegl/distort">GitHub</a>.</p>
Ground-truthing Phylotype Assignments for Antarctic Invertebrates
<p>Release from GitHub to Zenodo after proof correction (and branch merging). See article and README.md for further dataset descriptions.</p>
Ground Truth Dataset with mappings of companies to OpenCorporates legal entities
<p>Ground Truth Dataset with mappings of companies to OpenCorporates legal entities. Available data per company: company name, country of headquarters, state of headquarters (in case of US companies) and address of headquarters.</p>
Jingju a cappella singing pitch contour segmentation ground truth dataset
<p>The dataset used in the paper:</p> <blockquote> <p>Gong, Rong; Yang, Yile; Serra, Xavier; Pitch Contour Segmentation for Computer-aided Jingju Singing Training Sound and Music Computing (SMC 2016), 2016, Hamburg, Germany</p> </blockquote> <p>is in "dataset" folder. The a cappella singing audio recordings are not contained in this folder due to their large size, please contact the paper authors to request them (rong.gong@upf.edu). In the "dataset" folder you can find:</p> <ol> <li>ground truth</li> <li>Jinging singing scores in .xml format used for estimating the bigram note transition probabilities.</li> </ol> <p>The ground truth annotation is used for:</p> <ul> <li>melodic transcription (male_12_pos_1 missing)</li> <li>parameter optimization,</li> <li>evaluating the StdCdLe thresholding and the overall segmentation performance.</li> </ul> <p>The subfolder "groundtruth" contains the following annotation for each jingju a cappella audio:</p> <ul> <li>file name: description (format)</li> <li>*_melodicTrans.csv: melodic transcription ground truth used for the evaluation (start_time pitch duration -).</li> <li>*_coarseSeg.csv: StdCdLe ground truth used for the parameter optimization and the evaluation (segmentation points).</li> <li>*_refinedSeg.csv: ground truth used for optimizing other parameters and the evaluation (start_time - duration).</li> <li>*_pitchtrack.csv: pitch track (contour) extracted by pYIN pitch-tracking algorithm (filename time pitch).</li> <li>*_monoNoteOut.csv: notes estimated by pYIN note-tracking algorithm (filename start_time duration pitch).</li> </ul> <p> </p> <p> </p>
Ground truth data for ultrasound assessment of thoracolumbar fascia deformation/shearing
<div> <h2>Provided data/ultrasound videos</h2> </div> <div> <p>Here we provide ground truth data to validate ultrasound (US) measurement methods that assess shear or deformation of the thoracolumbar fascia (TLF). Studies that have done so are in the instance:</p> </div> <div> <ul> <li><strong>Langevin et al. (2011).</strong> Reduced thoracolumbar fascia shear strain in human chronic low back pain. BMC Musculoskeletal Disorders, 12, 203. <a href="https://doi.org/10.1186/1471-2474-12-203">https://doi.org/10.1186/1471-2474-12-203</a></li> <li><strong>Weber et al. (2022).</strong> The Influence of a Single Instrument-Assisted Manual Therapy (IAMT) for the Lower Back on the Structural and Functional Properties of the Dorsal Myofascial Chain in Female Soccer Players: A Randomised, Placebo-Controlled Trial. Journal of Clinical Medicine, 11(23), 7110. <a href="https://doi.org/10.3390/jcm11237110">https://doi.org/10.3390/jcm11237110</a></li> <li><strong>Brandl et al. (2023).</strong> Thoracolumbar fascia deformation during deadlifting and trunk extension in individuals with and without back pain. Frontiers in Medicine, 10, 1177146. <a href="https://doi.org/10.3389/fmed.2023.1177146">https://doi.org/10.3389/fmed.2023.1177146</a></li> <li><strong>Brandl et al. (2024).</strong> Quantifying thoracolumbar fascia deformation to discriminate acute low back pain patients and healthy individuals using ultrasound. Scientific Reports, 14(1), 20044. <a href="https://doi.org/10.1038/s41598-024-70982-7">https://doi.org/10.1038/s41598-024-70982-7</a></li> </ul> </div> <div> <p>The data contains ground truth of different velocities and distances from dynamic US measurements of the gel pad simulated the erector spinae muscle and upper tissue layers. For details regarding the gel pads, see <strong>Bartsch et al. (2023).</strong> Assessing reliability and validity of different stiffness measurement tools on a multi-layered phantom tissue model. Scientific Reports, 13(1), Article 1. <a href="https://doi.org/10.1038/s41598-023-27742-w">https://doi.org/10.1038/s41598-023-27742-w</a><br>For this reason, we have developed a customised device based on a linear guide driven by a stepper motor, with which the mimicked tissue layers can be moved with high precision. An explanatory video with detailed information on device setup can be found in the file: <strong>Device setup.mp4</strong></p> </div> <div> <p>Measurements for demographics were taken according to <strong>Brandl (2024).</strong> Ultrasound measurement of thoracolumbar fascia deformation. <a href="http://Protocols.io">Protocols.io</a>. <a href="https://dx.doi.org/10.17504/protocols.io.eq2lyjbmwlx9/v1">https://dx.doi.org/10.17504/protocols.io.eq2lyjbmwlx9/v1</a></p> </div> <div> <p>Measurement uncertainty for distance: <strong>+/- 0.002 mm k = 2 (95% confidence interval)</strong> and for speed: <strong>+/- 0.001 mm k = 2 (95% confidence interval)</strong> according to <strong>ISO (1993).</strong> Guide to the expression of uncertainty in measurement. International Organization for Standardization, Geneva.</p> </div> <div> <p>We further provide an example MATLAB script that demonstrate the use of the ground truth data with the open source image processing software Kinovea, <strong>Charmant, J., & contributors. (2023).</strong> Kinovea (Version 2023.1.1) Computer software. <a href="https://www.kinovea.org">https://www.kinovea.org</a></p> </div> <div> <p>A step-by-step validation protocol to facilitate the application of the ground truth data is provided along with a video showing the setup of the device in the root directory of the data.</p> <h2>Data structure</h2> </div> <div> <div> <table> <tbody> <tr> <th>Folder</th> <th>ODS Table</th> <th>Sheets</th> <th>Description</th> </tr> </tbody> <tbody> <tr> <td>TLFD_demographics</td> <td>TLFD_demographics</td> <td>demographics</td> <td>contains the data</td> </tr> <tr> <td> </td> <td> </td> <td>abbreviations</td> <td>abbreviations and units used in the table</td> </tr> <tr> <td>TLFD_GT\TLFD_GT_distances</td> <td>TLFD_GT_distances</td> <td>distances</td> <td>contains ground truth distance data and filenames of ultrasound</td> </tr> <tr> <td> </td> <td> </td> <td>abbreviations</td> <td>abbreviations and units used in the table</td> </tr> <tr> <td>TLFD_GT\TLFD_GT_speeds</td> <td>TLFD_GT_speeds</td> <td>speeds</td> <td>contains ground truth speed data and filenames of ultrasound</td> </tr> <tr> <td> </td> <td> </td> <td>abbreviations</td> <td>abbreviations and units used in the table</td> </tr> <tr> <td>MATLAB</td> <td> </td> <td> </td> <td>complete demo files for speckle tracking analysis with Kinovea</td> </tr> <tr> <td>MATLAB\Readme</td> <td> </td> <td> </td> <td>detailed instructions for using the MATLAB script</td> </tr> </tbody> </table> </div> </div> <p> </p>
EMG from Combination Gestures with Ground-truth Joystick Labels
<p>Dataset of surface EMG recordings from 11 subjects performing single and combination gestures, from "**A Multi-label Classification Approach to Increase Expressivity of EMG-based Gesture Recognition**" by Niklas Smedemark-Margulies, Yunus Bicer, Elifnur Sunger, Stephanie Naufel, Tales Imbiriba, Eugene Tunik, Deniz Erdogmus, and Mathew Yarossi.</p> <p>For more details and example usage, see the following:</p> <ul> <li>Paper pdf - <a href="https://arxiv.org/pdf/2309.12217.pdf">https://arxiv.org/pdf/2309.12217.pdf</a></li> <li>Experiment code - <a href="https://github.com/neu-spiral/multi-label-emg">https://github.com/neu-spiral/multi-label-emg</a></li> </ul> <h1>Contents</h1> <p>Dataset of single and combination gestures from 11 subjects. <br>Subjects participated in 13 experimental blocks.<br>During each block, they followed visual prompts to perform gestures while also manipulating a joystick.<br>Surface EMG was recorded from 8 electrodes on the forearm; labels were recorded according to the current visual prompt and the current state of the joystick.</p> <p>Experiments included the following blocks:</p> <ul> <li>1 Calibration block</li> <li>6 Simultaneous-Pulse Combination blocks (3 without feedback, 3 with feedback)</li> <li>6 Hold-Pulse Combination blocks (3 without feedback, 3 with feedback)</li> </ul> <p>The contents of each block type were as follows:</p> <ul> <li>In the Calibration block, subjects performed 8 repetitions of each of the 4 direction gestures, 2 modifier gestures, and a resting pose.<br>Each Calibration trial provided 160 overlapping examples, for a total of: 8 repetitions x 7 gestures x 160 examples = 8960 examples.</li> <li>In Simultaneous-Pulse Combination blocks, subjects performed 8 trials of combination gestures, where both components were performed simultaneously.<br>Each Simultaneous-Pulse trial provided 240 overlapping examples, for a total of: 8 trials x 240 examples = 1920 examples.</li> <li>In Hold-Pulse Combination blocks, subjects performed 28 trials of combination gestures, where 1 gesture component was held while the other was pulsed.<br>Each Hold-Pulse trial provided 240 overlapping examples, for a total of: 28 trials x 240 examples = 6720 examples.</li> </ul> <p>A single data example (from any block) corresponds a window 250ms of EMG recorded at 1926Hz (built-in 20–450 Hz bandpass filtering applied).<br>A 50ms step size was used between each window; note that neighboring data examples are therefore overlapping.</p> <p>Feedback was provided as follows:</p> <ul> <li>In blocks with feedback, a model pre-trained on the Calibration data was used to give realtime visual feedback during the trial.</li> <li>In blocks without feedback, no model was used, and the visual prompt was the only source of information about the current gesture.</li> </ul> <p>For more details, see the paper.</p> <h1>Labels</h1> <p>Two types of labels are provided: </p> <ul> <li>joystick labels were recorded based on the position of the joystick, and are treated as ground-truth.</li> <li>visual labels were also recorded based on what prompt was currently being shown to the subject.</li> </ul> <p>For both joystick and visual labels, the following structure applies. Each gesture trial has a two-part label.</p> <p>The first label component describes the direction gesture, and takes values in {0, 1, 2, 3, 4}, with the following meaning:</p> <ul> <li>0 - "Up" (joystick pull)</li> <li>1 - "Down" (joystick push)</li> <li>2 - "Left" (joystick left)</li> <li>3 - "Right" (joystick right)</li> <li>4 - "NoDirection" (absence of a direction gesture; none of the above)</li> </ul> <p>The second label component describes the modifier gesture, and takes values in {0, 1, 2}, with the following meaning:</p> <ul> <li>0 - "Pinch" (joystick trigger button)</li> <li>1 - "Thumb" (joystick thumb button)</li> <li>2 - "NoModifier" (absence of a modifier gesture; none of the above)</li> </ul> <h2>Examples of Label Structure</h2> <p>Single gestures have labels like (0, 2) indicating ("Up", "NoModifier") or (4, 1) indicating ("NoDirection", "Thumb").</p> <p>Combination gesture have labels like (0, 0) indicating ("Up", "Pinch") or (2, 1) indicating ("Left", "Thumb").</p> <h1>File layout</h1> <p>Data are provided in Numpy and MATLAB format. Descriptions below apply for both.</p> <p>Each experimental block is provided in a separate folder.<br>Within one experimental block, the following files are provided:</p> <ul> <li>`data.npy` - Raw EMG data, with shape (items, channels, timesteps).</li> <li>`joystick_direction_labels.npy` - one-hot joystick direction labels, with shape (items, 5).</li> <li>`joystick_modifier_labels.npy` - one-hot joystick modifier labels, with shape (items, 3).</li> <li>`visual_direction_labels.npy` - one-hot visual direction labels, with shape (items, 5).</li> <li>`visual_modifier_labels.npy` - one-hot visual modifier labels, with shape (items, 3).</li> </ul> <h1>Loading data</h1> <p>For example code snippets for loading data, see the associated code repository.</p>
Mind the leaf anatomy while taking ground truth with portable chlorophyll meters.
<p>Measurements of four chlorophyll meters — three transmittance-based (SPAD-502, Dualex-4 Scientific, and MultispeQ 2.0) and one fluorescence-based (CCM-300), were calibrated against biochemically assessed chlorophyll content (Chl) on three distinctive common leaf types differing in leaf anatomy: laminar (i.e., broadleaved woody species with different anthocyanin content verified by biochemical assay) dorsiventral leaves, narrow grass leaves, and conifer needles. Reflectance in the 400-2500 nm range was measured on the laminar leaf samples using a contact probe.</p> <p><strong>Methods</strong></p> <p>In the present study we investigated three distinctive leaf anatomical types: laminar (i.e., deciduous woody species) leaves, grass leaves, and needles. The three groups are defined as follows: 1) laminar leaves of woody angiosperms dorsiventrally flattened (i.e., bifacial) leaves with differentiated mesophyll to palisade and spongy parenchyma and reticulate anastomosing vasculature (Laminar leaves). <span>We further distinguished three anatomical subtypes of laminar leaves 1a) mesomorphic leaves of deciduous tree species, 1b) scleromorphic leaves of evergreen trees and shrubs, and 1c) scleromorphic leaves with pronounced hypodermis represented by <em>Ficus</em> species</span> 2) The second group included C3 grasses with bifacial strap-like leaves with undifferentiated mesophyll with longitudinally arranged vasculature (Grass leaves), and 3) gymnosperm equilateral needle-like leaves without differentiated mesophyll and vascular bundle in the central cylinder (Needles) represented only by Norway spruce (<em>Picea abies</em>) though with irradiance induced differentiation into sun and shaded ecotypes.</p> <p><strong><em>Collection of leaf samples</em></strong></p> <p>Leaves were collected at four different locations in the Czech Republic during the growing seasons 2019, 2020, and 2021. The dates (as DOY - day of the year are indicated in particular datasets). Woody plants with laminar bifacial leaves with differentiated mesophyll were collected in the Botanical Garden of Charles University in Prague (50.072N, 14.424E). Plants were selected to correspond to one of the following leaf subtypes: 1) mesomorphic leaves of deciduous species, 2) scleromorphic leaves of evergreen trees and shrubs and 3) scleromorphic leaves with pronounced hypodermis represented by indoor grown <em>Ficus</em> species. Usually, shaded leaves were sampled from the ground.</p> <p>For independent verification of the relationship of Chl content to chlorophyll meter reading, leaves were sampled in the floodplain forest at the confluence of the rivers Morava and Dyje, near the town of Lanžhot (48.682N, 16.946E) using deciduous woody plants with laminar bifacial leaves and differentiated mesophyll. Sunlit and shaded branches were cut by a tree climber from mature trees of <em>Acer campestre </em>L., <em>Carpinus betulus </em>L., <em>Fraxinus angustifolia </em>Vahl., <em>Populus alba </em>L., <em>Quercus cerris </em>L., <em>Quercus robur </em>L. and <em>Tilia cordata </em>Mill. </p> <p>Grass leaves were represented by four coexisting wild species from <em>Poaceae</em> family (<em>Calamagrostis villosa </em>(Chaix) J.F.Gmel., <em>Deschampsia cespitosa </em>(<a title="Carl Linnaeus" href="https://en.wikipedia.org/wiki/Carl_Linnaeus">L.</a>) <a title="Ambroise Marie François Joseph Palisot de Beauvois" href="https://en.wikipedia.org/wiki/Ambroise_Marie_Fran%C3%A7ois_Joseph_Palisot_de_Beauvois">P.Beauv.</a>, <em>Molinia caerulea </em>(<a title="Carl Linnaeus" href="https://en.wikipedia.org/wiki/Carl_Linnaeus">L.</a>) <a title="Conrad Moench" href="https://en.wikipedia.org/wiki/Conrad_Moench">Moench</a> and <em>Nardus stricta </em>L.) and were collected in relict alpine-arctic grass tundra in the Krkonoše (Giant Mountains) (50.734N, 15.696E). For each species, six plots with homogeneous canopy cover of the species were sampled. </p> <p>Needle leaves were represented by mature trees of Norway spruce (<em>Picea abies </em>(<a title="Carl Linnaeus" href="https://en.wikipedia.org/wiki/Carl_Linnaeus">L.</a>) <a title="Gustav Karl Wilhelm Hermann Karsten" href="https://en.wikipedia.org/wiki/Gustav_Karl_Wilhelm_Hermann_Karsten">H. Karst.</a>) collected at the experimental station Bílý Kříž, Beskydy Mountains, Czech Republic (49.503N, 18.539E). Sunlit and shaded branches were cut by a tree climber, and samples were taken from the current year's needles, the previous year's needles, and four-year-old needles. </p> <p><strong><em>Leaf sampling</em></strong></p> <p>Laminar leaves: leaves were measured immediately after being detached from the branch or stored in a refrigerator for no more than 30 minutes before processing. First, the reflectance of the leaves was measured using a spectroradiometer and a contact probe. Second, readings from all portable chlorophyll meters were recorded. Third, one disk (area = 68 mm<sup>2</sup>) was cut from each leaf for Chl and anthocyanin extraction. Finally, a square segment of the leaf was cut out and immersed in fixative solution for anatomical analysis. A second leaf of similar size, colour, position in the canopy, and developmental stage was removed from the branch, weighed, scanned, and later dried and weighed again. This "twin" was used to assess leaf mass per area (LMA), equivalent water thickness (EWT).</p> <p>Grass leaves: chlorophyll meter readings were taken on grass leaves attached to the plant using a chlorophyll meter (CCM<sub>CFR</sub>,) then leaves were collected immediately in the field for Chl extraction. A 2 cm long leaf segment was cut, flattened under a microscope glass, photographed for area assessment, and stored in plastic vials in a refrigerator before freezing. A subsample was weighed fresh, scanned, and dried for calculation of LMA and EWT.</p> <p>Needles: Shoots were separated from the branch, sorted by age, and stored in a refrigerator for no longer than 24 hours before processing. First, CCM<sub>CFR</sub> were taken from the middle part of three needles and the same needles were used for Chl extraction. A second parallel set of needles was immersed in fixative solution for anatomical analysis. The third set of needles was weighed fresh, scanned, and dried for calculation of LMA and EWT.</p> <p><strong><em>Optical assessment of Chl content using portable chlorophyll meters</em></strong></p> <p>Three transmittance-based chlorophyll meters: SPAD-502 SPAD), Dualex-4 Scientific (Dx) and MultispeQ (MSPQ), and one fluorescence chlorophyll meter: CCM-300 (CCM), were used for optical assessment of Chl content in leaves. For laminar leaves, three readings were taken on each leaf with each instrument from the adaxial leaf side. Measurements were taken in the central part of the leaf, avoiding the midrib and main veins. The three measurements were averaged, and the average was used as a representative value for the leaf. Measurements with all four chlorophyll meters (SPAD<sub>values</sub>, Dx<sub>values</sub>, MSPQ<sub>values</sub>, CCM<sub>CFR</sub>) were obtained for laminar leaves. For grass leaves, Chl values were measured at a single location in the apical third of the leaf blade. All four grass species were measured by CCM (CCM<sub>CFR</sub>), and three species with a wide enough lamina to cover the SPAD measurement area (<em>Calamagrostis villosa</em>, <em>Deschampsia cespitosa</em>, and <em>Molinia careulea</em>) were also measured by SPAD and SPAD<sub>values</sub> detected. For the spruce needles, three needles were measured only once with the CCM, always taking a reading in the central part of the needle. The average of these three needle measurements was used to relate to Chl.</p> <p><strong><em>Reflectance measurements and spectral processing</em></strong></p> <p>Reflectance was measured for laminar leaves collected in Botanical Garden of Charles University in Prague and deciduous trees from floodplain forest. Leaf reflectance from the adaxial side of the leaves was measured with an ASD FieldSpec 4 Wide-Res spectroradiometer with attached contact probe (ASD Inc., Boulder, CO, USA). Three measurements per leaf were always taken, when leaf size allowed. Measurements were placed at the same locations where chlorophyll meter readings were taken. Leaf reflectance spectra ranging from 350 to 2500 nm were normalized against a white reference spectrum (99% Spectralon white panel) to obtain relative reflectance spectra. The median of the spectral curve from three measurements was used as a representative value for the leaf. </p>
MC VIII SAMOIEDICA 2: JURAK-SAMOIEDICA 1: Line-aligned Ground Truth
<p>MC VIII SAMOIEDICA 2: JURAK-SAMOIEDICA 1: Line-aligned Ground Truth</p> <p>This dataset contains 172 microfilm scans Tundra Nenets materials, in which the text content is manually aligned line by line with the scanned images. This material has been created in collaboration between the Finno-Ugrian Society and the University of Innsbruck. It is intended specifically for handwritten text recognition experiments, training and benchmarking. For electronic materials and printed volumes that are intended to be used in linguistic, ethnographic and folkloric research, please refer to other publications in this Zenodo collection or [Manuscripta Castreaniana website](https://www.sgr.fi/manuscripta/).</p> <p>The materials were aligned in the University of Innsbruck with contributions by Günter Mühlberger and Günter Hackl. Other contributors are Karina Lukin and Niko Partanen. [Transkribus](https://readcoop.eu/transkribus/?sc=Transkribus) platform was extensively used in processing this dataset, and the file format is a direct Transkribus image and Page XML export.</p> <p> </p>
Castren 1844: Elementa grammaticae Syrjaenae, OCR Ground Truth
<p>Matthias Alexander Castrén published Elementa grammaticae Syrjaenae in 1844. This dataset contains two scans of this work, the higher quality version originating from the Internet Archive: </p> <p>https://archive.org/details/elementagrammati00cast/mode/2up</p> <p>All pages are layout detected with Transkribus, and there are 26 proofread pages. Three pages contain table layouts. </p>
Multiple Nuclei HeLa cell ground truth images with four labels (nuclear envelope, nucleus, rest of the cell, and background) for deep learning architecture training.
<p>This is a data set that contains <strong>labelled HeLa cell images</strong>, indicating the four different classes - nuclear envelope, nucleus, rest of the cell, and background. Similar ground truth have been published for this data set, but in this case, multiple nuclei have been labelled, whilst previous ones only focused on the central cell (https://doi.org/10.5281/zenodo.3874949)</p> <p>Details of the imaging, preparation and segmentation have been published in:</p> <ul> <li>Cefa Karabağ, Martin L. Jones, Christopher J. Peddie, Anne E. Weston, Lucy M. Collinson, Constantino Carlos Reyes-Aldasoro. Segmentation and Modelling of the Nuclear Envelope of HeLa Cells Imaged with Serial Block Face Scanning Electron Microscopy. <em>J. Imaging</em> <strong>2019</strong>, <em>5</em>(9), 75; <a href="https://doi.org/10.3390/jimaging5090075">https://doi.org/10.3390/jimaging5090075</a></li> <li>Cefa Karabağ, Martin L. Jones, Christopher J. Peddie, Anne E. Weston, Lucy M. Collinson, Constantino Carlos Reyes-Aldasoro. Semantic segmentation of HeLa cells: An objective comparison between one traditional algorithm and four deep-learning architectures, PLOS ONE, <strong>2020</strong>; <a href="https://doi.org/10.1371/journal.pone.0230605">https://doi.org/10.1371/journal.pone.0230605</a></li> <li> <p>Cefa Karabağ, Martin L. Jones, Constantino Carlos Reyes-Aldasoro, Segmentation of the Plasma Membrane of HeLa Cells,<em> J. Imaging</em> <strong>2021</strong>, <em>7</em>(6), 93; <a href="https://doi.org/10.3390/jimaging7060093">https://doi.org/10.3390/jimaging7060093</a></p> </li> </ul> <ul> <li>The data sets are freely available through EMPIAR: http://dx.doi.org/10.6019/EMPIAR-10094 EMPIAR.</li> </ul>
Eye blink events ground truth for UBFC video dataset
<p>In this dataset we annotate the timing of blinking events for all 42 videos in the UBFC dataset (available at https://sites.google.com/view/ybenezeth/ubfcrppg).</p> <p>There is one .txt file for each subject, for which each line corresponds to the frame number of a blinking event (videos @30fps).</p> <p>This dataset was created for testing the performance of blinking detection algorithms and is motivated by the potential of using blinking behaviour to assess the depressive condition through video teleconsultation.</p> <p> </p>
Biodiversity Metadata Ground Truth
<p>This repository contains ground truth data for 18 datasets that are collected from 7 Biodiversity data portals. We manually annotated the metadata fields according to the <a href="https://doi.org/10.5281/zenodo.6948519">Biodiversity Metadata Ontology (BMO)</a> ontology. This ground truth is used to evaluate the developed <a href="https://github.com/fusion-jena/Meta2KG">Meta2KG </a>approach that is used to transform raw metadata filed into RDF. </p>
OCR model for Pracalit for Sanskrit and Newar MSS 16th to 19th C., Ground Truth
<p>Ground truth data (png and xml files) for a an OCR model. Will be continually updated.</p> <p>Originally trained on Transkribus with a PyLaia model created from ground truth data based on transcripts into Pracalit Unicode of four Nepalese manuscripts. The manuscripts used to create this model are Staatsbibliothek zu Berlin's Hitopadeśa (MIK I 4851) (mixed Newar and Sanskrit dating to 1561) and Vetālapañcaviṃśati (HS. Or. 6414) (Newar dating to 1675) as well as Cambridge Digital Library's Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) (Sanskrit, 18th century) and the Royal Asiatic Society Online Collection's Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) (Newar and Sanskrit dating to c. 1800).</p> <p>The training was done on 441 pages and validation on 242 pages.</p> <p>This model does not recognise spacing, except for large gaps (i.e. for pictures or string holes). Newar word divider markers may not be represented or may be transcribed as virama. In general, the model is made for MSS with scriptio continua and will transcribe into scriptio continua into Pracalit Unicode.</p> <p>Transcription was performed by Dr Alexander O'Neill (SOAS University of London). Transcription of the Vetālapañcaviṃśati (HS. Or. 6414) and Madhyamasvayaṃbhūpurāṇa (RAS Hodgson MS 23) was aided by unpublished materials provided by Dr Felix Otter (Philipps-Universität Marburg), as well as the published transcription in Shakya, Min Bahadur, and Shanta Harsha Bajracharya, eds. "Svayambhū Purāṇa." Lalitpur: Nagarjuna Institute of Exact Methods, 2001. The transcription of Avalokiteśvaraguṇakāraṇḍavyūha (MS Add. 1322) was aided by the transcription provided by the Digital Sanskrit Buddhist Canon Project based on Lokesh Chandra, "Guṇakāraṇḍavyūhasūtram," New Delhi: International Academy of Indian Culture, 1999.</p>
6000 ground truth of VOC and notarial deeds 3.000.000 HTR of VOC, WIC and notarial deeds
<p>The National Archives of the Netherlands and Noord-Hollands Archief conducted a project using the Transkribus HTR (Handwritten Text Recognition) platform. The aim was to semi automatically transcribe 2 million pages of old Dutch texts.</p> <p>The transcribed archives are 17<sup>th</sup> and 18<sup>th</sup> century documents from the Dutch East-India Company (VOC). And 19th century notarial deeds from Noord-Hollands Archief and other archives in the provinces.</p> <p>In order to train the HTR software a team produced transcriptions of approximately 6000 scans. The scans are randomly selected from the dataset. With the transcriptions a model is trained that can recognize more than 90% of the characters correctly. Transkribus transcribed the 2 million scans automatically using the trained model.</p> <p>The following Transkribus HTR+ model has been trained for the text recognition: "IJsberg". More information about the model can be found <a href="https://readcoop.eu/transkribus/public-models/">here</a>. See the chapter "Dutch Handwriting". However, the Transkribus team retrained the model with <a href="https://readcoop.eu/transkribus/howto/how-to-train-pylaia-models-in-transkribus/">PyLaia</a> technology, which improved the HTR+ model. This PyLaia model is not publicly available.</p> <p>Later on, 1 million extra scans concerning the West India Company (WIC) were transcribed automatically without adding extra ground truth or training. These archives are from the 17<sup>th</sup> and 18<sup>th</sup> century.</p> <p>The <a href="https://github.com/knaw-huc/loghi">Loghi Handwritten Text Recognition Toolkit</a> has been added to the arsenal of the Nation Archives of the Netherlands. 1.05.11.14, Notarissen Suriname tot 1828 [digitaal duplicaat] has been processed with this tooling.</p> <p>The datasets published in Zenodo contain the ground truth (scans in JPG, transcription in PAGE XML) and the HTR results (in PAGE XML and TXT). See the overview below. Scroll to the bottom of the page to download the actual files.</p> <p>For more information on how the Dutch National Archive innovate on digital accessibility click <a href="https://www.nationaalarchief.nl/over-het-na/datalab-nationaal-archief">here</a>.</p> <p>For open data access of scans and inventories of the National Archives click <a href="https://www.nationaalarchief.nl/onderzoeken/open-data/open-data-archiefinventarissen-en-scans-van-archieven">here</a>.</p> <p><strong>Disclaimer</strong>: due to a variety of languages used and the bad state of the documents the HTR results of "1.05.21, Dutch series Guyana" can be of poor quality.</p> <p>--------------------------------------------------------------</p> <p><strong>Dataset HTR </strong><br>(Dataset, name archive, number archive, inventory numbers, link to inventory)<br><br><strong>The National Archives of the Netherlands</strong><br>HTR results VOC, VOC, 1.04.02, 7527-9540, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.04.02/invnr/%40Deel%20I?query=1.04.02&search-type=inventory">EAD</a><br>HTR results 1.04.02, Oost-Indische Testamenten, 1.04.02, 6847-6897, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.04.02/invnr/%40Deel%20I?query=1.04.02&search-type=inventory">EAD</a> <br>HTR results 1.05.01.01, Oude WIC, 1.05.01.01, 1-87, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.01.01/invnr/%40A..?query=1.05.01.01&search-type=inventory">EAD</a><br>HTR results 1.05.01.02, Tweede WIC, 1.05.01.02, 1-1382, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.01.02/invnr/%40VII~1324C2?query=1.05.01.01&search-type=inventory">EAD</a> <br>HTR results 1.05.02, Raad der Koloniën, 1.05.02, 1-192, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.02/invnr/%40A?query=1.05.02&search-type=inventory">EAD</a><br>HTR results 1.05.03, Sociëteit van Suriname, 1.05.03, 1-566, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.03/invnr/%40A?query=1.05.03&search-type=inventory">EAD</a><br>HTR results 1.05.05, Sociëteit van Berbice, 1.05.05, 1-445, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.05/invnr/%40I?query=1.05.05&search-type=inventory">EAD</a><br>HTR results 1.05.06, Verspreide West-Indische stukken, 1.05.06, 1-1413, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.06/invnr/%401?query=1.05.06&search-type=inventory">EAD</a><br>HTR results 1.05.21, Dutch series Guyana, 1.05.21, AB.1.1-BB.7.1, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.21/invnr/%401.?query=1.05.21&search-type=inventory">EAD</a><br>HTR results 2.01.28.01, West-Indisch comité, 2.01.28.01, 1-254, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/2.01.28.01/invnr/%40I?query=2.01.28.01&search-type=inventory">EAD</a><br>HTR results 2.01.28.02, Raad der Amerikaanse Bezittingen, 2.01.28.02, 1-264, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/2.01.28.02/invnr/%40I.?query=2.01.28.02&search-type=inventory">EAD</a><br>HTR results 1.05.11.14, Notarissen Suriname tot 1828 [digitaal duplicaat], <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.05.11.14/invnr/%401?query=1.05.11.14&search-type=inventory&start=0&searchAfter=1%2C%401">EAD</a><br>HTR results 2.10.02, Koloniën, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/2.10.02/invnr/%40A.?query=2.10.02&search-type=inventory&start=0&searchAfter=1%2C%40A.">EAD</a> (indices only)</p> <p><strong>Noord-Hollands archief</strong><br>HTR results NHA Notarial 1617, Oud notarieel archief Haarlem, 1617,1593-1805, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1617&milang=nl&miview=inv2">EAD</a><br>HTR results NHA Notarial 1972, Nieuw notarieel archief Haarlem, 1972, 5-813, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1972&milang=nl&miview=inv2">EAD</a></p> <p><strong>Brabants Historisch Informatie Centrum</strong><br>HTR results BHIC 7048 , Notarissen in Boxmeer, 1814-1935, 7048, 1-103, 162, <a href="https://proxy.archieven.nl/235/7203CD0B57624FD0B60EC6330DC7C083">EAD</a><br>HTR results BHIC 7128 , Notarissen in Grave, 1648-1935, 7128, 140-266, <a href="https://proxy.archieven.nl/235/73F9057F3EA34930829D1B55A9806B4D">EAD</a><br>HTR results BHIC 7637 , Notarissen in Sint-Oedenrode, 1642-1935, 7637, 17-78A, <a href="https://proxy.archieven.nl/235/4EBF145BB30A4CB3996DDD968914BAEE">EAD</a></p> <p><strong>Gelders Archief</strong><br>HTR results GA 0168, Notariële Archieven 1811-1925, 168, 64-69, 943-960, 1366-1395, 2472-2501, 3481-3485, 3904-3926, <a href="https://permalink.geldersarchief.nl/7A6A6A052F8A45EAAB6C7D4ECCA4041A">EAD</a></p> <p><strong>Groninger Archieven</strong><br>HTR results GRA 85, Notarissen te Appingedam (standplaats 1), 1811-1935, 85, 2-157, <a href="https://hdl.handle.net/21.12105/23FF3576C5AD4A97A847C6FB6E55323A">EAD</a><br>HTR results GRA 86, Notarissen te Appingedam (standplaats 2), 1812-1922, 86, 2-71, <a href="https://hdl.handle.net/21.12105/5C702F8661C6438EBC02025A9FC48446">EAD</a></p> <p><strong>Historisch Centrum Overijssel</strong><br>HTR results HCO 0122, Notarissen in Overijssel, 122, 5-48, 2044-2073, 3019-3047, 3733-3775, <a href="https://historischcentrumoverijssel.nl/archieven/?mivast=20&mizig=210&miadt=141&micode=0122&miview=inv2">EAD</a></p> <p><strong>The Utrecht Archives</strong><br>HTR results HUA 34-1, Notarissen in de provincie Utrecht, 1617-1895, 34-1, 928-930, 2209-2330, <a href="https://hetutrechtsarchief.nl/collectie/609C5BB45AB14642E0534701000A17FD">EAD</a></p> <p><strong>Regionaal Historisch Centrum Limburg</strong><br>HTR results RHCL 09.009, Notarissen in de Arrondissementen Maastricht en Roermond, 1896-1905, 09.009, 9147-9279, <a href="http://www.archieven.nl/mi/1540/?mivast=1540&mizig=210&miadt=38&micode=09.009&miview=inv2">EAD</a></p> <p><strong>Tresoar</strong><br>HTR results Tresoar 26, Notarieel archief, 26, 1001-9028 (met hiaten), <a href="https://www.archieven.nl/nl/zoeken?mivast=0&mizig=210&miadt=36&micode=26&miview=inv2">EAD</a></p> <p><strong>Zeeuws Archief</strong><br>HTR results ZA 13.2, Notariële Archieven Zeeland 1906-1915, (1886) 1906-1915 (1925), 13.2, 1152-1163, 1261-1320, <a href="https://hdl.handle.net/21.12113/DA16099B241C49C8843B404710612498">EAD</a></p> <p><strong>Drents Archief</strong><br>HTR results DA 114.10, Notaris jhr.mr. J.A.G.van der Wijck te Assen, 114.10, 4-7, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.10&miview=inv2">EAD</a><br>HTR results DA 114.11, Notaris mr. D.A.M.de Fremery te Assen, 114.11, 1, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.11&miview=inv2">EAD</a><br>HTR results DA 114.18, Notaris mr. Warmolt van Roijen te Borger, 114.18, 2-7, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.18&miview=inv2">EAD</a><br>HTR results DA 114.19, Notaris mr. Ernst Sigismund. Cornets de Groot te Borger, 114.29, 2-8, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.19&miview=inv2">EAD</a><br>HTR results DA 114.22, Notaris mr. Albertus Slingenberg te Coevorden, 114.22, 1-24, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.22&miview=inv2">EAD</a><br>HTR results DA 114.23, Notaris mr. Gozewienus Weys te Coevorden, 114.23, 4-13, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.23&miview=inv2">EAD</a><br>HTR results DA 114.28, Notaris mr. Johannes Beckeringh van Loenen te Dwingeloo, 114.29, 7-13, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.28&miview=inv2">EAD</a><br>HTR results DA 114.39, Notaris mr. Gerrit ten Raa ten Gieten, 114.39, 6-14, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.39&miview=inv2">EAD</a><br>HTR results DA 114.45, Notaris mr. Hendrik Jan Carsten te Hoogeveen, 114.45, 17-36, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.45&miview=inv2">EAD</a><br>HTR results DA 114.54, Notaris mr. Warmold Lunsingh Tonckens te Meppel, 114.54, 18-26, <a href="http://www.drentsarchief.nl/onderzoeken/archiefstukken?mivast=34&mizig=210&miadt=34&micode=0114.54&miview=inv2">EAD</a></p> <p> </p> <p><strong>Dataset Ground Truth</strong><br>(Name archive, number archive, inventory numbers, link to inventory, type of dataset)</p> <p>Dataset: Notarial deeds Ground Truths of the trainingset</p> <ul> <li>Oud notarieel archief Haarlem, 1617, 495 random scans from 1593-1805, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1617&milang=nl&miview=inv2">EAD</a>, GT Transcriptions</li> <li>Nieuw notarieel archief Haarlem, 1972, 952 random scans from 5-813, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1972&milang=nl&miview=inv2">EAD</a>, GT Transcriptions</li> <li>(And 168 transcripties from 7 other archives.)</li> </ul> <p>Dataset: Notarial deeds Images of the trainingset,</p> <ul> <li>Nieuw notarieel archief Haarlem, 1972, 952 random scans from 5-813, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1972&milang=nl&miview=inv2">EAD</a>, GT Scans</li> <li>Oud notarieel archief Haarlem, 1617, 495 random scans from 1593-1805, <a href="https://noord-hollandsarchief.nl/bronnen/archieven?mivast=236&mizig=210&miadt=236&micode=1617&milang=nl&miview=inv2">EAD</a>, GT Scans</li> <li>(And 168 scans from 7 other archives.)</li> </ul> <p><br>Dataset: VOC Ground Truths of the trainingset,<br>VOC, 1.04.02, 4735 random scans from 7527-9540, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.04.02/invnr/%40Deel%20I?query=1.04.02&search-type=inventory">EAD</a>, GT Transcriptions</p> <p><br>Dataset: VOC Images of the trainingset,<br>VOC, 1.04.02, 4735 random scans from 7527-9540, <a href="https://www.nationaalarchief.nl/onderzoeken/archief/1.04.02/invnr/%40Deel%20I?query=1.04.02&search-type=inventory">EAD</a>, GT Scans</p> <p>--------------------------------------------------------------</p> <p>Version 3.0: The first HTR results from the VOC-collection are available in .txt format, Inventory numbers 7527-9540.</p> <p>Version 3.1: The HTR results from the VOC-collection are also available in PAGE xml format. </p> <p>Version 4.0: About 30 missing inventory numbers have been added to the VOC transcriptions. The HTR results of the Notarial Deeds from the NHA archives have been added. An example on full text searchable research can be found here (Dutch): <a href="https://kia.pleio.nl/groups/view/55812425/htr-en-ocr/blog/view/55814752/reconstructie-van-een-verijdelde-slavenopstand-met-behulp-van-automatische-handschriftherkenning-en-text-mining">https://kia.pleio.nl/groups/view/55812425/htr-en-ocr/blog/view/55814752/reconstructie-van-een-verijdelde-slavenopstand-met-behulp-van-automatische-handschriftherkenning-en-text-mining</a></p> <p>Version 5.0: Around a million pages of HTR results of the following archives have been added.</p> <p>Version 6.0: The HTR results of Oost-Indische Testamenten have been added. </p> <p>Version 7.0: The HTR results of the Brabants Historisch Informatie Centrum, Gelders Archief, Groninger Archieven, Historisch Centrum Overijssel, The Utrecht Archives, Regionaal Historisch Centrum Limburg, Tresoar, Zeeuws Archief and Drents Archief have been added.</p> <p>Version 7.1: A spreadsheet "ijsberg train-val.xlsx" has been added. The division of the training- and validationset of Ground Truth of the IJsberg model can be found here</p> <p>Version 8.0: HTR results of 1.05.11.14 have been added. The scans have been inferenced with Loghi.</p> <p>Version 8.1: HTR results of 2.10.02 indices have been added. The scans have been inferenced with Loghi.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.