Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
95
datasets available to search
ShareScore release 0.9.0
Dataset results
95 results for “digitized collection”
Figure 1 from: Sikes DS, Copas K, Hirsch T, Longino JT, Schigel D (2016) On natural history collections, digitized and not: a response to Ferro and Flick. ZooKeys 618: 145-158. https://doi.org/10.3897/zookeys.618.9986
Figure 1 - Number of insect records in GBIF.org (triangles) between December 2007 and March 2016, in comparison to all records (circles).
LadiesDebating_HTO: A Knowledge Graph for representing the "Edinburgh Ladies' Debating Society Digital Collection" (1865 - 1880) following Heritage Textual Ontology
<p>This Knowlege Graph represents the information of the "Edinburgh Ladies’ Debating Society<strong>"</strong> (years: 1865 - 1880) collection in RDF (ttl format). This collection consists of the complete runs of two Edinburgh journals, <strong>‘The Attempt’ (10 volumes, 1865-74)</strong> and its successor ‘<strong>The Ladies’ Edinburgh Magazine’ (6 volumes, 1875-80)</strong>. These publications were produced by a leading Edinburgh women’s club, known during the period as the Edinburgh Essay Society or the Ladies’ Edinburgh Essay Society, but subsequently as the Ladies’ Edinburgh Debating Society. The Society existed from 1865 to 1935. The raw dataset is provided by the NLS in this <a href="https://data.nls.uk/data/digitised-collections/edinburgh-ladies-debating-society/">link</a>. As other NLS data collections, they are originally provided using two XMLs schemas: METS for descriptive, structural, technical and administrative metadata (Title, Author, Publisher, etc); and ALTO for encoding the OCR text of a page.</p> <p>In this work, we have extracted the information from METS and ALTO XMLS using <a href="https://github.com/francesNLP/defoe">defoe</a> tool. The KG uses the <a href="https://w3id.org/hto">HTO</a> to represent the information extracted. Furthermore, during the information extraction phase, we have employed several techniques to mitigate two common OCR errors: long-S and the line-break hyphenation.</p>
The survey of the Marega Collection at the Vatican Library and the construction of a digital open access database (マリオ・マレガ収集資料の調査とデータベース)
<p>Paper presented on Friday 11 June 2021 at the Digital Medievalist Global Symposium <em>The past, present, and future of Digital Medieval Studies</em> for the Asia & Oceania Panel, in the session Reading Indic and Japanese scripts.</p> <p>In 2011, about 14,000 documents related to the ban of Christianity in Japan and the surveillance over the family of former Christians were found at the Vatican Library. The documents, roughly spanning from the 17th to the 19th century, were originally collected by Father Mario Marega, a Salesian missionary, who resided in the Oita Prefecture, Kyushu, Japan, since 1929. Marega then sent the documents to the Vatican Library in the 1950s, where, for various reasons, they were set aside and forgotten until their rediscovery during the Library’s renovation works in 2010. Researchers from Italy and Japan have conducted research for a decade to catalogue and publish the collection, recently also made available in a digital database. In this presentation, we will introduce the surveying method and the guiding principles of the database structure and its data model of the <a href="https://base1.nijl.ac.jp/~marega/en/">Mario Marega Archive</a>. We will also elaborate on the results achieved and the issues that the research process and the database definition presented.</p> <p>2011年にバチカン図書館で17ー19世紀の日本におけるキリスト教統制に関する文書(約14,000点)が発見された。1929年より日本の大分県に滞在したイタリアのマリオ・マレガ神父が収集し、1950年代にバチカン図書館へ送ったものである。これらを研究・公開するため、イタリア・日本の研究者が共同で調査およびデータベース構築を行っている。今回の発表では、調査の方法やデータベースの考え方などを紹介し、その成果と課題を展望する。</p> <p> </p> <p> </p> <p> </p>
Figure 4 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 4 - Specimen image processing. Using Adobe Photoshop Lightroom software to process images. New York Botanical Garden.
Figure 3 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 3 - Custom specimen holder. Museum of Compartive Zoology (MCZ) Rhopalocera (Lepidoptera) Rapid Digitization Project.
Figure 3 from: Blagoderov V, Kitching I, Livermore L, Simonsen T, Smith V (2012) No specimen left behind: industrial scale digitization of natural history collections. ZooKeys 209: 133-146. https://doi.org/10.3897/zookeys.209.3178
Figure 3 - Metadata Creator software: a–c working areas a drawer image b specimen records c annotation fields d tool selector e unique IDs.
Figure 2 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 2 - Specimen image capture. Fossil specimen imaging, specimen label imaging. Two very different imaging set-ups. Yale Peabody Museum, University of Kansas - Entomology.
Figure 1 from: Blagoderov V, Kitching I, Livermore L, Simonsen T, Smith V (2012) No specimen left behind: industrial scale digitization of natural history collections. ZooKeys 209: 133-146. https://doi.org/10.3897/zookeys.209.3178
Figure 1 - SatScan imaging: a SatScan machine b specimens being imaged c individual frames aligned d fragment of a stitched image; final resolution of the stitched image ~11 lines/mm.
Figure 1 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 1 - Pre-digitization specimen curation and staging. Preparing barcodes and imaging labels, affixing barcodes, updating taxonomy. L to R: University of Kansas – Entomology, New York Botanical Garden and Yale Peabody Museum.
Figure 8 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 8 - Whole-drawer image of dragonfly specimens used for a pilot study investigating the error associated with direct and indirect measures of morphological characters, such as wing length.
Figure 2 from: Blagoderov V, Kitching I, Livermore L, Simonsen T, Smith V (2012) No specimen left behind: industrial scale digitization of natural history collections. ZooKeys 209: 133-146. https://doi.org/10.3897/zookeys.209.3178
Figure 2 - Image based digitization workflow consisting of four stages: Imaging, Metadata capture, Institutional databading and Publication.
Figure 5 from: Nelson G, Paul D, Riccardi G, Mast A (2012) Five task clusters that enable efficient and effective digitization of biological collections. ZooKeys 209: 19-45. https://doi.org/10.3897/zookeys.209.3135
Figure 5 - Electronic data capture. Entering data straight from the specimen label into the database. New York Botanical Garden.
Figure 5 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 5 - Inset from previous figure (Figure 4). Label data attached to small specimens is often almost completely readable. Therefore, specimen metadata could be extracted and digitised using specialised character recognition software.
Figure 6 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 6 - Specimen with QR Code containing label data. A smart phone with the appropriate software can read and access the label data for this specimen from the image.
Figure 7 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 7 - Ultra high-resolution image of Buforaniidae grasshoppers (Orthoptera) from the ANIC. Note that the specimens are arranged by species, and then by the State from which they were collected. In this example, Northern Territory specimens are pinned in the first and second columns, followed by Queensland specimens in columns three and four. The online version of this image is viewable at Morphbank-ALA.
Figure 4 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 4 - Whole-drawer image of unsorted Hemiptera specimens with identifications provided by a remotely located expert, Dr Murray Fletcher. This drawer was subsequently re-curated according to the identifications, with specimens accessioned into the appropriate locations within the ANIC Hemiptera collection. See Appendix 1 for full list of remote identifications.
Figure 4 from: Dietrich C, Hart J, Raila D, Ravaioli U, Sobh N, Sobh O, Taylor C (2012) InvertNet: a new paradigm for digital access to invertebrate collections. ZooKeys 209: 165-181. https://doi.org/10.3897/zookeys.209.3571
Figure 4 - Image of multiple pinned insect specimens in unit tray (left) and same specimens segmented into separate files (right) using customized ImageJ image processing protocol.
Figure 2 from: Schmidt S, Balke M, Lafogler S (2012) DScan – a high-performance digital scanning system for entomological collections. ZooKeys 209: 183-191. https://doi.org/10.3897/zookeys.209.3115
Figure 2 - Partial drawer images taken at the same position using three different file formats: captured as JPEG (a), captured as TIFF (b), and TIFF converted from RAW (c). A high resolution version of the image is available under media.zsm-entomology.de/suppl/zookeys_mass_digitisation_volume/Fig_2.png
Figure 3 from: Dietrich C, Hart J, Raila D, Ravaioli U, Sobh N, Sobh O, Taylor C (2012) InvertNet: a new paradigm for digital access to invertebrate collections. ZooKeys 209: 165-181. https://doi.org/10.3897/zookeys.209.3571
Figure 3 - Current version of InvertNet's Medici multimedia semantic content management system interface, accessible from InvertNet digital collections tab on homepage, showing taxonomic tree, drag and drop file upload space, and zoomable user interface for viewing gigapixel images.
Figure 1 from: Schmidt S, Balke M, Lafogler S (2012) DScan – a high-performance digital scanning system for entomological collections. ZooKeys 209: 183-191. https://doi.org/10.3897/zookeys.209.3115
Figure 1 - Schematic drawing of the DScan system. Flashes (not shown) are placed inside the scanner.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.