Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
65
datasets available to search
ShareScore release 0.9.0
Dataset results
65 results for “collections digitisation”
Inventory of criteria for prioritization of digitisation of collections focussed on scientific and societal needs
<p>Anno 2017 the task of mobilizing data from biocollections ahead of us is still enormous (data of 90% of the biocollections still needs to be mobilized). It is imperative for stakeholders, individual keepers of natural science collections, the community at large, and even for funding agencies, not only to tackle this backlog as quickly as possible, but do it in the best possible order. To establish the best possible order for digitizing biocollections a demand driven framework is required based among others on criteria used to digitize biocollections.</p>
Fig. 6.1. Shell digitised with different methods. The photogrammetry model was captured with a 100 in Handbook of best practice and standards for 2D+ and 3D imaging of natural history collections
Fig. 6.1. Shell digitised with different methods. The photogrammetry model was captured with a 100 mm Macro lens and processed with Agisoft Photoscan. The visual comparison of the mollusc shows a similar level of detail between photogrammetry and MechScan for the external surfaces, with still a bit more detail for the MechScan. The HDI Advance has a much lower resolution.
The Annotated Corpus of Classical Tibetan (ACTib), Part I - Segmented version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL.
<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, & Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., & Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p>
The Annotated Corpus of Classical Tibetan (ACTib), Part II - POS-tagged version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL
<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, & Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., & Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p> <p> </p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p>
Data from Wilson et al.: Applying computer vision to digitised natural history collections for climate change research: temperature-size responses in British butterflies
<p>This dataset supports the publication: Wilson et al. "Applying computer vision to digitised natural history collections for climate change research: temperature-size responses in British butterflies". These are the data used for the data figures (Fig 3-6, SI Figs 1-2) and the supplementary information tables.</p>
Supplementary material 1 from: Torralba-Burrial A, Merino-Sáinz I, Anadón A (2014) The relevance, biases, and importance of digitising opportunistic non-standardised collections: A case study in Iberian harvestmen fauna with BOS Arthropod Collection datasets (Arachnida, Opiliones). ZooKeys 404: 71-89. https://doi.org/10.3897/zookeys.404.6520
Harvestmen specimens included in this unplanned collection events subset.: Explanation note: Alternative link for download: http://hdl.handle.net/10651/24734
A collection of rendered videos that present the digitisation outcomes and the AR application of the Knossos Palace
<p>A collection of rendered videos that present the digitisation outcomes and the AR application of the Knossos Palace</p>
Fig. 6.7. Bone retoucher digitised with photogrammetry using a in Handbook of best practice and standards for 2D+ and 3D imaging of natural history collections
Fig. 6.7. Bone retoucher digitised with photogrammetry using a zoom lens (Ptg 18–55), a fixed focal macro lens (Ptg 100), focus stacking photogrammetry (FS-Ptg 60), structured light (SL) and microCT (µCT).
Figure 1d from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 1d A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Fossilised animal skin (Natural History Museum 2009)
Figure 1b from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 1b A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Pinned insect specimen (Natural History Museum 2018)
Figure 1c from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 1c A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Microscope slide (Natural History Museum 2017)
Figure 1a from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 1a A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Herbarium specimen (Natural History Museum 2007a)
Figure 11 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 11 The distribution of languages across the specimen and herbaria. EN=English, FR=French, LA=Latin, ET=Estonian, DE=German, NL=Dutch, PT=Portuguese, ES=Spanish, SV=Swedish, RU=Russian, FI=Finnish, IT=Italian, ZZ=Unknown. The codes for the contributing herbaria are listed in Table 11 (from Dillen et al. 2019).
Supplementary material 1 from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Appendices
Figure 1e from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789
Figure 1e A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Liquid preserved specimen (Natural History Museum 2010)
Figure 2 from: Willemse L, Runnel V, Saarenmaa H, Casino A, Gödderz K (2020) Digitisation of private collections. Research Ideas and Outcomes 6: e57767. https://doi.org/10.3897/rio.6.e57767
Figure 2 Lower part of the "Field Trip Report" form on www.laji.fi, where observations and their details can be entered.
Figure 1 from: Willemse L, Runnel V, Saarenmaa H, Casino A, Gödderz K (2020) Digitisation of private collections. Research Ideas and Outcomes 6: e57767. https://doi.org/10.3897/rio.6.e57767
Figure 1 Upper part of the "Field Trip Report" form on the Notebook Service of the FinBIF portal (www.laji.fi), where details of the gathering event can be entered.
Supplementary material 4 from: Dixey K, Woodburn M, Hardy H, Livermore L, Smith VS (2020) Identification of provisional Centres of Excellence for digitisation of European natural science collections. Research Ideas and Outcomes 6: e57750. https://doi.org/10.3897/rio.6.e57750
WP7 MS45 Centres of Excellence - Service Descriptions
Supplementary material 3 from: Dixey K, Woodburn M, Hardy H, Livermore L, Smith VS (2020) Identification of provisional Centres of Excellence for digitisation of European natural science collections. Research Ideas and Outcomes 6: e57750. https://doi.org/10.3897/rio.6.e57750
WP7 MS45 Centres of Excellence - Service Requirements
Supplementary material 2 from: Dixey K, Woodburn M, Hardy H, Livermore L, Smith VS (2020) Identification of provisional Centres of Excellence for digitisation of European natural science collections. Research Ideas and Outcomes 6: e57750. https://doi.org/10.3897/rio.6.e57750
WP7 MS45 Centres of Excellence - Service x Levels
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.