Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

130

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

130 results for “digitising”

Learn how ShareScore rates datasets ↗
zenodo40/100

Digitisation of Weather Records of Seungjeongwon Ilgi: A Historical Weather Dynamics Dataset of the Korean Peninsula (1623-1910)

<p><strong>Introduction</strong></p> <p>This study has exploited the daily weather records of Seungjeongwon Ilgi from the NIKH database (http://sjw.history.go.kr/main.do). Seungjeongwon Ilgi is a daily record of the Seungjeongwon, the Royal Secretariat of the Joseon Dynasty of Korea. These diaries span from 1623 to 1910 and generally involve daily weather records in the entry header. Their observational site would be located in Seoul (N37&deg;35&prime;, E126&deg;59&prime;). We have exploited the weather records from the NIKH database and classified the daily weather using text mining method. We have also converted the report dates from the traditional lunisolar calendar to the Gregorian calendar, to better contextualise our data into the contemporary daily measurements.</p> <p><strong>Data</strong></p> <p>We provide different formats (csv, xlsx, json) to facilitate the usage of data. The main contents of data are listed as below.</p> <ul> <li><strong>ID</strong>: The unique identifier of a specific record in the metadata, which can also serve as the identifier to merge with external data in the NIKH digital database.</li> <li><strong>Traditional calendar</strong>: The original lunar dates in the NIKH digital database, which are listed in data format &quot;YYYY-MM-DD&quot;. More specifically, &quot;L0&quot; implies the leap year and &quot;L1&quot; implies the common year.</li> <li><strong>Leap</strong>: The identifier of a leap year.</li> <li><strong>Gregorian calendar</strong>: The Gregorian calendar date that converted by the traditional calendar date.</li> <li><strong>Weather Text</strong>: The text that describe the weather conditions. Specifically, multiple weather descriptions of the same day have been put together.</li> <li><strong>Flag</strong>: The computed value that indicates different combinations of weather conditions.</li> <li><strong>Volume</strong>: The volume of text in the original record.</li> <li><strong>Herbal Volume</strong>: The volume of text in the herbal record.</li> <li><strong>Sunny</strong>: A dummy variable that represents whether the weather description contains the expression of sunny.</li> <li><strong>Cloudy</strong>: A dummy variable that represents whether the weather description contains the expression of cloudy.</li> <li><strong>Rainy</strong>: A dummy variable that represents whether the weather description contains the expression of rainy.</li> <li><strong>Snow</strong>: A dummy variable that represents whether the weather description contains the expression of snow.</li> <li><strong>Wind</strong>: A dummy variable that represents whether the weather description contains the expression of wind.</li> </ul> <p><strong>Import Data</strong></p> <pre><code class="language-python"># Python # CSV file import pandas as pd data=pd.read_csv('~/SJWilgi_Seoul_Weather_YR1623_1910.csv',encoding="utf-8") # JSON file data=pd.read_json('~/SJWilgi_Seoul_Weather_YR1623_1910.json',encoding="utf-8") # Excel file data=pd.read_excel('~/SJWilgi_Seoul_Weather_YR1623_1910.xlsx') # Excel file</code></pre> <pre><code class="language-bash"># R # CSV file library(readr) data&lt;- read_csv("~/SJWilgi_Seoul_Weather_YR1623_1910.csv") # Excel file library(readxl) data &lt;- read_excel("~/SJWilgi_Seoul_Weather_YR1623_1910.xlsx")</code></pre> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

The Annotated Corpus of Classical Tibetan (ACTib), Part I - Segmented version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL.

<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, &amp; Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., &amp; Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p>

opencc-by-4.0Jul 2017View details →
zenodo36/100

The Annotated Corpus of Classical Tibetan (ACTib), Part II - POS-tagged version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL

<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, &amp; Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., &amp; Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p> <p>&nbsp;</p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p>

opencc-by-4.0Jul 2017View details →
dryad36/100

3D data obtained with a MicroScribe digitising arm and photogrammetry to address bioarchaeological research questions

<p>Virtual methods for studying human remains are becoming increasingly popular in bioarchaeology, and the rate of technological innovation in the last few years has been such that we now have multiple options to choose from when collecting data. This raises the question of whether datasets generated with different methods are transposable. In the study reported here, we investigated whether it is valid to combine 3D data obtained with a MicroScribe digitising arm and 3D data collected via photogrammetry. We did so by simulating a population-based analysis similar to those commonly undertaken in bioarchaeology. Our sample comprised 19 crania from two ethnic groups, Ancient Egyptians and Guanches, and the landmarks we employed pertained to facial shape.</p> <p>The analyses yielded several findings. First, we found that photogrammetry was significantly more precise than the MicroScribe digitising arm. Second, the photogrammetry-based method revealed the existence of facial shape differences between the two ethnic groups that were not captured by the MicroScribe-based method. Third, we found that the two methods did not consistently capture the same facial shapes—they did for one of the ethnic groups but not for the other. Fourth, the analyses indicated that using the two methods can result in ethnic group-level differences in facial shape when they are applied to individuals from a single ethnic group. Lastly, the two methods of data collection yielded different patterns of variation in facial shape. Together, these findings suggest that combining 3D landmark coordinates collected with a MicroScribe and those obtained via photogrammetry may introduce considerable error into an analysis, and, consequently, bioarchaeologists should be cautious about doing so.</p>

opencc-zeroOct 2022View details →
zenodo36/100

Video recording of the digitisation of the Knossos Palace

<p>Video recording of the digitisation of the Knossos Palace composed by registered aerial and terrestrial scans&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Images and 3D digitisations of Branding Heritage #3

<p>These files are 3D digitisations and images of Branding Heritage</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Images and 3D digitisations of Branding Heritage #4

<p>These files are 3D digitisations and images of Branding Heritage</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Images and 3D digitisations of Branding Heritage #5

<p>These files are 3D digitisations and images of Branding Heritage</p>

opencc-by-4.0Oct 2023View details →
dryad36/100

3D data obtained with a MicroScribe digitising arm and photogrammetry to address bioarchaeological research questions

Open the record for dataset details and reuse information.

publicOct 2022View details →
zenodo32/100

Data from Wilson et al.: Applying computer vision to digitised natural history collections for climate change research: temperature-size responses in British butterflies

<p>This dataset supports the publication: Wilson et al. &quot;Applying computer vision to digitised natural history collections for climate change research: temperature-size responses in British butterflies&quot;. These are the data&nbsp;used for the data figures (Fig 3-6, SI Figs 1-2) and the supplementary information tables.</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

Sample of Digitised Books - Images identified as Embellishments. c. 1510 - c. 1900. JPG

<p>A subsample of Digitised Books - Images identified as Embellishments. c. 1510 - c. 1900. JPG https://bl.iro.bl.uk/concern/datasets/59d1aa35-c2d7-46e5-9475-9d0cd8df721e?locale=en</p>

opencc-zeroFeb 2022View details →
zenodo32/100

Supplementary material 1 from: Torralba-Burrial A, Merino-Sáinz I, Anadón A (2014) The relevance, biases, and importance of digitising opportunistic non-standardised collections: A case study in Iberian harvestmen fauna with BOS Arthropod Collection datasets (Arachnida, Opiliones). ZooKeys 404: 71-89. https://doi.org/10.3897/zookeys.404.6520

Harvestmen specimens included in this unplanned collection events subset.: Explanation note: Alternative link for download: http://hdl.handle.net/10651/24734

opencc-by-4.0Apr 2014View details →
zenodo32/100

A collection of rendered videos that present the digitisation outcomes and the AR application of the Knossos Palace

<p>A collection of rendered videos that present the digitisation outcomes and the AR application of the Knossos Palace</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Images and 3D digitisations of Branding Heritage #2

<p>These files are 3D digitisations and images of Branding Heritage</p>

opencc-by-4.0Aug 2023View details →
zenodo28/100

Fig. 6.7. Bone retoucher digitised with photogrammetry using a in Handbook of best practice and standards for 2D+ and 3D imaging of natural history collections

Fig. 6.7. Bone retoucher digitised with photogrammetry using a zoom lens (Ptg 18–55), a fixed focal macro lens (Ptg 100), focus stacking photogrammetry (FS-Ptg 60), structured light (SL) and microCT (µCT).

opencc-by-4.0Apr 2020View details →
zenodo28/100

Figure 3 from: Dupont S, Humphries J, Butcher AJ, Baker E, Balcells L, Price BW (2020) Ahead of the curve: three approaches to mass digitisation of vials with a focus on label data capture. Research Ideas and Outcomes 6: e53606. https://doi.org/10.3897/rio.6.e53606

Figure 3 A lateral image of ReVILE with the side panel removed to show the camera (a), turntable (b), stepper motor (c), Arduino Uno and motor driver controllers (d), front light panels (e), back light panel (f), light shield (g) and position of vial (arrow).

opencc-by-4.0May 2020View details →
zenodo28/100

Figure 1 from: Dupont S, Humphries J, Butcher AJ, Baker E, Balcells L, Price BW (2020) Ahead of the curve: three approaches to mass digitisation of vials with a focus on label data capture. Research Ideas and Outcomes 6: e53606. https://doi.org/10.3897/rio.6.e53606

Figure 1 MALICE vial setup showing the acrylic mirrors (a) LEGO mirror tilt arms (b), formex base (c) and LEGO friction joint (d)

opencc-by-4.0May 2020View details →
zenodo28/100

Figure 5 from: Dupont S, Humphries J, Butcher AJ, Baker E, Balcells L, Price BW (2020) Ahead of the curve: three approaches to mass digitisation of vials with a focus on label data capture. Research Ideas and Outcomes 6: e53606. https://doi.org/10.3897/rio.6.e53606

Figure 5 MALICE: Image output of MALICE including original output image (a) and the final processed image (b).

opencc-by-4.0May 2020View details →
zenodo28/100

Figure 7 from: Dupont S, Humphries J, Butcher AJ, Baker E, Balcells L, Price BW (2020) Ahead of the curve: three approaches to mass digitisation of vials with a focus on label data capture. Research Ideas and Outcomes 6: e53606. https://doi.org/10.3897/rio.6.e53606

Figure 7 ReVILE: Two Rollout photography outputs of ReVILE. Left to right: a frame from the video (rotated 90° clockwise) showing the vial itself; the uncropped rollout image, covering more than one full rotation; the cropped rollout image, showing only one 360° rotation; the cropped rollout image, "shifted" across (by transferring a manually-defined block of pixel columns from the left side of the image to the right) to show the complete label

opencc-by-4.0May 2020View details →
zenodo28/100

Figure 1d from: Owen D, Livermore L, Groom Q, Hardisty A, Leegwater T, van Walsum M, Wijkamp N, Spasić I (2020) Towards a scientific workflow featuring Natural Language Processing for the digitisation of natural history collections. Research Ideas and Outcomes 6: e55789. https://doi.org/10.3897/rio.6.e55789

Figure 1d A range of sample specimens that demonstrate the wide taxonomic range of specimens encountered in collections. They also demonstrate the diversity of label types, which include handwritten, typed, and printed labels. Note the presence of various barcodes, rulers, and a colour chart in addition to labels describing the origin of the specimen and its identity. - Fossilised animal skin (Natural History Museum 2009)

opencc-by-4.0Jul 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record