Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “digital curation”
The Curated Courier: Digital Text Corpora from the UNESCO Courier (1948–2020)
<p>Founded in 1948 as the official magazine of the United Nations Educational, Scientific and Cultural Organization, <i>The UNESCO Courier</i> represents an extraordinary resource for research on global themes in the humanities. The complete <a href="https://en.unesco.org/courier/archives">archive of the magazine</a> is available in PDF form through UNESCO. These files make it possible for users anywhere to read individual issues, but it does not allow for full-text searching, much less any of the computational text analysis methods that have recently made important advances in humanities research.</p><p>The Curated Courier 1.0 is a package of digital text corpora, text analysis tools, and supplementary materials that makes the complete archive of <i>The UNESCO Courier</i> from 1948 to 2020 machine-readable, accessible, and reusable for digital text analysis. </p><p>Here on Zenodo we publish two <i>Courier</i> corpora. The first corpus (curated_courier_article_corpus) consists of the texts of all articles published in the English-language edition of <i>The UNESCO Courier</i> between 1948 and 2020. For this corpus we have extracted and reconstructed the complete text of all articles, for example by pulling together non-contiguous pages where necessary and by removing non-article text (masthead, photo captions, letters to the editor, and so on). We have linked each article to a comprehensive curated metadata index, included in the download (document_index.csv).</p><p>The second corpus (curated_issues) compiles the complete text of all <i>Courier</i> issues (English-language edition), 1948-2020. To prepare this corpus we extracted text from <a href="https://en.unesco.org/courier/archives">the PDFs that UNESCO has made available</a>, used multiple modes of OCR, and rendered each issue as a simple text file. Our test of the OCR quality finds an average error rate of 0.7 %, which should be considered good quality.</p><p>Working data from the process can be found in our <a href="https://github.com/inidun/tagged_courier">GitHub repository "tagged Courier."</a> The products, text analysis tools, and additional documentation are in the <a href="https://github.com/inidun/curated_courier">repository "Curated Courier."</a></p><p>The text of <i>The UNESCO Courier</i> is <a href="https://courier.unesco.org/en/about">available in Open Access</a> under the Attribution-ShareAlike 3.0 IGO (CC-BY-SA 3.0 IGO) license, in the context of <a href="https://en.unesco.org/open-access/">UNESCO's open access publications policy</a>. This dataset is published under the most recent version of the same license: Attribution-ShareAlike 4.0 International (<a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0 Deed</a>).</p><p>These datasets was developed as part of the research project "International Ideas at UNESCO: Digital Approaches to Global Conceptual History" (INIDUN), led by Benjamin G. Martin at Uppsala University and funded by a grant from the Swedish Research Council (Vetenskapsrådet), 2020-2024. For more information, see: <a href="https://inidun.github.io">https://inidun.github.io</a>, as well as the<a href="https://github.com/inidun"> project repository on GitHub</a>, which includes documentation and files related to the curating process.</p>
Facebook posts for analyzing Content Strategies for Digital Consumer Engagement: a curated dataset
<p>This database contains public data from publications made on Facebook by small and medium companies in the tourism sector operating in the Amazon region in Brazil. The collection was carried out from January to June 2018 using the public API provided by the platform.</p> <p>These data were processed and classified by different evaluators according to the content categories proposed by <a href="https://doi.org/10.1080/00913367.2017.1405751">Gavilanes</a>. Thus, there is a column in the dataset called "category", which contains the classification of each publication, with each number is associated with a category, as described below:</p> <ul> <li> <p>“0” - No category</p> </li> <li> <p>“1” - New product announcement</p> </li> <li> <p>“2” - Sweepstakes and contest</p> </li> <li> <p>“3” - Sales</p> </li> <li> <p>“4” - Consumer Feedback</p> </li> <li> <p>“5” - Infotainment</p> </li> <li> <p>“6” - Organization Branding</p> </li> <li> <p>“N\A” - Non-agreement between evaluators</p> </li> </ul> <p>Therefore, the columns with data present in the CSV file are:</p> <ul> <li> <p>“status_id” - Identification of each publication in the social network. String Textual content of each publication;</p> </li> <li> <p>“status_message” - The textual content of each publication;</p> </li> <li> <p>“link_name” - Which part of the profile the post are related;</p> </li> <li> <p>“status_type” - The type of the publication is a ’photo’, ’video’, ’status’ and/or ’link’;</p> </li> <li> <p>“status_link” - URL to the publications on the platform;</p> </li> <li> <p>“status_published” - Publication date in the social network;</p> </li> <li> <p>“num_comments” - Numbers of comments made by users;</p> </li> <li> <p>Reactions - Numbers of emoticons reactions from users on post – these numbers are presents in the columns: "num_reactions", "num_shares", "num_likes", "num_loves", "num_wows", "num_hahas", "num_sads", "num_angrys", "num_special";</p> </li> <li> <p>“category” - The category of content that we mentioned before.</p> </li> </ul>
Figure 8 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 8 - Whole-drawer image of dragonfly specimens used for a pilot study investigating the error associated with direct and indirect measures of morphological characters, such as wing length.
Figure 5 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 5 - Inset from previous figure (Figure 4). Label data attached to small specimens is often almost completely readable. Therefore, specimen metadata could be extracted and digitised using specialised character recognition software.
Figure 6 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 6 - Specimen with QR Code containing label data. A smart phone with the appropriate software can read and access the label data for this specimen from the image.
Figure 7 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 7 - Ultra high-resolution image of Buforaniidae grasshoppers (Orthoptera) from the ANIC. Note that the specimens are arranged by species, and then by the State from which they were collected. In this example, Northern Territory specimens are pinned in the first and second columns, followed by Queensland specimens in columns three and four. The online version of this image is viewable at Morphbank-ALA.
Figure 4 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 4 - Whole-drawer image of unsorted Hemiptera specimens with identifications provided by a remotely located expert, Dr Murray Fletcher. This drawer was subsequently re-curated according to the identifications, with specimens accessioned into the appropriate locations within the ANIC Hemiptera collection. See Appendix 1 for full list of remote identifications.
Figure 3 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 3 - A whole-drawer image displayed in MorphbankALA for online for viewing, editing and download. Image properties: 17,003x16,425 pixels, 30 MB (JPEG), and 464 MB (LZW compressed TIFF).
Figure 2 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 2 - Workflow process in ANIC to Digitise whole drawers of insects and load images into Morphbank-ALA
Figure 1 from: Mantle B, LaSalle J, Fisher N (2012) Whole-drawer imaging for digital management and curation of a large entomological collection. ZooKeys 209: 147-163. https://doi.org/10.3897/zookeys.209.3169
Figure 1 - The SatScan imaging system used in ANIC. Shown here with the front cover removed.
digital curation
<p>digital curation</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.