Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
18
datasets available to search
ShareScore release 0.7.1
Dataset results
18 results for “Digital History”
Harnessing the power of digitized natural history collections to visualize spatiotemporal patterns in native and non-native bee flight phenology
<p>What time of year are bees flying, where are they flying, and how do biogeographical factors, sex, and native status affect flight phenology? Consistent monitoring along with creating spatially and temporally explicit visualizations using large openly available data sets enhance our understanding of trends in flight time phenology and shape our understanding of bee-plant interactions, including shifts in the phenology of bee pollinators.</p> <p>Species occurrence data from digitized collection networks (iNaturalist, Global Biodiversity Information Faculty (GBIF), Integrated Digitized Biocollections (iDigBio), Symbiota Collections of Arthropods Network (SCAN), and UC Santa Barbara Collection Network) are part of an effort to improve our understanding of bees in coastal Santa Barbara County, including the California Channel Islands. New inventory collections combined with historical data from over 11 natural history museums and 2 observation networks are used in an effort to examine patterns and changes in phenology of native and non-native bee species, and create updated species inventories.</p> <p>Synthesizing species observation data from digitized natural history collections makes use of a wealth of existing data and multiplies the analytical power of isolated observations, but it is not without limitations and challenges. By exploring novel techniques to generate clear and accurate visualizations to communicate bee flight time, we present our key initial findings and identify geographic, temporal, and taxonomic gaps, which will lead to further focused inventory projects of coastal Santa Barbara County, improved data quality for phenological analyses, and reusable methods for visualizing insect phenology data across taxa or geography.</p> <p><strong>The attached files include the R code and some of the .csv files used to produce the figures in my poster that was available on demand at the Entomology Society of America 2020 virtual meeting. </strong></p>
1805-1898 Census Records of Lausanne : a Long Digital Dataset for Demographic History
<p><strong>Context. </strong>This historical dataset stems from the project of automatic extraction of 72 census records of Lausanne, Switzerland. The complete dataset covers a century of historical demography in Lausanne (1805-1898), which corresponds to 18,831 pages, and nearly 6 million cells.</p> <p><strong>Content.</strong> The data published in this repository correspond to a first release, i.e. a diachronic slice of one register every 8 to 9 years. Unfortunately, the remaining data are currently under embargo. Their publication will take place as soon as possible, and at the latest by the end of 2023. In the meantime, the data presented here correspond to a large subset of 2,844 pages, which already allows to investigate most research hypotheses.</p> <p><strong>Description. </strong>The population censuses, digitized by the <a href="https://www.lausanne.ch/vie-pratique/culture/bibliotheques-et-archives/archives.html">Archives of the city of Lausanne</a>, continuously cover the evolution of the population in Lausanne throughout the 19th century, starting in 1805, with only one long interruption from 1814 to 1831. Highly detailed, they are an invaluable source for studying migration, economic and social history, and traces of cultural exchanges not only with Bern, but also with France and Italy. Indeed, the system of tracing family origin, specific to Switzerland, allows to follow the migratory movements of families long before the censuses appeared. The bourgeoisie is also an essential economic tracer. In addition, censuses extensively describe the organization of the social fabric into family nuclei, around which gravitate various boarders, workers, servants or apprentices, often living in the same apartment with the family.</p> <p><strong>Production. </strong>The structure and richness of censuses have also provided an opportunity to develop automatic methods for processing structured documents. The processing of censuses includes several steps, from the identification of text segments to the restructuring of information as digital tabular data, through Handwritten Text Recognition and the automatic segmentation of the structure using neural networks. Please note that the detailed extraction methodology, as well as the complete evaluation of performance and reliability is published in:</p> <ul> <li>Petitpierre R., Rappo L., Kramer M. (2023). <em>An end-to-end pipeline for historical censuses processing</em>. International Journal on Document Analysis and Recognition (IJDAR). doi: <a href="https://doi.org/10.1007/s10032-023-00428-9">10.1007/s10032-023-00428-9</a></li> </ul> <p><strong>Data structure.</strong> The data are structured in rows and columns, with each row corresponding to a household. Multiple entries in the same column for a single household are separated by vertical bars ⟨|⟩. The center point ⟨·⟩ indicates an empty entry. For some columns (e.g., street name, house number, owner name), an empty entry indicates that the last non-empty value should be carried over. The page number is in the last column.</p> <p><strong>Liability. </strong>The data presented here are not curated nor verified. They are the raw results of the extraction, the reliability of which was thoroughly assessed in the above-mentioned publication. We insist on the fact that for any reuse of this data for research purposes, the implementation of an appropriate methodology is necessary. This may typically include string distance heuristics, or statistical methodologies to deal with noise and uncertainty.</p>
Supplementary data for Plutniak, S. 2022. "What makes the identity of a scientific method? A history of the 'Structural and analytical typology' in the growth of evolutionary and digital archaeology in southwestern Europe (1950s–2000s)", Journal of Paleolithic Archaeology, vol. 5, 10.
<p>Supplementary data for Plutniak, S. 2022. “What makes the Identity of a Scientific Method? A History of the ‘Structural and analytical typology’ in the Growth of Evolutionary and Digital Archaeology in Southwestern Europe (1950s–2000s)”, <em>Journal of Paleolithic Archaeology</em>. vol 5, 10. DOI: <a href="https://doi.org/10.1007/s41982-022-00119-7">10.1007/s41982-022-00119-7</a>.</p> <p> </p> <p> </p> <p> </p>
Linked collectors and determiners for: Digitization PEN: Digital Data from the Terrestrial Polyneoptera at American Museum of Natural History.
Natural history specimen data linked to collectors and determiners held within, "Digitization PEN: Digital Data from the Terrestrial Polyneoptera at American Museum of Natural History". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/d475b3b1-1e5a-4d56-a547-f7afb810c4c5">https://bionomia.net/dataset/d475b3b1-1e5a-4d56-a547-f7afb810c4c5</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/d475b3b1-1e5a-4d56-a547-f7afb810c4c5">https://gbif.org/dataset/d475b3b1-1e5a-4d56-a547-f7afb810c4c5</a>. Formatted as a Frictionless Data package.
Linked collectors and determiners for: Historic Butterfly Specimens in the UVM Zadock Thompson Natural History Collection Digitized for the Vermont Butterfly Atlas.
Natural history specimen data linked to collectors and determiners held within, "Historic Butterfly Specimens in the UVM Zadock Thompson Natural History Collection Digitized for the Vermont Butterfly Atlas". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/a5d361b0-410b-44fd-9666-0a938c75abd1">https://bionomia.net/dataset/a5d361b0-410b-44fd-9666-0a938c75abd1</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/a5d361b0-410b-44fd-9666-0a938c75abd1">https://gbif.org/dataset/a5d361b0-410b-44fd-9666-0a938c75abd1</a>. Formatted as a Frictionless Data package.
Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History
<p>Throughout the history of art, the pose—as the holistic abstraction of the human body's expression—has proven to be a constant in numerous studies. However, due to the enormous amount of data that so far had to be processed by hand, its crucial role to the formulaic recapitulation of art-historical motifs since antiquity could only be highlighted selectively. This is true even for the now automated estimation of human poses, as domain-specific, sufficiently large data sets required for training computational models are either not publicly available or not indexed at a fine enough granularity. With the <em>Poses of People in Art</em> data set, we introduce the first openly licensed data set for estimating human poses in art and validating human pose estimators. It consists of 2,454 images from 22 art-historical depiction styles, including those that have increasingly turned away from lifelike representations of the body since the 19<sup>th</sup> century. A total of 10,749 human figures are precisely enclosed by rectangular bounding boxes, with a maximum of four per image labeled by up to 17 keypoints; among these are mainly joints such as elbows and knees. For machine learning purposes, the data set is divided into three subsets—training, validation, and testing—, that follow the established JSON-based Microsoft COCO format, respectively. Each image annotation, in addition to mandatory fields, provides metadata from the art-historical online encyclopedia WikiArt.</p>
Novel digital assessment technique used to describe peanut nodulation life history
<p>Biological nitrogen fixation (BNF) in legume root nodules is an important source of N globally. Typically, the quantification of BNF is an estimate of N fixation through techniques such as an acetylene reduction assay or isotopic signatures. However, these methods have large logistical limitations, and do not allow for assessments of the nodule morphology or the nodulation process itself, knowledge that is important to fully understand the BNF process. To provide an updated method of nodulation assessments, peanut (<i>Arachis hypogaea </i>L.) nodules were sampled from a two-year field trial and quantified using scanned images of excised nodules, allowing digital assessment of nodule number and size. Further, a digital color analysis method was developed to assess internal nodule color (a known proxy for nodule BNF activity and senescence). To apply and validate the new method, peanut <span>nodulation</span> trait data was used for an exploration of the relationship between nodulation traits, plant physiological processes, and nitrogen status in the plant, followed by a nodulation life history analysis. The nodulation assessment method was validated by strong correlation with measures of peanut tissue N, biomass, photosynthetic function, and yield. Nodulation life history data suggest novel findings regarding differences in the development of nodules between taproots and lateral roots, contrary to the previous understanding of peanut nodulation life history. This new nodulation assessment method can provide further insight into the BNF process by allowing the collection of detailed data in a relatively simplified manner in field conditions. This dataset represents the nodulation and other variable data collected for this study.</p>
Barbara Thiers of the New York Botanical Garden and president of the Society for the Preservation of Natural History Collections is among those leading the effort to harness the explosion of data being digitized by collections around the globe. Photograph: New York Botanical Garden, Bronx, NY. in The Evolution of Natural History Collections
Barbara Thiers of the New York Botanical Garden and president of the Society for the Preservation of Natural History Collections is among those leading the effort to harness the explosion of data being digitized by collections around the globe. Photograph: New York Botanical Garden, Bronx, NY.
Data Table - Digital Scholarship, PhD project "Bridging Data Science and Intellectual History: Computing the Nodes and Edges in the Old University of Louvain (1425-1797)"
<p>How were academic networks configured in the premodern world, and how did they change? This doctoral project seeks to offer a data-driven answer by computing networks at and around the Old University of Louvain (1425-1797), a crucial hub for the transfer of knowledge in late medieval and early modern Europe. Drawing upon datasets under construction from the teams of the PIs at KU Leuven and UCLouvain, this project sets out to plot and visualize networks of students, scholars and their ‘books’ over almost four centuries, thus integrating data from demographic, prosopographical and book historical datasets. This will lead to a better understanding of how academic communities evolved in the past, and it will help to assess how their organization and structure promoted or hindered the creation and transfer of knowledge in premodern Europe. Hence, the doctoral research project offers an innovative test case to develop novel understandings of networks (e.g.. new tested ways to define nodes and edges) in datasets on scholars and ‘literati’, and it will be able to compare these new results to interpretations of human capital indices in the past. As such, the project creates a pioneering pilot for integrating data science into the field of intellectual and early modern history. This enhanced collaboration between KU Leuven and UCLouvain on a theme related to their common past is timely in view of 600 years Leuven/Louvain in 2025.</p>
Natural History Study of COVID-19 Using Digital Wearables
ClinicalTrials.gov study NCT04927442. IPD Sharing: Not stated. Countries: 1. Publications: 4.
Novel digital assessment technique used to describe peanut nodulation life history
Open the record for dataset details and reuse information.
Figure 1 from: Sikes DS, Copas K, Hirsch T, Longino JT, Schigel D (2016) On natural history collections, digitized and not: a response to Ferro and Flick. ZooKeys 618: 145-158. https://doi.org/10.3897/zookeys.618.9986
Figure 1 - Number of insect records in GBIF.org (triangles) between December 2007 and March 2016, in comparison to all records (circles).
Figure 3 from: Blagoderov V, Kitching I, Livermore L, Simonsen T, Smith V (2012) No specimen left behind: industrial scale digitization of natural history collections. ZooKeys 209: 133-146. https://doi.org/10.3897/zookeys.209.3178
Figure 3 - Metadata Creator software: a–c working areas a drawer image b specimen records c annotation fields d tool selector e unique IDs.
Figure 1 from: Blagoderov V, Kitching I, Livermore L, Simonsen T, Smith V (2012) No specimen left behind: industrial scale digitization of natural history collections. ZooKeys 209: 133-146. https://doi.org/10.3897/zookeys.209.3178
Figure 1 - SatScan imaging: a SatScan machine b specimens being imaged c individual frames aligned d fragment of a stitched image; final resolution of the stitched image ~11 lines/mm.
Figure 2 from: Blagoderov V, Kitching I, Livermore L, Simonsen T, Smith V (2012) No specimen left behind: industrial scale digitization of natural history collections. ZooKeys 209: 133-146. https://doi.org/10.3897/zookeys.209.3178
Figure 2 - Image based digitization workflow consisting of four stages: Imaging, Metadata capture, Institutional databading and Publication.
Wednesday 6 May 2020: Digital Methods for Web History (Anne Helmond, University of Amsterdam)
<p>Wednesday 6 May 2020: Digital Methods for Web History (Anne Helmond, University of Amsterdam)</p>
Using Anonymous Data from a Digital Tool for Medical History-Taking to Improve Healthcare Services
ClinicalTrials.gov study NCT06761105. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.
Natural History, Biogeography and Evolution of the Iberian white oak syngameon (Quercus L. Sect. Quercus) - Digital Annex
<p>Digital Annex for the following thesis:</p> <p>Vila-Viçosa, C. (2023). <em>Natural History, Biogeography and Evolution of the Iberian white oak syngameon </em>(Quercus <em>L. Sect.</em> Quercus). Ph.D. Thesis, Faculdade de Ciências da Universidade do Porto, Portugal</p> <p>Abstract:</p> <p>The genus <em>Quercus</em> L. is one of the most diverse and important group of woody plants, particularly when considering that they are the trees that rule the Northern Hemisphere forests. Oaks have an intricate Biogeography that criss-crosses diverse climatic and edaphic gradients, encompassing a huge ambiguity in terms of species delimitation. Frequently, the taxonomic proposals brought by traditional Linnaean Botany are either insufficient or rather inflate the number of species and nomenclatural assignments, which are further diluted into inconsistent taxonomic ranks, varying from species to subspecies and varieties. The supremacy given to morphological characters that are inherently fragile and plastic, spread across the distribution areas of distinct lineages, may carry ambiguity on the identification and proper species delimitation. From the oaks that are distributed across the Western Palearctic region, the ones that are deciduous or brevi-deciduous present higher levels of ambiguity in terms of species number and their delimitation. This ambiguity is particularly strong in the circummediterranean region and in the transitional areas between the two major biogeographic Regions of the western Palearctic region, the Euro-Siberian and Mediterranean. This degree of uncertainty, which increases towards the Southern European Peninsulas, is amplified by the ease that the different species of oaks tend to hybridize among them.</p> <p>The present work provides a holistic framework that covers multiple areas, from the taxonomic and evolutive study of this genus, to biogeography and molecular characterization. Its major objective was to resolve the species delimitation of the Iberian deciduous and marcescent oaks and putative introgression among them, enhancing the available knowledge about species diversity, which can foster suitable species and forest conservation. A specific objective was to cross-reference the natural history revision and the different taxonomic treatments brought by distinct authors, with personal observations. These data were then incorporated into ecological modelling and molecular characterization, which in the end fed a newly updated taxonomic proposal. In Section A we obtained results from extensive field, herbaria, and literature review, updating the nomenclature of the Portuguese and western Mediterranean oaks. Section B was supported by Section A’s in-depth review and enabled finer species distribution models, nurturing both hindcast (since <em>ca.</em> 20 Kyr) and forecast (2070-2100) exercises of the range dynamics of Mediterranean oaks species. The study of past and future range shifts solved important pending biogeographic questions, especially related to past range-shifts. Such past-range shifts improved our knowledge on species responses to climate dynamics and allowed a better anticipation of future responses of range shifts driven by climate change. Section C encompassed the molecular characterization of Iberian white oak species and their hybrids, whose delimitation is often faltering when one intends to infer about species rank, or hypothesize about the participation of parent taxon in natural hybrid swarms. This work allowed us to solve the phylogenetic backbone of western Palearctic white oaks, suggesting a significant segregation of the Iberian pedunculate oaks and unveiling two subsections inside Section <em>Quercus</em>. These subsections are biogeographically well-segregated and present diverse levels of introgression among species. Results demonstrated the efficiency of RADSeq for rebuilding the reticulate phylogeny of the Eurasian white oaks, showcasing the significance of the Iberian Peninsula as a major hotspot for oak diversity.</p> <p>We implemented a circular approach to these methods, which retro-fed themselves in terms of insight generation, enabling a powerful strategy to solve the evolutionary history of this difficult groups of plants. We estimate that the reticulate historical biogeography of the western Palearctic white oaks deserves further scrutiny by adding vicariant oak populations from northern Africa, the Near East and southern European Peninsulas. Methods should again follow this similar additive and sequential process of adjoining deep Natural History examination, with extensive fieldwork in type populations and genome-wide molecular surveys, in order to solve this group of plants. With the present work, we were able to significantly improve on the depiction of the basic unit of Biodiversity (the Species), in the complex <em>Quercus</em> genus. We provided tools to enable further efforts for the conservation of the Mediterranean oak forests, which overwhelm one of the most important (and one of the most threatened) Biomes for plant conservation at the global scale.</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.