Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

322

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

322 results for “census”

Learn how ShareScore rates datasets ↗
edi44/100

Sapling Census:FAB 1: Forests and Biodiversity Experiment - High density diversity experiment

A forest biodiversity experiment (FAB) focused on trees of our region investigates the consequences of multiple dimensions of tree diversity for soil, food webs, plant communities and ecosystems. FAB is designed to unravel effects of three forms of biological diversity: species richness (SR), functional diversity (FD), and phylogenetic diversity (PD). We define FD as the representation of multiple traits of leaves, roots, seeds, and the whole organism that are correlated with species positions along gradients of resource supply, growth, and decomposition. PD is the representation of evolutionary lineages measured as the genetic distances between species. While PD and FD are often correlated, convergent evolution and adaptive differentiation can decouple them. When functional traits that drive specific ecosystem functions are not phylogenetically conserved, PD and FD may give contrasting predictions. SR, PD, and FD are not independent, and we posit that PD may help explain SR effects, and FD may help explain both PD and SR effects. Thus FAB is designed to examine the separate and combined effects of all three components of diversity for multiple ecosystem functions and to distinguish between ???sampling??? and ???complementarity??? effects of biodiversity. Due to the long lag between planting tree seedlings and determining effects of tree composition and diversity on ecosystem functioning, fewer experiments have been established to elucidate the role of biodiversity in the functioning of forest ecosystems than grassland experiments. FAB will contribute to this gap and is a member of the IDENT and TreeDiv network of forest biodiversity experiments (www.treedivnet.ugent.be). Hypotheses: 1. PD, FD, and SR will all contribute to increased productivity, stability, and diversity of other trophic levels (herbivores, predators, parasitoids, soil microbes, soil flora and fauna) as well as to greater soil C sequestration. 2. Because PD incorporates both the number of species a

openCC0Mar 2024View details →
edi44/100

Sapling Census:FAB 2 -DIV: Forests and Biodiversity Experiment - Low density diversity experiment

The long-term forest biodiversity experiment (FAB2), is designed to expand the capacity and scope of FAB1, established in 2012. Like its predecessor, FAB2 focuses on trees of our region, tested for survival and growth at Cedar Creek. The experiment provides the ability to examine the same diversity of plant, soil, decomposer, food web, and ecosystem responses that have been studied in the grassland biodiversity experiments but also allows us to pose novel questions regarding the effects of functional and phylogenetic diversity on ecosystem processes and the long-term consequences of species-specific effects on ecosystem properties.

openCC0Mar 2024View details →
zenodo40/100

Refined personal name data from the census book of Vodskaja pjatina

<p>The data contains approximately 36,000 personal names derived from medieval Russian documentation. More preciously,&nbsp;names are collected from an edited version of the census book of Vodskaja pjatina, which was one of the five administrative areas in the late 15<sup>th</sup> century Novgorod.</p> <p>Editions were compiled in parts and the first two, which cover the northernmost region, are called <em>Переписная окладная книга по новугороду вотской пятины</em>&nbsp;(1851, 1852)(POKV I‒II). The third part of the book series <em>Новгородские пистсовые книги</em>&nbsp;(1868)(NPK III) covers the southern and western parts of the study area.</p> <p>The process of obtaining the personal from the inscription has been following: First, editions of the census book were obtained as scanned PDF files. These were transformed as editable copies by using OCR (=Optical Character Recognition) software Abbyy. The program read the original mid-19<sup>th</sup> century Russian text adequately with its old Russian alphabet package.</p> <p>After the initial corrections, a Python script was written to harvest the personal names. This was based on exploiting the systematic formalities in how most of the names were presented in the census book. The script looked for abbreviations &ldquo;дв.&rdquo; and &ldquo;д.&rdquo; and extracted all following capitalized words until section end markers &ldquo;.&rdquo;, &ldquo;;&rdquo; or &ldquo;:&rdquo;. As an output, a name to pogost matrix was produced, which held the raw frequencies of each word in each pogost.</p> <p>The process of cleaning the name data, in turn, has been done mostly by data wrangling program OpenRefine in following manner: For starters, all name forms shorter than four characters were removed as there were no personal names consisting of three or less letters. Furthermore, nouns that were not names were removed. This meant discarding expressions that described person&rsquo;s special feature or profession, like such as being a widow (&ldquo;вдова&rdquo;) or working as a deacon (&ldquo;діакъ&rdquo;). For some reason, editors followed inconsistent conventions in capitalizing these non-name nouns.</p> <p>In addition, some orthographical and morphological harmonization was done on the data. The letter <em>ы </em>was cut from the end of bynames, where it denotes plurality. Similarity of so called soft and hard signs, <em>ь </em>and <em>ъ</em> caused some problems. As the latter one is not used in contemporary Russian and was not used in the original documents either (Неволин 1853 : 4 (in Appendix 1)) it was removed. The soft sign <em>ь </em>was also removed because it was absent in the original documents and it had been used inconsistently by the editors. The letter <em>ѣ</em> (yat) is rarely used in personal names but nevertheless, it was changed to <em>е </em>(like as it is in contemporary Russian) as since it was often confused with soft and hard signs (<em>ь </em>and <em>ъ</em>). Furthermore, the letter <em>ѳ </em>(fita) was often erroneously recognized as <em>о </em>or <em>е. </em>As it is only found in NPK III and only in the beginning of certain names, which all are also written with &ldquo;Ф&rdquo; (e.g. &ldquo;Ѳедко&rdquo; vs. &ldquo;Федко&rdquo;), it was replaced with <em>Ф</em>.</p> <p>In the second phase most of the erroneous orthographies were corrected. We do not detail herescribe all the OCR-errors here that were found, but in the following a short description is given of the most significant corrections. There were, for example, many letters whose similarity caused problems for the OCR-program (e.g.&nbsp; <em>и </em>/ <em>й </em>and <em>б </em>/ <em>в</em>). In these cases, the correct orthography was sought in the census book editions and accordingly, Openrefine was used to change erroneous forms to right correct ones.</p> <p>After the corrections were made, the number of name types (= name variants) was reduced from 4942 to 2748. The Overall overall number of name tokens was dropped as well: from 36,405 to 35,726. Of the name types, more than half (1484) have only one occurrence.</p> <p>The refined and harmonized data is published as pogost-by-name frequency tabulations (<em>pogost,</em> equivalent of English <em>parish</em>). The file is in tab-delimited file (.tsv) format.</p> <p>References:</p> <p>Неволин, К. А. 1853, О пятинах и погостах новгородских в XVI веке, с приложением карты,&nbsp; Санкт-Петербург (Из Записок Императорского русского географического общества, Кн. VIII).</p> <p>NPK III = Новгородские писцовые книги, Т. 3 : Переписная оброчная книга Вотской пятины, 1500 года, 1868, 1868, Санкт Петербург.</p> <p>POKV I, II = Переписная окладная книга по Новугороду Вотьской пятины, 1851, 1852, Имп. Моск. о-во истории и древностей рос., Москва.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Urban Redevelopment by Census Tract in New York City (2000 -2020)

<p>The annual urban redevelopment map for NYC was produced using the classification method proposed in this experiment to highlight the spatial and temporal distribution of urban reconstructions. The time-series gentrification risk maps illustrate areas that have faced gentrification risk since 2000. The data was aggregated to the Census tract level for displaying a visually friendly result. The raw building-level data is also provided.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Australian Statistical-Area (SA) Level Regions and Census Income Data (2011)

<p>The Australian Statistical Geography Standard (ASGS) defines a series of nested geographical areas in Australia known as Statistical Area (SA) Levels. SA3 regions are aggregations of SA2 regions, and SA2 regions are aggregations of SA1 regions. This data set contains the shapefiles of all SA1, SA2, and SA3 regions across Australia at the time of the 2011 census, originally downloaded from the Australian Bureau of Statistics (<a href="https://www.abs.gov.au/AUSSTATS/abs@.nsf/DetailsPage/1270.0.55.001July\%20201.">ABS</a>).</p><p>This data set also contains income information from the 2011 census, at the SA1 and SA2 level in New South Wales (NSW). Specifically, it contains the number of families of various types within a range of weekly income brackets.</p><p>Sainsbury-Dale et al. (2023) used a subset of this data set in a study on poverty levels in an area of &nbsp;(NSW) surrounding Sydney. &nbsp;</p><p>&nbsp;</p><p><strong>References</strong></p><p>Sainsbury-Dale, M., Zammit-Mangion, A., and Cressie, N. (2023) "Modelling Big, Heterogeneous, Non-Gaussian Spatial and Spatio-Temporal Data using FRK", <i>Journal of Statistical Software</i>, to appear.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

F I G U R E 1 in Targeted census of lionfishes (Scorpaenidae) reveals high densities in their native range

F I G U R E 1 Lionfishes observed during targeted surveys at three sites at Moorea, French Polynesia. A total of 90 transects (total area surveyed = 3600 m2) were conducted across three depths ["Reef Crest" (c. 1–3 m depth); "Mid" (3 m below the reef crest, c. 4–6 m depth); and "Deep" (6 m below the reef crest, c. 7–9 m depth)] at the three locations (a). Three species of lionfish were observed: clearfin lionfish Pterois radiata (b, e); spotfin lionfish Pterois antennata (c, f); and twinspot lionfish Dendrochirus biocellatus (d, g). Density is averaged (mean ± se) across the three locations. Site 1, Site 2, Site 3

opencc-by-4.0Dec 2022View details →
zenodo40/100

DAO Census Dataset

<p>A collection of DAOs, proposals and votes from aragon, daohaus, daostack, governor, realms and snapshot</p>

opencc-zeroAug 2023View details →
zenodo40/100

Vascular epiphyte census data at the San Lorenzo Crane plot

<p>Long-term community data of vascular epiphytes on&nbsp;a permanent vegetation study&nbsp;at the San Lorenzo Crane plot. The epiphyte community at the San Lorenzo crane plot&nbsp;on 207 tree individuals was first recorded between 1998 to 2000 and again between 2010 to 2012. Two files are available, one per each census.<br> <br> Headers in each file correspond to:</p> <p><strong>- empty header:</strong> row ID<br> <strong>- treenum:&nbsp;</strong>code of each tree, data taken from the CTFS-Tree Censuses and Inventories in Panama (STRI-CTFS) https://stricollections.org/portal/collections/misc/collprofiles.php?collid=27<br> <strong>- spp_code:</strong> epiphyte species code<br> <strong>- tsppc:</strong> host tree species code<br> <strong>- freq:</strong> number of individuals of epiphyte species</p> <p><br> <br> &nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

POPP Datasets : Datasets for handwriting recognition from French population census

<p><strong>POPP datasets</strong></p> <p>This repository contains 3 datasets created within the POPP project (<a href="https://popp.hypotheses.org/#ancre2">Project for the Oceration of the Paris Population Census</a>) for the task of handwriting text recognition. These datasets have been published in <a href="https://hal.science/hal-03675614/"><em>Recognition and information extraction in historical handwritten tables: toward understanding early 20th century Paris census</em> at DAS 2022.</a></p> <p>The 3 datasets are called &ldquo;Generic dataset&rdquo;, &ldquo;Belleville&rdquo;, and &ldquo;Chauss&eacute;e d&rsquo;Antin&rdquo; and contains lines made from the extracted rows of census tables from 1926. Each table in the Paris census contains 30 rows, thus each page in these datasets corresponds to 30 lines.</p> <p>The structure of each dataset is the following:</p> <ul> <li>double-pages : images of the double pages</li> <li>pages: <ul> <li>images: images of the pages</li> <li>xml: METS and ALTO files of each page containing the coordinates of the bounding boxes of each line</li> </ul> </li> <li>lines: contains the labels in the file <code>labels.json</code> and the line images splitted into the folders <em>train</em>, <em>valid</em> and <em>test</em>. The double pages were scanned at a resolution of 200dpi and saved as PNG images with 256 gray levels. The line and page images are shared in the TIFF format, also with 256 gray levels.</li> </ul> <p>Since the lines are extracted from table rows, we defined 4 special characters to describe the structure of the text:</p> <ul> <li>&curren; : indicates an empty cell</li> <li>/ : indicates the separation into columns</li> <li>? : indicates that the content of the cell following this symbol is written above the regular baseline</li> <li>! : indicates that the content of the cell following this symbol is written below the regular baseline</li> </ul> <p>We provide a script <code>format_dataset.py</code> to define which special character you want to use in the ground-truth.</p> <p>The split for the <em>Generic Dataset</em> and <em>Belleville</em> have been made at the double-page level so that each writer only appears in one subset among train, evaluation and test. The following table summarizes the splits and the number of writers for each dataset:</p> <table> <thead> <tr> <th>Dataset</th> <th>train - # of lines</th> <th>validation - # of lines</th> <th>test - # of lines</th> <th># of writers</th> </tr> </thead> <tbody> <tr> <td>Generic</td> <td>3840 (128 pages)</td> <td>480 (16 pages)</td> <td>480 (16 pages)</td> <td>80</td> </tr> <tr> <td>Belleville</td> <td>1140 (38 pages)</td> <td>150 (5 pages)</td> <td>180 (6 pages)</td> <td>1</td> </tr> <tr> <td>Chauss&eacute;e d&rsquo;Antin</td> <td>625</td> <td>78</td> <td>77</td> <td>10</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Generic dataset (or POPP dataset)</strong></p> <ul> <li>This dataset is made 4800 annotated lines extracted from 80 double pages of the 1926 Paris census.</li> <li>There is one double page for each of the 80 districts of Paris</li> <li>There is one writer per double page so the dataset contains 80 different writers.</li> </ul> <p>&nbsp;</p> <p><strong>Belleville dataset</strong></p> <p>This dataset is a mono-writer dataset made of 1470 lines (49 pages) from the <em>Belleville</em> district census of 1926.</p> <p>&nbsp;</p> <p><strong>Chauss&eacute;e d&rsquo;Antin dataset</strong></p> <p>This dataset is a multi-writer dataset made of 780 lines (26 pages) from the <em>Chauss&eacute;e d&rsquo;Antin</em> district census of 1926 and written by 10 different writers.</p> <p>&nbsp;</p> <p><strong>Error reporting</strong></p> <p>It is possible that errors persist in the ground truth, so any suggestions for correction are welcome. To do so, please make a merge request on the <a href="https://github.com/Shulk97/POPP-datasets">Github repository</a> and include the correction in both the labels.json file and in the XML file concerned.</p> <p>&nbsp;</p> <p><strong>Citation Request</strong></p> <p>If you publish material based on this database, we request you to include a reference to paper <a href="http://link.springer.com/chapter/10.1007/978-3-031-06555-2_10"><code>T. Constum, N. Kempf, T. Paquet, P. Tranouez, C. Chatelain, S. Br&eacute;e, and F. Merveille,Recognition and information extraction in historical handwritten tables: toward understanding early 20th century Paris census ,Document Analysis Systems (DAS), pp. 143- 157, La Rochelle, 2022.</code></a></p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Philippine Librarians Census

<p>The University of the Philippines School of Library and Information Studies (UPSLIS)&nbsp;conducted a study of Philippine Librarians from November 2018 to October 2019. This endeavor was the first-ever wide-scale study undertaken to collect the occupational profile of Librarians in the Philippines. The study was motivated by the fact that we do not have any baseline data about Philippine Librarians. While we may assume that more women practice librarianship in the Philippines than males, we do not have empirical data to support such a claim. UPSLIS hopes that this report, together with the raw dataset and manual on how to process and analyze the dataset, will:</p> <p>-&nbsp;&nbsp; &nbsp;Gain a deeper insight into the current state of librarians in the Philippines<br> -&nbsp;&nbsp; &nbsp;Enable other researchers to validate and extend this current project to create more follow-up and comprehensive studies<br> -&nbsp;&nbsp; &nbsp;Serve as a benchmarking tool for future LIS research<br> -&nbsp;&nbsp; &nbsp;Create opportunities for collaboration with other researchers not only in the LIS field but with other disciplines as well<br> -&nbsp;&nbsp; &nbsp;Assist policymakers and decision-makers in drafting and creating policies and guidelines that would affect the LIS profession in the Philippines.</p> <p>UPSLIS also endeavors to conduct another census in five to six years to provide a complete longitudinal picture of the changes and development in the profession.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

NAPS Continuous Data: across 52 Census Divisions from 1980-2019

<p>The National Air Pollution Surveillance (NAPS) Program is a national system of air pollution&nbsp;sampling sites managed by &nbsp;Environment and Climate Change Canada, with cooperation from Provincial, Territorial, and regional partners who manage the sites and report data.</p> <p>This dataset contains&nbsp;aggregated and imputed continuous monitoring pollutant concentrations from 52 Canadian census divisions spanning 1980-2019.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Figure S2 in Assessing structure and seasonal variations of a temperate shallow water fish assemblage through Snorkel Visual Census

Figure S2. – Diel variation in observation frequency of individual species. Gobiusculus flavescens and S. melops were both more frequent at daytime, while A. anguilla, M. scorpius, S. trutta, G. morhua, C. harengus and T.bubalis were, all more frequently encountered at night.

opencc-by-4.0Dec 2019View details →
zenodo40/100

Figure S1 in Assessing structure and seasonal variations of a temperate shallow water fish assemblage through Snorkel Visual Census

Figure S1. – Diel variations in assemblage structure was significant (X2 P &lt;0.05). Demersal fishes are most abundant at day, while benthic fishes are dominant at night. Abundance of pelagic fishes increase at night.

opencc-by-4.0Dec 2019View details →
zenodo40/100

Figure 2. – Monthly average species richness observed during diurnal counts from June 2013 in Assessing structure and seasonal variations of a temperate shallow water fish assemblage through Snorkel Visual Census

Figure 2. – Monthly average species richness observed during diurnal counts from June 2013 to August 2014. Error bars represent ±SD. Number of counts per month are, indicated at column bases.

opencc-by-4.0Dec 2019View details →
zenodo40/100

Figure 3 in Assessing structure and seasonal variations of a temperate shallow water fish assemblage through Snorkel Visual Census

Figure 3. – Correlation between temperature and species richness compared between diurnal and nocturnal counts. Dotted trendlines show the quadratic and linear relationships between species richness and temperature at diurnal and nocturnal counts respectively.

opencc-by-4.0Dec 2019View details →
zenodo40/100

Census of Artists and Artisans of the Municipality of Pasto (Colombia) 2018

<p>The present work exposes the results of the census of artists and artisans of the Municipality of Pasto, which had as objective to know the social conditions of this population. The document is structured in four fundamental parts: the socodemographic information, the conditions of life, the artistic activity and the culture of artists and craftsmen. The study is descriptive and determined as observational characteristics of artists and craftsmen, producers, cultivators and interpreters residing in the Municipality of Pasto.</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

Fig. 2 in Assessment Of Census Techniques For Estimating Density And Biomass Of Gibbons (Primates: Hylobatidae)

Fig. 2. Cumulative number of groups of Müller's gibbon, Hylobates muelleri, as observed during range mapping in Kayan Mentarang National Park (KMNP) and Sungai Wain protection forest (SWPF), East Kalimantan, Indonesia.

opencc-by-4.0Dec 2005View details →
zenodo40/100

Fig. 1 in Assessment Of Census Techniques For Estimating Density And Biomass Of Gibbons (Primates: Hylobatidae)

Fig. 1. The island of Borneo showing the location of the two study areas. Grey shading indicates the area in which range mapping of all gibbon groups was executed, the straight lines indicate the transects, and the triangles indicate the listening positions from which the fixed point counts were made.

opencc-by-4.0Dec 2005View details →
zenodo40/100

Fig. 1 in Census of dinosaur skin reveals lithology may not be the most important factor in increased preservation of hadrosaurid skin

Fig. 1. Part (A) and counterpart (B) skin of hadrosaurid Kritosaurus sp. (YPM PU 016969) showing the typical dinosaurian morphology of non-imbricating, polygonal tubercles. Courtesy of the Peabody Museum of Natural History, Yale University, New Haven, USA.

opencc-by-4.0Nov 2012View details →
zenodo40/100

Linked collectors and determiners for: Census of species recorded during phytosociological surveys in Benin.

Natural history specimen data linked to collectors and determiners held within, "Census of species recorded during phytosociological surveys in Benin". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/c89babc3-4e67-4cdd-84e5-dd62f1193053">https://bionomia.net/dataset/c89babc3-4e67-4cdd-84e5-dd62f1193053</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/c89babc3-4e67-4cdd-84e5-dd62f1193053">https://gbif.org/dataset/c89babc3-4e67-4cdd-84e5-dd62f1193053</a>. Formatted as a Frictionless Data package.

opencc-zeroJan 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record