Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

75

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

75 results for “Document Datasets”

Learn how ShareScore rates datasets ↗
zenodo32/100

Dataset for: Faculty Publication Trends in a Japanese National University: A Diachronic Document Analysis

<p>Dataset for paper titled:&nbsp;Faculty Publication Trends in a Japanese National University: A Diachronic Document Analysis</p>

opencc-by-4.0Jun 2023View details →
zenodo32/100

An Open Dataset of Scholarly Publications Referenced in Selected Policy Documents (POLIDOC_SCHOLAR)

<p>POLIDOC_SCHOLAR: &nbsp;An Open Dataset of Scholarly Publications Referenced in Selected Policy Documents</p> <p>This repository contains an open dataset of scholarly publications cited by selected policy documents.</p> <p><strong>1. Background:</strong></p> <ul> <li>We do not aim to create a dataset of references for all policy documents or millions of policy documents but rather from a carefully selected set of policy documents.</li> <li>The long-term plan is to facilitate the inclusion of citations of scholarly publications in open bibliometric databases (or at least to create inter-operable datasets).</li> <li>In the short-term, we plan to increase the number of policy documents included in the dataset and continue to monitor and increase the data quality (completeness of records, provided external identifiers).</li> <li>We will also document - in the next release - the reference extraction process (including code used)</li> </ul> <p>&nbsp;</p> <p><strong>2.&nbsp; Structure of the dataset:</strong></p> <p>The dataset is structured into two primary categories: &quot;<strong>Collections</strong>&quot; and &quot;<strong>Collection References</strong>.&quot;</p> <p><strong>Collections:</strong></p> <p>The metadata for selected policy documents is included the &quot;<em>collections.jsonl</em>&quot; file.</p> <p>The <strong>collection</strong> is a central feature of the POLIDOC_SCHOLAR dataset.&nbsp; The selected policy documents are listed in the &ldquo;<em>collections.jsonl</em>&rdquo;.</p> <p>For instance, a collection might include reports like the IPCC reports of the 6th Cycle (the &quot;IPCC_AR_6 collection&quot;) or the reports from IPBES (the &quot;IPBES collection&quot;).</p> <p>Within each collection, there are &quot;documents.&quot; These can be twofold:</p> <ul> <li>They represent individual reports within a collection (e.g., the IPCC_AR_6 collection contains 6 reports: 3 assessment reports and 3 special reports from the 6th Cycle of the IPCC assessment).</li> <li>They also denote specific sections of these reports that contain bibliographic references. These sections can be chapters or other segments like supplementary materials or annexes (any section which has a reference list). Each document has a unique code, and the relationships between a main document and its subdivisions are indicated in the &quot;is_part_of&quot; field.</li> </ul> <p><strong>Collection References:</strong></p> <p>To allow users to access only the collections they are interested in, we&#39;ve separated references by collection in files named &quot;<em>collection_reference_{&hellip;name of collection&hellip;}jsonl</em>.&quot;</p> <ul> <li>Each of these files includes bibliographic references for every document in a specific collection.</li> <li>Besides presenting these as &quot;reference strings&quot; (in their original format within the document), we also offer unique identifiers like DOI and OpenAlex ID to facilitate linkage to external databases.</li> </ul> <p>The documentation of the dataset is provided in the file &ldquo;<em>data_dictionary</em>&rdquo;</p> <p><strong>3, Content release v1:</strong></p> <p>This release (POLIDOC_SCHOLAR version 1) includes 2 collections:</p> <ol> <li>IPCC Assessment Cycle 6&nbsp;</li> <li>IPBES Assessment reports</li> </ol> <p>&nbsp;</p> <table> <tbody> <tr> <td> <p>&nbsp;</p> </td> <td> <p>collection</p> </td> <td> <p>Number of reports</p> </td> <td> <p>Number of documents (&ldquo;sections&rdquo; with reference)</p> </td> <td> <p>Number of references (strings, not unique)</p> </td> <td> <p>Number of references with DOI (unique)</p> </td> <td> <p>Number of references with DOI (unique)</p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p>IPCC Assessment Cycle 6</p> </td> <td> <p>6</p> </td> <td> <p>103</p> </td> <td> <p>94,958&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;</p> </td> <td> <p>51,713</p> </td> <td> <p>48,695</p> <p>&nbsp;</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>IPBES Assessment reports</p> </td> <td> <p>3</p> </td> <td> <p>27</p> </td> <td> <p>21,750&nbsp;&nbsp;&nbsp;&nbsp;</p> </td> <td> <p>12,100&nbsp;&nbsp;</p> </td> <td> <p>11,896</p> </td> </tr> </tbody> </table>

opencc-by-4.0Jul 2023View details →
zenodo28/100

Dataset for "Information Correspondence between Types of Documentation for APIs"

<p>This online appendix contains the coding guide and the data used in the paper&nbsp;<em>Information Correspondence between Types of Documentation for APIs</em>&nbsp;accepted for publication in the&nbsp;<em>Empirical Software Engineering</em>&nbsp;(EMSE) journal. The tutorial data was retrieved in October 2018.</p> <p>It contains the following files:</p> <p>1.&nbsp;<strong>CodingGuide.pdf</strong>: the coding guide to classify a sentence as&nbsp;<em>API Information</em>&nbsp;or&nbsp;<em>Supporting Text</em>.</p> <p>2.&nbsp;<strong>annotated_sampled_sentences.csv</strong>: the set of 332 sampled sentences and two columns of corresponding annotations &ndash; one by the first author of this work and the second by an external annotator. This data was used to calculate the agreement score reported in the paper.</p> <p>3.&nbsp;<strong>&lt;language&gt;-&lt;topic&gt;.csv</strong>: the data set of annotated sentences in the tutorial on &lt;topic&gt; in &lt;language&gt;. For example Python-REGEX.csv&nbsp;is the file containing sentences from the Python tutorial on regular expressions. This file contains the preprocessed sentences from the tutorial, their source files, and their annotation of sentence correspondence with reference documentation.</p> <p>For licensing reasons, we are unable to upload the original API reference documentation and tutorials, however these are available on request.</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

ICDAR 2021 Historical Document Classification Test Dataset for Task 1 - Scripts

<p>Test set for thescript classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The csv file indicates for each image from which document and page it corresponds to, as well as whether augmentations have been applied.</p>

opencc-by-4.0May 2021View details →
zenodo28/100

Figure 5 from: Bloom D, Thomer A, Vaidya G, Guralnick R, Russell L (2012) From documents to datasets: A MediaWiki-based method of annotating and extracting species observations in century-old field notebooks. ZooKeys 209: 235-253. https://doi.org/10.3897/zookeys.209.3247

Figure 5 - An example of how a location (Big Thompson Creek near Loveland), a date (Sunday, June 10, 1906), and a taxon (Cottonwood, genus Populus) are grouped from across multiple pages.

opencc-by-4.0Jul 2012View details →
zenodo28/100

Figure 4 from: Bloom D, Thomer A, Vaidya G, Guralnick R, Russell L (2012) From documents to datasets: A MediaWiki-based method of annotating and extracting species observations in century-old field notebooks. ZooKeys 209: 235-253. https://doi.org/10.3897/zookeys.209.3247

Figure 4 - Editing a notebook page on Wikisource. This screenshot shows side-by-side transcription and wiki markup syntax.

opencc-by-4.0Jul 2012View details →
zenodo28/100

Figure 2 from: Bloom D, Thomer A, Vaidya G, Guralnick R, Russell L (2012) From documents to datasets: A MediaWiki-based method of annotating and extracting species observations in century-old field notebooks. ZooKeys 209: 235-253. https://doi.org/10.3897/zookeys.209.3247

Figure 2 - Index page for Notebook #1. Each Index page corresponds to a multipage file. The Index page displays volume metadata and links to sections of the notebook, while also providing links out to each notebook page and color-coding to determine which pages have been already transcribed and proofed.

opencc-by-4.0Jul 2012View details →
zenodo28/100

Figure 1 from: Bloom D, Thomer A, Vaidya G, Guralnick R, Russell L (2012) From documents to datasets: A MediaWiki-based method of annotating and extracting species observations in century-old field notebooks. ZooKeys 209: 235-253. https://doi.org/10.3897/zookeys.209.3247

Figure 1 - Web browser view of a scanned page of Henderson's journal displayed side-by-side with transcriptions and annotations using the MediaWiki Proofread Page extension.

opencc-by-4.0Jul 2012View details →
zenodo28/100

Figure 3 from: Bloom D, Thomer A, Vaidya G, Guralnick R, Russell L (2012) From documents to datasets: A MediaWiki-based method of annotating and extracting species observations in century-old field notebooks. ZooKeys 209: 235-253. https://doi.org/10.3897/zookeys.209.3247

Figure 3 - Henderson's first sentence. "Boulder, Colo. July 28, 1905. Saw Say [sic] Phoebe and siskins, [American] Robins, [Northern] Flicker."

opencc-by-4.0Jul 2012View details →
zenodo28/100

Dataset to accompany Clustering and Visualising Documents using Word Embeddings

<p>Dataset to accompany <em>Clustering and Visualising Documents using Word Embeddings</em>, a lesson for the Programming Historian.</p>

opencc-by-4.0Jul 2023View details →
zenodo24/100

Document Visibilty Graph Threshold Estimation Dataset

<p>The aim of this dataset is to help with an estimation of thresholds used in geometrical algorithms for the creation of Visibility Graphs out of document content.</p> <p>For that purpose, the following thresholds are optimal values, leading to an maximal Area F1 for Table Region Detection tasks where those thresholds are the basis.</p> <p>Prediction target thresholds are:</p> <ul> <li>x_eps: alignment epsilon for vertical edges in points</li> <li>y_eps: alignment epsilon for horizontal edges in points</li> <li>page_ratio_x: maximal relative horizontal distance of two nodes where an edge can be created</li> <li>page_ratio_y: maximal relative vertical distance of two nodes where an edge can be created</li> <li>threshold_page_width: Indicating at maximal which width of a node the width should be added as an edge condition</li> <li>width_pct_eps: relative width difference of nodes as a condition for vertical edges</li> <li>font_eps: Font size difference between two nodes in points, acting again as an edge condition</li> </ul> <p>Independent variables here are:</p> <ul> <li>font_size_entropy: Shannon entropy of font sizes in a document, related to if a comparision of font sizes would be meaninful or is frequently present</li> <li>font_name_entropy: Shannon entropy of font names in a document</li> <li>bold_pct: percentage of bold texts in a document</li> <li>italic_pct: percentage of italic texts in a document</li> <li>x_var: deviation of the coordinate-based horizontal differences between nodes</li> <li>y_var: deviation of the coordinate-based vertical differences between nodes</li> <li>avg_width: average width of textual elements</li> </ul> <p>The corresponding PDF documents used will be referred in upcoming versions.</p>

opencc-by-4.0Aug 2020View details →
zenodo24/100

ArchiWood: dataset of legacy documents about wood anatomical, morphological, and architectural traits for plant species in Madagascar

<p><strong>Abstract</strong></p> <p>Digitization of a part of CIRAD's' xylotheque, one of the most important collections of tropical timber in the world, now allows researchers, experts in the field, trainers and archaeologists to have a new source of homogenized information. This corpus constitutes a unique database on the botanical characteristics of many species of Madagascar and could feed the research in both ecological and systematic fields by the possibility to define traits and criteria usable for example in phylogeny or functional ecology.</p> <p>The island of Madagascar presents a remarkable floristic diversity, with a very high rate of endemism, making it a permanent laboratory for the study of the mechanisms governing evolution.</p> <p>The digitization project was conducted in close coordination with other projects currently underway involving the same partners. 995 species are concerned by the ArchiWood project: 3 planes of anatomical sections for each species, associated with a total of 100 field notebooks (architectural sketches) and 20,000 slides.</p> <p><strong>Contents</strong></p> <p>This dataset is composed of digitized legacy documents, all related to Madagascar biodiversity:</p> <ul> <li>More than 2,300 illustrations and notes taken from field notebooks</li> <li>More than 1,200 anatomical cuts</li> <li>More than 70 photographs of plants and landscapes</li> </ul> <p><strong>Methods</strong></p> <p><em>Digitization of anatomical cuts</em><br> For each species, there are 3 anatomical cuts associated with the 3 planes of symmetry of the material: radial, tangential and longitudinal. Each cut was scanned using an Olympus DP71 digital camera fitted on an Olympus BX60 microscope. The shooting software that drives the camera is Archimed from Microvision Instruments. The image format is 24 bits colored TIFF. Each image is 1600 x 1200 pixels for 3 different magnifications of the microscope: x40 (overview of the cut), x200 (main features anatomical from IAWA), x400 (highlighting fine lines such as punctuations).</p> <p><em>Digitization of sketches (from field notebooks) </em><br> The digitized information in the field notebooks is made up of architectural sketches (with descriptive notes). The slides include additional remarkable visual information of field sketches (flowers, fruits, overview, geographical implantation conditions…). The sketches were scanned using a resolution of 300 dpi, 24 bits colors in JPEG format.</p> <p><em>Digitization of photographs (24 x 36 mm)</em><br> The photographs were scanned at 600 dpi, 24 bits colors in JPEG format.</p> <p><strong>File management and data quality control</strong></p> <p>Controlling images and adding metadata operations were carried out by the researchers concerned in their field of activity (anatomy, botany, architecture). The homogenization and final harmonization of the metadata and data was carried out by the data manager.<br>        </p>

opencc-by-nc-nd-4.0Dec 2015View details →
zenodo24/100

ICDAR 2021 Historical Document Classification Dataset for Task 3 - Location

<p>Test set for the localization classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and a ground truth CSV file. The metadata csv file contains information, such as where the image comes from.</p>

opencc-by-4.0May 2021View details →
zenodo20/100

Dataset related to the article "Aortic Valve Sclerosis Adds to Prediction of Short-Term Mortality in Patients with Documented Coronary Atherosclerosis"

<p>This record contains raw data related to the article &quot;Aortic Valve Sclerosis Adds to Prediction of Short-Term Mortality in Patients with Documented Coronary Atherosclerosis&quot;.</p> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p>Aims: Aortic valve sclerosis (AVSc), a non-uniform thickening of leaflets with an unrestricted opening, is characterized by inflammation, lipoprotein deposition, and matrix degradation. In the general population, AVSc predicts long-term cardiovascular mortality (+50%) even after adjustment for vascular risk factors and clinical atherosclerosis. We have hypothesized that AVSc is a risk-multiplier able to predict even short-term mortality. To address this issue, we retrospectively analyzed 90-day mortality of all patients who underwent isolated coronary artery bypass grafting (CABG) at Centro Cardiologico Monzino over a ten-year period (2006&ndash;2016). Methods: We analyzed 2246 patients and 90-day all-cause mortality was 1.5% (31 deaths). We selected only patients deceased from cardiac causes (n&nbsp;= 29) and compared to alive patients (n&nbsp;= 2215). A cardiologist classified the aortic valve as no-AVSc (n&nbsp;= 1352) or AVSc (n&nbsp;= 892). Cox linear regression and integrated discrimination improvement (IDI) analyses were used to evaluate AVSc in predicting 90-day mortality. Results: AVSc 90-day survival (97.6%) was lower than in no-AVSc (99.4%;&nbsp;p&nbsp;&lt; 0.0001) with a hazard ratio (HR) of 4.0 (95%CI: 1.78, 9.05;&nbsp;p&nbsp;&lt; 0.0001). The HR for AVSc, adjusted for propensity score, was 2.7 (95%CI: 1.17, 6.23;&nbsp;p&nbsp;= 0.02) and IDI statistics confirmed that AVSc significantly adds (p&nbsp;&lt; 0.001) to the identification of high-risk patients than EuroSCORE II alone. Conclusion: Our data supports the hypothesis that a risk stratification strategy based on AVSc, added to ESII, may allow better recognition of patients at high-risk of short-term mortality after isolated surgical myocardial revascularization. Results from this study warrant further confirmation.</p>

restrictedMar 2020View details →
zenodo20/100

Dataset fot the study of Documentation Practices of Machine Learning Resources

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record