Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

48

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

48 results for “Information Extraction”

Learn how ShareScore rates datasets ↗
zenodo36/100

Extracted Information for the Systematic Review of Survey Scales for Measuring Information Privacy Concerns on Social Network Sites

<p>The data set is part of a systematic literature review of&nbsp;survey scales for measuring information privacy concerns (IPCs) used in research on social network sites (SNSs).</p> <p>The results of this systematic literature review are available in Bartol, J., Vehovar, V., &amp; Petrovčič, A. (2023).&nbsp;Systematic review of survey scales measuring information privacy concerns on social network sites.&nbsp;<em>Telematics and Informatics</em>,&nbsp;102063. https://doi.org/10.1016/j.tele.2023.102063</p> <p>The article also includes a detailed description of the methods used in generating this data set.</p>

opencc-by-nc-4.0May 2022View details →
ClinicalTrials.gov36/100

Improving Information Extraction From EEG on Cerebral Anesthetic Drug Effects

ClinicalTrials.gov study NCT02043938. IPD Sharing: NO. Countries: 1. Publications: 5.

closedIPD-NOFeb 2026View details →
zenodo32/100

Towards Olfactory Information Extraction from Text: A Case Study on Detecting Smell Experiences in Novels

<p>Dataset accompanying &quot;Ryan Brate, Paul Groth and Marieke van Erp (2020) Towards Olfactory Information Extraction from Text: A Case Study on Detecting Smell Experiences in Novels. LaTeCH-CLfL 2020. Barcelona, December 2020.&quot;</p> <p>Abstract:</p> <p>Environmental factors determine the smells we perceive, but societal factors factors shape the importance, sentiment and biases we give to them. Descriptions of smells in text, or as we call them `smell experiences&#39;, offer a window into these factors, but they must first be identified. To the best of our knowledge, no tool exists to extract references to smell experiences from text. In this paper, we present two variations on a semi-supervised approach to identify smell experiences in English literature. The combined set of patterns from both implementations offer significantly better performance than a keyword-based baseline.</p>

opencc-by-4.0Nov 2020View details →
zenodo32/100

SympTEMIST Corpus: Gold Standard annotations for clinical symptoms, signs and findings information extraction

<p><strong>SympTEMIST</strong> stands for Symptoms TExt MIning Shared Task. It is a shared task and set of resources focused on the <strong>detection of mentions, normalization and indexing of symptoms, signs and findings in medical documents</strong> in <strong>Spanish</strong>. SympTEMIST is complementary to the DisTEMIST (https://temu.bsc.es/distemist) and MedProcNER/ProcTEMIST (https://temu.bsc.es/medprocner) corpora as they all use the same document collection.</p> <p>&nbsp;</p> <p><strong>Please cite if you use this dataset:</strong></p> <p>"Lima-L&oacute;pez, S., Farr&eacute;-Maduell, E., Gasco-S&aacute;nchez, L., Rodr&iacute;guez-Miret, J. and Krallinger, M. (2023). Overview of SympTEMIST at BioCreative VIII: corpus, guidelines and evaluation of systems for the detection and normalization of symptoms, signs and findings from text. In: <em>Proceedings of the BioCreative VIII Challenge and Workshop: Curation and Evaluation in the era of Generative Models</em>."</p> <p>&nbsp;</p> <p>This repository includes the:</p> <ul> <li><strong>Train </strong>and<strong> Test Set </strong>for the three subtasks</li> <li><strong>SYMPTEMIST gazetteer </strong>of <strong>SNOMED symptoms, signs &amp; findings&nbsp;</strong></li> <li><strong>Multilingual </strong>Silver Standard<strong> </strong>in <strong>9 languages:</strong> <ul> <li><em><strong>English</strong></em></li> <li><em><strong>Portuguese</strong></em></li> <li><em><strong>French</strong></em></li> <li><em><strong>Italian</strong></em></li> <li><em><strong>Romanian</strong></em></li> <li><em><strong>Catalan</strong></em></li> <li><em><strong>Swedish</strong></em></li> <li><em><strong>Dutch</strong></em></li> <li><em><strong>Czech</strong></em></li> </ul> </li> <li><strong>Background set</strong> of over 15,000 clinical cases.</li> <li><strong>Spanish</strong> Silver Standard (predictions by participants over the background set)</li> </ul> <p>&nbsp;</p> <p>Please read the README file attached for more information on folder structure and file format.</p> <p>&nbsp;</p> <p>SympTEMIST was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of BioCREATIVE 2023. For more information on the corpus, annotation scheme and task in general, please visit: https://temu.bsc.es/symptemist.</p> <p>&nbsp;</p> <p><strong>Resources:</strong></p> <ul> <li><a href="https://temu.bsc.es/symptemist/"><strong>Task web</strong></a></li> <li><a href="https://biocreative.bioinformatics.udel.edu/"><strong>BioCreative web</strong></a></li> <li><strong>Citation:&nbsp;</strong>Lima-L&oacute;pez, S., Farr&eacute;-Maduell, E., Gasco-S&aacute;nchez, L., Rodr&iacute;guez-Miret, J. and Krallinger, M. (2023). Overview of SympTEMIST at BioCreative VIII: corpus, guidelines and evaluation of systems for the detection and normalization of symptoms, signs and findings from text. In: <em>Proceedings of the BioCreative VIII Challenge and Workshop: Curation and Evaluation in the era of Generative Models</em>.</li> <li><a href="../records/8246440"><strong>Annotation guidelines</strong></a></li> <li><a href="../records/10103191"><strong>BioCreative/AMIA workshop proceedings</strong></a></li> <li><a href="../doi/10.5281/zenodo.10104546"><strong>Overview paper</strong></a></li> <li><a href="https://www.youtube.com/playlist?list=PLyLDDulunoofv79Pci0-y2rGP5GOBoZhp"><strong>Youtube videos (overview &amp; teams)</strong></a></li> <li><a href="https://www.slideshare.net/MartinKrallinger/symptemist-shared-task-on-symptoms-signs-and-findings-detection-and-normalization-task-overview-talk-at-biocreative-viii-workshop-amia-2023"><strong>SympTEMIST overview talk slides at BioCreative/AMIA workshop</strong></a></li> </ul> <p>&nbsp;</p> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong>Contact</strong></p> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-L&oacute;pez (&lt;salvador [dot] limalopez [at] gmail [dot] com&gt;)<br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p> <p><strong>Additional resources and corpora</strong></p> <p>If you are interested in SympTEMIST, you might want to check out these corpora and resources:</p> <ul> <li><a href="../records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, same document collection)</li> <li><a href="../records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, same document collection)</li> <li><a href="../records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization, same document collection)</li> <li><a href="../records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection with some overlapping documents)</li> <li><a href="../records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection with some overlapping documents)</li> <li><a href="../records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), modified synthetic verions of the document collection)</li> <li><a href="../records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li> <li><a href="../records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, same document collection)</li> <li><a href="../records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy, different document collection with some overlapping documents)</li> <li><a href="../records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, same document collection)</li> <li><a href="../records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), same document collection)</li> <li><a href="../records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, same document collection)</li> <li><a href="../records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts, different document collection)</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo32/100

Enhancing georeferenced biodiversity inventories: automated information extraction from literature records reveal the gaps

<p>Data and code supplement to our article revised submission to PeerJ.</p> <p>&nbsp;</p> <p>The file is compressed using standard&nbsp;zip.&nbsp;The uncompressed size is about 50 GB. There is a readme.md in the archive, which explains the structure of the contents.</p> <p>&nbsp;</p> <p>Abstract:</p> <p>We use natural language processing (NLP) to retrieve location data for cheilostome bryozoan species (text-mined occurrences [TMO]) in an automated procedure. We compare these results with data combined from two major public databases (DB): the Ocean Biogeographic Information System (OBIS), and the Global Biodiversity Information Facility (GBIF). Using DB and TMO data separately and in combination, we present latitudinal species richness curves using standard estimators (Chao2 and the Jackknife) and range-through approaches. Our combined DB and TMO species richness curves quantitatively document a bimodal global latitudinal diversity gradient for extant cheilostomes for the first time, with peaks in the temperate zones. 79% of the georeferenced species we retrieved from TMO (N = 1408) and DB (N = 4549) are non-overlapping. Despite clear indications that global location data compiled for cheilostomes should be improved with concerted effort, our study supports the view that many marine latitudinal species richness patterns deviate from the canonical latitudinal diversity gradient (LDG). Moreover, combining online biodiversity databases with automated information retrieval from the published literature is a promising avenue for expanding taxon-location datasets.</p>

opencc-by-4.0Apr 2022View details →
zenodo32/100

ARCHI4MOM: Using Tracing Information to Extract the Architecture of Microservice-based Systems from Message-oriented Middleware

<p>The data published&nbsp;here are relevant for&nbsp;ESCA_2022 publication.&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo32/100

CompactIE: Compact Facts in Open Information Extraction [Dataset]

<p>Dataset for the paper <a href="https://arxiv.org/abs/2205.02880">&quot;CompactIE: Compact Facts in Open Information Extraction&quot;</a>.&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo32/100

Information Extraction System: a case study in information extraction from the Polish Fire Service rescue reports

<p>The mind maps of the car accident ontologies.</p>

opencc-by-4.0Sep 2017View details →
zenodo32/100

Datasets for "Reading Order Independent Metrics for Information Extraction in Handwritten Documents"

<p>This repository includes the five datasets used for our paper entitled <em>Reading Order Independent Metrics for Information Extraction in Handwritten Documents</em>, in which we compare various metrics to evaluate end-to-end information extraction from scanned documents.</p> <h2>Datasets</h2> <p>Five datasets are released following the BIO format:</p> <ul> <li>IAM</li> <li>Simara</li> <li>POPP</li> <li>Esposalles</li> <li>French Military Records</li> </ul> <p>For each dataset, we provide the following data (on test sets):</p> <ul> <li>Ground truth annotations (<code>gt/</code>)</li> <li>Automatic predictions (<code>dan/</code>)</li> <li>Automatic predictions with entities appearing in random order (<code>dan_shuffled/</code>)</li> </ul> <p>The data is organized as follows:</p> <p><code>├── Dataset name/</code><br><code>│ &nbsp; ├── gt/</code><br><code>│ &nbsp; ├── dan/</code><br><code>│ &nbsp; └── dan_shuffled/</code></p> <h2>Metrics</h2> <p>To install the <a href="https://pypi.org/project/ie-eval/"><code>ie-eval</code></a> package, run <code>pip install ie-eval</code>.</p> <p>To compute all metrics on a specific dataset, run:<br><br><code>ie-eval all --label-dir IAM_paragraph/gt/ --prediction-dir IAM_paragraph/dan/</code><br><br></p> <p>To learn more about the various options, use the <code>--help</code> argument or read the <a href="https://ie-eval-ner-metrics-050f40e80b04480e2310d39ad338de778f6bec80e18.pages.teklia.com/">documentation</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Assisted Data Annotation for Business Process Information Extraction from Textual Documents

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
dryad32/100

Data from: HyRAD-X, a versatile method combining exome capture and RAD sequencing to extract genomic information from ancient DNA

Over the last decade, protocols aimed at reproducibly sequencing reduced-genome subsets in non-model organisms have been widely developed. Their use is however limited to DNA of relatively high molecular weight. During the last year, several methods exploiting hybridization capture using probes based on RAD-sequencing loci have circumvented this limitation and opened avenues to the study of samples characterized by degraded DNA, such as historical specimens. Here, we present a major update to those methods, namely Hybridization capture from RAD-derived probes obtained from a reduced eXome template (hyRAD-X), a technique applying RAD-sequencing to messenger RNA from one or few fresh specimens to elaborate bench-top produced probes, i.e., a reduced representation of the exome, further used to capture homologous DNA from a samples set. In contrast to previous hybridization-capture methods, the reference catalog on which reads are aligned does not rely on de novo assembly of anonymous RAD-sequencing loci, but on an assembled transcriptome obtained from RNAseq data, thus increasing the accuracy of loci definition and Single-Nucleotide-Polmorphisms (SNP) call, and targeting, specifically, expressed genes. Finally, the capture step of hyRAD-X relies on RNA probes, increasing stringency of hybridization, making it well suited for low-content DNA samples. As a proof of concept, we applied hyRAD-X to subfossil needles from the coniferous tree Abies alba, collected in lake sediments (Origlio, Switzerland) and dating back from 7200-5800 years before present (BP). More specifically we investigated genetic variation before, during, and after an anthropogenic perturbation that caused an abrupt decrease in Abies alba population size, 6500-6200 years BP. HyRAD-X produced a matrix encompassing 524 exome-derived SNPs. Despite a lower observed heterozygosity was observed during the 6.500-6.200 years BP time slice, genetic composition was nearly identical before and after the perturbation, indicating that re-expansion of the population after the decline was driven by autochthonous specimens. To the best of our knowledge, this is the first time a population genomic study incorporating ancient DNA samples of tree subfossils is conducted at a moderate cost using reproducible exome-reduced complexity.

opencc-zeroDec 2016View details →
dryad32/100

Extracting physiological information in experimental biology via Eulerian video magnification

<p><b>Background:</b></p> <p>Videographic material of animals can contain inapparent signals, such as color changes or motion that hold information about physiological functions, such as heart and respiration rate, pulse wave velocity and vocalization. Eulerian video magnification allows enhancement of such signals to enable their detection. The purpose of this study is to demonstrate how signals relevant to experimental physiology can be extracted from non-contact videographic material of animals.</p> <p><b>Results: </b></p> <p>We applied Eulerian video magnification to detect physiological signals in a range of experimental models and in captive and free ranging wildlife. Neotenic Mexican axolotls were studied to demonstrate the extraction of heart rate signal of non-embryonic animals from dedicated videographic material. Heart rate could be acquired both in single and multiple animal setups of leucistic and normally colored animals under different physiological conditions (resting, exercised or anesthetized) using a wide range of video qualities. Pulse wave velocity could also be measured in the low blood pressure system of the axolotl as well as in the high pressure system of the human being. Heart rate extraction was also possible from videos of conscious, unconstrained zebrafish and from non-dedicated videographic material of sand lizard and giraffe. This technique also allowed for heart rate detection in embryonic chickens <i>in ovo</i> through the eggshell and in embryonic mice <i>in utero,</i> and could be used as a gating signal to acquire two-phase volumetric micro-CT data of the beating embryonic chicken heart. Additionally, Eulerian video magnification was used to demonstrate how vocalization induced vibrations can be detected in infrasound producing Asian elephants. </p> <p><b>Conclusions:</b></p> <p>Eulerian video magnification provides a technique to extract inapparent temporal signals from videographic material of animals. This can be applied in experimental and comparative physiology where contact based recordings (e.g. heart rate) cannot be acquired.</p>

opencc-zeroDec 2018View details →
ClinicalTrials.gov32/100

Should Preoperative Information Before Impacted Third Molar Extraction?

ClinicalTrials.gov study NCT05548790. IPD Sharing: Not stated. Countries: 1. Publications: 4.

restrictedIPD-UNDECIDEDFeb 2026View details →
dryad32/100

Data from: HyRAD-X, a versatile method combining exome capture and RAD sequencing to extract genomic information from ancient DNA

Open the record for dataset details and reuse information.

publicApr 2018View details →
dryad32/100

Extracting physiological information in experimental biology via Eulerian video magnification

Open the record for dataset details and reuse information.

publicOct 2019View details →
dryad28/100

Data from: Robust extraction of quantitative structural information from high-variance histological images of livers from necropsied Soay sheep

Quantitative information is essential to the empirical analysis of biological systems. In many such systems, spatial relations between anatomical structures is of interest, making imaging a valuable data acquisition tool. However, image data can be difficult to analyse quantitatively. Many image processing algorithms are highly sensitive to variations in the image, limiting their current application to fields where sample and image quality may be very high. Here, we develop robust image processing algorithms for extracting structural information from a dataset of high-variance histological images of inflamed liver tissue obtained during necropsies of wild Soay sheep. We demonstrate that features of the data can be measured in a fully automated manner, providing quantitative information which can be readily used in statistical analysis. We show that these methods provide measures that correlate well with a manual, expert operator-led analysis of the same images, that they provide advantages in terms of sampling a wider range of information and that information can be extracted far more quickly than in manual analysis.

opencc-zeroDec 2016View details →
zenodo28/100

MySQL dump for finding optimal parameters for data augmentation techniques in publication "Leveraging Data Augmentation for Process Information Extraction"

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

M-POPP datasets: Datasets for full page text recognition and information extraction from French handwritten and printed marriage records

<h1><strong>M-POPP datasets</strong></h1> <p>This repository contains 2 datasets created within the <strong>EXO-POPP project</strong> (<a href="https://exopopp.hypotheses.org/">Optical EXtraction of handwritten named entities for marriage records of the POPulation of Paris</a>) for the task of text recognition and information extraction. These datasets have been published in <a href="https://arxiv.org/abs/2404.19329"><code>End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940</code></a> <code>[1]</code>at ICDAR 2024.</p> <p><strong>This version contains the labels for Handwritten Text Recognition and Handwritten Text Recognition + Information Extraction as used in our new paper "<a href="https://hal.science/hal-04555188">DANIEL: A fast Document Attention Network for Information Extraction and Labelling of handwritten documents</a>" [3].</strong></p> <p><strong>This version makes corrections to the handwritten dataset. </strong>More precisely, it corrects a few errors in transcription annotations and named entities.</p> <p><strong>The printed dataset is unchanged compared to version 2.</strong></p> <p><strong>The performances of the models described in [1] and [3] are detailled in the Leaderboard section.</strong></p> <h2><strong>General information</strong></h2> <p>The <strong>EXO-POPP project</strong> aims to establish a comprehensive database comprising 300,000 marriage records from Paris and its suburbs, spanning the years 1880 to 1940, which are preserved in over 130,000 scans of double pages. Each marriage record may encompass up to 118 distinct types of information that require extraction from plain text. The M-POPP corpus (which stands for Marriage records of the POPulation of Paris) is the corpus on which the EXO-POPP project focuses. This corpus was built by gathering the marriage records of Paris and its suburb regions (Hauts- de-Seine, Seine-Saint-Denis, Val-de-Marne).</p> <p>The M-POPP corpus are a subset of the M-POPP database with annotations for full-page text recognition and named entity recognition/information extraction from both handwritten and printed documents. The first dataset comprises handwritten marriage records, while the second dataset consists of typewritten marriage records. It should be noted that even in typewritten marriage records, some handwritten information occurs, especially concerning the names of the spouses, and notes in the margin.<br>The dataset contains single-page images obtained from the original scans of double pages via page segmentation.</p> <p>The structure of the files is the following:</p> <ul> <li>handwritten:&nbsp;<em>the handwritten dataset</em><br> <ul> <li>images: <em>images of the dataset divided following the split used in [1]</em><br> <ul> <li>train</li> <li>valid</li> <li>test</li> </ul> </li> <li>labels:&nbsp;<em>labels for joint handwritten text recognition and information extraction for each encoding tested in [1]</em></li> </ul> </li> <li>printed: <em>the printed dataset</em><br> <ul> <li>images:&nbsp;<em>images of the dataset divided following the split used in [1]</em><br> <ul> <li>train</li> <li>valid</li> <li>test</li> </ul> </li> <li>labels:&nbsp;<em>labels for joint handwritten text recognition and information extraction for each encoding tested in [1]</em></li> </ul> </li> <li>encoding-2-to-encoding-5.json:&nbsp;<em>a JSON file giving the correspondence between the symbols of encoding 2 and encoding 5.</em></li> </ul> <p>&nbsp;</p> <p>Table 1: Details on the split of the handwritten dataset.</p> <table> <tbody> <tr> <td>&nbsp;</td> <td>Train</td> <td>Validation</td> <td>Test</td> </tr> <tr> <td>Pages</td> <td>250</td> <td>32</td> <td>32</td> </tr> <tr> <td>Acts</td> <td>344</td> <td>51</td> <td>53</td> </tr> <tr> <td>Named entities</td> <td>16727</td> <td>2223</td> <td>2517</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Table 2: Details on the split of the printed dataset.</p> <table> <tbody> <tr> <td>&nbsp;</td> <td>Train</td> <td>Validation</td> <td>Test</td> </tr> <tr> <td>Pages</td> <td>116</td> <td>14</td> <td>13</td> </tr> <tr> <td>Acts</td> <td>363</td> <td>43</td> <td>30</td> </tr> <tr> <td>Named entities</td> <td>22036</td> <td>2559</td> <td>2405</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Table 3: Average annotation statistics per act for the two M-POPP datasets.</p> <table> <tbody> <tr> <td>Dataset</td> <td># of characters</td> <td># of words</td> <td># of named entities</td> </tr> <tr> <td>Handwritten</td> <td>1519</td> <td>231</td> <td>48</td> </tr> <tr> <td>Printed</td> <td>1328</td> <td>200</td> <td>60</td> </tr> </tbody> </table> <p>&nbsp;</p> <h3><strong>Document structure Annotation</strong></h3> <p>We employ the procedure applied in [2], which involves adding opening and closing tags to the character set for each text block we want to recognize.<br>In total, we define four types of text blocks.</p> <ul> <li>Block A is located in the margin and contains the last names of the married couple, possibly with their first names and the date of the marriage.</li> <li>Block B is the body of the text. Block B is the one that contains most of the information to be extracted.</li> <li>Block C is optional and corresponds to marginal notes used in various cases, such as the mention of a divorce or a correction made to the act.</li> <li>Block D corresponds to a set containing a block A and a block B, optionally with one or more blocks C.</li> </ul> <p>&nbsp;</p> <h3><strong>Information Extraction annotation</strong></h3> <p>The dataset contains 118 information categories. As explained in the paper, we broke down the named entities into sub-elements pertaining to 4 hierarchical levels, which reduces the total number of categories to 23 instead of 118. Notice that level 1, 2, and 3 categories do not encode named entities but rather the relations that may occur between some lower level categories for example: (day, birth, husband) encodes the fact that the annotated piece of text is the date of birth of the husband.&nbsp;</p> <p>For these datasets, we chose to represent these hierarchical elements with emojis. For instance, the information <em>first name</em> is represented by the emoji 💬.<br>The meaning of each emoji can be found in Table 4. To determine the best way to encode named entities in the ground truth, we compared in [1] 5 types of encoding. To illustrate these encodings, let&rsquo;s take for instance&nbsp;<em>Louis Alexandre MOUDEL</em> that we define as the father of the bride, where <em>Louis Alexandre</em> are his two first names, and <em>Moudel</em> is his last name.&nbsp;</p> <p>1) Single separate tags before each word: In this approach, each level of information is indicated by a dedicated tag, and the tags are placed before the word they encode information for. With this encoding, the ground truth for the example would be:</p> <p>💬👴👰Louis &nbsp; 💬👴👰Alexandre&nbsp; 🗨️👴👰MOUDEL</p> <p>2) Single separate tags after each word: Similar to the previous approach, except here the tags are placed after the word. With this encoding the previous example becomes:</p> <p>Louis👰👴💬&nbsp; Alexandre👰👴💬&nbsp; MOUDEL👰👴🗨️</p> <p>3) Open &amp; close separate tags: Here, each word presenting information to be extracted is surrounded by one or more opening and closing tags, where each tag encodes a level of information. So the example would be as:</p> <p>&lt;👰&gt; &lt;👴&gt; &lt;💬&gt; Louis &lt;\💬&gt; &lt;\👴&gt; &lt;\👰&gt;<br>&lt;👰&gt; &lt;👴&gt; &lt;💬&gt; Alexandre &lt;\💬&gt; &lt;\👴&gt; &lt;\👰&gt;<br>&lt;👰&gt; &lt;👴&gt; &lt;🗨️&gt; MOUDEL &lt;\🗨️&gt; &lt;\👴&gt; &lt;\👰&gt;</p> <p>4) Nested open &amp; close separate tags: Similar to the previous approach, but this time a tag is closed only when the encoded information is no longer the same for that level of information. We can see in the example below that the tags for wife and father are only used twice.</p> <p>&lt;👰&gt; &lt;👴&gt; &lt;💬&gt; Louis Alexandre &lt;\💬&gt; &lt;🗨️&gt; MOUDEL &lt;\🗨️&gt;</p> <p>5) Single combined tags after each word: In the last approach, one tag encodes all the hierarchical levels constituting information. The tags are located after the word they encode information for.&nbsp;</p> <p>Louis&lt;wife_father_first_name&gt;&nbsp; Alexandre&lt;wife_father_first_name&gt;&nbsp; MOUDEL&lt;wife_father_family_name&gt;</p> <p>NB: In the labels file of encoding 5, the information are still encoded with emojis but the chosen emojis do not have a semantic meaning due to the number of information categories to be represented. The correspondence between the symbols of encoding 2 and encoding 5 can be found in the file&nbsp;<em>encoding-2-to-encoding-5.json</em>.<em><br></em></p> <p>&nbsp;</p> <p>Table 4: Details of the hierarchical breakdown of named entities. Each tag is placed in the corresponding hierarchical level and associated with the emoji representing it.</p> <table> <tbody> <tr> <td>Level</td> <td>Tags</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>1</td> <td>Administrative 📖</td> <td> <pre>Husband<code> 👨</code></pre> </td> <td>Wife 👰</td> <td>Witness 🥸</td> </tr> <tr> <td>2</td> <td>Father 👴</td> <td>Mother 👵</td> <td>Ex-husband 💔</td> <td>&nbsp;</td> </tr> <tr> <td>3</td> <td>Birth 🏥</td> <td>Residence 🏠</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <td>4</td> <td>First name 💬</td> <td>Family name 🗨️</td> <td>Age ⌛</td> <td>Occupation 🔧</td> </tr> <tr> <td>5</td> <td>Street number 🔟</td> <td>Street type 🛣</td> <td>Street name 🔠</td> <td>City 🌆</td> </tr> <tr> <td>&nbsp;</td> <td>Department 🗺</td> <td>Country 🗺</td> <td>Day 🌞</td> <td>Month 📅</td> </tr> <tr> <td>&nbsp;</td> <td>Year 🗓</td> <td>Hour ⏰</td> <td>Minute ⏱</td> <td>&nbsp;</td> </tr> </tbody> </table> <p>&nbsp;</p> <h2><strong>Leaderboard</strong></h2> <h3><strong>Results on M-POPP handwritten</strong></h3> <p><strong>HTR<br></strong></p> <p>The following table contains the current leaderboard of M-POPP v3 for HTR on the handwritten dataset.</p> <p>In this configuration, layout block C is not considered.</p> <p>These results for DAN NER are given using the named entity encoding format 5 described above.</p> <p>HTR stands for Handwritten Text Recognition and HTR+IE for combined Handwritten Text Recognition and Information Extraction.</p> <p>Metrics are expressed in percentages.</p> <table> <tbody> <tr> <td>Method</td> <td>CER</td> <td>WER</td> <td>LOER</td> <td>mAP CER</td> </tr> <tr> <td>DAN - HTR [1]</td> <td>7.21</td> <td>16.42</td> <td>5.35</td> <td>83.03</td> </tr> <tr> <td>DAN NER - HTR + IE [1]</td> <td>6.52</td> <td>14.80</td> <td>3.79</td> <td>86.29</td> </tr> <tr> <td>DANIEL - HTR [3]</td> <td>5.72</td> <td>14.08</td> <td>1.34</td> <td>89.28</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>NER</strong></p> <p>The following table contains the current leaderboard of M-POPP v3 for NER on the handwritten dataset.</p> <p>In this configuration, layout block C is not considered.</p> <p>These results are given using the named entity encoding format 5 described above.</p> <p>Metrics are expressed in percentages.</p> <table> <tbody> <tr> <td>Method</td> <td>F1</td> </tr> <tr> <td>DAN NER [1]</td> <td>76.37</td> </tr> <tr> <td>DANIEL [3]</td> <td>76.37</td> </tr> </tbody> </table> <h3>&nbsp;</h3> <h3><strong>Results on M-POPP printed</strong></h3> <p><strong>HTR</strong></p> <p>The following table contains the current leaderboard of this version for TR on the printed dataset.</p> <p>The labels, and therefore the results, are the same as in version 2.</p> <p>In this configuration, only layout block B is considered.</p> <p>These results for DAN NER are given using the named entity encoding format 5 described above.</p> <p>TR stands for Text Recognition and TR+IE for combined Text Recognition and Information Extraction.</p> <p>Metrics are expressed in percentages.</p> <p><strong>NB:</strong> The model DANIEL from [3] was not evaluated on the printed dataset.</p> <table> <tbody> <tr> <td>Method</td> <td>CER</td> <td>WER</td> </tr> <tr> <td>DAN - TR [1]</td> <td>0.88</td> <td>3.17</td> </tr> <tr> <td>DAN NER - TR + IE [1]</td> <td>1.54</td> <td>3.55</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>NER</strong></p> <p>The following table contains the current leaderboard of this version for NER on the printed dataset.</p> <p>The labels, and therefore the results, are the same as in version 2.</p> <p>In this configuration, only layout block B is considered.</p> <p>These results are given using the named entity encoding format 5 described above.</p> <p>Metrics are expressed in percentages.</p> <p><strong>NB:</strong> The model DANIEL from [3] was not evaluated on the printed dataset.</p> <table> <tbody> <tr> <td>Method</td> <td>F1</td> </tr> <tr> <td>DAN NER [1]</td> <td>93.04</td> </tr> </tbody> </table> <h2>&nbsp;</h2> <h2><strong>Citation Request</strong></h2> <p>If you publish material based on this database, we request you to include a reference to the paper&nbsp;<code><a href="https://arxiv.org/abs/2404.19329">T. Constum, L. Preel, T. Paquet, P. Tranouez, S. Br&eacute;e, End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940, International Conference on Document Analysis and Recognition (ICDAR), Athens, Greece, 2024</a>.</code></p> <p>&nbsp;</p> <h2><strong>Bibliography</strong></h2> <p><a href="https://arxiv.org/abs/2404.19329">1: T. Constum, L. Preel, T. Paquet, P. Tranouez, S. Br&eacute;e: End-to-end information extraction in handwritten documents: Understanding Paris marriage records from 1880 to 1940, International Conference on Document Analysis and Recognition (ICDAR), Athens, Greece, 2024.</a></p> <p>2: D.Coquenet, C. Chatelain, T. Paquet: DAN: a Segmentation-free Document Attention Network for Handwritten Document Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence pp. 1&ndash;17 (2023).</p> <p><a href="https://hal.science/hal-04555188/">3: T. Constum, T. Paquet, P. Tranouez: DANIEL: A fast Document Attention Network for Information Extraction and Labelling of handwritten documents, preprint, 2024&nbsp;</a></p>

opencc-by-4.0Apr 2024View details →
zenodo28/100

RealKIE: Five Novel Datasets for Enterprise Key Information Extraction

<p>These are datasets to acompany the paper "RealKIE: Five Novel Datasets for Enterprise Key Information Extraction"</p> <p>We recommend following the instructions in our Github Repo https://github.com/IndicoDataSolutions/realkie to download from Wasabi. This copy on Zenodo is to increase accessibility and ensure that the data is available indefinitely.</p> <p>Resource Contracts is in 7z format due to it's size. Others are in Zip format. The dataset formats are described in depth in our Paper and the Github Repo.</p>

opencc-by-nc-4.0Aug 2024View details →
zenodo28/100

DBpedia Information Extraction Against Amharic Wikipedia

Open the record for dataset details and reuse information.

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record