Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
75
datasets available to search
ShareScore release 0.9.0
Dataset results
75 results for “Document Datasets”
Dataset for Paper: A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents
<p>Dataset for the paper: "A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents", P. Kaddas, K. Palaiologos, B. Gatos, V. Katsouros, K. Christopoulou, 17th International Conference on Document Analysis and Recognition (ICDAR), San Jose, California, USA</p> <p>The dataset consists of 57 pages from the third edition of the Greek New Testament published by Robert Estienne (1503–1559), who was appointed “Royal Typographer” by the King of France François I (1494–1547). Robert Estienne produced this edition in 1550 using the grecs du roi typeface, produced by Claude Garamont on the basis of the Greek minuscule style of the calligrapher Angelos Vergikios (1505–1569) from Crete, who active copying Greek manuscripts in Venice and France. The dataset consists of 2045 cropped text line images in .png format with their corresponding OCR in .txt format, where 1431 used for training, 204 for validation and 410 for test. Initial images acquired from: https://bibles-online.net/flippingbook/1550/</p>
Forschungsprojekt: Digitalisierung, Klassifikationen und Gesundheits-Apps. Dataset B - Document analysis for health app users and developers
<p><strong>Allgemeine Hinweise Data Set B</strong></p> <ol> <li>Titel des Forschungsprojekts</li> </ol> <p>Digitale Gesundheitsklassifikationen in Apps - Praktiken und Probleme ihrer Entwicklung und situativen Anwendung. Projektleitung: Prof. Dr. Rainer Diaz-Bone. Bearbeitung: Valeska Cappel Dipl. Soz., Miriam Kutt (Hilfsassistenz). Laufzeit: 2019-2023. Finanzierung: Schweizer Nationalfonds.</p> <p>2. PrimärforscherInnen:</p> <ul> <li>Rainer Diaz-Bone</li> <li>Valeska Cappel</li> <li>Miriam Kutt</li> </ul> <p>3. Publikationsjahr:</p> <ul> <li>2023</li> </ul> <p>4. Hinweise zur Verfügbarkeit</p> <p>Die Daten werden über LORY (Lucerne Open Repository) dauerhaft zugänglich gemacht.</p> <ul> <li>Alle Dateien die sich auf die App-Entwicklung beziehen beginnen mit E</li> <li>Alle Dateien die sich auf die App-Nutzung beziehen beginnen mit N</li> </ul> <p>5. Fachgebiet</p> <ul> <li>Soziologie</li> </ul> <p>6. Kategorie und Schlagwörter</p> <ul> <li>Gesundheitswesen</li> <li>Selbstvermessung</li> <li>Gesundheits-Apps</li> <li>Klassifikationen</li> <li>Pragmatismus</li> <li>Economics of convention</li> <li>Soziologie der Konventionen</li> <li>Digitalisierung</li> </ul> <p>7. Abstract, wozu die Daten erhoben wurden</p> <p>Die Daten wurden für eine Dokumentenanalyse erhoben. Ausgewählt wurden (1) Daten von forschungsrelevanten Akteuren oder Institutionen, die nicht für ein Interview gewonnen werden konnten und (2) Medienberichte, Anleitungen, Beschreibungen von präventiven Gesundheits-Apps, Wissenschaftliche Berichte (White Papers, Ausschreibungen) sowie Bewertungen von diesen Gesundheits-Apps aus dem „App-Store“ und „Google-Play-Store“ (Plattformen zum Download von Apps). Ziel war es mit diesen Dokumenten weitere Analysen durchzuführen, die diskursanlytische Aussagen über die Entstehung und Nutzung von präventiven Gesundheits-Apps sowie Entwicklungen im Feld der digitalen Gesundheit zulassen. Gleichzeitig wurden auch Dokumente, wie Nutzer-Bewertungen erhoben, um die Interviews zu ergänzen aus einer pragmatischen Perspektive Aushandlungs- und Problemlösungsprozesse im Umgang mit Gesundheits-Apps und der damit verbundenen Technologie zu untersuchen.</p> <p>8. Untersuchungsgebiet</p> <p>Das Untersuchungsgebiet liegt im Bereich der digitalen Gesundheit und beschränkt sich im Speziellen auf die Prozesse der Entwicklung von Gesundheits-Apps, sowie die Nutzung der Gesundheits-Apps. Untersuchungsgebiet waren öffentlich zugängliche Medien sowie die Gesundheits-App selbst. Dabei wurden Unternehmen, Zeitschriften, Blogs und Apps, die sich konkret mit der App-Entwicklung, der App-Nutzung oder der Berichterstattung über präventive Gesundheits-Apps beschäftigen ausgewählt zur Dokumentenerhebung ausgewählt. Bei der Auswahl wurden der Fokus darauf gelegt Inhalte auszuwählen, die sich nicht auf Gesundheits-App als Medizinprodukte konzentrieren, sondern auf präventive Gesundheits-Apps, die kein spezifisches Krankheitsbild adressieren, sondern einen allgemeinen positiven Gesundheitszustand herstellen oder erhalten sollen. </p> <p>9. Gesamtheit auf die generalisiert werden könnte („Grundgesamtheit“)</p> <p>Insgesamt wurden ca. 300 Dokumente erhoben.</p> <p>10. Auswahlverfahren und Stichproben</p> <p>Die Daten im Projekt wurden anhand qualitativer Methoden gewonnen. Die Fälle und Daten wurden über die Methode der „Theoretical-Sampling-Technik“ ausgewählt. Die genauen Begründungen zur Auswahl der Fälle wurden im Verlauf des Projektes theoretisch erarbeitet. Dabei wurde sich dem Forschungsfeld der präventiven Gesundheits-Apps mit heuristischen Vermutungen angenähert, die anhand der Konzepte der Theorie der Konventionen und einer machttheoretischen Perspektive Foucaults entwickelt wurden. Gesundheit und Gesundheitshandlungen wurden dabei aus einer pragmatischen Perspektive als ein Ergebnis von Koordinationsbemühungen zwischen Akteuren, Gegenständen, Technologien und Machtstrukturen verstanden. Die Auswahl der Dokumente stützte sich besonders auf die Memos und Inhalte der vorher geführten Interviews und den daraus entwickelten heuristische Fragestellungen während des Forschungsprozesses.</p> <p>11. Erhebungszeitraum</p> <p>Die Dokumente wurden in dem Zeitraum 2019-2023 erhoben. Das Datum in den Dateinamen bezieht sich immer auf den Erhebungszeitpunkt</p> <p>Sprache</p> <ul> <li>Deutsch und Englisch</li> </ul> <p>12. Größe des Datensatzes</p> <ul> <li>ca. 45 MB</li> </ul> <p>13. Verwendete Dateiformate und notwendige Software</p> <ul> <li>Dateiformat: PDF (Portable Document Format) und RTF (Rich Text Format)</li> </ul> <p>Software:</p> <ul> <li>RTF Standard-Textprogrammen auf unterschiedlichen Betriebssystemen (Bspw. Word, Wordpad, LibreOffice, OpenOffice)</li> <li>PDF: PDF-Programme/Reader oder auch ATLAS</li> </ul> <p> </p>
TexBiG Dataset for Analysing Complex Document Layouts in the Digital Humanities
<p>This is the dataset for the paper "A Dataset for Analysing Complex Document Layouts in the Digital Humanities and its Evaluation with Krippendorff ’s Alpha" in its second version, containing an update of the test images (without annotations) from the paper "Drawing the Same Bounding Box Twice? Coping Noisy Annotations in Object Detection with Repeated Labels". Organization of the dataset is also updated to make it easier to use.</p> <p>TexBiG (from the German Text-Bild-Gefüge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century. The dataset provides instance segmentation (bounding boxes and polygons/masks) annotations for 19 different classes with more then 52.000 instances. The added test images can be used to make submission on the leaderboard on <a href="https://eval.ai/web/challenges/challenge-page/2078/overview">EvalAI</a>. </p> <p>The <a href="https://zenodo.org/record/6885144/files/Annotations_Guideline.pdf?download=1">annotation guideline</a> can be found in the first of the dataset.</p>
Dataset and documented R code for "Nouns and verbs in the speech signal"
<p>The files available constitute supplementary material to the following article:</p> <p>Lohmann, Arne. Nouns and verbs in the speech signal: Are there phonetic correlates of grammatical category? <em>Linguistics</em> - <em>An Interdisciplinary Journal of the Language Sciences</em>.</p> <p>The article is to be published online in 2020, and in 2021 in the print version of the journal.</p>
Dataset of the paper: "How do Hugging Face Models Document Datasets, Bias, and Licenses? An Empirical Study"
<p>This replication package contains datasets and scripts related to the paper: "*How do Hugging Face Models Document Datasets, Bias, and Licenses? An Empirical Study*"</p><p> </p><p>## Root directory</p><p>- `statistics.r`: R script used to compute the correlation between usage and downloads, and the RQ1/RQ2 inter-rater agreements</p><p>- `modelsInfo.zip`: zip file containing all the downloaded model cards (in JSON format)</p><p>- `script`: directory containing all the scripts used to collect and process data. For further details, see README file inside the script directory.</p><p> </p><p>## Dataset</p><p>- `Dataset/Dataset_HF-models-list.csv`: list of HF models analyzed</p><p>- `Dataset/Dataset_github-prj-list.txt`: list of GitHub projects using the *transformers* library</p><p>- `Dataset/Dataset_github-Prj_model-Used.csv`: contains usage pairs: project, model</p><p>- `Dataset/Dataset_prj-num-models-reused.csv`: number of models used by each GitHub project</p><p>- `Dataset/Dataset_model-download_num-prj_correlation.csv` contains, for each model used by GitHub projects: the name, the task, the number of reusing projects, and the number of downloads</p><p> </p><p> </p><p>## RQ1</p><p>- `RQ1/RQ1_dataset-list.txt`: list of HF datasets</p><p>- `RQ1/RQ1_datasetSample.csv`: sample set of models used for the manual analysis of datasets</p><p>- `RQ1/RQ1_analyzeDatasetTags.py`: Python script to analyze model tags for the presence of datasets. it requires to unzip the `modelsInfo.zip` in a directory with the same name (`modelsInfo`) at the root of the replication package folder. Produces the output to stdout. To redirect in a file fo be analyzed by the `RQ2/countDataset.py` script</p><p>- `RQ1/RQ1_countDataset.py`: given the output of `RQ2/analyzeDatasetTags.py` (passed as argument) produces, for each model, a list of Booleans indicating whether (i) the model only declares HF datasets, (ii) the model only declares external datasets, (iii) the model declares both, and (iv) the model is part of the sample for the manual analysis</p><p>- `RQ1/RQ1_datasetTags.csv`: output of `RQ2/analyzeDatasetTags.py`</p><p>- `RQ1/RQ1_dataset_usage_count.csv`: output of `RQ2/countDataset.py`</p><p> </p><p> </p><p>## RQ2</p><p>- `RQ2/tableBias.pdf`: table detailing the number of occurrences of different types of bias by model Task</p><p>- `RQ2/RQ2_bias_classification_sheet.csv`: results of the manual labeling</p><p>- `RQ2/RQ2_isBiased.csv`: file to compute the inter-rater agreement of whether or not a model documents Bias</p><p>- `RQ2/RQ2_biasAgrLabels.csv`: file to compute the inter-rater agreement related to bias categories</p><p>- `RQ2/RQ2_final_bias_categories_with_levels.csv`: for each model in the sample, this file lists (i) the bias leaf category, (ii) the first-level category, and (iii) the intermediate category</p><p> </p><p> </p><p>## RQ3</p><p>- `RQ3/RQ3_LicenseValidation.csv`: manual validation of a sample of licenses</p><p>- `RQ3/RQ3_{NETWORK-RESTRICTIVE|RESTRICTIVE|WEAK-RESTRICTIVE|PERMISSIVE}-license-list.txt`: lists of licenses with different permissiveness</p><p>- `RQ3/RQ3_prjs_license.csv`: for each project linked to models, among other fields it indicates the license tag and name</p><p>- `RQ3/RQ3_models_license.csv`: for each model, indicates among other pieces of info, whether the model has a license, and if yes what kind of license</p><p>- `RQ3/RQ3_model-prj-license_contingency_table.csv`: usage contingency table between projects' licenses (columns) and models' licenses (rows)</p><p>- `RQ3/RQ3_models_prjs_licenses_with_type.csv`: pairs project-model, with their respective licenses and permissiveness level</p><p> </p><p>## scripts</p><p>Contains the scripts used to mine Hugging Face and GitHub. Details are in the enclosed README</p><p> </p>
Dataset and additional files/softwares required for the paper "LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents"
<p>This dump contains all files and softwares required for running the codes for the paper "LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents". Specifically, these codes are available at https://github.com/Law-AI/LeSICiN.</p> <p>LeSICiN is a deep neural network for the task of Legal Statute Identification which also uses graphical properties of the document-statute citation network for training and predictions.</p> <p>We have three datasets --- train, dev and test. These are all .jsonl files with each instance dict per line; each instance dict contains the unique id, list of sentences and cited labels of the particular instance. Also, there is a fourth file --- secs.jsonl, which stores the text of all the statutes in similar format.</p> <p>schemas.json list out the metapath schemas for fact and section type nodes, while type_map.json maps the id of each node to its type (Act/Chapter/Topic/Section/Fact). </p> <p>label_tree.json and citation_network.json list out the edges for the two parts of the network in the format of a 3-tuple ('source id', 'relationship type', 'target id')</p> <p>"ils2v.bin" is the pretrained sent2vec vectorizer that can generate a 200-dim vector for each sentence</p>
QAngaroo (MedHop + WikiHop) - Constructing Datasets for Multi-hop Reading Comprehension Across Documents
<p>Most Reading Comprehension methods limit themselves to queries which can be answered using a single sentence, paragraph, or document. Enabling models to combine disjoint pieces of textual evidence would extend the scope of machine comprehension methods, but currently no resources exist to train and test this capability. We propose a novel task to encourage the development of models for text understanding across multiple documents and to investigate the limits of existing methods. In our task, a model learns to seek and combine evidence — effectively performing multihop, alias multi-step, inference. We devise a methodology to produce datasets for this task, given a collection of query-answer pairs and thematically linked documents. Two datasets from different domains are induced, and we identify potential pitfalls and devise circumvention strategies. We evaluate two previously proposed competitive models and find that one can integrate information across documents. However, both models struggle to select relevant information; and providing documents guaranteed to be relevant greatly improves their performance. While the models outperform several strong baselines, their best accuracy reaches 54.5% on an annotated test set, compared to human performance at 85.0%, leaving ample room for improvement.</p>
Complementary dataset of Overton metadata on citing policy-related documents for the study "From intent to impact: Investigating the effects of open sharing commitments"
<p>This document provides the underlying dataset for the bibliometric component for the 2022 study "From intent to impact: Investigating the effects of open sharing commitments" by Research Consulting and Science-Metrix.</p> <p>Before reproducing the study findings or re-using the underlying datasets for other purposes, please cautiously review their limitations in the study's technical annex and main report, available at: https://zenodo.org/communities/data-sharing-in-public-health-emergencies/ </p> <p>Special thanks from the Science-Metrix / Elsevier teams to Euan Adie and Overton for this exceptional public release of Overton metadata, and for conducting extraordinary data collection to retrieve citations towards arXiv preprints.</p> <p> </p> <p>Scope: note that this file combines cited journal publications and preprints from the Covid19, HVRD, Zika and HVVD thematic sets.</p> <p>Data treatment: this data is intend foremost to provide manual validation or qualitative triangulation of our findings. No special efforts have been made to process and clean the data for its eventual re-use in secondary analysis or text mining approaches.</p> <p>Definitions used in this table:</p> <table> <tbody> <tr> <td>Column name </td> <td>Definition</td> </tr> <tr> <td>document_type</td> <td>preprint or journal publication</td> </tr> <tr> <td>doi</td> <td>digital object identifier</td> </tr> <tr> <td>arxiv_id</td> <td>arXiv preprint server's unique identifier for its preprints</td> </tr> <tr> <td>ssrn_id</td> <td>SSRN preprint server's unique identifier for its preprints. Note that some of these IDs are contained within the DOIs also assigned to some (but not all) SSRN preprints , in the form of "10.2139/ssrn." + 'ssrn_id'</td> </tr> <tr> <td>coalesce_id</td> <td>coalesce function applied to the DOI, arxiv_id and ssrn_id. Redundant for journal publications.</td> </tr> <tr> <td>policy_source_title</td> <td>name of the policy-related organization</td> </tr> <tr> <td>published_on</td> <td>publication date of the citing policy-related document</td> </tr> <tr> <td>title</td> <td>title of the citing policy-related document</td> </tr> <tr> <td>pdf_url</td> <td>URL for the online version of the policy-related document</td> </tr> <tr> <td>snippet</td> <td>Where available, excerpt of the text immediatly before and after the citation to a journal publication or preprint found in the citing policy-related document</td> </tr> </tbody> </table> <p> </p>
Dataset for an analysis of selected studies on software development and reuse in the field of software documentation
<p>The data record contains bibliographical metadata from 2012 to 2023, which have been extracted from the databases ScienceDirect, SpringerLink and IEEE Xplore. It was created in the context of a master thesis which will be published later. The aim was to make a selection of studies on software documentation from the perspective of software development based on the master thesis.</p>
Datasets to accompany twitter reproducible methods document
<p>Datasets to accompany twitter reproducible methods document</p>
Toy dataset for metaGEM documentation (Gut v1)
<p>This dataset was generated to test and benchmark the <a href="https://github.com/franciscozorrilla/metaBAGpipes">metaGEM</a> workflow.</p> <p>The 100 bp illumina WGS reads consist of ~10% subsets of 3 paired end sets of reads from the following <a href="https://www.nature.com/articles/nature12198">publication</a>:</p> <blockquote> <p>Karlsson, Fredrik H., et al. “Gut Metagenome in European Women with Normal, Impaired and Diabetic Glucose Control.”<em> Nature</em>, vol.498,no.7452,2013,pp.99–103., doi:10.1038/nature12198.</p> </blockquote> <p>SRA Accession code <a href="https://www.ncbi.nlm.nih.gov/sra?term=ERP002469">ERP002469</a>.</p> <p>The subsets were generated using the command line tool <a href="https://github.com/lh3/seqtk">seqtk</a>:</p> <pre><code class="language-bash">seqtk sample -s100 sample_X.fastq.gz 3000000 > subset_X.fastq.gz</code></pre> <p>The sample names in the original publication used to generate the toy dataset are ERR260162 (sample1), ERR260173 (sample2), ERR260184 (sample3).</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 5)
<p>This is part 5 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 1)
<p>This is part 1 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 4)
<p>This is part 4 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 3)
<p>This is part 3 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 2)
<p>This is part 2 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 7)
<p>This is part 7 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 6)
<p>This is part 6 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 8)
<p>This is part 8 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
Multi-layout Invoice Document Dataset (MIDD)
<p>Research Purpose/Goal of Multi-Layout Invoice Document Dataset (MIDD)</p> <p>· To provide the annotated and varied invoice layout documents in IOB format to identify and extract named entities (named entity recognition) from the invoice documents to the researchers working in this domain. Obtaining a high-quality and sufficient annotated corpus for automated information extraction from unstructured documents is the biggest challenge researchers face.</p> <p>· To overcome the limitations of rule-based and template-based named entity extraction from unstructured documents traditionally used so far in information extraction approaches. Template-free processing is the only key to processing, and managing a huge pile of unstructured documents in the recent digitized era.</p> <p>· To provide varied invoice layouts so that researchers can develop a generalized AI-based model that will train on various unstructured invoice layouts. Obtained structured output can later be utilized for integrating into information management application of the organization and used for the decision-making process.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.