Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

75

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

75 results for “Document Datasets”

Learn how ShareScore rates datasets ↗
zenodo40/100

Dataset for Paper: A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents

<p>Dataset for the paper: &quot;A System for Processing and Recognition of Greek Byzantine and Post-Byzantine Documents&quot;, P. Kaddas, K. Palaiologos, B. Gatos, V. Katsouros, K. Christopoulou, 17th&nbsp;International Conference on Document Analysis and Recognition (ICDAR), San Jose, California, USA</p> <p>The dataset consists&nbsp;of 57 pages from the third edition of the Greek New Testament published by Robert Estienne (1503&ndash;1559), who was appointed &ldquo;Royal Typographer&rdquo; by the King of France Fran&ccedil;ois I (1494&ndash;1547). Robert Estienne produced this edition in 1550 using the grecs du roi typeface, produced by Claude Garamont on the basis of the Greek minuscule style of the calligrapher Angelos Vergikios (1505&ndash;1569) from Crete, who active copying Greek manuscripts in Venice and France. The dataset consists of 2045 cropped text line images in .png format with their corresponding OCR in .txt format, where 1431 used for training, 204 for validation and 410 for test. Initial images acquired from: https://bibles-online.net/flippingbook/1550/</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Forschungsprojekt: Digitalisierung, Klassifikationen und Gesundheits-Apps. Dataset B - Document analysis for health app users and developers

<p><strong>Allgemeine Hinweise Data Set B</strong></p> <ol> <li>Titel des Forschungsprojekts</li> </ol> <p>Digitale Gesundheitsklassifikationen in Apps - Praktiken und Probleme ihrer Entwicklung und situativen Anwendung. Projektleitung: Prof. Dr. Rainer Diaz-Bone. Bearbeitung: Valeska Cappel Dipl. Soz., Miriam Kutt (Hilfsassistenz). Laufzeit: 2019-2023. Finanzierung: Schweizer Nationalfonds.</p> <p>2. Prim&auml;rforscherInnen:</p> <ul> <li>Rainer Diaz-Bone</li> <li>Valeska Cappel</li> <li>Miriam Kutt</li> </ul> <p>3. Publikationsjahr:</p> <ul> <li>2023</li> </ul> <p>4. Hinweise zur Verf&uuml;gbarkeit</p> <p>Die Daten werden &uuml;ber LORY (Lucerne Open Repository) dauerhaft zug&auml;nglich gemacht.</p> <ul> <li>Alle Dateien die sich auf die App-Entwicklung beziehen beginnen mit E</li> <li>Alle Dateien die sich auf die App-Nutzung beziehen beginnen mit N</li> </ul> <p>5. Fachgebiet</p> <ul> <li>Soziologie</li> </ul> <p>6. Kategorie und Schlagw&ouml;rter</p> <ul> <li>Gesundheitswesen</li> <li>Selbstvermessung</li> <li>Gesundheits-Apps</li> <li>Klassifikationen</li> <li>Pragmatismus</li> <li>Economics of convention</li> <li>Soziologie der Konventionen</li> <li>Digitalisierung</li> </ul> <p>7. Abstract, wozu die Daten erhoben wurden</p> <p>Die Daten wurden f&uuml;r eine Dokumentenanalyse erhoben. Ausgew&auml;hlt wurden (1) Daten von forschungsrelevanten Akteuren oder Institutionen, die nicht f&uuml;r ein Interview gewonnen werden konnten und (2) Medienberichte, Anleitungen, Beschreibungen von pr&auml;ventiven Gesundheits-Apps, Wissenschaftliche Berichte (White Papers, Ausschreibungen) sowie Bewertungen von diesen Gesundheits-Apps aus dem &bdquo;App-Store&ldquo; und &bdquo;Google-Play-Store&ldquo; (Plattformen zum Download von Apps). Ziel war es mit diesen Dokumenten weitere Analysen durchzuf&uuml;hren, die diskursanlytische Aussagen &uuml;ber die Entstehung und Nutzung von pr&auml;ventiven Gesundheits-Apps sowie Entwicklungen im Feld der digitalen Gesundheit zulassen. Gleichzeitig wurden auch Dokumente, wie Nutzer-Bewertungen erhoben, um die Interviews zu erg&auml;nzen aus einer pragmatischen Perspektive Aushandlungs- und Probleml&ouml;sungsprozesse im Umgang mit Gesundheits-Apps und der damit verbundenen Technologie zu untersuchen.</p> <p>8. Untersuchungsgebiet</p> <p>Das Untersuchungsgebiet liegt im Bereich der digitalen Gesundheit und beschr&auml;nkt sich im Speziellen auf die Prozesse der Entwicklung von Gesundheits-Apps, sowie die Nutzung der Gesundheits-Apps. Untersuchungsgebiet waren &ouml;ffentlich zug&auml;ngliche Medien sowie die Gesundheits-App selbst. Dabei wurden Unternehmen, Zeitschriften, Blogs und Apps, die sich konkret mit der App-Entwicklung, der App-Nutzung oder der Berichterstattung &uuml;ber pr&auml;ventive Gesundheits-Apps besch&auml;ftigen ausgew&auml;hlt zur Dokumentenerhebung ausgew&auml;hlt. Bei der Auswahl wurden der Fokus darauf gelegt Inhalte auszuw&auml;hlen, die sich nicht auf Gesundheits-App als Medizinprodukte konzentrieren, sondern auf pr&auml;ventive Gesundheits-Apps, die kein spezifisches Krankheitsbild adressieren, sondern einen allgemeinen positiven Gesundheitszustand herstellen oder erhalten sollen.&nbsp;</p> <p>9. Gesamtheit auf die generalisiert werden k&ouml;nnte (&bdquo;Grundgesamtheit&ldquo;)</p> <p>Insgesamt wurden ca. 300 Dokumente erhoben.</p> <p>10. Auswahlverfahren und Stichproben</p> <p>Die Daten im Projekt wurden anhand qualitativer Methoden gewonnen. Die F&auml;lle und Daten wurden &uuml;ber die Methode der &bdquo;Theoretical-Sampling-Technik&ldquo; ausgew&auml;hlt. Die genauen Begr&uuml;ndungen zur Auswahl der F&auml;lle wurden im Verlauf des Projektes theoretisch erarbeitet. Dabei wurde sich dem Forschungsfeld der pr&auml;ventiven Gesundheits-Apps mit heuristischen Vermutungen angen&auml;hert, die anhand der Konzepte der Theorie der Konventionen und einer machttheoretischen Perspektive Foucaults entwickelt wurden. Gesundheit und Gesundheitshandlungen wurden dabei aus einer pragmatischen Perspektive als ein Ergebnis von Koordinationsbem&uuml;hungen zwischen Akteuren, Gegenst&auml;nden, Technologien und Machtstrukturen verstanden. Die Auswahl der Dokumente st&uuml;tzte sich besonders auf die Memos und Inhalte der vorher gef&uuml;hrten Interviews und den daraus entwickelten heuristische Fragestellungen w&auml;hrend des Forschungsprozesses.</p> <p>11. Erhebungszeitraum</p> <p>Die Dokumente wurden in dem Zeitraum 2019-2023 erhoben. Das Datum in den Dateinamen bezieht sich immer auf den Erhebungszeitpunkt</p> <p>Sprache</p> <ul> <li>Deutsch und Englisch</li> </ul> <p>12. Gr&ouml;&szlig;e des Datensatzes</p> <ul> <li>ca. 45 MB</li> </ul> <p>13. Verwendete Dateiformate und notwendige Software</p> <ul> <li>Dateiformat: PDF (Portable Document Format) und RTF (Rich Text Format)</li> </ul> <p>Software:</p> <ul> <li>RTF Standard-Textprogrammen auf unterschiedlichen Betriebssystemen (Bspw. Word, Wordpad, LibreOffice, OpenOffice)</li> <li>PDF: PDF-Programme/Reader oder auch ATLAS</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo40/100

TexBiG Dataset for Analysing Complex Document Layouts in the Digital Humanities

<p>This is the dataset for the paper&nbsp;&quot;A Dataset for Analysing Complex Document Layouts in the Digital Humanities and its Evaluation with Krippendorff &rsquo;s Alpha&quot; in its second version, containing an update of the test images (without annotations) from the paper &quot;Drawing the Same Bounding Box Twice? Coping Noisy Annotations in Object Detection with Repeated Labels&quot;. Organization of the dataset is also updated to make it easier to use.</p> <p>TexBiG (from the German Text-Bild-Gef&uuml;ge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century. The dataset provides instance segmentation (bounding boxes and polygons/masks) annotations for 19 different classes with more then 52.000 instances. The added test images can be used to make submission on the leaderboard on <a href="https://eval.ai/web/challenges/challenge-page/2078/overview">EvalAI</a>.&nbsp;</p> <p>The <a href="https://zenodo.org/record/6885144/files/Annotations_Guideline.pdf?download=1">annotation guideline</a> can be found in the first of the dataset.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

Dataset and documented R code for "Nouns and verbs in the speech signal"

<p>The files available constitute supplementary material to the following article:</p> <p>Lohmann, Arne. Nouns and verbs in the speech signal: Are there phonetic correlates of grammatical category? <em>Linguistics</em> - <em>An Interdisciplinary Journal of the Language Sciences</em>.</p> <p>The article is to be published online in 2020, and in 2021 in the print version of the journal.</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Dataset of the paper: "How do Hugging Face Models Document Datasets, Bias, and Licenses? An Empirical Study"

<p>This replication package contains datasets and scripts related to the paper: "*How do Hugging Face Models Document Datasets, Bias, and Licenses? An Empirical Study*"</p><p>&nbsp;</p><p>## Root directory</p><p>- `statistics.r`: R script used to compute the correlation between usage and downloads, and the RQ1/RQ2 inter-rater agreements</p><p>- `modelsInfo.zip`: zip file containing all the downloaded model cards (in JSON format)</p><p>- `script`: directory containing all the scripts used to collect and process data. For further details, see README file inside the script directory.</p><p>&nbsp;</p><p>## Dataset</p><p>- `Dataset/Dataset_HF-models-list.csv`: list of HF models analyzed</p><p>- `Dataset/Dataset_github-prj-list.txt`: list of GitHub projects using the *transformers* library</p><p>- `Dataset/Dataset_github-Prj_model-Used.csv`: contains usage pairs: project, model</p><p>- `Dataset/Dataset_prj-num-models-reused.csv`: number of models used by each GitHub project</p><p>- `Dataset/Dataset_model-download_num-prj_correlation.csv` contains, for each model used by GitHub projects: the name, the task, the number of reusing projects, and the number of downloads</p><p>&nbsp;</p><p>&nbsp;</p><p>## RQ1</p><p>- `RQ1/RQ1_dataset-list.txt`: list of HF datasets</p><p>- `RQ1/RQ1_datasetSample.csv`: sample set of models used for the manual analysis of datasets</p><p>- `RQ1/RQ1_analyzeDatasetTags.py`: Python script to analyze model tags for the presence of datasets. it requires to unzip the `modelsInfo.zip` in a directory with the same name (`modelsInfo`) at the root of the replication package folder. Produces the output to stdout. To redirect in a file fo be analyzed by the `RQ2/countDataset.py` script</p><p>- `RQ1/RQ1_countDataset.py`: given the output of `RQ2/analyzeDatasetTags.py` (passed as argument) produces, for each model, a list of Booleans indicating whether (i) the model only declares HF datasets, (ii) the model only declares external datasets, (iii) the model declares both, and (iv) the model is part of the sample for the manual analysis</p><p>- `RQ1/RQ1_datasetTags.csv`: output of `RQ2/analyzeDatasetTags.py`</p><p>- `RQ1/RQ1_dataset_usage_count.csv`: output of `RQ2/countDataset.py`</p><p>&nbsp;</p><p>&nbsp;</p><p>## RQ2</p><p>- `RQ2/tableBias.pdf`: table detailing the number of occurrences of different types of bias by model Task</p><p>- `RQ2/RQ2_bias_classification_sheet.csv`: &nbsp;results of the manual labeling</p><p>- `RQ2/RQ2_isBiased.csv`: file to compute the inter-rater agreement of whether or not a model documents Bias</p><p>- `RQ2/RQ2_biasAgrLabels.csv`: &nbsp;file to compute the inter-rater agreement related to bias categories</p><p>- `RQ2/RQ2_final_bias_categories_with_levels.csv`: for each model in the sample, this file lists (i) the bias leaf category, (ii) the first-level category, and (iii) the intermediate category</p><p>&nbsp;</p><p>&nbsp;</p><p>## RQ3</p><p>- `RQ3/RQ3_LicenseValidation.csv`: manual validation of a sample of licenses</p><p>- `RQ3/RQ3_{NETWORK-RESTRICTIVE|RESTRICTIVE|WEAK-RESTRICTIVE|PERMISSIVE}-license-list.txt`: lists of licenses with different permissiveness</p><p>- `RQ3/RQ3_prjs_license.csv`: for each project linked to models, among other fields it indicates the license tag and name</p><p>- `RQ3/RQ3_models_license.csv`: for each model, indicates among other pieces of info, whether the model has a license, and if yes what kind of license</p><p>- `RQ3/RQ3_model-prj-license_contingency_table.csv`: usage contingency table between projects' licenses (columns) and models' licenses (rows)</p><p>- `RQ3/RQ3_models_prjs_licenses_with_type.csv`: pairs project-model, with their respective licenses and permissiveness level</p><p>&nbsp;</p><p>## scripts</p><p>Contains the scripts used to mine Hugging Face and GitHub. Details are in the enclosed README</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

Dataset and additional files/softwares required for the paper "LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents"

<p>This dump contains all files and softwares required for running the codes for the paper&nbsp;&quot;LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents&quot;. Specifically, these codes are available at&nbsp;https://github.com/Law-AI/LeSICiN.</p> <p>LeSICiN is a deep neural network for the task of Legal Statute Identification which also uses graphical properties of the document-statute citation network for training and predictions.</p> <p>We have three datasets --- train, dev and test. These are all .jsonl files with each instance dict per line; each instance dict contains the unique id, list of sentences and cited labels of the particular instance. Also, there is a fourth file --- secs.jsonl, which stores the text of all the statutes in similar format.</p> <p>schemas.json list out the metapath schemas for fact and section type nodes, while type_map.json maps the id of each node to its type (Act/Chapter/Topic/Section/Fact).&nbsp;</p> <p>label_tree.json and citation_network.json list out the edges for the two parts of the network in the format of a 3-tuple (&#39;source id&#39;, &#39;relationship type&#39;, &#39;target id&#39;)</p> <p>&quot;ils2v.bin&quot; is the pretrained sent2vec vectorizer that can generate a 200-dim vector for each sentence</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

QAngaroo (MedHop + WikiHop) - Constructing Datasets for Multi-hop Reading Comprehension Across Documents

<p>Most Reading Comprehension methods limit themselves to queries which can be answered using a single sentence, paragraph, or document. Enabling models to combine disjoint pieces of textual evidence would extend the scope of machine comprehension methods, but currently no resources exist to train and test this capability. We propose a novel task to encourage the development of models for text understanding across multiple documents and to investigate the limits of existing methods. In our task, a model learns to seek and combine evidence &mdash; effectively performing multihop, alias multi-step, inference. We devise a methodology to produce datasets for this task, given a collection of query-answer pairs and thematically linked documents. Two datasets from different domains are induced, and we identify potential pitfalls and devise circumvention strategies. We evaluate two previously proposed competitive models and find that one can integrate information across documents. However, both models struggle to select relevant information; and providing documents guaranteed to be relevant greatly improves their performance. While the models outperform several strong baselines, their best accuracy reaches 54.5% on an annotated test set, compared to human performance at 85.0%, leaving ample room for improvement.</p>

opencc-by-sa-3.0Jun 2018View details →
zenodo36/100

Complementary dataset of Overton metadata on citing policy-related documents for the study "From intent to impact: Investigating the effects of open sharing commitments"

<p>This document provides the underlying dataset for the bibliometric component for the 2022 study &quot;From intent to impact: Investigating the effects of open sharing commitments&quot; by Research Consulting and Science-Metrix.</p> <p>Before reproducing the study findings or re-using the underlying datasets for other purposes, please cautiously review their limitations in the study&#39;s technical annex and main report, available at: https://zenodo.org/communities/data-sharing-in-public-health-emergencies/&nbsp;</p> <p>Special thanks from the Science-Metrix / Elsevier teams to Euan Adie and Overton for this exceptional public release of Overton metadata, and for conducting extraordinary data collection to retrieve citations towards arXiv preprints.</p> <p>&nbsp;</p> <p>Scope: note that this file combines cited journal publications and preprints from the Covid19, HVRD, Zika and HVVD thematic sets.</p> <p>Data treatment: this data is intend foremost to provide manual validation or qualitative triangulation of our findings. No special efforts have been made to process&nbsp;and clean the data&nbsp;for its eventual re-use in secondary analysis or&nbsp;text mining approaches.</p> <p>Definitions used in this table:</p> <table> <tbody> <tr> <td>Column name&nbsp;</td> <td>Definition</td> </tr> <tr> <td>document_type</td> <td>preprint or journal publication</td> </tr> <tr> <td>doi</td> <td>digital object identifier</td> </tr> <tr> <td>arxiv_id</td> <td>arXiv preprint server&#39;s unique identifier for its preprints</td> </tr> <tr> <td>ssrn_id</td> <td>SSRN preprint server&#39;s unique identifier for its preprints. Note that some of these IDs are contained within the DOIs also assigned to some (but not all) SSRN preprints , in the form of &quot;10.2139/ssrn.&quot; + &#39;ssrn_id&#39;</td> </tr> <tr> <td>coalesce_id</td> <td>coalesce function applied to the DOI, arxiv_id and ssrn_id. Redundant for journal publications.</td> </tr> <tr> <td>policy_source_title</td> <td>name of the policy-related organization</td> </tr> <tr> <td>published_on</td> <td>publication date of the citing policy-related document</td> </tr> <tr> <td>title</td> <td>title of the citing policy-related document</td> </tr> <tr> <td>pdf_url</td> <td>URL for the online version of the policy-related document</td> </tr> <tr> <td>snippet</td> <td>Where available, excerpt of the text immediatly before and after the citation to a journal publication or preprint found in the citing policy-related document</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Dataset for an analysis of selected studies on software development and reuse in the field of software documentation

<p>The data record contains bibliographical metadata from 2012 to 2023, which have been extracted from the databases ScienceDirect, SpringerLink and IEEE Xplore. It was created in the context of a master thesis which will be published later. The aim was to make a selection of studies on software documentation from the perspective of software development based on the master thesis.</p>

opencc-zeroJul 2022View details →
zenodo36/100

Datasets to accompany twitter reproducible methods document

<p>Datasets to accompany twitter reproducible methods document</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Toy dataset for metaGEM documentation (Gut v1)

<p>This dataset was generated to test and benchmark the <a href="https://github.com/franciscozorrilla/metaBAGpipes">metaGEM</a> workflow.</p> <p>The 100 bp illumina WGS reads consist of ~10% subsets of 3 paired end sets of reads from the following <a href="https://www.nature.com/articles/nature12198">publication</a>:</p> <blockquote> <p>Karlsson, Fredrik H., et al. &ldquo;Gut Metagenome in European Women with Normal, Impaired and Diabetic Glucose Control.&rdquo;<em> Nature</em>, vol.498,no.7452,2013,pp.99&ndash;103., doi:10.1038/nature12198.</p> </blockquote> <p>SRA Accession code&nbsp;<a href="https://www.ncbi.nlm.nih.gov/sra?term=ERP002469">ERP002469</a>.</p> <p>The subsets were generated using the command line tool <a href="https://github.com/lh3/seqtk">seqtk</a>:</p> <pre><code class="language-bash">seqtk sample -s100 sample_X.fastq.gz 3000000 &gt; subset_X.fastq.gz</code></pre> <p>The sample names in the original publication used to generate the toy dataset are ERR260162 (sample1), ERR260173 (sample2), ERR260184 (sample3).</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 5)

<p>This is part 5 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 1)

<p>This is part 1 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 4)

<p>This is part 4 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 3)

<p>This is part 3 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 2)

<p>This is part 2 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 7)

<p>This is part 7 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 6)

<p>This is part 6 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 8)

<p>This is part 8 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

Multi-layout Invoice Document Dataset (MIDD)

<p>Research Purpose/Goal of Multi-Layout Invoice Document Dataset (MIDD)</p> <p>&middot;&nbsp;To provide the annotated and varied invoice layout documents in IOB format to identify and extract named entities (named entity recognition) from the invoice documents to the researchers working in this domain. Obtaining a high-quality and sufficient annotated corpus for automated information extraction from unstructured documents is the biggest challenge researchers face.</p> <p>&middot;&nbsp;To overcome the limitations of rule-based and template-based named entity extraction from unstructured documents traditionally used so far in information extraction approaches. Template-free processing is the only key to processing, and managing a huge pile of unstructured documents in the recent digitized era.</p> <p>&middot;&nbsp;To provide varied invoice layouts so that researchers can develop a generalized AI-based model that will train on various unstructured invoice layouts. Obtained structured output can later be utilized for integrating into information management application of the organization and used for the decision-making process.</p>

opencc-by-4.0Jul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record