Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,085

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,085 results for “Documentation”

Learn how ShareScore rates datasets ↗
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 2)

<p>This is part 2 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 7)

<p>This is part 7 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 6)

<p>This is part 6 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 8)

<p>This is part 8 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br>&nbsp; title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br>&nbsp; author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br>&nbsp; journal={arXiv preprint arXiv:2408.01690},<br>&nbsp; year={2024}<br>}</p>

opencc-zeroSep 2024View details →
zenodo36/100

Escape Game "Sortez du Cube", une initiation à la documentation en médecine

<p><strong>Pr&eacute;sentation</strong></p> <p>L&rsquo;Escape Game &ldquo;Sortez du Cube&rdquo; est un Escape Game physique &agrave; destination des &eacute;tudiants en sant&eacute; mais ouvert &agrave; tous les &eacute;tudiants. Il se compose d&rsquo;une succession lin&eacute;aire d&rsquo;&eacute;nigmes dont l&rsquo;objectif de leur accomplissement est de trouver l&rsquo;indice suivant puis, en dernier lieu, la cl&eacute; pour sortir du Cube, la salle d&rsquo;Innovation p&eacute;dagogique des biblioth&egrave;ques de l&rsquo;UVSQ,&nbsp; et donc finir le jeu. L&rsquo;Escape Game est chronom&eacute;tr&eacute; et doit &ecirc;tre termin&eacute; en 30 minutes maximum. Il peut &ecirc;tre jou&eacute; &agrave; 8 joueurs en m&ecirc;me temps en pleine jauge (6 est le chiffre optimal).</p> <p><strong>Objectif</strong></p> <p>L&rsquo;Escape Game porte autant un objectif p&eacute;dagogique que de valorisation documentaire. Les &eacute;nigmes permettent en effet de signaler et d&rsquo;apprendre &agrave; manipuler de la documentation en sant&eacute;, physique et &eacute;lectronique, de la DBIST de l&rsquo;UVSQ (la base de donn&eacute;es Visible Body est au centre de ce produit de valorisation). Elles permettent &eacute;galement aux &eacute;tudiants de travailler en collaboration et de se former entre eux (&ldquo;va voir le sommaire&rdquo; ; &ldquo;ScienceDirect ? Ce doit &ecirc;tre une base de donn&eacute;es&hellip;&rdquo;) plut&ocirc;t que de recevoir l&rsquo;information de mani&egrave;re verticale. De plus, leur posture de recherche est active et donc propice &agrave; une meilleure int&eacute;gration des connaissances.</p> <p><strong>D&eacute;roul&eacute;</strong></p> <p>On fait entrer le groupe dans la salle, puis on leur demande de mettre une tenue de m&eacute;decin (ou d&rsquo;infirmier) que l&rsquo;on met &agrave; leur disposition (Photo 1). La s&eacute;ance commence par l&rsquo;inscription des joueurs (nom, pr&eacute;nom, num&eacute;ro &eacute;tudiant, classe et adresse email). On leur explique ensuite qu&rsquo;ils n&rsquo;ont pas besoin de casser ou d&rsquo;arracher le mat&eacute;riel : les &eacute;nigmes sont documentaires.&nbsp;</p> <p>On commence ensuite le sc&eacute;nario (Document 1) puis on leur donne la lettre (Document 2). Ils doivent trouver le mot de passe sur un &eacute;cran pour acc&eacute;der &agrave; l&rsquo;&eacute;nigme suivante sur un Genially (Lien 1). Cela continue ainsi, d&rsquo;&eacute;nigmes en &eacute;nigmes jusqu'&agrave; la fin du jeu.</p> <p>A la fin du jeu, nous leur faisons un point sur le site de la BU et nos ressources &eacute;lectroniques, puis nous leur demandons de remplir une &eacute;valuation de l&rsquo;activit&eacute;. Enfin, nous leur remettons un sac de goodies.</p> <p><strong>Retours</strong></p> <p>Sur les 85 &eacute;valuations re&ccedil;ues pour le moment, 100% sont positives. Les &eacute;tudiants appr&eacute;cient autant le jeu en lui-m&ecirc;me que son utilit&eacute; dans le cadre de l&rsquo;apprentissage aux comp&eacute;tences informationnelles. Les commentaires des &eacute;tudiants du parcours sant&eacute; sont particuli&egrave;rement &eacute;logieux.</p> <p><strong>Organisation et ressources</strong></p> <p>Cet Escape Game a &eacute;t&eacute; cr&eacute;&eacute; par une &eacute;quipe de 6 agents. 2 personnels de cat&eacute;gorie A, 3 personnels de cat&eacute;gorie B, 1 personnel de cat&eacute;gorie C. Sa conception a demand&eacute; la ma&icirc;trise de l&rsquo;outil Genially, une technicit&eacute; pour la cr&eacute;ation d&rsquo;&eacute;nigmes (technicit&eacute; acquise en amont gr&acirc;ce &agrave; la conception d&rsquo;autres produits ludop&eacute;dagogiques) et beaucoup de bricolages. Le montant d&eacute;pens&eacute; pour cet Escape Game s&rsquo;&eacute;l&egrave;ve &agrave; moins de 50 euros (achat d&rsquo;un minuteur).</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Uncovering Hidden Inefficiencies in the Route Availability Document

Open the record for dataset details and reuse information.

openmit-licenseNov 2024View details →
dryad36/100

Data from: Open notes sounds great, but will a provider's documentation change?

<p><strong>Background</strong>: The effects of shared clinical notes on patients, care partners, and clinicians ("open notes") were first studied as a demonstration project in 2010. Since then, multiple studies have shown clinicians agree shared progress notes are beneficial to patients, and patients and care partners report benefits from reading notes. To determine if implementing open notes at a hematology/oncology practice changed providers' documentation style, we assessed the length and readability of clinicians' notes before and after open notes implementation at an academic medical center in Boston, MA.</p> <p><strong>Methods</strong>: We analyzed 143,888 notes from 60 hematology/oncology clinicians before and after the open notes debut at Beth Israel Deaconess Medical Center, from January 1, 2012, to September 1, 2016. We measured the providers' (medical doctor/nurse practitioner) documentation styles by analyzing character length, the number of addenda, note entry mode (dictated vs. typed) and note readability. Measurements used five different readability formulas and were assessed on notes written before and after the introduction of open notes on November 25, 2013.</p> <p><strong>Results</strong>: After the introduction of open notes, the mean length of progress notes increased from 6,174 characters to 6,648 characters (P&lt;0.001), and the mean character length of the "assessment and plan" (A&amp;P) increased from 1,435 characters to 1,597 characters (P&lt;0.001). The Average Grade Level Readability of progress notes decreased from 11.50 to 11.33, and overall readability improved by 0.17 (P=0.01). There were no statistically significant changes in the length or readability of "Initial Notes" or Letters, inter-doctor communication, nor in the modality of the recording of any kind of note.</p> <p><strong>Conclusions</strong>: After the implementation of open notes, progress notes and A&amp;P sections became both longer and easier to read. This suggests clinician documenters may be responding to the perceived pressures of a transparent medical records environment.</p>

opencc-zeroJul 2021View details →
zenodo36/100

Multi-layout Invoice Document Dataset (MIDD)

<p>Research Purpose/Goal of Multi-Layout Invoice Document Dataset (MIDD)</p> <p>&middot;&nbsp;To provide the annotated and varied invoice layout documents in IOB format to identify and extract named entities (named entity recognition) from the invoice documents to the researchers working in this domain. Obtaining a high-quality and sufficient annotated corpus for automated information extraction from unstructured documents is the biggest challenge researchers face.</p> <p>&middot;&nbsp;To overcome the limitations of rule-based and template-based named entity extraction from unstructured documents traditionally used so far in information extraction approaches. Template-free processing is the only key to processing, and managing a huge pile of unstructured documents in the recent digitized era.</p> <p>&middot;&nbsp;To provide varied invoice layouts so that researchers can develop a generalized AI-based model that will train on various unstructured invoice layouts. Obtained structured output can later be utilized for integrating into information management application of the organization and used for the decision-making process.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Two document-concept representations of the biomedical literature

<p>These two datasets represent the biomedical literature (Medline abstracts and PubMedCentral articles) in the &quot;document-concept matrix&quot; format produced by <a href="https://github.com/erwanm/tdc-tools">TDC Tools</a>.&nbsp; These datasets can be used in downstream IR applications such as Literature-Based Discovery.</p> <p>Each of the two datasets corresponds to a specific data extraction method, see details <a href="https://erwanm.github.io/tdc-tools/input-data-format/">here</a> and in the paper linked below.</p> <ul> <li>Paper: <em>pending </em></li> <li>Code: <a href="https://github.com/erwanm/tdc-tools">https://github.com/erwanm/tdc-tools</a> <ul> <li>Documentation:<a href="https://erwanm.github.io/tdc-tools/">https://erwanm.github.io/tdc-tools/</a></li> </ul> </li> </ul> <p><strong>Important:</strong> the raw data from which this data is derived was downloaded from <a href="https://www.nlm.nih.gov/medline/medline_overview.html">Medline</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/">PubMedCentral</a> and <a href="https://www.ncbi.nlm.nih.gov/research/pubtator/">PubTatorCentral</a>, provided <a href="https://www.nlm.nih.gov/databases/download/terms_and_conditions.html">courtesy of the U.S. National Library of Medicine (NLM)</a>. The data was extracted in January 2021 and do not reflect the most current/accurate data available from NLM. See the github repository above in order to generate similar datasets from up to date data.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

DocMine: A Software Documentation-Related Dataset of 950 GitHub Repositories

<p>DocMine dataset consists of textual information collated from multiple software artifacts, across 950 GitHub Repositories. It also consists of probable percentage contribution of text in each software artifact towards different documentation types in each repository, accompanied by metadata information about the repository such as stargazer count, number of pull requests, commits, issues and other files analyzed.</p>

opencc-by-nc-sa-4.0Jan 2021View details →
zenodo36/100

ICDAR 2021 Historical Document Classification Dataset for Task 2 - Dating

<p>Dataset for the localization classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and ground truth file (CSV format). The metadata csv file contains information, such as where the image comes from.</p>

opencc-by-4.0May 2021View details →
zenodo36/100

Nicosia, Bedestan. Cosmati work under the dome as documented in 1980-81 with accompanying reconstruction.

<p>Nicosia, Bedestan. Cosmati work under the dome as documented in 1980-81 with accompanying reconstruction. Drawing by M. Willis and Vicki Herring.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Data for Seed-driven Document Ranking for Systematic Reviews: A Reproducibility Study

<p>Data for Seed-driven Document Ranking for Systematic Reviews: A Reproducibility Study</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

Collection of documents of the IPCC DDC at WDCC

<p>Materials about the plans of the IPCC DDC at WDCC/DKRZ for AR6 circled among the participants of the First IPCC AR6 Data Workshop (19/20 September 2017) at DKRZ in Hamburg, Germany, prior to the meeting. It includes a talk given at the IPCC Expert Meeting on the future of TGICA (01/2016 in Geneva, Switzerland) and a draft list of variables for the CMIP6 data pool compiled from the DICAD project partners' data requests and from statistics of the AR5 variable usage (status: 06/2017).</p>

opencc-by-4.0Sep 2017View details →
zenodo36/100

Long document similarity dataset, Wikipedia excerptions for movies collections

<p>Movies-related articles&nbsp;extracted from Wikipedia.</p> <p>For all articles, the figures and tables have been filtered out, as well as the categories and &quot;see also&quot; sections.</p> <p>The article structure, and&nbsp;particularly the sub-titles and paragraphs are kept in these datasets</p> <p>&nbsp;</p> <p><strong>Movies</strong></p> <p>The Wikipedia Movies dataset consists of 100,371 articles describing various movies. Each article may consist of text passages describing the plot, cast, production, reception, soundtrack, and more.</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

Airborne eDNA documents a diverse and ecologically complex tropical bat and other mammal community

<p><span>Environmental (e)DNA has rapidly become a powerful biomonitoring tool, particularly in aquatic ecosystems. This approach has not been as widely adopted in terrestrial communities where the methods of vertebrate eDNA collection have varied from the use of secondary collectors such as blood-feeding parasites and spider webs to washing surfaces of leaves and soil sampling. Recent studies have demonstrated the potential of direct collection of eDNA from air sampling, but none have tested how effective airborne eDNA sampling might be in a </span><span>biodiverse environment.</span> <span>We used three prototype samplers to actively sample a mixed neotropical bat community in a partially controlled environment. We assess whether airborne eDNA can accurately characterize a high-diversity community with skewed abundances and to determine if filter design impacts DNA collection and taxonomic recovery. Our study provides evidence for the accuracy of airborne eDNA as a detection tool and highlights its potential for monitoring high-density, diverse assemblages such as </span><span>bat roosts. </span><span>Analysis of air samples recovered &gt;91% of the species present and some limited relationship between species abundance and read count. Our data suggest this method can accurately depict a diverse mixed mammal community, particularly when the location is contained (e.g., a roost, den or burrow) but also highlights the potential for secondary transfer of eDNA material on clothing and equipment. Our results also demonstrate that simple, inexpensive, battery-operated homemade air samplers can collect an abundance of eDNA from the air, opening the opportunity for sampling in remote environments. </span></p>

opencc-zeroDec 2022View details →
zenodo36/100

DOCUMENTED HUMAN OSTEOLOGICAL COLLECTIONS AS BIOBANKS: RELEVANCE FOR RARE DISEASES IDENTIFICATION IN THE PAST

<p><em><strong>Presented at:&nbsp;23rd Paleopathology Association European Meeting, Vilnius, Litu&acirc;nia, 25-29 Agosto. Paleopathology Association European (Vilnius, Litu&acirc;nia)</strong></em></p> <p>Disease identification in paleopathology relies on the exercise of differential diagnosis, and interpretation. Only a few diseases leave macroscopic pathognomonic traits in bone, and even in cases where microscopic, biochemical and biomolecular analyses are used, diagnosis is invariably inconclusive. Additionally, bone response to a variety of etiologies tends to be homogenous, with mosaic pattern(s) of bone formation and destruction. Therefore, access to pathological cases from human remains of Documented Human Osteological Collections (DHOC) is an exceptional approach. The access to biographical data of the individuals incorporated into the DHOC includes the cause of death, ancestry, sex, age, clinical data and other information akin to clinical data allowing for the possibility of hypothesis-driven research in which bones changes correlate with causes of death - hence providing tested and informed differential diagnosis. In this sense, DHOC may be viewed as a biobank equivalent, i.e. biorepository that stores biological samples for research in the identification of bone changes related to diseases associated with clinical and personal data. This paper will explore known cases of diseases&rsquo; diagnoses, such as lepra, neoplasias, tuberculosis, syphilis, and diffuse idiopathic skeletal hyperostosis that have used DHOC as diagnostic testing grounds, to explore bone changes and methodological advancements. The paper also introduces the idea of DHOC as biobanks dedicated to the study of rare diseases, as rarely reposted diseases, in paleopathology.&nbsp;</p> <p><strong>Keywords: </strong>Health, biorepository, biobanks, DHOC, differential diagnosis</p>

opencc-by-nc-nd-4.0Dec 2022View details →
zenodo36/100

Open-source Software Governance Documentation Dataset on GitHub

<p>This dataset contains 710 GitHub-hosted OSS projects, which contain a governance file in the root directory of the project. It also contains commits, issues, and comments on each project.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Aesthetic Responses to the Covid-19 Crisis: The corona performance No Problama Video/Audio Documentation

<p>The video files document the artistic performance&nbsp;<em>No Problama, a coronavirus story from Homborsund&nbsp;</em>created during the Covid-19 crisis in Grimstad, Norway in 2021. Each video presents components of the performance which are described and analysed in the article&nbsp;<strong>Aesthetic responses to the Covid-19 pandemic: The Corona-performance&nbsp;<em>No Problama 2021</em></strong>&nbsp;&nbsp;which belongs to the&nbsp;Routledge Open Access Collection<em> Art and Crisis.&nbsp;</em>All participants have given their written consent for the publication.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Mid.ru press documents 2003–2019, Russian

<p>Document collection scraped from the Russian governmental website mid.ru. Includes all items (transcripts, comments etc.) listed on:<br> - https://mid.ru/ru/press_service/minister_speeches/,<br> - https://mid.ru/ru/press_service/deputy_ministers_speeches/,<br> - https://mid.ru/ru/press_service/telefonnye-razgovory-ministra/,<br> - https://mid.ru/ru/press_service/spokesman/briefings/,<br> - https://mid.ru/ru/press_service/spokesman/official_statement/,<br> - https://mid.ru/ru/press_service/spokesman/answers/,<br> - https://mid.ru/ru/press_service/spokesman/kommentarii/,<br> and on their following pages (e.g. https://mid.ru/ru/press_service/minister_speeches/?PAGEN_1=2) from the first documents&nbsp;(4 January 2003) until the end of 2019.</p> <p>11,857 documents. One document in each row. Columns:</p> <p>- ID: format MID-1&nbsp;&nbsp;<br> - ID_no: format 1&nbsp;&nbsp;<br> - Date: Document date, format 2019-12-31&nbsp;<br> - Title: Document title<br> - Type: One of the following: Брифинги; Выступления заместителей Министра; Выступления Министра; Комментарии; Ответы на вопросы СМИ; Официальные заявления; Телефонные разговоры Министра<br> - Text: Document text including title&nbsp;&nbsp;<br> - URL: URL from which the document&nbsp;is downloaded&nbsp;&nbsp;<br> - Downloaded: Date of download, format 2019-12-31</p> <p>Formats: rds and json.</p> <p>Version 1.1: edited column names.</p> <p>---</p> <p>From www.mid.ru:</p> <p>Materials on the website of the Russian Ministry of Foreign Affairs are generally accessible and open for non-commercial use (personal, family, education, research, etc.).<br> Their reprinting, as well as any quoting in the mass media is allowed only with a reference to the website of the Russian Ministry of Foreign Affairs as a source of the information.</p>

openother-ncJan 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record