Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,085
datasets available to search
ShareScore release 0.9.0
Dataset results
1,085 results for “Documentation”
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 2)
<p>This is part 2 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 7)
<p>This is part 7 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 6)
<p>This is part 6 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection (part 8)
<p>This is part 8 of the IDNet dataset of our research paper "IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection". Here's a link to the paper: https://arxiv.org/pdf/2408.01690</p> <p>Citation:</p> <p>@article{guan2024idnet,<br> title={IDNet: A Novel Dataset for Identity Document Analysis and Fraud Detection},<br> author={Guan, Hong and Wang, Yancheng and Xie, Lulu and Nag, Soham and Goel, Rajeev and Swamy, Niranjan Erappa Narayana and Yang, Yingzhen and Xiao, Chaowei and Prisby, Jonathan and Maciejewski, Ross and Zou, Jia},<br> journal={arXiv preprint arXiv:2408.01690},<br> year={2024}<br>}</p>
Escape Game "Sortez du Cube", une initiation à la documentation en médecine
<p><strong>Présentation</strong></p> <p>L’Escape Game “Sortez du Cube” est un Escape Game physique à destination des étudiants en santé mais ouvert à tous les étudiants. Il se compose d’une succession linéaire d’énigmes dont l’objectif de leur accomplissement est de trouver l’indice suivant puis, en dernier lieu, la clé pour sortir du Cube, la salle d’Innovation pédagogique des bibliothèques de l’UVSQ, et donc finir le jeu. L’Escape Game est chronométré et doit être terminé en 30 minutes maximum. Il peut être joué à 8 joueurs en même temps en pleine jauge (6 est le chiffre optimal).</p> <p><strong>Objectif</strong></p> <p>L’Escape Game porte autant un objectif pédagogique que de valorisation documentaire. Les énigmes permettent en effet de signaler et d’apprendre à manipuler de la documentation en santé, physique et électronique, de la DBIST de l’UVSQ (la base de données Visible Body est au centre de ce produit de valorisation). Elles permettent également aux étudiants de travailler en collaboration et de se former entre eux (“va voir le sommaire” ; “ScienceDirect ? Ce doit être une base de données…”) plutôt que de recevoir l’information de manière verticale. De plus, leur posture de recherche est active et donc propice à une meilleure intégration des connaissances.</p> <p><strong>Déroulé</strong></p> <p>On fait entrer le groupe dans la salle, puis on leur demande de mettre une tenue de médecin (ou d’infirmier) que l’on met à leur disposition (Photo 1). La séance commence par l’inscription des joueurs (nom, prénom, numéro étudiant, classe et adresse email). On leur explique ensuite qu’ils n’ont pas besoin de casser ou d’arracher le matériel : les énigmes sont documentaires. </p> <p>On commence ensuite le scénario (Document 1) puis on leur donne la lettre (Document 2). Ils doivent trouver le mot de passe sur un écran pour accéder à l’énigme suivante sur un Genially (Lien 1). Cela continue ainsi, d’énigmes en énigmes jusqu'à la fin du jeu.</p> <p>A la fin du jeu, nous leur faisons un point sur le site de la BU et nos ressources électroniques, puis nous leur demandons de remplir une évaluation de l’activité. Enfin, nous leur remettons un sac de goodies.</p> <p><strong>Retours</strong></p> <p>Sur les 85 évaluations reçues pour le moment, 100% sont positives. Les étudiants apprécient autant le jeu en lui-même que son utilité dans le cadre de l’apprentissage aux compétences informationnelles. Les commentaires des étudiants du parcours santé sont particulièrement élogieux.</p> <p><strong>Organisation et ressources</strong></p> <p>Cet Escape Game a été créé par une équipe de 6 agents. 2 personnels de catégorie A, 3 personnels de catégorie B, 1 personnel de catégorie C. Sa conception a demandé la maîtrise de l’outil Genially, une technicité pour la création d’énigmes (technicité acquise en amont grâce à la conception d’autres produits ludopédagogiques) et beaucoup de bricolages. Le montant dépensé pour cet Escape Game s’élève à moins de 50 euros (achat d’un minuteur).</p>
Uncovering Hidden Inefficiencies in the Route Availability Document
Open the record for dataset details and reuse information.
Data from: Open notes sounds great, but will a provider's documentation change?
<p><strong>Background</strong>: The effects of shared clinical notes on patients, care partners, and clinicians ("open notes") were first studied as a demonstration project in 2010. Since then, multiple studies have shown clinicians agree shared progress notes are beneficial to patients, and patients and care partners report benefits from reading notes. To determine if implementing open notes at a hematology/oncology practice changed providers' documentation style, we assessed the length and readability of clinicians' notes before and after open notes implementation at an academic medical center in Boston, MA.</p> <p><strong>Methods</strong>: We analyzed 143,888 notes from 60 hematology/oncology clinicians before and after the open notes debut at Beth Israel Deaconess Medical Center, from January 1, 2012, to September 1, 2016. We measured the providers' (medical doctor/nurse practitioner) documentation styles by analyzing character length, the number of addenda, note entry mode (dictated vs. typed) and note readability. Measurements used five different readability formulas and were assessed on notes written before and after the introduction of open notes on November 25, 2013.</p> <p><strong>Results</strong>: After the introduction of open notes, the mean length of progress notes increased from 6,174 characters to 6,648 characters (P<0.001), and the mean character length of the "assessment and plan" (A&P) increased from 1,435 characters to 1,597 characters (P<0.001). The Average Grade Level Readability of progress notes decreased from 11.50 to 11.33, and overall readability improved by 0.17 (P=0.01). There were no statistically significant changes in the length or readability of "Initial Notes" or Letters, inter-doctor communication, nor in the modality of the recording of any kind of note.</p> <p><strong>Conclusions</strong>: After the implementation of open notes, progress notes and A&P sections became both longer and easier to read. This suggests clinician documenters may be responding to the perceived pressures of a transparent medical records environment.</p>
Multi-layout Invoice Document Dataset (MIDD)
<p>Research Purpose/Goal of Multi-Layout Invoice Document Dataset (MIDD)</p> <p>· To provide the annotated and varied invoice layout documents in IOB format to identify and extract named entities (named entity recognition) from the invoice documents to the researchers working in this domain. Obtaining a high-quality and sufficient annotated corpus for automated information extraction from unstructured documents is the biggest challenge researchers face.</p> <p>· To overcome the limitations of rule-based and template-based named entity extraction from unstructured documents traditionally used so far in information extraction approaches. Template-free processing is the only key to processing, and managing a huge pile of unstructured documents in the recent digitized era.</p> <p>· To provide varied invoice layouts so that researchers can develop a generalized AI-based model that will train on various unstructured invoice layouts. Obtained structured output can later be utilized for integrating into information management application of the organization and used for the decision-making process.</p>
Two document-concept representations of the biomedical literature
<p>These two datasets represent the biomedical literature (Medline abstracts and PubMedCentral articles) in the "document-concept matrix" format produced by <a href="https://github.com/erwanm/tdc-tools">TDC Tools</a>. These datasets can be used in downstream IR applications such as Literature-Based Discovery.</p> <p>Each of the two datasets corresponds to a specific data extraction method, see details <a href="https://erwanm.github.io/tdc-tools/input-data-format/">here</a> and in the paper linked below.</p> <ul> <li>Paper: <em>pending </em></li> <li>Code: <a href="https://github.com/erwanm/tdc-tools">https://github.com/erwanm/tdc-tools</a> <ul> <li>Documentation:<a href="https://erwanm.github.io/tdc-tools/">https://erwanm.github.io/tdc-tools/</a></li> </ul> </li> </ul> <p><strong>Important:</strong> the raw data from which this data is derived was downloaded from <a href="https://www.nlm.nih.gov/medline/medline_overview.html">Medline</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/">PubMedCentral</a> and <a href="https://www.ncbi.nlm.nih.gov/research/pubtator/">PubTatorCentral</a>, provided <a href="https://www.nlm.nih.gov/databases/download/terms_and_conditions.html">courtesy of the U.S. National Library of Medicine (NLM)</a>. The data was extracted in January 2021 and do not reflect the most current/accurate data available from NLM. See the github repository above in order to generate similar datasets from up to date data.</p>
DocMine: A Software Documentation-Related Dataset of 950 GitHub Repositories
<p>DocMine dataset consists of textual information collated from multiple software artifacts, across 950 GitHub Repositories. It also consists of probable percentage contribution of text in each software artifact towards different documentation types in each repository, accompanied by metadata information about the repository such as stargazer count, number of pull requests, commits, issues and other files analyzed.</p>
ICDAR 2021 Historical Document Classification Dataset for Task 2 - Dating
<p>Dataset for the localization classification task of the ICDAR 2021 Competition on Historical Document Classification competition. The tar.gz file contains the images and ground truth file (CSV format). The metadata csv file contains information, such as where the image comes from.</p>
Nicosia, Bedestan. Cosmati work under the dome as documented in 1980-81 with accompanying reconstruction.
<p>Nicosia, Bedestan. Cosmati work under the dome as documented in 1980-81 with accompanying reconstruction. Drawing by M. Willis and Vicki Herring.</p>
Data for Seed-driven Document Ranking for Systematic Reviews: A Reproducibility Study
<p>Data for Seed-driven Document Ranking for Systematic Reviews: A Reproducibility Study</p>
Collection of documents of the IPCC DDC at WDCC
<p>Materials about the plans of the IPCC DDC at WDCC/DKRZ for AR6 circled among the participants of the First IPCC AR6 Data Workshop (19/20 September 2017) at DKRZ in Hamburg, Germany, prior to the meeting. It includes a talk given at the IPCC Expert Meeting on the future of TGICA (01/2016 in Geneva, Switzerland) and a draft list of variables for the CMIP6 data pool compiled from the DICAD project partners' data requests and from statistics of the AR5 variable usage (status: 06/2017).</p>
Long document similarity dataset, Wikipedia excerptions for movies collections
<p>Movies-related articles extracted from Wikipedia.</p> <p>For all articles, the figures and tables have been filtered out, as well as the categories and "see also" sections.</p> <p>The article structure, and particularly the sub-titles and paragraphs are kept in these datasets</p> <p> </p> <p><strong>Movies</strong></p> <p>The Wikipedia Movies dataset consists of 100,371 articles describing various movies. Each article may consist of text passages describing the plot, cast, production, reception, soundtrack, and more.</p>
Airborne eDNA documents a diverse and ecologically complex tropical bat and other mammal community
<p><span>Environmental (e)DNA has rapidly become a powerful biomonitoring tool, particularly in aquatic ecosystems. This approach has not been as widely adopted in terrestrial communities where the methods of vertebrate eDNA collection have varied from the use of secondary collectors such as blood-feeding parasites and spider webs to washing surfaces of leaves and soil sampling. Recent studies have demonstrated the potential of direct collection of eDNA from air sampling, but none have tested how effective airborne eDNA sampling might be in a </span><span>biodiverse environment.</span> <span>We used three prototype samplers to actively sample a mixed neotropical bat community in a partially controlled environment. We assess whether airborne eDNA can accurately characterize a high-diversity community with skewed abundances and to determine if filter design impacts DNA collection and taxonomic recovery. Our study provides evidence for the accuracy of airborne eDNA as a detection tool and highlights its potential for monitoring high-density, diverse assemblages such as </span><span>bat roosts. </span><span>Analysis of air samples recovered >91% of the species present and some limited relationship between species abundance and read count. Our data suggest this method can accurately depict a diverse mixed mammal community, particularly when the location is contained (e.g., a roost, den or burrow) but also highlights the potential for secondary transfer of eDNA material on clothing and equipment. Our results also demonstrate that simple, inexpensive, battery-operated homemade air samplers can collect an abundance of eDNA from the air, opening the opportunity for sampling in remote environments. </span></p>
DOCUMENTED HUMAN OSTEOLOGICAL COLLECTIONS AS BIOBANKS: RELEVANCE FOR RARE DISEASES IDENTIFICATION IN THE PAST
<p><em><strong>Presented at: 23rd Paleopathology Association European Meeting, Vilnius, Lituânia, 25-29 Agosto. Paleopathology Association European (Vilnius, Lituânia)</strong></em></p> <p>Disease identification in paleopathology relies on the exercise of differential diagnosis, and interpretation. Only a few diseases leave macroscopic pathognomonic traits in bone, and even in cases where microscopic, biochemical and biomolecular analyses are used, diagnosis is invariably inconclusive. Additionally, bone response to a variety of etiologies tends to be homogenous, with mosaic pattern(s) of bone formation and destruction. Therefore, access to pathological cases from human remains of Documented Human Osteological Collections (DHOC) is an exceptional approach. The access to biographical data of the individuals incorporated into the DHOC includes the cause of death, ancestry, sex, age, clinical data and other information akin to clinical data allowing for the possibility of hypothesis-driven research in which bones changes correlate with causes of death - hence providing tested and informed differential diagnosis. In this sense, DHOC may be viewed as a biobank equivalent, i.e. biorepository that stores biological samples for research in the identification of bone changes related to diseases associated with clinical and personal data. This paper will explore known cases of diseases’ diagnoses, such as lepra, neoplasias, tuberculosis, syphilis, and diffuse idiopathic skeletal hyperostosis that have used DHOC as diagnostic testing grounds, to explore bone changes and methodological advancements. The paper also introduces the idea of DHOC as biobanks dedicated to the study of rare diseases, as rarely reposted diseases, in paleopathology. </p> <p><strong>Keywords: </strong>Health, biorepository, biobanks, DHOC, differential diagnosis</p>
Open-source Software Governance Documentation Dataset on GitHub
<p>This dataset contains 710 GitHub-hosted OSS projects, which contain a governance file in the root directory of the project. It also contains commits, issues, and comments on each project.</p>
Aesthetic Responses to the Covid-19 Crisis: The corona performance No Problama Video/Audio Documentation
<p>The video files document the artistic performance <em>No Problama, a coronavirus story from Homborsund </em>created during the Covid-19 crisis in Grimstad, Norway in 2021. Each video presents components of the performance which are described and analysed in the article <strong>Aesthetic responses to the Covid-19 pandemic: The Corona-performance <em>No Problama 2021</em></strong> which belongs to the Routledge Open Access Collection<em> Art and Crisis. </em>All participants have given their written consent for the publication.</p>
Mid.ru press documents 2003–2019, Russian
<p>Document collection scraped from the Russian governmental website mid.ru. Includes all items (transcripts, comments etc.) listed on:<br> - https://mid.ru/ru/press_service/minister_speeches/,<br> - https://mid.ru/ru/press_service/deputy_ministers_speeches/,<br> - https://mid.ru/ru/press_service/telefonnye-razgovory-ministra/,<br> - https://mid.ru/ru/press_service/spokesman/briefings/,<br> - https://mid.ru/ru/press_service/spokesman/official_statement/,<br> - https://mid.ru/ru/press_service/spokesman/answers/,<br> - https://mid.ru/ru/press_service/spokesman/kommentarii/,<br> and on their following pages (e.g. https://mid.ru/ru/press_service/minister_speeches/?PAGEN_1=2) from the first documents (4 January 2003) until the end of 2019.</p> <p>11,857 documents. One document in each row. Columns:</p> <p>- ID: format MID-1 <br> - ID_no: format 1 <br> - Date: Document date, format 2019-12-31 <br> - Title: Document title<br> - Type: One of the following: Брифинги; Выступления заместителей Министра; Выступления Министра; Комментарии; Ответы на вопросы СМИ; Официальные заявления; Телефонные разговоры Министра<br> - Text: Document text including title <br> - URL: URL from which the document is downloaded <br> - Downloaded: Date of download, format 2019-12-31</p> <p>Formats: rds and json.</p> <p>Version 1.1: edited column names.</p> <p>---</p> <p>From www.mid.ru:</p> <p>Materials on the website of the Russian Ministry of Foreign Affairs are generally accessible and open for non-commercial use (personal, family, education, research, etc.).<br> Their reprinting, as well as any quoting in the mass media is allowed only with a reference to the website of the Russian Ministry of Foreign Affairs as a source of the information.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.