Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,085

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,085 results for “Documentation”

Learn how ShareScore rates datasets ↗
zenodo36/100

Documentation and code for reproducing analyses presented in: vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework

<p>Documentation and code for reproducing analyses presented in: vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework. Please see the <strong>README.pdf</strong> for step-by-step instructions for reproducing the entire analysis described in the paper.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Complementary dataset of Overton metadata on citing policy-related documents for the study "From intent to impact: Investigating the effects of open sharing commitments"

<p>This document provides the underlying dataset for the bibliometric component for the 2022 study &quot;From intent to impact: Investigating the effects of open sharing commitments&quot; by Research Consulting and Science-Metrix.</p> <p>Before reproducing the study findings or re-using the underlying datasets for other purposes, please cautiously review their limitations in the study&#39;s technical annex and main report, available at: https://zenodo.org/communities/data-sharing-in-public-health-emergencies/&nbsp;</p> <p>Special thanks from the Science-Metrix / Elsevier teams to Euan Adie and Overton for this exceptional public release of Overton metadata, and for conducting extraordinary data collection to retrieve citations towards arXiv preprints.</p> <p>&nbsp;</p> <p>Scope: note that this file combines cited journal publications and preprints from the Covid19, HVRD, Zika and HVVD thematic sets.</p> <p>Data treatment: this data is intend foremost to provide manual validation or qualitative triangulation of our findings. No special efforts have been made to process&nbsp;and clean the data&nbsp;for its eventual re-use in secondary analysis or&nbsp;text mining approaches.</p> <p>Definitions used in this table:</p> <table> <tbody> <tr> <td>Column name&nbsp;</td> <td>Definition</td> </tr> <tr> <td>document_type</td> <td>preprint or journal publication</td> </tr> <tr> <td>doi</td> <td>digital object identifier</td> </tr> <tr> <td>arxiv_id</td> <td>arXiv preprint server&#39;s unique identifier for its preprints</td> </tr> <tr> <td>ssrn_id</td> <td>SSRN preprint server&#39;s unique identifier for its preprints. Note that some of these IDs are contained within the DOIs also assigned to some (but not all) SSRN preprints , in the form of &quot;10.2139/ssrn.&quot; + &#39;ssrn_id&#39;</td> </tr> <tr> <td>coalesce_id</td> <td>coalesce function applied to the DOI, arxiv_id and ssrn_id. Redundant for journal publications.</td> </tr> <tr> <td>policy_source_title</td> <td>name of the policy-related organization</td> </tr> <tr> <td>published_on</td> <td>publication date of the citing policy-related document</td> </tr> <tr> <td>title</td> <td>title of the citing policy-related document</td> </tr> <tr> <td>pdf_url</td> <td>URL for the online version of the policy-related document</td> </tr> <tr> <td>snippet</td> <td>Where available, excerpt of the text immediatly before and after the citation to a journal publication or preprint found in the citing policy-related document</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Historical Photograph and Document Collection

<p>After its establishment in 1929, TREC &ldquo;began business&rdquo; in 1930, with the construction of its first lab and office buildings and planting of the first crops. The historic photographs emphasize the first 10 years of TREC, when the most visible transition occurred from pineland to farmland, though also from the following decades (1940s-1970s), when TREC started to resemble the present campus. During this time and beyond, the nearly soiless oolitic limestone with native pine rockland transitioned into ornamental, fruit, and vegetable crops, while the pine rockland habitat was becoming vastly reduced and endangered.</p> <p>The catalog for this collection is contained as `Photo&amp;DocumentScans.xlsx` with the collection as the accompanying `Photo&amp;DocumentScansCollection.zip`.</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Dataset for an analysis of selected studies on software development and reuse in the field of software documentation

<p>The data record contains bibliographical metadata from 2012 to 2023, which have been extracted from the databases ScienceDirect, SpringerLink and IEEE Xplore. It was created in the context of a master thesis which will be published later. The aim was to make a selection of studies on software documentation from the perspective of software development based on the master thesis.</p>

opencc-zeroJul 2022View details →
zenodo36/100

Fig. 4 in The first documented record of Chvalaea Papp & Földvári, 2002 (Diptera, Hybotidae, Ocydromiinae) from the Australasian Region: a new species and its possible relationship to other members of the genus

Fig. 4. Geographical records of the Australian species of Chvalaea Papp &amp; Földvári, 2002.

opencc-by-4.0Sep 2022View details →
zenodo36/100

Citizen research landscape documentation

<p>This report reflects on existing citizen research activity in the UK heritage sector by summarising both published research and current citizen research activity at Independent Research Organisations (IROs) and other heritage organisations, addressing the following questions:<br> How do institutions and volunteers experience online citizen research; what are the motivations and rewards? Were projects successful in what they set out to do? What challenges have projects faced? What recommendations can be made based on the experiences of these projects?<br> In December 2020, we circulated a call for information from cultural heritage practitioners in the UK running citizen research projects online. This report draws on the responses generously sent in answer to our call, focusing primarily on online projects based in the UK. A full list of the featured projects can be found in the Appendix. In July 2021, members of the Engaging Crowds project team coordinated a workshop at the Discovering Collections, Discovering Communities (DCDC) conference to gather insights from practitioners and researchers with experience of enabling and supporting online volunteering. The discussions from this workshop are also summarised in this report.<br> A broad range of experience of citizen research in the UK is included in this report, from small-scale short-term projects, such as those set up to move pre-established volunteering projects online in response to the onset of the Covid-19 pandemic, to large-scale long-term citizen research projects. What is evident from all responses received to the call for information is that online citizen research projects bring many benefits to both volunteers and organisations, and also present a number of challenges. Successes relate to volunteer recruitment and engagement, as well as opportunities for increased data production and data quality. Data quality, however, was also reported by a number of respondents as an area of challenge, together with issues relating to chosen citizen research platforms and open access. In synthesising the experiences reported by responding projects and workshop participants, this report has identified a number of recommendations for future citizen research projects.</p>

opencc-by-4.0Sep 2022View details →
zenodo36/100

MEDDOPROF corpus: complete gold standard annotations for occupation detection in medical documents in Spanish

<p><strong>UPDATE 27/09/2022: </strong>A complete normalization of all mentions in the corpus to SNOMED CT has been added to the &#39;meddoprof-norm.tsv&#39; file.</p> <p><strong>Description</strong></p> <p>This repository contains the complete MEDDOPROF Gold Standard, a collection of 1,844 clinical cases in Spanish with annotations for occupations, working statuses and activities. MEDDOPROF is a Shared Task celebrated in 2021 that explores the application of natural language processing to occupational health. If you&#39;d like to learn more, please visit: <a href="https://temu.bsc.es/meddoprof">https://temu.bsc.es/meddoprof</a>.</p> <p><strong>Folder and File Structure</strong></p> <p>The corpus&#39; files are presented in the format used by the annotation tool brat. That is, for each clinical case there is a .txt file with the text and a .ann file with its corresponding annotations.</p> <p><em>- meddoprof-ner/</em></p> <p>Clinical cases annotated with these labels: PROFESION (PROFESSION), SITUACION_LABORAL (WORKING_STATUS) or ACTIVIDAD (ACTIVIDAD).</p> <p><em>- meddoprof-class/</em></p> <p>Clinical cases with the same annotations as &#39;meddoprof-ner&#39; but with these labels instead: PACIENTE (patient), FAMILIAR (family member), SANITARIO (health professional) or OTRO (other).</p> <p><em>- ner_class_joint/</em></p> <p>Clinical cases with both levels of annotation (ner and class) joint (that is, a mention classified as as PROFESOR in meddoprof-ner and as PACIENTE in meddoprof-class would be PROFESION-PACIENTE here).</p> <p><em>- meddoprof-norm.tsv</em></p> <p>Tab-separated file (.tsv) with the mapping of each mention in the corpus to ESCO and SNOMED CT. The file has five columns: filename, mention text, span, ESCO code and SNOMED code.</p> <p>Additionally, two files with the filenames of the train and test partitions are included.</p> <p>&nbsp;</p> <p><strong>Please cite if you use this resource:</strong></p> <p>Salvador Lima-L&oacute;pez, Eul&agrave;lia Farr&eacute;-Maduell, Antonio Miranda-Escalada, Vicent Briv&aacute;-Iglesias and Martin Krallinger. NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts. In Procesamiento del Lenguaje Natural, 67. 2021.</p> <pre><code>@article{meddoprof, title={NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts}, author={Lima-López, Salvador and Farré-Maduell, Eulàlia and Miranda-Escalada, Antonio and Brivá-Iglesias, Vicent and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, volume = {67}, year={2021}, issn = {1989-7553}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6393}, pages = {243--256} }</code></pre> <p><strong>Related Resources:</strong></p> <p>- <a href="http://temu.bsc.es/meddoprof">Web</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4694768">Training Data</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4889776">Test set</a></p> <p>- <a href="https://zenodo.org/record/4722741">Codes Reference List</a> (for MEDDOPROF-NORM)</p> <p>- <a href="https://zenodo.org/record/4720833">Annotation Guidelines</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4524658">Occupations Gazetteer</a></p> <p>&nbsp;</p> <blockquote> <p>MEDDOPROF is part of the IberLEF 2021 workshop, which is co-located with the SEPLN 2021 conference. For further information, please visit&nbsp;<a href="https://temu.bsc.es/meddoprof/">https://temu.bsc.es/meddoprof/</a>&nbsp;or email us at encargo-pln-life@bsc.es</p> <p>MEDDOPROF is promoted by the Plan de Impulso de las Tecnolog&iacute;as del Lenguaje de la Agenda Digital (Plan TL) and the Spanish government&#39;s 2020 Proyectos de I+D+i RTI Tipo A (AI4PROFHEALTH - DESCIFRANDO EL PAPEL DE LAS PROFESIONES EN LA SALUD DE LOS PACIENTES A TRAVES DE LA MINERIA DE TEXTOS (PID2020-119266RA-I00)).</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Datasets to accompany twitter reproducible methods document

<p>Datasets to accompany twitter reproducible methods document</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation

<p>This repository contains the following 3&nbsp;datasets for legal document summarization :</p> <p>- IN-Abs : Indian Supreme Court case documents &amp; their `abstractive&#39; summaries, obtained from http://www.liiofindia.org/in/cases/cen/INSC/<br> - IN-Ext : Indian Supreme Court case documents &amp; their `extractive&#39; summaries, written by two law experts (A1, A2).<br> - UK-Abs : United Kingdom (U.K.) Supreme Court case documents &amp; their `abstractive&#39; summaries, obtained from https://www.supremecourt.uk/decided-cases/</p> <p>Please refer to the paper and the README file for more details.</p>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Foundation document of King Rusa II

Foundation document of King Rusa II Basalt Kesis-Golu [Turkey], 7th cent. BCE Pergamon Museum, Berlin Source: Objaverse 1.0 / Sketchfab

opencc-byFeb 2020View details →
zenodo36/100

SUNFISH Platform Documentation - API Interfaces of Distributed Components

<p>This document summarizes the interfaces provided in the dataset used to generate and document APIs designed and applied in the SUNFISH components.</p>

opencc-by-4.0Dec 2017View details →
zenodo36/100

Supplementary Material for Requirements documentation containing natural language: A Systematic Tertiary Literature Review

<p>Context: Requirements documentation in natural language&nbsp;has diverse artifacts, but few studies address their suitability to types<br>of requirements or ease of communication.</p> <p>Methods: We conducted a&nbsp;systematic tertiary literature review (STLR) and identified 22 relevant&nbsp;review papers that address natural language artifacts used by practitioners to document software requirements. We also investigated which&nbsp;types of requirements are addressed by artifacts and if there are guidelines for each.</p> <p>Results: A variety of artifacts used for this purpose were&nbsp;identified, of which the most referenced in the literature were diagrams,<br>use cases, conceptual models, user stories, and prototypes. The analysis highlighted that artifacts are applied differently to functional and&nbsp;non-functional requirements. In general, diagrams, use cases, scenarios,&nbsp;and prototypes can be used for both types of requirements, depending&nbsp;on the content (usability, security, etc.). However, user stories and derived artifacts are more recommended for functional requirements and&nbsp;have limitations for non-functional requirements.</p> <p>Conclusion: Furthermore, the study explored different guidelines, structures, and formats used in documentation artifacts, reflecting the diversity in requirements documentation practices in software projects.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

Document

Description

opencc-by-4.0May 2024View details →
zenodo36/100

BBR and ACM-RT inputs for figures in ACMB-DF documenting paper (Barker et al., 2024).

<p><span>For the Earth Cloud, Aerosol, Radiation Explorer (EarthCARE) satellite mission there are a number of algorithms used to process the observations.&nbsp;&nbsp; One of these algorithms, called ACMB-DF, is designed to perform continuous radiative closure assessment of EarthCARE observations. This is done using radiance and flux measurements from the Broad-Band Radiometer (BBR) and forward solar and thermal radiative transfer calculations applied to retrieved geophysical properties. &nbsp;The ACMB-DF algorithm is documented in an Atmospheric Measurements and Techniques (AMT) article.&nbsp;</span></p> <p><span>To illustrate the methodology used for ACMB-DF, detailed calculations were performed for the &ldquo;Hawaii&rdquo; scene.<span>&nbsp; </span>These include detailed 3D Monte Carlo calculations of the BBR and Multi-Spectral Imager (MSI) radiances applied to select sections of the Hawaii test scene.<span>&nbsp; </span>These where then used in the chain of EarthCARE retrievals to produce geophysical retrievals used as input for the ACM-RT processor to compute radiative quantities needed for closure assessment and as input to generate BBR radiances and fluxes.</span></p> <p><span>The relevant publication for these calculations is:</span></p> <p><span>Barker, H. W., J. N. S. Cole, N. Villefranque, Z. Qu, A. Velazquez-Blazquez, C. Domenech, S. L. Mason, and R. J. Hogan : Radiative Closure Assessment of Retrieved Cloud and Aerosol Properties for the EarthCARE Mission: The ACMB-DF Product. Submitted to AMT, May 2024.</span></p>

opencc-by-4.0May 2024View details →
zenodo36/100

Circular economy standards for automotive electronics: from a state of the art analysis to a new pre-standardization document

<p>The Standardization Toolkit, developed by <a href="https://www.uni.com/en/" target="_blank" rel="noopener">UNI (Italian Standards Body)</a> within the TREASURE project, is an open reserach tool designed to simplify reserach of international and european standards (EN and ISO) relevant to circular economy strategies in the automotive sector.</p> <p>This dynamic tool, powered by Microsoft Power BI, provides an extensive list of national, European, and international standards, detailing key information such as standard numbers, URLs for document retrieval, titles, publication years, and scopes.</p> <p>The Toolkit maps out 73 standards, including 61 current standards and 12 works in progress, overseen by prominent technical committees such as ISO and CEN. It covers crucial areas like life-cycle assessment, substance determination, and decision support frameworks, offering a valuable resource for stakeholders aiming to implement effective and sustainable practices.</p> <p>The Toolkit is freely available at&nbsp;<a href="https://www.treasureproject.eu/standardization-toolkit/">this link</a>.</p> <p>This publication includes the dateset behind the Standardization Toolkit, while D8.4 "Standardization Toolkit" describes the previous version of the mapping and the methodology behind the state of the art analysis.</p> <p>It has also been uploaded the D8.5 "Strategic standardization roadmap" which includes a description of the Standardization Toolkit, and the new pre-standardization document (CEN Workshop Agreement - CWA) developed within Treasure project.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>Disclaimer:</strong> The standards retrieved from the Standardization Toolkit cannot be fully downloaded. The Toolkit serves to simplify the search for these standards and to provide an overview of what has been mapped during the TREASURE project&rsquo;s standardization activities.</p>

openMay 2024View details →
zenodo36/100

"Adaptive Radial Projection on Fourier Magnitude Spectrum for Document Image Skew Estimation"

<p>We build DISE2021 datasets from 95 images from DISEC2013 dataset [12], 70 images from RDCL dataset [22],&nbsp;and 324 images from RVL-CDIP dataset [14]. The composed datasets contains various types of documents, multiple languages, and typography features. Firstly, all the images are ensured and verified to be&nbsp;in a straight position. Secondly, we use the&nbsp;generating algorithm as in [12] to generate skew images in the&nbsp;range &minus;15 to +15 skew degree. The dataset is split into two development/test&nbsp;sets by a ratio of 0.7/0.3 that results in 3399 development images&nbsp;and 1491 testing images. When generating the skew dataset in the&nbsp;range from &minus;44.9 to 44.9 skew degree, we double the augmented image that&nbsp;results in 6980 development images and 2800 testing images.&nbsp;</p> <p>Note: This datasets are built upon three other datasets: DISEC 2013, RVL-CDIP, RDCL 2017. So I urge you to respect their LICENSE.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Fig. 4 in First record of Gracixalus quangi Rowley, Dau, Nguyen, Cao & Nguyen, 2011, from Hoa Binh Province, Vietnam, including the first documentation of advanced larval stages and an extended tadpole description

Fig. 4. Karst forest habitat in Hoa Binh Province. Photo credit C.T. Pham.

opencc-by-4.0Feb 2019View details →
zenodo36/100

Fig. 1 in First record of Gracixalus quangi Rowley, Dau, Nguyen, Cao & Nguyen, 2011, from Hoa Binh Province, Vietnam, including the first documentation of advanced larval stages and an extended tadpole description

Fig. 1. Adult Gracixalus quangi from Hoa Binh Province. Photography by T. Ziegler.

opencc-by-4.0Feb 2019View details →
zenodo36/100

CATCH-EyoU: Public Authorities engaging with youth: Policy Documents: UK

<p>The specific data collected are written public documents by UK governmental and civil society actors regarding youth at European, national and local levels, produced during a three-year period between 2012-2014.&nbsp;The aim of the collection of public documents is to characterize contemporary European Union policy discourses and examine how policies and pronouncements about youth are both conceptualized and enacted in the UK.&nbsp;The dataset contains 10 primary documents selected and summarized in an Excel spreadsheet.&nbsp; Full citation of data sources is provided via excel spreadsheet. The dataset consists of original source document summaries that are catalogued in an excel spreadsheet by identifying information (author/title/year) and summaries of the general character of each document.</p>

opencc-by-nc-4.0Mar 2018View details →
zenodo36/100

Contextual Documentation Referencing on Stack Overflow — Supplementary Material

<p>Supplementary material for our paper &quot;Contextual Documentation Referencing on Stack Overflow&quot;.</p>

opencc-by-sa-4.0Feb 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record