Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

227

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

227 results for “provenance”

Learn how ShareScore rates datasets ↗
edi52/100

Software Tools to Collect and Use Provenance in R

The software tools that scientists use to process and analyze data are typically optimized for performance and ease of use. Few if any such tools are designed to capture and record the details of what happens as the tool performs its task. This detailed information, and more generally the history of an item of data from its creation to its present state, is known as provenance. Provenance has the potential to make science more transparent, reliable, and reproducible. This project focused on collecting and using provenance for scripts written in the R statistical language, which is widely used by ecologists and environmental scientists for data analysis and visualization. Our tools include a provenance collector (rdtLite), which collects provenance as an R script executes (or during a console session), as well as other tools that use the collected provenance to document and visualize the execution or to support activites such as script debugging. The R packages included here are also available on CRAN. For more details, see the project website on GitHub (https://end-to-end-provenance.github.io).

openCC0Feb 2024View details →
zenodo48/100

OpenCitations Meta RDF dataset of agent roles metadata and its provenance information

<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>agent roles</strong> of bibliographic resources<strong>&nbsp;</strong>(<a href="http://purl.org/spar/pro/RoleInTime" target="_blank" rel="noopener">http://purl.org/spar/pro/RoleInTime</a>). These agents can be authors, editors, or publishers. It contains all the metadata and its provenance information, structured specifically around agent roles, in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /ar/06250/10000/1000/1000.zip, while information about provenance in /ar/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>

opencc-zeroApr 2024View details →
zenodo48/100

OpenCitations Meta RDF dataset of page numbers metadata and its provenance information

<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>page numbers</strong> of bibliographic resources, known as <strong>manifestations </strong>(<a href="http://purl.org/spar/fabio/Manifestation" target="_new">http://purl.org/spar/fabio/Manifestation</a>). It contains all the bibliographic metadata and its provenance information, structured specifically around manifestations (page numbers), in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /br/06250/10000/1000/1000.zip, while information about provenance in /br/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>

opencc-zeroApr 2024View details →
zenodo48/100

OpenCitations Meta RDF dataset of identifiers metadata and its provenance information

<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to&nbsp;<strong>identifiers </strong>(<a href="http://purl.org/spar/datacite/Identifier" target="_blank" rel="noopener">http://purl.org/spar/datacite/Identifier</a>) of bibliographic resources. It contains all the metadata and its provenance information, structured specifically around identifiers, in JSON-LD format.</p> <p>The inner folders are named through the&nbsp;<strong>supplier prefix</strong>&nbsp;of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to&nbsp;<strong>06*0</strong>).</p> <p>After that, the folders have&nbsp;<strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the&nbsp;<strong>zipped&nbsp;</strong>RDF data.</p> <p>At the same level, additional folders containing the&nbsp;<strong>provenance&nbsp;</strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called&nbsp;<strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /id/06250/10000/1000/1000.zip, while information about provenance in /id/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the&nbsp;<a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>

opencc-zeroApr 2024View details →
zenodo48/100

OpenCitations Meta RDF dataset of bibliographic resources metadata and its provenance information

<div> <p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to&nbsp;<strong>bibliographic resources&nbsp;</strong>(<a href="http://purl.org/spar/fabio/Expression" target="_blank" rel="noopener">http:///purl.org/spar/fabio/Expression</a>). It contains all the metadata and its provenance information, structured specifically around bibliographic resources, in JSON-LD format.</p> <p>The inner folders are named through the&nbsp;<strong>supplier prefix</strong>&nbsp;of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to&nbsp;<strong>06*0</strong>).</p> <p>After that, the folders have&nbsp;<strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the&nbsp;<strong>zipped&nbsp;</strong>RDF data.</p> <p>At the same level, additional folders containing the&nbsp;<strong>provenance&nbsp;</strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called&nbsp;<strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /br/06250/10000/1000/1000.zip, while information about provenance in /br/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the&nbsp;<a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p> <p>&nbsp;</p> </div>

opencc-zeroApr 2024View details →
zenodo44/100

TokTrack: A Complete Token Provenance and Change Tracking Dataset for the English Wikipedia

<p><strong>Fixes in version 1.1 (= Zenodo's "version 2")</strong></p> <p>*In 20161101-revisions-part1-12-1728.csv, missing first data line is added.</p> <p>*In Current_content and Deleted_content files, some token values ('str' column) which contain regular quotes ('"') are fixed.</p> <p>*In Current_content and Deleted_content files, some wrong revision ID values for 'origin_rev_id', 'in' and 'out' columns are fixed.</p> <p> ------</p> <p><strong>This dataset contains every instance of all tokens (≈ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article revision it was originally created in, and (ii) lists with all the revisions in which the token was ever deleted and (potentially) re-added and re-deleted from its article, enabling a complete and straightforward tracking of its history.</strong></p> <p>This data would be exceedingly hard to create by an average potential user as it is (i) very expensive to compute and as (ii) accurately tracking the history of each token in revisioned documents is a non-trivial task. <br> Adapting a state-of-the-art algorithm, we have produced a dataset that allows for a range of analyses and metrics, already popular in research and going beyond, to be generated on complete-Wikipedia scale; ensuring quality and allowing researchers to forego expensive text-comparison computation, which so far has hindered scalable usage.</p> <p>This dataset, its creation process and use cases are described in a dedicated dataset paper of the same name, published at the ICWSM 2017 conference. In this paper, we show how this data enables, on token level, computation of provenance, measuring survival of content over time, very detailed conflict metrics, and fine-grained interactions of editors like partial reverts, re-additions and other metrics.</p> <p>Tokenization used: https://gist.github.com/faflo/3f5f30b1224c38b1836d63fa05d1ac94</p> <p>Toy example for how the token metadata is generated: <br> https://gist.github.com/faflo/8bd212e81e594676f8d002b175b79de8</p> <p><strong>Be sure to read the ReadMe.txt or - even more detailed - the supporting paper which is referenced under "related identifiers".</strong></p>

opencc-by-sa-4.0Mar 2017View details →
zenodo44/100

Data to Support Predictive Models for Detrital Titanite Provenance with application to the Nanga Parbat syntaxial massif, western Himalaya."

<p>The files published here are metadata that are being used to support a manuscript currently (Mar, 2024) undergoing final reviews in Journal of Geophysical Research: Earth Surface.</p> <p>The intention of these data and code is to support a publication that is about generating a predictive categorisation scheme for the mineral titanite.</p> <p>The code to generate the titanite classification schemes was created in Python3, using Jupyter Notebook. The files also provide more motivation for why a predictive categorisation scheme for the mineral titanite is desirable, and other similar context. Chiefly, the dataset and random forest models published here will allow us to trace titanite in detritus.</p> <p>For info on running Jupyter Notebook, please visit (<a href="https://jupyter-notebook-beginner-guide.readthedocs.io/en/latest/execute.html">https://jupyter-notebook-beginner-guide.readthedocs.io/en/latest/execute.html</a>) to seek instructions. We also provide a readme file with some instructions. If you get really stuck, just email the authors.</p> <p>Our Model can be compared to similar previously published works (e.g.&nbsp;<a href="https://doi.org/10.1111/ter.12574">https://doi.org/10.1111/ter.12574</a>). Model was trained using skikit-learn v1.41.</p> <p>The supplementary file "Table_S4_Merged.csv" was used to train and generate the model.</p> <p>Your unknowns must contain the correct elements and labelling for the code to successfully run, these details are provided in the code (Titanite_Random_Forest_Model1_Mar24.ipynb). A template is also provided for you to paste your unknown data into (titanite_data_template.csv)</p> <p>Any new published data are titanite compositional or isotopic data collected by LA-ICP-MS. Description of how those data were collected is given in "OSullivan_et_al_Supp..." file.</p> <p>Some of the data, information and code in this submission has been subject to change after journal review, this is a second version of this content.</p> <p>References for the dataset compilation are provided in File S3.</p> <p>If you have any queries contact:<br>Gary O'Sullivan, Trinity College Dublin</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Dataset for manuscript Tracing Quartz Provenance: A Multi-Disciplinary Investigation of Luminescence Sensitisation Mechanisms of Quartz from Granite Source Rocks and Derived Sediments

<p><span>Quartz optically stimulated luminescence (OSL) sensitivity as well as some electron spin resonance (ESR) and cathodoluminescence (CL) signals have been empirically proposed as provenance indicators. Sensitivity is defined as luminescence emitted in response to a given dose per unit mass. While it is largely believed to be acquired by earth surface processes, recent studies bring evidence that sensitisation processes depend on source geology.</span></p> <p><span>Here we combine OSL and thermoluminescence (TL), ESR and CL analyses to understand the mechanisms of quartz OSL sensitisation. We investigate granites and their derived sediments from catchments draining simple lithologies of known age that display contrasting OSL sensitisation behaviour both in nature and during irradiation and light exposure laboratory experiments. The sample displaying increased OSL sensitisation is characterised by TL emission at intermediate temperatures (150-250 &deg;C), Ti-related signals in CL, and Ti and Ge lithium compensated signals in ESR. <span>The insensitive samples either lack or exhibit very weak such characteristics and contain several times less amount of trace titanium measured by </span></span><span>laser ablation inductively coupled plasma mass spectrometry (</span><span>LA-ICP-MS).</span></p> <p><span>We demonstrate that the OSL sensitisation results as an effect of the existence of certain defects and impurities in the quartz crystal in the parent rock, such as titanium and germanium. However, the degree of sensitisation reached in nature is significantly higher than in the laboratory. <span>&nbsp;</span>As such, the existence of this precursor represents the potential for sensitisation, which can later be amplified by environmental factors during sedimentary history.</span></p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

eDNAbyss samples provenance and environmental context - version 1

<p>This is a tabular&nbsp;version of the eDNAbyss sample provenance &amp; environmental context metadata submitted to BioSamples (<a href="https://www.ebi.ac.uk/biosamples/samples?text=eDNAbyss">https://www.ebi.ac.uk/biosamples/samples?text=eDNAbyss</a>).</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Reliquary of contacts for: A pragmatic approach to complex citations, closing the provenance gap between IPCC AR6 figures and CMIP6 simulations

<p>Photos and metadadata pannels of a "Reliquary of contacts for: A pragmatic approach to complex citations, closing the provenance gap between IPCC AR6 figures and CMIP6 simulations" produced to support the "A pragmatic approach to complex citations, closing the provenance gap between IPCC AR6 figures and CMIP6 simulations" presentation given at EGU 2024.</p> <p>------</p> <p>With ever growing abilities to process greater volumes of data the abiiity to sustain the citability and tracability of the underluing source data within outputs such as publications is becoming increasingly challenging. With a range of use-cases, work on how to handle complex citations from the perspective of those producing outputs, journals and those handling the knowledge graph and associated services, is exmaning a how to handle these situations in a sustainable and manageable fashion.<br><br>At the European Geophysical Union (EGU) General Assembly in Vienna, 2024, a pragmatic solution using Zenodo to store 'reliquary' objects was presented. The poster presentation demonstrated the use of existing strucutres within a Zenodo object to address the complex citation use-case around figure, the related data and the source datasets related to the IPCC's AR5 figure data. I.e. how to utulise the existing constructs of a Zenodo item and the range of available metadata fields to give an off-the-shelf solution to allow tracability to the specific datasets used (via their Handle identifiers) and citability of the higher level, DOI-ed dataset collections within which the specific Handle-ed datasets were selected from. Additionally, the connectivity between these two levels of PID objects was also captured within the stored files around which the rich metata was captured.<br><br>The concept of a complex citation 'reliquary' as a metadtata rich object, acting as a referencable nexus in the knowledge graph has been put forth as a solution to the complex citation challenge. It borrows the concept from its historical use, denoting a container or shrine, often richly embellished, for sacred relics (e.g. saints bones, artefacts etc). In the same way here we have both the rich metadata 'container' around the specific details (the 'bones in the box', with their preserved connectivity).<br><br>However, the term 'reliquary' is often a hard one to convey, being somewhat of an obscure term (likewise the term 'nexus' may also be one lacking wider recogniton). Thus, to aid the discussions around the presentation by Pascoe et al. (2024) at the EGU 2023 General Assembly, a physical representation of a metadata reliquary object was produced.<br><br>The purpose of this object was two fold:<br><br>&nbsp;- The first was to show how the reliquary container itself is metadata rich, detailing through the use of ORCIDS, RORs and a DOI, references to external items, complemented by further metadata concerning the specifics of the reliquary's own metadata (its title and the credit for the artist that created it). Futher more, the relationship between the reliquary and those referenced parties/objects was also captured. The contents were also used to demonstrate the importance of making the contents useful for onward users (in this case contact details on business cards). <br>&nbsp;- The second, and for the funder of this piece, arguably the most important aspect was a degree of outreach this provided, both to engage the audience of Pascoe et al (2024), and directly to the artist to demonstrate the importance of this work to the international research data management community and overall to aid engagemeng with the funder's work.<br><br>This resource is provided here as a repository of images of the reliquary itself and in context at the EGU 2024 event as a potential resource others may use to aid further discussions around the use of reliquaries with regards to complex citations. The slides provided of the reliquary box labels are also provided with some annotation to further expand on the metadata aspects of their content.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Data from: Possible provenance of IRD by tracing late Eocene Antarctic iceberg melting using a high-resolution ocean model

<p>This repository contains the data supplemented to&nbsp;<a href="https://doi.org/10.5194/cp-21-441-2025">Elbertsen et al. (2025)</a>&nbsp;based on Mark Elbertsen's MSc project in which he performed depth-integrated Lagrangian iceberg tracing around Antarctica during the late Eocene using high-resolution ocean model data. Using the OceanParcels framework, iceberg melting (or growth) was simulated using several kernels, including for the dominant iceberg melt terms: basal melt, buoyant convection and wave erosion. By defining kernels for five different order-of-magnitude iceberg size classes, the model was be used to determine the minimum iceberg size required for icebergs to survive the late Eocene warmth. The model output of these simulations can be found here.</p> <p>&nbsp;</p> <p>This research is funded by ERC Starting Grant 802835 (OceaNice) to Peter K. Bijl.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Bibliographic dataset based on Scientometrics, containing provenance information compliant with the OpenCitations Data Model and non disambigued authors

<p>The dataset contains bibliographical information about scholarly works in the journal Scientometrics only if the DOI is known.&nbsp;The data was extracted via Crossref.&nbsp;It is a temporal dataset in which provenance information and change-tracking have been managed by adopting the OpenCitations Data Model. Moreover, the dataset contains information on all the cited academic works.&nbsp;Journals and bibliographic resources always appear unambiguously, without duplicates.&nbsp;On the contrary, the authors have not been disambigued. Finally, heuristics have been applied to recover the DOI of the cited works in case Crossref did not provide such information.</p>

opencc-zeroJul 2021View details →
zenodo44/100

Datasets Cured and Enriched with Provenance from the National Vaccination Campaign Against COVID-19

<p>The COVID-19 pandemic is a global threat. If, on the one hand, weaccount for many losses, on the other hand, the generation of datasets and ur-gent analytical demands has accelerated. Among the combat strategies, vacci-nation and data-centered epidemiological investigations stand out. This datasetpaper presents the process of building cured and annotated datasets with prove-nance metadata. The main dataset is based on the registration data of the Vacci-nation Campaign against COVID-19 in Brazil. The dataset contains thousandsof records processed up to March 2021. The data were analyzed, investigated,treated and cross-checked with other sources, in order to correct and comple-ment them, resulting in cured datasets and aligned to the FAIR principles.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Bibliographic dataset based on Scientometrics, including provenance information compliant with the OpenCitations Data Model

<p>The dataset contains bibliographical information about scholarly works in the journal Scientometrics only if the DOI is known.&nbsp;The data was extracted via Crossref.&nbsp;It is a temporal dataset in which provenance information and change-tracking have been managed by adopting the OpenCitations Data Model. Moreover, the dataset contains information on all the cited academic works.&nbsp;Journals,&nbsp;bibliographic resources, and authors always appear unambiguously, without duplicates. Finally, heuristics have been applied to recover the DOI of the cited works in case Crossref did not provide such information.</p> <p>The dataset is distributed as two journal files, one for the data and one for the provenance, readable via the triplestore Blazegraph. There are 4,960,087 data triples and 19,348,027 provenance triples, which corresponds to 1,134,545 entities and 2,696,689 snapshots. Therefore, on average, each entity has two snapshots. Among the data, there are 231,217 agent roles, 221,602 responsible agents, 206,003 bibliographic resources, 142,472 citations, 141,555 bibliographical references, 108,112 identifiers, and 83,584 resource embodiments.</p> <p>The code to generate and modify such collections is available at&nbsp;<a href="https://doi.org/10.5281/zenodo.5579754">https://doi.org/10.5281/zenodo.5579754</a>.&nbsp;&nbsp;</p>

opencc-zeroOct 2021View details →
zenodo44/100

Flowchart on the methodology used to conduct the literature search on provenance representation models and change-tracking in RDF

<p>The image contains a flowchart on the methodology used to conduct the literature search on provenance representation models and change-tracking in RDF. This methodology was used in the Abstract submitted to the DH2023 conference in Graz.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Reconstructing dust provenance from quartz optically stimulated luminescence (OSL) and electron spin resonance (ESR) signals: Preliminary results on loess from around the world

<p>Dataset for publication</p> <p><strong>Reconstructing dust provenance from quartz optically stimulated luminescence (OSL) and electron spin resonance (ESR) signals: </strong></p> <p><strong>Preliminary results on loess from around the world</strong></p> <p>&nbsp;</p> <p>Quantitative provenance analysis studies are instrumental in understanding the tectonic and climatic processes that shape the earth&rsquo;s landscape. Although the most abundant mineral in the sedimentary system is quartz, almost all studies in provenance analysis investigate accessory minerals. Quartz crystals contain a vast number of point defects, intrinsic or due to impurities. For a signal to be an accurate indicator of provenance one needs to show that it is either dose independent or reaches a quantifiable steady state characteristic of the source rock. For signals used by trapped charge dating methods (optically stimulated luminescence (OSL) and electron spin resonance (ESR)), the latter option is the feasible one. By using quartz samples collected from the Chinese Loess Plateau (Luochuan loess-paleosol section), we show that the laboratory and natural dose response curves of E`<sub>1</sub> and and peroxy electron spin resonance signals of quartz (as defined later) overlap and reach a steady state for doses over about 1000 Gy. For E&rsquo;<sub>1</sub> signals we attribute this steady state to reaching an equilibrium state between diamagnetic oxygen vacancies (the oxygen deficiency centre (ODC), Si=Si<em>)</em> and paramagnetic oxygen vacancies (E&rsquo;<sub>1</sub>). For sedimentary quartz irradiated naturally or artificially in this dose range we show a strong linear relationship with zero intercept between E&rsquo;<sub>1</sub> and peroxy signals for samples worldwide, supporting the hypothesis that these defects are Frenkel pairs. Further, we show significant correlations between the optically stimulated (OSL) sensitivity and the above two mentioned ESR signals. The very strong correlations (Pearson`s r ˃0.9) between E&rsquo;<sub>1</sub>, peroxy and OSL sensitivity remain valid after the samples have been heated for 15 min to 350 ˚C for E&rsquo;<sub>1</sub> to reach its maximum value, believed to be a result of the conversion of diamagnetic oxygen vacancies to E&rsquo;<sub>1</sub>, clearly suggesting a relationship between OSL sensitivity and oxygen vacancies in general. Samples collected from different loess sites around the world can be distinguished based on both these OSL and ESR properties. An empirical increase in OSL sensitivity as well as oxygen related defect concentrations is observed in areas where the source material has components with older detrital zircon U-Pb ages, inferring a positive correlation between OSL sensitivity, as well as the signal intensity for E<sub>1</sub>` and peroxy defects and the age of the source rocks.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Packing provenance using CPM RO-Crate profile

<p>This dataset is an <a href="https://www.researchobject.org/ro-crate/">RO-Crate</a> that bundles artifacts of an AI-based computational pipeline execution. It is an example of application of the <a href="https://by-covid.github.io/cpm-ro-crate/0.1/">CPM RO-Crate profile</a>,&nbsp; which integrates&nbsp; the <a href="https://doi.org/10.1038/s41597-022-01537-6">Common Provenance Model</a> (CPM), and the <a href="https://www.researchobject.org/workflow-run-crate/profiles/process_run_crate">Process Run Crate profile</a>.</p> <p>As the CPM is a groundwork for the<em> ISO 23494 Biotechnology &mdash; Provenance information model for biological material and data</em> provenance standards&nbsp;series development, the resulting profile and the example is intended to be presented at one of the ISO TC275 WG5 regular meetings, and will become an input for the<em> ISO 23494-5 Biotechnology &mdash; Provenance information model for biological material and data &mdash; Part 5: Provenance of Data Processing</em> standard development.</p> <p><strong>Description of the AI pipeline</strong></p> <p>The goal of the AI pipeline whose execution is described in the dataset is to train an AI model to detect the presence of carcinoma cells in high resolution human prostate images. The pipeline is implemented as a set of python scripts that work over a filesystem, where the datasets, intermediate results, configurations, logs, and other artifacts are stored. In particular, the AI pipeline consists of the following three general parts:</p> <ul> <li> <p><strong>Image data preprocessing</strong>. Goal of this step is to prepare the input dataset &ndash; whole slide images (WSIs) and their annotations &ndash; for the AI model. As the model is not able to process the entire high resolution images, the preprocessing step of the pipeline splits the WSIs into groups (training and testing). Furthermore, each WSI is broken down into smaller overlapping parts called patches. The background patches are filtered out and the remaining tissue patches are labeled according to the provided pathologists&rsquo; annotations.</p> </li> <li> <p><strong>AI model training</strong>. Goal of this step is to train the AI model using the training dataset generated in the previous step of the pipeline. Result of this step is a trained AI model.</p> </li> <li> <p><strong>AI model evaluation</strong>. Goal of this step is to evaluate the trained model performance on a dataset which was not provided to the model during the training. Results of this step are statistics describing the AI model performance.</p> </li> </ul> <p>In addition to the above, execution of the steps results in generation of log files. The log files contain detailed traces of the AI pipeline execution, such as file paths, model weight parameters, timestamps, etc. As suggested by the CPM, the logfiles and additional metadata&nbsp;present on the filesystem are then used by a provenance generation step that transforms available information into the CPM compliant data structures, and serializes them into files.&nbsp;</p> <p>Finally, all these artifacts are packed together in an RO-Crate.</p> <p>For the purpose of the example, we have included only a small fragment of the input image dataset in the resulting crate, as this has no effect on how the Process Run Crate and CPM RO-Crate profiles are applied to the use case. In real world execution, the input dataset would consist of terabytes of data. In this example, we have selected a representative image for each of the input dataset parts. As a result, the only difference between the real world application and this example would be that the resulting real world crate would contain more input files.&nbsp;</p> <p><strong>Description of the RO-Crate</strong></p> <p><strong>Process Run Crate related aspects</strong></p> <p>The Process Run Crate profile can be used to pack artifacts of a computational workflow of which individual steps are not controlled centrally. Since the pipeline presented in this example consists of steps that are executed individually, and that the pipeline execution is not managed centrally by a workflow engine, the process run crate can be applied.&nbsp;</p> <p>Each of the computational steps is expressed within the crate&rsquo;s ro-crate-metadata.json file as a pair of elements: 1) SW used to create files; 2) specific execution of that SW. In particular, we use the SoftwareSourceCode type to indicate the executed python scripts and the CreateAction type to indicate actual executions.&nbsp;</p> <p>As a result, the crate consists the seven following &ldquo;executables&rdquo;:</p> <ul> <li> <p>Three python scripts, each corresponding to a part of the pipeline: preprocessing, training, and evaluation.</p> </li> <li> <p>Four provenance generation scripts, three of which implement the transformation of the proprietary log files generated by the AI pipeline scripts into CPM compliant provenance files. The fourth one is a meta provenance generation script.</p> </li> </ul> <p>For each of the executables, their execution is expressed in the resulting ro-crate-metadata.json using the CreateAction type. As a result, seven create-actions are present in the resulting crate.</p> <p>Input dataset, intermediate results, configuration files and resulting provenance files are expressed according to the underlying RO Crate specification.</p> <p><strong>CPM RO-Crate related aspects</strong></p> <p>The main purpose of the CPM RO-Crate profile is to enable identification of the CPM compliant provenance files within a crate. To achieve this, the CPM RO-Crate profile specification prescribes specific file types for such files: CPMProvenanceFile, and CPMMetaProvenanceFile.</p> <p>In this case, the RO Crate contains three CPM Compliant files, each documenting a step of the pipeline, and a single meta-provenance file. These files are generated as a result of the three provenance generation scripts that use available log files and additional information to generate the CPM compliant files. In terms of the CPM, the provenance generation scripts are implementing the concept of provenance finalization event. The three provenance generation scripts are assigned SoftwareSourceCode type, and have corresponding executions expressed in the crate using the CreateAction type.</p> <p><strong>Remarks</strong></p> <p>The resulting RO Crate packs artifacts of an execution of the AI pipeline. The scripts that implement individual steps of the pipeline and provenance generation are not included in the crate directly. The implementation scripts are hosted on github and just referenced from the crate&rsquo;s ro-crate-metadata.json file to their remote location.</p> <p>The input image files included in this RO-Crate are coming from the <a href="http://gigadb.org/dataset/100439">Camelyon16 dataset</a>.</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

FIG. 4 in The Paris Bloubok (Hippotragus leucophaeus (Pallas, 1766) [Bovidae]) and its provenance

FIG. 4. — Gordon illustration of the Bloubok. Courtesy of the Rijksmuseum. https://www.robertjacobgordon.nl/drawings/rp-t-1914-17-159

opencc-zeroFeb 2020View details →
zenodo40/100

FIG. 3 in The Paris Bloubok (Hippotragus leucophaeus (Pallas, 1766) [Bovidae]) and its provenance

FIG. 3. — Gordon illustration of a Roan antelope skin sent to The Hague in 1779. Courtesy of the Rijksmuseum. https://www.robertjacobgordon.nl/drawings/rp-t-1914-17-161

opencc-zeroFeb 2020View details →
zenodo40/100

FIG. 1 in The Paris Bloubok (Hippotragus leucophaeus (Pallas, 1766) [Bovidae]) and its provenance

FIG. 1. — Photograph of the Paris specimen of Hippotragus leucophaeus (Pallas, 1766). Photo: © MNHN - C. Lemzaouda / P. Lafaite.

opencc-zeroFeb 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record