Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
69
datasets available to search
ShareScore release 0.7.1
Dataset results
69 results for “traceability”
RESICE - Reusability-targeted Enriched Sea Ice Core Database - Interactive data2source Traceability
<div> <div> <p>The <em>interactive_data2source_traceability.svg</em> file illustrates data and metadata availability as well as their relation to their original sources for 287 sea ice cores that are part of the <a title="RESICE" href="https://doi.org/10.5281/zenodo.10866346" target="_blank" rel="noopener">Reusability-targeted Enriched Sea Ice Core Database (RESICE)</a> as it is published in Zenodo.</p> <p>(a) shows the availability of selected data and metadata features (x-axis) for each sea ice core as indexed on the y-axis. The color map indicates the type of availability. Primary availability refers to data or metadata extracted directly from the data set. Secondary availability refers to data or metadata are extracted from articles or reports that provide supplementary information about the data sets. Tertiary availability means that additional resources unrelated to the data sets, such as instrument manuals, were used to obtain the metadata of interest.</p> <p>(b) allows to trace the data and metadata collected per RESICE sea ice core back to their original data sources. Gray dots indicate the sea ice core - again as indexed on the y-axis. Green, orange and pink dots represent primary sources (data sets), secondary sources (articles and expedition reports) and tertiary sources (instrument manuals). <br>Each of the dots is interactive, so you can click on it, and it will link you to the original doi or url of the source or to the YAML file in the RESICE living database provided on gitLab.</p> <p>The code used to generate this figure is available in the <a title="pyresice " href="https://doi.org/10.5281/zenodo.11198658" target="_blank" rel="noopener">pyresice</a> Python package available on <a title="https://git.rwth-aachen.de/mbd/pyresice" href="https://git.rwth-aachen.de/mbd/pyresice">GitLab</a>.</p> </div> </div>
Advances in traceable calibration of cylinder pressure transducers
<p>"Pressure signals generated with VTT MIKES primary standard.xlsx". Dynamic pressure signals of VTT MIKES primary standard. </p> <p>"Calibration results of dynamic pressure sensors.xlsx". Calibration results for a commercial sensor and VTT cylinder pressure sensor (CPS) at different temperatures.</p>
Figure data for the manuscript "SI-traceable frequency dissemination at 1572.06 nm in a stabilized fiber network with ring topology"
<p>This file contains the data shown in Fig. 1, Fig 3, Fig. 4, Fig. 5 and Fig. 6 of the manuscript "SI-traceable frequency dissemination at 1572.06 nm in a stabilized fiber network with ring topology". Additional information on the data and processing procedure are available from the author upon reasonable request.</p>
TCTracer: Establishing Test-to-Code Traceability Links Using Dynamic and Static Techniques - Evaluation Data - Empirical Software Engineering 2021
<p>This repository provides the data artefacts for the experiments conducted using our tool TCTracer for the journal paper "TCTracer: Establishing Test-to-Code Traceability links Using Dynamic and Static Techniques" as submitted to the Empirical Software Engineering journal in 2021.</p>
Data and code for Renier et al. "Transparency, Traceability And Deforestation In The Ivorian Cocoa Supply Chain"
<p>Data and code required to reproduce figures and stats quoted in Renier C, Vandromme M, Meyfroidt P, Ribeiro V, Kalischek N, zu Ermgassen E K H J 2023. Transparency, traceability and deforestation in the Ivorian cocoa supply chain. Envir. Res. Let. doi:10.1088/1748-9326/acad8e</p> <p>These data refer to SEI-PCS Côte d’Ivoire Cocoa v1.1.0</p> <p>Find assets required for the GEE scripts at : https://doi.org/10.5281/zenodo.7586257</p>
Assets for Renier et al. "Transparency, Traceability And Deforestation In The Ivorian Cocoa Supply Chain"
<p>Assets to run the Google Earth Engine scripts provided here: https://doi.org/10.5281/zenodo.7503845</p> <p>Required to reproduce figures and stats quoted in Renier C, Vandromme M, Meyfroidt P, Ribeiro V, Kalischek N, zu Ermgassen E K H J 2023. Transparency, traceability and deforestation in the Ivorian cocoa supply chain. Envir. Res. Let. doi:10.1088/1748-9326/acad8e</p>
Dataset for Requirements Classification in Traceability Link Recovery Datasets
<p>The dataset contains a gold standard for classifying parts of requirements in five traceability link recovery benchmark datasets.</p> <p><strong>Classification</strong></p> <ul> <li>For aspect classification: <ul> <li>functional aspects (<strong>F</strong>)</li> <li>quality aspects (<strong>Q</strong>)</li> </ul> </li> <li>For concerns in functional requirements (c.f. <a href="https://doi.org/10.1109/RE48521.2020.00028">NoRBERT publication</a>): <ul> <li><strong>Function</strong>: A function that a system shall perform</li> <li><strong>Behavior</strong>: Behavior, the system displays or reactions that are triggered by one or more stimuli</li> <li><strong>Data</strong>: A data item or data structure that shall be part of a system's state</li> <li><strong>UserRelated</strong>: Behavior of the user or functionality of the system attributable to the user</li> </ul> </li> </ul> <p><strong>Datasets</strong></p> <p>The dataset comprises preprocessed requirements of the eTour, iTrust, SMOS, eAnci and LibEST datasets. As SMOS and eAnci's original requirements were written in Italian, the dataset comprises automatically translated versions of the requirements to English. The datasets were retrieved from the <a href="http://coest.org/">website</a> of the Center of Excellence for Software & Systems Traceability (CoEST). Attribution for the datasets:</p> <p>The original eTour dataset was provided for the TEFSE challenge at 6th International Workshop on Traceability in Emerging Forms of Software Engineering (TEFSE), 2011 and was retrieved from <a href="http://coest.org/">http://coest.org/</a></p> <p>The iTrust dataset was retrieved from <a href="http://coest.org/">http://coest.org/</a></p> <p>The original SMOS and eAnci datasets can be attributed to Gethers et al., On integrating orthogonal information retrieval methods to improve traceability recovery. In 2011 27th IEEE International Conference on Software Maintenance (ICSM), Sep. 2011 and were retrieved from <a href="http://coest.org/">http://coest.org/</a> </p> <p>The LibEST dataset can be attributed to Moran et al., Improving the Effectiveness of Traceability Link Recovery using Hierarchical Bayesian Networks. In 2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE), May 2020 and was retrieved from <a href="https://gitlab.com/SEMERU-Code-Public/Data/icse20-comet-data-replication-package">https://gitlab.com/SEMERU-Code-Public/Data/icse20-comet-data-replication-package</a></p> <p> </p>
Extracting the Architecture of Microservices: An Approach for Explainability and Traceability
<p>Dataset and replication package</p>
Early Requirements Traceability with Domain-Specific Taxonomies - A Pilot Experiment
<p>Experiment instrumentation (code), raw data, and R analysis scripts.</p> <p>Github repository: https://github.com/munterkalmsteiner/inception/tree/CoClassRecommender</p>
Cutting through the Jungle: Disambiguating Model-based Traceability Terminology — Terminology Mapping Table
<p>This dataset contains the mapping of model-based traceability terminology developed in "Cutting through the Jungle: Disambiguating Model-based Traceability Terminology" to the secondary and primary sources that have been used as part of the validation of the terminology.</p>
An Approach for Traceability Recovery between Bug Reports and Test Cases
<p>Scripts and data sets used in research for <em>An Approach for Traceability Recovery between Bug Reports and Test Cases</em>.</p> <p><em><strong>(Context)</strong></em> Automatic traceability recovery between software artifacts may promote early detection of issues.<br> Information Retrieval (IR) techniques have been proposed for the task, but they differ considerably in terms of input parameters and results. It is difficult to assess results when those techniques are applied in isolation, usually in small or medium-sized software projects. Also, an overview would be more comprehensive if a Deep Learning (DL) based technique is applied, in comparison with traditional IR techniques.<br> <em><strong>(Objective)</strong></em> We propose an approach to recover traceability links between bug reports and test cases, which can be instantiated with a set of IR and DL techniques.<br> <em><strong>(Method)</strong></em> For applying and evaluating our solution, we used historical data from the Mozilla Firefox quality assurance (QA) team, on which we assessed the following IR techniques: LSI, LDA, and BM25. We also experimented with a DL architecture called Convolutional Neural Networks (CNNs) through the use of Word Embeddings.<br> <em><strong>(Results)</strong></em> In this context of traceability, we noticed poor performances from three out of the four studied techniques. Only the LSI technique was effective, even standing out over the state-of-the-art BM25 technique.<br> <em><strong>(Conclusions)</strong></em> The obtained results suggest that the semi-automatic application of the LSI technique -- with an appropriate combination of thresholds -- is feasible for real-world software projects.</p>
Implementing Traceability Repositories as Graph Databases for Software Quality Improvement: Datasets used to test our methodology that is presented in the paper 10.1109/QRS.2018.00040
<p>The first dataset is the Event Based Traceability for Managing Evolutionary Change (EBT), it is a public dataset provided by CoEST, the original artifacts and trace links are represented in XML and text format. From the EBT dataset, we selected the 41<em> requirements </em>and 25<em> test case </em>artifacts, in addition to the answer set of 51 trace links which relates the <em>requirements </em>with the<em> test case. </em>Artifacts and trace links are prepared in XML format<strong>. </strong>The data set contains XML for each artifact such as RQ.xml, EBTrelations.xml is the answer set file, TradModel.xml which describes the defined model and TradTraceabilityRule.xml that includes the rules applied for trace link types.</p> <p> </p> <p>The second dataset AgileOERP is collected from commercial management tool to customize an open source ERP applying agile methodology. It contains 350<em> user stories (US), </em>1323<em> tasks (TS) </em>and 198<em> developer test (DT) </em>artifacts, in addition to answer set of trace links that manually generated by developers which relates the <em>user story </em>artifact with<em> task (</em>1304) artifact, as such relates the <em>task </em>artifact with<em> developer test </em>artifact (65). Artifacts and trace links are prepared in XML format<strong>. </strong>The data set contains XML for each artifact such as US.xml, ERPrelations.xml is the answer set file, AgileModel.xml which describes the defined model and AgileTraceabilityRule,xml that includes all rules applied for trace links type</p> <p> </p> <p>The original dataset of the last dataset is the Aqualush irrigation system which is used as a case study in “C. Fox, Introduction to Software Engineering Design: Processes, Principles and Patterns with UML2. Addison-Wesley, 2006”. The trace links are generated and provided in “E. Ben Charrada, D. Caspar, C. Jeanneret, and M. Glinz, towards a benchmark for traceability, in Joint EVOL and IWPSE 2011, pp. 21-30”, in HTML format. For our work, we selected the <em>software requirements specification (</em>396 SRS), <em>user level requirements (</em>48 ULR), <em>use case (</em>74 UC), <em>detailed design (85 DD) </em>and <em>software architecture(15 SArch) </em>artifacts in addition to the answer set of trace links that relate the SRS with other artifacts(4038) and thus relates the DD artifact with other artifacts (1719) . Artifacts and trace links are prepared in XML<strong>. </strong>The data set contains XML for each artifact such as SRS.xml, AqualushRelations.xml is the answer set file, TradModel.xml which describes the defined model and TradTraceabilityRule.xml that includes the rules applied for trace link types.</p>
Monitoring and traceability of genetically modified soybean event GTS 40-3-2 during soybean protein concentrate and isolate preparation
To evaluate DNA fragmentation and GMO quantification during soybean protein concentrate and isolate preparation, genetically modified soybean event GTS 40-3-2 (RRS) was blended with conventional soybeans at mass percentages of 0.9%, 2%, 3%, 5%, and 10%. Qualitative PCR and real-time PCR were used to monitor the taxon-specific lectin and exogenous cp4 epsps target levels in all of the main products and by-products, which has practical significance for RRS labelling threshold and traceability. Along the preparation chain, the majority of DNA was distributed in main products, and the DNA degradation was noticed. From a holistic perspective, the lectin target degraded more than cp4 epsps target during both of the two soybean proteins preparations. Therefore, the transgenic contents in the final protein products were higher than the actual mass percentages of RRS in raw materials. Our results are beneficial to the improvement of GMO labelling legislation and the protection of consumer rights.
Complete set of raw and processed datasets, as well as associated Jupyter notebooks for analysis, associated with manuscript entitled: "The MOUSE project: a practical approach for obtaining traceable, wide-range X-ray scattering information"
<p>This dataset is a complete set of raw, processed and analyzed data, complete with Jupiter notebooks, associated with the manuscript mentioned in the title. </p> <p>In the manuscript, we provide a ``systems architecture''-like overview and detailed discussions of the methodological and instrumental components that, together, comprise the "MOUSE" project (<strong>M</strong>ethodology <strong>O</strong>ptimization for <strong>U</strong>ltrafine <strong>S</strong>tructure <strong>E</strong>xploration). Through this project, we aim to provide a comprehensive methodology for obtaining the highest quality X-ray scattering information (at small and wide angles) from measurements on materials science samples. </p>
Dataset for "Extracting Traceability Links between Javadoc Sentences and JUnit Test Code Lines"
<p>Dataset for "Extracting Traceability Links between Javadoc Sentences and JUnit Test Code Lines"</p>
Data from: Outlier SNPs enable food traceability of the southern rock lobster, Jasus edwardsii
Recent advances in next-generation sequencing have enhanced the resolution of population genetic studies of non-model organisms through increased marker generation and sample throughput. Using double digest restriction site-associated DNA sequencing (ddRADseq), we investigated the population structure of the commercially important southern rock lobster, Jasus edwardsii, in Australia and New Zealand with the aim of identifying a panel of SNP markers that could be used to trace country of origin. Four ddRADseq libraries comprising a total of 88 individuals were sequenced on the Illumina MiSeq platform, and demultiplexed reads were used to create a reference catalog of loci. Individual reads were then mapped to the reference catalog, and variant calling was performed. We have characterized two single-nucleotide polymorphism (SNP) panels comprised in total of 656 SNPs. The first panel contained 535 neutral SNPs and the second, 121 outlier SNPs that were characteristic of being putatively under selection. Both neutral and outlier SNP panels showed significant differentiation between the two countries, with the outlier loci demonstrating much larger FST values (FST outlier SNP panel = 0.134, P < 0.0001; FST neutral SNP panel = 0.022, P < 0.0001). Assignment tests performed with the outlier SNP panel allocated 100 % of the individuals to country of origin, demonstrating the usefulness of these markers for food traceability of J. edwardsii.
The traceability of waste compounds from the orange juice industry by NMR and GC-MS techniques
<p>The traceability of waste compounds from the orange juice industry by NMR and GC-MS techniques</p>
Assessing Word Similarity Metrics for Traceability Link Recovery - Evaluation Dataset
<p>This dataset includes all data that was used for the evaluation of my bachelor's thesis:</p> <p><em>Assessing Word Similarity Metrics for Traceability Link Recovery</em></p> <p>The following files correspond to the following data sets from the evaluation:</p> <ul> <li>cc-en-300.tar.gz corresponds to fastText's cc.en.300.bin embedding</li> <li>crawl-300d-2M-subword.tar.gz corresponds to fastText's crawl-300d-2M-subword.bin embedding</li> <li>wiki-news-300d-1M-subword.tar.gz corresponds to fastText's wiki-news-300d-1M-subword.bin embedding</li> <li>wordnet.tar.gz corresponds to the WordNet 3.1 semantic network</li> <li>sewordsim.tar.gz corresponds to SEWordSimDB's vector similarity database</li> <li>glove_cc_840B_300d.tar.gz corresponds to GloVe's CC vector embedding</li> <li>glove_wikigiga_300d.tar.gz corresponds to GloVe's 300 dimensional WIGI vector embedding</li> <li>glove_wikigiga_200d.tar.gz corresponds to GloVe's 200 dimensional WIGI vector embedding</li> <li>glove_wikigiga_100d.tar.gz corresponds to GloVe's 100 dimensional WIGI vector embedding</li> <li>glove_wikigiga_50d.tar.gz corresponds to GloVe's 50 dimensional WIGI vector embedding</li> <li>glove_twitter_200d.tar.gz corresponds to GloVe's 200 dimensional TWTR vector embedding</li> <li>glove_twitter_100d.tar.gz corresponds to GloVe's 100 dimensional TWTR vector embedding</li> <li>glove_twitter_50d.tar.gz corresponds to GloVe's 50 dimensional TWTR vector embedding</li> <li>glove_twitter_25d.tar.gz corresponds to GloVe's 25 dimensional TWTR vector embedding</li> <li>eval_results.tar.gz contains the detailed evaluation results for each configuration of all measures</li> </ul> <p>The licenses of all data sets are included in their respective files.</p> <p>Some of these data sets are .sql files. To use these files to reproduce the evaluation, they need to be imported into a sqlite3 database. The version of ArDoCo used for the evaluation is only able to work with sqlite3 databases and not with sql files.</p>
Dataset of "Inferring Fine-grained Traceability Links between Javadoc Comments and JUnit Test Code"
<p>Dataset of "Inferring Fine-grained Traceability Links between Javadoc Comments and JUnit Test Code"</p> <p>- study object, true link, sentence, test code snippet, experiment result</p>
Exploring Architectural Design Decisions in Mailing Lists and their Traceability to Issue Trackers
<p>The online repository for the paper: Exploring Architectural Design Decisions in Mailing Lists and their Traceability to Issue Trackers, published in ECSA 2024.</p> <p>The online repo has the following folders:</p> <p>1) Classifier: it contains the source code of the classifier and quantitative analysis. Furthermore, there are sufficient details on the classifier performance, and how to replicate the results.</p> <p>2) Dataset: it contains the datasets we used to perform our analysis. This involves dataset before and after BERT classification, as well as exported json files used for training and analysis. The dataset in a zip file. This is a MySQL database in a zip file. It can be opened separately or opened using the search tool.</p> <p>3) Qualitative analysis: it contains coding book of design decisions in mailing lists as well as coding book of methods to discuss ADDs between emails and issues, as well as precision charts for the applied similarity algorithms.</p> <p>4) Searching tool: it contains the jar file and source code of the searching tool. The tool used to annotate and search for emails. It can be used to open the dataset from the zip directly. The tool is a jar file which can be run directly through a double click. The tool is tested on Windows and Linux. The folder also contains keywords to search for architectural emails, as well as documentation.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.