Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
29
datasets available to search
ShareScore release 0.9.0
Dataset results
29 results for “Systems Domain”
Domain-Driven Design for Microservices Architecture Systems Development: A Systematic Mapping Study
<p>This repository contains all artifacts related to the study: Domain-Driven Design for Microservices Architecture Systems Development: A<br> Systematic Mapping Study</p>
CrossDomainTypes4Py: A Python Dataset for Cross-Domain Evaluation of Type Inference Systems
<p>This dataset contains python repositories mined on GitHub on January 20, 2021. It allows a cross-domain evaluation of type inference systems. For this purpose, it consists of two sub-datasets, each containing only projects from the web or scientific calculation domain, respectively. Therefore we searched for projects with dependencies to either <a href="https://numpy.org/">NumPy</a> or <a href="https://flask.palletsprojects.com/en/2.0.x/">Flask</a>. Furthermore, only projects with dependencies to <a href="http://mypy-lang.org/">mypy</a> were considered, because this should ensure that at least parts of the projects have type annotations. These can be used later as ground truth. Further details about the dataset will be described in an upcoming paper, as soon as it is published it will be linked here.<br> The dataset consists of two files for the two sub-datasets. The web domain dataset contains 3129 repositories and the scientific calculation domain dataset contains 4783 repositories. The files have two columns with the URL to the GitHub repository and the used commit hash. Thus, it is possible to download the dataset using shell or python scripts, for example, the pipeline provided by <a href="https://github.com/saltudelft/many-types-4-py-dataset">ManyTypes4Py</a> can be used.<br> If repositories do not exist anymore or are private, you can contact us via the following email address: bernd.gruner@dlr.de. We have a backup of all repositories and will be happy to help you. </p>
Expression of immunoglobulin constant domain genes in neurons of the mouse central nervous system
<p>Data related to the publication "Expression of immunoglobulin constant domain genes in neurons of the mouse central nervous system". </p> <p>The fasta file (.fa) contains the sequence of neuronal FC-Ighm</p> <p>The .pdb files contain the model of the two different protein versions of Ighm</p> <p>The excel file contains the data of the quantification of co-expression and the ATG prediction results.</p>
Supplementary Material: Comparing Formal Tools for System Design: a Case Study from the Railway Domain
<p>The package includes a set of models for a railway moving-block system:</p> <p>(a) a PDF document named Moving-block Model and Requirements.pdf, which includes a UML model of a moving-block system together with a set of requirements for the system;</p> <p>(b) a set of 10 folders, each one associated to a formal or semi-formal development tool. Each folder contains one or more model of the moving-block system from (a), developed by means of the tool.</p>
Domain-Driven Design in Microservices-Based Systems Development: A Systematic Literature Review and Thematic Analysis [Dataset]
<p>This repository contains all artifacts related to the study: Domain-Driven Design in Microservices-Based Systems Development: A Systematic Literature Review and Thematic Analysis</p>
Research data supporting: "Detecting dynamic domains and local fluctuations in complex molecular systems via timelapse neighbors shuffling"
<p>This repository contains the set of data shown in the paper "Detecting dynamic domains and local fluctuations in complex molecular systems via timelapse neighbors shuffling" published on PNAS (DOI: 10.1073/pnas.2300565120).</p>
MultiCardioNER Corpus: Multilingual Adaptation of Clinical NER Systems to the Cardiology Domain
<h1><strong>MultiCardioNER</strong></h1> <p><strong>MultiCardioNER</strong> is a shared task about the adaptation of clinical NER systems to the cardiology domain. It uses a combination of two existing datasets (DisTEMIST for diseases and the newly-released DrugTEMIST for medications), as well as a new, smaller dataset of cardiology clinical cases annotated using the same guidelines.</p> <p>Participants are provided DisTEMIST and DrugTEMIST as training data to use as they see fit (1,000 documents, with the original partitions splitting them into 750 for training and 250 for testing). The cardiology clinical cases (cardioccc) are meant to be used as a development or validation set (258 documents), although participants are encourage to experiment with the documents and annotations as they see fit. The evaluation is done using a different collection of cardiology clinical cases (250).</p> <p>MultiCardioNER proposes two tracks:</p> <p>- Track 1: Spanish adaptation of disease recognition systems to the cardiology domain.<br>- Track 2: Multilingual (Spanish, English and Italian) adaptation of medication recognition systems to the cardiology domain.</p> <p>Please read the README file attached for more information on folder structure and file format.</p> <p><strong>MultiCardioNER</strong> was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of BioASQ 2024. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">https://temu.bsc.es/multicardioner</a>. This task is promoted by Spanish and European projects such as DataTools4Heart, AI4HF, BARITONE and AI4ProfHealth.</p> <p><strong>UPDATE MAY 28th 2024: </strong>The test set annotations are now out! We've also included the original background set files, as well as a file with the mappings from the masked filenames used during the evaluation phase to the original filenames. Please check the README for more information.</p> <h2><strong>Resources</strong></h2> <ul> <li><a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">MultiCardioNER website</a></li> <li><a href="http://bioasq.org/" target="_blank" rel="noopener">BioASQ website</a></li> <li><a href="../doi/10.5281/zenodo.6458078" target="_blank" rel="noopener">DisTEMIST Guidelines</a></li> <li><a href="../doi/10.5281/zenodo.11065432" target="_blank" rel="noopener">DrugTEMIST Guidelines</a></li> </ul> <h2><strong>License</strong></h2> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <h2><strong>Contact</strong></h2> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)<br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p> <h2><strong>Additional resources and corpora</strong></h2> <p>If you are interested in MultiCardioNER, you might want to check out these corpora and resources:</p> <ul> <li><a href="../records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT)</li> <li><a href="../records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT)</li> <li><a href="../records/10635215">SympTEMIST</a> (Corpus of clinical findings and normalization to SNOMED CT)</li> <li><a href="../records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization)</li> <li><a href="../records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization)</li> <li><a href="../records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization)</li> <li><a href="../records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI))</li> <li><a href="../records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization)</li> <li><a href="../records/3837305">CodiESP</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version)</li> <li><a href="../records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy)</li> <li><a href="../records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags)</li> <li><a href="../records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries))</li> <li><a href="../records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags)</li> <li><a href="../records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts)</li> </ul>
Neural network prediction of strong lensing systems with domain adaptation and uncertainty quantification
<p>This project combines the emerging field of Domain Adaptation with Uncertainty Quantification, working towards applying machine learning to real scientific datasets with limited labelled data. For this project, simulated images of strong gravitational lenses are used as source and target dataset, and the Einstein radius θ E and its uncertainty are determined through regression.</p> <p>Applying machine learning in science domains such as astronomy is difficult. With models trained on simulated data being applied to real data, models frequently underperform - simulations cannot perfectlty capture the true complexity of real data. Enter domain adaptation (DA). The DA techniques used in this work use Maximum Mean Discrepancy (MMD) Loss to train a network to being embeddings of labelled "source" data gravitational lenses in line with unlabeled "target" gravitational lenses. With source and target datasets made similar, training on source datasets can be used with greater fidelity on target datasets.</p> <p>Scientific analysis requires an estimate of uncertainty on measurements. We adopt an approach known as mean-variance estimation, which seeks to estimate the variance and control regression by minimizing the beta negative log-likelihood loss. To our knowledge, this is the first time that domain adaptation and uncertainty quantification are being combined, especially for regression on an astrophysical dataset.</p>
Testing Context-Aware Software Systems in the Automotive Domain: A Multi Vocal Literature Review Protocol and Dataset
<h2>Testing Context-Aware Software Systems in the Automotive Domain: A Multi Vocal Literature Review Protocol and Dataset</h2><p>A Multi-Vocal Literature Review (MVLR) is a form of a systematic review that includes grey literature in addition to peer-review literature (Garousi, Felderer, and Mäntylä 2019). The decision to justify an MVLR is drawn from the results of recent literature reviews, in particular the recent results from (Matalonga et al. 2022) where it is shown that there is little evidence in the white literature on the approaches to testing non-academic CASS Systems.</p><p>Previous academic works (including our Quasi-Systematic Literature Reviews and Rapid Reviews) operate under the following assumptions and observations:</p><ul><li>Assumption 1. CASS are widespread and being deployed for commercial or industrial use.</li><li>Observation 1. The software engineering and software testing communities have had time to adopt (or develop new) techniques to deal with the context and effects of testing software systems.</li><li>Observation 2. Academics have been able to work with software organizations to transfer or study the approaches used to test CASS, yet the published case studies we are aware of describe a partial picture of the overall adoption and approach of the problem.</li></ul><p>In spite of these assumptions and the availability of systematic literature review studies, there is little evidence of how software organizations are testing CASS.</p><h3>Research Goals</h3><h4><strong>Aim: </strong>To uncover evidence on how the automotive industry reports their working with the dynamic testing process regarding CASS.</h4><p>We use the term industry to broaden our scope to include stakeholders with an interest in the quality of CASS like NGOs, standard-setting organizations and regulation-setting organizations who can influence how software must be treated in different domains.</p><p>The following research questions convey the general interest of our enquiries. These are driven by our previous research and expectations on the sources.</p><ul><li><strong>RQ1</strong> Are there sources to support the understanding and indicate directions to deal with the problem of testing CASS?</li><li><strong>RQ2</strong> What are the challenges of using these dynamic testing process solutions?</li><li><strong>RQ3 </strong>How are the dynamic testing processes that deal with the context of CASS described in the sources?</li></ul><h3>Dataset </h3><p>This dataset contains the following artifacts:</p><ol><li><strong>Testing CASS MVLR Protocol.pdf</strong>: protocol containing the methodological details of performing the MVLR</li><li><strong>Sources Identification, Selection and Data Extraction.xlsx</strong>: spreadsheet used to record and control discovered and selected sources</li><li><strong>Extraction_Documents.zip</strong>: compressed file containing all data collection forms filled with data extracted from the</li><li><strong>MVLR_Analysis-Codebook.xlsx</strong>: listing of codes emerging from the collected data</li></ol><p> </p>
Introducing a transition domain for describing the solute exchange between macropores/fractures and matrix in dual-permeability system
<p>出口 1 至 5 的浓度值和出口处的总浓度记录在"实验数据"中。图 8 至 10 中的仿真结果数据记录在"仿真结果"中。构建数值模型的方法记录在"数值模拟"中,需要 COMSOL Multiphysics 将其打开。</p>
Adopting Microservices and DevOps in the Cyber-Physical Systems Domain: A Rapid Review and Case Study at Siemens AG
<p>This repository contains the artifacts for a rapid review and interview-based case study at Siemens AG. We analyzed challenges and practices related to microservices and DevOps in the context of the cyber-physical systems (CPS) domain.</p>
'The Canon' (Heidelbergensis Palatinus gr. 281, fol. 173 v—public domain image) and the structure of the Hypolydian Unchanging Perfect System.
<p>The Canon’ (Heidelbergensis Palatinus gr. 281, fol. 173 v—<a href="https://digi.ub.uni-heidelberg.de/diglit/cpgraec281/0356">public domain image</a>) and the structure of the Hypolydian Unchanging Perfect System –– Figure 6 from Lynch forthcoming, <em>Unlocking the Riddles of Imperial Greek Melodies I: the “Lydian” Metamorphosis of the Classical Harmonic System.</em></p> <p>Licensed<em> </em>under CC-BY-NC-ND</p>
Dataset: Maturity Evaluation of Domain-Specific Language Ecosystems for Cyber-Physical Production Systems
<p>This online repository contains the accompanying data for the ETFA publication "Maturity Evaluation of Domain-Specific Language Ecosystems for Cyber-Physical Production Systems".<br> <br> This involves all tables which contain the values of the maturity evaluation criteria of the five examined subject DSLs.</p>
Early Identification of Mental Disorders: Application of a Multi-modal & Domains System
ClinicalTrials.gov study NCT05939154. IPD Sharing: YES. Countries: 1. Publications: 8.
Self-control and Mindfulness Within Ambulatorily Assessed Network Systems Across Health Related Domains
ClinicalTrials.gov study NCT02647801. IPD Sharing: Not stated. Countries: 1. Publications: 7.
Dataset for the paper "Fast Multi-Distance Time-Domain NIRS and DCS System for Clinical Applications"
Open the record for dataset details and reuse information.
Supplementary material 6 from: Petrocelli A, Cecere E, Rubino F (2019) Successions of phytobenthos species in a Mediterranean transitional water system: the importance of long term observations. In: Mazzocchi MG, Capotondi L, Freppaz M, Lugliè A, Campanaro A (Eds) Italian Long-Term Ecological Research for understanding ecosystem diversity and functioning. Case studies from aquatic, terrestrial and transitional domains. Nature Conservation 34: 217-246. https://doi.org/10.3897/natureconservation.34.30055
: Data type: measurement
Supplementary material 2 from: Petrocelli A, Cecere E, Rubino F (2019) Successions of phytobenthos species in a Mediterranean transitional water system: the importance of long term observations. In: Mazzocchi MG, Capotondi L, Freppaz M, Lugliè A, Campanaro A (Eds) Italian Long-Term Ecological Research for understanding ecosystem diversity and functioning. Case studies from aquatic, terrestrial and transitional domains. Nature Conservation 34: 217-246. https://doi.org/10.3897/natureconservation.34.30055
: Data type: measurement
Supplementary material 5 from: Petrocelli A, Cecere E, Rubino F (2019) Successions of phytobenthos species in a Mediterranean transitional water system: the importance of long term observations. In: Mazzocchi MG, Capotondi L, Freppaz M, Lugliè A, Campanaro A (Eds) Italian Long-Term Ecological Research for understanding ecosystem diversity and functioning. Case studies from aquatic, terrestrial and transitional domains. Nature Conservation 34: 217-246. https://doi.org/10.3897/natureconservation.34.30055
: Data type: measurement
Supplementary material 7 from: Petrocelli A, Cecere E, Rubino F (2019) Successions of phytobenthos species in a Mediterranean transitional water system: the importance of long term observations. In: Mazzocchi MG, Capotondi L, Freppaz M, Lugliè A, Campanaro A (Eds) Italian Long-Term Ecological Research for understanding ecosystem diversity and functioning. Case studies from aquatic, terrestrial and transitional domains. Nature Conservation 34: 217-246. https://doi.org/10.3897/natureconservation.34.30055
: Data type: measurement
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.