Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

87

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

87 results for “Learning design”

Learn how ShareScore rates datasets ↗
zenodo48/100

Machine learning designs non-hemolytic antimicrobial peptides

<p>The upload contains additional primary data associated with the publication, including raw data in the original file format whenever possible.</p> <p>Data content: HRMS, HPLC-MS, CD, MD, TEM</p>

opencc-by-4.0Jul 2021View details →
zenodo48/100

Hybrid quantum-classical machine learning for generative chemistry and drug design: Generated molecules

<p>Deep generative chemistry models emerge as powerful tools to expedite drug discovery. How- ever, the immense size and complexity of the structural space of all possible drug-like molecules pose significant obstacles, which could be overcome with hybrid architectures combining quantum computers with deep classical networks.&nbsp;As the first step toward this goal, we built a compact discrete variational autoencoder (DVAE) with a Restricted Boltzmann Machine (RBM) of reduced size in its latent layer. The size of the proposed model was small enough to fit on a state-of-the-art D-Wave quantum annealer and allowed training on a subset of the ChEMBL dataset of biologically active compounds. Finally, we generated 2331 novel chemical structures with medicinal chemistry and synthetic accessibility properties in the ranges typical for molecules from ChEMBL.&nbsp;The pre- sented results demonstrate the feasibility of using already existing or soon-to-be-available quantum computing devices as testbeds for future drug discovery applications.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Improving students' privacy awareness – Analysis of a pilot survey to design a VR environment for self-paced learning

<p>In this research, we measured the knowledge of students at the University of Debrecen in the field of data privacy awareness, online and password security.</p> <p><strong>Description</strong></p> <ul> <li>In the questionnaire, green-highlighted answer signs the correct answer to each question.</li> <li>Total data set contains the answers to each question and the respondent&#39;s age.</li> <li>Correct/incorrect data set contains information if the answer is correct to each question, and it also contains the respondent&#39;s age.&nbsp;One means the answer was correct, and zero means the answer was incorrect.</li> </ul>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Replication package for paper: Insights on the Use of Software Design Principles in Machine Learning Pipelines

<p>This is the replication package of the paper "Insights on the Use of Software Design Principles in Machine Learning Pipelines".</p> <p>This replication package contains two files:</p> <ul> <li><a href="../api/records/13828806/draft/files/Data%20extraction.xlsx/content" target="_blank" rel="noopener noreferrer">Data extraction.xlsx</a>: file containing the details of the extracted data for each single ML project.&nbsp;</li> <li><a href="../api/records/13828806/draft/files/Source%20Code%20and%20Metadata.zip/content" target="_blank" rel="noopener noreferrer">Source Code and Metadata.zip</a>: zip file including the source code local copy analyzed and the repository metadata (.json) provided by GitHub API for each ML project .repository&nbsp;</li> </ul> <p>Reference: [1] Lidia L&oacute;pez, Cristina G&oacute;mez, and Claudia Ayala. Insights on the Use of Software Design Principles in Machine Learning Pipelines. <em>Accepted </em>in the 2024 edition of the International Conference on Product-Focused Software Process Improvement (PROFES 2024).</p> <p><strong>Note</strong>: The licence is applicable to the excel file. "Source Code and Metadata.zip" file contains source code repositories downloaded from GitHub, the license for each repository is defined in the corresponding GitHub repository by their authors.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Inverse design of metal-organic frameworks for direct air capture of CO2 via deep reinforcement learning

<p>The combination of several interesting characteristics makes metal-organic frameworks (MOFs) a highly sought-after class of nanomaterials for a broad range of applications like gas storage and separation, catalysis, drug delivery, and so on. However, the ever-expanding and nearly infinite chemical space of MOFs makes it extremely challenging to identify the most optimal materials for a given application. In this work, we present a novel approach using deep reinforcement learning for the inverse design of MOFs, our motivation being designing promising materials for the important environmental application of direct air capture of CO2&nbsp;(DAC). We demonstrate that the reinforcement learning framework can successfully design MOFs with critical characteristics important for DAC. Our top-performing structures populate two separate subspaces of the MOF chemical space: the subspace with high CO2&nbsp;heat of adsorption and the subspace with preferential adsorption of CO2&nbsp;from humid air, with few structures having both characteristics. Our model can thus serve as an essential tool for the rational design and discovery of materials for different target properties and applications.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Comprehensive Datasets for RNA Design, Machine Learning and Beyond

<p>This repository contains a comprehensive collection of RNA multi-loops extracted from major RNA databases, along with benchmark results for various RNA design algorithms. The resource is intended to facilitate research and development in RNA design, particularly for multi-loop structures.</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Fast Radio Map Estimation and Automated Radio Network Design Using Deep Learning App and Dataset

<p>In this research, we present the Deep Learning architecture encoder/fully-connected for estimate optimal desings WLANs in indoor scenarios. This architecture was implemented for WLAN structures consisting of 1, 2, 3, 4, and 5 access points, with the capability to perform the&nbsp;<a href="https://link.springer.com/chapter/10.1007/978-3-662-44415-3_4">Balanced k-means algortihm</a>, but in a fast manner.</p> <p><a href="https://github.com/johanflorez98/Fast-Radio-Map-Estimation-and-Automated-Radio-Network-Design-Using-Deep-Learning-App-and-Dataset/blob/main/README.md#general-dataset-structure">General Dataset Structure</a></p> <p>A major initial difficulty for starting the research was the lack of data, in this case, indoor scenario floor plans, users posisitons and optimal designs for training the architecture. Therefore, it was necessary to create an appropriate database that would facilitate the respective trainings. The dataset was created in base the&nbsp;<a href="https://link.springer.com/chapter/10.1007/978-3-662-44415-3_4">Balanced k-means algortihm</a>. This implementation was carried out in the MATLAB software&nbsp;<a href="https://github.com/johanflorez98/WLANs-optimal-designs-methodology">WLANs-optimal-designs-methodology</a>.</p> <p>Thus, this research provides a dataset that can be used for training multiple Deep Learning architectures and can facilitate future investigations into similar problems.</p> <p>We show a dataset&nbsp;composed by&nbsp;optimal designs for&nbsp;the 5GHz band WiFi&nbsp;in indoor scenarios: it&nbsp;has 102&nbsp;indoor constructions plans and around 20 users&nbsp;distributions per plan floor as image and APs positions per case as coordinates. These distributions are random and several WLAN's structures: 1 to 5 access points.</p> <p>The above explain that we got a total of 61000 RME&nbsp;and CME, this presents that is a model without interference between channels.</p> <p>The pictures have a depth of 8 bits and size of <em>256pixels X&nbsp;256pixels</em> equivalents to indoor constructions of <em>20 X 20</em> m<sup>2</sup>. These ones make reference to offices's spaces at&nbsp;general or classroom.</p> <p><a href="https://github.com/johanflorez98/Fast-Radio-Map-Estimation-and-Automated-Radio-Network-Design-Using-Deep-Learning-App-and-Dataset/blob/main/README.md#obtained-models">Obtained Models</a></p> <p>To evaluate the obtained models, the dataset consisting of floor plans 911 to 102 can be used, with available user distributions for each case of APs configurations as a test set, or other new images can be used.</p> <p>To manipulate the codes better, click here in the <a href="https://github.com/johanflorez98/Fast-Radio-Map-Estimation-and-Automated-Radio-Network-Design-Using-Deep-Learning-App-and-Dataset">repository</a>.</p>

openmit-licenseSep 2023View details →
zenodo40/100

Small dataset machine-learning approach for efficient design space exploration: engineering ZnTe-based high-entropy alloys for water splitting

<p>Atomic structure data used in the research article entitled "Small Dataset Machine-Learning Approaches to Explore the Design Space of High-Entropy Alloys: Engineering ZnTe-based Multicomponent Alloys for the Photo-Splitting of Water"</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Architectural Design Decisions for the Machine Learning Workflow: Dataset and Code

<p><strong>Title:</strong> Architectural Design Decisions for the Machine Learning Workflow: Dataset and Code</p> <p><strong>Authors:</strong> Stephen John Warnett; Uwe Zdun</p> <p><strong>About:</strong> This is the dataset and code artifact for the article entitled &quot;Architectural Design Decisions for the Machine Learning Workflow&quot;.</p> <p><strong>Contents:</strong> The &quot;_generated&quot; directory contains the generated results, including latex files with tables for use in publications and the Architectural Design Decision model in textual and graphical form. &quot;Generators&quot; contains Python applications that can be run to generate the above. &quot;Metamodels&quot; contains a Python file with type definitions. &quot;Sources_coding&quot; contains our source codings and audit trail. &quot;Add_models&quot; contains the Python implementation of our model and source codings. Finally, &quot;appendix&quot; contains a detailed description of our research method.</p> <p><strong>Article Abstract:&nbsp;</strong>Bringing machine learning models to production is challenging as it is often fraught with uncertainty and confusion, partially due to the disparity between software engineering and machine learning practices, but also due to knowledge gaps on the level of the individual practitioner. We conducted a qualitative investigation into the architectural decisions faced by practitioners as documented in gray literature based on Straussian Grounded Theory and modeled current practices in machine learning. Our novel Architectural Design Decision model is based on current practitioner understanding of the topic and helps bridge the gap between science and practice, foster scientific understanding of the subject, and support practitioners via the integration and consolidation of the myriad decisions they face. We describe a subset of the Architectural Design Decisions that were modeled, discuss uses for the model, and outline areas in which further research may be pursued.</p> <p><strong>Objective:</strong> This article aims to study current practitioner understanding of architectural concepts associated with data processing, model building, and Automated Machine Learning (AutoML) within the context of the machine learning workflow.</p> <p><strong>Method:</strong> Applying Straussian Grounded Theory to gray literature sources containing practitioner views on machine learning practices, we studied methods and techniques currently applied by practitioners in the context of machine learning solution development and gained valuable insights into the software engineering and architectural state of the art as applied to ML.</p> <p><strong>Results:</strong> Our study resulted in a model of Architectural Design Decisions, practitioner practices, and decision drivers in the field of software engineering and software architecture for machine learning.</p> <p><strong>Conclusions:</strong> The resulting Architectural Design Decisions model can help researchers better understand practitioners&#39; needs and the challenges they face, and guide their decisions based on existing practices. The study also opens new avenues for further research in the field, and the design guidance provided by our model can also help reduce design effort and risk. In future work, we plan on using our findings to provide automated design advice to machine learning engineers.</p>

openapache2.0Nov 2021View details →
zenodo40/100

Architectural Design Decisions for Machine Learning Deployment: Dataset and Code

<p><strong>Title:</strong> Architectural Design Decisions for Machine Learning Deployment: Dataset and Code</p> <p><strong>Authors:</strong>&nbsp;Stephen John Warnett; Uwe Zdun</p> <p><strong>About:</strong>&nbsp;This is the dataset and code artefact for the paper entitled &quot;Architectural Design Decisions for Machine Learning Deployment&quot;.</p> <p><strong>Contents:</strong>&nbsp;The &quot;_generated&quot; directory contains the generated results, including latex files with tables for use in publications and the Architectural Design Decision model in textual and graphical form. &quot;Generators&quot; contains Python applications that can be run to generate the above. &quot;Metamodels&quot; contains a Python file with type definitions. &quot;Sources_coding&quot; contains our source codings and audit trail. &quot;Add_models&quot; contains the Python implementation of our model and source codings. Finally, &quot;appendix&quot; contains a detailed description of our research method.</p> <p><strong>Paper Abstract:</strong>&nbsp;Deploying machine learning models to production is challenging, partially due to the misalignment between software engineering and machine learning disciplines but also due to potential practitioner knowledge gaps. To reduce this gap and guide decision-making, we conducted a qualitative investigation into the technical challenges faced by practitioners based on studying the grey literature and applying the Straussian Grounded Theory research method. We modelled current practices in machine learning, resulting in a UML-based architectural design decision model based on current practitioner understanding of the domain and a subset of the decision space and identified seven architectural design decisions, various relations between them, twenty-six decision options and forty-four decision drivers in thirty-five sources. Our results intend to help bridge the gap between science and practice, increase understanding of how practitioners approach the deployment of their solutions, and support practitioners in their decision-making.</p> <p><strong>Objective:</strong>&nbsp;This paper aims to study current practitioner understanding of architectural concepts associated with machine learning deployment.</p> <p><strong>Method:</strong>&nbsp;Applying Straussian Grounded Theory to gray literature sources containing practitioner views on machine learning practices, we studied methods and techniques currently applied by practitioners in the context of machine learning solution development and gained valuable insights into the software engineering and architectural state of the art as applied to ML.</p> <p><strong>Results:</strong>&nbsp;Our study resulted in a model of Architectural Design Decisions, practitioner practices, and decision drivers in the field of software engineering and software architecture for machine learning.</p> <p><strong>Conclusions:</strong>&nbsp;The resulting Architectural Design Decisions model can help researchers better understand practitioners&#39; needs and the challenges they face, and guide their decisions based on existing practices. The study also opens new avenues for further research in the field, and the design guidance provided by our model can also help reduce design effort and risk. In future work, we plan on using our findings to provide automated design advice to machine learning engineers.</p>

openapache2.0Jan 2022View details →
zenodo40/100

Design of experiment (DOE) used in the study: Lightweight design of variable-stiffness imperfection-insensitive cylinders enabled by continuous tow shearing and machine learning

<p>There are five input variables that are changed for this design of experiment (DOE) within the following range:</p> <p><span class="math-tex">\(\begin{eqnarray} 0.05 \leq r_{CTS} \leq 0.20 \nonumber \\ 1 \leq n \leq 12 \nonumber \\ 0 \leq {c_2}_{ratio} \leq 1 \\ 0 \leq \theta_1 \leq 75 \nonumber \\ 0 \leq \theta_2 \leq 75 \nonumber \end{eqnarray}\)</span></p> <p>For the sake of simplicity, these variables are respectively called v1, v2, v3, v4, v5.</p> <p>The DOE consist of 2000 points created with Latin Hyper-cube Sampling, as available in the LHS toolbox (Carnell, R. lhs: Latin Hypercube Samples, 2021. R package version 1.1.3.).</p> <p>The outputs evaluated with this design of experiment (DOE) are the critical buckling load $P_{critical}$, $b_{factor}$ obtained with Koiter&#39;s asymptotic approach, and the mass.</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

TocoDecoy: a new approach to design unbiased datasets for training and benchmarking machine-learning scoring functions

<p>This dataset file contains TocoDecoy datasets generated based on the targets and active ligands of LIT-PCBA.</p> <p>1_property_filtered.zip :</p> <ul> <li>TD set: the ligand file name, 2D T-sne vectors, Smiles, molecular weight (MW), Wildman-Crippen partition coefficient (log P), number of rotatable bonds (RB), number of hydrogen-bond acceptors (HBA), number of hydrogen-bond donors (HBD), number of halogens (HAL), topology similarities of decoys to the seed active ligands, active label (active or inactive) and training set label (whether belongs to training set or test set) <strong>OF active ligands and their topologically dissimilar decoys</strong></li> <li>CD set: the decoy conformations with low docking scores generated by docking active ligands into protein pockets using Glide, Schr&ouml;dinger.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

Figure 1. The design of the study-The Effect of English Learning Anxiety on Iranian High-School Students' English Language Achievement

<p>The dependent variable in this study was English<br> achievement, and the independent variable was language anxiety. The intervening variable is<br> English proficiency. The control variables are: age of subjects (second year high school students<br> with an average age of 17) and years of experience in English learning (a minimum of 5<br> consecutive years). The schematic design of the study is presented below (Figura 1).</p>

opencc-by-4.0Jun 2011View details →
zenodo40/100

Datasets for "Deep learning-based design of synthetic orthologs of SH3 signaling domains"

<p>Description of data for "Deep learning-based design of synthetic orthologs of SH3 signaling domains":<br><br>biochemistry_data.zip --&gt; contains the binding assay, melting temperature, and enthalpy measurements that reproduce table 1 in the main text.<br>sequence_data.zip --&gt; contains the sequences with relative enrichment measurements and other meta data information (e.g. latent embeddings, paralog labels, etc.).<br><br>sequence_data.zip &gt; SH3_Library_Natural.xlsx --&gt; contains the natural alleles with normalized relative enrichment scores, paralog labels, mmd latent coordinates, and among other meta data.<br>sequence_data.zip &gt; SH3_Library_Design.xlsx --&gt; contains the design alleles with normalized relative enrichment scores, mmd latent coordinates, and among other meta data.<br>sequence_data.zip &gt; paralog_mapping.xlsx --&gt; contains the mapping between paralog name and COG labels.<br><br>note: to access these spreadsheets for analysis, we recommend using pandas library in Python. To reproduce figures within the manuscript, please follow this github repo link: https://github.com/chemgeeklian/SH3_orthology_paper_analysis</p>

opencc-by-4.0Jul 2024View details →
zenodo40/100

Accompanying data for publication: "Learning the Optimal Power Flow: Environment Design Matters"

<p>All the data created for the publication "Learning the Optimal Power Flow: Environment Design Matters" by Wolgast and Nie&szlig;e. The dataset contains all training runs performed, including the final neural network weights, meta-data about the training run, and various metrics during the course of training, which were used to generate the results and plots. The source code to re-produce the plots for the publication (and everything else) can be found on GitHub: https://github.com/Digitalized-Energy-Systems/rl-opf-env-design</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Extended data integration of motivational aspects in gamification and game-based learning educational designs

<p>This is the extended data for a systematic review article about the integration of motivational aspects in gamification and game-based learning educational designs related to teacher&acute;s training and teacher&acute;s professional development.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

UNSUPERVISED MACHINE LEARNING AND VECTOR MODELS IN DESIGNING AND OPTIMIZATION OF TELECOM RETAIL CHANNELS

<p>This paper examines the use of unsupervised machine learning and vector models in the design and optimization of retail channels for telecommunications services. Unsupervised machine learning allows you to analyze and identify hidden patterns in large volumes of untagged data, which is especially important in a dynamically changing consumer market. Vector models, in turn, provide high accuracy of demand forecasting and inventory management, contributing to an increase in the efficiency of trading channels. The synergy of these technologies allows companies to improve customer experience, optimize operational processes and increase competitiveness in the market. The main focus of the work is on data processing methods, including correlation analysis, the use of the support vector machine (SVM) method and its adaptation to solve problems related to predicting customer behavior and optimizing logistics processes.</p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Datasets for the manuscript "In silico proof of principle of machine learning-based antibody design at unconstrained scale"

<p>The zip file contains dataset files for the manuscript &quot;In silico proof of principle of machine learning-based antibody design at unconstrained scale&quot;</p>

opencc-by-4.0Aug 2021View details →
zenodo40/100

NGS Data Accompanying "Deep Learning Enables Design of Multifunctional Synthetic Human Gut Microbiome Dynamics"

<p>NGS Data Accompanying &quot;Deep Learning Enables Design of Multifunctional Synthetic Human Gut Microbiome Dynamics&quot;, currently in review.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Challenges and Limitations in the Design and Implementation of Fair and Equitable Machine Learning Algorithms in Healthcare

<p>We provide the programs in Python, a CSV file with references, images used in the article and a text corpus generated with Python.</p>

opencc-by-4.0Jun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record