Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Quantification of salt stress in wheat leaves by Raman spectroscopy and machine learning
<p>Train and test datasets used in the manusicript "Quantification of salt stress in wheat leaves by Raman spectroscopy and machine learning". Trained models are included.</p>
A Dataset for Utility Prediction in Computational Persuasion with Machine Learning Techniques
<p>This dataset contains data for a new benchmark for the prediction of user's utilities with Machine Learning techniques for Computational Persuasion. This work has been accepted at AAAI-22, more information in the relative repository containing the source code: <a href="https://github.com/ivanDonadello/ML-Argument-Based-Computational-Persuasion">https://github.com/ivanDonadello/ML-Argument-Based-Computational-Persuasion</a></p>
Architectural Design Decisions for Machine Learning Deployment: Dataset and Code
<p><strong>Title:</strong> Architectural Design Decisions for Machine Learning Deployment: Dataset and Code</p> <p><strong>Authors:</strong> Stephen John Warnett; Uwe Zdun</p> <p><strong>About:</strong> This is the dataset and code artefact for the paper entitled "Architectural Design Decisions for Machine Learning Deployment".</p> <p><strong>Contents:</strong> The "_generated" directory contains the generated results, including latex files with tables for use in publications and the Architectural Design Decision model in textual and graphical form. "Generators" contains Python applications that can be run to generate the above. "Metamodels" contains a Python file with type definitions. "Sources_coding" contains our source codings and audit trail. "Add_models" contains the Python implementation of our model and source codings. Finally, "appendix" contains a detailed description of our research method.</p> <p><strong>Paper Abstract:</strong> Deploying machine learning models to production is challenging, partially due to the misalignment between software engineering and machine learning disciplines but also due to potential practitioner knowledge gaps. To reduce this gap and guide decision-making, we conducted a qualitative investigation into the technical challenges faced by practitioners based on studying the grey literature and applying the Straussian Grounded Theory research method. We modelled current practices in machine learning, resulting in a UML-based architectural design decision model based on current practitioner understanding of the domain and a subset of the decision space and identified seven architectural design decisions, various relations between them, twenty-six decision options and forty-four decision drivers in thirty-five sources. Our results intend to help bridge the gap between science and practice, increase understanding of how practitioners approach the deployment of their solutions, and support practitioners in their decision-making.</p> <p><strong>Objective:</strong> This paper aims to study current practitioner understanding of architectural concepts associated with machine learning deployment.</p> <p><strong>Method:</strong> Applying Straussian Grounded Theory to gray literature sources containing practitioner views on machine learning practices, we studied methods and techniques currently applied by practitioners in the context of machine learning solution development and gained valuable insights into the software engineering and architectural state of the art as applied to ML.</p> <p><strong>Results:</strong> Our study resulted in a model of Architectural Design Decisions, practitioner practices, and decision drivers in the field of software engineering and software architecture for machine learning.</p> <p><strong>Conclusions:</strong> The resulting Architectural Design Decisions model can help researchers better understand practitioners' needs and the challenges they face, and guide their decisions based on existing practices. The study also opens new avenues for further research in the field, and the design guidance provided by our model can also help reduce design effort and risk. In future work, we plan on using our findings to provide automated design advice to machine learning engineers.</p>
Dataset of the paper "Machine learning for expert-level image-based identification of very similar species in the hyperdiverse plant bug family Miridae (Hemiptera: Heteroptera)"
<p>This dataset contains 3792 images of 26 plant bug (Insecta: Heteroptera: Miridae: Mirini) species used to test the performance of a CNN in species recognition. All jpg files are 1920 pixels on the long size and additionally available as an archive file to facilitate download of the entire dataset. </p> <p>Bar code labels (unique specimen identifiers or USIs) were attached to all examined specimens used for this study. Further information such as additional photographs of habitus and genitalic structures, georeferenced coordinates of each locality, specimens dissected, notes, collecting method can be obtained from the Heteroptera Species Pages (http://research.amnh.org/pbi/heteropteraspeciespage/) which assembles available data from a specimen database and are also provided as an Excel spreadsheet (file _Adelphocoris_CNN_label_data.xlsx).</p>
TimeSpec4LULC: A Smart-Global Dataset of Multi-Spectral Time Series of MODIS Terra-Aqua from 2000 to 2021 for Training Machine Learning models to perform LULC Mapping
<p>TimeSpec4LULC is a smart open-source global dataset of multi-spectral time series for 29 Land Use and Land Cover (LULC) classes ready to train machine learning models. It was built based on the seven spectral bands of the MODIS sensors at 500 m resolution from 2000 to 2021 (262 observations in each time series). Then, was annotated using spatial-temporal agreement across the 15 global LULC products available in Google Earth Engine (GEE).</p> <p>TimeSpec4LULC contains two datasets: the original dataset distributed over 6,076,531 pixels, and the balanced subset of the original dataset distributed over 29000 pixels.</p> <p>The original dataset contains 30 folders, namely "Metadata", and 29 folders corresponding to the 29 LULC classes. The folder "Metadata" holds 29 different CSV files describing the metadata of the 29 LULC classes. The remaining 29 folders contain the time series data for the 29 LULC classes. Each folder holds 262 CSV files corresponding to the 262 months. Inside each CSV file, we provide the seven values of the spectral bands as well as the coordinates for all the LULC class-related pixels.</p> <p>The balanced subset of the original dataset contains the metadata and the time series data for 1000 pixels per class representative of the globe. It holds 29 different JSON files following the names of the 29 LULC classes.</p> <p>The features of the dataset are:</p> <p>- ".geo": the geometry and coordinates (longitude and latitude) of the pixel center.</p> <p>- "ADM0_Code": the GAUL country code.</p> <p>- "ADM1_Code": the GAUL first-level administrative unit code.</p> <p>- GHM_Index": the average of the global human modification index.</p> <p>- "Products_Agreement_Percentage": the agreement percentage over the 15 global LULC products available in GEE.</p> <p>- "Temporal_Availability_Percentage": the percentage of non-missing values in each band.</p> <p>- "Pixel_TS": the time series values of the seven spectral bands.</p>
Machine Learning based scratches on printed paper detection, in high-speed printing systems [Dataset]
<p>Printing industry rapidly is adopting digital technologies and the requirements in terms of speed and print quality are also becoming more demanding. The is a wide range of possible quality defects in printed paper. This makes it impossible to have humans inspect the printed paper for such a big amount of possible quality defects at the high-speeds the printouts are produced.</p> <p>Printing industry is not taking advantage of the Artificial Intelligence to detect defects in printed paper at speed without human intervention. It is possible to generate millions of images (captures) with printed content from a printing system every day. Most of these images will not have any defect but some other will and can be used to generate a data set to be used in a machine learning system.</p> <p>The intention of this research work is to find ways artificial intelligence can help on automatically detecting defects on printed paper in a printing system and classifying them, without human intervention. Focusing on scratches, I’ve explored what are the actual proposals and solutions, and how machine learning can help improving them by using datasets with different techniques, implementing possible solutions and comparing the obtained results.</p>
Data for the paper: Water-Food-Energy nexus: Learning from global cities using machine learning algorithms
<p>This data respository contains the datasets used to produce the "<strong>Water-Food-Energy nexus: Learning from global cities using machine learning algorithms" </strong>article. </p>
A database of MMS bow shock crossings compiled using machine learning
<p>We use a machine learning approach to automatically identify shock crossings from the Magnetospheric Multiscale (MMS) spacecraft. We compile a database of 2797 crossings including various spacecraft related and shock related parameters for each event. Furthermore, for each event we provide an overview plot containing key parameters of the shock crossing.</p> <p>A Technical report detailing the content of the database can be found at the DOI: http://dx.doi.org/10.1029/2022JA030454</p>
Companion for "Understanding Distributed Deep Learning Performance by Correlating HPC and Machine Learning Measurements"
<p>This is the Companion Material for the paper “Understanding Distributed Deep Learning Performance by Correlating HPC and Machine Learning Measurements”, by Ana Luisa Veroneze Solórzano and Lucas Mello Schnorr. The manuscript was approved for publication in the <a href="https://www.isc-hpc.com/research-papers-2022.html">ISC High Performance 2022</a> for the Research Papers session. A public companion is also availabl in GitLab: <a href="https://gitlab.com/anaveroneze/isc2022-companion/">https://gitlab.com/anaveroneze/isc2022-companion</a>.</p> <p> </p>
Home-based measurements of dystonia and choreoathetosis in cerebral palsy using smartphone-coupled inertial sensor technology and machine learning: A proof-of-concept study - dataset
<p>Home-based measurements of dystonia in cerebral palsy using smartphone-coupled inertial sensor technology and machine learning: A proof-of-concept study</p> <p> </p> <p>This project contains:</p> <p>- 1 main MATLAB script: MODYSathome_main.m<br> - 12 MATLAB functions:<br> - function_calc_mean_recall_precision.m<br> - function_create_dataframes.m<br> - function_deep_learning.m<br> - function_determine_best_ML_model.m<br> - function_display_DL_results.m<br> - function_display_ML_results.m<br> - function_index_extremities.m<br> - function_machine_learning.m<br> - function_oversample.m<br> - function_partition_data.m<br> - function_pick_best_models.m<br> - function_prepare_DL_data.m</p> <p>Downloading the Matlab scripts</p> <p> - Create a folder named 'MODYS' and create a subfolder named 'results'<br> - Download the zip file via <a href="https://zenodo.org/record/6379348">RehabAUmc/modys-at-home: v1.0 | Zenodo</a><br> - Unzip the zip file in the path MODYS\</p> <p>STEPS<br> 1. Open MATLAB<br> 2. In MATLAB, go to the 'HOME' tab and click on 'Set Path'<br> 3. Click on 'Add Folder' and browse to MODYS/RehabAUmc-modys-at-home-86b14c3/functions<br> 4. Click on 'Select Folder' and click on 'Save'<br> 5. Click on 'Browse to folder' and browse to a patients' data in MODYS/data/PatientXXX, then click on 'Select Folder'<br> 6. In the 'HOME' tab click on 'Open' and open MODYSathome.m in MODYS/RehabAUmc-modys-at-home-86b14c3<br> 7. In the 'EDITOR' tab click on 'Run Section' to run the script<br> 8. When the code has been run, the results are displayed in the Command Window and saved in MODYS/results/PatientXXX</p>
InpactorDB: A Plant classified lineage-level LTR retrotransposon reference library for free-alignment methods based on Machine Learning
<p>LTR retrotransposons are mobile elements that make up the major part of most plant genomes. Their identification and annotation via bioinformatics approaches represent a major challenge in the era of massive plant genome sequencing. In addition to their involvement in the variation in genome size, these elements are also associated in the function and structure of different chromosomal regions and in the alteration of the function of coding regions, among others. Several plant retrotransposon sequence databases of LTR retrotransposons are available with public access such as PGSB, RepetDB or restricted access such as Repbase. Although they are useful for approaches to identify LTR-RTs in new genomes by similarity, the elements of these databases are not classified down to the lineage/family level. with great depth. </p> <p>Here, we present InpactorDB a semi-curated dataset composed of 130,511 elements from 195 plant genomes (belonging to 108 plant species), classified down to the lineage level. This data set has been used to train two deep neural networks (one fully connected and one convolutional) for fast classification of elements. Used in lineage-level classification approaches, we obtain a score above 98% of F1-score, precision and recall. </p> <p>In order to classify elements of the ‘LTR_STRUC’ and ‘EDTA’ datasets, we used the methodology proposed by Inpactor, which uses homology-based strategy with known coding domains belonging to LTR-RTs. We utilized the RexDB domain library as reference. LTR-RTs were classified into superfamilies, Gypsy (RLG) or Copia (RLC) and sub-classified into lineages according to the similarities of five different amino acid reference domains (GAG, AP, RT, RNAseH, and INT domains). In addition, we applied filters to remove keep only intact elements:</p> <p>1) to remove predicted elements with domains from two different superfamilies (i.e. Gypsy and Copia),</p> <p>2) or elements with domains belonging to two or more different lineages,</p> <p>3) to remove elements with lengths different than those reported by Gypsy Database (GyDB) with a tolerance of 20%,</p> <p>4) to delete incomplete elements which has less than three identified domains, and</p> <p>5) to remove elements with insertions of TE class II (reported in Repbase). </p> <p>The final non-redundant version of InpactorDB consists of 67,305 LTR retrotransposons. Both redundant and non-redundant versions of InpactorDB are available in Fasta format in which sequences have identifiers with the following general Identification code:</p> <p>>Superfamily-Lineage-plant_family-specie-source-length-ID,</p> <p>Where Superfamily can is either RLC (for Copia) or RLG (for Gypsy), Lineage/family follows following the RexDB nomenclature, source (can be Repbase, RepetDB, PGSB, LTR_STRUC or EDTA datasets), length, and ID, is a unique number which identify each element inside the InpactorDB.</p>
Artifact: "If security is required": Engineering and Security Practices for Machine Learning-based IoT Devices
<p>Artifact for "If security is required": Engineering and Security Practices for Machine Learning-based IoT Devices</p>
Sticky Pi -- Machine Learning Data, Configuration and Models
<p><strong>Dataset for the Machine Learning section of the Sticky Pi project (https://doc.sticky-pi.com/)</strong></p> <p>Contains the dataset for the three algorithms described in the publication: Universal Insect Detector, Siamese Insect Matcher and Insect Tuboid Classifier.</p> <p><strong>Universal Insect Detector:</strong></p> <p>`universal_insect_detector/` contains training/validation data, configuration files to train the model, and the model as trained and used for publication.</p> <ul> <li>`data/` – A set of svg images that contain the embedded jpg raw image, and a set of non-intersecting polygon around the labelled insects</li> <li>`output/` <ul> <li>`model_final.pth` – the model as trained for the publication</li> </ul> </li> <li>`config/` <ul> <li>`config.yaml `– The configuration file defining the hyperparameters to train the model</li> <li>`mask_rcnn_R_101_C4_3x.yaml` – the base configuration file from which config is derived</li> </ul> </li> </ul> <p> </p> <p><strong>Siamese Insect Matcher</strong></p> <p>`siamese_insect_matcher/` contains training/validation data, configuration files to train the model, and the model as trained and used for publication.</p> <ul> <li>`data/` – a set of svg images that contain two embedded jpg raw images vertically stacked corresponding to two frames in a series. Each predicted insect is labelled as a polygon. Insects that are labelled as the same instance, between the two frames, are grouped (i.e. SVG group). The filename of each image is `<device>.<datetime_frame_1>.<datetime_frame_2>.svg`</li> <li>`output/` <ul> <li>`model_final.pth` – the model as trained for the publication</li> </ul> </li> <li>`config/` <ul> <li>`config.yaml` – The configuration file defining the hyperparameters to train</li> </ul> </li> </ul> <p><strong>Insect Tuboid Classifier:</strong></p> <p>`insect_tuboid_classifier/` contains images of insect tuboid, a database file describing their taxonomy, a configuration file to train the model, and the model as trained and used for publication.</p> <ul> <li>`data/` <ul> <li>`database.db`: a sqlite file with a single table `ANNOTATIONS`. The table maps a unique identifier of each tuboid (tuboid_id) to a set of manually annotated taxonomic variables.</li> <li>A directory tree of the form: `<series_id>/<tuboid_id>/`. Each terminal directory contains: <ul> <li> <ul> <li>`tuboid.jpg` – a jpeg image made of 224 x 224 tiles representing all the shots in a tuboid, left to right, top to bottom – might be padded with empty images</li> <li>`metadata.txt` – a csv text file with columns: <ul> <li> <ul> <li>parrent_image_id – <device>.<UTC_datetime></li> <li>X – the X coordinates of the object centroid</li> <li>Y – the Y coordinates of the object centroid</li> </ul> </li> </ul> </li> <li>scale – The scaling factor applied between the original and image and the 224 x 224 tile (>1 => image was enlarged)</li> <li>`context.jpg` – a representation of the first whole image of a series, with a box around the first tuboid shot (this is for debugging/labelling purposes)</li> </ul> </li> </ul> </li> </ul> </li> <li>`output/` <ul> <li>`model_final.pth` – the model as trained for the publication</li> </ul> </li> <li>config/ <ul> <li>`config.yaml` – The configuration file defining the hyperparameters to train the model as well as the taxonomic labels</li> </ul> </li> </ul>
Capturing functional relations in fluid-structure interaction via machine learning
<p>While fluid-structure interaction (FSI) problems are ubiquitous in various applications from cell-biology to aerodynamics, they involve huge computational overhead. In this paper, we adopt a machine learning (ML)-based strategy to bypass the detailed FSI analysis that requires cumbersome simulations in solving the Navier-Stokes (N-S) equations. To mimic the effect of fluid on an immersed beam, we have introduced dissipation into the beam model with time-varying forces acting on it. The forces in a discretized setup have been decoupled via an appropriate linear algebraic operation, which generates the ground truth force/moment data for the ML analysis. The adopted ML technique, symbolic regression, generates computationally tractable functional forms to represent the force/moment with respect to space and time. These estimates are fed into the dissipative beam model to generate the immersed beam's deflections over time, which are in conformity with the detailed FSI solutions. Numerical results demonstrate that the ML-estimated continuous force and moment functions are able to accurately predict the beam deflections under different discretizations.</p>
A new remote sensing benchmark dataset for machine learning applications : MultiSenGE
<p>[UPDATE] You can now access MultiSen (GE and NA) collection though this portal : <a href="https://doi.theia.data-terra.org/ai4lcc/?lang=en">https://doi.theia.data-terra.org/ai4lcc/?lang=en</a></p> <p>MultiSenGE is a new large-scale multimodal and multitemporal benchmark dataset covering one of the biggest administrative region located in the Eastern part of France. It contains 8,157 patches of 256 * 256 pixels for Sentinel-2 L2A, Sentinel-1 GRD and a regional LULC topographic regional database. </p> <p>Every file has a specific nomenclature :</p> <ul> <li>Sentinel-1 patches: {tile}_{date}_S1_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>Sentinel-2 patches: {tile}_{date}_S2_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>Ground reference patches: {tile}_GR_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>JSON Labels: {tile}_{x-pixel-coordinate}_{y-pixel-coordinate}.json</li> </ul> <p>where <em>tile</em> is the Sentinel-2 tile number, <em>date</em> the date of acquisition of the patch, <em>x-pixel-coordinate</em> and <em>y-pixel-coordinate</em> are the coordinates of the patch in the tile.</p> <p>In addition, you can find a set of useful python tools for extracting information about the dataset on Github : <a href="https://github.com/r-wenger/MultiSenGE-Tools">https://github.com/r-wenger/MultiSenGE-Tools</a></p> <p>First experiments based on this <em>dataset</em> is in press in ISPRS Annals : <strong>Wenger, R., </strong>Puissant, A., Weber, J., Idoumghar, L., and Forestier, G.: MULTISENGE: A MULTIMODAL AND MULTITEMPORAL BENCHMARK DATASET FOR LAND USE/LAND COVER REMOTE SENSING APPLICATIONS, ISPRS Ann. Photogramm. Remote Sens. Spatial Inf. Sci., V-3-2022, 635–640, https://doi.org/10.5194/isprs-annals-V-3-2022-635-2022, 2022.</p> <p>Due to the large size of the dataset, you will only find the associated JSON files on this Zenodo repository. To download the Sentinel-1, Sentinel-2 patches and the reference data, please do so via these links: </p> <ul> <li>Sentinel-1 temporal serie patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/s1.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/s1.tgz</a></li> <li>Sentinel-2 temporal serie patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/s2.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/s2.tgz</a></li> <li>Ground reference patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/ground_reference.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/ground_reference.tgz</a></li> <li>JSON files for each patch: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/labels.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/labels.tgz</a></li> </ul>
Studying the Practices of Deploying Machine Learning Projects on Docker
<p>This repository contains the dataset for our study titled above:</p> <p>Below is the abstract:</p> <p>Docker is a containerization service that allows for convenient deployment of websites, databases, applications' APIs, and machine learning (ML) models with a few lines of code. Studies have recently explored the use of Docker for deploying general software projects with no specific focus on how Docker is being used to deploy ML-based projects. In this study, we conducted an exploratory study to understand how Docker is being used to deploy ML-based projects. As the initial step, we examined the categories of ML-based project that use Docker. We then examined why and how these projects use Docker, and the characteristics of the resulting Docker images. Our results indicate that six categories of ML-based projects use Docker for deployment, including ML Applications, MLOps/ AIOps, Tookits, DL Frameworks, Models, and Documentation. We derived the taxonomy of 21 major categories representing the purposes of using Docker, including those specific to models such as model management tasks (e.g., testing, training), data management (e.g., migration, persistent storage, sharing), software testing, interactive development. We then showed that ML engineers use Docker images mostly to help with the platform portability, such as transferring the software across the operating systems, runtimes such as GPU-accelerated, and language constraints. However, we also found that more resources may be required to run the Docker images for building ML-based software projects due to the large number of files contained in the image layers with deeply nested directories. Through this study we hope to shed light on the emerging practices of deploying ML software projects using containers and highlight aspects that should be improved.</p>
Development of a machine learning model to predict non- durable response to anti-TNF therapy in Crohn's disease using transcriptome imputed from genotypes
<p>This is the expression value predicted using PrediXcan version 7 to find a gene feature that can distinguish between patients with and without effect on infliximab.</p> <p>Among the various tissue models provided by PrediXcan v7, three models were selected and used: whole blood, Colon transverse, and terminal ileum of small intestine, and the predicted gene counts of each model were 6,294, 5,612 and 3,107.</p> <p>For each of the three models, predicted gene expression values and phenotype information per sample were submitted.</p>
Design of experiment (DOE) used in the study: Lightweight design of variable-stiffness imperfection-insensitive cylinders enabled by continuous tow shearing and machine learning
<p>There are five input variables that are changed for this design of experiment (DOE) within the following range:</p> <p><span class="math-tex">\(\begin{eqnarray} 0.05 \leq r_{CTS} \leq 0.20 \nonumber \\ 1 \leq n \leq 12 \nonumber \\ 0 \leq {c_2}_{ratio} \leq 1 \\ 0 \leq \theta_1 \leq 75 \nonumber \\ 0 \leq \theta_2 \leq 75 \nonumber \end{eqnarray}\)</span></p> <p>For the sake of simplicity, these variables are respectively called v1, v2, v3, v4, v5.</p> <p>The DOE consist of 2000 points created with Latin Hyper-cube Sampling, as available in the LHS toolbox (Carnell, R. lhs: Latin Hypercube Samples, 2021. R package version 1.1.3.).</p> <p>The outputs evaluated with this design of experiment (DOE) are the critical buckling load $P_{critical}$, $b_{factor}$ obtained with Koiter's asymptotic approach, and the mass.</p>
TocoDecoy: a new approach to design unbiased datasets for training and benchmarking machine-learning scoring functions
<p>This dataset file contains TocoDecoy datasets generated based on the targets and active ligands of LIT-PCBA.</p> <p>1_property_filtered.zip :</p> <ul> <li>TD set: the ligand file name, 2D T-sne vectors, Smiles, molecular weight (MW), Wildman-Crippen partition coefficient (log P), number of rotatable bonds (RB), number of hydrogen-bond acceptors (HBA), number of hydrogen-bond donors (HBD), number of halogens (HAL), topology similarities of decoys to the seed active ligands, active label (active or inactive) and training set label (whether belongs to training set or test set) <strong>OF active ligands and their topologically dissimilar decoys</strong></li> <li>CD set: the decoy conformations with low docking scores generated by docking active ligands into protein pockets using Glide, Schrödinger.</li> </ul> <p> </p>
To what extent naringenin binding and membrane depolarization shape mitoBK channel gating - a machine learning approach (code and dataset)
<p>The dataset consists of dwell-time series (sampling frequency 100 kHz) of the mitoBK ion channel activation modulated by the naringenin binding and membrane<br> depolarization. It also contains the code written in Python, with the use of tslearn and scikit-learn packages, classifying the dwell-time subseries into right categories.</p> <p>The dataset is organized as follows. The mitoBK_ML.zip directory consists of two directories:</p> <ol> <li><strong>dwell times </strong>containing 5 subdirectories comprising groups of dwell-time subseries obtained at different pipette potentials and naringenin concentration. First number in the name od directory stands for the applied voltage in mV, whilst the second one denotes the naringenin concentration in µmol. For instance, directory named 20_3 means that the obtained dwell-times series were obtained at 20 mV (value of pipette potential) and 3 µmol (concentration of naringenin). These subdirectories are named as follows:</li> </ol> <ul> <li><strong>1group </strong>comprising dwell time series <strong>20_3, 40_1, 60_0</strong></li> <li><strong>2group</strong> comprising dwell-time series <strong>20_10, 60_1</strong></li> <li><strong>3group</strong> comprising dwell-time series <strong>40_10</strong>, <strong>60_3</strong></li> <li><strong>naringenina</strong> comprising dwell-time series <strong>60_0, 60_10</strong></li> <li><strong>voltage</strong> comprising dwell-time series <strong>20_10, 60_10</strong></li> </ul> <p><strong>1group, 2group and 3group</strong> contain the dwell-time series with approximately the same value of open-state probability of the ion channel.</p> <p>The <strong>naringenina</strong> contains the dwell-time series with the same value of potential (60 mV) and different values of naringenin concentration (0 µmol and 10 µmol). </p> <p>The <strong>voltage </strong>contains the dwell-time series with the same value of naringenin concentration (10 µmol) and different values of applied voltage (20 mV and 60 mV).</p> <p> 2. <strong>rslt </strong>is organized analogously to <strong>dwell times. </strong>The subdirectories are empty, but they will be filled with the results after launching the Python scripts placed in the <strong>knn_ion_channel.ipynb</strong> or <strong>shapelet_ion_channel.ipynb </strong>files.</p> <p>The Python code is placed in two files:</p> <ol> <li><strong>knn_ion_channel.ipynb </strong>containing kNN (<em>k-Nearest Neighbors</em>) algorithm classifying dwell-time series belonging to one of 5 different categories enumerated above: <strong>1group, 2group, 3group, naringenina, voltage</strong>. More detailed description of the code can be found inside uploaded Jupyter notebook.</li> <li><strong>shapelet_ion_channel.ipynb </strong>containing <em>shapelet-learning algorithm</em> classifying dwell-time series belonging to one of 5 different categories enumerated above. <strong>1group, 2group, 3group, naringenina, voltage. </strong>More detailed description of the code can be found inside uploaded Jupyter notebook.</li> </ol> <p> </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.