Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
285
datasets available to search
ShareScore release 0.9.0
Dataset results
285 results for “machine learning dataset”
Machine Learning Dataset for Poultry Diseases Diagnostics - PCR annotated
<p>The dataset of poultry disease diagnostics was annotated using Polymerase Chain Reaction (PCR). Polymerase Chain Reaction (PCR) is a molecular biology technique for rapid diagnostics. We gathered both the fecal images and fecal samples from layers, cross and indigenous breeds of chicken from poultry farms in Arusha and Kilimanjaro regions in Tanzania between September 2020 and February 2021. Each fecal sample collected was coded to its corresponding image during data collection. PCR method is used for detection and identification of pathogens through amplification of DNA sequences unique to the pathogen. We used existing primers from literature to amplify the target DNA/RNA on the poultry fecal samples for PCR. The targets were Coccidiosis, Newcastle disease and Salmonella. We used the primers for PCR diagnostics at the molecular laboratory of the Nelson Mandela African Institution of Science and Technology (NM-AIST). The fecal samples were stored at -80 degrees celsius. The PCR diagnostics were conducted using reagents and kits from Zymo Research and the protocol is summarized in these five stages: 1. DNA sample loading 2. DNA extraction 3. Amplification; 4. Quantification and 5. Detection.</p> <p>All the PCR annotated fecal images are in the <strong><strong>.zip files</strong></strong>; “pcrcocci.zip” has 373 images, “pcrhealthy.zip” has 347 images, “pcrsalmo.zip” has 349 images, "pcrncd.zip" has 186 images. A total of 1,255 image files are labeled.</p> <p>The research project is funded by the Organization for Women in Science for the Developing World (OWSD) with Grant Award Number: 4500406715.</p>
A Dataset for Utility Prediction in Computational Persuasion with Machine Learning Techniques
<p>This dataset contains data for a new benchmark for the prediction of user's utilities with Machine Learning techniques for Computational Persuasion. This work has been accepted at AAAI-22, more information in the relative repository containing the source code: <a href="https://github.com/ivanDonadello/ML-Argument-Based-Computational-Persuasion">https://github.com/ivanDonadello/ML-Argument-Based-Computational-Persuasion</a></p>
Architectural Design Decisions for Machine Learning Deployment: Dataset and Code
<p><strong>Title:</strong> Architectural Design Decisions for Machine Learning Deployment: Dataset and Code</p> <p><strong>Authors:</strong> Stephen John Warnett; Uwe Zdun</p> <p><strong>About:</strong> This is the dataset and code artefact for the paper entitled "Architectural Design Decisions for Machine Learning Deployment".</p> <p><strong>Contents:</strong> The "_generated" directory contains the generated results, including latex files with tables for use in publications and the Architectural Design Decision model in textual and graphical form. "Generators" contains Python applications that can be run to generate the above. "Metamodels" contains a Python file with type definitions. "Sources_coding" contains our source codings and audit trail. "Add_models" contains the Python implementation of our model and source codings. Finally, "appendix" contains a detailed description of our research method.</p> <p><strong>Paper Abstract:</strong> Deploying machine learning models to production is challenging, partially due to the misalignment between software engineering and machine learning disciplines but also due to potential practitioner knowledge gaps. To reduce this gap and guide decision-making, we conducted a qualitative investigation into the technical challenges faced by practitioners based on studying the grey literature and applying the Straussian Grounded Theory research method. We modelled current practices in machine learning, resulting in a UML-based architectural design decision model based on current practitioner understanding of the domain and a subset of the decision space and identified seven architectural design decisions, various relations between them, twenty-six decision options and forty-four decision drivers in thirty-five sources. Our results intend to help bridge the gap between science and practice, increase understanding of how practitioners approach the deployment of their solutions, and support practitioners in their decision-making.</p> <p><strong>Objective:</strong> This paper aims to study current practitioner understanding of architectural concepts associated with machine learning deployment.</p> <p><strong>Method:</strong> Applying Straussian Grounded Theory to gray literature sources containing practitioner views on machine learning practices, we studied methods and techniques currently applied by practitioners in the context of machine learning solution development and gained valuable insights into the software engineering and architectural state of the art as applied to ML.</p> <p><strong>Results:</strong> Our study resulted in a model of Architectural Design Decisions, practitioner practices, and decision drivers in the field of software engineering and software architecture for machine learning.</p> <p><strong>Conclusions:</strong> The resulting Architectural Design Decisions model can help researchers better understand practitioners' needs and the challenges they face, and guide their decisions based on existing practices. The study also opens new avenues for further research in the field, and the design guidance provided by our model can also help reduce design effort and risk. In future work, we plan on using our findings to provide automated design advice to machine learning engineers.</p>
Dataset of the paper "Machine learning for expert-level image-based identification of very similar species in the hyperdiverse plant bug family Miridae (Hemiptera: Heteroptera)"
<p>This dataset contains 3792 images of 26 plant bug (Insecta: Heteroptera: Miridae: Mirini) species used to test the performance of a CNN in species recognition. All jpg files are 1920 pixels on the long size and additionally available as an archive file to facilitate download of the entire dataset. </p> <p>Bar code labels (unique specimen identifiers or USIs) were attached to all examined specimens used for this study. Further information such as additional photographs of habitus and genitalic structures, georeferenced coordinates of each locality, specimens dissected, notes, collecting method can be obtained from the Heteroptera Species Pages (http://research.amnh.org/pbi/heteropteraspeciespage/) which assembles available data from a specimen database and are also provided as an Excel spreadsheet (file _Adelphocoris_CNN_label_data.xlsx).</p>
TimeSpec4LULC: A Smart-Global Dataset of Multi-Spectral Time Series of MODIS Terra-Aqua from 2000 to 2021 for Training Machine Learning models to perform LULC Mapping
<p>TimeSpec4LULC is a smart open-source global dataset of multi-spectral time series for 29 Land Use and Land Cover (LULC) classes ready to train machine learning models. It was built based on the seven spectral bands of the MODIS sensors at 500 m resolution from 2000 to 2021 (262 observations in each time series). Then, was annotated using spatial-temporal agreement across the 15 global LULC products available in Google Earth Engine (GEE).</p> <p>TimeSpec4LULC contains two datasets: the original dataset distributed over 6,076,531 pixels, and the balanced subset of the original dataset distributed over 29000 pixels.</p> <p>The original dataset contains 30 folders, namely "Metadata", and 29 folders corresponding to the 29 LULC classes. The folder "Metadata" holds 29 different CSV files describing the metadata of the 29 LULC classes. The remaining 29 folders contain the time series data for the 29 LULC classes. Each folder holds 262 CSV files corresponding to the 262 months. Inside each CSV file, we provide the seven values of the spectral bands as well as the coordinates for all the LULC class-related pixels.</p> <p>The balanced subset of the original dataset contains the metadata and the time series data for 1000 pixels per class representative of the globe. It holds 29 different JSON files following the names of the 29 LULC classes.</p> <p>The features of the dataset are:</p> <p>- ".geo": the geometry and coordinates (longitude and latitude) of the pixel center.</p> <p>- "ADM0_Code": the GAUL country code.</p> <p>- "ADM1_Code": the GAUL first-level administrative unit code.</p> <p>- GHM_Index": the average of the global human modification index.</p> <p>- "Products_Agreement_Percentage": the agreement percentage over the 15 global LULC products available in GEE.</p> <p>- "Temporal_Availability_Percentage": the percentage of non-missing values in each band.</p> <p>- "Pixel_TS": the time series values of the seven spectral bands.</p>
Machine Learning based scratches on printed paper detection, in high-speed printing systems [Dataset]
<p>Printing industry rapidly is adopting digital technologies and the requirements in terms of speed and print quality are also becoming more demanding. The is a wide range of possible quality defects in printed paper. This makes it impossible to have humans inspect the printed paper for such a big amount of possible quality defects at the high-speeds the printouts are produced.</p> <p>Printing industry is not taking advantage of the Artificial Intelligence to detect defects in printed paper at speed without human intervention. It is possible to generate millions of images (captures) with printed content from a printing system every day. Most of these images will not have any defect but some other will and can be used to generate a data set to be used in a machine learning system.</p> <p>The intention of this research work is to find ways artificial intelligence can help on automatically detecting defects on printed paper in a printing system and classifying them, without human intervention. Focusing on scratches, I’ve explored what are the actual proposals and solutions, and how machine learning can help improving them by using datasets with different techniques, implementing possible solutions and comparing the obtained results.</p>
Home-based measurements of dystonia and choreoathetosis in cerebral palsy using smartphone-coupled inertial sensor technology and machine learning: A proof-of-concept study - dataset
<p>Home-based measurements of dystonia in cerebral palsy using smartphone-coupled inertial sensor technology and machine learning: A proof-of-concept study</p> <p> </p> <p>This project contains:</p> <p>- 1 main MATLAB script: MODYSathome_main.m<br> - 12 MATLAB functions:<br> - function_calc_mean_recall_precision.m<br> - function_create_dataframes.m<br> - function_deep_learning.m<br> - function_determine_best_ML_model.m<br> - function_display_DL_results.m<br> - function_display_ML_results.m<br> - function_index_extremities.m<br> - function_machine_learning.m<br> - function_oversample.m<br> - function_partition_data.m<br> - function_pick_best_models.m<br> - function_prepare_DL_data.m</p> <p>Downloading the Matlab scripts</p> <p> - Create a folder named 'MODYS' and create a subfolder named 'results'<br> - Download the zip file via <a href="https://zenodo.org/record/6379348">RehabAUmc/modys-at-home: v1.0 | Zenodo</a><br> - Unzip the zip file in the path MODYS\</p> <p>STEPS<br> 1. Open MATLAB<br> 2. In MATLAB, go to the 'HOME' tab and click on 'Set Path'<br> 3. Click on 'Add Folder' and browse to MODYS/RehabAUmc-modys-at-home-86b14c3/functions<br> 4. Click on 'Select Folder' and click on 'Save'<br> 5. Click on 'Browse to folder' and browse to a patients' data in MODYS/data/PatientXXX, then click on 'Select Folder'<br> 6. In the 'HOME' tab click on 'Open' and open MODYSathome.m in MODYS/RehabAUmc-modys-at-home-86b14c3<br> 7. In the 'EDITOR' tab click on 'Run Section' to run the script<br> 8. When the code has been run, the results are displayed in the Command Window and saved in MODYS/results/PatientXXX</p>
A new remote sensing benchmark dataset for machine learning applications : MultiSenGE
<p>[UPDATE] You can now access MultiSen (GE and NA) collection though this portal : <a href="https://doi.theia.data-terra.org/ai4lcc/?lang=en">https://doi.theia.data-terra.org/ai4lcc/?lang=en</a></p> <p>MultiSenGE is a new large-scale multimodal and multitemporal benchmark dataset covering one of the biggest administrative region located in the Eastern part of France. It contains 8,157 patches of 256 * 256 pixels for Sentinel-2 L2A, Sentinel-1 GRD and a regional LULC topographic regional database. </p> <p>Every file has a specific nomenclature :</p> <ul> <li>Sentinel-1 patches: {tile}_{date}_S1_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>Sentinel-2 patches: {tile}_{date}_S2_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>Ground reference patches: {tile}_GR_{x-pixel-coordinate}_{y-pixel-coordinate}.tif</li> <li>JSON Labels: {tile}_{x-pixel-coordinate}_{y-pixel-coordinate}.json</li> </ul> <p>where <em>tile</em> is the Sentinel-2 tile number, <em>date</em> the date of acquisition of the patch, <em>x-pixel-coordinate</em> and <em>y-pixel-coordinate</em> are the coordinates of the patch in the tile.</p> <p>In addition, you can find a set of useful python tools for extracting information about the dataset on Github : <a href="https://github.com/r-wenger/MultiSenGE-Tools">https://github.com/r-wenger/MultiSenGE-Tools</a></p> <p>First experiments based on this <em>dataset</em> is in press in ISPRS Annals : <strong>Wenger, R., </strong>Puissant, A., Weber, J., Idoumghar, L., and Forestier, G.: MULTISENGE: A MULTIMODAL AND MULTITEMPORAL BENCHMARK DATASET FOR LAND USE/LAND COVER REMOTE SENSING APPLICATIONS, ISPRS Ann. Photogramm. Remote Sens. Spatial Inf. Sci., V-3-2022, 635–640, https://doi.org/10.5194/isprs-annals-V-3-2022-635-2022, 2022.</p> <p>Due to the large size of the dataset, you will only find the associated JSON files on this Zenodo repository. To download the Sentinel-1, Sentinel-2 patches and the reference data, please do so via these links: </p> <ul> <li>Sentinel-1 temporal serie patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/s1.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/s1.tgz</a></li> <li>Sentinel-2 temporal serie patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/s2.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/s2.tgz</a></li> <li>Ground reference patches: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/ground_reference.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/ground_reference.tgz</a></li> <li>JSON files for each patch: <a href="https://s3.unistra.fr/a2s_datasets/MultiSenGE/labels.tgz">https://s3.unistra.fr/a2s_datasets/MultiSenGE/labels.tgz</a></li> </ul>
TocoDecoy: a new approach to design unbiased datasets for training and benchmarking machine-learning scoring functions
<p>This dataset file contains TocoDecoy datasets generated based on the targets and active ligands of LIT-PCBA.</p> <p>1_property_filtered.zip :</p> <ul> <li>TD set: the ligand file name, 2D T-sne vectors, Smiles, molecular weight (MW), Wildman-Crippen partition coefficient (log P), number of rotatable bonds (RB), number of hydrogen-bond acceptors (HBA), number of hydrogen-bond donors (HBD), number of halogens (HAL), topology similarities of decoys to the seed active ligands, active label (active or inactive) and training set label (whether belongs to training set or test set) <strong>OF active ligands and their topologically dissimilar decoys</strong></li> <li>CD set: the decoy conformations with low docking scores generated by docking active ligands into protein pockets using Glide, Schrödinger.</li> </ul> <p> </p>
To what extent naringenin binding and membrane depolarization shape mitoBK channel gating - a machine learning approach (code and dataset)
<p>The dataset consists of dwell-time series (sampling frequency 100 kHz) of the mitoBK ion channel activation modulated by the naringenin binding and membrane<br> depolarization. It also contains the code written in Python, with the use of tslearn and scikit-learn packages, classifying the dwell-time subseries into right categories.</p> <p>The dataset is organized as follows. The mitoBK_ML.zip directory consists of two directories:</p> <ol> <li><strong>dwell times </strong>containing 5 subdirectories comprising groups of dwell-time subseries obtained at different pipette potentials and naringenin concentration. First number in the name od directory stands for the applied voltage in mV, whilst the second one denotes the naringenin concentration in µmol. For instance, directory named 20_3 means that the obtained dwell-times series were obtained at 20 mV (value of pipette potential) and 3 µmol (concentration of naringenin). These subdirectories are named as follows:</li> </ol> <ul> <li><strong>1group </strong>comprising dwell time series <strong>20_3, 40_1, 60_0</strong></li> <li><strong>2group</strong> comprising dwell-time series <strong>20_10, 60_1</strong></li> <li><strong>3group</strong> comprising dwell-time series <strong>40_10</strong>, <strong>60_3</strong></li> <li><strong>naringenina</strong> comprising dwell-time series <strong>60_0, 60_10</strong></li> <li><strong>voltage</strong> comprising dwell-time series <strong>20_10, 60_10</strong></li> </ul> <p><strong>1group, 2group and 3group</strong> contain the dwell-time series with approximately the same value of open-state probability of the ion channel.</p> <p>The <strong>naringenina</strong> contains the dwell-time series with the same value of potential (60 mV) and different values of naringenin concentration (0 µmol and 10 µmol). </p> <p>The <strong>voltage </strong>contains the dwell-time series with the same value of naringenin concentration (10 µmol) and different values of applied voltage (20 mV and 60 mV).</p> <p> 2. <strong>rslt </strong>is organized analogously to <strong>dwell times. </strong>The subdirectories are empty, but they will be filled with the results after launching the Python scripts placed in the <strong>knn_ion_channel.ipynb</strong> or <strong>shapelet_ion_channel.ipynb </strong>files.</p> <p>The Python code is placed in two files:</p> <ol> <li><strong>knn_ion_channel.ipynb </strong>containing kNN (<em>k-Nearest Neighbors</em>) algorithm classifying dwell-time series belonging to one of 5 different categories enumerated above: <strong>1group, 2group, 3group, naringenina, voltage</strong>. More detailed description of the code can be found inside uploaded Jupyter notebook.</li> <li><strong>shapelet_ion_channel.ipynb </strong>containing <em>shapelet-learning algorithm</em> classifying dwell-time series belonging to one of 5 different categories enumerated above. <strong>1group, 2group, 3group, naringenina, voltage. </strong>More detailed description of the code can be found inside uploaded Jupyter notebook.</li> </ol> <p> </p> <p> </p> <p> </p>
ImageNet-Patch: A Dataset for Benchmarking Machine Learning Robustness against Adversarial Patches
<p>Adversarial patches are optimized contiguous pixel blocks in an input image that cause a machine-learning model to misclassify it. However, their optimization is computationally demanding and requires careful hyperparameter tuning. To overcome these issues, we propose ImageNet-Patch, a dataset to benchmark machine-learning models against adversarial patches. It consists of a set of patches optimized to generalize across different models and applied to ImageNet data after preprocessing them with affine transformations. This process enables an approximate yet faster robustness evaluation, leveraging the transferability of adversarial perturbations.</p> <p>We release our dataset as a set of folders indicating the patch target label (e.g., `banana`), each containing 1000 subfolders as the ImageNet output classes.</p> <p>An example showing how to use the dataset is shown below.</p> <pre><code class="language-python"># code for testing robustness of a model import os.path from torchvision import datasets, transforms, models import torch.utils.data class ImageFolderWithEmptyDirs(datasets.ImageFolder): """ This is required for handling empty folders from the ImageFolder Class. """ def find_classes(self, directory): classes = sorted(entry.name for entry in os.scandir(directory) if entry.is_dir()) if not classes: raise FileNotFoundError(f"Couldn't find any class folder in {directory}.") class_to_idx = {cls_name: i for i, cls_name in enumerate(classes) if len(os.listdir(os.path.join(directory, cls_name))) > 0} return classes, class_to_idx # extract and unzip the dataset, then write top folder here dataset_folder = 'data/ImageNet-Patch' available_labels = { 487: 'cellular telephone', 513: 'cornet', 546: 'electric guitar', 585: 'hair spray', 804: 'soap dispenser', 806: 'sock', 878: 'typewriter keyboard', 923: 'plate', 954: 'banana', 968: 'cup' } # select folder with specific target target_label = 954 dataset_folder = os.path.join(dataset_folder, str(target_label)) normalizer = transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]) transforms = transforms.Compose([ transforms.ToTensor(), normalizer ]) dataset = ImageFolderWithEmptyDirs(dataset_folder, transform=transforms) model = models.resnet50(pretrained=True) loader = torch.utils.data.DataLoader(dataset, shuffle=True, batch_size=5) model.eval() batches = 10 correct, attack_success, total = 0, 0, 0 for batch_idx, (images, labels) in enumerate(loader): if batch_idx == batches: break pred = model(images).argmax(dim=1) correct += (pred == labels).sum() attack_success += sum(pred == target_label) total += pred.shape[0] accuracy = correct / total attack_sr = attack_success / total print("Robust Accuracy: ", accuracy) print("Attack Success: ", attack_sr) </code></pre> <p> </p>
DATASET - Improving Remote Sensing of Extreme Events with Machine Learning: Application to IASI LST Retrievals
<p>Data for experiments presented in the paper "Improving Remote Sensing of Extreme Events with Machine Learning: Application to IASI LST Retrievals" </p>
Global input datasets for use in constraints on global seafloor biogenic methane production from deterministic and machine learning modeling
<p>This dataset includes 9 grids used as model input for manuscript "Constraints on global seafloor biogenic methane production from deterministic and machine learning modeling". Additionally, there are four grids (heat flow, total organic carbon, porosity, and crust age) for which variable uncertainty was given.</p> <p>Grids here are available in xyz (longitude in decimal degrees, latitude in decimal degrees, and variable) ascii file format. Each reference is below is the grids native reference. For more information on the creation of these grids please visit the main manuscript.</p> <p>Below are respective file names and variable name/units:</p> <p>Dataset 1: Elevation in Meters (+ indicates above sea level, - below sea level)</p> <p>Tozer, B., Sandwell, D. T., Smith, W. H. F., Olson, C., Beale, J. R., & Wessel, P. (2019). Global bathymetry and topography at 15 arc sec: SRTM15+. <em>Earth and Space Science</em>, 6. https://doi.org/10.1029/ 2019EA000658</p> <p>Dataset 2: Seawater Density in Kilograms per Cubic Meter</p> <p>Boyer, T. P., Antonov, J. I., Baranova, O. K., Garcia, H. E., Johnson, D. R., Mishonov, A. V., … Grodsky, A. (2013). World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), Technical Ed.; <em>NOAA Atlas NESDIS</em> 72 (pp. 209).</p> <p>Dataset 3: Seawater Temperature in Degrees Celcius </p> <p>Boyer, T. P., Antonov, J. I., Baranova, O. K., Garcia, H. E., Johnson, D. R., Mishonov, A. V., … Grodsky, A. (2013). World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), Technical Ed.; <em>NOAA Atlas NESDIS</em> 72 (pp. 209).</p> <p>Dataset 4: Seawater Salinity in Percent Salinity Units</p> <p>Boyer, T. P., Antonov, J. I., Baranova, O. K., Garcia, H. E., Johnson, D. R., Mishonov, A. V., … Grodsky, A. (2013). World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), Technical Ed.; <em>NOAA Atlas NESDIS</em> 72 (pp. 209).</p> <p>Dataset 5: Heat Flow in Milliwatts per Square Meter</p> <p>Global Heat Flow Compilation Group (2013). Component parts of the World Heat Flow Data Collection. <em>PANGAEA</em>, https://doi.org/10.1594/PANGAEA.810104</p> <p>Hornbach, M. J., Harris, R. N. & Phrampus, B. J. (2020). Heat flow on the U.S. Beaufort Margin, Arctic Ocean: Implications for ocean warming, methane hydrate stability, and regional tectonics. <em>Geochemistry, Geophysics, Geosystems</em>, 21(5). e2020GC008933. https://doi.org/10.1029/2020GC008933</p> <p>Dataset 6: Sediment Thickness in Meters</p> <p>Straume, E. O., Gaina, C., Medvedev, S., Hochmuth, K., Gohl, K., Whittaker, J. M., … Hopper, J. R. (2019). GlobSed: updated total sediment thickness in the world’s oceans. <em>Geochemistry, Geophysics, Geosystems</em>, 20(4), 1756–1772.</p> <p>Dataset 7: Seafloor Porosity in Fraction</p> <p>Martin, K. M., Wood, W. T., & Becker, J. J. (2015). A global prediction of seafloor sediment porosity using machine learning. <em>Geophysical Research Letters</em>, 42(24), 2015GL065279. https://doi.org/10.1002/2015GL065279</p> <p>Dataset 8: Seafloor Total Organic Carbon in Percent Dry Weight</p> <p>Lee, T.R., Wood, W.T., & Phrampus, B.J. (2019). A machine learning (kNN) approach to predicting global seafloor total organic carbon. <em>Global Biogeochemical Cycles</em>. 33, 37–46, doi:10.1029/2018GB005992.</p> <p>Dataset 9: Crust Age in Million Years</p> <p>Müller, R. D., Sdrolias, M., Gaina, C., & Roest, W. R. (2008). Age, spreading rates, and spreading asymmetry of the world’s ocean crust. <em>Geochemistry, Geophysics, Geosystems</em>, 9, Q04006. https://doi.org/10.1029/2007GC001743</p> <p>Dataset 10: Seafloor Porosity Uncertainty in Fraction</p> <p>Dataset 11: Seafloor Total Organic Carbon Uncertainty in Percent Dry Weight</p> <p>Lee, T.R., Wood, W.T., & Phrampus, B.J. (2019). A machine learning (kNN) approach to predicting global seafloor total organic carbon. <em>Global Biogeochemical Cycles</em>. 33, 37–46, doi:10.1029/2018GB005992.</p> <p>Dataset 12: Heat Flow Uncertainty in Milliwatts per Square Meter</p> <p>Dataset 13: Crust Age Uncertainty in Million Years</p> <p>Müller, R. D., Sdrolias, M., Gaina, C., & Roest, W. R. (2008). Age, spreading rates, and spreading asymmetry of the world’s ocean crust. <em>Geochemistry, Geophysics, Geosystems</em>, 9, Q04006. https://doi.org/10.1029/2007GC001743</p>
Dataset for Integrated hydrodynamic and machine learning models
<p>The dataset is the supplement to our publication in <a href="https://www.nonlinear-processes-in-geophysics.net/">Nonlinear Processes in Geophysics</a> (https://doi.org/10.5194/npg-2021-36). To use this data, please give us credit by citing our article.</p>
The Dataset of Quantifying Alignment Deviations for Uniaxial Material Mechanical Testing via Automated Machine Learning
<p>The dataset consists of 4 alignment deviations of the uniaxial testing machine as well as 12 strain measurement points on cruciform specimens. A deep learning model is trained on the dataset to quantify 4 alignment deviations using 12 strain values on a thin plate specimen. The design of experiments includes Optimal Latin Hypercube, numerical modelling of Finite Element Methods. Using the Optimal Latin Hypercube, 12496 distinct groups of DOE simulation tests are constructed. Under the boundary conditions of 4 distinct deviations, 12 strain values at the required location on the cruciform specimen are obtained using Python scripts.</p> <p>The nine CSV files correspond to the nine analysis steps. The only difference among the nine analysis steps is the pretension force acting on RP1. Each CSV file contains 24 columns of data, and the corresponding contents of each column of data are as follows:</p> <ul> <li>Columns 1-6 are the freedoms of RP1 reference point, which are U1, U2, U3, ur1, UR2 and UR3 respectively;</li> <li>Columns 7-12 are the freedoms of RP2 reference points, which are U1, U2, U3, ur1, UR2 and UR3 respectively;</li> <li>Columns 13-24 are the strain values of the last 12 strain measurements of the thin plate rectangular specimen。</li> </ul>
Datasets from "Electrostatic Embedding of Machine Learning Potentials"
<p>Data required to reproduce results in "Electrostatic Embedding of Machine Learning Potentials" <a href="https://doi.org/10.26434/chemrxiv-2022-rknwt">article</a>. See <a href="https://github.com/emedio/embedding">https://github.com/emedio/embedding</a> for details.</p> <ul> <li>QM7_B3LYP_cc-pVTZ.tgz - outputs of single point B3LYP/cc-pVTZ calculations of structures in <a href="http://quantum-machine.org/data/qm7.mat">QM7 dataset</a> with ORCA 5. Include molecular dipolar polarizabilities.</li> <li>QM7_B3LYP_cc-pVTZ_horton.tgz - MBIS partitioning of the B3LYP/cc-pVTZ densities with <a href="https://github.com/theochem/horton">Horton 2.1.0</a>.</li> <li>mpro_xyz.tgz - coordinates of the ligand and surrounding point charges from 100 snapshots of SARS-CoV-2 Mpro complex with PF-00835231.</li> <li>mpro_*.tgz - DFT and semiempirical single point calculations with ORCA 5 for the coordinates from mpro_xyz.tgz, <em>in vacuo </em>and in presence of point charges.</li> <li>mlmm.mat - learned parameters and SOAP feature vectors of reference atomic environments</li> </ul>
CASM: A long-term Consistent Artificial-intelligence based Soil Moisture dataset based on machine learning and remote sensing
<p>Paper to cite: Skulovich, O., Gentine, P. A Long-term Consistent Artificial Intelligence and Remote Sensing-based Soil Moisture Dataset. <em>Sci Data</em> 10, 154 (2023). https://doi.org/10.1038/s41597-023-02053-x</p> <p> </p> <p>The Consistent Artificial Intelligence (AI)-based Soil Moisture (CASM) dataset is a global, consistent, and long-term, remote sensing soil moisture (SM) dataset created using machine learning. It is based on the NASA Soil Moisture Active Passive (SMAP) satellite mission SM data as a target and is aimed at extrapolating SMAP-like quality SM data back in time with previous satellite microwave platforms. Machine learning approach, such as neural network (NN) has the advantage of being both nonlinear, and state-dependent, and naturally imposing a global distribution matching between the source and the target data. Utilizing this, the new CASM dataset was created using high-quality SMAP SM as a target and Soil Moisture and Ocean Salinity (SMOS) or Advanced Microwave Scanning Radiometer - Earth Observing System (AMSR-E/2) brightness temperature as a source, which allowed extrapolating SM data 13 years back from before SMAP mission launch. CASM represents SM in the top soil layer, defined on a global 25 km EASE-2 grid and covers 2002-2020 with a 3-day temporal resolution. The resulting dataset exhibits excellent spatial and temporal homogeneity, without compromising the interannual variability, and is in excellent agreement with the SMAP data (with a mean correlation of 0.97 between the SMAP and CASM SM for the period when the two overlap). Moreover, the input and target datasets were divided into seasonal cycle and residuals, with the NN trained on the residuals. This approach ensures that the high performance does not mask a simple seasonal cycle matching but rather exemplifies the skill targeted at predicting extremes; with the NN achieving a correlation of 0.75 on the test data for the residuals. Comparison to 367 global in-situ SM monitoring sites shows a SMAP-like median correlation of 0.66 between station SM and CASM SM from the corresponding grid cell. Additionally, the SM product uncertainty was assessed, and both aleatoric and epistemic uncertainties were estimated and included in the dataset. Mean epistemic uncertainty, related to the NN model structure, ranges from 0.007 m<sup>3</sup>/m<sup>3</sup> to 0.014 m<sup>3</sup>/m<sup>3</sup> and on average is close to a desired SM product stability threshold of 0.01 m<sup>3</sup>/m<sup>3</sup> per year. Aleatoric uncertainty, defined as input noise propagated through the system, depends on the introduced level of noise. With 10% noise applied to the residuals, the resulting mean standard deviation of the model outputs rises from 0.005 to 0.007 m<sup>3</sup>/m<sup>3</sup>. </p>
The compiled 8-year dataset (2012-2019) consisting of weekly river water quality indicators (CODMn, DO, NH3-N and PH ) in majors 10 sub-basin of Yangtze river based on imputation of machine learning
<p>Water quality is significantly affected by global climate change and human activities, with diverse critical factors shaping its state in rivers and lakes. In the study, we utilized four indicators to characterize water quality: the physical water quality parameters included dissolved oxygen (DO, mg/L) and PH, while the chemical water quality parameters encompassed chemical oxygen demand (CODMn, mg/L) and ammonia nitrogen (NH3-N, mg/L). This study establishes weekly water quality models for typical 10 sub-basins along the Yangtze River using machine learning methods, which incorporate the impacts of hydro-meteorological and anthropogenic factors.These 10 sub-basins represent the principal tributaries of the Yangtze River basin and include Dongting Lake, the upper Han River, the lower Han River, the Jialing River, the Jinsha River, the Li River, the Min River, Poyang Lake, the Xiang River, and the Yuan River. This data collection was performed by National Environmental Monitoring Centre (http://www.cnemc.cn/sssj/szzdjczb/index_1.shtml). The water quality indicators discussed in this study are assessed in accordance with the national standard GB 3838-2002. Please refer to the paper for details.</p>
Unlabeled AnuraSet: A dataset for leveraging unlabeled data in machine learning models for passive acoustic monitoring
<p>The Unlabeled AnuraSet (U-AnuraSet) is an extension of the original AnuraSet dataset. It consists of soundscape recordings from passive acoustic monitoring conducted in Brazil. The recording sites are identical to those in the original AnuraSet. Each site comprises 2,666 one-minute raw audio files of unlabeled data. The U-AnuraSet is publicly available to encourage machine learning researchers to explore innovative methods for leveraging unlabeled data in the training of models aimed at solving problems such as anuran call identification.</p> <p>If you find the Unlabeled AnuraSet useful for your research, please consider citing it as follows:</p> <p>Cañas, J.S., Toro-Gómez, M.P., Sugai, L.S.M., et al. A dataset for benchmarking Neotropical anuran calls identification in passive acoustic monitoring. Sci Data 10, 771 (2023). https://doi.org/10.1038/s41597-023-02666-2</p>
Dataset for: Adapting Explainable Machine Learning to Study Mechanical Properties of Two-Dimensional Hybrid Halide Perovskites
<p>This archive contains the in plane and out of plane Young's moduli (complete with respective VASP in and outputs) for 154 n=1 and 30 n>1 2D hybrid organic and inorganic perovskites. The data was used in the publication "Adapting Explainable Machine Learning to Study Mechanical Properties of Two-Dimensional Hybrid Halide Perovskites".</p> <p>Computational settings for the calculations were:</p> <p>Perdew-Burke-Ernzerhof (PBE) exchange-correlation with Tkatchenko-Scheffler (TS) van der Waals (vdW) corrections<br>Projector augmented-wave (PAW) method for the description of interactions between core and valence electrons.<br>A plane wave cutoff energy of 520 eV<br>A Γ-centered Monkhorst-Pack k-point mesh with a grid spacing of 2π × 0.040 Å−1 <br>Geometry optimizations were performed until energy and residual forces fell below 10−6 eV and 0.001 eV/ Å, respectively. <br><br></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.