Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
China's transboundary water resources estimation using machine learning approaches
<p>China's transboundary water resources estimated using machine learning models (random forest, gradient boosting, and stacking).</p>
Initial dataset used in SIMON, an automated machine learning approach
<p>The 7-Zip file contains raw data in the CSV file downloaded from Stanford Data Miner and used for further analysis using mulset algorithm and SIMON, as described in the publication:</p> <p>Tomic A, Tomic I, Rosenberg-Hasson Y, Dekker CL, Maecker HT, and Davis MM. SIMON, an automated machine learning system reveals immune signatures of influenza vaccine responses. <em>JImmunol</em>, doi: 10.4049/jimmunol.1900033, 2019.</p> <p>File was compressed using 7-Zip available at https://www.7-zip.org/.</p>
EELS dataset used for machine learning in github.com/trygvrad/RNN-on-EELS-data
<p>EELS dataset used for machine learning.</p> <p>The dataset shows a GaN substrate with an AlGaN film and SiNX top layer</p> <p>The data is courtesy of Andrew Lang and Mitra Taheri, and was uploaded by Trygve M. Ræder</p> <p>Code available at doi:10.5281/zenodo.2580160<br> or github.com/trygvrad/RNN-on-EELS-data<br> Data available at doi:10.5281/zenodo.2580185<br> Presentation available at doi:10.5281/zenodo.2580206</p>
Hyperparameter tuning and performance assessment of statistical and machine-learning models using spatial data.
<p>This is a research compendium (RC) for the publication "Hyperparameter tuning and performance assessment of statistical and machine-learning algorithms using spatial data".</p> <p>The code (including figures, appendices and the manuscript) is packed in <strong>pathogen-modeling-3.zip </strong>or can be found directly in the <a href="https://github.com/pat-s/pathogen-modeling">Github repository</a>.</p> <ul> <li><strong>Publication figures</strong>: analysis/paper/submission/3/latex-source-files/</li> <li><strong>Appendices</strong>: analysis/paper/submission/3/</li> </ul> <p>This RC represents a static snapshot at the time of submission. The Github repository will receive changes after the publication was published.</p> <p><strong>Data sources</strong></p> <ul> <li>Atlas Climatico: <a href="http://opengis.uab.es/wms/iberia/index.htm">http://opengis.uab.es/wms/iberia/index.htm</a></li> <li>DEM: ftp://ftp.geo.euskadi.eus/lidar/MDE_LIDAR_2016_ETRS89/</li> <li>Lithology: <a href="http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home">http://www.geo.euskadi.eus/geonetwork/srv/spa/main.home</a></li> <li>pH: <a href="https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0">https://esdac.jrc.ec.europa.eu/content/soil-ph-europe#tabs-0-description=0</a></li> <li>soil: <a href="https://www.isric.org/explore/soilgrids">https://www.isric.org/explore/soilgrids</a></li> </ul> <p><strong>Licenses</strong></p> <p>All files are shared via the given license with the exception of "soil.tif" which is shared via the <strong>ODbL </strong>license<strong>.</strong></p>
DeepThought DPR: Distributed peer review enhanced with natural language processing and machine learning - Dataset I
<p>This is the anonymized dataset obtained from the DPR Experiment run at ESO in Fall 2018. If this dataset is used both this DOI as well as the main paper need to be cited. </p>
HDF5 Table for Machine Learning on Difference image analysis
<p>A table with the features.</p> <p>There exists several tables within this file.</p> <p>Each table is a set of transient candidates, with the column IS_REAL stating their status of real or bogus.</p>
Machine Learning for LTE Energy DetectionPerformance Improvement
<p>LTE signal generated using matlab. LTE signal is transmitted through fading channel with shadowing. Energy Detection is performed on received LTE data. Energy Detection results are saved in the text files. Those results can be used in Python programs to test out k-Nearest Neighbors and Random Forest Machine Learning algorithms. Results can be then used in matlab programs to show on plots improvement in terms of probability of detection and false alarm.</p>
WaivOps EDM-TR8: Open Audio Resources for Machine Learning in Music
<p><strong>EDM-TR8 Dataset</strong></p> <p>EDM-TR8 is an open audio dataset composed of a series of drum recordings in the style of electronic dance music (EDM). This dataset primarily focuses on the iconic sounds of the Roland TR-808 drum machine with additional electro synth drums. The dataset contains 3,790 audio loops recorded in uncompressed stereo WAV format, generated with custom audio samples and a MIDI dataset used for training symbolic music models.</p> <p><strong>Dataset</strong></p> <p>The primary objective of this dataset is to provide accessible content for machine learning applications in music and audio research. Some potential use cases for this dataset include tempo detection and classification, drum rhythm analysis, audio-to-MIDI conversion, source separation, automated mixing, music information retrieval, AI music generation, sound design and signal processing.</p> <p><strong>Specifications</strong></p> <ul> <li>3790 audio loops (approximately 9 hours)</li> <li>16-bit WAV format</li> <li>BPM labeled</li> <li>Tempo range: 95–130bpm</li> <li>Variational drum patterns</li> <li>Multi-genre rhythm styles</li> </ul> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed by the sound label company Patchbanks.. All recordings have been compiled by verified sources for copyright clearance.</p> <p>The EDM-TR8 dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-EDM-TR8">GitHub repository</a>.</p>
Dataset for Real-Time Indoor Localization System Based on Wearable Device, Bluetooth Low Energy (BLE) Beacons, and Machine Learning
<p>The dataset titled <strong>"Real-Time Indoor Localization System Based on Wearable Device, Bluetooth Low Energy (BLE) Beacons, and Machine Learning</strong><strong>"</strong> was collected to support the development of an indoor localization system that operates at the room level. The dataset includes measurements of Received Signal Strength Indication (RSSI) from Bluetooth Low Energy (BLE) beacons (specifically the iBKS105 model) recorded by an ESP32 device. These RSSI values were captured across various rooms, allowing for precise localization within an indoor environment. The dataset is particularly useful for research in indoor localization system including machine learning-based localization algorithms.</p>
WaivOps RGTM-PNO: Open Audio Resources for Machine Learning in Music
<p><strong>RGTM-PNO Dataset</strong></p> <p>RGTM-PNO is an open audio dataset featuring a collection of vintage piano songs in the style of ragtime, a genre that flourished around the turn of the 20th century. The dataset contains 262 audio tracks recorded in uncompressed stereo WAV format, synthetically generated using a custom soundfont and MIDI files sourced from public resources online.</p> <p><strong>Dataset</strong></p> <p>The primary objective of this dataset is to provide accessible content for machine learning applications in music and audio research. Some potential use cases for this dataset include audio classification, automatic music transcription (ADT), music information retrieval (MIR), melody analysis, AI music generation, sound design and signal processing.</p> <p><strong>Specifications</strong></p> <ul> <li>262 piano songs (approximately 13.5 hours)</li> <li>16-bit WAV format</li> <li>Tempo: 120bpm (live performance in absolute time)</li> <li>Variational chorus detuning (vintage piano sound)</li> <li>Paired audio and MIDI data</li> </ul> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed by the sound label company Patchbanks. The audio recordings were sonified from MIDI files containing historical musical compositions believed to be in the public domain and copyright free.</p> <p>The RGTM-PNO dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-EDM-TR8">GitHub repository</a>.</p>
Datasets, trained models and supporting results for machine learning tensorial properties of atomic systems via XPaiNN model.
Open the record for dataset details and reuse information.
Machine Learning based identification of putative coral pathogens in endangered Caribbean staghorn coral
<h1>Supplementary Files</h1> <p>SupplementaryFile1.csv.gz – Metadata for field collected samples with columns:</p> <ul> <li>“sample_id” – individual sample names.</li> <li>“health” – “H” healthy and “D” diseased fragments.</li> <li>“year” – year fragment collected.</li> <li> “season” – season fragment collected (“S” July and “W” January)</li> <li>“site” – location fragment collected from</li> <li>“lib.size” – total number of sequenced reads</li> <li>“norm.factors” – factor used to normalize read counts of ASVs</li> </ul> <p>SupplementaryFile2.csv.gz – Metadata for tank collected samples with columns:</p> <ul> <li> “sample_id”<a name="_Hlk163818341"></a> – individual sample names.</li> <li>“geno” – fragment genotype</li> <li> “fragment_id” – fragment identification tracked through repeated sampling</li> <li> “tank_id” – tank identification</li> <li>“time_treat” – concatenated metric for sampling time, exposure, and disease outcome separated by “_” <ul> <li> Time – 0, 2, 8</li> <li>Exposure – “D” Diseased, “N” Healthy</li> <li>Disease Outcome - “D” Diseased, “H” Healthy</li> </ul> </li> <li> “lib.size” – total number of sequenced reads</li> <li>“norm.factors” – factor used to normalize read counts of ASVs</li> </ul> <p>SupplementaryFile3.fasta – FASTA file including complete 16s sequences named with ASV identifier and taxonomy.</p> <p>SupplementaryFile4.csv.gz – Matrix of the number of reads of each ASV sequenced in each sample. Combined both field and tank samples.</p> <p>SupplementaryFile5.csv.gz – Matrix of the log2 CPM of each ASV sequenced in each sample. Combined both field and tank samples.</p> <p>SupplementaryFile6.csv.gz – Complete results for each ASV association.</p> <ul> <li> “top_classification” – lowest taxonomic classification with more than 80% confidence.</li> <li>“taxonomy” – Full taxonomy including confidence in each taxonomic level.</li> <li>“passedFilter” – indicates taxa filtered from analysis due to rarity and/or lack of observations across sample times.</li> <li> “rank_*” – machine learning model rankings, median ranking, and model estimated ranking along with standard error, confidence interval, and FDR adjusted p-value used to identify important ASVs.</li> <li>“ml_retained” – Indicates if the ASV was of above average importance to ML models. NA values indicate ASVs which were filtered prior to ML modelling.</li> <li>“fieldModel_*” – ANOVA table results for each ASV testing the effects of health, year, season and all possible interactions indicating: <ul> <li>Sums of squares, mean squares, numerator and denominator degrees of freedom, F statistic, p-value, and FDR corrected p-value.</li> <li>NA values are filled for ASVs filtered prior to differential abundance analysis.</li> </ul> </li> <li>“diffAbundance_healthAssociation” – Marks the health association of ASVs from differential abundance analysis of field samples: “H” health, “D” diseased, “N” none, NA – filtered prior to differential abundance analysis.</li> <li>“fieldLogFC_*” – Post-hoc contrasts for ML retained ASVs testing the significance of the log2 fold-change between disease and healthy fragments within each sampling time (year: 2016, 2017 & season: “S” July, “W” January) showing: <ul> <li>Mean estimate, standard error, degrees of freedom, lower and upper 95% confidence interval, t-statistic, p-value, FDR adjusted p-value.</li> <li>NA values are filled for ASVs which were not marked as important by ML models.</li> </ul> </li> <li> “field_consistent” – Indicates if the ASV was consistently healthy or disease associated across sampling times. NA values are filled for ASVs which were not marked as important by ML models.</li> <li>“tankModel_*” – ANOVA table results for each ASV testing the effect of the combination of time, disease exposure, and disease outcome, indicating: <ul> <li>Sums of squares, mean squares, numerator and denominator degrees of freedom, F statistic, p-value, and FDR corrected p-value.</li> <li>NA values are filled for ASVs filtered prior to tank experimental analysis.</li> </ul> </li> <li> “tankLogFC_*” – Post-hoc contrasts for ASVs tested in tank exposure experiments. <ul> <li>Contrasts include: <ul> <li>Post-exposure diseased vs healthy outcome regardless of exposure (DvH)</li> <li>Post-exposure diseased vs healthy exposure regardless of outcome (DvN)</li> <li>Post-exposure disease exposed corals with disease symptoms compared to disease exposed but still healthy corals (DDvDH)</li> <li>Post-exposure disease exposed corals with disease symptoms compared to healthy exposed and still healthy corals (DDvNH)</li> <li>Post-exposure disease exposed corals which stay healthy compared to healthy exposed and still healthy corals (DHvNH)</li> <li>Pre-exposure compared to Post-exposure in corals with the disease regardless of exposure (PostvPreD)</li> <li>Pre-exposure compared to Post-exposure in corals without the disease regardless of exposure (PostvPreH)</li> </ul> </li> <li>Mean estimate, standard error, degrees of freedom, lower and upper 95% confidence interval, t-statistic, p-value, FDR adjusted p-value.</li> <li>NA values are filled for ASVs which were not consistently associated with healthy or diseased corals in the field experiment.</li> </ul> </li> <li>“pathogen_classification” – Indicates the predicted microbial classification based on the tank results. Pathogen, Opportunist, Commensal <ul> <li>NA values are filled for ASVs which were not consistently associated with healthy or diseased corals in the field experiment.</li> </ul> </li> </ul>
FireSafetyNet: An Image-Based Dataset with Pretrained Weights for Machine Learning-Driven Fire Safety Inspection
<p>This dataset offers a diverse collection of images curated to support the development of computer vision models for detecting and inspecting Fire Safety Equipment (FSE) and related components. Images were collected from a variety of public buildings in Germany, including university buildings, student dormitories, and shopping malls. The dataset consists of self-captured images using mobile cameras, providing a broad range of real-world scenarios for FSE detection.</p> <p>In the journal paper associated with these image datasets, the open-source dataset FireNet (Boehm et al. 2019) was additionally utilized for training. However, to comply with licensing and distribution regulations, images from <a href="https://www.firenet.xyz/">FireNet</a> have been excluded from this dataset. Interested users can visit the FireNet repository directly to access and download those images if additional data is required. The provided weights (.pt), however, are trained on the provided self-made images and FireNet using YOLOv8.</p> <p>The dataset is organized into six sub-datasets, each corresponding to a specific FSE-related machine learning service:</p> <ol> <li> <p><strong>Service 1: FSE Detection</strong> - This sub-dataset provides the foundation for FSE inspection, focusing on the detection of primary FSE components like fire blankets, fire extinguishers, manual call points, and smoke detectors.</p> </li> <li> <p><strong>Service 2: FSE Marking Detection</strong> - Building on the first service, this sub-dataset includes images and annotations for detecting FSE marking signs.</p> </li> <li> <p><strong>Service 3: Condition Check - Modal</strong> - This sub-dataset addresses the inspection of FSE condition in a modal manner, focusing on instances where fire extinguishers might be blocked or otherwise non-compliant. This dataset includes semantic segmentation annotations of fire extinguishers. For upload reasons, this set is split into <em>3_1_FSE Condition Check_modal_train_data (containing training images and annotations) </em>and <em>3_1_FSE Condition Check_modal_val_data_and_weights (containing validation images, annotations </em>and<em> the best weights).</em></p> </li> <li> <p><strong>Service 4: Condition Check - Amodal</strong> - Extending the modal condition check, this sub-dataset involves amodal detection to identify and infer the state of FSE components even when they are partially obscured. This dataset includes semantic segmentation annotations of fire extinguishers. This dataset includes semantic segmentation annotations of fire extinguishers. For upload reasons, this set is split into <em>4_1_FSE Condition Check_amodal_train_data (containing training images and annotations) </em>and <em>4_1_FSE Condition Check_amodal_val_data_and_weights (containing validation images, annotations </em>and<em> the best weights).</em></p> </li> <li> <p><strong>Service 5: Details Extraction - Inspection Tags</strong> - This sub-dataset provides a detailed examination of the inspection tags on fire extinguishers. It includes annotations for extracting semantic information such as the next maintenance date, contributing to a thorough evaluation of FSE maintenance practices.</p> </li> <li> <p><strong>Service 6: Details Extraction - Fire Classes Symbols</strong> - The final sub-dataset focuses on identifying fire class symbols on fire extinguishers.</p> </li> </ol> <p>This dataset is intended for researchers and practitioners in the field of computer vision, particularly those engaged in building safety and compliance initiatives.</p>
Raw data for Machine learning approach for photocatalysis: An experimentally validated case study of photocatalytic dye degradation
<p>Specification of affiliations:</p> <ul> <li>Hassan Ali - Centre of Polymer Systems</li> <li>Muhammad Yasir - Centre of Polymer Systems</li> <li>Hamza Ul Haq - Laboratory of Alternative Fuel and Sustainability, School of Chemical and Materials Engineering,</li> <li>Ali Can Guler - Centre of Polymer Systems</li> <li>Milan Masar - Centre of Polymer Systems</li> <li>Muhammad Nouman Aslam Khan - Laboratory of Alternative Fuel and Sustainability, School of Chemical and Materials Engineering,</li> <li>Michal Machovsky - Centre of Polymer Systems</li> <li>Vladimir Sedlarik - Centre of Polymer Systems</li> <li>Ivo Kuritka - Centre of Polymer Systems</li> </ul> <p> </p> <p>Raw data for the research paper. Information on the data collection are described in the manuscript. </p>
Datasets used in "Assesing the quality of random number generators through neural networks", Machine Learning: Science and Technology 5 (2024) 025072
<p>Datasets corresponding to the bits generated by different random number generators used in J. L. Crespo et al, Machine Learning: Science and Technology 5 (2024) 025072.</p> <p>VCSEL_QRNG_postprocessed_bits.txt: postprocessed bits from the random generator based on gain-switching of VCSELs </p> <p>EC_LCG_bits.txt: bits from the linear congruential generator on elliptic curves</p> <p>LCG_32_bits.txt:: bits from the linear congruential generator with 32 bits</p> <p>VCSEL_QRNG_raw_bits.txt: raw bits from the random generator based on gain-switching of VCSELs</p> <p> </p>
GalaxiesML: an imaging and photometric dataset of galaxies for machine learning
<div># GalaxiesML README</div> <p> </p> <div>Version 6.1</div> <p> </p> <div>## Overview</div> <p> </p> <div>GalaxiesML is a machine learning-ready dataset of galaxy images, photometry, redshifts, and structural parameters. It is designed for machine learning applications in astrophysics, particularly for tasks such as redshift estimation and galaxy morphology classification. The dataset comprises **286,401 galaxy images** from the Hyper-Suprime-Cam (HSC) Survey PDR2 in five filters: g, r, i, z, y, with spectroscopically confirmed redshifts as ground truth.</div> <p> </p> <div>This dataset is particularly useful for developing machine learning models for upcoming large-scale surveys like **LSST** and **Euclid**.</div> <p> </p> <div>## Features</div> <p> </p> <div>- **286,401 galaxy images** in five photometric bands (g, r, i, z, y).</div> <div>- Spectroscopic redshifts for each galaxy, with redshift values ranging from **0.01 to 4**.</div> <div>- Morphological parameters derived from galaxy images, including **Sérsic index**, **half-light radius**, and **ellipticity**.</div> <div>- **Machine learning-friendly formats**: images are provided in **HDF5** format, along with CSV metadata.</div> <p><br><br></p> <div>## Examples of Using GalaxiesML</div> <p> </p> <div>Examples of uses of GalaxiesML are outlined in Do et al. (2024). The repository for example code are here:</div> <p> </p> <div><a href="https://github.com/astrodatalab/galaxiesml_examples">https://github.com/astrodatalab/galaxiesml_examples</a></div> <p> </p> <div>## Citation</div> <p> </p> <div>Please cite the following papers if you use this dataset in your work:</div> <p> </p> <div>1. **GalaxiesML Dataset**:</div> <div> <div> <div>Do, T. et al., *GalaxiesML: A Dataset of Galaxy Images, Photometry, Redshifts, and Structural Parameters for Machine Learning*. arXiv:2410.00271, <a href="https://arxiv.org/abs/2410.00271">https://arxiv.org/abs/2410.00271</a> (2024)</div> </div> </div> <p> </p> <div>2. **Hyper Suprime-Cam Subaru Strategic Program (HSC PDR2)**:</div> <div>- Aihara, H., et al., *Second Data Release of the Hyper Suprime-Cam Subaru Strategic Program*. Publications of the Astronomical Society of Japan, 71(6), 114 (2019). DOI: [10.1093/pasj/psz103](https://doi.org/10.1093/pasj/psz103)</div> <p> </p> <div>3. **Spectroscopic Surveys**:</div> <div>- Several publicly available spectroscopic redshift catalogs were used in creating this dataset. Notable sources include:</div> <div>- **zCOSMOS Survey**: Lilly, S. J., et al., *The zCOSMOS 10k-Bright Spectroscopic Sample*. The Astrophysical Journal Supplement Series, 184(2), 218-229 (2009). DOI: [10.1088/0067-0049/184/2/218](https://doi.org/10.1088/0067-0049/184/2/218)</div> <div>- **VIMOS Public Extragalactic Survey (VIPERS)**: Garilli, B., et al., *The VIMOS Public Extragalactic Survey (VIPERS): First Data Release of 57,204 Spectroscopic Measurements*. Astronomy & Astrophysics, 562, A23 (2014). DOI: [10.1051/0004-6361/201322790](https://doi.org/10.1051/0004-6361/201322790)</div> <div>- **DEEP2 Survey**: Newman, J. A., et al., *The DEEP2 Galaxy Redshift Survey: Design, Observations, Data Reduction, and Redshifts*. The Astrophysical Journal Supplement Series, 208(1), 5 (2013). DOI: [10.1088/0067-0049/208/1/5](https://doi.org/10.1088/0067-0049/208/1/5)</div> <p><br><br></p> <div>## How to Access</div> <p> </p> <div>The dataset is publicly available on **Zenodo** with the DOI: **[10.5281/zenodo.11117528](https://doi.org/10.5281/zenodo.11117528)**.</div> <p> </p> <div>## License</div> <p> </p> <div>This dataset is licensed under a **Creative Commons Attribution 4.0 International License (CC BY 4.0)**. You are free to share and adapt the dataset as long as appropriate credit is given. For more details, visit: **[CC BY 4.0 License](https://creativecommons.org/licenses/by/4.0/)**.</div> <p> </p> <div>Please cite the references mentioned above if you use this dataset in your work.</div>
Quantification of Fatty Acids in Hemp Seeds (Cannabis sativa L.) and Yield Prediction Using Machine Learning for Soxhlet and Ultrasound Extraction Methods
<p>This study focuses on the quantification of fatty acids present in hemp seeds (Cannabis sativa L.) cultivated in the Ecuadorian Andes using Soxhlet and ultrasound extraction methods. The aim is to evaluate and compare the extraction efficiency of these two techniques. Furthermore, machine learning models are applied to predict extraction yields based on experimental conditions. Using locally cultivated seeds provides valuable insights into the influence of regional agro-climatic conditions on the chemical composition. The integration of predictive algorithms offers a novel approach to optimizing the extraction process, enhancing both precision and efficiency. The findings could contribute to developing sustainable extraction methods for high-value bioactive compounds in the food and pharmaceutical industries.</p>
UNSUPERVISED MACHINE LEARNING AND VECTOR MODELS IN DESIGNING AND OPTIMIZATION OF TELECOM RETAIL CHANNELS
<p>This paper examines the use of unsupervised machine learning and vector models in the design and optimization of retail channels for telecommunications services. Unsupervised machine learning allows you to analyze and identify hidden patterns in large volumes of untagged data, which is especially important in a dynamically changing consumer market. Vector models, in turn, provide high accuracy of demand forecasting and inventory management, contributing to an increase in the efficiency of trading channels. The synergy of these technologies allows companies to improve customer experience, optimize operational processes and increase competitiveness in the market. The main focus of the work is on data processing methods, including correlation analysis, the use of the support vector machine (SVM) method and its adaptation to solve problems related to predicting customer behavior and optimizing logistics processes.</p>
MACHINE LEARNING ALGORITHMS FOR ANOMALY DETECTION IN PUBLIC DATA USING GITHUB AS AN EXAMPLE
<p>This study explores the application of machine learning algorithms for detecting anomalies in GitHub data to enhance the evaluation of technological projects. The research aims to develop a robust methodology for identifying data anomalies, such as artificial activity spikes, that can distort project assessments. Methods such as Isolation Forest, One-Class SVM, and advanced deep learning techniques like autoencoders and GANs are employed to analyze and identify irregular patterns in GitHub repositories. The findings demonstrate that these algorithms effectively detect both obvious and subtle anomalies, offering reliable insights into project authenticity. The proposed conceptual model integrates these methods into a scalable system, enhancing transparency and accuracy in technological project evaluation. The novelty of this work lies in its comprehensive approach to analyzing GitHub data, combining traditional and deep learning techniques to improve the reliability of assessments, making it a significant contribution to the field.</p>
MACHINE LEARNING APPROACHES FOR DEMAND FORECASTING: THE IMPACT OF CUSTOMER SATISFACTION ON PREDICTION ACCURACY
<p><span>This study investigates the effectiveness of various machine learning models in predicting product demand based on customer satisfaction data. Four models—Linear Regression, Random Forest, Gradient Boosting, and Support Vector Machine (SVM)—were evaluated using performance metrics, including Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R² score. The results indicate that Gradient Boosting achieved the highest accuracy, with an MAE of 2.56, MSE of 12.75, RMSE of 3.57, and R² score of 0.82, effectively capturing the complex, non-linear relationships inherent in customer satisfaction factors. Random Forest also demonstrated strong performance, while Linear Regression and SVM showed limitations in handling intricate datasets. These findings underscore the importance of utilizing advanced machine learning techniques for accurate demand forecasting, highlighting the critical role of customer satisfaction data in enhancing predictive capabilities. The insights gained from this research can guide organizations in optimizing inventory management and improving customer satisfaction in a rapidly evolving market.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.