Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Long-term trends of ambient nitrate (NO3-) concentrations across China based on ensemble machine-learning models
<p>The data is the monthly NO3- concentrations across China during 2005-2015. These data was obtained using a novel ensemble model combining random forest (RF), gradient boosting decision tree (GBDT), and extreme gradient boosting (XGBoost) algorithms based on satellite data, assimilated meteorology, and other geographical covariates.</p> <p>In the datasets, XX-YY denote the XX month in YY year.<br> For instance, January-05 denotes the January in 2005.<br> NaN in the data denote the missing values.</p>
Star formation and morphological properties of galaxies in the P\lowercase{an}-STARRS 3$\pi$ survey- I.\\ A machine learning approach to galaxy and supernova classification
<pre>This is the catalog presented in Baldeschi, et al (2020). The catalog is subdivided in 13 csv files. Files description: First column: Panstar ID (integer) Second column: Right ascension [deg] (float) Third column: Declination [deg] (float) Fourth column: Probability for a source of being a star (P_star) (float) Fifth column: Probability for a galaxy of being higly star-forming (P_HSFF) (float) Sixth column: Probability for a galaxy of being spiral (P_spiral) (float) seventh column: Compleatness flag (string) If using this catalog for publications, please cite Baldeschi, et al (2020). Fourth column values are from Tachibana & Miller (2018). For a detailed description of the columns we refer to the Appendix A of Baldeschi, et al (2020).</pre>
Machine learning for buildings' characterization and power-law recovery of urban metrics
<p>We focus on a critical component of the city: its building stock, which holds much of its socio-economic activities. In our case, the lack of a comprehensive database about their features and its limitation to a surveyed subset lead us to adopt data-driven techniques to extend our knowledge to the near-city-scale. Neural networks and random forests are applied to identify the buildings’ number of floors and construction periods’ dependencies on a set of shape features: area, perimeter, and height along with the annual electricity consumption, relying on a surveyed data in the city of Beirut. The predicted results are then compared with established scaling laws of urban forms, which constitutes a further consistency check and validation of our workflow.</p>
Machine-learning method for predicting the scanning parameters influence on random measurement error
<p>Data acompaniing the study.</p>
Machine Learning for Anomaly Detection in Cyanobacterial Fluorescence Signals
<p>Excel files containing chlorophyll a and phycocyanin fluorescence data imported from <a href="https://www.glerl.noaa.gov/res/HABs_and_Hypoxia/habTracker.html">https://www.glerl.noaa.gov/res/HABs_and_Hypoxia/habTracker.html</a> for buoys WE2, WE4, WE8, and WE13. The Python code used to manipulate the data is also included.</p>
Characterization of descriptors in machine learning for data-based sputtering yield prediction
<p>table I-V</p>
Identifying galaxies, quasars and stars with machine learning: a new catalogue of classifications for 111 million SDSS sources without spectra - parquet format
<p>This is the same as the published data available under 10.5281/zenodo.3768398, but in the format of parquet files. This means you can access it using Dask for convenience when using cloud compute facilities. </p> <p>Abstract: We used 3.1 million spectroscopically labelled sources from the Sloan Digital Sky Survey (SDSS) to train an optimised random forest classifier using photometry from the SDSS and the Widefield Infrared Survey Explorer (WISE). We applied this machine learning model to 111 million previously unlabelled sources from the SDSS photometric catalogue which did not have existing spectroscopic observations. Our new catalogue contains 50.4 million galaxies, 2.1 million quasars, and 58.8 million stars. We provide individual classification probabilities for each source, with 6.7 million galaxies (13%), 0.33 million quasars (15%), and 41.3 million stars (70%) having classification probabilities greater than 0.99; and 35.1 million galaxies (70%), 0.72 million quasars (34%), and 54.7 million stars (93%) having classification probabilities greater than 0.9. Precision, Recall, and F1 score were determined as a function of selected features and magnitude error. We investigate the effect of class imbalance on our machine learning model and discuss the implications of transfer learning for populations of sources at fainter magnitudes than the training set. We used a non-linear dimension reduction technique (Uniform Manifold Approximation and Projection: UMAP) in unsupervised, semi-supervised, and fully-supervised schemes to visualise the separation of galaxies, quasars, and stars in a two-dimensional space. When applying this algorithm to the 111 million sources without spectra, it is in strong agreement with the class labels applied by our random forest model.</p> <p>When using this dataset, please reference our paper via the journal (<a href="https://arxiv.org/abs/1909.10963">https://arxiv.org/abs/1909.10963</a>) and this DOI (10.5281/zenodo.4060257). If you make use of our scripts please reference our Github repository DOI (10.5281/zenodo.3855160).</p> <p>File descriptions:</p> <p>All of these files are Pandas Dataframes, saved as uncompressed parquet files for ease of access when using cloud compute such as Dask. df_spec_classprobs.parquet contains the spectroscopically observed sources used for training and testing. This has been cleaned, and has the results of the random forest classifier added as additional columns (sources used for training have NaNs in the class_pred column). SDSS-ML-all.parquet contains the 111 million photometrically observed sources, with our class labels and probabilities added.</p>
Kolberger Heide community compositions and machine learning results
<p>Bacteria are ubiquitous and live in complex microbial communities, which can react rapidly to changing environmental conditions. Their physiological variety enables communities to respond in specific ways to environmental drivers, potentially resulting in distinct microbial fingerprints for a given environmental state. Our goal was to assess the opportunities and limitations of machine learning to detect fingerprints indicating the presence of the munition compound 2,4,6-trinitrotoluene (TNT) in southwestern Baltic Sea sediments.</p> <p>Over 40 environmental variables including grain size distribution, elemental composition and concentration of munition compounds (mostly at pmol g<sup>-1</sup> levels) from 150 sediments collected at the near-to-shore munition dumpsite Kolberger Heide by the German city of Kiel were combined with 16S rRNA gene amplicon sequencing libraries. Prediction was achieved using Random Forests; the robustness of predictions was validated using Artificial Neural Networks. To facilitate machine learning with microbiome data we developed the R package phyloseq2ML.</p> <p>Using the most classification-relevant 25 bacterial genera exclusively, potentially representing a TNT-indicative fingerprint, TNT was predicted correctly with up to 81.5 % balanced accuracy. False positive classifications indicated that this approach has also the potential to identify samples where the original TNT contamination was no longer detectable. The sensitivity of this approach can be deduced from the fact that TNT presence was neither identified among the main drivers of the microbial community composition, nor did it correlate with sediment metal content, demonstrated by decreased prediction rates using environmental variables.</p> <p>Our results suggest that microbial communities can predict even minor influencing factors in complex environments, demonstrating the potential of this approach for the discovery of contamination events over an integrated period of time and for environmental monitoring in general.</p>
Development, evaluation, and validation of machine learning models for COVID-19 detection based on routine blood tests
<p>The .xlsx dataset includes all patients used for training, internal-external and external validation: these can be distinguished by looking at the ID (first column) in the dataset: those in format Axxxx-<Date> are the data used for the training, those in the format 20xx are the data used for the internal-external validation, while the remaining data were used for external validation.</p> <p>As regards the features: for the Target feature the value 1 stands for "Positive to COVID-19" while the value 0 stands for "Negative to COVID-19"; while for the Sex feature the value 1 stands for "Male" while the value 0 stands for "Female".</p> <p>The full article is available at: https://www.degruyter.com/view/journals/cclm/ahead-of-print/article-10.1515-cclm-2020-1294/article-10.1515-cclm-2020-1294.xml.</p> <p>A pre-print version of the article is also available on MedrXiv: https://www.medrxiv.org/content/10.1101/2020.10.02.20205070v1</p> <p><strong>ABSTRACT</strong></p> <p><strong>Background</strong> The rRT-PCR test, the current gold standard for the detection of coronavirus disease (COVID-19), presents with known shortcomings, such as long turnaround time, potential shortage of reagents, false-negative rates around 15–20%, and expensive equipment. The hematochemical values of routine blood exams could represent a faster and less expensive alternative. </p> <p><strong>Methods</strong> Three different training data set of hematochemical values from 1,624 patients (52% COVID-19 positive), admitted at San Raphael Hospital (OSR) from February to May 2020, were used for developing machine learning (ML) models: the complete OSR dataset (72 features: complete blood count (CBC), biochemical, coagulation, hemogasanalysis and CO-Oxymetry values, age, sex and specific symptoms at triage) and two sub datasets (COVID-specific and CBC dataset, 32 and 21 features respectively). 58 cases (50% COVID-19 positive) from another hospital, and 54 negative patients collected in 2018 at OSR, were used for internal-external and external validation.</p> <p><strong>Results</strong> We developed five ML models: for the complete OSR dataset, the area under the receiver operating characteristic curve (AUC) for the algorithms ranged from 0.83 to 0.90; for the COVID-specific dataset from 0.83 15 to 0.87; and for the CBC dataset from 0.74 to 0.86. The validations also achieved good results: respectively, AUC 16 from 0.75 to 0.78; and specificity from 0.92 to 0.96. </p> <p><strong>Conclusions</strong> ML can be applied to blood tests as both an adjunct and alternative method to rRT-PCR for the fast and cost-effective identification of COVID-19-positive patients. This is especially useful in developing countries, or in countries facing an increase in contagions.</p>
InSet: A Tool to Identify Architecture Smells Using Machine Learning
<p>Architectural smells (ASs) are architectural decisions that negatively affect the maintenance and evolution of software. Most of the existing tools able to identify AS rely on few metrics with fixed thresholds. However, it is not possible to define specific metrics and thresholds that meet all the cases, i.e., the classification of a piece of code in smell or not can depend on the domain, the experience of developers, organization patterns or even from a vast set of features - so there is a subjective ingredient in this decision.<br> Machine Learning (ML) can help to make these decisions/classifications more precise by taking into consideration a vast set of features and also feedback from experts.<br> This paper presents a machine learning-based tool to detect the architectural smells Unstable Dependency(UD) and God Component(GC). Our tool is able to take into consideration users' feedback to retrain the algorithms and constantly improve their performance. Our tool got good result in terms of accuracy, precision, recall, F-measure and Kappa's coefficient.</p>
Data related to: MACHINE LEARNING AND QUANTUM MECHANICS APPROACH TO MORE CHEMICALLY-AWARE MOLECULAR DESCRIPTORS FOR MEDICINAL CHEMISTRY APPLICATIONS
<p>DEMIN VS QM ELECTRONIC POTENTIAL CORRELATIONS FOR THE N:= ATOM TYPE AND THE N1 ATOM TYPE</p> <p>dEmin versus H-bond basicity scale and H-bond acidity scale </p>
Data from: Study of the accuracy of a machine learning muscle MRI-based tool for diagnosis the of muscular dystrophies
<p>Objective: Genetic diagnosis of muscular dystrophies (MDs) has classically been guided by clinical presentation, muscle biopsy and muscle MRI data. Muscle MRI suggests diagnosis based on the pattern of muscle fatty replacement. However, patterns overlap between different disorders and knowledge about disease-specific patterns is limited. Our aim was to develop a software-based tool that can recognize muscle MRI patterns and thus aid diagnosis of MDs. Methods: We collected 976 pelvic and lower limbs T1 weighted muscle MRIs from 10 different MDs. Fatty replacement was quantified using Mercuri score and files containing the numeric data were generated. Random forest unsupervised machine learning was applied to develop a model useful to identify the correct diagnosis. 2000 different models were generated and the one with higher accuracy was selected. A new set of 20 MRIs was used to test the accuracy of the model, and the results were compared with diagnoses proposed by 4 specialists in the field. Results: A total of 976 lower limbs MRIs from 10 different MDs were used. The best model obtained had a 95.7% accuracy, with 92.1% sensitivity and 99.4% specificity. When compared with experts on the field, the diagnostic accuracy of the model generated was significantly higher in a new set of 20 MRIs. Conclusion: Machine learning can help medical doctors in the diagnosis of muscle dystrophies by analyzing patterns of muscle fatty replacement in muscle MRI. This tool can be helpful for daily clinics but also in the interpretation of the results of next generation sequencing tests. Classification of Evidence: This study provides Class II evidence that a muscle MRI-based artificial intelligence tool accurately diagnosis muscular dystrophies.</p>
Evolutionary computing and machine learning for the discovering of low-energy defect configurations
<p>The archive "dataset.tar.gz" contains the structures obtained and reported in the study:<br> "Evolutionary computing and machine learning for the discovering of low-energy defect configurations".</p> <p>Read the README file for detailed information about the archive content and how to read the hdf5 files.</p>
Accuracy or novelty: what can we gain from target-specific machine learning-based scoring functions in virtual screening?
<p>Datasets, features, and some representative scripts utilized in the paper "Accuracy or novelty: what can we gain from target-specific machine learning-based scoring functions in virtual screening?" </p>
Data from: A machine learning method to monitor China's AIDS epidemics with data from Baidu Trends
Background: AIDS victims' unwillingness to report their disease, due to social discrimination against them, makes it hard for disease control departments to accurately monitor the disease's dynamics through traditional surveillance tools, such as over-the-counter drug sales and hospital or self-reported data. With the diffusion and adoption of the Internet, the 'big data' aggregated from Internet search engines, which contain users' information on the concern or reality of their health status, provide a new opportunity for AIDS surveillance. This paper uses search engine data to monitor and forecast AIDS in China. Methods: A machine learning method, artificial neural networks (ANNs), is used to forecast AIDS occurrences and deaths. Search trend data related to AIDS from the largest Chinese search engine, Baidu.com, are collected and selected as the input variables of ANNs, and officially reported actual AIDS occurrences and deaths are used for the output variable. Three criteria, the mean absolute percentage error, the root mean squared percentage error, and the index of agreement, are used to test the forecasting performance of the ANN method. Results: Based on the monthly time-series data from January 2011 to June 2017, this article finds that, under three criteria, the ANN method can lead to satisfactory forecasting of AIDS occurrences and deaths, regardless of the change of the number of search queries. Conclusions: Internet-based data should be adopted as a real-time, cost-effective complement to a traditional AIDS surveillance system.
Data from: Machine learning to classify animal species in camera trap images: applications in ecology
Motion‐activated cameras ("camera traps") are increasingly used in ecological and management studies for remotely observing wildlife and are amongst the most powerful tools for wildlife research. However, studies involving camera traps result in millions of images that need to be analysed, typically by visually observing each image, in order to extract data that can be used in ecological analyses. We trained machine learning models using convolutional neural networks with the ResNet‐18 architecture and 3,367,383 images to automatically classify wildlife species from camera trap images obtained from five states across the United States. We tested our model on an independent subset of images not seen during training from the United States and on an out‐of‐sample (or "out‐of‐distribution" in the machine learning literature) dataset of ungulate images from Canada. We also tested the ability of our model to distinguish empty images from those with animals in another out‐of‐sample dataset from Tanzania, containing a faunal community that was novel to the model. The trained model classified approximately 2,000 images per minute on a laptop computer with 16 gigabytes of RAM. The trained model achieved 98% accuracy at identifying species in the United States, the highest accuracy of such a model to date. Out‐of‐sample validation from Canada achieved 82% accuracy and correctly identified 94% of images containing an animal in the dataset from Tanzania. We provide an r package (Machine Learning for Wildlife Image Classification) that allows the users to (a) use the trained model presented here and (b) train their own model using classified images of wildlife from their studies. The use of machine learning to rapidly and accurately classify wildlife in camera trap images can facilitate non‐invasive sampling designs in ecological studies by reducing the burden of manually analysing images. Our r package makes these methods accessible to ecologists.
Dataset "Machine learning for natural hazard mapping"
Open the record for dataset details and reuse information.
The best performing landslide susceptibility maps using ensemble machine learning models and precipitation data on basin and regional level in Lombardy, Italy
<p>A selection of landslide susceptibility maps computed through ensemble machine learning models with included precipitation data for the basin of Valchiavenna, and the Lombardy region in Italy.</p> <p>A list of the used base machine learning methods:</p> <ul> <li>Neural Networks.</li> </ul> <p>A list of the precipitation data included in the models:</p> <ul> <li>Average hourly precipitation for the year of 2020,</li> <li>90<sup>th</sup> percentile for the hourly precipitation for the year of 2020 ,</li> <li>Averaged + 90<sup>th</sup> percentile for the hourly precipitation for the year of 2020.</li> </ul> <p>A full list of the model combinations can be found in the "Case Studies" document.</p> <p>The maps are in WGS 84/ UTM zone 32N (EPSG:32632).</p> <p>The map production process details are discussed in Xu et al. 2024. If you use the dataset, please, cite also the paper:</p> <p><em>Qiongjie Xu, Vasil Yordanov, Lorenzo Amici & Maria Antonia Brovelli (2024) Landslide susceptibility mapping using ensemble machine learning methods: a case</em><br><em>study in Lombardy, Northern Italy, International Journal of Digital Earth, 17:1, 2346263, DOI:10.1080/17538947.2024.2346263</em></p> <p>The maps are produced as part of the "Geoinformatics and Earth Observation for Landslide Monitoring" Italy-Vietnam.</p> <p>The work is partially funded by the Italian Ministry of Foreign Affairs and International Cooperation within the project “Geoinformatics and Earth Observation for Landslide Monitoring” CUP D19C21000480001.</p> <p> </p>
Landslide susceptibility maps using base machine learning models on basin and regional level in Lombardy, Italy
<p>A selection of landslide susceptibility maps computed through base machine learning models for the basins of Val Tartano, Upper Valtellina and Valchiavenna, and on a regional level for the Lombardy region in Italy.</p> <p>A list of the used machine learning methods:</p> <ul> <li>Bagging,</li> <li>Random Forest,</li> <li>AdaBoost,</li> <li>Gradient Tree Boosting,</li> <li>Neural Networks.</li> </ul> <p>A full list of the model combinations can be found in the "Case Studies" document.</p> <p>The maps are in WGS 84/ UTM zone 32N (EPSG:32632).</p> <p>The map production process details are discussed in Xu et al. 2024. If you use the dataset, please, cite also the paper:</p> <p><em>Qiongjie Xu, Vasil Yordanov, Lorenzo Amici & Maria Antonia Brovelli (2024) Landslide susceptibility mapping using ensemble machine learning methods: a case</em><br><em>study in Lombardy, Northern Italy, International Journal of Digital Earth, 17:1, 2346263, DOI:10.1080/17538947.2024.2346263</em></p> <p>The maps are produced as part of the "Geoinformatics and Earth Observation for Landslide Monitoring" Italy-Vietnam.</p> <p>The work is partially funded by the Italian Ministry of Foreign Affairs and International Cooperation within the project “Geoinformatics and Earth Observation for Landslide Monitoring” CUP D19C21000480001.</p> <p> </p>
Research data for "Hydrogen under Pressure as a Benchmark for Machine-Learning Potentials"
<p>This dataset supports the paper: "Hydrogen under Pressure as a Benchmark for Machine-Learning Potentials". The paper is online here: https://doi.org/XXXXXXXXXXXXXXXXX.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.