Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
285
datasets available to search
ShareScore release 0.9.0
Dataset results
285 results for “Machine learning dataset”
Dataset for a machine learning tool to improve lymph node staging with FDG-PET/CT
<p>This upload provides Open Data associated with the publication "A machine learning tool to improve prediction of mediastinal lymph node metastases in non-small cell lung cancer using routinely obtainable [<sup>18</sup>F]FDG-PET/CT parameters" by Rogasch JMM <em>et al.</em> (2022).</p> <p>The upload contains the anonymized dataset with 10 features necessary for the final GBM model that was presented in the publication. However, the original full dataset with 40 features was excluded from this Open Data repository because it may not comply with strict rules of data anonymization. The full dataset can be obtained from the corresponding author (julian.rogasch@charite.de) upon reasonable request.</p> <p>Besides the dataset, this upload provides the original python and R scripts that were used as well as their output.</p> <p>A description of all files can be found in "content_description_2022_11_19.txt".</p> <p>A user-friendly web tool that implements the final machine learning model can be found here: <a href="https://baumgagl.github.io/PET_LN_calculator/">PET_LN_calculator</a> </p>
Bioacoustic Dataset of African and Florida Manatee Vocalizations for Machine Learning Applications, 2020-2022
This data package presents a comprehensive acoustic library of manatee vocalizations for machine learning (ML) and classifier development. It includes recordings from two species, African and Florida manatees, sampled across four locations. The species are combined due to the acoustic similarity of their vocalizations, providing a diverse and representative training set for ML algorithms. The dataset consists of 0.5-second WAV clips categorized as either containing manatee vocalizations (MV, n=18,129 clips) or not (Noise, n=23,444 clips). MV clips may include multiple vocalizations or truncated calls. All clips were manually verified by two researchers with expertise in manatee acoustics. Recordings were collected using stationary hydrophones deployed in natural habitats, with variable signal-to-noise ratios (SNR) resulting from changes in distance between the vocalizing manatees and the recorders. Background noise across sites is relatively low, with minimal anthropogenic noise; caution is advised when applying models to noisier environments. No dolphin species are believed to be present at the recording sites, and models trained on this dataset should be used cautiously in dolphin-inhabited regions to avoid false positives. If you use this dataset, please reach out to the listed contacts, we are interested in learning how it supports your work.
Dataset Nucleation Patterns of Polymer Crystals Analyzed by Machine Learning Models
<p>This dataset contains the raw data (01_raw_data), processed data (02_processed_data), and plotting scripts (03_figures) related to the paper:</p> <p>"Nucleation Patterns of Polymer Crystals Analyzed by Machine Learning Models"<br>Atmika Bhardwaj, Jens-Uwe Sommer, Marco Werner</p> <p>Macromolecules <strong>2024</strong>; DOI: <a href="10.1021/acs.macromol.4c00920">10.1021/acs.macromol.4c00920</a></p> <p>Please refer to the README.md files in their respective folders.</p>
Dataset of "Advanced machine learning techniques for State-of-Health estimation in lithium-ion batteries: A comparative study"
This research focuses on State-of-Health (SOH) estimation of lithium-ion (Li-ion) batteries to enhance lifespan and reliability. Using Samsung INR18650-35E cells, 600 cycles were analyzed with machine learning (ML) techniques, including Gaussian Process Regression (GPR), Support Vector Regression (SVR), Feed-Forward Neural Network (FFNN) and Adaptive Neuro-Fuzzy Inference System (ANFIS). Input features from charging and discharging cycles were selected with Pearson Correlation Analysis (PCA) and Exhaustive Search (ES) to optimize inputs for each ML method. Models were tested on datasets of varying sizes to evaluate performance and overfitting, including an experiment where SOH estimation of one battery was performed using training data from another. The findings highlight each model's strengths and limitations, guiding their application in battery health prediction.
RADIT: A Machine Learning-Reconstructed Dataset of River Discharge, Temperature, and Heat Flux into the Arctic Ocean
<p>The Reconstructed Arctic-draining river DIscharge and Temperature (RADIT) dataset provides daily records of river discharge, temperature, and heat flux for 25 major Arctic-draining rivers from 1950 to 2023. Using machine learning methods and ERA5-Land reanalysis data, we reconstructed these key hydrological variables with high accuracy (most NSEs > 0.8).</p> <p>Due to licensing restrictions and to encourage adherence to the stated licenses of the original input data, this dataset only provides the reconstructed (filled) values. Users can obtain the complete historical observational data from their original publicly available sources as detailed in our documentation. By combining these original observations with our reconstructed data, a comprehensive and continuous daily dataset from 1950 to 2023 can be assembled. Clear instructions and links for downloading the original observational data used in this study can be found at: <a href="https://github.com/zhwang24/RADIT-Reconstructed-Arctic-River-Data" target="_blank" rel="noopener">https://github.com/zhwang24/RADIT-Reconstructed-Arctic-River-Data</a>. Should you encounter any issues or have questions, please feel free to contact the first author, Zihan Wang (zhwang2018@163.com).</p>
ML-TOMCAT V2.0: Machine-Learning-Based Satellite-Corrected Global Stratospheric Ozone Profile Dataset
<p>MLTOMCAT V2 is 46 years (1979-2024) of gap free ozone profile data sets that is created by correcting biases in a TOMCAT Chemical Transport Model (CTM) simulated ozone profiles. We use Random Forest regression model to correct model biases. </p> <p>Each file contain monthly mean zonal mean ozone profiles. There are 6 data files.</p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_vmr_V2.nc</a> contains ozone profiles on geometric height levels (1 to 60 km) in mixing ratio units, whereas <a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_ht_nd_V2.nc</a> contains ozone profile in number density units.</p> <p>Similarly, </p> <p><a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_vmr_V2.nc</a> contains ozone profiles on 43 MLS pressure levels (1000 to 0.1 hPa) in mixing ratio units, whereas <a href="https://zenodo.org/api/files/0416f9bd-c908-4d3f-9368-47ca2e04d7bd/MLTOMCAT_1979_2020_72_ht_vmr.nc">MLTOMCAT_1979_2024_72_plev_nd_V2.nc</a> contains ozone profile in number density units.</p> <p>Please note that data below 300 hPa (~8km) and 1 hPa (~50 km) should be used with caution.</p> <p>There are two straospheric column files</p> <p>ML-TOMCAT-SCO_120ppb_boundary_V2_197901-202412.nc and</p> <p>ML-TOMCAT-SCO_150ppb_boundary_V2_197901-202412.nc</p> <p>Stratospheric column files calculated using 120 ppb and 150 ppb as a chemical ozone boundaries.</p> <p>A manuscript describing MLTOMCAT would be published in EESD (Dhomse et al., 2021).</p>
Dataset for "Machine learning predictions on an extensive geotechnical dataset of laboratory tests in Austria"
<p>This dataset comprises over 20 years of geotechnical laboratory testing data collected primarily from Vienna, Lower Austria, and Burgenland. It includes 24 features documenting critical soil properties derived from particle size distributions, Atterberg limits, Proctor tests, permeability tests, and direct shear tests. Locations for a subset of samples are provided, enabling spatial analysis.</p> <p>The dataset is a valuable resource for geotechnical research and education, allowing users to explore correlations among soil parameters and develop predictive models. Examples of such correlations include liquidity index with undrained shear strength, particle size distribution with friction angle, and liquid limit and plasticity index with residual friction angle.</p> <p>Python-based exploratory data analysis and machine learning applications have demonstrated the dataset's potential for predictive modeling, achieving moderate accuracy for parameters such as cohesion and friction angle. Its temporal and spatial breadth, combined with repeated testing, enhances its reliability and applicability for benchmarking and validating analytical and computational geotechnical methods.</p> <p>This dataset is intended for researchers, educators, and practitioners in geotechnical engineering. Potential use cases include refining empirical correlations, training machine learning models, and advancing soil mechanics understanding. Users should note that preprocessing steps, such as imputation for missing values and outlier detection, may be necessary for specific applications.</p> <p><strong>Key Features</strong>:</p> <ul> <li><strong>Temporal Coverage</strong>: Over 20 years of data.</li> <li><strong>Geographical Coverage</strong>: Vienna, Lower Austria, and Burgenland.</li> <li><strong>Tests Included</strong>: <ul> <li>Particle Size Distribution</li> <li>Atterberg Limits</li> <li>Proctor Tests</li> <li>Permeability Tests</li> <li>Direct Shear Tests</li> </ul> </li> <li><strong>Number of Variables</strong>: 24</li> <li><strong>Potential Applications</strong>: Correlation analysis, predictive modeling, and geotechnical design.</li> </ul> <p><strong>Technical Details</strong>:</p> <ul> <li>Missing values have been addressed using K-Nearest Neighbors (KNN) imputation, and anomalies identified using Local Outlier Factor (LOF) methods in previous studies.</li> <li>Data normalization and standardization steps are recommended for specific analyses.</li> </ul> <p><strong>Acknowledgments</strong>:<br>The dataset was compiled with support from the European Union's MSCA Staff Exchanges project 101182689 Geotechnical Resilience through Intelligent Design (GRID).</p>
Network Digital Twin-Generated Dataset for Machine Learning-based Detection of Benign and Malicious Heavy Hitter Flows
<h3>Overview</h3> <p>This record provides a dataset created as part of the study presented in the following publication and is made <strong>publicly available for research purposes</strong>. The associated article provides a comprehensive description of the dataset, its structure, and the methodology used in its creation. If you use this dataset, please <strong>cite the following article </strong>published in the journal <strong>IEEE Communications Magazine</strong>:</p> <blockquote> <p><strong>A. Karamchandani, J. Nunez, L. de-la-Cal, Y. Moreno, A. Mozo, and A. Pastor, “On the Applicability of Network Digital Twins in Generating Synthetic Data for Heavy Hitter Discrimination,” IEEE Communications Magazine, pp. 2–8, 2025, DOI: 10.1109/MCOM.003.2400648.</strong></p> </blockquote> <p>More specifically, the record contains several synthetic datasets generated to differentiate between benign and malicious heavy hitter flows within a realistic virtualized network environment. Heavy Hitter flows, which include high-volume data transfers, can significantly impact network performance, leading to congestion and degraded quality of service. Distinguishing legitimate heavy hitter activity from malicious Distributed Denial-of-Service traffic is critical for network management and security, yet existing datasets lack the granularity needed for training machine learning models to effectively make this distinction.</p> <p>To address this, a Network Digital Twin (NDT) approach was utilized to emulate realistic network conditions and traffic patterns, enabling automated generation of labeled data for both benign and malicious HH flows alongside regular traffic.</p> <h3>Feature Set:</h3> <p>The feature set includes the following flow statistics commonly used in the literature on network traffic classification:</p> <ul> <li>The protocol used for the connection, identifying whether it is TCP, UDP, ICMP, or OSPF.</li> <li>The time (relative to the connection start) of the most recent packet sent from source to destination at the time of each snapshot.</li> <li>The time (relative to the connection start) of the most recent packet sent from destination to source at the time of each snapshot.</li> <li>The cumulative count of data packets sent from source to destination at the time of each snapshot.</li> <li>The cumulative count of data packets sent from destination to source at the time of each snapshot.</li> <li>The cumulative bytes sent from source to destination at the time of each snapshot.</li> <li>The cumulative bytes sent from destination to source at the time of each snapshot.</li> <li>The time difference between the first packet sent from source to destination and the first packet sent from destination to source.</li> </ul> <h3>Dataset Variations:</h3> <p>To accommodate diverse research needs and scenarios, the dataset is provided in the following variations:</p> <ol> <li> <p><strong><code>All at Once</code></strong>:</p> <ol> <li>Contains a synthetic dataset where all traffic types, including benign, normal, and malicious DDoS heavy hitter (HH) flows, are combined into a single dataset.</li> <li>This version represents a holistic view of the traffic environment, simulating real-world scenarios where all traffic occurs simultaneously.</li> </ol> </li> <li> <p><strong><code>Balanced Traffic Generation</code></strong>:</p> <ol> <li>Represents a balanced traffic dataset with an equal proportion of benign, normal, and malicious DDoS traffic.</li> <li>Designed for scenarios where a balanced dataset is needed for fair training and evaluation of machine learning models.</li> </ol> </li> <li> <p><strong><code>DDoS at Intervals</code></strong>:</p> <ol> <li>Contains traffic data where malicious DDoS HH traffic occurs at specific time intervals, mimicking real-world attack patterns.</li> <li>Useful for studying the impact and detection of intermittent malicious activities.</li> </ol> </li> <li> <p><strong><code>Only Benign HH Traffic</code></strong>:</p> <ol> <li>Includes only benign HH traffic flows.</li> <li>Suitable for training and evaluating models to identify and differentiate benign heavy hitter traffic patterns.</li> </ol> </li> <li> <p><strong><code>Only DDoS Traffic</code></strong>:</p> <ol> <li>Contains only malicious DDoS HH traffic.</li> <li>Helps in isolating and analyzing attack characteristics for targeted threat detection.</li> </ol> </li> <li> <p><strong><code>Only Normal Traffic</code></strong>:</p> <ol> <li>Comprises only regular, non-HH traffic flows.</li> <li>Useful for understanding baseline network behavior in the absence of heavy hitters.</li> </ol> </li> <li> <p><strong><code>Unbalanced Traffic Generation</code></strong>:</p> <ol> <li>Features an unbalanced dataset with varying proportions of benign, normal, and malicious traffic.</li> <li>Simulates real-world scenarios where certain types of traffic dominate, providing insights into model performance in unbalanced conditions.</li> </ol> </li> </ol> <p>For each variation, the output of the different packet aggregators is provided separated in its respective folder.</p> <p>Each variation was generated using the NDT approach to demonstrate its flexibility and ensure the reproducibility of our study's experiments, while also contributing to future research on network traffic patterns and the detection and classification of heavy hitter traffic flows. The dataset is designed to support research in network security, machine learning model development, and applications of digital twin technology.</p>
Machine learning for Gravity Spy: Glitch classification and dataset
<p>We present the first version of the training set used in the Gravity Spy citizen science project. This training set, discussed in detail <a href="https://www.sciencedirect.com/science/article/pii/S0020025518301634">here</a>, was utilized to train the convolutional neural network employed in the Gravity Spy project. We anticipate moving forward to release more labelled Gravity Spy data sets, including a refined version of this training set which can be found here <a href="https://doi.org/10.5281/zenodo.1476551">10.5281/zenodo.1476551</a>, and data sets containing the annotations provided by our citizen science volunteers.</p> <p><strong>Data Set Information</strong></p> <p>There are three files provided in this data set</p> <ul> <li><strong>trainingset_v1d0_metadata.csv</strong> <ul> <li>This file has three columns, <em>gravityspy_id, label, </em>and <em>sample_type.</em><em> gravityspy_id </em>is the unique 10 character hash given to every Gravity Spy sample. <em>label</em> is the string label of the sample. <em>sample_type </em>indicates whether this sample was used in the paper for testing training or validating the models. This is provided for those who would like to do direct comparisons to the network described in the paper.</li> </ul> </li> <li><strong>trainingsetv1d0.h5</strong> <ul> <li>This file contains the exact arrays used in the paper for every Gravity Spy sample. Each Gravity Spy sample is defined by four different images with varying temporal duration, <em>0.5, 1.0, 2.0, and 4.0</em> second, respectively. This also determines the naming conventions of the PNGs: <em>interferometer_gravityspyid_spectrogram_duration.png (e.g. H1_Fv3p6eROvA_spectrogram_0.5.png, H1_Fv3p6eROvA_spectrogram_1.0.png, H1_Fv3p6eROvA_spectrogram_2.0.png, H1_Fv3p6eROvA_spectrogram_4.0.png</em>).</li> <li>This file contains all the information needed for each sample in the Gravity Spy dataset (i.e. the label, the sample type of the sample, the unique id of the sample, and the image data for that sample. <ul> <li>/1080Lines/validation/xUEyaWr34c Group<br> /1080Lines/validation/xUEyaWr34c/0.5.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/1.0.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/2.0.png Dataset {1, 140, 170}<br> /1080Lines/validation/xUEyaWr34c/4.0.png Dataset {1, 140, 170}</li> </ul> </li> </ul> </li> <li><strong>trainingsetv1d0.tar.gz</strong> <ul> <li>Contains the raw PNGs of the Gravity Spy training set.</li> <li>The structure of the folder is <em>/"label"/"sample_type"/"pngs"</em></li> </ul> </li> </ul> <p><strong>Data Set Parsing Information</strong></p> <p>To read and crop out the plot axis and labels of the provided PNGs, the following small python code using scikit-image should work.</p> <p>from skimage import io</p> <p>image_data = io.imread("filename_of_image")</p> <p>x=[66, 532]; y=[105, 671]</p> <p>image_data = image_data[x[0]:x[1], y[0]:y[1], :3]</p>
Voxelized fragment dataset for machine learning
<p>One of the primary challenges inherent in utilizing deep learning models is the scarcity and accessibility hurdles associated with acquiring datasets of sufficient size to facilitate effective training of these networks. This is particularly significant in object detection, shape completion, and fracture assembly. Instead of scanning a large number of real-world fragments, it is possible to generate massive datasets with synthetic pieces. However, realistic fragmentation is computationally intensive in the preparation (e.g., pre-factured models) and generation. Otherwise, simpler algorithms such as Voronoi diagrams provide faster processing speeds at the expense of compromising realism. Hence, it is required to balance computational efficiency and realism for generating large datasets for marching learning.</p> <p>We proposed a GPU-based fragmentation method to improve the baseline Discrete Voronoi Chain aimed at completing this dataset generation task. The dataset in this repository includes voxelized fragments from high-resolution 3D models, curated to be used as training sets for machine learning models. More specifically, these models come from an archaeological dataset, which led to more than 1M fragments from 1,052 Iberian vessels. In this dataset, fragments are not stored individually; instead, the fragmented voxelizations are provided in a compressed binary file (.rle.zip). Once uncompressed, each fragment is represented by a different number in the grid. The class to which each vessel belongs is also included in <em>class.csv</em>. The GPU-based pipeline that generated this dataset is explained at <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.cag.2024.104104" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.cag.2024.104104</a>.</p> <p>Please, note that this dataset originally provided voxel data, point clouds and triangle meshes. However, we opted for including only voxel data because 1) the original dataset is too large to be uploaded to Zenodo and 2) the original intent of our paper is to generate implicit data in the form of voxels. If interested in the whole dataset (450GB), please visit the web page of our <a href="https://s5-ceatic.ujaen.es/fragment-dataset-uja/">research institute</a>.</p>
Improving machine-learning models in materials science through large datasets
<p>1. Image of the <a href="https://alexandria.icams.rub.de/"><strong>Alexandria database </strong></a> state corresponding to the paper "<strong>Improving machine-learning models in materials science through large datasets</strong>".</p> <ul> <li>Static pbe calculations for 1D, 2D, 3D compounds can be found in 1D_pbe.tar.gz, 2D_pbe.tar.gz, 3D_pbe.tar.gz in batches of 100k materials. The latter also contains a separate convex hull pickle with all compounds on the pbe convex hull (convex_hull_pbe_2023.12.29.json.bz2) and a list of prototypes in the database (prototypes.json.bz2). The systematic 3D calculations performed for the article <strong>Improving machine-learning models in materials science through large datasets </strong>(in the paper referred to as round 2 and 3) can be found by the location keyword in the data dictionary of each ComputedStructureEntry containing "<strong>cgat_comp/quaternaries</strong>" (round 2) and "<strong>cgat_comp2/</strong>" (round 3). Round 1 (10.1002/adma.202210788) can be found under "cgat_comp/ternaries", ""cgat_comp/binaries".</li> <li>Static pbesol calculations for 3D compounds can be found in 3D_ps.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the pbesol convex hull (convex_hull_ps_2023.12.29.json.bz2). </li> <li>Static scan calculations for 3D compounds can be found in 3D_scan.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the scan convex hull (convex_hull_scan_2023.12.29.json.bz2). </li> <li>Geometry relaxation curves for 1D and 2D and 3D compounds calculated with PBE can be found in geo_opt_1D.tar.gz, geo_opt_2D.tar.gz. and geo_opt_3D.tar. Each file in each folder contains a batch of up to 10k relaxation trajectories.</li> <li>PBESOL relaxation trajectories for 3D compounds can be found in geo_opt_ps.tar</li> </ul> <p>2. Crystal graph attention networks to predict the volume (<a href="https://zenodo.org/api/records/12582650/draft/files/volume_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">volume_round_3.tar.gz</a>) and distance to the convex hull (<a href="https://zenodo.org/api/records/12582650/draft/files/e_above_hull_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">e_above_hull_round_3.tar.gz</a>) trained for the paper "Improving machine-learning models in materials science through large datasets".</p> <p>Can be used with the code at https://github.com/hyllios/CGAT/tree/main/CGAT.<br><strong>Note will predict the distance to the convex hull not normalized per atom when using the code on the github.<br></strong></p> <p>3. Alignn models as well as m3gnet and mace models corresponding to the publication can be found in <a href="https://zenodo.org/api/records/12582650/draft/files/alexandria_v2.tar.gz/content" target="_blank" rel="noopener noreferrer">alexandria_v2.tar.gz</a></p> <p>4. scripts.tar.gz Some scripts used for generating CGAT input data/ performing parallel predictions and for relaxations with m3gnet/mace force fields</p>
Dataset for: Importance of satellite observations for high-resolution mapping of near-surface NO2 by machine learning
<p>Dataset for: Importance of satellite observations for high-resolution mapping of near-surface NO<sub>2 </sub>by machine learning</p> <p>This dataset is uploaded as a part of the article by Kim et al. (2021). The dataset is the hourly maps of near-surface nitrogen dioxide (NO<sub>2</sub>) concentrations at 100 m resolution for an Alpine domain (Switzerland and northern Italy, 6-12 °E, 42-48 °N). The dataset is provided per day (24 hours) in a netcdf (*.nc ~550MB). In this work, we have generated NO<sub>2 </sub>hourly maps for Feb. 2019 to May 2020 and, here, we upload for March 2019 only (~16 GB). If you need data for another period of time, please contact Gerrit Kuhlmann (gerrit.kuhlmann@empa.ch) or Minsu Kim (minsu.kim@empa.ch). </p>
ManyTypes4Py: A Benchmark Python Dataset for Machine Learning-Based Type Inference
<ul> <li>The dataset is gathered on Sep. 17th 2020 from GitHub.</li> <li>It has <em>clean</em> and <em>complete</em> versions (from v0.7): <ul> <li>The clean version has 5.1K <strong>type-checked </strong>Python repositories and 1.2M type annotations.</li> <li>The complete version has 5.2K Python repositories and 3.3M type annotations.</li> </ul> </li> <li>The dataset's source files are type-checked using <a href="https://mypy.readthedocs.io/">mypy</a> (clean version).</li> <li>The dataset is also de-duplicated using the <a href="https://github.com/saltudelft/CD4Py">CD4Py</a> tool.</li> <li>Check out the <strong>README.MD</strong> file for the description of the dataset.</li> <li>Notable changes to each version of the dataset are documented in <strong>CHANGELOG.md</strong>.</li> <li>The dataset's scripts and utilities are available on <a href="https://github.com/saltudelft/many-types-4-py-dataset">its GitHub repository</a>.</li> </ul>
An analyst-created, high-resolution seismic arrival time dataset for evaluating machine learning phase detectors
<p>This data set contains phase arrival times and phase labels for one hour of continuous seismic data recorded at the three-component broadband station WY.YNR on 2014 30 March from 13:00:00 to 14:00:00 UTC. This hour of data follows a M<sub>w</sub> 4.8 occurring at 12:34 UTC in the Yellowstone region and contains many small events close together in space and time. All 687 picks (404 P and 283 S) were made by a seismic analyst trained at the University of Utah Seismograph Stations. The goal was to pick as many reasonable arrivals as possible for evaluating the performance of machine-learning-based phase detectors on continuous data during high seismicity rates. Phase picks were made using the Seismic Analysis Code (SAC; Goldstein and Snoke, 2005) and Pyrocko (Heimann <em>et al</em>., 2017). In general, a 1 to 17 Hz bandpass filter was used.</p> <p>The csv file contains the pick id, phase arrival time in UTC and Unix epoch formats, and the phase labels (P or S).</p>
Datasets and codes for the peer review article "Human and natural impacts on the U.S. freshwater salinization and alkalinization: A machine learning approach"
<p>Ongoing salinization and alkalinization in U.S. rivers have been attributed to inputs of road salt and effects of human-accelerated weathering in previous studies. Salinization poses a severe threat to human and ecosystem health, while human derived alkalinization implies increasing uncertainty in the dynamics of terrestrial sequestration of atmospheric carbon dioxide. A mechanistic understanding of whether and how human activities accelerate weathering and contribute to the geochemical changes in U.S. rivers is lacking. To address this uncertainty, we compiled dissolved sodium (salinity proxy) and alkalinity values along with 32 watershed properties ranging from hydrology, climate, geomorphology, geology, soil chemistry, land use, and land cover for 226 river monitoring sites across the coterminous U.S. Using these data, we built two machine-learning models to predict monthly-aggregated sodium and alkalinity fluxes at these sites. The sodium-prediction model detected human activities (represented by population density and impervious surface area) as major contributors to the salinity of U.S. rivers. In contrast, the alkalinity-prediction model identified natural processes as predominantly contributing to variation in riverine alkalinity flux, including runoff, carbonate sediment or siliciclastic sediment, soil pH and soil moisture. Unlike prior studies, our analysis suggests that the alkalinization in U.S. rivers is largely governed by local climatic and hydrogeological conditions.</p>
Datasets for "Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Classification Algorithms in Reiner Gamma and Mare Ingenii"
<p>Final surface reflectance data at 2.6 m/pixel resolution with floating point values are available as GeoTiff and ASCII text files. Definition files for the K-Means and MLC algorithms in classifying swirl units are also available as ASCII text files. See README file for further details.</p> <p>Data used in the research article:</p> <p>Chuang, F.C., M.D. Richardson, J.R. Weirich, A.A. Sickafoose, and D.L. Domingue, 2022. Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Image Classification Algorithms in Reiner Gamma and Mare Ingenii. The Planetary Science Journal, 3:231. doi://10.3847/PSJ/ac8f43</p> <p> </p>
PFAM Protein Families Dataset for Machine Learning
<p>A cleaned dataset of protein sequences and protein families for classification. The dataset is exported from PFAM as of June 2023 and curated to achieve the following characteristics:</p> <ul> <li>only protein families included with >=100 sequences</li> <li>families with >2000 sequences are truncated and only represented by 2000 sequences (chosen randomly)</li> <li>only proteins with sequence lengths between 100 and 1000</li> <li>amino acid sequences are form PDB; chains are concatenated only if not similar</li> </ul> <p>The dataset is not balanced, numbers of sequences per family in PFAM and in in dataset are:</p> <pre><code>families: 62, sequences: 46872 total (in PFAM) -> included (in dataset) Number in family ALLERGEN: 122 -> 122 Number in family APOPTOSIS: 381 -> 381 Number in family BIOSYNTHETIC PROTEIN: 346 -> 346 Number in family BIOTIN BINDING PROTEIN: 165 -> 165 Number in family BLOOD CLOTTING: 138 -> 138 Number in family CALCIUM BINDING PROTEIN: 135 -> 135 Number in family CELL ADHESION: 1116 -> 1116 Number in family CELL CYCLE: 511 -> 511 Number in family CHAPERONE: 964 -> 964 Number in family CONTRACTILE PROTEIN: 158 -> 158 Number in family CYTOKINE: 191 -> 191 Number in family DE NOVO PROTEIN: 253 -> 253 Number in family DNA BINDING PROTEIN: 1008 -> 1008 Number in family ELECTRON TRANSPORT: 841 -> 841 Number in family FLUORESCENT PROTEIN: 348 -> 348 Number in family GENE REGULATION: 607 -> 607 Number in family HORMONE: 272 -> 272 Number in family HORMONE GROWTH FACTOR: 159 -> 159 Number in family HORMONE RECEPTOR: 121 -> 121 Number in family HYDROLASE: 19551 -> 2000 Number in family HYDROLASE ANTIBIOTIC: 120 -> 120 Number in family HYDROLASE HYDROLASE INHIBITOR: 2890 -> 2000 Number in family HYDROLASE INHIBITOR: 315 -> 315 Number in family IMMUNE SYSTEM: 3333 -> 2000 Number in family IMMUNOGLOBULIN: 155 -> 155 Number in family ISOMERASE: 2457 -> 2000 Number in family ISOMERASE ISOMERASE INHIBITOR: 139 -> 139 Number in family LECTIN: 139 -> 139 Number in family LIGASE: 1780 -> 1780 Number in family LIGASE LIGASE INHIBITOR: 163 -> 163 Number in family LIPID BINDING PROTEIN: 421 -> 421 Number in family LIPID TRANSPORT: 115 -> 115 Number in family LUMINESCENT PROTEIN: 221 -> 221 Number in family LYASE: 4150 -> 2000 Number in family LYASE LYASE INHIBITOR: 298 -> 298 Number in family MEMBRANE PROTEIN: 1338 -> 1338 Number in family METAL BINDING PROTEIN: 951 -> 951 Number in family METAL TRANSPORT: 409 -> 409 Number in family MOTOR PROTEIN: 195 -> 195 Number in family OXIDOREDUCTASE: 11531 -> 2000 Number in family OXIDOREDUCTASE OXIDOREDUCTASE INHIBITOR: 766 -> 766 Number in family OXYGEN STORAGE: 127 -> 127 Number in family OXYGEN STORAGE TRANSPORT: 260 -> 260 Number in family OXYGEN TRANSPORT: 414 -> 414 Number in family PHOTOSYNTHESIS: 173 -> 173 Number in family PLANT PROTEIN: 255 -> 255 Number in family PROTEIN BINDING: 1613 -> 1613 Number in family PROTEIN TRANSPORT: 693 -> 693 Number in family RECEPTOR: 108 -> 108 Number in family REPLICATION: 161 -> 161 Number in family RNA BINDING PROTEIN: 546 -> 546 Number in family SIGNALING PROTEIN: 2312 -> 2000 Number in family STRUCTURAL PROTEIN: 869 -> 869 Number in family SUGAR BINDING PROTEIN: 1250 -> 1250 Number in family TOXIN: 546 -> 546 Number in family TRANSCRIPTION REGULATION: 3283 -> 2000 Number in family TRANSFERASE: 14724 -> 2000 Number in family TRANSFERASE INHIBITOR: 126 -> 126 Number in family TRANSFERASE TRANSFERASE INHIBITOR: 2465 -> 2000 Number in family TRANSLATION: 370 -> 370 Number in family TRANSPORT PROTEIN: 2782 -> 2000 Number in family VIRAL PROTEIN: 2150 -> 2000</code></pre> <p>Files:</p> <ul> <li>families.csv: list of protein families with frequencies</li> <li>pfam_46872x62.csv: full dataset with amino acid sequences as string (one-letter code)</li> <li>pfam-trn-xy.csv: training dataset with amino acid sequences as tokens (1..25) and padded to a common length of 1000 with padding token 0:</li> </ul> <pre><code> Amino acid | Token | Description -------------------------------- C | 1 | Cysteine S | 2 | Serine T | 3 | Threonine A | 4 | Alanine G | 5 | Glycine P | 6 | Proline D | 7 | Aspartic acid E | 8 | Glutamic acid Q | 9 | Glutamine N | 10 | Asparagine H | 11 | Histidine R | 12 | Arginine K | 13 | Lysine M | 14 | Methionine I | 15 | Isoleucine L | 16 | Leucine V | 17 | Valine W | 18 | Tryptophan Y | 19 | Tyrosine F | 20 | Phenylalanine B | 21 | Aspartic acid or Asparagine Z | 22 | Glutamic acid or Glutamine J | 23 | Leucine or Isoleucine U | 24 | Selenocysteine X | 25 | Unknown amino acid . | 0 | padding token</code></pre> <p> </p> <ul> <li>pfam-trn-labels.csv: plain-text labels for training data</li> <li>pfam-tst-xy.csv</li> <li>pfam-tst-labels.csv: test data</li> <li>pfam-balanced-trn-xy.csv</li> <li>pfam-balanced-trn-labels.csv:</li> <li>pfam-balanced-tst-xy.csv</li> <li>pfam-balanced-tst-labels.csv: balanced datasets, created by oversampling.</li> </ul>
A Pan-European, Quantile Machine learning (QML) based, Total, Fine-Mode and Coarse-Mode Aerosol Optical Depth dataset (QML AOD))
<p>The V 1.1.0 product is an improved Aerosol Optical Depth (AOD) product based on Gap-filled MAIAC AOD, which provide first full-coverage, high-resolution monitoring of fine-mode and coarse-mode aerosols in Europe from 2003-20. This dataset has successfully rectified the previously identified issue of weak associations between satellite AOD and PM2.5 in Europe, which was primarily attributable to current limitations of AOD data. Our innovative approach has yielded stronger correlations with PM10, PM2.5, and PMcoarse than previous AOD product, laying a critical groundwork for improving PM10, PM2.5, and PMcoarse predictions in further epidemiological studies or environmental monitoring.</p> <p>We have uploaded three QML AOD datasets in Geotiff format, covering the region from -27° to 72° latitude and from -25° to 45° longitude. These datasets will be useful for researchers and policymakers to better understand the impacts of aerosols on the environment and human health.</p> <p> Note: v1.0.0 product do not include MAIAC AOD in their models.</p> <p>Please read more details in our paper </p> <h1><span>Estimation of pan-European, daily total, fine-mode and coarse-mode Aerosol Optical Depth at 0.1° resolution to facilitate air quality assessments</span></h1> <p><a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.scitotenv.2024.170593" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.scitotenv.2024.170593</span></a></p>
Dataset for Machine Learning Framework for Modeling Exciton-Polaritons in Molecular Materials
<p>The data consists of several NumPy arrays saved in the binary format (npy files) with a total size of 108 MB. The details of these files are listed below.</p> <table> <tbody> <tr> <td> <p><strong>Filename and path</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>training/azo_R.npy</p> </td> <td> <p>Coordinates for training</p> </td> </tr> <tr> <td> <p>training/azo_Z.npy</p> </td> <td> <p>Atomic indices for training</p> </td> </tr> <tr> <td> <p>training/azo_E.npy</p> </td> <td> <p>Molecular energies for training</p> </td> </tr> <tr> <td> <p>training/azo_D.npy</p> </td> <td> <p>Transition dipoles for training</p> </td> </tr> <tr> <td> <p>training/azo_ScaledNACR.npy</p> </td> <td> <p>Non-adiabatic coupling vectors scaled by energy difference for training</p> </td> </tr> <tr> <td> <p>scan/azo_R.npy</p> </td> <td> <p>Coordinates for PES scan</p> </td> </tr> <tr> <td> <p>scan/azo_Z.npy</p> </td> <td> <p>Atomic indices for PES scan</p> </td> </tr> <tr> <td> <p>scan/azo_E.npy</p> </td> <td> <p>Molecular energies for PES scan</p> </td> </tr> <tr> <td> <p>scan/azo_D.npy</p> </td> <td> <p>Transition dipoles for PES scan</p> </td> </tr> <tr> <td> <p>scan/azo_ScaledNACR.npy</p> </td> <td> <p>Non-adiabatic coupling vectors scaled by energy difference for PES scan</p> </td> </tr> <tr> <td> <p>spectrum/azo_R.npy</p> </td> <td> <p>Coordinates for spectrum calculations</p> </td> </tr> <tr> <td> <p>spectrum/azo_Z.npy</p> </td> <td> <p>Atomic indices for spectrum calculations</p> </td> </tr> <tr> <td> <p>spectrum/azo_E.npy</p> </td> <td> <p>Molecular energies for spectrum calculations</p> </td> </tr> <tr> <td> <p>spectrum/azo_D.npy</p> </td> <td> <p>Transition dipoles for spectrum calculation</p> </td> </tr> </tbody> </table>
Dataset for "Machine Learning Stability and Bandgaps of Lead-Free Perovskites for Photovoltaics"
<p>Datasets used in the publication "Machine Learning Stability and Bandgaps of Lead-Free Perovskites for Photovoltaics" [doi:10.1002/adts.201900178].</p> <p>All structures were relaxed with the following parameters using Quantumwise QATK 2017:</p> <p>- SG15-GGA norm-conserving (Vanderbilt) pseudopotentials employed in a LCAO-approach (200 Hartree cutoff)<br> - 2x1x2-cubic-perovskite-supercells, relaxed from cubic 11.4Åx5.7Åx11.4Å-structures (forces < 0.01eV/Å)<br> - 300K Fermi-Dirac-smearing<br> - a 6x12x6 k-point grid (Monkhorst-Pack)</p> <p><br> Specifically, the included files are:</p> <p><strong>db_2.data: </strong>the actual database used for model building (json-format)<br> <strong>lead_set.data:</strong> the "external" test set used to test predictive power with out of sample compounds (json-format)<br> <strong>load_stanley_c.py:</strong> a python script to parse the .json-files to a python-dictionary including the structures (relaxed and unrelaxed) as <a href="https://gitlab.com/ase/ase">ASE</a>-atoms</p> <p>The format of the datafiles is as follows (-1 generally denote values not parsed from the raw data):<br> {<br> "<idstring>" : {<br> "trajectory" : n/a,<br> "energy" : total DFT energy in eV,<br> "rstruc" : relaxed structure, 3-tuple: (cell-vectors, scaled_positions, elements),<br> "gaps" : { "opt_gap", "ind_gap } - both direct and indirect gap,<br> "effective_mass" : n/a,<br> "iterations" : number of relaxation steps,<br> "calc" : some calculation metadata,<br> "ustruc" : unrelaxed input structure,<br> <br> }<br> }<br> Missing ids relate to structures filtered out, because the calculation didn't converge.</p> <p>Some code which works with a different representation of this data can be found at https://github.com/jstanai/Machine-Learning-Perovskite-Properties-for-Photovoltaics</p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.