Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.7.1
Dataset results
1,943 results for “Machine Learning”
ALICE machine learning image set
<p>Pinned insect images and corresponding label outlines in JSON format. This image set is used for machine learning of label identification and segmentation for the ALICE project. https://doi.org/10.31219/osf.io/s2p73</p> <p> </p>
Voxelized fragment dataset for machine learning
<p>One of the primary challenges inherent in utilizing deep learning models is the scarcity and accessibility hurdles associated with acquiring datasets of sufficient size to facilitate effective training of these networks. This is particularly significant in object detection, shape completion, and fracture assembly. Instead of scanning a large number of real-world fragments, it is possible to generate massive datasets with synthetic pieces. However, realistic fragmentation is computationally intensive in the preparation (e.g., pre-factured models) and generation. Otherwise, simpler algorithms such as Voronoi diagrams provide faster processing speeds at the expense of compromising realism. Hence, it is required to balance computational efficiency and realism for generating large datasets for marching learning.</p> <p>We proposed a GPU-based fragmentation method to improve the baseline Discrete Voronoi Chain aimed at completing this dataset generation task. The dataset in this repository includes voxelized fragments from high-resolution 3D models, curated to be used as training sets for machine learning models. More specifically, these models come from an archaeological dataset, which led to more than 1M fragments from 1,052 Iberian vessels. In this dataset, fragments are not stored individually; instead, the fragmented voxelizations are provided in a compressed binary file (.rle.zip). Once uncompressed, each fragment is represented by a different number in the grid. The class to which each vessel belongs is also included in <em>class.csv</em>. The GPU-based pipeline that generated this dataset is explained at <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.cag.2024.104104" target="_blank" rel="noreferrer noopener">https://doi.org/10.1016/j.cag.2024.104104</a>.</p> <p>Please, note that this dataset originally provided voxel data, point clouds and triangle meshes. However, we opted for including only voxel data because 1) the original dataset is too large to be uploaded to Zenodo and 2) the original intent of our paper is to generate implicit data in the form of voxels. If interested in the whole dataset (450GB), please visit the web page of our <a href="https://s5-ceatic.ujaen.es/fragment-dataset-uja/">research institute</a>.</p>
Improving machine-learning models in materials science through large datasets
<p>1. Image of the <a href="https://alexandria.icams.rub.de/"><strong>Alexandria database </strong></a> state corresponding to the paper "<strong>Improving machine-learning models in materials science through large datasets</strong>".</p> <ul> <li>Static pbe calculations for 1D, 2D, 3D compounds can be found in 1D_pbe.tar.gz, 2D_pbe.tar.gz, 3D_pbe.tar.gz in batches of 100k materials. The latter also contains a separate convex hull pickle with all compounds on the pbe convex hull (convex_hull_pbe_2023.12.29.json.bz2) and a list of prototypes in the database (prototypes.json.bz2). The systematic 3D calculations performed for the article <strong>Improving machine-learning models in materials science through large datasets </strong>(in the paper referred to as round 2 and 3) can be found by the location keyword in the data dictionary of each ComputedStructureEntry containing "<strong>cgat_comp/quaternaries</strong>" (round 2) and "<strong>cgat_comp2/</strong>" (round 3). Round 1 (10.1002/adma.202210788) can be found under "cgat_comp/ternaries", ""cgat_comp/binaries".</li> <li>Static pbesol calculations for 3D compounds can be found in 3D_ps.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the pbesol convex hull (convex_hull_ps_2023.12.29.json.bz2). </li> <li>Static scan calculations for 3D compounds can be found in 3D_scan.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the scan convex hull (convex_hull_scan_2023.12.29.json.bz2). </li> <li>Geometry relaxation curves for 1D and 2D and 3D compounds calculated with PBE can be found in geo_opt_1D.tar.gz, geo_opt_2D.tar.gz. and geo_opt_3D.tar. Each file in each folder contains a batch of up to 10k relaxation trajectories.</li> <li>PBESOL relaxation trajectories for 3D compounds can be found in geo_opt_ps.tar</li> </ul> <p>2. Crystal graph attention networks to predict the volume (<a href="https://zenodo.org/api/records/12582650/draft/files/volume_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">volume_round_3.tar.gz</a>) and distance to the convex hull (<a href="https://zenodo.org/api/records/12582650/draft/files/e_above_hull_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">e_above_hull_round_3.tar.gz</a>) trained for the paper "Improving machine-learning models in materials science through large datasets".</p> <p>Can be used with the code at https://github.com/hyllios/CGAT/tree/main/CGAT.<br><strong>Note will predict the distance to the convex hull not normalized per atom when using the code on the github.<br></strong></p> <p>3. Alignn models as well as m3gnet and mace models corresponding to the publication can be found in <a href="https://zenodo.org/api/records/12582650/draft/files/alexandria_v2.tar.gz/content" target="_blank" rel="noopener noreferrer">alexandria_v2.tar.gz</a></p> <p>4. scripts.tar.gz Some scripts used for generating CGAT input data/ performing parallel predictions and for relaxations with m3gnet/mace force fields</p>
Toward a Generalizable Machine-Learned Potential for Metal-Organic Frameworks
<ul> <li>This repository contains the dataset used in the publication<br> `Toward Generalizable Machine Learned Potential for Metal-Organic Frameworks` Yue Yifei, Saad Aldin Mohammed, Loh Duane*, Jiang Jianwen*<br> <br> Please each the README.md within each subfolder. For brevity, the data is organized into three sections<br> <br> 1. The dataset in DATASET<br> - The training and testing dataset, including structures of MOFs in extxyz format<br> <br> 2. The training output files and logs in NEQUIP-TRAIN<br> - The conda environment details, training scripts and logs<br> - Also Training and testing metrics in csv files<br> - This is split into two zip files NEQUIP-TRAIN1 and NEQUIP-TRAIN2 due to their size<br> <br> 3. Examples of using the developed models in MD simulations<br> - Including LAMMPS scripts, data file and environment details used in our scalability tests<br> - The complied Nequip-patched LAMMPS version is also provided<br> - Details on how to use our models - we used a default model that is slower but more accurate in our study but faster models are also developed.</li> </ul>
Dataset for: Importance of satellite observations for high-resolution mapping of near-surface NO2 by machine learning
<p>Dataset for: Importance of satellite observations for high-resolution mapping of near-surface NO<sub>2 </sub>by machine learning</p> <p>This dataset is uploaded as a part of the article by Kim et al. (2021). The dataset is the hourly maps of near-surface nitrogen dioxide (NO<sub>2</sub>) concentrations at 100 m resolution for an Alpine domain (Switzerland and northern Italy, 6-12 °E, 42-48 °N). The dataset is provided per day (24 hours) in a netcdf (*.nc ~550MB). In this work, we have generated NO<sub>2 </sub>hourly maps for Feb. 2019 to May 2020 and, here, we upload for March 2019 only (~16 GB). If you need data for another period of time, please contact Gerrit Kuhlmann (gerrit.kuhlmann@empa.ch) or Minsu Kim (minsu.kim@empa.ch). </p>
Machine learning designs non-hemolytic antimicrobial peptides
<p>The upload contains additional primary data associated with the publication, including raw data in the original file format whenever possible.</p> <p>Data content: HRMS, HPLC-MS, CD, MD, TEM</p>
ManyTypes4Py: A Benchmark Python Dataset for Machine Learning-Based Type Inference
<ul> <li>The dataset is gathered on Sep. 17th 2020 from GitHub.</li> <li>It has <em>clean</em> and <em>complete</em> versions (from v0.7): <ul> <li>The clean version has 5.1K <strong>type-checked </strong>Python repositories and 1.2M type annotations.</li> <li>The complete version has 5.2K Python repositories and 3.3M type annotations.</li> </ul> </li> <li>The dataset's source files are type-checked using <a href="https://mypy.readthedocs.io/">mypy</a> (clean version).</li> <li>The dataset is also de-duplicated using the <a href="https://github.com/saltudelft/CD4Py">CD4Py</a> tool.</li> <li>Check out the <strong>README.MD</strong> file for the description of the dataset.</li> <li>Notable changes to each version of the dataset are documented in <strong>CHANGELOG.md</strong>.</li> <li>The dataset's scripts and utilities are available on <a href="https://github.com/saltudelft/many-types-4-py-dataset">its GitHub repository</a>.</li> </ul>
Data and code related to the paper: "Programmable Droplet Microfluidics Based on Machine Learning and Acoustic Manipulation"
<p>This archive contains the raw data and Matlab scripts to reproduce the plots and supplementary movies for the paper:</p> <p>Kyriacos Yiannacou, Vipul Sharma and Veikko Sariola, "Programmable Droplet Microfluidics Based on Machine Learning and Acoustic Manipulation", <em>Langmuir</em> 2022, 38, 38, 11557–11564.</p> <p><a href="https://doi.org/10.1021/acs.langmuir.2c01061">Link to the paper</a>.</p> <p>The scripts were tested on Matlab R2021a on Windows.</p> <p>The acoustofluidic controller software is the same as in our previous paper and is archived <a href="https://doi.org/10.5281/zenodo.4593021">here</a>.</p> <p>Generally speaking, there is a folder containing the plotting scripts for each figure(s) and/or movie(s). Within each folder, the raw data files are under the folder `data/`. Once ran, the scripts produce another folder called `output/`, to which they place the created plots and movies. Most folder contain a script name `plot_*.m` that makes the figure(s) and `video_*.m` that generates the video(s). You will need `ffmpeg` installed to convert the serial images into a video.<br> </p>
An analyst-created, high-resolution seismic arrival time dataset for evaluating machine learning phase detectors
<p>This data set contains phase arrival times and phase labels for one hour of continuous seismic data recorded at the three-component broadband station WY.YNR on 2014 30 March from 13:00:00 to 14:00:00 UTC. This hour of data follows a M<sub>w</sub> 4.8 occurring at 12:34 UTC in the Yellowstone region and contains many small events close together in space and time. All 687 picks (404 P and 283 S) were made by a seismic analyst trained at the University of Utah Seismograph Stations. The goal was to pick as many reasonable arrivals as possible for evaluating the performance of machine-learning-based phase detectors on continuous data during high seismicity rates. Phase picks were made using the Seismic Analysis Code (SAC; Goldstein and Snoke, 2005) and Pyrocko (Heimann <em>et al</em>., 2017). In general, a 1 to 17 Hz bandpass filter was used.</p> <p>The csv file contains the pick id, phase arrival time in UTC and Unix epoch formats, and the phase labels (P or S).</p>
Hybrid quantum-classical machine learning for generative chemistry and drug design: Generated molecules
<p>Deep generative chemistry models emerge as powerful tools to expedite drug discovery. How- ever, the immense size and complexity of the structural space of all possible drug-like molecules pose significant obstacles, which could be overcome with hybrid architectures combining quantum computers with deep classical networks. As the first step toward this goal, we built a compact discrete variational autoencoder (DVAE) with a Restricted Boltzmann Machine (RBM) of reduced size in its latent layer. The size of the proposed model was small enough to fit on a state-of-the-art D-Wave quantum annealer and allowed training on a subset of the ChEMBL dataset of biologically active compounds. Finally, we generated 2331 novel chemical structures with medicinal chemistry and synthetic accessibility properties in the ranges typical for molecules from ChEMBL. The pre- sented results demonstrate the feasibility of using already existing or soon-to-be-available quantum computing devices as testbeds for future drug discovery applications.</p>
Datasets and codes for the peer review article "Human and natural impacts on the U.S. freshwater salinization and alkalinization: A machine learning approach"
<p>Ongoing salinization and alkalinization in U.S. rivers have been attributed to inputs of road salt and effects of human-accelerated weathering in previous studies. Salinization poses a severe threat to human and ecosystem health, while human derived alkalinization implies increasing uncertainty in the dynamics of terrestrial sequestration of atmospheric carbon dioxide. A mechanistic understanding of whether and how human activities accelerate weathering and contribute to the geochemical changes in U.S. rivers is lacking. To address this uncertainty, we compiled dissolved sodium (salinity proxy) and alkalinity values along with 32 watershed properties ranging from hydrology, climate, geomorphology, geology, soil chemistry, land use, and land cover for 226 river monitoring sites across the coterminous U.S. Using these data, we built two machine-learning models to predict monthly-aggregated sodium and alkalinity fluxes at these sites. The sodium-prediction model detected human activities (represented by population density and impervious surface area) as major contributors to the salinity of U.S. rivers. In contrast, the alkalinity-prediction model identified natural processes as predominantly contributing to variation in riverine alkalinity flux, including runoff, carbonate sediment or siliciclastic sediment, soil pH and soil moisture. Unlike prior studies, our analysis suggests that the alkalinization in U.S. rivers is largely governed by local climatic and hydrogeological conditions.</p>
Data used in Machine learning reveals the waggle drift's role in the honey bee dance communication system
<p><strong>Data and metadata used in "Machine learning reveals the waggle drift’s role in the honey bee dance communication system" </strong></p> <p>All timestamps are given in ISO 8601 format.</p> <p><strong>The following files are included:</strong></p> <p><strong>Berlin2019_waggle_phases.csv, Berlin2021_waggle_phases.csv</strong></p> <p>Automatic individual detections of waggle phases during our recording periods in 2019 and 2021.</p> <ul> <li> <p>timestamp: Date and time of the detection.</p> </li> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> <li> <p>x_median, y_median: Median position of the bee during the waggle phase (for 2019 given in millimeters after applying a homography, for 2021 in the original image coordinates).</p> </li> <li> <p>waggle_angle: Body orientation of the bee during the waggle phase in radians (0: oriented to the right, PI / 4: oriented upwards).</p> </li> </ul> <p><strong>Berlin2019_dances.csv</strong></p> <p>Automatic detections of dance behavior during our recording period in 2019.</p> <ul> <li> <p>dancer_id: Unique ID of the individual bee.</p> </li> <li> <p>dance_id: Unique ID of the dance.</p> </li> <li> <p>ts_from, ts_to: Date and time of the beginning and end of the dance.</p> </li> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> <li> <p>median_x, median_y: Median position of the individual during the dance.</p> </li> <li> <p>feeder_cam_id: ID of the feeder that the bee was detected at prior to the dance.</p> </li> </ul> <p><strong>Berlin2019_followers.csv</strong></p> <p>Automatic detections of attendance and following behavior, corresponding to the dances in Berlin2019_dances.csv.</p> <ul> <li> <p>dance_id: Unique ID of the dance being attended or followed.</p> </li> <li> <p>follower_id: Unique ID of the individual attending or following the dance.</p> </li> <li> <p>ts_from, ts_to: Date and time of the beginning and end of the interaction.</p> </li> <li> <p>label: “attendance” or “follower”</p> </li> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> </ul> <p><strong>Berlin2019_dances_with_manually_verified_times.csv</strong></p> <p>A sample of dances from Berlin2019_dances.csv where the exact timestamps have been manually verified to correspond to the beginning of the first and last waggle phase down to a precision of ca. 166 ms (video material was recorded at 6 FPS).</p> <ul> <li> <p>dance_id: Unique ID of the dance.</p> </li> <li> <p>dancer_id: Unique ID of the dancing individual.</p> </li> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> <li> <p>feeder_cam_id: ID of the feeder that the bee was detected at prior to the dance.</p> </li> <li> <p>dance_start, dance_end: Manually verified date and times of the beginning and end of the dance.</p> </li> </ul> <p><strong>Berlin2019_dance_classifier_labels.csv</strong></p> <p>Manually annotated waggle phases or following behavior for our recording season in 2019 that was used to train the dancing and following classifier. Can be merged with the supplied individual detections.</p> <ul> <li> <p>timestamp: Timestamp of the individual frame the behavior was observed in.</p> </li> <li> <p>frame_id: Unique ID of the video frame the behavior was observed in.</p> </li> <li> <p>bee_id: Unique ID of the individual bee.</p> </li> <li> <p>label: One of “nothing”, “waggle”, “follower”</p> </li> </ul> <p><strong>Berlin2019_dance_classifier_unlabeled.csv</strong></p> <p>Additional unlabeled samples of timestamp and individual ID with the same format as Berlin2019_dance_classifier_labels.csv, but without a label. The data points have been sampled close to detections of our waggle phase classifier, so behaviors related to the waggle dance are likely overrepresented in that sample.</p> <p><strong>Berlin2021_waggle_phase_classifier_labels.csv</strong></p> <p>Manually annotated detections of our waggle phase detector (bb_wdd2) that were used to train the neural network filter (bb_wdd_filter) for the 2021 data.</p> <ul> <li> <p>detection_id: Unique ID of the waggle phase.</p> </li> <li> <p>label: One of “waggle”, “activating”, “ventilating”, “trembling”, “other”. Where “waggle” denoted a waggle phase, “activating” is the shaking signal, “ventilating” is a bee fanning her wings. “trembling” denotes a tremble dance, but the distinction from the “other” class was often not clear, so “trembling” was merged into “other” for training.</p> </li> <li> <p>orientation: The body orientation of the bee that triggered the detection in radians (0: facing to the right, PI /4: facing up).</p> </li> <li> <p>metadata_path: Path to the individual detection in the same directory structure as created by the waggle dance detector.</p> </li> </ul> <p><strong>Berlin2021_waggle_phase_classifier_ground_truth.zip</strong></p> <p>The output of the waggle dance detector (bb_wdd2) that corresponds to Berlin2021_waggle_phase_classifier_labels.csv and is used for training. The archive includes a directory structure as output by the bb_wdd2 and each directory includes the original image sequence that triggered the detection in an archive and the corresponding metadata. The training code supplied in bb_wdd_filter directly works with this directory structure.</p> <p><strong>Berlin2019_tracks.zip</strong></p> <p>Detections and tracks from the recording season in 2019 as produced by our tracking system. As the full data is several terabytes in size, we include the subset of our data here that is relevant for our publication which comprises over 46 million detections. We included tracks for all detected behaviors (dancing, following, attending) including one minute before and after the behavior. We also included all tracks that correspond to the labeled and unlabeled data that was used to train the dance classifier including 30 seconds before and after the data used for training.<br> We grouped the exported data by date to make the handling easier, but to efficiently work with the data, we recommend importing it into an indexable database.</p> <p>The individual files contain the following columns:</p> <ul> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> <li> <p>timestamp: Date and time of the detection.</p> </li> <li> <p>frame_id: Unique ID of the video frame of the recording from which the detection was extracted.</p> </li> <li> <p>track_id: Unique ID of an individual track (short motion path from one individual). For longer tracks, the detections can be linked based on the bee_id.</p> </li> <li> <p>bee_id: Unique ID of the individual bee.</p> </li> <li> <p>bee_id_confidence: Confidence between 0 and 1 that the bee_id is correct as output by our tracking system.</p> </li> <li> <p>x_pos_hive, y_pos_hive: Spatial position of the bee in the hive on the side indicated by cam_id. Given in millimeters after applying a homography on the video material.</p> </li> <li> <p>orientation_hive: Orientation of the bees’ thorax in the hive in radians (0: oriented to the right, PI / 4: oriented upwards).</p> </li> </ul> <p><strong>Berlin2019_feeder_experiment_log.csv</strong></p> <p>Experiment log for our feeder experiments in 2019.</p> <ul> <li> <p>date: Date given in the format year-month-day.</p> </li> <li> <p>feeder_cam_id: Numeric ID of the feeder.</p> </li> <li> <p>coordinates: Longitude and latitude of the feeder. For feeders 1 and 2 this is only given once and held constant. Feeder 3 had varying locations.</p> </li> <li> <p>time_opened, time_closed: Date and time when the feeder was set up or closed again.<br> sucrose_solution: Concentration of the sucrose solution given as sugar:water (in terms of weight). On days where feeder 3 was open, the other two feeders offered water without sugar.</p> </li> </ul> <p> </p> <ul> </ul> <p><strong>Software used to acquire and analyze the data:</strong></p> <ul> <li> <p><a href="https://github.com/BioroboticsLab/bb_pipeline">bb_pipeline: Tag localization and decoding pipeline</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_pipeline_models">bb_pipeline_models: Pretrained localizer and decoder models for bb_pipeline</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_binary">bb_binary: Raw detection data storage format</a></p> </li> <li> <p><a href="https://doi.org/10.5281/zenodo.4436419">bb_irflash: IR flash system schematics and arduino code</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_imgacquisition">bb_imgacquisition: Recording and network storage </a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_behavior">bb_behavior: Database interaction and data (pre)processing, feature extraction</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_tracking">bb_tracking: Tracking of bee detections over time</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_wdd2">bb_wdd2: Automatic detection and decoding of honey bee waggle dances</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_wdd_filter/">bb_wdd_filter: Machine learning model to improve the accuracy of the waggle dance detector</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_dance_networks/tree/master/bb_dance_networks">bb_dance_networks: Detection of dancing and following behavior from trajectories</a></p> </li> </ul> <p> </p>
Datasets for "Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Classification Algorithms in Reiner Gamma and Mare Ingenii"
<p>Final surface reflectance data at 2.6 m/pixel resolution with floating point values are available as GeoTiff and ASCII text files. Definition files for the K-Means and MLC algorithms in classifying swirl units are also available as ASCII text files. See README file for further details.</p> <p>Data used in the research article:</p> <p>Chuang, F.C., M.D. Richardson, J.R. Weirich, A.A. Sickafoose, and D.L. Domingue, 2022. Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Image Classification Algorithms in Reiner Gamma and Mare Ingenii. The Planetary Science Journal, 3:231. doi://10.3847/PSJ/ac8f43</p> <p> </p>
PFAM Protein Families Dataset for Machine Learning
<p>A cleaned dataset of protein sequences and protein families for classification. The dataset is exported from PFAM as of June 2023 and curated to achieve the following characteristics:</p> <ul> <li>only protein families included with >=100 sequences</li> <li>families with >2000 sequences are truncated and only represented by 2000 sequences (chosen randomly)</li> <li>only proteins with sequence lengths between 100 and 1000</li> <li>amino acid sequences are form PDB; chains are concatenated only if not similar</li> </ul> <p>The dataset is not balanced, numbers of sequences per family in PFAM and in in dataset are:</p> <pre><code>families: 62, sequences: 46872 total (in PFAM) -> included (in dataset) Number in family ALLERGEN: 122 -> 122 Number in family APOPTOSIS: 381 -> 381 Number in family BIOSYNTHETIC PROTEIN: 346 -> 346 Number in family BIOTIN BINDING PROTEIN: 165 -> 165 Number in family BLOOD CLOTTING: 138 -> 138 Number in family CALCIUM BINDING PROTEIN: 135 -> 135 Number in family CELL ADHESION: 1116 -> 1116 Number in family CELL CYCLE: 511 -> 511 Number in family CHAPERONE: 964 -> 964 Number in family CONTRACTILE PROTEIN: 158 -> 158 Number in family CYTOKINE: 191 -> 191 Number in family DE NOVO PROTEIN: 253 -> 253 Number in family DNA BINDING PROTEIN: 1008 -> 1008 Number in family ELECTRON TRANSPORT: 841 -> 841 Number in family FLUORESCENT PROTEIN: 348 -> 348 Number in family GENE REGULATION: 607 -> 607 Number in family HORMONE: 272 -> 272 Number in family HORMONE GROWTH FACTOR: 159 -> 159 Number in family HORMONE RECEPTOR: 121 -> 121 Number in family HYDROLASE: 19551 -> 2000 Number in family HYDROLASE ANTIBIOTIC: 120 -> 120 Number in family HYDROLASE HYDROLASE INHIBITOR: 2890 -> 2000 Number in family HYDROLASE INHIBITOR: 315 -> 315 Number in family IMMUNE SYSTEM: 3333 -> 2000 Number in family IMMUNOGLOBULIN: 155 -> 155 Number in family ISOMERASE: 2457 -> 2000 Number in family ISOMERASE ISOMERASE INHIBITOR: 139 -> 139 Number in family LECTIN: 139 -> 139 Number in family LIGASE: 1780 -> 1780 Number in family LIGASE LIGASE INHIBITOR: 163 -> 163 Number in family LIPID BINDING PROTEIN: 421 -> 421 Number in family LIPID TRANSPORT: 115 -> 115 Number in family LUMINESCENT PROTEIN: 221 -> 221 Number in family LYASE: 4150 -> 2000 Number in family LYASE LYASE INHIBITOR: 298 -> 298 Number in family MEMBRANE PROTEIN: 1338 -> 1338 Number in family METAL BINDING PROTEIN: 951 -> 951 Number in family METAL TRANSPORT: 409 -> 409 Number in family MOTOR PROTEIN: 195 -> 195 Number in family OXIDOREDUCTASE: 11531 -> 2000 Number in family OXIDOREDUCTASE OXIDOREDUCTASE INHIBITOR: 766 -> 766 Number in family OXYGEN STORAGE: 127 -> 127 Number in family OXYGEN STORAGE TRANSPORT: 260 -> 260 Number in family OXYGEN TRANSPORT: 414 -> 414 Number in family PHOTOSYNTHESIS: 173 -> 173 Number in family PLANT PROTEIN: 255 -> 255 Number in family PROTEIN BINDING: 1613 -> 1613 Number in family PROTEIN TRANSPORT: 693 -> 693 Number in family RECEPTOR: 108 -> 108 Number in family REPLICATION: 161 -> 161 Number in family RNA BINDING PROTEIN: 546 -> 546 Number in family SIGNALING PROTEIN: 2312 -> 2000 Number in family STRUCTURAL PROTEIN: 869 -> 869 Number in family SUGAR BINDING PROTEIN: 1250 -> 1250 Number in family TOXIN: 546 -> 546 Number in family TRANSCRIPTION REGULATION: 3283 -> 2000 Number in family TRANSFERASE: 14724 -> 2000 Number in family TRANSFERASE INHIBITOR: 126 -> 126 Number in family TRANSFERASE TRANSFERASE INHIBITOR: 2465 -> 2000 Number in family TRANSLATION: 370 -> 370 Number in family TRANSPORT PROTEIN: 2782 -> 2000 Number in family VIRAL PROTEIN: 2150 -> 2000</code></pre> <p>Files:</p> <ul> <li>families.csv: list of protein families with frequencies</li> <li>pfam_46872x62.csv: full dataset with amino acid sequences as string (one-letter code)</li> <li>pfam-trn-xy.csv: training dataset with amino acid sequences as tokens (1..25) and padded to a common length of 1000 with padding token 0:</li> </ul> <pre><code> Amino acid | Token | Description -------------------------------- C | 1 | Cysteine S | 2 | Serine T | 3 | Threonine A | 4 | Alanine G | 5 | Glycine P | 6 | Proline D | 7 | Aspartic acid E | 8 | Glutamic acid Q | 9 | Glutamine N | 10 | Asparagine H | 11 | Histidine R | 12 | Arginine K | 13 | Lysine M | 14 | Methionine I | 15 | Isoleucine L | 16 | Leucine V | 17 | Valine W | 18 | Tryptophan Y | 19 | Tyrosine F | 20 | Phenylalanine B | 21 | Aspartic acid or Asparagine Z | 22 | Glutamic acid or Glutamine J | 23 | Leucine or Isoleucine U | 24 | Selenocysteine X | 25 | Unknown amino acid . | 0 | padding token</code></pre> <p> </p> <ul> <li>pfam-trn-labels.csv: plain-text labels for training data</li> <li>pfam-tst-xy.csv</li> <li>pfam-tst-labels.csv: test data</li> <li>pfam-balanced-trn-xy.csv</li> <li>pfam-balanced-trn-labels.csv:</li> <li>pfam-balanced-tst-xy.csv</li> <li>pfam-balanced-tst-labels.csv: balanced datasets, created by oversampling.</li> </ul>
A Pan-European, Quantile Machine learning (QML) based, Total, Fine-Mode and Coarse-Mode Aerosol Optical Depth dataset (QML AOD))
<p>The V 1.1.0 product is an improved Aerosol Optical Depth (AOD) product based on Gap-filled MAIAC AOD, which provide first full-coverage, high-resolution monitoring of fine-mode and coarse-mode aerosols in Europe from 2003-20. This dataset has successfully rectified the previously identified issue of weak associations between satellite AOD and PM2.5 in Europe, which was primarily attributable to current limitations of AOD data. Our innovative approach has yielded stronger correlations with PM10, PM2.5, and PMcoarse than previous AOD product, laying a critical groundwork for improving PM10, PM2.5, and PMcoarse predictions in further epidemiological studies or environmental monitoring.</p> <p>We have uploaded three QML AOD datasets in Geotiff format, covering the region from -27° to 72° latitude and from -25° to 45° longitude. These datasets will be useful for researchers and policymakers to better understand the impacts of aerosols on the environment and human health.</p> <p> Note: v1.0.0 product do not include MAIAC AOD in their models.</p> <p>Please read more details in our paper </p> <h1><span>Estimation of pan-European, daily total, fine-mode and coarse-mode Aerosol Optical Depth at 0.1° resolution to facilitate air quality assessments</span></h1> <p><a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.scitotenv.2024.170593" target="_blank" rel="noreferrer noopener"><span>https://doi.org/10.1016/j.scitotenv.2024.170593</span></a></p>
Dataset for Machine Learning Framework for Modeling Exciton-Polaritons in Molecular Materials
<p>The data consists of several NumPy arrays saved in the binary format (npy files) with a total size of 108 MB. The details of these files are listed below.</p> <table> <tbody> <tr> <td> <p><strong>Filename and path</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>training/azo_R.npy</p> </td> <td> <p>Coordinates for training</p> </td> </tr> <tr> <td> <p>training/azo_Z.npy</p> </td> <td> <p>Atomic indices for training</p> </td> </tr> <tr> <td> <p>training/azo_E.npy</p> </td> <td> <p>Molecular energies for training</p> </td> </tr> <tr> <td> <p>training/azo_D.npy</p> </td> <td> <p>Transition dipoles for training</p> </td> </tr> <tr> <td> <p>training/azo_ScaledNACR.npy</p> </td> <td> <p>Non-adiabatic coupling vectors scaled by energy difference for training</p> </td> </tr> <tr> <td> <p>scan/azo_R.npy</p> </td> <td> <p>Coordinates for PES scan</p> </td> </tr> <tr> <td> <p>scan/azo_Z.npy</p> </td> <td> <p>Atomic indices for PES scan</p> </td> </tr> <tr> <td> <p>scan/azo_E.npy</p> </td> <td> <p>Molecular energies for PES scan</p> </td> </tr> <tr> <td> <p>scan/azo_D.npy</p> </td> <td> <p>Transition dipoles for PES scan</p> </td> </tr> <tr> <td> <p>scan/azo_ScaledNACR.npy</p> </td> <td> <p>Non-adiabatic coupling vectors scaled by energy difference for PES scan</p> </td> </tr> <tr> <td> <p>spectrum/azo_R.npy</p> </td> <td> <p>Coordinates for spectrum calculations</p> </td> </tr> <tr> <td> <p>spectrum/azo_Z.npy</p> </td> <td> <p>Atomic indices for spectrum calculations</p> </td> </tr> <tr> <td> <p>spectrum/azo_E.npy</p> </td> <td> <p>Molecular energies for spectrum calculations</p> </td> </tr> <tr> <td> <p>spectrum/azo_D.npy</p> </td> <td> <p>Transition dipoles for spectrum calculation</p> </td> </tr> </tbody> </table>
Course Materials for Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730)
In today's world, understanding environmental data and making informed decisions based on it is crucial for addressing complex environmental challenges. Yale School of the Environment's Environmental Data Science in R: Introduction to Data Integration and Machine Learning (ENV 730) course serves as an introduction to the integration of environmental data using R programming language, coupled with machine learning techniques. This dataset contains a zip file with all the data files used in this course, along with a README that has the metadata for those files.
WPE01 Assessing the value added of NEON for using machine learning to quantify vegetation mosaics and woody plant encroachment at Konza Prairie
Woody encroachment, or invasion of woody plants, is rapidly shifting tallgrass prairie into shrub and evergreen dominated ecosystems, mainly due to exclusion of fire. Tracking the pace and extent of woody encroachment is difficult because shrubs and small trees are much smaller than the coarse resolution (>10m2) of common remote sensed images. However, the US government has been investing in finer resolution (<2m2) remote sensing through USDA NAIP and the National Ecological Observatory Network (NEON), both of which cost multi-million dollars each year and contain different remote sensed products. We compared two methods of classification (random forests and support vector machines) with these two freely available remotely sensed aerial images to determine if and how much NEON adds to classification accuracy and determine which method of machine learning was more accurate. All models have very high overall classification accuracy (>91%), with the NEON image a few percent more accurate than NAIP. The NEON image significantly relies on canopy height (LiDAR) to make classifications, but the importance of bands is more evenly distributed during NAIP classification. Lastly, accuracy for Eastern Red Cedar specifically is high with NEON (78-84%), compared to the relatively low classification accuracy using NAIP imagery (55-61%).
Evaluation and Calibration of a Low-cost Particle Sensor in Ambient Conditions Using Machine Learning Methods
<p>Particle sensing technology has shown great potential for monitoring particulate matter (PM) with very few temporal and spatial restrictions because of its low-cost, compact size, and easy operation. However, the performance of low-cost sensors for PM monitoring in ambient conditions has not been thoroughly evaluated. Monitoring results by low-cost sensors are often questionable. In this study, a low-cost fine particle monitor (Plantower PMS 5003) was co-located with a reference instrument, named Synchronized Hybrid Ambient Real-time Particulate (SHARP) monitor, in Calgary Varsity air monitoring station from December 2018 to April 2019. The study evaluated the performance of this low-cost PM sensor in ambient conditions and calibrated its readings using simple linear regression (SLR), multiple linear regression (MLR), and two more powerful machine learning algorithms using random search techniques for the best model architectures. The two machine learning algorithms are XGBoost and feedforward neural network (NN). Field evaluation showed that the Pearson r between the low-cost sensor and the SHARP instrument was 0.78. Fligner and Killeen (F-K) test indicated a statistically significant difference between the variances of the PM<sub>2.5 </sub>values by the low-cost sensor and by the SHARP instrument. Large overestimations by the low-cost sensor before calibration were observed in the field and were believed to be caused by the variation of ambient relative humidity. The root mean square error (RMSE) was 9.93 when comparing the low-cost sensor with the SHARP instrument. The calibration by the feedforward NN had the smallest RMSE of 3.91 in the test dataset, compared to the calibrations by SLR (4.91), MLR (4.65), and XGBoost (4.19). After calibrations, the F-K test using the test dataset showed that the variances of the PM<sub>2.5</sub> values by the NN and the XGBoost and by the reference method were not statistically significantly different. From this study, we conclude that feedforward NN is a promising method to address the poor performance of the low-cost sensors for PM<sub>2.5</sub> monitoring. In addition, the random search method for hyperparameters was demonstrated to be an efficient approach for selecting the best model structure.</p>
RNAPosers: Machine Learning Classifiers For RNA-Ligand Poses [Data Set]
<ul> <li>This dataset contains the decoys poses used to train and test RNAPosers, a set of RNA-ligand pose classifiers.</li> <li>The folder of each RNA-ligand complex (identified using its PDB ID) contains: <ul> <li>Ligand SMILES: lig.smi</li> <li>Ligand coordinate: lig.sd</li> <li>Receptor coordinate: receptor.mol2</li> <li>Pose coordinates: poses.sd</li> <li>Pose similarity data: rmsd.txt</li> </ul> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.