Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,185
datasets available to search
ShareScore release 0.9.0
Dataset results
7,185 results for “learning”
Improving machine-learning models in materials science through large datasets
<p>1. Image of the <a href="https://alexandria.icams.rub.de/"><strong>Alexandria database </strong></a> state corresponding to the paper "<strong>Improving machine-learning models in materials science through large datasets</strong>".</p> <ul> <li>Static pbe calculations for 1D, 2D, 3D compounds can be found in 1D_pbe.tar.gz, 2D_pbe.tar.gz, 3D_pbe.tar.gz in batches of 100k materials. The latter also contains a separate convex hull pickle with all compounds on the pbe convex hull (convex_hull_pbe_2023.12.29.json.bz2) and a list of prototypes in the database (prototypes.json.bz2). The systematic 3D calculations performed for the article <strong>Improving machine-learning models in materials science through large datasets </strong>(in the paper referred to as round 2 and 3) can be found by the location keyword in the data dictionary of each ComputedStructureEntry containing "<strong>cgat_comp/quaternaries</strong>" (round 2) and "<strong>cgat_comp2/</strong>" (round 3). Round 1 (10.1002/adma.202210788) can be found under "cgat_comp/ternaries", ""cgat_comp/binaries".</li> <li>Static pbesol calculations for 3D compounds can be found in 3D_ps.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the pbesol convex hull (convex_hull_ps_2023.12.29.json.bz2). </li> <li>Static scan calculations for 3D compounds can be found in 3D_scan.tar (still zip compressed) in batches of 100k materials. The folder also contains a separate convex hull pickle with all compounds on the scan convex hull (convex_hull_scan_2023.12.29.json.bz2). </li> <li>Geometry relaxation curves for 1D and 2D and 3D compounds calculated with PBE can be found in geo_opt_1D.tar.gz, geo_opt_2D.tar.gz. and geo_opt_3D.tar. Each file in each folder contains a batch of up to 10k relaxation trajectories.</li> <li>PBESOL relaxation trajectories for 3D compounds can be found in geo_opt_ps.tar</li> </ul> <p>2. Crystal graph attention networks to predict the volume (<a href="https://zenodo.org/api/records/12582650/draft/files/volume_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">volume_round_3.tar.gz</a>) and distance to the convex hull (<a href="https://zenodo.org/api/records/12582650/draft/files/e_above_hull_round_3.tar.gz/content" target="_blank" rel="noopener noreferrer">e_above_hull_round_3.tar.gz</a>) trained for the paper "Improving machine-learning models in materials science through large datasets".</p> <p>Can be used with the code at https://github.com/hyllios/CGAT/tree/main/CGAT.<br><strong>Note will predict the distance to the convex hull not normalized per atom when using the code on the github.<br></strong></p> <p>3. Alignn models as well as m3gnet and mace models corresponding to the publication can be found in <a href="https://zenodo.org/api/records/12582650/draft/files/alexandria_v2.tar.gz/content" target="_blank" rel="noopener noreferrer">alexandria_v2.tar.gz</a></p> <p>4. scripts.tar.gz Some scripts used for generating CGAT input data/ performing parallel predictions and for relaxations with m3gnet/mace force fields</p>
Toward a Generalizable Machine-Learned Potential for Metal-Organic Frameworks
<ul> <li>This repository contains the dataset used in the publication<br> `Toward Generalizable Machine Learned Potential for Metal-Organic Frameworks` Yue Yifei, Saad Aldin Mohammed, Loh Duane*, Jiang Jianwen*<br> <br> Please each the README.md within each subfolder. For brevity, the data is organized into three sections<br> <br> 1. The dataset in DATASET<br> - The training and testing dataset, including structures of MOFs in extxyz format<br> <br> 2. The training output files and logs in NEQUIP-TRAIN<br> - The conda environment details, training scripts and logs<br> - Also Training and testing metrics in csv files<br> - This is split into two zip files NEQUIP-TRAIN1 and NEQUIP-TRAIN2 due to their size<br> <br> 3. Examples of using the developed models in MD simulations<br> - Including LAMMPS scripts, data file and environment details used in our scalability tests<br> - The complied Nequip-patched LAMMPS version is also provided<br> - Details on how to use our models - we used a default model that is slower but more accurate in our study but faster models are also developed.</li> </ul>
Dataset from: "Reward expectation facilitates context learning and attentional guidance in visual search"
<p>Dataset for Bergmann N, Koch D, Schubö A (2019). Reward expectation facilitates context learning and attentional guidance in visual search, <em>Journal of Vision</em>, 19(3). <a href="https://doi.org/10.1167/19.3.10">https://doi.org/10.1167/19.3.10</a></p>
Dataset for: Importance of satellite observations for high-resolution mapping of near-surface NO2 by machine learning
<p>Dataset for: Importance of satellite observations for high-resolution mapping of near-surface NO<sub>2 </sub>by machine learning</p> <p>This dataset is uploaded as a part of the article by Kim et al. (2021). The dataset is the hourly maps of near-surface nitrogen dioxide (NO<sub>2</sub>) concentrations at 100 m resolution for an Alpine domain (Switzerland and northern Italy, 6-12 °E, 42-48 °N). The dataset is provided per day (24 hours) in a netcdf (*.nc ~550MB). In this work, we have generated NO<sub>2 </sub>hourly maps for Feb. 2019 to May 2020 and, here, we upload for March 2019 only (~16 GB). If you need data for another period of time, please contact Gerrit Kuhlmann (gerrit.kuhlmann@empa.ch) or Minsu Kim (minsu.kim@empa.ch). </p>
Machine learning designs non-hemolytic antimicrobial peptides
<p>The upload contains additional primary data associated with the publication, including raw data in the original file format whenever possible.</p> <p>Data content: HRMS, HPLC-MS, CD, MD, TEM</p>
Data Report: "Health care of Persons Deprived of Liberty" Course from Brazil's Unified Health System Virtual Learning Environment
<p><strong>Dataset name: </strong>asppl-dataset.csv</p> <p><strong>Version: </strong>1.0</p> <p><strong>Dataset period: </strong>06/07/2018- 05/25/2021</p> <p><strong>Dataset Characteristics: </strong>Multivalued</p> <p><strong>Number of Instances: </strong>4861</p> <p><strong>Number of Attributes: </strong>33</p> <p><strong>Missing Values: </strong>Yes</p> <p><strong>Area(s): </strong>Health and education </p> <p><strong>Sources: </strong></p> <ul> <li> <p><strong>Primary</strong>: Unified Health System Virtual Learning Environment (AVASUS, in Portuguese: Ambiente Virtual de Aprendizagem do Sistema Único de Saúde) [1];</p> </li> <li> <p><strong>Secondary: </strong></p> <ol> <li> <p>Brazilian Classification of Occupations (CBO, in Portuguese: Classificação Brasileira de Ocupação) [2];</p> </li> <li> <p>National Registry of Health Establishments (CNES, in Portuguese: Cadastro Nacional de Estabelecimentos de Saúde) [3]; and </p> </li> <li> <p>Brazilian Institute of Geography and Statistics (IBGE, in Portuguese: Instituto Brasileiro de Geografia e Estatística) [4].</p> </li> </ol> </li> </ul> <p><strong>Description: </strong>The data contained on the asppl-dataset.csv dataset (see Table 1) originates from participants of the technology-based educational course “Health care of Persons Deprived of Liberty”. The course is available on the Unified Health System Virtual Learning Environment [1]. This dataset provides elementary data for analyzing the course’s impact and reach, as well as the profile of its participants.</p> <p> </p>
ManyTypes4Py: A Benchmark Python Dataset for Machine Learning-Based Type Inference
<ul> <li>The dataset is gathered on Sep. 17th 2020 from GitHub.</li> <li>It has <em>clean</em> and <em>complete</em> versions (from v0.7): <ul> <li>The clean version has 5.1K <strong>type-checked </strong>Python repositories and 1.2M type annotations.</li> <li>The complete version has 5.2K Python repositories and 3.3M type annotations.</li> </ul> </li> <li>The dataset's source files are type-checked using <a href="https://mypy.readthedocs.io/">mypy</a> (clean version).</li> <li>The dataset is also de-duplicated using the <a href="https://github.com/saltudelft/CD4Py">CD4Py</a> tool.</li> <li>Check out the <strong>README.MD</strong> file for the description of the dataset.</li> <li>Notable changes to each version of the dataset are documented in <strong>CHANGELOG.md</strong>.</li> <li>The dataset's scripts and utilities are available on <a href="https://github.com/saltudelft/many-types-4-py-dataset">its GitHub repository</a>.</li> </ul>
Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution
<p>Efficiency is essential to support ever-growing datasets, especially for Deep Learning (DL) systems. DL frameworks have traditionally embraced deferred execution-style DL code---supporting symbolic, graph-based Deep Neural Network (DNN) computation. While scalable, such development is error-prone, non-intuitive, and difficult to debug. Consequently, more natural, imperative DL frameworks encouraging eager execution have emerged but at the expense of run-time performance. Though hybrid approaches aim for the "best of both worlds," using them effectively requires subtle considerations. Our key insight is that, while DL programs typically execute sequentially, hybridizing imperative DL code resembles parallelizing sequential code in traditional systems. Inspired by this, we present an automated refactoring approach that assists developers in determining which otherwise eagerly-executed imperative DL functions could be effectively and efficiently executed as graphs. The approach features novel static imperative tensor and side-effect analyses for Python. Due to its inherent dynamism, analyzing Python may be unsound; however, the conservative approach leverages a speculative (keyword-based) analysis for resolving difficult cases that informs developers of any assumptions made. The approach is: (i) implemented as a plug-in to the PyDev Eclipse IDE that integrates the WALA Ariadne analysis framework and (ii) evaluated on nineteen DL projects consisting of 132 KLOC. The results show that 326 of 766 candidate functions (42.56%) were refactorable, and an average relative speedup of 2.16x on performance tests was observed with negligible differences in model accuracy. The results indicate that the approach is useful in optimizing imperative DL code to its full potential.</p>
Data and code related to the paper: "Programmable Droplet Microfluidics Based on Machine Learning and Acoustic Manipulation"
<p>This archive contains the raw data and Matlab scripts to reproduce the plots and supplementary movies for the paper:</p> <p>Kyriacos Yiannacou, Vipul Sharma and Veikko Sariola, "Programmable Droplet Microfluidics Based on Machine Learning and Acoustic Manipulation", <em>Langmuir</em> 2022, 38, 38, 11557–11564.</p> <p><a href="https://doi.org/10.1021/acs.langmuir.2c01061">Link to the paper</a>.</p> <p>The scripts were tested on Matlab R2021a on Windows.</p> <p>The acoustofluidic controller software is the same as in our previous paper and is archived <a href="https://doi.org/10.5281/zenodo.4593021">here</a>.</p> <p>Generally speaking, there is a folder containing the plotting scripts for each figure(s) and/or movie(s). Within each folder, the raw data files are under the folder `data/`. Once ran, the scripts produce another folder called `output/`, to which they place the created plots and movies. Most folder contain a script name `plot_*.m` that makes the figure(s) and `video_*.m` that generates the video(s). You will need `ffmpeg` installed to convert the serial images into a video.<br> </p>
Data: Learning Lattice Quantum Field Theories with Equivariant Continuous Flows
<p>Network parameters of continuous normalizing flows trained for the <span class="math-tex">\(\varphi^4\)</span> theory.</p> <p>Corresponding article: Learning Lattice Quantum Field Theories with Equivariant Continuous Flows [<a href="https://arxiv.org/abs/2207.00283">2207.00283</a>]</p> <p>Abstract: We propose a novel machine learning method for sampling from the high-dimensional probability distributions of Lattice Field Theories, which is based on a single neural ODE layer and incorporates the full symmetries of the problem. We test our model on the <span class="math-tex">\(\varphi^4\)</span> theory, showing that it systematically outperforms previously proposed flow-based methods in sampling efficiency, and the improvement is especially pronounced for larger lattices. Furthermore, we demonstrate that our model can learn a continuous family of theories at once, and the results of learning can be transferred to larger lattices. Such generalizations further accentuate the advantages of machine learning methods.</p>
Phase unwrapping using Deep Learning in Holographic Tomography - dataset.
<p>This dataset contains two types of data: phase images and trained model files.</p> <ol> <li> <p><strong>Real phase images</strong> - these phase images are contained with the files named with the prefix "real_". The type of the data files is "<strong>.npz</strong>", to be loaded with NumPy (<em>np.load()</em>), as a dictionary. The data is stored within the key ["<em>arr_0</em>"]. The images depict cells [1], organoids [2], phantoms [3-4] and regular 3D printed structures with high scattering properties [5]. The images have been augmented in order to expand the volume of the training dataset. All images are of shape (<em>256,256,1</em>). It is a big dataset containing 27,189 images of each type for training the unwrapping model:</p> <ul> <li> <p>unwrapped - continuous phase distribution (<em>float32</em>)</p> </li> <li> <p>wrapped - phase wrapped into mod2<span class="math-tex">\(\pi\)</span> (<em>float32</em>)</p> </li> <li> <p>wrapcount - wrap count phase maps coded in the integer form (0,1,2...) (<em>uint8</em>)</p> </li> </ul> </li> <li> <p><strong>Synthetic phase images</strong> - phase images in these files were generated algorithmically in the MATLAB programming language. The files containing this dataset have a prefix "synthetic_". The type of the data files is "<strong>.npz</strong>", to be loaded with NumPy (<em>np.load()</em>), as a dictionary. The data is stored within the key ["<em>arr_0</em>"]. Phase images contained in the synthetic dataset can be split into 3 types by their type: spherical distribution, simulated cells w/ spherical background and simulated cells w/ introduced linear tilt. All images are of shape (<em>256,256,1</em>). This dataset contains 10,000 images of each type for training the unwrapping and denoising models:</p> <ul> <li> <p>unwrapped - continuous phase distribution (<em>float32</em>)</p> </li> <li> <p>wrapped - phase wrapped into mod2<span class="math-tex">\(\pi\)</span> (<em>float32</em>)</p> </li> <li> <p>wrapcount - wrap count phase maps coded in the integer form (0,1,2...) (<em>uint8</em>)</p> </li> <li> <p>noised - wrapped phase images w/ synthetic noise (<em>float32</em>).</p> </li> </ul> </li> <li> <p><strong>Trained models</strong> - trained model files. These model files are in the format "<strong>.h5</strong>", which contains the model architecture and the weights. They have been developed and saved with the <em>keras</em> library, and are loaded with the <em>keras.models.load_model()</em> function. The models list:</p> <ul> <li> <p><em>Unet_Denoising.h5 </em>- U-Net model used for denoising as an image translation task. The input is a wrapped phase image with noise and the output is the same wrapped phase distribution, but denoised. Model is trained on the synthetic phase dataset.</p> </li> <li> <p><em>Attn_Unet_Unwrapping.h5 </em>- U-Net model with Attention Gates and Residual Blocks trained for the semantic segmentation task. The input of the model is the wrapped phase image and its output is the wrap count map. Model is trained on the real phase dataset.</p> </li> </ul> </li> </ol> <p> </p> <p> </p> <p>[1] M. Baczewska, W. Krauze, A. Kuś, P. Stępień, K. Tokarska, K. Zukowski, E. Malinowska, Z. Brzózka, and M. Kujawińska, “On-chip holographic tomography for quantifying refractive index changes of cells’ dynamics,” in Quantitative Phase Imaging VIII, vol. 11970 Y. Liu, G. Popescu, and Y. Park, eds., International Society for Optics and Photonics (SPIE, 2022), p. 1197008.<br> [2] P. Stępień, M. Ziemczonok, M. Kujawińska, M. Baczewska, L. Valenti, A. Cherubini, E. Casirati, and W. Krauze, “Numerical refractive index correction for the stitching procedure in tomographic quantitative phase imaging,” Biomed. Opt. Express 13, 5709–5720 (2022).<br> [3] M. Ziemczonok, A. Kuś, P. Wasylczyk, and M. Kujawińska, “3d-printed biological cell phantom for testing 3d quantitative phase imaging systems,” Sci. Reports 9, 1–9 (2019).<br> [4] M. Ziemczonok, A. Kuś, and M. Kujawińska, “Optical diffraction tomography meets metrology — measurement accuracy on cellular and subcellular level,” Measurement 195, 111106 (2022).<br> [5] W. Krauze, A. Kuś, M. Ziemczonok, M. Haimowitz, S. Chowdhury, and M. Kujawińska, “3d scattering microphantom sample to assess quantitative accuracy in tomographic phase microscopy techniques,” Sci. Reports 12, 1–9 (2022).</p>
An analyst-created, high-resolution seismic arrival time dataset for evaluating machine learning phase detectors
<p>This data set contains phase arrival times and phase labels for one hour of continuous seismic data recorded at the three-component broadband station WY.YNR on 2014 30 March from 13:00:00 to 14:00:00 UTC. This hour of data follows a M<sub>w</sub> 4.8 occurring at 12:34 UTC in the Yellowstone region and contains many small events close together in space and time. All 687 picks (404 P and 283 S) were made by a seismic analyst trained at the University of Utah Seismograph Stations. The goal was to pick as many reasonable arrivals as possible for evaluating the performance of machine-learning-based phase detectors on continuous data during high seismicity rates. Phase picks were made using the Seismic Analysis Code (SAC; Goldstein and Snoke, 2005) and Pyrocko (Heimann <em>et al</em>., 2017). In general, a 1 to 17 Hz bandpass filter was used.</p> <p>The csv file contains the pick id, phase arrival time in UTC and Unix epoch formats, and the phase labels (P or S).</p>
Hybrid quantum-classical machine learning for generative chemistry and drug design: Generated molecules
<p>Deep generative chemistry models emerge as powerful tools to expedite drug discovery. How- ever, the immense size and complexity of the structural space of all possible drug-like molecules pose significant obstacles, which could be overcome with hybrid architectures combining quantum computers with deep classical networks. As the first step toward this goal, we built a compact discrete variational autoencoder (DVAE) with a Restricted Boltzmann Machine (RBM) of reduced size in its latent layer. The size of the proposed model was small enough to fit on a state-of-the-art D-Wave quantum annealer and allowed training on a subset of the ChEMBL dataset of biologically active compounds. Finally, we generated 2331 novel chemical structures with medicinal chemistry and synthetic accessibility properties in the ranges typical for molecules from ChEMBL. The pre- sented results demonstrate the feasibility of using already existing or soon-to-be-available quantum computing devices as testbeds for future drug discovery applications.</p>
Datasets and codes for the peer review article "Human and natural impacts on the U.S. freshwater salinization and alkalinization: A machine learning approach"
<p>Ongoing salinization and alkalinization in U.S. rivers have been attributed to inputs of road salt and effects of human-accelerated weathering in previous studies. Salinization poses a severe threat to human and ecosystem health, while human derived alkalinization implies increasing uncertainty in the dynamics of terrestrial sequestration of atmospheric carbon dioxide. A mechanistic understanding of whether and how human activities accelerate weathering and contribute to the geochemical changes in U.S. rivers is lacking. To address this uncertainty, we compiled dissolved sodium (salinity proxy) and alkalinity values along with 32 watershed properties ranging from hydrology, climate, geomorphology, geology, soil chemistry, land use, and land cover for 226 river monitoring sites across the coterminous U.S. Using these data, we built two machine-learning models to predict monthly-aggregated sodium and alkalinity fluxes at these sites. The sodium-prediction model detected human activities (represented by population density and impervious surface area) as major contributors to the salinity of U.S. rivers. In contrast, the alkalinity-prediction model identified natural processes as predominantly contributing to variation in riverine alkalinity flux, including runoff, carbonate sediment or siliciclastic sediment, soil pH and soil moisture. Unlike prior studies, our analysis suggests that the alkalinization in U.S. rivers is largely governed by local climatic and hydrogeological conditions.</p>
Data used in Machine learning reveals the waggle drift's role in the honey bee dance communication system
<p><strong>Data and metadata used in "Machine learning reveals the waggle drift’s role in the honey bee dance communication system" </strong></p> <p>All timestamps are given in ISO 8601 format.</p> <p><strong>The following files are included:</strong></p> <p><strong>Berlin2019_waggle_phases.csv, Berlin2021_waggle_phases.csv</strong></p> <p>Automatic individual detections of waggle phases during our recording periods in 2019 and 2021.</p> <ul> <li> <p>timestamp: Date and time of the detection.</p> </li> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> <li> <p>x_median, y_median: Median position of the bee during the waggle phase (for 2019 given in millimeters after applying a homography, for 2021 in the original image coordinates).</p> </li> <li> <p>waggle_angle: Body orientation of the bee during the waggle phase in radians (0: oriented to the right, PI / 4: oriented upwards).</p> </li> </ul> <p><strong>Berlin2019_dances.csv</strong></p> <p>Automatic detections of dance behavior during our recording period in 2019.</p> <ul> <li> <p>dancer_id: Unique ID of the individual bee.</p> </li> <li> <p>dance_id: Unique ID of the dance.</p> </li> <li> <p>ts_from, ts_to: Date and time of the beginning and end of the dance.</p> </li> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> <li> <p>median_x, median_y: Median position of the individual during the dance.</p> </li> <li> <p>feeder_cam_id: ID of the feeder that the bee was detected at prior to the dance.</p> </li> </ul> <p><strong>Berlin2019_followers.csv</strong></p> <p>Automatic detections of attendance and following behavior, corresponding to the dances in Berlin2019_dances.csv.</p> <ul> <li> <p>dance_id: Unique ID of the dance being attended or followed.</p> </li> <li> <p>follower_id: Unique ID of the individual attending or following the dance.</p> </li> <li> <p>ts_from, ts_to: Date and time of the beginning and end of the interaction.</p> </li> <li> <p>label: “attendance” or “follower”</p> </li> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> </ul> <p><strong>Berlin2019_dances_with_manually_verified_times.csv</strong></p> <p>A sample of dances from Berlin2019_dances.csv where the exact timestamps have been manually verified to correspond to the beginning of the first and last waggle phase down to a precision of ca. 166 ms (video material was recorded at 6 FPS).</p> <ul> <li> <p>dance_id: Unique ID of the dance.</p> </li> <li> <p>dancer_id: Unique ID of the dancing individual.</p> </li> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> <li> <p>feeder_cam_id: ID of the feeder that the bee was detected at prior to the dance.</p> </li> <li> <p>dance_start, dance_end: Manually verified date and times of the beginning and end of the dance.</p> </li> </ul> <p><strong>Berlin2019_dance_classifier_labels.csv</strong></p> <p>Manually annotated waggle phases or following behavior for our recording season in 2019 that was used to train the dancing and following classifier. Can be merged with the supplied individual detections.</p> <ul> <li> <p>timestamp: Timestamp of the individual frame the behavior was observed in.</p> </li> <li> <p>frame_id: Unique ID of the video frame the behavior was observed in.</p> </li> <li> <p>bee_id: Unique ID of the individual bee.</p> </li> <li> <p>label: One of “nothing”, “waggle”, “follower”</p> </li> </ul> <p><strong>Berlin2019_dance_classifier_unlabeled.csv</strong></p> <p>Additional unlabeled samples of timestamp and individual ID with the same format as Berlin2019_dance_classifier_labels.csv, but without a label. The data points have been sampled close to detections of our waggle phase classifier, so behaviors related to the waggle dance are likely overrepresented in that sample.</p> <p><strong>Berlin2021_waggle_phase_classifier_labels.csv</strong></p> <p>Manually annotated detections of our waggle phase detector (bb_wdd2) that were used to train the neural network filter (bb_wdd_filter) for the 2021 data.</p> <ul> <li> <p>detection_id: Unique ID of the waggle phase.</p> </li> <li> <p>label: One of “waggle”, “activating”, “ventilating”, “trembling”, “other”. Where “waggle” denoted a waggle phase, “activating” is the shaking signal, “ventilating” is a bee fanning her wings. “trembling” denotes a tremble dance, but the distinction from the “other” class was often not clear, so “trembling” was merged into “other” for training.</p> </li> <li> <p>orientation: The body orientation of the bee that triggered the detection in radians (0: facing to the right, PI /4: facing up).</p> </li> <li> <p>metadata_path: Path to the individual detection in the same directory structure as created by the waggle dance detector.</p> </li> </ul> <p><strong>Berlin2021_waggle_phase_classifier_ground_truth.zip</strong></p> <p>The output of the waggle dance detector (bb_wdd2) that corresponds to Berlin2021_waggle_phase_classifier_labels.csv and is used for training. The archive includes a directory structure as output by the bb_wdd2 and each directory includes the original image sequence that triggered the detection in an archive and the corresponding metadata. The training code supplied in bb_wdd_filter directly works with this directory structure.</p> <p><strong>Berlin2019_tracks.zip</strong></p> <p>Detections and tracks from the recording season in 2019 as produced by our tracking system. As the full data is several terabytes in size, we include the subset of our data here that is relevant for our publication which comprises over 46 million detections. We included tracks for all detected behaviors (dancing, following, attending) including one minute before and after the behavior. We also included all tracks that correspond to the labeled and unlabeled data that was used to train the dance classifier including 30 seconds before and after the data used for training.<br> We grouped the exported data by date to make the handling easier, but to efficiently work with the data, we recommend importing it into an indexable database.</p> <p>The individual files contain the following columns:</p> <ul> <li> <p>cam_id: Camera ID (0: left side of the hive, 1: right side of the hive).</p> </li> <li> <p>timestamp: Date and time of the detection.</p> </li> <li> <p>frame_id: Unique ID of the video frame of the recording from which the detection was extracted.</p> </li> <li> <p>track_id: Unique ID of an individual track (short motion path from one individual). For longer tracks, the detections can be linked based on the bee_id.</p> </li> <li> <p>bee_id: Unique ID of the individual bee.</p> </li> <li> <p>bee_id_confidence: Confidence between 0 and 1 that the bee_id is correct as output by our tracking system.</p> </li> <li> <p>x_pos_hive, y_pos_hive: Spatial position of the bee in the hive on the side indicated by cam_id. Given in millimeters after applying a homography on the video material.</p> </li> <li> <p>orientation_hive: Orientation of the bees’ thorax in the hive in radians (0: oriented to the right, PI / 4: oriented upwards).</p> </li> </ul> <p><strong>Berlin2019_feeder_experiment_log.csv</strong></p> <p>Experiment log for our feeder experiments in 2019.</p> <ul> <li> <p>date: Date given in the format year-month-day.</p> </li> <li> <p>feeder_cam_id: Numeric ID of the feeder.</p> </li> <li> <p>coordinates: Longitude and latitude of the feeder. For feeders 1 and 2 this is only given once and held constant. Feeder 3 had varying locations.</p> </li> <li> <p>time_opened, time_closed: Date and time when the feeder was set up or closed again.<br> sucrose_solution: Concentration of the sucrose solution given as sugar:water (in terms of weight). On days where feeder 3 was open, the other two feeders offered water without sugar.</p> </li> </ul> <p> </p> <ul> </ul> <p><strong>Software used to acquire and analyze the data:</strong></p> <ul> <li> <p><a href="https://github.com/BioroboticsLab/bb_pipeline">bb_pipeline: Tag localization and decoding pipeline</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_pipeline_models">bb_pipeline_models: Pretrained localizer and decoder models for bb_pipeline</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_binary">bb_binary: Raw detection data storage format</a></p> </li> <li> <p><a href="https://doi.org/10.5281/zenodo.4436419">bb_irflash: IR flash system schematics and arduino code</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_imgacquisition">bb_imgacquisition: Recording and network storage </a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_behavior">bb_behavior: Database interaction and data (pre)processing, feature extraction</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_tracking">bb_tracking: Tracking of bee detections over time</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_wdd2">bb_wdd2: Automatic detection and decoding of honey bee waggle dances</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_wdd_filter/">bb_wdd_filter: Machine learning model to improve the accuracy of the waggle dance detector</a></p> </li> <li> <p><a href="https://github.com/BioroboticsLab/bb_dance_networks/tree/master/bb_dance_networks">bb_dance_networks: Detection of dancing and following behavior from trajectories</a></p> </li> </ul> <p> </p>
Datasets for "Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Classification Algorithms in Reiner Gamma and Mare Ingenii"
<p>Final surface reflectance data at 2.6 m/pixel resolution with floating point values are available as GeoTiff and ASCII text files. Definition files for the K-Means and MLC algorithms in classifying swirl units are also available as ASCII text files. See README file for further details.</p> <p>Data used in the research article:</p> <p>Chuang, F.C., M.D. Richardson, J.R. Weirich, A.A. Sickafoose, and D.L. Domingue, 2022. Mapping Lunar Swirls with Machine Learning: The Application of Unsupervised and Supervised Image Classification Algorithms in Reiner Gamma and Mare Ingenii. The Planetary Science Journal, 3:231. doi://10.3847/PSJ/ac8f43</p> <p> </p>
MEWL: Few-shot multimodal word learning with referential uncertainty
<p><strong>Dataset Release for <a href="https://arxiv.org/abs/2306.00503">MEWL: Few-shot multimodal word learning with referential uncertainty (ICML 2023) </a></strong></p> <p><strong>GitHub:</strong> <a href="https://github.com/jianggy/MEWL">https://github.com/jianggy/MEWL</a></p> <p><strong>Abstract: </strong>Without explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning. Despite recent advancements in multimodal learning, a systematic and rigorous evaluation is still missing for human-like word learning in machines. To fill in this gap, we introduce the MachinE Word Learning (MEWL) benchmark to assess how machines learn word meaning in grounded visual scenes. MEWL covers human's core cognitive toolkits in word learning: cross-situational reasoning, bootstrapping, and pragmatic learning. Specifically, MEWL is a few-shot benchmark suite consisting of nine tasks for probing various word learning capabilities. These tasks are carefully designed to be aligned with the children's core abilities in word learning and echo the theories in the developmental literature. By evaluating multimodal and unimodal agents' performance with a comparative analysis of human performance, we notice a sharp divergence in human and machine word learning. We further discuss these differences between humans and machines and call for human-like few-shot word learning in machines.</p>
learn2learn: Few-Shot Learning Datasets
<p>Few-shot learning datasets, including:</p> <ul> <li>mini-ImageNet</li> <li>tiered-ImageNet</li> <li>CIFAR-FS</li> <li>FC100</li> </ul> <p>New in 1.0.1:</p> <ul> <li>FGVC Fungi</li> <li>FGVC Aircrafts</li> <li>Describable Textures</li> <li>VGG Flowers</li> <li>CUB200</li> </ul> <p>Please cite the respective datasets if you use them, not this archive.</p>
MNIST-Federated-Learning
<p>Please find below the descriptions of the three configurations for partitioning <strong>the MNIST Train dataset into 10 clients and the MNIST Train data: </strong><br> </p> <ol> <li><strong>Balanced Distribution:</strong> In the first configuration, the MNIST dataset is partitioned among 10 clients in a balanced manner. This means that the data samples from each class are evenly distributed among the clients. Each client receives a roughly equal number of images from each digit class, ensuring that the distribution of samples across clients is proportional and representative of the overall dataset. <strong> [ Config 1]</strong></li> <li><strong>Heterogeneous Distribution (One Class per Client)</strong>: In the second configuration, the MNIST dataset is partitioned in a heterogeneous manner, where each client is assigned a single digit class exclusively. This means that one client will only receive images of the digit '0', another client will receive images of the digit '1', and so on. In this setup, each client becomes an expert in classifying a specific digit, allowing for specialized training and evaluation. <strong>[ Config 2]</strong></li> <li><strong>Mixed Distribution:</strong> In the third configuration, the MNIST dataset is partitioned using a mixed distribution approach. This means that the data samples from all digit classes are distributed among the 10 clients, but the distribution is not necessarily balanced. The number of samples assigned to each client may vary for different digit classes, resulting in an uneven distribution across the clients. This configuration aims to capture both the overall diversity of the dataset and the varying difficulty levels of classifying different digits. <strong>[ Config 3 ]</strong></li> </ol> <p> </p> <p>Mnist-dataset/<br> ├── config1/<br> │ ├── client-1/<br> │ │ └── data.csv<br> │ ├── client-2/<br> │ │ └── data.csv<br> │ ├── client-3/<br> │ │ └── data.csv<br> │ └── ...<br> ├── config2/<br> │ ├── client-1/<br> │ │ └── data.csv<br> │ ├── client-2/<br> │ │ └── data.csv<br> │ ├── client-3/<br> │ │ └── data.csv<br> │ └── ...<br> ├── config3/<br> │ ├── client-1/<br> │ │ └── data.csv<br> │ ├── client-2/<br> │ │ └── data.csv<br> │ ├── client-3/<br> │ │ └── data.csv<br> │ └── ...<br> └── mnist_test.csv<br> </p> <p>***</p> <p>License: Yann LeCun and Corinna Cortes hold the copyright of MNIST dataset, which is a derivative work from original NIST datasets. MNIST dataset is made available under the terms of the <a href="https://creativecommons.org/licenses/by-sa/3.0/">Creative Commons Attribution-Share Alike 3.0 license.</a></p> <p>***</p>
PFAM Protein Families Dataset for Machine Learning
<p>A cleaned dataset of protein sequences and protein families for classification. The dataset is exported from PFAM as of June 2023 and curated to achieve the following characteristics:</p> <ul> <li>only protein families included with >=100 sequences</li> <li>families with >2000 sequences are truncated and only represented by 2000 sequences (chosen randomly)</li> <li>only proteins with sequence lengths between 100 and 1000</li> <li>amino acid sequences are form PDB; chains are concatenated only if not similar</li> </ul> <p>The dataset is not balanced, numbers of sequences per family in PFAM and in in dataset are:</p> <pre><code>families: 62, sequences: 46872 total (in PFAM) -> included (in dataset) Number in family ALLERGEN: 122 -> 122 Number in family APOPTOSIS: 381 -> 381 Number in family BIOSYNTHETIC PROTEIN: 346 -> 346 Number in family BIOTIN BINDING PROTEIN: 165 -> 165 Number in family BLOOD CLOTTING: 138 -> 138 Number in family CALCIUM BINDING PROTEIN: 135 -> 135 Number in family CELL ADHESION: 1116 -> 1116 Number in family CELL CYCLE: 511 -> 511 Number in family CHAPERONE: 964 -> 964 Number in family CONTRACTILE PROTEIN: 158 -> 158 Number in family CYTOKINE: 191 -> 191 Number in family DE NOVO PROTEIN: 253 -> 253 Number in family DNA BINDING PROTEIN: 1008 -> 1008 Number in family ELECTRON TRANSPORT: 841 -> 841 Number in family FLUORESCENT PROTEIN: 348 -> 348 Number in family GENE REGULATION: 607 -> 607 Number in family HORMONE: 272 -> 272 Number in family HORMONE GROWTH FACTOR: 159 -> 159 Number in family HORMONE RECEPTOR: 121 -> 121 Number in family HYDROLASE: 19551 -> 2000 Number in family HYDROLASE ANTIBIOTIC: 120 -> 120 Number in family HYDROLASE HYDROLASE INHIBITOR: 2890 -> 2000 Number in family HYDROLASE INHIBITOR: 315 -> 315 Number in family IMMUNE SYSTEM: 3333 -> 2000 Number in family IMMUNOGLOBULIN: 155 -> 155 Number in family ISOMERASE: 2457 -> 2000 Number in family ISOMERASE ISOMERASE INHIBITOR: 139 -> 139 Number in family LECTIN: 139 -> 139 Number in family LIGASE: 1780 -> 1780 Number in family LIGASE LIGASE INHIBITOR: 163 -> 163 Number in family LIPID BINDING PROTEIN: 421 -> 421 Number in family LIPID TRANSPORT: 115 -> 115 Number in family LUMINESCENT PROTEIN: 221 -> 221 Number in family LYASE: 4150 -> 2000 Number in family LYASE LYASE INHIBITOR: 298 -> 298 Number in family MEMBRANE PROTEIN: 1338 -> 1338 Number in family METAL BINDING PROTEIN: 951 -> 951 Number in family METAL TRANSPORT: 409 -> 409 Number in family MOTOR PROTEIN: 195 -> 195 Number in family OXIDOREDUCTASE: 11531 -> 2000 Number in family OXIDOREDUCTASE OXIDOREDUCTASE INHIBITOR: 766 -> 766 Number in family OXYGEN STORAGE: 127 -> 127 Number in family OXYGEN STORAGE TRANSPORT: 260 -> 260 Number in family OXYGEN TRANSPORT: 414 -> 414 Number in family PHOTOSYNTHESIS: 173 -> 173 Number in family PLANT PROTEIN: 255 -> 255 Number in family PROTEIN BINDING: 1613 -> 1613 Number in family PROTEIN TRANSPORT: 693 -> 693 Number in family RECEPTOR: 108 -> 108 Number in family REPLICATION: 161 -> 161 Number in family RNA BINDING PROTEIN: 546 -> 546 Number in family SIGNALING PROTEIN: 2312 -> 2000 Number in family STRUCTURAL PROTEIN: 869 -> 869 Number in family SUGAR BINDING PROTEIN: 1250 -> 1250 Number in family TOXIN: 546 -> 546 Number in family TRANSCRIPTION REGULATION: 3283 -> 2000 Number in family TRANSFERASE: 14724 -> 2000 Number in family TRANSFERASE INHIBITOR: 126 -> 126 Number in family TRANSFERASE TRANSFERASE INHIBITOR: 2465 -> 2000 Number in family TRANSLATION: 370 -> 370 Number in family TRANSPORT PROTEIN: 2782 -> 2000 Number in family VIRAL PROTEIN: 2150 -> 2000</code></pre> <p>Files:</p> <ul> <li>families.csv: list of protein families with frequencies</li> <li>pfam_46872x62.csv: full dataset with amino acid sequences as string (one-letter code)</li> <li>pfam-trn-xy.csv: training dataset with amino acid sequences as tokens (1..25) and padded to a common length of 1000 with padding token 0:</li> </ul> <pre><code> Amino acid | Token | Description -------------------------------- C | 1 | Cysteine S | 2 | Serine T | 3 | Threonine A | 4 | Alanine G | 5 | Glycine P | 6 | Proline D | 7 | Aspartic acid E | 8 | Glutamic acid Q | 9 | Glutamine N | 10 | Asparagine H | 11 | Histidine R | 12 | Arginine K | 13 | Lysine M | 14 | Methionine I | 15 | Isoleucine L | 16 | Leucine V | 17 | Valine W | 18 | Tryptophan Y | 19 | Tyrosine F | 20 | Phenylalanine B | 21 | Aspartic acid or Asparagine Z | 22 | Glutamic acid or Glutamine J | 23 | Leucine or Isoleucine U | 24 | Selenocysteine X | 25 | Unknown amino acid . | 0 | padding token</code></pre> <p> </p> <ul> <li>pfam-trn-labels.csv: plain-text labels for training data</li> <li>pfam-tst-xy.csv</li> <li>pfam-tst-labels.csv: test data</li> <li>pfam-balanced-trn-xy.csv</li> <li>pfam-balanced-trn-labels.csv:</li> <li>pfam-balanced-tst-xy.csv</li> <li>pfam-balanced-tst-labels.csv: balanced datasets, created by oversampling.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.