Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Dataset for Machine Learning Assisted Citation Screening for Systematic Reviews
<p>The work "Machine Learning Assisted Citation Screening for Systematic Reviews" explored the problem of citation screening automation using machine-learning (ML) with an aim to accelerate the process of generating <a href="https://en.wikipedia.org/wiki/Systematic_review#:~:text=Systematic%20reviews%20are%20a%20type,synthesize%20findings%20qualitatively%20or%20quantitatively." rel="nofollow">systematic reviews</a>. Manual process of citation screening involve two reviewers manually screening the searched studies using a predefined inclusion criteria. If the study passes the "inclusion" criteria, it is included for further analysis or is excluded. As apparant through manual screening process, the work considered citation screening as a binary classification problem whereby any ML classifier could be trained to separate the searched studies into these two classes (include and exclude).</p> <p> </p> <p>A physiotherapy citation screening dataset was used to test automation approaches and the dataset includes the studies identified for citation screening in an update to the systematic review by Hilfiker <em>et al.</em> The dataset included titles and abstracts (citations) from 31,279 (deduplicated: 25,540) studies identified during the search phase of this SR. These studies were already manually assessed for relevance and labelled by two reviewers into two mutually exclusive labels. The uploaded file consists of 25,540 data samples, with each data sample separated by a new line. It is a tab separated file and the data in it is structured as shown below. This dataset was manually labelled into include and exclude by Hilfiker <em>et al.</em></p> <p> </p> <table> <tbody> <tr> <td><strong>Title</strong></td> <td><strong>PMID</strong></td> <td><strong>Abstract </strong></td> <td><strong>Class</strong></td> <td><strong>MeSH terms (separated by a pipe)</strong></td> </tr> <tr> <td>Structured exercise improves physical functioning in women with stages I and II breast cancer: results of a randomized controlled trial. </td> <td>11157015</td> <td>Abstract PURPOSE: Self-directed and supervised exercise were compared with usual care in a clinical trial designed to evaluate the effect of structured exercise on physical functioning and other dimensions of health-related quality of life in women with stages I and II breast cancer. PATIENTS AND METHODS: One hundred twenty-three women with stages I and II breast cancer completed baseline evaluations of generic and disease- and site-specific health-related quality of life, aerobic capacity, and body weight. Participants were randomly allocated to one of three intervention groups: usual care (control group), self-directed exercise, or supervised exercise. Quality of life, aerobic capacity, and body weight measures were repeated at 26 weeks...</td> <td>include or exclude</td> <td>Clinical Trial | Comparative Study | Randomized Controlled Trial | Research Support, Non-U.S. Gov't | Antineoplastic Combined Chemotherapy Protocols | Breast Neoplasms | Breast Neoplasms | Breast Neoplasms | Chemotherapy, Adjuvant | Exercise | Female | Humans | Middle Aged | Neoplasm Staging | Quality of Life | Radiotherapy, Adjuvant</td> </tr> </tbody> </table> <p> </p> <p>If you use this dataset in your research, please cite our papers.</p>
Stiffness Moduli Modelling and Prediction in Four-Point Bending of Asphalt Mixtures: A Machine Learning-Based Framework within Weave-UNISONO 2021 project, NCN project No 2021/03/Y/ST8/00079, and GACR project GA22-04047K
<div><strong>Summary:</strong></div> <div>Two selected mixtures were thoroughly investigated in an experimental trial carried out by means of a four-point bending test (4PBT) apparatus. The mixtures were prepared using spilite aggregate, a conventional 50/70 penetration grade bitumen, and limestone filler. Their stiffness moduli (SM) were determined while samples were exposed to 11 loading frequencies (from 0.1 to 50 Hz) and 4 testing temperatures (from 0 to 30 °C). Observations were recorded and used to develop a machine learning (ML) model. The main scope was the prediction of the stiffness moduli based on the volumetric properties and testing conditions of the corresponding mixtures, which would provide the advantage of reducing the laboratory efforts required to determine them.</div> <div> </div> <div><strong>The dataset includes:</strong></div> <div>Characteristics of bituminous binder, CSV raw data</div> <div> <ul> <li>bituminous binder.csv</li> </ul> </div> <div>Grading curves of tested asphalt mixtures</div> <ul> <li>AML16 Grading curves.csv</li> <li>AMP22 Grading curves.csv</li> </ul> <div>Volumetric characterizations of AML16 and AMP22 mixtures</div> <ul> <li>AML16 Volumetric characterizations.csv</li> <li>AMP22 Volumetric characterizations.csv</li> </ul> <div>Outcomes of the 4PBT experimental trial carried out on AML16 and AMP22 mixtures</div> <ul> <li>AML16 Stiffness Modulus 4PB.csv</li> <li>AMP22 Stiffness Modulus 4PB.csv</li> </ul>
Database for machine learning of hydrogen storage materials properties
<p><strong>Database for machine learning of hydrogen storage materials properties</strong></p> <p>Matthew Witman<sup>a</sup>, Mark Allendorf<sup>a</sup>, Vitalie Stavila<sup>a</sup></p> <p><sup>a</sup>Sandia National Laboratories, Livermore, CA</p> <p> </p> <p><strong>Description</strong></p> <p>This ML-HydPARK dataset provides a csv file of metal hydride compositions, capacities, and thermodynamic values that can be used as target properties for building, training, and testing machine learning models. It has been parsed and cleaned from the DOE’s original publicly available HydPARK database according to the procedure in [1] to make it more suitable for immediate use with data-driven models. Generally, this removed duplicate entries, removed entries missing critical data, and attempted to fix various entries with obvious errors in the data. It is continuously updated under version control as new metal alloy hydrides are published in the open literature. Most entries contain data on the enthalpy and entropy of the hydriding reaction, as well the maximum hydrogen capacity, for which compositional machine learning models can be trained [1,2].</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>The authors gratefully acknowledge research support from the U.S. Department of Energy, Office of Energy Efficiency and Renewable Energy, Fuel Cell Technologies Office through the Hydrogen Storage Materials Advanced Research Consortium (HyMARC). This work was supported by the Laboratory Directed Research and Development (LDRD) program at Sandia National Laboratories. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. This paper describes objective technical results and analysis. Any subjective views or opinions that might be expressed in the paper do not necessarily represent the views of the U.S. Department of Energy<br> or the United States Government.</p> <p> </p> <p><strong>References</strong></p> <ol> <li>Witman, M.; Ling, S.; Grant, D. M.; Walker, G. S.; Agarwal, S.; Stavila, V.; Allendorf, M. D. Extracting an Empirical Intermetallic Hydride Design Principle from Limited Data via Interpretable Machine Learning. <em>J. Phys. Chem. Lett</em>. <strong>2020</strong>, 11, 40–47.</li> <li>Witman, M.; Ek, G.; Ling, S.; Chames, J.; Agarwal, S.; Wong, J.; Allendorf, M. D.; Sahlberg, M.; Stavila, V. Data-Driven Discovery and Synthesis of High Entropy Alloy Hydrides with Targeted Thermodynamic Stability. <em>Chem. Mater</em>. <strong>2021</strong>, 33, 4067–4076.</li> </ol> <p> </p> <p><strong>Contact</strong></p> <p>Please email <a href="mailto:mwitman@sandia.gov">mwitman@sandia.gov</a> , <a href="mailto:mdallen@sandia.gov">mdallen@sandia.gov</a>, or <a href="mailto:vnstavi@sandia.gov">vnstavi@sandia.gov</a> for questions or to request addition of recent data from the literature to this dataset.</p>
Machine learning predicts earthquakes in the continuum model of a rate-and-state fault with frictional heterogeneities
<p>Numerical data used to make Figures in the manuscript entitled "Machine learning predicts earthquakes in the continuum model of a rate-and-state fault with frictional heterogeneities". We provide the data to create Figures 1 to 4 from the main text and Figures S1 to S9 from the supplementary information. We also provide Python scripts to plot them.</p>
3d Transition Metal K-edge XANES Dataset for Machine Learning Models
<p><strong>Data</strong><br><br>This dataset contains machine learning data for K-edge X-ray Absorption Near-Edge Structure (XANES) prediction models for eight 3d transition metals (Ti -Cu).</p> <ul> <li><strong>features_and_spectra:</strong> Material features (X) and corresponding XAS spectra (y) for each dataset split: training (train), validation (val), and test.</li> <li><strong> material_id_and_site:</strong> Material identifiers and site indices (according to <a href="https://github.com/AI-multimodal/Lightshow">Lightshow</a>) for each dataset split. </li> </ul> <p><strong>Funding</strong><br><br>This research is based upon work supported by the U.S. Department of Energy, Office of Science, Office Basic Energy Sciences, under Award Number FWP PS-030. This research also used theory and computational resources of the Center for Functional Nanomaterials, which is a U.S. Department of Energy Office of Science User Facility, and the Scientific Data and Computing Center, at Brookhaven National Laboratory under Contract No. DE-SC0012704 and by Brookhaven National Laboratory (BNL), Laboratory Directed Research and Development (LDRD) grant no. 24-004.</p> <p> </p>
EPL Points Prediction Using Machine Learning
<p>This project leverages machine learning techniques to predict match outcomes and the final standings of the English Premier League (EPL). By utilizing historical match data collected through web scraping, we built a robust prediction model that forecasts the results of upcoming fixtures and projects the season's points table.The data pipeline includes comprehensive data preprocessing, feature engineering, and the application of five machine learning algorithms: K-Nearest Neighbors (KNN), Random Forest, AdaBoost, XGBoost, and CatBoost. These models were evaluated using accuracy metrics, achieving up to 96% accuracy in predicting match outcomes. Key features such as team form, home/away advantage, and head-to-head records were engineered to enhance the model's predictive power.The best-performing models, XGBoost and CatBoost, were used to simulate the remaining fixtures, allowing us to generate a predicted points table for the EPL. This project demonstrates the potential of data-driven approaches in sports analytics, offering insights for football clubs, analysts, and enthusiasts looking to understand team performance dynamics.</p>
Machine learning-based quality assessment of Antarctic margins salinity - code, data and figures
<p>The submission contains the data, functions and code needed to reproduce the figures in Sohail et al., 2025</p>
Reference Mean and Low Streamflow for all Brazilian Catchments Generated Using Machine Learning Models
<p>This dataset provides comprehensive hydrological information for river networks in Brazil, focusing on reference streamflows, specifically long-term mean flows (qm) and low flows exceeded 95% of the time (q95). Covering over 400,000 ungauged river points, the dataset was developed using advanced machine learning models trained on environmental descriptors and validated against data from 1,069 gauging stations spread across the country. The machine learning pipeline evaluated six regression models to achieve high predictive accuracy (R² > 0.8 for qm and > 0.7 for q95). The 62 environmental descriptors - encompassing climate, topography, land cover, lithology, and water storage characteristics - that were used as features for the models are also included.</p> <p>Key features:</p> <ul> <li><strong>Spatial Coverage:</strong> Brazilian territory and the Amazon River basin, based on the <a href="https://metadados.snirh.gov.br/geonetwork/srv/api/records/f7b1fc91-f5bc-4d0d-9f4f-f4e5061e5d8f" target="_blank" rel="noopener">BHO 5k</a> dataset of officially adopted river networks.</li> <li><strong>Outputs:</strong> Predicted qm and q95 values for each river stretch, with 90% and 75% confidence intervals to account for prediction uncertainty.</li> <li><strong>Environmental Descriptors:</strong> Aggregated from upstream catchment area.</li> </ul> <p> </p>
Data supplement: Spatial autocorrelation in machine learning for modelling soil organic carbon
<p>Spatial autocorrelation in machine learning for modelling soil organic carbon: Data supplement</p> <p><br>Alexander Kmoch, Clay Taylor Harrison, Jeonghwan Choi, Evelyn Uuemaa</p> <p>Spatial autocorrelation, the relationship between nearby samples of a spatial<br>random variable, is often overlooked in machine learning models, leading to<br>biased results. This study investigates various methods to account for spa-<br>tial autocorrelation when predicting soil organic carbon (SOC) using random<br>forest models. Five models incorporating spatial structure were compared<br>against baseline models that did not have any added spatial components.<br>Cross-validation showed slight improvements in accuracy for models consid-<br>ering spatial autocorrelation, while Shapley Additive Explanations confirmed<br>the importance of spatial variables. However, no decrease in spatial autocor-<br>relation of residuals was observed. Raster-based models exhibited enhanced<br>prediction detail, but high-resolution validation data availability limited thor-<br>ough validation. The findings emphasize the value of incorporating spatial<br>autocorrelation for improved SOC prediction in machine learning models.<br>Considerations such as the distribution of predictions and computational<br>complexity should help guide the selection of suitable approaches for specific<br>spatial modelling tasks.</p>
Glossed Hittite Texts with German Translation for Machine Learning
<p>This dataset contains 7,099 processed Hittite texts from 143 CTH numbers, sourced with permission from the <a href="https://www.hethport.uni-wuerzburg.de/HPM/index.php" target="_blank" rel="noopener"><strong>Hethitologie Portal Mainz (HPM)</strong></a>, which is the comprehensive resource of Hittite texts and culture (modern Turkey, c. 1,650 - 1,200 BCE). These texts have been converted from XML format to a tabular structure for computational and linguistic analysis, as well as machine learning applications, while attempting to preserve the full complexity of the source material.</p> <p><strong>Key Features:</strong></p> <ul> <li><strong>txtid</strong>: Identifier for the specific text (e.g., "IBoT 1.30+"), referencing entries in the HPM Konkordanz (S. Košak, hethiter.net/: hetkonk (2.plus)).</li> <li><strong>lnr</strong>: Surface, column (if extant), and line number within the text.</li> <li><strong>cth_number</strong>: <em>Catalogue des textes hittites </em>(CTH) classification, organizing texts by genre and content (S. Košak – G.G.W. Müller – S. Görke – Ch.W. Steitler, hethiter.net/: CTH (2022-10-26)).</li> <li><strong>word</strong>: Orthographic form of the text as presented in the HPM, not in cuneiform but in a standard Latin script representation.</li> <li><strong>translit</strong>: Detailed transliteration, preserving all nuances such as diacritics, broken parts, and editorial markings to reflect the original condition on the text.</li> <li><strong>gloss</strong>: Linguistic glosses, including grammatical, morphological, and semantic information.</li> <li><strong>trans_de</strong>: German translation of the word or phrase.</li> </ul>
Antisemitism on Twitter: A Dataset for Machine Learning and Text Analytics
<h1><strong><span><span>Dataset from the Institute for the Study of Contemporary Antisemitism (ISCA) at Indiana University: </span></span></strong></h1> <p> </p> <div> <div> <p><span><span>The </span><span>Social Media</span><span> & Hate research lab at the Institute for the Study of Contemporary Antisemitism compiled this dataset using an annotation portal (Jikeli, Soemer, and Karali 2024), which was used to label tweets as either antisemitic or non-antisemitic, among other labels. Note that annotation was done on live data, including images and context, such as threads. All data was annotated by two experts, and all discrepancies were discussed</span><span> (Jikeli et al. 2023)</span><span>.</span></span><span> </span></p> </div> </div> <p><br><strong>Content: </strong></p> <p><span><span>This dataset </span><span>contains</span> <span>1</span><span>1</span><span>311</span><span> tweets </span><span>covering</span><span> a wide range of topics common in conversations about Jews, Israel, and antisemitism between January 2019 and </span><span>April 2023</span><span>. </span><span>The dataset consists of random samples of relevant keywords during this </span><span>time period</span><span>.</span><span> 1,</span><span>953</span><span> tweets (1</span><span>7</span><span>%) </span><span>are antisemitic </span><span>according to </span><span>the IHRA definition of antisemitism.</span><span> </span></span><span> </span></p> <div> <p><span><span>The distribution of tweets by year is as follows:</span><span> 1499 (</span><span>13</span><span>%) from 2019, 371</span><span>2</span><span> (</span><span>33</span><span>%) from 2020, </span><span>2591</span><span> (2</span><span>3</span><span>%) from 2021</span><span>, 2644 from 2022 </span><span>(23%)</span> <span>and 865 </span><span>(8%)</span> <span>f</span><span>rom 2023</span><span>. </span><span>6365</span><span> (</span><span>56</span><span>%) </span><span>contain</span><span> the keyword "Jews,"</span><span> 4134 </span><span>(</span><span>3</span><span>7</span><span>%) include "Israel," 529 (</span><span>5</span><span>%) feature the derogatory term "</span><span>ZioNazi</span><span>*," and 283 (</span><span>3</span><span>%) use the slur "K---s." Some tweets may </span><span>contain</span><span> multiple keywords. </span></span><span> </span></p> </div> <div> <p><span><span>725</span><span> out of the </span><span>6365</span><span> tweets with the keyword "Jews" (11%) and </span><span>664</span><span> out of the </span><span>4134</span><span> tweets with the keyword "Israel" (1</span><span>6</span><span>%) were classified as antisemitic. 97 out of the 283 tweets using the antisemitic slur "K---s" (34%) are antisemitic.</span> <span>Interestingly, many tweets featuring the slur "K---s" actually </span><span>call out</span><span> its u</span><span>s</span><span>e.</span><span> In contrast, </span><span>the majority of</span><span> tweets </span><span>using</span><span> the derogatory term "</span><span>ZioNazi</span><span>*" are antisemitic, with 467 out of 529 (88%) being classified as such. </span></span><span> </span></p> </div> <p> </p> <p><strong>File Description: </strong></p> <div> <div> <p><span><span>The dataset is provided in a csv file format, with each row </span><span>representing</span><span> a single message, including replies, quotes, and retweets. The file </span><span>contains</span><span> the following columns: </span></span><span> </span></p> </div> <div> <p><span><span> </span></span><span><span> </span><br></span><span><span>‘ID’:</span> <span>Represents</span><span> the tweet ID. </span></span><span> </span></p> </div> <div> <p><span><span>‘Username’: </span><span>Represents</span><span> the username </span><span>that posted </span><span>the tweet</span><span>. </span></span><span> </span></p> </div> <div> <p><span><span>‘Text’: </span><span>Represents</span><span> the full text of the tweet (not pre-processed).</span></span><span> </span></p> </div> <div> <p><span><span>‘</span><span>CreateDate</span><span>’: </span><span>Represents</span><span> the date </span><span>on which </span><span>the tweet was created</span><span>. </span></span><span> </span></p> </div> <div> <p><span><span>‘Biased’: </span><span>Represents</span><span> the label given by our annotations as to whether the tweet </span><span>is antisemitic or no</span><span>t</span><span>.</span></span><span> </span></p> </div> <div> <p><span><span>‘Keyword’: </span><span>Represents</span><span> the keyword that was used in the query. The keyword can be in the text, including </span><span>hashtags, </span><span>mentioned </span><span>users</span><span>, or the username</span><span> itself.</span><span> </span></span><span> </span></p> </div> </div> <p> </p> <p>Licences </p> <p>Data is published under the terms of the "Creative Commons Attribution 4.0 International" licence (https://creativecommons.org/licenses/by/4.0) </p> <p> </p> <p>Acknowledgements </p> <p>We are grateful for the support of Indiana University’s Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media & Hate Research Lab at Indiana University’s Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenhäusler, Sophie von Máriássy, Mabel Poindexter, Jenna Solomon, Clara Schilling, and Victor Tschiskale. </p> <p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.</p>
Parsimonious machine learning for the global mapping of aboveground biomass density
<p>This repository hosts data and code presented in the article "Parsimonious machine learning for the global mapping of aboveground biomass potential". The repository contains a compressed file containing all the code needed to reproduce the methodology that we developed and to analyse its results. We did not upload all the temporary and intermediate data files that are created during the execution of the method. We rather uploaded "milestone" data, i.e. final results or important intermediate ones. This includes the final training dataset, model calibration data, the final trained model, the global data for prediction, the final global map of potential aboveground biomass density (AGBD) at present times (raster files at 1km2 and 10km2 resolution), maps depicting regions where climatic conditions are outside of the training range of positive AGBD instances and maps depicting world regions without trees. </p> <p><strong>Files:</strong></p> <p><a href="../api/records/11580414/draft/files/code.zip/content" target="_blank" rel="noopener noreferrer">code.zip</a> : Compressed directory with all the code needed to reproduce the methodology presented in the manuscript. Contains a README file. Also contains temporary data generated in the process, the training dataset, the trained model, and model calibration data.</p> <p><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">potential_AGBD_Mgha_1km_present_climate_1980_2010.tif</a> : the predicted global potential AGBD under contemporary climate conditions and at a resolution of 1 squared kilometer.</p> <p><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">potential_AGBD_Mgha_10km_</a><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">present_climate_1980_2010.tif</a> : the predicted global potential AGBD under contemporary climate conditions downsampled at a resolution of 10 squared kilometers.</p> <p><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">potential_AGBD_Mgha_10km_model_difference.tif</a> : the difference between our prediction of potential AGBD and the prediction from a complex state-of-the-art model from Walker et al. (2022). </p> <p><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">potential_AGB_Mg_1km_</a><a href="../api/records/11580414/draft/files/potential_AGBD_Mgha_1km2_contemporary_climate.tif/content" target="_blank" rel="noopener noreferrer">present_climate_1980_2010.tif</a> : the predicted global potential pixel-level AGB under contemporary climate conditions downsampled at a resolution of 1 squared kilometers.</p> <p><a href="../api/records/11580414/draft/files/number_predictors_out_of_range.zip/content" target="_blank" rel="noopener noreferrer">number_predictors_out_of_range.zip</a> : tiled maps representing the number of climatic predictors outside of the training range before including 0 AGBD instances in the training dataset. </p> <p><a href="../api/records/11580414/draft/files/tree_absence_map.zip/content" target="_blank" rel="noopener noreferrer">tree_absence_map.zip</a> : tiled maps representing world regions without trees. Based on Crowther et al. (2015) (https://elischolar.library.yale.edu/yale_fes_data/1/).</p> <p><a href="../api/records/11580414/draft/files/potential_agbd_Mgha_climate_envelope.pkl/content" target="_blank" rel="noopener noreferrer">inference_pipeline_potential_agbd_Mgha_climate.pkl</a> : Calibrated model for the prediction of potential AGBD given bioclimatic conditions. </p> <p><a href="../api/records/11580414/draft/files/predictors_data_global.zip/content" target="_blank" rel="noopener noreferrer">predictors_data_global.zip</a> : Global predictors data to apply the model on.</p>
Machine Learning Predicts Meter-Scale Laboratory Earthquakes
<p>Numerical data and Python scripts used to make Figures in the manuscript entitled "Machine Learning Predicts Meter-Scale Laboratory Earthquakes". We provide the data and scripts to create Figures 1 to 6 from the main text and Figures S1 to S15 from the supplementary information. Note that you should download the experimental catalog of Yamashita et al. (Nature Comm, 2021) and address the request for shear force data (experimental number LB12-011) to Futoshi Yamashita, the organizer of the target experiment. This code is developed using Python 3.11.7.</p>
Simulated datasets for detector and particle flow reconstruction: CLIC detector, machine learning format
<p><strong>Synopsis</strong></p> <p>Machine-learning friendly format of tracks, clusters and target particles in electron-positron events, simulated with the CLIC detector. Ready to be used with <a href="https://zenodo.org/records/14930299">jpata/particleflow:v2.3.0</a>. Derived from the EDM4HEP ROOT files in <a href="https://zenodo.org/record/8260741">https://zenodo.org/record/8260741</a>.</p> <ul> <li>clic_edm_ttbar_pf.zip: e+e- -> ttbar, center of mass energy at 380 GeV</li> <li>clic_edm_qq_pf.zip: e+e- -> Z* -> qqbar, center of mass energy at 380 GeV</li> <li>clic_edm_ww_fullhad_pf.zip: e+e- -> WW -> W decaying hadronically, center of mass energy at 380 GeV</li> <li>clic-tfds.ipynb: an example notebook on how to load the files</li> </ul> <p><strong>Contents</strong></p> <p>Each .zip file contains the dataset in the <a href="https://github.com/tensorflow/datasets">tensorflow-datasets</a>, <a href="https://github.com/google/array_record">array_record</a> format. We have split the full datasets into 10 subsets, due to space considerations on zenodo, two subsets from each dataset are uploaded. Each dataset contains a train and test split of events.</p> <p><strong>Dataset semantics (to be updated)</strong></p> <p>Each dataset consists of events that can be iterated over using the tensorflow-datasets library and used in either tensorflow or pytorch. Each event has the following information available:</p> <ul> <li>X: the reconstruction input features, i.e. tracks and clusters</li> <li>ytarget: the ground truth particles with the features ["PDG", "charge", "pt", "eta", "sin_phi", "cos_phi", "energy", "jet_idx"], with "jet_idx" corresponding to the gen-jet assignment of this particle</li> <li>ycand: the baseline Pandora PF particles with the features ["PDG", "charge", "pt", "eta", "sin_phi", "cos_phi", "energy", "jet_idx"], with "jet_idx" corresponding to the gen-jet assignment of this particle</li> </ul> <p>The full semantics, including the list of features for X, are available at https://github.com/jpata/particleflow/blob/v2.3.0/mlpf/heptfds/clic_pf_edm4hep/utils_edm.py and https://github.com/jpata/particleflow/blob/v2.3.0/mlpf/data/key4hep/postprocessing.py.</p>
Machine-Learning of Aerosol-Cloud-Climate Interactions Reveals an increased Cloud Fraction
<p>Data presented in the manuscript "Machine-Learning of Aerosol-Cloud-Climate Interactions Reveals an increased Cloud Fraction" by Chen et al. (2022).</p>
Machine learning code and dataset for "Nowcasting thunderstorm hazards using machine learning: the impact of data sources on performance"
<p>This repository contains the code and dataset for the paper:</p> <p>Nowcasting thunderstorm hazards using machine learning: the impact of data sources on performance, Natural Hazards and Earth System Sciences, 2022, <a href="https://doi.org/10.5194/nhess-2021-171">https://doi.org/10.5194/nhess-2021-171</a></p> <p>The GitHub code repository at <a href="https://github.com/meteoswiss-mdr/ts-nowcast-datasources">https://github.com/meteoswiss-mdr/ts-nowcast-datasources</a> may contain a more up-to-date version of the code if bug fixes etc. have been necessary. The file <a href="https://zenodo.org/api/files/41faa1b7-17f6-4a75-be09-7743426ef13c/ts-nowcast-datasources-publication.zip">ts-nowcast-datasources-publication.zip</a> in this Zenodo release contains the status of the GitHub repository at the time of the publication of the paper.</p> <p>For instructions for using the data, please see the <a href="https://github.com/meteoswiss-mdr/ts-nowcast-datasources">code repository</a>.</p> <p> </p>
SMART - Self-adaptive Machine Learning Approach for Real-time Tuning of IEEE 802.11 PHY and MAC layers
<p><strong>Introduction</strong></p> <p>Worldwide the demand for wireless access networks providing very high throughputs has been increasing exponentially, namely due to bandwidth-hungry applications such as high definition video streaming and augmented reality. In order to fulfil these requirements, the Wi-Fi standard was enriched with new amendments, such as IEEE 802.11n, IEEE 802.11ac, and recently IEEE 802.11ax (Wi-Fi 6). New parameters have been proposed for both physical (PHY) and media access control (MAC) layers, including channel bonding, short guard interval (SGI), and advanced modulation and coding schemes (MCS).</p> <p>However, the high variability of the signal strength in the wireless radio channel, allied to the channel asymmetry, makes the selection of optimal configurations for these parameters a challenge. Typically, these parameters are configured with a default value. For runtime optimization, some algorithms have already been proposed. Still, they were designed considering legacy IEEE 802.11 releases and static scenarios. Besides, these parameters have their trade-offs that need to be properly managed. To help dealing with this, machine learning has been recently introduced in wireless networks, providing the intelligence that networks need in order to be smart and self-adaptive.</p> <p>SWOP (Smart Wireless Optimization) is a cross-layer optimization approach for Wi-Fi networks extending the current Rate Adaptation (RA) approach, for instance, followed by the well-known Minstrel algorithm widely used in practice. Our approach takes advantage of Deep Reinforcement Learning (DRL) in order to learn the optimal Wi-Fi link configuration. By considering the wireless channel as the environment, the transmitter node (the agent) chooses the best link parameters (the action) in order to maximize the throughput (the reward) based on the channel metrics captured from the environment (the state). In this work we propose a simple DRL-based Wi-Fi Rate Adaptation (RA) algorithm, named Data-driven Algorithm for Rate Adaptation (DARA), which is one of the modules of Smart Wireless Optimization (SWOP)</p> <p>SMART aimed to run a set of wireless experiments on top of w-iLab.t testbeds provided by the Fed4FIRE+ project to directly validate our DRL model and learn a policy from the wireless experiments executed in a controlled environment. However, after facing difficulties with the scenarios we could achieve on the real testbed, we decided to train and test DARA using a trace-based simulation approach. In simulation, we could train our model in scenarios that are more complex and diverse whilst easy to configure, when compared to real testbeds. The w-iLab.t testbeds were still used to capture data traces (e.g. Signal-to-Noise Ratio, position of nodes, transmission power and link distance) that were then injected in ns-3 for validating DARA.</p> <p>With this work, we concluded that DARA performance is impacted when operating in scenarios with asymmetric links, which is common in the highly dynamic and unpredictable wireless environments. Furthermore, the asymmetry offset varies between scenarios and it may also change for the same scenario, as time progresses. This randomness is not addressed when solely considering the SNR as the link metric, posing a challenge in the learning phase of DARA. Despite these limitations, the results obtained show that DARA still achieves up to 14.9% higher throughput higher than Minstrel [1] and slightly lower than Ideal [2] for most of the scenarios. The results obtained will serve as a basis to support our ongoing and future research.</p> <p> </p> <p><strong>Folder Organization</strong></p> <p>The following dataset presents the results of the SMART project, organized in different folders for each Rate Adaptation Algorithm, as well as the traces that were used to obtain such results:</p> <ul> <li><strong>DARA: </strong>Results obtained using our solution <strong>(Naming Convention #1, Folder Content #1)</strong></li> <li><strong>MIN: </strong>Results obtained using Minstrel-HT <strong>(Naming Convention #1, Folder Content #2)</strong></li> <li><strong>ID: </strong>Results obtained using Ideal <strong>(Naming Convention #1, Folder Content #2)</strong></li> <li><strong>TRACES: </strong>Trace files used to obtain the results present in this dataset <strong>(Naming Convention #2, Folder Content #3)</strong></li> </ul> <p><strong>Naming Convention #1 – RAA TID TP TO:</strong></p> <ul> <li>Rate Adaptation Algorithm<strong> (RAA) </strong> <ul> <li><strong>drl </strong>– Data Driven Algorithm for Rate Adaptation</li> <li><strong>min </strong>– MinstrelHTWifiManager</li> <li><strong>id </strong>– IdealWifiManager</li> </ul> </li> <li>Trace ID<strong> (TID)</strong> <ul> <li><strong>3 </strong>up to<strong> 8</strong></li> </ul> </li> <li>Transport Protocol<strong> (TP)</strong> <ul> <li><strong>udp </strong>– User Datagram Protocol</li> </ul> </li> <li>Traffic Orientation<strong> (TO)</strong> <ul> <li><strong>normal </strong>– A<strong>-></strong>B</li> <li><strong>reversed </strong>– B<strong>-></strong>A</li> </ul> </li> </ul> <p><strong>Naming Convention #2 – TID_TXP:</strong></p> <ul> <li>Trace ID<strong> (TID) </strong> <ul> <li><strong>3 </strong>up to<strong> 8</strong></li> </ul> </li> <li>Transmitting Power in dBm <strong>(TXP) </strong> <ul> <li><strong>3, 5, 7, 9, 12 dBm </strong></li> </ul> </li> </ul> <p><strong>Folder Content #1: </strong></p> <ul> <li><em>checkpoint_ RAA TID TP TO</em><strong> (Folder)</strong> <ul> <li><strong>Policy Checkpoint</strong> with which the results were obtained</li> </ul> </li> <li> <ul> <li><strong>Flowmonitor </strong>output for the configured scenario</li> </ul> </li> <li> <ul> <li>Column 1 – <strong>Step Counter</strong></li> <li>Column 2 – <strong>Reward Value</strong></li> <li>Column 3 – <strong>Observation Value</strong></li> <li>Column 4 – <strong>Action Value</strong></li> </ul> </li> <li> <ul> <li>Column 1 – <strong>Simulation Time </strong>(seconds)</li> <li>Column 2 – <strong>Throughput </strong>(Mbit/100ms)</li> </ul> </li> </ul> <p><strong>Folder Content #2: </strong></p> <ul> <li> <ul> <li><strong>Flowmonitor </strong>output for the configured scenario</li> </ul> </li> <li> <ul> <li>Column 1 – <strong>Simulation Time </strong>(seconds)</li> <li>Column 2 – <strong>Throughput </strong>(Mbit/100ms)</li> </ul> </li> </ul> <p><strong>Folder Content #3 - </strong>Source: <a href="https://zenodo.org/record/3713271#.YjjBVDXLdhE">https://zenodo.org/record/3713271#.YjjBVDXLdhE</a><strong>: </strong></p> <p>· <em>date_time</em><strong>.cfg </strong>configuration details of the experiment</p> <p>· <em>date_time_NodeID</em><a href="https://zenodo.org/record/3713271#_ftn1"><strong><em><sup>[1]</sup></em></strong></a><em>_SenderID</em><a href="https://zenodo.org/record/3713271#_ftn2"><strong><em><sup>[2]</sup></em></strong></a><em>_ReceiverID</em><a href="https://zenodo.org/record/3713271#_ftn3"><strong><em><sup>[3]</sup></em></strong></a><em>_FlowType</em><a href="https://zenodo.org/record/3713271#_ftn4"><strong><em><sup>[4]</sup></em></strong></a><em>_Params</em><a href="https://zenodo.org/record/3713271#_ftn5"><strong><em><sup>[5]</sup></em></strong></a><strong>.snr </strong>– logs of the Signal/Noise ratio (1 file per node/flow) </p> <p>· <em>date_time_NodeID_SenderID_ReceiverID_FlowType_Params</em><strong>.stats</strong> – logs of the packets received (1 file per node/flow) </p> <p><a href="https://zenodo.org/record/3713271#_ftnref1"><sub>[1]</sub></a><sub> ID of the node Logging node</sub></p> <p><a href="https://zenodo.org/record/3713271#_ftnref2"><sub>[2]</sub></a><sub> ID of the Sender node</sub></p> <p><a href="https://zenodo.org/record/3713271#_ftnref3"><sub>[3]</sub></a><sub> ID of the Receiver node</sub></p> <p><a href="https://zenodo.org/record/3713271#_ftnref4"><sub>[4]</sub></a><sub> Flow type: Unidirectional, Bidirectional or Unidirectional with Multiple Access</sub></p> <p><a href="https://zenodo.org/record/3713271#_ftnref5"><sub>[5]</sub></a><sub> Configurable parameters: Sender/Receiver Transmission Power and Data Rate (when applicable)</sub></p> <p> </p> <p><strong>References</strong></p> <p>1. F. FietKau, “Minstrel_HT: New rate control module for 802.11n [LWN.net]”. Mrt-2010.</p> <p>2. “ns-3: ns3::IdealWifiManager Class Reference,” Jan 2021, [Online; accessed 23. Jun. 2021]. Available: <a href="https://www.nsnam.org/docs/release/3.33/doxygen/classns3_1_1_ideal_wifi%20manager.html">https://www.nsnam.org/docs/release/3.33/doxygen/classns3_1_1_ideal_wifi manager.html</a></p>
First Three-dimensional Quantification of Planktic Food Chain lower levels (Copepods) for the Ross Sea region Marine Protected Area (RSRMPA), Antarctica: Using FAIR-inspired legacy data with Machine Learning, and Open Source GIS
<p>This dataset is relative to the paper entitled: "First Three-dimensional Quantification of Planktic Food Chain lower levels (Copepods) for the Ross Sea region Marine Protected Area (RSRMPA), Antarctica: Using FAIR-inspired legacy data with Machine Learning, and Open Source GIS" publishing in journal Diversity (MPDI).</p> <p>Abstract:</p> <p>Zooplankton is a fundamental group in all aquatic ecosystems located the base of the food chain. It forms a link between the lower trophic levels with secondary consumers and shows marked fluctuations of populations with environmental change, especially reacting to heating and water acidification. At sea copepod crustaceans account for app. 70% in abundance of zooplankton and are a target of monitoring activities in key areas such as the Southern Ocean. In this study we have used FAIR-inspired legacy data (dating back to the ‘80s) collected in the Ross Sea by the Italian National Antarctic Program in GBIF.org. Together with other open-access GIS data sources and tools it allows generating, for the first time, three-dimensional predictive distribution maps for twenty-six copepod species. These predictive maps were obtained by applying machine learning techniques to grey literature data, which were visualized in open-source GIS platforms. In a Species Distribution Modeling (SDM) framework we used machine learning with three types of algorithms (TreeNet, RandomForest and Ensemble) to analyze the presence and absence of copepods at different areas and depth classes in function of environmental descriptors obtained from the Polar Macroscope Layers present in Quantartica. The models allow for the first time to map-predict the food chain in quantitative terms showing the relative index of occurrence (RIO) and identified the presence for each copepod species analyzed in the Ross Sea. Our results show marked geographical preferences that vary with species and trophic strategy. This study demonstrates that machine learning is a successful method in accurately predicting Antarctic copepod presence, also providing useful data to orient future sampling and management of wildlife and conservation.</p>
Image-based & machine learning-guided multiplexed serology test for SARS-CoV-2
<p>Single-cell extracted imaging features created in project "Image-based & machine learning-guided multiplexed serology test for SARS-CoV-2". The dataset includes train (with annotations) and test features used in the manuscript. Four SARS-CoV-2 antigens (S, N, R, M) were imaged separately with serum samples presenting IgG, IgA and IgM antibodies.</p>
Improving The Diagnosis of Thyroid Cancer by Machine Learning and Clinical Data
<p>This repository contains the dataset used in the paper "Improving The Diagnosis of Thyroid Cancer by Machine Learning and Clinical Data" published in <strong>Scientific Reports</strong>. Please check our <a href="https://www.nature.com/articles/s41598-022-15342-z">formal publication</a> for the full details. The dataset contains 1232 nodules from 724 patients. Each row represents one nodule and each column represents one variable that describes the characteristics of the patient or nodule. The meaning of each variable is summarized below.</p> <ul> <li>id: the unique identity of the patient who carries the nodule</li> <li>age: the age of the patient</li> <li>FT3: triiodothyronine test result</li> <li>FT4: thyroxine test result</li> <li>TSH: thyroid-stimulating hormone test result</li> <li>TPO: thyroid peroxidase antibody test result</li> <li>TGAb: thyroglobulin antibodies test result</li> <li>site: the nodule location, 0: right, 1: left, 2: isthmus</li> <li>echo_pattern: thyroid echogenicity, 0: even, 1: uneven</li> <li>multifocality: if multiple nodules exist in one location, 0: no, 1: yes</li> <li>size: the nodule size in cm</li> <li>shape: the nodule shape, 0: regular, 1: irregular</li> <li>margin: the clarity of nodule margin, 0: clear; 1: unclear</li> <li>calcification: the nodule calcification, 0: absent, 1: present</li> <li>echo_strength: the nodule echogenicity, 0: none, 1: isoechoic, 2: medium-echogenic, 3: hyperechogenic, 4: hypoechogenic</li> <li>blood_flow: the nodule blood flow, 0: normal, 1: enriched</li> <li>composition: the nodule composition, 0: cystic, 1: mixed, 2: solid</li> <li>multilateral: if nodules occur in more than one location, 0: no, 1: yes</li> <li>mal: the nodule malignancy, 0: benign, 1: malignant</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.