Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
TCOM-CH4: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric methane profile dataset [1991-2021] constructed using machine-learning
<p>Methodology: </p> <p><span>he </span><strong><span>TOMCAT simulation</span></strong><span> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized </span><strong><span>ERA-5 reanalysis data</span></strong><span>.</span></p> <h3><span>CH4 Profile Processing and Bias Correction</span></h3> <p><strong><span>Collocated CH4 profiles</span></strong><span> are organized into five distinct latitude bins:</span></p> <ul> <li> <p><strong><span>NH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>NH mid-lat</span></strong><span>: </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>Tropics</span></strong><span>: </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>SH mid-lat</span></strong><span>: </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> <li> <p><strong><span>SH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> </ul> <p><span>Initially, </span><strong><span>differences between TOMCAT and satellite measurements</span></strong><span> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from </span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>). It is important to note that unlike previous versions that might have used both HALOE and ACE measurements, this version exclusively utilizes </span><strong><span>ACE-FTS data</span></strong><span>, which is why the dataset starts from 2000.</span></p> <p><strong><span>Separate XGBoost regression models</span></strong><span> are then trained for these CH4 differences at each height level within a given latitude bin. These trained models are subsequently used to estimate </span><strong><span>CH4 bias corrections</span></strong><span> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</span></p> <p><strong><span>Height-resolved CH4 profile data</span></strong><span> are then interpolated onto 28 standard pressure levels (from </span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</span></p> <h3><span>Data Files</span></h3> <p><span>The dataset includes two files containing daily mean zonal mean CH4 profiles:</span></p> <ul> <li> <p><code><span>zmch4_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>height level data</span></strong><span> (</span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> </li> <li> <p><code><span>zmch4_TCOM_plev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>pressure level data</span></strong><span> (</span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>).</span></p> </li> </ul> <h3><span>Reference Publication</span></h3> <p><span>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</span></p> <p><span>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105–5120, </span><a title="null" href="https://doi.org/10.5194/essd-15-5105-2023"><span>https://doi.org/10.5194/essd-15-5105-2023</span></a><span>, 2023.</span></p>
TCOM-N2O: TOMCAT CTM and Occultation Measurements based daily zonal stratospheric nitrous oxide profile dataset [1991-2021] constructed using machine-learning
<p>Methodology: </p> <p><span>The </span><strong><span>TOMCAT simulation</span></strong><span> was conducted at a T64L32 resolution, consistent with previous work by Dhomse et al. (2021, 2022), covering the period from 2000 to 2024. These simulations utilized </span><strong><span>ERA-5 reanalysis data</span></strong><span>.</span></p> <h3><span>N2O Profile Processing and Bias Correction</span></h3> <p><strong><span>Collocated N2O profiles</span></strong><span> are organized into five distinct latitude bins:</span></p> <ul> <li> <p><strong><span>NH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>NH mid-lat</span></strong><span>: </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N - </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>Tropics</span></strong><span>: </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>4</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>N</span></p> </li> <li> <p><strong><span>SH mid-lat</span></strong><span>: </span><span><span><span><span><span>7</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>2</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> <li> <p><strong><span>SH polar</span></strong><span>: </span><span><span><span><span><span>9</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S - </span><span><span><span><span><span>5</span><span>0<span><span><span><span><span><span><span>∘</span></span></span></span></span></span></span></span></span></span></span></span><span>S</span></p> </li> </ul> <p><span>Initially, </span><strong><span>differences between TOMCAT and satellite measurements</span></strong><span> (primarily ACE-FTS data) are calculated for each zonal bin across 51 height levels (ranging from </span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> <p><strong><span>Separate XGBoost regression models</span></strong><span> are then trained for these N2O differences at each height level within a given latitude bin. These trained models are subsequently used to estimate </span><strong><span>N2O bias corrections</span></strong><span> for all daytime TOMCAT grids (9132 days), specifically sampled at 1:30 PM local time at the equator. This yields grid-specific bias corrections that are applied to the original TOMCAT profiles.</span></p> <p><strong><span>Height-resolved N2O profile data</span></strong><span> are then interpolated onto 28 standard pressure levels (from </span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>), using pressure levels directly from the TOMCAT grids. For overlapping latitude bins, values are averaged to ensure smoother fields near boundary regions.</span></p> <h3><span>Data Files</span></h3> <p><span>The dataset includes two files containing daily mean zonal mean N2O profiles:</span></p> <ul> <li> <p><code><span>zmn2o_TCOM_hlev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>height level data</span></strong><span> (</span><span><span><span><span><span>10</span><span>,</span><span><span>km</span></span></span></span></span></span><span> to </span><span><span><span><span><span>60</span><span>,</span><span><span>km</span></span></span></span></span></span><span>).</span></p> </li> <li> <p><code><span>zmn2o_TCOM_plev_T2Dz_2000-2024_V1.1.nc</span></code><span>: Contains </span><strong><span>pressure level data</span></strong><span> (</span><span><span><span><span><span>300</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span> to </span><span><span><span><span><span>0.1</span><span>,</span><span><span>hPa</span></span></span></span></span></span><span>).</span></p> </li> </ul> <h3><span>Reference Publication</span></h3> <p><span>This methodology, incorporating only ACE-FTS data and various minor algorithmic developments, is based on the following publication:</span></p> <p><span>Dhomse, S. S. and Chipperfield, M. P.: Using machine learning to construct TOMCAT model and occultation measurement-based stratospheric methane (TCOM-CH4) and nitrous oxide (TCOM-N2O) profile data sets, Earth Syst. Sci. Data, 15, 5105–5120, </span><a title="null" href="https://doi.org/10.5194/essd-15-5105-2023"><span>https://doi.org/10.5194/essd-15-5105-2023</span></a><span>, 2023.</span></p>
Supplementary material for the publication: "Efficient Surrogate Models for Materials Science Simulations: Machine Learning-based Prediction of Microstructure Properties"
<p><span><span><span>This dataset contains supplementary code, images and models for the publication „Efficient Surrogate Models for Materials Science Simulations: Machine Learning-based Prediction of Microstructure Properties“.</span></span></span></p> <p> </p> <p><span><span><span>The content will be updated and additionally linked to the corresponding git repositories.</span></span></span></p>
EUNIS Habitat Maps: Enhancing Thematic and Spatial Resolution for Europe through Machine Learning
<p>The EUNIS habitat classification is essential for categorising European habitats and supporting European policy on nature conservation and to implement the Nature Restoration Law. As such, to meet the growing demand for detailed and accurate habitat information, we provide spatial predictions for 260+ EUNIS habitat types at EUNIS level 3, together with validation and uncertainty analyses. </p> <p>More specifically, using ensemble machine learning models together with high-resolution satellite imagery and other climatic, terrain and soil variables, we produced an European habitat map at a 100-m resolution indicating the most likely EUNIS habitat at level 3 for every location across Europe. Predictions were validated for three independent countries, namely for France, the Netherlands and Austria. We also provide information on uncertainty and the most probable habitats at level 3 within each EUNIS level 1 formation. Products can be further refined with accurate and local land cover data. This product is thus likely to be particularly useful for restoration but also conservation purposes. </p> <p>Figure: <strong>Wall-to-wall map of EUNIS habitats at level 3 - (color coded at level 2 for visibility)</strong></p> <p></p>
Human pan-body age- and sex-specific molecular phenomena inferred from public transcriptome data using machine learning - Data
<p>Expression data used in manuscript <i>Human pan-body age- and sex-specific molecular phenomena inferred from public transcriptome data using machine learning</i></p>
A Novel Approach to Impact Crater Mapping and Analysis on Enceladus, using Machine Learning: Supplemental data
<p>This dataset includes the crater map and equatorial crater depths and diameters presented in the paper: A Novel Approach to Impact Crater Mapping and Analysis on Enceladus, using Machine Learning.</p>
Radiomics and machine learning analysis by computed tomography and magnetic resonance imaging in colorectal liver metastases prognostic assessment
<p>We uploaded the raw data related to extracted features of the manuscript "Granata V, Fusco R, De Muzio F, Brunese MC, Setola SV, Ottaiano A, Cardone C, Avallone A, Patrone R, Pradella S, Miele V, Tatangelo F, Cutolo C, Maggialetti N, Caruso D, Izzo F, Petrillo A. Radiomics and machine learning analysis by computed tomography and magnetic resonance imaging in colorectal liver metastases prognostic assessment. Radiol Med. 2023 Nov;128(11):1310-1332. doi: 10.1007/s11547-023-01710-w. Epub 2023 Sep 11. PMID: 37697033."</p>
Machine Learning-Based Bridge Maintenance Optimization Model for Maximizing Performance within Available Annual Budgets
<p>Effective maintenance planning for bridges is crucial for maintaining their performance, safety, and minimizing maintenance costs. Timely implementation of interventions can improve the performance of bridges and avoid the need for costly interventions. However, bridge maintenance is often delayed due to inadequate planning and budget allocation, as well as resource constraints such as funding. With availability of historical condition data of bridges in databases such as the National Bridge Inventory (NBI) and National Bridge Elements (NBE), there is an opportunity to use data-driven methods to predict deterioration of bridge elements and optimize their maintenance interventions to maximize performance of bridges. This paper presents the development of a novel system that uses Machine Learning (ML) techniques to predict condition of concrete bridge elements and binary linear programming optimization method to identify the optimal selection of maintenance interventions and their timing to maximize the performance of bridges while complying with available annual budgets. Four ML methods are explored: decision tree, random forest, gradient boosting, and support vector machines. The results of the ML evaluation show that, while the values of the predictive performance metrics varied for different elements, random forest method had the best performance for all elements. A case study of a concrete bridge is analyzed to evaluate the performance of the system and demonstrate its new capabilities. The case study results show that the developed model identifies optimal maintenance interventions for various annual budgets over a 50-year study period. The primary contributions of this research to the body of knowledge are: (1) development of a novel system that integrates machine learning techniques and linear programming for predicting bridge element conditions and optimizing maintenance interventions; (2) modeling and predicting the deterioration of bridge elements based on health index metric; and (3) generating long-term maintenance plans for each of bridge elements to maximize the performance of bridges within available annual budgets. The present system is expected to support decision makers, such as highway agencies, in allocating limited financial resources for bridge maintenance more efficiently and cost-effectively.</p>
Gauging Size Resolved Ambient Particulate Matter Concentration Solely Using Biometric Observations: A Machine Learning and Causal Approach
<p>Notebook and data to accompany the (unpublished) paper titled "Gauging Size Resolved Ambient Particulate Matter Concentration Solely Using Biometric Observations: A Machine Learning and Causal Approach". This work expands a previous study, relating particulate matter concentrations and short-term biometric features across multiple participants. </p><p>Github link: https://github.com/mi3nts/DUEDARE_multiple_participants</p>
Global lithospheric thickness reconstruction using machine learning
<p><span>The lithosphere, as the outermost solid layer of our planet, preserves a progressively more fragmentary record of geological events and processes from Earth's history the further back in time one looks. Thus, the evolution of lithospheric thickness and its cascading impacts on Earth's tectonic system are presently unknown. Herein, we track the lithospheric thickness history using machine learning based on lithogeochemical big data of basalt. Our results demonstrate that four dramatic lithospheric thinning events occurred during the Paleoarchean, early Paleoproterozoic, Neoproterozoic, and Phanerozoic with intermediate thickening scenarios. These events respectively correspond to supercontinent breakup and assembly periods. Causality investigation further indicates that crustal metamorphic and deformation styles are the feedback of lithospheric thickness. Cross-correlation between lithospheric thickness and metamorphic thermal gradients record the transition from intra-oceanic subduction systems to continental margin plus intra-oceanic in the Paleoarchean and Mesoarchean, and a progressive emergence of large thick continents that allow supercontinent growth, which promoted assembly of the first supercontinent during the Neoarchean.</span></p>
Code and Source Data for "Knowledge-Guided Machine Learning can improve C cycle quantification in agroecosystems"
<p>Datasets for code and Source Data for the study "Knowledge-Guided Machine Learning can improve C cycle quantification in agroecosystems" https://doi.org/10.1038/s41467-023-43860-5. All files belong to Licheng Liu and Zhenong Jin at University of Minnesota. deposit_code_v2.zip contains packaged codes and sample runs for KGML-ag-Carbon training, validation and implementations. Source Data.zip contains data for generating the figures inside the study. </p> <p>Note: We used Pytorch 1.6.0 (<a href="https://pytorch.org/get-started/previous-versions/">https://pytorch.org/get-started/previous-versions/</a>, last access: 21 Oct 2023) and Python 3.7.11 (<a href="https://www.python.org/downloads/release/python-3711/">https://www.python.org/downloads/release/python-3711/</a>, last access: 21 Oct 2023) as the programming environment for model development. Statistical analysis, such as linear regression, was conducted using Statsmodels 0.14.0 (<a href="https://github.com/statsmodels/statsmodels/">https://github.com/statsmodels/statsmodels/</a>, last access: 21 Oct 2023) In order to use a GPU to speed-up the training process, we installed the CUDA Toolkit 10.1.243 (<a href="https://developer.nvidia.com/cuda-toolkit">https://developer.nvidia.com/cuda-toolkit</a>, last access: 21 Oct 2023). </p> <p><strong>To use the full kgml_lib function, please create a new environment with the same python and libs above.</strong></p>
Research data supporting: "Machine learning of microscopic structure-dynamics relationships in complex molecular systems"
<p>This repository contains the set of data and the code to reproduce the results shown in "Machine learning of microscopic structure-dynamics relationships in complex molecular systems" published on Machine Learning: Science and Technology (DOI: 10.1088/2632-2153/ad0fa5).</p>
Diversity-aware Fairness Testing of Machine Learning Classifiers through Hashing-based Sampling
<p>The experimental results of the evaluation of VBT-X.</p> <h2>Abstract</h2> <div> <h3>Context:</h3> <p>There are growing concerns about algorithmic fairness, as some machine learning (ML)-based algorithms have been found to exhibit biases against protected attributes such as gender, race, age and so on. Individual fairness requires an ML classifier to produce similar outputs for similar individuals. Verification Based Testing (<span>Vbt</span>) is a state-of-the-art black-box testing algorithm for individual fairness that leverages constraint solving to generate test cases.</p> </div> <div> <h3>Objective:</h3> <p>Generating diverse test cases is expected to facilitate efficient detection of diverse discriminatory data instances (i. e., cases that violate individual fairness). Hashing-based sampling techniques draw a sample approximately uniformly at random from the set of solutions of given Boolean constraints. We propose <span>Vbt</span>-X, which improves <span>Vbt</span> with hashing-based sampling, aiming to improve its testing performance.</p> </div> <div> <h3>Method:</h3> <p>We realize hashing-based sampling for <span>Vbt</span>. The challenge is that the off-the-shelf hashing-based sampling techniques cannot be integrated in a straightforward manner because the constraints in <span>Vbt</span> are generally not Boolean. Moreover, we propose several enhancement techniques to make <span>Vbt</span>-X more efficient.</p> </div> <div> <h3>Results:</h3> <p>To evaluate our method, we conduct experiments, where <span>Vbt</span>-X is compared to <span>Vbt</span>, <span>Sg</span> and ExpGA (other well-known fairness testing algorithms) over a set of configurations consisting of several datasets, protected attributes, and ML classifiers. The results show that, with each configuration, <span>Vbt</span>-X detects more discriminatory data instances with higher diversity than <span>Vbt</span> and <span>Sg</span>. <span>Vbt</span>-X detects discriminatory data instances with higher diversity than ExpGA, though the number of discriminatory data instances detected by <span>Vbt</span>-X is lesser than ExpGA.</p> </div> <div> <h3>Conclusion:</h3> <p>Our proposed method performs better than other state-of-the-art black-box fairness testing algorithms, particularly in terms of diversity. Our method can serve to efficiently identify flaws in ML classifiers with respect to individual fairness for subsequent improvements of an ML classifier. On the other hand, although our method is specific to individual fairness, it could work for testing other aspects of a software system such as security and counterfactual explanations with some technical adaptations, which remains for future work.</p> </div> <p> </p> <div> <h2>Acknowledgments</h2> <p>This paper is partly based on results obtained from a project, JPNP20006, commissioned by the New Energy and Industrial Technology Development Organization (NEDO). This paper is supported by JST SPRING, Grant Number JPMJSP2131.</p> </div>
WaivOps EDM-TR9: Open Audio Resources for Machine Learning in Music
<p><strong>EDM-TR9 Dataset</strong></p> <p>EDM-TR9 is an open audio dataset composed of a series of drum recordings in the style of electronic dance music (EDM). This dataset primarily focuses on the distinctive sounds and rhythm patterns of the Roland TR-909 drum machine within the subgenres of dance, house and techno music. The dataset contains 3780 audio loops recorded in uncompressed stereo WAV format, produced with custom drum samples and MIDI-programmed rhythms at various tempo rates.</p> <p><strong>Dataset</strong></p> <p>The primary objective of this dataset is to provide accessible content for machine learning applications in music and audio research. Some potential use cases for this dataset include tempo detection and classification, drum rhythm analysis, audio-to-MIDI conversion, source separation, automated mixing, music information retrieval, AI music generation, sound design and signal processing.</p> <p><strong>Specifications</strong></p> <ul> <li>3780 audio loops (approximately 8 hours)</li> <li>24-bit WAV format</li> <li>BPM labeled</li> <li>Tempo range: 120–140bpm</li> <li>Variational drum patterns</li> <li>EDM drum rhythms</li> </ul> <p><strong>License</strong></p> <p>This dataset was compiled by WaivOps, a crowdsourced music project managed by the sound label company Patchbanks. All recordings have been compiled by verified sources for copyright clearance.</p> <p>The EDM-TR9 dataset is licensed under Creative Commons Attribution 4.0 International <a href="https://creativecommons.org/licenses/by/4.0/">(CC BY 4.0)</a>.</p> <p><strong>Additional Info</strong></p> <p>For audio examples or more information about this dataset, please refer to the <a href="https://github.com/patchbanks/WaivOps-EDM-TR9">GitHub repository</a>.</p>
Machine learning reveals that climate, geography, and cultural drift all predict bird song variation in coastal Zonotrichia leucophrys
<p>Previous work has demonstrated that there is extensive variation in the songs of White-crowned Sparrow (<em>Zonotrichia leucophrys</em>) throughout the species range, including between neighboring (and genetically distinct) subspecies <em>Z. l. nuttalli </em>and <em>Z. l. pugetensis</em>. Using a machine learning approach to bioacoustic analysis, we demonstrate that variation in song is correlated with year of recording (representing cultural drift), geographic distance, and climatic differences, but the response is subspecies- and season-specific. Automated machine learning methods of bird song annotation can process large datasets more efficiently, allowing us to examine 1,913 recordings across ~60 years. We utilize a recently published artificial neural network to automatically annotate White-crowned Sparrow vocalizations. By analyzing differences in syllable usage and composition, we recapitulate the known pattern where <em>Z. l. nuttalli </em>and <em>Z. l. pugetensis </em>have significantly different songs. Our results are consistent with the interpretation that these differences are caused by the changes in characteristics of syllables in the White-crowned Sparrow repertoire. This supports the hypothesis that the evolution of vocalization behavior is affected by the environment, in addition to population structure.</p>
Towards metadata for machine learning - Crosswalk tables
<p><strong>Crosswalks for Machine Learning models and datasets used for training</strong></p> <p>Here we present a collection of crosswalks for ML models (in TSV and XLSX formats) and datasets used for training (in TSV and XLSX formats). These crosswalks were created during an <a href="https://www.nfdi4datascience.de/">NFDI4DataScience</a> hackathon organized by the Semantic Technologies team (SemTec) at<a href="https://www.zbmed.de/en/"> ZB MED Information Centre for Life Sciences (ZB MED)</a> with the aim of provindg a starting point for a common proposal towards a metadata schema for ML models based on <a href="http://schema.org">schema.org</a>.</p> <p><strong>Files</strong></p> <ul> <li>2023.11.23 Metadata for ML - ML dataset Union.tsv: Crosswalks for datasets in TSV format</li> <li>2023.11.23 Metadata for ML - ML dataset Union.xlsx: Crosswalks for datasets in XLSX format</li> <li>2023.11.23 Metadata for ML - ML model union.tsv: Crosswalks for ML models in TSV format</li> <li>2023.11.23 Metadata for ML - ML model union.xlsx: Crosswalks for ML models in XLSX format</li> </ul>
Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models
<p><span>Knowledge of individual age can help both in-situ and ex-situ conservation programs to design more efficient and suitable management plans for targeted wildlife species. DNA methylation is one of the epigenetic aging markers that has emerged as a promising tool that can estimate age with high accuracy using only a tiny amount of biological material, which can be collected in a minimally invasive way. Here, we sequenced five targeted genetic regions and used </span><span>8–23</span><span> selected CpG sites to build age estimation models with machine learning methods </span><span>with about only $3–7 per sample</span><span>, using blood samples of seven Felidae species—ranging from small to big, and domestic to endangered species: domestic cats (<em>Felis catus</em>, 139 samples), Tsushima leopard cats (<em>Prionailurus bengalensis euptilurus</em>, 84 samples), and five<em> Panthera </em>species (96 samples). </span><span>The models built achieved satisfactory accuracy—the mean absolute error of the best models was 1.966, 1.348, and 1.552 years in domestic cats, Tsushima leopard cats, and <em>Panthera</em> spp., respectively.</span><span> Our models in domestic cats and Tsushima leopard cats were applicable to individuals regardless of health conditions, indicating the high applicability of our models to samples collected from diverse situations, e.g., rescued individuals in the context of conservation. We also showed the possibility of developing universal age estimation models for the five<em> Panthera</em> spp. using two of the five genetic regions, suggesting an even lower cost to use our models for future applications.</span></p>
Machine learning links T cell function and spatial localization to neoadjuvant immunotherapy and clinical outcome in pancreatic cancer
<p>Data supporting the findings of "<a href="https://doi.org/10.1158/2326-6066.CIR-23-0873" target="_blank" rel="noopener">Machine learning links T cell function and spatial localization to neoadjuvant immunotherapy and clinical outcome in pancreatic cancer</a>" publication. Files include patient and tissue region metadata (in metadata folder) and output of multiplex immunohistochemistry computational image processing workflow for each tissue region (in mIHC_files folder). The code used to produce the results of this study is available at: <a href="https://github.com/kblise/PDAC_mIHC_paper">https://github.com/kblise/PDAC_mIHC_paper</a>.</p>
Images supporting: Nondestructive, quantitative viability analysis of 3D tissue cultures using machine learning image segmentation
<p>Two image datasets (as zip files) including all images analyzed in the manuscript Nondestructive, quantitative viability analysis of 3D tissue cultures using machine learning image segmentation. Images are of pancreatic adenocarcinoma (PDAC) cystic spheroid samples grown in either BME or Matrigel. Some images have background noise in the form of iron oxide nanoparticles introduced to them.</p>
Supplementary material from: Prediction of the Cold Flow Properties of Biodiesel using the FAME Distribution and Machine Learning Techniques
<p><span>The dataset is divided into three sections within the worksheet.</span></p> <p><span> </span><span>The first section contains the definition of the data's feedstock and its source reference. The reference includes the year, DOI (if available, as some are collected from books), publication journal, article title, and authors.</span></p> <p><span> </span><span>The second section describes the FAME distribution, starting from C4:0 up to C24:0, including a column of unidentified FAMEs.</span></p> <p><span><span>The third and final section describes the measured properties Cloud Point (CP), Cold Filter Plugging Point (CFPP) and Pour Point (PP).</span></span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.