Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
132
datasets available to search
ShareScore release 0.7.1
Dataset results
132 results for “Random Forests”
Data for "Random forest-based modeling of stream nutrients at national level in a data-scarce region"
<p>The aim of the study was to model annual total nitrogen (TN) and total phosphorus (TP) concentrations at national level using an ML approach. We used water quality data originating from the Environmental Monitoring Database KESE to train RF models for nutrient concentration prediction in 242 catchments across Estonia. A total of 82 environmental variables were used as predictors in the models. In order to yield the best results, a feature selection strategy along with hyperparameter optimization was performed when building the models. The models are applicable for predicting nutrient loads on an annual level, e.g. for the purpose of reporting national level water quality statistics in regional projects, such as HELCOM. The results showed that this relatively basic RF modeling approach can have a performance similar to process-based models. Moreover, these models are easier to reuse and apply on a larger scale, since the required inputs can be derived from freely available datasets (e.g. satellite imagery)</p> <p>This repository contains the input data used for building the RF models and the files describing the modeling results.</p> <p>The description of the files is given in the README.txt file.</p> <p>Virro, H., Kmoch, A., Vainu, M. and Uuemaa, E., 2022. Random forest-based modeling of stream nutrients at national level in a data-scarce region. Science of The Total Environment, 840, p.156613.</p> <p><a href="https://doi.org/10.1016/j.scitotenv.2022.156613">https://doi.org/10.1016/j.scitotenv.2022.156613</a></p>
Dataset - Identification of early abandonment in cropland through radar-based coherence data and application of a Random-Forest model
<p>This dataset accompanies the manuscript titled "Identification of early abandonment in cropland through radar-based coherence data and application of a Random-Forest model", submitted by co-authors to the journal Global Change Biology (GCB) Bioenergy.</p> <p>Wouter Meijninger<sup>1</sup>, Berien Elbersen<sup>1</sup>, Michiel van Eupen<sup>1</sup>, Stephan Mantel<sup>2</sup>, Pilar Ciria Ciria<sup>3</sup>, Andrea Parenti<sup>4</sup>, Marina Sanz Gallego<sup>3</sup> and Paloma Perez Ortiz<sup>3</sup>, Marco Acciai<sup>4</sup>,and Andrea Monti<sup>4</sup><br> Institutes: 1) Wageningen University & Research, 2) ISRIC, 3) CIEMAT, 4) Bologna University,</p> <p><strong>Abstract (Manuscript)</strong></p> <p>In the context of increased pressures on land for food and non-food production it is relevant to understand better, which land resources have become unused and abandoned and where these lands are. Data on where these lands are and what their extend is are not collected in regular statistics. In this paper we present an approach to detect signs of abandonment in cropping land using radar coherence data. The methodology was tested in the Spanish regions of Albacete and Soria where agricultural land abandonment is a common process. The results show that land abandonment detection using radar coherence data works well for the region of Albacete in arable lands. The radar-based analysis is a relatively simple method to detect land abandonment in an early to longer-term state and can therefore be applied once developed and tested further in other regions to larger areas of the EU where land abandonment is serious and needs monitoring and policy response. The applicability of the method to Soria and Emilia Romagna (Italy) regions show that there are still challenges to overcome to make the method more widely applicable for detecting land abandonment in other environmental zones of Europe. Lack of reliable training and validation data, like LPIS data, in regions is one of the challenges in this respect.</p> <p><strong>Readme data files</strong></p> <p><em>Coherence_quarterly_statisitcs_2017_to_2020.zip</em></p> <p>Radar coherence quarterly statistics - Albacete (Spain)</p> <p>Radar coherence data is based on Sentinel-1B<br> Period: 2017 to 2020</p> <p>File naming (.tif files) per year (<em>YYYY</em>):</p> <ul> <li>Mean coherence: <em>mean_YYYY_1to4.tif</em></li> <li>Standard deviation coherence: <em>std_YYYY_1to4.tif</em></li> <li>Range coherence: <em>range_YYYY_1to4.tif</em></li> <li>Mean delta coherence: <em>mean_delta_YYYY_1to4.tif</em></li> <li>Standard deviation delta coherence: <em>std_delta_YYYY_1to4.tif</em></li> <li>Maximum delta coherence: <em>max_delta_YYYY_1to4.tif</em></li> </ul> <p>Each file consists of 4 bands:</p> <ul> <li>band 1: 1st quarter [Jan-Feb-March]</li> <li>band 2: 2nd quarter [April-May-June]</li> <li>band 3: 3rd quarter [July-Aug-Sept]</li> <li>band 4: 4th quarter [Oct-Nov-Dec]</li> </ul> <p>Statistics are based on radar coherence data, which is scaled between >0 and 1<br> No data: 0-values</p> <p>Projection:<br> EPSG:32630 - WGS 84 / UTM zone 30N<br> Pixel size: 20m</p> <p><em>SIGPAC_data_Albacete_2018_to_2020.zip</em></p> <ul> <li>More than 5 year fallow (20m raster files)</li> <li>Land Use Land Cover LULC (20m raster files)</li> </ul> <p>More than 5 year fallow (according to SIGPAC)<br> Period: 2018 to 2020<br> File naming (ENVI files):</p> <ul> <li>Albacete_SIGPAC_MoreThan5YrsFallowAreas_2018_20m.dat (+ Albacete_SIGPAC_MoreThan5YrsFallowAreas_2018_20m.hdr)</li> <li>Albacete_SIGPAC_MoreThan5YrsFallowAreas_2019_20m.dat (+ Albacete_SIGPAC_MoreThan5YrsFallowAreas_2019_20m.hdr)</li> <li>Albacete_SIGPAC_MoreThan5YrsFallowAreas_2020_20m.dat (+ Albacete_SIGPAC_MoreThan5YrsFallowAreas_2020_20m.hdr)</li> </ul> <p>Pixel values:<br> 0: Not fallow<br> 1: Fallow more than 5 years</p> <p>Projection:<br> EPSG:32630 - WGS 84 / UTM zone 30N<br> Pixel size: 20m</p> <p>Land Use Land Cover LULC (according to SIGPAC)<br> Period: 2018 to 2020<br> File naming (ENVI files):</p> <ul> <li>LULC_SIGPAC_Albacete_2018_20m.dat (+ LULC_SIGPAC_Albacete_2018_20m.hdr)</li> <li>LULC_SIGPAC_Albacete_2019_20m.dat (+ LULC_SIGPAC_Albacete_2019_20m.hdr)</li> <li>LULC_SIGPAC_Albacete_2020_20m.dat (+ LULC_SIGPAC_Albacete_2020_20m.hdr)</li> </ul> <p>Pixel values:</p> <ul> <li>0 - Nan</li> <li>1 - Arable land</li> <li>2 - Vineyards</li> <li>3 - Olives</li> <li>4 - Fruits</li> <li>5 - Nuts</li> <li>6 - Citrus</li> <li>7 - Permanent grassland</li> <li>8 - Forest</li> <li>9 - Rest, small elements</li> <li>10 - Built-up areas</li> <li>11 - Water</li> <li>12 - Roads</li> <li>13 - Unproductive land</li> </ul> <p>Projection:<br> EPSG:32630 - WGS 84 / UTM zone 30N<br> Pixel size: 20m</p> <p><em>Annual_unused_used_land_maps_Albacete_2017_to_2020.zip</em></p> <p>Derived annual unused/used land maps - Albacete (Spain), based on Random-Forest model<br> Period: 2017-2020<br> File naming (ENVI files):</p> <ul> <li>predict_RF_Albacete_2017_quarterly_stats_LU1_v181920.dat (+ predict_RF_Albacete_2017_quarterly_stats_LU1_v181920.hdr)</li> <li>predict_RF_Albacete_2018_quarterly_stats_LU1_v181920.dat (+ predict_RF_Albacete_2018_quarterly_stats_LU1_v181920.hdr)</li> <li>predict_RF_Albacete_2019_quarterly_stats_LU1_v181920.dat (+ predict_RF_Albacete_2019_quarterly_stats_LU1_v181920.hdr)</li> <li>predict_RF_Albacete_2020_quarterly_stats_LU1_v181920.dat (+ predict_RF_Albacete_2020_quarterly_stats_LU1_v181920.hdr)</li> </ul> <p>Pixel values:<br> 0 - Used (and/or Nan)<br> 1 - Unused</p> <p>Projection:<br> EPSG:32630 - WGS 84 / UTM zone 30N<br> Pixel size: 20m</p> <p><em>Four_year_abandoned_land_Albacete_2017_to_2020.zip</em></p> <p>Four-year abandonment map is based on the 4 annual unused/used land maps<br> File naming (ENVI):</p> <ul> <li>Four_year_abandoned_land_Albacete_2017_to_2020.dat (+ Four_year_abandoned_land_Albacete_2017_to_2020.hdr)</li> </ul> <p>Pixel values:<br> 0 - (Nan)<br> 1 - Used (1 year unused in period 2017 - 2020)<br> 2 - Used (2 year unused in a row in period 2017 - 2020)<br> 3 - Abandoned (3 year unused in a row in period 2017 - 2020)<br> 4 - Abandoned (4 year unused in a row in period 2017 - 2020)</p> <p>Projection:<br> EPSG:32630 - WGS 84 / UTM zone 30N<br> Pixel size: 20m</p>
Sentiment Analysis of RUU PDP with Naive Bayes, Support Vector Machine, and Random Forest Classification Algorithm
<p>Dataset from the results of data crawling via Twitter which discusses the Rancangan Undang Undang Pelindungan Data Pribadi to be used in the sentiment analysis process. The dataset is divided into several parts according to the process executed on RapidMiner.</p>
Random forest climatic modeling of agricultural insurance loss across the inland Pacific Northwest region of the United States
<p>We compared climatic relationships to insurance loss across the inland Pacific Northwest region of the United States, using a design matrix methodology, to identify optimum temporal windows for climate variables by county in relationship to wheat insurance loss due to drought. The results of our temporal window construction for water availability variables (precipitation, temperature, evapotranspiration, and the Palmer drought severity index [PDSI]) identified spatial patterns across the study area that aligned with regional climate patterns, particularly with regards to drought-prone counties of eastern Washington. Using these optimum time-lagged correlational relationships between insurance loss and individual climate variables, along with commodity pricing, we constructed a regression-based random forest model for insurance loss prediction and evaluation of climatic feature importance. Our cross-validated model results indicated that PDSI was the most important factor in predicting total seasonal wheat/drought insurance loss, with wheat pricing and potential evapotranspiration having noted contributions. Our overall regional model had a R<sup>2</sup> of 0.49 and a RMSE of $30.8 million. Model performance typically underestimated annual losses, with moderate spatial variability in terms of performance between counties.</p>
A Methodology for the Fast Identification and Monitoring of Microplastics in Environmental Samples using Random Decision Forest Classifiers
<p>This short video shows the results of the application of a classifier for microplastics as described by Hufnagl et al. (2019).</p> <p> </p> <p>If you reuse this video please cite</p> <p> </p> <p>Hufnagl, B., Steiner, D., Renner, Löder, M. G. J., Laforsch, C. and Lohninger, H. <em>A Methodology for the Fast Identification and Monitoring of Microplastics in</em><em> Environmental Samples using Random Decision Forest Classifiers,</em> Analytical Methods, 2019, DOI:10.1039/C9AY00252A</p>
Identification of high-wind features within extratropical cyclones using a probabilistic random forest - Part 2: Climatology - Video Supplement
<p>These videos provide examples of application of RAMEFI (RAndom-forest based MEsoscale wind Feature Identification) for winter storm cases between 2000 and 2019 and an animation of cyclone-relative occurrence (cf. Fig. 7 of 10.5194/wcd-2023-10).</p>
Soil chemistry dataset from the work "Modelling and prediction of major soil chemical properties with Random Forest: machine learning as tool to understand soil-environment relationships in Antarctica"
<p>Bases sum, H+Al (potential acidity), pH, phosphorous, remaining P (P-rem), sodium and total organic carbon distribution in Antarctic soils modeled and predicted through Machine Learning approaches, legacy soil data and environmental covariates. The quantile and prediction interval data represent the spatial uncertainty of the predictions.</p> <p>As soon as the work "Modelling and prediction of major soil chemical properties with Random Forest: machine learning as tool to understand soil-environment relationships in Antarctica" is published, the paper will be cited here. </p> <p>The .zip file contains the following folders:</p> <p>1) soil_chemistry_antarctica: data containing the soil chemical attributes distribution</p> <p>2) soil_chemistry_prediction_interval: uncertainty from the prediction interval 90% (Q95% - Q5%) of the soil attributes prediction</p> <p>4) soil_texture_quantile05: quantile 5% of the soil attributes prediction</p> <p>5) soil_texture_quantile95: quantile 95% of the soil attributes prediction</p>
Random forest climatic modeling of agricultural insurance loss across the inland Pacific Northwest region of the United States
Open the record for dataset details and reuse information.
Global soil saturated hydraulic conductivity map using random forest in a Covariate-based GeoTransfer Functions (CoGTF) framework at 1 km resolution
<p>The global Ksat map at 1 km resolution was developed by harnessing the technological advances in machine learning and availability of remotely sensed surrogate information such as terrain, climate, vegetation, and soil covariates. We merge concepts of predictive soil mapping with a large data set of Ksat measurements and local information (soil, vegetation, climate) into covariate-based “Geo Transfer Functions'' (CoGTFs) to generate global estimates of Ksat values (to highlight the impact of Geo-referenced covariates including various remote sensing maps, we use the term Geotransfer function GTF and not pedotransfer function PTF; in the latter case, typically only soil properties are used to estimate Ksat).</p> <p>The Ksat dataset is provided in GeoTIFF format. A total of 4 files that represent different soil depths (0, 30, 60, and 100 cm) are provided. The Ksat values are log-transformed (log10 Ksat) and cm/day was selected as a standardized unit.</p> <p>The Global Ksat training dataset used for this study is available here:<br> <a href="https://doi.org/10.5281/zenodo.3752721">https://doi.org/10.5281/zenodo.3752721</a></p> <p>The R code used for this study is available here:<br> <a href="https://github.com/ETHZ-repositories/Ksat_mapping_2020">https://github.com/ETHZ-repositories/Ksat_mapping_2020</a></p> <p>For more details / to cite this dataset please use:</p> <ul> <li>Gupta, S., Lehmann, P., Bonetti, S., Papritz, A., and Or, D., (2020): <strong>Global prediction of soil saturated hydraulic conductivity using random forest in a Covariate-based Geo Transfer Functions (CoGTF) framework</strong>. Journal of Advances in Modeling Earth Systems,<strong> </strong>13(4), e2020MS002242. https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2020MS002242</li> </ul> <p>Other datasets related to this project:</p> <p>The Global vG training dataset is available here:</p> <p><a href="https://doi.org/10.5281/zenodo.5547338">10.5281/zenodo.5547338</a></p> <p>Examples of using this dataset to generate van Genuchten parameters maps can be found in <a href="https://doi.org/10.5281/zenodo.6343570">10.5281/zenodo.6343570</a>.</p> <p>The study was supported by ETH Zurich (Grant ETH-18 18-1). We would like to thank Zhongwang Wei, Samuel Bickel and Simone Fatichi (ETH Zurich) for insightful discussions.</p> <p> </p> <p> </p>
Three‐dimensional reconstruction of porous polymer films from FIB‐SEM nanotomography data using random forests
<p>Dataset and code used in M. Röding, et al, "Three-dimensional reconstruction of porous polymer films from FIB-SEM nanotomography data using random forests", published in Journal of Microscopy, 2020. In this work, we develop a segmentation method for focused ion beam scanning electron microscopy (FIB-SEM) data acquired by volumetric imaging of porous polymer films made from ethyl cellulose and hydroxypropyl cellulose (EC/HPC) polymer blends. This type of polymer films are used for controlled release applications. Based on manual segmentation of a fraction of the data, a random forest classifier is trained and applied to the full data set. Here, raw data, manual segmentations, and the Matlab code used for all steps in the analysis are supplied.</p>
Seafloor Density Measurements, Prediction, and Associated Uncertainty for "Predicting global marine sediment density using the random forest regressor machine learning algorithm"
<p>Global seafloor density prediction results using the random forest regressor machine learning algorithm. </p> <p>Dataset S1. Seafloor density measurements. Columns are labeled with a header and include associated drilling project and measurement type for each sample. File format: CSV text file</p> <p>Dataset S2. Seafloor density prediction results from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p> <p>Dataset S3. Seafloor density prediction standard deviation from the random forest regressor machine learning algorithm at 5×5-arc minute resolution. Units are g/cm^3. File format: netCDF (.nc)</p>
Random forests for predicting species identity of forensically important blow flies (Diptera: Calliphoridae) and flesh flies (Sarcophagidae) using geometric morphometric data: proof of concept
<p>Wing shape variation has been shown to be useful for delineating forensically important fly species in two Diptera families: Calliphoridae and Sarcophagidae. Compared to DNA-based identification, the cost of geometric morphometric data acquisition and analysis is relatively much lower because the tools required are basic, and stable softwares are available. However, to date, an explicit demonstration of using wing geometric morphometric data for species identity prediction in these two families remains lacking. Here, geometric morphometric data from 19 homologous landmarks on the left wing of males from seven species of Calliphoridae (<em>n</em>=55), and eight species of Sarcophagidae (<em>n</em>=40) were obtained and processed using Generalized Procrustes Analysis. Allometric effect was removed by regressing centroid size (in log10) against the Procrustes coordinates. Subsequently, principal component analysis of the allometry-adjusted Procrustes variables was done, with the first 15 principal components used to train a random forests model for species prediction. Using a real test sample consisting of 33 male fly specimens collected around a human corpse at a crime scene, the estimated percentage of concordance between species identities predicted using the random forests model and those inferred using DNA-based identification was about 80.6% (approximate 95% confidence interval = [68.9%, 92.2%]). In contrast, baseline concordance using naive majority class prediction was 36.4%. The results provide proof of concept that geometric morphometric data has good potential to complement morphological and DNA-based identification of blow flies and flesh flies in forensic work. </p>
Data from: Demographic model selection using random forests and the site frequency spectrum
Phylogeographic data sets have grown from tens to thousands of loci in recent years, but extant statistical methods do not take full advantage of these large data sets. For example, approximate Bayesian computation (ABC) is a commonly used method for the explicit comparison of alternate demographic histories, but it is limited by the "curse of dimensionality" and issues related to the simulation and summarization of data when applied to next-generation sequencing (NGS) data sets. We implement here several improvements to overcome these difficulties. We use a Random Forest (RF) classifier for model selection to circumvent the curse of dimensionality and apply a binned representation of the multidimensional site frequency spectrum (mSFS) to address issues related to the simulation and summarization of large SNP data sets. We evaluate the performance of these improvements using simulation and find low overall error rates (~7%). We then apply the approach to data from Haplotrema vancouverense, a land snail endemic to the Pacific Northwest of North America. Fifteen demographic models were compared, and our results support a model of recent dispersal from coastal to inland rainforests. Our results demonstrate that binning is an effective strategy for the construction of a mSFS and imply that the statistical power of RF when applied to demographic model selection is at least comparable to traditional ABC algorithms. Importantly, by combining these strategies, large sets of models with differing numbers of populations can be evaluated.
Figure 2 in Diagnosability of mtDNA with Random Forests: Using sequence data to delimit subspecies
Figure 2. Illustration of steps in constructing a Random Forests ensemble of classification trees.
Trained Random Forest Model and Scaler Parameters for (Phy+Man), 17 Feb, 2025
<p>This repository contains the trained models and scaler parameters. There are two trained random forest models. These models were trained on 5000 traces/class, The traces were of 40s (P-10, P+30), 110s (P-10, P+100) and 150s (P-50, P+100) bandpass filtered between 0.5-15 Hz, and resampled to 50 Hz. <br><br>The new version (17/02/2025) of the models were trained on 6000 traces per class. </p>
Random Forest model for PPI predictions
<p>The uploaded model contains the trained random forest model to predict protein-protein Interactions based on their amino acid sequence. The model is described in our preprint "ProteinPrompt: a webserver for predicting protein-protein interactions" which can be found on bioRxiv.org: https://doi.org/10.1101/2021.09.03.458859</p> <p>The model was trained with the scikit-learn package in version 0.20.3 under Python 3.7.12</p>
MetaComNet: A random forest-based framework for making spatial prediction of plant-pollinator interactions
<p>1. Predicting plant-pollinator interaction networks over space and time will improve our understanding of how environmental change is likely to impact the functioning of ecosystems. Here we propose a framework for producing spatially explicit predictions of the occurrence and number of pairwise plant-pollinator interactions and of the species richness, diversity, and abundance of pollinators visiting flowers. We call the framework 'MetaComNet' because it aims to link metacommunity dynamics to the assembly of ecological networks.</p> <p>2. To illustrate the MetaComNet functionality, we used a dataset on bee-flower networks sampled at 16 sites in southeast Norway along with random forest models to predict bee-flower interactions. We included variables associated with climatic conditions (elevation) and habitat availability within a 250m radius of each site. Regional commonness, site-specific distance to conspecifics, social guild, and floral preference were included as bee traits. Each plant species was assigned a score reflecting its site-specific abundance, and four scores reflecting the bee species that the plant family is known to attract. We used leave-one-out cross-validations to assess the models' ability to predict pairwise plant-bee interactions across the landscape.</p> <p>3. The relationship between observed occurrence or absence of interactions and the predicted probability of interactions was nearly proportional (GLMlogistic regression slope = 1.09), matching the data well (AUC = 0.88), and explained 30% of the variation. Predicted probability of interactions was also correlated with the number of observed pairwise interactions (r = 0.32). The sum of predicted probabilities of bee-flower interactions were positively correlated with observed species richness (r = 0.50), diversity (r = 0.48), and abundance (r = 0.42) of wild bees interacting with plant species within sites.</p> <p>4. Our findings show that the MetaComNet framework can be a useful approach for making spatially explicit predictions and mapping plant-pollinator interactions. Such predictions have the potential to identify areas where the pollination potential for wild plants is particularly high, and where conservation action should be directed to preserve this ecosystem function.</p>
Global maps of soil water characteristics parameters developed using the random forest in a Covariate-based GeoTransfer Functions (CoGTF) framework at 1 km resolution
<p>The global soil water characteristics parameters (<em>α, n, θ<sub>r</sub>, and θ<sub>s</sub></em>) maps based on van Genuchten (vG) model at 1 km resolution was developed by harnessing the technological advances in machine learning and availability of remotely sensed surrogate information such as terrain, climate, vegetation, and soil covariates. We merge concepts of predictive soil mapping with a large data set of vG parameters and local information (soil, vegetation, climate) into "Covariate-based GeoTransfer Functions'' (CoGTFs) to generate global estimates of vG parameters (to highlight the impact of Geo-referenced covariates including various remote sensing maps, we use the term Geotransfer Function GTF and not pedotransfer function PTF; in the latter case, typically only soil properties are used to predict vG parameters).</p> <p>The vG parameters (<em>α, n, θ<sub>r</sub>, and θ<sub>s</sub></em>) dataset is provided in GeoTIFF format. A total of 16 files that represent different soil depths (0, 30, 60, and 100 cm) are provided for each parameter.</p> <table> <caption>Description of vG parameters and their units</caption> <tbody> <tr> <td>vG Parameters</td> <td>Description</td> <td>units</td> </tr> <tr> <td><em>α</em></td> <td>Inverse air entry pressure</td> <td>Log<sub>10</sub><em>α (m<sup>-1</sup>)</em></td> </tr> <tr> <td><em>n</em></td> <td>Shape parameter</td> <td>Log<sub>10</sub><em>n </em>(dimensionless)</td> </tr> <tr> <td><em>θ<sub>r</sub></em></td> <td>Residual water content</td> <td>m<sup>3</sup>/m<sup>3</sup></td> </tr> <tr> <td><em>θ<sub>s</sub></em></td> <td>Saturated water content</td> <td>m<sup>3</sup>/m<sup>3</sup></td> </tr> </tbody> </table> <p> </p> <p>The Global vG training dataset used for this study is available here:</p> <p><a href="https://doi.org/10.5281/zenodo.5547338">10.5281/zenodo.5547338</a></p> <p>For more details / to cite this dataset please use:</p> <ul> <li>Gupta, S., Papritz, A., Lehmann, P., Hengl, T., Bonetti, S., & Or, D. (2022). Global Mapping of Soil Water Characteristics Parameters—Fusing Curated Data with Machine Learning and Environmental Covariates. <strong><em>Remote Sensing</em></strong>, <em>14</em>(8), 1947.</li> </ul> <p>The study was supported by ETH Zurich (Grant ETH-18 18-1). We thank Zhongwang Wei, Associate professor at Sun Yat-Sen University, for helping to collect the datasets and for insightful discussions. We would like to thank Andrea Carmintai, Professor at ETH Zurich, for the insightful discussions.</p>
IIb-RAD-seq coupled with random forest classification indicates regional population structuring and sex-specific differentiation in salmon lice (Lepeophtheirus salmonis)
<p><span>The aquaculture industry has been dealing with salmon lice problems forming serious threats to salmonid farming. Several treatment approaches have been used to control the parasite. Treatment effectiveness must be optimized, and the systematic genetic differences between sub-populations must be studied to monitor louse species and enhance targeted control measures. We have used IIb-RAD sequencing in tandem with a random forest classification algorithm to detect the regional genetic structure of the Norwegian salmon lice and identify important markers for sex differentiation of this species. We identified 19428 single nucleotide polymorphisms (SNPs) from 95 individuals of salmon lice. These SNPs, however, were not able to distinguish differential structure of lice populations. Using the random forest algorithm, we selected 91 SNPs important for geographical classification and 14 SNPs important for sex classification. The geographically important SNP data substantially improved the genetic understanding of the population structure and classified regional demographic clusters along the Norwegian coast. </span><span>We also uncovered SNP markers that could help determine the sex of the salmon louse. </span><span>A large portion of the SNPs identified to be under directional selection were also ranked highly important by random forest. According to our findings, there is a regional population structure of salmon lice associated with the geographical location along the Norwegian coastline.</span></p>
Objective identification of high-wind features within extratropical cyclones using a probabilistic random forest (RAMEFI). Part I: Method and illustrative case studies - Video Supplement
<p>These videos provide examples of application of RAMEFI (RAndom-forest based MEsoscale wind Feature Identification), a new objective identification of high-wind features within extratropical cyclones, for twelve selected case studies.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.