Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,943
datasets available to search
ShareScore release 0.9.0
Dataset results
1,943 results for “machine learning”
Machine learning algorithm evaluation
<p>This is the data management plan for the purpose of this report, to compare three different classifiers through supervised machine learning on two diverse datasets. The whole machine learning process was applied and conducted in different experiments. The exploration of the data sets as well as the preprocessing strategies are outlined in the following. Furthermore, the modelling processes and the performance measures on which their results are evaluated will be explained. Finally, different parameter adjustments and settings are compared and discussed which leads to a conclusion.</p>
Molecular similarity perception based on machine-learning models
<p>Molecular similarity is an particularly important notion for chemical legislation, specifically in the evaluation process for orphan drugs (i.e., drugs for rare diseases). A new molecule needs to be dissimilar from any other existing drug for a given disease to be assigned the financially advantageous status of orphan drug. Currently, there are many ways to define whether two molecules are similar or dissimilar. Thus far, the European Medicines Agency has used experts majority voting on discretional judgments of similarity when assessing new drugs for rare diseases. The decision of individual expert whether two compounds are similar is inherently subjective, depending on factors such as gender, age, state of mind, and previous experiences. It is therefore desirable, in this context, to benefit from an objective measure of similarity. To answer this need, we report a new dataset of molecular similarity assessments, that includes complex and difficult similarity scenarios. As a result, we propose new and improved models for similarity-prediction procedures, including 3D properties. These models are publicly available: <a href="https://chematlas.chimie.unistra.fr/ReadySim/">https://chematlas.chimie.unistra.fr/ReadySim/</a>.</p> <p>Software, 3D structures and pictures are available in the git related to this deposit: <a href="https://github.com/enricogandini/paper_similarity_prediction.git">https://github.com/enricogandini/paper_similarity_prediction.git</a></p> <p>The deposit contains two files.</p> <ul> <li>original_training_set.csv: this is one of the dataset published initially in [doi: 10.1186/1758-2946-6-5].</li> <li>new_dataset.csv: result from a new survey organized in 2020</li> </ul> <p>The columns are the following:</p> <ul> <li>id_pair: unique identifier of the compound pair</li> <li>curated_smiles_molecule_a: first compound of the pair</li> <li>curated_smiles_molecule_b: second compound of the pair</li> <li>tanimoto_cdk_Extended: ECFP similarity measure</li> <li>TanimotoCombo: ComboScore similarity measure</li> <li>pchembl_distance: difference of activity of the compound pair</li> <li>target_name: protein to which the compound pair is binding</li> <li>simil_2D: similar based on ECFP (0 or 1)</li> <li>simil_3D: similar based on ComboScore (0 or 1)</li> <li>dissimil_2D: dissimilar based on ECFP (0 or 1)</li> <li>dissimil_3D: dissimilar based on ComboScore (0 or 1)</li> <li>pair_type: pairs are classified based on ECFP and ComboScore as similar or dissimilar in 2D and 3D - Sim2DSim3D, Sim2DDis3D, Dis2D,Sim3D, Dis2DSim3D</li> <li>n_answers: number of answers from experts</li> <li>n_similar: number of answers labeling the pair as similar compounds</li> <li>frac_similar: n_similar/n_answers</li> </ul>
Prediction of runoff characteristics in ungauged basins in Central Europe with machine learning – files
<p><em><strong>English</strong></em></p> <p>This are the shapefiles accompanying the paper: Klingler et al. (2022), Prediction of runoff characteristics in ungauged basins with machine learning, published in the journal Österreichische Wasser- und Abfallwirtschaft: <a href="https://doi.org/10.1007/s00506-022-00891-4">https://doi.org/10.1007/s00506-022-00891-4</a></p> <p>The basic idea was to train a machine learning model with observed runoff characteristics of the hydrological years 2003 - 2017 (LamaH_observations, 859 features) and 90 different catchment characteristics to be able to predict runoff characteristics in unobserved catchments (OWK_predictions, 9533 features).</p> <p>We provide two shapefiles to download:<br> <strong>1) LamaH_observations</strong>, which contains attributes for 6 different runoff characteristics calculated from observed runoff timeseries from the LamaH-CE dataset (https://doi.org/10.5194/essd-13-4529-2021).<br> <strong>2)</strong> <strong>OWK_predictions</strong>, which includes additionally to the predicted 6 runoff characteristics also attributes for uncertainty quantification.<br> All attributes of the shapefiles are described in the associated metadata (.qmd files).</p> <p><strong>Disclaimer:</strong> We have created the shapefiles with care and checked the outputs for plausibility. By downloading the data, you agree that we nor the provider of the used source datasets (e.g. observed runoff time series) cannot be liable for the data provided.</p> <p><strong>License:</strong> This work is licensed with CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/). This means that you may freely use and modify the data (even for commercial purposes). But you have to give appropriate credit (associated ÖWAV paper, version of dataset), indicate if and what changes were made and distribute your work under the same public license as the original.</p> <p><strong>Contact:</strong> If you find any errors in the dataset or have any further questions, feel free to send us an email: info@baseflow.ai</p> <p>-------------</p> <p><strong><em>Deutsch</em></strong></p> <p>Dies sind die beiden Shapefiles, welche dem folgenden Fachartikel zugehörig sind: Klingler et al. (2022), Vorhersage von hydrologischen Abflusskennwerten in unbeobachteten Einzugsgebieten mit Machine Learning, veröffentlicht im Journal Österreichische Wasser- und Abfallwirtschaft: <a href="https://doi.org/10.1007/s00506-022-00891-4">https://doi.org/10.1007/s00506-022-00891-4</a></p> <p>Der Ansatz hinter dieser Arbeit war ein Machine Learning Modell mit beobachteten Abflusskennwerten der hydrologischen Jahre 2003 - 2017 (LamaH_observations, 859 Features) und 90 verschiedenen Einzugsgebietseigenschaften zu trainieren, um anschließend diese Abflusskennwerte in unbeobachteten Einzugsgebieten vorherzusagen (OWK_predictions, 9533 Features).</p> <p>Wir bieten zwei Shapefiles zum Download an:<br> <strong>1)</strong> <strong>LamaH_observations</strong>, welche Attribute für 6 verschiedene Abflusskennwerten (MJHQ, MQ, MJNQ, MJNQ7, Q95, Q98) enthält, die aus beobachteten Abflusszeitreihen aus dem LamaH-CE-Datensatz berechnet wurden (https://doi.org/10.5194/essd-13-4529-2021).<br> <strong>2)</strong> <strong>OWK_predictions</strong>, welche zusätzlich zu den vorhergesagten 6 Abflusskennwerten auch Attribute zur Quantifizierung der Unsicherheit enthält.<br> Alle Attribute der Shapefiles sind in den zugehörigen Metadaten (.qmd-Dateien) beschrieben.</p> <p><strong>Haftungsausschluss:</strong> Wir haben die Shapefiles mit Sorgfalt erstellt und die Ergebnisse auf Plausibilität geprüft. Mit dem Herunterladen der Daten erklären Sie sich damit einverstanden, dass weder wir noch der Anbieter der verwendeten Quelldatensätze (zB. beobachtete Abflusszeitreihen) für die bereitgestellten Daten haften.</p> <p><strong>Lizenz:</strong> Diese Arbeit ist lizenziert mit CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/). Dies bedeutet, dass Sie die Daten frei verwenden und verändern dürfen (auch für kommerzielle Zwecke). Sie müssen jedoch eine entsprechende Quellenangabe machen (zugehöriger ÖWAV-Artikel, Version des Datensatzes), angeben ob und welche Änderungen vorgenommen wurden, und Ihre Arbeit unter der gleichen Lizenz wie das Original veröffentlichen.</p> <p><strong>Kontakt:</strong> Wenn Sie Fehler im Datensatz finden oder weitere Fragen haben, können Sie uns gerne eine E-Mail schicken: info@baseflow.ai</p>
Weather data (forecast and observation) at three locations in France over 2021 for Machine Learning Training
<p>The data provided data are historical weather measurement and forecast at three location in France.</p> <p>Measurements are inside files named OBS_xxx</p> <p>Forecasts are inside files names YYY_xxx, with YYY is the name of the forecast simultion (GFS0.25, WRF12km or WRF3KM).</p> <p>In the two cases, xxx is the name of the site (Site 1, Site2 or Site3).</p> <p><br> <strong>Description of OBS_xxx files:</strong><br> - One line per measurement with hourly resolution<br> - columns are: Date(TU),Temperature2m_degC,WindSpeed10m_m/s,WindDirection10m_m/s<br> Date = date of measurement in TU and format DD/MM/YYYY HH:MM<br> Temperature2m_degC = air temperature at 2m height in °Celsius<br> WindSpeed10m_m/s = wind speed at 10m height in m/s<br> WindDirection10m_deg = wind direction at 10m height in deg. (0 or 360 = wind from north to south, 45°=wind from east to east, ....)<br> If measurement is not available for a specific hour for one parameter, the value "-999" is used.</p> <p>The observation data go:<br> from 16/04/2021 00H <br> to 31/01/2022 23H</p> <p><br> <strong>Description of YYY_xxx files:</strong><br> - One line per forecast with hourly resolution<br> - columns are: First date run (TU),forecast hour,Temperature2m_degC,WindSpeed10m_m/s,WindDirection10m_m/s<br> First date run (TU) = date of start of the forecast in TU and format DD/MM/YYYY HH:MM. HH could be 00 and 12 according to the cycle of forecast start.<br> forecast hour = forecast hour from the start of the forecast date. 00 = forecast for "first date run". 01 = forecast for "First date run" + 1 hour. .... 95 = forecast for "First date run" + 95 hours.<br> For GFS0.25, forecast hour go from 00 to 95<br> For WRF12km, forecast hour go from 00 to 95<br> For WRF3m, forecast hour go from 00 to 95<br> Temperature2m_degC = air temperature at 2m height in °Celsius<br> WindSpeed10m_m/s = wind speed at 10m height in m/s<br> WindDirection10m_deg = wind direction at 10m height in deg. (0 or 360 = wind from north to south, 45°=wind from east to east, ....)<br> If measurement is not available for a specific hour for one parameter, the value "-999" is used.</p> <p>The forecast data go:<br> from 13/04/2021 00H + 72H = first forecast for the 16/04/2021 00H<br> to 31/01/2022 12H + 11H = last forecast for the 31/01/2022 23H</p>
Machine Learning Assisted SSH Keys Extraction From The Heap Dump
<p>This dataset contains heap dump of OpenSSH that contains session keys.</p> <p>On the performance test data, we also include the PCAP file that contains the encrypted SSH network traffic. With the correct session keys, it can be decrypted.</p>
ImageNet-Patch: A Dataset for Benchmarking Machine Learning Robustness against Adversarial Patches
<p>Adversarial patches are optimized contiguous pixel blocks in an input image that cause a machine-learning model to misclassify it. However, their optimization is computationally demanding and requires careful hyperparameter tuning. To overcome these issues, we propose ImageNet-Patch, a dataset to benchmark machine-learning models against adversarial patches. It consists of a set of patches optimized to generalize across different models and applied to ImageNet data after preprocessing them with affine transformations. This process enables an approximate yet faster robustness evaluation, leveraging the transferability of adversarial perturbations.</p> <p>We release our dataset as a set of folders indicating the patch target label (e.g., `banana`), each containing 1000 subfolders as the ImageNet output classes.</p> <p>An example showing how to use the dataset is shown below.</p> <pre><code class="language-python"># code for testing robustness of a model import os.path from torchvision import datasets, transforms, models import torch.utils.data class ImageFolderWithEmptyDirs(datasets.ImageFolder): """ This is required for handling empty folders from the ImageFolder Class. """ def find_classes(self, directory): classes = sorted(entry.name for entry in os.scandir(directory) if entry.is_dir()) if not classes: raise FileNotFoundError(f"Couldn't find any class folder in {directory}.") class_to_idx = {cls_name: i for i, cls_name in enumerate(classes) if len(os.listdir(os.path.join(directory, cls_name))) > 0} return classes, class_to_idx # extract and unzip the dataset, then write top folder here dataset_folder = 'data/ImageNet-Patch' available_labels = { 487: 'cellular telephone', 513: 'cornet', 546: 'electric guitar', 585: 'hair spray', 804: 'soap dispenser', 806: 'sock', 878: 'typewriter keyboard', 923: 'plate', 954: 'banana', 968: 'cup' } # select folder with specific target target_label = 954 dataset_folder = os.path.join(dataset_folder, str(target_label)) normalizer = transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]) transforms = transforms.Compose([ transforms.ToTensor(), normalizer ]) dataset = ImageFolderWithEmptyDirs(dataset_folder, transform=transforms) model = models.resnet50(pretrained=True) loader = torch.utils.data.DataLoader(dataset, shuffle=True, batch_size=5) model.eval() batches = 10 correct, attack_success, total = 0, 0, 0 for batch_idx, (images, labels) in enumerate(loader): if batch_idx == batches: break pred = model(images).argmax(dim=1) correct += (pred == labels).sum() attack_success += sum(pred == target_label) total += pred.shape[0] accuracy = correct / total attack_sr = attack_success / total print("Robust Accuracy: ", accuracy) print("Attack Success: ", attack_sr) </code></pre> <p> </p>
DATASET - Improving Remote Sensing of Extreme Events with Machine Learning: Application to IASI LST Retrievals
<p>Data for experiments presented in the paper "Improving Remote Sensing of Extreme Events with Machine Learning: Application to IASI LST Retrievals" </p>
Machine learning the quantum flux-flux correlation function for catalytic surface reactions
<p>This dataset contains information on each of the 14 reactions used in the paper, the geometries for these reactions, the product of the quantum reaction rate constant and canonical reactant partition function and the flux-flux correlation function time series values for each reaction-temperature combination.</p> <p><strong>reaction_details.csv</strong></p> <p>This is a .csv file containing additional details on the reactions used in this paper. Each row contains one reaction/temperature combination, of which there are 55.</p> <p> </p> <p>Column descriptions:</p> <ul> <li>reaction_number: Reaction identifier number used in this work</li> <li>reaction: The chemical reaction equation</li> <li>metal_surface: atomic symbol of metal surface</li> <li>facet_number: Miller indices of surface</li> <li>reactants: Python dictionary object of reactants and their quantities</li> <li>products: Python dictionary object of products and their quantities</li> <li>reaction_energy [eV]: reaction energy in electron-volts</li> <li>activation_energy [eV]: activation energy of reaction in electron-volts</li> <li>temperature [K]: The randomly assigned temperature a calculation was run for</li> <li>kQ_Cff [1/au]: The calculated integrated reaction rate product at corresponding temperature {1,2,3,4} in units 1/(au time).</li> <li>reaction_split: Train/test placement of that reaction/temperature combination for reaction split</li> <li>temperature_split:<strong> </strong>Trian/test placement of that reaction/temperature combination for temperature split</li> <li>catalysishub_reactionID: Catalysis Hub reaction ID identifier for referencing catalysis hub database</li> <li>doi:<strong> </strong>digital object identifier of original publication for which DFT calculations were performed</li> </ul> <p> </p> <p> </p> <p><strong>Flux_flux_correlation_functions:</strong></p> <p>Directory containing flux-flux correlation function time series values for each reaction temperature combination. Values are organized in subdirectories, one for each of the 14 reaction. In each subdirectory .csv files are labeled by reaction number and temperature in Kelvin. Each csv file contains a column with time points [au of time] and the corresponding flux-flux correlation function value in units [1/(au of time)<sup>2</sup>].</p> <p> </p> <p><strong>Geometries:</strong></p> <p>Directory containing geometry files for each reaction. Geometries of reactants on the surface were shifted respect to those supplied by catalysis hub to create continuous reaction pathways where necessary. Geometry files are organized in subdirectories for each reaction. When complete nudged elastic band (NEB) minimum energy paths (MEP) were not available ,subdirectories contain a products.xyz, reactants.xyz, and TSstar.xyz file (reactions 1 to 11) otherwise the complete set of NEB MEP images labeled neb{n}.xyz is given (reactions 12, 13, 14).</p> <p> </p> <p> </p>
Replication Package for "Evaluating the layout quality of UML class diagrams using machine learning"
<p>Open Science material including dataset and replication instructions accompanying the article "Evaluating the layout quality of UML class diagrams using machine learning."</p>
Gregory-MS: list of research papers relevant/not relevant for Multiple Sclerosis research (for machine learning training)
<p>This dataset represents the list of research papers' data that was used to train and test different machine learning algorithms in the Gregory-MS project. The list includes the title and abstract (when available) of the research papers and an annotated field (relevant) that specifies if the given research paper is relevant or not for multiple sclerosis research.</p>
Global input datasets for use in constraints on global seafloor biogenic methane production from deterministic and machine learning modeling
<p>This dataset includes 9 grids used as model input for manuscript "Constraints on global seafloor biogenic methane production from deterministic and machine learning modeling". Additionally, there are four grids (heat flow, total organic carbon, porosity, and crust age) for which variable uncertainty was given.</p> <p>Grids here are available in xyz (longitude in decimal degrees, latitude in decimal degrees, and variable) ascii file format. Each reference is below is the grids native reference. For more information on the creation of these grids please visit the main manuscript.</p> <p>Below are respective file names and variable name/units:</p> <p>Dataset 1: Elevation in Meters (+ indicates above sea level, - below sea level)</p> <p>Tozer, B., Sandwell, D. T., Smith, W. H. F., Olson, C., Beale, J. R., & Wessel, P. (2019). Global bathymetry and topography at 15 arc sec: SRTM15+. <em>Earth and Space Science</em>, 6. https://doi.org/10.1029/ 2019EA000658</p> <p>Dataset 2: Seawater Density in Kilograms per Cubic Meter</p> <p>Boyer, T. P., Antonov, J. I., Baranova, O. K., Garcia, H. E., Johnson, D. R., Mishonov, A. V., … Grodsky, A. (2013). World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), Technical Ed.; <em>NOAA Atlas NESDIS</em> 72 (pp. 209).</p> <p>Dataset 3: Seawater Temperature in Degrees Celcius </p> <p>Boyer, T. P., Antonov, J. I., Baranova, O. K., Garcia, H. E., Johnson, D. R., Mishonov, A. V., … Grodsky, A. (2013). World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), Technical Ed.; <em>NOAA Atlas NESDIS</em> 72 (pp. 209).</p> <p>Dataset 4: Seawater Salinity in Percent Salinity Units</p> <p>Boyer, T. P., Antonov, J. I., Baranova, O. K., Garcia, H. E., Johnson, D. R., Mishonov, A. V., … Grodsky, A. (2013). World Ocean Database 2013. In S. Levitus, A. Mishonov (Ed.), Technical Ed.; <em>NOAA Atlas NESDIS</em> 72 (pp. 209).</p> <p>Dataset 5: Heat Flow in Milliwatts per Square Meter</p> <p>Global Heat Flow Compilation Group (2013). Component parts of the World Heat Flow Data Collection. <em>PANGAEA</em>, https://doi.org/10.1594/PANGAEA.810104</p> <p>Hornbach, M. J., Harris, R. N. & Phrampus, B. J. (2020). Heat flow on the U.S. Beaufort Margin, Arctic Ocean: Implications for ocean warming, methane hydrate stability, and regional tectonics. <em>Geochemistry, Geophysics, Geosystems</em>, 21(5). e2020GC008933. https://doi.org/10.1029/2020GC008933</p> <p>Dataset 6: Sediment Thickness in Meters</p> <p>Straume, E. O., Gaina, C., Medvedev, S., Hochmuth, K., Gohl, K., Whittaker, J. M., … Hopper, J. R. (2019). GlobSed: updated total sediment thickness in the world’s oceans. <em>Geochemistry, Geophysics, Geosystems</em>, 20(4), 1756–1772.</p> <p>Dataset 7: Seafloor Porosity in Fraction</p> <p>Martin, K. M., Wood, W. T., & Becker, J. J. (2015). A global prediction of seafloor sediment porosity using machine learning. <em>Geophysical Research Letters</em>, 42(24), 2015GL065279. https://doi.org/10.1002/2015GL065279</p> <p>Dataset 8: Seafloor Total Organic Carbon in Percent Dry Weight</p> <p>Lee, T.R., Wood, W.T., & Phrampus, B.J. (2019). A machine learning (kNN) approach to predicting global seafloor total organic carbon. <em>Global Biogeochemical Cycles</em>. 33, 37–46, doi:10.1029/2018GB005992.</p> <p>Dataset 9: Crust Age in Million Years</p> <p>Müller, R. D., Sdrolias, M., Gaina, C., & Roest, W. R. (2008). Age, spreading rates, and spreading asymmetry of the world’s ocean crust. <em>Geochemistry, Geophysics, Geosystems</em>, 9, Q04006. https://doi.org/10.1029/2007GC001743</p> <p>Dataset 10: Seafloor Porosity Uncertainty in Fraction</p> <p>Dataset 11: Seafloor Total Organic Carbon Uncertainty in Percent Dry Weight</p> <p>Lee, T.R., Wood, W.T., & Phrampus, B.J. (2019). A machine learning (kNN) approach to predicting global seafloor total organic carbon. <em>Global Biogeochemical Cycles</em>. 33, 37–46, doi:10.1029/2018GB005992.</p> <p>Dataset 12: Heat Flow Uncertainty in Milliwatts per Square Meter</p> <p>Dataset 13: Crust Age Uncertainty in Million Years</p> <p>Müller, R. D., Sdrolias, M., Gaina, C., & Roest, W. R. (2008). Age, spreading rates, and spreading asymmetry of the world’s ocean crust. <em>Geochemistry, Geophysics, Geosystems</em>, 9, Q04006. https://doi.org/10.1029/2007GC001743</p>
Machine Learning Models and New Computational Tool for the Discovery of Insect Repellents that Interfere with Olfaction
<ul> <li><strong>SI1_Supporting Information</strong> file (docx) brings together detailed information on the outstanding models obtained for each dataset analyzed in this study such as statistical and training parameters and outliers. There can be found the responses in spikes/s of the mosquito <em>Culex quinquefasciatus </em>to the 50 IRs. Besides, there is presented a full table of the up-to-date studies related to QSAR and insect repellency.</li> <li><strong>SI2_EXP1_50IRs from Liu et al (2013)</strong> SDF file presents the structures of each of the 50 IRs analyzed.</li> <li><strong>SI3_EXP2_Datasets</strong> gathers the four datasets as SDF files from Oliferenko <em>et al.</em> (2013), Gaudin<em> et al. </em>(2008), Omolo <em>et al.</em> (2004), and Paluch <em>et al.</em> (2009) used for the repellency modeling in <strong>EXP2</strong>.</li> <li><strong>SI4_EXP3_Prospective analysis </strong>provides Malaria Box Library (400 compounds) as an SDF file, which were analyzed in our virtual screening to prospect potential virtual hits.</li> <li><strong>SI5_QuBiLS-MIDAS MDs lists</strong> contain three TXT lists of 3D molecular descriptors used in QuBiLS-MIDAS to describe the molecules used in the present study.</li> <li><strong>SI6_EXP1_Sensillar Modeling</strong> comprises two subfolders: Classification and Regression models for each of the six sensilla. Models built to predict the physiological interaction experimentally obtained from Liu <em>et al.</em> (2013). All of the models are implemented in the software SiLiS-PAPACS. Every single folder compiles a DOCX file with the detailed description of the model, an XLSX file with the output obtained from the training in Weka 3.9.4, an ARFF, and CSV files with the MDs for each molecule, and the SDF of the study dataset.</li> <li><strong>SI7_EXP2_Repellency Modeling </strong>encompasses the four datasets in the study: Oliferenko <em>et al.</em> (2013), Gaudin<em> et al. </em>(2008), Omolo <em>et al.</em> (2004), and Paluch <em>et al.</em> (2009). Inside the subfolders, there are three models per type of MDs (duplex, triple, generic, and mix) selected that best predict each dataset. As well as the SI6 folder, each model includes six files: DOCX, XLSX, ARFF, CSV, and an SDF.</li> <li><strong>SI8_Virtual Hits </strong>includes the cluster analysis results and physico-chemical properties of new IR virtual leads.</li> </ul>
eDNAssay: a machine learning tool that accurately predicts qPCR cross-amplification
<p>Environmental DNA (eDNA) sampling is a highly sensitive and cost-effective technique for wildlife monitoring, notably through the use of qPCR assays. However, it can be difficult to ensure assay specificity when many closely related species cooccur. In theory, specificity may be assessed in silico by determining whether assay oligonucleotides have enough base-pair mismatches with nontarget sequences to preclude amplification. However, the mismatch qualities required are poorly understood, making in silico assessments difficult and often necessitating extensive in vitro testing—typically the greatest bottleneck in assay development. Increasing the accuracy of in silico assessments would therefore streamline the assay development process. In this study, we paired 10 qPCR assays with 82 synthetic gene fragments for 530 specificity tests using SYBR Green intercalating dye (n = 262) and TaqMan hydrolysis probes (n = 268). Test results were used to train random forest classifiers to predict amplification. The primer-only model (SYBR Green-based) and full-assay model (TaqMan probe-based) were 99.6% and 100% accurate, respectively, in cross-validation. We further assessed model performance using six independent assays not used in model training. In these tests the primer-only model was 92.4% accurate (n = 119) and the full-assay model was 96.5% accurate (n = 144). The high performance achieved by these models makes it possible for eDNA practitioners to more quickly and confidently develop assays specific to the intended target. Practitioners can access the full-assay model via eDNAssay (https://NationalGenomicsCenter.shinyapps.io/eDNAssay), a user-friendly online tool for predicting qPCR cross-amplification.</p>
Dataset for Integrated hydrodynamic and machine learning models
<p>The dataset is the supplement to our publication in <a href="https://www.nonlinear-processes-in-geophysics.net/">Nonlinear Processes in Geophysics</a> (https://doi.org/10.5194/npg-2021-36). To use this data, please give us credit by citing our article.</p>
Data from: Modeling the drivers of eutrophication in Finland with a machine learning approach
<p>The dataset contains data on characteristics of 1547 Finnish EU Water Framework Directive monitored lakes and their catchments from years 2016-2019.</p> <p> </p> <p><strong>Usage notes</strong></p> <p>The zip-file contains catchment and lake characteristics data to support the main analysis code in Github (link) (‘Analyses and visualization’) which produces results, figures and tables presented in the study ‘Modeling the drivers of eutrophication in Finland with a machine learning approach’. Detailed information about the data can be found in the ‘README.docx’ file.</p>
Classification of unstructured text in types of violence against women using text mining and Machine learning techniques
<p>These are the data used for the development of the investigation.</p> <p>This file was extracted from our mongoDB database. The data set contains real news of violence against women, which were organized with their date, the title and the body of the news.</p>
Machine Learning Ready Induced Seismicity Data
<p>This dataset contains previously published data on induced seismicity that has been processed to be machine learning ready.<br> <br> The data contains time series of the cumulative number of seismic events in certain areas and the corresponding pressures induced from injecting fluids into the ground. The natural task is to forecast future seismicity given past seismicity and pressures. These datasets aim to require as little seismology experience as necessary to prepare the data for forecasting algorithms. <br> <br> Data is provided for different locations. For Decatur Illinois, the seismic data was taken from Williams-Stroud et al., 2018 and the pressure data originated from Luu et al., 2022. Data aggregated over the whole region lies in the temporal_datasets/decatur_illinois/ folder. The region was further subdivided into subregions and the corresponding data stored separately (e.g. in loc1). </p> <p>The Kansas data originated from Cochran et al., 2018 and is further divided into subregions. </p> <p>The Cushing, Oklahoma is adapted from Skoumal et al., 2020. </p> <p>Each seismic file contains the following columns: epoch latitude longitude depth easting northing magnitude. The epoch corresponds to the number of seconds since a certain date (e.g. November 17, 2011 for Decatur). Each seismic event corresponds to one row in the file. <br> <br> Each pressure file contains the following columns: epoch pressure dpdt. dpdt is the derivative of pressure. </p>
Bee Tracker – an open-source machine-learning based video analysis software for the assessment of nesting and foraging performance of cavity-nesting solitary bees
<p>The foraging and nesting performance of bees can provide important information on bee health and is of interest for risk and impact assessment of environmental stressors. While radio-frequency identification (RFID) technology is an efficient tool increasingly used for the collection of behavioral data in social bee species such as honey bees, behavioral studies on solitary bees still largely depend on direct observations, which is very time-consuming.</p> <p>Here, we present a novel automated methodological approach of individually and simultaneously tracking and analyzing foraging and nesting behavior of numerous cavity-nesting solitary bees. The approach consists of monitoring nesting units by video recording and automated analysis of videos by a machine learning based software. This <i>Bee Tracker</i> software consists of four trained deep learning networks to detect bees that enter or leave their nest and to recognize individual IDs on the bees' thorax as well as the IDs of their nests according to their positions in the nesting unit.</p> <p>The software is able to identify each nest of each individual nesting bee, which permits to measure individual-based measures of reproductive success. Moreover, the software quantifies the number of cavities a female enters until it finds its nest as a proxy of nest recognition, and it provides information on the number and duration of foraging trips. By training the software on 8 videos recording 24 nesting females per video, the software achieved a precision of 96% correct measurements of these parameters.</p> <p>The software could be adapted to various experimental setups by training it to an according set of videos. The presented method allows to efficiently collect large amounts of data on cavity-nesting solitary bee species and represents a promising new tool for the monitoring and assessment of behavior and reproductive success under laboratory, semi-field and field conditions.</p>
The Dataset of Quantifying Alignment Deviations for Uniaxial Material Mechanical Testing via Automated Machine Learning
<p>The dataset consists of 4 alignment deviations of the uniaxial testing machine as well as 12 strain measurement points on cruciform specimens. A deep learning model is trained on the dataset to quantify 4 alignment deviations using 12 strain values on a thin plate specimen. The design of experiments includes Optimal Latin Hypercube, numerical modelling of Finite Element Methods. Using the Optimal Latin Hypercube, 12496 distinct groups of DOE simulation tests are constructed. Under the boundary conditions of 4 distinct deviations, 12 strain values at the required location on the cruciform specimen are obtained using Python scripts.</p> <p>The nine CSV files correspond to the nine analysis steps. The only difference among the nine analysis steps is the pretension force acting on RP1. Each CSV file contains 24 columns of data, and the corresponding contents of each column of data are as follows:</p> <ul> <li>Columns 1-6 are the freedoms of RP1 reference point, which are U1, U2, U3, ur1, UR2 and UR3 respectively;</li> <li>Columns 7-12 are the freedoms of RP2 reference points, which are U1, U2, U3, ur1, UR2 and UR3 respectively;</li> <li>Columns 13-24 are the strain values of the last 12 strain measurements of the thin plate rectangular specimen。</li> </ul>
Extracting Session Keys From the Main Memory Using Brute-force and Machine Learning
<p>This dataset contains:</p> <ol> <li>Heap dump of three version of OpenSSH (V_7_9_P1, V_8_0_P1 and V_8_1_P1)</li> <li>Heap dump of two applications that uses TLS (lynx and curl)</li> <li>Network recording in format of pcap</li> <li>JSON file that contains the keys' information</li> </ol> <p>The source code is available at: https://github.com/smartvmi/SSH-TLS-key-extraction</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.