Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,782
datasets available to search
ShareScore release 0.7.1
Dataset results
1,782 results for “algorithms”
Genes with transcripts predicted to be miR-195 or miR-26b targets by all 5 predictive algorithms included in starBase, or experimentally identified as targets by pulldown assay
<p>Genes with transcripts predicted to be miR-195 or miR-26b targets by all 5 predictive algorithms included in starBase, or experimentally identified as targets by pulldown assay</p>
Data and supplementary information for "A multiple time step algorithm for trajectory surface hopping simulations"
<p>This repository contains the supplementary information as well as the data necessary to reproduce the content of the article "A multiple time step algorithm for trajectory surface hopping simulations".</p>
Global occurrence of higher-bands ECH waves in radiation belts based on a novel noise reduction algorithm
<p>The database of ECH waves based on observations from Van Allen Probes between September 17, 2012 and October 13, 2019. The probe and the year of observations are marked in the data file names. The columns from left to right are: the mean value of the SuperMAG electrojet index in the previous hour, L-Shell, the magnetic local time, the magnetic latitude, the wave amplitude of band 1to band 4.</p>
A New Tracking Algorithm Based on Foshan Total Lightning Data and the Influence of Historical Velocity and Matching Method on the Results
<p>This paper introduces a new tracking algorithm (V2CM) based on total lightning data from the Foshan Total Lightning Location System (FTLLS), and the other four algorithms are used to test the influence of historical velocity and matching method on the results during the formation of V2CM. Different methods are used to quantitatively evaluate the performance of five tracking methods. Additionally, the number of splits and merges along a track is used as a new index to describe the complexity of the track. It is found that the best tracking method V2CM has the highest probability of detection (POD), the false alarm rate (FAR) and the critical success index (CSI) being 65.6%, 39.2%, 46.4%, respectively. It also produces longer and more coherent tracks than those of other four methods. The average number of splits and merges of the best method is 2.40 and lightning clusters matching method between two successive time intervals has a greater impact on the evaluation results than the historical velocity. In addition, lightning clusters move in the northeast direction, with a mean speed of 48.48 km/h and a median speed of 39.08 km/h. Complex tracks, involving splitting or merging, move faster than simple track.</p>
Anomaly Detection Algorithm Performance
<p>Anomaly detection algorithms performance metrics: AUC and Average precision; two sets of 298 + 13 algorithms; 9315 datasets.</p>
Optimization used in in the publication scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data
<p>OptimizationResults contain the expanded input model for each data set, in the C field are the indices of the core reactions in the expanded input model used to build the multi-cell population for the 121 different parameters, A contains the indices of reaction in the multi-cell population model for the 121 runs, thresh the parameter setting for the cover and REI. The REI is given in %. So 5 REI means 0.05. While A might change because of alternative optimals when run on a different computer, C and the expanded Input model will remain unchanged.</p> <p> </p> <p>Multi_cell_population_CRC and Multi_cell_population_NM are the models obtained by scFASTCORMICS with the optimal setting for Dataset1 and Dataset2, respectively. </p> <p>Please, check the publication (scFASTCORMICS: A contextualization algorithm to reconstruct metabolic multi-cell population models from single-cell RNAseq data) for more details.</p>
Algorithms for determining transposable genes in a genome
<p>Transposons are nucleotide sequences in DNA that can change their positions. Many transposons are shorter than a general gene. When we restrict to nucleotide sequences that form complete genes, we can still find genes that change their relative locations in a genome. Thus for different individuals of the same species, the orders of genes might be different. A practical problem is to determine such transposable genes in given gene sequences. Through an intuitive rule, we transform the biological problem of determining transposable genes into a rigorous mathematical problem of determining the longest common subsequence. Depending on whether the gene sequence is linear (each sequence has a fixed head and tail) or circular (we can choose any gene as the head, and the previous one is the tail), and whether genes have multiple copies, we classify the problem of determining transposable genes into four scenarios: (1) linear sequences without duplicated genes; (2) circular sequences without duplicated genes; (3) linear sequences with duplicated genes; (4) circular sequences with duplicated genes. With the help of graph theory, we design fast algorithms for different scenarios. Specifically, we study the situation where the longest common subsequence is not unique.</p> <p>This dataset contains code files for the corresponding algorithms. Besides, it has gene sequence data for certain <em>Escherichia</em> <em>coli</em> strains (from NCBI), which are used to test those algorithms.</p>
Apple orchard production estimation using deep learning strategies: a comparison of tracking-by-detection algorithms - SensitivityAnalysis
<p>The dataset "Sensitivity Analysis" consists of image sequences (videos) for apple detection and tracking and its corresponding ground truth. The ground truth is presented in MOT format. This dataset is part of the paper:</p> <p>Villacrés, J., Viscaino, M., Delpiano, J., Vougioukas, S. & Cheein, F. A. (2022). Apple orchard production estimation using deep learning strategies: a comparison of tracking-by-detection algorithms. <em>Computers and Electronics in Agriculture</em>.</p> <p>The article is currently accepted. For a better reference format, please refer to the journal's official website.</p> <p>If you have used the material presented in this data set, please cite the previous article.</p> <p>For more information regarding the dataset, please refer to the paper mentioned below.</p>
Shorelines dataset. Deliverable 3.2 - Algorithms for satellite derived shoreline mapping and shorelines dataset, ECFAS project (GA 101004211), www.ecfas.eu.
<p>The European Copernicus Coastal Flood Awareness System (ECFAS) project aimed at contributing to the evolution of the Copernicus Emergency Management Service (https://emergency.copernicus.eu/) by demonstrating the technical and operational feasibility of a European Coastal Flood Awareness System. Specifically, ECFAS provides a much-needed solution to bolster coastal resilience to climate risk and reduce population and infrastructure exposure by monitoring and supporting disaster preparedness, two factors that are fundamental to damage prevention and recovery if a storm hits.</p> <p>The ECFAS Proof-of-Concept development ran from January 2021 to December 2022. The ECFAS project was a collaboration between Scuola Universitaria Superiore IUSS di Pavia (Italy, ECFAS Coordinator), Mercator Ocean International (France), Planetek Hellas (Greece), Collecte Localisation Satellites (France), Consorzio Futuro in Ricerca (Italy), Universitat Politecnica de Valencia (Spain), University of the Aegean (Greece), and EurOcean (Portugal), and was funded by the <strong>European Commission H2020 Framework Programme</strong> within the call LC-SPACE-18-EO-2020 - Copernicus evolution: research activities in support of the evolution of the Copernicus services. </p> <p><em><strong>Reference literature:</strong></em></p> <p><em><strong>Palomar-Vázquez, J.; Pardo-Pascual, J.E.; Almonacid-Caballer, J.; Cabezas-Rabadán, C. Shoreline Analysis and Extraction Tool (SAET): A New Tool for the Automatic Extraction of Satellite-Derived Shorelines with Subpixel Accuracy. Remote Sens. 2023, 15, 3198. <a href="https://doi.org/10.3390/rs15123198">https://doi.org/10.3390/rs15123198</a></strong></em></p> <p><em><strong>J.E. Pardo-Pascual, J. Almonacid-Caballer, C. Cabezas-Rabadán, A. Fernández-Sarría, C. Armaroli, P. Ciavola, J. Montes, P.E. Souto-Ceccon, J. Palomar-Vázquez: Assessment of satellite-derived shorelines automatically extracted from Sentinel-2 imagery using SAET. Coastal Engineering, 2023, 104426, ISSN 0378-3839, <a href="https://doi.org/10.1016/j.coastaleng.2023.104426">https://doi.org/10.1016/j.coastaleng.2023.104426</a>.</strong></em></p> <p><em>Pardo-Pascual, J. E., Cabezas-Rabadán, C., Palomar-Vázquez, J., Fernández-Sarría, A., Almonacid-Caballer, J., Souto-Ceccon, P. E., Montes, J., Armaroli, C., and Ciavola, P.: Satellite-derived shorelines extracted using SAET for characterizing the effect of Storm Gloria in the Ebro Delta (W Mediterranean), EGU General Assembly 2022, Vienna, Austria, 23–27 May 2022, EGU22-9856, <a href="https://doi.org/10.5194/egusphere-egu22-9856">https://doi.org/10.5194/egusphere-egu22-9856</a>, 2022.</em></p> <p>In relation to <strong>SAET</strong>, additional information, instructions and the <strong>open-source code</strong> are available here <a href="https://doi.org/10.5281/zenodo.5807710"><strong>https://zenodo.org/records/10256957</strong></a> and also in <strong>GitHub </strong>here<strong> <a href="https://github.com/jpalomav/SAET_master">https://github.com/jpalomav/SAET_master</a></strong></p> <p> </p> <p><strong>Description of the containing files inside the Dataset</strong></p> <p>The deliverable includes two different files: the dataset of shorelines and the accompanying report.</p> <p>The report describes the structure of the dataset of shorelines produced in Task 3.2 - Shoreline mapping validation and calibration and described in Deliverable 3.2 - Algorithms for satellite derived shoreline mapping and shorelines dataset (DOI: 10.5281/zenodo.5807711).</p> <p>The dataset is composed of three folders.</p> <ol> <li>The first folder ("SDS_vs_VHR_shoreline") contains the SDSs extracted at each test site by the different algorithms tested in ECFAS (SHOREX, Cabezas-Rabadán et al., 2021; Sánchez-García et al., 2020; CoastSat, Vos et al., 2019a,b and SAET, Palomar-Vázquez et al., 2021; Pardo-Pascual et al., 2021) to assess their performance through the comparison with shorelines photo-interpreted on coincident VHR satellite images. The SDSs obtained at all the test sites using SAET and CoastSAT are included, as well the SDSs obtained using SHOREX at the Spanish test sites. The photo-interpreted shorelines are also provided. For all the test sites it is provided the line separating the instantaneous shoreline (water/land boundary), obtained by photo-interpretation. For the sites in the Netherlands, also the wet/dry line is provided.</li> <li>The second folder ("SDS_vs_video-monitored_shorelines") presents for each of the three test sites the SDSs obtained using the different algorithms and video-derived shorelines. The SDSs at the beaches of la Victoria and Cala Millor were obtained using SAET, SHOREX and CoastSAT. In the case of Gerakas Beach, the shorelines were obtained using SAET and CoastSAT.</li> <li>The third folder ("SDS_SAET_Storm_cases") includes eight examples in different European coasts in which the SDSs before and after a storm have been obtained.</li> </ol> <p>This <strong>ECFAS dataset of shorelines</strong> is made available under the <strong>Open Database License</strong>: <a href="http://opendatacommons.org/licenses/odbl/1.0/">http://opendatacommons.org/licenses/odbl/1.0/</a>. Any rights in individual contents of the ECFAS dataset of shorelines are licensed under the <strong>Open Database License</strong>: <a href="http://opendatacommons.org/licenses/dbcl/1.0/">http://opendatacommons.org/licenses/dbcl/1.0/</a>.</p> <p>This <strong>Report</strong> is made available under the <strong>Creative Commons Attribution 4.0 International License</strong>.</p> <p> </p> <p><em><strong>Disclaimer:</strong></em></p> <p>ECFAS partners provide the data "as is" and "as available" without warranty of any kind. The ECFAS partners shall not be held liable resulting from the use of the information and data provided.</p> <p>This project has received funding from the Horizon 2020 research and innovation programme under grant agreement No. 101004211</p> <p> </p>
Drivers of bat activity at wind turbines advocate for mitigating bat exposure using multicriteria algorithm-based curtailment
<p>data used for the paper</p>
Geomagnetic datasets of BJI station reconstructed through Artificial Neural Network improved by Genetic Algorithm in 2021
<p>Beijing station established in 1954 is one of the oldest geomagnetic observatories in China, which plays an important role in data exchange, and further provide data or standardization for satellite observation and geomagnetic model construction. With the development of urbanization, the observed data are greatly disturbed by subways, and data disturbed are almost unavailable. The dataset was reconstructed through Artificial Neural Network improved by Genetic Algorithm, including minutely data of three components (<em>D</em>, <em>H</em> and <em>Z</em>) in 2021. This reconstruction method has been proved to be effective.</p>
Screening of key risk SNPs for glioma based on machine learning algorithms
<p>Glioma is a common primary malignant brain tumor and is the most aggressive and lethal solid tumor, accounting for approximately 80% of all intracranial malignancies. Our aim was to screen key SNP by LASSO regression and random forest (a machine learning algorithm) and construct a model based on these SNP to predict the risk of glioma in Chinese Han population.</p>
When to be Discrete: Analyzing Algorithm Performance on Discretized Continuous Problems - Data and Code
<p>This repository contains the code and data for reproducibility of the paper 'Modular Differential Evolution'. </p> <p>The following files are included:</p> <p>- experiments: Python files which were used to create the discretized functions and collect the data (the experiment_runner files) and the code for each of the tested algorithms (two subfolders)</p> <p>- Process_discr: code which takes the raw IOH data and turns it into the relevant csv-files which are used to create the figures. CSV files are included 'csv_fb_*.zip'. ecdf_creation_script is used to extract ecdf and ert data for the ECDF_plot notebook. This data is added as 'other_csvs.zip'</p> <p>- Plot*: notebooks which are used to generate the figures from the papers</p>
Computer vision: algorithms to make sense of the world
<p><strong>The following video describes how computer vision is used by the ROMI platform for object and species detection in both 2D and 3D, and how it is integral to the weeding tool. Funded by EU Grant 773875.</strong></p> <p><em>Videos are available in:</em></p> <ul> <li>Hi-res (1080p Apple ProRes)</li> <li>Mid-res (1080p H265)</li> </ul> <p><strong>Video script:</strong></p> <p>(CAMPRODON) What's computer vision? Mmm … computer vision for me is making sense of pixels. I think computer vision has a profound effect in the way that we understand the world because as humans vision is so centric right. If dogs would be making computers, maybe they wouldn't talk that much about vision. But for us it's so centric in the way that we perceive the world and the way we learn about the world, that actually I think it's easier for that of course to program and think useful ways machines could get information you know through vision. And also especially because vision is one of the much more complex senses that we have.<br> <br> (SOLLAZZO) Image and videos that represent nowadays the 80 percent of the data that we produce and the introduction of computer vision and machine learning becomes necessary to start extrapolating information out of this new source of data.<br> <br> (COLLIAUX) So just a point of clarification because we often talk about AI and so just to be a bit more precise about what we do in ROMI. Because AI is quite a vague term, and so what we do mainly is robotics and computer vision.<br> <br> (SOLLAZZO) Computer vision is at the end a limited set of tools and systems that are basically based on mathematical representation and description of the pixel that represent the image, they are part of the image, and machine learning is based on a different approach of interpretation of those pixels.<br> <br> (COLLIAUX) So the rover is for weeding and to remove the weeds you need to detect the weeds first and so we use a computer vision algorithm to detect where are the weeds where are the salads.<br> <br> (SOLLAZZO) So let's see one by one which are the methods that we implemented, in our algorithm, in our system. So we start with the feature extraction in order to do that in fact we go one by one over the images and we understand which are the pixels in common between one and the other. From these method in fact it's possible to recreate an orthomosaic view, an orthomosaic image, but afterwards we need to align it to all the previous images that we've been creating in the previous analysis. So after the generation of the orthomosaic view, what we do is that we start to cut the main image into a portion into a series of smaller portions. This facilitates the execution of the machine learning algorithm and the possibility to recognise the presence or not, of lettuce in the scene. After the recognition has been performed we put together the images once again and we can reconstruct an orthomosaic view with a detected position of the different lettuce. This is necessary to understand not only the position but also the area of growth that these different lettuce are occupying over time. From the geolocation of every single plant we start to analyse the growing curve over time. This is possible thanks to the implementation of ‘Mask RCNN’. So thanks to the generation of all these different areas that during time, will tell us the growing pattern of every single lettuce, and this will be extremely useful to understand when is the moment to harvest the plant when the plant is in fact bolting, more or less this is ok.<br> <br> (COLLIAUX) So we do what I showed was about 2d computer vision, but we do a lot of 3d computer vision also in the project and so let me show you a bit what we do with a plant scanner. So it is uh used by biologists to study the geometry of the plants so they want to reconstruct the pre-architecture of a plant and study that architecture. So for this we take many images of a plant by turning a camera in a circle around the plant, we generate a mask but again a segmentation algorithm to detect where where the plant is and where the background is, and then we can generate a point cloud by an algorithm called ‘space carving’ or ‘shape from silhouette’ which based on the many silhouettes you collected it looks for it it carves the space for the shape which is the most compatible with all the projection of the shape.<br> <br> (CAMPRODON) So what we're doing in ROMI at the end, I would say in a way we hack existing technologies, we take advantage of the low cost cameras that exist in phones right we don't need to rely anymore in high-end industrial cameras, we take advantage of the low-cost computational power, computing cheaper than ever. So these images that we take we can process them with software in ways that was not possible before, we take advantage of software, of especially of open source software and free software and then we build the training models right, so this software is capable to detect on top of that images insights, to go from data to information.</p>
Data segmentation and analysis: developing algorithms to virtually dissect plants
<p><strong>The following video describes how biological data analysis and image segmentation is conducted at different scales and informs the ROMI data pipeline. Funded by EU Grant 773875.</strong></p> <p><em>Videos are available in:</em></p> <ul> <li>Hi-res (1080p Apple ProRes)</li> <li>Mid-res (1080p H265)</li> </ul> <p><strong>Video script:</strong></p> <p>(LEGRAND)<strong> </strong>I’m an engineer in a biological data analysis. So what I'm doing is to create tools and create a code helping biologists to go from these images they acquire to this labeled image from which we extract cellular features, and from these features we do statistical analysis. Okay so on this board we describe the pipeline we are trying to set up for this image analysis so we're starting with potentially five dimensional images so you have the XYZ spatial Dimension then you have channels and potentially time. So with these images they go to a reader this reader creates a specific data structure so for example if you have a multiple observation of the same object from different angle what you would like to do is to reconstruct and fuse together this multiple angle and then that gives you one big image that you may want to filter and from from these images you can then perform segmentation. For example nuclei detection or cellular segmentation for example this is a pull out transport pump being able to quantify how many pumps you have gives you an idea of the flow of protein or hormones so from that you will be able to build up models and to try and make a realistic model of flower development or phyllotaxis. If you talk about the flower arrangement around the the stem.<br> <br> So compared to my original work the ROMI project is for me we represent a change in scale so we're moving from the cellular scale or to tissular scale, to where we try to understand how flower or leaves are arranged around the stem to use a macroscopic scale where you see the plant in full. And you're trying to observe and also quantify also how flowers or leaves are arranged around the stem.<br> <br> (HÉTROY-WHEELER) So I'm a bit at the end of the pipeline so as input we take a 3D Point cloud which is a sample of the plant represented in a virtual way so we've got a points which each of them has three 3D coordinates. And only from the set of points we try to infer the geometry of the plant and from this geometry so basically to recover the shape of the plant we try then to segment the plant into its organs so the leaves the stems and in between stems and leaves the petals. So the idea is that only from geometry and maybe some colour information or other information which we try to use as less information as possible we would be able to detect the organs of the plants and then do some computation for example simple computation like computing the number of leaves but also more advanced computation like for example trying to guess what is the area of the leaves or the angles between the different stems and so on things that are useful for a biologist and also for people in Agronomy or agriculture.</p>
Dataset for testing the Adaptive Source Term Estimation algorithm
<p>The objective of this work is to study the Adaptive Source Term Estimation algorithm developed by the authors and that will be presented in a future article. Briefly, source term estimation (STE) is a mathematical tool aimed to characterize the source parameters of some gas hazard based on its measurements at given positions. The adaptive STE (ASTE) is aimed to perform this procedure iteratively, i.e. based on the received measurements at a given set of deployed detectors, the algorithm proposes the location for the next detector in order to optimally solve the STE problem.</p>
"Comparing Algorithm Selection Approaches on Black-Box Optimization Problems"
<p>Code and data for the paper "Comparing Algorithm Selection Approaches on Black-Box Optimization Problems".</p> <p>Details on the algorithm portfolios can be found in algorithm_portfolio.pdf</p>
Double shrinking (DOSH), a regression-based algorithm for gene regulatory network inference from co-expression data
<p>Data for the preprint "Double shrinking (DOSH), a regression-based algorithm for gene regulatory network inference from co-expression data". The preprint is live on ResearchSquare: <a href="http://t.researchsquare.com/track/click/31114617/doi.org?p=eyJzIjoiWVJQQUYtT09mMXFoWnRoMGk0SlpQZTZqWWpJIiwidiI6MSwicCI6IntcInVcIjozMTExNDYxNyxcInZcIjoxLFwidXJsXCI6XCJodHRwczpcXFwvXFxcL2RvaS5vcmdcXFwvMTAuMjEyMDNcXFwvcnMuMy5ycy0yNzM4NjgzXFxcL3YxXCIsXCJpZFwiOlwiZDMzODNjZGNhNWNiNGE2Yjk5NWRkY2UyNmIyODI5NTlcIixcInVybF9pZHNcIjpbXCIzZGQwZTAxMmExMzk4NDhkNTAzYjI4ZTBiZmU1Y2QxMDcxNzhlZTgwXCJdfSJ9">10.21203/rs.3.rs-2738683/v1</a>.</p>
Nevergrad algorithm performance on SBOX-COST suite
<p>The files "XD_benchmark_data.zip" contain the performance results for the respective dimensions of the algorithms Estimation of Multivariate Normal Algorithm (EMNA), Differential Evolution (DE), Constrained Optimization BY Linear Approximation (Cobyla) and Particle Swarm Optimization (PSO) run on the SBOX-COST benchmarking suite. The data is stored in IOHprofiler format.<br> <br> "data_collection_SBOX_ng.py" contains the code used to collect the data.<br> <br> The remaining two zip files contain plots generated by IOHanalyzer for all used dimensionalities {2, 3, 5, 10, 20, 40}.</p>
Supplementary data to 'Coupling between cohesive element method and Node-to-segment contact algorithm : Implementation and application'
<p>This archive contains the data and the scripts for the simulations presented in the paper</p> <p>Pundir, M., Guillaume, A. "Coupling between cohesive element method and Node-to-segment contact algorithm : Implementation and application'" (2020)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.