Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,782
datasets available to search
ShareScore release 0.7.1
Dataset results
1,782 results for “algorithms”
Implementation of Genetic Algorithms to Optimize Metal-Organic Frameworks for CO2 Capture
<p>Dataset associated with the publication "Implementation of Genetic Algorithms to Optimize Metal-Organic Frameworks for CO2 Capture".</p> <p> </p> <p>Changelog:</p> <p>- Include sample input files for GCMC using RASPA2 and geometry optimization using LAMMPS.</p>
Enhancement of the Ensemble Nonlinear Least Squares Algorithm for i4DVar
Open the record for dataset details and reuse information.
SEN2VENµS, a dataset for the training of Sentinel-2 super-resolution algorithms
<p><strong>1 Description</strong></p> <p><strong>SEN2VENµS</strong> is an open dataset for the super-resolution of Sentinel-2 images by leveraging simultaneous acquisitions with the VENµS satellite. The dataset is composed of 10m and 20m cloud-free surface reflectance patches from Sentinel-2, with their reference spatially-registered surface reflectance patches at 5 meters resolution acquired on the same day by the VENµS satellite. This dataset covers 29 locations with a total of 132 955 patches of 256x256 pixels at 5 meters resolution, and can be used for the training of super-resolution algorithms to bring spatial resolution of 8 of the Sentinel-2 bands down to 5 meters.</p> <p><strong>Changelog with respect to version 1.0.0</strong> (https://zenodo.org/records/6514159)</p> <ul> <li>All patches are now stored in indivual geoTiFF files with proper geo-referencing, regrouped in zip files per site and per category,</li> <li>The dataset now includes 20 meter resolution SWIR bands B11 and B12 from Sentinel-2 (L2A from Theia). Note that there is no HR reference for those bands, since the VENµS sensor has no SWIR band.</li> </ul> <p><strong>2 Files organization</strong></p> <p>The dataset is composed of separate sub-datasets embedded in separate zip files, one for each site, as described in table <a href="#org5e17b56">1</a>. Note that there might be slight variations in number of patches and number of pairs with respect to version 1.0.0, due do incorrect count of samples in previous version (an empty tensor was still accounted for).</p> <p>Table 1: Number of patches and pairs for each site, along with VENµS viewing zenith angle</p> <table> <tbody> <tr> <th>Site</th> <th>Number of patches</th> <th>Number of pairs</th> <th>VENµS Zenith Angle</th> </tr> </tbody> <tbody> <tr> <td>FR-LQ1</td> <td>4888</td> <td>18</td> <td>1.795402</td> </tr> <tr> <td>NARYN</td> <td>3813</td> <td>24</td> <td>5.010906</td> </tr> <tr> <td>FGMANAUS</td> <td>129</td> <td>4</td> <td>7.232127</td> </tr> <tr> <td>MAD-AMBO</td> <td>1442</td> <td>18</td> <td>14.788115</td> </tr> <tr> <td>ARM</td> <td>15859</td> <td>39</td> <td>15.160683</td> </tr> <tr> <td>BAMBENW2</td> <td>9018</td> <td>34</td> <td>17.766533</td> </tr> <tr> <td>ES-IC3XG</td> <td>8822</td> <td>34</td> <td>18.807686</td> </tr> <tr> <td>ANJI</td> <td>2312</td> <td>14</td> <td>19.310494</td> </tr> <tr> <td>ATTO</td> <td>2258</td> <td>9</td> <td>22.048651</td> </tr> <tr> <td>ESGISB-3</td> <td>6057</td> <td>19</td> <td>23.683871</td> </tr> <tr> <td>ESGISB-1</td> <td>2891</td> <td>12</td> <td>24.561609</td> </tr> <tr> <td>FR-BIL</td> <td>7105</td> <td>30</td> <td>24.802892</td> </tr> <tr> <td>K34-AMAZ</td> <td>1384</td> <td>20</td> <td>24.982675</td> </tr> <tr> <td>ESGISB-2</td> <td>3067</td> <td>13</td> <td>26.209776</td> </tr> <tr> <td>ALSACE</td> <td>2653</td> <td>16</td> <td>26.877071</td> </tr> <tr> <td>LERIDA-1</td> <td>2281</td> <td>5</td> <td>28.524780</td> </tr> <tr> <td>ESTUAMAR</td> <td>911</td> <td>12</td> <td>28.871947</td> </tr> <tr> <td>SUDOUE-5</td> <td>2176</td> <td>20</td> <td>29.170244</td> </tr> <tr> <td>KUDALIAR</td> <td>7269</td> <td>20</td> <td>29.180855</td> </tr> <tr> <td>SUDOUE-6</td> <td>2435</td> <td>14</td> <td>29.192055</td> </tr> <tr> <td>SUDOUE-4</td> <td>935</td> <td>7</td> <td>29.516127</td> </tr> <tr> <td>SUDOUE-3</td> <td>5363</td> <td>14</td> <td>29.998115</td> </tr> <tr> <td>SO1</td> <td>12018</td> <td>36</td> <td>30.255978</td> </tr> <tr> <td>SUDOUE-2</td> <td>9700</td> <td>27</td> <td>31.295256</td> </tr> <tr> <td>ES-LTERA</td> <td>1701</td> <td>19</td> <td>31.971764</td> </tr> <tr> <td>FR-LAM</td> <td>7299</td> <td>22</td> <td>32.054056</td> </tr> <tr> <td>SO2</td> <td>738</td> <td>22</td> <td>32.218481</td> </tr> <tr> <td>BENGA</td> <td>5857</td> <td>28</td> <td>32.587334</td> </tr> <tr> <td>JAM2018</td> <td>2564</td> <td>18</td> <td>33.718953</td> </tr> </tbody> </table> <p> </p> <p>Each site zip file contains a subfolder with the site name. This subfolder contains secondary zip files for each date, following this naming convention as the pair <code>id</code>: <code>{site_name}_{acquisition_date}_{mgrs_tile}</code>. For each date, 5 zip files are available, as shown in table <a href="#org504e2aa">2</a>.Each zip file contain subfolder <code>{bands}/{resolution}/</code> in which one GeoTiFF file per patch is stored, with the following naming convention: <code>{site_name}_{idx}_{acquisition_date}_{mgr_tile}_{bands}_{resolution}.tif</code>. Pixel values are encoded as 16 bits signed integers and should be converted back to floating point surface reflectance by dividing each and every value by 10 000 upon reading.</p> <p>Table 2: Naming convention for zip files associated to each date.</p> <table> <tbody> <tr> <th>File</th> <th>Content</th> </tr> </tbody> <tbody> <tr> <td><code>{id}_05m_b2b3b4b8.zip</code></td> <td>5m patches (\(256\times256\) pix.) for S2 B2, B3, B4 and B8 (from VENµS)</td> </tr> <tr> <td><code>{id}_10m_b2b3b4b8.zip</code></td> <td>10m patches (\(128\times128\) pix.) for S2 B2, B3, B4 and B8 (from Sentinel-2)</td> </tr> <tr> <td><code>{id}_05m_b5b6b7b8a.zip</code></td> <td>5m patches (\(256\times256\) pix.) for S2 B5, B6, B7 and B8A (from VENµS)</td> </tr> <tr> <td><code>{id}_20m_b5b6b7b8a.zip</code></td> <td>20m patches (\(64\times64\) pix.) for S2 B5, B6, B7 and B8A (from Sentinel-2)</td> </tr> <tr> <td><code>{id}_20m_b11b12.zip</code></td> <td>20m patches (\(64\times64\) pix.) for S2 B11 and B12 (from Sentinel-2)</td> </tr> </tbody> </table> <p> </p> <p>Each file comes with a master <code>index.csv</code> CSV (Comma Separated Values) file, with one row for each pair sampled in the given site. Columns are named after the <code>{bands}_{resolution}</code> pattern, and contains the full path to the corresponding GeoTiFF wihin the corresponding zip file:</p> <p><code>{site}_{acquisition_date}_{mgrs_tile}_{bands}_{resolution}.zip/{bands}/{resolution}/{site}_{idx}_{acquisition_date}_{mgrs_tile}_{bands}_{resolution}.tif</code></p> <p><strong>3 Licencing</strong></p> <p><strong>3.1 Sentinel-2 patches</strong></p> <p><strong>3.1.1 Copyright</strong></p> <p>Value-added data processed by CNES for the Theia data centre www.theia-land.fr using Copernicus products. The processing uses algorithms developed by Theia's Scientific Expertise Centres. Note: Copernicus Sentinel-2 Level 1C data is subject to this license: <a href="https://theia.cnes.fr/atdistrib/documents/TC_Sentinel_Data_31072014.pdf">https://theia.cnes.fr/atdistrib/documents/TC_Sentinel_Data_31072014.pdf</a></p> <p><strong>3.1.2 Licence</strong></p> <p>Files <code>*_b2b3b4b8_10m.tif</code>, <code>*_b5b6b7b8a_20m.tif</code> and <code>*_b11b12_20m.tif</code> are distributed under the the original licence of the Sentinel-2 Theia L2A products, which is the Etalab Open Licence Version 2.0 <sup><a href="#fn.2">2</a></sup>.</p> <p><strong>3.2 VENµS patches</strong></p> <p><strong>3.2.1 Copyright</strong></p> <p>Value-added data processed by CNES for the Theia data centre www.theia-land.fr using VENµS satellite imagery from CNES and Israeli Space Agency. The processing uses algorithms developed by Theia's Scientific Expertise Centres.</p> <p>3.2.2 <strong>Licence</strong></p> <p>Files <code>*_b2b3b4b8_05m.tif</code> and <code>*_b5b6b7b8a_05m.tif</code> are distributed under the original licence of the VENµS products, which is Creative Commons BY-NC 4.0 <sup><a href="#fn.3">3</a></sup>.</p> <p><strong>3.3 Remaining files</strong></p> <p>All remaining files are distributed under the Creative Commons BY 4.0 <sup><a href="#fn.4">4</a></sup> licence.</p> <p><strong>4 Note to users</strong></p> <p>Note that even if the VenµS2 dataset is sorted by sites and by pairs, we strongly encourage users to apply the full set of machine learning best practices when using it : random keeping separate pairs (or even sites) for testing purpose, and randomization of patches accross sites and pairs in the training and validation sets.</p> <p><strong>5 Citing</strong></p> <p>Please cite the following data paper (preprint, submitted to <em>MDPI Data</em>) and zenodo link when publishing work derived from this dataset:</p> <p>Michel, J.; Vinasco-Salinas, J.; Inglada, J.; Hagolle, O. SEN2VENµS, a Dataset for the Training of Sentinel-2 Super-Resolution Algorithms. <em>Data</em> <strong>2022</strong>, <em>7</em>, 96. https://doi.org/10.3390/data7070096</p> <p><a href="https://zenodo.org/deposit/6514159">10.5281/zenodo.14603764</a></p> <p><strong>Footnotes:</strong></p> <p><sup><a href="#fnr.1">1</a></sup></p> <p><a href="https://pytorch.org/">https://pytorch.org/</a></p> <p><sup><a href="#fnr.2">2</a></sup></p> <p><a href="https://theia.cnes.fr/atdistrib/documents/Licence-Theia-CNES-Sentinel-ETALAB-v2.0-en.pdf">https://theia.cnes.fr/atdistrib/documents/Licence-Theia-CNES-Sentinel-ETALAB-v2.0-en.pdf</a></p> <p><sup><a href="#fnr.3">3</a></sup></p> <p><a href="https://creativecommons.org/licenses/by-nc/4.0/">https://creativecommons.org/licenses/by-nc/4.0/</a></p> <p><sup><a href="#fnr.4">4</a></sup></p> <p><a href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</a></p>
Testing the Acoustic Localisation Positioning System-algorithm: Data sets
<p>The data is from 3 separate experimental setups: Tracking a device of constant speed using the Audio1 (single processor) software (B.1); and tracking a faster device using the al-Qt (multi processor) software (B.2) The experiments were also recorded on video. (see Video folder). For the experiments in B.1 a visualisation of the data has been made with the help of a MATLAB script which also included. The dimensions of the room and loudspeaker spacings were identical for both set-ups, namely 3 x 2 meters, with the loudspeakers at the corners of the rectangle. the loudspeakers were situated on the floor, approximately level with the tracked devices.</p> <p>Setup C collects debug log examples of experiments with very short capture cycles with 2 mics and 2 loudspeakers running ALPS al-Qt.</p> <p>The source code for Audio is available from: <a href="https://github.com/spatmus/alps/tree/master/Audio1">https://github.com/spatmus/alps/tree/master/Audio1</a><br> The source code for al-Qt is available from: <a href="https://github.com/spatmus/alps/tree/master/al-Qt">https://github.com/spatmus/alps/tree/master/al-Qt</a><br> A binary of al-Qt for macOS 11, can be found in the assets folder of the release <a href="https://github.com/spatmus/alps/releases">https://github.com/spatmus/alps/releases</a></p>
Data for "Heuristic algorithms in Evolutionary Computations and modular organization of biological macromolecules: applications to in vitro evolution"
<p>This publication contains data for the construction of the figures and tables for the paper "Heuristic algorithms in Evolutionary Computations and modular organization of biological macromolecules: applications to in vitro evolution" accepted to publication in PLOS ONE.</p>
Using cancer risk algorithms to improve risk estimates and referral decisions
<p><span><span><span><span><span><span><span><span><span><span><span><b>Background: </b></span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>Cancer risk algorithms were introduced to clinical practice in the last decade, but they remain underused. We investigated whether GPs change their referral decisions in response to an unnamed algorithm, if decisions improve, and if changing decisions depends on having information about the algorithm and on whether GPs overestimated or underestimated risk.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Methods: </b></span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>157 UK General Practitioners (GPs) were presented with 20 vignettes describing patients with possible colorectal cancer symptoms. GPs gave their risk estimates and inclination to refer. They then saw the risk score of an unnamed algorithm and could update their responses. Half of the sample was given information about the algorithm's derivation, validation, and accuracy. At the end, we measured their algorithm disposition. We analysed the data using multilevel regressions with random intercepts by GP and vignette.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Results: </b></span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>We find that, after receiving the algorithm's estimate, GPs' inclination to refer changes 26% of the time and their decisions switch entirely 3% of the time. Decisions become more consistent with the NICE 3% referral threshold (<i>OR</i> 1.45 [1.27, 1.65], <i>p</i><.001). The algorithm's impact is greatest when GPs have underestimated risk. Information about the algorithm does not have a discernible effect on decisions but it results in a more positive GP disposition towards the algorithm. GPs' risk estimates become better calibrated over time, i.e., move closer to the algorithm. </span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span><b>Conclusions: </b></span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>Cancer risk algorithms have the potential to improve cancer referral decisions. Their use as learning tools to improve risk estimates is promising and should be further investigated. </span></span></span></span></span></span></span></span></span></span></span></p>
BasqueRoads, a dataset to benchmark road selection algorithms
<p>This dataset contains two road networks, one used for 1:25k scale maps (roads_ini), and one used at the 1:80k scale (roads_final). This dataset can be used to benchmark road selection algorithm that seek to transform the initial dataset into the final dataset.</p>
Clear-sky profile database for the development of Land Surface Temperature algorithms
<p>This dataset includes clear sky atmospheric profiles from the European Centre for Medium Range Forecast (ECMWF) version-5 reanalysis (ERA5), specially selected to support the development of algorithms of Land Surface Temperature (LST) retrieval from Earth observation (EO) data. The profiles were re-sampled from an ERA5 dataset covering the 2009-2019 period, with a 1x1 degree spatial resolution, hourly sampling and using the full vertical resolution (137 model levels). The re-sampling technique is based on a dissimilarity criterion applied to profiles of temperature and specific humidity, in order to obtain regular distributions of atmospheric variables of relevance for LST retrieval in the Thermal Infrared (TIR) spectral range. The database is limited to clear-sky conditions over land, being therefore suitable for the development of satellite land products relying on optical and thermal infrared imagery in general, despite targeting especially LST.</p> <p><strong>Dataset description:</strong></p> <p>The dataset is divided in multiple netCDF4 files based on the range of skin temperature (Tskin; Kelvin) and the range of total column water vapour (TCWV; mm). Each file includes the following variables:</p> <ul> <li>Time</li> <li>Longitude</li> <li>Latitude</li> <li>2-m temperature (t2m)</li> <li>Surface pressure (sp)</li> <li>Total cloud cover (tcc)</li> <li>Total column water vapour (tcwv)</li> <li>Skin temperature (skt)</li> <li>Surface emissivity (emis)</li> <li>Land cover classification (lcc)</li> <li>Temperature profile (t)</li> <li>Specific humidity profile (q)</li> <li>Ozone profile (o3)</li> <li>Pressure profile (p)</li> </ul> <p>All profiles are provided on model levels. For each profile, 6 values of skin temperature and 25 values of emissivity are provided (see publication for details). Emissivity values correspond to the wavelengths of ~11 and ~12 µm.</p> <p><strong>Credit:</strong></p> <p>To use this data please cite this dataset and the respective journal publication:</p> <p>Ermida, S.L.; Trigo, I.F. (2022) A Comprehensive Clear-Sky Database for the Development of Land Surface Temperature Algorithms. <em>Remote Sens.</em>, <em>14</em>, 2329. <a href="https://doi.org/10.3390/rs14102329">https://doi.org/10.3390/rs14102329</a> </p> <p> </p> <p><strong>Access: </strong></p> <p>Currently, Zenodo does not provide a simple way to download datasets with a large number of files. We recomend trying the <a href="https://zenodo.org/record/3676567#.YnJAxtPMJhE">Zenodo_get</a> to simplify the download.</p>
Traffic images captured from UAVs for use in training Machine Vision Algorithms for traffic management
<p><strong>If you use this dataset please cite this paper: Bemposta Rosende, S.; Ghisler, S.; Fernández-Andrés, J.; Sánchez-Soriano, J. Dataset: Traffic Images Captured from UAVs for Use in Training Machine Vision Algorithms for Traffic Management. Data 2022, 7, 53. https://doi.org/10.3390/data7050053</strong></p> <p>A dataset of road traffic images taken from unmanned aerial vehicles (UAV) with the purpose of being used to train artificial vision algorithms, among which those based on convolutional neural networks stand out. </p> <p>Dataset is available and accessible in order to improve the performance of road traffic vision and management systems due to the lack of resources in this specific domain. The full description of the characteristics of the dataset, as well as its components and format, can be found here: <a href="https://www.mdpi.com/2306-5729/7/5/53">https://www.mdpi.com/2306-5729/7/5/53</a></p> <p>The dataset format is YOLO. Relevant data about the dataset:</p> <table> <tbody> <tr> <td> <p><strong>Scenes</strong></p> </td> <td> <p><strong>Frames</strong></p> </td> <td> <p><strong>Targets</strong></p> </td> <td> <p><strong>Cars</strong></p> </td> <td> <p><strong>Motorbikes</strong></p> </td> </tr> <tr> <td> <p>Regional road</p> </td> <td> <p>4.500</p> </td> <td> <p>24.858</p> </td> <td> <p>14.577</p> </td> <td> <p>10.281</p> </td> </tr> <tr> <td> <p>Urban intersection</p> </td> <td> <p>2.462</p> </td> <td> <p>10.759</p> </td> <td> <p>10.759</p> </td> <td> <p>0</p> </td> </tr> <tr> <td> <p>Rural road</p> </td> <td> <p>1.292</p> </td> <td> <p>746</p> </td> <td> <p>746</p> </td> <td> <p>0</p> </td> </tr> <tr> <td> <p>Split roundabout</p> </td> <td> <p>2.297</p> </td> <td> <p>3.107</p> </td> <td> <p>3.107</p> </td> <td> <p>0</p> </td> </tr> <tr> <td> <p>Roundabout (Far)</p> </td> <td> <p>1.814</p> </td> <td> <p>71.819</p> </td> <td> <p>64.844</p> </td> <td> <p>6.975</p> </td> </tr> <tr> <td> <p>Roundabout (Near)</p> </td> <td> <p>3.997</p> </td> <td> <p>4.4039</p> </td> <td> <p>43.569</p> </td> <td> <p>470</p> </td> </tr> <tr> <td> <p><strong>Total</strong></p> </td> <td> <p><strong>15.070</strong></p> </td> <td> <p><strong>155.328</strong></p> </td> <td> <p><strong>137.602</strong></p> </td> <td> <p><strong>17.726</strong></p> </td> </tr> </tbody> </table>
K-Nearest-Neighbor algorithm to predict the survival time and classification of various stages of Oral Cancer: A machine learning approach
<p>This project predicts the survival time of a cancer patient in terms of the number of days and also classifies the dataset into various stages of cancer</p> <p>This is executed on the SPYDER platform using Python 3.7 on Anaconda Navigator.</p> <p>The dataset includes the oral cancer patient's record of 4 countries</p> <p> </p>
ESM of determination of transdermal rate of metallic microneedle array through impedance measurements-based numerical check screening algorithm
<p>Microneedle systems have been widely used in health monitoring, painless drug delivery, and medical cosmetology. Although many studies on microneedle materials, structures, and applications have been conducted, the applications of microneedles often suffered from issues of inconsistent penetration rates due to the complication of skin-microneedle interface. In this study, we demonstrated a methodology of determination of transdermal rate of metallic microneedle array through impedance measurements-based numerical check screening algorithm. Metallic sheet microneedle array sensors with different sizes were fabricated to evaluate different transdermal rates. In vitro sensing of hydrogen peroxide confirmed the effect of transdermal rate on the sensing outcomes. A FEM simulation model of a microneedle array revealed the monotonous relation between the transdermal state and test current. Accordingly, two methods were first derived to calculate the transdermal rate from the test current. First, an exact logic method provided the number of unpenetrated tips per sheet, but it required more rigorous testing results. Second, a fuzzy logic method provided an approximate transdermal rate on adjacent areas, being more applicable and robust to errors. Real-time transdermal rate estimation may be essential for improving the performance of microneedle systems, and this study provides various fundaments toward that goal.</p>
No Algorithmization without Representation: Pilot Study on Regulatory Experiments in an Exploratory Sandbox
<p>Accompanying data and scripts for the paper.</p>
Theta: portfolio of CEGAR-based analyses with dynamic algorithm selection (Competition Contribution): Tool Archive
<p>This archive contains the tool archive of Theta for SV-COMP 2022, which was also added as a release to the tool (<a href="https://github.com/ftsrg/theta/releases/tag/svcomp22-v1">here</a>).</p>
FollowMe - A Pedestrian Following Algorithm for Agricultural Logistic Robots (video results)
<p>Video result of submitted paper "FollowMe - A Pedestrian Following Algorithm for Agricultural Logistic Robots" in ICARSC confrence</p>
Per-Run Algorithm Selection with Warm-starting using Trajectory-based Features - Data
<p>This repository contains the code and data for reproducing the results in the paper 'Per-Run Algorithm Selection with<br> Warm-starting using Trajectory-based Features'</p> <p>The file 'figure_generation.ipnb' contains the code used to create the heatmaps shown in the paper.</p> <p>The file 'Raw_data' contains the raw performance data, for both running A1 and recording all relevant features and for running the switching algorithms for each of the 6 algorithms in the portfolio while recording only performance. The processed version of the A2 data into relative performance is available in 'perf_relative'.</p> <p>The file 'data-collection' contains the code used to run the switching algorithms with the used warmstarting procedures.</p> <p>The file 'compute_features.ipnyb' contains the code used to extract the timeseries features, and those features themselves are included in 'TS_Features'.</p> <p>The file 'AlgorithmSelector' contains the scripts and data used to create the algorithm selector.</p> <p>The files 'selector_5_as_f' and 'selector_ng_f' contain the results of running the final models on the test-instances and nevergrad functions respectively.</p>
Networkx graph generation algorithm sample binary graphs
<p>Example graph binary object from Networkx's random graph generation algorithms</p> <ul> <li>gnp: Erdős-Rényi graph</li> <li>gnx: Random graph</li> <li>regular: Random regular graph</li> <li>barabasi: Graph from Barabási-Albert preferential attachment model</li> </ul> <p>Parameters Used:</p> <ul> <li>P = 0.015</li> <li>N = 200</li> <li>DEGREE = 20</li> <li>SEED = 0</li> </ul>
Binomial and toric ideal data for learning a performance metric of Buchberger's algorithm
<p>This data set consists of randomly generated binomial and toric ideals. It was used for predicting a certain complexity measure of Buchberger's algorithm for toric and binomial ideals in small number of variables. See also the corresponding code on <a href="https://github.com/Sondzus/LearningGBvaluemodel">GitHub</a> and the Involve journal paper available on <a href="https://arxiv.org/abs/2106.03676">arXiv</a>, which explains in detail the models used to generate the data. </p> <p> </p> <p><strong>From the article:</strong></p> <p>What can be (machine) learned about the performance of Buchberger's algorithm?</p> <p>Given a system of polynomials, Buchberger's algorithm computes a Gr\"obner basis of the ideal these polynomials generate using an iterative procedure based on multivariate long division. The runtime of each step of the algorithm is typically dominated by a series of polynomial additions, and the total number of these additions is a hardware-independent performance metric that is often used to evaluate and optimize various implementation choices. In this work we attempt to predict, using just the starting input, the number of polynomial additions that take place during one run of Buchberger's algorithm. Good predictions are useful for quickly estimating difficulty and understanding what features make a Gr\"obner basis computation hard. Our features and methods could also be used for value models in the reinforcement learning approach to optimize Buchberger's algorithm introduced in the second author's thesis. </p> <p>We show that a multiple linear regression model built from a set of easy-to-compute ideal generator statistics can predict the number of polynomial additions somewhat well, better than an uninformed model, and better than regression models built on some intuitive commutative algebra invariants that are more difficult to compute. We also train a simple recursive neural network that outperforms these linear models. Our work serves as a proof of concept, demonstrating that predicting the number of polynomial additions in Buchberger's algorithm is a feasible problem from the point of view of machine learning.</p>
LiDAR and thermal data for camera pose estimation using the depth-map correspondence algorithm
<p>Folder and file structure:</p> <ul> <li>lidar_roi.ply : ~360 MB mesh file which is a sub-part of the whole Orlova Chuka scan collected in [1]</li> <li>yyyy-mm-dd total of ~17 GB. All video data including raw data, exported video, digitised xy points and calibration results <ul> <li>2018-08-19</li> <li>2018-08-17</li> <li>2018-08-14</li> <li>2018-07-28</li> <li>2018-07-25</li> <li>2018-07-21</li> </ul> </li> </ul> <p><em>Thermal camera YYYY-MM-DD folder substructure</em>: Each of the yyyy-mm-dd dates is one recording session. Each session folder has the following structure:</p> <ul> <li>avi_files (present on some nights)</li> <li>cave_photos: (present on some nights)</li> <li>mic_and_wall_points</li> <li>tmc_files: (present on some nights) The TMC files is a proprietary format to store thermal camera video data (TeAx GmbH, Germany). on 2018-08-17, only P0000000 is provided as it doesnt' have humans blocking the scene. Each frame can be exported to csv using the ThermoViewer tool, downloadable at: https://thermalcapture.com/thermoviewer-download/</li> <li>video_calibration: results and associated data to get DLT coefficients estimated using the easyWand [2] workflow. <ul> <li>image : csv file with pixel values of the images used for annotations</li> <li>mics : 2D point locations of mics placed on the cave walls</li> <li>other_cave_surface : other points on the cave surface that were pointed at <ul> <li>calibration_output: results from easyWand runs. Choose the highest round number <ul> <li>yyyy-mm-dd_roundX_<wandscore>_cam1Tforms.mat (undistortion files)</li> <li>yyyy-mm-dd_roundX_<wandscore>_cam2Tforms.mat</li> <li>yyyy-mm-dd_roundX_<wandscore>_cam3Tforms.mat</li> <li>yyyy-mm-dd_roundX_<wandscore>_dltCoefs.csv (each column is one camera's DLT coefficients)</li> <li>yyyy-mm-dd_roundX_<wandscore>_easyWandData.mat (easyWand session file)</li> </ul> </li> <li>gravity: (mostly there) video and xy points for a falling object to align the calbiration to gravity. Output from DLTdv7 clicking session.</li> <li>wand: video and xy points of the 'wand' calibration object. Output from DLTdv7 clicking session.</li> <li>camera_intrinsic.txt or thermalcam_camprofiles_profile.txt : the camera instrinsics</li> </ul> </li> </ul> </li> </ul> <ul> <li>alignment_results: ~887 MB zipped folder. <ul> <li>dmcp_experiments: the results of DMCP alignment <ul> <li>round_01 : <em>ignore this folder</em></li> <li>round_03 : <em>ignore this folder</em></li> <li>round_05: here yyyy-mm-dd is short for all other nights. Each yyyy-mm-dd folder has multiple csv files. The 'transform.csv' is the most relevant file, as it holds the transformation matrix to move 3D points from camera triangulations into the LiDAR coordinate system. <ul> <li>2018-07-21--cam0</li> <li>2018-07-21--cam1</li> <li>2018-07-21--cam2</li> <li>yyyy-mm-dd--cam0</li> <li>yyyy-mm-dd--cam1</li> <li>yyyy-mm-dd--cam2</li> <li>...</li> <li>...</li> <li>...</li> <li>2018-08-19--cam0</li> <li>2018-08-19--cam1</li> <li>2018-08-19--cam2</li> </ul> </li> </ul> </li> </ul> </li> </ul> <p> </p> <p>CITATION: If you use this dataset for your research please cite this Zenodo dataset and the accompanying paper.</p> <p>This uploaded dataset is part of the <em>Ushichka</em> dataset [3]. The audio-video system was designed by Holger R. Goerlitz. The LiDAR data was collected by Asparuh Kamburov. Video data collected by Thejasvi Beleyur.</p> <p>References</p> <p>[1] : Kamburov, A., Goerlitz, H. R., Beleyur, T 2018, Geospatial modelling inside the "Orlova Chuka" cave in Bulgaria, <em>non-peer reviewed conference contribution</em>, XXVIII International Symposium on Modern Technologies and Professional Practise in Geodesy and related fields</p> <p>[2]: Theriault, D. H., Fuller, N. W., Jackson, B. E., Bluhm, E., Evangelista, D., Wu, Z., M., Betke & Hedrick, T. L. (2014). A protocol and calibration method for accurate multi-camera field videography. <em>Journal of Experimental Biology</em>, <em>217</em>(11), 1843-1848.</p> <p>[3]: Beleyur Thejasvi, 2021. Theoretical and empirical investigations of echolocation in bat groups, PhD dissertation, University of Konstanz (<a href="http://nbn-resolving.de/urn:nbn:de:bsz:352-2-q41u3qlu1em03">http://nbn-resolving.de/urn:nbn:de:bsz:352-2-q41u3qlu1em03</a>)</p>
Data from: Interpreting the FLOCK algorithm from a statistical perspective
We show that the algorithm in the program FLOCK (Duchesne & Turgeon 2009) can be interpreted as an estimation procedure based on a model essentially identical to the STRUCTURE (Pritchard et al. 2000) model with no admixture and non-correlated allele frequency priors. Rather than using MCMC, the FLOCK algorithm searches for the maximum-a-posteriori estimate of this STRUCTURE model via a simulated annealing algorithm with a rapid cooling schedule (namely, the exponent on the objective function --> ∞). We demonstrate the similarities between the two programs in a two step approach. First, to enable rapid batch processing of many simulated data sets, we modified the source code of STRUCTURE to use the FLOCK algorithm, producing the program FLOCKTURE. With simulated data we confirmed that results obtained with FLOCK and FLOCKTURE are very similar (though ockture is some 200 times faster). Second, we simulated multiple large data sets under varying levels of population differentiation for both microsatellite and SNP genotypes. We analyzed them with FLOCKTURE and STRUCTURE and assessed each program on its ability to cluster individuals to their correct subpopulation. We show that FLOCKTURE yields results similar to STRUCTURE albeit with greater variability from run to run. FLOCKTURE did perform better than STRUCTURE when genotypes were composed of SNPs and differentiation was moderate (FST = 0.022 - 0.032). When differentiation was low, STRUCTURE outperformed FLOCKTURE for both marker types. On large data sets like those we simulated, it appears that FLOCK's reliance on inference rules regarding its "plateau record" are not helpful. Interpreting FLOCK's algorithm as a special case of the model in STRUCTURE should aid in understanding the program's output and behavior.
data for drought prediction using deep learning algorithms
<p>This dataset is monthly SPEI from 1901 to 2015 with 0.5 spatial resolution. It is standardized to [0,1] as inputs/outputs of neural network.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.