Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,291

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,291 results for “artificial intelligence”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data files for figures in "Sensitivity of extreme precipitation to climate change inferred using artificial intelligence shows high spatial variability" by Bird et al.

<p>The data files for figures in&nbsp;<i>Sensitivity of extreme precipitation to climate change inferred using artificial intelligence shows high spatial variability</i> by Bird, Bodeker and Clem. The files required to create each figure in the paper and in the supplementary material are described in a readme.txt file which is also provided below:</p><p><strong>Figure 1</strong></p><p>The background image was obtained from the 'NaturalEarthFeature' function of the python Cartopy library (Figure1_background.png). The data required to generate the plots shown in Figure 1 are provided in the Figure1.nc file:</p><ul><li>The latitudes and longitudes for the 10,000 training sites are provided in Training_location_latitudes and Training_location_longitudes variables. &nbsp;</li><li>The latitudes and longitudes for the 8 sites used to demonstrate the ability of the CNN to generalise spatially are provided in the Validation_location_latitudes and Validation_location_longitudes variables. &nbsp;</li><li>The block maxima at each of the 8 sites are provided in the Location_1year_block_maxima variable.</li><li>The GEV fits at 0°C are provided in the GEV_fit_at_0.0C variable.</li><li>The GEV fits at 1.5°C are provided in the GEV_fit_at_1.5C variable.</li></ul><p><strong>Figure 2</strong></p><ul><li>The 1-in-100 year precipitation fields for values of T'Global from 0.0°C to 2.0°C in 0.1°C increments are provided in netCDF files named Figure2_&lt;region&gt;.nc where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'.</li><li>The ratios of the 1-in-100 year precipitation depths with respect to the long-term (20 years or more) average of the annual maximum 1-day precipitation are provided in text files named Figure2_Precipitation_mean_block_max_&lt;region&gt;.txt where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'.</li></ul><p><strong>Figure 3</strong></p><ul><li>The cumulative distribution functions (CDFs) shown in the lower four panels are provided as text files listing the ARI in years and the daily total precipitation depth in mm. These files are named CDF_&lt;lat&gt;_&lt;long&gt;.dat where &lt;lat&gt; is the latitude and &lt;long&gt; is the longitude. Files for each region are zipped into .7z files named Figure3_&lt;region&gt;.7z where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'.</li><li>The latitudes and longitudes for the upper panels can be inferred from the file names for each region.</li></ul><p><strong>Figure 4</strong></p><p>The data for each panel are provided in a netCDF file named Figure4_&lt;region&gt;.nc where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'.</p><p><strong>Figure 5</strong></p><p>The data are provided as text files named Figure5_&lt;region&gt; where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'. Each file lists the sensitivities at ARIs of 10, 20, 50, 100, and 200 years.</p><p><strong>Figure 6</strong></p><p>The data are provided as text files named Figure6_&lt;region&gt; where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'. Each file lists the precipitation depths and the sensitivities at ARIs of 10, 20, 50, 100, and 200 years. &nbsp;</p><p><strong>Figure 7</strong></p><p>The data for each panel are provided in a netCDF file named Figure7_&lt;region&gt;.nc where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'.</p><p><strong>Figure 8</strong></p><p>The data are provided as text files named Figure8_&lt;region&gt;.txt where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'. Each file lists the global surface temperature anomaly (°C) and the average negative log likelihood.</p><p><strong>Figure S1 and Figure S3</strong></p><p>The data are provided as text files named Figure_S1_and_S3_&lt;region&gt;.txt where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'. A header at the top of each file describes its contents.</p><p><strong>Figure S2</strong></p><ul><li>The 1-in-20 year precipitation fields for values of T'Global from 0.0°C to 2.0°C in 0.1°C increments are provided in netCDF files named FigureS2_&lt;region&gt;.nc where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'.</li><li>The ratios of the 1-in-20 year precipitation depths with respect to the long-term (20 years or more) average of the annual maximum 1-day precipitation are provided in text files named FigureS2_Precipitation_mean_block_max_&lt;region&gt;.txt where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'.</li></ul><p><strong>Figure S4</strong></p><ul><li>The block maxima for each site are provided in text files named FigureS4_blockmaxima_siteA.txt and FigureS4_blockmaxima_siteB.txt. A header at the top of each column described the column contents. &nbsp;</li><li>The GEV-derived curves for each site are provided in text files named FigureS4_gevcurves_siteA.txt and FigureS4_gevcurves_siteB.txt. A header at the top of each column described the column contents. &nbsp;</li></ul><p><strong>Figure S5</strong></p><p>There are no data associated with this figure. This figure was made using Microsoft Powerpoint.</p><p><strong>Figure S6</strong></p><p>The data are provided as text files named FigureS6_&lt;region&gt;.txt where &lt;region&gt; is either 'north_america', 'australia', 'europe', or 'new_zealand'. A header at the top of each file describes the files contents.</p>

opencc-by-4.0Oct 2023View details →
zenodo48/100

Artificial Intelligence for Quality Control of manufacturing operations: Macro-mechanical milling in the Pilot Line GAMHE 5.0.

<p>Quality is defined as the extent to which a product conforms to the design specifications and how it complies with the requirements of component functionality. For some industries, such as automotive and aeronautical, the quality of their parts is very important given the high requirements to which they are subject. However, difficulties arise from the fact that a measure of quality can only be evaluated &lsquo;&lsquo;out-of-process&rdquo;, resulting in losses because there is no alternative to removing defective parts from the production line. Therefore, it is necessary to apply Artificial Intelligence-based kits/solutions that provide in-process estimation to predict quality from some measured variables.&nbsp;</p> <p>The main goal of these datasets is to monitor the final quality of the manufactured components or parts by estimating surface roughness from vibration signals and cutting parameters information using Artificial Intelligence-based solutions. Surface roughness is an essential feature in quality control defined by the deviation in the direction of the normal vector of a real surface from its ideal form. Because the roughness measurement is an offline and post process procedure, being able to estimate this value online brings a series of benefits in terms of time and cost reduction in manufacturing lines, energy efficiency, unnecessary wear of tools and machines, etc. Once a part has been detected with a surface quality below what is desired, a series of corrective measures can be applied for the following operations, such as: reducing the feed rate percentage, increasing the percentage of spindle speed or reducing the axial depth per pass, etc.</p>

opencc-by-4.0Oct 2021View details →
zenodo48/100

A living catalogue of artificial intelligence datasets and benchmarks for medical decision making

<p>We provide&nbsp;a comprehensive curated catalogue of&nbsp;<strong>artificial intelligence datasets</strong> and <strong>benchmarks for medical decision making</strong>. At the time of first release (April 2021), the dataset contains more than 400&nbsp;biomedical and clinical datasets&nbsp;of which 252 are publicly available or available upon request.</p> <p>The dataset was compiled based on a systematic literature review covering both biomedical and computer science literature and&nbsp;grey literature data sources. All datasets were manually systematized and annotated for meta-information, such as:</p> <ul> <li>Availability and licensing information</li> <li>Type of source data</li> <li>Links to source publications, main references or dataset repositories</li> </ul> <p>Benchmark dataset were additionally annotated for the following information:</p> <ul> <li>Associated task</li> <li>Performance metrics commonly used for evaluation</li> <li>Clinical relevance</li> <li>The availability of data splits</li> </ul> <p>In addition to the versioned TSV file on Zenodo, the dataset can also be explored live via&nbsp;<a href="https://docs.google.com/spreadsheets/d/1QjUxxnZ3tuyW5dj6nkt_o5yJcWUZec4ttfJxO8Zlty4/edit?usp=sharing">this Google Spreadsheet</a>.&nbsp;The dataset is intended as a living, extendable resource. Edit suggestions and additions are encouraged and can be submitted via the comment function of the Google sheet.</p> <p>&nbsp;</p> <p><strong>File descriptions</strong></p> <p><em>annotated-datasets.tsv</em> -- contains the annotated datasets</p> <p><em>arXiv-literature-export.tsv</em> -- contains the original literature record export from arXiv</p> <p><em>pubmed-literature-export.tsv</em> -- contains the original literature record export from PubMed</p> <p><em>README.md</em> -- contains a detailed description of all annotation fields</p>

opencc-by-sa-4.0Apr 2021View details →
zenodo48/100

Artificial Intelligence for quality control in manufacturing operations: Micro-mechanical milling in the Pilot Line GAMHE 5.0

<p>Quality is defined as the extent to which a product conforms to the design specifications and how it complies with the requirements of component functionality. For some industries, such as automotive and aeronautical, the quality of of manufactured parts is very important due to the high requirements. However, difficulties arise from the fact that a measure of quality can only be evaluated &lsquo;&lsquo;out-of-process&rdquo;, resulting in losses because there is no alternative to removing defective parts from the production line. Therefore, it is necessary to incorporate AI-based kits/solutions that provide in-process estimation to predict quality from some measured variables.</p> <p>The main goal of these datasets is to enable monitoring of final quality of the manufactured components or parts by estimating surface roughness from vibration signals and cutting parameters information. Surface roughness is an essential feature in quality control defined by the deviation in the direction of the normal vector of a real surface from its ideal form. Because the roughness measurement is an offline and post process procedure, being able to estimate this value online brings a series of benefits in terms of time and cost reduction in manufacturing lines, energy efficiency, unnecessary wear of tools and machines, etc. Once a part has been detected with a surface quality below what is desired, a series of corrective measures can be applied for the following operations, such as: reducing the feed rate percentage, increasing the percentage of spindle speed or reducing the axial depth per pass, etc.</p> <p>Workstation 4 (WS4) of the GAMHE 5.0 pilot line is a Kern Evo high-precision machining centre, with a maximum spindle speed of 50 000 rpm and Blum laser system and is used to run micro-milling and micro-drilling operations. In this experimental dataset, five cutting parameters were considered in the processes: spindle speed, <em>n</em>; feed rate, <em>f</em>; and axial depth of cut, <em>a<sub>P</sub></em>. The radial depth of cut, <em>a<sub>e</sub></em>; was equal to the mill tool radius, <em>r</em>, in all of the slots.</p> <p>These experiments were micro-milling operations with 0.3 mm, 0.5 mm, 0.8 mm and 1 mm-diameter mills on a sintered tungsten-copper alloy (W78Cu22). The data collected for each micro milling operation was the rms and peak value of the vibrations in the three-machine axis. In addition, five cutting parameters were also collected: position in <em>X</em> of the last point of the sample, feed rate, spindle speed, tool radius and axial depth.</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Artificial Intelligence for EU Decision-Making: Effects on Citizens Perceptions of Input, Throughput and Output Legitimacy

<p>The uploaded dataset was used for the statistical analysis of the pre-print &quot;Artificial Intelligence for EU Decision-Making: Effects on Citizens&rsquo; Perceptions of Input, Throughput and Output Legitimacy&quot; (Permanent identifier: <a href="https://arxiv.org/abs/2003.11320">arXiv:2003.11320</a>)</p> <p>A lack of political legitimacy undermines the ability of the European Union (EU) to resolve major crises and threatens the stability of the system as a whole. By integrating digital data into political processes, the EU seeks to base decision-making increasingly on sound empirical evidence. In particular, artificial intelligence (AI) systems have the potential to increase political legitimacy by identifying pressing societal issues, forecasting potential policy outcomes, informing the policy process, and evaluating policy effectiveness. This paper investigates how citizens&rsquo; perceptions of EU input, throughput, and output legitimacy are influenced by three distinct decision-making arrangements: (1) independent human decision-making (HDM); (2) independent algorithmic decision-making (ADM) by AI-based systems; and (3) hybrid decision-making by EU politicians and AI-based systems together. The results of a pre-registered online experiment (n = 572) suggest that existing EU decision-making arrangements are still perceived as the most democratic (input legitimacy). However, regarding the decision-making process itself (throughput legitimacy) and its policy outcomes (output legitimacy), no difference was observed between the status quo and hybrid decision-making involving both ADM and democratically elected EU institutions. Where ADM systems are the sole decision-maker, respondents tend to perceive these as illegitimate. The paper discusses the implications of these findings for (a) EU legitimacy and (b) data-driven policy-making.</p>

opencc-by-4.0Mar 2020View details →
zenodo44/100

1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test using artificial intelligence and image-based pattern recognition techniques

<p><strong>The Great Isaiah Scroll (1QIsa<sup>a</sup>) data set for writer identification</strong></p> <p>This data set is collected for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497</p> <p>Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> <br> <strong>Copyright (c) </strong>&nbsp;&nbsp; &nbsp;University of Groningen, 2021. All rights reserved.<br> <strong>Disclaimer and copyright notice for all data contained on this .tar.gz file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the data for research purposes. It is not allowed to distribute this data for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this data.</p> <p><strong>4) </strong>the user should refer to the first public article on this data set:<br> <br> <em>Popović, M., Dhali, M. A., &amp; Schomaker, L. (2020). Artificial intelligence-based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.</em><br> <br> BibTeX:</p> <pre>@article{popovic2020artificial, title={Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)}, author={Popovi{\&#39;c}, Mladen and Dhali, Maruf A and Schomaker, Lambert}, journal={arXiv preprint arXiv:2010.14476}, year={2020} }</pre> <p><strong>5) </strong>the recipient should refrain from proliferating the data set to third parties external to his/her local research group. Please refer interested researchers to this site for obtaining their own copy.</p> <p><strong>Organisation of the data:</strong></p> <p>The .tar.gz file contains three directories: images, features, and plots. The included &#39;README&#39; file contains all the instructions.</p> <p>The &#39;images&#39; directory contains NetPBM images of the columns of 1QIsa<sup>a</sup>. The NetPBM format is chosen because of its simplicity. Additionally, there is no doubt about lossy compression in the processing chain. There are two images for each of the Great Isaiah Scroll columns: one is the direct binarized output from the BiNet (<em>arxiv.org/abs/1911.07930</em>) system, and the other one is the manually cleaned version of the binarized output. &nbsp; The file names for the direct binarized output are of the format &#39;1QIsaa_col&lt;columnnr&gt;.pbm&#39;, for example, &#39;1QIsaa_col15.pbm&#39;. And, for the cleaned version, the format is &#39;1QIsaa_col&lt;columnnr&gt;_cleaned.pbm&#39;, for example, &#39;1QIsaa_col15_cleaned.pbm&#39;. Note: the image files are not in a separate directory; they will be extracted in the same place. However, due to the unique naming, there is no problem extracting them in one single directory.</p> <p>The &#39;features&#39; directory contains feature files computed for each of the column images. There are two types of feature files: Hinge and Adjoined. They are distinguishable by their extension, for example, &#39;1QIsaa_col15_cleaned.hinge&#39; and &#39;1QIsaa_col15_cleaned.adjoined&#39;. They are also arranged in separate directories for ease of use.</p> <p>The &#39;plots&#39; directory contains a simple python script to perform PCA on the feature files and then visualize them in a 3D plot. The file takes the location of feature files as an input. The &#39;README_plot&#39; file contains examples of how-to-run in the terminal.</p> <p><strong>Brief description:</strong><br> According to ImageMagick&#39;s&#39; identify&#39; tool, the original images are in grayscale (.jpg) from Brill collection, in &#39;8-bit Gray 256c&#39;. &nbsp;These images pass through multiple preprocessing measures to become suitable for pattern recognition-based techniques. The first step in preprocessing is the image-binarization technique. In order to prevent any classification of the text-column images based on irrelevant background patterns, a specific binarization technique (BiNet) was applied, keeping the original ink traces intact. After performing the binarization, the images were cleaned further by removing the adjacent columns that partially appear on the target columns&#39; images. Finally, few minor affine transformations and stretching corrections were performed in a restrictive manner. These corrections are also targeted for aligning the texts where the text lines get twisted due to the leather writing surface&#39;s degradation. Hence, the clean images are there in the directory along with the direct binarized images. No effort has been made to obtain a balanced set in any way.</p> <p><strong>Tools:</strong><br> <strong>Binarization:</strong><br> The BiNet tool is available for scientific use upon request (m.a.dhal(at)rug.nl)</p> <p><strong>Image Morphing:</strong><br> In the original article, data augmentation was performed using image morphing. The tool is available on GitHub:<br> https://github.com/GrHound/imagemorph.c</p> <p><strong>Features for writer identification:</strong><br> Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> <em><strong>1.&nbsp;</strong>L. Schomaker &amp; M. Bulacu (2004). Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.<br> <strong>2. </strong>Bulacu, M. &amp; Schomaker, L.R.B. (2007). Text-independent Writer Identification and Verification Using Textural and Allographic Features, &nbsp;IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</em><br> &nbsp;<br> The features (hinge, fraglets) have been combined in a single MS Windows application, GIWIS, which is available for scientific use upon request (l.r.b.schomaker(at)rug.nl)</p> <p><strong>If you have any question, please contact us:</strong><br> Maruf A. Dhali &lt;m.a.dhali(at)rug.nl&gt;<br> Lambert Schomaker &lt;l.r.b.schomaker(at)rug.nl&gt;<br> Mladen Popović &lt;m.popovic(at)rug.nl&gt;</p> <p><strong>Please cite our papers if you use this data set:</strong><br> <em><strong>1.</strong> Popović, M., Dhali, M. A., &amp; Schomaker, L. (2020). Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.<br> <strong>2. </strong>Dhali, M. A., de Wit, J. W., &amp; Schomaker, L. (2019). Binet: Degraded-manuscript binarization in diverse document textures and layouts using deep encoder-decoder networks. arXiv preprint arXiv:1911.07930.</em></p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Artificial Intelligence Identifies Individuals with Prediabetes from Single-Lead Electrocardiograms

<h2>Contents</h2> <ul> <li><strong>codes.zip</strong> <ul> <li>For ECG feature extraction (this will need original ECG signal data) <ul> <li>ecg_feature_extraction.sh</li> <li>ecg_feature_extraction.py</li> <li>feature_extractor.py</li> </ul> </li> <li>For training with hyperparameter optimization <ul> <li>train.sh</li> <li>train.py</li> </ul> </li> <li>For prediction of prediabetes/diabetes from ECG feature <ul> <li>test.sh</li> <li>test.py</li> </ul> </li> </ul> </li> <li><strong>raw_ecg_data.zip</strong>: 16,766 ECG records used in our analyses. Each record is a 5,000 x 12 matrix in a CSV file. &nbsp;(In this dataset, value 1 represents&nbsp;4.88 &micro;V.)</li> <li><strong>external_ecg_data.zip</strong>: 2,456 ECG records used in our external validation. Each record is a 5,000 x 12 matrix in a CSV file. (In this dataset, value 1 represents 1 &micro;V.)</li> <li><strong>participant_characteristics.csv</strong>: Health check records of 16,766 participants where the information below are stored. <ul> <li>participant_id: IDs for participant. Some IDs are duplicated because the dataset contains multiple records from some of the participants.</li> <li>ecg_id: IDs for ECG records, all of which are unique</li> <li>age: The age of each participant at the time of the health checkup</li> <li>male_sex: If the participant is male, "True" is recorded</li> <li>smoking: if the participant smokes, "True" is recorded</li> <li>drinking: 1 for "rarely", 2 for "occasionally" and 3 for "regularly" is recorded according to the frequency of drinking</li> <li>height: participant's height in centimeters (cm)</li> <li>weight: participant's body weight in kilograms (kg)</li> <li>BMI: body mass index, calculated using the formula: weight (kg) / [height (m)]^2</li> <li>pulse_rate:&nbsp; pulse rate in pulse per minute (/min)&nbsp;&nbsp;</li> <li>sBP: systolic blood pressure in mmHg</li> <li>dBP: diastolic blood pressure in mmHg</li> <li>FPG: fasting plasma glucose levels measured in milligrams per deciliter (mg/dL)</li> <li>HbA1c: hemoglobin A1c levels in %</li> <li>dm_under_treatment: if the participant was undergoing treatment for known diabetes, "True" is recorded</li> <li>prediabetes_diabetes: classification label which is "True" if a participant meet either of the following criteria <ul> <li>FPG &ge; 110 mg/dL</li> <li>HbA1c &ge; 6.0%</li> <li>Undergoing treatment for diabetes</li> </ul> </li> <li>development_data: "True" in records used as development data in our study</li> </ul> </li> <li><strong>external_cohort_characteristics.csv</strong>: Health check records of 2,456 participants where the information below are stored. <ul> <li>ecg_id: IDs for ECG records, all of which are unique</li> <li>prediabetes_diabetes: classification label which is "True" if a participant meet either of the following criteria <ul> <li>FPG &ge; 110 mg/dL</li> <li>HbA1c &ge; 6.0%</li> <li>Undergoing treatment for diabetes</li> </ul> </li> <li>FPG: fasting plasma glucose levels measured in milligrams per deciliter (mg/dL)</li> <li>HbA1c: hemoglobin A1c levels in %</li> <li>dm_under_treatment: if the participant was undergoing treatment for known diabetes, "True" is recorded</li> </ul> </li> <li><strong>ecg_feature_data.zip</strong>: extracted ECG features (unprocessed), for 12-lead and 1-lead ECG <ul> <li>ecg_features_1-lead.csv&nbsp; &nbsp; [Single-lead (lead I) ECG]</li> <li>ecg_features_12-lead.csv&nbsp; [12-lead ECG]</li> <li>ecg_features_12-leads_external_cohort.csv &nbsp;[12-lead ECG of external cohort]</li> </ul> </li> <li><strong>feature_list.zip</strong>:&nbsp;List of ECG features used (to be used for ECG extraction for original data) <ul> <li>feature_list_269_12-lead.csv&nbsp; &nbsp;[269 features for 12-lead ECG analysis]</li> <li>feature_list_28_1-lead.csv &emsp;&nbsp; &nbsp;[28 features for single-lead (lead I) analysis]</li> </ul> </li> <li><strong>model_12-lead.zip, model_1-lead.zip</strong>: model trained with our 12-lead or single-lead (lead I) ECG data, and the classification thresholds, used for test<br> <ul> <li>model_fold_1.pkl - model_fold_10.pkl : model for each of 10-fold cross validation</li> <li>average_threshold.pkl : classification threshold, which is the average of 10-fold</li> </ul> </li> </ul> <p>Codes and data for demo are also available in (https://github.com/dkoga4116/diabetes_detector)</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Artificial Intelligence Enables Precision Diagnosis of Cervical Cytology Grades and Cervical Cancer

<p>This repository includes source data used to genrtate all tables and figures&nbsp; for published stduy "Artificial Intelligence Enables Precision Diagnosis of Cervical Cytology Grades and Cervical Cancer". Besides, a small set of digital images for different class of cervical smear samples are included.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Data and code associated with "The Observed Availability of Data and Code in Earth Science and Artificial Intelligence"

<p>Data and code associated with "The Observed Availability of Data and Code in Earth Science&nbsp;<br>and Artificial Intelligence" by Erin A. Jones, Brandon McClung, Hadi Fawad, and Amy McGovern.</p> <p>Instructions: To reproduce figures, download all associated Python and CSV files and place<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; in a single directory.<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Run BAMS_plot.py as you would run Python code on your system.</p> <p>Code:<br>BAMS_plot.py: Python code for categorizing data availability statements based on given data<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; documented below and creating figures 1-3.&nbsp;</p> <p>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Code was originally developed for Python 3.11.7 and run in the Spyder&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; (version 5.4.3) IDE.<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Libraries utilized:<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; numpy &nbsp; &nbsp; &nbsp;&nbsp; (version 1.26.4)&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; pandas &nbsp; &nbsp; &nbsp; (version 2.1.4)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; matplotlib &nbsp;(version 3.8.0)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; For additional documentation, please see code file.</p> <p>Data:<br>ASDC_AIES.csv: &nbsp; &nbsp; &nbsp;CSV file containing relevant availability statement data for Artificial&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Intelligence for the Earth Systems (AIES)<br>ASDC_AI_in_Geo.csv: CSV file containing relevant availability statement data for Artificial&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Intelligence in Geosciences (AI in Geo.)<br>ASDC_AIJ.csv: &nbsp; &nbsp; &nbsp; CSV file containing relevant availability statement data for Artificial&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Intelligence (AIJ)<br>ASDC_MWR.csv: &nbsp; &nbsp; &nbsp; CSV file containing relevant availability statement data for Monthly&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Weather Review (MWR)<br><br></p> <p><br>Data documentation:<br>All CSV files contain the same format of information for each journal. The CSV files above are&nbsp;<br>needed for the BAMS_plot.py code attached.</p> <p>Records were analyzed based on the criteria below.</p> <p>&nbsp; Records:<br>&nbsp; &nbsp; 1) Title of paper<br>&nbsp; &nbsp; &nbsp; &nbsp; The title of the examined journal article.<br>&nbsp; &nbsp; 2) Article DOI (or URL)<br>&nbsp; &nbsp; &nbsp; &nbsp; A link to the examined journal article. For AIES, AI in Geo., MWR, the DOI is&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; generally given. For AIJ, the URL is given.<br>&nbsp; &nbsp; 3) Journal name<br>&nbsp; &nbsp; &nbsp; &nbsp; The name of the journal where the examined article is published. Either a full<br>&nbsp; &nbsp; &nbsp; &nbsp; journal name (e.g., Monthly Weather Review), or the acronym used in the&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; associated paper (e.g., AIES) is used.<br>&nbsp; &nbsp; 4) Year of publication<br>&nbsp; &nbsp; &nbsp; &nbsp; The year the article was posted online/in print.<br>&nbsp; &nbsp; 5) Is there an ASDC?<br>&nbsp; &nbsp; &nbsp; &nbsp; If the article contains an availability statement in any form, "yes" is&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; recorded. Otherwise, "no" is recorded.<br>&nbsp; &nbsp; 6) Justification for non-open data?<br>&nbsp; &nbsp; &nbsp; &nbsp; If an availability statement contains some justification for why data is not&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; openly available, the justification is summarized and recorded as one of the&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; following options: 1) Dataset too large, 2) Licensing/Proprietary, 3) Can be&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; obtained from other entities, 4) Sensitive information, 5) Available at later&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; date. If the statement indicates any data is not openly available and no&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; justification is provided, or if no statement is provided is provided "None"&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; is recorded. If the statement indicates openly available data or no data&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; produced, "N/A" is recorded.<br>&nbsp; &nbsp; 7) All data available<br>&nbsp; &nbsp; &nbsp; &nbsp; If there is an availability statement and data is produced, "y" is recorded&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; if means to access data associated with the article are given and there is no&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; indication that any data is not openly available; "n" is recorded if no means&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; to access data are given or there is some indication that some or all data is&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; not openly available. If there is no availability statement or no data is&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; produced, the record is left blank.<br>&nbsp; &nbsp; 8) At least some data available<br>&nbsp; &nbsp; &nbsp; &nbsp; If there is an availability statement and data is produced, "y" is recorded&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; if any means to access data associated with the article are given; "n" is&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; recorded if no means to access data are given. If there is no availability&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; statement or no data is produced, the record is left blank.<br>&nbsp; &nbsp; 9) All code available<br>&nbsp; &nbsp; &nbsp; &nbsp; If there is an availability statement and data is produced, "y" is recorded&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; if means to access code associated with the article are given and there is no&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; indication that any code is not openly available; "n" is recorded if no means&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; to access code are given or there is some indication that some or all code is&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; not openly available. If there is no availability statement or no data is&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; produced, the record is left blank.<br>&nbsp; &nbsp; 10) At least some code available<br>&nbsp; &nbsp; &nbsp; &nbsp; If there is an availability statement and data is produced, "y" is recorded&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; if any means to access code associated with the article are given; "n" is&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; recorded if no means to access code are given. If there is no &nbsp;availability&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; statement or no data is produced, the record is left blank.<br>&nbsp; &nbsp; 11) All data available upon request<br>&nbsp; &nbsp; &nbsp; &nbsp; If there is an availability statement indicating data is produced and no data&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; is openly available, "y" is recorded if any data is available upon request to&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; the authors of the examined journal article (not a request to any other&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; entity); "n" is recorded if no data is available upon request to the authors&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; of the examined journal article. If there is no availability statement, any&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; data is openly available, or no data is produced, the record is left blank.<br>&nbsp; &nbsp; 12) At least some data available upon request<br>&nbsp; &nbsp; &nbsp; &nbsp; If there is an availability statement indicating data is produced and not all&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; data is openly available, "y" is recorded if all data is available upon&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; request to the authors of the examined journal article (not a request to any&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; other entity); "n" is recorded if not all data is available upon request to&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; the authors of the examined journal article. If there is no availability&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; statement, all data is openly available, or no data is produced, the record<br>&nbsp; &nbsp; &nbsp; &nbsp; is left blank.<br>&nbsp; &nbsp; 13) no data produced<br>&nbsp; &nbsp; &nbsp; &nbsp; If there is an availability statement that indicates that no data was<br>&nbsp; &nbsp; &nbsp; &nbsp; produced for the examined journal article, "y" is recorded. Otherwise, the<br>&nbsp; &nbsp; &nbsp; &nbsp; record is left blank.<br>&nbsp; &nbsp; 14) links work<br>&nbsp; &nbsp; &nbsp; &nbsp; If the availability statement contains one or more links to a data or code&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; repository, "y" is recorded if all links work; "n" is recorded if one or more&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; links do not work. If there is no availability statement or the statement&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; does not contain any links to a data or code repository, the record is left&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; blank.&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Can Artificial Intelligence help in the study of vegetative growth dynamics from herbarium collections? An evaluation of the tropical flora of the French Guiana forest

<p>Dataset was used for the article &quot;Can Artificial Intelligence help in the study of vegetative growth dynamics from herbarium collections? An evaluation of the tropical flora of the French Guiana forest&quot;.</p> <p>The related work proposes to study to what extent the use of automated visual analysis techniques, based on deep learning, can help not only to detect relatively rare vegetative structures in herbarium collections but also to automatically classify them by type of growing shoot (continuous or rhythmic).</p> <p>Abstract of the paper:</p> <p>A better knowledge of tree vegetative growth patterns and their relationship to environmental variables is crucial in understanding forest growth dynamics and how climate change may affect them. Generally less studied than reproductive structures, the phenology of tree vegetative growth mainly focuses on the analysis of growing shoots, from vegetative buds development to leaf fall. This growth process usually strongly differs between temperate and tropical regions. In temperate regions, this pattern is quite well known. Low winter temperatures impose a stop of the vegetative growth shoots and lead to the typical expression of an annual growth cycle for the vast majority of tree species. In moist tropical regions, on the other hand, the seasonality is much less marked. In addition, these regions contain a much wider variety of tree species. These two aspects lead to a tremendous diversity of phenological patterns that are still poorly known and understood. In particular, not much is known on the periodicity and timing of growth at individual trees, population, or community levels.</p> <p>The work carried out in this study aims to advance knowledge in this area, focusing more particularly on herbarium scans, as herbarium collections offer the promise of monitoring plant phenology over long time periods. However, such a study requires the ability to detect a sufficiently large number of growing shoots in herbarium collections to draw statistically relevant conclusions, which can be very costly if the work is done manually. Furthermore, herbarium collections traditionally focus on reproductive organs, and herbarium specimens showing growing shoots are pretty rare.</p> <p>We propose in this paper to study to what extent the use of automated visual analysis techniques, based on deep learning, can help not only to detect these relatively rare vegetative structures in herbarium collections but also to automatically classify them by type of growing shoot (continuous or rhythmic). Our results show the relevance of using herbarium data for vegetative phenology research, as well as the potential of deep learning approaches for growth shoot detection.</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI)

<p><strong>A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI)</strong></p> <p>This controlled vocabulary of keywords related to the field of Artificial Intelligence (AI) was built by SIRIS Academic in collaboration with ART-ER (the R&amp;I and sustainable development in-house agency of the Emilia-Romagna region in Italy) and the Generalitat de Catalunya (the regional government of Catalonia, Spain), in order to identify AI research, development and innovation activities. The work was carried out by consulting domain experts&#39; advice and it was ultimately applied to inform regional strategies on AI and research and innovation policy.</p> <p>The aim of this vocabulary is to enable one to retrieve texts (e.g. R&amp;D projects and scientific publications) featuring the concepts included in the present vocabulary in their titles and abstracts, assuming that these records have a certain contribution of applications, techniques and issues, in the domain of AI.</p> <p>The present effort was carried out because, despite the high number of contributions and technological developments in the field of AI, there is no closed or static vocabulary of concepts that allows to unequivocally define the boundaries of what should be considered &ldquo;an Artificial Intelligence intellectual product&rdquo; (or what should not). Indeed, the literature presents different definitions of the domain, with visions that could be contradictory. AI encompasses today a wide variety of subdomains, ranging from general purpose areas such as learning and perception to more specific ones such as autonomous vehicle driving, theorem proving, or industrial process monitoring. AI synthesises and automates intellectual tasks, and is therefore potentially relevant to any area of human intellectual activity. In this sense, it is a genuinely universal and multidisciplinary field. AI draws upon disciplines as diverse as cybernetics, mathematics, philosophy, sociology and economics.</p> <p>As a ground for the construction of the AI controlled vocabulary, an initial set of concepts was taken from different subdomains of the <em>ACM Computing Classification System 2012, </em>&nbsp;to define the boundaries of the AI domain. Notably, although some relevant AI subdomains have an independent category in the ACM taxonomy outside of AI, they have been included in the list of subdomains. In order to align the ACM taxonomical definition with the Catalan Strategy of AI, <em>CATALONIA.AI</em>, in <em>version 1 </em>of this resource the emerging area of AI Ethics was included in the vocabulary, while some other categories which are not relevant for the objectives were removed from the subdomains list. In the current <em>version 2</em>, the classification and the labels of the subdomains have been revised because of the evolution of the field. Some fields have been grouped in order to reduce the overlap between subdomains and to provide a taxonomy that makes more sense for the analysis of R&amp;I ecosystems.&nbsp;</p> <p>The different subdomains in the versions are presented in the following table:</p> <table> <tbody> <tr> <td><strong>Version&nbsp;&nbsp; </strong></td> <td><strong>Subdomains</strong></td> </tr> <tr> <td> <p><em>Version 2</em></p> </td> <td> <p>(1)&nbsp; Machine learning and deep learning; (2)&nbsp; Computer Vision; (3)&nbsp; Natural Language Processing and speech recognition; (4)&nbsp; Intelligent agents, planning, scheduling, problem-solving, control methods, and search; (5)&nbsp; Expert Systems, Knowledge representation and reasoning; (6)&nbsp; AI Ethics.</p> </td> </tr> <tr> <td><em>Version 1</em></td> <td>(1) General, (2) Machine Learning, (3) Computer Vision, (4) Natural Language Processing, (5) Knowledge Representation and Reasoning, (6) Distributed Artificial Intelligence, (7) Expert Systems, Problem-Solving, Control Methods and Search and (8) AI Ethics.</td> </tr> </tbody> </table> <p>Although a keyword rule-based approach suffers from the major shortcomings of not capturing all the lexical and linguistic variants of specific concepts nor the context of the words -&nbsp; namely, keyword-based approaches would miss relevant texts if the specific pattern is not matched during the search - the present vocabulary allowed us to obtain fairly good results, due to the specificity of the concepts describing the AI domain. Furthermore, an understandable and transparent controlled vocabulary allows a better control of the final results and the final definition of the domain borders. Also, a plain list of terms allows a much easier and interactive engagement of interested stakeholders with different degrees of knowledge (such as, for instance, domain experts, policy-makers and potential users) who can make use of vocabulary to retrieve pertinent literature or to enrich the resource itself.</p> <p>The vocabulary has been built taking advantage of advanced language models and resources from knowledge datasets such as arXiv, DBpedia and Wikipedia. The resulting vocabulary comprises 833 keywords, and has been validated by experts from several universities in Emilia-Romagna and Catalonia.</p> <p>The <em>version 0.5</em> of this resource was developed by the SIRIS Academic in 2019 in collaboration with ART-ER, Emilia-Romagna (Quinquill&aacute; et <em>al.</em>, 2020), the <em>version 1 </em>was the result of an update done&nbsp; in 2020 in collaboration with the Generalitat de Catalunya, and the current version (<em>version 2</em>) has resulted&nbsp; in 2021 from the collaboration with ART-ER and the integration of an additional set of keywords provided by the <em>Artificial Intelligence and Intelligence Systems (AIIS)</em> Laboratory of the CINI (<em>Consorzio interuniversitario nazionale per l&rsquo;informatica </em>based in Rome, Italy).</p> <p>The methodology for the construction of the controlled vocabulary is presented in the following steps:</p> <ol> <li> <p>An initial set of scientific publications was collected by retrieving the following records as a weakly-supervised (in the sense that records are linked to AI by their taxonomy and not by a manual label) dataset in the domain of Artificial Intelligence :</p> <ol> <li> <p>Publications from Scopus with the keyword &ldquo;Artificial Intelligence&rdquo;</p> </li> <li> <p>Publications from arXiv in the category &ldquo;Artificial Intelligence&rdquo;</p> </li> <li> <p>Publications in relevant journals in the scientific domain of &ldquo;Artificial Intelligence&rdquo;</p> </li> </ol> </li> <li> <p>An automated algorithm was used to retrieve, from the APIs of DBpedia, a series of terms that have some categorical relationships (i.e. those that are indexed as &ldquo;sub-categories of&rdquo;,&nbsp; &ldquo;equivalent to&rdquo;, among other relations in DBpedia) with the Artificial Intelligence concept and with the AI categories in the ACM taxonomy. The DBpedia tree has been exploited down to the level 3, and the relevant categories have been manually selected (for instance: <em>Classification algorithms</em>,<em> Machine learning</em> or <em>Evolutionary computation</em>) and others were ignored (for instance: <em>Artificial intelligence in fiction</em>, <em>Robots</em> or <em>History of artificial intelligence</em>) because they were not relevant, or not specifically in the domain.</p> </li> <li> <p>The keywords in publications in the dataset were extracted from the keyword sections and from the abstracts. The keywords with a higher <em>TF-IDF</em>, using an <em>IDF</em> matrix in the open domain, have been selected. The co-occurrence of keywords with categories in specific AI subdomain and a clusterization of the main keywords has been used for a categorization of the keywords at the thematic level.</p> </li> <li> <p>This list of keywords tagged by thematic category has been manually revised, removing the non-pertinent keywords and changing the wrong categorizations by fields.</p> </li> <li> <p>The weak-supervised dataset in the domain of Artificial Intelligence is used to train a Word2Vec (Mikolov <em>et al.</em>, 2013) word embedding model (a machine learning model based on neural networks).</p> </li> <li> <p>The terms&rsquo; list is then enriched by means of automatic methods, which are run in parallel: &nbsp;&nbsp;&nbsp;</p> <ol> <li> <p>The trained Word2Vec model is used to select, among the indexed keywords of the reference corpus, all terms &ldquo;semantically close&rdquo; to the initial set of words. This step is carried out to select terms that might not appear in the texts themselves, but that were deemed pertinent to label the textual records.</p> </li> <li> <p>Further, terms that are mentioned in the texts of the reference corpus and that are valued by the trained Word2Vec model as &ldquo;semantically close&rdquo; to the initial set of words are also retained. This step is performed to include in the controlled vocabulary a series of terms that are related to the focus of the SDGs and which are used by practitioners.</p> </li> </ol> </li> <li> <p>The final list produced by steps 2-6 is manually revised.</p> </li> </ol> <p>&nbsp;</p> <p>The definition of the vocabulary does not, per se, allow to identify STI contributions to AI: this activity in fact boils down to actually matching the terms in the controlled vocabulary to the content of the gathered STI textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance. Some relatively ambiguous keywords (which may match unwanted pieces of text), have a set of associated &ldquo;extra&rdquo; terms. These &ldquo;extra&rdquo; terms are defined as further terms that must co-appear, in the same sentence, together with their associated ambiguous keywords.</p> <p>Finally, each keyword in the vocabulary was assigned one or more AI subdomains, so that the vocabulary can also be used to tag collections of texts within narrower AI sub-domains.&nbsp; In order to complement the alignment between keywords and subdomains, a set of subdomain-specific keywords have been defined to better capture the scope of the subdomains. These allow better characterization of subdomains that are more difficult to define only by means of unambiguous specific concepts, or that overlap with the wide &ldquo;machine learning&rdquo; subdomain (example: machine learning applied to object recognition or text translation). The alignment between keywords and subdomains, and these keyword lists of each subdomain, have been applied to capture AI subdomains in research outputs. Through this classification process, we have identified projects and publications related to AI, with a focus on mapping the research competencies in the AI domain in Emilia-Romagna. The resulting research records have been reviewed by experts in the domain, given the occurrence of some false positives, which have been used to improve the approach.</p> <p>The final controlled vocabulary has been evaluated with an external test set, proposed by (Dunham <em>et al.,</em> 2020). The test set consists of the abstract of 10,606 papers published in the arXiv repository, of which 1,076 within the Artificial Intelligence subcategories and 9,530 in arXiv categories other than Artificial Intelligence. Evaluating the controlled vocabulary on this data set, we observe accuracy of .94. However, because the pertinence of these publications to the field of AI is based solely on their taxonomic classification (i.e., on whether they are classified in the arXiv within Artificial Intelligence and not on a manual labelling), this evaluation can only yield an orientative performance assessment.</p> <p>The version 2 includes new keywords extracted from the (1) re-training of the enrichment pipeline (steps 5-6 in the methodology) considering as initial set of terms the version 1 of the vocabulary on a reference corpus of new publications, and (2) from the flat keywords list provided by the <em>Artificial Intelligence and Intelligence Systems</em> <em>(AIIS)</em> Lab of CINI (Consorzio interuniversitario nazionale per l&rsquo;informatica). The keywords in (2) have been cleaned by calculating precision and f-measure on the dataset (Dunham et al., 2020), selecting those keywords with the highest scores, and being manually validated a posteriori.</p> <p>The AI controlled vocabulary has been applied in two practical cases, which have the purpose of identifying skills, stakeholders and capabilities, of a specific research ecosystem at the regional level. See the following references:</p> <ul> <li> <p>Quinquill&aacute;, Arnau, Duran-Silva, Nicolau, Massucci, Francesco Alessandro, Fuster, Enric, Rondelli, Bernardo, Bologni, Leda, &hellip; Moretti, Giorgio. (2020). Text mining to identify skills, stakeholders and capabilities: the case of Artificial Intelligence in Emilia-Romagna. Zenodo. <a href="http://doi.org/10.5281/zenodo.3606342">http://doi.org/10.5281/zenodo.3606342</a>. Poster presented at: World Open Innovation Conference 2019 (WOIC); 11th december 2019, Rome, Italy.</p> </li> <li> <p>Bigas, E., Duran, N., Fuster, E., Parra, C., Fern&aacute;ndez, T. (2021): &ldquo;An&agrave;lisi de l&rsquo;especialitzaci&oacute; en intel&middot;lig&egrave;ncia artificial&rdquo;. Col&middot;lecci&oacute; Monitoratge de la RIS3CAT, Generalitat de Catalunya <a href="http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf">http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf</a></p> </li> </ul> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <ul> <li> <p>Tatiana Fern&aacute;ndez (Direcci&oacute; General de Promoci&oacute; Econ&ograve;mica, Compet&egrave;ncia i Regulaci&oacute;, de la Generalitat de Catalunya),&nbsp;</p> </li> <li> <p>Daniel Marco, Daniel Santanach and Eduard Balbuena (Departament de Pol&iacute;tiques Digitals i Administraci&oacute; P&uacute;blica, de la Generalitat de Catalunya)&nbsp;</p> </li> <li> <p>Albert Sabater (Observatori d&rsquo;&Egrave;tica en Intel&middot;lig&egrave;ncia Artificial i Universitat de Girona)</p> </li> <li> <p>Leda Bologni, Lucia Mazzoni and Giorgio Moretti (Art-ER)</p> </li> <li> <p>Prof. RIta Cucchiara and Dr. Lorenzo Baraldi (Universit&agrave; degli Studi di Modena e Reggio Emilia)</p> </li> <li> <p>Artificial Intelligence and Intelligence Systems (AIIS) Lab of CINI (Consorzio interuniversitario nazionale per l&rsquo;informatica)</p> </li> </ul> <p>&nbsp;</p> <p><strong>Bibliography</strong></p> <p>Bigas, E., Duran, N., Fuster, E., Parra, C., Fern&aacute;ndez, T. (2021): &ldquo;An&agrave;lisi de l&rsquo;especialitzaci&oacute; en intel&middot;lig&egrave;ncia artificial&rdquo;. Col&middot;lecci&oacute; Monitoratge de la RIS3CAT, Generalitat de Catalunya <a href="http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf">http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf</a></p> <p>Dunham, J.W., Melot, J., &amp; Murdick, D. (2020). Identifying the Development and Application of Artificial Intelligence in Scientific Text. ArXiv, abs/2002.07143. Available at: <a href="https://arxiv.org/abs/2002.07143">https://arxiv.org/abs/2002.07143</a></p> <p>Mikolov, Tomas &amp; Corrado, G.s &amp; Chen, Kai &amp; Dean, Jeffrey. (2013). Efficient Estimation of Word Representations in Vector Space. 1-12.</p> <p>Quinquill&aacute;, Arnau, Duran-Silva, Nicolau, Massucci, Francesco Alessandro, Fuster, Enric, Rondelli, Bernardo, Bologni, Leda, &hellip; Moretti, Giorgio. (2020). Text mining to identify skills, stakeholders and capabilities: the case of Artificial Intelligence in Emilia-Romagna. Zenodo. <a href="http://doi.org/10.5281/zenodo.3606342">http://doi.org/10.5281/zenodo.3606342</a>. Poster presented at: World Open Innovation Conference 2019 (WOIC); 11th december 2019, Rome, Italy.</p>

opencc-by-sa-4.0Feb 2021View details →
zenodo44/100

Classification of Artificial Intelligence and eXplainable Artificial Intelligence publications in Air Traffic Management

<p>v1.0 version used and partially published in &quot;A Survey on Artificial Intelligence (AI) and eXplainable AI in Air Traffic Management: Current Trends and Development with Future Research Trajectory&quot;. In this version, it references mainly Transportation Reasearch Part C, ICRAT, Journal of ATM, and ATM Seminar, IEEE transaction on ITS, but not only</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Global Artificial Intelligence in Banking Market 2024–2033

<p><a href="https://www.custommarketinsights.com/report/artificial-intelligence-in-banking-market/" target="_blank" rel="noopener">Artificial Intelligence in Banking Market Size</a>, Trends and Insights By Component (Service, Solution), By Application (Fraud Detection and Prevention, Transaction Monitoring, Identity Verification, Customer Service, Virtual Assistants, Automated Customer Support, Risk Management, Credit Scoring, Market Risk Analysis, Personalized Banking, Customer Recommendations, Targeted Marketing, Compliance and Regulatory Reporting, Anti-Money Laundering (AML), Know Your Customer (KYC), Others), By Technology (Machine Learning, Supervised Learning, Unsupervised Learning, Reinforcement Learning, Natural Language Processing (NLP), Text Analysis, Speech Recognition, Chatbots and Virtual Assistants, Robotic Process Automation (RPA), Process Automation, Workflow Automation, Predictive Analytics, Risk Management, Customer Insights), By Enterprise Size (Large Enterprise, SMEs), and By Region - Global Industry Overview, Statistical Data, Competitive Analysis, Share, Outlook, and Forecast 2024&ndash;2033</p> <p><strong>VC investments in AI by country</strong></p> <table> <tbody> <tr> <td><strong>Country</strong></td> <td><strong>VC Investment</strong></td> </tr> <tr> <td><strong>US</strong></td> <td><strong>54,836</strong></td> </tr> <tr> <td><strong>China</strong></td> <td><strong>18,270</strong></td> </tr> <tr> <td><strong>EU</strong></td> <td><strong>7,921</strong></td> </tr> </tbody> </table> <p>For more details <strong>DOWNLOAD FREE SAMPLE</strong> Now at <a href="https://www.custommarketinsights.com/request-for-free-sample/?reportid=58189" target="_blank" rel="noopener">https://www.custommarketinsights.com/request-for-free-sample/?reportid=58189</a></p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Demonstration of the GOLDEN Artificial Intelligence (AI) GUI - Artificial Intelligence Platform for mine site monitoring (Open Pit Extraction, Valea Sesei and Roșia Poieni (Romania)).

<p>Demonstration of the GOLDEN Artificial Intelligence (AI) GUI - Artificial Intelligence Platform for mine site monitoring in the&nbsp;Open Pit Extraction (mine located at Valea Sesei and Roșia Poieni (Romania)) (3D view mode).</p> <p>Accessing the GOLDENAI GUI, please refer to the following link&nbsp;(<strong>login required</strong>): <a href="https://next-gui.goldenai.opt-net.eu/ ">https://next-gui.goldenai.opt-net.eu/&nbsp;</a></p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Demonstration of the GOLDEN Artificial Intelligence (AI) GUI - Artificial Intelligence Platform for mine site monitoring (Underground Extraction, Pyhäsalmi (Finland)).

<p>Demonstration of the GOLDEN Artificial Intelligence (AI) GUI - Artificial Intelligence Platform for mine site monitoring in the&nbsp;Underground Extraction (mine located at Pyh&auml;salmi (Finland)) (2D view mode).</p> <p>Accessing the GOLDENAI GUI, please refer to the following link&nbsp;(<strong>login required</strong>): <a href="https://next-gui.goldenai.opt-net.eu/ ">https://next-gui.goldenai.opt-net.eu/&nbsp;</a></p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Generative artificial intelligence predicts human performance

<p>Research data for a study that used generative artificial intelligence (i.e., ChatGPT with the GPT-4 and the Google Gemini 2.0 Flash models) to predict human performance in a language-based memory task. In particular, we studied the effects of context on the relatedness and memorability of garden-path sentences.&nbsp;</p>

opencc-by-4.0Feb 2023View details →
dryad40/100

Data from: Artificial intelligence enabled multi-purpose smart detection in active-matrix digital microfluidics

<p>Active-matrix digital microfluidics (AM-DMF), integrated with hundreds of thousands of active electrodes, can simultaneously realize multiple on-chip bio-chemical reactions at the single-cell level. An intelligent detection system is critical for fully automating manipulations of thousands of digitalized bio-samples and programming the subsequent experiments in real time. In this work, we developed a series of deep learning algorithms based on an AM-DMF system for sample detections. We used the U-net model to quantitatively evaluate different splitting methods on sample droplet generation uniformity. The results revealed that droplets generated using the "one-to-two" strategy exhibits optimal uniformity. We used the YOLOv5 model to monitor the droplet splitting success rates over 18 different AM-DMF chips, and a 97.7% splitting success rate was observed. The results indicated that the model precision was 99.980% and the model recall was 99.976% through manual verification. In addition, we used an improved YOLOv8 model to detect single cells in nanoliter droplets effectively. In comparison with manual verification, the results showed that the model achieved a precision of 99.260% and a recall of 99.193%. By leveraging an artificial intelligence enabled smart detection system, AM-DMF has shown great potential as a ubiquitous platform for true lab-on-a-chip.</p>

opencc-zeroOct 2023View details →
zenodo40/100

Canada's national artificial intelligence governance system: Dataset from interviews with 20 government leaders & subject matter experts

<p><strong>Summary</strong></p> <p>Anonymized aggregate data from interviews with 20 government leaders and subject matter experts. The data was collected as part of a study of Canada's national system of artificial intelligence governance. The data was collected from February 2023 to July 2023. The dataset contains 610 topics that emerged from thematic analysis of interview transcripts from July 2023 to October 2023. The contexts, actors, resources, networks, evaluations, logics, functional bounds, rules, ecosystem-level dynamics, opportunities for improvement, and other topics contained in the dataset collectively represent the most significant components of Canada's national AI governance system that emerged over the course of the interviews with the 20 participants.</p> <p>&nbsp;</p> <p><strong>Notes for interpreting this dataset</strong></p> <p>Topics in analytical dimensions 1, 3, and 6-11 contain counts of the frequency with which aggregate topics emerged across each of the interviews with the 20 participants. Topics in analytical dimensions 2, 4, and 5 contain categories instead of frequency counts: the topics in these dimensions represent every unique actor, resource, and network that emerged over the course of the interviews instead of aggregate topics.&nbsp;</p> <p>Column titles contain the following abbreviations:<br>LEAD: Interviews with leaders of public sector AI governance initiatives.<br>SME-PS: Interviews with subject matter experts employed in the private sector.<br>SME-CS: Interviews with subject matter experts employed in the academic or civil sectors.</p> <p>&nbsp;</p> <p><strong>Full report</strong></p> <p>A report containing more information about this dataset and about the findings of our study can be found on SSRN: <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4783525">https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4783525</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Supporting data for the AI education publication statistics in "An Experience Report of Executive-Level Artificial Intelligence Education in the United Arab Emirates"

<p>Supporting data for the AI education publication statistics presented in the paper &quot;An Experience Report of Executive-Level Artificial Intelligence Education in the United Arab Emirates&quot; to be published at the Twelfth AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-22). The data was used to plot the figure showing the cumulative number of publications from 1976 to 2020 relating to AI education.</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Multivocal Literature Review & Interviews: Continuous End-to-End Lifecycle Management Pipeline for Artificial Intelligence

<p>During a multivocal literature review 151 relevant formal and informal sources were extracted. Additional information was collected from nine semi-structured interviews.</p> <p>The extracted dataset provides the foundation for the presentation and comparison of different terminologies for DevOps and CI/CD for AI, MLOps, (end-to-end) lifecycle management, and CD4ML. Furthermore, the dataset comprises potential triggers for reiterating the pipeline and consolidated tasks necessary in a continuous end-to-end lifecycle pipeline categorized into four stages: Data, Model, Dev and Ops. Additionally, the dataset provides information regarding challenges of the lifecycle pipelines for AI.</p>

opencc-by-4.0Feb 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record