Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,025
datasets available to search
ShareScore release 0.7.1
Dataset results
2,025 results for “AI”
CKN Edge AI Dataset for Image inference at the Edge (CEAD)
<p>This synthetic workload models camera device requests for resource constrained inference requests at the Edge for Campaign Knowledge Network evaluation. </p> <p>The workload is a deterministic and pre-ordered set of time windows containing close to 5 million individual data points belonging to 1500 time windows, each time window with a number of requests between 100-1000. Composed of independent inference requests (events), the workload is structured to reflect sudden changes in need as reflected by the user-perceived quality of experience (e.g., accuracy and latency). </p>
Synthetic AIS Dataset of Vessel Proximity Events
<p>The Automatic Identification System (AIS) allows vessels to share identification, characteristics, and location data through self-reporting. This information is periodically broadcast and can be received by other vessels with AIS transceivers, as well as ground or satellite sensors. Since the International Maritime Organisation (IMO) mandated AIS for vessels above 300 gross tonnage, extensive datasets have emerged, becoming a valuable resource for maritime intelligence.</p> <p>Maritime collisions occur when two vessels collide or when a vessel collides with a floating or stationary object, such as an iceberg. Maritime collisions hold significant importance in the realm of marine accidents for several reasons:</p> <ol> <li>Injuries and fatalities of vessel crew members and passengers.</li> <li>Environmental effects, especially in cases involving large tanker ships and oil spills.</li> <li>Direct and indirect economic losses on local communities near the accident area.</li> <li>Adverse financial consequences for ship owners, insurance companies and cargo owners including vessel loss and penalties.</li> </ol> <p>As sea routes become more congested and vessel speeds increase, the likelihood of significant accidents during a ship's operational life rises. The increasing congestion on sea lanes elevates the probability of accidents and especially collisions between vessels.</p> <p>The development of solutions and models for the analysis, early detection and mitigation of vessel collision events is a significant step towards ensuring future maritime safety. In this context, a synthetic vessel proximity event dataset is created using real vessel AIS messages. The synthetic dataset of trajectories with reconstructed timestamps is generated so that a pair of trajectories reach simultaneously their intersection point, simulating an unintended proximity event (collision close call). The dataset aims to provide a basis for the development of methods for the detection and mitigation of maritime collisions and proximity events, as well as the study and training of vessel crews in simulator environments.</p> <p>The dataset consists of 4658 samples/AIS messages of 213 unique vessels from the Aegean Sea. The steps that were followed to create the collision dataset are:</p> <p>Given 2 vessels X (vessel_id1) and Y (vessel_id2) with their current known location (LATITUDE [lat], LONGITUDE [lon]): </p> <ol> <li>Check if the trajectories of vessels X and Y are spatially intersecting.</li> <li>If the trajectories of vessels X and Y are intersecting, then align temporally the timestamp of vessel Y at the intersect point according to X’s timestamp at the intersect point. The temporal alignment is performed so the spatial intersection (nearest proximity point) occurs at the same time for both vessels.</li> <li>Also for each vessel pair the timestamp of the proximity event is different from a proximity event that occurs later so that different vessel trajectory pairs do not overlap temporarily.</li> </ol> <p>Two csv files are provided. vessel_positions.csv includes the AIS positions vessel_id, t, lon, lat, heading, course, speed of all vessels. Simulated_vessel_proximity_events.csv includes the id, position and timestamp of each identified proximity event along with the vessel_id number of the associated vessels. The final sum of unintended proximity events in the dataset is 237. Examples of unintended vessel proximity events are visualized in the respective png and gif files.</p> <p>The research leading to these results has received funding from the European Union's Horizon Europe Programme under the CREXDATA Project, grant agreement n° 101092749. </p>
Carbon dioxide, methane, and chemical data from Batang Ai reservoir
<p>The dataset contains biogeochemical in situ field measurements taken in Batang Ai reservoir (located on the Borneo Island, Malaysia). Samples were taken over four sampling campaigns from 2016 to 2018. Data was used to analyse carbon dioxide and methane flux patterns and to calculate the carbon footprint of the reservoir in the paper: “The carbon footprint of a Malaysian tropical reservoir: measured versus modeled estimates highlight the underestimated key role of downstream processes” (<a href="https://doi.org/10.5194/bg-17-1-2020">https://doi.org/10.5194/bg-17-1-2020</a>).</p> <p>Data were also used to calculate budgets of CO2 and CH4 in the epilimnion of Batang Ai reservoir in the paper: “Changing sources and processes sustaining surface CO2 and CH4 fluxes along a tropical river to reservoir system” (<a href="https://doi.org/10.5194/bg-2020-258">https://doi.org/10.5194/bg-2020-258</a>).</p>
MultiCaRe: An open-source clinical case dataset for medical image classification and multimodal AI applications
<p>The dataset contains multi-modal data from over 70,000 open access and de-identified case reports, including metadata, clinical cases, image captions and more than 130,000 images. Images and clinical cases belong to different medical specialties, such as oncology, cardiology, surgery and pathology. The structure of the dataset allows to easily map images with their corresponding article metadata, clinical case, captions and image labels. Details of the data structure can be found in the file data_dictionary.csv.</p> <p>More than 90,000 patients and 280,000 medical doctors and researchers were involved in the creation of the articles included in this dataset. The citation data of each article can be found in the metadata.parquet file.</p> <p>Refer to the examples showcased in this <a href="https://github.com/mauro-nievoff/MultiCaRe_Dataset">GitHub repository</a> to understand how to optimize the use of this dataset.<br><br>The license of the dataset as a whole is CC BY-NC-SA. However, its individual contents may have less restrictive license types (CC BY, CC BY-NC, CC0). For instance, regarding image filess, 66K of them are CC BY, 32K are CC BY-NC-SA, 32K are CC BY-NC, and 20 of them are CC0.</p>
NeSy4VRD: A Multifaceted Resource for Neurosymbolic AI Research using Knowledge Graphs in Visual Relationship Detection
<p><strong>NeSy4VRD</strong></p> <p>NeSy4VRD is a multifaceted, multipurpose resource designed to foster neurosymbolic AI (NeSy) research, particularly NeSy research using Semantic Web technologies such as OWL ontologies, OWL-based knowledge graphs and OWL-based reasoning as symbolic components. The NeSy4VRD research resource pertains to the <em>computer vision</em> field of AI and, within that field, to the application tasks of <em>visual relationship detection (VRD) and scene graph generation</em>.</p> <p>Whilst the core motivation of the NeSy4VRD research resource is to foster computer vision-based NeSy research using Semantic Web technologies such as OWL ontologies and OWL-based knowledge graphs, AI researchers can readily use NeSy4VRD to either: 1) pursue computer vision-based NeSy research without involving Semantic Web technologies as symbolic components, or 2) pursue computer vision research without NeSy (i.e. pursue research that focuses purely on deep learning alone, without involving symbolic components of any kind). This is the sense in which we describe NeSy4VRD as being <em>multipurpose</em>: it can readily be used by diverse groups of computer vision-based AI researchers with diverse interests and objectives.</p> <p>The NeSy4VRD research resource in its entirety is distributed across two locations: Zenodo and GitHub.</p> <p> </p> <p><strong>NeSy4VRD on Zenodo: the NeSy4VRD dataset package</strong></p> <p>This entry on Zenodo hosts the <em>NeSy4VRD dataset package</em>, which includes the <em>NeSy4VRD dataset</em> and its companion <em>NeSy4VRD ontology</em>, an OWL ontology called VRD-World.</p> <p>The <em>NeSy4VRD dataset</em> consists of an image dataset with associated visual relationship annotations. The images of the <em>NeSy4VRD dataset</em> are the same as those that were once publicly available as part of the <a href="https://cs.stanford.edu/people/ranjaykrishna/vrd/">VRD</a> dataset. The NeSy4VRD visual relationship annotations are a highly customised and quality-improved version of the original VRD visual relationship annotations. The <em>NeSy4VRD dataset</em> is designed for computer vision-based research that involves detecting objects in images and predicting relationships between ordered pairs of those objects. A visual relationship for an image of the <em>NeSy4VRD dataset</em> has the form <'subject', 'predicate', 'object'>, where the 'subject' and 'object' are two objects in the image, and the 'predicate' describes some relation between them. Both the 'subject' and 'object' objects are specified in terms of bounding boxes and object classes. For example, representative annotated visual relationships are <'person', 'ride', 'horse'>, <'hat', 'on', 'teddy bear'> and <'cat', 'under', 'pillow'>.</p> <p>Visual relationship detection is pursued as a computer vision application task in its own right, and as a building block capability for the broader application task of scene graph generation. Scene graph generation, in turn, is commonly used as a precursor to a variety of enriched, downstream visual understanding and reasoning application tasks, such as image captioning, visual question answering, image retrieval, image generation and multimedia event processing.</p> <p>The <em>NeSy4VRD ontology</em>, VRD-World, is a rich, well-aligned, companion OWL ontology engineered specifically for use with the <em>NeSy4VRD dataset.</em> It directly describes the domain of the <em>NeSy4VRD dataset</em>, as reflected in the NeSy4VRD visual relationship annotations. More specifically, all of the object classes that feature in the NeSy4VRD visual relationship annotations have corresponding classes within the VRD-World OWL class hierarchy, and all of the predicates that feature in the NeSy4VRD visual relationship annotations have corresponding properties within the VRD-World OWL object property hierarchy. The rich structure of the VRD-World class hierarchy and the rich characteristics and relationships of the VRD-World object properties together give the VRD-World OWL ontology rich inference semantics. These provide ample opportunity for OWL reasoning to be meaningfully exercised and exploited in NeSy research that uses OWL ontologies and OWL-based knowledge graphs as symbolic components. There is also ample potential for NeSy researchers to explore supplementing the OWL reasoning capabilities afforded by the VRD-World ontology with Datalog rules and reasoning.</p> <p>Use of the <em>NeSy4VRD ontology</em>, VRD-World, in conjunction with the <em>NeSy4VRD dataset </em>is, of course, purely optional, however. Computer vision AI researchers who have no interest in NeSy, or NeSy researchers who have no interest in OWL ontologies and OWL-based knowledge graphs, can ignore the <em>NeSy4VRD ontology</em> and use the <em>NeSy4VRD dataset </em>by itself.</p> <p>All computer vision-based AI research user groups can, if they wish, also avail themselves of the other components of the NeSy4VRD research resource available on GitHub.</p> <p> </p> <p><strong>NeSy4VRD on GitHub: open source infrastructure supporting extensibility, and sample code</strong></p> <p>The NeSy4VRD research resource incorporates additional components that are companions to the <em>NeSy4VRD dataset package</em> here on Zenodo. These companion components are available at <a href="https://github.com/djherron/NeSy4VRD/">NeSy4VRD on GitHub</a>. These companion components consist of:</p> <ul> <li>comprehensive open source Python-based infrastructure supporting the extensibility of the NeSy4VRD visual relationship annotations (and, thereby, the extensibility of the <em>NeSy4VRD ontology</em>, VRD-World, as well)</li> <li>open source Python sample code showing how one can work with the NeSy4VRD visual relationship annotations in conjunction with the <em>NeSy4VRD ontology</em>, VRD-World, and RDF knowledge graphs.</li> </ul> <p>The NeSy4VRD infrastructure supporting extensibility consists of:</p> <ul> <li>open source Python code for conducting deep and comprehensive analyses of the <em>NeSy4VRD dataset</em> (the VRD images and their associated NeSy4VRD visual relationship annotations)</li> <li>an open source, custom-designed <em>NeSy4VRD protocol</em> for specifying visual relationship annotation customisation instructions declaratively, in text files</li> <li>an open source, custom-designed <em>NeSy4VRD workflow, </em>implemented using Python scripts and modules, for applying small or large volumes of customisations or extensions to the NeSy4VRD visual relationship annotations in a configurable, managed, automated and repeatable process.</li> </ul> <p>The purpose behind providing comprehensive infrastructure to support extensibility of the NeSy4VRD visual relationship annotations is to make it easy for researchers to take the <em>NeSy4VRD dataset</em> in new directions, by further enriching the annotations, or by tailoring them to introduce new or more data conditions that better suit their particular research needs and interests. The option to use the NeSy4VRD extensibility infrastructure in this way applies equally well to each of the diverse potential NeSy4VRD user groups already mentioned.</p> <p>The NeSy4VRD extensibility infrastructure, however, may be of particular interest to NeSy researchers interested in using the <em>NeSy4VRD ontology</em>, VRD-World, in conjunction with the <em>NeSy4VRD dataset. </em>These researchers can of course tailor the VRD-World ontology if they wish without needing to modify or extend the NeSy4VRD visual relationship annotations in any way. But their degrees of freedom for doing so will be limited by the need to maintain alignment with the NeSy4VRD visual relationship annotations and the particular set of object classes and predicates to which they refer. If NeSy researchers want full freedom to tailor the VRD-World ontology, they may well need to tailor the NeSy4VRD visual relationship annotations first, in order that alignment be maintained.</p> <p>To illustrate our point, and to illustrate our vision of how the NeSy4VRD extensibility infrastructure can be used, let us consider a simple example. It is common in computer vision to distinguish between <em>thing</em> objects (that have well-defined shapes) and <em>stuff</em> objects (that are amorphous). Suppose a researcher wishes to have a greater number of <em>stuff</em> object classes with which to work. Water is such a <em>stuff</em> object. Many VRD images contain water but it is not currently one of the annotated object classes and hence is never referenced in any visual relationship annotations. So adding a <em>Water</em> class to the class hierarchy of the VRD-World ontology would be pointless because it would never acquire any instances (because an object detector would never detect any). However, our hypothetical researcher could choose to do the following:</p> <ul> <li>use the analysis functionality of the NeSy4VRD extensibility infrastructure to find images containing water (by, say, searching for images whose visual relationships refer to object classes such as 'boat', 'surfboard', 'sand', 'umbrella', etc.);</li> <li>use free image analysis software (such as GIMP, at gimp.org) to get bounding boxes for instances of water in these images;</li> <li>use the <em>NeSy4VRD protocol</em> to specify new visual relationships for these images that refer to the new 'water' objects (e.g. <'boat', 'on', 'water'>);</li> <li>use the <em>NeSy4VRD workflow</em> to introduce the new object class 'water' and to apply the specified new visual relationships to the sets of annotations for the affected images;</li> <li>introduce class Water to the class hierarchy of the VRD-World ontology (using, say, the free Protege ontology editor);</li> <li>continue experimenting, now with the added benefit of the additional <em>stuff</em> object class 'water';</li> <li>contribute the enriched set of NeSy4VRD visual relationship annotations, and the enriched companion VRD-World ontology, to research communities.</li> </ul> <p> </p> <p><strong>Information pertaining to the VRD dataset</strong></p> <p>Information about the original VRD dataset is available <a href="https://cs.stanford.edu/people/ranjaykrishna/vrd/">here</a>. </p> <p>Public availability of the VRD images (via information accessible from that location) ceased sometime in the latter part of 2021. We thank Dr. Ranjay Krishna, one of the principals associated with the VRD dataset, for granting us permission to re-establish the public availability of the VRD images as part of NeSy4VRD.</p> <p>The original VRD visual relationship annotations are still publicly available from that location. But our deep analysis of those annotations, driven by our desire to design a robust companion ontology, revealed them to be highly problematic in many ways that made credible ontology modelling infeasible. They were also found to be replete with all manner of errors. The NeSy4VRD visual relationship annotations are far superior and we recommend them over the original VRD annotations to anyone contemplating conducting research using the VRD images. The NeSy4VRD annotations also have the added benefit of the rich, well-aligned companion <em>NeSy4VRD ontology</em>, VRD-World, for those whose research requires such a companion ontology.</p> <p>Researchers wishing to use the original VRD dataset may still do so. They can access the VRD images here, from within the <em>NeSy4VRD dataset</em> on Zenodo, and access the VRD visual relationship annotations from the location in the link.</p> <p><em>A note of caution</em>: the <em>NeSy4VRD ontology</em>, VRD-World, is <em>not</em><strong> </strong>compatible with the original VRD visual relationship annotations and cannot be used in conjunction with them. The VRD-World ontology has been engineered in relation to the highly customised and quality-improved NeSy4VRD visual relationship annotations. The customisations that were applied include ones that introduced many new object classes, merged some of the existing object classes, introduced one new predicate, and changed several predicate names.</p> <p>However, researchers can, if they wish, use the NeSy4VRD extensibility infrastructure (described above) to undertake their own customisation and quality-improvement exercise with respect to the original VRD visual relationship annotations. This is precisely how the NeSy4VRD visual relationship annotations were created in the first place. The primary intended use case of NeSy4VRD's extensibility infrastructure, however, is for researchers to use the NeSy4VRD visual relationship annotations as their starting point, and to take these annotations forward with onward customisations and extensions, as illustrated in the example use case given above.</p> <p> </p> <p> </p>
Intervju med fokusgrupp (AI och bibliotek)
<p>Pseudonymised transcript of interview with four Swedish library staff about libraries and AI.</p>
GLAB-VOD: Global L-band AI-Based Vegetation Optical Depth Dataset Based on Machine Learning and Remote Sensing
<p>GLAB VOD is a Global L-band Ai-Based vegetation optical depth dataset with 18-day temporal and 25 km spatial resolution, covering 2002 to 2020. The dataset is created using a neural network with SMOS-SMAP-INRAE-BORDEAUX (SMOSMAP-IB) VOD product as a target (over 2015-2020) and brightness temperatures (TB) from the SMOS, AMSR-E, and AMSR-2 spaceborne missions alongside with a novel soil moisture dataset (CASM) as inputs. The GLAB-VOD dataset was created using a recently developed methodology previously used to create a long-term consistent soil moisture dataset CASM, adapted to the VOD retrievals. First, the TB and VOD signals were divided into fixed seasonal cycle and residuals, where the residual part of the signal contains sub-seasonal periodic signals, trends, extremes, and noise. Then, a multi-staged neural network training scheme was used to achieve internally consistent predictions by merging data from different sources without introducing biases or compromising data distribution. A side-product of this project is GLAB TB - a global long-term brightness temperature dataset that matches SMOS TB quality and spawns back to 2002. GLAB TB has daily temporal resolution and 25 km spatial resolution. </p>
Data inputs and results from AI-supported title and abstract screening "Lack of evidence regarding markers identifying acute heart failure in patients with COPD: an AI-supported systematic review"
<p>These comma-separated data files were used to conduct the AI supported screening of [Lack of Evidence Regarding Markers Identifying Acute Heart Failure in Patients with COPD: An AI-supported Systematic Review (working title)], following the methodology described in the publication (URL/doi to be uploaded).</p> <p>These files provide insight into the AI-supported screening process and the choices made by the human reviewer.</p>
Human-AI Collaboration: A tool to enable AI model generation with human-in-the-loop
<p>Human-AI collaboration enables domain experts to contribute their expertise with the goal of enhancing the knowledge learned by the AI models from the patterns in the data. This enables the integration of domain-specific knowledge to enrich the data for further improvement of the models through retraining. The human-AI collaboration is composed of multiple sub-components and interfaces that enables communication with external systems such as data sources, model repositories, machine configurations and decision support systems.</p> <p>Human-AI Collaboration component is developed using Python programming language. The frontend is developed using Streamlit1. The backend is developed using python and the API is implemented using FastAPI2. The choice of the programming language was made because of its wide usage and vast user base. The frameworks Streamlit and FastAPI are chosen because of the rich features for functionality and documentation as well as suitability for data analysis tasks. The applications are packaged as docker images for deployment. The application runs as a web application served by nginx for reverseproxying and users can access it via client applications such as web browsers or REST clients like Postman.</p>
RGB pixels VALUES FOR APPLES/LETTUCE AI OPTICAL RECOGNITION - 5 categories of Freshness
<p>The Datasets include RGB color pallete per pixel values for optical recognition on apples/lettuce and freshness categorized using AI Algorithm . Those Datasets are for AI Training projects . It will be used on the stage of creation, verification or optimization for new optical AI models. The tables can be used direclty on the AI tools, inserted and using the pixels colors number for every category. The freshness categories are 5, from the highest- crop day (5) to the lowest - not for eating (1).</p> <p> </p>
Brick Kiln Dataset for Pakistan's IGP Region Using AI
<p>This dataset represents the first geospatial mapping of brick kiln sites in the IGP region of Pakistan, providing an invaluable resource for understanding the spatial distribution of these sites. Each data point captures a brick kiln's precise location, including coordinates, state, and other important information, standardized in Coordinate Reference System (CRS) EPSG:4326 (WGS 84). This dataset, to the best of our knowledge, is the first of its kind to consolidate and geolocate brick kiln operations across this region, where air pollution impacts from kiln emissions are a significant environmental and public health concern.</p> <p>In addition to the primary geolocation data, the dataset also includes an initial, secondary estimation of emissions (PM10, PM2.5, NOx, and SOx) from these sites. This supplementary information supports preliminary risk assessments, emphasizing proximity-based exposure for populations and sensitive areas (e.g., schools, hospitals) within a 1 km radius of each kiln site. </p> <p>The dataset is made available in multiple formats to facilitate wide usage across spatial analysis platforms:</p> <ul> <li><strong>Geojson (.geojson)</strong></li> <li><strong>Shapefile (.shp)</strong></li> <li><strong>Comma-Separated Values (.csv)</strong></li> </ul> <p><strong>Metadata:</strong></p> <ul> <li>Geographic Coverage: IGP - Pakistan</li> <li>CRS: EPSG:4326 (WGS 84)</li> </ul> <p><strong>Data Structure:</strong><br><br>"Main" files have <strong><em>Basic ID & Location </em></strong>information, while the "Emission" files have <strong><em>Basic ID & Location + Production & Emission </em></strong>information. </p> <p><strong>Basic ID & Location</strong></p> <ul> <li><code>id</code>: Unique identifier for the kiln.</li> <li><code>lat</code>: Latitude of the kiln location.</li> <li><code>lon</code>: Longitude of the kiln location.</li> <li><code>state</code>: State where the kiln is located.</li> <li><code>type</code>: Type of kiln (e.g., FCBK).</li> </ul> <p><strong>Production Data</strong></p> <ul> <li><code>avg_bricks</code>: Average number of bricks produced daily. Refer to the GitHub repository for detailed methodology.</li> <li><code>seasonprod(bricks)</code>: Seasonal production of bricks in kilograms (excluding monsoon and smog days, considering 215 operational days).</li> </ul> <p><strong>Daily Emissions</strong></p> <ul> <li><code>pm2.5d(kg)</code>: Daily PM2.5 emissions in kilograms.</li> <li><code>pm10d(kg)</code>: Daily PM10 emissions in kilograms.</li> <li><code>noxd(kg)</code>: Daily NOx emissions in kilograms.</li> <li><code>soxd(kg)</code>: Daily SOx emissions in kilograms.</li> </ul> <p><strong>Seasonal Emissions</strong></p> <ul> <li><code>pm2.5s(kg)</code>: Seasonal PM2.5 emissions in kilograms.</li> <li><code>pm10s(kg)</code>: Seasonal PM10 emissions in kilograms.</li> <li><code>noxs(kg)</code>: Seasonal NOx emissions in kilograms.</li> <li><code>soxs(kg)</code>: Seasonal SOx emissions in kilograms.</li> </ul> <p><strong>Emission Factors</strong></p> <ul> <li><code>emf_coal(pm2.5)</code>: Emission factor for PM2.5 using coal as fuel (kg).</li> <li><code>emf_coal(pm10)</code>: Emission factor for PM10 using coal as fuel (kg).</li> <li><code>emf_coal(nox)</code>: Emission factor for NOx using coal as fuel (kg).</li> <li><code>emf_coal(so2)</code>: Emission factor for SOx using coal as fuel (kg).</li> <li><code>emf_coal(pm2.5)</code>: Emission factor for PM2.5 using biomass as fuel (kg).</li> <li><code>emf_coal(pm10)</code>: Emission factor for PM10 using biomass as fuel (kg).</li> <li><code>emf_coal(nox)</code>: Emission factor for NOx using biomass as fuel (kg).</li> <li><code>emf_coal(so2)</code>: Emission factor for SOx using biomass as fuel (kg).</li> </ul> <p><strong>Seasonal Emissions by Fuel Type</strong></p> <ul> <li><code>pm2.5s_c(kg)</code>: Seasonal PM2.5 emissions in kilograms using coal as a fuel.</li> <li><code>pm10s_c(kg)</code>: Seasonal PM10 emissions in kilograms using coal as a fuel.</li> <li><code>noxs_c(kg)</code>: Seasonal NOx emissions in kilograms using coal as a fuel.</li> <li><code>so2s_c(kg)</code>: Seasonal SO2 emissions in kilograms using coal as a fuel.</li> <li><code>pm2.5s_b(kg)</code>: Seasonal PM2.5 emissions in kilograms using biomass as a fuel.</li> <li><code>pm10s_b(kg)</code>: Seasonal PM10 emissions in kilograms using biomass as a fuel.</li> <li><code>noxs_b(kg)</code>: Seasonal NOx emissions in kilograms using biomass as a fuel.</li> <li><code>so2s_b(kg)</code>: Seasonal SO2 emissions in kilograms using biomass as a fuel.</li> </ul> <p><strong>Funding Sources</strong>: This research and data collection were funded by Amazon Web Services (AWS) and Smith School of Enterprise and The Environment (University of Oxford). </p>
Concentrating solar power (CSP) plants AI-training dataset for flux density measurements.
<p>In this dataset, the tools required for the training of a neural net in the context of flux density measurements in concentrating solar power (CSP) plants are included. An Excel file with 931 meteorological conditions and the positions of the power plant and the receiver is included, as well as 15928 pairs of images resulting from ray-tracing in Solarturm Juelich (STJ) each of these conditions with 17 different combinations of heliostats. <br> <br>This dataset is part of the WP1 of TOPCSP european project (funded by HORIZON MSCA Doctoral Network, Project number 101072537).</p>
A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI)
<p><strong>A controlled vocabulary for research and innovation in the field of Artificial Intelligence (AI)</strong></p> <p>This controlled vocabulary of keywords related to the field of Artificial Intelligence (AI) was built by SIRIS Academic in collaboration with ART-ER (the R&I and sustainable development in-house agency of the Emilia-Romagna region in Italy) and the Generalitat de Catalunya (the regional government of Catalonia, Spain), in order to identify AI research, development and innovation activities. The work was carried out by consulting domain experts' advice and it was ultimately applied to inform regional strategies on AI and research and innovation policy.</p> <p>The aim of this vocabulary is to enable one to retrieve texts (e.g. R&D projects and scientific publications) featuring the concepts included in the present vocabulary in their titles and abstracts, assuming that these records have a certain contribution of applications, techniques and issues, in the domain of AI.</p> <p>The present effort was carried out because, despite the high number of contributions and technological developments in the field of AI, there is no closed or static vocabulary of concepts that allows to unequivocally define the boundaries of what should be considered “an Artificial Intelligence intellectual product” (or what should not). Indeed, the literature presents different definitions of the domain, with visions that could be contradictory. AI encompasses today a wide variety of subdomains, ranging from general purpose areas such as learning and perception to more specific ones such as autonomous vehicle driving, theorem proving, or industrial process monitoring. AI synthesises and automates intellectual tasks, and is therefore potentially relevant to any area of human intellectual activity. In this sense, it is a genuinely universal and multidisciplinary field. AI draws upon disciplines as diverse as cybernetics, mathematics, philosophy, sociology and economics.</p> <p>As a ground for the construction of the AI controlled vocabulary, an initial set of concepts was taken from different subdomains of the <em>ACM Computing Classification System 2012, </em> to define the boundaries of the AI domain. Notably, although some relevant AI subdomains have an independent category in the ACM taxonomy outside of AI, they have been included in the list of subdomains. In order to align the ACM taxonomical definition with the Catalan Strategy of AI, <em>CATALONIA.AI</em>, in <em>version 1 </em>of this resource the emerging area of AI Ethics was included in the vocabulary, while some other categories which are not relevant for the objectives were removed from the subdomains list. In the current <em>version 2</em>, the classification and the labels of the subdomains have been revised because of the evolution of the field. Some fields have been grouped in order to reduce the overlap between subdomains and to provide a taxonomy that makes more sense for the analysis of R&I ecosystems. </p> <p>The different subdomains in the versions are presented in the following table:</p> <table> <tbody> <tr> <td><strong>Version </strong></td> <td><strong>Subdomains</strong></td> </tr> <tr> <td> <p><em>Version 2</em></p> </td> <td> <p>(1) Machine learning and deep learning; (2) Computer Vision; (3) Natural Language Processing and speech recognition; (4) Intelligent agents, planning, scheduling, problem-solving, control methods, and search; (5) Expert Systems, Knowledge representation and reasoning; (6) AI Ethics.</p> </td> </tr> <tr> <td><em>Version 1</em></td> <td>(1) General, (2) Machine Learning, (3) Computer Vision, (4) Natural Language Processing, (5) Knowledge Representation and Reasoning, (6) Distributed Artificial Intelligence, (7) Expert Systems, Problem-Solving, Control Methods and Search and (8) AI Ethics.</td> </tr> </tbody> </table> <p>Although a keyword rule-based approach suffers from the major shortcomings of not capturing all the lexical and linguistic variants of specific concepts nor the context of the words - namely, keyword-based approaches would miss relevant texts if the specific pattern is not matched during the search - the present vocabulary allowed us to obtain fairly good results, due to the specificity of the concepts describing the AI domain. Furthermore, an understandable and transparent controlled vocabulary allows a better control of the final results and the final definition of the domain borders. Also, a plain list of terms allows a much easier and interactive engagement of interested stakeholders with different degrees of knowledge (such as, for instance, domain experts, policy-makers and potential users) who can make use of vocabulary to retrieve pertinent literature or to enrich the resource itself.</p> <p>The vocabulary has been built taking advantage of advanced language models and resources from knowledge datasets such as arXiv, DBpedia and Wikipedia. The resulting vocabulary comprises 833 keywords, and has been validated by experts from several universities in Emilia-Romagna and Catalonia.</p> <p>The <em>version 0.5</em> of this resource was developed by the SIRIS Academic in 2019 in collaboration with ART-ER, Emilia-Romagna (Quinquillá et <em>al.</em>, 2020), the <em>version 1 </em>was the result of an update done in 2020 in collaboration with the Generalitat de Catalunya, and the current version (<em>version 2</em>) has resulted in 2021 from the collaboration with ART-ER and the integration of an additional set of keywords provided by the <em>Artificial Intelligence and Intelligence Systems (AIIS)</em> Laboratory of the CINI (<em>Consorzio interuniversitario nazionale per l’informatica </em>based in Rome, Italy).</p> <p>The methodology for the construction of the controlled vocabulary is presented in the following steps:</p> <ol> <li> <p>An initial set of scientific publications was collected by retrieving the following records as a weakly-supervised (in the sense that records are linked to AI by their taxonomy and not by a manual label) dataset in the domain of Artificial Intelligence :</p> <ol> <li> <p>Publications from Scopus with the keyword “Artificial Intelligence”</p> </li> <li> <p>Publications from arXiv in the category “Artificial Intelligence”</p> </li> <li> <p>Publications in relevant journals in the scientific domain of “Artificial Intelligence”</p> </li> </ol> </li> <li> <p>An automated algorithm was used to retrieve, from the APIs of DBpedia, a series of terms that have some categorical relationships (i.e. those that are indexed as “sub-categories of”, “equivalent to”, among other relations in DBpedia) with the Artificial Intelligence concept and with the AI categories in the ACM taxonomy. The DBpedia tree has been exploited down to the level 3, and the relevant categories have been manually selected (for instance: <em>Classification algorithms</em>,<em> Machine learning</em> or <em>Evolutionary computation</em>) and others were ignored (for instance: <em>Artificial intelligence in fiction</em>, <em>Robots</em> or <em>History of artificial intelligence</em>) because they were not relevant, or not specifically in the domain.</p> </li> <li> <p>The keywords in publications in the dataset were extracted from the keyword sections and from the abstracts. The keywords with a higher <em>TF-IDF</em>, using an <em>IDF</em> matrix in the open domain, have been selected. The co-occurrence of keywords with categories in specific AI subdomain and a clusterization of the main keywords has been used for a categorization of the keywords at the thematic level.</p> </li> <li> <p>This list of keywords tagged by thematic category has been manually revised, removing the non-pertinent keywords and changing the wrong categorizations by fields.</p> </li> <li> <p>The weak-supervised dataset in the domain of Artificial Intelligence is used to train a Word2Vec (Mikolov <em>et al.</em>, 2013) word embedding model (a machine learning model based on neural networks).</p> </li> <li> <p>The terms’ list is then enriched by means of automatic methods, which are run in parallel: </p> <ol> <li> <p>The trained Word2Vec model is used to select, among the indexed keywords of the reference corpus, all terms “semantically close” to the initial set of words. This step is carried out to select terms that might not appear in the texts themselves, but that were deemed pertinent to label the textual records.</p> </li> <li> <p>Further, terms that are mentioned in the texts of the reference corpus and that are valued by the trained Word2Vec model as “semantically close” to the initial set of words are also retained. This step is performed to include in the controlled vocabulary a series of terms that are related to the focus of the SDGs and which are used by practitioners.</p> </li> </ol> </li> <li> <p>The final list produced by steps 2-6 is manually revised.</p> </li> </ol> <p> </p> <p>The definition of the vocabulary does not, per se, allow to identify STI contributions to AI: this activity in fact boils down to actually matching the terms in the controlled vocabulary to the content of the gathered STI textual records. To successfully carry out this task, a series of pattern matching rules must be defined to capture possible variants of the same concept, such as permutations of words within the concept and/or the presence of null words to be skipped. For this reason, we have carefully crafted matching rules that take into account permutations of words and that allow words within concept to be within a certain distance. Some relatively ambiguous keywords (which may match unwanted pieces of text), have a set of associated “extra” terms. These “extra” terms are defined as further terms that must co-appear, in the same sentence, together with their associated ambiguous keywords.</p> <p>Finally, each keyword in the vocabulary was assigned one or more AI subdomains, so that the vocabulary can also be used to tag collections of texts within narrower AI sub-domains. In order to complement the alignment between keywords and subdomains, a set of subdomain-specific keywords have been defined to better capture the scope of the subdomains. These allow better characterization of subdomains that are more difficult to define only by means of unambiguous specific concepts, or that overlap with the wide “machine learning” subdomain (example: machine learning applied to object recognition or text translation). The alignment between keywords and subdomains, and these keyword lists of each subdomain, have been applied to capture AI subdomains in research outputs. Through this classification process, we have identified projects and publications related to AI, with a focus on mapping the research competencies in the AI domain in Emilia-Romagna. The resulting research records have been reviewed by experts in the domain, given the occurrence of some false positives, which have been used to improve the approach.</p> <p>The final controlled vocabulary has been evaluated with an external test set, proposed by (Dunham <em>et al.,</em> 2020). The test set consists of the abstract of 10,606 papers published in the arXiv repository, of which 1,076 within the Artificial Intelligence subcategories and 9,530 in arXiv categories other than Artificial Intelligence. Evaluating the controlled vocabulary on this data set, we observe accuracy of .94. However, because the pertinence of these publications to the field of AI is based solely on their taxonomic classification (i.e., on whether they are classified in the arXiv within Artificial Intelligence and not on a manual labelling), this evaluation can only yield an orientative performance assessment.</p> <p>The version 2 includes new keywords extracted from the (1) re-training of the enrichment pipeline (steps 5-6 in the methodology) considering as initial set of terms the version 1 of the vocabulary on a reference corpus of new publications, and (2) from the flat keywords list provided by the <em>Artificial Intelligence and Intelligence Systems</em> <em>(AIIS)</em> Lab of CINI (Consorzio interuniversitario nazionale per l’informatica). The keywords in (2) have been cleaned by calculating precision and f-measure on the dataset (Dunham et al., 2020), selecting those keywords with the highest scores, and being manually validated a posteriori.</p> <p>The AI controlled vocabulary has been applied in two practical cases, which have the purpose of identifying skills, stakeholders and capabilities, of a specific research ecosystem at the regional level. See the following references:</p> <ul> <li> <p>Quinquillá, Arnau, Duran-Silva, Nicolau, Massucci, Francesco Alessandro, Fuster, Enric, Rondelli, Bernardo, Bologni, Leda, … Moretti, Giorgio. (2020). Text mining to identify skills, stakeholders and capabilities: the case of Artificial Intelligence in Emilia-Romagna. Zenodo. <a href="http://doi.org/10.5281/zenodo.3606342">http://doi.org/10.5281/zenodo.3606342</a>. Poster presented at: World Open Innovation Conference 2019 (WOIC); 11th december 2019, Rome, Italy.</p> </li> <li> <p>Bigas, E., Duran, N., Fuster, E., Parra, C., Fernández, T. (2021): “Anàlisi de l’especialització en intel·ligència artificial”. Col·lecció Monitoratge de la RIS3CAT, Generalitat de Catalunya <a href="http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf">http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf</a></p> </li> </ul> <p> </p> <p><strong>Acknowledgements</strong></p> <ul> <li> <p>Tatiana Fernández (Direcció General de Promoció Econòmica, Competència i Regulació, de la Generalitat de Catalunya), </p> </li> <li> <p>Daniel Marco, Daniel Santanach and Eduard Balbuena (Departament de Polítiques Digitals i Administració Pública, de la Generalitat de Catalunya) </p> </li> <li> <p>Albert Sabater (Observatori d’Ètica en Intel·ligència Artificial i Universitat de Girona)</p> </li> <li> <p>Leda Bologni, Lucia Mazzoni and Giorgio Moretti (Art-ER)</p> </li> <li> <p>Prof. RIta Cucchiara and Dr. Lorenzo Baraldi (Università degli Studi di Modena e Reggio Emilia)</p> </li> <li> <p>Artificial Intelligence and Intelligence Systems (AIIS) Lab of CINI (Consorzio interuniversitario nazionale per l’informatica)</p> </li> </ul> <p> </p> <p><strong>Bibliography</strong></p> <p>Bigas, E., Duran, N., Fuster, E., Parra, C., Fernández, T. (2021): “Anàlisi de l’especialització en intel·ligència artificial”. Col·lecció Monitoratge de la RIS3CAT, Generalitat de Catalunya <a href="http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf">http://catalunya2020.gencat.cat/web/.content/00_catalunya2020/Documents/estrategies/fitxers/analisi-especialitzacio-intelligencia-artificial.pdf</a></p> <p>Dunham, J.W., Melot, J., & Murdick, D. (2020). Identifying the Development and Application of Artificial Intelligence in Scientific Text. ArXiv, abs/2002.07143. Available at: <a href="https://arxiv.org/abs/2002.07143">https://arxiv.org/abs/2002.07143</a></p> <p>Mikolov, Tomas & Corrado, G.s & Chen, Kai & Dean, Jeffrey. (2013). Efficient Estimation of Word Representations in Vector Space. 1-12.</p> <p>Quinquillá, Arnau, Duran-Silva, Nicolau, Massucci, Francesco Alessandro, Fuster, Enric, Rondelli, Bernardo, Bologni, Leda, … Moretti, Giorgio. (2020). Text mining to identify skills, stakeholders and capabilities: the case of Artificial Intelligence in Emilia-Romagna. Zenodo. <a href="http://doi.org/10.5281/zenodo.3606342">http://doi.org/10.5281/zenodo.3606342</a>. Poster presented at: World Open Innovation Conference 2019 (WOIC); 11th december 2019, Rome, Italy.</p>
AI-TAM: a model to investigate user acceptance and collaborative intention in human-in-the-loop AI applications
<p>More and more frequently, digital applications make use of Artificial Intelligence (AI) capabilities<br> to provide advanced features; on the other hand, human-in-the-loop approaches are on the<br> rise to involve people in AI-powered pipelines for data collection, results validation and decision making.<br> Does the introduction of AI features affect user acceptance? Does the AI result quality<br> affect people’s willingness to use such applications? Does the additional user effort required in<br> human-in-the-loop mechanisms change the application adoption and use?<br> This study aims to provide a reference approach to answer those questions. We propose a model<br> that extends the Technology Acceptance Model (TAM) with further constructs explicitly related to<br> AI – user trust in AI and perceived quality of AI output, from explainable AI (XAI) literature – and<br> collaborative intention – willingness to contribute to AI pipelines.<br> We tested the proposed model with an application for car damage claim reporting with AI-powered<br> damage estimation for insurance customers. The results showed that the XAI related factors have<br> a strong and positive effect on behavioral intention, perceived usefulness, and ease of use of the<br> application. Moreover, there is a strong link between behavioral intention and collaborative intention,<br> indicating that indeed human-in-the-loop approaches can be successfully adopted in final user<br> applications.</p> <p>Users were invited to test the interactive prototype of the BumpOut application and to report the given car accident from start to finish. These are the two interactive prototypes experienced by users:</p> <ul> <li> <p><a href="https://bit.ly/bo-prototype-flawlessAI">FlawlessAI-Group prototype</a></p> </li> <li> <p><a href="https://bit.ly/bo-prototype-failingAI">FailingAI-Group prototype</a></p> </li> </ul> <p> </p> <p>This study is shared as a research object adopting the <a href="https://www.researchobject.org/ro-crate/1.0/">RO-Crate</a> specification.</p>
Experimental AI corpus from OpenAlex
<p>A corpus of AI research from OpenAlex. Includes:</p> <ul> <li>A works table with metadata about AI papers</li> <li>An authors table with information about the authors</li> <li>An institutions table with information about institutions</li> <li>A concepts table with information about concepts in works</li> <li>A MeSH table with information about MeSH terms in works</li> <li>A concepts json with the OpenAlex concept taxonomy</li> <li>An abstracts json with deinverted abstracts</li> <li>A citations json with citations from papers</li> </ul> <p>See `ai_openalex_description.md` for data dictionaries.</p> <p>See `ai_openalex_methodology.md` for a description of the method used to create the dataset.</p> <p>See here for additional information: <a href="https://github.com/nestauk/ai_genomics">https://github.com/nestauk/ai_genomics</a></p>
Explainable AI for unveiling deep learning pollen classification model - Pollen dataset
<p>Dataset consists automatic particle detector Rapid-E measurements of pollen grains from 12 classes: Acer, Alnus, Alopecurus, Carex, Cupressus, Dactylis, Juglans, Morus, Platanus, Populus, Salix and Ulmus. Data are available i json format.</p> <p>Dataset also contains preprocessed data packed into csv files of 3 modalities: spectrum, lifetime, scattering and additional features from scattering and lifetime data are also available. These are ready to be used with machine learning models. Labels 0, 1, 2, ... 11 correspond to alphabetical order of examined pollen classes Acer, Alnus, Alopecurus ... Ulmus.</p> <p> </p>
AI results complementing the 2021 Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Sweden
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2021, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
AI results complementing the 2021 Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Portugal
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2021, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
AI results complementing the 2021 Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Lithuania
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2021, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
AI results complementing the 2021 Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - United Kingdom (Northern Ireland)
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2021, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund</li> </ul> <p>Disclaimer: In accordance with the Agreement on the Withdrawal of the United Kingdom from the EU, and in particular with the Protocol on IE/NI, the EU requirements on data sampling are also applicable to Northern Ireland.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.