Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,025
datasets available to search
ShareScore release 0.7.1
Dataset results
2,025 results for “AI”
AI results complementing the 2021 Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Slovakia
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2021, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
AI results complementing the 2021 Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Norway
<p>This dataset contains the results of the surveillance activities conducted in 2021, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
AI results complementing the 2021 Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Iceland
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2021, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
Story Map of the AI Ethics Lab of the Austrian Institute of Technology (AIT)
<p>The Co-Change Lab at AIT, the Austrian Institute of Technology, focuses on addressing the promises and challenges associated with research work on and the application of machine learning and artificial intelligence. An interdisciplinary team of social and data scientists is working on AI ethics.</p>
Trypanosoma Epitope Dataset: Valid Epitopes and Randomly Generated Peptides with Biochemical Metrics and AI-Generated Scores
<p>This dataset contains information about valid linear B-Cell epitopes from the Trypanosoma genus, as well as randomly generated peptides. It includes biochemical metrics generated by the EpiBuilder-1.0 tool and scores generated by the BepiPred-3.0 software. The data was originally collected from the IEDB and UniProtKB platforms and has been processed and enhanced with these informations for researchers interested in understanding the molecular interactions between Trypanosoma protozoans and the immune system.</p>
SD4EO: AI-based synthetic satellite multispectral agricultural textures in Spain (Oct 2017 - Sep 2018)
<p>This dataset has been created as part of the deliverables for ESA’s <a title="https://eo4society.esa.int/projects/sd4eo/" href="https://eo4society.esa.int/projects/sd4eo/" target="_blank" rel="noopener">SD4EO project.</a> It consists of textures generated using a multispectral variant of a still unpublished high-order statistical constraint synthesis method for each of the following crop types:</p> <ul> <li> Barley.</li> <li> Wheat.</li> <li> Other grain leguminous.</li> <li> Peas.</li> <li> Fallow & Bare soil.</li> <li> Vetch.</li> <li> Alfalfa.</li> <li> Sunflower.</li> <li> Oats.</li> </ul> <p>The initial data was sampled from satellite images, specifically from Copernicus’ Sentinel-1 and Sentinel-2 satellites. The images were acquired over a period from October 2017 to September 2018 on the central-east region of northern Spain (Castile and León and Catalonia). From these images, the corresponding crops were extracted and used as samples for assembling large puzzles that have been applied as input reference images to generate the synthetic images that make up this dataset.</p> <p>The datasets of assembled crop field "puzzles" used as reference images combine the largest crop areas to create a square multispectral texture of the largest possible size that is a power of 2 (or nearly a power of 2). Each base image combines data from all available Sentinel-2 satellite passes for the same month and a previous monthly composition from Sentinel-1. Due to cloud masks influence, the shape and number of crops vary for each time sample, preventing the reuse of element disposition in the “puzzles” across different months. Therefore, we have a base image (puzzle) for each month and crop type, with a size dependent on the number and area of crops not covered by clouds. These base image sizes range between 256, 384, 512, 768, 1024, 1536, and 2048 pixels per side, influenced by weather conditions and crop type each year season.</p> <p>In <em>this</em> dataset, the synthetic texture sizes match the corresponding base image sizes to facilitate debugging the method implementation and enable subsequent comparisons. For crops with a base image size of 1536 pixels or larger, the generated synthetic images have been reduced to half their size to reduce computational costs and RAM requirements, thereby completing the synthesis faster. Consequently, there remains some diversity in file sizes, generally smaller for crop types with less cultivated area.</p> <p>Additionally, to increase the amount of available data, six variants have been synthesized from each base multispectral image. This number can be arbitrarily increased, as initialization with noise (random numbers) ensures the distinction among the generated data.</p> <p>File names are structured as follows:</p> <ul> <li>Prefix "HO" indicating the synthesis method</li> <li>The crop type name: <ul> <li>Barley</li> <li>Wheat</li> <li>OtherGrainLeguminous</li> <li>Peas</li> <li>FallowAndBareSoil</li> <li>Vetch</li> <li>Alfalfa</li> <li>Sunflower</li> <li>Oats</li> </ul> </li> <li>Year/Month/01 (representing the start of the month period)</li> <li>Side length of the multispectral texture in pixels (based on the highest precision instrument of Sentinel-2: 10m x 10m)</li> <li>Number of the synthesis variant</li> </ul> <p>The generation parameters for all images include:</p> <ul> <li>Normalized and weighted bands (VH band influence increased by a factor of 3 compared to others)</li> <li>4 levels of depth in the Steerable pyramid</li> <li>6 orientations in the Steerable pyramid</li> <li>14 joint statistics of the wavelet coefficients corresponding to basis functions at adjacent spatial locations, orientations, and scales. This parameter is crucial for capturing local dependencies between wavelet coefficients, essential for the visual perception of texture.</li> <li>30 iterations</li> </ul> <p>A significant effort has been made to stabilize the algorithm, and to eliminate artifacts in the generated textures, resulting in much more robust outcomes. However, in rare cases, the initial white noise distribution can be statistically unfavorable, leading to instabilities. Files have been left as generated, without correcting these effects, to make them visible despite their low frequency. Specifically, among the 657 generated multispectral textures, this phenomenon has occurred prominently in only two and is relatively noticeable in another two, leaving the rest free of this effect (affecting less than 1% of the syntheses).</p> <p>Thus, the following files can be considered partially failed syntheses:</p> <ul> <li>HO_Alfalfa_20180801_768_1.nc</li> <li>HO_FallowAndBareSoil_20180101_768_3.nc</li> <li>HO_OtherGrainLeguminous_20171201_256_4.nc</li> <li>HO_Vetch_20180301_384_3.nc</li> </ul> <p>Files are encoded in the standardized net4CDF format [<a href="https://unidata.github.io/netcdf4-python/">link</a>], each containing a single xarray with metadata corresponding to a 3D array with the synthesized texture of the indicated crop type and satellite passes for the regions of Castilla y León and Catalonia for the corresponding monthly period.</p> <p>The most important data structure is the 3D array, where the first two dimensions correspond to the pixel extent indicated in the file name as square textures ('x' and 'y' labels in the xarray). The third dimension denotes the spectral band of the satellite, ordered by constellation and pixel size:</p> <ul> <li>'B02' 10m (Sentinel-2)</li> <li>'B03' 10m (Sentinel-2)</li> <li>'B04' 10m (Sentinel-2)</li> <li>'B08' 10m (Sentinel-2)</li> <li>'B05' originally 20m, resampled to 10m (Sentinel-2)</li> <li>'B06' originally 20m, resampled to 10m (Sentinel-2)</li> <li>'B07' originally 20m, resampled to 10m (Sentinel-2)</li> <li>'B11' originally 20m, resampled to 10m (Sentinel-2)</li> <li>'B12' originally 20m, resampled to 10m (Sentinel-2)</li> <li>'B8A' originally 20m, resampled to 10m (Sentinel-2)</li> <li>'VH' also resampled to 10m (Sentinel-1)</li> </ul> <p>The original dynamic range is preserved in all bands, and they have been synthesized together using our multispectral algorithm variant. The new band combination may result in slightly unusual values in vegetation indices since restrictions were not considered in their transformed space, but in the latent space of the decorrelated Steerable pyramid.</p> <p>Additionally, the following metadata are stored as xarray attributes:</p> <ul> <li>"long_name": corresponding to the crop type name</li> <li>"date": the period of the original data used as the base image for synthesis</li> <li>"dataset": denotes the combination of the initial Castilla y León dataset and the extended 6 Tiles from Catalonia</li> <li>"synthetic_method": corresponds to the high-order constrained method</li> <li>"max_visible_value": a reference value to maintain the same dynamic range when comparing with base images, avoiding distortions in color space and contrast</li> </ul> <p>A total of:</p> <p><strong> 9</strong> types of crops x <strong>12</strong> months x <strong>6</strong> variants = <strong>648</strong> synthetized multispectral textures</p> <p>occupying <strong>34.5</strong>GB, have been organized and uploaded into 9 ZIP files (one per crop type) on the Zenodo website for distribution under Creative Commons Attribution 4.0 International license.</p> <p>The SD4EO Project is funded by the ESA’s FutureEO programme under contract no. 4000142334/23/I-DT and supervised by ESA Φ-lab.</p> <p> </p>
DGU-AI-LAB/Korean-Tourist-Spot-Dataset: Korean Tourist Spot Dataset
<p>The KTS dataset has four modals (image, text, hashtag, likes) and consists of 10 classes related to Korean tourist spots. All data were extracted from Instagram and preprocessed.</p>
AI validated plant observations from social media: Flickr images from central London 2011-2019
<p>This dataset is the result of using an AI image classifer to classify images of plants on social media. We believe this is the first AI validated dataset of biological records taken from social media. This represents the dawn of AI naturalists whose domain of exploration is not the outdoor world but the digital realm. These AI naturalists will trawl streams of data from all over the globe, identifying genuine images of species, and in so doing create valuable data sets that will further our understanding of the distribution of wildlife on our planet. </p> <p>This dataset contains 31,973 classifications of images taken in central London between May 2011 and September 2019 retrieved using the search term 'flower' on Flickr.com. Some images have very low classification confidence (7910 below 0.1), while others have very high confidence (3185 over 0.9). As expected given the spatial extent of the dataset many of the observations are of planted species in gardens and parks.</p> <p>August_et_al_2019.csv provides the data while metadata.txt contains a description of the data and its generation.</p> <p>An interactive visualisation of this data can be viewed at <a href="https://tomaugust.shinyapps.io/ai_flickr_data/">https://tomaugust.shinyapps.io/ai_flickr_data/</a></p>
Collaborazionisti o resistenti? L'accademia ai tempi della valutazione della ricerca (e della scienza aperta)
<p>Questo intervento tenta di rispondere alla domanda: ma una valutazione massiva della ricerca, come quella sviluppata in Italia con la VQR o nel Regno Unito con il RAE/REF, serve davvero? Nella prima parte discuto cinque argomenti usati per giustificare esercizi massivi di valutazione ex post della ricerca. Dopo aver mostrato che questi cinque argomenti non sono robusti, tento di spiegare a che serve realmente la valutazione. E suggerisco di abbandonare la retorica dell’eccellenza a favore di quella della solidità della ricerca.</p>
Regional Disparities in AI Awareness
<p>The dataset consists of variables linked to the benefits, risks, and impacts of AI and regional development in a variety of areas. Sources of the data: Office for National Statistics (ONS), UK.</p>
Replication Package for the Paper Titled "Emerging Results in Using Explainable AI to Improve Software Vulnerability Prediction"
<p>This is a replication package for the paper titled "Emerging Results in Using Explainable AI to Improve Software Vulnerability Prediction".</p>
First discovery and confirmation of PN candidates found from AI and deep learning techniques applied to VPHAS+ survey data
<div> <div> <div> <div> <div> <div> <p>Appendices: VPHAS+ detected PNG images labelled by red boxes, SHS Hα-Rband quotient images, VPHAS+ Hα-Rband quotient images, SAAO spectra with spectral lines labelled.</p> </div> </div> </div> <p>Table: Parameters of the observed PN candidates and any associated nebulosity or outflows.</p> <p>Reduced spectra.</p> </div> </div> </div>
Brick Kiln Dataset for Pakistan's IGP Region Using AI
<p>This dataset represents the first geospatial mapping of brick kiln sites in the IGP region of Pakistan, providing an invaluable resource for understanding the spatial distribution of these sites. Each data point captures a brick kiln's precise location, including coordinates, state, and other important information, standardized in Coordinate Reference System (CRS) EPSG:4326 (WGS 84). This dataset, to the best of our knowledge, is the first of its kind to consolidate and geolocate brick kiln operations across this region, where air pollution impacts from kiln emissions are a significant environmental and public health concern.</p> <p>In addition to the primary geolocation data, the dataset also includes an initial, secondary estimation of emissions (PM10, PM2.5, NOx, and SOx) from these sites. This supplementary information supports preliminary risk assessments, emphasizing proximity-based exposure for populations and sensitive areas (e.g., schools, hospitals) within a 1 km radius of each kiln site. </p> <p>The dataset is made available in multiple formats to facilitate wide usage across spatial analysis platforms:</p> <ul> <li><strong>Geojson</strong></li> <li><strong>Shapefile (SHP)</strong></li> <li><strong>Comma-Separated Values (CSV)</strong></li> </ul> <p><strong>Metadata:</strong></p> <ul> <li>Geographic Coverage: IGP - Pakistan</li> <li>CRS: EPSG:4326 (WGS 84)</li> <li>File Formats: GeoJSON, Shapefile (SHP), CSV<br>- Columns: <p>1. id: Unique identifier for each brick kiln site.</p> <p>2. lat: Latitude coordinate of the brick kiln location, in decimal degrees (CRS: EPSG:4326).</p> <p>3. lon: Longitude coordinate of the brick kiln location, in decimal degrees (CRS: EPSG:4326).</p> <p>4. type: Type or classification of the brick kiln (Either FCBK or ZigZag)</p> <p>5. state: Administrative state or region where the brick kiln is located.</p> <p>6. schools1km: Number of schools located within a 1 km radius of the brick kiln, indicating nearby educational facilities potentially exposed to emissions.</p> <p>7. hosp1km: Number of hospitals within a 1 km radius of the brick kiln, indicating health facilities potentially affected by air pollution exposure.</p> <p>8. pop1km: Estimated population within a 1 km radius of the brick kiln, indicating the number of individuals potentially at risk from emissions.</p> <p>9. avg_bricks: Average number of kilns in operation per day.</p> <p>10. dailyprod(kg): Estimated daily production weight of bricks in kilograms. The production is calculated based on the brick weight multiplied by a factor of 3.</p> <p>11. pm2.5d(kg): Daily emissions estimate of particulate matter (PM2.5) in kilograms, representing fine particles with a diameter of 2.5 micrometers or smaller, which are a health hazard.</p> <p>12. pm10d(kg): Daily emissions estimate of particulate matter (PM10) in kilograms, representing inhalable particles with a diameter of 10 micrometers or smaller.</p> <p>13. noxd(kg): Daily emissions estimate of nitrogen oxides (NOx) in kilograms, a pollutant contributing to respiratory issues and environmental harm.</p> <p>14. soxd(kg): Daily emissions estimate of sulfur oxides (SOx) in kilograms, a pollutant associated with respiratory problems and acid rain.</p> <p>15. seasonprod(kg): Seasonal production estimate in kilograms, representing total brick production over a defined seasonal period (Outside of Monsoon and Smog Season). </p> <p>16. pm2.5s(kg): Seasonal emissions estimate of particulate matter (PM2.5) in kilograms over the specified seasonal period.</p> <p>17. pm10s(kg): Seasonal emissions estimate of particulate matter (PM10) in kilograms over the specified seasonal period.</p> <p>18. noxs(kg): Seasonal emissions estimate of nitrogen oxides (NOx) in kilograms over the specified seasonal period.</p> <p>19. soxs(kg): Seasonal emissions estimate of sulfur oxides (SOx) in kilograms over the specified seasonal period.</p> <p> </p> </li> </ul> <p><strong>Funding Sources</strong>: This research and data collection were funded by Amazon Web Services (AWS) and Smith School of Enterprise and The Environment (University of Oxford). </p>
Environmental and AIS data collected during the EUMarineRobots Trans-National Access activities experiments using the NATO STO-CMRE Littoral Ocean Observatory Network testbed
<p>Environmental and AIS data collected during the H2020 project EUMarineRobots Trans-National Access activities experiments using the NATO STO-CMRE Littoral Ocean Observatory Network (LOON) testbed. Environmental data consists of temperature measured across the water column; sound velocity measured close to the surface and close to the sea bottom; meteorological data at the surface (i.e., pressure, temperature, wind speed and direction, humidity and rain). The environmental dataset is complemented with Automatic Identification System (AIS) data for the ships transiting close to the LOON area (Gulf of La Spezia, Italy)</p> <p>Temperature measured across the water column in the LOON area (Gulf of La Spezia, Italy). The dataset includes measurements for:<br> i) Nov 12, 19-20, 23-24 - 2020<br> ii) Dec 1-4, 14-20 - 2020<br> iii) Jan 12-13, 15, 18-24, 27-28 - 2021</p> <p><br> Meteorological data at the surface (i.e., pressure, temperature, wind speed and direction, humidity and rain) in the LOON area (Gulf of La Spezia, Italy). The dataset includes measurements for:<br> i) Nov 12, 19-20, 23-24 - 2020<br> ii) Dec 1-4, 14-20 - 2020<br> iii) Jan 12-13, 15, 18-24, 27-28 - 2021</p> <p><br> Sound velocity measured close to the surface (SVP1) and close to the sea bottom (SVP2) in the LOON area (Gulf of La Spezia, Italy). The dataset includes measurements for:<br> i) Nov 12, 19-20, 23-24 - 2020<br> ii) Dec 1-4, 14-20 - 2020<br> iii) Jan 12-13, 15, 18-24, 27-28 - 2021</p> <p>SVP2 data missing for Dec 14-20 (2020) and Jan 24, 27-28 (2021).</p> <p>Automatic Identification System (AIS) data for the ships transiting close to the LOON area (Gulf of La Spezia, Italy). The dataset includes AIS data for:<br> i) Nov 12, 19-20, 23-24 - 2020<br> ii) Dec 1-4, 14-20 - 2020<br> iii) Jan 12-13, 15, 18-24, 27-28 - 2021<br> </p> <p>For reference, see: "Environmental data collected on the CMRE LOON tested during the EUMR project: dataset description", Petroccia, Roberto; Zappa, Giovanni; Cimino, Giampaolo; Grati, Alberto; Alves, João. CMRE-DA-2021-001. July 2021, available at https://www.cmre.nato.int/research/publications/latest-techreports/1638-cmre-da-2021-001</p>
AI-related patents (WIPO, category G06N) and market capitalisation by companies registering at least 2 new ones in 2019, sorted into four global regions (China, USA, EEA, rest of the world)
<p>NOTE: for some reason the pptx and previews keep getting munged on this supposedly permanent arxiv, but the data is still there, unchanged, and you can see how the pptx should look in either the jpg, or the the article.</p> <p>Datasets and presentations concerning the strength of the EU and "the rest of the world" relative to China and the USA, for the purpose of illustrating and counteracting / better informing narratives concerning a "new AI cold war". The materials authored by us may be freely used under the terms of the MIT License, which appears in its entirety in both the dataset and the presentation. The other materials are only curated by us, taken from Twitter as examples of misinformation pertaining to this concern.</p> <p>As of 28 June, this work now also appears in a formal publication: Joanna J. Bryson, Helena Malikova; Is There an AI Cold War?. <em><em>Global Perspectives</em></em> 2021; 2 (1): 24803. doi: <a href="https://doi.org/10.1525/gp.2021.24803">https://doi.org/10.1525/gp.2021.24803</a></p> <p>Authors: The original analysis was conducted primarily by Malikova in collaboration with Bryson. An associated publication is anticipated where Bryson is the lead author.</p> <p>Contributors: independently followed Malikova's procedures to check her work. Inconsistencies were triple checked and resolved.</p>
Fighting COVID-19 with computational tools: an AI guided review of 17,000 studies - The CSCoV database.
<p>CSCoV (Computational Studies about COVID-19) is a dataset containing COVID-19 related studies extracted from PubMed, bioRxiv, medRxiv, and arXiv, together with article and author related metrics obtained from Semantic Scholar (plus page views from bioRxiv and medRxiv). Using machine learning, the articles are categorized in six topics (Pharmacology, Genomics, Epidemiology, Healthcare, Clinical Medicine, Clinical Imaging) and prioritized. The database is periodically updated.</p> <ul> <li>Publication: TBA</li> <li>Files included in this release: <ul> <li>cscov_09_2021.png: dataset statistics for the current CSCoV release.</li> <li>cscov_09_2021.tsv: CSCoV database.</li> <li>schema.json: metadata.</li> <li>cscov_09_2021.tar.gz: Doc2Vec and DeepWalk features used for the DL model</li> </ul> </li> <li> <p>Source code: <a href="https://github.com/SFB-KAUST/covid-review">https://github.com/SFB-KAUST/covid-review</a></p> </li> </ul>
GOLDEN-AI functionalities
<p>GOLDEN-AI functionalities.</p>
Demonstration of the GOLDEN Artificial Intelligence (AI) GUI - Artificial Intelligence Platform for mine site monitoring (Open Pit Extraction, Valea Sesei and Roșia Poieni (Romania)).
<p>Demonstration of the GOLDEN Artificial Intelligence (AI) GUI - Artificial Intelligence Platform for mine site monitoring in the Open Pit Extraction (mine located at Valea Sesei and Roșia Poieni (Romania)) (3D view mode).</p> <p>Accessing the GOLDENAI GUI, please refer to the following link (<strong>login required</strong>): <a href="https://next-gui.goldenai.opt-net.eu/ ">https://next-gui.goldenai.opt-net.eu/ </a></p>
Demonstration of the GOLDEN Artificial Intelligence (AI) GUI - Artificial Intelligence Platform for mine site monitoring (Underground Extraction, Pyhäsalmi (Finland)).
<p>Demonstration of the GOLDEN Artificial Intelligence (AI) GUI - Artificial Intelligence Platform for mine site monitoring in the Underground Extraction (mine located at Pyhäsalmi (Finland)) (2D view mode).</p> <p>Accessing the GOLDENAI GUI, please refer to the following link (<strong>login required</strong>): <a href="https://next-gui.goldenai.opt-net.eu/ ">https://next-gui.goldenai.opt-net.eu/ </a></p>
AIS nmea sample dataset
<p>AIS nmea sample dataset used for testing AIS decoding library. This dataset was originally published in <a href="https://github.com/aduvenhage/ais-decoder">https://github.com/aduvenhage/ais-decoder</a> as a sample test dataset. The original dataset can be found at <a href="https://github.com/aduvenhage/ais-decoder/blob/master/data/nmea-sample.txt">https://github.com/aduvenhage/ais-decoder/blob/master/data/nmea-sample.txt</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.