Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,025
datasets available to search
ShareScore release 0.9.0
Dataset results
2,025 results for “AIS”
Exploring the Emotional Impact of AI-Driven Simulations in University Education: A Gender Perspective
<p><span>This research evaluates the impact of AI-driven simulations using ChatGPT on university students' emotions, integrating educational psychology and gender studies frameworks. The study investigates how emotions such as interest, irritability, enthusiasm, discomfort, and disgust are affected by AI-enhanced active learning environments. Additionally, it explores gender differences in emotional responses to AI applications in education. Various statistical tests compare emotion scores, including gender comparisons and pre- and post-intervention analyses. The findings aim to enhance understanding of the effectiveness of AI in simulations and its influence on students' emotional experiences, thereby contributing to the design of more effective AI-based educational systems</span></p>
AI classification of normal and malignant cells on the basis of their viscoelastic properties
Open the record for dataset details and reuse information.
SD4EO: AI-based synthetic solar panel dataset on urban areas
<p>This dataset has been created as part of the deliverables for ESA’s <a href="https://eo4society.esa.int/projects/sd4eo/">SD4EO project</a>. It consists of aerial images of urban areas that mimic certain regions inside Poitiers, Bordeaux and Toulouse.<br>The images have been synthetized with a generative diffusion model, conditioned with schematic maps, and later augmented to include solar panels in the most sunlighted portions of the building roofs.</p> <p>Each entry has 4 types of files:</p> <ul> <li>A PNG image containing the binary mask with pixel-by-pixel segmentation of areas where solar panels have been installed.</li> <li>A PNG image which shows a variant of the synthetic image, each featuring differently placed panels with white support structures.</li> <li>A TXT file containing the locations of axis-aligned bounding boxes in YOLOv8 format (also compatible with YOLOv5). These are included only if panels have been added to the image, with one file per panel variant, to facilitate YOLO training without needing to modify its Dataset class for file reading.</li> <li>An NPZ file storing the coordinates of the bounding boxes aligned with the solar panels (not with the axes), intended for use in more advanced object detection models like YOLOv8.1 or Mask R-CNN.</li> </ul> <p>The dataset includes 46,872 synthetic images, which have been further enhanced by adding solar panels in strategic locations. The complete dataset with metadata has been subdivided into 18 ZIP files (each containing a different subfolder), totalling 6.5GB. This division was made to avoid overloading the storage system. By distributing the 156,127 files across multiple folders, we prevent potential issues on users' computers related to exceeding the maximum number of inodes in the file system and/or the operating system.</p> <p>The SD4EO Project is funded by the ESA’s FutureEO programme under contract no. 4000142334/23/I-DT and supervised by ESA Φ-lab.</p>
Final dataset of Trajectory Synopses over AIS kinematic messages in Brest area (ver. 0.8)
<p>Data provided by NARI (Institut de Recherche de l'�cole Navale) contains AIS kinematic messages from vessels sailing in the Atlantic Ocean around the port of Brest, Brittany, France and span a period from 1 October 2015 to 31 March 2016.</p> <p>Raw AIS messages - file: nari_dynamic.csv, 19,035,630 records.</p> <p>After deduplication of original AIS messages, this dataset yielded 18,495,677 point locations (kinematic AIS messages only), which was used as input for creating trajectory synopses.</p> <p>Attribute "MMSI" in the original data is used in processing as the identifier of each vessel (e.g., "244670495").</p>
RDF triples of raw and trajectory synopses over AIS kinematic messages in Brest
<p>The following list of data sets is derived from the final data set of trajectory synopses available at https://zenodo.org/record/2563256 . We have converted the original data into a) ESRI shapefiles and b) using RDF-Gen (https://zenodo.org/record/2556747) into RDF triples w.r.t. the datAcron ontology.</p>
Link Discovery on AIS and Contextual data sets
<p>This data set contains the links discovered between AIS synopses provided in https://zenodo.org/record/2576152 and contextual data sets available at https://zenodo.org/record/2576584 . Specifically, the detected relations and the contextual data sets used are:</p> <p>a) C1 ports of Brittany, World Port Index and SeaDataNet fishing ports for proximity relation "nearto", stored in AIS_nearto_ports.ttl.7z</p> <p>b) C4 fishing areas (European Commission) for "within" relation stored in AIS_within_fishingAreas.ttl.7z</p> <p>c) C5 fishing interdiction for "within" relation stored in AIS_within_FishingConstraints.ttl.7z</p> <p>d) C5 Natura2000 for "within" relation stored in AIS_within_natura2000.ttl.7z</p> <p>Please notice that the attached data sets contain only the detected links. Description of the resources is provided in the corresponding TTL files.</p>
Data for publication 'Recreational vessels without Automatic Identification System (AIS) dominate anthropogenic noise contributions to a shallow water soundscape' (Scientific Reports 2019)
<p>Data on vessel tracks and underwater noise levels presented in the publication Hermannsen, L., Mikkelsen, L., Tougaard, J., Beedholm, K., Johnson, M. and P. T. Madsen, "Recreational vessels without Automatic Identification System (AIS) dominate anthropogenic noise contributions to a shallow water soundscape", Scientific Reports 9:15477 (<a href="https://doi.org/10.1038/s41598-019-51222-9">https://doi.org/10.1038/s41598-019-51222-9</a>).</p>
Database of retrosynthesis, SA-score and Ei values for AI-Design of ET MALDI matrices
Open the record for dataset details and reuse information.
The Rising Influence of AI in Higher Education: Trends and Insights from a Bibliometric Analysis
<p>The purpose of this research was to examine the evolution, scope, and orientation of the scientific production on artificial intelligence applications in university students. The methodology, with a non-experimental design and qualitative approach, involved a search in Scopus, identifying 643 documents between 1975-2024, analyzed through VOSviewer and Bibliometrix. The results show an emerging field, but with rapid growth (4.59% per year), with notoriety of Kong, Abdulrahman and Chai. Research is predominantly in computer science (61%), social sciences (33%) and engineering (23%) from China, USA, Spain and Taiwan. Current applications focus on the use of AI in education, machine learning, support for academic decisions and student mental health. However, it is necessary to expand the approach towards ethical and regulatory aspects and the evaluation of multifaceted effects on different student profiles. In conclusion, although production is growing rapidly, more comprehensive perspectives are required to responsibly enhance the impact of these technologies on the university educational experience.</p>
Suplementary Dataset VirDetect-AI
<p>This is a Suplementary data of tool VirDetect-AI </p> <p>This repository contains a Deep Learning model for identifying partial virus protein sequences in metagenomic data. </p>
A topical review on AI-interlinked biodomain sensors for multi-purpose applications
<p>This repository consists the following four csv files:</p> <p><strong>acronyms.csv</strong>: The file consists of the acronyms for the terms that are associated with AI methods/AI-based technologies and are used within the article(Thapa et al., Measurement, 227 (2024) 114123) .</p> <p><strong>Fig2data.csv</strong>: The concepts and ideas that have been utilized to construct Fig. 2 are tabulated in this csv file.</p> <p><strong>Fig8data.csv</strong>: The information flow roadmap in an electronic tongue technology is presented in the csv file. These concepts and ideas have been utilized to draw Fig. 8.</p> <p> <br><strong>AI_combined_with_biodomain_sensors.csv</strong>: This file consists of the major words outlining the concept to formulate the graphical abstract.</p> <p> </p>
Data Science Tasks used in "AI Support for Data Scientists: An Empirical Study on Workflow and Alternative Code Recommendation"
<p>This entry contains the supplementary files for a scientific article. </p> <p>The dataset contains the necessary files for the two data science tasks used in the experiment study from scientific article "AI Support for Data Scientists: An Empirical Study on Workflow and Alternative Code Recommendation"</p> <p> </p>
OA-Concepts and Wikipedia-Links for "The different AI of Science and Wikipedia"
<p>The files contain the data for the VosViewer analyses in "The Different Artificial Intelligences of Science and Wikipedia" (Korte et al. 2024).</p> <p>OpenAlex:</p> <p>As described in the paper, all works from OpenAlex from 2001 and 2022 with the concept "Artificial Intelligence" and a concept score > 0.3 were downloaded. In August 23 using the OpenAlex API. The files contain for each work all concepts with a concept score > 0.3 in one line each separeted by dots. This allows co-occurrence analyses of concepts in VosViewer. For the analyses in the paper whitespaces and "(" in the concepts were removed.</p> <p>First line for 2001: <br>Random forest. Mathematics. AdaBoost. Statistics. Tree (set theory). Generalization. Support vector machine. Generalization error. Measure (data warehouse). Artificial intelligence. Pattern recognition (psychology). </p> <p>Wikipedia:</p> <p>As described in the paper, all Wikipedia pages of the category "Artificial Intelligence" and of all direct sibling categories were downloaded. In August 23 with the Wikipedia Periodic Revisions tool of El Baff and Hecking (https://github.com/DLR-SC/wikipedia-periodic-revisions): . The files contain for each page all hyperlinks to other Wikipedia pages for the dates of 12.31.2005 and 12.31.2021 in one line each separeted by dots. This allows co-occurrence analyses of links in Wikipedia.</p> <p>First line of 2006: <br>vehicleregistrationplate. corporation. googlesearch. googleplatform. googol. google(disambiguation). menlopark,california. erice.schmidt. sergeybrin. lawrencee.page. georgereyes. internet. [...]</p>
The Impact of Generative AI on Student Learning Outcomes: A Statistical Analytical Approach - Dataset
Open the record for dataset details and reuse information.
Trust in AI: Perspectives of C-Level Executives in Brazilian Organizations
<p>Context: With the advancement of Artificial Intelligence (AI) and its increasing integration into business processes, the trust of top management in Brazilian companies has become a crucial issue. Business leaders must be aware of the challenges and opportunities associated with adopting AI in their operations. A lack of understanding and knowledge about the capabilities and limitations of AI can lead to hesitations and concerns from top management regarding its use. Goal: This work aims to identify the main challenges preventing C-level executives from fully trusting AI and its applications within their organizations in the Brazilian context. Additionally, a reference guide is proposed to help top management better understand how AI can be effectively and ethically integrated into their business strategies. Method: We conducted a survey with business leaders from various sectors to understand their perceptions of trust in AI and their concerns regarding its implementation. Results: The results revealed that the main obstacles faced by top management in Brazilian companies were the lack of understanding about AI's capabilities and its ethical implications. Therefore, it is imperative for business leaders to invest in education and awareness about AI, seeking to understand its benefits and challenges. Only then will they be able to make informed decisions and fully trust AI solutions to drive innovation and sustainable growth in their organizations and the improvement of organizational processes.</p>
4th International Workshop on Camera Traps, AI, and Ecology - Photos
Open the record for dataset details and reuse information.
EMS3D-KITTI-Synthetic: A Synthetic 3D Dataset in KITTI Format with Balanced EMS Vehicle Distribution for Autonomous Driving AI Model Training
<p>A 3D synthetic dataset in KITTI format, focused on emergency vehicles such as ambulances and police cars. The dataset was generated across 8 towns within the CARLA simulator and converted into the KITTI format, ensuring compatibility for direct use in AI model training for autonomous driving applications.</p>
DESI 2023 details for AI indicator
<p>The following data has been selected from EUROSTAT isoc_eb_ai database:</p> <ul> <li>Enterprises don't use any AI system (of E_CHTB, E_BDAML, E_BDANL, E_RBTS)</li> <li>Enterprises use one AI system (of E_CHTB, E_BDAML, E_BDANL, E_RBTS)</li> <li>Enterprises use two AI systems (of E_CHTB, E_BDAML, E_BDANL, E_RBTS)</li> <li>Enterprises use three AI systems (of E_CHTB, E_BDAML, E_BDANL, E_RBTS)</li> <li>Enterprises use four AI systems (of E_CHTB, E_BDAML, E_BDANL, E_RBTS)</li> <li>Enterprises use AI technologies performing analysis of written language (text mining)</li> <li>Enterprises use AI technologies converting spoken language into machine-readable format (speech recognition)</li> <li>Enterprises use AI technologies generating written or spoken language (natural language generation)</li> <li>Enterprises use AI technologies identifying objects or persons based on images (image recognition, image processing)</li> <li>Enterprises use machine learning (e.g. deep learning) for data analysis</li> <li>Enterprises use AI technologies automating different workflows or assisting in decision making (AI based software robotic process automation)</li> <li>Enterprises use AI technologies enabling physical movement of machines via autonomous decisions based on observation of surroundings (autonomous robots, self-driving vehicles, autonomous drones)</li> <li>Enterprises use at least one of the AI technologies: AI_TTM, AI_TSR, AI_TNLG, AI_TIR, AI_TML, AI_TPA, AI_TAR</li> <li>Enterprises don't use any of the AI technologies: AI_TTM, AI_TSR, AI_TNLG, AI_TIR, AI_TML, AI_TPA, AI_TAR</li> <li>Enterprises use at least two of the AI technologies: AI_TTM, AI_TSR, AI_TNLG, AI_TIR, AI_TML, AI_TPA, AI_TAR</li> <li>Enterprises use at least three of the AI technologies: AI_TTM, AI_TSR, AI_TNLG, AI_TIR, AI_TML, AI_TPA, AI_TAR</li> <li>Enterprises use AI technologies for marketing or sales</li> <li>Enterprises use AI technologies for production processes</li> <li>Enterprises use AI technologies for organisation of business administration processes</li> <li>Enterprises use AI technologies for management of enterprises</li> <li>Enterprises use AI technologies for logistics</li> <li>Enterprises use AI technologies for ICT security</li> <li>Enterprises use AI technologies for human resources management or recruiting</li> <li>Enterprises use AI technologies for at least one of the purposes: AI_PMS, AI_PPP, AI_PBA, AI_PME, AI_PLOG, AI_PITS, AI_PHR</li> <li>Enterprises use AI technologies for at least two of the purposes: AI_PMS, AI_PPP, AI_PBA, AI_PME, AI_PLOG, AI_PITS, AI_PHR</li> <li>Enterprises use AI technologies for at least three of the purposes: AI_PMS, AI_PPP, AI_PBA, AI_PME, AI_PLOG, AI_PITS, AI_PHR</li> </ul>
AI-based real-time animal management system for flock monitoring in the husbandry of fattening turkeys
<p>The dataset was collected in a commercial turkey barn using a self-developed AI-based real-time animal management system to assess the behavior of turkeys. The data covers turkeys from the 6th to the 18th week of life.</p>
Geoparsing with Large Language Models: Leveraging the linguistic capabilities of generative AI to improve geographic information extraction
<h2>Geoparsing with Large Language Models</h2> <p>The .zip file included in this repository contains all the code and data required to reproduce the results from our paper. Note, however, that in order to run the OpenAI models, users will required an OpenAI API key and sufficient API credits.</p> <div> <h3>Data</h3> <p>The data used for the paper are in the <code>datasetst</code> and <code>results</code> folders.</p> <ul> <li> <p>**Datasets: **This contains the XML files (LGL and Geovirus) and Json files (News2024) used to benchmark the models. It also contains all the data used to fine-tune the gpt-3.5 model, the prompt templates sent to the LLMs, and other data used for mapping and data creation.</p> </li> <li> <p>**Results: **This contains the results for the models on the three datastes. The folder is separated by dataset, with a single <code>.csv</code> file giving the results for each model on each dataset separately. The <code>.csv</code> file is structured so that each row contains either a predicted toponym and an associated true toponym (along with assigned spatial coordinates), if the model correctly identified a toponym; otherwise the true toponym columns are empty for false positives and the predicted columns are empty for false negatives.</p> </li> </ul> <h3>Code</h3> <p>The code is split into two seperate folders <code>gpt_geoparser</code> and <code>notebooks</code>.</p> <ul> <li>**GPT_Geoparser: **this contains the classes and methods used process the XML and JSON articles (<code>data.py</code>), interact with the Nominatim API for geocoding (<code>gazetteer.py</code>), interact with the OpenAI API (<code>gpt_handler.py</code>), process the outputs from the GPT models (<code>geoparser.py</code>) and analyse the results (<code>analysis.py</code>).</li> <li><strong>Notebooks</strong>: This series of notebooks can be used to reproduce the results given in the paper. The file names a reasonably descriptive of what they do within the context of the paper.</li> </ul> <h3>Code/software</h3> <h3>Requirements</h3> <ul> <li>Numpy</li> <li>Pandas</li> <li>Geopy</li> <li>Scitkit-learn</li> <li>lxml</li> <li>openai</li> <li>matplotlib</li> <li>Contextily</li> <li>Shapely</li> <li>Geopandas</li> <li>tqdm</li> <li>huggingface_hub</li> <li>Gnews</li> </ul> <h3>Access information</h3> <p>Other publicly accessible locations of the data:</p> <ul> <li>The LGL and GeoVirus datasets can also be obtained <a href="https://github.com/milangritta/Pragmatic-Guide-to-Geoparsing-Evaluation" target="_blank" rel="noopener">here<span> (opens in new window)</span></a>.</li> </ul> <h3>Abstract</h3> <div> <p>Geoparsing- the process of associating textual data with geographic locations - is a key challenge in natural language processing. The often ambiguous and complex nature of geospatial language make geoparsing a difficult task, requiring sophisticated language modelling techniques. Recent developments in Large Language Models (LLMs) have demonstrated their impressive capability in natural language modelling, suggesting suitability to a wide range of complex linguistic tasks. In this paper, we evaluate the performance of four LLMs - GPT-3.5, GPT-4o, Llama-3.1-8b and Gemma-2-9b - in geographic information extraction by testing them on three geoparsing benchmark datasets: GeoVirus, LGL, and a novel dataset, News2024, composed of geotagged news articles published outside the models' training window. We demonstrate that, through techniques such as fine-tuning and retrieval-augmented generation, LLMs significantly outperform existing geoparsing models. The best performing models achieve a toponym extraction F1 score of 0.985 and toponym resolution accuracy within 161 km of 0.921. Additionally, we show that the spatial information encoded within the embedding space of these models may explain their strong performance in geographic information extraction. Finally, we discuss the spatial biases inherent in the models' predictions and emphasize the need for caution when applying these techniques in certain contexts.</p> </div> <h3>Methods</h3> <div> <p>This contains the data and codes required to reproduce the results from our paper. The LGL and GeoVirus datasets are pre-existing datasets, with references given in the manuscript. The News2024 dataset was constructed specifically for the paper. </p> <p>To construct the News2024 dataset, we first created a list of 50 cities from around the world which have population greater than 1000000. We then used the GNews python package <a href="https://pypi.org/project/gnews/" target="_blank" rel="noopener">https://pypi.org/project/gnews/<span> (opens in new window)</span></a> to find a news article for each location, published between 2024-05-01 and 2024-06-30 (inclusive). Of these articles, 47 were found to contain toponyms, with the three rejected articles referring to businesses which share a name with a city, and which did not otherwise mention any place names.</p> <p>We used a semi autonmous approach to geotagging the articles. The articles were first processed using a Distil-BERT model, fine tuned for named entity recognicion. This provided a first estimate of the toponyms within the text. A human reviewer then read the articles, and accepted or rejected the machine tags, and added any tags missing from the machine tagging process. We then used OpenStreetMap to obtain geographic coordinates for the location, and to identify the toponym type (e.g. city, town, village, river etc). We also flagged if the toponym was acting as a geo-political entity, as these were reomved from the analysis process. In total, 534 toponyms were identified in the 47 news articles. </p> </div> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.