Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,025
datasets available to search
ShareScore release 0.9.0
Dataset results
2,025 results for “AIS”
cultural-ai/wordsmatter: Words Matter: a knowledge graph of contentious terms
<p>The choice of words describing cultural heritage can cause debates. It is especially sensitive when artefacts relate to different cultures and peoples who have been historically marginalised. Words chosen by archivists or curators may transmit stereotypes. The cultural heritage community has produced knowledge on potentially stereotyping and offensive terminology in heritage collections. At the same time, their knowledge is difficult to incorporate into existing online collections unless this knowledge is structured and machine-readable.</p> <p>The Words Matter Knowledge Graph represents domain expert knowledge on discussions about contentious terminology in the cultural sector. In the knowledge graph, 75 English and 83 Dutch contentious terms are linked to explanations of their usage and suggested alternatives from domain experts. There are also related matches between contentious terms and sources from external datasets: Wikidata, Princeton WordNet, Open Dutch WordNet, and Getty Art & Architecture Thesaurus.</p> <p>This Zenodo publication includes the CULCO scheme used to model contentious terms in the knowledge graph. The scheme documentation is <a href="https://cultural-ai.github.io/wordsmatter/" target="_blank" rel="noopener">available on a separate page</a>.</p> <p>This knowledge graph is <a href="https://amsterdam.wereldmuseum.nl/en/about-wereldmuseum-amsterdam/research/words-matter-publication" target="_blank" rel="noopener">based</a> on the publication “Words Matter: An Unfinished Guide to Word Choices in the Cultural Sector” by the National Museum of World Cultures (NMVW). </p> <p><a href="https://doi.org/10.1007/978-3-031-33455-9_30" target="_blank" rel="noopener">Read more</a> about this work in the paper "A Knowledge Graph of Contentious Terminology for Inclusive Representation of Cultural Heritage" (2023) by Andrei Nesterov, Laura Hollink, Marieke van Erp & Jacco van Ossenbruggen.</p> <p>In this version:</p> <ul> <li>the CULCO scheme documentation is updated</li> <li>versioning is fixed</li> <li>typos are corrected</li> </ul>
Convergent approaches to AI Explainability for HEP muonic particles pattern recognition Dataset
<p>Dataset associated to the publication "Convergent approaches to AI Explainability for HEP muonic particles pattern recognition", Leandro Maglianella, Lorenzo Nicoletti, Stefano Giagu*, Christian Napoli, and Simone Scardapane, submitted to Computing and Software for Big Science.</p> <p>*corresponding author: stefano.giagu [AT] uniroma1.it</p> <p>Description:</p> <p>provided as a compressed zip file. Contains 7 numpy .npy files:</p> <ul> <li>train_images_with_noise.npy: numpy array containing 850003 "images" of muonic tracks with detector noise (shape (850003, 9, 384)). Each image contains 1 muonic track.</li> <li>train_images_without_noise.npy: numpy array containing 850003 "images" of muonic tracks w/o detector noise (shape (850003, 9, 384)). Each image contains 1 muonic track.</li> <li>train_labels.npy: labels associated to each image (shape (850003, 5)), corresponding to (pT, eta, phi, 0, nhits) of the muonic track, with pT: transverse momentum, eta: pseudo-rapidity, phi: azimuthal angle, and nhits: the number of pixels turned on by the muon</li> <li>test_images_with_noise.npy: same as above for a 94445 images test set</li> <li>test_images_without_noise.npy: same as above for a 94445 images test set</li> <li>test_labels.npy: same as above for a 94445 images test set</li> <li>images_only_noise.npy: numpy array containing 944448 "images" w/o muons, containing detector noise only (shape (944448, 9, 384))</li> </ul>
LungVis1.0: Active learning AI-powered 3D imaging ecosystem for spatial profiling of lung geometry and pulmonary nanoparticle delivery
<p>The imaging dataset was obtained by light sheet fluorescence microscopy on tissue cleared murine lungs. It includes whole lung autofluorence image, particle fluorescence image, and artifical intelligence nnU-Net generated lung airway segments. The dataset provides 78 healthy murine lung strucutre and airway geometry for C57BL/6 mice and offers comprehensive delivery features including qualitative and quantitative analysis on the temporal and spatial inter- and intra-acinar deposition patterns and NP regional dosimetry for four commonly-used routes of pulmonary delivery,namely intranasal liquid aspiration, intratracheal liquid instillation, ventilator-assisted and nose-only aerosol inhalation.</p> <p>Raw LSFM imaging data collection was carried out between 2017-2021, the AI code and generated airway segmention were performed in 2021-2022, the whole datasets were then compiled in 2023. </p> <p>Please ensure to cite our paper for any reuse or reanalysis. Yang, L., Liu, Q., Kumar, P. <em>et al.</em> LungVis 1.0: an automatic AI-powered 3D imaging ecosystem unveils spatial profiling of nanoparticle delivery and acinar migration of lung macrophages. <em>Nat Commun</em> <strong>15</strong>, 10138 (2024). https://doi.org/10.1038/s41467-024-54267-1</p> <p>For any inquiries, please feel free to contact us at lin.yang@helmholtz-munich.de </p>
Image and label patches used to train GLASS-AI
<p>This archive contains the paired image and label patches used to train our machine learning pipeline, Grading of Lung Adenocarcinoma with Simultaneous Segmentation by Artificial Intelligence (GLASS-AI). </p> <p>Image patches were generated from whole slide images of H&E-stained sections using an Aperio ScanScope AT2 Slide Scanner (Leica) at 20x magnification with a 0.5022 microns/pixel resolution. The individual tumors and airways were annotated by an expert human before being divided into 224x224 pixel patches of the H&E image and paired annotation layer.</p> <p>For more details regarding how these data were used to train GLASS-AI, please see our forthcoming manuscript. </p>
AI-derived annotations for the NLST and NSCLC-Radiomics computed tomography imaging collections
<p>Public imaging datasets are critical for the development and evaluation of automated tools in cancer imaging. Unfortunately, many of the available datasets do not provide annotations of tumors or organs-at-risk, crucial for the assessment of these tools. This is due to the fact that annotation of medical images is time consuming and requires domain expertise. It has been demonstrated that artificial intelligence (AI) based annotation tools can achieve acceptable performance and thus can be used to automate the annotation of large datasets. As part of the effort to enrich the public data available within NCI Imaging Data Commons (IDC) (<a href="https://imaging.datacommons.cancer.gov/">https://imaging.datacommons.cancer.gov/</a>) [1], we introduce this dataset that consists of such AI-generated annotations for two publicly available medical imaging collections of Computed Tomography (CT) images of the chest. For detailed information concerning this dataset, please refer to our publication <a href="https://www.nature.com/articles/s41597-023-02864-y">here</a> [2]. </p> <p>We use publicly available pre-trained AI tools to enhance CT lung cancer collections that are unlabeled or partially labeled. The first tool is the nnU-Net deep learning framework [3] for volumetric segmentation of organs, where we use a pretrained model (Task D18 using the SegTHOR dataset) for labeling volumetric regions in the image corresponding to the heart, trachea, aorta and esophagus. These are the major organs-at-risk for radiation therapy for lung cancer. We further enhance these annotations by computing 3D shape radiomics features using the pyradiomics package [4]. The second tool is a pretrained model for per-slice automatic labeling of anatomic landmarks and imaged body part regions in axial CT volumes [5].</p> <p>We focus on enhancing two publicly available collections, the Non-small Cell Lung Cancer Radiomics (NSCLC-Radiomics collection) [6,7], and the National Lung Screening Trial (NLST collection) [8,9]. The CT data for these collections are available both in The Cancer Imaging Archive (TCIA) [10] and in NCI Imaging Data Commons (IDC). Further, the NSLSC-Radiomics collection includes expert-generated manual annotations of several chest organs, allowing us to quantify performance of the AI tools in that subset of data.</p> <p>IDC is relying on the DICOM standard to achieve FAIR [10] sharing of data and interoperability. Generated annotations are saved as DICOM Segmentation objects (volumetric segmentations of regions of interest) created using the <em>dcmqi</em> [12], and DICOM Structured Report (SR) objects (per-slice annotations of the body part imaged, anatomical landmarks and radiomics features) created using <em>dcmqi </em>and <em>highdicom</em> [13]. 3D shape radiomics features and corresponding DICOM SR objects are also provided for the manual segmentations available in the NSCLC-Radiomics collection. </p> <p>The dataset is available in IDC, and is accompanied by our publication <a href="https://www.nature.com/articles/s41597-023-02864-y">here</a> [2]. This pre-print details how the data were generated, and how the resulting DICOM objects can be interpreted and used in tools. Additionally, for further information about how to interact with and explore the dataset, please refer to our <a href="https://github.com/ImagingDataCommons/nnU-Net-BPR-annotations/">repository</a> and accompanying <a href="https://github.com/ImagingDataCommons/nnU-Net-BPR-annotations/blob/main/usage_notebooks/scientific_data_paper_usage_notes.ipynb">Google Colaboratory notebook</a>. </p> <p>The annotations are organized as follows. For NSCLC-Radiomics, three nnU-Net models were evaluated ('2d-tta', '3d_lowres-tta' and '3d_fullres-tta'). Within each folder, the PatientID and the StudyInstanceUID are subdirectories, and within this the DICOM Segmentation object and the DICOM SR for the 3D shape features are stored. A separate directory for the DICOM SR body part regression regions ('sr_regions') and landmarks ('sr_landmarks') are also provided with the same folder structure as above. Lastly, the DICOM SR for the existing manual annotations are provided in the 'sr_gt' directory. For NSCLC-Radiomics, each patient has a single StudyInstanceUID. The DICOM Segmentation and SR objects are named according to the SeriesInstanceUID of the original CT files. </p> <ul> <li>nsclc <ul> <li>2d-tta <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_SEG.dcm</li> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>3d_lowres-tta <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_SEG.dcm</li> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>3d_fullres-tta <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_SEG.dcm</li> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_regions <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_regions_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_landmarks <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_landmarks_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_gt <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> </ul> </li> </ul> <p>For NLST, the '3d_fullres-tta' model was evaluated. The data is organized the same as above, where within each folder the PatientID and the StudyInstanceUID are subdirectories. For the NLST collection, it is possible that some patients have more than one StudyInstanceUID subdirectory. A separate directory for the DICOM SR body par regions ('sr_regions') and landmarks ('sr_landmarks') are also provided. The DICOM Segmentation and SR objects are named according to the SeriesInstanceUID of the original CT files. </p> <ul> <li>nlst <ul> <li>3d_fullres-tta <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_SEG.dcm</li> <li>ReferencedSeriesInstanceUID_features_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_regions <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_regions_SR.dcm</li> </ul> </li> </ul> </li> </ul> </li> <li>sr_landmarks <ul> <li>PatientID <ul> <li>StudyInstanceUID <ul> <li>ReferencedSeriesInstanceUID_landmarks_SR.dcm </li> </ul> </li> </ul> </li> </ul> </li> </ul> </li> </ul> <p>The query used for NSCLC-Radiomics is <a href="https://github.com/ImagingDataCommons/ai_medima_misc/blob/main/common/queries/NSCLC_Radiomics_query.txt">here</a>, and a list of corresponding SeriesInstanceUIDs (along with PatientIDs and StudyInstanceUIDs) is <a href="https://github.com/ImagingDataCommons/ai_medima_misc/blob/main/common/queries/zenodo_nsclc_radiomics_series_analyzed.csv">here</a>. The query used for NLST is <a href="https://github.com/ImagingDataCommons/ai_medima_misc/blob/main/common/queries/NLST_query.txt">here</a>, and a list of corresponding SeriesInstanceUIDs (along with PatientIDs and StudyInstanceUIDs) is <a href="https://github.com/ImagingDataCommons/ai_medima_misc/blob/main/common/queries/zenodo_nlst_series_analyzed.csv">here</a>. The two csv files that describe the series analyzed, <em>nsclc_series_analyzed.csv</em> and <em>nlst_series_analyzed.csv</em>, are also available as uploads to this repository. </p> <p><em>Version updates: </em></p> <p><em>Version 2: For the regions SR and landmarks SR, changed to use a distinct TrackingUniqueIdentifier for each MeasurementGroup. Also instead of using TargetRegion, changed to use FindingSite. Additionally for the landmarks SR, the TopographicalModifier was made a child of FindingSite instead of a sibling.</em></p> <p><em>Version 3: Added the two csv files that describe which series were analyzed </em></p> <p><em>Version 4: Modified the landmarks SR as the TopographicalModifier for the Kidney landmark (bottom) does not describe the landmark correctly. The Kidney landmark is the "first slice where both kidneys can be seen well." Instead, removed the use of the TopographicalModifier for that landmark. For the features SR, modified the units code for the Flatness and Elongation, as we incorrectly used mm units instead of no units. </em></p>
AI and IIoT system for soya beans production
<p>Animated video presenting an AI and IIoT system for soya beans production - process optimisation and equipment predictive maintenance.</p>
AIS data
<p>Terrestrial vessel automatic identification system (AIS) data was collected around Ålesund, Norway in 2020, from multiple receiving stations with unsynchronized clocks. Features are '<em>mmsi</em>', '<em>imo</em>', '<em>length</em>', '<em>latitude</em>', '<em>longitude</em>', '<em>sog</em>', '<em>cog</em>', '<em>true_heading</em>', '<em>datetime UTC</em>', '<em>navigational status</em>', and '<em>message number</em>'. Compact parquet files can be turned into data frames with python's pandas library. Data is irregularly sampled because of the <a href="https://imorules.com/GUID-D7E2DECA-C42B-419D-B613-9D03236FA4F1.html">navigational status</a>. The preprocessing script for training the machine learning models can be found <a href="https://github.com/WenjieDu/TSDB">here</a>. There you will find gathered dozen of trainable models and hundreds of datasets. Visit <a href="https://www.kystverket.no/navigasjonstjenester/ais/tilgang-pa-ais-data/">this</a> website for more information about the data. If you have additional questions, please find our information in the links below:</p> <ul> <li><a href="https://www.ntnu.no/ansatte/luka.grgicevic">Luka Grgičević</a></li> <li><a href="https://www.ntnu.no/ansatte/ottar.osen">Ottar Laurits Osen</a></li> </ul>
Dataset Worldwide Survey on the Impact of AI Chatbots and Large Language Models in Dental Education: Insights from Dental Educators
<p><strong>This dataset contains responses from participants regarding their awareness, knowledge, and perceptions of AI-powered tools in dental education. The data was collected during May-June 2023 to investigate the potential enhancement that AI can bring to dental education. The dataset includes variables related to participants' demographics, experiences, perceptions, and opinions.</strong></p> <p><strong>Details in the published protocol by Uribe, S. E., & Maldupa, I. (2023, June 2). Chatbots In Dental Education - Research Protocol. https://doi.org/10.17605/OSF.IO/3BSG2</strong></p>
Number of works grouped by AI technique employed
<p>Number of works grouped by AI technique employed. Part of the study "What do we mean by GenAI?"</p>
Number of works grouped by objective, domain and AI technique employed
<p>Number of works grouped by objective, domain and AI technique employed. Part of the study "What do we mean by GenAI?"</p>
Relationships among the generated content type, task, AI technique, and application domain in the retrieved works
<p>Relationships among the generated content type, task, AI technique, and application domain in the retrieved works. Part of the study "What do we mean by GenAI?"</p>
Test film for Dating ancient manuscripts using radiocarbon and AI-based writing style analysis
<p>This film is associated with the following article:<br> Title: <strong>Dating ancient manuscripts using radiocarbon and AI-based writing style analysis</strong><br> Authors: Mladen Popović, Maruf A. Dhali, Lambert Schomaker, Johannes van der Plicht, Kaare Lund Rasmussen, Jacopo La Nasa, Ilaria Degano, Maria Perla Colombini, and Eibert Tigchelaar<br> <em>(under review)</em></p> <p> </p> <p>This film is made for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497<br> Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> Videographer: <a href="https://videobrouwers.nl/">https://videobrouwers.nl/</a></p> <p><strong>Copyright (c) </strong> University of Groningen, 2023. All rights reserved.<br> <strong>Disclaimer and copyright notice for this file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the video for research purposes. It is not allowed to distribute this video for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this video.</p> <p><strong>4) </strong>the user should refer to the first public article mentioned above in this video.</p> <p><strong>5) </strong>the recipient should refrain from proliferating the video to third parties external to his/her local research group. Please refer interested researchers to this site to obtain their own copy.</p> <p> </p> <p><strong>Film description:</strong><br> A test was conducted on 6 July 2021. The test consisted of giving unseen <sup>14</sup>C results to the AI experts to see whether Enoch (date prediction model) would give date prediction estimates that match the <sup>14</sup>C results. However, at the start of the test, it was unknown to the AI experts that the samples were chosen because <sup>14</sup>C results were available for them. The <sup>14</sup>C results were taken from the 1990s <sup>14</sup>C dating of the Dead Sea Scrolls [1,2]. The assumption was that the manuscripts chosen were not contaminated with castor oil as these manuscripts were not handled by the original team of editors in the 1950s [3,4,5]. This applies to 1QIsa<sup>a</sup>, 1QpHab, 1QapGen, 1QS, 1QH<sup>a</sup>, 11Q19, Mas1l. Two more manuscripts were added for other reasons. 4Q53 was added because scholars assume that it was written by the same scribe as 1QS. 4Q319 was added because it is actually the same manuscript as 4Q259 [6], which was subjected to <sup>14</sup>C dating by our own project. The test was filmed. The film captures the whole process that was conducted in one go.</p> <p><strong>If you have any questions, please get in touch with us:</strong><br> Mladen Popović <m.popovic(at)rug.nl><br> Maruf A. Dhali <m.a.dhali(at)rug.nl><br> Lambert Schomaker <l.r.b.schomaker(at)rug.nl></p> <p><strong>References:</strong><br> <em>1.Bonani, G., Ivy, S., Wölfli, W., Broshi, M., Carmi, I., & Strugnell, J. (1992). Radiocarbon dating of fourteen Dead Sea scrolls. Radiocarbon, 34(3), 843-849.<br> 2. Jull, A. T., Donahue, D. J., Broshi, M., & Tov, E. (1995). Radiocarbon dating of scrolls and linen fragments from the Judean desert. Radiocarbon, 37(1), 11-19.<br> 3. Doudna, G., Flint, P. W., & VanderKam, J. C. (1998). Dating the Scrolls on the basis of radiocarbon analysis. The Dead Sea scrolls after fifty years. A comprehensive assessment. Volume one, 1, 430-471.<br> 4. Carmi, I. (2002). Are the 14C dates of the Dead Sea Scrolls affected by castor oil contamination? Radiocarbon, 44(1), 213-216.<br> 5. Rasmussen, K. L., van der Plicht, J., Doudna, G., Nielsen, F., Højrup, P., Stenby, E. H., & Pedersen, C. T. (2009). The effects of possible contamination on the radiocarbon dating of the Dead Sea Scrolls II: empirical methods to remove castor oil and suggestions for redating. Radiocarbon, 51(3), 1005-1022.<br> 6. Hempel, C. (2020). The Community Rules from Qumran: A Commentary (Vol. 183). Mohr Siebeck.</em></p>
HawaiiCoast_GT: Curated AIS for Hawaii's coast correlated with ground truth incidents
<p>Because of the high-risk nature of emergencies and illegal activities at sea, it is critical that algorithms designed to detect anomalies from maritime traffic data be robust. However, there exist no publicly available maritime traffic datasets with real-world labelled anomalies. As a result, most anomaly detection algorithms for maritime traffic are validated without ground truth. We introduce the HawaiiCoast_GT dataset, the first ever publicly available automatic identification system dataset with a large corresponding set of true anomalous incidents. This dataset—cleaned and curated from Bureau of Ocean Energy Management (BOEM) and National Oceanic and Atmospheric Administration (NOAA) automatic identification system (AIS) data--covers Hawaii’s coastal waters for four years (2017-2020) and contains 88,749,176 AIS points for a total of 2,622 unique vessels. 208 tracks are labelled corresponding to 154 labelled real-world incidents. The codebase used to curate the original AIS data is being made openly available on GitHub.</p>
AI-rmonizer Weights
<p><strong>Training results. AI-rmonizer unsupervised music composition system.</strong></p> <p>Weights calculated from training with the <a href="https://magenta.tensorflow.org/datasets/maestro#v300">MAESTRO v3.0.0 MIDI dataset</a>.</p> <p>Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck. "Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset." In International Conference on Learning Representations, 2019.</p>
HITS Inc.'s models and data for Dacon challenge, Jump AI 2023
<p>Here deposits model and data files developed during HITS Inc.'s participation in the Dacon challenge, Jump AI 2023:</p> <p><a href="https://dacon.io/competitions/official/236127/overview/description">https://dacon.io/competitions/official/236127/overview/description</a></p> <p>Note that the files here alone are less useful unless appropriate codes are employed.</p> <p><strong>File description:</strong></p> <ul> <li>pred_model_AutoGluon.tar.xz: AutoGluon model parameters for prediction.</li> <li>valid_model_AutoGluon.tar.xz: AutoGluon model parameters for validation.</li> <li>ckpts_original.tar.xz: fine-tuned <a href="https://github.com/yuyangw/MolCLR">MolCLR</a> model parameters.</li> <li>qc_out.tar.xz: molecular electronic structure files (.wfn).</li> <li>sdf_optimized.tar.xz: molecular structure files (.sdf).</li> <li>atomwfn.tar.xz: atomic electronic structure files (.wfn).</li> </ul>
Authorship Identification of SOurce COde 2020 (AI-SOCO)
<p>General authorship identification is essential to the detection of undesirable deception of others' content misuse or exposing the owners of some anonymous hurtful content. This is done by revealing the author of that content. <strong>A</strong>uthorship <strong>I</strong>dentification of <strong>SO</strong>urce <strong>CO</strong>de (AI-SOCO) focuses on uncovering the author who wrote some piece of code. This facilitates solving issues related to cheating in academic, work and open source environments. Also, it can be helpful in detecting the authors of malware softwares over the world.</p> <p>The detection of cheating in academic communities is significant to properly address the contribution of each researcher. Also, in work environments, credit sometimes goes to people that did not deserve it. Such issues of plagiarism could arise in open source projects that are available on public platforms. Similarly, this could be used in public or private online coding contests whether done in coding interviews or in official coding training contests to detect the cheating of applicants or contestants. A system like this could also play a big role in detecting the source of anonymous malicious softwares.</p> <p>The dataset is composed of source codes collected from the open submissions in the <a href="http://www.google.com/url?q=http%3A%2F%2Fcodeforces.com%2F&sa=D&sntz=1&usg=AFQjCNHKGPIjzjl6ujCm0t4EU_waJWvU-Q">Codeforces</a> online judge. Codeforces is an online judge for hosting competitive programming contests such that each contest consists of multiple problems to be solved by the participants. A Codeforces participant can solve a problem by writing a solution for it using any of the available programming languages on the website, and then submitting the solution through the website. The solution's result can be correct (accepted) or incorrect (wrong answer, time limit exceeded, etc.).</p> <p>In our dataset, we selected 1,000 users and collected 100 source codes from each one. So, the total number of source codes is 100,000. All collected source codes are correct, bug-free, compile-ready and written using the C++ programming language using different versions. For each user, all collected source codes are from unique problems.</p> <p>Given the pre-defined set of source codes and their authors, the task is to build a system to determine which one of these authors wrote a given unseen before source code.</p> <p>Dataset website: https://sites.google.com/view/ai-soco-2020.</p>
Ethical Perspectives in AI: A Two-folded Exploratory Study From Literature and Active Development Projects - Supplementary Material
<p>This is the Supplementary Material provided for the work accepted on HICSS 2021.</p> <p>Title of work: Ethical Perspectives in AI: A Two-folded Exploratory Study From Literature and Active Development Projects.</p> <p> </p> <p> </p> <p>This dataset depicts 589 GitHub README files explored in our article. We devised each repository into 4 categories. AI Applications (78), reference lists (486), Explainable AI tool (6), Ethical AI tool (15). Moreoveer, 4 repositories could not be found, and 181 of them had a Programming Language. In addition, we classified them according to our judgement of Interesting (82) or Great (33), regarding the objectives of our work. 75 repositories were not related to AI Ethics, and 7 of them mentioned COVID-19 in their repositories in some manner. Highlighted repositories in red pertain to categories 3 and 4, that is, tools for implementing AI Ethics, while those in yellow are marked as “Great”.</p> <p>Explainable AI tool and Ethical AI tool were added to 21 occurrences of tools for implementing AI ethics publicly available in repositories. Furthermore, it is seen that many papers open-sourced their codes on GitHub (as in https://github.com/lopusz/awesome-interpretable-machine-learning), meaning that the academic field has also made good progress in implementing ethics in AI, hence, the combined energy of both sources fosters an enhanced debate and stimulates progress towards AI ethics in practice.</p>
F.A.I.R. open dataset of brushed DC motor faults for testing of AI algorithms
<p>Practical research in AI often lacks of available and reliable datasets so the practitioners can try different algorithms. The field of predictive maintenance is particularly challenging in this aspect as many researchers don't have access to full-size industrial equipment or there is not available datasets representing a rich information content in different evolutions of faults.</p> <p>This dataset presents the evolution of typical faults (commutator, winding and brush wear) in inexpensive DC motors under extensive monitoring (vibration, temperature, voltage, current and noise). These motors exhibit a particularly short useful life when operating out of nominal conditions (from 30 minutes to 6 hours) which make them very interesting to test different signal processing algorithms and introduce students and researchers into signal processing, fault detection and predictive maintenance.</p> <p>The data-set comprises two main elements:</p> <ul> <li>A spreadsheet with the processing of each raw data file</li> <li>4 folders with raw data files in HDF5 format</li> </ul> <p>The spread sheet contains the following columns</p> <ul> <li>filename of the raw data file</li> <li>timestamp of the raw data file</li> <li>speed of the motor in rpm</li> <li>speed of the motor in Hz</li> <li>Current of the moter (A)</li> <li>Voltage supply (V)</li> <li>surface motor temperature (ºC)</li> <li>ambient temperature (ºC)</li> <li>For each measured signal (Vibration, current, voltage) the vibration of the main harmonic (at the speed of the motor) and it's first 10 multiple.</li> <li>For each measured signal (Vibration, current, voltage) the vibration in 4 bands: 0-4kHz, 4kHz-8kHz, 8kHz-16kHz, 16kHz-26kHz</li> </ul> <p>The raw data files in HDF5 format contains the instantaneous measured vibration (g), current (A) and voltage (volts) of the DC motor at 51.200 Hz.</p>
AI results complementing the Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Sweden
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2019, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
AI results complementing the Annual Report on surveillance for Avian Influenza in poultry and wild birds in Member States of the European Union - Slovakia
<p>This dataset contains the results of the EU co-funded surveillance activities conducted in 2019, which consisted of:</p> <ul> <li>Serological surveys to monitor the circulation of AIV subtypes H5 and H7 in poultry (active surveillance). These surveys should preferentially target poultry species or production systems with increased risk for introduction of avian influenza (AI).</li> <li>Passive surveillance aiming at the virological detection of AI in wild birds found dead or moribund.</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.