Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

9,300

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

9,300 results for “detection”

Learn how ShareScore rates datasets ↗
zenodo44/100

Today's cat is tomorrow's dog: accounting for time-based changes in the labels of ML vulnerability detection approaches (Replication Package Part 3: OpenSSL dataset)

<h1><strong>The Replication Package of</strong></h1> <h1><strong>"Today's cat is tomorrow's dog: accounting for time-based changes in the labels of ML vulnerability detection approaches"</strong></h1> <h3><strong>Part 3 (OPENSSL Dataset)</strong></h3> <div> <div>This repository includes:</div> <ol> <li><em><strong>Code.zip</strong></em> that contains the codes to replicate some parts of this study:<br>a.&nbsp;<em>1_generate_datasets</em> implements our methodology to generate the datasets.<br>b.&nbsp;<em>2_run_models</em> runs the ML models during the evaluation.<br>c.&nbsp;<em>3_result_replication </em>generates charts presented in the paper from the ML evaluation results.</li> <li><em><strong>Datasets.zip</strong></em> that contain 2 folders:<br>a.&nbsp;<em>original</em> datasets: 1 from <a href="https://github.com/CGCL-codes/VulDeePecker" target="_blank" rel="noopener">NVD Vuldeepecker</a> and 3 extracted from&nbsp;<a href="https://github.com/ZeoVan/MSR_20_Code_vulnerability_CSV_Dataset" target="_blank" rel="noopener">BigVul</a>.<br> <div> <div>b. <em>OPENSSL</em> datasets: train, validation, test sets for each time of observation extracted using our methodology from <a href="https://github.com/ZeoVan/MSR_20_Code_vulnerability_CSV_Dataset" target="_blank" rel="noopener">BigVul</a>&nbsp;dataset for project <em>openssl</em>.</div> </div> </li> <li><em><strong>Pretrained-models.zip</strong></em>&nbsp;that we generated during our evaluation (3 test results for each time point in the timeline [2013-2019]).</li> <li><em><strong>Results.zip</strong></em> of our evaluation, the folder <em>ALL</em> contains the overall results and other folders are results by model.</li> </ol> <p><strong>UPDATED version 5<br></strong>- added a GLOBAL_README.md which contains the 3 stages and how they are connected to each other<br>- updated LineVul.ipynb: import AdamW from torch.optim instead of transformers<br>- updated README.md in Code2Vec with the prerequisites of Java to run gradlew for astmine</p> <p><strong>UPDATED version 6<br></strong>- updated CodeBert.ipynb: import AdamW from torch.optim instead of transformers</p> <p>Documentations</p> <ol> <li><em><strong>INSTALL.pdf&nbsp;</strong></em>: how to install the codes</li> <li><em><strong>README.pdf</strong></em>: readme file</li> <li><em><strong>REQUIREMENTS.pdf</strong></em>: hardware and software requirements</li> <li><em><strong>STATUS.pdf</strong></em>&nbsp;: status for artifact submission</li> <li><em><strong>LICENSE.pdf</strong></em>: the license of this artifact</li> <li><em><strong>PAPER.pdf</strong></em>: the camera-ready version of the paper</li> </ol> </div> <div> <div>Please refer to the following repositories for the other datasets and pre-trained models:</div> <div>- Part 1 NVD Vuldeeepecker :&nbsp;<a href="https://doi.org/10.5281/zenodo.8207883" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.8207883</a></div> - Part 2 LINUX :&nbsp;<a href="https://doi.org/10.5281/zenodo.10960662" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10960662</a><br> <div>- Part 4 POPPLER : <a href="https://doi.org/10.5281/zenodo.14713143">https://doi.org/10.5281/zenodo.14713143</a></div> <div>&nbsp;</div> <div>This work was partly funded by the EU under the H2020 Program AssureMOSS (Grant n. 952647) and the Horizon Europe Program Sec4AI4Sec (Grant n. 101120393), by the Italian Ministry of University and Research (MUR) under the P.N.R.R. &ndash; NextGenerationEU grant n.\ PE00000014 (SERICS subproject COVERT), and by the Dutch Research Council (NWO) under the grant NWA.1215.18.006 (Theseus) and grant KIC1.VE01.20.004 (HEWSTI).&nbsp;</div> </div>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Challenges in Replay Detection by TDLM in Post-Encoding Resting State

<p>Data for the paper "Challenges in Replay Detection by TDLM in Post-Encoding Resting State".</p> <p>This extension of the previous dataset contains the resting state data. For each participant, a 8 minutes resting state was recorded before and after the main experiment (localizer plus learning), and before the final retrieval session.</p> <p>Two files are uploaded per participant, the pre-experiment resting state (RS1) and the post-learning resting state (RS2). All files are MaxFiltered and the head positioning has been realigned using MaxFilter movement correction to the head position during the initial localizer. The localizer data has been previously published and can be downloaded in v1 of this dataset at https://doi.org/10.5281/zenodo.8001755</p> <p>All relevant information can be found in the related publication. Behavioural data necessary to reproduce the results will be uploaded to GitHub at https://github.com/CIMH-Clinical-Psychology/DeSMRRest-TDLM-Simulation</p> <p>There are markers in the files as follows:</p> <p>###############################################################<br>## Port Trigger Table<br>## Port Code | Meaning<br>## ---------------------------------------------------<br>## 0 &nbsp; &nbsp; &nbsp;| don't send trigger<br>## 10 &nbsp; &nbsp; &nbsp; &nbsp; | start RS session<br>## 11 &nbsp; &nbsp; &nbsp; &nbsp; | end RS session<br>## 127 &nbsp; &nbsp;| button press has happened<br>## 255 &nbsp; &nbsp;| start and end of session<br>###############################################################</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

NeSy4VRD: A Multifaceted Resource for Neurosymbolic AI Research using Knowledge Graphs in Visual Relationship Detection

<p><strong>NeSy4VRD</strong></p> <p>NeSy4VRD is a multifaceted, multipurpose resource designed to foster neurosymbolic AI (NeSy) research, particularly NeSy research using Semantic Web technologies such as OWL ontologies, OWL-based knowledge graphs and OWL-based reasoning as symbolic components. The NeSy4VRD research resource pertains to the <em>computer vision</em> field of AI and, within that field, to the application tasks of <em>visual relationship detection (VRD) and scene graph generation</em>.</p> <p>Whilst the core&nbsp;motivation of the NeSy4VRD research resource is to foster computer vision-based NeSy research using Semantic Web technologies such as OWL ontologies and OWL-based knowledge graphs, AI researchers can readily use NeSy4VRD to either: 1) pursue computer vision-based NeSy research without involving Semantic Web technologies&nbsp;as symbolic components, or 2) pursue computer vision research&nbsp;without&nbsp;NeSy (i.e. pursue research that focuses purely on deep learning alone, without involving&nbsp;symbolic components of any kind).&nbsp; &nbsp;This is the sense in which we describe NeSy4VRD as being <em>multipurpose</em>: it can readily be used by diverse groups of computer vision-based&nbsp;AI researchers with diverse interests and objectives.</p> <p>The NeSy4VRD research resource in its entirety is distributed across two locations: Zenodo and GitHub.</p> <p>&nbsp;</p> <p><strong>NeSy4VRD on Zenodo: the NeSy4VRD dataset package</strong></p> <p>This entry on Zenodo hosts the <em>NeSy4VRD dataset package</em>, which includes the <em>NeSy4VRD dataset</em> and&nbsp;its companion <em>NeSy4VRD ontology</em>, an OWL ontology called VRD-World.</p> <p>The <em>NeSy4VRD dataset</em> consists of an image dataset with associated visual relationship annotations. The images of the <em>NeSy4VRD dataset</em> are the same as those that were once publicly available as part of the <a href="https://cs.stanford.edu/people/ranjaykrishna/vrd/">VRD</a> dataset. The NeSy4VRD visual relationship annotations&nbsp;are a highly customised and quality-improved version of the original VRD visual relationship annotations.&nbsp; The <em>NeSy4VRD dataset</em> is designed for computer vision-based research that involves detecting objects in images and predicting relationships between ordered pairs of those objects.&nbsp; A visual relationship for an image of the <em>NeSy4VRD dataset</em> has the form &lt;&#39;subject&#39;, &#39;predicate&#39;, &#39;object&#39;&gt;, where the &#39;subject&#39; and &#39;object&#39; are two objects in the image, and the &#39;predicate&#39; describes some relation between them.&nbsp; Both the &#39;subject&#39; and &#39;object&#39; objects are specified in terms of bounding boxes and object classes.&nbsp; For example, representative annotated visual relationships are &lt;&#39;person&#39;, &#39;ride&#39;, &#39;horse&#39;&gt;, &lt;&#39;hat&#39;, &#39;on&#39;, &#39;teddy bear&#39;&gt; and &lt;&#39;cat&#39;, &#39;under&#39;, &#39;pillow&#39;&gt;.</p> <p>Visual relationship detection is pursued as a computer vision application task in its own right, and as a building block capability for the broader application task of scene graph generation.&nbsp; Scene graph generation, in turn, is commonly used as a precursor to a variety of enriched, downstream visual understanding and reasoning application tasks, such as image captioning, visual question answering, image retrieval, image generation and multimedia event processing.</p> <p>The <em>NeSy4VRD ontology</em>, VRD-World, is a rich, well-aligned, companion OWL ontology engineered specifically for&nbsp;use with the <em>NeSy4VRD dataset.</em>&nbsp; It directly&nbsp;describes the domain of the <em>NeSy4VRD dataset</em>, as reflected in the NeSy4VRD visual relationship annotations.&nbsp; More specifically, all of the object classes that feature in the NeSy4VRD visual relationship annotations have corresponding classes within the VRD-World OWL class hierarchy, and all of the predicates that feature in the NeSy4VRD visual relationship annotations have corresponding properties within the VRD-World OWL object property hierarchy. The rich structure of the VRD-World class hierarchy and the rich characteristics and relationships of the VRD-World object properties together give the VRD-World OWL ontology rich inference semantics. These provide&nbsp;ample opportunity for OWL reasoning&nbsp;to be meaningfully exercised and exploited in NeSy research that uses OWL ontologies and OWL-based knowledge graphs as symbolic components.&nbsp; There is also ample potential for NeSy researchers to explore supplementing the OWL reasoning capabilities afforded by the VRD-World ontology with Datalog rules and reasoning.</p> <p>Use of the&nbsp;<em>NeSy4VRD ontology</em>, VRD-World, in conjunction with the&nbsp;<em>NeSy4VRD dataset </em>is, of course, purely optional, however.&nbsp; Computer vision AI researchers who have no interest in NeSy, or&nbsp;NeSy researchers who have no interest in OWL ontologies and&nbsp;OWL-based knowledge graphs, can ignore the <em>NeSy4VRD ontology</em>&nbsp;and use the&nbsp;<em>NeSy4VRD dataset </em>by itself.</p> <p>All computer vision-based&nbsp;AI research user groups can, if they wish, also avail themselves of the other components of the NeSy4VRD research resource available on GitHub.</p> <p>&nbsp;</p> <p><strong>NeSy4VRD on GitHub: open source infrastructure supporting extensibility, and sample code</strong></p> <p>The NeSy4VRD research resource incorporates additional components that are companions to the&nbsp;<em>NeSy4VRD dataset package</em> here on Zenodo.&nbsp; These companion components are available&nbsp;at <a href="https://github.com/djherron/NeSy4VRD/">NeSy4VRD on GitHub</a>. These companion components consist of:</p> <ul> <li>comprehensive open source Python-based&nbsp;infrastructure supporting the extensibility of the NeSy4VRD visual relationship annotations (and, thereby, the extensibility of the <em>NeSy4VRD ontology</em>, VRD-World, as well)</li> <li>open source Python sample code showing how one can work&nbsp;with the&nbsp;NeSy4VRD visual relationship annotations in conjunction with the <em>NeSy4VRD ontology</em>, VRD-World, and RDF knowledge graphs.</li> </ul> <p>The NeSy4VRD infrastructure supporting extensibility consists of:</p> <ul> <li>open source Python code for conducting deep and comprehensive analyses of the <em>NeSy4VRD dataset</em> (the VRD images and their associated NeSy4VRD visual relationship annotations)</li> <li>an open source, custom-designed <em>NeSy4VRD protocol</em> for specifying visual relationship annotation customisation instructions declaratively, in text files</li> <li>an open source, custom-designed <em>NeSy4VRD workflow,&nbsp;</em>implemented using&nbsp;Python scripts and modules,&nbsp;for applying small or large volumes of customisations or extensions to the NeSy4VRD visual relationship annotations in a configurable, managed, automated and repeatable process.</li> </ul> <p>The purpose behind providing comprehensive infrastructure to support extensibility of the NeSy4VRD visual relationship annotations is to make it easy for researchers to take the <em>NeSy4VRD dataset</em> in new directions, by further enriching the&nbsp;annotations, or by tailoring them&nbsp;to introduce new or more data conditions that better&nbsp;suit&nbsp;their particular research needs and interests.&nbsp; The option to use the NeSy4VRD extensibility infrastructure in this way applies equally well to each of the diverse potential NeSy4VRD user groups already mentioned.</p> <p>The NeSy4VRD extensibility infrastructure, however, may be of particular interest to NeSy researchers interested in&nbsp;using the <em>NeSy4VRD ontology</em>, VRD-World, in conjunction with the <em>NeSy4VRD dataset. </em>These researchers can of course&nbsp;tailor the VRD-World ontology if they wish&nbsp;without needing to modify&nbsp;or extend&nbsp;the NeSy4VRD visual relationship annotations in any way. But their degrees of freedom for doing so will be limited by the need to maintain alignment with the NeSy4VRD visual relationship annotations and the particular set of object classes and predicates to which they refer.&nbsp; If NeSy researchers want full freedom to tailor the VRD-World ontology, they may well need to tailor the NeSy4VRD visual relationship annotations first, in order that alignment be maintained.</p> <p>To illustrate our point, and to illustrate our vision of how the NeSy4VRD extensibility infrastructure can be used, let us consider a simple example.&nbsp;It is common in computer vision to distinguish between <em>thing</em> objects (that have well-defined shapes) and <em>stuff</em> objects (that are amorphous). Suppose a researcher&nbsp;wishes to have a greater number of <em>stuff</em> object classes with which to work.&nbsp; Water is such a <em>stuff</em> object.&nbsp; Many VRD images contain water but it is not currently one of the&nbsp;annotated object classes and hence is never&nbsp;referenced in any visual relationship annotations. So adding a <em>Water</em> class to the class hierarchy of the VRD-World ontology would be pointless because it would never acquire any instances (because an object detector would never detect any). However, our hypothetical researcher could choose to&nbsp;do the following:</p> <ul> <li>use the analysis functionality of the NeSy4VRD extensibility infrastructure to find images containing water (by, say, searching for images whose visual relationships refer to object classes such as &#39;boat&#39;, &#39;surfboard&#39;, &#39;sand&#39;, &#39;umbrella&#39;, etc.);</li> <li>use free image analysis software (such as GIMP, at gimp.org) to get bounding boxes for instances of water in these images;</li> <li>use the <em>NeSy4VRD protocol</em> to specify new visual relationships for these images&nbsp;that refer to the new &#39;water&#39; objects&nbsp;(e.g. &lt;&#39;boat&#39;, &#39;on&#39;, &#39;water&#39;&gt;);</li> <li>use the <em>NeSy4VRD workflow</em> to introduce&nbsp;the new object class &#39;water&#39;&nbsp;and to apply the&nbsp;specified&nbsp;new visual relationships to the sets of annotations for the affected&nbsp;images;</li> <li>introduce class Water to the class hierarchy of the VRD-World ontology (using, say, the free Protege ontology editor);</li> <li>continue experimenting, now with the added benefit of the additional <em>stuff</em> object class &#39;water&#39;;</li> <li>contribute the enriched set of NeSy4VRD visual relationship annotations, and the enriched companion VRD-World ontology, to research communities.</li> </ul> <p>&nbsp;</p> <p><strong>Information pertaining to the VRD dataset</strong></p> <p>Information about the original VRD dataset&nbsp;is available <a href="https://cs.stanford.edu/people/ranjaykrishna/vrd/">here</a>.&nbsp;</p> <p>Public availability of the VRD images (via information accessible from that location) ceased sometime in the latter part of 2021.&nbsp; We thank Dr. Ranjay Krishna, one of the principals associated with the VRD dataset, for granting us permission to re-establish the public availability of the VRD images as part of NeSy4VRD.</p> <p>The original VRD visual relationship annotations&nbsp;are still publicly available from that location.&nbsp; But our deep analysis of those annotations, driven by our desire to design a robust companion ontology,&nbsp;revealed them to be highly problematic in many ways that made credible ontology modelling infeasible.&nbsp; They were also found to be replete with all manner of&nbsp;errors.&nbsp; The NeSy4VRD visual relationship annotations are far superior and we recommend them over the original VRD annotations to anyone contemplating conducting research using the VRD images.&nbsp; The&nbsp;NeSy4VRD annotations also have the added benefit of the rich, well-aligned companion <em>NeSy4VRD ontology</em>, VRD-World, for those whose research requires such a companion ontology.</p> <p>Researchers wishing to use the original VRD dataset may still do so. They can access the VRD images here, from within the <em>NeSy4VRD dataset</em> on Zenodo, and access the VRD visual relationship annotations from the location in the link.</p> <p><em>A note of caution</em>: the <em>NeSy4VRD ontology</em>, VRD-World, is <em>not</em><strong>&nbsp;</strong>compatible with the original VRD visual relationship annotations and cannot be used in conjunction with them.&nbsp; The VRD-World ontology has been engineered in relation to the highly customised and quality-improved NeSy4VRD visual relationship annotations. The&nbsp;customisations that were applied&nbsp;include ones&nbsp;that introduced many new object classes, merged some of the existing object classes, introduced one new predicate,&nbsp;and changed several predicate names.</p> <p>However, researchers&nbsp;can, if they wish, use the NeSy4VRD&nbsp;extensibility infrastructure (described above) to undertake their own customisation and quality-improvement exercise&nbsp;with respect to the original VRD visual relationship annotations. This is precisely how the NeSy4VRD visual relationship annotations were created in the first place. The primary intended use case of NeSy4VRD&#39;s extensibility infrastructure, however, is for researchers to use the NeSy4VRD visual relationship annotations as their starting point, and to take these annotations forward with onward customisations and extensions, as illustrated in the example use case given above.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo44/100

TAMPAR: Visual Tampering Detection for Parcels Logistics in Postal Supply Chains

<p>TAMPAR is a real-world dataset of parcel photos for tampering detection with annotations in <a href="https://cocodataset.org/#format-data">COCO format</a>. For details see our paper and for visual samples our <a href="https://a-nau.github.io/tampar/">project page</a>. Features are:&nbsp;</p><ul><li>&gt;900 annotated real-world images with &gt;2,700 visible parcel side surfaces</li><li>6 different tampering types</li><li>6 different distortion strengths</li></ul><p>Relevant computer vision tasks:</p><ul><li>bounding box detection</li><li>classification</li><li>instance segmentation</li><li>keypoint estimation</li><li>tampering detection and classification</li></ul><p>If you use this resource for scientific research, please consider citing our WACV 2024 <a href="https://arxiv.org/abs/2311.03124">paper</a> <i>"TAMPAR: Visual Tampering Detection for Parcel Logistics in Postal Supply Chains".</i></p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Phlorest phylogeny derived from Hruschka et al. 2015 'Detecting regular sound changes in linguistics as events of concerted evolution'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Hruschka, D. J., Branford, S., Smith, E. D., Wilkins, J., Meade, A., Pagel, M., &amp; Bhattacharya, T. (2015). Detecting regular sound changes in linguistics as events of concerted evolution. Current Biology, 25(1), 1-9.</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Identification of biomarkers for the early detection of non-small cell lung cancer: a systematic review and meta-analysis

<p>We sought to identify the best biomarkers for the early diagnosis of LC, using a systematic review of seven databases. We identified 79 articles that focused on the identification and assessment of diagnostic biomarkers and then performed a meta-analysis. This work has been submitted for publication.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Towards an open pipeline for the detection of Critical Infrastructure from satellite imagery – A case study on electrical substations in The Netherlands

<p><strong>Abstract.</strong> Critical infrastructure (CI) are at risk of failure due to the increased frequency and magnitude of climate extremes related to climate change. It is thus essential to include them in a risk management framework to identify risk hotspots, develop risk management policies and support adaptation strategies to enhance their resilience. However, the lack of information on the exposure of CI prevents their incorporation in large-scale risk assessment studies. This study sets out to improve the representation of CI for risk assessment studies by building a neural network model to detect CI assets from optical remote sensing imagery. We present a pipeline that extracts CI from OpenStreetMaps, processes the imagery and assets' masks, and trains a Mask R-CNN model that allows for instance segmentation of CI at the asset level. This study provides an overview of the pipeline and tests it with the detection of electrical substations assets in the Netherlands. Several experiments are presented for different under-sampling percentages of the majority class (25%, 50% and 100%) and hyperparameters settings (batch size and learning rate). The best metrics achieved are an Average Precision at an Intersection over Union of 50% of 30.93 and a tile F-score of 89.88%. This allows us to confirm the feasibility of the method and invite disaster risk researchers to use this pipeline for other infrastructure types. We conclude by exploring the different avenues to improve the pipeline by addressing the class imbalance, Transfer Learning and Explainable AI.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Integrated disturbance mapping over the Tibetan Plateau based on multiple detection algorithms

<p>This dataset contains the mapped representation of vegetation disturbance across the Tibetan Plateau, rendered at a 30-meter spatial resolution. The dataset comprehensively illustrates the extent of vegetation disturbance on the Tibetan Plateau during the period spanning 1986 to 2020. Notably, the map employs a categorization system, with values assigned to distinct classes: 0 denotes areas of disturbed vegetation, 1 denotes regions characterized by undisturbed vegetation, and 2 denotes zones devoid of vegetation.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Spectrophotometric Assay for the Detection of 2,5-Diformylfuran and Its Validation through Laccase-Mediated Oxidation of 5-Hydroxymethylfurfural

<p>Modern biocatalysis requires fast, sensitive, and efficient high-throughput screening methods to screen enzyme libraries in order to seek out novel biocatalysts or enhanced variants for the production of chemicals. For instance, the synthesis of bio-based furan compounds like 2,5-diformylfuran (DFF) from 5-hydroxymethylfurfural (HMF) via aerobic oxidation is a crucial process in industrial chemistry. Laccases, known for their mild operating conditions, independence from cofactors, and versatility with various substrates, thanks to the use of chemical mediators, are appealing candidates for catalyzing HMF oxidation. Herein, Schiff-based polymers based on the coupling of DFF and 1,4-phenylenediamine (PPD) have been used in the set-up of a novel colorimetric assay for detecting the presence of DFF in different reaction mixtures. This method may be employed for the fast screening of enzymes (Z' values ranging from 0.68 to 0.72). The sensitivity of the method has been proved, and detection (8.4 μM) and quantification (25.5 μM) limits have been calculated. Notably, the assay displayed selectivity for DFF and enabled the measurement of kinetics in DFF production from HMF using three distinct laccase–mediator systems.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Ground truth and raw hyperspectral files of olive trees for plant stress detection

<p>This dataset contains raw hyperspectral images from Cubert S-185 collected on 13 May 2021 from an olive field in Halkidiki, Northern Greece. Included is also a matrix containing the id of each recorded olive tree (the samples) that also appears in the hyperspectral images. QGIS (ver.3.28.0) software plugin 'zonal statistics multiband' was used to compute zonal statistics for each of the 138 spectral bands available for each sample. Accompanying each sample is also the ground truthing data recorded, which addresses the present stress of 3 stressors (<i>Verticillium dahliae, Pleospora herbarum </i>and 'other stressors').</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Dataset: Submicron‐ and Nanoplastic Detection at Low Micro‐ to Nanogram Concentrations Using Gold Nanostar‐Based Surface‐Enhanced Raman Scattering (SERS) Substrates

<p>ABSTRACT</p> <p>The presence of submicron- (1 &micro;m &ndash; 100 nm) and nanoplastic (&lt; 100 nm) particles within various sample matrices, ranging from marine environments to foods and beverages, has become a topic of increasing interest in recent years. Despite this interest, very few analytical techniques remain that allow for the detection of these small plastic particles in the low concentration ranges that they are anticipated to be present at. Research focused on optimizing surface-enhanced Raman scattering (SERS) to enhance signal obtained in Raman spectroscopy has been shown to have great potential for the detection of plastic particles below conventional resolution limits. In this study, we produce SERS substrates composed of gold nanostars and assess their potential for submicron- and nanoplastic detection. The results show 33 nm polystyrene could be detected down to 1.25 &micro;g/mL while 36 nm poly(ethylene terephthalate) was detected down to 5 &micro;g/mL. These results confirm the promising potential of the gold nanostar-based SERS substrates for nanoplastic detection. Furthermore, combined with findings for 121 nm polypropylene and 126 nm polyethylene particles, they highlight potential differences in analytical performance that depend on the properties of the plastics being studied.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Fusion of Underwater Camera and Multibeam Sonar for Diver Detection and Tracking

<div><strong>Context</strong></div> <div>&nbsp;</div> <div>This dataset is related to previously published public dataset "Sonar-to-RGB Image Translation for Diver Monitoring in Poor Visibility Environments".&nbsp;<a href="../records/7728089">https://zenodo.org/records/7728089</a></div> <div>It contains ZED-right camera and sonar images collected from Hemmoor Lake and DFKI Maritime Exploration Hall.</div> <div>&nbsp;</div> <div>Sensors: Low Frq (1.2MHz) Blueprint Oculus M1200d Sonar and ZED Right Camera</div> <div>&nbsp;</div> <div><strong>Content</strong></div> <div>&nbsp;</div> <div>The dataset is created for Diver Detection and Diver Tracking applications.</div> <div>&nbsp;</div> <div>For the Diver Detection part, the dataset is prepared to train, validate and test YOLOv7 model.</div> <div>7095 images are used for training data, and 3095 images are used for validation data. These sets are augmented from originally captured and sampled ZED camera images.&nbsp;Augmentation methods are not applied to the Test data, which contains 822 images. Train and validation contain images from both the DFKI pool and Hemmor Lake, while the test data is only collected from the lake.</div> <div>&nbsp;</div> <div>To distinguish between the original image and the augmented image, check the name coding.&nbsp;</div> <div>Naming of object detection images:</div> <div>original_image_name.jpg</div> <div>if augmented:</div> <div>original_image_name_&lt;augmentation_number_of_the_same_image&gt;.jpg</div> <div>&nbsp;</div> <div>Object Detection Label Format:&nbsp;</div> <div>YOLO [(class), ((x_min + (x_max - x_min)/2)&nbsp; / image_width), ((y_min + (y_max - y_min)/2)&nbsp; / image_height), ((x_max - x_min) / image_width), ((y_max - y_min) / image_height)]</div> <div>&nbsp;</div> <div>Class: "diver", represented by "0" in object detection labels.</div> <div>&nbsp;</div> <div>Resolution of Object Detection Camera Images: 640x640</div> <div>Resolution of Object Tracking Camera Images: 1280x720</div> <div>Resolution of Object Tracking Low Frequency Sonar: 932x514</div> <div>&nbsp;</div> <div>About the Object Tracking on Sonar, the sampled data is the part where diver moves around the table and the platform.&nbsp;</div> <div>There are 4 cases shared in the dataset, which contain a sonar stream, and corresponding ZED-right camera images.&nbsp;</div> <div>Totally, 1193 points represent the diver on sonar images for the diver tracking application.</div> <div>&nbsp;</div> <div>For the tracking, "tracking_sonar_coordinates_&lt;number&gt;.csv" contains x,y coordinates of a point where the diver is in the sonar image.&nbsp;</div> <div>And "image_sonar_&lt;number&gt;.csv" file contains the matching between sonar and camera images.</div> <div>&nbsp;</div> <div><strong>Acknowledgements</strong></div> <div>&nbsp;</div> <div>The data in this repository were collected as a joint effort between the German Center for Artificial Intelligence (DFKI), the German Federal Agency for technical Relief (THW), and Kraken Robotics GmbH. This work is part of the project DeeperSense that received funding from the European Commission. Program H2020-ICT-2020-2 ICT-47-2020 Project Number: 101016958.</div> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Automated sarcomere detection Matlab tool

<p>The Matlab code presented was developed by Dr. Ian Estabrook in 2019-2023 to automatically detect sarcomeres in multi-channel z-stack images of developing myofibrils, as used in the preprint: "A tension-driven sarcomere division mechanism facilitates muscle growth" by Clement Rodier, Ian Estabrook, Vincent Loreau, Dirk G&ouml;rlich, Benjamin M. Friedrich, Frank Schnorrer. We thank Yasmin Magdy Emadeldin Mohamed Abdelghaffar for help with preparing this repository and the documentation of the code. An example data set with corresponding sarcomere tracking is included.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Toxic Content Detection in online social networks: a new dataset from Brazilian Reddit Communities

<p>This is new dataset of 2,500 manually annotated examples of comments extracted from the top 10 largest Brazilian subreddits on Reddit. The dataset has been annotated by crowd-sourcing efforts with contributions from the departments of computer science (DCC) and the linguistic group @ UFMG. As part of our contribution to the toxicity automatic detection and moderation of online social networks, we're making the dataset public for research.</p> <h3>Dataset</h3> <p>The dataset contains 2,500 manually annotated comments from the most popular brazilian communities on Reddit. The data sampling proccess was a stratified sampling by the number of generated publications by subreddit and the month of publication. The list of communities collected is presented below. The collected data period ranges from January 2022 to December 2022.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Subreddit</strong></td> <td><strong>Posts</strong></td> <td><strong>Comments</strong></td> </tr> <tr> <td>r/brasil</td> <td>110,829&nbsp;</td> <td>2,136,866</td> </tr> <tr> <td>r/desabafos</td> <td>115,876</td> <td>1,211,643</td> </tr> <tr> <td>r/futebol</td> <td>35,826</td> <td>1,214,412</td> </tr> <tr> <td>r/saopaulo</td> <td>7,308</td> <td>81,969</td> </tr> <tr> <td>r/eu_nvr</td> <td>12,631</td> <td>188,620</td> </tr> <tr> <td>r/botecodoreddit</td> <td>7,059</td> <td>57,298</td> </tr> <tr> <td>r/conversas</td> <td>21,967</td> <td>326,061</td> </tr> <tr> <td>r/investimentos</td> <td>9,756</td> <td>141,823</td> </tr> <tr> <td>r/tiodopave</td> <td>2,371</td> <td>11,584</td> </tr> <tr> <td>r/brasilivre</td> <td>67,301</td> <td>1,219265</td> </tr> <tr> <td>Total</td> <td>390,924</td> <td>6,589,541</td> </tr> </tbody> </table> <p>&nbsp;</p> <h3><strong>Annotation proccess</strong></h3> <p>The annotators were divided into groups of raters and each group was assigned a batch of comments to label. The raters were then asked to label a comment as <strong>Toxic</strong>, <strong>Non-toxic</strong>, <strong>I do not know</strong> and <strong>Missing info</strong>. During the annotation process, the raters were encouraged to assign one of the uncertain labels when they're not sure about the toxicity of a comment or the context is missing.&nbsp;</p> <h3>Available data</h3> <p>The dataset is available as csv file and the label was assigned as a majority vote among the raters. The available data are the original collected comment id and body. The label was created from the original classification from the annotators. No data processing has been done on this version of the dataset. The overall schema of the dataset if presented below.</p> <p>- <strong>id</strong>: The unique identifier of the comment on the Reddit platform<br>- <strong>body</strong>: The original comment text publication<br>- <strong>is_toxic</strong>: The final label of a given comment. The label is <strong>0</strong> for non-toxic comments, <strong>1</strong> for toxic comments and <strong>-1</strong> for comments where the raters disagreed about the toxicity.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Low-entropy Packed Binary Detection using Hardware Performance Counters

<p><span>Malware analysis faces a critical challenge in accurately identifying&nbsp;packed executables, especially those with low entropy. Existing&nbsp;software-based solutions often fail in detecting packers used by&nbsp;malware, resulting in inaccurate classifications. To address this&nbsp;shortcoming, in this study we introduce a novel method using<br>Hardware Performance Counters (HPCs) to facilitate the classification of binary packers due to HPCs&rsquo; minimal access overhead&nbsp;and ability to obviate the necessity for source code. We trained&nbsp;classic machine-learning models by selecting relevant hardware&nbsp;attributes associated with the unpacking procedure for detecting<br>packers used by low-entropy binary programs. Extensive experiments shows the substantial role played by Hardware Performance&nbsp;Counters in detecting binary packing characterized by low entropy,<br>offering a promising avenue for further exploration and refinement&nbsp;of techniques in malware analysis<br><br><br></span></p> <p><span>The following zip files are executables that represent low entropy versions of software packers using byte-padding. The name of the files are the names of the packers which are represened,&nbsp; Acprotect, Armadillo, Aspack, Nspack, Pecompact, Petite, UPX, and Zprotect. These can be used to measure the unpacking process using hardware performance counters in order to test &amp; train machine earning classifiers for accurate classification of low entropy packers.</span></p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Auditory Attention Detection Dataset KULeuven

<div>***************************************</div> <p>Please cite the original paper where this data set was presented:</p> <p>Biesmans, W., Das, N., Francart, T., &amp; Bertrand, A. (2016). Auditory-inspired speech envelope extraction methods for improved EEG-based auditory attention detection in a cocktail party scenario. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 25(5), 402-412.</p> <p>***************************************</p> <p><strong>IMPORTANT MESSAGE FROM THE AUTHORS (January 2024):</strong></p> <p>We have observed that this dataset is widely used in research, establishing it as a standard benchmark for evaluating novel decoding strategies in auditory attention decoding (AAD). We emphasize the critical importance of rigorous cross-validation in such studies. In particular, researchers should be aware of two common and significant validation pitfalls:</p> <ol> <li> <p><strong>Trial fingerprints</strong>: Avoid splitting data from the same experimental trial into training and testing segments. Classifiers can detect whether a test segment belongs to a specific trial if other, non-overlapping segments from that trial are in the training set.</p> </li> <li> <p><strong>Gaze bias</strong>: AAD algorithms that directly classify EEG signals (often referred to as spatial AAD or SpAAD) without explicitly correlating the decoder output signal to the speech stimulus, should <em>not </em>be evaluated on this dataset, as it is affected by gaze-related bias. Instead, use the gaze-controlled dataset of Rotaru et al. available at <a href="https://zenodo.org/records/11058711">https://zenodo.org/records/11058711</a></p> </li> </ol> <p>Further details on both issues are provided below.</p> <p>In the original study by Biesmans et al., which produced this dataset, linear correlation-based methods were employed, and a straightforward random cross-validation sufficed. However, with the recent surge in the application of machine learning techniques, particularly deep neural networks, in tackling the AAD challenge, a more stringent cross-validation approach becomes imperative. Deep networks are susceptible to overfitting to trial-specific patterns in EEG data, even from very brief segments (less than 1 second), leading to the ability to identify the trial source. Since a subject typically maintains attention to the same speaker throughout a trial, having knowledge of the trial effectively results in a perfect attention decoding.</p> <p>We observed that many research papers utilizing our dataset still adhere to the basic random cross-validation method, neglecting the separation of trials into training and testing sets. Consequently, these studies frequently report remarkably high AAD accuracies when using extremely short EEG segments (one or a few seconds). Nevertheless, research has demonstrated that such an approach yields inaccurate and excessively optimistic outcomes. Accuracies often plummet significantly, sometimes even falling below chance levels, when employing a proper cross-validation where this trial bias is removed (e.g., leave-one-trial-out, leave-one-story-out, or leave-one-subject-out cross-validation).</p> <p>This overfitting effect is described in:&nbsp;Corentin Puffay et al., "Relating EEG to continuous speech using deep neural networks: a review", Journal of Neural Engineering 20, 041003, 2023 DOI:10.1088/1741-2552/ace73f</p> <p>Moreover, it's important to note that AAD strategies which directly classify an EEG snippet, rather than explicitly computing a correlation between the decoder output and the corresponding speech envelopes, may be susceptible to an eye-gaze bias. This bias refers to the tendency of the subject to subtly and often unknowingly direct their gaze towards the attended speaker. Given that EEG equipment can inadvertently capture these gaze patterns, it becomes possible to leverage this gaze information, whether intentionally or unintentionally, to enhance AAD performance. It's crucial to highlight that there is a relatively strong eye gaze bias in this dataset (such gaze bias is present in the majority of public AAD datasets).</p> <p>This eye-gaze overfitting effects is discussed in:&nbsp;Rotaru et al. "What are we really decoding? Unveiling biases in EEG-based decoding of the spatial focus of auditory attention", Journal of Neural Engineering, vol. 21, 016017, DOI: https://doi.org/10.1088/1741-2552/ad2214. Also available on bioRxiv: https://doi.org/10.1101/2023.07.13.548824</p> <p>To test whether your method is not using gaze as a shortcut, use the Rotaru et al. data set available at <a href="https://zenodo.org/records/11058711">https://zenodo.org/records/11058711</a></p> <p>***************************************</p> <p>Explanation about the data set:</p> <p>This work was done at ExpORL, Dept. Neurosciences, KULeuven and Dept. Electrical Engineering (ESAT), KULeuven.<br>This dataset contains EEG data collected from 16 normal-hearing subjects. EEG recordings were made in a soundproof, electromagnetically shielded room at ExpORL, KULeuven. The BioSemi ActiveTwo system was used to record 64-channel EEG signals at 8196 Hz sample rate. The audio signals, low pass filtered at 4 kHz, were administered to each subject at 60 dBA through a pair of insert phones (Etymotic ER3A). The experiments were conducted using the APEX 3 program developed at ExpORL [1].</p> <p>Four Dutch short stories [2], narrated by different male speakers, were used as stimuli. All silences longer than 500 ms in the audio files were truncated to 500 ms. Each story was divided into two parts of approximately 6 minutes each. During a presentation, the subjects were presented with the six-minutes part of two (out of four) stories played simultaneously. There were two stimulus conditions, i.e., `HRTF' or `dry' (dichotic). &nbsp;An experiment here is defined as a sequence of 4 presentations, 2 for each stimulus condition and ear of stimulation, with questions asked to the subject after each presentation. All subjects sat through three experiments within a single recording session. An example for the design of an experiment is shown in Table 1 in [3]. The first two experiments included four presentations each. &nbsp;During a presentation, the subjects were instructed to listen to the story in one ear, while ignoring the story in the other ear. After each presentation, the subjects were presented with a set of multiple-choice questions about the story they were listening to in order to help them stay motivated to focus on the task. In the next presentation, the subjects were presented with the next part of the two stories. This time they were instructed to attend to their other ear. In this manner, one experiment involved four presentations in which the subjects listened to a total of two stories, switching attended ear between presentations. The second experiment had the same design but with two other stories. Note that the Table was different for each subject or recording session, i.e., each of the elements in the table were permuted between different recording sessions to ensure that the different conditions (stimulus condition and the attended ear) were equally distributed over the four presentations. Finally, the third experiment included a set of presentations where the first two minutes of the story parts from the first experiment, i.e., a total of four shorter presentations, were repeated three times, to build a set of recordings of repetitions. Thus, a total of approximately 72 minutes of EEG was recorded per subject.&nbsp;</p> <p>We refer to EEG recorded from each presentation as a trial. For each subject, we recorded 20 trials - 4 from &nbsp;the first experiment, 4 from the second experiment, and 12 from the third experiment (first 2 minutes of the 4 presentations from experiment 1 X 3 repetitions). The EEG data is stored in subject specific mat files of the format 'Sx', 'x' referring to the subject number. The audio data is stored as wav files in the folder 'stimuli'. Please note that the stories were not of equal lengths, and the subjects were allowed to finish listening to a story, even in cases where the competing story was over. Therefore, for each trial, we suggest referring to the length of the EEG recordings to truncate the ends of the corresponding audio data. This will ensure that the processed data (EEG and audio) contains only competing talker scenarios. Each trial was high-pass filtered &nbsp;(0.5 Hz cut off) and downsampled from the recorded sampling rate of 8192 Hz to 128 Hz. Artifacts were removed using the MWF-filtering method in [4]. Please get in touch with the team (of Prof. Alexander Bertrand or Prof. Tom Francart) if you wish to obtain the raw EEG data (without the mentioned high-pass filtering and artifact removal).</p> <p>Each trial (trial*.mat) contains the following information:&nbsp;</p> <p><strong>RawData.Channels</strong> : channel numbers (1 to 64)<br><strong>RawData.EegData</strong> : &nbsp; EEG data (samples X channels)<br><strong>FileHeader.SampleRate</strong> : Sampling frequency of the saved data<br><strong>TrialID</strong> : a number between 1 to 20, showing the trial number<br><strong>attended_ear</strong> : the direction of attention of the subject. 'L' for left, 'R' for right<br><strong>stimuli</strong> : cell array with stimuli{1} and stimuli{2} indicating the name of audio files presented in the left ear and the right ear of the subject respectively<br><strong>condition</strong> : stimulus presentation condition. 'HRTF' - stimuli were filtered with HRTF functions to simulate audio from 90 degrees to the left and 90 degrees to the right of the speaker, 'dry' - a dichotic presentation in which there was one story track each presented separately via the left and the right earphones.<br><strong>experiment</strong> : the number of the experiment (1,2 or 3)<br><strong>part</strong> : part of the story track being presented (can be 1 to 4 for experiments 1 and 2, and 1 to 12 for experiment 3)<br><strong>attended_track</strong> : the attended story track. '1' for track 1 and '2' for track 2. Each track maintains continuity of the story. In Experiment 1, attention is always to track 1, and in Experiment 2, attention is always to track 2.&nbsp;<br><strong>repetition</strong> : binary variable indicating where the trial is a repetition (of presented stimuli) or not.<br><strong>subject</strong> : subject id of the format 'Sx', 'x' being the subject number.</p> <p>The 'stimuli' folder contains wav files of the format: part{part number}_track{track number}_{condition}.wav. Although the folder contains stimuli with HRTF filtering as well, for the analysis, we have assumed knowledge of the original clean stimuli (i.e. stimuli presented under the 'dry' condition), and hence envelopes were extracted only from part{part number}_track{tracknumber}_dry.wav files.</p> <p>The Matlab file 'preprocess_data.m' gives an example of how the synchronization and preprocessing of EEG and audio data can be done as described in [5]. Dependency: AMToolbox.</p> <p>This dataset has been used in [3, 5-14] (not updated anymore).&nbsp;</p> <p>[1] Francart, T., Van Wieringen, A., &amp; Wouters, J. (2008). APEX 3: a multi-purpose test platform for auditory psychophysical experiments. <em>Journal of neuroscience methods</em>, 172(2), 283-293.<br>[2] Radioboeken voor kinderen, <a href="http://radioboeken.eu/kinderradioboeken.php?lang=NL">http://radioboeken.eu/kinderradioboeken.php?lang=NL</a>, 2007 (Accessed: 30 March 2015)<br>[3] Das, N., Biesmans, W., Bertrand, A., &amp; Francart, T. (2016). The effect of head-related filtering and ear-specific decoding bias on auditory attention detection.<em> Journal of neural engineering</em>, 13(5), 056014.<br>[4] Somers, B., Francart, T., &amp; Bertrand, A. (2018). A generic EEG artifact removal algorithm based on the multi-channel Wiener filter. <em>Journal of neural engineering</em>, 15(3), 036007.<br>[5] Das, N., Vanthornhout, J., Francart, T., &amp; Bertrand, A. (2019). Stimulus-aware spatial filtering for single-trial neural response and temporal response function estimation in high-density EEG with applications in auditory research. <em>bioRxiv</em> 541318; doi: <a href="https://doi.org/10.1101/541318">https://doi.org/10.1101/541318</a><br>[6] Biesmans, W., Das, N., Francart, T., &amp; Bertrand, A. (2016). Auditory-inspired speech envelope extraction methods for improved EEG-based auditory attention detection in a cocktail party scenario. <em>IEEE Transactions on Neural Systems and Rehabilitation Engineering</em>, 25(5), 402-412.<br>[7] Das, N., Van Eyndhoven, S., Francart, T., &amp; Bertrand, A. (2016, August). Adaptive attention-driven speech enhancement for EEG-informed hearing prostheses. In 2016 <em>38th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)</em> (pp. 77-80). IEEE.<br>[8] Van Eyndhoven, S., Francart, T., &amp; Bertrand, A. (2016). EEG-informed attended speaker extraction from recorded speech mixtures with application in neuro-steered hearing prostheses. <em>IEEE Transactions on Biomedical Engineering</em>, 64(5), 1045-1056.<br>[9] Das, N., Van Eyndhoven, S., Francart, T., &amp; Bertrand, A. (2017, August). EEG-based attention-driven speech enhancement for noisy speech mixtures using N-fold multi-channel Wiener filters. In 2017 <em>25th European Signal Processing Conference (EUSIPCO)</em> (pp. 1660-1664). IEEE.<br>[10] Narayanan, A. M., &amp; Bertrand, A. (2018, July). The effect of miniaturization and galvanic separation of EEG sensor devices in an auditory attention detection task. In 2018 <em>40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)</em> (pp. 77-80). IEEE.<br>[11] Deckers, L., Das, N., Ansari, A. H., Bertrand, A., &amp; Francart, T. (2018). EEG-based detection of the attended speaker and the locus of auditory attention with convolutional neural networks. <em>bioRxiv </em>475673; doi: <a href="https://doi.org/10.1101/475673">https://doi.org/10.1101/475673</a><br>[12] Narayanan, A. M., &amp; Bertrand, A. (2019). Analysis of miniaturization effects and channel selection strategies for EEG sensor networks with application to auditory attention detection. <em>IEEE Transactions on Biomedical Engineering</em>.<br>[13] Geirnaert, S., Francart, T., &amp; Bertrand, A. A New Metric to evaluate auditory attention detection performance based on a Markov chain. Accepted for publication in <em>Proc. European Signal Processing Conference (EUSIPCO)</em>, A Coruna, Spain, Sep. 2019.<br>[14] Geirnaert, S., Francart,T., Bertrand A. (2019). An &nbsp;Interpretable performance metric for auditory attention decoding algorithms in &nbsp;a &nbsp;context &nbsp;of &nbsp;neuro-steered &nbsp;gain &nbsp;control. <em>bioRxiv </em>745695; doi: <a href="https://doi.org/10.1101/745695">https://doi.org/10.1101/745695</a>&nbsp;</p>

opencc-by-nc-sa-4.0Aug 2019View details →
zenodo44/100

The Automotive Visual Inspection Dataset (AutoVI): A Genuine Industrial Production Dataset for Unsupervised Anomaly Detection

<p><strong>See the official website: <a href="https://autovi.utc.fr">https://autovi.utc.fr</a></strong></p> <p>Modern industrial production lines must be set up with robust defect inspection modules that are able to withstand high product variability. This means that in a context of industrial production, new defects that are not yet known may appear, and must therefore be identified.</p> <p>On industrial production lines, the typology of potential defects is vast (texture, part failure, logical defects, etc.). Inspection systems must therefore be able to detect non-listed defects, i.e. not-yet-observed defects upon the development of the inspection system. To solve this problem, research and development of unsupervised AI algorithms on real-world data is required.</p> <p>Renault Group and the Universit&eacute; de technologie de Compi&egrave;gne (Roberval and Heudiasyc Laboratories) have jointly developed the <em>Automotive Visual Inspection Dataset (AutoVI)</em>, the purpose of which is to be used as a scientific benchmark to compare and develop advanced unsupervised anomaly detection algorithms under real production conditions. The images were acquired on Renault Group's automotive production lines, in a genuine industrial production line environment, with variations in brightness and lighting on constantly moving components. This dataset is representative of actual data acquisition conditions on automotive production lines.</p> <p>The dataset contains 3950 images, split into 1530 training images and 2420 testing images.</p> <p>The evaluation code can be found at&nbsp;<a href="https://github.com/phcarval/autovi_evaluation_code">https://github.com/phcarval/autovi_evaluation_code</a>.</p> <p><strong>Disclaimer</strong><br>All defects shown were intentionally created on Renault Group's production lines for the purpose of producing this dataset. The images were examined and labeled by Renault Group experts, and all defects were corrected after shooting.</p> <p><strong>License</strong><br>Copyright &copy; 2023-2024 Renault Group</p> <p>This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License. To view a copy of the license, visit <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>.</p> <p>For using the data in a way that falls under the commercial use clause of the license, please contact us.</p> <p><strong>Attribution</strong><br>Please use the following for citing the dataset in scientific work:</p> <p>Carvalho, P., Lafou, M., Durupt, A., Leblanc, A., &amp; Grandvalet, Y. (2024). The Automotive Visual Inspection Dataset (AutoVI): A Genuine Industrial Production Dataset for Unsupervised Anomaly Detection [Dataset]. <a href="https://doi.org/10.5281/zenodo.10459003">https://doi.org/10.5281/zenodo.10459003</a></p> <p><strong>Contact</strong><br>If you have any questions or remarks about this dataset, please contact us at philippe.carvalho@utc.fr, meriem.lafou@renault.com, alexandre.durupt@utc.fr, antoine.leblanc@renault.com, yves.grandvalet@utc.fr.</p> <p><strong>Changelog</strong></p> <ul> <li><em>v1.0.0</em> <ul> <li>Cropped engine_wiring, pipe_clip and pipe_staple images</li> <li>Reduced tank_screw, underbody_pipes and underbody_screw image sizes</li> </ul> </li> <li><em>v0.1.1</em> <ul> <li>Added ground truth segmentation maps</li> <li>Fixed categorization of some images</li> <li>Added new defect categories</li> <li>Removed tube_fastening and kitting_cart</li> <li>Removed duplicates in pipe_clip</li> </ul> </li> </ul>

opencc-by-nc-sa-4.0Feb 2024View details →
zenodo44/100

Probiotics reshape the coral microbiome in situ without detectable off-targeted effects in the surrounding environment.

<p>The R code scripts and Supplementary data files from the paper: "Probiotics reshape the coral microbiome in situ without detectable off-targeted effects in the surrounding environment," accepted in Communications Biology. All R code and data necessary to reproduce the published results are available.&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Source Data for Manuscript "Sodium salicylate improves detection of amplitude-modulated sound in mice"

<p>This repository contains the source data for our papers <strong>Sodium salicylate improves detection of amplitude-modulated sound in mice </strong>(van den Berg*, Wong*, Houtak, Williamson, Borst). The code to generate figure panels can be found in our github repository at https://github.com/aaronbwong/salicylateonam</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Lipidomics LC-MS analysis support tools for outlier detection

<p>Identification of features with high levels of confidence in liquid chromatography-mass spectrometry (LC MS) lipidomics research is an essential part of biomarker discovery, but existing software platforms can give inconsistent results, even from identical spectral data. This poses a clear challenge for reproducibility in bioinformatics work, and highlights the importance of data-driven outlier detection in assessing spectral outputs &ndash; here demonstrated using a machine learning approach based on support vector machine regression combined with leave-one-out cross validation &ndash; as well as manual curation, in order to identify software-driven errors driven by closely related lipids and by co-elution issues.</p> <p>The lipidomics case study dataset used in this work analysed a lipid extraction of a human pancreatic adenocarcinoma cell line (PANC-1, Merck, UK, cat no. 87092802) analysed using an Acquity M-Class UPLC system (Waters, UK) coupled to a ZenoToF 7600 mass spectrometer (Sciex, UK). Raw output files are included alongside processed data using MS DIAL (v4.9.221218) and Lipostar (v2.1.4) and a Jupyter notebook with Python code to analyse the outputs for outlier detection.</p>

opencc-by-sa-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record