Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

73

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

73 results for “artificial dataset”

Learn how ShareScore rates datasets ↗
zenodo36/100

Dataset for the article "Polarizable Embedding without Artificial Boundary Polarization"

<p>This dataset contains additional material related to the article: &quot;Polarizable Embedding without Artificial Boundary Polarization&quot;.&nbsp;A preprint is freely available at https://doi.org/10.26434/chemrxiv-2023-fc0gk. See readme.md for further description.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

BurstDECONV, artificial datasets

<p>Artificial dataset used in the paper&nbsp;BurstDECONV: A signal deconvolution method to uncover mechanisms of transcriptional bursting in live cells.</p> <p>BurstDECONV is an innovative inference method&nbsp;that deconvolves live cell transcription imaging data to reconstruct each polymerase initiation event.</p> <p>This&nbsp;new version includes 297 synthetic datasets simulating&nbsp;live transcription imaging data for 2 states and 3 states promoters with various switching rate parameters.</p> <p>We &nbsp;provide the code used for generating&nbsp;these datasets and for&nbsp;producing a revised version of Figure 4 of the paper.&nbsp;</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

Habitat-driven vulnerability to nest predation in Arctic-breeding plovers based on artificial nest experiment and real shorebird nest monitoring – Survival dataset

Open the record for dataset details and reuse information.

publicFeb 2020View details →
zenodo32/100

Dataset for "The State of the ML-universe: 10 Years of Artificial Intelligence & Machine Learning Software Development on GitHub"

<p>Supplementary data to &quot;The State of the ML-universe: 10 Years of Artificial Intelligence &amp; Machine Learning Software Development on GitHub&quot; accepted for publication at MSR 2020.</p> <p>The data included in this package were used to conduct analyses to characterize the AI &amp; ML software development community hosted on GitHub. Please read the paper for a full understanding of what data was collected and how it was used.</p> <p>Questions and comments can be directed to Danielle Gonzalez dng2551@rit.edu</p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Finite element dataset and Artificial Neural Networks algorithms to predict the mechanical properties of innovative CLT

<p>This folder includes the data collected from the finite element simulations of the innovative CLT to compute its mechanical properties, the error of the closed-form solutions predicting the bending stiffness in the minor direction D22, the variation of the distance between the Reissner Mindlin and Bending Gradient theory in terms of spacing between lateral lamellas, the hyperparameters tuning of several Artificial Neural Networks algorithms with or without prior knowledge, the ML evaluations, the saved artificial neural network algorithms to predict each mechanical property of innovative CLT, and the ML application to use it.</p>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Dataset for the article Artificial intelligence for earthquake prediction: a preliminary system based on periodically trained neural networks using ionospheric anomalies

<p>Training and validation data sets along with the corresponding trained convolutional neural network in the article "Artificial intelligence for earthquake prediction: a preliminary system based on periodically trained neural networks using ionospheric anomalies" by Sergio Baselga published in&nbsp;<em>Appl. Sci.</em>&nbsp;<strong>2024</strong>,&nbsp;<em>14</em>(23), 10859; https://doi.org/10.3390/app142310859</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Dataset of publications focused on artificial light at night in fishes (up to January 2021)

<p>Summary of publications focused on artificial light in fishes (as identified by search in Web of Science Core Collection; January 2021) with designation of&nbsp;whether they pertain to light pollution or other contexts (e.g., bycatch reduction, aquaculture). The types of effects studied are also designated (e.g., behaviour, physiology, community structure).</p>

opencc-by-4.0Dec 2021View details →
zenodo32/100

An Artificial Intelligence Dataset for Solar Energy Locations in India

<p>To expedite development of solar energy, land use planners will need access to up-to-date and accurate geo-spatial information of PV infrastructure. In this work, we develop a machine learning model to map utility-scale solar projects across India using freely available satellite imagery. Model predictions were validated by human experts to obtain a total of 1438 solar farms. We also estimate the solar footprint across India and quantified the degree of land modification associated with land cover types that may cause conflicts. Our analysis indicates that over 74% of solar development in India was built on landcover types that have natural ecosystem preservation, and agricultural values. Our work increases the feasibility of long-term monitoring of renewable energy deployment targets.</p>

opencc-by-4.0Jan 2022View details →
zenodo32/100

artificial dataset of small engine parts for object detection-segmentation

<p>This dataset contains images and masks of&nbsp;parts used in engine assembly. The images were generated artificially from CAD models using gazebo simulator. The dataset consists of three classes: Large Bolt, Small Bolt and Rocker Arm. Annotations are represented as mask images with the same name as corresponding RGB images. 1080 images for each class, 3240 images total. Resolution of each image: 640x480 pixels. CAD models for each part are also included.</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Supporting Datasets and Notebook for the manuscript: "Humans program artificial delegates to accurately solve collective-risk dilemmas but lack precision"

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo32/100

Detailed dataset and code generation for Artificial intelligence-based modelling of compressive strength of slurry infiltrated fiber concrete

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo32/100

Dataset of images SfM - FRM - Lighting and Artificial texture

<p>This collection of images was employed to examine the impact of various configurations in the 3D modeling process for short-distance environments. This analysis established settings to achieve submillimeter accuracy in the RMSE values of the analyzed points.&nbsp;</p> <p>Set of images used to evaluate the use of light aids (softboxes) and artificial textures in a short-distance environment. This set was used to evaluate the configurations and the possibility of using the SfM technique in the 3D modeling of structural tests. Test specimens used (Concrete, Metal and Wood)&mdash;artificial texture in white Chalk (on Concrete and Wood) and red marker (on metal).<br>Texture patterns were drawn in a checkerboard fashion (T1) and a more closed checkerboard shape (T2).</p> <p>Images in CR3 - Conversion to TIFF or JPG required for use and processing in Agisoft Metashape.</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Artificial datasets for "Online Conformance Checking Using Behavioural Patterns"

<p>Dataset containing artificial datasets for the stress test of the online conformance prototype and for the correlation of the results of the online conformance checker with state of the art technique.</p> <p><strong>Stress test log</strong></p> <p>We randomly generated a BPMN model containing 64 activities and 26 gateways. The model was then used to simulate an event stream of 2 million events.</p> <p><strong>Correlation logs</strong></p> <p>We generated 12 random process models&nbsp;with number of activities according to a triangular distribution with lower bound 10, mode 20, and upper bound 30. We did not include duplicate labels, a probability of 0.2 for addition of silent activities, moreover, the probability of control-flow operator insertion was: 0.45 for sequence, 0.2 for parallel and xor-split operators, 0.05 for an inclusive-or operator and 0.1 for loop constructs. Incremental noise levels (both on a trace- and event-level) were introduced in the logs. Probability of trace- and event-level noises ranged from 0.1 to 0.5 with steps of 0.1.</p>

opencc-by-4.0Mar 2018View details →
zenodo32/100

Repository for the codes and raw dataset used in the paper: Testing driving mechanisms of megathrust seismicity with Explainable Artificial Intelligence.

<p>This repository contains the codes/notebooks and raw dataset used in the paper:</p> <p>Testing driving mechanisms of megathrust seismicity with Explainable Artificial Intelligence. by&nbsp;Juan Carlos Graciosa, Fabio A. Capitanio, Adam Beall, Mitchell Hargreaves, Thyagarajulu Gollapalli, Titus Tang, Mohd Zuhair</p> <h3>xai-megathrust:</h3> <p>This directory contains the following:</p> <p>1. helper_pkg: Package containing helper routines used during the creation of grids.<br>2. in-data: Contains the processed but non-standardized features. Standardization is done during runtime.<br>3. ml4szeq: Main set of codes used in the study.<br>4. ntbk: Notebooks used in the study. This includes the sampling of the raw data into grids (0_grid_sampling.ipynb), creation of classification maps (1_make_classification_maps.ipynb), and the creation of LRP heatmaps (2_make_lrp_heatmaps.ipynb).<br>5. vis_pkg: Package used for creating maps</p> <p>&nbsp;</p> <h3>xai-megathrust-raw-data:</h3> <p>This contains the raw dataset. Here, the data prefix indicates the convergent region it is a part of and are as follows:</p> <p>1. alu: Alaska-Aleutians<br>2. cam: Central America<br>3. izu: Izu-Bonin-Mariana<br>4. ker: Tonga-Kermadec<br>5. kur: Japan-Kuriles-Kamchatka<br>6. ryu: Ryukyu-Nankai<br>7. sam: South America<br>8. sum: Southeast Asia&nbsp;</p> <p>This was adapted from the notation used by Hayes et al., 2018.</p>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Geomagnetic datasets of BJI station reconstructed through Artificial Neural Network improved by Genetic Algorithm in 2021

<p>Beijing station established in 1954 is one of the oldest geomagnetic observatories in China, which plays an important role in data exchange, and further provide data or standardization for satellite observation and geomagnetic model construction. With the development&nbsp;of urbanization, the observed&nbsp;data are&nbsp;greatly disturbed&nbsp;by subways, and data disturbed are almost unavailable. The dataset&nbsp;was reconstructed through Artificial Neural Network improved by Genetic Algorithm, including minutely&nbsp;data&nbsp;of three components (<em>D</em>, <em>H</em>&nbsp;and <em>Z</em>) in&nbsp;2021. This reconstruction method has been proved to be effective.</p>

opencc-by-4.0Jan 2023View details →
zenodo28/100

Dataset for "Artificial neural network and SARIMA based models for power load forecasting in Turkish electricity market"

<p>This is the dataset for the manuscript "Artificial neural network and SARIMA based models for power load forecasting in Turkish electricity market" submitted to the journal PLOS ONE. </p>

opencc-by-nc-4.0Mar 2017View details →
dryad28/100

Semi-artificial datasets as a resource for validation of bioinformatics pipelines for plant virus detection

<p>In the last decade, High-Throughput Sequencing (HTS) has revolutionized biology and medicine. This technology allows the sequencing of huge amount of DNA and RNA fragments at a very low price. In medicine, HTS tests for disease diagnostics are already brought into routine practice. However, the adoption in plant health diagnostics is still limited. One of the main bottlenecks is the lack of expertise and consensus on the standardization of the data analysis. The Plant Health Bioinformatic Network (PHBN) is an Euphresco project aiming to build a community network of bioinformaticians/computational biologists working in plant health. One of the main goals of the project is to develop reference datasets that can be used for validation of bioinformatics pipelines and for standardization purposes.</p> <p>Semi-artificial datasets have been created for this purpose (Datasets 1 to 10). They are composed of a "real" HTS dataset spiked with artificial viral reads. It will allow researchers to adjust their pipeline/parameters as good as possible to approximate the actual viral composition of the semi-artificial datasets. Each semi-artificial dataset allows to test one or several limitations that could prevent virus detection or a correct virus identification from HTS data (<i>i.e.</i> low viral concentration, new viral species, non-complete genome).</p> <p>Eight artificial datasets only composed of viral reads (no background data) have also been created (Datasets 11 to 18). Each dataset consists of a mix of several isolates from the same viral species showing different frequencies. The viral species were selected to be as divergent as possible. These datasets can be used to test haplotype reconstruction software, the goal being to reconstruct all the isolates present in a dataset.</p> <p><span>A GitLab repository (<a href="https://gitlab.com/ilvo/VIROMOCKchallenge">https://gitlab.com/ilvo/VIROMOCKchallenge</a>) is available and provides a complete description of the composition of each dataset, the methods used to create them and their goals.</span></p>

opencc-zeroNov 2021View details →
zenodo28/100

Dataset for "Reprogramming macrophages with R848-loaded artificial protocells to modulate skin and skeletal wound healing"

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo28/100

Dataset for "A Fuel Cell Power Supply System Equipped with Artificial Gill Membranes for Underwater Applications"

<p>This data file contains the experimental data pertaining to the manuscript&nbsp;<br>"A Fuel Cell Power Supply System Equipped with Artificial Gill Membranes for Underwater Applications" by Lucas Merckelbach and Prokopios Georgopanos.</p> <p>Table of Contents</p> <p>-------------------</p> <p>experiments: data files and description/notes taken for each experiment reported in the manuscript, as well as the python script to create Figure 3 of the manuscript.</p> <p>The directory contains a README file with further information and instructions.</p>

opencc-by-4.0Jul 2024View details →
zenodo28/100

Dataset: Artificial light at night causes reproductive failure in clownfish

<p>Dataset on the reproductive output of <em>Amphiprion ocellaris</em>&nbsp;under control and ALAN conditions.</p>

opencc-by-4.0Jan 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record