Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

358

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

358 results for “dataset generation”

Learn how ShareScore rates datasets ↗
zenodo32/100

HoVer-NeXt: A Fast Nuclei Segmentation and Classification Pipeline for Next Generation Histopathology - Datasets

<p>This repository contains training and validation data for</p> <p><strong>HoVer-NeXt: A Fast Nuclei Segmentation and Classification Pipeline for Next Generation Histopathology&nbsp;</strong></p> <p><strong>Accepted for Oral Presentation at MIDL2024: <a href="https://openreview.net/pdf?id=3vmB43oqIO">https://openreview.net/pdf?id=3vmB43oqIO</a>&nbsp;<br></strong></p> <p><strong>More information and code are available at <a href="https://github.com/digitalpathologybern/hover_next_inference" target="_blank" rel="noopener">https://github.com/digitalpathologybern/hover_next_inference</a></strong></p> <p>Modified Lizard dataset to include mitosis (lizard_mitosis.zip), mitosis dataset (mitosis_ds.zip) and a holdout eosinophil validation set (eos_eval.zip)</p> <p>mitosis_ds.zip also contains the hold-out H&amp;E mitosis test set.</p> <p>The original lizard dataset was createdy by Simon Graham et al. and was shared under CC BY-NC-SA 4.0. The tile-based dataset can be downloaded from <a href="https://conic-challenge.grand-challenge.org/Data/">https://conic-challenge.grand-challenge.org/Data/</a> after registering for the challenge. We modify the dataset by including an additional mitosis class, however note that there are a number of mitosis which are still not (correctly annotated).</p>

opencc-by-nc-sa-4.0Feb 2024View details →
zenodo32/100

Image Datasets for "Virtual tissue microstructure reconstruction across species using generative deep learning"

<p>Training and velautaion image datasets used in the manuscript &nbsp;"Virtual tissue microstructure reconstruction across species using generative deep learning"</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

HLSFactory Pre-Generated Datasets

<p>This archive contains pre-generated datasets used for various case studies in the HLSFactory paper. We offer these pre-generated datasets so that users do not have to generate them from scratch (which can take more than 12 hours for large datasets and require FPGA vendor tools). However, these datasets can be reproduced using the HLSFactory framework if desired.</p>

opencc-by-sa-4.0Jul 2024View details →
zenodo32/100

Datasets for "Exploring the potential of neural machine translation for cross-language clinical NLP resource generation through annotation projection"

<p>This repository contains the data and additional resources used for the paper:</p> <p>"Exploring the Potential of Neural Machine Translation for Cross-Language Clinical NLP Resource Generation through Annotation Projection. Rodriguez Miret et al. Information (2024)".</p> <p>There are four different datasets included, namely:</p> <ul> <li>The (1)&nbsp; <strong><a href="https://temu.bsc.es/distemist" target="_blank" rel="noopener">DisTEMIST</a>, </strong>(2)&nbsp;<strong> <a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">DrugTEMIST</a> </strong>and<strong>&nbsp;</strong>(3)&nbsp;<strong><a href="https://temu.bsc.es/meddoprof" target="_blank" rel="noopener">MEDDOPROF</a> Spanish corpora and corresponding versions in 10 different languages</strong>, created through Machine Translation and annotation projection techniques. The Catalan annotations, used in the paper's experiments, were validated by bilingual expert annotators, who also provided alternative translations for the annotated terms in case they were wrongly translated. Thus, for Catalan we provide two different versions of the data: (i) the output of the annotation projection process as is, without any further validation, and (ii) the validated version of the data. For the rest of the languages (with the exception of the DrugTEMIST English and Italian data, used for the MultiCardioNER shared task), only an unvalidated version is provided.<strong><br></strong></li> <li>The (4) <strong>Catalan Clinical Case Corpus (CataCCC)</strong>, a collection of <em>200 clinical case reports in originally written in Catalan </em>covering a variety of clinical specialties. This corpus includes manually validated annotations for diseases, medications and professions created by the experts who annotated the corpora mentioned above, using the same guidelines and annottaion criteria. It can this be considered the first clinical Gold Standard corpus for diseases, medications and processions in Catalan.</li> </ul> <p>It is noteworthy that the MEDDOPROF-related data includes annotations for two labels, PROFESION and SITUACION_LABORAL, but only the former was used for training and evaluation in the paper.</p> <p>These are the <strong>10 languages</strong> included in the repository (along with their language codes):</p> <ul> <li><strong>Spanish</strong> (`es-gs`, with `gs` standing for Gold Standard)</li> <li><strong>Catalan</strong> (`cat`)</li> <li><strong>English</strong> (`en`)</li> <li><strong>French</strong> (`fr`)</li> <li><strong>Italian</strong> (`it`)</li> <li><strong>Dutch</strong> (`nl`)</li> <li><strong>Portuguese</strong> (`pt`)</li> <li><strong>Romanian</strong> (`ro`)</li> <li><strong>Swedish</strong> (`sv`)</li> <li><strong>Czech</strong> (`cz`)</li> </ul> <h2><strong>Related Links</strong></h2> <ul> <li><a href="../doi/10.5281/zenodo.13151039" target="_blank" rel="noopener">Validation and Correction Guidelines for the Multilingual Annotation Projection of Gold Standard Corpora</a></li> <li><a href="https://temu.bsc.es/multicardioner/" target="_blank" rel="noopener">MultiCardioNER Shared Task</a>&nbsp;</li> </ul> <h2><strong>License</strong></h2> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <h2><strong>Contact</strong></h2> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-L&oacute;pez (&lt;salvador [dot] limalopez [at] gmail [dot] com&gt;)<br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p>

opencc-by-4.0Jul 2024View details →
zenodo32/100

Automated rationale generation: a technique for explainable AI and its effects on human perceptions (Dataset)

<p>Explainable AI Dataset for paper published in the proceedings of IUI 2019 titled, &quot;<strong>Automated rationale generation: a technique for explainable AI and its effects on human perceptions</strong>&quot;. Consists of data collected from human participants via Mechanical Turk.<br> <strong>Abstract</strong>:<br> <em>Automated rationale generation</em> is an approach for real-time explanation generation whereby a computational model learns to translate an autonomous agent&#39;s internal state and action data representations into natural language. Training on human explanation data can enable agents to learn to generate human-like explanations for their behavior. In this paper, using the context of an agent that plays <em>Frogger</em>, we describe (a) how to collect a corpus of explanations, (b) how to train a neural rationale generator to produce different styles of rationales, and (c) how people perceive these rationales. We conducted two user studies. The first study establishes the plausibility of each type of generated rationale and situates their user perceptions along the dimensions of <em>confidence, humanlike-ness, adequate justification, and understandability.</em> The second study further explores user preferences between the generated rationales with regard to <em>confidence</em> in the autonomous agent, communicating <em>failure and unexpected behavior.</em> Overall, we find alignment between the intended differences in features of the generated rationales and the perceived differences by users. Moreover, context permitting, participants preferred detailed rationales to form a stable mental model of the agent&#39;s behavior.</p>

opencc-by-4.0Mar 2019View details →
zenodo32/100

Dataset of the Paper "Copilot-in-the-Loop: Fixing Code Smells in Copilot-Generated Python Code using Copilot"

<p>This dataset contains a list of 102 code smells detected from Copilot-generated Python code, along with Python code files generated by Copilot from the&nbsp;<em>Repositories</em>&nbsp;and&nbsp;<em>Code</em>&nbsp;label, respectively. This dataset also includes Copilot Chat&rsquo;s responses to fixing the 102 detected code smells. A brief description of each document and folder in the dataset is provided below:</p> <p><strong>1. files folder</strong></p> <p>contains 311 Python code files generated by Copilot. In the 311 Python files, 171 are retrieved under the&nbsp;<em>Repositories</em>&nbsp;label, indicating Python code files entirely generated by Copilot, and 140 are retrieved under the&nbsp;<em>Code</em>&nbsp;label, indicating Python code snippets generated by Copilot.</p> <p><strong>2. results of RQ1.xlsx</strong></p> <p>contains a list of 102 code smells detected from Copilot-generated Python code.</p> <p><strong>3. results of RQ2.xlsx</strong></p> <p>contains Copilot Chat&rsquo;s responses to fixing the detected 102 code smells instructed by three prompts of varying detail levels.</p>

opencc-by-4.0Jun 2024View details →
zenodo32/100

The Nucleosynthetic Yields of Core-collapse Supernovae: Prospects for the Next Generation of Gamma-Ray Astronomy Dataset

<p>Models used in "The Nucleosynthetic Yields of Core-collapse Supernovae: Prospects for the Next Generation of Gamma-Ray Astronomy"</p>

opencc-by-4.0Feb 2020View details →
zenodo32/100

Synthetic Datasets for "Binary Classification Optimisation with AI-Generated Data"

<p>Images of melanomas and Basal Cell Carcinoma generated with a stylegan2. Dataset corresponding to the article "Binary Classification Optimisation with&nbsp;AI-Generated Data"</p>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Datasets and code used to generate the figures in the article "Influence of Forest Cover Loss on Land Surface Temperature Differs by Drivers in China"

<p>We have provided the data and code used to generate the figures in the article "Influence of Forest Cover Loss on Land Surface Temperature Differs by Drivers in China" for reference and further reading. These data can be used to replicate the analyses presented in the paper. If you wish to use the data for other purposes, please contact the authors for permission. Thank you.</p>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Dataset for generation of LOD4 models for buildings towards the automated 3D modeling of BIMs and digital twins

<div> <div>This repository contains the dataset used for the automated image-based generation of LOD4 models for buildings, along with the corresponding results. The methodology utilizing this dataset was presented in the paper "Generation of LOD4 models for buildings towards the automated 3D modeling of BIMs and digital twins" by Pantoja-Rosero et., al. (2024) (https://doi.org/10.1016/j.autcon.2024.105822).</div> </div>

opencc-by-4.0Sep 2024View details →
zenodo32/100

Datasets and scripts related to the paper: "*Can Generative AI Help us in Open Coding of Software Engineering Data?*"

<p>This replication package contains datasets and scripts related to the paper: "<em>Can Generative AI Help us in Open Coding of Software Engineering Data?</em>"</p> <p>The replication package is organized into two directories:</p> <ul> <li> <p><code>manual_analysis</code>: This directory contains all sheets used to perform the manual analysis for RQ1, RQ2, and RQ3.</p> </li> <li> <p><code>stats</code>: This directory contains all datasets, scripts, and results metrics used for the quantitative analyses of RQ1 and RQ2.</p> </li> </ul> <p>In the following, we describe the content of each directory:</p> <h2>manual_analysis</h2> <ul> <li> <p><code>manual_analysis_rq1</code>: This directory contains all sheets used to perform manual analysis for RQ1 (independent and incremental coding).</p> <ul> <li> <p>The sub-directory <code>incremental_coding</code> contains .csv files for all datasets (<code>DL_Faults_COMMIT_incremental.csv</code>, <code>DL_Faults_ISSUE_incremental.csv</code>, <code>DL_Fault_SO_incremental.csv</code>, <code>DRL_Challenges_incremental.csv</code> and <code>Functional_incremental.csv</code>). All these .csv files contain the following columns:</p> <ul> <li><em>Link</em>: The link to the instances</li> <li><em>Prompt</em>: Prompt used as input to GPT-4-Turbo</li> <li><em>ID</em>: Instance ID</li> <li><em>FinalTag</em>: Tag assigned by the human in the original paper</li> <li><em>Chatgpt_output_memory</em>: Output of GPT-4-Turbo with incremental coding</li> <li><em>Chatgpt_output_memory_clean</em>: (only for the DL Faults datasets) output of GPT-4-Turbo considering only the label assigned, excluding the text</li> <li><em>Author1</em>: Label assigned by the first author</li> <li><em>Author2</em>: Label assigned by the second author</li> <li><em>FinalOutput</em>: Label assigned after the resolution of the conflicts</li> </ul> </li> <li> <p>The sub-directory <code>independent_coding</code> contains .csv files for all datasets (<code>DL_Faults_COMMIT_independent.csv</code>, <code>DL_Faults_ISSUE_ independent.csv</code>, <code>DL_Fault_SO_ independent.csv</code>, <code>DRL_Challenges_ independent.csv</code> and <code>Functional_ independent.csv</code>), containing the following columns:</p> <ul> <li><em>Link</em>: The link to the instances</li> <li><em>Prompt</em>: Prompt used as input to GPT-4-Turbo</li> <li><em>ID</em>: Specific ID for the instance</li> <li><em>FinalTag</em>: Tag assigned by the human in the original paper</li> <li><em>Chatgpt_output</em>: Output of GPT-4-Turbo with independent coding</li> <li><em>Chatgpt_output_clean</em>: (only for DL Faults datasets) output of GPT-4-Turbo considering only the label assigned, excluding the text</li> <li><em>Author1</em>: Label assigned by the first author</li> <li><em>Author2</em>: Label assigned by the second author</li> <li><em>FinalOutput</em>: Label assigned after the resolution of the conflicts.</li> </ul> </li> <li> <p>Also, the sub-directory contains sheets with inconsistencies after resolving conflicts. The directory <code>inconsistency_incremental_coding</code> contains .csv files with the following columns:</p> <ul> <li><em>Dataset</em>: The dataset considered</li> <li><em>Human</em>: The label assigned by the human in the original paper</li> <li><em>Machine</em>: The label assigned by GPT-4-Turbo</li> <li><em>Classification</em>: The final label assigned by the authors after resolving the conflicts. Multiple classifications for a single instance are separated by a comma &ldquo;,&rdquo;</li> <li><em>Final</em>: final label assigned after the resolution of the incompatibilities</li> </ul> </li> <li> <p>Similarly, the sub-directory <code>inconsistency_independent_coding</code> contains a .csv file with the same columns as before, but this is for the case of independent coding.</p> </li> </ul> </li> <li> <p><code>manual_analysis_rq2</code>: This directory contains .csv files for all datasets (<code>DL_Faults_redundant_tag.csv</code>, <code>DRL_Challenges_redundant_tag.csv</code>, <code>Functional_redundant_tag.csv</code>) to perform manual analysis for RQ2.</p> <ul> <li> <p>The <code>DL_Faults_redundant_tag.csv</code> file contains the following columns:</p> <ul> <li><em>Tags Redundant</em>: tags identified as redundant by GPT-4-Turbo</li> <li><em>Matched</em>: inspection by the authors to see if the tags are redundant matching or not</li> <li><em>FinalTag</em>: final tag assigned by the authors after the resolution of the conflict</li> </ul> </li> <li> <p>The <code>Functional_redundant_tag.csv</code> file contains the same columns as before</p> </li> <li> <p>The <code>DRL_Challenges_redundant_tag.csv</code> file is organized as follows:</p> <ul> <li><em>Tags Suggested</em>: The final tag suggested by GPT-4-Turbo</li> <li><em>Tags Redundant</em>: tags identified as redundant by GPT-4-Turbo</li> <li><em>Matched</em>: inspection by the authors to see if the tags redundant matching or not with the tags suggested</li> <li><em>FinalTag</em>: final tag assigned by the authors after the resolution of the conflict</li> </ul> </li> <li> <p>The sub-directory <code>code_consolidation_mapping_overview</code> contains .csv files (<code>DL_Faults_rq2_overview.csv</code>, <code>DRL_Challenges_rq2_overview.csv</code>, <code>Functional_rq2_overview.csv</code>) organized as follows:</p> <ul> <li><em>Initial_Tags</em>: list of the unique initial tags assigned by GPT-4-Turbo for each dataset</li> <li><em>Mapped_tags</em>: list of tags mapped by GPT-4-Turbo</li> <li><em>Unmatched_tags</em>: list of unmatched tags by GPT-4-Turbo</li> <li><em>Aggregating_tags</em>: list of consolidated tags</li> <li><em>Final_tags</em>: list of final tags after the consolidation task</li> </ul> </li> </ul> </li> <li> <p><code>prompt_for_each_rq</code>: This directory contains: - (i) the history of prompts used in each dataset (<code>prompts_history.txt</code>) -(ii) all final prompt used for the analysis of each dataset, prompt used for incremental coding, prompt used in rq2 to consolidate redundant codes, prompt used in rq3 to create taxonomy (<code>generic_prompt.txt</code>) -(iii) all .csv files in which there are indicate, for each dataset, the link and the prompt used (<code>prompt_DL_Faults_COMMIT.csv</code>, <code>prompt_DL_Faults_ISSUE.csv</code>, <code>prompt_DL_Faults_SO.csv</code>, <code>prompt_DRL_Challenges.csv</code>). For the Functional Dataset .csv file contains, instead, Question, Answer and Prompt used (<code>prompt_Functional.csv</code>)</p> </li> <li> <p><code>rq3</code>: This directory contains the taxonomies obtained from GPT-4-Turbo for the DL Faults and for the DRL Challenges (<code>taxonomy_DL_Faults.txt</code>,<code>taxonomy_DRL_Challenges.txt</code>)</p> </li> </ul> <h2>stats</h2> <ul> <li> <p><code>RQ1</code>: contains script and datasets used to perform metrics for RQ1. The analysis calculates all possible combinations between Matched, More Abstract, More Specific, and Unmatched.</p> <ul> <li><code>RQ1_Stats.ipynb</code> is a Python Jupyter nooteook to compute the RQ1 metrics. To use it, as explained in the notebook, it is necessary to change the values of variables contained in the first code block.</li> <li><code>independent-prompting</code>: Contains the datasets related to the independent prompting. Each line contains the following fields: <ul> <li><em>Link</em>: Link to the artifact being tagged</li> <li><em>Prompt</em>: Prompt sent to GPT-4-Turbo</li> <li><em>FinalTag</em>: Artifact coding from the replicated study</li> <li><em>chatgpt_output_text</em>: GPT-4-Turbo output</li> <li><em>chatgpt_output</em>: Codes parsed from the GPT-4-Turbo output</li> <li><em>Author1</em>: Annotator 1 evaluation of the coding</li> <li><em>Author2</em>: Annotator 2 evaluation of the coding</li> <li><em>FinalOutput</em>: Consolidated evaluation</li> </ul> </li> <li><code>incremental-prompting</code>: Contains the datasets related to the incremental prompting (same format as independent prompting)</li> <li><code>results</code>: contains files for the RQ1 quantitative results. The files are named <code>RQ1\_&lt;&lt;Dataset&gt;&gt;\_&lt;&lt;Prompt method&gt;&gt;\_&lt;&lt;ExcludingNegative&gt;&gt;\_&lt;&lt;MetricAggregation&gt;&gt;.csv</code>, where <em>Dataset</em> is the dataset name, <em>Prompt method</em> indicates whether results are for independent or incremental prompting, <em>Excluding Negatives</em> (for datasets where this applies) whether results have been obtained by excluding negative instances, and <em>MetricAggregation</em> (where it applies) how metrics have been aggregated (macro or weighted average). The files report columns indicating the <em>Dataset</em>, the <em>Matching type</em>, the <em>Accuracy</em>, <em>Precision</em>, <em>Recall</em>, <em>F1 Score</em>, and <em>Cohen's Kappa</em>.</li> </ul> </li> <li> <p><code>RQ2</code>: contains the script used to perform metrics for RQ2, the datasets it uses, and its output.</p> <ul> <li><code>RQ2_SetStats.ipynb</code> is the Python Jupyter notebook to perform the analyses. The scripts takes as input the following types of files, contained in the directory contains the script used to perform the metrics for RQ2. The script takes in input:</li> <li>RQ1 Data Files (<code>RQ1_DLFaults_Issues.csv</code>, <code>RQ1_DLFaults_Commits.csv</code>, and <code>RQ1_DLFaults_SO.csv</code>, joined in a single .csv <code>RQ1_DLFaults.csv</code>). These are the same files used in RQ1.</li> <li>Mapping Files (<code>RQ2_Mappings_DRL.csv</code>, <code>RQ2_Mappings_Functional.csv</code>, <code>RQ2_Mappings_DLFaults.csv</code>). These contain the mappings between human tags (<em>HumanTags</em>), GPT-4-Turbo tags (<em>Final Tags</em>), with indicated the type of matching (<em>MatchType</em>).</li> <li>Additional codes creating during the consolidation (<code>RQ2_newCodes_DRL.csv</code>, <code>RQ2_newCodes_Functional.csv</code>, <code>RQ2_newCodes_DLFaults.csv</code>), annotated with the matching: <em>new code</em>,<em>old code</em>,<em>human code</em>,<em>match type</em></li> <li>Set files (<code>RQ2_Sets_DRL.csv</code>, <code>RQ2_Sets_Functional.csv</code>, <code>RQ2_Sets_DLFaults.csv</code>). Each file contains the following columns: <ul> <li><em>HumanTags</em>: List of tags from the original dataset</li> <li><em>InitialTags</em>: Set of tags from RQ1,</li> <li><em>ConsolidatedTags</em>: Tags that have been consolidated,</li> <li><em>FinalTags</em>: Final set of tags (results of RQ2, used in RQ3)</li> <li><em>NewTags</em>: New tags created during consolidation</li> </ul> </li> <li><code>RQ2_Set_Metrics.csv</code>: Reports the RQ2 output metrics (Precision, Recall, F1-Score, Jaccard).</li> </ul> </li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo32/100

Datasets for reproducing the results in "True random number generators with flicker noise: stochastic model, min-entropy calculation and online test"

<p>Datasets for reproducing the results in "True random number generators with flicker noise: stochastic model, min-entropy calculation and online test"</p> <p>Includes scripts for generating raw results, postprocessing scripts, as well as scripts/notebooks for generating plots for the manuscript</p>

opencc-by-4.0Feb 2024View details →
zenodo32/100

Zebrafish capable of generating future state prediction error show improved active avoidance behavior in virtual reality [Dataset]

<p>The calcium imaging data of the telencephalon of head-tethered adult zebrafish during GO/NOGO tasks in the virtual reality environment and the behavior data were deposited.</p> <p>The codes to process the neural activity data by calcium imaging to&nbsp;perform Non-negative Matrix Factorization&nbsp;</p> <p>For details, see &quot;Zebrafish capable of generating future state prediction error show improved active avoidance behavior in virtual reality&quot; Torigoe et al., Nature Communications in press.</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Dataset of Waves generated by discrete and sustained gas eruptions with implications for submarine volcanic tsunamis

<p>This dataset contains all the data used to plot figures in the manuscript &quot;Waves generated by discrete and sustained gas eruptions with implications for submarine volcanic tsunamis&quot; by Yaxiong Shen, Colin Whittaker, Emily Lane, James White, William Power and Bruce Melville&nbsp;submitted to Geophysical Research Letters.</p> <p>The folder&nbsp;&quot;jetplumefountainquantification&quot; contains data for figure 1.</p> <p>The wavedata.xlsx contains wave data and plume rise velocity data.&nbsp;</p> <p>The MATLAB script fig1.m is used to visualise the jet-plume-fountain evolution.</p>

opencc-by-4.0Sep 2021View details →
zenodo32/100

Next Generation Spacecraft Pose Estimation Dataset (SPEED+)

<p>SPEED+ is the next-generation dataset for spacecraft pose estimation with specific emphasis on the robustness of Machine Learning (ML) models across the domain gap. Similar to its predecessor, SPEED+ consists of images of the Tango spacecraft from the PRISMA mission. SPEED+ consists of three different domains of imageries from two distinct sources. The first source is the <strong>OpenGL-based Optical Stimulator camera emulator</strong> software of <a href="https://damicos.people.stanford.edu">Stanford&rsquo;s Space Rendezvous Laboratory (SLAB)</a>, which is used to create the synthetic domain comprising 59,960 synthetic images. The labeled synthetic domain is split into 80:20 train/validation sets and is intended to be the main source of training of an ML model.</p> <p>The second source is the&nbsp;<strong>Testbed for Rendezvous and Optical Navigation (TRON)&nbsp;</strong>facility at SLAB, which is used to generate two simulated Hardware-In-the-Loop (HIL) domains with different sourcesof illumination: lightbox and sunlamp. Specifically, these two domains are constructed using realistic illumination conditions using lightboxes with diffuser plates for albedo simulation and a sun lamp to mimic direct high-intensity homogeneous light from the Sun.</p> <p>Compared to synthetic imagery, they capture corner cases, stray lights, shadowing, and visual effects in general which are not easy to obtain through computer graphics. The lightbox and sunlamp domains are <strong>unlabeled</strong> and thus intendeded mainly for testing, representing a typical scenario in developing a spaceborne ML model in which the labeled images from the target space domain are not available prior to deployment.</p> <p>SPEED+ is made publicly available to the&nbsp;aerospace community and beyond as part of the <strong>second international Satellite Pose Estimation Competition (SPEC2021)</strong> co-hosted by SLAB and the <a href="https://www.esa.int/gsp/ACT/">Advanced Concepts Team (ACT)</a> of the <a href="https://esa.int">European Space Agency</a>.</p> <p>The construction of the TRON testbed was partly funded by the U.S. Air Force Office of Scientific Research (AFOSR) through the Defense University Research InstrumentationProgram (DURIP) contract FA9550-18-1-0492, titled High-Fidelity Verification and Validation of Spaceborne Vision-Based Navigation. The SPEED+ dataset was&nbsp;created using the TRON testbed by SLAB at Stanford University. The post-processing of the raw images was&nbsp;reviewed by ACT to meet the quality requirement of SPEC2021.</p> <p>For more details on the dataset and the competition, please visit&nbsp;<a href="https://kelvins.esa.int/pose-estimation-2021/">https://kelvins.esa.int/pose-estimation-2021/</a></p> <p>&nbsp;</p>

opencc-by-4.0Oct 2021View details →
zenodo32/100

Example datasets generated with Stytra

<p>Example datasets accompanying Stytra with:</p> <p><strong>example_closed_loop_embedded.zip</strong></p> <p>Behavioral data from a closed loop oculomotor response paradigm. The fish was swimming in response to backward-moving gratings controlled in closed loop. The code for the protocol is <a href="https://github.com/portugueslab/stytra/blob/master/stytra/examples/closed_loop_exp.py">here</a>.</p> <p><strong>example_imaging.zip</strong></p> <p>Behavioral and calcium imaging data from a closed loop oculomotor response paradigm in a transgenic Huc:GCaMP6f fish.<br> The fish was swimming in response to backward-moving gratings controlled in closed loop, while a plane of the fish brain was scanned with a two-photon microscope. The code for the protocol is <a href="https://github.com/portugueslab/stytra/blob/master/stytra/examples/imaging_exp.py">here</a>.</p> <p><strong>replication_portugues2011.zip</strong></p> <p>Dataset acquired with the experimental protocol described in <a href="https://www.frontiersin.org/articles/10.3389/fnsys.2011.00072/full">Portugues and Engert, 2011</a>, reproduced within Stytra. The code for the protocol is <a href="https://github.com/portugueslab/stytra/blob/master/stytra/examples/portugues2011_exp.py">here</a>.</p> <p><strong>example_trace_cheap_setup.zip</strong></p> <p>Behavioral data from a closed loop oculomotor response paradigm in the low-cost behavioral rig described in the article. The code for the protocol is <a href="https://github.com/portugueslab/stytra/blob/master/stytra/examples/portugues2011_exp.py">here</a>.</p> <p><strong>example_eye_motion.zip</strong></p> <p>eye motion in response to sinusoidal rotation of black and white windmill stimuli. The code for the protocol is <a href="https://github.com/portugueslab/stytra/blob/master/stytra/examples/eye_tracking_exp.py">here</a>.</p> <p><strong>phototaxis.zip</strong></p> <p>Replication of the <a href="https://dx.doi.org/10.1016%2Fj.cub.2013.06.044">Huang et. al Current Biology. 2013 </a>study. For 10 s, the right side of the visual field is illuminated, while the left side is kept dark. Then, for 10 s the full field is bright, which repeats 60-80 times. The code for the protocol is <a href="https://github.com/portugueslab/stytra/blob/master/stytra/examples/phototaxis.py">here.</a></p> <p>All datasets were acquired with 6-8dpf larvae of Danio Rerio.</p> <p>Notebooks demonstrating analysis of this data are avilable at <a href="https://github.com/portugueslab/example_stytra_analysis">github</a></p>

opencc-by-4.0Nov 2018View details →
zenodo32/100

Dataset generated in support of nitric oxide assimilation by Heterosigma akashiwo when acclimated to growth on nitrate, ammonium, or a mix of nitrate and ammonium

<p>This dataset includes primary data generated during our investigation of the effects of competing nitrogen sources on the assimilation of nitric oxide (NO) into algal biomass by the harmful algal species, <em>Heterosigma akashiwo</em>. <em>H. akashiwo </em>was grown in continuous culture on artificial seawater and provided with either nitrate, ammonium or nitrate plus ammonium as a sole source of nitrogen. Cultures were spiked with NO (100 &micro;M), or NO<sub>3</sub><sup>-</sup>, or NH<sub>4</sub><sup>+ </sup>at concentrations equal to 10% ambient levels. Controls were spiked with an equal volume of water. Samples were collected for RT-qPCR analysis of gene expression at 15 and 60 minutes after spiking. Additional samples were collected for analysis of NR enzyme activity at 2 hours after spiking, and again at 4 hours and 24 hours to evaluate assimilation of NO into biomass.</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Automated and Reproducible Application Traces Generation for IoT Applications Dataset Lighting Application

<p>This data represents&nbsp;an IoT smart city application. It results&nbsp;from an experiment that&nbsp;runs a firmware on a set of representative nodes that have to exchange packets in a broadcast mode using the IEEE 802.15.4-2006 MAC layer and RPL routing protocol. Each application produces data according to 1 of the 3 following modes: periodic (Tx nodes produce data every x milliseconds), event based (modeled with an exponential law with occurrence rate lambda) and hybrid (combination of the two previous modes).</p> <p>Each application&nbsp;has the following parameters :<br> - Surveillance has 10 sensors and 3 routers that exchange packets with a length of 127B. The generation type is exponential with a lambda of 196.74.<br> - Emergency Response has 40 sensors and 5 routers that exchange packets with a length of 127B. The generation type is hybrid with a lambda of 0.03 and a period of 30 seconds.<br> - HVAC has 100 sensors and 5 routers that exchange packets with a length of 60B. The generation type is periodic with a period of 260 seconds.<br> - Lighting has 100 sensors and 5 routers that exchange packets with a length of 30B. The generation type is exponential with a lambda of 0.00208.<br> - VoIP has 10 sensors and 1 router that exchange packets with a length of 127B. The generation type is hybrid with a lambda of 15.74 and a period of 0.063532 seconds.</p> <p>As a result, this dataset has files containing the following data :<br> - Received data : name of the receiving node (node_name); message reception time (timestamp); message unique identifier (message_id); reception delay in milliseconds (reception_delay)<br> - Transmitted data : name of the transmitting node (node_name); message transmission time (timestamp); message unique identifier (message_id); success (transmission success)</p> <p>Furthermore, datasets of 5&nbsp; more IoT Applications are&nbsp; available at the following <a href="https://zenodo.org/record/7347970">link</a>&nbsp;<br> <br> For&nbsp;any&nbsp;questions,&nbsp;please&nbsp;contact Nina&nbsp;Santi&nbsp;(<a href="mailto:nina.santi@inria.fr">nina.santi@inria.fr</a>)</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Automated and Reproducible Application Traces Generation for IoT Applications Dataset

<p>This data represents&nbsp;an IoT smart city application. It results&nbsp;from an experiment that&nbsp;runs a firmware on a set of representative nodes that have to exchange packets in a broadcast mode using the IEEE 802.15.4-2006 MAC layer and RPL routing protocol. Each application produces data according to 1 of the 3 following modes: periodic (Tx nodes produce data every x millisecond), event based (modeled with an exponential law with occurrence rate lambda), and hybrid (combination of the two previous modes).</p> <p>Each application&nbsp;has the following parameters :<br> - Surveillance has 10 sensors and 3 routers that exchange packets with a length of 127B. The generation type is exponential with a lambda of 196.74.<br> - Emergency Response has 40 sensors and 5 routers that exchange packets with a length of 127B. The generation type is hybrid with a lambda of 0.03 and a period of 30 seconds.<br> - HVAC has 100 sensors and 5 routers that exchange packets with a length of 60B. The generation type is periodic with a period of 260 seconds.<br> - Lighting has 100 sensors and 5 routers that exchange packets with a length of 30B. The generation type is exponential with a lambda of 0.00208.<br> - VoIP has 10 sensors and 1 router that exchange packets with a length of 127B. The generation type is hybrid with a lambda of 15.74 and a period of 0.063532 seconds.</p> <p>As a result, this dataset has files containing the following data :<br> - Received data : name of the receiving node (node_name); message reception time (timestamp); message unique identifier (message_id); reception delay in milliseconds (reception_delay)<br> - Transmitted data : name of the transmitting node (node_name); message transmission time (timestamp); message unique identifier (message_id); success (transmission success)</p> <p>Furthermore, a dataset of an IoT Lighting Application is available at the following <a href="https://zenodo.org/record/7348232#.Y4D9VdJBxhE">link</a><br> For&nbsp;any&nbsp;questions,&nbsp;please&nbsp;contact Nina&nbsp;Santi&nbsp;(<a href="mailto:nina.santi@inria.fr">nina.santi@inria.fr</a>)</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Dataset for paper Pavel Perezhogin, Laure Zanna, Carlos Fernandez-Granda "Generative data-driven approaches for stochastic subgrid parameterizations in an idealized ocean model" submitted to JAMES.

<p>The dataset consists of the directory tree of .zarr archives. See <a href="https://github.com/m2lines/pyqg_generative/blob/master/Google-Colab/dataset.ipynb">Github repository</a>&nbsp;for the description of the dataset.</p> <p>The directory tree is:</p> <pre><code>├── eddy │ ├── 48 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 64 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 96 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ └── hires ├── jet │ ├── 48 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 64 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ ├── 96 │ │ ├── gauss │ │ ├── hires-gauss │ │ ├── hires-sharp │ │ ├── lores │ │ └── sharp │ └── hires</code></pre> <ul> <li>Every individual dataset is a&nbsp;<code>.zarr</code>&nbsp;<a href="https://zarr.readthedocs.io/en/stable/">archive</a></li> <li><code>eddy/jet</code>&nbsp;- configuration of the pyqg; eddy is default; See&nbsp;<a href="https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/2022MS003258">Ross2022</a>&nbsp;for description</li> <li><code>hires.zarr</code>&nbsp;- high-resolution simulation at 256x256 grid</li> <li><code>48/64/96</code>&nbsp;- resolution of the coarse models</li> <li><code>lores.zarr</code>&nbsp;- low-resolution simulation</li> <li><code>gauss.zarr</code>,&nbsp;<code>sharp.zarr</code>&nbsp;- training datasets for prediction of subgrid forcing obtained with Gaussian or Sharp filters</li> <li><code>hires-gauss.zarr</code>,&nbsp;<code>hires-sharp.zarr</code>&nbsp;- high-resolution simulation projected onto coarse grid with Gaussian or Sharp filters</li> </ul> <p>The directory tree is split into small tar.gz files each representing a separate .zarr archive. Download any required parts of the dataset and unpack with:</p> <p><strong>tar -xf *.tar.gz&nbsp;</strong></p> <p><strong>The directory tree will be restored automatically!</strong></p>

opencc-by-4.0Feb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record