Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

11,687

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

11,687 results for “training”

Learn how ShareScore rates datasets ↗
zenodo48/100

Dataset for Training Material - Galaxy Workflow - Analyse unaligned ncRNAs

<p>Input dataset for Galaxy Training Material for the Analyze unaligned ncRNAs workflow.</p> <p>See https://github.com/galaxyproject/training-material for more information.</p>

opencc-by-4.0Oct 2019View details →
zenodo48/100

LCZ-Generator Training Areas

<p>This dataset contains all training areas (TA) submitted to the <a href="https://lcz-generator.rub.de/">LCZ Generator</a> (<a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a>) since 2021-04-14. The LCZ Generator follows a crowdsourcing approach, making fast and easy LCZ-mapping available to the public, while collecting LCZ maps and TAs in a centralized, easy to access, location. The crowdsoucing approach overcomes the limitations of previous approaches where a manual review was mandatory before publication. While this improved the quality of individual LCZ-maps, the number of cities mapped during this period remained low. The LCZ Generator removed the manual review process, allowing for faster collection of LCZ maps and TAs, however, sacrificing some quality since any person can submit to the LCZ Generator without prior training or review.</p> <p>This dataset is based on crowdsourcing, hence LCZs may be mislabelled, polygon shapes may not be perfect etc. also city names may not be correct. Some contributors chose to name their city e.g. "..", "....amsa" etc. also some author names may not be correct. The automated quality control (see table below and section 2.3 in <a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a>) may help filter out some of the incorrect TAs. We intentionally included all available TAs to allow for (the development of) custom filtering.</p> <p>The data was extracted from the LCZ-Generator database taking into account:</p> <ol> <li>Whether or not the submitting author agreed to show their name (if not, it is also left blank in this dataset)</li> <li>The license the TA was submitted under (a license change happened with version 2.0.0 of the LCZ Generator)</li> <li>Duplicate geometries were dropped, since multiple (re-)submission may have the same geometries. Only the <strong>first</strong> submitted version is kept and attributed to the <code>submission_id</code> of the first submission.</li> </ol> <p>The data, up to December 2021, was used during creation of the global LCZ Map (<a href="https://doi.org/10.5194/essd-14-3835-2022">Demuzere et al. 2022</a>).</p> <p>Additional TAs were extracted from the <a href="https://wudapt.cs.purdue.edu">WUDAPT Portal</a> and processed using the LCZ-Generator.</p> <p><strong>Note</strong>: The data is updated periodically, but not on a fixed schedule.</p> <h2>Data Description</h2> <p>The data is provided as GeoPackage (<code>.gkpg</code>) which can be used with most GIS.</p> <table> <tbody> <tr> <th>Column Name</th> <th>Description</th> </tr> </tbody> <tbody> <tr> <td><code>geometry</code></td> <td>The polygon geometry of the training area (TA) in EPSG:4326</td> </tr> <tr> <td><code>submission_id</code></td> <td>The ID of the corresponding submission in the <a href="https://lcz-generator.rub.de/">LCZ Generator</a></td> </tr> <tr> <td><code>submission_date</code></td> <td>The date and time in <strong>UTC</strong> the TA was submitted to the LCZ Generator</td> </tr> <tr> <td><code>city</code></td> <td>The city the TAs are for. Note: This is sometimes incorrect due to users entering incorrect information and the LCZ Generator following a crowdsourcing approach.</td> </tr> <tr> <td><code>reference</code></td> <td>The submitting author may have provided (additional) references via this field. This can be a scientific paper or a citation of the original creator of this TA</td> </tr> <tr> <td><code>remarks</code></td> <td>General information: e.g. co-authors, information about the study/framework the TAs were generated for</td> </tr> <tr> <td><code>representative_date</code></td> <td>The date the TAs are representative for (i.e. the date the aerial image was taken)</td> </tr> <tr> <td><code>firstname</code></td> <td>First name of the submitting author (if the author did not agree to publish their name, this is left blank and the submission is treated anonymously)</td> </tr> <tr> <td><code>lastname</code></td> <td>Last name of the submitting author (if the author did not agree to publish their name, this is left blank and the submission is treated anonymously)</td> </tr> <tr> <td><code>license</code></td> <td>The license this specific polygon is licensed under. With version 2.0.0 of the LCZ Generator the license was changed from CC BY-SA to CC BY-NC-SA 4.0</td> </tr> <tr> <td><code>cite_as</code></td> <td>A suggestion how to cite the TA (-set) based on the name, year, and city information. If the author submitted anonymously, this is left blank</td> </tr> <tr> <td><code>version</code></td> <td>The version of the LCZ Generator the polygon was submitted to. Detailed information can be found in the <a href="https://github.com/RUBclim/LCZ-Generator-Issues?tab=readme-ov-file#changelog">Changelog</a></td> </tr> <tr> <td><code>class</code></td> <td>The LCZ Class the TA-polygon has been labelled (1 - 17)</td> </tr> <tr> <td><code>area</code></td> <td>The area of the TA-polygon in km<sup>2</sup></td> </tr> <tr> <td><code>perimeter</code></td> <td>The perimeter of the TA-polygon km</td> </tr> <tr> <td><code>shape</code></td> <td>The shape of the TA-polygon calculated as: (perimeter<sup>2</sup>) / (4 &pi; &middot; area)</td> </tr> <tr> <td><code>vertices</code></td> <td>The number of vertices of the TA-polygon</td> </tr> <tr> <td><code>qc_step1</code></td> <td>Whether the TA-polygon passed the automated quality control (QC) step 1: Surface area below 0.04 km<sup>2</sup> (too small) or a shape ratio 3 (too complex shape) are flagged. More information about the QC can be found in the <a href="https://lcz-generator.rub.de/faq#why-suspicious-tas">FAQ</a> and the corresponding paper <a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a></td> </tr> <tr> <td><code>qc_step2</code></td> <td>Whether the TA-polygon passed the automated quality control (QC) step 2: Average spectral value of a polygon of LCZ class is considered as an outlier compared to the average spectral values of all other polygons of that class. Note that this is done on a per-submission basis More information about the QC can be found in the <a href="https://lcz-generator.rub.de/faq#why-suspicious-tas">FAQ</a> and the corresponding paper <a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a></td> </tr> <tr> <td><code>qc_step3</code></td> <td>Whether the TA-polygon passed the automated quality control (QC) step 3: Considers all individual pixel values of all polygons in each LCZ class compared to the polygon average approach from QC Step 2. More information about the QC can be found in the <a href="https://lcz-generator.rub.de/faq#why-suspicious-tas">FAQ</a> and the corresponding paper <a href="https://doi.org/10.3389/fenvs.2021.637455">Demuzere et al. 2021</a></td> </tr> <tr> <td><code>oa</code></td> <td>The overall accuracy of the submission based on <a href="https://doi.org/10.1038/s41597-020-00605-z">Demuzere et al. 2020</a></td> </tr> <tr> <td><code>oau</code></td> <td>The overall accuracy for the urban LCZ classes only of the submission based on <a href="https://doi.org/10.1038/s41597-020-00605-z">Demuzere et al. 2020</a></td> </tr> <tr> <td><code>oabu</code></td> <td>The overall accuracy of the built versus natural LCZ classes only of the submission based on <a href="https://doi.org/10.1038/s41597-020-00605-z">Demuzere et al. 2020</a></td> </tr> <tr> <td><code>oaw</code></td> <td>A weighted accuracy taking the similarities of LCZs into account based on <a href="https://doi.org/10.1038/s41597-020-00605-z">Demuzere et al. 2020</a></td> </tr> <tr> <td><code>f1_1</code></td> <td>Class-wise metric F1 for LCZ class 1</td> </tr> <tr> <td><code>f1_2</code></td> <td>Class-wise metric F1 for LCZ class 2</td> </tr> <tr> <td><code>f1_3</code></td> <td>Class-wise metric F1 for LCZ class 3</td> </tr> <tr> <td><code>f1_4</code></td> <td>Class-wise metric F1 for LCZ class 4</td> </tr> <tr> <td><code>f1_5</code></td> <td>Class-wise metric F1 for LCZ class 5</td> </tr> <tr> <td><code>f1_6</code></td> <td>Class-wise metric F1 for LCZ class 6</td> </tr> <tr> <td><code>f1_7</code></td> <td>Class-wise metric F1 for LCZ class 7</td> </tr> <tr> <td><code>f1_8</code></td> <td>Class-wise metric F1 for LCZ class 8</td> </tr> <tr> <td><code>f1_9</code></td> <td>Class-wise metric F1 for LCZ class 9</td> </tr> <tr> <td><code>f1_10</code></td> <td>Class-wise metric F1 for LCZ class 10</td> </tr> <tr> <td><code>f1_11</code></td> <td>Class-wise metric F1 for LCZ class 11</td> </tr> <tr> <td><code>f1_12</code></td> <td>Class-wise metric F1 for LCZ class 12</td> </tr> <tr> <td><code>f1_13</code></td> <td>Class-wise metric F1 for LCZ class 13</td> </tr> <tr> <td><code>f1_14</code></td> <td>Class-wise metric F1 for LCZ class 14</td> </tr> <tr> <td><code>f1_15</code></td> <td>Class-wise metric F1 for LCZ class 15</td> </tr> <tr> <td><code>f1_16</code></td> <td>Class-wise metric F1 for LCZ class 16</td> </tr> <tr> <td><code>f1_17</code></td> <td>Class-wise metric F1 for LCZ class 17</td> </tr> </tbody> </table> <h2>Acknowledgements</h2> <p>We acknowledge all WUDAPT contributors and community members for providing the training areas via the LCZ Generator.</p>

opencc-by-nc-sa-4.0Sep 2024View details →
zenodo48/100

Supplementary Material for "Advancing quantum technology workforce: industry insights into qualification and training needs" and "Extending the European Competence Framework for Quantum Technologies: new proficiency triangle and qualification profiles"

<p>This is a file collection as supplementary material for the paper <em>Advancing quantum technology workforce: industry insights into qualification and training needs, <a href="https://doi.org/10.1140/epjqt/s40507-024-00294-2">doi 10.1140/epjqt/s40507-024-00294-2</a>.</em> It consists of:</p> <ol> <li>Interview guide: questions and more as guideline for the interviews conducted for the industry needs analysis documented in the publication.</li> <li>Interview transcript extracts: anonymised phrases from the interviews that are given as quotes (in a shortened/liguistically smoothed out form) in the publication as well as further phrases that are refered in the results sections of the publication.</li> <li>Dataset of the follow-up survey</li> </ol> <p>The results of this study were also used to update the <a href="https://doi.org/10.5281/zenodo.10976836" target="_blank" rel="noopener">European Competence Framework for Quantum Technologies Version 2.5</a>, which is documented in <em>Extending the European Competence Framework for Quantum Technologies: new proficiency triangle and qualification profiles, <a href="https://doi.org/10.1140/epjqt/s40507-024-00302-5">doi 10.1140/epjqt/s40507-024-00302-5</a></em>. In an additional sheet, the three&nbsp;draft versions of qualification profile descriptions (v2.1, v2.2, v2.3) are provided.</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Nephrops (Nephrops norvegicus) Burrow object detection simple training dataset from Irish Underwater TV surveys

<div> <div> <div> <div> <h1>Training dataset</h1> <p>Norway prawns (<em>Nephrops norvegicus</em>), also known as the Dublin Bay prawn, are common around the Irish coast. They are found in distinct sandy/muddy areas where the sediment is suitable for them to construct their burrows.&nbsp;<em>Nephrops&nbsp;</em>spend a great deal of time in their burrows and their emergence from these is related to time of year, light intensity and tidal strength. The Irish&nbsp;<em>Nephrops&nbsp;</em>fishery is extremely valuable with landings recently worth around &euro;55m at first sale, supporting an important Irish fishing industry.&nbsp;</p> <p><em>Nephrops</em> are managed in Functional Units (FUs). The Marine Institute has conducted under water television surveys since 2002 to independently estimate abundance, distribution and stock sizes of <em>Nephrops</em>&nbsp;<em>norvegicus&nbsp;</em>for:</p> <ul> <li>Irish Sea&nbsp;<em>Nephrops</em>&nbsp;Grounds (FU 14 and 15) in collaboration with&nbsp;<a title="Link to 'Fisheries and Aquatic Ecosystems' work in AFBI Northern Ireland" href="https://www.afbini.gov.uk/area-of-expertise/fisheries-and-aquatic-ecosystems">AFBI</a>&nbsp;an&nbsp;<a title="Link to Cefas (the Centre for Environment, Fisheries, and Aquaculture Science) in the UK" href="https://www.cefas.co.uk/">CEFAS</a>.</li> <li>Porcupine Bank&nbsp;<em>Nephrops</em>&nbsp;Grounds (FU16)</li> <li>Aran, Galway Bay and Slyne Head&nbsp;<em>Nephrops</em>&nbsp;Grounds (FU17)</li> <li>South and South west Ireland&nbsp;<em>Nephrops</em>&nbsp;Grounds (FU19)</li> <li>Labadie, Jones and Cockburn&nbsp;<em>Nephrops</em>&nbsp;Grounds (FU20 and 21)</li> <li>&ldquo;Smalls&rdquo;&nbsp;<em>Nephrops</em>&nbsp;Grounds (FU22)</li> </ul> <p>Each year during the summer months, on average 300 stations are surveyed each year, in three survey legs, covering all the FUs in depths from 20 to 650 metres.</p> <p>A high definition camera system is towed over the sea bed for 10 minutes travelling approx. 200m at 0.8 knots on a purpose built sledge.&nbsp;The UWTV survey follows survey protocols available&nbsp;<a title="Link to survey protocols" href="https://doi.org/10.17895/ices.pub.8014">here</a>&nbsp;agreed by International Council for the Exploration of the Sea (ICES) Working Group on&nbsp;<em>Nephrops&nbsp;</em>surveys (WGNEPS).&nbsp;</p> <p>As part of the iMagine project a selection of images from the Underwater TV survey Functional Units were annotated with bounding boxes and labels in YOLOv8 format to train an YOLOv8 Object Detection Models. The training dataset is saved in YOLOv8 format.&nbsp; It is intended to train a YOLOv8 Nephrrops burrow object detection model to assess the utility of an Object Detection model is assisting Prawn Survey work in the semi automated annotation of prawn burrow imagery.</p> </div> </div> </div> </div>

opencc-by-4.0Jul 2024View details →
zenodo48/100

AMFinder training dataset

<p>Soil fungi establish mutualistic interactions with the roots of most vascular land plants. Arbuscular mycorrhizal (AM) fungi are among the most extensively characterised mycobionts to date.</p> <p>The software <a href="https://github.com/SchornacklabSLCU/amfinder.git">AMFinder</a> allows for automatic computer vision-based identification and quantification of AM fungal colonisation and intraradical hyphal structures on ink-stained root images using convolutional neural networks.</p> <p><strong>This dataset contains ink-stained root images used for AMFinder training.</strong></p>

opencc-by-4.0Jul 2021View details →
zenodo48/100

DeepCytometer pipeline parameter files, Klf14 mouse white adipose tissue histology and hand-traced training contours

<p>Latest description of this data set:&nbsp;<a href="https://github.com/MRC-Harwell/cytometer/blob/main/DATA.md">Data.md at cytometer project</a></p> <pre># Publications related to the data The data associated to the DeepCytometer project (https://github.com/MRC-Harwell/cytometer) is available from Zenodo (doi: 10.5281/zenodo.5137433 and 10.5281/zenodo.5149005). The histology and mouse measures were generated as part of the Small et al. 2018 study: &gt; Small et al. &quot;Regulatory variants at KLF14 influence type 2 diabetes risk via a female-specific effect on adipocyte size and body composition&quot;. Nature Genetics, 50:572&ndash;580, 2018. The hand traced data set, colour maps, and automatic segmentations were generated for the Casero et al. 2021 paper: &gt; Casero et al. &quot;Phenotyping of Klf14 mouse white adipose tissue enabled by whole slide segmentation with deep neural networks&quot;. bioRxiv, 2021. doi: [10.1101/2021.06.03.444997](https://www.biorxiv.org/content/10.1101/2021.06.03.444997v1.full). # Data protocols ## Histology and laboratory measures To develop and evaluate our methods we used Klf14tm1(KOMP)Vlcg C57BL/6NTac (B6NTac) mice tissue samples and additional data generated as part of the Small et al. 2018 study(Small et al. 2018). It should be noted that the single exon Klf14 gene is imprinted and only expressed from the maternally inherited allele(Parker-Katiraee et al. 2007). This was taken into account by (Small et al. 2018) by crossing a Het parent with a WT parent, so that each offspring inherited a WT allele from the WT parent, and the Klf14 gene knockout or a WT allele from the other parent (from the father, PAT, or the mother, MAT). We also take Klf14 imprinting into account by using as controls the PAT mice and comparing them to the MAT WT and MAT Het (or functional KO, FKO) mice.&nbsp; We used a total of 76 Klf14-B6NTac mice (nfemale=nmale=38), of which 20 mice from the Control and FKO groups were used for training and testing the DeepCytometer pipeline, as well as the hand traced population experiment (summary in Table MICE). The histopathology screen involved fixing, processing and embedding in wax, sectioning and staining with Hematoxylin and Eosin (H&amp;E) both inguinal subcutaneous and gonadal adipose depots. For paraffin-embedded sections, all samples were fixed in 10% neutral buffered formalin (Surgipath) for at least 48 hours at RT and processed using an Excelsior&trade; AS Tissue Processor (Thermo Scientific). Samples were embedded in molten paraffin wax and 8 &mu;m sections were cut through the respective depots using a Finesse&trade; ME+ microtome (Thermo Scientific). Sampling was conducted at 2sxns per slide, 3 slides per depot block onto simultaneous charged slides, stained with haematoxylin Gill 3 and eosin (Thermo scientific) and scanned using an NDP NanoZoomer Digital pathology scanner (RS C10730 Series; Hamamatsu).&nbsp;Body weight (BW) and depot weight (DW) were measured with Satorius BAL7000 scales. ## White adipose tissue segmentation For cell area quantification, we applied DeepCytometer v8 to 75 inguinal subcutaneous and 72 gonadal whole histology slides with DeepCytometer (with the Corrected method), including the 20 slides sampled for the hand-traced data set, corresponding to 73 females and 74 males, to produce 2,560,067 subcutaneous and 2,467,686 gonadal cells (on average, 34,134 and 34,273 cells per slide, respectively). Full segmentation of all whole slides was performed with script [klf14_b6ntac_exp_0106_full_slide_pipeline_v8.py](https://github.com/MRC-Harwell/cytometer/blob/39358ed1d79df07d1d522b98728c7efd745513f7/scripts/klf14_b6ntac_exp_0106_full_slide_pipeline_v8.py). In this case, the segmentation contours were grouped by tiles in the output AIDA annotation `.json` file (one contour per cell, one file per slide). Non-white adipocyte contours were filtered out, and white adipocyte contours were aggregated into an AIDA annotation `.json` file with a single tile with script [klf14_b6ntac_exp_0106_annotations_postprocessing_v8.py](https://github.com/MRC-Harwell/cytometer/blob/39358ed1d79df07d1d522b98728c7efd745513f7/scripts/klf14_b6ntac_exp_0106_annotations_postprocessing_v8.py) (one contour per cell, one file per slide). # List of directories and files ## Casero et al. (2021) &quot;DeepCytometer pipeline parameter files, Klf14 mouse white adipose tissue histology and hand-traced training contours&quot; (doi: 10.5281/zenodo.5137433) ### `deepcytometer_pipeline_v8.zip` (60.6 MB) Weights, colourmaps, etc. necessary to run the pipeline (v8, with mode colour correction). This is the version of the pipeline described in the paper. There are 10 weight files per convolutional neural network (CNN), corresponding to 10-fold cross-validation * `klf14_b6ntac_exp_0086_cnn_dmap_model_fold_[0..9].h5`: Keras weights for the **EDT CNN** (Histology to Euclidean Distance Transform regression) * `klf14_b6ntac_exp_0089_cnn_segmentation_correction_overlapping_scaled_contours_model_fold_[0..9].h5`: Keras weights for the **Correction CNN** (Segmentation Correction regression) * `klf14_b6ntac_exp_0091_cnn_contour_after_dmap_model_fold_[0..9].h5`: Keras weights for the **Contour CNN** (EDT to Contour detection) * `klf14_b6ntac_exp_0095_cnn_tissue_classifier_fcn_model_fold_[0..9].h5`: Keras weights for the **Tissue CNN** (Pixel-wise tissue classifier) * `klf14_b6ntac_exp_0094_generate_extra_training_images.pickle`: training dataset description * **&#39;file_list&#39;**: list of SVG files with hand-traced contours for network training. Each SVG file has a corresponding TIFF file with the histology used for segmentation * **&#39;idx_test&#39;**: 10 lists with file indices for testing in 10-fold cross-validation * **&#39;idx_train&#39;**: 10 lists with file indices for training in 10-fold cross-validation * **&#39;fold_seed&#39;**: seed number used for the random number generator to assign file indices to folds * `klf14_b6ntac_exp_0098_filename_area2quantile.npz`: quantile colour maps calculated in `klf14_b6ntac_exp_0098_full_slide_size_analysis_v7.py` using the whole Klf14 data set with v7 of the pipeline, and used in earlier experiments, including some where v8 of the pipeline was used for segmentation. * `klf14_b6ntac_exp_0106_filename_area2quantile_v8.npz`: quantile colour maps calculated in `klf14_b6ntac_exp_0106_full_slide_pipeline_v8.py` using the whole Klf14 data set with v8 of the pipeline, and used in later experiments. * `klf14_training_colour_histogram.npz`: statistics from Klf14 histology images to be used in colour correction * **&#39;xbins_edge&#39;**, **&#39;xbins&#39;**: edges and centres of the bins used for histogram calculations * **&#39;hist_r_q1&#39;**, **&#39;hist_r_q2&#39;**, **&#39;hist_r_q3&#39;** * **&#39;hist_g_q1&#39;**, **&#39;hist_g_q2&#39;**, **&#39;hist_g_q3&#39;** * **&#39;hist_b_q1&#39;**, **&#39;hist_b_q2&#39;**, **&#39;hist_b_q3&#39;**: density quartiles (Q1, Q2, Q3) for RGB channels for each bin the histogram * **&#39;mode_r&#39;**, **&#39;mode_g&#39;**, **&#39;mode_b&#39;**: modes for RGB channels (this corresponds to the most typical background colour in the histology images) * **&#39;mean_l&#39;**, **&#39;mean_a&#39;**, **&#39;mean_b&#39;**: mean intensity for L*a*b channels of the image * **&#39;std_l&#39;**, **&#39;std_a&#39;**, **&#39;std_b&#39;**: intensity standard deviations for L*a*b channels of the image * `klf14_exp_0112_training_colour_histogram.npz`: other statistics from Klf14 histology images to be used in colour correction * **&#39;p&#39;**: vector of quantile values used in ECDF calculations * **&#39;val_r_klf14&#39;**, **&#39;val_g_klf14&#39;**, **&#39;val_b_klf14&#39;**: all intensity values for the RGB channels of Klf14 training images that contain at least a white adipocyte * **&#39;f_ecdf_to_val_r_klf14&#39;**, **&#39;f_ecdf_to_val_g_klf14&#39;**, **&#39;f_ecdf_to_val_b_klf14&#39;**: linear interpolation function that maps ECDF quantiles to intensity values in the Klf14 training data set. These functions can be used together with intensity-&gt;quantile interpolation functions calculated for a new histology image to perform histogram matching colour correction * **&#39;mean_klf14&#39;**, **&#39;std_klf14&#39;**: mean and standard deviation of the **&#39;val_r_klf14&#39;**, **&#39;val_g_klf14&#39;**, **&#39;val_b_klf14&#39;** vectors There are also weight files for the pipeline trained with all the data, instead of the 10-fold cross-validation partition. These were not used for the paper, but could be useful for future experiments * `klf14_b6ntac_exp_0101_cnn_dmap_model.h5`: Keras weights for the **EDT CNN** (Histology to Euclidean Distance Transform regression) * `klf14_b6ntac_exp_0104_cnn_segmentation_correction_overlapping_scaled_contours_model.h5`: Keras weights for the **Correction CNN** (Segmentation Correction regression) * `klf14_b6ntac_exp_0102_cnn_contour_after_dmap_model.h5`: Keras weights for the **Contour CNN** (EDT to Contour detection) * `klf14_b6ntac_exp_0103_cnn_tissue_classifier_fcn_model.h5`: Keras weights for the **Tissue CNN** (Pixel-wise tissue classifier) ### `histology.7z` (29.1 GB) 165 H&amp;E histology whole slides from Hamamatsu scanner (`.ndpi`). ### `klf14.7z` (2.3 GB) Mice metadata, training/testing data sets for the pipeline, intermediate files created during training, and neural network weights for multiple experiments. * `klf14_b6ntac_meta_info.csv`: Klf14 mice metadata * **Animal Identifier**, **id:** unique ID for each mouse * **ko_parent:** heterozygous parent of origin for the KO allele (father, PAT or mother, MAT) * **sex:** female or male * **genotype:** wild type (KLF14-KO:WT) or heterozygous (KLF14-KO:Het) * **BW:** body weight (g) * **SC:** subcutaneous depot weight (g) * **gWAT:** gonadal depot weight (g) * **Liver:** livel weight (g) * **cull_age:** age at time of culling (days) * **BW_alive:** body weight measured before culling * **BW_alive_date:** age at time of BW_alive measure * **mother:** unique ID for mouse&#39;s mother * **mother_genotype:** mouse&#39;s mother genotype * `klf14_b6ntac_training`: Directory with hand-traced segmentations of training histology windows. 131 windows sampled from 20 whole slides, plus hand-traced contours that were used for training DeepCytometer and compute population distributions. These segmentations were used for CNN training, but note that there&#39;s a cleaned-up version of these data below, and it was the cleaned-up version that was used for the paper experiments * `ndpifile_row_YYYYYY_col_XXXXXX[.tif/.xcf/.svg]`: * **ndpifile:** name of the whole slide file (e.g. `KLF14-B6NTAC 36.1c PAT 98-16 C1 - 2016-02-11 10.45.00`) * **row_YYYYYY:** Y-coordinate of the top-left corner of the sampling window, in pixels * **col_XXXXXX:** X-coordinate of the top-left corner of the sampling window, in pixels * **.tif:** TIFF file with the histology sampling window * **.xcf:** Gimp file with the histology and hand-traced contours (the contours were drawn in Gimp) * **.svg:** SVG (Scalable Vector Graphics) that contains the hand-traced contours in the XCF file * `klf14_b6ntac_training_v2`: Same as `klf14_b6ntac_training`, but the hand-traced data set was cleaned up to remove small contours of dubious cells, or cells that are fully overlapped by others * `klf14_b6ntac_training_non_overlap`: Directory with intermediate images to train the networks. These images are generated by script [`klf14_b6ntac_training_non_overlap`](https://github.com/MRC-Harwell/cytometer/blob/main/scripts/klf14_b6ntac_exp_0077_generate_non_overlap_training_images.py) * `klf14_b6ntac_training_augmented`: Directory with intermediate images used to train the networks (using augmentation to reduce overfitting). These images are generated by script [`klf14_b6ntac_exp_0078_generate_augmented_training_images.py`](https://github.com/MRC-Harwell/cytometer/blob/main/scripts/klf14_b6ntac_exp_0078_generate_augmented_training_images.py) * `klf14_b6ntac_seg`: Deprecated. Directory to store whole slide coarse segmentations in old experiments (e.g. `klf14_b6ntac_exp_0076_generate_training_images.py`). Of little interest for most users * `klf14_b6ntac_results`: Deprecated. Directory to store miscellanea output from some experiments. Of little interest for most users ## Casero et al. (2021). &quot;Klf14 mouse white adipose tissue histology DeepZoom files and AIDA annotations for visualisation of DeepCytometer white adipocyte segmentations&quot; (doi: 10.5281/zenodo.5149005) ### `aida_data_Klf14_v8_images.7z` (16.9 GB) Histology images converted to DeepZoom so that they can be visualised with [AIDA](https://github.com/alanaberdeen/AIDA). To use this, decompress this file and put the resulting `images` directory in your `AIDA/dist/data/` directory. ### `aida_data_Klf14_v8_annotations.7z` (18 GB) White adipocyte segmentations in AIDA annotation `.json` files (one contour per cell, one file per whole slide). Each slide has the following files: * `SLIDENAME.json`: Soft link to the annotations file that we want to associate to slide `SLIDENAME.ndpi`, e.g. `SLIDENAME` = `KLF14-B6NTAC-PAT-39.2d 454-16 B1 - 2016-03-17 12.16.06` * `SLIDENAME.lock`: Empty file used to tell the pipeline that `SLIDENAME.ndpi` has already been processed or is being currently processed * `SLIDENAME_coarse_mask.npz`: File with the coarse tissue segmentation of `SLIDENAME.ndpi` and the internal state of the pipeline (execution times, steps, etc) * `SLIDENAME_exp_0106_auto.json`: Annotations (all segmentations without filtering from the Auto algorithm, i.e. segmentation without object overlap). Contours are grouped by the tile they were processed in * `SLIDENAME_exp_0106_auto_aggregated.json`: Filtered annotations (non-white adipocytes removed) of the Auto algorithm. All contours aggregated into a single tile * `SLIDENAME_exp_0106_corrected.json`: Annotations (all segmentations without filtering from the Corrected algorithm, i.e. segmentation with object overlap). Contours are grouped by the tile they were processed in * `SLIDENAME_exp_0106_corrected_aggregated.json`: Filtered annotations (non-white adipocytes removed) of the Corrected algorithm. All contours aggregated into a single tile To use this, decompress this file and put the resulting `annotations` directory in your `AIDA/dist/data/` directory. </pre>

opencc-by-4.0Jul 2021View details →
zenodo48/100

Training data for neural network-based determination of nematic elastic constants

<p>Neural network training data packets (<strong><em>intensities_{i}.csv, K1K3_{i}.csv</em></strong>), each consisting of 1000 training data pairs, used in a machine learning-based method for determination of&nbsp;Frank elastic constants of nematic liquid crystals, experimental measurements of time-dependent light intensities&nbsp;(<strong><em>experimental_time</em></strong>_<strong><em>{i}.csv, experimental_intensity_{i}.csv</em></strong>), diode spectrum data (<strong><em>diode_lbd</em></strong><strong><em>.csv, diode_w.csv</em></strong>).</p> <p>These data sets are associated with the paper <a href="https://www.nature.com/articles/s41598-023-33134-x"><strong><em>[Zaplotnik et al. SciRep, 2023]</em></strong></a></p> <p>This is supplementary material for a Jupyter Notebook uploaded on&nbsp;<a href="https://zenodo.org/record/7368828">Zenodo</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo48/100

Merging Bioactivity Predictions from Cell Morphology and Chemical Fingerprint Models Using Similarity to Training Data

<p>The applicability domain of machine learning models trained on structural fingerprints for the prediction of biological endpoints is often limited by the lack of diversity of chemical space of the training data. In this work, we developed &ldquo;similarity-based merger models&rdquo; which combined the output of individual models trained on cell morphology (based on Cell Painting) and chemical structure (based on chemical fingerprints) and the structural and morphological similarities of the test compounds to training compounds. We applied these similarity-based merger models using logistic equations to weigh individual features and predicted assay hit calls of 177 assays from ChEMBL, PubChem and the Broad Institute, where the required Cell Painting annotations were available. We found that the similarity-based merger models outperformed other models with an additional 20% assays (79 out of 177 assays) with an AUC&gt;0.70 compared with 65 out of 177 assays using structural models and 50 out of 177 assays using Cell Painting models. Our results demonstrate that similarity-based merger models combining structure and cell morphology models can more accurately predict a wide range of biological assay outcomes and expand the applicability domain by better extrapolating to new structural and morphology spaces.</p>

opencc-by-4.0Jan 2023View details →
zenodo48/100

Open Soil Spectral Library (training data and calibration models)

<p><strong>Open Soil Spectral Library</strong> contains training MIR (91,631) and VisNIR (65,063) spectral scans + soil calibration data (&gt;60,000 unique locations) and calibration models. Key data set:</p> <ul> <li>ossl_all_L1_v1.2.qs: soil laboratory, site and spectra information;</li> </ul> <p>Important note: The data set spatially over-represents USA and European Union, with little training data in Asia, South America and Australia, hence calibration models reflect primarily soils of USA and Europe.</p> <p>To use the models and data please install <a href="https://hub.docker.com/r/opengeohub/r-geo">R and required packages</a>. Read more about the <strong><a href="https://github.com/traversc/qs">QS data format</a></strong> and how to convert it to CSV or similar. Modeling steps are explained in detail in: <a href="https://github.com/soilspectroscopy/ossl-models">https://github.com/soilspectroscopy/ossl-models</a>. To visualize database please use: <a href="https://explorer.soilspectroscopy.org/">https://explorer.soilspectroscopy.org/</a></p> <p>Complete OSSL documentation can be found at: <a href="https://soilspectroscopy.github.io/ossl-manual/">https://soilspectroscopy.github.io/ossl-manual/</a></p> <p><a href="https://soilspectroscopy.org/"><strong>Soil Spectroscopy for the Global Good</strong></a> is a Coordinated Innovation Network funded by USDA NIFA Food and Agriculture Cyberinformatics Tools Program (<a href="https://nifa.usda.gov/press-release/nifa-invests-over-7-million-big-data-artificial-intelligence-and-other">Award #2020-67021-32467</a>).</p> <p>Input datasets are property of the <a href="https://www.nrcs.usda.gov/wps/portal/nrcs/main/soils/research">USDA NRCS National Soil Survey Center &ndash; Kellogg Soil Survey Laboratory</a>, <a href="https://www.worldagroforestry.org/">ICRAF-World Agroforestry</a>, <a href="https://www.isric.org/">ISRIC-World Soil Information</a>, the <a href="http://africasoils.net/services/data/soil-databases/">Africa Soil Information Service</a> funded by the Bill and Melinda Gates Foundation, the <a href="https://esdac.jrc.ec.europa.eu/">European Soil Data Centre</a>, the <a href="https://www.neonscience.org/">National Ecological Observatory Network</a>, and <a href="https://sae.ethz.ch/">ETH Zurich</a>.&nbsp;</p> <p>For more advanced uses of the soil spectral libraries <strong>we advise to contact the original data producers</strong> especially to get help with using, extending and improving the original SSL data.</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

UDP Synthetic Dataset for training ML time series models

<p>The dataset available has been produced by the &quot;Next-Generation IoT solutions for the universal supply chain&quot; (iNGENIOUS) project&rsquo;s consortium under EC grant agreement 957216, &nbsp;made publicly available as part of the Horizon 2020 Open Research Data Pilot (<a href="https://www.openaire.eu/what-is-the-open-research-data-pilot">ORD pilot</a>).<br> The European Commission is not liable for any use that may be made of the information contained herein.</p> <p>The available dataset is in csv format and contains synthetic data of UDP packets received and sent by a single User Plane Function (UPF) covering a span of 6 weeks. The format of the datafile is:</p> <ul> <li>index</li> <li>timestamp&nbsp;</li> <li>UDP packets_rcvd - Total number of UDP packets received</li> <li>UDP packets sent - Total number of UDP packets sent</li> </ul> <p>The simulation was performed based on behavior of UPF and 5GC Network functions inferred from stress tests performed in the iNGENIOUS project&#39;s Automated Robots with Heterogeneous Networks Use Case, as well as patterns in urban mobility taken from available UE datasets [NCS+19].</p> <p>More information on the iNGENIOUS project can be found on the project&rsquo;s website: <a href="https://ingenious-iot.eu/">https://ingenious-iot.eu/</a></p> <p>[NCS+19] Noussan M, Carioni G, Sanvito FD, Colombo E. Urban Mobility Demand Profiles:<br> Time Series for Cars and Bike-Sharing Use as a Resource for Transport and Energy<br> Modeling. Data. 2019; 4(3):108. https://doi.org/10.3390/data4030108</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

QLKNN11D training set

<p><strong>QLKNN11D training set</strong></p> <p>This dataset contains a large-scale run of ~1 billion flux calculations of the quasilinear gyrokinetic transport model QuaLiKiz. QuaLiKiz is applied in numerous tokamak integrated modelling suites, and is openly available at <a href="https://gitlab.com/qualikiz-group/QuaLiKiz/">https://gitlab.com/qualikiz-group/QuaLiKiz/</a>. This dataset was generated with the &#39;QLKNN11D-hyper&#39; tag of QuaLiKiz, equivalent to 2.8.1 apart from the negative magnetic shear filter being disabled. See <a href="https://gitlab.com/qualikiz-group/QuaLiKiz/-/tags/QLKNN11D-hyper">https://gitlab.com/qualikiz-group/QuaLiKiz/-/tags/QLKNN11D-hyper</a> for the in-repository tag.</p> <p>The dataset is appropriate for the training of learned surrogates of QuaLiKiz, e.g. with neural networks. See <a href="https://doi.org/10.1063/1.5134126">https://doi.org/10.1063/1.5134126</a> for a Physics of Plasmas publication illustrating the development of a learned surrogate (QLKNN10D-hyper) of an older version of QuaLiKiz (2.4.0) with a 300 million point 10D dataset. The paper is also available on arXiv <a href="https://arxiv.org/abs/1911.05617">https://arxiv.org/abs/1911.05617</a> and the older dataset on Zenodo <a href="https://doi.org/10.5281/zenodo.3497066">https://doi.org/10.5281/zenodo.3497066</a>. For an application example, see Van Mulders et al 2021<a href="http://https://doi.org/10.1088/1741-4326/ac0d12"> https://doi.org/10.1088/1741-4326/ac0d12</a>, where QLKNN10D-hyper was applied for ITER hybrid scenario optimization. For any learned surrogates developed for QLKNN11D, the effective addition of the alphaMHD input dimension through rescaling the input magnetic shear (s) by s = s - alpha_MHD/2, as carried out in Van Mulders et al., is recommended.</p> <p>Related repositories:</p> <ul> <li>General QuaLiKiz documentation <a href="https://qualikiz.com">https://qualikiz.com</a></li> <li>QuaLiKiz/QLKNN input/output variables naming scheme <a href="https://qualikiz.com/QuaLiKiz/Input-and-output-variables">https://qualikiz.com/QuaLiKiz/Input-and-output-variables</a></li> <li>Training, plotting, filtering, and auxiliary tools <a href="https://gitlab.com/Karel-van-de-Plassche/QLKNN-develop">https://gitlab.com/Karel-van-de-Plassche/QLKNN-develop</a></li> <li>QuaLiKiz related tools <a href="https://gitlab.com/qualikiz-group/QuaLiKiz-pythontools">https://gitlab.com/qualikiz-group/QuaLiKiz-pythontools</a></li> <li>FORTRAN QLKNN implementation with wrapper for Python and MATLAB <a href="https://gitlab.com/qualikiz-group/QLKNN-fortran">https://gitlab.com/qualikiz-group/QLKNN-fortran</a></li> <li>Weights and biases of &#39;hyperrectangle style&#39; QLKNN <a href="https://gitlab.com/qualikiz-group/qlknn-hype">https://gitlab.com/qualikiz-group/qlknn-hype</a></li> </ul> <p><strong>Data exploration</strong></p> <p>The data is provided in 43 netCDF files. We advise opening single datasets using <a href="https://xarray.dev/">xarray </a>or multiple datasets out-of-core using <a href="https://www.dask.org/">dask</a>. For reference, we give the load times and sizes of a single variable that just depends on the scan size `dimx` below. This was tested single-core on a Intel Xeon 8160 CPU at 2.1 GHz and 192 GB of DDR4 RAM. Note that during loading, more memory is needed than the final number.</p> <table align="center"> <caption>Timing of dataset loading</caption> <thead> <tr> <th scope="col">Amount of datasets</th> <th scope="col">Final in-RAM memory (GiB)</th> <th scope="col"> <p>Loading time single var</p> (M:SS)</th> </tr> </thead> <tbody> <tr> <td>1</td> <td>10.3</td> <td>0:09</td> </tr> <tr> <td>5</td> <td>43.9</td> <td>1:00</td> </tr> <tr> <td>10</td> <td>63.2</td> <td>2:01</td> </tr> <tr> <td>16</td> <td>98.0</td> <td>3:25</td> </tr> <tr> <td>17</td> <td>Out Of Memory</td> <td>x:xx</td> </tr> </tbody> </table> <p><strong>Full dataset</strong></p> <p>The full dataset of QuaLiKiz in-and-output data is available on request. Note that this is 2.2 TiB of netCDF files!</p>

opencc-by-4.0Jun 2023View details →
zenodo48/100

Pre-training Audio Embeddings

<p>Pre-trained audio embeddings (VGGish, OpenL3, YAMNet) of <a href="https://www.upf.edu/web/mtg/irmas">IRMAS</a> and <a href="https://zenodo.org/record/1432913">OpenMIC-2018</a> datasets, released&nbsp;in the following paper:</p> <p>Changhong Wang, Brian McFee, and Ga&euml;l Richard. &quot;<strong>Transfer Learning and Bias Correction with Pre-trained Audio Embeddings</strong>&quot;.&nbsp;<em>Proceedings of the&nbsp;<a href="https://ismir2023.ismir.net/">International Society for Music Information Retrieval (ISMIR) Conference</a></em>, 2023.</p>

opencc-by-4.0Jul 2023View details →
zenodo48/100

Data for: Adaptive P300-Based Brain-Computer Interface for Attention Training

<p>The dataset contains EEG and behavioral data of 47 participants who completed 9 runs (i.e. copy-spelled 9 words) in a P300 speller task, as well as a random dot motion (RDM) task and questionnaires in a single experimental session. Details of the experimental protocol can be found here:</p> <p>Noble SC,&nbsp;Woods E,&nbsp;Ward T,&nbsp;Ringwood JV. &ldquo;Adaptive P300-Based Brain-Computer Interface for Attention Training: Protocol for a Randomized Controlled Trial.&rdquo; <em>JMIR Res Protoc</em> 2023, 12:e46135, doi:&nbsp;<a href="https://doi.org/10.2196/46135">10.2196/46135</a></p> <p>A journal article describing the results of the study can be found here:<br><br>Noble SC, Woods E, Ward T, Ringwood JV. &ldquo;Accelerating P300-Based Neurofeedback Training for Attention Enhancement Using Iterative Learning Control: A Randomised Controlled Trial.&rdquo; <em>J Neural Eng</em> 2024, 21(2), doi: <a href="https://doi.org/10.1088/1741-2552/ad2c9e" target="_blank" rel="noopener">10.1088/1741-2552/ad2c9e</a></p> <p>Please cite the results paper when using the data.</p> <p>Each participant folder contains:</p> <ul> <li>[xxx]-raw.[xxx] &ndash; unprocessed EEG signals (<strong>in</strong> <strong>mV</strong>) from 32 electrodes for all 9 P300 speller runs in Openvibe (.ov) and Matlab (.mat) file formats, see details of the runs below</li> <li>[xxx]-processed.[xxx] &ndash; contains 3 xDAWN components extracted by the xDAWN spatial filter according to the weights in &ldquo;spatial-filter.cfg&rdquo;</li> <li>classifier.cfg - LDA classifier weights</li> <li>spatial-filter.cfg - xDAWN spatial filter weights</li> <li>log.txt - contains the group assignment, start and end time of the experiment, and performance in the P300 speller and RDM tasks</li> </ul> <p>The&nbsp;file &ldquo;Subject Information.csv&rdquo; contains the age and gender of all participants.</p> <p>The file &ldquo;Questionnaire scores.csv&rdquo; contains the responses to the questionnaire described in the experimental protocol and the NASA Task Load Index (TLX) for all participants.</p> <p>The .ov and .mat files contain data from the following runs:</p> <table> <tbody> <tr> <th>Filename</th> <th>Word to be copy-spelled</th> <th>Number of flashes per row and column</th> <th>Feedback given to participant</th> </tr> </tbody> <tbody> <tr> <td>calibration-signal1</td> <td>THE</td> <td>12</td> <td>no</td> </tr> <tr> <td>calibration-signal2</td> <td>QUICK</td> <td>12</td> <td>no</td> </tr> <tr> <td>calibration-signals</td> <td>Concatenation of calibration-signal1 and calibration-signal2</td> </tr> <tr> <td>eval</td> <td>DOG</td> <td>12</td> <td>yes</td> </tr> <tr> <td>training-run-1</td> <td>BEAUTIFUL</td> <td>10</td> <td>yes</td> </tr> <tr> <td>training-run-2 to training-run-5</td> <td>BEAUTIFUL</td> <td>varying</td> <td>yes</td> </tr> <tr> <td>post-training-run</td> <td>DANCE</td> <td>12</td> <td>yes</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>This research is supported by the Irish Research Council under project ID GOIPG/2020/692 and Science Foundation Ireland under grant number 12/RC/2289_P2.</p>

opencc-by-4.0Jul 2023View details →
zenodo48/100

PANDEM-2 European COVID-19 training data set

<p>The PANDEM-2 COVID-19 European training dataset is a large collection of time series either of real or realistic synthetic (generated) data and indicators associated with the European pandemic response to the COVID-19 pandemic. It is intended to be used for training in pandemic management.</p> <p>&nbsp;</p> <p>This dataset is the result of a data gathering requirement process for pandemic management involving feedback and inputs from several public health and first responder professionals as well as researchers and military personnel directly involved in the European COVID-19 pandemic response. This work is part of the PANDEM-2 project funded by the <em>Horizon 2020 Secure Societies</em> program. To collect this data, an open source software was developed named PANDEM-Source allowing reproducibility and customisation of this dataset.&nbsp;</p> <p>&nbsp;</p> <p>The dataset includes indicators for cases, deaths, hospitalisation, testing and laboratory data including pathogen genomic information, vaccination, non-pharmaceutical interventions, participatory surveillance, social media, flights resources (human and material such as beds or vaccines), and contact tracing activities. When no open available data was found, realistic synthetic data and indicators were generated with the goal of producing a data set to be used for pandemic management&nbsp; training.&nbsp;</p> <p>The project received funding from the European Union&rsquo;s Horizon 2020 Research and Innovation programme under the Grant Agreement No. 883285. The material presented and views expressed here are the responsibility of the author(s) only. The EU Commission takes no responsibility for any use made of the information set out.<br> References</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo48/100

Pirate Illustrations for Research Data Management Training

<p>This collection of icons and comics was created to illustrate a workshop on Data Management Plans (<a href="https://doi.org/10.5281/zenodo.5575920">https://doi.org/10.5281/zenodo.5575920</a>). It is provided here to allow further reuse, for example to illustrate presentations.</p> <p>The theme of this collection is revolving around pirates, their accessories, and maritime items in general.</p> <p>Created by Jeanne Wilbrandt.</p>

opencc-by-4.0Sep 2023View details →
zenodo48/100

Pre-training with simulated ultrasound images for breast mass segmentation and classification - dataset

<p>Dataset assosiated with the MICCAI Workshop on Data Engineering in Medical Imaging paper: &quot;Pre-training with&nbsp;Simulated Ultrasound Images for&nbsp;Breast Mass Segmentation and&nbsp;Classification&quot;</p>

opencc-by-4.0Oct 2023View details →
OpenNeuro44/100

Cognitive Training

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo44/100

Hydraulic scale model experiments on the two-dimensional run-up of impulse wave trains on steep to vertical slopes

<p>This dataset includes the experimental data and videos, which were generated during the study on the run-up of impulse wave trains at the Laboratory of Hydraulics, Hydrology and Glaciology (VAW), ETH Zurich.</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

Benchmark and training data for replicating financial and insurance examples

<p>This dataset contains training, validation and out-of-sample test data for two&nbsp;European calls and two examples of portfolio of&nbsp;variable annuity guarantees.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

GIIRS RTTOV coefficient file using local training profiles

<p>GIIRS RTTOV coefficient file using local training profiles&nbsp;&nbsp;</p> <p>Reference:</p> <p>Di, D., Jun Li, Han Wei, W. Bai, C. Wu, and W. Paul Menzel, 2018: Enhancing the fast radiative transfer model for FengYun-4 GIIRS by using local training profiles, <em>Journal of Geophysical Research - Atmospheres</em>, DOI: 10.1029/2018JD029089.</p>

opencc-by-4.0Nov 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record