Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo36/100

A metadata-based approach for research discipline prediction using machine learning techniques and distance metrics

<p>The dataset is based on&nbsp;the paper:&nbsp;</p> <p>Hoang-Son Pham, Hanne Poelmans&nbsp;and Amr Ali-Eldin &lsquo;&rsquo;A metadata-based approach for research discipline prediction using machine learning techniques and distance metrics&rsquo;&rsquo;, IEEE Access (2023).</p> <p>The dataset includes:&nbsp;</p> <p>1. a list of project metadata extracted from FRIS portal</p> <p>2. a list of VODS disciplines</p> <p>3. a distance matrix</p> <p>&nbsp;</p> <p>* Kindly refer to our paper for more details on the dataset.</p> <p>https://ieeexplore.ieee.org/document/10156853</p>

opencc-by-4.0May 2023View details →
dryad36/100

Data from: Behaviour-specific spatiotemporal patterns of habitat use by sea turtles revealed using biologging and supervised machine learning

<ol> <li>Conservation of threatened species and anthropogenic threat mitigation commonly rely on spatially managed areas selected according to habitat preference. Since the impact of threats can be behaviour-specific, such information could be incorporated into spatial management to improve conservation outcomes. However, collecting spatially explicit behavioural data is challenging.</li> <li>Using multi-sensor biologging tags containing high-resolution movement sensors (e.g., accelerometer, magnetometer, GPS) and animal-borne video cameras, combined with supervised machine learning, we developed a method to automatically identify and geolocate typically ambiguous behaviours for the poorly understood flatback turtle <em>Natator depressus</em>. Subsequently, we evaluated behaviour-specific spatiotemporal patterns of habitat use.</li> <li>Boosted regression trees successfully identified the presence of foraging and resting in 7074 dives (AUC &gt; 0.9), using dive features representing characteristics of locomotory activity, body posture, and three-dimensional dive paths validated by ancillary video data. Foraging was characterised by dives with longer duration, variable depth, tortuous bottom phases; resting was characterised by dives with decreased locomotory activity and longer duration bottom phases.</li> <li>Foraging and resting showed minimal spatial segregation based on 50% and 95% utilisation distributions. Expected diel patterns of behaviour-specific habitat use were superseded by the extreme tides at the near-shore study site. Turtles rested in areas close to the subtidal and intertidal boundary within larger overlapping foraging areas, allowing efficient access to intertidal food resources upon inundation at high tides when foraging was ~25% more likely.</li> <li> <em>Synthesis and applications:</em><span> Using supervised machine learning and biologging tools, we show the potential for dynamic spatial management of flatback turtles to mitigate behaviour-specific threats by prioritising protection of important locations at pertinent times. Although results are a species-specific response to a super-tidal environment</span>, our approach can be generalised to a broad range of taxa and study systems, facilitating a conceptual advance in spatial management.</li> </ol>

opencc-zeroMay 2023View details →
zenodo36/100

Comparison between ribosomal assembly and machine learning tools for microbial identification of organisms with different characteristics

<p><strong>DNABERT+DeLUCS_notebooks.zip </strong></p> <ul> <li>Code notebooks for running DNABERT and DeLUCS</li> </ul> <p>&nbsp;</p> <p><strong>images-20230519T015235Z-001.zip </strong></p> <ul> <li>Heatmaps</li> <li>Factor plots</li> </ul> <p>&nbsp;</p> <p><strong>Data-20230518T191203Z-003.zip </strong></p> <ul> <li>MBARC and Hot Springs datasets <ul> <li>Reference genomes</li> <li>16S sequences (barrnap)</li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>hot-springs-reads.gz </strong></p> <ul> <li>Reads data for Hot Springs dataset</li> </ul> <p>&nbsp;</p> <p><strong>mbarc-reads-download.txt</strong></p> <ul> <li>Reads data for MBARC dataset <ul> <li>Link to download from NCBI</li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>assemblies.zip</strong></p> <ul> <li>Megahit and MetaSPAdes assemblies for both MBARC and Hot Springs</li> </ul>

opencc-by-4.0May 2023View details →
zenodo36/100

Generative Machine Learning for Detector Response Modeling with a Conditional Normalizing Flow

<p>The samples are datasets used for testing in the scenarios: baseline (corr0.hdf5), correlation of 0.5 (corr50.hdf5), correlation of 1.0 (corr100.hdf5), and asymmetric detector responses (asymmetric.hdf5). Each file contains the conditional variables, generated and simulated detector responses, and generated and simulated reconstruction-level variables.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Data and code example for the article: "Massively parallel hybrid quantum-classical machine learning for kernelized time-series classification"

<p>Data needed to reproduce the figures of&nbsp;<a href="https://arxiv.org/abs/2305.05881">https://arxiv.org/abs/2305.05881</a>&nbsp;and a simple code example of a quantum-convex-classical neural network&nbsp;used to train a sine versus cosine classification problem.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 3,001-4,000 (reference reconstructions and segmentations)

<p>This upload contains the reference reconstructions and segmentation of slices 3,001 &ndash; 4,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices OOD (reference reconstructions and segmentations)

<p>This upload contains the reference reconstructions and segmentation of the out-of-distribution slices (OOD) from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 1,001-2,000 (reference reconstructions and segmentations)

<p>This upload contains the reference reconstructions and segmentation of slices 1,001 &ndash; 2,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 4,001-5,000

<p>This upload contains slices 4,001 &ndash; 5,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 3,001-4,000

<p>This upload contains slices 3,001 &ndash; 4,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 1-1,000

<p>This upload contains slices 1 &ndash; 1,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 1-1,000 (reference reconstructions and segmentations)

<p>This upload contains the reference reconstructions and segmentation of slices 1 &ndash; 1,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices OOD

<p>This upload contains the out-of-distribution slices (OOD) from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 2,001-3,000 (reference reconstructions and segmentations)

<p>This upload contains the reference reconstructions and segmentation of slices 2,001 &ndash; 3,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 2,001-3,000

<p>This upload contains slices 2,001 &ndash; 3,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 1,001-2,000

<p>This upload contains slices 1,001 &ndash; 2,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning: Slices 4,001-5,000 (reference reconstructions and segmentations)

<p>This upload contains the reference reconstructions and segmentation of slices 4,001 &ndash; 5,000 from the data collection described in</p> <p>Maximilian B. Kiss, Sophia B. Coban, K. Joost Batenburg, Tristan van Leeuwen, and Felix Lucka &ldquo;&quot;2DeteCT - A large 2D expandable, trainable, experimental Computed Tomography dataset for machine learning&quot;, <a href="https://doi.org/10.1038/s41597-023-02484-6"><em>Sci Data</em> <strong>10</strong>, 576 (2023)</a> or&nbsp; <a href="https://arxiv.org/abs/2306.05907">arXiv:2306.05907 (2023)</a></p> <p>Abstract:<br> &quot;Recent research in computational imaging largely focuses on developing machine learning (ML) techniques for image reconstruction, which requires large-scale training datasets consisting of measurement data and ground-truth images. However, suitable experimental datasets for X-ray Computed Tomography (CT) are scarce, and methods are often developed and evaluated only on simulated data. We fill this gap by providing the community with a versatile, open 2D fan-beam CT dataset suitable for developing ML techniques for a range of image reconstruction tasks. To acquire it, we designed a sophisticated, semi-automatic scan procedure that utilizes a highly-flexible laboratory X-ray CT setup. A diverse mix of samples with high natural variability in shape and density was scanned slice-by-slice (5000 slices in total) with high angular and spatial resolution and three different beam characteristics: A high-fidelity, a low-dose and a beam-hardening-inflicted mode. In addition, 750 out-of-distribution slices were scanned with sample and beam variations to accommodate robustness and segmentation tasks. We provide raw projection data, reference reconstructions and segmentations based on an open-source data processing pipeline.&quot;</p> <p>The data collection has been acquired using a highly flexible, programmable and custom-built X-ray CT scanner, the FleX-ray scanner, developed by <a href="https://info.tescan.com/micro-ct">TESCAN-XRE NV,</a> located in the FleX-ray Lab at the <a href="https://www.cwi.nl/en/">Centrum Wiskunde &amp; Informatica (CWI)</a> in Amsterdam, Netherlands. It consists of a cone-beam microfocus X-ray point source (limited to 90 kV and 90 W) that projects polychromatic X-rays onto a 14-bit CMOS (complementary metal-oxide semiconductor) flat panel detector with CsI(Tl) scintillator (Dexella 1512NDT) and 1536-by-1944 pixels,&nbsp; <span class="math-tex">\(74.8\mu m^2\)</span> each. To create a 2D dataset, a fan-beam geometry was mimicked by only reading out the central row of the detector. Between source and detector there is a rotation stage, upon which samples can be mounted. The machine components (i.e., the source, the detector panel, and the rotation stage) are mounted on translation belts that allow the moving of the components independently from one another.</p> <p>Please refer to the paper for all further technical details.</p> <p>The complete dataset&nbsp;can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8014758">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8014766">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8014787">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8014829">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8014874">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8014907">OOD</a>.<br> The reference reconstructions and segmentations can be found via the following links: <a href="https://doi.org/10.5281/zenodo.8017583">1-1000</a>, <a href="https://doi.org/10.5281/zenodo.8017604">1001-2000</a>, <a href="https://doi.org/10.5281/zenodo.8017612">2001-3000</a>, <a href="https://doi.org/10.5281/zenodo.8017618">3001-4000</a>, <a href="https://doi.org/10.5281/zenodo.8017624">4001-5000</a>, <a href="https://doi.org/10.5281/zenodo.8017653">OOD</a>.</p> <p>The corresponding Python scripts for loading, pre-processing, reconstructing and segmenting the projection data in the way described in the paper can be found on <a href="https://github.com/mbkiss/2DeteCTcodes">github</a>. A machine-readable file with the used scanning parameters and instrument data for each acquisition mode as well as a script loading it can be found on the GitHub repository as well.</p> <p>Note: It is advisable to use the graphical user interface when decompressing the .zip archives. If you experience a zipbomb error when unzipping the file on a Linux system rerun the command with the UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE environment variable by setting in your .bashrc &ldquo;export UNZIP_DISABLE_ZIPBOMB_DETECTION=TRUE&rdquo;.</p> <p>For more information or guidance in using the data collection, please get in touch with</p> <p>&nbsp;&nbsp; &nbsp;Maximilian.Kiss [at] cwi.nl</p> <p>&nbsp;&nbsp; &nbsp;Felix.Lucka [at] cwi.nl</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

A dataset on in-situ electromagnetic wave integrity control of selective laser fusion printed parts using Machine Learning

<p><strong>DeepSLM Dataset</strong></p> <p>This dataset contains signals from metal parts printed in 3D using SLM (Selective laser melting) technology, produced as part of the DeepSLM project:<em> In-situ monitoring of Selective-Laser-Melted Ti6Al4V Parts Using Eddy Current Testing and Machine Learning</em>.</p> <p><em>Sensor and data collection</em></p> <p>An impedance-based non-destructive testing device is installed on the LPBF coating device. It is in the form of sensors with the following characteristics:</p> <p><em>Spectrum of analysis of the sensors</em></p> <ul> <li>Minimal length: 3-5 mm</li> <li>Minimal width acquisition: 2-3 mm (Without borders effects)</li> <li>Minimal number of layers: 50 layers</li> <li>Signal minimal penetration depth (TiAl<em>6</em>V<em>4</em>/875kHz) = 1.5 mm</li> </ul> <p>The sensors record the signal during printing, while simultaneously sending it to a computer via Bluetooth. The signals are first stored in the database. Next, a dataset is constructed from the extracted signals related to each part, and from the &nbsp;added semi-automatic annotations which describe the porosity of the part and layers.</p> <p><em>Annotations</em></p> <p>In the annotation process, there are 2 different types of labels and a metadata summary for these prints:</p> <p>1. Archimedean porosity measurement for the whole part<br>2. Layer porosity obtained after image processing applied on the resulting metallography images<br>3. Print-related metadata such as print parameters, part dimensions, layer thickness, etc</p> <p>Size: 66 parts<br>Set of printing parameters: 250</p> <p>&nbsp;--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p><strong>Dataset organization</strong></p> <p>The dataset is composed of 7 different printings and is organized as follows:<br>- signals: contains collected signals<br>- metallography: contains porosity extracted by metallography<br>- archimedean: contains porosity extracted by archimedean weighing<br>- metallography-images: contains metallography process images<br>- metadata.csv: contains print metadata (print parameter, part name, strategy, etc.)</p> <p>Each printed part has a unique ID ranging from 0 to 65. Each sub-part has an ID which is the ID of the base part suffixed with the sub-part position: 14_1 is the first sub-part of part 14. This nomenclature is uniform throughout the dataset.</p> <p>The README.md file in the dataset provides more detailed information on the structure of the dataset and the format of each of the files making up the dataset.</p> <p>&nbsp;</p> <p><strong>Relative article</strong></p> <p><a title="https://doi.org/10.1007/978-3-031-47784-3_18" href="https://doi.org/10.1007/978-3-031-47784-3_18" target="_blank" rel="noreferrer noopener">https://doi.org/10.1007/978-3-031-47784-3_18</a></p> <p>Sallem, H., Ghorbel, H., Goffinet, E., Cinna, A., Pralong, J., Wicht, J., &amp; Revaz, B. (2023, May). In-Situ Monitoring of Selective Laser Melted Ti&ndash;6Al&ndash;4V Parts Using Eddy Current Testing and Machine Learning. In <em>Advances In Additive Manufacturing Conference</em> (pp. 139-148). Cham: Springer Nature Switzerland.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Supporting data for "Machine Learning Made Easy (MLme): A Comprehensive Toolkit for Machine Learning-Driven Data Analysis"

<p>Machine learning (ML) has emerged as a vital asset for researchers to analyze and extract valuable information from complex datasets. However, developing an effective and robust ML pipeline can present a real challenge, demanding considerable time and effort, thereby impeding research progress. Existing tools in this landscape require a profound understanding of ML principles and programming skills. Furthermore, users are required to engage in the comprehensive configuration of their ML pipeline to obtain optimal performance.</p> <p>To address these challenges, we have developed a novel tool called <em>Machine Learning Made Easy </em>(MLme) that streamlines the use of ML in research, specifically focusing on classification problems at present. By integrating four essential functionalities, namely Data Exploration, AutoML, CustomML, and Visualization, MLme fulfills the diverse requirements of researchers while eliminating the need for extensive coding efforts. To demonstrate the applicability of MLme, we conducted rigorous testing on six distinct datasets, each presenting unique characteristics and challenges. Our results consistently showed promising performance across different datasets, reaffirming the versatility and effectiveness of the tool. Additionally, by utilizing MLme&#39;s feature selection functionality, we successfully identified significant markers for CD8+ na&iuml;ve (BACH2), CD16+ (CD16), and CD14+ (VCAN) cell populations.</p> <p>MLme serves as a valuable resource for leveraging machine learning (ML) to facilitate insightful data analysis and enhance research outcomes, while alleviating concerns related to complex coding scripts. The source code and a detailed tutorial for MLme are available at <a href="https://github.com/FunctionalUrology/MLme">https://github.com/FunctionalUrology/MLme</a>.</p>

opencc-zeroDec 2022View details →
zenodo36/100

Data sets and machine learning models for: Predicting critical properties and acentric factor of fluids using multi-task machine learning

<p>The experimental data sets, data splits, additional features, QM calculations, model predictions, and final machine learning models for the manuscript &quot;Predicting Critical Properties and Acentric Factor of Fluids Using Multi-Task Machine Learning&quot;.&nbsp;<strong>Citation should refer directly to the manuscript:</strong></p> <ul> <li> <p>Biswas, S.;&nbsp;Chung, Y.;&nbsp;Ramirez, J.;&nbsp;Wu, H.;&nbsp;Green, W. H.&nbsp;Predicting Critical Properties and Acentric Factors of Fluids Using Multitask Machine Learning. <em>Journal of Chemical Information and Modeling.</em>&nbsp;<strong>2023</strong>&nbsp;<em>63</em>&nbsp;(15), 4574-4588. DOI: <a href="https://doi.org/10.1021/acs.jcim.3c00546">10.1021/acs.jcim.3c00546</a></p> </li> </ul> <p>To use the machine learning&nbsp;models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a>.&nbsp;</p> <p>Detailed information&nbsp;can be found in README.md file.</p> <p>&nbsp;</p> <p><strong>Details on the properties considered</strong></p> <p>The data set includes the following 8 properties:</p> <ul> <li>Tc: critical temperature, in K</li> <li>Pc: critical pressure, in bar</li> <li>rhoc: critical density, in mol/L</li> <li>omega: acentric factor, unitless</li> <li>Tb: boiling point, in K</li> <li>Tm: melting point, in K</li> <li>dHvap: enthalpy of vaporization at boiling point, in kJ/mol</li> <li>dHfus: enthalpy of fusion at melting point, in kJ/mol</li> </ul> <p><strong>Details on the files</strong></p> <p>1. Data sets under CritProp_v1.1.0:</p> <ul> <li>all_data: includes the data sets used in this work. All data points are listed for each chemical compound as well as&nbsp;its corresponding data source. The details of the data sources can be found in the README.md file. The distribution of the data set is included in each folder. <ul> <li>estimated_data_for_pretraining:&nbsp;contains the estimated data from Yaws&#39; handbook that are used to pre-train our machine learning&nbsp;(ML) model.</li> <li>experimental_data: contains the experimental data (references 1 - 15) used to fine-tune our final ML model.</li> </ul> </li> <li>additional_features: includes the additional features tested for the ML model.&nbsp;The Abraham features are generated for all data (references 1 - 15)&nbsp;while the acsf, qm, and rdkit features are only generated for the data from references 1 - 9. <ul> <li>abraham: Abraham solute parameters (E, S, A, B, L). Molecular features.</li> <li>acsf: ACSF (atom-centered symmetry functions). Atomic features that are coverted from the 3D coordinates of the compound</li> <li>qm_atom: QM (quantum chemical) atomic feature.&nbsp;</li> <li>qm_mol: QM molecular feature.</li> <li>rdkit: Selected RDKit 2D molecular features.</li> </ul> </li> <li>data_splits_and_model_predictions: contains the training and test sets used to evaluate the model. It also&nbsp;contains the predicted values from our final ML model for each test set. <ul> <li>random and scaffold splits: training and test sets that include the data from references 1 - 9.</li> <li>external test set: a test set that includes the data from only references 10 - 15.</li> </ul> </li> </ul> <p>2. Machine learning (ML) model files:</p> <ul> <li>CritProp_ML_model_files_with_abraham_feat.zip: contains the Chemprop ML model files that are trained using Abraham features&nbsp;as additional molecular features. This gives the best results.</li> <li>CritProp_ML_model_files_without_additional_feat.zip: contains the Chemprop ML model files that are trained without&nbsp;any additional features. This gives the second best results.</li> </ul> <p>To use these ML models, please refer to the sample files and instructions on <a href="https://github.com/yunsiechung/chemprop/tree/crit_prop">https://github.com/yunsiechung/chemprop/tree/crit_prop</a></p> <p>3. QM (quantum chemical) calculations:</p> <ul> <li>QM_calculations.zip: contains the results of the QM calculations that are performed to compute QM features.</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record