Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “Vision Transformer”
Transformers Model Zoos and Soups: A Population of Language and Vision Models
<p>Model Zoos submitted to the NeurIPS 2024 Dataset & Benchmark track: "<em>Transformer Model Zoos and Soups: A Population of Language and Vision Models</em>"</p> <p>We generate two model zoos, one for computer vision built on the ViT-S architecture, and one for language modeling based on the BERT architecture. For each, we train several backbone models with varying hyperparameters, and further fine-tune them using multiple hyperparameter combinations. We further annotate every model with performance metrics. These include test accuracy and F1-score, as well as the generalization gap. For the vision models, we also include the robust accuracy after a FGSM attack.</p>
Vision-Transformer, ViT, model validation dataset
<p><span>The U.S. cotton industry is highly concerned with removing plastic contamination from cotton lint. A major source of this </span><span>contamination is the plastic used to wrap cotton modules produced by John Deere round module harvesters. A machine-vision </span><span>detection and removal system has been developed to address this problem, using low-cost color cameras to detect plastic in the </span><span>cotton stream and remove it. However, the system requires a lot of calibration and is difficult for cotton gin workers to operate due to </span><span>its reliance on custom machine-vision classifier running on low-cost ARM computers running Linux. This research aims to make the system more user-friendly by adding an </span><span>auto-calibration feature that can track cotton colors and avoid plastic images, reducing the need for skilled personnel to operate the </span><span>system and making it easier for the cotton ginning industry to adopt. This image dataset was created to validate several Vision-</span><span>Transformer, ViT, AI models that in combination provides the key enabling technology for the auto-calibration code.</span></p>
Unbiased single-cell morphology with self-supervised vision transformers -- Cell Painting
<p>The data necessary to reproduce the Cell Painting results in the paper <a href="https://www.biorxiv.org/content/10.1101/2023.06.16.545359v1">Unbiased single-cell morphology with self-supervised vision transformers</a>. </p>
Unbiased single-cell morphology with self-supervised vision transformers -- HPA FOV
<p>The data necessary to reproduce the HPA FOV results in the paper <a href="https://www.biorxiv.org/content/10.1101/2023.06.16.545359v1">Unbiased single-cell morphology with self-supervised vision transformers</a>. </p>
Unbiased single-cell morphology with self-supervised vision transformers -- HPA single cells
<p>The data necessary to reproduce the HPA single cells results in the paper <a href="https://www.biorxiv.org/content/10.1101/2023.06.16.545359v1">Unbiased single-cell morphology with self-supervised vision transformers</a>. </p>
Unbiased single-cell morphology with self-supervised vision transformers -- WTC11
<p>The data necessary to reproduce the WTC11 results in the paper <a href="https://www.biorxiv.org/content/10.1101/2023.06.16.545359v1">Unbiased single-cell morphology with self-supervised vision transformers</a>. </p> <p> </p>
Estimating Compositions and Nutritional Values of Seed Mixes based on Vision Transformers
<p>The cultivation of seed mixtures for local pastures is a traditional mixed cropping techniques of cereals and legumes for producing at a low production cost, a balanced animal feed in energy and protein in livestock systems. By considerably improving the autonomy and safety of agricultural systems, as well as reducing their impact on the environment, it is a type of crop that responds favorably both to the evolution of the European regulations on the use phyto-sanitary products, and the expectations of consumers who wish to increase their consumption of organic products. However, farmers find it difficult to adopt it because cereals and legumes do not ripen synchronously and the harvested seeds are heterogeneous, making it more difficult to assess their nutritional value. Many efforts therefore remain to be made to acquire and aggregate technical and economical references to evaluate to what extent the cultivation of seed mixtures could positively contribute to secure and reduce costs on herd feeding. The work presented in this paper proposes to evaluate recent deep learning techniques that could be transferred to an online or smartphone application to automatically estimate the nutritive value of harvested seed mixes to help farmers better managing the yield and thus engage them to promote and contribute to better knowledge of this type of cultivation. For this purpose, we have built an original image dataset containing 4,749 images of seed mixes, covering 11 seed varieties, with which we have compared 2 types of deep learning models. Our results highlight the potential of this method, and show that the best performing model is a recent state-of-the-art Vision Transformer pre-trained with self-supervision (BeiT). It allows an estimation of the nutritive value of seed mixtures with a coefficient of determination <span class="math-tex">\(R^2\)</span> Score of 0.91, which demonstrates the interest of this type of approach, for its possible use on a large scale.</p>
MMV_Im2Im: An Open Source Microscopy Machine Vision Toolbox for Image-to-Image Transformation
<p>This dataset contains trained deep learning models and sample data for the manuscript "MMV_Im2Im: An Open Source Microscopy Machine Vision Toolbox for Image-to-Image Transformation". Please find the software and more information including tutorials here: https://github.com/MMV-Lab/mmv_im2im.</p><p> </p><p>sample_data.zip includes the following datasets:</p><ul><li>Labelfree prediction of nuclear structure from 2D/3D brighteld images<ul><li>2D<ul><li><a href="https://zenodo.org/record/6139958#.Y78QJKrMLtU">https://zenodo.org/record/6139958#.Y78QJKrMLtU</a></li><li><a href="https://zenodo.org/record/6140064#.Y78YeqrMLtU">https://zenodo.org/record/6140064#.Y78YeqrMLtU</a></li><li>Both repositories have a Creative Commons Attribution 4.0 International License</li></ul></li><li>3D<ul><li><a href="https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset">https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset</a></li><li>Terms of use: <a href="https://www.allencell.org/terms-of-use.html">https://www.allencell.org/terms-of-use.html</a></li><li>"Your use of the Content, including creation of derivative works of the services, data and tools, must be for research or other noncommercial purposes unless it is otherwise set forth in these Terms or agreed to in writing by the Allen Institute."</li></ul></li></ul></li><li>2D semantic segmentation of tissues from H&E images<ul><li><a href="https://www.kaggle.com/datasets/sani84/glasmiccai2015-gland-segmentation">https://www.kaggle.com/datasets/sani84/glasmiccai2015-gland-segmentation</a></li><li>"<strong>The dataset used in this competition is provided for research purposes only. Commercial uses are not allowed.</strong><br>If you intend to publish research work that uses this dataset, you must cite our review paper to be published after the competition"</li></ul></li><li>Instance segmentation<ul><li>2D<ul><li><a href="https://bbbc.broadinstitute.org/BBBC010">https://bbbc.broadinstitute.org/BBBC010</a></li><li>Terms of use: <a href="https://bbbc.broadinstitute.org/">https://bbbc.broadinstitute.org/</a></li><li>"Researchers are encouraged to use these image sets as reference points when developing, testing, and publishing new image analysis algorithms for the life sciences."</li></ul></li><li>3D<ul><li><a href="https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset">https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset</a></li><li>Terms of use: <a href="https://www.allencell.org/terms-of-use.html">https://www.allencell.org/terms-of-use.html</a></li><li>"Your use of the Content, including creation of derivative works of the services, data and tools, must be for research or other noncommercial purposes unless it is otherwise set forth in these Terms or agreed to in writing by the Allen Institute."</li></ul></li></ul></li><li>Compare semantic segmentation and instance segmentation<ul><li><a href="https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset">https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset</a></li><li>Terms of use: <a href="https://www.allencell.org/terms-of-use.html">https://www.allencell.org/terms-of-use.html</a></li><li>"Your use of the Content, including creation of derivative works of the services, data and tools, must be for research or other noncommercial purposes unless it is otherwise set forth in these Terms or agreed to in writing by the Allen Institute."</li></ul></li><li>Unsupervised semantic segmentation<ul><li><a href="https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset">https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset</a></li><li>Terms of use: <a href="https://www.allencell.org/terms-of-use.html">https://www.allencell.org/terms-of-use.html</a></li><li>"Your use of the Content, including creation of derivative works of the services, data and tools, must be for research or other noncommercial purposes unless it is otherwise set forth in these Terms or agreed to in writing by the Allen Institute."</li></ul></li><li>Generating synthetic images<ul><li><a href="https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset">https://open.quiltdata.com/b/allencell/packages/aics/hipsc_single_cell_image_dataset</a></li><li>Terms of use: <a href="https://www.allencell.org/terms-of-use.html">https://www.allencell.org/terms-of-use.html</a></li><li>"Your use of the Content, including creation of derivative works of the services, data and tools, must be for research or other noncommercial purposes unless it is otherwise set forth in these Terms or agreed to in writing by the Allen Institute."</li></ul></li><li>Image denoising<ul><li><a href="https://csbdeep.bioimagecomputing.com/scenarios/">https://csbdeep.bioimagecomputing.com/scenarios/</a></li><li>Two datasets: "Denoising in 3D (Planaria nuclei)" and "Denoising in 3D (Tribolium nuclei)"</li><li>Terms of use: <a href="http://csbdeep.bioimagecomputing.com/">http://csbdeep.bioimagecomputing.com/</a></li><li>"The entire CSBDeep toolbox is fully open source and intended to be used from either Python or <a href="https://fiji.sc">Fiji</a>."</li></ul></li><li>Imaging modality transformation<ul><li><a href="https://zenodo.org/record/4624364#.Y9bWOoHMIqJ">https://zenodo.org/record/4624364#.Y9bWOoHMIqJ</a></li><li>Two datasets: "Confocal_2_STED.zip" (Microtubule and Nuclear_Pore_complex)</li><li>Repository has a Creative Commons Attribution 4.0 International License</li></ul></li><li>Staining transformation:<ul><li><a href="https://zenodo.org/record/4751737#.Y9gbv4HMLVZ">https://zenodo.org/record/4751737#.Y9gbv4HMLVZ</a></li><li>Dataset "BC-DeepLIIF_Training_Set.zip" and "BC-DeepLIIF_Validation_Set.zip"</li><li>Repository has a Creative Commons Attribution 4.0 International License</li></ul></li></ul><p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.