Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

17

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

17 results for “Weak supervision”

Learn how ShareScore rates datasets ↗
zenodo52/100

Sample data for "A weakly supervised framework for high resolution crop yield forecasts"

<p>This dataset includes sample data for the United States to run the weakly supervised framework as described in the paper titled&nbsp;<em>A weakly supervised framework for high resolution crop yield forecasts</em>, accessible at&nbsp;</p> <table summary="Additional metadata"> <tbody> <tr> <td><a href="https://doi.org/10.48550/arXiv.2205.09016">https://doi.org/10.48550/arXiv.2205.09016</a></td> </tr> </tbody> </table> <p>&nbsp;</p> <p>The updated paper (including results from the US) is&nbsp;published in Environmental Research Letters:</p> <p><a href="https://doi.org/10.1088/1748-9326/acf50e">https://doi.org/10.1088/1748-9326/acf50e</a></p> <p>&nbsp;</p> <p>The software implementation of the machine learning baseline is available at:&nbsp;https://github.com/BigDataWUR/MLforCropYieldForecasting/tree/weaksup.</p> <p>&nbsp;</p> <p>Data</p> <p>1. County data (county-data.zip)&nbsp;for county-level strongly supervised models:</p> <p>*&nbsp;CROP_AREA_COUNTY_US.csv: County crop production area statistics (acres). Source: NASS (USDA-NASS, 2022).</p> <p>*&nbsp;CSSF_COUNTY_US.csv: Crop productivity indicators including total above-ground production (kg ha<sup>-1</sup>), total weight of storage organs (kg ha<sup>-1</sup>), development stage (0-2). Source: de Wit et al. (2022).</p> <p>*&nbsp;METEO_COUNTY_US.csv: Meteo data including maximum, minimum, average daily air temperature (℃);&nbsp;sum of daily precipitation (PREC) (mm);&nbsp;sum of daily evapotranspiration of short vegetation (ET0) (Penman-Monteith, Allen et al., (1998)) (mm);&nbsp;climate water balance = (PREC - ET0) (mm). Source: Boogaard et al. (2022).</p> <p>*&nbsp;REMOTE_SENSING_COUNTY_US.csv: Fraction of Absorbed Photosynthetically Active Radiation (Smoothed) (FAPAR). Source: Copernicus GLS (2020).</p> <p>*&nbsp;SOIL_COUNTY_US.csv: Soil water holding capacity. Source: WISE Soil Property Database (Batjes, 2016).</p> <p>*&nbsp;YIELD_COUNTY_US.csv: County yield statistics (bushels/acre). Source: NASS (USDA-NASS, 2022).</p> <p>&nbsp;</p> <p>2. 10-km grid data (grid-data.zip) for grid-level strongly supervised models:</p> <p>* COUNTY_GRIDS_US.csv: Mapping between counties and grids.</p> <p>*&nbsp;CSSF_GRIDS_US.csv: Crop productivity indicators at 10km grid level (similar to county data above).</p> <p>*&nbsp;METEO_GRIDs_US.csv: Meteo data at 10km grid level&nbsp;(similar to county data above).</p> <p>*&nbsp;REMOTE_SENSING_GRIDS_US.csv: FAPAR at 10km grid level (similar to county data above).</p> <p>*&nbsp;SOIL_GRIDS_US.csv: Soil water holding capacity at 10km grid level (similar to county data above).</p> <p>*&nbsp;YIELD_GRIDS_US.csv: Grid-level modeled yields (t ha<sup>-1</sup>). Source: Deines et al. (2021), Lobell et al.&nbsp;(2020).</p> <p>&nbsp;</p> <p>3. County labels and 10-km grid inputs (dscale-US.zip) for weak supervision:</p> <p>* COUNTY_GRIDS_US.csv: Mapping between counties and grids.</p> <p>*&nbsp;CSSF_GRIDS_US.csv: Crop productivity indicators at 10km grid level.</p> <p>*&nbsp;METEO_GRIDs_US.csv: Meteo indicators at 10km grid level.</p> <p>*&nbsp;REMOTE_SENSING_GRIDS_US.csv: FAPAR at 10km grid level.</p> <p>*&nbsp;SOIL_GRIDS_US.csv: Soil water holding capacity at 10km grid level.</p> <p>*&nbsp;YIELD_GRIDS_US.csv: Grid-level modeled yields (t ha<sup>-1</sup>). Source: Deines et al. (2021).</p> <p>*&nbsp;YIELD_COUNTY_US.csv: County yield statistics (bushels/acre). Source: NASS (USDA-NASS, 2022).</p> <p>*&nbsp;CROP_AREA_COUNTY_US.csv: County crop production area statistics (acres). Source: NASS (USDA-NASS, 2022).</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Not So Weak-PICO: Leveraging weak supervision for Participants, Interventions, and Outcomes recognition for systematic review automation

<p>EBM-PICO is a widely used dataset with PICO annotations at two levels: span-level or coarse-grained and entity-level or fine-grained. Span-level annotations encompass the full information about each class. Entity-level annotations cover the more fine-grained information at the entity level, with PICO classes further divided into fine-grained subclasses. For example, the coarse-grained Participant span is further divided into participant age, gender, condition and sample size in the randomised controlled trial. This dataset comes pre-divided into a training set (n=4,933) annotated through crowd-sourcing and an expert annotated gold test set (n=191) for evaluation.</p> <p>The <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6174533/bin/NIHMS988059-supplement-Appendix.pdf">EBM-PICO annotation guidelines</a> caution about variable annotation quality. <a href="http://ceur-ws.org/Vol-2429/paper1.pdf">Abaho et al.</a> developed a framework to post-hoc correct EBM-PICO outcomes annotation inconsistencies. <a href="https://arxiv.org/pdf/1904.09557.pdf">Lee et al.</a> studied annotation span disagreements suggesting variability across the annotators. Low annotation quality in the training dataset is excusable, but the errors in the test set can lead to faulty evaluation of the downstream ML methods. We evaluate 1% of the EBM-PICO training set tokens to gauge the possible reasons for the fine-grained labelling errors and use this exercise to conduct an error-focused PICO re-annotation for the EBM-PICO gold test set.&nbsp;The file &#39;test_ebm_correctedlabels.tsv&#39; has error corrected EBM-PICO gold test set.</p> <p>&nbsp;</p> <p>The upload also contains two zip files containing labelling sources mentioned in the Distant-PICO paper.&nbsp;</p> <ol> <li>ds_cto_dict.zip: contains the four distant supervision dictionaries (P: &nbsp;participant.txt, I = intervention.txt, intervetion_syn.txt, O: outcome.txt) generated from clinicaltrials.gov using the methodology described in Distant-CTO.&nbsp;</li> <li>handcrafted_dictionaries.zip: contains three files <ul> <li>gender_sexuality.txt: contains a list of possible genders and sexual orientations found across the web. The list is not comprehensive.</li> <li>endpoints_dict.txt: contains outcome names and the names of questionnaires used to measure outcomes assembled from PROM questionnaires and PROMs.</li> <li>comparator_dict: contains a list of idiosyncratic comparator terms like a sham, saline, placebo, etc., compiled from the literature search. The list is not comprehensive.</li> </ul> </li> </ol>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Weakly Supervised Learning for Industrial Optical Inspection

<p><strong>Abstract</strong></p> <p>In the following, we present a synthetic benchmark corpus for detect detection on statistically textured surfaces.We hope that it facilitates to further develop and benchmark classification algorithms for applications of industrial optical inspection. All data is publicly available and can be downloaded from this page.</p> <p><strong>Competition at DAGM 2007 symposium</strong></p> <p>The <a href="https://www.dagm.de/">DAGM (Deutsche Arbeitsgemeinschaft f&uuml;r Mustererkennung e.V., German chapter of the IAPR (International Association for Pattern Recognition))</a> and the <a href="http://www.gnns.de/">GNSS (German Chapter of the European Neural Network Society)</a> offered an open competition on <em>Weakly Supervised Learning for Industrial Optical Inspection</em> held as part of the DAGM symposium in 2007.<br><br>The competition was inspired by the fact that automated optical inspection allows to reduce the cost of industrial quality control significantly. The competitors had to design a classification algorithm which:</p> <ul> <li>detects miscellaneous defects on various statistically textured backgrounds.</li> <li>learns to discern defects automatically from a weakly labelled training data.</li> <li>works on data whose exact characteristics are unknown at development time.</li> <li>adapts all parameters automatically and does not require any human intervention.</li> <li>has a moderate running time (in this competition 24 hours for training and 12 hours for the test phase).</li> <li>takes into account asymmetric costs for false positive and false negative decisions (1:20 was used for the competition).</li> </ul> <p><strong>Data description</strong></p> <p>Preview Image: <a href="../api/iiif/record:12750201:examples_small.jpg/full/!800,800/0/default.jpg" target="_blank" rel="noopener">https://zenodo.org/api/iiif/record:12750201:examples_small.jpg/full/!800,800/0/default.jpg</a></p> <p>The data is artificially generated, but similar to real world problems. The first six out of ten datasets, denoted as development datasets, are supposed to be used for algorithm development. The remaining four datasets, which are referred to as competition datasets, can be used to evaluate the performance. Researchers should consider not using or analyzing the competition datasets before the development is completed as a code of honour.<br>In the following we provide some details about the datasets:</p> <ul> <li>Each development (competition) dataset consists of 1000 (2000) 'non-defective' and of 150 (300) 'defective' images saved in grayscale 8-bit PNG format.</li> <li>Each dataset is generated by a different texture model and defect model.</li> <li>'Non-defective' images show the background texture without defects, 'defective' images have exactly one labelled defect on the background texture.</li> <li>All datasets has been randomly split into a training and testing sub-dataset of equal size.</li> <li>Weak labels are provided as ellipses roughly indicating the defective area. Technically, defective images are augmented with a separate grayscale 8-bit image in the PNG format located in a folder 'Label'. The values 0 and 255 denote background and defective area, respectively.</li> </ul> <p>All meta-data is subsumed in a separate ASCII textfile called 'Labels.txt' which is located in the 'Label' folder. The structure is as follows:<br>1 \n<br>[id of item no. 1] \t [0 if non-defective, 1 if defective] \t [filename of raw image no. 1] \t 0 \t [filename of label image no. 1 if defective, 0 otherwise] \n<br>...<br>[id of item no. N] \t [0 if non-defective, 1 if defective] \t [filename of raw image no. N] \t 0 \t [filename of label image no. N if defective, 0 otherwise] \n</p>

opencc-by-4.0Sep 2007View details →
zenodo44/100

TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak Supervision

<p>This repository contains the silver standard dataset and code for the paper &quot;TweetDIS: A Large Twitter Dataset for Natural Disasters Built using Weak Supervision&quot;.</p> <p>The file &quot;heuristic_uniq_terms_nd.txt&quot; contains the list of terms used as the heuristic and the file &quot;natural_disasters_ssd_tweetids.tsv&quot; contains the tweet ids in&nbsp;the silver standard dataset.&nbsp;</p> <p>To hydrate the tweets, you can use tools like twarc or Social Media Mining toolkit - https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7362951/</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

DoMars16k: A Diverse Dataset for Weakly Supervised Geomorphologic Analysis on Mars

<p>The dataset contains 16150 samples extracted from&nbsp;163 CTX images. Each sample depicts&nbsp;one of fifteen Martian surface landforms.&nbsp;One CTX image&nbsp;contributed with at least three and at most 1247 patches to the creation of the dataset. The dataset is subdivided into training, test, and validation sets, which contain seventy, ten, and twenty percent of the samples. The sets are mutually exclusive. Each sample has a size of 200 x 200 px or roughly 1.2km x 1.2km.</p> <p><strong>Contents</strong></p> <ul> <li>data.zip contains the dataset separated into&nbsp;training, validation, and test sets.&nbsp;</li> <li>models.zip contains pre-trained neural networks.</li> </ul> <p><strong>Classes</strong></p> <ul> <li><strong>Aeolian Bedforms</strong> <ul> <li>Aeolian Curved (ael)</li> <li>Aeolian Straight (aec)</li> </ul> </li> <li><strong>Topographic Landforms</strong> <ul> <li>Cliff (cli)</li> <li>Ridge (rid)</li> <li>Channel (fsf)</li> <li>Mounds (sfe)</li> </ul> </li> <li><strong>Slope Feature Landforms</strong> <ul> <li>Gullies (fsg)</li> <li>Slope Streaks (fse)</li> <li>Mass Wasting (fss)</li> </ul> </li> <li><strong>Impact Landforms</strong> <ul> <li>Crater (cra)</li> <li>Crater Field (sfx)</li> </ul> </li> <li><strong>Basic Terrain Landforms</strong> <ul> <li>Mixed Terrain (mix)</li> <li>Rough Terrain (rou)</li> <li>Smooth Terrain (smo)</li> <li>Textured Terrain (tex)</li> </ul> </li> </ul> <p><strong>Code</strong></p> <p>Python code to train, evaluate, and apply deep neural networks to Martian surface data is available at GitHub:&nbsp;<a href="https://github.com/thowilh/geomars">https://github.com/thowilh/geomars</a></p> <p><strong>Attribution</strong></p> <p>If you find this work useful please consider citing:</p> <p>Wilhelm, T.; Geis, M.; P&uuml;ttschneider, J.; Sievernich, T.; Weber, T.; Wohlfarth, K.; W&ouml;hler, C. DoMars16k: A Diverse Dataset for Weakly Supervised Geomorphologic Analysis on Mars.&nbsp;Remote Sens.&nbsp;2020,&nbsp;12, 3981.</p> <pre><code>@article{wilhelm2020domars16k, doi = {10.3390/rs12233981}, url = {https://doi.org/10.3390/rs12233981}, year = {2020}, month = dec, publisher = {{MDPI} {AG}}, volume = {12}, number = {23}, pages = {3981}, author = {Thorsten Wilhelm and Melina Geis and Jens P\"{u}ttschneider and Timo Sievernich and Tobias Weber and Kay Wohlfarth and Christian W\"{o}hler}, title = {{DoMars}16k: A Diverse Dataset for Weakly Supervised Geomorphologic Analysis on Mars}, journal = {Remote Sensing} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
dryad40/100

Not so weak-PICO: Leveraging weak supervision for Participants, Interventions, and Outcomes recognition for systematic review automation

<p><strong>Objective: </strong>PICO (Participants, Interventions, Comparators, Outcomes) analysis is vital but time-consuming for conducting systematic reviews (SRs). Supervised machine learning can help fully automate it, but a lack of large annotated corpora limits the quality of automated PICO recognition systems. The largest currently available PICO corpus is manually annotated, which is an approach that is often too expensive for the scientific community to apply. Depending on the specific SR question, PICO criteria are extended to PICOC (C-Context), PICOT (T-timeframe), and PIBOSO (B-Background, S-Study design, O-Other) meaning the static hand-labelled corpora need to undergo costly re-annotation as per the downstream requirements. We aim to test the feasibility of designing a weak supervision system to extract these entities without hand-labelled data.</p> <p><strong>Methodology:</strong> We decompose PICO spans into its constituent entities and re-purpose multiple medical and non-medical ontologies and expert-generated rules to obtain multiple noisy labels for these entities. These labels obtained using several sources are then aggregated using simple majority voting and generative modelling approaches. The resulting programmatic labels are used as weak signals to train a weakly-supervised discriminative model and observe performance changes. We explore mistakes in the currently available PICO corpus that could have led to inaccurate evaluation of several automation methods.</p> <p><strong>Results: </strong>We present Weak-PICO, a weakly-supervised PICO entity recognition approach using medical and non-medical ontologies, dictionaries and expert-generated rules. Our approach does not use hand-labelled data.</p> <p><strong>Conclusion: </strong>Weak supervision using weak-PICO for PICO entity recognition has encouraging results, and the approach can potentially extend to more clinical entities readily. </p>

opencc-zeroDec 2022View details →
dryad40/100

Not so weak-PICO: Leveraging weak supervision for Participants, Interventions, and Outcomes recognition for systematic review automation

Open the record for dataset details and reuse information.

publicDec 2022View details →
zenodo36/100

Weakly-Supervised Crack Detection Dataset

<p>This repo contains two files: crack detection dataset (weakly_sup_crackdet_dataset.zip), and pretrained TensorFlow model for Xception65 (pascal_voc_seg.zip).</p> <p>The&nbsp;dataset consists of&nbsp;rough annotations used in weakly-supervised crack detection. It contains roughly annotated ground truths for the following datasets:</p> <ul> <li> <p><a href="https://www.irit.fr/~Sylvie.Chambon/AigleRN_GT.html">Aigle</a></p> </li> <li> <p><a href="https://github.com/cuilimeng/CrackForest-dataset">Crack Forest Dataset</a></p> </li> <li> <p><a href="https://github.com/yhlleo/DeepCrack">DeepCrack</a></p> </li> </ul> <p>Annotations of different &quot;roughness&quot; are stored. Directories suffixed &quot;*_dil*&quot; are synthetically-generated annotations, while directories suffixed &quot;*_rough&quot; and &quot;*_rougher&quot; are manually-generated annotations. The detail of the dataset is described in [1]. Please also refer to our GitHub repo&nbsp;<a href="https://github.com/hitachi-rd-cv/weakly-sup-crackdet">https://github.com/hitachi-rd-cv/weakly-sup-crackdet</a> for more details.</p> <p>This dataset is made available by Hitachi, Ltd.</p> <p>The pretrained model is used by [1]. Please use it for comparison experiments. Please refer to our GitHub repor for more details.</p> <p>[1] Inoue, Y., Nagayoshi, H.: Crack detection as a weakly-supervised problem: Towards achieving less annotation-intensive crack detectors. In: International Conference on Pattern Recognition (ICPR) (2020)</p>

opencc-by-4.0Nov 2020View details →
zenodo36/100

Experimental Data for the Paper 'Hierarchical Fusion and Divergent Activation Based Weakly Supervised Learning for Object Detection from Remote Sensing Images'

<p><strong>Experimental Data for the Paper &#39;Hierarchical Fusion and Divergent Activation Based Weakly Supervised Learning for Object Detection from Remote Sensing Images&#39;</strong></p> <p>In this repository, we provide the implementation of the algorithms developed in the paper &#39;Hierarchical Fusion and Divergent Activation Based Weakly Supervised Learning for Object Detection from Remote Sensing Images&#39; along with the experimental results, and the methods used for comparison.<br> The goal is to provide the elements needed to validate and reproduce our research work as well as all the tools needed to reach the same conclusions as we did.<br> The data used in our experiments that we have the copyright of [<a href="http://doi.org/10.1109/ACCESS.2020.3019956">A</a>],[<a href="http://doi.org/10.5281/zenodo.3843229">B</a>]&nbsp;is already published <a href="http://doi.org/10.5281/zenodo.3843229">on zenodo</a>.<br> The licences valid for the elements of this repository are discussed under point &quot;3. Licenses&quot; below.</p> <p><strong>1. Structure</strong></p> <p>The repository contains the following items:</p> <ol> <li>&quot;CODE_AND_RESULTS.zip&quot;&nbsp;with the source codes and results of our method and the comparison methods,</li> <li>&quot;README&quot;&nbsp;-&nbsp;this text here.</li> <li>&quot;LICENSE&quot;&nbsp;-&nbsp;the <a href="https://mit-license.org/">MIT License</a></li> </ol> <p>We now focus on the structure of the file CODE_AND_RESULTS.zip.<br> It contains the following items:</p> <ol> <li>The directory &quot;new_methods&quot;&nbsp;contains the source code and results of the new methods proposed in our paper.</li> <li>The directory &quot;comparison&quot; contains the source code of the two approaches used for comparison: ACoL [<a href="https://doi.org/10.1109/CVPR.2018.00144">A</a>]&nbsp;and DANet [<a href="http://doi.org/10.1109/ICCV.2019.00669">B</a>].</li> <li>The folder &quot;tools_and_metrics&quot; holds additional libraries, software tools, and metrics using in our experiments.&nbsp;</li> <li>&quot;README&quot; - this text here.</li> <li>&quot;LICENSE&quot; -&nbsp;the <a href="https://mit-license.org/">MIT License</a></li> </ol> <p>Inside the folder &quot;new_methods,&quot; the following sub-folders are provided:</p> <ol> <li>&quot;data&quot; includes data loading code and code for how organizing the input data of the neural network.</li> <li>&quot;expr&quot; includes training code.</li> <li>&quot;model&quot; includes neural network model, basic network and additional modules, depending on the file name, including improved network, and comparison model.</li> <li>&quot;utils&quot; includes some used library functions and test codes when testing, including image segmentation, searching for the largest connected area and data visualization, etc. Verification on the WSADD dataset is done via test_airplane.py and on the DIOR dataset via val_model.py.</li> </ol> <p>In our experiments, we used two datasets:</p> <p>&quot;WSADD&quot; [<a href="http://doi.org/10.1109/ACCESS.2020.3019956">A</a>],[<a href="http://doi.org/10.5281/zenodo.3843229">B</a>], which is already published <a href="http://doi.org/10.5281/zenodo.3843229">on zenodo</a>&nbsp;under the <a href="https://creativecommons.org/licenses/by/4.0/legalcode">Creative Commons Attribution 4.0 International</a>&nbsp;license.<br> The &quot;<a href="https://doi.org/10.1109/CVPR.2018.00144">DIOR</a>&quot;&nbsp;proposed in [<a href="http://doi.org/10.1016/j.isprsjprs.2019.11.023">C</a>].</p> <p><strong>2. References</strong></p> <p>[<a href="http://doi.org/10.1109/ACCESS.2020.3019956">A</a>]&nbsp;Z.-Z. Wu, T. Weise, Y. Wang, Y. Wang, Convolutional neural network based weakly supervised learning for aircraft detection from remote sensing image, <em>IEEE Access</em>&nbsp;8 (2020) 158097-158106. doi:<a href="http://doi.org/10.1109/ACCESS.2020.3019956">10.1109/ACCESS.2020.3019956</a>. &nbsp;&nbsp;<br> [<a href="http://doi.org/10.5281/zenodo.3843229">B</a>]&nbsp;Z.-Z. Wu. Weakly Supervised Airplane Detection Dataset: WSADD. May 2020. zenodo.org. doi:<a href="http://doi.org/10.5281/zenodo.3843229">10.5281/zenodo.3843229</a>.<br> [<a href="http://doi.org/10.1016/j.isprsjprs.2019.11.023">C</a>]&nbsp;K. Li, G. Wan, G. Cheng, L. Meng, J. Han, Object detection in optical remote sensing images: A survey and a new benchmark, <em>ISPRS Journal of Photogrammetry and Remote Sensing</em>&nbsp;159 (2020) 296-307. doi:<a href="http://doi.org/10.1016/j.isprsjprs.2019.11.023">10.1016/j.isprsjprs.2019.11.023</a>. &nbsp;&nbsp;<br> [<a href="https://doi.org/10.1109/CVPR.2018.00144">D</a>]&nbsp;X. Zhang, Y. Wei, J. Feng, Y. Yang, T. S. Huang, Adversarial complementary learning for weakly supervised object localization, in: <em>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</em>&nbsp;(CVPR&#39;18), Jun. 18-22, 2018, Salt Lake City, UT, USA, IEEE Computer Society, 2018, pp. 1325-1334. doi:<a href="https://doi.org/10.1109/CVPR.2018.00144">10.1109/CVPR.2018.00144</a>. &nbsp;&nbsp;<br> [<a href="http://doi.org/10.1109/ICCV.2019.00669">E</a>] H. Xue, C. Liu, F. Wan, J. Jiao, X. Ji, Q. Ye, DANet: Divergent activation for weakly supervised object localization, in: <em>Proceedings of the IEEE/CVF International Conference on Computer Vision</em>&nbsp;(ICCV&#39;19), Oct. 27-Nov. 2, 2019, Seoul, Korea, IEEE, 2019, pp. 6588-6597. doi:<a href="http://doi.org/10.1109/ICCV.2019.00669">10.1109/ICCV.2019.00669</a>.</p> <p><strong>3. Licenses</strong></p> <p>The following licenses apply for the files and folders in the archive &quot;CODE_AND_RESULTS.zip&quot;:</p> <ul> <li>The files in the folder `new_methods` are under the <a href="https://mit-license.org/">MIT License</a>.</li> <li>The files in the folder `comparison/ACoL` have been obtained from https://github.com/xiaomengyc/ACoL, which is under the <a href="https://mit-license.org/">MIT License</a>.</li> <li>We put our code and data under the&nbsp;</li> <li>The files in the folder &quot;comparison/DANet&quot; have been obtained from <a href="https://github.com/xuehaolan/DANet">https://github.com/xuehaolan/DANet</a>, which is an open source project without associated license at the time of this writing. They will therefore remain under the copyright of the user <a href="https://github.com/xuehaolan/">https://github.com/xuehaolan/</a>.</li> <li>The files in the folder &quot;tools_and_metrics/detections_DIOR&quot; are related to the repository <a href="https://github.com/rafaelpadilla/Object-Detection-Metrics">https://github.com/rafaelpadilla/Object-Detection-Metrics</a>, which is under the <a href="https://mit-license.org/">MIT License</a>, and therefore are under the same license.</li> <li>The files in the folder &quot;tools_and_metrics/Nest-pytorch&quot; are based on the repository <a href="https://github.com/ZhouYanzhao/Nest">https://github.com/ZhouYanzhao/Nest</a>, which is under the <a href="https://mit-license.org/">MIT License</a>.</li> <li>The files in the folder &quot;tools_and_metrics/PRM-pytorch&quot; are based on the repository <a href="https://github.com/ZhouYanzhao/PRM">https://github.com/ZhouYanzhao/PRM</a>, which is an open source project without associated license at the time of this writing. They will therefore remain under the copyright of the user <a href="https://github.com/ZhouYanzhao/">https://github.com/ZhouYanzhao/</a>.</li> </ul> <p>The <a href="https://mit-license.org/">MIT License</a> is included here as file &quot;LICENSE&quot;.</p> <p><strong>4. Contact</strong></p> <p>1. Dr. <a href="http://iao.hfuu.edu.cn/146">Zhize WU</a>, wuzz@hfuu.edu.cn<br> 2. Dr. <a href="http://iao.hfuu.edu.cn/5">Thomas WEISE</a>, tweise@hfuu.edu.cn, tweise@ustc.edu.cn</p> <p>Institute of Applied Optimization, &nbsp;&nbsp;<br> School of Artificial Intelligence and Big Data, &nbsp;&nbsp;<br> Hefei University, South Campus 2, Jinxiu Dadao 99, &nbsp;&nbsp;<br> Hefei Economic and Technological Development Area, &nbsp;&nbsp;<br> Shushan District, Hefei 230601, Anhui, China<br> &nbsp;</p>

openmit-licenseJan 2021View details →
zenodo36/100

AbdomenCT-1K: Weakly Supervised Learning Benchmark

<p>This is the dataset of AbdomenCT-1K: Weakly Supervised Learning Benchmark.</p> <p>Related paper: <a href="https://ieeexplore.ieee.org/document/9497733/">https://ieeexplore.ieee.org/document/9497733/</a></p> <p>Benchmark homepage: https://abdomenct-1k-weaklysupervisedlearning.grand-challenge.org/</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Data, code, models for "Weakly Supervised Semantic Segmentation for Joint Key Local Structure Localization and Classification of Aurora Image"

<p>Data, code and models for https://ieeexplore.ieee.org/document/8410588/</p>

opencc-by-4.0Jul 2018View details →
zenodo32/100

Weakly Supervised Airplane Detection Dataset: WSADD

<p>The images in this dataset includes airport and nearby areas of different countries (mainly from China, the United States, the United Kingdom, France, Japan, and Singapore) taken from the Google Earth satellite. The dataset comprises 700 remote sensing images in total, of which 400 images contain an aircraft target (the &lsquo;&rsquo;positive sample set&rsquo;&rsquo;) and the other 300 do not (the &lsquo;&rsquo;negative sample set&rsquo;&rsquo;). In the process of dataset construction, the spatial resolution of the image was controlled between 0.3 m and 2 m, and the size was fixed to 768&times;768. We sought to collect images from different sensors during different daytimes, different seasons, and different light intensities to ensure that the dataset has a high diversity.</p>

opencc-by-4.0May 2020View details →
zenodo32/100

A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains

<p><strong>Content</strong></p> <p>This repository contains pre-trained computer vision models, data labels, and images used in the pre-print publication &quot;A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains&quot;:</p> <ol> <li><em>ADPdevkit</em>: a folder containing the 50 validation (&quot;tuning&quot;) set and 50 evaluation (&quot;segtest&quot;) set of images from the Atlas of Digital Pathology database formatted in the VOC2012 style--the full database of 17,668 images is available for download from the original website</li> <li><em>VOCdevkit</em>: a folder containing the relevant files for the PASCAL VOC2012 Segmentation dataset, with both the trainaug and test sets</li> <li><em>DGdevkit</em>: a folder containing the 803 test images of the DeepGlobe Land Cover challenge dataset formatted in the VOC2012 style</li> <li><em>cues</em>: a folder containing the pre-generated weak cues for ADP, VOC2012, and DeepGlobe datasets, as required for the SEC and DSRG methods</li> <li><em>models_cnn</em>: a folder containing the pre-trained CNN models</li> <li><em>models_wsss</em>: a folder containing the pre-trained SEC, DSRG, and IRNet models, along with dense CRF settings</li> </ol> <p><strong>More information</strong></p> <p>For more information, please refer to the following article.&nbsp;<strong>Please cite this article when using the data set.</strong></p> <p>@misc{chan2019comprehensive,<br> &nbsp; &nbsp; title={A Comprehensive Analysis of Weakly-Supervised Semantic Segmentation in Different Image Domains},<br> &nbsp; &nbsp; author={Lyndon Chan and Mahdi S. Hosseini and Konstantinos N. Plataniotis},<br> &nbsp; &nbsp; year={2019},<br> &nbsp; &nbsp; eprint={1912.11186},<br> &nbsp; &nbsp; archivePrefix={arXiv},<br> &nbsp; &nbsp; primaryClass={cs.CV}<br> }</p> <p>For the full code released on GitHub, please visit the repository at:&nbsp;<a href="https://github.com/lyndonchan/wsss-analysis">https://github.com/lyndonchan/wsss-analysis</a></p> <p><strong>Contact</strong></p> <p>For questions, please contact:<br> Lyndon Chan<br> lyndon.chan@mail.utoronto.ca<br> http://orcid.org/0000-0002-1185-7961</p>

opencc-by-4.0Jun 2020View details →
zenodo32/100

Weakly-Supervised Learning Significantly Reduces the Number of Labels Required for Intracranial Hemorrhage Detection on Head CT

<p>Modern machine learning pipelines, in particular those based on deep learning (DL) models, require large amounts of labeled data. For classification problems, the most common learning paradigm consists of presenting labeled examples during training, thus providing \emph{strong supervision} by directly presenting examples from the different classes, e.g. positive and negative samples. As a result, the adequate training of these models demands the curation of large datasets with high-quality labels. This constitutes a major obstacle for the development of DL models in radiology---in particular for cross-sectional imaging (e.g., computed tomography [CT] scans)---where labels must come from manual annotations by expert radiologists at the image or slice-level (as opposed to the examination level, such as could be obtained using natural language processing of radiology reports).&nbsp;<br> This work studies the question of what kind of labels&nbsp;should be collected for the problem of intracranial hemorrhage detection in brain CT. We investigate whether image-level annotations should be preferred to examination-level ones. By framing this task as a Multiple Instance Learning (MIL) problem, and employing modern attention-based DL architectures, we analyze the degree to which different levels of supervision improve the detection performance. We find that strong supervision (learning with local image-level annotations) and weak supervision (learning with only global examination-level labels) achieve comparable performance in both examination- and image-level hemorrhage detection, as well as in hemorrhage localization at the pixel-level (explainability). Furthermore, we study this behavior as a function of the number of labels available during training. Our results suggest that local labels may not be necessary at all, drastically reducing the time and cost involved in collecting and curating datasets.</p>

opencc-by-4.0Nov 2022View details →
zenodo32/100

Online Fusion of Multi-resolution Multispectral Images with Weakly Supervised Temporal Dynamics

<p>This is the data set for experiments of satellite images of&nbsp;Oroville dam site in paper:&nbsp;Online Fusion of Multi-resolution Multispectral Images with Weakly Supervised Temporal Dynamics. Inside the zip file, there are 3 folders: &#39;HD-IMG-Database-Landsat-8&#39;, &#39;HD-IMG-Database-Landsat-8-Qest&#39; and &#39;MODIS_250&#39;.</p>

opencc-by-4.0Jan 2023View details →
zenodo28/100

Multi-granular Software Classification using File-Level Weak Supervision - Data

<p>Data for our paper:&nbsp;Multi-granular Software Classification using File-Level Weak Supervision</p>

opencc-by-4.0May 2023View details →
zenodo24/100

A Weak Supervision-Based Approach to Improve Chatbots for Code Repositories

<p>The dataset and scripts used&nbsp;in &quot;A Weak Supervision-Based Approach to Improve Chatbots for Code Repositories&quot; paper.</p>

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record