Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,433

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,433 results for “masks”

Learn how ShareScore rates datasets ↗
edi60/100

Evaluation of Mask R-CNN Model for Counting Reproductive Structures of Six Plant Species 1895-2018

Phenology––the timing of life-history events––is a key trait for understanding responses of organisms to climate. The digitization and online mobilization of herbarium specimens is rapidly advancing our understanding of plant phenological response to climate and climatic change. The current common practice of manually harvesting data from individual specimens greatly restricts our ability to scale data collection to entire collections. Recent investigations have demonstrated that machine-learning models can facilitate data collection from herbarium specimens. However, present attempts have focused largely on simplistic binary coding of reproductive phenology (e.g., flowering or not). Here, we use crowd-sourced phenological data of numbers of buds, flowers, and fruits of more than 3000 specimens of six common wildflower species of the eastern United States (Anemone canadensis, A. hepatica, A. quinquefolia, Trillium erectum, T. grandiflorum, and T. undulatum} to train a model using Mask R-CNN to segment and count phenological features. A single global model was able to automate the binary coding of reproductive stage with greater than 90% accuracy. Segmenting and counting features were also successful, but accuracy varied with phenological stage and taxon. Counting buds was significantly more accurate than flowers or fruits. Moreover, botanical experts provided more reliable data than either crowd-sourcers or our Mask R-CNN model, highlighting the importance of high-quality human training data. Finally, we also demonstrated the transferability of our model to automated phenophase detection and counting of the three Trillium species, which have large and conspicuously-shaped reproductive organs. These results highlight the promise of our two-phase crowd-sourcing and machine-learning pipeline to segment and count reproductive features of herbarium specimens, providing high-quality data with which to study responses of plants to ongoing climatic change.

openCC0Dec 2023View details →
zenodo48/100

Mask at 300 m of water-body locations more than 5, 15 and 20 km distant from land

<p>Locations of water-body locations remote from land:&nbsp;This dataset is a latitude-longitude grid indicating&nbsp;the locations of water-body locations more distant from land than 5, 15 and 20 km. It is derived from Carrea et al., 2016, which in turn was derived from the ESA Climate Change Initiative for Land Cover Water Bodies product released in October 2014. 3 = distance greater than 20 km; &gt;=2 = distance greater than 15 km; &gt;=1 = distance greater than 5 km. Paper describing underlying distance-to-land dataset: Carrea, L., Embury, O., Merchant, C.J. (2016) Datasets related to inland water for limnology and remote sensing applications: distance-to-land, distance-to-water, water-body identifier and lake-centre co-ordinates. Geoscience Data Journal, 2(2). pp. 83-97. doi: https://doi.org/10.1002/gdj3.32. This work done within the project: ESA Climate Change Initiative Lakes, by University of Reading, UK. &nbsp;</p> <p>&nbsp;&#39;geospatial_lat_min&#39;: -90.0,\<br> &nbsp;&#39;geospatial_lat_max&#39;: 90.0,\<br> &nbsp;&#39;geospatial_lon_min&#39;: -180.0,\<br> &nbsp;&#39;geospatial_lon_max&#39;: 180.0,\<br> &nbsp;&#39;geospatial_lat_units&#39;: &#39;degrees_north&#39;,\<br> &nbsp;&#39;geospatial_lat_resolution&#39;: 0.0027777778,\<br> &nbsp;&#39;geospatial_lon_units&#39;: &#39;degrees_east&#39;,\<br> &nbsp;&#39;geospatial_lon_resolution&#39;: 0.0027777778,\<br> &nbsp;&#39;spatial_resolution&#39;: &#39;300m&#39;</p>

opencc-by-4.0Apr 2020View details →
zenodo48/100

Global Mangrove Watch: Mangrove Habitat Mask

<p>This is the habitat mask used to define the locations where mangroves can be found. It was used&nbsp;during the creation&nbsp;of the Global Mangrove Watch (GMW;&nbsp;<a href="https://www.globalmangrovewatch.org/?map=eyJiYXNlbWFwIjoibGlnaHQiLCJ2aWV3cG9ydCI6eyJsYXRpdHVkZSI6MjAsImxvbmdpdHVkZSI6MCwiem9vbSI6MiwiYmVhcmluZyI6MCwicGl0Y2giOjB9fQ%3D%3D">https://www.globalmangrovewatch.org</a>) extent products. Details of how this layer was originally produced are within Bunting et al., 2018 but it has subsequently been edited with further regions added as the GMW layers have been updated and improved. This is considered a living dataset&nbsp;which is edited, and new versions are produced&nbsp;when missing areas or improvements are identified. New&nbsp;versions will be uploaded here on zenodo.</p> <p><strong>Relevant publications:</strong></p> <p>Bunting, P., Rosenqvist, A., Lucas, R., Rebelo, L.-M., Hilarides, L., Thomas, N., Hardy, A., Itoh, T., Shimada, M., Finlayson, C., 2018. The Global Mangrove Watch&mdash;A New 2010 Global Baseline of Mangrove Extent. Remote Sens-basel 10, 1669. <a href="https://doi.org/10.3390/rs10101669" target="_blank" rel="noopener">https://doi.org/10.3390/rs10101669</a></p> <p>Bunting, Pete, Rosenqvist, Ake, Lucas , Richard, Rebelo, Lisa-Maria, Hilarides, Lammert, Thomas, Nathan, Hardy, Andy, Itoh, Takuya, Shimada, Masanobu, &amp; Finlayson, Max. (2019). Global Mangrove Watch (1996 - 2016) Version 2.0 (2.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.5658808" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.5658808</a></p> <p>Bunting, P., Rosenqvist, A., Hilarides, L., Lucas, R.M., Thomas, N., 2022. Global Mangrove Watch: Updated 2010 Mangrove Forest Extent (v2.5). Remote Sens-basel 14, 1034. <a href="https://doi.org/10.3390/rs14041034" target="_blank" rel="noopener">https://doi.org/10.3390/rs14041034</a></p> <p>Pete Bunting, Ake Rosenqvist, Lammert Hilarides, Richard M. Lucas, &amp; Nathan Thomas. (2022). Global Mangrove Watch 2010 Baseline (v2.5) (2.5) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.5828339" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.5828339</a></p> <p>Bunting, P., Rosenqvist, A., Hilarides, L., Lucas, R.M., Thomas, N., Tadono, T., Worthington, T.A., Spalding, M., Murray, N.J., Rebelo, L.-M., 2022. Global Mangrove Extent Change 1996&ndash;2020: Global Mangrove Watch Version 3.0. Remote Sens-basel 14, 3657. <a href="https://doi.org/10.3390/rs14153657" target="_blank" rel="noopener">https://doi.org/10.3390/rs14153657</a></p> <p>Bunting, P., Rosenqvist, A., Hilarides, L., Lucas, R., Thomas, N., Tadono, T., Worthington, T., Spalding, M., Murray, N., Rebelo, L.-M., (2022) Global Mangrove Watch (1996 - 2020) Version 3.0 Dataset.&nbsp;<a href="https://doi.org/10.5281/zenodo.6894273" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.6894273</a></p> <p><br><br>&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Hail Event on 2022-06-28 in Locarno-Monti (TI), Switzerland: Drone Photogrammetry Imagery, Mask R-CNN Model and Analysis Data of Hailstones

<p>This hail data collection belongs to a drone hail survey performed on 2022-06-28 in Locarno-Monti (TI, Switzerland). The supercell reached the location around 07:50 UTC in the morning. Only one photogrammetry flight could be performed and thus no estimation of the hail melting process is available. The orthophoto is masked to ignore parts where detection of hail is unwanted.</p> <p>&nbsp;</p> <p>Expert 1 (lai, mlainer), Expert 2 (jtm), Expert 3 (por, jportmann)</p>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Lake mask and distance to land dataset of 2024 lakes for the European Space Agency Climate Change Initiative Lakes v2

<p>This dataset contains the distance to land and the lake identifiers as a global netcdf file for all the water pixels at 1km (1/120 deg) lat/lon resolution of 2024 lakes distributed globally. It contains also the list of lakes as a csv file with information such as the lake center as defined in [1], and the coordinate of a box to easily locate the like in the global netcd file. The mask excludes islands on lakes and it has been derived from the GloboLakes high resolution limnology dataset [2]. The dateset have been&nbsp;further harmonized with the lake maximum extent lake polygons by PML [3]. The lake list with the plot of the mask and the polygons is available as a html file accessible also from the lake website at the University of Reading: http://www.laketemp.net/home_CCI/LMPolygons.php</p> <p>This dataset accompanies the <strong>ESA CCI Lakes v2 dataset</strong> [4].</p> <p>&nbsp;</p> <p>[1] Carrea, L.; Embury, O.; Merchant, C.J. (2015): High-resolution datasets related to in-land water for limnology and remote sensing applications: distance-to-land, distance-to-water, water-body identifier and lake-centre co-ordinates - Geoscience Data Journal, 2 (2). pp. 83-97. ISSN 2049-6060 doi: https://doi.org/10.1002/gdj3.32</p> <p>[2] Carrea, L.; Embury, O.; Merchant, C.J. (2015): GloboLakes: high-resolution global limnology dataset v1. Centre for Environmental Data Analysis. doi:10.5285/6be871bc-9572-4345-bb9a-2c42d9d85ceb. <a href="http://dx.doi.org/10.5285/6be871bc-9572-4345-bb9a-2c42d9d85ceb">http://dx.doi.org/10.5285/6be871bc-9572-4345-bb9a-2c42d9d85ceb</a></p> <p>[3] Simis, S.; Mata, A.; Selmes, N.; Carrea, L. (2021) Lake polygons dataset accompanying Calimnos v1.4.0 and ESA CCI Lakes Climate Research Data Package v2.0. zenodo https://doi.org/10.5281/zenodo.4899250</p> <p>[4] Carrea, L.; Cr&eacute;taux, J.-F.; Liu, X.; Wu, Y.; Berg&eacute;-Nguyen, M.; Calmettes, B.; Duguay, C.; Jiang, D.; Merchant, C.J.; Mueller, D.; Selmes, N.; Simis, S.; Spyrakos, E.; Stelzer, K.; Warren, M.; Yesou, H.; Zhang, D. (2022): ESA Lakes Climate Change Initiative (Lakes_cci): Lake products, Version 2.0.1. NERC EDS Centre for Environmental Data Analysis <a href="https://catalogue.ceda.ac.uk/uuid/03c935c6890c4b2ebf4aae4d84cd9472">https://catalogue.ceda.ac.uk/uuid/03c935c6890c4b2ebf4aae4d84cd9472</a></p>

opencc-by-4.0Apr 2022View details →
zenodo48/100

Kodaikanal Solar Observatory (KoSO) White-Light Sunspot Regions Masks (1904-2017)

<p>Regular observations at the Kodaikanal Solar Observatory (KoSO) began in 1904 using a white-light telescope with a 10-cm aperture lens and an f/15 light beam. Between 1912 and 1917, the objective lens was changed several times. In 1918, a 15-cm achromatic lens was installed. This new configuration produced a 20.4 cm size image of the Sun in the image plane. Photographic plates were used to capture the image. The same telescope has been used since 1918 up until 2017 to take regular white-light observations of the Sun. This data set provides the sunspot mask in HDF5 format for all the White Light Observations acquired at KoSO. Each HDF5 file contains the sunspot mask for all the observations for that year. The sunspot masks are provided in two different coordinate systems: (i) Full Disk as observed and (ii) Carrington heliographic coordinate, which is transformed from full disk using near point interpolation. Each data set also contains metadata in the form of HDF5 attributes. The Carrington co-ordinate data if full Sun map, hence the near-side &nbsp;of the Sun is the region where values in the mask in non-zero, where as sunspot regions are filled with value 2.&nbsp;</p> <p>A&nbsp;<strong>Python package (KoSOpy), which can be located on <a href="https://github.com/Kodaikanal-Solar-Observatory/kosopy" target="_blank" rel="noopener">GitHub</a>,</strong>&nbsp;is being developed which can be used to navigate through these data sets.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Masks for ISIMIP3 Agriculture (GGCMI phase 3) model runs

<p>This dataset consists of a netCDF file with a number of layers at half-degree global resolution. Each layer is a binary map representing whether each gridcell is included (1 if yes, 0 if no) in one or more input datasets used in the ISIMIP3 Agriculture (GGCMI phase 3) model runs. Individual-dataset masks:</p> <ul> <li>has_soil&nbsp;indicates&nbsp;inclusion in the&nbsp;<a href="https://data.isimip.org/10.48364/ISIMIP.942125">ISIMIP3 soil input dataset (Volkholz &amp; M&uuml;ller, 2020)</a>&nbsp;(ignoring &quot;gravel,&quot; which has some missing cells).</li> <li>has_cropcals indicates inclusion in the <a href="https://zenodo.org/record/5062513">J&auml;germeyr et al. (publication in prep.)</a>&nbsp;crop calendar dataset (ignoring second-season rice, which is not grown in all gridcells).</li> <li>has_lu indicates inclusion in all 15 area maps in the historical land use area dataset <a href="https://protocol.isimip.org/protocol/ISIMIP3b/index.html#socioeconomic-forcing">prepared for ISIMIP3</a> (landuse-totals_histsoc_annual_1850_2014.nc).</li> <li>has_crops indicates inclusion in all 15 area maps in the historical 15-crop dataset <a href="https://protocol.isimip.org/protocol/ISIMIP3b/index.html#socioeconomic-forcing">prepared for ISIMIP3</a> (landuse-15crops_histsoc_annual_1850_2014.nc). Note that there are two gridcells that are missing from this&nbsp;</li> <li>has_fertilizer indicates inclusion in every fertilizer_application_histsoc*.nc file in the&nbsp;<a href="http://doi.org/10.5281/zenodo.4954582">fertilizer and manure dataset prepared by Heinke et al. (2021) for GGCMI3</a>.</li> <li>has_all is a composite mask indicating inclusion in all of the above.</li> </ul> <p>Also included are a figure showing the masks and the MATLAB script used to generate the data and figure.</p> <ul> </ul>

opencc-by-4.0Jun 2021View details →
zenodo48/100

Mask of large scale river catchments

<p>The catchment mask provides information about the location of large scale river catchments on a global grid. Its purpose is the provision of a common reference for the computation of area averages, especially for the analysis of Earth System Model output.</p>

openbsd-3-clauseDec 2018View details →
OpenNeuro44/100

Multi-echo masking test dataset

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
zenodo44/100

Sample of facial mask N95 and FFP2 pricing on retail webs and time evolution per country

<p>We&#39;ve gathered - for a Data Science&nbsp;educational project&nbsp;- the pricing of several face mask for breathing protection in a given period of time.</p> <p>Countries : Spain&#39;, &#39;USA&#39;, &#39;France&#39;, &#39;UK&#39;, &#39;Germany&#39;,&nbsp;&#39;Italy&#39;, &#39;Netherlands&#39;, &#39;Australia&#39;</p> <p>&nbsp;</p> <p>&#39;asin&#39; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; type:&nbsp;STRING &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&quot;C&oacute;digo de identif&iacute;caci&oacute;n &uacute;nico de product equivalente de AMAZON&quot;</p> <p>&#39;description&#39;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;type:&nbsp;&nbsp;STRING &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &#39;Texto descriptivo del producto&#39;</p> <p>&#39;dateTime&#39; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; type:&nbsp;TIMESTAMP&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &#39;Cadena de carateres que contiene fecha y hora GMT&#39;</p> <p>&#39;date&#39; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Type:&nbsp;DATETIME &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &#39; Formato diferente de la misma fecha / hora de captura &#39;</p> <p>&#39;country&#39; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;type:&nbsp; STRING &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&#39;Pais al que pertenece a distribuci&oacute;n del producto &#39;Valores posibles: &#39;</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Virus+ Sequence Masked Mouse Reference Genome (GRCm38)

<p>A version of the mouse genome (<a href="https://www.ncbi.nlm.nih.gov/assembly/327618">GRCm38</a>)&nbsp;masked for all possible viral sequences.</p> <p>See&nbsp;<a href="https://zenodo.org/record/4116107#.X5B7ti9h3UI">Virus+ Masked Human Genome</a> for a masked human reference database.</p> <p>The following commands were used to generate the additional virus sequence masked reference database:</p> <p><strong>1) Download all RefSeq and Neighbor nucleotide records:</strong></p> <p><a href="https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])">https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])</a></p> <p><strong>2) Shred the downloaded viral genomes using shred.sh from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a>&nbsp;package</strong></p> <p>shred.sh in=refseq_virus_reformated.fasta out=virus_shred.fasta.gz length=85 minlength=75 overlap=30</p> <p><strong>3) Map shredded virus sequence to the GRCm38</strong><strong> genome using bbmap.sh&nbsp;from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a>&nbsp;package</strong></p> <p>bbmap.sh ref=GRCm38.fa.gz in=virus_shred.fasta.gz outm=map_mouse_all_viruses.sam minid=0.90</p> <p><strong>4) Mask virus sequenced mapped regions from the&nbsp;GRCm38 genome using bbmask.sh from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a>&nbsp;package</strong></p> <p>bbmask.sh in=GRCm38.fa.gz out=GRCm38_virus_masked.fasta.gz sam=map_mouse_all_viruses.sam</p> <p><strong>5) Remove all N&#39;s to further reduce file size using&nbsp;<a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></strong><br> seqkit -is replace -p &quot;n&quot; -r &quot;&quot; GRCm38_virus_masked.fasta.gz &nbsp;&gt;&nbsp;mouse_virus_masked.fasta_Ns_removed.gz</p> <p><strong>Additional References:</strong></p> <ol> <li><a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a></li> <li><a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></li> <li><a href="https://www.ncbi.nlm.nih.gov/genome/viruses/">NCBI Virus Genome RefSeq</a></li> </ol>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Virus+ Sequence Masked Human Reference Genome (hg19)

<p>A version of the human genome (hg19) originally masked for ribosomal, plant, animal, fungal and&nbsp;low-entropy sequences&nbsp;by Brian Bushnell (<a href="https://zenodo.org/record/1208052#.X5BuTy9h3UI">Bushnell Masked Human Genome</a>) additionally masked for all possible viral sequences.</p> <p>The following commands were used to generate the additional virus sequence masked reference database:</p> <p><strong>1) Download all RefSeq and Neighbor nucleotide records:</strong></p> <p><a href="https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])">https://www.ncbi.nlm.nih.gov/nuccore/?term=Viruses[Organism]%20NOT%20cellular%20organisms[ORGN]%20NOT%20wgs[PROP]%20NOT%20gbdiv%20syn[prop]%20AND%20(srcdb_refseq[PROP]%20OR%20nuccore%20genome%20samespecies[Filter])</a></p> <p><strong>2) Shred the downloaded viral genomes using shred.sh from the <a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>shred.sh in=refseq_virus_reformated.fasta out=virus_shred.fasta.gz length=85 minlength=75 overlap=30</p> <p><strong>3) Map shredded virus sequence to the hg19-masked human genome using bbmap.sh&nbsp;from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>bbmap.sh ref=hg19_main_mask_ribo_animal_allplant_allfungus.fa.gz in=virus_shred.fasta.gz outm=map_human_all_viruses.sam minid=0.90</p> <p><strong>4) Mask virus sequenced mapped regions from the hg19-masked human genome using bbmask.sh from the&nbsp;<a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a> package</strong></p> <p>bbmask.sh in=hg19_main_mask_ribo_animal_allplant_allfungus.fa.gz out=human_virus_masked.fasta.gz sam=map_human_all_viruses<br> .sam</p> <p><strong>5) Remove all N&#39;s to further reduce file size using <a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></strong><br> seqkit -is replace -p &quot;n&quot; -r &quot;&quot; human_virus_masked.fasta.gz &nbsp;&gt; human_virus_masked.fasta_Ns_removed.gz</p> <p><strong>Additional References:</strong></p> <ol> <li><a href="http://seqanswers.com/forums/showthread.php?t=42552">http://seqanswers.com/forums/showthread.php?t=42552</a> for additional information on the original masking of hg19</li> <li><a href="https://jgi.doe.gov/data-and-tools/bbtools/">bbtools</a></li> <li><a href="https://bioinf.shenwei.me/seqkit/">seqkit</a></li> <li><a href="https://www.ncbi.nlm.nih.gov/genome/viruses/">NCBI Virus Genome RefSeq</a></li> </ol>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Sentinel-2 Cloud Mask Catalogue

<p><strong>Overview</strong></p> <p>This dataset comprises cloud masks for 513 1022-by-1022 pixel subscenes, at 20m resolution, sampled random from the 2018 Level-1C Sentinel-2 archive. The design of this dataset follows from some observations about cloud masking: (i) performance over an entire product is highly correlated, thus subscenes provide more value per-pixel than full scenes, (ii) current cloud masking datasets often focus on specific regions, or hand-select the products used, which introduces a bias into the dataset that is not representative of the real-world data, (iii) cloud mask performance appears to be highly correlated to surface type and cloud structure, so testing should include analysis of failure modes in relation to these variables.</p> <p>The data was annotated semi-automatically, using the <a href="https://github.com/ESA-PhiLab/iris">IRIS toolkit</a>, which allows users to dynamically train a Random Forest (implemented using <a href="https://github.com/microsoft/LightGBM">LightGBM</a>), speeding up annotations by iteratively improving it&#39;s predictions, but preserving the annotator&#39;s ability to make final manual changes when needed. This hybrid approach allowed us to process many more masks than would have been possible manually, which we felt was vital in creating a large enough dataset to approximate the statistics of the whole Sentinel-2 archive.</p> <p>In addition to the pixel-wise, 3 class (CLEAR, CLOUD, CLOUD_SHADOW) segmentation masks, we also provide users with binary<br> classification &quot;tags&quot; for each subscene that can be used in testing to determine performance in specific circumstances. These include:</p> <ul> <li><strong>SURFACE TYPE</strong>: <em>11 categories</em></li> <li><strong>CLOUD TYPE</strong>: <em>7 categories</em></li> <li><strong>CLOUD HEIGHT</strong>: <em>low, high</em></li> <li><strong>CLOUD THICKNESS</strong>: <em>thin, thick</em></li> <li><strong>CLOUD EXTENT</strong>: <em>isolated, extended</em></li> </ul> <p>&nbsp;</p> <p>Wherever practical, cloud shadows were also annotated, however this was sometimes not possible due to high-relief terrain, or large ambiguities. In total, 424 were marked with shadows (if present), and 89 have shadows that were not annotatable due to very ambiguous shadow boundaries, or terrain that cast significant shadows. If users wish to train an algorithm specifically for cloud shadow masks, we advise them to remove those 89 images for which shadow was not possible, however, bear in mind that this will systematically reduce the difficulty of the shadow class compared to real-world use, as these contain the most difficult shadow examples.</p> <p>In addition to the 20m sampled subscenes and masks, we also provide users with shapefiles that define the boundary of the mask on the original Sentinel-2 scene. If users wish to retrieve the L1C bands at their original resolutions, they can use these to do so.</p> <p>Please see the README for further details on the dataset structure&nbsp;and more.</p> <p>&nbsp;</p> <p><strong>Contributions &amp; Acknowledgements</strong></p> <p>The data were collected, annotated, checked, formatted and published by Alistair Francis and John Mrziglod.</p> <p>Support and advice was provided by Prof. Jan-Peter Muller and Dr. Panagiotis Sidiropoulos, for which we are grateful.</p> <p>We would like to extend our thanks to Dr. Pierre-Philippe Mathieu and the rest of the team at <em>ESA PhiLab</em>, who provided the environment in which this project was conceived, and continued to give technical support throughout.</p> <p>Finally, we thank the <em>ESA Network of Resources</em> for sponsoring this project by providing ICT resources.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo44/100

Eddy Kinetic Energy and SST gradients global datasets and trends. Additionally, this dataset includes ocean basins and ocean processes masks.

<p>This dataset includes the post-processed data used for the paper titled &quot;Mesoscale kinetic energy response to changing oceans&quot;. The original data was obtained from AVISO+ SSH altimetry&nbsp;and NOAA optimal interpolated sea surface temperature (OISST):</p> <p>AVISO+ SSH:&nbsp;https://www.aviso.altimetry.fr/en/data/products/sea-surface-height-products/global/gridded-sea-level-heights-and-derived-variables.html</p> <p>NOAA-OISST:&nbsp;https://www.ncdc.noaa.gov/oisst</p> <p>From satellite observations of sea surface height (SSH) and sea surface temperature (SST) over the satellite record (1993 - 2019),&nbsp;EKE and SST gradients are derived.&nbsp;</p> <p>Then the fields are then temporally smoothed using a running average of 12 months. &nbsp;Trends and the&nbsp;significance of each field are finally computed with linear regression and a modified Mann&ndash;Kendall test (https://github.com/josuemtzmo/xarrayMannKendall).</p> <p>Geographical regions consist of the following ocean basins: the Southern Ocean, the Indian Ocean, the&nbsp;Pacific Ocean, and the Atlantic ocean. These ocean basins were expert-defined to capture ocean processes at all scales (ocean_basins_and_dynamical_masks.nc).</p> <p>Dynamical regions (Fig. 5d): the Antarctic Circumpolar Current (ACC), the boundary currents and their extensions, the tropics, the subtropical ocean gyres, and&nbsp;the remaining regions (ocean_basins_and_dynamical_masks.nc).</p> <p>Further information and scripts to reproduce the result of the manuscript can be found at:&nbsp;https://github.com/josuemtzmo/EKE_SST_trends</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Nonlinearity corrections and bad pixel masks for the WINTER sensors

<p>Nonlinearity corrections and bad pixel masks for the WINTER sensors to be used with https://github.com/winter-telescope/winternlc.&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Social Media Mask Dataset

<p>The Social Media Mask Dataset is a dataset made up of Twitter Images intended for the training of&nbsp;Convolutional Neural Networks to detect masks in images and video. In this case, the term masks refer to a device worn on the face intended to reduce the spread of respiratory illness. Due to Twitter&#39;s TOS, we cannot directly publish Twitter images. Instead, we publish Tweet keys along with information our script uses to download the target image. The file training.json is our suggested training set&nbsp;while testing.json is our suggested test set. Our script download_dataset.py can be used to download the full dataset with a Twitter developer account.&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Young people's media use and adherence to preventive measures in the "infodemic": Is it masked by political ideology?

<p>Data to replicate the publication &quot;Young people&#39;s media use and adherence to preventive measures in the &ldquo;infodemic&rdquo;: Is it masked by political ideology?&quot;. This publication examines the role of political ideology and political extremism for COVID-19 information seeking and preventive behaviour with data of the COVIDisc project. COVIDisc investigates how young people aged 15 to 34 years perceive the discussion in the Coronavirus Pandemic, which messages reach them, what media they use to inform themselves and how they experience the situation. en</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Snow-Cloud Validation Masks for Multispectral Satellite Data.

<p>Geotiffs of manually validated&nbsp;snow, cloud, &amp; clear-sky snow free pixels for&nbsp;13 Landsat 8 images. These acquisitions are of mid-latitude mountainous regions that contain both snow and cloud cover.&nbsp;&nbsp;Four spectral libraries of snow and cloud are also provided. These are the snow and cloud spectra extracted from both these 13 scenes and the 13 L8 SPARCS Cloud Validation Masks that contained both snow and cloud.&nbsp;1&amp;2.) Snow and cloud top-of-atmosphere reflectance for the eight Landsat 8 OLI 30 meter optical bands,&nbsp;aggregated from the 26&nbsp;scenes. 3&amp;4.) The top-of-atmosphere reflectance for the eight Landsat 8 OLI 30 meter optical bands of all snow misidentified as cloud and cloud misidentified as snow by CFMASK, the cloud mask that ships in the BQA file of Landsat 8 Collection 1.</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

Flood Masks Doñana 1984/2019

<p>Time Series of flooded areas derived from Landsat TM, ETM+ &amp; OLI in the Path 202 Row 34 (Do&ntilde;ana). Also, these products and its metadata are freely available to consult or downloaded in the LAST-EBD Cartography Server: http://mercurio.ebd.csic.es/imgs/</p> <p>Methodology is described in this paper: Remote Sensing 8(9):775 &middot; September 2016. DOI: 10.3390/rs8090775</p>

opencc-by-4.0Oct 2019View details →
zenodo44/100

Water Turbidity Masks Doñana 1984/2019

<p>Time Series water turbidity derived from Landsat TM, ETM+ &amp; OLI in the Path 202 Row 34 (Do&ntilde;ana). Also, the product and its metadata are freely available to consult or downloaded in the LAST-EBD Cartography Server:&nbsp;http://mercurio.ebd.csic.es/imgs/</p> <p>Teh methodology is described in the paper:&nbsp;Empirical models to estimate water turbidity from reflectance data from TM or ETM+ Landsat sensors in shallow wetlands such as Do&ntilde;ana marshes. See the reference: Bustamante, J. et al. 2009. Predictive models of turbidity and water depth in the Do&ntilde;ana marshes using Landsat TM and ETM+ images. Journal of Environmental Management. 90:2219-2225.https://doi.org/10.1016/j.jenvman.2007.08.021.</p>

opencc-by-4.0Oct 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record