Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,943

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,943 results for “machine learning”

Learn how ShareScore rates datasets ↗
zenodo36/100

Mapping discrete forest age classes of Mediterranean pinelands since the pre-satellite era using historical orthoimage mosaics and machine learning.

<p>Data and sripts for Journal of Forestry Research submitted manuscript. Authors: Vicent Agust&iacute; Ribas Costa, Andrew Trlica, &amp; Aitor Gast&oacute;n Gonz&aacute;lez.</p> <p>This data and scripts are part of Vicent's PhD project at Universidad Polit&eacute;cnica de Madrid, supervised by Dr. Aitor. Contact Vicent for any queries at va.ribas@upm.es.</p> <p>The PNOA LiDAR and 1956 and 2021 orthophotos are openly available at the Centro de Descargas of the Instituto Geogr&aacute;fico Nacional (<a href="https://centrodedescargas.cnig.es/CentroDescargas/index.jsp">https://centrodedescargas.cnig.es/CentroDescargas/index.jsp</a>). The 1989 orthophoto is available under request at the Institut Cartogr&agrave;fic i Geogr&agrave;fic de les Illes Balears (<a href="https://www.caib.es/webgoib/institut-cartografic-i-geografic-de-les-illes-balears-icgib-">https://www.caib.es/webgoib/institut-cartografic-i-geografic-de-les-illes-balears-icgib-</a>).</p> <p>Forest inventory data is property of the landowners and managed by Terrapi World Ltd.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations: Preprocessed Satellite and In-situ observation datasets

<p>This record includes all of the prepared data used in the manuscript, &quot;Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations&quot; (citation information forthcoming). As a part of this manuscript, we analyzed the ability for machine learning models to extract sea&nbsp;surface information (from salinity, temperature, sea height anomaly) to predict mixed layer depth. In this manuscript there are two experimental datasets: (1) info derived from CESM POP2 ocean model dataset (1989-1998), and (2) info derived from a combination of satellite sources and MLD from Argo profiles. More details below.&nbsp;</p> <p>All of these data files are preprocessed and organized to be used with the ml-ocean-bl github code found at&nbsp;https://github.com/NCAR/ml-ocean-bl/mloceanbl/.</p> <ul> <li><strong>CESM POP2 Ocean model dataset</strong></li> </ul> <p>Preprocessed sea surface salinity (SSS), temperature (SST), sea surface height anomalies (SSH), and ocean mixed layer depth (MLD, or HMXL) derived from the CESM POP2 Ocean model. Specifically,&nbsp;CESM POP2 model in a hindcast forced by JRA55do atmospheric reanalysis from 1958 to present and initialized with an oceanic climatology as in e.g. <a href="https://journals.ametsoc.org/view/journals/phoc/aop/JPO-D-20-0217.1/JPO-D-20-0217.1.xml">Deppenmeier et al. (2021)</a>. The model outputs include the ocean mixed layer depth (MLD), sea surface salinity (SSS), sea surface temperature (SST), and sea height anomaly (SSH) at a temporal frequency of 5-days and an approximate latitude and longitude resolution of 0.1 degrees.</p> <p>Relevant files:</p> <ol> <li>full_EPO.nc, full_SIO.nc <ul> <li>NetCDF4 containing SSS, SST, SSH, MLD for the equatorial Pacific Ocean (EPO) and southern Indian Ocean (SIO) (see manuscript for details). Data is regridded onto a 1/2 degree lat/lon 5 day grid to correspond with data used for Argo datasets (see below).</li> </ul> </li> <li>clim_EPO.nc, clim_SIO.nc, clim_std_EPO.nc, std_clim_EPO.nc, std_clim_SIO.nc <ul> <li>NetCDF4 containing mean and standard deviation climatologies of SSS, SST, SSH, and MLD for EPO and SIO.</li> </ul> </li> <li>std_anomalies_EPO.nc, std_anomalies_SIO.nc <ul> <li>NetCDF4 containing SSS, SST, SSH, and MLD standardized anomalies for EPO and SIO. This is the dataset directly used for training in aforementioned manuscript. Use with&nbsp;ml-ocean-bl/ml-ocean-test/data.&nbsp;</li> </ul> </li> </ol> <ul> <li><strong>Satellite and Argo datasets</strong></li> </ul> <p>Preprocessed satellite sea surface salinity (SSS), temperature (SST), and sea surface height anomalies (SSH) and Argo-based mixed layer depth (MLD) profiles. Original data can be found at:</p> <p>(SST):&nbsp;Remote Sensing Systems. 2017. MW optimum interpolated SST data set. Ver. 5.0. PO.DAAC, CA, USA.&nbsp; Further information available at at&nbsp;<a href="https://doi.org/10.5067/GHMWO-4FR05">https://doi.org/10.5067/GHMWO-4FR05</a>. Data can be accessed at&nbsp;https://podaac-tools.jpl.nasa.gov/drive/files/allData/ghrsst/data/GDS2/L4/GLOB/REMSS/mw_OI/v5.0/.</p> <p>(SSS):&nbsp;Oleg Melnichenko. 2018. Aquarius L4 Optimally Interpolated Sea Surface Salinity. Ver. 5.0. PO.DAAC, CA, USA. Further information at <a href="https://doi.org/10.5067/AQR50-4U7CS">https://doi.org/10.5067/AQR50-4U7CS</a>. Data can be accessed at&nbsp;https://podaac-tools.jpl.nasa.gov/drive/files/SalinityDensity/aquarius/L4/IPRC/v5/7day.&nbsp;</p> <p>(SSH):&nbsp;Zlotnicki, Victor; Qu, Zheng; Willis, Joshua. 2019. SEA_SURFACE_HEIGHT_ALT_GRIDS_L4_2SATS_5DAY_6THDEG_V_JPL1609. Ver. 1812. PO.DAAC, CA, USA. Information available at&nbsp;<a href="https://doi.org/10.5067/SLREF-CDRV2">https://doi.org/10.5067/SLREF-CDRV2</a>. Data can be accessed at&nbsp;https://podaac-tools.jpl.nasa.gov/drive/files/SeaSurfaceTopography/merged_alt/L4/cdr_grid</p> <p>(MLD)&nbsp;Argo-based ocean surface mixed layer depths using the buoyancy gradient definition of Whitt Nicholson and Carranza (2019) processed dataset available at https://doi.org/10.5281/zenodo.4291175.</p> <p>Relevant files:</p> <ol> <li>https://github.com/NCAR/ml-ocean-bl/mloceanbl/preprocess_mld.py and .../preprocess_sss_sst_ssh.py. <ul> <li>Preprocessing code</li> </ul> </li> <li>sss_sst_ssh_anomalies.nc. <ul> <li>Regridded and resampled SSS, SST, SSH onto a 1/2 degree lat/lon 7day grid. Contains preprocessed seasonal data along with anomalies.</li> </ul> </li> <li>&nbsp;mldb_climatology_climatologystd_binned.nc <ul> <li>Smoothed argo-based mixed layer depths are used to calculate climatologies and standardized climatologies. 4 degree lat/lon gridded&nbsp;climatologies.</li> </ul> </li> <li>mldb_full_anomalies_stdanomalies_climatology_stdclimatology.nc <ul> <li>Contains the Argo profile-derived&nbsp;MLD, anomalies, standard anomalies, climatologies, and standardized climatologies with corresponding argo locations, times, and corresponding weeks.&nbsp;</li> </ul> </li> <li>equatorial_pacific_model_oi_re.nc,&nbsp;&nbsp;southern_indian_model_oi_re.nc <ul> <li>Model outputs for the Equatorial Pacific Ocean and Southern Indian Ocean. These gridded files contain the model outputs (vlcnn, vlcnn variance, OI, OI&nbsp;variance, reanalysis, and reanalysis variance - see manuscript for nomenclature details) at each of the 200 weeks available. It should be noted that, in the equatorial Pacific Ocean, the lat/lon location of (-138.75,&nbsp;-9.75) is masked during the training and filled with a NaN in the .nc files.&nbsp;</li> </ul> </li> </ol> <p>&nbsp;</p> <p>&nbsp;</p> <p>Contact D. Foster with any questions.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Datatset: Machine-Learning Side-Channel Attacks on the GALACTICS Constant-Time Implementation of BLISS

<p>This dataset accompanies the paper &quot;Machine-Learning Side-Channel Attacks on the GALACTICS Constant-Time Implementation of BLISS&quot;. It was used to experimentally prove the presented attack strategies on real hardware. The corresponding source code for all three attacks is also publicly available.</p> <p>A detailed description of how the data was obtained can be found in the paper. Section 4 precisely describes the experimental setup.</p> <p>&nbsp;</p> <p>Prerequisites:</p> <pre><code class="language-bash">sudo apt-get install p7zip</code></pre> <p>&nbsp;</p> <p>Extract the data:</p> <pre><code class="language-bash">7z x galactics_attack_data.7z</code></pre> <p>&nbsp;</p> <p>Running the attacks:</p> <p>The source code to run the three presented attacks can be found on Github. The instructions on how to use the python code can be obtained from the corresponding README.</p> <p>&nbsp;</p> <p>Re-using the dataset:</p> <p>The dataset consists of <em>.pickle</em> and <em>.bin</em> files. The <em>.pickle</em> files can be read using <a href="https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_pickle.html">Pythons Pandas library</a>. Python access functions for the <em>.bin</em> files are also provided.</p>

opencc-by-4.0Jul 2021View details →
dryad36/100

Using machine learning to model nontraditional spatial dependence in occupancy data

<p>Spatial models for occupancy data are used to estimate and map the true presence of a species, which may depend on biotic and abiotic factors as well as spatial autocorrelation. Traditionally researchers have accounted for spatial autocorrelation in occupancy data by using a correlated normally distributed site-level random effect, which might be incapable of modeling nontraditional spatial dependence such as discontinuities and abrupt transitions. Machine learning approaches have the potential to model nontraditional spatial dependence, but these approaches do not account for observer errors such as false absences. By combining the flexibility of Bayesian hierarchal modeling and machine learning approaches, we present a general framework to model occupancy data that accounts for both traditional and nontraditional spatial dependence as well as false absences. We demonstrate our framework using six synthetic occupancy data sets and two real data sets. Our results demonstrate how to model both traditional and nontraditional spatial dependence in occupancy data which enables a broader class of spatial occupancy models that can be used to improve predictive accuracy and model adequacy.</p>

opencc-zeroJul 2021View details →
zenodo36/100

Life beneath the ice: jellyfish and ctenophores from the Ross Sea, Antarctica, with an image-based training set for machine learning

<p>This Zenodo dataset contain the Common Objects in Context (COCO) files linked to the following publication:</p> <p>Verhaegen, G, Cimoli, E, &amp; Lindsay, D (2021). Life beneath the ice: jellyfish and ctenophores from the Ross Sea, Antarctica, with an image-based training set for machine learning. Biodiversity Data Journal.</p> <p>Each COCO zip folder contains an &quot;annotations&quot; folder including a json file and an &quot;images&quot; folder containing the annotated images.</p> <p>Details on each COCO zip folders:</p> <ul> <li><strong>Beroe_sp_A_images-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Beroe&nbsp;</em>sp. A<em> </em>for the following 114 images:</p> <p>MCMEC2018_20181116_NIKON_Beroe_sp_A_c_1 to MCMEC2018_20181116_NIKON_Beroe_sp_A_c_16, MCMEC2018_20181125_NIKON_Beroe_sp_A_d_1 to MCMEC2018_20181125_NIKON_Beroe_sp_A_d_57, MCMEC2018_20181127_NIKON_Beroe_sp_A_e_1 to MCMEC2018_20181127_NIKON_Beroe_sp_A_e_2, MCMEC2019_20191116_SONY_Beroe_sp_A_a_1 to MCMEC2019_20191116_SONY_Beroe_sp_A_a_28, and MCMEC2019_20191127_SONY_Beroe_sp_A_f_1 to MCMEC2019_20191127_SONY_Beroe_sp_A_f_12</p> <ul> <li><strong>Beroe_sp_B_images-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Beroe&nbsp;</em>sp. B<em> </em>for the following 2 images:</p> <p>MCMEC2019_20191115_SONY_Beroe_sp_B_a_1 and MCMEC2019_20191115_SONY_Beroe_sp_B_a_2</p> <ul> <li><strong>Callianira_cristata_images-coco 1.0.zip </strong></li> </ul> <p>COCO annotations of <em>Callianira cristata</em><em> </em>for the following 21 images:</p> <p>MCMEC2019_20191120_SONY_Callianira_cristata_b_1 to MCMEC2019_20191120_SONY_Callianira_cristata_b_21</p> <ul> <li><strong>Diplulmaris_antarctica_images-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Diplulmaris antarctica</em> for the following 83 images:</p> <p>MCMEC2019_20191116_SONY_Diplulmaris_antarctica_a_1 to MCMEC2019_20191116_SONY_Diplulmaris_antarctica_a_9, and MCMEC2019_20191201_SONY_Diplulmaris_antarctica_c_1 to MCMEC2019_20191201_SONY_Diplulmaris_antarctica_c_74</p> <ul> <li><strong>Koellikerina_maasi_images-coco 1.0.zip </strong></li> </ul> <p>COCO annotations of <em>Koellikerina maasi</em><em> </em>for the following 49 images:</p> <p>MCMEC2018_20181127_NIKON_Koellikerina_maasi_b_1 to MCMEC2018_20181127_NIKON_Koellikerina_maasi_b_4, MCMEC2018_20181129_NIKON_Koellikerina_maasi_c_1 to MCMEC2018_20181129_NIKON_Koellikerina_maasi_c_29, and MCMEC2019_20191126_SONY_Koellikerina_maasi_a_1 to MCMEC2019_20191126_SONY_Koellikerina_maasi_a_16</p> <ul> <li><strong>Leptomedusa_sp_A-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of Leptomedusa sp. A<em> </em>for Figure 5 (see paper).</p> <ul> <li><strong>Leuckartiara_brownei_images-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Leuckartiara brownei</em> for the following 48 images:&nbsp;</p> <p>MCMEC2018_20181129_NIKON_Leuckartiara_brownei_b_1 to MCMEC2018_20181129_NIKON_Leuckartiara_brownei_b_27, MCMEC2018_20181129_NIKON_Leuckartiara_brownei_c_1 to MCMEC2018_20181129_NIKON_Leuckartiara_brownei_c_6, and MCMEC2019_20191116_SONY_Leuckartiara_brownei_a_1 to MCMEC2019_20191116_SONY_Leuckartiara_brownei_a_15</p> <ul> <li><strong>MCMEC2019_20191115_SONY_Mertensiidae_sp_A_a_3-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of Mertensiidae sp. A for the following video (total of&nbsp;1847 frames): MCMEC2019_20191115_SONY_Mertensiidae_sp_A_a_3 (<a href="https://youtu.be/0W2HHLW71Pw">https://youtu.be/0W2HHLW71Pw</a>)</p> <ul> <li><strong>MCMEC2019_20191116_SONY_Leuckartiara_brownei_a_3-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Leuckartiara brownei</em> for the following video (total of&nbsp;1367 frames): MCMEC2019_20191116_SONY_Leuckartiara_brownei_a_3 (<a href="https://youtu.be/dEIbVYlF_TQ">https://youtu.be/dEIbVYlF_TQ</a>)</p> <ul> <li><strong>MCMEC2019_20191122_SONY_Callianira_cristata_a_1-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Callianira cristata</em> for the following video (total of 2423 frames): MCMEC2019_20191122_SONY_Callianira_cristata_a_1 (<a href="https://youtu.be/30g9CvYh5JE">https://youtu.be/30g9CvYh5JE</a>)</p> <ul> <li><strong>MCMEC2019_20191122_SONY_Leptomedusa_sp_B_a_1-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of Leptomedusa sp. B for the following video (total of 1164 frames): MCMEC2019_20191122_SONY_Leptomedusa_sp_B_a_1 (<a href="https://youtu.be/hrufuPQ7F8U">https://youtu.be/hrufuPQ7F8U</a>)</p> <ul> <li><strong>MCMEC2019_20191126_SONY_Koellikerina_maasi_a_1-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Koellikerina maasi</em> for the following video (total of&nbsp;1643 frames): MCMEC2019_20191126_SONY_Koellikerina_maasi_a_1 (<a href="https://youtu.be/QiBPf_HYrQ8">https://youtu.be/QiBPf_HYrQ8</a>)</p> <ul> <li><strong>MCMEC2019_20191129_SONY_Mertensiidae_sp_A_b_1-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of Mertensiidae<em> </em>sp. A for the following video (total of&nbsp;239 frames): MCMEC2019_20191129_SONY_Mertensiidae_sp_A_b_1 (<a href="https://youtu.be/pvXYlQGZIVg">https://youtu.be/pvXYlQGZIVg</a>)</p> <ul> <li><strong>MCMEC2019_20191129_SONY_Pyrostephos_vanhoeffeni_b_2-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Pyrostephos vanhoeffeni</em> for the following video (total of 444 frames): MCMEC2019_20191129_SONY_Pyrostephos_vanhoeffeni_b_2 (<a href="https://youtu.be/2rrQCybEg0Q">https://youtu.be/2rrQCybEg0Q</a>)</p> <ul> <li><strong>MCMEC2019_20191129_SONY_Pyrostephos_vanhoeffeni_b_3-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Pyrostephos vanhoeffeni</em> for the following video (total of&nbsp;683 frames): MCMEC2019_20191129_SONY_Pyrostephos_vanhoeffeni_b_3 (<a href="https://youtu.be/G9tev_gdUvQ">https://youtu.be/G9tev_gdUvQ</a>)</p> <ul> <li><strong>MCMEC2019_20191129_SONY_Pyrostephos_vanhoeffeni_b_4-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Pyrostephos vanhoeffeni</em> for the following video (total of&nbsp;1127 frames): MCMEC2019_20191129_SONY_Pyrostephos_vanhoeffeni_b_4<strong> </strong>(<a href="https://youtu.be/NfJjKBRh5Hs">https://youtu.be/NfJjKBRh5Hs</a>)</p> <ul> <li><strong>MCMEC2019_20191130_SONY_Beroe_sp_A_b_1-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Beroe </em>sp. A for the following video (total of&nbsp;2171 frames): MCMEC2019_20191130_SONY_Beroe_sp_A_b_1<strong> </strong>(<a href="https://youtu.be/kGBUQ7ZtH9U">https://youtu.be/kGBUQ7ZtH9U</a>)</p> <ul> <li><strong>MCMEC2019_20191130_SONY_Beroe_sp_A_b_2-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Beroe </em>sp. A for the following video (total of&nbsp;359 frames): MCMEC2019_20191130_SONY_Beroe_sp_A_b_2 (<a href="https://youtu.be/Vbl_KEmPNmU">https://youtu.be/Vbl_KEmPNmU</a>)</p> <ul> <li><strong>Mertensiidae_sp_A_images-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of Mertensiidae sp. A for the following 49 images:</p> <p>MCMEC2018_20181127_NIKON_Mertensiidae_sp_A_c_1 to MCMEC2018_20181127_NIKON_Mertensiidae_sp_A_c_2, MCMEC2018_20181127_NIKON_Mertensiidae_sp_A_f_1 to MCMEC2018_20181127_NIKON_Mertensiidae_sp_A_f_8, MCMEC2018_20181129_NIKON_Mertensiidae_sp_A_d_1 to MCMEC2018_20181129_NIKON_Mertensiidae_sp_A_d_13, MCMEC2018_20181201_ROV_Mertensiidae_sp_A_e_1 to MCMEC2018_20181201_ROV_Mertensiidae_sp_A_e_15, and MCMEC2019_20191115_SONY_Mertensiidae_sp_A_a_1 to MCMEC2019_20191115_SONY_Mertensiidae_sp_A_a_11</p> <ul> <li><strong>Pyrostephos_vanhoeffeni_images-coco 1.0.zip</strong></li> </ul> <p>COCO annotations of <em>Pyrostephos vanhoeffeni</em>&nbsp; for the following 14 images: MCMEC2019_20191125_SONY_Pyrostephos_vanhoeffeni_a_1 to MCMEC2019_20191125_SONY_Pyrostephos_vanhoeffeni_a_8, MCMEC2019_20191129_SONY_Pyrostephos_vanhoeffeni_b_1 to MCMEC2019_20191129_SONY_Pyrostephos_vanhoeffeni_b_6</p> <ul> <li><strong>Solmundella_bitentaculata_images-coco 1.0.zip </strong></li> </ul> <p>COCO annotations of <em>Solmundella bitentaculata</em> for the following 13 images: MCMEC2018_20181127_NIKON_Solmundella_bitentaculata_a_1 to MCMEC2018_20181127_NIKON_Solmundella_bitentaculata_a_13</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Dataset for the publication of "WRF model parameter calibration to improve the prediction of tropicalcyclones over the Bay of Bengal using Machine Learning-basedMultiobjective Optimization"

<p>The dataset consists of the modified WRF model software, that can be extracted and used in any Linux system with preinstalled required software.</p> <p>The namelists_file.zip consists of the namelist.input files that are used for the default and calibration simulations with different driving data namely, FNL files at 1deg with two nested domains, ERA files at 1deg with two nested domains, ERA files at 0.25deg with a single domain, and the ERA files at 0.25deg with two nested domains.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Dynamic Evolution of Changbaishan Volcanism in Northeast China Illuminated by Machine Learning

<p><strong>This dataset is for our work entitled &quot;<em>Dynamic Evolution of Changbaishan Volcanism in Northeast China Illuminated by Machine Learning</em>&quot;.</strong></p> <p><strong>This dataset includes: 1)&nbsp;The Cenozoic basalts in Changbaishan area used for predicting, 2)&nbsp;The&nbsp;IAB (Island Arc Basalts) and OIB (Ocean Island Basalts) samples used for training machine learning models.</strong></p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Investigating automated bird detection from webcams using machine learning

<p>We provide a dataset of images(.jpeg) with their corresponding annotations files(.xml) used to train a bird detection deep learning model. These images were collected from the live stream feeds of&nbsp; Cornell Lab of Ornithology&nbsp;(https://www.allaboutbirds.org/cams/) situated in 6 unique locations around the world as follows:</p> <ul> <li>Treman bird feeding garden at the Cornell Ornithology Laboratory in Ithaca, New York. At this station, Axis P11448-LE cameras are used to capture the recordings from feeders perched on the edge of both Sapsucker Woods and its 10-acre ponds. This site mainly attracts forest species like chickadees (Poecile atricapillus), red-winged blackbirds (Agelaius phoeniceus), and woodpeckers (Picidae). A total of 2065 images were captured from this location.</li> <li>&nbsp;Fort Davis in Western Texas, USA. At this site, a total of 30&nbsp; hummingbird feeder cams are hosted at an elevation of over 5500 feet. From this site, 1440 images were captured.</li> <li>Sachatamia Lodge in Mindo, Ecuador. This site has a live hummingbird feed watcher that attracts over 132 species of hummingbirds including: Fawn-breasted Brilliant, White-necked Jacobin, Purple-bibbed Whitetip, Violet-tailed Sylph, Velvet-purple Coronet, and many others. A total of 2063 images were captured from this location.</li> <li>Morris County, New Jersey, USA. Feeders at this location attract over 39 species including Red-bellied Woodpecker, Red-winged Blackbird, Purple Finch, Blue Jay, Pine Siskin, Hairy Woodpecker, and others. Footage at this site is captured by an Axis P1448-LE Camera and Axis T8351 Microphone. A total of 1876 images were recorded from this site.</li> <li>Canopy Lodge in El Valle de Anton, Panama. Over 158 bird species visit this location annually and these include Gray-headed Chachalaca, Ruddy Ground-Dove, White-tipped Dove, Green Hermit, and others. A total of 1600 images were captured.</li> <li>Southeast tip of South Island, New Zealand. At this site, nearly 10000 seabirds visit this location annually and a total of 1548 images were captured.</li> </ul> <p>&nbsp;The Cornell Lab of Ornithology is an institute dedicated to biodiversity conversation with the main focus on birds through research, citizen science, and education. The autoscreen software was used to capture the images from the live feeds and images of approximately 1 Megapixel (Joint Photographic Experts Group) JPEG-coloured images of resolution 1366 X 768 X 3 pixels were collected (https://sourceforge.net/projects/autoscreen/). The software took a new image every 30 seconds and was captured during different times of the day in order to avoid a sample-biased dataset. In total, 10592 images were collected for this study.</p> <p><strong>Files provided</strong></p> <p>Train.zip &ndash; contains 6779 image files(.jpeg) and 6779 annotation files (.xml)</p> <p>Validation.zip &ndash; contains 1695 image files(.jpeg) and 1695 annotation files (.xml)</p> <p>Test.zip &ndash;contains 2118 image files(.jpeg)</p> <p>Scripts.zip - Contains scripts needed in manipulating the dataset like dataset partitioning, and creation of CSV and tfrecords files.&nbsp;</p> <p>This dataset was used in the MSc thesis titled &ldquo;Investigating automated bird detection from webcams using machine learning&rdquo; by Alex Mirugwe, University of Cape Town &ndash; South Africa.</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Thermally Averaged Magnetic Anisotropy Tensors via Machine Learning Based on Gaussian Moments

<ul> <li><strong>D-Tensor Data</strong></li> </ul> <p>The reference data for 3500 configurations of&nbsp;[Co(N<sub>2</sub>S<sub>2</sub>O<sub>4</sub>C<sub>8</sub>H<sub>10</sub>)<sub>2</sub>]<sup>2&minus;</sup> (CoSar),&nbsp;[Fe(tpa)<sup>Ph</sup>]<sup>&minus;</sup> (FeTPAPh), and&nbsp;[Ni(HIM<sub>2</sub>&minus;py)<sub>2</sub>NO<sub>3</sub>]<sup>+</sup> (NiComplex) is generated employing Molpro package [1]. For more details see the original publication [2]. The data is stored in python compressed array format (.npz) with the D-Tensor in cm<sup>-1</sup>. The data set contains four <span class="math-tex">\(np.ndarray\)</span></p> <pre><code class="language-python">import numpy as np data = np.load('CoSar.npz') R = data['R'] # Cartesian coordinates of nuclei in Ang. D = data['MAT'] # D-Tensor values in cm-1, D = (D11, D12, D13, D22, D23, D33) N = data['N'] # Number of atoms in each structure Z = data['Z'] # Nuclear charges</code></pre> <ul> <li><strong>AIMD Data</strong></li> </ul> <p>To propagate the periodic cell containing four CoSar molecules, for which D-Tensor was computed above, a data set containing total energies as well as atomic forces of 3500 structures was generated employing VASP package [3-6]. For more details see the original publication [2]. The data is stored in python compressed array format (.npz) with the total energy in kcal/mol and atomic forces in kcal/mol/Ang. The data set contains six <span class="math-tex">\(np.ndarray\)</span></p> <pre><code class="language-python">import numpy as np data = np.load('CoSar_bulk.npz') R = data['R'] # Cartesian coordinates of nuclei in Ang. C = data['C'] # Cell vectors in Ang. E = data['E'] # Total energy in kcal/mol F = data['F'] # Atomic forces in kcal/mol/Ang. N = data['N'] # Number of atoms in each structure Z = data['Z'] # Nuclear charges</code></pre> <ul> <li><strong>References</strong></li> </ul> <p>[1] H.-J. Werner, P. J. Knowles, G. Knizia, F. R. Manby, M. Sch&uuml;tz,et al.,&ldquo;Molpro, version 2020.0, a package of ab initio programs,&rdquo; (2020), see https://www.molpro.net.</p> <p>[2] V. Zaverkin, J. Netz, F. Zills, A. K&ouml;hn, and J. K&auml;stner, &ldquo;Thermally Averaged Magnetic Anisotropy Tensors via Machine Learning Based on Gaussian Moments,&rdquo;<strong> submitted</strong> (2021).</p> <p>[3]&nbsp;P. E. Bl&ouml;chl, &ldquo;Projector augmented-wave method,&rdquo; Phys. Rev. B 50, 17953 (1994).</p> <p>[4] G. Kresse and J. Hafner, &ldquo;Ab initio molecular dynamics for liquid metals,&rdquo; Phys. Rev. B 47, 558 (1993).</p> <p>[5] G. Kresse and J. Furthm&uuml;ller, &ldquo;Efficiency of ab-initio total energy calculations for metals and semiconductors using a plane-wave basis set,&rdquo; Comput. Mater. Sci. 6, 15 &ndash; 50 (1996).</p> <p>[6] G. Kresse and J. Furthm&uuml;ller, &ldquo;Efficient iterative schemes for ab initio total-energy calculations using a plane-wave basis set,&rdquo; Phys. Rev. B 54, 11169 (1996).</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Machine learning-based evidence and attribution mapping of 100,000 climate impact studies - Data

<p>Data for the paper&nbsp;Machine learning-based evidence and attribution mapping of 100,000 climate&nbsp;impact studies</p> <p><strong>Document Metadata</strong></p> <p>0c_doc_info.csv contains basic document metadata for each document considered in our study</p> <p><strong>Predictions</strong></p> <p>In each predictions file, 1 refers to a document hand-labelled as belonging to a category, and 0 refers to a document hand-labelled as not belonging to a category. All values in between are predicted values, where for values greater than 0.5, a document is considered likely to belong to the given category.</p> <p>1_document_relevance.csv contains the predicted relevance of a document to the study.</p> <p>1_driver_predictions.csv contains the predicted climate driver of each document.</p> <p>1_impact_predictions.csv contains the predicted impact type of each document</p> <p><strong>Geographical data</strong></p> <p>Place_df.csv contains a row for each geographical entity automatically extracted from each study</p> <p>Study_gridcell_2.5.csv contains a row matching each study with each grid cell covered by the study&rsquo;s smallest mentioned geographical entity</p> <p><strong>Merged data</strong></p> <p>2_study_da.csv contains a row for each study describing the aggregated detection and attribution characteristics of the grid cells the study refers to</p> <p>2_merged_da_data.csv contains a row for each grid cell describing the attribution categories and the number of weighted grid cells for each climate driver.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Supporting Information 1 to the paper "Human-machine-learning integration and task allocation in citizen science".

<p>This appendix - Supporting Information 1 - is a dataset excel file&nbsp;directly related to the following paper:</p> <p>Ponti, M., Seredko, A. <a href="http://doi.org/10.1057/s41599-022-01049-z">Human-machine-learning integration and task allocation in citizen science.</a>&nbsp;<em>Humanit Soc Sci Commun</em>&nbsp;<strong>9,&nbsp;</strong>48 (2022). https://doi.org/10.1057/s41599-022-01049-z</p> <p>The dataset in this excel file is a detailed result of the integrative literature review conducted for the manuscript.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Supplemental data for "Large-scale quantum machine learning"

<p>This data supports&nbsp;&quot;Large-scale quantum machine learning&quot; by Tobias Haug, Chris N. Self, M. S. Kim (arxiv:2108.01039) https://arxiv.org/abs/2108.01039</p> <p>Related code can be found in the GitHub repository:&nbsp;(https://github.com/chris-n-self/large-scale-qml).&nbsp;The&nbsp;&#39;studies&#39; folder&nbsp;here can be dropped inside the code repository in order to run the analysis scripts.</p> <p>Both &#39;processed&#39; and &#39;unprocessed&#39; data is provided. Unprocessed data is the qiskit measurement results for each case study, executed on&nbsp;the IBM Quantum device <em>ibmq_guadalupe</em> and <em>ibmq_toronto</em>. Processed is the Gram matrix evaluated from the measurements&nbsp;and the data vectors needed to fit support vector machine classifiers.</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Consensus machine-learning models for protein-ligand binding affinity estimation

<p><strong>Motivation:</strong> In structure-based virtual screening, machine learning based scoring function gained popularity in the last few years as they outperformed classical scoring function. The protein-ligand system can be encoded by a set of orthogonal descriptor spaces, which are then mined by machine learning algorithms to find a relationship with the binding affinity experimental value.</p> <p><strong>&nbsp;</strong></p> <p><strong>Results:</strong> In this work we propose our modelling approach to derive a new scoring function, derived from a combination of multiple descriptor spaces coupled with machine learning algorithms ensembled in consensus. The SF has been trained on the PDBbind v.2019 data and has been extensively internally and externally validated on a large set of complexes. When benchmarked on the PDBbind core set, it achieved better performance than state-of-the-art counterparts, scoring: R<sub>Pearson </sub>= 0.85-0.86 r<sup>2</sup> = 0.70-0.72 and RMSE = 1.15-1.21. As highlights: (i) an applicability domain definition has been implemented to delimit the SF&rsquo;s application boundaries, and (ii) a mechanistic interpretation is proposed by investigating the contribution of each protein-ligand atom pairs in the prediction of the binding affinity, which could provide a support in the lead-optimization process.</p> <p><strong>&nbsp;</strong></p> <p><strong>Availability and implementation:</strong> Our scoring function is freely available through the webportal: <a href="https://predictor.exscalate.eu/">https://predictor.exscalate.eu/</a></p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

Dataset for "Comparative Analysis of Machine Learning Models to Forecast Flaring Capability of Solar Active Regions: A Parameter Based Approach"

<p>This CSV file contains the values of 14 selected magnetic features along with the active region class for all the regions used in our study. As discussed in the paper, these 14 magnetic features characterize the properties of active regions. All these magnetic features are obtained from the HMI SHARP data series which provides open-sourced vector magnetic field information of solar active regions. the column named &#39;AR_class&#39; carries information about the class of active regions i.e., 1 for flaring regions and 0 for non-flaring regions.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Supporting Information for "Machine Learning interpretation of the correlation between infrared emission features of interstellar polycyclic aromatic hydrocarbons"

<p>This is the supporting information for the article&nbsp;&quot;Machine Learning interpretation of the correlation between infrared emission features of interstellar polycyclic aromatic hydrocarbons&quot;. It contains:</p> <p>In&nbsp;Supporting_Information.pdf:</p> <p>1. A map of&nbsp;the correlation between descriptors.</p> <p>2. A spectral distribution.</p> <p>3. An explanation of the file&nbsp;example_code.zip.</p> <p>In&nbsp;example_code.zip:</p> <p>An example code of a machine learning model based on ECFP and corresponding input files.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Randomized Cooperative Overtake Maneuvers for Use in Machine Learning Algorithms

<p>Randomized&nbsp;cooperative overtake maneuvers involving one host vehicle and up to three remote vehicles. Maneuver containers are defined as specified in H&auml;fner et al., (2020) &quot;CVIP: A Protocol for Complex Interactions Among Connected Vehicles.&quot;</p> <p>If for a maneuver container &quot;rel-target-lane&quot; is given, then the maneuver type is a lane change left/right.</p> <p>If, instead, &quot;rel-target-speed&quot; is given, then the maneuver type is &quot;change speed&quot;.</p> <p>Within the raw data, the &quot;validate_stats&quot; files contain validation data as specified in the accompanying conference paper.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Anomaly Detection and Machine Learning

<p>The datasets were preprocessed. Correlated features were removed.</p> <p>Related papers:</p> <p><strong>[1]</strong>&nbsp;Iman Sharafaldin, Arash Habibi Lashkari, and Ali A. Ghorbani, &ldquo;Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization&rdquo;, 4th International Conference on Information Systems Security and Privacy (ICISSP), Portugal, January 2018</p> <p><strong>[2]</strong>&nbsp;Nour Moustafa, October 16, 2019, &quot;UNSW_NB15 dataset&quot;, IEEE Dataport, doi: https://dx.doi.org/10.21227/8vf7-s525.</p> <p><strong>[3]</strong> &ldquo;Sebastian Garcia, Agustin Parmisano, &amp; Maria Jose Erquiaga. (2020). IoT-23: A labeled dataset with malicious and benign IoT network traffic (Version 1.0.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.4743746&rdquo;</p> <p><strong>[4]</strong>&nbsp; A. D. Kent, &ldquo;Comprehensive, Multi-Source Cybersecurity Events,&rdquo; Los Alamos National Laboratory, http://dx.doi.org/10.17021/1179829, 2015.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Artificial Intelligence and COVID-19 using chest CT scan and chest X-ray images: Machine Learning and Deep Learning Approaches for Diagnosis and Treatment

<p>We uploaded the Table of included articles in the systematic&nbsp; review &quot;Artificial Intelligence and COVID-19 using chest CT scan and chest X-ray images: Machine Learning and Deep Learning Approaches for Diagnosis and Treatment&quot;</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Landslide mapping using satellite imagery and machine learning algorithms

<p>Cyclone Idai&nbsp;made landfall on 15th March near Beira, Mozambique, and&nbsp;caused heavy rainfall across Mozambique, Malawi, Madagascar, and eastern Zimbabwe. Chimanimani District of Zimbabwe received 200 to 400 mm rainfall between 15th and 19th March, which caused widespread flooding and triggered thousands of landslides. This study aims to map the landslides in Chimanimani District and differentiate concurrent flooding from the landslides using high resolution PlanetScope imagery and DEM. Three machine learning algorithms&nbsp;namely, Random Forest, Artificial Neural Network, and Support&nbsp;Vector Machine have been deployed for the supervised landslide classification.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Generating a Labeled Dataset to Train Machine Learning Algorithms for Lithological Classification of Drill Cuttings

<p>This dataset contains 16,700 fully labeled&nbsp;SEM&nbsp;images of rock chips isolated from&nbsp;14 thin sections of drill cutting samples.&nbsp;These samples come from a low-permeability reservoir in western Canada.</p>

opencc-by-4.0Sep 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record