Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,250

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,250 results for “classification”

Learn how ShareScore rates datasets ↗
zenodo44/100

Blaze Fire Classification – Segmentation Dataset

<p>The dataset is destined to be used for wildfire image classification and burnt area segmentation tasks for Unmanned Aerial Vehicles. It is comprised of 5,408 frames of aerial views taken from 56 videos and 2 public datasets. From the D-Fire public dataset, 829 photographs were used; and from the Burned Area UAV public dataset 34 images were used. For the classification task, there are 5 classes (&lsquo;Burnt&rsquo;, &lsquo;Half-Burnt&rsquo;, &rsquo;Non-Burnt&rsquo;, &lsquo;Fire&rsquo;, &lsquo;Smoke&rsquo;). As for the segmentation task, 404 segmentation masks on a subset have been created, which assign to each pixel of the image the class &lsquo;burnt&rsquo; or the class &lsquo;non-burnt&rsquo;.</p> <p>Details on acquiring the dataset can be found <strong><a href="https://aiia.csd.auth.gr/blaze-fire-classification-segmentation-dataset/" target="_blank" rel="noopener">here</a></strong>.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

CLDF dataset derived from Gerardi and Reichert's "The Tupí-Guaraní Language Family: A Phylogenetic Classification" from 2021

<p>Cite the source of the dataset as:</p> <blockquote> <p>Ferraz Gerardi, Fabrício and Reichert, Stanislav (2021) The Tupí-Guaraní Language Family: A Phylogenetic Classification. Diachronica 38(2). 151--188. DOI: https://doi.org/10.1075/dia.18032.fer.</p> </blockquote>

opencc-by-4.0Sep 2024View details →
zenodo44/100

VCM Dataset for the Classification of Resident Space Objects

<p>The Vector Covariance Message (VCM) data comprise 22,303 RSOs over a period of six months (9/1/2022-2/28/2023). VCM data consist of Resident Space Objects (RSOs) ephemerides from a high-precision special perturbations orbit propagator and estimator using tracking observations. VCMs are issued by the US Space Force (USSF) Space Command (USSPACECOM) and were provided through an Orbital Data Request (ODR) the authors submitted to the 18th Space Defense Squadron (18th SDS).&nbsp;</p> <p>The dataset is organized into subfolders, each containing VCMs for a specific satellite. Filenames correspond to the satellite's NORAD ID (North American Aerospace Defense Catalog Number). A readme file provides details about the VCM content and format. Note that the full covariance matrix has been excluded for public release, whereas the standard deviation of error in satellite's position and velocity is provided.</p> <p>The VCM data have been used in the following work, "Early Classification of Space Objects based on Astrometric Time Series Data", presented at the 25th Advanced Maui Optical and Space Surveillance Technologies Conference (AMOS) in Maui, Hawaii, United States.</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

PhasAGE Training School 1 -Overview of bioinformatics tools for the life sciences & Classification and evolution of non-globular proteins- LECTUREs

<p>The Training School 1&nbsp;<strong>&ldquo;Computational Methods to Study Protein Phase Separation&rdquo;</strong>&nbsp;is the first edition of a series of PhasAGE training activities.</p> <p>The goal of this course is to provide participants with the basic knowledge to understand the phenomenon of&nbsp;<strong>Phase Separation</strong>, its role in biological processes and diseases. In addition, the course will provide&nbsp;<strong>an overview of the available computational resources</strong>&nbsp;to navigate this knowledge. Participants will have&nbsp;<strong>hands-on training</strong>&nbsp;in tools and resources available for life sciences, to collect information from the literature on biomolecular phase transitions, identify features triggering phase transitions, mutations associated with diseases, known or predicted PTMs and molecular interaction sites.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Dataset and Code for Manuscript "Multi-angle pulse shape detection of scattered light in flow cytometry for label-free cell cycle classification"

<p>Dataset of measurements for cell cycle analysis with description:</p> <ul> <li>ReadMe file with explanations on the data set and analysis</li> <li>exemplary Matlab script file for analysis</li> <li>binary data files conatining the pulse shapes in all channels</li> <li>FCS data files containing common flow cytometry parameters in each channel</li> </ul> <p>Data on unsorted HEK cells, HEK cells sorted for cell cycle phases, and unsorted Jurkat cell are included.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Bioclimatic classifications for Colombia

<pre>Bioclimatic classifications seek to divide a study region into geographic areas with similar bioclimatic characteristics. In this study we proposed two bioclimatic classifications for Colombia using machine learning techniques. We initially obtained a modified Lang classification using dimensionality reduction and classification techniques. We then integrated the modified Lang and Caldas classifications to derive a second bioclimatic classification for Colombia.</pre>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Classification of Tragedies and Comedies in Calderón de la Barca's Comedias Nuevas

<p>Data publication accompanying the research article "<a href="https://doi.org/10.17175/2022_012">Classification of Tragedies and Comedies in Calder&oacute;n de la Barca&rsquo;s Comedias Nuevas</a>"</p> <p>The data publication contains the R code used for the analysis, the full text of all 112 Calderonian Comedias Nuevas (only the spoken text, no stage directions; zip-archive "Fulltexts"), the verbs, nouns, and adjectives of the 112 Comedias Nuevas (zip-archive "POS-texts"), a zip-archive containing 36 identified tragedies ("Tragedies"), a zip-archive containing 25 identified comedies ("Comedies"), and a csv file containing the classifications of all 112 dramas achieved with four clustering methods ("Annex-112-dramas-and-their-classificiation.csv").</p> <p>All 112 Calderonian "Comedias Nuevas" have been published as part of the Calder&oacute;n Drama Corpus within the Drama Corpora Project (DraCor), see <a href="https://dracor.org/cal">https://dracor.org/cal</a>.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

DeepAstroUDA: Semi-Supervised Universal Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection

<p>We present the data used in &quot;DeepAstroUDA: Semi-Supervised Universal Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection&quot;. It was also used in the&nbsp;conference paper presented in&nbsp;Machine Learning and the Physical Sciences workshop at&nbsp;NeurIPS&nbsp;2022:&nbsp;&quot;Semi-Supervised Domain Adaptation for Cross-Survey Galaxy Morphology Classification and Anomaly Detection&quot;.</p> <p>A plethora of AI methods, has already shown huge promise&nbsp;in increasing quality and speed of work with astronomical&nbsp;datasets, but high complexity&nbsp;of AI methods leads to extraction of dataset-specific non-robust features, which&nbsp;leads to models that cannot work on multiple datasets at the same time. We develop a Universal Domain Adaptation method <em><strong>DeepAstroUDA</strong></em>,&nbsp;capable of performing&nbsp;<strong>semi-supervised domain adaptation, that can be applied&nbsp;to datasets with different data distributions and class overlap</strong>. Extra classes&nbsp;can be present in any of the two datasets, and the method can even be used&nbsp;in the presence of unknown classes. We&nbsp;apply our model to three examples&nbsp;of galaxy morphology classification tasks of different complexities (3-class and&nbsp;10-class&nbsp;problems), with anomaly detection i.e.&nbsp;in all our experiments we have one extra class in the unlabeled target dataset, which represents our anomaly class.</p> <p>&nbsp;</p> <p><strong>DATA:</strong></p> <p><strong>1) DA across two different data releases of the same survey (LSST 1&nbsp;and 10 years of observation):</strong> We use data from Ciprijanovic et al. 2022. which&nbsp;can also be found&nbsp;on Zenodoo:&nbsp;<a href="https://zenodo.org/record/5514180#.Y6SM7y-B2_w">https://zenodo.org/record/5514180</a>&nbsp;. Data contains three classes: spiral (0), elliptical (1)&nbsp;and merging galaxies (3, anomaly class).</p> <p><strong>2) DA across two surveys (SDSS and DeCALS): </strong>We create datasets using data and labels from the Galaxy Zoo project. Datasets contain&nbsp;10 classes (9 known classes present in both SDSS and DeCALS data, and one unknown anomaly class present only in DeCALS data):&nbsp;disturbed&nbsp;(0), merging (1), round smooth (2), cigar shaped&nbsp;smooth (3), barred spiral (4), unbarred tight spiral (5),&nbsp;unbarred loose spiral (6), edge-on without bulge (7),&nbsp;edge-on with bulge (8), lenses (9, unknown anomaly class).</p> <p>SDSS (wide filed): datasets is split into two files &nbsp;-&nbsp;sdss_1.h5, sdss_2.h5</p> <p>DeCALS:&nbsp; decals.zip</p> <p><strong>3) DA between wide and&nbsp;deep observing fields of the same survey (SDSS):</strong> We create&nbsp;datasets using data and labels from the Galaxy Zoo project. Datasets contain same 10 classes as in 2), with the final lens anomaly class being only present in the SDSS deep field.</p> <p>SDSS (wide filed):&nbsp;the same data as in 2)</p> <p>SDSS (Strip 82 deep field):&nbsp;sdss_stripe82.zip</p> <p>All SDSS and DECaLS files contain full datasets (train, validation and test). Exact split that we performed (0.6 : 0.2 : 0.2) can be done using the code that accompanies this publication:&nbsp;<a href="https://github.com/deepskies/DeepAstroUDA">https://github.com/deepskies/DeepAstroUDA</a>&nbsp;.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Data and Analyses for Defining Filler Particles: A Phonetic Account of the Terminology, Form, and Grammatical Classification of Filled Pauses.

<p>Aggregated data and analyses for the article Belz, Malte (2023): Defining Filler Particles: A Phonetic Account of the Terminology,<br> Form, and Grammatical Classification of &quot;Filled Pauses&quot;. Languages. <a href="https://www.mdpi.com/journal/languages/special_issues/Pauses_in_Speech">https://www.mdpi.com/journal/languages/special_issues/Pauses_in_Speech</a></p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

New Particle Search at CERN open classification data

<p>This is a reduced, anonymised dataset containing the classifications made in the New Particle Search at CERN demonstrator project available on Zooniverse during the implementation period - from the 19th of October, 2021, to the 23rd of October, 2023 - as part of the REINFORCE project.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Cosmic Muon Images open classification data

<p>This is a reduced, anonymised dataset containing the classifications made in the Cosmic Muon Images demonstrator project available on Zooniverse during the implementation period - from the 19th of October, 2021, to the 23rd of October, 2023 - as part of the REINFORCE project.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

GWitchHunters open classification data

<p>This is a reduced, anonymised dataset containing the classifications made in the GWitchHunters demonstrator project available on Zooniverse during the implementation period - from the 19th of October, 2021, to the 23rd of October, 2023 - as part of the REINFORCE project.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Deep Sea Explorers open classification data

<p>This is a reduced, anonymised dataset containing the classifications made in the Deep Sea Explorers demonstrator project available on Zooniverse during the implementation period - from the 19th of October, 2021, to the 23rd of October, 2023 - as part of the REINFORCE project.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Multimodala Dataset for multimodal contrastive learning for crop classification

<p>We developed this dataset using an existing dataset name DENETHOR developed by TUM <a href="https://openreview.net/forum?id=uUa4jNMLjrL">https://openreview.net/forum?id=uUa4jNMLjrL</a> to conduct our multi-modal contrastive learning experiments.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

The Effect of Soundscape Composition on Bird Vocalization Classification in a Citizen Science Biodiversity Monitoring Project

<p>This archive includes sound clips (.wav files) and associated mel-scale spectrograms of bird vocalizations for 54 species in Sonoma County, California, USA. These data were used for training and validating convolutional neural network (CNN) models for bird species detection. We also include xeno-canto training and validation mel spectrograms&nbsp;used to pretrain CNNs. Details on these data are explained in the paper by Clark et al. (2023) titled &quot;The effect of soundscape composition on bird vocalization classification in a citizen science biodiversity monitoring project&quot;. These data are available for use without restrictions, with no warranty on data quality or utility for a given application. We request that any work that does use these data cite the Clark et al. (2023) paper.<br> <br> Clark, M.L., Salas, L., Baligar, S., Quinn, C., Snyder, R.L., Leland, D., Schackwitz, W., Goetz, S.J., Newsam, S. (2023). The effect of soundscape composition on bird vocalization classification in a citizen science biodiversity monitoring project. <em>Ecological Informatics</em>.&nbsp;<a href="https://doi.org/10.1016/j.ecoinf.2023.102065">https://doi.org/10.1016/j.ecoinf.2023.102065</a></p> <p>Associated code for training CNN models,&nbsp;performing inference, and applying post-classification corrections can be found in the GitHub archive&nbsp;<a href="https://github.com/pointblue/Soundscapes2Landscapes/tree/master/CNN_Bird_Species">https://github.com/pointblue/Soundscapes2Landscapes/tree/master/CNN_Bird_Species</a></p> <p>Raw sound data from the Soundscapes to Landscapes project are available upon request: Dr. Matthew Clark, matthew.clark@sonoma.edu</p> <p>These data were collected as part of the&nbsp;Soundscapes to Landscapes project (<a href="https://soundscapes2landscapes.org/">soundscapes2landscapes.org</a>),&nbsp;funded by NASA&rsquo;s Citizen Science for Earth Systems Program (CSESP) 16-CSESP 2016-0009 under cooperative agreement 80NSSC18M0107.<br> <br> ----------------------------<br> This depository&nbsp;includes the following archives:</p> <ul> <li> <p>mel_specs.zip: contains 2-sec mel spectrograms split into training (&ldquo;tr&rdquo;), validation (&ldquo;val&rdquo;), testing (&ldquo;test&rdquo;) data for each target bird species (n = 54) used to fine-tune the CNNs. Select spectrogram files are appended with &ldquo;aug&rdquo; if they are augmented versions for the training data.</p> </li> <li> <p>wav.zip: contains the associated wav-format sound recordings used to generate the training, validation, testing mel spectrograms found in mel_specs.zip.</p> </li> <li> <p>Xeno-canto_pretrain.tar: contains 2-sec mel spectrograms split into training and validation data for 40 bird species used for CNN pre-training that were generated using a warbleR segmentation methodology described in the paper. The sound files used to generate these mel spectrograms came from the Kaggle competition,&nbsp;<a href="https://www.kaggle.com/datasets/imoore/xenocanto-bird-recordings-dataset">https://www.kaggle.com/datasets/imoore/xenocanto-bird-recordings-dataset</a><br> Mel spectrogram naming reflects the XC number used for cataloging on Xeno-canto in the format XC123456_2.png. The six numbers following the XC characters can be used to search for unique recordings on Xeno-canto (<a href="https://xeno-canto.org/">https://xeno-canto.org/</a>) using the search query &ldquo;nr:123456&rdquo; in the search tool or queried using the Xeno-canto API (<a href="https://xeno-canto.org/explore/api">https://xeno-canto.org/explore/api</a>). Unique recording names can be extracted from the mel spectrogram filenames.</p> </li> <li> <p>soundscape_test_wavs.zip: the wav-format&nbsp;sound recordings&nbsp;used to perform soundscape testing.</p> </li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo44/100

CRACK-CH: A Crack detection and classification dataset on complex stone masonry surfaces

<p><strong>Description</strong></p> <p>This dataset includes various images with cracks&nbsp;from the test sites of HYPERION H2020 project (Grant Agreement No. 821054). Specifically, square image patches of 224x224 pixels from the test sites of Naillac and&nbsp;St. Nikolaos Fort are included.</p> <p>The Saint Nikolaos Fort is an important part of the great fortifications of the Medieval City of Rhodes located at the entrance of the Mantraki port. At this location, there was just a chapel dedicated to Saint Nikolaos until 1464, when it was turned to a fortification. Since then it has undergone reinforcements and expansions in order to defend the city. The outer walls were built in 1480 AD and in 1863 AD it was finally transformed to a lighthouse. The second study area is the Naillac at Saint Paul&rsquo;s rampart where a monumental tower was located as part of the fortification of the Commercial Harbour of Rhodes. It was constructed around 1400 AD on the Hellenistic Pier, but it was destroyed in 1863 after a severe earthquake. In 2017, the Naillac Tower was graphically reconstructed and presented as it stood until 1863, during the Ottoman rule. The Rodini Roman Bridge is one of the few ancient bridges surviving in Greece and part of the Hellenistic fortification of the city, making it a monument of great importance. It was built across the stream of Rhodini, situated outside the Medieval City and has two arched openings. The Roman Bridge is in continuous use until today and its static efficiency has deteriorated, while the scaffoldings which now support the arches are gradually rusting and losing their efficiency.&nbsp;Those images were split into two separate categories, facilitating the later training and evaluation of the model: &ldquo;Cracks&rdquo; and &ldquo;No cracks&rdquo;.</p> <p>The dataset is used to train and evaluate the&nbsp;CNN models for crack detection on complex stone masonry surfaces.</p> <p><strong>Publication</strong></p> <p>The paper is availbale here:&nbsp;https://arxiv.org/abs/2303.17989</p> <p><strong>If you use this dataset please cite it as CRACK-CH&nbsp;[reference]</strong>.<br> [Reference] Agrafiotis, Panagiotis, Doulamis, Anastasios, and Georgopoulos, Andreas. (2023) &quot;Unsupervised crack detection on complex stone masonry surfaces&quot;,&nbsp;<em>arXiv preprint arXiv:2303.17989</em><br> <br> Bibtex entry:</p> <p>@misc{agrafiotis2023unsupervised,<br> &nbsp; &nbsp; &nbsp; title={Unsupervised crack detection on complex stone masonry surfaces},&nbsp;<br> &nbsp; &nbsp; &nbsp; author={Panagiotis Agrafiotis and Anastastios Doulamis and Andreas Georgopoulos},<br> &nbsp; &nbsp; &nbsp; year={2023},<br> &nbsp; &nbsp; &nbsp; eprint={2303.17989},<br> &nbsp; &nbsp; &nbsp; archivePrefix={arXiv},<br> &nbsp; &nbsp; &nbsp; primaryClass={cs.CV}<br> }</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Taxonomy, distribution and classification of ecosystem-types, integrating the recent IUCN function-based typology and local conceptualizations

<p>1. Introduction:</p> <p>This dataset is a work in progress. It compiles data gathered on ecosystem-types and their distribution based on a series of field studies led by the author, in Seychelles and West and Central Africa (Senterre 2014, Senterre &amp; Wagner 2014, Senterre 2016, Senterre et al. 2017, 2019, 2020, 2021a, 2022). The aims of this dataset are:</p> <p>a. To share in an explicit and transparent way data on proposed taxonomies of ecosystems, i.e. conceptualizations of ecosystem-types, including explicit ecosystem names and management of synonymies.</p> <p>b. To develop ecosystem red listing based on transparent and falsifiable distribution raw data, combining distribution modeling (maps) and in situ observation of individual stand occurrences.</p> <p>c. To illustrate in detail how to deal with ecosystem data following the approach described in Senterre et al. (2021b) (i.e. &quot;ecosystemology&quot; approach).</p> <p>d. To integrate the above approach with the newly developed function-based typology of ecosystems (Keith et al. 2022), therefore contributing to bridging the persistent gap between the global and the local scales in ecosystem descriptions and classifications.</p> <p>&nbsp;</p> <p>2. Context and versions:</p> <p>This dataset was initially planned for publication on GBIF (Global Biodiversity Information Facility), as part of a project developed for the review of Key Biodiversity Areas in Seychelles: &quot;Mainstreaming recent species and ecosystem distribution data into Key Biodiversity Areas assessments in Seychelles&quot; (<a href="https://www.gbif.org/dataset/f513fe98-b1c3-45ee-8e14-7f2a5b7890bf">https://www.gbif.org/dataset/f513fe98-b1c3-45ee-8e14-7f2a5b7890bf</a>).</p> <p>In the first version of the GBIF dataset (<a href="https://www.gbif.org/dataset/f513fe98-b1c3-45ee-8e14-7f2a5b7890bf">https://www.gbif.org/dataset/f513fe98-b1c3-45ee-8e14-7f2a5b7890bf</a>), we proposed an analysis of the potential &#39;core&#39; and &#39;extension&#39; files available in GBIF for a publication of ecosystem-type names (and synonymies) and their corresponding occurrences recorded from field observations. This is an original analysis of taxonomic principles managed entirely at the scale of local observable objects, and their history of identifications or interpretations.</p> <p>Toward the end of the above-mentioned GBIF project, considering the limitations and gaps currently present in GBIF, it was decided to restrict the GBIF dataset to a simple &#39;metadata&#39; entry and to publish the complete version of this dataset in Zenodo. This allows to include all tables needed, as well as all required fields without having to accommodate them within the limited GBIF structure (see metadata description on GBIF for more details). The fields of the tables published here are described in the GBIF metadata entry and in the ecosystemology paper (Senterre et al. 2021b).</p> <p>&nbsp;</p> <p>3. New development on typology aspects:</p> <p>In addition, considering that the new IUCN global typology of ecosystems is now published (Keith et al. 2022), we have reviewed in detail the possibility of integration of ecosystems conceptualized using our ecosystemology approach within the new IUCN typology. The result of this analysis is being considered for a publication, and this Zenodo dataset would then be published in full (i.e. including all typology aspects) as supplementary materials. In the meantime, I would be happy to discuss any of these aspects with whoever is interested.</p> <p>&nbsp;</p> <p>4. Access to ecosystem data for conservation actors:</p> <p>Finally, the actual data (published here) on ecosystem-types, their names, synonymies, classification, distribution, and red list status are compiled into a format that we designed to be useful to conservation actors in the form of interactive webpages (produced with R as shiny apps). This development is based on very limited resources, and the author is still quite new to R, so any help or feedback on ways to improve the scripts would be very much welcomed.</p> <p>The interactive page is available here (currently filtered to Seychelles&#39; data only, although the dataset contains data beyond the Seychelles): https://shiny.bio.gov.sc/bioeco/</p> <p>The R scripts are available on Github: https://github.com/bsenterre/ecosystemology</p> <p>&nbsp;</p> <p>5. Tables contained in this dataset:</p> <p>a. Ecosystem taxonomy tables:</p> <p>ecoSpecies: Contains the list of all ecosystem-type names with their unique identifier.</p> <p>ecoOccurrences: Contains the list of individual stand occurrences, including ecosystem characters as standardized in Senterre et al. (2021b; i.e. virtual ecosystem specimen).</p> <p>ecoSpeciesProfiles: Contains basic metadata on ecosystem-types, such as their Red List evaluations.</p> <p>ecoIdentifications: Contains all the different interpretations/identifications (referring to the table ecoSpecies or to higher levels of classification, see below) made on the stands observed in the ecoOccurrences table.</p> <p>&nbsp;</p> <p>b. Ecosystem typology tables (TO BE ADDED LATER):</p> <p>IUCNL3: This is just a transcription, as is, of the IUCN global typology version 2.1.</p> <p>IUCNL3BIOCrossover: This table defines and comments correspondences between BIOL2 (the level 2 of the typology used by us) and the IUCN typology L3 (level 3).</p> <p>BIOL2: This is a variation based on the IUCN typology, here our level 2.</p> <p>BIOL3: This is a variation based on the IUCN typology, here our level 3.</p> <p>BIOL4: This is a variation based on the IUCN typology, here our level 4.</p> <p>ecoGenus: This is a general type of stand (thus excluding any regional ecosystem connotation), defined at a local scale and never combined with any geographic connotation (see ecosystemology paper: Senterre et al. 2021b).</p> <p>ecoFamily: This is a generalized version of the ecoGenus (i.e. still excluding any regional, sub-regional or geographic aspect).</p> <p>ecoOrder: This is a further generalized version of the ecoGenus (see also Senterre et al. 2020).</p> <p>lifeZone: This is a basic and incomplete list of life zones as defined following the Holdridge (1967) approach, with some additional elements proposed in Senterre et al. (2021b).</p> <p>&nbsp;</p> <p>6. Literature cited:</p> <p>Holdridge, L. R. 1967. Life zone ecology. Tropical Science Center, San Jose, Costa Rica.</p> <p>Keith, D. A., J. R. Ferrer-Paris, E. Nicholson, M. J. Bishop, B. A. Polidoro, E. Ramirez-Llodra, M. G. Tozer, J. L. Nel, R. Mac Nally, E. J. Gregr, K. E. Watermeyer, F. Essl, D. Faber-Langendoen, J. Franklin, C. E. R. Lehmann, A. Etter, D. J. Roux, J. S. Stark, J. A. Rowland, N. A. Brummitt, U. C. Fernandez-Arcaya, I. M. Suthers, S. K. Wiser, I. Donohue, L. J. Jackson, R. T. Pennington, T. M. Iliffe, V. Gerovasileiou, P. Giller, B. J. Robson, N. Pettorelli, A. Andrade, A. Lindgaard, T. Tahvanainen, A. Terauds, M. A. Chadwick, N. J. Murray, J. Moat, P. Pliscoff, I. Zager, and R. T. Kingsford. 2022. A function-based typology for Earth&rsquo;s ecosystems. . Nature 610:513&ndash;518. doi:10.1038/s41586-022-05318-4.</p> <p>Senterre, B. 2014. Mapping habitat-types within the Hummingbird site at Dugbe (Liberia, West Africa). Consultancy Report, Missouri Botanical Garden. P. 56. https://doi.org/10.13140/RG.2.2.32628.48003.</p> <p>Senterre, B. 2016. Habitat-type ground-truthing and assessment of ecosystem conservation value in the Bel Air Alufer mining site (Guinea, West Africa), with recommendations for improving the draft map of land cover types. Consultancy Report, Missouri Botanical Garden, A study conducted for Alufer Mining Limited. P. 54.</p> <p>Senterre, B., E. Bidault, and T. St&eacute;vart. 2019. Identification et &eacute;valuation des &eacute;cosyst&egrave;mes menac&eacute;s du Mont Nimba. Rapport de consultance, Missouri Botanical Garden (MBG), Africa and Madagascar Department. P. 106. https://doi.org/10.13140/RG.2.2.13242.93129.</p> <p>Senterre, B., E. Bidault, T. St&eacute;vart, and P. P. Lowry II. 2020. Assessment of Key Biodiversity Areas in the Lofa-Gola-Mano &amp; Nimba complexes (West Africa) using ecosystem criteria. Final Report, Missouri Botanical Garden. P. 146. 10.13140/RG.2.2.17934.89924.</p> <p>Senterre, B., E. Bidault, T. St&eacute;vart, M. Wagner, and P. Lowry. 2017. Mapping habitat-types in south-east Kouilou (Republic of Congo). Consultancy Report, Missouri Botanical Garden (MBG), Africa and Madagascar Department, St. Louis, Missouri, USA. P. 163.</p> <p>Senterre, B., R. M. Bristol, G. Gendron, and E. Henriette. 2021a. Fine-tuning conservation priorities in Seychelles at the landscape scale, using global KBA guidelines with both species and ecosystem criteria. Consultancy Report, United Nations Development Programme, GOS/UNDP/GEF Programme Coordination Unit, Victoria, Seychelles.</p> <p>Senterre, B., P. P. Lowry II, E. Bidault, and T. St&eacute;vart. 2021b. Ecosystemology: a new approach toward a taxonomy of ecosystems. . Ecological Complexity 47:100945. doi:https://doi.org/10.1016/j.ecocom.2021.100945.</p> <p>Senterre, B., A.-H. Paradis, E. Bidault, T. St&eacute;vart, and P. P. Lowry II. 2022. Qualit&eacute; et distribution des savanes montagnardes du Nimba. Rapport de consultance, Missouri Botanical Garden (MBG), Africa and Madagascar Department. P. 73. http://dx.doi.org/10.13140/RG.2.2.13433.34401.</p> <p>Senterre, B., and M. Wagner. 2014. Mapping Seychelles habitat-types on Mah&eacute;, Praslin, Silhouette, La Digue and Curieuse. Consultancy Report, Government of Seychelles, United Nations Development Programme, Victoria, Seychelles. P. 119. https://doi.org/10.13140/RG.2.1.4558.6009.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Keras video classification example with a subset of UCF101 - Action Recognition Data Set (top 10 videos)

<p>Classify video clips with natural scenes of actions performed by people visible in the videos.</p> <p>See the UCF101 Dataset web page: <a href="https://www.crcv.ucf.edu/data/UCF101.php#Results_on_UCF101">https://www.crcv.ucf.edu/data/UCF101.php#Results_on_UCF101</a></p> <p>This example datasets consists of the 10 most numerous video from the UCF101 dataset. For the top 5 version, see: <a href="https://doi.org/10.5281/zenodo.7924745">https://doi.org/10.5281/zenodo.7924745</a>&nbsp;.</p> <p>Based on this code: <a href="https://keras.io/examples/vision/video_classification/">https://keras.io/examples/vision/video_classification/</a> (needs to be updated, if has not yet been already; see the issue: <a href="https://github.com/keras-team/keras-io/issues/1342">https://github.com/keras-team/keras-io/issues/1342</a>).</p> <p>Testing if data can be downloaded from figshare with `wget`, see: <a href="https://github.com/mojaveazure/angsd-wrapper/issues/10">https://github.com/mojaveazure/angsd-wrapper/issues/10</a></p> <p>For generating the subset, see this notebook: <a href="https://colab.research.google.com/github/sayakpaul/Action-Recognition-in-TensorFlow/blob/main/Data_Preparation_UCF101.ipynb">https://colab.research.google.com/github/sayakpaul/Action-Recognition-in-TensorFlow/blob/main/Data_Preparation_UCF101.ipynb</a> -- however, it also needs to be adjusted (if has not yet been already - then I will post a link to the notebook here or elsewhere, e.g., in the corrected notebook with Keras example).</p> <p>I would like to thank Sayak Paul for contacting me about his example at Keras documentation being out of date.&nbsp;</p> <p>Cite this dataset as:</p> <p>Soomro, K., Zamir, A. R., &amp; Shah, M. (2012). UCF101: A dataset of 101 human actions classes from videos in the wild.&nbsp;<em>arXiv preprint arXiv:1212.0402</em>.&nbsp;<a href="https://doi.org/10.48550/arXiv.1212.0402">https://doi.org/10.48550/arXiv.1212.0402</a></p> <p>To download the dataset via the command line, please use:</p> <pre><code class="language-bash">wget -q https://zenodo.org/record/7882861/files/ucf101_top10.tar.gz -O ucf101_top10.tar.gz tar xf ucf101_top10.tar.gz</code></pre> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

CAELUS: Classification of sky conditions from 1-min time series of global solar irradiance using variability indices and dynamic thresholds

<p>CAELUS, a novel classification algorithm that relies on various thresholds to separate all possible sky conditions into six classes, is presented in Ruiz-Arias and Gueymard (2023, doi: <a href="https://doi.org/10.1016/j.solener.2023.111895">10.1016/j.solener.2023.111895</a>).</p> <p>This dataset was used to develop, validate and benchmark CAELUS. It is made up by 1-min quality-assured observations of global horizontal irradiance (GHI)&nbsp;and diffuse horizontal irradiance&nbsp;at 54 stations of the Baseline Surface Radiation Network (BSRN) archive, which is publicly available (see download instructions in https://bsrn.awi.de/data). The dataset includes&nbsp;5 years of data per station, except in two of them (Petrolina, Brazil, and Solar Village, Saudi Arabia), combined with other variables that are required to run CAELUS, namely: solar zenith angle (sza), extraterrestrial horizontal solar irradiance (eth), clear-sky GHI (ghics) and GHI in a clean and dry atmosphere (ghicda). In addition, the dataset also provides the sky classification obtained with CAELUS.</p> <p>Further details about CAELUS and the dataset compilation is available in Ruiz-Arias and Gueymard (2023, doi:&nbsp;<a href="https://doi.org/10.1016/j.solener.2023.111895">10.1016/j.solener.2023.111895</a>). A Python implementation of CAELUS is available in&nbsp;https://github.com/jararias/caelus.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Keras video classification example with a subset of UCF101 - Action Recognition Data Set (top 5 videos)

<p>Classify video clips with natural scenes of actions performed by people visible in the videos.</p> <p>See the UCF101 Dataset web page:&nbsp;<a href="https://www.crcv.ucf.edu/data/UCF101.php#Results_on_UCF101">https://www.crcv.ucf.edu/data/UCF101.php#Results_on_UCF101</a></p> <p>This example datasets consists of the 5&nbsp;most numerous video from the UCF101 dataset. For the top 10 version see:&nbsp;<a href="https://doi.org/10.5281/zenodo.7882861">https://doi.org/10.5281/zenodo.7882861</a>&nbsp;.</p> <p>Based on this code:&nbsp;<a href="https://keras.io/examples/vision/video_classification/">https://keras.io/examples/vision/video_classification/</a>&nbsp;(needs to be updated, if has not yet been already; see the issue:&nbsp;<a href="https://github.com/keras-team/keras-io/issues/1342">https://github.com/keras-team/keras-io/issues/1342</a>).</p> <p>Testing if data can be downloaded from figshare with `wget`, see:&nbsp;<a href="https://github.com/mojaveazure/angsd-wrapper/issues/10">https://github.com/mojaveazure/angsd-wrapper/issues/10</a></p> <p>For generating the subset, see this notebook:&nbsp;<a href="https://colab.research.google.com/github/sayakpaul/Action-Recognition-in-TensorFlow/blob/main/Data_Preparation_UCF101.ipynb">https://colab.research.google.com/github/sayakpaul/Action-Recognition-in-TensorFlow/blob/main/Data_Preparation_UCF101.ipynb</a>&nbsp;-- however, it also needs to be adjusted (if has not yet been already - then I will post a link to the notebook here or elsewhere, e.g., in the corrected notebook with Keras example).</p> <p>I would like to thank Sayak Paul for contacting me about his example at Keras documentation being out of date.&nbsp;</p> <p>Cite this dataset as:</p> <p>Soomro, K., Zamir, A. R., &amp; Shah, M. (2012). UCF101: A dataset of 101 human actions classes from videos in the wild.&nbsp;<em>arXiv preprint arXiv:1212.0402</em>.&nbsp;<a href="https://doi.org/10.48550/arXiv.1212.0402">https://doi.org/10.48550/arXiv.1212.0402</a></p> <p>To download the dataset via the command line, please use:</p> <pre><code class="language-bash">wget -q https://zenodo.org/record/7924745/files/ucf101_top5.tar.gz -O ucf101_top5.tar.gz tar xf ucf101_top5.tar.gz</code></pre>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record