Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

45,411

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

45,411 results for “collection”

Learn how ShareScore rates datasets ↗
zenodo44/100

Harnessing the power of digitized natural history collections to visualize spatiotemporal patterns in native and non-native bee flight phenology

<p>What&nbsp;time&nbsp;of&nbsp;year&nbsp;are&nbsp;bees&nbsp;flying,&nbsp;where&nbsp;are&nbsp;they&nbsp;flying,&nbsp;and&nbsp;how&nbsp;do&nbsp;biogeographical&nbsp;factors,&nbsp;sex,&nbsp;and&nbsp;native&nbsp;status&nbsp;affect&nbsp;flight&nbsp;phenology?&nbsp;Consistent&nbsp;monitoring&nbsp;along&nbsp;with&nbsp;creating&nbsp;spatially&nbsp;and&nbsp;temporally&nbsp;explicit&nbsp;visualizations&nbsp;using&nbsp;large&nbsp;openly&nbsp;available&nbsp;data&nbsp;sets&nbsp;enhance&nbsp;our&nbsp;understanding&nbsp;of&nbsp;trends&nbsp;in&nbsp;flight&nbsp;time&nbsp;phenology&nbsp;and&nbsp;shape&nbsp;our&nbsp;understanding&nbsp;of&nbsp;bee-plant&nbsp;interactions,&nbsp;including&nbsp;shifts&nbsp;in&nbsp;the&nbsp;phenology&nbsp;of&nbsp;bee&nbsp;pollinators.</p> <p>Species&nbsp;occurrence&nbsp;data&nbsp;from&nbsp;digitized&nbsp;collection&nbsp;networks&nbsp;(iNaturalist,&nbsp;Global&nbsp;Biodiversity&nbsp;Information&nbsp;Faculty&nbsp;(GBIF),&nbsp;Integrated&nbsp;Digitized&nbsp;Biocollections&nbsp;(iDigBio),&nbsp;Symbiota&nbsp;Collections&nbsp;of&nbsp;Arthropods&nbsp;Network&nbsp;(SCAN),&nbsp;and&nbsp;UC&nbsp;Santa&nbsp;Barbara&nbsp;Collection&nbsp;Network)&nbsp;are&nbsp;part&nbsp;of&nbsp;an&nbsp;effort&nbsp;to&nbsp;improve&nbsp;our&nbsp;understanding&nbsp;of&nbsp;bees&nbsp;in&nbsp;coastal&nbsp;Santa&nbsp;Barbara&nbsp;County,&nbsp;including&nbsp;the&nbsp;California&nbsp;Channel&nbsp;Islands.&nbsp;New&nbsp;inventory&nbsp;collections&nbsp;combined&nbsp;with&nbsp;historical&nbsp;data&nbsp;from&nbsp;over&nbsp;11&nbsp;natural&nbsp;history&nbsp;museums&nbsp;and&nbsp;2&nbsp;observation&nbsp;networks&nbsp;are&nbsp;used&nbsp;in&nbsp;an&nbsp;effort&nbsp;to&nbsp;examine&nbsp;patterns&nbsp;and&nbsp;changes&nbsp;in&nbsp;phenology&nbsp;of&nbsp;native&nbsp;and&nbsp;non-native&nbsp;bee&nbsp;species,&nbsp;and&nbsp;create&nbsp;updated&nbsp;species&nbsp;inventories.</p> <p>Synthesizing species observation data from digitized natural history collections makes use of a wealth of existing data and multiplies the analytical power of isolated observations, but it is not without limitations and challenges. By exploring novel techniques to generate clear and accurate visualizations to communicate bee flight time, we present our key initial findings and identify geographic, temporal, and taxonomic gaps, which will lead to further focused inventory projects of coastal Santa Barbara County, improved data quality for phenological analyses, and reusable methods for visualizing insect phenology data across taxa or geography.</p> <p><strong>The attached files include the R code and some of the .csv files used to produce the figures in my poster that was available on demand at the Entomology Society of America 2020 virtual meeting.&nbsp;&nbsp;</strong></p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Collected recommendations and requirements for FAIR-enabling services

<p>Within FAIRsFAIR task 2.4, we carried out a structured literature review to extract&nbsp;requirements, recommendations and other desiderata for FAIR-enabling services. This document contains the full list of excerpts, structured and annotated.&nbsp;This work&nbsp;has been used as input for the basic framework on FAIRness of services developed by FAIRsFAIR task 2.4&nbsp;(see https://doi.org/10.5281/zenodo.4292599).</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

MAVIS Twitter dataset: A collection of tweets and sentiment analysis in Spanish about vaccines and diseases during the period 2015-2018

<p>MAVIS dataset comprises a full knowledge base regarding Twitter messages published in Spanish during the period 2015-2018, in the context of sentiment analysis of specific vaccines and their related diseases. Such diseases and vaccines are summarized as follows:</p> <ul> <li>Invasive meningococcal disease (&ldquo;EMI&rdquo; in Spanish): Bexsero, Trumenba, Nimenrix</li> <li>Invasive pneumococcal disease (&ldquo;ENI&rdquo; in Spanish)</li> <li>Influenza</li> <li>Hepatitis</li> <li>Rotavirus: Rotarix, Rotateq</li> <li>Measles (&ldquo;Sarampi&oacute;n&rdquo; in Spanish) and MMR (&ldquo;Triple v&iacute;rica&rdquo; in Spanish)</li> <li>Sepsis</li> <li>Whooping cough (&ldquo;Tosferina&rdquo; in Spanish)</li> <li>Chickenpox (&ldquo;Varicela&rdquo; in Spanish): Varivax, Varilrix; and Shingles (&ldquo;Zoster&rdquo; in Spanish)</li> <li>Human papillomavirus infection (&ldquo;VPH&rdquo; in Spanish): Cervarix, Gardasil</li> </ul> <p>Tweets have been manually classified as having a negative or non-negative sentiment by 5 experts. Moreover, an automatic classification has been performed by 3 different tools: IBM Watson (now Watson Tone Analyzer, <a href="https://www.ibm.com/watson/services/tone-analyzer/">https://www.ibm.com/watson/services/tone-analyzer/</a>), Google Cloud Natural Language (<a href="https://cloud.google.com/natural-language">https://cloud.google.com/natural-language</a>), and Meaning Cloud (<a href="https://www.meaningcloud.com/">https://www.meaningcloud.com/</a>). IBM Watson and Google Cloud Natural Language returned a numerical sentiment score ranging from -1 to 1, while Meaning Cloud returned a categorical variable with the values &lsquo;P+&rsquo;, &lsquo;P&rsquo;, &lsquo;NEU&rsquo;, &lsquo;N&rsquo; and &lsquo;N+&rsquo;, which were converted to 1, 2, 3, 4 and 5 respectively.</p> <p>With these variables (IBM Watson, Google Cloud Natural Language, and Meaning Cloud annotations and the experts&rsquo; classification as the target label), a machine learning metamodel was developed. Tweets were also annotated with the sentiment output given by this classifier. &nbsp;&nbsp;</p> <p>The provided data includes intrinsic tweets information, intrinsic information regarding the users that posted the tweets, the keywords mentioned in each tweet, and the annotations that the experts, the tools, and the model gave to each tweet.</p> <p><strong>Funding</strong>: This dataset was obtained with funding from&nbsp;MSD, Spain under MAVIS Study (VEAP ID: 7789).</p> <p><strong>Current studies using this dataset at the moment of the publication</strong>:</p> <ul> <li>Rodr&iacute;guez-Gonz&aacute;lez et al., &ldquo;Creating a metamodel based on machine learning to identify the sentiment of vaccine and disease-related messages in Twitter: the MAVIS study&rdquo; in 2020 IEEE 33st International Symposium on Computer-Based Medical Systems (CBMS), Jul. 2020, p. 6. DOI: 10.1109/CBMS49503.2020.00053</li> <li>Rodr&iacute;guez-Gonz&aacute;lez et al., &quot;Identifying Polarity in Tweets from an Imbalanced Dataset about Diseases and Vaccines Using a Meta-Model Based on Machine Learning Techniques&quot; in Applied Sciences, 2020, 10. DOI: 10.3390/app10249019</li> </ul>

opencc-by-4.0Dec 2020View details →
zenodo44/100

1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test using artificial intelligence and image-based pattern recognition techniques

<p><strong>The Great Isaiah Scroll (1QIsa<sup>a</sup>) data set for writer identification</strong></p> <p>This data set is collected for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497</p> <p>Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> <br> <strong>Copyright (c) </strong>&nbsp;&nbsp; &nbsp;University of Groningen, 2021. All rights reserved.<br> <strong>Disclaimer and copyright notice for all data contained on this .tar.gz file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the data for research purposes. It is not allowed to distribute this data for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this data.</p> <p><strong>4) </strong>the user should refer to the first public article on this data set:<br> <br> <em>Popović, M., Dhali, M. A., &amp; Schomaker, L. (2020). Artificial intelligence-based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.</em><br> <br> BibTeX:</p> <pre>@article{popovic2020artificial, title={Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)}, author={Popovi{\&#39;c}, Mladen and Dhali, Maruf A and Schomaker, Lambert}, journal={arXiv preprint arXiv:2010.14476}, year={2020} }</pre> <p><strong>5) </strong>the recipient should refrain from proliferating the data set to third parties external to his/her local research group. Please refer interested researchers to this site for obtaining their own copy.</p> <p><strong>Organisation of the data:</strong></p> <p>The .tar.gz file contains three directories: images, features, and plots. The included &#39;README&#39; file contains all the instructions.</p> <p>The &#39;images&#39; directory contains NetPBM images of the columns of 1QIsa<sup>a</sup>. The NetPBM format is chosen because of its simplicity. Additionally, there is no doubt about lossy compression in the processing chain. There are two images for each of the Great Isaiah Scroll columns: one is the direct binarized output from the BiNet (<em>arxiv.org/abs/1911.07930</em>) system, and the other one is the manually cleaned version of the binarized output. &nbsp; The file names for the direct binarized output are of the format &#39;1QIsaa_col&lt;columnnr&gt;.pbm&#39;, for example, &#39;1QIsaa_col15.pbm&#39;. And, for the cleaned version, the format is &#39;1QIsaa_col&lt;columnnr&gt;_cleaned.pbm&#39;, for example, &#39;1QIsaa_col15_cleaned.pbm&#39;. Note: the image files are not in a separate directory; they will be extracted in the same place. However, due to the unique naming, there is no problem extracting them in one single directory.</p> <p>The &#39;features&#39; directory contains feature files computed for each of the column images. There are two types of feature files: Hinge and Adjoined. They are distinguishable by their extension, for example, &#39;1QIsaa_col15_cleaned.hinge&#39; and &#39;1QIsaa_col15_cleaned.adjoined&#39;. They are also arranged in separate directories for ease of use.</p> <p>The &#39;plots&#39; directory contains a simple python script to perform PCA on the feature files and then visualize them in a 3D plot. The file takes the location of feature files as an input. The &#39;README_plot&#39; file contains examples of how-to-run in the terminal.</p> <p><strong>Brief description:</strong><br> According to ImageMagick&#39;s&#39; identify&#39; tool, the original images are in grayscale (.jpg) from Brill collection, in &#39;8-bit Gray 256c&#39;. &nbsp;These images pass through multiple preprocessing measures to become suitable for pattern recognition-based techniques. The first step in preprocessing is the image-binarization technique. In order to prevent any classification of the text-column images based on irrelevant background patterns, a specific binarization technique (BiNet) was applied, keeping the original ink traces intact. After performing the binarization, the images were cleaned further by removing the adjacent columns that partially appear on the target columns&#39; images. Finally, few minor affine transformations and stretching corrections were performed in a restrictive manner. These corrections are also targeted for aligning the texts where the text lines get twisted due to the leather writing surface&#39;s degradation. Hence, the clean images are there in the directory along with the direct binarized images. No effort has been made to obtain a balanced set in any way.</p> <p><strong>Tools:</strong><br> <strong>Binarization:</strong><br> The BiNet tool is available for scientific use upon request (m.a.dhal(at)rug.nl)</p> <p><strong>Image Morphing:</strong><br> In the original article, data augmentation was performed using image morphing. The tool is available on GitHub:<br> https://github.com/GrHound/imagemorph.c</p> <p><strong>Features for writer identification:</strong><br> Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> <em><strong>1.&nbsp;</strong>L. Schomaker &amp; M. Bulacu (2004). Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.<br> <strong>2. </strong>Bulacu, M. &amp; Schomaker, L.R.B. (2007). Text-independent Writer Identification and Verification Using Textural and Allographic Features, &nbsp;IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</em><br> &nbsp;<br> The features (hinge, fraglets) have been combined in a single MS Windows application, GIWIS, which is available for scientific use upon request (l.r.b.schomaker(at)rug.nl)</p> <p><strong>If you have any question, please contact us:</strong><br> Maruf A. Dhali &lt;m.a.dhali(at)rug.nl&gt;<br> Lambert Schomaker &lt;l.r.b.schomaker(at)rug.nl&gt;<br> Mladen Popović &lt;m.popovic(at)rug.nl&gt;</p> <p><strong>Please cite our papers if you use this data set:</strong><br> <em><strong>1.</strong> Popović, M., Dhali, M. A., &amp; Schomaker, L. (2020). Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.<br> <strong>2. </strong>Dhali, M. A., de Wit, J. W., &amp; Schomaker, L. (2019). Binet: Degraded-manuscript binarization in diverse document textures and layouts using deep encoder-decoder networks. arXiv preprint arXiv:1911.07930.</em></p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

ecoli_VF_collection: v0.1

<p>First release (v0.1) of ecoli_VF_collection, with 1069 <em>E. coli</em> virulence factor protein sequences.</p> <p>https://github.com/aleimba/ecoli_VF_collection</p>

opencc-by-4.0Jun 2016View details →
zenodo44/100

Archaeological Survey of India Collections: Burma Circle, 1907-13. Photo 1004/2 : 1907-1913 (.jpg format)

<p>This is a digitized collections of photos held by the British Library (catalog record photo 1004/2). These photos were taken by the Archeological Survey of India (Burma Circle) between 1907-1913. The digitization was conducted by the photography lab of the British Library as part of the Pyu epigraphy sub-project (PI, Nathan W. Hill of SOAS University of London) of the ERC synergy grant "Beyond Boundaries: Religion, Region, Language and the State" (Identifier: ASIA 609823). The original photographs are under crown copyright, which means that sufficient time has past for them to be distributed to the public. Here is the record from the BL--</p> <ul> <li><strong>Title:</strong> <p>Archaeological Survey of India Collections: Burma Circle, 1907-13. Photographer(s): Archaeological Survey of India</p> </li> <li><strong>Collection Area: </strong> Visual Arts</li> <li><strong>Reference: </strong> Photo 1004/2</li> <li><strong>Creation Date: </strong> 1907-1913</li> <li><strong>Extent and Access: </strong><br> <strong>Extent: </strong>321 items<br> <strong>Conditions of Use: </strong>Appointment Required to view these records. Please consult Asian and African Studies Print Room staff.</li> <li><strong>Language: </strong> Not applicable</li> <li><strong>Contents and Scope: </strong><br> <strong>Contents: </strong> <p>Blue half-leather album, 385x322mm, containing prints mounted on pages interleaved with tissue. The photographs were taken by the Archaeological Survey, Burma, between 1907-13, and listed in the annual Report of the Superintendent... Copies of each year's report precede the photographs for that year, as follows:</p> <p>Year Photo no</p> <p>1907-08 510-609</p> <p>1908-09 610-750</p> <p>1909-10 751-859</p> <p>1910-11 860-962</p> <p>1911-12 963-1053</p> <p>1912-13 1054-1182</p> <p>The photographs, including views of Burmese architecture, sculpture and relics, were taken under the direction of the Superintendent, a post held at the time by Taw Sein Ko.</p> <p>Process = Collodio-chloride and gelatine silver prints</p> <p>Photographers = Archaeological Survey of India., ;</p> </li> <li><strong>History: </strong><br> <strong>Immediate Source of Acquisition: </strong> <p>Official Deposit</p> </li> </ul>

opencc-by-4.0Jul 2017View details →
zenodo44/100

Archaeological Survey of India Collections: Burma Circle, 1913-1916. Photo 1004/3 : 1913-1916 (.jpg format)

<p>This is a digitized collections of photos held by the British Library (catalog record photo 1004/3). These photos were taken by the Archeological Survey of India (Burma Circle) between 1913-1916. The digitization was conducted by the photography lab of the British Library as part of the Pyu epigraphy sub-project (PI, Nathan W. Hill of SOAS University of London) of the ERC synergy grant "Beyond Boundaries: Religion, Region, Language and the State" (Identifier: ASIA 609823). The original photographs are under crown copyright, which means that sufficient time has past for them to be distributed to the public. Here is the record from the BL--</p> <ul> <li><strong>Title:</strong> <p>Archaeological Survey of India Collections: Burma Circle, 1913-16. Photographer(s): Archaeological Survey of India</p> </li> <li><strong>Collection Area: </strong> Visual Arts</li> <li><strong>Reference: </strong> Photo 1004/3</li> <li><strong>Creation Date: </strong> 1913-1916</li> <li><strong>Extent and Access: </strong><br> <strong>Extent: </strong>265 items<br> <strong>Conditions of Use: </strong>Appointment Required to view these records. Please consult Asian and African Studies Print Room staff.</li> <li><strong>Language: </strong> Not applicable</li> <li><strong>Contents and Scope: </strong><br> <strong>Contents: </strong> <p>Blue half-leather album, 385x322mm, containing prints mounted on pages interleaved with tissue. The photographs were taken by the Archaeological Survey, Burma, between 1913-16, and listed in the annual Report of the Superintendent... Copies of each year's report precede the photographs for that year, as follows:</p> <p>Year Photo no</p> <p>1913-14 1183-1317</p> <p>1914-15 1318-1490</p> <p>1915-16 1491-1593</p> <p>The photographs, including views of Burmese architecture, sculpture and relics, were taken under the direction of the Superintendent, a post held at the time by Taw Sein Ko.</p> <p>Process = Gelatine silver prints</p> <p>Photographers = Archaeological Survey of India., ;</p> </li> <li><strong>History: </strong><br> <strong>Immediate Source of Acquisition: </strong> <p>Official Deposit</p> </li> </ul>

opencc-by-4.0Jul 2017View details →
zenodo44/100

Archaeological Survey of India Collections: Burma Circle, 1916-22. Photo 1004/4 : 1916-22 (.jpg format)

<p>This is a digitized collections of photos held by the British Library (catalog record photo 1004/4). These photos were taken by the Archeological Survey of India (Burma Circle) between 1916-22. The digitization was conducted by the photography lab of the British Library as part of the Pyu epigraphy sub-project (PI, Nathan W. Hill of SOAS University of London) of the ERC synergy grant "Beyond Boundaries: Religion, Region, Language and the State" (Identifier: ASIA 609823). The original photographs are under crown copyright, which means that sufficient time has past for them to be distributed to the public. Here is the record from the BL--</p> <ul> <li><strong>Title:</strong> <p>Archaeological Survey of India Collections: Burma Circle, 1916-22. Photographer(s): Archaeological Survey of India</p> </li> <li><strong>Collection Area: </strong> Visual Arts</li> <li><strong>Reference: </strong> Photo 1004/4</li> <li><strong>Creation Date: </strong> 1916-1922</li> <li><strong>Extent and Access: </strong><br> <strong>Extent: </strong>238 items<br> <strong>Conditions of Use: </strong>Appointment Required to view these records. Please consult Asian and African Studies Print Room staff.</li> <li><strong>Language: </strong> Not applicable</li> <li><strong>Contents and Scope: </strong><br> <strong>Contents: </strong> <p>Blue half-leather album, 385x335mm, containing prints mounted on pages interleaved with tissue. The photographs were taken by the Archaeological Survey, Burma, between 1916-22 (except for 1917-18 and 1918-19), and listed in the annual Report of the Superintendent ... Copies of each year's report precede the photographs for that year, as follows:</p> <p>Year Photo no</p> <p>1916-17 1594-1743</p> <p>1917-18 No prints</p> <p>1918-19 No prints</p> <p>1919-20 1991-2094</p> <p>1920-21 2095-2203</p> <p>1921-22 2204-2280</p> <p>The photographs, including views of Burmese architecture, sculpture and relics, were taken under the direction of the Superintendent, a post held until 1919 by Taw Sein Ko, and thereafter by his successor Charles Duroiselle.</p> <p>Process = Gelatine silver prints</p> <p>Photographers = Archaeological Survey of India., ;</p> </li> <li><strong>History: </strong><br> <strong>Immediate Source of Acquisition: </strong> <p>Official Deposit</p> </li> </ul> <p> </p>

opencc-by-4.0Aug 2017View details →
zenodo44/100

Underwater images collected by Scuba diving in Ifaty, Madagascar - 2023-05-07

<i>This dataset was collected by Scuba diving in Ifaty, Madagascar - 2023-05-07.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 2.66 GB of MP4 files, which were trimmed into 1056 frames (at 2997/1000 fps). <br> The frames are not georeferenced. <br> 93.75% of these extracted images are useful and 6.25% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> No GPS. <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by Scuba diving in Nosy-Ve, Madagascar - 2023-04-30

<i>This dataset was collected by Scuba diving in Nosy-Ve, Madagascar - 2023-04-30.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 1.5 GB of JPG files, which were trimmed into 267 frames (at 4 fps). <br> The frames are georeferenced. <br> 99.63% of these extracted images are useful and 0.37% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> No GPS. <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-03

<i>This dataset was collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-03.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 15.07 GB of MP4 files, which were trimmed into 5136 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 59.54% of these extracted images are useful and 40.46% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 0.0 %, Q2: 0.0 %, Q5: 100.0 % <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by an Autonomous Surface Vehicle in Ifaty, Madagascar - 2023-05-06

<i>This dataset was collected by an Autonomous Surface Vehicle in Ifaty, Madagascar - 2023-05-06.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 17.33 GB of MP4 files, which were trimmed into 4251 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 100.0% of these extracted images are useful and 0.0% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> Base : No Base <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 0.0 %, Q2: 0.0 %, Q5: 100.0 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 0.2 m and 50.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. <br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.422 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by an Autonomous Surface Vehicle in Anakao, Madagascar - 2023-05-01

<i>This dataset was collected by an Autonomous Surface Vehicle in Anakao, Madagascar - 2023-05-01.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 13.58 GB of MP4 files, which were trimmed into 2767 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 80.88% of these extracted images are useful and 19.12% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 0.0 %, Q2: 0.0 %, Q5: 100.0 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 0.2 m and 50.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. <br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.335 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by an Autonomous Surface Vehicle in Ifaty, Madagascar - 2023-05-07

<i>This dataset was collected by an Autonomous Surface Vehicle in Ifaty, Madagascar - 2023-05-07.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 25.86 GB of MP4 files, which were trimmed into 5425 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 99.17% of these extracted images are useful and 0.83% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> Base : No Base <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 0.0 %, Q2: 0.0 %, Q5: 100.0 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 0.2 m and 20.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. <br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.313 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-04

<i>This dataset was collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-04.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 0.0 %, Q2: 0.0 %, Q5: 100.0 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 0.2 m and 35.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. <br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.492 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-05

<i>This dataset was collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-05.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 34.7 GB of MP4 files, which were trimmed into 9193 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 98.34% of these extracted images are useful and 1.66% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 0.0 %, Q2: 0.0 %, Q5: 100.0 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 0.2 m and 20.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. <br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.463 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

IAGOS-CARIBIC MS files collection (v2025.07.11)

<h2><strong>Content</strong></h2> <p><em><strong>IAGOS-CARIBIC_MS_files_collection_20250711</strong></em> contains merged IAGOS-CARIBIC data, on a 10s grid (CARIBIC-1 and CARIBIC-2; &lt;https://www.caribic-atmospheric.com/&gt;). There is one netCDF (version 4) file per IAGOS-CARIBIC flight. Files were generated from NASA Ames 1001 source files. For detailed content information, see global and variable attributes. Global attribute `na_file_header` contains the original NASA Ames file header as an array of strings.</p> <h2><strong>Data Coverage</strong></h2> <p>The data set covers 22 years of CARIBIC data from 1997 to 2020, flight numbers 1 to 591. There is no data available after 2020. Also, there is no data available for the following flight numbers within the [1..591] range:</p> <ul> <li>CARIBIC-1, 1997-2002, 1-97: 4, 5, 6, 21, 28, 30, 31, 32, 38, 39, 73, 81, 83, 91, 95</li> <li>98 to 109 do not exist</li> <li>CARIBIC-2, 2005-2020, 110-591: 217, 276, 277, 318, 320, 410, 411, 412, 425, 426, 427, 428, 434, 435 436</li> </ul> <h3>Special note on CARIBIC-1 data</h3> <p>CARIBIC-1 data only contains a subset of the variables found in CARIBIC-2 data files. To distinguish those two campaigns, use the global attribute 'mission'.</p> <h2>File format</h2> <p>netCDF v4, created with xarray, &lt;https://docs.xarray.dev/en/stable/&gt;. Compression: zlib, level 5. Metadata conventions: CF-1.10, ACDD-1.3 (see also 'comment'&nbsp; global attribute).</p> <h2><strong>Authors and Parameters Info</strong></h2> <p>See `CARIBIC-MS_files_species-and-contributors.csv` in the zip archive.</p> <h3><strong>Primary Contact</strong></h3> <ul> <li>Andreas Zahn, IAGOS-CARIBIC Coordinator , &lt;andreas.zahn@kit.edu&gt;</li> <li>Florian Obersteiner, data management, &lt;florian.obersteiner@kit.edu&gt;</li> </ul> <h2><strong>Data Availability<br></strong></h2> <p>This dataset is also available via the KIT-IMKASF THREDDS server, &lt;https://thredds.atmohub.kit.edu/thredds/catalog/iagos-caribic/catalog.html&gt;.</p> <h2><strong>Changelog</strong></h2> <ul> <li>`2025.07.11`: extend netCDF metadata (CF-1.10, ACDD-1.3), revise standard names and units. Data unchanged.</li> <li>`2025.06.02`: introduce netCDF file compression (zlib, level 5). Data unchanged.</li> <li>`2025.03.25`: variable name change (PosLat =&gt; lat, PosLong =&gt; lon, pstatic =&gt; p, Altitude =&gt; alt), add acetonitrile and acetone measurement precision columns</li> <li>`2024.10.28`: revise CARIBIC-1 data, all flights (note on lat/lon inaccuracy, range checks for static pressure and temperature). CARIBIC-2 data unchanged.</li> <li>`2024.07.17`: revise ozone data for flights 294 to 591</li> <li>`2024.01.12`: revise naming convention of (nc attributes), add POF IV funding reference (zenodo)</li> <li>`2023.11.15`: add CARIBIC-1 data, revise variable long names</li> <li>`2023.10.17`: remove duplicate flights ("MSA" is "MAA"), add previously missing flights (200, 201)</li> <li>`2023.09.26`: extend data; include soot photometer measurements</li> <li>`2023.07.26`: initial upload</li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Underwater images collected by Scuba diving in Nosy-Ve, Madagascar - 2023-04-30

<i>This dataset was collected by Scuba diving in Nosy-Ve, Madagascar - 2023-04-30.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 2.49 GB of JPG files, which were trimmed into 460 frames (at 4 fps). <br> The frames are not georeferenced. <br> 100.0% of these extracted images are useful and 0.0% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> No GPS. <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-03

<i>This dataset was collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-03.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 16.67 GB of MP4 files, which were trimmed into 8820 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 81.93% of these extracted images are useful and 18.07% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 0.0 %, Q2: 0.0 %, Q5: 100.0 % <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →
zenodo44/100

Underwater images collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-05

<i>This dataset was collected by an Autonomous Surface Vehicle in Sarodrano, Madagascar - 2023-05-05.</i> <br> <br><br>Underwater or aerial images collected by scientists or citizens can have a wide variety of use for science, management, or conservation. These images can be annotated and shared to train IA models which can in turn predict the objects on the images. We provide a set of tools (hardware and software) to collect marine data, predict species or habitat, and provide maps.<br><br> This dataset is part of larger collection referencing numerous underwater and aerial images <a href="https://doi.org/10.5281/zenodo.11125847" target="_blank">Seatizen Altas</a>. Methods, tools and scientific objectives are also described in a dedicated data paper.<br> <h2>Image acquisition</h2> This session has 20.22 GB of MP4 files, which were trimmed into 5846 frames (at 2997/1000 fps). <br> The frames are georeferenced. <br> 87.7% of these extracted images are useful and 12.3% are useless, according to predictions made by <a href="jacques-v0.1.0_model-20240513_v20.0" target="_blank">Jacques model</a>. <br> Multilabel predictions have been made on useful frames using <a href="https://huggingface.co/lombardata/DinoVdeau-large-2024_04_03-with_data_aug_batch-size32_epochs150_freeze" target="_blank">DinoVd'eau</a> model. <br> <h2> GPS information: </h2> The data was processed with a PPK workflow to achieve centimeter-level GPS accuracy. <br> Base : Files coming from rtk a GPS-fixed station or any static positioning instrument which can provide with correction frames. <br> Device GPS : Emlid Reach M2 <br> Quality of our data - Q1: 0.0 %, Q2: 0.0 %, Q5: 100.0 % <br> <h2> Bathymetry </h2> The data are collected using a single-beam echosounder <a href="https://www.echologger.com/products/single-frequency-echosounder-deep" target="_blank">ETC 400</a>. <br> We keep the points that are the waypoints.<br> We keep the raw data where depth was estimated between 0.2 m and 35.0 m deep. <br> The data are first referenced against the WGS84 ellipsoid. <br> At the end of processing, the data are projected into a homogeneous grid to create a raster and a shapefiles. <br> The size of the grid cells is 0.352 m. <br> The raster and shapefiles are generated by linear interpolation. The 3D reconstruction algorithm is ballpivot. <br> <h2> Generic folder structure </h2> YYYYMMDD_COUNTRYCODE-optionalplace_device_session-number <br> ├── DCIM : folder to store videos and photos depending on the media collected. <br> ├── GPS : folder to store any positioning related file. If any kind of correction is possible on files (e.g. Post-Processed Kinematic thanks to rinex data) then the distinction between device data and base data is made. If, on the other hand, only device position data are present and the files cannot be corrected by post-processing techniques (e.g. gpx files), then the distinction between base and device is not made and the files are placed directly at the root of the GPS folder. <br> │ ├── BASE : files coming from rtk station or any static positioning instrument. <br> │ └── DEVICE : files coming from the device. <br> ├── METADATA : folder with general information files about the session. <br> ├── PROCESSED_DATA : contain all the folders needed to store the results of the data processing of the current session. <br> │ ├── BATHY : output folder for bathymetry raw data extracted from mission logs. <br> │ ├── FRAMES : output folder for georeferenced frames extracted from DCIM videos. <br> │ ├── IA : destination folder for image recognition predictions. <br> │ └── PHOTOGRAMMETRY : destination folder for reconstructed models in photogrammetry. <br> └── SENSORS : folder to store files coming from other sources (bathymetry data from the echosounder, log file from the autopilot, mission plan etc.). <br> <h2> Software </h2> All the raw data was processed using our <a href="https://doi.org/10.5281/zenodo.15853010" target="_blank">worflow</a>. <br>All predictions were generated by our <a href="https://doi.org/10.5281/zenodo.15228535" target="_blank">inference pipeline</a>. <br>You can find all the necessary scripts to download this data in this <a href="https://github.com/SeatizenDOI/zenodo-tools" target="_blank">repository</a>. <br>Enjoy your data with <a href="https://github.com/SeatizenDOI" target="_blank">SeatizenDOI</a>! <br>

opencc-by-4.0May 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record