Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
7,031
datasets available to search
ShareScore release 0.7.1
Dataset results
7,031 results for “marine”
Supporting data for “Climate Intervention through Stratospheric Aerosol Injection may partially mitigate marine heatwaves"
Although climate intervention aims to lower the global average temperature, the potential impact of Stratospheric Aerosol Injection on marine heatwaves (MHW) has not been thoroughly examined. This spatial dataset provides global and regional MHW metrics—such as frequency, maximum intensity, and duration—from the Community Earth System Model, version 2 (CESM2), using the baseline scenario SSP2-4.5, referred to as a no climate intervention scenario, and the ARISE-SAI ensemble. The ARISE-SAI model uses the SSP2-4.5 scenario, introducing stratospheric aerosol injection at approximately 21 km in 2035, aiming to keep global mean surface air temperature near 1.5°C for ARISE-SAI-1.5 and near 1.0°C for ARISE-SAI-1.0 above pre-industrial levels. The dataset includes global MHW properties for the historical period (1990-2009), the current period under SSP2-4.5 emission scenario (2015-2034), and future scenarios under SSP2-4.5, ARISE-SAI-1.5, and ARISE-SAI-1.5 for 2050-2059 and 2060-2069.
Santa Barbara Channel Marine BON: Nearshore kelp forest integrated benthic cover, 1980-ongoing
The Santa Barbara Channel Marine Biodiversity Observation Network (SBCMBON) tracks long-term patterns in species abundance and diversity. This dataset contains cover of kelp forest sessile invertebrates, understory macroalgae, and substrate types by integrating data from four contributing projects working in the kelp forests of the Santa Barbara Channel, USA. Divers collect data on using either uniform point contact (UPC) or random point contact (RPC) methods. The four contributing projects are two research projects: The Santa Barbara Coastal LTER (SBC LTER) and the Partnership for Interdisciplinary Studies of Coastal Oceans (PISCO), the kelp forest monitoring program of the Santa Barbara Channel National Park, and the San Nicolas Island monitoring program supported by USGS. Together, these projects have recorded data for more than 200 species at approximately 100 sites on both the mainland coast and on the Santa Barbara Channel Islands. Sampling began in 1982 and is ongoing. Data were collected by human observation (divers using SCUBA) during regular surveys. Percent cover is recorded for taxa where individuals cannot be counted. Cover can be calculated from the data here as the fraction of total points at which the taxon was present x 100. With UPC and RPC methods, multiple species can be recorded at any given point. The total percent cover of all species combined using this method can exceed 100%; however, the percent cover of any single species cannot exceed 100%. See Methods for information on integration and data processing. MBON is funded by National Aeronautics and Space Administration (NASA), Bureau of Ocean Energy Management (BOEM), and National Oceanic and Atmospheric Administration (NOAA). For users who are interested in using all or part of this integrated datasets, please contact data owners to discuss your research interests, data-related issues or any other questions. A recommended citation for the data package is available from the download page. In
Santa Barbara Channel Marine BON: Nearshore kelp forest integrated fish, 1981-ongoing
The Santa Barbara Channel Marine Biodiversity Observation Network (SBCMBON) tracks long-term patterns in species abundance and diversity. This dataset contains counts of fish (including cryptic fish, which are deliberately sought out) produced by integrating data from four contributing projects working in the kelp forests of the Santa Barbara Channel, USA. The four contributing projects are two research projects, the Santa Barbara Coastal LTER (SBC LTER) and the Partnership for Interdisciplinary Studies of Coastal Oceans (PISCO), and the kelp forest monitoring program of the Santa Barbara Channel National Park, and the San Nicolas Island monitoring program supported by USGS. Together, these projects have recorded data for more than 200 species at approximately 100 sites on both the mainland coast and on the Santa Barbara Channel Islands. Sampling began in 1980 and is ongoing. Data were collected by human observation (divers using SCUBA) during regular surveys. This dataset includes five entities, three data tables and two R scripts. The main data table contains counts of organisms, the bottom-area and height of the water column over which the fish were surveyed. The column labeled “count” records the number of organisms found in each plot/transect at a given timestamp. A second data table contains place names and geolocation for sampling sites. Information is sufficient for the calculation of fish density, which is left to the user. The third data table contains the depths of each transect the for the fish survey. Sample R script is included to illustrate generation of a basic table of areal density by taxa and sampling site. A second R script is included to convert these data from their primary MBON structure to a Darwin Core Archive (and available through multiple sources). MBON is funded by National Aeronautics and Space Administration (NASA), Bureau of Ocean Energy Management (BOEM), and National Oceanic and Atmospheric Administration (NOAA). For users who are i
Santa Barbara Channel Marine BON: Nearshore kelp forest integrated quad and swath survey, 1980-ongoing
The Santa Barbara Channel Marine Biodiversity Observation Network (SBCMBON) tracks long-term patterns in species abundance and diversity. This dataset contains counts of algae and invertebrates (both sessile and mobile) by integrating data from four contributing projects working in the kelp forests of the Santa Barbara Channel, USA. The four contributing projects are two research projects: The Santa Barbara Coastal LTER (SBC LTER) and the Partnership for Interdisciplinary Studies of Coastal Oceans (PISCO), the kelp forest monitoring program of the Santa Barbara Channel National Park, and the San Nicolas Island monitoring program supported by USGS. Together, these projects have recorded data for more than 200 species at approximately 100 sites on both the mainland coast and on the Santa Barbara Channel Islands. Sampling began in 1982 and is ongoing. Data were collected by human observation (divers using SCUBA) during regular surveys. The data table documents the number of organisms and the area over which that number was counted for calculation of areal abundance. Data were collected by human observation (divers using SCUBA) during regular surveys. The algae and invertebrate counts record the number of taxa found in each plot, including quad (small square plots such as 1 or 2 m2) and swath (large linear plots such as 60 m2). See Method and protocol for information on integration and data processing. MBON is funded by National Aeronautics and Space Administration (NASA), Bureau of Ocean Energy Management (BOEM), and National Oceanic and Atmospheric Administration (NOAA). For users who are interested in using all or part of this integrated datasets, please contact data owners to discuss your research interests, data-related issues or any other questions. A recommended citation for the data package is available from the download page. In addition, any manuscript generated using this dataset is expected to be sent to the data owners before publication so we can be sure t
Santa Barbara Channel Marine BON: Nearshore kelp forest integrated taxa, 1980-ongoing
The Santa Barbara Channel Marine Biodiversity Network (SBC MBON) tracks long-term patterns in species abundance and diversity. By integrating research and monitoring efforts (both existing and de novo data), we provide a comprehensive view of biodiversity in the region. This dataset is a combined species list from four datasets currently handled and integrated by SBC MBON, with identifiers from an appropriate taxonomic registry included for each taxon. Because this species list reflects the integration of multiple collections, it will be updated as SBC MBON integration efforts develop. Santa Barbara Channel Marine BON: Integrated kelp forest/reef: Fish https://portal.edirepository.org/nis/mapbrowse?scope=edi&identifier=5&revision=newest Santa Barbara Channel Marine BON: Integrated kelp forest/reef: Quad and swath cover https://portal.edirepository.org/nis/mapbrowse?scope=edi&identifier=6&revision=newest Santa Barbara Channel Marine BON: Integrated kelp forest/reef: Benthic cover https://portal.edirepository.org/nis/mapbrowse?scope=edi&identifier=3&revision=newest As of 2018, this dataset contains approximately 400 taxa from combined observations beginning in 1980. Data include codes for the originating project and sampling method, plus basic taxonomic lineage. MBON is funded by National Aeronautics and Space Administration (NASA), Bureau of Ocean Energy Management (BOEM), and National Oceanic and Atmospheric Administration (NOAA). For users who are interested in using all or part of this integrated datasets, please contact data owners to discuss your research interests, data-related issues or any other questions. A recommended citation for the data package is available from the download page. In addition, any manuscript generated using this dataset is expected to be sent to the data owners before publication so we can be sure the data is used in the proper context and methods are reported accurately: Santa Barbara Coastal LTER (LTER): Dan Ree
LTER-Italy site Marine Area of Portofino Promontory - Italy figure
<p>Geographical representation of the LTER-Italy site Marine Area of Portofino Promontory - Italy (LTER_EU_IT_015) - DEIMS-ID <a href="https://deims.org/769556a6-0ee6-46a9-acbb-a1f2d51c07e8">https://deims.org/769556a6-0ee6-46a9-acbb-a1f2d51c07e8</a></p>
Marine magnetic anomaly data from high resolution surveys off the SW Portuguese coast
<p>This dataset contains <strong>magnetic anomaly grids</strong> that result from the full processing of marine magnetic data collected off the SW Portuguese coast between 2014 and 2019. A total area of ~4400 km<sup>2</sup> was surveyed with average line spacing of 1 nautic mile. Surveys covered the continental shelf and in some regions reaching up to 2500 m bathymetric levels. Total magnetic field data were acquired with a G882 Cesium vapor marine magnetometer towed, towed at sea surface.</p> <p><strong>Full processing</strong> of magnetic data included: layback correction; noise removal; IGRF subtraction; base station correction; line leveling; minimum curvature gridding. The resulting sea level magnetic anomaly grid was further processed for upward continuation and reduction to the pole, providing additional outputs. </p> <p>The following grids are provided in <strong>georeferenced geotiff format</strong>:</p> <ul> <li>Magnetic anomaly (sealevel)</li> <li>Magnetic anomaly reduced to the pole (sealevel)</li> <li>Magnetic anomaly upward continued to 200 m height </li> <li>Magnetic anomaly upward continued to 200 m height, reduced to the pole</li> <li>Magnetic anomaly upward continued to 3000 m height </li> <li>Magnetic anomaly upward continued to 3000 m height, reduced to the pole</li> </ul> <p><strong>Published in</strong>: Neres, M., P. Terrinha, J. Noiva, P. Brito, M. Rosa, L. Batista, C. Ribeiro (2023). <em>New Late Cretaceous and CAMP magmatic sources off West Iberia, from high-resolution magnetic surveys on the continental shelf.</em> <strong>Tectonics</strong>. doi: 10.1029/2022TC007637</p> <p> </p>
2-Dimensional habitat files for 47 representative marine species
<p><strong>2D marine species habitats in NetCDF format on 0.5*0.5 degree global regular grid.</strong></p> <p>Based on Close et al. (2006) and converted from .CSV format. </p> <p>Filename is in format 'presence_speciesnumber.nc' where species numbers are listed in the README.txt file (identifier for each species). The README.txt file further contains each species' species_group which is the assigned depth group for each species (1=0-200m epipelagic, 2=200-1000m mesopelagic, 3=sea floor demersal) and species_name which is the Latin name of each species with underscore in between.</p> <p>The species occurs where the variable 'presence' equals 1 (in the accompanying paper we assume this to be the 1995-2014 climatological mean distribution).</p> <p>In the NetCDF files, the variable 'presence' has as an attribute 'species' which contains the Latin species name without underscore.</p>
MS and NMR data of in situ Captured Marine Exometabolites
<p>This folder contains the raw data pertaining to the article <i><strong>In Situ</strong></i> <strong>Capture and Real Time Enrichment of Marine Chemical Diversity </strong></p><p><a href="https://doi.org/10.1021/acscentsci.3c00661">https://doi.org/10.1021/acscentsci.3c00661</a></p><p>Data are organized in folders corresponding to each figure. Briefly, this folder contains the raw mass spectrometry (MS) data, the cytoscape files of the full molecular network (Fig3), the xcel spreadsheets of annotated MS spectra related to each investigated specialized exometabolites from the Mediterranean sponges <i>Aplysina cavernicola </i>(AC, Fig4), <i>Spongia officinalis </i>(SO, Fig5)<i>, </i>and <i>Agelas oroides </i>(AO, Fig6)<i>, </i>the raw 1H NMR data from each sponge exometabolite (EM) extract with their corresponding crude extract (CR).</p><ul><li>All MS2 data were acquired on a Bruker Impact II qTOF (ESI positive, collision energy 20-40eV) also deposited here : MSV000091465</li><li>SIRIUS software and CANOPUS were used to further annotate the chemodiversity of captured marine EMs</li><li>All NMR data were acquired on a BRUKER avance II+ instrument (600 MHz, cryoprobe) in CD<i>3</i>OD</li></ul><p>-------------------------</p><p><strong>References related to in silico MS annotation tools:</strong></p><ul><li>Kai Dührkop, Louis-Félix Nothias, Markus Fleischauer, Raphael Reher, Marcus Ludwig, Martin A. Hoffmann, Daniel Petras, William H. Gerwick, Juho Rousu, Pieter C. Dorrestein and Sebastian Böcker <i>Systematic classification of unknown metabolites using high-resolution fragmentation mass spectra</i>. Nature Biotechnology, 2020. https://doi.org/10.1038/s41587-020-0740-8</li><li>Yannick Djoumbou Feunang, Roman Eisner, Craig Knox, Leonid Chepelev, Janna Hastings, Gareth Owen, Eoin Fahy, Christoph Steinbeck, Shankar Subramanian, Evan Bolton, Russell Greiner, David S. Wishart <i>ClassyFire: automated chemical classification with a comprehensive, computable taxonomy </i>J Cheminf, 8, 2016. https://doi.org/10.1186/s13321-016-0174-y</li><li>Kim, Hyun Woo and Wang, Mingxun and Leber, Christopher A. and Nothias, Louis-Félix and Reher, Raphael and Kang, Kyo Bin and van der Hooft, Justin J. J. and Dorrestein, Pieter C. and Gerwick, William H. and Cottrell, Garrison W. NPClassifier:<i> A Deep Neural Network-Based Structural Classification Tool for Natural Products. </i>Journal of Natural Products, 84, 2021. https://doi.org/10.1021/acs.jnatprod.1c00399</li></ul>
Annual Water Quality in Everglades National Park, Florida Bay, West Florida Shelf, and Florida Keys National Marine Sanctuary, Florida, USA: 1994-2019
Annual (water year basis) geometric mean concentrations of total phosphorus (TP), soluble-reactive phosphorus (SRP), total nitrogen (TN), dissolved inorganic nitrogen (DIN; calculated as nitrate + nitrite + ammonia), chlorophyll-a (Chl-a), and total organic carbon (TOC) concentrations across Everglades National Park (ENP), Florida Keys National Marine Sanctuary (FKNMS), West Florida Shelf and Florida Bay. This dataset is composed of data from multiple sources including Florida Coastal Everglades, Florida International University Southeast Research Center (FIU SERC), South Florida Water Management District (SFWMD), and National Oceanic and Atmospheric Administration Atlantic Oceanographic and Meteorological Laboratory (NOAA AOML). All values reported less than the laboratory minimum detection limit (MDL) were set to one-half the MDL. Annual geometric mean concentrations were computed for monitoring locations with greater than five years of data and four samples per year with a minimum of one sample in the wet and dry seasons. This dataset was created to evaluate long-term spatial and temporal trends in nutrients, chlorophyll-a, and total organic carbon at the landscape scale relative to freshwater and marine ecosystems.
MCR LTER: Coral Reef: Farmerfish gardens help buffer stony corals against marine heat waves, data for Honeycutt et al., PLOS One 2023
These data were generated in support of the manuscript: Honeycutt RC, Holbrook SJ, Brooks, AJ, and RJ Schmitt, PLOS One In Moorea, French Polynesia, we evaluated the response and fate of stony coral following a major thermal stress event in 2019 that caused a substantial amount of branching coral (dominantly Pocillopora) to bleach and die. We investigated whether Pocillopora colonies that occurred within territorial gardens protected by the farmerfish Stegastes nigricans were less susceptible to or survived bleaching better than Pocillopora on adjacent, undefended substrate. Bleaching prevalence and severity, which were quantified for >1,100 colonies shortly after they bleached, did not differ between colonies within or outside of defended gardens. By contrast, 399 focal colonies followed for one year revealed that a bleached coral within a garden was a third less likely to suffer complete colony death and, for survivors, about twice as likely to recover to its pre-bleaching cover of living tissue compared to Pocillopora outside of a farmerfish garden. Our findings indicate that while residing in a farmerfish garden may not reduce the bleaching susceptibility of a coral during thermal stress, it does help buffer a bleached coral against severe outcomes. This oasis effect of farmerfish gardens, where survival and recovery of thermally-damaged corals are enhanced, is another mechanism that helps explain why large Pocillopora colonies are far more abundant in farmerfish territories than elsewhere in the lagoons of Moorea, despite gardens being much less common. As such, farmerfish may have a growing role in maintaining the resilience of branching corals as the frequency and intensity of marine heat waves continue to increase. This material is based upon work supported by the U.S. National Science Foundation under Grant No. OCE 16-37396 (and earlier awards) as well as a generous gift from the Gordon and Betty Moore Foundation. Research was completed under permits issued
Estimates of nitrogen and phosphorus excretion rates in individual marine and estuarine animals
This dataset contains nitrogen and phosphorus excretion rate, as well as dry biomass, estimates for individual vertebrate and invertebrate animals in marine and estuarine environments. This dataset is a product of an LTER Synthesis Working Group aimed at evaluating the spatiotemporal variability in consumer nutrient dynamics in the wake of global change across eight long-term ecological research projects. These projects include seven long-term ecological research programs (LTER) funded by the National Science Foundation: (1) California Current Ecosystem, (2) Florida Coastal Everglades, (3) Moorea Coral Reef, (4) Northern Gulf of Alaska, (5) Plum Island Ecosystems, (6) Santa Barbara Coastal, and (7) Virginia Coast Reserve LTER projects. Additionally, the dataset includes data from (8) The Partnership for Interdisciplinary Science of Coastal Oceans (PISCO) research program. The temporal coverage of each time series data varies among projects, with the earliest record in 1997 and the most recent in 2023. This data package also includes two folders of R scripts used for data harmonization, identical to those in the LTER Synthesis Working Group: Consumer-Mediated Nutrient Dynamics Project, v2.0.0. You can find the release in GitHub here: https://github.com/lter/lterwg-marine-cnd/releases/tag/v2.0.0
Overview of the time series in the PALMOD 130k marine palaeoclimate data synthesis
<p>Palaeoclimate time series in the PALMOD 130k marine palaeoclimate data synthesis v1.0.1. This table lists the site names and location, parameters including additional information as well as the source of the data and the original publications where the data were presented.</p>
Data in support of 'ENSO influences subsurface marine heatwave occurrence in the Kuroshio Extension'
<p>Data in support of 'Chandler M, Sprintall J, Zilberman NV. (2025). ENSO influences subsurface marine heatwave occurrence in the Kuroshio Extension. <em>Journal of Geophysical Research: Oceans</em>. <a href="https://doi.org/10.1029/2025JC022899" target="_blank" rel="noopener">https://doi.org/10.1029/2025JC022899</a>'</p> <p> </p> <p>There are 2 netCDF files:</p> <ol> <li>p40tem1211_2312.nc</li> <li>synthetic_T_10day_px40_kuroshio_chandler2024.nc</li> </ol> <p><strong>p40tem1211_2312.nc </strong>contains the temperature sections from <a href="https://www-hrx.ucsd.edu/px40.html">HR-XBT transect PX40</a> objectively mapped onto a 10-m depth grid and a 0.1° longitudinal grid. <em>[LONGITUDE; LATITUDE; DEPTH; TIME; TEM]</em></p> <p><strong>synthetic_T_10day_px40_kuroshio_chandler2024.nc</strong> contains the synthetic temperature anomaly time series between the surface and 800-m deep at the western end of transect PX40 over the period from January-1993 to April-2023, as well as the temperature annual cycle needed for reconstructing the full synthetic temperature time series. <em>[time; depth; longitude; latitude; T_prime; T_ann]</em></p> <p> </p> <p>There is 1 MATLAB file:</p> <ol> <li>px40_synthetic_T.m</li> </ol> <p><strong>px40_synthetic_T.m</strong> is the MATLAB script used to produce the synthetic temperature anomaly time series saved in synthetic_T_10day_px40_kuroshio_chandler2024.nc.</p> <p> </p> <p>There is 1 Julia file:</p> <ol> <li>px40_synthetic_T_julia.jl</li> </ol> <p><strong>px40_synthetic_T_julia.jl</strong> is a Julia implementation of the MATLAB script px40_synthetic_T.m.</p> <p> </p> <p>There is 1 R file:</p> <ol> <li>px40_synthetic_T_R.R</li> </ol> <p><strong>px40_synthetic_T_R.R</strong> is an R implementation of the MATLAB script px40_synthetic_T.m.</p> <p> </p> <p><code>Version history:</code><br><code>v1.0.0 First uploaded (25-November-2024)</code><br><code>v1.0.1 Julia script uploaded (18-January-2025)</code><br><code>v1.0.2 R script uploaded (28-January-2025)</code><br><code>v1.1.0 Updated description of synthetic_T_10day_px40_kuroshio_chandler2024.nc to include reference to accepted publication (21-August-2025)</code></p>
Smartbay Marine Species Object Detection Training dataset
<h1>Training dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Species Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using species names.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 training dataset format. </p>
Smartbay Marine Types Object Detection Training dataset
<h1>Training Dataset</h1> <p>The SmartBay Observatory in Galway Bay is an important contribution by Ireland to the growing global network of real-time data capture systems deployed within the ocean – technology giving us new insights into the ocean which we have not had before.</p> <p>The observatory was installed on the seafloor 1.5km off the coast of Spiddal, County Galway, Ireland . The observatory uses cameras, probes and sensors to permit continuous and remote live underwater monitoring. This observatory equipment allows ocean researchers unique real-time access to monitor ongoing changes in the marine environment. Data relating to the marine environment at the site is transferred in real-time from the SmartBay Observatory through a fibre optic telecommunications cable to the Marine Institute headquarters and onwards onto the internet. The data includes a live video stream, the depth of the observatory node, the sea temperature and salinity, and estimates of the chlorophyll and turbidity levels in the water which give an indication of the volume of phytoplankton and other particles, such as sediment, in the water.</p> <p>The Smartbay Marine Types Object Detection training Dataset is an initial Bounding Box Annotated image dataset used in attempting to Train a YOLOv8 Object Detection Model to classify the Marine Fauna observed in the Smartbay Observatory Video footage using broad "Marine Type" classes.</p> <p>The imagery used in this training dataset consists of image frame captures from the <a href="https://smartbay.marine.ie">Smartbay</a> video Archive files, CC-BY imagery from the <a href="https://www.minka-sdg.org">www.minka-sdg.org</a> website and images taken by Eva Cullen in the "<a href="https://nationalaquarium.ie/">Galway Atlantaquaria</a>" Aquarium in Galway, Ireland.</p> <p>The imagery were annotated using CVAT, collated on <a href="https://www.roboflow.com/">Roboflow</a> and exported in YOLOv8 trainign dataset format. </p>
Targeted and untargeted LC-MS copepodamide datasets for marine and freshwater copepods
<p>This repository contains the datasets, analysis code and output generated and used in the scientific article titled "Mass spectroscopy reveals compositional differences in copepodamides from limnic and marine copepods" published in Scientific Reports (https://doi.org/10.1038/s41598-024-53247-1)<em>.</em></p> <p>Detailed information about the datasets are available in the README.txt.</p> <p>The source dataset created from the sampling effort, with targeted liquid chromatography coupled mass spectrometry (LC-MS) data, taxonomic information of individual copepods, their length measurements, estimated biomass etc is available in <em><strong>Masterfile_targeted_data_final.xlsx</strong></em>.</p> <p>The source dataset for precursor LC-MS scan data is available in <em><strong>Precursor_data_Deisotoped.xlsx</strong></em>. </p> <p>The resulting analysis data frames (last sheet in each .xlsx file) are available as separate .csv files (<strong>Arnoldt_targeted_analysis_data.csv</strong> & <strong>Arnoldt_targeted_analysis_data.xlsx</strong>). These files are denoted "Supplementary Data. 2" and "Supplementary Data. 1" respectively in the main article. Data files <strong>Chromatography.csv</strong> and <strong>zooplankton_composition_bulk.csv</strong> are used to create chromatograph line plots (Figures 3a & 3b in the article) and one of the supplementary figures (S1), respectively.</p> <p>A R-markdown file (<strong>Arnoldt_R_Code</strong><em><strong>.Rmd</strong></em>) with the code to analyse and visualise all data, and its html-output file (<strong>Arnoldt_R_Code_Output</strong><em><strong>.html</strong></em>) are also available here. The markdown files uses the four csv-files described in the paragraph above to generate all analyses and figures. The output (.html) file is denoted "Supplementary Code" in the main article.</p>
MarFERReT: an open-source, version-controlled reference library of marine microbial eukaryote functional genes
<p>Metatranscriptomics generates large volumes of sequence data about transcribed genes in natural environments. Taxonomic annotation of these datasets depends on availability of curated reference sequences. For marine microbial eukaryotes, current reference libraries are limited by gaps in sequenced organism diversity and barriers to updating libraries with new sequence data, resulting in taxonomic annotation of only about half of eukaryotic environmental transcripts. Here, we introduce version 1.0 of the Marine Functional EukaRyotic Reference Taxa (MarFERReT), an updated marine microbial eukaryotic sequence library with a version-controlled framework designed for taxonomic annotation of eukaryotic metatranscriptomes. We gathered 902 marine eukaryote genomes and transcriptomes from multiple sources and assessed these candidate entries for sequence quality and cross-contamination issues, selecting 800 validated entries for inclusion in the library. MarFERReT v1 contains reference sequences from 800 marine eukaryotic genomes and transcriptomes, covering 453 species- and strain-level taxa, totaling nearly 28 million protein sequences with associated NCBI and PR2 Taxonomy identifiers and Pfam functional annotations. An accompanying MarFERReT project repository hosts containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT.<br><br>MarFERReT is linked to a code repository hosting containerized build scripts, documentation on installation and use case examples, and information on new versions of MarFERReT here: <a href="https://github.com/armbrustlab/marferret">https://github.com/armbrustlab/marferret</a></p> <p>The raw source data for the 902 candidate entries considered for MarFERReT v1.1.1, including the 800 accepted entries, are available for download from their respective online locations. The source URL for each of the entries is listed here in MarFERReT.v1.1.1.entry_curation.csv, and detailed instructions and code for downloading the raw sequence data from source are available in the MarFERReT code repository (<a href="https://github.com/armbrustlab/marferret/blob/main/docs/process_clean_marmicrodb.log.sh">link</a>). </p> <p>This repository release contains MarFERReT database files from the v1.1.1 MarFERReT release using the following MarFERReT library build scripts: <strong>assemble_marferret.sh</strong>, <strong>pfam_annotate.sh</strong>, and <strong>build_diamond_db.sh</strong><br><br>The following MarFERReT data products are available in this repository:</p> <p><strong>MarFERReT.v1.1.1.metadata.csv</strong><br>This CSV file contains descriptors of each of the 902 database entries, including data source, taxonomy, and sequence descriptors. Data fields are as follows:</p> <ol> <li><strong>entry_id</strong>: Unique MarFERReT sequence entry identifier.</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y/N). The Y/N values can be adjusted to customize the final build output according to user-specific needs.</li> <li><strong>marferret_name</strong>: A human and machine friendly string derived from the NCBI Taxonomy organism name; maintaining strain-level designation wherever possible.</li> <li><strong>tax_id</strong>: The NCBI Taxonomy ID (taxID).</li> <li><strong>pr2_accession</strong>: Best-matching PR2 accession ID associated with entry</li> <li><strong>pr2_rank</strong>: The lowest shared rank between the entry and the pr2_accession</li> <li><strong>pr2_taxonomy</strong>: PR2 Taxonomy classification scheme of the pr2_accession</li> <li><strong>data_type</strong>: Type of sequence data; transcriptome shotgun assemblies (TSA), gene models from assembled genomes (genome), and single-cell amplified genomes (SAG) or transcriptomes (SAT).</li> <li><strong>data_source</strong>: Online location of sequence data; the Zenodo data repository (<a href="../">Zenodo</a>), the datadryad.org repository (<a href="http://datadryad.org/">datadryad.org</a>), MMETSP re-assemblies on Zenodo (MMETSP)17, NCBI GenBank (<a href="https://www.ncbi.nlm.nih.gov/genbank/">NCBI</a>), JGI Phycocosm (<a href="https://phycocosm.jgi.doe.gov/phycocosm/home">JGI-Phycocosm</a>), the TARA Oceans portal on Genoscope (<a href="http://www.genoscope.cns.fr/tara/">TARA</a>), or entries from the Roscoff Culture Collection through the METdb database repository (<a href="https://metdb.sb-roscoff.fr/metdb/">METdb</a>).</li> <li><strong>source_link</strong>: URL where the original sequence data and/or metadata was collected.</li> <li><strong>pub_year</strong>: Year of data release or publication of linked reference.</li> <li><strong>ref_link</strong>: Pubmed URL directs to the published reference for entry, if available.</li> <li><strong>ref_doi</strong>: DOI of entry data from source, if available.</li> <li><strong>source_filename</strong>: Name of the original sequence file name from the data source.</li> <li><strong>seq_type</strong>: Entry sequence data retrieved in nucleotide (nt) or amino acid (aa) alphabets.</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file.</li> <li><strong>source_name:</strong> Full organism name from entry source</li> <li><strong>original_taxID</strong>: Original NCBI taxID from entry data source metadata, if available</li> <li><strong>alias:</strong> Additional identifiers for the entry, if available</li> </ol> <p><br><strong>MarFERReT.v1.1.1.curation.csv</strong><br>This CSV file contains curation and quality-control information on the 902 candidate entries considered for incorporation into MarFERReT v1, including curated NCBI Taxonomy IDs and entry validation statistics. Data fields are as follows:</p> <ol> <li><strong>entry_id:</strong> Unique MarFERReT sequence entry identifier</li> <li><strong>marferret_name: </strong>Organism name in human and machine friendly format, including additional NCBI taxonomy strain identifiers if available.</li> <li><strong>tax_id</strong>: Verified NCBI taxID used in MarFERReT</li> <li><strong>taxID_status</strong>: Status of the final NCBI taxID (Assigned, Updated, or Unchanged)</li> <li><strong>taxID_notes</strong>: Notes on the original_taxID</li> <li><strong>n_seqs_raw</strong>: Number of sequences in the original sequence file</li> <li><strong>n_pfams</strong>: Number of Pfam domains identified in protein sequences</li> <li><strong>qc_flag</strong>: Early validation quality control flags for the following: LOW_SEQS; less than 1,200 raw sequences; LOW_PFAMS; less than 500 Pfam domain annotations.</li> <li><strong>flag_Lasek</strong>: Flag notes from Lasek-Nesselquist and Johnson (2019); contains the flag 'FLAG_LASEK' indicating ciliate samples reported as contaminated in this study.</li> <li><strong>VV_contam_pct</strong>: Estimated contamination reported for MMETSP entries in Van Vlierberghe et al., (2021).</li> <li><strong>flag_VanVlierberghe: </strong>Flag for a high level of estimated contamination, from 'flag_VanVlierberghe' values over 50%: FLAG_VV.</li> <li><strong>rp63_npfams</strong>: Number of ribosomal protein Pfam domains out of 63 total.</li> <li><strong>rp63_contam_pct</strong>: Percent of total ribosomal protein sequences with an inferred taxonomic identity in any lineage other than the recorded identity, as described in the Technical Validation section from analysis of 63 Pfam ribosomal protein domains.</li> <li><strong>flag_rp63</strong>: Flag for a high level of estimated contamination, from 'rp63_contam_pct' values over 50%: FLAG_RP63.</li> <li><strong>flag_sum: </strong>Count of the number of flag columns (`qc_flag`, `flag_Lasek`, `flag_VanVlierberghe`, and `flag_rp63`). All entries with one or more flag are nominally rejected ('accepted' = N); entries without any flags are validated and accepted ('accepted' = Y).</li> <li><strong>accepted: </strong>Acceptance into the final MarFERReT build (Y or N).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins.faa.gz</strong><br>This Gzip-compressed FASTA file contains the 27,951,013 final translated and clustered protein sequences for all 800 accepted MarFERReT entries. The sequence defline contains the unique identifier for the sequence and its reference (mftX, where 'X' is a ten-digit integer value). </p> <p> </p> <p><strong>MarFERReT.v1.1.1.taxonomies.tab.gz</strong><br>This Gzip-compressed tab-separated file is formatted for interoperability with the DIAMOND protein alignment tool commonly used for downstream analyses and contains some columns without any data. Each row contains an entry for one of the MarFERReT protein sequences in MarFERReT.v1.proteins.faa.gz. Note that 'accession.version' and 'taxid' are populated columns while 'accession' and 'gi' have NA values; the latter columns are required for back-compatibility as input for the DIAMOND alignment software and LCA analysis. </p> <p>The columns in this file contain the following information:</p> <ol> <li><strong>accession</strong>: (NA)</li> <li><strong>accession.version</strong>: The unique MarFERReT sequence identifier ('mftX').</li> <li><strong>taxid</strong>: The NCBI Taxonomy ID associated with this reference sequence.</li> <li><strong>gi</strong>: (NA).</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.proteins_info.tab.gz</strong><br>This Gzip-compressed tab-separated file contains a row for each final MarFERReT protein sequence with the following columns:</p> <ol> <li><strong>aa_id</strong>: the unique identifier for each MarFERReT protein sequence.</li> <li><strong>entry_id</strong>: The unique numeric identifier for each MarFERReT entry.</li> <li><strong>source_defline</strong>: The original, unformatted sequence identifier</li> </ol> <p> </p> <p><strong>MarFERReT.v1.1.1.best_pfam_annotations.csv.gz<br></strong>This Gzip-compressed CSV file contains the best-scoring Pfam annotation for intra-species clustered protein sequences from the 800 validated MarFERReT entries; derived from the hmmsearch annotations against Pfam 34.0 functional domains. This file contains the following fields:</p> <ol> <li><strong>aa_id</strong>: The unique MarFERReT protein sequence ID ('mftX').</li> <li><strong>pfam_name</strong>: The shorthand Pfam protein family name.</li> <li><strong>pfam_id</strong>: The Pfam identifier.</li> <li><strong>pfam_eval</strong>: hmm profile match e-value score</li> <li><strong>pfam_score:</strong> hmm profile match bitscore</li> </ol> <p><br><strong>MarFERReT.v1.1.1.dmnd</strong><br>This binary file is the indexed database of the MarFERReT protein library with embedded NCBI taxonomic information generated by the DIAMOND makedb tool using the build_diamond_db.sh script from the MarFERReT /scripts/ library. This can be used as the reference DIAMOND database for annotating environment sequences from eukaryotic metatranscriptomes. <br><br></p>
STREAM - Sub-THz Radar sensing of the Environment for future Autonomous Marine platforms: Multi-Perspective Sensing - Maritime Environment - Side-looking Perspective
<p>This dataset contains the files corresponding to which results have been included in the journal paper titled 'High-Resolution Multi-Modal Sensing of Distributed Radar Network'. The full description of the conducted trials and data structure is mentioned in the attached PDF document.</p> <p>The trials were conducted at the Gosport Marina, Portsmouth, UK with a sea state of approximately 3 according to the Douglas Scale.</p> <p>The experiments were performed with automotive radars operating in the 79 GHz band to investigate the Doppler and imaging capabilities of these radars. A multi-sensory suite distributed around Valkyrie VI was mounted in front, corner, side and backward-looking orientations.</p> <p>This dataset contains data from the side-looking orientation, where the installation angle of radar is 90 degrees respective to the platform velocity vector.</p> <p><strong>Radar Data:</strong></p> <p>The radar data is stored in the file 'GM2_Out1_240522_160925.h5'. The methodology to process the data in MATLAB is presented in the attached pdf. document.</p> <p><strong>Inertial Measurement Unit:</strong></p> <p>Three xSens 680G IMU were mounted on the roof, front and back of the boat. They have been included in the corresponding zip folders.</p> <p>PC3_Corner_RLG: IMU at the corner of the boat.</p> <p>PC4_Forward_RLG: IMU at the roof of the boat.</p> <p>PC5_Backward_RLG: IMU at the back of the boat.</p> <p>The IMU data is converted to .txt files that can be directly loaded into MATLAB.</p> <p><strong>Timestamped Velocity:</strong></p> <p>The file 'Corner_160925.mat' contains the time-stamped velocity for each radar frame. Here, the integration interval is 128 ms with 512 radar chirps.</p> <p>The file 'CommonFramesCornner_160925.mat' contains the timestamped velocity for the frames that are synchronised with the frames of front-looking radar.</p> <p>(The dataset for the front-looking radar is stored in another repository with DOI: 10.5281/zenodo.14215115)</p> <p><strong>Camera:</strong></p> <p>Each radar also has a camera for ground truth. The time-stamped camera frames for each radar frame are stored in 'CommonFramesCornner_160925.mat'.</p> <p>Processed camera frames and video of the scene are available in: 'GM2_Corner_240522_160925_CameraFrames.zip'.</p> <p> </p> <p>For more information, please contact:</p> <p>Anum Pirkani: a.a.a.pirkani@bham.ac.uk, anum.apirkani@gmail.com</p> <p>Marina Gashinova: m.s.gashinova@bham.ac.uk</p>
STREAM - Sub-THz Radar sensing of the Environment for future Autonomous Marine platforms: Multi-Perspective Sensing - Maritime Environment - Front-looking Perspective
<p>This dataset contains the files corresponding to which results have been included in the journal paper titled 'High-Resolution Multi-Modal Sensing of Distributed Radar Network'. The full description of the conducted trials and data structure is mentioned in the attached PDF document.</p> <p>The trials were conducted at the Gosport Marina, Portsmouth, UK with a sea state of approximately 3 according to the Douglas Scale.</p> <p>The experiments were performed with automotive radars operating in the 79 GHz band to investigate the Doppler and imaging capabilities of these radars. A multi-sensory suite distributed around Valkyrie VI was mounted in front, corner, side and backward-looking orientations.</p> <p>This dataset contains data from the front-looking orientation, where the installation angle of radar is 0 degrees respective to the platform velocity vector.</p> <p><strong>Radar Data:</strong></p> <p>The radar data is stored in the file 'GM2_Lab_240522_160943.h5'. The methodology to process the data in MATLAB is presented in the attached pdf. document.</p> <p><strong>Inertial Measurement Unit:</strong></p> <p>Three xSens 680G IMU were mounted on the roof, front and back of the boat. They have been included in the corresponding zip folders.</p> <p>PC3_Corner_RLG: IMU at the corner of the boat.</p> <p>PC4_Forward_RLG: IMU at the roof of the boat.</p> <p>PC5_Backward_RLG: IMU at the back of the boat.</p> <p>The IMU data is converted to .txt files that can be directly loaded into MATLAB.</p> <p><strong>Timestamped Velocity:</strong></p> <p>The file 'Front_160943.mat' contains the time-stamped velocity for each radar frame. Here, the integration interval is 128 ms with 512 radar chirps.</p> <p>The file 'CommonFramesFront_160943.mat' contains the timestamped velocity for the frames that are synchronised with the frames of side-looking radar.</p> <p>(The dataset for the side-looking radar is stored in another repository with DOI: 10.5281/zenodo.14174138)</p> <p><strong>Camera:</strong></p> <p>Each radar also has a camera for ground truth. The time-stamped camera frames for each radar frame are stored in 'CommonFramesFront_160943.mat'.</p> <p>Processed camera frames and video of the scene are available in: 'GM2_Front_240522_160943_CameraFrames.zip'.</p> <p> </p> <p>For more information, please contact:</p> <p>Anum Pirkani: a.a.a.pirkani@bham.ac.uk, anum.apirkani@gmail.com</p> <p>Marina Gashinova: m.s.gashinova@bham.ac.uk</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.