Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,764
datasets available to search
ShareScore release 0.9.0
Dataset results
4,764 results for “file”
A multi-year DAILY weather file for the Toolik Field Station at Toolik Lake, Alaska starting 1988 to present.
A multi-year DAILY weather file for the Arctic Tundra Long-Term Ecological Research (LTER) site at Toolik Lake, AK. Included are daily averages and/or maximums and minimums of air, wind speed, soil temperature, and sum of global radiation and precipitation. In 2008 Toolik Field Station took over maintenance of the main weather station. See http://toolik.alaska.edu/edc/index.php for current weather data. In addition to the main weather station the Arctic LTER maintains several stations that collect data on the experimental plots.
SBC LTER: Time series of quarterly NetCDF files of kelp biomass in the canopy from Landsat 5, 7 and 8, since 1984 (ongoing)
This data file represents a time series of canopy area of giant kelp, Macrocystis pyrifera, and bull kelp, Nereocystis luetkeana, and canopy biomass of giant kelp derived from Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper Plus (ETM+), Landsat 8 Operational Land Imager (OLI), and Landsat 9 Operational Land Imager 2 satellite imagery, along with relevant metadata. The kelp canopy is composed of the portions of fronds and stipes floating on the surface of the water. Canopy area (m) data are given for individual 30 x 30 meter pixels for all coastal areas of Baja California, Mexico, California, Oregon, and the outer coast of Washington (including offshore islands). Biomass data (wet weight, kg) are given for individual 30 x 30 meter pixels in the coastal areas extending from near Ano Nuevo, CA through the southern range limit in Baja California (including offshore islands), representing the range where giant kelp is the dominant canopy forming species. Data were derived from the three Landsat sensors listed above. Observations are made on a 16 day repeat cycle, for each instrument, but the temporal coverage is irregular because of cloud cover, instrument failure, and the mission length of each sensor (TM: 1984 – 2011, ETM+: 1999 – present, OLI: 2013 – present). Estimates of canopy area are derived from the fractional cover of kelp canopy determined from satellite surface reflectance. Estimates of kelp canopy biomass are derived from the relationship between giant kelp fractional cover determined from satellite surface reflectance and empirical measurements of giant kelp canopy biomass in long-term SBC LTER study plots obtained using SCUBA. The different Landsat sensors were calibrated to each other using simulated Landsat data derived from hyperspectral imagery. Missing data due to the ETM+ scan line corrector error were filled using a synchrony-based gap filling method. Data are organized into a single NetCDF file and contain the quarterly area and
Storm Database Files for CLIMK–WINDS: A New Database of Extreme European Winter Windstorms
<p>This database is comprised of the four netCDF files containing the 50 most extreme European winter windstorms identified within the four input sources, with one netCDF file per source: ERA5 reanalysis, CCLM_ERA5_EUR-11 regional climate model simulation, COSMO-REA6 reanalysis, and CCLM_ERA5_CEU-3 regional climate model. This database was created by Clare Marie Flynn and its creation is described in the following paper: Flynn, C. M., Moemken, J., Pinto, J., Schutte, M., and Messori, G.: CLIMK–WINDS: A New Database of Extreme European Winter Windstorms, under review for final submission, Earth System Science Data, 2025.</p>
S1000 corpus, large-scale tagging results and other supplementary files
<p>Data associated with the S1000 corpus</p><p>The tagger software for which the dictionary files in <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/tagger-organisms-dictionary-S1000.tar.gz">tagger-organisms-dictionary-S1000.tar.gz </a>can be used with can be found here: <a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a></p><p>The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/s1000-corpus-annotation-guidelines/</a></p><p>The S1000 corpus split in training, development and test sets in BRAT format can be found in <a href="https://zenodo.org/api/records/10285825/files/S1000-corpus.tar.gz">S1000-corpus.tar.gz</a><a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-corpus.tar.gz?versionId=ac7ce430-c265-49bb-8c8f-9b5f8e271cbe"> </a>and in CoNLL format here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/s1000-conll.tar.gz">s1000-conll.tar.gz</a></p><p>The tagging results of Jensenlab tagger for the S1000 test set are here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000-jensenlab-tagger.tar.gz?versionId=d8d9c9f5-ee3b-4738-aefa-a4a95475d25d">S1000-jensenlab-tagger.tar.gz</a></p><p>The result from the large scale run in entire PubMed and PMC Open Access articles for Jensenlab tagger is provided here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz?versionId=48825928-9fc9-423c-8a4c-4f8994e95805">Jensenlab_tagger_large_scale_matches_with_rank.tsv.gz</a></p><p>The model used for the large scale run of the transformer-based method is here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/S1000_Transformer_based_tagger_large_scale_model.tar.gz?versionId=8e974f64-9abc-4449-a377-e3f97e91d612">S1000_Transformer_based_tagger_large_scale_model.tar.gz</a> and the results from the large scale tagging here: <a href="https://zenodo.org/api/files/b8a0e221-3cc3-4db5-a2e9-f19a1bd2e5cb/Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip?versionId=dc21a6ba-9763-4130-9f02-0341a885c692">Transformer_based_tagger_large_scale_matches_with_rank.tsv.zip</a></p>
Dataset and program scripts for the reproducibility of the hierarchical data structure file. Related to the manuscript entitled: Hierarchical Representation of Measurement Data, Metrological Uncertainty and Metadata for Calibrated Battery Tests
<p>We present an interoperable hierarchical data representation for battery tests, leading to improved scalability of data transmission and enhanced data accessibility and comprehensibility for both human interpretation and machine processing. The hierarchical data format includes the raw trace electrical measurement data, the metrological calibration and uncertainty data, the metadata such as experimental settings, instruments and software versions, as well as post-processed data such as electrochemical model fit parameters. This data representation allows repetition of the battery test under the exact same conditions such that identical results are achieved within defined error bounds. This is in line with the general F.A.I.R. data approach and provides repeatability and traceability in the battery value chain. As an application of the hierarchical data representation, we show the classification of cells as pass/fail being performed with quantitative confidence levels. We demonstrate the complete workflow of establishing the hierarchical data structure for electrochemical impedance spectroscopy (EIS), starting from metrological traceability of the calibration and uncertainty analysis towards the storage of the structured data as a single integrated file that preserves the hierarchical data format.</p>
Example Microscopy Metadata JSON files produced using Micro-Meta App to document example microscopy experiments performed at individual core facilities
<p>Example <strong>Microscopy Metadata </strong>(Microscope.JSON and Settings.JSON)<strong> files </strong>produced using<strong> <a href="https://wu-bimac.github.io/MicroMetaApp.github.io/">Micro-Meta App</a> </strong>to document the <strong>Hardware Specifications</strong> of example Microscopes and the <strong>Image Acquisition Settings</strong> utilized to acquire example images as listed in the table below.</p> <blockquote> <p>For each facility, the dataset contains two JSON files:</p> <ol> <li><strong>Microscope.JSON file</strong> (e.g., 01_marcello_uliverpool_cci_zeiss_axioobserz1_lsm710.json)</li> <li><strong>Settings.JSON file</strong> (indicated with the name of the image and with the _AS suffix)</li> </ol> </blockquote> <p><strong>Micro-Meta App was</strong> developed as part of a <strong>global community initiative</strong> including the <a href="http://www.4dnucleome.org/"><strong>4D Nucleome (4DN)</strong> </a>Imaging Working Group, <strong>BioImaging North America (BINA)</strong> <a href="https://www.bioimagingna.org/qc-dm-wg">Quality Control and Data Management Working Group</a>, and <strong>QUAlity and REProducibility for Instrument and Images in Light Microscopy</strong> (<a href="https://quarep.org/"><strong>QUAREP-LiMi</strong></a>), to extend the <strong>Open Microscopy Environment (OME)</strong> <a href="https://www.openmicroscopy.org/Schemas/Documentation/Generated/OME-2016-06/ome.html">data model</a>.</p> <blockquote> <p>The works of this <strong>global community effort</strong> resulted in multiple publications featured on a recent <strong>Nature Methods FOCUS ISSUE </strong>dedicated to <a href="https://www.nature.com/collections/djiciihhjh">Reporting and reproducibility in microscopy</a>.</p> </blockquote> <blockquote> <p><strong>Learn More!</strong> For a thorough description of <strong>Micro-Meta App</strong> consult our recent <a href="https://doi.org/10.1038/s41592-021-01315-z">Nature Methods</a> and <a href="https://doi.org/10.1101/2021.05.31.446382">BioRxiv.org</a> publications!</p> </blockquote> <p> </p> <table> <tbody> <tr> <td><strong>Nr.</strong></td> <td><strong>Manufacturer</strong></td> <td><strong>Model</strong></td> <td><strong>Tier</strong></td> <td><strong>Εxperiment Type</strong></td> <td><strong>Facility Name</strong></td> <td><strong>Department and Institution</strong></td> <td><strong>URL</strong></td> <td><strong>References</strong></td> </tr> <tr> <td>1</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Observer Z1 (with LSM 710 scan head)</strong></td> <td>1</td> <td>3D visualization of superhydrophobic polymer-nanoparticles</td> <td>Centre for Cell Imaging (CCI)</td> <td>University of Liverpool</td> <td>https://cci.liv.ac.uk/equipment_710.html</td> <td>Upton et al., 2020</td> </tr> <tr> <td>2</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Observer (Axiovert 200M)</strong></td> <td>2</td> <td>Μeasurement of illumination stability on Chinese Hamster Ovary cells expressing Paxillin-EGFP</td> <td>Advanced BioImaging Facility (ABIF).</td> <td>McGill University</td> <td>https://www.mcgill.ca/abif/equipment/axiovert-1</td> <td>Kiepas et al., 2020</td> </tr> <tr> <td>3</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Observer Z1 (with Spinning Disk)</strong></td> <td>2</td> <td>Immunofluorescence imaging of cryosection of Mouse kidney</td> <td>Imagerie Cellulaire; Quality Control managed by Miacellavie (https://miacellavie.com/)</td> <td>Centre de recherche du Centre Hospitalier Université de Montréal (CR CHUM), University of Montreal</td> <td>https://www.chumontreal.qc.ca/crchum/plateformes-et-services (the web site is for all core facilities, not specifically for the core facility hosting this microscope)</td> <td>Pilliod et al., 2020</td> </tr> <tr> <td>4</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Imager Z2 (with Apotome)</strong></td> <td>2</td> <td>Immunofluorescence imaging of mitotic division in Hela cells using </td> <td>Bioimaging Unit</td> <td>Newcastle University</td> <td>https://www.ncl.ac.uk/bioimaging/</td> <td>Watson et al., 2020</td> </tr> <tr> <td>5</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Observer Z1</strong></td> <td>2</td> <td>Fluorescence microscopy of human skin fibroblasts from Glycogen Storage Disease patients.</td> <td>Life Imaging Center (LIC)</td> <td>Centre for Integrative Signalling Analysis (CISA), University of Freiburg</td> <td>https://miap.eu/equipments/sd-i-abl/</td> <td>Hannibal et al., 2020</td> </tr> <tr> <td>6</td> <td><strong>Leica Microsystems</strong></td> <td><strong>DMI6000B</strong></td> <td>2</td> <td>3D immunofluorescence imaging rhinovirus infected macrophages </td> <td>IMAG'IC Confocal Microscopy Facility</td> <td>Institut Cochin, CNRS, INSERM, Université de Paris</td> <td>https://www.institutcochin.fr/core_facilities/confocal-microscopy/cochin-imaging-photonic-microscopy/organigram_team/10054/view</td> <td>Jubrail et al., 2020</td> </tr> <tr> <td>7</td> <td><strong>Leica Microsystems</strong></td> <td><strong>DM5500B</strong></td> <td>2</td> <td>Immunofluorescence analysis of the colocalization of PML bodies with DNA double-strand breaks</td> <td>Bioimaging Unit</td> <td>Edwardson Building on the Campus for Ageing and Vitality, Newcastle University</td> <td>https://www.ncl.ac.uk/bioimaging/equipment/leica-dm5500/#overview</td> <td>da Silva et al., 2019; Nelson et al., 2012<br> </td> </tr> <tr> <td>8</td> <td><strong>Leica Microsystems</strong></td> <td><strong>DMI8-CS (with TCS SP8 STED 3X)</strong></td> <td>2</td> <td>Live-cell imaging of N. benthamiana leaves cells-derived protoplasts</td> <td>Center for Advanced Imaging (CAi)</td> <td>School of Mathematics/Natural Sciences, Heinrich-Heine-Universität Düsseldorf</td> <td>https://www.cai.hhu.de/en/equipment/super-resolution-microscopy/leica-tcs-sp8-sted-3x</td> <td>Singer et al., 2017; Hänsch et al., 2020</td> </tr> <tr> <td>9</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti</strong></td> <td>2</td> <td>Immunofluorescence analysis of the cytoskeleton structure in COS cells</td> <td>Advanced Imaging Center (AIC)</td> <td>Janelia Research Campus, Howard Hughes Medical Institute</td> <td>https://www.janelia.org/support-team/light-microscopy/equipment</td> <td>Abdelfattah et al., 2019; Qian et al., 2019; Grimm et al., 2020</td> </tr> <tr> <td>10</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti-E (HCA)</strong></td> <td>2</td> <td>Τime-lapse analysis of the bursting behavior of amine-functionalized vesicular assemblies</td> <td>Light Microscopy Facility (IALS-LIF)</td> <td>Institute for Applied Life Sciences, University of Massachusetts at Amherst</td> <td>https://www.umass.edu/ials/light-microscopy</td> <td>Fernandez et al., 2020</td> </tr> <tr> <td>11</td> <td><strong>Nikon Instruments/Coleman laboratory (customized)</strong></td> <td><strong>TIRF HILO Epifluorescence light Microscope (THEM)/ Eclipse Ti</strong></td> <td>2</td> <td>Single-particle tracking of Halo-tagged PCNA in Lox cells</td> <td>Coleman laboratory</td> <td>Anatomy and Structural Biology Department, The Albert Einstein College of Medicine</td> <td>https://einsteinmed.org/faculty/12252/robert-coleman/</td> <td>Drosopoulos et al., 2020</td> </tr> <tr> <td>12</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti (with Andor Dragon Fly Spinning Disk)</strong></td> <td>2</td> <td>Investigation of the 3D structure of cerebral organoids</td> <td>Montpellier Resources Imagerie</td> <td>Centre de Recherche de Biologie cellulaire de Montpellier (MRI-CRBM), CNRS, Univerity of Montpellier</td> <td>https://www.mri.cnrs.fr/en/optical-imaging/our-facilities/mri-crbm.html</td> <td>Ayala-Nunez et al., 2019</td> </tr> <tr> <td>13</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti2</strong></td> <td>2</td> <td>Ιmmunofluorescence imaging of cryosections of mouse hearth myocardium </td> <td>Neuroscience Center Microscopy Core</td> <td>Neuroscience Center, University of North Carolina</td> <td>https://www.med.unc.edu/neuroscience/core-facilities/neuro-microscopy/</td> <td>Aghajanian et al., 2021</td> </tr> <tr> <td>14</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti2</strong></td> <td>2</td> <td>Live-cell imaging of bacterial cells expressing GFP-PopZ</td> <td>Microscopy Resources on the North Quad (MicRoN)</td> <td>Harvard Medical School </td> <td>https://micron.hms.harvard.edu/</td> <td>Lim and Bernhardt 2019; Lim et al., 2019</td> </tr> <tr> <td>15</td> <td><strong>Olympus/Biomedical Imaging Group (customized)</strong></td> <td><strong>TIRF Epifluorescence Structured light Microscope (TESM)/IX71</strong></td> <td>3</td> <td>3D distribution of HIV-1 in the nucleus of human cells</td> <td>Biomedical Imaging Group</td> <td>Program in Molecular Medicine, University of Massachusetts Medical School</td> <td>https://trello.com/b/BQ8zCcQC/tirf-epi-fluorescence-structured-light-microscope</td> <td>Navaroli et al., 2012</td> </tr> <tr> <td>16</td> <td><strong>Olympus/Computer Vision Laboratory (customized)</strong></td> <td><strong>3D BrightField Scanner/IX71</strong></td> <td>3</td> <td>Transmitted light brightfield visualization of swimming spermatocytes</td> <td>Laboratorio Nacional de Microscopia Avanzada (LNMA) and Computer Vision Laboratory of the Institute of Biotechnology</td> <td>Universidad Nacional Autonoma de Mexico (UNAM)</td> <td>https://lnma.unam.mx/wp/</td> <td>Pimentel et al., 2012; Silva-Villalobos et al., 2014</td> </tr> </tbody> </table> <p><strong>Getting started</strong></p> <p>Use these videos to get started with using Micro-Meta App after installation into OMERO and downloading the example data files:</p> <ol> <li><a href="https://vimeo.com/562022222">Video 1</a></li> <li><a href="https://vimeo.com/562022281">Video 2</a></li> </ol> <p><strong>More information</strong></p> <blockquote> <p>For full information on how to use Micro-Meta App please utilize the following resources:</p> <ol> <li>Micro-Meta App <a href="https://wu-bimac.github.io/MicroMetaApp.github.io/">website</a></li> <li><a href="https://micrometaapp-docs.readthedocs.io/en/latest/index.html">Full documentation</a></li> <li><a href="https://micrometaapp-docs.readthedocs.io/en/latest/docs/intro/installation.html">Installation</a> instructions</li> <li><a href="https://micrometaapp-docs.readthedocs.io/en/latest/docs/tutorials/index.html#step-by-step-instructions">Step-by-Step Instructions</a></li> <li><a href="https://micrometaapp-docs.readthedocs.io/en/latest/docs/tutorials/VideoTutorials.html#micro-meta-app-video-tutorials">Tutorial Videos</a></li> </ol> </blockquote> <p><strong>Background</strong></p> <p>If you want to learn more about the importance of <strong>metadata and quality contro</strong>l to ensure full <strong>reproducibility, quality and scientific value</strong> in light microscopy, please take a look at our recent publications describing the development of community-driven light <strong>4DN-BINA-OME Microscopy Metadata</strong> specifications <a href="https://doi.org/10.1038/s41592-021-01327-9">Nature Methods</a> and <a href="https://doi.org/10.1101/2021.04.25.441198">BioRxiv.org</a> and our <a href="https://arxiv.org/abs/1910.11370">overview manuscript</a> entitled <strong>A perspective on Microscopy Metadata: data provenance and quality control</strong>.</p> <p> </p> <p> </p>
Example Microscopy Metadata JSON files produced using Micro-Meta App to document the acquisition of example images using a custom-built TIRF Epifluorescence Structured Illumination Microscope
<p><strong>Example Microscopy Metadata JSON files produced using the <a href="https://wu-bimac.github.io/MicroMetaApp.github.io/">Micro-Meta App</a> documenting an example raw-image file acquired using the custom-built TIRF Epifluorescence Structured Illumination Microscope.</strong></p> <p>For this use case, which is presented in Figure 5 of <a href="http://doi: https://doi.org/10.1101/2021.05.31.446382">Rigano et al., 2021</a>, Micro-Meta App was utilized to document:</p> <p>1) The <strong>Hardware Specifications</strong> of the custom build TIRF Epifluorescence Structured light Microscope (TESM; <a href="https://www.pnas.org/content/109/8/E471.long">Navaroli et al., 2010</a>) developed, built on the basis of the based on Olympus IX71 microscope stand, and owned by the Biomedical Imaging Group (http://big.umassmed.edu/) at the Program in Molecular Medicine of the University of Massachusetts Medical School. Because TESM was custom-built the most appropriate documentation level is <strong>Tier 3</strong> (<em>Manufacturing/Technical Development/Full Documentation</em>) as specified by the <a href="https://doi.org/10.5281/zenodo.4710731">4DN-BINA-OME</a> Microscopy Metadata model (<a href="https://doi.org/10.1101/2021.04.25.441198">Hammer et al., 2021</a>).</p> <p>The TESM Hardware Specifications are stored in: <strong>Rigano et al._Figure 5_UseCase_Biomedical Imaging Group_TESM.JSON</strong></p> <p>2) The <strong>Image Acquisition Settings</strong> that were applied to the TESM microscope for the acquisition of an example image (FSWT-6hVirus-10minFIX-stk_4-EPI.tif.ome.tif) obtained by Nicholas Vecchietti and Caterina Strambio-De-Castillia. For this image, TZM-bl human cells were infected with HIV-1 retroviral three-part vector (FSWT+PAX2+pMD2.G). Six hours post-infection cells were fixed for 10 min with 1% formaldehyde in PBS, and permeabilized. Cells were stained with mouse anti-p24 primary antibody followed by DyLight488-anti-Mouse secondary antibody, to detect HIV-1 viral Capsid. In addition, cells were counterstained using rabbit anti-Lamin B1 primary antibody followed by DyLight649-anti-Rabbit secondary antibody, to visualize the nuclear envelope and with DAPI to visualize the nuclear chromosomal DNA.</p> <p>The Image Acquisition Settings used to acquire the FSWT-6hVirus-10minFIX-stk_4-EPI.tif.ome.tif image are stored in: <strong>Rigano et al._Figure 5_UseCase_AS_fswt-6hvirus-10minfix-stk_4-epi.tif.JSON</strong></p> <p><em><strong>Instructional video tutorials on how to use these example data files:</strong></em><br> Use these videos to get started with using Micro-Meta App after downloading the example data files available here.</p> <ul> <li><a href="https://vimeo.com/562022222">Part 1/2</a></li> <li><a href="https://vimeo.com/562022281">Part 2/2</a></li> </ul>
Input files for simulation of potassium channels using the AMOEBA polarizable force field
<p>This dataset contains input Tinker xyz and key files for the simulation of KcsA potassium channels in DOPC bilayer, a simple script for converting CHARMM pdb file to Tinker xyz file, and modified Tinker source code to support one-dimensional position restraints.<br> "params.tar.gz" contains a description of the force field modifications.<br> <br> To use "mod2", add the following lines to the key file.</p> <pre><code>#compatible with amoebabio18.prm polarize 5 1.4500 0.3900 3 polarize 11 1.4500 0.3900 9 polarize 3 1.7500 0.3900 1 5 7 50 225 227 polarize 9 1.7500 0.3900 1 7 11 50 225 227</code></pre> <p> </p>
A novel approach to the detection of unusual mitochondrial protein change suggests hypometabolism of ancestral simians: Supplemental Files
<p><strong>Supplementary Fig. S1</strong>: θ<sub>evo</sub> calculated for each analyzed edge for specific OXPHOS complexes. Analyses were performed as in fig. 1F, except that SPCSs calculated from mtDNA-encoded protein positions in Complex I, Complex III, Complex IV, or Complex V were used to generate θevo values.</p> <p><strong>Supplementary Fig. S2</strong>: Mammalian orders differ in their propensity for potentially efficacious mitochondrial protein substitutions within specific OXPHOS complexes (median calculations). Analysis was performed as in fig. 2A, except that θ<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V polypeptides.</p> <p><strong>Supplementary Fig. S3</strong>: Mammalian orders differ in their propensity for potentially efficacious mitochondrial protein substitutions within specific OXPHOS complexes (median confidence intervals). Analysis was performed as in (<em>A</em>) fig. 2B or (<em>B</em>) fig. 2C, except that θ<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V proteins.</p> <p><strong>Supplementary Fig. S4</strong>: Mammalian families differ in their propensity for potentially efficacious mitochondrial protein substitutions at specific OXPHOS complexes (median calculations). Analysis was performed as in fig. 3A, except that θ<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V subunits.</p> <p><strong>Supplementary Fig. S5</strong>: Mammalian families differ in their propensity for potentially efficacious mitochondrial protein substitutions at specific OXPHOS complexes (median confidence intervals ordered by lower 90% median confidence limit). Analysis was performed as in fig. 3B, except that θ<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V proteins.</p> <p><strong>Supplementary Fig. S6</strong>: Mammalian families differ in their propensity for potentially efficacious mitochondrial protein substitutions at specific OXPHOS complexes (median confidence intervals ordered by upper 90% median confidence limit). Analysis was performed as in fig. 3C, except that θ<sub>evo</sub> values were obtained by analysis of mtDNA-encoded Complex I, Complex III, Complex IV, or Complex V polypeptides.</p> <p>---</p> <p><strong>Supplementary File 1</strong>: All predicted protein substitutions along all edges at positions containing less than 2% gaps across input and ancestral sequences are listed, along with associated taxonomy information, TSS, and branch length. All alignment positions refer to Bos taurus reference sequences.</p> <p><strong>Supplementary File 2</strong>: The TSS calculated for each mitochondrial protein alignment position. All alignment positions refer to Bos taurus reference sequences.</p> <p><strong>Supplementary File 3</strong>: SPCS and θevo outputs are provided for analyses across all mitochondria-encoded positions, as well as for focused analyses of specific OXPHOS complexes and individual proteins.</p> <p><strong>Supplementary File 4</strong>: A GenBank flat file containing RefSeq entries for mammalian mtDNAs, as well as the entry for the reptile Anolis punctatus.</p> <p><strong>Supplementary File 5</strong>: A maximum likelihood inferred tree generated by a RAxML-NG analysis of concatenated and aligned protein coding sequences from mammalian and Anolis punctatusmtDNAs.</p> <p><strong>Supplementary File 6</strong>: Bootstrap replicates were generated from the alignment of concatenated protein coding sequences. Felsenstein’s Bootstrap Proportions (Felsenstein 1985) were calculated and used to label the maximum likelihood inferred tree of mammalian mtDNAs.</p> <p><strong>Supplementary File 7</strong>: Bootstrap replicates were generated using concatenated mammalian mtDNA coding sequences. Transfer Bootstrap Expectations (Lemoine 2018) were calculated and used to label the maximum likelihood inferred tree of mammalian mtDNAs.</p> <p><strong>Supplementary File 8</strong>: PAGAN tree output produced using aligned amino acid sequences and the rooted maximum likelihood inferred tree as input.</p>
Malawi probabilistic seismic hazard analysis (PSHA) using the Malawi Seismogenic Source Model (MSSM). Supplementary Files v1.1
<p>Updated (October 2022) version of supplementary files for running probabilistic seismic hazard analysis (PSHA) MATLAB codes for Malawi. The PSHA codes themselves (v1.0) are available at: https://doi.org/10.5281/zenodo.7265781and the most recent version will be available on GitHub at: https://github.com/jack-williams1/Malawi_PSHA. Note the variables stored here are not stored on GitHub due to the file size.</p> <p>Includes both input files for performing PSHA and output ground motions for plotting PSHA results.</p> <p>Files are:</p> <ul> <li>malawi_Vs30_active.txt: Input USGS slope-based Vs30 values for Malawi (Wald and Allen 2007)</li> <li>EQCAT_comb.mat: MSSM Direct catalog for all possible rupture weightings (stored as MATLAB variable)</li> <li>GM_MSSM_em_20221027: Ground motions for plotting PSHA maps (stored as MATLAB variable)</li> <li>GM_MSSM_20221021.mat: Ground motions needed for plotting PSHA-site analysis figures (stored as MATLAB variable)</li> <li>mssm_comb.mat: Matlab file for combined MSSM Direct and Adapted MSSM catalogs (stored as MATLAB variable)</li> <li>MSSM_Catalog_Adapted_em.mat: Adapated MSSM event catalog (stored as MATLAB variable)</li> <li>syncat_bg.mat: Areal source stochastic event catalog (stored as MATLAB variable)</li> </ul> <p>Further descriptions of these files and how to use them are provided on Github. An open-access manuscript describing the PSHA is available at: </p> <p>Williams J. N., Werner M. J., Goda K., Wedmore L. N. J., De Risi R., Biggs J., Mdala H., Dulanya Z., Fagereng Å, Mphepo F., Chindandali P. (2023). Fault-based probabilistic seismic hazard analysis in regions with low strain rates and a thick seismogenic layer: a case study from Malawi, Geophysical Journal International, Volume 233, Issue 3, June 2023, Pages 2172–2206, <a href="https://doi.org/10.1093/gji/ggad060">https://doi.org/10.1093/gji/ggad060</a></p> <p>Please reference this publication along with this repository when using these data.</p> <p>USGS vs30 value compilation described in:</p> <p>Allen, T. I., and Wald, D. J., 2009, On the use of high-resolution topographic data as a proxy for seismic site conditions (Vs30), Bulletin of the Seismological Society of America, 99, no. 2A, 935-943.</p> <p> </p>
MesoLF demo data and auxiliary files
<p>Demo data and auxiliary files accompanying the article:</p> <p>Nöbauer, T., Zhang, Y., Kim, H. & Vaziri, A.<br> Mesoscale volumetric light-field (MesoLF) imaging of neuroactivity across cortical areas at 18 Hz.<br> <em>Nature Methods</em> 1–10 (2023). doi:<a href="https://doi.org/10.1038/s41592-023-01789-z">10.1038/s41592-023-01789-z</a><br> <br> The files provided here are required for running a demo of the MesoLF pipeline. These files will be downloaded automatically by the Matlab live notebook "mesolf_demo.mlx" that was published as part of "Supplementary Software 1" with the associated article. For installation instructions, see file "README.md" in "Supplementary Software 1". For future software updates, check <a href="https://github.com/vazirilab">https://github.com/vazirilab</a></p>
2-Dimensional habitat files for 47 representative marine species
<p><strong>2D marine species habitats in NetCDF format on 0.5*0.5 degree global regular grid.</strong></p> <p>Based on Close et al. (2006) and converted from .CSV format. </p> <p>Filename is in format 'presence_speciesnumber.nc' where species numbers are listed in the README.txt file (identifier for each species). The README.txt file further contains each species' species_group which is the assigned depth group for each species (1=0-200m epipelagic, 2=200-1000m mesopelagic, 3=sea floor demersal) and species_name which is the Latin name of each species with underscore in between.</p> <p>The species occurs where the variable 'presence' equals 1 (in the accompanying paper we assume this to be the 1995-2014 climatological mean distribution).</p> <p>In the NetCDF files, the variable 'presence' has as an attribute 'species' which contains the Latin species name without underscore.</p>
Arctic Grayling length, weight and tag data from Arctic LTER Streams project, Toolik Filed Station Alaska, 1985 to 2018
Since 1983, the Streams Project at the Toolik Field Station has monitored physical, chemical, and biological parameters in a 5-km, fourth-order reach of the Kuparuk River near its intersection with the Dalton Highway and the Trans-Alaska Pipeline. In 1989, similar studies were begun on a 3.5-km, third-order reach of a second stream, Oksrukuyik Creek. Fish were collected on each river. Station locations, representing kilomter values certain distances from original phosphorus dripper (see method) were noted. 1985 to 2012 long-term tagging file for Arctic Grayling (Thymallus arcticus) on the Kuparuk River. All grayling adults and juveniles captured during the field season are measured, weighed, tagged and released. Grayling were tagged originally with a colored tag with a number. In 1993, researchers started pit tagging the grayling. These pit tags can be read with an antenna to track the migration of the grayling throughout the Kuparuk River system. Arctic grayling young-of-the-year (YOY) were caught multiple times during each summer and measured and weighed as well. This file combines the data from the following data sets: Dataset ID Short name 10325 1985-2012_Kuparuk_Grayling_Tags 10327 1986-2012_Kuparuk_YOY 10329 1989-2011_Oksrukuyik_Grayling_Tags 10330 1989-2012_Oksrukuyik_YOY
Hourly weather data from the Arctic LTER Moist Acidic Tussock Experimental plots from 2011 to present, Toolik Filed Station, North Slope, Alaska.
Hourly weather data from the LTER Moist Acidic Tussock Experimental plots. The station was installed in 1990 in block 2 of the Toolik LTER experimental moist acidic tussock plots. The plots are located on a hillside near Toolik Lake (68 38' N, 149 36'W). Global solar radiation, photosynthetic active radiation, unfrozen precipitation, air temperature, relative humidity, wind speed, and wind direction are measured at 3 meters. Additional sensors in greenhouses and shade houses plots measure air temperature, relative humidity and photosynthetic active radiation during the growing season. The sensors are read every minute and averaged or totaled every hour.
Hubbard Brook Experimental Forest: Soil type prediction raster files
This dataset consists of raster files predicting spatial patterns in soils for the entire Hubbard Brook Experimental Forest. Eight soil units are used, following a hydropedologic approach, based on relationships between soil genetic horizon presence and thickness, and the frequency and depth of groundwater fluctuations. Nine raster files on a five-meter grid are presented, including one raster each showing the probability of presence of each of the eight soil units; the ninth raster represents the soil unit most likely to be present at each grid cell. The methods section of the metadata includes descriptions of the eight soil units and guidance for users of the model outputs. These data were gathered as part of the Hubbard Brook Ecosystem Study (HBES). The HBES is a collaborative effort at the Hubbard Brook Experimental Forest, which is operated and maintained by the USDA Forest Service, Northern Research Station.
RAPID Model Input Files for Mekong-Indus-Ganges-Brahmaputra-Megna (MIGBM) River Basins
<p>This database contains Inputs and intermediate files of the RAPID model pre-processor (RRR), and also outputs from the RRR (<em>i.e.</em>, Inputs for RAPID); which were used by <em>Sikder et al.</em> [2019] to assess the performance of available global LSM runoffs in South and Southeast Asian river basins. If you use this RAPID Model Input Files for Mekong-Indus-Ganges-Brahmaputra-Megna (MIGBM) River Basins in your work, please cite: <em>Sikder et al.</em>, [2019], Evaluation of Available Global Runoff Datasets Through a River Model in Support of Transboundary Water Management in South and Southeast Asia, Front. Environ. Sci., 7:171, <a href="https://doi.org/10.3389/fenvs.2019.00171">https://doi.org/10.3389/fenvs.2019.00171</a>.</p> <p>The database contains;</p> <ul> <li>Global River basin and Network Shapefiles: HydroSHEDS.tar.gz</li> <li>Extracted Basin Shapefile: MIGBM_basin.tar.gz</li> <li>Extracted River Network Shapefiles: MIGBM_<strong><em>res</em></strong>_ntwk.tar.gz (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Catchment Files: rapid_catchment_as_<strong><em>riv</em></strong>_res.csv (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Connectivity Files: rapid_connect_<strong><em>res</em></strong>_MIGBM.csv (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Coordinate Files: coords_<strong><em>res</em></strong>_MIGBM.csv (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Base Parameter Files: <strong><em>p</em></strong>fac_<strong><em>res</em></strong>_MIGBM_1km_hour.csv (Note: <strong><em>p</em></strong> = k or x; <strong><em>res</em></strong> = fine or coarse)</li> <li>Sort Files: sort_<strong><em>res</em></strong>_MIGBM_topo.csv (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Sorted Basin Files: riv_bas_id_<strong><em>res</em></strong>_MIGBM_topo.csv (Note: <strong><em>res</em></strong> = fine or coarse)</li> <li>Coupling Files: rapid_coupling.tar.gz</li> <li>Parameter Files: rapid_param.tar.gz</li> <li>Volume Files: m3_riv_<strong><em>res</em></strong>_MIGBM_20000101_20091231_<strong><em>prj</em></strong>_<strong><em>LSMsr</em></strong>_<strong><em>tr</em></strong>_utc.nc (Note: <strong><em>res</em></strong> = fine or coarse; <strong><em>prj</em></strong> = GLDAS or GLDAS.2.0 or GLDAS.2.1 or ECMWF; <strong><em>LSM</em></strong> = CLM, MOS, NOAH, VIC, ERAint; <strong><em>sr</em></strong> = 10 or 025; <strong><em>tr</em></strong> = 3H or D)</li> </ul> <p> </p> <p>Other necessary links associated with this database:</p> <p>RAPID model: <a href="https://github.com/c-h-david/rapid">https://github.com/c-h-david/rapid</a></p> <p>RAPID model pre-processor (rrr): <a href="https://github.com/c-h-david/rrr">https://github.com/c-h-david/rrr</a></p> <p>GLDAS outputs: <a href="https://disc.gsfc.nasa.gov/datasets?keywords=GLDAS">https://disc.gsfc.nasa.gov/datasets?keywords=GLDAS</a></p> <p>ECMWF outputs: <a href="https://www.ecmwf.int/en/forecasts/datasets/reanalysis-datasets/era-interim-land">https://www.ecmwf.int/en/forecasts/datasets/reanalysis-datasets/era-interim-land</a></p> <p> </p> <p>References:</p> <p>Balsamo, G., Albergel, C., Beljaars, A., Boussetta, S., Brun, E., Cloke, H., et al. [2015], ERA-Interim/Land: a global land surface reanalysis data set, Hydrol. Earth Syst. Sci., 19, 389–407, <a href="https://doi.org/10.5194/hess-19-389-2015">https://doi.org/10.5194/hess-19-389-2015</a></p> <p>David, C. H., D. R. Maidment, G. Y. Niu, Z. L. Yang, F. Habets, and V. Eijkhout [2011], River network routing on the NHDPlus dataset, J. Hydrometeorol., 12, 913–934, <a href="https://doi.org/10.1175/2011JHM1345.1">https://doi.org/10.1175/2011JHM1345.1</a></p> <p>Rodell, M., P. R. Houser, U. Jambor, J. Gottschalck, K. Mitchell, C.-J. Meng, et al. [2004], The global land data assimilation system, Bull. Am. Meteorol. Soc. 85, 381–394, <a href="https://doi.org/10.1175/BAMS-85-3-381">https://doi.org/10.1175/BAMS-85-3-381</a></p> <p>Sikder, M. S., C. H. David, G. H. Allen, X. Qiao, E. J. Nelson, and M. A. Matin [2019], Evaluation of Available Global Runoff Datasets Through a River Model in Support of Transboundary Water Management in South and Southeast Asia, Front. Environ. Sci., 7:171, <a href="https://doi.org/10.3389/fenvs.2019.00171">https://doi.org/10.3389/fenvs.2019.00171</a></p>
Input files for Dispa-SET for the JRC report "Power System Flexibility in a variable climate"
<p><strong>Input files for Dispa-SET for the JRC report "Power System Flexibility in a variable climate"</strong></p> <p>Here you can find the input files needed to reproduce the results of the <a href="https://doi.org/10.2760/75312">report</a>:</p> <pre><code>De Felice, M., Busch, S., Kanellopoulos, K., Kavvadias, K. and Hidalgo Gonzalez, I., Power system flexibility in a variable climate, EUR 30184 EN, Publications Office of the European Union, Luxembourg, 2020, ISBN 978-92-76-18183-5 (online), doi:10.2760/75312 (online), JRC120338. </code></pre> <p>The results in the report are generated with the Dispa-SET power system model, available and explained at <a href="https://www.dispaset.eu/">www.dispaset.eu</a>.</p> <p>A description of the data sources with the references can be found into the report.</p> <p><strong>How to use this dataset</strong></p> <p>This dataset can be used as input data for the Dispa-SET model. We refer to the <a href="https://doi.org/10.2760/75312">report</a> and the <a href="https://www.dispaset.eu">official model documentation</a> for information about the data and the model.</p> <p><strong>Description of the dataset</strong></p> <p>The file <code>EnVarClim.yml</code> is a template of the YAML configuration file used by Dispa-SET. To run a specific climate year the <code>XXXX</code> present in some input files must be replaced with the year.</p> <p><strong>Availability factors</strong></p> <p>In the folder <code>AvailabilityFactors</code> there are the availability factors (from 0 to 1) for the power plants and the renewable generation. There is a subfolder for each simulated zone and inside a file for each climate year: from <code>emh_and_cc_availability_1990.csv</code> to <code>emh_and_cc_availability_2015.csv</code>.</p> <p><strong>Cross-border transmission</strong></p> <p>In the folder <code>DayAheadNTC</code> there is the file <code>merged_constant_NTC.csv</code> containing the capacity (in MW).</p> <p><strong>NOTE</strong>: due to an error in the pre-processing code there are some additional lines for the Western Balkans countries ending with a <code>1</code> (e.g. <code>GR -> MK1</code>). Those lines are ignored by the model because are not associated to any simulated zone.</p> <p><strong>Cross-border historical flows</strong></p> <p>In the file <code>CC_L_flows.csv</code> under the folder <code>Flows</code> are contained the hourly flows between the simulated zones and their neighbours (RU, TR, UA).</p> <p><strong>Fuel prices</strong></p> <p>In the folder <code>FuelPrices</code> are contained a set of files containing the hourly prices for the fuels (biomass, coal, lignite, gas, oil) and CO2 emissions. It is worth noting that in spite of their hourly resolution the time-series are constant through the year.</p> <p><strong>Hourly load</strong></p> <p>In the folder <code>Load_RealTime</code> there are hourly load time-series for each zone considering a different climate year. For the Western Balkans countries we use the same time-series for each climate year.</p> <p><strong>Outage factors</strong></p> <p>The files <code>CC_L_outages.csv</code> in the folder <code>OutageFactors</code> contain the outage factor (from 1, full outage, to 0) for the various generation units. Whenever a simulation zone is missing the model assumes the absence of outages.</p> <p><strong>Power plants data</strong></p> <p>In the folder <code>PowerPlants</code> there is a file named <code>CC_L_plants.mip.csv</code> for each simulated zone. The CSV files contain the data <a href="http://www.dispaset.eu/en/latest/data.html#power-plant-data">needed by Dispa-SET</a>.</p> <p><strong>Water storage levels</strong></p> <p>The folder <code>ReservoirLevel</code> contains the storage level (values from 0 to 1 relative to the size of the storage) for all the simulated zones. The levels have been computed for each climate year using a different inflow using the <a href="http://www.dispaset.eu/en/latest/mid_term.html">mid-term scheduler</a> recently implemented in Dispa-SET. For the Western Balkans countries we use the same time-series for each climate year.</p> <p><strong>Hydro-power inflows</strong></p> <p>In the folder <code>ScaledInflows</code> are contained the inflows used for the hydro-power generation. The values in the CSV files describes how much energy is available for hydro-power generation compared to the installed capacity.</p> <p><strong>Linked resources</strong></p> <ul> <li>Model output files:<strong> </strong>https://zenodo.org/record/3778133</li> <li>Source code for the figures: https://github.com/energy-modelling-toolkit/figures-JRC-report-power-system-and-climate-variability</li> </ul>
IPBES Data Management Tutorials - Session 3.4: Data management report details: File formats
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The <em>IPBES data management reports </em>chapter provides an overview and discussion of specific elements of IPBES data management reports.</p> <p>This session, <em>Data management report details: File formats</em>, focuses on specific recommended file formats for text, tabular data, images, sound, and geospatial data. </p>
Evolution of software code at the level of fine-grained elements: data files
<p>The data files available here (68GB uncompressed) have been used for studying the evolution of code at the level of fine-grained elements. The data are associated with the processing of the 89 open source software repositories hosted on GitHub. Details regarding each individual GitHub project are stored in the repos folder under directories matching the owner and project name used on GitHub. For example, the files under repos/KDE/kdevelop correspond to the project hosted on https://github.com/KDE/kdevelop. Data associated with the statistical analysis of the processed repositories are stored in the statistical-analysis folder. The file project_details.txt contains the data used for selecting the processed projects.</p>
RAPID input and output files corresponding to "River Network Routing on the NHDPlus Dataset"
<p><strong>Corresponding peer-reviewed publication</strong></p> <p>This dataset corresponds to all the RAPID input and output files that were used in the study reported in:</p> <ul> <li>David, Cédric H., David R. Maidment, Guo-Yue Niu, Zong-Liang Yang, Florence Habets and Victor Eijkhout (2011), River Network Routing on the NHDPlus Dataset, Journal of Hydrometeorology, 12(5), 913-934. DOI: 10.1175/2011JHM1345.1. </li> </ul> <p> </p> <p>When making use of any of the files in this dataset, please cite both the aforementioned article and the dataset herein. </p> <p> </p> <p><strong>Time format</strong></p> <p>The times reported in this description all follow the ISO 8601 format. For example 2000-01-01T16:00-06:00 represents 4:00 PM (16:00) on Jan 1<sup>st</sup> 2000 (2000-01-01), Central Standard Time (-06:00). Additionally, when time ranges with inner time steps are reported, the first time corresponds to the beginning of the first time step, and the second time corresponds to the end of the last time step. For example, the 3-hourly time range from 2000-01-01T03:00+00:00 to 2000-01-01T09:00+00:00 contains two 3-hourly time steps. The first one starts at 3:00 AM and finishes at 6:00AM on Jan 1<sup>st</sup> 2000, Universal Time; the second one starts at 6:00 AM and finishes at 9:00AM on Jan 1<sup>st</sup> 2000, Universal Time.</p> <p> </p> <p><strong>Data sources</strong></p> <p>The following sources were used to produce files in this dataset:</p> <ul> <li>The National Hydrography Dataset Plus (NHDPlus) Version 1, obtained from http://www.horizon-systems.com/nhdplus. </li> <li>The National Water Information System (NWIS), obtained from http://waterdata.usgs.gov/nwis. </li> <li>Outputs from a simulation using the community Noah land surface model with multiparameterization options (Noah-MP, Niu et al. 2011, http://www.jsg.utexas.edu/noah-mp). The simulation was run by Guo-Yue Niu, and produced 3-hourly time steps from 2004-01-01T00:00+00:00 to 2008-01-01T00:00+00:00. Further details on the inputs and options used for this simulation are provided in David et al. (2011).</li> </ul> <p> </p> <p><strong>Software</strong></p> <p>The following software were used to produce files in this dataset:</p> <ul> <li>The Routing Application for Parallel computation of Discharge (RAPID, David et al. 2011, http://rapid-hub.org), Version 1.0.0. Further details on the inputs and options used for this series of simulations are provided below and in David et al. (2011).</li> <li>ESRI ArcGIS (http://www.arcgis.com). </li> <li>Microsoft Excel (https://products.office.com/en-us/excel). </li> <li>CUAHSI HydroGET (http://his.cuahsi.org/hydroget.html). </li> <li>The GNU Compiler Collection (https://gcc.gnu.org) and the Intel compilers (https://software.intel.com/en-us/intel-compilers). </li> </ul> <p> </p> <p><strong>Study domain</strong></p> <p>The files in this dataset correspond to two study domains:</p> <ul> <li>The combination of the San Antonio and Guadalupe River Basins, TX. RAPID can only use the river reaches of NHDPlus that have a known flow direction and focus is made on these reaches here (a total of 5,175). The temporal range corresponding to this domain is from 2004-01-01T00:00-06:00 to 2007-12-31 T00:00-06:00.</li> <li>The Upper Mississippi River Basin. RAPID can only use the river reaches of NHDPlus that have a known flow direction and focus is made on these reaches here (a total of 182,240). The temporal range corresponding to this domain spans 100 fictitious days.</li> </ul> <p> </p> <p><strong>Description of files for the San Antonio and Guadalupe River Basins</strong></p> <p>All files below were prepared by Cédric H. David, using the data sources and software mentioned above. </p> <ul> <li><em>rapid_connect_San_Guad.csv.</em> This CSV file contains the river network connectivity information and is based on the unique IDs of NHDPlus reaches (the COMIDs). For each river reach, this file specifies: the COMID of the reach, the COMID of the unique downstream reach, the number of upstream reaches with a maximum of four reaches, and the COMIDs of all upstream reaches. A value of zero is used in place of NoData. The river reaches are sorted in increasing value of COMID. The values were computed using a combination of the following NHDPlus fields: COMID, DIVERGENCE, FROMNODE and TONODE. This file was prepared using ArcGIS and Excel.</li> <li><em>m3_riv_San_Guad_2004_2007_cst.nc. </em>This netCDF file contains the 3-hourly accumulated inflows of water (in cubic meters) from surface and subsurface runoff into the upstream point of each river reach. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2007/12/31T18:00-06:00. The values were computed by superimposing a 900-m gridded map of NHDPlus catchments to the outputs of Noah-MP. This file was prepared using ArcGIS and a Fortran program.</li> <li><em>kfac_San_Guad_1km_hour.csv. </em>This CSV file contains a first guess of Muskingum k values (in seconds) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on the following NHDPlus fields: COMID, LENGTHKM, Equation (13) in David et al. (2011), and using a wave celerity of 1 km/h. This file was prepared using a Fortran program.</li> <li><em>kfac_San_Guad_celerity.csv. </em>This CSV file contains a first guess of Muskingum k values (in seconds) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on the following NHDPlus fields: COMID, LENGTHKM, Equation (13) in David et al. (2011), and using the wave celerity numbers of Table 2 in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>k_San_Guad_2004_1.csv. </em>This CSV file contains Muskingum k values (in seconds) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on the following NHDPlus fields: COMID, LENGTHKM, and using Equation (17) in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>k_San_Guad_2004_2.csv. </em>This CSV file contains Muskingum k values (in seconds) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on the following NHDPlus fields: COMID, LENGTHKM, and using Equation (18) in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>k_San_Guad_2004_3.csv. </em>This CSV file contains Muskingum k values (in seconds) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on the following NHDPlus fields: COMID, LENGTHKM, and using Equation (19) in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>k_San_Guad_2004_4.csv. </em>This CSV file contains Muskingum k values (in seconds) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on the following NHDPlus fields: COMID, LENGTHKM, and using Equation (21) in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>x_San_Guad_2004_1.csv. </em>This CSV file contains Muskingum x values (dimensionless) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on Equation (17) in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>x_San_Guad_2004_2.csv. </em>This CSV file contains Muskingum x values (dimensionless) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on Equation (18) in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>x_San_Guad_2004_3.csv. </em>This CSV file contains Muskingum x values (dimensionless) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on Equation (19) in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>x_San_Guad_2004_4.csv. </em>This CSV file contains Muskingum x values (dimensionless) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_San_Guad.csv</em>. The values were computed based on Equation (21) in David et al. (2011). This file was prepared using a Fortran program.</li> <li><em>basin_id_San_Guad_hydroseq.csv. </em>This CSV file contains the list of unique IDs of NHDPlus river reaches (COMID) in the San Antonio and Guadalupe River Basins. The river reaches are sorted from upstream to downstream. The values were computed using the following NHDPlus fields: COMID and HYDROSEQ. This file was prepared using Excel.</li> <li><em>Qout_San_Guad_1460days_p1_dtR=900s.nc.</em> This netCDF file contains the 3-hourly averaged outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_San_Guad_hydroseq.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2007-12-31-00:00-06:00. The values were computed using the Muskingum method with parameters of Equation (17) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> <li><em>Qout_San_Guad_1460days_p2_dtR=900s.nc. </em>This netCDF file contains the 3-hourly averaged outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_San_Guad_hydroseq.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2007-12-31-00:00-06:00. The values were computed using the Muskingum method with parameters of Equation (18) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> <li><em>Qout_San_Guad_1460days_p3_dtR=900s.nc. </em>This netCDF file contains the 3-hourly averaged outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_San_Guad_hydroseq.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2007-12-31-00:00-06:00. The values were computed using the Muskingum method with parameters of Equation (19) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> <li><em>Qout_San_Guad_1460days_p4_dtR=900s.nc. </em>This netCDF file contains the 3-hourly averaged outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_San_Guad_hydroseq.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2007-12-31-00:00-06:00. The values were computed using the Muskingum method with parameters of Equation (21) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> <li><em>QoutR_San_Guad_182days_p1_dtR=900s.nc. </em>This netCDF file contains the 15-min outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_San_Guad_hydroseq.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2004-07-01-00:00-06:00. The values were computed using the Muskingum method with parameters of Equation (17) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> <li><em>QoutR_San_Guad_182days_p2_dtR=900s.nc. </em>This netCDF file contains the 15-min outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_San_Guad_hydroseq.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2004-07-01-00:00-06:00. The values were computed using the Muskingum method with parameters of Equation (18) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> <li><em>QoutR_San_Guad_182days_p3_dtR=900s.nc. </em>This netCDF file contains the 15-min outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_San_Guad_hydroseq.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2004-07-01-00:00-06:00. The values were computed using the Muskingum method with parameters of Equation (19) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> <li><em>QoutR_San_Guad_182days_p4_dtR=900s.nc. </em>This netCDF file contains the 15-min outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_San_Guad_hydroseq.csv</em>. The time range for this file is from 2004-01-01T00:00-06:00 to 2004-07-01-00:00-06:00. The values were computed using the Muskingum method with parameters of Equation (21) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> <li><em>gage_id_San_Guad_2004_2007_full.csv. </em>This CSV file contains the list of COMIDs of rivers containing USGS gauges and with full daily data record. The river reaches are sorted in increasing value of COMID. The time range used for determining a full record is daily from 2004-01-01T00:00-06:00 to 2008-01-01T00:00-06:00. The values were computed using the following NHDPlus field: COMID. This file was prepared using ArcGIS, HydroGET, and Excel.</li> <li><em>Qobs_San_Guad_2004_2007_full.csv. </em>This CSV file contains daily averaged measured stream flow (in cubic meters per second). The river reaches have the same COMIDs and are sorted similarly to <em>gage_id_San_Guad_2004_2007_full.csv</em>. The time range for the daily values is from 2004-01-01T00:00-06:00 to 2008-01-01T00:00-06:00. The values were computed using the following NHDPlus field: COMID, and the observations from NWIS. This file was prepared using ArcGIS, HydroGET, and Excel.</li> </ul> <p> </p> <p><strong>Description of files for the Upper Mississippi River Basin</strong></p> <p>All files below were prepared by Cédric H. David, using the data sources and software mentioned above. </p> <ul> <li><em>rapid_connect_Reg07.csv. </em>This CSV file contains the river network connectivity information and is based on the unique IDs of NHDPlus reaches (the COMIDs). For each river reach, this file specifies: the COMID of the reach, the COMID of the unique downstream reach, the number of upstream reaches with a maximum of four reaches, and the COMIDs of all upstream reaches. A value of zero is used in place of NoData. The river reaches are sorted in increasing value of COMID. The values were computed using a combination of the following NHDPlus fields: COMID, DIVERGENCE, FROMNODE and TONODE. This file was prepared using ArcGIS and Excel. </li> <li><em>m3_riv_Reg07_100days_dummy.nc. </em>This netCDF file contains the 3-hourly accumulated inflows of water (in cubic meters) from surface and subsurface runoff into the upstream point of each river reach. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_Reg07.csv</em>. The time range for this file is for 100 fictitious days. The values were computed using a unique value of 1 cubic meter for all river reaches and all time steps. This file was prepared using a Fortran program.</li> <li><em>kfac_Reg07_2.5ms.csv. </em>This CSV file contains a first guess of Muskingum k values (in seconds) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_Reg07.csv</em>. The values were computed based on the following NHDPlus fields: COMID, LENGTHKM, and using Equation (22) in David et al. (2011). This file was prepared using a Fortran program. </li> <li><em>xfac_Reg07_0.3.csv. </em>This CSV file contains a first guess of Muskingum x values (dimensionless) for all river reaches. The river reaches have the same COMIDs and are sorted similarly to <em>rapid_connect_Reg07.csv</em>. The values were computed based on Equation (22) in David et al. (2011). This file was prepared using a Fortran program. </li> <li><em>basin_id_Reg07_hydroseq.csv. </em>This CSV file contains the list of unique IDs of NHDPlus river reaches (COMID) in the Upper Mississippi River Basin. The river reaches are sorted from upstream to downstream. The values were computed using the following NHDPlus fields: COMID and HYDROSEQ. This file was prepared using Excel.</li> <li><em>Qout_Reg07_100days_pfac_dtR900s.nc.</em> This netCDF file contains the 3-hourly averaged outputs (in cubic meters per second) from RAPID corresponding to the downstream point of each reach. The river reaches have the same COMIDs and are sorted similarly to <em>basin_id_Reg07_hydroseq.csv</em>. The time range for this file spans 100 fictitous days. The values were computed using the Muskingum method with parameters of Equation (22) in David et al. (2011). This file was prepared using RAPID v1.0.0 running with the preonly ILU solver on one core.</li> </ul> <p> </p> <p><strong>Known bugs and limitations in this dataset or the associated manuscript.</strong></p> <p>The confluence of the Missouri River and the Upper Mississippi River upstream of Saint Louis, MO was overlooked. The contribution from the Missouri River is therefore not accounted for in the network connectivity corresponding to the Upper Mississippi River Basin. This has no effect on the conclusions of David et al. (2011) since the Upper Mississippi River Basin was studied with synthetic data and solely to evaluate parallel performance of RAPID.</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was partially supported by the U.S. National Aeronautics and Space Administration under the Interdisciplinary Science Project NNX07AL79G; by the U.S. National Science Foundation under project EAR-0413265: CUAHSI Hydrologic Information Systems; by Ecole des Mines de Paris, France; and by the American Geophysical Union under a Horton (Hydrology) Research Grant.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.