Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

414

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

414 results for “Generative model”

Learn how ShareScore rates datasets ↗
zenodo56/100

Tropical Pacific SST and wind anomalies generated by a Nonlinear Inverse Model

<p>Tropical Pacific (40S-40N; 120E-50W) sea surface temperature (SST), zonal wind (U) and meridional wind (V) anomalies generated by the Nonlinear Inverse Model described in Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5). The data consists in 99 realizations (<a href="../api/records/10411023/draft/files/NLIM_output_085.nc/content" target="_blank" rel="noopener noreferrer">NLIM_output_XXX.nc</a>) of 1,000yrs each emulating SST, U, and V monthly anomalies conditions during 1980-2020 (<a href="../api/records/10411023/draft/files/Monthly_obs_1980_2020.nc/content" target="_blank" rel="noopener noreferrer">Monthly_obs_1980_2020.nc</a>) given in a 2.5deg-2.5deg grid. For observations, we used the NOAA Extended Reconstruction SST v5 reanalysis (SST; Huang et al., 2017) and NCEP-NCAR reanalysis (winds; Kalnay et al., 1996) The observed anomalies are calculated as described in Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5).</p> <p>Given that the stochastic forcing considered is white in time and space (https://doi.org/10.1038/s41612-024-00675-5; Methods, section "Offline simulation of SSH_{12}, PC2, and spatial patterns fron nonlinear inverse model output"), the spatial patterns and lead-lag relationships are better identified using composites. A modification of the methodology that allows for spatially coherent stochastic forcing will be implemented in a future article.</p> <p>When using the data please cite https://doi.org/10.5281/zenodo.10411023 (the data) and Martinez-Villalobos et al., 2024 (https://doi.org/10.1038/s41612-024-00675-5; for the methodology).&nbsp;</p> <p>Any question, please contact Cristian Martinez-Villalobos at his email cristian.martinez.v@uai.cl</p> <p>References</p> <p>Martinez-Villalobos, C., Dewitte, B., Garreaud, R.D.&nbsp;<em>et al.</em>&nbsp;Extreme coastal El Ni&ntilde;o events are tightly linked to the development of the Pacific Meridional Modes.&nbsp;<em>npj Clim Atmos Sci</em>&nbsp;<strong>7</strong>, 123 (2024). https://doi.org/10.1038/s41612-024-00675-5</p> <p>Huang, B. et al. Extended Reconstructed Sea Surface Temperature, Version 5 (ERSSTv5): Upgrades, Validations, and Intercomparisons. Journal of Climate 30, 8179&ndash;8205 (2017).</p> <p>Kalnay, E. et al. The NCEP/NCAR 40-Year Reanalysis Project. Bulletin of the American Meteorological Society 77, 437&ndash;471 (1996).</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Data from: "Deep Generative Modeling of Periodic Variable Stars Using Physical Parameters"

<p>This dataset was used for the training of a conditioned Variational Autoencoder that generates physically informed light curves of periodic variable stars. The light curves correspond to data obtained from The Optical Gravitational Lensing Experiment (<a href="https://ui.adsabs.harvard.edu/abs/1992AcA....42..253U/abstract">OGLE</a>), while ancillary information was obtained from the Gaia Data Release 2 (<a href="https://ui.adsabs.harvard.edu/link_gateway/2016A&amp;A...595A...1G/doi:10.1051/0004-6361/201629272">GAIA DR2</a>). This repository contains the preprocessed OGLE light curves and the GAIA measurements corresponding to each cross-matched source. We also provided a subsample of cross-matched sources that were carefully validated following several steps described in the companion article (paper reference).</p> <p>This dataset is realized in tandem with the corresponding&nbsp;<a href="https://github.com/jorgemarpa/PELS-VAE">GitHub</a>&nbsp;and&nbsp;<a href="https://arxiv.org/abs/2005.07773">article</a>.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo48/100

Data release for paper "Towards the routine use of subdominant harmonics in gravitational-wave inference: re-analysis of GW190412 with generation X waveform models"

<p>This data release for the paper &quot;Towards the routine use of subdominant harmonics in gravitational-wave inference: re-analysis of GW190412 with generation X waveform models&quot; [<a href="https://arxiv.org/abs/2010.05830">arXiv:2010.2010.05830</a>] contains posterior samples for the GW190412 binary black hole merger event obtained from public GWOSC data with the parallel bilby Bayesian inference package, dynesty nested sampler and a set of waveforms from the &quot;generation X&quot; of phenomenological waveform models: IMRPhenomXAS, IMRPhenomXHM, IMRPhenomXP, IMRPhenomXPHM, IMRPhenomT and IMRPhenomTHM. The provided file is a &quot;meta file&quot; that can be read with the <a href="https://lscsoft.docs.ligo.org/pesummary/">PESummary</a> python package. The posterior samples included correspond to runs [2,6,10,12,14,26] in Table III of the paper (standard settings for each waveform, standar priors and sampler settings of Nlive=2048 and Nact=10 or 50). If you make use of these samples, please cite both this data release and the paper.</p>

opencc-by-4.0Oct 2020View details →
zenodo48/100

Best-fitting Sea Level Curves generated from the Tidal Notch Generator model

<p>The following dataset contains the data produced by the TidalNotch Generator model found at: https://zenodo.org/badge/latestdoi/700386384 and is part of the publication entitled:&nbsp;<strong>Decoding the interplay between tidal notch geometry and sea-level variability during the Last Interglacial (Marine Isotopic Stage 5e) high stand.</strong></p> <p>Each folder&nbsp;name describes the Erosion Rate used for each simulation, the Linear Regression of each curve group, and the number of peaks: e.g. Filename: 05mm_Negative_3peak.&nbsp;</p> <p>Each of the subfolders contains&nbsp;the final clusters grouped based on the methodology followed, extensively described in the manuscript.</p> <p>Each txt file contains 15 columns, while the content of each one is described below:</p> <p>Column 1: Random Sea Level Curve (Years)</p> <p>Column 2: Random Sea Level Curve (Elevation)</p> <p>Column 3: Modeled Notch Geometry (Notch Depth)</p> <p>Column 4: Modeled Notch Geometry (Notch Elevation)</p> <p>Column 5: Measured&nbsp;Notch Geometry (Notch Depth)</p> <p>Column 6: Measured&nbsp;Notch Geometry (Notch Elevation)</p> <p>Column 7: Fitting score (e.g. 0.17 --&gt; 1-0.17=0.83--&gt;83%)</p> <p>Column 8: Polynomial Order used to Interpolate the Randomly generated Sea Level points</p> <p>Column 9: Erosion Rate used for the simulation</p> <p>Column 10: ID of measured notch profile</p> <p>Column 11: number&nbsp;of simulation&nbsp;</p> <p>Column 12: Inclination of the measured notch</p> <p>Column 13: Category of Inclination</p> <p>Column 14: second ID of measured notch profile</p> <p>Column 15:&nbsp;Linear Regression value</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

PanTaGruEl - a pan-European transmission grid and electricity generation model

<p>If you have any questions or comments, please write to <a href="mailto:laurent.vincent.pagnier@gmail.com">laurent.vincent.pagnier@gmail.com</a>.</p> <p>When publishing results based on this data set, please cite:</p> <p>L. Pagnier, P. Jacquod, &ldquo;Inertia location and slow network modes determine disturbance propagation in large-scale power grids&rdquo;, PLOS ONE 14(3): e0213550, 2019. <a href="https://doi.org/10.1371/journal.pone.0213550">PLOS ONE 14(3): e0213550</a>, 2019.</p> <p>and</p> <p>M. Tyloo, L. Pagnier, P. Jacquod, &ldquo;The Key Player Problem in Complex Oscillator Networks and Electric Power Grids: Resistance Centralities Identify Local Vulnerabilities&rdquo;, <a href="https://doi.org/10.1126/sciadv.aaw8359">Science Advances 5(11): eaaw8359</a>, 2019.</p> <p><strong>Description:</strong></p> <p>PanTaGruEl is a dynamical grid model designed to investigate the propagation of disturbances in the continental European transmission grid.</p> <p>The construction of the model is detailed <a href="https://doi.org/10.1371/journal.pone.0213550.s002">here</a>.</p> <p><strong>Features</strong>:</p> <ul> <li>Precise distribution of national demands to network buses.</li> <li>Realistic electrical parameters of transmission lines.</li> <li>Merit-Order based economic dispatch of generators.</li> <li>Dynamical parameters of generators and loads for transient stability investigations.</li> </ul> <p><strong>Files:</strong></p> <p>Data files:</p> <p>Our model is provided in an extended Matpower format and as csv raw data. For more information on Matpower format, see Appendix B of its <a href="https://matpower.org/docs/MATPOWER-manual.pdf">manual</a>.</p> <p>Script files:</p> <p><em>opf_ex.m </em>performs optimal power flow computations for two load configurations.<br> <em>spectral_ex.m</em> presents a basic spectral analysis.<br> <em>dynamics</em><em>_ex.m</em> give a minimal example of dynamical simulations.</p> <p><strong>Requirements:</strong></p> <p>Our model has been developed for use with <a href="https://matpower.org/">Matpower</a>. If you are interested in a port to another language, please <a href="mailto:laurent.vincent.pagnier@gmail.com?subject=Info%20on%20PanTaGruEl">contact us</a>.</p> <p><strong>Acknowledgement:</strong></p> <p>The authors thank M. Tyloo and K. Van Walstijn for their useful comments and remarks on the model.</p> <p><strong>Sources</strong>:</p> <p>B. Wiegmans, <a href="https://doi.org/10.5281/zenodo.55853">&ldquo;GridKit extract of ENTSO-E interactive map&rdquo;</a><br> Global Energy Observatory, <a href="http://globalenergyobservatory.org">&ldquo;GEO Power plants database&rdquo;</a><br> Siemens, <a href="http://siemens.com/power-engineering-guide">&ldquo;Power Engineering Guide&rdquo;</a></p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Next generation global ice-ocean-biogeochemistry coupled model with 13C-cycling (GFDL MOM5-BLING13C)

<p>&nbsp;</p> <p>======= &nbsp;DESCRIPTION &nbsp;=======</p> <p>This is the model output supporting our paper&nbsp;<em>A next generation ocean carbon isotope model for climate studies I: Steady state controls on ocean <sup>13</sup>C</em>&nbsp;(2021 Global Biogeochemical Cycles).</p> <p>This model output simulates the transient response of ocean carbon biogeochemistry to anthropogenic CO<sub>2</sub> and <sup>13</sup>CO<sub>2</sub>&nbsp;atmospheric emissions with a nominal lateral resolution of 1&deg; and 50 vertical levels. The model uses the NOAA&#39;s Geophysical Fluid Dynamics Laboratory (GFDL) MOM5 coupled to the NOAA-GFDL Biogeochemistry with Light Iron Nutrients and Gas (BLING) with <sup>13</sup>C-cycling. Atmospheric forcing is prescribed using the repeating annual cycle of the Common Ocean Reference Experiment version 2 normal year forcing dataset (COREv2-NYF). The implementation of <sup>13</sup>C-cycling applies isotopic fractionations during air-sea gas exchange, photosynthetic production of organic matter, and formation of calcium carbonate. The sensitivity of dissolved inorganic <sup>13</sup>C in the ocean to the CO<sub>2</sub> gas exchange rate is explored by repeating the simulation twice, once using the latest OMIP-CMIP6 protocol for the k-U<sub>10 </sub>parameterization (standard) and once using the previous OCMIP2 protocol (fast-gas-exchange).</p> <p>&nbsp;</p> <p>Files information:</p> <ul> <li><strong>ocean_static.nc</strong>: Static fields (longitude, latitude, area).</li> <li><strong>1990-2002.ocean_month.nc</strong>: Monthly output between 1990 and 2002 of ocean physical variables (temperature, salinity, averaged mixed layer depth, maximum mixed layer depth).</li> <li><strong>1990-2002.ocean_bling_trc_month_CMIP6.nc</strong>: Monthly output between 1990 and 2002 of biogeochemical variables* for the simulation using the OMIP-CMIP6 air-sea gas exchange protocol.</li> <li><strong>1970_1989_d13c_org_mldave_CMIP6.nc</strong>: Monthly output between 1970 and 1989 of d<sup>13</sup>C of organic matter averaged over the mixed layer.</li> <li><strong>1990-2002.ocean_bling_trc_month_OCMIP2.nc</strong>: Monthly output between 1990 and 2002 of biogeochemical variables* for the simulation using the OCMIP2 air-sea gas exchange protocol.</li> </ul> <p>* Biogeochemical variables are dissolved inorganic carbon, dissolved inorganic carbon-13, oxygen, and dissolved inorganic phosphate.</p> <p>&nbsp;</p> <p>======= &nbsp;HOW TO CITE &nbsp;=======</p> <p>This model output can be freely distributed, but please cite it using the following paper:</p> <p>Claret, M., Sonnerup, R. E., &amp; Quay, P. D. (2021). A next generation ocean carbon isotope model for climate studies I: Steady state controls on ocean <sup>13</sup>C. <em>Global Biogeochemical Cycles</em>, 35, e2020GB006757. <a href="https://doi.org/10.1029/2020GB006757">https://doi.org/10.1029/2020GB006757</a></p> <p>&nbsp;</p> <p>======= &nbsp;ACKNOWLEDGEMENTS&nbsp;=======</p> <p>This work was funded by the National Science Foundation (NSF-OCE 1356756 and NSF-OCE 1829796). We would also&nbsp;like to acknowledge high-performance computing support from Cheyenne (<a href="https://doi.org/10.5065/D6RX99HX">doi:10.5065/D6RX99HX</a>) provided by NCAR&#39;s Computational and Information Systems Laboratory, sponsored by the NSF.</p> <p>&nbsp;</p> <p>======= &nbsp;QUESTIONS AND REQUESTS? &nbsp;=======</p> <p>Please contact Mariona Claret (mclaret@uw.edu) or Rolf Sonnerup (rolf@uw.edu).</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Datasets of sequences, alignments and structural models generated for the structural prediction of complexes mediated by intrinsically disordered regions.

<p>This repository contains input and ouput files&nbsp;used and generated for the scanning of intrinsically disordered region and the prediction of their binding sites to receptor proteins using the <a href="https://github.com/i2bc/SCAN_IDR">SCAN_IDR</a> pipeline with AlphaFold2-Multimer.</p><p>It contains two archives:&nbsp;</p><ol><li><a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> dedicated to the analysis of a dataset of 42 protein complexes non redundant with the dataset used for AlphaFold2 training,</li><li><a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> dedicated to the analysis of 923 complexes from the ELM database.</li></ol><p>These data can be used to rerun specific sections of the pipeline and scripts provided in: <a href="https://github.com/i2bc/SCAN_IDR">https://github.com/i2bc/SCAN_IDR</a></p><h4><strong>Dataset of 42 non redundant complexes</strong></h4><p>The first archive <a href="https://zenodo.org/api/records/10068949/draft/files/scanidr_data_repository_corr6J08.tar/content"><i><strong>scanidr_data_repository_corr6J08.tar</strong></i></a> contains 3 compressed directories and a README file detailing their contents :</p><ul><li>the initial raw sequence and alignment data for every chain&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;-&gt; DIRECTORY <strong>fasta_msa/</strong></li><li>the input and output data of every Alphafold run for every complex&nbsp; &nbsp;-&gt; DIRECTORY <strong>af2_runs/</strong></li><li>the native reference structures&nbsp;&nbsp;&nbsp; -&gt; DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p>The protein-peptide complex cases have been assigned a distinct index number, from 1 to 42, consistent across the several directories of the archive. Their corresponding directories are labelled as <i>&lt;index&gt;_&lt;pdbcode&gt;</i>.</p><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.2</i></p><h4><strong>Dataset of 923 complexes selected from the ELM database</strong></h4><p>The second archive <a href="https://zenodo.org/api/records/10068949/draft/files/923_elm_cases_repository.tar.gz/content"><i><strong>923_elm_cases_repository.tar.gz</strong></i></a> contains input and ouput files used and generated for the analysis of 923 Eukaryotic Linear Motifs (ELM) database entries.</p><p>Each ELM entry is indexed with specific integer id and is composed of a receptor and a ligand protein. &nbsp;</p><p>The archive contains a Table associating ELM indexes with the ELM entry information, 5 directories and a README file detailing their contents:</p><ul><li>the table describing ELM entries -&gt; FILE <strong>Table_923ELM_uid_delimitations_info_for_archive.txt</strong></li><li>the initial raw sequence and multiple sequence alignment (MSA) data for every chain &nbsp; &nbsp; &nbsp; &nbsp;-&gt; DIRECTORY <strong>fasta_msa/</strong></li><li>the concatenated MSA model for every ELM complex and protocol used -&gt; DIRECTORY <strong>af2_elm_coali_inputs/</strong></li><li>the best model of every AF2 protocol for every complex according to the AF2 &nbsp; -&gt; DIRECTORY <strong>af2_elm_models/</strong></li><li>the best model cut in the ligand part to select only the ELM motifs as used for the evaluation of the models -&gt; DIRECTORY <strong>elm_cut_models/</strong></li><li>the reference structures used for the evaluation of the models &nbsp; -&gt; DIRECTORY <strong>ref_capri_curated/</strong></li></ul><p><i>The models in this archive were generated using AlphaFold2-Multimer v2.3</i></p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Auxiliary files and data to generate eddy flux and validate 2D model for MALTA

<p>This repository contains the following directories to accompany the manuscript 'A Zonally-Averaged Global Atmospheric Transport Model for Long-lived Trace Gases', submitted to JAMES:</p><p>1) <strong>GEOSChem&nbsp;</strong>This directory contains the run directory template and (slurm) runscript to generate the tracer fields used to generate the eddy fluxes. The GEOSChem model will have to be installed locally to run this, and the run directory&nbsp;built to your local area. It may be easiest to just copy the relevant bits&nbsp;in /Tracer_2D_template/&nbsp;(i.e., the .rc files, /RestartFiles/, input.geos, reset_restart.py and species_database.yml) into a GEOSChem Transport run directory and change the directories in the copied files. If using slurm on an HPC, just change the directories in the runtracers_inputs.sh script to match that of your own HPC. Else, a different script will have to be written copying the slurm functionality.</p><p>2)&nbsp; <strong>GEOSChem_SF6&nbsp;</strong>This directory contains the monthly mean SF6 mole fractions generated using GEOSChem used to validate the 2D model MALTA. Emissions come from the EDGAR&nbsp;v4.2 emissions inventory. Emissions after 2008 continue to use 2008 as the emissions value.</p><p>3)&nbsp;<strong>CFC11_inversion</strong>&nbsp;This directory contains the relevant script and files to quantify emissions of CFC-11 using an output mole fraction from the TOMCAT 3D model using MALTA, and compare these to the TOMCAT emissions used to generate the mole fractions. The directory paths at the beginning of the main script in CFC11_inversion.py must be changed to point to the remaining files in the /CFC11_inversion/ directory, and a save directory must be specified, before running locally. MALTA must be installed to run this.</p><p>4) <strong>singapore.dat </strong>This file contains the QBO winds above Singapore, taken from https://www.geo.fu-berlin.de/en/met/ag/strat/produkte/qbo/index.html</p><p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Dataset and neural network weights to the paper: "Generative diffusion for regional surrogate models from sea-ice simulations"

<p>All the needed code and data to reproduce the results from the paper: "Generative diffusion for regional surrogate models from sea-ice simulations".<br>While most of the code is a frozen clone of the original&nbsp;<a href="https://github.com/cerea-daml/diffusion-nextsim-regional">Repository</a>, this capsule also includes the dataset and neural network weights to train and apply the surrogate models.</p> <p>The <strong>dataset</strong> for training and evaluation can be found at&nbsp;<em>data/nextsim</em>, which includes three different Zarr folders for training/validation/testing. The dataset is based on neXtSIM simulation data and ERA5 forcing data and extracted from the <a href="https://ige-meom-opendap.univ-grenoble-alpes.fr/thredds/catalog/meomopendap/extract/catalog.html">SASIP shared data OpenDAP server</a>:</p> <ul> <li>The neXtSIM simulations were performed by Gauillaume Boutin and published in the paper "<a href="https://doi.org/10.5194/tc-17-617-2023">Arctic sea ice mass balance in a new coupled ice&ndash;ocean model using a brittle rheology framework</a>" (Boutin et al., 2023) and available as Zenodo <a href="../records/7277523">dataset</a> (Boutin et al., 2022).</li> <li>The forcing data is based on the ERA5 reanalysis dataset published in the paper: "<a href="https://doi.org/10.1002/qj.3803">The ERA5 global reanalysis</a>" (Hersbach et al., 2020) and available as dataset from the Copernicus Climate Change Service (C3S, Copernicus Climate Change Service, 2023). The here used forcing data is based on the <a href="https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-single-levels">hourly reanalysis data on single levels</a> and interpolated with nearest neighbors to the curvilinear grid as used in the output from the neXtSIM simulations. <strong>Disclaimer:</strong> The results contain modified Copernicus Climate Change Service information, 2023. Neither the European Commission nor ECMWF is responsible for any use that may be made of the Copernicus information or data it contains.</li> </ul> <p>The <strong>neural network weights</strong> are included under <em>data/models </em>and split into weights for the deterministic models and the diffusion models.<br>These neural network weights have been used to generate the results presented in the paper.</p> <p>In this capsule, the <em>notebooks</em> folder includes also the figures used within the paper and additional trajectory data used in the qualitative analysis of the paper.</p> <p>Generally, we recommend to just download the <em>data.tar.gz </em>file and use otherwise the original <a href="https://github.com/cerea-daml/diffusion-nextsim-regional">Repository</a>, since the here included code can be outdated. We further refer to the repository for additional information.</p> <p>&nbsp;</p> <p>Contained in this capsule:</p> <ul> <li>configs.tar.gz: The configuration files for the experiments.</li> <li>data.tar.gz: The dataset and neural network weights.</li> <li>diffusion_nextsim.tar.gz: The main code for the neural network etc.</li> <li>environment.yaml: The anaconda environment file, can be used to install the needed packages.</li> <li>notebooks.tar.gz: The notebooks that were used to create the figures in the paper. The figures from the paper and the data from the qualitative analysis are included as well.</li> <li>readme.md: The readme file from the repository.</li> <li>scripts.tar.gz: The scripts used for the experiments.</li> <li>setup.py: the file to install the <em>diffusion_nextsim</em> package in a python environment.</li> </ul> <p>References:</p> <p>Guillaume Boutin, Heather Regan, Einar &Oacute;lason, Laurent Brodeau, Claude Talandier, Camille Lique, &amp; Pierre Rampal. (2022). Data accompanying the article "Arctic sea ice mass balance in a new coupled ice-ocean model using a brittle rheology framework" (1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.7277523</p> <p>Boutin, G., &Oacute;lason, E., Rampal, P., Regan, H., Lique, C., Talandier, C., Brodeau, L., and Ricker, R.: Arctic sea ice mass balance in a new coupled ice&ndash;ocean model using a brittle rheology framework, The Cryosphere, 17, 617&ndash;638, https://doi.org/10.5194/tc-17-617-2023, 2023.</p> <p>Copernicus Climate Change Service (2023): ERA5 hourly data on single levels from 1940 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS), DOI:&nbsp;<a href="https://doi.org/10.24381/cds.adbb2d47">10.24381/cds.adbb2d47</a>.</p> <p>Hersbach H, Bell B, Berrisford P, et al. The ERA5 global reanalysis. <em>Q J R Meteorol Soc</em>. 2020; 146: 1999&ndash;2049. <a href="https://doi.org/10.1002/qj.3803">https://doi.org/10.1002/qj.3803</a></p> <p>&nbsp;</p>

openmit-licenseApr 2024View details →
zenodo44/100

Human-AI Collaboration: A tool to enable AI model generation with human-in-the-loop

<p>Human-AI collaboration enables domain experts to contribute their expertise with the goal of enhancing the knowledge learned by the AI models from the patterns in the data. This enables the integration of domain-specific knowledge to enrich the data for further improvement of the models through retraining. The human-AI collaboration is composed of multiple sub-components and interfaces that enables communication with external systems such as data sources, model repositories, machine configurations and decision support systems.</p> <p>Human-AI Collaboration component is developed using Python programming language. The frontend is developed using Streamlit1. The backend is developed using python and the API is implemented using FastAPI2. The choice of the programming language was made because of its wide usage and vast user base. The frameworks Streamlit and FastAPI are chosen because of the rich features for functionality and documentation as well as suitability for data analysis tasks. The applications are packaged as docker images for deployment. The application runs as a web application served by nginx for reverseproxying and users can access it via client applications such as web browsers or REST clients like Postman.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data

<p>This repository contains the training data and source code to reproduce the results of our paper:<br>scGraph2Vec: a deep generative model for gene embedding augmented by Graph Neural Network and single-cell omics data</p> <p>More description can be also found in GitHub (https://github.com/LPH-BIG/scGraph2Vec).</p>

opencc-zeroJun 2024View details →
zenodo44/100

Reference Mean and Low Streamflow for all Brazilian Catchments Generated Using Machine Learning Models

<p>This dataset provides comprehensive hydrological information for river networks in Brazil, focusing on reference streamflows, specifically long-term mean flows (qm) and low flows exceeded 95% of the time (q95). Covering over 400,000 ungauged river points, the dataset was developed using advanced machine learning models trained on environmental descriptors and validated against data from 1,069 gauging stations spread across the country. The machine learning pipeline evaluated six regression models to achieve high predictive accuracy (R&sup2; &gt; 0.8 for qm and &gt; 0.7 for q95). The 62 environmental descriptors - encompassing climate, topography, land cover, lithology, and water storage characteristics - that were used as features for the models are also included.</p> <p>Key features:</p> <ul> <li><strong>Spatial Coverage:</strong> Brazilian territory and the Amazon River basin, based on the <a href="https://metadados.snirh.gov.br/geonetwork/srv/api/records/f7b1fc91-f5bc-4d0d-9f4f-f4e5061e5d8f" target="_blank" rel="noopener">BHO 5k</a> dataset of officially adopted river networks.</li> <li><strong>Outputs:</strong> Predicted qm and q95 values for each river stretch, with 90% and 75% confidence intervals to account for prediction uncertainty.</li> <li><strong>Environmental Descriptors:</strong> Aggregated from upstream catchment area.</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

A Practical Tool-Chain for the Development of Coordination Scenarios - Graphical Modeler, DSL, Code Generators and Automaton-Based Simulator

<p>The Peer Model is a modeling tool for coordination based on blackboard-based collaboration.&nbsp;</p> <p>The tool-chain consists of a modeler, translator and simulator.</p> <p>Its goal is to help developers of distributed and concurrent coordination software better understand algorithms and identify deficiencies from the beginning.</p> <p><br> &nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Data supporting "Transformer Model Generated Bacteriophage Genomes are Compositionally Distinct from Natural Sequences"

<p>Sequence and composition data supporting doi: <a href="https://doi.org/10.1101/2024.03.19.585716" target="_blank" rel="noopener">10.1101/2024.03.19.585716</a>.&nbsp;Uncompressed file size is ~5.8GB.</p> <p>Data in zip files is organized by sequence provenance (generRNA, natural, or transformer (megaDNA)). Common file types between folders include:</p> <ul> <li>Multi-record fasta file: Sequence data for all sequences of a given provenance. For generRNA sequences, these are found within the `seq` column of file "MFE_distribution_Fig4a.csv"</li> <li>Composition files: Individual sequence level compositional metrics for sliding 120 bp windows. Only structural metrics were used in this study.</li> <li>Genomad: Results from the genomad pipeline (https://portal.nersc.gov/genomad/)</li> <li>Stats: Aggregate statistics for all sequences of a given provenance.</li> </ul> <p>The natural folder also has a metadata file detailing the taxonomy for all natural sequences.<br><br>Figure datasets are the cleaned (sometimes aggregated) datasets that underly specific figures in the manuscript. The figure designations are based on the order in: https://www.biorxiv.org/content/10.1101/2024.03.19.585716v1.</p>

opencc-by-4.0May 2024View details →
zenodo44/100

Task 3 Dataset for Dreaming of Electrical Waves: Generative Modeling of Cardiac Excitation Waves using Diffusion Models

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo44/100

Computational models for kaolinite nano-particles (Generations 1-3) and their comprehensive FTIR spectra

<p>The dataset contains a large number of computational models and detailed spectral comparison, fitting, and deconvolution of a large set of FTIR data for crystalline and exfoliation kaolinite, nano-kaolinite and halloysite, nano-halloysite samples.<br> The <strong>G1.xyz</strong>, <strong>G2.xyz</strong>, and <strong>G3.xyz</strong> files contain the initial structures for the first three generations of nano-kaolinite molecules.<br> The compressed folder&nbsp;<strong>SVP-def2TZVP.zip</strong> contains the structural information relevant for comparing and contrasting the performance a double-zeta (SVP) and triple-zeta (TZVP) basis sets.<br> The <strong>edge_protonation.zip</strong> folder guides the reader through the stepwise evaluation of various edge protonation models and shows the final converged results.<br> The <strong>full_optimization.zip</strong> folder summarizes&nbsp;the stationary structure calculations at various levels of theory carried out for the G2 model.<br> &nbsp;</p>

opencc-by-4.0Mar 2018View details →
zenodo44/100

Predictive models for off-target binding profiles generation

<p>Models for predicting off-target binding, built with Conformal Prediction, and the <a href="http://cpsign-docs.genettasoft.com">CPSign software</a>. The dataset is part of an upcoming publication (Manuscript in preparation), which will provide more details.</p> <p>The dataset is a GZipped Tar archive, with the models as Java Archive (JAR) files. For every JAR-file, there is also a corresponding audit log, with the extension &quot;.audit.json&quot;, produced by the workflow software (<a href="http://scipipe.org">SciPipe</a>) used to train the models. This audit file contains all the shell commands used in the workflow that produced the models.</p>

opencc-by-4.0Jun 2018View details →
zenodo44/100

Outputs of the next generation sea ice model (neXtSIM) for winter 2006 - 2007 saved for comparison with RGPS.

<p>NeXtSIM was run from 1 December 2006 to 15 April 2007 with the following parameters:</p> <p><code>[mesh]</code><br><code>filename=small_arctic_10km.msh</code></p> <p><code>[simul]</code><br><code>duration=150</code><br><code>time_init=2006-11-15</code><br><code>timestep=900</code></p> <p><code>[dynamics]</code><br><code>compression_factor=13800</code><br><code>C_lab=2675000</code><br><code>nu0=0.301</code><br><code>tan_phi=0.624</code><br><code>substeps=90</code><br><code>time_relaxation_damage=15</code><br><code>use_temperature_dependent_healing=true</code></p> <p><code>[output]</code><br><code>exporter_path=/cluster/work/users/akorosov/music/sa10free_mat00</code><br><code>output_per_day=4</code><br><code>variables=M_VT</code><br><code>variables=Concentration</code><br><code>variables=Thickness</code></p> <p><code>[setup]</code><br><code>atmosphere-type=era5</code><br><code>ice-type=topaz_osisaf_icesat</code><br><code>ocean-type=topaz</code><br><code>bathymetry-type=etopo</code><br><code>dynamics-type=bbm</code></p> <p><code>[thermo]</code><br><code>diffusivity_sss=0</code><br><code>diffusivity_sst=0</code><br><code>h_young_max=0.3</code><br><code>newice_type=1</code><br><code>hnull=0.5</code></p> <p><code>[debugging]</code><br><code>check_fields_fast=false</code></p> <p>The outputs (binary snapshots at every 3 hours) were then merged with RGPS data from the same period using this notebook:</p> <p>https://github.com/nansencenter/music_nextsim_tuning_paper/blob/main/02_process_nextsim.ipynb</p> <p>&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

From Ridge 2 Reef: An Interdisciplinary Model for Training the Next Generation of Environmental Problem Solvers

<p>This dataset contains the raw data from the evaluation instruments and accompanies the manuscript: "From Ridge 2 Reef: An Interdisciplinary Model for Training the Next Generation of Environmental Problem Solvers". It contains all trainee and advisor interviews from 2018 - 2022, as well as a select few partner interviews. It also contains pre and post-annual trainee survey data and the codebook to decipher the survey data. Rubric criteria and scores are included for trainees enrolled in the R2R Communication Skills course. The R script contains the statistical analyses reported in the manuscript and code used to generate figures.</p>

opencc-by-4.0Jun 2023View details →
zenodo44/100

Supplementary datasets for the manuscript "Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states" - Part 2

<p>Supplementary files containing datasets needed to reproduce the results of the manuscript &quot;Generative machine learning produces kinetic models that accurately characterize intracellular metabolic states&quot; by S. Choudhury et al.</p> <p>The code to use with these data and reproduce the manuscript results is available at&nbsp; https://github.com/EPFL-LCSB/renaissance and https://gitlab.com/EPFL-LCSB/renaissance. The execution of parts of this code is dependent on the SkimPy toolbox (https://github.com/EPFL-LCSB/skimpy). Refer to the readme files on the RENAISSANCE code repositories for more details.</p> <p>The dataset contains the following files:</p> <p>1. param_fixing.zip - self-explanatory (Figure 4 &amp; 5); contains an explanatory note for this part (experiment_details.txt), and the file containing Km values fetched from the BRENDA database (Km_database.csv).</p> <p>2. scripts.zip - scripts to generate figure 2-5 on toy data</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record