Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
12,072
datasets available to search
ShareScore release 0.7.1
Dataset results
12,072 results for “Global”
Global distribution of predicted soil types at 1 km resolution based on the WRB 2022 classification
<p>Global maps at 1 km spatial resolution of the predicted soil types (0–100% probabilities) at 1 km resolution based on the <a href="https://www.fao.org/soils-portal/data-hub/soil-classification/world-reference-base/en/">WRB 2022</a> (<strong>World Reference Base</strong> the international standard for soil classification) classification system. The training data comes from the following 3 main sources:</p> <ol> <li>WOSIS points available via: <a href="https://www.isric.org/explore/wosis">https://www.isric.org/explore/wosis</a>;</li> <li>HWSD v2 (random draw of cca 20,000 points): <a href="https://iiasa.ac.at/models-tools-data/hwsd">https://iiasa.ac.at/models-tools-data/hwsd</a>;</li> <li>Other national datasets / data from publications and projects.</li> </ol> <p>Predictions are based on using Rando Forest algorithm as implemented in the <a href="https://www.randomforestsrc.org/">randomForestSRC package</a> with cca 190 covariate layers representing soil forming factors (CHELSA Climate, Global Lithological DB GLiM, MODIS EVI and LST long-term derivatives, Digital Terrain model parameters and similar).</p> <p>All TIF files are provided as <a href="https://www.cogeo.org/">COGs</a>, which means that you can open them directly in QGIS or similar. Publication explaining all modeling steps is pending.</p> <p>Update of the predictions takes about 4–5 hrs and will be regularly run provided that new training points are available. Disclaimer: These are initial results with limited accuracy and possible issues with quality of training points, location errors and harmonization issues. Use at own risk.</p> <p>Note: original list of soil types have been subset to classes that appear at least 10 times and at least in 2 countries. If you notice an error or artifact <strong>please report via <a href="https://github.com/OpenGeoHub/SoilTypeMapping">the Github repository</a></strong>. Help us improve this dataset by contributing training points.</p>
CoCO2-MOSAIC 1.0: a global mosaic of regional, gridded, fossil and biofuel CO2 emission inventories
<p>CoCO2-MOSAIC 1.0 is a global mosaic of regional bottom-up inventories of anthropogenic CO2 emissions developed in the framework of the CoCO2 project (<a href="https://coco2-project.eu/">https://coco2-project.eu/</a>). CoCO2-MOSAIC 1.0 provides gridded (0.1˚×0.1˚) monthly emissions fluxes of CO2 fossil fuel (CO2ff, long cycle) and CO2 biofuel (CO2bf, short cycle) for the years 2015 to 2018 disaggregated in seven sectors: energy_s (super-emitting sources above 7.9e-6 kg/m2/s), energy_a (average emitters), manufacturing, settlements, transport, aviation land/take-off (LTO) and other. The regional inventories included are CAMS-GHG-REG 5.1 (Europe), DACCIWA 2.0 (Africa), GEAA-AEI 3.0 (Argentina), INEMA 1.0 (Chile), REAS 3.2.1 (South-East Asia) and VULCAN 3.0 (USA). EDGAR 6.0 and CAMS-GLOB-SHIP 3.1 are used for gap-filling missing sectors and regions. CAMS-GLOB-TEMPO 3.1 is used for temporal disaggregation of inventories providing annual emissions. Aviation emissions from climb, descent, and cruise are not covered by regional inventories and are provided as a separate file. Note that 2015 is the only year when all regional inventories are simultaneously available. </p> <p>Compared to global inventories, CoCO2-MOSAIC 1.0 includes all the regional information available without the limitation of providing spatially consistent emissions. Therefore, CoCO2-MOSAIC 1.0 can be used as a global baseline inventory due to the higher level of detail, higher spatial resolution, and country-specific information included by regional inventories. </p> <p>For further details see Urraca et al. 2023 (ESSD submitted). The paper (i) describes the CoCO2-MOSAIC methodology and (ii) uses the mosaic to inter-compare the most widely used global inventories: CAMS-GLOB-ANT 5.3, EDGAR 6.0/7.0, ODIAC v2020b, and CEDS v2020_04_24.</p>
Sparse observations induce large biases in estimates of the global ocean CO2 sink: an ocean model subsampling experiment
<p>Dataset underlying the analysis in Hauck et al., 2023: Sparse observations induce large biases in estimates of the global ocean CO<sub>2</sub> sink - an ocean model subsampling experiment, Philosophical Transactions A</p> <p>Surface ocean partial pressure of CO<sub>2 </sub>(pCO<sub>2</sub>) and air-sea CO<sub>2</sub> flux reconstructions, using two mapping methods (MPI-SOM-FFN, CarboScope) three different sampling masks: SOCAT, SOCAT+SOCCOM, IDEAL (based on bgcArgo, Roemmich et al., 2019).</p> <p>Also, all FESOM-REcoM output fields that were used in the reconstructions are provided.</p> <p>We further provide the three masks that were used for subsampling: SOCAT, SOCAT+SOCCOM, IDEAL (bgcArgo).</p> <p> </p>
Global Meteor Network observations of Crew-5 Dragon trunk re-entry 2023-04-27
<p>This dataset contains video observations by some stations of the Global Meteor Network of the re-entry of the Crew-5 dragon trunk above Arizona on 2023-04-27 around 08:52 UTC.</p> <p>There are several types of files:</p> <ul> <li>FF files: these are 10.24 second videos compressed in the four-frame format. They are just FITS files with four frames, containing per pixel 1) the maximum value over 256 frames 2) the frame nr (between 0 and 255) where the maximum occurred 3) the mean value of all 256 frames and 4) the RMS of the 256 values.</li> <li>FR files: compressed video recordings of detected fireballs. These can be read with the RMS software.</li> <li>MP4 files: rendered movies of combined FF and FR files for one station (more can be made with FR_binviewer from RMS software).</li> <li>Platepar-files: these contain astrometry corresponding to the FITS files. These can be interpreted by the RMS software.</li> <li>ECSV files: these contain manually picked points (with SkyFit2.py from RMS) along the track of the reentry. For each point, time and apparent coordinates are recorded. These files can be interpreted by the WesternMeteorPyLib trajectory solver.</li> <li>trajectory-points.txt: solutions from the trajectory solver.</li> <li>reentry-map-v4.png: a rendered map of the trajectory (made in QGIS).</li> <li>compilation.png: rendered version of the FF-files of most stations.</li> </ul> <p>The files can be processed with the software in https://github.com/CroatianMeteorNetwork/RMS and https://github.com/wmpg/WesternMeteorPyLib.</p>
Data set discussed in "Beyond Fortune 500: Women in a Global Network of Directors"
<p>Bipartite graph of directors and companies. Generated from information on the Financial Times website (<a href="https://markets.ft.com/data/equities/results">https://markets.ft.com/data/equities/results</a>), retrieved on 17 September 2016.</p> <p>Blank fields are used for missing data.</p> <p><strong>comp_nodes.csv:</strong></p> <ul> <li>id: unique identifier</li> <li>ft_country: name of the country</li> <li>ft_sector: segment of the economy in which a company operates</li> <li>ft_industry: specific business (i.e., subset of sector) in which a company operates</li> <li>ft_employees_num: number of company's employees. "NA" if the vertex represents a person or if the company's number of employees is unknown.</li> </ul> <p><strong>comp_people_edges.csv:</strong></p> <ul> <li>person_id:</li> <li>comp_id: company identifier. It matches the identifier in comp_nodes.csv</li> </ul> <p><strong>people_one_mode_edges.csv:</strong></p> <p>Edges in the one-mode projection, in which two directors are connected if and only if they sit together on at least one board. Numbers correspond to the identifiers in unique_people_nodes.csv.</p> <p><strong>unique_people_nodes.csv:</strong></p> <ul> <li>ID: unique identifier</li> <li>age: years of age</li> <li>gender_base: "Male" or "Female"</li> </ul>
MODIS MCD12Q1 Land Cover and Land Use Time Series Global Mosaics 2001-2022 (500 m)
<p><strong>General Description</strong></p> <p>The yearly land use and land cover dataset is derived from <a href="https://ladsweb.modaps.eosdis.nasa.gov/missions-and-measurements/products/MCD12Q1"><abbr title="MCD12Q1 MODIS/Terra+Aqua Land Cover Type Yearly L3 Global 500m">MCD12Q1 v061</abbr></a>. This data provides an yearly mosaics of land use and land cover data from 2001 to 2022 in cloud optimized Geotiff (COG) format. This dataset includes layers of land cover type 1 (t1), 2 (t2), and 5 (t5), land cover property 1 (p1) and 2 (p2), land cover property assessment 1 (p1a) and 2 (p2a), and land cover quality control (qc). </p> <p><strong>Data Details</strong></p> <ul> <li><strong>Time period:</strong> 2001–2022</li> <li><strong>Type of data:</strong> Land cover and land use</li> <li><strong>How the data was collected or derived:</strong> Derived from MCD12Q1 v061 using <a href="https://earthengine.google.com">Google Earth Engine</a>.</li> <li><strong>Statistical methods used:</strong> None</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Antarctica.</li> <li><strong>Coordinate reference system:</strong> EPSG:4326</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (-180.00000, -62.00081, 179.99994, 87.37000)</li> <li><strong>Spatial resolution:</strong> 1/120 d.d. = 0.008333333 (1km)</li> <li><strong>Image size:</strong> 86,400 x 35,849</li> <li><strong>File format:</strong> Cloud optimized Geotiff.</li> </ul> <p><strong>Support</strong></p> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/-/issues">GitLab Issues</a></li> <li>General questions and comments: <a href="https://disqus.com/home/forums/landgis">LandGIS Forum</a></li> </ul> <p><strong>Name convention</strong></p> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describes important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. The fields are: </p> <ol> <li>generic variable name: lc = Land cover</li> <li>variable procedure combination: mcd12q1v061.t1 = MCD12Q1 v061 LC Type1 band</li> <li>Position in the probability distribution / variable type: c = class | p = probability</li> <li>Spatial support: 500m</li> <li>Depth reference: s = surface</li> <li>Time reference begin time: 20010101 = 2001-01-01</li> <li>Time reference end time: 20011231 = 2001-12-31</li> <li>Bounding box: go = global (without Antarctica)</li> <li>EPSG code: epsg.4326 = EPSG:4326</li> <li>Version code: v20230818 = creation date</li> </ol>
Metabolism dataset: one year of high-frequency temperature, dissolved oxygen, wind, photosynthetically active radiation observations and low-frequency nutrient data for 58 lakes in the Global Lake Ecological Observatory Network
Understanding controls on primary productivity is essential for describing ecosystems and their responses to environmental change. Lake primary production is strongly controlled by inputs of nutrients and colored dissolved organic matter. While past studies have developed mathematical models of this nutrient-color paradigm, broad empirical tests of these models are scarce. We compiled data from 58 diverse and globally distributed and mostly temperate lakes to test such a model and improve understanding and prediction of the controls on lake primary production. These lakes varied widely in size (0.02-2300 km2), pelagic gross primary production (20-8000 mg C m-2 d-1), and other characteristics. The data package includes high-frequency dissolved oxygen, water temperature, wind speed, and solar radiation data as well as daily estimates of GPP and ER derived from those data. In addition, the data package includes median in-lake and stream concentrations of dissolved organic carbon and total phosphorus for a subset of 18 of those lakes.
Species cover, community biomass, and richness in global grasslands from NutNet (2007–2023): Dominant species predict plant richness and biomass in global grasslands
The Nutrient Network (NutNet) is a globally coordinated research initiative designed to investigate the impacts of human-driven alterations in nutrient availability and consumer presence on grassland ecosystems. Data were collected from over 130 herbaceous-dominated sites worldwide, spanning diverse environmental conditions from desert grasslands to arctic tundra. Standardized methodologies were employed across all sites to enable direct comparisons of productivity, diversity, and ecosystem responses. Experimental treatments included nutrient additions to assess co-limitation of plant growth by multiple nutrients, as well as grazer manipulations to examine their role in regulating biomass, species diversity, and community composition. By compiling these cross-site data, NutNet aims to enhance our understanding of productivity-diversity relationships and provide new insights into the ecological consequences of anthropogenic changes to nutrient cycles and food webs at a global scale.
Insights on global rangeland ecosystem services shaped by grazing and fertilization (2007-2021)
The Nutrient Network (NutNet) is a globally coordinated research initiative aimed at investigating the impacts of human-induced changes in nutrient availability and consumer presence on grassland ecosystems. In this study, we used data from 79 grassland sites participating in NutNet, which includes a factorial experiment involving herbivory exclusion and/or nutrient addition. Standardized methodologies were applied across all sites to facilitate direct comparisons of response variables. We used ecosystem variables to quantify three provisioning ecosystem services (forage quantity, forage chemical quality, and forage physical quality), three supporting services (forage stability, soil fertility, and soil stability), and eight regulating services (erosion control, control of soil acidification, regulation of water quantity and quality, carbon storage, resistance to plant invasion, pest control, and pollination). Additionally, we identified three plant biodiversity variables that are closely related to the provisioning of ecosystem services (alpha richness, beta diversity, and native diversity). Using this data, we quantified key ecosystem services provided by rangelands, assessed both short- and long-term impacts of grazing exclusion and fertilization on these services, and identified synergies and trade-offs between them.
A Comprehensive Global Aquatic N2O Emission Database (GANED): Unravelling N2O Emission Patterns from Different Water Bodies, 1980-2023
The Global Aquatic Nitrous Oxide Emission Database (GANED) is a comprehensive synthesis of empirical observations of N2O concentration measurements and flux records, spanning the period 1980-2023. GANED advances N2O research by providing the first global systematic emission mechanisms among the different aquatic system types, including rivers, streams, estuaries, reservoirs, ponds, lakes, open seas and coastal areas. The N2O data in GANED is further interconnected with biogeochemical metadata on dissolved oxygen, dissolved organic carbon, ammonium, nitrate, nitrite, total nitrogen, water temperature, salinity and pH, along with site data (latitude, longitude, codes of channel type, depth, surface area, elevation). The dataset explains the discrepancy that emission of N2O in aquatic bodies is determined mainly by substrate availability, and not by climatic factors, and reveals the systematic biases of concentration-only measurements, which can result in an underestimation of fluxes in effluent water of dynamically changing aquatic waters. Consequently, GANED constitutes a crucial transition “where” emissions occur to understanding “why” they differ across systems, and thus enabling targeted mitigation interventions. GANED includes 5130 records of N2O concentration and 7386 flux measurements from 3,002 unique sites, most of which are resolved to the daily time scale.
Global Climate Change Impacts on the Vegetation and Fauna of Mangrove Forested Ecosystems in Florida (FCE): Nekton Portion from March 2000 to April 2004
Depth is measured at 3 random locations within each net at time of set. All other variables (salinity, temperature, dissolved oxygen) are measured at the river bank adjacent to each net also at the time of set. Minimum and maximum values for sites were found to be: Salinity(ppt) = SRSMc-S2: 0.3-14.7, SRSMc-S3: 15.6-34.4, SRSMc-S4: 2.4-34; Water temp(degrees C)= SRSMc-S2: 22.2-31.5, SRSMc-S3: 16.6-31.1, SRSMc-S4: 21.1-30.6; DO(mg/l)= SRSMc-S2: 2.55-5.27, SRSMc-S3: 2.08-5.3, SRSMc-S4: 1.25-4.2; Mean depth(cm)= SRSMc-S2: 0.0-24.6, SRSMc-S3: 5.7-41.5, SRSMc-S4: 0.0-21.4
Global Climate Change Impacts on the Vegetation and Fauna of Mangrove Forested Ecosystems in Florida (FCE): Nekton Mass from March 2000 to April 2004
Bottomless lift nets are buried within the mangrove forest floor and raised remotely on slack high spring tides to enclose a 6m2 area. As the tide ebbs, fishes retreat into a subtidal refuge cleared when the tide has fallen. Three replicate nets have been sampled at 3 locations along a salinity gradient on Shark River for 4 years. Small resident forage fish and grass shrimp dominate the collections. Exotic species and estuarine transient species that use the estuary as a nursery are rare within the assemblage of fishes that routinely use the flooded forest.
Daily phenocam image data and derived timeseries for global change experiments at the Jornada Basin LTER site, 2014-2020
This dataset contains daily data extracted from phenocams installed at a global exchange experiment involving Chihuahuan desert plant communities at the Jornada Basin LTER site in southern New Mexico, U.S.A. Cycles of plant growth, termed phenology, are tightly linked to environmental controls, and our overarching objective in this study is to determine if temperature or precipitation are relatively more important for determining shrub and grass greenup date (start of season) and senescence date (end of season). At these camera locations, we experimentally manipulated incoming precipitation at the Jornada Basin LTER for over a decade and recorded plant leaf phenology at the daily scale for seven years using phenocams. The data included here comes from phenocams installed in two ongoing studies at the Jornada Basin LTER site, one studying ecosystem responses to long term changes in water and nitrogen availability, and one studying plant productivity and partitioning responses to water availability and herbivory (studies 349 and 456, respectively). Phenocams at the sites have collected images since 2014, and this dataset includes color values extracted from shrub and grass regions in these images. Further analyses, including daily values of calculated greenness (green chromatic coordinate), precipitation, and temperature, for all the plots included in the study are in EDI dataset knb-lter-jrn.210574002. This study is ongoing.
Globally distributed lake surface water temperatures collected in situ and by satellites; 1985-2009
Global environmental change has influenced lake surface temperatures, a key driver of ecosystem structure and function. Recent studies have suggested significant warming of water temperatures in individual lakes across many different regions around the world. However, the spatial and temporal coherence associated with the magnitude of these trends remains unclear. Thus, a global dataset of water temperature is required to understand and synthesize global, long-term trends in surface water temperatures of inland bodies of water. We assembled a database of summer lake surface temperatures for 291 lakes collected in situ and/or by satellites for the period 1985-2009. In addition, corresponding climatic drivers (air temperatures, solar radiation, and cloud cover) and geomorphometric characteristics (latitude, longitude, elevation, lake surface area, maximum depth, mean depth, and volume) that influence lake surface temperatures were compiled for each lake. This unique dataset offers an invaluable baseline perspective on global-scale lake thermal conditions as environmental change continues. This dataset accompanies a data publication in the journal Scientific Data
A Global database of methane concentrations and atmospheric fluxes for streams and rivers
This dataset, referred to as MethDB, is a collation of publicly available values of methane (CH4) concentrations and atmospheric fluxes for world streams and rivers, along with supporting information on location, geographic, physical, and chemical conditions of the study sites. The data set is composed of four linked tables, corresponding to the data sources (Papers_MethDB), the study sites (Sites_MethDB), concentrations (Concentrations_MethDB), and influx/efflux rates (Fluxes_MethDB). Information was extracted from journal articles, government reports, book chapters, and similar sources that were acquired before 15 September 2015. Concentrations and fluxes were converted to a standard unit (micromoles per liter for concentration and millimoles per square meter per day for flux) and both the author-reported and converted data are included in the database. MethDB was assembled as part of a larger synthesis effort on stream and river CH4 dynamics, and assembled data were used to identify large-scale patterns and potential drivers of fluvial CH4 and to generate an updated global-scale estimate of CH4 emissions from world rivers.
Supplementary Material for A Global Analysis of Dark Matter Signals from 27 Dwarf Spheroidal Galaxies using 11 Years of Fermi-LAT Observations
<p><strong>Description of the Supplementary Data</strong></p> <p>This record contains tabulated Bayesian and frequentist exclusion limits, profile likelihood maps and posterior probability maps for the publication S. Hoof, A. Geringer-Sameth, and R. Trotta, “<i>A Global Analysis of Dark Matter Signals from 27 Dwarf Spheroidal Galaxies using 11 Years of Fermi-LAT Observations</i>,” <a href="https://doi.org/10.1088/1475-7516/2020/02/012">JCAP 02 (2020) 012</a> (also available on the <a href="https://arxiv.org/abs/1812.06986">arXiv</a>). The dwarf spheroidal galaxies considered in this work are (in alphabetical order): Aquarius II, Boötes I, Canes Venatici I, Canes Venatici II, Carina, Carina II, Coma Berenices, Draco, Draco II, Fornax, Grus I, Hercules, Horologium I, Leo I, Leo II, Leo IV, Leo V, Pegasus III, Pisces II, Reticulum II, Sculptor, Segue 1, Sextans, Tucana II, Ursa Major I, Ursa Major II, and Ursa Minor.</p> <p>This record consists of the following files, which correspond to the limits presented Figures 9 and 10 of the paper. The files can be downloaded individually or obtained by downloading and unpacking the <code>record_2612268.zip</code>. In what follows,<code><strong>[CHANNEL]</strong></code> refers to the annihilation channel used, i.e. <i>e<sup>+</sup> e<sup>-</sup></i>, <i>μ<sup>+</sup> μ<sup>-</sup></i>, <i>τ<sup>+</sup> τ<sup>-</sup></i>, <i>b b̄</i>, <i>c c̄</i>, <i>t t̄</i>, <i>g g</i>, <i>W<sup>+</sup> W<sup>-</sup></i>, and <i>Z Z</i>. We also provide a simple plotting script for <code>Python</code>, named <code>plotting_script.py</code>, which provides basic plotting routines for all files.</p> <ul> <li>One-dimensional limits on <i><σ v></i>. The files <code>oneD_frequentist_limits_<strong>[CHANNEL]</strong>_channel.txt</code> contain the frequentist limits (at 95% confidence level, 1 degree of freedom) given the value of the WIMP mass <i>m<sub>χ</sub></i> tabulated there. The files <code>oneD_Bayesian_limits_<strong>[CHANNEL]</strong>_channel.txt</code> contain the Bayesian limit (95% credibility conditioned on the mass <i>m<sub>χ</sub></i> tabulated there).</li> <li>Two-dimensional grid of profile likelihood values. The files <code>twoD_profile_likelihood_map_<strong>[CHANNEL]</strong>_channel.txt</code> contain the natural logarithm of the profile likelihood w.r.t. the global best-fit likelihood value for that channel together with the corresponding values of <i>m<sub>χ</sub></i> and <i><σ v></i>. Note that for obtaining the limits in Fig. 10, which are conditioned on the WIMP mass, one needs to rescale the profile likelihood values with the maximum profile likelihood for a given WIMP mass.</li> <li>Two-dimensional grid of posterior probabilities for each combination of <i>m<sub>χ</sub></i> and <i><σ v></i>. The files <code>twoD_posterior_probability_map_<strong>[CHANNEL]</strong>_channel.txt</code> contain probabilities (obtained using a log-uniform prior on <i><σ v></i>) together with the corresponding values of <i>m<sub>χ</sub></i> and <i><σ v></i>. The tabulated values of <i>m<sub>χ</sub></i> and <i><σ v></i> correspond to the centres of the respective bins in <i>m<sub>χ</sub></i> and <i><σ v></i> and the posterior probability contained in them (the total posterior probability sums to 1).</li> </ul> <p>Please contact the authors if you require different data or have any questions regarding this data set.</p>
Data set for Global quantitative synthesis of ecosystem functioning across climatic zones and ecosystem types
<p>Dataset used in the publication: " Global quantitative synthesis of ecosystem functioning across climatic zones and ecosystem types". The dataset gathers estimates of ecosystem standing stocks (biomass, organic carbon, detritus), fluxes (GPP, ER, NEP) and process rates (decomposition and carbon uptake rates) for eight broad ecosystem types (forest, grassland, agroecosystem, desert, stream, lake, pelagic and benthic marine ecosystems) in five broad climatic zones (arctic, boreal, arid, temperate, tropical, arid).</p> <p>The scripts to produce the figures and the statistics of the publication are released along with the txt version of the data, which file is uploaded when running the script.</p>
Global consensus map of human transcription factor footprints
<p>Vierstra, J. <em>et al.</em> <strong>Global reference mapping of human transcription factor footprints.</strong> <em>Nature</em><strong> </strong>583, 729–736 (2020). <a href="https://doi.org/10.1038/s41586-020-2528-x">https://doi.org/10.1038/s41586-020-2528-x</a></p> <p>Preprint @ bioRxiv: <a href="https://doi.org/10.1101/2020.01.31.927798">https://doi.org/10.1101/2020.01.31.927798</a></p> <p><strong>Contact:</strong> Jeff Vierstra (<a href="mailto:jvierstra@altius.org?subject=Consensus%20DNase%20I%20footprints">jvierstra@altius.org</a>)</p> <p>Genomic DNase I footprinting enables quantitative, nucleotide-resolution delineation of sites of transcription factor occupancy within native chromatin. We combined sampling of >67 billion uniquely mapping DNase I cleavages from >240 human cell types and states to index, with unprecedented accuracy and resolution, human genomic footprints and thereby the sequence elements that encode transcription factor recognition sites.</p> <p>Please see <a href="http://vierstra.org/resources/dgf">http://vierstra.org/resources/dgf </a>for additional information and a complete set of raw DNase I data for individual datasets. Additionally, raw data can also be accessed via the ENCODE data portal (<a href="http://encodeproject.org">http://encodeproject.org</a>) using the dataset accessions found in Supplementary Table 1.</p> <p>Code for footprint analysis and tutorials on how to access and manipulate digital genomic footprint data can be found at <a href="https://footprint-tools.readthedocs.io/en/latest/">https://footprint-tools.readthedocs.io/en/latest/</a>.</p> <p>All files herein correspond to human genome build version GRCh38 (UCSC hg38).</p> <p><strong>Dataset contents:</strong></p> <ul> <li><strong>Biosample metadata</strong> – Supplementary_Table_1.xlsx</li> <li><strong>Motif clustering metadata </strong>– Supplementary_Table_2.xlsx</li> <li><strong>ChIP-seq validation metadata </strong>–<strong> </strong>Supplementary_Table_3.xlsx</li> <li><strong>Consensus footprint coordinates and assigned motif archetypes</strong><br> TSV file (BED-format) with consensus footprint (posterior probability>0.99) coordinates and overlaps with matches to motif model clusters. The legend file contains column definitions in detail. <ul> <li>consensus_footprints_and_motifs_hg38.bed.gz</li> <li>consensus_footprints_and_motifs_legend.txt</li> </ul> </li> <li><strong>Motif archetype matches overlapping consensus footprints</strong><br> TSV file (BED-format) containing the coordinates for clustered motif model matches that overlap consensus footprints <ul> <li>collapsed_motifs_overlaping_consensus_footprints.bed.gz</li> <li>collapsed_motifs_overlaping_consensus_footprints_legend.txt</li> </ul> </li> <li><strong>Footprint occupancy matrix of consensus footprints</strong><br> Rows are same order as the consensus footprint file and columns are same order as in the metadata files. <ul> <li>consensus_index_matrix_full_hg38.txt.gz (Values are –log(1-posterior))</li> <li>consensus_index_matrix_binary_hg38.txt.gz (binary occupancy matrix, where footprints with posterior footprint probability >0.99 are considered occupied)</li> </ul> </li> <li><strong>Single nucleotide variants tested for allelic imbalance </strong><br> The legend file contains column definitions in detail. <ul> <li>genotypes.vcf.gz - Genotyping and allelic read depth for each biosample (see header for more information)</li> <li>tested_snvs_padj.bed.gz - SNVs tested for imbalance (TSV, BED-format)</li> <li>tested_snvs_padj_legend.txt</li> </ul> </li> </ul>
Data and code release for Carleton, Cornetet, Huybers, Meng & Proctor (PNAS, 2020), "Global evidence for ultraviolet radiation decreasing COVID-19 growth rates"
<p>This upload contains all replication material for "Global evidence for ultraviolet radiation decreasing COVID-19 growth rates" (PNAS, 2020). Please note that previous versions of this upload provided data and code for the pre-print version of the article, which changed somewhat through the peer review process. </p> <p><strong>Authors:</strong> Tamma Carleton, Jules Cornetet, Peter Huybers, Kyle C. Meng, Jonathan Proctor.</p> <p><strong>Code is located within CCHMP_covid_climate_code_release.zip</strong>, and is written in R, Stata, and Matlab. The working directory should be set to the repository folder at the top of each script (all other filepaths are relative).</p> <p>Please find the code needed to replicate the main findings of the paper described below:</p> <ul> <li>Plots of data: R and Stata scripts to make figures 1B, 2A/B/C, S1, S2, and S3, can be found within “code/analysis/data_plots/”.</li> <li>Regression analysis: Stata scripts to run the distributed lag regressions and plot the results in figures 2, 3C, S5, S6, S7, S8, S10, and S14, as well as Table S1, can be found within “code/analysis/regressions/”. R scripts for data analysis and plotting for figures 3A/B and S9 are also within "code/analysis/regressions/".</li> <li>Seasonal simulations: R and Stata scripts to replicate the seasonal simulation shown in figures 4, S4 and S11 can be found within “code/analysis/seasonal_sim/”.</li> <li>SEIR simulations: Matlab scripts to replicate the SEIR simulations shown in figures S12 and S13 can be found within “code/analysis/SEIR/”.</li> </ul> <p><strong>Data are located within CCHMP_covid_climate_data_release.zip.</strong></p>
IPBES Data Management Tutorials - Session 6.2: Literature review from the Global Assessment chapter 4
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The chapter on<em> Examples of implementing the IPBES data management Policy</em> contains examples of how certain data management tasks and workflows were implemented within IPBES so that they follow the data management policy. <strong>Currently, this chapter contains legacy videos and the most recent examples can be found within the IPBES technical guidelines here:</strong> <a href="https://ict.ipbes.net/ipbes-ict-guide/data-management/technical-guidelines">https://ict.ipbes.net/ipbes-ict-guide/data-management/technical-guidelines</a></p> <p>This session <em>Literature review from the Global Assessment chapter 4 </em>walks you through each step of the data management of the systematic literature review from the Chapter 4 of the Global Assessment. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.