Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,386
datasets available to search
ShareScore release 0.9.0
Dataset results
1,386 results for “summary”
Locally verified monthly summary temperature and precipitation data from a NOAA weather station at USDA Jornada Experimental Range headquarters, southern New Mexico USA, 1914-1998
This data package contains locally verified monthly meteorological observations from a NOAA National Weather Service station located at the USDA Jornada Experimental Range headquarters in southern New Mexico, USA. Monthly summary data (based on daily observations) has been collected there by USDA staff since 1914 for minimum and maximum air temperature and daily accumulated precipitation using standard U.S. climatological service instrumentation and procedures. The included data were verified and transcribed directly from the original paper data sheets and have undergone quality control and assurance procedures different than those in place at NOAA. These data therefore differ from those directly downloadable from NOAA servers. Local verification and transcription of observations from the data sheets ceased in 1998 and data are now directly entered to the NOAA system. Therefore, this dataset is complete and will no longer be added to. All observations from this weather station have also undergone NOAA QA/QC procedures and those data are available by accessing the Jornada Experimental Range, NM US GHCN station through the National Climatic Data Center portal https://www.ncdc.noaa.gov/cdo-web/datasets/GSOM/stations/GHCND:USC00294426/detail - daily and monthly data are available).
Summary statistics and annual trends in chloride concentration in urban Minnesota lakes and streams
These data tables describe statistical summaries of chloride concentration, temporal trends in annual chloride, and projected risk of future chloride pollution in lakes and streams of the 17 most urban counties in Minnesota. Data were summarized separately for lakes and for streams, and include statistics (mean, median, standard deviation, upper and lower confidence intervals, and maximum) over the entire data record, over the warm season (May - October), and over the most recent 5 years (i.e., since 2018). Trends were computed on annual means and medians. For lakes, data were aggregated by lake basin (MN DNR Lake ID, or DOW) as well as by depth of sample (surface and deep). For streams, data were aggregated by individual site level as well as by stream reach (per MN Pollution Control Agency assessment units). Risk of chloride pollution was also determined for sites with longer records (10+ years) based on current concentration, number of exceedances of chronic standards, and projected chloride concentration based on current trends. Raw data were extracted from two sources: (1) the National Water Quality Portal (USGS & EPA) and (2) the Metropolitan Council Environmental Information Management System. The data were originally collected by a large number of entities, including watershed management authorities in the state of Minnesota, tribal groups, the Minnesota Pollution Control Agency, municipalities, university researchers, private consultants, and the Metropolitan Council. Some data records begin as early as the 1950's or 1960's, with many sites still including active data collection. A total of approximately 45,000 observations of chloride were included for lakes and wetlands, and approximately 70,000 observations for streams. Nearly 1600 stream sites and 700 lake/wetland sites were represented in the raw data, with 356 stream sites and 600 lakes represented in the summaries (after filtering out sites with less than 1 year of data collection). Primary data retr
CETAF-DiSSCo/COVID19-TAF biodiversity-related knowledge hub working group: indexed biotic interactions and review summary
<p>This data publication originated as part of developing a biodiversity-related knowledge hub on COVID-19 via COVID19-TAF - Communities Taking Action (https://cetaf.org/covid19-taf-communities-taking-action), a community-rooted initiative raised jointly by the Consortium of European Taxonomic Facilitaties (CETAF, https://cetaf.org) and Distributed Systems of Scientific Collections (DiSSCo, https://www.dissco.eu/).</p> <p>This archive contains the biodiversity datasets of interest identified in period 14 April-6 October 2020 through COVID19-TAF activities and subsequently indexed by Global Biotic Interactions (GloBI, https://globalbioticinteractions.org). GloBI provides open access to finding species interaction data (e.g., predator-prey, pollinator-plant, virus-host, parasite-host) by combining existing open datasets using open source software.</p> <p>These identified datasets (see references and reviews below) add to a growing collection of open species interaction datasets already indexed by GloBI. So, this data publication only includes a small subset of indexed datasets and include only datasets that were added as a direct consequence of COVID19-TAF activities of the biodiversity-related knowledge hub working group.</p> <p>If you have questions or comments about this publication, please open an issue at https://github.com/ParasiteTracker/tpt-reporting or contact the authors by email.</p> <p>Funding:<br> The creation of this archive was made possible in part by reporting software developed as part of the National Science Foundation award "Collaborative Research: Digitization TCN: Digitizing collections to trace parasite-host associations and predict the spread of vector-borne disease," Award numbers DBI:1901932 and DBI:1901926 . Also, this material is based upon work supported by the National Science Foundation under Grant No. DGE-1545433 .</p> <p>References:<br> Jorrit H. Poelen, James D. Simons and Chris J. Mungall. (2014). Global Biotic Interactions: An open infrastructure to share and analyze species-interaction datasets. Ecological Informatics. https://doi.org/10.1016/j.ecoinf.2014.08.005.</p> <p>GloBI Data Review Report</p> <p>Datasets under review:<br> - Geiselman, Cullen K. & Sarah Younger. 2020. Bat Eco-Interactions Database. www.batbase.org accessed via https://github.com/globalbioticinteractions/batbase/archive/9c65cfeee1a054f9db8cd8bf6892017fd1b3c840.zip on 2020-10-04T22:53:45.576Z<br> - Geiselman, Cullen K. and Tuli I. Defex. 2015. Bat Eco-Interactions Database. www.batplant.org accessed via https://github.com/globalbioticinteractions/batplant/archive/a2e1b57052244d5251d17e96ea61f58bea88975e.zip on 2020-10-04T22:54:28.727Z<br> - Daniel Becker, Gregory F Albery, Anna R Sjodin, Timothee Poisot, Tad Dallas, Evan A. Eskew, Maxwell J. Farrell, Sarah Guth, Barbara A Han, Nancy B Simmons, Colin J Carlson. 2020. Predicting wildlife hosts of betacoronaviruses for SARS-CoV-2 sampling prioritization. bioRxiv 2020.05.22.111344; doi: https://doi.org/10.1101/2020.05.22.111344 accessed via https://github.com/globalbioticinteractions/becker2020/archive/47c6ad28e1c5058f3c13ca69a59fdf21229e8d7f.zip on 2020-10-04T22:54:46.723Z<br> - Chen L, Liu B, Yang J, Jin Q, 2014. DBatVir: the database of bat-associated viruses. Database (Oxford). 2014:bau021. doi:10.1093/database/bau021 accessed via https://github.com/globalbioticinteractions/dbatvir/archive/a906d76e362484d3ca1edbe9683f672838ab70b0.zip on 2020-10-04T22:56:13.913Z<br> - Chen L, Liu B, Wu Z, Jin Q, Yang J, 2017. DRodVir: A resource for exploring the virome diversity in rodents. J Genet Genomics. 44(5):259-264. accessed via https://github.com/globalbioticinteractions/drodvir/archive/0346c0e8d4d66c6400e9965bd6a6aeed24cd7586.zip on 2020-10-04T23:06:04.368Z<br> - Agosti, Donat. 2020. Transcription of Linné, C. von, 1758. Systema naturae per regna tria naturae secundum classes, ordines, genera, species, cum characteribus, differentiis, synonymis, locis. Available at: http://dx.doi.org/10.5962/bhl.title.542 . accessed via https://github.com/globalbioticinteractions/linnaeus1758/archive/a818060080fa04a88dac6df1ae5b897304ae8877.zip on 2020-10-05T00:46:04.852Z<br> - Mollentze, Nardus, & Streicker, Daniel G. (2019). Viral zoonotic risk is homogenous among taxonomic orders of mammalian and avian reservoir hosts (Version 1.0.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.3516613 accessed via https://github.com/globalbioticinteractions/mollentze2019/archive/ad12dc74d03c3d992618f16c37cafb7f7ffd9d01.zip on 2020-10-05T00:50:55.878Z<br> - Eneida L. Hatcher, Sergey A. Zhdanov, Yiming Bao, Olga Blinkova, Eric P. Nawrocki, Yuri Ostapchuck, Alejandro A. Schäffer, J. Rodney Brister, Virus Variation Resource – improved response to emergent viral outbreaks, Nucleic Acids Research, Volume 45, Issue D1, January 2017, Pages D482–D490, https://doi.org/10.1093/nar/gkw1065 . accessed via https://github.com/globalbioticinteractions/ncbi-virus/archive/531a8d743d7adcf1153a19087e5d3c5b76750e3e.zip on 2020-10-05T00:53:53.646Z<br> - Olival, K. J., Hosseini, P. R., Zambrana-Torrelio, C., Ross, N., Bogich, T. L., & Daszak, P. (2017). Host and viral traits predict zoonotic spillover from mammals. Nature, 546(7660), 646–650. doi:10.1038/nature22975 accessed via https://github.com/globalbioticinteractions/olival2017/archive/f61070a5339d0e6c6e76d7eb4e2102decb52317d.zip on 2020-10-05T00:56:43.356Z<br> - Pensoft Darwin Core Archives with associateTaxa columns accessed via https://github.com/globalbioticinteractions/pensoft-dwca/archive/ee8831a2a391203f4fa8c05a0ddd927202b234bf.zip on 2020-10-05T00:56:51.868Z<br> - Pensoft Darwin Core Archives available via Integrated Publication Toolkit accessed via https://github.com/globalbioticinteractions/pensoft-ipt/archive/4ad4b47978324681289e36f8c2b247b1bcc97b1a.zip on 2020-10-05T00:58:01.912Z<br> - De Rojas M, Doña J, Dimov I (2020) A comprehensive survey of Rhinonyssid mites (Mesostigmata: Rhinonyssidae) in Northwest Russia: New mite-host associations and prevalence data. Biodiversity Data Journal 8: e49535. https://doi.org/10.3897/BDJ.8.e49535 accessed via https://github.com/globalbioticinteractions/pensoft-table/archive/3488e0397ca4e083d5eca6949951e426a75713e3.zip on 2020-10-05T00:58:03.647Z<br> - Marcus Guidoti, Tatiana Ruschel, Donat Agosti. 2020. Corona virus related biotic associations manually extracted from literature. Plazi. accessed via https://github.com/globalbioticinteractions/plazi-covid19/archive/326578b0d9f974760dcd2e962d86636a6487a6c0.zip on 2020-10-05T00:58:08.025Z<br> - Shaw, LP, Wang, AD, Dylus, D, et al. The phylogenetic range of bacterial and viral pathogens of vertebrates. Mol Ecol. 2020; 29: 3361– 3379. https://doi.org/10.1111/mec.15463 accessed via https://github.com/globalbioticinteractions/shaw2020/archive/bb9ab857b7fdbb4e931752d01b43d37b3ada77cf.zip on 2020-10-05T01:05:23.554Z<br> - OpenBiodiv. 2020. Annotated biotic interaction tables from Pensoft publications. accessed via https://github.com/pensoft/pensoft-interaction-tables/archive/bb7d1dc9f2eba220a61502e06e6114053fd30788.zip on 2020-10-05T03:03:23.372Z<br> - Quentin J. Groom. 2020. Bat interation data manually extracted from literature. accessed via https://github.com/qgroom/batinterations/archive/70108945f9014aa0ac1db920191867f7e151c793.zip on 2020-10-05T03:04:11.533Z</p> <p>Generated on:<br> 2020-10-06</p> <p>by:<br> GloBI's Elton 0.10.2<br> (see https://github.com/globalbioticinteractions/elton).</p> <p> </p> <p>Note that all files ending with .tsv are files formatted<br> as UTF8 encoded tab-separated values files.</p> <p>https://www.iana.org/assignments/media-types/text/tab-separated-values</p> <p><br> Included in this review archive are:</p> <p>README:<br> This file.</p> <p>review_summary.tsv:<br> Summary across all reviewed collections of total number of distinct review comments.</p> <p>review_summary_by_collection.tsv:<br> Summary by reviewed collection of total number of distinct review comments.</p> <p>indexed_interactions_by_collection.tsv:<br> Summary of number of indexed interaction records by institutionCode and collectionCode.</p> <p>review_comments.tsv.gz:<br> All review comments by collection.</p> <p>indexed_interactions_full.tsv.gz:<br> All indexed interactions for all reviewed collections.</p> <p>indexed_interactions_simple.tsv.gz:<br> All indexed interactions for all reviewed collections selecting only sourceInstitutionCode, sourceCollectionCode, sourceCatalogNumber, sourceTaxonName, interactionTypeName and targetTaxonName.</p> <p>datasets_under_review.tsv:<br> Details on the datasets under review.</p> <p>elton.jar:<br> Program used to update datasets and generate the review reports and associated indexed interactions.</p> <p><br> datasets.zip:<br> source datasets collected by elton in process of executing the generate_report.sh script.</p> <p>generate_report.sh:<br> program used to generate the report</p> <p>generate_report.log:<br> log file generated as part of running the generate_report.sh script</p>
Summary raw meteorological data from the Southern Ocean collected on board the Antarctic Circumnavigation Expedition (ACE) during the austral summer of 2016/2017.
<p><strong>Dataset abstract</strong></p> <p>A Vaisala MAWS240 meteorological station was installed on the R/V Akademik Tryoshnikov during a circumnavigation of Antarctica in the austral summer season of 2016/2017. This dataset contains the raw meteorological data that have been extracted from the original raw text data files. Data coverage is from 17th November 2016 until 11th April 2017, with gaps where the ship was in port.</p> <p>Air temperature, relative humidity, dew point, solar radiation, ultraviolet radiation, cloud level and sky cover were recorded with a resolution of 30 seconds. Averaged wind parameter data are provided.</p> <p>Date_time should be combined with TIMEDIFF to convert it to UTC. Latitude and longitude recorded are not corrected. Underway seawater measurements were recorded as null values.</p> <p>Data from this dataset have been corrected and quality-checked in another published dataset. We recommend these data for further use (Landwehr et al., 2019; DOI 10.5281/zenodo.3379590).</p> <p><strong>Dataset contents</strong></p> <ul> <li>metdata_all_YYYYMMDD_YYYYMMDD.csv, data file, comma-separated values</li> <li>data_file_header, metadata, text format</li> <li>README.txt, metadata, text format</li> <li>ace_meteorology_raw_summary_change_log.txt</li> </ul> <p>Data files contain data for each leg of the Antarctic Circumnavigation Expedition (ACE). Dates included in the file name are the start and end dates of the legs and therefore the data within the files as well.</p> <p><strong>Change log</strong></p> <p><strong>v1.2</strong> - Added missing data from 2017-02-05 - 2017-02-08 inclusive. Updated this change log file.</p> <p><strong>v1.1</strong> - Added additional data coverage from 2016-11-17 - 2016-11-22 inclusive, into the first data file. Updated README.txt with information about data coverage. Added this change log file.</p> <p><strong>v1.0</strong> - Initial release of raw summary meteorological data.</p> <p><strong>Dataset license</strong></p> <p>This raw meteorological dataset is made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0) whose full text can be found at https://creativecommons.org/licenses/by/4.0/</p>
State of Wildfires 2024-25: Regional Summaries of Burned Area, Fire Emissions, and Individual Fire Characteristics for National, Administrative and Biogeographical Regions
<p>This dataset supports the State of Wildfires 2024-25 report under review at <em>Earth System Science Data</em> (Kelley et al., <em>under review)</em>. It is an update of the State of Wildfires 2023-24 report (Jones et al. 2024). The dataset provides annual data and final-year anomalies in burned area (BA), fire carbon (C) emissions, and fire properties (e.g. distributional statistics for fire count, size, rate of growth). Annual data relate to the global fire season defined as March-February (e.g., March 2024-February 2025), aligning with an annuall lull in the global fire calendar (see Jones et al., 2024). The complete methodology is described by Kelley et al. (<em>under review</em>).</p> <h3>Citation</h3> <p>Work utilising our regional summaries should <strong>cite both Kelley et al. (under review) AND the primary reference for the variable(s) of interest</strong> as follows:</p> <ul> <li>Giglio et al. (2018) for MODIS MCD64A1 BA.</li> <li>van der Werf et al. (2017) for GFED4.1s fire C emissions.</li> <li>Kaiser er al. (2012) for GFAS fire C emissions.</li> <li>van der Werf et al. (2017) AND Kaiser er al. (2012) for the average of GFED4.1s and GFAS fire C emissions.</li> <li>Andela et al. (2019) for the Global Fire Atlas.</li> <li>Giglio et al. (2016) for the Fire Radiative Power (FRP) observations.</li> <li>Chuvieco et al. (2024) for FireCCIS311 BA.</li> <li>Giglio et al. (2024) for VIIRS VNP64A1 BA.</li> </ul> <h3>Input Data</h3> <p><strong>Burned Area (BA)</strong></p> <ul> <li>BA data from NASA’s MODIS BA product (MCD64A1) are extended from Giglio et al. (2018) and are available from <a href="https://lpdaac.usgs.gov/products/mcd64a1v061/">Giglio et al. (2021)</a>. <ul> <li>Period: 2002-February 2025</li> <li>Resolution: 500m, daily</li> </ul> </li> <li>BA data from ESA's Climate Change Initiative BA product (FireCCIS311) are extended from Lizundia-Loiola et al. (2022) and are available from <a href="Chuvieco,%20E.;%20Pettinari,%20M.L.;%20Lizundia-Loiola,%20J.;%20Khairoun,%20A.;%20Danne,%20O.;%20Boettcher,%20M.;%20Storm,%20T.%20(2024):%20ESA%20Fire%20Climate%20Change%20Initiative%20(Fire_cci):%20Sentinel-3%20SYN%20Burned%20Area%20Grid%20product,%20version%201.1.%20NERC%20EDS%20Centre%20for%20Environmental%20Data%20Analysis,%2029%20February%202024.%20https://catalogue.ceda.ac.uk/uuid/da8e669a74334c82a56e0b470bc4ef04">Chuvieco et al. (2024)</a>. <ul> <li>Period: 2019-February 2025</li> <li>Resolution: 300m, daily</li> </ul> </li> <li>BA data from NASA’s VIIRS BA product (VNP64A1) are available from <a href="https://lpdaac.usgs.gov/products/vnp64a1v002/">Giglio et al. (2024)</a>. <ul> <li>Period: 2012-February 2025 (only the data after 2019 are used for consistency in the comparisons between MCD64A1, FireCCIS311, and VNP64A1).</li> <li>Resolution: 500m, daily</li> </ul> </li> </ul> <p><strong>Fire Carbon (C) Emissions</strong></p> <ul> <li>GFED4.1s fire C emissions data are extended from van der Werf and are available at <a href="https://globalfiredata.org/">https://globalfiredata.org/</a>. <ul> <li>Period: 2003-February 2025</li> <li>Resolution: 0.25 degree, daily</li> </ul> </li> </ul> <ul> <li>GFAS fire C emissions data are extended from Kaiser et al. (2012) and are available from the <a href="https://confluence.ecmwf.int/display/CKB/CAMS+global+biomass+burning+emissions+based+on+fire+radiative+power+%28GFAS%29%3A+data+documentation">ECMWF Confluence Server</a>. <ul> <li>Period: 2003-February 2025</li> <li>Resolution: 0.1 degree, daily</li> </ul> </li> </ul> <p><strong>Global Fire Atlas (Individual Fire Properties)</strong></p> <ul> <li>Global Fire Atlas data are extended from Andela et al. (2019) and are available from the repository maintained by <a href="https://doi.org/10.5281/zenodo.11400062">Andela and Jones (2025)</a>. <br> <ul> <li>Period: 2002-February 2025</li> <li>Driven by 500m MODIS BA data (collection 6.1)</li> </ul> </li> </ul> <p><strong>Fire Intensities</strong></p> <ul> <li>FRP data are extended from MOD14A1 and MYD14A1 (Giglio et al., 2016) and are available at <a href="https://lpdaac.usgs.gov/products/mod14a1v061/">Giglio and Justice (2021)</a>.<br> <ul> <li>Period: 2002-February 2025</li> <li>Resolution: 1km, daily</li> </ul> </li> </ul> <h3>Regional Analysis</h3> <p>We performed "cookie-cutting" (spatial and temporal masking) of the above input data sets to features in each of the following regional layers (e.g. per country in the "Countries" layer). </p> <p>The statistics derived from cookie-cutting are listed below. Full details in Kelley et al. (2025).</p> <div> <table> <tbody> <tr> <td> <p>Layer</p> </td> <td> <p>Short Form </p> </td> <td> <p>Source</p> </td> </tr> <tr> <td> <p>Biomes</p> </td> <td> <p>NA</p> </td> <td> <p>Olson et al. (2001)</p> </td> </tr> <tr> <td> <p>Ecoregions</p> </td> <td> <p>NA</p> </td> <td> <p>Olson et al. (2001)</p> </td> </tr> <tr> <td> <p>Continents</p> </td> <td> <p>NA</p> </td> <td> <p>ArcGIS Hub (2024)</p> </td> </tr> <tr> <td> <p>Continental Biomes</p> </td> <td> <p>NA</p> </td> <td> <p>See above</p> </td> </tr> <tr> <td> <p>Countries</p> </td> <td> <p>NA</p> </td> <td> <p>EU Eurostat (2020)</p> </td> </tr> <tr> <td> <p>UC Davis Global Administrative Areas (GADM) Level 1</p> </td> <td> <p>GADM-L1</p> </td> <td> <p>UC Davis (2022)</p> <br><br></td> </tr> <tr> <td> <p>Intergovernmental Panel on Climate Change Sixth Assessment Report (AR6) Working Group I (WGI) Reference Regions </p> </td> <td> <p>IPCC AR6 WGI Regions</p> </td> <td> <p>Iturbide et al. (2020)</p> </td> </tr> <tr> <td> <p>Global C Project Regional C Cycle Assessment and Processes (RECCAP2) Reference Regions</p> </td> <td> <p>RECCAP2 Regions</p> </td> <td> <p>Ciais et al. (2022)</p> </td> </tr> <tr> <td> <p>Global Fire Emissions Database (GFED) Basis Regions</p> </td> <td> <p>GFED4.1s Regions</p> </td> <td> <p>van der Werf et al. (2006)</p> </td> </tr> </tbody> </table> </div> <h3> </h3> <h3>Regional Statistics and Anomalies</h3> <ul> <li><strong>Burned Area (BA)</strong> <ul> <li>Calculated regional totals for each fire season.</li> <li>Relative and standardized anomalies from historical data (since 2002).</li> <li>Ranking amongst all recorded fire seasons.</li> <li>Onset, peak, and cessation based on monthly deviations from climatological means.</li> </ul> </li> </ul> <ul> <li><strong>Carbon Emissions</strong> <ul> <li>Calculated regional totals for each fire season.</li> <li>Relative and standardized anomalies from historical data (since 2003).</li> <li>Ranking amongst all recorded fire seasons.</li> <li>Onset, peak, and cessation based on monthly deviations from climatological means.</li> <li>Statistics available for GFAS, GFED, and their mean.</li> </ul> </li> </ul> <ul> <li><strong>Individual Fire Properties</strong> <ul> <li>Based on values of individual fire size and rate of growth ignition from the ignition point vectors of the Global Fire Atlas.</li> <li>Calculated regional count.</li> <li>Calculated regional maxima and 95th percentiles of fire size and rate of growth for each fire season.</li> <li>Relative and standardized anomalies from historical data (since 2002).</li> <li>Ranked anomalies among all recorded fire seasons.</li> </ul> </li> </ul> <ul> <li><strong>Fire Intensity</strong> <ul> <li>Based on active fire observations of FRP, which are pooled within each fire of the Global Fire Atlas.</li> <li>For each fire, the 95th percentile value of all FRP observations is the assigned intensity value (i.e. a "peak fire intensity" omitting any spurious high-end values).</li> <li>Regionally, the peak fire intensity values are averaged across individual fires.</li> <li>Relative and standardized anomalies from historical data (since 2002).</li> <li>Ranked anomalies among all recorded fire seasons.</li> </ul> </li> </ul>
INFORMATE Project - CHORUS Report Summaries - 20231106
<p>These data provide a summary of the All, Author Affiliation, and Dataset Reports generated by the <a href="https://dashboard.chorusaccess.org/">CHORUS Dashboard</a> for three agencies: the U.S. National Science Foundation, U.S. Geological Survey, and the U.S. Agency for International Development. The reports summarized here was collected on November 6-7, 2023 as part of the INFORMATE Project funded by NSF.</p><p>The columns are:</p><p>Column Definition</p><p>agency The funding agency [NSF, USGS, or USAID]</p><p>date. The date of data retrieval (YYYYMMDD)</p><p>report. The report [all, authors, datasets]</p><p>Property Name of the column in the input file</p><p>count Number of values (rows) of the property</p><p>unique Number of unique values of the property</p><p>top Most common value of the property</p><p>freq Number of occurrences (frequency) of the most common value</p><p>Count % The percentage of rows that include the property</p>
GWAS Summary Statistics from "Sex and statin-related genetic associations at the PCSK9 gene locus – results of genome-wide association meta-analysis"
<p>GWAMA summary statistics of PCSK9 levels stratified by sex and statin useage in Europeans.</p> <p>When using this data, please cite:</p> <p>Pott, J., Kheirkhah, A., Gadin, J.R. <em>et al.</em> Sex and statin-related genetic associations at the <em>PCSK9</em> gene locus: results of genome-wide association meta-analysis. <em>Biol Sex Differ</em> <strong>15</strong>, 26 (2024). https://doi.org/10.1186/s13293-024-00602-6</p> <p>All txt files contain the following columns:</p> <ul> <li>markername (unique SNP ID)</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>EA (effect allele)</li> <li>OA (other allele)</li> <li>EAF (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (in case of double-stratified data: number of studies; in case of single-stratified data: 2, as it is a meta-analysis of the two double-stratified data sets)</li> <li>beta (effect estimate)</li> <li>SE (standard error)</li> <li>pval (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>invalidAssoc (TRUE/FALSE flag if this variant was excluded in our analysis)</li> <li>reason4exclusion (reason why this SNP was excluded)</li> <li>phenotype (phenotyp setting)</li> </ul>
Results complementing the European Union summary report on surveillance for the presence of transmissible spongiform encephalopathies (TSE) - Croatia
<p>This dataset contains TSE surveillance results in cattle, sheep, goats, cervids and other species, and genotyping in sheep, pursuant to Regulation (EC) 999/2001.</p> <p><strong>Reporting authorities contributing to each data collection</strong>:</p> <ul> <li>TSE_2023_HR: Ministry of Agriculture (MPS)</li> <li>TSE_2022_HR: Ministry of Agriculture (MPS)</li> <li>TSE_2021_HR: Ministry of Agriculture (MPS)</li> <li>TSE_2020_HR: Ministry of Agriculture (MPS)</li> <li>TSE_2019_HR: Ministry of Agriculture (MPS)</li> </ul>
Improving anxiety research novel approach to reveal trait anxiety through summary measures of multiple states - raw count data set - RNAseq
<p>Raw count data of the RNAseq analysis of a project and manuscript under the title "Improving anxiety research novel approach to reveal trait anxiety through summary measures of multiple states". The header of the table includes the subject identifiers except the first column "genes". The latter column includes all assessed gene identifiers.</p>
Results complementing the European Union summary report on surveillance for the presence of transmissible spongiform encephalopathies (TSE) - Austria
<p>This dataset contains TSE surveillance results in cattle, sheep, goats, cervids and other species, and genotyping in sheep, pursuant to Regulation (EC) 999/2001.</p> <p><strong>Reporting authorities contributing to each data collection</strong>:</p> <ul> <li>TSE_2023_AT: Austrian Agency for Health and Food Safety (AGES)</li> <li>TSE_2022_AT: Austrian Agency for Health and Food Safety (AGES)</li> <li>TSE_2021_AT: Austrian Agency for Health and Food Safety (AGES)</li> <li>TSE_2020_AT: Austrian Agency for Health and Food Safety (AGES)</li> <li>TSE_2019_AT: Austrian Agency for Health and Food Safety (AGES)</li> </ul>
Genome- and transcriptome-wide association summary statistics for outcome from traumatic brain injury
<p>The dataset contains summary statistics for the genome- and transcriptome-wide association studies (GWAS, TWAS) of genetic effects on outcome in traumatic brain injury (TBI). The study participants attended hospital within 24 hours of TBI, and underwent head computed tomography imaging.</p> <p><strong>Study participants</strong></p> <p>European ancestry data set contains 4710 individuals; multi-ethnic cohort 5268 individuals, including Europeans (n = 4710), Africans (n = 245) and Admixed Americans (n = 313).</p> <p>The largest European population contribution was from CENTER-TBI (Collaborative European NeuroTrauma Effectiveness Research, https://www.center-tbi.eu), where each participating center (60 centers from 20 countries in Europe) recruited patients between December 2013 and December 2017. The patients recruited in CENTER-TBI were supplemented by subjects from cohorts recruited at two European centres (Cambridge, UK, and Turku, Finland).</p> <p>The majority of patients in the US cohort were recruited between 2014 and 2018 to TRACK-TBI (Transforming Research and Clinical Knowledge in TBI, https://tracktbi.ucsf.edu) by the 18 US participant sites. The subjects recruited to the US cohort from TRACK-TBI were supplemented by patients recruited to an institutional research initiative at Mass General Brigham (MGB).</p> <p><strong>Outcome definition</strong></p> <p>Outcomes were measured using the extended Glasgow Outcome Scale (GOSE), ranging from 1 (dead) to 8 (upper good recovery), measured 6 months post-TBI. TBI severity was specified using the Glasgow Coma Score (GCS), with TBI classified as mild (GCS 13-15), moderate (GCS 9-12), or severe (GCS 3-8).</p> <p>To account for the effect of injury severity on outcome, sliding dichotomization was used to categorize outcome as favourable or unfavourable. A GOSE ≤ 4 was used to define an unfavourable outcome for patients with either moderate (GCS 9-12) or severe (GCS 3-8) TBI, while the unfavourable group was extended to patients with GOSE ≤ 7 if they had mild (GCS 13-15) TBI.</p> <p><strong>Genotype data and imputation</strong></p> <p>Genotyping was completed at FIMM Technology Center for CENTER-TBI, Cambridge, Turku patients and the Broad Institute for TRACK-TBI, using the Illumina Global Screening Array (GSA-24v2-0 + Multi-Disease). The MGB cohort were genotyped using Illumina’s Multi-Ethnic Global array (MEGA) and the pre-releases forms, including MEGA and MEGA-Ex arrays at Illumina at the MGB Translational Genomics Core.</p> <p>A unified quality control procedure was applied for each study cohort and the array-based genotypes were imputed using the Haplotype Reference Consortium panel. Autosomal chromosomes were considered, post-imputation data was filtered by imputation quality (INFO > 0.4 for CENTER-TBI, Cambridge and Turku; R2 > 0.4 for TRACK-TBI and MGB) and MAF > 1%.</p> <p><strong>Genome-wide association analysis and meta-analysis</strong></p> <p>Genome-wide single-marker scans were performed using a penalized likelihood-based Firth logistic regression, and implemented in PLINK v2.0. Using favourable outcome as reference, models were fitted on the basis of imputed allelic dosages. Age, sex, major extracranial injury, pupillary reactivity, and the first 10 principal components were included as covariates. Study cohort (CENTER-TBI, Cambridge, Turku) was an additional covariate in the CENTER-TBI GWAS.</p> <p>Fixed-effects meta-analysis of the three European ancestry GWAS was performed using METAL. For trans-ethnic meta-analysis, summary statistics of five GWASs in patients of European, African and Admixed Americans were aggregated via MR-MEGA.</p> <p><strong>Transcriptome-wide association study</strong></p> <p>Genetically regulated gene expression (GREx) was imputed using a regression model fitted on a separate gene expression database. Elastic net models provided by PrediXcan for all available GTEx brain tissues and whole blood were used. For TWAS, the same sliding dichotomy model for outcome with the same set of covariates as in the GWAS, but PCA components were replaced with the top five principal components of the respective gene expression data. </p> <p><strong>Column headers - GWAS</strong></p> <p>rsID: variant rsID<br> Chrom: chromosome<br> Pos: position (build GRCh38)<br> A1: effect allele<br> A2: reference allele<br> EAF: allele frequency of effect allele<br> Effect: effect size of effect allele<br> StdErr: standard error of effect size<br> P: p value of association (with genomic correction)<br> N: sample size</p> <p>Note. 'Effect' and 'StdErr' are only available for the European ancestry meta-analysis.</p> <p><br> <strong>Column headers - TWAS</strong></p> <p>tissue: GTEx tissue type<br> id: ensembl gene id<br> coef: model coefficient<br> se: model standard error for coefficient<br> p: model-based p value<br> symbol: gene symbol<br> name: gene name written out<br> chr: chromosome<br> start: gene start position (build GRCh38)</p>
Summary of the most important features for selected ABs
<p>These data summarizes the relevant findings and the identified limitations (in terms of "Category", "Technology", "Properties", "Limitation", and "Applicability to railway"), coming from the overview of different Alternative Bearers (ABs), carried out in deliverable D21 (AB4Rail project, www.ab4rail.eu).<br> The results have provided an overview of several technologies, each of them showing specific characteristics. The heterogeneous nature of different ABs allows to provide a plethora of available communication technologies to be potentially used by the Adaptable Communication System (ACS) for different railway scenarios. All the selected ABs provide the IP interconnection feature since they are Integrated within OSI reference model.<br> In this way, it collects the planned objectives of deliverable D2.1, expressed as a technological overview of selected ABs, as possible candidates coexisting with Traditional Bearers (TBs) for supporting railway applications.</p>
Summary statistics accompanying the article "Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency" in Scientific Reports (2022)
<p>Summary statistics for genome-wide association studies reported in:</p> <p>Bell, S., Tozer, D.J., & Markus H.S. (2022). Genome-wide association study of the human brain functional connectome reveals strong vascular component underlying global network efficiency. <em>Scientific Reports</em>, DOI: <a href="https://dx.doi.org/10.1038/s41598-022-19106-7">10.1038/s41598-022-19106-7</a>. </p> <p><strong>Abstract</strong></p> <p>Complex brain networks play a central role in integrating activity across the human brain, and such networks can be identified in the absence of any external stimulus. We performed 10 genome-wide association studies of resting state network measures of intrinsic brain activity in up to 36,150 participants of European ancestry in the UK Biobank. We found that the heritability of global network efficiency was largely explained by blood oxygen level-dependent (BOLD) resting state fluctuation amplitudes (RSFA), which are thought to reflect the vascular component of the BOLD signal. RSFA itself had a significant genetic component and we identified 24 genomic loci associated with RSFA, 157 genes whose predicted expression correlated with it, and 3 proteins in the dorsolateral prefrontal cortex and 4 in plasma. We observed correlations with cardiovascular traits, and single-cell RNA specificity analyses revealed enrichment of vascular related cells. Our analyses also revealed a potential role of lipid transport, store-operated calcium channel activity, and inositol 1,4,5-trisphosphate binding in resting-state BOLD fluctuations. We conclude that that the heritability of global network efficiency is largely explained by the vascular component of the BOLD response as ascertained by RSFA, which itself has a significant genetic component.</p> <p> </p> <p>Further information on the files uploaded here can be found in the README. Users interested in bulk downloading these summary statistics may find <a href="https://github.com/dvolgyes/zenodo_get">zenodo_get</a> helpful.</p>
IPBES Assessment of the diverse values and valuation of nature - Figures presented in the summary for policymakers
<p>These figures are an integral part of the Summary for policymakers of the Methodological assessment of the diverse values and valuation of nature of the Intergovernmental Science-Policy Platform on Biodiversity and Ecosystem Services. To see the full document visit the related links. </p>
Meningitis hGWAS results (summary statistics)
<p>Summary statistics for association between genetic variation and meningitis phenotypes. Contains human genome association and interaction effects (pGWAS.tar.bz2).</p> <p>Unpack with `tar xf `. Contents are described in the README.</p>
ERA5-derived daily temperature summary 1980-2018
<p>Hourly air temperature at surface data from the ERA5 reanalysis at 0.5˚grid resolution was summarised to produce daily mean, minimum, and maximum temperatures.</p>
Voxel-level summary statistics of hippocampus shape, white matter microstructure, and cortical surface curvature in UK Biobank (n=33,324)
<p>This deposit hosts GWAS summary statistics of hippocampus shape (n=33,324), white matter microstructure (n=33,324), and cortical surface curvature (n=15,752) using UKB unrelated white subjects. The data was generated by using the highly efficient imaging genetics (<a href="https://github.com/Zhiwen-Owen-Jiang/heig">HEIG v1.1.0</a>) framework where only the triplets - summary statistics of low-dimensional representations (LDRs), the functional bases, and the variance-covariance matrix LDRs - are shared, which is sufficient to recover all voxel-variant pairs as well as to conduct voxel-level heritability and (cross-trait) genetic correlation analysis. Check the <a href="https://github.com/Zhiwen-Owen-Jiang/heig/wiki">tutorial</a> and the <a href="../records/13770930">example data</a> used in the tutorial. </p> <p>The shared data includes:</p> <p>1. Triplets for hippocampus shape measured by the radial distance from the medial model for each vertex. The original images contain 30,000 vertices while the shared data contains 49 LDRs. Left and right hemispheres were analyzed separately, each with 15,000 vertices.</p> <p>2. Triplets for 21 white matter tracts measured by fractional anisotropy. The original images contain 32,217 voxels and each tract contains 88 ~ 3503 voxels while the shared data contains 1,034 LDRs. Tracts were analyzed separately.</p> <p>3. Triplets for cortical surface curvature. The original images contain 59,412 vertices while the shared data contains 1,750 LDRs. The entire brain was analyzed as a whole.</p> <p>4. LD matrix and its inverse for 22 chromosomes including 460k genotyped SNPs. LD matrix and its inverse were estimated by using two separate datasets each containing 8.4k white unrelated subjects in UKB. Two regularization levels are provided: {85%, 80%} for heritability and genetic correlations within images and {75%, 70%} for cross-trait genetic correlations.</p> <p>5. LD matrix and its inverse for 22 chromosomes including 1.2 million imputed HapMap3 SNPs. LD matrix and its inverse were estimated by using two separate datasets each containing 42k white unrelated subjects in UKB. Two regularization levels are provided: {98%, 95%} for heritability and genetic correlations within images and {90%, 85%} for cross-trait genetic correlations.</p>
GWAS summary statistics for waist-to-hip ratio and body principal components
<p>This dataset contains genome-wide association summary statistics for waist-to-hip ratio (WHR), as well as those for body principal components (PCs). A subset of 387,139 unrelated, white British individuals were analyzed for WHR. PCs were combined from the summary statistics for WHR and 13 other anthropometric traits (body mass index, standing height, weight, hip circumference, waist circumference, arm lean mass (left), arm fat mass (left), leg lean mass (left), leg fat mass (left), trunk lean mass, trunk fat mass, body fat percentage, basal metabolic rate) provided by the Neale lab (http://www.nealelab.is/uk-biobank). All traits were inverse-rank normal transformed (by the Neale lab or ourselves for WHR).</p> <p>All effect sizes, including those for PCs, are standardized, i.e. they represent the effects on a trait with variance 1.</p> <p>The zip files contain the data to run the sample pipeline and the shiny app, both available from <a href="https://github.com/JonSulc/PCA_Cross-sex_MR">https://github.com/JonSulc/PCA_Cross-sex_MR</a>.</p>
Base rates of food safety practices in European households: Summary data from the SafeConsume Household Survey
<p>This data set contains estimates of the base rates of 550 food safety-relevant food handling practices in European households. The data are representative for the population of private households in the ten European countries in which the SafeConsume Household Survey was conducted (Denmark, France, Germany, Greece, Hungary, Norway, Portugal, Romania, Spain, UK).</p> <p><em>Sampling design</em></p> <p>In each of the ten EU and EEA countries where the survey was conducted (Denmark, France, Germany, Greece, Hungary, Norway, Portugal, Romania, Spain, UK), the population under study was defined as the private households in the country. Sampling was based on a stratified random design, with the NUTS2 statistical regions of Europe and the education level of the target respondent as stratum variables. The target sample size was 1000 households per country, with selection probability within each country proportional to stratum size.</p> <p><em>Fieldwork</em></p> <p>The fieldwork was conducted between December 2018 and April 2019 in ten EU and EEA countries (Denmark, France, Germany, Greece, Hungary, Norway, Portugal, Romania, Spain, United Kingdom). The target respondent in each household was the person with main or shared responsibility for food shopping in the household. The fieldwork was sub-contracted to a professional research provider (Dynata, formerly Research Now SSI). Complete responses were obtained from altogether 9996 households.</p> <p><em>Weights</em></p> <p>In addition to the SafeConsume Household Survey data, population data from Eurostat (2019) were used to calculate weights. These were calculated with NUTS2 region as the stratification variable and assigned an influence to each observation in each stratum that was proportional to how many households in the population stratum a household in the sample stratum represented. The weights were used in the estimation of all base rates included in the data set.</p> <p><em>Transformations</em></p> <p>All survey variables were normalised to the [0,1] range before the analysis. Responses to food frequency questions were transformed into the proportion of all meals consumed during a year where the meal contained the respective food item. Responses to questions with 11-point Juster probability scales as the response format were transformed into numerical probabilities. Responses to questions with time (hours, days, weeks) or temperature (C) as response formats were discretised using supervised binning. The thresholds best separating between the bins were chosen on the basis of five-fold cross-validated decision trees. The binned versions of these variables, and all other input variables with multiple categorical response options (either with a check-all-that-apply or forced-choice response format) were transformed into sets of binary features, with a value 1 assigned if the respective response option had been checked, 0 otherwise.</p> <p><em>Treatment of missing values</em></p> <p>In many cases, a missing value on a feature logically implies that the respective data point should have a value of zero. If, for example, a participant in the SafeConsume Household Survey had indicated that a particular food was not consumed in their household, the participant was not presented with any other questions related to that food, which automatically results in missing values on all features representing the responses to the skipped questions. However, zero consumption would also imply a zero probability that the respective food is consumed undercooked. In such cases, missing values were replaced with a value of 0.</p>
QTL summary statistics from the DIRECT consortium
<p>These are the complete summary statistics for DIRECT genotype-phenotypes associations (QTLs). The project performed genotype-phenotype associations for gene expression (RNAseq), targeted proteins (Olink), targeted metabolites (Biocrates) and untargeted metabolites (Metabolon) derived from 3,029 blood and plasma samples from the DIRECT cohort. This submission includes supplementary files and nominal pvalues (as uncorrected pvalues) for all associations included in the manuscript. Trans associations included are typically limited to pvalues <1e-04. Network tables are also included, with information to load and use Cytoscape to visualize them. This is version 2, some files were missing on version 1.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.