Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
311
datasets available to search
ShareScore release 0.7.1
Dataset results
311 results for “Empirical data”
AMOC reconstruction between 1981 and 2016 from hydrographic data using an empirical linear regression model from Worthington, E. L., Moat, B. I., Smeed, D. A., Mecking, J. V., Marsh, R., and McCarthy, G. D.: A 30-year reconstruction of the Atlantic meridional overturning circulation shows no decline, Ocean Sci., 17, 285–299, https://doi.org/10.5194/os-17-285-2021, 2021.
<p>Dataset used to create Figure 8 in Worthington et al., 2021 (https://doi.org/10.5194/os-17-285-2021). Details of the data and methods can be found in the journal article.<br> <br> Worthington, E. L., Moat, B. I., Smeed, D. A., Mecking, J. V., Marsh, R., and McCarthy, G. D.: A 30-year reconstruction of the Atlantic meridional overturning circulation shows no decline, Ocean Sci., 17, 285–299, <a href="https://doi.org/10.5194/os-17-285-2021">https://doi.org/10.5194/os-17-285-2021</a>, 2021.</p>
Data for "On the Practice of Semantic Versioning for Ansible Galaxy Roles: An Empirical Study and a Change Classification Model"
<p>This dataset accompanies a replication package provided for a study on Semantic Versioning for Ansible Galaxy roles.</p> <p>The replication package is available at https://github.com/ROpdebee/ansible_semver_ext_replication</p>
Data supporting 'Empirical correction of systematic orthorectification error in Sentinel-2 velocity fields for Greenlandic outlet glaciers'
<p><strong>Note: An updated dataset covering the majority of Greenland's marine-terminating glaciers is available as part of the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) project through the National Snow and Ice Data Center (NSIDC) at <a href="https://doi.org/10.5067/B28FM2QVVYWY">https://doi.org/10.5067/B28FM2QVVYWY</a>. </strong></p> <p>Data supporting the paper:</p> <blockquote> <p>Chudley, T. R., Howat, I. M., Yadav, B. N., & Noh, M. J. (2022). Empirical correction of systematic orthorectification error in Sentinel-2 velocity fields for Greenlandic outlet glaciers. <em>The Cryosphere. </em>16, 2629–2642, https://doi.org/10.5194/tc-16-2629-2022</p> </blockquote> <p>Dataset consists of four netCDF files containing stacked Sentinel-2 velocity data of four Greenlandic outlet glaciers (Helheim Glacier, Jakobshavn Isbræ, Store Glacier, and Kangerlussuaq) between 2017 and 2021. Velocity data are derived and corrected following the methods outlined in Chudley <em>et al.</em> (2022). </p> <p>NetCDF files are created by, and tested to be readable by, Python's xarray package.</p> <p>The dimensions of the netCDF file are as follows:</p> <ul> <li><strong>X</strong> - <em>x </em>coordinates in NSDIC Sea Ice Polar Stereographic North (EPSG:3413).</li> <li><strong>Y</strong> - <em>y</em> coordinates in NSDIC Sea Ice Polar Stereographic North (EPSG:3413).</li> <li><strong>time</strong> - temporal midpoint of velocity field.</li> </ul> <p>The variables of the netCDF file are as follows:</p> <ul> <li><strong>dmag</strong> - the absolute magnitude of the velocity, in metres per day.</li> <li><strong>dx</strong> - the velocity in the <em>x</em> direction, in metres per day.</li> <li><strong>dy</strong> - the velocity in the <em>y</em> direction, in metres per day.</li> <li><strong>date1</strong> - the date and time of the first scene acquisition.</li> <li><strong>date2</strong> - the date and time of the second scene acquisition.</li> <li><strong>baseline</strong> - the temporal baseline, in days, between scene acquisitions.</li> <li><strong>orbit_pair</strong> - the combination of orbital pathways in the string format 'RXXX_RYYY', where XXX is relative orbit number of the first scene and YYY the relative orbit number of the second scene.</li> <li><strong>mag_rmse</strong> - the root mean square error of the absolute velocity of the off-ice area. </li> <li><strong>dx_mean</strong> - the mean velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dx_sd</strong> - the standard deviation of the velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dy_mean</strong> - the mean velocity of the off-ice area in the <em>x</em> direction.</li> <li><strong>dy_sd</strong> - the standard deviation of the velocity of the off-ice area in the <em>y</em> direction.</li> </ul>
AusEFlux: Empirical upscaling of OzFlux eddy covariance flux tower data over Australia
<p>AusEFlux (<strong>Aus</strong>tralian <strong>E</strong>mpirical <strong>Flux</strong>es) is a high resolution (500 metre) gridded estimate of Gross Primary Productivity (GPP), Ecosystem Respiration (ER), Net Ecosystem Exchange (NEE), and Evapotranspiration over the Australian continent for the period January 2003 to Present. These datasets provide a benchmark for assessment against Land Surface Model simulations, and a means for monitoring of Australia’s terrestrial carbon cycle at an unprecedented high-resolution.</p> <p><strong>Version 2.1 </strong>of AusEFlux has just been released (as of May 2025) and was created to <strong>operationalise</strong> the research datasets published in this <a href="https://doi.org/10.5194/bg-20-4109-2023">EGU Biogeosciences publication.</a> In order to operationalise these datasets, changes to the input datasets were required to align the data sources with datasets that are regularly and reliably updated, along with general improvements. The datasets provided on Zenodo have been reprojected to 5 km resolution to facilitate easier uploading and sharing, but<strong> full resolution datasets (both v1.1 and v2.1) can be accessed freely through <a href="https://thredds.nci.org.au/thredds/catalog/ub8/au/AusEFlux/catalog.html">NCI's THREDDS portal.</a></strong></p> <p><strong>Two Jupyter Notebooks</strong> have been created (one for GPP and one for NEE) that demonstrate the differences between the research datasets (v1.1) and the operational datasets (v2.1), including showing the differences in specifications and inputs. You can view/download these notebooks using the links below:</p> <p><a href="https://nbviewer.org/github/cbur24/AusEFlux/blob/master/notebooks/analysis/Compare_AusEFlux_versions_GPP.ipynb">GPP comparison between versions</a></p> <p><a href="https://nbviewer.org/github/cbur24/AusEFlux/blob/master/notebooks/analysis/Compare_AusEFlux_versions_NEE.ipynb">NEE comparisons between versions</a></p> <p>Each dataset contains three variables:</p> <ul> <li>"<flux>_median": represents the 'best-estimate' of a given flux, the units are gC/m<sup><sub>2</sub></sup>/mon<sup>-1</sup></li> <li>"<flux>_25th_percentile": represents the lower uncertainty bound, the units are gC/m<sup><sub>2</sub></sup>/mon<sup>-1</sup></li> <li>"<flux>_75th_percentile": represents the upper uncertainty bound, the units are gC/m<sup><sub>2</sub></sup>/mon<sup>-1</sup></li> </ul> <p><span><strong>Version Guide</strong>:</span></p> <p><em>v1.0:</em> DO NOT USE THIS VERSION. There was a mistake in the modelling of ecosystem respiration, so this version of the dataset should not be used. As of version 1.1, the error has been rectified.</p> <p><em>v1.1: </em>This version of the datasets are those used to inform the EGU Publication linked above. Its time range is 2003-July 2022, and its spatial resolution is 5 km on Zenodo, but the 1 km resolution datasets can be accessed through <a href="https://thredds.nci.org.au/thredds/catalog/ub8/au/AusEFlux/catalog.html">NCI's THREDDS portal</a>.</p> <p><em>v2.0: <strong>IMPORTANT NOTE:</strong> a bug in the modelling of vegetation height resulted in data artefacts in the NEE and ER fluxes over very tall mesic forests in this version. This resulted in unnaturally high ER and lower than expected NEE (less negative than would be expected). This issue has been rectified in version 2.1.</em> <strong>Version 2 datasets represent the operational version of the datasets</strong>, it includes several improvements over version 1.1. Its time-range is 2003-2024 (and will be updated annually), and its spatial resolution is 500m. A 5 km reprojected version of the dataset is included here on Zenodo, but the 500 metre datasets can be accessed through<a href="https://thredds.nci.org.au/thredds/catalog/ub8/au/AusEFlux/catalog.html"> NCI's THREDDS portal.</a></p> <p><strong>v2.1: </strong>This version is a patch to version 2.0 to remove a bug in the modelling of vegetation height. <strong>It is recommended to use this version </strong>over v2.0. 500 metre resolution datasets can be accessed through<a href="https://thredds.nci.org.au/thredds/catalog/ub8/au/AusEFlux/catalog.html"> NCI's THREDDS portal.</a></p>
Empirical data, qualitative codes, analysis: Schuur J.S. et al. Identifying levers of urban neighbourhood transformation. npj Urban Sustainability (2023)
<p>Please refer to the stand-alone "2023_SchuurJS_UrbanSustainabilityfinal.html" file where the analysis and results corresponding to the article titled: "Identifying levers of urban neighbourhood transformation using serious games" is presented. The underlying data sets and Rmarkdown script used for the analysis can be used to re-run the analysis. Ensure to read the "0_README.txt" file to build the appropriate folder structure to do so.</p>
Global Empirical Picture of Magnetospheric Substorms Inferred from Multi-Mission Magnetometer Data
<p>Data associated with Journal of Geophysical Research: Space Physics article titled: "Global Empirical Picture of Magnetospheric Substorms Inferred from Multi-Mission Magnetometer Data". This includes all the digital data that was used in constructing the Figures from the main and supplementary text, along with files containing the fit set of coefficients and parameters for the model, and files describing the subset of magnetometer used for fitting the model. </p>
An Empirical Study of Container Image Configurations and Their Impact on Start Times (Container Image Data)
<p>Dataset with the container image metadata used for our IEEE/ACM CCGRID 2023 paper "An Empirical Study of Container Image Configurations and Their Impact on Start Times".</p> <p>Abstract of the paper: A core selling point of application containers is their fast start times compared to other virtualization approaches like virtual machines. Predictable and fast container start times are crucial for improving and guaranteeing the performance of containerized cloud, serverless, and edge applications. While previous work has investigated container starts, there remains a lack of understanding of how start times may vary across container configurations. We address this shortcoming by presenting and analyzing a dataset of approximately 200,000 open-source Docker Hub images featuring different image configurations (e.g., image size and exposed ports). Leveraging this dataset, we investigate the start times of containers in two environments and identify the most influential features. Our experiments show that container start times can vary between hundreds of milliseconds and tens of seconds in the same environment. Moreover, we conclude that no single dominant configuration feature determines a container's start time and that hardware and software parameters must be considered together for an accurate assessment.</p> <p>Dataset description: Our images dataset contains 200,986 entries with 21 features associated to each container image. In the following, we describe the meaning of each feature. Further information is available in <a href="https://github.com/opencontainers/image-spec">OCI Image Specification</a> and the <a href="https://docs.docker.com/engine/reference/run/">Docker Run Documentation</a>. Besides the 20 features grouped in the five categories below, each dataset entry has a image_id, which is used to uniquely identify the dataset entry.</p> <p>Features</p> <p>Metadata features (prefix: meta)</p> <ul> <li><strong>meta_repo_digest</strong> : The repo digest is a SHA-256 hash which is used to uniquely identify and pull the image from Docker Hub</li> <li><strong>meta_architecture</strong> : The CPU architecture which the binaries in the image are built to run on</li> <li><strong>meta_os</strong> : The name of the operating system which the image is built to run on</li> <li><strong>meta_docker_version</strong> : The Docker version used to built this image</li> </ul> <p>I/O stream features (prefix: io)</p> <ul> <li><strong>io_attach_stdin</strong> : boolean setting to determine whether the console should be attached to the process stdin stream</li> <li><strong>io_attach_stdout</strong> : boolean setting to determine whether the console should be attached to the process stdout stream</li> <li><strong>io_attach_stderr</strong> : boolean setting to determine whether the console should be attached to the process stderr stream</li> <li><strong>io_tty</strong> : boolean setting to determine whether the console should pretend to be a TTY when attached</li> <li><strong>io_open_std_in</strong> : boolean setting to determine whether the process stdin stream should be kept open even if console not attached</li> <li><strong>io_std_in_once</strong> : boolean setting to determine whether the process retrieved input from the stdin stream at least once</li> </ul> <p>Start command features (prefix: cmd)</p> <ul> <li><strong>cmd_args</strong> : Length of list of arguments to use as the command to execute when the container starts</li> <li><strong>cmd_envvars</strong> : Environment variables set per default when the container starts</li> <li><strong>cmd_additional_args</strong> : Length of list for additional arguments to the containers entrypoint</li> </ul> <p>File system features (prefix: fs)</p> <ul> <li><strong>fs_volumes</strong> : Number of volumes to create/use by default</li> <li><strong>fs_size</strong> : Size of this image in bytes</li> <li><strong>fs_virtual_size</strong> : Virtual size of this image in bytes (equals size)</li> <li><strong>fs_graph_driver_name</strong> : Name of the image's graph driver</li> <li><strong>fs_root_fs_type</strong> : Name of the file system type used in the image</li> <li><strong>fs_layers</strong> : Number of root file system layers</li> </ul> <p>Networking features (prefix: net)</p> <ul> <li><strong>net_ports</strong> : Number of ports to expose per default</li> </ul> <p> </p> <p>Dataset acquisition: The dataset has been acquired from Docker Hub using a web crawler. We used substring matches with the <a href="https://hub.docker.com/explore">Docker Hub Explore function</a>. As search strings, we used all letter combination with sizes 1 to 3, meaning that our first search string was 'a' and our last was 'zzz'. We included both results from the 'recently updated' and the 'most popular' selection. We came up with an initial list of 286,294 image names. We then tested we could pull and start these images once. These tests have been conducted from April to June 2022. We sorted out all images that were either not pullable or startable and retrieved all total of 200,986 valid images. In the following, we describe the error types that we encountered and that let to the removal of the causing image from the dataset:</p> <ul> <li>The image manifest was unknown when we tried to download it meaning that is has been renamed or deleted from the time when our web crawler was running</li> <li>The entrypoint command required a dependency that was missing in the image and therefore the container could not be started</li> <li>The image did not specify an entrypoint command and could therefore not be started</li> <li>The image declared an invalid root file system type</li> <li>The image had a malformed root file system</li> <li>The image configuration was incomplete and therefore not all required data could be obtained</li> </ul> <p>See also our CodeOcean capsule with the processing scripts for our paper: https://doi.org/10.24433/CO.4595026.v2</p>
Data for "Revisiting the zonally asymmetric extratropical circulation of the Southern Hemisphere spring using complex empirical orthogonal functions"
<p>Data used in "Revisiting the zonally asymmetric extratropical circulation of the Southern Hemisphere spring using complex empirical orthogonal functions"</p>
Original dataset for "A validation of co-authorship credit models with empirical data from the contributions of PhD candidates"
<p><strong>Publication reference:</strong><br> Donner, P. (2020). A validation of co-authorship credit models with empirical data from the contributions of PhD candidates. Quantitative Science Studies, v. 1, i. 2, p. 551-564. <a href="https://doi.org/10.1162/qss_a_00048">https://doi.org/10.1162/qss_a_00048</a>.</p> <p> </p> <p>The file contains one row per authorship contribution statement. Rows of publications and theses are grouped.</p> <p><strong>Description of columns:</strong></p> <p>dissertation_id - an integer identifying each dissertation thesis</p> <p>university - university at which the dissertation thesis was written and PhD degree conferred</p> <p>year - publication year of the dissertation thesis</p> <p>author - dissertation thesis author name</p> <p>title - dissertation thesis title</p> <p>subject - the field of research</p> <p>publication_id - an integer identifying each publication; publication associated with more than one thesis have the same id across theses</p> <p>reference - bibliographic reference for the publication associated with the thesis</p> <p>author_count - number of authors of the publication</p> <p>author_position - position in the author byline of the credited author</p> <p>credit - claimed credit of the author in percent</p> <p>corresponding_author - flag for whether the publication author of this row is a orresponding author</p>
Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments -- part I
<p>Nearly all current Bayesian phylogenetic applications rely on Markov chain Monte Carlo (MCMC) methods to approximate the posterior distribution for trees and other parameters of the model. These approximations are only reliable if Markov chains adequately converge and sample from the joint posterior distribution. While several studies of phylogenetic MCMC convergence exist, these have focused on simulated datasets or select empirical examples. Therefore, much that is considered common knowledge about MCMC in empirical systems derives from a relatively small family of analyses under ideal conditions. To address this, we present an overview of commonly applied phylogenetic MCMC diagnostics and an assessment of patterns of these diagnostics across more than 18,000 empirical analyses. Many analyses appeared to perform well and failures in convergence were most likely to be detected using the average standard deviation of split frequencies, a diagnostic that compares topologies among independent chains. Different diagnostics yielded different information about failed convergence, demonstrating that multiple diagnostics must be employed to reliably detect problems. The number of taxa and average branch lengths in analyses have clear impacts on MCMC performance, with more taxa and shorter branches leading to more difficult convergence. We show that the usage of models that include both Γ-distributed among-site rate variation and a proportion of invariable sites are not broadly problematic for MCMC convergence but are also unnecessary. Changes to heating and the usage of model-averaged substitution models can both offer improved convergence in some cases, but neither are a panacea.</p>
Data from: Properties of Markov chain Monte Carlo performance across many empirical alignments --part II
<p>Nearly all current Bayesian phylogenetic applications rely on Markov chain Monte Carlo (MCMC) methods to approximate the posterior distribution for trees and other parameters of the model. These approximations are only reliable if Markov chains adequately converge and sample from the joint posterior distribution. While several studies of phylogenetic MCMC convergence exist, these have focused on simulated datasets or select empirical examples. Therefore, much that is considered common knowledge about MCMC in empirical systems derives from a relatively small family of analyses under ideal conditions. To address this, we present an overview of commonly applied phylogenetic MCMC diagnostics and an assessment of patterns of these diagnostics across more than 18,000 empirical analyses. Many analyses appeared to perform well and failures in convergence were most likely to be detected using the average standard deviation of split frequencies, a diagnostic that compares topologies among independent chains. Different diagnostics yielded different information about failed convergence, demonstrating that multiple diagnostics must be employed to reliably detect problems. The number of taxa and average branch lengths in analyses have clear impacts on MCMC performance, with more taxa and shorter branches leading to more difficult convergence. We show that the usage of models that include both Γ-distributed among-site rate variation and a proportion of invariable sites are not broadly problematic for MCMC convergence but are also unnecessary. Changes to heating and the usage of model-averaged substitution models can both offer improved convergence in some cases, but neither are a panacea.</p>
Data from: Temperature drives Zika virus transmission: evidence from empirical and mathematical models
<p>Temperature is a strong driver of vector-borne disease transmission. Yet, for emerging arboviruses we lack fundamental knowledge on the relationship between transmission and temperature. Current models rely on the untested assumption that Zika virus responds similarly to dengue virus, potentially limiting our ability to accurately predict the spread of Zika. We conducted experiments to estimate the thermal performance of Zika virus (ZIKV) in field-derived Aedes aegypti across eight constant temperatures. We observed strong, unimodal effects of temperature on vector competence, extrinsic incubation period, and mosquito survival. We used thermal responses of these traits to update an existing temperature-dependent model to infer temperature effects on ZIKV transmission. ZIKV transmission was optimized at 29oC, and had a thermal range of 22.7oC - 34.7oC. Thus, as temperatures move toward the predicted thermal optimum (29oC) due to climate change, urbanization, or seasonally, Zika could expand north and into longer seasons. In contrast, areas that are near the thermal optimum were predicted to experience a decrease in overall environmental suitability. We also demonstrate that the predicted thermal minimum for Zika transmission is 5oC warmer than that of dengue, and current global estimates on the environmental suitability for Zika are greatly over-predicting its possible range.</p>
Data and code corresponding to the article "Interaction network structure explains species temporal persistence in empirical plant-pollinator communities"
<p>This upload contains the Datasets and code to generate the results of the article "Interaction network structure explains species temporal persistence in empirical plant-pollinator communities".</p><p>The database comprises two files containing the abundances of plants and pollinators, and one containing the interaction networks among plants and pollinators. </p><p>The code folder contains the code to generate the results, and to generate the figures of the manuscript. </p>
Empirical data on growth and residency of juvenile Pacific salmon in North America estuaries
<p>Dataset compiling empirical data on estuarine growth and residency of Pacific salmon in North America. We conducted a systematic literature review to create this database, starting with a literature search in <i>Web of Science Core Collection</i> through Simon Fraser University's library proxy on April 24th, 2018, using the search parameters (salmon, Oncorhynchus) AND (estuary*) AND (residen* OR growth OR survival OR mortality), which returned 681 results (these results were presented in Arbeider 2018). We updated the search on March 29th, 2020, which produced 24 additional papers published after April 2018, and again on July 7th, 2022 which yielded another 31 papers published since March 2020. From these results, we extracted papers that research Pacific salmonids and whose study estuary was in North America. From this reduced list of papers, we extracted growth and residency parameters, as well as other relevant data. For complete methods please refer to:</p><p>Arbeider, M. et al. (In press). The estuarine growth and residency of juvenile Pacific salmon in North America: a compilation of empirical data. <i>Canadian Journal of Fisheries and Aquatic Sciences.</i></p>
Replication data for An Empirical Approximation of the Effects of Trade Sanctions with an Application to Russia
<p>This is the dataset to replicate all the tables and figures in the paper <a href="https://doi.org/10.1093/epolic/eiad027">"An Empirical Approximation of the Effects of Trade Sanctions with an Application to Russia"</a>, published in <i>Economic Policy</i>, 2023, by Jean Imbs and Laurent Pauwels. All data manipulations and programming are detailed on the GitHub site:<a href="https://github.com/laurentpauwels/sanctionpaper"> https://github.com/laurentpauwels/sanctionpaper</a>. The raw and processed data are in this <i>sanctionpaperdata_v1/matlab/data folder. </i>For convenience the simulation output <i>(simulationoutput.txt) </i>required to build the scatter plots in Figure 1 with STATA is available in<i> sanctionpaperdata</i>_v1<i>/matlab/output</i>.</p><p><strong>Instructions</strong> </p><p> If you clone the GitHub repository:</p><p>1. Place the downloaded <i>data</i> folder (located in <i>sanctionpaperdata_v1/matlab/)</i> in the <i>matlab</i> folder of the GitHub repository. </p><p>2. Place the downloaded <i>simulation_output.txt</i> I(located in <i>sanctionpaperdata_v1/matlab/output/) </i>in the <i>matlab/output </i>folder of the GitHub repository if you do not want to run the simulations as detailed on GitHub.</p><p><strong>Description</strong></p><p>The <i>matlab/data/raw</i> folder contains an <i>ICIO21</i> folder with the ICIO21 data, and a <i>WIOD</i> folder with the SEA16 data (in <i>data/raw/WIOD/SEA16</i>) and the WIOT16 data in CSV format (in <i>data/raw/WIOD/WIOT16</i>).</p><p>NOTE: WIOD provides the data in XLSB format. The XLSB WIOD data is in the <i>WIOT_in_EXCEL.zip</i> located in the <i>matlab/data/raw/WIOD/</i>. Python is used to convert XLSB into CSV files. See python code in GitHub repository for unzipping and conversion to CSV. The converted CSV files are provided for convenience.</p><p>The parsed and pre-processed ICIO21, SEA16, and WIOT16 data are stored in the <i>/matlab/data/processed</i> folder into three separate .mat structure files:</p><p><i>icio21_strc.mat</i> contains:</p><ul><li>the meta data (<i>icio21_text</i>), i.e., the information about the structure of the numerical data such as lists of countrycode, countries, industrycode, industries, isic_rev4 codes, years covered, name of final categories, etc.</li><li>the numerical data (<i>icio21_data</i>):<ul><li>Z (<i>icio21_data.Z</i>), the intermediate IO data for the listed industries (R), countries (N), and years (T). Its structure is 3-dimensionsal: (NxR)x(NxR)xT.</li><li>F (<i>icio21_data.F</i>), the final demand data for the same countries, industries and years. Its structure is 3-dimension: (NxR)x(NxC)xT. The columns are NxC where C are the number of final demand categories.</li></ul></li></ul><p><br><i>wiod16_strc.mat</i> has the same structure as <i>icio21_strc.ma</i>t with the meta data in <i>wiot16_text</i> and the numerical data in <i>wiot16_data</i>.</p><p><i>sea16_strc.mat</i> has the meta data in <i>sea16_text</i> and the numerical data in <i>sea16_data</i>. SEA16 contains 16 variables instead of Input-Output type data. The country, industry, and year coverage is not the same as ICIO21.</p><p>NOTE: <i>matlab/scripts/convertMatlabStruc2data.m</i> in the GitHub repository converts <i>MATLAB v7.3 </i>format ("structure data") to an updated format without structure so that it is more easily compatible with other software. All data parsing and preprocessing are done with MATLAB, see GitHub repository for details.</p><p><strong>Sources</strong></p><p>The raw data come from these sources:</p><p>1. OECD Inter-Country Input-Output (ICIO) data November 2021 release (downloaded on 2 July 2023)</p><p>- Source: OECD-ICIO 2021 release data is available at <a href="http://oe.cd/icio">http://oe.cd/icio</a></p><p>2. WIOD Socio-Economic Accounts (SEA) data 2016 release (downloaded on 30 May 2023)</p><p>- Source: <a href="https://www.rug.nl/ggdc/valuechain/wiod/wiod-2016-release">https://www.rug.nl/ggdc/valuechain/wiod/wiod-2016-release</a></p><p>3. WIOD World Input-Output Tables (WIOT) data November 2016 (downloaded on 23 June 2023)</p><p>- Source: <a href="https://www.rug.nl/ggdc/valuechain/wiod/wiod-2016-release">https://www.rug.nl/ggdc/valuechain/wiod/wiod-2016-release</a> </p>
Data from: Central Mongolian lake sediments reveal new insights on climate change and equestrian empires in the Eastern Steppes
<p>The data set includes the results of ICP-OES, CNS, biomarker, and stable isotope analyses published in the research paper:</p> <p><strong>Struck, J., Bliedtner, M., Strobel, P., Taylor, W., Biskop, S., Plessen, B., Klaes, B., Bittner, L., Jamsranjav, B., Salazar, G., Szidat, S., Brenning, A., Bazarradnaa, E., Glaser, B., Zech, M., Zech, R.: Central Mongolian lake sediments reveal new insights on climate change and equestrian empires in the Eastern Steppes. Scientific Reports, 12, 2829, (2022). DOI: https://doi.org/10.1038/s41598-022-06659-w</strong></p> <p>For further information, in particular, the analyses and methods applied, we refer the reader/user to the original research paper and the supporting information published in Scientific Reports.</p> <p> </p> <p> </p>
Data for empirical example in: An effect size for comparing the strength of morphological integration across studies
<p>Understanding how and why phenotypic traits covary is a major interest in evolutionary biology. Biologists have long sought to characterize the extent of morphological integration in organisms, but comparing levels of integration for a set of traits across taxa has been hampered by the lack of a reliable summary measure and testing procedure. Here we propose a standardized effect size for this purpose, calculated from the relative eigenvalue variance, Vrel. First we evaluate several eigenvalue dispersion indices under various conditions, and show that only Vrel remains stable across samples size and the number of variables. We then demonstrate that Vrel accurately characterizes input patterns of covariation, so long as redundant dimensions are excluded from the calculations. However, we also show that the variance of the sampling distribution of Vrel depends on input levels of trait covariation, making Vrel unsuitable for direct comparisons. As a solution, we propose transforming Vrel to a standardized effect size (Z-score) for representing the magnitude of integration for a set of traits. We also propose a two-sample test for comparing the strength of integration between taxa, and show that this test displays appropriate statistical properties. We provide software for implementing the procedure, and an empirical example illustrates its use.</p>
Supplementary material and supplementary data files for: Handling logical character dependency in phylogenetic inference: Extensive performance testing of assumptions and solutions using simulated and empirical data
<p>Logical character dependency is a major conceptual and methodological problem in phylogenetic inference of morphological datasets, as it violates the assumption of character independence that is common to all phylogenetic methods. It is more frequently observed in higher-level phylogenies or in datasets characterizing major evolutionary transitions, as these represent parts of the tree of life where (primary) anatomical characters either originate or disappear entirely. As a result, secondary traits related to these primary characters become "inapplicable" across all sampled taxa in which that character is absent. Various solutions have been explored over the last three decades to handle character dependency, such as alternative character coding schemes and, more recently, new algorithmic implementations. However, the accuracy of the proposed solutions, or the impact of character dependency across distinct optimality criteria, has never been directly tested using standard performance measures. Here, we utilize simple and complex simulated morphological datasets analyzed under different maximum parsimony optimization procedures and Bayesian inference to test the accuracy of various coding and algorithmic solutions to character dependency. This is complemented by empirical analyses using a recoded dataset on palaeognathid birds. We find that in small, simulated datasets, absent coding performs better than other popular coding strategies available (contingent and multistate), whereas in more complex simulations (larger datasets controlled for different tree structure and character distribution models) contingent coding is favored more frequently. Under contingent coding, a recently proposed weighting algorithm produces the most accurate results for maximum parsimony. However, Bayesian inference outperforms all parsimony-based solutions to handle character dependency due to fundamental differences in their optimization procedures—a simple alternative that has been long overlooked. Yet, we show that the more primary characters bearing secondary (dependent) traits there are in a dataset, the harder it is to estimate the true phylogenetic tree, regardless of the optimality criterion, owing to a considerable expansion of the tree parameter space.</p>
Data from: Can extreme climatic events induce shifts in adaptive potential? A conceptual framework and empirical test with Anolis lizards
<p>Multivariate adaptation to climatic shifts may be limited by trait integration that causes genetic variation to be low in the direction of selection. However, strong episodes of selection induced by extreme climatic pressures may facilitate future population-wide responses if selection reduces trait integration and increases adaptive potential (i.e., evolvability). We explain this counter-intuitive framework for extreme climatic events in which directional selection leads to increased evolvability and exemplify its use in a case study. We tested this hypothesis in two populations of the lizard <em>Anolis scriptus</em> that experienced hurricane-induced selection on limb traits. We surveyed populations immediately before and after the hurricane as well as the offspring of post-hurricane survivors, allowing us to estimate both selection and response to selection on key functional traits: forelimb length, hindlimb length, and toepad area. Direct selection was parallel in both islands and strong in several limb traits. Even though overall limb integration did not change after the hurricane, both populations showed a non-significant tendency toward increased evolvability after the hurricane despite the direction of selection not being aligned with the axis of most variance (i.e., body size). The population with comparably lower between-limb integration showed a less constrained response to selection. Hurricane-induced selection, not aligned with the pattern of high trait correlations, likely conflicts with selection occurring during normal ecological conditions that favor functional coordination between limb traits, and would likely need to be very strong and more persistent to elicit a greater change in trait integration and evolvability. Future tests of this hypothesis should use G-matrices in a variety of wild organisms experiencing selection due to extreme climatic events. </p>
Game Data Event Log from Age of Empire Interactions
<p><span>The event log describes players' behavior in the real-time strategy game Age of Empires. Each case describes the events that a player triggers in a game. </span><span>There are 185.094 cases that consist</span><span> of more than 18 million events. The timestamp represents the elapsed time since the start of the game.</span></p> <p><span>Each player is assigned an</span><span> Elo ranking that is higher, the better the player is. This allows us to study the implications of skill on players' behavior. Also, games can take place on different maps, influencing the situations the players find themselves in. Some games follow clear initial strategies, which are called build orders. These build orders are comparable to chess openings.</span></p> <p><span>The event log is split into ten parts to make the import feasible for smaller machines.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.