Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,600
datasets available to search
ShareScore release 0.9.0
Dataset results
1,600 results for “input”
Input data set for the statistical analsysis of rockfall reach probabilities
<p>These files contain reach probability values extracted from 3D rockfall simulations for field-mapped block deposits as well as a series of attributes characterising the deposits. They served for the statistical analysis of the reach probability values as a function of site, forest and rockfall characteristics. The results of the analysis are published in Dorren et al. 2022: Delimiting rockfall runout zones using reach probability values simulated with a Monte-Carlo based 3D trajectory model. Natural Hazards and Earth System Scienses.</p>
Semi-empirical methods SPT inputs for bearing capacity prediction
<p>These datasets presents inputs for bearing capacity prediction methods. These methods are four well-known semi-empirical models for predicting bearing capacity of piles. The data was collected from the works of Lobo (2005), Vianna (2000) and Jr. (1988) and includes 168 load tests and SPT measures taken from severam Brazilian regions. . The file Data.csv is composed only with the numeric values used in each method and the file Data_with_soils.csv includes the soil types for the piles.</p> <p>The suffix 'Dq', 'Mey', 'Av' and 'Tx' represents which method this input was obtained from, corresponding respectively to Decourt and Quaresma (1978), Meyerhof (1983), Aoki and Velloso (1975) and Teixeira (1996).</p> <p>The columns indexes represents:</p> <p>N_pile - Pile number (for reference);<br> SPT_L - SPT result for the pile lenght;<br> SPT_P - SPT result for the pile tip<br> Soil_L - Predominant soil type along the pile lenght;<br> Soil_P - Predominant soil type in the pile tip;<br> L - pile lenght;<br> D - pile diameter;<br> Qu - Pile bearing capacity, obtained through NBR 6122 load test.<br> <br> When using this dataset, please cite the following paper:</p> <p>//<a href="http://soilsandrocks.com/sr-2021-074921">soilsandrocks.com/sr-2021-074921</a></p> <p>DOI: 10.28927/SR.2021.074921</p>
Example dataset input for IgIDivA
<p>Example dataset input for the Immunoglobulin Intraclonal Diversification Analysis (IgIDivA) tool. (Publication of IgIDivA under revision)</p> <p>The data was retrieved from ENA (https://www.ebi.ac.uk/ena/browser/view/PRJEB36589?show=reads) under the accession number PRJEB36589, and subsequently processed with IMGT/HighV-QUEST (https://www.imgt.org/HighV-QUEST/home.action) and tripr (https://bioconductor.org/packages/release/bioc/html/tripr.html).</p> <p> </p>
MIDAS2 protocol example input dataset v1
<p>Example input dataset for MIDAS2 protocols. </p> <p>Include single-end sequencing reads for two HMP mock community samples: SRR172902 and SRR172903.</p>
RRR/RAPID input and output files corresponding to "Underlying Fundamentals of Kalman Filtering for River Network Modeling"
<p><strong>Corresponding peer-reviewed publication</strong></p> <p>This dataset corresponds to all the RRR/RAPID input and output files that were used in the study reported in:</p> <ul> <li> <p>Emery, C. M., C. H. David, K. M. Andreadis, M. J. Turmon, J. T. Reager, and J. M. Hobbs (2020), Underlying Fundamentals of Kalman Filtering for River Network Modeling, Journal of Hydrometeorology, 21, 453-474, DOI: 10.1175/JHM-D-19-0084.1.</p> </li> </ul> <p>When making use of any of the files in this dataset, please cite both the aforementioned article and the dataset herein. </p> <p><strong>Known bugs and limitations in this dataset or the associated manuscript.</strong></p> <p>In the final version of the published manuscript, Figure 5a, Figure 5b, Figure 5c, and Figure SF1 are inaccurate. The issue in these figures is that they were all prepared with an incorrect indexing relating observed and simulated discharge, hence observations at any one location were consistently being compared to simulations at another different location. As a result, all values of "measured" discharge errors (i.e. Bias, STDE, and RMSE) are incorrect. This issue did not affect the values of "estimated" errors, nor did it affect all values of the Nash-Sutcliffe effeciency that are presented. The figures published in the manuscript can all be recreated using the files in which "BUG_DO_NOT_USE" was appended to the name. Correct figures can also be created using corresponding file names that were not so appended. </p> <p>Note that corrected versions of Figure 5a, Figure 5b, Figure 5c, and Figure SF1 all retain the same strong linear relationships that are discussed in the paper. The slope of the daily discharge STDE trend initially reported as <span class="math-tex">\(\alpha = 0.3876\)</span> in Figure 5c changes to <span class="math-tex">\(\alpha = 0.4507\)</span> after correction. The resulting value of the ideal inflation factor hence changes from <span class="math-tex">\(I = {1 \over 0.3876} \approx 2.58\)</span> to <span class="math-tex">\(I = {1 \over 0.4507} \approx 2.22\)</span>. This updated ideal inflation factor has no impact on the conclusions reached in the manuscript because it remains closer to <span class="math-tex">\(I = 2.58\)</span> than to <span class="math-tex">\(I = 1\)</span> or <span class="math-tex">\(I = 5\)</span>, <em>i.e.</em> the three values that were evaluated.</p> <p>Additionally, a faulty version 1.3.1 of the Python toolbox netCDF4 led to incorrect interpretation of _FillValue in which every data point of value greater than _FillValue was interpreted as masked. This created discrepancies in the following three files, which were updated between V1 and V2 of this dataset: "timeseries_rap_exp01.csv", "timeseries_rap_exp18.csv", and "stats_rap_exp18.csv". Faulty versions of the same files have "BUG_NETCDF4" appended to their names. Correct files have been recreated with file names that were not so appended. </p>
Southern African Power Pool GridPath Model Input Data
<p>This data repository holds GridPath model input data for the paper Chowdhury, A.K., Deshmukh, R., Wu, G., Uppal, A., Mileva, A., Curry, T., Armstrong, L., Galelli, S., and Kudakwashe, N. (2022) “Enabling a low-carbon electricity system for Southern Africa”, Joule. See Readme for more details. </p>
HANZE v2.0 exposure model input data
<p>This dataset provides all input data needed to run HANZE v2.0 model. The two ZIP files need to be downloaded and unpacked in the same directory, which has to be defined in "get_file.py" (variable "main_path" at the beginning of the file). For detailed description of the files, see the documentation provided with the code.</p>
Input features of E. coli proteome for predicting and modeling protein-protein interactions with AF2Complex
<p>Input features to be used with AF2Complex for predicting protein-protein interactions among ~4400 E. coli proteins. A pickled feature file was generated by the feature data pipeline of AF2Complex for each E. coli protein. To reduce storage size, we limited up to 10,000 MSA sequences and up to 10 structural templates from the Protein Data Bank. The cutoff date for sequence libraries and the Protein Data Bank releases used for feature generation is no later than 11-30-2021.</p> <ul> <li>ecoli_af2c_fea.txt -- A list of all E coli protein with pre-generated input features</li> <li>af2c_fea_ecoli_220331_msa10ktem10.tar -- Input features named after the UniProt ID of each proteins. Note that after untar the tarball, you may use the gzipped feature pickle files directly with AF2Complex w/o gunzip.</li> </ul>
Inputs of the Jupyter Notebook - Met Office UKV high-resolution atmosphere model data
<p>The dataset contains the inputs of the notebook "Met Office UKV high-resolution atmosphere model data" published in The Environmental Data Science Book.</p> <p>The input data refer to a subset of single sample data file for 1.5 m temperature as part of the Met Office contribution to the COVID 19 modelling effort.</p> <p>The full dataset was available for download from the Met Office Azure (https://metdatasa.blob.core.windows.net/covid19-response-non-commercial/). The full dataset was available for download under the terms of non-commercial purposes.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Samantha V. Adams (author), Met Office Informatics Lab, <a href="https://github.com/svadams">@svadams</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <ul> <li> <p>Met Office Informatics Lab (creator)</p> </li> <li> <p>Microsoft (support)</p> </li> <li> <p>European Regional Development Fund (support)</p> </li> </ul> <p><em>Dataset authors</em></p> <ul> <li> <p>Met Office</p> </li> </ul> <p><em>Dataset documentation</em></p> <ul> <li> <p>Theo McCaie. Met office and partners offer data and compute platform for covid-19 researchers. URL: <a href="https://medium.com/informatics-lab/met-office-and-partners-offer-data-and-compute-platform-for-covid-19-researchers-83848ac55f5f">https://medium.com/informatics-lab/met-office-and-partners-offer-data-and-compute-platform-for-covid-19-researchers-83848ac55f5f</a>.</p> </li> </ul> <p><strong>Note this data should be used only for non-commercial purposes.</strong></p>
Example input files and output data for 1D hydrodynamic simulations of shock compressed iron
<p>Example input files and output data for 1D hydrodynamic simulations of shock compressed iron. Input files consists of 3 examples from the SIMEX github wiki page for a 50 micron CH ablator with 5 micro Fe foil (laser pulse is a 6 ns flat top pulse, 1064 nm with 0.3 TW/cm<sup>2</sup>). Output data are from Esther hydrocode in .txt format and the SIMEX opmd.h5 format.</p>
HipFT Sample Input Dataset for Convective Flows and Data Assimilation
<p>This file package is a sample data set for running <a href="https://www.github.com/predsci/hipft">HipFT</a> with convective flows and data assimilation. </p> <p>The convective flows were generated with the <a href="https://www.github.com/predsci/conflow">ConFlow</a> code (soon to be released), while the data assimilation maps were processed from HMI M720s LOS data using the <a href="https://www.github.com/predsci/MagMAP">MagMAP</a> package (also soon to be released).</p> <p>See the enclosed README file on how to run an example included in the HipFT package that uses the two data sets.</p>
Binding Affinity Prediction Workflow - Simulation Input Files and Absolute Binding Free Energies
<p>The Binding Affinity Prediction (BAP) workflow calculates absolute binding free energies for protein-ligand complexes by taking their crystal structures, converting them into input files for molecular dynamics (MD) simulations with GROMACS after they have passed extensive quality checks, and analysing the resulting trajectories with the Generalised Born model of implicit solvation as implemented in gmx_MMPBSA to obtain the free-energy estimates. The workflow was designed for soluble proteins without post-translational modifications, co-factors and non-standard amino acids, and it has limited support for coordinated ions.</p> <p>For the dataset published here, the BAP workflow was run on the PDBbind 2020 (http://www.pdbbind.org.cn/index.php) refined set. This entry contains the MD simulation input files (BAPSimulationInputFiles.tar.gz) and the ABFE estimates (BAPBindingFreeEnergyEstimates.csv) obtained from four 250 ns trajectories for each complex. The MD simulations for more than 4000 complexes were run on the Leonardo supercomputer while the implicit-solvent calculations were carried out on Galileo, both operated by Cineca (Italy). The MD trajectories will be stored at Cineca for approx. 1 year after publication of this entry; contact Cineca's user support if you are interested in the trajectories.</p> <p>The README file describes how to reproduce the MD trajectories and the subsequent implicit-solvent calculations yielding the free-energy estimates. The workflow scripts can be downloaded from GitHub (https://github.com/LigateProject/Binding-Affinity-Prediction-workflow). The MD simulations were run with GROMACS 2023.2 (https://manual.gromacs.org/2023.2/index.html), and the implicit-solvent calculations were carried out with gmx_MMPBSA 1.6.1 (https://valdes-tresanco-ms.github.io/gmx_MMPBSA/v1.6.1/).</p>
Input data files for the CPR-DINCAE project
<p>This dataset contains the netCDF files used as the input for [DINCAE](https://github.com/gher-uliege/DINCAE.jl)</p>
Turkiye, Regional Domestic Input Output Tables, 2022, 26 Region, 62 X 62 Matris
<p>Regional data is crucial for conducting inter-regional economic analyses, understanding productivity, wages, and minimum wage dynamics, assessing the impact of unforeseen events like natural disasters, comparing sectors, and evaluating development levels. Supply-use tables, employment, and demographic statistics are the most commonly used data sources for these analyses.</p> <p>To address these needs and for my own economic analyses, I have created regional supply-use tables at the province level. Additionally, I have generated domestic input-output tables using the NUTS 2 regional classification for conducting regional and inter-regional analyses in Turkiye.</p> <h2>Scope of the Research</h2> <p>This research presents domestic input-output tables for 62 industries in 26 NUTS 2 regions of Turkiye for the year 2022.</p> <h2>Methodology</h2> <p>Regional accounts are published at the A10 level for every province in Turkey. In the Input-Output tables, which consist of 62 sectors, the input-output distribution has been conducted on a sector basis to derive the total at the A10 level. The estimations at the provincial level are automatically derived from regional accounts data.</p> <p>The NUTS 2 level regional input-output data have been published and disseminated with an A62 matrix. Additionally, province-level input-output data have been published and disseminated with an A62 matrix through the application. After summations the regional input output tables have been estimated.</p> <h2><strong>Application</strong></h2> <p>The link to the relevant application is attached:</p> <h3><a href="https://turkiye.streamlit.app/" target="_blank" rel="noopener ugc nofollow">https://turkiye.streamlit.app/</a></h3> <p> </p>
Fine-scale anthropogenic nutrient input data for 8 watersheds along the Saint Lawrence River (1981, 2021)
<p>We quantified Net Anthropogenic Nitrogen and Phosphorus Inputs (NANI-NAPI) at two scales (the finest one, the municipality, and a coarser one, the county) for all municipalities and counties of 8 watersheds in Québec, Canada, for 1981 and 2021.</p> <p>The datasets here report 1) watershed NANI and NAPI values accounted from both scales for 1981 and 2021, 2) municipality-scale NANI values for each municipality in the 8 watersheds, and 3) certain components of the municipality-scale NANI for all municipalities in the Yamaska watershed.</p>
Input files for "Faster Simulations with a 5 fs Time Step for Lipids in the CHARMM Force Field"
<p>The performance of all-atom molecular dynamics simulations is limited by an integration time step of 2 fs, which is needed to resolve the fastest degrees of freedom in the system, namely, the vibration of bonds and angles involving hydrogen atoms. The virtual interaction sites (VIS) method replaces hydrogen atoms by massless virtual interaction sites to eliminate these degrees of freedom while keeping intact nonbonded interactions and the explicit treatment of hydrogen atoms. We have modified the existing VIS algorithm for most lipids in the popular CHARMM36 force field by increasing the hydrogen atom masses at regular intervals in the lipid acyl chains and obtained lipid properties and pore formation free energies in very good agreement with those calculated in simulations without VIS. Our modified VIS scheme enables a 5 fs time step resulting in a significant performance gain for all-atom simulations of membranes. The method has the potential to make longer time and length scales accessible in all-atom simulations of membrane–protein complexes.</p> <p>The file set contains individual lipid topologies for virtual interaction sites for standard CHARMM lipids, as well as a README file with instructions on how to implement the VIS algorithm for membranes or membrane-protein complexes</p> <p>Please Cite: <a href="//pubs.acs.org/doi/10.1021/acs.jctc.8b00267">10.1021/acs.jctc.8b00267</a></p> <p> </p>
Data set for "State-dependent cell-type-specific membrane potential dynamics and unitary synaptic inputs in awake mice"
<p>Data set for: Pala A, Petersen CCH (2018) State-dependent cell-type-specific membrane potential dynamics and unitary synaptic inputs in awake mice. eLife 7: e35869. DOI: https://doi.org/10.7554/eLife.35869.</p> <p>There are 12 files in this data upload:</p> <p>1. '2018_Pala_eLife.pdf' - this is a pdf version of the online publication: Pala & Petersen (2018).</p> <p>2. 'data.mat' - this is a Matlab data structure, which contains all the data for the publication.</p> <p>3. 'DataViewer.m' - this is a Matlab code for viewing the data.</p> <p>4. 'DataViewer.fig' - this is a Matlab figure file, which is the GUI layout for 'DataViewer.m'.</p> <p>5. 'PalaPetersen_Plot.m' - this is a Matlab code, which plots the figures for Pala & Petersen (2018).</p> <p>6. 'PalaPetersen_Analysis.m' - this is a Matlab code, which analyses the data for the figures of Pala & Petersen (2018).</p> <p>7. 'blankAPs.m' - this is a Matlab code, which blanks action potentials from the membrane potential trace.</p> <p>8. 'lowpassfilt.m' - this is a Matlab code, which low pass filters the LFP.</p> <p>9. 'medianFiltAPs.m' - this is a Matlab code, which median filters the membrane potential trace to remove action potentials.</p> <p>10. 'remTrialswithAPs.m' - this is a Matlab code, which removes trials with action potentials.</p> <p>11. 'retrieveSegDur.m' - this is a Matlab code, which retrieves chunks of the recording of a given length.</p> <p>12. 'suptitleAP.m' - this is a Matlab code, which puts titles above subplots.</p>
Collection of Schistosoma mansoni ChIP-Seq input fastq files
<p>These are fastq files of ChIP-Seq input files for different life cycle stages of <em>Schistosoma mansoni</em>.</p> <ul> <li>adult female worms</li> <li>pairs of adults</li> <li>female cercariae</li> <li>miracidia</li> <li>primary sporocysts (sp1)</li> </ul> <p>Produced at IHPE (http://ihpe.univ-perp.fr/)</p>
Dataset: Approximating input data to a snowmelt model using Weather Research and Forecasting model outputs in lieu of meteorological measurements
<p>The dataset presented is the companion data to the Journal of Hydrometeorology publication entitled “Approximating input data to a snowmelt model using Weather Research and Forecasting model outputs in lieu of meteorological measurements.” The data that follows contains everything needed to reproduce the spatial inputs for the meteorological station model run using the Spatial Modeling for Resources Framework (SMRF, Havens et al., 2017).</p> <p> </p> <p>Software versions used:</p> <ul> <li>Image Processing Workbench v2.2.0 (Marks et al., 2017)</li> <li>Spatial Modeling for Resources Framework v0.5.3 (Havens et al., 2019)</li> </ul> <p> </p> <p><strong>NOTE:</strong> Reproducing the spatial inputs will generate 10 netCDF files at ~80GB per file.</p> <p> </p> <p><strong>topo.nc</strong> – Contains multiple static layers that are required to run SMRF and iSnobal. The netCDF layers are:</p> <ul> <li>dem – digital elevation model at 100 meter resolution, aggregated from the 10 meter National Elevation Dataset (Archuleta et al., 2017)</li> <li>mask – basin mask for the Boise River Basin</li> <li>veg_height – vegetation height in meters from the National Land Cover Database (Homer et al., 2015)</li> <li>veg_type – vegetation type from the National Land Cover Database</li> <li>veg_tau – vegetation fractional transmissivity derived from the vegetation type</li> <li>veg_k – vegetation emissivity derived from the vegetation type</li> </ul> <p> </p> <p><strong>maxus.nc</strong> – maximum upwind slope netCDF that contains 72 images for all wind directions in 5 degree increments using the algorithm described in Winstral and Marks (2002)</p> <p> </p> <p><strong>Station data:</strong></p> <ul> <li>Contains hourly meteorological station data downloaded from Mesowest (Horel et al., 2002). Data was cleaned and filtered prior to running SMRF.</li> <li>metadata.csv – metadata for 40 stations</li> <li>air_temp.csv – 38 stations</li> <li>cloud_factor.csv – 7 stations</li> <li>precip.csv – 21 stations</li> <li>vapor_pressure.csv – 19 stations</li> <li>wind_direction.csv – 14 stations</li> <li>wind_speed.csv – 14 stations</li> </ul> <p> </p> <p><strong>smrf_config.ini</strong> – Configuration file needed to reproduce the spatial inputs using SMRF. The paths will need to be changed to reflect the data location.</p>
Vortex input files - Carroll et al. Scientific Reports
<p>Vortex input files associated with the manuscript by Carroll et al. "Biological and sociopolitical sources of uncertainty in population viability analysis for endangered species recovery planning", published in Scientific Reports http://doi.org/10.1038/s41598-019-45032-2</p> <p>v2013nodd.xml: Scenario without density-dependent reproduction adapted from 2013 Mexican wolf PVA, as presented initially in Carroll et al. 2014. https://doi.org/10.1111/cobi.12156<br> v2013dd.xml: Scenario with density-dependent reproduction adapted from 2013 Mexican wolf PVA.<br> v2017.xml: Scenario adapted from 2017 Mexican wolf PVA (Miller 2017). <br> MXWlivNoMex.txt: Pedigree input file for above scenarios.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.