Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
477
datasets available to search
ShareScore release 0.9.0
Dataset results
477 results for “input data”
Input data of the mixed meshes used in: J. Grošelj, M. Kapl, M. Knez, T. Takacs, V. Vitrih, C1-smooth isogeometric spline functions of general degree over planar mixed meshes: The case of two quadratic mesh elements, Applied Mathematics and Computation 460 (2024) 128278; DOI: 10.1016/j.amc.2023.128278
Open the record for dataset details and reuse information.
Demo of Impatto: A Static Analyzer for Quantitative Input Data Usage
<p><strong>Impatto</strong> is a sound fully-automatic and always-terminating static analysis tool based on the quantitative framework for input data usage properties proposed by Mazzucato (https://hal.science/hal-04339001).<strong>Impatto</strong> leverages an underlying backward analyzer to compute the set of input-output relations of the program under analysis. This backward analyzer is a parameter of the tool, allowing different kind of analyses such as program or neural network analysis. Furthermore, the choice of the impact definition is also a parameter of the tool to better suit several factors, such as the program structure, the environment, and the intuition of the researcher.</p> <p>GitHub repository at https://github.com/denismazzucato/impatto</p>
Luminoso Input Data for SemEval-2018 Task 10: "Capturing Discriminative Attributes"
<p>This is the data required to run Luminoso's entry to the SemEval-2018 task on Capturing Discriminative Attributes.</p> <p>This data includes:</p> <ul> <li>A recently-computed version of the <a href="https://github.com/commonsense/conceptnet-numberbatch">ConceptNet Numberbatch</a> word embeddings</li> <li>The output of an implementation of Semantic Matching Energy over ConceptNet</li> <li>A SQLite database containing the lead section of all articles on the <a href="http://en.wikipedia.org">English Wikipedia</a> on 2017-12-20</li> <li>The text file that that database is constructed from</li> <li>A SQLite database of words that co-occur in <a href="http://storage.googleapis.com/books/ngrams/books/datasetsv2.html">Google Books 2-grams</a></li> <li>The text file containing total counts of 2-grams in the Google Books data, which that database is constructed from</li> </ul> <p>For more information, see the paper "Luminoso at SemEval-2018 Task 10: Distinguishing Attributes Using Text Corpora and Relational Knowledge", by Robyn Speer and Joanna Lowry-Duda, to appear in the proceedings of the SemEval workshop at NAACL 2018.</p>
Input data from the paper: Small room for compromise between oil palm cultivation and primate conservation in Africa
<p>All the raw input data needed to replicate the analyses from the paper:</p> <p><strong>Strona G., S. D. Stringer, G. Vieilledent, Z. Szantoi, J. Garcia-Ulloa, S. Wich.</strong> Small room for compromise between oil palm cultivation and primate conservation in Africa.</p>
ConceptNet Vector Ensemble 16.04 input data
<p>This is the data required to build the paper "An Ensemble Method to Build High-Quality Word Embeddings", by Robyn Speer and Joshua Chin.</p> <p>The input data itself comes from:</p> <ul> <li> <p><a href="http://conceptnet5.media.mit.edu/">ConceptNet 5.4</a>, which contains data from Wiktionary, WordNet, and many contributors to Open Mind Common Sense projects, edited by Robyn Speer</p> </li> <li> <p><a href="http://nlp.stanford.edu/projects/glove/">GloVe</a>, by Jeffrey Pennington, Richard Socher, and Christopher Manning</p> </li> <li> <p><a href="https://code.google.com/archive/p/word2vec/">word2vec</a>, by Tomas Mikolov and Google Research</p> </li> <li> <p><a href="http://www.cis.upenn.edu/~ccb/ppdb/">PPDB</a>, by Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch</p> </li> </ul>
Input Data for paper "Energy Storage Profit Risk under Stochastic Fuel Prices"
<p>This is a supplementary information accompanying "Energy Storage Profit Risk under Stochastic Fuel Prices" paper submitted to <a href="https://www.journals.elsevier.com/energy-economics/">Energy Economics</a>.</p>
Input Data for "Distinguishing attributes using ConceptNet Numberbatch"
<p>In a post on blog.conceptnet.io, we're showing how to use ConceptNet Numberbatch alone to create a good solution to SemEval-2018 Task 10, Capturing Discriminative Attributes. This is an alternative, simplified implementation of a result presented in the SemEval paper <a href="http://aclweb.org/anthology/S18-1162">Distinguishing Attributes Using Text Corpora and Relational Knowledge</a>, by Robyn Speer and Joanna Lowry-Duda.</p> <p>This data repository contains the data necessary to make the simplified implementation work.</p>
Heatmap input data
<p>Dataset can be visualized by using the destair_heatmap.R script (https://github.com/destairdenbi/tools/tree/master/destair_heatmap).</p>
Input files and data for path generation of alanine dipeptide isomerization in virtual reality
<p>The input files and resulting data for the accelerated sampling of the isomerization of alanine dipeptide used in the thesis:</p> <p>"Accelerated Sampling Methods for High Dimensional Molecular Systems", Mike O'Connor, University of Bristol. </p> <p> </p> <p> </p>
Processed input data for vampire-analysis-1
<p>Input files for the analysis and plotting code for the paper "Deep generative models for T cell receptor protein sequences." </p>
Viet Nam Technology Catalogue - Technology data input for power system modelling in Viet Nam
<p>Today, innovations and technology improvements within renewable energy are taking place at a very rapid pace. Long-term energy planning is very dependent on cost and performance of future energy producing technologies.<br> This technology catalogue provides estimates of costs and performance for a wide range of power producing technologies, thereby building one of the key inputs to good energy planning in Vietnam.<br> Due to the multi-stakeholder involvement in the data collection process, the technology catalogue contains data that have been scrutinised and discussed by a broad range of relevant stakeholders including the Ministry of Industry and Trade – MOIT, Vietnam Electricity – EVN, independent power producers, local and international consultants, organizations, associations and universities. This is essential because a main objective is to produce a technology catalogue which is well anchored amongst all stakeholders.<br> The technology catalogue will assist the long-term energy modelling in Vietnam and support government institutions, private energy companies, think tanks and others with a common and broadly recognized set of data for electricity producing technologies in Vietnam in the future.</p>
Input data for soil erosion practical
<p>This dataset is to be used for the soil erosion practical published at: https://github.com/wieka29/Soil-erosion-practical</p> <p> </p>
Numerical model code, input files and output data for publication "Rapid mixing and exchange of deep-ocean waters in an abyssal boundary current"
<p>Contains numerical model data (code, input files, selected output, matlab diagnostic routines) to supplement publication ``Rapid mixing and exchange of deep-ocean waters in an abyssal boundary current'', by Naveiro Garabato and co-authors. All numerical model data, including any errors, is the responsibility of Sonya Legg. This data set will allow reproduction of simulations, and reproduction of diagnostics shown in plots in the above-referenced paper.</p>
Developing Implementable Climatic Input Data and Moisture Boundary Conditions for Pavement Analysis and Design
<p>Corresponding data set for Tran-SET Project No. 18POKS03. Abstract of the final report is stated below for reference:</p> <p>"The main objective of this study is to develop a practical and implementable numerical model for predicting the moisture (suction) regime within the pavement subgrade system. The research quality and uniformly-dispersed climate data over short distances from Oklahoma Mesonet and the Mitchell based moisture (suction) prediction methods establish the main background of the research study. The study involved numerical modeling and statistical analysis of climatic weather data. The proposed moisture variation model predicts the suction distribution throughout the soil subgrade by solving the diffusion equation and incorporates the measured suction from the Oklahoma Mesonet to estimate the diffusion coefficient. The research study resulted in a practical prediction model that could be used to determine the moisture boundary conditions within the pavement structure."</p>
RNA Pol III input data and output models
<p>This repository contains the input experimental data used in a tutorial on modeling of RNA Polymerase III and the largest cluster of output models.</p>
Input data for Coalispr
<p>Datasets to illustrate the use of <a href="https://coalispr.codeberg.page/README.html" target="_blank" rel="noopener">Coalispr</a>, a Python tool to clean up (small) RNA sequencing results. Coalispr can visualize over 100 bedgraphs in one panel and helps to retrieve read counts from associated alignment files without reliance on reference features like GTF annotations. The archives contain bedgraphs (also in processed form), reference data, bam-alignment files for counting, and descriptions for experiments, ncRNAs and genes.</p>
Data from: Riverine transport and nutrient inputs affect phytoplankton communities in a coastal embayment
<p>1. Rivers often transport phytoplankton to coastal embayments and introduce nutrients that can enrich coastal plankton communities. We investigated the effects of the Nottawasaga River on the nearshore (i.e., within 500 m of shore) phytoplankton composition along a 10 km transect of Nottawasaga Bay, Lake Huron in 2015 and 2016. Imaging flow cytometry was used to identify and enumerate algal taxa, which were resolved at sizes larger than small nanoplankton (i.e., > 5 mm). Multivariate analysis (perMANOVA and RDA) and a dilution model were used to examine how nutrients and the transport of algal taxa affected community composition in the bay.</p> <p>2. Sampling stations with different percentages of river water had significantly different phytoplankton communities. Phytoplankton community composition was also strongly associated with nutrients, including total phosphorus, which also varied with the percentage of river water. The majority of the 51 phytoplankton taxa identified in 2016 had numerical abundances in the bay that could be explained simply by the dilution of incoming river water.</p> <p>3. Phytoplankton transported from the river had a higher proportion of "edible-sized" cells (< 30 mm), particularly in summer when colonial cyanobacteria were numerically dominant in the bay. Six taxa were more abundant than expected from the dilution of river water and included some cyanobacteria with late summer maxima. Five of the taxa that were transported from the river were less abundant than expected in the bay.</p> <p>4. Whereas impacts of fertilization due to the characteristically higher nutrient concentration in the river are to be expected, the strong and highly correlated effects of transport within the narrow coastal band of this study largely concealed any distinct fertilization effects.</p> <p>5. Riverine inputs may strongly influence the near-shore assemblage of phytoplankton in oligotrophic embayments in large lakes, creating hotspots for productivity, species turnover and trophic dynamics.</p>
Input Data for "Assembly and Analysis of Cell-Scale Membrane Envelopes"
<p>Input structures for a manuscript, along with selected output data and structures. This directory structure contains a cut-down copy of the directories used to generate the simulation data and the analysis. In order to make this fit into the 50GB Zenodo limit, it was constructed with the following tar command: `tar -zcvf protocellmodeling.tar.gz --exclude="*BAK" --exclude="*#" --exclude="*xtc" --exclude="*gro" --exclude="*trr" --exclude="*js" --exclude="*[0-9].out" --exclude="*old" --exclude="*dcd" --exclude="*tmp" --exclude="*xst" --exclude="*edr" --exclude="*state_prev.cpt" --exclude="*.o[0-9]*" cgDracula`, which intentionally excludes large files. The full 4.8TB dataset that includes trajectories is available upon request.</p> <p>The data is split into multiple subdirectories and largely undocumented, however here are the highlights:</p> <ul> <li>The <strong>Analysis</strong> subdirectory is where the analysis in the paper lives. All other directories are related to building or running systems.</li> <li><strong>getsources.py</strong> in the main directory is the script that downloads the initial structure from MemProtMD.</li> <li><strong>transform.py</strong> builds the initial protein models from MemProtMD.</li> <li><strong>vesiclebuilder.py</strong> builds the lipid ball.</li> <li><strong>protpatchplacer.py</strong> sets up the ultra-coarse grained simulation, which is in the <strong>supercg</strong> directory.</li> <li><strong>movepatches.py</strong> takes the results from the ultra-coarse grained simulation, and builds the protein ball.</li> <li><strong>gendx.tcl</strong> generates the density maps from the protein ball.</li> <li>This is used in <strong>lipids/picklipids.py</strong>, which cuts out the pieces of the lipid that need to be removed.</li> <li>The water is added to the system with <strong>addwater/quicksolvate.sh</strong></li> <li>The system is ionized by <strong>ionize.py</strong></li> <li>And a topology is written by <strong>writetop.py</strong></li> </ul>
The corticospinal tract primarily modulates sensory inputs in the mouse lumbar cord (Raw data)
<p>Raw data of the eLife 2021 article</p> <p><strong>The corticospinal tract primarily modulates sensory inputs in the mouse lumbar cord.</strong></p> <p>Authors:</p> <p><strong>Yunuen Moreno-Lopez<sup>1</sup>, Charlotte Bichara<sup>1</sup>, Gilles Delbecq, Philippe Isope, Matilde Cordero-Erausquin</strong></p> <p>Methods are described in the article. Data is organized by figure.</p>
Input Data for "Protein Function Prediction for newly sequenced organisms"
<p>The input sequence files in FASTA format and the detailed list of all organisms excluded when testing each specific bacterium.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.