Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Raw motif mapping bedfile data and model training set class probabilities
<p>Leveraging prior viral genome sequencing data to make predictions on whether an unknown, emergent virus harbors a 'phenotype-of-concern' has been a long-sought goal of genomic epidemiology. A predictive phenotype model built from nucleotide-level information alone is challenging with respect to RNA viruses due to the ultra-high intra-sequence variance of their genomes, even within closely related clades. We developed a degenerate k-mer method to accommodate this high intra-sequence variation of RNA virus genomes for modeling frameworks. By leveraging a taxonomy-guided 'group-shuffle-split' cross validation paradigm on complete coronavirus assemblies from prior to October 2018, we trained multiple regularized logistic regression classifiers at the nucleotide k-mer level. We demonstrate the feasibility of this method by finding models accurately predicting withheld SARS-CoV-2 genome sequences as human pathogens and accurately predicting withheld Swine Acute Diarrhea Syndrome coronavirus (SADS-CoV) genome sequences as non-human pathogens. Feature selection using L1 regularization identified several degenerate nucleotide predictor motifs with high model coefficients for the human pathogen class that were present across widely disparate clades of coronaviruses. However, these motifs differed in which genes they were present in, what specific codons were used to encode them, and what the translated amino acid motif was. This emphasizes the importance of a phenetic view of emerging pathogenic RNA viruses, as opposed to the canonical phylogenetic interpretations most commonly used to track and manage viral zoonoses. Applying our model to more recent Orthocoronavirinae genomes deposited since October 2018 yields a novel contextual view of pathogen potential across bat-related, canine-related, porcine-related, and rodent-related coronaviruses and critical adaptations which may have contributed to the emergence of the pandemic SARS-CoV-2 virus. Finally, we discuss the next steps to achieve robust predictive ensembles and the utility of these models (and their associated predictor motifs) to novel biosurveillance protocols that substantially increase the 'pound-for-pound' information content of field-collected sequencing data and make a strong argument for the necessity of routine collection and sequencing of zoonotic viruses. </p>
Data repository in support of the article: Implementation and evaluation of updated photolysis rates in the EMEP MSC-W chemical transport model using Cloud-J v7.3e
<p>This dataset contains the measurement data, model outputs and Python (v3.10) scripts that are used to produce figures and tables in the paper: Implementation and evaluation of updated photolysis rates in the EMEP MSC-W chemical transport model using Cloud-J v7.3e.</p> <p>The newly created modules providing the interface with Cloud-J in the EMEP MSC-W and BoxChem models are called CloudJ_mod.f90.</p> <p>The measurement data and supporting MATLAB scripts used to create the ATom-1 data files read in by the Python scripts provided here, can be downloaded from https://doi.org/10.3334/ORNLDAAC/1651</p> <p> </p>
Data Repository for "Integrating Water Quality Data with a Bayesian Network Model to Improve Spatial and Temporal Phosphorus Attribution: Application to the Maumee River Basin"
<p>Data for "Integrating Water Quality Data with a Bayesian Network Model to Improve Spatial and Temporal Phosphorus Attribution: Application to the Maumee River Basin". This repository contains all the processed data used in the simulation (in "processed" folder), part of the raw data (in "raw" folder), and the SWAT simulation results (in "SWAT" folder). The code for processing the raw data, which are either provided here or publicly available online, is provided in the <a href="https://doi.org/10.5281/zenodo.8132662">code repository</a>. The links to the publicly available raw data are also provided in the code repository.</p>
Gridded forecast data over the Mediterranean and the North Sea from Met Office Models
<p>This repository contains gridded forcast data from Met Office models from the last three months of 2018 over the Mediterranean and the North Sea. Foe further information, please consult the README file and the headers within each individual file.</p> <p> </p> <p>The data are provided under the terms of the Non-Commercial Government Licence,<br> <a href="https://www.nationalarchives.gov.uk/doc/non-commercial-government-licence/version/2/">Non-commercial Government Licence.</a></p> <p> </p>
animal soup sample data, ground truth dataset, and pre-trained models
<p>- sample data and ground truth files for animal soup tests</p> <p>- pre-trained models</p>
Large-eddy simulation model data for article Study of surface layer characteristics in the presence of suspended snow particles using observational data and large-eddy simulation
<p>Large-eddy simulation model data [8. Mortikov E.V., Glazunov A.V., <br> Lykosov V.N. Numerical study of plane Couette flow: <br> turbulence statistics and the structure of pressure-strain correlations <br> // Russ. J. Numer. Analysis Math. Model. 2019. V. 34. № 2. P. 119–132.]<br> The setup of experiments was based<br> on the GABLS-1. The height, width and length of the domain was 4000 m<br> with spatial resolution of 11.7m. U18, U16 - geostrophic wind 18 and 16 m/s,<br> CR - cooling rate 0K/h, 1K/h, 2K/h.<br> Two series of experiments were performed: “NS“ and “SS“.<br> In the experiment “NS“ the surface layer was described according to the<br> Monin-Obukhov similarity theory. The experiments “SS“<br> utilized the parameterization, which takes into account the effect<br> snow particles.</p>
Data for "No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models"
<p>Datasets to reproduce the experiments associated with the paper: https://doi.org/10.48550/arXiv.2307.06440</p> <p>The readme contains instructions for how to use them: https://github.com/JeanKaddour/NoTrainNoGain/blob/main/bert/README.md</p> <p>c4-subset-random.tar.bz2 is a subset of the C4 dataset (https://arxiv.org/abs/1910.10683), licensed under ODC-BY 1.0.</p>
Model and Experimental Data
<p>修正岩石断裂面接触模型的数值和实验室数据</p>
Data from: Economic uncertainty, geopolitical risk and U.S. energy price risk spillover: An empirical study based on the risk spillover model
<p>This data is mainly used to analyze the risk correlation among economic uncertainly,geopolitical risk and energy price,and can also be applied to the TVP-VAR model to analyze the correlation between different variables using the time-varying parameter model,which has great potential for reuse. The risk relationship between economic uncertainly.At the same time,since it is macroeconomic data,it does not involve any moral and ethical issues.</p>
SRDTrans model for simulated SMLM data
<p>SRDTrans model for simulated SMLM data</p>
Data for "Validating Ionospheric Models Against Technologically Relevant Metrics"
<p>Data for "Validating Ionospheric Models Against Technologically Relevant Metrics"</p>
An Empirical Model for Mercury's Field-Aligned Currents Derived from MESSENGER Magnetometer Data
<p>Companion dataset for manuscript titled "An Empirical Model for Mercury’s Field-Aligned Currents Derived from MESSENGER Magnetometer Data" submitted to <em>JGR: Space Physics</em>.</p> <p>The file MESSENGER_MAG_Disturbance_Indices_v12.txt defines the Mercury disturbance index as a function of MESSENGER orbit number.</p> <p>The file rsuns.txt defines the Mercury’s Heliocentric distance as a function of time.</p> <p>The file merge_magrdr_11083_15120_actidx_kt17res.txt contain the processed MESSENGER magnetometer vectors. This file is further broken into quintiles based on the disturbance index.</p>
A case study: assessing the efficacy of the revised dosage regimen via prediction model for recurrent event rate using biomarker data
Open the record for dataset details and reuse information.
Supporting data to run the inverse model of coupled phosphorus, carbon and oxygen model.
<p>Supporting data to run the inverse model of coupled phosphorus, carbon and oxygen model. </p>
Data supporting "Modeling and evaluating the effects of irrigation on land-atmosphere interaction in southwestern Europe with the regional climate model REMO2020-iMOVE using a newly developed parameterization"
<p>This data supports the analysis of the manuscript Asmus et al. 2023 "Modeling and evaluating the effects of irrigation on land-atmosphere interaction in southwestern Europe with the regional climate model REMO2020-iMOVE using a newly developed parameterization".</p><p><strong>Simulation data</strong></p><p>The simulation data is created with REMO2020-iMOVE using the new irrigation parameterization. The results are saved as NetCDF files with monthly mean values and/or time series (hourly) of single variables for the analysis period. A list of the simulations can be found below.</p><p><strong>Observation data </strong></p><p>The observation data is published with the kind permission of ISPRA which hosts the SCIA database (www.scia.isprambiente.it). If you use this data, please make sure to include the following data source:<br>SCIA by ISPRA - Area Climatologia operativa - Via V. Brancati 48 00144 Roma. <br>We downloaded monthly mean values for the variables T2Max, T2Min and T2Mean from <a href="http://193.206.192.214/servertsutm/serietemporali100.php">http://193.206.192.214/servertsutm/serietemporali100.php </a>(last accessed on 14/10/2022) to verify the model results with and without irrigation parameterization. </p><p>By untarring the tarballs, the data structure is created that is necessary to execute the analysis scripts.</p><p>For more information or additional data please contact the author.</p><p> </p><p><strong>tarball | exp_number | description </strong></p><p>067015.tar.gz | 067015 | not irrigated</p><p>067016.tar.gz | 067016 | irrigated with "adaptive water application scheme"</p><p>067017.tar.gz | 067017 | irrigated with "adaptive water application scheme"</p><p>067019.tar.gz | 067019 | irrigated with "flexible time water application scheme"</p><p>067020.tar.gz | 067020 | irrigated with "prescribed water application scheme"</p><p>observation_scia.tar | - | observation data from SCIA </p><p> </p><p><strong>Data structure for simulation data</strong></p><p>\<exp_number><br> \monthly<br> \hourly<br> \var_series<br> \<variable><br><br>Note:<br>\067015 includes static variables<br> \irrifrac (irrigated fraction)<br> \bla (land-sea-mask)</p>
Data for: Impact of islands on tidally dominated river plumes: a high-resolution modelling study
Open the record for dataset details and reuse information.
Data of the Paper: Analytical Modeling and Empirical Validation of Performability of Service- and Cloud-Based Dynamic Routing Architecture Patterns
<p>The online artifacts for the following article accepted at 30th Asia-Pacific Software Engineering Conference (APSEC 2023): </p><p>"Analytical Modeling and Empirical Validation of Performability of Service- and Cloud-Based Dynamic Routing Architecture Patterns"</p><p>Abstract:</p><p>Many dynamic routing architectural patterns are available, including distributed routing, e.g., using the sidecar pattern, or centralized routing, e.g., using event stores or service buses. Different Quality-of-Service (QoS) factors influence routing schemas and technology selection, such as performance, reliability, scalability, and control properties offered by the patterns. An analytical model can formalize the QoS factors and facilitate the architectural decision-making when changing the routing scheme, i.e., to more distributed or centralized routing. So far, the impact of these architectural patterns on performability, i.e., the overall performance of a system with impeded reliability, has not been extensively studied. This is important because deciding to increase performance, e.g., by parallel processing of requests, may lead to decreased reliability because of the added points of a crash. We propose an analytical performability model during component crashes. For the empirical validation of our proposed model, we ran an extensive experiment of 2412 hours of runtime on a private cloud infrastructure and Google Cloud Platform. The low prediction error of 1.75\% indicates the high accuracy of our performability model. These results provide important insights when making architectural decisions regarding service- and cloud-based dynamic routing.</p>
Data for "Supercells and Tornado-like Vortices in an Idealized Global Atmosphere Model"
<p>Data used in a manuscript on supercells and tornado-like vortices, for submission to ESS.</p>
Data and code for, "Large language models design sequence-defined macromolecules via evolutionary optimization"
<div> <pre># Codes and data for "Large language models design sequence-defined macromolecules via evolutionary optimization"<br><br>Note this repository contains codes and data files for the manuscript. This is a snapshot of the repository, frozen at the time of submission.<br><br># Codes<br><br>## LLM codes<br>- `run_claude.py` - the routine for performing LLM-based rollouts; intended for command line execution using argparse<br>- `message_utils.py` - utilities for constructing and parsing messages for LLM I/O<br>- `model_utils.py` - lightweight utilities for retrieving formatted predictions from the RNN ensemble<br>- `target_defs.py` - defines the sequence, locations, and natural language descriptions of the target structures<br>- `ask_about_oracle.ipynb` - asks the LLM to speculate about the nature of the optimization task<br><br>## other algorithms<br>- `active_learning.ipynb` - use EI acquisition with RF surrogate to label new sequences; includes an unused tokenization scheme<br>- `evolutionary_algorithm.ipynb` - use DEAP library to perform evolutionary optimization<br>- `random_sampling.ipynb` - sample sequences randomly from all possible sequences<br><br>## postprocessing<br>- `process_aggregated_logs.py` - reads data from the raw log files and prepares them for visualization<br>- `process_sample_rollouts.py` - reads data from the raw log files and prepares individual rollouts<br><br>## visualization<br>- `figure1b.ipynb` - renders panel b of Fig. 1<br>- `figure1efg.ipynb` - renders the last row of Fig. 1 (panels e-g)<br>- `figure2.ipynb` - renders all of Fig. 2<br>- `figure_si.ipynb` - renders Figs. S1 and S2<br>- `figure_md_validation.ipynb` - renders Fig. S3<br><br># Data files<br><br>- `prompts/`<br> - `prompt-scientific-v4.4.yml` - the full text of the scientific prompt, to be read by `run_claude.py`<br> - `prompt-oracle-v4.4.yml` - the full text of the oracle prompt, to be read by `run_claude.py`<br>- `models/` - the TorchScript RNN models used to make predictions<br>- `data/`<br> - `embeddings` - calculated embeddings for a collection of sequences from our prior work<br> - `llm-logs` - the raw logs obtained from the Claude 3.5 Sonnet LLM (other algorithms made to look like the LLM logs after the fact)<br> - `llm-logs-opus` - the raw logs obtained from the Claude 3.0 Opus LLM (used in the first draft of the article, replaced by Claude 3.5 Sonnet) <br> - `all-rollouts-kltd.csv` - postprocessed logs for all the rollouts using the "top $k < d^*$" metric<br> - `all-rollouts-topkd.csv` - postprocessed logs for all the rollouts using the "mean $d$ for top $k$" metric<br> - `sample-rollout-membranes-x-3.csv` - postprocessed logs for a single rollout replica, `x` = each algorithm type<br> - `snapshots` - png snapshots of MD simulation results at different locations in the manifold</pre> </div>
Vascular Positioning System G4 Algorithm ECG Data Collection for Model Training Study
ClinicalTrials.gov study NCT05702515. IPD Sharing: NO. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.