Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
64
datasets available to search
ShareScore release 0.9.0
Dataset results
64 results for “Assembly Modeling”
A comparison among three ways to assemble wall-to-wall land-cover maps from distribution models of vegetation types
<p>Dataset accompanying manuscript <em>"A comparison among three ways to assemble wall-to-wall land-cover maps from distribution models of vegetation types". </em>Datasets contain a wall-to-wall map of vegetation types covering the study area of terrestrial Norway, produced using three methods for assembling individual predictions from Distribution models (<em>probability-based method</em>, <em>performance-based method</em> and <em>prevalence-based method</em>). </p>
Data for: Highly contiguous genome assembly of Drosophila prolongata – a model for evolution of sexual dimorphism and male-specific innovations
Open the record for dataset details and reuse information.
Data and scripts for "Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"
<p>The README explains how to reproduce the analyses presented in the paper <strong>"Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"</strong> by Abrego et al.</p> <p>The input data for the script pipeline is the file “Kilpisjarvi_plant_data.csv”. This file includes the data on the plants and their traits in the long format. Hence, each row of the data matrix corresponds to measurements on one plant species in one study plot. The joint species-trait distribution modelling (JSTDM) pipeline that analyses these data consists of the following R-scripts.</p> <p>· <strong>S1_define_JSTDM_models.R</strong>. This script defines the JSDTM models (null model and environmental model) that include five response types for each species: the presence-absence, abundance conditional on presence, and the plot-level trait values of specific leaf area (SLA), leaf area (LA) and mean height (MH). The model is defined in the Hierarchical Modelling of Species Communities (HMSC) framework utilizing the R-package Hmsc. The models are saved in the file “unfitted_models.RData”.</p> <p>· <strong>S2_fit_models.R. </strong>This script loads the unfitted models and fits them using the posterior sampling methods implemented in the R-package Hmsc. The models are fitted with increasing thinning until thin=100, which value was used to generate the results of the paper. The fitted models are saved in the file "models_thin_100_samples_250_chains_4.Rdata".</p> <p>· <strong>S3_plot_Omega_matrices.R. </strong>This script loads the fitted models and plots the association matrices (Fig. 2 of the paper). The csv file containing the values used to construct Fig. 2 is also given (figure2Cdata.csv and figure2Ddata.csv).</p> <p>· <strong>S4_show_VP_Beta_Gamma.R. </strong>This script loads the fitted models and extracts information on the variance partitionings (VP; Figs. S3 and S4 of the paper), the relationships between response types and environmental predictors (beta; Fig. S2 of the paper), and the relationships between response types and species-level traits (gamma; Fig. S5 of the paper). The csv file containing the values used to construct Fig. S2 (figureS2Adata.csv and figureS2Bdata.csv), Fig. S3 (figureS3data.csv), Fig. S4 (figureS4data.csv) and Fig S5 (figureS5data.csv) are also given.</p> <p>· <strong>S5_conditional_cross_validation.R</strong>. This script performs 10-fold cross validation to the data to test the predictive power related to the modelled plant traits. The script performs both regular (unconditional) cross-validation where all data are masked for the test fold, and conditional cross-validation where only the trait data (but not the abundance data) are masked for the test fold.</p> <p>· <strong>S6_show_conditional_cross_validation_results.R. </strong>This script plots the results of cross-validation (Fig. 3 of the paper). The csv file containing the values used to construct Fig. 3 is also given (figure3Adata.csv and figure3Bdata.csv).</p> <p>· <strong>S7_scenario_predictions.R</strong>. This script performs the scenario simulations described and shown in Fig. 4 of the paper. The csv file containing the values used to construct Fig. 4 is also given (figure4Bdata.csv and figure4Cdata.csv).</p>
Simulation data for paper "Evaluation of Fendiline Treatment in VP40 System with Nucleation-Elongation Process: A Computational Model of Ebola Virus Matrix Protein Assembly"
<p>This is the original simulation data sets for paper "Evaluation of Fendiline Treatment in VP40 System with Nucleation-Elongation Process: A Computational Model of Ebola Virus Matrix Protein Assembly".</p>
Dataset for the publication titlted "A computational mechanics model for producing molecular assembly using molecularly woven pantographs" in the journal Cell Reports Physical Science, authored by Byeonghwa Goh and Joonmyung Choi.
<p>Dataset for the publication titlted "A computational mechanics model for producing molecular assembly using molecularly woven pantographs" in the journal Cell Reports Physical Science, authored by Byeonghwa Goh and Joonmyung Choi.</p>
Long read genome assembly of Automeris io (Lepidoptera: Saturniidae) an emerging model for the evolution of deimatic displays
<p>Automeris moths are a morphologically diverse group with 145 described species that have a geographic range that spans from the New World temperate zone to the Neotropics. Many Automeris have hindwing eyespots that are thought to deter or disrupt the attack of potential predators, allowing the moth time to escape. Some species in the genus have vestigial eyespots or lack them completely, suggesting that this trait may provide a selective benefit. The Io moth (Automeris io), known for its striking eyespots, is the most widely studied species within the genus and is an emerging model system to study the evolution of deimatism, a predatory defense that combines visual stimuli and movement. Here we present a high-quality, PacBio HiFi genome assembly for Io moth to aid existing research on the molecular development of eyespots. Genomic research is needed to address questions involving antipredatory defenses and eyespot pattern development. BUSCO analysis for this genome shows a completeness of 98.4%, and N50 of 15.</p>
AF2 models for "Co-translational assembly promotes functional diversification of paralogous proteins" by Mallik, et al.
<p>AF2 models used for analyses presented in "Co-translational assembly promotes functional diversification of paralogous proteins" by Saurav Mallik, Angel F. Cisneros, Christian R. Landry, and Emmanuel D. Levy.</p> <p>These models were used to analyze the structural divergence of 3703 Obligatory Homomer, 697 Mixed, and 181 Obligatory Heteromer pairs.</p> <p>Folders are separated into different categories:<br>- Models of homomers (from Schweke et al., 2024. Cell):<br> . AF2_HM_full_models: Full structures of homodimeric models.<br> . AF2_HM_nodiso3: Core structures of homomeric models, trimmed using scripts from Schweke et al. (2023).</p> <p>- Models of heteromers (generated in this work):<br> . AF2_HET_full_models: Full structures of heterodimeric models.<br> . AF2_HET_nodiso3: Core structures of heteromeric models, trimmed using a modified version of the code from Schweke et al. (2023) to work with heterodimers.</p>
Test models and test results to evaluate CAD assembly modules capabilities to generate component interfaces
<p>Set of 3D CAD assembly models in STEP AP 203 format.</p> <p>Assembly test models are devoted to evaluations of interfaces between components. The interfaces can be of type surface, rectilinear contacts, circular contacts, or point contacts.</p> <p>Test results obtained from some commercially available CAD assembly modules are given as a set of tables organized in accordancce with contact categories (surface, rectilinear, circular, point).</p> <p>The content and use of the test models are described into the pdf document: Test models and test results to evaluate CAD assembly modules capabilities to generate component interfaces.</p> <p> </p>
Input datasets for modeling snapshot structures of the NPC post mitotic assembly pathway
<p>This folder contains a reference coarse-grained model of the mature NPC and electron microscopy density maps that are used to restrain snapshot structures in the NPC assembly pathway model. These data sets are used to parameterize the native contact model and restrain the overall shape of NPC pre-pores at each snapshot along the assembly pathway. This data is provided in preparation of the PDB-dev deposition of the assembly pathway, accession to be determined.</p>
PlzR regulates type IV pili assembly in Pseudomonas aeruginosa via PilZ binding - ColabFold models
<p>This dataset contains the ColabFold models from the article 'PlzR regulates type IV pili assembly in <em>Pseudomonas aeruginosa</em> via PilZ binding' by Hendrix, H. <em>et al</em> (Nat. Comm., doi: <span><a href="https://doi.org/10.1038/s41467-024-52732-5" target="_blank" rel="noopener">10.1038/s41467-024-52732-5</a></span>).</p> <p>The dataset consists of two zips (a_PlzR_PilZ_models.zip, b_PlzR_PilZ_PilB_models.zip), which contain the direct output of the different ColabFold runs for the PlzR:PilZ protein complex, and the PlzR:PilZ:PilB protein complex, respectively. </p> <p>All models were predicted using ColabFold's AlphaFold2_mmseqs2 v1.2 notebook. For PlzR:PilZ models, models were generated with relaxation and templates enabled and 3, 6 and 48 recycles (folders PA2560_PilZ_auto_rec3_AMBER_templ_4e025.result, PA2560_PilZ_auto_rec6_AMBER_templ_4e025.result and PA2560_PilZ_auto_rec48_AMBER_templ_4e025.result, respectively). For PlzR:PilZ:PilB models, models were generated with templates enabled and 3 recycles (folder PA2560_PilZ_PilB_auto_rec3_templ_6ea29.result). For PlzR:PilZ:PilB models with only the first 200 amino acids of PilB modeled, models were generated with relaxation and templates enabled and 3 and 6 recycles (folders PA2560_PilZ_PilB200_auto_rec3_AMBER_templ_cd74d.result and PA2560_PilZ_PilB200_auto_rec6_AMBER_templ_cd74d.result, respectively).</p>
Surrogate Model Optimisation of a 'micro core' PWR fuel assembly arrangement using deep learning models - Figures
<p>Figures for Physior 2020 paper</p>
Nontuberculous mycobacteria persistence in a cell model mimicking alveolar macrophages (genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (CDS, rRNA, tRNA, tmRNA, misc_RNA) (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) for the following Nontuberculous mycobacteria (NTM) strains: <em>Mycobacterium smegmatis </em>mc<sup>2</sup>155 (reference strain), <em>Mycobacterium avium</em> ATCC25921 (reference strain), <em>M. avium </em>60/08 (clinical strain), <em>Mycobacterium fortuitum</em> ATCC6841 (reference strain) and <em>M. fortuitum</em> 747/08 (clinical strain).</p> <p>All raw sequence reads used in this study were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB30455).</p> <p>The associated article can be found here: <a href="https://www.ncbi.nlm.nih.gov/pubmed/31035520">https://www.ncbi.nlm.nih.gov/pubmed/31035520</a></p> <p> </p>
Molecular modeling of self-assembling peptides MELD structures
<p>Top 10 MELD structures for each system used in the "Molecular modeling of self-assembling peptides" paper.</p>
Revealing the drivers of parasite community assembly: using avian haemosporidians to model global dynamics of parasite species turnover
<p>Why do some regions share more or fewer species than others? Community assembly relies on the ability of individuals to disperse, colonize, and thrive in new regions. Therefore, many distinct factors, such as geographic distance and environmental features, can determine the odds of a species colonizing a new environment. For parasites, host community composition (i.e., resources) also plays a key role in their ability to colonize a new environment as they rely on their hosts to complete their life cycle. Thus, variation in host community composition and environmental conditions should determine parasite turnover among regions. Here, we explored the global drivers of parasite turnover using avian malaria and malaria-like (haemosporidian) parasites. We compiled global databases on avian haemosporidian lineages distributions, environmental conditions, avian species distributions and functional traits and ran generalized dissimilarity models to uncover the main drivers of parasite turnover. We demonstrated that haemosporidian parasite turnover is mainly driven by geographic distance followed by host functional traits, environmental conditions, and host distributions. The main host functional traits associated with high parasite turnover were the predominance of resident (i.e., non-migratory) species and strong territoriality while the most important climatic drivers of haemosporidian turnover were mean temperature and temperature seasonality. Overall, we establish the importance of geographic distance as a key predictor of ecological dissimilarity and show that host resources influence parasite turnover more strongly than environmental conditions. We also evidenced that parasite turnover is most pronounced among tropical and less interconnected regions (i.e., regions with mostly territorial and non-migratory hosts). Our findings provide a robust foundation for the prediction of avian pathogen spread and the emergence of infectious diseases.</p>
Mathematical model results for: Dynamic fibronectin assembly and remodeling by leader neural crest cells prevents jamming in collective cell migration
<p>Collective cell migration plays an essential role in vertebrate development, yet the extent to which dynamically changing microenvironments influence this phenomenon remains unclear. Observations of the distribution of the extracellular matrix (ECM) component fibronectin during the migration of loosely connected neural crest cells (NCCs) lead us to hypothesize that NCC remodeling of an initially punctate ECM creates a scaffold for trailing cells, enabling them to form robust and coherent stream patterns. We evaluate this idea in a theoretical setting by developing an agent-based model that incorporates reciprocal interactions between NCCs and their ECM. ECM remodeling, haptotaxis, contact guidance, and cell-cell repulsion are sufficient for cells to establish streams in silico, however additional mechanisms, such as chemotaxis, are required to consistently guide cells along the correct target corridor. Further investigations of the model imply that contact guidance and differential cell-cell repulsion between leader and follower cells are key contributors to robust collective cell migration by preventing stream breakage. Global sensitivity analysis and simulated underexpression/overexpression experiments suggest that long-distance migration without jamming is most likely to occur when leading cells specialize in creating ECM fibers, and trailing cells specialize in responding to environmental cues by upregulating mechanisms such as contact guidance. This dataset contains summary statistics, movies, parameter values, and photos obtained from individual realizations of the mathematical model.</p>
A Transformer-based Function Symbol Name Inference Model from an Assembly Language for Binary Reversing
<p>This is a dataset and pre-trained model for the official implementation of <a href="https://github.com/agwaBom/AsmDepictor"><strong>AsmDepictor</strong></a>, "A Transformer-based Function Symbol Name Inference Model from an Assembly Language for Binary Reversing", In the 18th ACM Asia Conference on Computer and Communications Security <a href="https://asiaccs2023.org/">AsiaCCS '2023</a></p> <p> </p>
RNA 3D structure modeling by fragment assembly with Small Angle X-ray Scattering restraints
<p>Structure determination is a key step in the functional characterization of many non-coding RNA molecules. High-resolution RNA 3D structure determination efforts, however, are not keeping up with the pace of discovery of new non-coding RNA sequences. This increases the importance of computational approaches and low-resolution experimental data, such as from the Small Angle X-ray Scattering experiments. We present RNA Masonry, a computer program and a web service for a fully automated modeling of RNA 3D structures. It assemblies RNA fragments into geometrically plausible models that meet user-provided secondary structure constraints, restraints on tertiary contacts and Small Angle X-ray Scattering data. We illustrate the method description with detailed benchmarks and its application to structural studies of viral RNAs with SAXS restraints.</p>
Assembling the anaerobic gamma-butyrobetaine to TMA metabolic pathway in Escherichia fergusonii and confirming its role in TMA production from dietary L-carnitine in murine models
<p>GraphPad Prism files containing source data for figures included in the manuscript "Assembling the anaerobic gamma-butyrobetaine to TMA metabolic pathway in Escherichia fergusonii and confirming its role in TMA production from dietary L-carnitine in murine models", by Dwidar et al., published in mBio.</p>
Scale-dependence of ecological assembly rules: insights from empirical datasets and joint species distribution modelling
Open the record for dataset details and reuse information.
Mathematical model results for: Dynamic fibronectin assembly and remodeling by leader neural crest cells prevents jamming in collective cell migration
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.