Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Data from: An integrated population model for estimating the relative effects of natural and anthropogenic factors on a threatened population of steelhead trout
<p>This collection includes data on the abundance, age composition, and harvest of adult steelhead trout, as well as the numbers of juveniles released from a hatchery, for steelhead trout (<em>Oncorhyncus mykiss</em>) from the Skagit River in Washington, USA. It also includes estimates of the marine survival of hatchery-origin steelhead.</p>
PlaNet model data
<p>Data for publication.</p>
Models, Data, and Scripts Underlying "Robustification of RosettaAntibody and Rosetta SnugDock"
<p>This repository contains the models, data, and scripts underlying the publication "Robustification of RosettaAntibody and Rosetta SnugDock" by <a href="https://www.biorxiv.org/content/10.1101/2020.05.26.116210v1.abstract">Jeliazkov et al.</a> Brief descriptions of the individual compressed files follow.</p> <ul> <li>fkic_H3.zip/master_H3.zip contain the models underlying Figure 4A. The models are compared against their corresponding crystal structures and the backbone RMSD is calculated for the CDR H3 (residues 95-102, Chothia definition). Each zip file contains PDB-labeled directories with homology models in the base directory and loop models in the "models" subdirectory. The "master" zip used the standard H3 modeling approach in RosettaAntibody. The "fkic" directory used the standard method plus fragment insertion.</li> <li>fragment_comparison.zip contains the scripts and raw data for Figure 4B.</li> <li>fragments.tar.gz contains the fragments used to model the antibodies loops (Figure 4A) and for the comparison in Figure 4B. The traditional protein loop models and fragments in Figure 4B are published separately by <em>Pan et al.</em> [to do: update citation].</li> <li>grafting_results.zip contains the data and analysis code behind Figure 3.</li> <li>other_stray_code.zip contains the analysis code for Figures 1, 5, and the supplemental figures.</li> <li>snugdock_example.zip contains the raw data for Figure 5 which is also to be published separately by <em>Zhou et al.</em> [to do: update citation].</li> </ul>
Frog tongue segmentation data for a realistic vascular structure in modelling
<p>In the last decade, numerical models have been an increasingly important tool in biological and medical science both for the fundamental understanding of physiology as well as the potential for novel diagnostics and treatments tools in clinical application. In this paper, a nonlinear multi-scale model framework is developed for blood flow distribution in the full vascular system of an organ. We couple a quasi 1D vascular graph model to represent blood flow in larger vessels and a porous media model to describe flow in smaller vessels and capillary bed. The vascular model is based on Poiseuille's law, with pressure correction by elasticity and pressure drop estimation at vessels junctions. The porous capillary bed is modelled as a two-compartment domain (artery and venous) using Darcy's law. The fluid exchange between the artery and venous capillary bed compartments are defined as blood perfusion. </p> <p>The numerical experiments show that the proposed model for blood circulation: 1) is closely dependent on the structure and parameters of both the vascular vessels and of the capillary bed, and 2) it provides a realistic blood circulation in the organ. The advantage of the proposed model is that it is complex enough to reliably capture the main underlying physiological function, yet highly flexible as it offers the possibility of incorporating various local effects. Furthermore, the numerical implementation of the model is straightforward and allows for simulations on a regular desktop computer. </p>
Data from: Modelling species distributions limited by geographic barriers: a case study with African and American primates
<p><strong>Aim:</strong> The boundaries of species distributions are often shaped by natural barriers such as mountains and rivers, but species distribution models usually fail to include these constraints. We tested several approaches that include barriers as explanatory variables in species distribution models.</p> <p><strong>Location:</strong> Africa and South America.</p> <p><strong>Time period:</strong> Current</p> <p><strong>Major taxa studied:</strong> Primates</p> <p><strong>Methods:</strong> We modelled the ranges of pairs of species separated by a river taking into account three explanatory components: the environment (ecosystems, topo-hydrography, climate, human pressure), the spatial structure shaped by history and population dynamics (using a trend-surface approach), and rivers as naturals barriers to dispersal (using a binary cis-trans variable that describes both sides of the river). To assess how the addition of a spatial structure and the barrier could improve distribution models, we used a nested approach by comparing models based on: a) the environment; b) the environment and the spatial structure; and c) the environment, the spatial structure and the river. These models were constructed using the favourability functions.</p> <p><strong>Results:</strong> There was a decreased occurrence of high-favourability values in the opposite side of the rivers in models that included the spatial structure of distributions, compared to models based on environment alone. This decrease was more marked when the description of the spatial structure was made more flexible. However, model performance was significantly improved by the inclusion of cis-trans variables that identified areas on the opposite side as totally unfavourable.</p> <p><strong>Main conclusions:</strong> The performance of distribution models can improve by the use of approaches that describe barriers. Although adding the location of geographic units in relation to a river appears to be the most accurate way to define the presence of a barrier, defining this variable may be challenging. A suitable alternative is to analyse the spatial structure of distributions using a flexible approach.</p>
Data and R Code for publication: Should dispersers be fast learners? Modelling the role of cognition in dispersal
<p><span><span><span><span><span><span><span><span><span><span><span>Both cognitive abilities and dispersal tendencies can vary strongly between individuals. Since cognitive abilities may help dealing with unknown circumstances it is conceivable that </span></span></span></span></span></span></span></span></span></span></span><span><span><span><span><span><span><span><span><span><span><span>dispersers may rely more heavily on learning abilities than residents. However, cognitive abilities are costly and leaving a familiar place might result in losing the advantage of having learned to deal with local conditions. Thus, individuals which invested in learning to cope with local conditions may be more reluctant to leave their natal place. In order to disentangle the complex relationship between dispersal and learning abilities we implemented individual-based simulations. By allowing for developmental plasticity, individuals could either develop a "resident" or "dispersal" cognitive phenotype.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span><span>In line with our expectations, the correlation between learning abilities and dispersal could take any direction, depending how much time individuals had to recoup their investment in cognition. Both, longevity and the timing of dispersal within lifecycles determine the time individuals have to recoup that investment and thus crucially influence this correlation. We therefore suggest that species' life-history will strongly impact the expected cognitive abilities of dispersers, relative to their resident conspecifics, and that cognitive abilities might be an integral part of dispersal syndromes.</span></span></span></span></span></span></span></span></span></span></span></p>
Data from: Combining correlative and mechanistic niche models with human activity data to elucidate the invasive potential of a sub-Antarctic insect
<p>Aim</p> <p>Correlative Species Distribution Models (SDMs) are subject to substantial spatio-temporal limitations when historical occurrence records of data-poor species provide incomplete and outdated information for niche modelling. Complementary mechanistic modelling techniques can, therefore, offer a valuable contribution to underpin more physiologically-informed predictions of biological invasions, the risk of which is often exacerbated by climate change. In this study we integrate physiological and human pressure data to address the uncertainties and limitations of correlative SDMs and to better understand, predict, and manage biological invasions.</p> <p>Location</p> <p>Western archipelagos of the Southern Ocean and martime Antarctica</p> <p>Taxon</p> <p>Eretmoptera murphyi (Chironomidae), invertebrates.</p> <p>Methods</p> <p>Mahalanobis Distances were used for correlative SDM construction for a species with few records. A mechanistic SDM was built around different fitness components (larval survival and life stage progression) as a function of temperature. SDM predictions were combined with human activity levels in Antarctica to generate a site vulnerability index to the colonization of E. murphyi. Future scenarios of ecophysiological suitability were built around the warming trends in the region.<br> Results Both SDMs converge to predict high environmental suitability in the species' native and introduced ranges. However, the mechanistic model indicates a slightly larger invasive potential based on larval performance at different temperatures. Human activity levels across the Antarctic Peninsula play a key role in discerning site vulnerabilities. Niche suitability in Antarctica grows considerably under long-term climate scenarios, leading to a substantially higher invasive threat to the Antarctic ecosystems. In turn changing conditions result on growing physiological mismatches with the environment in the native range on South Georgia.</p> <p>Main conclusions</p> <p>Long-term studies of invasion potential under climate benefit from integrating correlative predictions with physiological experiments, as the invasion potential varies depending on the area and the timescale examined. This study also highlights a conservation paradox whereby the accidental introduction of an insect represents a threat to the Antarctic ecoystems that contrasts with its endangered status at the native range.</p>
Data set - Lagrangian observations and modelling of turbulence along a tidally influenced river - Refined model
<p>The 'Dataset_Kaipara_Lagrangian_refined.mat' file contains Lagrangian observations collected in the Kaipara river and corresponding model predictions, after refinement of the model. </p> <p> </p> <p>The 'Kaipara_model_refined.mat' files contains the grid and bathymetry of a model of the Kaipara River, New Zealand, created notably in order to study turbulence in a Lagrangian frame of reference. </p>
model data of 75 ma
<p>Modeling the effect of the Gangdese Mountains on the Asian climate during the Late Cretaceous.</p>
Simulation data for the model of the Sagittarius stream in the presence of the Large Magellanic Cloud
<p>This archive contains simulations of the disrupting Sagittarius galaxy in the combined potential of the Milky Way and the Large Magellanic Cloud.</p> <p><strong>Sgr_snapshot</strong><br> contains the final (present-day) snapshot from the fiducial simulation with a triaxial Milky Way halo and M_LMC=1.5e11 Msun (see the readme file in that folder for details).</p> <p><strong>Sgr_snapshot_noLMC</strong><br> contains the same data but for a model without the LMC (which does not reproduce some aspects of the observations, but is nevertheless useful for a comparison with the other one).</p> <p><strong>potentials_triax</strong><br> contains the initial and subsequently evolving potentials of both the Milky Way and LMC, represented by multipole expansions, as well as the trajectory of the LMC and the reflex motion-induced acceleration of the Milky Way -- everything that is needed to study the dynamics of the Sgr stream and other objects in a time-dependent potential of the interacting Milky Way and LMC. This model corresponds to the stream simulation in the previous folder.</p> <p><strong>potentials_axisym</strong><br> contains the same data, but for another Milky Way halo model, which is axisymmetric rather than triaxial (note that in either case, its axis ratios vary with radius). It may be more convenient in certain applications, and produces a stream that fits the observations almost as well as the triaxial model.</p> <p><strong>scripts</strong><br> contains the Python scripts illustrating how to integrate orbits in a time-dependent potential, and how to construct initial conditions for these simulations (the parameter files for various choices of Milky Way halo potentials are also included). These scripts use the Agama framework, available at http://agama.software</p>
xltu/GSL_Joint_Min_Entropy: First release of the GSL paper related data and model
<p>First release</p>
Data and model files of dynELMOD used in OSE project
<p>The upload provides all the model files and data assumptions used for the dynELMOD runs performed in the Open Source Energiewende project.</p>
Numerical model code, input files and output data for publication ``Mixing and Transformation in a Deep Western Boundary Current: a case study''
<p>Contains numerical model data (code, input files, selected output) to supplement publication ``Mixing and Transformation in a Deep Western Boundary current'', by Spingys and co-authors. All umerical model data, including any errors, is the responsibility of Sonya Legg. This data set will allow reproduction of simulations used in the above-referenced paper.</p>
Amory et al. (2021), Geoscientific Model Development : data, model outputs and source code
<p><strong>Data and model outputs for the replication of the analysis made in:</strong><br> (see the published version of this article in Geoscientific Model Development, 2021 - please cite this version if you use these data)<br> C. Amory, C. Kittel, L. Le Toumelin, C. Agosta, A. Delhasse, V. Favier, and X. Fettweis: Performance of MAR (v3.11) in simulating the drifting-snow climate and surface mass balance of Adelie Land, East Antarctica, Geoscientific Model Development, accepted, 2021. </p> <p>See README.txt for a full description of the dataset content</p> <p>Please contact me at amory.charles@live.fr if you need other half-hourly outputs or for more details on the dataset</p>
Data from: A model of developmental canalization, applied to human cranial form
<p>Developmental mechanisms that canalize or compensate perturbations of organismal development (targeted or compensatory growth) are widely considered a prerequisite of individual health and the evolution of complex life, but little is known about the nature of these mechanisms. It is even unclear if and how a "target trajectory" of individual development is encoded in the organism's genetic-developmental system or, instead, emerges as an epiphenomenon. Here we develop a statistical model of developmental canalization based on an extended autoregressive model. We show that under certain assumptions the strength of canalization and the amount of canalized variance in a population can be estimated, or at least approximated, from longitudinal phenotypic measurements, even if the target trajectories are unobserved. We extend this model to multivariate measures and discuss reifications of the ensuing parameter matrix. We apply these approaches to longitudinal geometric morphometric data on human postnatal craniofacial size and shape as well as to the size of the frontal sinuses. Craniofacial size showed strong developmental canalization during the first 5 years of life, leading to a 50% reduction of cross-sectional size variance, followed by a continual increase in variance during puberty. Frontal sinus size, by contrast, did not show any signs of canalization. Total variance of craniofacial shape decreased slightly until about 5 years of age and increased thereafter. However, different features of craniofacial shape showed very different developmental dynamics. Whereas the relative dimensions of the nasopharynx showed strong canalization and a reduction of variance throughout postnatal development, facial orientation continually increased in variance. Some of the signals of canalization may owe to independent variation in developmental timing of cranial components, but our results indicate evolved, partly mechanically induced mechanisms of canalization that ensure properly sized upper airways and facial dimensions.</p>
Combining Satellite Remote Sensing and Climate Data in Species Distribution Models to Improve the Conservation of Iberian White Oaks (Quercus L.)
<p>The Iberian Peninsula hosts a high diversity of oak species, being a hot-spot for the conservation of European White Oaks (Quercus) due to their environmental heterogeneity and its critical role as a phylogeographic refugium. Identifying and ranking the drivers that shape the distribution of White Oaks in Iberia requires that environmental variables operating at distinct scales are considered. These include climate, but also ecosystem functioning attributes (EFAs) related to energy–matter exchanges that characterize land cover types under various environmental settings, at finer scales. Here, we used satellite-based EFAs and climate variables in species distribution models (SDMs) to assess how variables related to ecosystem functioning improve our understanding of current distributions and the identification of suitable areas for White Oak species in Iberia. We developed consensus ensemble SDMs targeting a set of thirteen oaks, including both narrow endemic and widespread taxa. Models combining EFAs and climate variables obtained a higher performance and predictive ability (true-skill statistic (TSS): 0.88, sensitivity: 99.6, specificity: 96.3), in comparison to the climate-only models (TSS: 0.86, sens.: 96.1, spec.: 90.3) and EFA-only models (TSS: 0.73, sens.: 91.2, spec.: 82.1). Overall, narrow endemic species obtained higher predictive performance using combined models (TSS: 0.96, sens.: 99.6, spec.: 96.3) in comparison to widespread oaks (TSS: 0.80, sens.: 92.6, spec.: 87.7). The Iberian White Oaks show a high dependence on precipitation and the inter-quartile range of Normalized Difference Water Index (NDWI) (i.e., seasonal water availability) which appears to be the most important EFA variable. Spatial projections of climate–EFA combined models contribute to identify the major diversity hotspots for White Oaks in Iberia, holding higher values of cumulative habitat suitability and species richness. We discuss the implications of these findings for guiding the long-term conservation of IberianWhite Oaks and provide spatially explicit geospatial information about each oak species (or set of species) relevant for developing biogeographic conservation frameworks.</p>
Data-driven subgrid-scale modeling of forced Burgers turbulence using deep learning with generalization to higher Reynolds numbers via transfer learning
<p>These are the data files for use with the codes in https://github.com/envfluids/Burgers_DDP_and_TL.</p>
Data from: Numerical models for assessing the risk of leaflet thrombosis post-transcatheter aortic valve-in-valve implantation
<p>Leaflet thrombosis has been suggested as the reason for the reduced leaflet motion in cases of hypoattenuated leaflet thickening of bioprosthetic aortic valves. This work aimed to estimate the risk of leaflet thrombosis in two post-ViV configurations, using five different numerical approaches. Realistic ViV configurations were calculated by modeling the deployments of the latest version of transcatheter aortic valve devices (Medtronic Evolut PRO, Edwards SAPIEN 3) in the surgical Sorin Mitroflow. Computational fluid dynamics simulations of blood flow followed the dry models. Lagrangian and Eulerian measures of near-wall stagnation were implemented by particle and concentration tracking, respectively, to estimate the thrombogenicity and to predict the risk locations. Most of the numerical approaches indicate on a higher leaflet thrombosis risk in the Edwards SAPIEN 3 device because of its intra-annular implantation. The Eulerian approaches estimated high-risk locations in agreement with the WSS separation points. On the other hand, the Lagrangian approaches predicted high-risk locations at the proximal regions of the leaflets matching the low WSS magnitude regions of both TAVI models and reported clinical and experimental data. The proposed methods can help optimizing future designs of transcatheter aortic valves with minimal thrombotic risks.</p>
Data from: A proteomics approach for the identification of cullin-9 (CUL9) related signaling pathways in induced pluripotent stem cell models
<p>CUL9 is a non-canonical and poorly characterized member of the largest family of E3 ubiquitin ligases known as the Cullin RING ligases (CRLs). Most CRLs play a critical role in developmental processes, however, the role of CUL9 in neuronal development remains elusive. We determined that deletion or depletion of CUL9 protein causes aberrant formation of neural rosettes, an in vitro model of early neuralization. In this study, we applied mass spectrometric approaches in human pluripotent stem cells (hPSCs) and neural progenitor cells (hNPCs) to identify CUL9 related signaling pathways that may contribute to this phenotype. Through LC-MS/MS analysis of immunoprecipitated endogenous CUL9, we identified several subunits of the APC/C, a major cell cycle regulator, as potential CUL9 interacting proteins. Knockdown of the APC/C adapter protein FZR1 resulted in a significant increase in CUL9 protein levels, however, CUL9 does not appear to affect protein abundance of APC/C subunits and adapters or alter cell cycle progression. Quantitative proteomic analysis of CUL9 KO hPSCs and hNPCs identified protein networks related to metabolic, ubiquitin degradation, and transcriptional regulation pathways that are disrupted by CUL9 deletion in both hPSCs and hNPCs. The results of our study build on current evidence that CUL9 may have unique functions in different cell types potentially contributing to the difficulty of identifying CUL9 substrates.</p> <p><strong>Information on data/files</strong>:</p> <p><em>Please first unzip Figure_4_CUL9-IP-LCMSMS-data.zip and Figure_7_iTRAQ-data.zip, then follow the description below.</em></p> <p>CUL9 immunoprecipitation from whole lysates collected for hPSCs were analyzed using LC/MS-MS. IgG was used as a control to determine proteins specifically enriched in the CUL9 IP. This data correlates to Figure 4 of the associated manuscript.</p> <p>Initial CUL9 immunoprecipitation was performed by VG and analyzed by LC-MS/MS at UNC.</p> <p>Fig4-LCMSMS_IPCUL9_Replicate1_UNC</p> <p>Two more immunoprecipitations were performed by VG and analyzed by LC-MS/MS at Vanderbilt University MSRC Proteomics Core. The first run (replicate 1) was used to produce the STRING figure in the associated manuscript, as well as the volcano plot in Supplemental Figure 8.</p> <p>Two Scaffold files contain raw data and our analysis parameters:</p> <ul> <li>Fig4-LCMSMS_IPCUL9_Replicate2_Vanderbilt</li> <li>Fig4-LCMSMS_LCMSMS_IPCUL9_Replicate3_Vanderbilt</li> </ul> <p>Excel files exported from the above Scaffold files contain the raw data in excel format including spectral counts, peptide counts, protein probability, and detailed raw data as it pertains to each identified protein. See these data sets in the folders below:</p> <ul> <li>Fig4-LCMSMS_IPCUL9_Run2_Vanderbilt_CompleteDataSet</li> <li>Fig4-LCMSMS_IPCUL9_Run3_Vanderbilt_CompleteDataSet</li> </ul> <p>CUL9 KO clones were used to identify proteins increased or decreased compared to parental wild-type cell lines using iTRAQ. This data correlates to Figure 8 of the associated manuscript. Two iTRAQ experiments were performed to identify proteins altered in CUL9 KO iPSCs, neural stem cells (NSCs), and NPCs. Each set of experiments was duplicated – each in a different isogenic CUL9 KO clone (Clone 20 or Clone B)</p> <p><strong>iTRAQ labeling of experiment one is as follows:</strong></p> <ol> <li>Label 115: Parental WT NSCs</li> <li>Label 117: CUL9 KO Clone 20 or B NSCs</li> <li>Label 114: Parental WT NPCs</li> <li>Label 116: CUL9 KO Clone 20 or B NPCs</li> </ol> <p>Files containing raw and analyzed data of experiment one are labeled as follows:</p> <ul> <li>Fig7_Experiment1_Clone 002_compared-to-NSC-WT</li> <li>Fig7_Experiment1_Clone001-compared-to-NSC-WT</li> <li>Fig7_Experiment1_Clone001iPSC_compared-to-iPSCWT</li> <li>Fig7_Experiment1_Clone002_compared-to-iPSCWT</li> </ul> <p><strong>iTRAQ labeling of experiment two is as follows:</strong></p> <p>For Clone 20:</p> <ol> <li>Label 115: Parental WT NSCs</li> <li>Label 117: CUL9 KO Clone 20 or B NSCs</li> <li>Label 114: Parental WT iPSCs</li> <li>Label 116: CUL9 KO Clone 20 or B iPSCs</li> </ol> <p>Files containing raw and analyzed data of experiment two are labeled as follows:</p> <ul> <li>Fig7_Experiment2_Clone001NPC_Compared-to-NPC-WT</li> <li>Fig7_Experiment2_Clone001NPC_Compared-to-NSC-WT</li> <li>Fig7_Experiment2_Clone002NPC_Compared-to-NPC-WT</li> <li>Fig7_Experiment2_Clone002NPC_Compared-to-NSC-WT</li> </ul> <p>Please note that the first tab of each excel document contains the full raw data set before analysis. All other tabs contain analyzed data comparing the proteins identified in two cell lines. Analyzed tabs are labeled to indicate which cell lines are being compared. Please note that the CUL9 isogenic clones Clone 20=Clone#1 and Clone B=Clone#2 as referenced in the associated manuscript.</p> <p>Detailed information about the protocols used for collection and analysis of this data can be found in the supplemental information within the associated manuscript.</p>
Data from: Identifying 'useful' fitness models: balancing the benefits of added complexity with realistic data requirements in models of individual plant fitness
<p>Direct species interactions are commonly included in individual fitness models used for coexistence and local-diversity modeling. Though widely considered important for such models, direct interactions alone are often insufficient for accurately predicting fitness, coexistence or diversity outcomes. Incorporating higher-order interactions (HOIs) can lead to more accurate individual fitness models, but also adds many model terms, which can quickly result in model over-fitting. We explore approaches for balancing the trade-off between tractability and model accuracy that occurs when HOIs are added to individual fitness models. To do this, we compare models parameterized with data from annual plant communities in Australia and Spain, varying in the extent of information included about the focal and neighbor species. The best performing models for both datasets were those that grouped neighbors based on origin status and life form, a grouping approach that reduced the number of model parameters substantially while retaining important ecological information about direct interactions and HOIs. Results suggest that specific focal- or neighbor-species identity is not necessary for building well-performing fitness models that include HOIs. In fact, grouping neighbors by even basic functional information seems sufficient to maximize model accuracy, an important outcome for the practical use of HOI-inclusive fitness models.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.