Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,805
datasets available to search
ShareScore release 0.9.0
Dataset results
5,805 results for “Data model”
Experimental data for Dynamic cover effects in lateral bedrock channel bank abrasion: Experiment and model comparison
<p>Experimental data for bank erosion.xlsx contains the data used for the figures in the paper, and the distribution of bedrock bank erosion in the longitudinal direction in Run 1 - Run 18.</p>
Occupant Simulation Data based on Honda Accord 2024 Simplified Passenger Model and Full-factorial Sampling with 3,125 samples and HIII05F, HIII50M, HIII95M
<p>Database and FE-models with 9,375 Honda Accord 2014 passenger occupant simulations. </p> <p> </p>
Code and Data for Sturm and Silva (2024) A nudge to the truth: atom conservation as a hard constraint in models of atmospheric composition using an uncertainty-weighted correction
<p>This record contains the Julia photochemical model (https://doi.org/10.5281/zenodo.13385541) output in csv format used for training XGBoost in ProjectionConservationRF.py to emulate ozone photochemical formation. Nonphysical predictions that violate conservation of atoms are corrected using a closed-form, constrained least-squares approach that factors in uncertainty and scale using species-level weights. The file ozoneNOx_visualization.py contains an example and visualization for a smaller system, the primary photolytic cycle from which the Leighton relationship can be derived.</p> <p>The corresponding preprint is available here: <a href="https://doi.org/10.48550/arXiv.2408.16109">https://doi.org/10.48550/arXiv.2408.16109</a></p>
Paradise fish (Macropodus opercularis) as novel translational model for emotional and cognitive function - RAW DATA
<p>The data of the study under the title "Paradise fish (Macropodus opercularis) as novel translational model for emotional and cognitive function". The structure of the data follows the structure of the results section of the paper. Data sheets are labeled according to figure numbers of the papers.</p>
Gene Expression Transcriptomics Data for benchmarking Perturbation Models - Part 2
Open the record for dataset details and reuse information.
Data and Code from: Dysregulation of zebrin-II cell subtypes is a shared feature across polyglutamine ataxia mouse models and human patients
<div> <div> <div> <p>Abstract</p> <p>Spinocerebellar ataxia type 7 (SCA7) is a genetic neurodegenerative disorder caused by a CAG- polyglutamine repeat expansion. Purkinje cells (PCs) are central to the pathology of ataxias, but their low abundance in the cerebellum underrepresents their transcriptomes in sequencing assays. To address this issue, we developed a PC enrichment protocol and sequenced individual nuclei from mice and patients with SCA7. Single-nucleus RNA sequencing in SCA7-266Q mice revealed dysregulation of cell identity genes affecting glia and PCs. Specifically, genes marking zebrin-II PC subtypes accounted for the highest proportion of DEGs in symptomatic SCA7-266Q mice. These transcriptomic changes in SCA7-266Q mice were associated with increased numbers of inhibitory synapses as quantified by immunohistochemistry and reduced spiking of PCs in acute brain slices. Dysregulation of zebrin-II cell subtypes was the predominant signal in PCs of SCA7-266Q mice and was associated with the loss of zebrin-II striping in the cerebellum at motor symptom onset. We furthermore demonstrated zebrin-II stripe degradation in additional mouse models of polyglutamine ataxia and observed decreased zebrin-II expression in cerebellum of patients with SCA7. Our results suggest that a breakdown of zebrin subtype regulation is a shared pathological feature of polyglutamine ataxias.</p> <p>Data and Code Availability</p> <p>Here you will find data and code associated with our manuscript "Dysregulation of zebrin-II cell subtypes is a shared feature across polyglutamine ataxia mouse models and human patients", Bartelt et al., <em>Sci. Trans. Med. </em>16, eadn5449 (2024).</p> <p>The data file labeled "HuCb_filtered.rds" is a processed and annotated single-nucleus RNA-seq Seurat object, containing the gene-level count data for the multiplexed snRNA-seq experiment performed on post-mortem human cerebellar tissues from patients with SCA7 and unaffected controls. Data obtained from WT and SCA7-266Q mice as described in our paper can be accessed in the NIH Gene Expression Omnibus under accession number GSE269430.</p> <p>There are three code files numbered 00 through 02 which contain analysis code for snRNA-seq data applied to both the mouse and human datasets. These files are sequential and will take the user from CellRanger output, to filtered and annotated Seurat objects, and include details for subclustering analysis as well as our pseudobulk DEseq2 differential expression approach. There are places where the user may need to modify the code based on their computer system, version of R or Seurat, and whether they are processing the 5 week, 8 week, or human data sets; these locations in the code are marked with comments.</p> <ul> <li>The first file, 00_Preprocessing_MULTIseq, begins with CellRanger filtered_feature_barcode_matrix output, extracts cell barcodes, utilizes the MULTIseq deMULTIplex software to match cell barcodes to oligo barcodes from MULTIseq fastq files, and annotates the Seurat file with metadata. Cell type identification and annotation also takes place in this file. Note: the deMULTIplex step will likely need to be run on a high performance compute cluster.</li> <li>The second file, 01_Seurat_Analysis, uses the filtered and annotated Seurat file to calculate useful QC metrics, investigate disease signals, and perform cell type subclustering analyses.</li> <li>The third file, 02_Pseudobulk_DEseq2, contains custom analysis code to extract raw counts for each cell type and each animal from the Seurat file, and uses the DEseq2 package to calculate DEGs, taking into account biological replicates, and raw read count differences between control and SCA7 animals.</li> </ul> </div> </div> </div>
Temperature data corresponding to "Hybrid Phenology Modeling for Predicting Temperature Effects on Tree Dormancy"
<p>MERRA2 Temperature data corresponding to "Hybrid Phenology Modeling for Predicting Temperature Effects on Tree Dormancy"</p>
Improving generalisability of 3D binding affinity models in low data regimes
<p>Structures of the PDBBind dataset (general protein-ligand) prepared with CCDC protein preparation software. After preparation, 18310 structures out of the total 19443 remained (1133 failed).</p>
Prosit-XL Models Data
<p>Training data used to train fragment ion intensity Prosit-XL models.</p>
Three-point contact data for "Multi-contact statistics distinguish models of chromosome organization"
<p>Three-point contact data for the publication "Multi-contact statistics distinguish models of chromosome organization".</p>
Data and code for "Mapping hotspots of zoonotic pathogen emergence: an integrated model- and participatory-based approach"
<p>Code for "Mapping hotspots of zoonotic pathogen emergence: an integrated model- and participatory-based approach". Note shapefiles will need to be downloaded from GADM (https://gadm.org/), or from the gadm package in R (https://rdrr.io/github/rspatial/geodata/man/gadm.html). </p> <p>Data are available in Version 1.0 of this record</p>
Data from: A test of the hierarchical model of litter decomposition
Our basic understanding of plant litter decomposition informs the assumptions underlying widely applied soil biogeochemical models, including those embedded in Earth system models. Confidence in projected carbon cycle-climate feedbacks therefore depends on accurate knowledge about the controls regulating the rate at which plant biomass is decomposed into products such as CO2. Here, we test underlying assumptions of the dominant conceptual model of litter decomposition. The model posits that a primary control on the rate of decomposition at regional to global scales is climate (temperature and moisture), with the controlling effects of decomposers negligible at such broad spatial scales. Using a regional-scale litter decomposition experiment at six sites spanning from northern Sweden to southern France – and capturing both within and among site variation in putative controls – we find that contrary to predictions from the hierarchical model, decomposer (microbial) biomass strongly regulates decomposition at regional scales. Further, the size of the microbial biomass dictates the absolute change in decomposition rates with changing climate variables. Our findings suggest the need for revision of the hierarchical model, with decomposers acting as both local- and broad-scale controls on litter decomposition rates, necessitating their explicit consideration in global biogeochemical models.
Geodetic model of the March 2021 Thessaly seismic sequence inferred from seismological and InSAR data
<p>A selection of Sentinel-1 (S1) wrapped and unwrapped measurements used in this study (from "a" to "u" files in tiff format as indicated in the word file attached). S1 data were processed by using our own internally developed InSAR processing chain.<br> <br> Earthquakes data locations.</p> <p><br> </p> <p> </p>
Complex ecological phenotypes on phylogenetic trees: a Markov process model for comparative analysis of multivariate count data
The evolutionary dynamics of complex ecological traits – including multistate representations of diet, habitat, and behavior – remain poorly understood. Reconstructing the tempo, mode, and historical sequence of transitions involving such traits poses many challenges for comparative biologists, owing to their multidimensional nature. Continuous-time Markov chains (CTMC) are commonly used to model ecological niche evolution on phylogenetic trees but are limited by the assumption that taxa are monomorphic and that states are univariate categorical variables. A necessary first step in the analysis of many complex traits is therefore to categorize species into a pre-determined number of univariate ecological states, but this procedure can lead to distortion and loss of information. This approach also confounds interpretation of state assignments with effects of sampling variation because it does not directly incorporate empirical observations for individual species into the statistical inference model. In this study, we develop a Dirichlet-multinomial framework to model resource use evolution on phylogenetic trees. Our approach is expressly designed to model ecological traits that are multidimensional and to account for uncertainty in state assignments of terminal taxa arising from effects of sampling variation. The method uses multivariate count data for individual species to simultaneously infer the number of ecological states, the proportional utilization of different resources by different states, and the phylogenetic distribution of ecological states among living species and their ancestors. The method is general and may be applied to any data expressible as a set of observational counts from different categories.
Data from: Diversity, dynamics and effects of long terminal repeat retrotransposons in the model grass Brachypodium distachyon
<ul> <li><span>Transposable elements (TEs) are the main reason for the high plasticity of plant genomes, where they occur as communities of diverse evolutionary lineages. Because research has typically focused on single abundant families or summarized TEs at a coarse taxonomic level, our knowledge about how these lineages differ in their effects on genome evolution is still rudimentary. </span></li> <li><span>Here we investigate the community composition and dynamics of 32 long terminal repeat retrotransposon (LTR-RT) families in the 272 Mb genome of the Mediterranean grass <i>Brachypodium distachyon. </i></span></li> <li><span>We find that much of the recent transpositional activity in the <i>B. distachyon </i>genome is due to centromeric <i>Gypsy </i>families and <i>Copia </i>elements belonging to the Angela lineage. With a half-life as low as 66 ky, the latter are the most dynamic part of the genome and an important source of within-species polymorphisms. Second, GC-rich <i>Gypsy </i>elements of the Retand lineage are the most abundant TEs in the genome. Their presence explains more than 20 percent of the genome-wide variation in GC content and is associated with higher methylation levels. </span></li> <li><span>Our study shows how individual TE lineages change the genetic and epigenetic constitution of the host beyond simple changes in genome size. </span></li> </ul>
Saguaro recruitment data obtained by inverse-growth modelling
<p>Each year, an individual mature large saguaro cactus produces about one million seeds in attractive juicy fruits that lure seed predators and seed dispersers in a three-month feast. From the million seeds produced, however, only a few will persist into mature saguaros. A century of research on saguaro population dynamics has led to the conclusion that saguaro recruitment is an episodic event that depends on the convergence of suitable conditions for survival during the critical early stages. Because most data have been collected in Arizona, particularly in the surroundings of Tucson, most research has relied on a limited amount of environmental variation. In this study, we upscaled this knowledge on saguaro recruitment to a regional scale with a new method that used the inverse-growth modeling of 1,487 saguaros belonging to 13 populations in a latitudinal gradient ranging from arid desert to tropical thornscrub forest in Sonora, Mexico. Using generalized linear and additive mixed models, we created two 110-year-long saguaro recruitment curves: one driven only by previous size, and the second driven by size, drought, and soil structure. We found evidence that saguaro recruitment is indeed episodic with periodicities of 20–30 years possibly related to strong El Niño Southern Oscillation events. Our results suggest that saguaros rely on multidecadal periodic pulses of good beneficial years to incorporate new individuals into their populations. Inverse-growth modelling can be used in a wide variety of plant species to study their recruitment dynamics.</p>
Data from: Resolving relationships and phylogeographic history of the Nyssa sylvatica complex using data from RAD-seq and species distribution modeling
Nyssa sylvatica complex consists of several woody taxa occurring in eastern North America. These taxa were recognized as two or three species including three or four varieties by different authors. Due to high morphological similarities and complexity of morphological variation, classification and delineation of taxa in the group have been difficult and controversial. Here we employ data from RAD-seq to elucidate the genetic structure and phylogenetic relationships within the group. Using the genetic evidence, we evaluate previous classifications and delineate species. We also employ Species Distribution Modeling (SDM) to evaluate impacts of climatic changes on the ranges of the taxa and to gain insights into the relevant refugia in eastern North America. Results from Molecular Variance Analysis (AMOVA), STRUCTURE, phylogenetic analyses using Maximum likelihood, Bayesian Inference, and Splittree methods of RAD-seq data strongly support a two-clade pattern, largely separating samples of N. sylvatica from those of N. biflora-N. ursina mix. Divergence time analysis with BEAST suggests the two clades diverged in the mid Miocene. The ancestor of the present trees of N. sylvatica was suggested to be in the Pliocene and that of N. biflora-N. ursina mix in the end of the Miocene. Results from SDM predicted a smaller range in the southern part of the species present range of each clade during the Last Glacial Maximum (LGM). A northward expansion of the ranges during interglacial period and a northward shift of the ranges in the future under a model of global warming were also predicted. Our results support the recognition of two species in the complex, N. sylvatica and N. biflora, following the phylogenetic species concept. We found no genetic evidence supporting recognitions of intraspecific taxa. However, we propose subsp. ursina and subsp. biflora within N. biflora due to their distinction in habits, distributions, and habitats. Our results further support movements of trees in eastern North America in response to climatic changes. Finally, our study demonstrates that RAD-seq data and a combination of population genomics and SDM are valuable in resolving relationship and biogeographic history of closely related species that are taxonomically difficult.
IFC Building Models for Automated Extraction of Data from Balconies
<p>This is a set of example Industry Foundation Classes (IFC) building models to extract domain specific construction information. More speficially, to extract locations of potential placement sites for thermal bridges between balconies and their neighbouring floors.</p>
Dataset: Process Mining for Reliability Modeling of Smart Manufacturing Systems with Reduced Data Requirements
<p>Operational state logs from the Industry 4.0 Lab, University of Southern Denmark.</p> <p>"I4.0Lab_state_log.csv" -> without failures</p> <p>I4.0Lab_state_log_failures.csv -> with failures</p>
Data and simulations files for the article "Accurate modeling and characterization of photothermal forces in optomechanics"
<p>Data and simulations files for the article "Accurate modeling and characterization of photothermal forces in optomechanics".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.