Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
16
datasets available to search
ShareScore release 0.9.0
Dataset results
16 results for “linear mixed models”
Minimal dataset for "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models"
<p>This repository contains a minimal data set to reproduce all results that don't compromise the privacy concerns for the manuscript "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models".<br> <br> The repository contains the following data:</p> <ul> <li>adaptscore_acute.csv <ul> <li>A csv file that contains the estimated adaptation scores for the acute data set with HLA I model.</li> </ul> </li> <li>adaptscore_leftout.csv <ul> <li>A csv file that contains the estimated adaptation scores for the leftout data set with the joint HLA I and HLA II model</li> </ul> </li> <li>adaptscore_training.csv <ul> <li>A csv file that contains the estimated adaptation scores for the traininig data set with the joint HLA I and HLA II model</li> </ul> </li> <li>adaptscore_training_hla1_without_clin.csv <ul> <li>A csv file that contains the estimated adaptation scores for the training data set with the HLA I model (via cross-validation)</li> </ul> </li> <li>adaptscore_training_seed2.csv <ul> <li>A csv file that contains the estimated adaptation scores for the training data set with the joint HLA I and HLA II model via cross-validation with another seed</li> </ul> </li> </ul>
SI Figure 1: Dispersion values (a boxplot using distance to centroids based on Bray Curtis distance matrix) of external and internal bacterial microbiome composition for different hosts. In a mixed linear model, microinvertebrates did not significantly impact dispersion (P=0.44), but microbiome type did (P=0.03). Pairwise contrasts show that while external microbiomes of P. murrayi and Tardigrada are more variable than their internal microbiomes, E. antarcticus external and internal microbiomes are equally variable. in External and internal microbiomes of Antarctic nematodes are distinct, but more similar to each other than the surrounding environment
SI Figure 1: Dispersion values (a boxplot using distance to centroids based on Bray Curtis distance matrix) of external and internal bacterial microbiome composition for different hosts. In a mixed linear model, microinvertebrates did not significantly impact dispersion (P=0.44), but microbiome type did (P=0.03). Pairwise contrasts show that while external microbiomes of P. murrayi and Tardigrada are more variable than their internal microbiomes, E. antarcticus external and internal microbiomes are equally variable.
Simulated data for paper "Conditional non-parametric bootstrap for non-linear mixed effect models"
<p>Data was simulated according to an Emax model (scenarios 1 and 2) or a Hill model (scenarios 3 and 4) with a rich (scenarios 1 and 3) and a sparse design (scenarios 2 and 4). The archive contains 4 folders with the data simulated in the first 4 scenarios (N=200 simulated datasets in each folder):<br> - scenario 1 - pdemax.rich<br> - scenario 2 - pdemax.sparse<br> - scenario 3 - pdhillhigh.rich<br> - scenario 4 - pdhillhigh.sparse<br> The data used in scenarios 5 and 6 was a subset of the datasets simulated in scenarios 3 and 4 respectively. In scenario 5, 20 subjects were taken from each dataset (subjects 1-5, 26-30, 51-55, 76-80) from the datasets in folder pdhillhigh.rich. In scenario 6, the datasets were constituted by the first 20 subjects from each sampling group of the data simulated in pdhillhigh.sparse.</p>
Consensus nucleotide sequences for env and gag for paper: Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models
<p>This is the consensus sequence repository to the manuscript "Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models".<br> It contains the 10% consensus nucleotide sequences of the env and gag (only p24) protein of HIV-1 used for the training and leftout data set. The NGS sequences are available under BioProject ID PRJNA810303 and the corresponding BioSample Accession IDs are SAMN26241863:26242168 and SAMN28728524:SAMN28728529</p> <ul> <li>env_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the leftout data set</li> </ul> </li> <li>env_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the training data set</li> </ul> </li> <li>gag_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the leftout data set</li> </ul> </li> <li>gag_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the training data set</li> </ul> </li> </ul>
Protea repens whole transcriptome count data for control and drought treatment for 8 populations, climatic data for the 8 populations and phenotypic data collected, and data used for linear mixed models for climate gene expression/trait correlation testing
Open the record for dataset details and reuse information.
Data from: A new analytical approach to landscape genetic modeling: least-cost transect analysis and linear mixed models
Landscape genetics aims to assess the effect of the landscape on intraspecific genetic structure. To quantify interdeme landscape structure, landscape genetics mostly uses landscape resistance surfaces and least-cost paths or straight-line transects. However, both approaches have drawbacks. Parameterization of resistance surfaces is a subjective process, and least-cost paths represent a single migration route. A transect-based approach might oversimplify migration patterns by assuming rectilinear migration. To overcome these limitations, we combined these two methods in a new landscape genetic approach: least-cost transect analysis (LCTA). Habitat-matrix resistance surfaces were used to create least-cost paths, which were subsequently buffered to form transects in which the abundance of several landscape elements was quantified. To maintain objectivity, this analysis was repeated so that each landscape element was in turn regarded as migration habitat. The relationship between landscape predictor variables and genetic distances was then assessed following a mixed modeling approach to account for the non-independence of values in distance matrices. Subsequently, predictor variables were selected making use of the R_β^2 statistic. We applied LCTA and the mixed model approach to an empirical genetic dataset on the endangered damselfly, Coenagrion mercuriale. We compared the results to those obtained from traditional least-cost, effective and resistance distance analysis and showed that LCTA not only outperforms existing methods in a statistical way, but also provides more information about the migration ecology of the focal species. Although we believe the statistical approach to be an improvement for the analysis of distance matrices in landscape genetics, more stringent testing is needed.
Data from: A new analytical approach to landscape genetic modeling: least-cost transect analysis and linear mixed models
Open the record for dataset details and reuse information.
Raw data and results for the paper "Conditional non-parametric bootstrap for non-linear mixed effect models"
<p>*Data* (comets_condBoot_data.zip)</p> <p>Data was simulated according to an Emax model (scenarios 1 and 2) or a Hill model (scenarios 3 and 4). The archive contains 4 folders with the data simulated in the first 4 scenarios (N=200 simulated datasets in each folder):<br> - scenario 1 - pdemax.rich<br> - scenario 2 - pdemax.sparse<br> - scenario 3 - pdhillhigh.rich<br> - scenario 4 - pdhillhigh.sparse<br> The data used in scenarios 5 and 6 was a subset of the datasets simulated in scenarios 3 and 4 respectively. In scenario 5, 20 subjects were taken from each dataset (subjects 1-5, 26-30, 51-55, 76-80) from the datasets in folder pdhillhigh.rich. In scenario 6, the datasets were constituted by the first 20 subjects from each sampling group of the data simulated in pdhillhigh.sparse.</p> <p>*Results:* (comets_scenarioXXX_results.zip, XXX=1,.. 6)</p> <p>6 simulation scenarios were assessed in the paper. Each file corresponds to 1 of 6 folders, one for each scenario:<br> - scenario 1 - pdemax.rich/results<br> - scenario 2 - pdemax.sparse/results<br> - scenario 3 - pdhillhigh.rich/results<br> - scenario 4 - pdhillhigh.sparse/results<br> - scenario 5 - pdhillhigh.n20rich/results<br> - scenario 6 - pdhillhigh.n20sparse/results</p> <p>In each "results" subfolder, the results for each bootstrap method and each dataset are written to a separate file, eg for simulation 1 in the first scenario:<br> - case bootstrap: scenarioHill1_bootstrapCase_sim1.res <br> - non-parametric bootstrap: scenarioHill1_bootstrapNP_sim1.res<br> - conditional non-parametric bootstrap: scenarioHill1_bootstrapNPc_sim1.res<br> - parametric bootstrap: scenarioHill1_bootstrapPar_sim1.res<br> The folder also contains:<br> - the saemix estimates for the 200 simulations: scenarioHill1_fitOrig.res<br> - tables with the bias and SE for the different bootstraps over the set of simulations, used to evaluate the methods: rbiasSEboot200.res, rbiasSEboot.res, rbiasWRsampleEstimates.res</p> <p> </p>
Data from: Mixed linear model approach for mapping quantitative trait loci underlying crop seed traits
The crop seed is a complex organ that may be composed of the diploid embryo, the triploid endosperm and the diploid maternal tissues. According to the genetic features of seed characters, two genetic models for mapping quantitative trait loci (QTLs) of crop seed traits are proposed, with inclusion of maternal effects, embryo or endosperm effects of QTL, environmental effects and QTL-by-environment (QE) interactions. The mapping population can be generated either from double back-cross of immortalized F2 (IF2) to the two parents, from random-cross of IF2 or from selfing of IF2 population. Candidate marker intervals potentially harboring QTLs are first selected through one-dimensional scanning across the whole genome. The selected candidate marker intervals are then included in the model as cofactors to control background genetic effects on the putative QTL(s). Finally, a QTL full model is constructed and model selection is conducted to eliminate false positive QTLs. The genetic main effects of QTLs, QE interaction effects and the corresponding P-values are computed by Markov chain Monte Carlo algorithm for Gaussian mixed linear model via Gibbs sampling. Monte Carlo simulations were performed to investigate the reliability and efficiency of the proposed method. The simulation results showed that the proposed method had higher power to accurately detect simulated QTLs and properly estimated effect of these QTLs. To demonstrate the usefulness, the proposed method was used to identify the QTLs underlying fiber percentage in an upland cotton IF2 population. A computer software, QTLNetwork-Seed, was developed for QTL analysis of seed traits.
Data from: miRglmm: a generalized linear mixed model of isomiR-level counts improves estimation of miRNA-level differential expression and uncovers variable differential expression between isomiRs
<p>These datasets can be used to reproduce all analyses from the publication "miRglmm: a generalized linear mixed model of isomiR-level counts improves estimation of miRNA-level differential expression and uncovers variable differential expression between isomiRs" in conjunction with codes found at https://github.com/mccall-group/miRglmm_paper. </p> <p>"Monocyte_data_subset.rda", "monocyte_exact_subset_filtered2.rda" and "sims_N100_m2_s1_rtruncnorm.rda" can be used to reproduce the simulation analysis. </p> <p>"panel_B_SE.rda" and "ERCC_filtered.rda" can be used to reproduce the ERCC synthetic data analysis with known ground truth.</p> <p>"study89_data_subset.rda" and "study89_data_subset_filtered2.rda" can be used to reproduce the immune cell-type analysis. </p> <p>"bladder_testes_data_subset.rda" and "bladder_testes_data_subset_filtered2.rda" can be used to reproduce the bladder vs testes tissue analysis.</p>
The data from CropPol which are not included in my linear mixed models and the reasons.
<p>The data from CropPol which are not included in my linear mixed models and the reasons. Note_for_not_include column refers to the reason why these data was excluded. </p>
Data from: Mixed linear model approach for mapping quantitative trait loci underlying crop seed traits
Open the record for dataset details and reuse information.
Data from: Generalized linear mixed models for mapping multiple quantitative trait loci
Open the record for dataset details and reuse information.
A linear-mixed model approach to identify expressed SFPs using short oligo microarrays - Placenta vs. Fibroblast
GEO Series GSE10446. Sus scrofa. 6 samples. Type: Expression profiling by genome tiling array.
A linear-mixed model approach to identify expressed single nucleotide polymorphisms using short oligo microarrays
GEO Series GSE10447. Sus scrofa. 6 samples. Type: Expression profiling by genome tiling array.
Dataset of the paper "Minimal detectable change of gait and balance measures in older neurological patients: estimating the standard error of the measurement from before-after rehabilitation data thanks to the linear mixed-effects models"
<p>The Excel file contains the dataset whose analysis has been presented in the manuscript entitled <em>"Minimal detectable change of gait and balance measures in older neurological patients: estimating the standard error of the measurement from before-after rehabilitation data thanks to the linear mixed-effects models" </em>and published in J Neuroeng Rehabil (doi: 10.1186/s12984-024-01339-4; PMID: 38566189).</p> <p>This data set cannot be publicly available because of sensitive information. Please send your request to: a.caronni@auxologico.it.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.