Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9
datasets available to search
ShareScore release 0.9.0
Dataset results
9 results for “BLUP”
Data from SwAsp collection - environmental PCAs and bud set BLUPs
<p>Environmental variables at place of origin for individuals from the SwAsp population as well as the three first environmental PCAs calculated from this data. </p> <p>Best linear unbiased prediction of clone genetic values for bud set for clones form the SwAsp collection BLUPs are based on data from two common gardens (Sävar (63.89N, 20.54E) and Ekebo (55.95N, 13.12E)) and three years (2005, 2006 and 2007). More details about the common gardens can be found in Luquez, V., Hall, D., Albrectsen, B. R., Karlsson, J., Ingvarsson, P. K., & Jansson, S. (2008). Natural phenological variation in aspen (Populus tremula): the SwAsp collection. Tree Genetics & Genomes, 4(2), 279–292.</p>
Ensemble BLUP, Machine Learning, and Deep Learning Models Predict Maize Yield Better Than Each Model Alone.
<p>Data and scripts exploring ensembling strategies using the models developed in <a href="https://academic.oup.com/g3journal/advance-article/doi/10.1093/g3journal/jkad006/6982634">Kick et al., 2023</a> (see also <a href="https://zenodo.org/record/7401113">1</a>, <a href="https://zenodo.org/record/6916775">2</a>). Download all files to a single directory then run setup.sh or manually unzip using tar.</p> <p> </p> <table> <tbody> <tr> <td><strong>Filename</strong></td> <td><strong>Description</strong></td> </tr> <tr> <td>setup.sh</td> <td>Simple script that unzips zipped directories</td> </tr> <tr> <td>ext_data</td> <td>Reduced data from Kick et al. 2023</td> </tr> <tr> <td>ext_data_notebooks</td> <td>Contains python notebooks containing analysis and R markdown file containing visualization of results. Python and R data objects are written to allow results to be read in instead of re-generated.</td> </tr> <tr> <td>output</td> <td>Folder containing a placeholder file.</td> </tr> </tbody> </table> <p> </p> <p>This research used resources provided by the United States Department of Agriculture’s Agricultural Research Service (project number 5070-21000-041-000-D). The SCINet project of the USDA Agricultural Research Service (project number 0500-00093-001-00-D) was instrumental in the training of the models used in this work. In addition, we would like to acknowledge those presently and historically involved in generating data for the Genomes to Fields Initiative.</p> <p> </p> <p> </p> <p> </p>
Efficient weighting methods for genomic best linear unbiased prediction (BLUP) adaption to the genetic architectures of quantitative traits
<p><a name="_Hlk19877414"></a>Genomic best linear unbiased prediction (GBLUP) assumes equal variance for all marker effects, which is suitable for traits that conform to the infinitesimal model. For traits controlled by major genes, Bayesian methods with shrinkage priors or genome-wide association study (GWAS) methods can be used to identify <a name="_Hlk24974556">causal variants</a> effectively. The information from Bayesian/GWAS methods can be used to construct the weighted genomic relationship matrix (<b>G</b>). However, it remains unclear which methods perform best for traits varying in genetic architecture. Therefore, we developed several methods to <a name="_Hlk23592218">optimize</a> the performance of weighted GBLUP and compare them with other available methods using simulated and real datasets. First, two types of methods (marker effects with local-shrinkage or normal prior) were used to obtain test statistics and estimates for each marker effect. Second, three weighted <b>G</b> matrices were constructed based on the marker information from the first step: (1) the genomic-feature weighted <b>G</b> (GFWG), (2) the estimated marker-variance weighted <b>G</b> (EVWG), and (3) the absolute value of estimated marker-effect weighted <b>G</b> (AEWG). Following the above process, six different weighted GBLUP methods (local-shrinkage/normal prior GF/EV/AE-WGBLUP) were proposed for genomic prediction. Analyses with both simulated and real data demonstrated that these options offer flexibility for optimizing the weighted GBLUP for traits with a broad spectrum of genetic architectures. The advantage of weighting methods over GBLUP in terms of accuracy were trait dependent, ranging from 14.8% to marginal for simulated traits and from 44% to marginal for real traits. Local-shrinkage prior EVWGBLUP is superior for traits mainly controlled by loci of large effect. Normal prior AEWGBLUP performs well for traits mainly controlled by loci of moderate effect. For traits controlled by some loci with large effects (<a name="_Hlk49869847">explain 25%~50% genetic variance</a>) and a range of loci with small effects, GFWGBLUP has advantages. In conclusion, the optimal weighted GBLUP method for genomic selection should take both the genetic architecture and number of QTLs of traits into consideration carefully.</p>
Subspecies and Distribution. P. b. breviceps Waterhouse, 1839 — E mainland Australia in New South Wales, Victoria, and South Australia, extending into SE South Australia; also Tasmania (where may be introduced). P. b. ariel Gould, 1842 — N Western Australia through N Northern Territory, including Tiwi Is and Groote Eylandt. Pb. longicaudatus Longman, 1924 — N & E Queensland, including Normanby I (off Russell River, S of Cairns). P. b. papuanus Thomas, 1888 — Moluccas (Halmahera, Ternate I, Bacan Is, Gebe I, Kai Besar I); and New Guinea, including West Papuan Is (Misool, Salawati), Adi I (off S Bomberai = Fakfak Peninsula), islands in Cenderawasih (= Geelvink) Bay (Yapen, Numfoor), islands off N coast (Bagabag, Bam, Blup Blup, Kadovar, Karkar, Koil, Vokeo, Wei), Bismarck Archipelago (New Britain, Duke of York), and islands off SE peninsula (Fergusson, Goodenough, Sudest, Woodlark, Misima, Normanby). in Petauridae
Subspecies and Distribution. P. b. breviceps Waterhouse, 1839 — E mainland Australia in New South Wales, Victoria, and South Australia, extending into SE South Australia; also Tasmania (where may be introduced). P. b. ariel Gould, 1842 — N Western Australia through N Northern Territory, including Tiwi Is and Groote Eylandt. Pb. longicaudatus Longman, 1924 — N & E Queensland, including Normanby I (off Russell River, S of Cairns). P. b. papuanus Thomas, 1888 — Moluccas (Halmahera, Ternate I, Bacan Is, Gebe I, Kai Besar I); and New Guinea, including West Papuan Is (Misool, Salawati), Adi I (off S Bomberai = Fakfak Peninsula), islands in Cenderawasih (= Geelvink) Bay (Yapen, Numfoor), islands off N coast (Bagabag, Bam, Blup Blup, Kadovar, Karkar, Koil, Vokeo, Wei), Bismarck Archipelago (New Britain, Duke of York), and islands off SE peninsula (Fergusson, Goodenough, Sudest, Woodlark, Misima, Normanby).
Subspecies and Distribution. M.r.rufescensAlston,1877—NewBritainandNewIrelandIs. M.r.calidiorThomas,1911—lowlandsofSWNewGuinea. M.r.gracilisThomas,1906—SescarpmentoftheCentralCordilleraofNewGuineafromMtBosaviEtotheOwenStanleyRange. M.r.hageniTroughton,1937—mountainsofNewGuineafromCWestPapua(=IrianJaya])EtotheKratkeRange. M.r.niviventerTate,1951—lowerDigulandFlyriverbasins,SCNewGuinea. M.r.stalkeriThomas,1904—lowlandsofN&SENewGuinea. M. r. wisselensis Menzies, 1996 — known only from the Wissel Lakes area ofW New Guinea. Known also from islands of Waigeo, Salawati, Yapen, Blup Blup, Karkar, and Sideia (off NW, N & SE New Guinea), but subspecies involved not identified. in Muridae
Subspecies and Distribution. M.r.rufescensAlston,1877—NewBritainandNewIrelandIs. M.r.calidiorThomas,1911—lowlandsofSWNewGuinea. M.r.gracilisThomas,1906—SescarpmentoftheCentralCordilleraofNewGuineafromMtBosaviEtotheOwenStanleyRange. M.r.hageniTroughton,1937—mountainsofNewGuineafromCWestPapua(=IrianJaya])EtotheKratkeRange. M.r.niviventerTate,1951—lowerDigulandFlyriverbasins,SCNewGuinea. M.r.stalkeriThomas,1904—lowlandsofN&SENewGuinea. M. r. wisselensis Menzies, 1996 — known only from the Wissel Lakes area ofW New Guinea. Known also from islands of Waigeo, Salawati, Yapen, Blup Blup, Karkar, and Sideia (off NW, N & SE New Guinea), but subspecies involved not identified.
Efficient weighting methods for genomic best linear unbiased prediction (BLUP) adaption to the genetic architectures of quantitative traits
Open the record for dataset details and reuse information.
Data from: Genomic BLUP decoded: a look into the black box of genomic prediction
Genomic best linear unbiased prediction (BLUP) is a statistical method that uses relationships between individuals calculated from single-nucleotide polymorphisms (SNPs) to capture relationships at quantitative trait loci (QTL). We show that genomic BLUP exploits not only linkage disequilibrium (LD) and additive-genetic relationships, but also cosegregation to capture relationships at QTL. Simulations were used to study the contributions of those types of information to accuracy of genomic estimated breeding values (GEBVs), their persistence over generations without retraining, and their effect on the correlation of GEBVs within families. We show that accuracy of GEBVs based on additive-genetic relationships can decline with increasing training data size and speculate that modeling polygenic effects via pedigree relationships jointly with genomic breeding values using Bayesian methods may prevent that decline. Cosegregation information from half sibs contributes little to accuracy of GEBVs in current dairy cattle breeding schemes but from full sibs it contributes considerably to accuracy within family in corn breeding. Cosegregation information also declines with increasing training data size, and its persistence over generations is lower than that of LD, suggesting the need to model LD and cosegregation explicitly. The correlation between GEBVs within families depends largely on additive-genetic relationship information, which is determined by the effective number of SNPs and training data size. As genomic BLUP cannot capture short-range LD information well, we recommend Bayesian methods with t-distributed priors.
Data from: Genomic BLUP decoded: a look into the black box of genomic prediction
Open the record for dataset details and reuse information.
Subspecies and Distribution. E. k. kalubu Fischer, 1829 — W Papuan Is (Waigeo, Salawati, Misool), and most of New Guinea, including Yapen I and islands NE of mainland (Bagabag, Blup Blup, Kadovar, Karkar, Koil, Vokeo). E. k. cockerelli Ramsay, 1877 — New Britain, Manus, and adjacent islands, in Bismarck Archipelago. E. k. oriomo Tate & Archbold, 1936 — Fly River region, in S New Guinea. E. k. plaulipi Troughton, 1945 — Biak-Supiori and Owi I, in Cenderawasih (= Geelvink) Bay. in Peramelidae
Subspecies and Distribution. E. k. kalubu Fischer, 1829 — W Papuan Is (Waigeo, Salawati, Misool), and most of New Guinea, including Yapen I and islands NE of mainland (Bagabag, Blup Blup, Kadovar, Karkar, Koil, Vokeo). E. k. cockerelli Ramsay, 1877 — New Britain, Manus, and adjacent islands, in Bismarck Archipelago. E. k. oriomo Tate & Archbold, 1936 — Fly River region, in S New Guinea. E. k. plaulipi Troughton, 1945 — Biak-Supiori and Owi I, in Cenderawasih (= Geelvink) Bay.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.