Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,243
datasets available to search
ShareScore release 0.7.1
Dataset results
1,243 results for “Statistics”
Full summary statistics of mixQTL for GTEx v8 Minor_Salivary_Gland
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Breast_Mammary_Tissue
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Adipose_Visceral_Omentum
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Artery_Aorta
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Nerve_Tibial
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Pancreas
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Adrenal_Gland
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Brain_Cerebellar_Hemisphere
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Cells_EBV-transformed_lymphocytes
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Brain_Spinal_cord_cervical_c-1
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Skin_Sun_Exposed_Lower_leg
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Thyroid
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Stomach
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Whole_Blood
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Full summary statistics of mixQTL for GTEx v8 Uterus
The mixQTL method is described in paper doi.org/10.1101/2020.04.22.050666. Please cite the original paper if using the data.
Genome-wide association summary statistics for sex- and age-specific analysis of chronic back pain
<p>The dataset comprises summary-level statistics for age- and sex-specific genome-wide association study of chronic back pain (cBP) in individuals of European descent from UK Biobank (<a href="https://www.ukbiobank.ac.uk/">https://www.ukbiobank.ac.uk/</a>). The study was carried out under UK Biobank approved project #18219. </p> <p><strong>The dataset accompanies the paper (please cite if using the dataset):</strong></p> <p><a href="https://pubmed.ncbi.nlm.nih.gov/33021770/">Freidin, Maxim B.; Tsepilov, Yakov A.; Stanaway, Ian B.; Meng, Weihua; Hayward, Caroline; Smith, Blair H.; Khoury, Samar; Parisien, Marc; Bortsov, Andrey; Diatchenko, Luda; Børte, Sigrid; Winsvold, Bendik S.; Brumpton, Ben M.; Zwart, John-Anker; HUNT All-In Pain; Aulchenko, Yurii S.; Suri, Pradeep; Williams, Frances M.K. Sex- and age-specific genetic analysis of chronic back pain. Pain. 2020. doi:10.1097/j.pain.0000000000002100.</a></p> <p>The phenotype of cBP was defined as back pain for 3+ months. Linear mixed-effects additive model was fitted adjusting for age, genotyping array type, and 10 genetic PCs provided by UK Biobank. The following filters were applied: minor allele frequency >0.001, genotyping and individual call rates >0.98%, imputation quality score (INFO) >0.7. GWAS were carried out in males and females separately in the whole sample (<strong>allages</strong>) as well as in groups of younger than 65 years (<strong>under65</strong>) and 65+ years old (<strong>65plus</strong>) as detailed in the paper. Accordingly, 6 files are deposited here, corresponding to each group. </p> <p><strong>Column headers:</strong></p> <p>SNP, SNP rsID </p> <p>CHR, chromosome</p> <p>BP, genomic position (GRCh37 build)</p> <p>EA, effect allele (coded as "1")</p> <p>OTHER, other allele (coded as "0")</p> <p>A1FREQ, frequency of effect allele</p> <p>INFO, imputation quality</p> <p>BETA, effect size (for effect allele)</p> <p>SE, standard error of effect size</p> <p>PVAL, p-value for association</p>
Dsuite - fast D-statistics and related admixture evidence from VCF files
<p>Patterson's D, also known as the ABBA-BABA statistic, and related statistics such as the f4-ratio, are commonly used to assess evidence of gene flow between populations or closely related species. Currently available implementations often require custom file formats, implement only small subsets of the available statistics, and are impractical to evaluate all gene flow hypotheses across datasets with many populations or species due to computational inefficiencies. Here we present a new software package Dsuite, an efficient implementation allowing genome scale calculations of the D and f4-ratio statistics across all combinations of tens or hundreds of populations or species directly from a variant call format (VCF) file. Our program also implements statistics suited for application to genomic windows, providing evidence of whether introgression is confined to specific loci and it can also aid in interpretation of a system of f4-ratio results with the use of the 'f-branch' method. Dsuite is available at https://github.com/millanek/Dsuite, is straightforward to use, substantially more computationally efficient than comparable programs, and provides a convenient suite of tools and statistics, including some not previously available in any software package. Thus, Dsuite facilitates the assessment of evidence for gene flow, especially across larger genomic datasets.</p>
G2G-EBV GWAS summary statistics
<p>G2G results from manuscript titled:</p> <p><strong>The influence of human genetic variation on Epstein-Barr virus sequence diversity</strong></p> <p>The GWAS result files (*.mlma) contains:<br> chromosome, SNP, physical position, reference allele (the coded effect allele), the other allele, frequency of the reference allele, SNP effect, standard error and p-value (<a href="https://cnsgenomics.com/software/gcta/#MLMA">GCTA-MLMA</a>).<br> <br> The scripts used to generate the GWAS results, are available here: <a href="https://github.com/sinarueeger/G2G-EBV-manuscript">github.com/sinarueeger/G2G-EBV-manuscript</a>.</p>
Training dataset: Statistical analysis of a HEK/Ecoli Spike-in DIA dataset using MSstats
<p>The uploaded files serve as a concise but meaningful training data set in the Galaxy training network (https://galaxyproject.github.io/training-material/).</p> <p>HEK and E.coli cell pellets were lysed with 5 % SDS, 50 mM triethylammonium bicarbonate (TEAB), pH 7.55. The obtained protein extracts were reduced by adding f.c. 5 mM TCEP and alkylated by the addition of f.c. 10 mM iodacetamide. Protein digestion and purification was performed on S-Trap columns. To ensure protein binding to the S-Trap columns, samples were acidified to a final concentration of 1.2 % phosphoric acid (~ pH 2). Six times the sample volume S-Trap buffer (90% aqueous methanol containing a final concentration of 100 mM TEAB, pH 7.1) was added to the samples which were then loaded on the columns and washed with S-Trap buffer. Protein digestion was performed with trypsin and LysC for one hour at 47 °C. Peptides were eluted in three steps with (1) 50 mM TEAB, (2) 0.2 % aqueous formic acid and (3) 50 % acetonitrile containing 0.2 % formic acid. Eluted peptides of HEK and E.coli were mixed in two different ratios and four replicates of each Spike/in ratio were measured and analysed using OpenSwathWorkflow in Galaxy. Results were exported using PyProphet and can be used for the statistical analysis and detection of the two different Spike-in Ratios. The Spike-in ratios were the following:</p> <p>Sample HEK E.coli <br> Spike_in_1 2.5 0.15<br> Spike_in_2 2.5 0.80 </p> <p>Besides the two PyProphet export files, we uploaded a sample annotation file as well as a comparison matrix file.<br> Additionally, we uploaded the Galaxy MSstats training result files: MSstats_ComparisonResult_export_tabular and MSstats_ComparisonResult_msstats_input.</p>
Dataset for "Topography-based statistical modelling reveals high spatial variability and seasonal emission patches in forest floor methane flux"
<p>This dataset provides measured and upscaled forest floor methane (CH4) fluxes and soil moisture.</p> <p>This dataset is related to the following manuscript:</p> <p>Vainio et al., Topography-based statistical modelling reveals high spatial variability and seasonal emission patches in forest floor methane flux, Biogeosciences, in review. (The discussion preprint is available at https://doi.org/10.5194/bg-2020-263.)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.