Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
131
datasets available to search
ShareScore release 0.7.1
Dataset results
131 results for “eQTL”
Extended data: Impact of admixture and ancestry on eQTL analysis and GWAS colocalization in GTEx
<p>eQTL summary statistics and GWAS colocalization posterior probabilities from eQTL calling in an admixed subcohort of GTEx v8 with local and global ancestry adjustments. For the original, non-peer-reviewed preprint, see <a href="https://www.biorxiv.org/content/10.1101/836825v1">https://www.biorxiv.org/content/10.1101/836825v1</a>. </p> <p>For the related source code, see <a href="https://doi.org/10.5281/zenodo.3924788">https://doi.org/10.5281/zenodo.3924788</a> or <a href="https://github.com/nicolerg/gtex-admixture-la">https://github.com/nicolerg/gtex-admixture-la</a>. </p>
Summary statistics of eQTLs obtained from single-nuclei RNA-seq in 8 major brain cell-types for mendelian randomisation
<p>This dataset contains <em>cis</em>-eQTL summary statistics for 8 brain cell-types, generated on a snRNA-seq dataset on post-mortem brains from 391 individuals (full), as well as a controls-only (subset of full, 183 individuals).</p> <p>Genotype dosage matrices were obtained with <code>SeqArray</code>, where 0 = homozygous alt, 1 = heterozygous, 2 = homozygous ref https://bioconductor.org/packages/release/bioc/manuals/SeqArray/man/SeqArray.pdf . The eQTL models as applied by <code>MatrixEQTL</code> therefore use "ref" as the effect allele (as implemented in their additive model).</p> <p>The eQTL summary statistics for within the "full" and "controls-only" dataset have been packed into <code>.tar.gz</code> files for each cell-type, where unpacking will yield summary statistics by chromosome. Each file contains the following columns;</p> <p>1. <code>SNP</code> (in rsid format)</p> <p>2. <code>gene</code> (in symbol format)</p> <p>3. <code>t.stat</code> (t-statistic as determined by MatrixEQTL)</p> <p>4. <code>p.value</code> (linear model association p-value)</p> <p>5. <code>FDR</code> (false discovery rate as determined by MatrixEQTL)</p> <p>6. <code>beta</code> (effect size / slope of the linear model)</p> <p>7. <code>chrom</code> (chromosome in "chrN" format)</p> <p>8. <code>position</code> (SNP position, hg38 build)</p> <p>9. <code>effect_allele</code> (this is the "ref" allele as described above)</p> <p>10. <code>other_allele</code> (alternate allele)</p> <p>11. <code>maf</code> (minor allele frequency, as determined by SeqArray on this dataset)</p> <p> </p> <p>In addition, single-cell expression matrices in count format are available in the <code>single-cell_data.tar</code> archive for the full 391 individuals (2,348,438 cells). This archive contains processed single-cell counts as described in our manuscript for the 4 datasets included; "BRYOIS_192" (separated into "MS" and "AD" as per their publication), "MATTHEWS", "ROCHE_MPD92" and "MRC_60". In addition, a cell-level metadata file containing covariates and cell-type labels across all datasets is included (<code>cell_level_metadata.rds</code>). Cell barcodes and individual IDs have been renamed to preserve anonymity.</p> <p> </p> <p><strong>January 2025 update: </strong>Now published at <strong><em>Nature Genetics</em></strong>. <strong>https://www.nature.com/articles/s41588-024-02050-9</strong></p> <p><strong>May 2025 update: </strong>Added the aggregated pseudo "Bulk" eQTLs as seen in Fig 1. d. </p>
INTERVAL eQTL & sQTL summary statistics
<p>Summary statistics for eQTL & sQTL traits computed from the INTERVAL RNA-seq samples (n=4,372) as detailed in the following preprint: <a href="https://www.medrxiv.org/content/10.1101/2023.11.25.23299014v1">https://www.medrxiv.org/content/10.1101/2023.11.25.23299014v1</a></p>
Cell-type-specific cis-eQTLs data in SMR format from https://doi.org/10.1038/s41593-022-01128-z
<p>Cell-type-specific cis-eQTLs data in SMR format from https://doi.org/10.1038/s41593-022-01128-z</p> <p>The hg38 version of the coordinates is used</p>
Unveiling genetic signatures of immune response in immune-related diseases through single-cell eQTL analysis across diverse conditions
<p><em><strong>Unveiling genetic signatures of immune response in immune-related diseases through single-cell eQTL analysis across diverse conditions</strong></em></p> <p> </p> <p>Tools and scripts were used to generate results in Zhang et al 2024.</p> <p>Supplementary files that were not included in the initial submission.</p> <p>Full summary statistics of eQTLs including top eQTLs and all SNP-gene pairs of each cell type, consistent.tar.gz for consistent eQTLs per cell and response.tar.gz for response eQTLs, respectively.</p>
eQTL Catalogue mapping from unique variant ids to rsIDs
<p>eQTL Catalogue mapping from unique variant ids to rsIDs</p>
Single-cell eQTL mapping in yeast reveals a tradeoff between growth and reproduction
<p>Expression quantitative trait loci (eQTLs) provide a key bridge between noncoding DNA sequence variants and organismal traits. The effects of eQTLs can differ among tissues, cell types, and cellular states, but these differences are obscured by gene expression measurements in bulk populations. We developed a one-pot approach to map eQTLs in <em>Saccharomyces cerevisiae</em> by single-cell RNA sequencing (scRNA-seq) and applied it to over 100,000 single cells from three crosses. We used scRNA-seq data to genotype each cell, measure gene expression, and classify the cells by cell-cycle stage. We mapped thousands of local and distant eQTLs and identified interactions between eQTL effects and cell-cycle stages. We took advantage of single-cell expression information to identify hundreds of genes with allele-specific effects on expression noise. We used cell-cycle stage classification to map 20 loci that influence cell-cycle progression. One of these loci influenced the expression of genes involved in the mating response. We showed that the effects of this locus arise from a common variant (W82R) in the gene <em>GPA1</em>, which encodes a signaling protein that negatively regulates the mating pathway. The 82R allele increases mating efficiency at the cost of slower cell-cycle progression and is associated with a higher rate of outcrossing in nature. Our results provide a more granular picture of the effects of genetic variants on gene expression and downstream traits.</p>
Cross-tissue eQTL enrichment of associations in schizophrenia
<p>non-CNS data used in the study.</p> <p>the original data these were generated from have been deposited in the European Genome-phenome Archive (EGA) under accession EGAS00001000805.</p> <p>complete files (all associations):<br> org/blood_cis_all.gz (whole blood)<br> org/fat_cis_all.gz (adipose)<br> org/lcl_all_cis.gz (LCLs)<br> org/skin_cis_all.gz (epidermal)</p> <p>e-gene files (best association per gene at FDR < 0.01):<br> org/eQTL.B.Peer50.FDR01 (whole blood)<br> org/eQTL.F.Peer50.FDR01 (adipose)<br> org/eQTL.L.Peer50.FDR01 (LCLs)<br> org/eQTL.S.Peer50.FDR01 (epidermal)</p> <p>control variants (2m template) bed files:<br> fakeQTL_tw_blood.bed<br> fakeQTL_tw_fat.bed<br> fakeQTL_tw_lcl.bed<br> fakeQTL_tw_skin.bed</p> <p>control variants (2m template) rs names:<br> fakeQTL_tw_blood.mrk<br> fakeQTL_tw_fat.mrk<br> fakeQTL_tw_lcl.mrk<br> fakeQTL_tw_skin.mrk</p> <p>control variants (9m template) R data files:<br> fakeQTL9m_tw_blood.RData<br> fakeQTL9m_tw_fat.RData<br> fakeQTL9m_tw_lcl.RData<br> fakeQTL9m_tw_skin.RData</p> <p>control variants (9m template) matlab files:<br> fakeQTL9m_tw_blood.mat<br> fakeQTL9m_tw_fat.mat<br> fakeQTL9m_tw_lcl.mat<br> fakeQTL9m_tw_skin.mat</p> <p> </p> <p>if you use these data, please cite the data sources:</p> <p>Genetic interactions affecting human gene expression identified by variance association mapping. Brown, A.A., Buil, A., Viñuela, A., Lappalainen, T., Zheng, H.-F., Richards, J.B., Small, K.S., Spector, T.D., Dermitzakis, E.T. and Durbin, R. (2014). eLife 2014;3:e01381 doi: 10.7554/eLife.01381</p> <p>Gene-gene and gene-environment interactions detected by transcriptome sequence analysis in twins. Buil, A., Brown, A.A., Lappalainen, T., Vinuela, A., Davies, M.N., Zheng, H.-F., Richards, J.B., Glass, D., Small, K.S., Durbin, R. et al. (2015) Nat Genet, 47, 88-91.</p> <p> </p>
Reference annotations for the eQTL Catalogue
<p>Reference annotations for the eQTL Catalogue</p>
eQTL's in Prakriti replicated SNPs
<p>List of<em> Prakriti </em>replicated SNPs acting as eQTL in GTEx tissues</p>
The pan-gene, pan-genome, and SV-eQTL dataset for A2 and AD1 cotton.
Open the record for dataset details and reuse information.
Mixed model-based deconvolution of cell-state abundances along a one-dimensional trajectory [csd-eQTL]
<p><strong>README:</strong></p> <p>The full summary data of the cell-state-dependent eQTLs for GTEx Esophagus Mucosa (n=497) are stored in the .parquet format.</p> <p>An example of the file name:</p> <p><strong>"GTEx_Esophagus_Mucosa_bin1.cis_qtl_pairs.1.parquet.gz"</strong> means the summary data of csd-eQTLs for bin1 of chromosome 1.</p>
Results for eQTL analysis for each brain region
<p>This dataset is part of the manuscript: "<em>Atlas of genetic effects in human microglia transcriptome across brain regions, aging and disease pathologies</em>", by Lopes KP, Snijders GJL, Humphrey J, et al.</p> <p> </p> <p>Description of files:</p> <p><em>MFG_eur_expression_peer10.cis_qtl_nominal.txt.gz - </em>Full <strong>nominal eQTL</strong> summary statistics from medial frontal gyrus (<strong>MFG</strong>)<em> </em>(gzip-compressed)</p> <p><em>STG_eur_expression_peer10.cis_qtl_nominal.txt.gz </em>- Full <strong>nominal eQTL</strong> summary statistics from superior temporal gyrus (<strong>STG</strong>) (gzip-compressed)</p> <p><em>SVZ_eur_expression_peer5.cis_qtl_nominal.txt.gz - </em>Full <strong>nominal eQTL</strong> summary statistics from subventricular zone (<strong>SVZ</strong>)<em> </em>(gzip-compressed)</p> <p><em>THA_eur_expression_peer10.cis_qtl_nominal.txt.gz - </em>Full <strong>nominal eQTL</strong> summary statistics from thalamus (<strong>THA</strong>)<em> </em>(gzip-compressed)</p> <p><em>MFG_eur_expression_peer10.cis_qtl.txt.gz </em>- Full <strong>permuted eQTL</strong> summary statistics from medial frontal gyrus (<strong>MFG</strong>) (gzip-compressed)</p> <p><em>STG_eur_expression_peer10.cis_qtl.txt.gz -</em> Full <strong>permuted eQTL</strong> summary statistics from superior temporal gyrus (<strong>STG</strong>) (gzip-compressed)</p> <p><em>SVZ_eur_expression_peer5.cis_qtl.txt.gz - </em>Full <strong>permuted eQTL</strong> summary statistics from subventricular zone (<strong>SVZ</strong>) (gzip-compressed)</p> <p><em>THA_eur_expression_peer10.cis_qtl.txt.gz - </em>Full <strong>permuted eQTL</strong> summary statistics from thalamus (<strong>THA</strong>) (gzip-compressed)</p> <p> </p> <p>Nominal QTL results include all SNP-gene pairs tested (using a 1Mb window from each side of the transcription start site (TSS) of a gene). Table columns are formatted as follows:</p> <ul> <li>phenotype_id - ensembl ID of the gene tested (GENCODE v30)</li> <li>variant_id - SNP tested for association (rsid or chr:position:ref:alt)</li> <li>tss_distance - distance of the SNP to the gene transcription start site (TSS)</li> <li>maf - minor allele frequency in MiGA cohort</li> <li>ma_samples - number of samples carrying the minor allele</li> <li>ma_count - total number of minor alleles across individuals</li> <li>pval_nominal - nominal <em>P</em>-value from linear regression</li> <li>slope - slope of the linear regression</li> <li>slope_se - standard error of the slope</li> </ul> <p>Permuted QTL results include only the top SNP-gene association for each gene. Table columns are formatted as follows:</p> <ul> <li>phenotype_id - ensembl ID of the gene tested (GENCODE v30)</li> <li>num_var - total number of variants tested in <em>cis</em></li> <li>beta_shape1 - first parameter value of the fitted beta distribution</li> <li>beta_shape2 - second parameter value of the fitted beta distribution</li> <li>true_df - effective degrees of freedom the beta distribution approximation</li> <li>pval_true_df - empirical <em>P</em>-value for the beta distribution approximation</li> <li>variant_id - ID of the top variant (rsid or chr:position:ref:alt)</li> <li>tss_distance - distance of the SNP to the gene transcription start site (TSS)</li> <li>ma_samples - number of samples carrying the minor allele</li> <li>ma_count - total number of minor alleles across individuals</li> <li>maf -minor allele frequency in MiGA cohort</li> <li>ref_factor - flag indicating if the alternative allele is the minor allele in the cohort (1 if AF <= 0.5, -1 if not)</li> <li>pval_nominal - nominal <em>P</em>-value from linear regression</li> <li>slope - slope of the linear regression</li> <li>slope_se - standard error of the slope</li> <li>pval_perm - first permutation <em>P</em>-value directly obtained from the permutations with the direct method</li> <li>pval_beta - second permutation <em>P</em>-value obtained via beta approximation. This is the one to use for downstream analysis</li> <li>qval - Storey q-value derived from pval_beta (FDR adjusted)</li> <li>pval_nominal_threshold - nominal <em>P</em>-value threshold for calling a variant-gene pair significant for the gene</li> </ul> <p><strong>NOTE: </strong>The effect sizes of eQTLs and sQTL are defined as the effect of the alternative allele (ALT) relative to the reference (REF) allele in the human genome reference (GRCh38). A file containing that information for all alleles tested is available at 10.5281/zenodo.4301005</p>
Data from: Mediation analysis demonstrates that trans-eQTLs are often explained by cis-mediation: a genome-wide analysis among 1,800 South Asians
A large fraction of human genes are regulated by genetic variation near the transcribed sequence (cis-eQTL, expression quantitative trait locus), and many cis-eQTLs have implications for human disease. Less is known regarding the effects of genetic variation on expression of distant genes (trans-eQTLs) and their biological mechanisms. In this work, we use genome-wide data on SNPs and array-based expression measures from mononuclear cells obtained from a population-based cohort of 1,799 Bangladeshi individuals to characterize cis- and trans-eQTLs and determine if observed trans-eQTL associations are mediated by expression of transcripts in cis with the SNPs showing trans-association, using Sobel tests of mediation. We observed 434 independent trans-eQTL associations at a false-discovery rate of 0.05, and 189 of these trans-eQTLs were also cis-eQTLs (enrichment P<0.0001). Among these 189 trans-eQTL associations, 39 were significantly attenuated after adjusting for a cis-mediator based on Sobel P<10-5. We attempted to replicate 21 of these mediation signals in two European cohorts, and while only 7 trans-eQTL associations were present in one or both cohorts, 6 showed evidence of cis-mediation. Analyses of simulated data show that complete mediation will be observed as partial mediation in the presence of mediator measurement error or imperfect LD between measured and causal variants. Our data demonstrates that trans-associations can become significantly stronger or switch directions after adjusting for a potential mediator. Using simulated data, we demonstrate that this phenomenon is expected in the presence of strong cis-trans confounding and when the measured cis-transcript is correlated with the true (unmeasured) mediator. In conclusion, by applying mediation analysis to eQTL data, we show that a substantial fraction of observed trans-eQTL associations can be explained by cis-mediation. Future studies should focus on understanding the mechanisms underlying widespread cis-mediation and their relevance to disease biology, as well as using mediation analysis to improve eQTL discovery.
Data from: Hypothalamic transcriptomes of 99 mouse strains reveal trans eQTL hotspots, splicing QTLs and novel non-coding genes
Previous studies had shown that integration of genome wide expression profiles, in metabolic tissues, with genetic and phenotypic variance, provided valuable insight into the underlying molecular mechanisms. We used RNA-Seq to characterize hypothalamic transcriptome in 99 inbred strains of mice from the Hybrid Mouse Diversity Panel (HMDP), a reference resource population for cardiovascular and metabolic traits. We report numerous novel transcripts supported by proteomic analyses, as well as novel non coding RNAs. High resolution genetic mapping of transcript levels in HMDP, reveals both local and trans expression Quantitative Trait Loci (eQTLs) demonstrating 2 trans eQTL 'hotspots' associated with expression of hundreds of genes. We also report thousands of alternative splicing events regulated by genetic variants. Finally, comparison with about 150 metabolic and cardiovascular traits revealed many highly significant associations. Our data provides a rich resource for understanding the many physiologic functions mediated by the hypothalamus and their genetic regulation.
Significant eQTLs and expression TWAS reference panels (AMP-AD brain and EADB Belgian LCL cohorts)
<p>This dataset is part of the manuscript "<em><strong>New insights into the genetic etiology of Alzheimer’s disease and related dementias</strong></em>" by Bellenguez, Küçükali, et al. Nature Genetics 2022.</p> <p>Publication link: <a href="https://www.nature.com/articles/s41588-022-01024-z">https://www.nature.com/articles/s41588-022-01024-z</a></p> <p>GitHub repository of all QTL/TWAS data shared for this study: <a href="https://github.com/SleegersLab-VIBCMN/EADB_GWAS_NatureGenetics_QTL_TWAS">https://github.com/SleegersLab-VIBCMN/EADB_GWAS_NatureGenetics_QTL_TWAS</a></p> <p>For details, please see the publication. For any questions, please contact Fahri Küçükali (<a href="mailto:fahri.kucukali@uantwerpen.vib.be">fahri.kucukali@uantwerpen.vib.be</a>) and Kristel Sleegers (<a href="mailto:Kristel.Sleegers@uantwerpen.vib.be">Kristel.Sleegers@uantwerpen.vib.be</a>).</p> <p>Significant eQTL catalogues are compressed with <em>gzip </em>and tar achieve of expression TWAS reference panels are compressed with <em>bzip2</em>.</p> <p><strong>eQTL catalogues</strong></p> <p>The files show significant eQTL - gene pairs mapped in AMP-AD brain and EADB Belgian LCL cohorts. The catalogues are in hg38/GRCh38 human genome build. Most of the columns in the files are based on FastQTL output (<a href="http://fastqtl.sourceforge.net/">http://fastqtl.sourceforge.net/</a>).</p> <p><em>eQTL file columns:</em></p> <ol> <li>variant_id - ID of the significant eQTL variant based on the dbSNPv151 rsID annotation or hg38/GRCh38 CHR_POS_REF_ALT ID if rsID not available.</li> <li>gene_id - eGene ENSG gene ID based on GENCODEv24 (AMP-AD) or GENCODEv32 (EADB Belgian)</li> <li>tss_distance - Genomic distance between eQTL variant and eGene</li> <li>ma_samples - Number of samples carrying the minor allele</li> <li>ma_count - Total count of minor alleles</li> <li>maf - Minor allele frequency</li> <li>pval_nominal - Nominal <em>P</em>-value of the association</li> <li>slope - Slope of the association with respect to alternative (ALT) allele indicated on column 14</li> <li>slope_se - Standard error of the slope</li> <li>pval_nominal_threshold - Nominal <em>P</em>-value significant threshold for this eGene</li> <li>min_pval_nominal - Most significant <em>P</em>-value observed for this eGene</li> <li>pval_beta - permutation <em>P</em>-value obtained via beta approximation and later used to calculate Storey q-values</li> <li>gene_name - eGene name based on GENCODEv24 (AMP-AD) and GENCODEv32 (EADB Belgian)</li> <li>GRCh38_chr_pos - Genomic position of the variant, separated by underscore</li> <li>ref_alt - Reference (REF) and alternative (ALT) allele of the variant, separated by ">" sign. ALT is the tested (A1) allele</li> </ol> <p>Of note, we also mapped the significant sQTLs in the same datasets (please see the data availability section of the manuscript or the GitHub repository).</p> <p><strong>Expression TWAS reference panels</strong></p> <p>Custom expression TWAS reference panels prepared using FUSION pipeline (<a href="http://gusevlab.org/projects/fusion/">http://gusevlab.org/projects/fusion/</a>) in AMP-AD brain and EADB Belgian LCL cohorts. All data in hg38/GRCh38 genome build. In each directory, you will find ".pos", ."profile", and ".profile.err" files. These are explained in the FUSION website as well, but briefly these are:</p> <ol> <li><strong>.pos:</strong> This is a position file that describes the 1Mb extended gene start and end coordinates for each calculated weight file for each gene expression phenotype. Used for scanning the variants in those coordinates for TWAS.</li> <li><strong>.profile: </strong>This informs about all prediction weights calculated, in terms of number of variants in the model, heritability information, and R2 info for each prediction model used (top1, blup, enet, bslmm, lasso; bslmm was not used therefore has NA values).</li> <li><strong>.profile.err: </strong>This summarizes the reference panel in terms of average hsq (with SD), and which model is the best performing.</li> </ol> <p>Each TWAS weight is provided in a .RDat file under <strong>All_Expression_Weights</strong>, and in this data we included all calculated functional weights independent of the fact that they are heritable features or not. In our TWAS analyses, we included the heritable functional weights at a hsq <em>P</em>-value ≤ 0.05 level.</p> <p>Please also see the splicing TWAS reference panels we prepared in the same datasets (see the data availability section of the manuscript or the GitHub repository). If you need an LD reference data in hg38/GRCh38 genome build based on 1000 Genomes NFE samples (whose variant ID annotation are matching to these functional weights), suitable for running the TWAS/FUSION pipeline, please contact us.</p>
Data from: Sequence-based association analysis reveals an MGST1 eQTL with pleiotropic effects on bovine milk composition
[No abstract entered]
BovReg_eqTL RNAseq demo input data as counts and aligned bam files
Open the record for dataset details and reuse information.
Summary statistics for the InsPIRE consortium (pancreatic islets eQTLs)
<p>Summary statistics for the InsPIRE study: Influence of genetic variants on gene expression in human pancreatic islets – implications for type 2 diabetes</p> <p>This datasets includes eQTLs from 420 pancreatic islets (exon and gene level quantifications) and 27 beta-cells FAC sorted (exon quantifications) form RNA-Seq samples. Full summary statistics and independent associations are included.</p> <p>The full description of methods is currently available here: Bioxiv (<a href="https://www.biorxiv.org/content/10.1101/655670v1">https://www.biorxiv.org/content/10.1101/655670v1</a>).</p>
Bioinformatics workflow for the detection of eQTL in the cattle genome using Nextflow DSL2
<p>The <em>in silico</em> detection of expression quantitative trait loci (eQTL) demands high throughput processing from hundreds of samples, which is often a challenge to handle and run such large datasets. In order to focus on the core analysis, it is convenient to have simple coding and hassle-free installation of different software tools required for the bioinformatics workflow. In this context, the newly available technologies like workflow managers and software containers enabled to develop workflows with less complexity. In this study, we developed an eQTL bioinformatics pipeline with the workflow manager Nextflow and docker container software, for coding and installing the required software tools. This workflow can be portable to a different computer environment, and the results are reproducible. We tested the functionality of our workflow with a sample dataset and the runtime estimates from this demo run will provide important information in planning future analyses with much larger datasets.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.