Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

15

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

15 results for “1000 Genomes Project”

Learn how ShareScore rates datasets ↗
zenodo44/100

Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project

SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.

openmit-licenseApr 2024View details →
zenodo44/100

1000 Genomes Project Cleaned Dataset

<p>The first four authors performed standard quality control analysis on the 1000 Genomes (1KG) Project genotypes that were generated on the Illumina Omni2.5M chip, at the Broad and Sanger Institutes. The datasets were then posted on the website of The Centre for Applied Genomics at Sick Kids Hospital at https://www.tcag.ca/tools/1000genomes.html. The last two authors then looked for overlap between those datasets and the Hapmap3 datasets that had gene expression for Endoplasmic Reticulum Aminopeptidase 2 (ERAP2), and chose the Yoruban from Ibadan, Nigeria (YRI) and Utah residents with Northern and Western European ancestry (CEU) subpopulations. These two subpopulations had the largest overlap between the 1KG and HapMap3 datasets, with 91 YRI and 104 CEU samples. The text files provided in this repository contain the IDs of all invidividuals and&nbsp;phenotypes for the labelled populations e.g. ERAP2_CEU_YRI_phenotypes.txt has phenotypes for both populations. The two *_pc_outliers.txt contain the IDs of the individuals excluded from analysis due to extraneous principal components. In summary, 88 YRI and 102 CEU individuals were included in the analysis.&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Relate-estimated coalescence rates, allele ages, and selection p-values for the 1000 Genomes Project

<p><strong>Overview</strong></p> <p>Coalescence rates, allele ages, and p-values for evidence of positive selection calculated for 2478&nbsp;samples of the&nbsp;1000 Genomes Project&nbsp;using Relate.</p> <p>We estimated the joint genealogy of all 1000 GP populations and then extracted the embedded genealogy for each population.<br> For the genealogy of each population, we jointly estimated the population size history and branch lengths.&nbsp;<br> Variants segregating in more than one&nbsp;population&nbsp;therefore have&nbsp;correlated but different allele ages in each population.</p> <p>Please refer to&nbsp;<a href="https://www.nature.com/articles/s41588-019-0484-x">Speidel et al.&nbsp;Nature Genetics (2019)</a>&nbsp;for more details or email leo.speidel@outlook.com for any queries.</p> <p><strong>Coalescence rates</strong></p> <p>The zipped directory&nbsp;coalescence_rates.zip&nbsp;contains coalescence rates for 26 populations in the 1000 Genomes Project data set.</p> <ul> <li>The .coal files show the haploid coalescence rates, please refer to the&nbsp;<a href="https://myersgroup.github.io/relate/modules.html#PopulationSizeScript_FileFormats">Relate documentation</a>&nbsp;for the file format.</li> <li>The popsize.RData file is an R data frame storing the diploid population sizes (0.5/coalescence rate) calculated using the .coal files. The columns of this data frame, named &quot;pop_size&quot;,&nbsp;are <ul> <li>gens_ago: Time in generations at which epoch starts. (To get years from generations, we multiply by 28.)</li> <li>population_size: Diploid population size in this epoch.</li> <li>population: Name of population&nbsp;</li> <li>region: Name of region (AFR, AMR, EAS, EUR, SAS)</li> </ul> </li> </ul> <p><strong>Allele ages and selection p-values</strong></p> <p>The zipped directories&nbsp;allele_ages_*.zip&nbsp;contain&nbsp;R&nbsp;data frames for each 1000GP population storing allele ages and selection p-values.<br> Please note that only mutations that segregate in the population and map to a unique branch in the Relate-estimated marginal trees are included. Selection p-values are only provided for mutations of DAF &gt; 2 that pass quality filters (see Speidel et al., 2019).&nbsp;</p> <p>To get an age estimate for a neutral mutation, use&nbsp;0.5*(lower_age + upper_age). To get years from generations, we multiply by 28.</p> <p>The columns of these&nbsp;data frames, named &quot;allele_ages&quot;,&nbsp;are</p> <ul> <li>CHR: chromosome index</li> <li>BP: base-pair position (GRCh37)</li> <li>ID: id of SNP</li> <li>lower_age: Age in generations of coalescence event at the lower end of the branch onto which the mutation maps</li> <li>upper_age: Age in generations of coalescence event at the upper end of the branch onto which the mutation maps</li> <li>ancestral/derived: Ancestral/derived allele</li> <li>upstream: Upstream (5&#39;) allele</li> <li>downstream: Downstream (3&#39;) allele</li> <li>DAF: Derived-allele frequency</li> <li>pvalue: log10 p-value for selection evidence</li> </ul>

opencc-by-4.0May 2019View details →
zenodo44/100

A unified genealogy of modern and ancient genomes: Unified, inferred tree sequences of 1000 Genomes, Human Genome Diversity, and Simons Genome Diversity Projects

<p>Unified, inferred tree sequences built from&nbsp;the 1000 Genomes phase 3, Human Genome Diversity, and Simons Genome Diversity Projects. Each tree sequence is the arm of an autosome (the short arm of acrocentric chromosomes are not included).&nbsp;Tree sequences were inferred using&nbsp;<a href="https://tsinfer.readthedocs.io/">tsinfer</a>&nbsp;version 0.2.1,&nbsp;dated using&nbsp;<a href="https://tsdate.readthedocs.io/en/latest/">tsdate</a> version 0.1.4&nbsp;and compressed using&nbsp;<a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. All data is in GRCh38.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on&nbsp;<a href="https://github.com/awohns/unified_genealogy_paper">GitHub</a>. A description can be found in the Supplementary Material of <a href="https://www.biorxiv.org/content/10.1101/2021.02.16.431497v2">Wohns et al. (2021)</a>.</p> <p>Tree sequences can&nbsp; be decompressed as follows:</p> <pre><code>$ tsunzip hgdp_tgp_sgdp_chr1_p.dated.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed in Python using&nbsp;<a href="https://tskit.readthedocs.io/">tskit</a>.&nbsp;</p> <pre><code>import tskit ts = tskit.load("hgdp_tgp_sgdp_chr1_p.dated.trees") # ts is an instance of tskit.TreeSequence print("The short arm of chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with nodes contain&nbsp;the mean and variance of tsdate&#39;s posterior distribution on node time. To access these values, we can use:</p> <pre><code>import json node = ts.node(10000) metadata_dict = json.loads(node.metadata) print("The mean of the posterior distribution on the age of node 10000 is {} generations".format(metadata_dict["mn"])) print("The variance of the posterior distribution on the age of node 10000 is {} generations".format(metadata_dict["vr"]))</code></pre> <p>Age estimates for&nbsp;each variant site can be derived from the mean of the age estimates of the&nbsp;upper and lower bounding nodes of the oldest mutation associated with a site. tsdate includes <a href="https://tsdate.readthedocs.io/en/latest/python-api.html?highlight=sites_time_from_ts#tsdate.sites_time_from_ts">a function to find the age estimates of all sites in the tree sequence</a>:</p> <pre><code>import tsdate site_times = tsdate.sites_time_from_ts(ts, node_selection='arithmetic')</code></pre> <p>This returns a numpy array which has a length equal to the number of sites.</p> <p>Accessing variant sites in the tree sequence provides&nbsp;the position and id of variants:</p> <pre><code>site = ts.site(1000) site_metadata = json.loads(site.metadata) print("The position of site 1000 is {} and its ID is {}.".format(site.position, site_metadata["ID"]))</code></pre> <p>Metadata associated with individuals and populations was derived from the original sources (<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">TGP</a>, <a>HGDP</a>, and <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">SGDP</a>)&nbsp;and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code>ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain&nbsp;all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code>pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the&nbsp;ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

1000 Genomes Project Transposable Element database

<p>Multi-sample VCF with transposable elements across individuals in the 1KGP dataset. Transposable elements were called using RetroSeq&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Protein haplotype sequences obtained by ProHap from the 1000 Genomes Project data set

<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the 1000 Genomes Project, aligned with the GRCh38 genome build (<a href="https://www.internationalgenome.org/data-portal/data-collection/grch38">https://www.internationalgenome.org/data-portal/data-collection/grch38</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts. The complete configuration file for each ProHap run is attached to this repository.</p> <p>This data set contains six compressed directories, five representing the superpopulations included in the 1000 Genomes Project (<a href="https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project">https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project</a>), and one created using all the samples included in the 1000 Genomes data set:</p> <ul> <li>AFR - African</li> <li>AMR - American</li> <li>EUR - European</li> <li>SAS - South Asian</li> <li>EAS - East Asian</li> <li>ALL - all participants in the 1000 Genomes Project</li> </ul> <p>Each of the directories contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap, using alleles with at least 1 % frequency within the selected population</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the simplified fasta file.&nbsp;</li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to&nbsp;<a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Va&scaron;&iacute;ček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

1000 Genomes Project

<p>The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations.&nbsp; The genomes of 2,504 individuals from 26 populations were reconstructed using a combination of low-coverage whole-genome sequencing, deep exome sequencing, and dense microarray genotyping. A broad spectrum of genetic variation was characterised, in total over 88 million variants (84.7 million single nucleotide polymorphisms (SNPs), 3.6 million short insertions/deletions (indels), and 60,000 structural variants), all phased onto high-quality haplotypes. This resource includes &gt;99% of SNP variants with a frequency of &gt;1% for a variety of ancestries.</p>

opencc-by-nc-3.0Aug 2019View details →
zenodo40/100

A unified genealogy of modern and ancient genomes: Unified, inferred tree sequences of 1000 Genomes, Human Genome Diversity, and Simons Genome Diversity Projects with ancient samples

<p>Unified, inferred tree sequences built from&nbsp;the 1000 Genomes phase 3, Human Genome Diversity, and Simons Genome Diversity Projects with high coverage sequenced ancient samples. The ancient samples are the Altai, Chagyrskaya, and Vindija Neanderthals, the Denisovan, and a high-coverage family of four from the Afanasievo Culture.</p> <p>Each tree sequence is the arm of an autosome (the short arm of acrocentric chromosomes are not included).&nbsp;Tree sequences were inferred with&nbsp;<a href="https://tsinfer.readthedocs.io/">tsinfer</a>&nbsp;version 0.2.1 and&nbsp;<a href="https://tsdate.readthedocs.io/en/latest/">tsdate</a> version 0.1.4, as&nbsp;described in <a href="http://www.biorxiv.org/content/10.1101/2021.02.16.431497v2">Wohns et al. (2021)</a>. The files were&nbsp;compressed using&nbsp;<a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. All data is in GRCh38.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on&nbsp;<a href="https://github.com/awohns/unified_genealogy_paper">GitHub</a>. A description can be found in the Supplementary Material of <a href="https://www.biorxiv.org/content/10.1101/2021.02.16.431497v2">Wohns et al. (2021)</a>.</p> <p>Tree sequences can&nbsp;be decompressed as follows:</p> <pre><code>$ tsunzip hgdp_tgp_sgdp_high_cov_ancients_chr1_p.dated.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed in Python using&nbsp;<a href="https://tskit.readthedocs.io/">tskit</a>.&nbsp;</p> <pre><code>import tskit ts = tskit.load("hgdp_tgp_sgdp_high_cov_ancients_chr1_p.dated.trees") # ts is an instance of tskit.TreeSequence print("The short arm of chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Accessing variant sites in the tree sequence provides&nbsp;the position and id of variants:</p> <pre><code>import json site = ts.site(1000) site_metadata = json.loads(site.metadata) print("The position of site 1000 is {} and its ID is {}.".format(site.position, site_metadata["ID"]))</code></pre> <p>Metadata associated with individuals and populations was derived from the original sources (<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">TGP</a>, <a>HGDP</a>, and <a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">SGDP</a>)&nbsp;and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code>ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain&nbsp;all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code>pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the&nbsp;ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

SINGER-inferred targets for exceptional population differentiation in coalescence times in African populations in 1000 Genomes Project

<p>This repo saves the gene targets which shows the signal of population differentiation in coalescence times in African populations in 1000 Genomes Project.&nbsp;</p> <p>Here are the detailed explanations for the header of the files:</p> <p>chrom &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; The chromosome where the genomic window is located (e.g., chr1, chrX).<br>window_start &nbsp; &nbsp; &nbsp; &nbsp;The 0-based start position of the genomic window being analyzed.<br>window_end &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;The 1-based end position of the genomic window (exclusive).<br>score &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Ratio between the overall diversity and the population specific diversity, in the genomic window.<br>transcript_chrom &nbsp; &nbsp;The chromosome where the transcript is located.<br>transcript_start &nbsp; &nbsp;The 0-based start position of the transcript.<br>transcript_end &nbsp; &nbsp; &nbsp;The 1-based end position of the transcript (exclusive).<br>transcript_id &nbsp; &nbsp; &nbsp; The unique identifier for the transcript (e.g., Ensembl or RefSeq ID).<br>strand &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;The strand of the transcript: "+" for forward, "-" for reverse.<br>coding_start &nbsp; &nbsp; &nbsp; &nbsp;The start position of the coding region of the transcript (if applicable).<br>coding_end &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;The end position of the coding region of the transcript (if applicable).<br>RGB &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; The RGB color used for visualization in genome browsers (format: R,G,B).<br>block_count &nbsp; &nbsp; &nbsp; &nbsp; The number of exons in the transcript.<br>block_sizes &nbsp; &nbsp; &nbsp; &nbsp; Comma-separated list of exon lengths (in base pairs).<br>block_starts &nbsp; &nbsp; &nbsp; &nbsp;Comma-separated list of exon start positions relative to transcript_start.</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Fine-scale subpopulation detection via SNP-based unsupervised method: A case study on the 1000 Genomes Project Resources

<p>Here are the supplementary files to the paper &quot;Fine-scale subpopulation detection via SNP-based unsupervised method:<br> A case study on the 1000 Genomes Project Resources&quot;.<br> <br> The repository is organized as:</p> <ol> <li><strong>Supplementary information</strong>: the additional detailed information for the experiments in the paper <ul> <li>Supplementary_information_IPCAPS_Chaichoompu_v1.pdf</li> </ul> </li> <li><strong>Real-life dataset</strong>: the 1000 genome dataset, which is referred to in the paper and is filtered with the parameters explained in the paper.&nbsp;Reference:&nbsp;https://www.internationalgenome.org/ <ul> <li>1000genomes_with_filtering.zip</li> </ul> </li> </ol>

opencc-by-4.0Jul 2022View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on 1000 Genomes Project 2504 Whole-Genome Sequencing Data

<p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the 1000 Genomes Project 2504 Dataset. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on 1000 Genomes Project 2504 Dataset</strong>:&nbsp;Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000G_2504_high_coverage/1000G_2504_high_coverage.sequence.index</code>. Please note that the index files are there as well. You just have to append a <code>.crai</code>. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>dbf39d6ff0e4389b900f9d985f2e6c64</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>4d53ef60ec16f3e4b566c45fdf0fb977</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Variant Tables</td> <td>16925b546051d37cce27df8ec57ccc5e</td> </tr> <tr> <td>1000 Genomes Project 2504</td> <td>Mitochondrial Genome Copy Number</td> <td>365c1b360ea327795d981356064658a6</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div>

opencc-by-4.0Sep 2024View details →
zenodo36/100

EGP Mitochondrial Genome Analysis on 1000 Genomes Project 698 Related Whole-Genome Sequencing Data

<div> <p><strong>Summary:&nbsp;</strong>This dataset consists of running EGP version 1.3 on whole-genome sequencing data from the 1000 Genomes Project 698 Related Dataset. The link to EGP is here https://github.com/tycheleturner/ElGenomaPequeno.</p> <p><strong>Author: </strong>Tychele N. Turner, Ph.D.</p> <p><strong>Short Writeup: EGP version 1.3 on 1000 Genomes Project 698 Related Dataset</strong>:&nbsp;Short-read WGS CRAM files were downloaded through the paths present in this file <code>https://ftp-trace.ncbi.nlm.nih.gov/1000genomes/ftp/1000G_2504_high_coverage/additional_698_related/1000G_698_related_high_coverage.sequence.index</code>. The results are shown below:</p> <div> <table> <tbody> <tr> <td>Public Dataset</td> <td>EGP Result File Type</td> <td>MD5</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Fasta Files for MEGA</td> <td>322038d61b4da2e937b32410613c3532</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome MitoMaster Result File</td> <td>36c782c12245100478903f7fa191a402</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Variant Tables</td> <td>68b2a51361ffae4e7ad9d420b8becd38</td> </tr> <tr> <td>1000 Genomes Project 698 Related</td> <td>Mitochondrial Genome Copy Number</td> <td>1e83c8ae132b0a7ef33b090757b29063</td> </tr> </tbody> </table> <p>Please note: I have found that with Zenodo you must use "Download All" for the copy number table to properly open.</p> </div> <p>&nbsp;</p> </div> <h2>&nbsp;</h2>

opencc-by-4.0Sep 2024View details →
zenodo36/100

GWAS dataset with simulated binary phenotypes for 1000 Genome Project

<p>Phenotypes in this dataset are generated by simulating quantitative blood lipid level phenotype using the method introduced in by (https://app.terra.bio/#workspaces/amp-t2d-op/2019_ASHG_Reproducible_GWAS-V2), and binarized to be suitable as a test case for GWAS methods.</p> <p>ALL.shapeit2_integrated_v1a.GRCh38.20181129.phased.rsid files and allpopid.txt are the 1000 Genome Project reference data used for generating MDS plot in&nbsp;https://github.com/csbio/BridGE-Python</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo28/100

spVCF for the resequenced 1000 Genomes Project (chr21)

<p>chr21 <a href="https://github.com/mlin/spVCF">spvcf.gz</a> file derived from <a href="https://www.internationalgenome.org/data-portal/data-collection/30x-grch38">New York Genome Center&#39;s GATK joint variant call set</a> of <a href="https://www.nature.com/articles/nature15393">1000 Genomes Project</a> whole-genome resequencing (N=2,504), roughly 1/5 of the original vcf.gz file size.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2019View details →
geo24/100

RNA-seq from four balanced translocation carriers in 1000 Genomes Project

GEO Series GSE94043. Homo sapiens. 4 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJul 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record