Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

865

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

865 results for “population genomics”

Learn how ShareScore rates datasets ↗
zenodo48/100

Genomic vcf file for D.melanogaster, Sussex LHM population

<p>Data, logs and code for genomic vcf genotypes file for the Morrow lab, D.melanogaster LHM sequencing and genotyping project.</p>

opencc-by-4.0Dec 2016View details →
zenodo48/100

Population genomics of Sussex LHM Drosophila melanogaster

<p>Input data, code, output data, summary plots, and run logs for investigation of the population genetics of the Drosophila melanogaster Sussex LHM sample.</p> <p>For output graphs, see popgen_plots.png</p> <p>Input data are 'plink binary' format.</p> <p>Code is a unix/linux shell script containing commands for Plink to perform population genetic tests.</p> <p>The two R scripts contain i. a short command for making a subpopulation file, ii. commands for plotting the output data. Both are initated in the shell script.</p> <p>Platform and version information are available in the log files. Other information available in the shell and R scripts.</p> <p>Broad observations are that the allele frequency disibribution is normal except a few humps around MAF 0.2-0.3 in the autosomes.</p> <p>Linkage disequilbrium, on average, levels-out after ~200bp but there can still be some at distances of 300Kb.</p> <p>The population appears to be divided into four genetically distinct groups (on the IBD-PCA scatter plot), with Fst analysis indicating that this is caused by genetic variation around the centromeres. This is possibly caused by historic admixture, and low centromeric recombination.</p>

opencc-by-4.0Jun 2017View details →
zenodo48/100

Variant, Metabolite and Source Data for: Population genomics uncover loci for trait improvement in the indigenous African cereal tef (Eragrostis tef)

<p>These files contain the variant and metabolome for a collection of 220 tef (<em>Eragrsotis tef)</em> accessions from an ethiopian diversity panel. The accessions were assembled and managed by the Ethiopian Institute of Agricultural Research (EIAR, Ethiopia). The variant data was produced at the John Innes Centre (UK). The metabolome data was produced at Aberystwyth University (UK). These dataset are described in Jones et al. (2024), <em>bioRxiv</em>, https://doi.org/10.1101/2024.09.30.615331. The source data for main figures in the publication are also included.</p> <p>The submission contains</p> <ol> <li>EIAR_filtered.vcf.gz: This is the variant data obtained from alignment of Illumina reads from all 220 teff accessions to the reference assembly of tef (Dabbi). &nbsp;Low quality variants were filtered out. This variant data was used for constructing the phylogenetic relationship between the accessions. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>pooled_EIAR_filtered.vcf.gz: After the phylogentic analysis described above, reads from accessions that were found to be genetically redundant were pooled before variant calling. This file was used for the SNP GWAS analysis. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>&nbsp;Metabolite_Profile.xlxs (source data for Figure 5): This file contains m/z feature intensities from untargeted metabolite fingerprinting using Flow Infusion Electrospray High-resolution Mass Spectrometry (FIE-HRMS). The sample names contains a combination of Location code and Plot number in Supplementary Table S10 e.g AT plot 1, CD plot 1, DZ plot 1, where AT, CD and DZ represent Alem Tena, Chefe Donsa and Debre Zeit, respectively. The data was used for the partial least squares discriminant analysis and differentially accumulated metabolites analysis presented in Figure 5.</li> <li>Source data: Numerical source data for graphs and charts in Figures 3 - 7.</li> <li>Tsedey TT2 Sequence from Improved Assembly: The 4A and 4B sequences around the TT2 orthologue in tef from the improved PacBio-based chromosome-scale assembly of tef. These sequences were used for plotting the LTR Copia alignments presented in Supplementary Figure 9. We thank Corteva for pre-publication access to this improved Tsedey genome assembly.</li> </ol>

opencc-by-4.0Oct 2024View details →
zenodo48/100

The pan-genome of Aspergillus fumigatus provides a high-resolution view of its population structure revealing high-levels of lineage-specific diversity driven by recombination

<p><em>Aspergillus fumigatus </em>is a deadly agent of human fungal disease, where virulence heterogeneity is thought to be at least partially structured by genetic variation between strains. While population genomic analyses based on reference genome alignments offer valuable insights into how gene variants are distributed across populations, these approaches fail to capture intraspecific variation in genes absent from the reference genome. Pan-genomic analyses based on <em>de novo</em> assemblies offer a promising alternative to reference-based genomics, with the potential to address the full genetic repertoire of a species. Here, we use a combination of population genomics, phylogenomics, and pan-genomics to assess population structure and recombination frequency, phylogenetically structured gene presence-absence variation, evidence for metabolic specificity, and the distribution of putative antifungal resistance genes in <em>A. fumigatus</em>. &nbsp;We provide evidence for three distinct populations of <em>A. fumigatus</em>, structured by both gene variation (SNPs and indels) and distinct gene presence-absence variation with unique suites of accessory genes present exclusively in each clade. Accessory genes displayed functional enrichment for nitrogen and carbohydrate metabolism, hinting that populations may be stratified by environmental niche specialization. Similarly, the distribution of antifungal resistance genes and resistance alleles were often structured by phylogeny. Despite low levels of outcrossing, <em>A. fumigatus</em> demonstrated a large pan-genome including many genes unrepresented in the Af293 reference genome. These results highlight the inadequacy of relying on a single-reference based approach for evaluating intraspecific variation, and the power of combined genomic approaches to elucidate population structure, genetic diversity, and the putative ecological drivers of clinically relevant fungi.</p> <p>Accompanying manuscript is available as preprint at <a href="https://dx.doi.org/10.1101/2021.12.12.472145">https://dx.doi.org/10.1101/2021.12.12.472145</a>&nbsp;</p> <p>Lotus A.&nbsp;Lofgren,&nbsp;Brandon S.&nbsp;Ross,&nbsp;Robert A.&nbsp;Cramer,&nbsp;Jason E.&nbsp;Stajich. Combined Pan-, Population-, and Phylo-Genomic Analysis of&nbsp;<em>Aspergillus fumigatus</em>&nbsp;Reveals Population Structure and Lineage-Specific Diversity bioRxiv&nbsp;2021.12.12.472145;&nbsp;doi:&nbsp;https://doi.org/10.1101/2021.12.12.472145</p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat

<p>SNPs obtained by UNEAK pipeline for <em>Habromys schmidlyi </em>and <em>Reithrodontomys microdon</em>.&nbsp;</p> <p>Pleae cite as:&nbsp;</p> <p>Colunga-Salas P.,&nbsp;T Marines-Mac&iacute;as,&nbsp;G Hern&aacute;ndez-Canchola,&nbsp;S&nbsp;Barbosa,&nbsp;C&nbsp;Ram&iacute;rez,&nbsp;JB&nbsp;Searle,&nbsp;L&nbsp;Le&oacute;n-Paniagua. 2022.&nbsp;<strong>Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat</strong>. Mammalian Reasearch. Doi: 10.1007/s13364-022-00667-x</p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Inferring whole-genome histories in large population datasets: inferred tree sequences for 1000 Genomes

<p>Tree sequences inferred for the 1000 Genomes phase 3&nbsp;autosomes using&nbsp;<a href="https://tsinfer.readthedocs.io/">tsinfer</a>&nbsp;version 0.1.4 and compressed using&nbsp;<a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can&nbsp; be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip 1kg_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using&nbsp;<a href="https://tskit.readthedocs.io">tskit</a>.&nbsp;</p> <pre><code class="language-python">import tskit ts = tskit.load("1kg_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original&nbsp;<a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130606_sample_info/20130606_g1k.ped">source</a>&nbsp;and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("1kg_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain&nbsp;all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the&nbsp;ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Inferring whole-genome histories in large population datasets: inferred tree sequences for Simons Genome Diversity Project

<p>Tree sequences inferred for the SGDP autosomes using&nbsp;<a href="https://tsinfer.readthedocs.io/">tsinfer</a>&nbsp;version 0.1.4 and compressed using&nbsp;<a href="https://tszip.readthedocs.io/en/stable/">tszip</a>. Tree sequences can&nbsp; be decompressed as follows:</p> <pre><code class="language-bash">$ tsunzip sgdp_chr1.trees.tsz</code></pre> <p>Once decompressed, trees files can be loaded and processed using&nbsp;<a href="https://tskit.readthedocs.io">tskit</a>.&nbsp;</p> <pre><code class="language-python">import tskit ts = tskit.load("sgdp_chr1.trees") # ts is an instance of tskit.TreeSequence print("Chromosome 1 contains {} trees".format(ts.num_trees))</code></pre> <p>Metadata associated with individuals and populations was derived from the original&nbsp;<a href="https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/SGDP_metadata.279public.21signedLetter.samples.txt">source</a>&nbsp;and converted to JSON form. For example, to access individual metadata we can use:</p> <pre><code class="language-python">import tskit import json ts = tskit.load("sgdp_chr1.trees") ind = ts.individual(0) metadata_dict = json.loads(ind.metadata)</code></pre> <p>The metadata_dict variable will now contain&nbsp;all the metadata for the individual with ID 0 as a dictionary. Metadata associated with populations can be found in a similar way. Population IDs are associated with individuals via their constituent nodes. For example,</p> <pre><code class="language-python">pop_metadata = [json.loads(pop.metadata) for pop in ts.populations()] ind_node = ts.node(ind.nodes[0]) ind_pop_metadata = pop_metadata[ind_node.population]</code></pre> <p>After this, the&nbsp;ind_pop_metadata variable will contain the population level metadata for individual ID 0.</p> <p>The full data pipeline used to generate these tree sequences and associated metadata is available on <a href="https://github.com/mcveanlab/treeseq-inference/tree/master/human-data">GitHub</a>.</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Highly parallel genomic selection response in replicated Drosophila melanogaster populations with reduced genetic variation

<p>Many adaptive traits are polygenic and frequently more loci contributing to the phenotype are segregating than needed to express the phenotypic optimum. Experimental evolution with replicated populations adapting to a new controlled environment provides a powerful approach to study polygenic adaptation. Since genetic redundancy often results in non-parallel selection responses among replicates, we propose a modified Evolve and Resequence (E&amp;R) design that maximizes the similarity among replicates. Rather than starting from many founders, we only use two inbred&nbsp;<em>Drosophila melanogaster</em>strains and expose them to a very extreme, hot temperature environment (29&deg;C). After 20 generations, we detect many genomic regions with a strong, highly parallel selection response in 10 evolved replicates. The X chromosome has a more pronounced selection response than the autosomes, which may be attributed to dominance effects. Furthermore, we find that the median selection coefficient for all chromosomes is higher in our two-genotype experiment than in classic E&amp;R studies. Since two random genomes harbor sufficient variation for adaptive responses, we propose that this approach is particularly well-suited for the analysis of polygenic adaptation.</p> <p>See the README.txt file to get&nbsp;a description of the uploaded files.&nbsp;Scripts.zip contains annotated command lines and scripts for the project&nbsp;(see internal README.txt file).</p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

"Centenarians have a diverse population of gut bacteriophages that may promote healthy lifespan" - Genomes and annotation

<p>File-dump associated with the manuscript:</p> <p>&quot;<strong>Centenarians have a diverse population of gut bacteriophages that may promote healthy lifespan&quot; (Not yet published)</strong></p> <p>MGVs refer to the viral genome database in the publication:&nbsp;https://www.nature.com/articles/s41564-021-00928-6&nbsp;</p> <p>&nbsp;</p> <p>Following uploaded:</p> <p>File 1: VOG Markers in vOTUs/vMAGs and MGV genomes</p> <p>File 2: Viral Tree Newick&nbsp;file with vOTUs/vMAGs and MGV genomes</p> <p>File 3: All vOTUs/vMAGs genomes</p> <p>File 4: Master table annotation of vOTUs/vMAGs</p> <p>File 5: Centenarian bacterial isolate proviruses</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Tracking Selection using Temporal Population Genomics Data

<p>This repository contains the implementation of a pipeline to run the simulations and to produce a reference table for the ABC-RF inference of demography and selection. In its new release, this repository contains the whole-genome polymorphism of contemporary and museum specimens of <em>Apis mellifera</em> feral populations analyzed&nbsp;by Cridland et al. (2018).</p>

opengpl-3.0Mar 2021View details →
zenodo44/100

Variant dataset and code for "Population-level whole genome sequencing of Ascochyta rabiei identifies genomic loci associated with isolate aggressiveness"

<p>This dataset contains genetic variants (SNPs) of <em>Ascochyta rabiei</em> isolates and the R code used in their analysis to generate the results and figures described in the manuscript "<strong>Population-level whole genome sequencing of <em>Ascochyta rabiei</em> identifies genomic loci associated with isolate aggressiveness</strong>".</p> <div> <div>&nbsp;</div> </div>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Genome data and resources on the recombination landscape and population history of the harlequin fly

<p>This dataset contains phased vcf files of <em>Chironomus riparius,&nbsp;</em>ouput files of RepeatMasker, MELT, RepeatOBserver, MSMC2, iSMC and bedtools, such as supporting files.&nbsp;</p> <p>For further details also check the GitHub page: <a href="https://github.com/lpettrich/Crip_Recombination_PopHistory_Cla_2024" target="_blank" rel="noopener">https://github.com/lpettrich/Crip_Recombination_PopHistory_Cla_2024</a></p> <ul> <li><strong>phased-vcfs: </strong>Artificially phased vcf-files of five populations with four individuals each. Needed to generate multihetsep files. Input files for iSMC.<br> <ul> <li>Hesse in Germany =&nbsp; MG</li> <li>Rh&ocirc;ne-Alpes in France = MF</li> <li>Lorraine in France = NMF</li> <li>Piemont in Italy = SI</li> <li>Andalusia&nbsp;in Spain = SS</li> </ul> </li> <li><strong>multihetsep-files:&nbsp;</strong>Created with msmc-tools. Input files for MSMC2.&nbsp;</li> <li><strong>RepeatMasker:&nbsp;</strong>Raw output of RepeatMasker run. Summary file and file with filtered <em>Cla</em>-element (a transposable element) included.<strong><br></strong></li> <li><strong>MELT: </strong>MELT ouput with added info on population and numbered insertions reflecting all 441 detected <em>Cla </em>insertions.<strong><br></strong></li> <li><strong>RepeatOBserver: </strong>Summary files on centromere predictions based on histograms and Shannon Diversity from RepeatOBserver. Genome-wise Shannon Diversity per chromosome included. <strong><br></strong></li> <li><strong>MSMC2: </strong>Raw ouput of combined cross-coalescence and mean values if MSMC2 per populations. <strong><br></strong></li> <li><strong>iSMC: </strong>Recombination rate rho in 10 kb windows and 100 kb windows along the genome. <strong><br></strong></li> <li><strong>bedtools closest ismc 10 kb: </strong>Bedtools closest analysis of the distance of the next <em>Cla</em>-element to the recombination rate rho in 10 kb windows.<strong><br></strong></li> <li><strong>bedtools closest ismc 100 kb:&nbsp;</strong>Bedtools closest analysis of the distance of the next <em>Cla</em>-element to the recombination rate rho in 100 kb windows.</li> <li><strong>input-files figures: </strong>Supporting files needed to create figures.<strong><br></strong></li> </ul>

opencc-by-4.0Oct 2024View details →
zenodo44/100

GWAS to single cell: Intersecting single-cell transcriptomics and genome wide association studies identifies crucial cell-populations and candidate genes for atherosclerosis.

<p><strong>Background</strong></p> <p>Genome-wide association studies (GWAS) have discovered hundreds of common genetic variants for atherosclerotic disease and cardiovascular risk factors. The translation of susceptibility loci into biological mechanisms and targets for drug discovery remains challenging. Intersecting genetic and gene expression data has led to identification of candidate genes. However, the assayed tissues are often non-diseased and heterogeneous in cell composition confounding the candidate prioritization. We collected single-cell transcriptomics (scRNA-seq) from atherosclerotic plaques and aimed to identify cell-type-specific expression of disease-associated genes.&nbsp;</p> <p>&nbsp;</p> <p><strong>Methods and Results</strong></p> <p>To identify disease-associated candidate genes, we applied gene-based analyses using GWAS summary statistics from 46 atherosclerotic, cardiometabolic, and other traits. Next we intersected these candidates with single-cell transcriptomics (scRNA-seq) to identify those genes that are specifically expressed in individual cell (sub)populations of atherosclerotic plaques. We derive an enrichment score and show that loci that associated with coronary artery disease demonstrated a prominent substrate in plaque smooth muscle cells (<em>SKI</em>, <em>KANK2</em>, <em>SORT1</em>), endothelial cells (<em>SLC44A1</em>, <em>ATP2B1</em>), and macrophages (<em>APOE</em>, <em>HNRNPUL1</em>). Further sub clustering of SMC-subtypes revealed genes in risk loci for coronary calcification specifically enriched in a synthetic cluster of SMCs. To verify the robustness of our approach, we used liver-derived scRNAseq-data and showed enrichment of circulating lipids-associated loci in hepatocytes.</p> <p><br> <strong>Conclusion</strong></p> <p>We confirm known gene-cell pairs relevant for atherosclerotic disease, and discovered novel pairs pointing to new biological mechanisms amenable for therapy. We present an intuitive single-cell transcriptomics driven workflow rooted in human large-scale genetic studies to identify putative candidate genes and affected cells associated with cardiovascular traits.</p> <p>&nbsp;</p>

opencc-by-4.0May 2021View details →
zenodo44/100

REPIN population analysis in 42 Pseudomonas chlororaphis genomes

<p>This dataset is the output of RAREFAN (http://rarefan.evolbio.mpg.de/) a webserver to identify REPIN populations across an entire bacterial species. The data was created using the following command &quot;java -jar -Xmx10g rarefan.jar chlororaphis/in/ chlororaphis/out/ chlTAMOak81.fas 55 21 chlororaphis/in/yafM_SBW25.faa chlororaphis.nwk 1e-30 true 1&quot;</p> <p>All input files are located in the folder chlororaphis/in/, all output data is located in chlororaphis/out/.</p> <p>The input files include the 42 fasta formatted <em>P. chlororaphis</em> genome files (*.fas) and a RAYT protein sequence called yafM_SBW25.faa.</p> <p>The output files include the following:</p> <p>A phylogenetic tree &quot;chlororaphis.nwk&quot; of all genomes generated with andi (<a href="http://github.com/evolbioinf/andi/">http://github.com/evolbioinf/andi/</a>) and clustDist (http://guanine.evolbio.mpg.de/problemsBook/node1.html).</p> <p>A file containing the frequencies of all 21bp long sequences found in the TAMOak81 genome:&nbsp;chlTAMOak81.wfr</p> <p>A file containing all 21bp long sequences that occur more frequently than 55 times in the TAMOak81 genome:&nbsp;chlTAMOak81.overrep</p> <p>A file containing information on the RAYTs and their cooccurrence with different REPIN populations: prox.stats</p> <p>A file containing the nucleotide sequences of all yafM_SBW25.faa&nbsp;relatives identified with BLAST+&nbsp;in the <em>P. chlororaphis</em>&nbsp;species:&nbsp;yafM_relatives.fna</p> <p>maxREPIN_[0-5] Contains the most frequent REPIN identified for each sequence type in each <em>P. chlororaphis</em> strain.</p> <p>&nbsp;presAbs_[0-5].txt Contains for each strain information on the number of RAYTs, the number of REPINs, the master sequence, the number of master sequences, the entire REP/REPIN population size, the number of REPIN clusters that contain more than 10 sequences, all REPINs in the population as well as all all REPINs that differ to the master sequences in at most three nucleotides.</p> <p>rayt_[strain name].tab contains location information for each identified RAYT relative for each strain. The files can be viewed with artemis.</p> <p>results.txt contains for each strain the frequency of the six identified 21bp long seeds.</p> <p>There is one folder called groupSeedSequences, which includes the data for identifying the most common 21 bp long sequences in <em>P. chlororaphis</em> TAMOak81. All 21bp long sequences in the genome that occur more frequently than 55 times are sorted into 6 sequence groups. These sequence groups are stored in the files&nbsp;Group_chlTAMOak81_*.out and .out.fas. There is also a chlTAMOak81_words.tab file, which contains the locations of all overrepresented 21bp long sequences in the TAMOak81 genome. This file can be viewed in artemis (https://www.sanger.ac.uk/tool/artemis/) together with the TAMOak81 genome file. The most common sequence in each group is used as a seed sequence to determine REPIN populations across all 42 genomes.</p> <p>&nbsp;</p> <p>For each genome there are six&nbsp;output folders (ending in _0 to _5), for each sequence group one.</p> <p>Each folder contains the following files:</p> <p>*.dd: Degree distribution of the REPIN network, where each REPIN is a node. A REPIN is connected to another REPIN if they differ in exactly one position. The degree distribution is a histogram of the number of connections of all the nodes.&nbsp;</p> <p>*.hist For the largest sequence cluster determined by mcl that consists of REPINs (two REPs in inverted orientation) this file contains the number of REPINs in each sequence class. Sequence class 0 is the master sequence. By definition the most common REPIN in the sequence population. Sequence class 1 contains all REPINs differing in exactly one position to the master sequence. Sequence class 2 contains REPINs differing in 2 positions etc.</p> <p>*.mcl Contains the clustering output by mcl. Each line contains the member of a cluster. Lines are sorted by cluster size.</p> <p>*.mw Contains the most common 21bp long sequence and its frequency in the genome, which is the basis for identifying first all related REP sequences and from those the REPINs formed by these REP sequences.</p> <p>*.nodes The identity and frequency of all REPINs and REP sequences for&nbsp;either all sequences or only for the largest sequence cluster.</p> <p>*.ss Contains REPINs and REP sequences as well as their positions in fasta format. Position information starts with the location in genome fasta file (first sequence is 0...) followed by the start and end position of the entire REPIN/REP sequence.&nbsp;&nbsp;</p> <p>*.ss.REP REP sequence information in fasta format.</p> <p>*.tab Location in tab format. Can be used to display locations of REPs and REPINs in the genome via artemis.</p> <p>*_[0-9].ss Contains REPIN/REP sequence information for each subcluster separately.</p> <p>*_[0-9].tab Contains the location of REP/REPINs for each subcluster separately for viewing in artemis.</p> <p>*allSeed.nw Contains network connections between nodes of all sequences. Can be used to view network in for example R or cytoscape together with the nodes file.</p> <p>*largestCluster.nodes Information on nodes only from the largest REPIN cluster.</p> <p>*largestCluster.ss *.ss file for the largest REPIN cluster.</p> <p>*largestCluster.tab *.tab file for the largest REPIN cluster.</p> <p>*_rayt_repin_prox.txt shows which REPIN/REP cluster is in proximity to any of the RAYT genes identified in the genome (within 200bp).</p> <p>And a subfolder that contains the complete sequences (including the variable region) for all identified REPs and REPINs.</p> <p><strong>The dataset was generated using the following external tools:</strong></p> <p>andi for tree building:</p> <p>B Haubold, F Kl&ouml;tzl, and P Pfaffelhuber.&nbsp;<strong>andi: fast and accurate estimation of evolutionary distances between closely related genomes.</strong> Bioinformatics, 2015 vol. 31 (8) pp. 1169-1175.</p> <p>MCL for REPIN population clustering:</p> <p>A J Enright, S Van Dongen, and C A Ouzounis.&nbsp;<strong>An efficient algorithm for large-scale detection of protein families.</strong>&nbsp;Nucleic Acids Research, 2002 vol. 30 (7) pp. 1575-1584.</p> <p>BLAST+ for identifying RAYT relatives in the different genomes:</p> <p>C&nbsp;Camacho, G&nbsp;Coulouris, V&nbsp;Avagyan, N&nbsp;Ma, J&nbsp;Papadopoulos, K&nbsp;Bealer, and T&nbsp;L Madden.&nbsp;<strong>BLAST+: architecture and applications.</strong> BMC Bioinformatics, 2009 vol. 10 (1) pp. 421-9.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

REPIN population analysis in 130 Neisseria meningitidis and N. gonorrhoeae genomes

<p>This dataset is the output of RAREFAN (http://rarefan.evolbio.mpg.de/) a webserver to identify REPIN populations across an entire bacterial species. The data was created using the following command &quot;java -jar -Xmx10g rarefan.jar neisseria/in/ neisseria/out/ Nmen_2594.fas 55 21 neisseria/in/NMAA_0235.faa neisseria.nwk 1e-80 true 1&quot;</p> <p>All input files are located in the folder neisseria/in/, all output data is located in neisseria/out/.</p> <p>The input files include the 130 fasta formatted&nbsp;<em>Neisseria</em>&nbsp;genome files (*.fas) and a RAYT protein sequence called NMAA_0235.faa.</p> <p>The output files include the following:</p> <p>A phylogenetic tree &quot;neisseria.nwk&quot; of all genomes generated with andi (<a href="http://github.com/evolbioinf/andi/">http://github.com/evolbioinf/andi/</a>) and clustDist (http://guanine.evolbio.mpg.de/problemsBook/node1.html).</p> <p>A file containing the frequencies of all 21bp long sequences found in the Nmen_2594 genome:&nbsp;Nmen_2594.wfr</p> <p>A file containing all 21bp long sequences that occur more frequently than 55 times in the Nmen_2594 genome:&nbsp;Nmen_2594.overrep</p> <p>A file containing information on the RAYTs and their cooccurrence with different REPIN populations: prox.stats</p> <p>A file containing the nucleotide sequences of all NMAA_0235.faa&nbsp;relatives identified with BLAST+ in the <em>Neisseria&nbsp;</em>species:&nbsp;yafM_relatives.fna</p> <p>maxREPIN_[0-5] Contains the most frequent REPIN identified for each sequence type in each&nbsp;<em>Neisseria</em>&nbsp;strain.</p> <p>&nbsp;presAbs_[0-5].txt Contains for each strain information on the number of RAYTs, the number of REPINs, the master sequence, the number of master sequences, the entire REP/REPIN population size, the number of REPIN clusters that contain more than 10 sequences, all REPINs in the population as well as all REPINs that differ to the master sequences in at most three nucleotides.</p> <p>rayt_[strain name].tab contains location information for each identified RAYT relative for each strain. The files can be viewed with artemis.</p> <p>results.txt contains for each strain the frequency of the six identified 21bp long seeds.</p> <p>There is one folder called groupSeedSequences, which includes the data for identifying the most common 21 bp long sequences in <em>Neisseria meningitidis</em>&nbsp;WUE 2594. All 21bp long sequences in the genome that occur more frequently than 55 times are sorted into 6 sequence groups. These sequence groups are stored in the files&nbsp;Group_Nmen_2594_*.out and .out.fas. There is also a Nmen_2594_words.tab file, which contains the locations of all overrepresented 21bp long sequences in the Nmen_2594 genome. This file can be viewed in artemis (https://www.sanger.ac.uk/tool/artemis/) together with the Nmen_2594 genome file. The most common sequence in each group is used as a seed sequence to determine REPIN populations across all 130 genomes.</p> <p>&nbsp;</p> <p>For each genome there are six&nbsp;output folders (ending in _0 to _5), for each sequence group one.</p> <p>Each folder contains the following files:</p> <p>*.dd: Degree distribution of the REPIN network, where each REPIN is a node. A REPIN is connected to another REPIN if they differ in exactly one position. The degree distribution is a histogram of the number of connections of all the nodes.&nbsp;</p> <p>*.hist For the largest sequence cluster determined by mcl that consists of REPINs (two REPs in inverted orientation) this file contains the number of REPINs in each sequence class. Sequence class 0 is the master sequence. By definition the most common REPIN in the sequence population. Sequence class 1 contains all REPINs differing in exactly one position to the master sequence. Sequence class 2 contains REPINs differing in 2 positions etc.</p> <p>*.mcl Contains the clustering output by mcl. Each line contains the member of a cluster. Lines are sorted by cluster size.</p> <p>*.mw Contains the most common 21bp long sequence and its frequency in the genome, which is the basis for identifying first all related REP sequences and from those the REPINs formed by these REP sequences.</p> <p>*.nodes The identity and frequency of all REPINs and REP sequences for&nbsp;either all sequences or only for the largest sequence cluster.</p> <p>*.ss Contains REPINs and REP sequences as well as their positions in fasta format. Position information starts with the location in genome fasta file (first sequence is 0...) followed by the start and end position of the entire REPIN/REP sequence.&nbsp;&nbsp;</p> <p>*.ss.REP REP sequence information in fasta format.</p> <p>*.tab Location in tab format. Can be used to display locations of REPs and REPINs in the genome via artemis.</p> <p>*_[0-9].ss Contains REPIN/REP sequence information for each subcluster separately.</p> <p>*_[0-9].tab Contains the location of REP/REPINs for each subcluster separately for viewing in artemis.</p> <p>*allSeed.nw Contains network connections between nodes of all sequences. Can be used to view network in for example R or cytoscape together with the nodes file.</p> <p>*largestCluster.nodes Information on nodes only from the largest REPIN cluster.</p> <p>*largestCluster.ss *.ss file for the largest REPIN cluster.</p> <p>*largestCluster.tab *.tab file for the largest REPIN cluster.</p> <p>*_rayt_repin_prox.txt shows which REPIN/REP cluster is in proximity to any of the RAYT genes identified in the genome (within 200bp).</p> <p>And a subfolder that contains the complete sequences (including the variable region) for all identified REPs and REPINs.</p> <p>The dataset was generated using the following external tools:</p> <p>andi for tree building:</p> <p>B Haubold, F Kl&ouml;tzl, and P Pfaffelhuber.&nbsp;<strong>andi: fast and accurate estimation of evolutionary distances between closely related genomes.</strong> Bioinformatics, 2015 vol. 31 (8) pp. 1169-1175.</p> <p>MCL for REPIN population clustering:</p> <p>A J Enright, S Van Dongen, and C A Ouzounis.&nbsp;<strong>An efficient algorithm for large-scale detection of protein families.</strong>&nbsp;Nucleic Acids Research, 2002 vol. 30 (7) pp. 1575-1584.</p> <p>BLAST+ for identifying RAYT relatives in the different genomes:</p> <p>C&nbsp;Camacho, G&nbsp;Coulouris, V&nbsp;Avagyan, N&nbsp;Ma, J&nbsp;Papadopoulos, K&nbsp;Bealer, and T&nbsp;L Madden.&nbsp;<strong>BLAST+: architecture and applications.</strong> BMC Bioinformatics, 2009 vol. 10 (1) pp. 421-9.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

Datasets for 'Mandrake: visualising microbial population structure by embedding millions of genomes into a low-dimensional representation'

<p>Datasets for the paper &#39;<strong>Mandrake: visualising microbial population structure by embedding millions of genomes into a low-dimensional representation</strong>&#39;</p> <p>Files:</p> <ul> <li>616k* - Files for the analysis of 661k bacterial genomes from the SRA (note typo 616-661k). Includes mandrake output and input files (.npz)</li> <li>gps_acc - Files for the analysis of 20k S. pneumoniae accessory genomes from the GPS project. Original accessory matrix is&nbsp;gps_gene_presence_absence.Rtab</li> <li>sc2million_v1* - Files for the analysis of ~1M SARS-CoV-2 genomes.&nbsp;sc2million_v3.npz are the input distances.</li> <li>sce&lt;commit hash&gt;.qdrep - Nvidia systems profile of code at that commit hash</li> <li>sce&lt;commit hash&gt;.ncu-rep - Nvidia kernel profile of code at that commit hash</li> </ul>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Data to support Whitney JL, Coleman RR, Deakos MH "Genomic evidence indicates small island-resident populations and sex-biased behaviors of Hawaiian Reef Manta Rays"

<p>Datasets supporting the manuscript: Whitney JL, Coleman RR, Deakos MH &quot;Genomic evidence indicates small island-resident populations and sex-biased behaviors of Hawaiian Reef Manta Rays&quot;. <em>BMC Ecology and Evolution&nbsp;</em><strong>23</strong>, 31 (2023). https://doi.org/10.1186/s12862-023-02130-0</p> <p>Nuclear data:</p> <p>&quot;Mobula-alfredi_nuclear_reference_RAD_contigs.fasta&quot; is a fasta of 359,751 contigs that serve as the reference for nuclear alignment of genotypes to RAD loci. Contigs begin and end with GATC cut site.</p> <p>Mobula-alfredi_nuclear_all_2048snps_38genotypes.vcf is a VCF file with all 2048 nuclear SNPs in final filtered SNP dataset. 38 genotypes are included from Maui Nui and Hawaii Island. This 2048 SNPs includes both 2038 neutral and 10 outlier SNPs.&nbsp;</p> <p>Mobula-alfredi_nuclear_neutral_2038snps_38genotypes.vcf&nbsp;is a VCF file with 2038 neutral nuclear SNPs genotyped in 38&nbsp;individuals from Maui Nui and Hawaii Island.&nbsp;</p> <p>Mobula-alfredi_nuclear_outliers_10snps_38genotypes.vcf is a VCF file with 10 outlier SNPs genotyped in 38&nbsp;individuals from Maui Nui and Hawaii Island.&nbsp;</p> <p>Structure (.str) files are also provided in addition to VCFs.&nbsp;In all files Population prefixes M=Maui Nui and K=Hawaii Island.&nbsp;</p> <p>Mitochondrial data:</p> <p>Mobula-alfredi_mitogenome_34haplotypes_9sites_min4x.vcf is a VCF file with 9 variant sites across the mitogenome haplotyped in 34 individuals from Maui Nui and Hawaii Island.&nbsp;</p> <p>Mobula-alfredi_mitogenome_34haplotypes_allsites_min4x.fasta is a FASTA file with whole mitogenomes aligned to OP562409 [https://www.ncbi.nlm.nih.gov/nuccore/OP562409]. Sites with less than 4x coverage&nbsp;were masked with Ns.&nbsp;</p> <p>Mobula-alfredi_mitogenome_reference_OP562409.fasta is a FASTA file containing the <em>Mobula alfredi</em> reference mitogenome&nbsp;OP562409 [https://www.ncbi.nlm.nih.gov/nuccore/OP562409].</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Lepidoptera genomics based on 88 chromosomal reference sequences informs population genetic parameters for conservation

<p>This repository contains (1) germline mutations called by the DeepVariant (v1.1.0) pipeline in VCF format; (2) rejected substitution scores calculated by the Genomic Evolutionary Rate Profiling (GERP++) software on each species and chromosome; and (3) the phylogenetic tree used as guide tree in the Cactus alignment.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Data For: Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data

<p>Simulation output and Genome-wide scan for nIBD variants in UK10K data as reported in:</p> <p>Identifying rare variants inconsistent with identity-by-descent in population-scale whole-genome sequencing data</p> <p>Johnson KE, Adams CJ, Voight BF. Methods Ecol Evol 2022 Nov;13(11):&nbsp;2429&ndash;2442.</p> <p>Code available at:&nbsp;https://github.com/kelsj/EVICORD</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Genome-wide characterization of human minisatellite VNTRs: population-specific alleles and gene expression differences

<p>This repository consists of minisatellite VNTR genotypes for 2,800 samples (2,770 individuals). The raw VCF files were produced using <a href="https://github.com/yzhernand/VNTRseek">VNTRseek</a>&nbsp;on xxx data sources: <a href="http://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1000_genomes_project/">30 high coverage WGS datasets</a>&nbsp;from the 1000 Genomes Project phase 3, <a href="https://www.internationalgenome.org/data-portal/data-collection/30x-grch38">2,504 unrelated genomes</a> from New York Genome Center (NYGC), <a href="https://www.internationalgenome.org/data-portal/data-collection/sgdp">253 genomes from Simons Diversity Genome Project</a>&nbsp;(SGDP), <a href="https://www.illumina.com/products/by-type/informatics-products/basespace-sequence-hub/apps/tumor-normal.html">two tumor-normal breast cancer samples</a>&nbsp;from Illumina Basespace, haploid genomes <a href="https://www.ncbi.nlm.nih.gov/sra/SRX652547">CHM1 </a>and <a href="https://www.ncbi.nlm.nih.gov/sra/SRX1009644">CHM13</a>, and seven genomes from the Personal Genome Project from the Genome In A Bottle Consortium (GIAB). Raw VCF files are provided for each data source separately.</p> <p>The raw VCF files were preprocessed (preprocess.sh) to extract genotypes and provided in VNTRseek_preprocessed_data.tar.gz (uncompressed size 10G). The R Markdown code to analyze the preprocessed data and produce figures and tables is also provided (tables_and_figures.Rmd). For more information see the ReadMe file.</p> <p>This work was supported in part by NSF grants IIS-1423022 and DBI-1559829.</p>

opencc-by-4.0Nov 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record