Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

537

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

537 results for “structural genomics”

Learn how ShareScore rates datasets ↗
zenodo48/100

The pan-genome of Aspergillus fumigatus provides a high-resolution view of its population structure revealing high-levels of lineage-specific diversity driven by recombination

<p><em>Aspergillus fumigatus </em>is a deadly agent of human fungal disease, where virulence heterogeneity is thought to be at least partially structured by genetic variation between strains. While population genomic analyses based on reference genome alignments offer valuable insights into how gene variants are distributed across populations, these approaches fail to capture intraspecific variation in genes absent from the reference genome. Pan-genomic analyses based on <em>de novo</em> assemblies offer a promising alternative to reference-based genomics, with the potential to address the full genetic repertoire of a species. Here, we use a combination of population genomics, phylogenomics, and pan-genomics to assess population structure and recombination frequency, phylogenetically structured gene presence-absence variation, evidence for metabolic specificity, and the distribution of putative antifungal resistance genes in <em>A. fumigatus</em>. &nbsp;We provide evidence for three distinct populations of <em>A. fumigatus</em>, structured by both gene variation (SNPs and indels) and distinct gene presence-absence variation with unique suites of accessory genes present exclusively in each clade. Accessory genes displayed functional enrichment for nitrogen and carbohydrate metabolism, hinting that populations may be stratified by environmental niche specialization. Similarly, the distribution of antifungal resistance genes and resistance alleles were often structured by phylogeny. Despite low levels of outcrossing, <em>A. fumigatus</em> demonstrated a large pan-genome including many genes unrepresented in the Af293 reference genome. These results highlight the inadequacy of relying on a single-reference based approach for evaluating intraspecific variation, and the power of combined genomic approaches to elucidate population structure, genetic diversity, and the putative ecological drivers of clinically relevant fungi.</p> <p>Accompanying manuscript is available as preprint at <a href="https://dx.doi.org/10.1101/2021.12.12.472145">https://dx.doi.org/10.1101/2021.12.12.472145</a>&nbsp;</p> <p>Lotus A.&nbsp;Lofgren,&nbsp;Brandon S.&nbsp;Ross,&nbsp;Robert A.&nbsp;Cramer,&nbsp;Jason E.&nbsp;Stajich. Combined Pan-, Population-, and Phylo-Genomic Analysis of&nbsp;<em>Aspergillus fumigatus</em>&nbsp;Reveals Population Structure and Lineage-Specific Diversity bioRxiv&nbsp;2021.12.12.472145;&nbsp;doi:&nbsp;https://doi.org/10.1101/2021.12.12.472145</p>

opencc-by-4.0Oct 2022View details →
zenodo48/100

Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat

<p>SNPs obtained by UNEAK pipeline for <em>Habromys schmidlyi </em>and <em>Reithrodontomys microdon</em>.&nbsp;</p> <p>Pleae cite as:&nbsp;</p> <p>Colunga-Salas P.,&nbsp;T Marines-Mac&iacute;as,&nbsp;G Hern&aacute;ndez-Canchola,&nbsp;S&nbsp;Barbosa,&nbsp;C&nbsp;Ram&iacute;rez,&nbsp;JB&nbsp;Searle,&nbsp;L&nbsp;Le&oacute;n-Paniagua. 2022.&nbsp;<strong>Population genomics reveals differences in genetic structure between two endemic arboreal rodent species in threatened cloud forest habitat</strong>. Mammalian Reasearch. Doi: 10.1007/s13364-022-00667-x</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Genome-wide structure and function modeling of SARS-COV-2

<p>Homology models and function annotation for all proteins in the SARS-CoV-2 genome. For a description of each file, follow <a href="https://zhanglab.ccmb.med.umich.edu/COVID-19/">this link</a>.&nbsp;</p>

opencc-by-4.0Jan 2020View details →
zenodo44/100

Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project

SV analysis of the long-read sequencing data of 1,019 samples from the 1000 Genomes Project. The data is hosted at the International Genome Sample Resource (IGSR) in the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/">1KG_ONT_VIENNA</a> directory. Please see the <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA.md">README</a> and <a href="https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/data_collections/1KG_ONT_VIENNA/README_1KG_ONT_VIENNA_datareuse_statement.md">data reuse statement</a> for further information about this dataset.

openmit-licenseApr 2024View details →
zenodo44/100

Large structural variations in the haplotype-resolved African cassava genome

<p>Cassava TME7 haplotype resolved assemblies and annotation</p> <p>&nbsp;</p> <p>ABSTRACT:</p> <p>Cassava (<em>Manihot esculenta</em> Crantz, 2n=36) is a global food security crop. Cassava has a highly heterozygous genome, high genetic load, and genotype-dependent asynchronous flowering. It is typically propagated by stem cuttings and any genetic variation between haplotypes, including large structural variations, is preserved by such clonal propagation. Traditional genome assembly approaches generate a collapsed haplotype representation of the genome. In highly heterozygous plants, this results in artifacts and an oversimplification of heterozygous regions. We used a combination of Pacific Biosciences (PacBio), Illumina, and Hi-C to resolve each haplotype of the genome of a farmer-preferred cassava line, TME7 (Oko-iyawo). PacBio reads were assembled using the FALCON suite. Phase switch errors were corrected using FALCON-Phase and Hi-C read data. The ultra-long-range information from Hi-C sequencing was also used for scaffolding. Comparison of the two phases revealed more than 5,000 large haplotype-specific structural variants affecting over 8 Mb, including insertions and deletions spanning thousands of base pairs. The potential of these variants to affect allele specific expression was further explored. RNA-seq data from 11 different tissue types were mapped against the scaffolded haploid assembly and gene expression data are incorporated into our existing easy-to-use web-based interface to facilitate use by the broader plant science community. These two assemblies provide an excellent means to study the effects of heterozygosity, haplotype-specific structural variation, gene hemizygosity, and allele specific gene expression contributing to important agricultural traits and further our understanding of the genetics and domestication of cassava.</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Datasets for 'Mandrake: visualising microbial population structure by embedding millions of genomes into a low-dimensional representation'

<p>Datasets for the paper &#39;<strong>Mandrake: visualising microbial population structure by embedding millions of genomes into a low-dimensional representation</strong>&#39;</p> <p>Files:</p> <ul> <li>616k* - Files for the analysis of 661k bacterial genomes from the SRA (note typo 616-661k). Includes mandrake output and input files (.npz)</li> <li>gps_acc - Files for the analysis of 20k S. pneumoniae accessory genomes from the GPS project. Original accessory matrix is&nbsp;gps_gene_presence_absence.Rtab</li> <li>sc2million_v1* - Files for the analysis of ~1M SARS-CoV-2 genomes.&nbsp;sc2million_v3.npz are the input distances.</li> <li>sce&lt;commit hash&gt;.qdrep - Nvidia systems profile of code at that commit hash</li> <li>sce&lt;commit hash&gt;.ncu-rep - Nvidia kernel profile of code at that commit hash</li> </ul>

opencc-by-4.0Oct 2021View details →
zenodo44/100

ScRAPv20230731: Telomere-to-telomere assemblies of 142 strains characterize the genome structural landscape in Saccharomyces cerevisiae

<p><strong><em>Saccharomyces cerevisiae </em>Reference Assembly Panel (ScRAP) v20230731 </strong>&gt;</p> <p>The haplotype-resolved and/or collapsed T2T genome assemblies for 142 <em>S. cerevisiae</em> strains isolated from diverse geographical and ecological niches.</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Dataset for "Whole-genome de novo assemblies reveal structural variations and organelle-to-nucleus DNA transfers in Asian and African rice""

<p>DXCWR_O.rufipogon_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored&nbsp;genome assembly for <em>O. rufipogon</em>&nbsp; DXCWR.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly&nbsp;DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat&nbsp;annotation for the genome assembly&nbsp;DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored&nbsp;genome assembly for <em>O. glaberrima</em>&nbsp; IRGC104165.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly&nbsp;IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat&nbsp;annotation for the genome assembly&nbsp;IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored&nbsp;genome assembly for <em>O. barthii</em>&nbsp; W1411.</p> <p>W1411_O.barthii_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly&nbsp;W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat&nbsp;annotation for the genome assembly&nbsp;W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored&nbsp;genome assembly for <em>O. nivara</em>&nbsp; W2014.</p> <p>W2014_O.nivara_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly&nbsp;&nbsp;W2014_O.nivara_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat&nbsp;annotation for the genome assembly&nbsp;&nbsp;W2014_O.nivara_scaffolded_anchored.fa.</p>

opencc-by-4.0Feb 2020View details →
dryad40/100

Analysis of copy number variation in dogs implicates genomic structural variation in the development of anterior cruciate ligament rupture

<p>Anterior cruciate ligament (ACL) rupture is an important condition of the human knee. Second ruptures are common and societal costs are substantial. Canine cranial cruciate ligament (CCL) rupture closely models the human disease. CCL rupture is common in the Labrador Retriever (5.79% prevalence), ~100-fold more prevalent than in humans. Labrador Retriever CCL rupture is a polygenic complex disease, based on genome-wide association study (GWAS) of single nucleotide polymorphism (SNP) markers. Dissection of genetic variation in complex traits can be enhanced by studying structural variation, including copy number variants (CNVs). Dogs are an ideal model for CNV research because of reduced genetic variability within breeds and extensive phenotypic diversity across breeds. We studied the genetic etiology of CCL rupture by association analysis of CNV regions (CNVRs) using 110 case and 164 control Labrador Retrievers. CNVs were called from SNPs using three different programs (PennCNV, CNVPartition, and QuantiSNP). After quality control, CNV calls were combined to create CNVRs using ParseCNV and an association analysis was performed. We found no strong effect CNVRs but found 46 small effect (max(T) permutation P&lt;0.05) CCL rupture associated CNVRs in 22 autosomes; 25 were deletions and 21 were duplications. Of the 46 CCL rupture associated CNVRs, we identified 39 unique regions. Thirty four were identified by a single calling algorithm, 3 were identified by two calling algorithms, and 2 were identified by all three algorithms. For 42 of the associated CNVRs, frequency in the population was &lt;10% while 4 occurred at a frequency in the population ranging from 10-25%. Average CNVR length was 198,872bp and CNVRs covered 0.11 to 0.15% of the genome. All CNVRs were associated with case status. CNVRs did not overlap previous canine CCL rupture risk loci identified by GWAS. Associated CNVRs contained 152 annotated genes; 12 CNVRs did not have genes mapped to CanFam3.1. Using pathway analysis, a cluster of 19 homeobox domain transcript regulator genes was associated with CCL rupture (P=6.6E-13). This gene cluster influences cranial-caudal body pattern formation during embryonic limb development. Clustered genes were found in 3 CNVRs on chromosome 14 (HoxA), 28 (NKX6-2), and 36 (HoxD). When analysis was limited to deletion CNVRs, the association was strengthened (P=8.7E-16). This study suggests a component of the polygenic risk of CCL rupture in Labrador Retrievers is associated with small effect CNVs and may include aspects of stifle morphology regulated by homeobox domain transcript regulator genes.</p>

opencc-zeroDec 2020View details →
dryad40/100

Data from: Genomic data reveal deep genetic structure but no support for current taxonomic designation in a grasshopper species complex

<p>Taxonomy has traditionally relied on morphological and ecological traits to interpret and classify biological diversity. Over the last decade, technological advances and conceptual developments in the field of molecular ecology and systematics have eased the generation of genomic data and changed the paradigm of biodiversity analysis. Here we illustrate how traditional taxonomy has led to species designations that are supported neither by high throughput sequencing data nor by the quantitative integration of genomic information with other sources of evidence. Specifically, we focus on <em>Omocestus antigai </em>and<em> O. navasi</em>, two montane grasshoppers from the Pyrenean region that were originally described based on quantitative phenotypic differences and distinct habitat associations (alpine vs. Mediterranean-montane habitats). To validate current taxonomic designations, test species boundaries, and understand the factors that have contributed to genetic divergence, we obtained phenotypic (geometric morphometrics) and genome-wide SNP data (ddRADSeq) from populations covering the entire known distribution of the two taxa. Coalescent-based phylogenetic reconstructions, integrative Bayesian model-based species delimitation, and landscape genetic analyses revealed that populations assigned to the two taxa show a spatial distribution of genetic variation that do not match with current taxonomic designations and is incompatible with ecological/environmental speciation. Our results support little phenotypic variation among populations and a marked genetic structure that is mostly explained by geographic distances and limited population connectivity across the abrupt landscapes characterizing the study region. Overall, this study highlights the importance of integrative approaches to identify taxonomic units and elucidate the evolutionary history of species.</p>

opencc-zeroJul 2019View details →
zenodo40/100

The structure of simple satellite variation in the human genome and its correlation with centromere ancestry (Supplemental Data)

<p>Accompanying <a href="https://github.com/is-the-biologist/1KGP_SATS" target="_blank" rel="noopener">Github</a></p> <p><strong>Supplemental File 1.</strong> BLAST results of k-mer concatemers against T2T-CHM13-v2.0.</p> <p><strong>Supplemental File 2.</strong> Annotations of centromeres, and telomeres of T2T-CHM13-v2.20. Table of abundance of k-mers in annotated regions as numpy file from BLAST hits. Abundance of k-mers across genome in 100kb bins from BLAST hits as .npz files accessible by example:</p> <p>&nbsp; &nbsp; import numpy as np<br>&nbsp; &nbsp; dense = np.load("filename.npz")<br>&nbsp; &nbsp; dense["chr1"]<br>&nbsp;&nbsp;<br><strong>Supplemental File 3</strong>. Table of pairwise R2 between simple satellites and table of pairwise interspersion OR between simple satellites. Folder containing QQ plots of negative binomial fit of satellite copy number distribution used to qualitatively asses model fit.</p> <p><strong>Supplemental File 4. </strong>Materials and results of cenGRM analysis. Boundaries used for centromeric regions of each cenGRM, cenGRMs in GCTA format, and tables with the results of cenGRM GCTA runs. Also provide pdfs of the dendrograms/heatmaps produced from UPGMA clustering of each cenGRM.&nbsp;</p> <p><strong>Supplemental File 5</strong>&nbsp;Non-human significant BLAST hits from BLAST-ing k-mer concatamers to non-human sequences.</p> <p><strong>Supplemental Table 1.</strong> Copy number normalized to 1x depth given GC bias of 126 most abundant satellites analyzed in paper in each individual. Additional columns represent metadata of the individual:</p> <ul> <li>instrument: sequencer instrument name used to sequence library.</li> <li>run: sequencer run of the library.</li> <li>flow: flowcell ID of the ibrary.</li> <li>pop: 1,000 Genomes Project population ID.</li> <li>superpop: 1,000 Genomes Project superpopulation ID.</li> <li>reads: average autosomal read depth of the library.</li> </ul> <p><strong>Supplemental Table 2. </strong>Copy number normalized to 1x depth given GC bias of the top 126 most abundant satellites analyzed in paper in each individual of the 1KGP, plus estimates of the same satellites in CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p> <p><strong>Supplemental Table 3.</strong> Copy number normalized to 1x depth given GC bias of all tandem repeats with k-mer &lt;= 20 (6,309) found collectively in the CHM13 short-read libraries subsampled from 18x-0.5x, 18x depth simulated library of the T2T-CHM13v2.0 assembly analyzed using k-Seek, and Tandem Repeat Finder results of the T2T-CHM13v2.0 asembly <a href="https://doi.org/10.1126/science.abk3112" target="_blank" rel="noopener">Hoyt 2022</a>.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Supporting data for the manuscript "Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads"

<p>Supporting data for the manuscript "Severus: accurate detection and characterization of somatic structural variation in tumor genomes using long reads".</p> <p>The archive contains files that are necessary to reproduce the cell line benchmarks from the paper, including:</p> <ul> <li>Scripts and command lines</li> <li>Original VCF outpurs of all tools used in benchmarking</li> <li>Minda evaluations and truthset VCF files</li> <li>Full Severus outputs + visualizations</li> <li>truvari calls</li> </ul>

opencc-by-4.0Mar 2024View details →
dryad40/100

The chloroplast genomes of Sanicula (Apiaceae): plastome structure, comparative analyses, and phylogenetic relationships

<p><em>Sanicula</em> (Apiaceae subfamily Saniculoideae) is a taxonomically difficult genus of medicinal value. Its distribution center is in China, where there are 18 species (11 of which are endemic). To provide plastid genome resources, whole chloroplast genomes of five <em>Sanicula</em> species (<em>S. flavovirens</em>, <em>S. giraldii</em>, <em>S. lamelligera</em>, <em>S. odorata</em>, and <em>S. rubriflora</em>) were sequenced and compared to the previously published <em>S. orthacantha</em> plastome. These genomes exhibit a typical quadripartite structure. All contain 129 different genes, including 84 protein-coding, 37 tRNA, and 8 rRNA genes. Loci <em>rpl2</em>, <em>matK</em>, <em>psbA</em>, and <em>ycf1</em> are the most variable. Results of maximum likelihood analysis of 90 whole plastome sequences from Apioideae and Saniculoideae and the outgroup <em>Hydrocotyle</em> (Araliaceae) reveal sectional relationships in <em>Sanicula</em> different from the traditional classification system, support the monophyly of Apioideae and its sister group relationship to Saniculoideae, and show concordant topologies to nrDNA ITS and other plastome-based phylogenies. <em>Sanicula orthacantha</em> and <em>S. chinensis</em> form a clade sister group to <em>S. lamelligera</em> and <em>S. odorata</em>, consecutively. These four species comprise a clade sister group to the clade of <em>S. rubriflora</em> and <em>S. flavovirens</em>, with this entire group sister to <em>S. giraldii</em>. The plastid genome resources provided herein will be important for future systematic, evolutionary, phylogenomic, and population-level studies of <em>Sanicula</em>.</p>

opencc-zeroMay 2022View details →
dryad40/100

Discordant population structure among rhizobium divided genomes and their legume hosts

<p>Symbiosis often occurs between partners with distinct life history characteristics and dispersal mechanisms. Many bacterial symbionts have genomes comprised of multiple replicons with distinct rates of evolution and horizontal transmission. Such differences might drive differences in population structure between hosts and symbionts and among the elements of the divided genomes of bacterial symbionts. These differences might, in turn, shape the evolution of symbiotic interactions and bacterial evolution. Here we use whole-genome resequencing of a hierarchically-structured sample of 191 strains of <em>Sinorhizobium meliloti</em> collected from 21 locations in southern Europe to characterize the population structures of this bacterial symbiont and its host plant <em>Medicago truncatula</em>. <em>Sinorhizobium meliloti</em> genomes showed high local (within-site) variation and little isolation by distance. This was particularly true for the two symbiosis elements pSymA and pSymB, which have population structures that are similar to each other, but distinct from both the bacterial chromosome and the host plant. The differences in population structure may result from among-replicon differences in the extent of horizontal gene transfer, although given limited recombination of the chromosome, different levels of purifying or positive selection may also contribute to among-replicon differences. Discordant population structure between hosts and symbionts indicates that geographically and genetically distinct host populations in different parts of the range might interact with genetically similar symbionts, potentially minimizing local specialization.</p>

opencc-zeroSep 2022View details →
zenodo40/100

Genomic Data for "Structure of Anellovirus-like Particles Reveal a Mechanism for Immune EvasionAnellovirus-like Particles Reveal a Mechanism for Immune Evasion"

<p>All genomic data supporting the paper 'Structure of Anellovirus-like Particles Reveal a Mechanism for Immune EvasionAnellovirus-like Particles Reveal a Mechanism for Immune Evasion'. Include amino acid sequence alignment used to produce supplemental figure 6 in the manuscript.</p>

openDec 2024View details →
dryad40/100

Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure

<p>Predicting how tree populations will respond to climate change is an urgent societal concern. An increasingly popular way to make such predictions is the genomic offset (GO) approach, which aims to use genomic and climate data to identify populations that may experience climate maladaptation in the near future. More precisely, GO tries to represent the change in allele frequencies required to maintain the current gene-climate relationships under climate change. However, the GO approach has major limitations and, despite promising validation of its predictions using height data from common gardens, it still lacks broad empirical testing. In the present study, we evaluated the consistency and empirical validity of GO predictions in maritime pine (<em>Pinus pinaster</em> Ait.), a tree species from southwestern Europe and North Africa with a marked population genetic structure. First, gene-climate relationships were estimated using 9,817 SNPs genotyped in 454 trees from 34 populations; and candidate SNPs potentially involved in climate adaptation were identified. Second, GO was predicted using four methods, namely Gradient Forest (GF), Redundancy Analysis (RDA), latent factor mixed model (LFMM) and Generalised Dissimilarity Modeling (GDM), two sets of SNPs (candidate and control SNPs) and five climate general circulation models (GCMs) to account for uncertainty in future climate predictions. Last, the empirical validity of GO predictions was evaluated within a Bayesian framework by estimating the associations between GO predictions and two independent data sources: mortality data from National Forest Inventories (NFI), and mortality and height data from five common gardens in contrasting environments. We found high variability in GO predictions across methods, SNP sets and GCMs. Regarding validation, GO predictions with GDM and GF (and to a lesser extent RDA) based on the candidate SNPs showed the strongest and most consistent associations with mortality rates in common gardens and NFI plots. We found almost no association between GO predictions and tree height in common gardens, most likely due to the overwhelming effect of population genetic structure on tree height in this species. Our study demonstrates the imperative to validate GO predictions with a range of independent data sources before they can be used as informative and reliable metrics in conservation or management strategies.</p>

opencc-zeroMay 2024View details →
zenodo40/100

Dataset for the article: "Weak genetic structure despite strong genomic signal in lesser sandeel in the North Sea"

<p>Dataset used for the article: &quot;Weak genetic structure despite strong genomic signal in lesser sandeel in the North Sea&quot;.</p> <p>Dataset consists on a VCF&nbsp;file&nbsp;from 471 individuals of lesser sandeel, <em>Ammodytes marinus</em> (L.). This VCF is the end product of the bioinformatic analysis described in the paper Jimenez-Mena et al. (2019). Data was obtained from double-digest Restriction-site Associated DNA (ddRAD) sequencing. More information can be obtained in Methods of the article. The information of each of the individuals in the VCF is also included as a separate file, as well as the supplementary tables of the article.</p>

opencc-by-4.0Sep 2019View details →
zenodo40/100

Fig. 5 in The complete mitochondrial genome of Platygaster robiniae (Hymenoptera: Platygastridae): A novel tRNA secondary structure, gene rearrangements and phylogenetic implications

Fig. 5. Phylogenetic tree Note: (A): Maximum likelihood (ML) phylogenetic tree inferred from the mitochondrial genome based on the 13 PCGs dataset; (B): Bayesian inference (BI) phylogenetic tree inferred from the mitochondrial genome based on the 13 PCGs dataset.

opencc-by-4.0Aug 2022View details →
zenodo40/100

Fig. 4 in The complete mitochondrial genome of Platygaster robiniae (Hymenoptera: Platygastridae): A novel tRNA secondary structure, gene rearrangements and phylogenetic implications

Fig. 4. Mitochondrial genome organization of Platygaster robiniae and 11 species of Platygastroidea, compared with the ancestral pancrustacean mt genome organization. Note: tRNA genes are indicated by single letter amino acid codes, L1, L2, S1 and S2 denote tRNALeu(CUN), tRNALeu(UUR), tRNASer(AGN) and tRNASer(UCN), respectively. Genes are transcribed from left to right except those indicated by underlining. Gene movements, relative to the ancestral organization, are indicated with arrows.

opencc-by-4.0Aug 2022View details →
zenodo40/100

Fig. 2 in The complete mitochondrial genome of Platygaster robiniae (Hymenoptera: Platygastridae): A novel tRNA secondary structure, gene rearrangements and phylogenetic implications

Fig. 2. Amino acids (A) and relative synonymous codons (B) of protein-coding genes of the mitochondrial genome of Platygaster robiniae.

opencc-by-4.0Aug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record