Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,574

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,574 results for “genome sequence”

Learn how ShareScore rates datasets ↗
zenodo40/100

FIG. 2 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 2.—PCA result plot showing clustering of individual bats from four Hawaiian Islands using 21,808,031 SNPs. Sample information included in supplementary table S4, Supplementary Material online.

opencc-by-4.0Aug 2020View details →
zenodo40/100

FIG. 4 in Analysis of Genomic Sequence Data Reveals the Origin and Evolutionary Separation of Hawaiian Hoary Bat Populations

FIG. 4.—SNAPP-based phylogenetic tree inference. (A) The maximum clade credibility or consensus tree, showing approximate divergence of hoary bats across the Hawaiian archipelago. The axis on the bottom of the figure corresponds to million years before present (Ma), using the emergence of Hawai'i (~0.43 Ma) as a calibration point (95% confidence intervals were given in square brackets). (B) The drawing of all sampled trees showing all ingroup nodes were supported by maximum posterior probabilities (1.00).

opencc-by-4.0Aug 2020View details →
zenodo40/100

Genbank annotation and sequence of the genome of a Rickettsiales symbiont of Reticulomyxa filosa

<p>The genome sequence of a Rickettsiales symbiont was selectively from the genome sequencing reads of its host, Reticulomyxa filosa. Those sequences were obtained from a previous third-party study (doi: 10.1016/j.cub.2013.11.027). The selective assembly procedure was based on GC content and coverage of contigs obtained from the total read sets with SPAdes. Full details are provided in the manuscript file.<br>The obtained symbiont sequence was then annotated with Prokka, and the gbk output file selected.</p><p>This genome was obtained and analysed in the context of a larger genome comparative studies on the Rickettsiales, aimed to investigate the evolutionary patterns within the whole lineage.</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

apomixis_parallel_evolution, and Fortunella hindsii (Citrus hindsii )Genome sequencing and assembly

<p>##The assemble(genome) files</p> <p>Citrus hindsii (Mini citrus)Genome assembly</p> <p>Citrus hindsii (Mini citrus)Genome assembly gene model gff3 file</p> <p>Citrus hindsii (Mini citrus)Genome assembly function annotation</p> <p>Citrus hindsii (Mini citrus)Genome assembly TE gff3 file</p> <p>##The population dataset</p> <p>LD_pur.vcf.gz&nbsp; //The LD purning SNP vcfs (1.4 M sites) used in analysis</p> <p><br> log10_auxin.txt //The expression (log10) related to auxin pathway</p> <p><br> sjg.temergedref.fasta.gz //The TE insertion modify genome in popTE2 analysis</p> <p><br> SVs.vcf.gz //The SV vcfs&nbsp; used in the paper</p> <p><br> te-hierarchy.txt //The TE classfication in popTE2 analysis</p> <p><br> TPM_count.txt //All samples expression in TMP count</p> <p><br> unfiltered.vcf.gz //The unfiltered vcfs file (7.3 M sites, within 0.4 M indels)&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Genome-wide mapping of individual replication fork velocities using nanopore sequencing

<p>These data are related to the <strong>&quot;Genome-wide mapping of individual replication fork velocities using nanopore sequencing&quot;&nbsp;</strong>manuscript by Theulot, Lacroix et al. (https://doi.org/10.1038/s41467-022-31012-0) and to the GitHub repository NanoForkSpeed (https://github.com/LacroixLaurent/NanoForkSpeed).</p> <p>BT1_run4.tar.gz contains the nanopore reads from a typical experiment as fast5 files.</p> <p>BT1_run4_mega.zip contains the result of the BrdU basecalling as described on the GitHub page in a modified bam file format&nbsp;as described on the GitHub page NanoForkSpeed/BrdU_Basecalling..</p> <p>BT1_run4_Megalodon_00_smdata.rds contains the result of the parsing function for the modified bam file as described on the GitHub page NanoForkSpeed/BrdU_Basecalling.</p> <p>BT1_run4_Megalodon_00_NFS_data.rds contains the result of the fork detection procedure on the&nbsp;BT1_run4_Megalodon_00_smdata.rds file&nbsp;as described on the GitHub page NanoForkSpeed/Forks_Detection.</p> <p>BT1_run4_merged_NFS_data.rds contains the result of the NFS_merge function as described on the GitHub page NanoForkSpeed/Forks_Detection.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
dryad40/100

Genome-wide sequence data show no evidence of hybridization and introgression among pollinator wasps associated with a community of Panamanian strangler figs

<p>The specificity of pollinator host choice influences opportunities for reproductive isolation in their host plants. Similarly, host plants can influence opportunities for reproductive isolation in their pollinators. For example, in the fig and fig wasp mutualism, offspring of fig pollinator wasps mate inside the inflorescence that the mothers pollinate. Although often host specific, multiple fig pollinator species are sometimes associated with the same fig species, potentially enabling hybridization between wasp species. Here we study the 19 pollinator species (<em>Pegoscapus</em> spp.) associated with an entire community of 16 Panamanian strangler fig species (<em>Ficus</em> subgenus <em>Urostigma</em>, section <em>Americanae</em>) to determine whether the previously documented history of pollinator host switching and current host sharing predicts genetic admixture among the pollinator species, as has been observed in their host figs. Specifically, we use genome-wide ultraconserved element (UCE) loci to estimate phylogenetic relationships and test for hybridization and introgression among the pollinator species. In all cases, we recover well-delimited pollinator species that contain high interspecific divergence. Even among pairs of pollinator species that currently reproduce within syconia of shared host fig species, we found no evidence of hybridization or introgression. This is in contrast to their host figs, where hybridization and introgression have been detected within this community, and more generally, within figs worldwide. Consistent with general patterns recovered among other obligate pollination mutualisms (<em>e.g.</em>, yucca moths and yuccas), our results suggest that while hybridization and introgression are processes operating within the host plants, these processes are relatively unimportant within their associated insect pollinators.<br>  </p>

opencc-zeroFeb 2022View details →
zenodo40/100

CONGA: Copy number variation genotyping in ancient genomes and low-coverage sequencing data

<p>To date, ancient genome analyses have been largely confined to the study of single nucleotide polymorphisms (SNPs). Copy number variants (CNVs) are a major contributor of disease and of evolutionary adaptation, but identifying CNVs in ancient shotgun-sequenced genomes is hampered by (i) most published genomes being &lt;1x&nbsp;coverage, (ii) ancient DNA fragments being typically &lt;80 bps. These characteristics preclude state-of-the-art CNV detection software to be effectively applied to ancient genomes. Here we present CONGA, an algorithm tailored for genotyping deletion and duplication events in genomes with low depths of coverage. Simulations and down-sampling experiments show that CONGA can genotype deletions &gt;1 kbps with F-scores &gt;0.75 at &gt;=1x, and distinguish between heterozygous and homozygous states. Using CONGA, we analyse deletion events at 10,018 loci in 56 ancient human genomes spanning the last 50,000 years, with coverages 0.4x-26x. We show that inter-individual genetic diversity measured using deletions and SNPs are highly correlated, as in modern-day genomes, confirming that deletion frequencies broadly reflect demographic history. We also identify signatures of strong purifying selection on deletions in ancient-genomes, such as an excess of singletons compared to those in SNPs. CONGA paves the way for systematic studies of drift, mutation load, and adaptation in ancient and modern-day gene pools through the lens of CNVs.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Sequence Similarity Network (SSN) and Genome Neighbourhood Network (GNN) for Mycobacterium Cytochrome P450 enzymes

<p>This dataset was generated in the context of the Horizon 2020&nbsp;MSCA IF action deCrYPtion (Grant 839116). The aim of this project is to use comparative genomics in order to propose and then test the function of uncharacterised Cytochrome P450 enzymes that are present among Mycobacterium species.</p> <p>More information about this project can be found at:&nbsp;https://cordis.europa.eu/project/id/839116.</p> <p>This dataset contains:</p> <p>- The FASTA sequences files obtained from the UniProt database, for members of the PF00067 protein family (CYP).</p> <p>- A set of reference FASTA sequences, matching the supplementary material from the following publication:&nbsp;Parvez, M.&nbsp;<em>et al.</em>&nbsp;(2016) &lsquo;Molecular evolutionary dynamics of cytochrome P450 monooxygenases across kingdoms: Special focus on mycobacterial P450s&rsquo;,&nbsp;<em>Scientific Reports</em>, 6(1), p. 33099. doi:<a href="https://doi.org/10.1038/srep33099">10.1038/srep33099</a>.</p> <p>- A combined FASTA files of both previously described, that was used for the generation of SSNs</p> <p>- A PNG&nbsp;image&nbsp;produced from the analysis of the&nbsp;Sequence Similarity Networks generated at AST78 (corresponding to 40% identity, defining CYP families)</p> <p>- A PNG&nbsp;image&nbsp;produced from the analysis of the&nbsp;Sequence Similarity Networks generated at AST141&nbsp;(corresponding to 55% identity, defining CYP subfamilies)</p> <p>- A Cytoscape session for&nbsp;the&nbsp;Sequence Similarity Networks from the combined FASTA file&nbsp;generated using the Enzyme Function Initiative web tools (https://efi.igb.illinois.edu), at AST78</p> <p>- A Cytoscape session containing&nbsp;Sequence Similarity Networks and&nbsp;Genome Neighborhood Network from the combined FASTA file&nbsp;&nbsp;generated using the Enzyme Function Initiative web tools (https://efi.igb.illinois.edu), at AST141</p>

opencc-by-4.0May 2022View details →
zenodo40/100

Polymorphism-aware estimation of species trees and evolutionary forces from genomic sequences with RevBayes

<p>Supplementary files of Polymorphism-aware estimation of species trees and evolutionary forces from genomic sequences with RevBayes by Borges, Boussau, H&ouml;hna, Pereira and Kosiol<br> &nbsp;</p>

opencc-by-4.0May 2022View details →
dryad40/100

Phylogenetic analysis of policistronic amino-acid sequences encoded by 116 flavivirus genomes

<p><span>Recently, Genome Biology and Evolution (11:3341-3352) published three statistical tests for testing whether alignments of sequence data violate the phylogenetic assumption of evolution under homogeneous conditions. The tests extend the matched-pairs tests of symmetry, marginal symmetry, and internal symmetry for pairs of aligned homologous sequences to the case where a whole alignment is considered. Here we reveal that the new tests are misleading. We explain why this is so, reveal how the tests of whole alignments may be done, and release new bioinformatics tools that implement statistically sound methods of dealing with multiple comparisons (i.e., by controlling the family-wise error rate or the false discovery rate). Using the new software to analyse an alignment of amino acids encoded by 116 flavivirus genomes, we reveal, for the first time, that these genomes are unlikely to have evolved under stationary, reversible, and homogeneous Markovian conditions.</span></p>

opencc-zeroJun 2022View details →
zenodo40/100

Whole genome re-sequencing of B. dorsalis in the Indian Ocean

<p>Files containing essential data for the production of the paper:&nbsp;<em>Bactrocera dorsalis </em>in the Indian Ocean: a tale of two invasions (Deschepper et. al, 202X) DOI:####</p>

opencc-by-4.0May 2022View details →
dryad40/100

Deep-sequencing of viral genomes from treatment-naive HIV-infected persons shows positive association between intrahost genetic diversity and viral load

<p><strong><span>Background:</span></strong><span> Infection with human immunodeficiency virus type 1 (HIV) typically results from transmission of a small and genetically uniform viral population. Following transmission, the virus population becomes more diverse because of recombination and acquired mutations through genetic drift and selection. Viral intrahost genetic diversity remains a major obstacle to the cure of HIV; however, there is a disagreement whether intrahost viral genetic diversification associates positively or negatively with disease progression and progression markers. Viral load is a key progression marker and understanding its relationship to viral intrahost genetic diversity could help design future strategies for HIV monitoring and treatment.</span></p> <p><span><strong>Methods:</strong> </span><span>We analyzed deep-sequenced viral genomes from 2,650 treatment-naive HIV-infected persons to measure the intrahost genetic diversity of 2,447 genomic codon positions as calculated by Shannon entropy. We tested for associations between viral load (VL) and amino acid (AA) entropy accounting for sex, age, race, duration of infection, and HIV population structure.</span></p> <p><strong><span>Results:</span></strong><span><strong> </strong>We confirmed that the intrahost genetic diversity is highest in the <em>env</em> gene. Furthermore, we showed that mean Shannon entropy is significantly associated with VL, especially in infections of &gt;24 months duration. We identified 16 significant associations between VL (p-value&lt;2.0x10<sup>-5</sup>) and Shannon entropy at AA positions which in our association analysis explained 13% of the variance in VL.</span></p> <p><strong><span>Conclusions: </span></strong><span>Our results elucidate that viral intrahost genetic diversity is associated with VL and could be used as a better disease progression marker than HIV consensus sequence variants, especially in infections of longer duration. We emphasize that viral intrahost diversity should be considered when studying viral genomes and infection outcomes.</span></p>

opencc-zeroJun 2022View details →
zenodo40/100

Sei whole-genome sequence class annotations

<p>Sei sequence class whole-genome annotations are available in the following files:</p> <ul> <li> <p>sorted.hg38.tiling.bed.ipca_randomized_300.labels.merged.bed - The sorted, merged sequence class assignments from Louvain community clustering of the 30 million sequences, uniformly tiling the whole human genome. The fourth column is the sequence class number, with any sequence classes numbering 40-61 excluded from our analyses in the publication. Sequence classes 0-39 can be mapped to the following labels:&nbsp;<a href="https://github.com/FunctionLab/sei-framework/blob/main/model/seqclass.names">https://github.com/FunctionLab/sei-framework/blob/main/model/seqclass.names</a></p> </li> <li> <p>sorted.hg19.tiling.bed.ipca_randomized_300.labels.merged.bed - lifted over version of the hg38 BED file.</p> </li> </ul>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Data from : Genomic Sequence of Klebsiella pneumoniae IIEMP-3, a Vitamin B12-Producing Strain from Indonesian Tempeh

<p>Klebsiella pneumoniae&nbsp;strain IIEMP-3, isolated from Indonesian tempeh, is a vitamin B<sub>12</sub>-producing strain that exhibited a different genetic profile from pathogenic isolates. Here we report the draft genome sequence of strain IIEMP-3, which may provide insights on the nature of fermentation, nutrition, and immunological function of Indonesian tempeh.</p>

opencc-by-4.0Feb 2016View details →
zenodo40/100

Fig. 3 in Sequencing and analysis of the complete mitochondrial genome of the giant dobsonfly Acanthacorydalis orientalis (McLachlan) (Insecta: Megaloptera: Corydalidae)

Fig. 3. Predicted secondary structure of the rrnl in the Acanthacorydalis orientalis mt genome. Roman numerals denote the conserved Watson-Crick base pairing and dot (•) indicates G-U base pairing.

opencc-by-4.0Dec 2014View details →
zenodo40/100

Fig. 4 in Sequencing and analysis of the complete mitochondrial genome of the giant dobsonfly Acanthacorydalis orientalis (McLachlan) (Insecta: Megaloptera: Corydalidae)

Fig. 4. Predicted secondary structure of the rrns in the A. orientalis mt genome. Roman numerals denote the conserved domain structure. Dash (-) indicates Watson-Crick base pairing and dot (•) indicates G-U base pairing.

opencc-by-4.0Dec 2014View details →
zenodo40/100

Fig. 1 in Sequencing and analysis of the complete mitochondrial genome of the giant dobsonfly Acanthacorydalis orientalis (McLachlan) (Insecta: Megaloptera: Corydalidae)

Fig. 1. Mitochondrial genome map of Acanthacorydalis orientalis. The tRNAs are denoted by the color blocks and are labeled according to the IUPACIUB single-letter amino acid codes. Gene name without underline indicates the direction of transcription

opencc-by-4.0Dec 2014View details →
zenodo40/100

Fig. 2 in Sequencing and analysis of the complete mitochondrial genome of the giant dobsonfly Acanthacorydalis orientalis (McLachlan) (Insecta: Megaloptera: Corydalidae)

Fig. 2. Inferred secondary structure of 22 tRNAs of the Acanthacorydalis orientalis mt genome. The tRNAs are labeled with the abbreviations of their corresponding amino acids. Dash (-) indicates Watson-Crick bonds and dot (·) indicates GU bonds.

opencc-by-4.0Dec 2014View details →
zenodo40/100

Fig. 5 in Sequencing and analysis of the complete mitochondrial genome of the giant dobsonfly Acanthacorydalis orientalis (McLachlan) (Insecta: Megaloptera: Corydalidae)

Fig. 5. Phylogenetic relationships among the sequenced Megaloptera insects. Numbers at the nodes are Bayesian posterior probabilities (left) and ML bootstrap values (right).

opencc-by-4.0Dec 2014View details →
zenodo40/100

Fig. 6 in Complete mitochondrial genome sequence of Bonasa sewerzowi (Galliformes: Phasianidae) and phylogenetic analysis

Fig. 6. The phylogenetic relationship of Bonasa among Galliformes based on the complete mitogenome. Branch lengths and topologies were obtained from Maximum Likelihood analyses. The numbers were the bootstrap values of MP/ML/BI trees in turn. * indicates that MP or BI tree was inconsistent with ML tree.

opencc-by-4.0Dec 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record