Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,574

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,574 results for “genome sequencing”

Learn how ShareScore rates datasets ↗
zenodo32/100

FIGURE 8 in Molecular Systematics of Redband Trout from Genome-Wide DNA Sequencing Substantiates the Description of a New Taxon (Salmonidae: Oncorhynchus mykiss calisulat) from the McCloud River

FIGURE 8. McCloud River Redband Trout, Onchorhynchus mykiss calisulat, ssp. nov., Sheepheaven Creek. A. WFB 5020, holotype, 120 mm SL. B. same specimen as in A, radiograph. C. TCWC 28772.01, paratype, 144 mm SL. D. Illustration of O. m. calisulat, ssp. nov., showing life colors, © J. Tomelleri, used with permission.

opennotspecifiedMar 2023View details →
zenodo32/100

FIGURE 4 in Molecular Systematics of Redband Trout from Genome-Wide DNA Sequencing Substantiates the Description of a New Taxon (Salmonidae: Oncorhynchus mykiss calisulat) from the McCloud River

FIGURE 4. Admixture plots from McCloud river trout and other Redband Trout in population genetics dataset. Admixture results from NGSAdmix for genetic clusters (K) from 2-4 with the subset of samples collected as Redband Trout. Sample size of 204, optimal K = 2. The x-axis labels are labeled according to watershed.

opennotspecifiedMar 2023View details →
dryad32/100

Simulated eukaryotic genomic sequencing, long and short reads

<p><span>As accuracy and throughput of nanopore sequencing improves, it is increasingly common to perform long-read-first </span><em>de novo</em> genome assemblies followed by polishing with accurate short reads (Kim et al. 2021). We briefly introduce FMLRC2, the successor to the original FM-index Long Read Corrector (FMLRC), and illustrate its performance as a fast and accurate <em>de novo</em> assembly polisher for both bacterial and eukaryotic genomes.</p>

opencc-zeroJan 2023View details →
zenodo32/100

Genome sequences for 2031 Saccharomyces cerevisiae genome assemblies

<p>Genome sequences for 2031&nbsp;Saccharomyces cerevisiae genome assemblies</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Gene sequences for 2031 Saccharomyces cerevisiae genome assemblies

<p>Gene sequences for 2031&nbsp;Saccharomyces cerevisiae genome assemblies</p>

opencc-by-4.0Dec 2022View details →
zenodo32/100

Protein sequences for 2031 Saccharomyces cerevisiae genome assemblies

<p>Protein sequences for 2031&nbsp;Saccharomyces cerevisiae genome assemblies</p>

opencc-by-4.0Dec 2022View details →
dryad32/100

Complete organelle genomes of Korean fir, Abies koreana and phylogenomics of the gymnosperm genus Abies using nuclear and cytoplasmic DNA sequence data

<span>Background</span> <p><em><span>Abies koreana</span></em><span> E. H. Wilson is an endangered evergreen coniferous tree that is native to high altitudes in South Korea and susceptible to the effects of climate change. Hybridization and reticulate evolution have been reported in the genus; therefore, multigene datasets from nuclear and cytoplasmic genomes are needed to better understand its evolutionary history.</span></p> <span>Results</span> <p><span>Using Illumina NovaSeq6000 and Oxford Nanopore Technologies (ONT) PromethION platforms, we generated complete mitochondrial (1,174,803 bp) and plastid (121,341 bp) genomes from <em>A. koreana</em>. The mitochondrial genome is highly dynamic, transitioning from cis- to trans-splicing and breaking the conserved gene clusters. In the case of the plastome, the ONT reads revealed two structural conformations of <em>A. koreana</em>. The short inverted repeats (1,186 bp) of the <em>A. koreana</em> plastome are associated with the different structural types. Transcriptomic sequencing revealed 1,356 sites of C-to-U RNA editing in the 41 mitochondrial genes. Using <em>A. koreana</em> as a reference, we additionally produced nuclear ribosomal DNA and organelle genomic sequences from eight Abies species and generated multiple datasets for maximum likelihood and network analyses. Three sections (<em>Balsamea</em>, <em>Momi</em>, and <em>Pseudopicea</em>) were well grouped in the nuclear phylogeny, but the phylogenomic relationships showed conflicting signals in the mitochondrial and plastid genomes, indicating a complicated evolutionary history that may have included introgressive hybridization.</span></p> <span>Conclusions</span> <p><span>These results illustrate that phylogenomic analyses based on the sequences from differently inherited organelle genomes resulted in conflicting trees. Organellar capture, organellar genome recombination, and incomplete lineage sorting in an ancestral heteroplasmic individual can contribute to phylogenomic discordance. We provide strong support for the relationships within <em>Abies</em> and new insights into the phylogenomic complexity of this genus.</span></p>

opencc-zeroApr 2023View details →
zenodo32/100

Genome Alignment of Cancer Sequencing Data

<p>Part of the GTN Cancer Analysis learning Pathway based on the Bioinformatics.ca Cancer Workshop</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Annotated sequences extracted from bacterial genomes

<p>Three files containing sequences extracted from 1,049,210&nbsp;bacterial genomes available from GenBank (release 252). Protein coding sequences were annotated with IDTAXA (PMID:&nbsp;34541527) using taxon-specific KEGG groups (Bacteria_Protein_subset.fas.gz). These annotations were transferred to their corresponding (nucleotide) coding sequences (Bacteria_Nucleotide_subset.fas.gz). Intergenic regions were extracted from each genome and annotated by FindNonCoding (PMID:&nbsp;34636849) for their overlap with any&nbsp;of 25 common bacterial non-coding RNAs in Rfam (v14). Intergenic regions were required to be at least 100 nucleotides long and contain no ambiguities (Bacteria_Intergenic_subset.fas.gz). Each subset contains only distinct sequences&nbsp;randomly ordered.</p> <p><strong><em>Headers</em></strong></p> <p>Sequence headers contain the assembly accession followed by the annotation and separated by a &quot;|&quot; character. For example:</p> <p><strong>Bacteria_Intergenic_subset.fas.gz</strong></p> <p>&gt;GCA_022121725.1|RF00000<br> ATGTTACCTTCTTGAGTGATACGGGATGAA[...]</p> <p><strong>Bacteria_Protein_subset.fas.gz</strong></p> <p>&gt;GCA_014764685.1|K02049<br> MPRDLIRISGLEKTYADGSVHALSNIDLSIKD[...]</p> <p><strong>Bacteria_Nucleotide_subset.fas.gz</strong></p> <p>&gt;GCA_015948525.1|K02197<br> GTGAACCTGCGACGTAAAAACCGGCTAYG[...]</p> <p><strong><em>Annotations</em></strong></p> <p>Protein and protein coding sequences are labeled with their KEGG group, starting with &quot;K&quot;. Intergenic sequences are named by any overlapping Rfam families, starting with a &quot;RF&quot;, and separated by commas when multiple are predicted. &quot;RF00000&quot; is a placeholder for the absence of any predicted RF families.</p>

opencc-by-4.0May 2023View details →
zenodo32/100

Figure 3 in Higher-level phylogeny and evolutionary history of nonditrysians (Lepidoptera) inferred from mitochondrial genome sequences

Figure 3. Heterogeneity of mitogenome composition for different datasets: 13PCG, 13PCG + 2RNA, 13PCG_ codon12 + 2RNA, 13PCG_AA and 13PCG_codon3. The pairwise Aliscore values range from −1 indicating full random similarity, to +1 indicating nonrandom similarity. Species names are listed on top and on the right side of the matrix and are colour-coded to match their member superfamilies (lower right corner).

opennotspecifiedMay 2023View details →
zenodo32/100

Figure 4 in Higher-level phylogeny and evolutionary history of nonditrysians (Lepidoptera) inferred from mitochondrial genome sequences

Figure 4. Mitochondrial gene rearrangements in nonditrysia. Gene sizes are not drawn to scale. Abbreviations of gene names are as follows: ATP6 and ATP8, ATP synthase subunits 6 and 8; COI–COIII, cytochrome c oxidase subunits 1–3; Cytb, cytochrome b; ND1–6 and ND4L, NADH dehydrogenase subunits 1–6 and 4L; 16S and 12S, large and small rRNA subunits. tRNA genes are indicated by their one-letter corresponding amino acids. CR, control region/A + T-rich region. Genes are transcribed from left to right except for those that are underlined, which have the opposite transcriptional orientation. Rearrangements of tRNA genes are highlighted by colours (green: gene inversion; blue and orange: gene rearrangements).

opennotspecifiedMay 2023View details →
zenodo32/100

Figure 2 in Higher-level phylogeny and evolutionary history of nonditrysians (Lepidoptera) inferred from mitochondrial genome sequences

Figure 2. Scatter plot of AT- and GC-skews in the nonditrysian mitogenomes. Values were calculated for the majority strand of the entire mitogenome sequences. All the species are listed in the Supporting Information, Table S2. The legend indicates nonditrysian families and their corresponding taxa numbers. AT-skew = (A-T)/(A + T); GC-skew = (G-C)/(G + C).

opennotspecifiedMay 2023View details →
zenodo32/100

Figure 6 in Higher-level phylogeny and evolutionary history of nonditrysians (Lepidoptera) inferred from mitochondrial genome sequences

Figure 6. Chronogram showing nonditrysian phylogeny, divergence time estimation and ancestral state reconstruction of the tRNA gene arrangement. Phylogenetic tree presenting divergence dates produced by the Bayesian method of the 13PCG dataset using three fossil calibration points (grey star targets). Blue bars indicate the 95% mean confidence interval (CI) of each node. A geological timescale is shown at the top. Branch lengths are measured in Myr. The colour of pie and block charts represent different tRNA gene arrangements in the gene clusters MIQ and TP, respectively.

opennotspecifiedMay 2023View details →
zenodo32/100

Figure 5 in Higher-level phylogeny and evolutionary history of nonditrysians (Lepidoptera) inferred from mitochondrial genome sequences

Figure 5. Bayesian inference and the maximum likelihood estimate of phylogenetic relationships among nonditrysian Lepidoptera inferred from the combined molecular dataset (PCGRNA, 64 taxa). Numbers on each node from left to right correspond to the Bayesian posterior probability values and the bootstrap percentage values of ML analysis, respectively. '-' indicates support values &lt;0.5/50 or missing. All habitus photographs are taken by CQ Liao, except photographs of c and g taken by S Yagi, with scientific names as follows: a, Vietomartyria aeuyunjiena Liao, Hirowatari &amp; Huang, 2020 (Micropterigidae); b, Neopseustis fanjingshana Yang, 1988 (Neopseustidae); c, Eriocrania komaii Mizukawa, Hirowatari &amp; Hashimoto, 2006 (Eriocraniidae); d, Ogygioses maoershana Liao, Hirowatari &amp; Huang, 2021 (Palaeosetidae); e, Stigmella sp. (Nepticulidae); f, Nemophora fluorites (Meyrick, 1907) (Adelidae); g, Tischeria decidua Wocke, 1876 (Tischeriidae); h, GibboƲalƲa kobusi Kumata &amp; Kuroko, 1988 (Gracillariidae); i, Lethe helle Leech, 1891 (Nymphalidae)

opennotspecifiedMay 2023View details →
zenodo32/100

Figure 1 in Higher-level phylogeny and evolutionary history of nonditrysians (Lepidoptera) inferred from mitochondrial genome sequences

Figure 1. Previous hypotheses on relationships among nonditrysian lineages. A, most parsimonious tree for combined 18S rDNA plus morphological characters from Wiegmann et al. (2002). B, synopsis of relationships inferred from morphology by Kristensen et al. (2007). C, maximum likelihood tree of the nonditrysian portion based on eight genes and 350 taxa from Mutanen et al. (2010). D, relationships among nonditrysian superfamilies inferred from 19 genes and 86 taxa by Regier et al. (2015). E, maximum likelihood tree of over 500 morphological characters and eight molecular genes of 473 taxa combined by Heikkilä et al. (2015). F, maximum likelihood tree of phylogenetic relationships among nonditrysian lineages estimated on phylotranscriptomics data of 28 taxa by Bazinet et al. (2017). G, estimated phylogeny of nonditrysian superfamilies of Lepidoptera synthesized from multiple previous studies by Mitter et al. (2017). H, evolutionary tree derived from a maximum-likelihood analysis of 749 791 amino acid sites from transcriptomes of 186 species by Kawahara et al. (2019). I, phylogenetic tree inferred using 1835 CDS nucleotides of 172 taxa by Mayer et al. (2021). Thicker lines indicate better supported groupings.

opennotspecifiedMay 2023View details →
dryad32/100

Supporting data for: Whole genome sequencing reveals fine-scale environment associated divergence near the range limits of a temperate reef fish

<p>Environmental variation is increasingly recognized as an important driver of diversity in marine species despite the lack of physical barriers to dispersal and the presence of pelagic stages in many taxa. A robust understanding of the genomic and ecological processes involved in structuring populations is lacking for most marine species, often hindering management and conservation action. Cunner (<em>Tautogolabrus adspersus</em>), is a temperate reef fish with both pelagic early life history stages and strong site-associated homing as adults; the species is also of interest for use as a cleaner fish in salmonid aquaculture in Atlantic Canada. We aimed to characterize genomic and geographic differentiation of cunner in the Northwest Atlantic. To achieve this, a chromosome-level genome assembly for cunner was produced and used to characterize spatial population structure throughout Atlantic Canada using whole genome resequencing. The genome assembly spanned 0.72 Gbp and 24 chromosomes; whole genome resequencing of 803 individuals from 20 locations from Newfoundland to New Jersey identified approximately 11 million genetic variants. Principal component analysis revealed four regional Atlantic Canadian groups. Pairwise F<sub>ST</sub> and selection scans revealed signals of differentiation and selection at discrete genomic regions, including adjacent peaks on chromosome 10 across multiple pairwise comparisons (<em>i.e.</em>, F<sub>ST</sub> 0.5–0.75). Redundancy analysis suggested association of environmental variables related to benthic temperature and oxygen range with genomic structure. Results suggest regional scale diversity in this temperate reef fish and can directly inform the collection and translocation of cunner for aquaculture applications and the conservation of wild populations throughout the Northwest Atlantic.</p>

opencc-zeroMay 2023View details →
zenodo32/100

Enhancing Protein Sequence Annotation in Viral Genomics Using Large Language Models and Soft Alignments.

<p>List of 200 most abundant VOG descriptions.</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

FIGURE 4 in A new subgenus, Australixodes n. subgen. (Acari: Ixodidae), for the kiwi tick, Ixodes anatis Chilton, 1904, and validation of the subgenus Coxixodes Schulze, 1941 with a phylogeny of 16 of the 22 subgenera of Ixodes Latreille, 1795 from entire mitochondrial genome sequences

FIGURE 4. Ventral view of the gnathosoma of Ixodes (Endopalpiger) barkeri to illustrate the strongly salient (ss) palpal article 1 (I) of the subgenus Endopalpiger (I, palpal article 1). Scale-bar 0.2 mm.

opennotspecifiedAug 2023View details →
zenodo32/100

FIGURE 1 in A new subgenus, Australixodes n. subgen. (Acari: Ixodidae), for the kiwi tick, Ixodes anatis Chilton, 1904, and validation of the subgenus Coxixodes Schulze, 1941 with a phylogeny of 16 of the 22 subgenera of Ixodes Latreille, 1795 from entire mitochondrial genome sequences

FIGURE 1. Mitochondrial genomes of Ixodes (Australixodes) anatis Chilton, 1904 (kiwi tick); Ixodes (Coxixodes) ornithorhynchi Lucas, 1846 (platypus tick); Ixodes (Amerixodes) loricatus Neumann, 1899 (no common name); Ixodes (Ixodes) pacificus Cooley &amp; Kohls, 1943 (no common name); Ixodes (Multidentatus) kohlsi Arthur, 1955 (little penguin Ixodes) and I. (Eschatocephalus) vespertilionis (long-legged bat tick). Protein-coding genes are in green, tRNAs are in yellow, rRNAs are in red whereas the two control regions are in blue. Protein-coding genes are labelled with their four-character abbreviations, tRNAs are labelled with their one-letter amino-acid abbreviations whereas the control regions are labelled as CR1 and CR2. The sizes of the mt genomes are indicated in brackets.

opennotspecifiedAug 2023View details →
zenodo32/100

thus and genetic % 1 than less indicates Green . ) kb 15 . ca ( Ixodes of ) individuals 40 ( species bold 34 in of are genomes study present mitochondrial the in entire sequenced the Species among. differences reference for genetic, species ) % ( same Pairwise the. 3 from FIGURE sequences in A new subgenus, Australixodes n. subgen. (Acari: Ixodidae), for the kiwi tick, Ixodes anatis Chilton, 1904, and validation of the subgenus Coxixodes Schulze, 1941 with a phylogeny of 16 of the 22 subgenera of Ixodes Latreille, 1795 from entire mitochondrial genome sequences

thus and genetic % 1 than less indicates Green . ) kb 15 . ca ( Ixodes of ) individuals 40 ( species bold 34 in of are genomes study present mitochondrial the in entire sequenced the Species among. differences reference for genetic, species ) % ( same Pairwise the. 3 from FIGURE sequences

opennotspecifiedAug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record