Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,574

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,574 results for “genome sequencing”

Learn how ShareScore rates datasets ↗
dryad32/100

Mitochondrial genome sequencing of marine leukemias reveals cancer contagion between clam species in the Seas of Southern Europe

Open the record for dataset details and reuse information.

publicDec 2021View details →
dryad32/100

Sequences of Staudtia kamerunensis obtained through low coverage whole genome skimming

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad32/100

House mouse Mus musculus dispersal in East Eurasia inferred from 98 newly determined complete mitochondrial genome sequences

Open the record for dataset details and reuse information.

publicAug 2020View details →
zenodo28/100

Supplementary material 1 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820

Table S1

opencc-zeroApr 2020View details →
zenodo28/100

Figure 2 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820

Figure 2 Variation in length and base composition of each of the 13 core protein coding genes (PCGs) among eight centipedes' mitochondrial genomes A PCG length variation B GC content across PCGsC AT skew D GC skew.

opencc-by-4.0Apr 2020View details →
zenodo28/100

Figure 4 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820

Figure 4 A Molecular phylogeny of eight centipede species based on Maximum Likelihood inference analysis of 13 protein-coding genes (PCGs) B Traditional morphological classification based on the position of spiracles and the variation of larvae.

opencc-by-4.0Apr 2020View details →
zenodo28/100

Figure 1 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820

Figure 1 Mitochondrial genome map of the Scolopendra mutilans. Genes drawn inside the circle are transcribed clockwise, and those outside are counterclockwise. PCGs are shown as brown arrows, rRNA genes as green arrows, tRNA genes as pink arrows. The innermost circle shows the GC content. GC content is plotted as the deviation from the average value of the entire sequence.

opencc-by-4.0Apr 2020View details →
zenodo28/100

Figure 3 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820

Figure 3 Mitogenome synteny among eight centipede species. Synteny analyses were generated in Mauve 2.4.0. A total of six large homologous regions were identified among the eight mitogenomes, while the sizes and relative positions of the homologous fragments varied across the mitogenomes.

opencc-by-4.0Apr 2020View details →
zenodo28/100

Evaluation SIHUMI dataset: A sectioning and database enrichment approach for improved peptide spectrum matching in large, genome-guided protein sequence databases

<p>Dataset for&nbsp;evaluation of&nbsp;database sectioning method for generating an enriched database for mass-spectrometry-based proteomic approaches&nbsp;using large databases. Our evaluation demonstrates that this method helps to&nbsp;increase the sensitivity of PSMs while maintaining acceptable FDR statistics. This dataset was acquired from the protein samples containing proteins from eight microorganisms (<em>Anaerostipes caccae, Bacteroides thetaiotaomicron, Bifidobacterium longum, Blautia producta, Clostridium butyricum, Clostridium ramosum, Escherichia coli, Lactobacillus plantarum</em>) that were grown in a bioreactor. The MS-data for this dataset was acquired by Dr. Hettich&#39;s group at Oak Ridge National Laboratory, and it was made available by Dr. Nico Jehmlich from&nbsp;Helmholtz Center for Environmental Research through&nbsp;the 3rd International Metaproteome Symposium (<a href="https://www.ufz.de/index.php?en=44639">https://www.ufz.de/index.php?en=44639</a>).&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo28/100

Simulated data and results from "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data"

<p>This dataset contains all the simulated data and the results of all the considered methods in the benchmark presented in &quot;Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data&quot; [Zaccaria &amp; Raphael, 2018]. All the data in this dataset and the corresponding formats are fully described at&nbsp;<a href="https://github.com/raphael-group/hatchet-paper">https://github.com/raphael-group/hatchet-paper</a>. The folder <em>simulations</em>&nbsp;which contains the entire dataset has been compressed with standard <em>zip</em>.</p>

opencc-by-4.0May 2020View details →
zenodo28/100

Figure 3 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460

Figure 3 Undated 18S–26S nuclear DNA repeat region BEAST 2 phylogeny of Pennantia, under the Birth-Death model. The tree was rooted to make P. cunninghamii sister to the other species of Pennantia, in accordance with the chloroplast DNA tree and the ITS tree of Keeling et al. (2004). Node posterior probability is shown next to the corresponding node. The sequences downloaded from GenBank have their accession number in round brackets; the others were generated from the samples used in this study.

opencc-by-4.0Aug 2020View details →
zenodo28/100

Supplementary material 2 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460

BEAST2 and RAxML files

opencc-zeroAug 2020View details →
zenodo28/100

Figure 2 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460

Figure 2 Dated chloroplast DNA BEAST 2 phylogeny of Pennantia, under the Birth-Death model. Mean node age and 95% HPD (in My) is given in the table embedded in the figure under the corresponding letter code. 95% HPD is also represented by blue bars. All node posterior probabilities are equal to 1 except if indicated otherwise. The calibrated nodes (see text) are indicated by red dots.

opencc-by-4.0Aug 2020View details →
zenodo28/100

Supplementary material 1 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460

Figs S1–S5; Tables S1–S3

opencc-zeroAug 2020View details →
zenodo28/100

Figure 1 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460

Figure 1 General distribution of the four Pennantia species. TKI = Three Kings Islands. Generated in QGIS 3.0.1 from Google Satellite data obtained through the XYZ Tiles tool (https://mt1.google.com/vt/lyrs=s&amp;x={x}&amp;y={y}&amp;z={z}).

opencc-by-4.0Aug 2020View details →
dryad28/100

Pooled whole genome sequencing from year 2004 and Early-Late SNP data from year 2014

<p>Speciation underlies the generation of novel biodiversity. Yet, there is much to learn about how natural selection shapes genomes during speciation. Selection is assumed to act against gene flow at barrier loci, promoting reproductive isolation. However, evidence for gene flow and selection is often indirect and we know very little about the temporal stability of barrier loci. Here we utilize haplodiploidy to identify candidate male barrier loci in hybrids between two wood ant species. As ant males are haploid they are expected to reveal recessive barrier loci, which can be masked in diploid females if heterozygous. We then test for barrier stability in a sample collected ten years later and use survival analysis to provide a direct measure of natural selection acting on candidate male barrier loci. We find multiple candidate male barrier loci scattered throughout the genome. Surprisingly, a proportion of them are not stable after ten years, natural selection apparently switching from acting against to favoring introgression in the later sample. Instability of barrier effect and natural selection for introgressed alleles could be due to environment-dependent selection, emphasizing the need to consider temporal variation in the strength of natural selection and the stability of barrier effect at putative barrier loci in future speciation work.</p>

opencc-zeroAug 2020View details →
zenodo28/100

PacBio amplicon re-sequencing of 62 P. tricornutum genomic loci – processed datasets

<p>Processed datasets from&nbsp;PacBio amplicon re-sequencing of 62 P. tricornutum genomic loci.&nbsp;Raw data are available at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA658511.&nbsp;</p> <p>Available datasets:&nbsp;</p> <p>- reference file for regions selected for amplicon sequencing: <em>Phaeodactylum_tricornutum_amplicon_sequencing_loci.fa</em></p> <p>- final .bam files containing processed PacBio sequencing reads&nbsp;aligned to the reference:</p> <p><em>PacBio_amplicon_seq_T1.bam</em> &nbsp;and&nbsp;<em>PacBio_amplicon_seq_T6.bam</em>&nbsp;</p> <p>- . table files with the position, reference and alternative allele for reliable biallelic SNPs selected in ILLUMINA sequencing of the culture at T1:&nbsp;</p> <p><em>P_tricornutum_PacBio_amplicon_sequencing_T1_SNPs.table</em> and&nbsp;<em>P_tricornutum_PacBio_amplicon_sequencing_T6_SNPs.table</em></p> <p>Re-sequencing of 62 endogenous P. tricornutum loci selected in a genome-wide analysis of haplotype diversity. The goal was to determine the number of haplotypes per locus and the appearance of new haplotypes over time. The length of the sequenced loci was 2kb (+/- 5%).&nbsp; Loci were amplified by emulsion PCR on the same culture harvested in two time points: five loci were amplified one month (T1) and all loci were amplified 6 months (T6) after the start of the culture from a single cell. Plasmids containing cloned GFP or YFP were amplified separately as a control for random errors. Control reactions for artificial haplotypes detection consisted of mixed CFP with YFP or CFP with GFP. Amplicons were pooled together into two samples. Sample PacBio_AS_T1 contained five P. tricornutum endogenous amplicons from DNA harvested at T1 time point, GFP amplified separately and CFP+YFP amplified in one reaction. Sample PacBio_AS_T6 contained 63 P. tricornutum endogenous amplicons from DNA harvested at T6 time point, YFP amplified separately and CFP+GFP amplified in one reaction. Samples were mixed in 1:9 PacBio_AS_T1: PacBio_AS_T6 ratio before sequencing on one PacBio Sequel SMRT cell.</p>

opencc-by-4.0Aug 2020View details →
zenodo28/100

Genome sequencing of P. tricornutum mother and daughter cultures derived from single cell and separated by 30 days of proliferation - processed datasets

<p><strong>Genome sequencing of <em>P. tricornutum</em> mother and daughter cultures derived from single cell and separated by 30 days of proliferation - processed datasets.</strong></p> <p>Raw data&nbsp;for this experiment&nbsp;are available at&nbsp;https://www.ncbi.nlm.nih.gov/bioproject/PRJNA658224.</p> <p>&nbsp;</p> <p><strong>Please note that the naming of files differs from the general description on&nbsp;/www.ncbi.nlm.nih.gov/bioproject website and in related publication:</strong></p> <p>Instead of MC1-3, the processed datasets are labelled Sc1-3</p> <p>Instead of DC1.1; DC1.2 and DC1.3, the processed datasets are labelled Sc11, Sc12 and Sc14 respectively</p> <p>Instead of DC2.1; DC2.2 and DC2.3, the processed datasets are labelled Sc21, Sc22 and Sc24 respectively</p> <p>Instead of DC3.1; DC3.2 and DC3.3, the processed datasets are labelled Sc31, Sc32 and Sc33&nbsp;respectively</p> <p><strong>Available datasets:&nbsp;</strong></p> <p><em>.bam</em> files with ILLUMINA reads aligned to the reference P. tricornutum v2 genome&nbsp;used for SNP calling&nbsp;</p> <p><em>.vcf</em> files for individual samples with SNPs called using GATK3.7.0</p> <p><em>joint_genotyping_cohort.vcf</em> file with SNPs called jointly for all samples using&nbsp;GATK4.2.1&nbsp;</p> <p>&nbsp;</p> <p><strong>Description of the experiment:</strong>&nbsp;</p> <p>Whole-genome Illumina sequencing of mother and daughter cultures derived from single cell to reveal genomic changes occurring within 30 day time frame. Three independent single cells were isolated from CCAP 1055/1 culture (sample label: Pt1) to start&nbsp;mother cultures (MC1; MC2; MC3 . On day 30 after mother culture isolation (T1 time point), three daughter cells were isolated from each mother culture forming cultures DC11-DC33. Part of mother cultures and CCAP 1055/1 culture were harvested at T1 (Samples: Pt1T1; MC1T1; MC2T1; MC3T1) . After another 30 days (T2 time point), all cultures were harvested (Samples: Pt1T2; mother culture MC1T2 and respective daughter cultures DC11, DC12 and DC13 ; mother culture MC2T2 and respective daughter cultures DC21, DC22 and DC23; mother culture MC3T1 and respective daughter cultures DC31, DC32, DC33).</p>

opencc-by-4.0Aug 2020View details →
zenodo28/100

Simulated pairs of nucleotide sequences for testing (alignment-free) genome distance estimate methods

<p>This repository contains 24,000 pairs of nucleotide sequences (and associated parameters) that have been simulated for testing alignment-free genome distance estimates. Given an evolutionary distance <em>d</em> varying from 0.05 to 1.00 nucleotide substitutions per character (step = 0.05), the program <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> was used to simulate the evolution of 200 nucleotide sequence pairs with <em>d</em> substitution events per character under the models GTR and GTR+&Gamma;. Each model was adjusted with three different equilibrium frequencies:</p> <ul> <li><em>f</em><sub>1</sub>: equal frequencies, i.e. freq(A) = freq(C) = freq(G) = freq(T) = 0.25,</li> <li><em>f</em><sub>2</sub>: GC-rich, i.e. freq(A) = 0.1, freq(C) = 0.3, freq(G) = 0.4, freq(T) = 0.2,</li> <li><em>f</em><sub>3</sub>: AT-rich, i.e. freq(A) = freq(T) = 0.4, freq(C) = freq(G) = 0.1.</li> </ul> <p>For each simulated sequence pair, model parameters (i.e. GTR: six relative rates of nucleotide substitution; GTR+&Gamma;: six rates and one &Gamma; shape parameter) were randomly drawn from 142 sets of parameters derived from real-case data (see file <a href="https://zenodo.org/record/4034261/files/GTR.params.trees.tsv?download=1">GTR.params.trees.tsv</a> at <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p>&nbsp;</p> <p>For each of the 20 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 1.00, six XZ-compressed files containing 200 simulation data are available:</p> <ul> <li><code>data-d-f1-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f1-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f2-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f2-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f3-nogam.tsv.xz</code> &nbsp; data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>3</sub></li> <li><code>data-d-f3-gamma.tsv.xz</code> &nbsp; data simulated under the model GTR+&Gamma; with equilibrium frequencies <em>f</em><sub>3</sub></li> </ul> <p>&nbsp;</p> <p>Each file is tab-delimited and contains the 18 following fields:</p> <ul> <li><code>[1]&nbsp; &nbsp;</code>&nbsp;&nbsp; integer <em>seed</em> value specified to <em>INDELible</em>,</li> <li><code>[2-5]&nbsp;</code> &nbsp; frequencies of T, C, A, G, respectively, specified to <em>INDELible</em>,</li> <li><code>[6-10]&nbsp;</code> C-T, A-T, G-T, A-C, C-G rate parameters, respectivly (normalized such that A-G rate = 1), specified to <em>INDELible</em>,</li> <li><code>[11] &nbsp; </code> &nbsp; &Gamma; shape parameter <em>alpha</em> (= 0 in the <code>nogam</code> files, i.e. GTR substitution model without &Gamma;) specified to <em>INDELible</em>,</li> <li><code>[12] &nbsp; </code> &nbsp; length <em>lgt1</em> of the first sequence <em>seq1</em> (i.e. no. A, C, G, T in <em>seq1</em>),</li> <li><code>[13] &nbsp; </code> &nbsp; length <em>lgt2</em> of the second sequence <em>seq2</em> (i.e. no. A, C, G, T in <em>seq2</em>),</li> <li><code>[14] &nbsp; </code> &nbsp; no. <em>sites</em> in aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. A, C, G, T and gap character states in <em>seq1</em> or <em>seq2</em>),</li> <li><code>[15] &nbsp; </code> &nbsp; no. non-gapped sites (<em>core</em> sites) in aligned sequences <em>seq1</em> and <em>seq2</em>,</li> <li><code>[16] &nbsp; </code> &nbsp; observed <em>p-distance</em> between aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. nucleotide mismatches divided by no. <em>core</em> sites),</li> <li><code>[17] &nbsp; </code> &nbsp; aligned <em>seq1</em> (containing indel gaps),</li> <li><code>[18] &nbsp; </code> &nbsp; aligned <em>seq2</em> (containing indel gaps).</li> </ul> <p>Of note, <em>seq1</em> and <em>seq2</em> (fields <code>[17-18]</code>) being aligned, these two entries are two strings with identical no. <em>sites</em> (field <code>[14]</code>). Gap character states (<code>-</code>) should be removed from <em>seq1</em> and <em>seq2</em> to obtain the unaligned sequences.</p> <p>_____</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>

opencc-by-4.0Sep 2020View details →
zenodo28/100

The Genome sequences of Calonectria ilicicola (anamorph Cylindrocladium parasiticum) causing Cylindrocladium black rot of peanut and red crown rot of soybean

<p>The fungus <em>Calonectria ilicicola</em> (anamorph <em>Cylindrocladium parasiticum</em>) is an important plant pathogen causing Cylindrocladium black rot (CBR) on peanut and Red crown rot (RCR) on soybean. CBR infection of peanuts cause symptoms such as chlorosis of leaves, blackening of taproots, and wilting, while RCR infection of soybeans lead to root and interveinal necrosis of soybean. In the present study, we sequenced the genome of four CBR-related<em> Ca. ilicicola</em> strains and three RCR-related<em> Ca. ilicicola</em> strains. The draft genome of <em>Ca. ilicicola</em>, ranged from 68.63 Mb to 70.35 Mb in genome size, containing 18536 to 19361 protein-coding genes. The described genome sequences will provide insights into factors that contribute to pathogenicity toward peanut and soybean and will be useful for future research in population genomics and molecular diagnostic marker development to quickly detect this pathogen.</p>

opencc-by-4.0Oct 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record