Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Mitochondrial genome sequencing of marine leukemias reveals cancer contagion between clam species in the Seas of Southern Europe
Open the record for dataset details and reuse information.
Sequences of Staudtia kamerunensis obtained through low coverage whole genome skimming
Open the record for dataset details and reuse information.
House mouse Mus musculus dispersal in East Eurasia inferred from 98 newly determined complete mitochondrial genome sequences
Open the record for dataset details and reuse information.
Supplementary material 1 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820
Table S1
Figure 2 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820
Figure 2 Variation in length and base composition of each of the 13 core protein coding genes (PCGs) among eight centipedes' mitochondrial genomes A PCG length variation B GC content across PCGsC AT skew D GC skew.
Figure 4 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820
Figure 4 A Molecular phylogeny of eight centipede species based on Maximum Likelihood inference analysis of 13 protein-coding genes (PCGs) B Traditional morphological classification based on the position of spiracles and the variation of larvae.
Figure 1 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820
Figure 1 Mitochondrial genome map of the Scolopendra mutilans. Genes drawn inside the circle are transcribed clockwise, and those outside are counterclockwise. PCGs are shown as brown arrows, rRNA genes as green arrows, tRNA genes as pink arrows. The innermost circle shows the GC content. GC content is plotted as the deviation from the average value of the entire sequence.
Figure 3 from: Hu C, Wang S, Huang B, Liu H, Xu L, Hu Z, Liu Y (2020) The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes. ZooKeys 925: 73-88. https://doi.org/10.3897/zookeys.925.47820
Figure 3 Mitogenome synteny among eight centipede species. Synteny analyses were generated in Mauve 2.4.0. A total of six large homologous regions were identified among the eight mitogenomes, while the sizes and relative positions of the homologous fragments varied across the mitogenomes.
Evaluation SIHUMI dataset: A sectioning and database enrichment approach for improved peptide spectrum matching in large, genome-guided protein sequence databases
<p>Dataset for evaluation of database sectioning method for generating an enriched database for mass-spectrometry-based proteomic approaches using large databases. Our evaluation demonstrates that this method helps to increase the sensitivity of PSMs while maintaining acceptable FDR statistics. This dataset was acquired from the protein samples containing proteins from eight microorganisms (<em>Anaerostipes caccae, Bacteroides thetaiotaomicron, Bifidobacterium longum, Blautia producta, Clostridium butyricum, Clostridium ramosum, Escherichia coli, Lactobacillus plantarum</em>) that were grown in a bioreactor. The MS-data for this dataset was acquired by Dr. Hettich's group at Oak Ridge National Laboratory, and it was made available by Dr. Nico Jehmlich from Helmholtz Center for Environmental Research through the 3rd International Metaproteome Symposium (<a href="https://www.ufz.de/index.php?en=44639">https://www.ufz.de/index.php?en=44639</a>). </p>
Simulated data and results from "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data"
<p>This dataset contains all the simulated data and the results of all the considered methods in the benchmark presented in "Accurate quantification of copy-number aberrations and whole-genome duplications in multi-sample tumor sequencing data" [Zaccaria & Raphael, 2018]. All the data in this dataset and the corresponding formats are fully described at <a href="https://github.com/raphael-group/hatchet-paper">https://github.com/raphael-group/hatchet-paper</a>. The folder <em>simulations</em> which contains the entire dataset has been compressed with standard <em>zip</em>.</p>
Figure 3 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460
Figure 3 Undated 18S–26S nuclear DNA repeat region BEAST 2 phylogeny of Pennantia, under the Birth-Death model. The tree was rooted to make P. cunninghamii sister to the other species of Pennantia, in accordance with the chloroplast DNA tree and the ITS tree of Keeling et al. (2004). Node posterior probability is shown next to the corresponding node. The sequences downloaded from GenBank have their accession number in round brackets; the others were generated from the samples used in this study.
Supplementary material 2 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460
BEAST2 and RAxML files
Figure 2 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460
Figure 2 Dated chloroplast DNA BEAST 2 phylogeny of Pennantia, under the Birth-Death model. Mean node age and 95% HPD (in My) is given in the table embedded in the figure under the corresponding letter code. 95% HPD is also represented by blue bars. All node posterior probabilities are equal to 1 except if indicated otherwise. The calibrated nodes (see text) are indicated by red dots.
Supplementary material 1 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460
Figs S1–S5; Tables S1–S3
Figure 1 from: Maurin KJL (2020) A dated phylogeny of the genus Pennantia (Pennantiaceae) based on whole chloroplast genome and nuclear ribosomal 18S–26S repeat region sequences. PhytoKeys 155: 15-32. https://doi.org/10.3897/phytokeys.155.53460
Figure 1 General distribution of the four Pennantia species. TKI = Three Kings Islands. Generated in QGIS 3.0.1 from Google Satellite data obtained through the XYZ Tiles tool (https://mt1.google.com/vt/lyrs=s&x={x}&y={y}&z={z}).
Pooled whole genome sequencing from year 2004 and Early-Late SNP data from year 2014
<p>Speciation underlies the generation of novel biodiversity. Yet, there is much to learn about how natural selection shapes genomes during speciation. Selection is assumed to act against gene flow at barrier loci, promoting reproductive isolation. However, evidence for gene flow and selection is often indirect and we know very little about the temporal stability of barrier loci. Here we utilize haplodiploidy to identify candidate male barrier loci in hybrids between two wood ant species. As ant males are haploid they are expected to reveal recessive barrier loci, which can be masked in diploid females if heterozygous. We then test for barrier stability in a sample collected ten years later and use survival analysis to provide a direct measure of natural selection acting on candidate male barrier loci. We find multiple candidate male barrier loci scattered throughout the genome. Surprisingly, a proportion of them are not stable after ten years, natural selection apparently switching from acting against to favoring introgression in the later sample. Instability of barrier effect and natural selection for introgressed alleles could be due to environment-dependent selection, emphasizing the need to consider temporal variation in the strength of natural selection and the stability of barrier effect at putative barrier loci in future speciation work.</p>
PacBio amplicon re-sequencing of 62 P. tricornutum genomic loci – processed datasets
<p>Processed datasets from PacBio amplicon re-sequencing of 62 P. tricornutum genomic loci. Raw data are available at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA658511. </p> <p>Available datasets: </p> <p>- reference file for regions selected for amplicon sequencing: <em>Phaeodactylum_tricornutum_amplicon_sequencing_loci.fa</em></p> <p>- final .bam files containing processed PacBio sequencing reads aligned to the reference:</p> <p><em>PacBio_amplicon_seq_T1.bam</em> and <em>PacBio_amplicon_seq_T6.bam</em> </p> <p>- . table files with the position, reference and alternative allele for reliable biallelic SNPs selected in ILLUMINA sequencing of the culture at T1: </p> <p><em>P_tricornutum_PacBio_amplicon_sequencing_T1_SNPs.table</em> and <em>P_tricornutum_PacBio_amplicon_sequencing_T6_SNPs.table</em></p> <p>Re-sequencing of 62 endogenous P. tricornutum loci selected in a genome-wide analysis of haplotype diversity. The goal was to determine the number of haplotypes per locus and the appearance of new haplotypes over time. The length of the sequenced loci was 2kb (+/- 5%). Loci were amplified by emulsion PCR on the same culture harvested in two time points: five loci were amplified one month (T1) and all loci were amplified 6 months (T6) after the start of the culture from a single cell. Plasmids containing cloned GFP or YFP were amplified separately as a control for random errors. Control reactions for artificial haplotypes detection consisted of mixed CFP with YFP or CFP with GFP. Amplicons were pooled together into two samples. Sample PacBio_AS_T1 contained five P. tricornutum endogenous amplicons from DNA harvested at T1 time point, GFP amplified separately and CFP+YFP amplified in one reaction. Sample PacBio_AS_T6 contained 63 P. tricornutum endogenous amplicons from DNA harvested at T6 time point, YFP amplified separately and CFP+GFP amplified in one reaction. Samples were mixed in 1:9 PacBio_AS_T1: PacBio_AS_T6 ratio before sequencing on one PacBio Sequel SMRT cell.</p>
Genome sequencing of P. tricornutum mother and daughter cultures derived from single cell and separated by 30 days of proliferation - processed datasets
<p><strong>Genome sequencing of <em>P. tricornutum</em> mother and daughter cultures derived from single cell and separated by 30 days of proliferation - processed datasets.</strong></p> <p>Raw data for this experiment are available at https://www.ncbi.nlm.nih.gov/bioproject/PRJNA658224.</p> <p> </p> <p><strong>Please note that the naming of files differs from the general description on /www.ncbi.nlm.nih.gov/bioproject website and in related publication:</strong></p> <p>Instead of MC1-3, the processed datasets are labelled Sc1-3</p> <p>Instead of DC1.1; DC1.2 and DC1.3, the processed datasets are labelled Sc11, Sc12 and Sc14 respectively</p> <p>Instead of DC2.1; DC2.2 and DC2.3, the processed datasets are labelled Sc21, Sc22 and Sc24 respectively</p> <p>Instead of DC3.1; DC3.2 and DC3.3, the processed datasets are labelled Sc31, Sc32 and Sc33 respectively</p> <p><strong>Available datasets: </strong></p> <p><em>.bam</em> files with ILLUMINA reads aligned to the reference P. tricornutum v2 genome used for SNP calling </p> <p><em>.vcf</em> files for individual samples with SNPs called using GATK3.7.0</p> <p><em>joint_genotyping_cohort.vcf</em> file with SNPs called jointly for all samples using GATK4.2.1 </p> <p> </p> <p><strong>Description of the experiment:</strong> </p> <p>Whole-genome Illumina sequencing of mother and daughter cultures derived from single cell to reveal genomic changes occurring within 30 day time frame. Three independent single cells were isolated from CCAP 1055/1 culture (sample label: Pt1) to start mother cultures (MC1; MC2; MC3 . On day 30 after mother culture isolation (T1 time point), three daughter cells were isolated from each mother culture forming cultures DC11-DC33. Part of mother cultures and CCAP 1055/1 culture were harvested at T1 (Samples: Pt1T1; MC1T1; MC2T1; MC3T1) . After another 30 days (T2 time point), all cultures were harvested (Samples: Pt1T2; mother culture MC1T2 and respective daughter cultures DC11, DC12 and DC13 ; mother culture MC2T2 and respective daughter cultures DC21, DC22 and DC23; mother culture MC3T1 and respective daughter cultures DC31, DC32, DC33).</p>
Simulated pairs of nucleotide sequences for testing (alignment-free) genome distance estimate methods
<p>This repository contains 24,000 pairs of nucleotide sequences (and associated parameters) that have been simulated for testing alignment-free genome distance estimates. Given an evolutionary distance <em>d</em> varying from 0.05 to 1.00 nucleotide substitutions per character (step = 0.05), the program <a href="http://abacus.gene.ucl.ac.uk/software/indelible/"><em>INDELible</em></a> was used to simulate the evolution of 200 nucleotide sequence pairs with <em>d</em> substitution events per character under the models GTR and GTR+Γ. Each model was adjusted with three different equilibrium frequencies:</p> <ul> <li><em>f</em><sub>1</sub>: equal frequencies, i.e. freq(A) = freq(C) = freq(G) = freq(T) = 0.25,</li> <li><em>f</em><sub>2</sub>: GC-rich, i.e. freq(A) = 0.1, freq(C) = 0.3, freq(G) = 0.4, freq(T) = 0.2,</li> <li><em>f</em><sub>3</sub>: AT-rich, i.e. freq(A) = freq(T) = 0.4, freq(C) = freq(G) = 0.1.</li> </ul> <p>For each simulated sequence pair, model parameters (i.e. GTR: six relative rates of nucleotide substitution; GTR+Γ: six rates and one Γ shape parameter) were randomly drawn from 142 sets of parameters derived from real-case data (see file <a href="https://zenodo.org/record/4034261/files/GTR.params.trees.tsv?download=1">GTR.params.trees.tsv</a> at <a href="https://zenodo.org/record/4034261">https://zenodo.org/record/4034261</a>). Initial sequence length was 5 Mbs, and an indel rate of 0.01 was set with indel length drawn from [1, 50000] according to a Zipf distribution with parameter 1.5 (see <em>INDELible</em> <a href="http://abacus.gene.ucl.ac.uk/software/indelible/manual/model.shtml">manual</a>).</p> <p> </p> <p>For each of the 20 evolutionary distances <em>d</em> = 0.05, 0.10, ..., 1.00, six XZ-compressed files containing 200 simulation data are available:</p> <ul> <li><code>data-d-f1-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f1-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>1</sub></li> <li><code>data-d-f2-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f2-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>2</sub></li> <li><code>data-d-f3-nogam.tsv.xz</code> data simulated under the model GTR with equilibrium frequencies <em>f</em><sub>3</sub></li> <li><code>data-d-f3-gamma.tsv.xz</code> data simulated under the model GTR+Γ with equilibrium frequencies <em>f</em><sub>3</sub></li> </ul> <p> </p> <p>Each file is tab-delimited and contains the 18 following fields:</p> <ul> <li><code>[1] </code> integer <em>seed</em> value specified to <em>INDELible</em>,</li> <li><code>[2-5] </code> frequencies of T, C, A, G, respectively, specified to <em>INDELible</em>,</li> <li><code>[6-10] </code> C-T, A-T, G-T, A-C, C-G rate parameters, respectivly (normalized such that A-G rate = 1), specified to <em>INDELible</em>,</li> <li><code>[11] </code> Γ shape parameter <em>alpha</em> (= 0 in the <code>nogam</code> files, i.e. GTR substitution model without Γ) specified to <em>INDELible</em>,</li> <li><code>[12] </code> length <em>lgt1</em> of the first sequence <em>seq1</em> (i.e. no. A, C, G, T in <em>seq1</em>),</li> <li><code>[13] </code> length <em>lgt2</em> of the second sequence <em>seq2</em> (i.e. no. A, C, G, T in <em>seq2</em>),</li> <li><code>[14] </code> no. <em>sites</em> in aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. A, C, G, T and gap character states in <em>seq1</em> or <em>seq2</em>),</li> <li><code>[15] </code> no. non-gapped sites (<em>core</em> sites) in aligned sequences <em>seq1</em> and <em>seq2</em>,</li> <li><code>[16] </code> observed <em>p-distance</em> between aligned sequences <em>seq1</em> and <em>seq2</em> (i.e. no. nucleotide mismatches divided by no. <em>core</em> sites),</li> <li><code>[17] </code> aligned <em>seq1</em> (containing indel gaps),</li> <li><code>[18] </code> aligned <em>seq2</em> (containing indel gaps).</li> </ul> <p>Of note, <em>seq1</em> and <em>seq2</em> (fields <code>[17-18]</code>) being aligned, these two entries are two strings with identical no. <em>sites</em> (field <code>[14]</code>). Gap character states (<code>-</code>) should be removed from <em>seq1</em> and <em>seq2</em> to obtain the unaligned sequences.</p> <p>_____</p> <p>Criscuolo A (2020) <em>On the transformation of MinHash-based uncorrected distances into proper evolutionary distances for phylogenetic inference</em>. F1000Research, 9:1309. <a href="https://doi.org/10.12688/f1000research.26930.1">doi:10.12688/f1000research.26930.1</a></p>
The Genome sequences of Calonectria ilicicola (anamorph Cylindrocladium parasiticum) causing Cylindrocladium black rot of peanut and red crown rot of soybean
<p>The fungus <em>Calonectria ilicicola</em> (anamorph <em>Cylindrocladium parasiticum</em>) is an important plant pathogen causing Cylindrocladium black rot (CBR) on peanut and Red crown rot (RCR) on soybean. CBR infection of peanuts cause symptoms such as chlorosis of leaves, blackening of taproots, and wilting, while RCR infection of soybeans lead to root and interveinal necrosis of soybean. In the present study, we sequenced the genome of four CBR-related<em> Ca. ilicicola</em> strains and three RCR-related<em> Ca. ilicicola</em> strains. The draft genome of <em>Ca. ilicicola</em>, ranged from 68.63 Mb to 70.35 Mb in genome size, containing 18536 to 19361 protein-coding genes. The described genome sequences will provide insights into factors that contribute to pathogenicity toward peanut and soybean and will be useful for future research in population genomics and molecular diagnostic marker development to quickly detect this pathogen.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.