Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,574

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,574 results for “genome sequencing”

Learn how ShareScore rates datasets ↗
zenodo28/100

Genomic characterization of metastatic patterns from prospective clinical sequencing of 25,000 patients

<p>Metastatic progression is the main cause of death in cancer patients, whereas the underlying genomic mechanisms driving metastasis remain largely unknown. Here, we assembled MSK-MET, a pan-cancer cohort of over 25,000 patients with metastatic diseases. By analyzing genomic and clinical data from this cohort, we identified associations between genomic alterations and patterns of metastatic dissemination across 50 tumor types. We found that chromosomal instability is strongly correlated with metastatic burden in some tumor types, including prostate adenocarcinoma, lung adenocarcinoma and HR+/HER2+ breast ductal carcinoma, but not in others, including colorectal cancer and high grade serous ovarian cancer, where copy-number alteration patterns may be established early in tumor development. We also identified somatic alterations associated with metastatic burden and specific target organs. Our data offer a valuable resource for the investigation of the biological basis for metastatic spread and highlight the complex role of chromosomal instability in cancer progression.</p>

opencc-by-4.0Jan 2022View details →
dryad28/100

Data from: Diagnostic yield in epileptic encephalopathies is improved by genome sequencing and re-analysis

<p><b>Objective:</b> To assess the benefits and limitations of whole genome sequencing (WGS) compared to exome sequencing (ES) or multigene panel (MGP) in the molecular diagnosis of developmental and epileptic encephalopathies (DEE).</p> <p><b>Methods: </b>We performed WGS of 30 comprehensively phenotyped DEE patient trios that were undiagnosed after first-tier testing, including chromosomal microarray (CMA), and either research ES (n=15) or diagnostic MGP (n=15).</p> <p><b>Results</b>: 8 diagnoses were made in the 15 individuals who received prior ES (53%): 3 individuals had complex structural variants; 5 had ES-detectable variants which now had additional evidence for pathogenicity. 11 diagnoses were made in the 15 MGP-negative individuals (68%); the majority (n=10) involved genes not included in the panel, particularly in individuals with post-neonatal onset of seizures and those with more complex presentations including movement disorders, dysmorphic features and/or multi-organ involvement.  42% of diagnoses were autosomal recessive or X-chromosome linked.</p> <p><span><span><b>Conclusion:</b> WGS was able to improve diagnostic yield over ES primarily through the detection of complex structural variants (n=3). The higher diagnostic yield was otherwise better attributed to the power of re-analysis rather than inherent advantages of the WGS platform. Additional research is required to assist in the assessment of pathogenicity of novel non-coding and complex structural variants and further improve diagnostic yield for patients with DEE and other neurogenetic disorders.</span></span></p>

opencc-zeroJan 2022View details →
zenodo28/100

Draft Genome Sequences of Pseudomonas sp. Strain MWU13-2924, Isolated from a Wild Cranberry Bog in Truro, MA

<p>Annotated genome of Pseudomonas sp. MWU13.2924.</p>

opencc-by-4.0Mar 2022View details →
zenodo28/100

A phased chromosome-level genome and full mitochondrial sequence for the dikaryotic myrtle rust pathogen, Austropuccinia psidii

<p>The fungal plant pathogen <em>Austropuccinia psidii</em> is spreading globally and causing myrtle rust disease symptoms on plants in the family Myrtaceae. <em>A. psidii </em>is dikaryotic, with two nuclei that do not exchange genetic material during the dominant phase of its life-cycle. Phased and scaffolded genome resources for rust fungi are important for understanding heterozygosity, mechanisms of pathogenicity, pathogen population structure and for determining the likelihood of disease spread. We have assembled a chromosome-level phased genome for the pandemic biotype of <em>A. psidii </em>and, for the first time, show that each nucleus contains 18 chromosomes, in line with other distantly related rust fungi. We show synteny between the two haplo-phased genomes and provide a new tool, ChromSyn, that enables efficient comparisons between chromosomes based on conserved genes. Our genome resource includes a fully assembled and circularised mitochondrial sequence for the pandemic biotype. Please cite the following manuscript:&nbsp;https://www.biorxiv.org/content/10.1101/2022.04.22.489119v1</p>

opencc-by-4.0Apr 2022View details →
zenodo28/100

Direct access to millions of mutations by Whole Genome Sequencing of an oilseed rape mutant population

<p>Induced mutations are an essential source of genetic variation in plant breeding. EMS mutagenesis has been frequently applied, and mutants have been detected by phenotypic or genotypic screening of large populations. In this study, a rapeseed M<sub>2</sub>&nbsp;population was derived from M<sub>1</sub>&nbsp;parent cultivar &ldquo;Express&rdquo; treated with EMS. Whole genomes were sequenced from fourfold (4x) pools of 1,988 M<sub>2</sub>&nbsp;plants representing 497 M<sub>2</sub>&nbsp;families. Detected mutations were not evenly distributed and displayed distinct patterns across the 19 chromosomes with lower mutation rates towards the ends. Mutation frequencies ranged from 32/Mb to 48/Mb. On average, 284,442 single nucleotide polymorphisms per M<sub>2</sub>&nbsp;DNA pool were found resulting from EMS mutagenesis. 55% were C&rarr;T and G&rarr;A transitions, characteristic for EMS induced (&lsquo;canonical&rsquo;) mutations, whereas the remaining SNPs were &lsquo;non-canonical&rsquo; transitions (15%) or transversions (30%). Additionally, we detected 88,725 high confidence insertions and deletions (InDels) per pool. On average, each M<sub>2</sub>&nbsp;plant carried 39,120 canonical mutations, corresponding to a frequency of one mutation per 23.6 kb. Roughly 82% of such mutations were located either 5 kb upstream or downstream (~56%) of gene coding regions or within intergenic regions (26%). The remaining 18% were located within regions coding for genes. All mutations detected by whole-genome sequencing could be verified by comparison with known mutations. Furthermore, all sequences are accessible via the online tool &ldquo;EMS Brassica&rdquo; (<a href="http://www.emsbrassica.plantbreeding.uni-kiel.de/">http://www.emsbrassica.plantbreeding.uni-kiel.de/</a>), which enables direct identification of mutations in any target sequence. The sequence resource described here will further add value for functional gene studies in rapeseed breeding.</p>

opencc-by-4.0Dec 2021View details →
dryad28/100

Data from: Next-generation museum genomics: phylogenetic relationships among palpimanoid spiders using sequence capture techniques (Araneae: Palpimanoidea)

Historical museum specimens are invaluable for morphological and taxonomic research, but typically the DNA is degraded making traditional sequencing techniques difficult to impossible for many specimens. Recent advances in Next-Generation Sequencing, specifically target capture, makes use of short fragment sizes typical of degraded DNA, opening up the possibilities for gathering genomic data from museum specimens. This study uses museum specimens and recent target capture sequencing techniques to sequence both Ultra-Conserved Elements (UCE) and exonic regions for lineages that span the modern spiders, Araneomorphae, with a focus on Palpimanoidea. While many previous studies have used target capture techniques on dried museum specimens (for example, skins, pinned insects), this study includes specimens that were collected over the last two decades and stored in 70% ethanol at room temperature. Our findings support the utility of target capture methods for examining deep relationships within Araneomorphae: sequences from both UCE and exonic loci were important for resolving relationships; a monophyletic Palpimanoidea was recovered in many analyses and there was strong support for family and generic-level palpimanoid relationships. Ancestral character state reconstructions reveal that the highly modified carapace observed in mecysmaucheniids and archaeids has evolved independently.

opencc-zeroDec 2017View details →
dryad28/100

Data from: Ribosomal DNA sequence heterogeneity reflects intra-species phylogenies and predicts genome structure in two contrasting yeast species

The ribosomal RNA encapsulates a wealth of evolutionary information, including genetic variation that can be used to discriminate between organisms at a wide range of taxonomic levels. For example, the prokaryotic 16S rDNA sequence is very widely used both in phylogenetic studies and as a marker in metagenomic surveys and the ITS region, frequently used in plant phylogenetics, is now recognised as a fungal DNA barcode. However, this widespread use does not escape criticism, principally due to issues such as difficulties in classification of paralogous versus orthologous rDNA units and intragenomic variation, both of which may be significant barriers to accurate phylogenetic inference. We recently analysed datasets from the Saccharomyces Genome Resequencing Project, characterising rDNA sequence variation within multiple strains of the baker's yeast <i>Saccharomyces cerevisiae</i> and its nearest wild relative <i>Saccharomyces paradoxus</i> in unprecedented detail. Notably, both species possess single locus rDNA systems. Here, we use these new variation datasets to assess whether a more detailed characterisation of the rDNA locus can alleviate the second of these phylogenetic issues, sequence heterogeneity, while controlling for the first. We demonstrate that a strong phylogenetic signal exists within both datasets and illustrate how they can be used, with existing methodology, to estimate intra-species phylogenies of yeast strains consistent with those derived from whole-genome approaches. We also describe the use of partial Single Nucleotide Polymorphisms, a type of sequence variation found only in repetitive genomic regions, in identifying key evolutionary features such as genome hybridisation events and show their consistency with whole-genome Structure analyses. We conclude that our approach can transform rDNA sequence heterogeneity from a problem to a useful source of evolutionary information, enabling the estimation of highly accurate phylogenies of closely related organisms, and discuss how it could be extended to future studies of multi-locus rDNA systems.

opencc-zeroDec 2013View details →
dryad28/100

Data from: "Transcriptome sequences for Campanula gentilis" in Genomic Resources Notes accepted 1 April 2015 – 31 May 2015

In this report, we present the transcriptome of a single accession of Campanula gentilis Kovanda, obtained through the sequencing of both a normalized and a non-normalized cDNA library generated from stem and leaf tissue. The resources we provide include the raw sequence reads, the assembled contigs, the putative open reading frames, the contig/ORF annotations and the normalized as well as non-normalized expression levels.

opencc-zeroDec 2015View details →
dryad28/100

Data from: "RAD Sequencing for SNP Discovery in Two Populations of Bighorn Sheep (Ovis canadensis)" in Genomic Resources Notes accepted 1 April 2013 - 31 May 2013

In this work we present the development of a large set of single nucleotide polymorphisms (SNPs) discovered in two populations of bighorn sheep (Ovis canadensis). To do so we used restriction-site associated DNA (RAD) sequencing of four individuals from each population. Through alignment of reads to the domestic sheep (Ovis aries) genome we discovered &gt;83,000 SNPs, of which &gt;38,000 are suitable for assays such as an Illumina SNP chip. These loci will allow for fine-mapping of loci associated with horn size, and examination of the consequences of the genetic rescue including mapping genes underling differences in life-history characteristics.

opencc-zeroDec 2012View details →
dryad28/100

Data from: "Transcriptome sequencing of the Queensland fruit fly, Bactrocera tryoni (Diptera: Tephritidae)" in Genomic Resources Notes accepted 1 December 2013 to 31 January 2014

[No abstract filled]

opencc-zeroDec 2013View details →
zenodo28/100

Similarity matrix between the core genome sequences

<p>Data table.</p>

opencc-by-4.0Aug 2022View details →
zenodo28/100

Supplementary material 1 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909

Supplementary tables :

opencc-by-4.0Nov 2016View details →
zenodo28/100

Figure 4 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909

Figure 4 - Putative secondary structures of the 22 tRNA genes of Xystodesmus sp. Watson-Crick base-pairing is indicated by solid lines, and G–T pairs are indicated with plus signs.

opencc-by-4.0Nov 2016View details →
zenodo28/100

Figure 3 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909

Figure 3 - Putative secondary structures of the 22 tRNA genes of Asiomorpha coarctata. Watson-Crick base-pairing is indicated by solid lines, and G–T pairs are indicated with plus signs.

opencc-by-4.0Nov 2016View details →
zenodo28/100

Figure 2 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909

Figure 2 - Sequences of the non-coding region in Asiomorpha coarctata, primary structures of tandemly repeated regions (11.4 × 38 bp).

opencc-by-4.0Nov 2016View details →
zenodo28/100

Figure 1 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909

Figure 1 - Mitochondrial genomes of the two millipedes sequenced in this study. A Asiomorpha coarctata B Xystodesmus sp. Circular maps were drawn with Geneious v9.1.2. Arrows indicate the orientation of gene transcription. Abbreviations of gene names are: atp6 and atp8 for ATP synthase subunits 6 and 8; cox1–3 for cytochrome oxidase subunits 1–3; cob for cytochrome b, nad1–6 and nad4L for NADH dehydrogenase subunits 1–6 and 4L; and lrRNA and srRNA for large and small rRNA subunits. tRNA genes are indicated with their one-letter corresponding amino acids. CR for control region. The GC content was plotted using a green sliding window and the AT content was blue.

opencc-by-4.0Nov 2016View details →
zenodo28/100

Figure 6 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909

Figure 6 - Phylogenetic tree of the Arthropoda, including Myriapoda, Hexapoda, Crustacea and Chelicerata and outgroups reconstructed based on protein-coding genes from mtDNA genomes. Each group of four numbers indicates node confidence values (from top left): Bayesian posterior probabilities in percent (BPP) in amino acid and nucleotide datasets; maximum likelihood bootstrapping values (MLBP) in amino acid and nucleotide datasets.

opencc-by-4.0Nov 2016View details →
zenodo28/100

Figure 5 from: Dong Y, Zhu L, Bai Y, Ou Y, Wang C (2016) Complete mitochondrial genomes of two flat-backed millipedes by next-generation sequencing (Diplopoda, Polydesmida). ZooKeys 637: 1-20. https://doi.org/10.3897/zookeys.637.9909

Figure 5 - Comparison of gene arrangements in mtDNA of the arthropod ground pattern. Gene segments are not drawn to scale. Genes shaded gray have different relative positions compared to the ground pattern. Underlining indicates the gene is encoded on the opposite strand, and arrows indicate translocation of trnT. CR: putative control region. Gene arrangements of two diplopods, Narceus annularus and Thyropygus sp. are similar and represented as one.

opencc-by-4.0Nov 2016View details →
zenodo28/100

Figure 1 from: Yuhui X, Lijun Z, Yue H, Xiaoqi W, Chen Z, Huilun Z, Ruoran W, Da P, Hongying S (2017) Complete mitochondrial genomes from two species of Chinese freshwater crabs of the genus Sinopotamon recovered using next-generation sequencing reveal a novel gene order (Brachyura, Potamidae). ZooKeys 705: 41-60. https://doi.org/10.3897/zookeys.705.11852

Figure 1 - Mitochondrial genome sequenced in the present study. Gene order and sizes are shown relative to one another, including non-coding regions. Protein-coding genes encoded on the light strand are underlined. Transfer RNA (tRNA) genes encoded on the light strand are underlined. Each tRNA gene is designated by a single-letter amino acid code, except L1 (trnLeu (CUN)), L2 (trnLeu (UUR)), S1 (trnSer (AGN)) and S2 (trnSer (UCN)). Numbers inside circles represent the size of the non-coding region separating two adjacent genes or the amount of shared nucleotides between two overlapping genes. The translocations of gene or gene block are shaded gray.

opencc-by-4.0Oct 2017View details →
zenodo28/100

Figure 2 from: Yuhui X, Lijun Z, Yue H, Xiaoqi W, Chen Z, Huilun Z, Ruoran W, Da P, Hongying S (2017) Complete mitochondrial genomes from two species of Chinese freshwater crabs of the genus Sinopotamon recovered using next-generation sequencing reveal a novel gene order (Brachyura, Potamidae). ZooKeys 705: 41-60. https://doi.org/10.3897/zookeys.705.11852

Figure 2 - Phylogenetic analyses derived for brachyurans using the maximum likelihood (ML) analyses and Bayesian inferences (BI) using dataset A (13 PCGs) and dataset B (13 PCGs + two rRNAs). Branch lengths and topologies came from ML analysis. Values at the branches represent BP (Bootstrap value)/BPP (Bayesian posterior probability). 100/1.00 is denoted by an asterisk. The horizontal line stands for BP under 50 or BPP under 0.9 ML analyses. The gene rearrangement is denoted by the block on (A): (I) the translocation of trnH shared by the Brachyura taxa sampled; (II) the transposition of trnQ shared by potamid species; (III) the five-gene block, (trnM-nad2-trnW-trnC-trnY), translocation shared by three Sinopotamon crabs sampled.

opencc-by-4.0Oct 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record