Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Whole genome sequence and annotation of Penstemon davidsonii
<p><em><span>Penstemon</span></em><span> is the most speciose flowering plant genus endemic to North America. <em>Penstemon</em> species' diverse morphology and adaptation to various environments have made them a valuable model system for studying evolution, but the absence of publicly available reference genomes limits possible research directions. Here we report the first reference genome assembly and annotation for <em>Penstemon</em> <em>davidsonii</em>. Using PacBio long-read sequencing and Hi-C scaffolding technology, we constructed a de novo reference genome of 437,568,744 bases, with a contig N50 of 40 Mb and L50 of 5. The annotation includes 18,199 gene models, and both the genome and transcriptome assembly contain over 95% complete eudicot BUSCOs. This genome assembly will serve as a valuable reference for studying the evolutionary history and genetic diversity of the <em>Penstemon</em> genus.</span></p>
Genome sequencing of three Orestias species and the analysis of their phylogenetic relationships within the Cyprinodontiformes order
<p>Relevant files and datasets for the paper: <em>"Genomes of the Orestias pupfish from the Andean Altiplano shed light on their evolutionary history and phylogenetic relationships within Cyprinodontiformes." </em>by Morales <em>et al</em>.</p>
Genomic sequences of 343 Xanthomonas citri pv. citri strains from Florida
<p><i>Xanthomonas citri </i>pv.<i> citri </i>(Xcc) causes the devastating citrus canker disease. Xcc is known to have been introduced into Florida, USA in at least three different events in 1915, 1986 and 1995 with the first two claimed to be eradicated. Here, we investigated the population structure of Xcc to understand the mysteries related to its introduction, spread and eradication, and how Xcc strains have evolved in term of pathogenicity and copper resistance. We sequenced whole genome sequences of 343 Xcc strains collected from Florida groves over 20 years. </p>
Data from: Improved genome assembly of the whiteleg shrimp Penaeus (Litopenaeus) vannamei using long- and short-read sequences from public databases
Open the record for dataset details and reuse information.
Data from: Using transcriptome sequencing and pooled exome capture to study local adaptation in the giga-genome of Pinus cembra
Open the record for dataset details and reuse information.
Whole genome sequence and annotation of Penstemon davidsonii
Open the record for dataset details and reuse information.
Whole genome sequence data of Mycobacterium tuberculosis and Mycolicibacterium smegmatis mutants of the riboflavin biosynthetic pathway- Part 1
Open the record for dataset details and reuse information.
Genome report: Genome sequence of 1S1, a transformable and highly regenerable diploid potato for use as a model for gene editing and genetic engineering
Open the record for dataset details and reuse information.
Deep-sequencing of viral genomes from treatment-naive HIV-infected persons shows positive association between intrahost genetic diversity and viral load
Open the record for dataset details and reuse information.
Data corresponding to: Evaluation of sequencing and PCR-based methods for the quantification of the viral genome formula
Open the record for dataset details and reuse information.
Genome-wide sequence data show no evidence of hybridization and introgression among pollinator wasps associated with a community of Panamanian strangler figs
Open the record for dataset details and reuse information.
Genomic characterization and gene bank curation of Aegilops using genotyping-by-sequencing
Open the record for dataset details and reuse information.
Whole genome sequence data of Mycobacterium tuberculosis and Mycolicibacterium smegmatis mutants of the riboflavin biosynthetic pathway- Part 2
Open the record for dataset details and reuse information.
Whole-genome sequencing reveals asymmetric introgression between two sister species of cold-resistant leaf beetles
Open the record for dataset details and reuse information.
Whole genome sequences of 23 species from the Drosophila montium species group (Diptera: Drosophilidae): a resource for testing evolutionary hypotheses
Open the record for dataset details and reuse information.
Genome sequence and characterization of a freshwater photoarsenotroph, Cereibacter azotoformans strain ORIO, isolated from sediments capable of cyclic light-dark arsenic oxidation and reduction
Open the record for dataset details and reuse information.
Chloroplast genome sequencing reads from snow gum
<p>Tutorial data for chloroplast genome assembly: fastq reads from illumina and nanopore sequencing for the snow gum, <em>Eucalyptus pauciflora</em>.</p> <p>Data from: Wang, W., Schalamun, M., Morales-Suarez, A. et al. Assembly of chloroplast genomes with long- and short-read data: a comparison of approaches using Eucalyptus pauciflora as a test case. BMC Genomics 19, 977 (2018) doi:10.1186/s12864-018-5348-8</p> <p>Data hosted at NCBI under accession numbers: illumina (SRR7153063) and nanopore (SRR7153095). Additional illumina file SRR7153071 not used here. </p> <p>This is how the files have been changed from the original datasets: </p> <p>Using the Galaxy platform (usegalaxy.org): </p> <ul> <li> <p>Each dataset was separately mapped to the NCBI Reference Sequence for <em>Eucalyptus pauciflora </em>chloroplast NC_039597.1, using BWA-MEM. </p> </li> <li> <p>Unmapped reads were filtered out using a SAMtools flag. </p> </li> <li> <p>Bam files were converted to fastq files.</p> </li> <li> <p>Each fastq file was then reduced in size:</p> </li> <li> <p>snow-gum-illumina-cp-reduced: has the first 62,500 reads only. Note that original pairing of reads has not been preserved so consider these to be unpaired reads for this tutorial.</p> </li> <li> <p>snow-gum-nanopore-cp-reduced: has only reads that are longer than 90,000 bp.</p> </li> </ul>
Sex-linked markers by genome-wide RAD sequencing to identify XX/XY Sex Chromosomes in the spiny frog (Quasipaa boulengeri)
<p><span>We use genotyping by sequencing as an approach to identify sex-linked markers in the spiny frog <i>Quasipaa boulengeri</i> with 43 wild-collected adults from a single site. The GBS methodology identified 2 loci on sex differences in allele frequencies, 50 loci on sex differences in heterozygosity, and 523 loci on male-limited occurrence, altogether associated with males heterogamety, indicating an XX-XY system. The sex specificity of five markers was further validated by PCR amplification with a large number of additional individuals from 26 various populations in this species. A total of 27 sex linkage markers were matched to Dmrt1 gene, a ubiquitous role in sex determination and differentiation from flies and nematodes to mammals. Chromosome 1, that harboring Dmrt1, has further been assigned to a highly potential candidate sex chromosome in anurans. Five sex-linked SNP makers explored 3 sex reversals out of 133 individuals here, sparsely showing sex reversal detected in wild amphibian populations. </span></p>
A genomic data set of single‐nucleotide polymorphisms (SNPs) generated by ddRAD tag sequencing in Q. petraea (Matt.) Liebl. populations from Central-Eastern Europe and Balkan Peninsula
<p>This genomic dataset provides highly variable single-nucleotide polymorphism (SNP) markers from georeferenced natural <em>Quercus petraea</em> (Matt.) Liebl. populations collected in Bulgaria, Hungary, Romania, Serbia, Bosnia and Herzegovina, Kosovo and Albania. These SNP loci can be used to assess genetic diversity, differentiation, population structure, and can also be used to detect signatures of selection and local adaptation.</p>
Supplementary datasets for: Large-scale genome sequencing reveals the driving forces of viruses in microalgal evolution
<p>Microalgae are integral primary producers for global ecosystems whose genomes can be mined for ecological insights, but representative genome sequences are lacking for many phyla. We cultured and sequenced 107 microalgae species from 11 different phyla indigenous to varied geographies and climates. This genome collection was used to resolve genomic differences between saltwater and freshwater microalgae. Freshwater species showed domain-centric ontology enrichment for nuclear and nuclear membrane functions, while saltwater species were enriched in organellar and cellular membrane functions. Marine species contained significantly more viral families in their genomes (<span>p-value = 8 x 10(-4))</span>. Viral sequences were identified from Chlorovirus, Coccolithovirus, Pandoravirus, Marseillevirus, Tupanvirus, and others integrated into algal genomes. Algal, viral-origin sequences were found to be expressed and to code for a wide variety of functions. Our results clarify the poorly characterized occurrences of viral elements in algal genomes and define a unified adaptive strategy for algal halotolerance.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.