Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
128
datasets available to search
ShareScore release 0.7.1
Dataset results
128 results for “chromosome level”
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
<p><i>Telopea speciosissima, </i>the New South Wales waratah, is an Australian endemic woody shrub in the family Proteaceae. Waratahs have great potential as a model clade to better understand processes of speciation, introgression and adaptation, and are significant from a horticultural perspective. Here, we report the first chromosome-level genome for <i>T. speciosissima</i>. Combining Oxford Nanopore long-reads, 10x Genomics Chromium linked-reads and Hi-C data, the assembly spans 823 Mb (scaffold N50 of 69.0 Mb) with 97.8 % of Embryophyta BUSCOs 'Complete'. We present a new method in Diploidocus (<a href="https://github.com/slimsuite/diploidocus">https://github.com/slimsuite/diploidocus</a>) for classifying, curating and QC-filtering scaffolds, which combines read depths, <i>k</i>-mer frequencies and BUSCO predictions. We also present a new tool, DepthSizer (<a href="https://github.com/slimsuite/depthsizer">https://github.com/slimsuite/depthsizer</a>), for genome size estimation from the read depth of single-copy orthologues and estimate the genome size to be approximately 900 Mb. The largest 11 scaffolds contained 94.1 % of the assembly, conforming to the expected number of chromosomes (2<i>n</i> = 22). Genome annotation predicted 40,158<code> </code>protein-coding genes, 351 rRNAs and 728 tRNAs. We investigated <i>CYCLOIDEA </i>(<i>CYC</i>)<i> </i>genes, which have a role in determination of floral symmetry, and confirm the presence of two copies in the genome. Read depth analysis of 180 'Duplicated' BUSCO genes using a new tool, DepthKopy (<a href="https://github.com/slimsuite/depthkopy">https://github.com/slimsuite/depthkopy</a>), suggests almost all are real duplications, increasing confidence in the annotation and highlighting a possible need to revise the BUSCO set for this lineage. The chromosome-level <i>T. speciosissima</i> reference genome (Tspe_v1) provides an important new genomic resource of Proteaceae to support the conservation of flora in Australia and further afield.</p>
A chromosome-level genome assembly of the highly heterozygous sea urchin Echinometra sp. EZ reveals adaptation in the regulatory regions of stress response genes
<p><em>Echinometra</em> is the most widespread genus of sea urchin and has been the focus of a wide range of studies in ecology, speciation, and reproduction. However, available genetic data for this genus are generally limited to a few select loci. Here, we present a chromosome-level genome assembly based on 10x Genomics, PacBio, and Hi-C sequencing for <em>Echinometra</em> sp. EZ from the Persian/Arabian Gulf. The genome is assembled into 210 scaffolds totaling 817.8 Mb with an N50 of 39.5 Mb. From this assembly we determined that the <em>E</em>. sp. EZ genome consists of 2n = 42 chromosomes. BUSCO analysis showed that 95.3% of BUSCO genes were complete. ab initio and transcript-informed gene modeling and annotation identified 29,<span>405</span> genes, including a conserved Hox cluster. <em>E.</em> sp. EZ can be found in high-temperature and high-salinity environments, and we therefore compared gene families and transcription factors associated with environmental stress response ("defensome") with other echinoid species with similar high-quality genomic resources. While the number of defensome genes was broadly similar for all species, we identified strong signatures of positive selection in non-coding elements near genes involved in environmental response pathways as well as losses of transcriptions factors important for environmental response. These data provide key insights into the biology of <em>E</em>. sp. EZ as well as the diversification of <em>Echinometra</em> more widely and will serve as a useful tool for the community to explore questions in this taxonomic group and beyond.</p>
Chromosome-level assembly and annotation of Lates japonicus
<p>It is known that some endangered species have persisted for thousands of years despite their very small effective population sizes (<em>N</em><sub>e</sub>s) and low levels of genetic polymorphisms. To understand the importance of genome-wide genetic diversity for the long-term persistence of natural populations in threatened species, we determined the whole genome sequences of akame (<em>Lates japonicus</em>), which is considered to have survived a long time with extremely low genetic variations. Genome-wide single nucleotide variant heterozygosity in akame was estimated to be 3.3–3.4 × 10<sup>-4</sup> /bp, one of the smallest values in teleost fishes. Analysis of demographic history inferred that the <em>N</em><sub>e</sub> in akame was around 1,000 from 30,000 years ago to the recent past. However, a detailed analysis of genetic diversity in the akame genome revealed that multiple genomic regions containing genes involved in immunity, synaptic development, and olfactory sensory systems have retained relatively high nucleotide polymorphisms. This implies that the akame genome has preserved the functional genetic variations by balancing selection, to avoid a reduction in viability and loss of adaptive potential in fluctuating environments. Analysis of synonymous and nonsynonymous nucleotide substitution rates has detected signs of positive selection in many akame genes, indicating that adaptive evolution to the temperate waters has occurred after the speciation of akame and its close relative, barramundi (<em>L. calcarifer</em>). Our results indicated that the functional genetic diversity in akame likely contributed to avoiding the harmful effects of the reduced population size, despite the increased genetic load in this species.</p>
Chromosome-level genome assembly of Cyamophila willieti (Hemiptera: Psyllidae)
Open the record for dataset details and reuse information.
Chromosome-level genome assemblies of sunflower oilseed and confectionery cultivars
<p>In this study, we obtained high-quality genomes of two cultivar representatives of oil and confectionery common sunflower (Helianthus annuus L.) lineages in China at the chromosome level using the PacBio Revio system Circular Consensus Sequencing and high-throughput chromatin conformation capture (Hi-C) scaffolding sequencing technologies. OXS is an inbred oil-type sunflower line with high kernel rate (74.5% kernels), contains 42.9% oil and fat. It is highly susceptible to Verticillium wilt and moderately susceptible to several other diseases. YDS is an inbred non-oil line with high plant height (180-220 cm) and large plump seeds (19.23 grams per 100 seeds). The genome assembly of OXS, spans 3.03 Gb, with 99.58% of sequences anchored to 17 chromosomes and a contig N50 length of 154 Mb. Similarly, the assembly size of YDS, is 3.02 Gb, with 99.40% of sequences mapped to 17 chromosomes and a contig N50 length of 153 Mb. The gene completeness of BUSCO reached 98.2% for OXS and 98.4% for YDS, while the LTR Assembly Index (LAI) stood at 24.73 and 25.85 for OXS and YDS, respectively. A comparative genomics approach identified 6,535 (OXS) and 6,498 (YDS) gene families that have evolved rapidly and are associated with substance synthesis, cell growth, grain weight, and protective mechanisms against biotic and abiotic stresses. We discovered that the YDS genome assembly shows high collinearity with the OXS assembly, apart from three significant inversions on chromosomes 7 and 17. We also identified 15,056 large deletions and insertions between the OXS and YDS assemblies. The publication of these genomes has greatly contributed to the improvement of genetic breeding by integrating internal genetic and external environmental factors in Helianthus annuus L. crops.</p>
Chromosome-level Assemblies of Three Candidatus Liberibacter solanacearum Vectors: Dyspersa apicalis (Förster, 1848), Dyspersa pallida (Burckhardt, 1986), and Trioza urticae (Linnaeus, 1758) (Hemiptera: Psylloidea)
<p>Genomic datasets generated from three species of psyllid insect (Hemiptera: Psylloidea). This repository includes chromosome-scale genomic assemblies, mitochondrial genomes, co-assembled bacterial genomes, coding sequence annotations, transposable element annotations, and called SNPs, as well as files related to comparative genomics analyses. </p> <p><strong>Dataset contains:</strong><br><strong>From Trioza urticae genome assembly:</strong><br> - Genome assembly (fasta)<br> - Suspected contaminant seqeunces removed from the genome assembly (fasta)<br> - T. urticae derived Candidatus Carsonella ruddii primary endosymbiont co-assembled genome (fasta)<br> - Transposable element annotations from EarlgreyTE:<br> - - Transpoable element library (fasta)<br> - - Predicted TEs (bed and gff)<br> - - Figures (pdf)<br> - Gene predictions from braker3+ :<br> - - Braker gene predictions (gft and aa) <br> - - Longest isoforms (faa)<br> - - - Interproscan annotation of gene predicitions (tsv)</p> <p><strong>From Dyspersa pallida (Trioza anthrisci) genome assembly:</strong><br> - Genome assembly (fasta)<br> - Suspected contaminant seqeunces removed from the genome assembly (fasta)<br> - D. pallida mitochondrial genome assembly (fasta)<br> - D. pallida derived Candidatus Carsonella ruddii primary endosymbiont co-assembled genome (fasta)<br> - Transposable element annotations from EarlgreyTE:<br> - - Transpoable element library (fasta)<br> - - Predicted TEs (bed and gff)<br> - - Figures (pdf)<br> - Gene predictions from braker3+ :<br> - - Braker gene predictions (gft and aa) <br> - - Longest isoforms (faa)<br> - - - Interproscan annotation of gene predicitions (tsv)</p> <p><strong>From Dyspersa apicalis (Trioza apicalis) genome assembly:</strong><br> - Genome assembly (fasta)<br> - Suspected contaminant seqeunces removed from the genome assembly (fasta)<br> - D. apicalis mitochondrial genome assembly (fasta)<br> - D. apicalis derived Candidatus Carsonella ruddii primary endosymbiont co-assembled genome (fasta)<br> - Transposable element annotations from EarlgreyTE:<br> - - Transpoable element library (fasta)<br> - - Predicted TEs (bed and gff)<br> - - Figures (pdf)<br> - Gene predictions from braker3+ :<br> - - Braker gene predictions (gft and aa) <br> - - Longest isoforms (faa)<br> - - - Interproscan annotation of gene predicitions (tsv)</p> <p><strong>From comparative genomics analysis:</strong><br> - Orthofinder analysis<br> - - Output of orthofinder analysis comparing protein predictions from de novo psyllid assemblies with other hemiptera proteomes (tsv and fasta)<br> - Cafe5 analysis<br> - - Output of cafe analysis comparing protein predictions from de novo psyllid assemblies with other hemiptera proteomes (excel, png, tab)<br> - - Enrichment analysis of GO and KO terms associated with expanded/contracted gene families at the Dyspersa taxonomic node (excel and tiff)<br> - - Enrichment analysis of GO and KO terms associated with expanded/contracted gene families at the D. pallida taxonomic node (excel and tiff)<br> - - Enrichment analysis of GO and KO terms associated with expanded/contracted gene families at the D. apicalis taxonomic node (excel and tiff)<br> - - - Plots showing expansion/contraction of different orthogroups across the hemiptera phylogeny (png)<br> - Time calibrated phylogenetic tree of hemiptera including psyllids produced by iqtree2 (txt)<br> - Time calibrated phylogenetic tree of hemiptera including psyllids produced by astral (txt)<br> - C. Ca ruddii primary endosymbiont phylogenetic tree (txt)</p> <p><strong>From psyllid population resequencing:</strong><br> - Resequencing data<br> - - High confidence biallelic SNPs from D. pallida resequenced samples called against the de novo D. pallida genome assembly (vcf)<br> - - High confidence biallelic SNPs from D. apicalis resequenced samples called against the de novo D. apicalis genome assembly (vcf)<br> - - High confidence biallelic SNPs from resequenced samples called against the reference C. Ca ruddi endosymbiont genome assembly (vcf)<br> - - For suspected contanimant contigs removed from the D. pallida genome assembly; predicted identity, and coverage in each resequenced D. pallida sample (txt)<br> - - For suspected contanimant contigs removed from the D. apicalis genome assembly; predicted identity, and coverage in each resequenced D. apicalis sample (txt)<br> - - - Qualimap evaluation of resequencing data aligned to de novo psyllid genome for each resequenced sample (pdf)<br><br><br></p>
Chromosome-level genome assembly of Pterygoplichthys pardalis reveals its genetic basis of extensive invasion
<p>The catfish, <em>Pterygoplichthys</em> <em>pardalis</em>, which belongs to the Loricariidae family, an invasive species which has caused huge damage to the ecological environment. However, the high-quality reference genome for the catfish has not yet been reported. In this study, we successfully assembled the first chromosome-level high-quality genome of <em>P</em>. <em>pardalis</em> using the data we produced from multiple sequencing platforms, which contains 26 chromosomes and with a scaffold N50 of 49.47 Mb. Different evaluation methods all indicate the high connectivity and accuracy of the <em>P</em>. <em>pardalis</em> genome we got. We predicated 23,859 protein-coding genes in the <em>P</em>. <em>pardalis</em> genome, and 22,169 (~92.92%) coding genes could be functionally annotated in public databases. Phylogenetic relationship analysis found <em>P</em>. <em>pardalis</em> was clustered with all the catfishes we used and diverged with them 132.5 million years ago. Besides, whole-genome collinearity analysis found that chromosome 6 of <em>P</em>. <em>pardalis</em> was aligned to two distinct chromosomes both for <em>Ameiurus</em> <em>melas</em>, <em>Pangasianodon</em> <em>hypophthalmus</em> and <em>Ictalurus</em> <em>punctatus</em>, indicating that there may have been a chromosomal fusion/fission event occurred. Furthermore, many immune-system-related genes were large-scale expanded in <em>P</em>. <em>pardalis</em> genome, which may make great contributions to their adaptive traits, even for the highly polluted environmental conditions, and successful invasion. Taken together, this study not only provides insights into the genetic basis of the successful invasion of <em>P</em>. <em>pardalis</em>, but also provides important data resources for comparative genomic analysis of <em>P</em>. <em>pardalis</em> in Siluriformes in the future.</p>
Chromosome-level assemblies of the Pieris mannii butterfly genome suggest Z-origin and rapid evolution of the W chromosome
<p><span>The insect order Lepidoptera (butterflies and moths) represents the largest group of organisms with ZW/ZZ sex determination. While the origin of the Z chromosome predates the evolution of the Lepidoptera, the W chromosomes are considered younger, but their origin is debated. To shed light on the origin of the lepidopteran W, we here produce chromosome-level genome assemblies for the butterfly <em>Pieris</em> <em>mannii</em>, and compare the sex chromosomes within and between <em>P. mannii </em>and its sister species <em>P. rapae</em>. Our analyses clearly indicate a common origin of the W chromosomes of the two <em>Pieris</em> species, and reveal similarity between the Z and W in chromosome sequence and structure. This supports the view that the W in these species originates from Z-autosome fusion rather than from a redundant B chromosome. We further demonstrate the extremely rapid evolution of the W relative to the other chromosomes and argue that this may preclude reliable conclusions about the origins of W chromosomes based on comparisons among distantly related Lepidoptera. Finally, we find that sequence similarity between the Z and W chromosomes is greatest toward the chromosome ends, perhaps reflecting selection for the maintenance of recognition sites essential to chromosome segregation. Our study highlights the utility of long-read sequencing technology for illuminating chromosome evolution.</span></p>
Tspe_v1 (Telopea speciosissima) genome supplementary files for: Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C
Open the record for dataset details and reuse information.
Data from: Chromosome-level genome of the melon thrips yields insights into evolution of a sap-sucking lifestyle and pesticide resistance
Open the record for dataset details and reuse information.
Chromosome-level assemblies of the Pieris mannii butterfly genome suggest Z-origin and rapid evolution of the W chromosome
Open the record for dataset details and reuse information.
Data from: A prelude to conservation genomics: First chromosome-level genome assembly of a flying squirrel (Pteromyini: Pteromys volans)
Open the record for dataset details and reuse information.
Data supporting: Chromosome-level genome of the transformable northern wattle, Acacia crassicarpa
Open the record for dataset details and reuse information.
A chromosome-level genome assembly of the snow leopard, Panthera uncia
Open the record for dataset details and reuse information.
Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis
Open the record for dataset details and reuse information.
Chromosome-level genome assembly of Poropuntius huangchuchieni
Open the record for dataset details and reuse information.
Chromosome-level genome assembly of the bay scallop Argopecten irradians
Open the record for dataset details and reuse information.
Chromosome-level genome assembly of Dynastes reidi reveals structural evolution of autosomes and the sex chromosomes in Hercules Beetles
Open the record for dataset details and reuse information.
Chromosome-level genome of the peach fruit moth Carposina sasakii (Lepidoptera: Carposinidae) provides a resource for evolutionary studies on moths
Open the record for dataset details and reuse information.
Chromosome-level assembly and annotation of Lates japonicus
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.