Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

128

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

128 results for “chromosome level”

Learn how ShareScore rates datasets ↗
zenodo44/100

A chromosome-level genome resource for studying virulence mechanisms and evolution of the coffee rust pathogen Hemileia vastatrix

<p>Recurrent epidemics of coffee leaf rust, caused by the fungal pathogen <em>Hemileia vastatrix,</em> have constrained the sustainable production of Arabica coffee for over 150 years. The ability of <em>H. vastatrix </em>to overcome resistance in coffee cultivars and evolve new races is inexplicable for a pathogen that supposedly only utilizes clonal reproduction. Understanding the evolutionary complexity between <em>H. vastatrix</em> and its only known host, including determining how the pathogen evolves virulence so rapidly is crucial for disease management. Achieving such goals relies on the availability of a comprehensive and high-quality genome reference assembly. To date, two reference genomes have been assembled and published for <em>H. vastatrix</em> that, while useful, remain fragmented and do not represent chromosomal scaffolds. Here, we present a complete scaffolded pseudochromosome-level genome resource for <em>H. vastatrix </em>strain 178a (Hv178a). Our initial assembly revealed an unusually high degree of gene duplication (over 50% BUSCO basidiomycota_odb10 genes). Upon inspection, this was predominantly due to a single scaffold that itself showed 91.9% BUSCO Completeness. Taxonomic analysis of predicted BUSCO genes placed this scaffold in Exobasidiomycetes and suggests it is a distinct genome, which we have named Hv178a associated fungal genome (Hv178a AFG). The high depth of coverage and close association with Hv178a raises the prospect of symbiosis, although we cannot completely rule out contamination at this time. The main Ca. 546 Mbp Hv178a genome was primarily (97.7%) localised to 11 pseudochromosomes (51.5 Mb N50), building the foundation for future advanced studies of genome structure and organization. Citation:&nbsp;https://doi.org/10.1101/2022.07.29.502101</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

A chromosome-level genome assembly of the woolly apple aphid, Eriosoma lanigerum (Hausman) (Hemiptera: Aphididae)

<p><strong><em>Eriosoma lanigerum</em> v1.0 frozen release</strong></p> <p>Genome assembly: Eriosoma_lanigerum.v1.0.scaffolds.fa.gz</p> <p>BRAKER2 gene models: Eriosoma_lanigerum.v1.0.scaffolds.gff</p> <p>BRAKER2 protein sequences: Eriosoma_lanigerum.v1.0.scaffolds.gff.aa.fa</p> <p>BRAKER2 protein sequences (longest transcript per gene only): Eriosoma_lanigerum.v1.0.scaffolds.gff.aa.LTPG.fa</p> <p>BRAKER2 coding sequences: Eriosoma_lanigerum.v1.0.scaffolds.gff.cds.fa</p> <p><em>Buchnera aphidicola</em>&nbsp;scaffolds:&nbsp;Buchnera_aphidicola.scaffolds.fa</p> <p><strong>Aphid&nbsp;orthogroups</strong></p> <p>OrthoFinder&nbsp;run files (see for details&nbsp;<a href="https://github.com/davidemms/OrthoFinder/blob/master/OrthoFinder-manual.pdf">https://github.com/davidemms/OrthoFinder/blob/master/OrthoFinder-manual.pdf</a>):&nbsp;OrthoFinder_run.tar.gz</p>

opencc-by-4.0May 2020View details →
dryad40/100

Chromosomal-level genome assembly of the scimitar‐horned oryx: insights into diversity and demography of a species extinct in the wild

<p>Captive populations provide a valuable insurance against extinctions in the wild. However, they are also vulnerable to the negative impacts of inbreeding, selection and drift. Genetic information is therefore considered a critical aspect of conservation management. Recent developments in sequencing technologies have the potential to improve the outcomes of management programmes; however, the transfer of these approaches to applied conservation has been slow. The scimitar‐horned oryx (<i>Oryx dammah)</i> is a North African antelope that has been extinct in the wild since the early 1980s and is the focus of a large‐scale and long‐term reintroduction project. To enable the selection of suitable founder individuals, facilitate post‐release monitoring and improve captive breeding management, comprehensive genomic resources are required. Here, we used 10X Chromium sequencing together with Hi‐C contact mapping to develop a chromosomal‐level genome assembly for the species. The resulting assembly contained 29 chromosomes with a scaffold N50 of 100.4 Mb, and displayed strong chromosomal synteny with the cattle genome. Using resequencing data from six additional individuals, we demonstrated relatively high genetic diversity in the scimitar‐horned oryx compared to other mammals, despite it having experienced a strong founding event in captivity. Additionally, the level of diversity across populations varied according to management strategy. Finally, we uncovered a dynamic demographic history that coincided with periods of climate variation during the Pleistocene. Overall, our study provides a clear example of how genomic data can uncover valuable insights into captive populations and contributes important resources to guide future management decisions of an endangered species.</p>

opencc-zeroJun 2020View details →
zenodo40/100

Sequencing a botanical monument: a chromosome-level assembly of the 400-year-old Goethe's Palm (Chamaerops humilis L.) at the Botanical Garden of the University of Padua (Italy)

<p>The enclosed data pertains to the genome assemblies of the mitochondrion (final_mitogenome.fasta) and the plastid (plastid_genome.fasta) of the dwarf palm <em>Chamaerops humilis</em> L.</p> <p><strong><em>Please refer to the published paper for further details.</em></strong></p>

opencc-zeroOct 2024View details →
zenodo40/100

Chromosome-level genome assembly of a living fossil, the Atlantic Horseshoe Crab Limulus polyphemus

<p>Associated data for male Atlantic horseshoe crab <em>Limulus polyphemus&nbsp;</em>chromosome-scale genome and annotations, including genome (FASTA), structural gene annotations (GFF3), functional annotations (TSV), coding sequences (CDS), protein sequences (PEP), RepeatModeler library (FA.CLASSIFIED), and repeat annotations (OUT).</p> <p>qaLimPoly3.1 - Publication analyses were completed with this genome.&nbsp;</p> <p>qaLimPoly3.3 - This is the current reference assembly. Assembly updated with Sanger sequencing based edits of Chr11 and removal of adapter contamination. Gene annotations updated with curation of canonical proclotting genes. Repeat annotations updated with curation of repeat elements ltr-1_family-1, ltr-1_family-4, and ltr-1_family-26.&nbsp;</p> <p>HSC_Genomic_FacC_860F_PREMIX_CNNJ42_1.ab1 and HSC_Genomic_FacC_1544R_PREMIX_CNNJ43_2.ab1 are Sanger sequenced PCR products for Lp_g42129 (Factor C) from primers FacC_860F.fasta and FacC_1544R.fasta.</p>

opencc-by-4.0Aug 2024View details →
dryad40/100

Data from: A de novo chromosome-level genome assembly of Coregonus sp. "Balchen": one representative of the Swiss Alpine whitefish radiation

<p>Salmonids are of particular interest to evolutionary biologists due to their incredible diversity of life-history strategies and the speed at which many salmonid species have diversified. In Switzerland alone, over 30 species of Alpine whitefish from the subfamily Coregoninae have evolved since the last glacial maximum, with species exhibiting a diverse range of morphological and behavioural phenotypes. This, combined with the whole genome duplication which occurred in the ancestor of all salmonids, makes the Alpine whitefish radiation a particularly interesting system in which to study the genetic basis of adaptation and speciation and the impacts of ploidy changes and subsequent rediploidization on genome evolution. Although well curated genome assemblies exist for many species within Salmonidae, genomic resources for the subfamily Coregoninae are lacking. To assemble a whitefish reference genome, we carried out PacBio sequencing from one wild-caught <i>Coregonus sp. "Balchen" </i>from Lake Thun to ~90x coverage. PacBio reads were assembled independently using three different assemblers, Falcon, Canu and wtdbg2 and subsequently scaffolded with additional Hi-C data. All three assemblies were highly contiguous, had strong synteny to a previously published <i>Coregonus</i>linkage map, and when mapping additional short-read data to each of the assemblies, coverage was fairly even across most chromosome-scale scaffolds. Here, we present the first <i>de novo</i>genome assembly for the Salmonid subfamily Coregoninae. The final 2.2 Gb wtdbg2 assembly included 40 scaffolds, an N50 of 51.9 Mb, and was 93.3% complete for BUSCOs. The assembly consisted of ~52% TEs and contained 44,525 genes.</p>

opencc-zeroMay 2020View details →
zenodo40/100

Additional annotation, alignment, and results from Ka/Ks analysis for Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus)

<p><strong>Annotation files, alignments, and results summaries from&nbsp;Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus).</strong></p> <p>Pairwise genome alignments contain the .maf suffix</p> <p>FASTA alignments from stitched gene blocks&nbsp;contain the .fasta suffix</p> <p>CSV file containing the Ka/Ks results</p> <p>RepeatMasker .out file</p>

opencc-by-4.0Mar 2023View details →
dryad40/100

Supplementary data for: Chromosome-level genome assembly and circadian gene repertoire of the Patagonia blennie Eleginops maclovinus

<p>This dataset contains the genome assembly and associated annotation of the Patagonian Blennie (<em>Eleginops maclovinus</em>), the closest extant taxon to the Antarctic notothenioid radiation. In addition to the characterization of the <em>E. maclovinus </em>genome, the dataset includes a description of circadian rhythm orthologs for <em>E. maclovinus</em>, other notothenenioid taxa, and teleost outgroups, as well as a copy of the bioinformatic scripts used for the assembly, annotation, and other downstream analysis.</p>

opencc-zeroMay 2023View details →
zenodo40/100

Near-Chromosomal-Level Genome of the Red Palm Weevil (Rhynchophorus ferrugineus), a Potential Resource for Genome-Based Pest Control.

<p>Red palm weevil genome annotation data set</p>

opencc-by-4.0Sep 2023View details →
dryad40/100

Chromosome-level genome assembly and annotation of the emblematic silver-lipped pearl oyster, Pinctada maxima Jameson, 1901

Open the record for dataset details and reuse information.

publicJun 2025View details →
dryad40/100

Data from: High-resolution chromosome-level genome of Scylla paramamosain provides molecular insights into adaptive evolution in crab

Open the record for dataset details and reuse information.

publicNov 2024View details →
dryad40/100

A chromosome-level genome assembly of the beavertail cactus, Opuntia basilaris

Open the record for dataset details and reuse information.

publicJun 2025View details →
dryad40/100

Chromosomal-level genome assembly of the scimitar‐horned oryx: insights into diversity and demography of a species extinct in the wild

Open the record for dataset details and reuse information.

publicJun 2020View details →
dryad40/100

Supplementary data for: Chromosome-level genome assembly and circadian gene repertoire of the Patagonia blennie Eleginops maclovinus

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

Data from: A phased chromosome-level genome of the annelid tubeworm <em>Galeolaria caespitosa</em>

Open the record for dataset details and reuse information.

publicSep 2025View details →
dryad40/100

Data from: A de novo chromosome-level genome assembly of Coregonus sp. “Balchen”: one representative of the Swiss Alpine whitefish radiation

Open the record for dataset details and reuse information.

publicMay 2020View details →
dryad36/100

Data from: Chromosome-level genome of the melon thrips yields insights into evolution of a sap-sucking lifestyle and pesticide resistance

<p>Thrips are tiny insects from the order Thysanoptera (Hexapoda: Condylognatha), including many sap-sucking pests that are causing increasing damage to crops worldwide. In contrast to their closest relatives of Hemiptera (Hexapoda: Condylognatha) including numerous sap-sucking species, there are few genomic resources available for thrips. In this study, we assembled the first thrips genome at the chromosome level from the melon thrips, <i>Thrips palmi</i>, a notorious pest in agriculture, using PacBio long-read and Illumina short-read sequences. The assembled genome was 270.43 Mb in size with 4,120 contigs and a contig N50 of 426 kb. All contigs were assembled into 16 linkage groups assisted by the Hi-C technique. In total, 16,333 protein-coding genes were predicted, of which 88.13% were functionally annotated. Among sap-sucking insects, polyphagous species usually possess more detoxification genes than oligophagous species. The polyphagous thrips genomes characterized so far have relatively more detoxification genes in the GST and CCE families than polyphagous aphids, but they have fewer UGTs. HSP genes, especially from the Hsp70s group, have expanded in thrips compared to other hemipteran insects. These differences point to different genetic mechanisms associated with detoxification and stress responses in these two groups of sap-sucking insects. The expansion of these gene families may contribute to the rapid development of pesticide resistance in thrips, as supported by a transcriptome comparison of resistant and sensitive populations of <i>T. palmi</i>. The high-quality genome developed here provides an invaluable resource for understanding the ecology, genetics and evolution of thrips as well as their relatives more generally.</p>

opencc-zeroJun 2020View details →
dryad36/100

Chromosome-level genome assembly of Poropuntius huangchuchieni

<p><i>Poropuntius huangchuchieni</i> is a diploid species in the family cyprinid, widely distributed in Mekong and Red River basins. Previous study suggested that it is one of the most closely related diploid ancestral species to common carp, which has allotetraploidized genome generated by merging two diploid genomes during evolution. Therefore, <i>P. huangchuchieni</i> is an ideal diploid model for polyploid evolution study in Cyprinidae. Here, we report a high-quality chromosome-level genome assembly of <i>P. huangchuchieni</i> by the integrating of the Oxford Nanopore Technology and Hi-C technology. The assembled genome size was 1021.38 Mb with a scaffold N50 of 32.93 Mb. More than 47.61% of the genome was identified as repetitive elements, and 895.66 Mb sequences were anchored onto 25 chromosomes. Of the 24,099 predicted protein-coding genes, 97.57% were functional annotated. Approximately 95.9% of complete BUSCOs were detected in the genome.The high-quality genomic data of <i>P. huangchuchieni</i> provides an ancestral diploid reference for the evolution and adaptation of allotetraploid carps.</p>

opencc-zeroDec 2019View details →
dryad36/100

Chromosome-level genome of the peach fruit moth Carposina sasakii (Lepidoptera: Carposinidae) provides a resource for evolutionary studies on moths

<p>Here we provide scripts and parameters for genome assembly and annotation, as well as the manually annotated circadian genes of <i>period</i> (PER), <i>timeless</i> (TIM), <i>Clock</i> (CLK), <i>cycle</i> (CYC) and cryptochrome (CRY), five detoxification gene families of cytochrome P450 monooxygenase (P450s), glutathione S-transferase (GSTs), carboxyl/cholinesterases (CCEs), UDP-glycosyltransferases (UGTs) and ATP-binding cassette (ABC) transporters, IR, OR, OBP, GR genes from the genome of <span class="fontstyle01"><span>the peach fruit moth (PFM), </span></span><span class="fontstyle01"><span><i>Carposina sasakii</i></span></span><span class="fontstyle01"><span> Matsumura (Lepidoptera: Carposinidae, superfamily Copromorphoidea) and genomes of its related species.</span></span></p>

opencc-zeroOct 2020View details →
dryad36/100

Data supporting: Chromosome-level genome of the transformable northern wattle, Acacia crassicarpa

<p>The genus <em>Acacia</em> is a large group of woody legumes containing an enormous amount of morphological diversity in leaf shape. This diversity is at least in part the result of an innovation in leaf development where many <em>Acacia</em> species are capable of developing leaves of both bifacial and unifacial morphology. While not unique in the plant kingdom, unifaciality is most commonly associated with monocots, and its developmental genetic mechanisms have yet to be explored beyond this group. Here we identify an accession of <em>Acacia crassicarpa</em> with high regeneration rates and isolate a clone for genome sequencing. We generate a chromosome-level assembly of this readily transformable clone and using comparative analyses confirm a whole genome duplication unique to Caesalpinoid legumes. This resource will be important for future work examining genome evolution in legumes and the unique developmental genetic mechanisms underlying unifacial morphogenesis in <em>Acacia</em>.</p>

opencc-zeroNov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record