Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
199
datasets available to search
ShareScore release 0.7.1
Dataset results
199 results for “reference genome”
Construction of a chromosome-scale long-read reference genome assembly for potato
<p><span><strong>Background:</strong> Worldwide, the cultivated potato, <i>Solanum tuberosum </i>L<i>.</i>, is the number one vegetable crop and a critical food security crop. The genome sequence of DM1-3 516 R44, a doubled monoploid clone of S. <i>tuberosum </i>Group Phureja, was published in 2011 using a whole-genome shotgun sequencing approach with short read sequence data. Current advanced sequencing technologies now permit generation of near-complete, high-quality chromosome-scale genome assemblies at a minimal cost. </span></p> <p><span><strong>Findings: </strong>Here, we present an updated version of the DM1-3 516 R44 genome sequence (v6.1) using Oxford Nanopore Technologies long reads coupled with proximity-by-ligation scaffolding (Hi-C) yielding a chromosome-scale assembly. The new (v6.1) assembly represents 741.6 Mb of sequence (87.8 %) of the estimated 844 Mb genome, of which, 741.5 Mb is non-gapped with 731.2 Mb anchored to the 12 chromosomes. Use of Oxford Nanopore Technologies full-length cDNA sequencing enabled annotation of 32,917 high-confidence protein-coding genes encoding 44,851 gene models that had a significantly improved representation of conserved orthologs compared to the previous annotation. The new assembly has improved contiguity with a 595-fold increase in N50 contig size, 99% reduction in the numbersof contigs, a 44-fold increase in N50 scaffold size, and an LTR Assembly Index score of 13.56, placing it in the category of reference genome quality. The improved assembly also permitted annotation of the centromeres via alignment to sequencing reads derived from CENH3 nucleosomes. </span></p> <p><span><strong>Conclusions: </strong>Access to advanced sequencing technologies and improved software permitted generation of a high-quality, long-read, chromosome-scale assembly and improved annotation dataset for the reference genotype of potato that will facilitate research aimed at improving agronomic traits and understanding genome evolution.</span></p>
Reference flow VCF for pre-built genomes
<p>Pre-built genomes (in VCF format) for the RandFlow-LD and RandFlow-LD-26 methods in reference flow. The references can be built using the reference flow software (https://github.com/langmead-lab/reference_flow). An archival version of the software is available at http://doi.org/10.5281/zenodo.4287778</p> <p> </p> <p>The reference flow method is described at https://www.biorxiv.org/content/10.1101/2020.03.03.975219v4</p>
Data from: The plover neurotranscriptome assembly: transcriptomic analysis in an ecological model species without a reference genome
We assembled a de novo transcriptome of short-read Illumina RNA-Seq data generated from telencephalon and diencephalon tissue samples from the Kentish plover, Charadrius alexandrinus. This is a species of considerable interest in behavioural ecology for its highly variable mating system and parental behaviour, but it lacks genomic resources and is evolutionarily distant from the few available avian draft genome sequences. We assembled and identified over 21 000 transcript contigs with significant expression in our samples, showing high homology to exonic sequences in avian draft genomes. From these, we identified >31 000 high-quality SNPs and > 2500 simple sequence repeats (SSRs). We also analysed expression patterns in our data to identify potential candidate genes related to differences in male and female behaviour, identifying over 200 nonoverlapping putative autosomal transcripts that show significant expression differences between males and females. Gene ontology analysis revealed that female-biased transcripts were significantly enriched for cerebral functions related to learning, cognition and memory, and male-biased transcripts were mostly enriched for terms related to neural function such as neuron projection and synapses. This data set provides one of the first de novo transcriptome assemblies from non-normalized short-read next-generation data and outlines an effective strategy for measuring sequence and expression variability simultaneously without the aid of a reference genome.
Data from: A genomic reference panel for Drosophila serrata
Here we describe a collection of re-sequenced inbred lines of Drosophila serrata, sampled from a natural population situated deep within the species endemic distribution in Brisbane, Australia. D. serrata is a member of the speciose montium group whose members inhabit much of south east Asia and has been well studied for aspects of climatic adaptation, sexual selection, sexual dimorphism, and mate recognition. We sequenced 110 lines that were inbred via 17-20 generations of full-sib mating at an average coverage of 23.5x with paired-end Illumina reads. 15,228,692 biallelic SNPs passed quality control after being called using the Joint Genotyper for Inbred Lines (JGIL). Inbreeding was highly effective and the average levels of residual heterozygosity (0.86%) were well below theoretical expectations. As expected, linkage disequilibrium decayed rapidly, with r2 dropping below 0.1 within 100 base pairs. With the exception of four closely related pairs of lines which may have been due to technical errors, there was no statistical support for population substructure. Consistent with other endemic populations of other Drosophila species, preliminary population genetic analyses revealed high nucleotide diversity and, on average, negative Tajima's D values. A preliminary GWAS was performed on a cuticular hydrocarbon trait, 2-MeC28 revealing 4 SNPs passing Bonferroni significance residing in or near genes. One gene Cht9 may be involved in the transport of CHCs from the site of production (oenocytes) to the cuticle. Our panel will facilitate broader population genomic and quantitative genetic studies of this species and serve as an important complement to existing D. melanogaster panels that can be used to test for the conservation of genetic architectures across the Drosophila genus.
Data from: Genomics of Compositae crops: reference transcriptome assemblies, and evidence of hybridization with wild relatives
Although the Compositae harbours only two major food crops, sunflower and lettuce, many other species in this family are utilized by humans and have experienced various levels of domestication. Here we have used next generation sequencing technology to develop 15 reference transcriptome assemblies for Compositae crops or their wild relatives. These data allow us to gain insight into the evolutionary and genomic consequences of plant domestication. Specifically, we performed Illumina sequencing of Cichorium endivia, Cichorium intybus, Echinacea angustifolia, Iva annua, Helianthus tuberosus, Dahlia hybrida, Leontodon taraxacoides and Glebionis segetum, as well 454 sequencing of Guizotia scabra, Stevia rebaudiana, Parthenium argentatum and Smallanthus sonchifolius. Illumina reads were assembled using Trinity, and 454 reads were assembled using MIRA and CAP3. We evaluated the coverage of the transcriptomes using BLASTX analysis of a set of ultra-conserved orthologs (UCOs) and recovered most of these genes (88-98%). We found a correlation between contig length and read length for the 454 assemblies, and greater contig lengths for the 454 compared to the Illumina assemblies. This suggests that longer reads can aid in the assembly of more complete transcripts. Finally, we compared the divergence of orthologs at synonymous sites (Ks) between Compositae crops and their wild relatives and found greater divergence when the progenitors were self-incompatible. We also found greater divergence between pairs of taxa that had some evidence of post-zygotic isolation. For several more distantly related congeners, such as chicory and endive, we identified a signature of introgression in the distribution of Ks values.
Data from: Long read reference genome-free reconstruction of a full-length transcriptome from Astragalus membranaceus reveals transcript variants involved in bioactive compound biosynthesis
Astragalus membranaceus, also known as Huangqi in China, is one of the most widely used medicinal herbs in Traditional Chinese Medicine. Traditional Chinese Medicine formulations from Astragalus membranaceus have been used to treat a wide range of illnesses, such as cardiovascular disease, type 2 diabetes, nephritis and cancers. Pharmacological studies have shown that immunomodulating, anti-hyperglycemic, anti-inflammatory, antioxidant and antiviral activities exist in the extract of Astragalus membranaceus. Therefore, characterising the biosynthesis of bioactive compounds in Astragalus membranaceus, such as Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside, is of particular importance for further genetic studies of Astragalus membranaceus. In this study, we reconstructed the Astragalus membranaceus full-length transcriptomes from leaf and root tissues using PacBio Iso-Seq long reads. We identified 27 975 and 22 343 full-length unique transcript models in each tissue respectively. Compared with previous studies that used short read sequencing, our reconstructed transcripts are longer, and are more likely to be full-length and include numerous transcript variants. Moreover, we also re-characterised and identified potential transcript variants of genes involved in Astragalosides, Calycosin and Calycosin-7-O-β-D-glucoside biosynthesis. In conclusion, our study provides a practical pipeline to characterise the full-length transcriptome for species without a reference genome and a useful genomic resource for exploring the biosynthesis of active compounds in Astragalus membranaceus.
Data from: Chromosome-level reference genome assembly and gene editing of the dead-leaf butterfly Kallima inachus
<p class="p">The leaf resemblance of <i><span>Kallima</span></i> (Nymphalidae) butterflies<i> </i>is an important ecological adaptive mechanism that increases survival. However, the genetic mechanism underlying ecological adaptation remains unclear owing to a dearth of genomic information. Herein, we revealed the karyotype (n = 31) of the dead-leaf butterfly <i><span>Kallima inach</span></i><i><span>us</span></i>, assembled its high-quality chromosome-level reference genome (568.92 Mb; contig N50: 19.20 Mb), and identified its Z and candidate W chromosomes. To our knowledge, this is the first study to report on these aspects of this species. In the assembled genome, 15,309 protein-coding genes and 49.86% repeat elements were annotated. Phylogenetic analysis showed that <i><span>K. inachus</span></i> diverged from <i><span>Melitaea cinxia </span></i>(no leaf resemblance), both of which are in Nymphalinae, around 40 million years ago. Demographic analysis indicated that the effective population size of <i><span>K. inach</span></i><i><span>us</span></i> decreased during the last interglacial period in the Pleistocene. The wings of adults with the pigmentary gene <i><span>ebony</span></i> knocked out using CRISPR/Cas9 showed phenotypes in which the orange dorsal region and entire ventral surface darkened, suggesting its vital role in the ecological adaption of dead-leaf butterflies. Our results provide important genome resources for investigating the genetic mechanism underlying protective resemblance in dead-leaf butterflies and insights into the molecular basis of protective coloration.</p>
DeepCopy's reference genomes
<p>DeepCopy is an algorithm for determining copy number aberrations on single cell DNA sequencing data of cancer cells. Reference genomes data is one of the required inputs of DeepCopy. </p>
Chromosome-level Reference Genome of the Critically Endangered Tree Kmeria septentrionalis
<p><i>Kmeria septentrionalis, </i>a critically endangered tree endemic to Guangxi in China and listed on the International Union for Conservation of Nature's Red List, suffers from a lack of genetic information and a paucity of high-quality genome data. In our study, we construct and annotate a complete <i>K. septentrionalis</i> genome at the chromosome level, assess its quality, and contextualize it with the genomic data of other relative plants. The genome is measured at 2.57 Gb with a contig N50 of 11.93 Mb. Using Hi-C guided genome assembly, we assembled 496 out of the initial 705 raw contigs into 19 pseudochromosomes that have a scaffold N50 of 135.08 Mb. The assembled 2.54 Gb anchored genome has achieved 98.9% completeness, and contains 35,927 genes, of which 94.15% could be functional annotated. </p>
Phased T2T reference genome and pangenome reveal expanded resistance gene analogs in apple domestication
<p>Phased T2T reference genome and pangenome reveal expanded resistance gene analogs in apple domestication<br>The data contains two files, namely genome file and genome annotation file:<br>1. Genome:<br>This file contains 12 genomes for 5 species (cultivar), namely: <em>Malus domestica</em> cv. ‘Golden Delicious’ (GD), <em>Malus domestica</em> cv. ‘Gala’ (Gala), <em>Malus domestica</em> cv. ‘Honeycrisp’ (HC), <em>Malus domestica</em> cv. ‘HFTH1’ (HFTH1), <em>Malus baccata</em> (Mba), <em>Malus sieversii</em> (Msi), and <em>Malus sylvestris</em> (Msy)</p> <p>2. Genome annotation:<br>Corresponding to the genome file.<br>Note: Hap1 and hap2 represent two haplotypes of a species (cultivar), all data used for the construction of pan-genomes and comparative genomics analysis.</p>
The reference chromosome genome for Zhengitettix transpicula
<p>The reference chromosome genome for Zhengitettix transpicula. The HIC genome and the gff3 annotation file.</p>
Collated reference lichen genomes
<p>Reference genomes concatenated into files for each lichen family - sourced from NCBI and JGI. <br>Details on how databases were created fully available on the main <a href="https://github.com/Kamouyiaraki/DEFRALichens/tree/main/databases">project github</a></p>
Whole genome re-sequencing workshop data: fastq files and reference genomes
<p> </p> <p>Whole genome re-sequencing data analysis workshop datasets. The files are necessary inputs for the workshop in https://github.com/PoODL-CES/Genomics_learning_workshop</p> <p>Tools and scripts listed in the https://github.com/PoODL-CES/Genomics_learning_workshop repository.</p> <p> </p> <p>These are subsampled fastq files from:<br><br>Khan, A., Patel, K., Shukla, H., Viswanathan, A., van der Valk, T., Borthakur, U., Nigam, P., Zachariah, A., Jhala, Y.V., Kardos, M. and Ramakrishnan, U., 2021. Genomic evidence for inbreeding depression and purging of deleterious genetic variation in Indian tigers. <em>Proceedings of the National Academy of Sciences</em>, <em>118</em>(49), p.e2023018118.</p> <p>The reference genome is from :</p> <p>Shukla, H., Suryamohan, K., Khan, A., Mohan, K., Perumal, R.C., Mathew, O.K., Menon, R., Dixon, M.D., Muraleedharan, M., Kuriakose, B. and Michael, S., 2023. Near-chromosomal de novo assembly of Bengal tiger genome reveals genetic hallmarks of apex predation. <em>GigaScience</em>, <em>12</em>, p.giac112.</p> <p> </p> <p>The reference has been indexed using:</p> <p>bwa index <a target="_blank" rel="noopener noreferrer">GCA_021130815.1_PanTigT.MC.v3_genomic.fna</a></p> <p> </p>
Nicotiana benthamiana and virus genomic references
<p>Nicotiana benthamiana curated and modified genomic references.</p> <p>TSWV viral references.</p>
Reference genomes for WGSUniFrac experiments - simulated
<p>99_otus.fasta: reference genome for 16S data</p> <p>wgs_genome2: reference genome for wgs data</p>
Addiitional Files: The diagrams of population structure, highly divergent regions, GC content and Nanopore reads depth, SNP number and Nanopore reads depth, and analyses of co-linearity against Nipponbare reference genome in 251 accessions.
<p>Additional Files for " <strong>A Super Pan-Genomic Landscape of Rice".</strong></p> <p>Addtional File1: Supplementary File1.Population structure of 251 rice accessions inferred by ADMIXTURE from K=6 to K=15.</p> <p>Additional File2: Supplementary File2.The diagram of co-linearity for assembled genome against Nipponbare refercne genome in 251 rice accessions.</p> <p>Additional File3: Supplementary File3. Highly divergent regions based on SV.</p> <p>Additional File4: Supplementary File4. The diagram of SNP number and Nanopore reads depth per 100kb windows in 251 rice accessions.</p> <p>Additonal File5:Supplementary File5. The diagram of GC content and the Nanopore reads depth per 10kb windows in 251 rice accessions.</p> <p> </p>
Pre-computed MGCs from human microbiome reference genomes
<p>This dataset contains non-redundant metabolic gene clusters (MGCs) collected by running gutSMASH and BiG-MAP on a collection of unique high-quality reference genomes. This collection consist of MGCs predicted by gutSMASH using 1,520 genomes from the Culturable Genome Reference (CGR), 2,308 genomes from the Human Microbiome Project (HMP) and 414 Clostridia genomes as input and then filtered for redundancy using the family module of BiG-MAP. For more information: <a href="http://doi.org/10.1101/2021.02.25.432841">https://doi.org/10.1101/2021.02.25.432841</a></p> <p><strong>BiG-MAP_mg.pickle</strong> -> suitable for <strong>metagenome</strong> analyses</p> <p><strong>BiG-MAP_mt.pickle </strong>-> suitable for <strong>metatranscriptome </strong>analyses</p> <p>The files can be used as direct input in the third module of BiG-MAP (BiG-MAP.map.py: <a href="https://github.com/medema-group/BiG-MAP">https://github.com/medema-group/BiG-MAP</a>).</p>
1000 Genomes and S-LDSC reference files for CT-FM (Kim et al., 2024, in prep)
Open the record for dataset details and reuse information.
Supporting data for: "mtGrasp: Streamlined mitochondrial genome reference-grade assembly and standardization to enhance mitogenome resources and improve the development of environmental DNA assays"
<p>Here, we provide supporting data for the manuscript "mtGrasp: Streamlined mitochondrial genome reference-grade assembly and standardization to enhance mitogenome resources and improve the development of environmental DNA assays".</p> <p>Phylogenetic_analysis.tar.gz contains the script and fasta files used for the phylogenetic analysis, and Mitogenomes.tar.gz contains the mitochondrial sequences utilized that are not publicly available in GenBank.</p>
1K_genomes_reference_panel
<p>The 1K Genomes Reference Panel for Europeans was downloaded from the GitHub repository of LOGODetect (https://github.com/ghm17/LOGODetect).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.