Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

660

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

660 results for “genome assembly”

Learn how ShareScore rates datasets ↗
zenodo40/100

Nicotiana tabacum genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana tabacum</em>.</p> <p>The following files are available:</p> <ul> <li>ntab.fa.gz: reference genome sequence in fasta format</li> <li>ntab.gff3.gz: gene annotation in GFF3 format</li> <li>ntab.gtf.gz: gene annotation in GTF format</li> <li>ntab.tx.fa.gz: transcript sequences in fasta format</li> <li>ntab.cds.fa.gz: coding sequences in fasta format</li> <li>ntab.prot.fa.gz: protein sequences in fasta format</li> <li>ntab.tsv.gz: gene functional annotation in TSV format</li> <li>ntab.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>ntab.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>ntab.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>ntab.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Syzygium syzygioides genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Syzygium syzygioides</em>.</p> <p>The following files are available:</p> <ul> <li>ssyz.fa.gz: reference genome sequence in fasta format</li> <li>ssyz.gff3.gz: gene annotation in GFF3 format</li> <li>ssyz.gtf.gz: gene annotation in GTF format</li> <li>ssyz.tx.fa.gz: transcript sequences in fasta format</li> <li>ssyz.cds.fa.gz: coding sequences in fasta format</li> <li>ssyz.prot.fa.gz: protein sequences in fasta format</li> <li>ssyz.tsv.gz: gene functional annotation in TSV format</li> <li>ssyz.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Nicotiana sylvestris genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of <em>Nicotiana sylvestris</em>.</p> <p>The following files are available:</p> <ul> <li>nsyl.fa.gz: reference genome sequence in fasta format</li> <li>nsyl.gff3.gz: gene annotation in GFF3 format</li> <li>nsyl.gtf.gz: gene annotation in GTF format</li> <li>nsyl.tx.fa.gz: transcript sequences in fasta format</li> <li>nsyl.cds.fa.gz: coding sequences in fasta format</li> <li>nsyl.prot.fa.gz: protein sequences in fasta format</li> <li>nsyl.tsv.gz: gene functional annotation in TSV format</li> <li>nsyl.rt.fa.gz: retrotransposon sequences in fasta format</li> <li>nsyl.rt.gff3.gz: retrotransposon annotation on GFF3 format</li> <li>nsyl.rt.tsv.gz: retrotransposon annotation in TSV format</li> <li>nsyl.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Clove (Syzygium aromaticum) genome assembly and annotation

<p><em>De novo</em> genome assembly and annotation of clove (<em>Syzygium aromaticum</em>).</p> <p>The following files are available:</p> <ul> <li>saro.fa.gz: reference genome sequence in fasta format</li> <li>saro.gff3.gz: gene annotation in GFF3 format</li> <li>saro.gtf.gz: gene annotation in GTF format</li> <li>saro.tx.fa.gz: transcript sequences in fasta format</li> <li>saro.cds.fa.gz: coding sequences in fasta format</li> <li>saro.prot.fa.gz: protein sequences in fasta format</li> <li>saro.tsv.gz: gene functional annotation in TSV format</li> <li>saro.chr_to_id.tsv.gz: mapping of sequence names to ids in TSV format</li> </ul>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Supplementary data for: Chromosome-scale genome assemblies of aphids reveal extensively rearranged autosomes and long-term conservation of the X chromosome

<p><strong><em>Myzus persicae&nbsp;</em>clone O v2 frozen release</strong></p> <p>Genome assembly: Myzus_persicae_O_v2.0.scaffolds.fa.gz</p> <p>BRAKER2 gene models:&nbsp;Myzus_persicae_O_v2.0.scaffolds.braker2.gff3</p> <p>List of gene models containing internal stop codons (removed from the protein and cds fasta files):&nbsp;Myzus_persicae_O_v2.0.scaffolds.braker2.bad_genes.lst</p> <p>BRAKER2 protein&nbsp;sequences:&nbsp;Myzus_persicae_O_v2.0.scaffolds.braker2.gff3.filtered.aa.fa</p> <p>BRAKER2 protein sequences (longest transcript per gene only):&nbsp;Myzus_persicae_O_v2.0.scaffolds.braker2.gff3.filtered.aa.LTPG.fa</p> <p>BRAKER2 coding&nbsp;sequences:&nbsp;Myzus_persicae_O_v2.0.scaffolds.braker2.gff3.filtered.cds.fa</p> <p>BRAKER2 coding sequences (longest transcript per gene only):&nbsp;Myzus_persicae_O_v2.0.scaffolds.braker2.gff3.filtered.cds.LTPG.fa</p> <p><em>De novo </em>repeat library (ReapeatModeler merged with repbase insecta):&nbsp;Myzus_persicae_O_v2.0_repeat_lib.repeatmodeler_merged_repbase_insecta.fa</p> <p>RepeatMasker transposable element annotation using the <em>M. persicae de novo</em> repeat library: Myzus_persicae_O_v2.0.scaffolds.repeatmodeler_merged_repbase_insecta.repeatmasker.gff.out</p> <p>RepeatMasker transposable element annotation using the <em>M. persicae</em> <em>de novo r</em>epeat library (gff format): Myzus_persicae_O_v2.0.scaffolds.repeatmodeler_merged_repbase_insecta.repeatmasker.gff</p> <p><strong><em>Acyrthosiphon pisum</em> clone JIC1 v1&nbsp;frozen release</strong></p> <p>Genome assembly: Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.fa.gz</p> <p>BRAKER2 gene models:&nbsp;Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.braker2.gff</p> <p>List of gene models containing internal stop codons (removed from the protein and cds fasta files):&nbsp;Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.braker2.bad_genes.lst</p> <p>BRAKER2 protein&nbsp;sequences:&nbsp;Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.braker2.gff.filtered.aa.fa</p> <p>BRAKER2 protein sequences (longest transcript per gene only):&nbsp;Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.braker2.gff.filtered.aa.LTPG.fa</p> <p>BRAKER2 coding&nbsp;sequences:&nbsp;Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.braker2.gff.filtered.cds.fa</p> <p>BRAKER2 coding sequences (longest transcript per gene only):&nbsp;Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.braker2.gff.filtered.cds.LTPG.fa</p> <p><em>De novo </em>repeat library (ReapeatModeler merged with repbase insecta):&nbsp;Acyrthosiphon_pisum_JIC1_repeat_lib.repeatmodeler_merged_repbase_insecta.fa</p> <p>RepeatMasker transposable element annotation using the <em>A. pisum</em> <em>de novo</em> repeat library: Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.repeatmodeler_merged_repbase_insecta.repeatmasker.out</p> <p>RepeatMasker transposable element annotation using the <em>A. pisum&nbsp;de novo</em> repeat library (gff format): Acyrthosiphon_pisum_JIC1_v1.0.scaffolds.repeatmodeler_merged_repbase_insecta.repeatmasker.gff</p> <p><strong><em>Rhodnius prolixus</em> DNA zoo chromosome-scale genome assembly annotation</strong></p> <p><em>R. prolixus </em>chromosome-scale genome assembly was obtained here:&nbsp;<a href="https://www.dnazoo.org/assemblies/Rhodnius_prolixus">https://www.dnazoo.org/assemblies/Rhodnius_prolixus</a>.</p> <p>Genome assembly:&nbsp;Rhodnius_prolixus-3.0.3_HiC.fasta</p> <p>BRAKER2 gene models:&nbsp;Rhodnius_prolixus-3.0.3_HiC.braker2.gff</p> <p>BRAKER2 protein&nbsp;sequences:&nbsp;Rhodnius_prolixus-3.0.3_HiC.braker2.gff.aa.fa</p> <p>BRAKER2 protein sequences (longest transcript per gene only):&nbsp;Rhodnius_prolixus-3.0.3_HiC.braker2.gff.aa.LTPG.fa</p> <p>BRAKER2 coding&nbsp;sequences:&nbsp;Rhodnius_prolixus-3.0.3_HiC.braker2.gff.cds.fa</p> <p><strong><em>Triatoma rubrofasciata</em>&nbsp;chromosome-scale genome assembly annotation</strong></p> <p><em>T.&nbsp;rubrofasciata&nbsp;</em>chromosome-scale genome assembly was obtained here:&nbsp;<a href="http://dx.doi.org/10.5524/100614">http://dx.doi.org/10.5524/100614</a></p> <p>Genome assembly:&nbsp;zhuichun_assembly.fasta</p> <p>BRAKER2 gene models:&nbsp;zhuichun_assembly.braker2.gff</p> <p>BRAKER2 protein&nbsp;sequences:&nbsp;zhuichun_assembly.braker2.gff.aa.fa</p> <p>BRAKER2 protein sequences (longest transcript per gene only):&nbsp;zhuichun_assembly.braker2.gff.aa.LTPG.fa</p> <p>BRAKER2 coding&nbsp;sequences:&nbsp;zhuichun_assembly.braker2.gff.cds.fa</p> <p><strong>Hemiptera orthogroups and species tree</strong></p> <p>OrthoFinder was used to cluster proteomes of 14 Hemiptera into orthogroups for phylogenomic analysis. All proteomes were reduced to the longest transcript per gene. See here for full details:</p> <p>Species included, taxon IDs and data source:</p> <p>Mcer = Myzus cerasi v1.1 (<a href="https://bipaa.genouest.org/sp/myzus_cerasi/">https://bipaa.genouest.org/sp/myzus_cerasi/</a>)</p> <p>MperO = Myzus persicae clone O v2 (This study)</p> <p>Dnox = Diuraphis noxia Thorpe et. al. gene predictions (<a href="https://bipaa.genouest.org/sp/diuraphis_noxia/">https://bipaa.genouest.org/sp/diuraphis_noxia/</a>)</p> <p>Apis = Acyrthosiphon pisum JIC1 v1 (This study)</p> <p>Pnig = Pentalonia nigronervosa (This study)</p> <p>Rmai = Rhopalosiphum maidis v0.1 (<a href="http://gigadb.org/dataset/100572">http://gigadb.org/dataset/100572</a>)</p> <p>Rpad = Rhopalosiphum padi v1.0 (<a href="https://bipaa.genouest.org/sp/rhopalosiphum_padi/">https://bipaa.genouest.org/sp/rhopalosiphum_padi/</a>)</p> <p>Agly = Aphis glycines biotype 4 v2.1 (<a href="https://zenodo.org/record/3453468#.XnpL5JOgLRY">https://zenodo.org/record/3453468#.XnpL5JOgLRY</a>)</p> <p>BtabMEAM1 = Bemissia tabacci MEAM1 v1.2 (<a href="http://www.whiteflygenomics.org/cgi-bin/bta/index.cgi">http://www.whiteflygenomics.org/cgi-bin/bta/index.cgi</a>)</p> <p>Trub = Triatoma rubrofasciata (This study)</p> <p>Rpro = Rhodnius prolixus&nbsp;(This study)</p> <p>Ofas =&nbsp;Oncopeltus fasciatus OGS v1.0 (<a href="https://i5k.nal.usda.gov/Oncopeltus_fasciatus">https://i5k.nal.usda.gov/Oncopeltus_fasciatus</a>)</p> <p>Sfuc =&nbsp;Sogatella furcifera v1 (<a href="http://dx.doi.org/10.5524/100255">http://dx.doi.org/10.5524/100255</a>)</p> <p>Nlug =&nbsp;Nilaparvata lugens (<a href="https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0521-0#Sec42">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-014-0521-0#Sec42</a>)</p> <p>Files:</p> <p>Proteomes included in the analysis:&nbsp;proteomes.tar.gz</p> <p>Orthogroups:&nbsp;Orthogroups.txt</p> <p>Gene counts per orthogroup, per species:&nbsp;Orthogroups.GeneCount.csv</p> <p>Single copy conserved orthogroups used for species tree: SingleCopyOrthogroups.txt</p> <p>Species tree alignment:&nbsp;SpeciesTreeAlignment.fa</p> <p>r8s configuration file (includes time calibrations and OrthoFinder ML species tree with branch lengths):&nbsp;species_tree_rooted.r8s.nex</p> <p>r8s time calibrated species tree:&nbsp;r8s_tree.nwk</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Training material for flye genome assembly (Galaxy Training Network tutorial)

<p>Datasets are subsets of 3 public datasets (Mucor mucedo Fresen. NRRL 3635 Standard Draft genome sequencing with PacBio technology)</p> <p>https://www.ncbi.nlm.nih.gov/sra/SRX5336965[accn]</p> <p>https://www.ncbi.nlm.nih.gov/sra/SRX5336964[accn]</p> <p>https://www.ncbi.nlm.nih.gov/sra/SRX5336963[accn]</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

apomixis_parallel_evolution, and Fortunella hindsii (Citrus hindsii )Genome sequencing and assembly

<p>##The assemble(genome) files</p> <p>Citrus hindsii (Mini citrus)Genome assembly</p> <p>Citrus hindsii (Mini citrus)Genome assembly gene model gff3 file</p> <p>Citrus hindsii (Mini citrus)Genome assembly function annotation</p> <p>Citrus hindsii (Mini citrus)Genome assembly TE gff3 file</p> <p>##The population dataset</p> <p>LD_pur.vcf.gz&nbsp; //The LD purning SNP vcfs (1.4 M sites) used in analysis</p> <p><br> log10_auxin.txt //The expression (log10) related to auxin pathway</p> <p><br> sjg.temergedref.fasta.gz //The TE insertion modify genome in popTE2 analysis</p> <p><br> SVs.vcf.gz //The SV vcfs&nbsp; used in the paper</p> <p><br> te-hierarchy.txt //The TE classfication in popTE2 analysis</p> <p><br> TPM_count.txt //All samples expression in TMP count</p> <p><br> unfiltered.vcf.gz //The unfiltered vcfs file (7.3 M sites, within 0.4 M indels)&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo40/100

Chromosome-scale, haplotype-resolved genome assembly of Suaeda glauca

<p><em>Suaeda glauca</em>is an annual herb of Suaeda and an important saline-alkali plant resource, which is widespread on beaches and saline lands around the world. It is also a good candidate for food, feed, and drug development. There has been no publication of the&nbsp;<em>Suaeda glauca</em>genome assembly, limiting the evolutionary study of Amaranthaceae and the bioavailability of&nbsp;<em>Suaeda glauca</em>.</p> <p>Using PacBio HiFi and Hi-C sequencing data, we successfully generated chromosome-scale, haplotype-resolved assemblies of the&nbsp;<em>Suaeda glauca</em>genome. The size of the final primary assembly was 622.95 Mb, and the contig N50 was 19.42 Mb, which was successfully anchored to 9 chromosomes, accounting for 96.79% of the total assembly size. The repeat content and genome size of&nbsp;<em>Suaeda glauca</em>are much higher than those of the same genus&nbsp;<em>Suaeda aralocaspica</em>, presumably due to a recent burst of LTR insertions. Using HiFi reads, we assembled the complete circular chloroplast genome of&nbsp;<em>Suaeda glauca</em>. Through gene family and phylogenetic tree analysis, it was shown that&nbsp;<em>Suaeda glauca</em>and&nbsp;<em>Suaeda aralocaspica</em>differentiated at ~26.36 million years ago (MYA), and Amaranthaceae species began to differentiate at ~52.00 MYA.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

ATAC-seq processing resources for the GRCh38 (hg38) assembly of the human genome

<p>A collection of publicly available, but preprocessed, reference data for the analysis of ATAC-seq samples using the&nbsp;GRCh38 (hg38) assembly of the human genome&nbsp;using&nbsp;the&nbsp;<a href="https://doi.org/10.5281/zenodo.6323634">Ultimate ATAC-seq Data Processing &amp; Analysis Pipeline</a>&nbsp;(details in the documentation on GitHub).</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Metagenome-assembled Genomes of Scandinavium goeteborgense MCPNR19-05 and Erwinia aphidicola MCPNR19-06

<p>Here we provide two fasta files (MCPNR19-05.fa and MCPNR19-06.fa) which represent low-quality metagenome assembled genomes (MAGs) obtained from genomic DNA from Massospora cicadina isolate MCPNR19 azygospores collected from multiple seventeen-year cicada (Magicicada septendecim) June 2019 at Powdermill Nature Reserve, Rector, Pennsylvania.</p> <p>MCPNR19-05.fa = Scandinavium goeteborgense MCPNR19-05, a 1.82 Mb 45.61% complete MAG<br> MCPNR19-06.fa = Erwinia aphidicola MCPNR19-06, a 1.79 Mb 29.31% complete MAG<br> <br> <strong>Raw data availability</strong><br> Sequence reads are deposited under SRA project accessions <a href="https://ncbi.nlm.nih.gov/sra/SRR17553520">SRR17553520</a>-<a href="https://ncbi.nlm.nih.gov/sra/SRR17553526">SRR17553526</a> and BioProject <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA795459">PRJNA795459</a>. These MAGs are metagenomic assemblies obtained from the host Massospora cicadina (BioSample: SAMN24722893). &nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo40/100

Genome assembly and annotation of Pisum sativum cultivar ZW6 (PeaZW6)

<p>This reposity stores the genome assembly and gene annotation of Pisum sativm cultivar ZW6 (PeaZW6)</p> <p>Current Version : Release Candidate Version 2 (RC2)</p> <p>Associated NCBI BioProject :&nbsp;<strong>PRJNA730094</strong></p> <p>Correspondance&nbsp;: gaoshh@im.ac.cn</p> <p>&nbsp;</p> <p>pea.assembly.ZW6.RC2.fasta.gz&nbsp;- Full&nbsp;genome sequences</p> <p>pea.assembly.ZW6.RC2.chr.fasta.gz - Genome sequences with only chromosome molecules&nbsp;</p> <p>pea.assembly.ZW6.RC2.annotated.gff3 / gtf / bed - Gene annotation in GFF3 / GTF / BED formats</p> <p>pea.assembly.ZW6.RC2.annotated.cds.fasta - Gene coding sequences</p> <p>pea.assembly.ZW6.RC2.annotated.proteins.fasta - Gene protein sequences</p> <p>pea.assembly.ZW6.RC2.annotated.annotations.txt - Additional annotation information in tabular text format (TSV)</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
dryad40/100

Transcript- and annotation-guided genome assembly of the European starling

<p>The European starling, <em>Sturnus vulgaris</em>, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (<em>S. vulgaris</em> vAU), and a second North American genome (<em>S. vulgaris</em> vNA), as complementary reference genomes for population genetic and evolutionary characterisation. <em>S. vulgaris</em> vAU combined 10x Genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 1,628 scaffolds (72.5 Mb scaffold N50). Species-specific transcript mapping and gene annotation revealed high structural and functional completeness (94.6% BUSCO completeness). Further scaffolding against the high-quality zebra finch (<em>Taeniopygia guttata</em>) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Rapid, recent advances in sequencing technologies and bioinformatics software have highlighted the need for evidence-based assessment of assembly decisions on a case-by-case basis. Using <em>S. vulgaris</em> vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, SAAGA) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counter-intuitive behaviour in traditional BUSCO metrics, and present BUSCOMP, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. Finally, we present a second starling assembly, <em>S. vulgaris</em> vNA, to facilitate comparative analysis and global genomic research on this ecologically important species.</p>

opencc-zeroJul 2022View details →
zenodo40/100

Training material for Genome assembly quality control (Galaxy Training Network tutorial)

<p>This Zenodo repository includes the required datasets for following the GTN: Genome assembly quality control.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

SPAAM Summer School 2022: Introduction to Ancient Metagenomics - 4c Introduction to Genome Assembly

<p>Teaching data for&nbsp;practical session: &quot;4c&nbsp;Introduction to Genome Assembly&quot;&nbsp;of the 2022 SPAAM Summer School: Introduction to Ancient Metagenomics (Aug. 1-5 2022).</p> <p>See:&nbsp;<a href="https://spaam-community.github.io/wss-summer-school/#/2022/">https://spaam-community.github.io/wss-summer-school/#/2022/</a>&nbsp;or&nbsp;<a href="https://doi.org/10.5281/zenodo.6976711">https://doi.org/10.5281/zenodo.6976711</a>&nbsp;for slides.</p> <p>Once downloaded, run:</p> <pre><code>tar xvfz &lt;session&gt;.tar.gz</code></pre> <p>&nbsp;to decompress the data directory for&nbsp;the session.</p>

opencc-by-4.0Aug 2022View details →
zenodo40/100

Genome and annotation files for Blumeria graminis f. sp. tritici isolate ISR_7 (genome assembly: Bgt_ISR7_genome_v1_4)

<p>Genome and annotation files for Blumeria graminis f. sp. tritici isolate ISR_7 (genome assembly: Bgt_ISR7_genome_v1_4)</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Octopus bimaculoides genome assembly

<p>To facilitate identification of molecular cell types in the<em>&nbsp;Octopus bimaculoides</em>&nbsp;optic lobe, we conducted high fidelity long-read genomic and transcriptomic sequencing. Here, we provide open access to these datafiles, which resulted in a single-cell atlas of the&nbsp;<em>O. bimaculoides</em>&nbsp;visual system (<a href="https://doi.org/10.1016/j.cub.2022.10.015" target="_blank" rel="noopener noreferrer">https://doi.org/10.1016/j.cub.2022.10.015</a>). Below, we include a new genome assembly and annotation, phylogenetic trees for genes referenced in the manuscript, Cell Ranger outputs, and R scripts and files for Seurat cluster analysis. Raw single-cell sequence files are deposited to NCBI SRA at BioProject ID PRJNA854179. If any of these files are used, we ask that the paper is cited.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Genome annotations of Drosophila melanogaster and Drosophila simulans wild-type strains from long read sequencing assemblies

<p>Genome assemblies were performed for eight wild-type strains of Drosophila melanogaster and Drosophila simulans from Oxford Nanopore long read sequencing (please refer to Mohamed et al. Cells 2020 (doi:10.3390/cells9081776)). Assemblies were deposited in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB50024 (<a href="https://www.ebi.ac.uk/ena/browser/view/PRJEBxxxx">https://www.ebi.ac.uk/ena/browser/view/</a>PRJEB50024).</p> <p>Transposable Element annotations: we used RepeatMasker 4.1.0 (<a href="http://repeatmasker.org/">http://repeatmasker.org/</a>) -species Drosophila, followed by OneCodeToFindThemAll (Bailly-Bechet et al. 2014) with default parameters.</p> <p>Gene annotations: We retrieved gtf files from FlyBase : <a>ftp.flybase.net/genomes/D</a><a>rosophila_melanogaster/dmel_r6,46_FB2022_03/gft/dmel-all-r6.46.gtf.gz</a> and <a>ftp.flybase.net/genomes/Drosophila_simulans/dsim_r2,02_FB2017_04/gtf/dsim-all-</a><a>r2,02.gtf.gz</a>. The corresponding fasta files were also downloaded from FlyBase: <a>ftp.flybase.net/genomes/Drosophila_melanogaster/dmel_r6,46_FB2022_03/</a><a>fasta</a><a>/dmel-all-</a><a>chromosome-</a><a>r6.46.</a><a>fasta</a><a>.gz</a> and <a>ftp.flybase.net/genomes/Drosophila_simulans/dsim_r2,02_FB2017_04/</a><a>fasta</a><a>/dsim-all-</a><a>chromosome-</a><a>r2,02.</a><a>fasta</a><a>.gz</a>. We used Liftoff (Shumate and Salzberg, 2020) to lift over gene annotations from the references to our genome assemblies. We used -flank 0.2 and only kept the &ldquo;gene&rdquo; and &ldquo;exon&rdquo; terms.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Genomic insights into local adaptation and future climate-induced vulnerability of a keystone forest tree in East Asia (The genome assembly and annotion fiile)

<p>The&nbsp;genome assembly and annotion fiile used in the manuscript:&nbsp;<strong>Genomic insights into local adaptation and future climate-induced vulnerability of a keystone forest tree in East Asia</strong></p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Supporting Information for: Single-fly genome assemblies fill major phylogenomic gaps across the Drosophilidae Tree of Life

<p>This data repository contains supporting information, data, and code for figures and analysis pipelines the PLOS Biology article: "Single-fly genome assemblies fill major phylogenomic gaps across the Drosophilidae Tree of Life."</p> <ul> <li><strong>4d_full.treefile</strong>: Data underlying Figure 1 (note: tree was plotted as a cladogram and key groups collapsed on iToL; the treefile was not modified) and Figure S1.</li> <li><strong>S2_data.csv</strong>: Data underlying Figure 2.</li> <li><strong>S3_data.csv</strong>: Data underlying Figure 3.</li> <li>Data underlying Figure 4 is found in Table S4 of supplementary_tables.xlsx in the main manuscript</li> <li><strong>S5_data.csv</strong> Data underlying Figure 5</li> <li><strong>S6_data.csv</strong> Data underlying Figure S2&nbsp;</li> <li><strong>illumina_only_assms.tar.gz</strong>: Archive of Illumina-only assemblies (FASTA) based on publicy available data that we did not generate. Assemblies generated from our own short-read data have been submitted to NCBI GenBank.</li> <li><strong>illumina_vcfs.tar.gz</strong>: Illumina-based variant calls and BED tracks of masked bases.</li> <li><strong>genomes.tar.gz</strong>: Genome files, for archival purposes.</li> <li><strong>repeatModeler-lib.tar.gz</strong>: RepeatModeler2 libraries.</li> <li><strong>diploid_genomes.tar.gz</strong>: diploid genomes and BED tracks of phased regions.</li> <li><strong>trees.tar.gz</strong>: phylogenies</li> </ul>

opencc-by-4.0May 2024View details →
zenodo40/100

Supplemental Results for Assembly, Annotation, and Analysis from HiFi reads of Gulf Toadfish Genome and Transcriptome fOpsBet2.1

<p>This repository contains gzipped tarballs of the results of the various assembly, annotation, and analysis steps performed during the assembly of the fOpsBet2.1 genome assembly for Opsanus beta at the University of Miami Rosenstiel School of Marine, Atmospheric, and Earth Science for the McDonald Toadfish Lab. These results are too numerous to include as supplemental data for a journal publication and so are available here for review. In this repository you will find results for:</p><p>Scripts:</p><p>-all bash and LSF scheduler job scripts used as part of the analysis, both exploratory and final.&nbsp;</p><p>QC:</p><p>-GenomeScope2 estimation of genome metrics from HiFi Reads</p><p>-QUAST genome statistics for each assembly step</p><p>-BUSCO completeness assessments for each assembly step&nbsp;</p><p>-inspector logs for polishing of initial assembly</p><p>-logs from Kraken2 contaminant screen</p><p>Assembly and Scaffolding:</p><p>-ntLINKS logs and intermediates for initial scaffolding</p><p>-ragtag logs and metrics for super-scaffolding to the ThaAma1.1 T. amazonica reference assembly</p><p>-mitoHIFI results for mitogenome assembly from HiFi reads, primary assembly, and purged alternate assembly</p><p>Annotation:</p><p>-PASA directory with full input and output for SQLite PASA assembly of transcriptome for gene predictors</p><p>-Results folder for Funannotate::annotate for gene models, annotations, and CDS/mRNA/protein fastas</p><p>-InterProSCan5 results for protein annotation used as input into Funannotate</p><p>-ghostKOALA KEGG assignment results for predicted proteins from funannotate results</p><p>Repetitive Elements:</p><p>-tidk telomere repeat analysis results</p><p>-TRAH satellite DNA analysis with subsequent analysis with HiCAT and StainedGlass</p><p>-RepeatModeler results for de novo TE prediction</p><p>-repclassifier results for TE curation</p><p>Comparative Analysis:</p><p>-OrthoFinder ortholog search for O. beta to several other vertebrates</p><p>-CAFE5 gene family expansion and contraction of Orthogroups from OrthoFinder results</p><p>&nbsp;</p>

opencc-by-4.0Jul 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record