Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,439
datasets available to search
ShareScore release 0.9.0
Dataset results
2,439 results for “Assembly”
The Metagenome-Assembled Genome Inventory for Children (MAGIC)
<div> <div>Existing microbiota databases are biased towards adult samples, hampering accurate profiling of the infant gut microbiome. Here, we generated a **M**etagenome-**A**ssembled **G**enome **I**nventory for **C**hildren (**MAGIC**) from a large collection of bulk and viral-like particle-enriched metagenomes from 0-7 years of age, encompassing `3,299` prokaryotic and `139,624` viral species-level genomes, `8.5%` and `63.9%` of which are unique to MAGIC. MAGIC improves early-life microbiome profiling, with the greatest improvement in read mapping observed in Africans. We then identified `54` candidate keystone species, including several *Bifidobacterium spp.* and four phages, forming guilds that fluctuated in abundance with time. Their abundances were reduced in preterm infants and were associated with childhood allergies. By analyzing the *B. longum* pangenome, we found evidence of phage-mediated evolution and quorum sensing-related ecological adaptation. Together, the MAGIC database recovers genomes that enable characterization of dynamics of early-life microbiomes, identification of candidate keystone species, and strain-level study of target species.</div> </div>
Gene Annotations of 49 Bacillariophyta Genome Assemblies (Individual gff3 files)
<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div> </div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a> . It is a copy of the data hostet at <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a> , but instead of storing one archive will all gff3 files included, the gff3 files are here hosted, individually. This copy was made upon request from the RDA Working Group "FAIRification of Genomic Annotations – metadata harmonisation at scale".</p> <div> <h2>Files</h2> <div>The following gzip-compressed gff3-files with structural and functional genome annotation are included:</div> <div> </div> <div>Asterionella_formosa.gff3.gz<br>Asterionellopsis_glacialis.gff3.gz<br>Bacterosira_constricta.gff3.gz<br>Chaetoceros_muellerii.gff3.gz<br>concatenated_output.gff3.gz<br>Conticribra_guillardii.gff3.gz<br>Conticribra_weissflogii.gff3.gz<br>Craspedostauros_australis.gff3.gz<br>Cyclostephanos_invisitatus.gff3.gz<br>Cyclostephanos_tholiformis.gff3.gz<br>Cyclotella_atomus.gff3.gz<br>Cyclotella_baltica.gff3.gz<br>Cyclotella_choctawhatcheeana.gff3.gz<br>Cyclotella_cryptica.gff3.gz<br>Cylindrotheca_fusiformis.gff3.gz<br>Detonula_confervacea.gff3.gz<br>Discostella_pseudostelligera.gff3.gz<br>Discostella_stelligera.gff3.gz<br>Discostella_stelligeroides.gff3.gz<br>Epithemia_pelagica.gff3.gz<br>Fistulifera_pelliculosa.gff3.gz<br>Fistulifera_solaris.gff3.gz<br>Fragilaria_radians.gff3.gz<br>Fragilariopsis_cylindrus.gff3.gz<br>Licmophora_abbreviata.gff3.gz<br>Mediolabrus_comicus.gff3.gz<br>Nitzschia_palea.gff3.gz<br>Nitzschia_putrida.gff3.gz<br>Porosira_glacialis.gff3.gz<br>Psammoneis_japonica.gff3.gz<br>Pseudo-nitzschia_multiseries.gff3.gz<br>Pseudo-nitzschia_pungens.gff3.gz<br>Skeletonema_costatum.gff3.gz<br>Skeletonema_marinoi.gff3.gz<br>Skeletonema_menzelii.gff3.gz<br>Skeletonema_potamos.gff3.gz<br>Skeletonema_tropicum.gff3.gz<br>Stephanocyclus_meneghinianus.gff3.gz<br>Stephanodiscus_minutulus.gff3.gz<br>Stephanodiscus_triporus.gff3.gz<br>Thalassiosira_allenii.gff3.gz<br>Thalassiosira_delicatula.gff3.gz<br>Thalassiosira_exigua.gff3.gz<br>Thalassiosira_gravida.gff3.gz<br>Thalassiosira_livingstoniorum.gff3.gz<br>Thalassiosira_mediterranea.gff3.gz<br>Thalassiosira_oceanica.gff3.gz<br>Thalassiosira_ordinaria.gff3.gz<br>Thalassiosira_pacifica.gff3.gz<br>Thalassiosira_profunda.gff3.gz</div> <div> </div> <div>To extract individual files after download execute the following command:</div> <div> </div> <div><code>gunzip *.gff3.gz</code></div> <h2>Genome Assemblies</h2> <p> </p> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <p> </p> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p> </p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p> </p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p> </p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^>/ s/ .*//' genome.fasta > genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <p> </p> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <p> </p> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>The submission and release was made upon request of the RDA working group "FAIRification of Genomic Annotations – metadata harmonisation at scale". The contained data is identical to <a href="https://zenodo.org/records/13933292">https://zenodo.org/records/13933292</a></p> <h2>License</h2> <p> </p> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div> <p> </p> </div> </div>
Assembly of enterohemorrhagic Escherichia coli type IV pilin PpdD and its variants
<p>Assembly of the EHEC major type IV pilin PpdD was analyzed in a reconstituted TP assembly system described in LunaRico et al. Mol Microbiol. 2019 Mar;111(3):732-749. doi: 10.1111/mmi.14188. Bacteria of strain BW25113 F'tet harboring plasmids pMS41 and pCHAP8565 (or its variants) were grown for 2 days at 30°C on M9 plates containing 0.5% glycerol, amplicillin (100 ug/ml) chloramphenicol (25 ug/ml) and 1 mM IPTG.</p> <p>Bacteria were collected and fractionated as described in Luna Rico et al Methods Mol Biol. 2018;1764:291-305. doi: 10.1007/978-1-4939-7759-8_18. Cell and sheared fractions were analysed by electrophoresis on 10 % Tris-Tricin gels, transferred on nitrocellulose and probed with anti-MalE-PpdD polyclonal antibodies. The fluorescence signal was developed with ECL2 (Thermo) and recorded with Typhoon FLA9000 imager (GE).</p> <p>The signal was quantified using ImageJ. The fractions of PpdD assembled into pili were quantified and analysed using Prism9.</p> <p>The images uploaded here are the raw data used to produce the Fig. 4B of the article Karami et al., Structure, 2021.</p>
Fraxinus pennsylvanica genome assembly and annotation
<p>We report the first chromosome-level assembly for green ash (<em>Fraxinus pennsylvanica</em>) to assist in ash breeding efforts to propagate resistance to the emerald ash borer. The final haploid assembly consists of 23 chromosomes and 87 unplaced scaffolds of 10 kb or more. Over 99% of the bases anchored to the chromosomes. The assembly spans 757 Mb and consists of 49.43% repetitive DNA. Gene annotation yielded 35,470 high-confidence gene models, all located on the chromosomes and assigned to 22,976 Asterid Orthogroups.</p>
A database solutions for the type two assembly line balancing problems
<p> Assembly Line Balancing Problems have a significant impact on performance of manufacturing systems, specially for the cases of mass production. These problems are widely cited and treated in the literature. </p> <p>One from the most important variants of those problems is the “Task Restrictions Assembly Line Balancing Problem” of type 2. For this problem, a set of tasks need to be affected to a predefined number of stations m from the way that minimises the cycle time and respects a set of constraints related to precedence and compatibility between tasks (Triki et al., 2016).</p> <p>For this variant we suggest an innovative speed and effective approach based on the hybridisation of two powerful tools: the ant colony optimisation and the genetic algorithm. The effectiveness of this approach is evaluated through a set of instances collected from the literature (Thomas, 1990; Triki et al., 2016) .</p> <p>This document presents the best generated solutions for those problems. </p> <p> </p> <p> </p>
Assembled transcriptomes of ovary, testis, and brain (male and female) of Amphibolurus muricatus (jacky dragon) generated using Trinity v2.11.0
<p><strong><em>A. muricatus</em> transcriptome assemblies generated using Trinity v2.11.0 (Haas et al. 2013; Grabherr et al. 2011; Henschel et al. 2012)</strong><br> • Amphibolurus-muricatus_brain.fa.tar.gz: Combined Trinity assembly of <em>A. muricatus</em> brain (male and female).<br> • Amphibolurus-muricatus_combined.fa.tar.gz: Combined Trinity assembly of <em>A. muricatus</em> ovary, testis, and brain (male and female).<br> • Amphibolurus-muricatus_female_brain.fa.tar.gz: Trinity assembly of female <em>A. muricatus</em> brain.<br> • Amphibolurus-muricatus_male_brain.fa.tar.gz: Trinity assembly of male <em>A. muricatus</em> brain.<br> • Amphibolurus-muricatus_ovary.fa.tar.gz: Trinity assembly of <em>A. muricatus</em> ovary.<br> • Amphibolurus-muricatus_testis.fa.tar.gz: Trinity assembly of <em>A. muricatus</em> testis.</p> <p> </p> <p><strong>References</strong></p> <ul> <li>Grabherr, M.G., B.J. Haas, M. Yassour, J.Z. Levin, D.A. Thompson et al., 2011 Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat Biotechnol 29 (7):644-652.</li> <li>Haas, B.J., A. Papanicolaou, M. Yassour, M. Grabherr, P.D. Blood et al., 2013 De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat Protoc 8 (8):1494-1512.</li> <li>Henschel, R., M. Lieber, L.-S. Wu, P.M. Nista, B.J. Haas et al., 2012 Trinity RNA-Seq assembler performance optimization, pp. 45 in Proceedings of the 1st Conference of the Extreme Science and Engineering Discovery Environment: Bridging from the eXtreme to the campus and beyond. Association for Computing Machinery, Chicag, IL, USA.</li> </ul> <p> </p>
Pep_pel_assembly_MSU_V2_2021
<p>Genome-guided transcriptome assembly for <em>Peperomia pellucida. </em>Derived from whole seedlings, young, and mature leaves.</p>
Mycobacteroides abscessus subp. bolletii strain associated with a persistent infection (genome assembly and annotation dataset)
<p>This dataset includes the assembled contigs (.fasta and .gbk files), the nucleotide sequences of the prediction transcripts (.ffn files) and the respective amino acid sequences of the translated CDS sequences (.faa files) of a <strong><em>Mycobacteroides abscessus subp. bolletti </em></strong>strain associated with a persistente infection. (genome anotation was performed using Bakta v1.2.2 https://github.com/oschwengers/bakta)</p> <p>The raw sequence reads were deposited in the European Nucleotide Archive (ENA) (BioProject PRJEB57933; Run Accession: ERR10554471).</p>
Data for "Ecosystem size filters life-history strategies to shape community assembly in lakes"
<p>Dataset 1. List of 71 fish species collected from north temperate lakes in Wisconsin USA. Data include critical life-history data used for strategy classifications according to Winemiller and Rose (1992), principal component scores, and strategy classification according to the cluster analysis.</p> <p>Dataset 2. Species occurrence data in all study lakes along with results from the 'soft classification" according to Euclidean distance.</p> <p>Dataset 3. Limnological and fish community characteristics of study lakes including species richness, lake area, estimated lake volume, and convex hull statistics for the overall fish community and each life-history strategy type.</p>
Transcriptome assemblies of three diatom and three prymnesiophyte isolates from Station ALOHA and Kaneohe Bay
<p><strong>Culture ID/name</strong></p> <p>AT125A – Pseudo-nitzschia sp.</p> <p>AT125C – Pseudo-nitzschia sp.</p> <p>ATCH2 – Chaetoceros sp.</p> <p>Pn B2 – Pseudo-nitzschia sp.</p> <p>KB-HA01 – Chrysochromulina sp. (also called </p> <p>AL-TEMP-12 – Chrysochromulina sp. (also called </p> <p>NF-H275 – Chrysochromulina sp.<br> <br> </p> <p><strong>Growth Conditions</strong></p> <p>All cultures were grown at 27°C, 12:12 light:dark cycle, and with a light intensity of 100 µmol photons m<sup>-2</sup> sec<sup>-1</sup>. AT125C, AT125A, and Pn B2 were grown with Aquil media. ATCH2 was grown with F/20 media with the phosphate concentration modified to a final concentration of 0.5µM. KB-HA01 was grown with F/2 media and AL-TEMP-12 and NF-H275 were grown with K media. None of the cultures were axenic. All cultures were filtered in “light” and “dark” conditions and were in exponential phase when filtered. (Filter types and volumes filtered listed below.) After filtration, all filters were placed into 2mL screwcap tubes, flash frozen with liquid nitrogen, and stored at -80°C.</p> <p><strong>Growth Conditions</strong></p> <p>All cultures were grown at 27°C, 12:12 light:dark cycle, and with a light intensity of 100 µmol photons m<sup>-2</sup> sec<sup>-1</sup>. AT125C, AT125A, and Pn B2 were grown with Aquil media. ATCH2 was grown with F/20 media with the phosphate concentration modified to a final concentration of 0.5µM. KB-HA01 was grown with F/2 media and AL-TEMP-12 and NF-H275 were grown with K media. None of the cultures were axenic. All cultures were filtered in “light” and “dark” conditions and were in exponential phase when filtered. (Filter types and volumes filtered listed below.) After filtration, all filters were placed into 2mL screwcap tubes, flash frozen with liquid nitrogen, and stored at -80°C.<br> <br> [TRANSCRIPTOME SEQUENCING]<br> <br> [QC AND ASSEMBLY]<br> <br> [POST-ASSEMBLY PROCESSING]<br> Diamond v2.0.5.143 was used to blast (e-value: 1e-5) to a cross-kingdom reference sequence database (as described in Coesel et al., 2021). Diamond v2.0.5.143 was used to find the least common ancestor of each contig based upon the blast results. Contigs that were identified as bacteria, archaea, or viruses were excluded from the assemblies.<br> <br> </p>
The OHEJP BeONE Project – Escherichia coli genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 308 <em>Escherichia coli</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7120057">https://zenodo.org/record/7120057</a>), comprising genome assemblies of 1,999 <em>E. coli</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Ec_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers, in-silico Multi Locus Sequence Type and Serotype, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Ec_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>E. coli</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57098">PRJEB57098</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 308 isolates passed the dataset curation step and were included in the final dataset. In-silico serotyping was performed with <a href="https://github.com/B-UMMI/seq_typing">seq_typing</a> v2.2.</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. </p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
The OHEJP BeONE Project – Campylobacter jejuni genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 610 <em>Campylobacter jejuni </em>samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7120166">https://zenodo.org/record/7120166</a>), comprising genome assemblies of 3,076 <em>C. jejuni</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Cj_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers and <em>in-silico</em> Multi Locus Sequence Type, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Cj_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>C. jejuni</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57119">PRJEB57119</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 610 isolates passed the dataset curation step and were included in the final dataset.</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. </p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
The OHEJP BeONE Project – Listeria monocytogenes genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 1,426 <em>Listeria monocytogenes</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7116878">https://zenodo.org/record/7116878</a>), comprising genome assemblies of 1,874 <em>L. monocytogenes</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Lm_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers and <em>in-silico</em> Multi Locus Sequence Type, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Lm_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>L. monocytogenes </em>genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB57166">PRJEB57166</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,426 isolates passed the dataset curation step and were included in the final dataset.</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme.</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p> <p> </p>
The OHEJP BeONE Project – Salmonella enterica genome assembly dataset
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies of 1,540 <em>Salmonella enterica</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (<a href="https://onehealthejp.eu/jrp-beone/">https://onehealthejp.eu/jrp-beone/</a>). Additionally, a complementary dataset is also made available (<a href="https://zenodo.org/record/7119735">https://zenodo.org/record/7119735</a>), comprising genome assemblies of 1,434 <em>S. enterica</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA).</p> <p>File “<strong>BeONE_Se_metadata.xls</strong>x” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers, in-silico Multi Locus Sequence Type and Serotype, and information regarding year of sampling, country and source.</p> <p>The archive “<strong>BeONE_Se_assemblies.zi</strong>p” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>This anonymized dataset of <em>S. enterica</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number <a href="http://www.ebi.ac.uk/ena/browser/view/PRJEB57179">PRJEB57179</a>. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,540 isolates passed the dataset curation step and were included in the final dataset. In-silico serotyping was performed with SeqSero2 v1.2.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31540993/">Zhang et al. 2019</a>).</p> <p> </p> <p><strong>Funding</strong></p> <p>This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. </p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Assemblies of 269 Metagenomic Tara Pacific Sequencing Samples - part 2
<p>This data is the result of the metagenomic assembly of 269 sequencing samples reflecting a first subset of the Tara Pacific metagenomes. Assemblies are used in </p> <p>- Preprint: <a href="https://doi.org/10.1101/2022.04.11.487905">Endogenous viral elements reveal associations between a non-retroviral RNA virus and symbiotic dinoflagellate genomes</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.7839794">Part 1</a></p>
Assemblies of 269 Metagenomic Tara Pacific Sequencing Samples - part 1
<p>This data is the result of the metagenomic assembly of 269 sequencing samples reflecting a first subset of the Tara Pacific metagenomes. Assemblies are used in </p> <p>- Preprint: <a href="https://doi.org/10.1101/2022.04.11.487905">Endogenous viral elements reveal associations between a non-retroviral RNA virus and symbiotic dinoflagellate genomes</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.7840044">Part 2</a></p>
Concentration-, Temperature- and Solvent-Dependent Self-Assembly: Merocyanine Dimerization as a Showcase Example for Obtaining Reliable Thermodynamic Data
<p><strong>Abstract:</strong> Mathematical models for the concentration-, temperature- and solvent-dependent analysis of self-assembly equilibria are derived for the most simple case of dimer formation, to highlight the assumptions these models and the thus determined thermodynamic parameters are based on. The three models were applied to UV/Vis absorption data for the dimerization of a highly dipolar merocyanine dye in 1,4-dioxane. Isothermal titration calorimetry (ITC) dilution experiments were performed as an independent reference technique. While the concentration-dependent analysis is according to our studies the most reliable method, also the less time-consuming temperature-dependent evaluation can give accurate results in the present example, despite small thermochromic effects. In contrast, the strong negative solvatochromism of the merocyanine tampers with the results from the solvent-dependent evaluation. Even though the studies presented in this work are limited to the monomer-dimer equilibrium of a dipolar dye, the basic principles can be transferred to other chromophores and different self-assembly models, including those for supramolecular polymerization.</p>
Photonic crystals with rainbow colors by centrifugation-assisted assembly of colloidal lignin nanoparticles
<p>Source data (CSV files) associated with the publication titled <strong>Photonic crystals with rainbow colors by </strong><strong>centrifugation-assisted assembly </strong><strong>of colloidal lignin nanoparticles</strong>.</p>
Palleja et al. 2018 Metagenome Assemblies for Veseli et al. 2023
<p>A collection of anvi'o contigs databases for 57 human fecal metagenome assemblies generated for the study by Veseli et al. titled "High metabolic independence is a determinant of microbial resilience in the face of gut stress". These are publicly-available gut metagenomes originally obtained from the study by Palleja et al titled "Recovery of gut microbiota of healthy adults following antibiotic exposure" (https://doi.org/10.1038/s41564-018-0257-9). See `PALLEJA_ET_AL_SAMPLES_INFO.txt` file for sample SRA accessions.</p> <p>The metagenomes were assembled individually using IDBA-UD as part of the anvi'o metagenomics workflow in anvi'o v7.1-dev. As part of this workflow, they were annotated with KEGG KOfams using `anvi-run-kegg-kofams` and a KEGG snapshot from December 12, 2020 (modules database hash value `45b7cc2e4fdc`). See manuscript and its reproducible workflow for details.</p>
Genome assembly and annotation files for Corylus americana accessions 'Rush' and 'Winkler'
<p>The native shrub American hazelnut (<em>Corylus americana</em>) is currently used in breeding programs that are aiming to develop commercially viable hazelnut varieties for the U.S. Upper Midwestern U.S. This species provides significant ecological benefits as it is a perennial crop and well-adapted to this region. Breeding cycles for perennial species are long, and may benefit from the use of predictive methods such as genomic selection to reduce cycle time and increase the efficiency of field trials.</p> <p>High-quality reference genome assemblies are very useful for the implementation marker-assisted selection and genomic prediction, and we therefore developed the first chromosome-scale reference assemblies for <em>C. americana</em>, using the accessions 'Rush' and 'Winkler'. Initial draft assemblies were created using HiFi PacBio reads and Arima Hi-C sequencing to assemble genomes into 11 pseudomolecules. We then utilized Oxford Nanopore reads and a high-density genetic map in order to perform error correction. N50 scores were calculated to be 31.9 Mb and 35.3 Mb for 'Rush' and 'Winkler', respectively, while 97.1% (for 'Winkler') and 90.2% (for 'Rush') of the total genome was assembled into the 11 pseudomolecules. Gene prediction was performed using both RNAseq libraries as well as protein homology data. 'Rush' had a BUSCO score of 99.0 for its assembly and 99.0 for its annotation, while 'Winkler' had corresponding scores of 96.9 and 96.5, indicating extremely high-quality assemblies.</p> <p>These two independent, de novo assemblies enable unbiased assessment of structural variation across the genome, as well as patterns of syntenic relationships within C. americana and the <em>Corylus</em> genus. These assemblies are also an important first step in providing a resource for using next-generation sequencing data in the improvement of <em>C. americana</em>. We demonstrate this utility through the generation of high-density SNP marker sets from genotyping-by-sequencing data for 1,343 <em>C. americana</em>, <em>C. avellana</em>, and <em>C. americana</em> x <em>C. avellana</em> hybrids, in order to assess population structure in natural and breeding populations. Finally, the transcriptomes of these assemblies, as well as several other recently published <em>Corylus</em> genomes, were utilized to perform phylogenetic analysis of sporophytic self-incompatibility (SSI) in hazelnut, providing further evidence of unique molecular pathways governing self-incompatibility in Corylus not exhibited in other well-studied SSI systems. We hope these assemblies will aide in the application of modern breeding methods to the development of commercially viable hazelnut varieties for the U.S. Upper Midwest.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.