Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,439
datasets available to search
ShareScore release 0.9.0
Dataset results
2,439 results for “Assembly”
Gene annotation files for Fraxinus excelsior (European ash) genome assembly BATG-0.5
<p>Gene annotation files for <em>Fraxinus excelsior</em> genome assembly v. BATG-0.5, published in Nature (doi:10.1038/nature20786). These files were previously hosted on the Ash Tree Genomes website (http://www.ashgenome.org/transcriptomes) and first made available for download via that site on 2016-02-08.</p> <p>The following annotation files are available:</p> <p>### GFF file of all gene models (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3</p> <p>### FASTA file of all cDNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.pep.fa</p> <p>### Functional annotation for each gene model (all isoforms)<br> Fraxinus_excelsior_38873_TGAC_v2.gff3.functional_annotation.tsv</p> <p>### GFF file of all gene models (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3</p> <p>### FASTA file of all cDNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cdna.fa</p> <p>### FASTA file of all CDS DNA sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.cds.fa</p> <p>### FASTA file of all peptide sequences (longest isoform only)<br> Fraxinus_excelsior_38873_TGAC_v2.longestCDStranscript.gff3.pep.fa</p> <p>### GFF file for gene models identified as probable transposable element related sequences (excluded from the other files)<br> Fraxinus_excelsior_38873_TGAC_v2.transposable_elements.gff3</p> <p><br> NB: The annotation files include preliminary annotations for genes within the organellar scaffolds (gene models FRAEX38873_v2_000400370-FRAEX38873_v2_000401330), which were not reported in the publication of the BATG0.5 assembly (doi:10.1038/nature20786).</p>
Xanthene[n]arenes: Exceptionally Large, Bowl-Shaped Macrocyclic Building Blocks Suitable for Self-Assembly
<p>Data underlying the figures in the publication “Xanthene[<em>n</em>]arenes: Exceptionally Large, Bowl-Shaped Macrocyclic Building Blocks Suitable for Self-Assembly”, published in <em>J</em><em>ACS Au</em> 2021, 1, 11, 1885–1891. <a href="https://doi.org/10.1021/jacsau.1c00343">https://doi.org/10.1021/jacsau.1c00343</a></p>
Scans dataset for evaluation of AR-based assembly pipeline of half-timber dry-stone structures
<p><strong>Dataset description</strong></p> <p><em>The dataset is realized by the EPFL Laboratories of IBOIS, and EESD on the occasion of joint collaboration with the common goal of exploring augmented reality (AR)-based fabrication for half-timber dry-stone construction. It is open sourced and free to be used for anyone's own research.</em></p> <p>The current dataset has been employed to evaluate the designed digital fabrication pipeline. It consists of raw point clouds and reconstructed models of two digitized prototypes of one-layered dry-stone structures with a 1.7 m length, ~0.6 height, and 0.7 m width, each with a different composition of building units. The first wall presents 40 mineral by-products uniquely from by-products of quarry processes involving sawing and water-cutting. In contrast, the second wall features 30 mineral scraps issued from all mixed quarry's operations of transformations.</p> <p>The collection contains both the reconstructed models of the as-built artifacts and the corresponding recorded model of the AR-guided assembly processes for the two walls. The recorded model contains meshes of the placed stones accordingly to the developed geometric planner. By comparing the misalignment between each stone of both models (e.g. Root Mean Square Error distances between two set of points) is possible to obtain metrics about the performance of the proposed AR-assembly method.</p> <p> </p> <p><strong>Dataset labeling guide</strong></p> <table summary="test2"> <thead> <tr> <th scope="row">label</th> <th scope="col">content</th> </tr> </thead> <tbody> <tr> <th scope="row">01/02</th> <td>First or second wall.</td> </tr> <tr> <th scope="row">rs</th> <td>Raw data: unprocessed scans.</td> </tr> <tr> <th scope="row">ab</th> <td>As-built model: model presenting all the registered stones.</td> </tr> <tr> <th scope="row">rc</th> <td>Recorded model: collection of tracked stones during the construction. Each stone has being placed following the AR system is recorded in this sub-set.</td> </tr> <tr> <th scope="row">ev</th> <td>Evaluation set: documentation and working files for the evaluation of the developed AR assembly system.</td> </tr> <tr> <th scope="row">colored_stone</th> <td>Evaluation set outcome: a graphical representation of the misalignment between each stone of the as-built model (ab) with the recorded model (rc).</td> </tr> </tbody> </table> <p> </p> <p><strong>Equipement specs</strong></p> <p>All the raw scans present in the dataset have been obtained from a FARO Freestyle 2 equipped with a Mobile PC for live point cloud processing.</p> <pre><code>@manual{farofreestyle2, title = {FARO Freestyle 2 and Mobile PC}, year = 2022, month = feb, note = {User Manual}, organization = {FARO Technologies Inc.}, url = {https://downloads.faro.com/index.php/s/sqcRBipgSy9GaEq?dir=undefined&openfile=138985} } </code></pre> <p> </p> <p><strong>For extra info</strong></p> <p>The open-sourced code for the developed AR assembly:</p> <pre><a href="https://doi.org/10.5281/zenodo.7181087">https://doi.org/10.5281/zenodo.7181087</a></pre> <p>The complete dataset of digitized mineral scraps:</p> <pre><a href="https://doi.org/10.5281/zenodo.7189478">https://doi.org/10.5281/zenodo.7189478</a></pre> <p> </p> <p><strong>Version notes</strong></p> <p>- The current dataset is published as linked documentation to a future publication, yet not reviewed.</p> <p><strong>Change log</strong></p> <p>- Typos</p> <p>- Correct order of authors</p>
Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 3,076 Campylobacter jejuni isolates
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 2,794-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/4/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 3,076 <em>Campylobacter jejuni </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of Sequence Type [ST]). In total, 476 different STs are represented in this dataset, with ST21, ST50, ST48, ST45 and ST257 being the most represented ones and, together, corresponding to 29.1% of the dataset.</p> <p>File “Cj_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST.</p> <p>The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. </p> <p>The file “profiles/Cj_profiles_wgMLST.tsv” corresponds to a tab separated file with the 2,794-loci wgMLST profiles of each solate presented in the metadata file. The files “profiles/Cj_profiles_cgMLST_95.tsv”, “profiles/Cj_profiles_cgMLST_98.tsv” and “profiles/Cj_profiles_cgMLST_100.tsv” correspond to a 1,012-loci, 987-loci and 29-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>C. jejuni</em> genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://pubmlst.org/organisms/campylobacter-jejunicoli">PubMLST</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 3,539 samples. The majority of them are associated with the INNUENDO project (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>). The remaining ones are associated with five BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB31119">PRJEB31119</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB38253">PRJEB38253</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB40238">PRJEB40238</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB4165">PRJEB4165</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA350537">PRJNA350537</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 3,076 isolates passed this curation step and were included in the final dataset. wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 2,794-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/4">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31st, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mixão et al. 2022</a>) using the 2,794-loci wgMLST profiles of the 3,076 isolates as input and setting distinct “--site-inclusion” thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 1,012-loci, 987-loci and 29-loci allelic matrices, respectively).</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 1,999 Escherichia coli isolates
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 7,601-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/5/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 1,999 <em>Escherichia coli </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of serotype). In total, 411 different serotypes are represented in this dataset, with O157:H7 being the most represented one, corresponding to 37.1% of the dataset.</p> <p>File “Ec_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST and serotype.</p> <p>The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. </p> <p>The file “profiles/Ec_profiles_wgMLST.tsv” corresponds to a tab separated file with the 7,601-loci wgMLST profiles of each isolate presented in the metadata file. The files “profiles/Ec_profiles_cgMLST_95.tsv”, “profiles/Ec_profiles_cgMLST_98.tsv” and “profiles/Ec_profiles_cgMLST_100.tsv” correspond to a 2,826-loci, 2,704-loci and 465-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>E. coli </em>genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://enterobase.warwick.ac.uk/species/index/ecoli">Enterobase</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 2,688 samples associated with three BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA230969">PRJNA230969</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB27020">PRJEB27020</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA248042">PRJNA248042</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,999 isolates passed this curation step and were included in the final dataset. In-silico serotyping was performed with <a href="https://github.com/B-UMMI/seq_typing">seq_typing</a> v2.2. wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 7,601-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/5">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31st, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mixão et al. 2022</a>) using the 7,601-loci wgMLST profiles of the 1,999 isolates as input and setting distinct “--site-inclusion” thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 2,826-loci, 2,704-loci and 465-loci allelic matrices, respectively).</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 1,434 Salmonella enterica isolates
<p><strong>Dataset</strong></p> <p>This dataset comprises the genome assemblies and respective 8,558-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>) available in <a href="https://chewbbaca.online/species/8/schemas/1">chewie-NS</a> (<a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>)] of a final set of 1,434 <em>Salmonella enterica </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the <a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information</a> (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of serotype). In total, 125 different serotypes are represented in this dataset, with Typhimurium (including monophasic), Enteritidis and Infantis being the most represented ones and, together, corresponding to 56.2% of the dataset.</p> <p>File “Se_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST and serotype.</p> <p>The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. </p> <p>The file “profiles/Se_profiles_wgMLST.tsv” corresponds to a tab separated file with the 8,558-loci wgMLST profiles of each isolate presented in the metadata file. The files “profiles/Se_profiles_cgMLST_95.tsv”, “profiles/Se_profiles_cgMLST_98.tsv” and “profiles/Se_profiles_cgMLST_100.tsv” correspond to a 3,261-loci, 3,179-loci and 874-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below.</p> <p> </p> <p><strong>Dataset selection and curation</strong></p> <p>With the objective of creating a diverse dataset of <em>S. enterica</em> genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at <a href="https://enterobase.warwick.ac.uk/species/index/senterica">Enterobase</a> database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 1,779 samples associated with four BioProjects (<a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB16326">PRJEB16326</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB20997">PRJEB20997</a>, <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB30335">PRJEB30335</a> and <a href="https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB39988">PRJEB39988</a>). Their WGS data was downloaded from ENA/SRA with <a href="https://github.com/rpetit3/fastq-dl">fastq-dl</a> v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8145556/">Deneke et al. 2021</a>) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,434 isolates passed this curation step and were included in the final dataset. In-silico serotyping was performed with SeqSero2 v1.2.1 (<a href="https://pubmed.ncbi.nlm.nih.gov/31540993/">Zhang et al. 2019</a>). wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (<a href="https://pubmed.ncbi.nlm.nih.gov/29543149/">Silva et al. 2018</a>), using the 8,558-loci INNUENDO schema available in <a href="https://chewbbaca.online/species/8">chewie-NS</a> (<a href="https://efsa.onlinelibrary.wiley.com/doi/epdf/10.2903/sp.efsa.2018.EN-1498">Llarena et al. 2018</a>; <a href="https://academic.oup.com/nar/article/49/D1/D660/5929238">Mamede et al. 2022</a>) and downloaded on May 31<sup>st</sup>, 2022. Three cgMLST schemas were obtained with <a href="https://github.com/insapathogenomics/ReporTree">ReporTree</a> v1.0.0 (<a href="https://www.researchsquare.com/article/rs-1404655/v1">Mixão et al. 2022</a>) using the 8,558-loci wgMLST profiles of the 1,434 isolates as input and setting distinct “--site-inclusion” thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 3,261-loci, 3,179-loci and 874-loci allelic matrices, respectively).</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p>We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.</p>
Elliptical Alignment Holes Enabling Accurate Direct Assembly of Microchips to Standard Waveguide Flanges at sub-THz Frequencies - Dataset
<p>Current waveguide flange standards do not allow for the accurate fitting of microchips, due to the large mechanical tolerances of the flange alignment pins and the brittle nature of Silicon, requiring greatly oversized alignment holes on the chip to fit worst-case fabrication tolerances, resulting in unacceptably large misalignment error for sub-THz frequencies. This paper presents, for the first time, a new method for directly aligning micromachined Silicon chips to standard, i.e. unmodified, waveguide flanges with alignment accuracy significantly better than the waveguide-flange fabrication tolerances, through the combination of a tightly-fitting circular and an elliptical alignment hole on the chip. A Monte Carlo analysis predicts the reduction of the mechanical assembly margin by a factor of 5.5 compared to conventional circular holes, reducing the potential chip misalignment from 46 μm to 8.5 μm for a probability of fitting of 99.5%. For experimental verification, micromachined waveguide chips using either conventional (oversized) circular or the proposed elliptical alignment holes were fabricated and measured. A reduction in the standard deviation of the reflection coefficient by a factor of up to 20 was experimentally observed from a total of 200 measurements with random chip placement, exceeding the<br> expectations from the Monte Carlo analysis. To our knowledge, this paper presents the first solution for highly accurate assembly<br> of micromachined waveguide chips to standard waveguide flanges, requiring no custom flanges or other tailor-made split blocks.<br> </p>
Cladonema radiatum Alr and IgSF genes assembly and domain predictions
<p><span>This dataset is related to the </span><span>submitted paper</span><span> " A single gene determines allorecognition in hydrozoan jellyfish <em>Cladonema radiatum</em> inbred lines ".</span></p> <h1><span>Abstract:</span></h1> <p><strong><span> </span></strong></p> <p><span>Allorecognition—the ability of an organism to discriminate between self and non-self—is crucial to colonial marine animals to avoid invasion by other individuals in the same habitat. The cnidarian hydroid <em>Hydractinia</em> has long been a major research model in studying invertebrate allorecognition, establishing a rich knowledge foundation. In this study, we introduce a new cnidarian model <em>Cladonema radiatum</em> (<em>C. radiatum</em>). <em>C. radiatum</em> is a hydroid jellyfish which also forms polyp colonies interconnected with stolons. Allorecognition responses, fusion or regression of stolons, are observed when stolons encounter each other. By transmission electron microscopy, we observe rapid tissue remodelling contributing to gastrovascular system connection in fusion. Rejection responses are regulated by reconstruction of the chitinous exoskeleton perisarc, and induction of necrotic and autophagic cellular responses at cells in contact with the opponent. Genetic analysis identifies allorecognition genes: six<em> Alr </em>genes located on the putative Allorecognition Complex (ARC) and four immunoglobulin superfamily genes on a separate genome region. C. radiatum allorecognition genes show notable conservation with the <em>Hydractinia Alr</em> family. Remarkedly, stolon encounter assays of inbred lines reveal that genotypes of Alr1 solely determine allorecognition outcomes in <em>C. radiatum</em>.</span></p>
Structural conversion of the spidroin C-terminal domain during assembly of spider silk fibers
<p>GENERAL INFORMATION<br>- Dataset title: Structural conversion of the spidroin C-terminal domain during assembly of spider silk fibers<br>- Description: The dataset contains raw data associated with the publication with the same name, accepted for publication in Nature Communications.<br>- Authors: Danilo Hirabae De Oliveira, Vasantha Gowda, Tobias Sparrman, Linnea Gustafsson, Rodrigo Sanches Pires, Christian Riekel, Andreas Barth, Christofer Lendel, My Hedhammar </p> <p>ORGANIZATION<br>The folder contains zip-files for each figure in the publication. Each zip-file contains data and a .txt file describing the content, the methods for data acquisition and analysis, and the file types.</p> <p><br>DATA COLLECTION<br>Data collection and analysis is described in the paper and in the .txt files included in each zip-file.</p>
Assemblies, synapse clustering and network topology interact with plasticity to explain structure-function relationships of the cortical connectome
<p>Dataset linked to the article with the same title</p> <p>The model itself is very similar to its non-plastic counterpart under the following DOI: <a href="../record/7930275">10.5281/zenodo.7930275</a>, i.e. a 1.5 mm diameter cortical tissue comprising 211,712 neurons and their connectivity in the front limb and jaw subregions and the dysgranular zone of the Paxinos & Watson rat brain atlas. It's formatted in the open <a href="https://github.com/AllenInstitute/sonata">SONATA</a> standard and contains neuron locations and their properties (such as morphological types, cortical layer, etc.), their detailed morphologies, and synaptic connectivity (with all their anatomical and physiological parameters). The main difference from the non-plastic version is the addition of plasticity related parameters to <em>O1/S1nonbarrel_neurons__S1nonbarrel_neurons__chemical/edges.h5. </em>Extrinsic synaptic connections from the thalamus are included in this release, but for inputs from neurons in the remainder of non-barrel somatosensory cortex please see the non-plastic version of the circuit.</p> <p><strong>Analyzing the model</strong></p> <p>The model can be analyzed in terms of its anatomy, physiology and connectivity using the packages <a href="https://neurom.readthedocs.io/en/stable/">NeuroM</a>, <a href="https://bluebrainsnap.readthedocs.io/en/stable/">BlueBrain SNAP</a> and <a href="https://github.com/BlueBrain/ConnectomeUtilities">ConnectomeUtilities</a>. (see first Jupyter notebook)</p> <p><strong>Simulating the model</strong></p> <p>To simulate the model we'd recommend using out using our open-source simulator <a href="https://github.com/BlueBrain/neurodamus">Neurodamus</a>. The reference version is the branch <em>nbS1-2023</em>, which is archived under the following DOI: <a href="http://doi.org/10.5281/zenodo.8075202">10.5281/zenodo.8075202</a>. Instructions on how to use the simulator are provided on the GitHub page linked above. Briefly, you'll first have to <a href="https://github.com/BlueBrain/neurodamus#install-neurodamus">install Neurodamus</a>. Next, build a <em>"special"</em> executable that include compiled versions of ion channel and synapse models. To do that, follow <a href="https://github.com/BlueBrain/neurodamus#build-special-with-mod-files">these instructions</a>, where <em>mod-files-from-released-circuit </em>is replaced by the location of <em>O1/mods</em> on your system. Finally, <a href="https://github.com/BlueBrain/neurodamus#examples">run a simulation</a>. The specific simulation conditions and stimuli are specified in simulation configuration files. An exemplary simulation configuration is included in this release (<em>simulation_config.zip</em>).</p> <p><strong>Analyzing simulation results</strong></p> <p>Simulation results can be analyzed with <a href="https://bluebrainsnap.readthedocs.io/en/stable/">BlueBrain SNAP</a>, <a href="https://github.com/BlueBrain/ConnectomeUtilities">ConnectomeUtilities</a>, and <a href="https://github.com/BlueBrain/assemblyfire">assemblyfire</a>. Notebooks 2-5 go though these analysis and recreate some of the panels from our article. In most cases the notebooks can be run with the shared HDF5 files and don't require running any simulations.</p> <p><strong>Version 2</strong></p> <p>Bug fix in simulation_config.json and therefore new version of results (and corresponding notebooks). The underlying circuit model (O1.xz) did not change from v1.</p> <p>--</p> <p><em>The development of this dataset was supported by funding to the Blue Brain Project, a research center of the École polytechnique fédérale de Lausanne (EPFL), from the Swiss government’s ETH Board of the Swiss Federal Institutes of Technology.</em></p>
A Simulated Heterozygous Diploid Genome for Third-gen Sequencing, Assembly, and Curation
<p>A simulated heterozygous diploid genome based on <em>Saccharomyces</em> <em>cerevisiae</em>, and <em>S. paradoxus</em> homologous chromosomes.</p> <p>Simulated PacBio subreads were generated from both parent haplomes and mixed together. A phased assembly was produced using FALCON assembler and FALCON Unzip (doi:10.1038/nmeth.4035). This dataset and assembly were then used to validate the Purge Haplotigs pipeline (https://bitbucket.org/mroachawri/purge_haplotigs). See workflow.sh for commands, comments and file descriptions.</p>
Genome assemblies for "Versatile genome assembly evaluation with QUAST-LG"
<p>Supplementary data for A. Mikheenko, A. Prjibelski, V. Saveliev, D. Antipov, A. Gurevich. Versatile genome assembly evaluation with QUAST-LG. ISMB 2018 PROCEEDINGS (<em>Bioinformatics </em>journal)</p> <p>Reference genomes of</p> <ol> <li><em>Saccharomyces cerevisiae </em>(yeast) version R64-1-1 </li> <li><em>Caenorhabditis elegans </em>(worm) version WBcel235</li> <li><em>Drosophila melanogaster </em>(fruit fly) version BDGP6</li> </ol> <p>And<em> de novo</em> genome assemblies of</p> <ol> <li>Yeast_PB (<em>S. cerevisiae</em>, genome size: 12.1 Mb): Canu, FALCON, Flye, MaSuRCA, Miniasm (from Illumina pair-ends and PacBio SMRT)</li> <li>Yeast_NP (<em>S. cerevisiae</em>, genome size: 12.1 Mb): Canu, Flye, MaSuRCA, Miniasm (from Illumina pair-ends and Oxford Nanopore)</li> <li>Worm_PB (<em>C. elegans</em>, genome size: 100.3 Mb): Canu, FALCON, Flye, MaSuRCA, Miniasm (from Illumina pair-ends and PacBio SMRT)</li> <li>Fly_MP (<em>D. melanogaster</em>, genome size: 137.6 Mb): ABySS2, MaSuRCA, Meraculous, Platanus, SOAPdenovo2, SPAdes (from Illumina pair-ends and mate-pairs)</li> <li>Human_MP (<em>H. sapiens</em>, genome size: 3.1 Gb): UpperBound assembly only (from Illumina pair-ends and mate-pairs)</li> <li>Human_NP (<em>H. sapiens</em>, genome size: 3.1 Gb): UpperBound assembly only (from Illumina pair-ends and Oxford Nanopore)</li> </ol> <p>Each pack (items 1-4) is accompanied with the upper bound assembly created with QUAST-LG (for computing theoretical limits on assembly correctness and completeness for a particular genome and set of reads). For more information, interactive QUAST-LG reports, and links to <em>de novo</em> assemblies of the human datasets please visit http://cab.spbu.ru/software/quast-lg/ or write to <em>quast.support@cab.spbu.ru.</em></p>
Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes
<p> </p> <p><strong>Annotation of metagenome-assembled genomes retrieved from Amazon river basin metagenomes</strong></p> <p> </p> <p> RELEASE MAG-2018/01<br> --------------------------------------</p> <p> </p> <p>1. INTRODUCTION</p> <p>Here is deposited the genes and proteins annotation from metagenome-assembled genomes (MAGs) retrieved from Amazon river basin metaganomes (SRP044326, PRJEB25171 and SRP039390) were deposited under European Nucleotide Archive - ENA project PRJEB25176. Briefly, metagenomes were coassembled in groups by geographical location with Megahit v.1.0 and the contigs were used to a reads mapping and binning with BWA-MEM (version 0.7.12-r1039), SamTools (version 1.3.1) and Metabat (v2.12.1). MAGs with overall quality greater than 50, calculated with CheckM (version 1.0.11), were selected for refining precedures. Contigs outliers were eliminated by using RefineM (version 0.0.23). Finished MAGs were then annotated by Prokka (version 1.11) pipeline, and with the other most completes databases up to date (KEGG, UniProtKB, dbCAN, PFAM, eggNOG and COG).</p> <p> </p> <p>2. LOCATION</p> <p> </p> <p> MAGs sequences are available under ENA project PRJEB25176.</p> <p> </p> <p> ENA_accession Isolate<br> -------------------- --------------<br> ERZ494218 AM_0118<br> ERZ494219 AM_0219<br> ERZ494220 AM_0226<br> ERZ494221 AM_0228<br> ERZ494222 AM_0233<br> ERZ494223 AM_0240<br> ERZ494224 AM_0244<br> ERZ494225 AM_0256<br> ERZ494226 AM_0268<br> ERZ494227 AM_0275<br> ERZ494228 AM_0466<br> ERZ494229 AM_0507<br> ERZ494230 AM_0510<br> ERZ494231 AM_0519<br> ERZ494232 AM_0528<br> ERZ494233 AM_0546<br> ERZ494234 AM_0608<br> ERZ494235 AM_0615<br> ERZ494236 AM_0616<br> ERZ494237 AM_0619<br> ERZ494238 AM_0621<br> ERZ494239 AM_0630<br> ERZ494240 AM_0643<br> ERZ494241 AM_0729<br> ERZ494242 AM_0764<br> ERZ494243 AM_0832<br> ERZ494244 AM_0849<br> ERZ494245 AM_0854<br> ERZ494246 AM_0876<br> ERZ494247 AM_0902<br> ERZ494248 AM_0936<br> ERZ494249 AM_1003<br> ERZ494250 AM_1104<br> ERZ494251 AM_1111<br> ERZ494252 AM_1205<br> ERZ494253 AM_1312<br> ERZ494254 AM_1409<br> ERZ494255 AM_1503<br> ERZ494256 AM_1603<br> ERZ494257 AM_1606<br> ERZ494258 AM_1801<br> ERZ494259 AM_1811<br> ERZ494260 AM_2104<br> ERZ494261 AM_2116<br> ERZ494262 AM_2124<br> ERZ494263 AM_2202<br> ERZ494264 AM_2207<br> ERZ494265 AM_2208<br> ERZ494266 AM_2324<br> ERZ494267 AM_2502<br> ERZ494268 AM_2804<br> </p> <p>3. ACKNOWLEDGEMENTS<br> </p> <p>This work is a joint effort of Laboratory of molecular biology from Federal<br> University of São Carlos, São Paulo, Brazil (LBM/UFSCAR) and Protists group<br> of Institut del Ciencias del Mar, Barcelone, Spain (ICM). We are grateful to<br> Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq), as well as, the spanish funding organ Consejo Superior de Investigaciones Científicas (CSIC).</p> <p>This study was financed in part by the Coordenação de Aperfeiçoamento de Pessoal de Nível Superior - Brasil (CAPES) - Finance Code 001.</p> <p> </p> <p>4. CONTACT INFORMATION</p> <p> Current curators:</p> <p> - Célio Dias Santos Júnior (celio.diasjunior@gmail.com)<br> - Flavio Henrique-Silva (dfhs@ufscar.br)<br> - Ramiro R. Logares (ramiro.logares@icm.csic.es)<br> </p> <p>5. COPYRIGHT NOTICE</p> <p> Amazon River Basin Metagenome-Assembled Genomes Annotation - AM/MAGs<br> Copyright (C) 2018 The AMnrGC consortium.</p> <p> This database is provided “as is” and without any warranty of any kind,<br> of openly available. You can redistribute and/or modify it<br> as you wish, under the terms of Creative Commons CC BY 4.0:</p> <p> https://creativecommons.org/licenses/by/4.0/</p> <p>___________________<br> Barcelone, Feb/2018</p>
Simulated Arabidopsis thaliana sequencing datasets for chloroplast assembler benchmarking
<p><strong>Changes</strong></p> <ul> <li>Fixed non-circular sampling from chloroplast and mitochondrion in version 1.1.0</li> <li>Fixed off-by-one error in reverse read in version 1.0.0</li> </ul> <p><strong>Purpose and Documentation</strong></p> <p>See: <a href="https://github.com/chloroExtractorTeam/benchmark">github.com/chloroExtractorTeam/benchmark</a></p> <p><strong>Original data</strong><br> The original <em>Arabidopsis thaliana </em>sequences were downloaded from TAIR: </p> <p>The Arabidopsis Information Resource (<a>TAIR</a>) on www.arabidopsis.org, Mar 22, 2019 available under the <a href="http://www.arabidopsis.org/doc/about/tair_terms_of_use/417">TAIR Terms of Use</a> </p> <p><em>Tanya Z. Berardini, Leonore Reiser, Donghui Li, Yarik Mezheritsky, Robert Muller, Emily Strait and Eva Huala. "The Arabidopsis Information Resource: Making and mining the "gold standard" annotated reference plant genome." genesis 2015 <a href="https://doi.org/10.1002/dvg.22877">doi:10.1002/dvg.22877</a></em></p> <p><strong>Programs used to generate this data</strong><br> - <a href="https://github.com/shenwei356/seqkit">seqkit</a> (v0.10.1): Shen W, Le S, Li Y, Hu F (2016) "SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA/Q File Manipulation." PLOS ONE 11(10): e0163962. <a href="https://doi.org/10.1371/journal.pone.0163962">doi:10.1371/journal.pone.0163962</a></p> <p> </p>
GenoNet scores for human genome assembly GRCh37
<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files. </p> <p>Each row represents a genomic region with 131 columns. Please find the header line in "genonet.header.txt". </p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 10000 10025 chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37 <a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover <a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>
De novo genome assembly of the meadow brown butterfly, Maniola jurtina
<p>1. Whole-genome GFF file (raw and filtered for min. gene length) [<em>Maniola.jurtina.gff3</em>, <em>Maniola_jurtina_filtered.gff3</em>]</p> <p>2. List of <em>M. jurtina</em> proteins [<em>Mjurtina_proteins.fa</em>].</p> <p>3. Results of spot pattern genes BLAST against <em>M. jurtina</em> proteome [<em>Lepidoptera_MJ_protein_matches.xlsx</em>]. </p> <p>4. Annotations [blast2go_export.txt]</p>
ST131_4071_genome_assembly_annotation_files
<p>The genome assembly annotation files of 4,071 E. coli ST131 genomes (see Decano & Downing 2019).</p>
Sardinops sagax genome assemblies and annotations
<p>Files included are the genome assemblies for each haplotype (hap 1 and hap 2) of Sardinops sagax and the corresponding annotation files for each haplotype.</p> <p> </p> <p> </p>
LyBar v2.0 Genome Assembly and Annotation for Lycium barbarum
<p>LyBar v2.0 genome assembly and annotation files for <em>Lycium barbarum</em>.</p>
Software, Dataset, and Techreport: Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration
<p>This upload contains a techreport titled "Mixed-precision finite element kernels and assembly: Rounding error analysis and hardware acceleration" together with the software (with documentation) and dataset generating the results. The software is also available on GitHub at https://github.com/croci/mpfem-paper-experiments-2024/ . The GitHub version may be updated in the future. This upload corresponds to commit number 8506dd368b84655201c8c72b1307239b9b4e43fd . See README.md file for installation instructions. The manuscript is also available on the arXiv: https://arxiv.org/abs/2410.12614.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.