Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,439
datasets available to search
ShareScore release 0.9.0
Dataset results
2,439 results for “Assembly”
Tracking Down Chimeric Assemblies In The TrackIt DNA Ladder Using Nanopore Sequencing
<h2>Dataset Description</h2><p>These files represent two different LSK114 sequencing runs on a TrackIt 1kb Plus DNA Ladder sample, and associated data analysis.</p><ul><li>July 20 2023 Flongle Run (191 Mb; 465k reads)<ul><li>pod5_files_2023-Jul-20_DAE_DNA_Ladder.tar.gz<br>- raw POD5 format files</li><li>called_2023-Jul-20_DAE_DNA_Ladder_duplex.bam<br>- duplex called reads, called using dorado v4.0 with the 2023-09-22 bacterial methylation model</li><li>sequence_QC_2023-Jul-20_DAE_DNA_Ladder.pdf<br>- sequence length / quality QC plots</li><li>LAST_2023-Jul-20_DAE_DNA_Ladder_reads_vs_reference.tar.gz<br>- Alignment summary statistics from LAST mapping of reads to their associated reference</li><li>lengths_summary_2023-Jul-20_DAE_DNA_Ladder.txt<br>- Length / QC summary statistics</li></ul></li><li>October 12 2023 P2 Solo Run (1.95 Gb, 1.11M reads)<ul><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_fail.tar.gz<br>- raw POD5 format files (all failed reads)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_000-059.tar.gz<br>- raw POD5 format files (passed reads, bundle #000-059)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_060-119.tar.gz<br>- raw POD5 format files (passed reads, bundle #060-119)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_120-179.tar.gz<br>- raw POD5 format files (passed reads, bundle #120-179)</li><li>pod5_files_2023-Oct-12_DNA-Ladder-1kbplus_pass_180-222.tar.gz<br>- raw POD5 format files (passed reads, bundle #180-222)</li><li>called_2023-Oct-12_DNA-Ladder-1kbplus_duplex.bam<br>- duplex called reads [October 12, 2023], called using dorado v4.0 with the 2023-09-22 bacterial methylation model</li><li>sequence_QC_2023-Oct-12_DNA-Ladder-1kbplus.pdf<br>- sequence length / quality QC plots</li><li>LAST_2023-Oct-12_DNA-Ladder-1kbplus_reads_vs_reference.tar.gz<br>- Alignment summary statistics from LAST mapping of reads to their associated reference</li><li>lengths_summary_2023-Oct-12_DNA-Ladder-1kbplus.txt<br>- Length / QC summary statistics</li><li>ladder_seqs.fa<br>- assembled DNA ladder sequences, based on simplex reads</li></ul></li></ul><h3>Methods</h3><h3>Sample preparation</h3><p>Preparation of DNA for sequencing was carried out following the ONT Ligation Sequencing DNA V14 (SQK-LSK114) protocol, with modifications to exclude DNA repair, and keeping the sample in the same 1.5ml tube to reduce sample loss.</p><h4>Tris-buffered Saline (TBS) buffer preparation</h4><ol><li>1M stock of NaCl was made by adding 2.922g of NaCl into a 50 ml Falcon tube, then made up to 50 ml with MilliPore water</li><li>A 50 mM TBS stock was created by adding 750 μl 1M NaCl solution to a 15 ml Falcon tube, then made up to 15 ml using Qiagen Elution Buffer (EB, i.e. 10 mM Tris-HCl at pH 8.0)</li><li>The pH was confirmed to be 7.9-8.1 using a pH indicator strip (e.g. MColorpHast 6.5 - 10.0; MER1095430001)</li></ol><h4>End prep</h4><ol><li>1 μg DNA ladder (i.e. 10 μl of 0.1 μg / μl DNA ladder) was transferred into a 1.5ml Eppendorf DNA LoBind tube</li><li>The volume was topped up to 43.5 μl with TBS (i.e. 33.5 μl TBS)</li><li>3.5 μl Ultra II End-prep Reaction Buffer and 3 μl Ultra II End-prep Enzyme Mix was added</li><li>After mixing by gentle pipetting, the mixture was incubated at RT for 5 minutes, then 65 \degrees for 5 minutes</li></ol><h4>Bead cleanup</h4><ol><li>The mixture was combined with 60 μl Ampure XP beads, and incubated on a rotator mixer at RT for 5 minutes</li><li>The tube was transferred to a magnetic rack [https://www.printables.com/model/532085-open-walled-magnetic-rack]</li><li>After the supernatant became clear and colourless, supernatant was pipetted off</li><li>The magnetic beads were washed twice with 150 μl of an 80% ethanol solution</li><li>The sample was dried briefly for 30s, then eluted in 60 μl TBS</li></ol><h4>Adapter ligation and final bead cleanup</h4><ol><li>To the sample tube was added 25 μl ONT Ligation buffer (LNB), 5μl NEBNext Quick T4 DNA Ligase (reduced from the protocol-suggested 10μl because that was all that was left in the tube), and 5μl ONT Ligation Adapter (LA)</li><li>The tube was mixed by gentle pipetting, spun down for 1-3s on a mini centrifuge, then incubated for 10 minutes at RT</li><li>The mixture was combined with 40 μl Ampure XP beads (100μl Ampure XP beads were used for the Flongle sample), and incubated on a rotator mixer at RT for 5 minutes</li><li>The tube was transferred to a magnetic rack [https://www.printables.com/model/532085-open-walled-magnetic-rack]</li><li>After the supernatant became clear and colourless, supernatant was pipetted off</li><li>The magnetic beads were washed twice with 250 μl of ONT Long Fragment Buffer (LFB) for the P2 Solo run, and 250μl ONT Short Fragment Buffer <br>(SFB) for the Flongle run</li><li>The sample was dried briefly for 30s, then eluted for 10 minutes at 37 \degrees in 15 μl ONT Elution buffer (EB)</li></ol><h4>Addition of sequencing library buffers</h4><ol><li>A flow cell was prepared by flushing with ONT Flow Cell Flush (FCF) mixed with ONT Flow Cell Tether (FCT). For the P2 Solo, I used 500 μl of a 1170μl FCF solution that had 30μl FCT added to it; for the Flongle I used 60μl of a 117μl FCF solution that had 3 μl FCT added to it</li><li>1 μl of the eluted library was quantified on a Quantus Fluorometer, and approximately 50 fmol (assuming 1kb average length) was transferred to a new 1.5μl tube</li><li>For the P2 Solo run, the volume was topped up to 32 μl TBS; for the Flongle run, the volume was topped up to 12 μl TBS</li><li>To the sample tube was added ONT Sequencing Buffer (SB; P2 Solo - 100μl; Flongle - 30μl) and ONT Library Beads (LIB; P2 Solo - 68μl; Flongle - 20μl)</li><li>The flow cell was re-flushed with additional FCF/FCT mixture (500 μl for the P2 Solo; 30 μl for the Flongle)</li><li>The sequencing library was then added to the flow cell (200 μl for the P2 Solo; 30 μl for the Flongle)</li><li>The prepared flow cell was left for 10 minutes to allow the library to settle before starting sequencing</li></ol><h4>DNA Sequencing and basecalling</h4><ol><li>Sequencing was carried out using MinKNOW v23.04.6, sequencing in fast mode at 400 bases per second with a 20bp minimum sequence length and <br>5 kHz sampling rate, with reads output as POD5 files</li><li>The Flongle flow cell was run for a full standard run length (24h), whereas the PromethION flow cell was run for 1.5 hours (after which the <br>counts of 15kb reads exceeded 200)</li><li>Sequenced reads were recalled in standard (simplex) mode using Dorado v0.4.0 and the 2023-09-22 bacterial methylation model [res_dna_r10.4.1_e8.2_400bps_sup@2023-09-22_bacterial-methylation]</li></ol><h3>Bioinformatics Analysis of Ladder Sequences </h3><h4>Sequence assembly</h4><p>Assembly process for bands that are 3k in length and greater (done on LFB-depleted P2 Solo sequences): </p><ol><li>Filter >q20 reads for a 100bp region around the target length (e.g. 4950-5050bp for the 5k band) [High quality reads were not sufficient for the 15kb band; all reads were needed]</li><li>Chop the reads up with a 1000bp overlap (e.g. 3000bp for the 5k band). This works around a Canu expectation that any read overlaps should be less than X% of the read.</li><li>Assemble the reads with Canu v2.2 [#REF], treating them as "pacbio" reads (for correction and homopolymer compression), with the GenomeSize parameter set to the expected band length (e.g. GenomeSize=5000).</li><li>Extract the first reported assembled contig.</li><li>Map the contig to the nanopore adapter sequences, and trim to exclude any matching sequence.</li></ol><p>[Canu has a default genome size and read length cutoff of 1kb, and performs poorly on sequences shorter than this] <br><br>Assembly process for bands under 3k in length (done on LFB-depleted P2 Solo sequences):</p><ol><li>Filter >q20 reads for a 100bp region around the target length (e.g. 4950-5050bp for the 5k band) [High quality reads were not in sufficient abundance for the 100bp band; all reads were needed]</li><li>Assemble using a<a href="https://gitlab.com/gringer/bioinfscripts/-/blob/master/fastx-kassembler.pl"> kmer-based de-bruijn assembler</a>, trimming off low-count kmers</li><li>Extract the first reported trimmed assembled chain</li><li>Map the assembled chain to the nanopore adapter sequences, and trim to exclude any matching sequence</li><li>Use web BLASTn [#REF] to help trim any additional trailing non-matching sequence</li></ol><h4>Mapping</h4><ol><li>Use a <a href="https://gitlab.com/gringer/bioinfscripts/-/blob/master/fastx-kmapper.pl">kmer-based lightweight mapper</a> to map reads to assembled bands</li><li>Created LAST mismatch matrix using `last-train` on the 5k reads together, using the full assembled ladder sequences as a reference:<br>last-train -Q 1 ladder_seqs.fa 5k_reads.fq.gz</li><li>Mapped all reads to the assembled ladder sequences (only the reference corresponding to the most likely band source), retaining (for each read) the mapping that had the longest combined proportion of read and reference sequence mapped:<br>lastal -p bacterial.mat -P 10 ladder_seqs.fa reads_2023-Oct-12_DNA-Ladder-1kbplus_called_all.fq.gz | \ <br> ~/scripts/maf2csv.pl | \ <br> awk -F ',' '{print $0","($8/100 * $13/100)}' | \ <br> sort -t ',' -k 16rg,16 | sort -t ',' -k 1,1 -u | sort -t ',' -k 1r,1 | \ <br> perl -pe 's/,[^,]*$/\n/' > LAST_reads_vs_ladder_longestMatch.csv.gz</li></ol>
Extending the executability of assembly task poses by robot through end-effectors
<p><strong>Extending the executability of assembly task poses by robots through end-effectors</strong></p> <p><br><em>Aline Kluge-Wilkes, Presley Demuner Reverdito</em></p> <p>The following data set was created during the validation of a proposed method to evaluate assembly station formations considering the executability of assembly tasks, including the effects of equipped end-effectors on robots.</p> <p>The underlying paper can be found at: https://doi.org/10.1007/978-3-031-34821-1_58 </p> <p>The underlying source code can be found at: https://git-ce.rwth-aachen.de/wzl-mq-public/iop/ws-b2.iv_formation-planning-of-mobile-robots/end-effector-dependent-executability</p> <p>In dependence on the geometries and degrees of freedom of the equipped end-effectors on the robot, a single task pose is transferred into an area of feasible task poses. Determining the executability of the task, the resulting representation of the feasible task poses is overlapped with the robot's workspace. If an overlap occurs, it can be assumed that there is a feasible robot configuration to execute the allocated assembly task. The proposed method is exemplified on a UR10 equipped with a screwdriver, and distributed task poses. The proposed method quantifies the end-effector's effect on the executability of assembly tasks and provides a means of determining the feasible base placement of robots in changeable assembly stations. Therefore, the method lays the foundation for automated formation planning in assembly stations.</p> <p><strong>Design of Experiments:</strong></p> <p>The published data here results from a conducted series of experiments structured as a full-factorial design of experiments. The following variables and expressions / data points of those variables were chosen:</p> <p>Tasks ("GoalPose") poses as [x, y, z, x-quaternion, y-quaternion, z-quaternion, w-quaternion]:</p> <ul> <li> Task 1 = (0.416m ,-0.399m, 0.764m, 0.071, 0.703, 0.134, 0.694)</li> <li> Task 2 = (0.438m, -0.647m, 0.816m, 0.005, 0.707, 0.068, 0.704)</li> <li> Task 3 = (0.448m, -0.755m, 0.945m, -0.243, 0.664 ,-0.183, 0.683)</li> </ul> <p>Dimension of the tools ("DimensionOfTool") as [x, y, z]: </p> <ul> <li> (0.1m, 0.2m, 0m)</li> <li> (0.2m, 0.1m, 0m)</li> </ul> <p>Robot model ("RobotModel"): </p> <ul> <li> UR10 (https://www.universal-robots.com/de/produkte/ur10-roboter/)</li> <li> UR5 (https://www.universal-robots.com/products/ur5-robot/)</li> </ul> <p>Resolution of reachability map ("ReachMap"): </p> <ul> <li> 0.05m</li> <li> 0.08m</li> <li> 0.1m</li> </ul> <p>IK solver ("IKSolver"): </p> <ul> <li>TRAC-IK <ul> <li>documented in: http://docs.ros.org/en/kinetic/api/moveit_tutorials/html/doc/trac_ik/trac_ik_tutorial.html </li> <li>source: https://bitbucket.org/traclabs/trac_ik/src/master/</li> </ul> </li> <li>KDL solver <ul> <li>https://docs.orocos.org/kdl/overview.html</li> </ul> </li> </ul> <p>Based on these definitions, a fully factorial design was implemented, resulting in 72 required tests (2³ × 3²). These tests were randomly organized into a single experimental block.</p> <p><strong>Test execution:</strong></p> <p>The function "[calcForAllTasks](https://git-ce.rwth-aachen.de/wzl-mq-public/iop/ws-b2.iv_formation-planning-of-mobile-robots/end-effector-dependent-executability/-/blob/main/code/main.py?ref_type=heads)" on the file main.py was created to generate 72 files, each containing 10 poses of the circle of possible poses for the robot with a calculated reachability map. These poses were manually tested using MoveIt! (https://moveit.ros.org/) and using two different IK solvers, KDL and TRAC-IK.</p> <p>There were now four packages: ur5_kdl, ur5_trac_ik, ur10_kdl, and ur10_trac_ik. Each package had its own group_name, which was the same as the package name. For each test, it was necessary to run the MoveIt! file of the robot with the correspondent IK solver, and the [move_group_python_interface.py](https://git-ce.rwth-aachen.de/wzl-mq-public/iop/ws-b2.iv_formation-planning-of-mobile-robots/end-effector-dependent-executability/-/blob/main/code/move_group_python_interface.py?ref_type=heads) with the desired pose. On the Python code, the goal pose was set, and using the MoveGroupPythonIntefaceTutorial class, Moveit! attempt to move the robot to the desired pose. If the pose is reachable, the reachability index was set to 1. Otherwise, it was set to 0, and the terminal output would show "ABORTED: No motion plan found. No execution attempted."</p> <p>After simulating all poses, the reachability average of each test was calculated, along with the sum of reachable results. The reachability index ranged from 0 to 100, and the reachable column ranged from 0 to 10. This result can be found in the file DoE-tests-and-results.xlsx.</p> <p> </p> <p>The file "DoE-tests-and-results.xlsx" provides an overview of all experiments in its first table. The creation order ("StdOrder") and the order in which the experiments were conducted are given ("RunOrder"). The following columns in the first table indicate the expressions of the variables as explained above (GoalPose, DimensionOfTool, RobotMode, ReachMap, IKSolver). Next, the results of the experiments are given: the average reachability index per tested goal pose and the indication of how many of the ten discrete robot flange poses per goal pose are reachable by the robot (Reachability Index, Reachable). </p> <p>The following 72 tables each provide the ten robot flange poses per goal pose per experiment as [x, y, z, x-quaternion, y-quaternion, z-quaternion, w-quaternion]. </p> <p>The 72 .csv files contain the same information as the 72 tables in the Excel file but contain the information as it was transferred during execution. </p> <p> </p> <p> </p> <p>After all simulations, it is possible to conclude that the reachability depends on the position of the task and the model of the robot. For example, UR5 could not reach task 3, while the UR10 had a high value of reachability for it due to the fact that the UR10 has a bigger workspace.</p> <p>-------------------------------------------------------------------------------------------------------------------------------------------------------</p> <p>Acknowledgement:</p> <p>Funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany's Excellence Strategy - EXC-2023 Internet of Production - 390621612.<br> </p>
The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3
<p>The North Pacific Eukaryotic Gene Catalog consolidates eukaryotic metatranscriptome data from three latitudinal transects of the North Pacific transition zone and one cruise in the subtropical gyre. Metatranscriptomes were gathered from latitudinally-resolved surface samples, and diel-resolved temporal studies, with samples taken in triplicate or duplicate and collected on 0.2-100 μm, 0.2-3 μm, and 3 μm-100 or 200 μm size fractions. These metatranscriptome data were <em>de novo</em> assembled into 175 independent assemblies, totalling 182 million clustered nucleotide contigs. Assemblies were annotated by taxonomy and function. This catalog provides assembled environmental contigs, their translated peptide sequences, and their taxonomic and functional annotations with the aim of facilitating continued discoveries about the molecular ecology of microbial eukaryotes in the North Pacific.<br><br>A full description of this data is published in Scientific Data, available here: <a href="https://www.nature.com/articles/s41597-024-04005-5" target="_blank" rel="noopener">The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations</a>. Please cite this publication if your research uses this data:<br><br>Groussman, R. D., Coesel, S. N., Durham, B. P., Schatz, M. J., & Armbrust, E. V. (2024). The North Pacific Eukaryotic Gene Catalog of metatranscriptome assemblies and annotations. <em>Scientific Data</em>, <em>11</em>(1), 1161.</p> <div> <p>This dataset repository is associated with a codebase and documentation repository:<br><a href="https://github.com/armbrustlab/NPac_euk_gene_catalog" target="_blank" rel="noopener">https://github.com/armbrustlab/NPac_euk_gene_catalog</a><br>Please see this code repository for additional data and project updates<br><br>Translated and processed protein sequences and their annotations are available in this repository: <br><a href="../doi/10.5281/zenodo.10472589">https://zenodo.org/doi/10.5281/zenodo.10472589</a><br><br>99% identity clustered nucleotide sequences and kallisto enumerations are available here:<br><a href="../doi/10.5281/zenodo.10570448">https://zenodo.org/doi/10.5281/zenodo.10570448</a></p> </div> <div> <p>File contents: this repository contains five .tar.gz compressed tarballs with raw de novo Trinity assemblies of poly-A selected metatranscriptomes from the Gradients 1 through 3 cruises, and a plain-text file with the custom spike-in mRNA standards (CustomStandardSequences.txt)</p> </div> <div> <p><strong><br>Gradients1.KOK1606.PA.assemblies.tar.gz</strong><br>- Link to <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G1PA" target="_blank" rel="noopener">G1PA project github page</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/KOK1606" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KOK1606</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G1PA.process_short_reads.sh" target="_blank" rel="noopener">G1PA.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G1PA.trinity_assemblies.sh" target="_blank" rel="noopener">G1PA.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>Gradients2.MGL1704.PA.assemblies.tar.gz</strong><br>- Link to <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G2PA" target="_blank" rel="noopener">G2PA project github page</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/MGL1704" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/MGL1704</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G2PA.process_short_reads.sh" target="_blank" rel="noopener">G2PA.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G2PA.trinity_assemblies.sh" target="_blank" rel="noopener">G2PA.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>Gradients3.KM1906.PA.assemblies.tar.gz</strong><br>- Link go <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G3PA" target="_blank" rel="noopener">G3PA project github page</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/KM1906" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1906</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_UW.process_short_reads.sh" target="_blank" rel="noopener">G3PA_UW.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_UW.trinity_assemblies.sh" target="_blank" rel="noopener">G3PA_UW.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>G3_diel.KM1906.PA.assemblies.tar.gz</strong><br>- Link go <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/tree/main/projects/G3PA" target="_blank" rel="noopener">G3PA project github page</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/KM1906" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1906</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_diel.process_short_reads.sh" target="_blank" rel="noopener">G3PA_diel.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/G3PA_diel.trinity_assemblies.sh" target="_blank" rel="noopener">G3PA_diel.trinity_assemblies.sh</a></p> </div> <div> <p><strong><br>CustomStandardSequences.txt<br></strong>- Plain-text FASTA file with the spike-in standards used during mRNA extraction and sequencing prep<br>- Link to publication of spike-in standards methods: <a href="https://www.nature.com/articles/s41564-019-0507-5" target="_blank" rel="noopener">https://www.nature.com/articles/s41564-019-0507-5</a></p> </div> <div> <p>The 2015 SCOPE Diel metatranscriptome raw assemblies have been released in a previous Zenodo repository, and are not included again in this deposition. We provide the links to the Diel1 resources here:<br>- Diel1 raw metatranscriptome assembly Zenodo repository: <a href="../records/5009803" target="_blank" rel="noopener">https://zenodo.org/records/5009803</a><br>- Dataset DOI: <a href="https://doi.org/10.5281/zenodo.5009803" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.5009803</a><br>- Associated publication: <a href="https://www.frontiersin.org/articles/10.3389/fmicb.2021.682651/full" target="_blank" rel="noopener">https://www.frontiersin.org/articles/10.3389/fmicb.2021.682651/full</a><br>- Codebase: <a href="https://github.com/armbrustlab/diel_eukaryotes" target="_blank" rel="noopener">https://github.com/armbrustlab/diel_eukaryotes</a><br>- Simons CMAP cruise page and datasets: <a href="https://simonscmap.com/catalog/cruises/KM1513" target="_blank" rel="noopener">https://simonscmap.com/catalog/cruises/KM1513</a><br>- Short read processing code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/D1PA.process_short_reads.sh" target="_blank" rel="noopener">D1PA.process_short_reads.sh</a><br>- Trinity assembly code: <a href="https://github.com/armbrustlab/NPac_euk_gene_catalog/blob/main/scripts/D1PA.trinity_assemblies.sh" target="_blank" rel="noopener">D1PA.trinity_assemblies.sh</a></p> </div> <p><br><br></p>
Predicted genes from the Amblyomma americanum draft genome assembly
<p>Data for pub "Predicted genes from the <em>Amblyomma americanum </em>draft genome assembly."</p> <ul> <li>Amblyomma_americanum_filtered_assembly.fasta: Decontaminated A. americanum genome with bacterial contigs removed</li> <li>Amblyomma_americanum_bacterial_contigs_info.tsv: Information about contigs classified as bacteria that were removed</li> <li>Amblyomma_americanum_annotation_data.tar.gz: Directory of annotation data produced by EvidenceModeler as part of the nf-core/genomeannotator workflow. Includes files in FASTA format (predicted genes and proteins), set of proteins clustered at 99% identity in FASTA format, and annotations in both GFF3 and GTF formats. GTF file produced from the GFF3 file with AGAT.</li> <li>Amblyomma_americanum_transcriptome_assembly_data.tar.gz: Directory of data generated for the transcriptome assembly that was used for gene prediction</li> </ul>
Penicillium fuscoglaucum Pf_T2 Genome Assembly and Annotation
<p>During routine culturing on selective media in the lab, we obtained an isolate of P. fuscoglaucum Pf_T2 and sequenced its genome. The Pf_T2 genome is far superior to available genomic resources for the species. Our assembly exhibits a length of 35.1 Mb, a BUSCO score of 97.9% complete, and consists of 5 scaffolds/contigs representing the four expected chromosomes. It was determined that the Pf_T2 genome was colinear with a type specimen P. fuscoglaucum, and contained a lineage specific, intact cylcopaizonic acid (CPA) gene cluster.</p>
Metagenome-assembled genomes from Stordalen Mire, Sweden (2019) (MAGs from long-read, short-read, & hybrid assemblies)
<p>METHODS:</p> <p>Soil samples (6 total) were collected at the Stordalen Mire site in 2019 from two depths (1-5 & 20-24 cm below ground) across three habitats (Palsa, Bog, and Fen). DNA was extracted based on the protocol described by <a href="http://dx.doi.org/10.17504/protocols.io.yxmvm244bg3p/v1">Li et al. (2024)</a>. For short reads, libraries were prepared at the Joint Genome Institute (JGI) with the KAPA Hyperprep kit, and sequenced with Illumina NovaSeq 6000. For long reads, libraries were prepared with the SMRTbell Express Template Prep Kit 2.0 (PacBio), then sequenced using PacBio Sequel IIe at JGI. PacBio data was processed at JGI to form filtered CCS (Circular Consensus Sequencing) reads. </p> <p>Assemblies were generated with short-only, long-only, and hybrid read sources: <strong>Short-only</strong> was assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">metaSPAdes</a> (v3.15.4) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Long-only</strong> was assembled with <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768) using <a href="https://zenodo.org/records/10806928">Aviary</a> (v0.5.3) with default parameters. <strong>Hybrid</strong> assembly was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with default parameters. This involved a step-down procedure with long-read assembly through <a href="https://www.nature.com/articles/s41592-020-00971-x">metaFlye</a> (v2.9-b1768), followed by short-read polishing by <a href="https://genome.cshlp.org/content/27/5/737">Racon</a> (v1.4.3), <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0112963">Pilon</a> (v1.24) and then Racon again. Next, reads that didn't map to high-quality metaFlye contigs were hybrid assembled with <a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5411777/">SPAdes (--meta option)</a> and binned out with <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5). For each bin, the reads within the bin were hybrid assembled using <a href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1005595">Unicycler</a> (v0.4.8). The high-coverage metaFlye contigs and Unicycler contigs were then combined to form the assembly fasta file. Genome recovery was performed using <a href="https://zenodo.org/records/10806928">Aviary</a> v0.5.3 with samples chosen for differential abundance binning by <a href="https://zenodo.org/records/10939393">Bin Chicken</a> (v0.4.2) using <a href="https://zenodo.org/records/7130825">SingleM metapackage S3.0.5</a>. This involved initial read mapping through <a href="https://zenodo.org/records/10531254">CoverM</a> (v0.6.1) using <a href="https://academic.oup.com/bioinformatics/article/34/18/3094/4994778">minimap2</a> (v2.18) and binning by <a href="https://peerj.com/articles/1165/">MetaBAT</a>, <a href="https://peerj.com/articles/7359/">MetaBAT2</a> (v2.1.5), <a href="https://www.nature.com/articles/s41587-020-00777-4">VAMB</a> (v3.0.2), <a href="http://doi.org/10.1038/s41467-022-29843-y">SemiBin</a> (v1.3.1), <a href="https://zenodo.org/records/10460259">Rosella</a> (v0.4.2), <a href="https://www.nature.com/articles/nmeth.3103">CONCOCT</a> (v1.1.0) and <a href="https://academic.oup.com/bioinformatics/article/32/4/605/1744462">MaxBin2</a> (v2.2.7). Genomes were analyzed using <a href="https://www.nature.com/articles/s41592-023-01940-w">CheckM2</a> (v1.0.2) and clustered at 95% ANI using <a href="https://zenodo.org/records/10526086">Galah</a> (v0.4.0).</p> <p> </p> <p>FILES:</p> <ul> <li><strong>EMERGE_MAGs_2019_long-short-hybrid.tar.gz</strong> - Archive containing the MAG files (.fna).</li> <li><strong>metadata_MAGs_2019_EMERGE.tsv</strong> - Table containing source sample names and accessions, GTDB classifications, CheckM2 quality information, NCBI GenomeBatch- and MIMAG(6.0)-formatted attributes, and other metadata for the MAGs.</li> </ul> <p> </p> <p>FUNDING:</p> <p>This research is a contribution of the EMERGE Biology Integration Institute (<a href="https://emerge-bii.github.io/">https://emerge-bii.github.io/</a>), funded by the National Science Foundation, Biology Integration Institutes Program, Award # 2022070.</p> <p>This study was also funded by the Genomic Science Program of the United States Department of Energy Office of Biological and Environmental Research, grant #s DE-SC0004632. DE-SC0010580. and DE-SC0016440.</p> <p>We thank the Swedish Polar Research Secretariat and SITES for the support of the work done at the Abisko Scientific Research Station. SITES is supported by the Swedish Research Council's grant 4.3-2021-00164.</p> <p>Data from the Joint Genome Institute (JGI) was collected under BER Support Science Proposal 503530 (DOI: <a href="https://doi.org/10.46936/10.25585/60001148">10.46936/10.25585/60001148</a>), conducted by the U.S. Department of Energy Joint Genome Institute (<a href="https://ror.org/04xm1d337">https://ror.org/04xm1d337</a>), a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231.</p>
PB260 chromosome assembly
<p>Chromosome level assembly of <em>Hevea brasiliensis</em> (Mull.Arg.) clone PB260. Obtained in the frame of RUBIS project rubis-project.org</p>
Catalog of stool metagenome-assembled genomes from patients with different cancer types
<p><strong>A non-redundant catalog of 3,816 genomes with at least 75% completeness and no more than 15% contamination assembled from metagenomes. Samples of 976 metagenomes were obtained from patients receiving immunotherapy for the treatment of different types of cancers.</strong></p>
Data from: Genome assembly of Danaus chrysippus and comparison with the Monarch Danaus plexippus
<p><strong>GENOME ASSEMBLY DATA</strong></p> <p><strong>Dchry2_unaltered.fa.gz</strong><br> Unaltered version of Dchry2 before manual edits</p> <p><strong>Dchry2.haplotigs.fasta.gz</strong><br> Haplotypic contigs removed in the Purge_haplotigs step</p> <p><strong>Dchry2_to_Dchry2.2_transfers.txt</strong><br> Edits of the Dchry2 assembly to produce Dchry2.2 (contig breaks, reverse complements and name changes). Columns are: new contig name, new contig start, new contig end, original contig name, original contig start, original contig end, orientation. New contig numbers indicate the chromosome they correspond to.</p> <p><strong>mxv1.200520.ragoo.rnm.fa.gz</strong><br> Fasta file from of MEX_DaPlex assembly with mxdp_ fasta headers</p> <p><br> <strong>GENE AND REPEAT ANNOTATION</strong></p> <p><strong>danaus_plex_mex_braker_a002.sequences.tidy.gff3.gz</strong><br> Sorted and tidied gff3 from MEX_DaPlex re-annotation </p> <p><strong>danaus_plexv4_braker_a006.sequences.tidy.gff3.gz</strong><br> Sorted and tidied gff3 from Dplex_v4 re-annotation </p> <p><strong>mxv1.200520.ragoo.rnm.wDpv3.gff3.gz</strong><br> gff3 file from Cei of MEX_DaPlex annotation</p> <p><strong>dplex2_uniprot-proteome_UP000596680_and_dplex_mex.fasta.gz</strong><br> Protein set used for annotation of all three assemblies</p> <p><strong>functional_annotation_output.tar.gz</strong><br> all output from pannzer2 regarding functinal annotation of the Dchry2.2 genome</p> <p><strong>Lepidoptera_and_danaus_chrysippus2.2.repeatmasker.gz</strong><br> Custom repeat library (combination of lepidoptera and specific dchry2.2 library)</p>
Actinidia chinensis Red5 genome assembly (version 2) and annotation files
<p>We present version 2 of the genome assembly for <em>Actinidia chinensis</em> var. <em>chinensis</em> genotype Red5. The Red5 genome was originally assembled using short read Illumina data (Pilkington et al, 2018; <a href="https://doi.org/10.1186/s12864-018-4656-3">https://doi.org/10.1186/s12864-018-4656-3</a>). In version 2 we employed Pacific BioSciences Sequel Single Molecule Real Time (SMRT) sequencing technology in place of Illumina paired end read sequencing for the main assembly but leveraged that short read data (Pilkington et al, 2018) for post assembly base correction of long read assembly contigs. Additionally the Illumina long insert libraries from Pilkington et al (2018) were used for post assembly scaffolding of contigs. Scaffold assignment to linkage groups leveraged the genetic map described in Pilkington et al (2018) as well as consensus evidence from DNA synteny comparisons to existing whole genome sequences from <em>Actinidia</em>.</p> <p>To meet the file size restrictions some dataset components have been split into multiple parts.</p> <p><strong>Assembly</strong></p> <p>The assembly work flow used the FALCON/FALCON-unzip assembly suite is described in Red5_version_2_genome_assembly.md. The assembly yielded both primary and haplotig contig data sets, the metrics for which are documented in this file. The CDS and predicted peptide fasta and GFF3 gene annotation for the primary and haplotig sets are provided in separate files.</p> <p><strong>File Descriptions</strong></p> <ul> <li>Files named chr1.fasta to chr29.fasta represent the primary assembly linkage group level assembly units</li> <li>Files named haplotig_part_1.fasta to haplotig_part_10.fasta represent the haplotig contig sets split into 10 parts to meet upload file size restrictions</li> <li>Files named primary_assembly.primary.gff3 and haplotig.gff3 contain the gene model annotations for the primary and haplotig assembly datasets respectively</li> <li>primary_assembly.cds.fasta and primary_assembly.pep.fasta contain the CDS and peptide sequences for the annotations on the primary contigs</li> <li>haplotig.cds.fasta and haplotig.pep.fasta contain the CDS and peptide sequences for the annotations on the haplotig contigs</li> <li>haplotigs.placements.tsv and haplotigs.reassignments.tsv describe the placement of haplotigs relative to the primary contigs as derived from purge_haplotigs</li> <li>The file Red5_version_2_genome_assembly.md describes the assembly work flow and code steps used as well as assembly metrics</li> <li>Files HYV3_1.v.R5V2_1.png to HYV3_29.v.R5V2_29.png depict Circos plots of DNA:DNA synteny based on 1coords alignment filter of nucmer alignments using dnadiff</li> </ul> <p>See Red5_version_2_genome_assembly.md for description of assembly methods and assembly metrics.</p> <p><strong>Funding</strong></p> <p>This work was funded by Kiwifruit Royalty Investment Program by The New Zealand Institute for Plant & Food Research Ltd. with support from Zespri, and the CORE grant Endeavour Smart Idea Fund (UOOX1801) from the New Zealand Ministry of Business, Innovation and Employment (MBIE). The funding bodies had no role in the design of the study, the collection, analysis, or interpretation of data or writing this manuscript.</p>
Annotation of the the assembled genome of Fusarium oxysporum f. sp. albedinis strain 133, the causal agent of date palm dieback.
<p>Annotation of the the assembled genome of <em>Fusarium oxysporum f. sp. albedinis</em> strain 133 (Khayi et al., 2020). Gene prediction and annotation were carried out using funnotate pipeline v1.8.1 (Stajich, 2020), which includes masking, ab initio gene-prediction training, using Augustus and Genmark, with the EST dataset reported to the Ganoderma mycocosm repository, gene prediction, and the assignment of functional annotation to protein-coding gene models.</p>
Dataset for "CVD growth of self-assembled 2D and 1D WS2 nanomaterials for the ultrasensitive detection of NO2"
<p>This file contains the raw data used in the paper entitled CVD growth of self-assembled 2D and 1D WS2 nanomaterials for the ultrasensitive detection of NO2 published in Sensors and Actuators: B. Chemical 326 (2021) 128813</p> <p>DOI: <a href="https://doi.org/10.1016/j.snb.2020.128813">10.1016/j.snb.2020.128813</a></p>
Gila monster (Heloderma suspectum) genome assembly and annotation
<p><em>De novo</em> genome assembly and annotation of a male Gila monster (<em>Heloderma suspectum</em>). We annotated the genome using the Comparative Annotation Toolkit (CAT), and we have also included GFF3 files of the consensus gene set, output for each taxon included in this process.</p>
Myxococcus xanthus DZ2 Genome Assembly
<p><strong>We report a high-quality assembly and annotation of the <em>Myxococcus xanthus </em>strain DZ2 (CP080538) genome, using a combination of short Reads generated by the DNBSEQ™ (BGI Genomics), and Long High-Fidelity (HiFi) Reads generated by Pacific Biosciences (PacBio) Technologies.</strong></p>
GenoNet scores for human genome assembly GRCh38
<p>Predicting the functional consequences of genetic variants in non-coding regions is a challenging problem. We propose here a semi-supervised approach, GenoNet, to jointly utilize experimentally confirmed regulatory variants (labeled variants), millions of unlabeled variants genome-wide, and more than a thousand cell/tissue type-specific epigenetic annotations to predict functional consequences of non-coding variants.</p> <p><strong>Format</strong></p> <p>The GenoNet scores are stored in the tab-delimited text files. </p> <p>Each row represents a genomic region with 131 columns. Please find the header line in "genonet.header.txt". </p> <p>The first four columns are chromosome, start coordinate, end coordinate, and a region ID named by positions. Please note that the coordinates are counted in the 0-based UCSC Genome Browser BED format. For example, the following region with a start position 10000 and an end position 10025 includes 25 base pairs within chr1:10001-10025.</p> <p>chr1 10000 10025 chr1_10001_10025</p> <p>Columns 5-131 are the predicted tissue-specific functional effects (GenoNet scores) for the 127 Roadmap tissues. Each column is named by the corresponding epigenome ID. This <a href="https://docs.google.com/spreadsheet/ccc?key=0Am6FxqAtrFDwdHU1UC13ZUxKYy1XVEJPUzV6MEtQOXc&usp=sharing">online spreadsheet</a> includes the information about the 127 Roadmap tissues in detail.</p> <p><strong>Reference</strong><br> Zihuai He, Linxi Liu, Kai Wang, Iuliana Ionita-Laza. A semi-supervised approach for predicting cell type/tissue specific functional consequences of non-coding variation using massively parallel reporter assays. Nature Communications, 2018.</p> <p><strong>Release</strong></p> <p>GRCh37 <a href="https://zenodo.org/record/3336209">https://zenodo.org/record/3336209</a></p> <p>GRCh38 liftover <a href="https://zenodo.org/record/6484230">https://zenodo.org/record/6484230</a></p>
The first annotated genome assembly of Macrophomina tecta associated with charcoal rot of sorghum
<p>Raw reads of Macrophomina tecta were obtained from Nanopore, Illumina, and NextSeq (RNA). Files with information about the genome annotation, functional prediction, repeats, effectors and orthologous genes are included. </p>
Genome assemblies of Xanthomonas oryzae pv. oryzae (PXO35, FXO38, Huang604) and Xanthomonas oryzae pv. oryzicola (BAI35, MAI23)
<p>Genome assemblies of <em>Xanthomonas oryzae</em> pv. <em>oryzae</em> (<em>Xoo</em>) and <em>Xanthomonas oryzae</em> pv. <em>oryzicola</em> (<em>Xoc</em>). Genome assemblies of the <em>Xoo</em> strains PXO35, FXO38, Huang604 and the <em>Xoc</em> strains BAI35, MAI23, have been generated with Flye based on ONT reads. For each of these strains, we corrected the sequences encoding for transcription activator-like effectors (TALEs) with our TALE-correction pipeline (https://github.com/Jstacs/Jstacs/tree/master/projects/talecorrect). For Xoo PXO35, we additionally provide assemblies based on reads obtained from different sequencing methods (Illumina, PacBio, ONT) generated by a collection of (hybrid) assembly strategies and different polishing approaches applied to combinations of these.</p>
Supporting data: HiFi chromosome-scale diploid assemblies of the grape rootstocks 110R, Kober 5BB, and 101-14 Mgt
<p>Repository for supporting data to the paper: HiFi chromosome-scale diploid assemblies of the grape rootstocks 110R, Kober 5BB, and 101-14 Mgt</p>
Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction
<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the <em>P. teres </em>f.<em> maculata </em>isolate FGOB10Ptm-1. </p>
Four Reference Quality Genome Assemblies of Pyrenophora teres f. maculata: A Resource for Studying the Barley Spot Form Net Blotch Interaction
<p>Updated draft genome assembly (FASTA) and annotation (GFF) for the <em>P. teres </em>f.<em> maculata </em>isolate P-A14. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.