Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
11
datasets available to search
ShareScore release 0.9.0
Dataset results
11 results for “assembly graphs”
Supporting data for Novel functional sequences uncovered through a bovine multi-assembly graph
<p><strong>Description of the datasets</strong></p> <p>Data are organized as a folder and compressed with tar.gz.</p> <p>You need to unzip the folder using the command <em>tar -xz</em><em>v</em><em>f</em> data.tar.gz. Unzipping will output a folder named <em>data_tidy</em>, which is organized as follow:</p> <ul> <li>graph.gfa : Graph in GFA format constructed from 6 cattle assemblies</li> <li>nonref.fa : Non-reference sequences extracted from the graph</li> <li>nonref.fa.masked: Hard masked repetitive regions version of nonref.fa</li> <li>nonref_woflanking.fa: Nonref.fa without flanking sequences</li> <li>nonref_woflanking.fa.masked: Masked version of nonref_woflanking.fa</li> <li>augustus_predict.gtf: Annotated gene models of Augustus from non-ref sequences</li> <li>augustus_prot.fa: Protein fasta of the predicted gene models from Augustus</li> <li>breeds_assembled.gtf: Annotation of the StringTie assembled across-breed transcriptome</li> <li>breeds_expressed.tsv: Expression data of breeds_assembled.gtf</li> <li>de_assembled.gtf: Annotation of the StringTie assembled differentially-expressed transcriptome on non-ref sequences</li> <li>de_expression.tsv: Differential expression results from de_assembled.gtf</li> <li>variant_nonref.tsv: Variants called from non-ref sequences (-1, 0, 1, 2 indicates no call, hom ref, het, and hom alt respectively)</li> </ul>
Liftover of DGRP D.melanogaster genotypes to reference genome assembly v6.0, with QC graph
<p>Output are vcf and plink-format genotype files. Also provided are the run code in bash and R, the logs, summary statistics, and a graph showing how the positions of SNPs have changed.</p>
Supplementary files for the manuscript "Assembly of Long Error-Prone Reads Using Repeat Graphs"
<p>Supplementary files for the manuscript "Assembly of Long Error-Prone Reads Using Repeat Graphs"</p> <p> </p> <p>Contents<br> ---------</p> <p>* `human_assemblues` - Flye assemblies of the human ONT sequencing data + QUAST benchmarking<br> of Flye, Canu and MaSuRCA assemblies. Scripts for assembly graph analysis are also included.</p> <p>* `nctc_assemblis` - Flye assemblies of the NCTC 21 bacterial dataset.</p> <p>* `yeast_assemblies` - working directories Flye, Canu, Falcon, Hinge and Miniasm assemblies of <br> yeast PB and ONT datasets + final assemblies + quast report. Some large files <br> (such as read alignments) were deleted.</p> <p>* `worm_assemblies` - working directories Flye, Canu, Falcon, Hinge and Miniasm assemblies of <br> the c. elegans dataset + final assemblies + quast report. Some large files <br> (such as read alignments) were deleted. `tandem_misassemblies` directory contain<br> the detailed analysis of nine tandem misassemblies. We recommend "gepard" dot-plotter for visualization.</p> <p>* `metagenome_assemblies` - Flye and Canu assemblies of a PacBio mock metagenome dataset.<br> In addition to metagenome assemblies, each bacteria was reassembled separately to<br> estimate the rate of divergence between the target genomes and the available references.</p> <p>* `simulated_data` - two assemblies of the simulated data illustrating Figure 1 (from Appendix I),<br> as well as simulated unbridged repeats benchmark.</p> <p><br> Software versions and parameters<br> --------------------------------</p> <p>* Flye - 2.3.5 (commit 20afeda)<br> * Canu - 1.7.1 (commit dfa60b8)<br> * Falcon - 0.3.0 (FALCON-Integrate commit 7498ef9)<br> * HINGE - 0.5.0 (commit 79fdf66)<br> * Miniasm - 0.2-r168-dirty (commit 40ec280) / Minimap2 2.8-r711-dirty (commit 8fc5f8d)<br> * Quast - 5.0.0 (commit de6973bb)</p> <p>Flye and Canu were run with the default parameters. The config files / scripts for<br> Falcon, HINGE and Miniasm could be found in the 'asm_config' archive folder.</p> <p>The HUMAN (but not the HUMAN+) assembly was generated with the earlier <br> Flye version 2.3.2 (released on Feb 20 2018) to provide a fair comparison <br> with the Canu and MaSuRCA assemblies (which were not updated since the release of Flye 2.3.2).<br> We note that the HUMAN assembly using the latest Flye version 2.3.5 has <br> NGA50 = 7.3 Mb and improves over the Flye 2.3.2 assembly (NGA50 = 6.3Mb). <br> HUMAN+ was assembled using the latest Flye and Canu versions (as of September 2018).</p> <p>The code for unbridged repeat resolution is currently available <br> in a separate 'flye-trestle' branch (commit 6100d32)</p>
Sequence data from viral assembly graph analysis of the SERC virome
<p>Various sequence data and details that were used in publishing the manuscript describing assembly graph binning in the SERC viral metagenome. Including: assembled contigs, FASTG, predicted ORFs, two tables describing the assembled contigs and graphs.</p> <p>Also PolA and RNR sequences that were mined from the SERC assembly.</p>
SPAligner: alignment of long error-prone reads to assembly graphs
<p>This repository contains benchmarking datasets and scripts for the manuscript "SPAligner: alignment of long error-prone reads to assembly graphs". </p> <p>Graph representation of genome assemblies has been recently used in different applications — from gene finding to haplotype separation. While many of these applications are based on aligning DNA and protein sequences to assembly graphs, existing software tools for finding such alignments have important limitations. We present a novel SPAligner (Saint Petersburg Aligner) tool for aligning long reads to assembly graphs and demonstrate that it generates accurate alignments.</p> <p> </p> <p> </p>
metaFlye: scalable long-read metagenome assembly using repeat graphs
<p>Supplementary files for the manuscript titled "metaFlye: scalable long-read metagenome assembly using repeat graphs".</p> <p>The archive includes generated assemblies and the corresponding metaQUAST evaluations.</p>
SUAG Single-Use Assembly Graph-Structured Dataset
<p>The SUAG (Single-Use Assembly Graph) dataset is a comprehensive collection designed to advance research in the domain of single-use assembly systems and graph-based representations. This dataset encompasses a rich repository of structured information, bringing together detailed insights into pharmaceutical Single-Use Assemblies (SUAs) and their corresponding graph structures.</p> <p>Key Features:</p> <ul> <li><strong>Single-Use Assembly Data:</strong> Detailed records of diverse SUAs, including bioreactors, tubing sets, and related components, providing a holistic view of assembly configurations.</li> <li><strong>Graph Representations:</strong> Graph structures capture the interconnectivity and relationships among components within each SUA, offering a unique perspective on their functional and spatial arrangements.</li> <li><strong>Variety of Applications:</strong> The dataset caters to a wide range of applications, from biopharmaceutical production to system optimization, facilitating diverse research endeavors.</li> </ul> <p> </p> <p>Dataset Structure: </p> <ul> <li>The dataset containes three classes of SUA graphs, named: Bioreactor, Mixer, and conection.</li> <li>Classes containe 236,196 Graphs, 1,088 Graphs, 300,327 Graphs respectively. </li> <li>Items named sequentially in each class and saved in their row nature in npy format.</li> <li>Each graph file consists of a list of an edges list showing nodes' names and nodes' connectivity.</li> </ul> <p>graph structure example, file m_0.npy in Mixer class: </p> <p><code>[['M' 'L1S1'], ['L1S1' 'L1P1'], ['M' 'L2S1'], ['L2S1' 'L2A2'], ['M' 'L5A1'], ['M' 'L3T1']]</code></p> <p>The 'L' letter when followed by a number such as the case in L1, indicates the components level. components/nodes in different levels are not the same component/nodes. </p> <p>The node attribute/component is what follows for example in L1S1, the node/component is S. digit 1 indicates that the component S1 is not the same node as S2. </p> <p>We decided on this numerical way to reduce ambiguity when two nodes of the same components appear in the graph but they are different items in the pipeline. </p> <p> </p> <p>Component list:</p> <p> 'B' -> 'bioreactorBag'<br> 'H' -> 'hydrophobicFilter'<br> 'F' -> 'hydrophilicFilter'<br> 'Q' -> 'quickCoupler'<br> 'Y' -> 'yConnector'<br> 'X' -> 'xConnector'<br> 'P' -> 'plug'<br> 'A' -> 'asepticConnector'<br> 'D' -> 'asepticDisconnector'<br> 'S' -> 'straightFitting'<br> 'M' -> 'mixerBag'<br> 'R' -> 'couplerReducer'<br> 'T' -> 'triclampConnector'<br> 'N' -> 'tConnector'<br> 'G' -> 'twoDimensionalBag'<br> 'I' -> 'threeDimensionalBag'<br> 'E' -> 'triclampConnector'<br> 'O' -> 'sipConnector'<br> 'J' -> 'bottle'<br> 'C' -> 'mechanicDisconnector'<br> 'Z' -> 'lConnector'</p> <p> </p> <p> </p> <p>Researchers, practitioners, and educators can leverage the SUAG dataset to explore innovative solutions, develop algorithms for graph-based analyses, and enhance understanding of the evolving landscape of single-use assembly systems.</p>
SheepGut assembly graph used in the strainFlye paper
<p>SheepGut assembly graph, generated with metaFlye version 2.8 as described in the paper. This is a GFA file.</p> <p>We provide this graph here (in addition to the raw SheepGut sequencing data, which is located elsewhere) in order to simplify reproduction of the analyses shown in the report.</p>
Supporting data for the manuscript "metaFlye: scalable long-read metagenome assembly using repeat graphs"
<p>Genome assemblies, simulated datasets and evaluations described in the manuscript "metaFlye: scalable long-read metagenome assembly using repeat graphs".</p>
Assembly graphs and reference Verkko assembly for HG002 dataset
Open the record for dataset details and reuse information.
Data for "Capturing variation in metagenomic assembly graphs with MetaCortex".
<p>Data for paper "Capturing variation in metagenomic assembly graphs with MetaCortex". Includes all assemblies, simulated reads, and simulated genomes.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.