Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,439
datasets available to search
ShareScore release 0.9.0
Dataset results
2,439 results for “Assembly”
Formation and cycling data for Na-ion batteries from high-throughput synthesis, coating, and assembly
<p>Formation and cycling data from a combinatorial/high-throughput upscaling process for the production and characterization of sodium-ion batteries. The process involves batch synthesis, screen printing of electrodes, robotic cell assembly, and battery cycling. The goal of this study was to test how fast a new chemistry (to the group) could be introduced into the workflow and if we are able to enhance efficiency, accuracy, and reproducibility. The cathode material, Na0.9[Cu0.22Fe0.30Mn0.48]O2, was synthesized through a solid-state reaction (Na2CO3 (purity 99.5 %), CuO (purity 99.7 %), Fe2O3 (purity 99.9 %) and Mn2O3 (purity 98 %) at 850°C for 15h) in a pressed pellet (10 MPa) that was ground up again to make a slurry. The electrodes were prepared using screen printing, which offers simplicity, low cost, and quick coating of large areas in a reproducible manner. The binder was sodium carboxymethyl cellulose to make the electrodes water processable in air. The assembled batteries utilized the synthesized cathode material and hard carbon as the anode, with a glass fiber separator and a 1M NaPF6 EC:EMC 3:7 with 2 wt% FEC electrolyte.</p>
Genome Assembly of Fusarium avenaceum
<p>Here, we present a complete genome assembly for <em>Fusarium avenaceum</em> and associated annotation using long-read sequencing generated from the Oxford Nanopore Technologies (ONT;London, UK) platform for both DNA and RNA obtained from fruit-sampled cultures. </p>
Cortical cell assemblies and their underlying connectivity: an in silico study
<p>Dataset linked to the article with the same title</p> <ul> <li>simulation_config.zip: contains SONATA config files needed to re-run an exemplary simulation (after downloading the circuit from <a href="https://zenodo.org/record/7930275">10.5281/zenodo.7930275</a>). In order to run it, paths in circuit_config.json, and simulation_config.json have to be updated!</li> <li>assemblies.h5 is a dataset produced (and can be easily opened) by: <a href="https://zenodo.org/record/8112725">assemblyfire</a> (see GitHub README for more documentation) and serves as a basis for the manuscript. As the assemblies are the results of an unsupervised clustering (of high activity time bins) the resulting labels are not necessary meaningful. In the manuscript we have ordered the assemblies (from early to late responding ones, and from pattern A to J responsive ones) but the HDF5 file still stores the original labels. The mapping from the "random" labels to the ones presented in our article is stored in the config files on GitHub.</li> </ul> <p><em>The development of this dataset was supported by funding to the Blue Brain Project, a research center of the École polytechnique fédérale de Lausanne (EPFL), from the Swiss government’s ETH Board of the Swiss Federal Institutes of Technology.</em></p>
ScRAPv20230731: Telomere-to-telomere assemblies of 142 strains characterize the genome structural landscape in Saccharomyces cerevisiae
<p><strong><em>Saccharomyces cerevisiae </em>Reference Assembly Panel (ScRAP) v20230731 </strong>></p> <p>The haplotype-resolved and/or collapsed T2T genome assemblies for 142 <em>S. cerevisiae</em> strains isolated from diverse geographical and ecological niches.</p>
New Soil Metagenome-Assembled Genomes Catalogue Boosts Genetic Resources
<p><strong>Soil harbors a vast expanse of unidentified microbes, termed as microbial dark matter, presenting an untapped reservoir of microbial biodiversity and genetic resources, but has yet to be fully explored. In this study, we conducted the first large-scale excavation of soil microbial dark matter by reconstructing 40,039 metagenome-assembled genome bins (the SMAG catalog) from 3,304 soil metagenomes. We identified 16,530 of 21,077 species-level genome bins (SGBs) as unknown SGBs (uSGBs), which greatly expand archaeal and bacterial diversity across the tree of life. We also illustrate the pivotal role of uSGBs in augmenting soil microbiome's functional landscape and intra-species genome diversity, providing large proportions of the 43,169 biosynthetic gene clusters and 8,545 CRISPR-Cas genes. Additionally, we determined that uSGBs contributed 84.6% of novel viral-host associations identified from the SMAG catalog. Our results propose the SMAG catalog, a novel and expansive genomic resource that brings the soil microbial biodiversity and novel genetic resources to light.</strong></p>
Draft de novo genome assemblies of a male and female Amphibolurus muricatus (jacky dragon)
<p>Four de novo nuclear genome assemblies of <em>Amphibolurus muricatus</em></p> <p><strong>Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> • AmpMurF_1.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_1.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 1.1: Further scaffolding of assembly 1.0 using RNA-seq data</strong><br> • AmpMurF_1.1.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_1.1.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 2.0: Further scaffolding of assembly 1.0 using SLR-superscaffolder</strong><br> • AmpMurF_2.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_2.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing assembly</strong><br> • AmpMurF_3.0.fa.tar.gz (female <em>A. muricatus</em>)<br> • AmpMurM_3.0.fa.tar.gz (male <em>A. muricatus</em>)</p> <p><strong>Methods<br> Assembly 1.0: A 10x Genomics linked-read sequencing assembly</strong><br> Male and female <em>A. muricatus</em> genome sequencing libraries were constructed on the Chromium system (10x Genomics, Pleasanton, CA, USA) by the Ramaciotti Centre for Genomics (Sydney, Australia). The Chromium instrument enables unique barcoding of long stretches of DNA on gel beads. The barcodes allow later reconstruction of long DNA fragments from a series of short DNA fragments with the same barcode (i.e., linked-reads). After barcoding, DNA was sheared into smaller fragments and sequenced on the NovaSeq 6000 platform (Illumina, CA, USA) to generate 151 bp paired-end (PE) reads. A total of 904.9 M raw 10x Genomics Chromium linked-reads were generated. Raw 10x data were assembled with Supernova v2.1.1 (Weisenfeld et al., 2017) and a FASTA file was generated using the ‘pseudohap style’ option in Supernova mkoutput. All female (~450 M) and male (~550 M) read pairs were utilised (female sequencing depth ca 50.3×; male, ca 47.8×). The resulting assemblies was further scaffolded with ARKS v1.0.3 (Coombe et al., 2018), reusing the 10x reads, and the companion LINKS program (v1.8.7) (Warren et al., 2015). ARKS employs a <em>k</em>-mer approach to map linked barcodes to the contigs in the initial Supernova assembly to generate a scaffold graph with estimated distances for LINKS input. These assemblies were denoted AmpMurF_1.0 (female) and AmpMurM_1.0 (male). We used GapCloser v1.12 (part of SOAPdenovo2) (Luo et al., 2012) to fill gaps in the assembly. GapCloser was run using the parameter -l 150) and clean 10x Genomics reads PE reads. </p> <p><strong>Assembly 1.1: Further scaffolding using RNA-seq data</strong><br> We attempted to improve the v1.0 genome assemblies’ contiguity using RNA-sequencing reads. RNA-seq reads (from brain, ovary, and testis; see below) were filtered (i.e., cleaned) to remove adapters and low-quality reads using Flexbar v3.4.0 and used to further re-scaffold the v1.0 assemblies (FASTA files before gapclosing) with P_RNA_scaffolder (Zhu et al., 2018). The default Flexbar settings discards all reads with any uncalled bases. A final round of scaffolding was performed on the resulting assemblies using L_RNA_scaffolder (Xue et al., 2013). These assemblies were denoted AmpMurF_1.1 (female) and AmpMurM_1.1 (male). As before, GapCloser and clean 10x Genomics reads were used to fill gaps. </p> <p><strong>Assembly 2.0: Further scaffolding using SLR-superscaffolder</strong><br> As an alternative approach, we attempted to improve the v1.0 genome assemblies’ contiguity using SLR-superscaffolder (Guo et al., 2021). Briefly, SLR-superscaffolder employs single tube long fragment read (stLFR) sequencing (Wang et al., 2019) reads (see section below) to generate hybrid genome assemblies. The software was run with default parameters except for PE_SEED_MIN=300 (minimum contig size to fill; default 1000). These assemblies were denoted AmpMurF_2.0 (female) and AmpMurM_2.0 (male). GapCloser and clean stLFR reads (with the barcode removed using https://github.com/BGI-Qingdao/stLFR_barcode_split) were used to fill gaps. </p> <p><strong>Assembly 3.0: An stLFR linked-read sequencing and supernova assembly</strong><br> We also generated independent assemblies for the individuals sequenced on the 10x Genomics Chromium system using single tube long fragment read (stLFR) sequencing (Wang et al., 2019). BGI (Brisbane, Australia) generated ~100×-coverage 100-bp paired-end reads (plus a 42-bp stLFR barcode on the right/_2 read). Low-quality reads, PCR duplicates, and adaptors were removed using SOAPnuke v1.5 (Chen et al. 2018). The stLFRdenovo pipeline (<a href="https://github.com/BGI-biotools/stLFRdenovo">https://github.com/BGI-biotools/stLFRdenovo</a>), which is based on Supernova and customized for stLFR data, was used to generate a <em>de novo</em> genome assembly. The stLFRdenovo tool ‘FillGaps’ was used to fill gaps.</p> <p><strong>References</strong><br> Chen, Y., Chen, Y., Shi, C., Huang, Z., Zhang, Y., Li, S., Li, Y., Ye, J., Yu, C., Li, Z., et al. (2018). SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience 7, 1-6.<br> Coombe, L., Zhang, J., Vandervalk, B.P., Chu, J., Jackman, S.D., Birol, I., and Warren, R.L. (2018). ARKS: chromosome-scale scaffolding of human genome drafts with linked read kmers. BMC Bioinformatics 19, 234.<br> Guo, L., Xu, M., Wang, W., Gu, S., Zhao, X., Chen, F., Wang, O., Xu, X., Seim, I., Fan, G., et al. (2021). SLR-superscaffolder: a de novo scaffolding tool for synthetic long reads using a top-to-bottom scheme. BMC Bioinformatics 22, 158.<br> Luo, R., Liu, B., Xie, Y., Li, Z., Huang, W., Yuan, J., He, G., Chen, Y., Pan, Q., Liu, Y., et al. (2012). SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 1, 18.<br> Wang, O., Chin, R., Cheng, X., Wu, M.K.Y., Mao, Q., Tang, J., Sun, Y., Anderson, E., Lam, H.K., Chen, D., et al. (2019). Efficient and unique cobarcoding of second-generation sequencing reads from long DNA molecules enabling cost-effective and accurate sequencing, haplotyping, and de novo assembly. Genome Res 29, 798-808.<br> Warren, R.L., Yang, C., Vandervalk, B.P., Behsaz, B., Lagman, A., Jones, S.J., and Birol, I. (2015). LINKS: Scalable, alignment-free scaffolding of draft genomes with long reads. Gigascience 4, 35.<br> Weisenfeld, N.I., Kumar, V., Shah, P., Church, D.M., and Jaffe, D.B. (2017). Direct determination of diploid genome sequences. Genome Res 27, 757-767.<br> Xue, W., Li, J.T., Zhu, Y.P., Hou, G.Y., Kong, X.F., Kuang, Y.Y., and Sun, X.W. (2013). L_RNA_scaffolder: scaffolding genomes with transcripts. BMC Genomics 14, 604.<br> Zhu, B.H., Xiao, J., Xue, W., Xu, G.C., Sun, M.Y., and Li, J.T. (2018). P_RNA_scaffolder: a fast and accurate genome scaffolder using paired-end RNA-sequencing reads. BMC Genomics 19, 175.</p>
Dataset Changes in structure and assembly of a species-rich soil natural community with contrasting nutrient availability upon establishment of a plant-beneficial Pseudomonas in the wheat rhizosphere
<p>This dataset is related to the paper "<strong>Changes in structure and assembly of a species-rich soil natural community with contrasting nutrient availability upon establishment of a plant-beneficial <em>Pseudomonas </em>in the wheat rhizosphere</strong>" (Garrido-Sanz et al., 2023, doi: 10.1186/s40168-023-01660-5) and contains the data obtained from bacterial competition asays and plant-growth measurements.</p> <p>Sequencing data used in this study has been deposited in the NCBI Sequence Read Archive (RSA) under the BioProject accession number <a href="https://www.ncbi.nlm.nih.gov/bioproject/PRJNA948847">PRJNA948847</a>.</p> <p>The R script used to analyze the data generated in the paper is available at <a href="https://github.com/dgarrs/Pprotegens_proliferation_NatComs">GitHub </a>and <a href="https://doi.org/10.5281/zenodo.8322086">Zenodo</a>.</p>
Bibliographic data for the systematic review on tilting table tests of masonry assemblies
<p>This database contains all the bibliographic information about the records found after applying the Search Strategy used for the Systematic Review on Tilting Table Tests of Masonry Assemblies. The search was conducted in Scopus, Web of Science, IEEE Explore, Engineering Village, and Wiley Online Library databases. It was performed on 13/09/2023. The bibliographic data of the records found in the different databases is presented in .ris, .bib, and .csv format.</p>
Genome assemblies of four MDR B. fragilis isolates using PacBio data - supporting the PhD Thesis
<p>Unicycler and Canu assemblies using PacBio data of four MDR B. fragilis isolates.</p> <p>Data supporting the PhD Thesis <em>Epidemiology and Genomics of antimicrobial resistance in the Bacteroides fragilis group </em>by Thomas Vognbjerg Sydenham, The research unit of Clinical Microbiology, Department of Clinical Research, Faculty of Health Sciences, University of Southern Denmark September 2019.</p> <p> </p>
Metagenomic assembly and bin3C clustering result for a healthy human faecal microbiome transplant donor
<p>Metagenomic WGS assembly and Hi-C deconvolution of a healthy human faecal microbiome transplant donor.</p> <p>Metagenomic assembly was produced using Spades (v3.13.1).</p> <p>Extracted MAGs were produced using bin3C (v0.3.3) and QC'd using CheckM (v1.0.18).</p>
Rehti रहती (सिहोर जिला Madhya Pradesh). Hero stone used as part of re-assembled plinth.
<p>Rehti रहती (सिहोर जिला Madhya Pradesh). Hero stone used as part of re-assembled plinth.</p>
Rehti रहती (सिहोर जिला Madhya Pradesh). Temple remains re-assembled as plinth.
<p>Rehti रहती (सिहोर जिला Madhya Pradesh). Temple remains re-assembled as plinth.</p>
Rehti रहती (सिहोर जिला Madhya Pradesh). Temple remains re-assembled as plinth.
<p>Rehti रहती (सिहोर जिला Madhya Pradesh). Temple remains re-assembled as plinth.</p>
Rehti रहती (सिहोर जिला Madhya Pradesh). Temple remains re-assembled as plinth.
<p>Rehti रहती (सिहोर जिला Madhya Pradesh). Temple remains re-assembled as plinth.</p>
CAT 4.6 taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly
<p>Taxonomic binning of the gold standard pooled assembly<br> <strong>Software: </strong>CAT<br> <strong>SoftwareVersion: </strong>4.6<br> <strong>DataURL: </strong> https://data.cami-challenge.org/participate<br> <strong>SoftwareURL:</strong> https://github.com/dutilh/CAT<br> <strong>ReferenceDatabase:</strong> prebuilt 2018-12-12<br> <strong>Taxonomy:</strong> NCBI 2018-12-12<br> <strong>ShortReadsUsed:</strong> False<br> <strong>LongReadsUsed:</strong> False<br> <strong>CommandUsed:</strong> CAT contigs -c anonymous_gsa_pooled.fasta -d CAT_prepare_20181212/2018-12-12_CAT_database/ -t CAT_prepare_20181212/2018-12-12_taxonomy/ --tmpdir tmp --nproc 16</p>
PhyloPythiaS+ 1.4 taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly
<p>Taxonomic binning of the gold standard pooled assembly<br> <strong>Software: </strong>PhyloPythiaS+<br> <strong>SoftwareVersion: </strong>1.4<br> <strong>DataURL: </strong> https://data.cami-challenge.org/participate<br> <strong>SoftwareURL:</strong> https://github.com/algbioi/ppsp<br> <strong>DockerImage:</strong> cami/ppsp:1.4<br> <strong>IsBiobox:</strong> False<br> <strong>ReferenceDatabase:</strong> RefSeq 93, SILVA 132<br> <strong>Taxonomy:</strong> NCBI 2018-02-26<br> <strong>ShortReadsUsed:</strong> False<br> <strong>LongReadsUsed:</strong> False<br> <strong>CommandUsed:</strong> run_ppsp.py --pipelineDir ppsp_pipepline --inputFastaFile anonymous_gsa_pooled.fasta --databaseFile ncbi_taxonomy --refSeq refseq93 --s16Database SILVA_132 --mgDatabase reference_NCBI201502/mg5</p>
MetaBAT 2.12.1 genome binning of the CAMI 2 Mouse Gut Toy data set, samples 0-63, gold standard pooled assembly
Genome binning of the gold standard pooled assembly <br><strong>Software: </strong>MetaBAT<br><strong>SoftwareVersion: </strong>2.12.1<br><strong>DataURL: </strong> https://data.cami-challenge.org/participate<br><strong>SoftwareURL:</strong> https://bitbucket.org/berkeleylab/metabat<br><strong>ShortReadsUsed:</strong> True<br><strong>LongReadsUsed:</strong> False<br><strong>CommandUsed:</strong> bowtie2-build anonymous_gsa_pooled.fasta anonymous_gsa_pooled.fasta<br>for i in {0..63}; do bowtie2 -q --threads 30 --fr -x anonymous_gsa_pooled.fasta --interleaved sample_${i}/anonymous_reads.fq -S anonymous_reads_sample_${i}.sam ; done<br>for i in {0..63}; do samtools view -b sample_${i}.sam -o anonymous_reads_sample_${i}.bam & done<br>for i in {0..63}; do samtools sort anonymous_reads_sample_${i}.bam -o anonymous_reads_sample_${i}.sorted.bam ; done<br>for i in {0..63}; do samtools index anonymous_reads_sample_${i}.sorted.bam ; done<br>runMetaBat.sh -l anonymous_gsa_pooled.fasta anonymous_reads_sample_*.sorted.bam
Kraken 2.0.8 beta taxonomic binning of the CAMI 2 Mouse Gut Toy data set, gold standard pooled assembly
<p>Taxonomic binning of the gold standard pooled assembly<br> <strong>Software: </strong>Kraken<br> <strong>SoftwareVersion: </strong>2.0.8 beta<br> <strong>DataURL: </strong> https://data.cami-challenge.org/participate<br> <strong>SoftwareURL:</strong> https://ccb.jhu.edu/software/kraken2/<br> <strong>ReferenceDatabase:</strong> built 2019-05-22<br> <strong>Taxonomy:</strong> NCBI 2019-05-22<br> <strong>ShortReadsUsed:</strong> False<br> <strong>LongReadsUsed:</strong> False<br> <strong>CommandUsed:</strong> kraken2-build --standard --db kraken2db_std --use-ftp<br> kraken2 --db kraken2db_std --threads 16 --output 19122017_mousegut_scaffolds.kraken --report 19122017_mousegut_scaffolds.kreport anonymous_gsa_pooled.fasta<br> cat 19122017_mousegut_scaffolds | awk '{print $2 "\t" $3}' > 19122017_mousegut_scaffolds.cami</p>
GC-MS data set for Generation of a chromosome-scale genome assembly of the insect-repellant terpenoid-producing Lamiaceae species, Callicarpa americana
<p>RAW GC/MS data set for characterization of class II terpene synthases from <em>Callicarpa americana </em></p>
Dataset for "Whole-genome de novo assemblies reveal structural variations and organelle-to-nucleus DNA transfers in Asian and African rice""
<p>DXCWR_O.rufipogon_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. rufipogon</em> DXCWR.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>DXCWR_O.rufipogon_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly DXCWR_O.rufipogon_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. glaberrima</em> IRGC104165.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>IRGC104165_O.glaberrima_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly IRGC104165_O.glaberrima_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. barthii</em> W1411.</p> <p>W1411_O.barthii_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W1411_O.barthii_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W1411_O.barthii_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored.fa.gz</p> <p>--Scaffolded and anchored genome assembly for <em>O. nivara</em> W2014.</p> <p>W2014_O.nivara_scaffolded_anchored.gff.gz</p> <p>--Gene annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p> <p>W2014_O.nivara_scaffolded_anchored_repeatmasker.gff.gz</p> <p>--Repeat annotation for the genome assembly W2014_O.nivara_scaffolded_anchored.fa.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.