Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
373
datasets available to search
ShareScore release 0.7.1
Dataset results
373 results for “Nanopore”
Dataset for Multiplex-PCR detection and Nanopore-based genotyping of fish pathogens
<p>This is a revised zip file contains scripts, initial fastq files, assembled amplicon (public and from this study) as well as bioinformatics intermediate files used for this study.</p> <p>Changelog:</p> <p>1. Fixed a bug in the 02_consensus.sh to enable proper removal of amplicons with zero depth</p> <p>2. Added a script (06_unclassified_read.sh) to extract and annotate reads that previously could not align to the 4 reference gene segment. Now the previously unclassified reads will be re-align (raw fastq) back to the gene segments as well as an additional tilapia genome assembly to gauge amount of reads mapping to the host genome. Furthermore, any read that still fail to align with minimap2 was subsequently aligned using blastn (-word_size 15 -evalue 0.01) against the same sequences.</p> <p>File Structure and Descriptions</p> <p>├── 01_process.sh : primer trimming, length-based filtering, read alignment, alignment filtering (unique hit) and extraction of uniquely hit reads for consensus generation<br> ├── 02_consensus.sh : [need artic conda env] Generation of consensus based on uniquely-mapped reads and minimal read depth of 20x required to call a variant (or it will be masked)<br> ├── 03_cleanup.sh: General folder and intermediate file re-organization<br> ├── 04_filter.sh: [need quast conda env] statistic of consensus generated and filtering of consensus with one or more ambiguous base (N), not suitable for haplotype<br> ├── 05_cluster.sh: clustering of consensus based on 100% identity threshold to generate putative haplotype<br> ├── 06_unclassified_read.sh: Extraction and annotation of unclassified reads using lenient criteria and with host reference genome as added reference<br> ├── Amplicon_FastQ folder: uniquely mapped fastq files for consensus generation<br> ├── BAM: alignment files generated from minimap2 used as input for the artic pipeline to identify variants<br> ├── Cluster_Rep.txt: Consensus sequences that were chosen to represent each haplotype<br> ├── Consensus folder: consensus fasta files generated for each sample containing sequences for each specific pathogen<br> ├── Coverage folder: coverage and base-level read depth for each sample and each pathogen reference genes<br> ├── Filter: individual fasta sequences (only 1 sequence per file) for each pathogen and each sample without any ambiguous base for subsequent clustering analysis<br> ├── Full_Haplotype.fasta: all possible haplotype sequences generated for each pathogen<br> ├── Gap_Analysis.tsv: Table with percentage of gap (0-100%) for each consensus sequence generated (used for filtering)<br> ├── Haplotype folder: Intermediate file and sample-level haplotype used to infer final haplotype and generate haplotype summary<br> ├── Haplotype_summary.tsv: Table with sample ID and their respectively pathogen haplotype<br> ├── Minimap2_PAF: Intermediate alignment generated from minimap2 used to generate the count table<br> ├── FailMinimap2 folder: FastQ files that didn't align using minimap2. Will be subsequently aligned using blastN (more sensitive) against the same reference sequences as minimap2<br> ├── Host_4Pathogen.fasta: Fasta file containing the tilapia genome and 4 pathogen (primer binding site included)<br> ├── Original: fastq with original naming prior to renaming based on sampleID. a script (rename.sh) was included to show renaming scheme<br> ├── primer.fasta: Primer sequences used for identifying and trimming reads with flanking primer sequence<br> ├── primer.fasta.fai: the index file for primer.fasta<br> ├── PrimerTrim folder: Primer-trimmed reads<br> ├── quast_results: consensus statistics generated by quast<br> ├── RawCount.tsv: Count table generated that can used as a input to generate figure<br> ├── RawFastq folder: Raw reads that have been renamed to reflect sample information<br> ├── readme.md: The current readme file<br> ├── ref_full_latest.fasta: Reference sequence of (gene segments) 4 pathogens e.g. TilV, ISKNV, SAG (Streptococcus agalactiae), FNO (Francisella noatunensis subsp. orientalis)<br> ├── ref_full_latest.primer.fasta: Same as above but with their primer binding sequence trimmed similar to the processed reads<br> ├── ref_full_latest.primer.fasta.fai<br> ├── RenameHaplotype: Script to perform reorganization of cdhit output<br> ├── Seq.stat.tsv: Sequencing statistics<br> ├── Uniq_PAF: Minimap2 alignment file for raw reads that initially failed quality check (no primer present and/or less than 80% query coverage / not unique alignment)<br> ├── Unmap: Raw reads that initial failed quality check (no primer on both ends / less than 80% query coverage / not unique alignment) <br> └── VCF: VCF files from medaka variant calling used to generate the final consensus</p>
BeMAGIC_Nanoporous films and ultra-thin films for magneto-ionics and surface charging experiments
<p>BeMAGIC ITN (GA861145) Nanoporous films and ultra-thin films for magneto-ionics and surface charging experiments. Results from TUC, UCAM, AALTO and KIT</p>
Methylation-free E.coli nanopore sequencing (ONT R9.4.1) data set
<p>The data set consists of fast5 files divided into 5 zip files (fast5_[1-5].zip), a genome record (Ecoli_K12_MG1655.fasta), an Illumina assembly genome (illumina_contigs.fasta) and a fastq file from Guppy 5 (guppy_basecalled.fastq.gz). We sequenced the Ecoli non-methylated genomic DNA (D5016, Zymo Research) with an ONT MinION device. The sequencing libraries were prepared by fragmenting the genomic DNA using Covaris g-TUBE and a Ligation sequencing kit (SQK-LSK109, Oxford Nanopore) with Flow Cell chemistry R9.4.1. We also performed short-read Illumina sequencing on the same sample using the TruSeq PCR-free library preparation on the MiSeq sequencing platform (Illumina, USA), and constructed a draft assembly from the Illumina sequencing results using SPAdes v3.6.0. We also upload a reference genome directly obtained from the E.coli sample producer website. </p> <p>In addition, the data set contains two fastq files that produced by the Lokatt basecaller (lokatt_basecalled.fasta.gz) and local-trained Bonito basecaller (bonito_local_basecalled.fastq.gz ), respectively, which are used for benchmarking in the Lokatt basecaller paper.</p>
Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads
<p>In the paper associated with this dataset, a workflow is presented which enables the identification of single-/low-copy nuclear molecular markers for a plant group of interest, by mining data from a small representative target capture experiment done using a commercial probe kit and Oxford Nanopore long-read sequencing. The proposed pipeline first assesses sequence variability contained in the data from targeted loci and assigns reads to their respective genes, via a combined BLAST/clustering procedure. Cluster consensus sequences are then examined based on four pre-defined criteria presumably indicative for absence of paralogy. This is done by calculating four specialized indices; loci are ranked according to their performance in these indices, and top-scoring loci are considered putatively single- or low-copy. The approach can be applied to any probe set. As it relies on long reads, the contribution also provides template workflows for processing Nanopore-based target capture data. Identified loci can be used for NGS amplicon sequencing. For detection of possibly remaining paralogy in these data, which might occur in groups with rampant paralogy, the long-read assembly tool CANU is employed. The presented workflow can be useful for researchers dealing with reticulate or polyploidization phylogenetic histories in plants.</p> <p>The present dataset contains several documents supplementing the original paper. Its most important elements are a detailed description (alongside two graphical workflow figures) of all methods employed in the study, suitable for reproducing the steps of the workflow and also the wet-lab work. The workflow employs a collection of BASH, Python and R scripts which is available here, together with a detailed account on command line use in Linux. Also, reference sequences for the identified markers can be found as well as sequence alignments derived from the amplicon sequencing.</p>
Protein sizing with 15-nm conical biological nanopore YaxAB
<p>This dataset belongs to the article: "Protein sizing with 15-nm conical biological nanopore YaxAB" and contains the raw electrophysiology data of protein capture by YaxAB, conductance distributions and reverse potential experiments. A Matlab script describing the analysis is also added. This dataset also includes the molecular dynamics simulations and analysis (see below)</p> <p>The Data_exp.zip contains experimental files; Data_MD.zip contains MD simulation files. </p> <p>For Data_exp.zip:</p> <p>- Each electrophysiology file contains measurement info as follows:<br> [Date of measurement]_[Buffer conditions]_[Pore type]_[added analyte(s)]_[operator initials]</p> <p>- Each electrophys trace is accompannied by Excel file with Clampfit analysis in Sheet1 (named [voltage], with the columns corresponding to: Trace; Search; Level; State; Event Start Time (ms); Event End Time (ms); Amplitude (pA); Amp S.D. (pA); Dwell Time (ms); Inst. Freq. (Hz); Interevent Interval (ms); [empy]; open pore current (pA)</p> <p>- This Excel file was analyzed with inhouse Matlab script (matlab_script_protein_capture_SAP); and event data (Ires, dwell time, Amp S.D.) was used in scatter plots. </p> <p> </p> <p><br> Data_exp_zip: Folder containing raw electrophysiology data and result after Clampfit analysis, with each folder containing the following:</p> <p>Figure 1: Current distributions of YaxAB and YaxA<sub>d40</sub>B</p> <p>Figure 2: I/V curves and reverse potential of YaxAB and YaxA<sub>d40</sub>B</p> <p>Figure 3: Protein capture by YaxA<sub>d40</sub>B</p> <p>Figure 4 and Figure S25-S26: Depleted serum capture and titrated CRP capture by YaxA<sub>d40</sub>B</p> <p>Figure S7: I/V curves of reverse potential experiments of YaxAB, YaxA<sub>d40</sub>B, YaxA<sub>d40</sub>B RRR, YaxA<sub>d40</sub>B NNN</p> <p>Figure S8: Current distributions of YaxA<sub>d40</sub>B RRR and YaxA<sub>d40</sub>B NNN</p> <p>FIgure S12: Protein capture by YaxAB</p> <p>FIgure S13: Protein capture by YaxA<sub>d40</sub>B</p> <p>Figure S14-S17, and S19-S21: Voltage-dependent protein capture by YaxA<sub>d40</sub>B</p> <p>Figure S18: Protein capture by YaxA<sub>d40</sub>B with different unitary conductance</p> <p>Figure S23: Concentration-dependent protein capture by YaxA<sub>d40</sub>B</p> <p>Figure S24: Mixed protein capture by YaxA<sub>d40</sub>B with different unitary conductance</p> <p>Figure S26: CRP capture in absence/presence of depleted serum by YaxA<sub>d40</sub>B</p> <p>matlab_script_protein_capture_SAP: Matlab script used to obtain event data in scatter plots </p> <p> </p> <p>Data_MD: Folder containing structures and<strong> </strong>MD analysis script and tools. Trajectories are not uploaded as they are more than 2Tb of data and can be easily reproduced using the provided files and described methods.</p> <p>Figure 1 (MD) - PDBs of the modeled didecameric YaxAB (<strong>yaxab-codine.pdb</strong>) and YaxA_{D40}B (<strong>yaxab-d40.pdb</strong>)</p> <p>Figure 2 (MD) - <strong>1)</strong> FORTRAN source code to obtain the axial-symmetric maps and resistance-per -length profile from the VMD Volmap density maps .dx (<strong>dx2radial.zip</strong>); <strong>2)</strong> VMD TCL script to compute the velocity map from the MD trajectories (.dcd) and FORTRAN code (<strong>get_vField.zip</strong>).</p> <p>Figure 3 (MD) - FORTRAN and bash codes to obtain the protein-pore resistance from the VMD Volmap occupancy 3D maps (<strong>Hindrance-Yax.zip</strong>)</p> <p>Figure S2 (MD) - VMD TCL script for MODELLER (<strong>Modella.tcl</strong>) and output PDB for the initial YaxAB nanopore model (<strong>yaxab-codine-init_MODELLER.pdb</strong>)</p> <p>Figure S4 (MD) - VMD TCL scripts to obtain the ionic fluxes, EOF and charge maps .dx (<strong>EFIELD_analysis.zip</strong>). The filtered map is obtained by the <strong>dx2radial </strong>FORTRAN<strong> </strong>program.</p> <p>Figure S5 (MD) - Nothing to upload<br> <br> Figure S6 (MD) - PDBs of mutated pores (<strong>yaxab-RRR.pdb</strong>, <strong>yaxab-NNN.pdb</strong>)</p> <p>Figure S22 (MD) - Analysis of the resistance were performed with <strong>dx2radial </strong>program; nothing additional to update.</p> <p> </p>
De novo nanopore sequencing overrepresents RNA modification landscape
<p>RNA modifications are critical to the functional diversity and regulatory complexity of the transcriptome. With increasing frequency, direct nanopore RNA sequencing is applied to identify RNA modifications de novo. Here, we directly compare the MS2 phage genome RNA modification profiles determined using nanopore to orthogonal LC-MS/MS assays. The results reveal very different views of the modification landscape, suggesting caution when calling new RNA modifications using nanopore alone.</p>
Supplementary materials to: Nano-Strainer: a workflow for identification of single-copy nuclear loci for plant systematic studies, using target capture kits and Oxford Nanopore long reads
Open the record for dataset details and reuse information.
Raw Nanopore data for "Nanopore Long-Read Guided Complete Genome Assembly of Hydrogenophaga intermedia, and Genomic Insights into 4-Aminobenzenesulfonate, p-Aminobenzoic Acid and Hydrogen Metabolism in the Genus Hydrogenophaga"
<p>This is the raw Nanopore dataset (fast5) for Hydrogenophaga intermedia PBC. The gDNA was prepared using the now obsolete SQK-NSK007 kit and sequenced on a MINION R9 Flowcell. </p>
Dataset for "Nanopore-led long-read genome assembly of the Australian yabby, Cherax destructor"
<p>Intermediate_Assemblies.tar.gz: Intermediate genome assemblies e.g. raw wtdbg assembly (CD.raw.fa), polished wtdbg assembly (CD.cns.fa), 1st pilon polished assembly (CDF2_pilon1.fasta), 2nd pilon polished assembly (CDF2_pilon2.fasta) and RNA-scaffolded assembly (CDF2_pilon2_prna.fasta). Folders with run_"assembly name" are BUSCO output for each of the assembly.</p> <p>BRAKER2.tar.gz: BRAKER2 genome annotation output containing the initial set of predicted protein-coding genes as well as training intermediate files.</p> <p>BUSCOv3.tar.gz: BUSCO assessment of publicly available Decapod crustacean genome assemblies</p> <p>Cdes.filtered.codingseq: Filtered set of protein-coding sequences</p> <p>Cdes.filtered.faa: Translation of the filtered protein-coding sequences</p> <p>CDF2.NCBI.fasta.masked.gz: Repeat-masked (softmasked) Cherax destructor genome</p> <p>Cqua_transcriptome.tar.gz: rnaSPAdes output (combined fasta) of all Cherax quadricarinatus transcriptomes and its reduced dataset generated by EvidentialGene. </p> <p>Quast.tar.gz: Quast output of all Decapod crustacean genome assemblies assessed in this study</p> <p>Repeat_Annotation.tar.gz: Repeat annotation (.gff3) based on Cherax destructor-specific de novo repeat library and its summary (.tbl)</p> <p>RepeatLibrary.tar.gz: Cherax destructor-specific de novo repeat library generated by RepeatModeler</p> <p>Wtdbg2_assembly.log: Wtdbg2.5 log file showing exact command used, kmer distribution, memory usage and assembly duration.</p> <p>CAZY_Annotation.tar.gz: dbCAN2 Identification of CAZy in the selected crustacean proteomes as well as list of cellulase-associated GH groups (glycoside hydrolase). </p> <p>Orthofinder.tar.gz: Orthofinder2 output and proteomes of each crustacean used to infer orthologous clustering.</p> <p>GH9_Analysis.tar.gz: Selected GH9-associated protein sequences, amino acid alignment and IQTree output. </p> <p>Cdes_mito.gbf: GenBank file of the annotated complete mitogenome</p> <p>Cdes.filtered.codingseq: Cherax destructor protein-coding genes with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping. </p> <p>Cdes.filtered.faa: Cherax destructor proteins with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping. </p> <p>Cdes.ortholog.list: List of predicted Cherax destructor proteins with homology to other crustacean proteomes based on Orthofinder2 orthologous grouping. </p>
Raw Fast5 data for "Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon" - PART I
<p>Raw Fast5 data for "Microbiota profiling with long amplicons using Nanopore sequencing: full-length 16S rRNA gene and the 16S-ITS-23S of the rrn operon". See Supplementary Table 2 for associating each sample to its barcode.</p> <p>- FC1_1 includes data for the HM mock community from BEI resources and skin microbiome of the chin in dogs.</p> <p>- FC1_2 includes data for the dorsal skin samples</p> <p>- FC2 includes data for the Zymobiomics mock community and Staphylococcus pseudintermedius isolate</p> <p> </p> <p> </p>
Rapid and real-time identification of fungi up to the species level with long amplicon Nanopore sequencing from clinical samples
<p>Samples collected from fungal cultures, skin of dogs and ZymoBIOMICS<sup>TM </sup>mock community (which includes <em>Saccharomyces cerevisiae</em> and <em>Cryptococcus neoformans</em>). The amplicons length of the fungal cultures and ZymoBIOMICS<sup>TM </sup>mock community is 3,5 Kb and 6 Kb, while the <em>Malassezia spp</em> samples used as control is 3,5 Kb. The amplicons length of the four samples from the skin is 3,5 Kb.</p>
Data from: The preparation and characterization of uniform nanoporous structure on glass
<p>A novel fabrication method of uniform porous structures on the glass surface is proposed. The hydrofluoric acid fog formed by air-jet atomization etches the glass surface to fabricate nanoporous structure (NPS) on glass surface. This NPS shows the enhanced average light transmittance of ~92.9% and the superhydrophilic property with a contact angle less than 1° which presents an excellent anti-fog property. Passivated by fluorosilane, the NPS shows nearly the superhydrophobic property with a contact angle of 141.2°. This fabrication method has shown promising application prospects due to its simplicity, low cost and efficiency, which can be easily applied to large-scale industrial production.</p>
CAMISIM Default Nanopore Community
<p>Dataset generated with CAMISIM, using the mini_config.ini on Dec 20, 2020. The read simulator was swapped for NanoSim (<a href="https://github.com/abremges/NanoSim">https://github.com/abremges/NanoSim</a>), and the output set to 0.1Gb.</p>
Data set related to the manuscript "Simulations of ionic liquids confined in surface-functionalized nanoporous carbons: Implications for energy storage"
<p>Graphical files in the agr format for all the figures in the manuscript entitled "Simulations of ionic liquids confined in surface-functionalized nanoporous carbons: Implications for energy storage". Examples of input files for the three systems simulated.</p>
SMAdd-seq: Probing chromatin accessibility with small molecule DNA intercalation and nanopore sequencing
<p>Studies of in vivo chromatin organization have relied on the accessibility of the underlying DNA to nucleases or methyltransferases, which is limited by their requirement for purified nuclei and enzymatic treatment. Here, we introduce a nanopore-based sequencing technique called Small-Molecule Adduct sequencing (SMAdd-seq), where we profile chromatin accessibility by treating nuclei or intact cells with a small molecule, angelicin. Angelicin reacts with thymine bases in linker DNA not bound to core nucleosomes after UV light exposure, thereby labeling accessible DNA regions. By applying SMAdd-seq in Saccharomyces cerevisiae, we demonstrate that angelicin-modified DNA can be detected by its distinct nanopore current signals. To systematically identify angelicin modifications and analyze chromatin structure, we developed a neural network model, NEural network for mapping MOdifications in nanopore long-reads (NEMO). NEMO accurately called expected nucleosome occupancy patterns near transcription start sites at both bulk and single-molecule levels. We observe heterogeneity in chromatin structure and identify clusters of single-molecule reads with varying configurations at specific yeast loci. Furthermore, SMAdd-seq performs equivalently on purified yeast nuclei and intact cells, indicating the promise of this method for in vivo chromatin labeling on long single molecules to measure native chromatin dynamics and heterogeneity.</p>
CRAFTED: An exploratory database of simulated adsorption isotherms of nanoporous materials
<p><strong>Overview</strong></p><p>The files in this repository compose the <strong>C</strong>harge-dependent, <strong>R</strong>eproducible, <strong>A</strong>ccessible, <strong>F</strong>orcefield-dependent, and <strong>T</strong>emperature-dependent <strong>E</strong>xploratory <strong>D</strong>atabase (<strong>CRAFTED</strong>) of adsorption isotherms. This dataset contains the simulation of CO2 and N2 adsorption isotherms on 690 metal-organic frameworks taken from the CoRE-MOF-2014 database and 667 covalent organic frameworks taken from the CURATED-COFs database. The simulations were performed with two force fields (UFF and DREIDING), six partial charge schemes (no charges, Qeq, EQeq, DDEC, MPNN, and PACMOF), and three temperatures (273, 298, 323 K).</p><p><strong>Contents</strong></p><ul><li>CIF_FILES/ contains 6 folders (NEUTRAL, DDEC, EQeq, Qeq, MPNN, and PACMOF), each one with 1357 CIF files;</li><li>FORCEFIELDS/ contains 2 folders (UFF and DREIDING) with the definition of the forcefields;</li><li>INPUT_FILES/ contains 97,704 input files for the GCMC simulations;</li><li>ISOTHERM_FILES/ contains 97,704 adsorption isotherms resulting from the GCMC simulation;</li><li>ENTHALPY_FILES/ contains 97,704 enthalpies of adsorption from the isotherms;</li><li>RAC_DBSCAN/ contains the RAC and geometrical descriptors to perform the t-NSE + DBSCAN analysis;</li></ul><p><strong>Licenses</strong></p><p>The 690 MOF-related CIF files in the DDEC folder were downloaded from <a href="https://doi.org/10.5281/zenodo.3986573">CoRE-MOF-2014</a> and are licensed under the terms of the Creative Commons Attribution 4.0 International license (<a href="https://creativecommons.org/licenses/by/4.0/legalcode">CC-BY-4.0</a>). The 667 COF-related CIF files in the NEUTRAL folder were downloaded from <a href="https://github.com/danieleongari/CURATED-COFs">CURATED-COFs</a> and are licensed under the terms of the MIT license (<a href="https://github.com/danieleongari/CURATED-COFs/blob/master/LICENSE">MIT</a>).</p><blockquote><p>Dalar Nazarian, Jeffrey S. Camp, & David S. Sholl. (2016). Computation-Ready Experimental Metal-Organic Framework (CoRE MOF) 2014 DDEC Database [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.3986573">https://doi.org/10.5281/zenodo.3986573</a></p><p>Ongari, Daniele, et al. "Building a consistent and reproducible database for adsorption evaluation in covalent–organic frameworks." ACS Central Science 5.10 (2019): 1663-1675. <a href="https://doi.org/10.1021/acscentsci.9b00619">https://doi.org/10.1021/acscentsci.9b00619</a></p><p>Ongari, Daniele, Leopold Talirz, and Berend Smit. "Too many materials and too many applications: An experimental problem waiting for a computational solution." ACS Central Science 6.11 (2020): 1890-1900. <a href="https://doi.org/10.1021/acscentsci.0c00988">https://doi.org/10.1021/acscentsci.0c00988</a></p></blockquote><p>The CO2.def and N2.def forcefield files were downloaded from <a href="https://github.com/iRASPA/RASPA2/tree/master/molecules/ExampleDefinitions">RASPA</a> and are licensed under the terms of the <a href="https://github.com/iRASPA/RASPA2/blob/master/COPYING">MIT</a> license.</p><blockquote><p>Dubbeldam, David, et al. "RASPA: molecular simulation software for adsorption and diffusion in flexible nanoporous materials." Molecular Simulation 42.2 (2016): 81-101. <a href="https://doi.org/10.1080/08927022.2015.1010082">https://doi.org/10.1080/08927022.2015.1010082</a></p></blockquote><p>The remaining MOF-related CIF files in the PACMOF, MPNN, Qeq, EQeq and NEUTRAL folders were derived from those in the DDEC folder and are licensed under the terms of the Creative Commons Attribution 4.0 International license (<a href="https://creativecommons.org/licenses/by/4.0/legalcode">CC-BY-4.0</a>) from the CoRE-MOF-2014 subset. The remaining COF-related CIF files in the PACMOF, MPNN, Qeq, EQeq and DDEC folders were derived from those in the NEUTRAL folder and are licensed under the terms of the MIT license (<a href="https://github.com/danieleongari/CURATED-COFs/blob/master/LICENSE">MIT</a>) from the CURATED-COFs subset.</p><p>All remaining files were created by us, and are licensed under the terms of the <a href="https://cdla.dev/sharing-1-0/">CDLA-Sharing-1.0</a> license.</p><p><strong>Software requirements</strong></p><p>In order to create a Python environment capable of running the Jupyter notebooks, please install <a href="https://docs.conda.io/en/latest/miniconda.html">conda</a> and execute</p><p>conda env create --file environment.yml</p><p><strong>Usage instructions</strong></p><p>Execute the command below to run JupyterLab in the appropriate Python environment.</p><p>conda run --name crafted jupyter-lab</p><p> </p>
Hydrophobically gated memristive nanopores for neuromorphic computing
<p>The file named "experimental.zip" has the .abf files of the different electrophysiology experiments realized on the engineered FraC.</p><p>The file named "model_pore.zip" has the initial condition, the LAMMPS file to run the RMD simulations as well as the files required to compute the free energy, P1 and P2.</p><p>The file named "frac_md.zip" has the files to run the FraC simulations, as well as the files necessary to compute the free energy, P1 and P2.</p><p> </p>
De novo genome assembly of rice varieties using Nanopore long reads
<p>Genome sequences for Sugimura et al. (2024) of the rice (O. sativa) varieties 'Hitomebore' and 'Arroz da Terra.'</p> <p>Yusaku Sugimura, Kaori Oikawa, Yu Sugihara, Hiroe Utsushi, Eiko Kanzaki, Kazue Ito, Yumiko Ogasawara, Tomoaki Fujioka, Hiroki Takagi, Motoki Shimizu, Hiroyuki Shimono, Ryohei Terauchi, Akira Abe. Impact of rice GENERAL REGULATORY FACTOR14h (GF14h) on low-temperature seed germination and its application to breeding. PLoS Genet 20(8): e1011369. https://doi.org/10.1371/journal.pgen.1011369</p> <p>bioRxiv doi: https://doi.org/10.1101/2024.02.16.580620</p>
Data for: Sensitive and specific detection of tumor-derived exosomes using nanopore-crystal microchips
<p>Tumor-derived exosomes (tExos) have emerged as promising circulating biomarkers for early cancer diagnosis. However, the sensitivity and specificity of existing assays often limit the clinical translation of tExos. In this study, we present a highly versatile microfluidic platform, termed the nanopore-crystal microchip (NC-Chip), for specific isolation and ultrasensitive detection of tExos in as little as 0.5 μL of plasma samples from pancreatic cancer (PC) patients. The NC-Chip incorporates a herringbone-patterned hierarchical porous hydrogel scaffold, enabling fluid manipulation, size exclusion, immunoaffinity capture, and signal amplification. These integrated features significantly enhance the sensitivity and specificity of tExos assays in complex clinical scenarios. Utilizing an eight-protein signature, the NC-Chip can distinguish patients with pancreatitis and non-metastatic PC with 100% accuracy in the training cohort and 94.9% accuracy in the validation cohort. This platform is sensitive, specific, inexpensive, and only needs small-volume samples, offering a powerful exosome-based liquid biopsy tool for early PC diagnosis.</p>
Comprehensive, targeted eukaryotic metagenomics analysis of environmental DNA biodiversity using Oxford Nanopore sequencing
<p><span>Metagenomics has become a prominent technology for studying the functional potential of all organisms in a microbial and eukaryotic community. The study of symbiotic organisms from different classes or kingdoms, including those previously unknown, is possible with simultaneous and equally efficient metagenomic analysis of these species. A variety of targeted primer sets are used for eukaryotic metagenomic biodiversity, including those that are universal for specific families, classes</span><span>,<span> or kingdoms. The most universal for all existing cellular organisms is the presence of ribosomal RNA encoding gene sequences. For eukaryotic sequences, these are 16S and 23s rDNA, </span>and <span>for eukaryotic sequences of nuclear (18S and 28S) and mitochondrial (12S and 16S) ribosomal RNA. Here we present the application of the eukaryotic metagenomics approach to the simultaneous, quantitative</span>,<span> and unbiased identification of most eukaryotic species. To achieve this, we have developed a universal PCR assay that targets the most conservative nuclear regions of the ribosomal gene for all cellular organisms, including plants, algae, fungi, protists, insects</span>,<span> and animals. The amplification product contains polymorphic regions of both ribosomal genes and the intergenic spacer. The size of the PCR products varies by class, kingdom</span>,<span> or domain, ranging from 2 kb for fungi to 7 kb for birds. This assay is also adapted for use with the Oxford Nanopore Rapid Barcoding Library Kit, which enables metagenomic biodiversity analysis. Our approach provides a rapid, sensitive</span>,<span> and equally efficient way to study the composition of eDNA from mixed species in the environment. This protocol reduces the time and cost of metagenomic biodiversity analysis using Oxford Nanopore sequencing. We can efficiently analyze the biodiversity of mixed species present in environmental samples.</span></span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.