Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
468
datasets available to search
ShareScore release 0.7.1
Dataset results
468 results for “Deep sequencing”
Overcoming Limitation of AlphaFold2 by Deep-mutational Scanning and Stability-Selection of Protein Sequences
<p>This repository contains the processed datasets and corresponding code used in our study. While AlphaFold2 revolutionizes protein structure prediction, its accuracy critically depends on evolutionary information from natural homologs—limiting applications for proteins with sparse sequence families. Here, we bypass this bottleneck by employing deep mutational scanning and stability-guided selection to generate artificial homologs. Fed into AlphaFold2, these synthetic sequences match the accuracy achieved on well-predicted proteins with rich natural homology, while providing highly accurate predictions for difficult targets—including orphan proteins previously deemed "unpredictable." Our approach achieves high accuracy (<3 Å RMSD for 5/8 and <2 Å RMSD for 7/8 targets after excluding intrinsically flexible regions). Thus, integrating simple, scalable molecular biology (mutagenesis/selection) with high-throughput sequencing can deliver the accuracy similar to but at a fraction of the cost and time of traditional experimental structure-determination methods. This hybrid framework could democratize high-resolution structural biology, opening avenues to determine structures of protein complexes, modified proteins, and condition-dependent conformations. </p>
Deep learning-based earthquake catalog of the 2022 MW 6.9 Chihshang, Taiwan, earthquake sequence
<p>On 18 September 2022, the MW 6.9 Chihshang earthquake struck the southern Longitudinal Valley, Taiwan. We use SeisBlue, a deep-learning platform/package, to extract the two-month earthquake sequence from September to October 2022, including the MW 6.5 Guanshan foreshock, the MW 6.9 mainshock, over 14,000 aftershocks, and 866 focal mechanisms from two sets of broadband networks. For more details, please refer to our research article published at TAO (Sun et al., 2024; https://doi.org/10.1007/s44195-024-00063-9). The refined SeisBlue earthquake, FMS, and 20-year M3+ relocated CWA earthquake catalogs obtained in this study are listed here.</p> <ol> <li>The refined, deep-learning-based earthquake catalog of the 2022 Mw 6.9 Chihshang, Taiwan, earthquake sequence contains 5,151 seismic events with event time, location and error information, local and moment magnitudes, and hypoDD location. </li> <li>The FMS (focal mechanism solution) catalog is obtained by the P-wave polarities of 14 broadband stations and the FPFIT program (Reasenberg & Oppenheimer, 1985). 865 out of 1629 FMSs with at least six readings of P-wave polarity, F-fator <span>≤ </span>0.1 (F <span>< </span>0.5 for a good fit), and errors of strike, dip, and rake are all <span>< </span>20<span>°</span>, respectively, are listed in the attached FMS catalog.</li> <li>The 2001-2020 3D-hypoDD-relocated M3+ CWA earthquake catalog: We applied the HypoDD program (Waldhauser & Ellsworth, 2000) to the CWA (Central Weather Administration (CWA, Taiwan), 2012) catalog with P- and S-wave arrivals and obtained 5862 M3+ events between 2001 and 2020 for eastern Taiwan. The 3D velocity models used for this catalog are the local models from Kuo-Chen et al. (2012).</li> </ol> <p> </p>
VirHunter: a deep learning-based method for detection of novel RNA viruses in plant sequencing data
<p>This storage contains 2 archives: toy datasets to test the training of the VirHunter and weights of the fully trained VirHunter models for 3 host species (peach, grapevine, sugar beet) and for fragment sizes 500 and 1000. .</p> <p>The toy dataset consists of 3 archived files: 'viruses.fasta', 'host.fasta', 'bacteria.fasta'.</p> <p>'viruses.fasta' contains 10000 randomly selected plant viruses from the virus dataset described in the paper.</p> <p>'host.fasta' consists of peach chromosome 2.</p> <p>'bacteria.fasta' consists of 10 bacterial genomes selected randomly: GCF_000284415, GCF_000590555, GCF_001548055, GCF_002795265, GCF_003330825, GCF_003957805, GCF_005845345, GCF_009176625, GCF_010748935, GCF_014681765</p> <p> </p>
Data supporting publication: Ultra-deep Sequencing of Hadza Hunter-Gatherers Recovers Vanishing Gut Microbes
<p>Genomes, mapping databases, and data supporting the publication "Ultra-deep Sequencing of Hadza Hunter-Gatherers Recovers Vanishing Gut Microbes".</p> <p><strong>Please see the README on GitHub for descriptions of files and tutorials on how to use them: <a href="https://github.com/MrOlm/ZenodoREADME/blob/main/README.md">https://github.com/MrOlm/ZenodoREADME/blob/main/README.md</a></strong></p>
Deep sequencing datasets from: RNA-catalyzed evolution of catalytic RNA
<p>This dataset includes raw and processed sequencing data from evolving RNA populations described in Nikos Papastavrou, David P Horning, Gerald F Joyce. "RNA-Catalyzed Evolution of Catalytic RNA" (submitted). Briefly, directed evolution of a hammerhead ribozyme sequence was carried out over eight rounds of three steps each: 1) templated synthesis of the reverse-complement hammerhead RNA by a polymerase ribozyme; 2) templated synthesis of a new copy of the hammerhead RNA from the reverse-complement by a polymerase ribozyme; 3) selective recovery of hammerhead RNA that cleaved an attached RNA substrate. Cleaved RNA was reverse transcribed, PCR amplified, and archived for sequencing, while a portion was in vitro transcribed with T7 RNA polymerase to initiate the next round of evolution. Two distinct branches of evolution were carried out for 8 rounds, using the '52-2' or '71-89' polymerase ribozymes to replicate RNA, respectively. Sequenced RNA populations were analyzed to determine polymerase ribozyme fidelity and study the evolution of hammerhead sequences replicated by polymerases with low or high RNA copying fidelities. The dataset includes raw sequence files, processed tables of mutations by position along the sequence for each polymerase, processed tables of the sequence frequency distribution from each round in the evolving populations, and spreadsheets containing final processed data used directly in manuscript figures and tables.</p>
Deep-sequencing of viral genomes from treatment-naive HIV-infected persons shows positive association between intrahost genetic diversity and viral load
<p><strong><span>Background:</span></strong><span> Infection with human immunodeficiency virus type 1 (HIV) typically results from transmission of a small and genetically uniform viral population. Following transmission, the virus population becomes more diverse because of recombination and acquired mutations through genetic drift and selection. Viral intrahost genetic diversity remains a major obstacle to the cure of HIV; however, there is a disagreement whether intrahost viral genetic diversification associates positively or negatively with disease progression and progression markers. Viral load is a key progression marker and understanding its relationship to viral intrahost genetic diversity could help design future strategies for HIV monitoring and treatment.</span></p> <p><span><strong>Methods:</strong> </span><span>We analyzed deep-sequenced viral genomes from 2,650 treatment-naive HIV-infected persons to measure the intrahost genetic diversity of 2,447 genomic codon positions as calculated by Shannon entropy. We tested for associations between viral load (VL) and amino acid (AA) entropy accounting for sex, age, race, duration of infection, and HIV population structure.</span></p> <p><strong><span>Results:</span></strong><span><strong> </strong>We confirmed that the intrahost genetic diversity is highest in the <em>env</em> gene. Furthermore, we showed that mean Shannon entropy is significantly associated with VL, especially in infections of >24 months duration. We identified 16 significant associations between VL (p-value<2.0x10<sup>-5</sup>) and Shannon entropy at AA positions which in our association analysis explained 13% of the variance in VL.</span></p> <p><strong><span>Conclusions: </span></strong><span>Our results elucidate that viral intrahost genetic diversity is associated with VL and could be used as a better disease progression marker than HIV consensus sequence variants, especially in infections of longer duration. We emphasize that viral intrahost diversity should be considered when studying viral genomes and infection outcomes.</span></p>
DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing
<p>Curated dataset for the manuscript named "DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing".</p> <p>Five publicly available datasets sequenced on ONT MinION/GridION were used for the experiments (see below for original sources). These datasets contained raw signal data in single-FAST5 format (one file per each read), which were converted to BLOW5 format using slow5tools to enable convenient and efficient file manipulation. Then, 40,000 reads containing at least 4500 signal samples were extracted from each dataset. From each dataset, 20,000 reads are for training (<species>/train-<species>.blow5) and the rest for testing (<species>/test-<species>.blow5). Basecalled reads for the dataset used for testing are also available (test-<species>.fastq). Guppy version 6.1.3 under dna_r9.4.1_450bps_hac mode was used. The reference genomes are also given (<species>/<species>-ref.fasta)</p> <p>Original datasets are from the following sources:<br> SARS-CoV-2: https://community.artic.network/t/links-to-raw-fast5-fastq-data-for-artic-protocol/17<br> Zymo Metagenome: https://github.com/LomanLab/mockcommunity<br> Chlamydomonas: https://sra-download.ncbi.nlm.nih.gov/traces/era20/ERZ/003237/ERR3237140/Chlamydomonas_0.tar.gz<br> Saccharomyces cerevisiae: https://www.ncbi.nlm.nih.gov/bioproject/PRJNA510813</p>
Deep Learning Neural Network Development for the Classification of Bacteriocin Sequences Produced by Lactic Acid Bacteria
<p>This project contains the following underlying data:</p> <h3> Software-Related Files</h3> <p><strong> · </strong><strong>BacLABNet_script.ipynb</strong> (Deep Learning Neural Network for classification of Bacteriocin Sequences) </p> <p><strong> · </strong><strong>embed_proteins.py </strong>(Recurrent Neural Network to obtained the embedding vectors)</p> <p> · <strong> </strong><strong>model_I22.h5</strong> (This file contains the trained weights of the trained model)</p> <p> · <strong>model_I22.json</strong> (This file contains the structure of the trained model)</p> <p><strong> · </strong><strong>rnn_gru.pt</strong> (Initial weights of the Recurrent Neural Network to obtain embedding vectors)</p> <p><strong> · </strong><strong>List_kmers.csv</strong> (List of 5-mers and 7-mers obtained from dataset after it filtered sequences shorter than 50 aa and longer than 2000 aa)</p> <p><strong> </strong></p> <p> <strong>Files Used for Training, Testing, and Validation of the Neural Network</strong></p> <p> <strong>· </strong> <strong>data_nonBacLAB.csv</strong> (25000 nonBacLAB amino acid sequences retrieved from Uniprot)</p> <p> <strong>· data_BacLAB.csv</strong> (24964 BacLAB amino acid sequences retrieved from Uniprot)</p> <p> <strong> </strong></p> <p><strong> Additional Files</strong></p> <p><strong> · data_BacLAB_and_nonBacLAB.csv </strong>(Combination of sequences from data_BacLAB.csv and data_nonBacLAB.csv)</p> <p> <strong> · </strong><strong>all k.mers list.xlsx </strong>(Table of all k-mers obtained for k=3,5,7,15,20)</p> <p> </p> <p><strong>Note:</strong> Codes are additionally available on GitHub. </p> <p><span>Data are available under the terms of the Creative Commons Zero "No rights reserved" data waiver (CC0 1.0 Public domain dedication)(http://creativecommons.org/publicdomain/zero/1.0/)</span></p>
Nanopore deep sequencing as a tool to characterize and quantify aberrant splicing caused by variants in inherited retinal dystrophy genes
Open the record for dataset details and reuse information.
Sewage - deep sequencing data
Open the record for dataset details and reuse information.
Deep learning model for characterizing protein-RNA interactions from sequence at single-base resolution
<p> </p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip.h5</a> - This file contains the training, validation, and test data for the Reformer model.</p> <p><a href="https://zenodo.org/api/records/14021440/draft/files/encode_eclip_bc.h5/content" target="_blank" rel="noopener noreferrer">encode_eclip_bc.h5</a> - This file contains the training, validation, and test data for the Reformer-BC model.</p> <p><a href="https://zenodo.org/api/records/14027315/draft/files/Reformer-code.zip/content" target="_blank" rel="noopener">Reformer-code.zip</a> - This file contains the training code of Reformer.</p>
CNN models and training, validation and test datasets for "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"
<p>Convolutional neural network (CNN) models and their respective training, validation and test datasets used in manuscript:</p> <p>Tuomo Hartonen, Teemu Kivioja and Jussi Taipale, "PlotMI: interpretation of pairwise interactions and positional preferences learned by a deep learning model from sequence data"</p>
Supplementary data for: DNA sequences are as useful as protein sequences for inferring deep phylogenies
<p>Inference of deep phylogenies has almost exclusively used protein rather than DNA sequences, based on the perception that protein sequences are less prone to homoplasy and saturation or to issues of compositional heterogeneity than DNA sequences. Here we analyze a model of codon evolution under an idealized genetic code and demonstrate that those perceptions may be misconceptions. We conduct a simulation study to assess the utility of protein versus DNA sequences for inferring deep phylogenies, with protein-coding data generated under models of heterogeneous substitution processes across sites in the sequence and among lineages on the tree, and then analyzed using nucleotide, amino acid, and codon models. Analysis of DNA sequences under nucleotide-substitution models (possibly with the third codon positions excluded) recovered the correct tree at least as often as analysis of the corresponding protein sequences under modern amino acid models. We also applied the different data-analysis strategies to an empirical dataset to infer the metazoan phylogeny. Our results from both simulated and real data suggest that DNA sequences may be as useful as proteins for inferring deep phylogenies and should not be excluded from such analyses. Analysis of DNA data under nucleotide models has a major computational advantage over protein-data analysis, potentially making it feasible to use advanced models that account for among-site and among-lineage heterogeneity in the nucleotide-substitution process in inference of deep phylogenies.</p>
Deep-sequencing of viral genomes from treatment-naive HIV-infected persons shows positive association between intrahost genetic diversity and viral load
Open the record for dataset details and reuse information.
Deep sequencing datasets from: RNA-catalyzed evolution of catalytic RNA
Open the record for dataset details and reuse information.
Supplementary data for: DNA sequences are as useful as protein sequences for inferring deep phylogenies
Open the record for dataset details and reuse information.
A rapid and versatile tool for HIV-1 Drug Resistance Genotyping by Deep Sequencing: supporting dataset
<p>See file `viral_mixes.md` and [manuscript](http://www.sciencedirect.com/science/article/pii/S0166093416301987).</p> <div class="grammarly-disable-indicator"> </div> <div class="grammarly-disable-indicator"> </div>
Ultra-deep sequencing of HIV-1 near full-length and partial proviral genomes reveals high genetic diversity among Brazilian blood donors
<p>Here, we aimed to gain a comprehensive picture of the HIV-1 diversity in the northeast and southeast part of Brazil. To this end, a high-throughput sequencing-by-synthesis protocol and instrument were used to characterize the near full length (NFLG) and partial HIV-1 proviral genome in 259 HIV-1 infected blood donors at four major blood centers in Brazil: Pro-Sangue foundation (São Paulo state (SP), n 51), Hemominas foundation (Minas Gerais state (MG), n 41), Hemope foundation (Recife state (PE), n 96) and Hemorio blood bank (Rio de Janeiro (RJ), n 70).</p>
ProTInSeq: transposon insertion tracking by ultra-deep DNA sequencing applied to identify small and large translated ORFs
<p>ProTInSeq is a novel -omics technique designed to characterize proteomes by using DNA ultra-deep sequencing. The technique is based on transposons engineered to have a positive or negative protein selection marker expressed when the transposon is inserted in-frame into a protein-coding gene. In the genome-reduced bacterium Mycoplasma pneumoniae, ProTInSeq identifies 80% of known expressed proteins, as well as 5 new open reading frames (ORFs; >100 amino acids); and 153 novel small ORF-encoded proteins (SEPs; ≤100 aa) that represent up to 18% of this bacterium’s proteome. ProTInSeq can be used to detect translational noise, for protein quantification and to provide insight into functional protein aspects such as relative half-life, stability, and membrane topology. Herein, we describe a methodology that can be easily implemented in any living system and allows the deep understanding of proteomes and more importantly the identification of small proteins by DNA ultra-sequencing.</p> <p>We include the following files:</p> <p>- processed_inscalling.zip: output obtain after running FASTQINS transposon calling tool over the datasets found at <a href="https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee">https://www.ebi.ac.uk/biostudies/arrayexpress/studies/E-MTAB-10380?key=5f54209d-ce59-490a-9bcd-7084e9c619ee</a>. This include every genome position in <em>M. pneumoniae</em>, the number of times an insertion has been mapped to that position and the total read count value.</p> <p>- separated_library_metrics.zip: insertion and read count processed from processed_inscalling files associated to every ORF and intergenic region in <em>M. penumoniae. </em>Columns include frame measured (0 - whole gene, 1 - in-frame, 2 and 3 for following positions) and metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant).</p> <p>- allmetrics.xlsx: merged table with the combination of results from separated_library_metrics.zip tab-delimited files.</p> <p>- selective_metrics_allannotations.xlsx: this table includes all the available information about the 30,112 sequences that could encode for a coding sequence in <em>M. pneumoniae</em>. For each identifier (column B), we include coordinates information and nucleotide and amino acid length information (columns C-H). Column I includes the gene name when the entry is found annotated in <em>M. pneumoniae</em>. Localization and function are described in columns J and K. Column L includes the operon number in which the annotation would be expressed. We also included transcription-related information average expression (column M; as log2(gene read count/gene length) and estimated average RNA copies per cell (column N) considering 4 RNA sequencing samples covering different growth times (6, 24 and 48 hours, ArrayExpress identifier E‐MTAB‐6203). Column O accounts for the number of mass spectrometry experiments detecting that entry (to a maximum of 116) and column P accounts for the total number of unique tryptic peptides detected. This comes <a href="https://paperpile.com/c/BImj5N/eljz">[5]</a>, available for 12,426 sequences that present an amino acid length ≥19 (from 116 mass spectrometry experiments, ID PRIDE: PXD008243). Columns Q to T recapitulate protein copies per cell under different conditions (overall, extracting with urea, extracting with SDS and mean, respectively). Column U includes half-lives of the proteins. Columns V and W describe the reference density of insertion and essentiality assigned in previous studies. Column X-AA includes the predicted RanSEPs score, ribosome binding site presence, homology into seven groups: 0—no hits passed the thresholds defined; 1—conserved with an annotated function; 2—conserved as an annotated SEP in NCBI but no associated function; 3—conserved in a different species but target and homologous sequence not found in NCBI; 4—sequence is completely or partially (> 75%) repeated ≥ 3 times in the reference genome; 5—potential pseudogene; and 6—to depict those annotations that are found in the reference NCBI annotation file. , and function expected by homology, respectively. Columns AB to AD cover the output provided by Phobius, including the number of transmembrane segments, presence of signal peptide and transmembrane topology predicted by TM-HMM. Column AE includes the complex information where 1 implies that entry is functional as a monomer, 2 as dimer, and so on. Finally, columns AF-AH will be 1 if the protein is a Lon protease target, a lipoprotein, and/or a truncated gene or pseudogene, respectively, 0 otherwise. Following columns include for every sample presenting selective insertion rates in-frame using the following identifiers separated by underscores: marker (BarnB, Cm or Ery), type (control-AC or selection-BD, antibiotic concentration, sample replicate, frame measured, metric. Metrics account for number insertions in-frame (<em>I</em>), read count (<em>R</em>), linear density from non-coding regions used in the Poisson evaluation (<em>rNC</em>), probability measured (<em>sfNC</em>) and a binary for prediction (<em>pred</em>; 0 - no significant, 1 - significant). Last columns combine the number of samples each annotation has been identified. Notice for barnase library the results need to be interpreted considering it is a negative selection marker inverting the 0 and 1 meaning.</p>
(Extended Data) Amplicon deep sequencing of ama1 and mdr1 to track within-host P. falciparum diversity throughout treatment in a clinical drug trial
<p>These extended data accompany the manuscript: Targeted Amplicon deep sequencing of ama1 and mdr1 to track within-host <em>P. falciparum</em> diversity throughout treatment in a clinical drug trial</p> <p><strong>Table S1: Concentration ratios and resulting parasitemia in artificial dna mixtures of P. falciparum Lab Isolates 3D7 and Dd2.</strong> This table presents the parasitemia for the artificial mixtures of P. falciparum lab isolates 3D7 and Dd2. Each mixture was prepared at varying ratios of 3D7 to Dd2, starting from equal proportions to a complete presence of only 3D7. The original concentration of each isolate was approximately 50,000 parasites per microliter (pf/μl), and the table displays the proportion of each strain in the mixture and the resulting total parasitemia concentration.</p> <p><strong>Table S2. List of PCR and deep sequencing primers.</strong> This table shows the list of forward and reverse primers used for deep sequencing. In boldface are the MID tags, while in the regular face are the forward primers</p> <p><strong>Table S3. The relative frequencies of each ama1 variant and the number of samples with each variant.</strong> The relative frequencies (%) of the 33 AMA1 variants in pre-and post-treatment samples (n = 330) are shown as a 33 amino acid sequence. The frequencies were calculated by dividing the number of reads of each microhaplotype by the total number of reads obtained per sample (116,187,131).</p> <p><strong>Table S4. Distribution of microhaplotypes among samples.</strong> This table shows the occurrence of microhaplotypes across all participants, both with monoclonal and multiclonal ama1 infections. It presents the ama1 clonality – monoclonal or multiclonal (column 1) - participant IDs (column 2), microhaplotype IDs (column 3), and the relative frequencies of these microhaplotypes across timepoints from 0 to 1008 hours (day 42) (column 3). Dashes represent time points where microhaplotypes were missing or were not detected.</p> <p><strong>Table S5. Distribution of rare microhaplotypes among samples.</strong> This table shows the occurrence of rare microhaplotypes in various samples. It presents participant IDs (column 1), microhaplotype IDs (column 2), and the relative frequencies of these microhaplotypes across time points from 0 to 1008 hours (day 42) (column 3). Samples containing rare microhaplotypes - specifically from PID10, PID32, PID38, PID40, PID49, PID60, PID63, and PID65 - are shown in orange, along with the corresponding rare microhaplotypes and their time points of occurrence. Furthermore, participants are categorised by shared microhaplotypes to indicate instances of rarity and commonality. Except for one microhaplotype unique to PID30, rare microhaplotypes were detected in several samples, frequently exceeding a 5% relative frequency. Dashes represent time points where microhaplotypes were missing or were not detected.</p> <p><strong>Table S6. The parasitemia levels associated with each ama1 microhaplotype per timepoint.</strong> This table shows the parasitemia for each ama1 microhaplotype per timepoint and each participant. “Patient ID” represents the patient ID, “AMA1 COI at 0h” represents the complexity of infection (COI) for each participant at baseline, based on ama1 while subsequent columns represent the parasitemia for each ama1 microhaplotype from timepoint 0h to 1008h. Parasitemia was back-calculated using the COI and total parasitemia for each time point. For time points with a COI > 1, parasitemia for the respective ama1 microhaplotypes are separated by commas, cells in red indicate timepoints without sequencing data (ND = not determined). In contrast, cells in grey indicate time points where microhaplotypes were detected below 10 parasites/μl, hence at risk of falling below the sampling limit.</p> <p><strong>Figure S1. Performance of AmpSeq in the sequencing controls.</strong> Six aliquots were prepared for each control set to ensure sufficient control data in case of PCR or sequencing failure. The median read depth in the lab controls was 5,658 (range 4,310 – 12,603) and 704 (291 – 1,676). The x-axis represents the aliquot identifier across the five mixtures, starting from 1 to 6, while the y-axis represents the proportions of each variant across all aliquots. For ama1 (A), two variants (3D7 and Dd2) were detected, whereas in mdr1 (B), two variants were detected YY, FY and NY following amplification of Dd2 Copy I, Dd2 Copy II and 3D7, respectively. For ama1, sequencing failed for aliquot 6 of control set 1, while for mdr1, sequencing failed for aliquot 2 and 6 of control set 3, aliquots 1 and 6 of control set 4 and aliquots 1 and 5 of control set 5. Under the mdr1 control set 4, the Dd2 copy II (86F, 184Y) was not identified, possibly due to having very low concentrations that were not picked up in this aliquot. Based on our control mixtures, the minimum variant frequency we could detect was 5%.</p> <p><strong>Figure S2. Heatmaps of the successfully PCR amplified and sequenced samples for ama1 (A) and mdr1 (B).</strong> The rows represent the study participants, while the columns represent time in hours. Successfully sequenced samples are shown in blue, those that failed PCR are shown in red and those that failed sequencing are in black. The timepoint “ Rec” represents unscheduled visits where a recurrent sample was collected. The unshaded areas with "-" are time points where samples were not collected. For each time point, the number of samples successfully sequenced (n Successful) is indicated in the last row of each panel. The table in panel C shows the groupings of samples based on parasitemia, high (> 5,000), moderate (100-5,000) and low (< 100 parasites per microlitre). Many samples collected between 0h-12h had high parasitemia, samples collected between 18h–30h had moderate parasitemia, while samples collected after 30h were primarily of low parasitemia.</p> <p><strong>Figure S3. The mean complexity of infection (COI) by AMA1 throughout treatment.</strong> The mean COI (red diamonds) appeared to be stable (between 1.5 - 2) from baseline (0h) up to 72h and thereafter fluctuated due to the small sample sizes (<5) in the post-treatment samples. The black dots represent the COI per sample.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.