Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,574
datasets available to search
ShareScore release 0.9.0
Dataset results
1,574 results for “genome sequencing”
Fig. 4 in Complete mitochondrial genome sequence of Bonasa sewerzowi (Galliformes: Phasianidae) and phylogenetic analysis
Fig. 4. The structure of CR in Bonasa sewerzowi mitochondrial genome and comparasion with B. bonasia.
Fig. 5 in Complete mitochondrial genome sequence of Bonasa sewerzowi (Galliformes: Phasianidae) and phylogenetic analysis
Fig. 5. Nucleotide composition of different partitions from two Bonasa mitogenomes. AT-skew, (A-T)/(A+T); GC-skew, (G-C)/(G+C); PCG-1st, the first codon positions of PCGs; PCG-2nd, the second codon positions of PCGs; PCG-3rd, the third codon positions of PCGs.
Fig. 3 in Complete mitochondrial genome sequence of Bonasa sewerzowi (Galliformes: Phasianidae) and phylogenetic analysis
Fig. 3. The srRNA secondary structure of Bonasa sewerzowi mitogenome and comparasion with B. bonasia. The different nucleotides in B. bonasia was pointed out.
Fig. 1 in Complete mitochondrial genome sequence of Bonasa sewerzowi (Galliformes: Phasianidae) and phylogenetic analysis
Fig. 1. Gene map of the B. sewerzowi mitochondrial genome. Transfer RNA genes are designated by single-letter amino acid codes. L1, L2, S1, and S2 denote trnL (uur), trnL (cun), trnS (ucn) and trnS (agy), respectively.
Fig. 2 in Complete mitochondrial genome sequence of Bonasa sewerzowi (Galliformes: Phasianidae) and phylogenetic analysis
Fig. 2. The lrRNA secondary structure of Bonasa sewerzowi mitogenome and comparasion with B. bonasia. The different nucleotides in B.
Genome annotations of Drosophila melanogaster and Drosophila simulans wild-type strains from long read sequencing assemblies
<p>Genome assemblies were performed for eight wild-type strains of Drosophila melanogaster and Drosophila simulans from Oxford Nanopore long read sequencing (please refer to Mohamed et al. Cells 2020 (doi:10.3390/cells9081776)). Assemblies were deposited in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB50024 (<a href="https://www.ebi.ac.uk/ena/browser/view/PRJEBxxxx">https://www.ebi.ac.uk/ena/browser/view/</a>PRJEB50024).</p> <p>Transposable Element annotations: we used RepeatMasker 4.1.0 (<a href="http://repeatmasker.org/">http://repeatmasker.org/</a>) -species Drosophila, followed by OneCodeToFindThemAll (Bailly-Bechet et al. 2014) with default parameters.</p> <p>Gene annotations: We retrieved gtf files from FlyBase : <a>ftp.flybase.net/genomes/D</a><a>rosophila_melanogaster/dmel_r6,46_FB2022_03/gft/dmel-all-r6.46.gtf.gz</a> and <a>ftp.flybase.net/genomes/Drosophila_simulans/dsim_r2,02_FB2017_04/gtf/dsim-all-</a><a>r2,02.gtf.gz</a>. The corresponding fasta files were also downloaded from FlyBase: <a>ftp.flybase.net/genomes/Drosophila_melanogaster/dmel_r6,46_FB2022_03/</a><a>fasta</a><a>/dmel-all-</a><a>chromosome-</a><a>r6.46.</a><a>fasta</a><a>.gz</a> and <a>ftp.flybase.net/genomes/Drosophila_simulans/dsim_r2,02_FB2017_04/</a><a>fasta</a><a>/dsim-all-</a><a>chromosome-</a><a>r2,02.</a><a>fasta</a><a>.gz</a>. We used Liftoff (Shumate and Salzberg, 2020) to lift over gene annotations from the references to our genome assemblies. We used -flank 0.2 and only kept the “gene” and “exon” terms.</p>
Figure 5 in Genetic difference between two Schistosoma japonicum isolates with contrasting cercarial shedding patterns revealed by whole genome sequencing
Figure 5. Gene ontology enrichment analysis of genes from genome regions with strong selective signals.
Figure 2 in Genetic difference between two Schistosoma japonicum isolates with contrasting cercarial shedding patterns revealed by whole genome sequencing
Figure 2. Box plots of diversity, Tajima's D of S. japonicum and their FST. (A) Nucleotide diversity estimated in 100-kb windows sliding in 10-kb steps throughout the genome. (B) Tajima's D estimated within a nonoverlapping 100-kb window throughout the genome. (C) pairwise FST computed in 100-kb windows sliding in 10-kb steps throughout the genome.
Figure 4 in Genetic difference between two Schistosoma japonicum isolates with contrasting cercarial shedding patterns revealed by whole genome sequencing
Figure 4. Distribution of p ratios (pST/pHX) and FST values calculated in 100-kb windows sliding in 10-kb steps throughout the genome. Data points colored red and blue were identified as selected regions in ST (red dots) and in HX (blue dots), respectively. These points correspond to the 5% left and right tails of the empirical p ratio distribution, where the p ratios are 0.092 and 1.643, respectively (vertical dashed lines), and the 5% right tail of the empirical FST distribution, where FST is 0.965 (horizontal dashed line).
Figure 3 in Genetic difference between two Schistosoma japonicum isolates with contrasting cercarial shedding patterns revealed by whole genome sequencing
Figure 3. Box plots of diversity and relationship of two sample groups. (A) Nucleotide diversity estimated in 100-kb windows sliding in 10-kb steps throughout the genome. (B) Tajima's D estimated within a nonoverlapping 100-kb window throughout the genome. (C) FST between computed in 100-kb windows sliding in 10-kb steps throughout the genome.
Protein haplotype sequences obtained by ProHap from the 1000 Genomes Project data set
<p>Database of protein sequences obtained using ProHap (<a href="https://github.com/ProGenNo/ProHap">https://github.com/ProGenNo/ProHap</a>) on the data set of phased genotypes published by the 1000 Genomes Project, aligned with the GRCh38 genome build (<a href="https://www.internationalgenome.org/data-portal/data-collection/grch38">https://www.internationalgenome.org/data-portal/data-collection/grch38</a>). We used Ensembl v.110 for the mapping of coordinates between genes, exons, and transcripts. The complete configuration file for each ProHap run is attached to this repository.</p> <p>This data set contains six compressed directories, five representing the superpopulations included in the 1000 Genomes Project (<a href="https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project">https://catalog.coriell.org/1/NHGRI/Collections/1000-Genomes-Project-Collection/1000-Genomes-Project</a>), and one created using all the samples included in the 1000 Genomes data set:</p> <ul> <li>AFR - African</li> <li>AMR - American</li> <li>EUR - European</li> <li>SAS - South Asian</li> <li>EAS - East Asian</li> <li>ALL - all participants in the 1000 Genomes Project</li> </ul> <p>Each of the directories contains the following files:</p> <ul> <li>F1: The concatenated fasta file ready to be used with search engines, contains the following: <ul> <li>Protein haplotype sequences obtained by ProHap, using alleles with at least 1 % frequency within the selected population</li> <li>Reference proteome as per Ensembl v. 110</li> <li>Contaminant sequences from the cRAP project (<a href="https://www.thegpm.org/crap/">https://www.thegpm.org/crap/</a>)</li> <li>The file is provided in two formats - full and simplified. The simplified fasta contains only the artificial protein identifier and the matching gene name, and is optimised for compatibility with a wide range of tools. For annotation of peptides using the PeptideAnnotator, please provide the header (F1.2) in addition to the simplified fasta file. </li> </ul> </li> <li>F2: Additional information about the haplotype sequences, to be used for mapping identified peptides to the original haplotypes</li> <li>F3: Translations of haplotype cDNA sequences, before merging with the reference proteome</li> </ul> <p>For further description of the files, please refer to <a href="https://github.com/ProGenNo/ProHap/wiki/Output-files">https://github.com/ProGenNo/ProHap/wiki/Output-files</a>.</p> <p>For the usage of these databases with search engines, and downstream anaylsis of identified peptides, please refer to the project's wiki page: <a href="https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches">https://github.com/ProGenNo/ProHap/wiki/Using-the-database-for-proteomic-searches</a>.</p> <p>When using these databases in your publication, please cite: Vašíček, J., Kuznetsova, K.G., Skiadopoulou, D. <em>et al.</em> ProHap enables human proteomic database generation accounting for population diversity. <em>Nat Methods</em> (2024). <a href="https://doi.org/10.1038/s41592-024-02506-0">https://doi.org/10.1038/s41592-024-02506-0</a></p>
Human intestinal Bacteria Collection (HiBC): Genome sequences
<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the genome sequences of the isolates in the FASTA nucleotide format. Plasmids sequences when present are located at the very end of the file.</p> <p><strong>UPDATE v3</strong>: The genome of one of our isolate had been unfortunately swapped. This mistake has been now corrected on Zenodo and Coscine. The genome of <em>Segatella sinensis</em> CLA-AA-H117 should be considered correct with 103 contigs and 3 671 232 nt. Please note that the genome available at the NCBI is the correct one (GCA_040324585.2). Two typos regarding taxonomy have been corrected as well: <em>Maccoya intestinihominis</em> has been corrected to <em>Maccoyia intestinihominis</em> and <em>Faecousia faecis</em> to <em>Faecousia intestinalis</em>.</p>
Рис. 1. ФиΛогенетические Αеревья хантавируса AMRV и его прироΑного носитеΛя восточноазиатской мыши Apodemus peninsulae Thomas, 1906. А. ФиΛогенетическое Αерево восточноазиатской мыши Apodemus peninsulae, построенное метоΑом «максимаΛьного правΑопоΑобия» (ML) и поΛученное на основе анаΛиза участка гена цитохрома b мтΔНК (744 п.н.). В узΛах ветвΛения указаны бутстреп-поΑΑержки, рассчитанные ΑΛя 1000 повторов. Цветными Λиниями обозначены фиΛогенетические Λинии: Αве Китайские (зеΛеный), Корейская «Korea» (синий), Амурская «Amur» (красный). ПоΛужирным шрифтом выΑеΛены собственные образцы. Названия образцов из GenBank/NCBI быΛи сокращены; B. ФиΛогенетическое Αерево из работы Α. Н. Яшиной с ΑопоΛнениями, построенное метоΑом «бΛижайшего сосеΑа» (NJ) на основе посΛеΑоватеΛьностей фрагмента М-сегмента (2737–2980 н.п.) генома хантавирусов. В узΛах ветвΛения указаны бутстреппоΑΑержки, рассчитанные ΑΛя 1000 повторов. Жирным выΑеΛены иссΛеΑованные РНК изоΛяты (Яшина 2012; Яшина и Αр. 2019) Fig. 1. Phylogenetic trees of AMRV and its natural reservoir host — the Korean field mouse Apodemus peninsulae Thomas, 1906. A. Phylogenetic tree of the Korean field mouse Apodemus peninsulae constructed by the "maximum likelihood" method (ML). The data are obtained from the analysis of the cytochrome b mtDNA gene fragments (744 bp). Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. Colored lines indicate phylogenetic lines: two Chinese (green), Korea (blue), and Amur (red). Own samples are highlighted in bold. The names of the samples from GenBank/NCBI have been shortened; B. Phylogenetic tree from L. N. Yashina's work with additions constructed by the neighbour joining method (NJ). It is based on the sequences of an M-segment fragment (2737–2980 bp) of the hantavirus genome. Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. The researched RNA isolates are highlighted in bold (Yashina 2012; Yashina et al. 2019) in Variability of the gene cyt b in the Korean field mouse Apodemus peninsulae Thomas, 1906 - a reservoir host of AMRV in the Khasansky District of Primorsky Krai
Рис. 1. ФиΛогенетические Αеревья хантавируса AMRV и его прироΑного носитеΛя восточноазиатской мыши Apodemus peninsulae Thomas, 1906. А. ФиΛогенетическое Αерево восточноазиатской мыши Apodemus peninsulae, построенное метоΑом «максимаΛьного правΑопоΑобия» (ML) и поΛученное на основе анаΛиза участка гена цитохрома b мтΔНК (744 п.н.). В узΛах ветвΛения указаны бутстреп-поΑΑержки, рассчитанные ΑΛя 1000 повторов. Цветными Λиниями обозначены фиΛогенетические Λинии: Αве Китайские (зеΛеный), Корейская «Korea» (синий), Амурская «Amur» (красный). ПоΛужирным шрифтом выΑеΛены собственные образцы. Названия образцов из GenBank/NCBI быΛи сокращены; B. ФиΛогенетическое Αерево из работы Α. Н. Яшиной с ΑопоΛнениями, построенное метоΑом «бΛижайшего сосеΑа» (NJ) на основе посΛеΑоватеΛьностей фрагмента М-сегмента (2737–2980 н.п.) генома хантавирусов. В узΛах ветвΛения указаны бутстреппоΑΑержки, рассчитанные ΑΛя 1000 повторов. Жирным выΑеΛены иссΛеΑованные РНК изоΛяты (Яшина 2012; Яшина и Αр. 2019) Fig. 1. Phylogenetic trees of AMRV and its natural reservoir host — the Korean field mouse Apodemus peninsulae Thomas, 1906. A. Phylogenetic tree of the Korean field mouse Apodemus peninsulae constructed by the "maximum likelihood" method (ML). The data are obtained from the analysis of the cytochrome b mtDNA gene fragments (744 bp). Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. Colored lines indicate phylogenetic lines: two Chinese (green), Korea (blue), and Amur (red). Own samples are highlighted in bold. The names of the samples from GenBank/NCBI have been shortened; B. Phylogenetic tree from L. N. Yashina's work with additions constructed by the neighbour joining method (NJ). It is based on the sequences of an M-segment fragment (2737–2980 bp) of the hantavirus genome. Bootstrap supports calculated for 1,000 repeats are indicated in the branching nodes. The researched RNA isolates are highlighted in bold (Yashina 2012; Yashina et al. 2019)
Adapter sequences used for trimming of genomic sequences in the assembly of the Northern Spotted Owl (<i>Strix occidentalis caurina</i>) genome assembly version 1.0
<p>These files provide the sequences of the adapters used in the construction of the genomic libraries Hanna et al. (2017a) sequenced and used to assemble the Northern Spotted Owl (<em>Strix occidentalis caurina</em>) genome assembly version 1.0 (Hanna et al. 2017b). These files also contain relevant supplemental adapter sequences from the adapter files included with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011595_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011595. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" and "NexteraPE-PE.fa" files distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011596_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011596. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" and "NexteraPE-PE.fa" files distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011597_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011597. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" and "NexteraPE-PE.fa" files distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011614_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011614. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" and "NexteraPE-PE.fa" files distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011615_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011615. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" file distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p> <p><strong>SRR4011616_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011616. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" file distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).<br> <br> <strong>SRR4011617_adapters.fa</strong> : This FASTA format file contains the full length sequences of the adapters Hanna et al. (2017a) used to construct the genomic library they sequenced to produced the data uploaded as NCBI Sequence Read Archive (SRA) run accession SRR4011617. I have also included the partial adapter sequences provided in the "TruSeq3-PE-2.fa" file distributed with Trimmomatic version 0.36 (Bolger, Lohse & Usadel, 2014).</p>
Draft genome assemblies of killifish from the Fundulus genus with ONT and Illumina sequencing platforms
<p>Four species from the genus Fundulus were selected for genome sequencing to study the physiological and genetic mechanisms that diverge between euryhaline and stenohaline freshwater species within this cyprinodontiform order of ray-finned fishes.</p>
The Caedibacter taeniospiralis genome sequence and annotation
<p>Interest in host-symbiont interactions is continuously increasing, not only due to the relevance and<br> prevalence of microbiomes. Started with the detection and description of novel symbionts, attention<br> moves to the molecular consequences and innovations of symbioses. However, molecular analysis requires<br> genome data which is dicult to obtain from obligate intracellular and uncultivated bacteria.We describe<br> here the Caedibacter taeniospiralis genome and transcriptome, identified by DNA and RNA Dual-Seq<br> of infected paramecia.</p> <p>Data inlcudes a genome FASTA file, annotations of genes and operons in gff format, and the compelte annotation in Geneious format.</p>
Supplemental data for: Evaluation of SARS-CoV-2 response at the University of North Carolina (UNC) at Charlotte using percent positivity data and viral genomic sequence data.
<p>Supplemental data for:</p> <p>Evaluation of SARS-CoV-2 response at the<br>University of North Carolina (UNC) at Charlotte<br>using percent positivity data and viral genomic<br>sequence data.</p> <p>Submitted to Biocarla 2024</p> <p>https://carla2024.org/portfolios/biocarla/</p> <p>Authors:</p> <p>Daniel Janies 1,2,3 [0000−0002−7890−9906], Shirish Yasa 1,2,3 [0000−0003−3217−4921],<br>Colby T. Ford 1,3,4 [0000−0002−7859−3622] Jannatul Ferdous 2,3 [0000−0003−3053−9616],<br>William Taylor 2,3 [0009−0000−6204−1172], April Harris 2,3 [0009−0009−2557−7926],<br>Sam Kunkleman 2,3 [0000−0002−2309−6418], Juan Bolanos 2,3, Kevin Lambirth 2,3 [0000-0002-6568-543X], Denis<br>Jacob Machado 1,2,3 [0000−0001−9858−4515], Cynthia Gibas 1,2,3 [0000−0002−1288−9543],<br>and Jessica Schlueter 1,2,3 [0000−0002−6490−0580]</p> <p>Affiliations:</p> <p>1) Center for Computational Intelligence to Predict Health and Environmental Risks<br>(CIPHER), University of North Carolina at Charlotte 28223, USA<br>Correspondence to: djanies@charlotte.edu<br>https://cipher.charlotte.edu<br>2) Department of Bioinformatics and Genomics, University of North Carolina at<br>Charlotte 28223, USA https://cci.charlotte.edu/departments/<br>department-of-bioinformatics-and-genomics/<br>3) College of Computing and Informatics, University of North Carolina at Charlotte<br>28223, USA https://cci.charlotte.edu<br>4) School of Data Science, University of North Carolina at Charlotte 28223, USA<br>https://sds.charlotte.edu<br>5) Division of Research, University of North Carolina at Charlotte 28223, USA<br>https://research.charlotte.edu/</p> <p> </p>
Supplementary data to the African Yam Bean Whole Genome Sequencing Project ENA Project_ID:PRJEB57813
<p>The first chromosome-scale assembly of the African yam bean, Sphenostylis stenocarpa (Hochst. ex. A. Rich.) Harms, an original African tuberous legume producing both pods and protein-rich tubers.</p> <p>ENA Project_ID: <strong>PRJEB57813</strong></p> <p>Genome Assembly Accession: <strong>GCA_963425845</strong></p>
Code and data associated with Christiansen et al. 2021 "Facilitating population genomics of non-model organisms through optimized experimental design for reduced representation sequencing"
<p>All code and data input and output files (except reference genome and raw sequencing data) needed to reproduce the results of Christiansen et al. 2021 as released on <a href="https://github.com/notothen/radpilot">https://github.com/notothen/radpilot</a> alongside journal publication. See published paper:</p> <p>Christiansen, H., Heindler, F.M., Hellemans, B. <em>et al.</em> Facilitating population genomics of non-model organisms through optimized experimental design for reduced representation sequencing. <em>BMC Genomics</em> <strong>22, </strong>625 (2021). <a href="https://doi.org/10.1186/s12864-021-07917-3">https://doi.org/10.1186/s12864-021-07917-3</a></p>
Data release: Whole-genome sequencing of Schistosoma mansoni reveals extensive diversity with limited selection despite mass drug administration
<p>Source data used in the publication: Berger et al. (2021) - Provisional title: 'Whole-genome sequencing of <em>Schistosoma mansoni</em> reveals extensive diversity with limited selection despite mass drug administration'. These data were used to generate all figures used in the publication and all files are organised and labelled specifically to run with the custom code that uses these data can be found at: http://doi.org/10.5281/zenodo.4975908. </p> <p><br> <strong>File descriptions:</strong></p> <p><strong>SOURCE DATA.zip - All source data for all figures. </strong></p> <p><strong>Figure 1b:</strong></p> <ul> <li>supplementary_data_9.txt - Metadata</li> </ul> <p><strong>Figure 2a&b:</strong></p> <ul> <li>207_PCA.eigenvec - PCA eigenvectors</li> <li>207_PCA.eigenval - PCA eigenvalues</li> </ul> <p><strong>Figure 2c:</strong></p> <ul> <li>autosomes.mdist - PLINK distance matrix used to build the neighbour joining phylogeny</li> </ul> <p><strong>Figure 2d:</strong></p> <ul> <li>all.pi.pixy.schools.txt - Nucleotide diversity results for each school subpopulation.</li> </ul> <p><strong>Figure 2e:</strong></p> <ul> <li>autosomes.dxy.5kb.schools.txt - Autosomal D<sub>XY</sub> results between school subpopulations. </li> <li>autosomes.fst.5kb.schools.txt - Autosomal F<sub>ST</sub> results between school subpopulations.</li> </ul> <p><strong>Figure 2f:</strong></p> <ul> <li>admixture_all.txt - ADMIXTURE results for each sample and population sizes, column 1 represents number of populations (K), columns 3-8 represent admixture values for each population. </li> </ul> <p><strong>Figure 3a, Supplementary figure 10a:</strong></p> <ul> <li>sfs.csv - Site frequency spectra (allelic proportions at each frequency bin) for each school. </li> </ul> <p><strong>Figure 3b:</strong></p> <ul> <li>TD.all.txt - Tajima's D values calculated in 5 kb windows for each school subpopulation. </li> </ul> <p><strong>Figure 4a, Supplementary figures 13-18: </strong></p> <ul> <li>ALL.MAYUGE.IHS.ihs.out.100bins.norm.txt.zip - Normalised iHS scores for the Mayuge district parasite populations (Selscan output).</li> </ul> <p><strong>Figure 4b, Supplementary figures 13-18: </strong></p> <ul> <li>ALL.TORORO.IHS.ihs.out.100bins.norm.txt.zip -<strong> - </strong>Normalised iHS scores for the Tororo district parasite populations (Selscan output).</li> </ul> <p><strong>Figure 4c, Supplementary figures 13-18: </strong></p> <ul> <li>ALL.MAYUGEvsTORORO.xpehh.xpehh.out.norm.txt.zip - - Normalised XP-EHH scores between Mayuge and Tororo parasite populations.</li> </ul> <p><strong>Figure 4d, Supplementary figures 13-18:</strong></p> <ul> <li>MAYUGE_TORORO_2000.windowed.weir.txt.zip - F<sub>ST</sub> values calculated between Mayuge and Tororo populations in 2kb windows. </li> </ul> <p><strong>Figure 4e, Supplementary figures 12a&c:</strong></p> <ul> <li>MAYUGE_PI.windowed.pi.zip - Nucleotide diversity values calculated in 2 kb windows for Mayuge populations. </li> <li>TORORO_PI.windowed.pi.zip - Nucleotide diversity values calculated in 2 kb windows for Kocoge populations (Tororo district).</li> </ul> <p><strong>Figure 5a:</strong></p> <ul> <li>all.pi.treat.fix.txt.zip - Nucleotide diversity results for each treatment subpopulation</li> </ul> <p><strong>Figure 5b</strong></p> <ul> <li>autosomes.dxy.5kb.treatment.txt - <strong> </strong>- Autosomal D<sub>XY</sub> results between clearance phenotype subpopulations. </li> <li>autosomes.fst.5kb.treatment.txt<strong> </strong>- Autosomal F<sub>ST</sub> results between clearance phenotype subpopulations. </li> </ul> <p><strong>Figure 5c:</strong></p> <ul> <li>fst.windows.2kb.treatment.txt.zip - F<sub>ST</sub> values for comparisons between different treatment groups (Pre-treatment, post-treatment (good clearers), post-treatment (poor clearers))</li> </ul> <p><strong>Figure 5d: </strong></p> <ul> <li>assoc_err_binary.txt.zip - Results of binary trait association between miracidia sampled from hosts with good clearance phenotypes (where treatment appeared to be highly effective) and miracidia isolated post-treatment from hosts with poor clearance phenotypes (where miracidia are potentially derived from parasites that survived treatment.</li> </ul> <p><strong>Figure 5e:</strong></p> <ul> <li>assoc_err_linear.txt.zip - - Results of linear regression genome-wide association study with the ERR estimates for all 198 samples, using the mean of the posterior ERR estimates from Crellen et al. (2016) as a quantitative trait.</li> </ul> <p><strong>Supplementary figure 1:</strong></p> <ul> <li>median.coverage.txt - Normalised depth of read coverage (column 4) calculated in 25 kb windows (columns 2&3) across all samples for all chromosomes (column 1).</li> </ul> <p><strong>Supplementary figure 2a-f: </strong></p> <ul> <li>cohort.genotyped.txt.zip - <strong> </strong>- Variant quality site values (used to inform variant site retention or removal). </li> </ul> <p><strong>Supplementary figure 2g:</strong></p> <ul> <li>hard_filtered.imiss.txt - Per sample variant missingness (used to inform quality control).</li> </ul> <p><strong>Supplementary figure 2h:</strong></p> <ul> <li>hard_filtered_filtindv.lmiss.txt.zip - Per site missingness (used to inform quality control).</li> </ul> <p><strong>Supplementary figure 3a, 4a, 4b:</strong></p> <ul> <li>prunedData.eigenvec - PCA eigenvectors</li> <li>prunedData.eigenval - PCA eigenvalues</li> </ul> <p><strong>Supplementary figure 3b:</strong></p> <ul> <li>pruned_data.mdist.csv - Distance matrix used as the basis for the neighbour joining phylogeny.</li> </ul> <p><strong>Supplementary figure 5:</strong></p> <ul> <li>cv_scores.txt - ADMIXTURE coefficient of variation scores (column 2) for each population size (1).</li> </ul> <p><strong>Supplementary figure 6:</strong></p> <ul> <li>*_SMC_SE.csv - SMC++ results (from 25 subsampled replicates) for each school subpopulation and outgroup samples. </li> </ul> <p><strong>Supplementary Figure 7:</strong></p> <ul> <li>smcpp.csv - SMC++ results for each school subpopulation and outgroup samples. </li> </ul> <p><strong>Supplementary Figure 8a-d</strong></p> <ul> <li>pi.per_host.txt.zip - Nucleotide diversity values for each host infrapopulation. </li> </ul> <p><strong>Supplementary Figure 9:</strong></p> <ul> <li>sexing.csv - inferred sex (based on differential read coverage over pseudoautosomal and Z-specific regions of the Z chromosome). </li> </ul> <p><strong>Supplementary Figure 10b:</strong></p> <ul> <li>sfs_res.csv - residuals for the SFS analysis in 3a/10a.</li> </ul> <p><strong>Supplementary Figure 11:</strong></p> <ul> <li>MAYUGE_TAJIMA_D.Tajima.D.2kb.txt.zip - Tajima's D values calculated for the Mayuge population in 2kb windows. </li> <li>Tororo_TAJIMA_D.Tajima.D.2kb.txt.zip - Tajima's D values calculated for the Tororo population in 2kb windows. </li> </ul> <p><strong>Supplementary Figures 13-18:</strong></p> <ul> <li>genes.bed - Coordinates of gene models (<em>S. mansoni </em>v7 annotation).</li> <li>KOCOGE_SITE_PI.sites.pi.txt.zip - Per site nucleotide diversity values</li> <li>MAYUGE_TORORO_sites.weir.fst.txt.zip - Per site F<sub>ST</sub> values between Mayuge and Tororo populations. </li> <li>coverage_5kb.windows.txt.zip - Per sample depth of read coverage in 5 kb windows. Columns 4,5,6 represent the median, mean and sstev of coverage for each 5kb window (columns 2&3) along each chromosome (column 1). </li> <li>median.sample.coverage.txt - Median chromosomal depth of read coverage for each sample. </li> </ul> <p><strong>Supplementary Figure 19:</strong></p> <ul> <li>kocoge_median.ld.txt.zip - <strong> </strong>- The decay of linkage disequilibrium with genomic distance between all sites within 50 kb for the Kocoge parasite samples. Chromosomes are shown in column 1, distance in column 2, median values in column 3. </li> <li>mayuge_median.ld.txt.zip - The decay of linkage disequilibrium with genomic distance between all sites within 50 kb for the Mayuge parasite samples. Chromosomes are shown in column 1, distance in column 2, median values in column 3. </li> </ul> <p><strong>Misc files:</strong></p> <p>schools.list - List of samples and schools where they were sampled. </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.