Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,293
datasets available to search
ShareScore release 0.9.0
Dataset results
1,293 results for “gene sequencing”
Figure 2 in Investigation of genetic variation among Turkish populations of Andricus lignicola using mitochondrial cytochrome b gene sequence data
Figure 2. Bayesian analysis tree. Posterior probability values are given on the branches. Outgroup haplotypes: Ac (Andricus caliciformis) and Ak (Andricus kollari).
Figure. Phylogram showing phylogenetic relationships estimated using maximum likelihood analysis of 16S rRNA and COXI gene revealed the grouping of Orthochirus iranus, O. farzanpay, O. stockwelli, O. zagrosensis, O. innesi (JQ514244.1 Morocco), and O. bicolor (KT716038.1 India), with the outgroup species Androctonus crassicauda (FJ217732). in A study of genetic diversity among different population of Orthochirus sp. based on cytochrome C oxidase subunit I and 16srRNA sequencing
Figure. Phylogram showing phylogenetic relationships estimated using maximum likelihood analysis of 16S rRNA and COXI gene revealed the grouping of Orthochirus iranus, O. farzanpay, O. stockwelli, O. zagrosensis, O. innesi (JQ514244.1 Morocco), and O. bicolor (KT716038.1 India), with the outgroup species Androctonus crassicauda (FJ217732).
Fig. 2 in Diversity of fecal parasitomes of wild carnivores inhabiting Korea, including zoonotic parasites and parasites of their prey animals, as revealed by 18S rRNA gene sequencing
Fig. 2. Relative abundance of all parasite genera detected from fecal samples of wild carnivores in Korea. The relative abundance of each parasite is defined as the ratio of the number of sequence reads assigned to that parasite to the total number of sequence reads assigned to all target parasites.
Fig. 1 in Diversity of fecal parasitomes of wild carnivores inhabiting Korea, including zoonotic parasites and parasites of their prey animals, as revealed by 18S rRNA gene sequencing
Fig. 1. Diversity of fecal parasitomes of wild carnivores in Korea. The results shown are based on the diversity of zero-radius operational taxonomic units (ZOTUs) that were taxonomically assigned to parasites. (a) Comparison of richness and diversity of parasite ZOTUs between host animals estimated by the Chao1 estimator and Shannon index, respectively. (b) Non-metric multidimensional scaling (NMDS) plots showing the structure and membership of parasite ZOTUs represented by the Bray–Curtis dissimilarity and Jaccard index, respectively. In the panel (a), one asterisk (*) and two asterisks (**) represent p <0.05 and p <0.01, respectively, by the post hoc Wilcoxon rank-sum test. The abbreviation "ns" represents no statistical difference.
Nanopore deep sequencing as a tool to characterize and quantify aberrant splicing caused by variants in inherited retinal dystrophy genes
Open the record for dataset details and reuse information.
A scalable CRISPR-Cas9 gene editing system facilitates CRISPR screens in the malaria parasite Plasmodium berghei - sequencing data
<p>This holds raw sequencing data, and extracted sgRNA counts </p>
Machine learning models, and training, validation and test datasets for: "Sequence determinants of human gene regulatory elements"
<p>This record contains the training, test and validation datasets used to train and evaluate the machine learning models in manuscript:</p> <p><strong>Sahu, Biswajyoti, et al. "Sequence determinants of human gene regulatory elements." (2021).</strong></p> <p><br> This record contains also the final hyperparameter-optimized models for each training dataset/task combination described in the manuscript. The README-files provided with the record describe the datasets and models in more detail. The datasets deposited here are derived from the original raw data (GEO accession: GSE180158) as described in the Methods of the manuscript.</p>
Sanger sequencing traces of specific exons of the Sm.TRPM_PZQ gene from schistosome field samples
<p>Praziquantel (PZQ) is the only drug available to treat schistosomiasis, which is caused by schistosome blood flukes. In <em>Schistosoma mansononi</em>, the transient receptor potential (TRP) channel Sm.TRPM<sub>PZQ</sub> is strongly suspected to be the target of PZQ. Our <a href="https://doi.org/10.1101/2021.06.09.447779">genetic analysis of <em>S. mansoni</em> response to PZQ</a> revealed a QTL on its chromosome 3 which contains the <em>Sm.TRPM<sub>PZQ</sub></em> gene, strongly suggesting that <em>Sm.TRPM<sub>PZQ</sub></em> could be responsible for PZQ resistance. Therefore, understanding the natural variation in this gene and identifying potential resistance alleles will be a valuable tool for monitoring mass treatment programs aimed at schistosomiasis elimination.</p> <p>We investigated our schistosome collection to examine mutations present in <em>Sm.TRPM<sub>PZQ</sub></em> in natural schistosome populations. We analyzed exome sequencing data from 259 miracidia, cercariae or adult parasites from 3 African countries (Senegal, Niger, Tanzania), the Middle East (Oman) and South America (Brazil). We were able to sequence 36/41 exons of <em>Sm.TRPM<sub>PZQ</sub></em> from 122/259 parasites on average (s.e. = 18.65). We identified several mutations in critical areas of the channel. However, these mutations were supported by a limited number of reads only and required confirmation by Sanger sequencing.</p> <p>The present dataset corresponds to the sequencing effort done on specific exons which carried the mutations to be confirmed. We generated PCR products which were sequenced on on ABI sequencer using Eurofins Genomics services. The SCF files were then analyzed using PolyPhred (see manuscript for details about PCR conditions and data analysis). The trace files are available in the traces folder. Each filename carries a barcode which corresponds to a combination of sample, exon, and primer. All the combinations and corresponding barcodes are listed in the barcode_list.tsv file.</p> <p>Table header details of the barcode list:</p> <ul> <li><em>Sample</em>: the name of sample. The sample coding is as follows: species.country_patientID. BR: Brazil, SN: Senegal, NE: Niger, TZ: Tanzania, OM: Oman.</li> <li><em>Exon</em>: the exon targeted. The exon number corresponds to the exon number of isoform 5 and not the exon number of the gene.</li> <li><em>Barcode</em>: the barcode provided by Eurofins Genomics.</li> <li><em>Primer</em>: the primer used for sequencing. F: forward, R: reverse.</li> </ul>
Fig. 3 in Phylogenetic relationships of Eurema butterflies from Peninsular Malaysia inferred from CO1 and 28S gene sequences with emphasis on Eurema hecabe
Fig. 3. Maximum Likelihood output phylogram for CO1-28S concatenated analysis showing seven major clades representing the seven Eurema species obtained from this study. Bootstrap scores are shown at the branching points. The tree was rooted with the genus Graphium. The butterfly figures show the comparison of morphology among the species corresponding to their respective clades. Figures of butterflies provided as upperside of the wings (left) and downside of wings (right).
Fig. 1 in Phylogenetic relationships of Eurema butterflies from Peninsular Malaysia inferred from CO1 and 28S gene sequences with emphasis on Eurema hecabe
Fig. 1. The geographical sites where samplings have been conducted in Peninsular Malaysia. N, northern area; E, eastern area; W, western area; S, southern area. The dots indicate the distribution of various sampling sites in this study. Triplet letter represents the site code.
Fig. 2 in Phylogenetic relationships of Eurema butterflies from Peninsular Malaysia inferred from CO1 and 28S gene sequences with emphasis on Eurema hecabe
Fig. 2. Phylogenetic tree of Maximum-Likelihood method showing the comparison of phylogram as inferred from partial sequences of mtDNA CO1 and 28S rDNA genes. The bootstrap scores obtained from 1,000 replicates for ML/MP analyses are shown at the branching point. The trees were rooted with the genus Graphium.
Figure 1 in Phylogenetic analyses suggest that Psammomitra (Ciliophora, Urostylida) should represent an urostylid family, based on small subunit rRNA and alpha-tubulin gene sequence information
Figure 1. Morphology and infraciliature of Psammomitra retractilis (F–J, from Song & Warren, 1996). A, B, F, individuals in extended states to show the typical body shapes. Arrowheads in (A) mark the long, dominant membranelles. C, lateral view of a contracted specimen. D, posterior part, to demonstrate the long dorsal cilia. E, anterior part. Arrowheads indicate the long membranelles, whereas arrows mark the dorsal cilia. G, H, dorsal and lateral views of contracted cells. I, J, ventral and dorsal views to show the infraciliature and macronuclear nodules. Scale bars: A, C, D, F = 40 Mm; E = 30 Mm.
Figure 3 in Phylogenetic analyses suggest that Psammomitra (Ciliophora, Urostylida) should represent an urostylid family, based on small subunit rRNA and alpha-tubulin gene sequence information
Figure 3. Maximum parsimony phylogeny of small subunit rRNA genes. Psammomitra is highlighted in black, and holostichids are enclosed in rectangles. Thick branches and arrows denote position of investigated species. Numbers on branches are values generated from 1000 bootstrap replicates.
Figure 2 in Phylogenetic analyses suggest that Psammomitra (Ciliophora, Urostylida) should represent an urostylid family, based on small subunit rRNA and alpha-tubulin gene sequence information
Figure 2. Phylogenetic tree based on small subunit rRNA sequences showing the position of Psammomitra retractilis, by Bayesian inferences applying the GTR + G + I model. '-' reflects disagreement between a method and the reference Bayesian tree at a given node. The fully supported (1.00/100%/100%) branches are marked with solid circles. Psammomitra is shaded black, and holostichids are enclosed in rectangles. Thick branches and arrows denote position of investigated species. The scale bar corresponds to five substitutions per 100 nucleotide positions. Infraciliature of Oxytricha and Uroleptus (from Foissner et al., 2004), Amphisiella (from Li et al., 2007), Trachelostyla (from Gong et al., 2006), and Holosticha (from Hu & Song, 2001) are also shown.
Figure 4 in Phylogenetic analyses suggest that Psammomitra (Ciliophora, Urostylida) should represent an urostylid family, based on small subunit rRNA and alpha-tubulin gene sequence information
Figure 4. Bayesian trees based on different data sets showing phylogenetic relationships amongst Spirotrichea. '-' reflects disagreement between the maximum likelihood/ maximum parsimony method and the reference Bayesian tree at a given node. The fully supported (1.00/100%/100%) branches are marked with solid circles. Species sequenced in the present study are shown in bold type. The scale bar corresponds to 10/2 substitutions per 100 nucleotide positions. A, phylogenetic analyses inferred from alpha-tubulin gene sequences data set. B, phylogenetic analyses inferred from alpha-tubulin amino acids data set.
Figure 2 in Phylogenetic structure of the Sphaeriinae, a global clade of freshwater bivalve molluscs, inferred from nuclear (ITS-1) and mitochondrial (16S) ribosomal gene sequences
Figure 2. Strict consensus of the 1040 equally most parsimonious trees (L = 445; CI = 0.724; RI = 0.886) obtained from the phylogenetic analysis of sphaeriid nuclear ITS1 rDNA sequences. The inferred evolutionary gain and loss of a ~160 nt fragment are indicated. Two Eupera species, E. cubensis and E. platensis, were designated as outgroups and inferred sequence gaps were considered as missing data. Numbers above the branches represent bootstrap values and numbers below indicate decay index values.
Figure 3 in Phylogenetic structure of the Sphaeriinae, a global clade of freshwater bivalve molluscs, inferred from nuclear (ITS-1) and mitochondrial (16S) ribosomal gene sequences
Figure 3. The single most-parsimonious tree (L = 951; CI = 0.568; RI = 0.793) obtained from the maximum parsimony analysis of combined (16S + ITS1) sequence dataset. Maximum likelihood analysis produced a largely congruent topology (HKY model; Ln likelihood = - 7034.61154) with the only difference being Pisidium dubium sister to Sphaerium/Musculium clade. Taxonomic names are arranged according to suggested sphaeriinid taxonomy in the present study and five major monophyletic lineages are indicated. Two Eupera species, E. cubensis and E. platensis, were designated as outgroups. MP bootstrap values are shown to the left of the slash and decay index values to the right above the branches. Numbers below the branches indicate ML bootstrap values.
Figure 1 in Phylogenetic structure of the Sphaeriinae, a global clade of freshwater bivalve molluscs, inferred from nuclear (ITS-1) and mitochondrial (16S) ribosomal gene sequences
Figure 1. Strict consensus of the four equally most parsimonious trees (L = 526; CI = 0.447; RI = 0.743) obtained from the phylogenetic analysis of sphaeriid mitochondrial 16S rDNA sequences. Two Eupera species, E. cubensis and E. platensis, were designated as outgroups and inferred sequence gaps were considered as missing data. Numbers above the branches represent bootstrap values and numbers below indicate decay index values.
Genome report: Genome sequence of 1S1, a transformable and highly regenerable diploid potato for use as a model for gene editing and genetic engineering
<p>Generation of a genomic resource for a readily transformable diploid potato would provide a resource for high throughput functional analysis in potato. The heterozygous <em>Solanum tuberosum</em> Group Phureja clone 1S1 has a high regeneration rate, self-fertility, desirable tuber traits and is amenable to <em>Agrobacterium</em>-mediated transformation. To create a contiguous genome assembly, a homozygous doubled monoploid of 1S1 (DM1S1) was sequenced using 44 Gbp of long reads generated from Oxford Nanopore Technologies (ONT), yielding a 736 Mb assembly that encoded 31,145 protein-coding genes. The final assembly for DM1S1 represents a nearly complete genic space, shown by the presence of 99.6% (C:99.5%[S:97.8%, D:1.7%],F:0.1%,M:0.4%,n:1614) of the Benchmarking Universal Single Copy Orthologs. Variant analysis with Illumina reads from 1S1 was used to deduce its alternate haplotype using the variant calling tools Strelka2 (v2.9.10), GATK's Haplotypecaller (v4.1.4.1), and Freebayes (v1.3.2). These variants were used to create consensus fasta sequences with the DM1S1 assembly using bcftools (v1.9.64).</p>
Genome sequences and gene annotations for two Ophryocystis lineages
<p>Assembly, annotation, and gene sequences for the <em>Ophryocystis </em>lineages sequenced in "Genome sequence of <em>Ophryocystis elektroscirrha</em>, an apicomplexan parasite of monarch butterflies: cryptic diversity and response to host-sequestered plant chemicals." Each of the two lineages has three associated files: a genome sequence file (.fa), an annotation in .gff3 format, and gene sequences in .fna format. Sequences generated for <em>Ophryocystis elektroscirrha </em>come from direct DNA extraction and sequencing effort and are hosted elsewhere on NCBI as well. The other lineage, prefixed Ophryocystis-elektroscirrha_like, was bioinformatically extracted from the genome of an infected host. As such, we are less confident in its completeness and it is not archived elsewhere. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.