Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

704

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

704 results for “nucleotides”

Learn how ShareScore rates datasets ↗
zenodo36/100

A genomic data set of single‐nucleotide polymorphisms (SNPs) generated by ddRAD tag sequencing in Q. petraea (Matt.) Liebl. populations from Central-Eastern Europe and Balkan Peninsula

<p>This genomic dataset provides highly variable single-nucleotide polymorphism&nbsp;(SNP) markers from georeferenced natural <em>Quercus petraea</em> (Matt.) Liebl. populations collected in Bulgaria, Hungary, Romania, Serbia, Bosnia and Herzegovina, Kosovo and Albania. These SNP loci can be used to assess genetic diversity, differentiation, population structure, and can also be used to detect signatures of selection and local adaptation.</p>

opencc-by-4.0Jun 2020View details →
dryad36/100

Data from: Structure, gene order, and nucleotide composition of mitochondrial genomes in parasitic lice from Amblycera

<p>Parasitic lice have unique mitochondrial (mt) genomes characterized by rearranged gene orders, variable genome structures, and less AT content compared to most other insects. However, relatively little is known about the mt genomes of Amblycera, the suborder sister to all other parasitic lice. Comparing among nine different genera (including representative of all seven families), we show that Amblycera have variable and highly rearranged mt genomes. Some genera have fragmented genomes that vary considerably in length, whereas others have a single mt chromosome. Notably, these genomes are more AT-biased than most other lice. We also recover genus-level phylogenetic relationships among Amblycera that are consistent with those reported from large nuclear datasets, indicating that mt sequences are reliable for reconstructing evolutionary relationships in Amblycera. However, gene order data cannot reliably recover these same relationships. Overall, our results suggest that the mt genomes of lice, already know to be distinctive, are even more variable than previously thought.</p>

opencc-zeroNov 2020View details →
dryad36/100

Raw nucleotide counts for clinal Misty Lake and stream stickleback samples

<div> <p class="x_MsoNormal"><span>How ecological divergence causes strong reproductive isolation between populations in close geographic contact remains poorly understood at the genomic level. We here study this question in a stickleback fish population pair adapted to contiguous, ecologically different lake and stream habitats. Clinal whole-genome sequence data reveal numerous genome regions (nearly) fixed for alternative alleles over a distance of just a few hundred meters. This strong polygenic adaptive divergence must constitute a genome-wide barrier to gene flow because a steep cline in allele frequencies is observed across the entire genome, and because the cline center co-localizes with the habitat transition. Simulations confirm that such strong divergence can be maintained by polygenic selection despite high dispersal and small per-locus selection coefficients. Finally, comparing samples from near the habitat transition before and after an unusual ecological perturbation demonstrates the fragility of the balance between gene flow and selection. Overall, our study highlights the efficacy of divergent selection in maintaining reproductive isolation without physical isolation, and the analytical power of studying speciation at a fine eco-geographic and genomic scale.</span></p> </div> <p> </p>

opencc-zeroDec 2020View details →
dryad36/100

Data from: Discovery and characterization of single nucleotide polymorphisms in Chinook salmon, Oncorhynchus tshawytscha

Molecular population genetics of non-model organisms has been dominated by the use of microsatellite loci over the last two decades. The availability of extensive genomic resources for many species is contributing to a transition to the use of single nucleotide polymorphisms (SNPs) for the study of many natural populations. Here we describe the discovery of a large number of SNPs in Chinook salmon, one of the world's most important fishery species, through large-scale Sanger sequencing of expressed sequence tag (EST) regions. More than 3MB of sequence was collected in a survey of variation in more than 131KB of unique genic regions, from more than 225 separate ESTs, in a diverse ascertainment panel of 24 salmon. This survey yielded 117 TaqMan (5' nuclease) assays, almost all from separate EST regions, which were validated in population samples from 5 major stocks of salmon from the three largest basins on the Pacific coast of the coterminous United States: the Sacramento, Klamath and Columbia Rivers. The proportion of these loci that was variable in each of these stocks ranged from 86.3 to 90.6% and the mean minor allele frequency ranged from 0.194 to 0.236. There was substantial differentiation between populations with these markers, with a mean FST estimate of 0.107, and values for individual loci ranging from 0 to 0.592. This substantial polymorphism and population-specific differentiation indicates that these markers will be broadly useful, including for both pedigree reconstruction and genetic stock identification applications.

opencc-zeroDec 2009View details →
dryad36/100

Data from: Discovery and characterization of single nucleotide polymorphisms in two anadromous alosine fishes of conservation concern

Freshwater habitat alteration and marine fisheries can affect anadromous fish species, and populations fluctuating in size elicit conservation concern and coordinated management. We describe the development and characterization of two sets of 96 single nucleotide polymorphism (SNP) assays for two species of anadromous alosine fishes, alewife and blueback herring (collectively known as river herring), that are native to the Atlantic coast of North America. We used data from high-throughput DNA sequencing to discover SNPs and then developed molecular genetic assays for genotyping sets of 96 individual loci in each species. The two sets of assays were validated with multiple populations that encompass both the geographic range and the known regional genetic stocks of both species. The SNP panels developed herein accurately resolved the genetic stock structure for alewife and blueback herring that was previously identified using microsatellites and assigned individuals to regional stock of origin with high accuracy. These genetic markers, which generate data that are easily shared and combined, will greatly facilitate ongoing conservation and management of river herring including genetic assignment of marine caught individuals to stock of origin.

opencc-zeroDec 2016View details →
dryad36/100

Data from: Evaluation of a single nucleotide polymorphism baseline for genetic stock identification of Chinook Salmon (Oncorhynchus tshawytscha) in the California Current Large Marine Ecosystem

Chinook Salmon is an economically and ecologically important species, and populations from the west coast of North America are a major component of fisheries in the North Pacific Ocean. The anadromous life history strategy of this species generates populations (or stocks) that typically are differentiated from neighboring populations. In many cases, it is desirable to discern the stock of origin of an individual fish or the stock composition of a mixed sample to monitor the stock-specific effects of anthropogenic impacts and alter management strategies accordingly. Genetic stock identification (GSI) provides such discrimination, and we describe here a novel GSI baseline composed of genotypes from more than 8000 individual fish from 69 distinct populations at 96 single nucleotide polymorphism (SNP) loci. The populations included in this baseline represent the likely sources for more than 99% of the salmon encountered in ocean fisheries of California and Oregon. This new genetic baseline permits GSI with the use of rapid and cost-effective SNP genotyping, and power analyses indicate that it provides very accurate identification of important stocks of Chinook Salmon. In an ocean fishery sample, GSI assignments of more than 1000 fish, with our baseline, were highly concordant (98.95%) at the reporting unit level with information from the physical tags recovered from the same fish. This SNP baseline represents an important advance in the technologies available to managers and researchers of this species.

opencc-zeroDec 2013View details →
zenodo36/100

Drosophila simulans VCF: The set of single nucleotide polymorphisms and insertion/deletions in a population of 170 Drosophila simulans lines.

<p>Heritable phenotypic variation in natural populations exceeds the levels predicted under mutation-selection balance where purifying selection removes variation. Balancing selection, inefficient or weak selection, polygenic adaptation, and non-equilibrium populations are all possible explanations for excess variation. Yet, available genomic data indicate an abundance of directional selection. One potential explanation is that fleeting directional selection drives beneficial mutations to high frequency in rapid waves resulting in many intermediate frequency haplotypes. This hypothesis is supported by the genomic data from a panel of 170 D. simulans genotypes established from a single stable population which show evidence for an abundance of incomplete soft sweeps. Demography, admixture, and balancing selection cannot entirely explain the patterns in these data, while transient selective sweeps can account for all the patterns of variation observed in this population. One interpretation is that constant environmental shifts rapidly change the optimal phenotype within Drosophila populations, leaving a signature of adaptive responses.</p>

opencc-by-4.0Sep 2016View details →
zenodo36/100

Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 2. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 4. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using COI and 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

Figure 3. - Phylogenetic relationships among Dicronocephalus species reconstructed with Bayesian inference using 16S rRNA sequences. Numbers above branches indicate ML bootstrap values and Bayesian posterior probabilities. Numbers below branches are bootstrap, symmetric resampling, and jacknife support from parsimony searches, respectively. Scale bar represents 10% nucleotide mutation rate.

opencc-by-4.0Feb 2017View details →
zenodo36/100

Dissection of core promoter syntax through single nucleotide resolution modeling of transcription initiation (CLIPNET data)

<div>This contains data necessary to reproduce the figures in the CLIPNET paper (preprint <a href="https://www.biorxiv.org/content/10.1101/2024.03.13.583868">here</a>) as well as processed data used to train and evaluate CLIPNET. To preserve subdirectory structure, we've packaged the data into tar archives. Please refer to the README documents in our manuscript GitHub repo for more details on file contents:&nbsp;<a href="https://github.com/Danko-Lab/clipnet_paper/">https://github.com/Danko-Lab/clipnet_paper/</a></div> <div>&nbsp;</div> <div>Pretrained CLIPNET models are archived separately at <a href="../doi/10.5281/zenodo.10408622">DOI 10.5281/zenodo.10408622</a></div> <div>&nbsp;</div> <div>V5: Fixed bug in calculation of profile attribution scores causing them to be off by a factor of exactly 500. Genome-wide DeepSHAP tracks &amp; TF-MoDISco tracks have been accordingly updated. I have not updated the individual examples, as these can be quickly fixed by simply multiplying by 500 when plotting. Additionally, I have uploaded profile and quantity motif calls, which contain genome-wide seqlet annotations. The columns in these files are [chrom, start, end, peak_idx, motif_annotation].</div> <div>V4: Uploaded individual bigWigs. These have been lifted over using CrossMap from the original hg19 (GSE110638) to hg38 and RPM normalized.</div> <div>V3: Final version prior to journal submission. Don't recall exact details of what's changed.</div> <div>V2: evaluation_metrics.tar.gz and evaluation_data.tar.gz have been replaced. Previously, we benchmarked the models by treating each peak in each individual as a separate data point. Here, we instead predicted from the reference genome and compared against the averaged bigWigs.</div>

openmit-licenseJan 2024View details →
zenodo36/100

The nucleotides absent in genes of SARS-CoV-2 non-canonical subgenomic RNAs generate new Programmed -1 Ribosomal Frameshifting

<p>The data correspond to the article entitled:&nbsp;"dNTPs and adjuvant reagent solutions in 3&rsquo; RACE improve the characterization of noncanonical RNA SARS-CoV-2 genomes"</p> <p>R1. RACE 3&rsquo; Primer Blast Alignment. Contains BLAST alignments against the GenBank database using the consensus nucleotide sequence from the 3&rsquo; end of the SARS-CoV-2 genome and the polylinker. In addition, an illustration of the restriction enzyme pattern of the 3' RACE primer RV30AkCOVID19 and its synthesis by MALDI-TOF is included. The red box indicates the nucleotide sequence of the polylinker and the yellow box represents the 3' RACE primer along with the result of primer synthesis and purification.</p> <p>Graphic representation of the procedure for SARS-CoV-2 genome cDNA synthesis and design of the 3&rsquo; RACE RV30AkCOVID19 primer. The rectangle with vertical lines and the dots represents the 3&rsquo; RACE RV30AkCOVID19 primer and the polylinker, respectively, in the region complementary to the 3&rsquo; UTR end. The arrow represents the reverse transcriptase during complementary strand synthesis. The scissors represent RNases used in purification. The black spheres and magnets indicate the purification process using magnetism.</p> <p>R2. Reads and assembles SARS-CoV-2 genomes.</p> <p>The folder "1) Reads - Ion torrent" contains the reads obtained from sequencing via Ion Torrent technology and the reagents used in this study.</p> <p>The folder named "2) FastQC" contains the results of Ion Torrent sequencing. In the file name, the number indicates the sample, and the letters "RNA" indicate the sequencing according to the IonTorrent protocol. The cDNA synthesis procedures for this study correspond to the following nomenclature: dNTPs-R = dNTPs SARS-CoV-2 solution, DES-R = denaturation reagent, and COM PRO = commercial procedure.</p> <p>The folders named "3) IRMA" and "4) Bowtie2" contain the assemblies of the genomes.</p> <p>Regions and/or codons with loss of genomes 07dN120320 and 27sT122620.</p> <p>Mutations and amino acid substitutions of the SARS-CoV-2 genomes.</p> <p>In addition, an Excel document with the nucleotide ratios of each characterized genome is included from SARS-CoV-2.</p> <p>R3. BLAST alignment of assembled SARS-CoV-2 genomes. Contains two folders named "BLAST - IRMA" and "BLAST - Bowtie2," which contain plain text documents with the results of the BLAST alignment for the genomes obtained with each of the assemblies.</p> <p>R4. Pangolin v1.16 and Nextclade v2.9.1 lineages for SARS-CoV-2 genomes. Contains the folders "Pangolin and Nextclade (Bowtie2)" and "Pangolin and Nextclade (IRMA)." Each folder shows the data obtained with the Pangolin v1.16 and Nextclade v2.9.1 software for the classification of the genomes reported in this study, which were assembled with the IRMA and Bowtie2 software.</p> <p>R5. Reference genome alignment and assembled genomes. Contains the folders "1) IRMA genomes," "2) Bowtie2 genomes," and "3) Genomes 07dN120320 and 27St122620." The files show the sequences and alignments of the examined genomes (the file name indicates the analyzed genome) relative to the SARS-CoV-2 reference genome both in FASTA and Clustal W formats.</p> <p>R6. Programmed &minus;1 Ribosomal Frameshifting Structure. The folder "1) Gibbs free energy 2D" contains a plain text document indicating the secondary structures of the open reading frame stimulation element in dot-bracket format. The folder "2) modeling Data Modeling 3D" contains the information for generating the structure of folder 1 in 3D.</p> <p>R7. SARS-CoV-2 Database.</p> <p>1) GISAID_sequences.zip contains a Zip file that contains a folder named GISAID, which in turn contains plain text documents with the genomes of each variant indicated in the filename of each document.</p> <p>2) The depuration of sequences_GISAID contains two subfolders. The first subfolder, named "1) SARS-CoV-2 complete genome" contains plain text documents with the genomes downloaded from GISAID without undetermined nucleotides. The file name of each document corresponds to the analyzed variant. The subfolder "2) SARS-CoV-2 eliminate genome" contains the sequences eliminated from subfolder 1 because they differed from the majority of the analyzed sequences.</p> <p>3) SARS-CoV-2 consensus variants. Contains plain text documents with consensus sequences for each variant, with frequency thresholds of 20 and 100 indicated in the file name of each document.</p> <p>4) SARS-CoV-2 alignment consensus variants. Contains two subfolders, with the number indicating the alignment frequency threshold. The "Alignment 20_" subfolder contains four documents named "with Ns," which correspond to fasta and Clustal formats with undetermined nucleotides, whereas the files named "without" do not have undetermined nucleotides. The "100_" folder has the same file pattern as the previous folder.</p> <p>5) SARS-CoV-2 codons alignment consensus variants and nc-sgRNA. Contains a document with the alignment of the genomes characterized in this study with the reference genome of SARS-CoV-2. A subfolder named &ldquo;SARS-CoV-2 codons nc-sgRNA&rdquo; shows each of the nc-sgRNA obtained in this study with the reference genome, and the file name corresponds to the nc-sgRNAs. The subfolder &ldquo;SARS-CoV-2 Geneious Prime&rdquo; contains 4 documents. Each document includes the graphical representation of the alignment of the nc-sgRNA obtained with each treatment for the synthesis of SARS-CoV-2 cDNA with respect to the reference genome. The following three documents indicated with the numbers 25, 50, and 100 correspond to the percentage of identity with respect to the number of annotations relative to the reference genome, which is indicated in the title of each document.</p> <p>6) Variant Alignment &ndash; Ns. Contains eight documents corresponding to the fasta and clustal formats with SARS-CoV-2 genomes obtained in this study from the reference genome and from genomes containing undetermined nucleotides of the Gamma, Lambda, Mu and Omicron variants.</p> <p>R8. Phylogeny SARS-CoV-2. Contains two subfolders with the results of the phylogenetic analyses conducted via the maximum likelihood method of the genomes characterized in this study compared to the variants. The subfolder named "Phylogeny with Ns" indicates the analysis of genomes containing undetermined nucleotides, whereas "Phylogeny without Ns" corresponds to the analysis of complete genomes.</p>

opencc-by-4.0Jun 2023View details →
dryad36/100

Data from: Distances and their visualization in studies of spatial-temporal genetic variation using single nucleotide polymorphisms (SNPs)

<p>Distance measures are widely used for examining genetic structure in datasets that comprise many individuals scored for a very large number of attributes. Genotype datasets composed of single nucleotide polymorphisms (SNPs) typically contain bi-allelic scores for tens of thousands if not hundreds of thousands of loci.</p> <p>We examine the application of distance measures to SNP genotypes and sequence tag presence-absences (SilicoDArT) and use real datasets and simulated data to illustrate pitfalls in the application of genetic distances and their visualization.</p> <p>The datasets used to illustrate points in the associated review are provided here together with the R script used to analyse the data. Data are either simulated internal to this script or are SNP data generated as part of other studies and included as compressed binary files readily accessable by reading into R using R base function readRDS(). Refer to the analysis script for examples.</p>

opencc-zeroJan 2024View details →
dryad36/100

Data from: Self-amplifying RNA generated with the modified nucleotides 5-methylcytidine and 5-methyluridine mediate strong expression and immunogenicity in vivo

<p>When utilized in therapeutic applications, synthetic self-amplifying RNA can lead to higher and more sustained expression than standard messenger RNA. This feature is particularly important for gene replacement therapy applications where prolonged expression could reduce the dose and frequency of treatments. The inclusion of modified nucleotides in synthetic non-amplifying mRNA has been shown to increase RNA stability, reduce immune activation and enhance gene expression. Preclinical and clinical studies with self-amplifying RNA (saRNA) have so far exclusively relied on RNA containing the canonical nucleotides adenosine, cytidine, guanosine and uridine. For the first time, we show that non-canonical nucleotides, such as m5C and m5U, are sufficiently compatible with a replicon derived from Venezuelan equine encephalitis alphavirus mediating protein translation <em>in vitro</em>, while those containing m1ψ in place of uridine show no detectable expression. When administered <em>in vivo</em>, saRNA generated with m5C or m5U mediate sustained gene expression of the luciferase reporter gene with those incorporating m5U appearing to lead to more prolonged expression. Finally, distinct antigen-specific humoral and cellular immune responses were induced by modified saRNA encoding the model antigen ovalbumin. The use of modified nucleotides with saRNA-based platforms could enhance their potential to be used effectively in a variety of applications.</p>

opencc-zeroApr 2024View details →
zenodo36/100

Crystal structure of the tandem kinase & triphosphate tunnel metalloenzyme domain module of the TTM1 protein from Arabidoposis thaliana in complex with an adenosine nucleotide analog.

<p>bzip2ed tar archive containing the diffraction images (Pilatus 2M-F detector, SLS beamline PXIII, collected on 19.12.2016) and the associated data processing files (xds)&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Consensus nucleotide sequences for env and gag for paper: Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models

<p>This is the consensus sequence repository to the manuscript &quot;Insights to HIV-1 coreceptor usage by estimating HLA adaptation with Bayesian generalized linear mixed models&quot;.<br> It contains the 10% consensus nucleotide sequences of the env and gag (only p24) protein of HIV-1 used for the training and leftout data set. The NGS sequences are available under BioProject ID PRJNA810303 and the corresponding BioSample Accession IDs are SAMN26241863:26242168 and SAMN28728524:SAMN28728529</p> <ul> <li>env_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the leftout data set</li> </ul> </li> <li>env_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the env protein for the training data set</li> </ul> </li> <li>gag_leftout.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the leftout data set</li> </ul> </li> <li>gag_nt_274.fasta <ul> <li>A fasta file that contains the consensus nucleotide sequences for the gag protein for the training data set</li> </ul> </li> </ul>

openJun 2023View details →
zenodo36/100

Native mass spectrometry and structural studies reveal modulation of MsbA-nucleotide interactions by lipids

<p>The native MS data for paper <strong>"Native mass spectrometry and structural studies reveal modulation of MsbA-nucleotide interactions by lipids"</strong></p>

opencc-by-4.0Apr 2024View details →
dryad36/100

Genome-wide single nucleotide polymorphisms reveal the genetic diversity and population structure of Creole goats from northern Peru

<p>Goat farming constitutes a significant source of income for farmers in northern Peru. There is currently an absence of information about the genetics of Peruvian Creole goats that would enable us to understand their origins and genetic spread. The objective of this study was to estimate the genetic diversity of Creole goats from northern Peru using SNP markers. This study involved the collection of 192 male Creole goats from three key goat production regions in northern Peru. These goat samples were genotyped using the GGPGoat70k SNP panel. To explore the genetic influence of other breeds on Peruvian Creole goats, our dataset was combined with previously published SNP genotypes. External data set includes multiple breeds genotypes sampled from Argentina, Brazil, Spain, and Alpine breed from Italy, France, and Switzerland. After quality control 52,832 autosomal SNPs were used to assess genetic diversity in the Peruvian goats. For the population structure analysis of the merged data 20,513 common SNPs were used. Estimations for expected heterozygosity (H<sub>e</sub>), observed heterozygosity (H<sub>o</sub>), and inbreeding coefficient (F<sub>IS</sub>) were computed for the Peruvian groups. AMOVA, principal component analysis and ADMIXTURE were conducted to evaluate the population structure in the two data sets, Peru and merged. The results revealed a considerable genetic diversity, with H<sub>o</sub> values ranging from 0.40 to 0.41 for the Peruvian sampling groups, and inbreeding coefficient was notably low for Peruvian goat. The population structure analysis demonstrated a distinction (p&lt; 0.05) from other breeds. These findings suggest a level of genetic differentiation of the Peruvian goat population among other breeds, although further research is needed considering samples from other Peruvian areas. We expect this study will contribute to define genetic management strategies to prevent the loss of genetic diversity in Peruvian goat populations and for upcoming advancements in this field.</p>

opencc-zeroMay 2024View details →
dryad36/100

Selection pressure analysis of dengue virus complete genome and E gene nucleotide sequences from Pakistan

<p>This dataset comprises 43 E gene and 44 complete genome nucleotide sequences of the dengue virus from serotypes DENV-1 to DENV-4, representing all documented sequences in Pakistan to date, sourced from the Virus Pathogen Resource (ViPR) database and NCBI. The E gene is critical as it is involved in serotype changes of the dengue virus, making it a pivotal target for understanding shifts in viral pathogenicity and immune escape mechanisms. The aim of compiling this dataset is to facilitate comprehensive genetic analysis and enhance understanding of the evolutionary dynamics of the dengue virus within the region. To assess the evolutionary pressures acting on these sequences, we conducted a selection pressure analysis utilizing computational methods. These methods include the Single Likelihood Ancestor Counting (SLAC), Fixed Effects Likelihood (FEL), adaptive Branch Site Random Effects Likelihood (aBSREL), Mixed Effects Model of Evolution (MEME), and the Genetic Algorithm for Recombination Detection (GARD), all implemented in the HyPhy software package. Our analysis focused on identifying genomic sites under both positive and negative selection pressures, providing insights into the adaptive evolutionary processes affecting the E gene of the dengue virus in Pakistan. Understanding the molecular evolution of this gene is crucial for predicting serotype evolution, potentially aiding in the development of effective vaccines and therapeutic strategies.</p>

opencc-zeroMay 2024View details →
zenodo36/100

Nucleotide sequence database of Copper-containing membrane monooxygenases genes for analysing primer pairs targeting the ammonia monooxygenase subunit A gene of complete ammonia oxidising Nitrospira

<p>Nucleotide sequences of 487 Cu-mmo genes, including amoA comammox clade A and clade B, amoA ammonia oxidizing bacteria as well as other Cu-mmo genes.</p>

opencc-by-4.0Dec 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record