Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
dryad32/100

Data from: Fin whale (Balaenoptera physalus) mitogenomics: a cautionary tale of defining sub-species from mitochondrial sequence monophyly

The advent of massive parallel sequencing technologies has resulted in an increase of studies based upon complete mitochondrial genome DNA sequences that revisit the taxonomic status within and among species. Spatially distinct monophyly in such mitogenomic genealogies, i.e., the sharing of a recent common ancestor among con-specific samples collected in the same region has been viewed as evidence for subspecies. Several recent studies in cetaceans have employed this criterion to suggest subsequent intraspecific taxonomic revisions. We reason that employing intra-specific, spatially distinct monophyly at non-recombining, clonally inherited genomes is an unsatisfactory criterion for defining subspecies based upon theoretical (genetic drift) and practical (sampling effort) arguments. This point was illustrated by a re-analysis of a global mitogenomic assessment of fin whales, Balaenoptera physalus spp., published by Archer et al. (2013), which proposed to further subdivide the Northern Hemisphere fin whale subspecies, B. p. physalus. The proposed revision was based upon the detection of spatially distinct monophyly among North Atlantic and North Pacific fin whales in a genealogy based upon complete mitochondrial genome DNA sequences. The extended analysis conducted in this study (1,676 mitochondrial control region, 162 complete mitochondrial genome DNA sequences and 20 microsatellite loci genotyped in 358 samples) revealed that the apparent monophyly among North Atlantic fin whales reported by Archer et al. (2013) to be due to low sample sizes. In conclusion, defining sub-species from monophyly (i.e., the absence of para- or polyphyly) can lead to erroneous conclusions due to relatively "trivial" aspects, such as sampling. Basic population genetic processes (i.e., genetic drift and migration) also affect the time to the most recent common ancestor and hence the probability that individuals in a sample are monophyletic.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Phylogenetic relationships and timing of diversification in gonorynchiform fishes inferred using nuclear gene DNA sequences (Teleostei: Ostariophysi)

The Gonorynchiformes are the sister lineage of the species-rich Otophysi and provide important insights into the diversification of ostariophysan fishes. Phylogenies of gonorynchiforms inferred using morphological characters and mtDNA gene sequences provide differing resolutions with regard to the sister lineage of all other gonorynchiforms (Chanos vs. Gonorynchus) and support for monophyly of the two miniaturized lineages Cromeria and Grasseichthys. In this study the phylogeny and divergence times of gonorynchiforms are investigated with DNA sequences sampled from nine nuclear genes and a published morphological character matrix. Bayesian phylogenetic analyses reveal substantial congruence among individual gene trees with inferences from eight genes placing Gonorynchus as the sister lineage to all other gonorynchiforms. Seven gene trees resolve Cromeria and Grasseichthys as a clade, supporting previous inferences using morphological characters. Phylogenies resulting from either concatenating the nuclear genes, performing a multispecies coalescent species tree analysis, or combining the morphological and nuclear gene DNA sequences resolve Gonorynchus as the living sister lineage of all other gonorynchiforms, strongly support the monophyly of Cromeria and Grasseichthys, and resolve a clade containing Parakneria, Cromeria, and Grasseichthys. The morphological dataset, which includes 13 gonorynchiform fossil taxa that range in age from Early Cretaceous to Eocene, was analyzed in combination with DNA sequences from the nine nuclear genes and a relaxed molecular clock to estimate times of evolutionary divergence. This "tip dating" strategy accommodates uncertainty in the phylogenetic resolution of fossil taxa that provide calibration information in the relaxed molecular clock analysis. The estimated age of the most recent common ancestor (MRCA) of living gonorynchiforms is slightly older than estimates from previous node dating efforts, but the molecular tip dating estimated ages of Kneriinae (Kneria, Parakneria, Cromeria, and Grasseichthys) and the two paedomorphic lineages, Cromeria and Grasseichthys, are considerably younger.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Batch effects in a multi-year sequencing study: false biological trends due to changes in read lengths

High-throughput sequencing is a powerful tool, but suffers biases and errors that must be accounted for to prevent false biological conclusions. Such errors include batch effects, technical errors only present in subsets of data due to procedural changes within a study. If overlooked and multiple batches of data are combined, spurious biological signals can arise, particularly if batches of data are correlated with biological variables. Batch effects can be minimized through randomisation of sample groups across batches. However, in long-term or multi-year studies where data are added incrementally, full randomisation is impossible and batch effects may be a common feature. Here we present a case study where false signals of selection were detected due to a batch effect in a multi-year study of Alpine ibex (Capra ibex). The batch effect arose because sequencing read length changed over the course of the project and populations were added incrementally to the study, resulting in non-random distributions of populations across read lengths. The differences in read length caused small misalignments in a subset of the data, leading to false variant alleles and thus false SNPs. Pronounced allele frequency differences between populations arose at these SNPs because of the correlation between read length and population. This created highly statistically significant, but biologically spurious, signals of selection and false associations between allele frequencies and the environment. We highlight the risk of batch effects and discuss strategies to reduce the impacts of batch effects in multi-year high-throughput sequencing studies.

opencc-zeroDec 2017View details →
dryad32/100

Data from: RAD sequencing yields a high success rate for westslope cutthroat and rainbow trout species-diagnostic SNP assays

Hybridization with introduced rainbow trout threatens most native westslope cutthroat trout populations. Understanding the genetic effects of hybridization and introgression requires a large set of high-throughput, diagnostic genetic markers to inform conservation and management. Recently, we identified several thousand candidate single nucleotide polymorphism (SNP) markers based on RAD sequencing of 11 westslope cutthroat trout and 13 rainbow trout individuals. Here we used flanking sequence for 56 of these candidate SNP markers to design high-throughput genotyping assays. We validated the assays on a total of 92 individuals from 22 populations and seven hatchery strains. Forty-six assays (82%) amplified consistently and allowed easy identification of westslope cutthroat and rainbow trout alleles as well as heterozygote controls. The 46 SNPs will provide high power for early detection of population admixture and improved identification of hybrid and non-hybridized individuals. This technique shows promise as a very low-cost, reliable, and relatively rapid method for developing and testing SNP markers for non-model organisms with limited genomic resources.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Exploitation of a turbot (Scophthalmus maximus L.) immune-related expressed sequence tag (EST) database for microsatellite screening and validation

In this study, we identified and characterized 160 microsatellite loci from an expressed sequence tag (EST) database generated from immune-related organs of turbot (Scophthalmus maximus). A final set of 83 new polymorphic microsatellites were validated after the analysis of 40 individuals from Atlantic origin including both wild and farmed individuals. The allele number and the expected heterozygosity ranged from 2 to 18 and from 0.021 to 0.951, respectively. Evidences of null alleles at moderate-high frequencies were detected at six loci using population data. None of the analyzed loci showed deviations from Mendelian segregation after analysis of five full-sib families including ~92 individuals/family. The markers are used to consolidate the turbot genetic map and, since they are mostly EST-derived, they will be very useful for comparative genomic studies within flatfishes and with model fish species. Using an in silico approach, we detected significant homologies of microsatellite sequences with the EST databases of the flatfish species with highest genomic resources (Senegalese sole, Atlantic halibut, bastard halibut) at 31% of these turbot markers. The conservation of these microsatellites within Pleuronectiformes will pave the way for anchoring genetic maps of different species and identifying genomic regions related to productive traits.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Genome sequence of dwarf birch (Betula nana) and cross-species RAD markers

New sequencing technologies allow development of genome-wide markers for any genus of ecological interest, including plant genera such as Betula (birch) that have previously proved difficult to study due to widespread polyploidy and hybridisation. We present a de novo reference genome sequence assembly, from 67X short read coverage, of Betula nana (dwarf birch) – a diploid that is the keystone woody species of sub-arctic scrub communities but of conservation concern in Britain. We also present 100bp PstI RAD markers for B. nana and closely related Betula tree species. Assembly of RAD markers in 15 individuals by alignment to the reference B. nana genome yielded 44k-86k RAD loci per individual, whereas de novo RAD assembly yielded 64k-121k loci per individual. Of the loci assembled by the de novo method, 3k homologous loci were found in all 15 individuals studied, and 35k in 10 or more individuals. Matching of RAD loci to RAD locus catalogs from the B. nana individual used for the reference genome, showed similar numbers of matches from both methods of RAD locus assembly but indicated that the de novo RAD assembly method may over-assemble some paralogous loci. In 12 individuals hetero-specific to B. nana 37k-47k RAD loci matched a catalog of RAD loci from the B. nana individual used for the reference genome, whereas 44k-60k RAD loci aligned to the B. nana reference genome itself. We present a preliminary study of allele sharing among species, demonstrating the utility of the data for introgression studies and for the identification of species-specific alleles.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Two new species of Limbodessus diving beetles from New Guinea - short verbal descriptions flanked by online content (digital photography, μCT scans, drawings and DNA sequence data)

Background: To date only one species of Limbodessus diving beetles has been reported from the Island of New Guinea, L. compactus (Clark, 1862), which is widerspread in the Australian region. New information: We describe two new species of microendemic New Guinea Limbodessus and use a compact descriptive format flanked by enriched online content in wiki powered species pages. Limbodessus baliem sp.n. is described from ca. 1,600 m altitude in the Baliem Valley of Papua and Limbodessus alexanderi sp.n. from >3,000 m altitude north of Sugapa, Papua. Based on our analysis, we also transfer three species from other genera to Limbodessus Guignot, 1939, with the following changes: Limbodessus deflectus (Ordish, 1966), new combination; Limbodessus leveri (J. Balfour-Browne, 1944), new combination; and Limbodessus plicatus (Sharp, 1882), new combination.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Genotyping-by-sequencing for Populus population genomics: an assessment of genome sampling patterns and filtering approaches

Continuing advances in nucleotide sequencing technology are inspiring a suite of genomic approaches in studies of natural populations. Researchers are faced with data management and analytical scales that are increasing by orders of magnitude. With such dramatic advances comes a need to understand biases and error rates, which can be propagated and magnified in large-scale data acquisition and processing. Here we assess genomic sampling biases and the effects of various population-level data filtering strategies in a genotyping-by-sequencing (GBS) protocol. We focus on data from two species of Populus, because this genus has a relatively small genome and is emerging as a target for population genomic studies. We estimate the proportions and patterns of genomic sampling by examining the Populus trichocarpa genome (Nisqually-1), and demonstrate a pronounced bias towards coding regions when using the methylation-sensitive ApeKI restriction enzyme in this species. Using population-level data from a closely related species (P. tremuloides), we also investigate various approaches for filtering GBS data to retain high-depth, informative SNPs that can be used for population genetic analyses. We find a data filter that includes the designation of ambiguous alleles resulted in metrics of population structure and Hardy-Weinberg equilibrium that were most consistent with previous studies of the same populations based on other genetic markers. Analyses of the filtered data (27,910 SNPs) also resulted in patterns of heterozygosity and population structure similar to a previous study using microsatellites. Our application demonstrates that technically and analytically simple approaches can readily be developed for population genomics of natural populations.

opencc-zeroDec 2013View details →
dryad32/100

Data from: A high-density exome capture genotype-by-sequencing panel for forestry breeding in Pinus radiata

Development of genome-wide resources for application in genomic selection or genome-wide association studies, in the absences of full reference genomes, present a challenge to the forestry industry, where longer breeding cycles could benefit from the accelerated selection possible through marker-based breeding value predictions. In particular, large conifer megagenomes require a strategy to reduce complexity, whilst ensuring genome-wide coverage is achieved. Using a transcriptome-based reference template, we have successfully developed a high density exome capture genotype-by-sequencing panel for radiata pine (Pinus radiata D.Don), capable of capturing in excess of 80,000 single nucleotide polymorphism (SNP) markers with a minor allele frequency above 0.03 in the population tested. This represents approximately 29,000 gene models from a core set of 48,914 probes. A set of 704 SMP markers capable of pedigree reconstruction and differentiating individual genotypes were tested within two full-sib mapping populations. While as few as 70 markers could reconstruct parentage in almost all cases, the impact of missing genotypes was noticeable in several offspring. Therefore, sets of 60 sets of 110 randomly selected SNP markers were compared for both parentage reconstruction and clone differentiation. The performance in parentage reconstruction showed little variation over 60 iterations. However, there was notable variation in discriminatory power between closely related individuals, indicating a higher density SNP marker panel may be required to elucidate hidden relationships in complex pedigrees.

opencc-zeroOct 2019View details →
dryad32/100

Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates

Restriction-enzyme-based sequencing methods enable the genotyping of thousands of single nucleotide polymorphism (SNP) loci in non-model organisms. However, in contrast to traditional genetic markers, genotyping error rates in SNPs derived from restriction-enzyme-based methods remain largely unknown. Here, we estimated genotyping error rates in SNPs genotyped with double digest RAD sequencing from Mendelian incompatibilities in known mother-offspring dyads of Hoffman's two-toed sloth (Choloepus hoffmanni) across a range of coverage and sequence quality criteria, for both reference-aligned and de novo-assembled datasets. Genotyping error rates were more sensitive to coverage than sequence quality and low coverage yielded high error rates, particularly in de novo-assembled datasets. For example, coverage ≥5 yielded median genotyping error rates of ≥0.03 and ≥0.11 in reference-aligned- and de novo-assembled datasets, respectively. Genotyping error rates declined to ≤0.01 in reference-aligned datasets with a coverage >30, but remained >0.04 in the de novo-assembled datasets. We observed approximately 10- and 13-fold declines in the number of loci sampled in the reference-aligned and de novo-assembled datasets when coverage was increased from >5 to >30 at quality score ≥30, respectively. Finally, we assessed the effects of genotyping coverage on a common population genetic application, parentage assignments, and showed that the proportion of incorrectly assigned maternities was relatively high at low coverage. Overall, our results suggest that the tradeoff between sample size and genotyping error rates be considered prior to building sequencing libraries, reporting genotyping error rates become standard practice, and that effects of genotyping errors on inference be evaluated in restriction-enzyme-based SNP studies.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Targeted next-generation sequencing panels in the diagnosis of Charcot Marie Tooth disease

Objective: To investigate the effectiveness of targeted NGS panels in achieving a molecular diagnosis in CMT and related disorders in a clinical setting Methods: We prospectively enrolled 220 patients from two tertiary referral centres, one in London, UK (n=120) and one in Iowa, US (n=100) in whom a targeted CMT NGS panel had been requested as a diagnostic test. PMP22 duplication/deletion was previously excluded in demyelinating cases. We reviewed the genetic and clinical data upon completion of the diagnostic process. Results: After targeted NGS sequencing a definite molecular diagnosis, defined as a pathogenic or likely pathogenic variant, was reached in 30% of cases (n=67). The diagnostic rate was similar in London (32%) and Iowa (29%). Variants of unknown significance were found in an additional 33% of cases. Mutations in GJB1, MFN2, MPZ accounted for 39% of cases who received genetic confirmation, while the remainder of positive cases had mutations in diverse genes, including SH3TC2, GDAP1, IGHMBP2, LRSAM1, FDG4, GARS and another 12 less common genes. Copy number changes in PMP22, MPZ, MFN2, SH3TC2 and FDG4 were also accurately detected. A definite genetic diagnosis was more likely in cases with an early onset, a positive family history of neuropathy or consanguinity and a demyelinating neuropathy. Conclusions: NGS panels are effective tools in the diagnosis of CMT leading to the genetic confirmation in one third cases negative for PMP22 duplication/deletion, thus highlighting how rarer and previously undiagnosed subtypes represent today a relevant part of the genetic landscape of CMT.

opencc-zeroDec 2019View details →
dryad32/100

Data from: Phylogenetic systematics of subtribe Spiranthinae (Orchidaceae: Orchidoideae: Cranichideae) based on nuclear and plastid DNA sequences of a nearly complete generic sample

Subtribe Spiranthinae is the most species-rich lineage of terrestrial Neotropical orchids, encompassing > 500 species and 40 genera. We conducted maximum parsimony and maximum likelihood phylogenetic analyses of DNA sequence data of plastid matK-trnK and trnL-trnF and nuclear ribosomal ITS sequences for 36 genera and 182 species of Spiranthinae plus appropriate outgroups. The results strongly support monophyly of Spiranthinae (minus Discyphus, Discyphinae and Galeottiella, Galeottiellinae) and five major lineages, namely monospecific Cotylolabium (sister to the remaining Spiranthinae) and the Eurystyles, Pelexia, Spiranthes and Stenorrhynchos clades. Eighteen of the 27 genera of Spiranthinae for which more than one species was included in our analyses are monophyletic. Paraphyly of large genera, such as Cyclopogon and Sarcoglottis, resulted from segregation of particular species or groups of species exhibiting minor modifications of structures directly involved in pollination (e.g. nectary, rostellum and viscidium). Conversely, polyphyly has resulted from convergent evolution of floral attributes in distantly related species (e.g. Mesadenus). Some of the morphological characters used traditionally for generic delimitation and in non-molecular cladistic analyses of Spiranthinae are discussed against the evolutionary framework set by our molecular trees, emphasizing putative synapomorphies and problems derived from inappropriate character coding or incorrect homology assessments. Our ancestral area analysis indicates that Spiranthinae originated in eastern South America, with subsequent migrations and secondary radiations in Mesoamerica and North America, plus a derived migration from the latter region to the Old World (Spiranthes).

opencc-zeroDec 2017View details →
dryad32/100

Data from: Identification and characterization of sex-associated loci in sockeye salmon using genotyping-by-sequencing and comparison with a sex-determining assay based on the sdY gene

Loci that can be used to screen for sex in salmon can provide important information for study of both wild and cultured populations. Here, we tested for associations between sex and genotypes at thousands of loci available from a genotyping-by-sequencing (GBS) dataset to discover sex-associated loci in sockeye salmon (Oncorhynchus nerka). We discovered seven sex-associated loci, developed high-throughput assays for two loci, and tested the utility of these two assays in eight collections of sockeye salmon sampled throughout North America. We also screened an existing assay based on the master sex-determining gene in salmon (sdY) in these collections. The ability of GBS-derived loci to assign fish to their phenotypic sex varied substantially among collections suggesting that recombination between the loci that we discovered and the sex-determining gene has occurred. Assignment accuracy to phenotypic sex was much higher with the sdY assay but was still less than 100%. Alignment of sequences from GBS-derived loci to draft genomes for two salmonids provided strong evidence that many of these loci are found on the chromosome orthologous to the known sex chromosome in sockeye salmon. Our study is the first to describe the approximate location of the sex-determining region in sockeye salmon and indicates that sdY is also the master sex-determining gene in this species. However, discordances between sdY genotypes and phenotypic sex and the variable performance of GBS-derived loci warrant more research.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Phylogenomics of horned lizards (genus: Phrynosoma) using targeted sequence capture data

New genome sequencing techniques are enabling phylogenetic studies to scale-up from using a handful of loci to hundreds or thousands of loci from throughout the genome. In this study, we use targeted sequence capture (TSC) data from 540 ultraconserved elements and 44 protein-coding genes to estimate the phylogenetic relationships among all 17 species of horned lizards in the genus Phrynosoma. Previous molecular phylogenetic analyses of Phrynosoma based on a few nuclear genes, restriction site associated DNA (RAD) sequencing, or mitochondrial DNA (mtDNA) have produced conflicting relationships. Some of these conflicts are likely the result of rapid speciation at the start of Phrynosoma diversification, whereas other examples of gene tree discordance appear to be caused by active and residual traces of hybridization. Concatenation and coalescent-based species tree phylogenetic analyses of these new TSC data support the same topology, and a divergence dating analysis suggests that the Phrynosoma crown group is up to 30 million years old. The new phylogenomic tree supports the recognition of four main clades within Phrynosoma, including Anota (P. mcallii, P. solare, and the P. coronatum complex), Doliosaurus (P. modestum, P. goodei, and P. platyrhinos), Tapaja (P. ditmarsi, P. douglasii, P. hernandesi, and P. orbiculare), and Brevicauda (P. braconnieri, P. sherbrookei, and P. taurus). The phylogeny provides strong support for the relationships among all species of Phrynosoma and provides a robust new framework for conducting comparative analyses.

opencc-zeroDec 2014View details →
dryad32/100

Data from: High-throughput sequencing of nematode communities from total soil DNA extractions

Background: Nematodes are extremely diverse and numbers of species are predicted to be more than a million. Studies on nematode diversity are difficult and laborious using standard methods such as identification based on morphology and therefore high-throughput sequencing is an attractive alternative. Generally, primers that have been used for generating amplicons for sequencing are not nematode specific and also amplify other groups such as fungi and plantae. Thus a nematode enrichment step must be included that may introduce biases. Results: An amplification strategy, including a new primer, which selectively amplifies nematodes and other metazoans was developed. When this strategy was tested on DNA templates from a set of 22 agricultural soils, we obtained 64.4 % sequences of nematode origin in total, whereas the remaining sequences were almost entirely metazoan. The nematode sequences were derived from a broad taxonomic range and most sequences were from nematode taxa that have previously been found to be abundant in soil such as Tylenchida, Rhabditida, Dorylaimida, Triplonchida and Araeolaimida. Conclusions: This amplification and sequencing strategy for assessing nematode diversity was demonstrated to be able to collect a broad taxonomy of nematodes without prior enrichment and thus the method will be highly valuable in ecological studies of nematodes. Keywords: nematode, community, next-generation sequencing, SSU, diversity, 18S, rDNA

opencc-zeroDec 2014View details →
dryad32/100

Data from: Transatlantic secondary contact in Atlantic salmon, comparing microsatellites, a SNP array, and Restriction Associated DNA sequencing for the resolution of complex spatial structure

Identification of discrete and unique assemblages of individuals or populations is central to the management of exploited species. Advances in population genomics provide new opportunities for re-evaluating existing conservation units but comparisons among approaches remain rare. We compare the utility of RAD-seq, a single nucleotide polymorphism (SNP) array and a microsatellite panel to resolve spatial structuring under a scenario of possible trans-Atlantic secondary contact in a threatened Atlantic Salmon, Salmo salar, population in southern Newfoundland. Bayesian clustering indentified two large groups subdividing the existing conservation unit and multivariate analyses indicated significant similarity in spatial structuring among the three data sets. mtDNA alleles diagnostic for European ancestry displayed increased frequency in southeastern Newfoundland and were correlated with spatial structure in all marker types. Evidence consistent with introgression among these two groups was present in both SNP data sets but not the microsatellite data. Asymmetry in the degree of introgression was also apparent in SNP data sets with evidence of gene flow towards the east or European type. This work highlights the utility of RAD-seq based approaches for the resolution of complex spatial patterns, resolves a region of trans-Atlantic secondary contact in Atlantic Salmon in Newfoundland and demonstrates the utility of multiple marker comparisons in identifying dynamics of introgression.

opencc-zeroDec 2014View details →
dryad32/100

Data from: "Diagnostic SNPs for inferring population structure in American mink (Neovison vison) identified through RAD sequencing" in Genomic Resources Notes accepted 1 October 2014 to 30 November 2014

The article documents the public availability of RAD sequencing data and generated SNPs for the American mink (Neovison vison). 224,095 polymorphic loci were identified from 14 mink from which primers were designed for a subset of 380 SNPs. The panel was tested on 211 mink. Fisher's F-statistics (Fis, FIT and FST) as well as observed (HO), expected (HE) and unbiased expected (uHE) heterozygosity was calculated for the SNPs and 194 SNPs was validated as being useful for population genetic studies.

opencc-zeroDec 2014View details →
dryad32/100

Data from: Genome sequence and population declines in the critically endangered greater bamboo lemur (Prolemur simus) and implications for conservation

Background: The greater bamboo lemur (Prolemur simus) is a member of the Family Lemuridae that is unique in their dependency on bamboo as a primary food source. This Critically Endangered species lives in small forest patches in eastern Madagascar, occupying a fraction of its historical range. Here we sequence the genome of the greater bamboo lemur for the first time, and provide genome resources for future studies of this species that can be applied across its distribution. Results: Following whole genome sequencing of five individuals we identified over 152,000 polymorphic single nucleotide variants (SNVs), and evaluated geographic structuring across nearly 19k SNVs. We characterized a stronger signal associated with a north-south divide than across elevations for our limited samples. We also evaluated the demographic history of this species, and infer a dramatic population crash. This species had the largest effective population size (estimated between ~900,000 to one million individuals) between approximately 60,000-90,000 years before present (ybp), during a time in which global climate change affected terrestrial mammals worldwide. We also note the single sample from the northern portion of the extant range had the largest effective population size around 35,000 ybp. Conclusions: From our whole genome sequencing we recovered an average genomic heterozygosity of 0.0037%, comparable to other lemurs. Our demographic history reconstructions recovered a probable climate-related decline (60-90,000 ybp), followed by a second population decrease following human colonization, which has reduced the species to a census size of approximately 1,000 individuals. The historical distribution was likely a vast portion of Madagascar, minimally estimated at 44,259 km2, while the contemporary distribution is only ~1,700 km2. The decline in effective population size of 89-99.9% corresponded to a vast range retraction. Conservation management of this species is crucial to retain genetic diversity across the remaining isolated populations.

opencc-zeroDec 2017View details →
dryad32/100

Data from: A NGS approach to the encrusting Mediterranean sponge Crella elegans (Porifera, Demospongiae, Poecilosclerida): transcriptome sequencing, characterization and overview of the gene expression along three life cycle stages

Sponges can be dominant organisms in many marine and freshwater habitats where they play essential ecological roles. They also represent a key group to address important questions in early metazoan evolution. Recent approaches for improving knowledge on sponge biological and ecological functions as well as on animal evolution have focused on the genetic toolkits involved in ecological responses to environmental changes (biotic and abiotic), development and reproduction. These approaches are possible thanks to newly available, massive sequencing technologies–such as the Illumina platform, which facilitate genome and transcriptome sequencing in a cost-effective manner. Here we present the first NGS (next-generation sequencing) approach to understanding the life cycle of an encrusting marine sponge. For this we sequenced libraries of three different life cycle stages of the Mediterranean sponge Crella elegans and generated de novo transcriptome assemblies. Three assemblies were based on sponge tissue of a particular life cycle stage, including non-reproductive tissue, tissue with sperm cysts and tissue with larvae. The fourth assembly pooled the data from all three stages. By aggregating data from all the different life cycle stages we obtained a higher total number of contigs, contigs with blast hit and annotated contigs than from one stage-based assemblies. In that multi-stage assembly we obtained a larger number of the developmental regulatory genes known for metazoans than in any other assembly. We also advance the differential expression of selected genes in the three life cycle stages to explore the potential of RNA-seq for improving knowledge on functional processes along the sponge life cycle.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Characterization of the transcriptome, nucleotide sequence polymorphism, and natural selection in the desert adapted mouse Peromyscus eremicus

As a direct result of intense heat and aridity, deserts are thought to be among the most harsh of environments, particularly for their mammalian inhabitants. Given that osmoregulation can be challenging for these animals, with failure resulting in death, strong selection should be observed on genes related to the maintenance of water and solute balance. One such animal, Peromyscus eremicus, is native to the desert regions of the southwest United States and may live its entire life without oral fluid intake. As a first step toward understanding the genetics that underlie this phenotype, we present a characterization of the P. eremicus transcriptome. We assay four tissues (kidney, liver, brain, testes) from a single individual and supplement this with population level renal transcriptome sequencing from 15 additional animals. We identified a set of transcripts undergoing both purifying and balancing selection based on estimates of Tajima's D. In addition, we used the branch-site test to identify a transcript—Slc2a9, likely related to desert osmoregulation—undergoing enhanced selection in P. eremicus relative to a set of related non-desert rodents.

opencc-zeroDec 2013View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record