Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
25
datasets available to search
ShareScore release 0.7.1
Dataset results
25 results for “non-model species”
Data from: "Transcriptome resources for two non-model freshwater crustacean species" in Genomic Resources Notes accepted 1 October 2014 to 30 November 2014
Open the record for dataset details and reuse information.
Data from: Recommendations for population and individual diagnostic SNP selection in non-model species
Open the record for dataset details and reuse information.
Development of a panel of SNP loci in the emblematic southern damselfly (Coenagrion mercuriale) using a hybrid method: Pitfalls and recommendations for large-scale SNP genotyping in a non-model endangered species
Open the record for dataset details and reuse information.
Data from: Exploring a Pool-seq only approach for gaining population genomic insights in non-model species
<p>Developing genomic insights is challenging in non-model species for which resources are often scarce and prohibitively costly. Here, we explore the potential of a recently established approach using Pool-seq data to generate a de novo genome assembly for mining exons, upon which Pool-seq data is used to estimate population divergence and diversity. We do this for two pairs of sympatric populations of brown trout (Salmo trutta); one naturally sympatric set of populations and another pair of populations introduced to a common environment. We validate our approach by comparing the results to those from markers previously used to describe the populations (allozymes and individual based SNPs) and from mapping the Pool-seq data to a reference genome of the closely related Atlantic salmon (Salmo salar). We find that genomic differentiation (FST) between the two introduced populations exceeds that of the naturally sympatric populations (FST = 0.13 and 0.03 between the introduced and the naturally sympatric populations, respectively), in concordance with estimates from the previously used SNPs. The same level of population divergence is found for the two genome assemblies but estimates of average genic diversity differ (π ≈0.002 and π ≈0.001 when mapping to S. trutta and S. salar, respectively), although the relationships between population values are largely consistent. This discrepancy might be attributed to biases when mapping to a haploid condensed assembly made of highly fragmented read data compared to using a high-quality reference assembly from a divergent species. We conclude that the Pool-seq only approach can be suitable for detecting and quantifying genome wide population differentiation, and for comparing genomic diversity in populations of non-model species where reference genomes are lacking.</p>
Data from: A priori and a posteriori approaches for finding genes of evolutionary interest in non-model species: osmoregulatory genes in the kidney transcriptome of the desert rodent Dipodomys spectabilis (banner-tailed kangaroo rat)
One common goal in evolutionary biology is the identification of genes underlying adaptive traits of evolutionary interest. Recently next-generation sequencing techniques have greatly facilitated such evolutionary studies in species otherwise depauperate of genomic resources. Kangaroo rats (Dipodomys sp.) serve as exemplars of adaptation in that they inhabit extremely arid environments, yet require no drinking water because of ultra-efficient kidney function and osmoregulation. As a basis for identifying water conservation genes in kangaroo rats, we conducted a priori bioinformatics searches in model rodents (Mus musculus and Rattus norvegicus) to identify candidate genes with known or suspected osmoregulatory function. We then obtained 446,758 reads via 454 pyrosequencing to characterize genes expressed in the kidney of banner-tailed kangaroo rats (Dipodomys spectabilis). We also determined candidates a posteriori by identifying genes that were overexpressed in the kidney. The kangaroo rat sequences revealed nine different a priori candidate genes predicted from our Mus and Rattus searches, as well as 32 a posteriori candidate genes that were overexpressed in kidney. Mutations in two of these genes, Slc12a1 and Slc12a3, cause human renal diseases that result in the inability to concentrate urine. These genes are likely key determinants of physiological water conservation in desert rodents.
Data from: "Genome-wide microsatellite marker development from next-generation sequencing of two non-model bat species impacted by wind turbine mortality: Lasiurus borealis and L. cinereus (Vespertilionidae)" in Genomic Resources Notes accepted 1 October 2013 to 30 November 2013
Tree-roosting bats in the genus Lasiurus are widespread, migratory species that have not been well characterized for population genetic diversity and structure due to a lack of genetic resources. Generating genetic resources in Lasiurus is made pressing by the need for conservation genetic assessments of demographic trends in this genus, which comprise a large percentage of bat mortalities at wind turbine sites across North America. We report on marker development from whole-genome Illumina sequencing of the red bat (Lasirus borealis) and the hoary bat (L. cinereus). We generated paired-end libraries for a single individual of each species, sequenced on the Illumina HiSeq platform. We mapped a total of 46.6 million reads to the Myotis lucifigus reference genome, and used bioinformatics searches to identify tends of thousands of simple sequence repeats (SSRs) distributed across the bat genome. We selected 48 candidate microsatellite loci to develop cross-species primer sequences for Lasiurus, assembled these into multiplex combinations, and tested for amplification and polymorphism levels in a sample of 23 individuals from each of L. borealis and L. cinereus. In total, we identified 42 highly polymorphic loci that could be robustly amplified and scored, the majority of which (39) were also combinable into highly multiplexed assays of 4-8 loci each. The combination of new genomic sequence assemblies, a large set of highly polymorphic microsatellite loci, and the ability to efficiently multiplex represents a significant contribution to the genetic resources available for population and comparative genetic studies of bats.
An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.
<p>This archive is associated with the article “An evaluation of pool-sequencing transcriptome-based exon capture for population genomics of non-model species.”. Authors: Emeline Deleury, Thomas Guillemaud, Aurelie Blin & Eric Lombaert.</p> <p>The archive contains :<br> - The sequences of the 5,717 Harmonia axyridis randomly selected CDS (5717-targeted-CDS-sequences.gff3, sequence in FASTA format at the end of the file)<br> - For the subset of 3,161 targeted CDS that have a genomic match over their entire length, the positions of exons on transcripts (3161-targeted-CDS-EXON-POSITIONS.csv)</p>
Haploid, diploid, and pooled exome capture recapitulate features of biology and paralogy in two non-model tree species
<p>Despite their suitability for studying evolution, many conifer species have large and repetitive giga-genomes (16-31Gbp) that create hurdles to producing high coverage SNP datasets that capture diversity from across the entirety of the genome. Due in part to multiple ancient whole genome duplication events, gene family expansion and subsequent evolution within <i>Pinaceae</i>, false diversity from the misalignment of paralog copies creates further challenges in accurately and reproducibly inferring evolutionary history from sequence data. Here, we leverage the cost-saving benefits of pool-seq and exome-capture to discover SNPs in two conifer species, Douglas-fir (<i>Pseudotsuga menziesii</i> var. <i>menziesii </i>(Mirb.) Franco, <i>Pinaceae</i>) and jack pine (<i>Pinus banksiana</i> Lamb., <i>Pinaceae</i>). We show, using minimal baseline filtering, that allele frequencies estimated from pooled individuals show a strong positive correlation with those estimated by sequencing the same population as individuals (r > 0.948), on par with such comparisons made in model organisms. Further, we highlight the utility of haploid megagametophyte tissue for identifying sites that are likely due to misaligned paralogs. Together with additional minor filtering, we show that it is possible to remove many of the loci with large frequency estimate discrepancies between individual and pooled sequencing approaches, improving the correlation further (r > 0.973). Our work addresses bioinformatic challenges in non-model organisms with large and complex genomes, highlights the use of megagametophyte tissue for the identification of paralog sites, and suggests the combination of pool-seq and exome capture to be robust for further evolutionary hypothesis testing in these systems.</p>
Data from: Targeted re-sequencing of coding DNA sequences for SNP discovery in non-model species
Open the record for dataset details and reuse information.
Data from: "Genome-wide microsatellite marker development from next-generation sequencing of two non-model bat species impacted by wind turbine mortality: Lasiurus borealis and L. cinereus (Vespertilionidae)" in Genomic Resources Notes accepted 1 October 2013 to 30 November 2013
Open the record for dataset details and reuse information.
Data from: A multiscale approach to detect selection in non-model tree species: widespread adaptation despite population decline in Taxus baccata L.
Open the record for dataset details and reuse information.
Haploid, diploid, and pooled exome capture recapitulate features of biology and paralogy in two non-model tree species
Open the record for dataset details and reuse information.
Data from: Transposable element annotation in non-model species - on the benefits of species specific repeat libraries using semi-automated EDTA and DeepTE de novo pipelines
Open the record for dataset details and reuse information.
Data from: Exploring a Pool-seq only approach for gaining population genomic insights in non-model species
Open the record for dataset details and reuse information.
Data from: A priori and a posteriori approaches for finding genes of evolutionary interest in non-model species: osmoregulatory genes in the kidney transcriptome of the desert rodent Dipodomys spectabilis (banner-tailed kangaroo rat)
Open the record for dataset details and reuse information.
Data from: Finding candidate genes under positive selection in non-model species: examples of genes involved in host specialization in pathogens
Numerous genes in diverse organisms have been shown to be under positive selection, especially genes involved in reproduction, adaptation to contrasting environments, hybrid inviability, and host-pathogen interactions. Looking for genes under positive selection in pathogens has been a priority in efforts to investigate coevolution dynamics and to develop vaccines or drugs. To elucidate the functions involved in host specialization, here we aimed at identifying candidate sequences that could have evolved under positive selection among closely related pathogens specialized on different hosts. For this goal, we sequenced ca. 17,000-32,000 ESTs from each of four Microbotryum species, which are fungal pathogens responsible for anther smut disease on host plants in the Caryophyllaceae. Forty-two of the 372 predicted orthologous genes showed significant signal of positive selection, which represents a good number of candidate genes for further investigation. Sequencing 16 of these genes in 9 additional Microbotryum species confirmed that they have indeed been rapidly evolving in the pathogen species specialized on different hosts. The genes showing significant signals of positive selection were putatively involved in nutrient uptake from the host, secondary metabolite synthesis and secretion, respiration under stressful conditions and stress response, hyphal growth and differentiation, and regulation of expression by other genes. Many of these genes had transmembrane domains and may therefore also be involved in pathogen recognition by the host. Our approach thus revealed fruitful and should be feasible for many non-model organisms for which candidate genes for diversifying selection are needed.
Data from: De novo transcriptomic analyses for non-model organisms: an evaluation of methods across a multi-species data set
High-throughput sequencing (HTS) is revolutionizing biological research by enabling scientists to quickly and cheaply query variation at a genomic scale. Despite the increasing ease of obtaining such data, using these data effectively still poses notable challenges, especially for those working with organisms without a high-quality reference genome. For every stage of analysis – from assembly to annotation to variant discovery – researchers have to distinguish technical artefacts from the biological realities of their data before they can make inference. In this work, I explore these challenges by generating a large de novo comparative transcriptomic data set data for a clade of lizards and constructing a pipeline to analyse these data. Then, using a combination of novel metrics and an externally validated variant data set, I test the efficacy of my approach, identify areas of improvement, and propose ways to minimize these errors. I find that with careful data curation, HTS can be a powerful tool for generating genomic data for non-model organisms.
Data from: Finding candidate genes under positive selection in non-model species: examples of genes involved in host specialization in pathogens
Open the record for dataset details and reuse information.
Data from: De novo transcriptomic analyses for non-model organisms: an evaluation of methods across a multi-species data set
Open the record for dataset details and reuse information.
Data from: BsRADseq: screening DNA methylation in natural populations of non-model species
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.