Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

317

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

317 results for “gene structure”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: Impact of a putative riverine barrier on genomic population structure and gene flow in the presence of sexual selection

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Population studies of the wild tomato species Solanum chilense reveal geographically structured major gene-mediated pathogen resistance

Open the record for dataset details and reuse information.

publicDec 2020View details →
dryad36/100

Data from: Structure, gene order, and nucleotide composition of mitochondrial genomes in parasitic lice from Amblycera

Open the record for dataset details and reuse information.

publicNov 2020View details →
dryad36/100

Data from: Genetic analysis of red deer (Cervus elaphus) administrative management units in a human-dominated landscape - patterns of genetic diversity, population structure and gene flow

Open the record for dataset details and reuse information.

publicApr 2024View details →
dryad36/100

Data from: Population structure, ancestral admixture, gene flow, and landscape association of blacklegged ticks during range expansion in the Midwestern U.S.

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad36/100

Data from: A genomic assessment of population structure and gene flow in an aquatic salamander identifies the roles of spatial scale, barriers, and river architecture

Open the record for dataset details and reuse information.

publicFeb 2019View details →
dryad36/100

Structural evolution drives diversification of the large LRR-RLK gene family

Open the record for dataset details and reuse information.

publicJan 2020View details →
dryad32/100

Data from: Nearly complete rRNA genes from 371 Animalia: updated structure-based alignment and phylogenetic analysis

This study presents a manually constructed alignment of nearly complete rRNA genes from most animal clades (371 taxa from ∼33 of the ∼36 metazoan phyla), expanded from the 197 sequences in a previous study. This thorough, taxon-rich alignment, available at http://www.wsu.edu/≃jmallatt/research/rRNAalignment.html and in the Dryad Repository (doi: http://dx.doi.org/10.5061/dryad.1v62kr3q), is based rigidly on the secondary structure of the SSU and LSU rRNA molecules, and is annotated in detail, including labeling of the erroneous sequences (contaminants). The alignment can be used for future studies of the molecular evolution of rRNA. Here, we use it to explore if the larger number of sequences produces an improved phylogenetic tree of animal relationships. Disappointingly, the resolution did not improve, neither when the standard maximum-likelihood method was used, nor with more sophisticated methods that partitioned the rRNA into paired and unpaired sites (stem, loop, bulge, junction), or accounted for the evolution of the paired sites. For example, no doublet model of paired-site substitutions (16-state, 16A and 16B, 7A–F, or 6A–C models) corrected the placement of any rogue taxa or increased resolution. The following findings are from the simplest, standard, ML analysis. The 371-taxon tree only imperfectly supported the bilaterian clades of Lophotrochozoa and Ecdysozoa, and this problem remained after 17 taxa with unstably positioned sequences were omitted from the analysis. The problem seems to stem from base-compositional heterogeneity across taxa and from an overrepresentation of highly divergent sequences among the newly added taxa (e.g., sequences from Cephalopoda, Rotifera, Acoela, and Myxozoa). The rogue taxa continue to concentrate in two locations in the rRNA tree: near the base of Arthropoda and of Bilateria. The approximately uncertain (AU) test refuted the monophyly of Mollusca and of Chordata, probably due to long-branch attraction of the highly divergent cephalopod and urochordate sequences out of those clades. Unlikely to be correct, these refutations show for the first time that rRNA phylogeny can support some 'wrong' clades. Along with its weaknesses, the rRNA tree has strengths: It recovers many clades that are supported by independent evidence (e.g., Metazoa, Bilateria, Hexapoda, Nonoculata, Ambulacraria, Syndermata, and Thecostraca with Malacostraca) and shows good resolution within certain groups (e.g., in Platyhelminthes, Insecta, Cnidaria). As another strength, the newly added rRNA sequences yielded the first rRNA-based support for Carnivora and Cetartiodactyla (dolphin + llama) in Mammalia, for basic subdivisions of Bryozoa ('Gymnolaemata + Stenolaemata' versus Phylactolaemata), and for Oligostraca (ostracods + branchiurans + pentastomids + mystacocarids). Future improvement could come from better sequence-evolution models that account for base-compositional heterogeneity, and from combining rRNA with protein-coding genes in phylogenetic reconstruction.

opencc-zeroDec 2011View details →
dryad32/100

Data from: Spatial scales of genetic structure and gene flow in Calochortus albus (Liliaceae)

Calochortus (Liliaceae) displays high species richness, restriction of many individual taxa to narrow ranges, geographic coherence of individual clades, and parallel adaptive radiations in different regions. Here we test the first part of a hypothesis that all of these patterns may reflect gene flow at small geographic scales. We use amplified fragment length polymorphism variation to quantify the geographic scales of spatial genetic structure and apparent gene flow in Calochortus albus, a widespread member of the genus, at Henry Coe State Park in the Coast Ranges south of San Francisco Bay. Analyses of 254 mapped individuals spaced 0.001–14.4 km apart show a highly significant decline in genetic identity with ln distance, implying a root-mean-square distance of gene flow σ of 5–43 m. STRUCTURE analysis implies the existence of 2–4 clusters over the study area, with frequent reversals among clusters over short distances (<200 m) and a relatively high frequency of admixture within individuals at most sampling sites. While the intensity of spatial genetic structure in C. albus is weak, as measured by the Sp statistic, that appears to reflect low genetic identity of adjacent plants, which might reflect repeated colonizations at small spatial scales or density-dependent mortality of individual genotypes by natural enemies. Small spatial scales of gene flow and spatial genetic structure should permit, under a variety of conditions, genetic differentiation within species at such scales, setting the stage ultimately for speciation and adaptive radiation as such scales as well.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Landscape genetic structure of Scirpus mariqueter reveals a putatively adaptive differentiation under strong gene flow in estuaries

Estuarine organisms grow in highly heterogeneous habitats and their genetic differentiation are driven by selective and neutral processes as well as population colonization history. However, the relative importance of the processes that underlie genetic structure is still puzzling. Scirpus mariqueter is a perennial grass almost limited in the Changjiang River estuary and it's adjacent Qiantang River estuary. Here, using amplified fragment length polymorphism (AFLP), a moderate-high level of genetic differentiation among populations (range FST: 0.0310-0.3325) was showed despite large ongoing dispersal. FLOCK assigned all individuals to 13 clusters and revealed a complex genetic structure. Some genetic clusters were limited in peripheries compared with very mixing constitution in centre populations, suggesting local adaptation was more likely to occur in peripheral populations. 21 candidate outliers under positive selection were detected, and further, the differentiation patterns correlated with geographic distance, salinity difference and colonization history were analyzed with or without the outliers. Combined results of AMOVA and IBD based on different dataset, it was found that the effects of geographic distance and population colonization history on isolation seemed to be promoted by divergent selection. However, none-liner IBE pattern indicates the effects of salinity were overwhelmed by spatial distance or other ecological process in certain areas, and also suggests that salinity was not the only selective factor driving population differentiation. These results together indicate that geographic distance, salinity difference and colonization history co-contributed in shaping the genetic structure of S. mariqueter, and that their relative importance was correlated with spatial scale and environment gradient.

opencc-zeroDec 2018View details →
dryad32/100

Data from: Intercontinental genetic structure and gene flow in Dunlin (Calidris alpina), a potential vector of avian influenza

Waterfowl (Anseriformes) and shorebirds (Charadriiformes) are the most common wild vectors of influenza A viruses. Due to their migratory behavior, some may transmit disease over long distances. Migratory connectivity studies can link breeding and nonbreeding grounds while illustrating potential interactions among populations that may spread diseases. We investigated Dunlin (Calidris alpina), a shorebird with a subspecies (C. a. arcticola) that migrates from nonbreeding areas endemic to avian influenza in eastern Asia to breeding grounds in northern Alaska. Using microsatellites and mitochondrial DNA, we illustrate genetic structure among six subspecies: C. a. arcticola, C. a. pacifica, C. a. hudsonia, C. a. sakhalina, C. a. kistchinski, and C. a. actites. We demonstrate that mitochondrial DNA can help distinguish C. a. arcticola on the Asian nonbreeding grounds with >70% accuracy depending on their relative abundance, indicating that genetics can help determine if C. a. arcticola occurs where they may be exposed to highly pathogenic avian influenza (HPAI) during outbreaks. Our data reveal asymmetric intercontinental gene flow, with some C. a. arcticola short-stopping migration to breed with C. a. pacifica in western Alaska. Because C. a. pacifica migrates along the Pacific Coast of North America, interactions between these subspecies and other taxa provides route for transmission of HPAI into other parts of North America.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Fragmentation can increase spatial genetic structure without decreasing pollen-mediated gene flow in a wind-pollinated tree

Fragmentation reduces population sizes, increases isolation between habitats, and can result in restricted dispersal of pollen and seeds. Given that diploid seed dispersal contributes more to shaping fine-scale spatial genetic structure than haploid pollen flow, we tested whether fine-scale SGS can be sensitive to fragmentation even if extensive pollen dispersal is maintained. Castanopsis sclerophylla (Lindley & Paxton) Schottky (Fagaceae), a wind-pollinated and gravity seed-dispersed tree, was studied in an area of Southeast China where its populations have been fragmented to varying extents by human activity. Using different age classes of trees in areas subject to varying extents of fragmentation, we found no significant difference in genetic diversity between pre- and post-fragmentation C. sclerophylla subpopulations. Genetic differentiation among post-fragmentation subpopulations was also only slightly lower than among post-fragmentation subpopulations. In the most fragmented habitat, selfing rates were significantly higher than zero in pre-fragmentation, but not post-fragmentation, cohorts. These results suggest that fragmentation had not decreased gene flow among these populations and that pollen flow remains extensive. However, significantly greater fine-scale SGS was found in post-fragmentation subpopulations in the most fragmented habitat, but not in less fragmented habitats. This alteration in SGS reflected more restricted seed dispersal, induced by changes in the physical environments and the prevention of secondary seed dispersal by rodents. An increase in SGS can therefore result from more restricted seed dispersal, even in the face of extensive pollen flow, making it a sensitive indicator of the negative consequences of population fragmentation.

opencc-zeroDec 2010View details →
dryad32/100

Data from: Comparative landscape genetics of pond-breeding amphibians in Mediterranean temporal wetlands: the positive role of structural heterogeneity in promoting gene flow

Comparative landscape genetics studies can provide key information to implement cost-effective conservation measures favoring a broad set of taxa. These studies are scarce, particularly in Mediterranean areas, which include diverse but threatened biological communities. Here we focus on Mediterranean wetlands in central Iberia and perform a multi-level, comparative study of two endemic pond-breeding amphibians, a salamander (Pleurodeles waltl) and a toad (Pelobates cultripes). We genotyped 411 salamanders from 20 populations and 306 toads from 16 populations at 18 and 16 microsatellite loci, respectively, and identified major factors associated with population connectivity through the analysis of three sets of variables potentially affecting gene flow at increasingly finer levels of spatial resolution. Topographic, land use/cover, and remotely sensed vegetation/moisture indices were used to derive optimized resistance surfaces for the two species. We found contrasting patterns of genetic structure, with stronger, finer-scale genetic differentiation in Pleurodeles waltl, and notable differences in the role of fine-scale patterns of heterogeneity in vegetation cover and water content in shaping patterns of regional genetic structure in the two species. Overall, our results suggest a positive role of structural heterogeneity in population connectivity in pond-breeding amphibians, with habitat patches of Mediterranean scrubland and open oak woodlands ("dehesas") facilitating gene flow. Our study highlights the usefulness of remotely sensed continuous variables of land cover, vegetation and water content (e.g., NDVI, NDMI) in conservation-oriented studies aimed at identifying major drivers of population connectivity.

opencc-zeroDec 2016View details →
dryad32/100

Data from: Geographic population structure of the African malaria vector Anopheles gambiae suggests a role for the forest-savannah biome transition as a barrier to gene flow

The primary Afrotropical malaria mosquito vector Anopheles gambiae sensu stricto has a complex population structure. In western Africa, this species is split into two molecular forms and displays local and regional variation in chromosomal arrangements and behaviours. To investigate patterns of macro-geographic population substructure, 25 An. gambiae samples from 12 African countries were genotyped at 13 microsatellite loci. This analysis detected the presence of additional population structuring, with the M-form being subdivided into distinct west, central and southern African genetic clusters. These clusters are coincident with the central African rainforest belt and northern and southern savannah biomes, which suggests restrictions to gene flow associated with the transition between these biomes. By contrast geographically patterned population substructure appears much weaker within the S-form.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Population structure, relatedness and ploidy levels in an apple gene bank revealed through genotyping-by-sequencing

In recent years, new genome-wide marker systems have provided highly informative alternatives to low density marker systems for evaluating plant populations. To date, most apple germplasm collections have been genotyped using low-density markers such as simple sequence repeats (SSRs), whereas only a few have been explored using high-density genome-wide marker information. We explored the genetic diversity of the Pometum gene bank collection (University of Copenhagen, Denmark) of 349 apple accessions using over 15,000 genome-wide single nucleotide polymorphisms (SNPs) and 15 SSR markers, in order to compare the strength of the two approaches for describing population structure. We found that 119 accessions shared a clonal relationship with at least one other accession in the collection, resulting in the identification of 272 (78%) unique accessions. Of these unique accessions, over half (52%) share a first-degree relationship with at least one other accession. There is therefore a high degree of clonal and family relatedness in the Danish apple gene bank. We find significant genetic differentiation between Malus domestica and its supposed primary wild ancestor, M. sieversii, as well as between accessions of Danish origin and all others. Overall, we found strong concordance between analyses based on the genome-wide SNPs and the 15 SSR loci. However, we argue that GBS is superior to traditional SSR approaches because it allowed the estimation of ploidy levels that were in accordance with flow cytometry results, and can be further exploited in genome-wide association studies (GWAS). Finally, we compare GBS with SSR for the purposes of characterizing a diverse apple gene bank and discuss the advantages and constraints of the two approaches.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Recombination-dependent replication and gene conversion homogenize repeat sequences and diversify plastid genome structure

PREMISE OF THE STUDY: There is a misinterpretation in the literature regarding the variable orientation of the small single copy region of plastid genomes (plastomes). The common phenomenon of small and large single copy inversion, hypothesized to occur through intramolecular recombination between inverted repeats (IR) in a circular, single unit-genome, in fact more likely occurs through recombination-dependent replication (RDR) of linear plastome templates. If RDR can be primed through both intra- and intermolecular recombination, then this mechanism could not only create inversion isomers of so-called single copy regions, but also an array of alternative sequence arrangements. METHODS: We used Illumina paired-end and PacBio single-molecule real-time (SMRT) sequences to characterize repeat structure in the plastome of Monsonia emarginata L'Hér. (Geraniaceae). We used OrgConv and inspected nucleotide alignments to infer ancestral nucleotides and identify gene conversion among repeats and mapped long (>1 kb) SMRT reads against the unit-genome assembly to identify alternative sequence arrangements. RESULTS: Although M. emarginata lacks the canonical IR, we found that large repeats (>1 kilobase; kb) represent ~22% of the plastome nucleotide content. Among the largest repeats (>2 kb) we identified GC-biased gene conversion and mapping filtered, long SMRT reads to the M. emarginata unit-genome assembly revealed alternative, substoichiometric sequence arrangements. CONCLUSION: We offer a model based on RDR and gene conversion between long repeated sequences in the M. emarginata plastome, and provide support that both intra-and intermolecular recombination between large repeats, particularly in repeat-rich plastomes, varies unit-genome structure while homogenizing the nucleotide sequence of repeats.

opencc-zeroDec 2016View details →
dryad32/100

Data from: The influence of population structure on gene expression and flowering time variation in the ubiquitous weed Capsella bursa-pastoris (Brassicaceae)

Population structure is a potential problem when testing for adaptive phenotypic differences among populations. The observed phenotypic differences among populations can simply be due to genetic drift, and if the genetic distance between them is not considered, the differentiation may be falsely interpreted as adaptive. Conversely, adaptive and demographic processes might have been tightly associated and correcting for the population structure may lead to false negatives. Here, we evaluated this problem in the cosmopolitan weed Capsella bursa-pastoris. We used RNA-Seq to analyse gene expression differences among 24 accessions, which belonged to a much larger group that had been previously characterized for flowering time and circadian rhythm and were genotyped using genotyping-by-sequencing (GBS) technique. We found that clustering of accessions for gene expression retrieved the same three clusters that were obtained with GBS data previously, namely Europe, the Middle East and Asia. Moreover, the three groups were also differentiated for both flowering time and circadian rhythm variation. Correction for population genetic structure when analysing differential gene expression analysis removed all differences among the three groups. This may suggest that most differences are neutral and simply reflect population history. However, geographical variation in flowering time and circadian rhythm indicated that the distribution of adaptive traits might be confounded by population structure. To bypass this confounding effect, we compared gene expression differentiation between flowering ecotypes within the genetic groups. Among the differentially expressed genes, FLOWERING LOCUS C was the strongest candidate for local adaptation in regulation of flowering time.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Haplotype structure, adaptive history and associations with exploratory behaviour of the DRD4 gene region in four great tit (Parus major) populations

The assessment of genetic architecture and selection history in genes for behavioural traits is fundamental to our understanding of how these traits evolve. The dopamine receptor D4 (DRD4) gene is a prime candidate for explaining genetic variation in novelty seeking behaviour, a commonly assayed personality trait in animals. Previously we showed that a single nucleotide polymorphism in exon 3 of this gene is associated with exploratory behaviour in at least one of four Western European great tit (Parus major) populations. These heterogeneous association results were explained by potential variable linkage disequilibrium (LD) patterns between this marker and the causal variant or by other genetic or environmental differences among the populations. Different adaptive histories are further hypothesized to have contributed to these population differences. Here, we genotyped 98 polymorphisms of the complete DRD4 gene including the flanking regions for 595 individuals of the four populations. We show that the LD structure, specifically around the original exon 3 SNP is conserved across the four populations and does not explain the heterogeneous association results. Study-wide significant associations with exploratory behaviour were detected in more than one haplotype block around exon 2, 3 and 4 in two of the four tested populations with different allele effect models. This indicates genetic heterogeneity in the association between multiple DRD4 polymorphisms and exploratory behaviour across populations. The association signals were in or close to regions with signatures of positive selection. We therefore hypothesize that variation in exploratory and other dopamine-related behaviour evolves locally by occasional adaptive shifts in the frequency of underlying genetic variants.

opencc-zeroDec 2012View details →
dryad32/100

Data from: Intra-population genomics in a model mutualist: population structure and candidate symbiosis genes under selection in Medicago truncatula

Bottom-up evolutionary approaches, including geographically-explicit population genomic analyses, have the power to reveal the mechanistic basis of adaptation. Here we conduct a population genomic analysis in the model legume, Medicago truncatula, in order to characterize population genetic structure and identify symbiosis-related genes showing evidence of spatially-variable selection. Using RAD-seq, we generated over 26,000 SNPs from 191 accessions from within three regions of the native range in Europe. Results from STRUCTURE analysis identify 5 distinct genetic clusters with divisions that separate east and west regions in the Mediterranean basin. Much of the genetic variation is maintained within sampling sites, and there is evidence for isolation by distance. Extensive linkage disequilibrium was identified, particularly within populations. We conducted genetic outlier analysis with FST-based genome scans and a bayesian modeling approach (PCAdapt). There were 70 core outlier loci shared between these distinct methods with one clear candidate symbiosis related gene, DMI1. This work sets that stage for functional experiments to determine the important phenotypes that selection has acted upon and complementary efforts in rhizobium populations.

opencc-zeroDec 2015View details →
dryad32/100

Data from: Genetic diversity and structure of Lolium perenne ssp. multiflorum in California vineyards and orchards indicates potential for spread of herbicide resistance via gene flow

Management of agroecosystems with herbicides imposes strong selection pressures on weedy plants leading to the evolution of resistance against those herbicides. Resistance to glyphosate in populations of Lolium perenne L. ssp. multiflorum is increasingly common in California, USA, causing economic losses and the loss of effective management tools. To gain insights into the recent evolution of glyphosate resistance in L. perenne in perennial cropping systems of northwest California and to inform management, we investigated the frequency of glyphosate resistance and the genetic diversity and structure of 14 populations. The sampled populations contained frequencies of resistant plants ranging from 10% to 89%. Analyses of neutral genetic variation using microsatellite markers indicated very high genetic diversity within all populations regardless of resistance frequency. Genetic variation was distributed predominantly among individuals within populations rather than among populations or sampled counties, as would be expected for a wide-ranging outcrossing weed species. Bayesian clustering analysis provided evidence of population structuring with extensive admixture between two genetic clusters or gene pools. High genetic diversity and admixture, and low differentiation between populations, strongly suggests the potential for spread of resistance through gene flow and the need for management that limits seed and pollen dispersal in L. perenne.

opencc-zeroDec 2016View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record