Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
134
datasets available to search
ShareScore release 0.9.0
Dataset results
134 results for “marker development”
Data from: Genomic exploration and molecular marker development in a large and complex conifer genome using RADseq and mRNAseq
Open the record for dataset details and reuse information.
Data from: Conserved genetic regions across angiosperms as tools to develop single copy nuclear markers in gymnosperms: an example using cycads
Open the record for dataset details and reuse information.
Data from: Identifying differentially expressed genes under heat stress and developing molecular markers in orchardgrass (Dactylis glomerata L.) through transcriptome analysis
Open the record for dataset details and reuse information.
A diversity arrays technology developed high-density DArT markers as a tool for efficient and rapid genomic assessment in Ethiopian Mustard (Brassica carinata A. Braun)
Open the record for dataset details and reuse information.
Data from: Genome- and transcriptome-assisted development of nuclear insertion/deletion markers for Calanus species (Copepoda: Calanoida) identification
Open the record for dataset details and reuse information.
Data from: Transcriptional markers of sub-optimal nutrition in developing Apis mellifera nurse workers
Open the record for dataset details and reuse information.
SNPs markers of Trypoxylus dichotomus developed by using specific‐locus amplified fragment sequencing (SLAF‐seq) techniques
Open the record for dataset details and reuse information.
Data from: Evolutionary factors affecting the cross-species utility of newly developed microsatellite markers in seabirds
Open the record for dataset details and reuse information.
Data from: Rapid microsatellite marker development for African mahogany (Khaya senegalensis, Meliaceae) using next-generation sequencing and assessment of its intra-specific genetic diversity.
Open the record for dataset details and reuse information.
Data from: Developing nuclear DNA phylogenetic markers in the angiosperm genus Leucadendron (Proteaceae): a next-generation sequencing transcriptomic approach
Open the record for dataset details and reuse information.
QTL mapping and marker development for tolerance to sulfur phytotoxicity in melon (Cucumis Melo)
Open the record for dataset details and reuse information.
Data from: Low-coverage, whole-genome sequencing of Artocarpus camansi (Moraceae) for phylogenetic marker development and gene discovery
Premise of the study: We used moderately low-coverage (17×) whole-genome sequencing of Artocarpus camansi (Moraceae) to develop genomic resources for Artocarpus and Moraceae. Methods and Results: A de novo assembly of Illumina short reads (251,378,536 pairs, 2 × 100 bp) accounted for 93% of the predicted genome size. Predicted coding regions were used in a three-way orthology search with published genomes of Morus notabilis and Cannabis sativa. Phylogenetic markers for Moraceae were developed from 333 inferred single-copy exons. Ninety-eight putative MADS-box genes were identified. Analysis of all predicted coding regions resulted in preliminary annotation of 49,089 genes. An analysis of synonymous substitutions for pairs of orthologs (Ks analysis) in M. notabilis and A. camansi strongly suggested a lineage-specific whole-genome duplication in Artocarpus. Conclusions: This study substantially increases the genomic resources available for Artocarpus and Moraceae and demonstrates the value of low-coverage de novo assemblies for nonmodel organisms with moderately large genomes.
Data from: Transcriptome-wide mining, characterization, and development of microsatellite markers in Lychnis kiusiana (Caryophyllaceae)
Background: Lychnis kiusiana Makino is an endangered perennial herb native to wetland areas in Korea and Japan. Despite its conservational and evolutionary significance, population genetic resources are lacking for this species. Next-generation sequencing has been accepted as a rapid and cost-effective solution for the identification of microsatellite markers in nonmodel plants. Results: Using Illumina HiSeq 2000 sequencing technology, we assembled 67,498,600 reads into 91,900 contigs and identified 11,403 microsatellite repeat motifs in 9,563 contigs. A total of 4,510 microsatellite-containing transcripts had Gene Ontology (GO) annotations, and Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis identified 124 pathways with significant scores. Many microsatellites in the L. kiusiana leaf transcriptome were linked to genes involved in the plant response to light intensity, salt stress, temperature stimulus, and nutrient and water deprivation. A total of 12,486 single-nucleotide polymorphisms (SNPs) were identified on transcripts harboring microsatellites. The analysis of nucleotide substitution rates for 2,389 unigenes indicated that 39 genes were under strong positive selection. The primers of 6,911 microsatellites were designed, and 40 of 50 selected primer pairs were consistently and successfully amplified from 51 individuals. Twenty-five of these were polymorphic, and the average number of alleles per SSR locus was 6.96, with a range from 2 to 15. The observed and expected heterozygosities ranged from 0.137 to 0.902 and 0.131 to 0.827, respectively, and locus-specific FIS estimates ranged from -0.116 to 0.290. Eleven of the 25 primer pairs were successfully amplified in three additional species of Lychnis: 56% in L. wilfordii, 64% in L. cognata and 80% in L. fulgens. Conclusions: The transcriptomic SSR markers of Lychnis kiusiana provide a valuable resource for understanding the population genetics, evolutionary history, and effective conservation management of this species. Furthermore, the identified microsatellite loci linked to the annotated genes should be useful for developing functional markers of L. kiusiana. The developed markers represent a potentially valuable source of transcriptomic SSR markers for population genetic analyses with moderate levels of cross-taxon portability.
Data from: Development of conserved microsatellite markers of high cross-species utility in bat species (Vespertilionidae, Chiroptera, Mammalia)
Comparative ecological and behavioural studies of the widespread and diverse Vespertilionidae, which comprise almost 400 of the 1,100 bat species, have been limited by the availability of markers. The potential of new methods for developing conserved microsatellite markers which possess enhanced cross-species utility has recently been illustrated in studies of birds. We have applied these methods to develop enhanced microsatellite markers for vespertilionid bats, in particular for the genus Myotis (103 species). We compared published bat microsatellites with their homologs in the genome sequence of the little brown bat, Myotis lucifugus to create consensus sequences which were used to design candidate primer sets. Primer sets were then tested for amplification and polymorphism in 22 species of bat from nine of the largest families (including 11 Vespertilionidae). Of 46 loci tested, 33 were polymorphic, on average, for each of seven Myotis species tested, 20 in each of four non-Myotis vespertilionid species, and two in 11 non-vespertilionid species.
Data from: A new resource for the development of SSR markers: millions of loci from a thousand plant transcriptomes
Premise of the study: The One Thousand Plant Transcriptomes Project (1KP, 1000+ assembled plant transcriptomes) provides an enormous resource for developing microsatellite loci across the plant tree of life. We developed loci from these transcriptomes and tested their utility. Methods and Results: Using software packages and custom scripts, we identified microsatellite loci in 1KP transcriptomes. We assessed the potential for cross-amplification and whether loci were biased toward exons, as compared to markers derived from genomic DNA. We characterized over 5.7 million simple sequence repeat (SSR) loci from 1334 plant transcriptomes. Eighteen percent of loci substantially overlapped with open reading frames (ORFs), and electronic PCR revealed that over half the loci would amplify successfully in conspecific taxa. Transcriptomic SSRs were approximately three times more likely to map to translated regions than genomic SSRs. Conclusions: We believe microsatellites still have a place in the genomic age—they remain effective and cost-efficient markers. The loci presented here are a valuable resource for researchers.
Data from: Marker development for phylogenomics: the case of Orobanchaceae, a plant family with contrasting nutritional modes
Phylogenomic approaches, employing next-generation sequencing (NGS) techniques, have revolutionized systematic and evolutionary biology. Target enrichment is an efficient and cost-effective method in phylogenomics and is becoming increasingly popular. Depending on availability and quality of reference data as well as on biological features of the study system, (semi-)automated identification of suitable markers will require specific bioinformatic pipelines. Here, we established a highly flexible bioinformatic pipeline, BaitsFinder, to identify putative orthologous single copy genes (SCGs) and to construct bait sequences in a single workflow. Additionally, this pipeline has been constructed to be able to cope with challenging data sets, such as the nutritionally heterogeneous plant family Orobanchaceae. To this end, we used transcriptome data of differing quality available for four Orobanchaceae species and, as reference, SCG data from monkeyflower (Erythranthe guttata, syn. Mimulus g.; 1,915 genes) and tomato (Solanum lycopersicum; 391 genes). Depending on whether gaps were permitted in initial blast searches of the four Orobanchaceae species against the reference, our pipeline identified 1,307 and 981 SCGs with average length of 994 bp and 775 bp, respectively. Automated bait sequence construction (using 2× tiling) resulted in 38,170 and 21,856 bait sequences, respectively. In comparison to the recently published MarkerMiner 1.0 pipeline BaitsFinder identified about 1.6 times as many SCGs (of at least 900 bp length). Skipping steps specific to analyses of Orobanchaceae, BaitsFinder was successfully used in a group of non-parasitic plants (three Asteraceae species and, as reference, SCG data from Arabidopsis thaliana based on previously compiled SCGs). Thus, BaitsFinder is expected to be broadly applicable in groups, where only transcriptomes or partial genome data of differing quality are available.
Data from: Transcriptome sequencing and marker development for four underutilized legumes
Premise of the study: Combating threats to food and nutrition security in the context of climate change and global population increase is one of the highest priorities of major international organizations. Hundreds of species are grown on a small scale in some of the most drought/flood-prone regions of the world and as such may harbor some of the most environmentally tolerant crops (and alleles). Methods and Results: In this study, transcriptomes were sequenced, assembled, and annotated for four underutilized legume crops. Microsatellite markers were identified in each species, as well as a conserved orthologous set of markers for cross-family phylogenetics and comparative mapping, which were ground-truthed on a panel of diverse legume germplasm. Conclusions: An understanding of these underutilized legumes will inform crop selection and breeding by allowing the investigation of genetic variation and the genetic basis of adaptive traits to be established.
Figure 1 from: Acquah ME, Aboagye F, Ashong Y, Mosi L (2024) Development of a field diagnostic tool for Schistosoma mansoni Praziquantel resistant markers in selected endemic communities. Research Ideas and Outcomes 10: e120899. https://doi.org/10.3897/rio.10.e120899
Figure 1 Graphical Abstract. Adapted from Summers et al. (2022), available under the terms of the Creative Commons Attribution License 4.0 (CC BY 4.0).
Figure 1 from: Darschnik S, Leese F, Weiss M, Weigand H (2019) When barcoding fails: development of diagnostic nuclear markers for the sibling caddisfly species Sericostoma personatum (Spence in Kirby & Spence, 1826) and Sericostoma flavicorne Schneider, 1845. ZooKeys 872: 57-68. https://doi.org/10.3897/zookeys.872.34278
Figure 1 Schematic overview of the different multiplex and single RFLP fingerprints for S. flavicorne (Sf) and S. personatum (Sp). Fragment lengths are given above/below bands. Red: EcoRV1, EcoRV5, NdeI4, NdeI8, and PvuII2; Green: EcoRV4, EcoRV2, NdeI1, NdeI5.
Supplementary material 2 from: Darschnik S, Leese F, Weiss M, Weigand H (2019) When barcoding fails: development of diagnostic nuclear markers for the sibling caddisfly species Sericostoma personatum (Spence in Kirby & Spence, 1826) and Sericostoma flavicorne Schneider, 1845. ZooKeys 872: 57-68. https://doi.org/10.3897/zookeys.872.34278
: Data type: molecular data
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.