Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
318
datasets available to search
ShareScore release 0.7.1
Dataset results
318 results for “Data mining”
Data from: Mining the stable quantitative trait loci for agronomic traits in wheat (Triticum aestivum L.) based on an introgression line population
Open the record for dataset details and reuse information.
Legacies of historic charcoal production affect the forest flora in a Swedish mining district: survey data
Open the record for dataset details and reuse information.
Data from: Are dominant plant species more susceptible to leaf-mining insects? A case study at Saihanwula Nature Reserve, China
Open the record for dataset details and reuse information.
Data from: Anthropogenic mining alters macroinvertebrate size spectra in streams
Open the record for dataset details and reuse information.
Data from: Molecular phylogeny, revised higher classification, and implications for conservation of endangered Hawaiian leaf-mining moths (Lepidoptera: Gracillariidae: Philodoria)
Open the record for dataset details and reuse information.
Reconstruction of magnetospheric storm-time dynamics using cylindrical basis functions and multi-mission data mining
<p>This zip file contains data used to create figures and tables, describing the results of the paper "Reconstruction of magnetospheric storm-time dynamics using cylindrical basis functions and multi-mission data mining", by N. A. Tsyganenko, V. A. Andreeva, and M. I. Sitnov.</p>
Data from: Does mining waste concentration in the soil interfere with leaf selection by Acromyrmex subterraneus (Formicidae)
<p>Revegetation programs are proposed to recover the soil and biodiversity of disturbed sites, this being the case of the Rio Doce basin, Brazil. This region was hugely affected by a mining waste dam disruption, whose leakage on the soil altered its chemical and physical characteristics, and consequently the physiology and performance of plants. The expected alterations of the plants can make them more attractive for leaf-cutting ants, as lower water content induces an increase of non-structural carbohydrates. In this context, we evaluated whether <em>Acromyrmex subterraneus</em> workers differentiate among plants grown on soil with different mining waste concentrations. Leaf discs from plants grown in soil containing 0, 25, 50, 75 and 100% of mining waste were simultaneously offered to ant colonies in a foraging arena. The number of transported discs from each mining waste concentration was recorded until all discs of any concentration had been transported. Leaf selection assays were repeated after 30 days due to the novelty effect phenomenon. Leaf thickness, water, starch and total soluble carbohydrate contents were determined. Leaf discs from plants grown in soil with 100% of mining waste concentration were preferentially selected in both selection assays. Leaf thickness and water content were significantly lower in plants from the aforementioned treatment, while starch and total soluble carbohydrates were higher. Results suggest that seedlings implanted in sites with high mining waste concentration are under high predation risk. Revegetation programs must measure the impact of leaf-cutting ants as both herbivorous and soil ecosystem engineers, for the best management of these insects.</p>
Data from: Transcriptome-wide mining, characterization, and development of microsatellite markers in Lychnis kiusiana (Caryophyllaceae)
Background: Lychnis kiusiana Makino is an endangered perennial herb native to wetland areas in Korea and Japan. Despite its conservational and evolutionary significance, population genetic resources are lacking for this species. Next-generation sequencing has been accepted as a rapid and cost-effective solution for the identification of microsatellite markers in nonmodel plants. Results: Using Illumina HiSeq 2000 sequencing technology, we assembled 67,498,600 reads into 91,900 contigs and identified 11,403 microsatellite repeat motifs in 9,563 contigs. A total of 4,510 microsatellite-containing transcripts had Gene Ontology (GO) annotations, and Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis identified 124 pathways with significant scores. Many microsatellites in the L. kiusiana leaf transcriptome were linked to genes involved in the plant response to light intensity, salt stress, temperature stimulus, and nutrient and water deprivation. A total of 12,486 single-nucleotide polymorphisms (SNPs) were identified on transcripts harboring microsatellites. The analysis of nucleotide substitution rates for 2,389 unigenes indicated that 39 genes were under strong positive selection. The primers of 6,911 microsatellites were designed, and 40 of 50 selected primer pairs were consistently and successfully amplified from 51 individuals. Twenty-five of these were polymorphic, and the average number of alleles per SSR locus was 6.96, with a range from 2 to 15. The observed and expected heterozygosities ranged from 0.137 to 0.902 and 0.131 to 0.827, respectively, and locus-specific FIS estimates ranged from -0.116 to 0.290. Eleven of the 25 primer pairs were successfully amplified in three additional species of Lychnis: 56% in L. wilfordii, 64% in L. cognata and 80% in L. fulgens. Conclusions: The transcriptomic SSR markers of Lychnis kiusiana provide a valuable resource for understanding the population genetics, evolutionary history, and effective conservation management of this species. Furthermore, the identified microsatellite loci linked to the annotated genes should be useful for developing functional markers of L. kiusiana. The developed markers represent a potentially valuable source of transcriptomic SSR markers for population genetic analyses with moderate levels of cross-taxon portability.
Data from: A strengths-based data capture model: mining data-driven and person-centered health assets
With health care policy directives advancing value-based care, risk assessments and management have permeated health care discourse. The conventional problem-based infrastructure defines what data are employed to build this discourse and how it unfolds. Such a health care model tends to bias data for risk assessment and risk management toward problems and does not capture data about health assets or strengths. The purpose of this article is to explore and illustrate the incorporation of a strengths-based data capture model into risk assessment and management by harnessing data-driven and person-centered health assets using the Omaha System. This strengths-based data capture model encourages and enables use of whole-person data including strengths at the individual level and, in aggregate, at the population level. When aggregated, such data may be used for the development of strengths-based population health metrics that will promote evaluation of data-driven and person-centered care, outcomes, and value.
Data from: Mining microsatellite markers from public expressed sequence tags databases for the study of threatened plants
Background: Simple Sequence Repeats (SSRs) are widely used in population genetic studies but their classical development is costly and time-consuming. The ever-increasing available DNA datasets generated by high-throughput techniques offer an inexpensive alternative for SSRs discovery. Expressed Sequence Tags (ESTs) have been widely used as SSR source for plants of economic relevance but their application to non-model species is still modest. Methods: Here, we explored the use of publicly available ESTs (GenBank at the National Center for Biotechnology Information-NCBI) for SSRs development in non-model plants, focusing on genera listed by the International Union for the Conservation of Nature (IUCN). We also search two model genera with fully annotated genomes for EST-SSRs, Arabidopsis and Oryza, and used them as controls for genome distribution analyses. Overall, we downloaded 16 031 555 sequences for 258 plant genera which were mined for SSRsand their primers with the help of QDD1. Genome distribution analyses in Oryza and Arabidopsis were done by blasting the sequences with SSR against the Oryza sativa and Arabidopsis thaliana reference genomes implemented in the Basal Local Alignment Tool (BLAST) of the NCBI website. Finally, we performed an empirical test to determine the performance of our EST-SSRs in a few individuals from four species of two eudicot genera, Trifolium and Centaurea. Results: We explored a total of 14 498 726 EST sequences from the dbEST database (NCBI) in 257 plant genera from the IUCN Red List. We identify a very large number (17 102) of ready-to-test EST-SSRs in most plant genera (193) at no cost. Overall, dinucleotide and trinucleotide repeats were the prevalent types but the abundance of the various types of repeat differed between taxonomic groups. Control genomes revealed that trinucleotide repeats were mostly located in coding regions while dinucleotide repeats were largely associated with untranslated regions. Our results from the empirical test revealed considerable amplification success and transferability between congenerics. Conclusions: The present work represents the first large-scale study developing SSRs by utilizing publicly accessible EST databases in threatened plants. Here we provide a very large number of ready-to-test EST-SSR (17 102) for 193 genera. The cross-species transferability suggests that the number of possible target species would be large. Since trinucleotide repeats are abundant and mainly linked to exons they might be useful in evolutionary and conservation studies. Altogether, our study highly supports the use of EST databases as an extremely affordable and fast alternative for SSR developing in threatened plants.
Data from: Phylogenomic mining of the mints reveals multiple mechanisms contributing to the evolution of chemical diversity in Lamiaceae
The evolution of chemical complexity has been a major driver of plant diversification, with novel compounds serving as key innovations. The species-rich mint family (Lamiaceae) produces an enormous variety of compounds that act as attractants and defense molecules in nature and are used widely by humans as flavor additives, fragrances, and anti-herbivory agents. To elucidate the mechanisms by which such diversity evolved, we combined leaf transcriptome data from 48 Lamiaceae species and four outgroups with a robust phylogeny and chemical analyses of three terpenoid classes (monoterpenes, sesquiterpenes, iridoids) that share and compete for precursors. Our integrated chemical-genomic-phylogenetic approach revealed that: 1) gene family expansion rather than increased enzyme promiscuity of terpene synthases is correlated with mono- and sesqui-terpene diversity; 2) differential expression of core genes within the iridoid biosynthetic pathway is associated with iridoid presence/absence; 3) generally, production of iridoids and canonical monoterpenes appeared to be inversely correlated; and 4) iridoid biosynthesis was significantly associated with expression of geraniol synthase, which diverts metabolic flux away from canonical monoterpenes, suggesting that competition for common precursors can be a central control point in specialized metabolism. These results suggest that multiple mechanisms contributed to the evolution of chemodiversity in this economically important family.
Data from: An experimental study of the influence of lithology on compaction behaviour of broken waste rock in coal mine backfill
The research aims to explore the influences of lithology on the compaction behaviours of broken waste rocks. For this purpose, a WAW1000D servo test machine and a self-made bidirectional loading test system for granular materials were used to conduct axial and lateral compaction tests on four typical types of broken waste rocks: sandstone, mudstone, limestone, and shale. On this basis, we analysed the relationships between lateral and axial stress with the strain in, and porosity of, the four types of broken waste rocks. In addition, the relationship of axial stress with lateral stress and lateral pressure coefficient, and the changes in the particle size distribution of broken waste rocks before, and after, compaction were discussed. The test results demonstrated that the samples of higher strength were found to have low lateral and axial strains as well as a lower porosity in axial and lateral loading tests; while samples of lower strength showed low lateral stress and lateral pressure coefficient under axial load. After being compacted, the samples of the four types of broken waste rocks were found to have a higher proportion of small particles, indicating some particle crushing. Moreover, the samples of lower strength were broken to a greater extent.
Data from: Looking at cerebellar malformations through text-mined interactomes of mice and humans
We have generated and made publicly available two very large networks of molecular interactions: 49,493 mouse-specific and 52,518 human-specific interactions. These networks were generated through automated analysis of 368,331 full-text research articles and 8,039,972 article abstracts from the PubMed database, using the GeneWays system. Our networks cover a wide spectrum of molecular interactions, such as bind, phosphorylate, glycosylate, and activate; 207 of these interaction types occur more than 1,000 times in our unfiltered, multi-species data set. Because mouse and human genes are linked through an orthological relationship, human and mouse networks are amenable to straightforward, joint computational analysis. Using our newly generated networks and known associations between mouse genes and cerebellar malformation phenotypes, we predicted a number of new associations between genes and five cerebellar phenotypes (small cerebellum, absent cerebellum, cerebellar degeneration, abnormal foliation, and abnormal vermis). Using a battery of statistical tests, we showed that genes that are associated with cerebellar phenotypes tend to form compact network clusters. Further, we observed that cerebellar malformation phenotypes tend to be associated with highly connected genes. This tendency was stronger for developmental phenotypes and weaker for cerebellar degeneration.
Data from: Mining for single nucleotide polymorphisms and insertions / deletions in expressed sequence tag libraries of oil palm
The oil palm is a tropical oil bearing tree. Recently EST-derived SNPs and SSRs are a free by-product of the currently expanding EST (Expressed Sequence Tag) data bases. The development of high-throughput methods for the detection of SNPs (Single Nucleotide Polymorphism) and small indels (insertion / deletion) has led to a revolution in their use as molecular markers. Available (5452) Oil palm EST sequences were mined from dbEST of NCBI. CAP3 program was used to assemble EST sequences into contigs. Candidate SNPs and Indel polymorphisms were detected using the perl script auto_snip version 1.0 which has used 576 ESTs for detecting SNPs and Indel sites. We found 1180 SNP sites and 137 indel polymorphisms with frequency 1.36 SNPs / 100 bp. Among the six tissues from which the EST libraries had been generated, mesocarp had high frequency of 2.91 SNPs and indels per 100 bp whereas the zygotic embryos had lowest frequency of 0.15 per 100 bp. We also used the Shannon index to analyze the proportion of ten possible types of SNP/indels. ESTs from tissues of normal apex showed highest values of Shannon index (0.60) whereas abnormal apex had least value (0.02). The present report deals the use of Shannon index for comparing SNP/ indel frequencies mined from ESTlibraries and also confirm that the frequency of SNP occurrence in oil palm to use them as markers for genetic studies.
Data from: An exceptionally high nucleotide and haplotype diversity and a signature of positive selection for the eIF4E resistance gene in barley are revealed by allele mining and phylogenetic analyses of natural populations.
In barley, the eukaryotic translation initiation factor 4E (eIF4E) gene situated on chromosome 3H is recognised as an important source of resistance to the bymoviruses Barley yellow mosaic virus and Barley mild mosaic virus. In modern barley cultivars two recessive eIF4E alleles, rym4 and rym5, confer different isolate-specific resistances. In this study the sequence of eIF4E was analysed in 1090 barley landraces and non-current cultivars originating from 84 countries. An exceptionally high nucleotide diversity was evident in the coding sequence of eIF4E but not in either the adjacent MCT-1 gene or the sequence related eIF(iso)4E gene situated on chromosome 1H. Surprisingly, all nucleotide polymorphisms detected in the coding sequence of eIF4E resulted in amino acid changes. A total of 47 eIF4E haplotypes were identified and phylogenetic analysis using maximum likelihood provided evidence of strong positive selection acting on this barley gene. The majority of eIF4E haplotypes were found to be specific to distinct geographic regions. Furthermore, the eIF4E haplotype diversity (uh) was found to be considerably higher in East Asia, whereas SNP genotyping identified a comparatively low degree of genome-wide genetic diversity in 16 out of 17 tested accessions (each carrying a different eIF4E haplotype) from this same region. In addition, selection statistic calculations using coalescent simulations showed evidence of non neutral variation for eIF4E in several geographic regions, including East Asia, the region with a long history of the bymovirus-induced yellow mosaic disease. Together these findings suggest eIF4E may play a role in barley adaptation to local habitats.
Data from: Using plant functional distances to select species for restoration of mining sites
1. Plant facilitation, an ecological interaction that benefits at least one species without harming the other, is increasingly used as a restoration tool. To restore degraded habitats under a facilitation framework, practitioners must correctly select both the benefactor (nurse) and the beneficiary (facilitated) species. 2. Based on community assembly and species coexistence theory, we propose selecting plant species that largely differ in a suite of functional traits so that competition is minimized and facilitation maximized due to functional complementarity. To apply this guideline in a pilot restoration experiment performed in metalliferous mine tailings in South-Eastern Spain, we first built the plant-plant facilitative interaction network naturally occurring in a set of 12 tailings. After characterizing each species with 20 morphological and physiological traits, we verified that facilitative interactions were predominantly established between functionally distant species. 3. Then, we designed a sowing experiment combining 50 nurse-facilitated species pairs separated by a wide range of functional distances. The success of seedling establishment significantly increased with the functional distance between the nurse and the facilitated plant species. 4. Synthesis and applications. We encourage to use ecological facilitation together with trait-based species selection to design restoration programmes based on the principle of increasing functional distance between target species. This method may not only promote the restoration of the plant cover but also impact paramount ecosystem functions, thus being an efficient low cost restoration practice in abiotically stressful ecosystems.
Data from: Mining of expressed sequence tag libraries of cacao for microsatellite markes using five computational tools
Expressed Sequence Tags (ESTs) provide researchers with a quick and inexpensive route for discovering new genes, and data on gene expression and regulation and provide genic markers that help in constructing genome maps. Cacao is an important perennial crop of humid tropics. Cacao EST sequences as available in public domain were downloaded and made into contigs. A total of 769 contigs were made using contigs assembly program pharp. Puative information of contigs were identified using NCBI and ExPASy tools such as BlastX, tblastn.
Data from: Parasite infection of public databases: a data mining approach to identify apicomplexan contaminations in animal genome and transcriptome assemblies
Background: Contaminations from various exogenous sources are a common problem in next-generation sequencing. Another possible source of contaminating DNA are endogenous parasites. On the one hand, undiscovered contaminations of animal sequence assemblies may lead to erroneous interpretation of data; on the other hand, when identified, parasite-derived sequences may provide a valuable source of information. Results: Here we show that sequences deriving from apicomplexan parasites can be found in many animal genome and transcriptome projects, which in most cases derived from an infection of the sequenced host specimen. The apicomplexan sequences were extracted from the sequence assemblies using a newly developed bioinformatic pipeline (ContamFinder) and tentatively assigned to distinct taxa employing phylogenetic methods. We analysed 920 assemblies and found 20,907 contigs of apicomplexan origin in 51 of the datasets. The contaminating species were identified as members of the apicomplexan taxa Gregarinasina, Coccidia, Piroplasmida, and Haemosporida. For example, in the platypus genome assembly, we found a high number of contigs derived from a piroplasmid parasite (presumably Theileria ornithorhynchi). For most of the infecting parasite species, no molecular data had been available previously, and some of the datasets contain sequences representing large amounts of the parasite's gene repertoire. Conclusion: Our study suggests that parasite-derived contaminations represent a valuable source of information that can help to discover and identify new parasites, and provide information on previously unknown host-parasite interactions. We, therefore, argue that uncurated assembly data should routinely be made available in addition to the final assemblies.
FIGURE 1 in New species of leaf-mining Nepticulidae (Lepidoptera) from the Neotropical and Ando-Patagonian regions, with new data on host plants
FIGURE 1. Distribution map to the species treated in the current paper.
Data from: PMSeeker: A Scheme for paternity marker set mining
<p>The paternity test is a genetic test that analyzes genetic characteristics (mostly molecular markers) to identify whether two individuals have a parent-child relationship. It is frequently employed in judicial identification and economic species breeding. Because a single marker's discriminability is restricted, numerous markers are typically utilized to produce an accurate result. Obviously, having an excessive number of redundant markers wastes time and resources, and adequate approaches are required to screen reduced and efficient paternity marker sets (PMS). This study established a non-redundant PMS-screening scheme based on the exhaustive algorithm and greedy algorithm. When screening PMS, the greedy algorithm selects markers based on the parental dispersity index (PDI), a uniquely defined metric that outperforms polymorphic information content (PIC) and probability of exclusion (PE). With the conjunctive use of the two algorithms the optimal solutions were found for more than 99.7% of solvable cases in three groups of random sample experiments in this study. This scheme effectively reduces the number of markers in PMS, so conserving people and experimental resources and laying the groundwork for the widespread implementation of paternity assignment technology in economic species breeding.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.