Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
743
datasets available to search
ShareScore release 0.7.1
Dataset results
743 results for “clone”
Limited evidence of cloning and selfing within wild populations of coral-eating crown-of thorns seastar (Acanthaster cf. solaris)
<p>Population outbreaks of crown-of-thorns seastars (CoTS; <i>Acanthaster</i> spp.) are contributing to extensive coral loss and reef degradation throughout the Indo west-Pacific, but the causes and underlying mechanisms of population maintenance and outbreaks are equivocal. Two recent publications suggest that, in addition to outbreeding sexual reproduction, asexual reproduction through larval fission and selfing may contribute to rapid increases in the local abundance of <i>Acanthaster</i> spp. We re-analysed two large microsatellite datasets (collectively representing 3,714 individuals) that investigated connectivity in the Great Barrier Reef and Pacific region to investigate if potential cloning or selfing can be evidenced in the population genetic structure. Within this dataset we identified only a small number (18, < 0.5%) of putative clones (repeated multi locus genotypes). We argue that several of these are due to sampling and processing errors rather than direct evidence of cloning. Analysis of the population genetic structure (i.e., pairwise genetic differences between individuals, deviations from Hardy-Weinberg-Equilibrium, and linkage disequilibrium) also yielded no genetic evidence for asexual reproduction. There was a tendency towards slight heterozygote deficits, so we cannot refute that selfing does occur, but this is mostly likely attributable to sampling artefacts. Although we cannot exclude that asexual reproduction occurs to some extent in <i>Acanthaster</i> populations, we find no evidence that these processes make <span>a</span> contribution to population structure or directly enhance larval supply.</p>
Data from: Widespread generalist clones are associated with range and niche expansion in allopolyploids of Pacific Northwest Hawthorns (Crataegus L.)
Range and niche expansion are commonly associated with transitions to asexuality, polyploidy, and hybridity (allopolyploidy) in plants. The ability of asexual polyploids to colonize novel habitats may be due to widespread generalist clones, multiple ecologically specialized clones, or may be a neutral byproduct of multiple, independent origins of asexual polyploids throughout the range. We have quantified niche size and divergence for hawthorns of the Pacific Northwest using data from herbarium vouchers with known cytotypes. We find that all polyploid niches diverge from that of the diploid range, and allopolyploids have the broadest niches. Allotetraploids have the largest niche and the widest geographic distribution. We then assessed the genetic mechanism of range expansion by surveying the ecological and geographic distribution of genotypes within each cytotype from sites in which fine-scale habitat assessments were completed. We find no isolation by either geographic or ecological distance in allopolyploids, suggesting high dispersal and colonization ability. In contrast, autotriploids and diploids show patterns of isolation by geographic distance. We also compared the geographic and ecological distributions of clonal genotypes with those of randomly drawn sites of the most widespread cytotype. We found that most clones are geographically widespread and occur in a variety of habitats. We interpret these findings to suggest that patterns of range and niche expansion in Pacific Northwest Hawthorns may stem from these widespread, ecologically generalist clones of hybrid origin.
Data from: A multi-genome analysis approach enables tracking of the invasion of a single Russian wheat aphid (Diuraphis noxia) clone throughout the New World
This study investigated the population genetics, demographic history and pathway of invasion of the Russian wheat aphid (RWA) from its native range in Central Asia, the Middle East and Europe to South Africa and the Americas. We screened microsatellite markers, mitochondrial DNA, and endosymbiont genes in 504 RWA clones from nineteen populations worldwide. Following pathway analyses of microsatellite and endosymbiont data, we postulate that Turkey and Syria were the most likely sources of invasion to Kenya and South Africa, respectively. Furthermore, we found that one clone transferred between South Africa and the Americas was most likely responsible for the New World invasion. Finally, endosymbiont DNA was found to be a high resolution population genetic marker, extremely useful for studies of invasion over a relatively short evolutionary history timeframe. This study has provided valuable insights into the factors that may have facilitated the recent global invasion by this damaging pest.
Data from: Attack of the PCR clones: rates of clonality have little effect on RAD-seq genotype calls
Interpretation of high-throughput sequence data requires an understanding of how decisions made during bioinformatic data processing can influence results. One source of bias that is often cited is PCR clones (or PCR duplicates). PCR clones are common in restriction site associated sequencing (RAD-seq) datasets, which are increasingly being used for molecular ecology. To determine the influence PCR clones and the bioinformatic handling of clones have on genotyping, we evaluate four RAD-seq datasets. Datasets were compared before and after clones were removed to estimate the number of clones present in RAD-seq data, quantify how often the presence of clones in a dataset cause genotype calls to change compared to when clones were removed, investigate the mechanisms that lead to genotype call changes, and test if clones bias heterozygosity estimates. Our RAD-seq datasets contained 30 – 60% PCR clones, but 95% of RAD-tags had five or fewer clones. Relatively few genotypes changed once clones were removed (5-10%), and the vast majority of these changes (98%) were associated with genotypes switching from a called to no-call state or vice versa. PCR clones had a larger influence on genotype calls in individuals with low read depth but appeared to influence genotype calls at all loci similarly. Removal of PCR clones reduced the number of called genotypes by 2% but had almost no influence on estimates of heterozygosity. As such, while steps should be taken to limit PCR clones during library preparation, PCR clones are likely not a substantial source of bias for most RAD-seq studies.
Data from: Maladaptation to acute metal exposure in resurrected Daphnia ambigua clones after decades of increasing contamination
Human environmental impacts have driven some of the strongest and fastest phenotypic changes recorded in wild animal populations. Across populations, this variation is often adaptive, as populations evolve fitness advantages in response to human-modified environments. Yet some populations fail to adapt to changing environments. Evidenced by declines in relative fitness, such seemingly maladaptive outcomes are less common, but may be more likely in human modified contexts. Further, our ability to investigate the dynamics of these adaptive and maladaptive responses over time is typically limited in natural systems. I combined resurrection ecology and paleolimnology approaches to examine evolutionary responses of the freshwater zooplankter Daphnia to exposure to heavy metal contamination over the past 50-75 years using animals hatched from diapausing egg banks. In contrast to the predicted trend of adaptation to metal exposure over time, I observed an increase in sensitivity to both copper and cadmium exposure associated with increasing historic contamination. This potentially maladaptive trend occurred in Daphnia populations in three lakes. Given that the release of toxicants such as heavy metals is widespread and other researchers have observed local maladaptation to toxicant exposure, it is important to understand the drivers and implications of this pattern.
Data from: Positional cloning of rp2 QTL associates the P450 genes CYP6Z1, CYP6Z3 and CYP6M7 with pyrethroid resistance in the malaria vector Anopheles funestus
Pyrethroid resistance in Anopheles funestus is threatening malaria control in Africa. Elucidation of underlying resistance mechanisms is crucial to improve the success of future control programs. A positional cloning approach was used to identify genes conferring resistance in the uncharacterised rp2 QTL previously detected in this vector using F6 Advanced Intercross Lines (AIL). A 113 kb BAC clone spanning rp2 was identified and sequenced revealing a cluster of fifteen P450 genes and one salivary protein gene (SG7-2). Contrary to An. gambiae, AfCYP6M1 is triplicated in An. funestus while AgCYP6Z2 ortholog is absent. 565 new SNPs were identified for genetic mapping from rp2 P450s and other genes revealing high genetic polymorphisms with 1 SNP every 36bp. A significant genotype/phenotype association was detected for rp2 P450s but not for a cluster of cuticular protein genes previously associated with resistance in An. gambiae. QTL mapping using F6 AIL confirms the rp2 QTL with an increase logarithm of odds (LOD) score of 5. Multiplex gene expression profiling of 15 P450s and other genes around rp2 followed by individual validation using qRT-PCR indicated a significant over-expression in the resistant FUMOZ-R strain of the P450s AfCYP6Z1, AfCYP6Z3, AfCYP6M7 and the glutathione-s-transferase GSTe2 with respective fold-change of 11.2, 6.3, 5.5 and 2.8. Polymorphisms analysis of AfCYP6Z1 and AfCYP6Z3 identified amino acid changes potentially associated with resistance further indicating that these genes are controlling the pyrethroid resistance explained by the rp2 QTL. The characterisation of this rp2 QTL significantly improves our understanding of resistance mechanisms in An. funestus.
Data from: Clone configuration and spatial genetic structure of two Halophila ovalis populations with contrasting internode lengths
Fine-scale spatial genetic structure (SGS) is predominantly determined by gene flow. While sexually reproducing plants can disperse their genes through pollen and seed grains, clonal plants can additionally disperse genes through clonal growth. Plants' clonal reproduction strategy, however, often varies within and between species. Still, the effect of differential clonal reproduction strategy on fine-scale SGS remains somewhat unclear. Halophila ovalis is a fast-growing clonal seagrass, whose internode length (which defines a species' clonal reproduction strategy) varies among populations. Using eight polymorphic microsatellites, here we compare the genetic diversity, clonal structure and fine-scale SGS of two H. ovalis populations with contrasting internode lengths (Yingluo versus Xialongwei populations). We found moderate to high genotypic and allelic richness and heterozygosities in both populations. Compared to Xialongwei population, genetic and genotypic diversity was significantly lower in Yingluo population. Although their internode length was relatively short, clones of Yingluo population spread farther than those of Xialongwei population. Sexual-to-vegetative dispersal variance ratios were 34.6 and 445.5 in Yingluo and Xialongwei populations, respectively. In both populations, clonal growth significantly intensified the SGS, especially in short distance classes. The SGS at small distance classes were weaker in Yingluo than Xialongwei, in part, due to more intermingled distribution of genets and more extensive clonal expansion in the former population. Our results indicate that vegetative dispersal variance/distance, rather than internode length, plays a crucial role in shaping the fine-scale genetic structure.
Design, cloning and test expression of huntingtin domain constructs (2016/12/02)
<p>Open lab notebook huntingtin structure function project.<br> </p>
Cloning, eukaryotic expression and purification of full-length huntingtin Q23 (2017/06/01)
<p>Huntingtin structure function open lab notebook project</p>
Supplementary Materials for: Discovery of a novel merbecovirus DNA clone contaminating agricultural rice sequencing datasets from Wuhan, China
<p>Supplementary Materials for</p><p><strong>Discovery of a novel merbecovirus DNA clone contaminating agricultural rice sequencing datasets from Wuhan, China</strong></p><p>Adrian Jones, Daoyu Zhang, Steven E. Massey, Yuri Deigin, Louis R. Nemzer, Steven C. Quay</p>
Characterizing Code Clones from Large Language Models Dataset and Scripts
<p>characterizing_code_clones_data.zip: <br><br>This dataset contains a collection of code snippets generated by Large Language Models (LLMs) such as GPT-3.5 and GPT-4 in response to specific programming prompts derived from LeetCode. Each sub-directory within the dataset corresponds to a particular LLM version and contains code snippets, preprocessed data, and SLACC input files. </p><p>characterizing_code_clones_project.zip: </p><p>This zipped directory encompasses the core scripts and results used in the "Characterizing Code Clones of LLMs" research. It features the Python script <strong>collect_samples.py</strong> for collecting LLM-generated code snippets, as well as a suite of scripts in the <strong>slacc_scripts</strong> sub-directory for processing and analyzing the data using SLACC. The directory also includes the results of the LeetCode test suites, providing insights into the correctness and efficiency of the code generated by GPT-3.5 and GPT-4. </p>
Clone sequences
Open the record for dataset details and reuse information.
GBS and phenotype data for MASPOT population, a panel of tetraploid potato clones
Open the record for dataset details and reuse information.
Long-read genome sequencing of bread wheat facilitates disease resistance gene cloning
<p>Cloning agronomically important genes from large, complex crop genomes remains challenging. Here, we generate a 14.7-gigabase chromosome-scale<i> </i>assembly of the South African bread wheat (<i>Triticum aestivum</i>) cultivar Kariega by combining high-fidelity long reads, optical mapping, and chromosome conformation capture. The resulting assembly is an order of magnitude more contiguous than previous wheat assemblies. Kariega shows durable resistance against the devastating fungal stripe rust disease. We identified the race-specific disease resistance gene <i>Yr27</i>, encoding an intracellular immune receptor, as a major contributor to this resistance. <i>Yr27</i> is allelic to the leaf rust resistance gene <i>Lr13,</i> with the Yr27 and Lr13 proteins sharing 97% sequence identity. Our results thus demonstrate the feasibility of generating chromosome-scale wheat assemblies to clone genes and also exemplify that highly similar alleles of a single-copy gene can confer resistance to different pathogens, which might provide a basis for engineering <i>Yr27</i> alleles with multiple recognition specificities in future.</p>
Clones found with the oracle
<p>Here you can find the output of Simian and the clones we found between the training set and the oracle (i.e., the block we expected T5 to predict)</p>
Clones with predictions
<p>This contains the output of simian and the found clones with the prediction generated by the T5 model</p>
GPT-J Code Clone Detection
<p>This is the dataset of GPT-J code clone detection.</p> <p> </p> <p>'data/' folder contains the data to replicate our study. </p> <p>'code/' folder contains the code to replicate our study.</p> <p>`rebuttal/` folder contains the results of the rebuttal.</p> <p>Each folder has its own README, please check the details there.</p>
blase II: PHOENIX Subset Clone Archive
<p>As part of our study, we cloned 1,314 individual PHOENIX spectra (Husser et al. 2013) with blase and optimized their line shapes with ML. These interpretable clones have been saved as state dictionaries in .pt files, which we upload here to be readily accessible to others.</p> <p> </p> <p>You can download and unzip the file in order to access the clone state dictionaries. They can be loaded from disk using torch.load() (Paszke et al. 2019), and input into blase's SparseLinearEmulator (Gully-Santiago & Morley 2022) with its init_state_dict constructor argument.</p>
Analyzing the Dependability of Large Language Models for Code Clone Generation.
<div> <p>data.zip: <br><br>This dataset includes a collection of ten LeetCode programming problems used in the study "Analyzing the Dependability of Large Language Models for Code Clone Generation". At the top level, you will find a CSV file containing all the initial LeetCode data. Each subdirectory at this level represents a specific LeetCode problem. Within these subdirectories, you will find the original solutions, their behaviors, the input corpus, as well as folders dedicated to various temperatures, models, and code cloning tasks. Additionally, within the "repeated" folder, you will find the original LLM-generated snippets, the preprocessed snippets with the snippet behavior, and the results.</p> <p>characterizing_code_clones_project.zip: </p> <p>This zipped directory encompasses the core scripts and results used in the "Characterizing Code Clones of LLMs" research. The common folder was used to run the whole pipeline, the various parts of the pipeline are each in a folder as well as the various data analysis scripts! </p> <p> </p> </div>
Analyzing the Dependability of Large Language Models for Code Clone Generation
<p>data.zip: <br><br>This dataset includes a collection of ten LeetCode programming problems used in the study "Analyzing the Dependability of Large Language Models for Code Clone Generation". At the top level, you will find a CSV file containing all the initial LeetCode data. Each subdirectory at this level represents a specific LeetCode problem. Within these subdirectories, you will find the original solutions, their behaviors, the input corpus, as well as folders dedicated to various temperatures, models, and code cloning tasks. Additionally, within the "repeated" folder, you will find the original LLM-generated snippets, the preprocessed snippets with the snippet behavior, and the results.</p> <p>characterizing_code_clones_project.zip: </p> <p>This zipped directory encompasses the core scripts and results used in the "Characterizing Code Clones of LLMs" research. The common folder was used to run the whole pipeline, the various parts of the pipeline are each in a folder as well as the various data analysis scripts! </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.