Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13
datasets available to search
ShareScore release 0.9.0
Dataset results
13 results for “code completion”
Complete datasets and code for "Hungry or angry? Experimental evidence for the effects of food availability on two measures of stress in developing wild raptor nestlings"
<p><strong>Abstract</strong></p> <p>Food shortage challenges the development of nestlings; yet, to cope with this stressor, nestlings can induce stress responses to adjust metabolism or behaviour. Food shortage also enhances the antagonism between siblings, but it remains unclear whether the stress response induced by food shortage operates via the individual nutritional state or via the social environment experienced. In addition, the understanding of these processes is hindered by the fact that effects of food availability often co-vary with other environmental factors. We used a food supplementation experiment to test the effect of food availability on two complementary stress measures, feather corticosterone (CORTf) and Heterophil/Lymphocyte-ratio (H/L) in developing red kite (Milvus milvus) nestlings, a species with competitive brood hierarchy. By statistically controlling for the effect of food supplementation on the nestlings’ body condition, we disentangled the effects of food and ambient temperature on nestlings during development. Experimental food supplementation increased body condition, and both CORTf and H/L were reduced in nestlings of high body condition. Additionally, CORTf decreased with age in non-supplemented nestlings. H/L decreased with age in all nestlings and was lower in supplemented last-hatched nestlings compared to non-supplemented ones. Ambient temperature showed a negative effect on H/L. Our results indicate that food shortage increases the nestlings’ stress levels through both, a reduced food intake affecting nutritional state and the nestlings’ social environment. Thus, food availability in conjunction with ambient temperature shape between- and within nest differences in stress load, which may have carry-over effects on behaviour and performance in further life-history stages.</p>
YaCoS: a Complete Infrastructure to the Design and Exploration of Code Optimization Sequences
<p>The growing popularity of machine learning frameworks and algorithms has greatly contributed to the design and exploration of good code optimization sequences. Yet, in spite of this progress, mainstream compilers still provide users with only a handful of fixed optimization sequences. Finding optimization sequences that are good in general is challenging because the universe of possible sequences is potentially infinite. This paper describes a infrastructure that provides developers with the means to explore this space. Said infrastructure, henceforth called YaCoS, consists of benchmarks,search algorithms, metrics to estimate the distance between programs, and compilation strategies. YaCoS’s features let users build learning models that predict, for unknown pro-grams, optimization sequences that are likely to yield good results for them. In this paper, as a case study, we have used YaCoS to find good optimization sequences for LLVM, using code size as the objective function. Such study lets us evaluate three feature sets: two variations of the feature vectors proposed by Namolaru at al in 2010, plus the optimization statistics produced by LLVM. Our results show that YaCoSis able to find sequences that improve onto clang -Oz by3.75% on average. Our experiments do not indicate a dominant feature set out of the three approaches that we have investigated—it is possible to find programs in which one of them is strictly better than the others.</p>
Sample Stripped Pre-supernova Progenitors for open-source code CHIPS (Complete History for Interaction-Powered Supernovae)
<p>Inlists, mainly based on the test suite "example_make_pre_ccsn" in r12778, with slight amendments for removal of hydrogen (and helium, for Ic progenitors) envelope at core hydrogen (helium) exhausion.</p><p>For details: https://ui.adsabs.harvard.edu/abs/2023arXiv230810785T/abstract</p>
Data and code from: Complete metamorphosis promotes morphological and functional diversity in Caudata
Open the record for dataset details and reuse information.
Complete codes for "Linking functional traits and diversity-invasibility hypothesis in submerged macrophyte communities under eutrophication", by Li
Open the record for dataset details and reuse information.
Data and code for: The complete Kaposi Sarcoma-associated herpesvirus genome induces early-onset, metastatic angiosarcoma in transgenic mice
Open the record for dataset details and reuse information.
[Replication Package] Why Personalizing Deep Learning-Based Code Completion Tools Matters
<p>This repository contains scripts, datasets and results of the work: <em>Why Personalizing Deep Learning-Based Code Completion Tools Matters</em></p> <p><strong>Note: you can access our scripts also in the official GitHub repository: <a href="https://github.com/Devy99/comp-personalization">https://github.com/Devy99/comp-personalization</a></strong></p>
DeepSEA complete datasets and code
<p><strong>Surveying antimicrobial resistance (AMR) is essential to track the evolution and spread of resistant genes/proteins since AMR is a major cause of death, being responsible for more deaths than HIV and malaria combined. Alignment-based annotation tools use strict similarity (>70%) cutoffs to distinguish between potential and AMR real sequences and only annotate proteins similar to those in their databases. DeepARG and AMRFinderPlus use artificial neural networks (ANN) and Hidden Markov Models (HMM) to annotate AMR proteins with remote homology. However, DeepARG needs a pre-processing step that aligns the query data and selects the most probable proteins, although the filtering uses looser cutoffs. HMMs also depend on multi-sequence alignment (MSA) and are focused on a single AMR class. Here we present DeepSEA, an alignment-free tool fitted on antimicrobial resistant proteins (APR) and non-resistant proteins (NRP) aligned and unaligned to ARP. Our results show that DeepSEA outperforms the current multi-class AMR classifiers. Furthermore, DeepSEA’s model can cluster AMR by resistant mechanisms, showing that the model's latent variables successfully captured distinguishing features of antibiotic resistance. Our tool annotated functionally validated tetracycline destructases (TDases) and confirmed the identification of a novel TDase found by HMM. </strong></p>
Data from: Targeted capture of complete coding regions across divergent species
Despite continued advances in sequencing technologies, there is a need for methods that can efficiently sequence large numbers of genes from diverse species. One approach to accomplish this is targeted capture (hybrid enrichment). While these methods are well established for genome resequencing projects, cross-species capture strategies are still being developed and generally focus on the capture of conserved regions, rather than complete coding regions from specific genes of interest. The resulting data is thus useful for phylogenetic studies, but the wealth of comparative data that could be used for evolutionary and functional studies is lost. Here we design and implement a targeted capture method that enables recovery of complete coding regions across broad taxonomic scales. Capture probes were designed from multiple reference species and extensively tiled in order to facilitate cross-species capture. Using novel bioinformatics pipelines we were able to recover nearly all of the targeted genes with high completeness from species that were up to 200 myr divergent. Increased probe diversity and tiling for a subset of genes had a large positive effect on both recovery and completeness. The resulting data produced an accurate species tree, but importantly this same data can also be applied to studies of molecular evolution and function that will allow researchers to ask larger questions in broader phylogenetic contexts. Our method demonstrates the utility of cross-species approaches for the capture of full length coding sequences, and will substantially improve the ability for researchers to conduct large-scale comparative studies of molecular evolution and function.
The Hidden Cost of Code Completion: Understanding the Impact of the Recommendation-list Length on its Efficiency
<p>This contains the data set and the code we take advantage of to process data and make visualization. The data set needs to be unzipped before the program could run on it. Besides, we choose another visualization tool to make the results more explicit.</p>
DeepSEA project complete datset and code
Open the record for dataset details and reuse information.
Data from: Targeted capture of complete coding regions across divergent species
Open the record for dataset details and reuse information.
When Large Language Models Meet Fragile Code Completion Dataset
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.