Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
100
datasets available to search
ShareScore release 0.7.1
Dataset results
100 results for “genetic code”
Robust genetic codes enhance protein evolvability
<p>The repository contains all data for our manuscript on protein evolvability under rewired genetic codes: https://www.biorxiv.org/content/10.1101/2023.06.20.545706v1</p> <p>The corresponding code is available on GitHub: https://github.com/parizkh/rewired_codes_landscapes</p>
Experimental data to the publication "Genetic-optimised aperiodic code for distributed optical fibre sensors"
<p>The source data underlying Figs. 3-5 and Supplementary Figs. 6, 8-14 are provided as a Source Data file.</p>
Local chromatin context dictates the genetic determinants of the heterochromatin spreading reaction. Analysis Code, Numerical and Primary data.
<p>Uploaded under this Zenodo DOI is the following:</p> <p>1. the Analysis Code used for Flow Cytometry analysis in the paper, GO complex analysis (Figure 3) and Hit visualization (Figure 1, 2 S1, S4 Figs).</p> <p>2. The primary Flow Cytometry data from both the initial screen (ScreenFlowFCS) and validation experiments (ValidationFlowFCS) are included as .zip files.</p> <p>3. a .zip folder is uploaded that contains all the analysis code for the ChIP-Seq experiments. </p> <p>4. Excel worksheets that contain the numerical source data for all qPCR bar plots.</p>
Data and codes from "Daniel et al. What can optimized cost distances based on genetic distances offer? A simulation study on the use and misuse of ResistanceGA"
<p><span>Data and codes used for </span><span>“Daniel et al. What can optimized cost distances based on genetic distances offer? A simulation study on the use and misuse of ResistanceGA”</span></p>
Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure
<p>Predicting how tree populations will respond to climate change is an urgent societal concern. An increasingly popular way to make such predictions is the genomic offset (GO) approach, which aims to use genomic and climate data to identify populations that may experience climate maladaptation in the near future. More precisely, GO tries to represent the change in allele frequencies required to maintain the current gene-climate relationships under climate change. However, the GO approach has major limitations and, despite promising validation of its predictions using height data from common gardens, it still lacks broad empirical testing. In the present study, we evaluated the consistency and empirical validity of GO predictions in maritime pine (<em>Pinus pinaster</em> Ait.), a tree species from southwestern Europe and North Africa with a marked population genetic structure. First, gene-climate relationships were estimated using 9,817 SNPs genotyped in 454 trees from 34 populations; and candidate SNPs potentially involved in climate adaptation were identified. Second, GO was predicted using four methods, namely Gradient Forest (GF), Redundancy Analysis (RDA), latent factor mixed model (LFMM) and Generalised Dissimilarity Modeling (GDM), two sets of SNPs (candidate and control SNPs) and five climate general circulation models (GCMs) to account for uncertainty in future climate predictions. Last, the empirical validity of GO predictions was evaluated within a Bayesian framework by estimating the associations between GO predictions and two independent data sources: mortality data from National Forest Inventories (NFI), and mortality and height data from five common gardens in contrasting environments. We found high variability in GO predictions across methods, SNP sets and GCMs. Regarding validation, GO predictions with GDM and GF (and to a lesser extent RDA) based on the candidate SNPs showed the strongest and most consistent associations with mortality rates in common gardens and NFI plots. We found almost no association between GO predictions and tree height in common gardens, most likely due to the overwhelming effect of population genetic structure on tree height in this species. Our study demonstrates the imperative to validate GO predictions with a range of independent data sources before they can be used as informative and reliable metrics in conservation or management strategies.</p>
Using the Genetic Algorithm for the Optimization of Dynamic School Bus Routing Problem-Figure 4. Example of a chromosome structure with permutation coding
<p>Each chromosome found in the population formed in the GA is structurally an equal-length coded series. The chromosomes are made of genes. For coding purposes, binary, permutation, and value coding methods are widely used. In the travelling salesman or other similar VRPs, permutation coding technique is preferred over the other techniques. Using the permutation coding technique, each chromosome found in the population is expressed in terms of the numbers of each stop to be followed in the route, as shown in Figure 4.</p>
Data and code relating to Becher, Jackson & Charlesworth. Patterns of genetic variability in genomic regions with low rates of recombination.
<p>Data and code relating to Becher, Jackson & Charlesworth. Patterns of genetic variability in genomic regions with low rates of recombination.</p> <p>Contains genotype data, R code for analysis and visualisation, a SLiM simulation script, README, etc.</p>
Link to Dataset related to article "Interpreting Non-coding Genetic Variation in Multiple Sclerosis Genome-Wide Associated Regions"
<p>Link to Dataset related to article "Interpreting Non-coding Genetic Variation in Multiple Sclerosis Genome-Wide Associated Regions"</p> <p>Multiple sclerosis (MS) is the most common neurological disorder in young adults. Despite extensive studies, only a fraction of MS heritability has been explained, with association studies focusing primarily on protein-coding genes, essentially for the difficulty of interpreting non-coding features. However, non-coding RNAs (ncRNAs) and functional elements, such as super-enhancers (SE), are crucial regulators of many pathways and cellular mechanisms, and they have been implicated in a growing number of diseases. In this work, we searched for possible enrichments in non-coding elements at MS genome-wide associated loci, with the aim to highlight their possible involvement in the susceptibility to the disease. We first reconstructed the linkage disequilibrium (LD) structure of the Italian population using data of 727,478 single-nucleotide polymorphisms (SNPs) from 1,668 healthy individuals. The genomic coordinates of the obtained LD blocks were intersected with those of the top hits identified in previously published MS genome-wide association studies (GWAS). By a bootstrapping approach, we hence demonstrated a striking enrichment of non-coding elements, especially of circular RNAs (circRNAs) mapping in the 73 LD blocks harboring MS-associated SNPs. In particular, we found a total of 482 circRNAs (annotated in publicly available databases) vs. a mean of 194 ± 65 in the random sets of LD blocks, using 1,000 iterations. As a proof of concept of a possible functional relevance of this observation, we experimentally verified that the expression levels of a circRNA derived from an MS-associated locus, i.e., hsa_circ_0043813 from the <em>STAT3</em> gene, can be modulated by the three genotypes at the disease-associated SNP. Finally, by evaluating RNA-seq data of two cell lines, SH-SY5Y and Jurkat cells, representing tissues relevant for MS, we identified 18 (two novel) circRNAs derived from MS-associated genes. In conclusion, this work showed for the first time that MS-GWAS top hits map in LD blocks enriched in circRNAs, suggesting circRNAs as possible novel contributors to the disease pathogenesis.</p> <p>GEO database</p> <p>URL: <a href="https://www.ncbi.nlm.nih.gov/geo/">https://www.ncbi.nlm.nih.gov/geo/</a></p> <p>Numero di accesso del dataset: GSE110525</p>
Genetic analysis of mycobacteria isolated from suspect bovine tuberculosis lesions in Wolaita, Ethiopia - code and datasets.
<p>Bovine tuberculosis (bTB), caused by Mycobacterium bovis and other members of the Mycobacterium tuberculosis complex (MTBC), is a significant concern for livestock and public health in Ethiopia. This study aimed to assess the prevalence and causative agents of bTB in cattle from four abattoirs in the Wolaita region of Ethiopia. </p>
Supporting data and code for: Demographic and genetic impacts of powdery mildew in a young oak cohort
<p>This is a new release following the submission of the related PCI recommended manuscript to the <em>Annals of Forest Science</em> journal. It contains the necessary scripts to produce most of the analyses and figures of the manuscript. Apart from minor modifications following the recommendation in '<em>PCI Forest and Wood Sciences</em>', the main change is the addition of an extra dataset "Data_S2.txt" to the additional datasets. This dataset was previously included as a table in the 'supplementary material' file.</p>
Data and code for: Plastic and quantitative genetic divergence mirror environmental gradients among wild, fragmented populations of Impatiens capensis
<p><strong>Premise of the study:</strong> Habitat fragmentation generates molecular genetic divergence among isolated populations but few studies have assessed phenotypic divergence and fitness in populations where the genetic consequences of habitat fragmentation are known. Phenotypic divergence could reflect plasticity, local adaptation, and/or genetic drift.</p> <p><strong>Methods:</strong> We examined patterns and potential drivers of phenotypic divergence among 12 populations of jewelweed (<em>Impatiens capensis </em>Meerb.) that show strong molecular genetic signals of isolation and drift among fragmented habitats. We measured morphological and reproductive traits in both maternal plants within natural populations and their self-fertilized progeny grown together in a common garden. We also quantified environmental divergence between home sites and the common garden.</p> <p><strong>Key results: </strong>Populations with less molecular genetic variation expressed less maternal phenotypic variation. Progeny in the common garden converged in phenotypes relative to their wild mothers but retained among-population differences in morphology, survival, and reproduction. Among-population phenotypic variance was 3-10x greater in home sites than in the common garden for 6 of 7 morphological traits measured. Patterns of phenotypic divergence paralleled environmental gradients in ways suggestive of adaptation. Progeny resembled their mothers less as the environmental distance between their home site and the common garden increased.</p> <p><strong>Conclusions: </strong>Despite strong molecular signatures of isolation and drift, phenotypic differences among these <em>Impatiens </em>populations appear to reflect both adaptive quantitative genetic divergence and plasticity. Quantifying the extent of local adaptation and plasticity and how these covary with molecular and phenotypic variation help us predict when populations may lose their adaptive capacity. </p>
Supplementary data and code to "Known allosteric proteins have central roles in genetic disease" by G. Abrusan, D. Ascher and M. Inouye, PLOS Computational Biology 18(2):e1009806.
<p>Scripts and data to reproduce the figures and supplementary figures of "G. Abrusan, D. Ascher and M. Inouye (2022) Known allosteric proteins have central roles in genetic disease." PLOS Computational Biology 18(2):e1009806. https://doi.org/10.1371/journal.pcbi.1009806.</p>
Code and initial metapopulation data for model construction and simulation analyses for: Genetic rescue from protected areas is modulated by migration, hunting rate and timing of harvest
<p>Migrants from protected areas may buffer the risk of harvest-induced evolutionary changes in exploited populations that face strong selective harvest pressures in both terrestrial and marine ecosystems. Understanding the mechanisms favouring genetic rescue through migration could help ensure sustainable harvest outside protected areas and conserve genetic diversity inside those areas. We developed a stochastic individual-based metapopulation model to evaluate the potential for migration from protected areas to mitigate the evolutionary consequences of selective harvest. We parameterized the model with detailed data from individual monitoring of two populations of bighorn sheep subjected to trophy hunting. We tracked horn length through time in a metapopulation including large protected and trophy-hunted populations connected through male breeding migrations. We quantified and compared declines in horn length and rescue potential under various combinations of migration rate, hunting rate in hunted areas and temporal overlap in timing of harvest and migrations, which affects the migrants' survival and chances to breed within exploited areas. Our simulations suggest that the effects of size-selective harvest on male horn length in hunted populations can be dampened or avoided if harvest pressure is low, migration rate is substantial, and migrants have a low risk of being shot. Intense size-selective harvest impacts the phenotypic and genetic diversity in horn length, and population structure through changes in proportions of large-horned males, sex ratio and age structure. When hunting pressure is high and overlaps with male migrations, effects of selective removal also emerge in the protected population, so that instead of a genetic rescue of hunted populations, our model predicts undesirable effects inside protected areas. Our results stress the importance of a metapopulational approach to management, to promote genetic rescue from protected areas and limit ecological and evolutionary impacts of harvest on both harvested and protected populations.</p>
Data and Code for Publication "Inferring human neutral genetic variation from craniodental phenotypes"
<p>Data and code for publication: H. Rathmann et al., Inferring human neutral genetic variation from craniodental phenotypes. PNAS Nexus.</p> <p>The repository contains:</p> <ul> <li>“<em>R code for DP-DG analysis.txt</em>”: R code for testing levels of neutral evolutionary signals preserved in five craniodental data types: cranial metrics, dental metrics, cranial non-metric traits, dental non-metric traits, and craniodental metrics and non-metric traits combined.</li> </ul> <ul> <li>“<em>Cranial metric data.csv</em>”: Dataset consisting of 37 cranial metric variables for 26 worldwide modern populations, provided in a comma-separated values file format. The data were collected by T. Hanihara and originally presented in the publication titled: T. Hanihara, Comparison of craniofacial features of major human groups. <em>Am. J. Phys. Anthropol.</em> 99, 389–412 (1996) (<a href="https://doi.org/10.1002/(SICI)1096-8644(199603)99:3%3c389::AID-AJPA3%3e3.0.CO;2-S">https://doi.org/10.1002/(SICI)1096-8644(199603)99:3<389::AID-AJPA3>3.0.CO;2-S</a>).</li> </ul> <ul> <li>“<em>Dental metric data.csv</em>”: Dataset comprising 28 dental metric variables for 26 worldwide modern populations, provided in a comma-separated values file format. The data were collected by T. Hanihara and originally presented in the publication titled: T. Hanihara, H. Ishida, Metric dental variation of major human populations. <em>Am. J. Phys. Anthropol.</em> 128, 287–298 (2005) (<a href="https://doi.org/10.1002/ajpa.20080">https://doi.org/10.1002/ajpa.20080</a>).</li> </ul> <ul> <li>“<em>Cranial non-metric trait data.csv</em>”: Dataset consisting of 24 cranial non-metric trait variables for 26 worldwide modern populations, provided in a comma-separated values file format. The data were collected for the most part by T. Hanihara and presented in the publication titled: T. Hanihara, H. Ishida, Y. Dodo, Characterization of biological diversity through analysis of discrete cranial traits. <em>Am. J. Phys. Anthropol.</em> 121, 241–251 (2003) (<a href="https://doi.org/10.1002/ajpa.10233">https://doi.org/10.1002/ajpa.10233</a>).</li> </ul> <ul> <li>“<em>Dental non-metric trait data.csv</em>”: Dataset comprising 25 dental non-metric trait variables for 26 worldwide modern populations, provided in a comma-separated values file format. The data were collected by C. G. Turner II, G. R. Scott, and J. D. Irish. This individual-level dataset was artificially created from population-level trait frequency information presented in the publications: G. R. Scott, J. D. Irish, <em>Human Tooth Crown and Root Morphology </em>(Cambridge University Press, 2017) (<a href="https://doi.org/10.1017/9781316156629">https://doi.org/10.1017/9781316156629</a>); and: J. D. Irish, A. Morez, L. Girdland Flink, E. L. W. Phillips, G. R. Scott, Do dental nonmetric traits actually work as proxies for neutral genomic data? Some answers from continental- and global-level analyses. <em>Am. J. Phys. Anthropol. </em>172, 347–375 (2020) (<a href="https://doi.org/10.1002/ajpa.24052">https://doi.org/10.1002/ajpa.24052</a>).</li> </ul> <ul> <li>“<em>SNP data.txt</em>”: Dataset comprising 8,821 SNP markers for 26 worldwide modern populations, provided in a genepop file format. The data were obtained from various published sources: I. Lazaridis et al., Ancient human genomes suggest three ancestral populations for present-day Europeans. <em>Nature </em>513, 409–413 (2014) (<a href="https://doi.org/10.1038/nature13673">https://doi.org/10.1038/nature13673</a>); P. Qin, M. Stoneking, Denisovan ancestry in east Eurasian and native American populations. <em>Mol. Biol. Evol. </em>32, 2665–2674 (2015) (<a href="https://doi.org/10.1093/molbev/msv141">https://doi.org/10.1093/molbev/msv141</a>); P. Skoglund et al., Genomic insights into the peopling of the Southwest Pacific. <em>Nature </em>538, 510–513 (2016) (<a href="https://doi.org/10.1038/nature19844">https://doi.org/10.1038/nature19844</a>); M. R. Nelson et al., The Population Reference Sample, POPRES: a resource for population, disease, and pharmacological genetics research. <em>Am. J. Hum. Genet. </em>83, 347–358 (2008) (<a href="https://doi.org/10.1016/j.ajhg.2008.08.005">https://doi.org/10.1016/j.ajhg.2008.08.005</a>); J. K. Pickrell, J. K. Pritchard, Inference of population splits and mixtures from genome-wide allele frequency data. <em>PLoS Genet. </em>8, e1002967 (2012) (<a href="https://doi.org/10.1371/journal.pgen.1002967">https://doi.org/10.1371/journal.pgen.1002967</a>); A. Bergström et al., Insights into human genetic variation and population history from 929 diverse genomes. <em>Science </em>367 (2020) (<a href="https://doi.org/10.1126/science.aay5012">https://doi.org/10.1126/science.aay5012</a>); B. M. Henn et al., Genomic ancestry of North Africans supports back-to-Africa migrations. <em>PLoS Genet. </em>8, e1002397 (2012) (<a href="https://doi.org/10.1371/journal.pgen.1002397">https://doi.org/10.1371/journal.pgen.1002397</a>); S. Mallick et al., The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. <em>Nature </em>538, 201–206 (2016) (<a href="https://doi.org/10.1038/nature18964">https://doi.org/10.1038/nature18964</a>); Lao et al., Correlation between genetic and geographic structure in Europe. <em>Curr. Biol. </em>18, 1241–1248 (2008) (<a href="https://doi.org/10.1016/j.cub.2008.07.049">https://doi.org/10.1016/j.cub.2008.07.049</a>); and M. Lipson et al., Population Turnover in Remote Oceania Shortly after Initial Settlement. <em>Curr. Biol. </em>28, 1157-1165.e7 (2018) (<a href="https://doi.org/10.1016/j.cub.2018.02.051">https://doi.org/10.1016/j.cub.2018.02.051</a>).</li> </ul> <p>For population and variable names and abbreviations, see Supplementary Information in: H. Rathmann et al., Inferring human neutral genetic variation from craniodental phenotypes. PNAS Nexus.</p>
Data and codes from "How does dispersal shape the genetic structure of animal populations in European cities? A simulation approach"
<p>Codes and data used for "Savary et al. How does dispersal shape the genetic structure of animal populations in European cities? A simulation approach".</p> <p> </p>
Supplementary data files for Manzano-Marín et. al. 2023 "Evolution of an alternative genetic code in the Providencia symbiont of the haematophagous leech Haementeria acuecueyetzin"
<p>The data set consists of siz folders:</p> <p><strong>1)</strong> "genome_data": GenBank-formatted annotation files for newly sequenced <em>Providencia siddallii</em> endosymbionts.</p> <p><strong>2)</strong> "orthoMCL_data": Flat-text output files from the OrthoMCL pipeline.</p> <p><strong>3)</strong> "phylongey": MrBayes run files for <em>Providencia</em> spp. Bayesian phylogenetic inference.</p> <p><strong>4)</strong> "UGA_proteins_and_genes": FASTA-formatted alignments of UGA-containing protein-coding gene sequences from figure 3 and table 3.</p> <p><strong>5)</strong> "RepeatModeler_Prsiddallii": Log files for RepeatModeler runs of <em>P. siddallii</em> genomes.</p> <p><strong>6)</strong> "checkM2_Psiddallii": checkM2 input and output files for <em>P. siddallii</em> protein sets.</p> <p><strong>7)</strong> "breseq_runs": breseq output folders for variant calling on newly assembled genomes.</p> <p><strong>8)</strong> "GSAlign_Prsiddallii_GTOCOR": output files of variant calling using GSAlign between <em>P. siddallii</em> strain GTOCOR1 (sampled in 2015 and reported in Manzano-Marín <em>et. al.</em> 2015 <em>GBE</em>) and GTOCOR2 (sampled in 2019 and reported in the associated work).</p>
Data, code, and supplementary materials for Pearman P. B., Broennimann, O., et al. Monitoring species genetic diversity in Europe varies greatly and overlooks potential climate change impacts. Nature Ecology & Evolution
<p>The repository contains several archives of digital materials that were used and/or produced in the analyses presented in Pearman, P. B. and Broennimann et al. Monitoring species genetic diversity in Europe varies greatly and overlooks potential climate change impacts. <strong>Nature Ecology & Evolution</strong>, likely 2023. These archives include (1) Supplementary Materials files ; (2) Data and code to generate country-level maps and plots; and (3) data and code to generate all maps of species and joint climate niche marginality, all in G-zipped tar archives. Readme files are available in each archive to guide running of the scripts and identification of objects in the Supplementary Materials. Please see the paper for all co-authors names, and the methods, the results obtained, and discussion of their implications.</p> <p>This work is dedicated to the memory of our friend and colleague Michael Bruford (1963-2023).</p>
Data and code for: Plastic and quantitative genetic divergence mirror environmental gradients among wild, fragmented populations of Impatiens capensis
Open the record for dataset details and reuse information.
Data and code for: Realized genetic gains via recurrent selection in a tropical maize haploid inducer population and optimizing simultaneous selection for the next cycles
Open the record for dataset details and reuse information.
Data and code from: Evaluating genomic offset predictions in a forest tree with high population genetic structure
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.