Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5,153

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5,153 results for “Genetic data”

Learn how ShareScore rates datasets ↗
zenodo52/100

Frictionless Tabular Data Package for GC-MS Rose scent profile data for Data published in Nature genetics, June, 2018 & Science, July 2015

<p>This dataset, in the form of a Frictionless Tabular Data Package (https://frictionlessdata.io/specs/tabular-data-package/), holds the measurements of 35 known metabolites(all annotated with resolvable CHEBI identifiers and InChi strings), measured by gas chromatography mass-spectrometry (GC-MS) in one Rose cultivars (all annotated with resolvable NCBITaxonomy Identifiers) and one organism part (annotated with resolvable Plant Ontology identifiers). The quantitation types are annotated with resolvable STATO terms. The measurements over these metabolites, which were made in 2 distinct experiments, were extracted from: a supplementary material table, available from https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip and published alongside the Nature Genetics manuscript identified by the following doi: https://doi.org/10.1038/s41588-018-0110-3, published in June 2018 a supplementary material table available as a pdf from &#39;Biosynthesis of monoterpene scent compounds in roses&#39; by Magnard et al, Science 03 Jul 2015 identified by the following doi: https://doi.org/10.1126/science.aab0696. This dataset is used to demonstrate how to make data Findable, Accessible, Discoverable and Interoperable (FAIR)and how Frictionless Tabular Data Package representations can be easily mobilised for reanalysis and data science.It is associated to the following project: https://github.com/proccaserra/rose2018ng-notebook with all the necessaryinformation, executable code and tutorials in the form of Jupyter notebooks.</p>

opencc-by-4.0Apr 2019View details →
zenodo48/100

Data from: Visual pigment chromophore usage in Nicaraguan Midas cichlids: Phenotypic plasticity and genetic assimilation of cyp27c1 expression

<p>Code and Data associated with "Visual pigment chromophore usage in Nicaraguan Midas cichlids: Phenotypic plasticity and genetic assimilation of&nbsp;<em>cyp27c1</em> expression"</p> <h2><span>Abstract</span></h2> <p><span>The wide-ranging photic conditions found across aquatic habitats may act as selective pressures potentially driving rapid evolution and diversity in the visual system of teleost fishes. Fine-tuning of visual sensitivities in many fish species relies on regulating the two components of visual pigments, the opsin protein and the chromophore. Many studies have focused on opsin gene expression or opsin sequence divergence in fishes inhabiting contrasting habitats. However, variation in chromophore usage across photic habitats has received less attention. Species from the Nicaraguan Midas cichlid complex, <em>Amphilophus </em>cf <em>citrinellus </em>[G&uuml;nther 1864], have independently colonized seven isolated crater lakes of varying photic conditions resulting in repeated examples of small adaptive radiations. Here, we investigate variation in <em>cyp27c1</em>, the main enzyme involved in chromophore exchange, in response to photic environments in the wild, we measure its genetic component using laboratory-reared fish and test the effect of different rearing light conditions on <em>cyp27c1</em> expression. We found that photic environments significantly predict variation in <em>cyp27c1</em> expression in wild populations and that this variation seems to be genetically assimilated in two populations. We found that light-induced <em>cyp27c1</em> expression is variable across populations (i.e., genotype-by-environment interactions) and correlated with local photic conditions thus highlighting <em>cyp27c1</em> as a key factor of visual ecology in cichlid fishes.</span></p> <p><span>Keywords: <em>cyp27c1 </em>gene expression, sensory ecology, visual plasticity, Neotropical cichlids </span></p>

opencc-by-4.0Mar 2024View details →
zenodo48/100

Data from 'Tracability of Forest Reproductive Material with the quality label 'Plant van Hier': A DNA database with genetic profiles of native autochthonous tree and shrub species of Flanders, Belgium'

<h2>Background</h2> <p>Indigenous trees and shrubs play an important role in multifunctional forest management. They form a significant part of the biodiversity in our forests. Forest reproductive material (FRM) of autochthonous Flemish origin is sold under the quality label &lsquo;Plant van Hier&rsquo;, a certification mark of the Agency for Nature and Forests. To ensure the provenance of the seedlings, we developed a DNA-database of genetic profiles of potential parent trees, using species-specific genetic markers. This database enables the traceability of FRM of the &lsquo;Plant van Hier&rsquo; label throughout the entire production chain; from seed harvesting and cultivation to planting by the end user.</p> <p>This database contains the genetic profiles of almost all possible parent trees present within 27 Flemish autochthonous seed orchards of eight ecologically important tree and shrub species: <em>Carpinus betulus</em>, <em>Corylus avellana</em>, <em>Frangula alnus</em>, <em>Populus tremula</em>, <em>Sorbus aucuparia</em>, <em>Tilia cordata</em>, <em>Tilia platyphyllos,</em> and <em>Ulmus laevis</em>. The profiles were established using microsatellite markers (11 to 24 markers per species).&nbsp;&nbsp;New genetic markers were developed for&nbsp;<em>Carpinus betulus</em> and <em>Ulmus laevis</em>. PCR products were run on an ABI 3500 Genetic Analyser (Thermo Fisher Scientific).</p> <h2>Files</h2> <p>The files will be updated when new genotypes are added to the seed orchards. The current data files contain data from genotypes collected in the period 2018-2023.&nbsp;</p> <h3>Species_genotypes</h3> <p>These files contain the genetic fingerprints of the parent trees of autochthonous Flemish seed orchards. Missing data is indicated as &lsquo;MD&rsquo;. For <em>Carpinus betulus</em>, an octoploid species, the allelic phenotype is given instead of the genotype as the number of times that an allele occurs on a specific locus is not known.</p> <p>The next metadata is additionally given:<br>- Species: the Latin name of the species<br>- Seed_orchard: the name of the seed orchard in which the genotypes are located<br>- Code_seed_orchard: the code of the seed orchard in which the genotypes are located as given in the Register of Flemish Forest Reproductive Material (&lsquo;Register bosbouwkundig uitgangsmateriaal&rsquo;; inbo.be)<br>- Genotype: the fieldname given to the genotype<br>- Origin: the location where the genotype was collected in Flanders, Belgium. Genotypes were collected from natural stands which are assumed to have an autochthonous origin. When the specific location is unknown, the location &lsquo;Flanders&rsquo; is given.&nbsp;<br>- Year_sampled: the year in which the genotypes were sampled in the respective seed orchard for genetic analysis.</p> <h3>Species_binsets</h3> <p>These files contain the binsets and allele names that are used to score the alleles of the genotypes in the programme Geneious Prime 2019.3.2 (<a href="https://www.geneious.com">https://www.geneious.com</a>). For <em>Tilia platyphyllos </em>and <em>Tilia cordata</em>, the same binsets were used.</p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Data for: Increasing plant group productivity through latent genetic variation for cooperation

<p>Historic yield advances in the major crops have to a large extent been achieved by selection for improved productivity of groups of plant individuals such as high-density stands. Research suggests that such improved group productivity depends on &ldquo;cooperative&rdquo; traits (e.g., erect leaves, short stems) that &ndash; while beneficial to the group &ndash; decrease individual fitness under competition. This poses a problem for some traditional breeding approaches, especially when selection occurs at the level of individuals, because &ldquo;selfish&rdquo; traits will be selected for and reduce yield in high-density monocultures. One approach, therefore, has been to select individuals based on ideotypes with traits expected to promote group productivity. However, this approach is limited to architectural and physiological traits whose effects on growth and competition are relatively easy to anticipate.</p> <p>Here, we developed a general and simple method for the discovery of alleles promoting cooperation in plant stands. Our method is based on the game-theoretical premise that alleles increasing cooperation benefit the monoculture group but are disadvantageous to the individual when facing non-cooperative neighbors. Testing the approach using the model plant <em>Arabidopsis thaliana</em><em>, </em>we found a major effect locus where the rarer allele was associated with increased cooperation and productivity in high-density stands. The allele likely affects a pleiotropic gene, since we find that it is also associated with reduced root competition but higher resistance against disease. Thus, even though cooperation is considered evolutionarily unstable except under special circumstances, conflicting selective forces acting on a pleiotropic gene might maintain latent genetic variation for cooperation in nature. Such variation, once identified in a crop, could rapidly be leveraged in modern breeding programs and provide efficient routes to increase yields.</p>

opencc-by-4.0May 2019View details →
zenodo48/100

Using machine learning to integrate genetic and environmental data to model genotype-by-environment interactions

<p>Files generated from the study described in&nbsp;<a href="https://doi.org/10.1101/2024.02.08.579534">Fernandes et. al (2024)</a> .</p> <p>The file "cvs_h2s.csv" comprises the coefficient of variation and the Cullis heritability for each environment.</p> <p>The file "all_predictions.csv" contains the predictions from all the models evaluated, in different cross-validation (CV) scenarios.</p> <p>The file "coincidence_index.csv" has the Coincidence Index (CI) for each CV and models evaluated in our study.</p> <p>Our study used the multi-environment maize yield trials data from the Genomes to Fields 2022 initiative (<a href="https://doi.org/10.1186/s13104-023-06421-z">Lima et. al 2024</a>).</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Data from Neutral genetic structuring of pathogen populations during rapid adaptation

<p><strong>Datasets and temporary dataframes relating to the article "Neutral genetic structuring of pathogen populations during rapid adaptation".</strong></p> <p>These datasets and temporary dataframes are necessary to run the scripts from the public GitLab repository: <a href="https://gitlab.com/saubin.meline/neutral-genetic-structuring-adaptation">https://gitlab.com/saubin.meline/neutral-genetic-structuring-adaptation</a>. Please refer to this public GitLab repository for the latest version of the codes and to perform all analyses presented in the article.</p> <p>Original datasets from the demogenetic model:</p> <ul> <li>Output_RandomDesign.txt</li> <li>Output_RegularDesign_With_host_alternation.txt</li> <li>Output_RegularDesign_Without_host_alternation.txt</li> <li>Output_RandomDesign_Mnull_Medoid_With_host_alternation.txt</li> <li>Output_RandomDesign_Mnull_Medoid_Without_host_alternation.txt</li> </ul> <p>All remaining files correspond to temporary dataframes generated by the scripts in the GitLab repository, provided here for reproducibility of the results and to save time at certain time-consuming scripts.</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

Frictionless Tabular Data Package for GC-MS data from the 'Rose Genome' article published in Nature genetics, June, 2018

<p>This dataset, in the form&nbsp;of a Frictionless Tabular Data Package (<a href="https://frictionlessdata.io/specs/tabular-data-package/">https://frictionlessdata.io/specs/tabular-data-package/)</a>, holds the measurements of 61&nbsp;known metabolites (all annotated with resolvable CHEBI identifiers and InChi strings), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with&nbsp;resolvable NCBITaxonomy Identifiers) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The quantitation types are annotated with resolvable&nbsp;<a href="https://github.com/ISA-tools/stato">STATO</a> terms. &nbsp;</p> <p>The data was extracted from a supplementary material table,&nbsp;available from&nbsp;<a href="https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip">https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip</a>&nbsp; and published alongside the Nature Genetics manuscript identified by the following doi:&nbsp;<a href="https://doi.org/10.1038/s41588-018-0110-3">https://doi.org/10.1038/s41588-018-0110-3</a>, published in June 2018. This supplementary material table was deposited to Zenodo and is identified by the following doi: <a href="https://doi.org/10.5281/zenodo.2598799">https://doi.org/10.5281/zenodo.2598799</a></p> <p>This dataset is used to demonstrate how to make data Findable, Accessible, Discoverable and Interoperable (FAIR) and how Frictionless Tabular Data Package representations can be easily mobilised for reanalysis and data science.</p> <p>It is associated to the following project: <a href="https://github.com/proccaserra/rose2018ng-notebook">https://github.com/proccaserra/rose2018ng-notebook</a>&nbsp;with&nbsp;all the necessary information, executable code&nbsp;and tutorials in the form of Jupyter notebooks.</p>

opencc-by-4.0Feb 2019View details →
zenodo48/100

Frictionless Tabular Data Package for GC-MS Rose scent profile data for Data published in Nature genetics, June, 2018 & Science, July 2015

<p>This dataset, in the form&nbsp;of a Frictionless Tabular Data Package (<a href="https://frictionlessdata.io/specs/tabular-data-package/">https://frictionlessdata.io/specs/tabular-data-package/)</a>, holds the measurements of 61&nbsp;known metabolites (all annotated with resolvable CHEBI identifiers and InChi strings), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with&nbsp;resolvable NCBITaxonomy Identifiers) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The quantitation types are annotated with resolvable&nbsp;<a href="https://github.com/ISA-tools/stato">STATO</a>&nbsp;terms. &nbsp;</p> <p>The data were extracted from:</p> <ul> <li>a supplementary material table,&nbsp;available from&nbsp;<a href="https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip">https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip</a>&nbsp; and published alongside the Nature Genetics manuscript identified by the following doi:&nbsp;<a href="https://doi.org/10.1038/s41588-018-0110-3">https://doi.org/10.1038/s41588-018-0110-3</a>, published in June 2018</li> <li>a supplementary material table available as a pdf from &quot;Biosynthesis of monoterpene scent compounds in roses&quot; by Magnard et al, Science&nbsp;&nbsp;03 Jul 2015 identified by the following doi: <a href="https://doi.org/10.1126/science.aab0696">https://doi.org/10.1126/science.aab0696</a></li> </ul> <p>This dataset is used to demonstrate how to make data Findable, Accessible, Discoverable and Interoperable (FAIR) and how Frictionless Tabular Data Package representations can be easily mobilised for reanalysis and data science.</p> <p>It is associated to the following project:&nbsp;<a href="https://github.com/proccaserra/rose2018ng-notebook">https://github.com/proccaserra/rose2018ng-notebook</a>&nbsp;with&nbsp;all the necessary information, executable code&nbsp;and tutorials in the form of Jupyter notebooks.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2019View details →
zenodo48/100

Frictionless Tabular data package for GC-MS data from Rose Genome article published in Nature genetics, June, 2018

<p>This dataset, in the form of a Frictionless Tabular Data Package (https://frictionlessdata.io/specs/tabular-data-package/), holds the measurements of 61 known metabolites (all annotated with resolvable CHEBI identifiers and InChi), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with resolvable NCBITaxId) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The data was extracted from a supplementary material table, available from https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip and published alongside the Nature Genetics manuscript identified by the following doi: https://doi.org/10.1038/s41588-018-0110-3, published in June 2018. This dataset is used to demonstrate how to make data Findeable, Accessible, Discoverable and Interoperable(FAIR) and how Tabular Data Package representations can be easily mobilized for re-analysis and data science. It is associated to the following project available from github at: https://github.com/proccaserra/rose2018ng-notebook with all necessary information and Jupyter notebooks.</p>

opencc-by-4.0Feb 2019View details →
zenodo48/100

Supplementary dataset to publication: Oxford nanopore technologies - a valuable tool to generate whole-genome sequencing data for in silico serotyping and the detection of genetic markers in Salmonella, Thomas et al 2023

<p>Bacteria of the genus&nbsp;<em>Salmonella</em>&nbsp;pose a major risk to livestock, the food economy, and public health.&nbsp;<em>Salmonella</em>&nbsp;infections are one of the leading causes of food poisoning. The identification of serovars of&nbsp;<em>Salmonella</em>&nbsp;achieved by their diverse surface antigens is essential to gain information on their epidemiological context. Traditionally, slide agglutination has been used for serotyping. In recent years, whole-genome sequencing (WGS) followed by&nbsp;<em>in silico</em>&nbsp;serotyping has been established as an alternative method for serotyping and the detection of genetic markers for&nbsp;<em>Salmonella</em>. Until now, WGS data generated with Illumina sequencing are used to validate&nbsp;<em>in silico</em>&nbsp;serotyping methods. Oxford Nanopore Technologies (ONT) opens the possibility to sequence ultra-long reads and has frequently been used for bacterial sequencing. In this study, ONT sequencing data of 28&nbsp;<em>Salmonella</em>&nbsp;strains of different serovars with epidemiological relevance in humans, food, and animals were taken to investigate the performance of the&nbsp;<em>in silico</em>&nbsp;serotyping tools SISTR and SeqSero2 compared to traditional slide agglutination tests. Moreover, the detection of genetic markers for resistance against antimicrobial agents, virulence, and plasmids was studied by comparing WGS data based on ONT with WGS data based on Illumina. Based on the ONT data from flow cell version R9.4.1,&nbsp;<em>in silico</em>&nbsp;serotyping achieved an accuracy of 96.4 and 92% for the tools SISTR and SeqSero2, respectively. Highly similar sets of genetic markers comparing both sequencing technologies were identified. Taking the ongoing improvement of basecalling and flow cells into account, ONT data can be used for&nbsp;<em>Salmonella in silico</em> serotyping and genetic marker detection.</p>

opencc-by-4.0Feb 2023View details →
zenodo48/100

Data to reproduce analysis in "Systematic analysis of transcriptional and epigenetic effects of genetic variation in Kupffer cells enables discrimination of cell intrinsic and environment-dependent mechanisms"

<p>Here you can find the datasets necessary to reproduce all analyses described in the Glass lab paper by <a href="https://www.biorxiv.org/content/10.1101/2022.09.22.509046v1">Bennett et al</a>. The python and R code for reproducing analysis and figures can be found on our linked&nbsp;<a href="https://github.com/HunterBennett/KupfferCell_NaturalGeneticVariation">github repository.</a></p> <p>Briefly, this paper explores the effect of natural genetic variation&nbsp;<em>in vivo</em>, using Kupffer cells as a model cell type. We collect and analyze transcriptional and epigenetic data (ATAC-seq, H3K27Ac ChIP-seq) to identify putative&nbsp;<em>trans</em>&nbsp;regulators driving differential gene expression across inbred strains of mice. Additionally, we provide evidence that&nbsp;<em>trans</em>&nbsp;effects control a majority of strain differential genes at homeostasis while&nbsp;<em>cis</em>&nbsp;effects dominate the transcriptional response to an external signal (lipopolysaccharide).</p> <p>References:</p> <p>Hunter Bennett, Ty D. Troutman, Enchen Zhou, Nathanael J. Spann, Verena M. Link, Jason S. Seidman, Christian K. Nickl, Yohei Abe, Mashito Sakai, Martina P. Pasillas, Justin M. Marlman, Carlos Guzman, Mojgan Hosseini, Bernd Schnabl, Christopher K. Glass bioRxiv 2022.09.22.509046; doi:&nbsp;<a href="https://doi.org/10.1101/2022.09.22.509046">https://doi.org/10.1101/2022.09.22.509046</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo48/100

Supplementary data: Agro-morphological and molecular characterization reveal deep insights in promising genetic diversity and marker-trait associations in Fagopyrum esculentum and F. tataricum

<p>Our study focuses on the global/European buckwheat germplasm collected as part of the ECOBREDD project. The potential of this highly diverse collection for organic buckwheat breeding was evaluated at two complementary levels: phenotypic and genetic. Here, we characterized the phenotypic and genetic diversity of a global collection of the two cultivated buckwheat species <em>Fagopyrum esculentum</em> and <em>F. tataricum</em> (190 and 51 accessions, respectively) using 37 agro-morphological traits and 24 SSR markers (Simple Sequence Repeats) (see publication and info sheet of the data).</p>

opencc-by-4.0Jun 2023View details →
edi48/100

Delta Smelt (Hypomesus transpacificus) biomarker and genetic data from supplemental release into the San Francisco Estuary, 2022

Delta Smelt (Hypomesus transpacificus) is an endangered fish that is endemic to the San Francisco Estuary. As a conservation strategy, hatchery-reared Delta Smelt have been released into the San Francisco Estuary to supplement the wild population. State and federal agencies surveyed the abundance of Delta Smelt, collected fish specimens, and recorded associated environmental data from sampling sites. Delta Smelt specimens were preserved and transported to the University of California, Davis where a variety of biomarkers were assessed on individual fish. This project, including the supplementation of hatchery-reared Delta Smelt and data collection, is ongoing.

openCC (other)Aug 2024View details →
zenodo44/100

Data from: The genetic legacy of extreme exploitation in a polar vertebrate

<p>Microsatellite data (39 loci) from Antarctic fur seals and Subantarctic fur seals, used in the paper: &quot;The genetic legacy of extreme exploitation in a polar vertebrate&quot;</p> <p>&nbsp;</p> <p><strong>Abstract</strong></p> <p>Understanding the effects of human exploitation on the genetic composition of wild populations is important for predicting species persistence and adaptive potential.&nbsp; We therefore investigated the genetic legacy of large-scale commercial harvesting by reconstructing on a global scale the recent demographic history of the Antarctic fur seal (<em>Arctocephalus gazella</em>), a species that was hunted to the brink of extinction by 18<sup>th</sup> and 19<sup>th</sup> century sealers.&nbsp; Molecular genetic data from over 2,000 individuals, sampled from all eight major breeding colonies across the species᾿ circumpolar geographic distribution, show that at least four relict populations around Antarctica survived commercial hunting.&nbsp; Coalescent simulations suggest that all of these populations experienced severe bottlenecks down to effective population sizes of around 150&ndash;200.&nbsp; Nevertheless, comparably high levels of neutral genetic variability were retained as these declines are unlikely to have been strong enough to deplete allelic richness by more than around 15%.&nbsp; These findings suggest that even dramatic short-term declines need not necessarily result in major losses of diversity, and explain the apparent contradiction between the high genetic diversity of this species and its extreme exploitation history.</p> <p>&nbsp;</p> <p><strong>Funding</strong></p> <p>This research was supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) in<br> the framework of a Sonderforschungsbereich (project numbers 316099922 and 396774617&ndash;TRR 212) and the<br> priority programme &quot;Antarctic Research with Comparative Investigations in Arctic Ice Areas&quot; SPP 1158 (project<br> number 424119118). It was also funded by Norwegian Antarctic Research Expeditions (NARE) programme.<br> This work contributes to the Ecosystems project of the British Antarctic Survey, Natural Environmental Research<br> Council, and is part of the Polar Science for Planet Earth Programme. The Department of Environmental Affairs<br> provided logistical support for research at Marion Island and the Department of Science and Technology of<br> South Africa provided funding through the National Research Foundation (NRF). We are grateful to Caroline<br> Bonin, Debbie Baird-Bower and Iain Staniland together with the seal biologists working within the Marion<br> Island Marine Mammal Programme for sample collection and logistics. We acknowledge support for the Article<br> Processing Charge by the Deutsche Forschungsgemeinschaft and the Open Access Publication Fund of Bielefeld<br> University.</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Data from: Genetic admixture increases phenotypic diversity in the nectar yeast Metschnikowia reukaufii,

<p>Raw data and supplementary files for the manuscript &quot;Genetic admixture increases phenotypic diversity in the nectar yeast <em>Metschnikowia reukaufii</em>.&quot;</p> <p>-------------------</p> <p><strong>Table S5.xlsx </strong>-- Pairwise correlations between phenotypic traits of <em>Metschnikowia reukaufii</em>.</p> <p><strong>Table S6.xlsx</strong>&nbsp;--&nbsp;Detailed results obtained in tests of phylogenetic signal for different phenotypic traits and indices of overall performance of <em>Metschnikowia reukaufii</em>.</p> <p><strong>Table S7.xlsx</strong>&nbsp;--&nbsp;Detailed model fitting results obtained for phenotypic traits and indices of overall performance of <em>Metschnikowia reukaufii</em>.</p> <p><strong>mronlyvcf-renamed.vcf</strong> -- High coverage SNPs obtained from whole genome mapping of 73 <em>Metschnikowia reukaufii</em> strains to diploid reference (mean coverage = 47.9&times;, range 23 &ndash; 116&times;).</p> <p><strong>MR_phenotypes.xlsx</strong>&nbsp;-- Phenotypic data obtained for 73 <em>Metschnikowia reukaufii</em> strains.</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

Phenotypic data related to genetic architecture of transmission stage production and virulence in schistosome parasites

<p>These data were generated related to the study of the <strong>Genetic architecture of transmission stage production and virulence in schistosome parasites</strong>.</p> <p><strong>Abstract:</strong> Both theory and experimental data from multiple pathogens suggest that the production of transmission stages should be strongly associated with virulence, but the genetic bases of parasite transmission/virulence traits are poorly understood. In the blood fluke <em>Schistosoma mansoni</em>, parasite genotypes show extensive variation in numbers of cercariae larvae shed from infected snails. Furthermore, high shedding parasites cause high mortality to snails while low shedding parasites cause low mortality, consistent with expected trade-offs between parasite transmission and virulence. To understand the genetic basis of transmission stage production/virulence, we conducted reciprocal crosses between schistosomes from two laboratory populations that differ 8-fold in cercarial shedding and in their virulence to inbred snail hosts. Each parasite generation, we determined four-week cercarial shedding profiles in inbred <em>Biomphalaria glabrata</em> snails infected with single parasite larvae. We sequenced the whole genome of the F0 parents and the exome of the F1 progeny and 188 F2 progeny from each cross, and used linkage mapping to reveal quantitative trait loci (QTLs) underlying transmission stage production. Cercarial production is polygenic: we found three major QTLs on chromosome 1, 3 and 5 (Log-of-the-odds (LOD) = 5.61, 8.19, 6.25) and two minor QTLs on chromosome 2 and 4. These QTLs act additively and explained 28.56% of the phenotypic variation in cercarial shedding. Alleles inherited from the high and low shedding parents were co-dominant at all QTLs, except for chr. 1 and chr. 4 where the &ldquo;high cercarial shedding&rdquo; allele is recessive. These results demonstrate that the genetic architecture of key traits directly relevant to schistosome ecology can be dissected using classical linkage mapping approaches, and set the stage for fine mapping and functional validation of the genes involved using the growing armory of functional and cell biology tools available for this parasite.</p> <p>&nbsp;</p> <p>This dataset is made of 4 tables:</p> <ul> <li>F0_parental_populations.csv</li> <li>F1.csv</li> <li>F2.csv</li> <li>sex.tsv</li> </ul> <p>&nbsp;</p> <p><strong>F0_parental_populations.csv</strong></p> <p>&nbsp;</p> <p>This table contains the number of cercariae produced by each individual <em>Biomphalaria glabrata</em> Bg26 snails infected with single genotypes of <em>Schistosoma mansoni</em> parasite. We have compared the transmission stage production between two different populations of <em>S. mansoni</em> parasite. This dataset was originally published in Le Clec&#39;h et al., 2019 (Striking differences in virulence, transmission and sporocyst growth dynamics between two schistosome populations. Parasites and Vectors. 2019 Oct 16;12(1):485. doi: 10.1186/s13071-019-3741-z).</p> <p>&nbsp;</p> <p>This table is made of 9 columns:</p> <ul> <li><strong>id</strong>: the unique identifier of each sample.</li> <li><strong>schistosoma_population</strong>: the population of schistosome used for the infection of the snail. Each snail was infected with a single parasite genotype. We have used SmLE (high shedder/highly virulent population) and SmBRE (low shedding/low virulent population).</li> <li><strong>Shed.1</strong>: the number of cercariae produced by each parasite genotype at the first shedding week (4 weeks after exposure to parasite).</li> <li><strong>Shed.2</strong>: the number of cercariae produced by each parasite genotype at the second shedding week (5 weeks after exposure to parasite).</li> <li><strong>Shed.3</strong>: the number of cercariae produced by each parasite genotype at the third shedding week (6 weeks after exposure to parasite).</li> <li><strong>Shed.4</strong>: the number of cercariae produced by each parasite genotype at the fourth shedding week (7 weeks after exposure to parasite).</li> <li><strong>sum</strong>: the sum of the cercariae produced by each parasite genotype over the 4 weeks of shedding (Shed.1 + Shed.2 + Shed.3 + Shed.4).</li> <li><strong>average</strong>: the average number of cercariae produced by each parasite genotype over the 4 weeks of shedding.</li> <li><strong>sex</strong>: the sex of each parasite genotype determined by PCR <sup>1</sup>.</li> </ul> <p>&nbsp;</p> <p><strong>F1.csv</strong></p> <p>&nbsp;</p> <p>This table contains the number of cercariae produced by each individual <em>Biomphalaria glabrata</em> Bg26 snails infected with single genotypes of F1 progeny from the cross SmLE x SmBRE (see the manuscript for details).</p> <p>&nbsp;</p> <p>This table is made of 11 columns:</p> <ul> <li><strong>id</strong>: the unique identifier of each sample.</li> <li><strong>cross</strong>: F1A or F1B cross. Each snail was infected with a single parasite genotype from either F1A or F1B progeny.</li> <li><strong>Shed.1</strong>: the number of cercariae produced by each parasite genotype at the first shedding week (4 weeks after exposure to parasite).</li> <li><strong>Shed.2</strong>: the number of cercariae produced by each parasite genotype at the second shedding week (5 weeks after exposure to parasite).</li> <li><strong>Shed.3</strong>: the number of cercariae produced by each parasite genotype at the third shedding week (6 weeks after exposure to parasite).</li> <li><strong>Shed.4</strong>: the number of cercariae produced by each parasite genotype at the fourth shedding week (7 weeks after exposure to parasite).</li> <li><strong>sum</strong>: the sum of the cercariae produced by each parasite genotype over the 4 weeks of shedding (Shed.1 + Shed.2 + Shed.3 + Shed.4).</li> <li><strong>average</strong>: the average number of cercariae produced by each parasite genotype over the 4 weeks of shedding.</li> <li><strong>PO</strong>: the total phenoloxidase activity in infected snail hemolymph, measured at 7.5 weeks post-exposure <sup>2</sup>.</li> <li><strong>Hb</strong>: the hemoglobin rate in infected snail hemolymph, measured at 7.5 weeks post-exposure <sup>3</sup>.</li> <li><strong>sex</strong>: the sex of each parasite genotype determined by PCR <sup>1</sup>.</li> </ul> <p>&nbsp;</p> <p><strong>F2.csv</strong></p> <p>This table contains the number of cercariae produced by each individual <em>Biomphalaria glabrata</em> Bg26 snails infected with single genotypes of F2 progeny from the cross SmLE x SmBRE (see the manuscript for details).</p> <p>&nbsp;</p> <p>This table is made of 10 columns:</p> <ul> <li><strong>id</strong>: the unique identifier of each sample.</li> <li><strong>cross</strong>: F2A or F2B cross. Each snail was infected with a single parasite genotype from either F2A or F2B progeny.</li> <li><strong>Shed.1</strong>: the number of cercariae produced by each parasite genotype at the first shedding week (4 weeks after exposure to parasite).</li> <li><strong>Shed.2</strong>: the number of cercariae produced by each parasite genotype at the second shedding week (5 weeks after exposure to parasite).</li> <li><strong>Shed.3</strong>: the number of cercariae produced by each parasite genotype at the third shedding week (6 weeks after exposure to parasite).</li> <li><strong>Shed.4</strong>: the number of cercariae produced by each parasite genotype at the fourth shedding week (7 weeks after exposure to parasite).</li> <li><strong>sum</strong>: the sum of the cercariae produced by each parasite genotype over the 4 weeks of shedding (Shed.1 + Shed.2 + Shed.3 + Shed.4)</li> <li><strong>average</strong>: the average number of cercariae produced by each parasite genotype over the 4 weeks of shedding.</li> <li><strong>PO</strong>: the total phenoloxidase activity in infected snail hemolymph, measured at 7.5 weeks post-exposure <sup>2</sup>.</li> <li><strong>Hb</strong>: the hemoglobin rate in infected snail hemolymph, measured at 7.5 weeks post-exposure <sup>3</sup>.</li> </ul> <p>&nbsp;</p> <p><strong>sex.csv</strong></p> <p>&nbsp;</p> <p>This table contains the <em>in silico</em> sexing of F0 parents, F1 parents and F2 progeny of <em>S. mansoni</em> parasites.</p> <p>This table is made of 4 columns:</p> <ul> <li><strong>id</strong>: the unique identifier of each sample</li> <li><strong>read_depth</strong>: the read depth ratio between the Z-linked and pseudo-autosomal regions.</li> <li><strong>ratio</strong>: computed ratio between the Z-linked and pseudo-autosomal regions.</li> <li><strong>sex</strong>: the sex of each parasite genotype determined <em>in silico</em>: a ratio around 1 corresponds to a male carrying two Z chromosomes while a ratio around 0.5 corresponds to a female carrying only one Z chromosome.</li> </ul> <p><strong>Notes:</strong></p> <p><sup>1</sup>. Le Clec&rsquo;h W, Chevalier F et al. Real-time PCR for sexing Schistosoma mansoni cercariae. Mol Biochem Parasitol. Jan-Feb 2016; 205(1-2):35-8.doi: 10.1016/j.molbiopara.2016.03.010. Epub 2016 Mar 26.</p> <p><sup>2</sup>. Le Clec&rsquo;h W et al. Characterization of hemolymph phenoloxidase activity in two Biomphalaria snail species and impact of Schistosoma mansoni infection. Parasit Vectors. 2016 Jan 22; 9:32.doi: 10.1186/s13071-016-1319-6.</p> <p><sup>3</sup>. Le Clec&#39;h et al. Striking differences in virulence, transmission and sporocyst growth dynamics between two schistosome populations. Parasit Vectors. 2019 Oct 16; 12(1):485. doi: 10.1186/s13071-019-3741-z.</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Data from: Evolutionary potential and constraints in an aposematic species: Genetic correlations between warning coloration and fitness components in wood tiger moths

<p>Phenotypic data and pedigrees of two laboratory populations of wood tiger moths (<em>Arctia plantaginis</em>) of Finnish (=FIN) and Estonian (=EST) ancestry.</p> <p><strong>Pedigree:&nbsp;</strong><br>ID: individual identifier<br>sire = Father<br>dam=mother</p> <p><strong>Pheno.data:&nbsp;</strong><br>ID: individual identifier<br>Sex: 1=male; 2=female<br>hatchingdate: date when larva hatched<br>pupadate: date of pupation<br>adultdate: date of exclusion<br>Pupa.Weight: weight of pupa [mg]<br>Female.Colour = hindwing colour of females. In this species hindwing colour in females varies continuously from yellow to red. It was quantified by visual matching of hinwdings against a colour scale ranging from &nbsp;1 = yellow to 6 = red.&nbsp;<br>Signal.Size = larva signal size. Larvae show an orange patch of variable size on the back of their black body. The size is given as number of segments<br>Egg.N = egg number produced by the individual<br>Off.N = offspring number. Larvae were counted 2-3 weeks after egg laying</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

GWAS Summary Statistics for Publication: Identifying novel genetic and phenotypic associations to genomic features by leveraging off-target reads in exome sequencing data

<p>This dataset contains summary statistics for genome-wide association studies (GWAS) conducted on genomic features derived from off-target reads in whole-exome sequencing (WES) data. The study utilized tools like Seeing Beyond the Target (SBT) and ImReP to construct novel phenotypic features from unmapped reads in ~50,000 participants in the UK Biobank. Features include mitochondrial DNA (mtDNA) copy number, ribosomal DNA (rDNA) copy number (5S, 18S, 28S), immune repertoire metrics (e.g., T-cell receptor alpha diversity), and microvial genome load (viral and fungal).</p> <p>Summary statistics can be used for replication studies, meta-analyses, or further exploration of these phenotypes.</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

Data from: Enamel proteins reveal biological sex and genetic variability within southern African Paranthropus

<p>This dataset contains the sequences of Paranthropus robustus, first described in 'Enamel proteins reveal biological sex and genetic variability within southern African Paranthropus', as well as the reference data and all the results from the analysis of those sequences.</p> <p><strong>Folders and Sub-Folders:</strong></p> <p><strong>-&nbsp;Paranthropus_Raw_AA_Sequences_Unaligned:&nbsp;</strong>Contains 2 fasta files.&nbsp;Paranthropus_Unaligned.fasta contains all the Paranthropus robustus sequences that were used for all of the analyses.&nbsp;Paranthropus_Unaligned_UNFILTERED.fasta contains all the Paranthropus robusts sequences&nbsp;<strong>before&nbsp;</strong><strong>filtering&nbsp;</strong>for SAP quality/confidence. These sequences were not used in any of the analyses, but are provided here for openness.&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>-</strong>&nbsp;<strong>Reference_Datasets</strong>: Contains 3 fasta files. Each fasta file is a reference dataset used in at least one analysis. The identity and origin of each sample is described in the supplementary document of the publication.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>- Phylogenetic_Analysis_Datasets_and_Trees: </strong>Contains the following <strong>five folders</strong></p> <p>&nbsp; &nbsp; -&nbsp;<strong>Paranthropus_Alignments_All_Datasets</strong>: Contains three folders. Each folder contains the aligned and I/L corrected MSAs (Multiple Sequence Alignments) of Paranthropus robustus and a reference dataset.</p> <p>&nbsp; &nbsp; - <strong>Paranthropus_Diversity_Dataset_Trees_Results</strong>: Contains all analysis done using the 'diversity' reference dataset. Contains one folder for each protein, which includes the protein alignment and the phylogenetic tree of that protein. Additionally a folder named 'CONCATENATED' contains the concatenated alignemnts and trees. The BEAST2-STARBEAST3 folder contains the Starbeast3 analysis, including the xml, output log file, output trees and the input taxon set file.</p> <p>&nbsp; &nbsp; -&nbsp;<strong>Paranthropus_Representative_Dataset_Trees_Results:</strong>&nbsp;Contains all analysis done using the 'representative' reference dataset. Contains one folder for each protein, which includes the protein alignment and the phylogenetic tree of that protein. Additionally a folder named 'CONCATENATED' contains the concatenated alignemnts and trees. The BEAST2 folder contains the time-calibrated BEAST2 analysis, including the xml, output log file, output trees. The folder Distance_Matrix contains the generated distance matrix and the Rscript used to generate the heatmap from it.</p> <p>&nbsp; &nbsp; - <strong>Paranthropus_Independent_Dataset_Trees_Results:&nbsp;</strong>Contains all nexus files and tree-figures&nbsp;used in the analysis of the 'independent' reference dataset.&nbsp;</p> <p>&nbsp; &nbsp;- <strong>Tree_Figures:&nbsp;</strong>Contains three sub-folders and an additional figure. Each sub-folder contains the phylogenetic tree figures generated using one of the three reference datasets.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Data for "Unfolding the structural stability of nanoalloys via symmetry-constrained genetic algorithm and neural network potential"

<p><strong>PtNi_alloy_eam.db</strong> is the dataset (ase.db object) consisting of 55982 intially sampled Pt-Ni alloy structures with EAM energies and forces.</p> <p><strong>PtNi_alloy_dft.db</strong>&nbsp;is the dataset (ase.db object) consisting of the final 6828 resampled&nbsp;Pt-Ni alloy structures&nbsp;with DFT energies and forces calculated by VASP. This is the&nbsp;training set for the NNP, and could be very useful for fitting other machine learning models.</p> <p><strong>PtNi_nanoalloy_vertices_nnp.db</strong> is the dataset (ase.db object) consisting of all the vertices (stable structures) on the convex hulls obtained from NNP-based SCGA runs on 36 Pt-Ni nanoalloy systems. The energies are given by the NNP. Additional information such as mixing energy, motif and&nbsp;symmetry axis are also saved in the dataset and can be queried by the &#39;data&#39;&nbsp;keyword. An&nbsp;xyz format trajectory of these stable structures&nbsp;is also uploaded.</p> <p>All the input files and scripts for hybrid MC-MD&nbsp;simulations, QBC resampling, DFT&nbsp;calculations, NNP training, NNP-based SCGA runs&nbsp;and convex hull analysis are provided in&nbsp;<strong>inputs_and_scripts.zip</strong>.</p>

opencc-by-4.0Aug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record