Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

22,157

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

22,157 results for “Genomics”

Learn how ShareScore rates datasets ↗
edi56/100

Examining genome size and nutrient influence on plant damage patterns

Data was collected to examine whether and how plant genome size (GS) interacts with environmental nutrient additions to influence the amount and patterns of damage plants sustain from invertebrate herbivores and fungal pathogens. Plants were selected based on visual abundance in treatment plots in which nitrogen (N), phosphorus (P), or NP combined had been annually added (Cont. is the abbreviation we used for the control plot with no nutrients added). Additionally, plant traits of percent foliar carbon (% C), percent foliar nitrogen (% N), and specific leaf area (SLA) were measured from all the same plants that damage values were observed from. Data was collected from 847 plants (626 forb individuals, 221 grass individuals) in eight grassland sites that are part of the Nutrient Network (https://nutnet.org), a globally distributed experiment in which plots have different nutrient amendment treatments that are administered identically to allow cross-site comparisons of the effects of nutrients on biodiversity patterning. The sites chosen varied along a north-south latitude, longitude, mean annual precipitation (MAP) and mean annual temperature (MAT) gradient in the United States. All field data was collected between May 2022 and August 2022. The sites included in this study are listed below with their respective Nutrient Network site codes. churn.us= Churning Rapids in Hancock, MI spin.us= Spindletop Farm in Lexington, KY temple.us= Temple in Temple, TX kbs.us= Kellogg Biological Station in Hickory Corners, MI konz.us= Konza Prairie Biological Station in Manhattan, KS cgbg.us= Chichaqua Bottoms Greenbelt in Maxwell, IA cdcr.us= Cedar Creek in East Bethel, MN msum.us= Minnesota State University at Moorhead in Moorhead, MN

openCC (other)Nov 2025View details →
zenodo52/100

Draft genome assembly version 1 of the meadow spittlebug Philaenus spumarius (Linnaeus, 1758) (Hemiptera, Aphrophoridae)

<p>We sequenced the genome of the meadow spittlebug, <em>Philaenus spumarius </em>(Linnaeus, 1758), the main insect vector of <em>Xylella fastidiosa </em>Wells et al. 1987 in Europe (Saponari et al., 2014), using 10x Chromium linked-reads. A single <em>P. spumarius</em> adult female from Portugal (Fontanelas, Sintra; GPS location: 38&deg;50&#39;15.75&quot;N; 9&deg;25&#39;20.77&quot;W), collected in September of 2018, was selected for genome sequencing. This population was initially surveyed for colour polymorphism in 1988 (Quartau &amp; Borges, 1997) and was later included in phylogeographic and population genomic studies of this species (Rodrigues et al., 2014; Seabra et al., unpublished). It is also geographically close to the population from which the individual used for the first partial genome assembly was collected (Rodrigues et al., 2016). The availability of this previous genetic information contributed to the choice of this population as the source of genomic material for whole genome sequencing. A subset of males from the same collection date were analysed for genitalia morphology to confirm species identification, as the best diagnostic characters are the appendages of the aedeagus (Drosopoulos &amp; Quartau, 2002).</p> <p>The genomic DNA of the <em>P. spumarius</em> adult from Sintra was extracted using Illustra Nucleon Phytopure kit according to the manufacturer&rsquo;s instructions (GE Healthcare). We assessed the quality and concentration of the DNA using Femto fragment analyser (Agilent). 10x Chromium library preparation and Illumina genome sequencing (HiSeq X, 150bp paired-end) were performed by Novogene Bioinformatics Technology Co, Beijing, China, in accordance with standard protocols.</p> <p>To create the <em>de novo</em> 10x Chromium assembly we ran Supernova 2.1.1 (Weisenfeld et al., 2017) on the 10x Chromium linked-read data with default parameters, using 1.0 billion reads corresponding to 56X coverage. To improve the initial supernova assembly, we performed iterative scaffolding using all of the 10x raw data (2.3 billion of reads). We ran two rounds of Scaff10x (https://github.com/wtsi-hpag/Scaff10X), followed by mis-assembly detection and correction with Tigmint (Jackman et al., 2018). This was followed by a final round of scaffolding with ARCS (Yeo et al., 2018). The assembly was checked for contamination using the BlobTools pipeline (version 0.9.19; Laetsch and Blaxter 2017;&nbsp;Kumar et al., 2013) and k-mer content was analysed with the KAT comp tool (Mapleson et al., 2017). In order to perform these analyses, it was necessary to remove the 10x linked barcodes from the reads with the script process_10xReads.py (https://github.com/ucdavis-bioinformatics/proc10xG).&nbsp;We assessed the quality of our draft genome assembly by searching for conserved, single copy, arthropod genes (n=1,066) with Benchmarking Universal Single-Copy Orthologs (BUSCO) v3.0 (Waterhouse et al., 2018).</p> <p>With the above assembly procedure, we obtained a final assembly of 2.7 Gb, having a scaffold N50 length of 116 Kb (contig N50 = 18 Kb) and the longest scaffold was 3.7 Mb. The length of the assembly was consistent with the genome size estimated by flow cytometry (Rodrigues et al., 2016). The k-mer distribution indicated high heterozygosity, estimated at 2.3%. BlobTools analyses revealed the presence of contigs assigned to <em>Sodalis </em>spp. (Enterobacteriaceae), a symbiont in members of tribe Philaenini (Koga et al., 2013). These contigs were filtered from the final assembly. Gene completeness assessment shows that 956 (89.6%) among 1,066 BUSCOs were &nbsp;found as complete copies, with only 26 (2.4%) missing. Of the BUSCOs that were detected, 878 (82.4%) were complete and single-copy, 78 (7.3%) were complete and duplicated and 84 (7.9%) were fragmented.</p> <p>In conclusion, due in part to high (2.3%) heterozygosity levels, the <em>P. spumarius</em> version 1 genome assembly is highly fragmented. Nonetheless, the assembly is considered complete and is likely to contain the majority of the gene content of <em>P. spumarius.</em></p>

opencc-by-4.0Jan 2020View details →
zenodo52/100

Genome evolution and introgression in the New Zealand mud snails Potamopyrgus estuarinus and Potamopyrgus kaitunuparaoa

<p>We have sequenced, assembled, and analyzed the nuclear and mitochondrial genomes and transcriptomes of <i>Potamopyrgus estuarinus</i> and <i>Potamopyrgus kaitunuparaoa</i>, two prosobranch snail species native to New Zealand that together span the continuum from estuary to freshwater.<i> </i>These two species are the closest known relatives of the freshwater species <i>P. antipodarum—</i>a model for studying the evolution of sex, host-parasite coevolution, and biological invasiveness—and thus provide key evolutionary context for understanding its unusual biology. The <i>P. estuarinus</i> and <i>P. kaitunuparaoa </i>genomes are very similar in size and overall gene content. Comparative analyses of genome content indicate that these two species harbor a near-identical set of genes involved in meiosis and sperm functions, including seven genes with meiosis-specific functions. These results are consistent with obligate sexual reproduction in these two species and provide a framework for future analyses of <i>P. antipodarum—</i>a species comprising both obligately sexual and obligately asexual lineages, each separately derived from a sexual ancestor. Genome-wide multigene phylogenetic analyses indicate that <i>P. kaitunuparaoa</i> is likely the closest relative to <i>P. antipodarum. </i>We nevertheless show that there has been considerable introgression between <i>P. estuarinus</i> and <i>P. kaitunuparaoa.</i> That introgression does not extend to the mitochondrial genome, which appears to serve as a barrier to hybridization between <i>P. estuarinus </i>and <i>P. kaitunuparaoa.</i> Nuclear-encoded genes whose products function in joint mitochondrial-nuclear enzyme complexes exhibit similar patterns of non-introgression, indicating that incompatibilities between the mitochondrial and the nuclear genome may have prevented more extensive gene flow between these two species.<i>&nbsp;</i>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo52/100

Supplementary dataset to publication: Complete Genome Sequence of Ovine Mycobacterium avium subsp. paratuberculosis Strain JIII-386 (MAP-S/type III) and Its Comparison to MAP-S/type I, MAP-C, and M. avium Complex Genomes.

<p>This is the modified supplemented material to the publication &ldquo;Complete genome sequence of ovine Mycobacterium avium subsp. paratuberculosis strain JIII-386 (MAP-S/type III) and its comparison to MAP-S/type I, MAP-C, and M. avium complex genomes&rdquo;.</p> <p>The complete circular genome of Mycobacterium avium subsp. paratuberculosis (MAP) strain JIII-386 from Germany, closed by Nanopore technology in this study, was presented and compared with the draft genome of JIII-386, previously published in [doi:10.1093/gbe/ew154], the closed genome of the MAP-S/type I strain Telford, the MAP-S/type III draft genome of strain S397, twelve closed MAP-C (type II) strains and eight closed Mycobacterium avium (M. a.) strains of subsp. hominissuis (MAH) and subsp. avium (MAA). Structural comparisons clearly revealed the mosaic nature of MAP genomes, the differences between MAP subtypes I, II and III, and the higher diversity of MAP-S compared to MAP-C genomes.&nbsp;</p> <p>The material provides a wealth of detailed results from these analyses and comparisons. These include a list of identified ncRNA and Riboswitches, as well as additional genes in finished JIII-386, the gene content of identified prophage regions, copy number of identified transposable elements and a list of selected virulence-associated genes in the different MAP-type (I - III) strains. The genomic islands identified and included genes along with their predicted functions were presented for six MAP genomes (belonging to MAP-S/type I and III, and MAP-C), one MAH genome and one MAA genome. One table shows the corresponding genomic islands in the genomes of JIII-386, Telford and three MAP-C genomes. Furthermore, homologous genes of known MAP-S specific Large Sequence Polymorphisms regions (LSP<sup>S</sup> = LSP-S) were recorded in different MAP-S type strains, one MAH and one MAA strain, as well as genes of deletions #1 (LSP<sup>A</sup>-20), #2, and s-delta-1, previously described as MAP-S-specific deletions, their presence or absence in 3 MAP-S, 12 MAP-C, 4 MAH, and 4 MAA strains were listed. Different presence or absence of genes, but also identified frameshifts or disruptions of various virulence-associated genes could lead to the different MAP-type specific phenotypic characteristics. Comprehensive core and pan genome analyses (results listed in six tables) revealed unique genes and genes likely to have been acquired by horizontal gene transfer in different MAP types and subtypes, but also emphasized the highly conserved and close relationship, and the complex evolution of M. a. strains.</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

A genome-wide study of ruminants reveals two endogenous retrovirus families still active in goats

<p>Additional file - A genome-wide study of ruminants reveals two endogenous retrovirus families still active in goats&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo52/100

Genome-wide association summary statistics for human blood plasma glycome

<p>The dataset&nbsp;contains results of genome-wide association study of human blood plasma&nbsp;glycome. The 113 files contain association summary statistics for 113 glycome traits, of which 36 were directly measured by UPLC technology and 77 were derived glycome traits. Description of each glycome trait can be found in the <strong>Additional notes</strong> section. This&nbsp;dataset is also available for graphical exploration in the genomic context at <a href="http://gwasarchive.org">http://gwasarchive.org</a>.&nbsp;</p> <p>The data are provided on an &quot;AS-IS&quot; basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Sharapov, S. Z., Tsepilov, Y. A., Klaric, L., Mangino, M., Thareja, G., Shadrina, A. S., &hellip; Aulchenko, Y. (2019). Defining the genetic control of human blood plasma N-glycome using genome-wide association study. <em>Human Molecular Genetics</em>. http://doi.org/10.1093/hmg/ddz054</li> <li>Sodbo Sharapov, Yakov Tsepilov, Lucija Klaric, Massimo Mangino, Gaurav Thareja, Mirna Simurina, Concetta Dagostino, Julia Dmitrieva, Marija Vilaj, FranoVuckovic, Tamara Pavic, Jerko Stambuk, Irena Trbojevic-Akmacic, Jasminka Kristic, Jelena Simunovic, Ana Momcilovic, Harry Campbell, Malcolm Dunlop, Susan Farrington, Maria Pucic-Bakovic, Christian Gieger, Massimo Allegri, Edouard Louis, Michel Georges, Karsten Suhre, Tim Spector, Frances MK Williams, Gordan Lauc, Yurii Aulchenko. (2018). Genome-wide association summary statistics for human blood plasma glycome (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1298406</li> </ol> <p><strong>Funding</strong></p> <p>This work was supported by the European Community&rsquo;s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736) and by the European Structural and Investments funding for the &quot;Croatian National Centre of Research Excellence in Personalized Healthcare&quot; (contract #KK.01.1.1.01.0010).</p> <p>The work of SSh was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Programme.</p> <p>The work of YT was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017).</p> <p>Karsten Suhre and Gaurav Thareja are supported by &lsquo;Biomedical Research Program&rsquo; funds at Weill Cornell Medicine - Qatar, a program funded by the Qatar Foundation. We thank all staff at Weill Cornell Medicine - Qatar and Hamad Medical Corporation, and especially all study participants who made the QMDiab study possible.</p> <p>The SOCCS study was supported by grants from Cancer Research UK (C348/A3758, C348/A8896, C348/ A18927); Scottish Government Chief Scientist Office (K/OPR/2/2/D333, CZB/4/94); Medical Research Council (G0000657-53203, MR/K018647/1); Centre Grant from CORE as part of the Digestive Cancer Campaign (<a href="http://www.corecharity.org.uk">http://www.corecharity.org.uk</a>).</p> <p>TwinsUK is funded by the Wellcome Trust, Medical Research Council, European Union, the National Institute for Health Research (NIHR)-funded BioResource, Clinical Research Facility and Biomedical Research Centre based at Guy&rsquo;s and St Thomas&rsquo; NHS Foundation Trust in partnership with King&rsquo;s College London.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNP: SNP rsID</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build)&nbsp;</li> <li>OTHER_ALLELE: reference allele (coded as &quot;0&quot;)</li> <li>EFFECT_ALLELE: effective allele (coded as &quot;1&quot;)</li> <li>EAF: effective allele frequency&nbsp;</li> <li>N: sample size</li> <li>BETA: effect size of effective allele</li> <li>SE: standard error of effect size</li> <li>PVAL: P-value of association (without GC correction)</li> <li>IMPUTATION: imputation quality</li> </ol>

opencc-by-4.0Jun 2018View details →
zenodo52/100

Genome-wide association summary statistics for human healthspan

<p>The dataset contains genome-wide association summary statistics computed for heathspan. The UKB sub-population of 300,447 genetically Caucasian, British individuals were analyzed. For more details see [1].</p> <p>The data are provided on an &quot;AS-IS&quot; basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Zenin, A., Tsepilov, Y., Sharapov, S., Getmantsev, E., Menshikov, L. I., Fedichev, P. O., &amp; Aulchenko, Y. (2019). Identification of 12 genetic loci associated with human healthspan. <em>Communications Biology</em>, <em>2</em>(1), 41. http://doi.org/10.1038/s42003-019-0290-0</li> <li>Aleksandr Zenin, Yakov Tsepilov, Sodbo Sharapov, Evgeny Getmantsev, Leonid Menshikov, Peter Fedichev, &amp; Yurii Aulchenko. (2018). Genome-wide association summary statistics for human healthspan (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1302861</li> </ol> <p><strong>Funding</strong></p> <p>The work was supported by Russian Ministry of Science and Education under 5-100 Excellence Programme.&nbsp;<br> The work was supported by the Federal Agency of Scientific Organizations via the Institute of Cytology and Genetics (project #0324-2018-0017).&nbsp;<br> This research has been conducted using the UK Biobank Resource.&nbsp;<br> The study has been funded by Gero LLC.</p> <p><strong>Column headers:</strong></p> <ol> <li>SNPID - SNP rsID</li> <li>chr - chromosome</li> <li>pos - position (GRCh37 build / hg19)</li> <li>EA - effective allele (coded as &quot;1&quot;)</li> <li>RA - reference allele (coded as &quot;0&quot;)</li> <li>EAF - effective allele frequency</li> <li>beta - effect size of effective allele</li> <li>se - standard error of effect size</li> <li>Z - Z-value of association</li> <li>-log10(p-value) - minus log10(P-value) of association</li> </ol>

opencc-by-4.0Jul 2018View details →
zenodo52/100

Genome-wide association summary statistics for back pain

<p>The dataset contains results of a genome-wide association study of back pain. Two files contain association summary statistics for discovery GWAS based on the analysis of 350,000 white British individuals from the UK Biobank and meta-analysis GWAS based on the meta-analysis of the same 350,000 individuals and additional 103,862 individuals of European Ancestry from the UK biobank (total N = 453,862). The phenotype of back pain was defined by the answer provided by the UK biobank participants to the following question: &quot;Pain type(s) experienced in last month&quot;. Those who reported &ldquo;Back pain&rdquo;, were considered as cases, all the rest were considered as controls. Individuals who did not reply or replied: &quot;Prefer not to answer&quot; or &quot;Pain all over the body&quot; were excluded. This&nbsp;dataset is also available for graphical exploration in the genomic context at&nbsp;<a href="http://gwasarchive.org/">http://gwasarchive.org</a>.&nbsp;</p> <p>The data are provided on an &quot;AS-IS&quot; basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant. This research has been conducted using the UK Biobank Resource and the use of the data is guided by the principles formulated by the UK Biobank.</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li>Insight into the genetic architecture of&nbsp;back pain&nbsp;and its risk factors from a study of 509,000 individuals.&nbsp;Freidin, Maxim; Tsepilov, Yakov; Palmer, Melody; Karssen, Lennart; Suri, Pradeep; Aulchenko, Yurii; Williams, Frances MK,# CHARGE Musculoskeletal Working Group.&nbsp;PAIN: February 06, 2019 - Volume Articles in Press - Issue - p<br> doi: 10.1097/j.pain.0000000000001514</li> <li>Maxim B Freidin, Yakov A Tsepilov, Melody Palmer, Lennart Karssen, CHARGE Musculoskeletal Working Group, Pradeep Suri, &hellip; Frances MK Williams. (2018). Genome-wide association summary statistics for back pain (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1319332</li> </ol> <p><strong>Funding:</strong></p> <p>This study was supported by the European Community&rsquo;s Seventh Framework Programme funded project PainOmics (Grant agreement # 602736).&nbsp;<br> The research has been conducted using the UK Biobank Resource (project # 18219).</p> <p>The development of software implementing SMR/HEIDI test and database for GWAS results was&nbsp;supported by the Russian Ministry of Science and Education under the&nbsp;5-100 Excellence Program&rdquo;.</p> <p>Dr. Suri&rsquo;s time for this work was supported by VA Career Development Award # 1IK2RX001515 from the United States (U.S.) Department of Veterans Affairs Rehabilitation Research and Development Service. The contents of this work do not represent the views of the U.S. Department of Veterans Affairs or the United States Government.</p> <p>Dr. Tsepilov&rsquo;s time for this work was supported in part by the Russian Ministry of Science and Education under the 5-100 Excellence Program.</p> <p><strong>Column headers - discovery (350K)</strong></p> <ol> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build)&nbsp;</li> <li>ID: SNP rsID</li> <li>REF: reference allele (coded as &quot;0&quot;)</li> <li>ALT: effect allele (coded as &quot;1&quot;)</li> <li>CASE_ALLELE_CT: allele observation count in cases</li> <li>CTRL_ALLELE_CT: allele observation count in controls</li> <li>ALT_FREQ: effect allele frequency&nbsp;</li> <li>MACH_R2: imputation quality</li> <li>TEST: model of association test (additive)</li> <li>OBS_CT: sample size</li> <li>BETA: effect size of effect allele</li> <li>SE: standard error of effect size</li> <li>T_STAT: Z-value of effect allele</li> <li>P: P-value of association (without GC correction)</li> <li>MAF: minor allele frequency</li> </ol> <p><strong>Column headers - meta-analysis&nbsp;(450K)</strong></p> <ol> <li>MarkerName: SNP rsID</li> <li>Allele1: effect allele (coded as &quot;1&quot;)</li> <li>Allele2: reference allele (coded as &quot;0&quot;)</li> <li>Freq1: effect allele frequency</li> <li>FreqSE: standard error of effect allele frequency</li> <li>Effect: effect size of effect allele</li> <li>StdErr: standard error of effect size</li> <li>P-value: P-value of association (without GC correction)</li> <li>Direction: sign of effect in discovery and replication samples</li> <li>n_total: Total sample size</li> <li>CHR: chromosome</li> <li>POS: position (GRCh37 build)&nbsp;</li> <li>MACH_R2_discovery: imputation quality in discovery sample</li> </ol>

opencc-by-4.0Jul 2018View details →
edi52/100

Data in support of 'Mechanistic insights into plant community responses to environmental variables: genome size, cellular nutrient investments, and metabolic trade-offs.'

Data was collected to examine whether and how the plant genome size (GS) influences traits (stomata size, stomata density, cellular and tissue level carbon (C), nitrogen (N), and phosphorus (P) contents) and metabolic-tradeoffs (of photosynthesis, evapotranspiration, water-use, efficiency) of plants in treatment plots in which nothing, N, P, or NP had been annually added. Data was collected from ~500 plants from seven grassland sites that are all part of the Nutrient Network (https://nutnet.org), a globally distributed experiment in which plots have different nutrient amendment treatments that are administered identically to allow cross-site comparisons of the effects of nutrients on biodiversity patterning. The sites chosen varied along a North-South latitude, longitude, mean annual precipitation (MAP) and mean annual temperature (MAT) gradient.

openCC (other)Sep 2024View details →
edi52/100

Arctic grayling neutral genomic microsatellite loci from the Kuparuk, the Sagavanirktok (primarily Oksrukuyik Creek) and the Itkillik (primarily the I-Minus outlet stream) watersheds, 2010-2014

Since 2009, The FISHSCAPE Project (National Science Foundation grants: 1719267, 1417754, and 0902153), based at Toolik Field Station, has monitored physical, chemical, and biological parameters within three watersheds: The Kuparuk (including Toolik Lake and Toolik outlet stream), The Sagavanirktok (primarily Oksrukuyik Creek, but also including sections of the Atigun River and Tea and Galbraith Lakes), and Itkillik (primarily the I-Minus outlet stream a tributary that that feeds into the Itkilik River). Goals of the FISHSCAPE project are to understand and predict the adaptability and persistence of a key Arctic species, the Arctic grayling (Thymallus arcticus), to changing climate and hydrology. Research questions include: (1) Does landscape structure determine movement within and among watersheds; (2) do populations adapt to stream characteristics at local and regional scales; and (3) will the relative adaptability of populations determine their persistence under future climate change. We used genetics to investigate population structure and landscape genetics for Arctic grayling. Adult and young-of-the-year fish were captured at sampling locations and coordinates and/or specific station locations were noted. Fin clip samples (adults) or whole fish (young-of-the-year) were collected and preserved in 95% ethanol until Deoxyribonucleic acid (DNA) was extracted. Polymerase chain reaction (PCR) products from neutral genomic microsatellite loci were scored and used to assess population genetic structure and other population parameters. Adult capture and movement data, including length, weight and Passive Integrated Transponder (PIT) tag information, can be found in a separate data package.

openCC (other)Jan 2020View details →
zenodo48/100

Effects of crown gall disease on natural microbiota of Vitis vinifera - genome annotations

<p>Young grapevines (Vitis vinifera) frequently die due to the crown gall (CG) disease induced by the plant pathogen Allorhizobium vitis (Rhizobiaceae). Virulent members of A. vitis harbour a tumor-inducing (Ti) plasmid and cause formation of CGs due to genes encoded on the T-DNA. Expression of the oncogenes by transformed host cells induce cell proliferation, metabolic and physiological changes. The CG produces opines uncommon to plants, which provide an important nutrient source for A. vitis harbouring opine catabolism enzymes. CGs host a defined bacterial community and the mechanisms establishing a CG-specific bacterial community are currently unknown. Thus, we were interested in whether genes homologous to those of the Ti-plasmid coexist in the genomes of the microbial species coexisting in CGs. We isolated eight bacterial strains from grapevine CGs, sequenced their genomes and tested their virulence and opine utilization ability in bioassays. In addition, the eight genome sequences were aligned to the sequences of a Ti-plasmid and seven published bacterial genomes, including closely related plant associated bacteria but not from CGs. Homologous genes for virulence and opine anabolism were only present in the virulent Rhizobiaceae. By contrast, homologs of the opine catabolism genes were present in all strains including the non-virulent members of the Rhizobiaceae and non-Rhizobiaceae, indicating horizontal gene transfer of the opine degradation cluster from virulent to non-virulent strains. These results along with those of the opine utilization assay support the important role of opine utilization for co-colonization of virulent and non-virulent bacteria in CGs, thereby shaping the CG community.</p> <p>This dataset contains the prokka annotations of the genomes as used in &quot;Opportunistic bacteria of grapevine crown galls are equipped with the genomic repertoire for opine utilization&quot;</p>

opencc-by-4.0Dec 2021View details →
zenodo48/100

Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study (Supplementary Data)

<p>The zip file contains supplementary data for the publication - Genome-Wide DNA Methylation in Peripheral Blood and Long-Term Exposure to Source-Specific Transportation Noise and Air Pollution: The SAPALDIA Study, accepted for publication in Environmental Health Perspectives (DOI: 10.1289/EHP6174).</p> <p>The description of the files are noted below:</p> <p><strong>1. Readme File for SAPALDIA Noise and Air Pollution EWAS Single Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_SingleExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_SingleExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_SingleExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p>&nbsp;</p> <p><strong>General footnote for all files:</strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter &lt;2.5 &micro;m. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 &micro;g/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from single exposure epigenome-wide linear mixed models, with random intercept at the level of participant. Each model was adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator (for Lden models) and leukocyte composition. In a preliminary step, DNA methylation &beta;-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites.</p> <p>Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (&ldquo;modified winsorization&rdquo;). The &ldquo;winsorized&rdquo; data were then used as the dependent variables in the epigenome-wide association study.</p> <p>&nbsp;</p> <p><strong>2. Readme File for SAPALDIA Noise and Air Pollution EWAS Multi Exposure.zip </strong></p> <p>This zip file contains all the results of the association between source-specific transportation noise (aircraft, railway and road traffic), air pollution (NO<sub>2</sub> and PM<sub>2.5</sub>), and genome-wide DNA methylation, derived from multi-exposure models.</p> <p><strong>SAPALDIA_EWAS_MultiExposure_AircraftLden.txt</strong> contains the results for aircraft noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RailwayLden.txt</strong> contains the results for railway noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_RoadtrafficLden.txt</strong> contains the results for road traffic noise</p> <p><strong>SAPALDIA_EWAS_MultiExposure_NO2.txt</strong> contains the results for nitrogen dioxide</p> <p><strong>SAPALDIA_EWAS_MultiExposure_PM25.txt</strong> contains the results for fine particulate matter</p> <p><strong>General table footnotes: </strong>SAPALDIA: Swiss cohort study on air pollution and lung and heart diseases in adults. CpG: Cytosine-phosphate-Guanine. CHR: chromosome. SE: standard error. Lden: day-evening-night noise level. NO<sub>2</sub>: nitrogen dioxide. PM<sub>2.5</sub>: particulate matter with aerodynamic diameter &lt;2.5 &micro;m. Beta coefficients represent increase or decrease in DNA methylation per 10 dB increase in aircraft, railway or road traffic Lden or 10 &micro;g/m<sup>3</sup> increase in NO<sub>2</sub> or PM<sub>2.5</sub>. All estimates were from multi-exposure epigenome-wide linear mixed models, with random intercept at the level of participant, and were adjusted for age, sex, educational level, area, and neighborhood socio-economic status, greenness index, smoking status and pack years, exposure to passive smoke, consumption of fruits, vegetables and alcohol, nested study, asthma status, survey, source-specific noise truncation indicator and leukocyte composition. Multi-exposure models included all five exposures (Aircraft, railway, road traffic Lden and respective truncation indicators, NO<sub>2</sub> and PM<sub>2.5</sub>) at the same time. In a preliminary step, DNA methylation &beta;-values were regressed on the Illumina control probe-derived first 30 principal components to correct for correlation structures and technical bias, and residuals of these regressions covering 430,477 CpGs were used as the technical bias-corrected methylation level at the CpG sites. Extreme values of the residuals (lying beyond three times the interquartile range below the first quartile and above the third quartile at each CpG site) were replaced with their corresponding detection threshold value (&ldquo;modified winsorization&rdquo;). The &ldquo;winsorized&rdquo; data were then used as the dependent variables in the epigenome-wide association study.</p>

opencc-by-4.0May 2020View details →
zenodo48/100

PopDel identifies medium-size deletions jointly in tens of thousands of genomes - Variant call sets

<p>This data set contains the variant calls sets generated by different tools for the benchmarks in the paper <a href="https://www.nature.com/articles/s41467-020-20850-5">PopDel identifies medium-size deletions simultaneously in tens of thousands of genomes</a>. It includes the VCFs/BCFs for the following test cases:</p> <ul> <li>Random deletion simulation on up to 1000 chromosome 21 samples</li> <li>1000 Genomes Project deletions inserted into simulated chromosomes 17 to 22 of up to 500 samples</li> <li>HG001 (NA12878)</li> <li>Trio of <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG002_NA24385_son/NIST_HiSeq_HG002_Homogeneity-10953946/">HG002</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG003_NA24149_father/NIST_HiSeq_HG003_Homogeneity-12389378/">HG003</a> + <a href="https://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/AshkenazimTrio/HG004_NA24143_mother/NIST_HiSeq_HG004_Homogeneity-14572558/">HG004</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Diversity-Cohort">Polaris Diversity cohort</a></li> <li><a href="https://github.com/Illumina/Polaris/wiki/HiSeqX-Kids-Cohort">Polaris Kids cohort</a></li> </ul> <p>Further, the long and short read reference call sets for HG001 are provided. For HG002 the reference call set and the high confidence regions by the Genome in a Bottle consortium are provided.</p> <p>For details on how the files have been created, please refer to the paper and the script repository on <a href="https://github.com/kehrlab/PopDel-scripts">GitHub</a>.</p>

opencc-by-4.0Aug 2020View details →
zenodo48/100

Genomic vcf file for D.melanogaster, Sussex LHM population

<p>Data, logs and code for genomic vcf genotypes file for the Morrow lab, D.melanogaster LHM sequencing and genotyping project.</p>

opencc-by-4.0Dec 2016View details →
zenodo48/100

Population genomics of Sussex LHM Drosophila melanogaster

<p>Input data, code, output data, summary plots, and run logs for investigation of the population genetics of the Drosophila melanogaster Sussex LHM sample.</p> <p>For output graphs, see popgen_plots.png</p> <p>Input data are 'plink binary' format.</p> <p>Code is a unix/linux shell script containing commands for Plink to perform population genetic tests.</p> <p>The two R scripts contain i. a short command for making a subpopulation file, ii. commands for plotting the output data. Both are initated in the shell script.</p> <p>Platform and version information are available in the log files. Other information available in the shell and R scripts.</p> <p>Broad observations are that the allele frequency disibribution is normal except a few humps around MAF 0.2-0.3 in the autosomes.</p> <p>Linkage disequilbrium, on average, levels-out after ~200bp but there can still be some at distances of 300Kb.</p> <p>The population appears to be divided into four genetically distinct groups (on the IBD-PCA scatter plot), with Fst analysis indicating that this is caused by genetic variation around the centromeres. This is possibly caused by historic admixture, and low centromeric recombination.</p>

opencc-by-4.0Jun 2017View details →
zenodo48/100

Variant, Metabolite and Source Data for: Population genomics uncover loci for trait improvement in the indigenous African cereal tef (Eragrostis tef)

<p>These files contain the variant and metabolome for a collection of 220 tef (<em>Eragrsotis tef)</em> accessions from an ethiopian diversity panel. The accessions were assembled and managed by the Ethiopian Institute of Agricultural Research (EIAR, Ethiopia). The variant data was produced at the John Innes Centre (UK). The metabolome data was produced at Aberystwyth University (UK). These dataset are described in Jones et al. (2024), <em>bioRxiv</em>, https://doi.org/10.1101/2024.09.30.615331. The source data for main figures in the publication are also included.</p> <p>The submission contains</p> <ol> <li>EIAR_filtered.vcf.gz: This is the variant data obtained from alignment of Illumina reads from all 220 teff accessions to the reference assembly of tef (Dabbi). &nbsp;Low quality variants were filtered out. This variant data was used for constructing the phylogenetic relationship between the accessions. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>pooled_EIAR_filtered.vcf.gz: After the phylogentic analysis described above, reads from accessions that were found to be genetically redundant were pooled before variant calling. This file was used for the SNP GWAS analysis. The samples names corresponds to the DNA code in Supplementary Table S10 (Jones et al, 2024).</li> <li>&nbsp;Metabolite_Profile.xlxs (source data for Figure 5): This file contains m/z feature intensities from untargeted metabolite fingerprinting using Flow Infusion Electrospray High-resolution Mass Spectrometry (FIE-HRMS). The sample names contains a combination of Location code and Plot number in Supplementary Table S10 e.g AT plot 1, CD plot 1, DZ plot 1, where AT, CD and DZ represent Alem Tena, Chefe Donsa and Debre Zeit, respectively. The data was used for the partial least squares discriminant analysis and differentially accumulated metabolites analysis presented in Figure 5.</li> <li>Source data: Numerical source data for graphs and charts in Figures 3 - 7.</li> <li>Tsedey TT2 Sequence from Improved Assembly: The 4A and 4B sequences around the TT2 orthologue in tef from the improved PacBio-based chromosome-scale assembly of tef. These sequences were used for plotting the LTR Copia alignments presented in Supplementary Figure 9. We thank Corteva for pre-publication access to this improved Tsedey genome assembly.</li> </ol>

opencc-by-4.0Oct 2024View details →
zenodo48/100

Human intestinal Bacteria Collection (HiBC): Isolates and genomes metadata

<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the taxonomy of the isolates, as well as metadata regarding their cultivation and isolation. We also provide metadata regarding the sequencing, genome assembly process and the biological sequences.</p> <p><strong>UPDATE v7</strong>: INSDC accession for <em>Segatella sinensis</em> CLA-AA-H117 was a missing value and is now the correct value of GCA_040324585.2.</p> <p><strong>UPDATE v6:&nbsp;</strong>The growth atmosphere is now indicated by anaerobic or aerobic instead of "Anaerobe/Aerobe" that was a misleading term. The risk group of these two isolates went from 1 to 2:</p> <ul> <li>CLA-AA-H205: <em>Anaerostipes caccae&nbsp;</em></li> <li>CLA-AA-H83: <em>Bacteroides fragilis</em></li> </ul> <p>The risk group of the following isolates has been updated (usually from unknown to 1, or from 2 to 1):</p> <ul> <li>CLA-SR-H026: <em>Aedoeadaptatus acetigenes</em></li> <li>CLA-KB-H139:<em> Bacteroides xylanisolvens</em></li> <li>CLA-SR-H015: <em>Bacteroides xylanisolvens</em></li> <li>CLA-AA-H187: <em>Blautia fusiformis</em></li> <li>CLA-AA-H274: <em>Brotaphodocola catenula</em></li> <li>CLA-AA-H286: <em>Butyricimonas faecihominis</em></li> <li>CLA-AA-H278:<em> Clostridium fessum</em></li> <li>CLA-AA-H147: <em>Dorea ammoniilytica</em></li> <li>CLA-SR-H027: D<em>orea formicigenerans</em></li> <li>CLA-KB-H89: <em>Dorea longicatena</em></li> <li>CLA-KB-H94: <em>Dorea longicatena</em></li> <li>CLA-SR-H022: <em>Enterococcus lactis</em></li> <li>CLA-AA-H250: <em>Hominenteromicrobium mulieris</em></li> <li>CLA-AA-H232: H<em>ominilimicola fabiformis</em></li> <li>CLA-AA-H246: <em>Hominisplanchenecus faecis</em></li> <li>CLA-AA-H276:<em> Hominiventricola filiformis</em></li> <li>CLA-AA-H213:<em> Oliverpabstia intestinalis</em></li> <li>CLA-AA-H241: <em>Oliverpabstia intestinalis</em></li> <li>CLA-AA-H58: <em>Pilosibacter fragilis</em></li> <li>CLA-KB-H110: <em>Ruthenibacterium lactatiformans</em></li> <li>CLA-AA-H174: <em>Segatella sinensis</em></li> <li>CLA-AA-H2: <em>Veillonella parvula</em></li> <li>CLA-AA-H273: <em>Waltera acetigignens</em></li> </ul> <p>Typos in media list have been fixed.&nbsp;</p> <p><strong>UPDATE v5</strong>: The accessions number for the genomes on INSDC databases are added under the column Accession. Plus two typos in the risk group column have been corrected as follow:</p> <ul> <li>CLA-AA-H173: from Risk Group 4 (!) to 2 like the other strain of <em>Sutterella wadsworthensis</em></li> <li>CLA-AA-H198: from Risk Group 4 (!) to 1 like the other <em>Bifidobacterium&nbsp;</em>species.</li> </ul> <p><strong>UPDATE v4</strong>: Only the taxonomy of a couple of isolates has been changed, as follow:</p> <ul> <li>CLA-ER-H4: <em>Collinsella sp900547855</em> instead of <em>Collinsella sp900544645</em></li> <li>CLA-AA-H142: <em>Pilosibacter fragilis</em> (<em>f__Clostridiaceae</em>) instead of <em>Sakamotonia hominis gen. nov.</em> (<em>f__Lachnospiraceae</em>)</li> <li>CLA-AA-H58: <em>Pilosibacter fragilis&nbsp;</em>(<em>f__Clostridiaceae</em>)&nbsp;instead of <em>Sakamotonia hominis gen. nov.&nbsp;</em>(<em>f__Lachnospiraceae</em>)</li> <li>CLA-AA-H89B: <em>Lachnospira intestinalis sp. nov.</em> instead of <em>Lachnospira hominis sp. nov.</em></li> <li>CLA-JM-H10: <em>Lachnospira hominis sp. nov.</em> instead of <em>Lachnospira intestinalis sp. nov.</em></li> <li>CLA-JM-H7B: <em>Faecalibacterium taiwanense</em> instead of <em>Faecalibacterium faecis sp. nov.</em></li> <li>CLA-JM-H45: <em>Merdimmobilis hominis</em> instead of <em>Hominicola intestinalis gen. nov.</em></li> </ul> <p><strong>UPDATE v3</strong>: The genome of one of our isolate had been unfortunately swapped. This mistake has been now corrected on Zenodo and Coscine. The genome of <em>Segatella sinensis</em> CLA-AA-H117 should be considered correct with 103 contigs and 3 671 232 nt. Please note that the genome available at the NCBI is the correct one (GCA_040324585.2). Two typos regarding taxonomy have been corrected as well: <em>Maccoya intestinihominis</em> has been corrected to <em>Maccoyia intestinihominis</em> and <em>Faecousia faecis</em> to <em>Faecousia intestinalis</em>.</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Gene Annotations of 49 Bacillariophyta Genome Assemblies

<div>Contact: katharina.hoff@uni-greifswald.de.</div> <div>&nbsp;</div> <div> <h2>Manuscript</h2> <p>The data hosted here is associated with the preprint <a href="https://doi.org/10.48550/arXiv.2410.05467">https://doi.org/10.48550/arXiv.2410.05467</a></p> </div> <h2>Files</h2> <div>The following gff3-files with structural and functional genome annotation are included in the compressed archive Bacillariophyta_annotations.tar.gz:</div> <div>&nbsp;</div> <div>Asterionella_formosa.gff3<br>Asterionellopsis_glacialis.gff3<br>Bacterosira_constricta.gff3<br>Chaetoceros_muellerii.gff3<br>concatenated_output.gff3<br>Conticribra_guillardii.gff3<br>Conticribra_weissflogii.gff3<br>Craspedostauros_australis.gff3<br>Cyclostephanos_invisitatus.gff3<br>Cyclostephanos_tholiformis.gff3<br>Cyclotella_atomus.gff3<br>Cyclotella_baltica.gff3<br>Cyclotella_choctawhatcheeana.gff3<br>Cyclotella_cryptica.gff3<br>Cylindrotheca_fusiformis.gff3<br>Detonula_confervacea.gff3<br>Discostella_pseudostelligera.gff3<br>Discostella_stelligera.gff3<br>Discostella_stelligeroides.gff3<br>Epithemia_pelagica.gff3<br>Fistulifera_pelliculosa.gff3<br>Fistulifera_solaris.gff3<br>Fragilaria_radians.gff3<br>Fragilariopsis_cylindrus.gff3<br>Licmophora_abbreviata.gff3<br>Mediolabrus_comicus.gff3<br>Nitzschia_palea.gff3<br>Nitzschia_putrida.gff3<br>Porosira_glacialis.gff3<br>Psammoneis_japonica.gff3<br>Pseudo-nitzschia_multiseries.gff3<br>Pseudo-nitzschia_pungens.gff3<br>Skeletonema_costatum.gff3<br>Skeletonema_marinoi.gff3<br>Skeletonema_menzelii.gff3<br>Skeletonema_potamos.gff3<br>Skeletonema_tropicum.gff3<br>Stephanocyclus_meneghinianus.gff3<br>Stephanodiscus_minutulus.gff3<br>Stephanodiscus_triporus.gff3<br>Thalassiosira_allenii.gff3<br>Thalassiosira_delicatula.gff3<br>Thalassiosira_exigua.gff3<br>Thalassiosira_gravida.gff3<br>Thalassiosira_livingstoniorum.gff3<br>Thalassiosira_mediterranea.gff3<br>Thalassiosira_oceanica.gff3<br>Thalassiosira_ordinaria.gff3<br>Thalassiosira_pacifica.gff3<br>Thalassiosira_profunda.gff3</div> <div>&nbsp;</div> <div>To extract the dataset, execute the following command:</div> <div>&nbsp;</div> <div><code>tar -xvf Bacillariophyta_annotations.tar.gz</code></div> <h2>Genome Assemblies</h2> <div>The files in this folder attain to genome assemblies are publicly available at NCBI datasets (https://www.ncbi.nlm.nih.gov/datasets/). We used the following versions:</div> <div>&nbsp;</div> <div>Asterionella formosa GCA_002256025.1</div> <div>Asterionellopsis glacialis GCA_014885115.2</div> <div>Bacterosira constricta GCA_037356235.1</div> <div>Chaetoceros muellerii GCA_019693545.1</div> <div>Conticribra guillardii GCA_036939335.1</div> <div>Conticribra weissflogii GCA_036940025.1</div> <div>Craspedostauros australis GCA_026770025.1</div> <div>Cyclostephanos invisitatus GCA_036939675.1</div> <div>Cyclostephanos tholiformis GCA_036939975.1</div> <div>Cyclotella atomus GCA_036939935.1</div> <div>Cyclotella baltica GCA_036939635.1</div> <div>Cyclotella choctawhatcheeana GCA_036939855.1</div> <div>Cyclotella cryptica GCA_013187285.1</div> <div>Cylindrotheca fusiformis GCA_019693525.1</div> <div>Detonula confervacea GCA_036939415.1</div> <div>Discostella pseudostelligera GCA_036940085.1</div> <div>Discostella stelligera GCA_036939735.1</div> <div>Discostella stelligeroides GCA_036939555.1</div> <div>Epithemia pelagica GCA_946965045.2</div> <div>Fistulifera pelliculosa GCA_026008555.1</div> <div>Fistulifera solaris GCA_030295235.1</div> <div>Fragilaria radians GCA_900642245.1</div> <div>Fragilariopsis cylindrus GCA_900095095.1</div> <div>Licmophora abbreviata GCA_900291995.1</div> <div>Mediolabrus comicus GCA_036940125.1</div> <div>Nitzschia palea GCA_019593585.1</div> <div>Nitzschia putrida GCA_016586335.1</div> <div>Porosira glacialis GCA_036939395.1</div> <div>Psammoneis japonica GCA_008632985.1</div> <div>Pseudo-nitzschia multiseries GCA_037355745.1</div> <div>Pseudo-nitzschia pungens GCA_037355855.1</div> <div>Skeletonema costatum GCA_018806925.1</div> <div>Skeletonema marinoi GCA_030544225.1</div> <div>Skeletonema menzelii GCA_036940005.1</div> <div>Skeletonema potamos GCA_036940105.1</div> <div>Skeletonema tropicum GCA_037178625.1</div> <div>Stephanocyclus meneghinianus GCA_036940045.1</div> <div>Stephanodiscus minutulus GCA_036939435.1</div> <div>Stephanodiscus triporus GCA_036939755.1</div> <div>Thalassiosira allenii GCA_036939655.1</div> <div>Thalassiosira delicatula GCA_036939835.1</div> <div>Thalassiosira exigua GCA_036939895.1</div> <div>Thalassiosira gravida GCA_037356215.1</div> <div>Thalassiosira livingstoniorum GCA_036939595.1</div> <div>Thalassiosira mediterranea GCA_036939795.1</div> <div>Thalassiosira oceanica GCA_019693575.1</div> <div>Thalassiosira ordinaria GCA_036939695.1</div> <div>Thalassiosira pacifica GCA_036939875.1</div> <div>Thalassiosira profunda GCA_036939355.1</div> <p>&nbsp;</p> <h2>Converting to Protein FASTA and Coding Sequences FASTA</h2> <p>&nbsp;</p> <div>To save storage place at Zenodo, we did not upload the protein FASTA and coding sequence FASTA files. They can easily be generated from the genome FASTA file in combination with the respective GFF3 file. To do this, you can use the following commands:</div> <p>&nbsp;</p> <div><code># assume that genome.fa ist you respective genome FASTA file downloaded from NCBI datasets</code></div> <div><code>sed '/^&gt;/ s/ .*//' genome.fasta &gt; genome_short_headers.fasta</code></div> <div><code># assume that file.gff is the respective GFF3 file</code></div> <div><code>getAnnoFastaFromJoingenes.py -g genome_short_headers.fasta -3 file.gff -o nameStem</code></div> <div>&nbsp;</div> <div>This will produce the following files: nameStem.aa (protein FASTA file) and nameStem.codingseq (coding sequence FASTA file).</div> <div>&nbsp;</div> <div>The getAnnoFastaFromJoingenes.py script is available at https://raw.githubusercontent.com/Gaius-Augustus/Augustus/master/scripts/getAnnoFastaFromJoingenes.py . It is part of the AUGUSTUS software package.</div> <h2>Release notes</h2> <p>This release contains a gene set where a results of an OrthoFinder run that did not include genes on contigs that are suspected to be contaminants or horizontal gene transfer candidates were used to filter single exon genes. This means the gene and transcript counts changed compared to the previous release.</p> <h2>License</h2> <div>The genome annotation files are licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0). To view a copy of this license, visit http://creativecommons.org/licenses/by/4.0/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA.</div>

opencc-by-4.0Sep 2024View details →
zenodo48/100

Supplementary data for draft genome of a member of the ascomycotal fungal genus Pseudopithomyces (family Didymosphaeriaceae)

<p><span lang="EN-US">We update our previous draft genome of a member of genus <em>Pseudopithomyces</em> (previously annotated as <em>Pseudopithomyces maydicus</em> strain SBW1, now reannotated as <em>Pseudopithomyces sp</em>. strain SBW1. The new draft genome is based on a hybrid assembly utilising both ONT and Illumina data. The draft genome is comprised of 43 contigs with a total length of 39.65Mbp. We predict 13,669 protein coding gene models, of which 4241 (31%) were annotated to KEGG Orthology. Taxonomic assignment to <em>Pseudopithomyces sp.</em> was supported by comparative analysis of extracted ITS regions, mitochondrial DNA sequence and whole genome comparisons using <em>k</em>-mer sketches. </span></p> <p>&nbsp;</p> <p><span lang="EN-US">The following items of Additional Data Files are made available in this repository:</span></p> <p><span lang="EN-US">Additional Data File 1: contigs.fasta</span></p> <p><span lang="EN-US">FASTA file of entire assembly.&nbsp;</span><span lang="EN-US">&nbsp;</span></p> <p>&nbsp;</p> <p><span lang="EN-US">Additional Data File 2: draft_genome.fasta</span></p> <p><span lang="EN-US">FASTA file of draft whole genome sequence.</span></p> <p>&nbsp;</p> <p><span lang="EN-US">Additional Data File 3: ITS_full.fasta</span></p> <p><span lang="EN-US">FASTA file containing full length ITS sequences from contig 23 and contig 42.</span></p> <p><span lang="EN-US">&nbsp;</span></p> <p><span lang="EN-US">Additional Data File 4: 2NJ47W4U013-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for ITS region contained on contig 23</span></p> <p><span lang="EN-US">&nbsp;</span></p> <p><span lang="EN-US">Additional Data File 5: 2NJXH2GE013-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for ITS region contained on contig 42</span></p> <p><span lang="EN-US">&nbsp;</span></p> <p><span lang="EN-US">Additional Data File 6: 2NM79EUG016-Alignment.txt</span></p> <p><span lang="EN-US">Text file containing BLASTN alignments for the mitochondrial genome from contig 40 </span></p> <p><span lang="EN-US">&nbsp;</span></p> <p><span lang="EN-US">Additional Data File 6: sourmash_bc10_hy_pm1_3.txt</span></p> <p><span lang="EN-US">Text file containing the MASH similarities of the draft genome compared to 18,883 fungal genomes.</span></p>

opencc-by-4.0Nov 2022View details →
zenodo48/100

Introduction to Ancient Metagenomics Textbook (Edition 2025): de novo Genome Assembly

<p>Data and conda software environment file for the chapter &#39;<em>de novo</em> Genome Assembly&#39; of the SPAAM Community&#39;s textbook: Introduction to Ancient Metagenomics (https://www.spaam-community.org/intro-to-ancient-metagenomics-book).</p>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record