Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
50
datasets available to search
ShareScore release 0.7.1
Dataset results
50 results for “genomic alignment”
KMA Mapping and alignment statistics : livestock fecal metagenomes against ResFinder and genomes
<p>Three zip archives are included used in the analysis of the European livestock resistome.</p> <p>Two of them contain 'mapstat' files produced by the KMA software using the 'extended features' flag.<br> Each mapstat file thus summarize the mapping and alignment statistics when using KMA on a metagenome against a database.</p> <p>The last archive contains the 'refdata' file used to annotate the genomic mapstat hits. It encodes the taxonomic affilication of sequences hit by one or more samples.<br> </p>
Genome alignments for the project "Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol" - Protocol optimization
<p>Genome alignments for data generated in the project "<em>Whole transcriptome analysis of thousands of FACS-sorted single cells with the single cell nanoCAGE protocol – Optimization of the protocol.</em>" Files names indicate unique identifiers of MOIRAI workflow runs, with the following structure: library name, dot, workflow ID (OP-WORKFLOW-CAGEscan-short-reads-v2.0.), dot, timestamp. The raw (FASTQ) data of each library is also deposited in Zenodo (<a href="https://doi.org/10.5281/zenodo.250156">10.5281/zenodo.250156</a>). Library names correspond to the following runs:</p> <ul> <li> NC33: 151007_M00528_0161_000000000-AEBDC</li> <li> NC37: 151204_M00528_0173_000000000-AEBEF</li> <li> NC38: 151211_M00528_0175_000000000-AE9PJ</li> <li> NC39: 160122_M00528_0185_000000000-AEB18</li> <li> NC42: 160302_M00528_0192_000000000-AELYK</li> </ul> <p>This data can be analysed using the "CAGEr" software package available from Bioconductor. The "multiplex_files.zip" file contains tables indicating which samples are biological replicates of each other or negative controls.</p>
Alignments from "Caecilian genomes reveal molecular basis of adaptation and convergent evolution of limblessness in vertebrates"
<p>Compressed file containing the alignments at both nucleotide and amino acid level for the manuscript "Caecilian genomes reveal molecular basis of adaptation and convergent evolution of limblessness in vertebrates" </p>
Genome alignments for the article "Targeted reduction of highly abundant transcripts with pseudo-random primers"
<p>Sequence alignment (Moirai workflow management) for the article "Targeted reduction of highly abundant transcripts with pseudo-random primers". File names indicate unique run identifiers. In the manuscript, shorter names are used:</p> <ul> <li>NC12: NC12_1.CAGEscan_short-reads.20150629125015</li> <li>NC17: NC16-17_1.CAGEscan_short-reads.20150625154740</li> <li>NC22b: NC22b.CAGEscan_short-reads.20150625152335</li> <li>NCki: NCms10058_1.CAGEscan_short-reads.20150625154711</li> </ul>
Raw data used for COI delineation of the Eupolybothrus species: Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013
<p>Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar</p>
Dataset for the Galaxy Training Network (GTN) Tutorial "Viewing Cancer Alignments in a Genome Browser"
<p>Datasets for the Galaxy Training Network (GTN) Tutorial "Viewing Cancer Alignments in a Genome Browser"</p>
Additional annotation, alignment, and results from Ka/Ks analysis for Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus)
<p><strong>Annotation files, alignments, and results summaries from Chromosomal-level reference genome assembly of the African Spiny Mouse (Acomys cahirinus).</strong></p> <p>Pairwise genome alignments contain the .maf suffix</p> <p>FASTA alignments from stitched gene blocks contain the .fasta suffix</p> <p>CSV file containing the Ka/Ks results</p> <p>RepeatMasker .out file</p>
CASTER: Direct species tree inference from whole-genome alignments
Open the record for dataset details and reuse information.
The alignments of chloroplast genome sequences and nuclear ribosomal DNA fragments of six oak species sampled in the hot-dry valley of the Jinsha River, southwestern China
<p>Both chloroplast (cp) genome sequences and nuclear ribosomal (nr) DNA were assembled using GetOrganelle v.1.7.6.1 for 18 oak trees sampled in the Panzhihua Cycad National Nature Reserve, Sichuan Province, China. These trees belong to six oak species, including Quercus cocciferoides, Q. dolicholepis, Q. franchetii, Q. griffithii, Q. longispica, and Q. variabilis. We used PhyloSuite v.1.1.152 to extract coding sequences (CDSs), tRNA genes, rRNA genes, introns, and intergenic spacers (IGSs) of the 18 oak cp genomes. These sequences were aligned separately using MAFFT v.7.3.13 and manually adjusted with BioEdit v.7.2.5. Length variations in mononucleotide repeats were excluded and inversions were replaced with their reverse complements because of their tendency for homoplasy. Other indels were coded as binary characters according to the simple gap coding method using GapCoder. Separate assignments were concatenated according to their respective positions in the cp genome to obtain the alignments of LSC, SSC, IRb, and the whole cp genome.</p>
Supplemental Material for Genome Editing in Crop Plant Research - Alignment of expectations and current developments
<p>Supplemental Material for Paper "Genome Editing in Crop Plant Research - Alignment of expectations and current developments" as submitted to Plants</p>
45 fish genome alignment
<p>A visual representation of an all-to-all alignment of 45 fish genomes. Generated in two days on a single 48-core compute node using wfmash v0.7.0+2bbd135. Plot generated with pafplot -s 20000. Represented genomes:</p> <p>fArcCen1<br> fNotCel1.pri<br> sPriPec2.1.pri<br> fCycLum1.pri<br> fHipHip1.pri<br> fPerMag1.pri<br> fEsoLuc1.pri<br> fAngAng1.pri<br> fAntMac1.pri<br> fEleEle1.pri<br> fMegCyp1.pri<br> fAnaAna1.pri<br> fXenCan1.pri<br> fPygNat1.pri<br> fSebUmb1.pri<br> sCarCar2.pri<br> fAplTae1.pri<br> fMelBoe1.pri<br> fCheRos1.pri<br> fToxJac2.pri<br> fAstCal1.3<br> fAnaTes1.3<br> fMasArm1.3<br> fCotGob3.1<br> fParRan2.2<br> fGouWil2.2<br> fBetSpl5.3<br> fDenClu1.2<br> fErpCal1.2<br> fSpaAur1.2<br> fEcheNa1.2<br> fSclFor1.1<br> fTakRub1.3<br> fSalTru1.2<br> fSynAcu1.2<br> fSalaFa1.1<br> fSphaOr1.1<br> fMyrMur1.1<br> gadMor3.0<br> fChaCha1.1<br> fThaAma1.1<br> fDreTuH1.2<br> fDreABH1.1<br> fAcaLat1.1<br> fTraTra1.1<br> </p>
Corallium rubrum genome and transcripts alignment
<p><em>Corallium rubrum</em>, the precious red coral, is an octocoral endemic to the western Mediterranean Sea. It plays a key role in the well-known Mediterranean coralligenous ecosystem, a biodiversity hotspot. We sequenced its genome for further investigations into its biology (skeleton formation and color, genetic diversity), ecology, and evolutionary history.</p>
List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)
<p>The profiling of epigenetic marks like DNA methylation has become a central aspect of studies in evolution and ecology. Bisulfite sequencing is commonly used for assessing genome-wide DNA methylation at single nucleotide resolution but these data can also provide information on genetic variants like single nucleotide polymorphisms (SNPs). However, bisulfite conversion causes unmethylated cytosines to appear as thymines, complicating the alignment and subsequent SNP calling. Several tools have been developed to overcome this challenge, but there is no independent evaluation of such tools for non-model species, which often lack genomic references. Here, we used whole-genome bisulfite sequencing (WGBS) data from four female great tits (<i>Parus major</i>) to evaluate the performance of seven tools for SNP calling from bisulfite sequencing data. We used SNPs from whole-genome resequencing data of the same samples as baseline SNPs to assess common performance metrics like sensitivity, precision, and the number of true positive, false positive, and false negative SNPs for the full range of variant and genotype quality values. We found clear differences between the tools in either optimizing precision (Bis-SNP), sensitivity (biscuit), or a compromise between both (all other tools). Overall, the choice of SNP caller strongly depends on which performance parameter should be maximized and whether ascertainment bias should be minimized to optimize downstream analysis, highlighting the need for studies that assess such differences.</p>
Genome alignments for 'Machine-driven parameter-space exploration of biochemical reactions'
<p>The development of complex, multi-step <em>omics</em> methods in molecular biology is a laborious, costly, iterative and often intuition-bound process where an optimum is sought in a parameter space through step-by-step optimisations. The the difficulty of miniaturising assays and the cost of the experiments limit the dynamic range and the number of parameters that can be explored. However, because of non-linearities of the response of biochemical systems to their reagent concentrations, a broad dynamic range is necessary. Here we demonstrate the use of a high-performance nanoliter handling platform (Labcyte Echo 525) and computer generation of liquid transfer programs to explore in quadruplicates more than 600 combination of 4 parameters of a biochemical reaction, which lead us to uncover non-linear responses, parameter interactions and novel mechanical insights. With the increased availability of « <em>cloud biology</em> » computer-driven laboratory platforms, our results participate in changing methods development for biotechnology towards reproducible, computer-aided exhaustive characterisation of biochemical systems.</p> <p>This dataset contains the sequence alignments and other processing files produced by running the raw data (10.5281/zenodo.1680999) through a processing pipeline using the MOIRAI workflow manager. The most important output is the "CAGEscan_fragments" directories and represent the alignment of single mRNA molecules, which can be further analysed using the "CAGEr" software package available from Bioconductor.</p> <p>Run IDs: 171227_M00528_0321_000000000-B4GLP, 180123_M00528_0325_000000000-B4PCK, 180326_M00528_0346_000000000-B4GJR, 180403_M00528_0348_000000000-B4GP8, 180411_M00528_0351_000000000-BN3BL, 180501_M00528_0359_000000000-B4PJY, 180517_M00528_0364_000000000-BRGK6 180606_M00528_0367_000000000-BN3FG, 180607_M00528_0368_000000000-BN9KM</p>
Selecting a window size for the analysis of whole genome alignments using AIC
Open the record for dataset details and reuse information.
Corallium rubrum genome and transcripts alignment
Open the record for dataset details and reuse information.
List of known SNP positions (based on SNP chip data) for base quality score recalibration of alignments for whole-genome resequencing and whole-genome bisulfite sequencing data from great tits (Parus major)
Open the record for dataset details and reuse information.
Supplementary tables S5, S7, S9, S10, original protein models fasta files used for alignments, aligned and manually curated protein modes files used for phylogenies (PHYLIP format), and phylogenetic trees of plant cell wall decomposition gene families from 44 basidiomycete genomes (.tre files)
<p><span><span><span><span><span><span><span><span><span><span><span>Litter-decomposing Agaricales play key role in terrestrial carbon cycling, but little is known about their decomposition mechanisms. We assembled datasets of 42 gene families involved in plant-cell-wall decomposition from seven newly sequenced litter decomposers and 35 other Agaricomycotina members, mostly white-rot and brown-rot species. Using sequence similarity and phylogenetics, we split the families into phylogroups and compared their gene composition across nutritional strategies. Subsequently, we used Raman spectroscopy to examine the ability of litter decomposers, white-rot fungi, and brown-rot fungi to decompose crystalline cellulose. Both litter decomposers and white-rot fungi share the enzymatic cellulose decomposition, whereas brown-rot fungi possess a distinct mechanism that disrupts cellulose crystallinity. However, litter decomposers and white-rot fungi differ with respect to hemicellulose and lignin degradation phylogroups, suggesting adaptation of the former group to the litter environment. Litter decomposers show high phylogroup diversity, which is indicative of high functional versatility within the group, whereas a set of white-rot species shows adaptation to bulk-wood decomposition. In both groups, we detected species that have unique characteristics associated with hitherto unknown adaptations to diverse wood and litter substrates. Our results suggest that the terms white-rot fungi and litter decomposers mask a much larger functional diversity.</span></span></span></span></span></span></span></span></span></span></span></p>
Nucleotide alignments of eight meiosis genes under extreme selection following whole genome duplication in Arabidopsis lyrata/A.arenosa.
<p>In this study we performed a genotype-phenotype association analysis of meiotic stability in 10 autotetraploid <em>Arabidopsis lyrata</em> and <em>A</em>. <em>lyrata/A</em>. <em>arenosa</em> hybrid populations collected from the Wachau region and East Austrian Forealps. The aim was to determine the effect of eight meiosis genes under extreme selection upon adaptation to whole genome duplication. Individual plants were genotyped by high-throughput sequencing of the eight meiosis genes (<em>ASY1</em>, <em>ASY3</em>, <em>PDS5b</em>, <em>PRD3</em>, <em>REC8</em>, <em>SMC3</em>, <em>ZYP1a/b</em>) implicated in synaptonemal complex formation and phenotyped by assessing meiotic metaphase I chromosome configurations. Our results reveal that meiotic stability varied greatly (20–100%) between individual tetraploid plants and associated with segregation of a novel <em>ASYNAPSIS3</em> (<em>ASY3</em>) allele derived from <em>A</em>. <em>lyrata</em>. The <em>ASY3</em> allele that associates with meiotic stability possesses a putative in-frame tandem duplication (TD) of a serine-rich region upstream of the coiled-coil domain that appears to have arisen at sites of DNA microhomology. The frequency of multivalents observed in plants homozygous for the <em>ASY3 TD</em> haplotype was significantly lower than in plants heterozygous for <em>ASY3 TD/ND</em> (non-duplicated) haplotypes. The chiasma distribution was significantly altered in the stable plants compared to the unstable plants with a shift from proximal and interstitial to predominantly distal locations. The number of HEI10 foci at pachytene that mark class I crossovers was significantly reduced in a plant homozygous for <em>ASY3 TD</em> compared to a plant heterozygous for <em>ASY3 ND/TD</em>. Fifty-eight alleles of the 8 meiosis genes were identified from the 10 populations analysed, demonstrating dynamic population variability at these loci. Widespread chimerism between alleles originating from <em>A</em>. <em>lyrata/A</em>. <em>arenosa</em> and diploid/tetraploids indicates that this group of rapidly evolving genes may provide precise adaptive control over meiotic recombination in the tetraploids, the very process that gave rise to them.</p>
Genome reduction is associated with bacterial pathogenicity across different scales of temporal and ecological divergence - between species core gene alignments
<p><span>Emerging bacterial pathogens threaten global health and food security, and so it is important to ask whether these transitions to pathogenicity have any common features. We present a systematic study of the claim that pathogenicity is associated with genome reduction and gene loss. We compare broad-scale patterns across all bacteria, with detailed analyses of <i>Streptococcus suis</i>, an emerging zoonotic pathogen of pigs, which has undergone multiple transitions between disease and carriage forms. We find that pathogenicity is consistently associated with reduced genome size across three scales of divergence (between species within genera, and between and within genetic clusters of <i>S. suis</i>). While genome reduction is also found in mutualist and commensal bacterial endosymbionts, genome reduction in pathogens cannot be solely attributed to the features of their ecology that they share with these species, i.e. host restriction or intracellularity. Moreover, other typical correlates of genome reduction in endosymbionts (reduced metabolic capacity, reduced GC content, and the transient expansion of non-functional elements) are not consistently observed in pathogens. Together, our results indicate that genome reduction is a predictive marker of pathogenicity in bacteria.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.