Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,848
datasets available to search
ShareScore release 0.9.0
Dataset results
2,848 results for “sequence data”
Data set for the paper "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences"
<p>Data set for the paper "Temperatures and cooling rates recorded by the New Caledonia ophiolite: implications for cooling mechanisms in young forearc sequences" by Secchiari et al.</p>
Raw mass-standardized ionomic data from seven fish species and raw transcriptome sequences for mosquitofish inhabiting the Tar Creek Superfund Site in OK, USA
<p>Our understanding of the mechanisms mediating the resilience of organisms to environmental change remains lacking. Heavy metals negatively affect processes at all biological scales, yet organisms inhabiting contaminated environments must maintain homeostasis to survive. Tar Creek in Oklahoma, USA, contains high concentrations of heavy metals and an abundance of Western mosquitofish (<em>Gambusia affinis)</em>, though several fish species persist at lower frequency. To test hypotheses about the mechanisms mediating the persistence and abundance of mosquitofish in Tar Creek, we integrated ionomic data from seven resident fish species and transcriptomic data from mosquitofish to test hypotheses about the mechanisms mediating the persistence of mosquitofish in Tar Creek. We predicted that mosquitofish minimize uptake of heavy metals more than other Tar Creek fish inhabitants and induce transcriptional responses to detoxify metals that enter the body, allowing them to persist in Tar Creek at higher density than species that may lack these responses. Tar Creek populations of all seven fish species accumulated heavy metals, suggesting mosquitofish cannot block uptake more efficiently than other species. We found population-level gene expression changes between mosquitofish in Tar Creek and nearby unpolluted sites. Gene expression differences primarily occurred in the gill, where we found upregulation of genes involved with lowering transfer of metal ions from the blood into cells and mitigating free radicals. However, many differentially expressed genes were not in known metal response pathways, suggesting multifarious selective regimes and/or previously undocumented pathways could impact tolerance in mosquitofish. Our systems-level study identified well characterized and putatively new mechanisms that enable mosquitofish to inhabit heavy metal-contaminated environments.</p>
Complete Genome Sequence of an Aeromonas rivuli Strain Isolated from Ready-to-Eat Food - Data Files
<p>This dataset contains input and intermediate files of the bcgTree analysis described in the Schwartz <em>et al</em>. MRA manuscript entitled “Complete Genome Sequence of an <em>Aeromonas rivuli</em> Strain Isolated from Ready-to-Eat Food”.</p> <table> <tbody> <tr> <td> <p><strong>File name</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>‘<em>Aeromonadaceae</em> identifier’.fa</p> </td> <td> <p>Amino acid FASTA file of the translated CDS sequences of a strain X (bcgTree input file)</p> </td> </tr> <tr> <td> <p>full_alignment.concat.fa</p> </td> <td> <p>Alignment of the concatenated amino acid sequences of 107 single-copy core genes that is used for phylogenetic tree calculation in bcgTree (bcgTree intermediate file)</p> </td> </tr> </tbody> </table> <p> </p> <p>In the bcgTree files, the <em>Aeromonadaceae</em> sequences were named/abbreviated as follows:</p> <table> <tbody> <tr> <td> <p><strong>Sequence name in the bcgTree file</strong></p> </td> <td> <p><strong>Description</strong></p> </td> </tr> <tr> <td> <p>AAZUK01-1</p> </td> <td> <p><em>Tolumonas lignilytica </em>BRL6-1</p> </td> </tr> <tr> <td> <p>Aeromonas-caviae</p> </td> <td> <p><em>Aeromonas caviae </em>NCTC 12244</p> </td> </tr> <tr> <td> <p>Aeromonas-dhakensis</p> </td> <td> <p><em>Aeromonas dhakensis </em>CIP 107500</p> </td> </tr> <tr> <td> <p>Aeromonas-hydrophila</p> </td> <td> <p><em>Aeromonas hydrophila </em>ATCC 7966</p> </td> </tr> <tr> <td> <p>Aeromonas-rivuli</p> </td> <td> <p><em>Aeromonas rivuli </em>DSM 22539</p> </td> </tr> <tr> <td> <p>Aeromonas-rivuli-20-VB00005</p> </td> <td> <p><em>Aeromonas rivuli </em>20-VB00005</p> </td> </tr> <tr> <td> <p>Aeromonas-veronii</p> </td> <td> <p><em>Aeromonas veronii </em>CECT 4257</p> </td> </tr> <tr> <td> <p>CP001616-1</p> </td> <td> <p><em>Tolumonas auensis </em>DSM 9187</p> </td> </tr> <tr> <td> <p>JACHGR01-1</p> </td> <td> <p><em>Tolumonas osonensis </em>DSM 22975</p> </td> </tr> </tbody> </table>
FIGURE. Phylogenetic tree of specimens on Poaceae and related host plants constructed by MP method based on ITS+28S regions of rDNA. Bootstrap values of MP and ML are followed by the Bayesian posterior probabilities (Bpp) on the nodes in the topology. Asterisk (*) represents bootstrap values or Bpp less than 50% in the topology. Sample data are shown with voucher specimen number or GenBank accession number, and host plant. Sequence data determined in this study are shown in color. Teliospore shapes are shown in each clade detected, and new species are shown by asterisk (*) on clades. 0, I: Spermogonial and aecial host genus. Asterisk (*) on host plants: Spermogonial and aecial host plants. in Phylogenetic approach for identification and life cycles of Puccinia (Pucciniaceae) species on Poaceae from northeastern China
FIGURE. Phylogenetic tree of specimens on Poaceae and related host plants constructed by MP method based on ITS+28S regions of rDNA. Bootstrap values of MP and ML are followed by the Bayesian posterior probabilities (Bpp) on the nodes in the topology. Asterisk (*) represents bootstrap values or Bpp less than 50% in the topology. Sample data are shown with voucher specimen number or GenBank accession number, and host plant. Sequence data determined in this study are shown in color. Teliospore shapes are shown in each clade detected, and new species are shown by asterisk (*) on clades. 0, I: Spermogonial and aecial host genus. Asterisk (*) on host plants: Spermogonial and aecial host plants.
Data from: Genome-wide sequence-based genotyping supports a nonhybrid origin of Castanea alabamensis
<p>The genus Castanea in North America contains multiple tree and shrub taxa of conservation concern. The two species within the group, American chestnut (Castanea dentata) and chinquapin (C. pumila sensu lato), display remarkable morphological diversity across their distributions in the eastern United States and southern Ontario. Previous investigators have hypothesized that hybridization between C. dentata and C. pumila has played an important role in generating morphological variation in wild populations. A putative hybrid taxon, Castanea alabamensis, was identified in northern Alabama in the early 20th century; however, the question of its hybridity has been unresolved. We tested the hypothesized hybrid origin of C. alabamensis using genome-wide sequence-based genotyping of C. alabamensis, all currently recognized North American Castanea taxa, and two Asian Castanea species at >100,000 single-nucleotide polymorphism (SNP) loci. With these data, we generated a high-resolution phylogeny, tested for admixture among taxa, and analyzed population genetic structure of the study taxa. Bayesian clustering and principal components analysis provided no evidence of admixture between C. dentata and C. pumila in C. alabamensis genomes. Phylogenetic analysis of genome-wide SNP data indicated that C. alabamensis forms a distinct group within C. pumila sensu lato. Our results are consistent with the model of a nonhybrid origin for C. alabamensis. Our finding of C. alabamensis as a genetically and morphologically distinct group within the North American chinquapin complex provides further impetus for the study and conservation of the North American Castanea species.</p>
Data from: Genotyping-in-Thousands by sequencing panel development and application to inform kokanee salmon (Oncorhynchus nerka) fisheries management at multiple scales
<p>The ability to differentiate life history variants is vital for estimating fisheries management parameters, yet traditional survey methods can be inaccurate in mixed-stock fisheries. Such is the case for kokanee, the resident freshwater form of sockeye salmon (<i>Oncorhynchus nerka</i>), which exhibits various reproductive ecotypes (stream-, shore-, deep-spawning) that co-occur with each other and/or anadromous <i>O. nerka</i> in some systems across their pan-Pacific distribution. Here, we developed a multi-purpose Genotyping-in-Thousands by sequencing (GT-seq) panel of 288 targeted single nucleotide polymorphisms (SNPs) to enable accurate kokanee stock identification by geographic basin, migratory form, and reproductive ecotype across British Columbia, Canada. The GT-seq panel exhibited high self-assignment accuracy (93.3%) and perfect assignment of individuals not included in the baseline to their geographic basin, migratory form, and reproductive ecotype of origin. The GT-seq panel was subsequently applied to Wood Lake, a valuable mixed-stock fishery, revealing high concordance (>98%) with previous assignments to ecotype using microsatellites and TaqMan<span> </span>SNP genotyping assays, while improving resolution, extending a long-term time-series, and demonstrating the scalability of this approach for this system and others.</p>
Deep sequencing data for document titled: Rolling circle RNA synthesis catalysed by RNA
<p>RNA-catalysed RNA replication is widely considered a key step in the emergence of life's first genetic system. However, RNA replication can be impeded by the extraordinary stability of duplex RNA products, which must be dissociated for re-initiation of the next replication cycle. Here we have explored rolling circle synthesis (RCS) as a potential solution to this strand separation problem. RCS on small circular RNAs - as indicated by molecular dynamics simulations - induces a progressive build-up of conformational strain with destabilisation of nascent strand 5' and 3' ends. At the same time, we observe sustained RCS by a triplet polymerase ribozyme on small circular RNAs over multiple orbits with strand displacement yielding concatemeric RNA products. Furthermore, we show RCS of a circular Hammerhead ribozyme capable of self-cleavage and re-circularisation. Thus, all steps of a viroid-like RNA replication pathway can be catalysed by RNA alone. Our results have implications for the emergence of RNA replication and for understanding the potential of RNA to support complex genetic processes.</p>
FIGURE 3 in Two new species of Hypoxylon (Hypoxylaceae) from China based on morphological and DNA sequence data analyses
FIGURE 3. Hypoxylon jianfengense (Holotype FACATAS 845). A. Stromata on wood. B. Stromatal surface and ostioles. C, D. Stroma in vertical section showing the perithecia and tissue below the perithecial layer. E. KOH-extractable pigments. F. Stromatal granules in water. G. Mature and immature asci in water. H. Asci in Melzer's reagent. I. Immature asci in water. J, K. Mature asci in water. L. Apical apparatus in Melzer' s reagent. M. Ascospore in water showing germ slit. N. Ascospores in water. O. Ascospores in 10% KOH. P. Ascospore under SEM. Bars: A = 5 mm; B = 0.3 mm; C = 0.5 mm; D = 0.1 mm; G, I–K = 20 µm; H, L–O = 10 µm; P = 2.5 µm.
FIGURE 1 in Two new species of Hypoxylon (Hypoxylaceae) from China based on morphological and DNA sequence data analyses
FIGURE 1. Phylogenetic tree of Hypoxylon based on the multigene alignment of ITS-LSU-RPB2-TUB2 in the Maximum Likelihood analyses (RaxML). Support values of Maximum Likelihood (ML), Maximum Parsimony (MP) and Bayesian (B) analyses (bootstrap support above 50%, posterior probability value above 0.95) are displayed above or below the respective branches (ML/MP/BA).
FIGURE 2 in Two new species of Hypoxylon (Hypoxylaceae) from China based on morphological and DNA sequence data analyses
FIGURE 2. Hypoxylon larissae (Holotype FACATAS 844). A. Stromata on wood. B, C. Stromatal surface and ostioles. D, E. Stroma in vertical section showing the perithecia and tissue below the perithecial layer. F. KOH-extractable pigments. G. Stromatal granules in water. H. Mature and immature asci in Melzer's reagent. I. Asci in Melzer's reagent. J. Immature asci in water. K. Mature asci in water. L. Ascospores in water. M. Ascospores in 10% KOH. N. Ascospores in water showing germ slit. O. Apical apparatus in Melzer's reagent. P. Ascospore under SEM. Bars: A = 5 mm;B, C, E = 0.4 mm; D = 1 mm; H–K = 20 µm; L–O = 10 µm; P = 5 µm.
Raw data of sequencing results of our study: Bovine milk microbiota: Evaluation of different DNA extraction protocols in challenging samples
<p>Clean reads of the repeated milk samples with used Primer Pairs V1V2 and V3V4</p> <p>Raw data of sequencing results (amplicon single variants)</p>
Data for "Generative and interpretable machine learning for aptamer design and analysis of in vitro sequence selection"
<p>Once decompressed, the file contains a folder which contains:</p> <ul> <li>The files "s100_Nth.fasta" (where "N" is 5, 6, 7 or 8), which are the output of the SELEX experiment described in the paper with DOI: <a href="https://doi.org/10.1002/cbic.201900265">10.1002/cbic.201900265</a>. They are standard fasta files, and the descriptor of each sequence is of the form "seqX-Y", where "X" is an increasing label, and "Y" is the number of times "seqX" has been obtained (number of counts of "seqX").</li> <li>The file "Aptamer_Exp_Results.csv", which contains the sequences tested experimentally for the paper "Generative and interpretable machine learning for aptamer design and analysis of in vitro sequence selection" (preprint available at https://doi.org/10.1101/2022.03.12.484094), with the following experimental results for each sequence: (i) whether the sequence was able to bind thrombin ('B' for binders, 'NB' for non-binders); (ii) the thrombin exosite used for binding ('I' for exosite I, 'II' for exosite II, 'n/a' for sequences not tested).</li> </ul> <p>Examples of usage of the data are available at https://github.com/adigioacchino/RBMsForAptamers.</p>
Magnetic bead epicPCR with dMLA probes, sequencing data, March 11 2022
<p>Samples:</p> <p>1. Mocks, 16S, 55°C<br> 2. HAMBIs+Mocks, 16S, 55°C<br> 3. HAMBIs, 16S, 55°C<br> 4. Mocks, 16S, 60°C<br> 5. HAMBIs+Mocks, 16S, 60°C<br> 6. HAMBIs, 16S, 60°C<br> 7. Mocks, 16S, 65°C<br> 8. HAMBIs+Mocks, 16S, 65°C<br> 9. HAMBIs, 16S, 65°C<br> 10. Mocks, dMLA, 55°C<br> 11. HAMBIs+Mocks, dMLA, 55°C<br> 12. HAMBIs, dMLA, 55°C<br> 13. Mocks, dMLA, 60°C<br> 14. HAMBIs+Mocks, dMLA, 60°C<br> 15. HAMBIs, dMLA, 60°C<br> 16. Mocks, dMLA, 65°C<br> 17. HAMBIs+Mocks, dMLA, 65°C<br> 18. HAMBIs, dMLA, 65°C<br> 19. Chilomonas, 16S<br> 20. Chilomonas+WW, 16S<br> 21. WW, 16S<br> 22. Chilomonas, 18S<br> 23. Chilomonas+WW, 18S<br> 24. WW, 18S</p>
Color and W-chromosome sequence data from study on maternal inheritance of egg mimicry
<div>This dataset supports a study demonstrating that host-specific egg mimicry in the brood-parasitic African cuckoo finch <em>Anomalospiza imberbis</em> is maternally inherited. It includes egg reflectance spectra for the background colour of 188 cuckoo finch eggs from four host species in Zambia, and consensus sequences for 68 W-linked ddRAD-seq loci derived from 80 female cuckoo finches belonging to four different host-specific maternal lineages from three host species in Zambia. These data derive from two partially overlapping samples of eggs: some eggs with genetic data lacked egg spectral data, and vice versa. W-linked genetic data were all of offspring origin as they derived from embryonic or nestling tissue. Additional phenotypic data (host nest species and descriptions of egg phenotype), date and location data associated with each egg spectrum are provided in a separate file. Data on the origin of the individuals contributing to the W-linked loci are provided in Table S1 of the associated publication.</div>
Substitution mutational signatures in whole-genome-sequenced cancers in the UK population, Mutational Signatures Data
<p>This uploads contains the mutational signature data from the article <strong>Substitution mutational signatures in whole-genome-sequenced cancers in the UK population</strong>,<strong> </strong><em>Science</em>, doi:10.1126/science.abl9283, 2022.</p>
Source data and code for manuscript 'An executive network for the control of sequence-behavior in pigeons'
<p>The contents of this folder are part of the submission of the manuscript entitled 'An executive network for the control of sequence-behavior in pigeons', by Lukas Alexander Hahn & Jonas Rose</p> <p>Contact: lukas.hahn@ruhr-uni-bochum.de</p> <p>Data and code have been compressed into a .zip folder each. Unpack the contents of the folders to use the dataset. The dataset is split into two main folders and one Matlab file:</p> <p>'code'<br> contains all analysis code to produce all figures and reported statistics of the manuscript (refer to the<br> MATLAB live script 'manuscriptResultsLiveScript.mlx' to run the analysis, please adjust the path information of where the data is stored on your computer).</p> <p>'RESULTSSTATISTICS.mat'<br> contains all reported statistical values (generated by 'manuscriptResultsLiveScript.mlx')</p> <p>'sourceData'<br> Contains all required source data files (i.e. pre-processed data) required to run the analyses stored in 'code'.</p> <p>Data related to animal behavior was recorded using MATLAB (R2016b). Electrophysiological data was recorded by NeuroNexus microelectrodes and an INTAN RHD2000 headstage on an INTAN USB-Interface board, with a sampling rate of 30 kHz and was subsequently filtered for spike sorting at bandpass 0.5 - 7.5 kHz.</p> <p>Data format is the MATLAB '.mat' type (which can be loaded in by MATLAB, or alternatively by the freely available Octave Software (https://www.gnu.org/software/octave)).<br> Data is organized in MATLAB structures, one file per session for behavioral results, one file per neuron for different alignments and preprocessing conditions (refer to manuscriptResultsLiveScript).<br> Structures contain individual matrices (labelled by a descriptive name) that contain numerical values or character strings.<br> Matrices labelled by the keyword 'Info' contain character strings that give a brief description of the loaded data.<br> Source data contains two separate folders containing data of animal 1 ('P855'), and animal 2 ('T1003').</p> <p>Data was sorted into different subsets, for analysis of individual task phases. Subfolder 'NCL' refers to 'nidopallium caudolaterale', 'NIML' refers to 'nidopallium intermedium mediale pars laterale', the recorded brain regions.</p> <p> </p>
Data from: Exploring the phylogeography of a hexaploid freshwater fish by RAD sequencing
The KwaZulu-Natal yellowfish (Labeobarbus natalensis) is an abundant cyprinid, endemic to KwaZulu-Natal Province, South Africa. In this study we developed a Single Nucleotide Polymorphism (SNP) dataset from double-digest Restriction-site Associated DNA (ddRAD) sequencing of samples across the distribution. We addressed several hidden challenges, primarily focussing on proper filtering of RAD data and selecting optimal parameters for data processing in polyploid lineages. We used the resulting high-quality SNP dataset to investigate the population genetic structure of L. natalensis. A small number of mitochondrial markers present in these data had disproportionate influence on the recovered genetic structure. The presence of singleton SNPs also confounded genetic structure. We found a well-supported division into northern and southern lineages, with further subdivision into five populations, one of which reflects north-south admixture. Approximate Bayesian Computation scenario testing supported a scenario where an ancestral population diverged into northern and southern lineages, which then diverged to yield the current five populations. All river systems showed similar levels of genetic diversity, which appears unrelated to drainage system size. Nucleotide diversity was highest in the smallest river system, the Mbokodweni, which, together with adjacent small coastal systems, should be considered as a key catchment for conservation.
Genome wide mRNA sequencing data in macrophages without and with CX-5461 treatment
<p class="MsoNormal"><span>CX-5461, a novel selective RNA polymerase I inhibitor, shows potential anti-inflammatory and immunosuppressive activities. However, the molecular mechanisms underlying the inhibitory effects of CX-5461 on macrophage-mediated inflammation remain to be clarified. In the present study, we attempted to identify the systemic biological processes which were modulated by CX-5461 in inflammatory macrophages. Primary peritoneal macrophages were isolated from normal Sprague Dawley rats, and primed with lipopolysaccharide or interferon-gamma. Genome-wide RNA sequencing was performed. The study suggests that limiting cell proliferation predominates in the inhibitory effects of CX-5461 on macrophage-mediated inflammation. </span></p>
Data for for Detecting cell-of-origin and cancer-specific methylation features of cell-free DNA from Nanopore sequencing
<p>Datasets accompanying the paper https://doi.org/10.1101/2021.10.18.464684</p>
Integration of single-cell RNA-sequencing data across tissues and cancer types towards immune cell characterization
<p>To better understand dendritic cell states and subtypes, we collected individual single-cell RNAseq datasets from various studies and further integrated, batch corrected, and reprocessed the data using Besca (https://github.com/bedapub/besca).</p> <p>The following files are included:<br> 1) study_table_integrated_DCs.xlsx - contains a list of studies from where the datasets were gathered.<br> 2) int_dcs.raw.h5ad - An anndata object file containing the combined raw single-cell counts for DCs from individual studies. The datasets were joined based on the union of variables.<br> 3) intersection_genes_integrated_dcs.tsv - List of genes if the datasets were joined based on the intersection of variables. These variables were used in the subsequent analyses.</p> <p>4) int_dcs.annotated.h5ad - An anndata object file containing single-cell logarithmized counts for DCs data that have been integrated and reprocessed. The rows of the file contain cells, and the columns contain highly variable genes. A sparse matrix containing the logarithmized counts from all the genes (from the intersection genes integrated dcs.tsv file) can also be found (adata.raw.X) in the object. In the observations, cell-type annotation is available at three different hierarchal levels.<br> <br> This data was further used to produce results for the publication (https://jitc.bmj.com/content/10/6/e004268) on the effects of Toll-like receptor 8 agonists on conventional DCs.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.