Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,848

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,848 results for “sequence data”

Learn how ShareScore rates datasets ↗
zenodo28/100

Fig. 2A–O in Phylogenetic Analyses on the Tintinnid Ciliates (Protozoa, Ciliophora) Based on Multigene Sequence Data

Fig. 2A–O. Secondary structure of the internal transcribed spacer 2 (ITS2) RNA transcript of: A – Strombidinopsis sp.; B – Amphorellopsis acuta; C – Eutintinnus pectinis; D – Stenosemella nivalis; E – Codonellopsis nipponica; F – Tintinnopsis lohmanni; G – T. cylindrica; H – T. tubulosoides; I – Tintinnopsis sp. 2; J – Tintinnopsis sp. 1; K – Favella taraikaensis; L – F. ehrenbergii; M – F. campanula; N – Metacylis angulata and O – Tintinnopsis sp. 3. The diagram illustrates that all these species have a similar ITS2 secondary structure model – one palm with two fingers. Tintinnopsis sp. 1 and Tintinnopsis sp. 3 have the same ITS2 secondary structure, so are shaded together. Positions labeled II that lack a bulge are marked with arrows. Note that a bulge is present in this position in other species.

opencc-by-4.0Dec 2012View details →
zenodo28/100

YUTO MMS Dataset, Sequence B (Lidar data, Calibration files, Groundtruth, GPS, IMU, Timestamp)

Open the record for dataset details and reuse information.

opencc-by-4.0Jan 2024View details →
zenodo28/100

Analysis of Public Short-Read RNA-Sequencing Data: PRJNA543316

<p><strong><span>PRJNA543316</span></strong></p> <p><span>Metformin is a front-line drug in the treatment of type-2 diabetes mellitus (T2DM). In addition to its antigluconeogenic and insulin-sensitizing properties, it has emerged as a potent inhibitor of the inflammatory response of macrophages. Specifically, metformin has been shown to reduce transcript levels of Il1b, the gene encoding the pro-inflammatory cytokine interleukin (IL)-1b, during long-term exposure of macrophages to the bacterial cell-wall component lipopolysaccharide (LPS). However, the extent to which metformin affects the early transcriptional response to LPS has never been investigated. Here, we show that metformin affects transcript levels of a large yet selective subset of LPS-responsive genes after only two hours of LPS exposure, mostly counteracting the effect of LPS rather than enhancing it. The affected genes are implicated in a variety of biological functions, in particular cellular movement and trafficking. Intriguingly, metformin affects transcript levels of Il1b at this early time point as well, but through a molecular mechanism fundamentally different from the regulation observed after longer exposure. While down-regulation of Il1b by metformin during the late stages of the LPS response has been shown to rely on stabilization of hypoxia-inducible factor (HIF)-1&alpha; and production of IL-10, Il1b inhibition at the early stage requires AMP-activated protein kinase (AMPK) activation but is independent of HIF-1&alpha; and IL-10. These results reveal an unexpected complexity in the anti-inflammatory properties of metformin and demonstrate that Il1b is down-regulated by distinct mechanisms in the early and late stages of the LPS response. Overall design: Bone-marrow-derived macrophages from WT C57Bl/6 mice were either left untreated, stimulated with 100 ng/ml LPS for 2 h or pretreated with 5 mM metformin for 6 h and then stimulated with 100 ng/ml LPS for 2 h.</span></p> <p><span>Pipeline: FastQ -&gt; FastQC -&gt; fastp -&gt; STAR -&gt; samtools -&gt; multiQC</span></p>

openAug 2024View details →
zenodo28/100

FIGURE 6 in Senecio kumaonensis (Asteraceae, Senecioneae) is a Synotis based on evidence from karyology and nuclear ITS/ETS sequence data

FIGURE 6. Distribution of Synotis penninervis (= Senecio kumaonensis) (●).

opennotspecifiedJan 2017View details →
zenodo28/100

The Sanger sequencing data of HCAPV-1

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo28/100

Sequence data and structural data utilized in the study and analysis of grain protein function prediction.

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo28/100

Mass cytometry and integration sequencing data and code from "Quantification of intrinsic regulatory factors refines human hematopoietic progenitor definitions and reveals early erythroid lineage priming"

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo28/100

Linked collectors and determiners for: Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae).

Natural history specimen data linked to collectors and determiners held within, "Three new species and DNA sequence data of the rare South American water beetle genus Adelphydraena Perkins, 1989 (Coleoptera: Hydraenidae)". Claims or attributions were made on Bionomia by volunteer Scribes, <a href="https://bionomia.net/dataset/fa172759-c7a8-4682-b992-8871c155eb3a">https://bionomia.net/dataset/fa172759-c7a8-4682-b992-8871c155eb3a</a> using specimen data from the dataset aggregated by the Global Biodiversity Information Facility, <a href="https://gbif.org/dataset/fa172759-c7a8-4682-b992-8871c155eb3a">https://gbif.org/dataset/fa172759-c7a8-4682-b992-8871c155eb3a</a>. Formatted as a Frictionless Data package.

opencc-zeroJan 2024View details →
dryad28/100

Data from: Assessment of a 16S rRNA amplicon Illumina sequencing procedure for studying the microbiome of a symbiont-rich aphid genus

The bacterial communities inhabiting arthropods are generally dominated by a few endosymbionts that play an important role in the ecology of their hosts. Rather than comparing bacterial species richness across samples, ecological studies on arthropod endosymbionts often seek to identify the main bacterial strains associated with each specimen studied. The filtering out of contaminants from the results and the accurate taxonomic assignment of sequences are therefore crucial in arthropod microbiome studies. We aimed here to validate an Illumina 16S rRNA gene sequencing protocol and analytical pipeline for investigating endosymbiotic bacteria associated with aphids. Using replicate DNA samples from 12 species (Aphididae: Lachninae, Cinara) and several controls, we removed individual sequences not meeting a minimum threshold number of reads in each sample and carried out taxonomic assignment for the remaining sequences. With this approach, we show that: i) contaminants accounted for a negligible proportion of the bacteria identified in our samples; ii) the taxonomic composition of our samples and the relative abundance of reads assigned to a taxon were very similar across PCR and DNA replicates for each aphid sample; in particular, bacterial DNA concentration had no impact on the results. Furthermore, by analysing the distribution of unique sequences across samples rather than aggregating them into operational taxonomic units (OTUs), we gained insight into the specificity of endosymbionts for their hosts. Our results confirm that Serratia symbiotica is often present in Cinara species, in addition to the primary symbiont, Buchnera aphidicola. Furthermore, our findings reveal new symbiotic associations with Erwinia and Sodalis-related bacteria. We conclude with suggestions for generating and analysing 16S rRNA gene sequences for arthropod endosymbiont studies.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Differentiating founder and chronic HIV envelope sequences

Significant progress has been made in characterizing broadly neutralizing antibodies against the HIV envelope glycoprotein Env, but an effective vaccine has proven elusive. Vaccine development would be facilitated if common features of early founder virus required for transmission could be identified. Here we employ a combination of bioinformatic and operations research methods to determine the most prevalent features that distinguish 78 subtype B and 55 subtype C founder Env sequences from an equal number of chronic sequences. There were a number of equivalent optimal networks (based on the fewest covarying amino acid (AA) pairs or a measure of maximal covariance) that separated founders from chronics: 13 pairs for subtype B and 75 for subtype C. Every subtype B optimal solution contained the founder pairs 178–346 Asn-Val, 232–236 Thr-Ser, 240–340 Lys-Lys, 279–315 Asp-Lys, 291–792 Ala-Ile, 322–347 Asp-Thr, 535–620 Leu-Asp, 742–837 Arg-Phe, and 750–836 Asp-Ile; the most common optimal pairs for subtype C were 644–781 Lys-Ala (74 of 75 networks), 133–287 Ala-Gln (73/75) and 307–337 Ile-Gln (73/75). No pair was present in all optimal subtype C solutions highlighting the difficulty in targeting transmission with a single vaccine strain. Relative to the size of its domain (0.35% of Env), the α4β7 binding site occurred most frequently among optimal pairs, especially for subtype C: 4.2% of optimal pairs (1.2% for subtype B). Early sequences from 5 subtype B pre-seroconverters each exhibited at least one clone containing an optimal feature 553–624 (Ser-Asn), 724–747 (Arg-Arg), or 46–293 (Arg-Glu).

opencc-zeroDec 2016View details →
dryad28/100

Data from: Next-generation polyploid phylogenetics: rapid resolution of hybrid polyploid complexes using PacBio single-molecule sequencing

Difficulties in generating nuclear data for polyploids have impeded phylogenetic study of these groups. We describe a high-throughput protocol and an associated bioinformatics pipeline (PURC: "Pipeline for Untangling Reticulate Complexes") that is able to generate these data quickly and conveniently, and demonstrate its efficacy on accessions from the fern family Cystopteridaceae. We conclude with a demonstration of the downstream utility of these data by inferring a multilabeled species tree for a subset of our accessions. We amplified four ~1kb-long nuclear loci and sequenced them in a parallel-tagged amplicon sequencing approach using the PacBio platform. PURC infers the final sequences from the raw reads via an iterative approach that corrects PCR and sequencing errors and removes PCR-mediated recombinant sequences (chimeras). We generated data for all gene copies (homeologs, paralogs, and segregating alleles) present in each of three sets of 50 mostly-polyploid accessions, for four loci, in three PacBio runs (one run per set). From the raw sequencing reads PURC was able to accurately infer the underlying sequences. This approach makes it easy and economical to study the phylogenetics of polyploids, and in conjunction with recent analytical advances, facilitates investigation of broad patterns of polyploid evolution.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Sequencing of the needle transcriptome from Norway spruce (Picea abies Karst L.) reveals lower substitution rates, but similar selective constraints in gymnosperms and angiosperms

BACKGROUND: A detailed knowledge about spatial and temporal gene expression is important for understanding both the function of genes and their evolution. For the vast majority of species, transcriptomes are still largely uncharacterized and even in those where substantial information is available it is often in the form of partially sequenced transcriptomes. With the development of next generation sequencing, a single experiment can now simultaneously identify the transcribed part of a species genome and estimate levels of gene expression. RESULTS: mRNA from actively growing needles of Norway spruce (Picea abies) was sequenced using next generation sequencing technology. In total, close to 70 million fragments with a length of 76 bp were sequenced resulting in 5 Gbp of raw data. A de novo assembly of these reads, together with publicly available expressed sequence tag (EST) data from Norway spruce, was used to create a reference transcriptome. Of the 38,419 PUTs (putative unique transcripts) longer than 150 bp in this reference assembly, 83.5% show similarity to ESTs from other spruce species and of the remaining PUTs, 3,704 show similarity to protein sequences from other plant species, leaving 4,167 PUTs with limited similarity to currently available plant proteins. By predicting coding frames and comparing not only the Norway spruce PUTs, but also PUTs from the close relatives Picea glauca and Picea sitchensis to both Pinus taeda and Taxus mairei, we obtained estimates of synonymous and non-synonymous divergence among conifer species. In addition, we detected close to 15,000 SNPs of high quality and estimated gene expression differences between samples collected under dark and light conditions. CONCLUSIONS: Our study yielded a large number of single nucleotide polymorphisms as well as estimates of gene expression on transcriptome scale. In agreement with a recent study we find that the synonymous substitution rate per year (0.6 x 10-09 and 1.1 x 10-09) is an order of magnitude smaller than values reported for angiosperm herbs. However, if one takes generation time into account, most of this difference disappears. The estimates of the dN/dS ratio (non-synonymous over synonymous divergence) reported here are in general much lower than 1 and only a few genes showed a ratio larger than 1.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Practical low-coverage genomewide sequencing of hundreds of individually barcoded samples for population and evolutionary genomics in nonmodel species

Today most population genomic studies of nonmodel organisms either sequence a subset of the genome deeply in each individual or sequence pools of unlabelled individuals. With a step-by-step workflow, we illustrate how low-coverage whole-genome sequencing of hundreds of individually barcoded samples is now a practical alternative strategy for obtaining genomewide data on a population scale. We used a highly efficient protocol to generate high-quality libraries for ~6.5 USD from each of 876 Atlantic silversides (a teleost fish with a genome size ~730 Mb) that we sequenced to 1–4× genome coverage. In the absence of a reference genome, we developed a bioinformatic pipeline for mapping the genomic reads to a de novo assembled reference transcriptome. This provides an 'in silico' method for exome capture that avoids the complexities and expenses of using wet chemistry for target isolation. Using novel tools for analysis of low-coverage data, we extracted population allele frequencies, individual genotype likelihoods and polymorphism data for 2 504 335 SNPs across the exome for the 876 fish. To illustrate the use of the resulting data, we present a preliminary analysis of geographical patterns in the exome data and a comparison of complete mitochondrial genome sequences for each individual (constructed from the low-coverage data) that show population colonization patterns along the US east coast. With a total cost per sample of less than 50 USD (including sequencing) and ability to prepare 96 libraries in only 5 h, our approach adds a viable new option to the population genomics toolbox.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Sequence capture of ultraconserved elements from bird museum specimens

New DNA sequencing technologies are allowing researchers to explore the genomes of the millions of natural history specimens collected prior to the molecular era. Yet, we know little about how well specific next-generation sequencing (NGS) techniques work with the degraded DNA typically extracted from museum specimens. Here, we use one type of NGS approach, sequence capture of ultraconserved elements (UCEs), to collect data from bird museum specimens as old as 120 years. We targeted 5060 UCE loci in 27 western scrub-jays (Aphelocoma californica) representing three evolutionary lineages that could be species, and we collected an average of 3749 UCE loci containing 4460 single nucleotide polymorphisms (SNPs). Despite older specimens producing fewer and shorter loci in general, we collected thousands of markers from even the oldest specimens. More sequencing reads per individual helped to boost the number of UCE loci we recovered from older specimens, but more sequencing was not as successful at increasing the length of loci. We detected contamination in some samples and determined that contamination was more prevalent in older samples that were subject to less sequencing. For the phylogeny generated from concatenated UCE loci, contamination led to incorrect placement of some individuals. In contrast, a species tree constructed from SNPs called within UCE loci correctly placed individuals into three monophyletic groups, perhaps because of the stricter analytical procedures used for SNP calling. This study and other recent studies on the genomics of museum specimens have profound implications for natural history collections, where millions of older specimens should now be considered genomic resources.

opencc-zeroDec 2014View details →
dryad28/100

Data from: Molecular phylogeny and SNP variation of polar bears (Ursus maritimus), brown bears (U. arctos) and black bears (U. americanus) derived from genome sequences

We assessed the relationships of polar bears (Ursus maritimus), brown bears (U. arctos), and black bears (U. americanus) with high throughput genomic sequencing data with an average coverage of 25X for each species. A total of 1.4 billion 100-bp paired-end reads was assembled using the polar bear and annotated giant panda (Ailuropoda melanoleuca) genome sequences as references. We identified 13.8 million single nucleotide polymorphisms (SNP) in the three species aligned to the polar bear genome. These data indicate that polar bears and brown bears share more SNP with each other than either does with black bears. Concatenation and coalescence-based analysis of consensus sequences of approximately one million base pairs of ultra-conserved elements (UCE) in the nuclear genome resulted in a phylogeny with black bears as the sister group to brown and polar bears, and all brown bears are in a separate clade from polar bears. Genotypes for 162 SNP loci of 336 bears from Alaska and Montana showed that the species are genetically differentiated and there is geographic population structure of brown and black bears but not polar bears.

opencc-zeroDec 2012View details →
dryad28/100

Data from: On the optimal trimming of high-throughput mRNA sequence data

The widespread and rapid adoption of high-throughput sequencing technologies has afforded researchers the opportunity to gain a deep understanding of genome level processes that underlie evolutionary change, and perhaps more importantly, the links between genotype and phenotype. In particular, researchers interested in functional biology and adaptation have used these technologies to sequence mRNA transcriptomes of specific tissues, which in turn are often compared to other tissues, or other individuals with different phenotypes. While these techniques are extremely powerful, careful attention to data quality is required. In particular, because high-throughput sequencing is more error-prone than traditional Sanger sequencing, quality trimming of sequence reads should be an important step in all data processing pipelines. While several software packages for quality trimming exist, no general guidelines for the specifics of trimming have been developed. Here, using empirically derived sequence data, I provide general recommendations regarding the optimal strength of trimming, specifically in mRNA-Seq studies. Although very aggressive quality trimming is common, this study suggests that a more gentle trimming, specifically of those nucleotides whose Phred score &lt; 2 or &lt; 5, is optimal for most studies across a wide variety of metrics.

opencc-zeroDec 2013View details →
dryad28/100

Data from: Protein structure determination using metagenome sequence data

Despite decades of work by structural biologists, there are still ~5200 protein families with unknown structure outside the range of comparative modeling. We show that Rosetta structure prediction guided by residue-residue contacts inferred from evolutionary information can accurately model proteins that belong to large families and that metagenome sequence data more than triple the number of protein families with sufficient sequences for accurate modeling. We then integrate metagenome data, contact-based structure matching, and Rosetta structure calculations to generate models for 614 protein families with currently unknown structures; 206 are membrane proteins and 137 have folds not represented in the Protein Data Bank. This approach provides the representative models for large protein families originally envisioned as the goal of the Protein Structure Initiative at a fraction of the cost.

opencc-zeroDec 2016View details →
zenodo28/100

Supplementary material 1 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245

Tables S1–S3

opencc-zeroJun 2021View details →
zenodo28/100

Figure 4 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245

Figure 4 Maximum likelihood tree for concatenated matrix of all genes. Scale bar: 0.1 units, as estimated by RAXML.

opencc-by-4.0Jun 2021View details →
zenodo28/100

Chart 1 from: Kavanaugh DH, Maddison DR, Simison WB, Schoville SD, Schmidt J, Faille A, Moore W, Pflug JM, Archambeault SL, Hoang T, Chen J-Y (2021) Phylogeny of the supertribe Nebriitae (Coleoptera, Carabidae) based on analyses of DNA sequence data. In: Spence J, Casale A, Assmann T, Liebherr JК, Penev L (Eds) Systematic Zoology and Biodiversity Science: A tribute to Terry Erwin (1940-2020). ZooKeys 1044: 41-152. https://doi.org/10.3897/zookeys.1044.62245

Chart 1 Support for or against various clades. All columns provide maximum likelihood bootstrap values for or against a particular clade, except for column "8G B," which shows the Bayesian posterior probability estimates for the eight-gene matrix. "8GML" shows the bootstrap values for the eight-gene concatenated matrix, "Nuc G" for the concatenated nuclear genes, "NPC G" for the concatenated nuclear protein-coding genes, and "Mito G" for the concatenated mitochondrial genes. The remaining eight columns provide values for the single gene analyses. All values are expressed as percentages, with positive numbers indicating support for a clade and negative numbers indicating support for a contradictory clade having the highest support. Specific contradictory clades from alternative trees are highlighted in medium grey. Cells with bootstrap values ≥ 90 are shown in black, with values between 75 and 89 in dark grey, and values from 50 to 74 in light grey. Cells in white indicate clades present in the ML tree, but with bootstrap values &lt; 50. Cells in red have bootstrap values for a contradictory clade ≥ 50. Cells in pink have bootstrap values for or against a clade &lt; 50, and the clade is not present in the ML tree. A "-" in a cell indicates that taxon sampling for that gene was not sufficient to assess monophyly of that clade. "#g" shows the number of single-gene analyses (maximum of eight) that support a clade with bootstrap values of 50 or more.

opencc-by-4.0Jun 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record