Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

140

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

140 results for “Illumina Sequencing”

Learn how ShareScore rates datasets ↗
zenodo32/100

Sequence and assembly artifacts: GRCr8 and Illumina reads from BN

<p>GRCr8 rat reference genome assembly polishing data.</p>

opencc-by-4.0Jul 2024View details →
dryad32/100

SARA module 3: NGS epitope sequencing: Illumina FASTQ files

<p>We investigate the accumulated microbial and autoantigen antibody repertoire in adult-onset dermatomyositis patients sero-positive for TIF1γ (TRIM33) autoantibodies. We use an untargeted high-throughput approach which combines immunoglobulin disease-specific epitope-enrichment and identification of microbial and human antigens. We observe antibodies recognizing a wider repertoire of microbial antigens in dermatomyositis. Antibodies recognizing viruses and Poxviridae family species are significantly enriched. The identified autoantibodies recognise a large portion of the human proteome, including interferon regulated proteins; these proteins cluster in specific biological processes. In addition to TRIM33, we identify autoantibodies against eleven further TRIM proteins, including TRIM21. Some of these TRIM proteins share epitope homology with specific viral species including poxviruses. Our data suggest antibody accumulation in dermatomyositis against an expanded diversity of microbial and human proteins and evidence of non-random targeting of specific signalling pathways. Our findings indicate that molecular mimicry and epitope spreading events may play a role in dermatomyositis pathogenesis.</p>

opencc-zeroJan 2022View details →
zenodo32/100

16S Illumina Sequencing of Aguamiel Microbiota

<p>Dataset of the 16S gene V4 region Illumina Sequencing of&nbsp;<em>Agave salmiana&nbsp;</em>sieve&#39;s (aguamiel) microbial community.</p>

opencc-by-4.0Nov 2022View details →
dryad32/100

Data from: A cost-efficient and simple protocol to enrich prey DNA from extractions of predatory arthropods for large-scale gut content analysis by Illumina sequencing

Open the record for dataset details and reuse information.

publicOct 2016View details →
dryad32/100

Data from: Development and characterization of thirty-three microsatellite markers for the Patagonian sprat, Sprattus fuegensis (Jenyns, 1842), using paired-end Illumina shotgun sequencing

Open the record for dataset details and reuse information.

publicApr 2014View details →
dryad32/100

Data from: Evaluating metabarcoding to analyse diet composition of species foraging in anthropogenic landscapes using Ion Torrent and Illumina sequencing

Open the record for dataset details and reuse information.

publicSep 2019View details →
dryad32/100

LNCRNA Illumina sequencing results associated with renal fibrosis disease progression

Open the record for dataset details and reuse information.

publicJun 2024View details →
dryad32/100

SARA module 3: NGS epitope sequencing: Illumina FASTQ files

Open the record for dataset details and reuse information.

publicSep 2021View details →
dryad32/100

Data from: Massively parallel multiplex DNA sequencing for specimen identification using an Illumina MiSeq platform

Open the record for dataset details and reuse information.

publicFeb 2016View details →
dryad32/100

Valenzuela phylogenomic dataset from: Illumina whole genome sequencing indicates ploidy level differences within the Valenzuela flavidus (Psocodea: Psocomorpha: Caeciliusidae) species complex

Open the record for dataset details and reuse information.

publicNov 2021View details →
dryad28/100

Data from: Adapterama II: universal amplicon sequencing on Illumina platforms (TaggiMatrix)

Next-generation sequencing (NGS) of amplicons is used in a wide variety of contexts. In many cases, NGS amplicon sequencing remains overly expensive and inflexible, with library preparation strategies relying upon the fusion of locus-specific primers to full-length adapter sequences with a single identifying sequence or ligating adapters onto PCR products. In Adapterama I, we presented universal stubs and primers to produce thousands of unique index combinations and a modifiable system for incorporating them into Illumina libraries. Here, we describe multiple ways to use the Adapterama system and other approaches for amplicon sequencing on Illumina instruments. In the variant we use most frequently for large-scale projects, we fuse partial adapter sequences (TruSeq or Nextera) onto the 5' end of locus-specific PCR primers with variable-length tag sequences between the adapter and locus-specific sequences. These fusion primers can be used combinatorially to amplify samples within a 96-well plate (eight forward primers + 12 reverse primers yield 8 x 12 = 96 combinations), and the resulting amplicons can be pooled. The initial PCR products then serve as template for a second round of PCR with dual-indexed iTru or iNext primers (also used combinatorially) to make full-length libraries. The resulting quadruple-indexed amplicons have diversity at most base positions and can be pooled with any standard Illumina library for sequencing. The number of sequencing reads from the amplicon pools can be adjusted, facilitating deep sequencing when required or reducing sequencing costs per sample to an economically trivial amount when deep coverage is not needed. We demonstrate the utility and versatility of our approaches with results from six projects using different implementations of our protocols. Thus, we show that these methods facilitate amplicon library construction for Illumina instruments at reduced cost with increased flexibility. A simple web page to design fusion primers compatible with iTru primers is available at: http://baddna.uga.edu/tools-taggi.html. A fast and easy to use program to demultiplex amplicon pools with internal indexes is available at: https://github.com/lefeverde/Mr_Demuxy.

opencc-zeroSep 2020View details →
dryad28/100

Data from: Using Illumina Next Generation Sequencing Technologies to sequence multigene families in de novo species

The advent of Next Generation Sequencing Technology (NGST) has revolutionized molecular biology research, allowing for rapid gene/genome sequencing from a multitude of diverse species. As high throughput sequencing becomes more accessible, more efficient workflows must be developed to deal with the amounts of data produced and better assemble the genomes of de novo lineages. We combine traditional laboratory methods with Illumina NGST to amplify and sequence the largest mammalian multigene family, the Olfactory Receptor gene family, for species with and without a reference genome. We develop novel assembly methods to annotate and filter these data, which can be utilized for any gene family or any species. We find no significant difference between the ratio of genes within their respective gene families of our data compared with available genomic data. Using simulated data we explore the limitations of short-read sequence data and our assembly in recovering this gene family. We highlight the benefits and shortcomings of these methods. Compared with data generated from traditional polymerase chain reaction, cloning and Sanger sequencing methodologies, sequence data generated using our pipeline increases yield and sequencing efficiency without reducing the number of unique genes amplified. A cloning step is not required, therefore shortening data generation time. The novel downstream methodologies and workflows described provide a tool to be utilized by many fields of biology, to access and analyze the vast quantities of data generated. By combining laboratory and in silico methods, we provide a means of extracting genomic information for multigene families without complete genome sequencing.

opencc-zeroDec 2012View details →
dryad28/100

Data from: Characterization of 42 polymorphic microsatellite loci in Mimulus ringens (Phrymaceae) using Illumina sequencing

Premise of the study: Microsatellite markers were isolated and characterized in Mimulus ringens (Phrymaceae), a herbaceous wetland perennial, to facilitate studies of mating patterns and population genetic structure. Methods and Results: A total of 42 polymorphic loci were identified from a sample of 24 individuals from a single popula- tion in Ohio, USA. The number of alleles per locus ranged from two to nine, and median observed heterozygosity was 0.435. Conclusions: This large number of polymorphic loci will enable researchers to quantify male fitness, patterns of multiple pa- ternity, selfing, and biparental inbreeding in large natural populations of this species. These markers will also permit detailed study of fine-scale patterns of genetic structure.

opencc-zeroDec 2011View details →
dryad28/100

Data from: Successful recovery of nuclear protein-coding genes from small insects in museums using illumina sequencing

In this paper we explore high-throughput Illumina sequencing of nuclear protein-coding, ribosomal, and mitochondrial genes in small, dried insects stored in natural history collections. We sequenced one tenebrionid beetle and 12 carabid beetles ranging in size from 3.7 to 9.7 mm in length that have been stored in various museums for 4 to 84 years. Although we chose a number of old, small specimens for which we expected low sequence recovery, we successfully recovered at least some low-copy nuclear protein-coding genes from all specimens. For example, in one 56-year-old beetle, 4.4 mm in length, our de novo assembly recovered about 63% of approximately 41,900 nucleotides in a target suite of 67 nuclear protein-coding gene fragments, and 70% using a reference-based assembly. Even in the least successfully sequenced carabid specimen, reference-based assembly yielded fragments that were at least 50% of the target length for 34 of 67 nuclear protein-coding gene fragments. Exploration of alternative references for reference-based assembly revealed few signs of bias created by the reference. For all specimens we recovered almost complete copies of ribosomal and mitochondrial genes. We verified the general accuracy of the sequences through comparisons with sequences obtained from PCR and Sanger sequencing, including of conspecific, fresh specimens, and through phylogenetic analysis that tested the placement of sequences in predicted regions. A few possible inaccuracies in the sequences were detected, but these rarely affected the phylogenetic placement of the samples. Although our sample sizes are low, an exploratory regression study suggests that the dominant factor in predicting success at recovering nuclear protein-coding genes is a high number of Illumina reads, with success at PCR of COI and killing by immersion in ethanol being secondary factors; in analyses of only high-read samples, the primary significant explanatory variable was body length, with small beetles being more successfully sequenced.

opencc-zeroDec 2015View details →
dryad28/100

Data from: Characterization of microsatellite loci for the Gulf Coast waterdog (Necturus beyeri) using paired-end Illumina shotgun sequencing and cross-amplification in other Necturus

[No abstract filled]

opencc-zeroDec 2017View details →
dryad28/100

Data from: Assessment of a 16S rRNA amplicon Illumina sequencing procedure for studying the microbiome of a symbiont-rich aphid genus

The bacterial communities inhabiting arthropods are generally dominated by a few endosymbionts that play an important role in the ecology of their hosts. Rather than comparing bacterial species richness across samples, ecological studies on arthropod endosymbionts often seek to identify the main bacterial strains associated with each specimen studied. The filtering out of contaminants from the results and the accurate taxonomic assignment of sequences are therefore crucial in arthropod microbiome studies. We aimed here to validate an Illumina 16S rRNA gene sequencing protocol and analytical pipeline for investigating endosymbiotic bacteria associated with aphids. Using replicate DNA samples from 12 species (Aphididae: Lachninae, Cinara) and several controls, we removed individual sequences not meeting a minimum threshold number of reads in each sample and carried out taxonomic assignment for the remaining sequences. With this approach, we show that: i) contaminants accounted for a negligible proportion of the bacteria identified in our samples; ii) the taxonomic composition of our samples and the relative abundance of reads assigned to a taxon were very similar across PCR and DNA replicates for each aphid sample; in particular, bacterial DNA concentration had no impact on the results. Furthermore, by analysing the distribution of unique sequences across samples rather than aggregating them into operational taxonomic units (OTUs), we gained insight into the specificity of endosymbionts for their hosts. Our results confirm that Serratia symbiotica is often present in Cinara species, in addition to the primary symbiont, Buchnera aphidicola. Furthermore, our findings reveal new symbiotic associations with Erwinia and Sodalis-related bacteria. We conclude with suggestions for generating and analysing 16S rRNA gene sequences for arthropod endosymbiont studies.

opencc-zeroDec 2014View details →
zenodo28/100

Supporting data for Comparison of Oxford Nanopore and Illumina sequencing for SARS-CoV-2 variant monitoring in wastewater

<p>This dataset contains the FASTQ files used in the paper&nbsp;<em>Comparison of Oxford Nanopore and Illumina sequencing for SARS-CoV-2 variant monitoring in wastewater</em>.&nbsp;</p> <p>The FASTQ files are classified in 2 different folders: MinION and MiSeq. Inside the MinION folder, there are two more folders, one for the R9.4.1 flow cell data and the other for the R10.4.1 flow cell data.&nbsp;</p> <p>Twist synthetic RNA mixtures are named as mix1, mix2, etc.&nbsp;</p> <p>Twist synthetic RNA corresponding to the Wuhan sequence is named as mix11_WH</p> <p>Wastewater samples are named as WWTP_1, WWTP_2, etc.&nbsp;&nbsp;</p>

openApr 2023View details →
dryad28/100

Data from: Using Illumina Next Generation Sequencing Technologies to sequence multigene families in de novo species

Open the record for dataset details and reuse information.

publicFeb 2013View details →
dryad28/100

Data from: Characterization of 42 polymorphic microsatellite loci in Mimulus ringens (Phrymaceae) using Illumina sequencing

Open the record for dataset details and reuse information.

publicDec 2012View details →
dryad28/100

Data from: Characterization of microsatellite loci for the Gulf Coast waterdog (Necturus beyeri) using paired-end Illumina shotgun sequencing and cross-amplification in other Necturus

Open the record for dataset details and reuse information.

publicAug 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record