Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.9.0
Dataset results
15 results for “sequence reliability”
Data from: Skin swabbing of amphibian larvae yields sufficient DNA for efficient sequencing and reliable microsatellite genotyping
Skin swabbing, a minimally invasive DNA sampling method recently developed on adult amphibians, was tested on larvae of fire salamanders (Salamandra salamandra). The quality and quantity of the sampled DNA was evaluated by (i) measuring DNA concentration in DNA extracts, (ii) sequencing part of the mtDNA cytochrome b gene (692 bp) and (iii) genotyping eight polymorphic nuclear microsatellite loci. The multiple-tubes approach was used for calculating allelic dropout (ADO) and false allele (FA) rates to evaluate the reliability of the genotypes. DNA extracts from tissue samples of road-killed individuals were included in the study as positive controls. Our results showed that skin swabs of fire salamander larvae can provide DNA in sufficient quantity and quality, as sequencing was successful and no allelic dropouts or false alleles were detected. This method, tested for the first time on amphibian larvae, has proven to be an efficient and reliable alternative to the controversial tail fin clipping procedure.
Data from: Skin swabbing of amphibian larvae yields sufficient DNA for efficient sequencing and reliable microsatellite genotyping
Open the record for dataset details and reuse information.
Data from: Large-scale genotyping of highly polymorphic loci by next generation sequencing: how to overcome the challenges to reliably genotype individuals?
Studying the different roles of adaptive genes is still a challenge in evolutionary ecology and requires reliable genotyping of large numbers of individuals. Next-generation sequencing (NGS) techniques enable such large-scale sequencing, but stringent data processing is required. Here, we develop an easy to use methodology to process amplicon-based NGS data and we apply this methodology to reliably genotype four major histocompatibility complex (MHC) loci belonging to MHC class I and II of Alpine marmots (Marmota marmota). Our post-processing methodology allowed us to increase the number of retained reads. The quality of genotype assignment was further assessed using three independent validation procedures. A total of 3069 high-quality MHC genotypes were obtained at four MHC loci for 863 Alpine marmots with a genotype assignment error rate estimated as 0.21%. The proposed methodology could be applied to any genetic system and any organism, except when extensive copy-number variation occurs (that is, genes with a variable number of copies in the genotype of an individual). Our results highlight the potential of amplicon-based NGS techniques combined with adequate post-processing to obtain the large-scale highly reliable genotypes needed to understand the evolution of highly polymorphic functional genes.
Data from: Development of highly reliable in silico SNP resource and genotyping assay from exome capture and sequencing: an example from black spruce (Picea mariana)
Picea mariana is a widely distributed boreal conifer across Canada and the subject of advanced breeding programs for which population genomics and genomic selection approaches are being developed. Targeted sequencing was achieved after capturing P. mariana exome with probes designed from the sequenced transcriptome of Picea glauca, a distant relative. A high capture efficiency of 75.9% was reached although spruce has a complex and large genome including gene sequences interspersed by some long introns. The results confirmed the relevance of using probes from congeneric species to perform successfully interspecific exome capture in the genus Picea. A bioinformatics pipeline was developed including stringent criteria that helped detect a set of 97 075 highly reliable in silico SNPs. These SNPs were distributed across 14 909 genes. Part of an Infinium iSelect array was used to estimate the rate of true positives by validating 4267 of the predicted in silico SNPs by genotyping trees from P. mariana populations. The true positive rate was 96.2%, for in silico SNPs compared to a genotyping success rate of 96.7% for a set 1115 P. mariana control SNPs recycled from previous genotyping arrays. These results indicate the high success rate of the genotyping array and the relevance of the selection criteria used to delineate the new P. mariana in silico SNP resource. Furthermore, in silico SNPs were generally of medium to high frequency in natural populations, thus providing high informative value for future population genomics applications.
Figure 2 from: Nilsson R, Tedersoo L, Abarenkov K, Ryberg M, Kristiansson E, Hartmann M, Schoch C, Nylander J, Bergsten J, Porter T, Jumpponen A, Vaishampayan P, Ovaskainen O, Hallenberg N, Bengtsson-Palme J, Eriksson K, Larsson K, Larsson E, Kõljalg U (2012) Five simple guidelines for establishing basic authenticity and reliability of newly generated fungal ITS sequences. MycoKeys 4: 37-63. https://doi.org/10.3897/mycokeys.4.3606
Figure 2 - A reverse complementary sequence (bottom) aligned to its nine best BLAST matches, all of which were nearly identical to the query sequence based on BLAST scores, and all of which were given in the correct orientation by their respective authors.
Figure 6 from: Nilsson R, Tedersoo L, Abarenkov K, Ryberg M, Kristiansson E, Hartmann M, Schoch C, Nylander J, Bergsten J, Porter T, Jumpponen A, Vaishampayan P, Ovaskainen O, Hallenberg N, Bengtsson-Palme J, Eriksson K, Larsson K, Larsson E, Kõljalg U (2012) Five simple guidelines for establishing basic authenticity and reliability of newly generated fungal ITS sequences. MycoKeys 4: 37-63. https://doi.org/10.3897/mycokeys.4.3606
Figure 6 - An assembly chimera. The black dashed lines indicate breaks in the BLAST alignment and should always be taken to mean that manual examination is needed.
Figure 1 from: Nilsson R, Tedersoo L, Abarenkov K, Ryberg M, Kristiansson E, Hartmann M, Schoch C, Nylander J, Bergsten J, Porter T, Jumpponen A, Vaishampayan P, Ovaskainen O, Hallenberg N, Bengtsson-Palme J, Eriksson K, Larsson K, Larsson E, Kõljalg U (2012) Five simple guidelines for establishing basic authenticity and reliability of newly generated fungal ITS sequences. MycoKeys 4: 37-63. https://doi.org/10.3897/mycokeys.4.3606
Figure 1 - An ITS alignment featuring five random species each of the fungal phyla Ascomycota, Basidiomycota, Glomeromycota, Chytridiomycota, and Zygomycota s.l. The left half of the screen represents the ITS1 and the right half the 5.8S. Whereas the ITS1 alignment appears more or less chaotic, the 5.8S stands out as a very conserved element throughout these five phyla. The 5.8S starts at position 803 (indicated by the black cursor in the uppermost sequence). Seaview (Gouy et al. 2010) was used to display the alignment.
Figure 5 from: Nilsson R, Tedersoo L, Abarenkov K, Ryberg M, Kristiansson E, Hartmann M, Schoch C, Nylander J, Bergsten J, Porter T, Jumpponen A, Vaishampayan P, Ovaskainen O, Hallenberg N, Bengtsson-Palme J, Eriksson K, Larsson K, Larsson E, Kõljalg U (2012) Five simple guidelines for establishing basic authenticity and reliability of newly generated fungal ITS sequences. MycoKeys 4: 37-63. https://doi.org/10.3897/mycokeys.4.3606
Figure 5 - a Graphical overview of the BLAST results of a regular sequence b BLAST results of a chimeric sequence where the ITS1 comes from another species, such that the ITS1 is not involved in the alignment featuring the 5.8S+ITS2 (hence the lack of a match for the first ca. 180 bp.). Obviously, a severely compromised sequence that is already in INSD will always find a perfect match through BLAST in INSD: itself. In that case, the presence of a 100% similar reference sequence cannot be used as a testimony to the authenticity of the query sequence.
Figure 8 from: Nilsson R, Tedersoo L, Abarenkov K, Ryberg M, Kristiansson E, Hartmann M, Schoch C, Nylander J, Bergsten J, Porter T, Jumpponen A, Vaishampayan P, Ovaskainen O, Hallenberg N, Bengtsson-Palme J, Eriksson K, Larsson K, Larsson E, Kõljalg U (2012) Five simple guidelines for establishing basic authenticity and reliability of newly generated fungal ITS sequences. MycoKeys 4: 37-63. https://doi.org/10.3897/mycokeys.4.3606
Figure 8 - Untrimmed sequences tend to look like this when run through BLAST. Note how the first ca. 20 bp., and the last ca. 30 bp., of the query sequence (represented by the red bar with scale marks every 100 bp.) do not align to any of the BLAST hits. The use of different but closely situated primers may give a similar pattern, however, pointing at the need to also look at the BLAST alignments for start and end positions of the reference sequences.
Figure 4 from: Nilsson R, Tedersoo L, Abarenkov K, Ryberg M, Kristiansson E, Hartmann M, Schoch C, Nylander J, Bergsten J, Porter T, Jumpponen A, Vaishampayan P, Ovaskainen O, Hallenberg N, Bengtsson-Palme J, Eriksson K, Larsson K, Larsson E, Kõljalg U (2012) Five simple guidelines for establishing basic authenticity and reliability of newly generated fungal ITS sequences. MycoKeys 4: 37-63. https://doi.org/10.3897/mycokeys.4.3606
Figure 4 - A multiple alignment where the topmost sequence is chimeric and the remaining sequences represent its best BLAST matches. The alignment is fine in ITS1 and 5.8S (a; the 5.8S starts at position 479), but the alignment in ITS2 (b; position 637 and on) falls far short of scientific rigour. Alignments like these bespeak chimeric unions.
Figure 3 from: Nilsson R, Tedersoo L, Abarenkov K, Ryberg M, Kristiansson E, Hartmann M, Schoch C, Nylander J, Bergsten J, Porter T, Jumpponen A, Vaishampayan P, Ovaskainen O, Hallenberg N, Bengtsson-Palme J, Eriksson K, Larsson K, Larsson E, Kõljalg U (2012) Five simple guidelines for establishing basic authenticity and reliability of newly generated fungal ITS sequences. MycoKeys 4: 37-63. https://doi.org/10.3897/mycokeys.4.3606
Figure 3 - "Strand=Plus/Minus" indicates that the query and reference sequence come in opposing read directions. Another hint comes from the observation that the alignment starts at the first base (1) in the query sequence and progresses upwards to base 60 in the first alignment line; however, for the reference sequence, the alignment starts at base 635 and progresses downwards to base 578.
Figure 7 from: Nilsson R, Tedersoo L, Abarenkov K, Ryberg M, Kristiansson E, Hartmann M, Schoch C, Nylander J, Bergsten J, Porter T, Jumpponen A, Vaishampayan P, Ovaskainen O, Hallenberg N, Bengtsson-Palme J, Eriksson K, Larsson K, Larsson E, Kõljalg U (2012) Five simple guidelines for establishing basic authenticity and reliability of newly generated fungal ITS sequences. MycoKeys 4: 37-63. https://doi.org/10.3897/mycokeys.4.3606
Figure 7 - An assembly chimera. An extraneous sequence segment was assembled into a position where it should not have been, such as in the middle of the 5.8S. The white area in the reference sequences indicates the absence of sequence data for this particular part of the query sequence. Manual examination is always needed in cases like this.
Data from: Large-scale genotyping of highly polymorphic loci by next generation sequencing: how to overcome the challenges to reliably genotype individuals?
Open the record for dataset details and reuse information.
Data from: Development of highly reliable in silico SNP resource and genotyping assay from exome capture and sequencing: an example from black spruce (Picea mariana)
Open the record for dataset details and reuse information.
CRP-seq: a reliable and fast method for sequencing RNA G-quadruplexes transcriptome-wide
GEO Series GSE260479. Homo sapiens. 18 samples. Type: Expression profiling by high throughput sequencing.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.