Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
709
datasets available to search
ShareScore release 0.9.0
Dataset results
709 results for “Coverage”
[Artifacts] Colosseum: Regression Test Prioritization by Delta Displacement in Test Coverage
<p>Replication package for [Colosseum: Regression Test Prioritization by Delta Displacement in Test Coverage].</p>
Predicting Prime Path Coverage Using Regression Analysis
<p>SBES - Research Track - 2020 - Presentation Video.</p>
Estimation of the relative abundance of species in artificial mixtures of insects using low-coverage shotgun metagenomics
<p>Amplicon metabarcoding is an established technique to analyse the taxonomic composition of communities of organisms using high-throughput DNA sequencing, but there are doubts about its ability to quantify the relative proportions of the species, as opposed to the species list. Here, we bypass the enrichment step and avoid the PCR-bias, by directly sequencing the extracted DNA using shotgun metagenomics. This approach is common practice in prokaryotes, but not in eukaryotes, because of the low number of sequenced genomes of eukaryotic species. We tested the metagenomics approach using insect species whose genome is already sequenced and assembled to an advanced degree. We shotgun-sequenced, at low-coverage DNA, 18 species of insects in 22 single-species and 6 mixed-species libraries and mapped the reads against 110 reference genomes of insects. We used the single-species libraries to calibrate the process of assignation of reads to species and the libraries created from species mixtures to evaluate the ability of the method to quantify the relative species abundance. Our results showed that the shotgun metagenomic method is easily able to set apart closely-related insect species, like four species of <i>Drosophila</i> included in the artificial libraries. However, to avoid the counting of rare misclassified reads in samples, it was necessary to use a rather stringent detection limit of 0.001, so species with a lower relative abundance are ignored. We also identified that approximately half the raw reads were informative for taxonomic purposes. Finally, using the mixed-species libraries, we showed that it was feasible to quantify with confidence the relative abundance of individual species in the mixtures.</p>
Data from: Finding the right coverage: The impact of coverage and sequence quality on SNP genotyping error rates
Restriction-enzyme-based sequencing methods enable the genotyping of thousands of single nucleotide polymorphism (SNP) loci in non-model organisms. However, in contrast to traditional genetic markers, genotyping error rates in SNPs derived from restriction-enzyme-based methods remain largely unknown. Here, we estimated genotyping error rates in SNPs genotyped with double digest RAD sequencing from Mendelian incompatibilities in known mother-offspring dyads of Hoffman's two-toed sloth (Choloepus hoffmanni) across a range of coverage and sequence quality criteria, for both reference-aligned and de novo-assembled datasets. Genotyping error rates were more sensitive to coverage than sequence quality and low coverage yielded high error rates, particularly in de novo-assembled datasets. For example, coverage ≥5 yielded median genotyping error rates of ≥0.03 and ≥0.11 in reference-aligned- and de novo-assembled datasets, respectively. Genotyping error rates declined to ≤0.01 in reference-aligned datasets with a coverage >30, but remained >0.04 in the de novo-assembled datasets. We observed approximately 10- and 13-fold declines in the number of loci sampled in the reference-aligned and de novo-assembled datasets when coverage was increased from >5 to >30 at quality score ≥30, respectively. Finally, we assessed the effects of genotyping coverage on a common population genetic application, parentage assignments, and showed that the proportion of incorrectly assigned maternities was relatively high at low coverage. Overall, our results suggest that the tradeoff between sample size and genotyping error rates be considered prior to building sequencing libraries, reporting genotyping error rates become standard practice, and that effects of genotyping errors on inference be evaluated in restriction-enzyme-based SNP studies.
Data from: Conservation genetics of Australasian sailfin lizards: flagship species threatened by coastal development and insufficient protected area coverage
Despite rampant coastal development throughout Southeast Asia and the Pacific, studies of conservation genetics and ecology of vulnerable, coastal species are rare. Large bodied vertebrates with highly specialized habitat requirements may be at particular risk of extinction due to habitat degradation and fragmentation, especially if these habitats are naturally patchily distributed, marginal, or otherwise geographically limited, or associated in space with high human population densities or heavy anthropogenic disturbance. Particularly telling examples of these conservation challenges are large Australasian reptiles with obligate habitat requirements for lowland, coastal and mangrove forests. Plagued by habitat destruction due to high human densities along coastlines, sprawling rural development, and rapidly developing estuarine fisheries industry, coastal forest reptiles are experiencing rapid declines. And yet studies of population biology, genetics, and habitat requirements of species depending on these environments are few. We undertook the present study in order to take a multifaceted approach to understanding a poignant conservation problem. We identify significant evolutionary units for conservation in large-bodied Sailfin Lizards (genus Hydrosaurus), model suitable habitat in the Philippines from extensive occurrence data and evaluate the efficacy of the current protected area network, and identify the source of hydrosaurs in the illegal pet trade. We determine that the extent of the species' habitat coincident with protected areas is low. Our forensic evaluation of the illegal pet trade in the Philippines determines the existence of a natural population that is at risk of systematic exploitation by traders. Together, this integrative study characterizes a conservation urgency of particular significance: the genetically distinct Sailfin lizards of the Bicol faunal region, with suitable habitat virtually unprotected, and clear evidence of heavy exploitation for illegal trade. To the best of our knowledge, our study is the first conservation genetic study to evaluate the potential effectiveness of the protected landscape coverage in the Philippines, a Megadiverse nation and Biodiversity Hotspot.
Data from: Development of genomic tools in a widespread tropical tree, Symphonia globulifera L.f.: a new low-coverage draft genome, SNP and SSR markers
Population genetic studies in tropical plants are often challenging because of limited information on taxonomy, phylogenetic relationships and distribution ranges, scarce genomic information and logistic challenges in sampling. We describe a strategy to develop robust and widely applicable genetic markers based on a modest development of genomic resources in the ancient tropical tree species Symphonia globulifera L.f. (Clusiaceae), a keystone species in African and Neotropical rainforests. We provide the first low-coverage (11X) fragmented draft genome sequenced on an individual from Cameroon, covering 1.027 Gbp or 67.5% of the estimated genome size. Annotation of 565 scaffolds (7.57 Mbp) resulted in the prediction of 1046 putative genes (231 of them containing a complete open reading frame) and 1523 exact simple sequence repeats (SSRs, microsatellites). Aligning a published transcriptome of a French Guiana population against this draft genome produced 923 high-quality single nucleotide polymorphisms. We also preselected genic SSRs in silico that were conserved and polymorphic across a wide geographical range, thus reducing marker development tests on rare DNA samples. Of 23 SSRs tested, 19 amplified and 18 were successfully genotyped in four S. globulifera populations from South America (Brazil and French Guiana) and Africa (Cameroon and São Tomé island, FST = 0.34). Most loci showed only population-specific deviations from Hardy–Weinberg proportions, pointing to local population effects (e.g. null alleles). The described genomic resources are valuable for evolutionary studies in Symphonia and for comparative studies in plants. The methods are especially interesting for widespread tropical or endangered taxa with limited DNA availability.
Data from: Probabilistic interfractional motion carbon ion radiation therapy dose distribution for prostate cancer shows rectum sparing with moderate target coverage degradation
Purpose: This observational study investigates the influence of interfractional motion on clinical target volume (CTV) coverage, planning target volume (PTV) margins, and rectum tissue sparing in carbon ion radiation therapy (CIRT). It reports dose coverage to target structures and organs at risk in the presence of interfractional motion, investigates rectal tissue sparing, and provides recommendations for further lowering the rate of toxicity. We also propose probabilistic DVH for consideration in treatment planning to represent probable dose to the clinic's patient population. Methods: At Gunma University Hospital intensity-modulated x-ray therapy (IMXT, aka IMRT) prostate cancer patients are positioned on a table which is shifted twice based on cone-beam computed tomography (CBCT) to align bones and then align prostate tissue to isocenter. These shifts thereby contain interfractional motion. 1306 such tableshifts from 85 patients were collected. Normal probability distributions were fit to the difference between bone-matching and prostate-matching CBCT-to-planning CT tableshifts (i.e. interfractional motion). Between 2011 and 2016 CIRT prostate patients were treated with PTV1 and PTV2 margins as follows: PTV1 extends the prostate contour by 10/10, 5/10, 6/6 mm in the right/left, posterior/anterior, and superior/inferior directions, respectively, and the proximal seminal vesicles contour by 5 mm superiorly and inferiorly, 3 mm right and left. PTV2 reduces PTV1 posteriorly along a straight line to exclude the rectum and reduces the superior and inferior margins by 6 mm. From those treated with these margins, 40 patients' beam data were selected to create probable interfractional motion: The previously fit normal probability distributions were randomly sampled 2000 times per patient and beams shifted to simulate this motion. These shifted dose distributions were scaled down proportionately in magnitude and summed to obtain probable blurred dose distributions. Results: Probable dose to rectum is substantially less than planned for doses higher than 10 Gy(RBE). Absolute DVH show that mean clinical target volumes are about 138-670 cm3 smaller for a given probable dose than planned doses higher than 57 Gy(RBE) after accounting for standard error. Cumulative DVH show mean CTV fraction receiving a given probable dose is less than the mean fraction receiving the corresponding planned dose for doses larger than 52 Gy(RBE), up to 19% less at 57.4 Gy(RBE). Our PTV1 margins generally cover 95% of interfractional motion but seminal vesicles and inferior prostate receive less dose than planned due to insufficient PTV2 margins. Conclusion: Assuming rigidly shifting interfractional motion around the prostate region and neglecting minor changes in soft tissue stopping power, interfractional motion resulted in underdosing or tissue sparing in all cases. Given our low rates of relapse and recurrence, it appears less curative dose is needed than previously thought or else the target may be smaller than previously thought. In-room CT may be useful to lower dose, shrink target margins, conduct PET auto-activation dose verification studies and account for interfractional motion.
Data from: Newspaper coverage of maternal health in Bangladesh, Rwanda, and South Africa: a quantitative and qualitative content analysis
Objective: To examine newspaper coverage of maternal health in three countries that have made varying progress towards Millennium Development Goal 5 (MDG 5): Bangladesh (on track), Rwanda (making progress, but not on track) and South Africa (no progress). Design: We analysed each country's leading national English-language newspaper: Bangladesh's The Daily Star, Rwanda's The New Times/The Sunday Times, and South Africa's Sunday Times/The Times. We quantified the number of maternal health articles published from 1 January 2008 to 31 March 2013. We conducted a content analysis of subset of 190 articles published from 1 October 2010 to 31 March 2013. Results: Bangladesh's The Daily Star published 579 articles related to maternal health from 1 January 2008 to 31 March 2013, compared to 342 in Rwanda's The New Times/The Sunday Times and 253 in South Africa's Sunday Times/The Times over the same time period. The Daily Star had the highest proportion of stories advocating for or raising awareness of maternal health. Most maternal health articles in The Daily Star (83%) and The New Times/The Sunday Times (69%) used a 'human-rights' or 'policy-based' frame compared to 41% of articles from Sunday Times/The Times. Conclusions: In the three countries included in this study, which are on different trajectories towards MDG 5, there were differences in the frequency, tone and content of their newspaper coverage of maternal health. However, no causal conclusions can be drawn about this association between progress on MDG 5 and the amount and type of media coverage of maternal health.
Data from: Likelihood-based inference of population history from low coverage de novo genome assemblies
Short-read sequencing technologies have in principle made it feasible to draw detailed inferences about the recent history of any organism. In practice, however, this remains challenging due to the difficulty of genome assembly in most organisms and the lack of statistical methods powerful enough to discriminate among recent, non-equilibrium histories. We address both the assembly and inference challenges. We develop a bioinformatic pipeline for generating outgroup-rooted alignments of orthologous sequence blocks from de novo low-coverage short-read data for a small number of genomes, and show how such sequence blocks can be used to fit explicit models of population divergence and admixture in a likelihood framework. To illustrate our approach, we reconstruct the Pleistocene history of an oak-feeding insect (the oak gallwasp Biorhiza pallida) which, in common with many other taxa, was restricted during Pleistocene ice ages to a longitudinal series of southern refugia spanning theWestern Palaearctic. Our analysis of sequence blocks sampled from a single genome from each of three major glacial refugia reveals support for an unexpected history dominated by recent admixture. Despite the fact that 80% of the genome is affected by admixture during the last glacial cycle, we are able to infer the deeper divergence history of these populations. These inferences are robust to variation in block length, mutation model, and the sampling location of individual genomes within refugia. This combination of de novo assembly and numerical likelihood calculation provides a powerful framework for estimating recent population history that can be applied to any organism without the need for prior genetic resources.
Data from: Comparison of 454 pyrosequencing methods for characterizing the major histocompatibility complex of nonmodel species and the advantages of ultra deep coverage
Characterization and population genetic analysis of multilocus genes, such as those found in the major histocompatibility complex (MHC) is challenging in nonmodel vertebrates. The traditional method of extensive cloning and Sanger sequencing is costly and time-intensive and indirect methods of assessment often underestimate total variation. Here, we explored the suitability of 454 pyrosequencing for characterizing multilocus genes for use in population genetic studies. We compared two sample tagging protocols and two bioinformatic procedures for 454 sequencing through characterization of a 185-bp fragment of MHC DRB exon 2 in wolverines (Gulo gulo) and further compared the results with those from cloning and Sanger sequencing. We found 10 putative DRB alleles in the 88 individuals screened with between two and four alleles per individual, suggesting amplification of a duplicated DRB gene. In addition to the putative alleles, all individuals possessed an easily identifiable pseudogene. In our system, sequence variants with a frequency below 6% in an individual sample were usually artefacts. However, we found that sample preparation and data processing procedures can greatly affect variant frequencies in addition to the complexity of the multilocus system. Therefore, we recommend determining a per-amplicon-variant frequency threshold for each unique system. The extremely deep coverage obtained in our study (approximately 5000×) coupled with the semi-quantitative nature of pyrosequencing enabled us to assign all putative alleles to the two DRB loci, which is generally not possible using traditional methods. Our method of obtaining locus-specific MHC genotypes will enhance population genetic analyses and studies on disease susceptibility in nonmodel wildlife species.
Data from: Analysis of transposable elements in the genome of Asparagus officinalis from high coverage sequence data
Asparagus officinalis is an economically and nutritionally important vegetable crop that is widely cultivated and is used as a model dioecious species to study plant sex determination and sex chromosome evolution. To improve our understanding of its genome composition, especially with respect to transposable elements (TEs), which make up the majority of the genome, we performed Illumina HiSeq2000 sequencing of both male and female asparagus genomes followed by bioinformatics analysis. We generated 17 Gb of sequence (12×coverage) and assembled them into 163,406 scaffolds with a total cumulated length of 400 Mbp, which represent about 30% of asparagus genome. Overall, TEs masked about 53% of the A. officinalis assembly. Majority of the identified TEs belonged to LTR retrotransposons, which constitute about 28% of genomic DNA, with Ty1/copia elements being more diverse and accumulated to higher copy numbers than Ty3/gypsy. Compared with LTR retrotransposons, non-LTR retrotransposons and DNA transposons were relatively rare. In addition, comparison of the abundance of the TE groups between male and female genomes showed that the overall TE composition was highly similar, with only slight differences in the abundance of several TE groups, which is consistent with the relatively recent origin of asparagus sex chromosomes. This study greatly improves our knowledge of the repetitive sequence construction of asparagus, which facilitates the identification of TEs responsible for the early evolution of plant sex chromosomes and is helpful for further studies on this dioecious plant.
Figure 5. Per-base coverage plots for the 16S in Target-enriched DNA sequencing from historical type material enables a partial revision of the Madagascar giant stream frogs (genus Mantidactylus)
Figure 5. Per-base coverage plots for the 16S fragment in four Mantidactylus type specimens from the MNHN and BMNH collections. (a) BMNH 1947.2.25.48 (paralectotype of Rana guttulata); (b) BMNH 1947.2.25.51 (paralectotype of Rana guttulata); (c) MNHN 1895.255 (syntype of M. grandidieri); (d) MNHN 1883.520 (syntype of M. grandidieri).
MS coverage HTT sequence and assessment of the data so far (2016/02/23)
<p>Open lab notebook for project: huntingtin structural studies</p> <p> </p>
Written news coverage by CNN and FOX on China's COVID-19 epidemic from January 1, 2020, to May 31, 2021.
<p><span><span> </span>The researchers applied Python software to FOX's health section, CNN's health section, with "Coronavirus + China" </span><span>、</span><span>"Covid-19 + China" as keywords, to capture 4272 and 4167 articles on CNN and FOX from January 1, 2020 to May 31, 2021, respectively.</span></p>
Supplementary data and microkinetic model for 'Key Role of CO Coverage for Chain Growth in Co-Based Fischer-Tropsch Synthesis'
<p>This repository contains:</p> <ol> <li>The DFT data (energies, frequencies and structures) of all important intermediates of the microkinetic model constructed for the publication ‘Key Role of CO Coverage for Chain Growth in Co-Based Fischer-Tropsch Synthesis’.</li> <li>Input files for the high CO coverage microkinetic model in Chemkin.</li> <li>(update 2025-06-19) Input files for the high CO coverage microkinetic model in Chemkin with CO2 activation (sim2.zip).</li> </ol> <p>DOI: <u>10.1021/acscatal.5c03024</u> and <u>10.1021/acscatal.3c04844</u></p>
Annual Change in Hard Coral Coverage (%) in Different Coastal Regions
<p>The average annual change in hard coral coverage in coastal regions: East Asian Sea, Caribbean, Pacific, South Asia, and Gulf of Oman, from 1978 to 2019.</p>
Genotype likelihoods for low-coverage whole-genome sequencing data of yellow warblers
<p>The following datasets include the required input files used to empirically test population assignment in WGSassign on Yellow Warbler data. The file "yewa.known.ind105.ds_2x.beagle.gz" includes the filtered variants of 105 Yellow Warbler individuals output as genotype likelihoods and stored in a Beagle-formatted file. The ID file, "yewa.known.ind105.reference.IDs.txt", is a tab-delimited file with 2 columns, the first being the sample ID, and the second being the known reference population. The sample order in the ID file should match that of the input beagle file. To measure the assignment accuracy of WGSassign, we used leave-one-out cross validation using the input beagle file and our ID file.</p>
Supplemental Material for "ATNwalk: A Novel Approach for Grammar-Based Coverage-Guided Fuzzing"
<p>Contains source code, original data, additional graphs, and other material to reproduce the experiments, which were described in the paper.</p> <ul> <li><strong>corpus.tar.gz</strong> contains the seed corpus for every fuzzing campaign</li> <li><strong>crashes.tar.gz</strong> contains the crashes of each fuzzing campaign, filtered according to the technique described in the paper</li> <li><strong>data.zip</strong> contains CSV files that contain AFL++ and GCOV metrics that were used to generate the plots in the paper; it also contains the additional plots of other metrics which were referenced as supplemental material</li> <li><strong>fuzzing_20221116.tar.gz</strong> is the docker image that was used to perform the fuzzing campaigns and serves as a runtime environment</li> <li><strong>home.rocky.tar.gz</strong> contains the home folder that is mounted inside the container (see README.md for details)</li> <li><strong>plots.zip</strong> contains all plots from the paper and additional ones, like branches over time, or other box plots for AFL++ covered bits</li> </ul> <p>Consult the README.md to read on how to repeat the experiments.</p>
Mitigating the Uncertainty and Imprecision of Log-Based Code Coverage Without Requiring Additional Logging Statements (Replication Package)
<p>Replication package for Mitigating the Uncertainty and Imprecision of Log-Based Code Coverage Without Requiring Additional Logging Statements</p>
Optimized atomic coordinates of CO covered Ru(0001) at high coverage
<p>Optimized atomic coordinates of CO covered Ru(0001) from VASP calculations. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.