Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
709
datasets available to search
ShareScore release 0.9.0
Dataset results
709 results for “Coverage”
Average fold coverage and annotation results of those ORF that were increasing over treatment time
<p>Table listing those ORF that presented an increased frequency across the treatment time</p> <p>This a supplementary material for the doctoral thesis entitle: <em><strong>Understanding microbiome supra-metabolism responses under strong selective pressures as a resource for designing synthetic gene arrangements encoding key co-selected functions for environmental biotechnology applications. </strong></em>Villegas-Plazas M, 2020. Universidad del Valle. Cali, Colombia</p>
Data from: Low coverage genomic data resolve the population divergence and gene flow history of an Australian rain forest fig wasp
Population divergence and gene flow are key processes in evolution and ecology. Model-based analysis of genome-wide datasets allows discrimination between alternative scenarios for these processes even in non-model taxa. We used two complementary approaches (one based on the blockwise site frequency spectrum (bSFS), the second on the Pairwise Sequentially Markovian Coalescent (PSMC)) to infer the divergence history of a fig wasp, Pleistodontes nigriventris. Pleistodontes nigriventris and its fig tree mutualist Ficus watkinsiana are restricted to rain forest patches along the eastern coast of Australia, and are separated into northern and southern populations by two dry forest corridors (the Burdekin and St. Lawrence Gaps). We generated whole genome sequence data for two haploid males per population and used the bSFS approach to infer the timing of divergence between northern and southern populations of P. nigriventris, and to discriminate between alternative isolation with migration (IM) and instantaneous admixture (ADM) models of post divergence gene flow. Pleistodontes nigriventris has low genetic diversity (π = 0.0008), to our knowledge one of the lowest estimates reported for a sexually reproducing arthropod. We find strongest support for an ADM model in which the two populations diverged ca. 196kya in the late Pleistocene, with almost 25% of northern lineages introduced from the south during an admixture event ca. 57kya. This divergence history is highly concordant with individual population demographies inferred from each pair of haploid males using PSMC. Our analysis illustrates the inferences possible with genome-level data for small population samples of tiny, non-model organisms and adds to a growing body of knowledge on the population structure of Australian rain forest taxa.
Dataset for the submission entitled: Journal article publishing in the social sciences and humanities: a comparison of Web of Science coverage for five European countries
<p>Dataset for the manuscript submission entitled: Journal article publishing in the social sciences and humanities: a comparison of Web of Science coverage for five European countries.</p>
Data from: Age-appropriate vaccination coverage and its associated factors for pentavalent 1-3 and measles vaccine doses, in northeast Ethiopia: A community-based cross-sectional study
Background: In Ethiopia, there are no studies on age-appropriate vaccinations that children received at the recommended specific ages. Therefore, we assessed age appropriate vaccination coverage and its associated factors among children 12 to 23 months of age in Menz Lalo district, northeast Ethiopia. Methods: A community based cross sectional study was conducted in Menz Lalo district from March to April, 2018 among 417 mothers/caregivers with children 12 to 23 months of age using simple random sampling technique. Data were collected using pretested structured Amharic questionnaire. Age appropriate vaccination coverage was measured using World Health Organization vaccination schedule recommendation. Information about children vaccination status was collected from children vaccination cards. Data was entered into Epi-Info7 software and exported to SPSS-20 for analysis. Logistic regression analysis was carried out to identify factors associated with age inappropriate vaccinations. A p-value of < 0.05 was considered to sate statistically significant association. Results: Age appropriate vaccination coverage were 39.1% (95% CI: 34.1-43.6) for Pentavalent 1, 36.3% (95% CI: 30.5-40) for Pentavalent 2, 30.3% (95% CI: 23.5-32.4) for Pentavalent 3 and 26.4% (95% CI: 18-29.4) for measles vaccine doses. Age inappropriate Pentavalent 1-3 vaccinations was associated with being male child (AOR: 0.47, 95% CI: 0.29-0.74), absence of telephone (AOR: 2.2, 95% CI: 1.4-3.6), absence of usual caretaker (AOR: 2.6, 95% CI: 1.3-5.2), unplanned pregnancy (AOR: 1.9, 95% CI: 1.1-3.5), missing antenatal conference participation (AOR: 2.7, 95% CI: 1.3-5.7), first birth order (AOR: 0.34, 95% CI: 0.17-0.68) and insufficient knowledge (AOR: 2.7, 95% CI: 1.6-4.4). Conclusion: The proportion of age appropriate vaccinations coverage was low in the study area. Modifiable factors were associated with age inappropriate vaccinations. Vaccination interventions should consider identified modifiable factors to improve age appropriate vaccinations coverage.
Newspaper review and analysis data on media coverage of budget issues in Nigeria
<p>The data was collected from analysis of six newspaper publications to establish citizen engagement and media coverage of the budget discourse from 2009 to 2013 as part of the investigation of the use of the online national budget of Nigeria. The work was part of the 'Exploring the Emerging Impacts of Open Data in Developing Countries' (ODDC) research project.</p>
coverage and density of eeg source localization
<p>The file contains the lead field matrices of atlas and subject specific head models and coordinate files of HydroCel EEG nets. </p>
Running a Red Light: An Investigation into Why Software Engineers (Occasionally) Ignore Coverage Checks --- Appendix
<p>This online appendix contains the anonymised responses to our survey, which we have analysed in our report. In addition, it also includes the mapping of axial coding codes to larger groups, and (if not self-evident) a further explanation of these groups.</p>
Dataset for Aberration-Corrected STEM to Determine the Surface Coverage and Distribution of Immobilized Molecular Complexes
<p>Raw image data used for the paper "Aberration-Corrected STEM to Determine the Surface Coverage and Distribution of Immobilized Molecular Complexes".</p>
UAV Spraying Parameters-Coverage in Vineyards
<p>A set of UAV spraying data collected using WSPs, across different parameter configurations. All tests were conducted in an experimental vineyard, under field conditions. All WSP samples were analysed using the DepositScan software developed by USDA (doi: 10.1016/j.compag.2011.01.003), and all coverage and VMD values are estimated by this software.</p>
Full-coverage near surface CO2 over the global continent from 2015 to 2021 generated by Deep Forest
<p>These are full-coverage near surface CO2 distribution maps over the global continent from 2015 to 2021. We fused CO<sub>2</sub> in situ measurements with multiple variables including OCO-2 XCO<sub>2</sub> retrieval and auxiliary data such as meteorological factors, vegetation parameters, and human activities by MA-DF-LGB (Mix Attention-Deep Forest-LightGBM) to generate the dataset.</p>
LYCEUM: Learning to call copy number variants on low coverage ancient genomes
<p>Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduce additional noise into sequencing data; and finally, (iii) the typically low coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high coverage read-depth signals, underperform under such conditions. To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high- confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection.</p>
Outdoor NB-IoT and 5G coverage and channel information data in urban environments
<p>This dataset includes data for NB-IoT and 5G networks as collected in two cities: Oslo, Norway (NB-IoT only) and Rome, Italy (both NB-IoT and 5G).</p> <p>Data were collected using the Rohde & Schwarz TSMA6 mobile network scanner. 7 measurement campaigns are provided for Oslo, and 6 for Rome. Additional data collected in Rome are provided in the following large-scale dataset, focusing on the two major mobile network operators: <a href="https://ieee-dataport.org/documents/large-scale-dataset-4g-nb-iot-and-5g-non-standalone-network-measurements">https://ieee-dataport.org/documents/large-scale-dataset-4g-nb-iot-and-5g-non-standalone-network-measurements</a> </p> <p>The dataset includes a metadata file providing the following information for each campaign: </p> <ul> <li>date of collection;</li> <li>start time and end time of collection;</li> <li>length;</li> <li>type (walking/driving).</li> </ul> <p>Two additional metadata files are provided: two .kml files, one for each city, allowing the import of coordinates of data points organized by campaign in a GIS engine, such as Google Earth, for interactive visualization.</p> <p>The dataset contains the following data for NB-IoT:</p> <ul> <li>Raw data for each campaign, stored in two .csv files. For a generic campaign <X>, the files are: <ul> <li>NB-IoT_coverage_C<X>.csv including a geo-tagged data entry in each row. Each entry provides information on a Narrowband Physical Cell Identifier (NPCI), with data related to the time stamp the NPCI was detected, GPS information, network (NPCI, Operator, Country Code, eNodeB-ID) and RF signal (RSSI, SINR, RSRP and RSRQ values);</li> <li> NB-IoT_RefSig_cir_C<X>.csv, also including a geo-tagged data entry in each row. Each entry provides information on a NPCI, with data related to the time stamp the NPCI was detected, GPS information, network (NPCI, Operator ID, Country Code, eNodeB-ID) and Channel Impulse Response (CIR) statistics, including the maximum delay.</li> </ul> </li> <li>Processed data, stored in a Matlab workspace (.mat) file for each city: data are grouped in data points, identified by <Latitude, longitude> pairs. Each data point provides RF and CIR maximum delay measurements for each <NPCI, Operator ID, eNodeB-ID> unique combination detected at the coordinates of the data point.</li> <li>Estimated positions of eNodeBs, stored in a csv file for each city;</li> <li>A matlab script and a function to extract and generate processed data from the raw data for each city.</li> </ul> <p>The dataset contains the following data for 5G:</p> <ul> <li>Raw data for each campaign, stored in two .xslx files. For a generic campaign <X>, the files are: <ul> <li>5G_coverage_C<X>.xslx including a geo-tagged data entry in each row. Each entry provides information on a Physical Cell Identifier (PCI), with data related to the time stamp the PCI was detected, GPS information, network (PCI, Beamforming Index, Operator, Country Code) and RF data (SSB-RSSI, SSS-SINR, SSS-RSRP and SSS-RSRQ values, and similar information for the PBCH signal);</li> <li> 5G_RefSig_cir_C<X>.csv, also including a geo-tagged data entry in each row. Each entry provides information on a PCI, with data related to the time stamp the PCI was detected, GPS information, network (PCI, Beamforming Index, Operator ID, Country Code) and Channel Impulse Response (CIR) statistics, including the maximum delay.</li> </ul> </li> <li>Processed data, stored in a Matlab workspace (.mat) file: data are grouped in data points, identified by <Latitude, longitude> pairs. Each data point provides RF and CIR maximum delay measurements for each <PCI, Beamforming Index, Operator ID> unique combination detected at the coordinates of the data point.</li> <li>A matlab script and a supporting function to extract and generate processed data from the raw data.</li> </ul> <p>In addition, in the case of the Rome data additional matlab workspaces are provided, containing interpolated data in the feature dimensions according to two different approaches:</p> <ul> <li>A campaign-by-campaign linear interpolation (both NB-IoT and 5G);</li> <li>A bidimensional interpolation on all campaigns combined (NB-IoT only).</li> </ul> <p>A function to interpolate missing data in the original data according to the first approach is also provided for each technology. The interpolation rationale and procedure for the first approach is detailed in:</p> <p>L. De Nardis, G. Caso, Ö. Alay, U. Ali, M. Neri, A. Brunstrom and M.-G. Di Benedetto, "Positioning by Multicell Fingerprinting in Urban NB-IoT networks," Sensors, Volume 23, Issue 9, Article ID 4266, April 2023. <span>DOI: </span><a href="https://doi.org/10.3390/s23094266" target="_blank" rel="noopener"><span>10.3390/s23094266</span></a>.</p> <p>The second interpolation approach is instead introduced and described in:</p> <p>L. De Nardis, M. Savelli, G. Caso, F. Ferretti, L. Tonelli, N. Bouzar, A. Brunstrom, O. Alay, M. Neri, F. Elbahhar and M.-G. Di Benedetto, " Range-free Positioning in NB-IoT Networks by Machine Learning: beyond WkNN", under major revision in IEEE Journal of Indoor and Seamless Positioning and Navigation.</p> <p>Positioning using the 5G data was furthermore in investigated in: </p> <p>K. Kousias, M. Rajiullah, G. Caso, U. Ali, Ö. Alay, A. Brunstrom, L. De Nardis, M. Neri, and M.-G. Di Benedetto, "A Large-Scale Dataset of 4G, NB-IoT, and 5G Non-Standalone Network Measurements," <span>IEEE Communications Magazine, Volume 62, Issue 5, pp</span><span>. 44-49, May</span><span> 202</span><span>4</span><span>. DOI: </span><a href="https://doi.org/10.1109/MCOM.011.2200707" target="_blank" rel="noopener"><span>10.1109/MCOM.011.2200707</span></a><span>.</span></p> <p><span>G. Caso, M. Rajiullah, K. Kousias, U. Ali, N. Bouzar, L. De Nardis, A. Brunstrom, Ö. Alay, M. Neri and M.-G. Di Benedetto,"The Chronicles of 5G Non-Standalone: An Empirical Analysis of Performance and Service Evolution", IEEE Open Journal of the Communications Society, Volume 5, pp. 7380 - 7399, 2024. DOI: <a href="https://doi.org/10.1109/OJCOMS.2024.3499370" target="_blank" rel="noopener"><span>10.1109/OJCOMS.2024.3499370</span></a>.</span></p> <p>Please refer to the above publications when using and citing the dataset. </p>
Full-coverage daily 0.1° XCO2 in China
<p>This dataset is generated by a hybrid deep learning model.</p> <p> </p>
Excluding Code from Test Coverage (dataset)
<p>This is the dataset for the paper "What Code Is Deliberately Excluded from Test Coverage and Why?", IEEE/ACM International Conference on Mining Software Repositories (MSR), 2021.</p> <p>This is the dataset for the paper "Excluding Code from Test Coverage: Practices, Motivations, and Impact", Empirical Software Engineering, 2021.</p>
VaMoS 2022 Submission 12 - Continuous T-Wise Sampling: Increasing Coverage over Time
<p>Evaluation results for the VaMoS 2022 submission 12. For more details regarding the dataset please refer to the included readme.md file.</p>
High-resolution and full coverage AOD downscaling based on the bagging model over the arid and semi-arid areas, NW China
<p>High-resolution and full coverage 250 m monthly AOD product over the arid and semi-arid areas, NW China. the scale factor is 1000.</p>
Data for manuscript "Reciprocal Radicalization: The Rise of Culture War Terminology in British and American News Coverage"
<p>This data set contains frequency counts of target words in 16 million news and opinion articles from 10 popular news media outlets in the United Kingdom: The Guardian, The Times, The Independent, The Daily Mirror, BBC, Financial Times, Metro, Telegraph, The and The Daily Mail plus a few additional American-based outlets used for comparison reference. The target words are listed in the associated manuscript and are mostly words that denote some type of prejudice, social justice related terms or counterreaction to it. A few additional words are also available since they are used in the manuscript for illustration purposes.</p> <p>The textual content of news and opinion articles from the outlets listed in Figure 3 of the main manuscript is available in the outlet's online domains and/or public cache repositories such as Google cache (https://webcache.googleusercontent.com), The Internet Wayback Machine (https://archive.org/web/web.php), and Common Crawl (https://commoncrawl.org). We derived relative frequency counts from these sources. Textual content included in our analysis is circumscribed to articles headlines and main body of text of the articles and does not include other article elements such as figure captions.</p> <p>Targeted textual content was located in HTML raw data using outlet specific xpath expressions. Tokens were lowercased prior to estimating frequency counts. To prevent outlets with sparse text content for a year from distorting aggregate frequency counts, we only include outlet frequency counts from years for which there is at least 1 million words of article content from an outlet. </p> <p>Yearly frequency usage of a target word in an outlet in any given year was estimated by dividing the total number of occurrences of the target word in all articles of a given year by the number of all words in all articles of that year. This method of estimating frequency accounts for variable volume of total article output over time.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript </p> <p>-targetWordsInArticlesCounts.rar contains counts of target words in outlets articles as well as total counts of words in articles</p> <p>-targetWordsInArticlesCountsGuardianExampleWords contains counts of target words in outlets articles as well as total counts of words in articles for illustrative Figure 1 in main manuscript</p> <p>Usage Notes</p> <p>In a small percentage of articles, outlet specific XPath expressions can fail to properly capture the content of the article due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. As a result, the total and target word counts metrics for a small subset of articles are not precise. In a random sample of articles and outlets, manual estimation of target words counts overlapped with the automatically derived counts for over 90% of the articles.</p> <p>Most of the incorrect frequency counts were minor deviations from the actual counts such as for instance counting the word "Facebook" in an article footnote encouraging article readers to follow the journalist’s Facebook profile and that the XPath expression mistakenly included as the content of the article main text. To conclude, in a data analysis of 16 million articles, we cannot manually check the correctness of frequency counts for every single article and hundred percent accuracy at capturing articles’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our frequency metrics are representative of word prevalence in print news media content (see Figure 1 of main manuscript for supporting evidence).</p>
Blood Exposome Database Compounds with DDA spectra coverage
<p>The last two columns in the CSV file has 1) DDA spectra (publicly available) count for blood exposome database compound list and 2) coverage (y/n) in the NIST 2020 database. First block of the InchiKey was used to query the spectral data that was downloaded from the MONA, GNPS, MS-DIAL repositories. </p> <p>It can be useful for prioritizing compounds that need to be purchased for new DDA data collection, or they can be candidates for in-silico fragmentation analyses. </p>
ANANSE: REMAP genome coverage
<p><strong>hg38: </strong>Genome coverage in bigwig format from ReMap2022 TF ChIP-seq database (hg38). All peaks in bed format (<a href="https://remap.univ-amu.fr/storage/remap2022/hg38/MACS2/remap2022_all_macs2_hg38_v1_0.bed.gz">https://remap.univ-amu.fr/storage/remap2022/hg38/MACS2/remap2022_all_macs2_hg38_v1_0.bed.gz</a>) were computed to bigwig format using:</p> <pre><code>zcat remap2022_all_macs2_hg38_v1_0.bed.gz | sed '/chrEBV/d' | cut -f 1,7,8 | bedtools slop -i - -g /hg38.fa.sizes -b 25 | sort -k 1,1 | bedtools genomecov -bg -g /hg38.fa.sizes -i - > tmp.bg && bedGraphToBigWig tmp.bg / hg38.fa.sizes remap2022.hg38.w50.bw && rm tmp.bg</code></pre> <p> </p> <p><strong>hg19: </strong>Genome coverage in bigwig format from ReMap2022 TF ChIP-seq database (hg19). All peaks in bed format (<a href="https://remap.univ-amu.fr/storage/remap2022/hg19/MACS2/remap2022_all_macs2_hg19_v1_0.bed.gz">https://remap.univ-amu.fr/storage/remap2022/hg19/MACS2/remap2022_all_macs2_hg19_v1_0.bed.gz</a>) were computed to bigwig format using:</p> <pre><code>zcat remap2022_all_macs2_hg19_v1_0.bed.gz | sed '/chrEBV/d' | cut -f 1,7,8 | bedtools slop -i - -g /hg19.fa.sizes -b 25 | sort -k 1,1 | bedtools genomecov -bg -g /hg19.fa.sizes -i - > tmp.bg && bedGraphToBigWig tmp.bg / hg19.fa.sizes remap2022.hg19.w50.bw && rm tmp.bg</code></pre> <p> </p>
E.bieneusi sequence coverage
<p>The supplementary file shows the sequence coverage of three E. bieneusi hypothetical proteins (B7XJ00, B7XHE2 and B7XIN4) detected in all seven individual plasma samples using Mascot search engine.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.