Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

709

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

709 results for “Coverage”

Learn how ShareScore rates datasets ↗
zenodo44/100

Test shapes for ultrasonic testing coverage path planning

<p>This data set contains different geometric objects. The main intention of these it to test robotic coverage path planning with an ultrasound sensor.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

Synthetic Escherichia coli mixture samples with variable coverage

<p>This dataset contains the synthetic mixture samples and reference sequences - as well as the appropriate metadata - that were originally used in the 2021 revision of the mSWEEP manuscript.<br> <br> There are 87 samples in total, each containing 100bp paired-end Illumina sequencing reads from 10 different&nbsp;<em>Escherichia coli&nbsp;</em>strains from 10 different lineages. The number of reads is set so that the sequencing coverage of the individual strains varies between 50x and 0.10x and sums up to 100x.</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Coverage of DOAJ journals' citations through OpenCitations - Result DataSet

<p>The dataset contains:&nbsp;</p> <ul> <li><strong>by_journal.json</strong>: a file containing all information extracted by Open Citations about DOAJ journals divide by year and journal name. Inside the file, the metadata about the journal are:&nbsp; <ul> <li>ISSN</li> <li>EISSN</li> <li>number of articles overall in the journal</li> <li>subject(s)&nbsp;</li> <li>number of citations received</li> <li>number of citations done</li> <li>ratio between citations done and received</li> <li>number of citations received from DOAJ journals</li> <li>number of citations done to DOAJ journals</li> <li>ratio between citations done to and received from DOAJ journals.</li> </ul> </li> </ul> <ul> <li><strong>normal.json</strong>: a file containing all information extracted from Open Citations about DOAJ journals divided only by year. Inside the file, the data by year are: <ul> <li>number of citations received.</li> <li>number of citations done.</li> <li>ratio between citations done and received.</li> <li>number of self-citations made by DOAJ inside Open Citations.</li> <li>ratio between the self-citation and the total citations received and done by DOAJ.</li> </ul> </li> </ul> <ul> <li><strong>errors.json</strong>: a file containing the count of all errors obtained from computations. Inside the file: <ul> <li>errors about records that don&#39;t have any specified date (null dates).</li> <li>errors about records that have impossible dates (wrong dates).</li> <li>errors about articles that don&#39;t have any specified Dois.</li> <li>errors about Open Citations records that don&#39;t have any Dois in the citing or cited fields.</li> </ul> </li> <li><strong>DOAJ_metrics.json</strong>: a file containing metrics about DOAJ and Open Citations, obtained by computations. Inside the file are these fields: <ul> <li>number of journals with dois.</li> <li>number of articles which have been processed during computations.</li> <li>number of used Dois. All dois (with no repetition) which are used for the adding journal operation.</li> <li>number of repeated Dois. All dois which are repeated inside the same or in another journal.</li> <li>number of accepted Dois. All articles (with repetition) which have both a defined journal and a defined doi.</li> </ul> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

[Raw data] Media Coverage of 3D Visual Tools Used in Urban Participatory Planning

<p>Raw&nbsp;information on the articles used for the publication:&nbsp;Media Coverage of 3D Visual Tools Used in Urban Participatory Planning</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Variation of and associations with the depth and evenness of sequencing coverage in a sample of archived plastid genomes

<p>Depth and evenness of sequencing coverage are considered potential indicators of genome assembly quality. In plastid genomics, where new data generation has outpaced the development of suitable assembly quality indicators, these coverage metrics could offer insights into the quality of plastomes of different sizes, structures, or taxonomic origins. However, the typical variation of sequencing depth and evenness among archived plastid genomes, their variability between plastome partitions, and any association with methodological factors have yet to be evaluated. This study explores the variation of sequencing depth and evenness across a sample of publicly accessible plastid genomes and their potential associations with plastome structure, assembly accuracy, and the methodological provenance of the genome data using statistical tests. Our results indicate significant differences in sequencing depth across the four structural partitions as well as between the coding and non-coding sections of the genomes, a significant correlation between sequencing evenness and the number of ambiguous nucleotides, and a significant difference in sequencing evenness between several DNA sequencing platforms. These findings highlight that many publicly accessible plastid genomes are based on sequence data with highly variable sequencing depth and evenness and that this variation is influenced, at least partially, by genome structure and methodological factors.</p>

opencc-by-4.0May 2024View details →
edi44/100

Herb coverage on southern pine beetle and non-southern pine beetle impacted permanent plots in Coweeta white pine watershed 1 from 2001 to 2003

Percent cover of herb layer species was compared between beetle-impacted and non-beetle-impacted white pine plots in watershed 1.

openCustomJan 2020View details →
zenodo40/100

A dataset based on two graph coverage criteria: prime-path and edge coverage

<p>This repository contains 462 instances from 6 projects. The dataset structure contains 43 columns, in which 18 columns are the source code metrics of the application methods under test, 18 columns are the source code metrics of test methods, and seven columns are the test case metrics.&nbsp;</p>

opencc-by-4.0Feb 2020View details →
zenodo40/100

Broad-Coverage German Sentiment Classification Model and Dataset for Dialog Systems

<p><a href="http://www.lrec-conf.org/proceedings/lrec2020/pdf/2020.lrec-1.202.pdf"><strong>Training a Broad-Coverage German Sentiment Classification Model for Dialog Systems</strong></a></p> <p>This paper describes the training of a general-purpose German sentiment classification model. Sentiment classification is an important aspect of general text analytics. Furthermore, it plays a vital role in dialogue systems and voice interfaces that depend on the ability of the system to pick up and understand emotional signals from user utterances. The presented study outlines how we have collected a new German sentiment corpus and then combined this corpus with existing resources to train a broad-coverage German sentiment model. The resulting data set contains 5.4 million labelled samples. We have used the data to train both, a simple convolutional and a transformer-based classification model and compared the results achieved on various training configurations. The model and the data set will be published along with this paper.</p> <p>You can find the code for training testing the models, that was published along with the paper in this <a href="https://github.com/oliverguhr/german-sentiment">repository</a>.</p> <p>The <a href="https://github.com/oliverguhr/german-sentiment-lib"><em>germansentiment</em></a> Python package contains a easy to use interface for the model that was published with this paper.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo40/100

Biological soil covers: data on lichen, bryophyte and algae coverage in soils gathered by SoilSkin citizen science program using eBryoSoil app for smartphones

<p>Biological soil covers (BSC) are small-sized topsoil communities composed mainly by lichens, bryophytes and algae that cover the terrestrial surface and play an essential role in maintaining the quality of the soil. However, little is known about their distribution, conservation, and ecosystem functions. The SoilSkin citizen science project aims to expand the scientific knowledge about the distribution of biological soil covers as an important step to evaluate the vulnerability of soil ecosystems of the Iberian Peninsula in the face of global change.</p> <p>The project has a dedicated free of charge app for smartphones (eBryoSoil, available at Google Play <a href="https://play.google.com/store/apps/details?id=com.omarfiz.ebryosoil&amp;hl=ca&amp;gl=US">https://play.google.com/store/apps/details?id=com.omarfiz.ebryosoil&amp;hl=ca&amp;gl=US</a>) that is designed to obtain information about the coverage of the BSC communities. To use this app, users must select a sampling location and capture the three soil pictures required to complete a transect. These photographs are taken at a 27 cm distance from the soil, in a straight line with 15 meters of distance between each picture. After the acquisition of each image, users can quantify the coverage percentage of biological soil covers and select the type of habitat where the transect took place. The transect is complete when all three pictures and their respective information are uploaded.</p> <p>The data presented here contains the records from SoilSkin participants, which mainly include a characterization of the cover patterns of biological soil covers, the type of habitat and the coordinates where each record was taken. The data set is composed by 279 unique records taken by 37 unique users from 28/11/2019 to 12/12/2020, across the Iberian Peninsula. These records specifically detail the percentage of cover occupied by three types of lichen growth forms (crustose, foliose and fruticose); liverworts; two types of moss growth forms (acrocarpous and pleurocarpous); algae; and soil. Moreover, each record also contains a description of the main type of habitat where the transect took place, that was selected from a list contained in the app with the following habitats:</p> <ul> <li>Dense forest - Habitat characterized by trees of more than 2 meters tall and canopy over 60%.</li> <li>Open forest &ndash; Habitat characterized by trees with more than 2 meters tall and a canopy below 60%.</li> <li>Shrubland &ndash; Habitat characterized by woody vegetation with less than 2 meters tall.</li> <li>Grassland &ndash; Habitat characterized by herbaceous plants.</li> <li>Agricultural land &ndash; Habitat characterized by temporary or woody crops.</li> <li>Coastal habitat &ndash; Habitat characterized by a landscape where land is in contact with the sea, creating a visibly different landscape from inner terrestrial one&rsquo;s.</li> <li>Urban green spaces &ndash; Habitat characterized by a landscape in which man-made structures are present.</li> </ul> <p>The database was revised to correct any possible mistakes (e.g., miscalculation of total percentages; habitat missing in some registers; removal of invalid registers).</p> <p>The data file contains the following columns:</p> <ul> <li>Date: numerical variable indicating the &ldquo;day&rdquo;/&rdquo;month&rdquo;/&rdquo;year&rdquo; when the register was generated.</li> <li>User_ID: &nbsp;categorical variable with the identification number of the user who gathered the record.</li> <li>Transect: categorical variable with the identification of the number of the transect.</li> <li>Photo_number: numeric variable that takes values of 1, 2 or 3 and corresponds with the identification of the photographs within each transect.</li> <li>Photo_label: character string with the identification of the photograph from each record.</li> <li>Register_localization: categorical variable with the identification of the geographic area where the record was done.</li> <li>Latitude: integer, variable indicating the latitude of the sampling location&nbsp;in decimal degrees.</li> <li>Longitude: integer, variable indicating the longitude of the sampling location&nbsp;in decimal degrees.</li> <li>Accuracy: integer, variable indicating the accuracy of the coordinates given by the GPS.</li> <li>Habitat_type: categorical variable with the description of the main type of habitat of the sampling location.</li> <li>Lichen_Crustose: integer, variable indicating the percentage of crustose lichen cover quantified in the record.</li> <li>Lichen_Foliose: integer, variable indicating the percentage of foliose lichen cover quantified in the record.</li> <li>Lichen_Fruticose: integer, variable indicating the percentage of fruticose lichen cover quantified in the record.</li> <li>Total_lichen: integer, variable indicating the sum of all lichen coverage quantified in the record.</li> <li>Liverwort: integer, variable indicating the percentage of liverwort cover quantified in the record.</li> <li>Moss_Acrocarpous: integer, variable indicating the percentage of acrocarpous moss cover quantified in the record.</li> <li>Moss_Pleurocarpous: integer, variable indicating the percentage of pleurocarpous moss cover quantified in the record.</li> <li>Total_ moss: integer, variable indicating the sum of all moss coverage quantified in the record.</li> <li>Algae: integer, variable indicating the percentage of algae cover quantified in the record.</li> <li>Soil: integer, variable indicating the percentage of soil visible in the record.</li> </ul> <p>&nbsp;&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

MONROE_Profiling_Mobile_Broadband_Coverage

<p>Dataset for TMA&#39;16 paper Profiling Mobile Broadband Coverage.&nbsp;</p> <p>The dataset consists of grid blocks traversed by the train routes me measure in Norway, more specifically Oslo-Stavanger, Oslo-Voss, Oslo-Trondheim, Trondheim- Bod&oslash;.</p> <p>For each of the grids and for each run on a route, we measure the Radio Access Technology an end-user could access while in the train for two different Mobile Broadband providers, namely Telenor and Netcom (Telia) in Norway.&nbsp;The dataset csv files we upload here are organized per operator and&nbsp;per route.&nbsp;</p> <p>Each row in one file consists of:</p> <p>grid_id = unique ID of the grid block that delimits a portion of the route</p> <p>lat1 =&nbsp;latitude of the grid &nbsp; &nbsp; &nbsp;</p> <p>lon1 =&nbsp;longitude of the grid&nbsp; &nbsp;</p> <p>avg_speed = average speed of the train when traversing the grid block&nbsp; &nbsp; &nbsp; &nbsp;</p> <p>start = timestamp of when the train enters the grid</p> <p>end = timestamp when the train exits&nbsp;the grid&nbsp; &nbsp; &nbsp;</p> <p>ccu_desig =&nbsp;unique ID of the NSB passenger train&nbsp; &nbsp; &nbsp; &nbsp;</p> <p>total =&nbsp;total number of datapoints within the grid block&nbsp; &nbsp;</p> <p>4G &nbsp;=&nbsp;number of points within the grid where the RAT is 4G &nbsp; &nbsp;</p> <p>3G =&nbsp;number of points within the grid where the RAT is&nbsp;3G &nbsp; &nbsp;</p> <p>2G = number of points within the grid where the RAT is 2G&nbsp; &nbsp; &nbsp;</p> <p>nos = number of points within the grid where the RAT is No Service&nbsp; &nbsp; &nbsp;</p> <p>4gd = 4G distribution in the grid block&nbsp; &nbsp;</p> <p>3gd = 3G distribution in the grid block&nbsp; &nbsp; &nbsp;</p> <p>2gd = 2G distribution in the grid block</p> <p>nosd &nbsp;= No Service distribution in the grid block&nbsp;&nbsp;</p> <p>run_id &nbsp;= the ID of the run</p> <p>static &nbsp;= 1 if the train stops at any point in the grid, 0 is the train doesn&#39;t stop in the grid</p> <p>mobile &nbsp;= 1 if the train is mobile, 0 is it is not</p> <p>full_mobile = 1 if the train is moving at all times, 0 if it is not&nbsp; &nbsp;&nbsp;</p> <p>tunnel &nbsp;= 1 if the train traverses a tunnel within the grid block, 0 if there are no tunnels</p> <p>full_tunnel = &nbsp;1 if the train is in &nbsp;train the whole time it is in the respective grid block&nbsp; &nbsp;&nbsp;</p> <p>way = route direction&nbsp; &nbsp;</p> <p>route = train route&nbsp;</p> <p>&nbsp;</p>

opencc-zeroMar 2016View details →
zenodo40/100

Bare Chested Men - German Press Coverage Corpus

<p>A list of articles published in German gaming magazines in the 1980s and 1990s about the following games:</p><ul><li><a href="https://www.mobygames.com/game/1033/death-sword/">Barbarian I</a> (1987)</li><li><a href="https://www.mobygames.com/game/12167/axe-of-rage/">Barbarian II</a> (1988)</li><li><a href="https://research.swissdigitization.ch/?p=613">DragonSlayer</a> (1989, unreleased)</li><li><a href="https://www.mobygames.com/game/54344/torvak-the-warrior/">Torvak the Warrior</a> (1990)</li><li><a href="https://www.mobygames.com/game/6182/conan-the-cimmerian/">Conan the Cimmerian</a> (1991)</li><li><a href="https://www.mobygames.com/game/1618/commando/">Commando</a> (1985)</li><li><a href="https://www.mobygames.com/game/6739/ikari-warriors/">Ikari Warriors</a> (1986)</li><li><a href="https://www.mobygames.com/game/23105/leatherneck/">Leatherneck</a> (1988)</li><li><a href="https://www.mobygames.com/game/16149/dogs-of-war/">Dogs of War</a> (1989)</li></ul>

opencc-by-4.0Oct 2023View details →
zenodo40/100

Impact Craters on Mimas, Rhea, and Iapetus (incomplete surface coverage)

<p>Comma-separated values (CSV) files of impact crater data from Iapetus, Rhea, and Mimas, based on identification of features on individual images tied to Schenk (circa 2012–2014) basemaps. &nbsp;Data include latitude, longitude, and diameter in units of decimal degrees (location) and kilometers (size). &nbsp;These data cover approximately 27%, 11%, and 73% of the surface area of each body, respectively.</p>

opencc-by-4.0Oct 2023View details →
zenodo40/100

End-To-End (Selenium and Gatling) Test Coverage in Microservices

<p>It contains an analysis to TrainTicket benchmark, selenium&nbsp;test suites, and Gatling tests. It lists the endpoints and their intersections with each others. It is generated for&nbsp;the approach that calculates three levels of coverage (microservice coverage, test suite coverage, and overall coverage).</p> <p>Checkout more details and explanations in the paper titled: On&nbsp;End-To-End&nbsp;Test Coverage in Microservices.</p>

opencc-by-4.0Jan 2024View details →
zenodo40/100

Hydraulic geometry and whitewater coverage for a steep proglacial stream -- data sets and scripts

<p>Data sets and scripts used in the analyses for the following article:</p> <p>Dufficy, A.L., Eaton, B.C. and Moore, R.D. <span>Quantifying hydraulic geometry and whitewater coverage for steep proglacial streams to support stream temperature modelling.&nbsp;<em>Hydrological Processes</em>, DOI: 10.1002/hyp.70003.<br></span></p> <p><span>The number in the file names for the R scripts indicates the order in which the scripts should be run.</span></p> <p><span>The study was funded by the Natural Sciences and Engineering Research Council of Canada and the Faculty of Arts, University of British Columbia.</span></p> <p>&nbsp;</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Data for: Mapping the Limits of Passive Samplers in Water: Chemical Space Coverage Using Nontargeted LC-HRMS Analysis

<p>This dataset provides files for passive samplers nad blanks analyzed by LC-HRMS fullscan DIA MS2.</p> <p>Excel file provides information about passive samplers, sampling site and sample files.</p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Global Pasture Watch - Grassland sampling design derived by Feature Space Coverage Sampling (FSCS) at 1-km spatial resolution

<p>Sampling design used in the production of the <strong>global maps of grassland dynamics 2000&ndash;2022 at 30 m spatial resolution</strong> in the scope of the Global Pasture Wath initiative. The sampling desing was based in Feature Space Coverage Sampling and resulted in 10,000 sample tiles (1x1 km) distributed across the World, which were visual interpreted in Very-High Resolution imagery thorugh the QGIS plugin&nbsp;<a href="https://plugins.qgis.org/plugins/qgis-fgi-plugin/">QGIS Fast Grid Inspection</a>.</p> <p>FSCS steps include:</p> <ul> <li>Short vegetation mask that includes all pixels mapped as mosaic, shrubland, grassland, and sparse vegetation in at least one year from 1993 to 2021 according to <a href="https://www.esa-landcover-cci.org/">ESA/CCI global land cover</a> (<code>gpw_short.veg.mask_esacci.lc_p_1km_s_19920101_20201231_go_epsg.3857_v1.tif</code>),</li> <li>87 input raster layers (including vegetation indices, terrain, land temperature, climate and water variable),</li> <li>Principal Components Analysis (PCA) using all input layers,</li> <li>Selection of the 10 first components (explaining 75% of variance),</li> <li>&nbsp;K-Means with 10,000 clusters (targeted number of samples - &nbsp; &nbsp;&nbsp;<br><code>gpw_grassland_fscs.kmeans.cluster_c_1km_20000101_20221231_go_epsg.3857_v1.tif</code>)</li> <li>Calculation of euclidean distance (in the principal component space) of all 1-km pixels to the centre of each cluster,</li> <li>Selection of the pixel with the shortest distance for each cluster,</li> <li>Conversion of the selected pixels into sample tiles ()</li> </ul> <p>The file&nbsp;<code>gpw_grassland_fscs_tile.samples_1km_20000101_20221231_go_epsg.3857_v1.gpkg</code> provides the sample tiles and include the follow collumns:</p> <ul> <li><strong>X</strong>: Latitude in Web Mercator projection (EPSG:3857),</li> <li><strong>Y</strong>: Longitude in Web Mercator projection (EPSG:3857),</li> <li><strong>cluster_id</strong>: K-Means output ranging from 0&mdash;9999,</li> <li><strong>cluster_distance</strong>: Distance from the selected sample to the centre of the cluster,</li> <li><strong>cluster_size</strong>: Number o 1-km pixels inside the K-Means cluster, estimated using Web Mercator projection (<a href="https://epsg.io/3857">EPSG:3857</a>)</li> <li><strong>cluster_size_equal_area</strong>: Number o 1-km pixels inside the K-Means cluster, estimated using Goode Homolosine Land projection (<a href="https://epsg.io/54052">ESRI:54052</a>)</li> <li><strong>cluster_size_corr</strong>: Correction factor to adjust the area distortion due to Web Mercator projection, estimated by the difference in normalized propotional values of cluster_size and cluster_size_equal_area.</li> <li><strong>rf_n_pred</strong>: Number of pixels predicted by a RF model trained to estimate probability to select the pixel closer to the centre of the KMeans cluster. The RF models were trained individually per each cluster using the 10 first components derived by PCA (<code>gpw_comps_fscs.pca_m_1km_20000101_20221231_go_epsg.3857_v1.tar.gz</code>).</li> <li><strong>rf_samp_prob</strong>: Sampling probability based on RF model (<em>rf_n_pred / cluster_size</em>)</li> <li><strong>rf_samp_wei</strong>: Sampling weight estimated in Web Mercator projection.</li> <li><strong>rf_samp_wei_coor</strong>: Corrected sampling weight estimated in Goode Homolosine Land projection.</li> </ul> <h3>Related resources</h3> <ul> <li><strong>Maps of dominant grassland:</strong><br><a href="https://zenodo.org/records/13890400">2000-2002</a> <a href="https://zenodo.org/records/13890402">2003-2005</a> <a href="https://zenodo.org/records/13890404">2006-2008</a> <a href="https://zenodo.org/records/13890408">2009-2011</a> <a href="https://zenodo.org/records/13890410">2012-2014</a> <a href="https://zenodo.org/records/13890412">2015-2017</a> <a href="https://zenodo.org/records/13890414">2018-2020</a> <a href="https://zenodo.org/records/13890416">2021-2022</a></li> <li><strong>Probability maps of cultivated grassland:</strong><br><a href="https://zenodo.org/records/13890401/files/ggc-30m.csv?download=1">2000-2022 (All URLs)</a></li> <li><strong>Probability maps of natural/semi-natural grassland:</strong><br><a href="https://zenodo.org/records/13890401/files/ggc-30m.csv?download=1">2000-2022 (All URLs)</a></li> <li><strong>Grassland reference samples based on VHR imagery (2000&ndash;2022):</strong><br><a href="https://doi.org/10.5281/zenodo.11281157">GeoPackage files</a></li> <li><strong>Global machine learning models (Random Forest):</strong><br><a href="https://doi.org/10.5281/zenodo.13952806">Parquet and joblib python files</a></li> <li><strong>Reference sampling design derived by FSCV:</strong><br><a href="https://doi.org/10.5281/zenodo.11391517">GeoPackage and raster files</a></li> <li><strong>Harmonized reference samples based on existing LULC dataset:</strong><br><a href="https://doi.org/10.5281/zenodo.13951976">GeoPackage and raster files</a></li> <li><strong>Source code for reproducibility:<br></strong><a href="https://doi.org/10.5281/zenodo.13952867">GitHub release</a><strong><br></strong></li> <li><strong>Mapping feedback tool:</strong><br><a href="https://geo-wiki.org">GeoWiki</a></li> <li><strong>Data catalogues:</strong><br><a href="https://stac.openlandmap.org/gpw_ggc-30m/collection.json?.language=en">OpenLandMap STAC</a> <a href="https://global-pasture-watch.projects.earthengine.app/view/ggc-30m">Google Earth Engine</a></li> </ul> <h3>Support</h3> <p>For questions of bugs/inconsistencies related to the dataset raise a GitHub issue in&nbsp;<a href="https://github.com/wri/global-pasture-watch">https://github.com/wri/global-pasture-watch</a></p>

opencc-by-4.0Nov 2024View details →
zenodo40/100

Alignment files for coverage benchmarks: Illumina and Nanopore sequencing datasets

<ul> <li><strong>cpara-illumina-noseq.bam</strong> and <strong>cpara-ont-noseq.bam</strong>:&nbsp;BAM files produced aligning the raw reads produced respectively by Illumina NextSeq and ONT Nanopore sequencing of an isolate of <em>C. parapsilosis</em>&nbsp;to evaluate the coverage calculations using real datasets.*</li> <li><strong>HG00258.bam</strong>: Exome sequencing from the 1000 Genomes Project (Clarke et al 2016&nbsp;<a href="https://doi.org/10.1093/nar/gkw829">https://doi.org/10.1093/nar/gkw829</a>).</li> <li><strong>panel_01.bam</strong>: targeted sequencing of a Human gene panel of 16 genes.*</li> </ul> <p>* Sequences and qualities have been removed</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

Data and Analysis for "On the Reliability of Coverage-based Fuzzer Benchmarking"

<pre><strong>Data and Analysis for &quot;On the Reliability of Coverage-based Fuzzer Benchmarking&quot;</strong> <strong>## Cite</strong> </pre> <pre><code>@inproceedings{benchmarking, author = {B{\"o}hme, Marcel and Szekeres, L{\'a}szl{\'o} and Metzman, Jonathan}, title = {On the Reliability of Coverage-based Fuzzer Benchmarking}, year = {2022}, booktitle = {Proceedings of the 44th International Conference on Software Engineering}, series = {ICSE '22}, pages = {1-13},  doi = {10.1145/3510003.3510230} }</code></pre> <pre> <strong>## Data Analysis</strong> The Jupyter notebook generating all tables and figures can be found in fuzzbench.manual.ipynb <strong>## Generated Images and Tables</strong> The generated data analysis artifacts are also available in this artifact. <strong>## Data</strong> All the data is available in the FuzzBench Reports and will be automatically downloaded. * 20 trials of 23 hours with 15 programs and 10 fuzzers. * Experiment name: 2021-02-17-bug-paper * Report: https://www.fuzzbench.com/reports/2021-02-17-bug-paper/index.html * Data: https://www.fuzzbench.com/reports/2021-02-17-bug-paper/data.csv.gz * Fuzzbench Commit: [38e344fef2f1079579391a0d9dcb52319f7051f2](https://github.com/google/fuzzbench/commits/38e344fef2f1079579391a0d9dcb52319f7051f2) * 30 trials of 23 hours with 11 programs and 10 fuzzers. * Experiment name: 2021-08-19-crash-s * Report: https://www.fuzzbench.com/reports/2021-08-19-crash-s/index.html and * Data: https://www.fuzzbench.com/reports/2021-08-19-crash-s/data.csv.gz * Fuzzbench Commit: db192b60815ac87f69ee0f7f3e37aeac71949e1b * 30 trials of 23 hours with 11 programs and 10 fuzzers. * Experiment name: 2021-08-19-crash-s2 * Report: https://www.fuzzbench.com/reports/2021-08-19-crash-s2/index.html and * Data: https://www.fuzzbench.com/reports/2021-08-19-crash-s2/data.csv.gz * Fuzzbench Commit: db192b60815ac87f69ee0f7f3e37aeac71949e1b The deduplicated data can be found in * 2021-02-17-bug-paper-fixed2.csv.gz * 2021-08-19-crash-s-fixed2.csv.gz * 2021-08-19-crash-s2-fixed2.csv.gz <strong>## Reproducibility</strong> </pre> <pre><code class="language-bash"># Download the precise version of FuzzBench used for the experiment git clone https://github.com/google/fuzzbench.git cd fuzzbench git checkout &lt;Fuzzbench Commit&gt; # Download the internal config file. curl https://storage.googleapis.com/[experiment-name]/config/experiment.yaml &gt; /tmp/experiment-config.yaml make install-dependencies # Launch the experiment using paramters from the internal config file. PYTHONPATH=. python experiment/reproduce_experiment.py -c /tmp/experiment-config.yaml -e &lt;new_experiment_name&gt;</code></pre> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Supplementary Information for Coverage of in situ climatological observations in the world's mountains (Thornton et al.)

<p>Supplementary Information for &quot;Coverage of in situ climatological observations in the world&#39;s mountains&quot; (Thornton et al., Frontiers in Climate).</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

CONGA: Copy number variation genotyping in ancient genomes and low-coverage sequencing data

<p>To date, ancient genome analyses have been largely confined to the study of single nucleotide polymorphisms (SNPs). Copy number variants (CNVs) are a major contributor of disease and of evolutionary adaptation, but identifying CNVs in ancient shotgun-sequenced genomes is hampered by (i) most published genomes being &lt;1x&nbsp;coverage, (ii) ancient DNA fragments being typically &lt;80 bps. These characteristics preclude state-of-the-art CNV detection software to be effectively applied to ancient genomes. Here we present CONGA, an algorithm tailored for genotyping deletion and duplication events in genomes with low depths of coverage. Simulations and down-sampling experiments show that CONGA can genotype deletions &gt;1 kbps with F-scores &gt;0.75 at &gt;=1x, and distinguish between heterozygous and homozygous states. Using CONGA, we analyse deletion events at 10,018 loci in 56 ancient human genomes spanning the last 50,000 years, with coverages 0.4x-26x. We show that inter-individual genetic diversity measured using deletions and SNPs are highly correlated, as in modern-day genomes, confirming that deletion frequencies broadly reflect demographic history. We also identify signatures of strong purifying selection on deletions in ancient-genomes, such as an excess of singletons compared to those in SNPs. CONGA paves the way for systematic studies of drift, mutation load, and adaptation in ancient and modern-day gene pools through the lens of CNVs.</p>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record