Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
257
datasets available to search
ShareScore release 0.7.1
Dataset results
257 results for “data package”
HiSim Data Package for PIEG-Strom
<p><strong>HiSim Data Package for PIEG-Strom</strong></p> <p>Collection with weather data as well as thermal and electrical load profiles for the HiSim Framework (<a href="https://github.com/FZJ-IEK3-VSA/HiSim">https://github.com/FZJ-IEK3-VSA/HiSim</a>) regarding the simulations for <em>VDI 4657 guideline</em> and the webtool <em>PIEG-Strom Online</em></p> <p><strong>Folders</strong></p> <ul> <li>electrical-loadprofiles</li> <li>thermal-loadprofiles</li> <li>weather</li> <li>photovoltaic</li> </ul> <p><strong>Information</strong></p> <ul> <li>Every folder contains a <em>readme.md</em> with information about source and license of the files in <em>data_raw</em>.</li> <li>Final data files were produced with the <em>process_data.ipynb</em> notebooks and saved into <em>data_processed</em> folder.</li> </ul>
Data from: imageseg: An R package for deep learning-based image segmentation
<p>1. Convolutional neural networks (CNNs) and deep learning are powerful and robust tools for ecological applications, and are particularly suited for image data. Image segmentation (the classification of all pixels in images) is one such application and can for example be used to assess forest structural metrics. While CNN-based image segmentation methods for such applications have been suggested, widespread adoption in ecological research has been slow, likely due to technical difficulties in implementation of CNNs and lack of toolboxes for ecologists.</p> <p>2. Here, we present R package imageseg which implements a CNN-based workflow for general-purpose image segmentation using the U-Net and U-Net++ architectures in R. The workflow covers data (pre)processing, model training, and predictions. We illustrate the utility of the package with image recognition models for two forest structural metrics: tree canopy density and understory vegetation density. We trained the models using large and diverse training data sets from a variety of forest types and biomes, consisting of 2877 canopy images (both canopy cover and hemispherical canopy closure photographs) and 1285 understory vegetation images.</p> <p>3. Overall segmentation accuracy of the models was high with a Dice score of 0.91 for the canopy model and 0.89 for the understory vegetation model (assessed with 821 and 367 images, respectively). The image segmentation models performed significantly better than commonly used thresholding methods, and generalized well to data from study areas not included in training. This indicates robustness to variation in input images and good generalization strength across forest types and biomes.</p> <p>4. The package and its workflow allow simple yet powerful assessments of forest structural metrics using pre-trained models. Furthermore, the package facilitates custom image segmentation with single or multiple classes and based on color or grayscale images, e.g. for applications in cell biology or for medical images. Our package is free, open source, and available from CRAN. It will enable easier and faster implementation of deep learning-based image segmentation within R for ecological applications and beyond.</p>
DATASET: Package data for CERF
<p>The zipped file `cerf_package_data.zip` contains the following files that support the data package for the CERF python package (see https://github.com/IMMM-SFA/cerf). All spatial data is in the projected coordinate system described here: https://immm-sfa.github.io/cerf/user_guide.html#preparing-suitability-rasters. </p> <ul> <li>cerf_conus_boundary_albers.zip: A shapefile containing polygons for each state in the conterminous United States (CONUS)</li> <li>cerf_conus_states_albers_1km.tif: A rasterized version of the CONUS states at a 1 km resolution</li> <li>config_<year>.yml: Sample configuration files for CERF</li> <li>costs_gas_pipeline.yml: YAML file containing gas pipeline costs</li> <li>costs_per_kv_substation.yml: YAML file containing the costs of interconnection for each substation class</li> <li>eia_natural_gas_pipelines_conus_albers.zip: A shapefile containing the natural gas pipeline polylines for the CONUS as retrieved from https://www.eia.gov/maps/layer_info-m.php per the 4/28/2020 update</li> <li>hifld_substations_conus_albers.zip: A shapefile containing HIFLD substations for the CONUS as retrieved from https://hifld-geoplatform.opendata.arcgis.com/datasets/geoplatform::electric-substations/about per the 07/08/2020 update</li> <li>illustrative_lmp_8760-per-zone_dollars-per-mwh.zip: A CSV file of fake illustrative locational marginal pricing (LMP) for 8760 hours per zone</li> <li>lmp_zones_1km.img: A raster of LMP zones per 1 km grid for the CONUS</li> <li>region-abbrev_to_region-name.yml: YAML file containing state abbreviation to state name</li> <li>region-name_to_region-id.yml: YAML file containing state name to state id</li> <li>suitability_<techname>.sgrd: rasters for each technology suitability from <a href="http://doi.org/10.5334/jors.227">http://doi.org/10.5334/jors.227</a></li> </ul>
spectre: An R package to estimate spatially-explicit community composition using sparse data
<p>An understanding of how biodiversity is distributed across space is key to much of ecology and conservation. Many predictive modelling approaches have been developed to estimate the distribution of biodiversity over various spatial scales. Community modelling techniques may offer many benefits over single-species modelling. However, techniques capable of estimating precise species makeups of communities are highly data intensive and thus often limited in their applicability. Here we present an R package, spectre, which can predict regional community composition at a fine spatial resolution using only sparsely sampled biological data. The package can predict the presence and absence of all species in an area, both known and unknown, at the sample site scale. Underlying the spectre package is a min-conflicts optimisation algorithm that predicts species' presences and absences throughout an area using estimates of α-, β-, and γ-diversity. We demonstrate the utility of the spectre package using a spatially-explicit simulated ecosystem to assess the accuracy of the package's results. spectre offers a simple-to-use tool with which to accurately predict community compositions across varying scales, facilitating further research and knowledge acquisition into this fundamental aspect of ecology.</p>
Data from: hespdiv: an R package for spatially constrained, hierarchical and contiguous regionalization in palaeobiogeography
<p>This is data for the '"hespdiv": an R package for spatially constrained, hierarchical, and contiguous regionalization in palaeobiogeography' paper. It contains datasets used, their metada, dataset processing scripts, a list of references to data contributors, and R files containing some of the results presented in the paper.</p>
Data providers package for reporting Chemical Contaminants (official data reporting phase) SSD2
<p>In the framework of Articles 23 and 33 of Regulation (EC) No 178/2002 EFSA has received from the European Commission a mandate (<a href="http://registerofquestions.efsa.europa.eu/roqFrontend/mandateLoader?mandate=M-2010-0374">M-2010-0374</a>) to collect all available data on the occurrence of chemical contaminants in food and feed. These data are used in EFSA’s scientific opinions and reports on contaminants in food and feed.</p> <p>This data providers package provides the data collection configuration and supporting materials for reporting <strong>Chemical Contaminants in SSD2</strong>. These are to be used for the official data reporting phase.</p> <p>The package includes:</p> <p>CHECK advice on the values to be reported for the mandatory fields specified for this data collection.</p> <p>The Standard Sample Description Version 2 XSD schema definition for CONTAMINANTS reporting.</p> <p>The STX transformation file which automatically assigns sampEventId and sampAnId when this information is not provided.</p> <p>The general and CONTAMINANTS SSD2 specific business rules applied for the automatic validation of the submitted datasets.</p> <p>Excel Mapping tool to convert excel files after mapping into XML document.</p> <p>Guidance on how to use the Excel Mapping tool.</p> <p>Guidance on how to run the validation report after submitting data to the DCF.</p>
Data providers package for reporting Chemical Contaminants (official data reporting phase) SSD1
<p>In the framework of Articles 23 and 33 of Regulation (EC) No 178/2002 EFSA has received from the European Commission a mandate (<a href="http://registerofquestions.efsa.europa.eu/roqFrontend/mandateLoader?mandate=M-2010-0374">M-2010-0374</a>) to collect all available data on the occurrence of chemical contaminants in food and feed. These data are used in EFSA’s scientific opinions and reports on contaminants in food and feed.</p> <p>This data providers package provides the data collection configuration and supporting materials for reporting <strong>Chemical Contaminants in SSD1</strong>. These are to be used for the official data reporting phase.</p> <p>The package includes:</p> <p>The Standard Sample Description Version 2 XSD schema definition for CONTAMINANTS reporting.</p> <p>The general and CONTAMINANTS SSD1 specific business rules applied for the automatic validation of the submitted datasets.</p> <p>Excel Mapping tool to convert excel files after mapping into XML document.</p> <p>Please follow the instructions below for the correct use of the mapping tool to avoid compromising its functionalities:</p> <ol> <li>Download and save the MS Excel® Standard Sample Description file to your computer (do not open the file before saving and do not change the file name)</li> <li>Download and save the file MS Excel® Simplified Reporting Format (do not open the file before saving)</li> <li>Keep both Excel files in the same folder</li> <li>Open both Excel files and enable the macros</li> <li>Keep both files open in the same Excel instance when filling in the data</li> </ol> <p>Guidance on how to run the validation report after submitting data to the DCF.</p>
Data providers package for reporting monitoring results for veterinary medicinal product residues (official data reporting phase)
<p>This data providers package provides the data collection configuration and supporting materials for reporting veterinary medicinal product residues (VMPR) results according to Council Directive 96/23/EC of 29 April 1996 on measures to monitor certain substances and residues thereof in live animals and animal products and repealing Directives 85/358/EEC and 86/469/EEC and Decisions 89/187/EEC and 91/664/EEC. These are to be used for the official data reporting phase.</p> <p>The package includes:</p> <p>Advice on the values to be reported for the mandatory fields specified for this data collection.</p> <p>The Standard Sample Description Version 2 XML schema definition for VMPR reporting.</p> <p>The STX transformation file which automatically assigns sampEventId and sampAnId when this information is not provided.</p> <p>The general and VMPR specific business rules applied for the automatic validation of the submitted datasets.</p> <p>The VMPR specific terminologies to be used for reporting analytical methods and the residues included in the scope of the analytical methods.</p> <p>An example of a reportable dataset following the Quick Start Reporting guide.</p> <p>Excel Mapping tool to convert excel files after mapping into XML document.</p> <p>Guidance on how to use the Excel Mapping tool.</p> <p>Guidance on how to run the validation report after submitting data to the DCF.</p>
Data for Matlab package locFISH to simulate realistic 3d smFISH images
<p>Different data-sets needed by the Matlab package locFISH. locFISH allows the simulation and analysis of realistic single molecule FISH (smFISH) images.</p> <p><strong>data_simulation.zip</strong><br> Contains all necessary data to simulated smFISH images. Specifically, the zip archive contains a library of 3D cell shapes, realistic imaging background, and a simulated PSF (Point Spread Function). </p> <p><strong>GAPDH.zip</strong><br> Contains the smFISH data of GAPDH and the corresponding analysis results, which were used to create the library of cell shapes provided in data_simulation.zip </p> <p>For more details on these data and how do to use them, please consult the detailed user-manual provided with <strong>locFISH</strong>, available at</p> <p>https://bitbucket.org/muellerflorian/fish_quant</p>
R data objects for the HumanDEU package
<p>This submission contains several R data objects that are part of the R package HumanDEU available through Github (https://github.com/areyesq89/HumanTissuesDEU). The objects correspond to processed data needed to reproduce the statistics, tables and figures presented in the manuscript:<br> <br> A Reyes and W Huber. Alternative start and termination sites of transcription drive most transcript isoform differences across human tissues. Nucleic Acids Research, 2017. doi: https://www.doi.org/10.1093/nar/gkx1165<br> <br> For more details, please visit the Github repository.</p>
breakpointR: an R/Bioconductor package to localize strand state changes in Strand-seq data
<p>Simulated single-end Strand-seq data with various fraction of the genome covered from 0.01 to 0.075 (COV01 etc.). All files have simulated constant level of background (0.01). Simulated reads were aligned GRCH38 using bwa-mem and duplicate reads have been marked using sambamba. (See main publication for more details).</p> <p> </p>
Metabarcoding data package for QIIME2 workshop
<p>Raw data files and intermediate results of the feature tabulation and taxonomic classification of amplicon sequences using QIIME2. This data package is used during the Metabarcoding analyses using QIIME2 workshop, which is part of the 2021 Amsterdam Science Park Study Group summer school. The complete workflow is detailed in this tutorial: https://ejongepier.github.io/metabarcoding-qiime2-workshop/genintro.html.</p>
Replication Data for the retroharmonize R Package Case Study: Working With Arab Barometer Surveys
<p>Replication datasets for the <a href="https://retroharmonize.dataobservatory.eu/articles/arabbarometer.html">retroharmonize Case Study: Working With Arab Barometer Surveys</a></p>
NCQ Dataset (NPM Package Data)
<p>Dataset for use in <a href="https://github.com/damorimRG/node_code_query">Node Code Query</a>, contains package information in a tab separated csv file. The unzipped size is ~700MB.</p> <p>You do not need to manually download this file for use in NCQ, the setup scripts will handle this for you automatically.</p> <p>The dataset contains the following fields:</p> <p><strong>Mined from the NPM registry:</strong></p> <ul> <li>Package name</li> <li>Description</li> <li>Keywords</li> <li>License</li> <li>repositoryUrl</li> <li>timeModified</li> </ul> <p><strong>Derived from data on the NPM registry:</strong></p> <ul> <li>Array of Node.js code snippets extracted from the package README using https://github.com/Brittany-Reid/npm-code-snippets</li> <li>Number of markdown code blocks in the README (number may be larger than node.js snippets, these are non-filtered)</li> <li>Number of lines in the README</li> <li>If an install example exists in the README (if a code block exists with `npm install` or a `install` header exists)</li> <li>If a run example exists in the README (if a code block exists with `npm run` or a usage header exists)</li> </ul> <p><strong>Mined from GitHub for packages with a GitHub repository (values will be 0 or false for packages missing this data)</strong></p> <ul> <li>Number of stars</li> <li>Is a fork?</li> <li>Number of forks</li> <li>Number of watchers</li> <li>If a test directory exists (if the top level directory contains a folder called `test` or `tests`)</li> </ul> <p><strong>This updated version also contains:</strong></p> <ul> <li>The confidence value of an 'able to install' prediction.</li> </ul>
Data and scripts for: track2KBA: An R package for identifying important sites for biodiversity from tracking data
<p>Data derivates and analysis scripts (in R) used for the companion paper for the R package track2KBA.</p>
Replication package for Workshop on Software Engineering 22' - What does the pytest plugins data say?
<p>This database stores the information used to run the experiment in the article: <strong>What does the pytest plugins data say?</strong></p>
Supplemental Data and Code for "An exact version of Life Table Response Experiment analysis, and the R package exactLTRE"
<p>This dataset enables the user to repeat the analyses presented in the manuscript "An exact version of Life Table Response Experiment analysis, and the R package exactLTRE." It is comprised of two compressed archives: one which contains code, and one which contains data.</p>
Data from: aniMotum, an R package for animal movement data: rapid quality control, behavioural estimation and simulation
<p>1. Animal tracking data are indispensable for understanding the ecology, behaviour and physiology of mobile or cryptic species. Meaningful signals in these data can be obscured by noise due to imperfect measurement technologies, requiring rigorous quality control as part of any comprehensive analysis. </p> <p>2. State-space models are powerful tools that separate signal from noise. These tools are ideal for quality control of error-prone location data and for inferring where animals are and what they are doing when they record or transmit other information. However, these statistical models can be challenging and time-consuming to fit to diverse animal tracking data sets. </p> <p>3. The R package <em><span>aniMotum</span></em> eases the tasks of conducting quality control on and inference of changes in movement from animal tracking data. This is achieved via: 1) a simple but extensible workflow that accommodates both novice and experienced users; 2) automated processes that alleviate complexity from data processing and model specification/fitting steps; 3) simple movement models coupled with a powerful numerical optimization approach for rapid and reliable model fitting. </p> <p>4. We highlight <em>aniMotum</em>'s<em> </em>capabilities through three applications to real animal tracking data. Full R code for these and additional applications are included as Supporting Information so users can gain a deeper understanding of how to use <em>aniMotum</em> for their own analyses. </p>
Data package for TreeSAPP's Current Protocols
<p>This dataset contains all data required for a new tutorial for the TreeSAPP. TreeSAPP is a Python package for gene-centric taxonomic and functional classification of organismal (e.g. genome) and multi-organismal (e.g. metagenome) datasets. Here, this dataset includes a reference package for the McrA gene and associated annotation files, genomic and proteomic data to query and update the reference package with, as well as metagenomes and metatranscriptomes from Sakinaw Lake, BC. Together, these data help tutorial users create and update reference package, classify query sequences, and calculate relative abundances, thereby experiencing the primary workflow modules TreeSAPP offers.</p>
Data package for Poromechanical analysis of oil well cements in CO2-rich environments
<p>Data package for the results obtained in Poromechanical analysis of oil well cements in CO2-rich environments.</p> <p>Juan Cruz Barría, Mohammadreza Bagheri, Diego Manzanal, Seyed M. Shariatipour, Jean-Michel Pereira,<br> Poromechanical analysis of oil well cements in CO2-rich environments,<br> International Journal of Greenhouse Gas Control, Volume 119, 2022, 103734, ISSN 1750-5836, https://doi.org/10.1016/j.ijggc.2022.103734.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.