Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
257
datasets available to search
ShareScore release 0.7.1
Dataset results
257 results for “data package”
Data for: rtrees: An R package to assemble phylogenetic trees from megatrees
<p>Despite the increasingly available phylogenetic hypotheses for multiple taxonomic groups, most of them do not include all species. In phylogenetic ecology, there is still strong demand to have phylogenies with all species in a study included. The existing software tools to graft species to backbone megatrees, however, are mostly limited to a specific taxonomic group such as plants or fishes. Here, I introduce a new user-friendly R package `rtrees` that can assemble phylogenies from existing or user-provided megatrees. For most common taxonomic groups, users can only provide a vector of species' scientific names to get a phylogeny or a set of posterior phylogenies from megatrees. It is my hope that `rtrees` can provide an easy, flexible, and reliable way to assemble phylogenies from megatrees, facilitating the progress of phylogenetic ecology.</p>
Data package for paper "Functional diversity can facilitate the collapse of an undesirable ecosystem state"
<p>Data package accompanying the paper "Functional diversity can facilitate the collapse of an undesirable ecosystem state". The data package includes:</p> <ul> <li>Results and parameter of the experiments</li> <li>Measures extracted from the results for the paper</li> <li>intermediate data used for plotting</li> </ul> <p>The code is available at <a href="https://doi.org/10.5281/zenodo.7744094">10.5281/zenodo.7744094</a></p> <p>The paper is available at ENTER DOI</p>
amazonULC Data Package
<p>The Amazon-ULC Data Package, available as an R package, provides Urban Land Cover (ULC) classifications for selected cities in the Brazilian Amazon. The study areas cover approximately 1,200 km², including the municipal seats of Altamira (153 km²), Cametá (44 km²), Marabá (164 km²), Santarém (143 km²), and part of the Metropolitan Area of Belém (614 km²), all located in the state of Pará.These land cover maps have significant value in urban planning for Amazonian cities, as they can aid in monitoring urban sprawl, restricting construction in environmental protection areas, assisting in urban zoning, and identifying high-density areas, among other uses. Our classification model used images from the WPM sensor of the CBERS-4A satellite, and combined the GEOBIA approach, data mining techniques, and the random machine learning algorithm. </p>
sager package test data
<p>This repository contains the data files used in the <a href="https://uclouvain-cbio.github.io/sager/index.html">sager</a> package. The <a href="https://uclouvain-cbio.github.io/sager/reference/sagerData.html">sagerData()</a> manual page describes the functions that download, cache the files and returns them to the user, where the data were originally <a href="https://www.ebi.ac.uk/pride/archive/projects/PXD016766">retrieved from</a> and how they were processed. </p> <p><strong>ChangeLog:</strong></p> <ul> <li>version 2: subset data files updates and added config file</li> <li>version 3: provide 3 separate subsetted mzML files, and update quant and id files (generated from re-running sage on the mzML subsets).</li> <li>version 4: udpate subset files, and remove the prefix from three subsetted mzML files.</li> </ul>
Reproduction package for the paper "The Apertif Radio Transient System (ARTS): Design, Commissioning, Data Release, and Detection of the first 5 Fast Radio Bursts"
<p>This is a basic reproduction package for the paper "The Apertif Radio Transient System (ARTS): Design, Commissioning, Data Release, and Detection of the first 5 Fast Radio Bursts" by van Leeuwen et al. (2023).</p> <p>* arXiv:<a href="https://arxiv.org/abs/2205.12362">arXiv:2205.12362</a><br> * DOI: <a href="https://doi.org/10.1051/0004-6361/202244107">10.1051/0004-6361/202244107</a></p> <p> </p>
Data from: Rtapas: An R package to assess cophylogenetic signal between two evolutionary histories
<p class="MsoNormal"><span>Cophylogeny represents a framework to understand how ecological and evolutionary process influence lineage diversification. The recently developed algorithm Random Tanglegram Partitions provides a directly interpretable statistic to quantify the strength of cophylogenetic signal and incorporates phylogenetic uncertainty into its estimation, and maps onto a tanglegram the contribution to cophylogenetic signal of individual host-symbiont associations. We introduce </span><span>Rtapas</span><span>, an R package to perform Random Tanglegram Partitions. </span><span>Rtapas</span><span> </span><span>applies a given global-fit method to random partial tanglegrams of a fixed size to identify the associations, terminals, and internal nodes that maximize phylogenetic congruence. This new package extends the original implementation with a new algorithm that examines the contribution to phylogenetic incongruence of each host-symbiont association and adds ParaFit, a method designed to test for topological congruence between two phylogenies, to the list of global-fit methods than can be applied. </span><span>Rtapas</span><span> </span><span>facilitates and speeds up cophylogenetic analysis, as it can handle large phylogenies (100+ terminals) in affordable computational time as illustrated with two real-world examples. </span><span>Rtapas</span><span> </span><span>can particularly cater for the need for causal inference in cophylogeny in two domains: (i) Analysis of complex and intricate host-symbiont evolutionary histories and (ii) assessment of topological (in)congruence between phylogenies produced with different DNA markers and specifically identify subsets of loci for phylogenetic analysis that are most likely to reflect gene-tree evolutionary histories.</span></p>
Input data for the case study reported in "DREAM: an R package for druggability evaluation of human complex diseases".
<p>The data included in this record constituted the input for the case study reported in the manuscript "DREAM: an R package for druggability evaluation of human complex diseases", by Antonio Federico, Michele Fratello, Alisa Pavel, Lena Möbus, Giusy del Giudice, Angela Serra, Dario Greco. The data derive from transcriptomics experiments executed on lesional skin from atopic dermatitis patients and unaffected skin counterparts. The data consists of two files in ".txt" format reporting gene expression data in tabular format, where on the rows are reported genes and on the columns are reported samples. The data is an aggregated and batch-corrected collection of datasets originally downloaded by Gene Expression Omnibus (GEO, https://www.ncbi.nlm.nih.gov/geo/). The file "GE_Mic_AD_Pamr_MAARS.txt" reports gene expression estimates of lesional skin of atopic dermatitis patients, while the file "GE_Mic_AD_Pamr_nl_MAARS.txt" reports gene expression estimates of non-lesional skin of atopic dermatitis patients.</p>
Supplementary Datasets for: 'A processing and analytics system for microscopy data workflows: the Pycroscopy ecosystem of packages'
<p>The repository contains four independent datasets that are a part of the publication (<a href="https://arxiv.org/abs/2302.14629">arXiv:2302.14629</a>), which delineates the capabilities of the Pycroscopy ecosystem of packages. The details of the individual datasets can be found below. </p> <p>1) bfo_iv_final.hf5: Dataset of I-V curves captured by conductive atomic force microscopy on a BiFeO3 sample. The data has been transformed so that we plot not the log of the current density (J) as a function of the square root of the electric field. The dataset was originally presented in the paper 10.1038/s41467-017-01334-5 </p> <p>2) bto_atomic.dm3: Atomically resolved data BaTiO3 thin film acquired with scanning transmission electron microscopy. These were originally captured in the dm3 file format. This dataset was a part of the publication: doi.org/10.1002/adma.202106426</p> <p>3) EELS_STO.dm3: Scanning transmission electron microscope (STEM)-Electron energy loss spectroscopy (EELS) dataset of SrTiO3.</p> <p>4) STO-stack.h5: High-angle annular dark-field imaging (HAADF) scanning transmission electron microscope (STEM) image stack of SrTiO3. This image stack contains 25 images.</p> <p> </p>
HeatResilientCity II - work package 2.3: Interactions between buildings and open space adaptation measures – Meteorological input data for building performance simulation
<p>This repository contains <strong>meteorological</strong> <strong>data</strong> from urban climate simulations that were carried out in districts of the cities of Dresden and Erfurt as part of the <a href="http://heatresilientcity.de/">HeatResilientCity II</a> project. The data was extracted at specific points (receptors) of the urban climate model. In addition to the data, a <strong>script </strong>is attached that can be utilized to generate a time series for IDA ICE building performance simulations using IceWeather.exe. Therefore, a Microsoft Windows operating system is required. To create a time series, simply use the function <em>createIdaIceInput()</em> at the end of the script <em>createTimeSeries.py</em>. Further explanations can be found at the beginning of the script. Information about the ENVI-met data used to create the IDA ICE input can be found in <em>README_RawENVImetOutput_DD.txt</em> and <em>README_RawENVImetOutput_EF.txt</em>.</p> <p>Some input <strong>data files have already been generated</strong><strong> </strong>and can be directly used for<strong> thermal building performance simulations with IDA ICE</strong>. These files can be found in the folder <em>0.3_Input_Timeseries (Climate) for IDA ICE</em>.</p> <p>The <strong>naming convention</strong> of the final input data files for IDA ICE is as follows:</p> <ul> <li>TOWN_SCENARIO_RECEPTOR_AVERAGING_INTERFACE_LATITUDE_LONGITUDE_VERSION</li> <li>TOWN: Choose between 'Erfurt' and 'Dresden'</li> <li>SCENARIO: See further information in <em>README_RawENVImetOutput_DD.txt</em> and <em>README_RawENVImetOutput_EF.txt</em></li> <li>RECEPTOR: Location in the modelled area (ENVI-met simulation) where data was extracted.</li> <li>AVERAGING: Information about averaging the hourly values of the urban climate simulation (see <em>createTimeSeries.py and READMEs)</em></li> <li>INTERFACE: Information on how single days were joined together (see <em>createTimeSeries.py</em>).</li> <li>LATITUDE: Default values for Dresden and Erfurt are set in the script. Add additional values in the function <em>setIceWeatherParams()</em> if you are using other cities/custom ENVI-met simulation data.</li> <li>LONGITUDE: Default values for Dresden and Erfurt are set in the script. Add additional values in the function <em>setIceWeatherParams()</em> if you are using other cities/custom ENVI-met simulation data.</li> <li>VERSION: The version number can be set in the script.</li> </ul> <p>Example: <em>Dresden_2y_A1_a_timeSeries_24-24_51.0468_13.6707_v11.prn</em></p> <p><strong>Folder overview:</strong></p> <ul> <li>The ENVI-met raw data is stored in <em>0.1_Input_RawENVImetOutput</em>.</li> <li>The script is stored in <em>0.2_Input_ScriptsToCreateTimeSeries</em>.</li> <li>The final datasets ready for simulation with IDA ICE are stored in <em>0.3_Input_Timeseries(Climate)ForIDAICE</em>. This folder also contains some weather data time series that have already been created and can be used for IDA ICE (subfolders Erfurt_v11 and Dresden_v11).</li> </ul>
Reproduction package for the paper "The Impact of Hard and Easy Negative Training Data on Vulnerability Prediction Performance"
<p>This Reproduction package contains the datasets, code and results for the paper "The Impact of Hard and Easy Negative Training Data on Vulnerability Prediction Performance" for other researchers to use for reproducing or improving our work. </p>
Data for: rtrees: An R package to assemble phylogenetic trees from megatrees
Open the record for dataset details and reuse information.
Data from: Ppgm: an R package for integrating neontological, palaeontological, and climate data in a phylogenetic comparative framework
Open the record for dataset details and reuse information.
Data from: Cell size, photosynthesis and the package effect: an artificial selection approach
Open the record for dataset details and reuse information.
specleanr: An R package for automated flagging of environmental outliers in ecological data for modeling workflows
Open the record for dataset details and reuse information.
spectre: An R package to estimate spatially-explicit community composition using sparse data
Open the record for dataset details and reuse information.
Data from: hespdiv: an R package for spatially constrained, hierarchical and contiguous regionalization in palaeobiogeography
Open the record for dataset details and reuse information.
Data from: The article Euclimatch: An R package for climate matching with Euclidean distance metrics
Open the record for dataset details and reuse information.
Data from: Rtapas: An R package to assess cophylogenetic signal between two evolutionary histories
Open the record for dataset details and reuse information.
Data from: aniMotum, an R package for animal movement data: rapid quality control, behavioural estimation and simulation
Open the record for dataset details and reuse information.
Data from: imageseg: An R package for deep learning-based image segmentation
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.