Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

257

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

257 results for “data package”

Learn how ShareScore rates datasets ↗
dryad40/100

Data for: rtrees: An R package to assemble phylogenetic trees from megatrees

<p>Despite the increasingly available phylogenetic hypotheses for multiple taxonomic groups, most of them do not include all species. In phylogenetic ecology, there is still strong demand to have phylogenies with all species in a study included. The existing software tools to graft species to backbone megatrees, however, are mostly limited to a specific taxonomic group such as plants or fishes. Here, I introduce a new user-friendly R package `rtrees` that can assemble phylogenies from existing or user-provided megatrees. For most common taxonomic groups, users can only provide a vector of species' scientific names to get a phylogeny or a set of posterior phylogenies from megatrees. It is my hope that `rtrees` can provide an easy, flexible, and reliable way to assemble phylogenies from megatrees, facilitating the progress of phylogenetic ecology.</p>

opencc-zeroFeb 2023View details →
zenodo40/100

Data package for paper "Functional diversity can facilitate the collapse of an undesirable ecosystem state"

<p>Data package accompanying the paper &quot;Functional diversity can facilitate the collapse of an undesirable ecosystem state&quot;. The data package includes:</p> <ul> <li>Results and parameter of the experiments</li> <li>Measures extracted from the results for the paper</li> <li>intermediate data used for plotting</li> </ul> <p>The code is available at <a href="https://doi.org/10.5281/zenodo.7744094">10.5281/zenodo.7744094</a></p> <p>The paper is available at ENTER DOI</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

amazonULC Data Package

<p>The Amazon-ULC Data Package, available as an R package, provides Urban Land Cover (ULC) classifications for selected cities in the Brazilian Amazon. The study areas cover approximately 1,200 km&sup2;, including the municipal seats of Altamira (153 km&sup2;), Camet&aacute; (44 km&sup2;), Marab&aacute; (164 km&sup2;), Santar&eacute;m (143 km&sup2;), and part of the Metropolitan Area of Bel&eacute;m (614 km&sup2;), all located in the state of Par&aacute;.These land cover maps have significant value in urban planning for Amazonian cities, as they can aid in monitoring urban sprawl, restricting construction in environmental protection areas, assisting in urban zoning, and identifying high-density areas, among other uses. Our classification model used images from the WPM sensor of the CBERS-4A satellite, and combined the GEOBIA approach, data mining techniques, and the random machine learning algorithm.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

sager package test data

<p>This repository contains the data files used in the <a href="https://uclouvain-cbio.github.io/sager/index.html">sager</a> package. The <a href="https://uclouvain-cbio.github.io/sager/reference/sagerData.html">sagerData()</a>&nbsp;manual page describes the functions that download,&nbsp;cache the files and returns them to the user, where&nbsp;the data were originally <a href="https://www.ebi.ac.uk/pride/archive/projects/PXD016766">retrieved from</a> and how they were processed.&nbsp;</p> <p><strong>ChangeLog:</strong></p> <ul> <li>version 2: subset data files updates and added config file</li> <li>version 3: provide 3 separate subsetted mzML files, and update quant and id files (generated from re-running sage on the mzML subsets).</li> <li>version 4: udpate subset files, and remove the&nbsp;prefix from three subsetted mzML files.</li> </ul>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Reproduction package for the paper "The Apertif Radio Transient System (ARTS): Design, Commissioning, Data Release, and Detection of the first 5 Fast Radio Bursts"

<p>This is a basic reproduction package for the paper &quot;The Apertif Radio Transient System (ARTS): Design, Commissioning, Data Release, and Detection of the first 5 Fast Radio Bursts&quot; by van Leeuwen et al. (2023).</p> <p>* arXiv:<a href="https://arxiv.org/abs/2205.12362">arXiv:2205.12362</a><br> * DOI: <a href="https://doi.org/10.1051/0004-6361/202244107">10.1051/0004-6361/202244107</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
dryad40/100

Data from: Rtapas: An R package to assess cophylogenetic signal between two evolutionary histories

<p class="MsoNormal"><span>Cophylogeny represents a framework to understand how ecological and evolutionary process influence lineage diversification. The recently developed algorithm Random Tanglegram Partitions provides a directly interpretable statistic to quantify the strength of cophylogenetic signal and incorporates phylogenetic uncertainty into its estimation, and maps onto a tanglegram the contribution to cophylogenetic signal of individual host-symbiont associations. We introduce </span><span>Rtapas</span><span>, an R package to perform Random Tanglegram Partitions. </span><span>Rtapas</span><span> </span><span>applies a given global-fit method to random partial tanglegrams of a fixed size to identify the associations, terminals, and internal nodes that maximize phylogenetic congruence. This new package extends the original implementation with a new algorithm that examines the contribution to phylogenetic incongruence of each host-symbiont association and adds ParaFit, a method designed to test for topological congruence between two phylogenies, to the list of global-fit methods than can be applied. </span><span>Rtapas</span><span> </span><span>facilitates and speeds up cophylogenetic analysis, as it can handle large phylogenies (100+ terminals) in affordable computational time as illustrated with two real-world examples. </span><span>Rtapas</span><span> </span><span>can particularly cater for the need for causal inference in cophylogeny in two domains: (i) Analysis of complex and intricate host-symbiont evolutionary histories and (ii) assessment of topological (in)congruence between phylogenies produced with different DNA markers and specifically identify subsets of loci for phylogenetic analysis that are most likely to reflect gene-tree evolutionary histories.</span></p>

opencc-zeroMay 2023View details →
zenodo40/100

Input data for the case study reported in "DREAM: an R package for druggability evaluation of human complex diseases".

<p>The data included in this record constituted the input for the case study reported in the manuscript &quot;DREAM: an R package for druggability evaluation of human complex diseases&quot;, by Antonio Federico, Michele Fratello, Alisa Pavel, Lena M&ouml;bus, Giusy del Giudice, Angela Serra, Dario Greco. The data derive from transcriptomics experiments executed on lesional skin from atopic dermatitis patients and unaffected skin counterparts. The data consists of two files in &quot;.txt&quot; format reporting gene expression data in tabular format, where on the rows are reported genes and on the columns are reported samples. The data is an aggregated and batch-corrected collection of datasets originally downloaded by Gene Expression Omnibus (GEO, https://www.ncbi.nlm.nih.gov/geo/). The file &quot;GE_Mic_AD_Pamr_MAARS.txt&quot; reports gene expression estimates of lesional skin of atopic dermatitis patients, while the file &quot;GE_Mic_AD_Pamr_nl_MAARS.txt&quot; reports gene expression estimates of non-lesional skin of atopic dermatitis patients.</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Supplementary Datasets for: 'A processing and analytics system for microscopy data workflows: the Pycroscopy ecosystem of packages'

<p>The repository contains four independent datasets that are a part of the publication (<a href="https://arxiv.org/abs/2302.14629">arXiv:2302.14629</a>), which delineates the capabilities of the Pycroscopy ecosystem of packages. The details of the individual datasets can be found below.&nbsp;</p> <p>1) bfo_iv_final.hf5: Dataset of I-V curves captured by conductive atomic force microscopy&nbsp;on a BiFeO3 sample. The data has been transformed so that we plot not the log of the current density (J)&nbsp;as a function of the square root of the electric field. The dataset was originally presented in the paper&nbsp;10.1038/s41467-017-01334-5&nbsp;</p> <p>2) bto_atomic.dm3: Atomically resolved data BaTiO3 thin film acquired with scanning transmission electron microscopy. These were originally captured in the dm3 file format. This dataset was a part of the publication:&nbsp;doi.org/10.1002/adma.202106426</p> <p>3) EELS_STO.dm3: Scanning transmission electron microscope&nbsp;(STEM)-Electron energy loss spectroscopy (EELS) dataset of&nbsp;SrTiO3.</p> <p>4) STO-stack.h5:&nbsp;High-angle annular dark-field imaging&nbsp;(HAADF) scanning transmission electron microscope (STEM) image stack of SrTiO3. This image stack contains 25 images.</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo40/100

HeatResilientCity II - work package 2.3: Interactions between buildings and open space adaptation measures – Meteorological input data for building performance simulation

<p>This repository contains <strong>meteorological</strong> <strong>data</strong> from urban climate simulations that were carried out in districts of the cities of Dresden and Erfurt as part of the <a href="http://heatresilientcity.de/">HeatResilientCity II</a> project. The data was extracted at specific points (receptors) of the urban climate model. In addition to the data, a <strong>script </strong>is attached that can be utilized to generate a time series for IDA ICE building performance simulations using IceWeather.exe. Therefore, a Microsoft Windows operating system is required. To create a time series, simply use the function <em>createIdaIceInput()</em> at the end of the script <em>createTimeSeries.py</em>. Further explanations can be found at the beginning of the script. Information about the ENVI-met data used to create the IDA ICE input can be found in <em>README_RawENVImetOutput_DD.txt</em> and <em>README_RawENVImetOutput_EF.txt</em>.</p> <p>Some input <strong>data files have already been generated</strong><strong> </strong>and can be directly used for<strong> thermal building performance simulations with IDA ICE</strong>. These files can be found in the folder <em>0.3_Input_Timeseries (Climate) for IDA ICE</em>.</p> <p>The <strong>naming convention</strong> of the final input data files for IDA ICE is as follows:</p> <ul> <li>TOWN_SCENARIO_RECEPTOR_AVERAGING_INTERFACE_LATITUDE_LONGITUDE_VERSION</li> <li>TOWN: Choose between &#39;Erfurt&#39; and &#39;Dresden&#39;</li> <li>SCENARIO: See further information in <em>README_RawENVImetOutput_DD.txt</em> and <em>README_RawENVImetOutput_EF.txt</em></li> <li>RECEPTOR: Location in the modelled area (ENVI-met simulation) where data was extracted.</li> <li>AVERAGING: Information about averaging the hourly values of the urban climate simulation (see <em>createTimeSeries.py and READMEs)</em></li> <li>INTERFACE: Information on how single days were joined together (see <em>createTimeSeries.py</em>).</li> <li>LATITUDE: Default values for Dresden and Erfurt are set in the script. Add additional values in the function <em>setIceWeatherParams()</em> if you are using other cities/custom ENVI-met simulation data.</li> <li>LONGITUDE: Default values for Dresden and Erfurt are set in the script. Add additional values in the function <em>setIceWeatherParams()</em> if you are using other cities/custom ENVI-met simulation data.</li> <li>VERSION: The version number can be set in the script.</li> </ul> <p>Example: <em>Dresden_2y_A1_a_timeSeries_24-24_51.0468_13.6707_v11.prn</em></p> <p><strong>Folder overview:</strong></p> <ul> <li>The ENVI-met raw data is stored in <em>0.1_Input_RawENVImetOutput</em>.</li> <li>The script is stored in <em>0.2_Input_ScriptsToCreateTimeSeries</em>.</li> <li>The final datasets ready for simulation with IDA ICE are stored in <em>0.3_Input_Timeseries(Climate)ForIDAICE</em>. This folder also contains some weather data time series that have already been created and can be used for IDA ICE (subfolders Erfurt_v11 and Dresden_v11).</li> </ul>

opencc-by-4.0Jul 2023View details →
zenodo40/100

Reproduction package for the paper "The Impact of Hard and Easy Negative Training Data on Vulnerability Prediction Performance"

<p>This Reproduction package contains the datasets, code and results for the paper &quot;The Impact of Hard and Easy Negative Training Data on Vulnerability Prediction Performance&quot; for other researchers to use for reproducing or improving our work.&nbsp;</p>

opencc-by-4.0Feb 2023View details →
dryad40/100

Data for: rtrees: An R package to assemble phylogenetic trees from megatrees

Open the record for dataset details and reuse information.

publicMar 2023View details →
dryad40/100

Data from: Ppgm: an R package for integrating neontological, palaeontological, and climate data in a phylogenetic comparative framework

Open the record for dataset details and reuse information.

publicOct 2025View details →
dryad40/100

Data from: Cell size, photosynthesis and the package effect: an artificial selection approach

Open the record for dataset details and reuse information.

publicFeb 2023View details →
dryad40/100

specleanr: An R package for automated flagging of environmental outliers in ecological data for modeling workflows

Open the record for dataset details and reuse information.

publicNov 2025View details →
dryad40/100

spectre: An R package to estimate spatially-explicit community composition using sparse data

Open the record for dataset details and reuse information.

publicOct 2022View details →
dryad40/100

Data from: hespdiv: an R package for spatially constrained, hierarchical and contiguous regionalization in palaeobiogeography

Open the record for dataset details and reuse information.

publicMay 2024View details →
dryad40/100

Data from: The article Euclimatch: An R package for climate matching with Euclidean distance metrics

Open the record for dataset details and reuse information.

publicDec 2024View details →
dryad40/100

Data from: Rtapas: An R package to assess cophylogenetic signal between two evolutionary histories

Open the record for dataset details and reuse information.

publicMay 2023View details →
dryad40/100

Data from: aniMotum, an R package for animal movement data: rapid quality control, behavioural estimation and simulation

Open the record for dataset details and reuse information.

publicDec 2022View details →
dryad40/100

Data from: imageseg: An R package for deep learning-based image segmentation

Open the record for dataset details and reuse information.

publicAug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record