Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,481 results for “data set”

Learn how ShareScore rates datasets ↗
zenodo32/100

UOC Data Science M2.851.PR1 data set

<p>Dataset for testing</p> <p>Film Affinity Top Movies 2024</p>

opencc-zeroNov 2024View details →
dryad32/100

Data from: A set of plastid loci for use in multiplex fragment length genotyping for intraspecific variation in Pinus (Pinaceae)

Premise of the study: Recently released Pinus plastome sequences support characterization of 15 plastid Simple Sequence Repeat (ptSSR) loci originally published for P. contorta and P. thunbergii. This allows selection of loci for single-tube PCR multiplexed genotyping in any subsection of the genus. Methods: Unique placement of primers and primer conservation across the genus were investigated, and a set of six loci were selected for single-tube multiplexing. We compare interspecific variation between ptSSRs and nucleotide sequences of ycf1 then test intraspecific variation for ptSSRs using 911 samples in the P. ponderosa species complex. Results: The ptSSR loci contain mononucleotide and complex repeats with additional length variation in flanking regions. They are not located in hypervariable regions and most primers are conserved across the genus. A single PCR per sample multiplexed for six loci yielded 45 alleles in 911 samples. Discussion: The protocol allows efficient genotyping of many samples. The ptSSR loci are too variable for Pinus phylogenies but are useful for the study of genetic structure within and among populations. The multiplex method could easily be extended to other plant groups by choosing primers for ptSSR loci in a plastome alignment for the target group.

opencc-zeroDec 2013View details →
dryad32/100

Data from: Evaluating statistical multiple sequence alignment in comparison to other alignment methods on protein data sets

The estimation of multiple sequence alignments of protein sequences is a basic step in many bioinformatics pipelines, including protein structure prediction, protein family identification, and phylogeny estimation. Statistical co-estimation of alignments and trees under stochastic models of sequence evolution has long been considered the most rigorous technique for estimating alignments and trees, but little is known about the accuracy of such methods on biological benchmarks. We report the results of an extensive study evaluating the most popular protein alignment methods as well as the statistical co-estimation method BAli-Phy on 1192 protein data sets from established benchmarks as well as on 120 simulated data sets. Our study (which used more than 230 CPU years for the BAli-Phy analyses alone) shows that BAli-Phy has better precision and recall (with respect to the true alignments) than the other alignment methods on the simulated data sets, but has consistently lower recall on the biological benchmarks (with respect to the reference alignments) than many of the other methods. In other words, we find that BAli-Phy systematically under-aligns when operating on biological sequence data, but shows no sign of this on simulated data. There are several potential causes for this change in performance, including model misspecification, errors in the reference alignments, and conflicts between structural alignment and evolutionary alignments, and future research is needed to determine the most likely explanation. We conclude with a discussion of the potential ramifications for each of these possibilities.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Testing the benefits of conservation set-asides for improved habitat connectivity in tropical agricultural landscapes

1. Habitat connectivity is important for tropical biodiversity conservation. Expansion of commodity crops, such as oil palm, fragments natural habitat areas, and strategies are needed to improve habitat connectivity in agricultural landscapes. The Roundtable on Sustainable Palm Oil (RSPO) voluntary certification system requires that growers identify and conserve forest patches identified as High Conservation Value Areas (HCVAs) before oil palm plantations can be certified as sustainable. We assessed the potential benefits of these conservation set-asides for forest connectivity. 2. We mapped HCVAs and quantified their forest cover in 2015. To assess their contribution to forest connectivity, we modelled range expansion of forest-dependent populations with five dispersal abilities spanning those representative of poor dispersers (e.g., flightless insects) to more mobile species (e.g., large birds or bats) across 70 plantation landscapes in Borneo. 3. Because only 21% of HCVA area was forested in 2015, these conservation set-asides currently provide few connectivity benefits. Compared to a scenario where HCVAs contain no forest (i.e., a no-RSPO scenario), current HCVAs improved connectivity by ~3% across all dispersal abilities. However, if HCVAs were fully reforested, then overall landscape connectivity could improve by ~16%. Reforestation of HCVAs had the greatest benefit for poor to intermediate dispersers (0.5-3 km per generation), generating landscapes that were up to 2.7 times better connected than landscapes without HCVAs. By contrast, connectivity benefits of HCVAs were low for highly mobile populations under current and reforestation scenarios, because range expansion of these populations was generally successful regardless of the amount of forest cover. 4.Synthesis and applications. The RSPO requires that HCVAs be set aside to conserve biodiversity, but HCVAs currently provide few connectivity benefits because they contain relatively little forest. However, reforested HCVAs have the potential to improve landscape connectivity for some forest species (e.g., winged insects), and we recommend active management by plantation companies to improve forest quality of degraded HCVAs (e.g., by enrichment planting). Future revisions to the RSPO's Principles and Criteria (P&amp;C) should also ensure that large (i.e., with a core area &gt;2 km2) HCVAs are reconnected to continuous tracts of forest to maximise their connectivity benefits.

opencc-zeroAug 2019View details →
dryad32/100

Data from: Selective sets of mRNAs localize to extracellular paramural bodies in the rice glup6 mutant line

The transport of rice glutelin storage proteins to the storage vacuoles requires Rab5 and its cognate guanine nucleotide exchange factor (Rab5-GEF). Loss of function of these membrane vesicular trafficking factors results in the initial secretion of storage proteins and later their partial engulfment by the plasma membrane to form an extracellular paramural body (PMB); an aborted endosome complex. Here we show that in the Rab5-GEF mutant glup6 line, glutelin RNAs are specifically mis-localized from their normal location on the cisternal-ER to the protein body-ER and also apparently translocated to the PMBs. We substantiated the association of mRNAs with this aborted endosome complex by RNA-seq of PMBs purified by flow cytometry. Two PMB-associated RNA groups are readily resolved: those that are specifically enriched in this aborted complex and those that are highly expressed in the cytoplasm. Examination of the PMB-enriched RNAs indicates that they were not a random sampling of the glup6 transcriptome but, instead, encompassed only a few functional mRNA classes. Although specific autophagy is also an alternative mechanism, these studies support the view that RNA localization may co-opt membrane vesicular trafficking and that many RNAs, which share function or intracellular location, are co-transported in developing rice seeds.

opencc-zeroDec 2017View details →
zenodo32/100

FIGURE 3. Discriminant analysis. Data set included 140 in Morphometric analysis to differentiate taxonomically seven species of Eleutherodactylus (Amphibia: Anura: Leptodactylidae) from an Andean cloud forest of Colombia

FIGURE 3. Discriminant analysis. Data set included 140 of four species (Eleutherodactylus douglasi, E. merostictus, E. miyatai and E. prolixodiscus) and 11 quantitative variables. The model utilized stepwise discrimination in which all variables were included in the model and then, at each step, the variable that contributed least to the prediction of group memberships was eliminated.

opennotspecifiedJul 2005View details →
zenodo32/100

Assemblies of the CAMI 2 challenge data sets

<p>Assembly submissions, including participants&#39; submissions, of the CAMI 2 challenge data sets.</p> <p>The gold standards can be found either on the CAMI 2 challenge website:&nbsp;<a href="https://data.cami-challenge.org/participate">https://data.cami-challenge.org/participate</a>&nbsp;or in publisso alongside the data sets:&nbsp;https://repository.publisso.de/resource/frl:6425521</p> <p>Also see&nbsp;<a href="https://www.microbiome-cosi.org/cami">https://www.microbiome-cosi.org/cami</a>&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Architectural Uncertainty Analysis for Access Control Scenarios in Industry 4.0 - Data Set

<p>This data set contains additional information to the master&#39;s thesis of Nicolas Boltz. Included are the implementation, tests, and model instances of sample scenarios.</p>

openepl-2.0Jul 2021View details →
zenodo32/100

Data set for phase-field studies of crystal growth in single-seal syntaxial veins in limestone

<p>The numerical data in this repository consists of the simulation data of single-seal syntaxial calcite vein formation in limestone. The simulations were performed using&nbsp; the software package named &quot;Pace3D (v. 2.5.1)&quot;.</p> <p>The data is organized, the way it appears in the figures in the manuscript and the folders are named accordingly.</p> <ul> <li>The simulation data shows intermediate growth stages and was converted from Pace3D output data format to VTK data format. The VTK files can be visualized using open source software packages like Paraview. The data files in each subfolder are also compressed (file format *.gz). For visualization the data has to be decompressed (e.g. with gzip, 7zip).</li> </ul>

opencc-by-4.0Mar 2021View details →
zenodo32/100

CASINO V2 data sets

<p>Raw data (*.dat format, basically text file) files directly exported from simulations of CASINO V2.</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Data set of the article "Comparative analysis of the unbinding pathways of antiviral drug Indinavir from HIV and HTLV1 proteases by supervised molecular dynamics simulation"

<p>Data set of the article &quot;Comparative analysis of the unbinding pathways of antiviral drug Indinavir from HIV and HTLV1 proteases by supervised molecular dynamics simulation&quot;</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Data set for 'Remote control of neural function by X-ray-induced scintillation'

<p>Data set for: Matsubara T, Yanagida T, Kawaguchi N, Nakano T, Yoshimoto J, Sezaki M, Takizawa H, Tsunoda SP, Horigane S, Ueda S, Takemoto-Kimura S, Kandori H, Yamanaka A,&nbsp;Yamashita T&nbsp;(2021) Remote control of neural function by X-ray-induced scintillation. Nature Communications 12: 4478, doi 10.1038/s41467-021-24717-1</p> <p>There are 2 files in this upload:</p> <p>1. The file named &quot;Matsubara_data.zip&quot; (~18 GB) is a zipped version of a folder &quot;Dataset&quot; (~27 GB), which contains images,&nbsp;movies,&nbsp;and other&nbsp;data analyzed&nbsp;in the study. When unzipped, the folder contains subfolders categorized based on experiment&nbsp;types.&nbsp; A &#39;read me&#39; text file is&nbsp;associated with each small dataset.</p> <p>2.&nbsp;The file named &quot;Matsubara_published.zip&quot; (~5 MB) contains&nbsp;the Open Access pdf&nbsp;of the paper,&nbsp;the Supplementary Information file, and the Source Data file.</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Data and code for "Characterizing mid-circuit measurements on a superconducting qubit using gate set tomography" v1.0.2

<p>This is the supplemental code and data for arXiv:2103.03008 &quot;Characterizing mid-circuit measurements on a superconducting qubit using gate set tomography&quot;.</p>

opencc-by-4.0Jul 2021View details →
zenodo32/100

Test data set for Aquaria

<p>Data sets for testing zenodo in Aquaria.</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Data Set S1. Microplastic properties and Data Set S2. Sediment grain size

<p>The data for microplastic properties (including abundance, shape, colour, polymer type and size) and sediment grain size (including the grain size and the content of clay, silt and sand) in Core CCYY1.</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

Data set

<p>Data set for article &quot;Application of piezoelectric fast tool servo for turning non-circular shapes made of 6082 aluminum alloy&quot; in&nbsp;Applied Sciences</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

data set related to article Suicidal ideation, suicidal attempts and non-suicidal self-injuries in referred adolescents

<p>This record contains raw data to article Suicidal ideation, suicidal attempts and non-suicidal self-injuries in referred adolescents</p>

opencc-by-4.0Aug 2021View details →
zenodo32/100

data set related to article Suicidal ideation and suicidal attempts in referred adolescents with high functioning autism spectrum disorder and comorbid bipolar disorder: A pilot study

<p>This record contains raw data related to article Suicidal ideation and suicidal attempts in referred adolescents with high functioning autism spectrum disorder and comorbid bipolar disorder: A pilot study</p>

opencc-by-4.0Aug 2021View details →
dryad32/100

Data from: Soil dynamics in forest restoration: a data set for temperate and tropical regions

<p>Restoring forest ecosystems has become a global priority. Yet, soil dynamics is still poorly assessed among restoration studies and lacks knowledge on how soil is affected by forest restoration process. Here, we compile information on soil dynamics in forest restoration based on soil physical, chemical and biological attributes in temperate and tropical forest regions.  It encompasses 50 scientific papers across 17 different countries and contains 1,469 quantitative information of soil attributes between reference (e.g., old-growth forest) and restored ecosystems (e.g. forests in their initial or secondary stage of succession) within the same study. To be selected, studies had to be conducted in forest ecosystems, to include multiple sampling sites (replicates) in both restored and reference ecosystems, and to encompass quantitative data of soil attributes for both reference and restored ecosystems.</p> <p>We recorded in each study the following information: (i) study year; (ii) country; (iii) forest region (tropical or temperate); (iv) latitude; (v) longitude; (vi) soil class; (vii) past disturbance; (viii) restoration strategy (active or passive); (ix) restoration age; (x) soil attribute type (physical, chemical or biological); (xi) soil attribute; (xii) soil attribute unit; (xiii) soil sampling (procedures); (xiv) date of sampling; (xv) soil depth sampled; (xvi) soil analysis; (xvii) quantitative values of soil attributes for both restored and reference ecosystems; (xviii) type of variation (standard error ou deviation) for both restored and reference ecosystems; and (xix) quantitative values of the variation for both restored and reference ecosystems. These were the most common data available in the selected studies.</p> <p>This extensive database on the extent soil physical, chemical and biological attributes differ between reference and restored ecosystems can fill part of the existing gap on both soil science and forest restoration in terms of: (i) which are the critical soil attributes to be monitored during forest restoration? and (ii) how do environmental factors affect soil attributes in forest restoration? The data will be made available to the scientific community for further analyses on both soil science and forest restoration. Soil information gap during the forest restoration process and its general patterns can be addressed using this data set.</p>

opencc-zeroAug 2021View details →
zenodo32/100

Figure 2 in Biogeography, land snails and incomplete data sets: the case of three island groups in the Aegean Sea

Figure 2. Map of the island group of Astypalaia, with islands and sites studied. 13: Diapori island.

opennotspecifiedFeb 2008View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record