Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3,481

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3,481 results for “data set”

Learn how ShareScore rates datasets ↗
zenodo36/100

Data set for D6.7 with questionnaire responses

<p>Response from a survey of LTA industry partners on the use of monitoring in their activities.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Data set to: Mapping a brain parasite: occurrence and spatial distribution in fish encephalon

<p>Data for the manuscript &quot;Mapping a brain parasite: occurrence and spatial distribution in fish encephalon&quot;, doi:&nbsp;10.1016/j.ijppaw.2023.03.004. Description of the distribution of metacercariae from the trematode species <em>Cardiocephaloides longicollis</em> in the brain of fish.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Data Set "Efficient automatic construction of atom-economical QM regions with point-charge variation analysis"

<p>This data set accompanies the publication &quot;Efficient automatic construction of atom-economical QM regions with point-charge variation analysis&quot;&nbsp;by Felix Brandt and Christoph R. Jacob (TU Braunschweig, Germany)&nbsp;</p> <p>It contains the following files:</p> <p>- PDB files of the reactant and product starting structure</p> <p>- modified AMBER95 force field file</p> <p>- AMS fragment files for the ligands and ions</p> <p>- AMS input files for all geometry optimizations and single point calculations</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Experimental data set for the article entitled "Mathematical Model of Steam Reforming in the Anode Channel of a Molten Carbonate Fuel Cell"

<p>Experimental data for the article: Szablowski, L.; Dybinski, O.; Szczesniak, A.; Milewski, J. Mathematical Model of Steam Reforming in the Anode Channel of a Molten Carbonate Fuel Cell. Energies 2022, 15, 608.&nbsp;The experiments were performed by the first two authors.<br> These data set refer to experiments carried out on a stand used to test high-temperature fuel cells. The subject of the study was a molten carbonate fuel cell fueled with a mixture of methane and steam with steam to carbon ratio of 2.0, 2.5, 3.0 and 3.5 and at the cell operating temperature of 550&deg;C and 650&deg;C. Additionally, in the anode channel of the cell, there was a catalyst in the amount of 2 g. The active area of the cell was 20.25 cm<sup>2</sup>. The article that uses these research results is published in an open access journal with a CC-BY license. This research was funded by the National Science Center, Poland (Grant number 2020/39/D/ST8/02021).</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

GRAND-SLAM analysis of simulated nucleotide conversion in Illumina TruSeq data sets for grandRescue

<p>These are processed data sets from the simulation of nucleotide conversions (T&gt;C) in single-end and paired-end Illumina TruSeq reads for the purpose of investigating 4sU-induced mapping impairment by read lengths and library preparation methods and the potential of grandRescue to alleviate these effects.</p> <p>The original data set is from: Sarantopoulou, D. <em>et al. </em>(https://doi.org/10.1038/s41598-019-49889-1)</p> <p>GEO Accession:GSE124167 (samples: GSM3523316 - GSM3523318)</p> <p>&nbsp;</p> <p>The zip files contain the full output from the processing pipeline (including the mapped reads, the scripts to run the pipeline and the output) for single-end (R1) and paired-end before and after rescue. The *.tsv.gz files are the GRAND-SLAM output tables.</p> <p><br> To generate the GRAND-SLAM output yourself, first prepare the mouse genomes. Then run the following command with the respective cit-files, prefixes (*.cit) and genome:</p> <p>gedi -e Slam -trim5p 15 -reads *.cit -genomic m.ens102 -prefix grandslam_t15/* -plot&nbsp; -D -modelall</p> <p>To generate the cit file you have to modify the first lines in start.bash to match the paths on your file system, and then run it.</p> <p>&nbsp;</p> <p>Software versions:</p> <p>&nbsp;&nbsp;&nbsp; gedi toolkit 1.0.5<br> &nbsp;&nbsp;&nbsp; GRAND-SLAM 2.0.7<br> &nbsp;&nbsp;&nbsp; STAR version 2.7.10b</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

GRAND-SLAM analysis of simulated nucleotide conversion in QuantSeq data sets for grandRescue

<p>These are processed data sets from the simulation of nucleotide conversions (T&gt;C) in QuantSeq reads for the purpose of investigating 4sU-induced mapping impairment by read lengths and library preparation methods and the potential of grandRescue to alleviate these effects.</p> <p>The original data set is from: Lee, J. W. <em>et al. </em>(https://doi.org/10.1038/s41586-019-1004-y)</p> <p>GEO Accession: GSE109480 (Samples: GSM2944116 &ndash; GSM2944120)</p> <p>&nbsp;</p> <p>The zip files contain the full output from the processing pipeline (including the mapped reads, the scripts to run the pipeline and the output) before and after rescue. The *.tsv.gz files are the GRAND-SLAM output tables.</p> <p><br> To generate the GRAND-SLAM output yourself, first prepare the mouse genome. Then run the following command with the respective cit-files, prefixes (*.cit) and genome:</p> <p>gedi -e Slam -trim5p 15 -reads *.cit -genomic m.ens102 -prefix grandslam_t15/* -plot&nbsp; -D -modelall</p> <p>To generate the cit file you have to modify the first lines in start.bash to match the paths on your file system, and then run it.</p> <p>&nbsp;</p> <p>Software versions:</p> <p>&nbsp;&nbsp;&nbsp; gedi toolkit 1.0.5<br> &nbsp;&nbsp;&nbsp; GRAND-SLAM 2.0.7<br> &nbsp;&nbsp;&nbsp; STAR version 2.7.10b</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Digital Commensality data-set

<p>The Digital Commensality Data-set consists of facial activity data of 11 pairs of persons sharing a meal online through a videoconferencing software and self-reported qualitative and qualitative measures of their commensal experience (Computer-Mediated Communication questionnaire and Digital Commensality questionnaire). Facial activity data is extracted using the OpenFace tool.</p> <p>If you use this data-set for the uses allowed in the license, e.g., research purposes, please add the following citation:</p> <p>Ceccaldi, E., Niewiadomski, R., Mancini, M., &amp; Volpe, G. (2022). What&#39;s on your plate? Collecting multimodal data to understand commensal behavior. <em>Frontiers in Psychology</em>, <em>13</em>.</p> <p>https://doi.org/10.3389/fpsyg.2022.911000</p>

openother-ncSep 2022View details →
zenodo36/100

polyOne Data Set - 100 million hypothetical polymers including 29 properties

<p><strong>polyOne Data Set</strong></p> <p>The data set contains 100 million hypothetical polymers each with 29 predicted properties using&nbsp;machine learning models. We use&nbsp;PSMILES strings to represent&nbsp;polymer structures,&nbsp;see <a href="https://www.polymergenome.org/guide/index.php?m=3">here</a> and <a href="https://github.com/Ramprasad-Group/psmiles">here</a>. The polymers are generated by decomposing previously synthesized polymers into unique chemical fragments.&nbsp;Random and enumerative compositions of these fragments yield 100 million hypothetical PSMILES strings. All PSMILES strings are chemically valid polymers but, mostly, have never been synthesized before. More information can be found in the paper. Please note the&nbsp;license agreement in the LICENSE file.</p> <p><strong>Full data set including the properties</strong></p> <p>The data files are in Apache&nbsp;Parquet format. The files start with `polyOne_*.parquet`.</p> <p>I recommend using dask (`pip install dask`) to load and process the data set.&nbsp;&nbsp;Pandas also works but is slower.</p> <p>Load sharded data set with dask<br> ```python<br> import dask.dataframe as dd<br> ddf = dd.read_parquet(&quot;*.parquet&quot;, engine=&quot;pyarrow&quot;)<br> ```</p> <p>For example, compute the description of data set<br> ```python<br> df_describe = ddf.describe().compute()<br> df_describe</p> <p>```</p> <p><strong>PSMILES strings only</strong></p> <ul> <li>generated_polymer_smiles_train.txt -&nbsp;80 million PSMILES strings for training&nbsp;polyBERT.&nbsp;One string per line.</li> <li>generated_polymer_smiles_dev.txt - 20 million PSMILES strings&nbsp;for testing&nbsp;polyBERT. One string per line.</li> </ul>

openother-ncSep 2022View details →
zenodo36/100

Data set for "Tectonic evolution of the Tibetan Plateau during the late Cretaceous to early Eocene: Insights from geochemical records in the Fenghuoshan Group, Hoh Xil Basin"

<p>The mineral compositions, major&nbsp;and trace element gechemical data for the sediements from Fenghuoshan Group, Hoh Xil Basin.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Data set of simulated rimed aggregates for "A riming-dependent parameterization of scattering by snowflakes using the self-similar Rayleigh-Gans approximation"

<p><strong>Simulated rimed aggregates</strong> generated with https://github.com/jleinonen/aggregation in setting &quot;aggregation followed by riming&quot;.</p> <p>Aggregates were built from between 10 to 700 monomer crystals of <strong>columns, dendrites, needles, plates or rosettes</strong> with mean sizes of 100 or 200 micrometer. Then they were exposed to ELWP = 2.0 kg m⁻&sup2;. Monomer crystals are composed of cubical elements with resolution 20 micrometer. Frozen rime droplets are also represented by 20 micrometer cubes.</p> <p>The data set contains folders with <strong>evolution (evol) and shape files for each monomer crystal type</strong>. For each particle one evolution and one corresponding shape file exists. The evolution (evol) file contains particle mass, rime mass, area, size, fall speed (Heymsfield&amp;Westbrook, 2010), fall speed (Khvorostyanov&amp;Curry, 2005) for each step during the aggregation and riming process. The corresponding shape file contains the x,y,z positions of the cubical elements that compose the particle for each step. <strong>For further documentation see readme.</strong></p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

XRF analysis of Gotlandic box brooch data set

<p>XRF analasis of Gotlandic box brooch, performed in order to evaluate the material, composition of alloy, and where there are different areas of silver. The information is valuable when conservating the object and cleaning up original surfaces. Results show that brooch is multi coloured with brass, bronze and silver.</p> <p>Raw data set in .rtx format.&nbsp;ARTAX soft ware will be needed to open files.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Case study result data set for Energy Systems (submitted) article "Influence of hydrogen import prices on hydropower systems in climate-neutral Europe"

<p>The data set contains result data for the European system in a long term climate-neutral European energy system (scenario year 2050) as described in the publication &quot;Influence of hydrogen import prices on hydropower systems in climate-neutral Europe&quot;. The results have been generated with the model SCOPE SD of Fraunhofer Institute for Energy Economics and Energy System Technology IEE.&nbsp;</p> <p><strong>Abbreviations:</strong></p> <ul> <li>BEV - Battery Electric Vehicles</li> <li>CCGT - Combined Cycle Gas Turbine</li> <li>CHP - Combined heat and power</li> <li>con - consumption</li> <li>gen - generation</li> <li>HighCLEQ - High import prices / clustered-equivalent hydropower units</li> <li>HighEQ - High import prices / equivalent hydropower units</li> <li>LowCLEQ - Low import prices / clustered-equivalent hydropower units</li> <li>LowEQ - Low import prices / equivalent hydropower units</li> <li>MedCLEQ - Medium import prices / clustered-equivalent hydropower units</li> <li>MedEQ - Medium import prices / equivalent hydropower units</li> <li>OCGT - Open Cycle Gas Turbine</li> <li>PHEV - Plug-In Hybrid Vehicles</li> <li>PS - Pumped Storage</li> <li>w/ - with</li> <li>w/o - without</li> <li>yr - year</li> </ul>

opencc-by-4.0Mar 2023View details →
zenodo36/100

CoCon: A Data Set on Combined Contextualized Research Artifact Use

<p>CoCon is a large graph data set reflecting the combined use of research artifacts, contextualized in academic publications&rsquo; full-text. It comprises 35 k artifacts (data sets, methods, models, and tasks) and 340 k publications.</p> <p>The data set is generated from <a href="https://github.com/paperswithcode/paperswithcode-data">Papers With Code</a> and <a href="https://github.com/IllDepence/unarXive">unarXive</a>.</p> <p>You can find a Python package for loading the data as a NetworkX or Pytorch Geometric graph <a href="https://github.com/IllDepence/contextgraph">in this GitHub repository</a><br> &nbsp;</p>

opencc-by-sa-4.0Mar 2023View details →
zenodo36/100

Data set from a representative survey on artificial intelligence in Germany

<p>Data on the perception of medical artificial intelligence in Germany is currently lacking. Two online surveys were launched in Germany in 2021 to assess the knowledge and perception of artificial intelligence in general and in medicine, including the management of data in medicine. A total of 1,001 and 1,000 adults participated in the surveys. The data collected and the questionnaires will be published.</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

Porcine cell-free system mass spectrometry compiled data sets

<p><span>The degradation of sperm-borne mitochondria after fertilization is a conserved event. This process known as post-fertilization sperm mitophagy, ensures exclusively maternal inheritance of the mitochondria-harbored mitochondrial DNA genome. This mitochondrial degradation is in part carried out by the ubiquitin proteasome system. In mammals, ubiquitin-binding pro-autophagic receptors such as SQSTM1 and GABARAP have also been shown to contribute to sperm mitophagy. These systems work in concert to ensure the timely degradation of the sperm-borne mitochondria after fertilization. We hypothesize that other receptors, cofactors, and substrates are involved in post-fertilization mitophagy. <span>Mass spectrometry was used in conjunction with a porcine cell-free system to identify other autophagic cofactors involved in post-fertilization sperm mitophagy. This porcine cell-free system is able to recapitulate early fertilization proteomic interactions.  Altogether, 185 proteins were identified as statistically different between control and cell-free treated spermatozoa. Six of these proteins were further investigated, including MVP, PSMG2, PSMA3, FUNDC2, SAMM50, and BAG5. These proteins were phenotyped using porcine <em>in vitro </em>fertilization, cell imaging, proteomics, and the porcine cell-free system. The present data confirms the involvement of known mitophagy determinants in the regulation of mitochondrial inheritance and provides a master list of candidate mitophagy co-factors to validate in the future hypothesis-driven studies.</span></span></p>

opencc-zeroMar 2023View details →
zenodo36/100

Content Analysis Data set

<p><strong>Assessment of the Global Healthcare Industry during COVID-19 pandemic: A Content Analysis Approach</strong></p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Code and data for: Setting sustainable limits on anchoring to improve the resilience of coral reefs

<p>This dataset relates to the journal article: Mason, R. A. B., Bozec, Y.-M., Mumby, P. J. (2023)&nbsp;&quot;Setting sustainable limits on anchoring to improve the resilience of coral reefs&quot;&nbsp;<em>Marine Pollution Bulletin</em> 189: 114721 and is comprised of&nbsp;the following components:</p> <p>- The code and data used to perform the modelling that is described in the journal article.&nbsp;</p> <p>- The data outputs&nbsp;that resulted from running the code.</p> <p>- The code used to make figures using these data outputs, which&nbsp;are illustrated&nbsp;in the journal article</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Data sets used in Lee et al. (2023)

<p>Model simulations and processed data sets&nbsp;used in&nbsp;Lee, H., Jung, M., Carvalhais, N., Trautmann, T., Kraft, B., Reichstein, M., Forkel, M., and Koirala, S.: Diagnosing modeling errors of global terrestrial water storage interannual variability, Hydrol. Earth Syst. Sci. Discuss. [preprint], https://doi.org/10.5194/hess-2022-284, in review, 2022.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Historical data set for the Flumendosa case study in Sardinia, Italy.

<p>The files COUT, WTMP, EVAP, CTMP contain the outflow, outflow water temperature, evapotranspiration and corrected air temperature of the sub basins respectively. CCIN, CCON, CCTN, CCTP, CCPP, CCSP, CCSS, CCTS are the concentrations of inorganic ang organic nitrogen; the total nitrogen and phosphorus; &nbsp;the particulate and soluble phosphorus and the suspended and total sediment. CPRC is the corrected precipitation and UPVAP the evapotranspiration from the upstream area of each sub basin.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Data sets for "Updated MS²PIP web server supports cutting-edge proteomics applications"

<p>Data sets and code&nbsp;used to train and evaluate new MS&sup2;PIP models.&nbsp;</p>

opencc-by-4.0Feb 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record