Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

867

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

867 results for “repositories”

Learn how ShareScore rates datasets ↗
zenodo36/100

Top-100 liked repositories from HFH

<p>List of the 100 most liked repositories of HFH used in the &quot;Is Hugging Face Hub ready for Empirical Studies?&quot; paper</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

Supplementary web page for the paper "SEAL: Integrating Program Analysis and Repository Mining"

<p>This is an archive of the supplementary material for the paper &ldquo;SEAL: Integrating Program Analysis and Repository Mining&rdquo; including the website and dataset. The website can also be viewed here: <a href="https://se-sic.github.io/paper-SEAL/">https://se-sic.github.io/paper-SEAL/</a></p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

WorldCereal open global harmonized reference data repository (CC-BY-NC licensed data sets)

<p>Within the <strong>ESA funded</strong> WorldCereal project we have built an open harmonized reference data repository at global extent&nbsp;for model training or product validation&nbsp;in support of land cover and crop type mapping. Data from 2017 onwards were collected from many different sources and then&nbsp;harmonized, annotated and evaluated. These steps are explained in the harmonization protocol (10.5281/zenodo.7584463). This protocol also clarifies the naming convention of the shape files and the WorldCereal attributes&nbsp;(LC, CT, IRR, valtime and sampleID) that were added to the original data sets.</p> <p>This publication&nbsp;includes those harmonized&nbsp;data sets of which the original data set was&nbsp;published under the CC-BY-NC license or a license similar to CC-BY-NC. See document &quot;_In-situ-data-World-Cereal - license - CC-BY-NC.pdf&quot; for an overview of the original data sets. Currently this publication only includes a few small data sets for Tanzania originating from a disease monitoring program of the International Maize and Wheat Improvement Center (CIMMYT). CIMMYT made more data available for&nbsp;countries like Kenya, Ethiopia, Rwanda and, Malawi. However due project contraints these data sets were not yet harmonized.</p>

opencc-by-nc-4.0Dec 2022View details →
zenodo36/100

Repository

<p>The input and&nbsp;output files&nbsp;are separated in the different zip&nbsp;files.</p> <p><strong>Contents:</strong></p> <p>ADN.zip : Docking results and PLIP input, output files for adenosine</p> <p>AMP.zip&nbsp; : Docking results and PLIP input, output files for adenosine monophosphate</p> <p>ADP.zip&nbsp; : Docking results and PLIP input, output files for adenosine 5&rsquo;-diphosphate</p> <p>ATP.zip&nbsp; : Docking results and PLIP input, output files for&nbsp;adenosine 5&rsquo;-triphosphate</p> <p><strong>Notes:</strong>&nbsp;Each folder contains INFO.txt&nbsp;with further detailed information&nbsp;on respective types of data.</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Preparing your thesis for an Open Access Repository

<p>The presenters gave participants&nbsp;a helpful overview prior to submitting their PhD thesis to LSE Theses Online and making it open access. Their presentation answered such questions as:</p> <ul> <li>What do you need to know about using copyright material in your PhD thesis?</li> <li>How does making your PhD thesis available on LSETO benefit you?</li> <li>What are the implications for your publishing plans?</li> </ul>

opencc-by-4.0Nov 2022View details →
zenodo36/100

Raw data for the Github repository "Collection of scripts to download & process hydropower generation data in Argentina, Bolivia, Brazil, Uruguay"

<p>This is the raw data downloaded from the websites of the following system operators:</p> <p>&nbsp; - CNDC, Comit&eacute; Nacional de Despacho de Carga (Bolivia): https://www.cndc.bo/home/index.php<br> &nbsp; - ONS, Operador Nacional do Sistema El&eacute;trico (Brazil): https://www.ons.org.br/<br> &nbsp; - UTE, Usinas y Trasmisiones El&eacute;ctricas (Uruguay): https://www.ute.com.uy/<br> &nbsp; - CAMMESA, Compa&ntilde;&iacute;a Administradora del Mercado Mayorista El&eacute;ctrico Sociedad An&oacute;nima (Argentina): https://cammesaweb.cammesa.com</p> <p>Hydropower generation data is extracted using the R scripts available here: https://github.com/matteodefelice/hydro-sam</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

How to choose a research data repository software? Experience report. Table of requirements.

<p>In the age of digital transformation, scientific and social interest for data and data products is constantly on the rise. The quantity as well as the variety&nbsp;of digital research data is increasing significantly. This raises the question about the governance of this data. For example, how to store the data so that it is presented transparently, freely accessible and subsequently available for re-use in the context of good scientific practice. Research data repositories provide solutions to these issues.</p> <p>Considering the variety of repository software, it is sometimes difficult to identify a fitting solution for a specific use case. For this purpose a detailed analysis of existing software is needed. Presented table of requirements can serve as a starting point and decision-making guide for choosing the most suitable for your purposes repository software.&nbsp;This table is dealing as a supplementary material for the paper &quot;How to choose a research data repository software? Experience report.&quot; (persistent identifier to the paper will be added as soon as paper is published).</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Online Repository for "Sorting lithium-ion battery electrode materials using dielectrophoresis"

<p>Please see the readme file.</p> <p>The matlab script for evaluating the measurements is called &ldquo;Eval_Fluoro.m&rdquo; and can be found in this repository.</p> <p>The excel sheet &ldquo;20221028_photometric_iron.xlsx&rdquo; contaiins the data from the chemical analysis.</p> <p>The manufacturing data for the electrodes is provided in the zip folder: PCB_boards_Giesler.zip and can be uploaded to a manufacturer of choice.</p>

opencc-by-4.0Jan 2023View details →
zenodo36/100

The comparison of the AlphaFold and SwissModel Repository databases

<p>This dataset supplements the code at&nbsp;<a href="https://github.com/aozalevsky/alphafold2_vs_swissmodel">https://github.com/aozalevsky/alphafold2_vs_swissmodel</a> for the comparison of the AlphaFold2 database (<a href="https://alphafold.ebi.ac.uk/">https://alphafold.ebi.ac.uk</a>) with the SwissModel Repository (<a href="https://swissmodel.expasy.org/repository">https://swissmodel.expasy.org/repository</a>). Results of the analysis were published as part of the AlphaFold community review&nbsp;<a href="https://www.nature.com/articles/s41594-022-00849-w">https://www.nature.com/articles/s41594-022-00849-w</a>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

GitHub Repository: Datasets, Experimental Setups, and Code for Exploiting Relations Between Commits

<p>This release contains all previously missing results and notebooks.</p>

openother-openMar 2023View details →
zenodo36/100

Design Patterns for AI-based Systems: A Multivocal Literature Review and Pattern Repository

<p>The data for a multivocal literature review on design patterns for AI-based systems.</p> <ul> <li>mlr-search-and-selection.xlsx: the results from the queried databases and search engines, the inclusion/exclusion process, and the backward and forward snowballing results</li> <li>mlr-results.xlsx: the final set of selected resources, the patterns extracted from them, and some analysis</li> <li>query-strings-google-and-google-scholar.txt: the individual terms of the search query (broken up for Google Scholar and Google Search)</li> </ul>

opencc-by-4.0Jan 2023View details →
zenodo36/100

CES Collections in the UC Open Access Repository: 2021

<p>This is a test file.</p>

opencc-by-4.0Feb 2022View details →
zenodo36/100

Glass-Like Random Catalogues for Two-Point Estimates on the Light Cone (Data Repository)

<p>This is the data repository for the article&nbsp;Glass-Like Random Catalogues for Two-Point Estimates on the Light Cone.</p> <p>arxiv link:&nbsp;https://arxiv.org/abs/2304.02040</p> <p>DOI:</p> <p>&nbsp;</p> <p>The file grlic_data.tar.gz contains three directories: /data, /correlations, and /randoms.</p> <p>Inside the /data folder, the data catalogues used in the article are stored: in &quot;cat_part_1162568&quot; the three columns correspond to the redshift, the cosine of the polar angle, and&nbsp;the azimuthal angle, respectively. The first three columns of the &quot;cat_high_1164853&quot; and &quot;cat_low_1164853&quot; catalogues correspond to the same properties for the high-mass and low-mass halos, respectively. In addition to that, these files also contain additional information about the halos: the number of&nbsp;particles (4th column), their M_200b mass (5th column) and their parent ID (6th colum), which is -1 if the halo is not a subhalo.</p> <p>Inside the /randoms folder, the random catalogues for each of the data catalogues in /data are stored. Files beginning with &quot;part&quot; refer to randoms based on the particle catalogue, in a similar fashion the files beginning with &quot;high&quot; correspond to randoms based on the high-mass halo catalogue and those beginning with &quot;low&quot; refer to the randoms based on the low-mass halo catalogue.&nbsp;Files with &quot;..._glass&lt;x&gt;...&quot; correspond to the glass-like random catalogues, and files with &quot;..._rand&lt;x&gt;...&quot; correspond to the Poisson-sampled randoms, where &lt;x&gt; is the value of&nbsp;<span class="math-tex">\(\alpha\)</span> used. <span class="math-tex">\(\alpha\)</span> is the factor by which the number of objects in the data catalogue is multiplied to get the number of objects in the random catalogues,&nbsp;<span class="math-tex">\(N_R = \alpha N_D\)</span>. For the glass-like randoms based on the high-mass halo catalogue, the additional suffix,&nbsp;&quot;..._deltagrid&lt;y&gt;...&quot;, refers to the number of grid-cells used in the Zeldovich approximation, where &lt;y&gt; is the number of grid cells, and &quot;..._Niter&lt;z&gt;...&quot; refers to the number of Zeldovich iterations perfomed,&nbsp;which is given by&nbsp;&lt;z&gt;. Similarly to the&nbsp;data catalogues, the three columns represent the redshift, cosine of the polar angle, and the azimuthal angle of each object in the catalogue.</p> <p>Inside the /correlations folder, there are two subdirectories: /correlations/full and /correlations/multipoles. The /correlations/full directory&nbsp;contains the raw output from CUTE&nbsp;for each correlated pair of catalogues. For example, the subdirectory /correlations/full/low_low contains the outputs for the low-mass halo autocorrelation, or the subdirectory /correlations/full/high_part contains the outputs for the cross-correlation between the high-mass halo catalogue and the particle catalogue. The naming convention of these files is &quot;&lt;type&gt;&lt;alpha&gt;_&lt;catalogues&gt;_&lt;i&gt;&quot;, where &lt;type&gt; can either be &quot;rand&quot; or &quot;glass&quot;, for either the Poisson-sampled or glass-like random catalogues, &lt;alpha&gt; is the factor&nbsp;<span class="math-tex">\(\alpha\)</span>&nbsp;already introduced above, &lt;catalogues&gt; again describes which two data catalogues have been cross- (or auto-) correlated, e.g. &quot;highlow&quot; refers to a cross-correlation between the high-mass halo catalogue and the low-mass halo catalogue, and finally &lt;i&gt; is a number between 0 and 19&nbsp;for the 20 independent measurements of the correlations.&nbsp;For the high-mass halo catalogue autocorrelations involving glass-like randoms, the additional suffix,&nbsp;&quot;..._deltagrid&lt;y&gt;...&quot;, refers to the number of grid-cells used in the Zeldovich approximation, where &lt;y&gt; is the number of grid cells, and &quot;..._Niter&lt;z&gt;...&quot; refers to the number of Zeldovich iterations perfomed,&nbsp;which is given by&nbsp;&lt;z&gt;. The format of the files is the standard CUTE format for the 3D correlation using binning in (mu,r), i.e. see the readme of&nbsp;https://github.com/damonge/CUTE/tree/master/CUTE, section 5. It reads:</p> <pre>For the 3-D correlation functions the output file has 7 columns with x1 x2 xi(x1,x2) D1D2(x1,x2) D1R2(x1,x2) R1D2(x1,x2) R1R2(x1,x2) where (x1,x2) is either (pi,sigma) or (mu,r). </pre> <p>The /correlations/multipoles subdirectory contains the estimated mean multipoles and their variance derived from the 20 individual correlation measurements for each data catalogue and type of random catalogue. Similarly to the /correlations/full subdirectory, it contains a separate directory for each pair of data catalogue for which the correlation was estimated. The files are named according to &quot;&lt;multipole&gt;_&lt;type&gt;&lt;alpha&gt;_&lt;catalogues&gt;_mean_var&quot;, where &lt;multipole&gt; is either l0 for the monopole, l1 for the dipole, or l2 for the quadrupole. The &lt;type&gt;, &lt;alpha&gt; and &lt;catalogues&gt; are identical to what was described above for the /correlations/full subdirectory.&nbsp;&nbsp;Again, for the high-mass halo catalogue autocorrelations involving glass-like randoms, the additional suffix,&nbsp;&quot;..._deltagrid&lt;y&gt;...&quot;, refers to the number of grid-cells used in the Zeldovich approximation, where &lt;y&gt; is the number of grid cells, and &quot;..._Niter&lt;z&gt;...&quot; refers to the number of Zeldovich iterations perfomed,&nbsp;which is given by&nbsp;&lt;z&gt;. The columns in each of these files are the comoving separation d, the mean multipole&nbsp;<span class="math-tex">\(\xi_l\)</span>, and its variance&nbsp;<span class="math-tex">\(\sigma^2\)</span>of each bin.</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Open metadata of the Institutional Repository (O2) of the UOC (Dataset in English)

<pre>Dataset of the metadata of all the academic and scientific production generated by the university community of the UOC.</pre>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Metadades obertes del Repositori Institucional (O2) de la UOC (Dataset en català)

<p>Dataset de les&nbsp;&nbsp;metadades de tota la producci&oacute; acad&egrave;mica i cient&iacute;fica generada per la comunitat universit&agrave;ria de la UOC.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo36/100

A Compact and Efficient fNIRS Design - Repository

<p>This repository contains the data needed to fully recreate the work performed as part of my Master&#39;s Thesis at Villanova University. The thesis, titled &quot;A Compact and Efficient fNIRS Design&quot;, was submitted to the faculty of the Department of Electrical and Computer Engineering at Villanova University in partial fulfillment of the requirements for the degree of Master of Science in Electrical Engineering in May 2023.</p>

opencc-by-nc-sa-4.0Apr 2023View details →
zenodo36/100

D-DUST Analysis Ready Data Repository

<p>Analysis-ready data repository (<em>D22_ARD_repository_v1_24022022.zip</em>) and Data Management Plan (<em>D-DUST_DMP_v1_22122022.pdf</em>) developed within the D-DUST Project (Data-driven moDelling of particUlate with Satellite Technology aid)</p>

opencc-by-4.0Jul 2022View details →
zenodo36/100

Experimental Repository for "Certified Core-Guided MaxSAT Solving"

<p>Experimental repository for the paper &quot;Certified Core-Guided MaxSAT Solving&quot;</p> <p>&nbsp;</p> <p>Directory structure:</p> <p>- `examples`: Some example MaxSAT instances in WCNF format with proofs.</p> <p>- `plots`: Plots generated from our experiments; also contain the plots used in the paper.</p> <p>- `raw_data`: Raw data from the experiments and scripts to analyze the raw data.</p> <p>- `source_code`: The source code for the certifying version of CGSS (`certified-cgss`), vanilla CGSS with the bugs fixed (`cgss`) and the pseudo-Boolean proof check VeriPB (`VeriPB`) used to run the experiments.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Reproduction package for "Revisiting the reproducibility of empirical software engineering studies based on data retrieved from development repositories"

<p>Reproduction package for "Revisiting the reproducibility of empirical software engineering studies based on data retrieved from development repositories", published in Information and Software Technology, Volume 164, December 2023. DOI: <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.infsof.2023.107318" target="_blank" rel="noreferrer noopener"><span><span>https://doi.org/10.1016/j.infsof.2023.107318</span></span></a></p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

almost 3000 Networks (unweighted, undirected, simple, connected) from Network Repository

<p><strong>Data</strong></p> <p>All networks from <a href="https://networkrepository.com"><code>networkrepository.com</code></a> [1] with at most 1M edges (fall 2020) with the following modifications:</p> <ul> <li>weights and edge directions have been ignored</li> <li>multi-edges and self loops have been removed, i.e., the graphs are simple</li> <li>each graph has been reduced to its largest connected component</li> <li>for isomorphic graphs, only one copy has been kept</li> </ul> <p>[1] Ryan A. Rossi and Nesreen K. Ahmed, <em>The Network Data Repository with Interactive Graph Analytics and Visualization</em> (AAAI 2015)</p> <p><strong>Format</strong></p> <p>The data format is a simple edge list:</p> <ul> <li>each row contains two numbers <em>u</em> and <em>v </em>separated by a space representing an edge <em>{u, v}</em></li> <li>for a graph with <em>n</em> vertices, the numbers range from <em>0 </em>to <em>n - 1</em></li> <li>for each edge <em>{u, v} </em>only one of the pairs <em>u v</em> or <em>v u</em> is present, i.e., if the graph has <em>m</em> edges, the file contains <em>m</em> rows</li> </ul>

opencc-by-4.0Jun 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record