Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
867
datasets available to search
ShareScore release 0.9.0
Dataset results
867 results for “repositories”
Top-100 liked repositories from HFH
<p>List of the 100 most liked repositories of HFH used in the "Is Hugging Face Hub ready for Empirical Studies?" paper</p>
Supplementary web page for the paper "SEAL: Integrating Program Analysis and Repository Mining"
<p>This is an archive of the supplementary material for the paper “SEAL: Integrating Program Analysis and Repository Mining” including the website and dataset. The website can also be viewed here: <a href="https://se-sic.github.io/paper-SEAL/">https://se-sic.github.io/paper-SEAL/</a></p>
WorldCereal open global harmonized reference data repository (CC-BY-NC licensed data sets)
<p>Within the <strong>ESA funded</strong> WorldCereal project we have built an open harmonized reference data repository at global extent for model training or product validation in support of land cover and crop type mapping. Data from 2017 onwards were collected from many different sources and then harmonized, annotated and evaluated. These steps are explained in the harmonization protocol (10.5281/zenodo.7584463). This protocol also clarifies the naming convention of the shape files and the WorldCereal attributes (LC, CT, IRR, valtime and sampleID) that were added to the original data sets.</p> <p>This publication includes those harmonized data sets of which the original data set was published under the CC-BY-NC license or a license similar to CC-BY-NC. See document "_In-situ-data-World-Cereal - license - CC-BY-NC.pdf" for an overview of the original data sets. Currently this publication only includes a few small data sets for Tanzania originating from a disease monitoring program of the International Maize and Wheat Improvement Center (CIMMYT). CIMMYT made more data available for countries like Kenya, Ethiopia, Rwanda and, Malawi. However due project contraints these data sets were not yet harmonized.</p>
Repository
<p>The input and output files are separated in the different zip files.</p> <p><strong>Contents:</strong></p> <p>ADN.zip : Docking results and PLIP input, output files for adenosine</p> <p>AMP.zip : Docking results and PLIP input, output files for adenosine monophosphate</p> <p>ADP.zip : Docking results and PLIP input, output files for adenosine 5’-diphosphate</p> <p>ATP.zip : Docking results and PLIP input, output files for adenosine 5’-triphosphate</p> <p><strong>Notes:</strong> Each folder contains INFO.txt with further detailed information on respective types of data.</p>
Preparing your thesis for an Open Access Repository
<p>The presenters gave participants a helpful overview prior to submitting their PhD thesis to LSE Theses Online and making it open access. Their presentation answered such questions as:</p> <ul> <li>What do you need to know about using copyright material in your PhD thesis?</li> <li>How does making your PhD thesis available on LSETO benefit you?</li> <li>What are the implications for your publishing plans?</li> </ul>
Raw data for the Github repository "Collection of scripts to download & process hydropower generation data in Argentina, Bolivia, Brazil, Uruguay"
<p>This is the raw data downloaded from the websites of the following system operators:</p> <p> - CNDC, Comité Nacional de Despacho de Carga (Bolivia): https://www.cndc.bo/home/index.php<br> - ONS, Operador Nacional do Sistema Elétrico (Brazil): https://www.ons.org.br/<br> - UTE, Usinas y Trasmisiones Eléctricas (Uruguay): https://www.ute.com.uy/<br> - CAMMESA, Compañía Administradora del Mercado Mayorista Eléctrico Sociedad Anónima (Argentina): https://cammesaweb.cammesa.com</p> <p>Hydropower generation data is extracted using the R scripts available here: https://github.com/matteodefelice/hydro-sam</p>
How to choose a research data repository software? Experience report. Table of requirements.
<p>In the age of digital transformation, scientific and social interest for data and data products is constantly on the rise. The quantity as well as the variety of digital research data is increasing significantly. This raises the question about the governance of this data. For example, how to store the data so that it is presented transparently, freely accessible and subsequently available for re-use in the context of good scientific practice. Research data repositories provide solutions to these issues.</p> <p>Considering the variety of repository software, it is sometimes difficult to identify a fitting solution for a specific use case. For this purpose a detailed analysis of existing software is needed. Presented table of requirements can serve as a starting point and decision-making guide for choosing the most suitable for your purposes repository software. This table is dealing as a supplementary material for the paper "How to choose a research data repository software? Experience report." (persistent identifier to the paper will be added as soon as paper is published).</p>
Online Repository for "Sorting lithium-ion battery electrode materials using dielectrophoresis"
<p>Please see the readme file.</p> <p>The matlab script for evaluating the measurements is called “Eval_Fluoro.m” and can be found in this repository.</p> <p>The excel sheet “20221028_photometric_iron.xlsx” contaiins the data from the chemical analysis.</p> <p>The manufacturing data for the electrodes is provided in the zip folder: PCB_boards_Giesler.zip and can be uploaded to a manufacturer of choice.</p>
The comparison of the AlphaFold and SwissModel Repository databases
<p>This dataset supplements the code at <a href="https://github.com/aozalevsky/alphafold2_vs_swissmodel">https://github.com/aozalevsky/alphafold2_vs_swissmodel</a> for the comparison of the AlphaFold2 database (<a href="https://alphafold.ebi.ac.uk/">https://alphafold.ebi.ac.uk</a>) with the SwissModel Repository (<a href="https://swissmodel.expasy.org/repository">https://swissmodel.expasy.org/repository</a>). Results of the analysis were published as part of the AlphaFold community review <a href="https://www.nature.com/articles/s41594-022-00849-w">https://www.nature.com/articles/s41594-022-00849-w</a> </p>
GitHub Repository: Datasets, Experimental Setups, and Code for Exploiting Relations Between Commits
<p>This release contains all previously missing results and notebooks.</p>
Design Patterns for AI-based Systems: A Multivocal Literature Review and Pattern Repository
<p>The data for a multivocal literature review on design patterns for AI-based systems.</p> <ul> <li>mlr-search-and-selection.xlsx: the results from the queried databases and search engines, the inclusion/exclusion process, and the backward and forward snowballing results</li> <li>mlr-results.xlsx: the final set of selected resources, the patterns extracted from them, and some analysis</li> <li>query-strings-google-and-google-scholar.txt: the individual terms of the search query (broken up for Google Scholar and Google Search)</li> </ul>
CES Collections in the UC Open Access Repository: 2021
<p>This is a test file.</p>
Glass-Like Random Catalogues for Two-Point Estimates on the Light Cone (Data Repository)
<p>This is the data repository for the article Glass-Like Random Catalogues for Two-Point Estimates on the Light Cone.</p> <p>arxiv link: https://arxiv.org/abs/2304.02040</p> <p>DOI:</p> <p> </p> <p>The file grlic_data.tar.gz contains three directories: /data, /correlations, and /randoms.</p> <p>Inside the /data folder, the data catalogues used in the article are stored: in "cat_part_1162568" the three columns correspond to the redshift, the cosine of the polar angle, and the azimuthal angle, respectively. The first three columns of the "cat_high_1164853" and "cat_low_1164853" catalogues correspond to the same properties for the high-mass and low-mass halos, respectively. In addition to that, these files also contain additional information about the halos: the number of particles (4th column), their M_200b mass (5th column) and their parent ID (6th colum), which is -1 if the halo is not a subhalo.</p> <p>Inside the /randoms folder, the random catalogues for each of the data catalogues in /data are stored. Files beginning with "part" refer to randoms based on the particle catalogue, in a similar fashion the files beginning with "high" correspond to randoms based on the high-mass halo catalogue and those beginning with "low" refer to the randoms based on the low-mass halo catalogue. Files with "..._glass<x>..." correspond to the glass-like random catalogues, and files with "..._rand<x>..." correspond to the Poisson-sampled randoms, where <x> is the value of <span class="math-tex">\(\alpha\)</span> used. <span class="math-tex">\(\alpha\)</span> is the factor by which the number of objects in the data catalogue is multiplied to get the number of objects in the random catalogues, <span class="math-tex">\(N_R = \alpha N_D\)</span>. For the glass-like randoms based on the high-mass halo catalogue, the additional suffix, "..._deltagrid<y>...", refers to the number of grid-cells used in the Zeldovich approximation, where <y> is the number of grid cells, and "..._Niter<z>..." refers to the number of Zeldovich iterations perfomed, which is given by <z>. Similarly to the data catalogues, the three columns represent the redshift, cosine of the polar angle, and the azimuthal angle of each object in the catalogue.</p> <p>Inside the /correlations folder, there are two subdirectories: /correlations/full and /correlations/multipoles. The /correlations/full directory contains the raw output from CUTE for each correlated pair of catalogues. For example, the subdirectory /correlations/full/low_low contains the outputs for the low-mass halo autocorrelation, or the subdirectory /correlations/full/high_part contains the outputs for the cross-correlation between the high-mass halo catalogue and the particle catalogue. The naming convention of these files is "<type><alpha>_<catalogues>_<i>", where <type> can either be "rand" or "glass", for either the Poisson-sampled or glass-like random catalogues, <alpha> is the factor <span class="math-tex">\(\alpha\)</span> already introduced above, <catalogues> again describes which two data catalogues have been cross- (or auto-) correlated, e.g. "highlow" refers to a cross-correlation between the high-mass halo catalogue and the low-mass halo catalogue, and finally <i> is a number between 0 and 19 for the 20 independent measurements of the correlations. For the high-mass halo catalogue autocorrelations involving glass-like randoms, the additional suffix, "..._deltagrid<y>...", refers to the number of grid-cells used in the Zeldovich approximation, where <y> is the number of grid cells, and "..._Niter<z>..." refers to the number of Zeldovich iterations perfomed, which is given by <z>. The format of the files is the standard CUTE format for the 3D correlation using binning in (mu,r), i.e. see the readme of https://github.com/damonge/CUTE/tree/master/CUTE, section 5. It reads:</p> <pre>For the 3-D correlation functions the output file has 7 columns with x1 x2 xi(x1,x2) D1D2(x1,x2) D1R2(x1,x2) R1D2(x1,x2) R1R2(x1,x2) where (x1,x2) is either (pi,sigma) or (mu,r). </pre> <p>The /correlations/multipoles subdirectory contains the estimated mean multipoles and their variance derived from the 20 individual correlation measurements for each data catalogue and type of random catalogue. Similarly to the /correlations/full subdirectory, it contains a separate directory for each pair of data catalogue for which the correlation was estimated. The files are named according to "<multipole>_<type><alpha>_<catalogues>_mean_var", where <multipole> is either l0 for the monopole, l1 for the dipole, or l2 for the quadrupole. The <type>, <alpha> and <catalogues> are identical to what was described above for the /correlations/full subdirectory. Again, for the high-mass halo catalogue autocorrelations involving glass-like randoms, the additional suffix, "..._deltagrid<y>...", refers to the number of grid-cells used in the Zeldovich approximation, where <y> is the number of grid cells, and "..._Niter<z>..." refers to the number of Zeldovich iterations perfomed, which is given by <z>. The columns in each of these files are the comoving separation d, the mean multipole <span class="math-tex">\(\xi_l\)</span>, and its variance <span class="math-tex">\(\sigma^2\)</span>of each bin.</p>
Open metadata of the Institutional Repository (O2) of the UOC (Dataset in English)
<pre>Dataset of the metadata of all the academic and scientific production generated by the university community of the UOC.</pre>
Metadades obertes del Repositori Institucional (O2) de la UOC (Dataset en català)
<p>Dataset de les metadades de tota la producció acadèmica i científica generada per la comunitat universitària de la UOC. </p>
A Compact and Efficient fNIRS Design - Repository
<p>This repository contains the data needed to fully recreate the work performed as part of my Master's Thesis at Villanova University. The thesis, titled "A Compact and Efficient fNIRS Design", was submitted to the faculty of the Department of Electrical and Computer Engineering at Villanova University in partial fulfillment of the requirements for the degree of Master of Science in Electrical Engineering in May 2023.</p>
D-DUST Analysis Ready Data Repository
<p>Analysis-ready data repository (<em>D22_ARD_repository_v1_24022022.zip</em>) and Data Management Plan (<em>D-DUST_DMP_v1_22122022.pdf</em>) developed within the D-DUST Project (Data-driven moDelling of particUlate with Satellite Technology aid)</p>
Experimental Repository for "Certified Core-Guided MaxSAT Solving"
<p>Experimental repository for the paper "Certified Core-Guided MaxSAT Solving"</p> <p> </p> <p>Directory structure:</p> <p>- `examples`: Some example MaxSAT instances in WCNF format with proofs.</p> <p>- `plots`: Plots generated from our experiments; also contain the plots used in the paper.</p> <p>- `raw_data`: Raw data from the experiments and scripts to analyze the raw data.</p> <p>- `source_code`: The source code for the certifying version of CGSS (`certified-cgss`), vanilla CGSS with the bugs fixed (`cgss`) and the pseudo-Boolean proof check VeriPB (`VeriPB`) used to run the experiments.</p>
Reproduction package for "Revisiting the reproducibility of empirical software engineering studies based on data retrieved from development repositories"
<p>Reproduction package for "Revisiting the reproducibility of empirical software engineering studies based on data retrieved from development repositories", published in Information and Software Technology, Volume 164, December 2023. DOI: <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.infsof.2023.107318" target="_blank" rel="noreferrer noopener"><span><span>https://doi.org/10.1016/j.infsof.2023.107318</span></span></a></p>
almost 3000 Networks (unweighted, undirected, simple, connected) from Network Repository
<p><strong>Data</strong></p> <p>All networks from <a href="https://networkrepository.com"><code>networkrepository.com</code></a> [1] with at most 1M edges (fall 2020) with the following modifications:</p> <ul> <li>weights and edge directions have been ignored</li> <li>multi-edges and self loops have been removed, i.e., the graphs are simple</li> <li>each graph has been reduced to its largest connected component</li> <li>for isomorphic graphs, only one copy has been kept</li> </ul> <p>[1] Ryan A. Rossi and Nesreen K. Ahmed, <em>The Network Data Repository with Interactive Graph Analytics and Visualization</em> (AAAI 2015)</p> <p><strong>Format</strong></p> <p>The data format is a simple edge list:</p> <ul> <li>each row contains two numbers <em>u</em> and <em>v </em>separated by a space representing an edge <em>{u, v}</em></li> <li>for a graph with <em>n</em> vertices, the numbers range from <em>0 </em>to <em>n - 1</em></li> <li>for each edge <em>{u, v} </em>only one of the pairs <em>u v</em> or <em>v u</em> is present, i.e., if the graph has <em>m</em> edges, the file contains <em>m</em> rows</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.