Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
867
datasets available to search
ShareScore release 0.9.0
Dataset results
867 results for “repositories”
Issues of the Eclipse Platform Repository with Resolution set to FIXED
<p>In this dataset we can find all the issues of the Eclipse Platform repository where the Resolution attribute is set to FIXED. <br>The dataset contains information about the ID, Title, Description, StartDate( Date when the issue was asigned to its resolutor) and EndDate (Date when the issue was set to FIXED) of the issues in this repository.<br><br>This file does not include issues with a Duration ( EndDate-StartDate) is less than 5 minutes or issues that do not have its title or description set.</p>
FIGURE 14 in Synopsis of the Snakes of the Philippines A Synthesis of Data from Biodiversity Repositories, Field Studies, and the Literature
FIGURE 14. Ahaetulla prasina preocularis (Zambales Prov., Luzon Id.) (TNHC 62721). Photo © RMB.
UCSCXenaShiny Extra Data Repository
<p>This is the extra data repository for <a href="https://github.com/openbiox/UCSCXenaShiny">UCSCXenaShiny</a> project. All datasets are obtained either from the public database UCSC Xena/GDC portal, or publications. The datasets are cleaned and stored in RData format and most data source can be obtained by `attr(obj, "data_source)"` in R.</p> <p>If you use the data in your academic project, please cite regarding data source and our project paper:</p> <p>Wang, S., Xiong, Y., Zhao, L., Gu, K., Li, Y., Zhao, F., ... & Liu, X. S. (2022). UCSCXenaShiny: an R/CRAN package for interactive analysis of UCSC Xena data. <em>Bioinformatics</em>, <em>38</em>(2), 527-529.</p>
Repository for Jupiter Equatorial Zone Disturbance Data
<p>Antuñano et al. (2018) used Jupiter infrared data captured by 8 different ground-based instruments mounted on the Infrared Telescope Facility (IRTF) on Maunakea and Very Large Telescope (VLT) in Paranal between 1984 and 2017, to investigate a pattern of cloud clearance events at Jupiter's equatorial zone at 5 µm.This repository contains (i) cylindrical maps of Jupiter's 5-µm images from 1999-2000 and 2006-2007 captured by NSFCam and NSFCam2 instruments, respectively, mounted on the IRTF, showing Jupiter's Equatorial Zone disturbance, and (ii) a data file showing the brightness scans from ±7° latitude as a function of time used to build Figure 2.</p> <p>DATA</p> <p>As descibed in Antuñano et al. (2018), the normalised average brightness from 1984-2017 was used to charcterise three different cloud clearance events of Jupiter's equatorial zone at 5 µm. The normalised average brightness of the equatorial zone between ±7° latitude is given in "Brightness_scans_1984-2017.dat", where the first column represents the Julian date, the second column represents the latitude in planetocentric degrees, the third column represents the normalised average brightness and the fourth column gives the standard deviation of the normalised average brightness.</p>
Commits modifications from 61 projects hosted in Github's repositories
<p>This dataset contains data from 61 opensource projects that can be found in github repositories.The data contains informations about:</p> <ul> <li>committer name</li> <li>committer email</li> <li>modification type - changed, removed or created</li> <li>modified files</li> <li>commit comments</li> <li>commit date</li> </ul> <p> </p> <p>This dataset references 299 major version software, developed in five different programming languages. These languages are:</p> <ul> <li>C++</li> <li>Java</li> <li>Javascript</li> <li>Python</li> <li>Ruby</li> </ul> <p>The dataset was used in my undergraduate thesis</p>
Unix History Repository
<p>The history and evolution of the Unix operating system is made available as a revision management repository, covering the period from its inception in 1970 as a 2.5 thousand line kernel and 26 commands, to 2018 as a widely-used 30 million line system. The 1.5GB repository contains about half a million commits and more than two thousand merges. The repository employs Git system for its storage and is hosted on GitHub. It has been created by synthesizing with custom software 24 snapshots of systems developed at Bell Labs, the University of California at Berkeley, and the 386BSD team, two legacy repositories, and the modern repository of the open source FreeBSD system. In total, about one thousand individual contributors are identified, the early ones through primary research. The data set can be used for empirical research in software engineering, information systems, and software archaeology.</p>
Data showcase papers published in the Mining Software Repositories (MSR) conference
<p>Data regarding data showcase papers published in the Mining Software Repositories (MSR) conference.</p> <p>The following data files are included.</p> <p>citation-table.csv: SWEBOK areas of citing studies<br> citations.bib: Bibliographic details of citing studies<br> citing_dp_dois_citations.txt: Citations of citing studies<br> data_papers.bib: MSR data papers<br> dp_dois_citations.txt: Citations of data papers<br> false_citations.bib: Citing studies that don't actuall use data papers<br> msr-all: Bibliographic details of all MSR papers<br> ndp_dois_citations.txt: Citations of non-data papers<br> ndp_rand_dois_citations.txt: Citations of a randomly chose non-data paper weighted sample<br> self-citations.csv: Data papers citations by their authors</p> <p> </p>
Data repository for the paper titled: "The intriguing co-distribution of the copepods Calanus hyperboreus and Calanus glacialis in the subsurface chlorophyll maximum of Arctic seas"
<p>Data and R-code for publication: “The intriguing co-distribution of the copepods <em>Calanus hyperboreus </em>and <em>Calanus glacialis </em>in the subsurface chlorophyll maximum of Arctic seas”.</p> <p> </p> <p>By Moritz S Schmid and Louis Fortier</p> <p> </p> <p>Attached are R data files of copepod lipids, copepod vertical distributions, and chl <em>a </em>profiles, as well as R-code.</p> <p>The following sea ice data was used: Nimbus-7 SMMR and DMSP SSM/I-SSMIS passive microwave data, available at : https://nsidc.org/data/nsidc-0051.</p> <p>The following ocean color data was used: MODerate resolution Imaging Spectroradiometer (MODIS), Aqua satellite, available here: https://oceandata.sci.gsfc.nasa.gov/MODIS-Aqua</p> <p> </p> <p>Regards</p>
The Effects of Cycle and Treadmill Desks on Work Performance and Cognitive Function in Sedentary Workers: data repository of a review and meta-analysis.
<p>This repository contains additional files related to the review and meta-analysis. The first dataset contains list of search terms. The second dataset contains list od studies included in the meta-analysis. The third dataset contains study evaluation using PEDro scale tool. The fourth dataset contains two additional forrest plots (Effect of cycle and treadmill desks on typing errors and Effect of cycle and treadmill desks on congruent Eriksen Flanker test).</p>
Data repository of the paper "The First Terrestrial Electron Beam Observed by The Atmosphere-Space Interactions Monitor" by D. Sarria et al.
<p>Data repository / Supporting information of the paper "The First Terrestrial Electron Beam Observed by The Atmosphere-Space Interactions Monitor" (2019) by D. Sarria et al.</p> <p>Access to the article: <a href="https://doi.org/10.1029/2019JA027071">https://doi.org/10.1029/2019JA027071</a></p>
Data repository - Spatial reconstruction of single enterocytes uncovers broad zonation along the intestinal villus axis
<p>Data associated with the manuscript entitled "Spatial reconstruction of single enterocytes uncovers broad zonation along the intestinal villus axis".</p> <p>Files:</p> <p>table_A_LCM_TPM_values.tsv: Gene expression levels of microdissected villus quintiles. First column is the ensemble gene id. Next 15 columns are the raw Kallisto TPM values for villus segments 1 (bottom) to 5 (top) for three different mice (a to c). Additional columns include the external gene name, description, and gene biotype.</p> <p>table_B_scRNAseq_UMI_counts.tsv: Raw UMI counts of cells that were utilized in this study. Analysis is based on raw data from the NCBI GEO datasets GSM2644349 and GSM2644350. Each of the columns represents a single cell, column headers are the corresponding cell barcodes and enable retrieval of tSNE coordinates from table_C_scRNAseq_tsne_coordinates_zones.tsv. Values represent raw UMI counts.</p> <p>table_C_scRNAseq_tsne_coordinates_zones.tsv: tSNE coordinates and reconstructed zones of cells that were utilized in this study. Tab separated text file. Analysis is based on raw data from the NCBI GEO datasets GSM2644349 and GSM2644350. Columns: cell_id: cell barcode, corresponds to column header of Table S2. Seurat tSNE coordinate 1 and tSNE coordinate 2. Last column is the inferred zone (Crypt, V1..V6).</p> <p>table_D_zonation_reconstruction.tsv: Zonation table of reconstructed scRNAseq data. Tab separated text file. Columns: Gene names: gene id, mean expression in each of the crypt zone and 6 villus zones, standard error of the means in the Crypt zone and 6 villus zones, p-value and q-value for the zonation profiles.</p> <p>raw_data.zip: The raw and intermediary data for runnning the scripts in <a href="https://github.com/aemoor/Code_spatial_reconstruction_enterocytes">https://github.com/aemoor/Code_spatial_reconstruction_enterocytes</a></p>
Macro-scale analysis of biodiversity-ecosystem functioning relationships in lakes (github repository)
<p>National Lakes Assessment biodiversity-ecosystem functioning relationships across the continental United States. This is an archived version of a github repository, containing data and R code. The repository can also be found online https://github.com/cont-limno/NLA-Diversity-.</p>
Data Usage Metrics at Repositories: A Survey
<p>Results of a survey undertaken by the Research Data Alliance Data (RDA) Usage Metrics Working Group during February and March 2019 and presented at the 13th RDA Plenary Meeting in Philadelphia on 3 April 2019.</p>
Repository of speech features from speakers with and without Parkinson's Disease. Neurovoz - Rasta PLP - V2 - Scientific Reports Publication: Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease
<p>This repository contains the Rasta-PLP features of six different speech recordings (sentences) from Neurovoz corpus (47 parkinsonian and 32 control speakers whose mother tongue is Spanish Castillian.)<br> Number of PLP coefficients: [6, 8, 10, 12, 14, 16, 18, 20].<br> Delta coefficients: Yes<br> Delta Delta coefficients: Yes<br> Sampling rate: 16 kHz<br> Frame size: 15 ms<br> Frame overlapping: 50%</p> <p>This subset of the Neurovoz corpus was recorded between 2015 and 2017 by Universidad Politécncia de Madrid and Hospital General Universitario Gregorio Marañón.</p> <p>This version includes the same files as the previous version and information about UPDRS, H&Y, years since diagnosis and age of each participant.</p> <p>The sentences were:</p> <p>BARBAS: "Cuando las barbas de tu vecino veas pelar, pon las tuyas a remojar"</p> <p>CALLE: "De la calle vendrá quien de tu casa te echará"</p> <p>DIABLO: " Cuando el diablo no sabe qué hacer, con el rabo mata moscas "</p> <p>PETACA BLANCA: " La petaca blanca es mía"</p> <p>PIDIO: "No pidas a quien pidió ni sirvas a quien sirvió"</p> <p>SOMBRA: " El que a buen árbol se arrima, buena sombra le cobija "</p> <p> </p> <p>How to cite:<br> [1] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Grandas-Perez, F., Shattuck-Hufnagel, S. Yagüe-Jimenez, V., and Dehak, N. (2019). Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson’s disease.Scientific reports 9, 19066.</p> <p><br> [2] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Villalba, J., Rusz, J., Shattuck-Hufnagel, S. and Dehak, N. (2019). A forced Gaussians based methodology for the differential evaluation of Parkinson's Disease by means of speech processing. Biomedical Signal Processing and Control, 48, 205-220.</p> <p>BibTeX:</p> <pre><code>@article{moro2019phonetic, title={Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge A. and Godino-Llorente, Juan I. and Grandas-Perez, Francisco and Shattuck-Hufnagel, Stefanie and Yague-Jimenez, Virginia and Dehak, Najim}, journal={Scientific Reports}, volume={9}, pages={19066}, year={2019}, publisher={Nature Research Publishing} } @article{moro2019forced, title={A forced Gaussians based methodology for the differential evaluation of Parkinson's Disease by means of speech processing}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge Andres and Godino-Llorente, Juan Ignacio and Dehak, Najim}, journal={Biomedical Signal Processing and Control}, pages={205--220}, volume={48}, year={2019}, publisher={Elsevier} } </code></pre> <p> </p>
Data repository for manuscript "A new approach to Health Benefits Package design: an application of the Thanzi La Onse model in Malawi"
<p>Dataset to accompany the publication <em>“A new approach to Health Benefits Package design: an application of the Thanzi La Onse model in Malawi”</em> by Margherita Molaro, Sakshi Mohan, Bingling She, Martin Chalkley, Tim Colbourn, Joseph H. Collins, Emilia Connolly, Matthew M. Graham, Eva Janoušková, Ines Li Lin, Gerald Manthalu, Emmanuel Mnjowe, Dominic Nkhoma, Pakwanja D. Twea, Andrew N. Phillips, Paul Revill, Asif U. Tamuri, Joseph Mfutso-Bengo, Tara Mangal, and Timothy B. Hallett.</p> <p>The Thanzi La Onse (TLO) model used to produce this data is open source and available for review and usage at<a href="https://github.com/UCL/TLOmodel"> https://github.com/UCL/TLOmodel</a>. In particular, the outputs analysed in this study can be reproduced from model tag "Molaro_et_al_2024_HBP_design" (accessible at https://github.com/UCL/TLOmodel/tags) using the scenario file src/scripts/healthsystem/impact_of_policy/scenario_impact_of_policy.py. All analysis scripts used to generate the plots in the manuscript are located in the same directory and have filenames beginning with "analysis_impact_of_policy_".</p> <p>This repository contains post-processed simulation outputs, which were generated using the script src/scripts/healthsystem/impact_of_policy/analysis_extract_data.py (available from the same tag). The data included have the following structure:</p> <p>"Draw": Represents a specific prioritisation-policy, identified by the acronyms listed in Table 1 of the publication.</p> <p>"Run": Represents a single simulation instance of a draw. Each draw was simulated 10 times, each with independent random sampling, resulting in 10 "runs" per draw.</p> <p>The data files included in this repository are:</p> <p><strong>DALYS_by_cause_with_time.csv</strong>: DALYs (as defined in the publication) incurred on a given year due to each of the causes of DALYs considered.</p> <p><strong>HSIs_requested_by_type_and_facility_level_with_time.csv</strong>: total number of requested HSIs on a given year, broken down by HSI type and the facility level at which they were requested.</p> <p><strong>HSIs_delivered_by_type_and_facility_level_with_time.csv</strong>:total number of HSIs delivered on a given year broken down by HSI type and the facility level at which they were delivered.</p> <p><strong>Population_with_time.csv</strong>:total population size on a given year. </p> <p> </p> <p> </p>
Data and code repository for the research "Assessing the use of Airborne Electromagnetic Data for nitrate vulnerability assessment in the Central Valley, California"
<p>Data and code for "Assessing the use of Airborne Electromagnetic Data for nitrate vulnerability assessment in the Central Valley, California". This repository contains the analysis and post-processing code for generating results and figures used in the manuscript.</p>
Repository for the Method Article "Oriented artificial nanofibers and laser induced periodic surface structures as substrates for Schwann cells alignment"
<p>Repository containg the underlaying and extended data for the paper "Oriented artificial nanofibers and laser induced periodic surface structures as substrates for Schwann cells alignment".</p>
Data and Analysis Files Repository: Repurposing Large-Format Microarrays for Scalable Spatial Transcriptomics
<p>Data and Analysis Files from "Repurposing Large-Format Microarrays for Scalable Spatial Transcriptomics"</p> <p>ArraySeq_Method.zip contains the following folder and contents:</p> <ul> <li>STARSolo: All code and count matrix output from fastq spatial barcode demultiplexing. </li> <li>Images: All resolution-downsampled H&E image scans from analyzed tissues</li> <li>Space_Ranger: All 10x Space Ranger output from Visium datasets generated in the paper. </li> <li>Analysis: All scripts for analyzing and plotting Array-seq and Visium datasets generated in this paper. Also contains output h5ad files. </li> </ul> <p>ArraySeq_Barcode_generation_n12.rmd: The script used to generate the Array-seq probes with 12-mer spatial barcodes. </p>
Repository for "Light-induced cortical excitability reveals programmable shape dynamics in starfish oocytes"
<p>Data and code repository for paper "Light-induced cortical excitability reveals programmable shape dynamics in starfish oocytes". DOI tbd.</p>
Repository contents of 5 GitHub Collections
<p><a href="https://github.com/collections" target="_blank" rel="noopener">Github Collections</a> are a section within the GitHub platform where repositories are collected and organized into thematic collections curated by GitHub.</p> <p>These collections group popular or prominent projects within specific areas such as Machine Learning, Web Development, Data Science, Computer Security, and many other categories of technological interest. The goal is to provide users with repository data that allows to analyze different metrics, such as repository size, filetype used or popularity, among others.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.