Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
867
datasets available to search
ShareScore release 0.9.0
Dataset results
867 results for “repositories”
BEIRUT: Repository Mining for Defect Prediction
<p>This artifact includes two CSV files.</p> <p>The sample_metrics.csv contains a sample of 153 metrics extracted from project <a href="https://ratis.apache.org/">ratis</a>.</p> <p>The sample_prediction.csv contains a sample of the prediction results from the application of defect prediction to the extracted metrics.</p>
Appendix-A: Online Repositories Available for Text Mining
<p>Appendix A is associated with Chapter 2: Text data and where to find them of the book: Manika Lamba and Margam Madhusudhan (2021) Text Mining for Information Professionals: An Uncharted Territory, SpringerNature.</p>
Parameter estimation data repository - "Spatial discordances between mRNAs and proteins in the intestinal epithelium"
<p>The repository contains data associated with the estimation of protein translation and decay rates in the manuscript "Spatial discordances between mRNAs and proteins in the intestinal epithelium". Specifically, it includes MCMC-chains approximating the posterior parameter distribution and figures showing the model's fit to the data for each gene as well as a summary table of all fit results for two different models, the constant translation-rate model ("constant_rate_model") and the declining translation-rate model ("declining_rate_model") as explained in the manuscript.</p> <p>Code associated with the parameter estimation is available at https://github.com/LiBuchauer/spatial_MP_discordances .</p>
GREED: GitHub Repositories and Descriptions
<p><strong>GREED </strong>is dataset with metadata extracted from GitHub repositories (data collected through GitHub REST API) with Python and Jupyter Notebook as programming languages, more than five stars, and created in 2019 and 2020. In the <strong>final </strong>folder, the repositories descriptions are cleaned and filtered using natural language processing (NLP) techniques. The data is separated by programming language and year.</p> <p>In summary, this dataset contains:</p> <ul> <li><strong>102,358 </strong>repositories (<strong>original</strong>); <strong>58,911 </strong>repositories (<strong>final</strong>)</li> <li><strong>2 </strong>programming languages</li> <li><strong>2 </strong>creation years </li> </ul>
Repository Analytics and Metrics Portal (RAMP) 2018 data
<p>The Repository Analytics and Metrics Portal (RAMP) is a web service that aggregates use and performance use data of institutional repositories. The data are a subset of data from RAMP, the Repository Analytics and Metrics Portal (<a href="http://ramp.montana.edu/">http://rampanalytics.org</a>), consisting of data from all participating repositories for the calendar year 2018. For a description of the data collection, processing, and output methods, please see the "methods" section below. Note that the RAMP data model changed in August, 2018 and two sets of documentation are provided to describe data collection and processing before and after the change.</p>
Repository Analytics and Metrics Portal (RAMP) 2017 data
<p>The Repository Analytics and Metrics Portal (RAMP) is a web service that aggregates use and performance use data of institutional repositories. The data are a subset of data from RAMP, the Repository Analytics and Metrics Portal (<a href="http://ramp.montana.edu/">http://rampanalytics.org</a>), consisting of data from all participating repositories for the calendar year 2017. For a description of the data collection, processing, and output methods, please see the "methods" section below.</p>
Panel: Dataverse Community and CoreTrustSeal: Certifying generalist data repositories
<p>CoreTrustSeal Certification is an important tool that helps researchers and practitioners evaluate the trustworthiness of a dataset and a data repository, yet its certification model can be challenging for generalist repositories to meet. This panel will discuss strategies for meeting certification standards for Trustworthy Data Repositories (TDR) across a variety of repositories with different methods of appraisal, curation, preservation, and organization. The audience will be encouraged to add to the discussion and invited to provide feedback on the concepts presented by the panelists.</p>
24x7 Sessions 2 and 3 - Open Repositories 2021 (OR2021)
<p>This video contains presentations from two 24x7 sessions.</p> <p><strong>Elevating Open Data: Building an Accessible Environment for Data Stewardship in Research Libraries with CADRE</strong><br> <em>Wilkinson, Jaci (1); Wittenberg, Jamie V. (2); Mabry, Patricia L. (3); Pentchev, Valentin (4); Van Rennes, Robert (5); Fridmanski, Ethan (1)</em><br> <em>1: Indiana University Libraries; 2: University Libraries, University of Colorado Boulder; 3: HealthPartners; 4: Indiana University Network Science Institute; 5: Big Ten Academic Alliance</em><br> <br> <strong>Using Research Profiles as a Service (RPaS) to Populate an Institutional Repository</strong><br> <em>Gambill, Agnes</em><br> <em>Appalachian State University, United States of America</em><br> <br> <strong>Swallow Beyond the Repository: Development of a Custom Metadata Management Software</strong><br> <em>Neugebauer, Tomasz; Berrizbeitia, Francisco</em><br> <em>Concordia University, Canada</em><br> <br> <strong>Repository Migration: How it started, How it’s going</strong><br> <em>Homenda, Nicholas; Hardesty, Juliet L.</em><br> <em>Indiana University Libraries, United States of America</em><br> <br> <strong>Increasing the A in OA: How accessibility work in repositories should influence publisher agreements</strong><br> <em>Roosa, Sadie</em><br> <em>MIT Libraries, United States of America</em><br> <br> <strong>We’ve got a Digital Repository: now what do we do with it?</strong><br> <em>Reynolds, Kent Douglas; <em>Parisi,</em> Bianca; <em>Rigg</em>, Cindy</em><br> <em>Niagara College Canada, Canada</em></p> <p><strong>Navigating OA eBook usage data stakeholder interests to facilitate cross-platform data exchange</strong><br> <em>Drummond, Christina</em><br> <em>Educopia Institute, United States of America</em><br> <br> <strong>DataCite Service Providers program: Bringing the power of DOIs to a system near you</strong><br> <em>Krznarich, Liz</em><br> <em>DataCite, United States of America</em><br> <br> <strong>Reopening the Repository: Redeveloping ScholarSphere to support Open Access</strong><br> <em>Coughlin, Daniel; Erickson, Seth; Wead, Adam</em><br> <em>Penn State University, United States of America</em><br> <br> <strong>Sustaining Open Infrastructure Communities: Evaluating Fiscal and Administrative Service Options</strong><br> <em>Greer Klein, Heather (1); Metz, Rosalyn (2)</em><br> <em>1: Samvera, United States of America; 2: Emory University, United States of America</em><br> <br> <strong>An Engaged Campus Repository in Practice: The University of Minnesota’s Institutional Repository and the Pandemic Response</strong><br> <em>Moore, Erik A.; Collins, Valerie M.; Johnston, Lisa R.</em><br> <em>University of Minnesota, United States of America</em><br> <br> <strong>Taking (small) steps towards accessible repository content</strong><br> <em>Maistrovskaya, Mariya</em><br> <em>University of Toronto Libraries, Canada</em><br> </p>
DocMine: A Software Documentation-Related Dataset of 950 GitHub Repositories
<p>DocMine dataset consists of textual information collated from multiple software artifacts, across 950 GitHub Repositories. It also consists of probable percentage contribution of text in each software artifact towards different documentation types in each repository, accompanied by metadata information about the repository such as stargazer count, number of pull requests, commits, issues and other files analyzed.</p>
How are software repositories mined? A systematic literature review of workflows, methodologies, reproducibility, and tools
<p>This is the excel spreadsheet dataset containing our analysis of papers performing mining software repositories research from the conferences ICSE, ESEC/FSE, and MSR from the years 2018 - 2020. The data is broken into columns and can be explained at a high-level as follows:</p> <p>Column Content</p> <p>1 The paper being analyzed</p> <p>2 Does the paper state the data they analyzed is available</p> <p>3 Does the paper perform some sort of data analysis or sampling using data others have compiled in the past</p> <p>4 Does the paper state a timestamp for when they begin their work</p> <p>5 Does the paper state the use of systems pre-built to help with MSR work</p> <p>6 - 18 Forms of sampling researchers may have employed to select their data</p> <p>19 What datasets (if any) were used in the analysis</p> <p>20 What tools (if any) were used in the analysis</p> <p>21 How they performed their data sampling workflow</p> <p>22 How they performed their data filtering workflow</p> <p>23 How they performed their data retrieval workflow</p> <p>24 Did they create any scripts in each of these workflows</p> <p>25 - 33 Did they publish a replication package and what is contained within</p> <p>34 Is the paper describing a tool for research or not</p> <p>35 Short description of the paper read</p> <p>36 A high-level category of the work performed in each paper</p>
The EGFRvIII Transcriptome in glioblastoma - public data repository
<p>Compiled dataset from a large omics EGFRvIII study.</p>
Model data repository of "The role of sediment accretion and buoyancy on subduction dynamics and geometry"
<p>This dataset contains the code and data used in Brizzi et al. (2021): The role of sediment accretion and buoyancy on subduction dynamics and geometry</p>
Readme files in 16,000,000 public GitHub repositories (October 2016)
<p>Format</p> <p>index.csv.gz - CSV comma separated file with 3 columns: <repository name>, <flag>,<readme file name> For example: src-d/go-git,s,README.md</p> <p>The flag is either "s" (readme found) or "r" (readme does not exist on the root directory level). Readme file name may be any from the list:</p> <p>"README.md", "readme.md", "Readme.md", "README.MD", "README.txt", "readme.txt", "Readme.txt", "README.TXT", "README", "readme", "Readme", "README.rst", "readme.rst", "Readme.rst", "README.RST"</p> <p>100 part-r-00xxx files are in "new" Hadoop API format with the following settings:</p> <ol> <li> <p>inputFormatClass is org.apache.hadoop.mapreduce.lib.output.SequenceFileOutputFormat</p> </li> <li> <p>keyClass is org.apache.hadoop.io.Text - repository name</p> </li> <li> <p>valueClass is org.apache.hadoop.io.BytesWritable - gzipped readme file</p> </li> </ol>
Sets of mutually similar public GitHub repositories (October 2016)
<p>The format is JSON, the list of lists. Each list is the group of very similar repositories (Weighted Jaccard Similarity threshold 0.8~0.9).</p>
Programming language keyword frequencies extracted from 16,000,000 public GitHub repositories (October 2016)
<p>Origin</p> <p>16,000,000 repositories on GitHub as of October 2016, classified with github/linguist and parsed with Pygments. Token.Keyword tokens were filtered and MapReduce-d. Fuzzy duplicate repositories were discarded.</p> <p>Some languages, e.g. Haskell, are parsed wrong, resulting in <strong>many</strong> keywords. Still they were not removed since we are not familiar with such languages.</p> <p>Format</p> <p>Triples [language name]\t[keyword]\t[frequency]</p> <p>Tabs and new lines in keywords are escaped as \t and \n respectively.</p>
Repository Analytics and Metrics Portal (RAMP) Production Snapshot Dataset, 2018-11-01
<p>The data are publicly available via Globus: <a href="https://app.globus.org/file-manager?origin_id=a40be90c-f8c2-11e8-9340-0e3d676669f4&origin_path=%2F">https://app.globus.org/file-manager?origin_id=a40be90c-f8c2-11e8-9340-0e3d676669f4&origin_path=%2F</a></p> <p>The data consist of a snapshot of the production RAMP Elasticsearch instance [http://ramp.montana.edu/](http://ramp.montana.edu/). The snapshot was taken on November 1, 2018, and consists of 51 indices (one index each for 50 participating institutional repositories (IR) plus one master index or alias that provides computational access to all indices at once). In addition to the snapshot itself, the published dataset includes documentation describing data collection and processing, separate documentation of the requirements and steps to restore the snapshot to a working instance of Elasticsearch, a CSV file listing participating IR and their corresponding Elasticsearch index names, and a Jupyter Notebook with sample Python code for accessing Elasticsearch.</p> <p>The snapshot ID needed to restore the snapshot to a working index is '2018-11-01.' Please see the included file, 'restore_RAMP_snapshots.pdf' for more info.</p> <p>Because of the large file size, download via high speed network is recommended.</p> <p>RAMP development was funded by the Institute of Museum and Library Services (IMLS) as part of the "Measuring Up" project: IMLS: LG-06-14-0090<br> </p>
ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data (repository for simulated ONT RNA-seq data)
<p>Simulated ONT direct RNA and 1D cDNA sequencing data of varying sequencing depths (0.5 million, 1 million, 3 million, and 5 million simulated reads) used for benchmark evaluations of transcript discovery and quantification in our paper "ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data". All details can be found in the <strong>Materials and Methods</strong> section of the paper. </p> <p><em>HEK293T_DirectRNA.transcriptome_quantification.tsv</em> and <em>HEK293T_DirectRNA.transcriptome_quantification.tsv </em>are tab-separated files containing estimated raw read counts and normalized abundance values (in TPM) of transcripts annotated in GENCODE v34lift37. Transcript quantification was done using NanoSim (version 3.1.0). </p> <p><em>HEK293T_DirectRNA.NanoSim_500k.fastq.gz</em>,<em> </em><em>HEK293T_DirectRNA.NanoSim_1M.fastq.gz</em>, <em>HEK293T_DirectRNA.NanoSim_3M.fastq.gz</em>, and<em> HEK293T_DirectRNA.NanoSim_5M.fastq.gz </em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT direct RNA sequencing reads respectively. </p> <p><em>HEK293T_1DcDNA.NanoSim_500k.fastq.gz</em>,<em> HEK293T_1DcDNA.NanoSim_1M.fastq.gz</em>, <em>HEK293T_1DcDNA.NanoSim_3M.fastq.gz</em>, and<em> HEK293T_1DcDNA.NanoSim_5M.fastq.gz </em>are gzip compressed FASTQ files containing 0.5 million, 1 million, 3 million, and 5 million simulated ONT 1D cDNA sequencing reads respectively. </p>
Data repository of FDTD GPR antenna optimization by Sam Stadler
<p>This is the data repository for the article by Sam Stadler and Jan Igel by the name "Developing realistic FDTD GPR antenna surrogates via full-waveform inversion by means of particle swarm optimization". In this repository, all the measurement data, simulated data, and high-resolution figures are stored for further use.</p>
Data repository for "Phonon-mediated room-temperature quantum Hall transport in graphene"
<p>This is the data presented in the manuscript "Phonon-mediated room-temperature quantum Hall transport in graphene", Nat Commun 14, 318 (2023). https://doi.org/10.1038/s41467-023-35986-3</p>
colab_zirc_dims: full results, datasets, and replication code repository
<p>Repository for results, full datasets (training and grain measurement), and replication code for the colab_zirc_dims manuscript, with additional copy of the "czd large" training dataset. Updated for manuscript resubmission. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.