Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

764

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

764 results for “Usage”

Learn how ShareScore rates datasets ↗
zenodo36/100

OAPEN usage report for open access books funded by the Austrian Science Fund (FWF) 2014 - 2016

<p>For several years now, the Stand-Alone Publications Programme by the Austrian Science Fund (FWF) has been funding the production and simultaneous open-access publication of academic books, see: <a href="http://www.fwf.ac.at/en/research-funding/fwf-programmes/stand-alone-publications/">http://www.fwf.ac.at/en/research-funding/fwf-programmes/stand-alone-publications/</a></p> <p>All FWF-funded books are accessible free of charge in the FWF E-Book Library <a href="https://e-book.fwf.ac.at/">https://e-book.fwf.ac.at/</a> and in the OAPEN Library <a href="http://www.oapen.org/search">http://www.oapen.org/search</a>.  </p> <p>The OAPEN Library’s annual data on book usage for 2014 to 2016 shows that the number of archived open-access books is rising steadily, as is the number of downloads per book.</p>

opencc-by-4.0Mar 2017View details →
zenodo36/100

Data and Scripts for Looking into Pandora's Box: The Content of Sci-Hub and its Usage

<p>The data and scripts used in <em>Looking into Pandora's Box: The Content of Sci-Hub and its Usage.</em> See README.md for details.</p>

opencc-zeroApr 2017View details →
zenodo36/100

Branch Policies in CI/CD: Categorization, Adoption, and Usage - Supplemental for replicability

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo36/100

NANCY SNS-JU Project - Fronthaul network of fixed topology Usage Scenario - Dataset 1

<p>In the context of the NANCY project (https://nancy-project.eu/), this Dataset provides input data for the development of the B-RAN and attacks models for the NANCY framework, to model training and model inference functions. The data collected plays the role of ML algorithm-specific data preparation. The dataset contains time-series, collected transmitting a video content through the Italtel "VTU - video streaming and transcoding application", that can convert audio and video streams from one format to another, at multiple encodings schemes, changing resolution, bitrate, and video parameters. The data collected are related to the observation of some of the resources involved in the Usage Scenario: &ldquo;Fronthaul network of fixed topology &ndash; Direct Connectivity&rdquo;. In the Italtel Italian in-lab testbed, a MEC assisted 5G network scenario with a video streaming application for generating traffic is provided. Two different scenarios were set-up, related to downstream and upstream video flows. The variety of collected features ranges from radio front-end metrics to physical server operating system and network function metrics. The dataset consists of raw network traffic and extracted flow-based data captured in separate files. Each file captured is associated to a 10min video streaming of the &ldquo;Big Buck Bunny&rdquo; video. This video was transmitted on two different bands, N3 and N78, with different resolutions, 480p, 720p, 1080p; both in uplink (UL) and in downlink (DL); the type of protocol monitored is &ldquo;HTTP protocol&rdquo;; in case of N78 band, data related to the resource usage were also captured, for a total of more that 100 data files.</p> <p>The collected dataset is representative resource-intensive video traffic that has the greatest impact on 5G/B5G network planning and provisioning. The video streaming dataset includes data directly measured while watching the video on the mobile devices and data directly measured while generating downstream video stream traversing the gNB (i.e., downstream scenario), and vice versa (i.e., upstream scenario). In each experiment, we fixed the location of the UE and the gNB.</p> <p>The NANCY project has received funding from the Smart Networks and Services Joint Undertaking (SNS JU) under the European Union's Horizon Europe research and innovation programme under Grant Agreement No 101096456.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Resource Usage and Optimization Opportunities in Workflows of GitHub Actions - Artifact

<p>This package contains the data, the code used for collection and the analysis notebooks used to obtain the results presented in the paper: &quot;Resource Usage and Optimization Opportunities in Workflows of GitHub Actions&quot; published at ICSE2024.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Exploring the Automatic Classification of Usage Information in Feedback

<p>Datasets and list of relevant papers for the publication "Exploring the Automatic Classification of Usage Information in Feedback".</p><p>Please read the included readMe for more information on formatting of Json-Files</p>

opencc-by-4.0Nov 2023View details →
zenodo36/100

Estimating Usage Of Open Source Projects - Flutter Telemetry Case Study - MSR' 24

<p>This dataset (CSV) was assembled to support analysis within a case study that will be published in the proceedings of the <a href="https://conf.researchr.org/home/msr-2024">Mining Software Repositories</a> conference (MSR &lsquo;24) April 14-15 2024: <a title="Estimating Usage Of Open Source Projects" href="https://doi.org/10.1145/3643991.3645066">Estimating Usage Of Open Source Projects.</a></p> <p>This case study explored whether publicly available metrics could serve as proxies to estimate usage of an open source project. Using the <a href="https://flutter.dev/">Flutter</a> project as our case study, we collected monthly proxy metrics from GitHub, StackOverflow and Slack to compare with Flutter&rsquo;s monthly active user count over the same time period: January 2018 through February 2021.</p> <p>All metrics correspond to the last day of the month and/or represent aggregate activity in that month. Data from GitHub shows aggregate activity counts across the entire <a href="https://github.com/flutter">Flutter GitHub organization</a> (up to 33 repositories).&nbsp;</p> <p>Our specific metrics include:</p> <div> <table> <tbody> <tr> <td> <p>Source</p> </td> <td> <p>Metric</p> </td> <td> <p>Aggregation method</p> </td> <td> <p>Details</p> </td> </tr> <tr> <td> <p>Flutter</p> </td> <td> <p>Monthly Active Users (MAU)</p> </td> <td> <p>Google internal tooling</p> </td> <td> <p>Flutter users active in the last 30 days, collected on the last day of each month</p> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>PullRequest Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td> <p>As Google governs changes to this code base, we excluded known Google employees in code change related metrics&nbsp;</p> </td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Issues Created in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Issue Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Fork events in month</p> </td> <td> <p><a href="http://gharchive.org">gharchive.org</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>GitHub</p> </td> <td> <p>Fork cumulative count at the end of month</p> </td> <td> <p><a href="http://gharchive.org">gharchive.org</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>StackOverflow</p> </td> <td> <p>Question Authors in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>StackOverflow</p> </td> <td> <p>Questions in month</p> </td> <td> <p><a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a>, hosted by <a href="http://bitergia.com">bitergia.com</a></p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>Slack</p> </td> <td> <p>Claimed (cumulative) members at the end of the month</p> </td> <td> <p>fluttercommunity.slack</p> </td> <td>&nbsp;</td> </tr> <tr> <td> <p>Slack</p> </td> <td> <p>Cumulative messages at the end of the month</p> </td> <td> <p>fluttercommunity.slack</p> </td> <td>&nbsp;</td> </tr> </tbody> </table> </div> <p>Tools:&nbsp;</p> <ul> <li> <p>We used an instance of <a href="https://chaoss.github.io/grimoirelab/">GrimoireLab</a> hosted by <a href="https://bitergia.com/">Bitergia</a> to aggregate GitHub PullRequest Authors, GitHub Issue Authors, GitHub Issues, across all repositories under the Flutter Organization, and StackOverflow Question Authors, and StackOverflow Questions for questions that mention Flutter. We used the <a href="http://github.com/chaoss/grimoirelab-sortinghat">Sorting Hat</a> of feature GrimoireLab to identify Google employees in this sample.</p> </li> </ul> <ul> <li> <p><a href="http://gharchive.org">GHArchive</a> via <a href="https://cloud.google.com/blog/topics/public-datasets/github-on-bigquery-analyze-all-the-open-source-code">BigQuery</a> was used to count GitHub Star and Fork events across all repositories under the Flutter Organization</p> </li> <li> <p>We pulled Slack activity directly from&nbsp;<a href="http://fluttercommunity.slack.com">Flutter's slack channel </a>dashboard</p> </li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Survey on Cloud Computing usage in Montenegrin SMEs, 2017-2023

<p>Data collected among 100 SMEs in Montenegro related to their persepctvies on usage of cloud computing services in their businesses. Comprehesive questionnaire prepared on the bases of European Union Agency for Cybersecurity: Cloud Computing - SME Survey, conducted in 2017 and 2023</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

The Sampling Threat when Mining Generalizable Inter-Library Usage Patterns

<div> <div> <div> <p>Tool support in software engineering often relies on relationships, regularities, patterns, or rules mined from other users&rsquo; code. Examples include approaches to bug prediction, code recommendation, and code autocompletion. Mining is typically performed on samples of code rather than the entirety of available software projects. While sampling is crucial for scaling data analysis, it might influence the generalization of the mined patterns. This paper focuses on sampling software projects filtered for specific libraries and frameworks, and on mining patterns that connect different libraries. We call these inter-library patterns.</p> <p>We observe that limiting the sample to a specific library may hinder the generalization of inter-library patterns, posing a threat to their use or interpretation. Using a simulation and a real case study, we demonstrate this threat for different sampling methods. Our simulation shows that only when sampling for the disjunction of both libraries involved in the implication of a pattern, the implication generalizes well. Additionally, we demonstrate that real empirical data sampled using the GitHub search API does not behave as expected from our simulation. This identifies a potential threat relevant for many studies that use the GitHub search API for studying inter-library patterns.</p> </div> </div> </div>

opencc-by-4.0Oct 2024View details →
zenodo36/100

DiaWUG: Diatopic Word Usage Graphs for Spanish

<p>This data collection contains diatopic Word Usage Graphs (WUGs) for Spanish. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>Note:</p> <ul> <li>The date given for each word use does not correspond to the exact date of the document from which the use was sampled but only to the midpoint of the rough time period covered by the Corpus del Espa&ntilde;ol (~2000-2014).</li> <li>The numbers given as grouping for each word use map to Spanish variants in the following way: 0: Spain (ES), 1: Cuba (CU), 2: Colombia (CO), 3: Argentina (AR), 4: Peru (PE), 6: Venezuela (VE).</li> </ul> <p>Please find more information on the provided data in the paper referenced below.</p> <p>Version: 1.1.2, 11.1.2025. Update description. Update Reference. Normalize filenames. Assign noise uses the cluster label '-1' instead of removing them. Update plots. Additional removal of wrongly copied graphs.</p> <h3>Reference</h3> <p>Gioia Baldissin, Dominik Schlechtweg, Sabine Schulte im Walde. 2022. <a href="https://aclanthology.org/2022.lrec-1.278/">DiaWUG: A Dataset for Diatopic Lexical Semantic Variation in Spanish</a>. Proceedings of the Thirteenth Language Resources and Evaluation Conference.</p>

opencc-by-nd-4.0Sep 2021View details →
zenodo36/100

DWUG ES: Diachronic Word Usage Graphs for Spanish

<p>This data collection contains diachronic Word Usage Graphs (WUGs) for Spanish. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>Please find more information on the provided data in the papers referenced below.</p> <p>The annotation was funded by</p> <ul> <li>ANID FONDECYT grant 11200290, U-Inicia VID Project UI-004/20,</li> <li>ANID - Millennium Science Initiative Program - Code ICN17 002 and</li> <li>SemRel Group (DFG Grants SCHU 2580/1 and SCHU 2580/2).</li> </ul> <p>Version: 4.0.2, 7.1.2025. <strong>Full data</strong>. Quoting issues in uses resolved. Target word and target sentence indices corrected. One corrected context for word 'metro'. Judgments anonymized. Annotator 'gecsa' removed. Issues with special characters in filenames resolved. Additional removal of wrongly copied graphs.</p> <h3>Reference</h3> <p>Frank D. Zamora-Reina, Felipe Bravo-Marquez, Dominik Schlechtweg. 2022. <a href="https://aclanthology.org/2022.lchange-1.16/">LSCDiscovery: A shared task on semantic change discovery and detection in Spanish</a>. In Proceedings of the 3rd International Workshop on Computational Approaches to Historical Language Change. Association for Computational Linguistics.</p> <p>Dominik Schlechtweg, Tejaswi Choppa, Wei Zhao, Michael Roth. 2025. <a href="https://aclanthology.org/2025.comedi-1.4/">The CoMeDi Shared Task: Median Judgment Classification &amp; Mean Disagreement Ranking with Ordinal Word-in-Context Judgments</a>. In Proceedings of the 1st Workshop on Context and Meaning--Navigating Disagreements in NLP Annotations.</p>

opencc-by-nd-4.0Mar 2022View details →
zenodo36/100

Files for training purposes - Cluster usage training session @BIOI2

<p>3 sets of inputs to go with our cluster usage training session @BIOI2:</p> <p>- fastq extract top 1000 from SRR9732589</p> <p>- full-length homologs outputted by a BLAST search with NCBI of human ASF1A protein sequence (<a href="https://www.uniprot.org/uniprotkb/Q9Y294/entry#sequence">Q9Y294</a>)</p> <p>- 5 AlphaFold2 models of yeast Protein transport protein SEC39 (<a href="http://www.uniprot.org/uniprotkb/Q6CWC7/entry#sequences">Q6CWC7</a>) (with simplified names) and its X-ray structure <a href="https://www.rcsb.org/structure/8FTU">8FTU</a></p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

An example of critical thinking (CT) snippets usage at Fourah Bay College, University of Sierra Leone

<p>Recorded students &amp; lecturer WhatsApp group conversation while going through the &#39;CT snippets&#39; which are based on INASP&#39;s online course &#39;Questioning as we learn: An introduction to critical thinking&#39;</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Supplementary materials for User Acceptance Factors of Usage-Based Insurance

<p><strong>Overview</strong></p> <p>This is supplementary material for the paper <em>User Acceptance Factors of Usage-Based Insurance</em>.</p> <p>What is included:</p> <ul> <li>01-survey-questionnaire.pdf - questionnaire used in the user study.</li> <li>02-ubci-video-explanations.mp4 - a video with an explanation about Usage-Based Car Insurance, which we provided to participants in our user study.</li> <li>03-SEM_Iteration-1-standardized.pdf - SEM analysis of the theoretical model, including the standardized coefficient.</li> <li>04-SEM_Iteration-1-unstandardized.pdf - SEM analysis of the theoretical model, including the unstandardized coefficient.</li> <li>05-SEM_Iteration-22-standardized.pdf - SEM analysis of the refined model, including the standardized coefficient.</li> <li>06-SEM_Iteration-22-unstandardized.pdf - SEM analysis of the refined model, including the unstandardized coefficient.</li> <li>07-linear-regression.pdf - interactors / moderators analysis.</li> </ul>

opencc-by-4.0Nov 2021View details →
zenodo36/100

DURel: Diachronic Usage Relatedness

<p>This data collection contains diachronic semantic relatedness judgments for German word usage pairs. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>Please find more information on the provided data in the paper referenced below.</p> <p>See previous versions for additional plots, tables and testsets.</p> <p>Version: 3.0.0, 15.12.2021.</p> <p><strong>Reference</strong></p> <p>Dominik Schlechtweg, Sabine Schulte im Walde, Stefanie Eckmann. 2018. <a href="https://aclanthology.org/N18-2027/">Diachronic Usage Relatedness (DURel): A Framework for the Annotation of Lexical Semantic Change</a>. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT). New Orleans, Louisiana USA.</p>

opencc-by-nd-4.0Apr 2018View details →
zenodo36/100

Data for a mapping study about the usage of MDE in Safety and Security Domain

<p>Data for a mapping study. Contains a listing of papers and information about their content.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

An Empirical Study on the Usage and Availability of Machine Learning Libraries in Open-Source Python Projects - Dataset

<p>This repository contains the dataset of the manuscript:</p> <p>&quot;An Empirical Study on the Usage and Availability of Machine Learning Libraries in Open-Source Python Projects&quot;</p>

opencc-by-4.0Dec 2021View details →
zenodo36/100

RefWUG: Diachronic Reference Word Usage Graphs for German

<p>This data collection contains diachronic Word Usage Graphs (WUGs) for German created with reference use sampling. Find a description of the data format, code to process the data and further datasets on the <a href="https://www.ims.uni-stuttgart.de/data/wugs">WUGsite</a>.</p> <p>Please find more information on the provided data in the paper referenced below.</p> <p>Version: 1.1.0, 15.12.2021.</p> <p><strong>Reference</strong></p> <p>Dominik Schlechtweg and Sabine Schulte im Walde. submitted. Clustering Word Usage Graphs: A Flexible Framework to Measure Changes in Contextual Word Meaning.</p>

opencc-by-nd-4.0Sep 2021View details →
zenodo36/100

Dataset and Questionnaire: Creation and Usage of Acceptance Criteria in Practice

<p>In 2021 we conducted an interview study and supplementary assessed datasets of user stories with acceptance criteria. The questionnaire for the interview study and the assessed open source data are in this upload.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

CAGE and txrevise promoter usage QTL summary statistics

<p>This dataset contains promoter usage QTL summary statistics obtained using <a href="https://github.com/eQTL-Catalogue/qtlmap">QTLmap</a>. QTLmap_summary_CAGE.txt contains summary statistics for CAGE, QTLmap_summary_txrevise-25.txt for regular <a href="https://github.com/kauralasoo/txrevise">txrevise</a> and QTLmap_summary_txrevise-supplemented-20.txt for txrevise with the new generated annotations.</p> <p>This dataset contains the promoter annotations obtained with CAGE as part of the FANTOM5 project, cleaned and reformatted&nbsp;using <a href="https://github.com/andreasvija/cage/blob/master/qtlmap_prep/clean_promoters.R">this code</a> for the purpose of QTL mapping, that were used to obtain the aforementioned summary statistics. They can be found in Cleaned_FANTOM5_promoter_annotations.tsv.</p> <p>&nbsp;</p> <p>QTL summary statistics column descriptions:</p> <p>molecular_trait_object_id - gene</p> <p>molecular_trait_id - top promoter for the gene</p> <p>n_traits - how many promoters were tested for the gene</p> <p>n_variants - how many genetic variants were tested for the promoter</p> <p>variant - top genetic variant for the promoter</p> <p>chromosome - chromosome</p> <p>position - promoter&#39;s top genetic variant start position</p> <p>pvalue - nominal p-value of the association between the promoter and genetic variant</p> <p>beta - the fitted linear model&#39;s regression coefficent</p> <p>p_perm - empirical p-value calculated using permutations</p> <p>p_beta - estimated empirical p-value based on the beta distribution, this is the recommended p-value column</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record