Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
16
datasets available to search
ShareScore release 0.9.0
Dataset results
16 results for “Open Licenses”
Prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, number of journals per country and publisher
<p>An analysis on the prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, country and publisher according to the number of journals.</p>
Numbers and shares of Open Access Journals in Sociology using Creative Commons Licenses, June 2014
<p>a) The data for the year 2014 was retrieved as a CSV-file (doaj_2014-05-07_1330_utf8.csv) from the Directory of Open Access Journals (DOAJ) homepage at 2014-06-08.<br> b) A subset of journals assigned to the subject category Sociology was generated (n=109).<br> c) I manually checked the information on CC-licenses for each of the 109 journals<br> d) Where necessary I added correct information on licenses, see column k in the CSV-file for the updated information. Column l marks entries that were updated.</p> <p> </p>
Numbers and shares of Open Access Journals from all disciplines and from the discipline Sociology using Creative Commons Licenses as listed by the Directory of Open Access Journals (2014-06-08)
<p>These files contains the data on frequencies and shares of Open Access journals using Creative Commons Licenses (comparing journals from all disciplines and sociologial journals). Date of data collection: 2014-06-08.</p>
Dataset for Earth Sciences at Freie Universität Berlin: Open Access, Licenses and Persistent Identifiers Monitoring
<p>In <em>Version 4</em>, <strong>publishers </strong>and <strong>journals</strong> names has been extended.</p> <p>In <em>Version 3</em>, new entries have been added for both <strong>journal </strong>and <strong>non-journal article outputs</strong>, specifically including data from the year <strong>2023</strong>. Minor adjustments were also made to URLs and open access (OA) statuses.</p> <p><em>Note</em>: Data for journal and non-journal article outputs from the year 2023 were unavailable at the time of preparing the <strong>short paper</strong> presenting the results, findable under <a href="https://doi.org/10.5281/zenodo.14170751" target="_blank" rel="noopener">10.5281/zenodo.14170751</a> [1]).</p> <p><br>Started in 2021, Berlin University Alliance (BUA) Open Science Dashboards, followed by the BUA Open Science Magnifiers projects, seek to investigate Open Science (OS) practices across different research domains and communities. A primary focus of these initiatives lies in the development of OS indicators, tailored to discipline specific ones, alongside their visualisation for monitoring.</p> <p>Collaborating closely with the Department of Earth Sciences at Freie Universität Berlin (FU), one of the project's key objectives is the implementation of an Open Science Dashboard for Earth Sciences FU. The visualisation of the first OS metrics is already available under <a href="https://quest-open-earthsciences.charite.de/">https://quest-open-earthsciences.charite.de/</a>.</p> <p>The datasets utilized include the outputs from the Department of Earth Sciences at FU, i.a. on Open Access (OA) categorisations and statuses, persistent identifiers (PIDs) and Open Licences (Creative Commons) availability, published between 2016-2023. These datasets consist of (i) <strong>"journal_articles_v3.csv"</strong> and (ii) <strong>"non_journal_articles_outputs_v3.csv"</strong>, the latter including “book”, “book chapter”, “conference paper”, “conference abstract”, and “other research outputs” (e.g. book reviews, project reports, book chapters in school books, or electronic supplementary material).</p> <p>Data for the dashboard was obtained from the FU university bibliography (<a href="https://frub-berlin.primo.exlibrisgroup.com/">https://frub-berlin.primo.exlibrisgroup.com/</a>), but coverage of PID information was incomplete, OA category information was incomplete and often erroneous, and copyright/open licence information was missing in this data set. Therefore, the data set was <strong>enriched with manually researched information</strong>. Data enrichment was different for journal articles and for non-journal-article publications. For <strong><em>journal articles</em></strong>, <em>copyright/open licence</em> information was added, and <em>open access category</em> information was checked and added or corrected. For <strong><em>non-journal-article outputs</em></strong>, missing <em>PIDs</em> were added and <em>open access category</em> information was checked and added or corrected. </p> <p>The "<em>data_dictionary_earth_sciences_v3.csv"</em> table documents all variables of each data file containing here.</p> <p>Both for the dashboard, and in our following publications, we categorized <strong>"bronze"</strong> OA outputs as closed access. Although such publications are openly available on the publisher's websites, they lack licence information and thus cannot be openly reused, and presumably even change its openness status at any time. Following the methodology of Charité Dashboard on Responsible Research (<a href="https://quest-dashboard.charite.de/#tabStart">https://quest-dashboard.charite.de/#tabStart</a>) we only include "gold", "hybrid" and "green" OA as true OA. Further details about the enrichment process conducted on these datasets can be found under <a href="https://doi.org/10.5281/zenodo.1099821" target="_blank" rel="noopener">10.5281/zenodo.1099821</a>9 [2]</p> <p> </p> <p>[1] Duine, M., Iarkaeva, A., & Hübner, A. (2024, November 15). Initiating discipline-specific Open Science Monitoring with the Open Science Dashboard for Earth Sciences. 28th International Conference on Science, Technology and Innovation Indicators (STI2024), Berlin, Germany. <a href="https://doi.org/10.5281/zenodo.14170751" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.14170751</a><br>[2] Duine, M., Hübner, A., & Iarkaeva, A. (2024). Enrichment of university bibliography data for open science monitoring. Zenodo. <a href="https://doi.org/10.5281/zenodo.10998219">https://doi.org/10.5281/zenodo.10998219</a></p>
A Large-scale Dataset of (Open Source) License Text Variants
<p>We introduce a large-scale dataset of the complete texts of free/open source software (FOSS) license variants. To assemble it we have collected from the Software Heritage archive—the largest publicly available archive of FOSS source code with accompanying development history—all versions of files whose names are commonly used to convey licensing terms to software users and developers.<br> The dataset consists of 6.5 million unique license files that can be used to conduct empirical studies on open source licensing, training of automated license classifiers, natural language processing (NLP) analyses of legal texts, as well as historical and phylogenetic studies on FOSS licensing.<br> Additional metadata about shipped license files are also provided, making the dataset ready to use in various contexts; they include: file length measures, detected MIME type, detected SPDX license (using ScanCode), example origin (e.g., GitHub repository), oldest public commit in which the license appeared.<br> The dataset is released as open data as an archive file containing all deduplicated license blobs, plus several portable CSV files for metadata, referencing blobs via cryptographic checksums.</p> <p>For more details see the included README file and companion paper:</p> <ul> <li>Stefano Zacchiroli. <a href="https://doi.org/10.1145/3524842.3528491"><em>A Large-scale Dataset of (Open Source) License Text Variants</em></a>. In proceedings of the <a href="https://conf.researchr.org/home/msr-2022">2022 Mining Software Repositories Conference (MSR 2022)</a>. 23-24 May 2022 Pittsburgh, Pennsylvania, United States. ACM 2022.</li> </ul> <p>If you use this dataset for research purposes, please acknowledge its use by citing the above paper.</p> <ul> </ul>
Support of Creative Commons Licenses in German disciplinary and institutional open access repositories
<p>Which of the following open content licenses can be chosen for the metadata description of the open access full-texts (apart from the deposit license)? n=81* </p> <p>* This survey question is part of the 2014 Census of Open Access Repositories in Germany, Austria and Switzerland, see: http://nbn-resolving.de/urn:nbn:de:kobv:11-100222687 For the research data see: http://doi.org/10.5281/zenodo.10734 </p>
Total numbers and shares of Open Access Journals using Creative Commons Licenses as listed by the Directory of Open Access Journals
<p>This table provides information on the number and percentage of Open Access Journals listed by the Directory of Open Access Journals using a Creative Commons License. It identifies also the number and share of Journals using CC-Licenses that are compatible to the Open Definition.</p>
MJFF Data Community - Creative Commons Training: Copyright and Open Licensing
<p>This training, provided by Shanna Hollich, the Learning and Training Manager of Creative Commons (CC), was hosted by the Michael J. Fox Foundation's Data Community of Practice (DCOP). For more information on the DCOP, please contact: researchcommunity@michaeljfox.org.</p> <p>In the ever-evolving digital landscape, the management of research outputs, including data licensing and copyright, is of utmost importance. This 1.5-hour training provided a forum for participants to learn more about open licensing and copyright. It also aimed to equip participants with the knowledge and best practices they need to effectively navigate the complexities of CC licensing and copyright when using and generating research outputs such as scholarly publications, datasets, and white papers.<br><br>By the end of the workshop, participants developed an understanding of the basic principles of copyright, how it works, and where it applies. Participants are now able to describe the benefits of open licensing and the basics of how Creative Commons licenses work, have a deeper understanding of how research outputs interact with copyright and open licensing, and know where to find additional information and resources.</p> <p>To access a stream of this video with variable resolution, please <a href="https://share.vidyard.com/watch/5vMMsxsZK6q48yNHsDTkPe" target="_blank" rel="noopener">visit this link</a>.</p>
Numbers of Articles, Books and Dissertation theses indexed in BASE and percentages of items published Open Access, under Creative Commons Licenses and under Open Licenses (2013-2017)
<p>A look at the data provided by the Open Access search engine BASE (http://base-search.net) shows that the Open Science compliance among dissertation theses stagnates. BASE knows three categories of accessibility: Open Access, Unknown, Non-Open Access. In the following tables and graphs, figures reported as "Open Access" have been categorised by BASE as Open Access. The tables and graphics show data from BASE (as of 06.03.2018) as follows:</p> <p>a) Indexed theses, books and journal articles</p> <p>b) Indexed theses, books and journal articles published by Open Access</p> <p>c) indexed theses, books and journal articles under Creative Commons licenses.</p> <p>d) indexed theses, books and journal articles, which are published under open licenses in the sense of the Open License, i. e. reflect terms of use of the Open Source.</p> <p> </p> <p>Although doctoral theses already had a high share of open access by 2013 (43%), by 2017 it had risen by only 5% (2017:48%). At the same time, the proportion of books published in open access rose by 14% (from 20% to 34%) and articles by 17% from 44% (2013) to 61% (2017). The same effect can be seen in the proportion of CC-licensed items: Their share rose by 4% (from 9% to 13%) for doctoral theses, by 9% for books (from 4% to 13%) and 8% for articles (from 10% to 18%) between 2013 and 2017. However, the share of openly licensed items is most pronounced: it did not increase for doctoral theses, but remained at 2% between 2013 and 2017; in the same period it increased by 5% (from 1% to 6%) for books, and by 5% (from 5% to 10%) for articles.</p>
WorldCereal open global harmonized reference data repository (CC-BY-SA licensed data sets)
<p>Within the<strong> ESA funded</strong> WorldCereal project we have built an open harmonized reference data repository at global extent for model training or product validation in support of land cover and crop type mapping. Data from 2017 onwards were collected from many different sources and then harmonized, annotated and evaluated. These steps are explained in the harmonization protocol (10.5281/zenodo.7584463). This protocol also clarifies the naming convention of the shape files and the WorldCereal attributes (LC, CT, IRR, valtime and sampleID) that were added to the original data sets.</p> <p>This publication includes those harmonized data sets of which the original data set was published under the CC-BY-SA license or a license similar to CC-BY-SA. See document "_In-situ-data-World-Cereal - license - CC-BY-SA.pdf" for an overview of the original data sets.</p>
WorldCereal open global harmonized reference data repository (CC-BY licensed data sets)
<p>Within the <strong>ESA funded </strong>WorldCereal project we have built an open harmonized reference data repository at global extent for model training or product validation in support of land cover and crop type mapping. Data from 2017 onwards were collected from many different sources and then harmonized, annotated and evaluated. These steps are explained in the harmonization protocol (10.5281/zenodo.7584463). This protocol also clarifies the naming convention of the shape files and the WorldCereal attributes (LC, CT, IRR, valtime and sampleID) that were added to the original data sets.</p> <p>This publication includes those harmonized data sets of which the original data set was published under the CC-BY license or a license similar to CC-BY. See document "_In-situ-data-World-Cereal - license - CC-BY.pdf" for an overview of the original data sets. </p>
Replication Package for "Ensuring Open Source Integrity: The Intersection of Copy-Based Reuse and License Compliance"
<p>Replication Package for "Ensuring Open Source Integrity: The Intersection of Copy-Based Reuse and License Compliance"<br><br>Includes datasets, R and bash code.</p>
WorldCereal open global harmonized reference data repository (CC-BY-NC licensed data sets)
<p>Within the <strong>ESA funded</strong> WorldCereal project we have built an open harmonized reference data repository at global extent for model training or product validation in support of land cover and crop type mapping. Data from 2017 onwards were collected from many different sources and then harmonized, annotated and evaluated. These steps are explained in the harmonization protocol (10.5281/zenodo.7584463). This protocol also clarifies the naming convention of the shape files and the WorldCereal attributes (LC, CT, IRR, valtime and sampleID) that were added to the original data sets.</p> <p>This publication includes those harmonized data sets of which the original data set was published under the CC-BY-NC license or a license similar to CC-BY-NC. See document "_In-situ-data-World-Cereal - license - CC-BY-NC.pdf" for an overview of the original data sets. Currently this publication only includes a few small data sets for Tanzania originating from a disease monitoring program of the International Maize and Wheat Improvement Center (CIMMYT). CIMMYT made more data available for countries like Kenya, Ethiopia, Rwanda and, Malawi. However due project contraints these data sets were not yet harmonized.</p>
BioFlow-Insight Workflow Corpus Open License
<p>This corpus describes a collection of open license Nextflow workflows which we're gathered from Github in February 2024.</p>
Dataset: Open Access: An Analysis of Publisher Copyright and Licensing Policies in Europe, 2020
<p>This dataset supports the SPARC Europe study that investigates the copyright retention policy amongst publishers, self-archiving policies and records publisher policies on open licensing, also as relating to the Plan S requirements on rights and licensing. The report can be found here: 10.5281/zenodo.4046624)</p> <p><strong>Data reuse information</strong></p> <p>These data were compiled by Chris Morrison and Jane Secker as part of the above-mentioned research project. For more information on how the data were produced and verified see 3.3 of the report 10.5281/zenodo.4046624. There are two sets of data with supporting information that are made available under different reuse terms.</p> <p><strong>1. DOAJ dataset</strong></p> <p>The file "DOAJ data May2020" was extracted from the Directory of Open Access Journals on 10 May 2020 and represents an analysis of the data (which is (c) 2020 DOAJ) by Chris Morrison and Jane Secker. The dataset is licensed under CC BY-SA following the DOAJ data licence.</p> <p><strong>2. 10 large publishers dataset</strong></p> <p>The file "Academic Publishers Copyright Policies and Practices Table" was created by Chris Morrison and Jane Secker and is provided under CC0 dedication.</p> <p>The remaining files provide information from publisher websites and directly from publisher representatives and remain (c) of each respective organisation.</p>
Dataset from "What do developers talk about open source software licensing? " - SEAA2020
<p>This is the dataset used in the respective research work. The abstract is available below.</p> <p>If you want to cite this work, please use:</p> <p> </p> <p>Georgia M. Kapitsaki, Maria Papoutsoglou, Daniel German and Lefteris Angelis, What do developers talk about open source software licensing?, to appear in the Proceedings of the Euromicro Conference on <a href="https://dsd-seaa2020.um.si/seaa/index.html">Software Engineering and Advanced Applications</a>, SEAA 2020.</p> <p>Free and open source software has gained a lot of momentum in the industry and the research community. Open source<br> licenses determine the rules, under which the open source software can be further used and distributed. Previous works<br> have examined the usage of open source licenses in the framework of specific projects or online social coding platforms, examining developers specific licensing views for specific software. However, the questions practitioners ask about licenses and licensing as captured in Question and Answer websites also constitute an important aspect toward understanding practitioners general licenses and licensing concerns. In this paper, we investigate open source license discussions using data from the Software Engineering, Open Source and Law Stack Exchange sites that contain relevant data. We describe the process used for the data collection and analysis, and discuss the main results. Our results indicate that clarifications about specific licenses and specific license terms are required. The results can be useful for developers, educators and license authors.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.