Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
289
datasets available to search
ShareScore release 0.9.0
Dataset results
289 results for “License”
Sentinel2GlobalLULC: A dataset of Sentinel-2 georeferenced RGB imagery annotated for global land use/land cover mapping with deep learning (License CC BY 4.0)
<p>Sentinel2GlobalLULC is a deep learning-ready dataset of RGB images from the Sentinel-2 satellites designed for global land use and land cover (LULC) mapping. Sentinel2GlobalLULC v2.1 contains 194,877 images in GeoTiff and JPEG format corresponding to 29 broad LULC classes. Each image has 224 x 224 pixels at 10 m spatial resolution and was produced by assigning the 25th percentile of all available observations in the Sentinel-2 collection between June 2015 and October 2020 in order to remove atmospheric effects (i.e., clouds, aerosols, shadows, snow, etc.). A spatial purity value was assigned to each image based on the consensus across 15 different global LULC products available in Google Earth Engine (GEE). </p> <p> </p> <p>Our dataset is structured into 3 main zip-compressed folders, an Excel file with a dictionary for class names and descriptive statistics per LULC class, and a python script to convert RGB GeoTiff images into JPEG format. The first folder called "Sentinel2LULC_GeoTiff.zip" contains 29 zip-compressed subfolders where each one corresponds to a specific LULC class with hundreds to thousands of GeoTiff Sentinel-2 RGB images. The second folder called "Sentinel2LULC_JPEG.zip" contains 29 zip-compressed subfolders with a JPEG formatted version of the same images provided in the first main folder. The third folder called "Sentinel2LULC_CSV.zip" includes 29 zip-compressed CSV files with as many rows as provided images and with 12 columns containing the following metadata (this same metadata is provided in the image filenames): </p> <ul> <li>Land Cover Class ID: is the identification number of each LULC class</li> <li>Land Cover Class Short Name: is the short name of each LULC class</li> <li>Image ID: is the identification number of each image within its corresponding LULC class </li> <li>Pixel purity Value: is the spatial purity of each pixel for its corresponding LULC class calculated as the spatial consensus across up to 15 land-cover products </li> <li>GHM Value: is the spatial average of the Global Human Modification index (gHM) for each image</li> <li>Latitude: is the latitude of the center point of each image</li> <li>Longitude: is the longitude of the center point of each image</li> <li>Country Code: is the Alpha-2 country code of each image as described in the ISO 3166 international standard. To understand the country codes, we recommend the user to visit the following website where they present the Alpha-2 code for each country as described in the ISO 3166 international standard:https: //www.iban.com/country-codes</li> <li>Administrative Department Level1: is the administrative level 1 name to which each image belongs</li> <li>Administrative Department Level2: is the administrative level 2 name to which each image belongs</li> <li>Locality: is the name of the locality to which each image belongs</li> <li>Number of S2 images : is the number of found instances in the corresponding Sentinel-2 image collection between June 2015 and October 2020, when compositing and exporting its corresponding image tile</li> </ul> <p>For seven LULC classes, we could not export from GEE all images that fulfilled a spatial purity of 100% since there were millions of them. In this case, we exported a stratified random sample of 14,000 images and provided an additional CSV file with the images actually contained in our dataset. That is, for these seven LULC classes, we provide these 2 CSV files:</p> <ul> <li>A CSV file that contains all exported images for this class </li> <li>A CSV file that contains all images available for this class at spatial purity of 100%, both the ones exported and the ones not exported, in case the user wants to export them. These CSV filenames end with "including_non_downloaded_images".</li> </ul> <p>To clearly state the geographical coverage of images available in this dataset, we included in the version v2.1, a compressed folder called "Geographic_Representativeness.zip". This zip-compressed folder contains a csv file for each LULC class that provides the complete list of countries represented in that class. Each csv file has two columns, the first one gives the country code and the second one gives the number of images provided in that country for that LULC class. In addition to these 29 csv files, we provided another csv file that maps each ISO Alpha-2 country code to its original full country name.</p> <p>© <a href="https://doi.org/10.5281/zenodo.5055632">Sentinel2GlobalLULC Dataset </a>by Yassir Benhammou, Domingo Alcaraz-Segura, Emilio Guirado, Rohaifa Khaldi, Boujemâa Achchab, Francisco Herrera & Siham Tabik is marked with Attribution 4.0 International (CC-BY 4.0)</p>
EU License Plates Images
<p>A collection of cropped vehicle license plates from across the EU (primarily Germany) for training automated license plate detection and extraction ML and OCR models. German plates are further sourced from a variety of states, allowing for sticker detection, extraction, and state classification models to be further developed.</p> <p>This dataset is used in the SODALITE Vehicle IoT use case for training automated license plate recognition models.</p>
Prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, number of journals per country and publisher
<p>An analysis on the prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, country and publisher according to the number of journals.</p>
Numbers and shares of Open Access Journals in Sociology using Creative Commons Licenses, June 2014
<p>a) The data for the year 2014 was retrieved as a CSV-file (doaj_2014-05-07_1330_utf8.csv) from the Directory of Open Access Journals (DOAJ) homepage at 2014-06-08.<br> b) A subset of journals assigned to the subject category Sociology was generated (n=109).<br> c) I manually checked the information on CC-licenses for each of the 109 journals<br> d) Where necessary I added correct information on licenses, see column k in the CSV-file for the updated information. Column l marks entries that were updated.</p> <p> </p>
Numbers and shares of Open Access Journals from all disciplines and from the discipline Sociology using Creative Commons Licenses as listed by the Directory of Open Access Journals (2014-06-08)
<p>These files contains the data on frequencies and shares of Open Access journals using Creative Commons Licenses (comparing journals from all disciplines and sociologial journals). Date of data collection: 2014-06-08.</p>
Dataset for Earth Sciences at Freie Universität Berlin: Open Access, Licenses and Persistent Identifiers Monitoring
<p>In <em>Version 4</em>, <strong>publishers </strong>and <strong>journals</strong> names has been extended.</p> <p>In <em>Version 3</em>, new entries have been added for both <strong>journal </strong>and <strong>non-journal article outputs</strong>, specifically including data from the year <strong>2023</strong>. Minor adjustments were also made to URLs and open access (OA) statuses.</p> <p><em>Note</em>: Data for journal and non-journal article outputs from the year 2023 were unavailable at the time of preparing the <strong>short paper</strong> presenting the results, findable under <a href="https://doi.org/10.5281/zenodo.14170751" target="_blank" rel="noopener">10.5281/zenodo.14170751</a> [1]).</p> <p><br>Started in 2021, Berlin University Alliance (BUA) Open Science Dashboards, followed by the BUA Open Science Magnifiers projects, seek to investigate Open Science (OS) practices across different research domains and communities. A primary focus of these initiatives lies in the development of OS indicators, tailored to discipline specific ones, alongside their visualisation for monitoring.</p> <p>Collaborating closely with the Department of Earth Sciences at Freie Universität Berlin (FU), one of the project's key objectives is the implementation of an Open Science Dashboard for Earth Sciences FU. The visualisation of the first OS metrics is already available under <a href="https://quest-open-earthsciences.charite.de/">https://quest-open-earthsciences.charite.de/</a>.</p> <p>The datasets utilized include the outputs from the Department of Earth Sciences at FU, i.a. on Open Access (OA) categorisations and statuses, persistent identifiers (PIDs) and Open Licences (Creative Commons) availability, published between 2016-2023. These datasets consist of (i) <strong>"journal_articles_v3.csv"</strong> and (ii) <strong>"non_journal_articles_outputs_v3.csv"</strong>, the latter including “book”, “book chapter”, “conference paper”, “conference abstract”, and “other research outputs” (e.g. book reviews, project reports, book chapters in school books, or electronic supplementary material).</p> <p>Data for the dashboard was obtained from the FU university bibliography (<a href="https://frub-berlin.primo.exlibrisgroup.com/">https://frub-berlin.primo.exlibrisgroup.com/</a>), but coverage of PID information was incomplete, OA category information was incomplete and often erroneous, and copyright/open licence information was missing in this data set. Therefore, the data set was <strong>enriched with manually researched information</strong>. Data enrichment was different for journal articles and for non-journal-article publications. For <strong><em>journal articles</em></strong>, <em>copyright/open licence</em> information was added, and <em>open access category</em> information was checked and added or corrected. For <strong><em>non-journal-article outputs</em></strong>, missing <em>PIDs</em> were added and <em>open access category</em> information was checked and added or corrected. </p> <p>The "<em>data_dictionary_earth_sciences_v3.csv"</em> table documents all variables of each data file containing here.</p> <p>Both for the dashboard, and in our following publications, we categorized <strong>"bronze"</strong> OA outputs as closed access. Although such publications are openly available on the publisher's websites, they lack licence information and thus cannot be openly reused, and presumably even change its openness status at any time. Following the methodology of Charité Dashboard on Responsible Research (<a href="https://quest-dashboard.charite.de/#tabStart">https://quest-dashboard.charite.de/#tabStart</a>) we only include "gold", "hybrid" and "green" OA as true OA. Further details about the enrichment process conducted on these datasets can be found under <a href="https://doi.org/10.5281/zenodo.1099821" target="_blank" rel="noopener">10.5281/zenodo.1099821</a>9 [2]</p> <p> </p> <p>[1] Duine, M., Iarkaeva, A., & Hübner, A. (2024, November 15). Initiating discipline-specific Open Science Monitoring with the Open Science Dashboard for Earth Sciences. 28th International Conference on Science, Technology and Innovation Indicators (STI2024), Berlin, Germany. <a href="https://doi.org/10.5281/zenodo.14170751" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.14170751</a><br>[2] Duine, M., Hübner, A., & Iarkaeva, A. (2024). Enrichment of university bibliography data for open science monitoring. Zenodo. <a href="https://doi.org/10.5281/zenodo.10998219">https://doi.org/10.5281/zenodo.10998219</a></p>
A Large-scale Dataset of (Open Source) License Text Variants
<p>We introduce a large-scale dataset of the complete texts of free/open source software (FOSS) license variants. To assemble it we have collected from the Software Heritage archive—the largest publicly available archive of FOSS source code with accompanying development history—all versions of files whose names are commonly used to convey licensing terms to software users and developers.<br> The dataset consists of 6.5 million unique license files that can be used to conduct empirical studies on open source licensing, training of automated license classifiers, natural language processing (NLP) analyses of legal texts, as well as historical and phylogenetic studies on FOSS licensing.<br> Additional metadata about shipped license files are also provided, making the dataset ready to use in various contexts; they include: file length measures, detected MIME type, detected SPDX license (using ScanCode), example origin (e.g., GitHub repository), oldest public commit in which the license appeared.<br> The dataset is released as open data as an archive file containing all deduplicated license blobs, plus several portable CSV files for metadata, referencing blobs via cryptographic checksums.</p> <p>For more details see the included README file and companion paper:</p> <ul> <li>Stefano Zacchiroli. <a href="https://doi.org/10.1145/3524842.3528491"><em>A Large-scale Dataset of (Open Source) License Text Variants</em></a>. In proceedings of the <a href="https://conf.researchr.org/home/msr-2022">2022 Mining Software Repositories Conference (MSR 2022)</a>. 23-24 May 2022 Pittsburgh, Pennsylvania, United States. ACM 2022.</li> </ul> <p>If you use this dataset for research purposes, please acknowledge its use by citing the above paper.</p> <ul> </ul>
Analysis of the type of license for medical dissertation (Lublin 2011-2019)
<p>Analysis of the type of license for medical dissertation (Lublin 2011-2019).</p> <p>The data comes from the Digital Library of the Medical University of Lublin and the Internal Digital Library (2011-2019 to number 90/2019).</p> <p>Percentage of authors who both shared their dissertation and consented to its copying by users - (authors potentially open to sharing work under CC or similar licenses) - "extended" permitted use; percentage of authors who shared their dissertation, but without permission to copy it - only limited use complying with fair use doctrine; percentage of authors who chose access to their doctoral dissertation only in the Library (maximum access limitation, no online access)</p>
Dataset for the study of Potential Code Borrowing and License Violations in Java Projects on GitHub
<p>This is the dataset for the study of Potential Code Borrowing and License Violations in Java Projects on GitHub. The dataset is based on the Public Git Archive and consists of projects on GitHub that have at least 50 stars and have at least one line in Java. A total of 23,378 projects are listed here that we downloaded for analysis on June 1st, 2019.</p>
MSR Licenses
<p>For the purposes of replicating the license study on free software JavaScript ecosystem projects, all results generated in the study, including the script code used to perform the experiments and data analysis, as well as the data generated by the tool used are available here. .</p>
List of licenses discovered in the study of Potential Code Borrowing and License Violations in Java Projects on GitHub
<p>This is the list of licenses discovered in the study of Potential Code Borrowing and License Violations in Java Projects on GitHub. The licenses are ranged by the amount of files that they cover, there are a total of 94 different licenses. Where possible, the names are presented as identifiers at https://spdx.org/licenses/. "GitHub" stands for no license in the file or the project.</p>
Support of Creative Commons Licenses in German disciplinary and institutional open access repositories
<p>Which of the following open content licenses can be chosen for the metadata description of the open access full-texts (apart from the deposit license)? n=81* </p> <p>* This survey question is part of the 2014 Census of Open Access Repositories in Germany, Austria and Switzerland, see: http://nbn-resolving.de/urn:nbn:de:kobv:11-100222687 For the research data see: http://doi.org/10.5281/zenodo.10734 </p>
Total numbers and shares of Open Access Journals using Creative Commons Licenses as listed by the Directory of Open Access Journals
<p>This table provides information on the number and percentage of Open Access Journals listed by the Directory of Open Access Journals using a Creative Commons License. It identifies also the number and share of Journals using CC-Licenses that are compatible to the Open Definition.</p>
Vehicle to Vehicle Comms Using Licensed & Unlicensed Frequecies
<p>V2V communications by selecting an appropriate frequency band through the selection of available licensed and unlicensed frequency bands for vehicles.</p>
GitHub project dataset for license analysis
<p>This dataset consists of a number of GitHub repositories that cover the following programming languages: PHP, Java, JavaScript, C, C++, C#, Python, Visual Basic. The repositories were used by license extraction tools, i.e. FOSSology Nomos, Ninka, to see which open source software licenses exist in the source code. They were also used to perform analysis on the README.md file and discover licenses used in libraries as described in the README.</p>
MJFF Data Community - Creative Commons Training: Copyright and Open Licensing
<p>This training, provided by Shanna Hollich, the Learning and Training Manager of Creative Commons (CC), was hosted by the Michael J. Fox Foundation's Data Community of Practice (DCOP). For more information on the DCOP, please contact: researchcommunity@michaeljfox.org.</p> <p>In the ever-evolving digital landscape, the management of research outputs, including data licensing and copyright, is of utmost importance. This 1.5-hour training provided a forum for participants to learn more about open licensing and copyright. It also aimed to equip participants with the knowledge and best practices they need to effectively navigate the complexities of CC licensing and copyright when using and generating research outputs such as scholarly publications, datasets, and white papers.<br><br>By the end of the workshop, participants developed an understanding of the basic principles of copyright, how it works, and where it applies. Participants are now able to describe the benefits of open licensing and the basics of how Creative Commons licenses work, have a deeper understanding of how research outputs interact with copyright and open licensing, and know where to find additional information and resources.</p> <p>To access a stream of this video with variable resolution, please <a href="https://share.vidyard.com/watch/5vMMsxsZK6q48yNHsDTkPe" target="_blank" rel="noopener">visit this link</a>.</p>
Numbers of Articles, Books and Dissertation theses indexed in BASE and percentages of items published Open Access, under Creative Commons Licenses and under Open Licenses (2013-2017)
<p>A look at the data provided by the Open Access search engine BASE (http://base-search.net) shows that the Open Science compliance among dissertation theses stagnates. BASE knows three categories of accessibility: Open Access, Unknown, Non-Open Access. In the following tables and graphs, figures reported as "Open Access" have been categorised by BASE as Open Access. The tables and graphics show data from BASE (as of 06.03.2018) as follows:</p> <p>a) Indexed theses, books and journal articles</p> <p>b) Indexed theses, books and journal articles published by Open Access</p> <p>c) indexed theses, books and journal articles under Creative Commons licenses.</p> <p>d) indexed theses, books and journal articles, which are published under open licenses in the sense of the Open License, i. e. reflect terms of use of the Open Source.</p> <p> </p> <p>Although doctoral theses already had a high share of open access by 2013 (43%), by 2017 it had risen by only 5% (2017:48%). At the same time, the proportion of books published in open access rose by 14% (from 20% to 34%) and articles by 17% from 44% (2013) to 61% (2017). The same effect can be seen in the proportion of CC-licensed items: Their share rose by 4% (from 9% to 13%) for doctoral theses, by 9% for books (from 4% to 13%) and 8% for articles (from 10% to 18%) between 2013 and 2017. However, the share of openly licensed items is most pronounced: it did not increase for doctoral theses, but remained at 2% between 2013 and 2017; in the same period it increased by 5% (from 1% to 6%) for books, and by 5% (from 5% to 10%) for articles.</p>
Cdt1 inhibits CMG helicase in early S phase to separate origin licensing from DNA synthesis
A fundamental concept in eukaryotic DNA replication is the temporal separation of G1 origin licensing from S phase origin firing. Re-replication and genome instability ensue if licensing occurs after DNA synthesis has started. In humans and other vertebrates, the E3 ubiquitin ligase CRL4Cdt2 starts to degrade the licensing factor Cdt1 after origins fire, raising the question of how cells prevent re-replication in early S phase. Here, using quantitative microscopy, we show that Cdt1 inhibits DNA synthesis during an overlap period when cells fire origins while Cdt1 is still present. Cdt1 inhibits DNA synthesis by suppressing CMG helicase progression at replication forks through the MCM-binding domain of Cdt1, and DNA synthesis commences once Cdt1 is degraded. Thus, instead of separating licensing from firing to prevent re-replication in early S phase, cells separate licensing from DNA synthesis through Cdt1-mediated inhibition of CMG helicase after firing.
Text-fig. 1. Geographic position of the beaver-bearing sites discussed in this paper. Red triangles – records of Castor, blue dots – records of Trogontherium. Bilz II – Bilzingsleben II, Ehr – Weimar-Ehringsdorf, Mosb 2 – Mosbach 2, Taub – Weimar-Taubach, Teg – Tegelen. (This map was created using ArcGIS® software by Esri. ArcGIS® and ArcMap™ are the intellectual property of Esri and are used herein under license. Copyright © Esri). in Mortality Profiles Of Castor And Trogontherium (Mammalia: Rodentia, Castoridae), With Notes On The Site Formation Of The Mid-Pleistocene Hominin Locality Bilzingsleben Ii (Thuringia, Central Germany)
Text-fig. 1. Geographic position of the beaver-bearing sites discussed in this paper. Red triangles – records of Castor, blue dots – records of Trogontherium. Bilz II – Bilzingsleben II, Ehr – Weimar-Ehringsdorf, Mosb 2 – Mosbach 2, Taub – Weimar-Taubach, Teg – Tegelen. (This map was created using ArcGIS® software by Esri. ArcGIS® and ArcMap™ are the intellectual property of Esri and are used herein under license. Copyright © Esri).
WorldCereal open global harmonized reference data repository (CC-BY-SA licensed data sets)
<p>Within the<strong> ESA funded</strong> WorldCereal project we have built an open harmonized reference data repository at global extent for model training or product validation in support of land cover and crop type mapping. Data from 2017 onwards were collected from many different sources and then harmonized, annotated and evaluated. These steps are explained in the harmonization protocol (10.5281/zenodo.7584463). This protocol also clarifies the naming convention of the shape files and the WorldCereal attributes (LC, CT, IRR, valtime and sampleID) that were added to the original data sets.</p> <p>This publication includes those harmonized data sets of which the original data set was published under the CC-BY-SA license or a license similar to CC-BY-SA. See document "_In-situ-data-World-Cereal - license - CC-BY-SA.pdf" for an overview of the original data sets.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.