Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

15

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

15 results for “Open metadata”

Learn how ShareScore rates datasets ↗
zenodo44/100

National Open Access Monitor, Ireland - Research product metadata

<p>This dataset contains the foundational data for the National Open Access Monitor under an open license. OpenAIRE will routinely provide monthly data dumps to Zenodo, encompassing a comprehensive set of data and indicators related to the Open Access Monitor. This collaborative effort ensures accessibility and openness in sharing the data, promoting transparency and facilitating its use for research and analysis purposes.<br>The dataset comprises of the metadata of the research products metadata stored in the parquet format. The file contains two columns ("id", "xml"), the first of which is the OpenAIRE identifier of the research product and the the second the xml representation of the metadata. The schema of the metadata can be found in https://www.openaire.eu/schema/1.0/oaf-1.0.xsd and a detailed description of the contents can be found in https://graph.openaire.eu/docs/ and https://zenodo.org/records/2643199</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

SARS-CoV-2 wastewater surveillance data and metadata in the Open Data Model format. Part 1: Québec City

<p>SARS-CoV-2 wastewater surveillance data and metadata in the Open Data Model format. Part 1: Qu&eacute;bec City Authors</p> <ul> <li>Therrien, J-D<sup>1</sup></li> <li>Maere, T.<sup>1</sup></li> <li>Sanchez-Quete, F.<sup>2</sup></li> <li>Tsitouras, A.<sup>2</sup></li> <li>Goitom, E.<sup>3</sup></li> <li>Cloutier, F.<sup>4</sup></li> <li>Dufour, D.<sup>4</sup></li> <li>Proulx, F. <sup>4</sup></li> <li>Nicola&iuml;, N.<sup>1</sup></li> <li>Philippe, R.<sup>1</sup></li> <li>Tohidi, M.<sup>1</sup></li> <li>Dorner, S.<sup>3</sup></li> <li>Frigon, D.<sup>2</sup></li> <li>Vanrolleghem, P.A.<sup>1</sup></li> </ul> <p>Affiliations</p> <ul> <li><sup>1</sup> model<em>EAU</em>, D&eacute;partement de g&eacute;nie civil et de g&eacute;nie des eaux, Universit&eacute; Laval</li> <li><sup>2</sup> Microbial Community Engineering Lab (MiCEL), Department of Civil Engineering, McGill University</li> <li><sup>3</sup> Polytechnique Montr&eacute;al</li> <li><sup>4</sup> Ville de Qu&eacute;bec</li> </ul> <p>General Remarks</p> <p>Wastewater-based surveillance of SARS-CoV-2 virus can detect between 1 and 30 infected individuals per 100,000 (including asymptomatic ones) by analyzing the population&#39;s sewage. As such, this method is very attractive since it costs only a fraction of clinical testing (as low as 1%). Human faeces may contain the virus a few days before a person becomes ill. Thus, this approach allows for detection of outbreaks 2-7 days before the increase in reported cases stemming from clinical screening tests (Bibby et al., 2021). Wastewater-based surveillance complements clinical testing by geolocating outbreaks, which may help targeting intensive screening programs. Moreover, it provides a quick indication of whether new public health measures (e.g., masks, social distancing, confinement, and curfew) are effective.</p> <p>Sampling</p> <p>The reported dataset contains open data collected in the province of Qu&eacute;bec as part of the SARS-CoV-2 wastewater-based surveillance program <a href="https://www.centreau.ulaval.ca/en/covid/">CentrEau</a>-COVID. Four of the largest cities in the province (Montr&eacute;al, Laval, Qu&eacute;bec City, and Trois-Rivi&egrave;res), as well as the municipalities of four rural regions (Mauricie, Centre-du-Qu&eacute;bec, Bas-St-Laurent, and Gasp&eacute;sie) participated in the program. The entire dataset includes 31 sampling sites covering approximately half the population of the province of Qu&eacute;bec (population size of 8.5 million). The timeframe covered by the dataset varies for each site. The earliest surveillance program was launched in March 2020, others followed soon after. Samples were collected using various methods, such as 24h composite samples, grab samples, and passive sampling using variations on the Moore swab method (Schang et al., 2020)</p> <p>Analysis</p> <p>Prior to the analysis of the samples for SARS-CoV-2, physiochemical parameters such as total suspended solids (TSS), turbidity, conductivity, ammonium concentration, and pH were measured. The samples were subsequently concentred by filtration using a MEC filter (0.45 um), followed by total RNA extraction using the Qiagen AllPrep PowerViral DNA/RNA Kit (Qiagen, USA) with some modifications (beta-mercaptoethanol concentration raised to 10% and lysis performed at 55 &deg;C for 30 minutes) (Ahmed et al., 2020). SARS-CoV-2 viral RNA was detected by a one-step RT-qPCR. To assess the RNA recovery rate of the procedure, samples were spiked before extraction with a known concentration of Bovine Respiratory Syncytial Virus (BRSV) using the Zoetis INFORCE 3 vaccine (Zoetis, USA). In addition to SARS-CoV-2, samples were assessed for Pepper Mild Mottle Virus (PMMoV), the daily load of which is hypothesized to represent the fecal load contributions to the samples at a given site and time. PCR conditions and primer used to collect viral data are described in the files <code>primers.md</code> and <code>PCR conditions.md</code>.</p> <p>Compilation</p> <p>The measurements on wastewater samples carried out by the participating laboratories of this study are found in the <code>WWMeasure</code> table. The values provided by municipalities come from laboratories accredited by the Centre d&#39;expertise en analyse environnementale du Qu&eacute;bec (CEAEQ), in compliance with the latter&#39;s quality assurance protocols. The COVID-19-related public health data found in the <code>CPHD</code> table were collected from the Institut National de Sant&eacute; Publique du Qu&eacute;bec (INSPQ)&#39;s public reports. Wastewater data taken in-situ at the sampling sites (e.g., the flow at pumping stations or water resource recovery facilities (WRRFs)) are found in the <code>SiteMeasure</code> table and were taken by the institutions responsible for managing the sites. All of the data, stemming from multiple sources, were combined into the <a href="https://github.com/Big-Life-Lab/ODM">Open Data Model (ODM)</a> standard format using the <a href="https://github.com/modelEAU/ODM-Import">ODM-Import python package</a> (see also Structure).</p> <p>Validation</p> <p>Wastewater and sample data were manually assessed for quality by our research collaborators. Data points for which the quality appeared to be uncertain were tagged with the value <code>True</code> in the <code>qualityFlag</code> column. Conversely, data deemed of good quality have a quality flag of <code>False</code>. Data that were not checked have a quality flag of <code>NA</code>. Textual comments describing the issues with the data points in more detail are also included in the dataset using the <code>notes</code> column of the relevant tables. Note that data validation was carried out by the data custodians responsible for each city in the dataset according to available resources. As the project continues and data validation is undertaken on more sections of the dataset, data may be re-analyzed, flagged, or commented as needed. Revisions to the dataset will be reported to the best of our ability.</p> <p>Structure</p> <p>The data contained in this dataset has been structured according to the <a href="https://github.com/Big-Life-Lab/ODM">Open Data Model (ODM) for Wastewater-Based Surveillance</a>. This model provides a standardized dictionary to collect and share data and metadata stemming from wastewater-based surveillance programs. By convention, it splits all data into 10+ thematic tables with each record representing a unique measurement, i.e., long format. For convenience, the <code>wide</code> folder presents the data found in all the other tables in a wide format, i.e., multiple measurements are aligned by <code>timestamp</code>, with each column representing a different parameter.</p> <p>Acknowledgements</p> <p>The authors would like to acknowledge that this dataset was collected thanks to the financial support of the Fonds de Recherche du Qu&eacute;bec, the Molson Foundation, the Trottier Family Foundation, CentrEau and NSERC. The authors would also like to acknowledge the efforts of Douglas Manuel (Ottawa Hospital) and Howard Swerdfeger (Public Health Agency of Canada) for their original idea for the Open Data Model and continued development.</p> <p>References</p> <ol> <li> <p>Ahmed, W., Bertsch, P.M., Bivins, A., Bibby, K., Farkas, K., Gathercole, A., Haramoto, E., Gyawali, P., Korajkic, A., McMinn, B.R., Mueller, J.F., Simpson, S.L., Smith, W.J.M., Symonds, E.M., Thomas, K. v., Verhagen, R., Kitajima, M., 2020. Comparison of virus concentration methods for the RT-qPCR-based recovery of murine hepatitis virus, a surrogate for SARS-CoV-2 from untreated wastewater. Science of the Total Environment 739. <a href="https://doi.org/10.1016/j.scitotenv.2020.139960">https://doi.org/10.1016/j.scitotenv.2020.139960</a></p> </li> <li> <p>Bibby, K., Bivins, A., Wu, Z., North, D., 2021. Making waves: Plausible lead time for wastewater based epidemiology as an early warning system for COVID-19. Water Research 202, 117438. <a href="https://doi.org/10.1016/j.watres.2021.117438">https://doi.org/10.1016/j.watres.2021.117438</a></p> </li> <li> <p>Schang, C., Crosbie, N., Nolan, M., Poon, R., Wang, M., Jex, A., Scales, P., Schmidt, J., Thorley, B.R., Henry, R., Kolotelo, P., Langeveld, J., Schilperoort, R., Shi, B., Einsiedel, S., Thomas, M., Black, J., Wilson, S., McCarthy, D.T., 2020. Passive sampling of viruses for wastewater-based epidemiology: a case-study of SARS-CoV-2 [WWW Document]. URL <a href="https://www.researchgate.net/publication/347103410\_Passive\_sampling\_of\_viruses\_for\_wastewater-based\_epidemiology\_a\_case-study\_of\_SARS-CoV-2?channel=doi&amp;linkId=5fd800f392851c13fe892393&amp;showFulltext=true">https://www.researchgate.net/publication/347103410\_Passive\_sampling\_of\_viruses\_for\_wastewater-based\_epidemiology\_a\_case-study\_of\_SARS-CoV-2?channel=doi&amp;linkId=5fd800f392851c13fe892393&amp;showFulltext=true</a> (accessed 1.18.21).</p> </li> </ol>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Coverage and quality of open metadata for Dutch research output - dataset

<p>Record level data underlying the figures and tables in the report:<strong><br><br>Coverage and quality of open metadata for Dutch research output - report<br></strong></p> <p><a href="https://doi.org/10.5281/zenodo.10629457" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10629457</a><br><br>The current dataset contains 3 csv files:</p> <ul> <li><em>rpo_nl_list_long_20240201.csv</em> - list of identifiers (ROR ID, OpenAlex ID, OpenAIRE ID) of Dutch research performing organisations.&nbsp; <p>Identifiers were collected for the following groups of Dutch RPOs (see Appendix A):</p> <ul> <li> <p>Universities, organised in Universities of the Netherlands (UNL, n=14);</p> </li> <li> <p>University Medical Centres, organised in the Dutch Federation of University Medical Centres (NFU, n=9);&nbsp;</p> </li> <li> <p>National research institutes under the umbrella organisation of the Foundation for Dutch Scientific Research Institutes (NWO-i, n= 9);</p> </li> <li> <p>Research institutes of the Royal Netherlands Academy of Arts and Sciences (KNAW, n=10);</p> </li> <li> <p>Universities of Applied Sciences affiliated to the Netherlands Association of Universities of Applied Sciences (Vereniging Hogescholen) (VH, n=35 of 37)<br><br></p> </li> </ul> </li> <li><em>openalex_works_20231223_rpo_nl_2022&nbsp;</em>- record-level data of OpenAlex records retrieved for all Dutch RPOs in scope of the pilot (UNL/NFU, NWO-i, KNAW, VH) for publication year 2022<br><br></li> <li><em>openaire_products_20240116_rpo_nl_2022 - </em>record-level data of OpenAlex records retrieved for all Dutch RPOs in scope of the pilot (UNL/NFU, NWO-i, KNAW, VH) for publication year 2022</li> </ul>

opencc-zeroMay 2024View details →
zenodo44/100

Resource Metadata Harvested from Government and Research Open Data Portals

<p>This dataset consists of resource metadata harvested from the APIs of hundreds of government and research data portals from all over the world. This dataset was harvested between the 13<sup>th</sup> and 15<sup>th</sup> of September 2018. The metadata harvested from these portals was translated to a single metadata format (see <em>metadata_format.odt</em>). An overview of all harvested domains&nbsp;is given in <em>portal_list.txt</em>.</p> <p>The harvested data is divided into five gzipped&nbsp;json-lines files, based on the &lsquo;type&rsquo; of the resource that is derived from the data of the APIs:</p> <ul> <li><em>dataset_metadata.jsonl.gz</em>: Resources classified as a Dataset, or subsets of dataset (e.g. Dataset:Image and Dataset:Audio) [6 246 250 resources]</li> <li><em>document_metadata.jsonl.gz</em>: Resources classified as a Document, or subset of document (e.g. Document:Paper:Conference and Document:Book) [15 626 541 resources]</li> <li><em>software_metadata.jsonl.gz</em>: Resources classified as Sofware (including Software:Model) [42 036 resources]</li> <li><em>service_metadata.jsonl.gz</em>: Resources classified as a service (e.g. WMS, APIs) [1257 resources]</li> <li><em>other_metadata.jsonl.gz</em>: Resources of which the &lsquo;type&rsquo; could not be determined from the data the API returned. This set still contains many datasets [1 502 979 resources]</li> </ul>

opencc-by-4.0Sep 2018View details →
zenodo44/100

PLOS Open Science Indicators & Zotero Romania Metadata

<p>Matched metadata from PLOS Open Science Indicators and Zotero export</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

The LOTUS Initiative for Open Natural Products Research: frozen dataset union wikidata (with metadata)

<p>Dataset present on Wikidata used in the frame of the LOTUS Initiative: <a href="https://doi.org/10.7554/eLife.70780">https://doi.org/10.7554/eLife.70780</a></p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

The LOTUS Initiative for Open Natural Products Research: metadata

<p>Metadata of each of the three objects (structures, organisms, references) used in the frame of the LOTUS Initiative: <a href="https://doi.org/10.7554/eLife.70780">https://doi.org/10.7554/eLife.70780</a></p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Zenodo Open Metadata snapshot - Training dataset for records and communities classifier building

<p>This dataset contains Zenodo&#39;s published open access records&nbsp;and communities&nbsp;metadata, including entries marked by the Zenodo staff as spam and deleted.</p> <p>The datasets are&nbsp;gzipped compressed&nbsp;JSON-lines&nbsp;files, where each line is a JSON object representation of a Zenodo record or community.</p> <p><strong>Records dataset</strong></p> <p>Filename:<strong> </strong>zenodo_open_metadata_{ date of export }.jsonl.gz</p> <p>Each object&nbsp;contains the terms:&nbsp;<em>part_of,&nbsp;thesis, description, doi, meeting, imprint, references, recid, alternate_identifiers, resource_type, journal, related_identifiers,&nbsp;title, subjects, notes, creators, communities, access_right,&nbsp;keywords, contributors, publication_date</em></p> <p>which correspond&nbsp;to the fields with the same name available&nbsp;in Zenodo&#39;s record JSON Schema at&nbsp;<a href="https://zenodo.org/schemas/records/record-v1.0.0.json">https://zenodo.org/schemas/records/record-v1.0.0.json</a>.</p> <p>In addition, some terms have been altered:</p> <ul> <li>The term <strong>files</strong>&nbsp;contains a list of dictionaries containing <strong>filetype</strong>, <strong>size,</strong>&nbsp;and <strong>filename&nbsp;</strong>only.</li> <li>The term <strong>license</strong>&nbsp;contains a short Zenodo ID of the license (e.g.&nbsp;&quot;cc-by&quot;).</li> </ul> <p><strong>Communities dataset</strong></p> <p>Filename:<strong> </strong>zenodo_community_metadata_{ date of export }.jsonl.gz</p> <p>Each object&nbsp;contains the terms: <em>id, title, description, curation_policy, page&nbsp;</em></p> <p>which&nbsp;correspond&nbsp;to the fields with the same name available&nbsp;in Zenodo&#39;s community creation form.</p> <p><strong>Notes for all&nbsp;datasets</strong></p> <p>For each object the term <strong>spam</strong>&nbsp;contains a boolean value, determining whether a given record/community was marked as&nbsp;spam content&nbsp;by Zenodo staff.</p> <p>Some values for the top-level terms, which were missing in the metadata may contain a&nbsp;<strong>null</strong> value.</p> <p>A smaller uncompressed random sample of 200 JSON lines is&nbsp;also included for each dataset&nbsp;to test and get familiar with the format without having to download the entire dataset.</p>

opencc-by-4.0Dec 2022View details →
zenodo36/100

Complementary dataset of Overton metadata on citing policy-related documents for the study "From intent to impact: Investigating the effects of open sharing commitments"

<p>This document provides the underlying dataset for the bibliometric component for the 2022 study &quot;From intent to impact: Investigating the effects of open sharing commitments&quot; by Research Consulting and Science-Metrix.</p> <p>Before reproducing the study findings or re-using the underlying datasets for other purposes, please cautiously review their limitations in the study&#39;s technical annex and main report, available at: https://zenodo.org/communities/data-sharing-in-public-health-emergencies/&nbsp;</p> <p>Special thanks from the Science-Metrix / Elsevier teams to Euan Adie and Overton for this exceptional public release of Overton metadata, and for conducting extraordinary data collection to retrieve citations towards arXiv preprints.</p> <p>&nbsp;</p> <p>Scope: note that this file combines cited journal publications and preprints from the Covid19, HVRD, Zika and HVVD thematic sets.</p> <p>Data treatment: this data is intend foremost to provide manual validation or qualitative triangulation of our findings. No special efforts have been made to process&nbsp;and clean the data&nbsp;for its eventual re-use in secondary analysis or&nbsp;text mining approaches.</p> <p>Definitions used in this table:</p> <table> <tbody> <tr> <td>Column name&nbsp;</td> <td>Definition</td> </tr> <tr> <td>document_type</td> <td>preprint or journal publication</td> </tr> <tr> <td>doi</td> <td>digital object identifier</td> </tr> <tr> <td>arxiv_id</td> <td>arXiv preprint server&#39;s unique identifier for its preprints</td> </tr> <tr> <td>ssrn_id</td> <td>SSRN preprint server&#39;s unique identifier for its preprints. Note that some of these IDs are contained within the DOIs also assigned to some (but not all) SSRN preprints , in the form of &quot;10.2139/ssrn.&quot; + &#39;ssrn_id&#39;</td> </tr> <tr> <td>coalesce_id</td> <td>coalesce function applied to the DOI, arxiv_id and ssrn_id. Redundant for journal publications.</td> </tr> <tr> <td>policy_source_title</td> <td>name of the policy-related organization</td> </tr> <tr> <td>published_on</td> <td>publication date of the citing policy-related document</td> </tr> <tr> <td>title</td> <td>title of the citing policy-related document</td> </tr> <tr> <td>pdf_url</td> <td>URL for the online version of the policy-related document</td> </tr> <tr> <td>snippet</td> <td>Where available, excerpt of the text immediatly before and after the citation to a journal publication or preprint found in the citing policy-related document</td> </tr> </tbody> </table> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Dataset: The availability and completeness of open funder metadata - Case study for publications funded by the Dutch Research Council

<p>Data and code&nbsp;belonging to the manuscript:&nbsp;<strong>The availability and completeness of open funder metadata - Case study for publications funded by the Dutch Research Council&nbsp;</strong></p> <p><strong>Abstract:</strong><br> Research funders spend considerable efforts collecting information on outcomes of the research they fund. To help funders track publication output associated with their funding, Crossref initiated FundRef in 2013, enabling publishers to register funding information using persistent identifiers. However, it is hard to assess the coverage of funder metadata because it is unknown how many articles are the result of funded research and therefore should include funder metadata.&nbsp;</p> <p>In this paper we looked at 5,004 publications reported by researchers to be the result of funding by a specific funding agency: the Dutch Research Council NWO. Only 67% of these articles contain funding information in Crossref, with a subset acknowledging NWO as funder name and/or Funder IDs linked to NWO (53% and 45%, respectively).&nbsp;</p> <p>Web of Science (WoS), Scopus and Dimensions are all able to infer additional funding information from funding statements in the full text of the articles. Funding information in Lens largely corresponds to that in Crossref, with some additional funding information likely taken from PubMed. &nbsp;</p> <p>We observe interesting differences between publishers in the coverage and completeness of funding metadata in Crossref compared to proprietary databases, highlighting potential to increase the quality of open metadata on funding.&nbsp;</p> <p><strong>This dataset contains the following files:</strong></p> <ul> <li><strong>DOIs_unique_CR_Lens_Wos_Scopus_Dim.csv</strong><em> -&nbsp;</em>Dataset of unique DOIs (n= 5,004) with collected information from Crossref and presence/absence of funder information in Lens, Web of Science, Scopus &nbsp;and Dimensions&nbsp;</li> <li><strong>NWO_funder_names_Crossref.txt</strong><em>&nbsp;</em>- List of funder name variants for NWO found in Crossref</li> <li><strong>Google_Apps_Script.js</strong> - Google Apps Script for retrieving information from Crossref and processing Dimensions results&nbsp;</li> <li><strong>DOI_cleaning.R</strong><em> - </em>R script for cleaning DOIs</li> </ul>

opencc-pddcJul 2022View details →
zenodo36/100

Reti Medievali Open Archive: Metadata in RDF-XML

<p>This&nbsp;RDF-XML dump contains metadata about the items&nbsp;deposited in <em>RM Open Archive</em>&nbsp;converted in Linked Open Data.</p> <p>RM&nbsp;<em>Open Archive</em>&nbsp;is an Open Access scholarly repository, which covers the whole range of medieval studies: social, economic, political and institutional history, as well as cultural, religious and gender representations and practices.</p> <p>RM&nbsp;<em>Open Archive</em>&nbsp;was realised in the frame of the PRIN 2010-2011 project&nbsp;<a href="http://www.medievistica.unina.it/"><em>Concepts, Practices and Institutions of a Discipline: Italian Medieval Studies in 19th and 20th Centuries</em></a>, coordinated by Prof. Roberto Delle Donne at &quot;Federico II&quot; University of Naples.&nbsp;<br> It is under the aegis of the following learned societies, which invite their members to deposit publications.&nbsp;</p> <p><em><a href="http://www.sismed.eu/it/">Societ&agrave; italiana degli storici medievisti</a></em>&nbsp;(Italian Society of Medievalists)</p> <p><em>Consulta per il Medioevo e l&#39;Umanesimo latini</em>&nbsp;(Council for Latin Middle Ages and Humanism)</p> <p><em><a href="http://www.sifr.it/">Societ&agrave; Italiana di Filologia Romanza</a></em>&nbsp;(Italian Society of Romance Philology)</p> <p><em><a href="http://www.paleografi-diplomatisti.org/">Associazione Italiana dei Paleografi e Diplomatisti</a></em>&nbsp;(Italian Association of Paleography and Diplomatics)</p> <p><em><a href="http://www.mediaevistenverband.de/">Medi&auml;vistenverband e.V.</a></em>&nbsp;(German Association of Medievalists)</p>

opencc-zeroDec 2018View details →
zenodo36/100

Open metadata of the Institutional Repository (O2) of the UOC (Dataset in English)

<pre>Dataset of the metadata of all the academic and scientific production generated by the university community of the UOC.</pre>

opencc-by-4.0Apr 2023View details →
zenodo36/100

Open Source Software Package, Version and Dependency Metadata

<p>This dataset contains metadata about 7 million open source packages, versions and dependencies from 34 different software ecosystems in 4 different csv files exported from <a href="https://packages.ecosyste.ms/">https://packages.ecosyste.ms</a>. This can be seen as a spiritual successor to https://zenodo.org/records/3626071, up to date for 2023, including more packages and an expanded array of metadata fields.</p><p>Full copies of the postgresql database this data was extracted from are also available: <a href="https://packages.ecosyste.ms/open-data">https://packages.ecosyste.ms/open-data</a>&nbsp;</p><h4>Files</h4><p>Expanded csv sizes:</p><ul><li>packages-1.0.0-2023-10-17.csv - 2.6GB - 7,015,326 rows</li><li>packages_with_repository_fields-1.0.0-2023-10-17.csv - 4.3GB - 7,019,492 rows</li><li>versions-1.0.0-2023-10-17.csv - 14GB - 84,953,382 rows</li><li>dependencies2-1.0.0-2023-10-17.csv - 124GB - 976,944,543 rows</li></ul><p>Contact</p><p>If you would like any help, support or more data from Ecosyste.ms please do get in touch via email: hello@ecosyste.ms or open an issue on GitHub: <a href="https://github.com/ecosyste-ms/packages/issues">https://github.com/ecosyste-ms/packages/issues</a></p>

opencc-by-sa-4.0Oct 2023View details →
zenodo32/100

Libraries.io Open Source Repository and Dependency Metadata

<p><strong>What is in this release?</strong></p> <p>In this release you will find data about software distributed and/or crafted publicly on the Internet. You will find information about its development, its distribution and its relationship with other software included as a dependency. You will not find any information about the individuals who create and maintain these projects.</p> <p>Further information and documentation on this data set can be found at https://libraries.io/data</p> <p>For enquiries please contact data@libraries.io</p> <p>This dataset contains seven csv files:</p> <p><strong>Projects</strong></p> <p>A project is a piece of software available on any one of the 34 package managers supported by Libraries.io.</p> <p><strong>Versions</strong></p> <p>A Libraries.io version is an immutable published version of a Project from a package manager. Not all package managers have a concept of publishing versions, often relying directly on tags/branches from a revision control tool.</p> <p><strong>Tags</strong></p> <p>A tag is equivalent to a tag in a revision control system. Tags are sometimes used instead of Versions where a package manager does not use the concept of versions. Tags are often semantic version numbers.</p> <p><strong>Dependencies</strong></p> <p>Dependencies describe the relationship between a project and the software it builds upon. Dependencies belong to Version. Each Version can have different sets of dependencies. Dependencies point at a specific Version or range of versions of other projects.</p> <p><strong>Repositories</strong></p> <p>A Libraries.io repository represents a publically accessible source code repository from either github.com, gitlab.com or bitbucket.org. Repositories are distinct from Projects, they are not distributed via a package manager and typically an application for end users rather than component to build upon.</p> <p><strong>Repository dependencies</strong></p> <p>A repository dependency is a dependency upon a Version from a package manager has been specified in a manifest file, either as a manually added dependency committed by a user or listed as a generated dependency listed in a lockfile that has been automatically generated by a package manager and committed.</p> <p><strong>Projects with related Repository fields</strong></p> <p>This is an alternative projects export that denormalizes a projects related source code repository inline to reduce the need to join between two data sets.</p> <p><strong>Licence</strong></p> <p>This dataset is released under the Creative Commons Attribution-ShareAlike 4.0 International Licence.</p> <p>This licence provides the user with the freedom to use, adapt and redistribute this data. In return the user must publish any derivative work under a similarly open licence, attributing Libraries.io as a data source. The full text of the licence is included in the data.</p> <p><strong>Access, Attribution and Citation</strong></p> <p>The dataset is available to download from Zenodo at&nbsp;https://zenodo.org/record/2536573.</p> <p>Please attribute Libraries.io as a data source by including the words &lsquo;Includes data from Libraries.io, a project from Tidelift&rsquo; and reference the Digital Object identifier:&nbsp;10.5281/zenodo.3626071</p>

opencc-by-4.0Dec 2018View details →
zenodo32/100

Dump of metadata from the Czech National Open Data Catalog, 2020-04-20, State Administration of Land Surveying and Cadastre datasets removed

<p>Dump of metadata using DCAT-AP&nbsp;from the Czech National Open Data Catalog, 2020-04-20, in RDF TriG.<br> Used in the paper &quot;Modular Framework for Similarity-based Dataset Discovery using External Knowledge&quot;<br> State Administration of Land Surveying and Cadastre datasets were removed from the dump for practical reasons.</p>

opencc-by-4.0Jan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record