Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,146

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,146 results for “collaboration;”

Learn how ShareScore rates datasets ↗
zenodo44/100

Global trends and collaborations in electrochemical methods: a dataset on etching and deposition research

<p><span>This dataset supports the study "Electrochemical Etching vs. Electrochemical Deposition: A Comparative Bibliometric Analysis," which examines scientific publications on electrochemical etching and electrochemical deposition from 1970 to 2023. The dataset is derived from the Science Citation Index Expanded (SCIE) database and includes bibliometric information on publication trends, leading contributors, research areas, and keyword co-occurrences in both fields.</span></p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Secondary Data from Insights from Publishing Open Data in Industry-Academia Collaboration

<h1>Secondary Data from Insights from Publishing Open Data in Industry-Academia Collaboration</h1> <h2>Authors</h2> <p>Per Erik Strandberg [1], Philipp Peterseil [2], Julian Karoliny [3], Johanna Kallio [4], and Johannes Peltola [4].</p> <p>[1] Westermo Network Technologies AB (Sweden).<br>[2] Johannes Kepler University Linz (Austria)<br>[3] Silicon Austria Labs GmbH (Austria).<br>[4] VTT Technical Research Centre of Finland Ltd. (Finland).</p> <h2>Description</h2> <p>This data is to accompany a paper submitted to Elsevier's data in brief in 2024, with the title <em>Insights from Publishing Open Data in Industry-Academia Collaboration</em>.</p> <p><em>Tentative Abstract:</em> Effective data management and sharing are critical success factors in industry-academia collaboration. This paper explores the motivations and lessons learned from publishing open data sets in such collaborations. Through a survey of participants in a European research project that published 13 data sets, and an analysis of metadata from almost 281 thousand datasets in Zenodo, we collected qualitative and quantitative results on motivations, achievements, research questions, licences and file types. Through inductive reasoning and statistical analysis we found that planning the data collection is essential, and that only few datasets (2.4%) had accompanying scripts for improved reuse. We also found that authors are not well aware of the importance of licences or which licence to choose. Finally, we found that data with a synthetic origin, collected with simulations and potentially mixed with real measurements, can be very meaningful, as predicted by Gartner and illustrated by many datasets collected in our research project.</p> <h2>Secondary data from Survey</h2> <p>The file <code>survey.txt</code> contains secondary data from a survey of participants that published open data sets in the 3-year European research project InSecTT.</p> <h2>Secondary data from Zenodo</h2> <p>The file <code>secondary_data_zenodo.json</code> contains secondary data from an analysis of data sets published in Zenodo. It is accompanied with a <code>py</code>-file and a <code>ipynb</code>-file to serve as examples.</p> <h2>License</h2> <p>This data is licenced with the Creative Commons Attribution 4.0 International license. You are free to use the data if you attribute the authors. Read the license text for details.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

S120 | DUSTCT2024 | Substances from Second NORMAN Collaborative Dust Trial

<p>This is the collection associated with list S120 DUSTCT2024 Substances from Second NORMAN Collaborative Dust Trial on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>List of substances detected via GC-MS and LC-MS from the second NORMAN collaborative dust trial initiated in 2020, including classification and detection information as described in Haglund et al (2024) DOI: <a href="https://doi.org/10.1016/j.scitotenv.2024.177639">10.1016/j.scitotenv.2024.177639</a>.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Data and figures for "A half-century of global collaboration in science and the 'Shrinking World'"

<p>This package supplements the paper entitled <i>"A half-century of global collaboration in science and the 'Shrinking World'"</i> published in <i>Quantitative Science Studies</i> (doi: <a href="https://doi.org/10.1162/qss_a_00268">10.1162/qss_a_00268</a>).</p><p>It contains the datasets and figures used in the original paper&nbsp;based on bibliometric data from a broad set of scientific publications (works), including journal articles, preprints and datasets; see the subfolder named "all_works".</p><p>In addition, for reference,&nbsp;it also contains&nbsp;datasets and figures based on bibliometric data from&nbsp;journal articles only; see the subfolder named "journal_only". The bottom-level files in this subfolder are suffixed with '_j' for identification.</p><p>&nbsp;</p><p><strong>Contents and Instructions</strong></p><p>The datasets and figures in this package are based on the&nbsp;data obtained via <a href="https://docs.openalex.org/api/">OpenAlex API</a>. See the original paper for details.&nbsp;The following file and folders are found at the next level of the subfolders named "all_works"&nbsp;or "journal_only".</p><p>&nbsp;</p><p><strong>- nworks_intlrate_master</strong> (.csv file)</p><ul><li>This file contains information on the number of works&nbsp;('nworks_all') produced in each of the 15 research disciplines ('discipline' and 'disc_ID'; see below) by 18 countries (Australia, Canada, China, France, Germany, India, Indonesia, Iran, Italy, Japan, Netherlands, Poland, Russia, South Korea, Spain, Switzerland, UK and&nbsp;US) ('country' and 'country_code') from 1970 to 2021 ('year'), the number of international collaborative works among them ('nworks_intl'), and the international collaboration rate ('intlrate') calculated from the ratio of the two.</li><li>The 15 disciplines are Artificial Intelligence ('disc_ID' = 1; 'ai'), Quantum Science (2; 'quantum'), Biotechnology (3; 'bio'), Nanotechnology (4; 'nano'), Agricultural Engineering (5; 'agri'), Particle Physics (6; 'particle'), Aerospace Engineering (7; 'aerospace'), Nuclear Engineering (8; 'nuclear'), Marine Engineering (9; 'marine'), Neuroscience (10; 'neuro'), Condensed Matter Physics (11; 'condensed'), Environmental Engineering (12; 'envi'), Earth Science (13; 'earth'), Astronomy (14; 'astro') and Pure Mathematics (15; 'math').&nbsp;See the original paper for the definitions of these disciplines.</li><li>The figures contained in the folders '[line]_nworks' and '[line]_intlrate' are based on this dataset.</li></ul><p>&nbsp;</p><p><strong>&nbsp;- [line]_nworks</strong> (Folder)</p><ul><li>This folder contains line plots (.pdf/.png) representing the trends in the number of works by discipline and country, corresponding to the left-hand side diagrams of Fig. 1 and Suppl. Fig. S2 in the v1 preprint.</li></ul><p>&nbsp;</p><p><strong>- [line]_intlrate</strong> (Folder)</p><ul><li>This folder contains line plots (.pdf/.png) representing the trends in the international collaboration rate by discipline and country, corresponding to the right-hand side diagrams of Fig. 1 and Suppl. Fig. S2 in the v1 preprint.</li></ul><p>&nbsp;</p><p><strong>- [chord]_bilateral</strong> (Folder)</p><ul><li>This folder contains chord diagrams (.pdf/.png) representing the bilateral collaborative relationships by discipline and period, corresponding to Fig. 2 and Suppl. Fig. S4 in the v1 preprint. The number at the end of the file name indicates the period represented by the diagram;&nbsp;specifically, '1' = 1971–1990, '2' = 1991–2000, '3' = 2001–2010 and '4' = 2011–2020.</li><li>The raw data (.xlsx) to reproduce the contained diagrams are also provided by discipline in the accompanied 'Data' folder. The file named '[list]_nworks_(discipline name).xlsx' shows, for the top 30 countries ('country' and 'country_code') in work production during the period indicated by the sheet name, their work production ('nworks_all'), the number of international collaborative works among them ('nworks_intl'), and the international collaboration rate ('intlrate') calculated from the ratio of the two. The file named '[mat]_bilat_nworks_(discipline name)' shows the number of works produced by each country pair during the period indicated by the sheet name.&nbsp;Country names are abbreviated by two-letter country codes (ISO 3166-1 alpha-2).</li></ul><p>&nbsp;</p><p><strong>- [dend]_hcluster</strong> (Folder)</p><ul><li>This folder contains circularised dendrograms (.pdf/.png) representing the international research collaboration clusters by discipline and period, corresponding to Fig. 3 and Suppl. Fig. S5 in the v1 preprint.</li><li>The raw data (.xlsx) to reproduce the contained diagrams are also provided by discipline in the accompanied 'Data' folder. The file named '[mat]_bilat_dist_(discipline name)' shows the distance between each country pair for the period indicated by the sheet name, calculated based on the formula presented in the original paper. Country names are abbreviated by two-letter country codes (ISO 3166-1 alpha-2).</li></ul>

opencc-by-4.0Nov 2022View details →
zenodo44/100

The e-NDP project : collaborative digital edition of the Chapter registers of Notre-Dame of Paris (1326-1504). Ground-truth for handwriting text recognition (HTR) on late medieval manuscripts.

<p>The <a href="https://endp.hypotheses.org/">e-NDP project</a>, funded by the ANR, is led by the <a href="https://lamop.hypotheses.org/6870">LaMOP</a> (Julie Claustre and Darwin Smith).</p> <p>The project&#39;s partners are the Archives nationales, the&nbsp;Biblioth&egrave;que nationale de France (Department of Manuscripts, Biblioth&egrave;que de l&#39;Arsenal), the &Eacute;cole nationale des chartes and the Biblioth&egrave;que Mazarine.</p> <p>The e-NDP project aims at renewing our knowledge on <strong>Notre-Dame de Paris cathedral</strong> through the creation of a collaborative digital edition of the registers of its Chapter (1326-1504, <em>AN LL 105-128</em>), the community of 51 canons meeting three times a week on set days to take all administrative, financial and practical decisions pertaining to the cathedral, its estate and the society living in its cloister. This corpus has never been the object of a comprehensive study to understand the workings and history of this urban enclave and powerful community. The collaborative digital edition is based on a process of<strong> handwriting text recognition (HTR)</strong>, tested and supervised by scholars, researchers and engineers combining expertise in Medieval history, paleography, philology and digital humanities. The edition shall allow a better insight into the Chapter&rsquo;s administration, into its economical and political power within Paris, and the relationships it maintained with other institutions in the city.</p> <p>&nbsp;</p> <p><strong>Section 1 : The e-NDP ground-truth dataset for Handwriting text recognition.</strong></p> <p>The full e-NDP corpus kept today in the French National Archives and was entirely digitized and described in its&nbsp;<a href="https://www.siv.archives-nationales.culture.gouv.fr/siv/rechercheconsultation/consultation/ir/consultationIR.action?formCaller=GENERALISTE&amp;irId=FRAN_IR_059635">catalog</a>&nbsp;in 2022.</p> <p>The first major goal of the&nbsp;e-NDP projet is to propose a first automatic transcription of the 14k pages composing the 26 chapter registers. To achieve this goal representative samples from&nbsp;each one of the volumes were selected and transcribed in order to train a specialized HTR model able to propose a high quality automatic transcription. The collected ground-truth released on this repository currently has <strong>512 pages from the 26 registers</strong> of the cathedral chapter preserved in the National Archives (LL105 - LL128, <strong>1326-1504</strong>). The transcriptions were manually completed in <strong>two rounds</strong> by a group of 12 contributors, historians and paleographers, over the course of 2021-2022 using <a href="https://escriptorium.paris.inria.fr/">eScriptorium </a>as annotation environment.&nbsp;&nbsp;</p> <p>&nbsp;</p> <p><strong>Ground-truth features :</strong></p> <p><br> <em>Number of hands </em>: according to our estimates no fewer than 18&nbsp;main hands were involved in the writing of the registers during the medieval period.&nbsp;</p> <p><em>Language</em> : More than 98% of the content of the registers was written in Latin, the rest in French. The exact percentage is hard to estimate because the vernacular language is often used in formulae, notes and comments. It is rare to find entire pages or blocks written in French.&nbsp;</p> <p><em>Script family</em> : The registers were written using a Cursive script (ca. late XIIIe - XVIe).</p> <p><em>Documental typology</em> : The volumes containing the chapter conclusions were conceived to serve&nbsp;as memorial&nbsp;records, but above all as documents for regular use and consultation in the daily practice of administration and management. In diplomatics the notion of &quot;documentary manuscripts&quot; is used to describe this kind of sources&nbsp;also by opposition to books and litterary or&nbsp;normative&nbsp;manuscripts.</p> <table align="center"> <caption><strong>Ground truth statistics</strong></caption> <tbody> <tr> <th>Text units</th> <th>Count</th> </tr> <tr> <td>Pages</td> <td>512</td> </tr> <tr> <td>Annotated regions (see section 2)</td> <td>2448</td> </tr> <tr> <td>Lines of text</td> <td>34231</td> </tr> <tr> <td>Tokens</td> <td>205083</td> </tr> <tr> <td>Characters</td> <td>3320407</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Rules of transcription :</strong></p> <ul> <li>The abbreviations have been resolved, both those by suspension (<code>facimꝰ</code> ---&gt; <code>facimus</code>) and by contraction (<code>d&ntilde;i</code> --&gt; <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --&gt; <code>et</code> ; <code>ꝓ</code> --&gt; <code>pro</code>) have been resolved.&nbsp;</li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the notary are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the text: <code>.</code> and <code>/</code> have been transcribed, but the transcription has not been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> <li>More specific transcription rules can be found into the file <code>transcription_guidelines.pdf</code></li> </ul> <p>&nbsp;</p> <p><strong>Section 2. e-NDP Layout Segmentation.</strong></p> <p>Layout segmentation is a compulsory step before HTR recognition in order to distinguish sections and regions inside a document. This process intend to separate interdependant page zones to produce a recognition in a section-sequence order and not in a line-sequence order which mix textual and peri-textual content.</p> <p>The regions of 364&nbsp;pages (see <code>GT-layout_list</code>) of the e-NDP corpus were annotated using a 5 sections vocabulary (see <code>endp_layout_regions</code>) in order to describe&nbsp;the page distribution in all the 26 volumes :</p> <ol> <li><em>Block</em>&nbsp;: All the central text blocks, that normally corresponds to the main content called &quot;conclusions&quot; in registers.</li> <li><em>Liste</em>&nbsp;: List of names of the canons who were present during the meeting. Normally located before the <em>conclusions</em>.</li> <li><em>Entr&eacute;e</em>&nbsp;: Marginal notes or entries to inform about the content of <em>conclusions</em>.</li> <li><em>Date</em>&nbsp;: Paragraph contending the date. Normally at the head of a <em>conclusion</em>, but separate of the main body.</li> <li><em>Num&eacute;rotation</em>&nbsp;: Page numbers in roman or arabic. Usually appear in the top corners of the pages.</li> </ol> <table align="center"> <caption><strong>Layout GT statistics</strong></caption> <tbody> <tr> <th>Region</th> <th>Count</th> </tr> <tr> <td>block</td> <td>833</td> </tr> <tr> <td>liste</td> <td>431</td> </tr> <tr> <td>date</td> <td>448</td> </tr> <tr> <td>entr&eacute;e</td> <td>205</td> </tr> <tr> <td>num&eacute;rotation</td> <td>531</td> </tr> </tbody> </table> <p>&nbsp;</p> <p><strong>Section 3. The e-NDP HTR modeling.</strong></p> <p>The e-NDP project has progressively trained several HTR models adapted to work on late medieval cursive in order to accelerate the production of ground truth. Currently the best model delivers an average&nbsp;<strong>CER (Character error ratio) of 9.7%</strong> in handwriting recognition on&nbsp;the 26 registers (see <code>endp_learning_curve</code>) and can serve as generalist model&nbsp;for other manuscripts of the same period and similar script family. These models and their training implementation details can be found in the project&#39;s github <a href="https://github.com/chartes/e-NDP_HTR">repository</a>.&nbsp;</p> <p>Additionally, the automatic HTR transcriptions of the 26 registers (14k pages, 4.5M tokens) enriched with lexical and semantical information has been the subject of a first <a href="https://nosketch-engine.lamop.fr/#dashboard?corpname=endp">online publication</a> using the NoSketch engine that allows advanced data mining based on the combination of data, metadata and NLP features.&nbsp;</p> <p>&nbsp;</p> <p><strong>Section 4. Dataset content.</strong></p> <p>This zip dataset contains :</p> <p>- <code>HTR_ground_truth</code> : Two folders containing the jpg / jpeg images and their curated transcriptions in PAGE XML format.</p> <p>- <code>images_docs</code> : 4 files illustrating the different phases of the project (list of GT for layout segmentation, layout ontologie, transcription guideline and HTR evaluation curves)</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Journal and Data Archive Collaboration Forum [online event recording]

<p>The availability of research data underlying articles published in journals is becoming a common practice in scientific communication. The European Commission and other funders of scientific research have set high expectations for scientists towards openness and availability of scientific work and results. Scientific publishers, through journals and scholarly publications are the main point of realising open science in practice.<br> <br> This event was part of the continuous Journals Outreach initiative (<a href="https://www.cessda.eu/Training/Journals-outreach">https://www.cessda.eu/Training/Journals-outreach</a>), bringing together CESSDA service providers (SPs) with Social Science &amp; Humanities Journals. <strong>Its target audiences were publishers, editors, researchers, and CESSDA Service providers.&nbsp;</strong>The event was also an opportunity for publishers/journals to highlight new initiatives in research data services linked to scientific publications.<br> <br> The video is available on<a href="https://www.youtube.com/watch?v=zCKoyzLifkg"> the&nbsp;CESSDA Training&nbsp;YouTube channel</a>.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

CS3MESH4EOSC Final Event Science Mesh - Unlocking Open Science and Collaborative Research Landscape

<p>The recap video of CS3MESH4EOSC final event. The CS3MESH4EOSC final event, took place on 22 June 2023, at the EGI Conference in Poznan (Poland), showcased how the Science Mesh is contributing to an easier and more robust open science across Europe, thanks to novel approaches for data sharing and synchronisation.</p> <p><strong>The first half of the event</strong> will count with live demonstrations, where each data service from the Science Mesh will be presented from a user-perspective point of view. Event attendees will get practical information on how they can join the Science Mesh as a researcher, a software developer or a service provider. The event will also bring together representatives of Science Mesh use cases, who will explain how Science Mesh is making a difference in their lives, thanks to easier data sharing and synchronisation. A panel discussion with representatives of different sectors, from research to industry and education, will discuss the most urgent trends &amp; priorities for cross-border science collaboration between different sciences.</p> <p><strong>The second half of the event</strong> will be focused on the technical novelties within the Science technical foundation. A series of demonstrations will be presented, followed by a panel discussion, with representatives of CS3MESH4EOSC members that are part of the EOSC Task Forces, on how the Science Mesh contributes to EOSC&rsquo;s success and the overall EOSC Strategic Research and Innovation Agenda (SRIA).</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Data and figures for "Atlas of Science Collaboration, 1971–2020"

<p><strong>Abstract</strong></p><p>The evolving landscape of interinstitutional collaborative research across 15 natural science disciplines is explored using the open data sourced from OpenAlex.&nbsp;This extensive exploration spans the years from 1971 to 2020, facilitating a thorough investigation of leading scientific output producers and their collaborative relationships based on coauthorships.&nbsp;The findings are visually presented on world maps and other diagrams, offering a clear and insightful portrayal of notable variations in both national and international collaboration patterns across various fields and time periods.&nbsp;These visual representations serve as valuable resources for science policymakers, diplomats and institutional researchers, providing them with a comprehensive overview of global collaboration and aiding their intuitive grasp of the evolving nature of these partnerships over time.</p><p>&nbsp;</p><p><strong>Intended Readership</strong></p><ul><li>The booklet, entitled<i>&nbsp;'</i><a href="https://arxiv.org/abs/2308.16810"><i>Atlas of Science Collaboration</i></a><i>'</i>, aims to offer a broad overview of international and interinstitutional research collaboration, shedding light on its present status and evolution on a global scale. While it might not delve into intricate scholarly or academic data analysis, it remains a valuable resource for those seeking a general understanding of the collaborative relationships that have been established between research institutions in the world of science.</li><li>The intended readership including science and technology (S&amp;T) policymakers and diplomats, government research and development (R&amp;D) agencies, international organisations, S&amp;T think tanks, as well as institutional research divisions of universities or R&amp;D institutions.</li></ul><p>&nbsp;</p><p><strong>Data Source</strong></p><ul><li>The<i> </i><a href="https://arxiv.org/abs/2308.16810"><i>Atlas of Science Collaboration</i></a><i>&nbsp;</i>is based on data retrieved from&nbsp;<a href="https://docs.openalex.org/">OpenAlex</a>, a free and open (the CC0 license) catalogue of the world's scholarly papers, researchers, journals and institutions. Launched in January 2022, OpenAlex replaced&nbsp;<a href="https://www.microsoft.com/en-us/research/project/microsoft-academic-graph/">Microsoft Academic Graph (MAG)</a>, which retired at the beginning of 2022.</li><li>OpenAlex collects information on scientific publications, including journal articles, non-journal articles, preprints, conference papers, books and datasets—hereafter collectively referred to as 'works'—from various platforms such as&nbsp;<a href="https://www.crossref.org/">Crossref</a>,&nbsp;<a href="https://orcid.org/">ORCID</a>,&nbsp;<a href="https://ror.org/">ROR</a>,&nbsp;<a href="https://pubmed.ncbi.nlm.nih.gov/">PubMed</a>, preprint servers like&nbsp;<a href="https://arxiv.org/">arXiv</a>, and institutional or disciplinary repositories like&nbsp;<a href="https://zenodo.org/">Zenodo</a>. For comparison with other scholarly data sources such as&nbsp;<a href="https://www.scopus.com/">Scopus</a>,&nbsp;<a href="https://clarivate.com/products/scientific-and-academic-research/research-discovery-and-workflow-solutions/webofscience-platform/">Web of Science</a>&nbsp;and&nbsp;<a href="https://www.dimensions.ai/">Dimensions</a>, please refer to&nbsp;<a href="https://openalex.org/about#comparison">OpenAlex's website</a>.</li><li>OpenAlex offers extensive coverage of meta-information across a diverse spectrum of works, encompassing not only journal publications but also non-journal works, non-English works and contributions from the Global South. This attribute proves beneficial by providing a more precise augmentation of the extent of R&amp;D activities, along with their associated scholarly outputs. This is especially crucial in fields where journals are not the predominant channel for disseminating research outcomes. Furthermore, OpenAlex effectively captures outputs in the preprint format, which might persist for varying durations, spanning from months to years or even indefinitely, without necessarily transitioning into journal publications.</li><li>The present edition (August 2023) of&nbsp;the<i> </i><a href="https://arxiv.org/abs/2308.16810"><i>Atlas of Science Collaboration</i></a><i>&nbsp;</i>was compiled using data obtained via the&nbsp;<a href="https://docs.openalex.org/how-to-use-the-api/api-overview">OpenAlex API</a>&nbsp;during the period from the 12th to the 15th of August 2023. It is essential to note that OpenAlex is an ongoing project, continuously updating its data and improving its system. Consequently, the visualisations in this booklet may not provide the most comprehensive view or accurate data. Expect more accurate results when acquiring data in the future as OpenAlex undergoes further upgrades. Revised editions of&nbsp;the<i> Atlas of Science Collaboration&nbsp;</i>may be made available on&nbsp;<a href="https://zenodo.org/">Zenodo</a>&nbsp;or other open platforms beyond this release.</li></ul><p>&nbsp;</p><p><strong>R&amp;D Disciplines</strong></p><ul><li>In this current edition, the primary focus centres around the level-1 'concepts' listed in the following table&nbsp;sourced from the OpenAlex classification, as previously explored in <a href="https://doi.org/10.48550/arXiv.2211.04429">Okamura (2023)</a>. Each level-1 concept is accompanied by 'related concepts', which can offer a finer or broader delineation compared to the level-1 concept. Using this characteristic, an enhanced notion of R&amp;D discipline is constructed by including all associated subconcepts of level 2 or higher for each of the 15 level-1 concepts. For instance, our defined discipline of 'Artificial Intelligence' includes OpenAlex's level-2 concepts of '<a href="https://explore.openalex.org/concepts/C50644808">Artificial Neural Network</a>' and '<a href="https://explore.openalex.org/concepts/C108583219">Deep Learning</a>', but not the level-0 concepts of '<a href="https://explore.openalex.org/concepts/C41008148">Computer Science</a>' or '<a href="https://explore.openalex.org/concepts/C33923547">Mathematics</a>'.</li></ul><p>&nbsp;</p><p>&nbsp; OpenAlex Concept / Identifier / Discipline Code&nbsp;</p><ol><li>Artificial intelligence&nbsp;/ <a href="https://explore.openalex.org/concepts/C154945302">C154945302</a> / "ai"</li><li>Quantum mechanics&nbsp;/ <a href="https://explore.openalex.org/concepts/C62520636">C62520636</a> / "quantum"</li><li>Biotechnology&nbsp;/ <a href="https://explore.openalex.org/concepts/C150903083">C150903083</a> / "bio"</li><li>Nanotechnology&nbsp;/ <a href="https://explore.openalex.org/concepts/C171250308">C171250308</a> / "nano"</li><li>Agricultural engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/C88463610">C88463610</a> / "agri"</li><li>Particle physics&nbsp;/ <a href="https://explore.openalex.org/concepts/C109214941">C109214941</a> / "particle"</li><li>Aerospace engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/C146978453">C146978453</a> / "aerospace"</li><li>Nuclear engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/C116915560">C116915560</a> / "nuclear"</li><li>Marine engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/c199104240">C199104240</a> / "marine"</li><li>Neuroscience&nbsp;/ <a href="https://explore.openalex.org/concepts/c169760540">C169760540</a> / "neuro"</li><li>Condensed matter physics&nbsp;/ <a href="https://explore.openalex.org/concepts/C26873012">C26873012</a> / "condensed"</li><li>Environmental engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/C87717796">C87717796</a> / "envi"</li><li>Earth science&nbsp;/ <a href="https://explore.openalex.org/concepts/c1965285">C1965285</a> / "earth"</li><li>Astronomy&nbsp;/ <a href="https://explore.openalex.org/concepts/c1276947">C1276947</a> / "astro"</li><li>Pure mathematics&nbsp;/ <a href="https://explore.openalex.org/concepts/C202444582">C202444582</a> / "math"</li></ol><p>&nbsp;</p><p><strong>Analysis and Visualisation</strong></p><ul><li>First,&nbsp;the<i> World Map of Science Collaboration</i> ('<strong>wmap_bilat</strong>' folder)&nbsp;divides the period from 1971 to 2020 into four intervals: 1971–1990, 1991–2000, 2001–2010 and 2011–2020. For each period and discipline, bubbles represent the top 199 research institutions in terms of work production. Additionally, for the top 50 research institutions, their locations are connected on the world map using great circle curves (the shortest route between them) to illustrate bilateral coauthorship relationships. Coauthorship relationships with fewer than five coauthored papers are not displayed. The background world map utilises the&nbsp;world&nbsp;data from the&nbsp;<a href="https://cran.r-project.org/package=maps">maps</a>&nbsp;package&nbsp;in R. The connection visualisation between two research institutions leverages the&nbsp;gcIntermediate()&nbsp;function from the&nbsp;<a href="https://cran.r-project.org/package=geosphere">geosphere</a>&nbsp;package&nbsp;in R. The sizes of the bubbles are proportional to the volume of work and can be compared across the different period panels.</li><li>Second,&nbsp;the<i> Top 30 Productive Institutions on the World Map&nbsp;</i>('<strong>wmap_topinst</strong>' folder)&nbsp;displays the leading 30 institutions in terms of work production on the World Map for each discipline and the three respective periods: 1991–2000, 2001–2010 and 2011–2020. The background world map employs the&nbsp;world&nbsp;data from the&nbsp;<a href="https://cran.r-project.org/package=maps">maps</a>&nbsp;package in R along with the&nbsp;<a href="https://cran.r-project.org/package=ggplot2">ggplot2</a>&nbsp;package16&nbsp;in R. The sizes of the bubbles are proportional to the volume of work, standardised within each period panel, and cannot be compared across panels.</li><li>Third,&nbsp;the<i> Interregional Collaboration Matrix Diagram&nbsp;</i>('<strong>halfmat</strong>' folder)&nbsp;exhibits a half-matrix diagram at the country level for each discipline and the three respective periods: 1991–2000, 2001–2010 and 2011–2020. It counts the number of bilateral coauthorship relationships represented on the World Map. Each bubble's size (area) displayed in the matrix cell is proportional to the number of bilateral coauthorship relationships.&nbsp;This edition particularly focuses on five pivotal parties: the US, China, EU27, the UK and Japan.&nbsp;These parties were specifically selected due to their substantial contributions to work production across all scientific fields from 1971 to 2020.&nbsp;These choices also align with the nations acclaimed as the 'Big 5' science nations&nbsp;(the US, China, Germany, the UK and Japan) in the <a href="https://www.nature.com/articles/d41586-022-00569-7"><i>Nature Index</i></a>.&nbsp;Please note that the Matrix Diagram&nbsp;only takes into account the top 50 institutions in terms of work production for each period and discipline.&nbsp;Therefore, if a cell shows zero (as small dots), it does not necessarily imply the absence of coauthorship relationships for the corresponding bilateral pair.</li><li>Forth,&nbsp;the<i> Interinstitutional Collaboration Dendrogram&nbsp;</i>('<strong>cdend</strong>' folder)&nbsp;elucidates the development and evolution of interinstitutional research collaboration clusters spanning the last five decades. This is accomplished through hierarchical clustering analysis of institutions, considering the top 50 institutions in terms of work production across the four periods: 1971–1990, 1991–2000, 2001–2010 and 2011–2020.<ul><li>The method used for hierarchical clustering analysis is the same as developed in <a href="https://doi.org/10.48550/arXiv.2211.04429">Okamura (2023)</a>. The distance between institutions X and Y is defined as the number of works with nationalities from both X and Y divided by the total number of works with nationalities from at least one of X and Y, subtracted from 1. Hierarchical clustering analysis was performed on the distance matrix using the&nbsp;hclust&nbsp;function implemented in R with the&nbsp;ward.D2&nbsp;option (i.e. the original Ward's method) specified.</li><li>The method of dendrogram visualisation is primarily derived from an example detailed on the&nbsp;<a href="https://cran.r-project.org/web/packages/dendextend/vignettes/dendextend.html">dendextend&nbsp;website</a>. Circular dendrograms were created using the&nbsp;<a href="https://cran.r-project.org/package=dendextend">dendextend</a>&nbsp;and&nbsp;<a href="https://cran.r-project.org/package=circlize">circlize</a>&nbsp;packages in R. As one moves inward from the outer edge of the circle towards its centre, institutions or clusters of institutions that are in closer proximity to each other merge earlier.</li><li>To indicate the country where the institutions are located, the country names are included at the beginning of the terms of research institutions, using the two-letter&nbsp;<a href="https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2">ISO3166-1alpha-2</a>&nbsp;code.&nbsp;The accompanied circularised bar graphs represent the number of works for the institutions&nbsp;involved.&nbsp;<a href="https://ror.org/">ROR</a>s are used as the canonical identifiers of the research institutions. Readers of this booklet in PDF format can click on the ROR-based URL ('https://ror.org/...') in the diagrams to view the corresponding ROR webpage from their browser.</li></ul></li><li>Additionally, for each discipline and the respective periods of 1971–1990, 1991–2000, 2001–2010 and 2011–2020, the top 100 institutions in terms of work production are displayed in tabular format ('<strong>table</strong>' folder), showing their respective country codes and production volumes. If multiple research institutions have equal production volumes during each period, they are organised alphabetically by country codes and then by organisation names. Even if distinct rankings are shown, they lack significance and are treated as ties.</li></ul><p>&nbsp;</p><p><strong>Important Notes</strong></p><ul><li>It is worth reiterating that the data from OpenAlex used to compile&nbsp;the <a href="https://arxiv.org/abs/2308.16810"><i>Atlas of Science Collaboration</i></a>, even when incorporating bibliometric data related to past works, lacks consistent finality. As of the data acquisition for this version (August 2023), OpenAlex encompassed information regarding approximately 240 million works, with an additional influx of about 50,000 new data entries related to works being added daily.&nbsp;Furthermore, for a substantial portion of these works, information regarding the corresponding institution to which the authors belong remains unknown. As a result, should the same analyses as those embedded within this booklet be replicated in the future, although the qualitative extent of change remains uncertain, it is undeniable that quantitatively distinct data will be acquired. Nonetheless, for individuals seeking an understanding of the global scope and evolution of international and interinstitutional collaborative research, the potential availability of this booklet or an enhanced, continuously updated evidence base holds inherent value.</li><li>Further, it is worth reiterating that the term 'works' encompasses a wide variety of scholarly publications. The analyses conducted in the compilation of this booklet do not take into consideration whether these works are peer-reviewed articles or not, nor do they encompass considerations of their prominence, impact or quality. It is emphasised that the primary intent behind the visualisations in this booklet is to quantitatively capture the momentum of scholarly knowledge production outputs from diverse research institutions, and to identify how productive institutions collaborate internationally and interinstitutionally. Caution must be exercised, with acknowledgment that relying solely on the quantity of scholarly output produced by institutions falls short in encompassing discussions about their research potential, contributions to academia, or their relative superiority or inferiority. Further, it is recommended to consider the limitations discussed in <a href="https://doi.org/10.48550/arXiv.2211.04429">Okamura (2023)</a> when using this booklet.</li></ul><p>&nbsp;</p><p><strong>Miscellaneous</strong></p><ul><li>It is important to note that some research institutions may encounter difficulties in accurately assessing the actual production volume at the institutional level within each analysis period due to challenges related to name disambiguation and the influence of historical organisational changes in bibliometric databases.</li><li>For the Interinstitutional Collaboration Dendrograms and the rankings of the top 100 productive institutions, entities like universities and R&amp;D institutions are primarily identified using the nomenclature employed in OpenAlex. However, certain portions have been presented through abbreviations or acronyms, both for illustrative purposes and to effectively accommodate limited space. For instance, 'University of' is abbreviated as 'U.', 'Institution' and 'Institute' as 'Inst', 'National Laboratory' as 'NL', and 'Science' and 'Technology' as 'Sci' and 'Tech', correspondingly, among others. Should readers possess more fitting suggestions for abbreviations specific to particular organisations, or any other ideas aimed at enhancing the content of this booklet, we would greatly appreciate their input.</li></ul><p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Dataset - Survey results - Applying Model-based Requirements Engineering in AIDOaRt Collaborative Project

<p>This dataset and its associated report contain the results of an online survey on using a&nbsp;model-based requirements engineering approach in AIDOaRT project in 2022.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

The World Asellidae database and phylogeny: a collaborative backbone resource for comparative studies of subterranean life evolution

<p>Supplementary material for the article &quot;The World Asellidae database and phylogeny: a collaborative backbone resource for comparative studies of subterranean life evolution&quot;</p> <p>-&nbsp;SI Figure 5: The World Asellidae phylogeny with credibility Intervals for the age of the nodes.&nbsp;Node labels of the phylogeny indicate the 95% credibility intervals of the estimated dates.</p> <p>-&nbsp;SI Table 1: Metadata for the 2093 COI sequences used in the study.</p> <p>- SI Table 4: Alignment of the 2093 COI sequences used for the delimitation of MOTUs.</p> <p>-&nbsp;SI Table 5: Alignment of the 424 COI sequences used for the four-gene dated phylogeny.</p> <p>- SI Table 6: Alignment of the 424 16S sequences used for the four-gene dated phylogeny.</p> <p>- SI Table 7: Alignment of the 424 FASTKD4 sequences used for the four-gene dated phylogeny.</p> <p>-&nbsp;SI Table 8: Alignment of the 424 28S sequences used for the four-gene dated phylogeny.</p> <p>-&nbsp;SI Table 9: Metadata for the DNA sequences used for the 4-gene dated phylogeny.</p> <p>-&nbsp;SI Table 11: Data on body size, sexual body size dimorphism, habitat specialization and habitat size used in comparative analyses.</p> <p>- SI Table 12: Metadata for the DNA sequences deposited in NCBI as part of this study.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Structural Gender Imbalances in Ballet Collaboration Networks

<p>Data contains node list of company artists and their artist type and gender. Edge list provides collaboration network of each company.&nbsp;</p> <p>Null model data provides metrics obtained from null models where&nbsp;assortativity&nbsp;preferences are removed by shuffling collaborations (edges) or artists&#39; attributes (gender) in the collaboration network.</p> <p>For more details, please see documentation in&nbsp;<a href="/api/files/aa096f3e-bc2f-400e-94a9-6bd76ea4a324/Ballet_data_dict.rtf?versionId=7e58ba3a-caae-490b-9da6-9285bd7e4807">Ballet_data_dict.rtf</a></p> <p>For company abbreviations:</p> <p>ABT: American Ballet Theater; NYBC: New York City Ballet; NBC: National Ballet of Canada; ROH: The Royal Ballet of The Royal Opera House.&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Benchmark movement data set for trust assessment in human robot collaboration

<p>In the Drapebot project, a worker is supposed to collaborate with a large industrial manipulator in two tasks: collaborative transport of carbon fibre patches and collaborative draping. To realize data-driven trust assessement, the worker is equipped with a motion tracking suit and the body movement data is labeled with the trust scores from a standard Trust questionnaire (Trust perception scale - HRI, Schaefer 2016).</p> <p>Data has been collected in the transport and draping tasks (counterbalanced) from 20 participants,&nbsp; 7 female and 13 male, average age 25 (SD = 4.0). Average height was 1.74 meters (SD = 0.1). One session consists of 24 trials on average for the transport and draping task resulting in 951 trials across all conditions. For all sessions, body tracking was performed using the Xsens MVN Awinda tracking suit. It consists of a tight-fitting shirt, gloves, headband, and a series of straps used to attach 17 IMUs to the participant. After calibration the system uses inverse kinematics to track and log the movements of the participant at a rate of 60 Hz. The measurements include linear and angular speed, velocity, and acceleration of every skeleton tracking point (see <a href="https://www.xsens.com/hubfs/Downloads/Manuals/MVN_real-time_network_streaming_protocol_specification.pdf">XSENS manual</a> for a detailed description of avaiable measurements).</p> <p><strong>Data organization</strong></p> <p>There are 20 files for 20 participants of each task accordingly (transport and draping). The name of the files is P01SD, where the number 01 is the participant the D stands for draping. Accordingly, P01ST stands for transport. Each file contains all the data that was generated from the XSENS motion capture system. The files are xlsx files and for each sheet inside the excel file there are different types of data:</p> <ul> <li>Segment Orientation - Quat</li> <li>Segment Orientation - Euler</li> <li>Segment Position</li> <li>Segment Velocity</li> <li>Segment Acceleration</li> <li>Segment Angular Velocity</li> <li>Segment Angular Acceleration</li> <li>Joint Angles ZXY</li> <li>Joint Angles XZY</li> <li>Ergonomic Joint Angles ZXY</li> <li>Ergonomic Joint Angles XZY</li> <li>Center of Mass</li> <li>Sensor Free Acceleration</li> <li>Sensor Magnetic Field</li> <li>Sensor Orientation - Quat</li> <li>Sensor Orientation - Euler</li> </ul> <p>See also: <a href="https://base.movella.com/s/article/Output-Parameters-in-MVN-1611927767477?language=en_US">https://base.movella.com/s/article/Output-Parameters-in-MVN-1611927767477?language=en_US</a></p> <p>For more information on each specific data and/or sensors please see the xsens manual (Link above)</p> <p><strong>Data Annotation</strong></p> <p>For each procedure there is an annotation file called sorted_draping.xlsx and sorted_transport.xlsx. In these files the first column is the frame and from column 2 until column 21 are the annotations for each procedure for each participant. The annotations describe the different phases during the procedures for each data frame recorded by xsens:</p> <ul> <li>Transport phases: pick, transport, drop, return</li> <li>Draping phases: approach, draping, return</li> </ul> <p>The file trustscores.xlsx includes some demographic data as well as the results of the trust questionaire for each participant and each task, including the scores for the individual items as well as the calculated trust score. The different columns are:</p> <ul> <li>Subject: participant number for crossreferencing with annotation and movement data</li> <li>Transport.Speed: denoting the robot speed (fast or slow)</li> <li>Age: age of the participant</li> <li>Gender: gender of the participant</li> <li>DominantHand: dominant hand of the participant (left or right)</li> <li>Height: height of the participant</li> <li>Score for answers of the participant in related questions category.</li> </ul> <p>This is followed by the trust questionaire items:</p> <ul> <li>Which % of time does the robot <ul> <li>Function successfully</li> <li>Act consistently</li> <li>Communicate with people</li> <li>Provide feedback</li> <li>Malfunction</li> <li>Follow directions</li> <li>Meet the needs of the mission</li> <li>Perform exactly as instructed</li> <li>Have errors</li> </ul> </li> <li>Which % of the time is the robot: <ul> <li>Unresponsive</li> <li>Dependable</li> <li>Reliable</li> <li>Predictable</li> </ul> </li> </ul> <p>The last two columns are</p> <ul> <li>TrustScore &ndash; Final trust score calculated from all questions</li> <li>Task &ndash; Which task is being performed (Transport/Draping)</li> </ul>

opencc-by-4.0Oct 2023View details →
edi44/100

Coastal SEES Collaborative Research: Coastal Sustainability: A cross-site comparison of salt marsh persistence in response to sea-level rise and feedbacks from social adaptations

Coastal ecosystems are often valued for decision-making purposes based on monetized market and non-market values of goods and services, and associated economic impacts. Examples include values of fishery landings, price changes for waterfront homes, and tourism revenues. Monetized quantities such as these do not provide a comprehensive characterization of the values provided by these ecosystems. Human reliance on the goods and services provided by ecosystems and the global decline in the health of many of these ecosystems suggests the need for ecosystem valuation to help inform decision-making and conservation policy. However, traditionally employed economic valuation methods are rarely able to capture the full scope of the benefits ecosystems provide, including benefits provided by "cultural" ecosystem services. Qualitative methods such as focus groups can provide insight on these values not available through quantitative methods alone. This research explores public perceptions of salt marsh value through the use of semi-structured focus groups in marsh-adjacent communities in Massachusetts, Virginia, and Georgia. The data include de-identified focus group transcripts from three 90-minute focus groups held in each state. Initial questions were drawn from the same semi-structured question list in each focus group, with exploratory follow-up questions based on participant responses. Results of text analysis suggest that in case study communities, outdoor experiences in salt marshes inspire serenity in Massachusetts, influence shore identities in Virginia, and promote stewardship cultivation in Georgia. Perceived threats to these benefits, such as the threat of residential development, industrial pollution, and increasing flood risk, together constitute the context for various community responses related to marsh protection. Results supplement information from extant economic valuations and show the importance of utilizing diverse methods to elicit information on soci

openCustomJun 2017View details →
zenodo40/100

Agile Accelerator Program: From Industry-Academia Collaboration to Effective Agile Training

<p>The agile accelerator program takes place in a Brazilian technology park, as a collaboration between a university and a world-renowned technology company, specialized in agile development and consulting. This partnership has 8-year long with the main goal of preparing undergraduate students to work in high-performance agile teams. This partnership created a culturally rich environment for student learning while influencing other companies to follow the same initiative within this technology park. We conducted a Case Study aiming to characterize this partnership (explaining how it works) and the resulting program, understanding the benefits to the program students. Our results point out the importance of the kind of partnership that provides an immersive learning environment to students, where students can learn empirically, with real projects and real stakeholders and how important it was for the program&#39;s former students to enter the job market. This successful enhanced students&#39; training program on agile software development through the blending of culture between institutions can be of inspiration to those interested in aiming to bridge the gap between academia and industry.</p>

opencc-by-4.0Oct 2020View details →
zenodo40/100

The Software Sustainability Institute's Collaborations Workshop 2015 (CW15) attendees computational tools dataset

<p>Contains the question, raw data, and cleaned data for producing the most used software word cloud for those who attended the Software Sustainability Institute&#39;s Collaborations Workshop 2015 (CW15) held at the Oxford e-Research Institute, Oxford, UK from 25-27 March 2015</p>

opencc-by-4.0Jul 2015View details →
zenodo40/100

Inbred Strain Variant Database (ISVdb): A repository for probabilistically informed sequence differences among the Collaborative Cross strains and their founders

<p>Data files for the development of a database for storing (and a GUI for retrieving) the imputed variants for 72 Collaborative Cross strains of mice. Files include the inputs for the imputation, as well as the final results. See File_S1_Readme for more details on the included files.</p>

opencc-by-4.0Apr 2017View details →
zenodo40/100

The Yelp Collaborative Knowledge Graph

<p>This is the&nbsp;The Yelp Collaborative Knowledge Graph (YCKG) - a transformation of the Yelp Open Dataset into RDF format using Y2KG.&nbsp;</p> <p>The full YCKG dataset can be found in <code>yelp.ttl.gz</code> and <code>yckg.tar.xz</code></p> <p><strong>Paper Abstract</strong></p> <p>The Yelp Open Dataset (YOD) contains data about businesses, reviews, and users from the Yelp website and is available for research purposes. This dataset has been widely used to develop and test Recommender Systems (RS), especially those using Knowledge Graphs (KGs), e.g., integrating taxonomies, product categories, business locations, and social network information. Unfortunately, researchers applied naive or wrong mappings while converting YOD in KGs, consequently obtaining unrealistic results. Among the various issues, the conversion processes usually do not follow state-of-the-art methodologies, fail to properly link to other KGs and reuse existing vocabularies. In this work, we overcome these issues by introducing Y2KG, a utility to convert the Yelp dataset into a KG. Y2KG consists of two components. The first is a dataset including (1) a vocabulary that extends Schema.org with properties to describe the concepts in YOD and (2) mappings between the Yelp entities and Wikidata. The second component is a set of scripts to transform YOD in RDF and obtain the Yelp Collaborative Knowledge Graph (YCKG). The design of Y2KG was driven by 16 core competency questions. YCKG includes 150k businesses and 16.9M reviews from 1.9M distinct real users, resulting in over 244 million triples (with 144 distinct predicates) for about 72 million resources, with an average in-degree and out-degree of 3.3 and 12.2, respectively.</p> <p><strong>Links</strong></p> <p>Latest GitHub release:&nbsp;<a href="https://github.com/MadsCorfixen/The-Yelp-Collaborative-Knowledge-Graph">https://github.com/MadsCorfixen/The-Yelp-Collaborative-Knowledge-Graph/releases/latest</a></p> <p>PURL domain:&nbsp;<a href="https://purl.prod.archive.org/domain/yckg">https://purl.archive.org/domain/yckg</a></p> <p><strong>Files</strong></p> <ul> <li>Graph Data Triple Files <ul> <li><code>yelp.ttl.gz</code> full dataset</li> <li><code>yckg.tar.xz</code> full dataset</li> </ul> </li> <li>One sample file for each of the Yelp domains (Businesses, Users, Reviews, Tips and Checkins),&nbsp; each&nbsp;containing 20 entities. <ul> <li><code>yelp_schema_mappings.nt.gz</code>&nbsp;containing the mappings from Yelp categories to Schema things.</li> <li><code>schema_hierarchy.nt.gz</code>&nbsp;containing the full hierarchy of the mapped Schema things.</li> <li><code>yelp_wiki_mappings.nt.gz</code>&nbsp;containing the mappings from Yelp categories to Wikidata entities.</li> <li><code>wikidata_location_mappings.nt.gz</code>&nbsp;containing the mappings from Yelp locations to Wikidata entities.</li> </ul> </li> <li>Graph Metadata Triple Files <ul> <li><code>yelp_categories.ttl</code>&nbsp;contains metadata for all Yelp categories.</li> <li><code>yelp_entities.ttl</code>&nbsp;contains metadata regarding the dataset</li> <li><code>yelp_vocabulary.ttl</code>&nbsp;contains metadata on the created Yelp vocabulary and properties.</li> </ul> </li> <li>Utility Files <ul> <li><code>yelp_category_schema_mappings.csv</code>. This file contains the 310 mappings from Yelp categories to Schema types. These mappings have been manually verified to be correct.</li> <li><code>yelp_predicate_schema_mappings.csv</code>. This file contains the 14 mappings from Yelp attributes to Schema properties. These mappings are manually found.</li> <li><code>ground_truth_yelp_category_schema_mappings.csv</code>. This file contains the ground truth, based on 200 manually verified mappings from Yelp categories to Schema things. The ground truth mappings were used to calculate precision and recall for the semantic mappings.</li> <li><code>manually_split_categories.csv</code>. This file contains all Yelp categories containing either a &amp; or /, and their manually split versions. The split versions have been used in the semantic mappings to Schema things.</li> </ul> </li> </ul>

opencc-by-4.0May 2023View details →
zenodo40/100

PROCRAFT Final Meeting - Collaborative work between institutions and volunteers by Marie Grima - National Museum of Flight

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo40/100

PROCRAFT Final Meeting - Case study: conservation of the Dornier 17 wreck (collaboration between institutes and volunteers) by Darren Priday, RAF Museum

Open the record for dataset details and reuse information.

opencc-by-4.0Dec 2023View details →
zenodo40/100

Figure 1 in The TeaComposition initiative: Unleashing the power of international collaboration to understand litter decomposition

Figure 1. TeaComposition initiative. (A) An illustration to the asymptotic model (cf Berg 2014) for decomposition of standard tea bags (0.25mm mesh) of Green (Camelia sinensis) and Rooibos (Asphalantus linearis) tea, representing litter decomposition (here as mass remaining) over a period of three years. Dashed horizontal line shows the limit value (stabilized residue); (B) site distribution in terrestrial (red dots) and aquatic (blue dots) ecosystems across nine world biomes in 2020; (C) Number of participating sites per biome; (D) Networking activities of the initiative include active and potential collaborations with the following global research networks: Soil Biodiversity Observation Network (SoilBON), International Co-operative Programme on Assessment and Monitoring of Air Pollution Effects (ICP), Greenhouse gas inventory (GHG), Sustainable Development Goals (SDGs), Tree Diversity Network (TreeDivNet), Detrital Input and Removal Treatments (DIRT), and Terrestrial Environmental Observatories (TERENO).

opencc-by-4.0Dec 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record