Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
794
datasets available to search
ShareScore release 0.7.1
Dataset results
794 results for “publishing”
Frictionless Tabular data package for GC-MS data from Rose Genome article published in Nature genetics, June, 2018
<p>This dataset, in the form of a Frictionless Tabular Data Package (https://frictionlessdata.io/specs/tabular-data-package/), holds the measurements of 61 known metabolites (all annotated with resolvable CHEBI identifiers and InChi), measured by gas chromatography mass-spectrometry (GC-MS) in 6 different Rose cultivars (all annotated with resolvable NCBITaxId) and 3 organism parts (all annotated with resolvable Plant Ontology identifiers). The data was extracted from a supplementary material table, available from https://static-content.springer.com/esm/art%3A10.1038%2Fs41588-018-0110-3/MediaObjects/41588_2018_110_MOESM3_ESM.zip and published alongside the Nature Genetics manuscript identified by the following doi: https://doi.org/10.1038/s41588-018-0110-3, published in June 2018. This dataset is used to demonstrate how to make data Findeable, Accessible, Discoverable and Interoperable(FAIR) and how Tabular Data Package representations can be easily mobilized for re-analysis and data science. It is associated to the following project available from github at: https://github.com/proccaserra/rose2018ng-notebook with all necessary information and Jupyter notebooks.</p>
Publishing Reproducible Research Outputs - Interviewees and interview questions
<p>The table '<strong>Interview questions</strong>' shows the focus of our investigation and stakeholder engagement activities. It should be noted that not all interview questions were asked to all stakeholder groups based on appropriateness and time available. Some questions in the table may appear to be repeated: this is because slightly different phrasing was used based on the stakeholder interviewed.</p> <p>Legend:</p> <ul> <li>Research Funding Organisations: RFO</li> <li>Research Performing Organisations: RPO</li> <li>Infrastructure Providers: IP </li> <li>Academic Publishers: AP</li> <li>Researchers and research groups: RRG</li> </ul> <p>The table '<strong>List of interviewees</strong>' includes all stakeholders engaged in the context of this research.</p>
3D and assay data published in "XRF and 3D modelling on a composite Etruscan helmet"
<p>The data presented here are published as part of the publication Emmitt, J.J., McAlister, A., Bawden, N., and J. Armstrong "XRF and 3D modelling on a composite Etruscan helmet" <em>Applied Sciences</em>. <em>11</em>(17): 8026. DOI: 10.3390/app11178026. The methodology for the creation of the photogrammetry model is presented Emmitt et al. (2021a), and further information about the methods used to collect the pXRF data can be found in Emmitt et al. (2021b). The interpolation analysis is done using PyVista by Sullivan and Kaszynski (2019)</p> <p>The model is are published as a .ply file, the assay data is in a csv file with the corresponding location on the model, and a Juypter notebook for running the analysis. The PyVista Python package will be required (Sullivan and Kaszynski 2019). Contained here are:</p> <ul> <li>Negau Helmet, Doug Gold Collection - 1x .ply</li> <li>Helmet assay points and data - 1x .csv</li> <li>Juypter Notebook - 1x .ipynb</li> </ul> <p>Data are published with permission of Museo Nazionale Etrusco di Villa Giulia e Villa Poniatowski di Roma (Director Valentino Nizzo).</p>
Publishing Reproducible Research Outputs - Thematic coding of interview findings
<p>The spreadsheet in the present dataset (CSV format) includes the anonymised thematic coding that has been applied to our interview findings. A list of interviewees and interview questions is available <a href="https://doi.org/10.5281/zenodo.5141665">here</a>.</p> <p>The thematic coding has been applied by using <a href="https://www.qsrinternational.com/nvivo-qualitative-data-analysis-software/home">NVivo</a>, a professional qualitative analysis software, and then exported in spreadsheet form for public sharing. The findings of this analysis have been used to inform our final report, which is available in our <a href="https://zenodo.org/communities/ke-prro/?page=1&size=20">Zenodo project Community</a>.</p>
Dataset for the published article "Quantum version of the integral equation theory based dielectric scheme for strongly coupled electron liquids"
<p>The data contained in the zip file constitute the main research data of the article entitled as "<em>Quantum version of the integral equation theory based dielectric scheme for strongly coupled electron liquids</em>", published in the Journal of Chemical Physics as a Communication. In this article, a novel dielectric scheme is proposed for strongly coupled electron liquids that handles quantum mechanical effects beyond the random phase approximation level and treats electronic correlations within the integral equation theory of classical liquids. This self-consistent scheme features a complicated dynamic local field correction functional and yields unprecedently accurate results for the static structure factor without featuring any adjustable or empirical parameters.</p> <p>In particular, the datasets contain the static structure factors of the paramagnetic electron liquid as computed by four schemes of the self-consistent dielectric formalism and as extracted from state-of-the-art path integral Monte Carlo (PIMC) simulations. The dielectric schemes of interest are all tailor-made for the strongly coupled regime of the finite temperature uniform electron fluid (UEF; also known as jellium or quantum one-component plasma). These are the newly proposed quantum version of the integral equation theory based scheme (qIET), the newly proposed quantum version of the hypernetted-chain based scheme (qHNC), the integral equation theory based scheme (IET) [1,2] and the hypernetted-chain based scheme (HNC) [3,4].</p> <p>The static structure factors are provided for 20 paramagnetic UEF state points defined by (r<sub>s</sub>,Θ)={(50,0.50),(60,0.50),(70,0.50),(80,0.50),(90,0.50),(100,0.50),(100,0.75),(100,1.00),(100,2.00),(100,4.00),(110,0.50),(125,0.50),(125,0.75),(125,1.00),(125,1.50),(125,2.00),(150,0.50),(150,1.00),(200,0.50),(200,1.00)} where r<sub>s</sub> is the quantum coupling parameter and Θ is the degeneracy parameter. </p> <p>In the qIET, qHNC, IET and HNC datasets; the first column corresponds to the wavenumber normalized to the Fermi wavenumber and the second column corresponds to the static structure factor value. In the PIMC datasets, the first column corresponds to the wavenumber multiplied by the first Bohr radius, the second column corresponds to the static structure factor value and the third column corresponds to the associated error bars.</p> <p>[1] P. Tolias, F. Lucco Castello and T. Dornheim, J. Chem. Phys. 155, 134115 (2021).<br> [2] F. Lucco Castello, P. Tolias and T. Dornheim, EPL 138, 44003 (2022).<br> [3] S. Tanaka, J. Chem. Phys. 145, 214104 (2016).<br> [4] T. Dornheim, T. Sjostrom, S. Tanaka and J. Vorberger, Phys. Rev. B 101, 045129 (2020).</p>
Quality evaluation criteria, best practices, and assessment systems for Institutional Publishing Service Providers (IPSPs): dataset
<p>The dataset contains tabular information on the elements of best practice in scholarly publishing found in a set of documents (high-level recommendations and principles, indexation criteria and specific assessment guidelines used on the national and institutional levels). The set of documents subject to analysis (58 items) were identified by the DIAMAS project team members (bibliographic metadata are provided in IPSP-best-practice-documents.xml and IPSP-best-practice-documents.ris).</p> <p>The dataset was compiled by the DIAMAS project team using an analysis matrix that included the general information about the documents (title, issuing entity, scope and purpose, etc.) and the the seven core components of scholarly publishing identified in the Diamond Open Access Action Plan (2022) and revised by the DIAMAS project team.</p> <p>More information about the data collection methodology can be found in the report D3.1 IPSP Best Practices Quality evaluation criteria, best practices, and assessment systems for Institutional Publishing Service Providers (IPSPs) (<a href="https://doi.org/10.5281/zenodo.7859172">https://doi.org/10.5281/zenodo.7859172</a>), which is based on this dataset.</p> <p> </p> <p><strong>****Dataset contents****</strong></p> <p>IPSPs_best-practices-overview.csv</p> <p>IPSPs_best-practices-overview.ods</p> <p>IPSP-best-practice-documents.xml</p> <p>IPSP-best-practice-documents.ris</p> <p>README.txt</p> <p> </p> <p><strong>****Column headers and field types***</strong></p> <p>Title (original) (text)</p> <p>Title (English) (text)</p> <p>Publication date (date, DD/MM/YY)</p> <p>Last accessed (date, DD/MM/YY)</p> <p>URL (text-web address)</p> <p>Scope (text, controlled)</p> <p>Type of document (text, controlled)</p> <p>Original language (text)</p> <p>Other languages (text)</p> <p>Entity issuing the document (text)</p> <p>Entity responsible for the assessment (text)</p> <p>Scope of the assessment (text, controlled)</p> <p>Scope of assessment: region or country (text)</p> <p>Disciplines’ coverage (text)</p> <p>Periodicity of the assessment (text)</p> <p>Reassessment frequency? If yes: periodicity (text)</p> <p>Benefits linked to the assessment (text)</p> <p>(1) Funding (text)</p> <p>(2) Ownership and governance (text)</p> <p>(3) Open science practices (text)</p> <p>(4) Editorial quality, editorial management and research integrity (text)</p> <p>(5) Technical service efficiency (text)</p> <p>(6) Visibility (including indexation), communication, marketing and impact (text)</p> <p>(7) Diversity, Equity and Inclusion (text)</p>
Raman Spectra of K2ReCl6 and K2SnCl6, published in PRB 107, 214301 (2023)
<p>Raman Spectra of K2ReCl6 from 5 K to room temperature, as well as K2SnCl6 at room temperature, in c(aa)c' and c(ab)c' configuration. Spectra shown and discussed in Phys. Rev. B <strong>107 </strong>214301 (2023), also available as preprint https://arxiv.org/abs/2209.05866. </p>
State-of-the-art review of near-term freshwater forecasting literature published between 2017 and 2022
This data publication includes code and results from a systematic literature review on the current state of near-term forecasting of freshwater quality. The review aimed to address the following questions: (1) Freshwater variables, scales, models, and skill: Which freshwater variables and temporal scales are most commonly targeted for near-term forecasts, and what modeling methods are most commonly employed to develop these forecasts? How is the accuracy of freshwater quality forecasts assessed, and how accurate are they? How is uncertainty typically incorporated into water quality forecast output? (2) Forecast infrastructure and workflows: Are iterative, automated workflows commonly employed in near-term freshwater quality forecasting? How are forecasts validated and archived? (3) Human dimensions: What is the stated motivation for development of most near-term freshwater quality forecasts, and who are the most common end users (if any)? How are end users engaged in forecast development? An initial search was conducted for published papers presenting freshwater quality forecasts from 1 January 2017 to 17 February 2022 in the Web of Science Core Collection. Results were subsequently analyzed in three stages. First, paper titles were screened for relevance. Second, an initial screen was conducted to assess whether each paper presented a near-term freshwater quality forecast. Third, papers that passed the initial screen were analyzed using a standardized matrix to assess the state of near-term freshwater quality forecasting and identify areas of recent progress and ongoing challenges. Additional details regarding the systematic literature search and review are presented in the Methods section of the metadata.
A database of published mangrove articles for coastal Louisiana, USA
Mangroves are being increasingly recognized as natural climate solutions for the range of ecosystem services they provide. In North America, one of the northern range limits of mangroves is found in coastal Louisiana, USA, where in recent decades, mangroves have been expanding into wetlands formerly dominated by salt marsh primarily due to decreases in the frequency and severity of winter freeze events. While reviews focused on mangrove ecology that include coastal Louisiana within a broader geographic scope have been conducted, no systematic review has focused on what is known about mangrove ecology across coastal Louisiana, a region that contains the expansive Mississippi River Delta. To fill this knowledge gap, we conducted a systematic review to highlight the breadth of mangrove research topics that have been studied in coastal Louisiana and identify emerging and future research opportunities. We identified four main research topics: (1) mangrove expansion, (2) freeze tolerance, (3) coastal restoration, and (4) disturbance. We also identified geographic biases in where mangrove research has been conducted, with a focus around the heavily industrialized Port Fourchon/Grand Isle area.
Animal Gut Microbiome (AGM) Data from 91 Published Studies for 224 Animal Species
Diversity and heterogeneity often are conflated but are fundamentally different. An aphorism proposed by Shavit and Ellison (2021; J. Phil. 118: 525–548) for distinguishing them is that “a zoo is diverse whereas an ecosystem is heterogeneous.” That is, a zookeeper measuring diversity simply enumerates the different types of animals; interactions are not expected to occur between animals separated by fences or other barriers. In contrast, measures of heterogeneity ought to include both interspecific interactions and relationships between species and their heterogeneous habitats. Here, we use cross-scale, dual scaling-law analyses of heterogeneity and diversity of animal gut microbiomes (AGMs) to address three objectives: (i) estimate the spatial heterogeneity and diversity of animal-gut microbiomes; (ii) analyze influences of phylogeny and diets on scaling of diversity and heterogeneity; (iii) explore mechanistic differences between diversity and heterogeneity in AGMs. From 4903 AGM samples collected from 318 animal species covering all six classes of vertebrates and four major classes of invertebrates, we estimated that ≈640,000 operational taxonomic units (OTUs or “species”) make up the pool of microbial species that could inhabit animal guts, among which ≈8000 are relatively common and ≈800 are dominant. The gut of any single animal, however, includes only 0.01–0.5% of the total species pool. We extended Ma’s diversity-area relationship for scaling diversity and extend Taylor’s Power Law and Luna et al.’s (2020; Diversity 12: 86) interaction diversity for scaling heterogeneity. At the community scale, phylogeny significantly influenced heterogeneity, but diets did not. Phylogeny and diets had limited influence on diversity at both community and landscape scales. Although two common measures of diversity—beta diversity and unevenness—commonly are synonymized with heterogeneity, our data lead us to conclude that diversity and heterogeneity measure two very different
Publishing reproducible logbooks explainer comic strip
<p>This comic strip explains at a high level how to publish reproducible<br> notebooks using tools and services such as Jupyter and Binder.</p> <p>Files:<br> - reproducible_logbook.png: main picture<br> - reproducible_logbook_scenario.png: zoom on the scenario part of the picture<br> - reproducible_logbook.kra: original Krita source file<br> - reproducible_logbook_texts.svg: svg export (just the texts)<br> - reproducible_logbook_wo_text.png: png export without the texts (e.g. for translations)</p>
Prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, number of journals per country and publisher
<p>An analysis on the prevalence of Creative Commons licenses in the Directory of Open Access Journals by discipline, author fees, country and publisher according to the number of journals.</p>
Data for "Formation of very large 'blocky alpha' grains in Zircaloy-4" by V. Tong and T.B. Britton published in Acta Materialia (2017)
<p>Data for "Formation of very large ‘blocky alpha’ grains in Zircaloy-4"</p> <p>Vivian S Tong, T Ben Britton<br> Department of Materials, Imperial College London, Prince Consort Road, London, SW7 2AZ, UK</p> <p>For more information please contact: b.britton@imperial.ac.uk (Ben Britton)</p> <p>---</p> <p>Figures_data.xlsx contains the data for line graphs in the following figures on separate labelled sheets:<br> Figure 2(a)<br> Figure 2(b)<br> Figure 4(c)<br> Figure 6.</p> <p>Figures_data.xlsx also contains the HR-EBSD GND density data in Figures 3(b) and 3(d), which have been plotted on a log10 colour scale in the published figure.</p> <p>The EBSD orientation data have been exported as text files (.ctf) directly from Bruker Esprit 2.1 software.</p> <p>Orientations are described using Bruker EBSD software conventions, described in the paper "Tutorial: Crystal orientations and EBSD — Or which way is up?" by Britton et al.(http://dx.doi.org/10.1016/j.matchar.2016.04.008).</p> <p><br> EBSD data is provided for the following figures:<br> Figure 3(a)<br> Figure 3(c)<br> Figure 4(b), Figure 5(c), Figure 7(c) -- these are all the same dataset<br> Figure 5(b)<br> Figure 6 - EBSD maps of these two datsets were not shown, but this is the raw data from which twin fractions were calculated.<br> Figure 7(a)<br> Figure 7(b)</p>
ReliSA/dataset_optimal-set-ilp-2015-07: Published results
<p>The dataset as used for the results reported in the paper "<a href="https://doi.org/10.1007/978-3-662-49192-8_37">Jakub Danek, Premek Brada: Finding Optimal Compatible Set of Software Components Using Integer Linear Programming</a>. SOFSEM 2016: 457-468" (https://link.springer.com/chapter/10.1007%2F978-3-662-49192-8_37).</p>
Understanding the Publish-Review-Curate (PRC) Model of Scholarly Communication - Data and Code
<p>Summary data for the number of articles submitted to publish-review-curate platforms as of August 2024 (Figure 1) [Update 14 Nov 2024: Added JMIRx. Data still from August 2024]</p> <p>Summary data for the number of articles reviewed by review platforms (Figure 2)</p> <p>Analysis code to produce Figures 1 and 2</p> <p>Code to extract articles for inclusion in data</p>
Dataset of behavioral and neurophysiological data of a virtual sailing task published in: "Providing task instructions during motor training enhances performance and modulates attentional brain networks"
<p>Dataset belonging to the behavioral and neurophysiological data of the publication: "Providing task instructions during motor training enhances performance and modulates attentional brain networks". The two uploaded Zip files contain kinematic and electroencephalographic data of 36 participants for the Obstacle and HorizonTask.</p>
Brazilian papers on Building Information Modeling published until 2021
<p>Characterization of Brazilian research in BIM through articles published in Brazilian journals: PARC Research in Architecture and Construction, Project Management and Technology and Built Environment. It covered articles published until 07/2021.</p> <p>The characterization of the articles is by the fields:<br> ITEM: article number in the survey;<br> JOURNAL: Name of the Brazilian journal between Ambiente Construéido, Gestão & Tecnologia de Projetos and PARC Research in Architecture and Construction;<br> TITLE: Title of the article;<br> YEAR: year of publication;<br> INSTITUTION: name by extension of the Institution of the main author;<br> TYPE OF INSTITUTION: values between Public Education Institution, Private Education Institution, Private Company, State Public Company …;<br> LOCATION: Federative Unit of the main author's institution;<br> USE CATEGORY: chosen from MODEL USE SERIES in Category II as in https://bimexcellence.org/wp-content/uploads/211in-Model-Uses-Table.pdf ;<br> SUBCLASSIFICATION: chosen from MODEL USE in Category II as in https://bimexcellence.org/wp-content/uploads/211in-Model-Uses-Table.pdf;<br> BIBLIOGRAPHIC REFERENCE: ABNT format ;<br> LINK: url to the full article.</p>
A dataset of published journal papers using neural networks for seismological tasks.
<p>This is a dataset of 637 journal papers applying neural networks for various tasks in seismology spanning January 1988 to January 2022. The dataset mainly includes peer reviewed papers and does not contain duplicated works. It follows a hierarchical classification of papers based on seismological tasks (i.e. category, sub_category_I, sub_category_II, task, and sub_task). For each paper following information are provided: 1) first author's last name, 2) publication year, 3) paper's title, 4) journal 's name, 5) machine learning method used, 6) the type of used neural network, 7) the name of neural network architecture, 8) the number of neurons/kernels in each hidden layer, 9) type of training process, i.e. supervised, semi-supervised, etc, 10) input data into the network, 11) output data, 12) data domain, i.e. time, frequency, feature, etc, 13) the type of data used for training, e.g. synthetic or real data, 14) the size of training set, 15) the metrics used to measure the performance, 16) performance scores, 17) the baseline method used for evaluation, and 18) a short note summarizing the paper's objective, its approach, and its significance. </p> <p>An updating version of the dataset can be find from here: https://smousavi05.github.io/dl_seismology/ and here:https://github.com/smousavi05/dl_seismology/tree/main/docs. </p> <p>An updating glossary of seismological tasks and relevant machine learning techniques and papers are provided here: https://smousavi05.gitbook.io/mlseismology/</p>
Kerosene freeze data for paper to be published:
<p>This dataset originates in an oil refinery producing, among other products, kerosene. The freeze point of the kerosene is an important specification. In the paper to be published the authors use data quality assessment methods to define periods of the data suitable for the derivation of an inferential model.</p>
Monitoring open access publishing of NWO funded research (2015-2021) data set
<p>This is the dataset underlying the report "Monitoring open access publishing of NWO funded research" (<a href="https://doi.org/10.5281/zenodo.7041897">https://doi.org/10.5281/zenodo.7041897</a>)</p> <p>The report presents statistics on the extent to which publications from the period 2015–2021 funded by NWO are available in Open Access. The analyses presented in this report also cover publications funded by the Netherlands Organisation for Health Research and Development ZonMw. This report builds on two earlier reports, published in <a href="https://zenodo.org/record/4446042">2020</a> and <a href="https://zenodo.org/record/5056043">2021</a>, covering publications from the period 2015–2018 and 2015-2020, respectively.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.