Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,499
datasets available to search
ShareScore release 0.9.0
Dataset results
13,499 results for “researcher”
Research data for "[2.2.2.2]Paracyclophanetetraenes (PCTs): cyclic structural analogues of poly(p-phenylene vinylene)s (PPVs)"
<p>This dataset contains the underlying experimental (<sup>1</sup>H NMR, <sup>13</sup>C NMR, <sup>31</sup>P NMR, high-resolution mass spectrometry (HRMS), UV-vis absorption, photoluminescence (PL), cyclic voltammetry) and computational (molecular geometries, input/output files for Q-Chem, Gaussian, and TheoDORE) research data for the article “[2.2.2.2]Paracyclophanetetraenes (PCTs): cyclic structural analogues of poly(p-phenylene vinylene)s (PPVs)”.</p> <p>Content (names of folders and files are given in <strong>bold face</strong>):</p> <p>The folder for each molecule (<strong>Br-P2</strong>, <strong>O-3</strong>, <strong>O-P1</strong>, <strong>O-PCT</strong>, <strong>PCT</strong>, <strong>Quinine</strong>, <strong>S-2</strong>, <strong>S-3</strong>, <strong>S-P1</strong>, <strong>S-P2</strong>, <strong>S-PCT</strong>) contains the molecular structure as .mol and .cdxml file, as well as subfolders for the <strong>Experimental </strong>research data and, if available, the <strong>Computational </strong>research data.</p> <p>The <strong>Experimental </strong>research data<strong> </strong>folder for each molecule contains the measurement files. The file names are composed of the acronym of the measured molecule, the type of measurement, and further details about the measurement (if needed to distinguish the files). Measurements were performed as described in the article.</p> <p>Computations were performed on the molecules <strong>O-PCT</strong> and <strong>S-PCT</strong> as described in the article.<br> The <strong>Computational </strong>research data folder for each molecule contains subfolders for geometry optimisations of the individual electronic states and rotamers:</p> <ul> <li><strong>c2_neut</strong>: neutral singlet state for C<sub>2</sub> rotamer</li> <li><strong>c2_trip</strong>: T<sub>1</sub> state of C<sub>2</sub> rotamer</li> <li><strong>c2_2P</strong>: charged state (2+) of C<sub>2</sub> rotamer</li> <li><strong>c2_2M</strong>: charged state (2-) of C<sub>2</sub> rotamer</li> <li><strong>cs_neut</strong>: neutral singlet state for C<sub>s</sub> rotamer</li> <li>etc.</li> </ul> <p>Additional content:</p> <ul> <li><strong>TS</strong>: Full transition state optimisation</li> <li><strong>TS_constrained</strong>: constrained transition state optimisation</li> <li><strong>VIST</strong>: Data for VIST plots</li> <li><strong>FROZEN</strong>: Computations for frozen PCT structure (denoted <strong>O/S-PCT</strong>@<strong>PCT</strong> in the article)</li> </ul> <p>Each optimisation folder contains the following:</p> <ul> <li><strong>qchem.[in,out]</strong>: input/output for geometry optimisation</li> <li><strong>final.xyz</strong>: optimised geometry</li> <li><strong>SOLV/qchem.[in,out]</strong>: input/output for solvated single-point computation</li> <li><strong>tNICS/gaussian.[com,log]</strong>: input/output for NMR shielding tensors</li> </ul>
Data quality assurance at research data repositories: Survey data
<p>This dataset documents findings form a survey on the status quo of data quality assurance practices at research data repositories.</p> <p>The personalized online survey was conducted among repositories indexed in re3data in 2021. It covered the scope of the repository, types of data quality assessment, quality criteria, responsibilities, details of the review process, and data quality information, and yielded 332 complete responses.</p> <p>The dataset comprises a documentation file, the data file, a codebook, and the survey instrument.</p> <p>The <strong>documentation file</strong> (documentation.pdf) outlines details of the survey design and administration, survey response, and data processing. The <strong>data file</strong> (01_survey_data.csv) contains all 332 complete responses to 19 survey questions, fully anonymized. The <strong>codebook</strong> (02_codebook.csv) describes the variables, and the <strong>survey instrument</strong> (03_survey_instrument.pdf) comprises the questionnaire that was distributed to survey participants.</p>
Virtual Research Environments Ethnography: a Preliminary Study
<p>Datasets accompanying the paper “Virtual Research Environments Ethnography: a Preliminary Study”, a systematic mapping study on the literature about Science gateways, Virtual Research Environments, and Virtual Laboratories.</p> <p>While for legal reasons we can not share the original datasets obtained by querying the databases, since they include copyrighted data, we can share the two datasets derived from the query results and the two topic modelling datasets.</p> <p>The dataset “<strong>main_dataset.csv</strong>” consists of the merged query results from ACM Digital Library, IEEEXplore, ScienceDirect, Scopus, and SpringerLink databases. It is structured into six columns: (i) doi; (ii) title; (iii) content_type; (iv) publication year; (v) keyword_search; (vi) DB.</p> <p>The ‘<strong>doi</strong>’, ‘<strong>title</strong>’, and ‘<strong>publication_year</strong>’ labels are self-describing, and are used for the DOIs, titles, and publication years (in the yyyy format) respectively.</p> <p>The ‘<strong>content_type</strong>’ label refers to the different and normalised typologies of resources: (a) Article; (b) Book, (c) Book Chapter; (d) Chapter; (e) Chapter ReferenceWorkEntry; (f) Conference Paper; (g) Conference Review; (h) Early Access Articles; (i) Editorial; (j) Erratum; (k) Letter; (l) Magazines; (m) Masters Thesis; (n) Note; (o) Ph.D. Thesis; (p) Retracted; (q) Review; (r) Short Survey; (s) Standards. (c) and (d) refer to the same type of entry (they are used in different databases), while in the case of (e) we observed that it is used in the Springer database to refer mainly to encyclopaedic entries.</p> <p>The ‘<strong>keyword_search</strong>’ label is used for identifying the keyword group used for formulating the query: (a) science gateway | scientific gateway; (b) virtual laboratory | Vlab; or (c) virtual research environment.</p> <p>The ‘<strong>DB</strong>’ label indicates the provenance of the entries from one of the five databases we selected for our study: (a) ACM; (b) IEEE; (c) ScienceDirect; (d) scopus; and (e) Springer, identifying the ACM Digital Library, IEEEXplore, ScienceDirect, Scopus, and SpringerLink respectively.</p> <p>The dataset “<strong>filtered_dataset.csv</strong>” consists of the deduplicated and filtered entries (journal articles and conference papers from 2010 onward, with a DOI assigned) from the “main_dataset.csv” we used as the final dataset for answering our research questions. It is structured into ten columns: (i) doi; (ii) title; (iii) venue; (iv) publication_year; (v) content_type; (vi) abstract; (vii) keywords; (viii) science gateway | scientific gateway; (ix) virtual laboratory | Vlab; and (x) virtual research environment.</p> <p>As for the previous dataset, the ‘<strong>doi</strong>’, ‘<strong>title</strong>’, and ‘<strong>publication_year</strong>’ labels are self-describing, and are used for the DOIs, titles, and publication years (in the yyyy format) respectively.</p> <p>The ‘<strong>venue</strong>’ label is used for indicating the conference or the journal the entries refer to. The values derive from the original query results.</p> <p>The ‘<strong>abstract</strong>’ and ‘<strong>keyword</strong>’ labels are used for the abstracts and the keywords associated with the entries. The values are mainly derived from the original query results, as we integrated the missing ones by querying OpenAIRE.</p> <p>The ‘<strong>science gateway | scientific gateway</strong>’, ‘<strong>virtual laboratory | Vlab</strong>’ and ‘<strong>virtual research environment</strong>’ labels indicate the connection between the entries and the keyword group used for denoting them. The values are binary (1 if the keywords belong to the group, 0 if they do not).</p> <p>The datasets “<strong>sg_vlab_vre_topics_datasets.csv</strong>” and “<strong>sgvlabvre_topics_dataset.csv</strong>” consist of the three datasets and of the unique dataset resulting from topic modelling, the first (corpus divided into three datasets) and the second analysis (corpus as a whole) respectively. They share the same structure: (i) Topic; (ii) #studies; (iii) Representative word; (iv) Representative word weight.</p> <p>The ‘<strong>Topic</strong>’ label is used for the topic denomination and the values consist of an alphanumeric string indicating the dataset and the progressive topic number: (a) SG, for the scientific gateway dataset; (b) VRE, for the virtual research environment dataset; (c) VLAB, for the virtual laboratory dataset; and (d) A, for the corpus as a whole.</p> <p>The ‘<strong>#studies</strong>’ label indicates the number of studies contributing to each topic.</p> <p>The ‘<strong>Representative word</strong>’ and ‘<strong>Representative word weight</strong>’ labels are used for denoting the keywords describing each topic and their weights respectively.</p>
Multi-stakeholder research data management training as a tool to improve the quality, integrity, reliability and reproducibility of research: Quantitative data of the post-course surveys
<p>Data contains doctoral students' and postdoc researchers' (n=168) self-ratings of their RDM competencies before and after the 3 ECTS credits "Basics of Research Data Management" (BRDM) trainings held 2019-2021 in the University of Turku and Åbo Akademi University, Finland. Moreover, data contains respondents' self-reported further learning needs.</p>
Data from Phenocam (PHE) measurements of in-canopy vegetation (hartheim2) at Hartheim Forest Research Site (DE-Har) from 2018-11-05 to 2018-12-31 [RAW]
<p>Phenocam images from "hartheim2" at DE-Har separated into near-infrared (NIR) and visible (VIS) for the year 2022. The phenocam "hartheim2" was put into operation on November 5, 2018. There are no phenocam images before that date at this site.</p> <p>Phenocam "hartheim2" shows the view from the main tower at 8.4 m height towards N at the <a href="https://www.meteo.uni-freiburg.de/en/infrastructure/hartheim-forest-research-site?set_language=en">ICOS Associate Ecosyste Site DE-Har, Germany</a> recording the phenology and state of in-canopy vegetation.</p>
Descriptions of SNSF-funded research projects
<p>This repository contains the data to replicate the analyses performed in Meier, D. S., Mata, R., & Wulff, D. U. (2021). text2sdg: An R package to Monitor Sustainable Development Goals from Text. <em>arXiv preprint arXiv:2110.05856</em>.</p> <p>The <a href="../api/records/11060662/draft/files/GrantWithAbstracts.csv/content" target="_blank" rel="noopener noreferrer">GrantWithAbstracts.csv</a> data is originally provided by the Swiss National Science Foundation and can also be downloaded from their <a href="https://data.snf.ch/datasets">website</a>. This data contains information on research projects funded by the Swiss National Science Foundation between 1975 and 2022. Among other things, the data provides information on what the funded projects were about and how much funding was provided.</p> <p>The <a href="../api/records/11060662/draft/files/backtrans_table.RDS/content" target="_blank" rel="noopener noreferrer">backtrans_table.RDS</a> data contains 1,500 randomly selected projects from the <a href="../api/records/11060662/draft/files/GrantWithAbstracts.csv/content" target="_blank" rel="noopener noreferrer">GrantWithAbstracts.csv</a> data that were translated from English to German and then from German back to English.</p> <p>The <a href="../api/records/11060662/draft/files/benchmark_table_revision.rds/content" target="_blank" rel="noopener noreferrer">benchmark_table_revision.rds</a> data contains runtimes from benchmarking the text2sdg R package. </p>
Data from Phenocam (PHE) measurements of in-canopy vegetation (hartheim2) at Hartheim Forest Research Site (DE-Har) from 2020-01-01 to 2020-12-31 [RAW]
<p>Phenocam images from "hartheim2" at DE-Har separated into near-infrared (NIR) and visible (VIS) for the year 2020. </p> <p>Phenocam "hartheim2" shows the view from the main tower at 8.4 m height towards N at the <a href="https://www.meteo.uni-freiburg.de/en/infrastructure/hartheim-forest-research-site?set_language=en">ICOS Associate Ecosyste Site DE-Har, Germany</a> recording the phenology and state of in-canopy vegetation.</p>
Data from Phenocam (PHE) measurements of in-canopy vegetation (hartheim2) at Hartheim Forest Research Site (DE-Har) from 2021-01-01 to 2021-12-31 [RAW]
<div> <p>Phenocam images from "hartheim2" at DE-Har separated into near-infrared (NIR) and visible (VIS) for the year 2021. </p> <p>Phenocam "hartheim2" shows the view from the main tower at 7m height towards N at the <a href="https://www.meteo.uni-freiburg.de/en/infrastructure/hartheim-forest-research-site?set_language=en">ICOS Associate Ecosystem Site DE-Har, Germany</a> recording the phenology and state of in-canopy vegetation.</p> </div>
Data from Phenocam (PHE) measurements of in-canopy vegetation (hartheim2) at Hartheim Forest Research Site (DE-Har) from 2022-01-01 to 2022-12-31 [RAW]
<div> <p>Phenocam images from "hartheim2" at DE-Har separated into near-infrared (NIR) and visible (VIS) for the year 2022. </p> <p>Phenocam "hartheim2" shows the view from the main tower at at 8.4 m height towards N at the <a href="https://www.meteo.uni-freiburg.de/en/infrastructure/hartheim-forest-research-site?set_language=en">ICOS Associate Ecosystem Site DE-Har, Germany</a> recording the phenology and state of in-canopy vegetation.</p> <p> </p> </div>
GraspOS landscape survey on Reforming Research Assessment
<p><span>This dataset is related to GraspOS Deliverable D2.1 "OS-aware RRA approaches landscape report" (<a title="https://zenodo.org/records/11098095" href="../records/11098095" target="_blank" rel="noreferrer noopener">https://zenodo.org/records/11098095</a>), Annex 1. GraspOS landscape survey on Reforming Research Assessment.</span></p>
Beyond the Digital Divide: Sharing Research Data across Developing and Developed Countries
<p>The primary data collection element of this project related to observational based fieldwork at four universities in Kenya and South Africa undertaken by Louise Bezuidenhout (hereafter ‘LB’) as the award researcher. The award team selected fieldsites through a series of strategic decisions. First, it was decided that all fieldsites would be in Africa, as this continent is largely missing from discussions about Open Science. Second, two countries were selected – one in southern (South Africa) and one in eastern Africa (Kenya) – based on the existence of the robust national research programs in these countries compared to elsewhere on the continent. As country background, Kenya has 22 public universities, many of whom conduct research. It also has a robust history of international research collaboration – a prime example being the long-standing KEMRI-Wellcome Trust partnership. While the government encourages research, financial support for it remains limited and the focus of national universities is primarily on undergraduate teaching. South Africa has 25 public universities, all of whom conduct research. As a country, South Africa has a long history of academic research, one which continues to be actively supported by the government. </p> <p>Third, in order to speak to conditions of research in Africa, we sought examples of vibrant, “homegrown” research. While some of the researchers at the sites visited collaborated with others in Europe and North America, by design none of the fieldsites were formally affiliated to large internationally funded research consortia or networks. Fourth, within these two countries four departments or research groups in academic institutions were selected for inclusion based on their common discipline (chemistry/biochemistry) and research interests (medicinal chemistry). These decisions were to ensure that the differences in data sharing practices and perceptions between disciplines noted in previous studies would be minimized. </p> <p>Within Kenya, site 1 (KY1) and Site 2 (KY2) were both chemistry departments of well-established universities. Both departments had over 15 full time faculty members, however faculty to student ratios were high and the teaching loads considerable. KY1 had a large number of MSc and PhD candidates, the majority of whom were full-time and a number of whom had financial assistance. In contrast, KY2 had a very high number of MSc students, the majority of whom were self-funded and part-time (and thus conducted their laboratory work during holidays). In both departments space in laboratories was at a premium and students shared space and equipment. Neither department had any postdoctoral researchers. </p> <p>Within South Africa, site 1 (SA1) was a research group within the large chemistry department of a well-established and comparatively well-resourced university with a tradition of research. Site 2 (SA2) was the chemistry/biochemistry department of a university that had previously been designated a university for marginalized population groups under the Apartheid system. Both sites were the recipients of numerous national and international grants. SA2 had one postdoctoral researcher at the time, while SA1 had none.</p> <p>Empirical data was gathered using a combination of qualitative methods including embedded laboratory observations and semi-structured interviews. Each site visit took between three and six weeks, during which time LB participated in departmental activities, interviewed faculty and postgraduate students, and observed social and physical working environments in the departments and laboratories. Data collection was undertaken over a period of five months between November 2014 and March 2015, with 56 semi-structured interviews in total conducted with faculty and graduate students. Follow-on visits to each site were made in late 2015 by LB and Brian Rappert to solicit feedback on our analysis. </p>
R script and data files for Oakley et al (2017) Journal of Proteome Research. DOI: 10.1021/acs.jproteome.6b00797
<p>This R script and data replicates the analysis of Oakley et al (2017) Thermal shock induces host proteostasis disruption and endoplasmic reticulum stress in the model symbiotic Cnidarian <em>Aiptasia</em>. <em>Journal of Proteome Research</em>. 16:2121-2134. DOI: 10.1021/acs.jproteome.6b00797. </p>
Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions
<p>This data set corresponds to the paper: Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions [1] (Experiment: Comprehensive reporting).</p> <p>The key research questions corresponding to this data set were:</p> <p>RQ1: What is the role of challenges for the field of biomedical image analysis (e.g. How many challenges conducted to date? In which fields? For which algorithm categories? Based on which modalities?)</p> <p>RQ2: What is common practice related to challenge design (e.g. choice of metric(s) and ranking methods, number of training/test images, annotation practice etc.)? Are there common standards?</p> <p>RQ3: Does common practice related to challenge reporting allow for reproducibility and adequate interpretation of results?</p> <p>To address these research questions, we aimed to capture all biomedical image analysis challenges that have been conducted up to 2016. To acquire the data, we analyzed the websites hosting/representing biomedical image analysis challenges, namely grand-challenge.org, dreamchallenges.org and kaggle.com as well as websites of main conferences in the field of biomedical image analysis, namely Medical Image Computing and Computer Assisted Intervention (MICCAI), International Symposium on Biomedical Imaging (ISBI), International Society for Optics and Photonics (SPIE) Medical Imaging, Cross Language Evaluation Forum (CLEF), International Conference on Pattern Recognition (ICPR), The American Association of Physicists in Medicine (AAPM), the Single Molecule Localization Microscopy Symposium (SMLMS) and the BioImage Informatics Conference (BII). This yielded a list of 150 challenges with 549 tasks.</p> <p>Next, a tool for instantiating the challenge parameter list introduced in [1] was used by some of the authors (engineers and medical student) to formalize all challenges that met our inclusion criteria as follows: (1) Initially, each challenge was independently formalized by two different observers. (2) The formalization results were automatically compared. In ambiguous cases, when the observers could not agree on the instantiation of a parameter - a third observer was consulted, and a decision was made. When refinements to the parameter list were made, the process was repeated for missing values. Based on the formalized challenge data set, a descriptive statistical analysis was performed to characterize common practice related to challenge design and reporting.</p> <p>[1] Maier-Hein, L., Eisenmann, M., Reinke, A., Onogur, S., Stankovic, M., Scholz, P., Arbel, T., Bogunovic, H., Bradley, A. P., Carass, A., Feldmann, C., Frangi, A. F., Full, P. M., van Ginneken, B., Hanbury, A., Honauer, K., Kozubek, M., Landman, B. A., März, K., Maier, O., Maier-Hein, K., Menze, B. H., Müller, H., Neher, P. F., Niessen, W., Rajpoot, N., Sharp, G. C., Sirinukunwattana, K., Speidel, S., Stock, C., Stoyanov, D., Aziz Taha, A., van der Sommen, F., Wang, C.-W., Weber, M.-A., Zheng, G., Jannin, P., Kopp-Schneider, A.: Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions. arXiv preprint arXiv:1806.02051 (2018).</p>
Emissions Database for Global Atmospheric Research, version v4.3.2 part I Greenhouse gases
<p>The Emissions Database for Global Atmospheric Research (EDGAR) v4.3.2, partim Greenhouse gases compiles anthropogenic emissions data for CO2, CH4 and N2O based on international statistics and emission factors. The version v4.3.2 of the EDGAR emission inventory provides global estimates, broken down to IPCC-relevant source-sector levels, from 1970 (the year of EU’s first Air Quality Directive) to 2012 (the end year of the first commitment period of the Kyoto Protocol (KP)). Strengths of EDGAR v4.3.2 include global geo-coverage (226 countries), continuity in time, and comprehensiveness in activities. Emission sources of the multiple gases include all human activities except the land-use, land-use change and forestry sector and are compiled following a bottom-up and IPCC-compliant approach. The dataset provides in addition to the complete timeseries 1970-2012 also annual and global gridmaps of 0.1 degree by 0.1 degree resolution for each source-sector and each year. For 2010 also 12 monthly gridmaps per source-sector are provided.</p>
Quantitative assessment of research data management practice - University of Bordeaux
<p>This survey was run at the University of Bordeaux in January 2019 using the questionnaire "Quantitative assessment of research data management practice" :</p> <p>Teperek, M., Krause, J., Lambeng, N., Blumer, E., van Dijck, J., Eggermont, R., … der Velden, Y. T. (2019). Quantitative assessment of research data management practice. Retrieved from : <a href="https://osf.io/mz3fx/">https://osf.io/mz3fx/</a></p> <p>The questionnaire included all the primary and secondary common questions, institution-specific questions regarding services and file sharing (EPFL questions), institution-specific questions for profile information.</p> <p>Data from the 425 responses collected are published here.</p> <p>Details regarding data collection and curation are included in the README file.</p> <p> </p>
Dataset to "Persistent Identifiers for File Formats: enabling preservation and re-use of research data"
<p>This fileset includes a "preprint" and the main dataset <em>fileformatRecognizer</em> (as .xlsx and .csv) to the paper "Persistent identifiers for file formats: enabling preservation and re-use of research data" submitted to iPRES 2019, but subsequently rejected after peer review. For the sake of transparency, permission to make available here the anonymous reviews motivating the rejection (<em>ReviewsPIDs4fileFormats.odt</em>) was asked, but was left without response. Some images (screendumps) and text result files from file identification tools tested are included. Further, a simple xquery command file (BaseX) for <em>fetch:content-type</em>()<em>, </em>used for getting MIME-types for files, is also provided.</p>
Open Access in developing countries – attitudes and experiences of researchers Dataset
<p>A survey was conducted of 507 researchers from the developing world and connected to INASP’s AuthorAID project to ascertain experiences and attitudes to Open Access publishing. This file is the raw output from the survey, with names and email addresses removed to preserve anonymity. </p>
GitHub Profiles (users/organisations) and Repositories (research/non-research) of Potsdam Researchers and Research Organisations: An annotated dataset of with howfairis and software quality variables.
<p>This dataset accompanies the paper <em>"Software FAIRness, Documentation and Development Practices in Potsdam Researchers' GitHub Repositories"</em> It includes 3 CSV files that contain data related to github profiles of users/organisations, their repositories annotated as research/non-research repositories and followed by FAIRness and other software qualtiy variables. The data were collected using <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP">SWORDS-template-UP</a> (v1.0.0) methods (collect_users, collect_repositories, collect_variables) which is extended version of <a href="https://github.com/UtrechtUniversity/SWORDS-template">SWORS-template</a> adopted according our needs and detailed in the paper.</p> <p><strong>GitHub (research) user/organisation profiles. ( <em>github_profiles.csv )</em></strong></p> <table> <tbody> <tr> <td><strong>Column name</strong></td> <td><strong>Description </strong></td> </tr> <tr> <td>user_id</td> <td>GitHub username </td> </tr> <tr> <td>html_url </td> <td>URL of the GitHub profile </td> </tr> <tr> <td>type </td> <td>Type of profile (user or organization)</td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the organization </td> </tr> </tbody> </table> <p><strong>GitHub repositories <em>(github_repositories.csv)</em></strong></p> <p>This file contains the repositories scraped from the GitHub profiles of research users and organizations.</p> <table> <tbody> <tr> <td><strong>Column name </strong></td> <td><strong>Description </strong></td> </tr> <tr> <td>html_url </td> <td>URL link to the repository </td> </tr> <tr> <td>description</td> <td>GitHub project description </td> </tr> <tr> <td>project</td> <td>Specifies if the project is research or non-research</td> </tr> <tr> <td>language</td> <td>Programming language used in the project </td> </tr> <tr> <td>organisation</td> <td>Acronym or name of the university, institution, or research organization</td> </tr> <tr> <td>research_group</td> <td>Acronym or name of the research group the repository belongs to</td> </tr> </tbody> </table> <p><strong>Research repositories filtered and annotated <em>(github_research_repositories_filtered_annotated.csv)</em></strong></p> <p>This file contains filtered and annotated information about research repositories.</p> <table> <tbody> <tr> <td><strong>Column Name </strong></td> <td><strong>Description </strong></td> <td><strong>Collection Method </strong></td> </tr> <tr> <td>html_url </td> <td>Repository URL </td> <td> </td> </tr> <tr> <td>howfairis_repository</td> <td>Indicates if the repository is public or private (True/False) </td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_license </td> <td>Indicates if the repository has a license (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_registry</td> <td>Indicates if the repository has implemented community registry (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_citation</td> <td>Indicates if the repository has a .cff file (True/False) </td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>howfairis_checklist</td> <td>Indicates if the repository has implemented OpenSSF best practices badge (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/tree/main/collect_variables#usage">howfairis_variable.py</a>) is a wrapper for <a href="https://pypi.org/project/howfairis/">howfairis</a> pypi library that checks the 5 recommendations of <a href="https://fair-software.nl">FAIR</a></td> </tr> <tr> <td>fair_score</td> <td>Score based on howfairis variables (0-5) </td> <td> </td> </tr> <tr> <td>dlr_soft_class</td> <td>Name of the university, company, research institute, or research organization</td> <td>(Manual) Annotated the repository based on <a href="https://core.ac.uk/reader/211557820">DLR software engineering guideline.</a> There are no specific definitions on metrics how to categorise them (github repositories) into application classes. Which were needed to do a comparitive analysis. </td> </tr> <tr> <td>installation_instruction</td> <td>Presence of installation instruction (True/False) </td> <td>(Manual) Checked the presense of Installation Instruction in the readme or in the project wiki pages. </td> </tr> <tr> <td>project_information </td> <td>Presence of basic project information in README (True/False) </td> <td>(Manual) Checked if the readme have basic information about the project. </td> </tr> <tr> <td>usage_guide</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td>(Manual) Checked the presense of Usage Guide in the readme or in the project wiki pages. For command line tools checked if they have help command which guides how to use the tool. </td> </tr> <tr> <td>test_folder</td> <td>Presence of folder named test/tests in the root directory (True/False)</td> <td> <p>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/docs/collect_variables/scripts/soft_dev_pract/test_folder.py">test_folder.py</a>) Checks the folder names test/tests in the root directory of the repository.</p> </td> </tr> <tr> <td>requirements_explicit </td> <td>Explicit requirements for Python, R, C++ repositories (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/requirement_explicit.py">requirement_explicit.py</a>) Checks the files (requirements.txt, DESCRIPTION, CMakeLists.txt) in the root directory. </td> </tr> <tr> <td>continuous_integration</td> <td>Indicates if the repository uses continuous integration (True/False)</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>ci_tool </td> <td>Name of the continuous integration tool used</td> <td>(Script- <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/continious_integration.py">continious_integration.py</a>) Checks the presence of folder .github (github actions) same for other continious integration (travisCI, CircleCI, Jekins, azure pipeline)</td> </tr> <tr> <td>add_lint_rule </td> <td>Indicates if additional linting rules are present (True/False)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the <br>.github/workflows directory to detect the presence of (linters) Python, R, and C++.</td> </tr> <tr> <td>add_test_rule</td> <td>Indicates if additional testing rules are present (True/False) </td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/add_ci_rules.py">add_ci_rules.py</a>) - it scans the YAML files in the <br>.github/workflows directory to detect the presence of (testing libraries) Python, R, and C++.</td> </tr> <tr> <td>comment_at_start</td> <td>Indicates the level of comments at the start of the program (most, more, some, less)</td> <td>(Script - <a href="https://github.com/Software-Engineering-Group-UP/SWORDS-template-UP/blob/main/collect_variables/scripts/soft_dev_pract/comment_at_start.py">comment_at_start.py</a>) Checks the presence of brief comments at the start at source code files in GitHub repositories.</td> </tr> <tr> <td>language </td> <td>Programming language used in the repository </td> <td> </td> </tr> <tr> <td>type </td> <td>Specifies if the profile is a user or organization </td> <td>Github organisation or user profiles.</td> </tr> <tr> <td>organisation </td> <td>Name of the university, company, research institute, or research organization</td> <td>Oraganisation name (from where the user was found)</td> </tr> <tr> <td>research_group</td> <td>Name or acronym of the research group </td> <td> </td> </tr> </tbody> </table> <p> </p> <p>Data for publication - https://github.com/Software-Engineering-Group-UP/potsdam-research-repos</p>
Research data supporting "Block copolymer-directed single diamond hybrid structures derived from X-ray nanotomography"
<p>Research data supporting "Block copolymer-directed single diamond hybrid structures derived from X-ray nanotomography"</p>
Dataset for 'A Matter of Culture? Conceptualising and Investigating 'Evidence Cultures' within Research on Evidence-Informed Policymaking'
<p><strong><span>Introduction</span></strong><strong><span><br></span></strong><span>This document describes the data collection and datasets used in the manuscript "A Matter of Culture? Conceptualising and Investigating ‘Evidence Cultures’ within Research on Evidence-Informed Policymaking" <span>[1].</span></span></p> <p><strong><span>Data Collection</span></strong></p> <p><span>To construct the citation network analysed in the manuscript, we first designed a series of queries to capture a large sample of literature exploring the relationship between evidence, policy, and culture from various perspectives. Our team of domain experts developed the following queries based on terms common in the literature. These queries search for the terms included in the titles, abstracts, and associated keywords of WoS indexed records (i.e. ‘TS=’). While these are separated below for ease of reading, they combined into a single query via the OR operator in our search. Our search was conducted on the Web of Science’s (WoS) Core Collection through the University of Edinburgh Library subscription on 29/11/2023, returning a total of <strong><u>2,089 records</u></strong>.</span></p> <p><em><span>TS = ((“cultures of evidence” OR “culture of evidence” OR “culture of knowledge” OR “cultures of knowledge” OR “research culture” OR “research cultures” OR “culture of research” OR “cultures of research” OR “epistemic culture” OR “epistemic cultures” OR “epistemic community” OR “epistemic communities” OR “epistemic infrastructure” OR “evaluation culture” OR “evaluation cultures” OR “culture of evaluation” OR “cultures of evaluation” OR “thought style” OR “thought styles” OR “thought collective” OR “thought collectives” OR “knowledge regime” OR “knowledge regimes” OR “knowledge system” OR “knowledge systems” OR “civic epistemology” OR “civic epistemologies”) AND (“policy” OR “policies” OR “policymaking” OR “policy making” OR “policymaker” OR “policymakers” OR “policy maker” OR “policy makers” OR “policy decision” OR “policy decisions” OR “political decision” OR “political decisions” OR “political decision making”))</span></em></p> <p><em><span>OR</span></em></p> <p><em><span>TS = ((“culture” OR “cultures”) AND ((“evidence-based” OR “evidence-informed” OR “evidence-led” OR “science-based” OR “science-informed” OR “science-led” OR “research-based” OR “research-informed” OR “evidence use” OR “evidence user” OR “evidence utilisation” OR “evidence utilization” OR “research use” OR “researcher user” OR “research utilisation” OR “research utilization” OR “research in” OR “evidence in” OR “science in”) NEAR/1 (“policymaking” OR “policy making” OR “policy maker” OR “policy makers”)))</span></em></p> <p><em><span>OR</span></em></p> <p><em><span>TS = ((“culture” OR “cultures”) AND (“scientific advice” OR “technical advice” OR “scientific expertise” OR “technical expertise” OR “expert advice”) AND (“policy” OR “policies” OR “policymaking” OR “policy making” OR “policymaker” OR “policymakers” OR “policy maker” OR “policy makers” OR “political decision” OR “political decisions” OR “political decision making”))<span> </span></span></em></p> <p><em><span>OR</span></em></p> <p><em><span>TS = ((“culture” OR “cultures”) AND (“post-normal science” OR “trans-science” OR “transdisciplinary” OR “transdisiplinarity” OR “science-policy interface” OR “policy sciences” OR “sociology of knowledge” OR “sociology of science” OR “knowledge transfer” OR “knowledge translation” OR “knowledge broker” OR “implementation science” OR “risk society”) AND (“policymaking” OR “policy making” OR “policymaker” OR “policymakers” OR “policy maker” OR “policy makers”))</span></em></p> <p><strong><span>Citation Network Construction</span></strong></p> <p><span>All bibliographic metadata on these 2,089 records were downloaded in five batches in plain text and then merged in R. We then parsed these data into network readable files. All unique reference strings are given unique node IDs. A node-attribute-list (‘CE_Node’) links identifying information of each document with its node ID, including authors, title, year of publication, journal WoS ID, and WoS citations. An edge-list (‘CE_Edge’) records all citations from these documents to their bibliographies – with edges going <em>from</em> a citing document <em>to</em> the cited – using the relevant node IDs. These data were then cleaned by (a) matching DOIs for reference strings that differ but point to the same paper, and (b) manual merging of obvious duplicates caused by referencing errors.</span></p> <p><span>Our initial dataset consisted of 2,089 <em>retrieved</em> documents and 123,772 <em>unretrieved</em> cited documents (i.e. documents that were cited within the publications we retrieved but which were not one of these 2,089 documents). These documents were connected by 157,229 citation links, but ~87% of the documents in the network were cited just once. To focus on relevant literature, we filtered the network to include <em>only</em> documents with at least three citation or reference links. We further refined the dataset by focusing on the main connected component, resulting in 6,650 nodes and 29,198 edges. <strong><u>It is this dataset that we publish here</u></strong>, and it is this network that underpins Figure 1, Table 1, and the qualitative examination of documents (see manuscript for further details). </span></p> <p><span>Our final network dataset contains 1,819 of the documents in our original query (~87% of the original retrieved records), and 4,831 documents not retrieved via our Web of Science search but cited by at least three of the retrieved documents. We then clustered this network by modularity maximization via the Leiden algorithm <span>[2]</span>, detecting 14 clusters with Q=0.59. Citations to documents within the same cluster constitute ~77% of all citations in the network. </span></p> <p><strong><span>Citation Network Dataset Description</span></strong></p> <p><span>We include two network datasets: (i) ‘CE_Node.csv’ that contains 1,819 retrieved documents, 4,831 unretrieved referenced documents, making for a total of 6,650 documents (nodes); (ii)’CE_Edge.csv’ that records citations (edges) between the documents (nodes), including a total of 29,198 citation links. These files can be used to construct a network with many different tools, but we have formatted these to be used in Gephi 0.10<span>[3]</span>. </span></p> <p><strong><span>‘CE_Node.csv’</span></strong><span> is a comma-separate values file that contains two types of nodes: </span></p> <p><span><span>i.<span> </span></span></span><span>Retrieved documents – these are documents captured by our query. These include full bibliographic metadata and reference lists. </span></p> <p><span><span>ii.<span> </span></span></span><span>Non-retrieved documents – these are documents referenced by our retrieved documents but were not retrieved via our query. These only have data contained within their reference string (i.e. first author, journal or book title, year of publication, and possibly DOI). </span></p> <p><span>The columns in the .csv refer to:</span></p> <p><span><span>-<span> </span></span></span><em><span>Id</span></em><span>, the node ID</span></p> <p><span><span>-<span> </span></span></span><em><span>Label</span></em><span>, the reference string of the document</span></p> <p><span><span>-<span> </span></span></span><em><span>DOI</span></em><span>, the DOI for the document, if available</span></p> <p><span><span>-<span> </span></span></span><em><span>WOS_ID</span></em><span>, WoS accession number</span></p> <p><span><span>-<span> </span></span></span><em><span>Authors</span></em><span>, named authors</span></p> <p><span><span>-<span> </span></span></span><em><span>Title</span></em><span>, title of document</span></p> <p><span><span>-<span> </span></span></span><em><span>Document_type</span></em><span>, variable indicating whether a document is an article, review, etc.</span></p> <p><span><span>-<span> </span></span></span><em><span>Journal_book_title, </span></em><span>journal of publication or title of book</span></p> <p><span><span>-<span> </span></span></span><em><span>Publication year</span></em><span>, year of publication.</span></p> <p><span><span>-<span> </span></span></span><em><span>WOS_times_cited</span></em><span>, total Core Collection citations as of 29/11/2023</span></p> <p><span><span>-<span> </span></span></span><em><span>Indegree</span></em><span>, number of <strong><em>within</em></strong> network citations to a given document</span></p> <p><span><span>-<span> </span></span></span><em><span>Cluster</span></em><span>, provides the cluster membership number as discussed in the manuscript (Figure 1)</span></p> <p><strong><span>‘CE_Edge.csv’</span></strong><span> is a comma-separated values file that contains edges (citation links) between nodes (documents) (<em>n</em>=29,198). The columns refer to:</span></p> <p><span><span>-<span> </span></span></span><em><span>Source</span></em><span>, node ID of the <em>citing</em> document</span></p> <p><span><span>-<span> </span></span></span><em><span>Target, </span></em><span>node ID of the <em>cited</em> document</span></p> <p><strong><span>Cluster Analysis</span></strong></p> <p><span>We qualitatively analyse a set of publications from seven of the largest clusters in our manuscript. For this, we calculated the within cluster indegree of nodes, and read through the 10 most cited retrieved documents and 10 most cited unretrieved documents. To generate these lists, sub-graphs for each cluster needed to be generated, and then indegree was measured (i.e. counting the number of citations from papers within a cluster to other papers in that same cluster).</span></p> <p><strong><span>Notes</span></strong></p> <p><a href="https://zenodo.org/records/6615221#_ftnref1"><span>[1]</span></a><span> Bandola-Gill, J., Andersen, N., Leng, R. I., Pattyn, V., & Smith, K. E. (forthcoming). A Matter of Culture? Conceptualising and Investigating ‘Evidence Cultures’ within Research on Evidence-Informed Policymaking. Policy and Society</span></p> <p><a href="https://zenodo.org/records/6615221#_ftnref6"><span>[2]</span></a><span> Traag, V. A., Waltman, L., & van Eck, N. J. (2019). From Louvain to Leiden: guaranteeing well-connected communities. Scientific reports, 9(1), 5233. </span><a href="https://doi.org/10.1038/s41598-019-41695-z"><span>https://doi.org/10.1038/s41598-019-41695-z</span></a></p> <p><a href="https://zenodo.org/records/6615221#_ftnref5"><span>[3]</span></a><span> Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media. Gephi is available via </span><a href="https://gephi.org/"><span>https://gephi.org/</span></a></p> <p><span> </span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.