Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,499
datasets available to search
ShareScore release 0.9.0
Dataset results
13,499 results for “researcher”
Dataset for: Exploring the experiences of academic libraries with research data management: a meta-ethnographic analysis of qualitative studies
<p><strong>Overview</strong></p> <p>This dataset contains the raw data for the manusript:<br> Perrier L, Blondal E, MacDonald H. Exploring the experiences of academic libraries with research data management: a meta-ethnographic analysis of qualitative studies. 2018; 40(3-4): 173-183. doi: 10.1016/j.lisr.2018.08.002</p> <p>Full-text available at: <a href="https://doi.org/10.1016/j.lisr.2018.08.002">https://doi.org/10.1016/j.lisr.2018.08.002</a> </p> <p><strong>Data and Documentation Files</strong></p> <p>Five files make up the dataset:</p> <ol> <li>Data Dictionary: RDMMetaEthnography_DataDictionary_v1.pdf</li> <li>Data Abstraction Sheet: RDMMetaEthnography_StudyCharacteristics.csv</li> <li>Data Abstraction Sheet: RDMMetaEthnography_ParticipantCharacteristics.csv</li> <li>Data Abstraction Sheet: RDMMetaEthnography_Outcomes.csv</li> <li>Data Abstraction Sheet: RDMMetaEthnography_COREQ,csv</li> </ol> <p>Contact: Laure Perrier: <a href="https://orcid.org/0000-0001-9941-7129">orcid.org/0000-0001-9941-7129</a></p>
Resource Metadata Harvested from Government and Research Open Data Portals
<p>This dataset consists of resource metadata harvested from the APIs of hundreds of government and research data portals from all over the world. This dataset was harvested between the 13<sup>th</sup> and 15<sup>th</sup> of September 2018. The metadata harvested from these portals was translated to a single metadata format (see <em>metadata_format.odt</em>). An overview of all harvested domains is given in <em>portal_list.txt</em>.</p> <p>The harvested data is divided into five gzipped json-lines files, based on the ‘type’ of the resource that is derived from the data of the APIs:</p> <ul> <li><em>dataset_metadata.jsonl.gz</em>: Resources classified as a Dataset, or subsets of dataset (e.g. Dataset:Image and Dataset:Audio) [6 246 250 resources]</li> <li><em>document_metadata.jsonl.gz</em>: Resources classified as a Document, or subset of document (e.g. Document:Paper:Conference and Document:Book) [15 626 541 resources]</li> <li><em>software_metadata.jsonl.gz</em>: Resources classified as Sofware (including Software:Model) [42 036 resources]</li> <li><em>service_metadata.jsonl.gz</em>: Resources classified as a service (e.g. WMS, APIs) [1257 resources]</li> <li><em>other_metadata.jsonl.gz</em>: Resources of which the ‘type’ could not be determined from the data the API returned. This set still contains many datasets [1 502 979 resources]</li> </ul>
Long-term moss monitoring network for atmospheric deposition in Germany, link to research data and scientific software
<p>Research data and scientific software related to a study that aims to restructure a long-term monitoring network using moss as biomonitor for atmospheric deposition in Germany. Data from the European Moss Survey 2005 and a statistically based methodology including a decision support system were used to design the spatial network for the 2005 survey.</p>
Estimating heavy metal deposition in Germany using model calculations and biomonitoring data, link to research data and scientific software
<p>Research data and scientific software related to an investigation dealing with modelled data on Cd and Pb deposition (LOTOS-EUROS, EMEP/MSC-East) and monitoring data from the International Cooperative Programme on Effects of Air Pollution on Natural Vegetation and Crops (ICP Vegetation Moss Survey) and the German Environmental Specimen Bank (ESB) providing corresponding parameters on HM concentration in various biota. The study aimed at examining, whether an integrated use of model calculations and monitoring data can extend established methods for estimating and evaluating spatial patterns of atmospheric Pb and Cd deposition across Germany.</p>
Fuzzy modelling and mapping soil moisture in Germany, link to research data and scientific software
<p>Research data and scientific software related to spatio-temporal estimations of ecological soil moisture with available data covering the whole territory of Germany and the Kellerwald National Park (Hesse). Temporal trends of modelled soil moisture for the time period 1961–2070 were statistically analyzed. Soil moisture changes (drying-out) at both national and regional levels were mapped.</p>
research-software/resosuma-data: 0.4.1
<p><strong><em>resosuma-data</em></strong> represents activities in the research software sustainability space in the CSV format, where column 1 contains actors in the space, column 2 contains activities, and column 3 contains actees.</p> <p>resosuma-0.4.1.csv fixes issues <a href="https://github.com/research-software/resosuma-data/issues/3">#3</a> and <a href="https://github.com/research-software/resosuma-data/issues/4">#4</a>.</p>
A decade of Semantic Web research through the lenses of a mixed methods approach (Resources)
<p>This work has been submitted to <a href="http://www.semantic-web-journal.net/content/decade-semantic-web-research-through-lenses-mixed-methods-approach">Semantic Web Journal</a>. We provide here resources to reproduce our approach.</p> <p>In this paper, we aim to provide a broader and more complete picture of Semantic Web topics and trends by adopting a mixed methods methodology, which allows a combined use of both qualitative and quantitative approaches. Concretely, we build on a qualitative analysis of the main seminal papers, which adopt a top-down approach, and on quantitative results derived with three bottom-up data-driven approaches (<a href="https://technologies.kmi.open.ac.uk/Rexplore/">Rexplore</a>, <a href="http://saffron.insight-centre.org/">Saffron</a>, <a href="https://www.poolparty.biz/">PoolParty</a>), on a corpus of Semantic Web papers published in the last decade. In this process, we both use the latter for “fact-checking” on the former and also to derive key findings in relation to the strengths and weaknesses of top-down and bottom-up approaches to research topic identification.</p> <p>Please access the full set of resources at: <a href="https://aic.ai.wu.ac.at/qadlod/SW/">https://aic.ai.wu.ac.at/qadlod/SW/</a></p>
Derived data and analysis code accompanying Deines et al. 2019, Environmental Research Letters
<p>This codebase accompanies the paper:</p> <p>Deines, JM, AD Kendall, JJ Butler, Jr., & DW Hyndman. 2019. Quantifying irrigation adaptation strategies in response to stakeholder-driven groundwater management in the US High Plains Aquifer. Environmental Research Letters. DOI: <a href="https://doi.org/10.1088/1748-9326/aafe39">https://doi.org/10.1088/1748-9326/aafe39</a></p> <p>Data and code at time of publication.</p>
GEDII Survey on Research Teams Dataset
<p>This dataset contains the result of the GEDII cross-country survey carried out among EU based research teams during the year 2017. It contains research teams from Austria, Belgium, Czech Republic, Denmark, Finland, France, Germany, Italy, Lithuania, the Netherlands, Norway, Poland, Portugal, Spain, Sweden, Switzerland and the UK. Overall, the Consortium recruited 159 research groups submitting a total of 1501 online questionnaires - out of which 1357 are complete.</p> <p>The dataset <strong>is unique</strong> in that it combines individual level responses from research group members regarding team climate, communication patterns, or leadership style among others with bibliometric performance data for each particular group retrieved from Web of Science (Clarivate Analytics).</p> <p>The data is distributed as R package "gediiccs", including extensive documentation.</p> <p>NOTE: The open access version of the "gediiccs" package contains the <strong>team level data only</strong>. This restriction is due to confidentiality concerns: given the large quantity of variables in relation to the relatively small research groups, it would be easy for group members to "reverse engineer" the identity of their own group as well as the responses from individual group members. Hence, only<strong> team level aggregated data</strong> has been released.</p>
Dataset: Approximating input data to a snowmelt model using Weather Research and Forecasting model outputs in lieu of meteorological measurements
<p>The dataset presented is the companion data to the Journal of Hydrometeorology publication entitled “Approximating input data to a snowmelt model using Weather Research and Forecasting model outputs in lieu of meteorological measurements.” The data that follows contains everything needed to reproduce the spatial inputs for the meteorological station model run using the Spatial Modeling for Resources Framework (SMRF, Havens et al., 2017).</p> <p> </p> <p>Software versions used:</p> <ul> <li>Image Processing Workbench v2.2.0 (Marks et al., 2017)</li> <li>Spatial Modeling for Resources Framework v0.5.3 (Havens et al., 2019)</li> </ul> <p> </p> <p><strong>NOTE:</strong> Reproducing the spatial inputs will generate 10 netCDF files at ~80GB per file.</p> <p> </p> <p><strong>topo.nc</strong> – Contains multiple static layers that are required to run SMRF and iSnobal. The netCDF layers are:</p> <ul> <li>dem – digital elevation model at 100 meter resolution, aggregated from the 10 meter National Elevation Dataset (Archuleta et al., 2017)</li> <li>mask – basin mask for the Boise River Basin</li> <li>veg_height – vegetation height in meters from the National Land Cover Database (Homer et al., 2015)</li> <li>veg_type – vegetation type from the National Land Cover Database</li> <li>veg_tau – vegetation fractional transmissivity derived from the vegetation type</li> <li>veg_k – vegetation emissivity derived from the vegetation type</li> </ul> <p> </p> <p><strong>maxus.nc</strong> – maximum upwind slope netCDF that contains 72 images for all wind directions in 5 degree increments using the algorithm described in Winstral and Marks (2002)</p> <p> </p> <p><strong>Station data:</strong></p> <ul> <li>Contains hourly meteorological station data downloaded from Mesowest (Horel et al., 2002). Data was cleaned and filtered prior to running SMRF.</li> <li>metadata.csv – metadata for 40 stations</li> <li>air_temp.csv – 38 stations</li> <li>cloud_factor.csv – 7 stations</li> <li>precip.csv – 21 stations</li> <li>vapor_pressure.csv – 19 stations</li> <li>wind_direction.csv – 14 stations</li> <li>wind_speed.csv – 14 stations</li> </ul> <p> </p> <p><strong>smrf_config.ini</strong> – Configuration file needed to reproduce the spatial inputs using SMRF. The paths will need to be changed to reflect the data location.</p>
Data Structure of Clinical Research
<p><em>Bro</em><em>nchial asthma is one of the most common respiratory pathologies in children, characterized by rising incidence around the world. Early disease onset, severe clinical signs of bronchial asthma, the ineffectiveness of high doses of hormone therapy reduce the quality of life in patients and lead to disability. Analysis of bronchial asthma heterogeneity is now possible in virtue of computer technology and big data processing onrush.</em></p> <p><em>The information on 70 children suffering from bronchial asthma and 20 children from the control group was analyzed in this study.</em></p> <p><em>Gender, age, duration of disease, associated diseases, family history of allergic diseases, clinical blood and urine test, spirography and blood immunoassay results, total IgE, thymic stromal lymphopoietin and results of skin allergy tests were taken into account.</em></p>
Research institutions clustering based on the intensity of academic collaboration
<p>The clustering of research institutions has been conducted using the Louvain modularity algorithm. The Louvain modularity is a state-of-the-art method of identifying communities (clusters) in large networks. Modularity is a value between -1 and 1 that measures the density of edges inside communities to edges outside communities. Optimizing this value results in the best possible grouping of the nodes of a given network.</p> <p>In our exercise, the Louvain methods were applied to identify clusters of institutions within ACM and SSRN networks. In the network, nodes are constituted of institutions, and edges are represented by the intensity of research collaboration measured by number of papers co-authored by authors affiliated with the institutions.</p> <p>As an example, including the paper: <em>Fast unfolding of communities in large networks</em>, written by V. D. Blondel (Universite Catholique de Louvain), J-L. Guillaume (Universite Pierre et Marie Curie), R. Lambiotte (Imperial College London) and Etienne Lefebvre (Universite Catholique de Louvain) would impact the number of edges in our analysis in the following way:</p> <p>“Universite catholique de Louvain” ⇔ “Imperial College London” =+1</p> <p>“Universite catholique de Louvain” ⇔ “Universite Pierre et Marie Curie” =+1</p> <p>“Imperial College London” ⇔ “Universite Pierre et Marie Curie” =+1</p> <p>In our largest network we analyse 5362 institution nodes with 147 482 edges. The number of identified clusters highly depends on the resolution parameter. Resolution is a parameter for the Louvain community detection algorithm that affects the size of the recovered clusters. Smaller resolutions recover smaller, and therefore a larger number of clusters, and conversely, larger values recover clusters containing more data points. In all clusterizations, we have used a default resolution (1.0) tuned in the popular Gephi software for network analysis. Resolutions equal to one result in a moderate number of clusters, characterised by satisfactory statistical distribution. </p> <p><strong>Source:</strong></p> <p>- Association for Computing Machinery (ACM)</p> <p>Characteristics of the ACM Data Set following geographical classification</p> <p>Number of institutions: 5477</p> <p>Number of papers: 674684</p> <p>Number of countries: 122</p> <p>Years: 2011-2018</p> <p>As ACM contains publications across various areas of computer science, a more in-depth analysis requires the classification of papers into fields of interests. During the analysis, we looked at 3 wide areas:</p> <ul> <li> <p>Artificial intelligence and machine learning</p> </li> <li> <p>Technology (hardware, emerging technologies, infrastructure)</p> </li> <li> <p>Social issues</p> </li> </ul> <p>The 3 categories were set following expert analysis of the 1000 most frequent keywords in the dataset. If a term from the following list appeared among the paper’s keywords, the paper was assigned to that group, allowing a paper to assign to more than one group.</p> <p><strong>Files:</strong></p> <p>mod_ai.csv (based on keywords related to artificial intelligence)</p> <p>mod_tech.csv (based on keywords related to technologies)</p> <p>mod_soc.csv (based on keywords related to social issues)</p> <p>mod_all.csv (based on all papers)</p> <p> </p>
Dataset: Publication cultures and Dutch research output: a quantitative assessment
<p>Dataset belonging to the report: <a href="https://doi.org/10.5281/zenodo.2643360">Publication cultures and Dutch research output: a quantitative assessment</a></p> <p> </p> <p>On the report:</p> <p>Research into publication cultures commissioned by VSNU and carried out by Utrecht University Library has detailed university output beyond just journal articles, as well as the possibilities to assess open access levels of these other output types. For all four main fields reported on, the use of publication types other than journal articles is indeed substantial. For Social Sciences and Arts & Humanities in particular (with over 40% and over 60% of output respectively not being regular journal articles) looking at journal articles only ignores a significant share of their contribution to research and society. This is not only about books and book chapters, either: book reviews, conference papers, reports, case notes (in law) and all kinds of web publications are also significant parts of university output.</p> <p>Analyzing all these publication forms and especially determining to what extent they are open access is currently not easy. Even combining some the largest citation databases (Web of Science, Scopus and Dimensions) leaves out a lot of non-article content and in some fields even journal articles are only partly covered. Lacking metadata like affiliations and DOIs (either in the original documents or in the scholarly search engines) makes it even harder to analyze open access levels by institution and field. Using repository-harvesting databases like BASE and NARCIS in addition to the main citation databases improves understanding of open access of non-article output, but these routes also have limitations. The report has recommendations for stakeholders, mostly to improve metadata and coverage and apply persistent identifiers.</p>
Research data for: Preventing the coffee-ring effect and aggregate sedimentation by in situ gelation of monodisperse materials
<p>Raw data for the publication: Preventing the coffee-ring effect and aggregate sedimentation by in situ gelation of monodisperse materials</p>
Experimental result to investigate the influence of user's tweets and diversification on serendipitous research paper recommendations
<p>This is a raw dataset of the experiment result to investigate the influence of user's tweets and diversification on serendipitous research paper recommendations.</p> <p> </p>
Quantitative assessment of research data management practice
<p>This survey aims to investigate research data management practices in academic institutions. The survey comprises questions common to all institutions as well as institution-specific ones. Common questions were drafted in the frame of a collaboration between several RDM services: Tu Delft (team effort), EPFL (team effort), University of Cambridge (notably Marta Busse) and University of Illinois (notably Heidi Imker). The first survey was run by TU Delft and EPFL only end of 2017. In total, 1263 responses where collected (680 from TU Delft, 235 from EPFL and 348 from the University of Cambridge) and are published here. The results of each institution are provided in Microsoft Excel 2007 (XLSX) format. Consolidated results are provided in CSV format</p> <p>The first lines of the CSV file contains the question asked to researchers. Each further line contains the response of a researcher; answers to institution-specific questions are set to N/A for researchers of the other institutions. Column delimiters are commas(,), quote chars are double-quotes (") and subfield separators are semi-columns (;). The text encoding is UTF-8.</p> <p>More information about this survey as well as the exact survey questions and a detailed description how the survey might be re-used by other institutions is available on the project page on the Open Science Framework: htts://osf.io/mz3fx/ For any questions contact datastewards@tudelft.nl or researchdata@epfl.ch</p>
Toolkit on Open Access for Research Project Coordinators
<p>The materials in this toolkit were created by Romain Féret as a resource for training on how to help project coordinators to comply with their open access requirements. The slides of the training are available on Zenodo at 10.5281/zenodo.3381783. This training day took place on Wednesday the 5th of June 2019, at the University of Lille. It was organized with the support of Couperin as a part of its activities in the project OpenAIRE-Advanced.</p> <p>The tutorials are divided into two folders. The ‘Coordinator’ folder contains documents that can be sent directly to the researchers, while the ‘Support staff’ folder contains tutorials for support staff (librarians, project managers) who help the coordinators to manage their project. Each tutorial is in .pdf and .docx format for easy reuse and modification. Each document is available in French and in English.</p>
Science ready spectra and their best-fitting models described in the research paper ``Internal dynamics and stellar content of nine ultra-diffuse galaxies in the Coma cluster prove their evolutionary link with dwarf early-type galaxies'' by Chilingarian et al.
<p>Science ready spectra of nine ultra-diffuse galaxies in the Coma cluster collected with the Binospec multi-object spectrograph and their best-fitting PEGASE.HR templates obtained using the NBursts full spectrum fitting code. These spectra were presented in the paper ``Internal dynamics and stellar content of nine ultra-diffuse galaxies in the Coma cluster prove their evolutionary link with dwarf early-type galaxies'' by Chilingarian et al. accepted for publication in the Astrophysical Journal on Sep/3/2019 (arXiv:1901.05489).</p> <p>Each spectrum is presented as a binary FITS table, which contains a spectrum (wavelength, flux, uncertainties), best-fitting template, best-fitting parameters (radial velocity, age, metallicity), and a pixel mask used in the fitting procedure. For six galaxies there are two files provided: (i) one-dimensional optimally extracted integrated spectrum and (ii) two dimensional spectrum for spatially resolved radial velocity information. For the remaining three galaxies, only spatially resolved spectra are provided.</p>
NOAA NCCOS Assessment: Prioritizing Areas for Future Seafloor Mapping, Research, and Exploration Offshore of California, Oregon, and Washington from 2019-03-01 to 2019-04-01
<p>Spatial information about the seafloor is critical for decision-making by marine resource science, management and tribal organizations. Coordinating data needs can help organizations leverage collective resources to meet shared goals. To help enable this coordination, the National Oceanic and Atmospheric Administration (NOAA) National Centers for Coastal Ocean Science (NCCOS) developed a spatial framework, process and online application to identify common data collection priorities for seafloor mapping, sampling and visual surveys offshore of the West Continental United States Coast (WCC). Twenty-six participants from NOAA’s West Coast Deep Sea Coral Initiative (WCDSCI) and Expanding Pacific Research and Exploration of Submerged Systems (EXPRESS) entered their priorities in an online application, using virtual coins to denote their priorities in 10x10 minute grid cells. Grid cells with more coins were higher priorities than cells with fewer coins. Participants also reported why these locations were important and what data types were needed. Results were analyzed and mapped using statistical techniques to identify significant relationships between priorities, reasons for those priorities and data needs. Ten high priority locations were broadly identified for future mapping, sampling and visual surveys. These locations were distributed throughout the WCC, primarily in depths less than 1,000 m. Participants consistently selected (1) Exploration, (2) Biota/Important Natural Area and (3) Research as their top reasons (i.e., justifications) for prioritizing locations, and (1) Benthic Habitat Map and (2) Bathymetry and Backscatter as their top data or product needs. This ESRI shapefile summarizes the results from this spatial prioritization effort. This information will enable NOAA WCDSCI, EXPRESS and other WCC organization to more efficiently leverage resources and coordinate their mapping of high priority locations along California, Oregon and Washington. </p> <p>This effort was funded by NOAA’s Deep Sea Coral Research and Technology Program (DSCRTP) through its WCDSCI. The overall goal of the project was to systematically gather and quantify suggestions for seafloor mapping, sampling and visual surveys for the WCDSCI and EXPRESS. The results are expected to help WCDSCI, EXPRESS and other organizations on the WCC to identify locations where their interests overlap with other organizations, to coordinate their data needs and to leverage collective resources to meet shared goals.</p> <p>There were four main steps in the WCC spatial prioritization process. The first step was to identify the technical advisory team, which included the 11 members of the DSCRTP WCDSCI Steering Committee and all of the participants involved in the EXPRESS campaign. This advisory team invited 37 participants for the prioritization. Step two was to develop the spatial framework and an online application. To do this, the WCC was divided into five subregions and 3,265 square grid cells approximately 10x10 minutes in size. Existing relevant spatial datasets (<em>e.g.</em>, bathymetry, protected area boundaries, etc.) were compiled to help participants understand information and data gaps and to identify areas they wanted to prioritize for future data collections. These spatial datasets were housed in the online application, which was developed using Esri’s Web AppBuilder. In step three, this online application was used by 26 participants to enter their priorities in each subregion of interest. Participants allocated virtual coins in the 10x10 minute grid cells to denote their priorities. Grid cells with more coins were higher priorities than cells with fewer coins. Participants also reported why these locations were important and what data types were needed. Coin values were standardized across the subregions and used to identify spatial patterns across the WCC region as a whole. The number of coins were standardized because each subregion had a different number of grid cells and participants. Standardized coin values were analyzed and mapped using statistical techniques, including hierarchical cluster analysis, to identify significant relationships between priorities, reasons for those priorities and data needs. This ESRI shapefile contains the 10x10 minute grid cells used in this prioritization effort and associated the standardized coin values overall, as well as by organization, justification and product. For a complete description of the process and analyses please see: Costa <em>et al</em>. 2019.</p>
Global trends and collaborations in electrochemical methods: a dataset on etching and deposition research
<p><span>This dataset supports the study "Electrochemical Etching vs. Electrochemical Deposition: A Comparative Bibliometric Analysis," which examines scientific publications on electrochemical etching and electrochemical deposition from 1970 to 2023. The dataset is derived from the Science Citation Index Expanded (SCIE) database and includes bibliometric information on publication trends, leading contributors, research areas, and keyword co-occurrences in both fields.</span></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.