Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
13,499
datasets available to search
ShareScore release 0.9.0
Dataset results
13,499 results for “researcher”
3D models (true color, TIF): Towards a spatial data repository for archaeological research in the Romanian Mostiștea Basin and Danube Valley
<p>Spatial data are crucial in archaeological research, where orthophotos, digital elevation models, and 3D models are widely used for mapping, documenting, and monitoring archaeological sites. The introduction of affordable and compact unmanned aerial vehicles (UAVs) has significantly advanced the use of UAV-based photogrammetry in the past 20 years. Recently, compact airborne systems have also enabled the capture of thermal, multispectral, and aerial laser scanning data. This study presents the data acquired with different platforms and sensors at Chalcolithic archaeological sites in Romania's Mostiștea Basin and Danube Valley. Since laser scanning and photogrammetry generate large data volumes, data storage and dissemination must also be carefully considered. Based on a thorough study of system performance, data acquisition and processing methods, and data outputs, a workflow for the systematic mapping and documentation of sites has been proposed. Given the experience obtained in the last 5 summer campaigns (2018-2023), 19 sites have been accurately mapped, of which 5 sites are mapped using airborne laser scanning. 18 sites are documented using multispectral photogrammetry, and for 17 sites, interactive image-based 3D models are acquired using true-color photogrammetry. All data are stored on a publicly accessible website for visualization, as well as on an open-data platform for data exchange. For the multispectral data, a raster tile service has been implemented, allowing the use of the data in a GIS environment.</p>
LLM Research Repository
<p><strong>Overview</strong></p> <p>Welcome to the Large Language Models (LLM) Repository, a curated collection aimed at researchers, practitioners, and enthusiasts in the field of Natural Language Processing (NLP). This repository offers resources related to Large Language Models, including research papers, theses, tools, datasets, courses, open-source models, and benchmarks. </p> <p><strong>1. Research Papers</strong></p> <p>A compilation of seminal and cutting-edge research papers that shape the field of Large Language Models. This section includes:</p> <ul> <li>Foundational Papers: Groundbreaking papers that laid the framework for LLM research.</li> <li>Recent Advances: Latest research examining novel architectures, training techniques, and applications.</li> <li>Survey and Review Articles: Comprehensive surveys and reviews that aggregate findings and offer insightful analysis on various aspects of LLMs.</li> </ul> <p><strong>2. Theses</strong></p> <p>A collection of master's and doctoral theses that focus on various facets of LLMs, providing in-depth explorations of core concepts, novel methodologies, and empirical studies from all over the world.</p> <p><strong>3. Tools</strong></p> <p>A list of tools and libraries essential for working with Large Language Models. This section encompasses:</p> <ul> <li>Development Frameworks: Popular libraries and frameworks.</li> <li>Utilities: Tools for data preprocessing, model deployment, and inference acceleration.</li> </ul> <p><strong>4. Datasets</strong></p> <p>A collection of datasets tailored for training and evaluating Large Language Models. This section includes:</p> <ul> <li>Text Corpora: Large-scale text datasets from diverse domains such as news articles, books, and social media.</li> <li>Annotated Datasets: Datasets with human annotations for tasks such as named entity recognition, sentiment analysis, and machine translation.</li> </ul> <p><strong>5. Courses</strong></p> <p>A list of university courses related to Large Language Models.</p> <p><strong>6. Open Source Models</strong></p> <p>Access to state-of-the-art open-source Large Language Models, allowing you to leverage pre-trained models for various applications. This section includes:</p> <ul> <li>Model Repositories: Links to GitHub repositories and model zoos hosting popular LLMs such as GPT, BERT, T5, and their variants.</li> <li>Pre-trained Models: Ready-to-use models available through platforms like the Hugging Face Model Hub, including detailed usage instructions and licensing information.</li> <li>Customized Implementations: Specialized versions and fine-tuned models tailored for specific tasks or domains.</li> </ul> <p><strong>7. Benchmarks</strong></p> <p>A suite of benchmarks designed to evaluate the performance and robustness of Large Language Models. This section features:</p> <ul> <li>Standard Benchmarks: Widely-accepted benchmarks like GLUE, SuperGLUE, and the LAMBADA dataset.</li> <li>Challenge Sets: Curated datasets that test specific capabilities of models, such as commonsense reasoning, multilingual understanding, or adversarial robustness.</li> </ul>
Bibliometric Analysis on Recent Topics in ILS Research
<p><strong><span>Bibliometric Analysis on Recent Topics in ILS Research</span></strong></p> <p><span><span> </span></span></p> <p><span>Pablo Rosser (Corresponding Author)<span> </span></span></p> <p><span>Faculty of Education, International University of La Rioja, Spain<span> </span></span></p> <p><span>Avenida de la Paz 137, 26006 Logroño, La Rioja, Spain - pablo.rosser@unir.net<span> </span></span></p> <p><span>Phone: +34 607321642<span> </span></span></p> <p><span>ORCID: 0000-0002-9802-7169</span></p> <p><span><br><br></span></p> <p><span>Seila Soler<span> </span></span></p> <p><span>Faculty of Humanities and Social Sciences, Isabel I University, Spain<span> </span></span></p> <p><span>C. de Fernán González 76, 09003 Burgos, Spain<span> </span></span></p> <p><span>seilaaixa.soler@ui1.es<span> </span></span></p> <p><span>Phone: +34 687704604<span> </span></span></p> <p><span>ORCID: 0009-0001-9125-3382</span></p> <p><span><span> </span></span></p> <p><span> </span></p> <p><strong><span>ABSTRACT</span></strong></p> <p><span><span> </span></span></p> <p><span>This article presents a bibliometric analysis of recent topics in ILS (Individual Learning Styles) research. The methodology employed was based on a systematic review of documents and scientific publications related to the study object from 2004 to the present utilizing the Scopus database and the 'bibliometrix' package in the R programming language. Several decision steps were followed to ensure the reliability and validity of the analysis. The findings allowed for the identification of the most cited journals, the most prominent authors, and the most common affiliations among the authors of the analyzed documents. Additionally, different models of learning style preferences and their impact on the teaching process were explored. The findings contribute significantly to understanding the dynamics of research in the studied discipline. Regarding the most cited journals, "Computers & Education" was found to be the leading journal in terms of the total number of citations, followed by "Educational Research Review" and "Learning and Instruction". As for the most prominent authors, several relevant names were identified, such as David Kolb, Richard Felder, and Peter Honey. Regarding the most common affiliations among the authors, it was found that several Spanish universities lead this list. On the other hand, this study explored different learning style models and their impact on the educational process. It was found that there is a wide variety of theoretical models on learning style, but there is no clear consensus on which model is best to apply in an educational context. However, the importance of considering students' learning styles to adapt the educational process to their individual needs was highlighted. In conclusion, this bibliometric analysis provides an overview of the current state of ILS research and can be useful for researchers, educators, and professionals interested in this field. The obtained results can also be used to identify current trends in ILS research and to guide future investigations in this area.</span></p>
MIRA-KG: A Knowledge Graph of Hypotheses and Findings for Social Demography Research
<p>A shift in scientific publishing from paper-based to knowledge-based practices promotes reproducibility, machine actionability and knowledge discovery. This is important for disciplines like social science, as study indicators are often social constructs such as race or education; hypothesis tests are challenging to compare in demographic research due to their limited temporal and spatial coverage; and natural language in research papers is often imprecise and ambiguous. Therefore, we present the MIRA-KG, consisting of: (1) an ontology for capturing social demography research, which links hypotheses and findings to evidence, (2) annotations of papers on health inequality in terms of the ontology, gathered by (i) prompting a Large Language Model to annotate paper abstracts using the ontology, (ii) mapping concepts to terms from NCBO BioPortal ontologies and GeoNames, and (iii) refining the final graph by a set of SHACL constraints, developed according to data quality criteria. The utility of the resource lies in its use for formally representing social demography research hypotheses, discovering research biases, discovery of knowledge, and the derivation of novel questions.<br><br>This dataset was generated using the code available on Github at <a href="https://w3id.org/mira/">https://w3id.org/mira/</a> at version v1.0. It uses the following ontology: <a href="https://w3id.org/mira/ontology/">https://w3id.org/mira/ontology/</a>. </p>
Coverage and quality of open metadata for Dutch research output - dataset
<p>Record level data underlying the figures and tables in the report:<strong><br><br>Coverage and quality of open metadata for Dutch research output - report<br></strong></p> <p><a href="https://doi.org/10.5281/zenodo.10629457" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10629457</a><br><br>The current dataset contains 3 csv files:</p> <ul> <li><em>rpo_nl_list_long_20240201.csv</em> - list of identifiers (ROR ID, OpenAlex ID, OpenAIRE ID) of Dutch research performing organisations. <p>Identifiers were collected for the following groups of Dutch RPOs (see Appendix A):</p> <ul> <li> <p>Universities, organised in Universities of the Netherlands (UNL, n=14);</p> </li> <li> <p>University Medical Centres, organised in the Dutch Federation of University Medical Centres (NFU, n=9); </p> </li> <li> <p>National research institutes under the umbrella organisation of the Foundation for Dutch Scientific Research Institutes (NWO-i, n= 9);</p> </li> <li> <p>Research institutes of the Royal Netherlands Academy of Arts and Sciences (KNAW, n=10);</p> </li> <li> <p>Universities of Applied Sciences affiliated to the Netherlands Association of Universities of Applied Sciences (Vereniging Hogescholen) (VH, n=35 of 37)<br><br></p> </li> </ul> </li> <li><em>openalex_works_20231223_rpo_nl_2022 </em>- record-level data of OpenAlex records retrieved for all Dutch RPOs in scope of the pilot (UNL/NFU, NWO-i, KNAW, VH) for publication year 2022<br><br></li> <li><em>openaire_products_20240116_rpo_nl_2022 - </em>record-level data of OpenAlex records retrieved for all Dutch RPOs in scope of the pilot (UNL/NFU, NWO-i, KNAW, VH) for publication year 2022</li> </ul>
Assessing Quality Variations in Early Career Researchers' Data Management Plans: Quantitative Data of the Content Analysis
<p>The data includes the numerical results of the ranking of the data management plans created during the Basics of Research Data Management (BRDM) courses worth 3 ECTS credits in the years 2020 - 2022. The ranking was made using the Finnish DMP Evaluation Guidance (https://doi.org/10.5281/zenodo.4729831). Additionally, the data contains the results of the analysis of the best RDM practices included in the DMPs.</p> <p>Note 1: The comma-separated coded CSV version 1 (5.2.2024) may not open correctly on MacOS. You can use the comma-delimited CSV file version 2 or 3 (31.5.2024).</p> <p>Note 2: Versions 1 (Quality_variations_in_ECRs_DMPs_data) and 3 (Quality_variations_in_ECRs_DMPs_data_ver_3) contain evaluations of DMPs, best practices for data management, as well as methods for data sharing, storage, and preservation. In version 2 (Quality_variations_in_ECRs_DMPs_data_ver_2), the methods for data sharing, storage, and preservation are missing.</p> <p>Data is related to the research article https://doi.org/10.2218/ijdc.v18i1.873.</p>
The evolution of gender monitoring and its challenges in Research and Innovation in Europe: She Figures reports analysis dataset
<p><span>The article delves into the European Commission's flagship initiative on gender monitoring in science and innovation, offering a responsible metrics perspective and drawing on equality policy literature. Over two decades, the initiative has evolved from competitiveness-related justifications to more transformative objectives related to equality policy evaluation, with the measurement areas and policy focus also undergoing changes. While there has been notable progress, the article points out a logic of invisibility in how dimensions and indicators are conceptualised and their data sources and interpretation. However, it also highlights a significant improvement in the information available. The article suggests that the contextualisation of the process could be enhanced to better integrate it into the policy-making cycle, a crucial area for further research. It concludes with proposals for future gender monitoring science and innovation. The aim is to offer an encouraging vision of monitoring that counts more on who is monitored and in opening up the debates instead of closing them.</span></p>
Research data for "Phase transitions in NiO during the Oxygen Evolution Reaction assessed via electrochromic phenomena through operando UV-Vis spectroscopy"
<p>This is the dataset supporting the publication "Phase transitions in NiO during the Oxygen Evolution Reaction assessed via electrochromic phenomena through operando UV-Vis spectroscopy", published by the authors in Electrochimica Acta (2024), 144626 under <a title="Persistent link using digital object identifier" href="https://doi.org/10.1016/j.electacta.2024.144626" target="_blank" rel="noreferrer noopener">doi.org/10.1016/j.electacta.2024.144626</a>.</p> <p>The data is ordered according to Figure numbers in the main manuscript and the electronic supplementary information.</p>
The Coverage of Basic and Applied Research in Press Releases in EurekAlert!
<p>This dataset contains data from our research into the coverage of basic and applied research in EurekAlert! press releases. The data covers the years 2015 to 2022 and includes a detailed record of press releases and their associated academic research.</p> <p>The following fields are described.</p> <p>EurekAlert_basic_applied_dataset:<br>EUID: unique identifier for each press release.<br>DOI: digital identifier for the associated research paper.<br>post time: The publish year of press release<br>RL: research level of research paper<br>LR_main_field: The academic field or discipline of the research.<br>Press Release Title:Title of the press release.<br>Paper Title: Title of the associated research paper. Abstract<br>Abstract:Abstract of the related research paper.<br>cosine_similarity:similarity score between press release title and paper title.<br>eu_flesch:Ease of reading of the full press release score<br>paper_flesch: Ease of reading score for abstracts of papers</p> <p>Institutional origins of press releases:<br>EUID: unique identifier for each press release.<br>DOI: digital identifier for the associated research paper.<br>RL: research level of research paper<br>LR_main_field: The academic field or discipline of the research.<br>Institution: The institutions issuing press releases<br>Affiliation of authos: The affiliations of paper author<br>Journal: The journal of paper</p> <p>Source:EurekAlert! and OpenAlex.</p> <p>Purpose: This dataset can be used to analyse the coverage of basic and applied research in press releases, as well as the types and fields of scientific research disseminated.</p>
Data gathered during the first and second stage of carrying out the NCN research project "Odmieńcy. Performances of otherness in the Polish transition culture" (year 2022 and 2023) - PI
<p>Data gathered during the first and secon phase of the project "Odmieńcy. Performances of otherness in Polish transition culture" by Dorota Sosnowska used as a basis for two papers: <a href="https://open.icm.edu.pl/items/9117fe29-528d-43ff-99c7-2e1fe0a8a93a">Blasted 1999. Sarah Kane’s Body Against the Archive (icm.edu.pl)</a> and <a href="https://open.icm.edu.pl/items/0aa7964c-b690-42af-8f7c-3d4fc62084b5">Brzydkie uczucia. O nudzie w sztuce i teatrze lat’ 90 (icm.edu.pl)</a></p>
Data gathered during the first and second stage of carrying out the NCN research project "Odmieńcy. Performances of otherness in the Polish transition culture" (year 2022 and 2023)
<p>Data gathered during the first and secon phase of the project "Odmieńcy. Performances of otherness in Polish transition culture" by Łukasz Kiełpiński used as a basis for two papers: <a href="../records/10625768">Zarządzanie ambiwalencją. Polski dyskurs ekspercki wokół HIV/AIDS na przełomie lat osiemdziesiątych i dziewięćdziesiątych XX wieku (zenodo.org)</a> and <a href="../records/10625805">Gra o sumie zerowej. Ekonomia wstydu w filmie "Pora na czarownice" (zenodo.org)</a></p>
Data from: Collaborative Research: Influence of phosphorus deficiency on enigmatic biological methane production in oxic freshwater lakes
<p>Data from: Collaborative Research: Influence of phosphorus deficiency on enigmatic biological methane production in oxic freshwater lakes</p> <p>NSF Projects 1951002 (PI: Matthew J. Church), 1950963 (PI: John E. Dore)</p>
Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research
<p><strong>This is the dataset of the report: Data Echoes: Tracking Data Availability and Integrity in Software Engineering Research</strong></p> <p>It contains the following information of all the papers from ASE, FSE, and ICSE in 2023:</p> <ul> <li>Paper title</li> <li>Keyword</li> <li>Is the source data available and accessible in the paper?</li> <li>If the source data is not available, do the authors explain why?</li> <li>Hosting platforms</li> <li>Access mode</li> <li>License</li> <li>Is their experiment data reused from previous work, or newly generated specifically for this study, or combination of both? </li> <li>Do the authors change/modify their experiment data before experiment?</li> <li>What modifications do they perform?</li> <li>Does the link provide detailed instructions about how to replicate their paper?</li> <li>Does the link contains their complete experiment data, their source code or other materials that are necessary to replicate their experiments?</li> <li>What's the data format inside the link?</li> <li>What's the content of the link?</li> </ul> <p> </p> <p>This is a course project and I collect the data in a rush.</p> <p>If you want to use this dataset and find any error, please contact me ;-)</p> <p>My email: echo.xiangchen@gmail.com</p>
CLDF dataset derived from Castro's "Sui Dialect Research" from 2015
<p>Cite the source of the dataset as:</p> <blockquote> <p>Castro, Andy and Pan, Xingwen (2015): Sui dialect research. SIL: Guiyang.</p> </blockquote>
Data and Statistical analysis for: "Predator in the pool? A quantitative evaluation of non-indexed open access journals in aquaculture research"
<p>Data and Statistical analysis for: "Predator in the pool? A quantitative evaluation of non-indexed open access journals in aquaculture research" published in <em>Frontiers in Marine Science</em></p>
Science ready spectra of star clusters and their best-fitting models described in the research paper "Using Star Clusters as Tracers of Star Formation and Chemical Evolution: the Chemical Enrichment History of the Large Magellanic Cloud" by Chilingarian & Asa'd
<p>Science ready spectra of star clusters in the Large Magellanic Cloud and their best-fitting templates (alpha-enhanced MILES based simple stellar population models) obtained using the NBursts full spectrum fitting code. Each spectrum is presented as a binary FITS table, which contains a spectrum (wavelength, flux, uncertainties), best-fitting template, best-fitting parameters (radial velocity, age, metallicity), and a pixel mask used in the fitting procedure. For each cluster, 5 spectra are provided, which correspond to [alpha/Fe] values from 0.0 to 0.4 dex with a step of 0.1 dex. The only exception is NGC2249, for which only 3 models are provided. The alpha-enhancement value of a model grid used in the fitting procedure is given in the FITS keyword MGFEGRID.</p>
Spectral irradiance at Lammi Biological Station Research Forest 2015: for assessing scale-wise similarity of curves with a thick pen
<p>This dataset contains records of the solar spectral energy irradiance (W m<sup>-2</sup> nm<sup>-1</sup>) in the understorey of forest stands at Lammi Biological Station, southern Finland (61◦ 3.24’ N, 25◦ 118 2.23’ E) during the spring of 2015. These spectra allow the change in spectral energy irradiance to be followed through the period of canopy leaf flush. Records are the average of recorded spectra from four points recorded at 40-cm above the forest floor using a Maya 2000 Pro array spectrometer. Spectra were recorded from exactly the same location on three dates, 2015-04-25, 2015-05-22, and 2015-06-05, before, during and after leaf flush. Data were recorded from the understorey of a young Betula stand, an old Betula stand, an old mixed Betula stand, a Quercus stand, and a Picea stand, in three positions: shade, semi-shade from leaves, and full sun in a sunfleck. On each occasion control measurements of spectral energy irradiance in full sun of an open field were also recorded at the beginning, middle and end of each measurement period. All measurements were made during the 2 hours either side of solar noon, on clear-sky days. Details of the sampling method and interpretation are given in the paper, Hartikainen et al., (2018) in Ecology and Evolution, which showcases the use of Thick Pen Transform to compare spectra.</p>
Identified Charcoal Hearths from "Slope Analysis of 'Digital Elevation Model for Blue Mountain Charcoal Research Project'"
<p>This is a GeoJSON file that lists all of the potential charcoal hearths along the Blue Mountain of eastern Pennsylvania. For a detailed description of how this data was produced, please see:</p> <p>Carter, Benjamin. (2018, May 29). Description of Methods for Identifying Charcoal Hearths along the Blue Mountain of Pennsylvania. (Version 0.1.0). Zenodo. http://doi.org/10.5281/zenodo.1255101</p> <p>These hearths were identified using this data:</p> <p>Carter, Benjamin. (2018). Slope Analysis of "Digital Elevation Model for Blue Mountain Charcoal Research Project" (Version 0.1.0). Zenodo. http://doi.org/10.5281/zenodo.1252977</p> <p>The above is derived from:</p> <p>Carter, Benjamin P. (2018). Digital Elevation Model for Blue Mountain Charcoal Research Project (Version 0.1.0). Zenodo. http://doi.org/10.5281/zenodo.1252441</p> <p> </p>
Blue Mountain Charcoal Project Research Area
<p>This GEOJSON polygon identifies the area in which Dr. Benjamin Carter (Muhlenberg College) and students have focused their efforts in an attempt to use remote sensing and field work to identify charcoal hearths from the 19th century. The main destination for the charcoal was the furnaces and forge of the Balliet family (Lehigh and East Penn Furances and East Penn Forge) in East Penn, Carbon County and Washington, Lehigh County (Pennsylvania, USA). This polygon defines an area that includes much of Pennsylvania State Gamelands #217 but also surrounding privately owned areas that show limited recent impacts by people (especially avoiding homes and roads). It extends from the Lehigh Gap to Pennsylvania Route 309. While there is limited evidence of hearths to the east of the Lehigh Gap, there are definitely more hearths to the southwest of route 309.</p>
WSSSPE5.1 presentations: Research software sustainability activity analysis
<p>Distribution of actors, actions and actees from <a href="https://www.slideshare.net/danielskatz/research-software-sustainability-wssspe-urssi">Daniel S. Katz' schematic of research software sustainability</a> over presentations given at the Workshop on Sustainable Software for Science: Practice and Experiences (WSSSPE5.1) on 6 September 2017 in Manchester, UK (Proceedings: <a href="https://doi.org/10.6084/m9.figshare.c.3869782.v3">https://doi.org/10.6084/m9.figshare.c.3869782.v3</a>).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.