Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,025
datasets available to search
ShareScore release 0.9.0
Dataset results
6,025 results for “science”
iGEM: a model system for team science and innovation
<p>This dataset is extracted from the <strong>international Genetically Engineered Machine (iGEM) competition </strong>between years 2008 and 2018, and can be used as a model system for studying team science and innovation. It is described at length in <a href="https://arxiv.org/abs/2310.19858">this article</a>.</p> <p>The dataset encompass detailed records from the iGEM competition, capturing various aspects of team participation and achievements. Specifically, the <strong>Team Information</strong> dataset (<strong>teams_table.csv</strong>) provides insights into team characteristics and achievements, including medal status and region of origin. <strong>User Information</strong> (<strong>users_table.csv</strong>) offers a look into individual participants, detailing their roles in the team. <strong>Awards Information</strong> (<strong>awards_table.csv</strong>) and <strong>Medal Criteria</strong> (<strong>medals_criteria.csv</strong>) lay out the awards teams have garnered and the standards for medal attainment. The <strong>BioBricks Information</strong> (<strong>biobricks_table.csv</strong>) corresponds to the BioBrick sequences associated with each team, while <strong>Wiki Edits</strong> (<strong>wikis_table.csv</strong>) tracks the changes made by users on their team's (wiki) lab notebook. Finally, the <strong>Collaboration Network</strong> (<strong>collaboration_network.csv</strong>) corresponds to the weighted directed inter-team collaboration network collected using team mentions across team wikis.</p> <p>In addition to the structured dataframes above, we provide in <strong>team_wikis_full_text.zip</strong> the full texts of the wiki pages from the digital laboratory notebooks collaboratively edited by iGEM teams in the forms of wiki instances. There is a folder for each year from 2008-2018 and within which there are individual folders for each team. Each team folder has a file denoting the pagelist and two files for each page. One file is the html content, and the other the text content, extracted using the "KeepEverythingExtractor" option in the <em>boilerpipe.extract</em> library for processing and removing boilerplate content after webscraping.</p> <p> </p>
Related Works for the National Science Foundation of Sri Lanka and the National Sleep Foundation retrieved from the DataCite Commons.
<p>These data were retrieved in order to help understand the use of funder identifiers associated with common acronyms like NSF. They include the following fields: doi, registrationAgency, type, publisher, publicationYear, funderName, funderIdentifier, and awardNumber retrieved from DataCite Commons using the query</p> <div> <div>{organization(id: "' + ror + '") {id name works(first:2000) { totalCount pageInfo {endCursor hasNextPage} nodes {doi type registrationAgency {name} publisher {name} publicationYear fundingReferences {funderName funderIdentifier awardNumber}}}}}'</div> <div> </div> <div>The files are identifier with RORs:</div> <div><a href="../api/records/11116776/draft/files/00zc1hf95_relatedWorks_20240505_10.csv/content" target="_blank" rel="noopener noreferrer">00zc1hf95_relatedWorks_20240505_10.csv</a> are data for the National Sleep Foundation</div> <div><a href="../api/records/11116776/draft/files/010xaa060_relatedWorks_20240505_10.csv/content" target="_blank" rel="noopener noreferrer">010xaa060_relatedWorks_20240505_10.csv</a> are data for the National Science Foundation of Sri Lanka</div> <div> </div> <div>A blog post describing this work is at https://metadatagamechangers.com/blog/2024/4/12/funder-acronyms-are-still-not-enough</div> <div> </div> </div>
Research data management in the German-speaking Sports Sciences - Survey on the Status Quo
<p>The data set contains survey data on the status quo of research data management within the German-speaking sports science community. The survey was conducted as an online survey in the period from August 16<sup>th</sup> to September 30<sup>th</sup>, 2023.</p>
Birdwatching, eBird and citizen science in India: qualitative interviews with participants, practitioners and ecologists
<h1>Abstract</h1> <p>This study consists of qualitative interviews about birdwatching, citizen science, and the use of the birdwatching data platform <em>eBird </em>in India. Interview partners are birdwatchers, citizen science practitioners, and ecologists who have used eBird data. Some of the main topics covered include: the nature of the birdwatching community and styles of birdwatching in India; the history of the adoption of eBird in India; the value of birdwatching and citizen science; challenges involved in conducting or participating in citizen science; opportunities and limitations of using data from eBird and citizen science; processes of data collection and quality control in eBird; and ecological research, conservation priorities, and environmental activism in India. This study is part of the project A Philosophy of Open Science for Diverse Research Environments (PHIL_OS).</p> <h1>Methods</h1> <p>The data in this study was collected using semi-structured qualitative interviews.</p> <p>Interview partners were recruited by snowball sampling through their engagement with eBird India and related organisations. There were 17 interview partners, interviewed either once or several times. 19 interviews were conducted in total.</p> <p>Interview guides/questionnaires were designed for each interviewee depending on their status as birdwatchers, citizen science coordinators, and eBird data users.</p> <p>Interviews were conducted between April 2022 and June 2023. The interviews took place online using Zoom videoconferencing software. Interviews lasted 35-70 minutes. When participants provided their written consent, interviews were audio-recorded and transcribed smart verbatim using otter.ai and manual proofreading. Sensitive information was removed before publishing transcripts.</p> <p>Transcripts were analysed using semi-grounded coding. Codes were organised into parent codes using an inductive approach based on emergent categories.</p> <h1>Description of the data and file structure</h1> <p>Documentation files include interview guides, the information sheet and consent form, ethics approval, and the data narrative. Documentation files are named according to the structure: authorname_filename_DOCUMENTATION.</p> <p>Data files consist of a summary of participants, 17 of the interview transcripts, and a code list. Interview transcript files are named according to the structure: authorname_interviewnumber_date.</p> <p>A full list of files is provided in the README file.</p> <h1>Notes</h1> <p>This study was conducted as part of the project A Philosophy of Open Science for Diverse Research Environments (PHIL_OS). More information can be found at <a href="https://opensciencestudies.eu/">https://opensciencestudies.eu</a></p> <p>This project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (grant agreement No. 101001145).</p>
Reliance on Science
<p>This dataset contains <strong>patent-to-paper citations</strong> through 2023 as well as <strong>patent-paper</strong> <strong>pairs</strong>.</p> <p> <em>If you use the citations data, please cite </em>these two articles:</p> <p><strong>1. M. Marx & A. Fuegi, "Reliance on Science by Inventors: Hybrid Extraction of In-text Patent-to-Article Citations." </strong> <em>forthcoming in Journal of Economics and Management Strategy. </em>(<a href="http://doi.org/10.1111/jems.12455">http://doi.org/10.1111/jems.12455</a>)</p> <p><strong>2. M. Marx, & A. Fuegi, "Reliance on Science: Worldwide Front-Page Patent Citations to Scientific Articles" (2020), </strong><em><strong>Strategic Management Journal 41(9):1572-1594</strong></em><strong>. (</strong><a href="https://onlinelibrary.wiley.com/doi/full/10.1002/smj.3145">https://onlinelibrary.wiley.com/doi/full/10.1002/smj.3145</a><strong>) </strong></p> <p> <em>If you use the patent-paper-pairs data, please cite </em>this article:</p> <p><strong>1. M. Marx & E. Scharfmann, "Does Patenting Promote the Progress of Science? Evidence from Patent-Paper Pairs." </strong> </p> <p>The datafile containing the citations is <strong>_pcs_oa.csv. </strong> Each citation has the applicant/examiner flag, confidence score (1-10), whether the reference was a) only on the front page, b) only in the body text, or c) in both.</p> <p>The datafile containing the patent-paper pairs (PPPs) is <strong>_patent_paper_pairs.csv</strong>. These are USPTO only, through 2022. Each PPP has a confidence score and the count of days between the publication of the paper and the filing of the patent. (If the patent is a continuation of another patent, the filing date of the original patent is used.) Also, when a paper is paired with multiple patents, an indicator variable reports whether those patents are continuations or otherwise identical. </p> <p>The above is documented in greater detail in <strong>__relianceonscience2024.pdf.</strong></p> <p>These data are provided under a Creative Commons Attribution Non-Commercial license. Please contact us regarding commercial use. Questions & feedback to <a href="mailto:support@relianceonscience.org">support@relianceonscience.org</a><em>.</em></p> <p><em><strong>This work is sponsored by the Alfred P. Sloan Foundation grant #G-2021-16822.</strong></em></p>
Community science approach reveals temporal and eutrophication-related spatial patterns in bladderwrack-associated invertebrate fauna
<p>Data related to the "Community science approach reveals temporal and eutrophication-related spatial patterns in bladderwrack-associated invertebrate fauna" paper by Salo, Nieminen, Salovius-Laurén and Rinne published in Estuarine, Coastal and Shelf Science in 2024. <a href="https://doi.org/10.1016/j.ecss.2024.108822">https://doi.org/10.1016/j.ecss.2024.108822</a></p> <p>The data describes the community data collected with the community science method described in the paper. </p>
Recordings Q&A and matchmaking sessions for call for proposals 'Open Science Infrastructure'
<p>On Thursday, July 11, and Tuesday, July 16, 2024, Open Science NL organised two online Q&A sessions combined with a matchmaking opportunity for the Open Science NL call 'Open Science Infrastructure'.</p> <p>These are the two recordings of the two Q&A sessions. The Open Science NL team has drafted a Frequently Asked Questions document addressing all the questions that came up during the meetings. This is added as a seperate text-file (PDF). The slides presented during both meetings are shared as well as a PDF.</p> <p>For more information about the call and how to apply, please go to: <a href="https://www.openscience.nl/en/calls/open-science-infrastructure" target="_blank" rel="noopener">https://www.openscience.nl/en/calls/open-science-infrastructure</a></p>
Interviews with editors of library science journals on transitioning to open access
<p>These three files are related to qualitative, semi-structured interviews conducted in Fall 2023 with editors of Library and Information Science (LIS) journals on transitioning to open access. One subgroup consisted of participants who were editors at the time of an LIS journal when it transitioned (or flipped) to an open access model that does not charge a fee to either readers or authors (which this study refers to as equitable open access), and the other subgroup consisted of current editors (at the time) of LIS journals that have not yet transitioned (or unflipped) to an equitable open access model. Two of the files are the interview protocols for each group of flipped and unflipped editors, and the third file is the codebook the researchers used to analyze the interview transcripts. Interview transcripts are not being publicly shared to ensure confidentiality for interview participants.</p> <p>The interview protocols were created based on the findings of a prior research study:</p> <p>Borchardt, R., Dawson, D., & Schultz, T. (2024). Financial and other perceived barriers to transitioning to an equitable no-publishing fee open access model: A survey of LIS journal editors. College & Research Libraries, 85(1). <a href="https://doi.org/10.5860/crl.85.1.96">https://doi.org/10.5860/crl.85.1.96</a></p> <p>The codebook was created iteratively based on the researchers' review and analysis of the interview transcripts.</p>
[DATA_SCIENCE] Interviews PomBase Users, January-February 2016
<p>Here you find the transcripts of interviews collected by Sabina Leonelli as part of the ERC project "The Epistemology of Data-Intensive Science". You also find the information sheet provided to interviewees, which gives you the context for this project. Further information and related publications can be found at www.datastudies.eu. One paper that specifically makes use of these interviews was published by Sabina Leonelli in the journal Philosophy of Science in 2018, under the title "Data in Time: Time-Scales of Data Use in the Life Sciences." The transcripts document yeast researchers' attitudes to data curation and the use of databases in their field. Researchers have consented to have these transcripts made available as Open Data. Other interviewees did not give consent, so those transcripts are held securely by the research team in Exeter.</p>
Hypertension - Florida Annotated Corpus for Translational Science (FACTS), Vital Sign Ontology Annotations
<p>Florida Annotated Corpus for Translational Science (FACTS), which currently consists of 20 case reports about hypertension annotated with Vital Sign Ontology (VSO) classes (version 2012-04-25). </p>
Science ready spectra of star clusters and their best-fitting models described in the research paper "Using Star Clusters as Tracers of Star Formation and Chemical Evolution: the Chemical Enrichment History of the Large Magellanic Cloud" by Chilingarian & Asa'd
<p>Science ready spectra of star clusters in the Large Magellanic Cloud and their best-fitting templates (alpha-enhanced MILES based simple stellar population models) obtained using the NBursts full spectrum fitting code. Each spectrum is presented as a binary FITS table, which contains a spectrum (wavelength, flux, uncertainties), best-fitting template, best-fitting parameters (radial velocity, age, metallicity), and a pixel mask used in the fitting procedure. For each cluster, 5 spectra are provided, which correspond to [alpha/Fe] values from 0.0 to 0.4 dex with a step of 0.1 dex. The only exception is NGC2249, for which only 3 models are provided. The alpha-enhancement value of a model grid used in the fitting procedure is given in the FITS keyword MGFEGRID.</p>
Aurorasaurus Real-Time Citizen Science Aurora Data
<p>Aurorasaurus citizen science data is a collection of auroral sightings submitted to the project via its website (aurorasaurus.org) or apps and mined from social media. It is a robust data set and particularly abundant during strong geomagnetic storms. This data is offered to the scientific community for research use through an open-access database in its raw and scientific formats for the 2015-2016 period, each of which is described in detail in the following technical report:</p> <p>Kosar, B. C., MacDonald, E. A., Case, N. A., & Heavner, M. (2018). Aurorasaurus Database of Real‐Time, Crowd‐Sourced Aurora Data for Space Weather Research. <em>Earth and Space Science</em>, <em>5</em>(12), 970-980.</p> <p>For more information on the project, please contact the project leaders at aurorasaurus.info@gmail.com.</p> <p> </p>
Data processing scripts and images from e-MERLIN project CY6213 used in Ghirlanda et al. 2019, Science
<p>Data processing scripts and images from e-MERLIN project CY6213 used in Ghirlanda et al. 2019, Science</p> <p> </p> <ul> <li>info.txt contains a general description of how the data was processed and the main results.</li> <li>pipeline.tar contains the data pipeline used to process the e-MERLIN observations</li> <li>imaging.py is the script used to produce the final images</li> <li>CY6213_images.tar contain the final images of the target source (not corrected by calibration factor, described in the imaging script).</li> </ul> <p> </p>
Data Potential Bias in Peer Review of Grant Applications at the Swiss National Science Foundation
<p>Potential biases in the peer review of grant applications at the Swiss National Science Foundation.</p> <p> </p>
Science ready spectra and their best-fitting models described in the research paper ``Internal dynamics and stellar content of nine ultra-diffuse galaxies in the Coma cluster prove their evolutionary link with dwarf early-type galaxies'' by Chilingarian et al.
<p>Science ready spectra of nine ultra-diffuse galaxies in the Coma cluster collected with the Binospec multi-object spectrograph and their best-fitting PEGASE.HR templates obtained using the NBursts full spectrum fitting code. These spectra were presented in the paper ``Internal dynamics and stellar content of nine ultra-diffuse galaxies in the Coma cluster prove their evolutionary link with dwarf early-type galaxies'' by Chilingarian et al. accepted for publication in the Astrophysical Journal on Sep/3/2019 (arXiv:1901.05489).</p> <p>Each spectrum is presented as a binary FITS table, which contains a spectrum (wavelength, flux, uncertainties), best-fitting template, best-fitting parameters (radial velocity, age, metallicity), and a pixel mask used in the fitting procedure. For six galaxies there are two files provided: (i) one-dimensional optimally extracted integrated spectrum and (ii) two dimensional spectrum for spatially resolved radial velocity information. For the remaining three galaxies, only spatially resolved spectra are provided.</p>
Citizen Science projects in Argentina
<p>Citizen science activities recognized in Argentina.</p> <p>The code for the Actions column are: </p> <ul> <li>e -Collect <p>c - Hypothesis design</p> <p>d - Design collection strategies</p> f - Sample analysis</li> <li>g - Data analysis <p>h - Generate conclusions</p> <p>j - Generate new questions</p> k - Digitalization <p>i - Disseminate conclussions</p> </li> </ul>
Supporting Material for article "The ELIXIR Core Data Resources: fundamental infrastructure for the life sciences"
<p>This data set is the Supporting Material referred to in the Supplementary Data for the article "The ELIXIR Core Data Resources: fundamental infrastructure for the life sciences" (Drysdale, et al.) submitted for publication in April 2019.</p> <p> </p>
AirHeritage Datalake: Multi-site, Multi-season, Multi Unit dataset including Fixed and Mobile Citizen science data from networked Air Quality Low-Cost Multi-Sensors devices and reference stations
<p>This datalake comprises several datasets from <strong>37 networked low cost air quality multisensors</strong> (<strong>30</strong> <strong>mobile</strong> ENEA MONICA(tm) + <strong>7</strong> <strong>fixed</strong>) along with <strong>3</strong> (fixed) + <strong>1</strong> (mobile) <strong>reference stations</strong> operated by Campania Regional Envronmental Protection Agency. The datalake is organized in 3 main directories respectively related to fixed nodes, mobile nodes and nearby reference stations including a mobile laboratory used for colocation campaigns; each subdirectory include its own metadata description file.</p> <p>Data, curated by Energy and Data Science Laboratory of ENEA, include multi-weeks colocation periods when low cost devices have been colocated with reference stations as well as operational periods during which sensors are deployed for fixed or mobile monitoring campaigns. Data have been recorded during 2021 and 2022 in a<strong> pervasive, multi-site, multi-seasonal deployment</strong> in Portici, a densely populated small area city (4km2, 55k + inhabitants) located 7km south of Naples, Italy.</p> <p>The datalake consists in actual sensors and reference intrumentations timeseries along with metadata description files with deployment dates and location data. The dataset files include high sampling frequency raw sensor data of quality-controlled sensor network along with co-located reference stations data sets. Sensor data include electrochemical sensors data (intended target pollutants: NO2, O3, CO), Optical sensor data (PM2.5, PM10, PM1) readings along with meteorological parameters. .</p> <p>Further description of sensors and reference instruments are reported in the accompanying paper (see citation request).</p> <p>The dataset can be used for </p> <ul> <li> <strong>advanced (remote/universal/in field) data driven calibration strategies</strong> test or development including <strong>machine learning </strong>models</li> <li><strong>mobile opportunistic data fusion</strong> methods development</li> <li><strong>geomatics and data assimilation</strong> models studies</li> </ul> <p>as well as low cost sensor characterization performance studies. </p>
Electronic representation of Russian journals on Earth Sciences
<p>The dataset describes the deprth of digital archives of the most authoritative Russian journal on Earth Sciences. The five figures demostrate the results of graphical processing while preparing digital archives of Geologiya i Geofizika and Zapiski Gornogo Instituta journals, as well as the forms of metadata presentation in electronic archives.</p>
Participant survey data from the citizen science project FLOW, 2021
<p>This dataset is linked to the following publication:</p> <p>von Gönner, J., Masson, T., Köhler, S., Fritsche, I., Bonn, A. (in press): Citizen science promotes knowledge, skills and collective action to monitor and protect freshwater streams. People and Nature.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.