Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
191
datasets available to search
ShareScore release 0.9.0
Dataset results
191 results for “Open Research”
Supplementary material 1 from: Heuer K, Ghosh S, Robinson Sterling A, Toro R (2016) Open Neuroimaging Laboratory. Research Ideas and Outcomes 2: e9113. https://doi.org/10.3897/rio.2.e9113
Animation showing the different functionalities of BrainBox: opening a Magnetic Resonance Imaging volume, viewing it, and editing it collaboratively online.
Dataset of the study "Open Access and Data Sharing in Cancer Stem Cells research"
<p>Dataset of the study "Open Access and Data Sharing in Cancer Stem Cells research"</p>
Perceived Barriers to Open Science among Researchers in Mathematics, Natural Sciences, and Cognitive Sciences
<p><strong>General Information:</strong></p> <p>This dataset contains artifacts related to Riedel et al. (2024) "Perceived Barriers to Open Science among Researchers in Mathematics, Natural Sciences, and Cognitive Sciences", presented at the Hochschule 2034 Workshop (DOI WILL BE ADDED). Here, we investigate the perceived barriers to open science practices by researchers. This dataset contains the a XLSX file containing the questionnaire’s answers, and a Jupyter notebook script to evaluate the given data. </p>
pyKNEEr: An image analysis workflow for open and reproducible research on femoral knee cartilage - Validation data
<p>Image data used in the paper introducing pyKNEEr. Explanations about these data are in the <a href="https://github.com/sbonaretti/pyKNEEr/tree/master/publication">GitHub</a> repository</p> <p>Changes in version 0.2.0: </p> <p>- Added inHouse images </p> <p>- Segmented images casted to int16 for smaller file size</p>
Results from the Open Call: How Citizens can participate in solar energy research?
<p>Results from the public consultation that GRECO Project (H2020-787289) launched in November 2018. These responses were collected from around 70 research teams along the world. The document is structured in three sections to think beyond researchers' needs. Any kind of successful collaboration implies that both parts must have a benefit and the consultation was also focused on such perspective.</p>
Demystifying Open Science and Research Data Management: A practical workshop for researchers
<p>Have you ever wondered why is everyone discussing about Open Science and Research Data Management (RDM) these days? Good Research Data Management (RDM) is crucial for reproducible and robust scientific research. Consequently, more and more funding bodies, governments, research institutions and other agencies have emphasised the value and importance of good data management and introduced policies on data management and sharing. However, it is easier said than done.<br> Like most researchers you might have more questions than answers about the topic. You are not alone! Come join us on the 24th at an interactive workshop where you can understand the why and how of open science and research data management in practical terms.</p> <p> </p> <p>This zenodo entry is a recording of the workshop described above. </p>
Online Appendix for the HICSS2025 Paper: Bridging Citizens and Public Sector Employees through an Open Employee-driven Innovation Process - A Design Science Research Study
<p>Online Appendix for the HICSS2025 Paper: Bridging Citizens and Public Sector Employees through an Open Employee-driven Innovation Process - A Design Science Research Study</p>
Open Science - Group 02 - Research Integrity
<p>Provision of content on Research Integrity in the open science discipline at the State University of Maringá (UEM).</p>
Dataset for How open is innovation research?–An empirical analysis of data sharing among innovation scholars
<p>This dataset comprises responses from 241 innovation researchers on their personal data sharing behavior as well as their perceptions of and attitudes towards open research data. This dataset is the supplementary material to Barczak et al. (2021) (<a href="https://doi.org/10.1080/13662716.2021.1967727" target="_blank" rel="noopener">https://doi.org/10.1080/13662716.2021.1967727</a>).</p>
MonkeyPox2022Tweets: A Large-Scale Twitter Dataset on the 2022 Monkeypox Outbreak, Findings from Analysis of Tweets, and Open Research Questions
<p><strong>Please cite the following paper when using this dataset:</strong></p> <p>N. Thakur, “MonkeyPox2022Tweets: A large-scale Twitter dataset on the 2022 Monkeypox outbreak, findings from analysis of Tweets, and open research questions,” Infect. Dis. Rep., vol. 14, no. 6, pp. 855–883, 2022, DOI: https://doi.org/10.3390/idr14060087</p> <p><strong>Abstract</strong></p> <p>The mining of Tweets to develop datasets on recent issues, global challenges, pandemics, virus outbreaks, emerging technologies, and trending matters has been of significant interest to the scientific community in the recent past, as such datasets serve as a rich data resource for the investigation of different research questions. Furthermore, the virus outbreaks of the past, such as COVID-19, Ebola, Zika virus, and flu, just to name a few, were associated with various works related to the analysis of the multimodal components of Tweets to infer the different characteristics of conversations on Twitter related to these respective outbreaks. The ongoing outbreak of the monkeypox virus, declared a Global Public Health Emergency (GPHE) by the World Health Organization (WHO), has resulted in a surge of conversations about this outbreak on Twitter, which is resulting in the generation of tremendous amounts of Big Data. There has been no prior work in this field thus far that has focused on mining such conversations to develop a Twitter dataset. Therefore, this work presents an open-access dataset of <strong>571,831 Tweets</strong> about monkeypox that have been posted on Twitter since the first detected case of this outbreak on May 7, 2022. The dataset complies with the privacy policy, developer agreement, and guidelines for content redistribution of Twitter, as well as with the FAIR principles (Findability, Accessibility, Interoperability, and Reusability) principles for scientific data management.</p> <p> <strong>Data Description</strong></p> <p>The dataset consists of a total of <strong>571,831 Tweet IDs</strong> of the same number of tweets about monkeypox that were posted on Twitter from 7th May 2022 to 11th November (the most recent date at the time of uploading the most recent version of the dataset). The Tweet IDs are presented in 12 different .txt files based on the timelines of the associated tweets. The following represents the details of these dataset files.</p> <ul> <li>Filename: TweetIDs_Part1.txt (No. of Tweet IDs: 13926, Date Range of the associated Tweet IDs: May 7, 2022, to May 21, 2022)</li> <li>Filename: TweetIDs_Part2.txt (No. of Tweet IDs: 17705, Date Range of the associated Tweet IDs: May 21, 2022, to May 27, 2022)</li> <li>Filename: TweetIDs_Part3.txt (No. of Tweet IDs: 17585, Date Range of the associated Tweet IDs: May 27, 2022, to June 5, 2022)</li> <li>Filename: TweetIDs_Part4.txt (No. of Tweet IDs: 19718, Date Range of the associated Tweet IDs: June 5, 2022, to June 11, 2022)</li> <li>Filename: TweetIDs_Part5.txt (No. of Tweet IDs: 46718, Date Range of the associated Tweet IDs: June 12, 2022, to June 30, 2022)</li> <li>Filename: TweetIDs_Part6.txt (No. of Tweet IDs: 138711, Date Range of the associated Tweet IDs: July 1, 2022, to July 23, 2022)</li> <li>Filename: TweetIDs_Part7.txt (No. of Tweet IDs: 105890, Date Range of the associated Tweet IDs: July 24, 2022, to July 31, 2022)</li> <li>Filename: TweetIDs_Part8.txt (No. of Tweet IDs: 93959, Date Range of the associated Tweet IDs: August 1, 2022, to August 9, 2022)</li> <li>Filename: TweetIDs_Part9.txt (No. of Tweet IDs: 50832, Date Range of the associated Tweet IDs: August 10, 2022, to August 24, 2022)</li> <li>Filename: TweetIDs_Part10.txt (No. of Tweet IDs: 39042, Date Range of the associated Tweet IDs: August 25, 2022, to September 19, 2022)</li> <li>Filename: TweetIDs_Part11.txt (No. of Tweet IDs: 12341, Date Range of the associated Tweet IDs: September 20, 2022, to October 9, 2022)</li> <li>Filename: TweetIDs_Part12.txt (No. of Tweet IDs: 15404, Date Range of the associated Tweet IDs: October 10, 2022, to November 11, 2022) </li> </ul> <p>Please note: The dataset contains only Tweet IDs in compliance with the terms and conditions mentioned in the privacy policy, developer agreement, and guidelines for content redistribution of Twitter. The Tweet IDs need to be hydrated to be used. For hydrating this dataset, the <a href="https://github.com/DocNow/hydrator/releases">Hydrator application</a> may be used (a step-by-step process on how to use Hydrator to hydrate this dataset is explained in the above-mentioned paper). </p>
The LOTUS Initiative for Open Natural Products Research: biological and chemical trees
<p>Biological and chemical trees made from frozen metadata (<a href="https://doi.org/10.5281/zenodo.5794106">10.5281/zenodo.5794106</a>) (for example, for PubChem)</p>
Open access and international co-authorship: a longitudinal study of the United Arab Emirates research output
<p>Enriched data from Scopus used for the article: Open access and international co-authorship: a longitudinal study of the United Arab Emirates research output</p>
Open Research
<p>Open Research on "Advantages of GSMaP data for Multi-timescale Precipitation Estimation in Luzon"</p>
Open Science resources by research discipline
<p>This is the underlying data set to our map of Open Science resources by discipline: <a href="https://kumu.io/access2perspectives/open-science#disciplines">https://kumu.io/access2perspectives/open-science#disciplines</a></p> <p>For context and related resources please refer to <a href="https://access2perspectives.pubpub.org/open-science">https://access2perspectives.pubpub.org/open-science</a></p> <p>Explore the map by clicking on individual nodes and see descriptions, related websites, and other information about the resources.<br> With the navigation buttons to the right, you can zoom in and out, and select and focus on specific elements.</p> <p>Learn more about our work at <a href="http://access2perspectives.org">access2perspectives.org</a><br> If you have comments, questions or suggestions for improvements on this map email us at info@access2perspectives.org.</p> <p>LICENSE: <strong>Creative Commons <a href="https://creativecommons.org/licenses/by-sa/4.0/">Attribution-ShareAlike 4.0 International</a></strong></p>
Hamouda GRL Open Research
<p>DATA STATEMENT FOR GRL ARTICLE: On Convection Precipitation during Vb-events in present and Warmer Climate.</p> <p>*run_CPS_spate_dkrz_SPATEsetup.sh: is the namelist to downscale ERA-Interim to convective scale (CPS) using COSMO-CLM<br> *run_int2lm_SPATE_CPS: is the namelist to prepare input data for CPS downscaling using COSMO-CLM<br> *run_int2lm_CMIP6.sh: is the namelist to prepare input data for downscaling with convection parameterization using COSMO-CLM<br> *f1 and f2 are from Poujol et al. 2019 to detect convection.<br> *clusterfx.py is a python function to cluster GPH data using k-means.<br> * Vb-tracks are provided in the text files for ERAI and CMIP6 historical and SSP585.<br> </p>
Data from: tableone: An open source Python package for producing summary statistics for research papers
Open the record for dataset details and reuse information.
AndroZooOpen: Collecting Large-scale Open Source Android Apps for the Research Community
<p>The raw collected data for the paper</p>
COVID-19 Open Research Dataset (CORD-19)
<p><strong>Important</strong>: This dataset is updated regularly and the latest version for download can be found <a href="https://www.semanticscholar.org/cord19/download">here</a>.</p> <p>In response to the COVID-19 pandemic, the <a href="https://allenai.org/">Allen Institute for AI</a> has partnered with leading research groups to prepare and distribute the COVID-19 Open Research Dataset (CORD-19), a free resource of scholarly articles, including full text content, about COVID-19 and the coronavirus family of viruses for use by the global research community.</p> <p>This dataset is intended to mobilize researchers to apply recent advances in natural language processing to generate new insights in support of the fight against this infectious disease. The corpus will be updated weekly as new research is published in peer-reviewed publications and archival services like <a href="https://www.biorxiv.org/">bioRxiv</a>, <a href="https://www.medrxiv.org/">medRxiv</a>, and others.</p> <p>By downloading this dataset you are agreeing to the Dataset license. Specific licensing information for individual articles in the dataset is available in the metadata file.</p> <p>Additional licensing information is available on the <a href="https://www.ncbi.nlm.nih.gov/pmc/tools/openftlist/">PMC website</a>, <a href="https://www.medrxiv.org/submit-a-manuscript">medRxiv website</a> and <a href="https://www.biorxiv.org/about-biorxiv">bioRxiv website</a>.</p> <p><strong>Dataset content:</strong></p> <ul> <li>Commercial use subset</li> <li>Non-commercial use subset</li> <li>PMC custom license subset</li> <li>bioRxiv/medRxiv subset (pre-prints that are not peer reviewed)</li> <li>Metadata file</li> <li>Readme</li> </ul> <p>Each paper is represented as a single JSON object (see schema file for details).</p> <p><strong>Description:</strong></p> <p>The dataset contains all COVID-19 and coronavirus-related research (e.g. SARS, MERS, etc.) from the following sources:</p> <ul> <li>PubMed's PMC open access corpus using this <a href="https://www.ncbi.nlm.nih.gov/pmc/?term=%22COVID-19%22+OR+Coronavirus+OR+%22Corona+virus%22+OR+%222019-nCoV%22+OR+%22SARS-CoV%22+OR+%22MERS-CoV%22+OR+%E2%80%9CSevere+Acute+Respiratory+Syndrome%E2%80%9D+OR+%E2%80%9CMiddle+East+Respiratory+Syndrome%E2%80%9D">query</a> (COVID-19 and coronavirus research)</li> <li>Additional COVID-19 research articles from a corpus maintained by the <a href="https://www.who.int/emergencies/diseases/novel-coronavirus-2019/global-research-on-novel-coronavirus-2019-ncov">WHO</a></li> <li>bioRxiv and medRxiv pre-prints using the same query as PMC (COVID-19 and coronavirus research)</li> </ul> <p>We also provide a comprehensive metadata file of coronavirus and COVID-19 research articles with links to <a href="https://www.ncbi.nlm.nih.gov/pmc/?term=%22COVID-19%22+OR+Coronavirus+OR+%22Corona+virus%22+OR+%222019-nCoV%22+OR+%22SARS-CoV%22+OR+%22MERS-CoV%22+OR+%E2%80%9CSevere+Acute+Respiratory+Syndrome%E2%80%9D+OR+%E2%80%9CMiddle+East+Respiratory+Syndrome%E2%80%9D">PubMed</a>, <a href="https://aka.ms/AA7q3eb">Microsoft Academic</a> and the <a href="https://www.who.int/emergencies/diseases/novel-coronavirus-2019/global-research-on-novel-coronavirus-2019-ncov">WHO COVID-19 database of publications</a> (includes articles without open access full text).</p> <p>We recommend using metadata from the comprehensive file when available, instead of parsed metadata in the dataset. Please note the dataset may contain multiple entries for individual PMC IDs in cases when supplementary materials are available.</p> <p>This repository is linked to the WHO database of publications on coronavirus disease and other resources, such as Microsoft Academic Graph, PubMed, and Semantic Scholar. A coalition including the <a href="https://chanzuckerberg.com/">Chan Zuckerberg Initiative</a>, Georgetown University’s <a href="https://cset.georgetown.edu/">Center for Security and Emerging Technology</a>, <a href="https://www.microsoft.com/en-us/research/">Microsoft Research</a>, and the <a href="https://www.nlm.nih.gov/">National Library of Medicine</a> of the National Institutes of Health came together to provide this service.</p> <p><strong>Citation:</strong></p> <p>When including CORD-19 data in a publication or redistribution, please cite our <a href="https://www.semanticscholar.org/paper/263db91cea260ca775cdbc482bca5392815c0533">arXiv pre-print</a>.</p> <p>The <a href="https://allenai.org/">Allen Institute for AI</a> and particularly the Semantic Scholar team will continue to provide updates to this dataset as the situation evolves and new research is released.</p>
Figure 2 from: Haak L, Greene S, Ratan K (2020) A New Research Economy: Socio-technical framework to open up lines of credit in the academic community. Research Ideas and Outcomes 6: e60477. https://doi.org/10.3897/rio.6.e60477
Figure 2 Collaborative Workflow of an Incremental Dataset in a Consortium Setting (click to view enlarged slide in Present mode). The workflow of an investigator's incremental dataset is shown as it is incorporated into the Facilitated Living Review (FLR), moving along a continuum from closed to open review. (1) After ideation and hypothesis formation by the team, early experimentation creates an incremental dataset. (2) The dataset is shared and discussed with the Project X research consortium, then iterated. (3) The v.3 dataset is shared with the consortium for further analysis, positioning, and citation in the context of the latest published evidence in the FLR. (4) The FLR with its cited, organically peer reviewed dataset is posted on a preprint server under the authorship of the FLR consortium. (5) Feedback from the broader community may lead to further iteration of the dataset at the discretion of the project team in the next round of the FLR revisions.
Data from: Who shares? Who doesn't? Factors associated with openly archiving raw research data
Many initiatives encourage investigators to share their raw datasets in hopes of increasing research efficiency and quality. Despite these investments of time and money, we do not have a firm grasp of who openly shares raw research data, who doesn't, and which initiatives are correlated with high rates of data sharing. In this analysis I use bibliometric methods to identify patterns in the frequency with which investigators openly archive their raw gene expression microarray datasets after study publication. Automated methods identified 11,603 articles published between 2000 and 2009 that describe the creation of gene expression microarray data. Associated datasets in best-practice repositories were found for 25% of these articles, increasing from less than 5% in 2001 to 30%-35% in 2007-2009. Accounting for sensitivity of the automated methods, approximately 45% of recent gene expression studies made their data publicly available. First-order factor analysis on 124 diverse bibliometric attributes of the data creation articles revealed 15 factors describing authorship, funding, institution, publication, and domain environments. In multivariate regression, authors were most likely to share data if they had prior experience sharing or reusing data, if their study was published in an open access journal or a journal with a relatively strong data sharing policy, or if the study was funded by a large number of NIH grants. Authors of studies on cancer and human subjects were least likely to make their datasets available. These results suggest research data sharing levels are still low and increasing only slowly, and data is least available in areas where it could make the biggest impact. Let's learn from those with high rates of sharing to embrace the full potential of our research output.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.