Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
6,025
datasets available to search
ShareScore release 0.9.0
Dataset results
6,025 results for “Science of science”
RNA-seq analysis of bead-purified B cells of wild-type BALB/c mice stimulated in vitro with TLR9 agonists (CpG ODN 1826, Hokkaido System Science) or left unstimulated.
GEO Series GSE269458. Mus musculus. 6 samples. Type: Expression profiling by high throughput sequencing.
RNA sequencing analysis of postmortem human Brodmann Area 9 in the University of Texas Health Science Center at Houston Brain Collection
GEO Series GSE182321. Homo sapiens. 41 samples. Type: Expression profiling by high throughput sequencing.
Small RNA sequencing of postmortem human Brodmann Area 9 and postmortem human blood from Opioid Use Disorder Subjects and Controls in the University of Texas Health Science Center at Houston Brain Col
GEO Series GSE221515. Homo sapiens. 65 samples. Type: Non-coding RNA profiling by high throughput sequencing.
Generated datasets for Yue et al. (2020, Earth and Space Science): "Combining In-situ and Satellite Observations to Understand the Vertical Structure of Tropical Anvil Cloud Microphysical Properties During the TC4 Experiment"
<p>This archive contains the data sets generated from the research conducted by Yue et al. (2020) titled "Combining In-situ and Satellite Observations to Understand the Vertical Structure of Tropical Anvil Cloud Microphysical Properties During the TC4 Experiment" published in Earth and Space Science. The method to generated the following data sets is described in Yue et al. (2020) and stored as Matlab .mat files.</p> <p>CombiningTC4_Satellite_eof_cov_mat.mat contains the correlation matrix shown in Figure 1a.</p> <p>TC4_processed.mat contains the correlation matrix shown in Figure 1b.</p> <p>RO_processed.mat contains the correlation matrix shown in Figure 2a.</p> <p>RVOD_processed.mat contains the correlation matrix shown in Figure 2b.</p> <p>ICE_processed.mat contains the correlation matrix shown in Figure 2c.</p> <p> </p> <p> </p>
Results of the Nanofront open science survey
<p>This dataset contains the results of the survey performed during the talk I gave about Open Science during the NanoFront Winter Retreat 2019 in Courchevel, France</p> <p>The talk, along with the questions and the context, can be found here:</p> <pre>https://doi.org/10.5281/zenodo.2600527</pre>
Towards Continuous Privacy Compliance: A Design Science Study in a Small Organization
<p>This data contains the interview question template we used during our interviews with our study participants. </p> <p>This page contains supplementary material for our paper submission: 'Towards Continuous Privacy Compliance: A Design Science Study in a Small Organization'.</p>
Datasets corresponding to publication: AIDeveloper: deep learning image classification in life science and beyond
<p>Datasets and videos corresponding to publication:<br> AIDeveloper: deep learning image classification in life science and beyond</p>
Figure 1 from: Dixey K, Woodburn M, Hardy H, Livermore L, Smith VS (2020) Identification of provisional Centres of Excellence for digitisation of European natural science collections. Research Ideas and Outcomes 6: e57750. https://doi.org/10.3897/rio.6.e57750
Figure 1 Heat map matrix of Center of Excellence services versus organizational levels.
Matériel supplémentaire pour Paulin & Charlat (2020, Raisons éducatives, "L'épistémologie des sciences biologiques et géologiques : une occasion d'enseigner l'incertitude ?")
<p>Ce matériel accompagne l'article intitulé "L'épistémologie des sciences biologiques et géologiques : une occasion d'enseigner l'incertitude ?", de Fabienne PAULIN et Sylvain CHARLAT, en cours de publication dans la revue <em>Raisons éducatives</em>. Le fichier déposé contient les sujets des séances d'évaluation de Travaux Pratiques considérés dans cette analyse.</p>
The Varying Openness of Digital Open Science Tools
<p>Dataset accompanying the paper submitted to F1000 entitled <strong>The Varying Openness of Digital Open Science Tools.</strong></p> <p> </p> <p>Abstract of paper</p> <p>Digital tools that support Open Science practices play a key role in the seamless accumulation, archiving and dissemination of scholarly data, outcomes and conclusions. Despite their integration into Open Science practices, the providence and design of these digital tools are rarely explicitly scrutinized. This means that influential factors, such as the funding models of the parent organizations, their geographic location, and the dependency on digital infrastructures are rarely considered. Suggestions from literature and anecdotal evidence already draw attention to the impact of these factors, and raise the question of whether the Open Science ecosystem can realise the aspiration to become a truly “unlimited digital commons” in its current structure. </p> <p> </p> <p>In an online research approach, we compiled and analysed the geolocation, terms and conditions as well as funding models of 242 digital tools increasingly being used by researchers in various disciplines. Our findings indicate that design decisions and restrictions are biased towards researchers in North American and European scholarly communities. In order to make the future Open Science ecosystem inclusive and operable for researchers in all world regions including Africa, Latin America, Asia and Oceania, those should be actively included in design decision processes. </p> <p>Digital Open Science Tools carry the promise of enabling collaboration across disciplines, world regions and language groups through responsive design. We therefore encourage long term funding mechanisms and ethnically as well as culturally inclusive approaches serving local prerequisites and conditions to tool design and construction allowing a globally connected digital research infrastructure to evolve in a regionally balanced manner.</p>
Science Education Research Topic Modeling Dataset
<p>This dataset contains scraped and processed text from roughly 100 years of articles published in the Wiley journal <em>Science Education </em>(formerly <em>General Science Quarterly</em>). This text has been cleaned and filtered in preparation for analysis using natural language processing techniques, particularly topic modeling with <a href="https://dl.acm.org/doi/10.5555/944919.944937">latent Dirichlet allocation</a> (LDA). We also include a Jupyter Notebook illustrating how one can use LDA to analyze this dataset and extract latent topics from it, as well as analyze the rise and fall of those topics over the history of the journal.</p> <p>The articles were downloaded and scraped in December of 2019. Only non-duplicate articles with a listed author (according to the <a href="https://www.crossref.org/">CrossRef metadata</a> database) were included, and due to missing data and text recognition issues we excluded all articles published prior to 1922. This resulted in 5577 articles in total being included in the dataset. The text of these articles was then cleaned in the following way:</p> <ul> <li>We removed duplicated text from each article: prior to 1969, articles in the journal were published in a magazine format in which the end of one article and the beginning of the next would share the same page, so we developed an automated detection of article beginnings and endings that was able to remove any duplicate text.</li> <li>We removed the reference sections of the articles, as well headings (in all caps) such as “ABSTRACT”.</li> <li>We reunited any partial words that were separated due to line breaks, text recognition issues, or British vs. American spellings (for example converting “per cent” to “percent”) </li> <li>We removed all numbers, symbols, special characters, and punctuation, and lowercased all words.</li> <li>We removed all <em>stop words</em>, which are words without any semantic meaning on their own—“the”, “in,” “if”, “and”, “but”, etc.—and all single-letter words.</li> <li>We lemmatized all words, with the added step of including a part-of-speech tagger so our algorithm would only aggregate and lemmatize words from the same part of speech (e.g., nouns vs. verbs).</li> <li>We detected and create <em>bi-grams</em>, sets of words that frequently co-occur and carry additional meaning together. These words were combined with an underscore: for example, “problem_solving” and “high_school”.</li> </ul> <p>After filtering, each document was then turned into a list of individual words (or tokens) which were then collected and saved (using the python pickle format) into the file scied_words_bigrams_V5.pkl.</p> <p>In addition to this file, we have also included the following files:</p> <ol> <li>SciEd_paper_names_weights.pkl: A file containing limited metadata (title, author, year published, and DOI) for each of the papers, in the same order as they appear within the main datafile. This file also includes the weights assigned by an LDA model used to analyze the data</li> <li>Science Education LDA Notebook.ipynb: A notebook file that replicates our LDA analysis, with a written explanation of all of the steps and suggestions on how to explore the results.</li> <li>Supporting files for the notebook. These include the requirements, the README, a helper script with functions for plotting that were too long to include in the notebook, and two HTML graphs that are embedded into the notebook. </li> </ol> <p>This dataset is shared under the terms of the <a href="https://olabout.wiley.com/WileyCDA/Section/id-826542.html">Wiley Text and Data Mining Agreement,</a> which allows users to share text and data mining output for non-commercial research purposes. Any questions or comments can be directed to Tor Ole Odden, t.o.odden@fys.uio.no.</p>
Survey Data Set - Would you call this Citizen Science?
<p>Data set of survey outcomes.</p>
Designing Computer Science Competency Statements: A Process and Curriculum Model for the 21st Century
<p>Companion materials - Based on CS2013 Knowedge Areas - Competency Statements for CS Curricula:</p> <p>Produced by members of the ITiCSE 2020 Working Group Report:</p> <p>Designing Computer Science Competency Statements: A Process and Curriculum Model for the 21st Century</p> <p>Clear, A., Clear, T., Vichare, A., Charles, T., Frezza, S., Gutica, M., . . . Pitt, F. (2020). Designing Computer Science Competency Statements: A Process and Curriculum Model for the 21st Century [In Press]. In Proceedings of the 2020 ACM Conference on Innovation and Technology in Computer Science Education (pp. TBA): ACM.</p>
Pesquisa de Opinião sobre Open Science na UEM
<p>Pesquisa de Opinião sobre Open Science na UEM</p>
Data from: Science-graphic art partnerships to increase research impact
Graphics are becoming increasingly important for scientists to effectively communicate their findings to broad audiences, but most researchers lack expertise in visual media. We suggest collaboration between scientists and graphic designers as a way forward and discuss the results of a pilot project to test this type of collaboration.
THE UNAMBIGUOUS 1000 ELEMENTARY TERMS, DEFINITIONS, AND DESCRIPTIONS IN ELEMENTARY SCIENCE, COURTROOMS, AND SCHOOLS
<p>At least in Elementary Science, Courtrooms, and Schools all professionally used elementary terms, definitions, and descriptions should be unambiguous. This is a train-your-brain list by going through this prescriptive list. Search functions make an alphabetical list unnecessary. More detailed background: All used words are consistent with the Oxford Dictionary in the context of elementary science unless set in bold. When inconsistent the Oxford Dictionary is amended. No scientific dictionary makes this distinction. Within the elementary context and goal, lexicographers and scientists should go about improving i.e. correcting this list. Here I show by systematically splitting the terms 'basic' and 'elementary' that it's possible in a reductio ad absurdum proof that one simple law of nature makes one-simple law of human nature based on five exclusive objective axiomatic assumptions based on five subjective wise goals for homo sapiens. This can be reduced to one goal of the bare survival of homo sapiens and the assumption of one consistent cosmos. Not accepting this grand axiom and grand goal is per logical definition anti-scientific pseudo-science. Accepting this prevents premature extinction which is inevitable without this scientific paradigm change.</p>
Figure 1 from: Filter M, Candela L, Guillier L, Nauta M, Georgiev T, Stoev P, Penev L (2019) Open Science meets Food Modelling: Introducing the Food Modelling Journal (FMJ). Food Modelling Journal 1: e46561. https://doi.org/10.3897/fmj.1.46561
Figure 1 - Workflow for conversion of FSK-ML metadata into Model manuscripts.
Figure 2 from: Tilley LJ, Woodburn M, Vincent S, Casino A, Addink W, Berger F, Bogaerts A, De Smedt S, French L, Islam S, Mergen P, Nivart A, Papp B, Petersen M, Santos C, Schiller EK, Semal P, Smith VS, Wiltschke K (2024) Systematic Design of a Natural Sciences Collections Digitisation Dashboard. Research Ideas and Outcomes 10: e118244. https://doi.org/10.3897/rio.10.e118244
Figure 2 CDD relational data model.
Figure 1 from: Tilley LJ, Woodburn M, Vincent S, Casino A, Addink W, Berger F, Bogaerts A, De Smedt S, French L, Islam S, Mergen P, Nivart A, Papp B, Petersen M, Santos C, Schiller EK, Semal P, Smith VS, Wiltschke K (2024) Systematic Design of a Natural Sciences Collections Digitisation Dashboard. Research Ideas and Outcomes 10: e118244. https://doi.org/10.3897/rio.10.e118244
Figure 1 A simplified conceptual view of the TDWG Collections Description data model.
Figure 3 from: Tilley LJ, Woodburn M, Vincent S, Casino A, Addink W, Berger F, Bogaerts A, De Smedt S, French L, Islam S, Mergen P, Nivart A, Papp B, Petersen M, Santos C, Schiller EK, Semal P, Smith VS, Wiltschke K (2024) Systematic Design of a Natural Sciences Collections Digitisation Dashboard. Research Ideas and Outcomes 10: e118244. https://doi.org/10.3897/rio.10.e118244
Figure 3 First page of the Pilot CDD showing a collection overview (Licence: CC-BY).
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.