Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

46

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

46 results for “computer science”

Learn how ShareScore rates datasets ↗
zenodo48/100

Unlocking the power of computer modelling and simulation across the life sciences product lifecycle

<p><strong>Unlocking the Power of Computer Modelling and Simulation Across the Life Sciences Product Lifecycle</strong></p> <p>In an era where technology continuously reshapes the boundaries of research and development, the field of life sciences stands at the cusp of a transformative shift. The potent combination of computer modelling and simulation has begun to unlock unprecedented opportunities across the product lifecycle in life sciences, promising to revolutionize everything from medicinal product development to clinical research. Let's delve into how these technological advancements are paving the way for groundbreaking progress in medicine and healthcare.</p> <p><strong>The Fusion of Technology and Life Sciences</strong></p> <p><em>In Silico Methods: A New Frontier in Medicine</em></p> <p>The term 'in silico' refers to computer simulations used in the study of biological and chemical processes. The video highlights the growing importance of in silico methods in the life sciences sector, particularly in the United Kingdom. These methods allow for the virtual testing of new medicinal products, significantly reducing the need for costly and time-consuming physical trials.</p> <p><em>Bridging the Gap with Computational Modeling</em></p> <p>Computational modeling is another key aspect discussed in the presentation. It involves the use of computer algorithms and mathematical models to simulate real-world medical data. This approach enables researchers to predict how medicinal products will behave in various scenarios, including their interaction with different types of patient data. As a result, computational modeling is instrumental in enhancing the precision of clinical research and improving medical equitability by considering a broader range of patient profiles.</p> <p><strong>The Impact on Clinical Research and Patient Care</strong></p> <p><em>Enhancing Precision and Efficiency</em></p> <p>One of the most notable benefits of integrating computer modelling and simulation into the life sciences is the enhanced precision and efficiency it brings to clinical research. By leveraging real-world medical data, researchers can obtain more accurate predictions about the efficacy and safety of new medicinal products. This not only accelerates the development process but also ensures that treatments are more tailored to individual patient needs.</p> <p><em>Promoting Medical Equitability</em></p> <p>The video underscores the role of these technologies in promoting medical equitability. Through the use of patient data simulations, it becomes possible to account for a wider array of genetic, environmental, and lifestyle factors that influence health outcomes. This inclusive approach ensures that the benefits of medical advancements are accessible to a diverse population, addressing disparities in healthcare access and treatment efficacy.</p> <p><strong>Conclusion: The Future is Now</strong></p> <p>The integration of computer modelling and simulation in the life sciences heralds a new era of medical research and patient care. As we continue to explore the potential of these technologies, it's clear that they hold the key to unlocking more efficient, precise, and equitable healthcare solutions. The journey towards fully realizing this potential is just beginning, but the promise it holds is immense. As we stand on the brink of this technological revolution, one thing is certain: the future of medicine and healthcare is being shaped here and now, and it's brighter than ever.</p>

opengpl-3.0-or-laterApr 2024View details →
zenodo48/100

Dataset for the publication "The TACS Model: Understanding Teachers' Adoption of Computer Science Pedagogical Content in Primary School"

<p>This dataset contains the quantitative teacher data used to analyse an in service teacher training program for Computer Science that took place from September 2019 to March 2020 in the Canton Vaud in Switzerland. Approximately 180 teachers from the the 5th and 6th grade in primary school (ages 9-11) participated in 3 days of training sessions. At the end of each training session, teachers were asked to fill in a web-based questionnaire providing information relating to their perception of the training sessions and adoption of the computer science activities. The surveys were analysed from three perspectives which are detailed in the corresponding article (the professional development program&#39;s&nbsp;perspective, the activities&#39;&nbsp;perspective, the teacher&#39;s perspective). The present repository thus contains three csv files, one per analysis. A README is included and provides additional information regarding :</p> <p>- the requirements for re-use.&nbsp;</p> <p>- the survey instrument used</p> <p>- the specific content of the 3 csv files</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Data of the article Analysis of the self-archiving policies of journals in the highest rank category of the Finnish journal classification system within computer science, physics and electronic engineering

<p>The publication forum level three journals representing the three fields of science of computer science, computer science and electrical engineering were identified by utilizing the MinEdu field search filter while searching for the top-ranked journals from the publication channel search (https://www.tsv.fi/julkaisufoorumi/haku.php?lang=en), which is based on Field of Science, Statistics Finland classification (https://www.stat.fi/meta/luokitukset/tieteenala/001-2010/index_en.html). The data were extracted during august 2017 consists of total of 127 individual journals. It is worth noting that circa 30 journals were classified into more than one fields of sciences under scrutiny. First, the journals were divided into representing gold and hybrid model journals. Second, green open access policies of the identified hybrid journals were analyzed using Laakso&rsquo;s (2014) publisher policy coding framework. Also publishers of the individual journals were identified and subsequently added to the data.</p> <p>NOTE!&nbsp;The data includes the shortest embargo to either institutional or subject repositories. For example, Elsevier had no embargo to opening accepted manuscripts from arXiv subject repository and thus no embargoes to Elsevier&#39;s journals are included within this datasheet.</p> <p>Data is in CSV. format</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2017View details →
zenodo44/100

Dataset for Paper "Towards Increased Diversity in STEM Education: Five archetypes Derived through a Data-Driven Approach Examining a Computer Science Student Cohort

<p># Dataset for Paper &quot;Towards Increased Diversity in STEM Education: Five archetypes Derived through a Data-Driven Approach Examining a Computer Science Student Cohort&quot; - Rev #1</p> <p>This is the dataset for the paper titled &quot;Towards Increased Diversity in STEM Education: Five archetypes Derived through a Data-Driven Approach Examining a Computer Science Student Cohort&quot;.</p> <p>In case of questions, feel free to contact the authors, *anonymised*, ORCID: https://orcid.org/*anonymised*, current affiliation and email: *anonymised*</p> <p>## Survey 2019 ##<br> The raw survey data for the initial 2019 survey is available in the file *survey2019_anon.csv*. Note that the data is anonymised as free-text comments have been removed. Explanations on the variables and their levels are given in the files *variables_survey2019.csv* and *values_survey2019.csv*.<br> The questionnaire for the 2019 survey is contained in *survey2019_instrument.pdf*.</p> <p>## Survey 2020 ##<br> The raw survey data for the 2020 survey is available in the file *rdata_anon_survey2020.csv*. Additional scripts are supplied to reproduce the exploratory factor analysis. The main entry is the file *EFA.R*, which imports the data. The file contains some comments on the process.<br> The questionnaire for the 2020 survey is contained in *survey2020_instrument.pdf*.</p> <p>## Interviews ##<br> The interview guide used for the five interviews is available in the file *interview_instrument.pdf*.</p>

opencc-by-4.0May 2021View details →
zenodo44/100

Multiscale continuum figures from Tratnyek et al. (2017) "In silico environmental chemical science: Properties and processes from statistical and computational modelling"

<p>Accessible versions of selected figures from&nbsp;Tratnyek et al. (2017) &quot;In silico environmental chemical science: Properties and processes from statistical and computational modelling&quot; Environ. Sci. Processes Impacts 19(3): 188-202. DOI: 10.1039/C7EM00053G.</p> <p>The Abstract Art figure shows&nbsp;a classification of variables for predictive/diagnostic models used in silico environmental chemical science, in terms of system scales and variable types. Figure 3 shows&nbsp;a continuum of system scales encompassing the whole scope of predictive/diagnostic modelling for in silico environmental chemical sciences, juxtaposing earth and biological scales.</p> <p>The published version of Figure 3 is tall, for two-column page-layouts, but a wide version of Figure 3 is provided for landscape oriented formats. The 300 dpi versions of each figure should be adequate resolution for most purposes, and therefore are recommended.&nbsp;The large versions of the figures may take significant time to download, but may be useful for high resolution applications.</p> <p>This work is from the perspectives/review paper at the beginning of a themed issue on &quot;Quantitative Structure-Activity Relationships (QSARs) and Computational Chemistry Methods in the Environmental Chemical Sciences&quot;, published in the March 2017 issue of the Royal Society of Chemistry journal Environmental Sciences: Process and Impacts. The whole collection of papers can be accessed at rsc.li/qsars.</p>

opencc-by-4.0Aug 2017View details →
zenodo44/100

BIP! NDR (NoDoiRefs): a dataset of citations from papers without DOIs in computer science conferences and workshops

<h2>Overview</h2> <p>In the field of Computer Science, conference and workshop papers serve as important contributions, carrying substantial weight in research assessment processes, compared to other disciplines. However, a considerable number of these papers are not assigned a Digital Object Identifier (DOI), hence their citations are not reported in widely used citation datasets like OpenCitations and Crossref, raising limitations to citation analysis. While the Microsoft Academic Graph (MAG) previously addressed this issue by providing substantial coverage, its discontinuation&nbsp; has created a void in available data.</p> <p>BIP! NDR aims to alleviate this issue and enhance the research assessment processes within the field of Computer Science. To accomplish this, it leverages a workflow that identifies and retrieves Open Science papers lacking DOIs from the DBLP Corpus, and by performing text analysis, it extracts citation information directly from their full text.</p> <p>The current version of the dataset contains&nbsp;<em>~4.3M citations</em> made by approximately <em>211K open access Computer Science conference or workshop papers</em> that, according to DBLP, do not have a DOI. The DBLP snapshot used for this version was the one released on <em>September 2025</em>.&nbsp;</p> <h2>Dataset files</h2> <h3>1. Core Non-DOI Citation Dataset - bip_ndr_{version}.tar.gz</h3> <p>The dataset is formatted as a JSON Lines (JSONL) file (one JSON Object per line) to facilitate file splitting and streaming.&nbsp;</p> <p>Each JSON object has three main fields:</p> <ul> <li> <p>&ldquo;_id&rdquo;: a unique identifier,</p> </li> <li> <p>&ldquo;citing_paper&rdquo;, the &ldquo;dblp_id&rdquo; of the citing paper,</p> </li> <li> <p>&ldquo;cited_papers&rdquo;: array containing the objects that correspond to each reference found in the text of the &ldquo;citing_paper&rdquo;; each object may contain the following fields:</p> <ul> <li> <p>&ldquo;dblp_id&rdquo;: the &ldquo;dblp_id&rdquo; of the cited paper. Optional - this field is required if a &ldquo;doi&rdquo; is not present.</p> </li> <li> <p>&ldquo;doi&rdquo;: the doi of the cited paper. Optional - this field is required if a &ldquo;dblp_id&rdquo; is not present.</p> </li> <li> <p>&ldquo;bibliographic_reference&rdquo;: the raw citation string as it appears in the citing paper.</p> </li> </ul> </li> </ul> <p>Changes from previous version:</p> <ul> <li>Added more papers from DBLP.</li> </ul> <h3>2. Citation Intents Dataset - bip_ndr_ci_{version}.tar.gz</h3> <p>This file enriches the BIP! NDR dataset with citation-level intent classification.<br>It preserves the same base structure of the previous file, while adding a nested array of "citations" with each element of "cited_papers".</p> <p>Each "citation" provides the local textual context, section, and intent of the citation in the following format:</p> <ul> <li>"citation_id": Unique identifier in the format {citing_id}&gt;{cited_id}_CIT{index} linking the citing and cited entities.</li> <li>"section": The section of the citing paper where the citation occurs (e.g., Introduction, Methods, Results).</li> <li>"intent": Inferred purpose of the citation based on textual context (see classification schema below).</li> </ul> <p>The "intent" field follows the SciCite classification schema, which categorizes citations into three high-level functional types:</p> <ol> <li>background information: The citation states, mentions, or points to the background information giving more context about a problem, concept, approach, topic, or importance of the problem in the field.</li> <li>method: Making use of a method, tool, approach or dataset.</li> <li>results comparison: Comparison of the paper's results/findings with the results/findings of other work.</li> </ol> <p>The classification is done with the <a href="https://huggingface.co/sknow-lab/Qwen2.5-14B-CIC-SciCite">Qwen2.5-14B-CIC-SciCite fine-tuned Large Language Model, published by Athena RC</a>.&nbsp;</p> <p>Changes from previous version:&nbsp;</p> <ul> <li>Added more papers with intent</li> </ul>

opencc-zeroMay 2023View details →
zenodo44/100

arXiv abstracts and titles from 1,469 single-authored papers (100 unique authors) in computer science

<p>This dataset is meant to be used for experiments of Authorship Analysis. The dataset&nbsp;consists of abstracts of single-author papers from arXiv crawled using the arXiv&#39;s API by querying a list of computer-science-related keywords (&quot;deep learning&quot;, &quot;machine learning&quot;, &quot;information retrieval&quot;, &quot;computer science&quot;, &quot;data mining&quot;, &quot;support vector&quot;, &quot;logistic regression&quot;, &quot;artificial intelligence&quot;, &quot;supervised learning&quot;&#39;).&nbsp;The corpus somehow follows a power-law distribution, with few prolific authors and many authors accounting for very few papers each: we retained authors with at&nbsp;least 10 papers, resulting in a total of 1,469 documents from 100 authors. The most prolific authors (Peter D. Turney and Subhash Kak) have 34 abstracts to their names, the 10 most prolific authors have written 22 or more articles, while 50% of the authors have no more than 12 abstracts to their names. In order to divide the corpus into a training set and a test set we perform a stratified split, with the production of each author being split into a training set (70%) and a test set (30%). We use these documents as examples of &quot;scientific communication&quot;, characterised by a precise and compact style, with an abundance of technical terminology.</p>

opencc-by-4.0Dec 2022View details →
zenodo40/100

BRAIN Journal-Personality Questionnaires as a Basis for Improvement of University Courses in Applied Computer Science and Informatics-Figure 5. Hierarchical clustering by scores across the EPQ–R scales for data about all the participants

<p>The clusters were generated using an implementation of a hierarchical clustering algorithm available in the R environment (R, n.d.). The top three clusters were extracted from a hierarchical cluster tree shown in Figure 5, while the color of data points in the visualization shown in figure 4 was determined based on cluster labels. Hierarchical clusters could be used when investigating which students in the analyzed sample share similar personality traits. This could be especially useful for smaller student groups as the teacher may manually inspect the cluster tree and its leaves, which designate individual students. For instance, there are three students in cluster 3, who are represented within the tree in Figure 5 by identifiers 14, 22, and 24. The students with identifiers 14 and 22 are more closely linked and more similar to each other than to the student with identifier 24.&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

BRAIN Journal-Personality Questionnaires as a Basis for Improvement of University Courses in Applied Computer Science and Informatics-Figure 2. Comparison of mean scores on the EPQ–R scales

<p>The data for the workshop participants were loaded from the data warehouse, while the summary data from the original EPQ&ndash;R study were loaded from a CSV file. The bar chart featured in Figure 2 shows mean scores on the EPQ&ndash;R scales for the selected workshop participants (denoted by blue bars) and the selected participants of the original EPQ&ndash;R study (denoted by yellow bars). The mean scores on the P scale agree between the two samples, but the overall scores for the other scales vary.&nbsp;</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

BRAIN Journal-Personality Questionnaires as a Basis for Improvement of University Courses in Applied Computer Science and Informatics-Figure 4. Radial visualization of scores across the EPQ–R scales for clustered data about all the participants

<p>On the other hand, the division of data points by gender might not be the only useful strategy when visually inspecting the analyzed sample in a coordinate system. Numerous clustering algorithms may be used to determine which data points share similar scores across the EPQ&ndash;R scales, i.e., which data points belong to the same cluster of similar entities based on their corresponding EPQ&ndash;R scores. A radial visualization in which data points were organized into three clusters is given in Figure 4. Each cluster is marked by a different color: cluster 1 by red, cluster 2 by green, and cluster 3 by blue.&nbsp;</p>

opencc-by-4.0Jul 2017View details →
zenodo40/100

BRAIN Journal-Personality Questionnaires as a Basis for Improvement of University Courses in Applied Computer Science and Informatics-Figure 3. Radial visualization of scores across the EPQ–R scales for the male and female participants

<p>The radial visualization in Figure 3 depicts each participating student as a dot whose color indicates the gender of the student, blue for male students (M) and red for female students (F). The position of a dot in the visualization is determined by the scores of the associated student on the four EPQ&ndash;R scales. The radial overview may provide a much clearer outline of clustering within the analyzed group. Although there are only five female students, they are concentrated in a relatively narrow area within the radial coordinate system</p>

opencc-by-4.0Mar 2018View details →
zenodo40/100

All Computer Science Papers @ arXiv.org -- A High-Quality Gold Standard for Citation-based Tasks

<p>We propose a newly-created gold standard <strong>data set for citation-based tasks</strong>. This gold standard is based on <strong>all computer science papers in arXiv.org</strong>.</p> <p><strong>Abstract</strong>. Analyzing and recommending citations with their specific citation contexts have recently received much attention due to the growing number of available publications. Although data sets such as CiteSeerX have been created for evaluating approaches for such tasks, those data sets exhibit striking defects. This is understandable if one considers that both information extraction and entity linking as well as entity resolution need to be performed. In this paper, we propose a new evaluation data set for citation-dependent tasks based on arXiv.org publications. Our data set is characterized by the fact that it exhibits almost zero noise in the extracted content and that all citations are linked to their correct publications. Besides the pure content, available on a sentence-basis, cited publications are annotated directly in the text via global identifiers. As far as possible, referenced publications are further linked to DBLP. Our data set consists of over 15M sentences and is freely available for research purposes. It can be used for training and testing citation-based tasks, such as recommending citations, determining the functions or importance of citations, and summarizing documents based on their citations.</p> <p>&nbsp;</p> <p>More information can be found in our <strong>publication &quot;<a href="http://www.lrec-conf.org/proceedings/lrec2018/pdf/283.pdf">A High-Quality Gold Standard for Citation-based Tasks</a>&quot; (LREC&#39;18)</strong>.</p> <p>You can cite the data set as follows:</p> <pre><code>@inproceedings{DBLP:conf/lrec/0001TJ18, author = {Michael F{\"{a}}rber and Alexander Thiemann and Adam Jatowt}, title = "{A High-Quality Gold Standard for Citation-based Tasks}", booktitle = "{Proceedings of the Eleventh International Conference on Language Resources and Evaluation}", series = "{LREC'18}", location = "{Miyazaki, Japan}", year = {2018}, url = {http://www.lrec-conf.org/proceedings/lrec2018/summaries/283.html} } </code></pre> <p>&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

Conceptual Change Texts in Computer Science to Expand Students' Conceptions on the Topic of Artificial Intelligence

<p><strong><span>General information on the survey and cohort</span></strong></p> <p><span>The present data were collected in 2023 at a secondary school in Hamburg. A questionnaire was used as a pre- and post-test in computer science courses in years 10 and 11 to investigate the agreement with various items and how this changed. The questionnaires were transferred to SPSS for analysis.</span></p> <p><span>76 students aged 15-17 took part in the survey and the teaching intervention. This data set only contains the data that is available in full. It therefore includes the responses of 69 students, of which a total of 22 felt assigned to the female gender and 47 to the male gender.</span></p> <p>&nbsp;</p> <h3><span>Data set</span></h3> <p><span>The data is purely quantitative, as this is only the evaluation of the pre- and post-test and not the evaluation of the conceptual change texts themselves. The item list is made up of 20 items relating to artificial intelligence, which were drawn up according to the Big Ideas of the AI4K12 initiative [1]. The students' perceptions of the items come from various past studies that have surveyed students' conceptions of AI.</span></p> <p><span>The data has a pseudonymized code that can be linked to the conceptual change text of the experimental group and the data can be assigned accordingly. The data is also divided into the experimental group (abbreviation 2) and the control group (abbreviation 1) so that a comparison between the groups is quickly possible. The abbreviation W in front of the items stands for a &ldquo;true&rdquo; statement and SV for &ldquo;student conception&rdquo;. A Likert scale from 0 - does not apply at all to 3 - applies completely was used.</span></p> <p><span>The data set is cleansed data. All data sets that were not complete or that were given different codes in the pre-test and post-test and therefore could no longer be clearly assigned were removed.</span></p> <p><strong><span>Literature</span></strong></p> <p><span>[1]</span><span>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; </span>AI<span> </span>for<span> </span>K12<span> </span>(Hrsg.)<span> </span><span>(2020):<span> </span><em>AI4K12.org.<span> </span></em>Available online at:<span> </span>https://ai4k12.org/<span> </span>[last checked 12.03.2023]</span></p>

opencc-by-4.0Oct 2024View details →
zenodo40/100

Data sets - The attitude of computer science teachers to inclusive education, Motivation to teach, Perception of the possible impact of computer science on students with mental disabilities

<p>Data sets&nbsp;</p> <p>The attitude of computer science teachers to inclusive education, Motivation to teach, Perception of the possible impact of computer science on students with mental disabilities.&nbsp;<br>In the period from February to October 2024, a survey of 112 computer science teachers in Kazakhstan (Pavlodar region) was conducted to determine attitudes to inclusive education, motivation to teach, and perception of the possible impact of computer science on students with mental disabilities.</p> <p>Questionnaire&nbsp;<br>https://docs.google.com/document/d/1LzukKSqW_mHMZXbMtN0ecmmU4cKJiwgf0laTWBHQSng/edit?usp=sharing</p> <p><strong>This research has been funded by the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grant No. AP14872400).</strong></p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Mirror of data from NOAA U.S. Climate Reference Network for Research Computing in Earth Science

<p>This is a mirror of data from the NOAA U.S. Climate Reference Network (https://www.ncei.noaa.gov/products/land-based-station/us-climate-reference-network).</p> <p>It was created because outbound FTP access is not allowed from some cloud-based JupyterHub setups.</p>

opencc-by-4.0Oct 2021View details →
zenodo40/100

ICITS'23 - Understanding the Success Factors of Research Software: Interviews with Brazilian Computer Science Academic Researchers

<p>Artifacts used for data collection and analysis of the article accepted for publication in ICITS&#39;23.</p> <p>Mour&atilde;o, E., Trevisan, D., Viterbo, J. (2022).Understanding the Success Factors of Research Software: Interviews with Brazilian Computer Science Academic Researchers. In:&nbsp;ICITS&#39;23 - 6th International Conference on Information Technology &amp; Systems. Advances in Intelligent Systems and Computing,&nbsp;Springer, Cham.</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Dataset for the evaluation of student-level outcomes of a primary school Computer Science curricular reform

<p>Dataset for the evaluation of student-level outcomes of a primary school Computer Science curricular reform<br> =======================================================</p> <p>&bull; If you publish material based on this dataset, please cite the following :</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&bull; The Zenodo repository : Laila El-Hamamsy, Barbara Bruno,&nbsp;Jessica Dehler Zufferey, and Francesco Mondada (2023). Dataset for the evaluation of student-level outcomes of a primary school Computer Science curricular reform&nbsp;[Data set]. Zenodo. https://doi.org/10.5281/zenodo.7489244</p> <p>&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;&bull; The associated peer reviewed article that will appear in the International Journal of STEM education : El-Hamamsy, L., Bruno, B., Audrin, C., Chevalier, M., Avry S., Dehler Zufferey, J., and Mondada, F. (2023). How are Primary School Computer Science Curricular Reforms Contributing to Equity? Impact on Student Learning, Perception of the Discipline, and Gender Gaps. arXiv, to appear in the International Journal of STEM Education. https://doi.org/10.48550/arXiv.2306.00820</p> <p>&bull; License: This work is licensed under a Creative Commons Attribution 4.0 International license (CC-BY-4.0)</p> <p>&bull; Creator: El-Hamamsy, L., Bruno, B., Dehler Zufferey, J., and Mondada, F.</p> <p>&bull; Date: May 2nd 2023</p> <p>&bull; Subject: Computer Science; Curricular Reform; Elementary Education; Learning Achievement; Computational Thinking; Perception Survey; Equity; Gender Gaps</p> <p>&bull; Dataset format: CSV</p> <p>&bull; Dataset collection: January 2021 to May 2022</p> <p>&bull; Dataset size : &lt; 100 kB</p> <p>&bull; Dataset content : three excel files. We provide the detailed description of each of the files below. The original questions are available in the associated publication [1]. Please note that these datasets contain missing values due to students either not being present for all data collections or not having the associated teacher-related data.</p> <p>&bull; Abbreviations :<br> &nbsp; - CS : Computer Science<br> &nbsp; - CT : Computational Thinking<br> &nbsp; - PD : Professional Development</p> <p>&bull; Funding : This work was funded by the the NCCR Robotics, a National Centre of Competence in Research, funded by the Swiss National Science Foundation (grant number 51NF40_185543)</p> <p>&nbsp;</p> <p># References</p> <p>[1] El-Hamamsy, L., Bruno, B., Audrin, C., Chevalier, M., Avry S., Dehler Zufferey, J., and Mondada, F. (2023). How are Primary School Computer Science Curricular Reforms Contributing to Equity? Impact on Student Learning, Perception of the Discipline, and Gender Gaps. arXiv, to appear in the International Journal of STEM Education. https://doi.org/10.48550/arXiv.2306.00820</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Knowledge Freedom in computational science: a two stage peer-review process with KF Eligibility Access Review

<p>What is explained below is the&nbsp;<em>peer review process</em>&nbsp;applied by&nbsp;<em>Notes on Transdisciplinar</em><em>y Modelling for Environment&nbsp;</em>(NTMe). Authors and readers are encouraged to understand not only the process, but also the rationale behind the process, which&nbsp;is closely connected with the peculiar challenges faced by the scientific problems (computational-science modelling under uncertainty, in broad and heterogeneous contexts) in the scope of the&nbsp;journal.&nbsp;</p> <p>&nbsp;</p> <p><span class="math-tex">\(\mathsf{\Large\text{The rationale behind the peer-review process}}\)</span></p> <p><strong>Computational research problems for more than one discipline. </strong>Research specific to a particular disciplinary domain is typically authored, and then studied, by experts in that domain. Several conventions, assumptions, fundamental protocols and methodology practices, are shared by domain experts as a common ground. Given that proficiency in these aspects is a precondition for domain experts, this kind of knowledge is often implicit, and not well communicated within domain-specific literature. &quot;Typical&quot; data preprocessing, model settings, quality-assessment assumptions (and their &quot;default&quot; simplifications and shortcuts), in this literature are often minimally reported or omitted, because the domain-specific community of researchers and practitioners knows them so well that they may resemble a sort of acquired conditioned response. However, when in a wider research more than one disciplinary domain is connected among each other, then implicit unexpressed knowledge may lead to ambiguity, misuse of results, and avoidable errors - sometimes challengly difficult to detect. The more and more diverse the domains, the higher this risk might be.</p> <p><strong>Useful cross-domain building blocks. </strong>&quot;Transdisciplinary&quot; research may risk being misinterpreted as just the final summary list of results grabbed by autonomous disciplinary silos run in parallel. However, a poor ability to&nbsp;communicate computational methods and data - and their limitations - between multiple disciplinary domains&nbsp;would&nbsp;jeopardise the collective (cross-domain) ability to distill truly integrated computational science. Computational science for transdisciplinary problems requires a pragmatic, engineered, but reasonably simple and modular approach to break and connect disciplinary silos - when appropriate for the problems investigated. Publishing <em>useful</em> cross-domain building blocks of this approach is the aim of the peer-review process in NTMe.</p> <p><strong>Connecting disciplines, space and time scales under uncertainty: the need for a shared semantics.&nbsp;</strong>Wide-scale transdisciplinary modelling (WSTM) relies on methods proper to computational science for connecting multiple scientific disciplines, and shedding light on broad, complex problems. Several of the most pressing problems we face as a human society are only partially known in their chain of consequences, and hopelessly difficult to address by a single discipline and a narrow perspective. Some of these problems capture the specific interest by particularly exposed regions, but their mechanism is more general, with essential knowledge often lying in other spatial areas -&nbsp;or even other time periods - hence calling for a wider perspective. Multiple&nbsp;spatial and temporal scales, with dissimilar data resolution, are frequent in WSTM problems, since they are often dictated by the complexity of reality and the available data. A wide-scale extent typically implies uneven quality and availability of data, and many sources of uncertainty - which increases when assumptions, methods and semantics by different disciplines or research institutions do not easily merge. A truly transdisciplinary communication of this essential semantics requires the <em>freedom </em>to communicate scientific knowledge, accurately and transparently, between disciplinary and corporate barriers.</p> <p><strong>Real knowledge sharing in computational science.&nbsp;</strong>As a consequence, wide-scale transdisciplinary modelling demands a focus on reproducible research and real scientific knowledge freedom. Data and software freedom are essential aspects of knowledge freedom in computational science. Therefore, ideally published articles should also provide the readers with the data and source code of the described mathematical modelling. To maximise transparency, replicability,&nbsp;reproducibility and reusability, published data should be made available as open data while source code should be made available as free software. Communicating semantics even to non-experts in a given domain requires that mathematical assumptions, otherwise obvious within a specific discipline, are duly annotated in a portable and concise way. Accordingly, brief but semantically clear documentation of data, methods and software is a key precondition. Here, a two-stage peer review process is described in which scientific knowledge freedom is considered with a dedicated Eligibility Access Review. This new peer review process is applied by <em>Notes on Transdisciplinary Modelling for Environment</em>&nbsp;with a focus on WSTM for environment.</p> <p>&nbsp;</p> <p><span class="math-tex">\(\mathsf{\Large\text{The peer-review process}}\)</span></p> <p><strong>A&nbsp;two-stage peer review process to avoid single-use disposable&nbsp;computational science.&nbsp;</strong>The two-stage peer review process requires discussion papers&nbsp;to be published so as to receive feedback from the scientific community before their possible finalisation. Initial manuscript submission is subject to the soundness review outlined above, also ensuring eligibility criteria to be fulfilled so as to support&nbsp;scientific knowledge freedom. Although this concept is multifaceted, some few dimensions might be emphasised which broadly apply in computational science and engineering (CSE). Among the many possible eligibility criteria in CSE, it should be highlighted at least the need for:&nbsp;</p> <ul> <li>free software to have been&nbsp;published so as for it to be persistently available;</li> <li>appropriate licensing and source code review to have been&nbsp;done (portable modularisation, with semantics of mathematical data-transformation methods clearly annotated);</li> <li>free data&nbsp;to have been&nbsp;published so as for it to be persistently available (with semantics of quantities clearly annotated);</li> <li>a minimal share of free-access core references to be selected&nbsp;in order for scientists and research organisations not to be discriminated on the basis of their funding availability, when they try to access the core literature cited in the manuscript.</li> </ul> <p>Before acceptance for discussion, a manuscript (and its potential data, parameters and software)&nbsp;may be revised, without exposing to the public immature versions, until it becomes able to fulfil the eligibility criteria. These iterations between author(s) and editors/reviewers remain confidential, and a potential manuscript rejection does not preclude resubmission. Acceptance of the manuscript to the discussion stage is followed by the permanent publication of the accepted version. This does not preclude, and actually encourages, further revised versions to be resubmitted. Any new revision is subject to eligibility access review, which takes into account the submission history of the manuscript and the already accepted public older versions.</p> <p><strong>Not a static publication. </strong>The discussion stage of the peer review process allows short comments to be submitted by&nbsp;referees and the scientific community, while authors are encouraged to interact with pending&nbsp;comments by providing their responses. During this stage, the paper accepted for discussion is already citable. Depending on the specific goals of each research problem, the discussion may remain open for a short time interval, or instead for a noticeably longer period of time (even indefinitely, if appropriate: for example, whenever authors do consider publishing cumulative milestone versions as important). There is not an a-priori limit to the duration of this first stage, and to the number of intermediate public revisions (which are optional), for a paper accepted for discussion.&nbsp;The second stage of peer review concludes the discussion stage with the&nbsp;submission of a revised manuscript, with final review&nbsp;and corrections. Fulfilment of the eligibility criteria is required over all the&nbsp;publication stages.<br> The published materials might be updated from time to time (at the authors&#39; discretion), and - after peer-reviewing the changes - published so that the evolution is clear and available to others. Although preferred and encouraged in the discussion stage, updates are possible even after the final stage. Revisions offering a noticeably different content, compared with the previous versions, might sometime be recommended by editors/reviewers as deserving the status of a new publication - not to confuse the readers. This implies the full two-stage review process has to be applied to the new publication (as an independent publication), while the previous accepted publication is considered as the final permanent version. Recalling an analogy with the typical evolution of free-software, a software package is frequently subject to future improvements and corresponding new versions. If a new version is too different from the original package, it may become a new independent package.&nbsp;</p> <p><strong>Scientific opinions,&nbsp;perspectives, and overviews.</strong> Manuscripts focusing on computational science methods, data, parameters, and software are expected to contribute practical components of scientific knowledge which can be adapted, modified, or generally integrated within the future body of knowledge - potentially in unexpected ways. Therefore, the aforementioned two-stage review is, overall, meant to ease the future invention of derivative works. On the other hand, manuscripts offering expert opinions on these topics do not contribute &quot;practical&quot; components to be directly modified. Instead, they contribute organised ideas to support the evolution of scientific knowledge. Therefore, the peer review process for this typology of manuscripts focuses on:</p> <ul> <li>the broad understandability and potential interest of the expert opinions, and</li> <li>the&nbsp;factual&nbsp;correctness of the objective elements included in the opinion (with adequate bibliographic, tabular, visual support where appropriate).&nbsp;</li> </ul> <p><strong>Preregistration, protocols and methods. </strong>NTMe accepts a third typology of manuscripts, focusing on anticipated results (hence not yet computed) which are expected following a modelling procedure to transform input data and parameters into the desired output. As data and software are only anticipated in this typology of manuscripts, the peer-review process focuses on:</p> <ul> <li>the broad understandability and potential interest of the proposed protocol or methodology,</li> <li>the correctness and clarity of the mathematical formulation of the proposed computational-modelling steps (data-transformation modules),</li> <li>the discussion of anticipated results, pitfalls, sources of uncertainty and their proper management, bifurcations of the methodology depending on the potential variants that a user may desire to apply (depending on available data and their quality, and on the desired specialisation of the proposed general method), and</li> <li>the adequacy of bibliographic, tabular, visual supports where appropriate.&nbsp;</li> </ul> <p><strong>Abstract, and plain language summary. </strong>Each article includes a technical abstract. Although the topics of published content might deal with domain-specific aspects of computational/environmental science, their focused analysis is expected to support wider research in a transdisciplinary context. Therefore, the abstract is expected to be accessible beyond the domain-specific research community. Each article, either in the discussion or in the final stage, in addition to a technical abstract needs a <em>plain language summary</em> so that non-experts and readers with a basic scientific and technical literacy can understand the main contributions of the article. The aim of the plain language summary is to support educational dissemination. The summary is not a simpler abstract, and typically offers a longer, more comprehensive description. It is realised with an editorial support to authors, so as to respect a general structure (unless otherwise agreed):</p> <ul> <li>first, the general context of the work is presented (the general topics in which the work is situated);</li> <li>second, the general open questions are introduced to which the work aims to contribute;</li> <li>third, the specific way is described in which the work contributes to progress on the general open questions;</li> <li>fourth, the implications, new opportunities or suggested lines of research, and potential limitations and unaswered questions are summarised.</li> </ul> <p><strong>Constructive peer review.</strong> The ultimate aim of the peer review process is to support authors, and future readers, in sharing useful, durable, and understandable notes of transdisciplinary knowledge. Given the very nature of the topics in NTMe, the need for a cogent peer review on a multiplicity of expertise fields is essential. Rarely a single reviewer is able to cover all the aspects of transdiciplinary computational science. Even so, different opinions by different reviewers may emerge on specific points. Therefore, a <em>constructive </em>approach is promoted during each step of the peer review process, so that authors are guided to understand the peer review, and a compendium of recommandations may be presented in a consistent way.</p> <p>&nbsp;</p> <p><sup>&copy; 2012-2020 Daniele de Rigo. This work is licensed under a </sup><a href="http://creativecommons.org/licenses/by-nd/4.0/legalcode"><sup>Creative Commons Attribution No Derivatives 4.0 International</sup></a><sup> license. Please, cite as<br> de Rigo, D., 2012. </sup><strong><sup>Knowledge Freedom in computational science: a two stage peer-review process with KF eligibility access review</sup></strong><sup>. </sup><em><sup>Zenodo, </sup></em><sup>CERN. </sup><a href="https://doi.org/10.5281/zenodo.7578"><sup>https://doi.org/10.5281/zenodo.7578</sup></a><sup> </sup></p> <p>&nbsp;</p>

opencc-by-nd-4.0Dec 2012View details →
zenodo36/100

Glosario: A multilingual glossary for computing and data science terms.

<p><code>glosario</code> is an open-source glossary of terms used in data science that is available online and also as a library in both&nbsp;<a href="https://github.com/carpentries/glosario-r/">R</a>&nbsp;and&nbsp;<a href="https://github.com/carpentries/glosario-py/">Python</a>. By adding glossary keys to a lesson&rsquo;s metadata, authors can indicate what the lesson teaches, what learners ought to know before they start, and where they can go to find that knowledge. Authors can also use the library&rsquo;s functions to insert consistent hyperlinks for terms and definitions in their lessons in any of several languages. The master copy of the glossary lives in the <code>glossary.yml</code> file.&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Dataset for the publication titlted "A computational mechanics model for producing molecular assembly using molecularly woven pantographs" in the journal Cell Reports Physical Science, authored by Byeonghwa Goh and Joonmyung Choi.

<p>Dataset for the publication titlted "A computational mechanics model for producing molecular assembly using molecularly woven pantographs" in the journal Cell Reports Physical Science, authored by Byeonghwa Goh and Joonmyung Choi.</p>

opencc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record