Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
42
datasets available to search
ShareScore release 0.9.0
Dataset results
42 results for “Jupyter Notebook”
Outputs of the Jupyter Notebook - Deep learning and variational inversion to quantify and attribute climate change (CIRC23)
<p>The dataset contains the outputs of the notebook "Deep learning and variational inversion to quantify and attribute climate change (CIRC23)" published in The Environmental Data Science Book.</p>
Outputs of the Jupyter Notebook - Variational data assimilation with deep prior (CIRC23)
<p>The repository contains the outputs of the notebook "Variational data assimilation with deep prior (CIRC23)" published in The Environmental Data Science Book.</p>
Jupyter Notebook and comprising data for GRL2023GL106264R: Understanding the Cascade: Removing GCM biases improves dynamically downscaled climate projections
<p>This notebook and attendant files allows users to interface with a small subset of the data used to create the data in GRL2023GL106264R. Also feel free to check out the overall description of the non-bias corrected dynamically downscaled GCMs in WUS-D3 here: https://zenodo.org/records/10635867. This DOI also contains version of WRF 4.1.3 allowing for yearly CH4, CO2, and N2O updates, as well as a 360-day calendar version.</p>
Datasets for the paper "ReSplit: Improving the Structure of Jupyter Notebooks by Re-Splitting Their Cells"
<p>In this archive, you can find all the data used in the paper "ReSplit: Improving the Structure of Jupyter Notebooks by Re-Splitting Their Cells".</p> <p><strong>sklearn_full_cells.csv</strong> is the dataset from the paper of Pimentel et al. filtered with only Data Science notebooks.<br> <strong>complete.csv</strong> is the dataset obtained after the full run of ReSplit on the dataset: both merging and splitting.<br> <strong>split.csv</strong> is the dataset obtained after running only the splitting part of our dataset.<br> <strong>merged.csv</strong> is the dataset obtained after running only the merging part of our dataset.<br> <strong>duplicates_id.csv</strong> contains the IDs of the duplicate notebooks for deduplication.<br> <strong>changes.csv</strong> contains the IDs of the datasets, as well as their length before and after running ReSplit.<br> <strong>survey.csv</strong> is the table with the results of the survey.</p> <p>In the dataset CSVs, each line is a cell that has a unique identifier and an identifier of the corresonding notebook.</p>
Jupyter Notebooks for "Evaluating CephFS Performance vs. Cost on High-Density Commodity Disk Servers" 10.1007/s41781-021-00071-1
<p>Jupyter notebooks used to create plots in article DOI 10.1007/s41781-021-00071-1</p> <p>Title "Evaluating CephFS Performance vs. Cost on High-Density Commodity Disk Servers"</p> <p>Journal "Computing and Software for Big Science"</p>
Biotechnology data analysis training with Jupyter Notebooks
<p>Biotechnology has experienced innovations in analytics and data processing. As the volume of data and its complexity grows, new computational procedures for extracting information are developed. However, the rate of change outpaces the adaptation of biotechnology curricula, necessitating new teaching methodologies to equip biotechnologists with data analysis abilities. To simulate experimental data, we created a virtual organism simulator (<em>silvio</em>) by combining diverse cellular and sub-cellular microbial models. With the <em>silvio </em>Python package, we constructed a computer-based instructional workflow to teach growth curve data analysis, promoter sequence design, and expression rate measurement. The instructional workflow is a Jupyter Notebook with background explanations and Python-based experiment simulations combined. The data analysis is either conducted within the Notebook in Python or externally with Excel. This instructional workflow was separately implemented in two distance courses for Master's students in biology and biotechnology with assessment of the pedagogic efficiency. The concept of using virtual organism simulations that generate coherent results across different experiments can be used to construct consistent and motivating case studies for biotechnological data literacy.</p> <p>Here, the supplementary material is provided.</p> <table> <tbody> <tr> <td>2207_BLS-RecExpSim.mbz</td> <td>Moodle backup file for import as new moodle function.</td> </tr> <tr> <td>BLS_RecExpSim_PerformanceEvaluation Rubric.docx</td> <td>Expected learning outcomes with associated performance levels.</td> </tr> <tr> <td>BLS_SurveryQuestions.docx</td> <td>Survey questions to evaluate the educational approach.</td> </tr> <tr> <td>RecExpSim.html</td> <td>Html-Export of the Jupyter Notebook to teach biotechnology data analysis. This only serves as visual impression of the course because the dynamic Python-evaluations are not functioning.</td> </tr> <tr> <td>RecExpSim_Lecture.pdf</td> <td>Static pdf of preparatory lecture to cover the theoretical aspects in the simulations and to get student on comparable level.</td> </tr> <tr> <td>RecExpSim_Lecture.pptx</td> <td>Adjustable pptx of preparatory lecture to cover the theoretical aspects in the simulations and to get student on comparable level.</td> </tr> </tbody> </table> <p> </p> <p> </p>
Outputs of the Jupyter Notebook - Sea ice forecasting using the IceNet Library
<p>The dataset contains the outputs of the notebook "Sea ice forecasting using the IceNet library" published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li>James Byrne (author), British Antarctic Survey, <a href="https://github.com/JimCircadian">@JimCircadian</a></li> <li>Bryn Noel Ubald (author), British Antarctic Survey, <a href="https://github.com/tom-andersson">@tom-andersson</a></li> <li>Wei Ji (reviewer), Development Seed, <a href="https://github.com/weiji14">@weiji14</a></li> <li>William Gregory (reviewer), Princeton University, <a href="https://github.com/William-gregory">@William-gregory</a></li> <li>Anne Fouilloux (editor), Simula Research Laboratory, <a href="https://github.com/annefou">@annefou</a></li> </ul> <p><em>Modelling codebase</em></p> <ul> <li>James Byrne (Code author)</li> <li>Tom Andersson (Science author)</li> <li>Bryn Noel Ubald (Code maintainer and contributor)</li> </ul>
Vibrational-EELS Dataset and data processing routine (Jupyter Notebook) (Laforet et al.)
<p>Contains all the vibrational-EELS data presented in the article, accompanied with the python HyperSpy processing routine used (Jupyter Notebook). </p>
Python and Jupyter Notebook for Medical Image Analysis - OpenMRBenelux 2020
<p>Dataset for the workshop "Python and Jupyter Notebook for Medical Image Analysis" at OpenMRBenelux - January 22, 2020 - Nijmegen (The Netherlands)</p>
Complete set of raw and processed datasets, as well as associated Jupyter notebooks for analysis, associated with manuscript entitled: "The MOUSE project: a practical approach for obtaining traceable, wide-range X-ray scattering information"
<p>This dataset is a complete set of raw, processed and analyzed data, complete with Jupiter notebooks, associated with the manuscript mentioned in the title. </p> <p>In the manuscript, we provide a ``systems architecture''-like overview and detailed discussions of the methodological and instrumental components that, together, comprise the "MOUSE" project (<strong>M</strong>ethodology <strong>O</strong>ptimization for <strong>U</strong>ltrafine <strong>S</strong>tructure <strong>E</strong>xploration). Through this project, we aim to provide a comprehensive methodology for obtaining the highest quality X-ray scattering information (at small and wide angles) from measurements on materials science samples. </p>
Jupyter Notebook Activity Dataset (rsds-20241113)
<h2>List of data</h2> <ul> <li>rsds-20241113.zip: Collection of SQLite database files</li> <li>image.tar.gz: Docker image provided in our data collection experiment</li> <li>redspot-341ffa5.zip: Redspot source code (<a href="https://github.com/tomokinakamaru/redspot/tree/341ffa56cb941b6f1ad74bd50d23fcf0ee96b270" target="_blank" rel="noopener">redspot@341ffa5</a>)</li> </ul> <div> <h2>Extended version of Section 2D of our paper</h2> Redspot is a Jupyter extension (i.e., Python package) that records activity signals. However, it also offers interfaces to read recorded signals. The following shows the most basic usage of its command-line interface:<br> <div> </div> <div><code>redspot replay <path-to-db></code></div> <br> <div>This command generates snapshots (.ipynb files) restored from the signal records. Note that this command does not produce a snapshot for every signal. Since the change represented by a single signal is typically minimal (e.g., one keystroke), generating a snapshot for each signal results in a meaninglessly large number of snapshots. <em><strong>However, we want to obtain signal-level snapshots for some analyses. In such cases, one can analyze them using the application programming interfaces:</strong></em></div> <br> <div><code>from redspot import database</code></div> <div><code>from redspot.notebook import Notebook</code></div> <div><code>nbk = Notebook()</code></div> <div><code>for signal in database.get("path-to-db"):</code></div> <div><code> time, panel, kind, args = signal</code></div> <div><code> nbk.apply(kind, args) # apply change</code></div> <div><code> print(nbk) # print notebook</code></div> <br> <div>To record activities, one needs to run the Redspot command in the recording mode as follows:</div> <br> <div><code>redspot record</code></div> <br> <div>This command launches Jupyter Notebook with Redspot enabled. Activities made in the launched environment are stored in an SQLite file named ``redspot.db'' under the current path.</div> <br> <div>To launch the environment we provided to the participants, one first needs to download and import the image (image.tar.gz). One can then run the image with the following command:</div> <br> <div><code>docker run --rm -it -p8888:8888 <image-name></code></div> <br> <div>Note that the SQLite file is generated in the running container. The file can be downloaded into the host machine via the file viewer of Jupyter Notebook.</div> </div>
Dataset and code for "ReSplit: Improving the Structure of Jupyter Notebooks by Re-Splitting Their Cells"
<pre>This package contains the data and code for the paper "ReSplit: Improving the Structure of Jupyter Notebooks by Re-Splitting Their Cells".</pre> <pre> </pre> <pre>`code.zip` contains the source code of our approach, as well as the details about the conducted survey.</pre> <pre>`data.zip` contains all the data used in our work, both before and after running ReSplit, as well as the results of the survey.</pre> <pre> </pre> <pre>You can find more details in READMEs in each archive. Please note that running the tool requires unzipping the data and placing it into the correct folder, according to the inner README.</pre>
Outputs of the Jupyter Notebook - MODIS MOD021KM and FIRMS
<p>The dataset contains the outputs of the notebook "MODIS MOD021KM and FIRMS" published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Samuel Jackson (author), Science & Technology Facilities Council, <a href="https://github.com/samueljackson92">@samueljackson92</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <p>MOD021KM</p> <ul> <li> <p>MODIS Characterization Support Team (MCST)</p> </li> <li> <p>MODIS Adaptive Processing System (MODAPS)</p> </li> </ul> <p>Firms</p> <ul> <li> <p>University of Maryland</p> </li> </ul> <p><em>Dataset authors</em></p> <p>MOD021KM</p> <ul> <li> <p>MODIS Science Data Support Team (SDST)</p> </li> </ul> <p>Firms</p> <ul> <li> <p>NASA’s Applied Sciences Program</p> </li> </ul> <p><em>Dataset documentation</em></p> <ul> <li> <p>Louis Giglio, Wilfrid Schroeder, Joanne V. Hall, and Christopher O. Justice. MODIS Collection 6 Active Fire Product User’s Guide Revision B. Technical Report, NASA, 2018. URL: <a href="https://modis-fire.umd.edu/files/MODIS_C6_Fire_User_Guide_B.pdf">https://modis-fire.umd.edu/files/MODIS_C6_Fire_User_Guide_B.pdf</a>.</p> </li> </ul> <p> </p>
Additional 1000 notebooks for the paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts"
<p>Additional 1000 notebooks for the review of the paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts".</p> <p>The notebooks can be processed using Matroskin tool from the supplementary materials. To do that, place:</p> <p>- the folder "1k_notebooks_dataset" into the folder ".../databases/datasets"</p> <p>- the file "mapping_of_1k_notebooks.json" into the folder ".../databases/mappings"</p>
Outputs of the Jupyter Notebook - Cosmos-UK soil moisture
<p>The dataset contains the outputs of the notebook "Cosmos-UK soil moisture" published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Alejandro Coca-Castro (author), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></p> </li> <li> <p>Doran Khamis (reviewer), UK Centre for Ecology & Hydrology, <a href="https://github.com/dorankhamis">@dorankhamis</a></p> </li> <li> <p>Matt Fry (reviewer), UK Centre for Ecology & Hydrology, <a href="https://github.com/mattfry-ceh">@mattfry-ceh</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <ul> <li> <p>UK Centre for Ecology & Hydrology (creator)</p> </li> <li> <p>Natural Environment Research Council (support)</p> </li> </ul> <p><em>Dataset reference and documentation</em></p> <ul> <li> <p>S. Stanley, V. Antoniou, A. Askquith-Ellis, L.A. Ball, E.S. Bennett, J.R. Blake, D.B. Boorman, M. Brooks, M. Clarke, H.M. Cooper, N. Cowan, A. Cumming, J.G. Evans, P. Farrand, M. Fry, O.E. Hitt, W.D. Lord, R. Morrison, G.V. Nash, D. Rylett, P.M. Scarlett, O.D. Swain, M. Szczykulska, J.L. Thornton, E.J. Trill, A.C. Warwick, and B. Winterbourn. Daily and sub-daily hydrometeorological and soil data (2013-2019) [cosmos-uk]. 2021. URL: <a href="https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185">https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185</a>, <a href="https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185">doi:10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185</a>.</p> </li> </ul> <p><strong>Further references</strong></p> <ul> <li> <p>Jonathan G. Evans, H. C. Ward, J. R. Blake, E. J. Hewitt, R. Morrison, M. Fry, L. A. Ball, L. C. Doughty, J. W. Libre, O. E. Hitt, D. Rylett, R. J. Ellis, A. C. Warwick, M. Brooks, M. A. Parkes, G. M.H. Wright, A. C. Singer, D. B. Boorman, and A. Jenkins. Soil water content in southern england derived from a cosmic-ray soil moisture observing system – cosmos-uk. <em>Hydrological Processes</em>, 30:4987–4999, 12 2016. <a href="https://doi.org/10.1002/hyp.10929">doi:10.1002/hyp.10929</a>.</p> </li> <li> <p>M. Zreda, W. J. Shuttleworth, X. Zeng, C. Zweck, D. Desilets, T. Franz, and R. Rosolem. Cosmos: the cosmic-ray soil moisture observing system. <em>Hydrology and Earth System Sciences</em>, 16(11):4079–4099, 2012. URL: <a href="https://hess.copernicus.org/articles/16/4079/2012/">https://hess.copernicus.org/articles/16/4079/2012/</a>, <a href="https://doi.org/10.5194/hess-16-4079-2012">doi:10.5194/hess-16-4079-2012</a>.</p> </li> </ul>
Dataset and Jupyter notebook for "pyDARN: A Python Software for Visualizing SuperDARN Radar Data"
<p>SuperDARN radar dataset and Jupyter notebook used to generate figures for "pyDARN: A Python Software for Visualizing SuperDARN Radar Data".</p>
Density field data, VTK files and Jupyter notebook for dislocation embedding analysis (with periodic boundary condition)
<p><strong>pbc_dataset.zip</strong></p> <p>The density field data is only for 0 degrees of misorientation with loading directions along [100], [110], [111], [234]. </p> <p>Folder paths for low density and low resolution simulation are:</p> <ol> <li>0deg/dir100/2.5e+13/config3/10x10x10</li> <li>0deg/dir110/2.5e+13/config3/10x10x10</li> <li>0deg/dir111/2.5e+13/config3/10x10x10</li> <li>0deg/dir234/2.5e+13/config3/10x10x10</li> </ol> <p>Folder paths for high density and high resolutions are:</p> <ol> <li>0deg/dir100/1e+14/config1/20x20x20</li> <li>0deg/dir110/1e+14/config1/20x20x20</li> <li>0deg/dir111/1e+14/config1/20x20x20</li> <li>0deg/dir234/1e+14/config1/20x20x20</li> </ol> <p>Each set of simulation has 2000 density field data files. Total files : 16000</p> <p>Please ensure that above paths are entered in the Jupyter notebook script file.</p> <p> </p> <p><strong>vtk.zip</strong></p> <p>VTK files for each simulation. Each simulation has 2000 files. </p> <p><strong>Dislocation_embeddings_pbc.ipynb</strong></p> <p>Jupyter notebook to generate dislocation embeddings for uploaded dataset.</p> <p> </p>
The code and data for paper: "Observing Fine-Grained Changes in Jupyter Notebooks During Development Time"
<div> <p>This package represents supplementary materials for the paper "Observing Fine-Grained Changes in Jupyter Notebooks During Development Time". Please refer to README in the archive for details.</p> </div>
Datasets, scripts and Jupyter Notebook for "Two distinct magma storage regions at Ambrym volcano detected by satellite geodesy", Geophysical Research Letters
<p>This repository includes scripts and files necessary to create Figure 1 (<strong>S1.zip </strong>and <strong>plot_TS_Ambrym_2019_2022.py</strong>) in "Two distinct magma storage regions at Ambrym volcano detected by satellite geodesy", <em>Geophysical Research Letters</em>. We also include the Jupyter Notebook used to run the EnKF data assimilation (<strong>enkf_notebook.zip) </strong>and produce Figures 2 and 3. The Jupyter Notebook and files used to produce Figure 4b,c can be found on <a href="http://github.com/tshreve/jupyterNBs/">GitHub</a>.</p> <p>This version corrects a bug in the code used to plot the cross-sections in Figure 3 with <strong>enkf_notebook.zip</strong>.</p>
tashley/particle_tracking_data: Added Jupyter notebook
<p>Data and code related to the manuscript "Probability distributions of particle hop distance and travel time over equilibrium mobile bedforms" (Ashley et al, in revision)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.