Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

42

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

42 results for “Jupyter Notebook”

Learn how ShareScore rates datasets ↗
zenodo48/100

Datasets and Jupyter notebook for the structural analysis of protein-RNA interface evolution

<p>The present repository contains data and code related to our manuscript "Structural comparison of protein-RNA homologous interfaces reveals widespread overall conservation contrasted with versatility in polar contacts". In the manuscript, we analyze the evolution of protein-RNA interfaces by building a dataset of protein-RNA interologs (homologous interfaces) and exploring how interface contacts are conserved between homologous interfaces, as well as possible explanations for non-conserved contacts.</p> <p>This repository contains the following files:</p> <ul> <li>DataAnalysisNotebook.ipynb is a Jupyter notebook to reproduce contact conservation analysis and all figures from our manuscript, and to explore data</li> <li>env.yaml is an environment file in order to build a Conda/Mamba environment to run the Jupyter notebook&nbsp;</li> <li>2022-02-21-PDB.csv contains data from the PDB about 3D structures of complexes containing interacting protein and RNA chains (PDB structure identifier, chain identifiers, experimental technique and resolution)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.tsv contains more detailed information about interacting protein and RNA chains from these complexes (PDB and chain identifiers, protein and RNA size, interface size and number of contacts)</li> <li>2022-02-21-PDB_proteinchainscontactingRNAchains.groupbp.txt.selectXE_2.50_p30_r10_pi5_ri5_rep_bc-100.out_RNAcl_0.99.tsv contains the same detailed information, restricted to the filtered dataset used as a starting point in our interolog search pipeline</li> <li>PDBinterfaceAlign.csv contains information about the structural alignment of pairs of protein-RNA interactions (structural alignment TM-scores, sequence identity and coverage)</li> <li>DataInterologsParam.tsv contains information about a pre-filtered set of 2587 potential interologs (including interface RMSD, sequence identity and coverage and interface size)</li> <li>DataInterologsContactsFixedSASA.tsv contains detailed information about conserved and non-conserved contacts in the final set of 2022 interologs (atomic contacts, apolar contacts, hydrogen bonds, salt bridges and stacking information for aminoacid-nucleotide pairs, as well as information about whether each belongs to the interface, secondary structures, and the aminoacid surface accessibility and evolutionary conservation metrics) - compared to version 1, the calculation of solvent accessibility was fixed for a number of interolog pairs</li> <li>DataCons.csv contains precomputed contact conservation metrics for each of the 2022 interolog pairs, for fast reproduction of manuscript figures</li> <li>DataInterologsContactsResampledMaintainStructSeqId.tsv, DataInterologsContactsShuffled.tsv and DataInterologsShuffled.tsv relate to baselines computed for contact conservation assessment</li> <li>clan.txt, clan_membership.txt, ecod.latest.domains.uniq.txt, rfam_interfaces_977.txt, DataGroupsECOD.tsv, DataGroupesRFAM.tsv, DataGroupsRFAMClan.tsv, DataInterfaceGroupsECOD.tsv and DataInterfaceGroupsRFAM.tsv relate to the ECOD (respectively Rfam) classification of protein domains (respectively RNA) in protein-RNA interfaces from our dataset</li> <li>ListeIntraHbonds.pkl and ListeIntraSaltBridges.pkl are pickle-format data files containing intra-molecular hydrogen bonds and salt bridges (respectively) that are used to analyse scenarii of compensation for non-conserved polar contacts.</li> </ul>

opencc-by-4.0May 2024View details →
zenodo44/100

Outputs of the Jupyter Notebook - Detecting floating objects using Deep Learning and Sentinel-2 imagery

<p>The dataset contains the outputs of the notebook &quot;Detecting floating objects using Deep Learning and Sentinel-2 imagery&quot;&nbsp;published in the ocean modelling section of The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Modelling codebase</em></p> <ul> <li> <p>Jamila Mifdal (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/jmifdal">@jmifdal</a></p> </li> <li> <p>Raquel Carmo (author), European Space Agency &Phi;-lab,&nbsp;<a href="https://github.com/raquelcarmo">@raquelcarmo</a></p> </li> <li> <p>Marc Ru&szlig;wurm (author), EPFL-ECEO,&nbsp;<a href="https://github.com/MarcCoru">@marccoru</a></p> </li> </ul>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Outputs of the Jupyter Notebook - Met Office UKV high-resolution atmosphere model data

<p>The dataset contains the outputs of the notebook &quot;Met Office UKV high-resolution atmosphere model data&quot;&nbsp;published in the urban&nbsp;sensors section of The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Samantha V. Adams (author), Met Office Informatics Lab,&nbsp;<a href="https://github.com/svadams">@svadams</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <ul> <li> <p>Met Office Informatics Lab (creator)</p> </li> <li> <p>Microsoft (support)</p> </li> <li> <p>European Regional Development Fund (support)</p> </li> </ul> <p><em>Dataset authors</em></p> <ul> <li> <p>Met Office</p> </li> </ul> <p><em>Dataset documentation</em></p> <ul> <li> <p>Theo McCaie. Met office and partners offer data and compute platform for covid-19 researchers. URL:&nbsp;<a href="https://medium.com/informatics-lab/met-office-and-partners-offer-data-and-compute-platform-for-covid-19-researchers-83848ac55f5f">https://medium.com/informatics-lab/met-office-and-partners-offer-data-and-compute-platform-for-covid-19-researchers-83848ac55f5f</a>.</p> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Outputs of the Jupyter Notebook - Tree crown detection using DeepForest

<p>The dataset contains the outputs of the notebook &quot;Tree crown detection using DeepForest&quot;&nbsp;published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Alejandro Coca-Castro (author), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></p> </li> <li> <p>Matt Allen (reviewer), Department of Geography - University of Cambridge,&nbsp;<a href="https://github.com/mja2106">@mja2106</a></p> </li> </ul> <p><em>Modelling codebase</em></p> <ul> <li> <p>Ben Weinstein (maintainer &amp; developer), University of Florida,&nbsp;<a href="https://github.com/bw4sz">@bw4sz</a></p> </li> <li> <p>Henry Senyondo (support maintainer), University of Florida,&nbsp;<a href="https://github.com/henrykironde">@henrykironde</a></p> </li> <li> <p>Ethan White (PI and author), University of Florida,&nbsp;<a href="https://github.com/ethanwhite">@weecology</a></p> </li> <li> <p>Other contributors are listed in the&nbsp;<a href="https://github.com/weecology/DeepForest/graphs/contributors">GitHub repo</a></p> </li> </ul> <p><em>Modelling publications</em></p> <ul> <li> <p>Ben&nbsp;G Weinstein, Sergio Marconi, M&eacute;laine Aubry-Kientz, Gregoire Vincent, Henry Senyondo, and Ethan&nbsp;P White. Deepforest: a python package for rgb deep learning tree crown delineation.&nbsp;<em>Methods in Ecology and Evolution</em>, 11:1743&ndash;1751, 2020. URL:&nbsp;<a href="https://besjournals.onlinelibrary.wiley.com/doi/abs/10.1111/2041-210X.13472">https://besjournals.onlinelibrary.wiley.com/doi/abs/10.1111/2041-210X.13472</a>,&nbsp;<a href="https://doi.org/https://doi.org/10.1111/2041-210X.13472">doi:https://doi.org/10.1111/2041-210X.13472</a>.</p> </li> <li> <p>Ben&nbsp;G Weinstein, Sergio Marconi, Stephanie Bohlman, Alina Zare, and Ethan White. Individual tree-crown detection in rgb imagery using semi-supervised deep learning neural networks.&nbsp;<em>Remote Sensing</em>, 2019. URL:&nbsp;<a href="https://www.mdpi.com/2072-4292/11/11/1309">https://www.mdpi.com/2072-4292/11/11/1309</a>,&nbsp;<a href="https://doi.org/10.3390/rs11111309">doi:10.3390/rs11111309</a>.</p> </li> <li> <p>Ben&nbsp;G Weinstein, Sergio Marconi, Stephanie&nbsp;A Bohlman, Alina Zare, and Ethan&nbsp;P White. Cross-site learning in deep learning rgb tree crown detection.&nbsp;<em>Ecological Informatics</em>, 56:101061, 2020. URL:&nbsp;<a href="https://www.sciencedirect.com/science/article/pii/S157495412030011X">https://www.sciencedirect.com/science/article/pii/S157495412030011X</a>,&nbsp;<a href="https://doi.org/https://doi.org/10.1016/j.ecoinf.2020.101061">doi:https://doi.org/10.1016/j.ecoinf.2020.101061</a>.</p> </li> </ul>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Outputs of the Jupyter Notebook - SEVIRI Level 1.5

<p>The dataset contains the outputs of the notebook &quot;SEVIRI Level 1.5&quot;&nbsp;published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Samuel Jackson (author), Science &amp; Technology Facilities Council,&nbsp;<a href="https://github.com/samueljackson92">@samueljackson92</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a>, 18/01/22 (latest revision)</p> </li> </ul> <p><em>Dataset originator/creator</em></p> <p>SEVIRI Level 1.5 Image Data - MSG - 0 degree</p> <ul> <li> <p>European Organisation for the Exploitation of Meteorological Satellites (EUMETSAT)</p> </li> </ul> <p>FRPPIXEL</p> <ul> <li> <p>Land Surface Analysis, Satellite Application Facility on Land Surface Analysis (LSA SAF)</p> </li> </ul> <p><em>Dataset authors</em></p> <p>SEVIRI Level 1.5 Image Data - MSG - 0 degree</p> <ul> <li> <p>European Organisation for the Exploitation of Meteorological Satellites (EUMETSAT)</p> </li> </ul> <p>FRPPIXEL</p> <ul> <li> <p>Land Surface Analysis, Satellite Application Facility on Land Surface Analysis (LSA SAF)</p> </li> </ul> <p><em>Dataset documentation</em></p> <ul> <li> <p>Martin Wooster, Jiangping He, Weidong Xu, and Alessio Lattanzio. Frp - product user manual. URL:&nbsp;<a href="https://nextcloud.lsasvcs.ipma.pt/s/pnDEepeq8zqRyrq">https://nextcloud.lsasvcs.ipma.pt/s/pnDEepeq8zqRyrq</a>&nbsp;(visited on 2021-11-18).</p> </li> <li> <p>MJ&nbsp;Wooster, G&nbsp;Roberts, PH&nbsp;Freeborn, W&nbsp;Xu, Y&nbsp;Govaerts, R&nbsp;Beeby, J&nbsp;He, A&nbsp;Lattanzio, D&nbsp;Fisher, and R&nbsp;Mullen. Lsa saf meteosat frp products&ndash;part 1: algorithms, product contents, and analysis.&nbsp;<em>Atmospheric Chemistry and Physics</em>, 15(22):13217&ndash;13239, 2015.</p> </li> </ul>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Outputs of the Jupyter Notebook - Tree crown delineation using detectreeRGB

<p>The dataset contains the outputs of the notebook &quot;Tree crown detection using DeepForest&quot;&nbsp;published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li>Sebastian H. M. Hickman (author), University of Cambridge,&nbsp;<a href="https://github.com/shmh40">@shmh40</a></li> <li>Alejandro Coca-Castro (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></li> </ul> <p><em>Modelling codebase</em></p> <ul> <li>Sebastian H. M. Hickman (author), University of Cambridge&nbsp;<a href="https://github.com/shmh40">@shmh40</a></li> <li>James G. C. Ball (contributor), University of Cambridge&nbsp;<a href="https://github.com/PatBall1">@PatBall1</a></li> <li>David A. Coomes (contributor), University of Cambridge</li> <li>Toby Jackson (contributor), University of Cambridge</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo44/100

Outputs of the Jupyter Notebook - Sea ice forecasting using IceNet

<p>The dataset contains the outputs of the notebook &quot;Sea ice forecasting using IceNet&quot;&nbsp;published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li>Alejandro Coca-Castro (author), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></li> <li>Tom R. Andersson (reviewer), British Antarctic Survey,&nbsp;<a href="https://github.com/tom-andersson">@tom-andersson</a></li> <li>Nick Barlow (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/nbarlowATI">@nbarlowATI</a></li> </ul> <p><em>Modelling codebase</em></p> <ul> <li>Tom R. Andersson (author), British Antarctic Survey,&nbsp;<a href="https://github.com/tom-andersson">@tom-andersson</a></li> <li>James Byrne (contributor), British Antarctic Survey,&nbsp;<a href="https://github.com/JimCircadian">@JimCircadian</a></li> <li>Tony Phillips (contributor), British Antarctic Survey</li> </ul>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Inputs of the Jupyter Notebook - Met Office UKV high-resolution atmosphere model data

<p>The dataset contains the inputs of the notebook &quot;Met Office UKV high-resolution atmosphere model data&quot;&nbsp;published in The Environmental Data Science Book.</p> <p>The input data refer to a subset of&nbsp;single sample data file for 1.5 m temperature as part of the Met Office&nbsp;contribution to the COVID 19 modelling effort.</p> <p>The full dataset was&nbsp;available for download from the Met Office Azure (https://metdatasa.blob.core.windows.net/covid19-response-non-commercial/).&nbsp;The full dataset was available for&nbsp;download&nbsp;under the terms of non-commercial purposes.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Samantha V. Adams (author), Met Office Informatics Lab,&nbsp;<a href="https://github.com/svadams">@svadams</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <ul> <li> <p>Met Office Informatics Lab (creator)</p> </li> <li> <p>Microsoft (support)</p> </li> <li> <p>European Regional Development Fund (support)</p> </li> </ul> <p><em>Dataset authors</em></p> <ul> <li> <p>Met Office</p> </li> </ul> <p><em>Dataset documentation</em></p> <ul> <li> <p>Theo McCaie. Met office and partners offer data and compute platform for covid-19 researchers. URL:&nbsp;<a href="https://medium.com/informatics-lab/met-office-and-partners-offer-data-and-compute-platform-for-covid-19-researchers-83848ac55f5f">https://medium.com/informatics-lab/met-office-and-partners-offer-data-and-compute-platform-for-covid-19-researchers-83848ac55f5f</a>.</p> </li> </ul> <p><strong>Note this data should be used only for non-commercial purposes.</strong></p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Outputs of the Jupyter Notebook - Exploring Land Cover Data (Impact Observatory)

<p>The dataset contains the outputs of the notebook &quot;Exploring Land Cover Data (Impact Observatory)&quot;&nbsp;published in The Environmental Data Science Book.</p>

opencc-by-4.0Sep 2022View details →
zenodo44/100

Supporting Jupyter Python notebook for "A new class of efficient randomized benchmarking protocols"

<p>Python notebook containing the code used to generate the data for figure 2&nbsp;in the appendix of &quot;A new class of efficient randomized benchmarking protocols&quot; (arXiv:1806.02048).</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

Outputs of the Jupyter Notebook - Learning the Underlying Physics of a Simulation Model of the Ocean's Temperature (CIRC23)

<p>The dataset contains the outputs of the notebook &quot;Learning the Underlying Physics of a Simulation Model of the Ocean&#39;s Temperature (CIRC23)&quot;&nbsp;published in The Environmental Data Science Book.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

Modelica Models and Jupyter Notebooks for System Analysis of Glucose Insulin Regulation

<p>This dataset contains source code of Modelica models of Glucose-Insulin regulation using different techniques.</p> <p>Accompanying Jupyter notebook is demo for system analysis (parameter estimation) of artificial data and to match model simulation able to be used in Teaching class.</p> <ul> <li><strong>ModelicaIdentification.ipynb</strong> - default notebook - code contains ellipsis which needs to be replaced as per instruction in text</li> <li><strong>ModelicaIdentificationResolution.ipynb - </strong>notebook - code with exemplar solution to default notebook</li> <li><strong>glucoseinsulin.mo - </strong>Modelica source code</li> <li><strong>PatientInsulinConcentration.csv</strong> - sample data to be fitted against model</li> <li><strong>seminar11hw.GIExperiment.fmu</strong> - FMU exported from Modelica in order to run simulation in Python and PyFMI library</li> </ul> <p>Thanks to the MYBINDER service, the Jupyter notebook can be viewed and executed as</p> <ul> <li><a href="https://mybinder.org/v2/zenodo/10.5281/zenodo.3633324/">https://mybinder.org/v2/zenodo/10.5281/zenodo.3633324/</a> note that you need to launch terminal first in Jupyter -&gt; New -&gt; Terminal and install pyfmi and matplotlib by:</li> </ul> <pre><code class="language-bash">conda install -c conda-forge pyfmi matplotlib</code></pre> <ul> <li>Most recent version with other models and notebooks <a href="https://mybinder.org/v2/gh/creative-connections/Bodylight-notebooks/master?filepath=Seminar11GlucoseInsulinIdentification/">https://mybinder.org/v2/gh/creative-connections/Bodylight-notebooks/master?filepath=Seminar11GlucoseInsulinIdentification/</a></li> </ul>

opencc-by-4.0Jan 2020View details →
zenodo40/100

DistilKaggle: a distilled dataset of Kaggle Jupyter notebooks

<h2><strong>Overview</strong></h2> <p>DistilKaggle is a curated dataset extracted from Kaggle Jupyter notebooks spanning from September 2015 to October 2023. This dataset is a distilled version derived from the download of over 300GB of Kaggle kernels, focusing on essential data for research purposes. The dataset exclusively comprises publicly available Python Jupyter notebooks from Kaggle. The essential information for retrieving the data needed to download the dataset is obtained from the MetaKaggle dataset provided by Kaggle.</p> <h2><strong>Contents</strong></h2> <p>The DistilKaggle dataset consists of three main CSV files:</p> <p><strong>code.csv:</strong> Contains over 12 million rows of code cells extracted from the Kaggle kernels. Each row is identified by the kernel's ID and cell index for reproducibility.</p> <p><strong>markdown.csv:</strong> Includes over 5 million rows of markdown cells extracted from Kaggle kernels. Similar to <strong>code.csv</strong>, each row is identified by the kernel's ID and cell index.</p> <p><strong>notebook_metrics.csv:</strong> This file provides notebook features described in the accompanying paper released with this dataset. It includes metrics for over 517,000 Python notebooks.</p> <h2><strong>Directory Structure</strong></h2> <p>The <strong>kernels</strong> directory is organized based on Kaggle's Performance Tiers (PTs), a ranking system in Kaggle that classifies users. The structure includes PT-specific directories, each containing user ids that belong to this PT, download logs, and the essential data needed for downloading the notebooks.</p> <p>The <strong>utility</strong> directory contains two important files:</p> <p><strong>aggregate_data.py:</strong> A Python script for aggregating data from different PTs into the mentioned CSV files.</p> <p><strong>application.ipynb:</strong> A Jupyter notebook serving as a simple example application using the metrics dataframe. It demonstrates predicting the PT of the author based on notebook metrics.</p> <p><strong>DistilKaggle.tar.gz: </strong>It is just the compressed version of the whole dataset. If you downloaded all of the other files independently already, there is no need to download this file.</p> <h2><strong>Usage</strong></h2> <p>Researchers can leverage this distilled dataset for various analyses without dealing with the bulk of the original 300GB dataset. For access to the raw, unprocessed Kaggle kernels, researchers can request the dataset directly.</p> <h2><strong>Note</strong></h2> <p>The original dataset of Kaggle kernels is substantial, exceeding 300GB, making it impractical for direct upload to Zenodo. Researchers interested in the full dataset can contact the dataset maintainers for access.</p> <h2><strong>Citation</strong></h2> <p>If you use this dataset in your research, please cite the accompanying paper or provide appropriate acknowledgment as outlined in the documentation.</p> <p>If you have any questions regarding the dataset, don't hesitate to contact me at <a href="mailto:mohammad.abolnejadian@gmail.com">mohammad.abolnejadian@gmail.com</a></p> <p>Thank you for using DistilKaggle!</p>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Dataset of Jupyter Notebooks from the paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts"

<pre>This archive contains the dataset of properly-licensed Jupyter notebooks from the MSR&#39;22 paper &quot;A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts&quot;. The dataset contains 847,881 notebooks stored in the PostgreSQL dump file. You can find the details about the database in the README file. To transform the notebooks into this convenient format and to calcuate the structural metrics, we used our library called Matroskin, which can be found here: <a href="https://github.com/JetBrains-Research/Matroskin">https://github.com/JetBrains-Research/Matroskin</a>. </pre>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Inputs of the Jupyter Notebook - Cosmos-UK soil moisture

<p>The dataset contains the inputs of the notebook &quot;Cosmos-UK soil moisture&quot;&nbsp;published in The Environmental Data Science Book.</p> <p>The input data refer to a subset of the public 2013-2019 COSMOS-UK dataset, daily and subhourly observations and metadata for four stations:&nbsp;WYTH1,&nbsp;WADDN,&nbsp;SHEEP and&nbsp;CHIMN.&nbsp;These stations represent the first sites to prototype COSMOS sensors in the UK, see further details in Evans et al.&nbsp;(2016) and they are situated in human-intervened areas (grassland and cropland), except for one in a woodland land cover site.</p> <p>Data from COSMOS-UK up to the end of 2019 are available for download from the UKCEH Environmental Information Data Centre (EIDC). The data are accompanied by documentation that describes the site-specific instrumentation, data and processing including quality control. The full dataset is available for <a href="https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185">download</a>&nbsp;under the terms of the Open Government License.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Alejandro Coca-Castro (author), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></p> </li> <li> <p>Doran Khamis (reviewer), UK Centre for Ecology &amp; Hydrology,&nbsp;<a href="https://github.com/dorankhamis">@dorankhamis</a></p> </li> <li> <p>Matt Fry (reviewer), UK Centre for Ecology &amp; Hydrology,&nbsp;<a href="https://github.com/mattfry-ceh">@mattfry-ceh</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <ul> <li> <p>UK Centre for Ecology &amp; Hydrology (creator)</p> </li> <li> <p>Natural Environment Research Council (support)</p> </li> </ul> <p><em>Dataset reference and documentation</em></p> <ul> <li> <p>S.&nbsp;Stanley, V.&nbsp;Antoniou, A.&nbsp;Askquith-Ellis, L.A. Ball, E.S. Bennett, J.R. Blake, D.B. Boorman, M.&nbsp;Brooks, M.&nbsp;Clarke, H.M. Cooper, N.&nbsp;Cowan, A.&nbsp;Cumming, J.G. Evans, P.&nbsp;Farrand, M.&nbsp;Fry, O.E. Hitt, W.D. Lord, R.&nbsp;Morrison, G.V. Nash, D.&nbsp;Rylett, P.M. Scarlett, O.D. Swain, M.&nbsp;Szczykulska, J.L. Thornton, E.J. Trill, A.C. Warwick, and B.&nbsp;Winterbourn. Daily and sub-daily hydrometeorological and soil data (2013-2019) [cosmos-uk]. 2021. URL:&nbsp;<a href="https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185">https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185</a>,&nbsp;<a href="https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185">doi:10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185</a>.</p> </li> </ul> <p><strong>Further references</strong></p> <ul> <li> <p>Jonathan&nbsp;G. Evans, H.&nbsp;C. Ward, J.&nbsp;R. Blake, E.&nbsp;J. Hewitt, R.&nbsp;Morrison, M.&nbsp;Fry, L.&nbsp;A. Ball, L.&nbsp;C. Doughty, J.&nbsp;W. Libre, O.&nbsp;E. Hitt, D.&nbsp;Rylett, R.&nbsp;J. Ellis, A.&nbsp;C. Warwick, M.&nbsp;Brooks, M.&nbsp;A. Parkes, G.&nbsp;M.H. Wright, A.&nbsp;C. Singer, D.&nbsp;B. Boorman, and A.&nbsp;Jenkins. Soil water content in southern england derived from a cosmic-ray soil moisture observing system &ndash; cosmos-uk.&nbsp;<em>Hydrological Processes</em>, 30:4987&ndash;4999, 12 2016.&nbsp;<a href="https://doi.org/10.1002/hyp.10929">doi:10.1002/hyp.10929</a>.</p> </li> <li> <p>M.&nbsp;Zreda, W.&nbsp;J. Shuttleworth, X.&nbsp;Zeng, C.&nbsp;Zweck, D.&nbsp;Desilets, T.&nbsp;Franz, and R.&nbsp;Rosolem. Cosmos: the cosmic-ray soil moisture observing system.&nbsp;<em>Hydrology and Earth System Sciences</em>, 16(11):4079&ndash;4099, 2012. URL:&nbsp;<a href="https://hess.copernicus.org/articles/16/4079/2012/">https://hess.copernicus.org/articles/16/4079/2012/</a>,&nbsp;<a href="https://doi.org/10.5194/hess-16-4079-2012">doi:10.5194/hess-16-4079-2012</a>.</p> </li> </ul>

opencc-by-4.0May 2022View details →
zenodo40/100

Outputs of the Jupyter Notebook - Concatenating a gridded rainfall reanalysis dataset into a time series

<p>The dataset contains the outputs of the notebook &quot;Concatenating a gridded rainfall reanalysis dataset into a time series&quot;&nbsp;published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Timothy Lam (author), University of Exeter,&nbsp;<a href="https://github.com/timo0thy">@timo0thy</a></p> </li> <li> <p>Marlene Kretschmer (author), University of Reading,&nbsp;<a href="https://github.com/MarleneKretschmer">@MarleneKretschmer</a></p> </li> <li> <p>Samantha Adams (author), Met Office Informatics Lab,&nbsp;<a href="https://github.com/svadams">@svadams</a></p> </li> <li> <p>Rachel Prudden (author), Met Office Informatics Lab,&nbsp;<a href="https://github.com/RPrudden">@RPrudden</a></p> </li> <li> <p>Elena Saggioro (author), University of Reading,&nbsp;<a href="https://github.com/ESaggioro">@ESaggioro</a></p> </li> <li> <p>Nick Homer (reviewer), University of Edinburgh,&nbsp;<a href="https://github.com/NHomer">@NHomer</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute,&nbsp;<a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <ul> <li> <p>NOAA National Center for Environmental Prediction (creator)</p> </li> </ul> <p><em>Dataset authors</em></p> <ul> <li> <p>Eugenia Kalnay, Director, NCEP Environmental Modeling Center</p> </li> </ul> <p><em>Dataset documentation</em></p> <ul> <li> <p>E.&nbsp;Kalnay, M.&nbsp;Kanamitsu, R.&nbsp;Kistler, W.&nbsp;Collins, D.&nbsp;Deaven, L.&nbsp;Gandin, M.&nbsp;Iredell, S.&nbsp;Saha, G.&nbsp;White, J.&nbsp;Woollen, Y.&nbsp;Zhu, M.&nbsp;Chelliah, W.&nbsp;Ebisuzaki, W.&nbsp;Higgins, J.&nbsp;Janowiak, K.&nbsp;C. Mo, C.&nbsp;Ropelewski, J.&nbsp;Wang, A.&nbsp;Leetmaa, R.&nbsp;Reynolds, Roy Jenne, and Dennis Joseph. The ncep/ncar 40-year reanalysis project.&nbsp;Bulletin of the American Meteorological Society, 77(3):437 &ndash; 472, 1996. URL:&nbsp;<a href="https://journals.ametsoc.org/view/journals/bams/77/3/1520-0477_1996_077_0437_tnyrp_2_0_co_2.xml">https://journals.ametsoc.org/view/journals/bams/77/3/1520-0477_1996_077_0437_tnyrp_2_0_co_2.xml</a>,&nbsp;<a href="https://doi.org/10.1175/1520-0477(1996)077%3C0437:TNYRP%3E2.0.CO;2">doi:10.1175/1520-0477(1996)077&lt;0437:TNYRP&gt;2.0.CO;2</a>.</p> </li> </ul> <p><em>Pipeline documentation</em></p> <ul> <li> <p>Marlene Kretschmer, Samantha&nbsp;V. Adams, Alberto Arribas, Rachel Prudden, Niall Robinson, Elena Saggioro, and Theodore&nbsp;G. Shepherd. Quantifying causal pathways of teleconnections.&nbsp;Bulletin of the American Meteorological Society, 102(12):E2247 &ndash; E2263, 2021. URL:&nbsp;<a href="https://journals.ametsoc.org/view/journals/bams/102/12/BAMS-D-20-0117.1.xml">https://journals.ametsoc.org/view/journals/bams/102/12/BAMS-D-20-0117.1.xml</a>,&nbsp;<a href="https://doi.org/10.1175/BAMS-D-20-0117.1">doi:10.1175/BAMS-D-20-0117.1</a>.</p> </li> </ul>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Python scripts / Jupyter Notebooks and data for training segmentation models on slide scans of diatom preparations from river Menne

<p>This archive contains the Jupyter Notebooks and data used for the deep learning experiments published in Kloster et al. 2022: Improving deep learning-based segmentation of diatoms in gigapixel-sized virtual slides by object-based tile positioning and object integrity constraint.</p> <p>The notebooks are numbered according to the order in which they are to execute. Please refer to the comments and documentation within the notebooks as well as to the manuscript for details. The data (image data, mask data &amp; segmentation ground truth in COCO format for several different tiling strategies) is stored in separate subfolders corresponding with data usage (model training, validation, test) and tiling strategy. Please refer to the &quot;readme&quot; files for detailed information.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2022View details →
zenodo40/100

Data for "Interactive maps in the Jupyter notebook"

<p>Dataset used for the lesson &quot;<a href="https://annefou.github.io/jupyter_maps/index.html">Interactive maps in the Jupyter notebook</a>&quot;&nbsp;</p> <p>&nbsp;</p> <p>Taught at CarpentryConnect, Manchester 2019.&nbsp;</p>

opencc-by-4.0Jun 2019View details →
zenodo40/100

Small-angle Scattering Data Analysis Round Robin: anonymized results, figures and Jupyter notebook

<p>The intent of this round robin was to find out how comparable results from different researchers are, who analyse exactly the same processed, corrected dataset.</p> <p>This zip file contains the anonymized results and the jupyter notebook used to do the data processing, analysis and visualisation. Additionally, TEM images of the samples are included.&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Dataset of a Study of Computational reproducibility of Jupyter notebooks from biomedical publications

<p>This repository contains the dataset for the study of <a href="https://doi.org/10.1093/gigascience/giad113">computational reproducibility of Jupyter notebooks from biomedical publications</a>. Our focus lies in evaluating the extent of reproducibility of Jupyter notebooks derived from GitHub repositories linked to publications present in the biomedical literature repository, PubMed Central. We analyzed the reproducibility of Jupyter notebooks from GitHub repositories associated with publications indexed in the biomedical literature repository PubMed Central. The dataset includes the metadata information of the journals, publications, the Github repositories mentioned in the publications and the notebooks present in the Github repositories.</p> <p><strong>Data Collection and Analysis</strong></p> <p>We use the code for reproducibility of Jupyter notebooks from the study done by <a href="../record/2592524">Pimentel et al., 2019</a> and adapted the code from <a href="https://github.com/fusion-jena/ReproduceMeGit">ReproduceMeGit</a>. We provide code for collecting the publication metadata from PubMed Central using <a href="https://biopython.org/docs/1.76/api/Bio.Entrez.html">NCBI Entrez utilities via Biopython</a>.</p> <p>Our approach involves searching PMC using the esearch function for Jupyter notebooks using the query: ``(ipynb OR jupyter OR ipython) AND github''. We meticulously retrieve data in XML format, capturing essential details about journals and articles. By systematically scanning the entire article, encompassing the abstract, body, data availability statement, and supplementary materials, we extract GitHub links. Additionally, we mine repositories for key information such as dependency declarations found in files like requirements.txt, setup.py, and pipfile. Leveraging the GitHub API, we enrich our data by incorporating repository creation dates, update histories, pushes, and programming languages.</p> <p>All the extracted information is stored in a SQLite database. After collecting and creating the database tables, we ran a pipeline to collect the Jupyter notebooks contained in the GitHub repositories based on the code from Pimentel et al., 2019.</p> <p>Our reproducibility pipeline was started on 27 March 2023.</p> <p><strong>Repository Structure</strong></p> <p>Our repository is organized into two main folders:</p> <ul> <li><strong>archaeology</strong>: This directory hosts scripts designed to download, parse, and extract metadata from PubMed Central publications and associated repositories. There are 24 database tables created which store the information on articles, journals, authors, repositories, notebooks, cells, modules, executions, etc. in the db.sqlite database file.</li> <li><strong>analyses</strong>: Here, you will find notebooks instrumental in the in-depth analysis of data related to our study. The db.sqlite file generated by running the archaelogy folder is stored in the analyses folder for further analysis. The path can however be configured in the config.py file. There are two sets of notebooks: one set (naming pattern N[0-9]*.ipynb) is focused on examining data pertaining to repositories and notebooks, while the other set (PMC[0-9]*.ipynb) is for analyzing data associated with publications in PubMed Central, i.e.\ for plots involving data about articles, journals, publication dates or research fields. The resultant figures from the these notebooks are stored in the 'outputs' folder.</li> <li><strong>MethodsWorkflow</strong>: The MethodsWorkflow file provides a conceptual overview of the workflow used in this study.</li> </ul> <p><strong>Accessing Data and Resources:</strong></p> <ul> <li>All the data generated during the initial study can be accessed at https://doi.org/10.5281/zenodo.6802158</li> <li>For the latest results and re-run data, refer to this link.</li> <li>The comprehensive SQLite database that encapsulates all the study's extracted data is stored in the db.sqlite file.</li> <li>The metadata in xml format extracted from PubMed Central which contains the information about the articles and journal can be accessed in pmc.xml file.</li> </ul> <p><strong>System Requirements:</strong></p> <ul> <li>Centos 7 (Documentation: https://www.centos.org/)</li> <li>Conda 4.9.4 (Installation Guide: https://docs.anaconda.com/anaconda/install/linux/)</li> <li>Python 3.7.6 (Download Link: https://www.python.org/downloads/)</li> <li>GitHub account (Get Started: https://github.com/, Requires GitHub Username and Token)</li> <li>gcc 7.3.0 (Installation Guide: https://gcc.gnu.org/install/)</li> <li>lbzip2 (Command: `conda install -c conda-forge lbzip2')</li> </ul> <p><strong>Running the pipeline:</strong></p> <ul> <li>Clone the computational-reproducibility-pmc repository using Git:<br>git clone https://github.com/fusion-jena/computational-reproducibility-pmc.git<br>&nbsp;</li> <li>Navigate to the computational-reproducibility-pmc directory:<br>cd computational-reproducibility-pmc/computational-reproducibility-pmc</li> <li>Configure environment variables in the config.py file:<br>GITHUB_USERNAME = os.environ.get("JUP_GITHUB_USERNAME", "add your github username here")<br>GITHUB_TOKEN = os.environ.get("JUP_GITHUB_PASSWORD", "add your github token here")</li> <li>Other environment variables can also be set in the config.py file.<br>BASE_DIR = Path(os.environ.get("JUP_BASE_DIR", "./")).expanduser() # Add the path of directory where the GitHub repositories will be saved<br>DB_CONNECTION = os.environ.get("JUP_DB_CONNECTION", "sqlite:///db.sqlite") # Add the path where the database is stored.</li> <li>To set up conda environments for each python versions, upgrade pip, install pipenv, and install the archaeology package in each environment, execute:<br>source conda-setup.sh</li> <li>Change to the archaeology directory<br>cd archaeology</li> <li>Activate conda environment. We used py36 to run the pipeline.<br>conda activate py36</li> <li>Execute the main pipeline script (r0_main.py):<br>python r0_main.py</li> </ul> <p><strong>Running the analysis:</strong></p> <ul> <li>Navigate to the analysis directory.<br>cd analyses</li> <li>Activate conda environment. We use raw38 for the analysis of the metadata collected in the study.<br>conda activate raw38</li> <li>Install the required packages using the requirements.txt file.<br>pip install -r requirements.txt</li> <li>Launch Jupyterlab<br>jupyter lab</li> <li>Refer to the Index.ipynb notebook for the execution order and guidance.</li> </ul> <p><strong>References:</strong></p> <ul> <li>Sheeba Samuel, Daniel Mietchen. (2024). Computational reproducibility of Jupyter notebooks from biomedical publications, https://doi.org/10.1093/gigascience/giad113, GigaScience</li> <li>Sheeba Samuel, Daniel Mietchen. (2022). Computational reproducibility of Jupyter notebooks from biomedical publications, https://arxiv.org/pdf/2209.04308.pdf, CoRR abs/2209.04308</li> <li>Sheeba Samuel, &amp; Daniel Mietchen. (2022). Dataset of a Study of Computational reproducibility of Jupyter notebooks from biomedical publications [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6802158</li> </ul> <p>&nbsp;</p>

opencc-zeroJul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record