Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
134
datasets available to search
ShareScore release 0.9.0
Dataset results
134 results for “Notebook”
The value of increased spatial resolution of pesticide usage data for assessing risk to endangered species: Data, notebooks, and results
<p>Decision makers often cite data quality as a limitation in environmental management. Value of information approaches evaluates the benefit of new data collection for management outcomes. Pesticide exposure risk assessment for endangered species is one context where data limitations may affect decisions and a value of information type approach could be useful for identifying optimal data <span><span><span><span><span><span><span><span><span><span><span><span><span><span><span>quality and resolution. Under the U.S. Federal Insecticide, Fungicide and Rodenticide Act, the U.S. Environmental Protection Agency (EPA) is responsible for registering pesticides before they can be sold and regularly reviewing pesticides. Section 7 of the Endangered Species Act requires that the EPA consider potential impacts of pesticides to listed endangered species and critical habitats in this process, and for the Services—U.S. Fish and Wildlife Service and National Marine Fisheries Service—to complete a formal Section 7 consultation if the EPA deems it necessary. The current process is time‐intensive, lacks transparency and confidence among stakeholders, and leaves hundreds of unreviewed pesticides on the market. Increasing the resolution of pesticide usage data could address these concerns by improving estimated overlaps between species ranges and pesticide usage. Thus, we evaluated the relative importance of different resolutions of pesticide usage data for assessing expected carbaryl exposure to endangered plant species endemic to California. We found that spatially explicit, township resolution usage data (~36</span></span></span></span></span></span></span></span></span></span></span></span></span></span></span> <span><span><span><span><span><span><span><span><span><span><span><span><span><span><span>mile</span></span></span></span></span></span></span></span></span></span></span></span></span></span></span><sup>2</sup><span><span><span><span><span><span><span><span><span><span><span><span><span><span><span>) excluded 33% of terrestrial plants (55/168) and 51% their critical habitats (27/53) from requiring a Section 7 consultation, while coarser resolution data excluded none. In contrast, the EPA's biological evaluation for carbaryl only excludes 4% of terrestrial plants (nationally) from requiring formal Section 7 consultation. This suggests high‐resolution data could increase pesticide review efficiency and decrease the amount of time pesticides remain on the market without a formal evaluation.</span></span></span></span></span></span></span></span></span></span></span></span></span></span></span></p>
Dataset and code for "ReSplit: Improving the Structure of Jupyter Notebooks by Re-Splitting Their Cells"
<pre>This package contains the data and code for the paper "ReSplit: Improving the Structure of Jupyter Notebooks by Re-Splitting Their Cells".</pre> <pre> </pre> <pre>`code.zip` contains the source code of our approach, as well as the details about the conducted survey.</pre> <pre>`data.zip` contains all the data used in our work, both before and after running ReSplit, as well as the results of the survey.</pre> <pre> </pre> <pre>You can find more details in READMEs in each archive. Please note that running the tool requires unzipping the data and placing it into the correct folder, according to the inner README.</pre>
Workflow notebooks for deriving parameter of SUEWS v2020 based on FLUXNET2015 dataset
<p>This archive includes all supportive files (processing and simulation scripts and derived parameters) for revision of <a href="https://doi.org/10.5194/gmd-2020-148">GMD-2020-148</a>.</p> <p>1. ana.zip: scripts for parameter derivation and and result analysis.</p> <p>2. data.zip: preprocessed FLUXNET data and derived parameters.</p> <p>3. sim.zip: scripts for conducting SUEWS simulations.</p> <p>Please note: due to the Zenodo upload restriction, the simulation results are archived separately at: </p>
MultiVI - Intermediate datasets, notebooks, and scripts
<p>All intermediate analyses, datasets, scripts, and notebooks to reproduce all results in the MultiVI manuscript.</p>
Outputs of the Jupyter Notebook - MODIS MOD021KM and FIRMS
<p>The dataset contains the outputs of the notebook "MODIS MOD021KM and FIRMS" published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Samuel Jackson (author), Science & Technology Facilities Council, <a href="https://github.com/samueljackson92">@samueljackson92</a></p> </li> <li> <p>Alejandro Coca-Castro (reviewer), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <p>MOD021KM</p> <ul> <li> <p>MODIS Characterization Support Team (MCST)</p> </li> <li> <p>MODIS Adaptive Processing System (MODAPS)</p> </li> </ul> <p>Firms</p> <ul> <li> <p>University of Maryland</p> </li> </ul> <p><em>Dataset authors</em></p> <p>MOD021KM</p> <ul> <li> <p>MODIS Science Data Support Team (SDST)</p> </li> </ul> <p>Firms</p> <ul> <li> <p>NASA’s Applied Sciences Program</p> </li> </ul> <p><em>Dataset documentation</em></p> <ul> <li> <p>Louis Giglio, Wilfrid Schroeder, Joanne V. Hall, and Christopher O. Justice. MODIS Collection 6 Active Fire Product User’s Guide Revision B. Technical Report, NASA, 2018. URL: <a href="https://modis-fire.umd.edu/files/MODIS_C6_Fire_User_Guide_B.pdf">https://modis-fire.umd.edu/files/MODIS_C6_Fire_User_Guide_B.pdf</a>.</p> </li> </ul> <p> </p>
Additional 1000 notebooks for the paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts"
<p>Additional 1000 notebooks for the review of the paper "A Large-Scale Comparison of Python Code in Jupyter Notebooks and Scripts".</p> <p>The notebooks can be processed using Matroskin tool from the supplementary materials. To do that, place:</p> <p>- the folder "1k_notebooks_dataset" into the folder ".../databases/datasets"</p> <p>- the file "mapping_of_1k_notebooks.json" into the folder ".../databases/mappings"</p>
Outputs of the Jupyter Notebook - Cosmos-UK soil moisture
<p>The dataset contains the outputs of the notebook "Cosmos-UK soil moisture" published in The Environmental Data Science Book.</p> <p><strong>Contributions</strong></p> <p><em>Notebook</em></p> <ul> <li> <p>Alejandro Coca-Castro (author), The Alan Turing Institute, <a href="https://github.com/acocac">@acocac</a></p> </li> <li> <p>Doran Khamis (reviewer), UK Centre for Ecology & Hydrology, <a href="https://github.com/dorankhamis">@dorankhamis</a></p> </li> <li> <p>Matt Fry (reviewer), UK Centre for Ecology & Hydrology, <a href="https://github.com/mattfry-ceh">@mattfry-ceh</a></p> </li> </ul> <p><em>Dataset originator/creator</em></p> <ul> <li> <p>UK Centre for Ecology & Hydrology (creator)</p> </li> <li> <p>Natural Environment Research Council (support)</p> </li> </ul> <p><em>Dataset reference and documentation</em></p> <ul> <li> <p>S. Stanley, V. Antoniou, A. Askquith-Ellis, L.A. Ball, E.S. Bennett, J.R. Blake, D.B. Boorman, M. Brooks, M. Clarke, H.M. Cooper, N. Cowan, A. Cumming, J.G. Evans, P. Farrand, M. Fry, O.E. Hitt, W.D. Lord, R. Morrison, G.V. Nash, D. Rylett, P.M. Scarlett, O.D. Swain, M. Szczykulska, J.L. Thornton, E.J. Trill, A.C. Warwick, and B. Winterbourn. Daily and sub-daily hydrometeorological and soil data (2013-2019) [cosmos-uk]. 2021. URL: <a href="https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185">https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185</a>, <a href="https://doi.org/10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185">doi:10.5285/b5c190e4-e35d-40ea-8fbe-598da03a1185</a>.</p> </li> </ul> <p><strong>Further references</strong></p> <ul> <li> <p>Jonathan G. Evans, H. C. Ward, J. R. Blake, E. J. Hewitt, R. Morrison, M. Fry, L. A. Ball, L. C. Doughty, J. W. Libre, O. E. Hitt, D. Rylett, R. J. Ellis, A. C. Warwick, M. Brooks, M. A. Parkes, G. M.H. Wright, A. C. Singer, D. B. Boorman, and A. Jenkins. Soil water content in southern england derived from a cosmic-ray soil moisture observing system – cosmos-uk. <em>Hydrological Processes</em>, 30:4987–4999, 12 2016. <a href="https://doi.org/10.1002/hyp.10929">doi:10.1002/hyp.10929</a>.</p> </li> <li> <p>M. Zreda, W. J. Shuttleworth, X. Zeng, C. Zweck, D. Desilets, T. Franz, and R. Rosolem. Cosmos: the cosmic-ray soil moisture observing system. <em>Hydrology and Earth System Sciences</em>, 16(11):4079–4099, 2012. URL: <a href="https://hess.copernicus.org/articles/16/4079/2012/">https://hess.copernicus.org/articles/16/4079/2012/</a>, <a href="https://doi.org/10.5194/hess-16-4079-2012">doi:10.5194/hess-16-4079-2012</a>.</p> </li> </ul>
dataset of Kaggle Notebooks
<p>These are some dataset of Kaggle Notebooks</p>
Python notebooks as a pedagogical tool to teach NOT data reduction
<p>We will present a series of seven Python Jupyter Notebooks designed to teach master-level students the basic steps of data reduction for observations with the Alhambra Faint Object Spectrograph and Camera (ALFOSC). This pedagogical tool, which translates IRAF tasks into the widespread Python language, has been successfully deployed for the course “Observational Astrophysics II” at the Department of Astronomy of Stockholm University. Each notebook introduces the students to one specific task of the data reduction explaining the purpose of that task, how it is implemented and guiding the students through the completion of the task. With a hands-on approach, this allows the students to understand the reason behind each step and to test directly what is the effect of each step, thanks to interactive plots of the data and the intermediate products. In addition to its educational value, this material can be expanded to reach the quality needed for scientific works. Therefore it could offer a starting point for developing a personalised data reduction for a given scientific problem. A complete version of the material is publicly available at https://github.com/astrojuggler/data-reduction-Obs-II-course. This project has been financed by Stockholm University - Department of Astronomy with a grant earned by Professor Matthew Hayes.</p>
Dataset and Jupyter notebook for "pyDARN: A Python Software for Visualizing SuperDARN Radar Data"
<p>SuperDARN radar dataset and Jupyter notebook used to generate figures for "pyDARN: A Python Software for Visualizing SuperDARN Radar Data".</p>
Supporting Datasets and Notebook for the manuscript: "Humans program artificial delegates to accurately solve collective-risk dilemmas but lack precision"
Open the record for dataset details and reuse information.
Protein sequences for Dephosphorylation sites and fine-tuning notebook
<p>Protein sequences for Dephosphorylation sites</p>
Density field data, VTK files and Jupyter notebook for dislocation embedding analysis (with periodic boundary condition)
<p><strong>pbc_dataset.zip</strong></p> <p>The density field data is only for 0 degrees of misorientation with loading directions along [100], [110], [111], [234]. </p> <p>Folder paths for low density and low resolution simulation are:</p> <ol> <li>0deg/dir100/2.5e+13/config3/10x10x10</li> <li>0deg/dir110/2.5e+13/config3/10x10x10</li> <li>0deg/dir111/2.5e+13/config3/10x10x10</li> <li>0deg/dir234/2.5e+13/config3/10x10x10</li> </ol> <p>Folder paths for high density and high resolutions are:</p> <ol> <li>0deg/dir100/1e+14/config1/20x20x20</li> <li>0deg/dir110/1e+14/config1/20x20x20</li> <li>0deg/dir111/1e+14/config1/20x20x20</li> <li>0deg/dir234/1e+14/config1/20x20x20</li> </ol> <p>Each set of simulation has 2000 density field data files. Total files : 16000</p> <p>Please ensure that above paths are entered in the Jupyter notebook script file.</p> <p> </p> <p><strong>vtk.zip</strong></p> <p>VTK files for each simulation. Each simulation has 2000 files. </p> <p><strong>Dislocation_embeddings_pbc.ipynb</strong></p> <p>Jupyter notebook to generate dislocation embeddings for uploaded dataset.</p> <p> </p>
FIGURE 1. A. David W. Mitchell examining a twig for myxomycetes. B. Notebook pages. C. Taped matchboxes for specimens. D. Slide boxes. E. Mitchell notebooks. F in The type material in the David W. Mitchell myxomycete collection at the Royal Botanic Garden, Kew
FIGURE 1. A. David W. Mitchell examining a twig for myxomycetes. B. Notebook pages. C. Taped matchboxes for specimens. D. Slide boxes. E. Mitchell notebooks. F. Page of indexes in notebook.
FIGURE 9. Library trolley with all the separately boxed Lister notebooks. FIGURE 10 in Typification of the myxomycete taxa described by the Listers and preserved at the Natural History Museum, London (BM)
FIGURE 9. Library trolley with all the separately boxed Lister notebooks. FIGURE 10. Some of the metal and wooden herbarium cabinets. FIGURE 11. Hermetic wooden slide cabinets. FIGURE 12. One of the trays with slides from a slide cabinet.
Alevin notebook for Google Collab & backup Alevin output
<p>A Jupyter notebook containing the workflow presented in the GTN tutorial <a href="https://training.galaxyproject.org/training-material/topics/single-cell/tutorials/alevin-commandline/tutorial.html">Generating a single cell matrix using Alevin and combining datasets (bash + R)</a> that can be used in Google Collab. <br><br>Also, a folder of Alevin outputs resulting from running salmon alevin as shown in the tutorial.</p>
The code and data for paper: "Observing Fine-Grained Changes in Jupyter Notebooks During Development Time"
<div> <p>This package represents supplementary materials for the paper "Observing Fine-Grained Changes in Jupyter Notebooks During Development Time". Please refer to README in the archive for details.</p> </div>
Manually validated PageXML files for images in notebook "L'oddysée d'un garde civique"
<p>Transcription of a notebook of a Belgian member of paramilitary militia in World War I (70 pages in total). Transcription contains pages 4 to 65 in PageXML format, useful for training a handwritten text recognition model. The PageXML files were created by applying a non-public Transkribus model (The Text Titan I) on the images at https://europeana.transcribathon.eu/documents/story/?story=138144 and by manually validating the result.</p>
Facilitating Sensemaking in Computational Notebooks
Open the record for dataset details and reuse information.
Additional File 1 and Notebook Code for "Machine learning approaches for hospital acquired pressure injuries: a retrospective study of electronic medical records"
<p>Supplementary materials (Pressure_Injuries_Additional_File_1_final_double_blind.pdf) and Jupyter Notebook code (HAPI_Prediction_Script.pdf) developed as supplement for study "Machine learning approaches for hospital acquired pressure injuries: a retrospective study of electronic medical records".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.