Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
10
datasets available to search
ShareScore release 0.9.0
Dataset results
10 results for “kaggle”
DistilKaggle: a distilled dataset of Kaggle Jupyter notebooks
<h2><strong>Overview</strong></h2> <p>DistilKaggle is a curated dataset extracted from Kaggle Jupyter notebooks spanning from September 2015 to October 2023. This dataset is a distilled version derived from the download of over 300GB of Kaggle kernels, focusing on essential data for research purposes. The dataset exclusively comprises publicly available Python Jupyter notebooks from Kaggle. The essential information for retrieving the data needed to download the dataset is obtained from the MetaKaggle dataset provided by Kaggle.</p> <h2><strong>Contents</strong></h2> <p>The DistilKaggle dataset consists of three main CSV files:</p> <p><strong>code.csv:</strong> Contains over 12 million rows of code cells extracted from the Kaggle kernels. Each row is identified by the kernel's ID and cell index for reproducibility.</p> <p><strong>markdown.csv:</strong> Includes over 5 million rows of markdown cells extracted from Kaggle kernels. Similar to <strong>code.csv</strong>, each row is identified by the kernel's ID and cell index.</p> <p><strong>notebook_metrics.csv:</strong> This file provides notebook features described in the accompanying paper released with this dataset. It includes metrics for over 517,000 Python notebooks.</p> <h2><strong>Directory Structure</strong></h2> <p>The <strong>kernels</strong> directory is organized based on Kaggle's Performance Tiers (PTs), a ranking system in Kaggle that classifies users. The structure includes PT-specific directories, each containing user ids that belong to this PT, download logs, and the essential data needed for downloading the notebooks.</p> <p>The <strong>utility</strong> directory contains two important files:</p> <p><strong>aggregate_data.py:</strong> A Python script for aggregating data from different PTs into the mentioned CSV files.</p> <p><strong>application.ipynb:</strong> A Jupyter notebook serving as a simple example application using the metrics dataframe. It demonstrates predicting the PT of the author based on notebook metrics.</p> <p><strong>DistilKaggle.tar.gz: </strong>It is just the compressed version of the whole dataset. If you downloaded all of the other files independently already, there is no need to download this file.</p> <h2><strong>Usage</strong></h2> <p>Researchers can leverage this distilled dataset for various analyses without dealing with the bulk of the original 300GB dataset. For access to the raw, unprocessed Kaggle kernels, researchers can request the dataset directly.</p> <h2><strong>Note</strong></h2> <p>The original dataset of Kaggle kernels is substantial, exceeding 300GB, making it impractical for direct upload to Zenodo. Researchers interested in the full dataset can contact the dataset maintainers for access.</p> <h2><strong>Citation</strong></h2> <p>If you use this dataset in your research, please cite the accompanying paper or provide appropriate acknowledgment as outlined in the documentation.</p> <p>If you have any questions regarding the dataset, don't hesitate to contact me at <a href="mailto:mohammad.abolnejadian@gmail.com">mohammad.abolnejadian@gmail.com</a></p> <p>Thank you for using DistilKaggle!</p>
Derived Protein Sequence Data From Kaggle
<p>It is the dataset containing protein sequence and its associated go-terms, used for priliminary training process</p>
Solution #4 for Predicting Molecular Properties Kaggle Competition
<p>Code and additional data for solution #4 in <a href="https://www.kaggle.com/c/champs-scalar-coupling/overview">Predicting Molecular Properties</a> competition, described in <a href="https://www.kaggle.com/c/champs-scalar-coupling/discussion/106534#latest-613848">#4 Solution [Hyperspatial Engineers]</a>.</p>
dogs_vs_cats_subset_kaggle
<p>This is a subset of the well known image classification dataset for cats and dogs made by kaggle.</p> <p>https://www.microsoft.com/en-us/download/details.aspx?id=54765</p> <p>This dataset contains 1000 training images (500 cats & 500 dogs), 200 validation images (100 cats/100 dogs) and 100 unlabelled test images.</p> <p>There are other similar dataset already on zenodo like this https://zenodo.org/doi/10.5281/zenodo.5226944, but that dataset is not balanced even though it claims to be.</p> <p> </p> <p> </p>
dataset of Kaggle Notebooks
<p>These are some dataset of Kaggle Notebooks</p>
kaggle-whats-cooking
<h3><strong>Overview</strong></h3><p>Hypergraph where nodes are food ingredients, hyperedges are recipes made from combining multiple ingredients, and categories indicate cuisine (e.g., "Southern-US", "Indian", "Spanish"). Some summary statistics of the hypergraph are:</p><ul><li>Number of nodes: 6,714</li><li>Number of hyperedges: 39,774</li><li>Number of edge label categories: 20</li><li>Maximum hyperedge size: 65</li></ul><h4><strong>Source of original data</strong></h4><p>Sources:</p><ul><li><a href="https://drive.google.com/open?id=16BTDEv3AC9l81FeU2US5nG30_CGsRTqk">cat-edge-Cooking.zip</a></li><li><a href="https://www.kaggle.com/c/whats-cooking">"What's Cooking?" Kaggle competition</a></li><li><a href="https://www.yummly.com/">Yummly</a></li></ul><h4><strong>References</strong></h4><p>If you use this data, please cite the following paper:</p><ul><li><a href="https://doi.org/10.1145/3366423.3380152">Clustering in graphs and hypergraphs with categorical edge labels</a>. Ilya Amburg, Nate Veldt, and Austin R. Benson. Proceedings of the Web Conference (WWW), 2020.</li></ul>
Kaggle Data Reuse Community Survey Feb 11 2021
<p>Kaggle Data Reuse Community Survey Dataset</p> <p>Date: Feb 2021</p> <p> </p>
Subset of Kaggle Flowers dataset annotated for flower type and flower color (AHM 2024 MTL example data)
Open the record for dataset details and reuse information.
Kaggle notebook names of top 1100 'hottest' notebooks on May 01 2020
<p>A (sorted) list of the 1100 'hottest' notebooks from Kaggle.com on May 01 2020. The corresponding notebooks themselves can be downloaded with the Kaggle API.</p>
Ramen Ratings (Kaggle)
<p>from kaggle <a href="https://www.kaggle.com/residentmario/ramen-ratings">https://www.kaggle.com/residentmario/ramen-ratings</a></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.