Snapshot Testing Dataset - Repositories using Jest
<p>This is a dataset of GitHub repositories that were tagged with Jest, for JavaScript and TypeScript languages, that used Snapshot Testing. Information on all repositories is available in the file "<a href="https://zenodo.org/api/files/68da2ad8-feaa-4a41-af84-278194516f0a/0_Snapshot%20Testing%20Dataset.xlsx?versionId=a2423c05-1628-4661-85cc-96d615f4ddf2">0_Snapshot Testing Dataset.xlsx</a>" (named to be the very first file). Most files represent the repository packed in targz format as "<user>_<repository_name>.tar.gz". We split large repositories (>50MB) using the "split" command on Unix (use cat to rejoin them). </p> <p>In total there are 686 repositories. We collected only public repositories that were tagged with the Jest keyword, for JavaScript and TypeScript, had at least 1 star, and at least 1 snapshot file. The spreadsheet data was collected on July 13, 2022.</p> <p>We also have all scripts used to gather this data. Here, "<a href="https://zenodo.org/api/files/00168583-5071-4ce2-8dc6-2d44a5f1bf77/python_scripts.zip">python_scripts.zip</a>" has all python scripts to find repositories based on queries and save their attributes, and "<a href="https://zenodo.org/api/files/00168583-5071-4ce2-8dc6-2d44a5f1bf77/node_and_shell_scripts.zip">node_and_shell_scripts.zip</a>" contain the node and shell scripts to download a tarball of the repository. Therefore you should first use the python scripts to collect repositories & their attribute, and later use the node & shell to download a copy of the repositories. Moreover, inside each script folder/zip there is a Readme file with instructions and examples.</p> <p>Our GitHub repository is an exact copy of this dataset <<a href="https://github.com/hscrocha/SnapshotTestingDataset">https://github.com/hscrocha/SnapshotTestingDataset></a>, but it is much better organized into folders and the README files for the scripts will be nicely displayed on it.</p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 20
- Reuse readiness
- 8
- Engagement
- 4