Exploring Data at Scale with Arkouda: A Practical Introduction to Scalable Data Science
<p><span>Data scientists can be thought of as modern-day explorers, venturing into the vast unknown of information. However, this exciting journey is not without its hurdles. One of the biggest challenges they face is the sheer immensity of data they encounter. Modern datasets cannot fit in laptop memory, containing terabytes or even petabytes of information. Working with such massive data requires specialized tools and techniques to extract meaningful insights. As data sets are growing ever larger, data science demands interactivity, where scientists can learn while working with the data. At the same time, data science demands scalability, where scientists are able to work with data sets in their entirety. Data scientists have naturally been drawn to Python as it provides interactivity through its read, evaluate, print loop and performance through its utilization of libraries written in other languages, like C and Fortran. These libraries typically are not designed for HPC and run into problems when attempting to scale. The gap that Arkouda fills in the data science landscape is a library that is both interactive, providing a familiar Python API, and scalable, leveraging a scalable Chapel server in the backend. Arkouda is a framework for scalable Python packages for interactive data science and has applications ranging from oceanography to net flow analysis.</span></p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 0