?
Efficiency of Machine Learning Tasks on HPC Devices
Accurate benchmarking is critical for selecting computing architectures optimized for machine learning (ML) tasks. Conventional benchmarks such as High-Performance Linpack (HPL) and High Performance Conjugate Gradients (HPCG) often fail to capture the diversity and complexity of modern ML workloads. This study investigates the correlation between hardware parameters (e.g., processor architecture, cache size, frequency) and ML performance across various devices, including CPUs and accelerators. By studying a wide range of ML tasks, we identify key performance bottlenecks and explore whether a correlation-based approach can guide the selection of optimal hardware for an entire class of ML tasks. Our findings offer practical recommendations for future computing architectures and advance the efficient use of high-performance computing for diverse ML applications.