Dense matrix multiply in four parallel models — hand-tiled OpenMP with roofline, Cannon's algorithm, hybrid MPI+OpenMP reaching 45 TFLOP/s on 16 nodes, and multi-GPU cuBLAS
-
Updated
Sep 16, 2026 - C++
Dense matrix multiply in four parallel models — hand-tiled OpenMP with roofline, Cannon's algorithm, hybrid MPI+OpenMP reaching 45 TFLOP/s on 16 nodes, and multi-GPU cuBLAS
To associate your repository with the cannons-algorithm topic, visit your repo's landing page and select "manage topics."