To write a CUDA Matmul Kernel That Reaches 97% of cuBLAS Performance
August 10, 2026
Introduction
In this article, I will walk you through how to write a CUDA matrix multiplication naive kernel and optimizing it to reach ~97% of SGEMM cuBLAS kernel.
Learning CUDA has been an incredible and exciting journey. While learning CUDA has become easier over the past few years, most resources remain difficult to digest.
The intended audience for this article is someone who knows absolutely nothing about CUDA or GPU optimization. My goal is to bring them up to speed and teach them how to write a performant CUDA kernels from scratch.
This article teaches you how to optimize a matrix multiplication kernel; it is the single most important operation within modern deep learning. Therefore there is no better way to learn about CUDA than to write an optimized version of matrix multiplication that performs near performance of cuBLAS.
I've previously thought that writing performant CUDA kernels was about coming up with most very clever algorithms to squeeze out all of the performance of a GPU. I was very wrong. While learning how to optimize a the kernel myself, I've noticed the majority of optimization comes from understanding the GPU memory hierachy at the most fundamental level and learning how to feed data into computation units efficiently without wasting any resources.
Its all about the memory hierachy and the movement of data with the CUDA software. The idea is simple, but the execution can be difficult. Within this article, we will iteratively walk through the hardware components of a CUDA architecture all the way to optimizing a CUDA kernel.
Below is the table of contents, feel free to skip around although reading this sequentially is advised.
To write a CUDA Matmul Kernel That Reaches 97% of cuBLAS Performance
- From CUDA Hardware Architecture to CUDA Kernels
- The Naive Kernel: Establishing a Baseline
- Global Memory Coalescing
- Flattening the Block
- Shared Memory Tiling
- 1D Register Tiling: One Thread, TM Outputs
- 2D Register Tiling: The Outer Product
- Vectorized Memory Access: 128-bit Loads and Stores
- Warp Tiling: A Third Level of Tiling
- Autotuning: Searching the Tile-Shape Space
- Padded Shared Memory: Eliminating Bank Conflicts
- Double Buffering: Software Pipelining the K-Loop
- The Final Kernel: Retuning Against the Pipeline
This writing was inspired by blogs such as Simons writing on this exact topic1. Abhik's2 and Robert's3 on this topic is what cemented these concepts for me. The reset of my knowledge gap was filed with reading PMPP 4, Modals' GPU Glossary5, and a lot of tokens.