FLASH ATTENTION EXPLAINED
You know self attention mechanism right? Or multi head attention. But do you think we are using them now? Not the generic ones! It has…
FLASH ATTENTION EXPLAINED
You know self attention mechanism right? Or multi head attention. But do you think we are using them now? Not the generic ones! It has O(n²) complexity which is a bottleneck for GPUs.

Hierarchy of memory.
DRAM(dynamic ram) has big size but low bandwith (12.8GB/S,>1TB) . HBM (high bandwith memory) has mid size mid bandwith(1.5TB/S,40GB) . SRAM(static ram) has low size but big bandwith (19TB/S,20MB) . When you try to implement generic attention mechanism:
Q,K,V matrices stored in the HBM . Then you have to load them and compute them , S=Q.K^T(-2,-1) . Then these results (S) have go back to HBM. Then you have to apply softmax for EACH row. A=SOFTMAX(S)(dim=-1) then you have to write A to HBM . Final step load A and V by blocks from HBM then O=A.V compute it. Then write back to HBM. That is too much travel right? It comes and goes back again and again. That makes our inferences really slow. That is the reason why the FLASH ATTENTION is a very efficient way to compute attention mechanism.
Let’s remember the GPUs execution mechanism.
1-Loading data from memory to SRAM . Computations at SRAM. Write output back to HBM. It basicaly means you have to acces HBM repeatedly .Once for reading data,once for storing the results…

What flash attention does? It does not reduce the computational cost. It prevents going back to HBM again and again. Splitting tokens into batches then appropriately product them like CNNs convolution layer. Basically
Tile comes →Write SRAM →compute →delete tile →result appends O.
The tiles must be fit in to SRAM. And it never creates a NxN matrices.Thus our memory complexity goes O(N²) to O(N). It is very crucial for long length sequences.

Comparison of standart and flash attention mechanisms
Conclusion and Other Approaches
FlashAttention- 1 (2022): Introduced tiling and recomputation to minimize HBM access, reducing memory complexity.
FlashAttention- 2 (2023): Optimized work partitioning between GPU warps and reduced non matmul overhead, achieving close to 2x speedup over version 1.
FlashAttention- 3 (2024): Leverages Asynchronous execution and FP8 precision on Hopper architectures (H100 GPUs), pushing the limits of hardware utilization.
These are just the different versions of Flash Attention . Of course we have other applications like : paged attention,ring attention.
Also , do not forget to check out these links,you can get $100 Azure credit and GitHub copilot for students subscription for FREE! :
https://learn.microsoft.com/copilot/?wt.mc_id=studentamb_510772
https://azure.microsoft.com/free/students/?wt.mc_id=studentamb_510772
https://www.microsoft.com/microsoft-fabric?wt.mc_id=studentamb_510772
ERDEM YILMAZ
메타데이터
- post_id
- 8d09c8cf5965
- slug
- flash-attention-explained-8d09c8cf5965
- url
- https://medium.com/@erdemyy13/flash-attention-explained-8d09c8cf5965
- canonical_url
- https://medium.com/@erdemyy13/flash-attention-explained-8d09c8cf5965
- author_url
- https://medium.com/@erdemyy13
- status
- ok
- fetched_at
- 2026-06-09 15:37:30