← Back to list

FLASH ATTENTION EXPLAINED

You know self attention mechanism right? Or multi head attention. But do you think we are using them now? Not the generic ones! It has…

Erdem · 2026-05-08 09:36 · 1 claps · 2.1 min read
#flash-attention #attention #transformers #ai #artificial-intelligence
Open on Medium ↗
Wiki topics: AI · AI · General

FLASH ATTENTION EXPLAINED

You know self attention mechanism right? Or multi head attention. But do you think we are using them now? Not the generic ones! It has O(n²) complexity which is a bottleneck for GPUs.

Hierarchy of memory.

Hierarchy of memory.

DRAM(dynamic ram) has big size but low bandwith (12.8GB/S,>1TB) . HBM (high bandwith memory) has mid size mid bandwith(1.5TB/S,40GB) . SRAM(static ram) has low size but big bandwith (19TB/S,20MB) . When you try to implement generic attention mechanism:

Q,K,V matrices stored in the HBM . Then you have to load them and compute them , S=Q.K^T(-2,-1) . Then these results (S) have go back to HBM. Then you have to apply softmax for EACH row. A=SOFTMAX(S)(dim=-1) then you have to write A to HBM . Final step load A and V by blocks from HBM then O=A.V compute it. Then write back to HBM. That is too much travel right? It comes and goes back again and again. That makes our inferences really slow. That is the reason why the FLASH ATTENTION is a very efficient way to compute attention mechanism.

Let’s remember the GPUs execution mechanism.

1-Loading data from memory to SRAM . Computations at SRAM. Write output back to HBM. It basicaly means you have to acces HBM repeatedly .Once for reading data,once for storing the results…

What flash attention does? It does not reduce the computational cost. It prevents going back to HBM again and again. Splitting tokens into batches then appropriately product them like CNNs convolution layer. Basically

Tile comes →Write SRAM →compute →delete tile →result appends O.

The tiles must be fit in to SRAM. And it never creates a NxN matrices.Thus our memory complexity goes O(N²) to O(N). It is very crucial for long length sequences.

Comparison of standart and flash attention mechanisms

Comparison of standart and flash attention mechanisms

Conclusion and Other Approaches

FlashAttention- 1 (2022): Introduced tiling and recomputation to minimize HBM access, reducing memory complexity.

FlashAttention- 2 (2023): Optimized work partitioning between GPU warps and reduced non matmul overhead, achieving close to 2x speedup over version 1.

FlashAttention- 3 (2024): Leverages Asynchronous execution and FP8 precision on Hopper architectures (H100 GPUs), pushing the limits of hardware utilization.

These are just the different versions of Flash Attention . Of course we have other applications like : paged attention,ring attention.

Also , do not forget to check out these links,you can get $100 Azure credit and GitHub copilot for students subscription for FREE! :

https://learn.microsoft.com/copilot/?wt.mc_id=studentamb_510772

https://azure.microsoft.com/free/students/?wt.mc_id=studentamb_510772

https://www.microsoft.com/microsoft-fabric?wt.mc_id=studentamb_510772

ERDEM YILMAZ


메타데이터
post_id
8d09c8cf5965
slug
flash-attention-explained-8d09c8cf5965
url
https://medium.com/@erdemyy13/flash-attention-explained-8d09c8cf5965
canonical_url
https://medium.com/@erdemyy13/flash-attention-explained-8d09c8cf5965
author_url
https://medium.com/@erdemyy13
status
ok
fetched_at
2026-06-09 15:37:30