← Back to list

I Spent $4M on H100s. I Got 60% of the Throughput I Paid For.

256 cards. Four bottlenecks I didn’t see in the proposal. The 40% I left on the table.

The Speed Engineer in Beyond Localhost · 2026-05-15 14:01 · 76 claps · 4.7 min read paywalled
#ai #aws #machine-learning #computer-science #cuda
Open on Medium ↗
Wiki topics: OPS · LLMOps & Inference ML · Machine Learning AI · AI · General EDU · Education & Learning ☁️ · DevOps & Cloud 🔬 · Science · General

I Spent $4M on H100s. I Got 60% of the Throughput I Paid For.

256 cards. Four bottlenecks I didn’t see in the proposal. The 40% I left on the table.

256 of these on the floor. We were getting 153 cards’ worth of work. The other 103 were on, drawing power, doing nothing.

256 of these on the floor. We were getting 153 cards’ worth of work. The other 103 were on, drawing power, doing nothing.

The proposal said our 256-H100 cluster would deliver throughput equivalent to 256 H100s. We deployed it. We were getting throughput equivalent to about 153 H100s. The other 103 cards were drawing power — about $32K a month in electricity alone — and producing nothing.

The 40% gap split four ways. Three of them were predictable in retrospect. One was a CUDA-allocator quirk I had never seen documented in a way that would have warned us before we wrote the check.

The Setup

256 H100 SXM5s across 32 nodes (8 per node). NVLink within node, NVIDIA Quantum-2 InfiniBand between nodes (400 Gbps per port). Training a roughly 70B-parameter dense transformer with mixed precision (BF16 weights, FP32 master), using FSDP with hybrid sharding.

At-spec throughput: 256 cards × ~750 sustained TFLOPS = roughly 192 PFLOPS. Initial sustained: ~115 PFLOPS. Gap: 40%. The proposal had not been wrong about peak. It had been silent about everything that prevents you from reaching peak.

Where 28% Went: NCCL Topology Was Wrong

The single biggest leak was NCCL collective overhead at the wrong topology. We were using ring-allreduce by default, which is what NCCL picks if you don’t tell it otherwise on a heterogeneous fabric.

Ring is fine when intra-node and inter-node bandwidth are similar. Our intra-node NVLink was 900 GB/s; our inter-node InfiniBand was 50 GB/s effective. A ring algorithm spends most of its wall-clock time in the slow segment. Tree-allreduce uses the topology — reduces within-node first over fast NVLink, then between nodes over the slower fabric, exploiting the bandwidth hierarchy.

# we'd been running with the default. switching to tree-allreduce + tuning the threshold
# moved sustained throughput from 115 to 168 PFLOPS.
export NCCL_ALGO=Tree
export NCCL_PROTO=Simple
export NCCL_MIN_NCHANNELS=8
export NCCL_MAX_NCHANNELS=16
# also: set the buffer to match our gradient size class
export NCCL_BUFFSIZE=8388608

The improvement was 28 percentage points. We had paid for the cluster for two months before noticing. Roughly $640K of compute spent on the wrong NCCL config.

Where 8% Went: CPU-Bound Data Loading

The next leak was on the CPU side. PyTorch’s DataLoader was choking. We had 8 dataloader workers per process; each worker decompressing image-text pairs from a webdataset. The 8 workers contended for the same memory bandwidth, and the host CPUs were Sapphire Rapids with mediocre memory bandwidth per socket.

The GPU was idle for ~8% of every step waiting for the next batch. We saw it on nvidia-smi's GPU Utilization averaging 92%, which we'd celebrated until we ran dcgmi dmon and saw the GPU was actually idle a measurable fraction of every step.

Two fixes. Pre-decode the dataset to packed BF16 tensors stored on local NVMe — removed the decompression cost from the hot path. CPU pressure dropped dramatically. Then use pinned memory + non_blocking=True on the host-to-device copy, so the next batch transfer overlaps with the previous step's compute.

GPU idle waiting on data dropped to under 1%. 8 percentage points back.

Where 3% Went: Tensor Parallelism Was Off-Tier

At our scale we’d configured tensor_parallel_size=8 (intra-node) and pipeline_parallel_size=4. The reasoning was "TP within node uses NVLink." Correct intuition. Wrong constant.

The model’s hidden dimension was 8192. With TP=8, each card held a tensor with the partition dimension of 1024. The matmul kernels in cuBLAS for hidden dimension 1024 are less efficient than those tuned for 2048 or 4096 — there’s a sweet spot where the tensor cores hit best utilization. We were below it.

Dropping to tensor_parallel_size=4 (each card held 2048) and increasing pipeline_parallel_size=8 rebalanced the work. The pipeline depth cost us a tiny bit of bubble overhead but the matmul efficiency win was bigger. 3 percentage points.

Where 1% Went: GPU Memory Fragmentation

The smallest leak was the most surprising. Our training run had highly variable activation memory because of dynamic-length sequences. Over an hour, the CUDA caching allocator’s free list would fragment — allocations of size 47 MB couldn’t find a contiguous 47 MB hole, even though total free memory was several GB.

The allocator would fall back to cudaMalloc, which synchronizes with the device. Synchronization meant the GPU stalled for a millisecond on every fragmented allocation. With ~600 allocations per step, the stall added up to a measurable fraction of step time.

# the env var nobody told us about until production was already fragmenting.
# expandable_segments dramatically reduces fragmentation; max_split_size limits
# how aggressively the allocator splits large blocks.
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:512

1 percentage point back. Smaller than the others, but the one I had never seen documented in a way I’d have caught during planning.

What 60% to 96% Bought Us

After all four fixes, sustained throughput was about 184 PFLOPS — 96% of theoretical, roughly the practical ceiling. The training run that had been projected at 19 days finished in 12. We saved about $1.4M in compute against the original projection.

The Lesson That Cost $1.4M to Learn

The proposal told me what the hardware could do. It did not tell me what my software would let it do. Those are two different numbers, and the gap is where the money goes.

Every AI infra team I have consulted with since has the same gap. They buy the cards, deploy the cluster, run the training, look at the loss curve, ship the model. The “is the cluster delivering its specced throughput” question never gets asked because the loss curve is the user-visible metric and the model trained eventually and the bill got paid.

Your dashboard says GPU utilization is 92%. Run dcgmi dmon and look at SM occupancy. The number you see is the number you actually paid for. Everything between that and 100% is the hardware doing nothing while drawing power.

Three days with nvprof/nsys and nvidia-bug-report.sh gets you most of the gap. That is the cheapest 30% throughput win in modern infrastructure.

Where the Industry Is Today

Every team I’ve reviewed in the last six months has been running at 55–75% of theoretical. The ones at 75% have done one or two of the four fixes. Nobody I’ve seen has been at 90%+ without a dedicated performance engineer who has done all four plus another half-dozen smaller things.

If you’re running an AI training cluster and you can’t tell me your sustained PFLOPS as a percentage of theoretical, you don’t know what you bought. The cards are the same. The difference is the three days somebody spent profiling.

Enjoyed the read? Let’s stay connected!

  • 🚀 Follow The Speed Engineer for more Rust, Go and high-performance engineering stories.
  • 💡 Like this article? Follow for daily speed-engineering benchmarks and tactics.
  • ⚡ Stay ahead in Rust and Go — follow for a fresh article every morning & night.

Your support means the world and helps me create more content you’ll love. ❤️


메타데이터
post_id
81ece16d7ba8
slug
i-spent-4m-on-h100s-i-got-60-of-the-throughput-i-paid-for-81ece16d7ba8
url
https://medium.com/beyond-localhost/i-spent-4m-on-h100s-i-got-60-of-the-throughput-i-paid-for-81ece16d7ba8
canonical_url
https://medium.com/beyond-localhost/i-spent-4m-on-h100s-i-got-60-of-the-throughput-i-paid-for-81ece16d7ba8
author_url
https://medium.com/@speed_enginner
status
ok
fetched_at
2026-06-09 15:37:30