← Back to list

Building Lossless Network: NADDOD’s Solutions for AI Inference Workloads

As AI inference scales from single-node testing to large-scale production deployment, the optimization for AI inference workloads continues…

NADDOD · 2026-05-27 07:49 · 0 claps · 8.8 min read
#ai #networking #cloud-infrastructure
Open on Medium ↗
Wiki topics: OPS · LLMOps & Inference AI · AI · General

Building Lossless Network: NADDOD’s Solutions for AI Inference Workloads

As AI inference scales from single-node testing to large-scale production deployment, the optimization for AI inference workloads continues to shift. Inference workloads are typically oriented toward real-time requests — users expect immediate responses, not background task completion. As a result, network congestion, packet loss, retransmissions, and latency jitter are rapidly amplified at the service-experience level.

NADDOD’s AI inference network solution does not focus on improving any single parameter in isolation. Instead, it takes a layered design approach across the compute network, management network, and storage network — combining InfiniBand, RoCE, 800G/1.6T optical interconnects, 51.2T switches, and end-to-end test validation to help inference clusters build a more stable lossless network foundation.

Why Does AI Inference Require a Lossless Network?

AI inference differs from AI training — and so do their networking requirements. A single inference request may traverse multiple GPU nodes, model shards, KV Cache, storage systems, and front-end services. Significant congestion on any one segment of this path can increase overall response time.

In distributed inference, multiple GPUs must often collaborate on the same model task. When packet loss or retransmissions occur, data arrival times become unpredictable, causing some GPUs to stall while waiting for other nodes to complete synchronization — ultimately degrading GPU utilization. For high-concurrency inference workloads, this stalling accumulates continuously, manifesting as elevated tail latency, reduced throughput, and fluctuating quality of service.

The value of a lossless network lies in reducing packet loss and unnecessary retransmissions through more effective congestion control, careful link planning, and robust physical-layer reliability. It does not mean a network will never experience pressure; rather, it means the network can maintain more predictable transmission performance even under high load.

NADDOD Lossless Network Architecture

NADDOD structures its AI inference network solution around three distinct layers, each with specific responsibilities:

1. Compute Network Layer

The compute network layer carries GPU-to-GPU high-speed communication and is the most performance-sensitive part of an AI inference cluster. For inference workloads, the compute network must address three key challenges: whether there is sufficient bandwidth between nodes, whether the path is stable, and whether packet loss and retransmissions can be minimized during congestion.

NADDOD’s compute network solution supports either InfiniBand or RoCE architecture depending on cluster scale. InfiniBand is better suited for inference scenarios that demand low latency and deterministic communication; RoCE is appropriate for projects that aim to build a high-performance inference network within an Ethernet ecosystem. Neither option is inherently superior — the right choice depends on the customer’s cluster scale, budget, existing network infrastructure, and operational capabilities.

2. Management Network Layer

The management network layer does not directly carry GPU compute data, but it is critical to the long-term stable operation of an inference cluster. When management traffic shares the same fabric as compute traffic, it can impact core business links during peak periods and increase the difficulty of troubleshooting. NADDOD builds an isolated management network using devices such as the N6300–48Y8C and N6300–32C, separating operations traffic from compute traffic. The key benefit is that even when the compute network is under high load, the operations team retains full visibility into device state.

3. Storage Network Layer

AI Inference workloads do not rely solely on GPU compute. Model weight loading, vector database access, cache reads, log writes, and result callbacks all place demands on the storage network. In large-model inference scenarios in particular, model files are large and access frequency is high — if storage network bandwidth is insufficient, GPUs may stall waiting for data, dragging down overall inference efficiency.

NADDOD offers flexible bandwidth options on the storage network layer — including 400G, 200G, and 100G — allowing customers to configure capacity based on model size, request concurrency, and storage architecture. Compared to simply stacking high-speed ports, a three-layer separation of compute, management, and storage provides cleaner bandwidth planning and simplifies future capacity expansion.

NADDOD Typical AI Inference Lossless Network Architecture

NADDOD InfiniBand Network Solutions

InfiniBand XDR Solution for Rubin & Blackwell

NADDOD employs the NVIDIA Quantum-X800 Q3400-RA switch as both the Spine (×32) and Leaf (×18) switches, forming a 2-tier Fat-Tree topology.

This solution supports clusters of up to 2,304 nodes composed of 32 NVIDIA DGX Rubin NVL72 / GB300 compute nodes, with per-node XDR 800G connectivity.

To accommodate various link distances and deployment scenarios, the following interconnect combinations are available:

Switch-to-Switch:

Switch-to-Server:

This solution is designed for high-performance inference and large-scale GPU collective communication. Its primary focus is addressing path consistency and bandwidth scaling in large GPU clusters. The 2-tier Fat-Tree topology maintains relatively balanced communication paths between nodes, reducing the performance fluctuations caused by individual link overload. For tensor parallelism, pipeline parallelism, or high-concurrency model serving, this architecture helps improve cross-node communication efficiency.

InfiniBand NDR Solution for Blackwell & Hopper

NADDOD employs NVIDIA Quantum-2 NDR switches as both the Spine (×16) and Leaf (multiple groups) switches, forming a 2-tier Fat-Tree topology.

This solution supports clusters of up to 1,024 nodes composed of 256 NVIDIA DGX H200 / B200 compute nodes, with flexible per-node connectivity of either 1×400G or 2×400G.

Switch-to-Switch:

Switch-to-Server:

This architecture is suited for high-performance inference clusters targeting the Blackwell and Hopper platforms — particularly deployments requiring stable 400G per-node connectivity, low-latency communication, and strong synchronization efficiency. The 2-tier Fat-Tree topology creates multi-path interconnects between Spine and Leaf layers, reducing the impact of single-path congestion on overall inference task performance. For inference clusters evolving from 400G high-speed networks, this solution strikes a stable balance among performance, deployment complexity, and future scalability.

NADDOD RoCE Networking Solutions

1.6T RoCE Solution for Rubin

NADDOD employs the SN6600-LD switch as both the Spine (×18) and Leaf (×32) switches, forming a 2-tier Fat-Tree topology.

This solution targets the Rubin NVL72 inference cluster, supporting up to 2,304 nodes with 800G-class RoCE connectivity per node. Spine-to-Leaf uplinks use RoCE 2×800G, while Leaf-to-compute-node downlinks provide RoCE 800G access.

Switch-to-Switch:

Switch-to-Server:

This solution is designed for customers who wish to build large-scale AI inference networks on top of an Ethernet ecosystem. Rather than simply increasing port speeds, its focus is on leveraging a high-bandwidth RoCE fabric to support horizontal scaling of high-density GPU systems such as the Rubin NVL72. For high-concurrency inference, model-shard communication, and cross-node GPU collaborative computing, this architecture helps alleviate bandwidth bottlenecks, link congestion, and inter-node communication imbalance — while retaining Ethernet’s flexibility in operational tooling, automation, and ecosystem compatibility.

800G RoCE Solution for Rubin & Blackwell

NADDOD employs the N9500–64OC switch as both the Spine (×16) and Leaf (×32) switches, forming a 2-tier Fat-Tree topology.

This solution supports clusters of up to 2,048 nodes composed of 256 DGX Rubin NVL8 / B300 compute nodes, with RoCE 400G per-node connectivity. Spine-to-Leaf interconnects use RoCE 2×400G; Leaf-to-node downlinks use RoCE 400G access. Compatible optics include OSFP-800G-2×SR4 and OSFP-800G-2×DR4 modules, covering multimode short-reach and single-mode medium-reach scenarios.

Switch-to-Switch:

Switch-to-Server:

This architecture is well suited for medium-to-large AI inference clusters, especially projects that need to balance 400G per-node access with 800G high-density uplink ports. It primarily addresses the challenges of uplink bandwidth, port density, and cabling complexity as GPU node counts grow. Through a 2-tier Fat-Tree topology, the cluster can maintain clear east-west communication paths at scale, reducing the impact of link overload on inference throughput.

Back-End Networking Solution

NADDOD employs the N9520–64OC switch as both the Spine (×16) and Leaf (×32) switches, forming a 2-tier Fat-Tree back-end compute network.

This solution supports clusters of up to 2,048 nodes composed of 256 DGX B300 compute nodes, with RoCE 400G per-node connectivity. The N9520–64OC provides 64×800G OSFP ports and 51.2 Tbps switching capacity, enabling a high-bandwidth, low-latency RoCE back-end compute fabric.

This solution is primarily designed for high-speed inter-GPU communication, addressing compute-side bandwidth aggregation, east-west traffic forwarding, and lossless Ethernet transmission in inference clusters. In large-model inference, multiple GPUs must exchange intermediate data at high frequency; if back-end network bandwidth is insufficient or paths are unbalanced, GPU stalls and elevated tail latency can result. Through high-density 800G switches and a 2-tier Fat-Tree topology, NADDOD’s back-end compute network provides a more stable communication foundation for RoCE inference clusters.

Front-End Networking Solution

NADDOD’s front-end network solution primarily carries storage, in-band Ethernet, business access, and operations management traffic — forming a layered separation from the back-end GPU compute network.

Unlike the back-end compute network, the front-end network does not handle large-scale GPU synchronization communication directly. Instead, it provides data access, system management, monitoring, maintenance, and external connectivity for inference workloads.

This architecture primarily addresses traffic isolation and operational visibility. In AI inference clusters, model loading, log writes, storage access, and device management all generate sustained traffic flows. If these flows share the same network plane as GPU compute traffic, congestion and troubleshooting complexity increase significantly. Through an independent front-end network design, NADDOD helps customers isolate storage, management, and business access traffic from the back-end compute fabric — enabling the inference compute network to operate in a more focused communication environment while improving daily operations efficiency and fault localization.

Why Pre-Deployment Testing and Validation Matter for AI Inference Networks

The reliability of an AI inference network cannot be determined from theoretical specifications alone. Optical modules, cables, switch ports, firmware versions, and link distances all affect real-world deployment outcomes. Particularly in large-scale inference clusters, a single unstable link can cause localized retransmissions, port errors, or service jitter.

NADDOD’s capabilities include multi-version firmware testing, full-port live connectivity testing, and real-time Bit Error Rate (BER) testing, supported by a quality control system and project-level technical services. For inference networks, this type of validation helps customers identify compatibility and link stability issues before go-live, reducing post-deployment troubleshooting costs.

Conclusion: Lossless Networking Is the Foundation for Scalable AI Inference

Scaling AI inference is not simply a matter of adding more GPUs. Only when the network can stably carry GPU communication, model loading, storage access, and management traffic will an inference cluster’s compute capacity reliably translate into consistent business capability.

The value of NADDOD’s solution lies in combining InfiniBand, RoCE, high-speed switches, 800G/1.6T optical interconnects, management networking, and test validation into a comprehensive network architecture. For enterprises deploying or expanding AI inference clusters, a lossless network is not a single technology point — it is a systems engineering discipline centered on low latency, high throughput, reliable transmission, and long-term operational manageability.


메타데이터
post_id
f4c1d50eb889
slug
building-lossless-network-naddods-solutions-for-ai-inference-workloads-f4c1d50eb889
url
https://medium.com/@naddod/building-lossless-network-naddods-solutions-for-ai-inference-workloads-f4c1d50eb889
canonical_url
https://medium.com/@naddod/building-lossless-network-naddods-solutions-for-ai-inference-workloads-f4c1d50eb889
author_url
https://medium.com/@naddod
status
ok
fetched_at
2026-06-11 17:15:47