← Back to Blog
For: AI Engineers, ML Engineers, Platform Engineers, AI Systems Architects

Inside the LLM Inference Engine: How Serving Works in 2026

Prefill, decode, paged KV cache, continuous batching and chunked prefill are on by default now. The real decision is which optimizations to add on top, and on what measured condition.

Updated
#llm-inference#llm-serving#inference-engine#inference-architecture#kv-cache#pagedattention#continuous-batching#chunked-prefill#prefix-caching#speculative-decoding

State of the field as of 2026-10-01.

Rewritten 2 October 2026. The first version of this article, from February 2025, surveyed six serving frameworks and listed speculative decoding and mixture-of-experts serving as future directions. Since then Hugging Face archived Text Generation Inference (TGI), the techniques it called future became engine defaults, and a new class of hardware shipped. This version explains how the engine works, concept by concept, with a diagram for each, and says which optimizations to leave alone and which to turn on.

Here is a deployment that looks fine on paper. Llama 3.3 70B, quantized to 8-bit floating point (FP8), needs about 73 GB for its weights. An NVIDIA H100 has 80 GB. The team starts vLLM with an 8,192-token context, and it refuses to start: ValueError: No available memory for the cache blocks. They raise --gpu-memory-utilization, as the error suggests, and the next error says the KV cache needed for one 8,192-token request is larger than the memory left for it. So they lower --max-model-len until the server boots. The startup log now reports a maximum concurrency barely above one request, and under real traffic the preemption counter climbs while throughput collapses. The model fits. The workload does not.

That failure is about memory, not compute, and almost every important idea inside a modern LLM inference engine is about the same thing. This article walks through those ideas in the order a request meets them.

What this article decides, and what it leaves out

It answers one question: once you run a modern LLM inference engine, which of its optimizations should you rely on, which should you add, and what measured condition justifies each addition?

It explains each mechanism well enough that you can predict what a setting will do before you change it. Every concept gets a diagram.

Out of scope: which engine to run. That decision has its own guide, Choosing an LLM Inference Framework in 2026: vLLM by Default, and this article agrees with it. Also out of scope are training, fine-tuning, and model-side compression such as pruning and distillation, which change the model rather than how it is served. So is choosing which model to run, which is a systems decision of its own.

Leave the scheduling defaults on, size the key-value (KV) cache before anything else, use FP8 weights on GPUs that run FP8 natively, run one replica per GPU when the model fits on one, and use tensor parallelism inside one NVLink domain when it does not. Add speculative decoding, disaggregated serving, KV offload, or 4-bit weights only when a condition in the deviation table below is true for your traffic, and keep a change only if your own measurements move.

In vLLM v0.30.0 (22 September 2026), the defaults you are keeping are continuous batching, a paged KV cache, chunked prefill, automatic prefix caching, and CUDA graphs with torch.compile at optimization level 2. SGLang v0.5.20 ships equivalents, with one difference in chunked prefill noted below. You do not have to turn these on. You have to avoid turning them off by accident, and you have to understand them well enough to read the metrics they produce.

How an LLM inference engine works, and why each default exists

Prefill and decode: two phases with different bottlenecks

Every request runs in two phases. Prefill processes the whole prompt in one forward pass. Because every prompt token goes through the layers together, the work is large matrix-matrix multiplication, and the limit is the GPU's arithmetic throughput. Prefill writes one key vector and one value vector per token, per layer, into the KV cache, and it produces the first output token. Its duration sets the time to first token (TTFT).

Decode then produces one token per sequence per forward pass, and repeats until the model emits an end-of-sequence token or reaches max_tokens. Each decode step reads every weight in the model and the whole KV cache of every sequence in the batch, to produce only one new token per sequence. How long each of those tokens takes is the time per output token (TPOT).

d2
direction: down

prompt: "Prompt: 2,000 tokens" {
  style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"}
}

prefill: "PREFILL\nOne forward pass over all 2,000 tokens\nLarge matrix-matrix multiplies\nLimited by GPU compute" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}

kv: "KV cache\nOne key and one value vector\nper token, per layer" {
  shape: cylinder
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}

decode: "DECODE STEP\nOne forward pass, one new token per sequence\nReads every weight and the whole KV cache\nLimited by memory bandwidth" {
  style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}
}

out: "Token streamed to the client" {
  style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}
}

prompt -> prefill
prefill -> kv: "writes 2,000 entries"
prefill -> out: "first token\n(sets TTFT)"
kv -> decode: "read in full\nevery step"
decode -> kv: "appends 1 entry"
decode -> out: "each later token\n(sets TPOT)"
decode -> decode: "repeat until end-of-sequence\nor max_tokens"

Decode is the expensive phase to serve, and the diagram shows why. Moving 73 GB of weights from memory takes far longer than the arithmetic on one token, so the GPU's compute units mostly wait. Databricks put it plainly in 2023: memory bandwidth "is a better predictor of speed of token generation than their peak compute performance". They proposed model bandwidth utilization (MBU), the bandwidth a deployment achieves divided by the hardware peak, as the number to track. NVIDIA's inference optimization guide says the same of decode: "this is a memory-bound operation."

This gives you a lower bound you can compute before you buy anything. At batch size 1, one decode step must read every weight at least once:

text
TPOT  >=  weight bytes / memory bandwidth

For Llama 3.3 70B in FP8 (about 73 GB, because the embedding and output layers stay 16-bit) on an H100 SXM (3.35 TB/s), that is at least 22 ms per token, or at most about 46 tokens per second for a single user. No kernel or framework beats that number. There are only three ways past it: read fewer bytes per token (quantization), share each read across more sequences (batching), or get more than one token from each read (speculative decoding). Every section below is one of those three moves. One correction applies to mixture-of-experts (MoE) models: at batch size 1 a decode step reads only the experts each token is routed to, so use the bytes of the active parameters. As the batch grows, more experts are touched, until every step reads nearly all of them.

The request path inside the engine

Before the details, here is where each part sits. At the center is the scheduler: on every step it decides which sequences run, asks the KV cache manager for memory, and hands one batch to the GPU workers.

d2
direction: down

client: "Client" {
  shape: person
  style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}
}

api: "API server and tokenizer\nauth, chat template, text to token IDs" {
  style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"}
}

sched: "Scheduler\nwaiting and running queues,\na token budget for each step" {
  style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}
}

kvm: "KV cache manager\nblock table, free list,\nprefix-cache index" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}

fwd: "Model executor, one worker per GPU\nforward pass; attention reads KV blocks" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}

samp: "Sampler\ntemperature, top-p, output masks" {
  style: {fill: "#C2185B"; stroke: "#8E1243"; font-color: "#FFFFFF"}
}

detok: "Detokenizer and streamer" {
  style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}
}

client -> api: "1. request"
api -> sched: "2. token IDs"
sched <-> kvm: "3. allocate or\nreuse blocks"
sched -> fwd: "4. this step's batch"
fwd -> samp: "5. logits"
samp -> sched: "6. next token\n(loop)"
samp -> detok: "7. token IDs"
detok -> client: "8. streamed text"

The API server applies the model's chat template, which turns a message list into one string, before tokenizing it. That template ships with the model weights and differs between engines, which is why it is worth pinning as its own layer. The sampler turns logits into a token. It applies temperature and top-p, and also the masks that force valid JSON in structured output. It is also where logit biasing can embed a detectable watermark. Steps 3 to 6 run once per generated token, for every sequence in the batch at the same time.

The KV cache, and why it is paged

For every token, the KV cache holds one key vector and one value vector per layer, per KV head. For a model whose layers all use full attention, its size per token is:

text
KV bytes per token = 2 x layers x KV heads x head dimension x bytes per value

Llama 3.3 70B has 80 layers, 8 KV heads and a head dimension of 128 (from its config.json). In 16-bit precision, that is 2 x 80 x 8 x 128 x 2 = 327,680 bytes, or 320 KiB, for every token. One 8,192-token request needs 2.68 GB. One request at the model's full 131,072-token context needs 40 GiB, half of an H100.

Model architecture already fights this number. Grouped-query attention (GQA) is the reason Llama 3.3 70B stores 8 KV heads instead of 64. DeepSeek's multi-head latent attention (MLA) goes further: the DeepSeek-V2 paper reports that it cuts the KV cache by 93.3%. DeepSeek-V3's config caches a 512-value latent plus a 64-value positional key per layer. That works out to about 69 KiB per token in 16-bit, less than a quarter of Llama 3.3 70B's figure, for a model with 671B total parameters. Two other model families break the formula. gpt-oss-120b alternates full-attention layers with 128-token sliding-window layers, which cache only their window. Hybrid models that replace most attention layers with Mamba-style state, now supported in SGLang v0.5.20 and vLLM, keep a fixed-size state instead of a growing cache. For these models, trust the KV capacity vLLM logs at startup over the formula.

Before 2023, engines reserved one contiguous slab of memory per request, sized for the longest possible output. Most of it stayed empty. In the systems it compared against, the PagedAttention paper (SOSP 2023) measured that "only 20.4% - 38.2% of the KV cache memory is used to store the actual token states". Its fix borrowed the operating-system idea of virtual memory.

d2
direction: down

logical: "What each request sees: contiguous logical blocks (16 tokens each)" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 2
  grid-columns: 4
  grid-gap: 8
  a0: "A block 0" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  a1: "A block 1" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  a2: "A block 2\n5 of 16 used" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  a3: "" {style: {opacity: 0}}
  b0: "B block 0" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  b1: "B block 1\n9 of 16 used" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  b2: "" {style: {opacity: 0}}
  b3: "" {style: {opacity: 0}}
}

table: "Block table: logical block to physical block\nA: 0 to #5, 1 to #1, 2 to #6\nB: 0 to #3, 1 to #0" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}

physical: "GPU memory: one pool of fixed-size physical blocks, in any order" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 2
  grid-columns: 4
  grid-gap: 8
  p0: "#0\nB block 1" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  p1: "#1\nA block 1" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  p2: "#2\nfree" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  p3: "#3\nB block 0" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  p4: "#4\nfree" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  p5: "#5\nA block 0" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  p6: "#6\nA block 2" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  p7: "#7\nfree" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
}

logical -> table: "attention kernel looks up\nwhere each block lives"
table -> physical: "blocks are allocated one at a time,\nas tokens are generated"

Each request sees its cache as a row of logical blocks. A block table maps each logical block to any free physical block, and a block is allocated only when the previous one fills. Waste is limited to the unused tail of each request's last block: 11 slots for request A in the diagram, 7 for B. vLLM's default block size is 16 tokens. Instead of assuming contiguous memory, the attention kernel follows the block table, and that one change is what lets the rest of this list work.

Two things can happen when memory runs short. If the pool cannot hold even one request at --max-model-len, vLLM V1 refuses to start, with the ValueError from the opening. If it runs out mid-run, the scheduler preempts a sequence: it frees the sequence's blocks and recomputes them later. V1 does not log each preemption. It counts them, in the vllm:num_preemptions metric and a Preemptions: field in the periodic stats log. A counter that keeps climbing means the KV budget is too small for the concurrency you allowed, not that the GPU is too slow.

Continuous batching: scheduling per step, not per request

Batching is how decode escapes the bandwidth floor. One weight read serves every sequence in the batch, so 32 sequences cost little more per step than one. The question is what happens when sequences finish at different times.

d2
grid-rows: 1
grid-gap: 30

static: "Static batching: slots sit idle until B finishes" {
  grid-rows: 5
  grid-columns: 9
  grid-gap: 4
  style: { fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 17 }
  h0: "slot" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h1: "t1" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h2: "t2" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h3: "t3" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h4: "t4" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h5: "t5" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h6: "t6" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h7: "t7" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h8: "t8" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  r1: "1" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  c1_1: "A" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_2: "A" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_3: "A" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_4: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_5: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_6: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_7: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_8: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  r2: "2" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  c2_1: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_2: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_3: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_4: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_5: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_6: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_7: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_8: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  r3: "3" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  c3_1: "C" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_2: "C" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_3: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_4: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_5: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_6: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_7: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_8: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  r4: "4" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  c4_1: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_2: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_3: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_4: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_5: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_6: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_7: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_8: "idle" {width: 46; height: 40; style: { fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 15 }}
}

cont: "Continuous batching: a freed slot takes the next request" {
  grid-rows: 5
  grid-columns: 9
  grid-gap: 4
  style: { fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 17 }
  h0: "slot" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h1: "t1" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h2: "t2" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h3: "t3" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h4: "t4" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h5: "t5" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h6: "t6" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h7: "t7" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  h8: "t8" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  r1: "1" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  c1_1: "A" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_2: "A" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_3: "A" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c1_4: "E" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c1_5: "E" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c1_6: "E" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c1_7: "E" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c1_8: "E" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  r2: "2" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  c2_1: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_2: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_3: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_4: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_5: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_6: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_7: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c2_8: "B" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  r3: "3" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  c3_1: "C" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_2: "C" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c3_3: "F" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c3_4: "F" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c3_5: "F" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c3_6: "F" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c3_7: "F" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c3_8: "F" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  r4: "4" {width: 46; height: 40; style: { fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"; font-size: 15 }}
  c4_1: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_2: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_3: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_4: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_5: "D" {width: 46; height: 40; style: { fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 15 }}
  c4_6: "G" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c4_7: "G" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
  c4_8: "G" {width: 46; height: 40; style: { fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 15 }}
}

With static batching, a batch runs until its longest member, B, finishes. Request A finished after three steps and request C after two, and their slots sat empty while E, F and G waited in the queue. In the left grid, 15 of 32 slot-steps are idle. With continuous batching the scheduler re-forms the batch on every step, so a slot freed at the end of one step takes a waiting request on the next. Orca (OSDI 2022) introduced this idea as "iteration-level scheduling". It reported a 36.9x throughput gain over FasterTransformer on GPT-3 175B at the same latency.

Continuous batching needs the paged cache. A request that joins mid-batch needs memory immediately, in whatever pieces are free. Even offline jobs use it: vLLM's offline LLM.generate API and Ray Data run the same per-step scheduler.

Chunked prefill: keep long prompts from stalling everyone else

Continuous batching created a new problem. When a 16,000-token prompt joins, its prefill is one huge forward pass, and every sequence in the decode batch waits for it. Users in the middle of a response see the stream freeze, however carefully the client side streams tokens to the browser.

d2
direction: down

without: "Without chunked prefill: a long prompt takes a whole step" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 1
  grid-gap: 6
  a: "step 1\n64 decode\ntokens" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  b: "step 2\n8,192-token prefill\nall 64 decodes wait" {width: 330; style: {fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"}}
  c: "step 3\n65 decode\ntokens" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
}

with: "With chunked prefill: every step mixes decodes with one slice of the prompt" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 2
  grid-columns: 4
  grid-gap: 6
  d1: "decode\n64" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  d2: "decode\n64" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  d3: "decode\n64" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  d4: "decode\n64" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  p1: "prefill chunk 1\n1,984 tokens" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  p2: "prefill chunk 2\n1,984 tokens" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  p3: "prefill chunk 3\n1,984 tokens" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  p4: "chunk 4, then\nthe rest next step" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
}

note: "Each column is one step with a 2,048-token budget.\nThe 64 running requests get a token every step,\nso their time per output token stays flat." {
  style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"}
}

without -> with: "same work, re-sliced"
with -> note

Chunked prefill gives every step a token budget. First the scheduler fills the budget with one decode token per running sequence, then it spends the rest on a slice of the waiting prompt. Sarathi-Serve (OSDI 2024), the paper that built this, calls the result "stall-free schedules". It reported 2.6x more serving capacity for Mistral-7B on one A100 than the engine it was compared with. In vLLM V1, chunked prefill is "enabled by default whenever possible", and the scheduler "prioritizes decode requests". The diagram shows vLLM's behaviour. SGLang also chunks long prefills, but it mixes decode tokens into the same batch only when --enable-mixed-chunk is set, and that flag is off by default.

Only one setting here is worth knowing: the budget, --max-num-batched-tokens, which trades one latency for another. vLLM's tuning guide says it directly: "Smaller values (e.g., 2048) achieve better ITL", where ITL is inter-token latency, while "Higher values achieve better time to first token". For throughput it recommends a value above 8,192 "especially for smaller models on large GPUs". Chat traffic that users watch token by token wants a small budget. A retrieval pipeline that only cares when the answer starts wants a large one.

Prefix caching: never compute the same prompt prefix twice

Most production prompts start the same way: a long system prompt, a fixed set of tool definitions, or the same retrieved document. Prefix caching keeps the KV blocks of a finished prefill and lets the next request reuse them.

d2
direction: down

r1: "Request 1 (first to arrive)" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 1
  grid-gap: 6
  s0: "system\nblock 0" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  s1: "system\nblock 1" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  s2: "system\nblock 2" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  u: "user A\nblock" {style: {fill: "#FFA07A"; stroke: "#D9785A"; font-color: "#2C2C2A"}}
}

index: "Prefix-cache index\nkey = hash(parent block hash, this block's tokens)\nvalue = physical block in GPU memory" {
  style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}
}

r2: "Request 2 (same system prompt, different question)" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 1
  grid-gap: 6
  s0: "system 0\nHIT" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  s1: "system 1\nHIT" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  s2: "system 2\nHIT" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  u: "user B\nMISS" {style: {fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"}}
}

result: "Request 2 runs prefill only on its last block.\nThe three shared blocks are read, not recomputed." {
  style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"}
}

r1 -> index: "full blocks are hashed\nand registered after prefill"
index -> r2: "lookup walks the chain\nuntil the first miss"
r2 -> result

vLLM hashes each full block "by the tokens in the block and the tokens in the prefix before the block". In practice, a block's key includes its parent block's hash, so two requests share a block only if everything before it is identical too. That has a practical consequence the diagram makes visible. Put a timestamp or a user name at the top of the prompt, and every block after it misses. Moving variable text to the end of the prompt is usually the cheapest prefix-caching improvement.

vLLM turned this on by default in its V1 engine, after measuring that it costs "less than 1% decrease in throughput even when the cache hit rate is 0%". SGLang's version, RadixAttention, stores prefixes as a tree, which suits agent traffic that branches from a shared history. One limit is worth repeating from vLLM's own docs: prefix caching shortens prefill only. It "does not reduce the time of generating new tokens".

Attention kernels and CUDA graphs

Two more defaults sit below the scheduler. First is the attention kernel, which computes attention in tiles that fit in on-chip memory, so the full attention matrix never has to be written out. FlashAttention-3 reached about 75% utilization of an H100. FlashAttention-4, published in March 2026, is "optimized for Hopper and Blackwell GPUs" per its repository, and Together AI reports up to 1,605 TFLOPs/s in 16-bit on a B200. vLLM and SGLang pick among these and FlashInfer (MLSys 2025) automatically per GPU. You rarely choose a kernel by hand. You choose hardware, and the engine chooses the kernel.

CUDA graphs record a whole decode step once and replay it, which removes per-step CPU launch overhead. That matters most at small batch sizes, where each step is short. vLLM's --performance-mode (default balanced) adjusts this: interactivity uses finer-grained graphs and latency-oriented kernels for small batches, and throughput uses larger graphs and more aggressive batching. --enforce-eager turns graphs off, which is useful when debugging and expensive in production.

FP8 weights and an FP8 KV cache: the default on current hardware

Quantization reads fewer bytes per token, which goes straight to the decode bound. On NVIDIA Ada (such as the L40S), Hopper and Blackwell GPUs and AMD MI300-series hardware, FP8 runs natively in the tensor cores, so it halves both the weight memory and the bandwidth per decode step with no unpacking.

d2
direction: down

bars: "Weight memory for a 70.6B-parameter model, by number format" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 5
  grid-columns: 3
  grid-gap: 8

  l1: "BF16\n16 bits per weight" {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}}
  s1: "141 GB" {style: {fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"; font-size: 20}}
  n1: "Needs two 80 GB GPUs\nbefore any KV cache" {style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}}

  l2: "FP8 (E4M3)\n8 bits per weight" {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}}
  s2: "71 GB" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"; font-size: 20}}
  n2: "Native on Ada (L40S), Hopper,\nBlackwell, MI300 and newer" {style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}}

  l3: "NVFP4\n4 bits + 1 FP8 scale\nper 16 values" {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}}
  s3: "about 40 GB" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 20}}
  n3: "Native on NVIDIA Blackwell" {style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}}

  l4: "MXFP4\n4 bits + 1 power-of-two\nscale per 32 values" {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}}
  s4: "about 38 GB" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"; font-size: 20}}
  n4: "Open MX standard;\nBlackwell and MI355X" {style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}}

  l5: "INT4 (AWQ, GPTQ)\n4 bits + group scales" {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}}
  s5: "about 38 GB" {style: {fill: "#FFA07A"; stroke: "#D9785A"; font-color: "#2C2C2A"; font-size: 20}}
  n5: "Weights only: unpacked to 16-bit\nfor every multiply, on any GPU" {style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-color: "#2C2C2A"}}
}

The grid prices every weight in the given format. Real FP8 checkpoints keep the embedding and output layers in 16-bit, which brings Llama 3.3 70B to about 73 GB, the figure the rest of this article uses. Two kinds of 4-bit format appear in the grid. NVFP4 and MXFP4 are microscaling formats: each small block of 4-bit values shares one scale factor, and Blackwell and AMD MI355X tensor cores compute on them directly. INT4 methods such as AWQ and GPTQ store 4-bit weights but unpack them to 16-bit before each multiply, so they save memory and bandwidth but add work at large batch sizes.

Some models now ship in these formats. DeepSeek-V3 was trained with "FP8 mixed precision" and its checkpoint is FP8. OpenAI's gpt-oss-120b ships its expert weights in MXFP4.

FP8's accuracy cost is small enough to make it the default. In a large 2025 study (ACL 2025, "Give Me BF16 or Give Me Death"?), 8-bit quantization of 70B models recovered 99.75% of 16-bit accuracy on the Open LLM Leaderboard. A September 2026 Red Hat guide calls FP8 weights-and-activations on "Hopper or newer, serving real concurrency" "the production default". You can put the KV cache in FP8 too: --kv-cache-dtype fp8 halves the 320 KiB per token and doubles how many requests fit. Check output quality on your own evaluation set when you change it, since KV precision affects every later token. vLLM v0.30.0 also accepts 4-bit KV formats, but its own May 2026 study concluded that FP8 "remains the best default for KV-cache quantization". The 4-bit TurboQuant variants gave only "modest KV-cache savings (2.4x vs 2x)", at a consistent cost in throughput and latency.

The alternatives: 4-bit weights, speculative decoding, parallelism, disaggregation and KV offload

Each technique below is one of the three moves against the decode floor, pushed further than the defaults push it. 4-bit weights read even fewer bytes per token. Speculative decoding gets more than one token from each weight read. Tensor parallelism adds bandwidth by splitting the reads across GPUs. Disaggregation and KV offload protect the second move, batching, from the things that break it: prefill interference and cache evictions. Each one is a real improvement for some workloads and a cost or a regression for others, which is why none of them is a default.

4-bit weights

The strongest case. When one user waits for one answer, decode is purely bandwidth-bound, and 4-bit weights move half the bytes of FP8. The ACL 2025 study found that 16-bit-activation INT4 weights (W4A16) "is the most cost-efficient for synchronous setups". On a 70B model on one H100, a code-completion request finished in 28.6 seconds with INT4 against 32.8 with FP8. 4-bit also fits models onto fewer cards: the same paper ran the 405B model on 4 GPUs where 16-bit needed 16. Accuracy holds up well at this size. 70B models recovered 99.36% on the Open LLM Leaderboard at INT4. Red Hat reports NVFP4 at about 99% of 16-bit accuracy for 70B to 235B models. 4-bit also works on the capacity side of the budget: against FP8 it frees more than 30 GB on a 70B model, which becomes KV cache when concurrency is limited by memory rather than compute.

Why it is not the default. Under real concurrency the advantage reverses. In the same study's multi-user tests on 4 H100s, FP8 served 18.4 multi-turn chat queries per second against 16.1 for INT4, and 4.0 code-fixing queries against 3.1. In the paper's summary, "W8A8 dominates in asynchronous continuous batching". Smaller models, which SLM-first designs route much of their traffic to, also lose more: Red Hat reports 95% to 98% recovery for 7B to 14B models in NVFP4. Those throughput tables were measured on a 2024 vLLM, so trust the direction more than the exact ratios.

Speculative decoding

The strongest case. Decode leaves compute unused, and speculative decoding spends it. A cheap drafter proposes several tokens, and the target model checks all of them in one forward pass, which costs about the same weight read as producing one token.

d2
direction: down

ctx: "Text so far: \"The cache was\"" {
  style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}
}

draft: "Drafter: EAGLE head, MTP module or n-gram lookup\nCheap: proposes k = 4 tokens one after another" {
  style: {fill: "#FFA07A"; stroke: "#D9785A"; font-color: "#2C2C2A"}
}

target: "Target model: ONE forward pass scores all 4 proposals together\nThe same weight read that would have produced 1 token" {
  style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}
}

verify: "Verification, left to right" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 1
  grid-gap: 6
  t1: "\"evicted\"\naccepted" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  t2: "\"after\"\naccepted" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  t3: "\"one\"\nrejected" {style: {fill: "#E74C3C"; stroke: "#B0392D"; font-color: "#FFFFFF"}}
  t4: "\"hour\"\ndiscarded" {style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}}
}

out: "Output this step: \"evicted after\" + the target's own token \"ten\"\n3 tokens for one target pass. The text is the same as without speculation." {
  style: {fill: "#98D8C8"; stroke: "#5FA898"; font-color: "#2C2C2A"}
}

ctx -> draft
draft -> target: "4 draft tokens"
target -> verify: "compare with the target's\nown choice at each position"
verify -> out

The output does not change. Leviathan et al. (ICML 2023) report "identical outputs", and Chen et al. (2023) show that the method preserves the target model's distribution. Drafters have improved a great deal since then. EAGLE-3 (2025) trains a small head on the target's own hidden states and reports speedups up to 6.5x. Multi-token prediction (MTP) modules, which DeepSeek-V3 trains as part of the model, act as a built-in drafter. SGLang's documentation measures Llama 3.1 8B on one H100 going from 158 to 373 tokens per second with EAGLE-3. For agents and code editing, which repeat long spans of earlier text, n-gram and suffix drafters need no extra model at all. SuffixDecoding (NeurIPS 2025) reports up to 5.3x on SWE-Bench-style agent traces. New drafters keep arriving: EAGLE 3.1 (May 2026) makes EAGLE-3 more robust inside vLLM, and vLLM's August 2026 guide adds block drafters such as DFlash and DSpark to the list.

Why it is not the default. Speculation spends spare compute, and at high load there is none to spare. vLLM's 2024 measurements showed 1.4x and 1.8x slowdowns at high request rates. Red Hat's 2025 EAGLE-3 tests on Llama 3.3 70B found latency "reduced by up to 1.6X at low request rates, but latency increases at higher request rates due to compute saturation". It also hurts on text the drafter predicts poorly, such as translation, where the best draft length was "1, or even 0". vLLM's docs scope the feature to "medium-to-low QPS", queries per second. Its dynamic speculative decoding lowers the draft length as the batch grows, down to zero above a batch size you set. One 2026 result goes the other way: Red Hat measured EAGLE-3 on gpt-oss-120b still gaining up to 200 concurrent requests. That is one sparse mixture-of-experts (MoE) model on one H200, so measure your own model before you assume it.

Scaling out: tensor, pipeline and expert parallelism

When weights plus the KV budget do not fit on one GPU, the model has to be split. There are three ways to split it, and they use the interconnect very differently.

d2
grid-columns: 1
grid-gap: 40

tp: "Tensor parallel: every layer split across the GPUs of one node" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  g0: "GPU 0\n1/4 of every\nweight matrix" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  g1: "GPU 1\n1/4 of every\nweight matrix" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  g2: "GPU 2\n1/4 of every\nweight matrix" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  g3: "GPU 3\n1/4 of every\nweight matrix" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  ar: "all-reduce after every layer\nneeds NVLink-class bandwidth" {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}}
  g0 -> ar
  g1 -> ar
  g2 -> ar
  g3 -> ar
}

pp: "Pipeline parallel: whole layers per stage, often one stage per node" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  grid-rows: 1
  grid-gap: 10
  s0: "Stage 0 on node 1\nlayers 1-40" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  hop: "activations passed on,\nonce per micro-batch" {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}}
  s1: "Stage 1 on node 2\nlayers 41-80" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
}

ep: "Expert parallel (MoE): experts spread out, tokens travel to them" {
  label.near: top-center
  style: {fill: "#FFFFFF"; stroke: "#95A5A6"; font-size: 18}
  router: "router picks the top-k\nexperts for each token" {style: {fill: "#C2185B"; stroke: "#8E1243"; font-color: "#FFFFFF"}}
  e0: "GPU 0\nexperts 0-63" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  e1: "GPU 1\nexperts 64-127" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  e2: "GPU 2\nexperts 128-191" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  e3: "GPU 3\nexperts 192-255" {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  router -> e0: "all-to-all"
  router -> e1
  router -> e2
  router -> e3
}

Data parallelism (DP) comes first. When the model fits on one GPU with room for its KV cache, run one full replica per GPU instead of splitting it, and put a load balancer in front. Replicas need no traffic between GPUs at all.

Once there are several replicas, the router decides your prefix-cache hit rate, because a request only hits the cache on the replica that served its prefix. llm-d measured this on 8 vLLM pods (16 H100s) serving Qwen3-32B. Precise prefix-cache-aware scheduling reached 8,730 output tokens per second, "over double the throughput of the cache-blind configurations", and returned first tokens in 0.542 seconds where the cache-blind schedulers took over 90 seconds. llm-d's scheduler, Dynamo's KV router and Ray Serve LLM all route this way.

Tensor parallelism (TP) splits every weight matrix across GPUs, for a model that does not fit on one. Each GPU reads only its share of the weights per step, so TP also multiplies the bandwidth that bounds decode. Its price is an all-reduce after every layer, which is why vLLM's guidance is to use TP inside a node and set tensor_parallel_size "to the number of GPUs per node". On GB200 and GB300 NVL72 racks, that NVLink domain is 72 GPUs rather than 8. TP is part of the default.

TP has one limit that matters for KV sizing: it splits the KV cache only up to the number of KV heads. An MLA model keeps one shared latent per layer, so under TP every GPU holds a full copy of it, and TP adds no KV capacity. For those models SGLang runs attention data-parallel (--enable-dp-attention), so each GPU holds the cache of its own requests, combined with expert parallelism.

For long contexts, vLLM adds decode context parallelism (DCP), which splits each sequence's cache across GPUs. On 8 B200s serving Kimi K2.6, TP alone filled the KV cache at 64 concurrent requests and plateaued near 1,863 tokens per second per GPU, while DCP reached 6,091 at 512 concurrent requests.

Pipeline parallelism (PP) places whole layers on different GPUs and passes activations forward once per micro-batch. It needs far less bandwidth than TP. Where it wins: across nodes, and inside nodes whose GPUs lack NVLink. vLLM's docs name the L40S and say to use PP "instead of tensor parallelism for higher throughput" there. Why it is not the default: each request still passes through every stage in order, and the docs warn that raising PP "may cause latency penalties".

Expert parallelism (EP) is for mixture-of-experts models such as DeepSeek-V3 (671B total parameters, 37B active per token, 256 routed experts) and gpt-oss-120b (128 experts, 4 per token). Every expert must be resident in memory, but each token uses only a few. EP spreads the experts across GPUs and sends each token to the GPUs that hold its experts. Where it wins: large MoE models at scale. SGLang's team ran DeepSeek on 96 H100s with EP plus disaggregated serving, and measured up to 5x the output throughput of plain TP on the same hardware. Why it is not the default: EP is off unless you pass --enable-expert-parallel; without it, a single-node MoE model splits its experts with tensor parallelism. Across nodes, the all-to-all traffic needs a fast fabric.

Disaggregated prefill and decode

The strongest case. Prefill and decode compete for the same GPU, and chunked prefill only shares the time more fairly between them. Disaggregation runs them on separate pools of workers, each sized and tuned for its own bottleneck.

d2
direction: down

client: "Client" {
  shape: person
  style: {fill: "#95A5A6"; stroke: "#6B7B7C"; font-color: "#2C2C2A"}
}

router: "KV-aware router\npicks a prefill worker that already caches the prefix,\nand a decode worker with free KV memory" {
  style: {fill: "#C2185B"; stroke: "#8E1243"; font-color: "#FFFFFF"}
}

pre: "Prefill pool: tuned for compute" {
  label.near: top-right
  style: {fill: "#EAF2FB"; stroke: "#2C6FB0"; font-size: 18}
  p1: "Prefill worker 1\nprocesses the whole prompt,\nproduces the first token" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  p2: "Prefill worker 2" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
}

dec: "Decode pool: tuned for memory bandwidth and large batches" {
  label.near: bottom-right
  style: {fill: "#F0EDFD"; stroke: "#5A4BC4"; font-size: 18}
  d1: "Decode worker 1\nholds the KV cache,\ngenerates the rest" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  d2: "Decode worker 2" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  d3: "Decode worker 3" {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
}

client -> router: "1. request"
router -> pre.p1: "2. prompt"
pre.p1 -> dec.d1: "3. KV cache moves GPU to GPU\n(NIXL over NVLink or RDMA)"
dec.d1 -> client: "4. tokens streamed"

Research on this came from both sides of the industry in the same few months. DistServe (OSDI 2024) reported serving up to 7.4x more requests within the same latency targets. Microsoft's Splitwise (ISCA 2024) reported 1.4x higher throughput at 20% lower cost. Mooncake, the serving system behind Moonshot AI's Kimi, is the strongest production evidence. Its FAST 2025 paper, which won Best Paper, reports 59% to 498% more effective request capacity within latency targets, and says the system processes "over 100 billion tokens daily". Today the pattern ships as NVIDIA Dynamo (v1.5.0, 21 September 2026) and llm-d (v0.10.0, 29 September 2026), both running on top of vLLM, SGLang or TensorRT-LLM.

Why it is not the default. Every request now moves its KV cache across a network, and adds a router, two worker pools and a transfer layer you have to operate. Dynamo's own documentation is unusually direct: "It is not automatically better. For small models, short prompts, low concurrency, or clusters without a fast KV-transfer fabric, an aggregated deployment is simpler and often faster." Across nodes without remote direct memory access (RDMA), transfers fall back to TCP and "KV movement can dominate TTFT". llm-d's guide names its target as medium-to-large models with long inputs, "e.g 10k ISL | 1k OSL, not 200 ISL | 200 OSL" (input and output sequence lengths). Configuration mistakes also fail quietly: Dynamo warns that prefill and decode workers with different dtypes, block sizes or KV layouts produce "transfer errors or silently corrupt output".

KV cache offload to CPU memory or SSD

The strongest case. Prefix caching only helps while the blocks stay in GPU memory, and long multi-turn conversations or repeated documents evict each other. Offloading moves evicted blocks to CPU memory or SSD instead of deleting them. vLLM's native offloading connector (January 2026) reported TTFT reductions of 2x to 22x on Llama 3.1 8B, depending on prompt size, when blocks were reloaded from CPU memory instead of recomputed. It also helps without any prefix sharing, because a preempted request can reload its blocks instead of recomputing them. LMCache, the most-used external layer, reported 3.0x lower average TTFT against GPU-only caching on 739 Claude Code agent traces, with 32 users and 100,000-token contexts on AMD MI300X. That figure is its own benchmark of one model on one setup. The direction of the field favours offload inside the engine: Dynamo v1.5.0 deprecated its own KV block manager (KVBM), with removal targeted for v1.6.0.

Why it is not the default. Its value grows with the hit rate and falls toward zero without one. At low load, the same LMCache benchmark found that GPU-only prefix caching wins. Per the LMCache paper, context truncation, which many applications do, "can greatly reduce prefix cache hit ratio by half". LMCache also has an operational trap: without PYTHONHASHSEED=0, cache keys differ between processes and the hit rate is 0% "even on bit-identical prompts". In vLLM it stays off until you set --kv-offloading-size.

Same mechanisms everywhere: vLLM, SGLang, TensorRT-LLM, and what is no longer an option

Across vLLM, SGLang and TensorRT-LLM, the mechanisms above are the same. All three ship paged KV, continuous batching, chunked prefill, prefix caching, FP8, speculative decoding and disaggregation hooks. That is why the engine choice moves throughput by tens of percent rather than multiples. The engine guide covers it in full: vLLM by default, SGLang or TensorRT-LLM on a measured condition, llama.cpp for machines with no GPU, and Ollama (ollama run llama3.2) for one developer on a laptop. Two engines the previous version of this article recommended are no longer options for new work. Hugging Face archived TGI on 21 March 2026 and points users to vLLM and SGLang. DeepSpeed-MII has had no release since v0.3.3 in March 2025. Ray Serve LLM and KServe sit above the engine, not beside it: they handle routing, autoscaling and multi-model deployments, and run vLLM or SGLang underneath. Ray Serve LLM, for example, adds prefix-aware and KV-cache-aware routing across replicas, which raises the hit rate that the prefix cache inside each engine depends on.

When to deviate from the defaults

Each row is a condition you can check against your own traffic, and the technique that wins when it holds.

ConditionTurn onWhy
Weights plus your KV budget do not fit on one GPU, and the GPU runs FP8 natively (Ada, Hopper, Blackwell, MI300-class)FP8 weights, then --kv-cache-dtype fp8Halves weight and KV memory with about 99.7% accuracy recovery on 70B-class models
The model fits on one GPU with room for KV cacheOne replica per GPU (data parallel) behind a load balancerNo inter-GPU traffic; scales throughput linearly
More than one replica, and requests share prefixesPrefix-cache-aware routing (llm-d scheduler, Dynamo KV router, Ray Serve LLM)A prefix only hits on the replica that cached it; llm-d measured over 2x throughput
Still does not fit, GPUs in one NVLink domainTensor parallel, sized to the domain (8 GPUs on HGX, 72 on NVL72)Splits weights and multiplies decode bandwidth
An MLA model, or TP size larger than the KV-head count, and you are KV-boundData-parallel attention plus expert parallel; decode context parallel for long contextsTP copies the KV cache instead of splitting it
Spans nodes, or the GPUs lack NVLink (L40S-class)Pipeline parallel across nodes or linksFar less traffic than TP's per-layer all-reduce
A large MoE model across many GPUsExpert parallelExperts must all be resident; tokens move to them
Few concurrent users and latency-bound, or KV-bound on Blackwell or MI355X with your evaluation set passing at 4-bit4-bit weights (INT4, or NVFP4/MXFP4 where native)Fewest bytes per token and more than 30 GB more KV room on 70B; loses throughput to FP8 under heavy batching
Decode latency matters and concurrency is low to mediumSpeculative decoding: EAGLE-3 or MTP; n-gram or suffix for agents and codeUses compute that decode leaves idle; slows down at high load unless you configure dynamic draft lengths
Long prompts relative to outputs (llm-d's example is 10,000 input and 1,000 output tokens), high concurrency, medium-to-large models, and an RDMA or NVLink fabricDisaggregated prefill and decode (Dynamo, llm-d)Separate pools stop prefill from interfering with decode
Your prefix-cache hit rate drops because blocks are evicted, not because prompts differKV offload (--kv-offloading-size, or LMCache)Turns evictions into reloads instead of recomputation
Users watch tokens stream and complain about stutterLower --max-num-batched-tokensSmaller chunks improve inter-token latency
Only time-to-answer matters (batch retrieval, offline jobs)Raise --max-num-batched-tokens above 8,192Larger chunks improve TTFT and throughput

The wrong way and the right way to size a deployment

The opening failure comes from a sizing check that every team writes once. It compares the weights with the GPU memory and stops there:

python
# Wrong: checks only that the weights fitfits = weight_bytes(model, bytes_per_param) < gpu_gb * GB * n_gpus

Weights that fit tell you the server will start. They do not tell you how many requests the server can run at once, and that number is set by whatever memory is left for the KV cache. Below is a script that does both checks. It uses Llama 3.3 70B's real architecture values, keeps the embedding and output layers in 16-bit as FP8 quantization does, vLLM's default --gpu-memory-utilization of 0.92, and a reserve for activations and CUDA graphs. vLLM measures that reserve at startup, so here it is a stated assumption of 4 GB per GPU:

python
# kv_budget.py - how many full-length requests fit after the weights are loadedfrom dataclasses import dataclass@dataclassclass Model:    name: str    params_billion: float    unquantized_billion: float  # embeddings and output head: FP8 leaves them 16-bit    layers: int    kv_heads: int    head_dim: int# Values from the model's config.json on Hugging Face;# 2 x 128,256 vocab x 8,192 hidden = 2.1B embedding + output-head paramsLLAMA_70B = Model("Llama-3.3-70B", 70.6, 2.1, layers=80, kv_heads=8, head_dim=128)GB = 1e9def weight_bytes(m: Model, bytes_per_param: float) -> float:    quantized = (m.params_billion - m.unquantized_billion) * 1e9 * bytes_per_param    return quantized + m.unquantized_billion * 1e9 * 2def kv_bytes_per_token(m: Model, kv_bytes: float) -> float:    # one key and one value vector, per KV head, per layer    return 2 * m.layers * m.kv_heads * m.head_dim * kv_bytesdef wrong_way(m: Model, gpu_gb: float, n_gpus: int, bytes_per_param: float) -> str:    # Wrong: checks only that the weights fit    fits = weight_bytes(m, bytes_per_param) < gpu_gb * GB * n_gpus    return "fits" if fits else "does not fit"def right_way(m: Model, gpu_gb: float, n_gpus: int, bytes_per_param: float,              kv_bytes: float, context: int, mem_util: float = 0.92,              reserve_gb_per_gpu: float = 4.0) -> int:    # Right: what is left for the KV cache decides how many requests run at once.    # mem_util mirrors vLLM's --gpu-memory-utilization default; the reserve    # stands in for activations and CUDA graphs, which vLLM measures at startup.    budget = n_gpus * (gpu_gb * GB * mem_util - reserve_gb_per_gpu * GB)    free_for_kv = budget - weight_bytes(m, bytes_per_param)    per_request = kv_bytes_per_token(m, kv_bytes) * context    return max(0, int(free_for_kv // per_request))if __name__ == "__main__":    m = LLAMA_70B    print(f"KV cache per token, BF16: {kv_bytes_per_token(m, 2) / 1024:.0f} KiB")    print(f"KV cache per 8K-token request, BF16: "          f"{kv_bytes_per_token(m, 2) * 8192 / GB:.2f} GB")    print()    print(f"{'setup':<38}{'wrong way':<14}{'8K requests that fit'}")    for label, n, kvb in [        ("1x H100 80GB, FP8 weights, BF16 KV", 1, 2),        ("2x H100 80GB, FP8 weights, BF16 KV", 2, 2),        ("2x H100 80GB, FP8 weights, FP8 KV", 2, 1),        ("2x H100 80GB, BF16 weights, BF16 KV", 2, 2),    ]:        bpp = 2 if "BF16 weights" in label else 1        print(f"{label:<38}{wrong_way(m, 80, n, bpp):<14}"              f"{right_way(m, 80, n, bpp, kvb, 8192)}")

Running it with python kv_budget.py prints:

text
KV cache per token, BF16: 320 KiBKV cache per 8K-token request, BF16: 2.68 GBsetup                                 wrong way     8K requests that fit1x H100 80GB, FP8 weights, BF16 KV    fits          02x H100 80GB, FP8 weights, BF16 KV    fits          242x H100 80GB, FP8 weights, FP8 KV     fits          492x H100 80GB, BF16 weights, BF16 KV   fits          0

The wrong check says "fits" four times. The right check shows two deployments that cannot serve one full 8K request, including the 16-bit model on two H100s, where the weights need 141 GB and only 139 GB is left after the engine's reserve. It also shows that FP8 KV cache doubles concurrency on the same hardware. The single-H100 row is the opening failure.

Next, the right way sets the engine to match the arithmetic, and the row with 49 requests becomes this command. It uses the same two H100s, FP8 weights quantized at load time, an FP8 KV cache, an 8K context, and a concurrency cap at the computed number:

bash
vllm serve meta-llama/Llama-3.3-70B-Instruct \  --host 127.0.0.1 \  --tensor-parallel-size 2 \  --quantization fp8 \  --kv-cache-dtype fp8 \  --max-model-len 8192 \  --max-num-seqs 44

--max-num-seqs 44 keeps a margin under the computed 49, because the real activation reserve is measured at startup and may differ from the script's 4 GB. vLLM logs its own version of this number when it starts, as a line reading GPU KV cache size: ... tokens, Maximum concurrency for 8,192 tokens per request: .... If that figure is far from your estimate, trust the log and update your reserve. Continuous batching, chunked prefill and prefix caching are not on the command line, because they are already on. --host 127.0.0.1 follows the engine guide's security advice, since vLLM otherwise listens on every interface.

Then measure, with load shaped like your traffic. vLLM ships a benchmark client:

bash
vllm bench serve --backend vllm \  --model meta-llama/Llama-3.3-70B-Instruct \  --dataset-name random --random-input-len 2048 --random-output-len 512 \  --ignore-eos \  --max-concurrency 32 --num-prompts 500

--ignore-eos makes every response reach its full length, so TPOT is measured on the output length you asked for. The client reports mean, median and p99 for TTFT, TPOT and ITL, plus request and token throughput. Two cautions come from its own documentation. Repeated runs against the same server "can reuse prompts left in the prefix cache and inflate throughput", so restart between runs or vary the prompts. And for production load testing it recommends GuideLLM; the bench client is meant for comparing features and catching regressions.

Which optimization to turn on: the decision as a diagram

This is the deviation table above drawn as an order of questions. Each row on the left is a check, and the box on its right is what to turn on if the check calls for it.

d2
grid-rows: 7
grid-columns: 2
grid-gap: 26

classes: {
  q: {style: {fill: "#7B68EE"; stroke: "#5A4BC4"; font-color: "#FFFFFF"}}
  act: {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
  opt: {style: {fill: "#FFA07A"; stroke: "#D9785A"; font-color: "#2C2C2A"}}
  done: {style: {fill: "#6BCF7F"; stroke: "#3F9E52"; font-color: "#2C2C2A"}}
  blank: {label: ""; style: {opacity: 0}}
}

start: "1. Start from the engine defaults:\ncontinuous batching, paged KV, chunked prefill,\nprefix caching and CUDA graphs are all on" {class: done}
start_note: "Change nothing yet.\nMeasure TTFT, TPOT and throughput first." {class: done}

fit: "2. At your target concurrency and context,\ndo weights + KV cache fit on one GPU?" {class: q}
fp8: "If it fits: one replica per GPU.\nIf not: FP8 weights, then FP8 KV cache\n(Ada, Hopper, Blackwell, MI300 or newer)" {class: act}

fit2: "3. Still does not fit?" {class: q}
tp: "Tensor parallel inside one NVLink domain;\npipeline parallel across nodes;\nEP and DP attention for MoE and MLA" {class: act}

lat: "4. Is decode latency the problem,\nat low or medium concurrency?" {class: q}
spec: "If yes: speculative decoding\n(EAGLE-3, MTP or n-gram)" {class: opt}

pre: "5. Long prompts, high concurrency,\na medium or large model,\nand a fast KV-transfer fabric?" {class: q}
dis: "If yes: disaggregate prefill and decode\n(NVIDIA Dynamo, llm-d)" {class: opt}

reuse: "6. Prefixes repeat, but the GPU\nKV cache keeps evicting them?" {class: q}
off: "If yes: offload KV to CPU or SSD\n(LMCache, native offloading)" {class: opt}

measure: "7. Re-measure after every change.\nKeep it only if the numbers moved." {class: done}
m_blank: {class: blank}

start -> fit
fit -> fit2
fit2 -> lat
lat -> pre
pre -> reuse
reuse -> measure

The order matters. Memory comes first, because no other change helps a deployment that preempts. Latency techniques come before topology changes, because speculative decoding is one flag and disaggregation is a new cluster layout.

Hardware: what memory bandwidth means for decode

Because decode is bandwidth-bound, two numbers on a spec sheet predict most of a chip's serving behavior: memory capacity, which sets how much KV cache fits, and memory bandwidth, which sets the decode floor. Applying the batch-1 bound from the first section to the 73 GB FP8 Llama 3.3 70B on one chip gives the last column. For an MoE model, the bound uses its active parameters instead. Specifications come from each vendor's own page, read on 1 October 2026 and listed in the References. The ceiling is a number no software reaches, and it is only a starting point for comparing chips. Batching and tensor parallelism raise total throughput far above it.

AcceleratorMemoryBandwidthLow-precision formatsBatch-1 ceiling, dense 70B FP8 on one chip
NVIDIA H100 SXM80 GB3.35 TB/sFP846 tok/s (no room left for KV)
NVIDIA H200141 GB4.8 TB/sFP866 tok/s
NVIDIA B200up to 192 GB (180 GB per GPU in DGX B200)8 TB/sFP8, FP6, NVFP4110 tok/s
NVIDIA B300 (Blackwell Ultra)up to 288 GB8 TB/sFP8, FP6, NVFP4110 tok/s
AMD Instinct MI300X192 GB5.3 TB/sFP873 tok/s
AMD Instinct MI325X256 GB6 TB/sFP883 tok/s
AMD Instinct MI355X288 GB8 TB/sFP8, MXFP6, MXFP4110 tok/s
Google TPU v6e (Trillium)32 GB1.64 TB/sBF16, INT8does not fit on one chip
Google TPU7x (Ironwood)192 GiB7.38 TB/sFP8102 tok/s
AWS Inferentia232 GiB820 GiB/scFP8, INT8does not fit on one chip
AWS Trainium296 GiB2.9 TB/sFP840 tok/s
AWS Trainium3144 GiB4.9 TB/sMXFP8, MXFP467 tok/s
Intel Gaudi 3128 GB3.7 TB/sFP851 tok/s
Apple M5 Maxup to 128 GB unified614 GB/snot stated by Apple8 tok/s
Apple M5 Ultra (Mac Studio)up to 512 GB unified1.2 TB/snot stated by Apple17 tok/s

Three things stand out. First, capacity now matters as much as bandwidth: an H100 and a B200 differ by 2.4x in bandwidth but by more than 2x in capacity. Capacity is what turns the single-H100 "fits, serves zero" row from the sizing script into a working deployment. Second, AMD's MI300-series parts carry more memory per chip than the NVIDIA parts of the same generation, which is their strongest serving argument. Third, a Mac with 512 GB of unified memory can hold models no single data-center GPU can, but at 17 tokens per second for a 70B model it is a tool for one user, not a server. Graphcore, which the previous version of this article listed, was acquired by SoftBank in July 2024.

Decision checklist

  • Leave continuous batching, chunked prefill, prefix caching and CUDA graphs on. Search your launch flags for --no-enable-prefix-caching, --no-enable-chunked-prefill and --enforce-eager, and remove any you cannot justify.
  • Before you choose hardware, compute KV bytes per token from the model's config.json, multiply by your context length and target concurrency, and add the weights. Run kv_budget.py with your numbers. For sliding-window and hybrid models, use the KV capacity vLLM logs at startup instead.
  • On GPUs with native FP8 (Ada, Hopper, Blackwell, MI300-class), start with FP8 weights. Add --kv-cache-dtype fp8 when concurrency is limited by KV memory, and check quality on your own evaluation set.
  • If the model fits on one GPU, run replicas and route by prefix. If it does not, use tensor parallelism inside one NVLink domain, and add pipeline parallelism only across nodes or across GPUs without NVLink. For MLA models, check that TP is not copying the KV cache.
  • Watch the vllm:num_preemptions counter and the Preemptions: field in the stats log. A climbing count means the KV budget is too small for the concurrency you allowed.
  • Put variable text (timestamps, user names, retrieved results) at the end of the prompt, and watch the prefix-cache hit rate.
  • Set --max-num-batched-tokens by what your users see: lower for streaming chat, higher than 8,192 for throughput.
  • Add speculative decoding only after measuring that you are latency-bound at low to medium concurrency, and re-measure at your peak load.
  • Consider disaggregation only for long prompts relative to outputs, at high concurrency, on medium or large models, with an RDMA or NVLink fabric between workers.
  • Benchmark with vllm bench serve for feature comparisons and GuideLLM for production load, and restart the server between runs.

What would change this

The default above is a snapshot. These events would move it, and here is where to watch for them.

  • Disaggregation becomes simpler than aggregation. If Dynamo or llm-d can run prefill and decode pools that size themselves, with no extra configuration surface, the 8-GPU threshold falls and disaggregation becomes a default. Watch the Dynamo releases and llm-d releases.
  • Speculative decoding stops slowing down at high load. Dynamic draft lengths already exist in vLLM, which removes part of this condition. If engines turn speculation on by default with an MTP head shipped in the model, it moves from the deviation table into the defaults. Watch vLLM's speculative decoding docs and new model cards with built-in MTP modules.
  • 4-bit becomes the production default. Native NVFP4 and MXFP4 compute on Blackwell and MI355X removes INT4's unpacking cost. If independent evaluations show 4-bit recovery near FP8's on 30B-class models and below, the format default moves from 8 bits to 4. Watch quantized model evaluations from model vendors and Red Hat's quantization guides.
  • Architectures shrink the KV cache further. MLA already cut it by 93% against DeepSeek's earlier model. Hybrid models with Mamba-style layers, which keep a fixed-size state instead of a growing cache, already ship and run in vLLM and SGLang. If they become the dominant frontier architecture, KV sizing stops being the first question. Watch new open frontier model architectures.
  • A new memory technology changes the bandwidth ratio. Every number in the hardware table is bandwidth-bound. A large jump in memory bandwidth per dollar would move the decode floor more than any software change. Watch high-bandwidth memory (HBM) generations in the next accelerator announcements.

References

Genai

Llms

More Articles

Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:

Books by Ranjan Kumar

Harness Engineering for Production AI Systems cover

Harness Engineering

The 7 GenAI Architectures cover

The 7 GenAI Architectures

Building Real-World Agentic AI Systems with LangGraph cover

Building Real-World Agentic AI Systems

The ChatML Handbook cover

The ChatML Handbook

The Chat Templates Handbook cover

The Chat Templates Handbook

Comments