← Back to Blog
For: AI Engineers, ML Engineers, Platform Engineers, AI Systems Architects

Choosing an LLM Inference Framework in 2026: vLLM by Default

Pick vLLM, lock down its defaults, and switch only on a condition you can measure. The engine moves your throughput by 13 to 40 percent. Where you rent the GPU moves the price by up to 3.6x.

Updated
#best-framework-for-llm-production#best-llm-inference-engine#vllm#sglang#tensorrt-llm#vllm-vs-sglang#vllm-vs-tensorrt-llm#llama-cpp#ollama#cpu-llm-inference

State of the field as of 2026-09-25.

Rewritten 25 September 2026. The first version of this guide, published in December 2025, recommended Hugging Face TGI to teams already on Hugging Face tooling. Three months later Hugging Face archived TGI. This version replaces the old six-framework tour with one recommended default, the conditions that should move you off it, and GPU prices re-checked on the day above. Recheck it by 25 December 2026, because vLLM alone ships a new minor version every two weeks.

The short answer: use vLLM as the default LLM inference engine for self-hosted open-weight models, change three of its defaults, and switch to SGLang, TensorRT-LLM, llama.cpp, Ollama, MLX or a managed endpoint only when a condition in the deviation table below holds for you.

That recommendation is the failure this guide exists to prevent. In December 2025, "best for teams on the Hugging Face ecosystem" was a reasonable-sounding reason to pick an engine. On 21 March 2026 the repository became read-only, and its README now tells users to move to vLLM, SGLang, llama.cpp or MLX. A team that followed the advice owns a migration it did not plan for. The engine was chosen on a feature list, and feature lists do not show you which projects will still be maintained next year.

The second failure is quieter, and it happens to teams that pick the right engine. vLLM, the engine this guide recommends, binds to every network interface by default and ships with no API key. Between December 2025 and January 2026, Pillar Security's honeypots logged 35,000 attack sessions in a campaign that went after "every unauthenticated vLLM server" and OpenAI-compatible APIs on port 8000, then resold the stolen access.

What this guide decides, and what it leaves out

One decision: which engine you use to serve an open-weight LLM that you host yourself, judged on inputs you can check today. Do you have a GPU, and whose? How many requests arrive at once? How often does the model change? What does your team have time to operate?

Out of scope: which model to run (that is a systems decision of its own, and a small model may be the right answer, as in design patterns for SLM-first systems), training, fine-tuning, and how the engines work inside. For the internals, read Inside the LLM Inference Engine.

If you run TGI today: it still works, and nothing forces a move this week. It gets only minor bug fixes now, so new model architectures and new kernels will not arrive. Plan the move to vLLM, and budget time to recheck output quality. Chat templates render differently across serving engines, so the prompt TGI built for you is not guaranteed to match the prompt vLLM builds.

Make three decisions, in this order, because each one moves cost more than the one after it:

  1. Should you host the model at all? A managed endpoint can make the rest of this guide irrelevant.
  2. Where will you rent the GPU? For an H100-class GPU, the price gap between providers is larger than any gap between engines.
  3. Which engine? This decision matters, but it moves throughput by tens of percent, not by multiples.

Use vLLM. Pin the image tag and the model revision. Bind it to localhost. Put a reverse proxy in front of it that allows only the routes your clients call.

Two versions matter as you read: vLLM v0.30.0 shipped on 22 September 2026, and SGLang v0.5.20 on 18 September.

Why vLLM is the default inference engine

The case does not rest on benchmarks. It rests on adoption, hardware coverage, and how much there is to operate.

Adoption. An August 2026 study of open-source serving code, LLM Serving in the Wild, compared vLLM, SGLang, TensorRT-LLM, LMDeploy and FlashInfer. It found that "vLLM is the most visible framework in popularity and adoption", and that repositories rarely use more than one serving framework. Adoption measured in open-source code is not the same as production market share, and the paper does not claim it is. Production evidence exists as well: Amazon runs the Rufus shopping assistant on vLLM across multiple Trainium nodes, and the core vLLM maintainers raised a $150M seed round in January 2026 to build a company on it. SGLang has the same kind of backing: RadixArk, a company that maintains SGLang, works with Google on its TPU support. So funding does not separate the two leading engines. It separates both of them from a project like TGI, which depended on one sponsor. A stronger signal comes from NVIDIA itself: its packaged inference product, NIM for LLMs, "is built on vLLM", not on NVIDIA's own TensorRT-LLM.

Hardware coverage. vLLM's README lists NVIDIA, AMD and Intel GPUs, x86, ARM and PowerPC CPUs, and plugins for TPU, Gaudi, Ascend and Apple Silicon. If your GPU supplier changes next year, your engine does not have to. Two of those plugins matter for the rest of this guide. On TPUs, Google's GKE documentation now says "vLLM is now the recommended solution for serving LLMs on TPUs in GKE". On Macs, the official vllm-metal plugin, announced on 22 September 2026, brings "vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon".

Sensible performance defaults. Automatic prefix caching reuses the computed attention state for prompts that share a beginning, such as a long system prompt. It is on by default: enable_prefix_caching: bool = True in vllm/config/cache.py at v0.30.0. That fact decides how to read half the benchmarks published this year, which the next section covers.

Benchmarks are close, and that is enough. Jarvislabs ran H100 tests in May 2026 with both engines at their default launch flags, which leaves both prefix caches on. That is the fair setup, although the post does not record engine versions. On Qwen2.5-7B chat traffic with no concurrency cap, vLLM produced 23,523 output tokens per second against SGLang's 16,787. Their summary: "vLLM was the strongest default in our Qwen serving benchmarks." SGLang had lower decode latency on 16K-token prompts, and on unique prompts other tests put the two within a few percent. Neither engine wins everywhere. So the default goes to the one that costs least to run and least to replace.

Why many published comparisons are wrong about prefix caching

Several 2026 vendor benchmarks describe vLLM's prefix caching as "opt-in". Spheron's vLLM vs SGLang comparison labels its table "TTFT reduction vs vLLM no-APC", where APC is automatic prefix caching and TTFT is time to first token. It reports SGLang at 195 ms against vLLM at 310 ms on traffic where 80% of the prompt is shared. Spheron's own article admits that with prefix caching enabled, "the gap narrows to roughly 15-18%" at 10 concurrent requests. RunPod's KV-cache comparison says the same feature needs "manual configuration". Neither claim holds for current vLLM. Before you trust a vLLM-versus-SGLang number, check that the vLLM run did not have --no-enable-prefix-caching set, because only that flag turns the feature off.

The alternatives considered, and why they lost

Every alternative below has a real case. Each one lost only on the question this guide asks: what should a team with no measured reason to choose otherwise run?

First, the same four facts for every engine, read from each project's source, docs or release page on 25 September 2026:

EngineRuns onListens on by defaultConcurrency out of the boxLatest release
vLLMNVIDIA, AMD and Intel GPUs, CPUs, TPU and others via pluginsevery interfaceContinuous batching, prefix cache onv0.30.0, 22 Sep 2026
SGLangNVIDIA, AMD, Intel Xeon, Ascend; Google TPU through SGL-JAX127.0.0.1Continuous batching, radix cache on, first-come-first-served schedulingv0.5.20, 18 Sep 2026
TensorRT-LLMNVIDIA onlylocalhostBatches requests, but without max_tokens they run one after anotherStable v1.2.1, 20 Apr 2026; v1.3.0rc28, 23 Sep 2026
llama.cpp (llama-server)CPU, CUDA, Metal, ROCm, Vulkan127.0.0.1Parallel slots set automatically, continuous batching onb11177, 25 Sep 2026
OllamaNVIDIA, AMD, Apple127.0.0.11 request at a timev0.34.4, 23 Sep 2026
MLX (mlx_lm.server)Apple Silicon onlylocalhostOne request at a time when the KV cache is quantizedmlx-lm 0.31.3, 22 Apr 2026
LMDeployNVIDIA, Ascend, AMD ROCm, macOSevery interface (0.0.0.0)Not measured for this guidev0.17.0, 1 Sep 2026

No single comparison states the finding in the third column. Five of these seven servers listen only on the local machine unless you tell them otherwise. Only LMDeploy and vLLM listen on every interface, and vLLM is the engine this guide recommends. So the safest engine to choose is among the least safe to start with its defaults, which is why the recommendation above changes them.

SGLang

The strongest case. SGLang is the one engine that competes with vLLM across the whole range. Its maintainers describe it as a serving framework "designed to deliver low-latency and high-throughput inference across a wide range of setups, from a single GPU to large distributed clusters", and it runs on NVIDIA, AMD MI300 and MI355 and Ascend NPUs, with Google TPUs served through a separate project, SGL-JAX. RadixArk, a company that maintains SGLang, stands behind it. Its RadixAttention cache stores prompt prefixes as a tree, which suits traffic that branches: agents that fork a conversation, multi-turn chats, many users sharing one long system prompt. In the v0.5.20 release, the cache hit rate on DeepSeek-V4-Flash with a shared system prompt rose from 43.8% to 60.8%, and mean TTFT fell from 1.57 s to 1.07 s. New mixture-of-experts (MoE) models often run on SGLang the day they are released: its README lists day-0 support for DeepSeek-V4 and Kimi K3. It is also the rollout engine inside several reinforcement learning (RL) post-training frameworks, including verl, slime and AReaL. LinkedIn runs it in production for LLM-based ranking in job and people search, presented at GTC 2026 as a 2-3x throughput gain on H100s, with the change contributed back to SGLang.

Why it is not the default. The difference is how security reports get handled. CERT/CC published VU#326070 (CVE-2026-14890, an unauthenticated socket on the expert-parallel path) with "no patch available at this time", and VU#281278 on 30 July 2026, covering six vulnerabilities that include remote code execution through /load_lora_adapter_from_tensors and model-weight theft when no API keys are set. That advisory also says "no patches are available". Some of these need setups that are off by default, such as expert-parallel serving or prefill/decode disaggregation. Others, such as weight theft when no API key is set, apply to a plain launch that anyone can reach. For VU#326070, CERT records that "no response was obtained from the project maintainers". Earlier deserialization flaws (VU#665416) were fixed in 0.5.10. Several of these advisories recommend turning off the SGLANG_USE_PICKLE_IPC setting, and it is still on by default. Its cache-hit gauge, sglang:cache_hit_rate, is overwritten on every prefill batch and reads 0 under mixed traffic, so do not alert on it. And in the Jarvislabs tests its advantage on long prompts came with slower first tokens and lower throughput at saturation.

TensorRT-LLM

The strongest case. NVIDIA's engine usually gets NVIDIA's new kernels first, such as FP4 on Blackwell, wide expert parallelism and new speculative decoding methods, and NVIDIA publishes its Blackwell results on this stack. On a single H100 running Llama 3.3 70B in FP8, Spheron's three-way test measured 2,100 tokens per second against vLLM's 1,850 at 50 concurrent requests, a 13% lead. Setup is also much simpler than it was in 2024: the quick start runs trtllm-serve "nvidia/Qwen3-8B-FP8" against a Hugging Face model directly. NVIDIA's performance overview says the default PyTorch backend "does not require an engine to be built", and the latest release candidate removes the old TensorRT engine-serving path. Both the "weeks of expert setup" this guide used to quote and the 28-minute compile in Spheron's test describe that old path. Baseten runs TensorRT-LLM in production.

Why it is not the default. It runs only on NVIDIA GPUs, and it is not keeping a stable release current. The latest stable release, v1.2.1, came out on 20 April 2026, and its release note says it "Fixed an issue that caused KV cache corruption". Since then there have been 28 release candidates and no stable release. The v1.3.0rc28 notes list 11 known issues, including disaggregated serving that "may return content associated with another prompt in the same request batch". So you choose between staying on a five-month-old stable release and running a release candidate in production. The 1.3 line also collects anonymous telemetry by default: set TRTLLM_NO_USAGE_STATS=1 or pass --no-telemetry if your data residency rules forbid it. The 13% lead also comes with a caveat. Spheron measured it on the old compiled-engine path, and no one has published a same-hardware comparison of the PyTorch backend against current vLLM.

llama.cpp

The strongest case. It is the engine for machines with no data-centre GPU: CPU, Apple Metal, AMD ROCm and Vulkan, all served from GGUF model files (llama.cpp's single-file quantized format) by llama-server, which speaks the OpenAI API. Nothing else in this list runs as well on a laptop or on a server with no GPU at all. It has strong backing too: ggml.ai, the team behind it, joined Hugging Face in February 2026 and still has "full autonomy" over its direction. On a 64-core AMD EPYC 9554 with 12 channels of DDR5, a practitioner benchmark measured 49.6 tokens per second generating from Llama 3.1 8B at Q4_K_M quantization, and 7.1 from the 70B model.

Why it is not the default. Those are single-stream numbers on a CPU with unusually high memory bandwidth, and they leave out prompt processing, which dominates on long retrieval-augmented generation (RAG) prompts. One stream at 50 tokens per second serves a handful of users. Under concurrency the engine batches well, but the server needs tuning: in a December 2025 discussion, batched decoding on an RTX 5090 rose from 295 to 1,430 tokens per second at batch 32, while the user's real llama-server deployment stopped scaling beyond about 4 parallel slots.

Ollama

The strongest case. Ollama is the fastest way to get a model running on a developer's machine: install, pull, run. It manages downloads and quantization, keeps models loaded between requests, and exposes an OpenAI-compatible API. For a single developer, a local tool, or a prototype, nothing is simpler. It binds to 127.0.0.1:11434 by default, which is the safe choice.

Why it is not the default. It is built for one user at a time. Its FAQ sets OLLAMA_NUM_PARALLEL to 1 by default, notes that memory "will scale by OLLAMA_NUM_PARALLEL * OLLAMA_CONTEXT_LENGTH", and returns a 503 once 512 requests are queued. Red Hat, which sells a vLLM-based product, measured a peak of 793 tokens per second from vLLM against 41 from Ollama on an A100 with Llama 3.1 8B. That test ran Ollama 0.9.2 at FP16, not its usual Q4, so treat the size of the gap with care, not its direction. And when teams expose it, they tend to do so without authentication: SentinelLABS and Censys found 175,000 publicly reachable Ollama hosts in 130 countries, with India and Singapore among the largest groups.

MLX

The strongest case. On Apple Silicon, Apple's own framework is the fastest option. A study of local inference on an M2 Ultra, arXiv 2511.05502, found that "MLX achieves the highest sustained generation throughput" among the runtimes it tested. Ollama agrees in practice: its v0.40.0 release candidate, a pre-release on 25 September 2026, makes MLX the default engine on Apple Silicon.

Why it is not the default. It runs only on Apple hardware, and the same study found that Apple runtimes "still trail NVIDIA GPU-based systems such as vLLM". MLX's bundled server is explicit about its limits: its documentation says it "is not recommended for production as it only implements basic security checks". Serving on your own hardware does not remove that risk either: on-device AI ships its attack surface with it. The last mlx-lm release, 0.31.3, was in April 2026.

For several concurrent users on Macs there is now a third route, and it keeps you on the default. vllm-metal runs vllm serve on Apple Silicon, with MLX and Metal doing the execution underneath and vLLM's scheduler and paged KV cache on top. Its first official release is days old (v0.30.0 shipped on 23 September 2026), so pin the version and test it on your own traffic.

LMDeploy

The strongest case. LMDeploy's TurboMind engine is a C++ engine built for quantized models. It is also the engine behind the benchmark number most often misquoted in this field: in BentoML's June 2024 test of Llama 3 70B quantized to 4 bits on one A100, LMDeploy delivered "up to 700 tokens when serving 100 users while keeping the lowest TTFT across all levels". Earlier versions of this guide credited that figure to TensorRT-LLM. It is LMDeploy's. The project ships every two to five weeks (v0.17.0 on 1 September 2026) and has broad support for DeepSeek, Qwen, GLM and InternVL models.

Why it is not the default. Its best independent evidence is from 2024, on older hardware and older models. Its README claims "up to 1.8x higher request throughput than vLLM" without a date or a test setup. It has roughly a tenth of vLLM's GitHub following, and we found few documented production users outside the InternLM and ModelScope ecosystem.

A managed endpoint, instead of any engine

The strongest case. Per-token pricing for open-weight models is now low enough that a self-hosted GPU often cannot compete with it. On 25 September 2026, gpt-oss-120B cost $0.15 per million input tokens and $0.60 per million output tokens on both Together and Fireworks, and $0.10 / $0.50 on Baseten. Llama 3.3 70B cost $1.04 per million, input or output, on Together.

Why it is not the default for this guide. It is not self-hosting, and some teams cannot use it: data that cannot leave a region, models with custom weights, or traffic heavy enough to keep a GPU busy. The next section shows exactly where that line is.

The decision that moves cost more than the engine

The table below shows on-demand prices per GPU-hour, fetched on 25 September 2026.

OfferUnited StatesSingaporeIndia
Azure NC40ads H100 v5 (1x H100 NVL)$6.98$9.07$9.77 (Central India)
Azure NC24ads A100 v4 (1x A100 80GB)$3.67$4.78$5.14 (Central India)
GCP a3-highgpu-8g, per H100$11.06$14.28$11.52 (Mumbai)
E2E Networks H100, per GPU--₹255.55, about $2.67
Lambda H100 SXM, 1 GPU$4.29--

Sources: the Azure retail prices API (read directly), E2E Networks (read directly, before GST), Lambda, and a third-party mirror of GCP list prices updated 21 September 2026, because Google's own page renders its prices in the browser. Rupees are converted at 95.87 per dollar, the Federal Reserve rate for 18 September 2026. We could not verify AWS prices for Mumbai or Singapore, so they are not shown.

Two things stand out. Azure charges more for an H100 in India than in the US. And an Indian GPU cloud rents an H100-class card for about a quarter of what Azure's Indian region charges. The cards are not identical: Azure's is the 94 GB H100 NVL, and E2E lists an H100 80GB.

Now combine prices with throughput. GPUStack measured 6,095 tokens per second, input plus output, serving gpt-oss-120B on one H100 with vLLM v0.10.2 and no tuning. That is an older vLLM, so current numbers are probably higher. At full load around the clock, the ceiling is 6,095 × 86,400 = 526.6 million tokens a day. The managed price depends on your mix of input and output tokens, so we show two mixes. At 1 input token per output token, Together's blended price is (0.15 + 0.60) / 2 = $0.375 per million. At 3 input tokens per output token, it is (3 × 0.15 + 0.60) / 4 = $0.2625. Neither mix is measured from real traffic, so use your own.

  • Azure East US costs $6.98 × 24 = $167.52 a day. Breaking even needs 167.52 / 0.375 = 447 million tokens a day at the 1:1 mix, which is 85% of the GPU's capacity every hour of every day. At the 3:1 mix it needs 638 million, more than this untuned setup can produce. Almost no real traffic holds a GPU that busy around the clock.
  • E2E Networks costs about $2.67 × 24 = $64 a day. Breaking even needs 171 million tokens a day at 1:1, or 32% of capacity, and 244 million at 3:1, or 46%.

Tuning moves both lines. The same GPUStack page reports 16,042 tokens per second on two H100s after tuning, about 8,021 per GPU. That lifts the daily ceiling to 693 million tokens, which puts even the 3:1 Azure break-even within reach at 92% utilization.

These figures count only the GPU. They leave out the engineer who keeps the server running, idle hours, and network egress. The model calls themselves can cost more than the hardware, which is why agentic RAG systems often cost 10x more than they should. Swapping one production engine for another moved throughput by 13% to 40% in the tests cited above. Swapping one rental provider for another moves the price by up to 3.6x. So decide where to rent before you compare engines.

When to deviate

Stay on vLLM unless one of these conditions is true for you. Each condition is something you can check, and each names the engine that wins when it holds.

ConditionSwitch toWhy
You replayed your production traffic against both engines, SGLang showed a higher prefix-cache hit rate, and p50 and p99 TTFT improved at your real concurrency without lower throughput at saturationSGLangIts radix cache pays off on branching and multi-turn traffic, which is exactly what the replay measures
You need a new MoE model on release day, or you need an RL rollout backendSGLangDay-0 model support and RL framework integrations are what its maintainers build for. On TPUs, stay on vLLM, which Google recommends, unless the same replay test favours SGL-JAX
The fleet is NVIDIA only, one model version will stay pinned for 3 months or more, and your own A/B test shows at least 10% lower cost per token at your TTFT p95 targetTensorRT-LLMNVIDIA's kernels arrive there first, and a pinned model makes up for the release candidate churn
nvidia-smi and rocm-smi find no GPUllama.cppIt is the engine designed for CPU inference
Apple Silicon only, one user at a timeMLX, or Ollama 0.40 once it is stableMLX is the fastest runtime on Apple hardware, and Ollama's 0.40 pre-release uses it by default
Apple Silicon only, several concurrent usersvLLM with the vllm-metal pluginIt brings vLLM's scheduler and paged KV cache to Macs. It is new, so pin and test it
One user on a developer machineOllamaSetup takes minutes. Its default serves one request at a time, so raise OLLAMA_NUM_PARALLEL only if memory allows, since memory grows with it
You serve quantized DeepSeek, Qwen or GLM models, and your own A/B test shows TurboMind beating vLLM on your trafficLMDeployQuantized throughput is the one thing its benchmarks show
Your measured sustained load, multiplied by the managed per-token price, costs less than the GPU rentalA managed endpointThe break-even arithmetic in the previous section

One more tool fits here, but it is not an alternative to an engine. NVIDIA Dynamo v1.5.0 runs on top of vLLM, SGLang or TensorRT-LLM. It splits prefill and decode onto separate GPUs and routes each request to the worker that already caches its prefix. Its own documentation says it "is not automatically better. For small models, short prompts, low concurrency, or clusters without a fast KV-transfer fabric, an aggregated deployment is simpler and often faster." As a rule of thumb, consider it at about one full 8-GPU node serving long prompts at high concurrency. Below that, it adds components without adding speed. On Kubernetes, the same question has a second answer: llm-d (v0.9.0, August 2026), a Cloud Native Computing Foundation (CNCF) sandbox project founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA, adds prefix-aware routing and disaggregation above vLLM or SGLang. The same rule of thumb applies to it.

The wrong way and the right way to deploy vLLM

Most tutorials, including the example in vLLM's own Docker documentation, start vLLM like this:

bash
# Wrong: open to every network the host can reach, and pinned to nothingdocker run -d --gpus all -p 8000:8000 \  vllm/vllm-openai:latest \  Qwen/Qwen3-8B --api-key "$VLLM_API_KEY"

This command has three defects.

  1. -p 8000:8000 publishes the port on every interface of the host. Docker's own documentation says that published-port traffic "gets diverted before it goes through the ufw firewall settings", so a firewall rule you wrote for port 8000 does not apply. Running vllm serve without Docker is no safer: at v0.30.0 the launcher binds (args.host or "", args.port), which also means every interface, and api_key defaults to None.
  2. --api-key protects less than it appears to. vLLM's security documentation says the key applies "only for endpoints under the /v1, /v2, /inference, and /cohere path prefixes". /invocations runs the same inference without a key, which the documentation calls "particularly concerning". /pause, /abort_requests and /update_weights are open too. A bug report, #58144, filed on 22 September 2026, treats the /invocations gap as a defect, and a fix is in review in PR #58341. Once it ships, /invocations should return 401, so run the audit below again after you upgrade.
  3. :latest and an unpinned model both float. vLLM releases a minor version about every two weeks, and each recent one removed something: v0.29.0 dropped ten model architectures, and v0.30.0 removed GPTQ activation ordering. Security fixes usually land in new minor versions. In GHSA-x6mc-67gf-chw4, on servers running Qwen2-VL or Qwen3-VL video models, 74 bytes of JSON sent to the unauthenticated /tokenize route pushed server memory from 2,271 MiB to 13,629 MiB, is fixed only in 0.30.0. An older advisory shows the same pattern: CVE-2026-48746, rated critical, was an authentication bypass on the OpenAI API routes in every version from 0.3.0 until 0.22.0 fixed it. You want to control when you upgrade, and a floating tag takes that decision away from you.

You can check your own deployment. This script sends only read-only requests and empty request bodies, so a route that answers has done no work. It needs Python 3.10 or later. If your proxy uses a certificate from an internal certificate authority, point SSL_CERT_FILE at that certificate first, or the script will report the proxy as unreachable:

python
# audit_vllm.py - list the vLLM routes that answer a caller with no API keyimport sysimport urllib.errorimport urllib.requestPROBES = [    ("GET", "/v1/models", None),    ("GET", "/version", None),    ("POST", "/invocations", b"{}"),    ("POST", "/tokenize", b"{}"),]def probe(base_url: str, method: str, path: str, body: bytes | None) -> str:    req = urllib.request.Request(        base_url + path,        data=body,        method=method,        headers={"Content-Type": "application/json"},    )    try:        with urllib.request.urlopen(req, timeout=5) as resp:            code = resp.status    except urllib.error.HTTPError as err:        code = err.code    except urllib.error.URLError as err:        return f"unreachable ({err.reason})"    if code in (401, 403):        return f"blocked ({code})"    if code == 404:        return "not served (404)"    return f"OPEN without a key ({code})"if __name__ == "__main__":    base = sys.argv[1].rstrip("/")    for method, path, body in PROBES:        print(f"{method:<5}{path:<14}{probe(base, method, path, body)}")

Here is the output from a real vLLM v0.30.0 server, the CPU build serving Qwen2.5-0.5B-Instruct, started with --api-key and bound to the loopback port. From another machine you would use the server's address instead:

text
$ python audit_vllm.py http://127.0.0.1:8000GET  /v1/models    blocked (401)GET  /version      OPEN without a key (200)POST /invocations  OPEN without a key (400)POST /tokenize     OPEN without a key (400)

A 400 on the two POST routes means the request reached the handler and failed validation, so the key was never checked. Any code other than 401, 403 or 404 means the route is open. The server's own access log showed one more: GET /metrics returned 200 without the key, so anyone who can reach the port can read your traffic and cache statistics.

The right way changes three things. It binds to localhost, pins both versions, and puts an allowlist in front:

bash
# Right: loopback only, versions pinned, key read from the environmentexport VLLM_API_KEY="$(openssl rand -hex 32)"MODEL_REVISION="<commit hash from the model's Hugging Face page>"docker run -d --name vllm --gpus all --ipc=host \  -p 127.0.0.1:8000:8000 \  -e VLLM_API_KEY \  -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \  vllm/vllm-openai:v0.30.0 \  Qwen/Qwen3-8B --revision "$MODEL_REVISION" \  --max-model-len 16384
nginx
# /etc/nginx/conf.d/vllm.conf - only the routes your clients call get throughupstream vllm {    server 127.0.0.1:8000;    keepalive 32;}server {    listen 443 ssl;    server_name llm.internal.example.com;    ssl_certificate     /etc/nginx/tls/llm.crt;    ssl_certificate_key /etc/nginx/tls/llm.key;    location ~ ^/v1/(chat/completions|completions|models)$ {        proxy_pass http://vllm;        proxy_http_version 1.1;        proxy_set_header Connection "";        proxy_buffering off;      # stream tokens as they are generated        proxy_read_timeout 300s;  # allow long generations        client_max_body_size 20m; # the 1m default rejects long prompts and images    }    location / {        return 403;    }}

vLLM still checks the API key on the three routes the proxy lets through, and the proxy refuses every other route. We ran these exact location rules in nginx, in front of the same CPU vLLM server, listening on a loopback port without TLS. Through the proxy, the audit prints:

text
$ python audit_vllm.py http://127.0.0.1:18080GET  /v1/models    blocked (401)GET  /version      blocked (403)POST /invocations  blocked (403)POST /tokenize     blocked (403)

In the same run, /v1/models with the correct key returned 200 through the proxy, so the allowlist does not break legitimate clients, and /metrics returned 403. Check your own proxy with the script after you deploy it.

Finally, check that prefix caching is doing its job. Prometheus scrapes /metrics on the loopback port, and this query gives the share of prompt tokens served from cache:

code
rate(vllm:prefix_cache_hits_total[5m]) / rate(vllm:prefix_cache_queries_total[5m])

Both counters are measured in tokens, not requests. If the ratio stays near zero on traffic you believe shares a system prompt, the prompts are not identical from the first token. A timestamp or user name placed before the shared instructions is enough to break the match.

Decision flowchart: which LLM inference engine to use

This diagram follows the order this guide argues for: the hosting and rental decisions come before the engine decision, because they move cost more.

mermaid
flowchart TD
    A[Data may leave your region,<br/>and load is below break-even?] -->|Yes| M[Managed endpoint]
    A -->|No| B[NVIDIA, AMD or TPU accelerator?<br/>nvidia-smi or rocm-smi]
    B -->|No| C[Apple Silicon?]
    C -->|No| L[llama.cpp]
    C -->|Yes| G[Several concurrent users?]
    G -->|No| X[MLX]
    G -->|Yes| VM[vLLM with vllm-metal]
    B -->|Yes| D[One user on a<br/>developer machine?]
    D -->|Yes| O[Ollama]
    D -->|No| E[NVIDIA only, one pinned model,<br/>and your A/B test shows 10%+?]
    E -->|Yes| T[TensorRT-LLM]
    E -->|No| F[Replay shows SGLang winning,<br/>or day-0 MoE or RL rollout?]
    F -->|Yes| S[SGLang]
    F -->|No| V[vLLM - the default]

    style A fill:#7B68EE,color:#FFFFFF
    style B fill:#7B68EE,color:#FFFFFF
    style C fill:#7B68EE,color:#FFFFFF
    style D fill:#7B68EE,color:#FFFFFF
    style E fill:#7B68EE,color:#FFFFFF
    style F fill:#7B68EE,color:#FFFFFF
    style G fill:#7B68EE,color:#FFFFFF
    style VM fill:#6BCF7F,color:#2C2C2A
    style V fill:#6BCF7F,color:#2C2C2A
    style M fill:#98D8C8,color:#2C2C2A
    style X fill:#FFA07A,color:#2C2C2A
    style L fill:#FFA07A,color:#2C2C2A
    style O fill:#FFA07A,color:#2C2C2A
    style T fill:#4A90E2,color:#FFFFFF
    style S fill:#4A90E2,color:#FFFFFF

Decision checklist

  • Estimate sustained tokens per day, and compare the GPU's daily rental with the managed per-token price. If the endpoint is cheaper and your data may leave the region, stop here.
  • Price the same GPU from at least one hyperscaler and one specialist provider in your region before you compare engines.
  • Default to vLLM. Pin the image tag (vllm/vllm-openai:v0.30.0, not :latest) and the model --revision.
  • Publish the port on 127.0.0.1 only, set VLLM_API_KEY, and allowlist routes in a reverse proxy. Run audit_vllm.py from another host to confirm.
  • Watch the prefix-cache hit ratio. Put shared instructions at the start of the prompt, before anything that changes per request.
  • Plan an upgrade every 4 to 6 weeks. Security fixes usually ship in new minor versions.
  • Switch engines only when a row in the deviation table matches, and prove it with your own traffic replay. Do not switch because of a vendor benchmark.

What would change this

The default is a snapshot. These are the events that would move it, and where to watch for them.

  • SGLang patches its open CERT advisories and turns SGLANG_USE_PICKLE_IPC off by default. Its operational risk would then match vLLM's, and the choice between them would come down to traffic shape. Watch kb.cert.org VU#281278 and the SGLang releases.
  • A fair benchmark, with prefix caching on in both engines, shows SGLang or TensorRT-LLM ahead by 20% or more on common chat and RAG traffic across several model families. One result on one model does not move a default. Watch independent benchmarks that publish their engine versions and launch flags.
  • The NVIDIA purchase of Hugging Face closes. NVIDIA announced the deal on 3 September 2026. Hugging Face also sponsors llama.cpp, whose main value is running on hardware that is not NVIDIA's. Watch whether its CPU, Metal, ROCm and Vulkan backends stay first-class in the llama.cpp releases.
  • TensorRT-LLM ships a stable 1.3. Its biggest weakness right now is the gap between stable releases. Watch the release page.
  • vLLM's release pace or its funding falters. A two-week cadence with roughly 280 to 315 contributors per release is part of the case for it. Watch the vLLM releases.
  • Managed per-token prices fall further. Every cut moves more workloads to the top branch of the flowchart. Watch the Together, Fireworks and Baseten pricing pages.

References

Genai

Llms

Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:


Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:


Books by Ranjan Kumar

Harness Engineering for Production AI Systems cover

Harness Engineering

The 7 GenAI Architectures cover

The 7 GenAI Architectures

Building Real-World Agentic AI Systems with LangGraph cover

Building Real-World Agentic AI Systems

The ChatML Handbook cover

The ChatML Handbook

The Chat Templates Handbook cover

The Chat Templates Handbook

Comments