State of the field as of 2026-09-25.
Rewritten 25 September 2026. The first version of this guide, published in December 2025, recommended Hugging Face TGI to teams already on Hugging Face tooling. Three months later Hugging Face archived TGI. This version replaces the old six-framework tour with one recommended default, the conditions that should move you off it, and GPU prices re-checked on the day above. Recheck it by 25 December 2026, because vLLM alone ships a new minor version every two weeks.
The short answer: use vLLM as the default LLM inference engine for self-hosted open-weight models, change three of its defaults, and switch to SGLang, TensorRT-LLM, llama.cpp, Ollama, MLX or a managed endpoint only when a condition in the deviation table below holds for you.
That recommendation is the failure this guide exists to prevent. In December 2025, "best for teams on the Hugging Face ecosystem" was a reasonable-sounding reason to pick an engine. On 21 March 2026 the repository became read-only, and its README now tells users to move to vLLM, SGLang, llama.cpp or MLX. A team that followed the advice owns a migration it did not plan for. The engine was chosen on a feature list, and feature lists do not show you which projects will still be maintained next year.
The second failure is quieter, and it happens to teams that pick the right engine. vLLM, the engine this guide recommends, binds to every network interface by default and ships with no API key. Between December 2025 and January 2026, Pillar Security's honeypots logged 35,000 attack sessions in a campaign that went after "every unauthenticated vLLM server" and OpenAI-compatible APIs on port 8000, then resold the stolen access.
What this guide decides, and what it leaves out
One decision: which engine you use to serve an open-weight LLM that you host yourself, judged on inputs you can check today. Do you have a GPU, and whose? How many requests arrive at once? How often does the model change? What does your team have time to operate?
Out of scope: which model to run (that is a systems decision of its own, and a small model may be the right answer, as in design patterns for SLM-first systems), training, fine-tuning, and how the engines work inside. For the internals, read Inside the LLM Inference Engine.
If you run TGI today: it still works, and nothing forces a move this week. It gets only minor bug fixes now, so new model architectures and new kernels will not arrive. Plan the move to vLLM, and budget time to recheck output quality. Chat templates render differently across serving engines, so the prompt TGI built for you is not guaranteed to match the prompt vLLM builds.
Make three decisions, in this order, because each one moves cost more than the one after it:
- Should you host the model at all? A managed endpoint can make the rest of this guide irrelevant.
- Where will you rent the GPU? For an H100-class GPU, the price gap between providers is larger than any gap between engines.
- Which engine? This decision matters, but it moves throughput by tens of percent, not by multiples.
The recommended default LLM inference engine: vLLM, with its defaults changed
Use vLLM. Pin the image tag and the model revision. Bind it to localhost. Put a reverse proxy in front of it that allows only the routes your clients call.
Two versions matter as you read: vLLM v0.30.0 shipped on 22 September 2026, and SGLang v0.5.20 on 18 September.
Why vLLM is the default inference engine
The case does not rest on benchmarks. It rests on adoption, hardware coverage, and how much there is to operate.
Adoption. An August 2026 study of open-source serving code, LLM Serving in the Wild, compared vLLM, SGLang, TensorRT-LLM, LMDeploy and FlashInfer. It found that "vLLM is the most visible framework in popularity and adoption", and that repositories rarely use more than one serving framework. Adoption measured in open-source code is not the same as production market share, and the paper does not claim it is. Production evidence exists as well: Amazon runs the Rufus shopping assistant on vLLM across multiple Trainium nodes, and the core vLLM maintainers raised a $150M seed round in January 2026 to build a company on it. SGLang has the same kind of backing: RadixArk, a company that maintains SGLang, works with Google on its TPU support. So funding does not separate the two leading engines. It separates both of them from a project like TGI, which depended on one sponsor. A stronger signal comes from NVIDIA itself: its packaged inference product, NIM for LLMs, "is built on vLLM", not on NVIDIA's own TensorRT-LLM.
Hardware coverage. vLLM's README lists NVIDIA, AMD and Intel GPUs, x86, ARM and PowerPC CPUs, and plugins for TPU, Gaudi, Ascend and Apple Silicon. If your GPU supplier changes next year, your engine does not have to. Two of those plugins matter for the rest of this guide. On TPUs, Google's GKE documentation now says "vLLM is now the recommended solution for serving LLMs on TPUs in GKE". On Macs, the official vllm-metal plugin, announced on 22 September 2026, brings "vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon".
Sensible performance defaults. Automatic prefix caching reuses the computed attention state for prompts that share a beginning, such as a long system prompt. It is on by default: enable_prefix_caching: bool = True in vllm/config/cache.py at v0.30.0. That fact decides how to read half the benchmarks published this year, which the next section covers.
Benchmarks are close, and that is enough. Jarvislabs ran H100 tests in May 2026 with both engines at their default launch flags, which leaves both prefix caches on. That is the fair setup, although the post does not record engine versions. On Qwen2.5-7B chat traffic with no concurrency cap, vLLM produced 23,523 output tokens per second against SGLang's 16,787. Their summary: "vLLM was the strongest default in our Qwen serving benchmarks." SGLang had lower decode latency on 16K-token prompts, and on unique prompts other tests put the two within a few percent. Neither engine wins everywhere. So the default goes to the one that costs least to run and least to replace.
Why many published comparisons are wrong about prefix caching
Several 2026 vendor benchmarks describe vLLM's prefix caching as "opt-in". Spheron's vLLM vs SGLang comparison labels its table "TTFT reduction vs vLLM no-APC", where APC is automatic prefix caching and TTFT is time to first token. It reports SGLang at 195 ms against vLLM at 310 ms on traffic where 80% of the prompt is shared. Spheron's own article admits that with prefix caching enabled, "the gap narrows to roughly 15-18%" at 10 concurrent requests. RunPod's KV-cache comparison says the same feature needs "manual configuration". Neither claim holds for current vLLM. Before you trust a vLLM-versus-SGLang number, check that the vLLM run did not have --no-enable-prefix-caching set, because only that flag turns the feature off.
The alternatives considered, and why they lost
Every alternative below has a real case. Each one lost only on the question this guide asks: what should a team with no measured reason to choose otherwise run?
First, the same four facts for every engine, read from each project's source, docs or release page on 25 September 2026:
| Engine | Runs on | Listens on by default | Concurrency out of the box | Latest release |
|---|---|---|---|---|
| vLLM | NVIDIA, AMD and Intel GPUs, CPUs, TPU and others via plugins | every interface | Continuous batching, prefix cache on | v0.30.0, 22 Sep 2026 |
| SGLang | NVIDIA, AMD, Intel Xeon, Ascend; Google TPU through SGL-JAX | 127.0.0.1 | Continuous batching, radix cache on, first-come-first-served scheduling | v0.5.20, 18 Sep 2026 |
| TensorRT-LLM | NVIDIA only | localhost | Batches requests, but without max_tokens they run one after another | Stable v1.2.1, 20 Apr 2026; v1.3.0rc28, 23 Sep 2026 |
llama.cpp (llama-server) | CPU, CUDA, Metal, ROCm, Vulkan | 127.0.0.1 | Parallel slots set automatically, continuous batching on | b11177, 25 Sep 2026 |
| Ollama | NVIDIA, AMD, Apple | 127.0.0.1 | 1 request at a time | v0.34.4, 23 Sep 2026 |
MLX (mlx_lm.server) | Apple Silicon only | localhost | One request at a time when the KV cache is quantized | mlx-lm 0.31.3, 22 Apr 2026 |
| LMDeploy | NVIDIA, Ascend, AMD ROCm, macOS | every interface (0.0.0.0) | Not measured for this guide | v0.17.0, 1 Sep 2026 |
No single comparison states the finding in the third column. Five of these seven servers listen only on the local machine unless you tell them otherwise. Only LMDeploy and vLLM listen on every interface, and vLLM is the engine this guide recommends. So the safest engine to choose is among the least safe to start with its defaults, which is why the recommendation above changes them.
SGLang
The strongest case. SGLang is the one engine that competes with vLLM across the whole range. Its maintainers describe it as a serving framework "designed to deliver low-latency and high-throughput inference across a wide range of setups, from a single GPU to large distributed clusters", and it runs on NVIDIA, AMD MI300 and MI355 and Ascend NPUs, with Google TPUs served through a separate project, SGL-JAX. RadixArk, a company that maintains SGLang, stands behind it. Its RadixAttention cache stores prompt prefixes as a tree, which suits traffic that branches: agents that fork a conversation, multi-turn chats, many users sharing one long system prompt. In the v0.5.20 release, the cache hit rate on DeepSeek-V4-Flash with a shared system prompt rose from 43.8% to 60.8%, and mean TTFT fell from 1.57 s to 1.07 s. New mixture-of-experts (MoE) models often run on SGLang the day they are released: its README lists day-0 support for DeepSeek-V4 and Kimi K3. It is also the rollout engine inside several reinforcement learning (RL) post-training frameworks, including verl, slime and AReaL. LinkedIn runs it in production for LLM-based ranking in job and people search, presented at GTC 2026 as a 2-3x throughput gain on H100s, with the change contributed back to SGLang.
Why it is not the default. The difference is how security reports get handled. CERT/CC published VU#326070 (CVE-2026-14890, an unauthenticated socket on the expert-parallel path) with "no patch available at this time", and VU#281278 on 30 July 2026, covering six vulnerabilities that include remote code execution through /load_lora_adapter_from_tensors and model-weight theft when no API keys are set. That advisory also says "no patches are available". Some of these need setups that are off by default, such as expert-parallel serving or prefill/decode disaggregation. Others, such as weight theft when no API key is set, apply to a plain launch that anyone can reach. For VU#326070, CERT records that "no response was obtained from the project maintainers". Earlier deserialization flaws (VU#665416) were fixed in 0.5.10. Several of these advisories recommend turning off the SGLANG_USE_PICKLE_IPC setting, and it is still on by default. Its cache-hit gauge, sglang:cache_hit_rate, is overwritten on every prefill batch and reads 0 under mixed traffic, so do not alert on it. And in the Jarvislabs tests its advantage on long prompts came with slower first tokens and lower throughput at saturation.
TensorRT-LLM
The strongest case. NVIDIA's engine usually gets NVIDIA's new kernels first, such as FP4 on Blackwell, wide expert parallelism and new speculative decoding methods, and NVIDIA publishes its Blackwell results on this stack. On a single H100 running Llama 3.3 70B in FP8, Spheron's three-way test measured 2,100 tokens per second against vLLM's 1,850 at 50 concurrent requests, a 13% lead. Setup is also much simpler than it was in 2024: the quick start runs trtllm-serve "nvidia/Qwen3-8B-FP8" against a Hugging Face model directly. NVIDIA's performance overview says the default PyTorch backend "does not require an engine to be built", and the latest release candidate removes the old TensorRT engine-serving path. Both the "weeks of expert setup" this guide used to quote and the 28-minute compile in Spheron's test describe that old path. Baseten runs TensorRT-LLM in production.
Why it is not the default. It runs only on NVIDIA GPUs, and it is not keeping a stable release current. The latest stable release, v1.2.1, came out on 20 April 2026, and its release note says it "Fixed an issue that caused KV cache corruption". Since then there have been 28 release candidates and no stable release. The v1.3.0rc28 notes list 11 known issues, including disaggregated serving that "may return content associated with another prompt in the same request batch". So you choose between staying on a five-month-old stable release and running a release candidate in production. The 1.3 line also collects anonymous telemetry by default: set TRTLLM_NO_USAGE_STATS=1 or pass --no-telemetry if your data residency rules forbid it. The 13% lead also comes with a caveat. Spheron measured it on the old compiled-engine path, and no one has published a same-hardware comparison of the PyTorch backend against current vLLM.
llama.cpp
The strongest case. It is the engine for machines with no data-centre GPU: CPU, Apple Metal, AMD ROCm and Vulkan, all served from GGUF model files (llama.cpp's single-file quantized format) by llama-server, which speaks the OpenAI API. Nothing else in this list runs as well on a laptop or on a server with no GPU at all. It has strong backing too: ggml.ai, the team behind it, joined Hugging Face in February 2026 and still has "full autonomy" over its direction. On a 64-core AMD EPYC 9554 with 12 channels of DDR5, a practitioner benchmark measured 49.6 tokens per second generating from Llama 3.1 8B at Q4_K_M quantization, and 7.1 from the 70B model.
Why it is not the default. Those are single-stream numbers on a CPU with unusually high memory bandwidth, and they leave out prompt processing, which dominates on long retrieval-augmented generation (RAG) prompts. One stream at 50 tokens per second serves a handful of users. Under concurrency the engine batches well, but the server needs tuning: in a December 2025 discussion, batched decoding on an RTX 5090 rose from 295 to 1,430 tokens per second at batch 32, while the user's real llama-server deployment stopped scaling beyond about 4 parallel slots.
Ollama
The strongest case. Ollama is the fastest way to get a model running on a developer's machine: install, pull, run. It manages downloads and quantization, keeps models loaded between requests, and exposes an OpenAI-compatible API. For a single developer, a local tool, or a prototype, nothing is simpler. It binds to 127.0.0.1:11434 by default, which is the safe choice.
Why it is not the default. It is built for one user at a time. Its FAQ sets OLLAMA_NUM_PARALLEL to 1 by default, notes that memory "will scale by OLLAMA_NUM_PARALLEL * OLLAMA_CONTEXT_LENGTH", and returns a 503 once 512 requests are queued. Red Hat, which sells a vLLM-based product, measured a peak of 793 tokens per second from vLLM against 41 from Ollama on an A100 with Llama 3.1 8B. That test ran Ollama 0.9.2 at FP16, not its usual Q4, so treat the size of the gap with care, not its direction. And when teams expose it, they tend to do so without authentication: SentinelLABS and Censys found 175,000 publicly reachable Ollama hosts in 130 countries, with India and Singapore among the largest groups.
MLX
The strongest case. On Apple Silicon, Apple's own framework is the fastest option. A study of local inference on an M2 Ultra, arXiv 2511.05502, found that "MLX achieves the highest sustained generation throughput" among the runtimes it tested. Ollama agrees in practice: its v0.40.0 release candidate, a pre-release on 25 September 2026, makes MLX the default engine on Apple Silicon.
Why it is not the default. It runs only on Apple hardware, and the same study found that Apple runtimes "still trail NVIDIA GPU-based systems such as vLLM". MLX's bundled server is explicit about its limits: its documentation says it "is not recommended for production as it only implements basic security checks". Serving on your own hardware does not remove that risk either: on-device AI ships its attack surface with it. The last mlx-lm release, 0.31.3, was in April 2026.
For several concurrent users on Macs there is now a third route, and it keeps you on the default. vllm-metal runs vllm serve on Apple Silicon, with MLX and Metal doing the execution underneath and vLLM's scheduler and paged KV cache on top. Its first official release is days old (v0.30.0 shipped on 23 September 2026), so pin the version and test it on your own traffic.
LMDeploy
The strongest case. LMDeploy's TurboMind engine is a C++ engine built for quantized models. It is also the engine behind the benchmark number most often misquoted in this field: in BentoML's June 2024 test of Llama 3 70B quantized to 4 bits on one A100, LMDeploy delivered "up to 700 tokens when serving 100 users while keeping the lowest TTFT across all levels". Earlier versions of this guide credited that figure to TensorRT-LLM. It is LMDeploy's. The project ships every two to five weeks (v0.17.0 on 1 September 2026) and has broad support for DeepSeek, Qwen, GLM and InternVL models.
Why it is not the default. Its best independent evidence is from 2024, on older hardware and older models. Its README claims "up to 1.8x higher request throughput than vLLM" without a date or a test setup. It has roughly a tenth of vLLM's GitHub following, and we found few documented production users outside the InternLM and ModelScope ecosystem.
A managed endpoint, instead of any engine
The strongest case. Per-token pricing for open-weight models is now low enough that a self-hosted GPU often cannot compete with it. On 25 September 2026, gpt-oss-120B cost $0.15 per million input tokens and $0.60 per million output tokens on both Together and Fireworks, and $0.10 / $0.50 on Baseten. Llama 3.3 70B cost $1.04 per million, input or output, on Together.
Why it is not the default for this guide. It is not self-hosting, and some teams cannot use it: data that cannot leave a region, models with custom weights, or traffic heavy enough to keep a GPU busy. The next section shows exactly where that line is.
The decision that moves cost more than the engine
The table below shows on-demand prices per GPU-hour, fetched on 25 September 2026.
| Offer | United States | Singapore | India |
|---|---|---|---|
| Azure NC40ads H100 v5 (1x H100 NVL) | $6.98 | $9.07 | $9.77 (Central India) |
| Azure NC24ads A100 v4 (1x A100 80GB) | $3.67 | $4.78 | $5.14 (Central India) |
| GCP a3-highgpu-8g, per H100 | $11.06 | $14.28 | $11.52 (Mumbai) |
| E2E Networks H100, per GPU | - | - | ₹255.55, about $2.67 |
| Lambda H100 SXM, 1 GPU | $4.29 | - | - |
Sources: the Azure retail prices API (read directly), E2E Networks (read directly, before GST), Lambda, and a third-party mirror of GCP list prices updated 21 September 2026, because Google's own page renders its prices in the browser. Rupees are converted at 95.87 per dollar, the Federal Reserve rate for 18 September 2026. We could not verify AWS prices for Mumbai or Singapore, so they are not shown.
Two things stand out. Azure charges more for an H100 in India than in the US. And an Indian GPU cloud rents an H100-class card for about a quarter of what Azure's Indian region charges. The cards are not identical: Azure's is the 94 GB H100 NVL, and E2E lists an H100 80GB.
Now combine prices with throughput. GPUStack measured 6,095 tokens per second, input plus output, serving gpt-oss-120B on one H100 with vLLM v0.10.2 and no tuning. That is an older vLLM, so current numbers are probably higher. At full load around the clock, the ceiling is 6,095 × 86,400 = 526.6 million tokens a day. The managed price depends on your mix of input and output tokens, so we show two mixes. At 1 input token per output token, Together's blended price is (0.15 + 0.60) / 2 = $0.375 per million. At 3 input tokens per output token, it is (3 × 0.15 + 0.60) / 4 = $0.2625. Neither mix is measured from real traffic, so use your own.
- Azure East US costs $6.98 × 24 = $167.52 a day. Breaking even needs 167.52 / 0.375 = 447 million tokens a day at the 1:1 mix, which is 85% of the GPU's capacity every hour of every day. At the 3:1 mix it needs 638 million, more than this untuned setup can produce. Almost no real traffic holds a GPU that busy around the clock.
- E2E Networks costs about $2.67 × 24 = $64 a day. Breaking even needs 171 million tokens a day at 1:1, or 32% of capacity, and 244 million at 3:1, or 46%.
Tuning moves both lines. The same GPUStack page reports 16,042 tokens per second on two H100s after tuning, about 8,021 per GPU. That lifts the daily ceiling to 693 million tokens, which puts even the 3:1 Azure break-even within reach at 92% utilization.
These figures count only the GPU. They leave out the engineer who keeps the server running, idle hours, and network egress. The model calls themselves can cost more than the hardware, which is why agentic RAG systems often cost 10x more than they should. Swapping one production engine for another moved throughput by 13% to 40% in the tests cited above. Swapping one rental provider for another moves the price by up to 3.6x. So decide where to rent before you compare engines.
When to deviate
Stay on vLLM unless one of these conditions is true for you. Each condition is something you can check, and each names the engine that wins when it holds.
| Condition | Switch to | Why |
|---|---|---|
| You replayed your production traffic against both engines, SGLang showed a higher prefix-cache hit rate, and p50 and p99 TTFT improved at your real concurrency without lower throughput at saturation | SGLang | Its radix cache pays off on branching and multi-turn traffic, which is exactly what the replay measures |
| You need a new MoE model on release day, or you need an RL rollout backend | SGLang | Day-0 model support and RL framework integrations are what its maintainers build for. On TPUs, stay on vLLM, which Google recommends, unless the same replay test favours SGL-JAX |
| The fleet is NVIDIA only, one model version will stay pinned for 3 months or more, and your own A/B test shows at least 10% lower cost per token at your TTFT p95 target | TensorRT-LLM | NVIDIA's kernels arrive there first, and a pinned model makes up for the release candidate churn |
nvidia-smi and rocm-smi find no GPU | llama.cpp | It is the engine designed for CPU inference |
| Apple Silicon only, one user at a time | MLX, or Ollama 0.40 once it is stable | MLX is the fastest runtime on Apple hardware, and Ollama's 0.40 pre-release uses it by default |
| Apple Silicon only, several concurrent users | vLLM with the vllm-metal plugin | It brings vLLM's scheduler and paged KV cache to Macs. It is new, so pin and test it |
| One user on a developer machine | Ollama | Setup takes minutes. Its default serves one request at a time, so raise OLLAMA_NUM_PARALLEL only if memory allows, since memory grows with it |
| You serve quantized DeepSeek, Qwen or GLM models, and your own A/B test shows TurboMind beating vLLM on your traffic | LMDeploy | Quantized throughput is the one thing its benchmarks show |
| Your measured sustained load, multiplied by the managed per-token price, costs less than the GPU rental | A managed endpoint | The break-even arithmetic in the previous section |
One more tool fits here, but it is not an alternative to an engine. NVIDIA Dynamo v1.5.0 runs on top of vLLM, SGLang or TensorRT-LLM. It splits prefill and decode onto separate GPUs and routes each request to the worker that already caches its prefix. Its own documentation says it "is not automatically better. For small models, short prompts, low concurrency, or clusters without a fast KV-transfer fabric, an aggregated deployment is simpler and often faster." As a rule of thumb, consider it at about one full 8-GPU node serving long prompts at high concurrency. Below that, it adds components without adding speed. On Kubernetes, the same question has a second answer: llm-d (v0.9.0, August 2026), a Cloud Native Computing Foundation (CNCF) sandbox project founded by Red Hat, Google Cloud, IBM Research, CoreWeave and NVIDIA, adds prefix-aware routing and disaggregation above vLLM or SGLang. The same rule of thumb applies to it.
The wrong way and the right way to deploy vLLM
Most tutorials, including the example in vLLM's own Docker documentation, start vLLM like this:
# Wrong: open to every network the host can reach, and pinned to nothingdocker run -d --gpus all -p 8000:8000 \ vllm/vllm-openai:latest \ Qwen/Qwen3-8B --api-key "$VLLM_API_KEY"This command has three defects.
-p 8000:8000publishes the port on every interface of the host. Docker's own documentation says that published-port traffic "gets diverted before it goes through the ufw firewall settings", so a firewall rule you wrote for port 8000 does not apply. Runningvllm servewithout Docker is no safer: at v0.30.0 the launcher binds(args.host or "", args.port), which also means every interface, andapi_keydefaults toNone.--api-keyprotects less than it appears to. vLLM's security documentation says the key applies "only for endpoints under the/v1,/v2,/inference, and/coherepath prefixes"./invocationsruns the same inference without a key, which the documentation calls "particularly concerning"./pause,/abort_requestsand/update_weightsare open too. A bug report, #58144, filed on 22 September 2026, treats the/invocationsgap as a defect, and a fix is in review in PR #58341. Once it ships,/invocationsshould return 401, so run the audit below again after you upgrade.:latestand an unpinned model both float. vLLM releases a minor version about every two weeks, and each recent one removed something: v0.29.0 dropped ten model architectures, and v0.30.0 removed GPTQ activation ordering. Security fixes usually land in new minor versions. In GHSA-x6mc-67gf-chw4, on servers running Qwen2-VL or Qwen3-VL video models, 74 bytes of JSON sent to the unauthenticated/tokenizeroute pushed server memory from 2,271 MiB to 13,629 MiB, is fixed only in 0.30.0. An older advisory shows the same pattern: CVE-2026-48746, rated critical, was an authentication bypass on the OpenAI API routes in every version from 0.3.0 until 0.22.0 fixed it. You want to control when you upgrade, and a floating tag takes that decision away from you.
You can check your own deployment. This script sends only read-only requests and empty request bodies, so a route that answers has done no work. It needs Python 3.10 or later. If your proxy uses a certificate from an internal certificate authority, point SSL_CERT_FILE at that certificate first, or the script will report the proxy as unreachable:
# audit_vllm.py - list the vLLM routes that answer a caller with no API keyimport sysimport urllib.errorimport urllib.requestPROBES = [ ("GET", "/v1/models", None), ("GET", "/version", None), ("POST", "/invocations", b"{}"), ("POST", "/tokenize", b"{}"),]def probe(base_url: str, method: str, path: str, body: bytes | None) -> str: req = urllib.request.Request( base_url + path, data=body, method=method, headers={"Content-Type": "application/json"}, ) try: with urllib.request.urlopen(req, timeout=5) as resp: code = resp.status except urllib.error.HTTPError as err: code = err.code except urllib.error.URLError as err: return f"unreachable ({err.reason})" if code in (401, 403): return f"blocked ({code})" if code == 404: return "not served (404)" return f"OPEN without a key ({code})"if __name__ == "__main__": base = sys.argv[1].rstrip("/") for method, path, body in PROBES: print(f"{method:<5}{path:<14}{probe(base, method, path, body)}")Here is the output from a real vLLM v0.30.0 server, the CPU build serving Qwen2.5-0.5B-Instruct, started with --api-key and bound to the loopback port. From another machine you would use the server's address instead:
$ python audit_vllm.py http://127.0.0.1:8000GET /v1/models blocked (401)GET /version OPEN without a key (200)POST /invocations OPEN without a key (400)POST /tokenize OPEN without a key (400)A 400 on the two POST routes means the request reached the handler and failed validation, so the key was never checked. Any code other than 401, 403 or 404 means the route is open. The server's own access log showed one more: GET /metrics returned 200 without the key, so anyone who can reach the port can read your traffic and cache statistics.
The right way changes three things. It binds to localhost, pins both versions, and puts an allowlist in front:
# Right: loopback only, versions pinned, key read from the environmentexport VLLM_API_KEY="$(openssl rand -hex 32)"MODEL_REVISION="<commit hash from the model's Hugging Face page>"docker run -d --name vllm --gpus all --ipc=host \ -p 127.0.0.1:8000:8000 \ -e VLLM_API_KEY \ -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \ vllm/vllm-openai:v0.30.0 \ Qwen/Qwen3-8B --revision "$MODEL_REVISION" \ --max-model-len 16384# /etc/nginx/conf.d/vllm.conf - only the routes your clients call get throughupstream vllm { server 127.0.0.1:8000; keepalive 32;}server { listen 443 ssl; server_name llm.internal.example.com; ssl_certificate /etc/nginx/tls/llm.crt; ssl_certificate_key /etc/nginx/tls/llm.key; location ~ ^/v1/(chat/completions|completions|models)$ { proxy_pass http://vllm; proxy_http_version 1.1; proxy_set_header Connection ""; proxy_buffering off; # stream tokens as they are generated proxy_read_timeout 300s; # allow long generations client_max_body_size 20m; # the 1m default rejects long prompts and images } location / { return 403; }}vLLM still checks the API key on the three routes the proxy lets through, and the proxy refuses every other route. We ran these exact location rules in nginx, in front of the same CPU vLLM server, listening on a loopback port without TLS. Through the proxy, the audit prints:
$ python audit_vllm.py http://127.0.0.1:18080GET /v1/models blocked (401)GET /version blocked (403)POST /invocations blocked (403)POST /tokenize blocked (403)In the same run, /v1/models with the correct key returned 200 through the proxy, so the allowlist does not break legitimate clients, and /metrics returned 403. Check your own proxy with the script after you deploy it.
Finally, check that prefix caching is doing its job. Prometheus scrapes /metrics on the loopback port, and this query gives the share of prompt tokens served from cache:
rate(vllm:prefix_cache_hits_total[5m]) / rate(vllm:prefix_cache_queries_total[5m])
Both counters are measured in tokens, not requests. If the ratio stays near zero on traffic you believe shares a system prompt, the prompts are not identical from the first token. A timestamp or user name placed before the shared instructions is enough to break the match.
Decision flowchart: which LLM inference engine to use
This diagram follows the order this guide argues for: the hosting and rental decisions come before the engine decision, because they move cost more.
flowchart TD
A[Data may leave your region,<br/>and load is below break-even?] -->|Yes| M[Managed endpoint]
A -->|No| B[NVIDIA, AMD or TPU accelerator?<br/>nvidia-smi or rocm-smi]
B -->|No| C[Apple Silicon?]
C -->|No| L[llama.cpp]
C -->|Yes| G[Several concurrent users?]
G -->|No| X[MLX]
G -->|Yes| VM[vLLM with vllm-metal]
B -->|Yes| D[One user on a<br/>developer machine?]
D -->|Yes| O[Ollama]
D -->|No| E[NVIDIA only, one pinned model,<br/>and your A/B test shows 10%+?]
E -->|Yes| T[TensorRT-LLM]
E -->|No| F[Replay shows SGLang winning,<br/>or day-0 MoE or RL rollout?]
F -->|Yes| S[SGLang]
F -->|No| V[vLLM - the default]
style A fill:#7B68EE,color:#FFFFFF
style B fill:#7B68EE,color:#FFFFFF
style C fill:#7B68EE,color:#FFFFFF
style D fill:#7B68EE,color:#FFFFFF
style E fill:#7B68EE,color:#FFFFFF
style F fill:#7B68EE,color:#FFFFFF
style G fill:#7B68EE,color:#FFFFFF
style VM fill:#6BCF7F,color:#2C2C2A
style V fill:#6BCF7F,color:#2C2C2A
style M fill:#98D8C8,color:#2C2C2A
style X fill:#FFA07A,color:#2C2C2A
style L fill:#FFA07A,color:#2C2C2A
style O fill:#FFA07A,color:#2C2C2A
style T fill:#4A90E2,color:#FFFFFF
style S fill:#4A90E2,color:#FFFFFF
Decision checklist
- Estimate sustained tokens per day, and compare the GPU's daily rental with the managed per-token price. If the endpoint is cheaper and your data may leave the region, stop here.
- Price the same GPU from at least one hyperscaler and one specialist provider in your region before you compare engines.
- Default to vLLM. Pin the image tag (
vllm/vllm-openai:v0.30.0, not:latest) and the model--revision. - Publish the port on
127.0.0.1only, setVLLM_API_KEY, and allowlist routes in a reverse proxy. Runaudit_vllm.pyfrom another host to confirm. - Watch the prefix-cache hit ratio. Put shared instructions at the start of the prompt, before anything that changes per request.
- Plan an upgrade every 4 to 6 weeks. Security fixes usually ship in new minor versions.
- Switch engines only when a row in the deviation table matches, and prove it with your own traffic replay. Do not switch because of a vendor benchmark.
What would change this
The default is a snapshot. These are the events that would move it, and where to watch for them.
- SGLang patches its open CERT advisories and turns
SGLANG_USE_PICKLE_IPCoff by default. Its operational risk would then match vLLM's, and the choice between them would come down to traffic shape. Watch kb.cert.org VU#281278 and the SGLang releases. - A fair benchmark, with prefix caching on in both engines, shows SGLang or TensorRT-LLM ahead by 20% or more on common chat and RAG traffic across several model families. One result on one model does not move a default. Watch independent benchmarks that publish their engine versions and launch flags.
- The NVIDIA purchase of Hugging Face closes. NVIDIA announced the deal on 3 September 2026. Hugging Face also sponsors llama.cpp, whose main value is running on hardware that is not NVIDIA's. Watch whether its CPU, Metal, ROCm and Vulkan backends stay first-class in the llama.cpp releases.
- TensorRT-LLM ships a stable 1.3. Its biggest weakness right now is the gap between stable releases. Watch the release page.
- vLLM's release pace or its funding falters. A two-week cadence with roughly 280 to 315 contributors per release is part of the case for it. Watch the vLLM releases.
- Managed per-token prices fall further. Every cut moves more workloads to the top branch of the flowchart. Watch the Together, Fireworks and Baseten pricing pages.
References
- Hugging Face. text-generation-inference (archived 21 March 2026). https://github.com/huggingface/text-generation-inference
- Cohen, E., Fogel, A. (2026). Operation Bizarre Bazaar: First Attributed LLMjacking Campaign with Commercial Marketplace Monetization. Pillar Security. https://www.pillar.security/blog/operation-bizarre-bazaar-first-attributed-llmjacking-campaign-with-commercial-marketplace-monetization
- Majidi, F., Morovati, M. M., Khomh, F., Li, H. (2026). LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs. arXiv:2608.03036. https://arxiv.org/abs/2608.03036
- AWS Machine Learning Blog (2025). How Amazon scaled Rufus by building multi-node inference using AWS Trainium chips and vLLM. https://aws.amazon.com/blogs/machine-learning/how-amazon-scaled-rufus-by-building-multi-node-inference-using-aws-trainium-chips-and-vllm/
- Deutscher, M. (2026). Inferact launches with $150M in funding to commercialize vLLM. SiliconANGLE. https://siliconangle.com/2026/01/22/inferact-launches-150m-funding-commercialize-vllm/
- vLLM. Release v0.30.0 (22 September 2026). https://github.com/vllm-project/vllm/releases/tag/v0.30.0
- vLLM. Security documentation, v0.30.0. https://docs.vllm.ai/en/v0.30.0/usage/security/
- vLLM.
vllm/config/cache.pyat v0.30.0. https://raw.githubusercontent.com/vllm-project/vllm/v0.30.0/vllm/config/cache.py - vLLM. Security advisory GHSA-x6mc-67gf-chw4. https://github.com/vllm-project/vllm/security/advisories/GHSA-x6mc-67gf-chw4
- vLLM. Security advisory GHSA-94f4-hr76-p5j6 (CVE-2026-48746), OpenAI API auth bypass. https://github.com/vllm-project/vllm/security/advisories/GHSA-94f4-hr76-p5j6
- Kwon, W. et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. SOSP. https://arxiv.org/abs/2309.06180
- SGLang. README and release v0.5.20 (18 September 2026). https://github.com/sgl-project/sglang/releases/tag/v0.5.20
- SGLang. Issue #26608, cache hit rate gauge. https://github.com/sgl-project/sglang/issues/26608
- LMSYS (2026). SGLang at GTC 2026. https://www.lmsys.org/blog/2026-03-25-gtc2026/
- CERT/CC. VU#665416, VU#326070, VU#281278. https://kb.cert.org/vuls/id/665416 ; https://kb.cert.org/vuls/id/326070 ; https://kb.cert.org/vuls/id/281278
- Tonde, J. (2026). SGLang vs vLLM: H100 Benchmarks, with TensorRT-LLM. Jarvislabs. https://jarvislabs.ai/blog/vllm-sglang-trtllm-comparison
- Spheron (2026). vLLM vs SGLang 2026. https://www.spheron.network/blog/vllm-vs-sglang-2026/
- Spheron (2026). vLLM vs TensorRT-LLM vs SGLang: H100 benchmarks. https://www.spheron.network/blog/vllm-vs-tensorrt-llm-vs-sglang-benchmarks/
- RunPod (2026). When to choose SGLang over vLLM. https://www.runpod.io/blog/sglang-vs-vllm-kv-cache
- NVIDIA. TensorRT-LLM Quick Start Guide. https://nvidia.github.io/TensorRT-LLM/quick-start-guide.html
- NVIDIA. TensorRT-LLM releases (v1.2.1, v1.3.0rc28). https://github.com/NVIDIA/TensorRT-LLM/releases
- Baseten (2025). High performance ML inference with NVIDIA TensorRT. https://www.baseten.co/blog/high-performance-ml-inference-with-nvidia-tensorrt/
- Hugging Face (2026). GGML and llama.cpp join Hugging Face. https://huggingface.co/blog/ggml-joins-hf
- Huang, J. (2026). NVIDIA to Acquire Hugging Face. NVIDIA Blog. https://blogs.nvidia.com/blog/nvidia-to-acquire-hugging-face/
- neoX (2025). LLM inference benchmarks with llama.cpp and AMD EPYC 9554 CPU. https://ahelpme.com/ai/llm-inference-benchmarks-with-llamacpp-with-amd-epyc-9554-cpu/
- ggml-org. llama.cpp discussion #18308. https://github.com/ggml-org/llama.cpp/discussions/18308
- Ollama. FAQ. https://docs.ollama.com/faq
- Ollama. Releases. https://github.com/ollama/ollama/releases
- Umesh, H. (2025). Ollama vs. vLLM: A deep dive into performance benchmarking. Red Hat Developer. https://developers.redhat.com/articles/2025/08/08/ollama-vs-vllm-deep-dive-performance-benchmarking
- Lakshmanan, R. (2026). Researchers Find 175,000 Publicly Exposed Ollama AI Servers. The Hacker News. https://thehackernews.com/2026/01/researchers-find-175000-publicly.html
- Rajesh, V. et al. (2025). Production-Grade Local LLM Inference on Apple Silicon. arXiv:2511.05502. https://arxiv.org/abs/2511.05502
- Apple. mlx-lm server documentation. https://github.com/ml-explore/mlx-lm/blob/main/mlx_lm/SERVER.md
- Jiang, B., Zhao, L., Zhou, R., Sheng, S. (2024). Benchmarking LLM Inference Backends. BentoML. https://www.bentoml.com/blog/benchmarking-llm-inference-backends
- InternLM. LMDeploy releases. https://github.com/InternLM/lmdeploy/releases
- NVIDIA. Dynamo releases and disaggregated serving guide. https://github.com/ai-dynamo/dynamo/releases ; https://docs.nvidia.com/dynamo/user-guides/disaggregated-serving
- GPUStack. Optimizing GPT-OSS-120B throughput on NVIDIA H100. https://docs.gpustack.ai/latest/performance-lab/gpt-oss-120b/h100/
- Docker. Packet filtering and firewalls. https://docs.docker.com/engine/network/packet-filtering-firewalls/
- vLLM. Using Docker, v0.30.0. https://docs.vllm.ai/en/v0.30.0/deployment/docker/
- vLLM. Issue #58144,
--api-keyis bypassed by/invocations, and PR #58341. https://github.com/vllm-project/vllm/issues/58144 ; https://github.com/vllm-project/vllm/pull/58341 - vLLM (2026). Announcing vllm-metal: Concurrent Serving on Apple Silicon. https://vllm.ai/blog/2026-09-22-vllm-metal-v0-28-0 ; releases https://github.com/vllm-project/vllm-metal/releases
- Google Cloud. Serve Gemma using TPUs on GKE with JetStream (note recommending vLLM on TPUs). https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/serve-gemma-tpu-jetstream
- NVIDIA. NVIDIA NIM for LLM and VLM Documentation. https://docs.nvidia.com/nim/large-language-models/latest/about-nim-llm/overview.html
- LMSYS (2026). RadixArk Joins Forces with Google to Bring Full SGLang Features to TPUs. https://www.lmsys.org/blog/2026-07-30-sglang-google-tpu/
- llm-d. Repository and releases (v0.9.0). https://github.com/llm-d/llm-d
- ggml-org. llama.cpp releases. https://github.com/ggml-org/llama.cpp/releases
- Pricing pages, fetched 25 September 2026: Azure retail prices API https://prices.azure.com/api/retail/prices ; E2E Networks https://www.e2enetworks.com/gpus/nvidia-h100 ; Lambda https://lambda.ai/pricing ; GCP list prices via https://gcloud-compute.com/ ; Together https://www.together.ai/pricing ; Fireworks https://docs.fireworks.ai/serverless/pricing ; Baseten https://www.baseten.co/pricing/
- Federal Reserve Bank of St. Louis. DEXINUS, Indian rupees to one US dollar. https://fred.stlouisfed.org/series/DEXINUS
Related Articles
- Building Agents That Remember: State Management in Multi-Agent AI Systems
- Building Production-Ready Agentic AI: The Infrastructure Nobody Talks About
- Inside the LLM Inference Engine: Architecture, Optimizations, Tools, Key Concepts and Best Practices
- Stop Pasting Screenshots: How AI Engineers Document Systems with Mermaid




