Open source
Free
- Full inference engine under Apache 2.0, self-hosted on your own hardware
vLLM is an open-source inference and serving engine for LLMs, licensed under Apache 2.0. It started in UC Berkeley's Sky Computing Lab in 2023 and is now a PyTorch Foundation-hosted project. In January 2026, its creators and core maintainers founded Inferact to commercialize vLLM, with continued investment in the open-source project as the company's stated priority.
vLLM manages attention memory with PagedAttention and batches incoming requests continuously, which lets one set of GPUs serve many concurrent users. It supports quantization formats from FP8 to GPTQ, AWQ, and GGUF, speculative decoding, and tensor, pipeline, and expert parallelism for models too large for a single GPU. The server exposes an OpenAI-compatible API, plus the Anthropic Messages API and gRPC, and runs on NVIDIA, AMD, and Intel GPUs, CPUs, Google TPUs, and other accelerators through hardware plugins.
vLLM is a serving engine rather than a platform. You run and scale the servers yourself, it serves only the models you host, and evals, prompt management, and routing to hosted APIs come from other tools. It's also built for GPU servers, which makes it heavier to set up for single-user local work than desktop tools like Ollama or LM Studio.
Core capabilities this platform advertises.
What this tool does well, and the limitations to keep in mind.
Pros
Cons
What's included in each plan, and how the tiers compare.
Free
ML and platform engineers serving open models on their own GPUs to many concurrent users
Respan's gateway can route to a self-hosted vLLM server as a custom model. Register the server's OpenAI-compatible endpoint as a custom provider, and calls to it are traced, priced, and scored in the same place as calls to 1,000+ hosted models.
Top companies in Inference & Compute you can use instead of vLLM.
Side-by-side comparisons with other tools in this category.
Companies from adjacent layers in the AI stack that work well with vLLM.
Register your vLLM server as a custom model in Respan, and every call is traced, priced, and scored next to 1,000+ hosted models. Start free.
llama.cpp
GGUF universal model format (weights + tokenizer + metadata in one file)
llama.cpp vs vLLM