Apps & Tools
vllm.cpp
A lightweight C++ inference engine porting vLLM features with no Python dependency.
bymudler
Vllm cloneOpen source
vllm.cpp media is blocked
Allow external media to connect to the provider and play this content.
Description
vllm.cpp is a from-scratch C++20 inference engine that ports the serving core of vLLM—including continuous batching and paged KV—into a standalone library with no Python or PyTorch requirements. It offers a 140x smaller deployment footprint than the original vLLM while maintaining token-exact parity across over 25 architectures and hardware backends like CUDA, CPU, Metal, and Vulkan.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.