Apps & Tools
vllm.cpp
A C++ inference engine mirroring vLLM performance with a 140x smaller footprint.
bymudler
Vllm cloneOpen source
Description
A from-scratch C++20 inference engine that implements continuous batching and block-paged KV without Python or PyTorch dependencies. It supports multiple backends including CUDA, Metal, and Vulkan, and provides a flat C ABI for easy integration into C, Go, and Rust applications.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.