Apps & Tools
Qwen3.8-27B Optimized Marlin Serving Stack
High-performance Qwen 3.8 serving stack with Marlin kernels and MTP.
Playable
Qwen3.8-27B Optimized Marlin Serving Stack media is blocked
Allow external media to connect to the provider and play this content.
Description
An optimized implementation of Qwen3.8-27B designed for high-performance CUDA serving, achieving 110+ tok/s on a single A800 80GB GPU. The stack supports AWQ Marlin kernels, MTP speculative decoding, and multimodal vision-text inputs out of the box via vLLM and SGLang.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.