Apps & Tools
NInfer-3090
Fast Qwen3.8-27B inference on one RTX 3090 with MTP3 and ReplaySSM.
byDon-Chad
Open source
NInfer-3090 media is blocked
Allow external media to connect to the provider and play this content.
Description
A specialized C++20/CUDA inference engine designed to run Qwen3.8-27B and Qwen3.6 models efficiently on a single 24GB NVIDIA RTX 3090. It supports OpenAI/Anthropic APIs, speculative decoding (MTP3), ReplaySSM state transactions, and reasoning-effort controls for models with deliberation features. The engine includes native Windows and Linux builds with support for vision-enabled artifacts.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.