Apps & Tools
Qwen3.8-Flash-Next-GGUF
Experimental open-weight hybrid attention model with GGUF quants by Unsloth
Playable
Qwen3.8-Flash-Next-GGUF media is blocked
Allow external media to connect to the provider and play this content.
Description
Qwen3.8-Flash-Next is an experimental preview of the architecture for Qwen4, featuring a rethinking of core LLM components. It introduces Hybrid Attention with QSA (Qwen Sparse Attention) and Gated DeltaNet to reduce long-context latency, alongside n-gram embeddings for efficient parameter scaling.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.