Apps & Tools
Big Moe on Edge
Run massive MoE models (up to 284B) on low-RAM phones and PCs via flash-streaming.
Open source
Description
An inference engine built on llama.cpp that enables running Mixture-of-Experts (MoE) models far exceeding available RAM. It works by streaming only the required experts from flash storage (UFS/SSD) just-in-time for each token, allowing a 284B DeepSeek V4 model to run on a 12GB phone.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.