Apps & Tools
Big Moe on Edge
Run massive MoE models (up to 284B) on low-RAM phones and PCs via flash-streaming.
Open source
Big Moe on Edge media is blocked
Allow external media to connect to the provider and play this content.
Description
An inference engine built on llama.cpp that enables running Mixture-of-Experts (MoE) models far exceeding available RAM. It works by streaming only the required experts from flash storage (UFS/SSD) just-in-time for each token, allowing a 284B DeepSeek V4 model to run on a 12GB phone.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.