
Description
An open-source training infrastructure designed specifically for large Mixture-of-Experts (MoE) models. It optimizes how data is moved across hundreds of GPUs, reportedly achieving training speeds 2.7x faster than previous systems by placing experts more efficiently on individual chips.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.