Apps & Tools
Inkling · MLX
Run the 975B Inkling multimodal model on Apple Silicon with MLX.
Open source
Description
A from-scratch MLX port and streaming quantizer for Thinking Machines Lab's Inkling model family (up to 975B parameters). It enables running massive MoE models on Mac hardware via 4-bit quantization and REAP expert pruning, featuring custom implementations for hybrid local/global attention, interleaved gate weights, and multimodal vision/audio towers.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.