Apps & Tools
FreeToken
Edge-native MoE serving engine for running massive LLMs on low VRAM GPUs
PlayableOpen source
FreeToken media is blocked
Allow external media to connect to the provider and play this content.
Description
FreeToken is an inference engine designed to run large frontier models like Qwen 3.8 27B on consumer hardware with as little as 4GB VRAM. It features a bandwidth-adaptive hybrid backend, dynamic VRAM re-allocation, and provides a drop-in OpenAI-compatible API endpoint for local agent workflows.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.