Black Forest Labs' multimodal foundation model unifies image, video, audio and action prediction in a single architecture. GA release enables text-to-video, image-to-video, video continuation and keyframe control with optional synchronized audio (speech, effects, ambience). Supports draft mode for fast HD previews that can be enhanced to full FHD quality. Available via BFL API and dashboard on pay-as-you-go basis.