Multimodal2026-07-24VentureBeat

Black Forest Labs Launches FLUX 3 Multimodal Model

Black Forest Labs has launched FLUX 3, a multimodal frontier model that pushes the boundaries of generative AI by producing images, videos, and audio from a single text prompt. Unlike previous models that specialized in one modality, FLUX 3 can generate coherent 20-second video clips complete with synchronized audio, all derived from a simple textual description. The model extends the company's existing architecture to handle combined audio and video clips, enabling outputs that feel more immersive and natural. For example, a prompt like 'a thunderstorm over a forest at dusk' could yield a video with rolling clouds, flashes of lightning, and the sound of rain and thunder. Currently, FLUX 3 is available in a limited release, suggesting that Black Forest Labs is carefully managing access while gathering feedback. This launch positions the company as a competitor to other multimodal AI leaders like OpenAI and Google, but with a focus on creative and entertainment applications. The ability to generate synchronized audiovisual content from text could revolutionize fields such as advertising, film pre-production, game development, and virtual reality. As the model matures and becomes more widely available, it may redefine how creators approach content generation.

Related news