Black Forest Labs has released FLUX 3, a new multimodal AI model that adds video generation and robotics capabilities to its existing text and image tools. The company, known for its work in generative AI, says the model is designed to handle multiple data types within a single framework.
Multimodal expansion
FLUX 3 builds on earlier versions by integrating video generation and robotics control. The model can produce short video clips from text prompts and interpret visual data for robotic systems. This marks a shift from purely text-and-image models toward broader real-world applications.
The company did not disclose specific performance benchmarks or training data details. But the release signals growing competition in the multimodal AI space, where companies like OpenAI and Google have also introduced models that combine language, vision, and action.
Video generation and robotics
The video generation feature allows users to create short clips by describing scenes in natural language. Early demonstrations show the model handling simple motion and object interactions. For robotics, FLUX 3 can process visual input and generate commands for tasks like object manipulation or navigation.
Black Forest Labs has not yet announced commercial partnerships or a public API for the robotics component. The video generation tool is available through the company's developer platform, with pricing based on usage.
The launch positions Black Forest Labs to compete in the growing market for AI systems that go beyond text and images. The company has not set a timeline for expanding FLUX 3's capabilities or releasing a larger version. Developers and researchers can now access the model through the company's website.




