FLUX 3: Generation Unified. Storage Didn't.
Summary
Black Forest Labs' FLUX 3 is one network that jointly learns images, video, and audio: 20 second clips with sound in a single pass, preference wins over Runway Gen-4.5 (77%) and Luma Ray 3.2 (93%), and the same backbone powers FLUX-mimic, a video-action model for robots built with mimic robotics. This video walks the announcement, then the downstream gap: generation collapsed into one architecture while most storage stacks still split video, audio, and images into three systems that cannot search each other. Mixpeek is the warehouse side of that pipe: every modality indexed together, so generated media stays findable. Announcement: bfl.ai/blog/flux-3
About this video
Black Forest Labs' FLUX 3 is one network that jointly learns images, video, and audio: 20 second clips with sound in a single pass, preference wins over Runway Gen-4.5 (77%) and Luma Ray 3.2 (93%), and the same backbone powers FLUX-mimic, a video-action model for robots built with mimic robotics. This video walks the announcement, then the downstream gap: generation collapsed into one architecture while most storage stacks still split video, audio, and images into three systems that cannot search each other. Mixpeek is the warehouse side of that pipe: every modality indexed together, so generated media stays findable. Announcement: bfl.ai/blog/flux-3