Top
Best
New

Posted by ThouYS 15 hours ago

Flux 3(bfl.ai)
535 points | 125 commentspage 2
pwillia7 10 hours ago|
Awesome -- glad they're going to release the open weight version! I've been waiting for an excuse to re jump into AI OS image gen!
zmmmmm 14 hours ago||
> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.

I'm confused, videos contain images and audio ...?

ibotty 14 hours ago||
That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.
EricBurnett 10 hours ago|||
Video contains images (frames), but not every image would reasonably be found in the frames of a video, or interpreted spatially. In the space of world model synthesis, consider blueprints, relationship diagrams, pages of instructions, sheet music, or a boarding pass.
PxldLtd 13 hours ago||
It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.
AmbroseBierce 7 hours ago||
Amazing. This will be fundamental for the future for robots to distinguish the sound of the poors getting close to Besos/Musk's/Zuckerberg bunkers and quickly adapt to any new kind of attack by the masses, robots will quickly learn to adapt to the behavior of the attackers, quickly infer where they are grouped, their numbers and so for.

Of course there will be feuds from robots of different family groups but they will be minimal as it quickly becomes symmetrical robot conflict with high casualties as they learn too fast from each other, it's likely those will be avoided, it will be after all much easier to confront humans for any given resources.

Truly a pinnacle for technology, albeit perhaps not for mankind.

vrganj 7 hours ago|
When they hide in their bunkers, who's to stop people from pouring concrete down the ventilation shafts?
isoprophlex 3 hours ago||
Or aerosolized LSD...
abdusco 13 hours ago||
I wonder if this also creates people with huge heads and short necks like Flux Klein does.
yangcheng 10 hours ago||
I hope flux will include a 3D generation model. right now the open-weights version of 3D is failing behind closed source by a big margin. Hopefully the improved spatial ability helps with robotics too
saejox 13 hours ago||
i don't get why they are investing money on image/video gen. All generations i have see looked blurry, lacking in fine details and missing the artistic touch (lacks meaning? lifeless?)
mdp2021 13 hours ago||
> why they are investing money

Maybe they presume that after a series of "good enough to some" they may be getting near the Real Thing?

vitalyan8184 10 hours ago||
now look at AI images/videos from 5 years ago.
mattmanser 14 hours ago||
Open-weight plans are near the bottom (Launch section):

    - Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
    - Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
    - Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
    - Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
spidercob 12 hours ago|
[flagged]
rekpero 13 hours ago||
I have a feeling open-weight models ought to be outperforming proprietary ones by now, but that still hasn’t happened. So far, Nano Banana and GPT-2 Image seem to be the best in class, and Flux still isn’t crossing that quality bar.
ex-aws-dude 9 hours ago|
If you think about it why is language/image even separate from video?

Isn’t video + audio all you need?

doubleorseven 8 hours ago|
video is just a group of images (GOP) if the gist is what you're after
More comments...