Top
Best
New

Posted by minimaxir 1 day ago

FLUX 3 Image(bfl.ai)
240 points | 54 comments
vunderba 4 hours ago|
One of the things they seem to be emphasizing here is the UX around being able to place specific elements where you want them in an image. If the positions of the components in the overall composition are very important, this seems to make that a lot easier and kind of reminds me of InvokeAI.

Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.

I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.

[1] - https://docs.ideogram.ai/using-ideogram/getting-started/prom...

kranke155 4 hours ago||
You just get an LLM to do the bounding box stuff or use the ComfyUI node that provides a GUI for bounding box generation
CuriouslyC 4 hours ago|||
I always hated Comfy's node based UI, but agents make it tolerable. Now I just have them set up a workflow and I go in and tweak it manually if the results aren't where I want them. I even have agents cherry doing multiple runs and cherry picking the best outputs, models have gotten good enough that it's a real time saver, assuming you have references they can and a rubric to check against.
jarjoura 2 minutes ago|||
It's definitely not for me.

From where I'm sitting, it's just turning python functions into boxes and instead of write the function yourself, you drag from the output of one box to the input of another. For 2 or 3 boxes, this is cool, but I opened up a professional workflow and was taken into a view with 100s of boxes and wires all over the place. Uhh, ok?

For myself, I'd rather just create my own python environment, write some quick pytorch or mlx calls, wire up some cli to it and share that in a GitHub.

swiftcoder 2 hours ago||||
> I always hated Comfy's node based UI

It's one of the most uniquely hostile user experiences I've ever had the (dis)pleasure of working with

hdjrudni 2 hours ago|||
How do you get agents to set up a workflow? You just get them to modify the JSON directly and then import it, or do you have a tighter integration (e.g. in the UI)?
CuriouslyC 2 hours ago||
The agents can interact with Comfy via API pretty well, which afaik ends up being directly with JSON.
vunderba 4 hours ago|||
Yeah, given how much better Ideogram v4 outputs are when you use the proper structured JSON (background, elements, etc.) I think most users probably have stuck some kind of Qwen/Gemma-based LLM between their raw prompt and the CLIP encoder.
NBJack 2 hours ago|||
Adobe has had something like this for a while with Generative Fill. Drag a box, specify contents.
popalchemist 4 hours ago||
Ideogram 4.5 released 2 days ago with more features along these lines

https://www.youtube.com/watch?v=2mecWZgbaEg

arnaudsm 5 hours ago||
The UX looks amazing and very steerable, congrats to the team for focusing on the interface.

Chats can be awful user interfaces.

gAI 3 hours ago||
Love to see an AI lab outside of US/China releasing good models.
vergessenmir 4 hours ago||
I think we are all waiting for the open weights or local model releases.
pixelesque 4 hours ago|
Has that been announced?

The website mentions:

> FLUX 3 Image is available under a commercial weights license for companies running image generation at scale. Fine-tune and deploy it on your own infrastructure. Reach out to us to learn more.

I guess the open ones would be non-commercial?

JimDabell 3 hours ago|||
> Open Weights version of FLUX 3 Image is launching in the coming weeks.

— https://x.com/bfl_ai/status/2105734605621825738

vunderba 4 hours ago|||
If it is anything like the previous release Flux.2 [dev] - then yeah it'll probably be a non-commercial license.

https://bfl.ai/legal/non-commercial-license-terms

LordDragonfang 3 hours ago||
There's been an increasing trend of previously-open-weight models going closed-source once they reach a certain size (and size is proportional to capital investment). That the latest release will be open-weight is not necessarily a given just because the previous one was.
vunderba 2 hours ago|||
It was literally in the announcement from Black Forest Labs:

"Open Weights version of FLUX 3 Image is launching in the coming weeks."

https://nitter.cf/bfl_ai/status/2105734605621825738

trentor 2 hours ago|||
What trend? BFL was always non commercial. The "only" trend would be qwen image not being apache anymore.
KazaNLP 4 hours ago||
Agreed with other comments about the UX. I'm more interested in that than the model itself. Would like to start seeing UI like this where you get to choose the model and compare different models. Can't jump all over the internet to each model developers sandbox just to test their models. Doing it from one place would be nice.
imgbenchdude 1 hour ago|
[dead]
armcat 4 hours ago||
Does anyone know if it can be used to generate accurate frame-by-frame sprite sequences? I found that no image model can do this well (with sufficient fidelity) - neither with one shot (full spritesheet), nor single frame conditioning. It would be great if an imagegen model could do this. What I do now (I use my own tool https://github.com/acatovic/ai-game-studio) is basically generate a reference image, then condition on that image to generate a very short video, then extract and prune frames. Then I get indie-level sprite fidelity about 90% of the time.
popalchemist 3 hours ago|
The task you're describing is a video model task, not an image model task. It's inherently temporal.

Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.

armcat 3 hours ago||
Sure and that's what I do, but a video can be seen as a causal generation on discreet sequence of images, each image conditioned on the one before it. It can also be seen as a series of image editing tasks. It would be cool to get this working in imagegen because of the amount of control you would get. Right now with video generation you can at best specify start and end frame and hope for the best.
popalchemist 2 hours ago||
Image edit models can probably do a grid, but the temporal accuracy / coherence will never match what a video model, which is really a world model, can do.
armcat 1 hour ago||
Regarding your world model statement. This is completely FALSE. Learning the visual statistics of a physical world is NOT the same thing as learning its causal dynamics. The difference is observational likelihood versus intervention-dependent dynamics. There have been great studies disproving video models as world models, like this ICML paper: https://proceedings.mlr.press/v267/kang25g.html. Unfortunately lot of people treat them as world models, mostly because of their ability to reproduce increasingly convincing physical behaviour without ever discovering the underlying physical laws. This is due to many things that I could write an essay about, but better conditioning, latent space represtnation, scaling etc, all make them look awesome.

I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.

popalchemist 1 hour ago||
They are proto world models (lots written about this - flux being an example of a video model whose weights also power world-action-engines used in robots) in that they attempt to model causality in time, the thing that is required for what OP is asking for and which image models will never do because it is out of domain.
skybrian 3 hours ago||
This isn't much of a test, but I bought $10 in credits on their playground and generated a test image. Not bad, but it didn't get the accordion keyboard right. Haven't tried editing yet.

https://pages.skybrian.com/flux3-image-test/

swiftcoder 2 hours ago||
I wish they would work on fixing the "studio lighting" sheen that all these AI-generated humans have
imgbenchdude 1 hour ago||
[dead]
Doohickey-d 2 hours ago|||
I also don't think there is a park that looks like that, in that location relative ton the Eiffel Tower (although I could be wrong).
htx619 2 hours ago||
[dead]
neals 5 hours ago||
It's this a new model or a new ui?
swiftcoder 2 hours ago|
Porque no los dos? Presumably you need a model conditioned on the bounding box input to make effective use of the new UI
minimaxir 3 hours ago||
Of note is the OpenRouter endpoint has a promotional 50% discount, which is rare on image models: https://openrouter.ai/black-forest-labs/flux-3-image
vunderba 2 hours ago||
I think that’s less a product of OpenRouter’s generosity and more a result of promotional pricing coming directly from BFL, since other third-party vendors have it as well (Fal.ai, etc.).

https://bfl.ai/pricing

pwillia7 3 hours ago||
open weights or bust
Trufa 4 hours ago|
So much negativity as usual and so little talk about the product, this is pretty impressive, well done, it seems to be filling decently a gap that everyone that has worked enough generating images with AI has faced.
doctorpangloss 3 hours ago|
there basically isn't any authentic use for image generation, i would hardly say the negativity is unfounded
More comments...