Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.
I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.
[1] - https://docs.ideogram.ai/using-ideogram/getting-started/prom...
From where I'm sitting, it's just turning python functions into boxes and instead of write the function yourself, you drag from the output of one box to the input of another. For 2 or 3 boxes, this is cool, but I opened up a professional workflow and was taken into a view with 100s of boxes and wires all over the place. Uhh, ok?
For myself, I'd rather just create my own python environment, write some quick pytorch or mlx calls, wire up some cli to it and share that in a GitHub.
It's one of the most uniquely hostile user experiences I've ever had the (dis)pleasure of working with
Chats can be awful user interfaces.
The website mentions:
> FLUX 3 Image is available under a commercial weights license for companies running image generation at scale. Fine-tune and deploy it on your own infrastructure. Reach out to us to learn more.
I guess the open ones would be non-commercial?
"Open Weights version of FLUX 3 Image is launching in the coming weeks."
Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.
I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.