Top
Best
New

Posted by ilreb 14 hours ago

Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge(qwen.ai)
519 points | 207 commentspage 2
timedude 10 hours ago|
Zooming in on mobile on that website causes a large white area to obstruct the page. Might wanna look into that.

As for the image model, wow...

Mashimo 13 hours ago||
> Supports up to 4.5k token input, effortlessly generating complex layouts such as newspapers, storyboards, and exam papers.

Impressive.

Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?

kroaton 44 minutes ago||
It depends on what you need, but Krea/Klein9b/Ideogram4/Z-Image are among the best right now for text2image and Qwen Edit and Klein are probably still the best at editing.
woadwarrior01 13 hours ago||
Krea-2-Turbo. I've even got it working locally on my M5 iPad Pro.
Izmaki 11 hours ago|||
> We implemented safety measures across the full model development lifecycle.

Any suggestions for the best open, non-opinionated model?

woadwarrior01 10 hours ago|||
It is a reasonably non-opinionated model. My usual test is to ask these models to generate comic book and cartoon characters that hosted image generation models refuse to generate. I think that text is just CYA legalese.
Izmaki 5 hours ago||
What's your favourite (online or offline) model for image generation?
bitexploder 9 hours ago|||
You can also ablate it.
Mashimo 13 hours ago|||
thanks mate. Sadly not supported yet by Invoke, but I will take a look.
zzleeper 8 hours ago||
Random question, but has there been any improvement in OCR/document understanding in these newer models? Last time I checked (1mo ago) SOTA was still sadly Gemini, unless you wanted to pay $$$ for e.g. Sol
hawtads 5 hours ago|
Not sure about OCR specifically, but the newer (past quarter) vision language models all have a lot of post training on detecting garbled text specifically. You can feed some of the old stable diffusion outputs into a modern model and they can figure out pretty reliably if the text gen is mangled. I think the feature is probably used as part of the RL for the image gen to correct for bad text rendering.
gchokov 12 hours ago||
It failed to create a simple overlay on a map - something ChatGPT had no issues with.
jcattle 8 hours ago||
What I can not wrap my head around: How are these models trained?

What training mechanism or model architecture provides the glue to go from human text to images?

Don't you need to have millions of really descriptively labelled images?

nucleative 8 hours ago||
That's exactly how they do it.

There are ML models that do the reverse and output image to text, which assist quite a lot.

The better the text represents the unique thing in the photo, the better the model understands what that text means.

vonneumannstan 8 hours ago|||
Short answer yes.

Slightly longer answer for older text to image models you teach them how to encode images and text into the same latent space. Then you simply do a conversion, take a text input, put it into latent space and then extract the image that latent space represents.

mistercheph 8 hours ago||
I found this 3b1b guest video on diffusion helpful: https://www.youtube.com/watch?v=iv-5mZ_9CPY
Oarch 12 hours ago||
I assume Van Gogh didn't paint enough hands to train from!
ninjagoo 11 hours ago||
The examples posted on their launch blog page are quite impressive, especially for fine details, multi-panel/multi-page and text rendering.

But: not open-source/open-weights, and no indication that weights/source will be released either.

pal9000i 13 hours ago||
How long until we get rid of the AI "plasticness" in portrait kind of generated images?
jrs100000 12 hours ago||
The right models and LORAs can get rid of it right now. People apparently really like everyone to look like over exposed over filtered mannequins, so the big companies target that look.
lifeofpi331144 6 hours ago|||
what are the right models and loras?
cubefox 7 hours ago|||
I think it's unintentional. People usually dislike any recognizable "AI look", but they do like other aspects which might have an unrealistic AI look as a side effect.

For example, Google's Imagen 3 usually looked a lot less fake than the newer Imagen 4, but the latter still scored higher on most benchmarks because it made fewer mistakes and had better prompt following capabilities.

A similar thing happened with Dalle-E 2 and Dell-E 3: The new model was better but also more fake looking.

aitchnyu 12 hours ago|||
There is an IRL phenomena, glass skin skincare and makeup, where the person's skin is evenly flat and toned and glossy. Did you think they are more plasticky than IRL models?
numpad0 11 hours ago||
That's supposed to make skin appear lively. Human skins in AI images tend to look clouded, opaque, and overall un-alive, so to speak.
geooff_ 7 hours ago||
No pricing table? No benchmarks? This is just marketingslop.

Image token pricing has been fairly steady while text token prices fall, yet image model release discussion seems to be more focused on how beautiful the women the model generates are versus any sort of substantive discussion.

dsrtslnd23 12 hours ago|
Seems that this will not be open weights?
More comments...