Top
Best
New

Posted by ilreb 16 hours ago

Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge(qwen.ai)
539 points | 211 commentspage 4
lalith_c 7 hours ago|
didn’t anthropic accuse them of stealing fable?
saltysalt 15 hours ago||
It will be interesting to compare this to Flux 2.
flakiness 5 hours ago||
It's a bit shocking to see them showing off blatantly disinformation-al examples. See the Japanese manga one. It says "原作 監修 三浦建太郎" which says its original is written by, and it is supervised the by, the famous author (who died a few years ago so who can he possibly supervise?) and "共同制作 白泉社" saying: collaborated by a (famous Japanese) publisher.

I hope these stakeholders have good enough layers.

xiaoyu2006 16 hours ago||
The blog write-up style is so casual haha.
jdw64 16 hours ago||
Wow, it displays Korean properly without breaking. But there are still a lot of typos. Haha, it's good that Korean displays properly, but there are a lot of incorrect sentences
sheept 15 hours ago|
I find it mildly interesting how your comment repeats itself but with different phrasing
jdw64 15 hours ago||
It's because of the structure of Korean. I think in Korean first, so I end up translating it directly.

When I want to emphasize something, I tend to repeat it

m3kw9 10 hours ago||
Impressive, but their woman face generation always use very similar, too perfect, same prettiness faces, it's very obvious.
gpjanik 15 hours ago||
The real performance is nowhere close to what is presented in the marketing materials, which is pretty annoying. Especially text rendering and accuracy.

Try asking it for a plot of Polish GDP growth over the past 20 years. It's slop.

viraptor 15 hours ago|
This model isn't supposed to contain all the numerical data. It will give you a (usually) matching graph transformed from one you provide or from a table of information you provide. Or you can pipeline from an LLM doing research on that data first. But expecting an image gen model to get you GDP info has got to be one of the worst possible approaches.

> Especially text rendering

That's true though. I still got some completely fried letters in headings.

yorwba 14 hours ago|||
Giving it numerical data in an LLM-generated prompt doesn't seem to help much: https://imgur.com/a/KFhczOd

It included the table verbatim and even managed to hallucinate a reasonable heading for it, but then the graph doesn't even manage to align the data points with the time axis, leading to an unfortunate collision in the middle.

I guess you should use a traditional graphing library for your presentation slides for now.

spwa4 13 hours ago||
... but if you want an accurate graph, why not ask the LLM model to put the data points into a graphing library?
gpjanik 11 hours ago|||
I am not expecting it, it's the Qwen team is claimingthey can do much harder tasks than this, like rendering a consistent page of a maths paper, or creating true to fact explainers.

They can't.

mahimai 15 hours ago||
interesting
rvz 16 hours ago||
Midjourney already knew that image generation was going to zero. Again yet another reason why the model was never a moat in the first place.
amelius 7 hours ago|
The moat is the training data. But somehow we've collectively decided that it's not.
spwa4 16 hours ago|
Appears to be closed-weights entirely. No word at all on any weights release.
More comments...