Posted by ilreb 14 hours ago
As for the image model, wow...
Impressive.
Btw, what is currently the best model to run locally on a 16GB Vram? Is it Z-Image Turbo?
Any suggestions for the best open, non-opinionated model?
What training mechanism or model architecture provides the glue to go from human text to images?
Don't you need to have millions of really descriptively labelled images?
There are ML models that do the reverse and output image to text, which assist quite a lot.
The better the text represents the unique thing in the photo, the better the model understands what that text means.
Slightly longer answer for older text to image models you teach them how to encode images and text into the same latent space. Then you simply do a conversion, take a text input, put it into latent space and then extract the image that latent space represents.
But: not open-source/open-weights, and no indication that weights/source will be released either.
For example, Google's Imagen 3 usually looked a lot less fake than the newer Imagen 4, but the latter still scored higher on most benchmarks because it made fewer mistakes and had better prompt following capabilities.
A similar thing happened with Dalle-E 2 and Dell-E 3: The new model was better but also more fake looking.
Image token pricing has been fairly steady while text token prices fall, yet image model release discussion seems to be more focused on how beautiful the women the model generates are versus any sort of substantive discussion.