I always enjoy these comparisons between models, especially when they demonstrate the actual costs in addition to the outputs.
sorokod 4 hours ago||
Same model over 11 days, one prompt, different results?
plumb_samji 6 hours ago||
Interesting exploratory comparison, but I be cautious about treating it as a model benchmark
With only three runs per model, the results are highly sensitive to randomness
nullzzz 3 hours ago||
This would be interesting, if I was into building coffee shop sites and todo apps from scratch. For the HN crowd tho, I’d say these are toy examples. No offense!
chrisjj 5 hours ago||
> Vector graphics actually require a lot of work from the models
How so? Surely they can just steal such generic graphics off existing web sites.