Top
Best
New

Posted by toddmorey 5 hours ago

Choosing an AI model: one prompt, 11 models, different results(www.netlify.com)
115 points | 56 commentspage 2
s4i 4 hours ago|
In my opinion, this kind of a benchmark doesn't tell much about the models' capabilities on normal software development tasks. It's fun to look at the differences in the output of course, but how often would anyone prompt with very brief instructions, without even hinting the model about caring about any of the details in the outcome nor the implementation?

When there are implicit boundaries, negotiable tradeoffs, taste, whatever, in the mix, then the differences in the model capabilities become way more interesting.

tealmoonx 4 hours ago||
You betta lose yoself in the context

It’s not theft, you own it

You betta neva use Go (go) (go)

You only get 1 prompt

Do not use canvas (No!)

Cause opportunity comes once in a lifetime

feor 4 hours ago||
I like the output of the cheaper/older models better, surprisingly, say Gemini 3.1 or DeepSeek V4. They're mostly no frills and just text, and ironically look less AI-generated to me because of the lack of hip slogans and graphics. Definitely closer to what I'd want for my own site, but I don't claim to know what people want from a website for a coffeeshop.
pedrosbmartins 5 hours ago||
Pretty interesting how DeepSeek V4 Flash 0731 has such disparate (and cheap!) results. I would never guess they come from the same model and prompt.
edgyquant 4 hours ago||
Benchmarks are so difficult with ai because as soon as one gets popular it enters the dataset so the next iteration of the model is trained on the solution. I’m not sure if there’s any potential work around here or anyone doing interesting work but would love to hear about it if so
Schlagbohrer 4 hours ago|
But if this effect were really that strong, the models should be getting nearly 100% on the common benchmarks. But for many of the benchmarks, even after being public for more than a year, the new models only get 60-80%.
neom 5 hours ago||
I'd be curious to see Terra xhigh vs Sol low, only in that the visual languages are actually kinda different, that Terra work was a lot less "llm" feeling than a lot of the other designs, wonder if pumping up the effort would result in a more in-depth design but within that style.FWIW I enjoyed reading this way more than any usual benchmark posts we see.
throwa356262 4 hours ago||
This seems to be the easiest way to get on HN front page:

Come up with an arbitrary test, let a bunch of LLMs work on it. Make some very subjective judgement about the result...

jannishan 5 hours ago||
I think we should establish a professional evaluation organization for models; otherwise, it will be difficult for informal evaluations to form standards.
horsawlarway 5 hours ago||
Approaching this from the perspective of a potential customer and not a designer - I find that I actually like the smaller/open model output quite a bit more.

Kimi, GLM, and Deepseek all absolutely run away with the "Can I quickly read the menu and find the address" challenge.

Most of the rest of the pages are stylistic, but hard to parse.

If I were in a car on a mobile phone trying to find the address of the place to meet a friend for coffee... I don't want a bunch of fluff and stylistic design that makes it hard to parse the information on the site.

sorokod 3 hours ago|
Same model over 11 days, one prompt, different results?
More comments...