Top
Best
New

Posted by Retro_Dev 16 hours ago

Mercury 2.5 LLM hits 770 tokens per second(artificialanalysis.ai)
133 points | 77 commentspage 2
sharktheone 13 hours ago|
this feels like "we got the same benches as gpt-oss-120b but are also potentially slower while saying it is great"
whalesalad 11 hours ago||
I used this a few days ago and thought something must be wrong with how fast it was responding. "Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price." this is so funny. So when you have a stupid model that is fast - what do you use it for?
entrope 12 hours ago||
> Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price.

Well priced when compared to other models of similar price, eh?

Are we allowed to call this slop, even if the output is not directly from an LLM?

verdverm 9 hours ago|
The LLMs learned it somewhere...
Unified-Mentor 1 hour ago||
[dead]
pushpendraw 8 hours ago||
[dead]
rvz 16 hours ago|
The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.
copperx 15 hours ago||
Ah, the old "good, fast, or cheap; pick two" proves true once again.
downrightmike 15 hours ago||
Give it a few months.
timClicks 13 hours ago|||
It means something, because it an iterative workflow. If you're willing to burn tokens, it's possible for weaker models to implement tasks by incrementally improving drafts.
glouwbug 15 hours ago|||
Some of us want fast food
voiceeh 14 hours ago||
Not if your use case needs speed. For one of my products I can't use an LLM that has a p99 of >700ms for TTFT.
hansvm 14 hours ago||
If it could output 1k tokens per second but needed 4 seconds to produce the first batch of 4k, would that not be viable?