This should make an excellent choice for arbiter in llm-consortium, mercury-2 was pretty good. One of the main drawbacks of the multi-model system is the added latency of the llm judge, but having a model run at 1100tps goes a long a way to alleviate that.
momojo 11 hours ago||
Anyone here use Mercury 2.0? Curious what your experience with the model is.
nowittyusername 11 hours ago|
I used it for testing my voice agent. It was basically what I expected. Good fast model but "generic" or "vanilla" is how i would describe its personality emulation capability as. Gemma models still outperform it in that department. As far as technicals, one thing i found annoying is cash use was not that good, it missed more then i liked, i contacted support and they were fast and responsive and said they were working on that issue, maybe they solved it with 2.5? Anyways, im prolly gonna try 2.5 again see if anything different, but cant deny the speed, thats the biggest thing this company has going for this offering as if you are in the business of classical cascaded voice agent systems, latency is number one priority and this thing is fast....
nostrebored 6 hours ago||
latency, instruction adherence, reliable tool cools, conversationality are all in tension.
it's great when you can get a 170ms ttft. but if you have 700 ms endpointing on the stt side and 300ms ttfb on the voice side, then you haven't really made something super snappy.
WarmWash 13 hours ago||
I'd imagine at this point they are likely an acquisition target if they can get a halfway decent model. I can't imagine having diffusion sub-agents (or sub-sub-agents) in an orchestration wouldn't be beneficial.
refulgentis 10 hours ago|
I'm less bullish, this model and previous are halfway decent and you can do diffusion sub-agents and sub-agent-agents today, and there hasn't been a sea change, or anything noticeable or well-known. Economically, there's ~no moat, diffusion models aren't a mysterious untame-able force.
cevheribozoglan 11 hours ago||
interesting:
Quality: 40% increase in intelligence from Mercury 2. Comparable to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
available : https://openrouter.ai/inception/mercury-2.5
muppetman 6 hours ago||
Oh great a new model annoncemzzzzzz ZZZZZZZZZ
onel 9 minutes ago||
I think we should be happy when this happens as it encourages more players to be active in AI, and us not rely on only two companies
ghshephard 6 hours ago||
Not just a model - It's a diffusion model. Instead of Next Token prediction it builds the entire page at once and then denoises it. Kind of mind blowing when you watch it happen.
muppetman 2 hours ago|||
Ok well that is actually interesting, thank you. Instead of OpenAI beta 6.3pre3 “fairyfloss”
Cilvic 4 hours ago|||
I would have missed this, thanks
ndgold 11 hours ago||
I like the text output on logical and historical content that I sampled so far
ltbarcly3 11 hours ago||
They are comparing it to 2 and 3 version old flash/fast versions of models but purely for tok/s. Then only comparing it to Mercury 2 on intelligence. This is very misleading and I suspect this model is basically useless.
swiftcoder 4 hours ago|
Even pretty dumb models are useful for running subtasks (especially at this sort of speed). Note that in the coding section they only mention using it as a subagent for a smarter model
ashing 6 hours ago||
This token speed is too fast.
casualwriter 7 hours ago|
the output is good and fast. like it, not only fast, but a new architectures.