Posted by espeed 21 hours ago
For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.
I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
But also, everyone says "it's the harness" and almost nobody ever gives good examples, it gets a bit tiring to read everywhere, as if everyone wants to sell a harness to us.
Co-Authored By: Haiku 4.5
If I ask something like: "I'm building a simple X as a learning exercise, I'm writing the code so please only answer the question I'm asking and don't try to solve the problem directly. How does ..." There's a 30% chance it starts reading and writing code immediately and a 20% chance it argues with a "design decision" that will bite me in the non-existent future of my learning exercise. If I ask a follow up question, naively assuming that the context from my original question still stands without repeating, it will almost assuredly start making modifications to my code.
But here’s the standard question: At what speeds/other limiting factors?
Yeah, expecting the world when all you have is a 8GB graphics card? You're going to be disappointed.
16GB is table stakes (IQ3_XSS). 32 GB is better.
Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?
I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.
I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.
So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?
The problem is they can't fit any frontier level open models.
I'm asking, not arguing, because I'd like to understand. Is Astra so much more expensive for those tasks, and are they frequent?
Where'd you get 30 years from? Show your work.
How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.
I'm getting really sick of the constant rot and "magic breakthrough" cycle, so im going full local, at expense on paper but being able to trust something which I need to understand the reliability of is priceless.
I like predictable. I'll take slightly less capable over unreliably capable, since with reliable i can calibrate my expectations and learn what aspects of my workflows to entrust and trust it will work. You simply cannot do that with models you don't control and in my experience they will all errode after the initual marketing wave passes, likely you eventually get fed heavily quantized versions and are expected to accept degraded service when what convinced you to pay was a far superior product. No such issues with local.
Yet… even Altman called out Anthropic for serving dumbed down models.
Shits weird man
Even Altman called out Anthropic? Isn't Anthropic the biggest competitor Sam Altman has?
Not by a mile. They're even (probably illegally and there's apparently a class action lawsuit oncoming: at least something to that extent was posted on HN today) teaming up, as a duopoly, to push for the same bullshit regulations / "we need to slow down AI research".
The reason they're teaming up is the real competition is, as in many other domains, China.
Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.
This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.
Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?
Sorry if it’s a dumb question, I don’t really know much about the topic.
Opus 5 came out with better benchmark results than Fable, but it really did not feel better to use at all.
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
Bis repetitae.
Where?
I see so many accusations of this happening and it's so easy to check, but nobody ever proves it.
I imagine some people have their own personal in depth benchmarks they could do this for.
Those software settings are being changed in real-time by an algorithm and those algorithms are being tweaked and A/B tested daily by the ~~performance~~ revenue optimization teams. On the hardware side the footprint a particular model is running on is materially changing, growing or being re-distributed across DCs ~weekly.
Your benchmark didn't get 100 ? It's normal, it's not deterministic, and also your harness is wrong, and also you didn't do it when US users were offline, and also you got it wrong, and also we don't care about your results, the hivemind is speaking louder than you (also our bots are spamming more than you and drowning you out).
This very website has, at all times, a group of people saying "<Previous model> was never good enough for coding, but <current model> is the best thing and a game changer!" while the other goes "<current model> bad, <previous model> was better!". It's all vibes.
GPT5.6-Sol on Max thinking just became regarded as of a few days ago.
The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).
The cycle repeats.
We are being A/B tested on and there is nothing you can do about it.
Oh, yes there is. DeepSeek 4.1 Flash on max thinking can simply be dropped into Claude Code. Close your eyes as the chain-of-thought traffic scrolls by and you can easily fool yourself into thinking you're still running Opus, in terms of both cognition and throughput.
To be fair, matching Opus's throughput costs about as much as a new car, but cars suck nowadays and you didn't want a new one anyway, right...? Failing that, rent a cloud server, one that you control.
Thanks for your insight
They might but multiple competitors engaging in ongoing deception as an intentional corporate strategy isn't required to explain what we're seeing. It's entirely possible to get the same clearly unethical outcome without any employees knowingly participating in an explicitly unethical plan of record.
Instead it happens without overt coordination when individuals and groups within an org each pursue their local metrics and incentives. In isolation, no individual action seems obviously unethical on its own. They just look like 'optimizing performance', 'maintaining ASP or ARPU targets' or 'achieving operating margin', etc. Customers are still getting deceived and receiving less for their money than they think. The difference is most of the people involved in enabling it get to not feel bad about themselves.
Sam Altman believes he can train a model entirely on synthetic data, which he admits would not have human world knowledge but is interesting none the less, which likely led to their mathematical models.
Overall models have become cheaper to run and smarter per token.
Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.
Change is fast and abundant, and at the same time, it is hilariously more mundane than the dangerous-AGI-in-six-months tune we've been reading daily for years.
I would define it as a quick-moving market, but not nearly moving enough for the fantastic claims they make to justify ever-increasing funding.
The best case scenario is a situation like Y2K: a ton of people coordinate and work hard to produce no perceptible change, because unlike catastrophe, averting catastrophe feels boring.
Absolutely none of this points to “stagnant.” Stagnant is a terrible description of the AI industry.
I’m not saying it is stagnant. I’m saying for a hypothetical industry that was (maybe that fits AI, maybe not, I have zero authority to say myself)…
The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).
For my use cases, we are definitely on the flatter part of the curve at the moment.
In addition to the dozens of opaque model parameters and hardware variables that can nerf or buff model intelligence, speed and profit, there's also the very real possibility that models aren't just training on benchmarks but could be evaluating if they are being benchmarked in real-time and applying more resources adaptively. 'Driver optimizations' that detected benchmarks in real-time were deployed in the first 'GPU Wars'.
> I have no idea is the actual frontier is stagnating.
Like a lot of complex, rapidly evolving tech, the truth is it's probably rapidly accelerating on some measures for a few and stagnating on many others for most - hence the divergence in user reports. It's depends on how you use it, for what problems, how rigorously you assess the output and whether you happen to be on a server bank, RAM pool or shard at this moment which hasn't yet been sufficiently 'cost optimized' by the margin algorithms. They don't call them load balancers anymore. They're Margin Balancers.
Fable seems to be following the same enshittification arc of other Anthropic models.
Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.
I have no way to judge myself. I don’t even use them. But it’s interesting to follow by just reading stories and comments
On the other hand, New Coke was a resounding success. Well, it, itself wasn't, but in the aftermath, Coke outsold Pepsi 2:1.
Kind of reminds me of that, but with more smoke and mirrors
I would do a test to verify my suspicions.
I’m actually so far removed from this tech that I couldn’t run such a test myself lol
HN: Well, must be a stagnant industry...
Not unless your competitors do the same, or else you will only be perceived as falling behind others.
If one lab makes a breakthrough, can the other labs just distill to bear parity anyways and then set a new baseline industry wide.
I’m quite ignorant on this topic, so if any of this sounds moronic, forgive me
Could be what happens next, though.
A very small raise in prices may cause a very large loss in customers that you risk never getting back.
For example if customers figure out that the Chinese models are just as good, they are gone because they are so much cheaper.
Are you really saying AI is a stagnant industry?
Unlike you, however...
Otherwise, even if one company did something like this, everyone would notice because the other companies would be pulling ahead. Are all the AI developers coordinating a "dumbing down" of models? i.e. Are Open AI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, and so on, all working together?
So we might be approaching some limit (the "there's only so round a sphere can get" argument). But I very much doubt there is some massive conspiracy between all the AI developers.
That was one of the main points of the movie 'A Beautiful Mind' - that actors can coordinate without any explicit communication.
This isn't what we see in benchmarks.
Yes, the AI technology is known primarily for how stagant it is.
I have no idea, just had a thought and put it out there
there's also probably load balancers that downgrade models during high use.
AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.
There are several projects that repeat benchmarks on published models. None has ever found significant fluctations
Here's one example https://marginlab.ai/trackers/claude-code/
Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.
This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.
If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.
Click the part at the end that says "View historical performance". They wait to collect more data about a new model before adding it to the overall charts.
The overall solution rate continues to climb when new models are considered.
> If you look at the graphs, the latest benchmarks are showing a pretty significant dip, and they match pretty well with some horrible experiences I've had in recent weeks. You can see token usage steadily going down, matching exactly what the author measured on his own.
The y-axis is amplified to make differences look larger than they are.
Hover over the dots to see the confidence interval. A 1-2% change means nothing.
30-day average 83% [71-91], last result is 79% [66-88], and it dipped to 75% [61-85] a week ago.
> We always use the latest available Claude Code release and the SOTA model
They've stepped up the benchmark for each new model. Presumably they're gathering Fable data now.
They also link to Anthropic's public blog post about some degradations, their cause, and how they fixed them. The time period sounds like the "recent weeks" you experienced: https://www.anthropic.com/engineering/a-postmortem-of-three-...
The Anthropic post points to the latest fix on Sept 12, and the issues mentioned also only affected Sonnet 4 and Haiku. Opus was misbehaving just last week. You are choosing to not see the evidence of degration, 85% -> 75% is a generational dip in intelligence.
A lot of these daily-benchmarking sites popped up earlier this year. Most of them have faded away after they all failed to produce the smoking gun that everyone expected. This site survived because it kind of caught a dip one day, maybe.
The results are really rather flat and daily benchmark runs are expensive, so most of these projects give up after a while.
A properly run OWM enables you to stop wondering if you're being ripped off and that's what AI needs because I think it's incredibly easy to just assume we're being ripped off because these companies are all built on a foundation of wonton theft. (Not that I really care about that — I think all information should be free, but still.)
> 12. General terms
> Changes to the Services. Our Services are novel and will change. We may sometimes add or remove features, increase or decrease capacity limits, offer new Services, or stop offering certain Services.
> Unless we specifically agree otherwise in a separate agreement with you, we reserve the right to modify, suspend, or discontinue the Services or your access to the Services, in whole or in part, at any time without notice to you. Although we will strive to provide you with reasonable advance notice if we stop offering a Service, there may be urgent situations—such as preventing abuse, responding to legal requirements, or addressing security and operability issues—where providing advance notice is not feasible. We will not be liable for any change to or any suspension or discontinuation of the Services or your access to them.
You're not buying a gallon of milk or a pound of flour. You're buying hosted software that the host reserves the right to modify.
The only regulation that we need right now is the model that's on tap
... But what exactly is the "weight" metric you have in mind?
Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."
The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.
Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.
Self run infrastructure won’t have this cost but you have to manage the capacity and rollouts yourself, at which point it’s more obvious what’s happening, but the effects will be the same. The not knowing makes it harder, but also harder to plan your own work around.
If this is duress of competition or at gunpoint of regulators is up for the population to decide.
Honestly I lost patience with Anthropic both clearly messing around with things like this and their agitation over regulation. They aren't good actors, and quite why so many blindly trust them with their company crown jewels is a mystery.
Opus 5.5 is being served under opus 5 right now.
On what basis are you claiming this?
However, I believe that runtime model quantization is possible with some publicly-available inference engines (e.g. vLLM), so its not beyond belief that the closed labs do quantize at runtime, either to allocate compute, or to nudge users towards a preferred model (e.g. make the incumbent model dumber to push people to use the latest-and-greatest model, or vice versa to ease the load on the latest model, which is typically larger than the old one).
They can limit how hard the model thinks for a given effort. Suddenly xhigh only thinks as hard as high did, and high shifts down to medium effort, and so on.
They can also serve quantized models. And this has the benefit of practically not showing up in benchmarks at all even if the user experience is obviously degraded.
The other major thing the labs do is silently drop the usage limits. This has become very noticeable for codex users who are suddenly burning through their weekly usage in a few hours.
If anyone reading has a GPU it's worthwhile just messing with a smaller model for a bit to watch how the settings affect output.
I have no reason to doubt the claims of the employees at OpenAI and Anthropic who have told us personally multiple times, including here on HN, that they do not degrade the models in order to reduce load.
As for the endlessly long analysis in the OP, it appears it's based on analyzing their random usage data rather than any fixed benchmark. I don't think it makes much sense.
I think that’s the claim in the post, that even though no one can see the true chain of thought, that even the “thinking” text that does get exposed to the user is shorter given the same prompts over time. Not saying it’s true but I think that’s the claim. I’ve personally never noticed the alleged “nerfing” with my enterprise use at work or my subscription use at home which is only during off hours.
That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.
But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.
However, I also agree that I haven't seen any robust data from someone tracking it daily/weekly. The handful of sites purporting to do this aren't even running it enough times to hit stat sig.
edit: someone linked one elsewhere in this thread called AI Stupid Level - they "run 7 trials instead of just 1". I don't blame them. Doing this in a statistically sound manner would cost a small fortune.
In any case, I am willing to believe it's possible something degraded, but so far I have not seen any empirical evidence of it since the previous incident with the inference and harness bugs. I lean towards Anthropic probably not intentionally doing anything like this without disclosing it beforehand.
For example "We didn't change any settings, but when GPU use gets high the run time of a prompt is lessened. But you must remember this is always in effect so nothing changed at all. This happens occasionally on random prompts some of the time, and when it's busy it happens all of the time".
In someones eye this would fit the letter of the law but not the spirit of the law that you hold.
Fable is effectively worse than Opus 4.6 now. They severely messed with the model.
I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.
If you have a bank of rigorously tested benchmarks that you run every few days, with enough trials to know what your standard deviation is, and you are getting significant trends over time with those, that would be interesting.
But "I have a feeling" and "Seems like" really isn't a reliable signal at all, humans just can't handle perceiving these things reliably. On top of that changes in your work environment can easily pollute LLMs and change quality of results. Are things getting added to your memory or claude.md files that you don't realize? Is your project growing in size and thus claude is performing worse as more context is needed to work with it? etc etc
The implication is that humans are unreliable and shouldn't be trusted.
Or humans have certain shorthands when they complain on reddit, but their diagnoses are accurate for the specific context? If my AI does something stupid, am I not allowed to call it out? A NS-solving AI is still capable of not satisfying the abstract thing called the user experience. People have intelligent thoughts without compiling to lean.
OK, you say. Then let's get an aggregate benchmark for "intelligence". That doesn't prove that AI didn't flounder a specific use case that the user requested.
Classic moves: Humans are unreliable, converge to some "objective" benchmark that necessarily will quotient out the special cases, etc. Wonder how we'll be solving these issues in the AGI era - well, if you have an AGI that just replicates itself, dominates everybody because it's a machine and humans are soft fleshy creatures, and agrees with itself, fine. But part of the beauty of human experience is the messy part, and providing value is in the messy part.
The need to have a mental measure of competence for your fellow man is, most likely, a pre-human skill, probably with a dedicated bit of neurons for it. I think the problem is that those instincts were co-evolved with our fellow man, and, as you say, don't apply at all to a more non-deterministic system that, fundamentally, lacks some logic faculties that even small children have (simple riddle modifications, car wash question, etc).
At the end of the day the highest quality of benchmark tells you nothing if the man behind the curtain is constantly changing variables on you. You have no idea if you're really testing the same thing at all. So when you run your test at the top level on their system you're seeing lets say a 30% difference in quality most of the time, you have no idea if you should really only see a 5% difference in quality if you were running a local model with stable settings.
For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"
I had to tell it to ssh into the server and run journlctl to check it
Anecdotal, I know, but they all seem to be less capable with time.
_edit_ I use the same reasoning level of `medium`
As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.
Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.
Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.
If for example AWS sells you a 2xLarge server there may be some variability in performance but it's going to be averaged out very well.
When it comes to AI services executing your model there is absolutely no information on what and with what settings your model is being executed. Hell, you have no idea if it even is the model you're paying for. Add that models are not deterministic so variability can be pretty large.
This leads to a common set of dynamics that induce cheating behavior in humans. For example, is there a mix of different hardware. Does lessor hardware use different settings? How do you know xhigh is what your prompt ran under. Anthropic has a proven history of running your prompt silently under different models.
This is a huge mess that needs and will be regulated or sued heavily over. Hell, with as many people out there that hate AI it might be easier than one thinks to have a state sue the providers on this and elicit a huge amount of discovery.
I wonder what their official explanation for this behavior is.
(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
As far as the API goes, it would be really obvious. I run a small service that uses LLMs extensively, and if a model suddenly dropped in performance it would be straightforward for us to prove it. We regularly run comparisons where we generate completions with alternative models to e.g. see if we could get away with using cheap models for easy cases, if the baseline outputs deteriorated it would be all over our metrics.
...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no?
And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it.
nothing on the fine print tells you what the weights are, you're just getting Fable 5, whatever that is
Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business.
And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.
The reputational hit, if this was to be confirmed, would also be massive. And I do think it would leak! Some employee would say something.
> What is actually stopping these model companies
You can say this about any company in the world, selling anything.
It's trivially measurable, and there are people running the same benchmark on the leading models every day and measuring if they degrade. Spoiler: they don't.
But you can always say "the conspiracy goes higher", and that the companies know about these daily benchmarks and are routing them to "quality" envs.
Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude https://www.wired.com/story/anthropic-responds-to-backlash-o...
But it's still happening: https://github.com/anthropics/claude-code/issues/81759
Do you know why nobody outside the companies knows what's going on? Because they sell a black box with magic inside while steadfastly refusing to tell you if they are pushing buttons on said box while it is running.
Can you imagine how much fraud would exist in the gambling industry if the gambling commission didn't exist at all? Everytime an industry is unregulated and has high costs of entry the entities in the industry abuse their customers. The incentives are much too high for them not to.
This still leaves an absurd amount of wiggle room for arguments like "oh no, our evals show that this quantization has no detectable effect on performance (in the eval distribution) therefore running the quant doesn't degrade quality"
At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.
also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order
The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.
Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI
Are those dials set the same across all hardware configurations and clusters? Does model behavior average out the same across different hardware?
There are just too many different buttons that can be set to really trust a provider either not to directly commit fraud, or indirectly commit fraud with system complexity affecting the output.
But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets (which give positive PR).
The progress however is such that the number of tasks that you can do with >p% automated and X=1 keeps increasing. So many times just waiting works. Of course, here also it changes from field to field. There are some tasks at which AI hasn't even gotten started, others where it has already peaked, others where it's increasing slowly, and others where it's increasing fast.
I've seen bugs in the harnesses, sure, but never anything in the actual model where I could have said with any certainty that it's not just regular variation or me having a bad day myself.
No idea where people get the confidence from to make such claims every other week.
I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.
Their random collection of prompts for each day is not a benchmark.
The site you linked is a much better example of a real benchmark being repeated over time.
Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.
Buy decent coffee instead of your $200 subscription and sidestep all the scams.
1: "Our model will bring about the end of all things. Flee, flee for your lives"
2: "Our model is basically AGI"
3: "Our model will be available in limited release next week"
4: "Everybody who subscribes at the $200 level gets access now"
5: "Everybody who subscribes at the $20 level gets access now"
6, at least at Google: "Our model will be shoved down your throat every time you do a search, whether you want it or not"
I have quite strongly told them, in no uncertain terms, that they are going to kill themselves doing that.
Evidence that vendors are being misleading in what they are delivering is important to share, whether or not you personally approve of that product.
Tests their intelligence, not their diligence.
Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.
EDIT: I forgot (and am shocked) that HN still doesn't seem to support Markdown-style links.
Its similar to other data services I see around my F500 company.