Top
Best
New

Posted by espeed 22 hours ago

Fable 5 – Median thinking declined in August(twitter.com)
402 points | 283 commentspage 3
Andaith 14 hours ago|
Surely it's easy to verify by simply doing Benchmark tests every 2 or 3 days but not publishing them so they can't be gamed, then releasing all at once?
zerof1l 18 hours ago||
I for sure felt that this was the case for a while now, but couldn’t explain it. Newly released feels great for the first couple of weeks, but then it starts to get worse.
sarfaraznaushad 11 hours ago||
I'm always excited to run models on my local machine. I use Ollama for that.
ThoAppelsin 19 hours ago||
https://www.youtube.com/watch?v=BzNzgsAE4F0
dachworker 21 hours ago||
Makes sense, no? Test time compute is something you can vary, so it makes sense that you start covertly reducing it once the model has already made it's splash.
CamperBob2 22 hours ago||
How do you measure thinking tokens? They don't send those back to the client.
ivanbakel 22 hours ago|
They tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.
CamperBob2 22 hours ago||
Good point. I suppose watching the number go up is useful information in itself.

I have been using CC with DeepSeek 4.1 Flash lately, and it's nice to see how the sausage is being made (even if it's partly illusory, as CoT always is.)

Morkeeth 18 hours ago||
The NERF is finally established, this should be part of the ever growing benchmark maxxing.
parasti 8 hours ago||
I stopped reading when I realized the article reminds me of my own Claude-generated solutions at work - just an endless maze of special business logic on top of special business logic. You need an LLM to understand it. You need an LLM help write the documentation. You need an LLM to help read the documentation.
vb-8448 21 hours ago||
They want transparency from everyone else but not for them ... you don't say.
llmslave 22 hours ago|
I strongly believe that the real Fable is the one we had for a few days in June. Then they nerfed the model a bit after the government pulled it off the market. What we have now is something less, but still good
roncesvalles 21 hours ago|
I also believe this. Fable post-ban was never the same. At the least, whatever system prompt munging or pre/post filtering they did to strengthen the guardrails nerfed it.
llmslave 20 hours ago||
question is if they ever let the general public access borderline AGI
More comments...