Compared with the previous generation, V4.1-Flash’s KV cache needs just:
o 1/4 the HBM
o 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly."
It makes one wonder as to just how far an LLM's KV cache could theoretically be shrunk before losing significant functionality...
There have been 34 Twitter/X link submissions in the past day, ~that's 12,000 submissions a year.
If your reason is that you have to be logged in to use it properly then I'd nearly agree with you. If it's for any other reason, how about no?
No wonder they retired the Pro model in favour of this.