Top
Best
New

Posted by victormustar 19 hours ago

LensVLM: Compressing long context as images, expanding only relevant pages(huggingface.co)
81 points | 7 comments
rao-v 15 hours ago|
I really like this approach! I sort of think of the vision encoder here as an expensive high fidelity RAG encoder.

The thing I’d love to do with a system like this is train it to be KV cache ordering independent (ie permutation invariant at the page level). Basically each page’s KV cache should be understandable by the model in any ordering - which would allow you to go one step further and treat the KV cache of the vision encoded page as the chunk for the model to reason over.

Then all these zoom in for more detail tricks will extend naturally.

taylorfinley 14 hours ago||
Oh My Pi has done this for a while now, they call it Snap compact.
dvt 13 hours ago||
I remember reading a paper entitled "A Picture is Worth a Thousand Tokens" or something similar like 2-3 years ago. The reality is that no one really wants/needs contexts that big, anyway. It's hard enough making LLMs truly useful even with a small/medium context.
wangii 13 hours ago||
yep, deepseek
2muchtime 13 hours ago||
Ha! Didn’t realize that’s what it was doing, I’d compact and it would say snap compact with a little icon of a camera, so this all makes sense now.
himata4113 11 hours ago||
I always found it weird that we don't have glacial type input for llms or any kind of active-working memory.

There's no reason why we shouldn't be able to expose active relevant information that is only relevant for the next request: current agents running, time, etc.

There's also no reason why we shouldn't have a cheaper lossy input which uses way less bytes per token - see deepseek flash 4.1.

lathoa 15 hours ago||
Interesting approach. thanks
lohr13 13 hours ago||
[flagged]
aidiveyt 7 hours ago|
[dead]