Posted by damaru2 13 hours ago
First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.
But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.
And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!
> But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
Perhaps, but getting to the saturated market on this might be something LLM labs simply don’t have time for. They are haemorrhaging money, no path to profitability and OpenAI especially has made ridiculous promises on data centre -spending for the coming year. They need money now.
Not only this, many will continue to help the company by bullying anyone who decides to stop using the drugs.
Apparently they have not learned, or they learned the wrong lesson. I do not think they leak it intentionally as hoarding is typically more profitable than selling but anyone who has been in the industry for some time knows that the move fast and break things attitude has caused enormous amounts of data leaks.
The lesson they've learned is that if the profit from an activity is X, and the sanctions/reputational hit costs less than X, it's full steam forward.
So, you could get away with Qwen3.8 quantized to 1 bit which is available here: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF.
Now, depending on how much RAM you have, you can offload some of the inference to the RAM, but this is a performance bottle neck (less tokens per second)
I haven't tried the 1 bit quantization yet, but I often run the Q4 quantization on a 3090 machine I have Qwen3.8-27B-UD-Q4_K_XL.gguf which is 17.6GB and fits nicely on a 3090 with 24GB of VRAM.
To run these models you need to download them (usually from hugging face) and run them with something like llama.cpp (I personally recommend this over ollama). If you are doing that you need a GGUF file which is available from many people on hugging face, most famously the account named "unsloth" which takes the Safetensors weights that a company publishes and turns it into these GGUF files which can be run on a desktop with llama.cpp.
With the 5060 you are limited. If you can get a 3090 you can do really great local work with Q4 Qwen3.8-27B
In my personal set up (which I will release fully open source soon), my qwen model can actually browse the internet as if it was me (you can actually watch it view pages, its pretty cool), which means your little local model can scrape up to date info off the internet.
Hopefully that was helpful. Open source is the future!
Not good at all.
The amount of sensitive information that accidentally gets left on screenshots is pretty large. This is a pretty massive security issue
Access through: https://ai.ivx.run/chat/