Top
Best
New

Posted by damaru2 13 hours ago

A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf](jorgegarciaherrero.com)
398 points | 126 commentspage 2
zug_zug 9 hours ago|
time for somebody to make a quick script to poison your chat history by starting 1000 fake conversations with contradictory identifying details "I worry as a lesbian woman, my tween daughter doesn't blah blah Shabbat blah blah move home to Australia"
Madmallard 7 hours ago||
probably will lead to account punishments at some point
nick486 9 hours ago||
ask another llm to generate those prompts for you
dhanushnehru 9 hours ago||
It’s not a leak if it’s the business model.
skybrian 7 hours ago||
In case anyone finds it helpful, I asked ChatGPT to break down what they found by app:

https://chatgpt.com/s/t_6abbbee386bc8191a26717b4f1442657

Traster 11 hours ago||
I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.

First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.

But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.

The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.

delis-thumbs-7e 10 hours ago||
> First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.

And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!

> But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.

Perhaps, but getting to the saturated market on this might be something LLM labs simply don’t have time for. They are haemorrhaging money, no path to profitability and OpenAI especially has made ridiculous promises on data centre -spending for the coming year. They need money now.

sodapopcan 8 hours ago||
> And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!

Not only this, many will continue to help the company by bullying anyone who decides to stop using the drugs.

jeltz 11 hours ago|||
> a whole bunch of their execs come from Meta, and all those guys learned the hard way.

Apparently they have not learned, or they learned the wrong lesson. I do not think they leak it intentionally as hoarding is typically more profitable than selling but anyone who has been in the industry for some time knows that the move fast and break things attitude has caused enormous amounts of data leaks.

gorbachev 9 hours ago|||
> I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.

The lesson they've learned is that if the profit from an activity is X, and the sanctions/reputational hit costs less than X, it's full steam forward.

coliveira 8 hours ago||
The only thing they learned from Meta is that this does work!
lukehandcool 8 hours ago||
You should really assume that any information you give to a private model is going directly into a database. These companies don't care about your privacy. Local, open weight models are the only way to truly protect your data.
oreally 6 hours ago|
you got any recommendations for local models setups? lightweight enough for a 5060 if possible
lukehandcool 4 hours ago||
Yes I do! 5060 is tough because you have 8GB of VRAM. That means any model you load onto it will have to be smaller than that (a model that is 6GB on hugging face will take up 6 GB of VRAM just loaded onto your GPU, then need some more headroom for the actual inference).

So, you could get away with Qwen3.8 quantized to 1 bit which is available here: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF.

Now, depending on how much RAM you have, you can offload some of the inference to the RAM, but this is a performance bottle neck (less tokens per second)

I haven't tried the 1 bit quantization yet, but I often run the Q4 quantization on a 3090 machine I have Qwen3.8-27B-UD-Q4_K_XL.gguf which is 17.6GB and fits nicely on a 3090 with 24GB of VRAM.

To run these models you need to download them (usually from hugging face) and run them with something like llama.cpp (I personally recommend this over ollama). If you are doing that you need a GGUF file which is available from many people on hugging face, most famously the account named "unsloth" which takes the Safetensors weights that a company publishes and turns it into these GGUF files which can be run on a desktop with llama.cpp.

With the 5060 you are limited. If you can get a 3090 you can do really great local work with Q4 Qwen3.8-27B

In my personal set up (which I will release fully open source soon), my qwen model can actually browse the internet as if it was me (you can actually watch it view pages, its pretty cool), which means your little local model can scrape up to date info off the internet.

Hopefully that was helpful. Open source is the future!

Coeur 12 hours ago||
"multiple providers disclose sensitive conversation-derived artifacts — including titles, prompts, and screenshots — to third parties, often alongside persistent user identifiers that enable user attribution. We also find that some providers publicly expose conversation permalinks without access controls, allowing trackers to read the entire conversation."

Not good at all.

Aboutplants 11 hours ago||
“Screenshots”

The amount of sensitive information that accidentally gets left on screenshots is pretty large. This is a pretty massive security issue

baggachipz 9 hours ago||
I'm shocked... SHOCKED! that these companies would sell this data to advertisers and violate the privacy of their users. Who could have seen this coming??
davsti4 9 hours ago||
No way will your future spouse be an AI entomed robot, with a percentage of your salary endowed to the corporate owner of the hardware.
DrMandalay 12 hours ago||
The word is "sell" not "leak". This title takes away all agency from the thieves selling private data to advertisers.
otabdeveloper4 11 hours ago|
The personal data fell off the back of a truck, man.
0xcrypto 11 hours ago||
This is why I built my own chat interface. https://ai.ivx.run/

Access through: https://ai.ivx.run/chat/

pluc 11 hours ago|
You thought... they didn't?
More comments...