Posted by damaru2 13 hours ago
I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.
This partial prompt data might potentially be used to "pre-warm" some kind of cache.
But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.
OpenAI’s statements in response to the Millenium Prize (and related) disputes I think are a pretty obvious example of this in practice. One man’s “user prompts” is not another’s “reasoning trace scratchpad”.
This comment by Falserum on the mathematics research post articulates it well:
I'm pretty sure it is used for this; but rather than for anything nefarious, my guess is that this info is then fed to a classifier model to ensure that users of ChatGPT-the-service (as opposed to the OpenAI inference API) are actual humans, rather than agents trying to circumvent having to pay API pricing.
Perplexity does this. Visiting a past perplexity search url exposes your full conversation.
I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link
Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped
There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines
This is shockingly lax approach to data security and privacy by design
If you click "provide a shareable link" you should decide (and behave) as though that made it public.
I'm not saying it's good privacy posture on the side of the companies, but how else do you think that would work if there isn't any authentication step for the person viewing it? Even with authentication, "three may keep a secret, if two of them are dead."
It’s not like someone’s gonna guess that URL… right?
I also accidentally paste random stuff into input boxes all the time.
Although I think at the point of some on-device program reading your browser history against your will, you’re gonna have bigger problems.
On android it is possible to give permissions "once" "while using the app" "always".
But whatsapp now only accepts the full camera access. If you set "ask every time" to indeed only make a picture once and then no camera access anymore, it refuses and sends you to the permission dialoge.
however - agree that this is not great - espeically if chat TTL is long. someone who gets your URL can read everything you're asking (eg. by sniffing your network/accessing your browser history)
We have all become Milhouse now.
People love inequality when it's a celebrity they adore living large. Someone who hasn't scammed them and has demonstrably improved their life and the lives of others in a tangible way. The deeper the con, like Trump and Elon, the more damage to trust.
All we've chosen is free markets.
Weren't Mr. Huang, his cofounders, and his peers, rewarded based on merit? Where did socialism step in?
Who is spending the kind of money that makes regular NVIDIA employees millionaires? You have to consider how those people attained it. And at the scale of NVIDIA's growing value downstream of a rapid investment in AI infrastructure, you can trace a good amount of it to Elon Musk, who has interfered illegally in an election to purchase favor with a party that ended several serious investigations into his companies. Now, the government is using Elon's AI which trails in several benchmarks. Elon's storied history of gaming systems to keep his companies alive only begins there.
Unless you become a target of the government. Then people a lot worse than any teacher you ever had will be looking through it, and they, unfortunately, do care.
My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.
Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).
Anyway I'm a student so I'm not looking for something stable. BTW now it's getting better but if I were the client paying for those services I'd be pretty fucking furious for what's going on - guess they'll never know tho
If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).
I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.
Read the T&Cs. If there's even the tiniest bit of "we might provide your data to third parties for the purposes of...", it's not an accident. Virtually all the AI company T&Cs I've looked at had weasel words that open that door, because it's obvious to anyone paying attention that cramming ads into AI products is the next frontier in AI revenue streams.
>The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.
You are currently a cost to them¹ - using your data as free training/refinement, and potentially selling², it is a way to offset that a little.
----
[1] https://isaiprofitable.com/ - some of the green bars are creeping up a bit, but not much unless you count the shovel sellers
[2] sorry, leaking³
[3] Though it could of course be incompetence rather than malice, a mistake they are not actually making anything out of, as per Grey's amendment⁴ to Hanlon's razor.
[4] Any sufficiently advanced incompetence is indistinguishable from malice.
I thought I was the lazy, irrational, inferior human filth whose job they were replacing.
Who thinks OpenAI or any big company for that matter give a crap about them? This isn't a popular sentiment at all, it's just patently false.
In the modern world, it's just critically important that you exert full control over your SO. You wouldn't want them to run off and leak your most intimate secrets to others without your knowledge or control. I'm sure a lot of us could tell stories about our SOs that we didn't fully control and how they ended up leaking a lot of critical information about us. So full control is definitely mandatory in any SO relationship. It's really just prudent
wait something's gone wrong here
For example, we now have self driving cars, we have cameras everywhere. Based on your chat with a cloud AI, you can automatically trigger an automatic monitoring event that follows and tracks you in the real world with the fleet of cameras, cars, GPU, cell signal. Your tracking due to AI has moved into the real world and eventually, a self driving car will take you in to be "processed" against your will, not even for what you posted in a public forum, but for your private ribbing and chatting with some cloud AI.
So definitely put local AI into the mix and keep personal stuff and thoughts local only.
I have a friend/client who understandably doesn't trust the existing privacy policies of the major providers. They have the same problem many of us do: We want the most powerful models, we're willing to pay for them, but we see over and over how much of a frontier the frontier actually is. Frontiers are ugly if you don't have guns.
So for now maybe platform tools like Open WebUI and TypingMind are a good workaround since the big boys don't (apparently) train on API data (for now) or (probably) send that data to advertisers. It would be interesting to confirm that.
There's really no reason to think that enterprise plans are immune. While this analysis didn't test enterprise plans, and only focused on third party tracking, the people behind these companies have already demonstrated that they are willing to break the law to get what they want, are willing to lie to their users, willing to lie to the public, and even willing to lie to congress. Yet somehow people seem convinced that they'd never dare to lie to Random Corp LLC