Perplexity does this. Visiting a past perplexity search url exposes your full conversation.
I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link
Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped
There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines
This is shockingly lax approach to data security and privacy by design
If you click "provide a shareable link" you should decide (and behave) as though that made it public.
I'm not saying it's good privacy posture on the side of the companies, but how else do you think that would work if there isn't any authentication step for the person viewing it? Even with authentication, "three may keep a secret, if two of them are dead."
It’s not like someone’s gonna guess that URL… right?
I also accidentally paste random stuff into input boxes all the time.
Although I think at the point of some on-device program reading your browser history against your will, you’re gonna have bigger problems.
On android it is possible to give permissions "once" "while using the app" "always".
But whatsapp now only accepts the full camera access. If you set "ask every time" to indeed only make a picture once and then no camera access anymore, it refuses and sends you to the permission dialoge.
however - agree that this is not great - espeically if chat TTL is long. someone who gets your URL can read everything you're asking (eg. by sniffing your network/accessing your browser history)
We have all become Milhouse now.
People love inequality when it's a celebrity they adore living large. Someone who hasn't scammed them and has demonstrably improved their life and the lives of others in a tangible way. The deeper the con, like Trump and Elon, the more damage to trust.
All we've chosen is free markets.
Weren't Mr. Huang, his cofounders, and his peers, rewarded based on merit? Where did socialism step in?
Who is spending the kind of money that makes regular NVIDIA employees millionaires? You have to consider how those people attained it. And at the scale of NVIDIA's growing value downstream of a rapid investment in AI infrastructure, you can trace a good amount of it to Elon Musk, who has interfered illegally in an election to purchase favor with a party that ended several serious investigations into his companies. Now, the government is using Elon's AI which trails in several benchmarks. Elon's storied history of gaming systems to keep his companies alive only begins there.
Unless you become a target of the government. Then people a lot worse than any teacher you ever had will be looking through it, and they, unfortunately, do care.
My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.
Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).
Anyway I'm a student so I'm not looking for something stable. BTW now it's getting better but if I were the client paying for those services I'd be pretty fucking furious for what's going on - guess they'll never know tho
If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).
I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.
Read the T&Cs. If there's even the tiniest bit of "we might provide your data to third parties for the purposes of...", it's not an accident. Virtually all the AI company T&Cs I've looked at had weasel words that open that door, because it's obvious to anyone paying attention that cramming ads into AI products is the next frontier in AI revenue streams.
>The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.
You are currently a cost to them¹ - using your data as free training/refinement, and potentially selling², it is a way to offset that a little.
----
[1] https://isaiprofitable.com/ - some of the green bars are creeping up a bit, but not much unless you count the shovel sellers
[2] sorry, leaking³
[3] Though it could of course be incompetence rather than malice, a mistake they are not actually making anything out of, as per Grey's amendment⁴ to Hanlon's razor.
[4] Any sufficiently advanced incompetence is indistinguishable from malice.
I thought I was the lazy, irrational, inferior human filth whose job they were replacing.
Who thinks OpenAI or any big company for that matter give a crap about them? This isn't a popular sentiment at all, it's just patently false.
In the modern world, it's just critically important that you exert full control over your SO. You wouldn't want them to run off and leak your most intimate secrets to others without your knowledge or control. I'm sure a lot of us could tell stories about our SOs that we didn't fully control and how they ended up leaking a lot of critical information about us. So full control is definitely mandatory in any SO relationship. It's really just prudent
wait something's gone wrong here
For example, we now have self driving cars, we have cameras everywhere. Based on your chat with a cloud AI, you can automatically trigger an automatic monitoring event that follows and tracks you in the real world with the fleet of cameras, cars, GPU, cell signal. Your tracking due to AI has moved into the real world and eventually, a self driving car will take you in to be "processed" against your will, not even for what you posted in a public forum, but for your private ribbing and chatting with some cloud AI.
So definitely put local AI into the mix and keep personal stuff and thoughts local only.
I have a friend/client who understandably doesn't trust the existing privacy policies of the major providers. They have the same problem many of us do: We want the most powerful models, we're willing to pay for them, but we see over and over how much of a frontier the frontier actually is. Frontiers are ugly if you don't have guns.
So for now maybe platform tools like Open WebUI and TypingMind are a good workaround since the big boys don't (apparently) train on API data (for now) or (probably) send that data to advertisers. It would be interesting to confirm that.
There's really no reason to think that enterprise plans are immune. While this analysis didn't test enterprise plans, and only focused on third party tracking, the people behind these companies have already demonstrated that they are willing to break the law to get what they want, are willing to lie to their users, willing to lie to the public, and even willing to lie to congress. Yet somehow people seem convinced that they'd never dare to lie to Random Corp LLC
First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.
But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.
And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!
> But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
Perhaps, but getting to the saturated market on this might be something LLM labs simply don’t have time for. They are haemorrhaging money, no path to profitability and OpenAI especially has made ridiculous promises on data centre -spending for the coming year. They need money now.
Not only this, many will continue to help the company by bullying anyone who decides to stop using the drugs.
Apparently they have not learned, or they learned the wrong lesson. I do not think they leak it intentionally as hoarding is typically more profitable than selling but anyone who has been in the industry for some time knows that the move fast and break things attitude has caused enormous amounts of data leaks.
The lesson they've learned is that if the profit from an activity is X, and the sanctions/reputational hit costs less than X, it's full steam forward.
So, you could get away with Qwen3.8 quantized to 1 bit which is available here: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF.
Now, depending on how much RAM you have, you can offload some of the inference to the RAM, but this is a performance bottle neck (less tokens per second)
I haven't tried the 1 bit quantization yet, but I often run the Q4 quantization on a 3090 machine I have Qwen3.8-27B-UD-Q4_K_XL.gguf which is 17.6GB and fits nicely on a 3090 with 24GB of VRAM.
To run these models you need to download them (usually from hugging face) and run them with something like llama.cpp (I personally recommend this over ollama). If you are doing that you need a GGUF file which is available from many people on hugging face, most famously the account named "unsloth" which takes the Safetensors weights that a company publishes and turns it into these GGUF files which can be run on a desktop with llama.cpp.
With the 5060 you are limited. If you can get a 3090 you can do really great local work with Q4 Qwen3.8-27B
In my personal set up (which I will release fully open source soon), my qwen model can actually browse the internet as if it was me (you can actually watch it view pages, its pretty cool), which means your little local model can scrape up to date info off the internet.
Hopefully that was helpful. Open source is the future!
Not good at all.
The amount of sensitive information that accidentally gets left on screenshots is pretty large. This is a pretty massive security issue
Access through: https://ai.ivx.run/chat/
Someone should have to investigate, but I suppose it's all "legal"?
Google is an ad company.
or A Privacy Analysis of Web and Mobile Conversational AI Agents
Site link with context: https://jorgegarciaherrero.com/en/prompt-like-a-butterfly-st...
My stance used to be that the invention of the filter bubble combined with targeted advertising is the most dangerous invention for human society of the last 100 years.
AI turns that up to 11
What happened? Oh no!
How terrible! That’s just, that’s just awful!
How terrible! Oh no!
This shit needs to be stamped down by law.
People take up pitchforks and torches against AI data centers,
well how much time, money, and resources have been wasted on advertisements over the centuries?
How many people, including children, have been deceived by ads?
How much privacy has been violated in the name of "sErViNg YoU rElEvAnT aDs"?
Have you ever tried browsing YouTube from a poor connection, and noticed how long it takes to serve ads? How much bandwidth has been wasted on ads so far?
Who[m] does it all benefit?
Screenshots of conversations. TikTok received screenshots of Grok chats during sharing, exposing the actual visible conversation content.
Conversation-derived content tied to persistent identifiers, including prompts and automatically generated chat titles revealing sensitive facts. "Salary 85k NYC: mortgage 280–350k".
It's not a 'leak'. It's in the T&Cs you agreed to. This is the business plan.
if you have an individual who has been sexually harassing others regularly the problem exists both with the individual and with there not being clear reporting standards, remediation processes, and person-by-person training opportunities (or only offering lackluster, barely enforced versions of the training)
organizations, especially ones as large and well-funded as LLM model providers that can hire whole fleets of trainers, ethics watchdogs, etc, that then produce large-scale, systematic incompetence are and should be the primary culprit and the one targeted by accountability processes
outside of someone just straight up embezzling funds or grossly misusing their power, organizations should also bear the brunt of the culpability especially when it relates to harms to the public good, especially if those harms benefit the organizations directly. and even in those cases the question becomes why weren't there accountability practices? checks on power? etc
I would still bet they aren’t making any money from data leaking to advertisers
I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.
This partial prompt data might potentially be used to "pre-warm" some kind of cache.
But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.
OpenAI’s statements in response to the Millenium Prize (and related) disputes I think are a pretty obvious example of this in practice. One man’s “user prompts” is not another’s “reasoning trace scratchpad”.
This comment by Falserum on the mathematics research post articulates it well:
https://news.ycombinator.com/item?id=49649992
I'm pretty sure it is used for this; but rather than for anything nefarious, my guess is that this info is then fed to a classifier model to ensure that users of ChatGPT-the-service (as opposed to the OpenAI inference API) are actual humans, rather than agents trying to circumvent having to pay API pricing.