The cases are different in their details but identical in their outcome. A developer in Israel gets inbound messages from strangers because Gemini listed his personal number as a business contact. A PhD researcher in Washington types a colleague's name into Gemini and gets back her private cell number. A person on Reddit has his phone flooded for weeks by callers looking for a locksmith, a lawyer, a product designer. None of these people did anything wrong. Their information simply existed somewhere on the internet at some point.
This is what happens when AI companies hoover up the entire public web to train their products. Your name mentioned in a forum post from 2015, a workshop registration from last year, a public comment on a local news site. All of it is potential training material. And once it is in, it does not come out easily.
The scale of the problem is almost certainly much larger than reported. A company called DeleteMe, which helps people remove personal data from the web, says queries from customers specifically worried about AI chatbots have grown 400% in the past seven months. Most people who receive strange calls or messages will never trace the source back to a chatbot. They will just assume it is spam.
The friction between how AI works and how our privacy laws work is the core issue here. Europe's GDPR gives people the right to request deletion of their personal data. The US has patchwork state-level rules. But both frameworks were designed for traditional databases, where your data sits in a row in a table and can be deleted on request. An AI model does not work like that. Once personal data is absorbed into a model's internal settings during training, it becomes woven into billions of interdependent parameters. Removing one person's phone number without retraining the entire model is, right now, not technically straightforward. Researchers are working on methods called machine unlearning that might eventually solve this, but they are not yet in use at scale by any major AI company.
There is also a supply chain problem that most people do not know about. AI companies are not just scraping the public internet. They are also buying data from data brokers, the companies that collect, package, and sell personal information as a business. California's state registry revealed that 31 registered data brokers had sold consumer data specifically to developers of generative AI systems in the past year. These brokers collect phone numbers, addresses, family relationships, financial details, and more. Your information may have entered an AI model not from something you posted yourself, but from a broker who assembled a profile about you without your knowledge.
The companies involved have not been responsive. One person who submitted a formal removal request to Google heard nothing for weeks. The Israeli developer who contacted Google's customer service in March did not receive any reply until early May, nearly seven weeks later, and that reply just asked for documents he had already provided. OpenAI has a privacy portal for removal requests but reserves the right to decline them in the public interest. Anthropic has no clear removal process at all.
This also raises a practical business question that goes beyond individual privacy. Many professionals have personal contact details scattered across the internet from old conference registrations, professional directories, or industry forums. Any of that could surface through a chatbot query. The same applies to employees, suppliers, or clients whose data your organisation holds. If a chatbot can surface someone's home address by asking a few follow-up questions about their neighbourhood, the traditional idea of what counts as private information needs to be reconsidered.
The gap between what the law promises and what the technology can actually deliver is only going to grow. More AI models are being trained on more data every month. Data brokers are a growing industry, not a shrinking one. And the companies building these products have a commercial incentive to make them as helpful as possible, which means surfacing information when asked. The current safeguards were never designed for a world where anyone can type a name into a chat interface and get back a home address.