Every photo you post, every message you send, every search you type — it might already be feeding an AI model. Here is what is actually happening with your data, who profits from it, and what you can realistically do about it.
The Quiet Harvest: How Your Data Ends Up in AI Models
Sometime in October 2025, Google flipped a switch. Without fanfare or explicit permission, the company activated its Gemini AI across Gmail, Google Chat, and Google Meet for all users. The AI could now read your emails, parse your conversations, and analyze your video calls. Most people never noticed. A 2025 study found that 84% of people interacting with AI systems don’t even realize they’re doing so.
That should concern you. Not in an abstract, dystopian-future kind of way. Right now, in March 2026, your data is likely being harvested, processed, and fed into machine learning models by companies whose names you’d recognize instantly.
The mechanism is deceptively simple. When you agree to a terms-of-service update — the kind that pops up while you’re trying to check your morning messages — you’re often granting broad rights to use your content for “service improvement.” That phrase, buried in paragraph forty-seven of a document nobody reads, is the legal gateway through which your photos, your writing, your voice recordings, and your behavioral patterns flow into training datasets.
And the scale is staggering. Meta used billions of public Facebook and Instagram posts to train its Llama models. Stability AI scraped roughly 5.8 billion images from the internet, many of them personal photographs, to build Stable Diffusion. OpenAI trained GPT models on vast swaths of the open web — including forum posts, blog entries, and product reviews written by real people who had no idea their words would one day help a machine generate text.
The awareness gap is real: 99% of Americans use AI-powered products weekly, but only 36% realize those products use AI at all. Meanwhile, 43% of workers admit to sharing sensitive workplace information with AI tools without their employer’s knowledge.
Who Profits — and Who Doesn’t
The economics of AI training data tell a stark story. The companies building large language models and image generators are now worth trillions of dollars collectively. The people whose data makes those models possible receive nothing. Not a cent. Not even a notification.
Consider a freelance photographer whose portfolio sits on a personal website. Those images may have been scraped by Common Crawl, included in LAION-5B, and used to train models that now compete directly with that same photographer for commercial work. The photographer trained the machine that replaced them, involuntarily and for free.
More than 50 AI-related data lawsuits are currently pending in U.S. federal courts alone. The cases read like a catalog of modern tech grievances: Khan v. Figma alleges the design platform used creators’ work to train AI features. The class-action suit Thele v. Google challenges the secret Gemini activation. In Riganian v. LiveRamp, consumers allege a data broker used AI to combine and sell personal information drawn from both online and offline sources.
These lawsuits are testing a fundamental question: does scraping publicly available data for AI training constitute fair use, or is it theft dressed in legal language? Courts haven’t given a definitive answer yet — and the ambiguity benefits the companies doing the scraping.
Where Your Data Goes
You create the value. You receive none of the profit. That is the current arrangement.
The Legal Landscape: Catching Up, Slowly
Regulation is arriving, but it moves at the pace of legislation while AI moves at the pace of venture capital. Here is where things stand in early 2026.
California’s AB 2013, the Generative Artificial Intelligence Training Data Transparency Act, took effect on January 1, 2026. It requires developers of public-use generative AI systems to publish high-level information about their training data. It’s a disclosure law, not a consent law — companies must tell you what they used, but they don’t need your permission to use it.
The California Privacy Protection Agency (CPPA) also adopted new CCPA regulations covering automated decision-making technology, mandating consumer opt-out rights and requiring annual cybersecurity audits for covered businesses. These rules add real teeth to what was previously a largely theoretical right to say “no.”
In Europe, the picture is both clearer and more severe. The EU AI Act’s high-risk system requirements take full effect on August 2, 2026, with penalties that exceed even the GDPR: up to 35 million euros or 7% of global annual turnover for the most serious violations. The GDPR itself requires companies to prove a lawful basis for using personal data in AI training — and several European data protection authorities have argued that broad scraping doesn’t qualify.
| Regulation | Region | Key Requirement | Effective |
|---|---|---|---|
| AB 2013 | California | Training data transparency disclosure | Jan 2026 |
| CCPA ADMT Rules | California | Consumer opt-out for automated decisions | 2026 |
| EU AI Act (Annex III) | EU | High-risk system compliance, up to 7% turnover fines | Aug 2026 |
| GDPR Art. 6/9 | EU | Lawful basis required for personal data use | Active |
| Illinois HB 3773 | Illinois | AI notification and anti-discrimination | Jan 2026 |
What You Can Actually Do About It
Let’s be honest: the options available to an individual are limited. You can’t un-train a model. Once your data has been absorbed into the weights of a neural network, it doesn’t exist as a discrete, removable unit anymore. It’s been blended with billions of other data points into something new. That’s both the technical reality and the convenient excuse.
But “limited” is not the same as “nothing.” Here are concrete steps that actually matter.
Audit your privacy settings aggressively. Most major platforms now offer some form of AI training opt-out, though they rarely make it easy to find. On Meta, check Settings > Privacy > AI data usage. On Google, look for the Gemini activity controls buried in your Google Account dashboard. OpenAI lets ChatGPT users disable training on their conversations in the app settings. Reddit added opt-out toggles in late 2025 after user backlash.
Exercise your CCPA or GDPR rights. If you’re in California or the EU, you have a legal right to request deletion and to opt out of data sales. Companies must honor these requests within specific timeframes. The Global Privacy Control browser signal, which many sites are now legally required to recognize, automates opt-out requests across the web.
Be deliberate about what you share. This isn’t victim-blaming — you shouldn’t need a law degree to post a photo. But until the law catches up, treating public platforms as public broadcasts is a reasonable precaution. Anything you post publicly on the internet should be assumed to be training data.
Support the lawsuits. Class actions against major tech companies are the most realistic path toward systemic change. The outcomes of cases like Thele v. Google and the ongoing wave of AI privacy litigation will define whether companies can continue treating your data as free raw material.
The Deeper Question: Consent in the Age of AI
The real problem isn’t technical. It’s philosophical. Our entire framework of digital consent was designed for a simpler era — one where “using your data” meant showing you a targeted ad, not feeding your creative work into a system that generates competing creative work.
Terms-of-service agreements are legal fictions. Nobody reads them. Studies have consistently shown that users would need roughly 76 working days per year just to read the privacy policies of the sites they visit. The “consent” they grant is not informed consent in any meaningful sense. It’s capitulation to a system designed to extract maximum permissions with minimum friction.
Some researchers advocate for a new model: data dividends, where individuals receive compensation proportional to the value their data contributes to AI systems. Others push for collective bargaining structures, similar to unions, that would negotiate data rights on behalf of large groups. Neither idea has gained legislative traction yet, but both acknowledge something the current system refuses to: your data has value, and you deserve a say in how it’s used.
The window for shaping these norms is narrowing. Every month that passes, more models are trained, more data is absorbed, and the status quo becomes harder to reverse. The question isn’t whether AI should use human data — some level of data use is inevitable and even beneficial. The question is whether the people generating that data will have any voice in the process, or whether their silence will continue to be mistaken for consent.
Frequently Asked Questions
You can request deletion under GDPR (right to erasure) or CCPA (right to delete), and companies must acknowledge and respond. However, removing your specific data from an already-trained model is technically extremely difficult — the data has been transformed into statistical patterns across billions of parameters. In practice, companies typically commit to excluding your data from future training runs rather than retroactively removing it from existing models. The EU AI Act may eventually force more robust solutions.
It depends on the service and its terms. Google’s activation of Gemini across Gmail and Meet in October 2025 showed that the line between “private” and “training data” is thinner than most people assume. Generally, end-to-end encrypted services like Signal cannot access your content for training. Cloud storage providers vary — check whether your provider’s terms include language about using uploaded content for AI improvement. When in doubt, assume any unencrypted data on a commercial platform could potentially be used.
This is the central legal question currently being litigated across more than 50 federal cases. In the U.S., there is no definitive ruling yet. Companies argue that publicly available data is fair game; plaintiffs argue that public visibility does not equal consent for commercial AI training. The EU position is stricter — GDPR applies regardless of whether data is publicly accessible, and web creators can assert a machine-readable opt-out under EU copyright law. The resolution of these cases over 2026-2027 will set lasting precedent.