General

Thread: Follow the Data โ€” Where AI Training Data Actually Comes From

Published March 12, 2026

1/ Everyone talks about AI models. Almost nobody talks about where the training data comes from. Follow the data and you'll find the biggest redistribution story in AI. ๐Ÿงต

2/ Large language models are trained on text from the internet. But "the internet" means specific things: books (often pirated), Wikipedia (volunteer labor), Reddit posts (unpaid), news articles (paywalled content scraped), academic papers, code repositories.

3/ Image models: trained on billions of images scraped from the web. Artists' portfolios, photographers' work, designers' projects โ€” hoovered up without consent, payment, or credit. LAION-5B, a key training dataset, was assembled by scraping and contains identifiable people's photos.

4/ The labor layer nobody sees: AI training data requires human labeling. Millions of workers โ€” disproportionately in Kenya, India, Philippines, Venezuela โ€” label images, rate model outputs, flag harmful content. Pay: often $1-2/hour. Working conditions: often traumatic (content moderation).

5/ RLHF (the technique that makes chatbots useful) depends on human feedback. Those humans are largely invisible contract workers rating thousands of model outputs daily. When people praise ChatGPT's helpfulness, they're praising these workers' judgment โ€” routed through a model, stripped of attribution.

6/ The economics: Companies took freely available and copyrighted content, processed it with underpaid global labor, produced models worth billions, and captured essentially all the value. The content creators and data workers got nothing โ€” or pennies.

7/ This is extractive by structure, not by accident. The legal frameworks (fair use arguments for training data), the labor structures (contractor arrangements that avoid employment protections), and the technical architecture (models that can't attribute sources) all point the same way.

8/ Scale matters here. If one person scraped your blog to learn from it, you wouldn't mind. When a company scrapes the entire web to build a $100B product, the dynamic changes completely. The legal question is whether scale changes the ethics. The redistribution answer is obvious.

9/ The copyright cases (NYT v. OpenAI, visual artists v. Stability AI, etc.) will set precedents for decades. The core question: Can you build a commercial product by processing others' creative work without consent or compensation? Every other industry says no.

10/ What fair data sourcing would look like: Consent mechanisms for training data. Compensation for creators whose work trains models. Living wages and protections for data labelers. Transparent data documentation (datasheets for datasets). Attribution where possible.

11/ The redistribution lens makes this simple. Billions of people created the raw material. A handful of companies processed it. The value flows one direction. Redistribution means creating mechanisms โ€” legal, economic, technical โ€” to flow some of it back.

12/ Next time you use an AI tool, remember: behind the model are millions of uncredited creators and thousands of underpaid workers. The technology is impressive. The supply chain is extractive. Both things are true. @redistributed


LinkedIn version:

The biggest redistribution story in AI isn't about models or compute โ€” it's about data.

Large language models are trained on books (often pirated), Wikipedia (volunteer labor), Reddit posts (unpaid), news articles (paywalled content scraped), and code repositories. Image models were trained on billions of images scraped without consent from artists, photographers, and designers.

Then there's the invisible labor layer. Millions of workers โ€” disproportionately in Kenya, India, the Philippines, and Venezuela โ€” label data and rate model outputs for $1-2/hour, often under traumatic conditions (content moderation). RLHF, the technique that makes chatbots feel helpful, depends entirely on these workers' judgment, routed through a model and stripped of attribution.

The economics are stark: companies took freely available and copyrighted content, processed it with underpaid global labor, built models worth billions, and captured essentially all the value. Content creators and data workers got nothing โ€” or pennies.

This isn't extractive by accident. It's extractive by structure. The legal frameworks (fair use arguments), the labor structures (contractor arrangements), and the technical architecture (models that can't attribute sources) all point the same direction.

The ongoing copyright cases (NYT v. OpenAI, visual artists v. Stability AI) will set precedents for decades. Fair data sourcing would include consent mechanisms, creator compensation, living wages for data workers, and transparent data documentation.

Behind every AI model are millions of uncredited creators and thousands of underpaid workers. The technology is impressive. The supply chain is extractive. Both things are true.

#AIData #AILabor #DataRights #Redistribution #AIEthics #Copyright

On this page