General

The Myth of 'Open Source' AI โ€” What Open Weights Actually Mean

Published March 12, 2026

1/ Meta calls Llama "open source." It isn't โ€” not by any traditional definition. The confusion isn't accidental. What "open" means in AI is a redistribution question, and the answer matters. ๐Ÿงต

2/ Traditional open source (Linux, Firefox, Apache) means: open code, free to modify, free to redistribute, governed by a community. You can fork it, change it, compete with it. That's the deal.

3/ Llama's license prohibits use by companies with 700M+ monthly active users. You can't use the outputs to train competing models. You must display "Built with Meta Llama." This is not open source. It's a marketing license.

4/ The Open Source Initiative โ€” the organization that literally defines "open source" โ€” has been clear: most "open" AI models don't qualify. In 2024, OSI published its Open Source AI Definition. Most models fail it.

5/ What most "open" AI models actually release: the weights (the learned parameters). Not the training data. Not the data curation code. Not the RLHF pipeline. Not the compute budget. You get the cake, not the recipe.

6/ Why this matters: Without training data and full methodology, you can't reproduce the model. You can't audit it. You can't meaningfully improve it. You can fine-tune at the margins. That's dependency, not empowerment.

7/ Yochai Benkler's vision of commons-based peer production requires genuine openness: shared resources, collective governance, ability to fork. "Open weights" gives you a product, not a commons.

8/ Who benefits from the confusion? Large companies. "Open source" buys goodwill, attracts developers to your ecosystem, and creates an army of fine-tuners who add value you capture. It's a platform strategy dressed as generosity.

9/ The compute barrier is the quiet part. Even truly open models require millions in compute to train from scratch. Openness without compute access is like publishing a recipe that requires a $100M kitchen.

10/ Real openness would mean: open training data, open methodology, open weights, permissive license, community governance, AND funded compute access for researchers. Almost no project delivers all of these.

11/ Some projects try harder than others. EleutherAI and BLOOM published training data. Allen AI's OLMo released the full pipeline. These matter โ€” they prove genuine openness is possible, just not profitable enough for Big Tech.

12/ Next time someone says "open source AI," ask: Open weights or open source? Can I see the training data? Can I reproduce it? Who governs the project? The answers reveal who actually benefits from "openness." /end


LinkedIn version:

Meta calls Llama "open source." By any traditional definition, it isn't.

Traditional open source โ€” Linux, Apache, Firefox โ€” means open code, freedom to modify, freedom to redistribute, and community governance. Llama's license prohibits use by large companies, restricts using outputs to train competitors, and requires Meta branding. The Open Source Initiative, which defines the term, published its Open Source AI Definition and most "open" models fail it.

What most "open" AI releases actually include: the weights (learned parameters). Not the training data. Not the curation code. Not the RLHF methodology. Not the compute budget. You get the cake, but not the recipe.

This matters because without training data and full methodology, you can't reproduce, audit, or meaningfully improve the model. You can fine-tune at the margins. That's dependency, not empowerment.

Who benefits from the confusion? Large companies. "Open source" buys goodwill, attracts developers to your ecosystem, and creates fine-tuners who add value you capture. It's a platform strategy dressed as generosity.

The compute barrier compounds the problem. Even genuinely open models require millions to train. Openness without compute access is like publishing a recipe that requires a $100M kitchen.

Real openness means: open data, open methodology, open weights, permissive license, community governance, and funded compute. Projects like EleutherAI and Allen AI's OLMo prove this is possible โ€” just not profitable enough for Big Tech.