General

Thread: What 'Open' Really Means in AI โ€” A Practical Guide

Published March 12, 2026

1/ "Open source AI" is the most misused term in the industry right now. Here's what people mean vs. what's actually happening. Thread ๐Ÿงต

2/ Level 1 โ€” TRULY OPEN: Model weights, training code, training data, evaluation code all released. Anyone can reproduce, modify, and redistribute. Examples: very few models actually meet this bar.

3/ Level 2 โ€” OPEN WEIGHTS: Model weights released, but training data and full training process are not. You can use and fine-tune the model, but you can't reproduce it from scratch. This is what most "open source AI" actually is.

4/ Level 3 โ€” RESTRICTED OPEN: Weights released with usage restrictions (no military use, no competing products, etc.). Meta's Llama models fall here. You get access, but not freedom. This is "open" in marketing, not in principle.

5/ Level 4 โ€” API-ONLY: You can use the model through an API but have no access to weights, architecture details, or training data. OpenAI, Anthropic, Google โ€” this is closed, even if the API is publicly available.

6/ Why does this matter for redistribution? Because "open" determines who can build on AI and who remains a customer. The difference between owning infrastructure and renting it.

7/ The compute problem: Even truly open models require massive compute to train. Llama 3 training cost estimated at $100M+. "Open" weights don't help if you can't afford to fine-tune, let alone retrain. Openness without compute access is a partial freedom.

8/ The data problem: Most "open" models don't release training data. This means you can use the model but can't audit what it learned, verify its biases, or understand its limitations. It's open like a black box with the lid off โ€” you can see the output, not the mechanism.

9/ The governance problem: When Meta releases Llama, Meta still decides the license terms, the acceptable use policies, the next version's direction. Community can fork, but forking a 70B parameter model requires the resources of a well-funded company.

10/ What actual open AI infrastructure would look like: Public compute access (like CERN for AI). Community-governed training data (like Wikipedia for datasets). Transparent, reproducible training (like open-source software). Shared evaluation standards.

11/ The honest assessment: "Open weight" models are better than fully closed ones. Meaningfully. But calling them "open source" implies a freedom and accessibility that doesn't exist when training costs $100M and you can't see the data. Language matters.

12/ Next time someone says "open source AI," ask three questions: Can I see the training data? Can I afford to retrain it? Who decides the license terms? The answers tell you how "open" it really is. Follow @redistributed for more.


LinkedIn version:

The term "open source AI" is being used to describe everything from truly reproducible models to restricted-license weight releases to API access. This imprecision matters because "open" determines who can build on AI and who remains a customer โ€” it's the difference between owning infrastructure and renting it.

Here's a practical framework:

Truly open: Weights, training code, training data, and evaluation all released. Anyone can reproduce. Very few models meet this bar.

Open weights: Weights released but training data withheld. You can fine-tune but can't reproduce or fully audit.

Restricted open: Weights with usage limitations. Meta's Llama models fall here โ€” access, but not freedom.

API-only: Closed, regardless of public availability.

The deeper issue: even truly open models require $50-100M+ in compute to train. Openness without compute access is a partial freedom. And when the releasing company still controls license terms and development direction, the community's agency is limited.

What would genuine open AI infrastructure look like? Public compute (like CERN for AI), community-governed training data (like Wikipedia for datasets), transparent and reproducible training, and shared evaluation standards.

Open weights are better than closed models. But let's be precise about what "open" means โ€” and what it doesn't.

#OpenSourceAI #AIPolicy #Redistribution #AIInfrastructure

On this page