Open Model Adoption vs Attention: What the Download Data Means for Your Cost Model
Hugging Face's summer 2026 data shows the models the field talks about and the models it actually runs are almost entirely different sets, and your cost model should follow the second one.
Aug 17, 2026 · 5 min read
Key takeaways
Of the top 25 Hugging Face repositories by 2026 downloads and the top 25 by likes, exactly one appears in both lists.
Models under 1B parameters take 83% of all-time downloads, and only 3% of 2026 download volume goes to models above 70B.
Qwen now has 151,448 derivatives on the Hub, 2.6x Meta's total footprint, and leads local inference with 39.6 million GGUF downloads a month against Gemma's 20.8 million and Llama's 7.5 million.
The runtime layer is growing far faster than the model layer: repositories declaring the gguf library rose 464% and mlx 148%, against 21.5% growth in model repositories overall.
Permissive licences on frontier weights are a distribution strategy, not a promise, and some of the largest releases have started attaching commercial restrictions.
Why do downloads and likes tell opposite stories?
Because they record different acts. A like says a release matters, and it clusters on frontier models in the weeks after launch. A download says something is wired into a pipeline that runs on a schedule, and it accrues to small, stable models over years.
The numbers are stark. all-MiniLM-L6-v2 was pulled 1.55 billion times in seven months against 5,156 likes. Kimi-K3 was pulled roughly 60 times per like received. Not one model published in 2026 reaches the download top 25, while thirteen of those twenty-five date from 2022.
For cost planning, that asymmetry has a direct consequence. Roadmap conversations are dominated by frontier models; production token volume is dominated by small ones. If your COGS model is built around the model you demo, it is probably not built around the model that generates most of your inference bill.
What does the size distribution mean for inference spend?
Models under 1B parameters take 83% of all-time downloads. Everything above 100B takes 1%. Restricting to 2026 downloads changes almost nothing: 3% of the volume goes to models above 70B.
The practical read is that most real workloads are already tiered, whether or not anyone planned it that way. Classification, routing, embedding, and enrichment run on small models. Frontier models get reserved for the small share of requests that genuinely need them.
That is also the cheapest architecture available. The observation worth adding: tiering is usually described as an optimization you adopt later, but the download data suggests it is the default behaviour of the ecosystem, and teams that route everything through one large model are the outliers paying for it.
Where is the ecosystem actually growing?
Not in the model layer. Model repositories grew 21.5% over seven months. Repositories declaring the gguf library rose 464%, lerobot 194%, and Apple's mlx 148%, against 16% for transformers and peft.
That layer decides where a model can physically run: local inference formats, Apple silicon, robot control stacks. Its growth is three to seven times the platform average, which is why a trillion-parameter mixture-of-experts model can now be a realistic local deployment rather than a press release.
Qwen sits at the centre of it. 151,448 derivatives on the Hub, 2.6x Meta's total footprint and 4.7x the Llama repositories specifically, and 39.6 million GGUF downloads a month against Gemma's 20.8 million and Llama's 7.5 million. Llama-derived GGUF repositories slightly outnumber Qwen's, so that gap is demand, not supply.
Is a permissive licence a cost guarantee?
No, and this is the risk line to watch. Of 178 Chinese releases above 20B parameters this year, 59% carry Apache 2.0 and 22% carry MIT, more permissive than the American side of the same band, where 29% is Apache or MIT, 41% sits under custom terms, and 30% declares nothing.
But the trend has started to move. Recent very large releases including Kimi K3 and Qwen3.8 have begun attaching non-commercial restrictions and revenue-share requirements. The weights are given away because the return comes from API and cloud business, hardware positioning, or ecosystem placement, and those incentives can change.
So treat licence terms as a per-release cost variable, not a fixed property of the family. A self-hosting plan built on last year's licence assumption is a plan with an unpriced dependency in it.
The short version: price your stack against what the ecosystem runs, not what it upvotes, and re-check the licence every time you upgrade.
If you want to see how a small-model tier changes your cost per user against a frontier-only setup, you can compare both at calcaas.com.
Frequently asked questions
Do downloads or likes better predict which model to build on?
Downloads. Likes measure community attention and cluster on frontier models in the weeks after launch, while downloads measure what is wired into pipelines that run on a schedule. Of the top 25 Hugging Face repositories by 2026 downloads and the top 25 by likes, exactly one appears in both.
What share of open-model usage goes to large models?
Very little. Models under 1B parameters take 83% of all-time downloads and everything above 100B takes 1%. Restricting to downloads accumulated in 2026, only 3% of volume goes to models above 70B.
Which open model family has the strongest ecosystem position?
Qwen, by a wide margin. It has 151,448 derivatives on the Hub, 2.6x Meta's total footprint and 4.7x the Llama repositories specifically, and it leads local inference with 39.6 million GGUF downloads a month against Gemma's 20.8 million and Llama's 7.5 million.
Are open weights safe to build a cost model on?
Only if you re-check the terms per release. Of 178 Chinese releases above 20B parameters this year, 59% carry Apache 2.0 and 22% carry MIT, but recent very large releases including Kimi K3 and Qwen3.8 have started adding non-commercial restrictions and revenue-share requirements. Place the JSON-LD above inside a <script type="application/ld+json"> tag in the page head.