
What happened
Nvidia develops its own open Nemotron models and launched the Nemotron Coalition in March 2026 with Mistral AI, Perplexity, Cursor, and Thinking Machines Lab among 8 founding members. Alibaba Cloud gives away Qwen models freely, with cumulative downloads exceeding 1 billion by April 2026.
Why it matters
The prior model had frontier AI locked behind closed APIs, so only hyperscalers bought GPUs. Nvidia CEO Jensen Huang calls open models 'the lifeline of innovation' because open model performance is the precondition for its AI Factory vision — enterprises building their own GPU servers to generate AI tokens themselves.
What to watch
The outcome hinges on whether free distribution converts into paid cloud and GPU consumption, or becomes a permanent giveaway. Watch the US-China dimension: China, restricted by US export controls on advanced GPUs, spreads weights freely worldwide since published weights cannot be recalled, while the US in August 2026 notified companies that US open-weight models would be exempt from pre-deployment government safety review.
WHO IT HITSEnterprise IT teams deciding where to run generative AI workloads are the primary audience: the article's core advice is that local LLMs are not a replacement for cloud AI but a negotiating chip giving companies leverage over cloud vendors and a business continuity option for data that cannot leave the premises.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The article frames the local LLM boom as the intersection of three trends that accelerated and crossed by 2026: the maturation of models, hardware, and software; real anxiety that cloud AI providers hold life-or-death control over enterprise AI usage and could cut off access; and the strategic ambitions of both vendors and nations that want open models to win.
On the vendor side, Nvidia's position is structural. Its revenue currently rests on a handful of hyperscalers — Microsoft, Google, Amazon, and Meta — and its AI Factory concept depends on ordinary enterprises buying their own GPU servers to 'manufacture' AI tokens with their own data. That only works if enterprises have strong models they can run themselves. Hence Nvidia develops the Nemotron series, releases not just weights but datasets and training recipes, and in March 2026 formed the Nemotron Coalition with 8 founding members including Mistral AI, Perplexity, Cursor, and Thinking Machines Lab. Alibaba Cloud takes a different route: it keeps top-tier 'Max' models closed for API revenue while freely licensing lower tiers for commercial use, aiming to make Alibaba Cloud the default ecosystem and capture the computing demand that follows. The Qwen download total — over 1 billion cumulative by April 2026 — is the evidence that this funnel is filling.
The geopolitical layer makes the picture sharper. China, restricted by US export controls on leading-edge GPUs, cannot out-compute America's closed labs on raw hardware. Distributing weights freely worldwide is therefore a way to win ecosystem and mindshare, since published weights cannot be recalled. The US has moved in parallel: the July 2025 AI Action Plan acknowledged the geopolitical value of open-weight AI, a July 2026 public letter signed by Nvidia, Microsoft, Meta, IBM, Hugging Face, Mistral AI, and Palantir argued for spreading rather than restricting open weights (OpenAI and Google signed later), and in August 2026 the government told companies that US open-weight models would be exempt from pre-deployment safety review.
The practical read for businesses is that this is not a story about replacing cloud AI. The article's own guidance is that closed frontier models still win on performance and cost-effectiveness, and the realistic answer is task-by-task selection: frontier cloud models for the hardest tasks, local models for data that cannot leave or workflows that must not stop, and cheaper clouds hosting large open models in between. Holding the option to run AI yourself is the leverage — and a business continuity plan for a moment when a model provider might go dark.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Michael Burry, famous for The Big Short, is short Nvidia, Palantir and Tesla, and in his Substack newsletter s…

Investors have three creative routes to Anthropic exposure before its expected IPO: buying Alphabet, Amazon, o…

DeepSeek launched V4.1-Flash, a 763B-parameter open-weight model with a causal encoder-decoder architecture

Much of the attention on AI infrastructure buildouts is now tied to sheer compute power, with dominance define…

Barron's reported September 10 that Kepler Computing emerged from stealth with a memory architecture using fer…

Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…