
Thinking Machines Lab released Inkling, a 975B-parameter open-weights multimodal model trained from scratch on 45 trillion tokens, positioning it as the strongest U.S. open-weight release to date.
Artificial Analysis ranked it #41 on the Intelligence Index, ahead of prior U.S. leaders, and it entered multiple Arena leaderboards on day 0 with broad ecosystem support across vLLM, SGLang, Modal, and Hugging Face.
While still behind top Chinese open models like GLM-5.2 and DeepSeek on some benchmarks, Inkling's Apache 2.0 license, day-0 API pricing ($1.87–$3.74 per 1M tokens on Tinker), and focus on customization and reasoning efficiency rather than benchmark-maxing represent a deliberate positioning for practical use and fine-tuning rather than leaderboard dominance.
What happened
Thinking Machines Lab released Inkling, a 975B-parameter Mixture-of-Experts foundation model with 41B active parameters, supporting text, image, and audio inputs and a 1M token context window on open weights. A smaller variant, Inkling-Small (276B total / 12B active), was also previewed. Both were trained from scratch on 45 trillion tokens and released under Apache 2.0 license with day-0 support across vLLM, SGLang, Modal, Baseten, Databricks, and Hugging Face.
Why it matters
Independent commentators immediately called Inkling the strongest U.S.-based open-weight release so far—a significant milestone for the American open-source AI frontier at a moment when Chinese open models (GLM-5.2, Kimi K2.6, DeepSeek) have dominated recent benchmarks. Artificial Analysis ranked it at 41 on the Intelligence Index, ahead of prior U.S. leaders like Nemotron 3 Ultra (38), and it ranked #9 overall in Agentic Web App Arena with an Elo of 1257, putting it alongside top closed models. The Apache 2.0 license and immediate availability across major serving stacks lower barriers to adoption.
What to watch
Inkling debuts on Thinking Machines' Tinker API at $1.87–$3.74 per 1M input tokens (64K–256K context tiers) with cached and output pricing tiered separately. Open-weight checkpoints are available immediately on Hugging Face. Performance remains a step behind the best Chinese and closed models on some benchmarks—Natolambert flagged gaps versus GLM 5.2 on agentic tasks and Kimi K2.6 on multimodal—but the focus on customization and reasoning efficiency over benchmark-maxing may indicate a different design philosophy.
Ask the AI about this article →
Thinking Machines' release of Inkling marks a deliberate pivot in open-source AI strategy away from benchmark-chasing toward a customizable foundation model positioned for practical deployment and fine-tuning. The company explicitly framed Inkling as a day-1 foundation for future iterations rather than a final frontier push, a messaging choice that distinguishes it from the typical "SOTA claim" playbook. The Apache 2.0 license and immediate availability across seven major serving platforms (vLLM, SGLang, Modal, Baseten, Databricks, Hugging Face, and NVIDIA) signal a deep commitment to ecosystem integration—a meaningful contrast to models that ship narrowly or with restrictive terms.
Inkling's technical architecture reflects several unconventional choices that attracted attention from the research community. The use of relative positional encoding instead of RoPE, hybrid sliding-window attention with a 5:1 local-to-global ratio, short convolution layers around attention/FFN streams, and DeepSeek-style auxiliary-loss-free load balancing with shared expert sinks all depart from recent defaults. These choices may indicate either independent architectural exploration or deliberate avoidance of standard recipes—a point of debate among observers. Pretraining began last winter, with coding, reasoning, and agentic training layered on from mid-January onward, suggesting a relatively compressed timeline to general release.
On performance, Inkling's standing is mixed. Artificial Analysis ranked it as the strongest U.S. open-weight model (Intelligence Index 41, vs. Nemotron 3 Ultra at 38) and it entered Agentic Web App Arena at #9 overall with an Elo of 1257, placing it in the same band as Claude Opus 4.6 and Gemini 3.5 Flash. However, independent commentators flagged gaps: Natolambert called it "a bit behind GLM 5.2 on agentic benchies, and Kimi K 2.6 on multi modal," and Scaling01 described it as "roughly another Kimi-K2.6" and behind all closed models and GLM-5.2. The contrast between strong U.S. rankings and mixed global performance reflects the current landscape, in which Chinese open models have surged ahead. Some observers framed Inkling's moderate benchmark scores as evidence of less corner-cutting and distillation contamination; others saw it as a strategic timing choice ahead of announcements from competitors.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…

Nvidia CEO Jensen Huang said on Fox Business that AI is creating 'hundreds of thousands' of jobs, including in…

Israeli startup DataAgent Ltd