AIToday
Large Language ModelsLatent SpacePublished: Jul 16, 2026, 16:01 JST4 min read

Thinking Machines releases Inkling, strongest U.S. open-weight model to date

Thinking Machines releases Inkling, strongest U.S. open-weight model to date

Key takeaway

  • Thinking Machines Lab released Inkling, a 975B-parameter open-weights multimodal model trained from scratch on 45 trillion tokens, positioning it as the strongest U.S. open-weight release to date.

  • Artificial Analysis ranked it #41 on the Intelligence Index, ahead of prior U.S. leaders, and it entered multiple Arena leaderboards on day 0 with broad ecosystem support across vLLM, SGLang, Modal, and Hugging Face.

  • While still behind top Chinese open models like GLM-5.2 and DeepSeek on some benchmarks, Inkling's Apache 2.0 license, day-0 API pricing ($1.87–$3.74 per 1M tokens on Tinker), and focus on customization and reasoning efficiency rather than benchmark-maxing represent a deliberate positioning for practical use and fine-tuning rather than leaderboard dominance.

3 Key Points

  1. What happened

    Thinking Machines Lab released Inkling, a 975B-parameter Mixture-of-Experts foundation model with 41B active parameters, supporting text, image, and audio inputs and a 1M token context window on open weights. A smaller variant, Inkling-Small (276B total / 12B active), was also previewed. Both were trained from scratch on 45 trillion tokens and released under Apache 2.0 license with day-0 support across vLLM, SGLang, Modal, Baseten, Databricks, and Hugging Face.

  2. Why it matters

    Independent commentators immediately called Inkling the strongest U.S.-based open-weight release so far—a significant milestone for the American open-source AI frontier at a moment when Chinese open models (GLM-5.2, Kimi K2.6, DeepSeek) have dominated recent benchmarks. Artificial Analysis ranked it at 41 on the Intelligence Index, ahead of prior U.S. leaders like Nemotron 3 Ultra (38), and it ranked #9 overall in Agentic Web App Arena with an Elo of 1257, putting it alongside top closed models. The Apache 2.0 license and immediate availability across major serving stacks lower barriers to adoption.

  3. What to watch

    Inkling debuts on Thinking Machines' Tinker API at $1.87–$3.74 per 1M input tokens (64K–256K context tiers) with cached and output pricing tiered separately. Open-weight checkpoints are available immediately on Hugging Face. Performance remains a step behind the best Chinese and closed models on some benchmarks—Natolambert flagged gaps versus GLM 5.2 on agentic tasks and Kimi K2.6 on multimodal—but the focus on customization and reasoning efficiency over benchmark-maxing may indicate a different design philosophy.

Ask the AI about this article →

Context & Analysis

Thinking Machines' release of Inkling marks a deliberate pivot in open-source AI strategy away from benchmark-chasing toward a customizable foundation model positioned for practical deployment and fine-tuning. The company explicitly framed Inkling as a day-1 foundation for future iterations rather than a final frontier push, a messaging choice that distinguishes it from the typical "SOTA claim" playbook. The Apache 2.0 license and immediate availability across seven major serving platforms (vLLM, SGLang, Modal, Baseten, Databricks, Hugging Face, and NVIDIA) signal a deep commitment to ecosystem integration—a meaningful contrast to models that ship narrowly or with restrictive terms.

Inkling's technical architecture reflects several unconventional choices that attracted attention from the research community. The use of relative positional encoding instead of RoPE, hybrid sliding-window attention with a 5:1 local-to-global ratio, short convolution layers around attention/FFN streams, and DeepSeek-style auxiliary-loss-free load balancing with shared expert sinks all depart from recent defaults. These choices may indicate either independent architectural exploration or deliberate avoidance of standard recipes—a point of debate among observers. Pretraining began last winter, with coding, reasoning, and agentic training layered on from mid-January onward, suggesting a relatively compressed timeline to general release.

On performance, Inkling's standing is mixed. Artificial Analysis ranked it as the strongest U.S. open-weight model (Intelligence Index 41, vs. Nemotron 3 Ultra at 38) and it entered Agentic Web App Arena at #9 overall with an Elo of 1257, placing it in the same band as Claude Opus 4.6 and Gemini 3.5 Flash. However, independent commentators flagged gaps: Natolambert called it "a bit behind GLM 5.2 on agentic benchies, and Kimi K 2.6 on multi modal," and Scaling01 described it as "roughly another Kimi-K2.6" and behind all closed models and GLM-5.2. The contrast between strong U.S. rankings and mixed global performance reflects the current landscape, in which Chinese open models have surged ahead. Some observers framed Inkling's moderate benchmark scores as evidence of less corner-cutting and distillation contamination; others saw it as a strategic timing choice ahead of announcements from competitors.

FAQ

What are Inkling's core specifications?
Inkling is a Mixture-of-Experts model with 975B total parameters and 41B active parameters per token. It supports text, image, and audio inputs with text output, has a 1M token context window on open weights (256K on the Tinker API), and was trained on 45 trillion tokens.
How much does Inkling cost on Tinker, and where can I access it?
Tinker pricing is $1.87 per 1M input tokens for 64K context and $3.74 for 256K context, with separate cached and output token pricing ($0.374 / $0.748 and $4.68 / $9.36 respectively). Open weights are available on Hugging Face, and the model is also accessible via Databricks, Baseten, Modal, and vLLM/SGLang on day 0.
How does Inkling rank compared to other open-weight models?
Artificial Analysis ranked Inkling at 41 on the Intelligence Index, making it the leading U.S. open-weights release, ahead of Nemotron 3 Ultra (38), Gemma 4 31B (29), and gpt-oss-120b (24). However, independent commentators noted it is still behind top Chinese open models like GLM-5.2 on agentic benchmarks and Kimi K2.6 on multimodal tasks.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 2h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 2h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleApplied Computing raises $20M for AI model to streamline oil & gas plant operations