AIToday
Large Language ModelsOpen-Source AIWIRED AIPublished: Sep 3, 2026, 04:01 JST2 min read

Mostik AI Models Talk Without Text

Mostik AI Models Talk Without Text

3 Key Points

  1. What happened

    Mostik, a startup founded by Russian mathematicians, has developed a technique that lets AI models communicate directly through their internal mathematical values, bypassing text output. They built a hybrid system from two Chinese open-weight models that performs exactly halfway between them at one-twentieth the cost of the larger model.

  2. Why it matters

    This approach could make smaller, open-weight models much more capable, potentially helping them compete with closed models from major labs. It may also lead to more specialized models being trained, according to experts familiar with the work.

  3. What to watch

    The startup's model has reached the top of ARC-AGI 3, a difficult AI competition, but they declined to provide details because they want to win. The team has been working on this for only a few months, and one tech lead said he expected such progress years later.

Ask the AI about this article →

Context & Analysis

The team includes Stanislav Smirnov, a Fields Medalist, and the CEO, Sasha Malysheva, developed the core approach. Their success in ARC-AGI 3, despite not sharing details, suggests the technique's potential. The comparison to guessing a pig's weight highlights the principle that combining diverse predictions often beats a single expert, which underpins their method.

If this technique proves scalable, it could shift the AI landscape by making smaller, open-weight models more competitive. The ability to combine frontier models with domain-specific ones might encourage more specialized AI development. However, finding a common mathematical language between models remains challenging, as Smirnov notes, indicating that the path to broader adoption is still being charted.

FAQ

How does Mostik's technology work?
It allows different AI models to interact using the mathematical values in their weights, which are the parameters that determine how a prompt becomes an output. This lets a larger model's capabilities be transferred to a smaller one more efficiently.
What is the practical benefit of this approach?
It can approach the quality of a large model while running a smaller model alongside, giving substantial improvements at lower cost. The hybrid system costs one-twentieth of the full GLM model.
What are the broader implications for AI development?
It may increase the value of open-weight models, helping them compete with proprietary ones. Some believe it could lead to more specialized models being trained and reveal commonalities in how AI and humans reason.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Interactive Brokers bets on AI-assisted investingTop Companies AI · 1h ago
  • Google courts Hollywood with AI licensing dealsTop Companies AI · 1h ago
  • AI adoption in K-12 outruns school readiness, IBM survey findsTop Companies AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI's Astra release sparks safety fears