AIToday
Large Language ModelsAI Business & IndustryGIGAZINE AIPublished: Sep 28, 2026, 19:00 JST

Fireworks AI unveils Ember-1, a Kimi K3-based model that cuts tokens 35%

Fireworks AI unveils Ember-1, a Kimi K3-based model that cuts tokens 35%

3 Key Points

  1. What happened

    Fireworks AI announced Ember-1, a Kimi K3-based model built for its project to create specialized models developers need.

  2. Why it matters

    Tokens are what customers pay for, so a model that does the same work with far fewer of them could lower the cost of running AI agents.

  3. What to watch

    The 35% savings comes from a two-customer A/B test, so wider use hinges on whether it holds up; API pricing is $3 per 1M input tokens and $15 per 1M output tokens.

WHO IT HITSEnterprise teams running AI agents, who pay per token, would see lower bills if the token savings hold outside the tests Fireworks AI ran.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Fireworks AI describes its new Ember-1 as the first release from a project to build the specialized models developers need, targeted at a specific complaint: that cutting-edge models spend so much of their output on reasoning that costs balloon, especially when an AI agent reloads the whole context on every turn.

The approach was to add training to Kimi K3, keeping the parts of its reasoning that help complete the task, such as feedback, while trimming the rest. On the Bedside Bench medical task benchmark, Ember-1 held Kimi K3's score at a lower cost, and its processing time per task was shorter. Kimi K3 had been seen as costing more while performing below GPT-5.6 Sol; Ember-1 sits as the cheaper option that still trails GPT-5.6 Sol on performance.

What hinges on the outcome is whether the roughly 35% token reduction, confirmed in an A/B test with two customers, holds beyond those tests and for Fireworks AI's own staff. For developers and companies paying per token to run agents, that gap between price and performance is the number that decides whether they switch.

FAQ
How much does Ember-1 cost to use?
Fireworks AI lists the API price per 1 million tokens at $3 for input, $0.3 for cached input, and $15 for output.
How was Ember-1 tested before release?
Fireworks AI tested it on its own employees with no reported issues, and on two customers in an A/B test that confirmed the roughly 35% token reduction.
What benchmark was used to compare Ember-1 with other models?
The medical task benchmark Bedside Bench was used, measuring cost per task against score. Ember-1 reached the same score as Kimi K3 at a lower cost.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia's DGX Spark push: 'Data centers in the house'DIGITIMES Asia · 41m ago
  • Google adds Live Avatar to Gemini 3.8 Live, live in Gemini EnterpriseAI Watch (Impress) · 41m ago
  • Julia 1 matches Jev on 73.15% benchmark, loses 87% vs 61%Qiita 機械学習 · 41m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAI hack laws fall short as SB 53 threshold unmet