AIToday
Large Language ModelsOpen-Source AIQiita 機械学習Published: Sep 28, 2026, 19:00 JST

Julia 1 matches Jev on 73.15% benchmark, loses 87% vs 61%

Julia 1 matches Jev on 73.15% benchmark, loses 87% vs 61%

3 Key Points

  1. What happened

    On the 2,000-question typed-decisions benchmark, Supersonic Labs' free Julia 1 scored 73.15% vs. Jev's 72.70% (and 72.55% on CPU), but on the 72-way Banking77 task Jev won 87% vs. Julia 1's 61-64%.

  2. Why it matters

    Julia 1 runs locally on CPU with no fee, so teams with data they cannot send outside the company could try it first, but its weakness on many-option tasks is likely to force a shift to Jev for fine-grained labels.

  3. What to watch

    The benchmark comparison is not apples-to-apples, since Jev's numbers come from third-party records measured on different days and Julia 1's README does not say whether the same task type was seen in training.

WHO IT HITSEngineers prototyping AI text-routing and classification on a local Mac, and internal tooling teams that must keep inquiry data on-premises, gain a free CPU option; call-center and support-ops teams handling more than 20 fine-grained labels may still need the paid cloud service.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The comparison matters because Jev and Julia 1 approach the same "judgment model" task from opposite directions. Jev, from TypeSafe AI, is a cloud API opened to early access on September 15, 2026; Julia 1, from Brazil's Supersonic Labs, was published on Hugging Face on September 25-26, 2026, with weights under Apache-2.0. Both take a state, a question and options, and return a probability per option, but Jev is only reachable through a paid cloud API while Julia 1 runs on a CPU.

The reproduction itself carries weight. The author re-ran the 2,000-question typed-decisions test on a MacBook Pro with an Apple M5 Max and got 72.55%, matching the official CPU re-measurement exactly; the gap from Jev's 72.70% is within noise. That said, Julia 1's README does not say whether its training overlapped with the benchmark tasks, and Jev's figures were transcribed from third-party records measured on different days rather than a head-to-head run.

The author also reports that Julia 1's training and experiments used about $104 (540 reais) in cloud GPU costs, and that Julia 2 is planned to be built from the ground up rather than on mmBERT-small. The outcome likely hinges on whether Julia 1's accuracy holds on each user's own data once the question wording and option descriptions are adjusted, which is why the author's own recommendation is to measure 50-100 real inquiries before choosing.

FAQ
How much does Julia 1 cost to run?
Running Julia 1 locally is free, since its weights are released under Apache-2.0. A hosted API is planned at $0.025 per 1 million input tokens, while Jev charges $0.042 per 1 million input tokens.
Can Julia 1 handle a 72-choice classification like Banking77?
Julia 1 accepts only 2-20 options per call, so a 72-choice task must be split into a preliminary round and a final round. In testing, its accuracy dropped to 61% on Banking77, versus Jev's 87%.
Does Julia 1 work in Japanese?
On the author's 48-item Japanese inquiry set, 6-way team routing reached 77.1%. But open-ended wording failed even with descriptions, with an AUC of 0.47, close to chance, so the author recommends testing a 2-option choice rewrite first.
Qiita 機械学習Read Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia's DGX Spark push: 'Data centers in the house'DIGITIMES Asia · 1h ago
  • Google adds Live Avatar to Gemini 3.8 Live, live in Gemini EnterpriseAI Watch (Impress) · 1h ago
  • Fireworks AI unveils Ember-1, a Kimi K3-based model that cuts tokens 35%GIGAZINE AI · 1h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleFireworks AI unveils Ember-1, a Kimi K3-based model that cuts tokens 35%